Commit Graph

233 Commits

Author SHA1 Message Date
Tom Boucher
682eaae3f0 enh(#2876): retire the dead and pass-through exports from bin/install.js (#3615)
* enh(#2876): retire the dead and pass-through exports from bin/install.js

The installer exported 197 names and had zero production consumers - every
non-test require of it repo-wide sits inside a comment. Its interface was
shaped by test access, not by callers.

Removes 9 dead exports and 61 pass-throughs, repointing their tests onto the
extracted modules' own interfaces. 197 down to 127.

Every count in the issue was wrong: 197 exports not 188, 9 dead not 12, 61
pass-throughs not 49, 44 test files not 42 - and the audit itself then missed
7 more consumer files. restoreUserArtifacts was on the dead list but ceased to
exist in phase 6, and two _GSD_EFFORT_MANIFEST_* names listed as dead are now
genuinely asserted, so acting on that list would have deleted live exports.

7 of the 9 dead names collide with an independent declaration that install.js
delegates TO. Each removal was justified by which declaration a reference
resolves to, never by whether the name appears somewhere.

Coverage parity was the gate rather than test greenness: per-file counts were
captured before any edit and diffed after. 44 of 45 files are byte-identical;
the single delta is one added assertion, not a loss.

The sweep for scattered require sites found two forms static grep misses -
require(VARIABLE) and multi-line require() - plus tests asserting that
install.js re-exports the SAME object, which now assert retirement instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2876): close review findings — restore the duplicate-body guard, sweep orphaned code

Both review engines found real defects in the first cut.

The DEFECT.GENERATIVE-FIX single-owner guard from #1511 had been repointed
from a reference-identity check to install.X === undefined. Those are not
equivalent: the guard exists to catch a duplicate function body reintroduced
into install.js, and the replacement passes cleanly if that duplicate is used
internally and never exported. It now walks bin/install.js's real top-level
bindings, so it catches a duplicate under either shape, exported or not -
strictly stronger than the check it replaced. Proved by injecting a duplicate
and watching it go red.

That weakening survived the coverage-parity gate because the assertion count
never moved. The gate compares counts, so an assertion that changes meaning
rather than number is invisible to it.

Removing the exports had orphaned their wrapper bodies: 14 dead wrappers, 9
consts and 9 destructure entries, several pre-existing and found by the same
sweep. Dead code left in the file this phase exists to shrink.

Three more comments claimed re-exports this phase removed, and tests were
reading Cursor and Windsurf hook constants from install.js's local copy while
calling functions from the hooks surface - equal today, with nothing holding
them equal. The local consts now reference the owning module.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2876): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 09:23:44 -04:00
Tom Boucher
3ab0007164 enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19)

preserveUserArtifacts held user files only in an in-memory Map across the
wipe, so any process death between preserve and restore lost them outright.

Seven call sites, not the four the issue records. Three of them never called
the helper at all - they open-coded the same read/wipe/write - so searching
for callers under-counted by construction; the extra sites were found by
sweeping for the pattern instead.

The worst is the mainline install path, where the crash window spans the
entire gsd-core tree copy rather than a single rmSync.

Adds src/user-artifact-staging.cts: durable on-disk staging with a record
written after the copies land as the commit point, plus recovery of orphaned
batches on the next run - without recovery the staged bytes survive but the
user's file is still gone, which would pass its own test while delivering
nothing.

Routes copyPreservingSymlink through installFs() so staging cannot bypass the
install fs seam, and reunites its symlink-safety docblock with the function it
documents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-3574 with four claims disproved by implementation

Implementing Phase 6 disproved four statements the ADR rests on. The central
decision - no single materializer - is unaffected and stands.

Corrected: decision 3 was already satisfied, so nothing was extracted; the
agents-bypass runtime set omitted claude, kilo and opencode, and closing it
needed three new pieces of descriptor contract rather than proceeding on its
own terms; three of the four blockers the layout comment names were already
stale; and F19 is seven call sites, not four.

Records the generalizable lesson: the defect is the pattern of holding user
data in memory across a wipe, not the helper, so searching for callers of the
helper under-counts by construction.

Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and
notes that copyPreservingSymlink needed routing through the install fs seam
before it could be reused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close dangling-symlink blind spot and harden staging recovery

An adversarial review found the F19 staging work shipped red and unsafe.

Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween
missed dangling symlinks in both its root check and its per-segment walk,
because it probed with existsSync, which is false for a link whose target does
not exist. Fixing only the new module would have reused a guard that was
itself blind. This guard protects the whole install tree.

Recovery no longer throws: it degrades per entry and per file, so one bad
batch cannot block the others. Previously an unrecoverable entry propagated
out of the first statement of install and uninstall, before the cleanup that
would have removed it - wedging the installer permanently.

Partial fs adapters now throw on any omitted method instead of silently
reaching the real filesystem, closing the trap that let a test poison list
pass while real IO happened.

Staged names must be flat, recovery refuses a dangling destination symlink,
and a batch whose recovery genuinely failed is no longer swept - it was
discarding the only durable copy of the file it had just failed to restore.

Replaces three tests that could not fail, including the one labelled negative
proof.

Known limitation, documented not closed: concurrent installs sharing a staging
key can still lose a batch. A real fix needs a cross-process lock.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enh(#2875): make the descriptor authoritative for the agents kind

Deletes the inline agent-staging loop in bin/install.js and the
_DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from
its capability descriptor instead of an inline hostBehaviors dispatch.

Closing it needed three pieces of contract the descriptor pipeline never had,
all reducible to one missing input - per-agent resolution context: a
frontmatter-extensions step for claude's effort and disallowedTools, per-agent
model-override resolution for kilo and opencode, and a named branding
converter for hermes, whose rewrite data was already declared.

Seven runtimes were on the loop, not the six the design recorded - kimi-code
was found by a golden fixture, not by analysis. claude-local and kimi-code
both silently lost their agents mid-change; the fixtures caught both and the
cause was fixed rather than the fixtures regenerated.

A parity harness gates the migration: both pipelines over identical inputs,
byte-identical output including filenames, per runtime. It is demonstrated
red before being trusted. Surface and install paths converge for all seven,
which also fixes surface previously writing no agents for these runtimes.

Codex's config.toml strip stays put - it mutates host config, which no
descriptor kind models.

Also routes install-model-override-resolver and install-effort-resolver
through the install fs seam. Both leaked real filesystem IO from the install
call tree; the stricter adapter is what exposed them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): record the agents-descriptor migration and correct the ADR count

The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host
integration guide told readers to join a set that is gone. Replaces that with
what is now true - declare an agents entry and it installs, on the surface
path as well as install - and points anyone needing a per-agent transform at
the three extension points rather than at a new inline branch.

Corrects the ADR amendment: seven runtimes were on the inline loop, not six.
kimi-code was found by a golden fixture going red, not by reading. That is the
third short count this phase, all from enumerating by symbol or set membership
when the thing that matters is a behavior.

Adds the Changed changeset for the surface-path convergence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-2866 - claude global always wrote agents on disk

The claude row's global=[skills] described what capability.json declared, not
what the installer wrote. bin/install.js's inline agent-staging loop was never
scope-gated and never consulted the descriptor, so a claude --global install
has always written agents/gsd-*.md.

Phase 6 closes the gap by deleting that loop and declaring agents on claude's
descriptor at global scope. On-disk bytes are unchanged - the golden fixtures
did not move, which is the evidence that the descriptor, not the installer,
was incomplete.

#2218 is unaffected: agents are not trigger-bearing, so the wider row does not
introduce a new shadowing case.

Records the warning that an incomplete descriptor is invisible while a second
code path silently does its work, and only surfaces when the two are forced
into agreement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close review findings across staging, agents and the parity harness

Two independent reviews of this branch found defects the local gates missed.

Security: a dangling symlink at a migration destination allowed writing
outside configDir - the same class this change claimed to close, missed at the
terminal write of the flow being added. The staging-root resolver threw as the
first statement of install and uninstall, so a hostile symlink bricked both,
and symlinked-configDir users lost uninstall as well as install; it now
degrades instead of aborting. Recovery gained a source-side symlink check and
now refuses a relative destDir, which resolved against cwd. Converter dispatch
gained a runtime allowlist - lint-time validation stopped mattering once this
branch promoted that dispatch from the surface path to real installs.

Correctness: claude --local --minimal exited 1 because the minimal profile
legitimately yields zero agents and the new path treated that as a failure.
cline --local silently lost its agents - its descriptor declared none while
the deleted loop wrote them unconditionally. The agents prune was widened to
any gsd-* entry and destroyed user files it never owned.

The parity harness, on which the migration's safety argument rested, drove a
synthetic registry and never byte-compared the shipped descriptors; two of its
trap rows could not fail. It now drives the real registry across 13
runtime-scope rows including kimi-code and cline-local, and its red-proof is
demonstrated by corrupting a live capability.json. Three goldens that had
encoded the cline regression as expected behavior were corrected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close findings from both mandated review engines

/security-review found the staging source-side walk honouring
GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write
destination. A symlinked files/ component dereferenced because
copyPreservingSymlink lstats the leaf only, so an intermediate link is
followed. The source walk no longer honours the opt-in; the destination check
still does.

/code-review spec axis found this branch had reintroduced its own bug:
migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after
the legacy dir was wiped and before the staged batch was restored, so a
planted symlink bricked uninstall permanently and orphaned the batch. Refusal
kept, abort removed.

kimi-code local silently lost its agents, the same class as the cline bug, and
the parity harness recorded that exclusion as intentional - the third test in
this branch to pin a regression as correct.

--minimal now creates an empty agents/ dir that never existed. Behaviour
restored rather than softening the changeset, so its byte-identical claim
stays true.

Standards axis: try/finally removed from twelve test bodies, fast-check
properties added for parseOwnerPid, boundary coverage at the grace window and
the ancestor-probe depth, a parity assertion for the staging-root helper
duplicated across two files, and the 8-deep config walk deduplicated.

Records 60-review.json with every finding and disposition from five passes,
including the smells left unfixed and why.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): prune stale agents unconditionally in minimal mode

The previous round stopped an empty agents/ directory being created when the
resolved profile yields no agents. That was implemented by skipping the agents
kind entirely, which also skipped its stale-agent prune - so a full to minimal
downgrade left stale gsd-* agents behind.

The deleted inline loop pruned unconditionally and only skipped writing. Those
are three separate conditions, not one: prune always, write only when there is
something to write, create the directory only when writing.

Both call sites now run _removeGsdEntries before the empty-staged early exit.
The symlink-escape guard moved with it, since the prune also touches dest.
Codex .toml agents and the config.toml stanzas are cleaned again, and
user-owned agents are still preserved.

The agents/ directory is left in place after a prune empties it, matching
every sibling kind - none of them remove the destination directory itself.

Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh
install, so fixture generation is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): document interrupted-install recovery for user-owned files

The durable-staging fix is invisible to the user it protects. Someone whose
install died mid-flight has no way to know USER-PROFILE.md was staged before
the delete, that the next run restores it, or that recovery happens at the
start of that run rather than in the background.

Written as the task the user has - finish the interrupted command - rather
than as a description of the mechanism, and states what it will not do:
overwrite a file already present, or touch staging belonging to another
install still running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2875): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2875): assert the J8 model override without building a regex

CodeQL flagged incomplete string escaping: the assertion interpolated the
override value into a RegExp while escaping only forward slashes, which is
meaningless in a constructor, leaving real metacharacters unescaped.

The failure direction was the dangerous one - a metacharacter would have made
the match more permissive, so the row would pass when it should fail. That
matters here because J8 exists precisely because an earlier revision was a
tautology; the rewrite reintroduced a different way for the same assertion to
stop discriminating.

Replaced with a line-wise exact match, so no regex is constructed at all.
Swept the other test files this branch adds; no sibling instances.

lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full
metachar-escape copy, so a single slash replace slipped under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 17:25:53 -04:00
Tom Boucher
8a56595700 docs(#3574): record ADR-3574 install materialization primitives (#3575)
* docs(#3574): record ADR-3574 install materialization primitives

Epic #2866 phase 6 was scoped on the premise that materializing a layout
is implemented three times and should become one module. Measured against
the tree, the premise does not hold: the three sites overlap in shape and
diverge in mechanism.

applySurface prunes by allow-list precisely so it structurally cannot
delete a user's files. installRuntimeArtifacts wipes a prefix-scoped set
and restores a snapshot. A single writer has to pick one, and picking
either trades a working guarantee for a different one.

So the ADR declines phase 6's first acceptance criterion and says why,
because the next reader who notices three similar loops should find this
file rather than rediscover the conflict. What is extracted instead is the
genuinely shared part: durable user-artifact staging for #1874-F19, reusing
the migration primitive that copies strictly before delete and never
dereferences a symlink, plus the retired-kind prune both callers already
share. The agents bypass closes on its own terms.

Two of the issue's premises were also stale: ten runtimes have already
migrated off the inline agent dispatch, and the duplication comment's
deliberate-until condition is partly met.

Closes #3574

* chore(#3574): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-16 14:19:56 -04:00
Tom Boucher
c5b83cb050 chore(#3560): delete two unreachable workflows, gate workflow reachability in lint (#3564)
* chore(#3560): delete two unreachable workflows, gate reachability in lint

discovery-phase.md and plan-milestone-gaps.md shipped to all 19 runtime
install trees with no command, agent, or skill referencing them.
plan-milestone-gaps' command was deleted by #2790 and the workflow was
left behind; discovery-phase's own header claimed a caller in
plan-phase.md's mandatory_discovery step, and that step does not exist —
plan-phase.md contains zero occurrences of "discovery".

docs/INVENTORY.md asserted discovery-phase.md was an alternate entry for
/gsd-new-project. new-project.md never referenced it. The row and the
matching note sentence are removed across all five locales rather than
corrected.

Adds rule 6 to lint-command-contract: every shipped workflow must be
reachable from a loader, walking the transitive closure over the three
reference shapes this repo uses. The closure seeds ONLY from
commands/agents/skills, so a workflow that references only itself and a
pair that reference only each other are both correctly reported rather
than satisfying themselves; a visited set makes reference cycles
terminate. The measure is a mention in a LOADER — docs/ and install-tree
fixtures deliberately do not count, because scan.md proved a file can be
documented and shipped while entirely unreached.

Ships blocking, not report-only: #3561 is in this branch's base, so the
tree reports 0 unreachable from the start.

Closes #3560

* test(#3560): drive rule 6 end-to-end, sweep a stale allowlist, update ADR-0002

Review findings.

Rule 6 had no end-to-end coverage: the tests exercised the pure closure
with in-memory data, so the wiring — file collection, exit code,
diagnostic — was unproven, and #3560's acceptance list explicitly wants
a fixture showing the rule FAILS on a planted orphan. Adds an optional
--root to lint-command-contract (default behavior unchanged) and four
tests driving the real CLI through the process seam against a temp
fixture: clean=0, planted orphan=1, orphan referenced only from docs/=1,
orphan reachable transitively=0. The docs/ case is what pins the
Goodhart defense — a mention outside a loader must not confer
reachability.

Deletes two tests that were byte-identical to a third and could not
assert anything loader-specific, since the closure is source-agnostic by
design; that distinction lives in the lint script's file collection and
is now covered above.

Removes a stale ALLOWLIST entry for discovery-phase.md in
planner-language-regression — the exact sweep-miss class rule 6 exists
to catch, found in the PR that adds the rule.

ADR-0002 described five per-file frontmatter checks; rule 6 is a
repo-level reachability graph, so the Decision section now says so.

Refs #3560

* test(#3560): cut the bug-3298 test pin on the deleted plan-milestone-gaps workflow

The remote runner went red with four failures: tests/phase.test.cjs
asserted the plan-milestone-gaps workflow exists and checked its mkdir
patterns, so deleting the file broke the test that pinned it. This is the
fence the epic describes — the content-sync test IS what keeps an
unreachable file alive — and cutting the coupling is what makes the
deletion safe.

Removes only that arm. The bug-3298 block guards three workflows against
phase-dir prefix drift; the import and add-backlog arms and both shared
mkdir-pattern helpers are untouched.

Worth recording where the sweep failed: my reachability walk covered
commands, agents, skills, gsd-core and docs, and lint-removed-but-needed
covers .github/workflows, gsd-core, docs and package.json. Neither looks
at tests/, so a test-pinned deletion is invisible to both and surfaces
only on the remote runner. The how-to added by this PR names that gap
explicitly so the next deletion searches tests/ by hand.

Refs #3560

* docs(#3560): add a how-to for resolving unreachable-workflow findings

* chore(#3560): backfill changeset pr number to 3564

---------

Co-authored-by: sim <sim@local>
2026-08-15 23:29:30 -04:00
Tom Boucher
1591454357 feat(#3409): reject shell guards that cannot observe their own failure arm (#3558)
* test(#3409): failing-first regression tests for unreachable shell guard arms

Drives the three live defects fail-first, executing the shipped workflow
snippets rather than a re-typed copy:

- G1/G2 plan-phase.md Walking Skeleton gate reads `--pick summaries_total`,
  a field that does not exist, so PRIOR_SUMMARIES is always "" and the gate
  has never fired (#3365). G2 is the load-bearing negative-space case: it
  rejects a fix that treats "no answer" as "zero" and fires unconditionally.
- G3 plan-phase.md PHASE_REQ_IDS resolves "" instead of the TBD sentinel on
  a phase with zero requirements.
- G4 complete-milestone.md's bare `cat <glob>` blocks on stdin under a
  nullglob left set by an earlier block (measured hang).

Skipped on Windows for G4 only: the FIFO-blocked-stdin mechanism is POSIX
only, and a weakened assertion there would pass vacuously.

Refs #3409

* fix(#3409): make nine shell guards observe their own failure arm

`--pick` coerces a missing field to empty string and exits 0, so the
`|| echo <default>` fallback after it fires only on a verb typo, never on
the field absence it was written for. Nine sites relied on that arm.

- plan-phase.md walking-skeleton gate: `--pick summaries_total` names a
  field that does not exist under any flag combination, so the gate has
  never fired on any project (#3365). Repointed at the existing single
  owner, `phases.list --type summaries --pick count`, which returns a real
  integer in every case including a project with no `.planning` directory.
  No new counter is added: a second one would duplicate the ownership
  ADR-3180 Decision 1 forbids. The gate now fires only on a literal "0",
  so an unanswerable query fails safe instead of entering skeleton mode.
- plan-phase.md phase_req_ids: now falls back to the documented TBD.
- The remaining seven convert to an explicit empty test.
- complete-milestone.md read all phase summaries through a bare
  `cat <glob>`; under a nullglob left set by an earlier block that is zero
  operands, so cat blocks on stdin. Guarded with the array shape the
  #3300 fix already established in review.md.

Refs #3409

* fix(#3409): guard eleven more globs that defeat their own fallback arm

The nullglob audit this issue asks for turned up the same class in files
#3300 never touched.

- Eight bare `cat <glob>` reads (transition, complete-milestone, planner x4,
  verifier, phase-researcher). With nullglob set that is zero operands, so
  cat reads stdin and blocks; measured rc=137 at 3s.
- Three `ls <glob> || echo "<message>"` sites (session-report,
  review-backlog and its generated skill). nullglob makes ls succeed
  listing the cwd, so the message never prints and the user gets a
  directory listing instead.

Guarded with `[ -e "${_ARR[0]}" ]` rather than `[ ${#_ARR[@]} -gt 0 ]`.
The count form is correct only when nullglob is set, and six of these
seven files never set it: without it the array holds the unmatched literal
pattern, so the count is 1 and the guard passes wrongly. `-e` is correct
in both worlds. review.md keeps its count guards — that block sets
nullglob two lines above them.

skills/gsd-review-backlog regenerated from commands/, never hand-edited.

Refs #3409

* feat(#3409): add the unreachable-shell-guard drift lint

A sibling of lint-planning-prompt-drift.cjs, consuming the shared
scripts/lib/drift-scan.cjs rather than copying it, wired into lint:ci.

Both detectors are one shape — a fallback arm defeated by a legitimate
success-on-empty:

- Detector A: `--pick` and `|| echo` on one line. `--pick` is the
  discriminator because "missing field renders empty at exit 0" is a
  documented CLI contract, not a heuristic. A rule keyed on gsd_run
  matched 111 lines, ~132 of them legitimate, and was rejected.
- Detector B: `cat <glob>` in command position, and `ls <glob>` whose
  exit code feeds a real fallback or an if/while head. Informational
  `ls <glob>` whose stdout is consumed (97 sites) and `|| true` failure
  suppression (~15) are not guards and never fire.

Shrink-only ratchet keyed on (file, trimmed text) with a per-pair count,
POSIX-normalized unconditionally so Windows CI cannot report everything
fresh and stale at once. Ships with a ZERO-entry baseline: every site it
can find is fixed. Exemption is the per-line `# gsd-scan-ignore: #NNN`
marker whose reason must name an issue or URL; a malformed reason reports
a distinct error rather than silently exempting. No file allowlists.

ADR-3409 records the invariant, the measurements behind both detectors,
and why the upstream `--pick` contract fix belongs to #3473.

Refs #3409

* fix(#3409): resolve review findings — typed surface, sanitized reports, tighter marker

Standards axis (blocker): the guard's tests asserted on human-readable
stdout/stderr and on free-form baseline-load prose, which CONTRIBUTING
prohibits by name. Added the typed surface it prescribes instead of
weakening the tests: a frozen REASON enum, a --json report mode,
structured loadBaseline errors, and a test locking Object.keys(REASON)
so a new reason stays three coordinated changes.

Security axis: sanitizeForReport covered every violation field but not
the baseline-load error path, which embeds raw JSON.stringify output --
that escapes nothing above 0x1f, so bidi and C1 controls reached CI logs
unfiltered. Routed through the sanitizer at the output seam.

Security axis: the scan-ignore marker accepted `#0` and a bare
`http://`. Tightened to a positive issue number and a URL with a host.
This diverges deliberately from the sibling in
tests/commit-files-pathspec.test.cjs, whose looser form was copied
verbatim; the header now records the divergence.

Security axis: G4 built its FIFO with `mktemp -u`, reserving a name
without creating it. Now created inside a `mktemp -d` directory.

Spec axis: ADR-3409 claimed a ninth site landed after the issue was
filed. git blame disproves it -- all nine predate it; the issue's hand
count missed one. Corrected. The design and test matrix still specified
B9 as a FLAG after implementation reversed it to PASS; both now record
the reversal and why.

Refs #3409

* docs(#3409): add the how-to for resolving unreachable-guard findings

Reference and Explanation are carried by ADR-3409; this is the
task-oriented quadrant CI cannot check for.

The page exists mainly for one thing the lint structurally cannot catch:
both `[ -e "${_ARR[0]}" ]` and `[ ${#_ARR[@]} -gt 0 ]` remove the glob
from the command and therefore both pass, but the count form is correct
only when nullglob is set — and nullglob is usually set in a different
block of the same file. A reference table cannot carry that; a how-to can.

Also documents the reason codes, so a reader can tell "nothing to report"
from "could not look".

No tutorial: this is a gate inside an existing CI loop, not a new entry
point a newcomer starts from.

Refs #3409

* fix(#3409): bring the touched prompt files back under their size gates

The remote run was red on 14 tests, all size/attribution, none of them
the regression suite.

- agents/gsd-planner.md was 194 chars over a 49152 cap enforced by four
  separate tests, each of which says the remedy is extraction, not a bump.
  It had 41 chars of headroom before this branch. Its `## Checkpoint
  Types` section was an unlinked, condensed duplicate of
  references/checkpoints.md, which already carries all three types and
  their XML shapes; the section now points there and keeps the three
  names and percentages inline. Net -969, margin 1010.
- gsd-core/workflows/execute-phase.md sat 2 chars under a comfortable
  margin assertion. Dropped the AUTO_MODE default: the `|| echo "false"`
  it replaced was unreachable, so the value was already sometimes empty
  on next, and its only consumer compares against `true`. Net -16.
  Left plan-phase.md's AUTO_CHAIN default alone -- that file names an
  explicit `false` branch, so empty would match neither branch.
- Acknowledged the seven prompt files that genuinely grew, one specific
  reason each. Five of those paths were already claimed by spent
  fragments identical to next, which blocks a second source naming the
  same path; removed just the colliding key from each, deleting the two
  that this emptied.

Refs #3409

* test(#3409): extract the whole PHASE_REQ_IDS block, not just its first line

G3 failed on the remote runner with '' !== 'TBD'. The test was wrong, not
the workflow.

The shipped contract is now two consecutive lines -- the capture and the
`${PHASE_REQ_IDS:-TBD}` default -- but the helper's `^PREFIX=.*$` regex
returns only the first match, so the test executed half the contract and
correctly observed the empty string. Renamed to extractAssignmentBlockFor
and taught it to consume the contiguous run of lines sharing the prefix.

The assertion is untouched: TBD is the right expectation, and weakening
it to accept the empty string would have reinstated exactly the class
this suite exists to catch -- a check that cannot observe the thing it
is checking.

extractFencedBashAfterAnchor is unaffected: it is fence-delimited rather
than line-anchored, so G1/G2/G4 still capture their full blocks.

Refs #3409

* chore(#3409): drop a spent ack fragment that collided on complete-milestone.md

#3458 landed on next while this branch was in flight and its fragment
claims complete-milestone.md, which this branch also grows. Two ack
sources may never name the same path.

Its entry is spent: the +9163 it explains is already absorbed at base, so
it can no longer clear anything, and the checker's own guidance for spent
entries is to delete them. Removing the key emptied the fragment, so the
file goes too -- an empty one signals nothing.

Refs #3409

* chore(#3409): backfill changeset pr number 3558

* test(#3409): hoist a regex subject out of exec() to clear the injection scan

CI's prompt-injection scan flagged `MARKER_RE.exec('# gsd-scan-ignore: ...')`.
The pattern `exec[[:space:]]*\(["']` is receiver-blind on purpose, so it
catches `require('child_process').exec('...')` -- and the scanner's own
header records that RegExp.prototype.exec is collateral, to be handled by
its allowlist.

Allowlisting the file would blind it to the real exec vector permanently,
so the subject is hoisted into a const instead: same assertion, scanner
left at full strength, no security surface widened.

Refs #3409

---------

Co-authored-by: sim <sim@local>
2026-08-15 21:09:05 -04:00
Tom Boucher
6badb839a0 fix(#3514): deny internal fetch hosts; disclose unverified integrity (#3516)
* test(#3514): add failing-first denylist and integrity suites

* fix(#3514): deny internal fetch hosts; disclose unverified integrity

* docs(#3514): trust-model, glossary, and changeset entries

* fix(#3514): scope v6 checks to literals; exact pin kinds in prompt

* chore(#3514): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:34:29 -04:00
Tom Boucher
411196bc3a refactor(#3471): one enforcement point for the empty case, and reports that match the disk (#3519)
* refactor(#3471): one enforcement point for the empty case, and reports that match the disk

Implements ADR-3408 section 8.5 and section 8.4's residue (folded in when Phase
3 closed as subsumed). Four items, and two findings the design did not predict.

FINDING 1 — the guards could not simply be deleted, as the design instructed.
state sync and REGENERATE_STATE never run applyStatePreservation at all, so
those six conditions were their ONLY empty-field fallback. A baseline probe on
the unedited tree confirmed unconditional deletion drops current_phase,
current_phase_name, current_plan, stopped_at and paused_at from a blank-body
STATE.md on state sync — breaking the byte-identical requirement section 8.3
grants those two sanctioned-permanent exceptions. They are now GATED, not
deleted: on for the exceptions, off for the write seam, where an empty derived
value finally reaches the executor unmolested.

FINDING 2, the more serious one — there was a FOURTH encoding of this policy.
The pre-existing #2202 unknown-key carry-forward loop independently restored
the same six fields whenever derivedFm lacked the key, completely neutralizing
the fix. It is named nowhere in the ADR, the design, or three prior phases. It
was found only because a probe that should have passed did not: the first
attempt reported divergedFields: [] and silently restored both fields,
reproducing the exact bug this phase exists to close.

That is worth stating plainly. This epic's thesis is 'policy declared in one
table, enforcement hand-rolled per call site.' The final phase found one more
call site than anyone had counted — which is the fourth consecutive time a copy
count in this epic proved to be a lower bound.

Also: divergedFields could only observe fields the executor actively RESTORED,
by diffing postFm. A discard-to-empty is absent both before and after, so it
was invisible. A second pass now reports it, which is what makes section 8.5's
'preservation is visible' true for the delete-the-body-line case rather than
aspirational.

cmdPhaseComplete now reports what it preserved — #3374 was filed against that
command and its complaint was warnings: [], silence.

cmdStateJson's private third copy of the guards is routed onto the executor's
preserve-when-unchanged rule. A read is definitionally not a write, so the
#1230 delta is 'unchanged' and curated wins over a stale annotation.
shouldPreserveExistingProgress is a different rule and is untouched.

Report reconciliation is ONE shared helper across seven commands, not five
copies of fix(#3351)'s block. Five copies of a reconciliation is precisely the
shape this epic removes, and introducing it in the final phase would have been
a poor joke. Both untraced commands were traced rather than assumed:
cmdStatePlannedPhase matched cmdStateBeginPhase exactly; cmdStateCompletePhase
turned out to be a different legacy hand-rolled path reporting a mix of field
names AND a section name, where the naive helper would have dropped 'Current
Position' as a false negative every time.

* test(#3471): characterization coverage for one enforcement point and reconciled reports

Matrix sections A-E, asserted at the consumer's output per ADR-3180 Decision
4(b)/(c) — this phase owes Decision 5's outcome metric, the one the drift
guard's zero may never be reported without.

Three walls matter more than the new coverage:

  A2 is SIX separately named tests, one per gated guard, not one parameterised
  assertion over a list. A list is trivially shortened later; six named tests
  are not, and six guards is exactly where a field gets silently dropped.

  A6 pins what Phases 1-3 already fixed — non-empty stale body, delta
  unchanged, losing to fresher curated frontmatter, with the divergence
  reported. If A6 reddens, this phase broke the thing the epic was for.

  D1/D2 pin state sync byte-identical. The implementation had to GATE the six
  guards rather than delete them precisely because state sync has no executor,
  and a baseline probe showed unconditional deletion drops five fields.
  Nothing else in the suite would notice that regression.

E6 covers #3345's direction — a field preservation restored that the intent
never named IS reported. Nothing has ever tested that direction.

Assertions were empirically verified against the compiled lib and the real CLI
before being written, since the suite cannot be executed locally. That caught
two type bugs in the draft: fm.current_phase after a quoted-YAML round-trip is
the string '5', not the number 5.

E5 is recorded as structurally unreachable rather than weakened or faked. Those
four commands report body Title-Case labels, which cannot string-collide with a
frontmatter snake_case key the way cmdStatePatch's arbitrary field names can —
which is why fix(#3351) targeted only cmdStatePatch. Testing it directly would
need reconcileReportedFields exported from private scope; the helper is
exercised through E6 and all seven commands instead.

* docs(#3471): amend ADR-3408 section 8.5 — a fourth enforcement point, and guards that could not be deleted

Amendment 3. The contract held; two of section 8.5's own statements did not.

It said the six empty-only guards are DELETED. They cannot be. writeStateMd is
the sole path for both section 8.3 sanctioned-permanent exceptions and never
runs applyStatePreservation, so those guards were their only empty-field
fallback. A baseline probe on the unedited tree confirmed unconditional
deletion drops five fields from a blank-body STATE.md on state sync, breaking
the byte-identical guarantee section 8.3 grants it. They are gated instead.

It also mis-located cmdStateJson's guards, describing them as living in
syncStateFrontmatter. They were a separate private copy on the read path with
no delta check at all, so a stale body annotation always beat fresher curated
frontmatter in state.json — #3395's shape entirely outside the write seam.

THE FINDING: a fourth enforcement point nobody had counted. The pre-existing
#2202 unknown-key carry-forward loop independently restored the same six
fields, silently neutralizing the fix. It is named nowhere in this ADR, in the
phase design, or in three prior phases, and was found only because a probe that
should have passed did not.

Fourth consecutive time a copy count in this epic proved a lower bound: 2
write-seam bypasses became 4, three preservation encodings became four, and the
estimate was wrong every time. ADR-3180's standing rule has earned itself in
every phase — read the code, not the write-up.

Records the Row 2 decision (a discard-to-empty wins per the delta rule and is
reported, not silent — the sharpest Hyrum exposure in the epic), section 8.4's
residue landing as ONE shared reconcileReportedFields across seven commands
rather than five copies, and the parity assertion added because
FRONTMATTER_KEY_TO_BODY_LABEL was itself a second table that failed silently —
this epic's shape in miniature, in its final phase.

* fix(#3471): repair four regressions the checkpoint caught

Checkpoint returned 16 failures of 34389: six real regressions in pre-existing
tests, plus seven of my own test bugs.

My hypothesis was wrong and is recorded as such. I predicted the #2202
carry-forward skip was the cause, reasoning it had removed a load-bearing
fallback the way the six guards nearly were. It was not implicated in any of
the six. Three unrelated causes:

#2111 — current_phase came back undefined from milestone complete, which is
the epic's own defect class reintroduced by its final phase. Root cause is
Row 2 working exactly as designed: milestoneCompleteCore rewrites the body
Phase: line to a closure message, so current_phase's #1230 delta reads
CHANGED and the new rule correctly discards the curated value. The transition
never declared any intent to touch that field. Fixed by re-asserting
current_phase and current_phase_name through authoritativeFm — the existing
#2736 mechanism beginPhaseCore and completePhaseCore already use — rather
than by weakening Row 2, which A5 pins.

That interaction is worth naming: a rule that keys on 'did this write change
the body source' will fire on a transition that moves the body line for an
entirely unrelated reason. The design did not anticipate it.

#1264 / #3242 / the state.patch progress report — reconcileReportedFields
folded EVERY divergedFields entry into updated, including preserve-always
progress restores no caller asked about. Now scoped to preserve-when-unchanged
rows only.

#1162 / case-insensitive table fields — valueOf checked frontmatter before
body, so a lowercase table field name exact-matched the lowercase frontmatter
key sync always derives, comparing stale pre-sync body text against a
post-sync frontmatter enum. Flipped to body-first.

That last one is the SAME lesson as Phase 2's patchCore, recurring in a
different function two phases later: in this model the body is authoritative
and frontmatter is the projection, so a name that could mean either resolves
body-first. Twice now.

Test bugs: a stray unused parameter shifted every argument at six call sites,
so body arrived undefined; and A4 compared nested progress scalars against
numbers when extractFrontmatter returns raw YAML strings. The string-vs-number
YAML round-trip has now been caught three times in this phase alone.

* test(#3471): one helper for the progress coercion that bit four times

A2f failed on the string-vs-number YAML round-trip: extractFrontmatter returns
nested progress scalars as raw YAML strings, so a comparison against numeric
literals can never pass.

This is the FOURTH time this exact class has been caught in this phase — twice
during test authoring, once as A4 in the previous checkpoint, now as A2f.
Patching it a fourth time by hand would guarantee a fifth.

Added numericProgress() with a comment saying why it exists, and routed every
progress-reading assertion in the #3471 block through it. Swept the block:
C3 needed no change, because cmdStateJson's output already runs through
normalizeProgressNumbers.

Deliberately NOT shared with frontmatter.test.cjs's readPersistedProgress:
that one is path-based and re-reads from disk, while these assert on an
in-memory string that is never written. Sharing would have meant either a
disk round-trip these tests do not do, or duplicating half the helper — so
the coercion pattern is mirrored locally and the reason recorded, rather
than manufacturing a dependency to satisfy the letter of consolidation.

* chore(#3471): backfill pr number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-08-14 21:06:01 -04:00
Tom Boucher
fba7c90327 chore(#3484): adr-0174 behavior carry-forward amendment and merge gate (#3507)
* chore(#3484): adr-0174 behavior carry-forward amendment and merge gate

* chore(#3484): regen example context index for new ruleset predicates

* chore(#3484): review fixes - amendment heading per contributor-standards, helper-based fixtures

---------

Co-authored-by: sim <sim@local>
2026-08-14 19:31:48 -04:00
Tom Boucher
e2f4c16d9e refactor(#3469): one composition for the STATE.md write seam (#3501)
* docs(#3469): amend ADR-3408 section 8.3 — the pipeline has sanctioned exceptions

Section 8.3 read 'Every STATE.md write applies the pipeline.' That is false by
design for two commands, and acting on it would have inverted a shipped
feature.

Preservation makes curated frontmatter win over a re-derived body value.
state sync exists to do the opposite — #905's 'body annotation beats existing
frontmatter when both are present'; it re-derives frontmatter FROM the body.
REGENERATE_STATE is a factory reset that rebuilds STATE.md from scratch.
Applying the pipeline to either would re-lock exactly what the command was
invoked to replace.

This issue's own scope line, inherited from the epic, said to route the direct
writeStateMd callers through the pipeline. For cmdStateSync that would have
shipped silently, with every gate green, because no test asserts that sync
LETS the body win. Caught by reading the helper's docstring and then verifying
the claim against the code — a stale comment had already misdirected this epic
once.

Both commands are now named in a closed exception list and are permanent
ratchet entries.

Consequence recorded rather than left to bite Phase 4: the 'drive the ratchet
to 0 and delete the file' target in this ADR and in #3471 is wrong. Two
entries are permanent, so the correct end state is 2, and the honest report is
'0 removable bypasses, 2 sanctioned'. A guard reaching 0 here would only do so
by having stopped looking at two real writers.

* refactor(#3469): one composition for the write seam, not one per caller

Implements ADR-3408 section 8.3 as amended.

syncAndPreserveStateMd is now the single composition of syncStateFrontmatter
and applyPostSyncPreservation. readModifyWriteStateMd and cmdPhaseComplete
both CALL it instead of each assembling the two steps themselves.
cmdPhaseComplete keeps its own writePlanningFileSet envelope — the
composition returns content, it does not take over the write, so STATE.md
still commits atomically with ROADMAP and REQUIREMENTS.

Assembling the stages at a call site is a re-derivation even when every step
calls an owner. Upstream's fix(#3374) routed cmdPhaseComplete through
applyPostSyncPreservation but left it calling syncStateFrontmatter directly
first, so the composition was duplicated and free to diverge with both guards
green. That is ADR-3180 Amendment 2's finding repeating on the write side.

cmdMilestoneComplete gains preservation. It wrote through writeStateMd, so it
got sync and no preservation — the identical shape #3374 reported for
phase.complete, and flagged upstream as a follow-up in the helper's own
docstring. This is that follow-up.

Divergence is now visible: preservation_warnings names each field restored
over a disagreeing derived value. Deliberately NOT named warnings —
cmdPhaseComplete already exposes warnings as a prose string array, and two
sibling commands carrying that name with different element types is
Generative Fix Divergence, the class this epic exists to remove.

patchCore stops running stateReplaceField over the whole document. One
observable consequence, intended per design row 9: a frontmatter-shaped patch
key with no body counterpart now reports failed instead of silently
succeeding, because the old whole-document match was literally hitting the
YAML line case-insensitively.

The guard closes Phase 1's DECLARED KNOWN GAP as promised rather than
re-deferring it: section 8.3(b) detection is tractable now the composition
exists. Scoped by two factors to avoid Phase 1's measured 29-to-1 false
positive rate — a variable field-name argument AND a content argument whose
nearest preceding assignment is not stripFrontmatter. Verified 0 findings and
0 false positives across all 33 call sites, plus 5 synthetic shapes. It also
detects the re-assembly shape above.

Ratchet: 4 entries to 2, both sanctioned-permanent. cmdStateSync's owner
changes from #3471 to sanctioned-permanent per Amendment 2 — routing it
through preservation would invert the #905 contract.

Also fixed inline rather than deferred: cmdMilestoneComplete's STATE.md read
now happens inside withStateLock. It previously read outside any lock before
writeStateMd took its own, leaving a TOCTOU window under concurrent writers.

* test(#3469): characterization coverage for the single write seam

Matrix sections A-E. Criterion 6 was amended by maintainer decision — all five
instances closed by point fixes while Phase 1 was in flight — so these are
characterization tests at the consumer's output per ADR-3180 Decision 4(b)/(c),
paired with the drift guard's count, never either alone.

Section C is the one that earns its keep. cmdStateSync is a sanctioned
permanent exception: state sync exists to re-derive frontmatter FROM the body,
so preservation there re-locks exactly what the command was invoked to
replace. C1 pins that the body wins; C4 pins that this phase left the command
byte-identical. Nothing else in the suite would notice if a future change made
sync start preserving, and the natural reading of 'one write seam' is to make
precisely that change.

Section E pins the guard's false-positive scoping. E4 (updateCore's
strip-then-replace) and E5 (sectionBody-scoped calls) must NOT be reported —
the naive detector measured 29 false positives to 1 true positive in Phase 1.
E7 is the inverse: a sanctioned-permanent entry disappearing must FAIL,
because a guard reaching zero here would only do so by having stopped looking
at two real writers.

Also corrects a stale test that asserted patchCore's old whole-document
behavior, which this phase deliberately changes.

One honest limitation, flagged rather than papered over: A1's 'byte-identical
to pre-refactor' cannot be diffed against real pre-refactor bytes from inside
the suite. It is implemented as the seeded fast-check property that
cmdPhaseComplete's composed output equals readModifyWriteStateMd's for the
same inputs — the strongest available proxy, not the literal claim.

* docs(#3469): refresh the seam glossary entry and add the changeset

Two spec-review gaps, both real.

CONTEXT.md's STATE.md Transition Module entry named three direct writeStateMd
callers including cmdMilestoneComplete. This phase routed that one through the
composition, so the line was false the moment the refactor landed.

Worth recording plainly: I wrote that sentence in Phase 0, correcting an
older stale pointer in it, and my own Phase 2 change invalidated it again
within the same epic. That is the exact drift this epic exists to remove,
demonstrated on the epic's own documentation — and it is why the entry now
ends by saying the whole-repo drift guard, not this line, is the authoritative
count.

The entry now records the composition (syncAndPreserveStateMd) and states that
exactly two direct callers remain, both SANCTIONED PERMANENT rather than debt.

Changeset: type Changed, because milestone complete's observable output moves.
Tier-2 per ADR-3180 Decision 3 — a stale body line no longer wins over fresher
frontmatter, and the command gains preservation_warnings. Docs requirement is
met by the ADR amendment already in this diff.

* test(#3469): register property-test temp-dir cleanup at creation time

Standards review, minor but real: the new fast-check property cleaned up its
temp dirs in a loop AFTER fc.assert returned. A genuine property failure
throws, so that line never ran and every dir from the failing run — including
all of fast-check's shrinking iterations — leaked.

The failure path is exactly when a littered machine hurts most, and a failing
property test is the case the test exists for.

Cleanup is now registered with t.after() at dir-creation time, so teardown
happens however the test exits. Not try/finally — CONTRIBUTING.md:356 bans it
inside test bodies, which is why the after-the-assertion shape existed in the
first place.

Swept the rest of the branch's test diff for the same shape; phase.test.cjs
already uses registered teardown and nothing else matched.

* fix(#3469): patchCore routes frontmatter writes instead of dropping them

Checkpoint returned 10 failures of 33880. One implementation defect, three
test defects, one stale test — all fixed, and the implementation defect is the
one that matters.

patchCore stripped frontmatter and then reconstructed it VERBATIM, applying no
patches to it. An arbitrary custom frontmatter key with no body counterpart and
no FIELD_CLASSIFICATION row — risk_level in the upstream fix(#3351) test —
therefore always reported failed and silently never wrote. It worked before,
via the old whole-document match on the raw YAML line.

That is a regression against this phase's own design row 9, which requires
frontmatter changes to ROUTE THROUGH the seam — still work, policy-governed —
not to stop working. Removing a capability is not routing it. An upstream test
caught it, which is the argument for running the checkpoint before believing
the refactor.

patchCore now partitions by frontmatter shape, decided structurally from the
parsed frontmatter's own keys rather than a naming heuristic:
  - classified keys still report failed — policy owns them and a raw patch may
    not bypass it;
  - unclassified keys apply to the frontmatter object and report updated —
    Phase 1's behavior-table row 19, a field with no row is not this contract's
    business;
  - body-shaped keys are unchanged.

The property 'failure' was my own test breaking the repo's Clock Seams rule.
The two paths agree byte-for-byte; the only difference was last_updated,
stamped from the wall clock on two invocations milliseconds apart, so it could
never pass. Time is now frozen with mock.timers across both — not by excluding
last_updated from the comparison, which would have silently stopped comparing
a field the composition writes.

B4's fixture could not discriminate: normalizeStateStatus maps any text
containing 'complete' to 'completed', and milestone complete's own new body
value derives to exactly that — which was also the fixture's stale value. The
stale value is now 'executing' so the assertion can tell 'body correctly won'
from 'stale survived'.

B5's fixture tripped a pre-existing unstarted-phase guard before reaching any
write-seam code; it now has the matching phase directory.

D9 asserted the old exempt set. readModifyWriteStateMd now calls one symbol
rather than assembling two, so it needs no exemption; syncAndPreserveStateMd
is the sole legitimate composition site.

* fix(#3469): patchCore resolves body-first, so the body wins a name collision

Re-verification returned 2 failures of 33880, both D4 — the hostile row for a
key that exists as BOTH a frontmatter key and a body field.

The partition checked frontmatter first, so 'status' — classified in
FIELD_CLASSIFICATION and also present as a body 'Status:' line — routed to the
frontmatter branch, was rejected as classified, and reported failed.

Wrong order. Patching 'status' means the body field, and upstream fix(#3351)
says so in its own comment: 'the legitimate working case for state.patch is
display-cased BODY fields — Status, Current Plan, Phase.' The body is
authoritative in this model; frontmatter is the projection. D4 asserted
exactly that and was right.

Resolution order is now body, then frontmatter:
  1. resolves to a body field -> apply to body, updated
  2. else an own key of the frontmatter:
       classified   -> failed  (policy owns it)
       unclassified -> apply to frontmatter, updated
  3. else -> failed

Verified by probe against the compiled lib for all four cases rather than
asserted: risk_level (frontmatter-only, unclassified) still lands;
current_phase still fails; display-cased Status unchanged; D4's lower-cased
status now lands via the body with the frontmatter untouched.

The current_phase case was the one that could have regressed silently, so its
fixture was read rather than assumed — D1's body carries 'Phase: 3 (alpha)'
and no 'Current Phase:' line, so body-first cannot reach it.

* chore(#3469): backfill pr number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-08-14 16:04:09 -04:00
Tom Boucher
1218d76d62 refactor(#3468): dispatch state preservation on the declared policy, not the field (#3495)
* test(#3468): add write-path drift guard, ratcheted at its measured baseline

Guard-first, per ADR-3180 Amendment 3's standing rule that a phase builds
and runs its guard BEFORE its scope is fixed, and states its copy count as
'N found by the guard', never 'N per the epic'.

Measured, not assumed:

  Axis 1 (policy dispatch, ADR-3408 section 8.1) — 7 violations, RED by
  design. 5 field-name-keyed getFieldClassification('literal') branches
  plus 2 declared FieldPreservation members with no executor at all
  (derive, clear). This is the fail-first evidence for the refactor.

  Axis 2 (write seam, section 8.3) — 4 bypasses, ratcheted. Epic #3408
  scoped this at two writers; the whole-repo scan found four, and one the
  epic named (patchCore) is not among them because it bypasses via
  stateReplaceField rather than the seam calls. Fourth consecutive time an
  epic's copy count proved a lower bound.

Two detectors were written and removed again before this commit, both
recorded in the file header rather than silently dropped:

  - A prompt-layer detector that reported 5 backticked prose mentions as
    drift. That is ADR-3180 Amendment 3's recorded false-positive class,
    and CONTRIBUTING.md already settles it: a backticked command reference
    is a mention. Now gated on inline-code spans.

  - A stateReplaceField co-occurrence detector for section 8.3(b). Measured
    at 29 false positives to 1 true positive — it matched the function's own
    definition and ~20 calls on frontmatter-free body slices. Banking 29
    non-defects to catch one is the 'ratchet as a parking lot' gaming route
    Decision 5 names, so it is a DECLARED KNOWN GAP owned by Phase 2
    (#3469), which both fixes it and makes its detection tractable.

* test(#3468): failing-first coverage for policy dispatch and the loud failure

Matrix sections A, B and C from 50-test-matrix.md.

Expected RED against this tree, confirmed by static trace rather than
assumed:

  B1, B2, B3 — an unwired declared preserve-when-unchanged row must throw
  with code STATE_PRESERVATION_UNWIRED_ROW and a structured .field. Today
  src/state-transition.cts:314 silently continues.

  A4 — a whitespace-only snapshot is restored today, because the guard is
  .length > 0. Required behavior is skip.

Everything else is characterization, locking in behavior the refactor must
preserve. C1 is table-driven over every FIELD_CLASSIFICATION key; C2 pins
current_phase_name's exact outputs as literals, because its row is being
reclassified preserve-always to preserve-when-unchanged as a
behavior-preserving change and nothing else would catch a drift. C3 is a
seeded fast-check property (seed 3468, 200 runs, replay data on failure).

A22 is deliberately NOT a behavioral test. Whether 'derive' has an explicit
executor is not observable through applyStatePreservation's public API — it
is a structural property, and the drift guard's unimplemented_policy axis is
what enforces it. That split is ADR-3408 Decision 5's own pairing: the lint
is the structural metric, the test is the outcome metric, and neither is
reported alone.

* refactor(#3468): dispatch preservation on the declared policy, not the field

Implements ADR-3408 sections 8.1, 8.2 and 8.6.

applyStatePreservation is now one loop over FIELD_CLASSIFICATION dispatching
on the row's preservation value, with four small executors — one per
FieldPreservation member. No branch is selected by field name. Zero
literal-argument getFieldClassification calls remain.

Behavior-preserving for 16 of 20 input classes. The four that change:

  - An unwired declared preserve-when-unchanged row now THROWS
    (code STATE_PRESERVATION_UNWIRED_ROW, structured .field) instead of
    silently continuing. This fires only on an internal invariant violation
    with both ends in our own source; a drifted, malformed or unparseable
    user STATE.md must never reach it, which is section 8.2's bright line
    and what test B8 proves through the real CLI.
  - derive gained an explicit no-op executor. That is what makes the throw
    decidable: 'policy says do nothing' is now distinguishable from 'nobody
    wired this'.
  - current_phase_name's row is corrected from preserve-always to
    preserve-when-unchanged. The row was wrong, not the code — it has always
    been delta-gated on the body Phase line, so preserve-always had two
    divergent implementations. Behavior is unchanged and test C2 pins it.
  - A whitespace-only snapshot is no longer restored; the check is trimmed.

clear is deleted from the FieldPreservation union — no row used it and no
executor existed. Speculative Generality: a policy invented for a need that
never arrived. Verified zero dependents.

The caller folds six dedicated pre/post parameters into one bodyDeltas map
keyed by field, so all seven preserve-when-unchanged rows travel one channel
instead of two. Two shapes for one kind of data is why the executor needed
per-field branches at all.

Also fixed, found while reviewing the refactor rather than deferred:

  - applyPreserveIfPlaceholder opened with a field-name literal test, which
    section 8.1 forbids outright. The executor is idempotent, so the test
    bought nothing. The drift guard could not see it, so Axis 1 is widened
    to catch field-variable comparisons against literals — the guard
    reported zero while a violation sat in the file it polices, which is
    Goodhart's gaming-by-indirection.
  - loadBaseline conflated an unreadable baseline with an absent one. A
    guard whose own diagnostic collapses two states into one identical
    result reproduces the exact failure shape this epic exists to remove.

* docs(#3468): record Phase 1 validation as ADR-3408 Amendment 1

Amendment 1 records what Phase 1 found, per ADR-3408 section 8's rule that a
behavior it does not state is not decided:

- preserve-always had TWO divergent implementations; current_phase_name's
  row was wrong and is reclassified, behavior unchanged.
- section 8.6 resolved: clear is deleted, zero dependents.
- the closed guard vocabulary is real and has exactly one true member,
  because stopped_at's scoping turned out to be caller-side extraction.
- copy count found by the guard: 4 write-seam bypasses where the epic
  scoped 2, and patchCore — one of the two it named — is not among them.
- two detectors built and removed again, with their measured false-positive
  rates, so nobody re-attempts them.
- a DECLARED KNOWN GAP for section 8.3(b), owned by Phase 2.
- Decision 5's anti-gaming list earned itself twice in one phase.

Also adds the changeset fragment.

* test(#3468): fix review findings — try/finally, stale clear allowlist, ratchet owners

Standards axis, both hard violations:

  - tests/state-write-path-drift-guard.test.cjs wrapped stdout/argv/exitCode
    restoration in try/finally inside the test body. CONTRIBUTING.md:356
    forbids it outright, and the correct t.after() pattern was already in
    use two lines up in the same test.

  - tests/state-transition.test.cjs still listed 'clear' as an allowed
    FieldPreservation value in the row-enumeration test AND the
    getFieldClassification property test, after this PR deleted it. A stale
    allowlist weakens the property's negative space — it would accept a
    resurrected clear row as valid.

Contract tension, resolved rather than left:

  ADR-3408 section 8.3 requires each ratchet entry carry the issue owning
  its removal. All four shipped with owner: null. The guard was right not to
  INVENT one, but the owners are known from the phase plan, so recording
  them is not inventing: phase.cts -> #3469, state.cts and milestone.cts ->
  #3471, health-diagnostic.cts -> sanctioned-permanent.

  Rather than a JSDoc caveat, --baseline now MERGES prior owner values on
  the (file, source) key, so a mechanical regeneration can no longer
  silently discard curated provenance. Verified by regenerating twice.

* fix(#3468): sanitize attacker-controlled fields on every guard output path

Isolated security review, MEDIUM, confidence 8/10.

findSeamBypasses and findPromptSeamUses built findings with an UNSANITIZED
`file`, while the co-located `source` on the same object was correctly
wrapped in sanitizeForReport. On a fork PR a filename is exactly as
attacker-controlled as a source fragment — a repo can legally track a
filename carrying C1 control bytes or bidi overrides.

The raw value reached two paths: --json stdout, and the COMMITTED baseline
JSON via buildBaselineEntries. JSON.stringify neutralizes C0 controls but
does NOT escape C1 (0x7f-0x9f) nor the bidi/zero-width range
sanitizeForReport exists to strip — which is the precise threat the guard's
own header names. Only the human formatter was safe.

Sanitization now happens at CONSTRUCTION, so every consumer inherits it
rather than each output path having to remember. The same defect was present
on `field` and `policy` and is fixed alongside. Double-sanitization in the
formatter is left in place, verified idempotent: escaped output is ASCII and
cannot re-match the control/bidi classes.

Also: the guard was not referenced anywhere in package.json, so nothing ran
it. A drift guard nobody runs is not a guard, and ADR-3408 Decision 5 assumes
it runs. Wired into lint:ci beside its sibling drift guards; it was already
green on this tree, so the chain stays green.

* chore(#3468): re-curate ratchet after an upstream rewording of a tracked bypass

The rebase onto origin/next turned the guard red on its first real day, which
is the ratchet working rather than a defect.

c90ae479f fix(#3350) reworded cmdPhaseComplete's syncStateFrontmatter call
onto one line and changed its third argument. Because entries are keyed on
(file, trimmed source text) rather than a line number, that single upstream
edit registered as BOTH a stale acknowledgment and an unrecorded site — the
two-sided signal the design intends, forcing a human to look rather than
letting a tracked bypass drift out of view.

The owner-preserving merge behaved exactly as designed: three owners survived
because their keys were unchanged, and phase.cts's dropped to null because its
source text is genuinely a different key. Re-curated to #3469, the phase that
owns its removal.

Note for Phase 2: c90ae479f is #3350's fix landing independently on next —
one of the two instances Phase 2 was scoped to drive fail-first. Surfaced to
the epic rather than absorbed silently.

* test(#3468): derive B1's fixture from the table so it cannot go stale

Checkpoint 2 came back with 2 failures of 33803, both B1:

  actual   'current_phase_name'
  expected 'current_plan'

The implementation was right and the test was stale. B1 hand-built a
bodyDeltas literal intending current_plan to be the ONLY unwired row, but it
also omitted status, stopped_at and current_phase_name — all three of which
became preserve-when-unchanged rows in THIS PR. Table order puts
current_phase_name first, so the throw correctly named it.

B1 now builds from neutralBodyDeltas() and deletes exactly one key, which is
what its own comment always claimed it did. A future table change can no
longer silently make it assert the wrong field.

Audited every other bodyDeltas literal in the file: four exist, all correct —
two enumerate all seven rows explicitly, two pass {} where the emptiness is
the point of the test. Roughly thirty other sites already derive from the
helper.

Also renames the local unchchangedChanged to lastActivityDescChangedDeltas.
A typo'd identifier that happens to work is still a Mysterious Name; noted
during research and fixed now that this change touches the file.

* chore(#3468): re-curate ratchet and fold the seam channel into the shared helper

The rebase onto be9329b10 fix(#3374) was a true semantic conflict, resolved
rather than handed back, because the resolution was determinable:

That PR extracted the post-sync preservation pass into a shared
applyPostSyncPreservation helper — which is ADR-3408 section 8.3, i.e. a
piece of Phase 2's own deliverable, landing upstream. Its structure is kept
wholesale; this branch's contribution is applied INSIDE it.

That combination had to be checked rather than assumed. Upstream's helper
wires only FOUR bodyDeltas keys and still passes status / stopped_at /
current_phase_name through six dedicated parameters. This branch reclassifies
current_phase_name to preserve-when-unchanged, deletes those six parameters
from StatePreservationInput, and makes an unwired declared row THROW. Taking
upstream's file as-is would therefore have thrown on EVERY STATE.md write.

The helper now wires all seven rows through the single channel. Verified
7-to-7 against FIELD_CLASSIFICATION, with a clean tsc — which is the real
proof the dedicated parameters are gone, since they no longer exist on the
input type.

The ratchet also caught the same phase.cts call being reworded a second time,
reporting it as both a stale acknowledgment and an unrecorded site. Re-curated
to #3469. Recording the tradeoff plainly: keying on (file, source text) means
an upstream reword of a tracked line needs re-curation, where keying on line
numbers would churn on every unrelated edit. ADR-3180 Decision 4(e) chose
source text deliberately, and the owner-preserving merge added earlier covers
the common case where the text is unchanged.

* chore(#3468): backfill pr number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-08-14 13:25:30 -04:00
Tom Boucher
dbc8b4077a docs(#3467): adr-3408 state.md write-path behavior contract (#3474)
* docs(#3467): adr-3408 state.md write-path behavior contract

* docs(#3467): correct write-seam caller list and adr heading depth

Review findings from the Standards and Spec axes, fixed in place:

- ADR behavior contract demoted from H2 to `### 8` with `#### 8.x`
  subsections, matching ADR-3180's `### 7` / `#### 7.1` precedent the
  front matter claims to follow.
- CONTEXT.md placed the STATE.md factory-reset primitive at
  verify.cts:1925. It moved to health-diagnostic.cts:337 when
  cmdValidateHealth migrated onto the rule table (#3309); verify.cts
  now has no writeStateMd call. The design intent was correct — only
  the address was stale.
- A repo-wide scan found three direct writeStateMd callers, not two:
  cmdStateSync, cmdMilestoneComplete, and the REGENERATE_STATE remedy.
  The last is documented as a sanctioned permanent exception — it is
  a factory reset, so preservation would restore the values it was
  invoked to discard.
- phase.cts comment citation corrected to :2953-2957.
- Amendment 4 attribution corrected: the recorded owner-file exemption
  failure is roadmap-parser.cts; the state.cts transfer is this ADR's
  own extrapolation.

---------

Co-authored-by: sim <sim@local>
2026-08-14 11:03:48 -04:00
Tom Boucher
d30c99bc92 chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference

* test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md

* chore(#1892): reword retired-workflow mentions for removed-but-needed lint

* test(#1892): correct stale surface labels in retargeted suites

* docs(#1892): add verifier-phase-gates row to locale inventories

* chore(#3421): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-13 21:22:03 -04:00
sim
97876a77a4 docs(#3308): ADR-3180 §8.1 Enforced, Amendment 9 — Phase 10 validation
Updates docs/adr/3180-planning-semantic-model-single-owner.md to
reflect Phase 10 shipping: §8.1 status Required -> Enforced (Phase 10,
#3308), guard-roster row contract only -> enforced, phase-index table
issue/status backfilled, and a new Amendment 9 recording the guard's
real baseline (15 distinct raw-read sites, 21 total acknowledged
occurrences in cmdValidateHealth) against the issue's own vaguer
estimate, per Amendment 4a's standing "N found by the guard, never
per the epic" rule. Also records the intended reading of an absent
STATE.md as UNREADABLE-without-diagnostic, symmetric with every other
§7 owner's absence-vs-corruption distinction.
2026-08-12 21:52:50 -04:00
sim
1bc7f7e6b0 test(#3339): fold the state/phase/dispatch & model-profile issue-* cluster — Wave 7
Folds 9 legacy issue-*.test.cjs regression files (140 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
LAST of 4 issue-* waves — closes out the 74-file fix-*/issue-* backlog
(pending BUG_FILE_RE extension, held for a follow-up commit until Wave 6
is confirmed merged, per the epic's own zero-backlog precondition).

- issue-2828-flat-roadmap-total-phases.test.cjs (1) + issue-3204-state-
  writer-phase-count.test.cjs (21): both target state-document.cjs
  buildStateFrontmatter via different CLI entrypoints — merged jointly
  into state-document.test.cjs, 0 dropped.
- issue-2945-phase-complete-checkbox-rollback.test.cjs (4) + issue-2949-
  phase-complete-stage3-sentinel.test.cjs (4): both target phase.cts
  cmdPhaseComplete; issue explicitly warned of overlap — verified
  disjoint fixtures/assertions, 0 dropped, merged into phase.test.cjs.
- issue-2927-reviewer-lane-overlay-invocation.test.cjs (10) merged into
  review-lane-descriptor.test.cjs.
- issue-2939-dispatch-flatten-maxdepth.test.cjs (9) merged into
  host-integration.test.cjs, 2 dropped as verified exact duplicates.
- issue-2977-frontmatter-bom.test.cjs (5) merged into frontmatter.test.cjs.
- issue-2045-third-party-skills-surface.test.cjs (6) merged into
  capability-loader.test.cjs.
- issue-2517-runtime-aware-profiles.test.cjs (80, the largest single
  fold in the epic) merged into model-resolver.test.cjs, 1 dropped as a
  verified true duplicate (checked against src/model-resolver.cts logic,
  not just title similarity).

Fixed a genuine eslint irregular-whitespace finding: a literal BOM
character embedded in a doc comment (pre-existing content from the
original #2977 source, illustrating what a BOM looks like) — replaced
with a readable U+FEFF notation.

3 stale doc references found and fixed (docs/adr/2313, 3180, 443).

Zero net test-coverage loss. No production code changed.
2026-08-12 08:25:14 -04:00
Tom Boucher
aee83c8e9e test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5 (#3378)
* test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5

Folds 6 legacy issue-*.test.cjs regression files (120 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
Second of 4 issue-* waves.

- issue-766-plugin-manifest.test.cjs (50 tests): pure rename (git mv) into
  plugin-manifest.test.cjs, sole comprehensive suite for its module.
- issue-844-manifest-version-sync.test.cjs (14 tests, rename basis) +
  issue-1855-marketplace-manifest.test.cjs (17 tests, merged in): both
  target scripts/sync-manifest-versions.cjs via non-overlapping describe
  blocks (generic engine vs. marketplace.json-specific), now
  manifest-version-sync.test.cjs, 31 tests, 0 dropped.
- issue-498-identity-drift-lint.test.cjs (8 tests, rename basis) +
  issue-498-package-identity.test.cjs (17 tests, merged in): both target
  package-identity.cjs's surface from different angles (lint-side drift
  detection vs. derive/slugify), now package-identity.test.cjs, 25 tests,
  0 dropped.
- issue-607-cache-lineage.test.cjs (14 tests) merged into the existing
  gsd-statusline.test.cjs, matching its already-established fold-wrapper
  convention from prior waves.

Learned from Wave 4: fold agents checked src/** (not just docs/ and
gsd-core/references/) for stale filename references — found and fixed 5
across 3 ADR docs (766, 2121, 457), zero in src/ this time.

Zero net test-coverage loss. No production code changed.

* test(#3337): fix orthogonal-review findings — Wave 5 fold

Standards-axis review found real issues in the just-folded
manifest-version-sync.test.cjs, both fixed here:

- The folded:issue-1855-marketplace-manifest block omitted its own
  test/describe/assert/fs/path/os requires, silently closing over the
  #844 basis section's outer-scope bindings instead of declaring its own
  dependencies — inconsistent with every other fold block in this wave.
  Restored the local requires the original issue-1855 file declared.
- Merging two independently-lettered legacy files (each A-F) left
  duplicate top-level describe() labels (two "A:", two "B:", two "C:").
  Renamed the folded-in #1855 labels (A2/B4/C2) to make every top-level
  label in the file unique.

No test() count changed (31). No production code touched.

---------

Co-authored-by: sim <sim@local>
2026-08-11 23:55:11 -04:00
Tom Boucher
ad07f76a31 test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4 (#3376)
* test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4

Folds 10 legacy issue-*.test.cjs regression files (79 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
First of 4 issue-* waves (following the 3 fix-* waves, all merged).

- 1 file with no prior target coverage: renamed (git mv) into
  legacy-cleanup.test.cjs (sole comprehensive suite for that module).
- 9 files merged into 6 pre-existing suites: golden-parity-single-source,
  runtime-artifact-layout-surface, codex-config (4 sources merged jointly
  in one pass per the issue's own instruction, to catch overlap between the
  4 sources themselves, not just against the pre-existing target — zero
  overlap found, all 20 blocks additive), runtime-config-adapter-registry
  (1 of 10 source blocks dropped as a proven subset of existing coverage),
  cline-install, install.test.cjs.

Incidental fixes required to keep this wave's own ratchets green:
- Fixed a stale ADR doc reference (docs/adr/1235) to a folded-away filename.
- scripts/lint-allow-test-rule-refs: pruned 4 stale allowlist entries for
  renamed/merged-away files, cited 2 previously-uncited allow-test-rule
  comments that surfaced as "new" only because their file path changed,
  added 1 fresh allowlist entry for a pre-existing uncited comment that
  predates this PR, and tightened the exemption-file ceiling 309 -> 305
  to match the real post-fold high-water mark.

Zero net test-coverage loss. No production code changed.

* test(#3336): fix orthogonal-review findings — Wave 4 fold

Standards-axis review + Memtrace graph pass found real issues in the
just-folded suites, all fixed here:

- Standardized the fold-wrapper convention (block-scoped __foldDescribe)
  across golden-parity-single-source.test.cjs, runtime-artifact-layout-
  surface.test.cjs, runtime-config-adapter-registry.test.cjs, and
  cline-install.test.cjs to match the pattern already used by
  codex-config.test.cjs and install.test.cjs in this same wave (and by
  earlier folds elsewhere in the epic) — repeats the exact inconsistency
  Wave 3 (#3335) already fixed once in this epic.
- Fixed a stale allowlist entry's alphabetical position (cosmetic, not
  tool-gated, caught by review anyway).
- Fixed two stale test-filename references in PRODUCTION code comments
  (src/capability-writer.cts, src/runtime-config-adapter-registry.cts)
  caught by lint-removed-but-needed — a class of stale reference this
  wave's fold agents didn't check for, since they were scoped to docs/
  and gsd-core/references/ only, not src/. First fix attempt wrongly
  edited the gitignored gsd-core/bin/lib/*.cjs BUILD OUTPUT instead of
  the tracked .cts source; caught and corrected before commit.
- Fixed one remaining stale doc reference in docs/adr/1235 (a prior
  partial fix in this same wave missed it).

No test() count changed in any file. No production code BEHAVIOR
changed — comment-only fixes in src/.

---------

Co-authored-by: sim <sim@local>
2026-08-11 23:35:44 -04:00
0xdhx
0396d9cab1 enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection

The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.

That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.

Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.

review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.

* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines

Two CI failures from the first push, both mine:

1. lint-tests: the new regression test split readFileSync content on a
   literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
   "\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).

2. golden-install-parity / workflow-size-budget / workflow-compat: editing
   gsd-core/workflows/review.md changes its content hash and byte size, and
   both are pinned in committed baselines. Regenerated via the repo's own
   generators (npm run size:baseline, npm run gen:golden).

The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.

Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).

* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape

The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.

* enhance(#2483): also guard the claude leg against auto-memory injection

CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).

* enhance(#2483): match the claude binary in command position, not argument position

The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as 920a5f3f)
reshaped the effort-args lookup from

  --host claude 2>/dev/null | jq -r '.effort_argv_string // ""'

to

  --host claude --pick effort_argv_string

which put a flag immediately after `claude` and made the config query
read as a third claude dispatch, failing the count assertion.

The defect class is a binary name in *argument* position being read as a
command. Fixed at the class rather than the instance: tokenise the line
and skip any `claude` whose preceding token is a flag. That also covers
the latent sibling one line away in review.md (`command -v claude`),
which escaped today only because its next token is a redirect.

Negative-controlled four ways: stripping CLAUDE_CODE_DISABLE_AUTO_MEMORY=1
fails, stripping the whole env guard fails, adding a genuine third
unguarded dispatch (`timeout 900 claude --output-format text -p -`) still
fails — so the narrowing did not blind the matcher to reshapes, which is
the property the count assertion exists for — and the pre-#2589 jq form of
the lookup still passes, so the matcher is not pinned to today's base.

* enhance(#2483): carry the claude reviewer's memory guard as declared lane data

ADR-2782 Phase 5b replaced the hand-authored per-CLI dispatch legs in
review.md with the declared lane table, so the two `env`-prefixed shell
lines this PR previously added no longer have a surface to live on. The
guard is reimplemented where the lane contract now lives.

`SpawnInvoke` gains an optional `env`, the claude lane declares the pair,
the resolver folds own string-valued entries into `SpawnPlan.env` (absent
or empty resolves to null, so the runner has one shape to test), and the
runner passes it to spawn. Production merges it OVER `process.env` into a
fresh object for that one child, so nothing reaches the orchestrating
session or any other lane in the run.

Declared data rather than a handler (D6): the pairs are static per lane,
which is precisely what the manifest vocabulary is for. The capability
manifest carries the same field, because the lane-fidelity test compares
manifest and descriptor over the union of `invoke`'s keys.

The regression test is rewritten against the resolver and runner rather
than review.md's text. It gains the property the source-text assertions
could only approximate: that `process.env` is never mutated.

Scope boundary, asserted rather than left in prose: `env` is not part of
the trust-disclosure surface, which is safe only while no manifest body
reaches the resolver — the registry's reviewer bodies contribute slugs to
the parity check and execution resolves from `REVIEWER_LANES`. The new
test fails first if that ever changes.

* enhance(#2483): restate the guard's mechanism in the docs and changeset

Both described the fix as two `env`-prefixed dispatch lines, which is the
surface ADR-2782 Phase 5b removed. The user-visible behaviour is
unchanged; the carrier is not, and a changeset that ships a description
of a mechanism the tree does not have is a CHANGELOG entry nobody can
verify against the code.

* enhance(#2483): cover the production spawn wiring end to end

The unit tests stop at the runner's `deps.spawn` seam — every one injects
a spy. Production supplies that seam in `gsd-core/bin/gsd-tools.cjs` as a
hand-written object no test constructs, so the chain could be correct all
the way to `SpawnPlan.env` and the merge could still be wrong or absent
with the suite green. Deleting those four lines was the one mutation that
left every other control silent.

This runs the real `spawnSync` through `gsd-tools review-lane invoke`,
with a `claude` shim on PATH that records the environment it was handed.
It asserts both halves in one test: the pair arrives, and an unrelated
inherited variable survives — a wiring that REPLACED the environment
rather than merging over it would satisfy the first and break every
lane's PATH and HOME.

POSIX-only; mediating a Windows `.cmd` shim is a separate concern the
repo already tests on its own.

Noted rather than fixed: `timeout`, `killSignal`, `maxBuffer` and
`shell: false` on that same object are equally uncovered. That is the
epic's gap, not this change's, and closing it is not in scope here.

* enhance(#2483): validate the invoke.env shape and register it as spawn-only

`env` was the one spawn-invoke field with no shape enforcement: every sibling in
`validateSpawnInvoke` is checked, and a manifest declaring `env` as an array, a
string, a number, or an object with non-string values passed validation in
silence. That matters more than an ordinary schema gap here, because
`resolveLanePlan` DROPS a non-string value rather than coercing it — so an
unvalidated manifest declares a pair that never reaches the spawn, which is the
failure a memory guard can least afford.

Two registrations, not one. `env` was also absent from
`SPAWN_ONLY_INVOKE_FIELDS`, which is the list the openai-http arm rejects
against — so `invoke.env` was accepted on a transport that issues an HTTP POST
and has no child environment at all. It was the only spawn-shaped field accepted
there; the other six each produce two errors. Self-found while sweeping the
class, not raised in review.

Keys are held to the portable POSIX environment-name grammar. That is a policy,
not a claim about what an environment can hold: measured, only NUL is actually
rejected by `spawnSync`, while `=`, a leading digit, a dash and a space are all
carried through to the child (an `A=B` key arrives as the raw entry `A=B=value`).
They are refused because a name outside the grammar is not portably addressable
by the program meant to read it.

`__proto__` is refused for a different and concrete reason. It passes that
grammar and is a real own key once a manifest is JSON-parsed, but assigning it
onto a plain accumulator goes through the inherited `__proto__` setter rather
than creating an own property — and for the string values this field permits the
setter is a no-op that does not even change the prototype. The pair would
validate and then simply vanish before the spawn. (An environment CAN carry a
literal `__proto__` entry; this is about the resolver's accumulator, and the
error message says so.)

Deliberately narrower than the sibling reserved-name guards in this file, which
also reject `constructor`/`prototype`: those guard bracket lookups that resolve
prototype members, whereas this reads via `Object.keys` plus an own-value read,
where `constructor` assigns as an ordinary key the spawn could carry.

`effortChannel` is deliberately left in neither field list: ADR-2782 D2 defines
it for both transports, so it is shared rather than spawn-only.

Reversion-controlled, three mutations, all three fire a named test: dropping
`env` from the discriminator fails `httpTransportRejectsEnv`; removing the
`__proto__` arm fails `envRejectsProtoKeyThatWouldSilentlyVanish`; disabling
the block fails four.

(#2483)

* enhance(#2483): amend ADR-2782 D2 for the invoke.env vocabulary widening

D2 records the spawn `invoke` shape as a closed vocabulary, and its Amendments
section carries a dated entry for every prior widening (Phase 1 #2794, Phase 2
corrections #2795, Phase 5b #2799). This change extended that vocabulary in code
without touching the ADR governing it, so the ADR contradicted the
implementation — and the repo's own convention, recorded in CONTEXT.md, is that
the ADR is amended in the same PR precisely because the prior widenings did it
correctly.

Adds the `invoke.env` row to the D2 table and a dated Amendments entry.

The entry also corrects the authority this change cited. The source comment
pointed at D6, which governs the closed `handler` enum — imperative behavior
admitted first-party — and says nothing about the `invoke` field vocabulary.
That is D2's territory, so the citation never covered the gap.

Two claims are corrected rather than restated, both about the trust boundary
that justifies leaving `env` out of the D5 disclosure signature:

- The regression test does not enforce that boundary. On one forged lane it
  shows the resolver folds whatever it is handed, so a future path feeding it
  manifest lanes would not make any assertion in that test fail. Its comment
  claimed it "will fail first"; that was wrong, and both the comment and the
  ADR now say the boundary is a property of the production call chain instead.
- The ADR is internally inconsistent on whether third-party manifest lanes
  execute at all: Consequences says adding a reviewer needs "no core patch",
  while `gsd-tools.cjs` rejects every slug absent from the first-party
  REVIEWER_LANES map. CONTEXT.md, `workflows/review.md` and the resolver's own
  header take the first view. #2483 did not create that inconsistency and does
  not resolve it; the entry records it rather than settling it in its own favour.

(#2483)

* enhance(#2483): document invoke.env in the capability-manifest reference

ADR-2782 points capability and plugin authors at
`docs/reference/capability-manifest.md` as where the lane vocabulary must be
visible, and its `invoke` row enumerates the spawn sub-shape field by field.
`env` was absent from that table while being part of the real shape, so the one
document a third-party capability author would actually consult to learn the
field exists did not mention it.

Squarely Diataxis reference material — a field-by-field schema description — so
it goes here rather than in the user-facing prose, which was already updated.
States the constraints a manifest author can actually trip, and is explicit that
the name grammar is a portability policy rather than an OS limit, so a reader
does not take it for a claim about what an environment can hold.

(#2483)

* enhance(#2483): disclose and sign the reviewer lane's env and residual invoke fields

`invoke.env` was undisclosed at install time. That was defensible while manifest
lanes could not execute — the premise this PR's own ADR amendment recorded — and
#2927/#3062 retired it: `routeReviewLane` now merges installed overlay `reviewer`
bodies into its lane map via `mergeReviewerLanes`, which is a field-identical merge
by ADR-2782 D1 and deliberately does not deep-validate. An overlay's whole `invoke`
therefore reaches `resolveLanePlan`, and `env` reaches the spawned child. A consented
third-party capability could set `NODE_OPTIONS=--require ./evil.js` on a reviewer lane
with no install-time disclosure and no re-consent.

The same file already decided what `env` means in a manifest: MCP servers fold it into
the disclosure signature and render each key and value in the consent prompt, with an
inline rationale naming this exact shape. Reviewer lanes get the identical treatment.

`env` was the ninth unsigned invoke field, not the first. `defaultHost` (the manifest's
OWN fallback egress host, used whenever the config key resolves to nothing),
`path`, `outputChannel`/`outputArg`, `modelArg`, `effortChannel` and `modelDiscovery`
all reach `resolveLanePlan` and none was bound. Enumerating a ninth name leaves the
tenth open, so the lane signature carries a RESIDUAL of every other declared `invoke`
key — the completeness backstop `rawConfig` already gives the MCP line (#1459 finding 5),
and the "sign the whole object" remedy the recorded decision on this class prefers.

`defaultHost` is also rendered: `resolvedHost` comes from user config, so a lane whose
key is unset displayed "(unresolved …)" — which reads as "no destination" — while the
runtime egresses the plan and review text to the address the manifest picked.

D4.5 is preserved one level down: the extra element is appended ONLY when the lane
declares something beyond the eight already-bound fields, so an env-free lane's
signature stays byte-identical and no already-consented capability is re-prompted for
a field it does not use. A lane that does declare one re-consents, which is the point.

Execution-primitive env names are FLAGGED in the prompt, not refused in the validator.
A denylist cannot be the boundary here: `PATH` alone is a complete execution primitive
for a spawn lane and can never be refused, the child is an arbitrary third-party binary
so the true set spans every interpreter's injection vars, and the MCP `env` this mirrors
refuses nothing and discloses everything. Missing a name costs a quieter line, never a
boundary.

Refs #2483.

* enhance(#2483): exercise the real overlay merge path in the guard test

The test named for the manifest/first-party boundary did not test it. It built a
forged lane locally, handed it straight to `resolveLanePlan`, and asserted that
`REVIEWER_LANES` did not contain it — so no assertion in it depended on the claim its
name made, and a code path that fed manifest lanes to the resolver would not have made
it fail. Its own comment said as much, and named the production chain as the real
carrier of the guarantee: "gsd-tools.cjs builds its lane map solely from REVIEWER_LANES".

That sentence is now false. #3062 merged overlay reviewer bodies into that map, so the
test's premise and its subject both moved.

The replacement routes through `mergeReviewerLanes` — the real helper the production
path calls — and asserts the overlay lane is admitted, resolves, and carries its `env`
into `SpawnPlan.env`. That makes the security property falsifiable instead of narrated.
It then asserts what now backs it: the env is disclosed on the surface, rendered key
and value in the consent prompt, flagged when the name is an execution primitive, and
bound to the signature so a value change, an addition, or a removal each force
re-consent.

Three further cases, because the finding's generative half is what stops it recurring:
the residual backstop is asserted against five fields including one that does not exist
(`aFieldThatDoesNotExistYet`), so a future vocabulary widening cannot silently re-open
this; a fully-enumerated lane is pinned to its original 8-tuple, which is what keeps the
fix from re-prompting every consented capability; and an http lane's manifest-declared
`defaultHost` is asserted to reach both the prompt and the signature.

Reversion-controlled, seven mutations, all seven fail a named test: env dropped from the
surface, the prompt's env line removed, the execution-primitive warning removed, the
signature's extra element never appended, the residual emptied, the defaultHost line
removed, and the declares-something test un-widened. The last of those was SILENT on its
first run and its test was written in response, then the control re-run.

Refs #2483.

* enhance(#2483): correct the ADR amendment's manifest-lane premise

The amendment argued `env` needed no D5 disclosure because a manifest's `invoke`
fields never reach `resolveLanePlan`. That was true when written and #3062 retired it
22 hours after this branch's last commit: `routeReviewLane` now builds its lane map
from `mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true}))`, and
D1's no-translation-layer rule makes that a field-identical merge, so an overlay's
whole `invoke` reaches the resolver and executes.

The entry had named this exact trigger — "were manifest lanes ever made executable,
`env` must join the disclosed surface in that change, and nothing here will trip if it
does not." Nothing tripped. The premise is rewritten to current truth rather than
annotated, because an ADR is read in fragments and a superseded paragraph left standing
reads as live reasoning to the next author; a one-line dated tombstone points at git for
the withdrawn text.

The rewritten entry records four things the first draft could not: that the enumeration
itself was the defect (`env` was the ninth unbound `invoke` field, and `defaultHost` and
`path` are egress-relevant on their own), that the residual is what closes the class,
that D4.5's byte-identical-signature property is preserved by appending the residual only
when a lane declares something beyond the eight bound fields, and that consent — not
shape validation — is the boundary, since no honest env denylist can exclude `PATH`.

It also closes the internal inconsistency the previous entry could only record. This ADR,
`CONTEXT.md`, `gsd-core/workflows/review.md` and `resolveLanePlan`'s own header all said
overlay lanes reach the resolver while the runtime said otherwise; #3062 resolved that in
the documents' favour, which is what makes the disclosure mandatory rather than defensive.

Refs #2483.

* enhance(#2483): record in the manifest reference that invoke fields are consent-bound

`docs/reference/capability-manifest.md` is the field table ADR-2782 points capability
authors at, and it described `invoke` purely as a schema. A third-party author reading it
could not learn that everything they declare there is shown to the user at install and
bound to the consent signature — which is exactly what they need to know now that an
overlay reviewer lane executes (#2927/#3062).

States the two things the schema alone cannot: that `env` and `defaultHost` are named in
the consent prompt and the rest is covered by a residual, so any change to a declared
`invoke` field forces re-consent; and that `env`'s validation is a portability policy
rather than a safety boundary, since `PATH` is a complete execution primitive and cannot
be refused. Names that are execution primitives are highlighted in the prompt instead.

Refs #2483.

* enhance(#2483): add a Security changeset for the reviewer-lane disclosure

The existing fragment describes the enhancement this PR was opened for and stays as it
is. The disclosure fix is a separate user-visible change of a different type: a
capability declaring `invoke.env` or `invoke.defaultHost` will ask for consent once
more, and users are entitled to read why in the changelog rather than discover it as an
unexplained prompt.

Type is `Security` rather than `Changed` because the entry describes a closed
code-execution disclosure gap, not a behaviour adjustment.

Refs #2483.

* enhance(#2483): correct this round's own claim about who gets re-prompted

Self-found while auditing the round's claims before publishing them. The changeset and
the ADR entry both stated that a capability declaring `invoke.env` or `defaultHost`
"will ask for consent once more". That is wrong, and it overstated the cost of the fix
in the one direction a maintainer would have had to take on trust.

A code change to `disclosureSignature` re-prompts nobody. `hasProjectConsent` matches on
the recomputed bundle `contentHash` — the signature has not been the security binding
since #1459 CB-1/CB-2 — and the upgrade path's `executableSetChanged(old, new)` compares
two disclosures both computed by the CURRENT code, so widening the signature moves both
sides of that comparison equally. First-party capabilities never reach the path at all:
the install flow blocks a first-party id before trust evaluation.

What the widening actually buys is forward-looking, and is the real argument for it: an
upgrade whose manifest edits a declared `invoke` field now registers as an
executable-surface change and re-consents, where before it could change what the lane
runs in silence.

Also measured and recorded, because the D4.5 property was stated more strongly than it
deserved: of the twelve first-party reviewer capabilities, ZERO are in the
byte-identical-signature class — every real lane declares at least `effortChannel`. The
property is a guarantee about minimal lanes, not a description of the fleet, and the ADR
now says so.

Refs #2483.

* enhance(#2483): sign and disclose the probe binary and the lane's outer fields

Found by this round's own adversarial review, and it is the same defect one level out:
the `invoke` residual cannot reach the lane body's OUTER fields, and `probeLane` SPAWNS
`probe.binary` with `--help` before dispatch (`review-lane-runner.cts`, the
`command-exists`/`command-capability` arms). An overlay naming an arbitrary probe binary
therefore executes it — unsigned and undisclosed, exactly as `invoke.env` was, and
reachable on the same #3062 path.

The lane element now carries a second residual over the outer fields, and the probe
binary is shown in the consent prompt when it differs from the dispatch binary — it is a
program that runs, and the user is entitled to see it.

TWO fields stay excluded, and that is a decision rather than an omission:
`reviewsSection` and `timeoutFloorMs` are ADR-2782's cosmetic carve-outs (matrix
A10/A13), where re-consenting would present a prompt carrying no security information.
A test pins that they remain excluded, so a later widening cannot quietly reverse D4.5
while claiming to complete this fix.

Also corrects a miscount introduced by the previous commit: the source comment said the
enumeration had fallen behind by "seven fields" and omitted `fallbackModel`, while
asserting `env` was the ninth. `resolveLanePlan` reads twelve `inv.*` fields and four
were bound, so the number is eight. The comment now states the derivation rather than
just the total.

Reversion-controlled: emptying the outer residual fails "repointing the probe binary must
force re-consent"; removing the render line fails its own named assertion.

Refs #2483.

* enhance(#2483): refuse execution-primitive env names as defence in depth

Adopts the review's B5 after this round's own adversarial pass refuted my reason for
declining it. I had argued a denylist was worthless because `PATH` can never be refused.
That was wrong on the facts: no shipped reviewer manifest declares `PATH`, so it can be
refused, and it is the most complete primitive in the set — repoint it at a directory
holding a fake binary and the declared `invoke.binary` is irrelevant. A list that cannot
be exhaustive can still close the highest-confidence, lowest-legitimacy routes.

So the validator now rejects `PATH`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`,
`BASH_ENV`, `PYTHONPATH`, `PERL5OPT`, `RUBYOPT`, `GIT_SSH_COMMAND`, `JAVA_TOOL_OPTIONS`
and their siblings on a reviewer lane. A lane needing a specific executable declares an
absolute `invoke.binary` instead of reshaping the child's environment.

The comment states plainly that this is defence in depth and NOT the boundary — the
boundary is install-time consent, which discloses every declared pair and binds it to the
signature, so an unlisted name is still SEEN before it runs. That framing is load-bearing:
a future reader who mistakes the denylist for the control will under-invest in the one
that is, which is the failure mode I was trying to avoid by declining it outright.

Two tests: the rejection itself across ten names, and a guard asserting no shipped
reviewer capability declares a denied key — so if the list ever outgrows its evidence,
that surfaces as a decision rather than a silent removal.

Refs #2483.

* enhance(#2483): fix two stale D5 enumerations elsewhere in the ADR

The previous commit rewrote the amendment's premise but swept only the amendment. Two
normative passages earlier in the same ADR still enumerated the old closed field list and
now contradicted it: the `executableSetChanged` trigger list, and the split-binding note
asserting the seven manifest-derived fields were "everything that is SHA-pinned".

That is the failure the rewrite-don't-annotate rule exists to prevent, one section over —
an ADR is read in fragments, and a fragment carries no supersession marker, so a reader
landing on either passage would have taken the superseded enumeration as current.

Both now name the residual as the mechanism rather than restating a list, which is also
what stops them going stale the next time the vocabulary widens.

Found by this round's adversarial review, which grepped the whole document rather than
the section under edit.

Refs #2483.

* enhance(#2483): stop the probe disclosure claiming a spawn that does not happen

The probe line added one commit ago rendered "probes by running: <binary> --help" for
every lane. That is false for `kind: "command-exists"`, which only calls `hasBinary` — a
PATH/filesystem scan that starts no process. Only `command-capability` spawns.

A false statement in a consent prompt is worse than a missing one: the prompt is the
surface a user is asked to trust, and this one overstated what a lane does. Worse, the
test I wrote to prove the fix used `command-exists` — the kind that does NOT spawn — so
it pinned the wrong claim and would have kept the error green forever.

The surface now carries `probeKind` and the two kinds render differently: a spawn is
described as a spawn, a presence check as a presence check. The test exercises both, and
asserts the `command-exists` path never emits the spawn wording.

Also corrects the field-count parenthetical to state its derivation unambiguously —
`resolveLanePlan` reads thirteen `inv.*` fields including `env` (twelve before this PR),
four were bound, so eight were unbound before `env` and nine including it. The bare
"twelve" was true only of the pre-PR tree and read as a claim about the current one.

And retires two comments that argued AGAINST the validator denylist this round then
shipped. Leaving them would have handed the next reader the reasoning for removing it.

Reversion-controlled: conflating the two probe kinds fails a named test.

Refs #2483.

* enhance(#2483): match the reviewer-lane env denylist case-insensitively

The denylist added one commit ago compared exact case, so `Path`, `path`, `node_options`
and `Node_Options` all passed it. Windows environment lookup is case-insensitive, so
those reach the child as `PATH` and `NODE_OPTIONS` — the exact inputs the list names.

An exactly-cased denylist is worse than none: it reads as a control while admitting the
input it was written to refuse, and the next reader has no reason to doubt it. Members
are stored uppercase and the key is folded before lookup; the name grammar already
constrains keys to ASCII, so a plain fold is sufficient.

Reversion-controlled: restoring the exact-case compare fails `envDenylistIsCaseInsensitive`
on `Path`.

Refs #2483.

* enhance(#2483): correct the docs that still described the denylist as absent

Both the ADR and the manifest reference still said `env` carries no denylist and that
`PATH` "can never be refused" — written when that was this round's position, and left
standing after the round reversed it. A reader landing on either passage would have taken
the superseded argument as current, which is precisely the failure the rewrite-don't-
annotate rule exists to prevent.

Both now describe the denylist, name `PATH`'s inclusion and the case-insensitive match,
and keep the limit explicit: the list cannot be complete against an arbitrary child and
disclosure runs before validation, so consent remains the boundary.

The ADR's byte-identical-signature claim is also corrected rather than softened. With the
outer residual in place, a lane producing no residual is one the validator rejects — it
declares no `flags`, `probe`, `emptyOutput`, `evidenceClass`, `requiresBinaries` or
`promptBudgetKey`. So the property is about the ENCODING, not a claim that any real
signature is unchanged, and it is not the argument for the change being safe. That
argument is that consent binds to the bundle contentHash and no existing consent is
invalidated at all.

Refs #2483.

* test(#2483): cover the three new lane disclosure fields in the injection-safety parity guard

The PARITY test in section N exists to catch a renderer field that skips
`renderValueForPrompt` (#3248). Its payload manifest is hand-maintained, so it
covers the fields that existed when it was written — slug, binary, args,
hostConfigKey, handler — and none of the fields this PR adds.

This PR renders three further manifest-supplied values into consent-prompt
lines: `invoke.env` (keys and values), `invoke.defaultHost` and `probe.binary`.
The gap was silent rather than theoretical: with the lane env line reverted to
the pre-#3248 raw form, the whole 948-test lane/capability/trust-disclosure
suite stayed green.

Two manifests, because the shapes render disjoint lines — `defaultHost` only on
the openai-http branch, `env`/`probe` only where declared, and the probe line
only when the probe binary differs from the dispatch binary.

Non-vacuity is asserted on the typed disclosure object and on structural line
counts, not by substring-matching rendered prose: CONTRIBUTING.md forbids raw
text matching on test output, and this section's own header promises structural
assertions only, so a prose match here would have made that promise false.

Negative-controlled three ways against the merged tree, each producing exactly
one named failure: env rendered raw, defaultHost rendered raw, probe binary
rendered raw.

* docs(#2483): extend the #3248 render-site comment to the fields this PR adds

The comment enumerates every manifest-supplied value that must pass through
`renderValueForPrompt`, and it stopped at `handler` — the reviewer-lane fields
that existed when #3248 landed. This PR renders three more (`defaultHost`, the
probe binary, and the env keys and values), so the list understated its own
contract in the one place a future author would check before adding a fourth.

A comment enumerating a closed set is a set that can silently fall behind the
code it describes; the parity test added alongside is what makes the omission
fail loudly rather than read as deliberate.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:42:41 -04:00
Tom Boucher
9341d8b8d3 test(#3334): fold the workflow-dispatch & review-lane fix-* cluster — Wave 2 (#3342)
Folds 15 tests/fix-*.test.cjs regression files (191 test() blocks) into
their module's main suite, per the wave decomposition of #3315 (H3 of
epic #3053). 187 blocks land in 8 existing suites (4 exact-duplicate
cases dropped, documented inline); 4 blocks move via git mv into 2 new
suite files with no prior coverage to merge into. Zero production
behavior change.

Also tightens two H1 (#3313) ratchets that the fold's own file-count
reduction moved past their grace window, per the ratchets' documented
dual failure mode (a stale/too-loose baseline fails exactly like a
novel violation):
- lint-test-file-count.allowlist.json: removes the stale "audit" entry
  (folding fix-2766 into tests/uat.test.cjs drops that module back to
  its 2-file cap).
- lint-allow-test-rule-refs.ceiling.json: lowers maxFiles 314 -> 309,
  the real post-fold high-water mark (gsd-test's own repo-baseline
  test caught this — CI, not a human, found it).

Two orthogonal review passes (Standards+Spec code-review, isolated
security-review) found and this commit fixes two issues before push:
a genuinely-distinct #2287 test case (file-absent vs. file-present-
resolved) that a prior fold pass had wrongly dropped as a duplicate —
restored verbatim into tests/uat.test.cjs; and a missing same-line
allow-test-rule citation on the #2196 block in
tests/debug-session-management.test.cjs, added for consistency with
its sibling #2257 block.

lint-removed-but-needed also caught two stale doc references to the
now-folded-away fix-2285-claude-orchestration-wiring.test.cjs filename
(docs/adr/1143-claude-orchestration-capability.md,
gsd-core/references/execute-phase-response-language.md) — updated both
to point at tests/claude-orchestration.test.cjs, its new home.

Co-authored-by: sim <sim@local>
2026-08-10 20:51:40 -04:00
Tom Boucher
80734a9694 chore(#3053): clock-seam ADR-456 amendment + backfill — H2 (#3332)
* test(#3314): backfill deterministic clock-seam coverage (failing-first)

Replaces loose regex/range assertions with exact-value and boundary
tests for the CLI-subprocess and in-process clock-touching call sites
identified by H2's audit (epic #3053): cmdCurrentTimestamp,
_wsParseRetryAfter (commands.cts), cmdInitManager's is_active gate and
cmdInitQuick's quick_id generation (init.cts), and reapStaleTempFiles
(io.cts). The CLI-subprocess-pinned tests are expected RED until a
follow-up commit routes those call sites through realClock so
GSD_TEST_MODE+GSD_NOW_MS can reach them.

* refactor(#3314): route CLI-subprocess clock reads through realClock

cmdCurrentTimestamp, cmdInitManager's is_active gate, and cmdInitQuick's
quick_id generation read Date directly, which the GSD_TEST_MODE+GSD_NOW_MS
subprocess pin cannot reach (it only fires inside realClock.now()).
Behavior-preserving: realClock.now() falls through to Date.now() whenever
GSD_TEST_MODE is unset, which is every real invocation.

* docs(#3314): amend ADR-456 with reachability-based clock-control rule

ADR-456 §(a) documented one mechanism (injected {clock=Date} +
t.mock.timers). Adds the two this repo already relies on: t.mock.timers
for in-process direct-Date reads, and the GSD_TEST_MODE+GSD_NOW_MS
subprocess pin (routed through realClock) for CLI-spawned code. Updates
TESTING-STANDARDS.md's matching passages, which already referenced this
issue by number as the no-elapsed-assertion promotion precondition.

* fix(#3314): address orthogonal review findings

Spec-axis findings: file and link the no-elapsed-assertion promotion
follow-up (#3331) instead of leaving TESTING-STANDARDS.md pointing at
a dead #1885, and ship the module-by-module audit table in the ADR
itself rather than only in a gitignored phase artifact.

Standards-axis finding: pin the "hour-old file = not active" test via
GSD_TEST_MODE+GSD_NOW_MS for consistency with the reachability rule
this PR's own ADR amendment now documents.

---------

Co-authored-by: sim <sim@local>
2026-08-10 15:24:39 -04:00
Tom Boucher
aceea3ce4a refactor(#3217): withhold a percentage when its scope is not complete (#3318)
* wip(#3217): rule-4 scope withholding — parked, two open findings

Implemented but NOT shippable. An isolated review found buildStateFrontmatter
still hardcodes SCOPE.COMPLETE, so state json reports percent 0 where roadmap
analyze, stats and query progress all correctly report null on the same disk
state - rule 4 reintroduced at a site this phase claims to close. Also: roadmap
analyze emits scope complete beside progress_percent null with nothing
explaining it.

Parked to build Phase 4 (#3186) first, which is unblocked. Findings recorded in
.gsd/phase/refactor-3217-completion-ratio-scoping/60-review.json.

* fix(#3217): withhold the sync percentage on a non-complete scope

The parked blocker is fixed - buildStateFrontmatter no longer hardcodes
SCOPE.COMPLETE, and the prose Progress fallback is gated too, which was a second
leak found while tracing the first. roadmap analyze exposes progress_scope so a
consumer can tell WHY a percentage is absent from the JSON alone.

Then a residual gap was reproduced rather than assumed. cmdStateSync carried the
same hardcode behind a written reason claiming it did not reproduce. It did: on a
TRUNCATED window and on UNSCOPED row 4, state sync wrote Progress 0 percent to 100
percent while state json, roadmap analyze, stats and query progress all withheld -
and it persisted a self-contradictory file, body claiming 100 percent while its own
frontmatter correctly omitted percent.

The excuse was also wrong. syncRoadmapRaw is already parsed in that function and is
exactly what produces a real scope, so there was a scope to pass. Threaded through
listMilestonePhaseDirs; a non-complete scope now skips the write with a reason in
changes. milestoneBounded stays as the orthogonal 1761 guard for row 5.

Second time this epic a does-not-reproduce claim was too generous. Recorded in
ADR Amendment 8 as a correction rather than a quiet rewrite.

Verified on the remote runner.

* test(#3217): give the withholding fixtures a resolvable scope

40 matrix failures, all fixture drift - no code regression. My own hypothesis
that this was over-withholding was wrong and is recorded as such: the worry case,
a plain ROADMAP with Phase entries and no version heading, resolves to complete
exactly as ADR 7.1 says it should.

The real causes were two fixture shapes. Most had no ROADMAP.md at all, which is
unreadable via a pre-existing graceful path, and asserted a numeric percent. The
five vscode, pi-extension, mcp-server and shell-projection failures were that
shape - bare temp dirs using progress json as a reachability proxy while
asserting typeof percent is number, which under rule 4 is now null.

The rest had a version token in a title or heading with no STATE.md milestone
pointer to resolve it, which is classification row 4, versioned but unresolved,
so withholding is correct per the contract.

Verified on the remote runner.

* test(#3217): make the LM-tools reachability tests dispatch against their fixture

The gsd_progress reachability test was never testing its fixture. invoke()
resolves cwd from vscode.workspace.workspaceFolders by design (the real
LanguageModelToolInvocationOptions has no cwd field, per the 2103 fix in
extension.js), the mock had no workspace at all, and the test passed a cwd option
nothing reads - so it dispatched against the repo working directory. Writing a
ROADMAP into the temp dir had no effect. Rule 4 only made it visible.

Fixed by mocking workspaceFolders. The two siblings in the same file carried the
identical dead cwd and were dispatching against the repo too; they were not
failing only because their assertions did not touch scope-dependent output. Both
now use their own fixture with assertions unchanged - the no-planning fallback
paths already satisfy them honestly.

Re-scanned the other five reachability files: no further instances. They thread
cwd into parameters that genuinely read it, not through an options shape that
ignores it.

Verified on the remote runner.

* chore(#3217): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3318 exists.

* ci(#3217): give the coverage merge enough heap for the merged shards

The coverage gate OOMed at exit 134. c8 report merges three shard artifacts,
roughly 358MB of V8 dumps in coverage/tmp, and died holding their per-file
position maps at the ~4GB default heap. Verified as this branch's delta rather
than pre-existing: the same job succeeded on next at 14:18, after phases 4 and 5
merged.

Both coverage-gate steps get the bump because both re-slice the same merged data.
8192 doubles what failed and leaves headroom on a 16GB ubuntu runner, matching
the idiom the shard step already uses at 6144.

This is a memory bound, not a change to what is measured. No threshold was
touched. The test file was checked for gratuitous subprocess spawning and is
already reasonable at 43 spawns, each a distinct fixture-by-surface pairing.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-10 11:59:51 -04:00
Tom Boucher
e201cde73c refactor(#3186): one shared phase-completion predicate, disk-strict (#3306)
* docs(#3186): record the disk-strict completion decision in ADR-3180 7.4

The maintainer decided #2957 on 2026-08-08: disk state is authoritative and a
ROADMAP checkbox is a human annotation with no machine authority. Section 7.4
still carried the OPEN QUESTION and was marked blocked, so the contract said one
thing and the tracker another.

Recorded per section 7's own rule - a behavior not stated there is not decided,
and amending a rule is an ADR amendment rather than a code change with a comment.
The decision comment names Phase 4's PR as the carrier of this edit and makes it
an acceptance criterion that the text be in the tree before implementation
begins, so this lands first, alone, ahead of any code.

Also clears the stale blocked-on-2957 row in the guard roster.

* refactor(#3186): one shared phase-completion predicate, disk-strict

isPhaseComplete in verification.cts becomes the single owner. It calls
readVerificationStatus UNCONDITIONALLY - plan count is not a precondition - so a
zero-plan phase with a passing VERIFICATION.md is complete. That is #3168: init
gated the read on a plan count and synthesized a not_required sentinel, so
phase.complete succeeded while init.manager reported incomplete for the same
phase.

The guard, built and run before scope was fixed per Amendment 3, found 9
re-derivations where the ADR named 3. Four were unnamed, including one in the
prompt layer: mvp-phase.md ORed a ticked checkbox with disk status, which under
disk-strict is the divergence itself.

Per the #2957 decision, a ticked ROADMAP checkbox is a human annotation with no
machine authority. The overrides in roadmap analyze and init manager are deleted
rather than generalized; the user's checkbox stays in ROADMAP.md, only its
authority goes.

scanPhasePlans.completed and buildWorkstreamInventory are deliberately NOT folded
- they answer 'are all plans summarized', which is a different question, and
folding them would either over-report completion or invert the dependency
direction between Phase 1's owner and this one.

Verified on the remote runner.

* fix(#3186): close seven review findings and record the missing-verdict rule

The isolated review reproduced a write-path regression I introduced: migrating
cmdRoadmapUpdatePlanProgress dropped its summaryCount>=planCount gate, so a phase
with a fresh passing verification plus a newly-added unsummarized plan reported
complete AND wrote a checkbox into ROADMAP.md while phase complete refused. The
owner stays right per 7.4 - plan count is not a completion precondition - so the
gate is restored at the write site as an explicit composition, mirroring the
separate 2648 unexecuted-plan gate cmdPhaseComplete already carries.

The spec axis was right that my 0.x-split reasoning was too permissive. The 2957
decision names buildStateFrontmatter as one of the three that must converge, and
buildWorkstreamInventory combined a summaries-met local with verification data to
decide the same verdict - Decision 4(c)'s named bypass, and it reproduced 3168 in
a third surface. Both now route through the owner. The raw scanPhasePlans helper
stays: it answers are-plans-summarized, which genuinely is a different question.

Maintainer decision recorded in 7.4: a missing verdict is not a passing one, so
an absent VERIFICATION.md means not complete everywhere. That retires 2645's
verifier-disabled tolerance and inverts its Goodhart incentive - deleting the
evidence now lowers completion instead of raising it.

Guard hardened: block-form count gates and algebraic restatements are caught, and
the header now discloses its remaining limits instead of overclaiming.

Verified on the remote runner.

* fix(#3186): route state sync through the owner and catch bare completed reads

The matrix found 52 failures. 51 were fixtures asserting the old semantics: a
phase with plans and summaries but no VERIFICATION.md used to count complete and
correctly no longer does. Each fixture now carries a passing verification where
that is what the test was actually about, rather than having its assertion
weakened.

The 52nd was a real 10th re-derivation the guard could not see. cmdStateSync
destructured scanPhasePlans().completed directly - a bare field read, not a
comparison - and used it as a completion verdict, so state sync and state json
disagreed on completed_phases for identical disk state. Routed through the owner.

Guard gains shape (d): any read of .completed off a scanPhasePlans() result
outside plan-scan.cts, in chained, destructured and indirect forms, function
scoped with no line window. It cannot tell a summaries-met read from a completion
read - that is data flow - so it flags every one and requires a written-reason
exemption, which is the same discipline shapes a-c already use. The blind spot is
disclosed in the header rather than overclaimed.

The emitted-attribution failure was also mine, not pre-existing: the mvp-phase.md
checkbox-OR removal moves emitted bytes, acknowledged in tests/emitted-drift-acks.

Verified on the remote runner.

* test(#3186): give the nested-plans sync fixture a passing verification

Last 3 matrix failures were one failure echoing up two describe levels. Phase
01-alpha had plans and summaries but no VERIFICATION.md, so under disk-strict
completed stayed 0 and no Progress change was emitted - correct new behavior, not
a regression.

Added the passing verification rather than dropping the Progress expectation, so
the test still covers what #3257 is about: that a nested plans/ layout is counted
and not undercounted. Probe against the built lib confirms
Progress: 0% -> 50% alongside Total Plans in Phase: 0 -> 3.

* chore(#3186): backfill changeset PR number

pr:0 placeholder replaced with the real number now that #3306 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-10 10:18:29 -04:00
Tom Boucher
95d0da9060 docs(#3287): add ADR-3180 decision 8, the diagnostic contract (#3292)
Co-authored-by: sim <sim@local>
2026-08-09 23:32:11 -04:00
Tom Boucher
693f12ad56 refactor(#3187): give state field extraction one canonical owner (#3283)
* refactor(#3187): give state field extraction one canonical owner

stateFieldValue in state-document.cts becomes the single owner of the #1760
frontmatter-then-body fallback chain. The new whole-repo guard found 14
independent re-derivations where the epic scoped 5, all now routed through it:
cmdStateSnapshot (11), cmdStatePrune (2) and smart-entry fmScalar (1).

state validate was a gate that could not fail. Every warning it could emit sat
behind a phase resolved without the frontmatter tier, so a STATE.md whose phase
lives only in frontmatter skipped the drift scan entirely and returned
valid:true. It also read unstripped content, letting a frontmatter status: key
shadow the body field (#1255 class). Both fixed; output gains a scope field so
could-not-look stops being output-identical to looked-and-clean.

Verified on the remote runner.

* docs(#3187): document the state validate scope field and its reason codes

Adds docs/how-to/interpret-state-validate-results.md so a reader can tell
nothing-to-report from could-not-look, updates the COMMANDS.md and USER-GUIDE.md
entries, corrects the CONTEXT.md glossary overstatement about Current Position
sole ownership, and drops the changeset fragment.

* fix(#3187): close three drift-guard evasion shapes and test the refuse path

The isolated adversarial review found the ladder detector was evadable by
ordinary reformatting, not just deliberately: a member or computed operand
(fm.key / fm[key]) missed the bare-identifier backreference, a swapped tier
order missed a hardcoded number-then-boolean sequence, and a ladder wrapped
across lines missed single-line detection. All three now caught, each with its
own test plus a proven boundary control.

The frontmatter-parse refuse path on the destructive complete-phase route was
unreachable and therefore untested. It is now driven by an injected parse
failure and asserts STATE.md is byte-identical after the refusal, rather than
shipping untested defensive code on a path that rewrites user state.

Verified on the remote runner.

* fix(#3187): widen the drift guard to the prompt layer and disclose tier-2 changes

The code-review spec axis found the guard's scan surface was src/ only, which is
Decision 4(d)'s forbidden allowlist one directory wide - and it had a live miss:
gsd-core/workflows/smart-entry.md tells an agent to read status from frontmatter
or the body, a prose expression of this same chain. The surface now covers the
prompt layer. That one site carries a permanent written exemption rather than a
ratchet: it is the gsd-tools-is-down fallback, so it cannot call the owner by
construction, and a ratchet would imply removable debt that does not exist.

Two tier-2 output changes shipped undisclosed and are now named in the changeset
and docs: complete-phase's idempotency guard consulting frontmatter, and the
workstream inventory resolving frontmatter-only fields. docs/COMMANDS.md gains a
state complete-phase entry, which it never had.

Also records Amendment 5 on ADR-3180, extracts the duplicated frontmatter-parse
block the epic's own thesis forbids, and re-points two assertions from free-form
warning prose onto the structured drift object.

Verified on the remote runner.

* chore(#3187): backfill changeset PR number

pr:0 placeholder replaced with the real PR number now that #3283 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:49:39 -04:00
Tom Boucher
cf6de5e1c0 feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution

23 tests over the 50-test-matrix rows. RED by construction:
resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet,
and the validator silently ignores triggerPrecedence today.

Written in the per-runtime describe idiom the other four
runtime-artifact-layout suites use, not a table.

The rows that carry the weight: windsurf must NOT report a shadow it
does not have, since its global scope emits only agents and agents are
not trigger-bearing; agents and kimi-agents must be absent from the
output for every runtime; and reordering a runtime's triggerPrecedence
must flip the winner, which is the only assertion that proves the axis
is read rather than decorative.

Stems are injected, never scanned, so the surface is assertable with no
filesystem.

* feat(#2871): resolve triggers and host precedence, not just placement

resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger
a runtime emits, with the scope and kind that produced it, whether the
host registers it directly or only through a router, and which artifact
shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers
need placement only and the issue requires them unchanged.

AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The
host-integration matrix models command and dispatch as separate interface
points: an agent is invoked through the Agent tool's subagent_type, not
by typing a slash trigger, and _copyStaged never applies the kind prefix
to an agents entry. So agents and kimi-agents are excluded from the
surface entirely, and this commit amends ADR-2866 with a dated
correction. #2218's conclusion is unchanged -- the collision is strictly
commands-vs-skills, and claude's local /gsd-* trigger surface is still
fully shadowed -- but the ADR implied the local agents surface was lost
too, and it is not.

That correction is what makes windsurf come out right. Its global scope
emits only agents, so it has no global trigger and its local commands
are unshadowed. Model agents as trigger-bearing and windsurf falsely
reports a full shadow.

The triggerPrecedence axis lands on all 19 descriptors as an ordered
kind list, one value with one owner, rather than a numeric rank spread
across N kind entries with nothing keeping them consistent. Validation
uses a required-with-default shape that has no precedent in this
validator -- every existing axis is hard-required -- so a third-party
capability.json omitting the field still validates, which is what
ADR-894's additive-only contract promises.

Winner resolution reads Phase 1's scope rank first, then the kind
ordering. A test reorders the axis and asserts the winner flips, since
an axis that is added, validated and never consulted would pass every
other assertion.

shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this
issue's out-of-scope note.

Verified via the remote runner.

* fix(#2871): single-source namespacedByDir and close two test gaps

Four findings from the isolated adversarial review.

The namespacedByDir rule had reached three copies -- install-engine,
surface, and the new trigger resolver -- one of which carried a
hand-written keep-in-sync comment and no assertion. That is this repo's
generative-fix-divergence class. Extracted to one exported predicate all
three now call. Verified by diverging one copy deliberately: the existing
#816 parity test failed, and passes again on revert.

The omission test was vacuous. Row 16 asserted that a descriptor without
triggerPrecedence still validates, but built its fixture from claude's
shipped descriptor -- which this PR had just added the axis to. It now
clones and deletes the key, following the shippedDescriptorWithout
pattern, and asserts both that validation passes and that the resolver
still picks the right winner from the default. The second half is what
makes it prove anything.

resolveTriggerSurface silently dropped an unrecognized scope while every
sibling in this epic throws. Two phases of one epic should not disagree
about whether an invalid scope is an error, so it now rejects through the
same shared validator; an empty scope list still returns empty rather
than throwing.

The ADR amendment had been spliced into the middle of the References
list, orphaning its last bullet. Moved to the top, after the header
block, which is where ADR-3660 and ADR-1016 both put dated amendments.
No lint checks markdown structure, so this was green while malformed.

* fix(#2871): single-source the command filename composition too

The earlier fix shared the namespacedByDir boolean but left the
filename composition around it written twice -- once in _copyStaged as
what actually gets written, once in resolveTriggerSurface as what gets
predicted. The predictor could go stale silently.

One exported helper now composes it for both. The entry.name asymmetry
that looked like it would block extraction does not: entry.name is
filtered to end in .md and stem is entry.name minus those three
characters, so the two branches are the same string by construction.

Divergence proven to fail: injecting a marker into the helper broke the
trigger-surface suite; reverting restored 25/25. The four sibling layout
suites hold at 227 unchanged.

* docs(#2871): correct the ADR timing notes that this phase makes stale

The Amended by back-links on ADR-3660 and ADR-1016 were written in
Phase 0, when the widenings they describe had not shipped. Each carried
a forward-looking clause -- "the module changes at Phase 2, not before,
until then this module resolves placement only" -- which becomes false
the moment this PR merges. ADR-2866's own Amends header and its
reciprocal-notes section carried the same tense.

All four now describe what shipped. This is a tense and status
correction on Accepted ADRs, not a change to any decision.

Worth stating because it is the failure mode this epic keeps meeting:
gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI
would have caught either the missing back-link in Phase 0 or these stale
clauses now. They stay correct only because someone checks.

* chore(#2871): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:25:42 -04:00
Tom Boucher
b901d1e06f feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook

60 behavioral cases against src/complexity-trigger.cts, which does not exist yet:
decision-point counting, the comment/literal stripping leak surface, threshold and
jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics,
and fs fault injection via mock.method. Two fast-check properties assert that
stripping never manufactures a decision point and that comments and string
literals are score-neutral.

Also registers the refactor-trigger capability manifest (inert until
refactor.trigger_enabled) and regenerates the capability registry and matrix.

Verified RED on the remote runner before any implementation exists.

* feat(#1953): complexity-triggered refactor extension point

Adds the opt-in refactor-trigger capability. After a phase executes, an
execute:post step measures per-function complexity for the files the phase
touched and writes a scoped refactor proposal when a function crosses the
configured threshold or drifts past its recorded anchor.

Design notes worth carrying:

- The signal is computed in-core (decision-point counting over comment- and
  literal-stripped source, Node builtins only) rather than via Memtrace or a
  shelled-out analyzer. The hook fires as a deterministic CLI, not an agent
  with MCP tools, and core takes no external dependencies — this is the only
  option a behavioral test can bind to. The metric sits behind a seam.
- The baseline is a stable anchor, not a rolling value: set on first
  observation, moved only on disposition. A rolling baseline makes the delta
  the single-phase change, so a function creeping +2 per phase never trips a
  delta of 5 and the jump check adds nothing over the absolute threshold.
- Strict mode records an open deviation window in the broken-windows ledger
  rather than declaring its own ship:pre gate. ship.md has no generic ship:pre
  gate dispatch — only two hardcoded branches — so a third gate of any kind
  would be declared and never evaluated.
- The gate clears on the proposal being dispositioned, never on the score
  improving. A blocking complexity number is one an executor can satisfy by
  splitting a coherent function in two.

execute-phase.md gains a generic execute:post step-dispatch contract; it
previously matched only ref.skill == "code-review", so any other step
registered there was declared and never run. The code-review branch is
unchanged.

Full rationale in ADR-1953.

Closes #1953

* fix(#1953): close git option injection and symlink escape in the refactor hook

Three findings from the isolated security review, all fixed inline.

HIGH — changedFilesSince interpolated the --since value into a revision
token placed before the -- separator. A -- only stops PATHSPEC parsing of
arguments after it; git still option-parses what comes before. So
--since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts
as --output=<file> and uses to redirect diff output — an arbitrary write.
Fixed with --end-of-options before the revision range plus a conservative
ref validator. The validator deliberately permits ~ ^ @ { } because those
are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct
from ref-NAME syntax; --end-of-options is the actual barrier. The doc
comment asserting the trailing -- was sufficient was wrong and is corrected.

MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink
committed inside the repo passed the check (its own path is under cwd) and
readFileSync then followed it outside the root. Now lstat-checks for a
regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one
bad path skips one file and the run continues.

LOW — the new execute:post dispatch contract showed the gsd_run example
before the rule requiring ref.command be validated first. That prose is
executed by an agent, so textual order is execution order. Reordered.

Refs #1953

* fix(#1953): make the analyzer able to see TypeScript at all

Found by running the shipped analyzer over its own source: it reported
functions=1 for a 940-line module with 24 function forms. A return-type
annotation or a generic parameter list made a function invisible —
`function f(a): number {}` and `function f<T>(a: T): T {}` both detected as
zero. Since gsd-core is written in .cts and the capability declares
.ts/.cts/.mts analyzable, the feature silently found nothing in this repo's
own primary language while reporting success. A safety net that reports
"all clear" because it cannot see is worse than no safety net.

All 98 tests passed over this, because every fixture was plain JS — the
exact failure the test matrix's own "assert against the shape production
uses" warning describes. Adds a TypeScript-shapes suite covering return
types (including unions, generics, object literals and type predicates),
generic parameter lists (constrained and defaulted), export/async/generator
combinations, annotated arrows, class-method modifiers, and optional/
default/rest params — plus the two traps: an overload signature has no body
and must not count, and `a < b && c > d` is a comparison, not a generic.
Detection now reports 24/37/21 functions for the three source files, which
matches a hand count exactly.

Also from review:

- The strict-mode ledger dedup identified entries by parsing a prose
  description string. That is banned by CONTRIBUTING's raw-text-matching
  rule and was a real bug: the "exactly one window per untriaged proposal"
  guarantee rested on prose matching, so rewording a description or editing
  WINDOWS.md by hand silently produced duplicates. Now matches structurally
  on kind + phase + file + line.
- A property test asserted on the stripper's output text. Reframed to
  assert the same invariant through analyzeSource's score.
- nextBaseline's `candidates` parameter has been dead since the anchor
  change; removed from the signature and all call sites.
- Extracted the duplicated require-or-degrade and capability-check
  boilerplate.
- ADR-1953's Implementation bullet still named a `refactor.ship-gate` in
  check-command-router.cts — a leftover from the design cut D6 rejects.
  That file is untouched and no such gate exists. Removed.

Refs #1953

* fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test

Five of the seven remote-runner failures were one cause: the execute:post
dispatch contract, written out inline, grew execute-phase.md 1876 bytes
(93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack
does not clear that — tests/phase6-capstone-conformance.test.cjs and
tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is
literally under the cap.

The contract now lives in gsd-core/references/loop-hook-dispatch.md, which
already claimed to be the point-agnostic dispatch reference and already
documented ref.skill and ref.agent. It gains the ref.command shape, its
in-context validation rule, the advisory-by-construction statement, and a
note that a point whose workflow hand-rolls one kind is not implementing
this contract. execute-phase.md now defers to it in one line: 145 bytes of
growth, 55 B of headroom under the cap. Better placement than the first cut
— the reference was overstating its coverage, and this makes the claim true
rather than duplicating prose next to it.

Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather
than a new 1953-*.json: two ack sources may never name the same path, and
that fragment is already the accumulating ack for this file.

Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own
vacuity guard — the fixture generated ~480 KB against a `> 500000` assert,
so the guard fired and the three assertions after it never ran. The test
has been vacuous since it was written. The matrix row specifies ~1 MB, so
N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000.
Verified by reproducing the exact body against the compiled module: 1168888
bytes, 118 ms, all four assertions hold.

Refs #1953

* fix(#1953): fold the execute:post step deferral into the existing resolve line

The remaining two failures were one test: execute-phase.md carries a SECOND,
tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its
current size. The file cannot grow by a single byte. My previous fix got it
under 93,600 but not under 93,400, so it still failed. ("H." in the report is
just the parent describe of that same test, not a separate defect.)

Rather than add a paragraph, the deferral now REPLACES the existing hook
resolution line. It read:

  Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where
  `kind == "step"` and `ref.skill == "code-review"`.

which is the bug itself written down — only code-review was ever dispatched.
It now reads:

  Dispatch each `kind == "step"` hook per
  @gsd-core/references/loop-hook-dispatch.md. For `code-review`:

The following prose already begins "If no active code-review step hook
exists", so it reads correctly and the code-review handling is untouched.
Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin
assertion rather than merely under the ceiling.

That also removes the need for a drift-ack: the file shrank, so there is no
growth to acknowledge, and the append to the shared 2930-*.json fragment is
reverted. Leaving it would have shipped a claim of "145 bytes of growth"
that is no longer true, on a file six other issues share.

The test's own comment states the principle this ended up honoring: "the host
loop must stay small — optional-feature detail belongs in the capability
fragment, not the host workflow." Putting the dispatch contract in the
reference rather than inline is that rule, applied.

Refs #1953

* fix(#1953): keep the code-review hook literal the workflow test requires

tests/code-review.test.cjs extracts the <step name="code_review_gate"> block
and asserts it contains `ref.skill == "code-review"` verbatim. The previous
commit replaced the line carrying that literal, so the token vanished and the
test went red — a fair assertion: code-review IS the bespoke branch there and
the workflow should still name it.

Restored inside the same one-line deferral, which now reads:

  Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md.
  `ref.skill == "code-review"`:

93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below
the base, so the file continues to shrink rather than grow.

Because three consecutive runs were each reddened by a different assertion on
this one file, this change was verified by sweeping ALL of them at once rather
than one run at a time: every test under tests/ that reads execute-phase.md or
references/loop-hook-dispatch.md was located by resolving its path constants,
and each content/size assertion was evaluated directly against the working
tree — 22 assertions, plus two real executions (gen-section-manifest --check,
and emitted-attribution's full real-tree differential). All pass.

That sweep also confirms the earlier judgement call: the net change to
execute-phase.md is a SHRINK, and the size ratchet only gates growth, so
reverting the append to the shared 2930-*.json ack fragment was correct — an
ack would have been both unnecessary and factually wrong.

Refs #1953

* chore(#1953): backfill changeset pr number to 3261

* docs(#1953): add the missing how-to for acting on a refactor proposal

Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md,
FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and
that is the one a user reaches for. CONTRIBUTING's required-docs table is
'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap.

Enabling this feature is genuinely multi-step and no single page walked it:
turn it on, tune the threshold, understand advisory vs strict, discover
that strict needs a SECOND toggle on a DIFFERENT capability, and know what
to do when a proposal appears. The two-toggle subtlety in particular was a
footnote in a config table; here it is a section with both commands.

Follows the shape of its closest siblings, resolve-edge-coverage-findings
and resolve-prohibition-findings — both 'the loop surfaced a finding, here
is what to do with it'. Includes a reason-code table for the silent cases,
since the analyzer is deliberately quiet in six situations and a user who
expected a proposal needs to tell 'nothing to report' from 'could not look'.

Indexed from docs/README.md beside the other loop how-tos.

Docs-only: exempt from the push gate, no re-verification, pass marker on
2af188b4 untouched.

Refs #1953

* feat(#1953): warn when strict mode is on but nothing will actually block

Closes acceptance criterion 5, which I had wrongly marked satisfied.

refactor.trigger_strict records an untriaged proposal as an open deviation
window, but a ship only STOPS if workflow.windows_enforce is also on — a
toggle owned by the broken-windows capability that this feature neither sets
nor requires. So a user could enable strict, believe ship was gated, and find
out otherwise at ship time.

The split itself stays: requires:["broken-windows"] would force-install the
ledger on advisory users who never enable strict, and a ship:pre gate of our
own would never fire because ship.md has no generic ship:pre gate dispatch.
What was missing was discoverability, so that is what this fixes.

`refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning,
naming the exact remediation command, whenever strict is on and either
workflow.windows_enforce is off or broken-windows is unavailable. It fires
only on a run that produced a candidate — with nothing to block on there is
nothing to warn about, and warning every run would be noise.

Reads workflow.windows_enforce through the same resolveConfigKey walk the
router already uses for its own keys rather than a second config reader.
Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does
not, strict+ledger-absent warns, strict-off never warns.

Also corrects a user-facing message in this same file that told the user to
run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the
broken-windows capability both use `gsd config-set`, and gsd-tools is invoked
as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent
messages in this file now agree.

Refs #1953

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:52:47 -04:00
Tom Boucher
653f95e39f chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias (#3272)
* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal

Inverts the Phase 5a rows that assert the derived legacy alias still
contributes a reviewer slug, and adds the removal-warning coverage the
alias's exit needs (ADR-2782 D9).

RED against unmodified production code, by design: the six shipped
manifests still declare the key and collectReviewerWarnings emits nothing
for hostBehaviors.

Refs #2801

* chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias

ADR-2782 D9, Phase 7 — the final phase of epic #2782.

The derived legacy alias survived one release (Phase 5a shipped in 1.9.0;
1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer
body is now the only route onto the reviewer roster.

- deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli
- the key is stripped from the six manifests that carried it; each already
  declares a reviewer body whose slug equals its capability id, so the
  derived roster is unchanged at the same twelve slugs
- collectReviewerWarnings emits a presence-based, non-fatal removal notice
  for any manifest still declaring the key, reaching both the build-time
  registry generation and the third-party overlay load path. The check runs
  before the reviewer-body early-return, because the manifest it exists for
  is the alias-only one that has no body.
- hostBehaviors stays an open, unvalidated bag for its other 59 keys; this
  adds one keyed removal notice, not general validation

Refs #2801

* refactor(#2801): give the reviewer-warning channel a typed IR

Review finding: the new tests asserted with String#includes() on the
warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on
Test Outputs' bans in favor of a typed intermediate representation.

Adds the IR beside the renderer rather than replacing it, which is the
shape that section prescribes and bin/verify-reapply-patches.cjs already
models:

- REVIEWER_WARNING, a frozen code enum
- REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one
  symbol instead of duplicating a literal
- collectReviewerWarningRecords(cap), returning typed records

collectReviewerWarnings(cap) keeps its exact string[] contract as a thin
map over the records, so both production consumers are untouched. Every
section-K row now asserts on record.code/field/capId and none on the
rendered message. Locks the code surface, asserts the renderer stays
one-to-one with the records, and migrates the pre-existing Phase 2 test
on the same channel off prose matching.

Refs #2801

* test(#2801): invert the section F alias fall-through regression row

Caught by the remote runner: 2 unique failures on both Node lanes out of
31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's
isolated-security-review regressions — asserted that a blank reviewer.slug
falls through to the hostBehaviors.reviewerCli alias rather than dropping
the lane. That is the direct inverse of this phase's contract.

The original rationale held only while the alias existed. With it gone
there is nothing to fall through to: a blank body is not a declaration,
and a declaration is the only route onto the roster.

Inverted rather than deleted — the row carries the adversarial-review
provenance for the slug trim, and removing a security regression guard to
make a change pass is backwards. The duplicate row added earlier in
section C is dropped instead; section F is its canonical home.

Also corrects two count strings Phase 5b left at eleven while asserting
twelve, which would misreport on failure.

Refs #2801

* docs(#2801): give the removed reviewerCli flag a migration path

The Reference edit alone satisfied CI — a file under docs/ moved, so
lint-docs-required.cjs was green — while the task-oriented quadrant said
nothing about the removal. A maintainer whose lane had just gone silent
would have found the field documented as removed and no page telling them
what to do about it.

Adds a migration section to the how-to: the symptom, the verbatim warning
they will see, the before/after manifest, and the note to keep the
reviewer slug equal to the capability id so existing
review.default_reviewers entries and --<slug> flags survive.

Refs #2801

* chore(#2801): backfill changeset pr number to 3272

* feat(#2801): close the runtime.hostBehaviors vocabulary

ADR-1016 closes twelve descriptor axes and rejects an open escape hatch
in the descriptor. It never mentioned runtime.hostBehaviors, and that
silence was read as permission: 59 keys across 18 manifests, 39 of them
set by a single capability, validated by nothing. The reference docs went
further and attributed the open seam to ADR-1016, which does not mention
the field at all.

KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a
non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the
alias removal notice, reaching both build-time generation and overlay
install.

Warning, never error, for the reason this phase exists: an error would
hard-break an out-of-tree descriptor carrying a bespoke key with no
deprecation window, which is what reviewerCli was given a release to
avoid. Escalation is a separate decision.

reviewerCli is excluded from the unknown-key sweep so it keeps its own
notice with the migration pointer rather than drawing two records.

A parity test binds the vocabulary to the shipped manifests in both
directions, and a second asserts no shipped capability draws a notice, so
the closure is provably inert in-tree.

Records the decision and the miscitation as an ADR-1016 amendment.

Refs #2801

* fix(#2801): bound and sanitize the unknown-key diagnostics

Two findings from an isolated adversarial review of the closure commit,
both proven by execution rather than asserted.

MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep
had no ceiling. An installed third-party manifest is bounded only by
MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced
800,000 records and ~139MB of message text, retained for the registry's
lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a
summary carrying omittedCount, mirroring capability-loader's existing
slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars.

MINOR, newly reachable: manifest-supplied key names were interpolated raw.
Unlike cap.id, which validateCapability gates on KEBAB_RE before these
diagnostics run, hostBehaviors keys have no grammar check anywhere, so
ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New
describeKey replaces C0/C1 controls and clips at 80 chars. The file
already had describeValue for this and applied it only to values.

Both fixes land on the pre-existing reviewer.* sweep too — it carried the
identical pair, and fixing only the new copy would leave the same defect
one screen from its own fix.

Refs #2801

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:18:56 -04:00
Tom Boucher
58d73dd220 enhance(#3241): omit the codex per-agent model by default (#3276)
* test(#3241): failing-first suite for the codex passive model posture

Locks ADR-2313's D1-D5 before any production code exists, so the tests
bind to the behavior rather than to whatever the implementation happens
to do.

Red-first (fail against the current tree):
  - the resolver path emits no `model` and no `model_reasoning_effort`
  - a whitespace-only model_overrides value yields no pin
  - isAnthropicFlavoredModel / CLAUDE_AGENT_ALIASES on model-catalog
  - the one-time install notice, and its once-per-install dedupe

Regression guards (pass today, must keep passing): resolver-null via
`inherit` and via absent runtime; a resolver that resolves to nothing;
empty-string and non-string overrides; and the light-tier
service_tier/model_verbosity fields, which are NOT coupled to the model
pin and would silently regress if the implementation coupled them.

Classifying each test as red-first or regression guard is deliberate.
A test that passes on both sides of the change proves nothing, and this
epic has already shipped two such rows before catching them.

The whitespace case is a live defect, not a quirk: `'   '` is truthy,
survives the type guard, is not Anthropic-flavored, and is embedded
verbatim as `model = "   "` — the same class the #2310 guard exists to
stop. Same function, same path, fixed in this phase per CLAUDE.md §3.

Two matrix rows were dropped as vacuous rather than shipped green: a
64-char truncation case (the pinned notice interpolates no
user-controlled value, so it cannot exhibit truncation) and a newline
hazard that the input surface cannot reach.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3241): omit the codex per-agent model by default

Implements ADR-2313 D1-D5. generateCodexAgentToml no longer embeds the
runtime resolver's per-tier Codex model, so an agent inherits the
always-available session model instead of a pin a ChatGPT-account Codex
may not expose. model_reasoning_effort disappears with it via the
existing hasPinnedModel coupling (#838) — no logic change needed there.

Supersedes #2517's embedding on the default path only. An explicit
real-Codex model_overrides pin is still embedded verbatim, and the #2310
Anthropic-flavored guard is retained: the model_overrides route to it is
still live even though the resolver route is now unreachable.

The shared rule moves down a layer. CLAUDE_AGENT_ALIASES leaves
model-resolver for model-catalog — a genuine leaf importing only
node:path and its own JSON — with isAnthropicFlavoredModel defined beside
it, and is re-exported from model-resolver so every existing importer is
untouched. This is what lets Phase 2's install-check and Phase 3's sync
consume the rule without taking the config-loader dependency
model-resolver would have dragged into a module documented as pure
read/verify with 33 dependents. A parity test fails if the two ever fork.

Also fixes a live defect surfaced while writing the tests: a
whitespace-only model_overrides value was truthy, survived the type
guard, was not Anthropic-flavored, and so was embedded verbatim as
`model = "   "` — the same class the #2310 guard exists to stop, reached
by a different route. Trimmed before the truthiness test. It is
deliberately not routed through _warnCodexModelOverrideDropped, whose
text would misdescribe a blank field as a mis-typed model.

Adds the one-time install notice (maintainer direction, recorded as an
ADR-2313 amendment): one stderr line naming model_overrides and the
session model, deduped per install rather than per agent, and emitted
only for the population that actually loses a pin.

service_tier and model_verbosity stay decoupled from the model (#774);
a regression guard asserts they still emit with nothing pinned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3241): amend ADR-2313, add the model-catalog glossary entry

ADR-2313 gains two dated amendments rather than edits to its merged
text, since ADRs here are append-only.

The first records that a deprecation notice IS offered, reversing the
Migration section's "no deprecation window" position, and states why
that position was wrong rather than just superseding it: the ADR
identified the API-key population as losing something real and then
declined to warn it, in the same document. Hyrum's guidance was applied
to the recourse and not to the notice.

The second records the whitespace-only model_overrides defect and notes
that D2 always implied the fix — the implementation simply never
enforced it and no test covered the case.

CONTEXT.md gains a Model Catalog Module entry. The module had none,
which is why the glossary gate passed without one: check-glossary-refs
verifies that references resolve, not that modules are documented. The
entry records why the Anthropic-flavored rule lives there rather than in
model-resolver, so a later reader does not "helpfully" move it back. The
Model Resolver entry is updated to point at its new home and note the
back-compat re-export.

docs/CONFIGURATION.md carried a claim that is now false: that the
resolved tier ID is embedded into agent frontmatter at install time on
codex and opencode. Corrected to name codex as the exception, with the
400 symptom and the model_overrides recourse.

Changeset leads with the user-visible change and the migration line
rather than the implementation, per the ADR's Hyrum's-Law analysis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3241): only notice a lost pin when one was actually embeddable

Review finding from an isolated reviewer. The deprecation notice gated
on whether the runtime resolver would have returned *any* model, but the
question that matters is whether that model would have been *embedded*.

Those differ. The #2310 safety gate already rejected an Anthropic-
flavored model arriving from the resolver path before Phase 1 — so for a
mixed-runtime config resolving to a claude-* id against a Codex install
target, the user never had that pin. The notice told them they lost
something they never got, and pointed them at model_overrides for no
reason.

The existing #2310 test drives exactly that path but asserts only the
emitted `model` line, never stderr, which is why it slipped through. Now
covered.

Deliberately unchanged: an Anthropic-flavored model_overrides value plus
a legal resolver model fires BOTH the override warning and the notice.
That is correct — pre-Phase-1 the guard dropped the override, execution
fell through to the resolver, and the resolver's model was embedded, so
that user did lose a pin. Two messages, two distinct true facts, and the
prefixes differ (`gsd: warning — ` vs `gsd: notice — `) so the
one-notice-per-install contract holds. A regression test now pins that
behavior so it does not get "simplified" away.

Of the three tests added, only the first is red-first; the other two
pass on both sides by design and are labelled as guards — one against
over-correcting the fix into silence, one against removing the
intentional double message.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3241): reset the notice dedupe via a seam, not a require.cache bust

The remote runner caught a regression I introduced: the #2760
post-write-validation test began failing with the validator override no
longer intercepting.

Cause, confirmed by trace rather than guessed: the new #3241 review
tests deleted require.cache for bin/install.js and re-required it mid
suite, to clear the notice's module-level dedupe flag. But
runCodexInstall destructures `install` at file load, closing over the
ORIGINAL module's exports. After the cache bust a second instance
existed, so the test's `installModule.__codexSchemaValidator = ...`
mutated the new object while the code under test still called the old
one. The override silently stopped intercepting, the real validator ran
and passed on GSD-emitted output, and the abort-and-restore path was
never exercised.

Cache-busting a module mid-suite breaks every later test that assumes a
single instance, which every other test in the file is entitled to. So
the fix is a seam, not a workaround: bin/install.js exports
_resetCodexNoticeDedupeForTests(), and the three tests call it directly
instead of reloading the module.

The flag is module-level by design — the dedupe is per-install and
install() already resets it — so a unit test driving
generateCodexAgentToml directly needs an explicit way to reset it. That
is now what it has.

Swept the rest of the #3241 diff for the same hazard; this flag was the
only shared module-level state introduced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3241): reset both codex dedupe stores, not just the notice flag

Second incomplete fix, same class one layer down. bin/install.js keeps
TWO module-level dedupe stores and the require.cache bust I removed had
been papering over both; my replacement seam cleared only one.

_codexModelOverrideDroppedWarned is a Set keyed `${agent}::${value}`.
tests/codex-config.test.cjs:558 already emits for `gsd-executor::sonnet`,
so by the time the review test using the same agent and value ran,
_warnCodexModelOverrideDropped was a silent no-op and the expected
warning never appeared.

The seam now clears both stores and is renamed to say so. Its comment
records that per-install dedupe lives in module scope deliberately and
that this is the single sanctioned way for a unit test to clear it.

Swept bin/install.js for every other module-scope mutable a test could
latch. Two are inert (capability registries assigned once at require
time; selectedRuntimes computed once from argv). One is a genuine latent
hazard and is deliberately NOT folded in: attributionCache (:1654)
memoizes getCommitAttribution by runtime name for process lifetime, so
two in-process installs of one runtime with differing attribution config
would collide. It is unreachable from any current test and is a
different concern from Codex warning dedupe, so it stays out of this PR
rather than widening it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3241): correct the codex tier-routing how-to

The docs gate forced the task-oriented quadrant and found the worst
defect in this change's documentation surface.

docs/how-to/configure-model-profiles.md carried a section titled "If you
want tiered models on Codex" telling users to set runtime:codex +
model_profile:balanced, promising "GSD resolves each tier alias to the
Codex-native model and reasoning effort defined in the runtime tier
map." That is exactly the behavior this PR removes — a how-to page
confidently instructing users to do something that no longer works,
which is worse than a missing page because it fails at the moment of
use.

Rewritten to state that Codex does no tier routing, give the
model_overrides pin as the supported alternative, and name the two
constraints on what may be pinned: it must be a real Codex model id, and
the account must actually expose it — GSD cannot verify the second, so
the honest advice when unsure is to omit the pin. Carries an upgrade
note for both account types, since the change is a no-op for ChatGPT
accounts and a real loss for API-key ones.

Also tightened the same page's claim that Codex "embeds the resolved
model" at install time — now true only of an explicit override. The
re-install instruction it supports is still correct and still needed, so
only the premise moved.

Both the required-docs set (COMMANDS.md + FEATURES.md) and
lint-docs-required.cjs would have passed before this commit, since
CONFIGURATION.md and the ADR had already moved. Neither checks the
quadrant a user in trouble actually opens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3241): backfill changeset pr number (#3276)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 19:16:04 -04:00
Tom Boucher
2e2b8ba4a7 enhance(#2704): resolve documentation links and compare H1 status brackets in the ADR gate (#3266)
* test(#2704): failing-first coverage for ADR link resolution and H1 status brackets

Binds the gate to two assertions it does not yet make: every relative markdown
link under docs/adr/ must resolve, and an H1 trailing status bracket must agree
with the Status: field instead of being silently stripped.

Covers all 51 rows of the phase test matrix across two altitudes - the pure
extractLinks/maskCode IR for fence and inline-code-span boundaries, hostile
input and the fast-check totality properties, and the real CLI verdict for the
end-to-end classes. Includes the DEFECT.GENERATIVE-FIX parity test that iterates
the exported STATUSES array so a sixth status is covered the day it is added.

* feat(#2704): resolve ADR documentation links and compare H1 status brackets

The ADR gate validated naming, relation symmetry and index freshness but never
resolved a link target, and it stripped an ADR's trailing H1 status bracket for
display rather than comparing it against that ADR's own Status: field. Both
classes were structurally invisible: #2691 found five dangling references by
manual audit roughly a year after they were introduced, one of which reached the
published npm payload, while CI reported green throughout.

Both are now assertions on the same --check path, using only node:fs and
node:path - no dependency and no subprocess.

Fenced blocks and inline code spans are masked before scanning, because markdown
does not render a link inside code. That is not a policy choice: the corpus
contains exactly two such sequences today and both are ordinary JavaScript.
Masking preserves length and column positions so findings still name a real line.

Resolution is case-exact on every platform - a link that resolves only through
macOS or Windows case-folding still 404s on github.com and still fails the Linux
lane - and a destination resolving outside the repository is reported before any
filesystem call is made.

Also single-sources two duplicated surfaces this change would otherwise have
extended: the H1 bracket vocabulary (a second hand-written copy of STATUSES with
nothing asserting agreement, a DEFECT.GENERATIVE-FIX instance) and the docs/adr
directory traversal. Two tests added by #2691 that reimplemented link resolution
and bracket comparison inside the test file are removed for the same reason; the
corpus assertion is now made by running the real gate against the real corpus.

* fix(#2704): reject symlinks that leave the repository and linearize code masking

Four defects from the isolated adversarial security review, plus one it noted.

BLOCKER - a symlink defeated path containment. path.relative(ROOT, abs) is
purely lexical, but the case-exact walk then calls readdirSync, which follows
symlinks at the OS level: a contributor-committed docs/adr/x -> /etc together
with a link through it passed containment and listed the real external
directory, and a wrong-case probe echoed a real external filename through the
"Did you mean" hint into publicly-readable fork-PR logs. Every segment is now
lstat'd before descent; a symlink is realpathed and re-checked against
realpath(ROOT) - realpath on both sides, so a root under /var does not produce
false escapes - and an escape emits no hint and reads nothing further.

The same rule now governs which FILES are read: an ADR entry that is a symlink
out of the repository is excluded and reported rather than parsed, closing the
vector this change had widened by newly reading README.md, naming-violation
files, and full bodies rather than only header fields.

MAJOR - inline-span masking rescanned the line remainder per backtick run,
roughly O(n^1.6) on adversarial input: 1.76s for an 800KB line. Rewritten as a
single linear pass pairing runs through forward-only per-length cursors. Same
input now takes 3.31ms, with behavior unchanged.

MINOR - an unreadable or broken entry threw, and the generic handler wrote a
raw stack trace carrying absolute CI paths to stderr. The scan is now
fault-tolerant and reports excluded entries as ordinary violations. The status
vocabulary is escaped before being interpolated into a dynamic RegExp -
defence-in-depth, not a live bug.

The containment predicate had reached three hand-written copies while fixing
this; it is now the single escapesRoot() helper used by all four call sites.

* feat(#2704): add a --json report so the gate's tests assert on typed values

Maintainer-directed addition. CONTRIBUTING.md's "Prohibited: Raw Text Matching
on Test Outputs" requires that a system under test producing text also expose a
structured intermediate representation, and that tests assert on that IR rather
than on rendered prose. This gate had no such surface, so its verdict tests
matched on stderr.

--json runs exactly the same validation as --check and writes a report to stdout
with the same exit code, following the frozen-REASON-enum pattern already used
by verify-reapply-patches.cjs. Every violation carries a stable reason code plus
the fields a consumer needs, so nothing has to pattern-match an error message.
Adding a reason stays three coordinated changes - the enum, the emitting site,
and the test locking Object.keys(REASON).sort().

The human output is unchanged, deliberately: a large pre-existing suite asserts
on it and migrating that is not this PR's concern. Verified by running the
pre-change and post-change scripts against an identical violating corpus and
diffing their stderr - character-for-character identical.

This PR's own verdict tests now assert on parsed --json. Absence checks improve
the most: "no bracket violation" is now a reason-code predicate rather than a
negative regex over prose, which could pass for the wrong reason. The security
assertions were strengthened rather than translated - no leaked filename may
appear in ANY field of the serialized report.

Unknown flags are now rejected instead of silently falling through to printing
the index.

* test(#2704): fix the status-parity fixture and guard hooks/dist before overlay builds

Two failures from the matrix run of 79b29909.

The status-parity fixture was mine. It built, per status token, an ADR whose H1
bracket and Status field both carried that token - but Superseded carries an
obligation beyond the bracket: it must name its successor as a file link and be
symmetric with it. The fixture declared a bare Superseded, tripped that
unrelated invariant, and the test reported a bracket-parity failure for a reason
that had nothing to do with bracket parity. The fixture now satisfies each
token's own obligations in both the agreeing and contradicting corpora, derived
from the status actually declared rather than special-cased on one name, so a
future token carrying obligations is handled rather than silently skipped.

The second failure was not mine but is fixed here rather than deferred.
mcp-catalog-parity.install.test.cjs hardlinks hooks/dist/* while building its
overlay, but hooks/dist is a gitignored build artifact produced only by
build:hooks. The suite had no guard, so it passed only when some other suite
happened to build it first - an execution-order dependency, which is why it
failed on node22 and passed on node24 for identical code. install.test.cjs
already documents this exact hazard and guards it.

Six behaviorally identical copies of that guard existed across three files.
Rather than add a seventh, they are now one canonical
tests/helpers/hooks-dist.cjs - idempotent and bounded by the shared
BUILD_TIMEOUT_MS class norm - which is the same single-sourcing this PR applies
to the ADR gate itself.

* docs(#2704): add a how-to for contributors the ADR gate rejects

The reference and explanation quadrants were covered by Lifecycle rules 5 and 6,
but the task-oriented one was thin: a contributor meets this gate because it
failed on their PR, under pressure, and the rules told them what is checked
without telling them what to do about it.

Adds the command to reproduce the CI failure locally and a message-to-remedy
table covering every reason code that can be hit - unresolved target, wrong case
with the did-you-mean hint, repository escape, symlinked ADR file, bracket
contradiction - plus the backtick escape hatch for illustrative links and the
caveat that indented code blocks are not skipped.

The table is itself written in backticked inline code, so the gate skips it: the
escape hatch demonstrated on the page that documents it.

* chore(#2704): backfill changeset PR number

pr:0 placeholder replaced with the real PR number now that #3266 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-09 17:08:46 -04:00
Tom Boucher
2ac21c7fdb docs(#2980): ratify the payload-carried error idiom as a degraded result (#3270)
* docs(#2980): ratify the payload-carried error idiom as a degraded result

Records ADR-2980: an `error` key in a gsd-tools result payload on stdout with
exit 0 is a ratified contract meaning the command ran to completion and is
reporting a condition, not a process failure. Faults keep stderr + exit 1 +
--json-errors.

Normalizing the 42 output({error}) sites to exit 1 was declined on measured
blast radius (get_impact rates cmdStateSnapshot CRITICAL; output has 170 direct
callers) — a Hyrum's Law break with no versioning escape hatch for an exit code.

Adds the "Degraded results vs faults" section to docs/json-errors.md with a
correct-caller recipe, indexes that page from docs/README.md, and records the
two-channel contract in the CONTEXT.md I/O Module entry. No code change.

Closes #2980

* docs(#2980): correct the site count and cross-refs after review

The isolated review found the population figure was the answer to a regex,
not to the question. `output\(\{\s*error:` only matches literals whose FIRST
key is `error`; re-deriving it by brace-matching output()'s first argument
gives 60 sites across 9 modules (42 error-first + 18 error-not-first), adding
workstream.cts, phase.cts and gsd2-import.cts. roadmap.cts:260 is the case in
point — it carries `error` alongside `found:false` and is the site that
actually produces the documented `roadmap get-phase` output.

Also from review: correct the --raw claim (11 sites pass a rawValue, not 2),
reconcile the missing-required-argument count to the 7 verified sites, link
the bare ADR-2966 references per the ADR lifecycle rule, and fix 'licence'
to American spelling.

---------

Co-authored-by: sim <sim@local>
2026-08-09 16:44:29 -04:00
Tom Boucher
c07297cd50 docs(#2869): record ADR-2866 install-surface resolution (#3265)
Phase 0 of epic #2866. Records the decision that the install pipeline
resolves surface identity — (runtime × scope × trigger) — as a value
instead of implying it from destination paths.

The ADR makes four things explicit:

- Amends ADR-3660 (placement -> placement + trigger resolution) and says
  why placement-only stopped paying: the /gsd-<name> trigger two
  artifacts collide on is not a value anywhere, so #2218 cannot be
  stated by any module or test.
- Adds one axis to ADR-1016's deliberately-closed descriptor vocabulary
  (host trigger precedence), required-with-default so ADR-894's
  additive-only stability contract holds.
- Records the @-include constraint (expands ~, does NOT expand env
  vars, no conditional syntax) as the reason #2218 triage option 1 is
  REFUTED rather than merely deprioritized.
- Notes the non-conflicts: completes ADR-58 rather than revising it,
  and preserves ADR-1508's dependency direction.

ADR-3660 and ADR-1016 each gain the reciprocal Amended by back-link,
matching the corpus convention ADR-2782 already set on ADR-1016. Each
states that the decision is recorded now while the modules change at
Phase 2 (#2871), so no reader is told a widening has already shipped.

Also corrects CONTEXT.md's Installer Module entry: bin/install.js is
hand-authored, not generated. ADR-1508 states this verbatim and no
build step emits it; the stale annotation invites contributors to look
for a generator that does not exist.

Docs-only. Verified via the remote runner.

Closes #2869

Co-authored-by: sim <sim@local>
2026-08-09 16:15:42 -04:00
Tom Boucher
b183317abd docs(#3256): ratify ADR-2363 to Accepted (#3260)
Both phases of epic #2363 have landed and the epic is closed as
completed, so ADR-2363's own stated ratification bar is met: #3248
merged, the consent summary renders instruction surfaces, and a passing
test pins the D4 signature behavior.

Adds the dated Ratification section the corpus requires, naming the
files, symbols and tests for each of D1-D5, and restores the reciprocal
back-link on ADR-1244 - owed only on ratification, which is why the
premature flip in 4d26887e correctly withdrew it.

Records the judgment call the ratification rests on rather than burying
it: D3 classifies instruction surfaces as skills and agents, and the
mechanism discloses skills only. Third-party agents are never staged
into the instruction context, so disclosing them would have named a
surface that does not exist. Everything actually staged is disclosed,
which is what the decision requires; whether agents should be staged is
recorded in D5 as an open maintainer question.

Docs-only. No behavior change.

Closes #3256

Co-authored-by: sim <sim@local>
2026-08-09 14:24:05 -04:00
Tom Boucher
9f57fa43ed docs(#3240): record the codex passive/session-only model posture (#3251)
* docs(#3240): record the codex passive/session-only model posture

ADR-2313 locks the install-time contract for epic #2313: omit the
per-agent model from generated ~/.codex/agents/<agent>.toml by default
so the agent inherits the always-available Codex session model, embed
one only for an explicit real-Codex model_overrides pin, and keep
model_reasoning_effort coupled to a pinned model (#838). Supersedes
#2517's per-tier embedding on the default path only.

Also records the reader/writer boundary the downstream phases need
(strict writer, liberal-but-visible readers, never partially rewrite an
unparseable .toml), the migration path for API-key Codex users, and the
Phase 5 the coverage gate found unowned.

Amends ADR-1239 with a dated section: its effortSurface amendment
described this ADR as "not yet written", and the install-time vs
invocation-time boundary is now stated from both sides.

Docs-only. The posture is not real until Phase 1 (#3241) merges.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3240): remove the ADR index count cells that race between PRs

The generated region of docs/adr/README.md carried three numeric cells —
a per-group `### <heading> (N)` and a `_N ADRs._` footer — that every
ADR-adding PR must rewrite. Two PRs adding different ADRs merge their
table rows cleanly, since those are distinct lines, but both rewrite the
same count lines, so whichever lands second gets a green local
`gen-adr-index.cjs --check` and a red CI one: CI evaluates the PR merged
with next, where the count reflects both ADRs.

That is not hypothetical. It reddened this PR: ADR-2313 regenerated the
index at 75 while #3249 landed ADR-3247 concurrently, making the merged
tree 76.

The counts carry no verification value — --check regenerates and diffs
the whole region regardless — and are derivable by reading the table, so
they are removed rather than tolerated. Loosening --check to ignore them
would have let genuine staleness through. This is the shared-mutable-cell
problem CHANGELOG.md and the drift acks already solved with per-PR
fragment files; here removing the cell is enough.

The regression test locks the invariant rather than the symptom: adding
an ADR only INSERTS lines, so render(N) is a line-subsequence of
render(N+1). That is the property that makes concurrent PRs merge, and
unlike asserting the absence of one count format it fails for a count
reintroduced in any shape. Covered at append, lowest-id, middle-id,
empty-corpus, new-status-group, and hazardous-title positions; each names
the pre-fix line that would have failed it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 13:53:41 -04:00
Tom Boucher
f96cb44f85 enhance(#3248): disclose capability skills as an instruction surface (#3253)
* test(#3248): failing-first suite for instruction-surface disclosure

28 matrix rows from 50-test-matrix.md. Rows requiring the new
Disclosure.instructionSurfaces field fail today; rows 18-20/23-25 (the
ADR-2363 D4 signature invariants) pass today by construction because the
current code never reads skills/agents at all, and stand as regression
guards for the implementation commit.

Refs #3248

* feat(#3248): disclose capability skills and agents as an instruction surface

ADR-2363 D5. A capability whose only contribution was skills disclosed
nothing at install: summarizeDisclosure early-returned "ships no executable
surfaces (declarative only)" because hasExecutable was false, while each
SKILL.md body landed verbatim in the agent's instruction context.

discloseExecutableSurfaces gains a fifth, NON-executable class,
instructionSurfaces, collecting declared skills/agents stems through the same
safeCollect wrapper as the four existing collectors, so a hostile value
degrades only this class and the function stays total for any manifest shape.
Nothing existing is edited: the collectors, hasExecutable, disclosureSignature
and missingArtifacts are untouched. get_impact rates the symbol CRITICAL at
196 affected, which is why the design is strictly additive.

D4 is implemented by omission and pinned rather than left incidental: adding,
changing or removing skills/agents leaves disclosureSignature byte-identical,
so no stored consent record is perturbed and no spurious re-consent fires.
ADR-2782's conditional-append trick is deliberately NOT reused - it worked
because no manifest could declare a reviewer body before that class existed,
whereas skills predate this one, so a conditional append would re-sign every
already-consented skill-bearing capability.

The renderer is extracted as summarizeInstructionSurfaces and called from BOTH
branches of summarizeDisclosure. Appending only at the end would never render
for skill-only capabilities - the ones that need it - since those take the
early return. That branch's "declarative only" claim is now conditional on
there being no instruction surface either. The renderer iterates rather than
spreading into push, so an unbounded stem count cannot throw RangeError, and
tolerates the bare {} the CLI edge passes via `res.disclosure || {}`.

Scope note: #3248's prose says "skill stems"; ADR-2363 D3 classifies
instruction surfaces as "skills, agents". Shipping skills alone would leave an
ADR deliverable owned by no phase, and the epic has no Phase 2. Agents are the
same shape at no extra cost. Narrowing back is a two-line change.

Ratifies ADR-2363 (Proposed -> Accepted) and adds the owed ADR-1244 back-link.

Closes #3248

* fix(#3248): escape consent-prompt values and narrow disclosure to skills

Two review findings, both of which made the previous commit wrong.

BLOCKER (isolated adversarial review). Every manifest-supplied value
interpolated into a consent-prompt line was rendered unescaped. Those lines
are joined with \n and written RAW to stderr on the needs-consent path
(capability-command-router -> cli-exit runMain), so a stem carrying a newline
forged additional lines indistinguishable from genuine GSD disclosure text,
and an ANSI escape could clear or rewrite lines already printed. That defeats
the informed-consent guarantee this change exists to provide, and is a
prompt-injection vector against any agent that reads the stderr text to decide
whether to retry with --yes.

The hole was not unique to the new class - hook event/script, command
family/module/router, every MCP field, and every reviewer-lane field were
equally unescaped. Fixing only the new one would have created the
generative-fix divergence this repo tracks, so renderValueForPrompt is applied
to all five classes through one helper, guarded by a parity test that fails if
a future class skips it. Escaping is identity for ordinary names, so no
well-formed manifest's output changes. The disclosure OBJECT stays verbatim -
only the rendered LINE is escaped - because the signature and every consumer
reasoning about identity depend on the declared value.

NARROWED to skills only. The previous commit also collected agents, arguing
ADR-2363 D3 classifies instruction surfaces as "skills, agents". Verified
against staging: stageSkillsForRuntimeAsSkills takes a registry and unions
third-party skills in via readInstalledCapabilitySkill, while
stageAgentsForRuntimeWithConverter takes only a source directory and has no
registry-aware path. Third-party agents are never staged into the instruction
context, so disclosing them would have put a false claim in a security prompt -
worse than the scope creep two reviewers flagged it as. D3's classification
stands; D5 now records that Phase 1 implements the skills half and that
whether agents should be staged at all is an open maintainer question.

Also reverts the premature ADR-2363 ratification. The previous commit flipped
it to Accepted and asserted "#3248 merged" while this branch IS #3248 and is
unmerged. Status returns to Proposed, and the ADR-1244 back-link - owed only on
ratification - is withdrawn.

Adds the fast-check property suite CLAUDE.md requires and the direct precedent
(reviewer-trust-disclosure) already had: totality, D4 signature invariance, D3
hasExecutable invariance, and renderer totality over adversarial manifests.

Refs #3248

* chore(#3248): correct changeset scope claim and backfill pr number

The fragment was written against the pre-narrowing commit and still
advertised 'skills and agents'. 4d26887e narrowed disclosure to skills
only - third-party agents are never staged into the instruction context -
but did not touch the fragment, so the release notes would have carried a
claim the code does not implement.

Also backfills pr:0 -> 3253 and names the prompt-escaping fix, which is
user-visible and was absent from the original body.

Changeset-only; no code or test changed, so the gsd-test pass recorded for
4d26887e still describes this tree's behavior.

Refs #3248

---------

Co-authored-by: sim <sim@local>
2026-08-09 13:52:29 -04:00
Tom Boucher
bc5619dd27 docs(#3247): record the capability instruction-surface trust model (#3249)
* docs(#3247): record the capability instruction-surface trust model

ADR-2363 records the trust posture for third-party capability SKILL.md
bodies, which #2322/#2340 made agent-invocable without any content-level
control. The path-level protections that fix shipped are all present; no
content scanner exists, and external-descriptor-trust.cts never had one.
Nothing was bypassed - the control did not exist and the boundary was
never written down.

D1 records the posture: skill bodies are trusted, unscanned agent
instructions. D2 rejects content scanning on Kerckhoffs (a shipped rule
set is readable by the adversary who installs it), on threat-model
non-transfer from ADR-1577 (there, instructions are anomalous inside
data; here they are the payload's legitimate form), and on Goodhart (a
scanned-OK line displaces the judgment the consent prompt exists to
provoke). D3 replaces the executable/non-executable binary with three
classes, adding instruction surface.

D4 keeps instruction surfaces out of the v1 disclosureSignature. The
signature is NOT the activation binding - hasProjectConsent compares
contentHash only, and a global install carries no consent record at all.
What re-encoding would do is perturb the signature of every skill-bearing
capability and fire a spurious re-consent prompt on its next upgrade,
which is what ADR-2782 D4 rule 5 already forbids. If instruction surfaces
ever need to be signature-bound, that lands as a versioned v2 signature
with a migration, never an in-place re-encoding.

Corrects capability-trust-model.md, which claimed skills get lighter
consent because they do not execute code - true, and not the relevant
property, since the agent is the interpreter. Adds the author-side
boundary to develop-a-capability.md and links it from
publish-a-capability.md. Both state that per-skill disclosure at the
consent prompt lands with #3248 and does not happen today.

Docs-only. No behavior change; no consent record perturbed. D5's
mechanism is Phase 1 (#3248), which is why the ADR is Proposed.

Refs #2363

* chore(#3247): backfill changeset pr number to 3249

---------

Co-authored-by: sim <sim@local>
2026-08-09 11:47:54 -04:00
Tom Boucher
2a73f53cb3 fix(#3204): milestone sectioning is vocabulary, not heading position (#3230)
* test(#3204): failing-first suite for the clobbered phase count

A project declaring six phases with four phase directories on disk had
state.record-session write progress.total_phases: 4 — #2828 regressing at
1.9.1, reported in #3204 with a deterministic reproduction.

Before the fix in the following commit, these rows FAILED (wrote 4, expected
6): a flat roadmap carrying `## Progress`; one carrying `## Overview` and
`## Phase Details`; the CRLF variant of the first. Two more, found by
adversarial review and added after the first fix attempt, failed against that
attempt: structural headings interleaved among flat phase headings, and this
repo's own bundled-template shape (a `## Phases` wrapper around a single
nested milestone).

The #1761 control — sibling milestone sections must keep falling back to the
disk count — passes both before and after, so the fix has something it must
not break.

Assertions read progress.total_phases through the product's own frontmatter
parser via `state json --raw`, never a regex over STATE.md. Rows 12 and 13
are hostile: a phase heading carrying a version token, and a version heading
inside a fenced code block; neither may count as milestone sectioning.

Refs #3185, #3204

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3204): milestone sectioning is vocabulary, not heading position

buildStateFrontmatter chooses total_phases between the ROADMAP's declared
phase count and the on-disk directory count, and refuses the roadmap count
when hasMilestoneSectioning says the document is milestone-sectioned — because
a whole-document count would then conflate sibling milestones (#1761).

That predicate returned true for ANY non-Phase level-2/3 heading, so a flat
roadmap carrying an ordinary `## Progress` was called sectioned and the disk
count clobbered the declared one: six declared phases, four directories,
total_phases written as 4, converging on the truth only once the last
directory happened to exist. That is #2828 regressing at 1.9.1, and it came
from this epic — #3184 replaced state.cts's hand-rolled #2828 guard with this
predicate, and the replacement is strictly more permissive than the guard it
retired.

Three position-based models were tried and all failed, because position does
not carry milestone-ness:

  - any non-Phase heading (shipped) — over-detects, giving #3204;
  - strict nesting/ownership — misses same-level siblings, regressing #1761,
    and false-positives on the bundled template, where `## Phases` wraps a
    single `### v1.1`;
  - adjacency — reproduced live: `## Overview` and `## Notes` interleaved
    among six phase headings are two owning candidates, so a 6-phase roadmap
    with 2 directories wrote 2.

A heading is now a milestone heading iff it is a non-Phase heading carrying a
milestone signal: a version token, a status marker, or the word Milestone.
Sectioning means two or more, since one cannot conflate siblings.

Known limit, recorded in the doc comment rather than hidden: two milestone
sections carrying none of those three signals are not detected.

Also drops buildStateFrontmatter's local dedup-key regex, flagged in-source as
diverging from the canonical token rule, for phaseKeyFromDir — the remainder
of #3185, since #3222 had already routed the enumeration itself through
listMilestonePhaseDirs.

#1514, #2445 and #3017 are preserved untouched.

Closes #3185
Fixes #3204

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3185): changeset, glossary entry and ADR status for the phase-count fix

CONTEXT.md's Roadmap Parser Module entry never named hasMilestoneSectioning,
so the predicate whose semantics this change reverses had no glossary presence
at all — a PR gate for a module/seam change. Added, covering the vocabulary
model, the three position-based models that failed, and the residual limit.

ADR-3180 recorded the fifth enumeration copy as unowned in four places. It is
owned now. Amendment 4's scope table row 1 also carried an error worth keeping
visible rather than rewriting: it claimed Phase 3 merged without routing the
state writers, when #3222 had in fact routed the enumeration — the audit read
Amendment 3's silence about the symbol names as absence of the work. The real
gap was the trust discriminator one layer above, which is what #3204 was.

Changeset is Fixed and leads with the symptom a user sees — a phase count that
shrinks to match how many phase directories happen to exist yet — and carries
the known limit forward rather than leaving it in a source comment.

Refs #3185, #3204

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3185): stop quoting the retired phase-token regex in a comment

The remote runner failed tests/phase-id-drift-guard.test.cjs: the comment
explaining that the local dedup regex had been replaced by phaseKeyFromDir
quoted that regex verbatim, and scripts/lint-phase-id-drift.cjs scans for the
literal token without caring whether it sits in code or in prose.

That is the guard being right, not over-eager — a quoted pattern is one paste
away from being live again, which is exactly how the copy it replaced spread.
Described in prose instead.

Worth recording: this guard is check:phase-id-drift, which lint:ci does not
run — it is enforced by tests/phase-id-drift-guard.test.cjs. A green lint:ci
is therefore not evidence the drift guards pass.

Refs #3185

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3185): backfill changeset PR number (#3230)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 22:01:20 -04:00
Tom Boucher
86bebcefa2 refactor(#3216): bind milestone identity to the canonical locator (#3226)
* refactor(#3216): widen milestone-window guard to literal-## matchers

The guard keyed only on the `#{N,M}` quantifier plus a literal version or
phase-lookahead token. getMilestoneInfo hand-rolls its milestone-heading match
with a literal `^##`/`## ` and an interpolated ${escapedVer}, so it satisfied
neither token and the guard reported a clean zero on a file carrying live
re-derivations (#3171, #3197) — a zero it did not earn.

Widen token (a) to a literal 2-6 `#` run, admitted ONLY inside a heading-MATCHER
literal (a regex literal, or a string/template handed to new RegExp) so a
heading-BUILDING template is not mistaken for a re-derivation. Widen token (b)
with the grouped `v(\d+(?:\.\d+)+)` shape and an interpolated version
placeholder.

Ships BEFORE the consolidation per ADR-3180 s7.2: a guard widened afterwards
measures an already-cleaned surface. It is expected to be RED until the
consolidation lands.

* test(#3216): failing-first milestone-identity single-owner suite

63 tests across two files, from the matrix in .gsd/phase/. Section H of
milestone-window-single-owner.test.cjs covers the 21 input classes of the
design's behavior table plus its negative space; milestone-window-drift-guard
covers the widened tokens and proves the exemption is function-scoped, not
file-scoped.

Copy count is 3 found by the guard, not 1 per the epic (ADR-3180 Amendment 3's
standing rule, holding for the fourth consecutive phase): both getMilestoneInfo
sites plus cmdRoadmapAnalyze's milestone enumeration at roadmap.cts:454, which
carries the same #3171 truncation and #3197 phase-heading confusion.

Expected RED until the consolidation lands.

* refactor(#3216): bind milestone identity to the canonical locator

getMilestoneInfo hand-rolled two milestone-heading regexes inside the owner's
own file. Both were wrong, differently: the STATE-version site's ^## anchor is
level-blind so [^\n]* absorbs a third #, and the fallback site had no anchor at
all, so '## ' matched from the second # of '###'. Against
'### Phase 7: Close v3.3 gaps' the fallback returned {v3.3, gaps} (#3197). Both
captured names with [^\n(], truncating at a parenthetical (#3171).

Bind both to the canonical grammar. locateMilestoneHeadings becomes a
version-filtered view over one shared source, and a new version-agnostic
listMilestoneHeadings enumerates milestone headings for callers that need all
of them. getMilestoneInfo returns ScopedResult<MilestoneInfo|null>; the
{v1.0,'milestone'} default, which was output-identical to a real v1.0 project,
is deleted. The #2245 never-throws invariant is preserved.

Copy count: 3 found by the guard, not 1 per the epic. The third was
cmdRoadmapAnalyze's own milestone enumeration (roadmap.cts:454), carrying both
defects in the implementation the epic blessed.

buildStateFrontmatter and archivePhaseDirectories branch on scope: the first
writes null rather than a fabricated identity, the second falls through to its
dated-label fallback. A fabricated v3.3 passes ARCHIVE_VERSION_LABEL_RE, so it
would otherwise misfile phase history.

Also fixes an unsafe cast in init.cts that masked these type errors across five
call sites, which would have shipped undefined milestone fields under green tsc.

* fix(#3216): restore the #1761 unbounded guard and bullet precedence

Review and the first full-matrix run surfaced five real defects in the
consolidation, all fixed here rather than by relaxing the tests that caught
them:

- buildStateFrontmatter gated its isMilestoneBoundedInRoadmap check on the
  scope-gated milestone value, which is null on any non-COMPLETE scope, so the
  #1761 unbounded guard was silently skipped and state json reported a percent
  it must omit. It now gates on the STATE-asserted version, independent of
  identity scope.
- The rewrite lost #2135's precedence: the name-bearing progress-marker bullet
  is consulted before the heading again.
- A single-segment version (v3, no dot) did not resolve; the name-extraction
  fallback now accepts it.
- A version carrying regex metacharacters, or a $& / $1 replacement pattern,
  is matched literally.
- listMilestoneHeadings' heading field trimmed, so a CRLF roadmap no longer
  leaks a trailing carriage return into roadmap analyze's output.

Also emits milestone_version / milestone_name / current_milestone as explicit
null rather than omitting the key, so the prompt layer cannot render a bare
placeholder, and corrects an init.cts comment plus a cast left inconsistent.

* test(#3216): update milestone-identity expectations to the scoped contract

getMilestoneInfo returns ScopedResult<MilestoneInfo|null> and the
{v1.0,'milestone'} default is deleted, so the suites asserting the old shape
assert removed behavior. Updated rather than weakened: every touched call site
now asserts the scope explicitly against the frozen SCOPE enum.

roadmap-parser.test.cjs: 20 expectations moved to {value,scope}. The #1881
unreadable-vs-absent diagnostic assertions are untouched and still prove their
original point — only the return shape moved. One pre-existing assert.ok(info)
is now a specific UNSCOPED assertion, so that case is stronger than before.

new-milestone-clear-phases.test.cjs: the test asserting phases clear archives
under the v1.0 default now asserts the dated archived-<YYYYMMDD> fallback,
which is the deliberate consequence of deleting that default.

Two of this branch's own tests were also corrected after they drove the
implementation the wrong way: the parity test compared raw heading text and so
pushed a stray ## prefix into roadmap analyze's public output, and the hostile
metacharacter row demanded a pathological version resolve, which pushed a
widening of the ADR-locked \b boundary. Both now assert what the contract
actually requires.

* docs(#3216): document milestone identity and correct the CONTEXT.md entry

ADR-3180 s7.2 moves to Enforced and gains two rules that were unstated: the
name derives from the heading's own version token and drops a trailing status
marker, and a free-form legacy ROADMAP with no version anywhere is UNSCOPED
with no identity rather than a defaulted v1.0 (decided by the maintainer before
implementation, per s7's own rule that an unstated behavior is not decided).
Amendment 4 records Phase 6's validation, including that the copy count was a
lower bound for the fourth consecutive phase.

CONTEXT.md's Roadmap Parser entry described locateMilestoneHeadings as
boundary-matched with (?![\w.-]) — the alternative Amendment 2 tried and
REVERTED. The code uses \b and says so, and the ADR agrees; the revert updated
code and ADR and missed CONTEXT.md, which is the epic's own fixed-on-one-copy
failure class in the docs layer, on a file that is itself a PR gate.

* fix(#3216): persist the real version on a truncated identity

buildStateFrontmatter wrote null for BOTH milestone and milestone_name on any
non-COMPLETE scope, discarding a real version. ADR-3180 s7.2 rule 6: a version
known with no resolvable name is TRUNCATED carrying {version, name: null} —
'the version is a real answer, the name is a non-answer, and collapsing the two
is the failure this contract exists to prevent.'

The two fields are now gated by what is actually known: the version whenever one
exists (COMPLETE or TRUNCATED), the name only on COMPLETE. Never fabricated.

Caught by this phase's own Decision 4(c) consumer-output test, which is the
argument for asserting at the consumer rather than the owner — the owner was
correct throughout; only the consumer collapsed its answer.

* refactor(#3216): extract helpers and make cmdCommit's scope gate explicit

From the two-axis code review:

- init.cts repeated the identical getMilestoneInfo cast at five sites with
  copy-pasted comments — duplication inside a PR whose thesis is that duplicates
  get deleted. Extracted milestoneRecord(cwd); the one site-specific comment is
  kept, the four generic copies removed.
- getMilestoneInfo hand-built its { value, scope } literal at ten return points;
  a local scoped() constructor now does it once. Every per-branch rationale
  comment is preserved and no returned value or scope changed.
- cmdCommit gated the milestone branch name on plain truthiness, which is also
  true for TRUNCATED, so an unresolved identity drove branch creation
  incidentally rather than deliberately. It now gates on the SCOPE enum,
  accepting COMPLETE or TRUNCATED because both carry a real version, and the
  comment records why that differs from archivePhaseDirectories — which demands
  COMPLETE because it uses the value as a filesystem path component.

* test(#3216): cover the bare-version-in-prose truncated path

The spec review found the bareVersionMatch path — no STATE version, no
milestone heading, a version token only in prose — returning TRUNCATED with no
test exercising that exact shape, violating Decision 4's boundary-coverage
requirement.

* docs(#3216): record the missed Tier-2 surfaces and rule 5's corollary

Decision 3 requires an explicit call-out for EVERY Tier-2 change, and Amendment
4's first draft named eight surfaces while the change touched thirteen. Adds
cmdCommit's branch-name construction and the four init JSON bundles, an
incomplete list being the same defect in miniature that this epic removes.

s7.2 rule 5 gains a corollary separating two cases the original wording ran
together: no version token ANYWHERE is UNSCOPED, while a bare version token in
prose or a non-milestone heading is weak but real evidence and yields TRUNCATED
under rule 6.

* chore(#3216): set changeset fragment pr to 3226

---------

Co-authored-by: sim <sim@local>
2026-08-08 19:06:13 -04:00
Tom Boucher
b9f51836e6 refactor(#3180): ADR-3180 behavior contract + cross-surface drift guardrails (#3223)
* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract

The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a
lower bound for the third consecutive time, and that two derivation families
had never been named at all.

ADR-3180 gains Decision 7 — a normative behavior contract that says what the
right answer IS for each derivation, not merely who owns it. A reviewer with
no written rule can only ask "does this look like the others", which is how a
fifth copy passes review. Decision 4 gains (d) scan surface is every authored
surface and an owner FILE is never exempt, only its named functions; and (e)
a surface that cannot be consolidated today ships ratcheted, never unguarded.

Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined
copies of its own body across five modules. All six now route through it;
`clampPercentFromFraction` is added for the one caller that already held a
fraction. Every migration is behaviour-identical — clampPercent's first line IS
the `total > 0 ? … : 0` ternary each copy carried. Guarded by
lint-completion-ratio-drift.cjs, which reports zero re-derivations with no
file-level exemption.

Prompt layer: workflow markdown re-derives live-plan counting in raw shell
(#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs
scans it with a shrink-only baseline of the 7 sites that exist today — new
sites fail, and a baseline entry that stops firing fails too, so an
acknowledgment can never outlive the thing it describes.

lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only
the four named canonical functions are exempt now. The blanket exemption was
pointed at the one file most likely to grow the next copy, and it had.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage

Five findings from the two orthogonal review passes, all fixed.

Decision 4(c) breach: the completion-ratio identity test asserted at the
OWNER, which is exactly the bypass that decision exists to close — a consumer
can call clampPercent and then post-process locally, leaving both the lint and
an owner-level test green. It now drives `roadmap analyze`, `query progress`
and `stats` and asserts on their own output, over a fixture containing a
`status: superseded` plan so a consumer that re-counted raw files would report
60 where the owner reports 75.

Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the
issue that removes them. They name Phase 8 (#3218) now.

The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical
sites were one indistinguishable key and migrating either would have left the
guard green with the other alive. Entries carry an occurrence count; fewer than
acknowledged fails as a partial migration, more fails as a new copy.

Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's
test already had, and the fast-check property tests CONTRIBUTING requires for
clamp/budget-limit functions.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes)

`tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs`
under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`.
A fixed wall-clock budget around a double spawn, running inside a container
that is concurrently executing the full ~31k-test suite, fails by construction
under load.

Confirmed against three full matrix runs. Every failure was shaped
`null !== 0` — the child was KILLED, never an assertion about the thing under
test. One captured probe had already printed the correct resolution
(`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It
reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The
victim subset varies by run and by lane.

What these tests are actually about is suite-token RESOLUTION — `unit` as a
bare token in --files/--files-from. Executing the seeded trivial files is
incidental and is the entire timeout surface, so the assertions move
in-process against the same functions `main()` calls, in the same order.
`parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are
exported for that; no behavior, signature or logic changed.

No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the
harness for real and asserts exit codes end to end, on a 120s budget.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: delete the three elapsed-time assertions

CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all
three are load-sensitive: on a saturated bench each can fail while the code
under test is correct. In every case the load-bearing assertion sits on the
line above and the timing line adds no discrimination.

run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s
harness backstop?" — is already answered by the assertion above it. A backstop
kills by signal, which surfaces as status null, never 124. Observed directly
this session: three matrix runs produced exactly that null shape from killed
children.

normalize-test-command and context-predicates: both bounded a ReDoS check.
A threshold only ever separates "fast" from "slightly slow", which is bench
load, not correctness — catastrophic backtracking on 800 KB of input does not
take 251ms, it does not finish at all. A real regression therefore shows up as
the suite being killed on that test, which is louder and more reliable than a
number. The structural assertions (returned unchanged; cleanly rejected) are
what actually carry those tests, and they stay.

The sweep now reports zero elapsed-time assertions in tests/. The remaining
Date.now() uses are unique-path suffixes, barrier deadlines, fixture
timestamps and fake mtimes — none of them assertions.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3180): backfill changeset PR number (#3223)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows

The baseline keys on (file, trimmed text). `file` came from scanTree's
`path.relative()`, which uses NATIVE separators, while the committed baseline
stores POSIX. On Windows every violation was therefore unmatched — reported as
FRESH — and every baseline entry matched nothing — reported as STALE. The guard
failed 100% of the time there, on both CI shards:

  ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale
    + { file: 'gsd-core\\workflows\\execute-plan.md', ... }

The remote runner this repo gates on is Linux-only and cannot see this class at
all; the GitHub Actions Windows lane is what caught it.

Normalization is unconditional — never gated on process.platform. A
platform-conditional normalizer makes the POSIX path the special case and
leaves the Windows branch unexercised on every other OS, which is the same
blind spot in a different place. It is applied at one seam inside
findPromptDrift, which builds `file` on every returned violation, so the
baseline key, the --update writer, the stderr report and the tests all consume
one normalized value.

The regression tests drive a Windows-shaped relPath directly and run on every
OS rather than skipping off-Windows — a test that only runs on the platform
where the bug lives is why this escaped. They include a sanity check that
un-normalized input does NOT match, so the assertion cannot pass vacuously.

Audited the three sibling guards: none keys against a committed cross-platform
baseline, and their exemption keys are path.join-built, so producer and
consumer share the native convention. Left correct code alone rather than
making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing
there would break those three on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 16:05:17 -04:00
Tom Boucher
636ec92107 refactor(#3185): phase enumeration has one owner and a decidable scope (#3222)
* test(#3185): failing-first phase-enumeration single-owner suite

Covers the enumeration rows with direct code evidence: 999.* backlog dirs
listed by progress/stats, the phase-0 sentinel divergence, the #1324
letter-prefixed-decimal negative space, and the destructive-path find —
cmdPhasesClear carries a fifth sentinel copy (/^999(?:\.|$)/) that excludes
999 but not 0, so a 0-* directory roadmap.analyze preserves is deleted there.

Also covers the pass-all degrade, which is where the defect actually lives:
when the milestone window declares no phases the filter becomes a literal
() => true and its heading-side sentinel exclusion is unreachable. A fixture
carrying phase headings keeps the filter active and never reaches that path.

Named for the derivation, not a module: the suite drives commands, phase,
milestone, workstream-inventory and state, and both the phase and
phase-locator buckets are already at the per-module test-file cap.

Committed alone so the remote runner records the failure before the fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): phase enumeration has one owner and a decidable scope

Adds phase-locator.cts::listMilestonePhaseDirs as the single canonical owner
of "which phase directories belong to the current milestone". It applies the
milestone window AND the sentinel filter and returns a ScopedResult, so a
caller can tell a genuinely-empty milestone from an enumeration that could
not be scoped.

The sentinel test now runs against DIRECTORY NAMES and is unconditional.
getMilestonePhaseFilter excludes sentinels from its ROADMAP heading set, but
degrades to a literal () => true pass-all predicate when that set is empty --
at which point the heading set is never consulted and its sentinel exclusion
is unreachable exactly when it is needed. That degrade is the #3167 path, and
it is why stats already used the filter and still listed backlog directories.
The narrowing is sentinel-only: pass-all stays over-inclusive otherwise.

Sentinel copies deleted, canonical isSentinelPhaseId adopted:
  - cmdRoadmapAnalyze's local closure (parseInt === 0 || === 999), 2 call sites
  - cmdPhasesClear's /^999(?:\.|$)/ -- the DESTRUCTIVE path, which excluded
    999 but not 0, so a 0-* directory roadmap.analyze preserves was deleted

cmdStats also seeded rows from ROADMAP headings with no sentinel filter, so a
999 heading produced a row with no directory; that seed is filtered now.

cmdPhasesList routes only its ENUMERATION. --phase lookup searches the
physical set (scoping it would report an out-of-window phase as not found) and
--include-archived still merges archived dirs (they are by definition from
other milestones). Both exempt by documented reason, never a file allowlist.

Fixed inline, found while building: isDirInMilestone could not match a #1324
letter-prefixed-decimal directory (P0.0-foundation) to its own Phase P0.0
heading, so stats reported the phase with plans: 0 while its directory held
plan files. Defers to phase-id's extractPhaseToken rather than widening a
fourth bespoke regex; additive, so it can only admit directories.

Refs #3180. Closes #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): route the last two enumeration re-derivations

workstream-inventory countRoadmapPhases counted every `Phase` heading across
the whole ROADMAP -- no window, no sentinel filter -- so it counted 999.*
backlog and Phase 0 and spanned every milestone the document ever had. Its own
caller already resolved a currentVersion and passed it to getMilestonePhaseFilter
elsewhere in the same file; this was the sibling copy that never got the fix.

state.cts phaseInventoryProvider enumerated phase dirs with its own
/^(\d+)-(.+)$/ convention regex and neither filter, so a rebuilt STATE.md
inventory carried backlog and sentinel directories as current-milestone phases.
A non-COMPLETE enumeration scope now throws to the outer catch as a real scan
failure rather than reporting a confident undercount, mirroring the per-phase
scanPhasePlans contract beside it.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): consolidate 23 sentinel re-derivations onto one predicate

The whole-repo drift guard (ADR-3180 Decision 4a, no file allowlist) found the
sentinel rule re-implemented 23 times across 8 modules, in three regex variants
plus four integer-comparison forms. Most tested 999 only, so Phase 0 slipped
through them while roadmap.analyze and the engine-wide convention (#1580) both
treat 0 and 999 alike. That disagreement is the defect class this epic removes.

All 23 now call phase-id's isSentinelPhaseId (SENTINEL_RANGES [0,999]). Sites:
init recommended-actions and backlog counts, milestone phase scan, the
phase-lifecycle progress table, phase.cts used-number collection and the four
renumber-on-remove guards, roadmap-parser's heading and bullet milestone
counts, roadmap get-phase fallbacks, and state's heading denominator.

Excluding Phase 0 at these sites is a deliberate behavior change and the point
of the consolidation — several carried comments already saying 0 should be
excluded while the literal beside them caught only 999.

Adds scripts/lint-phase-enumeration-drift.cjs, wired into lint:ci. It scans the
whole src/ tree with no file allowlist and reports both shapes: an independent
phases-dir enumeration, and an independent sentinel literal. Exemptions are
function-scoped with a written reason. The guard is comment-aware — its first
pass flagged JSDoc and a comment documenting that the code below uses the
canonical owner, which would have trained readers to exempt prose.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): resolve every phases-dir enumeration; drift guard reports zero

Per-site triage of the 31 remaining whole-repo guard hits, applying the rule
generalized from #3183's Amendment 1: a LOOKUP, DIAGNOSTIC, ARCHIVAL or
MUTATION pass wants the physical set; only "which phases belong to this
milestone" wants the scoped set.

Routed (10): init new-milestone phase_dir_count, init milestone-op fallback
count, init manager, init progress, milestone complete stats/dry-run/archive
move, phase complete's next-phase scan, state update-progress, state
frontmatter stats, and uat audit's active set.

Exempt with a written function-scoped reason (never a file allowlist): the
audit/UAT/verification sweeps that deliberately scan every directory to report
gaps, phase create/insert/rename/renumber mutations, single-phase lookups,
roadmap-upgrade's cross-milestone migration, cmdPhasesClear's whole-tree
destructive pass, and the reads that list a phase dir's FILES rather than
enumerating the phases dir at all.

Latent defects fixed by the routing: sentinel directories leaked into
cmdInitNewMilestone's phase_dir_count, cmdMilestoneComplete's stats, dry-run
AND ARCHIVE MOVE, cmdStateUpdateProgress, buildStateFrontmatter and
cmdAuditUat's active set — every one of those hand-rolled an isDirInMilestone
filter with no sentinel exclusion, so `milestone complete` was archiving
backlog directories.

scripts/lint-phase-enumeration-drift.cjs now reports 0 re-derivations and
npm run lint:ci is green.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* docs(#3185): document milestone-scoped enumeration and record ADR Amendment 3

Changeset fragment (Changed), CLI-TOOLS/COMMANDS/USER-GUIDE updates for the
scoped output of progress, stats, phases list, phases clear and milestone
complete, the CONTEXT.md Phase Locator glossary entry naming
listMilestonePhaseDirs, and ADR-3180 Amendment 3.

Amendment 3 records: the SCOPE contract held unchanged; the declared deviation
from Decision 1's provisional signature (the window needs cwd/ws, which the
locked roadmapContent parameter cannot supply); the copy count being a lower
bound for the third consecutive phase (4 scoped vs 54 found); the load-bearing
finding that the sentinel exclusion sat on the heading set and was unreachable
under the pass-all degrade; the two destructive-path defects; and the
generalized exemption rule.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): wire scope to consumers; revert two wrong routings the suite caught

Review + remote runner findings, all fixed:

The three consumers computed the enumeration scope and threw it away, so
TRUNCATED/UNSCOPED/UNREADABLE collapsed into the same output as COMPLETE --
reproducing this epic's own output-identical-failure defect one layer up.
progress, stats and phases list now emit phase_scope (null on the phases list
--phase lookup path, which performs no enumeration).

Two routings were wrong and the suite proved it:

roadmap-parser's two milestone phase-count scans are reverted to the 999-only
literal. isSentinelPhaseId is BROADER than what it replaced: its legacy branch
runs /^0*(\d+)/ over "00.1", which backtracks to capture 0, so it read #2554's
decimal phase ids as sentinel milestone 0 and stopped counting them.

state.cts phaseInventoryProvider is reverted to the physical disk scan.
`state rebuild` is a RECONCILIATION pass -- scoping it made it throw on healthy
trees whose fixture resolves no window, swallowed the raw readdirSync fault
message #3057 B1 requires verbatim, and stopped it dropping orphan STATE.md
rows, which is the job.

Both are now function-scoped guard exemptions with written reasons, not
silent reverts. This is the consolidation trap named in the epic: a canonical
rule can cover MORE than the copy it replaces, and only real inputs show it.

Adds phases list coverage, a scope-branch test, and a drift-guard unit suite;
backports comment-awareness to the milestone-window and plan-count guards so
all three siblings share one false-positive profile; names #3161 alongside
#3167 in Amendment 3's subsumption record.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): correct isSentinelPhaseId's decimal-zero misclassification

An isolated security review caught this branch committing the epic's own sin:
the over-broad predicate was worked around at ONE call site and left live at
the destructive ones.

isSentinelPhaseId's legacy branch ran /^0*(\d+)/, which backtracks so any id
whose leading digit run is all zeros before a non-digit captures 0 -- "0.1",
"00.1" and "0.2554" all read as sentinel milestone 0. Two pinned contracts
disagree with that: #2554 requires "00.1" to be counted as a real phase, and
the 999 icebox is a whole reserved milestone so "999.1" must stay sentinel.

The rule is asymmetric and now says so explicitly: 999 is sentinel with or
without a decimal part; 0 is sentinel only when bare. A decimal phase under
either is a real phase for 0 and reserved for 999, because 999 reserves a
MILESTONE while 0 reserves a PHASE.

Fixing the owner lets the earlier workaround go: getMilestonePhaseFilter's two
scans route through isSentinelPhaseId again and the guard exemption that
existed only to accommodate the defect is deleted. The state.cts cmdStateRebuild
exemption stays -- that one is a genuine reconciliation-wants-the-physical-set
case.

Also corrects tests/adr-612-bracket-grammar.test.cjs, which asserted
isSentinelPhaseId('0.1') === true and so had encoded the defect as expected
behavior.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): keep isSentinelPhaseId's semantics — 0.x is layered, not wrong

Reverts the previous commit. The remote suite failed six tests proving it
wrong, and the reason is the sharpest finding of this phase.

An isolated security review observed that isSentinelPhaseId reads 0.1 and 00.1
as sentinel milestone 0 and judged that a defect against #2554. Correcting the
canonical predicate broke #2949. Both contracts are pinned and both are right,
because they ask different questions:

  #2554  is this dir part of the current milestone's phase SET?  -> count 00.1
  #2949  must this phase COMPLETE before the milestone closes?   -> 0.x sentinel

No single global predicate answers both. isSentinelPhaseId keeps its semantics
(0.x IS a sentinel, #2949), and the milestone-window layer keeps a narrower
999-only rule (#2554) as a function-scoped guard exemption with a written
reason — not a second silent copy.

That corrects how Decision 1 reads: "one owner per derivation" governs who
computes an answer, not how many questions share it. An over-broad canonical
rule is as much a defect as a divergent copy and fails worse, because it looks
like consolidation. Recorded in Amendment 3 as the lesson for Phases 4 and 5.

Where a review's inference about intent conflicts with a pinned contract, the
pinned contract wins; the finding is adjudicated, not fixed.

The boundary tables in the enumeration suite are corrected to assert 0.x IS a
sentinel, with the layering explained.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* chore(#3185): set changeset fragment pr to 3222

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 14:22:10 -04:00
Tom Boucher
bc3a3f170f docs(#3215): add ADR-3212 lexical seam consolidation (Phase 0) (#3221)
Design-lock ADR for epic #3212 — the lexical layer beneath the #1372
markdown-sectionizer and #2143 table/mutation seams.

Locks seven decisions Phases 1-4 execute against:
- src/pattern.cts as sole owner of dynamic regex construction,
  delegating to the built-in RegExp.escape; the ten private
  escapeRegex/escapeRegExp/escapeRe copies are deleted, not merged
- engines.node floor raised to the Active LTS line (>=24), which is
  what makes RegExp.escape reachable; delegating-shim alternative
  recorded and rejected
- src/text-lines.cts as sole owner of line-terminator handling —
  the primitive the existing no-crlf-fragile-split prohibition lacks;
  brings frontmatter.cts in from the #1372 exclusion
- tokenizer-first for stateful grammars, with a decidable five-condition
  test, generalizing the proven hooks/lib/git-cmd.js token-walk (#3129)
- bounded quantifiers over caller-supplied content
- extend-never-mutate (inherited from ADR-2143 §2)
- prohibition with teeth: no-adhoc-regex-escape, no-unbounded-quantifier,
  no-crlf-fragile-split widened to src/, plus a parity assertion

Explicit non-goal: no wholesale regex-to-parser rewrite. A census found
2,113 regex literals across 317 files; most are correct and stay.

Docs-only. No production code.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 13:29:45 -04:00
Tom Boucher
342590c70e refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite

Covers the 50 input classes in the phase test matrix: scope classification
(genuinely-empty vs truncated vs unscoped vs unreadable), the section-end
owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c),
the milestone.complete refusal with negative proof that no directory moved,
the version-token boundary defect, drift-guard behavior, and three fast-check
properties over document-shaped generators.

Committed alone so the remote runner records the failure before the fix lands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* refactor(#3184): milestone windowing routes through one owner

Three copies of the milestone section-end walk lived in roadmap-parser.cts —
two distinct computeSectionEnd function nodes plus an inline third in
getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is
now the sole owner and the other two are deleted, not kept in sync by comment.

The whole-repo drift guard found what the epic did not: state.cts held three
more re-derivations of the same vocabulary — two byte-identical milestone
bounding checks carrying a defect neither reported copy has (no boundary after
the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning
predicate. All three route through the owner now.

A composition-level duplicate appeared inside this change's own first pass:
getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out
of the owner's primitives, and had already diverged on whether to skip a closed
milestone heading. sliceMilestoneWindow is the one composition.

Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is
distinguishable from a genuinely empty milestone — those were output-identical,
which is the whole failure class. roadmap analyze emits it (#3165), and
milestone complete refuses to archive on anything but COMPLETE rather than
pass-all moving every phase directory on disk (#3166). The pass-all degrade is
preserved where its premise holds: making the filter deny-all would trade a
silent over-inclusive answer for a silent under-inclusive one on the read paths
that count with it.

extractCurrentMilestone keeps its signature — 200+ affected symbols across 41
files and 25 process flows — and is a one-line wrapper over the scoped owner.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): fence-aware phase detection and one heading-selection owner

Review fixes from the two orthogonal passes.

The blocker: hasPhaseEntries matched ATX phase headings fence-aware via
tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown,
so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely
empty milestone then classified TRUNCATED and milestone complete refused a
legitimate archive — a false positive in the destructive direction, worse than
the defect this phase set out to fix. Both that path and getMilestonePhaseFilter
own pre-existing bullet scan now run on stripFencedCode, since leaving one meant
the owner file gave two different answers to the same question.

The selection rule — locate, prefer the non-closed heading, else the first — had
been written three more times inside the file whose thesis is single ownership.
selectMilestoneHeading owns it; all three sites route through it. The copies were
behaviorally identical, so this is de-duplication with no observable change,
verified by probing that all three paths select the same heading.

roadmap analyze emitting a scope no consumer read left #3165's actual symptom
alive, so Route 0 in next.md now treats a non-complete scope as scan-failed
rather than as a clean empty scan, and the ADR amendment no longer overstates
what shipped.

Also: the scope refusal moved above the archive-directory create, so a refusal
leaves nothing on disk; the versionOverride comment names all four consumers;
COMMANDS.md documents the new guard beside its sibling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#2658): exclude the changelog from the malformed-path scan

The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts
none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships
into that tree, and its #2658 entry quotes both malformed paths while describing
the fix that removed them — so the release note documenting the fix trips the
fix's own regression test. Red on next before this branch.

The installer is correct: a probe over a real --trae --local install found 621
emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope
was the defect, not the product.

Excluded by exact relative path rather than by loosening the patterns or skipping
all markdown — the emitted agent and command markdown is precisely what #2658 was
about, so the gate stays strong everywhere it matters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#3184): regenerate install-tree fixtures for the shared drift scanner

scripts/lib/ ships in the npm package and installer, so extracting the shared
tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's
install tree. Regenerated via npm run gen:install-tree; the delta is exactly
that one path per fixture.

The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded),
so only the extracted library moves. This matches the existing
scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only
helper carried in the shipped tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal

The remote runner caught two regressions this branch introduced. Both were mine,
and neither review pass found them — only running the existing suite did.

The version-token boundary. I replaced locateMilestoneHeadings' \b with
(?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562
fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone
state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a
closed v8.0-A sibling (#730), and \b is what allows it while the stricter
boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to
\b; the state.cts consolidation is now a straight merge with no behavior change,
and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and
the design doc no longer claim otherwise.

The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is
about the TRUNCATED window specifically — the heading is found and the section
closes before the phase region, so pass-all archives everything. UNREADABLE and
UNSCOPED are pre-existing, legitimately handled states, and refusing on them
broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed
to TRUNCATED; docs corrected to match.

One of the new tests was also wrong: its fixture gave the shipped and current
milestones' phases the same numeric id, and the filter matches on that id, so it
could not have distinguished the two windows. Fixture corrected to exercise what
it claims to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): enumerate drift-scan.cjs for uninstall

The installer copies scripts/lib/ wholesale, but uninstall removes an explicit
set — deliberately, so a user's own helpers in that directory survive. The
extracted drift-scan.cjs was copied in and never enumerated, so it outlived
uninstall, left the directory non-empty, and the rmdir that follows failed.

Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is
likewise a lint-only helper that ships there and is enumerated. Verified with a
real install-then-uninstall into a temp target: scripts/lib/ held exactly the
three GSD files and was gone afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset

Found while shipping this phase, and fixed here rather than noted.

install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE —
the comment at the copy site literally says "and any future lib helpers".
uninstall() removes them by hardcoded enumeration, deliberately, so a user's own
helpers in those directories survive. A wholesale writer paired with an
enumerated remover cannot stay in sync by construction: any file added to either
directory ships to every user and is then orphaned in their repo forever, since
it survives uninstall, leaves the directory non-empty, and the rmdir that follows
fails. Nothing reported this. 31,225 tests were green over it.

That is the same divergence class this epic exists to delete, sitting in the
installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity
assertion that fails the moment the two surfaces disagree. The test compares each
directory's real contents against its enumeration and names the offending file
plus the constant to add it to.

Both enumerations are hoisted to module scope and exported, so the test asserts
on the actual arrays rather than pattern-matching the installer's source — no
allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on
the current tree, correct report when an unenumerated file is injected.

scripts/changeset/ turned out to carry the identical defect and is covered too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* chore(#3184): backfill changeset PR number

Also narrows the wording to match the shipped behavior: the refusal fires on a
truncated window specifically, not on any non-complete scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 09:50:35 -04:00
Tom Boucher
343835facc refactor(#3183): route live-plan counting through scanPhasePlans (#3199)
* refactor(#3183): route live-plan counting through scanPhasePlans

scanPhasePlans becomes the sole owner of the live-plan derivation. Twenty-one
independent re-derivations across seven modules now route through it, and
scripts/lint-plan-count-drift.cjs reports zero, scanning the whole repo rather
than an allowlist (ADR-3180 Decision 4a).

The epic scoped this at three copies. A whole-repo guard found twenty-six sites
across nine files, so Phase 1 absorbs every live-plan re-derivation and Phase 3
narrows to window plus sentinel enumeration.

Two sites are exempt with a documented reason rather than a bare allowlist:
audit.cts scans one quick task's own directory for a single completion record,
and gsd2-import.cts reads a foreign GSD-2 tasks/ layout during a one-time
import. Neither is a phase directory.

scanPhasePlans gains allPlanFiles (pre-supersession) alongside planFiles so one
owner answers both questions: verify.cts's numbering-gap check wants every plan
on disk, its pairing check wants the live set. Both fields are additive.

Highest-severity fix: cmdPhasePlanIndex, which feeds execute-phase wave
scheduling, was scheduling status:superseded plans into waves and reporting zero
plans for the post-#3139 nested layout.

filterPlanFiles and filterSummaryFiles are deleted; getPhaseFileStats orphaned
them and only their own tests still called them.

New leaf module src/planning-scope.cts carries the frozen SCOPE discriminator,
with its six-gate ripple closed: gitignore, inventory manifest, INVENTORY.md and
the CONTEXT.md glossary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* docs(#3183): amend ADR-3180 for the Phase 1/3 boundary re-slice

The contract held; the phase boundary did not. The whole-repo drift guard found
26 re-derivations across 9 files against the epic's estimate of 3, and
cmdProgressRender re-derives both enumeration and plan counting on adjacent
lines, so DW4 was unsatisfiable within Phase 1's original file scope.

Records the amended scope, scanPhasePlans's new allPlanFiles field,
findOrphanSummaries, the two documented exemptions, the re-derived Tier-2
table, and the describeNonCanonicalPlans trap for later phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): complete the canonical pairing rule and gate the naming diagnostic

The remote runner went red with 13 deterministic failures on both lanes,
and they were right: replacing verify.cts's canonicalPlanStem pairing with
summaryCandidates dropped a case the bespoke rule covered. A plan carrying a
descriptive slug after its id (68-01-scaffolding-PLAN.md) pairs with its
canonical-stem summary (68-01-SUMMARY.md), and summaryCandidates generated no
such candidate, so the plan read unsummarized.

The fix is to complete the one rule rather than restore a second:
summaryCandidates gains a canonical-id candidate, narrowed to fire only when an
id pair was actually extracted. countMatchedSummaries, findUnsummarizedPlans
and findOrphanSummaries all inherit it. The two-plans-one-summary collision
behaviour of the original rule is preserved deliberately and documented in
place.

Second defect, independently root-caused while verifying: routing the #2893
naming diagnostic through scanPhasePlans exposed it to the loose /PLAN/i
fallback, which is correct for counting and wrong for a naming check — a
non-canonically-named file was accepted as a valid plan and the diagnostic
went silent. cmdPhasesList, cmdFindPhase and cmdPhasePlanIndex now intersect
with a strict isCanonicalPlanFile predicate before reporting names.

Same class as the describeNonCanonicalPlans trap already recorded in ADR-3180:
a question about file naming wants the physical, strictly-matched set; only a
question about outstanding work wants the live set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): register planning-scope.cjs in the eslint migration list

tests/repo-invariants.test.cjs asserts every bin/lib/*.cjs is linted xor
ignored per its ADR-457 migration state. The new planning-scope module closed
five of the six .cts ripple gates - gitignore, inventory manifest, INVENTORY.md
and the CONTEXT.md glossary - but not eslint, because that one is enforced by a
test rather than by lint:ci, so the local pipeline stayed green while it was
missing.

Generated from src/planning-scope.cts, so the .cjs is ignored and the .cts is
linted, matching every other migrated module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): replace the plan-count drift detector with a literal tokenizer

CodeQL reported 4 high-severity js/redos alerts on REGEX_LITERAL_MD_RE, the
backtracking regex that finds "a regex literal mentioning PLAN/SUMMARY and an
escaped \.md". Five review rounds found it had two defects, not one:

  - EXPONENTIAL, then CUBIC. Its "any char" atom `(?:\\.|[^/\r\n])` let a `\.`
    pair be consumed either as one escape or as two class characters, which is
    exponential backtracking: 27,464ms on `"/\.mdplan" + "\.".repeat(28) + "X"`.
    Excluding `\` from the class killed that but left a cubic path — 23ms at
    N=200, 172ms at N=400, 1362ms at N=800 on `"/" + "PLAN\.md".repeat(N)` with
    no closing `/`. This guard is the last stage of `npm run lint:ci`, which CI
    runs on fork pull requests, so a crafted src/*.cts could stall the job.
  - A DETECTION HOLE. A character class holding a bare, unescaped `/` — e.g.
    `/SUMMARY[^/]*\.md$/`, an ordinary path-excluding filter — terminated the
    literal at that `/`, so the scan never reached `\.md` and the guard missed
    it entirely. (Classes holding an ESCAPED `\/` were already matched; the
    tests cover those separately as parity, not as regressions.)

Both defects have one root cause: regex-literal grammar — `\x` escapes, and
`/` inside `[...]` not terminating — is not expressible in a backtracking
regex. So the detector is now a tokenizer, not a regex.

readRegexLiteralAt reads the literal at a given `/` in a single left-to-right
pass with no backtracking, treating escapes as two-character units and
suppressing the `/` terminator inside a character class. findRegexLiteralMdMatch
restarts it at every `/` on the line, preserving the old "find anywhere"
behaviour; MAX_REGEX_LITERAL_LEN (400) bounds each read — including the
trailing-flag scan — which keeps the whole-line cost linear.

Results: cubic shape flat at 0.06-0.39ms out to N=3200 (25KB), exponential
shape 0.01ms at 28 reps and 0.00ms at 64, and the bare-`/` class shapes are now
caught. Differential against the old regex over 28,474 lines (those matching
FILENAME_TEST_RE but not PLAN_SUMMARY_LITERAL_RE, across src/tests/scripts/
gsd-core/bin/eslint-rules, excluding 265 lines with >6 backslashes on which the
old regex hangs): 6 differences, all the tokenizer returning the fuller or
newly-correct literal, 0 old-only misses. The `\.md` token stays
case-insensitive, matching the `/i` the old regex carried.

Also closes three holes in the same new file:

  - walk() tested entry.isFile(), false for a symlink, so a symlinked
    src/*.cts was silently unscanned — an evasion of a guard whose stated
    principle (ADR-3180 Decision 4a) is whole-repo discovery with no allowlist.
    It now resolves symlinks, but confined: file links must resolve inside the
    repo root, directory links inside the scanned dir itself. Every sibling
    drift guard in scripts/ uses the Dirent classification and never follows
    links, so following them unconfined would have made this the only linter
    able to read outside the tree — on fork PRs an arbitrary out-of-repo read
    whose matched fragments reach a public CI log. The narrower directory rule
    additionally stops `src/up -> ..` from sweeping the whole repo, and the
    skip list is now checked against resolved paths so `src/g -> ../.git`
    cannot reach .git/** or node_modules/**. Real paths are de-duplicated and
    files reported canonically, so a symlink alias cannot shift which
    FUNCTION_SCOPED_EXEMPTIONS key applies.
  - Both the reported fragment and the reported FILE PATH are attacker-
    controlled source text written straight to a CI log, and git permits
    control bytes in a filename. Both are now escaped — C0/C1/DEL plus the
    bidi and zero-width controls — so a crafted literal or filename cannot
    recolour the log, overwrite a line with CR, or fabricate a line that looks
    like this guard's own success output.

Regression coverage in tests/plan-count-single-owner.test.cjs: a child-process
probe over both pathological shapes (catastrophic backtracking is synchronous
and would freeze the suite rather than fail one test), the bare-`/` class
shapes verified to fail against the parent-commit blob, root-confinement tests
covering the outside-file, outside-directory, cycle, broken-link and duplicate
cases, direct isInsideRoot coverage including the sibling-prefix case that a
bare startsWith would let through, sanitizeForReport coverage, and
limit-1/limit/limit+1 coverage of MAX_REGEX_LITERAL_LEN derived from the
exported constant. The earlier structural assertion was dropped — it checked
for the substring `[^/`, which respelling the class as `[^\r\n/]` defeats
while staying exponential.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): backfill changeset PR number

Restores b77931869, which a force-push during the ReDoS remediation dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 01:20:55 -04:00
Tom Boucher
664d49e513 docs(#3182): ADR-3180 — planning semantic model single owner (#3196)
* docs(#3182): ADR-3180 — planning semantic model single owner

Phase 0 design lock for epic #3180. Names one canonical owner per
semantic derivation, specifies the frozen-enum scope contract that
distinguishes a genuinely-empty computation from a truncated or
unscoped one, and locks the drift-guard contract.

Ships no production code. Phases 1-5 execute against this ADR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* test(#3182): prune real issue 3182 from phantom-ref guard, fix empty-list regex

The guard's own header documents that its list rots: entries are phantom
only until the repo's shared issue/PR counter reaches them, and once the
counter passes an entry it must be deleted. Creating the Phase-0 sub-issue
advanced the counter past 3182, so the guard began rejecting a legitimate
citation of a real issue - the failure its header already records happening
twice, with PRs 2551 and 2361.

3182 was the last entry, and removing it exposed a latent bug: the regex
builder interpolated the list unconditionally, so an empty list yields
(?:#(?:)\b)|(?:issues/(?:)\b), whose empty alternation matches every issue
reference in the repo. Following the file's own maintenance instruction
would have turned a green guard into one failing on nearly every file.

buildRefRe() now returns null for an empty list and is exported, with
boundary coverage at 0/1/2 entries plus word-boundary and bare-digit
negative cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 20:25:12 -04:00
Tom Boucher
27aa40f65e fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir

pi reserves <configDir>/hooks as its deprecated extension location and warns
on every startup when it exists. Assert a pi install stages the shared hook
bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/.

Also adds pi to the local-scope dir table in install-shared.cjs: pi was in
RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved
path.join(root, undefined) and no local pi install could be exercised.

Fails before the fix. Verified via the remote runner.

* fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir

pi reserves <configDir>/hooks as its now-deprecated extension location and
warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs()
guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged
its shared hook bundle exactly there, and pi's advised remediation (move it to
extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's
extension auto-discovery.

The bundle directory name is now runtime-descriptor-driven: hostBehaviors
.sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are
byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path
segment — separators, dot-only segments, trailing dots, absolute paths, NUL,
and Windows reserved device names all fall back to the default, because the
value is joined onto a user's config root and written to.

Renamed in place rather than relocated: hook scripts resolve siblings via
__dirname/.., so a depth change would silently break them.

- install / uninstall / manifest sites all read the resolved name
- pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded
  trees still resolve; the never-throws contract is preserved
- new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new
  non-recursive remove-empty-dir engine primitive (rmdirSync only,
  symlink-refusing, containment-guarded); ADR-0008 amended accordingly
- fixes two latent name-dependencies the rename exposed: the stale-hook scan
  and the injection scanner's self-exclusion both hardcoded 'hooks'

Verified on the remote runner.

Closes #3023

* fix(#3023): close review findings and align emitted provenance with the rename

Adversarial review found two defects, and the remote runner found four
failure clusters. All fixed here.

Review BLOCKER — detect-custom-files was blind to the renamed bundle.
GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the
whole gsd-hooks/ tree was invisible to the custom-file scan and user-added
files there were never backed up before the next update's clean-install wipe.
The dir set now resolves via the .gsd-runtime marker plus the shipped
capability registry (never bin/install.js, which is not shipped into installed
trees), and falls back to scanning every known candidate when the runtime
cannot be determined — over-scanning is safe, under-scanning is the data loss.

Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir
accepted any directory, so an interrupted install left gsd-hooks/ winning over
a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now
qualifies only if it is non-empty.

Remote-runner clusters:
- emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped
  rules pointing at the same sources the existing hooks/ rules use. The table is
  total, so an unattributed family is a hard failure by design.
- pi tests in install-minimal-hooks and the install integration suite asserted
  the old layout; updated to derive the dir name from the descriptor rather than
  hardcoding either name.
- 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after()
  runs in registration order, cleanup was registered before mock.restoreAll(),
  and node22's JS rimraf calls the public fs.rmdirSync while node24's native
  path does not — so the EACCES stub leaked process-wide on one lane. Restore
  now runs first.

Verified on the remote runner.

* fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde

pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent
(packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty
configHome.env, so a user with that variable set had GSD installed where pi
never looks. Added the env name; the dot-home-nested resolver already handled
the override, so no resolver logic changed.

Also fixes expandTilde in the shared runtime-homes resolver, found while adding
that: it hardcoded os.homedir() and ignored the opts.home every caller threads,
so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi)
silently resolved against the real home. That is a correctness bug and a
test-escape hazard — a sandboxed test asserting on a tilde override reached the
developer's actual home directory. Now threaded through every branch; behavior
with no injected home is unchanged.

Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location
moved with the rename. The provenance rules satisfy the totality gate; the
differential gate needs the ack because the hook sources are byte-unchanged —
only the installer's target directory moved. The two hook files this branch
genuinely edits stay attributed and are not double-acked.

Note on piConfig.configDir: it is read from pi's OWN installed package.json
(getPackageDir walks up from pi's __dirname), alongside piConfig.name — a
white-label setting for a redistributed pi fork, not a per-project user setting.
Documented accordingly rather than treated as an unsupported override.

Verified on the remote runner.

* fix(#3023): reject blank env overrides, pin adapter/descriptor parity

Three review findings, all fixed.

A whitespace-only config-dir override was accepted verbatim: the guard was
`if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR='   '
resolved to a literal three-space directory name instead of falling back to the
descriptor default. Fixed across every env-consuming branch — dot-home,
dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's.
Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working.

pi/gsd.cjs's probe list and the descriptor were two independent sources of truth
for the bundle directory name; a future rename would have desynced them silently
and left every pi hook quiet with no error. The probe list stays deliberate — it
must resolve in a dev checkout and a half-upgraded tree, where the registry's
answer would be wrong — so this adds the parity assertion the repo's
generative-fix-divergence rule calls for: the descriptor value must be the FIRST
candidate, and the default must remain present.

Changeset body rewritten to cover the two later user-facing fixes it had not
caught up with.

Verified on the remote runner.

* chore(#3023): backfill changeset PR number

* fix(#3023): anchor injection-scan patterns and fix a macOS detection hole

CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the
same fact as a genuinely empty or absent one'. The match was the 'act as a'
INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending
in act tripped it (fact, impact, contract, artifact, interact, redact,
abstract). My four-line CONTEXT.md edit dragged the latent false positive into
this PR because the scan is diff-scoped by file but reads whole files. Anchored
with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would
have left the class alive for the next PR touching any file saying 'fact as a'.

Auditing the rest of the list for the same class surfaced a real detection hole:
the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex
escape. BSD/macOS grep reads it as four literal characters, so single-quoted
eval('...')/exec('...') payloads were NEVER detected there while passing on
GNU-grep CI. Replaced with a literal apostrophe class.

Boundaries were added only where a real word-suffix collision exists; exec,
jailbreak, developer mode and the role-manipulation family were audited and
deliberately left unanchored. 22 new cases cover both directions — the false
positives now scan clean, and every real payload still fires, including the
quote/punctuation/start-of-line boundary forms.

Also builds this branch's injection test fixture at runtime instead of carrying
the literal phrase, so the payload keeps its teeth without tripping the scan.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-07 13:41:21 -04:00
sim
c643320cef docs(#3155): ADR-3128 — shipped default is off, not adaptive
The maintainer decided the default after the ADR first merged, by
consistency with the shipped workflow config: verification gates default
on (research, plan_check, verifier, nyquist_validation,
security_enforcement), agent autonomy defaults off (auto_advance,
research_before_questions, plan_bounce, cross_ai_execution). Installing
probes into tracked source without a second confirmation is autonomy,
not a gate.

Adds Decision 8, flips the legacy/absent-section default and the
precedence tail to off, and marks Open question 2 resolved.

Decision 1's justification is corrected rather than deleted. It rested
on 'adaptive carries no flag', which the new default makes false. The
conclusion is unchanged and the real reason is stronger: precedence
includes the saved session policy, so a resumed session that persisted
adaptive passes no flag either, and a flag-keyed atom would exclude the
section from exactly the sessions already running the protocol.

Closes #3155

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 11:24:10 -04:00
sim
55136a19e9 docs(#3155): ADR-3128 adaptive runtime evidence — Phase 0 design lock
Records the design decisions #3128's maintainer approval made a
condition: schema v1, the probe/artifact ownership model, and the
cleanup state machine that gates terminal transitions.

Also amends ADR-1671 with a RESERVED atom rather than a widening. The
vocabulary stays at 29 until #3128's implementation lands; the
reservation exists so the widening is a coordinated decision rather than
an organic edit found in review.

The load-bearing decision is the atom's shape. #3128's probe policy is
tri-state (adaptive|force|off), so gating on flag:--runtime-probes would
exclude the protocol section from every default invocation -- adaptive
carries no flag -- and the feature's primary mode could never activate.
That is admission gate (2)'s silent-exclusion failure arriving through a
different door: not a fact nobody computes, but a fact computed for only
one of three policies. The atom is therefore a resolved boolean folded
in cmdInitDebug.

Closes #3155

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 11:09:53 -04:00
Tom Boucher
7203011400 feat(#3072): ship the deferred MCP served catalog (resources + prompts) (#3083)
* test(#3072): add failing-first coverage for the mcp served catalog

55 input-class rows from the phase test matrix, across four suites: the
catalog module over injected readFile/readDir seams, the protocol surface
through handleMessage, the install-vs-catalog parity gate, and fast-check
properties for uri round-trip, traversal refusal, and pagination partition.

src/mcp-catalog.cts lands as a skeleton whose functions throw, so the suites
fail on BEHAVIOR rather than on a missing module. The REASON enum is real so
tests assert typed codes instead of message prose.

Hostile coverage for the one client-controlled path surface (resources/read):
dot-dot and backslash traversal, percent- and double-encoded traversal,
absolute posix and windows paths, file:// scheme, null byte, symlink escape,
unindexed sibling, non-string and empty uri, wrong root segment.

IO faults are injected by monkeypatching the seam, never chmod 0o000 - root
bypasses mode bits, so a permission-based test silently passes with zero
coverage in root CI.

Refs #3072

* feat(#3072): serve the mcp catalog as resources and prompts

gsd-mcp-server now serves GSD's own content alongside its three tools: the
workflow, reference and command tree as MCP resources (resources/list, cursor
paginated, and resources/read over gsd://<segment>/<relpath> uris) and the 71
commands/gsd/*.md as MCP prompts keyed by bare command name. initialize
advertises resources and prompts, and deliberately does not advertise
subscribe or listChanged - the catalog is fixed for a server process lifetime,
so declaring a notification we never send would be a lie a host acts on.

Composition scope is SHARED, not re-declared. shouldCompose lives in
src/mcp-catalog.cts and bin/install.js now imports it instead of carrying its
own regex, so the served catalog and the installed file floor cannot drift on
what gets composed. Proven behavior-preserving across all 2871 tracked paths
plus windows-backslash, absolute and near-miss-prefix cases: zero mismatches.
tests/mcp-catalog-parity.test.cjs asserts served text equals the installer
composition-stage text over the real tree, with anti-vacuity guards requiring
both a marker-bearing workflow and a non-composed file in the comparison set.

Two measurements corrected the literal issue text. Composition is scoped to
gsd-core/workflows/ only, because a reference or command that documents marker
syntax with an unfenced example would otherwise be parsed as carrying a real
marker and have that line lossily dropped. And parity is asserted at the
composition stage rather than against an emitted runtime tree, since install
applies per-runtime path rewrites afterwards and the catalog is host-agnostic,
so byte equality with any one runtime would be false by construction.

resources/read is the one client-controlled path surface and is guarded in two
independent layers: the uri must be an exact key in the prebuilt index, which
defeats every traversal string by construction, and the mapped path is then
re-checked with validatePath so a symlink planted inside a root after indexing
is still refused.

Also fixes a real drift defect found while here: SERVER_VERSION was hardcoded
1.7.0 while the package is at 1.9.1. It now resolves lazily from VERSION or
package.json, reusing the precedent in runtime-artifact-conversion.

Closes #3072

* test(#3072): make the catalog parity gate drive the real installer

Review found the parity gate vacuous: it never imported or spawned
bin/install.js, and recomputed the installer side with the SAME shouldCompose
and composeWorkflow the catalog calls internally. It therefore proved only
that src/mcp-catalog.cts is self-consistent. The old row 52 compared
shouldCompose against a regex literal frozen in the test file rather than
against the installer at all. An inline divergent regex re-added to
bin/install.js - the exact regression ADR-1671 asks this gate to catch - would
have left the suite green.

The gate now spawns a real bin/install.js and compares the composition
DECISION, observed as gsd:section marker survival, against what the catalog
serves for the same files. Marker presence is the right observable because the
installer applies per-runtime path rewrites after composing while the catalog
applies none, so raw byte equality between the two surfaces is false by
construction and must not be asserted.

Sensitivity was proven, not assumed: overlaying the shouldCompose export that
bin/install.js imports so it always returns false makes a real spawned install
leave autonomous.md's markers in place while the catalog still strips them,
and the row 48 assertion diverges.

Anti-vacuity guards are kept and extended - the comparison set must be
non-empty, must contain a workflow that actually carries markers, must contain
a file the predicate declines to compose, and the install must have emitted a
non-zero file count. The marker-documenting reference case has no instance in
the real tree, so it uses an overlay fixture built with the same technique
workflow-fragments-emission.install.test.cjs already uses.

Renamed to .install.test.cjs so it lands in the install suite it now belongs to.

Refs #3072

* test(#3072): retarget the unknown-method assertion off a now-implemented method

tests/gsd-mcp-server.test.cjs used 'resources/read' as its example of an
UNKNOWN JSON-RPC method. The served catalog implements that method, so it now
returns -32602 (no uri supplied) rather than -32601. The remote runner caught
it deterministically on both linux lanes: -32602 !== -32601.

The test's intent is still correct and worth keeping, so it is corrected
rather than deleted or weakened. It now uses 'resources/subscribe', which the
server deliberately does not implement and deliberately does not advertise in
initialize's capabilities, because it never sends the corresponding
notification. That turns the assertion into a real contract - the advertised
capability surface and the implemented method surface agree - instead of an
arbitrary method name a future feature could invalidate the same way.

Swept the rest of the suite for other assertions pinning the newly implemented
methods; this was the only one.

Refs #3072

* chore(#3072): backfill changeset PR number 3083

* test(#3072): make the catalog fake fs separator-agnostic for windows

CI caught this on windows-latest (22 and 24): every catalog fixture indexed
ZERO entries, surfaced by the anti-vacuity guards as 'fixture catalog must
actually index resources for this property to mean anything'.

Mechanism: makeFakeFs keyed its dirMap/fileMap on POSIX-joined paths
(${root}/${rel}), while production buildCatalog looks paths up with
path.join, which is backslash-separated on Windows. Every lookup missed,
tryReadDir returned null, and the catalog came back empty.

Production is NOT at fault and is unchanged. The same CI run proves it: on
windows-latest the real-filesystem tests all passed, including 'installer
composition decision matches the served catalog for every file in the real
installed tree' and the row-51 non-vacuity proof against a real spawned
installer. A real Windows fs accepts both separators; the FAKE did not, so the
fake was the unfaithful one and is what changed.

Lookup keys are now normalized unconditionally with .replace(/\\/g,'/') in
readDir and readFile - never path.sep-conditional, never platform-gated. The
row-42/43 injected-fault wrappers got the same treatment, since they compared
raw production paths against POSIX-literal fixtures.

No assertion was weakened, and the anti-vacuity guards that caught this are
untouched - they are the reason this surfaced as a loud failure instead of a
suite that silently asserted nothing on Windows.

Refs #3072

---------

Co-authored-by: sim <sim@local>
2026-08-05 13:30:55 -04:00
Tom Boucher
d2e727d3b3 docs(#3074): correct adr-1671 mcp citation and stale runtime counts (#3080)
ADR-1671 attributed its MCP deferral to "ADR-857 §7 / #956" at five sites.
Neither source supports it: docs/adr/857-capability-system.md contains zero
MCP references (its Decision 7 is third-party code-loading, Decision 8 is
Runtime-as-Capability), and #956 is the closed first-party MemPalace plugin
pre-proposal that ADR-1239 explicitly disclaims in its own header.

The deferral itself is sound on ADR-1671's own runtime-partial reasoning and
never needed the borrowed citation. Ground it there, cross-reference ADR-1239
as the ADR that owns GSD's MCP surface, and record that a companion MCP server
shipped 2026-06-28 with three tools - so "MCP is deferred" is not misread as
"GSD has no MCP server". The deferral narrows to the served resources and
prompts catalog (#3072).

Also record that "deferred-tools", named alongside resources and prompts, is
not a deferred surface but an unbuildable one: MCP defines three server
primitives and the tools surface is tools/list plus tools/call, so schema
deferral is host behavior, not a server capability (#3075).

Correct two stale counts: 15 runtimes -> 19 (of 44 capability descriptors).

The ADR-857 reference in Open questions is a genuine Phase-6 completion
property and is deliberately left untouched.

Closes #3074

Co-authored-by: sim <sim@local>
2026-08-05 09:10:55 -04:00
Tom Boucher
c899f5ada3 chore(#3065): build the deterministic load-bearing-fragment contract gate (#3068)
* test(#3065): build the load-bearing contract gate ADR-1671 promised

Epic #1671 Phase 7. A post-merge audit of every promise in ADR-1671 against the
merged tree found one mitigation asserted-but-absent and two stale records.

ADR-1671 names exactly one correctness risk — trimming a load-bearing fragment,
with the recorded history of a paraphrased META.RULE causing agent violations —
and #2931 amended its mitigation to a deterministic contract gate that proves no
load-bearing fragment was omitted or shrunk, treats a floored fragment as a
success, and asserts the isolate prefix survives byte-identical, with an explicit
anti-vacuity rule.

That gate did not exist. What existed was tests/context-composer.test.cjs:
synthetic unit tests of the composeWithinBudget primitive over invented
fragments, asserting nothing about real declared strategies. The ADR asserted a
mitigation that was never built, which is the promised-but-not-built shape the
epic's own coverage discipline exists to catch.

The gate derives its load-bearing set from declared verbatim strategies rather
than a hand-maintained list, so it cannot go stale as upstream changes. It sweeps
budgets from 4x total down to a quarter of total and asserts at every step that
no load-bearing id appears in omitted or shrunk, that isolatePrefix is
byte-identical, and that hardFailed is surfaced rather than silently passed.

Both anti-vacuity guards are EXECUTABLE, not comments. One proves the empty
load-bearing set guard actually throws. The other proves a sweep that never
applies pressure is rejected — because a gate that only ever runs unpressured is
exactly how the original mitigation went missing without anyone noticing.
Measured: underPressure true at 6 of 7 budgets, false only at 4x total.

Three ADR records corrected in the same change, all doc-vs-reality drift:

  - Decision item 2 describes a composer that trims by priority to fit a measured
    per-runtime cap. composeWorkflow in fact passes MAX_SAFE_INTEGER with every
    fragment verbatim (both verified in source), so no trimming happens there;
    the emitted-byte cap is a separate measure-and-fail gate and Windsurf's limit
    a bespoke truncation. The wording described an option as shipped behavior.
  - flag:--converge never reached a terminal state. #2992 withheld six atoms;
    five were resolved explicitly. This one was resolved in code by reusing
    state:plan-strategy-converge but recorded nowhere — the same gap #2995 closed
    for flag:--verify-only, and I closed five of six.
  - The open-questions list enumerated three questions while two Resolved-by
    blocks resolved an unlisted Question 4. It is now listed.

Refs #3065

* fix(#3065): make the gate assert over production, not a copy of it

The isolated review found a blocker, and it was fatal to the gate's purpose: it
hand-copied applyBudget's fragment array into the test, so flipping a strategy in
src/prompt-budget.cts — say roadmap from verbatim to drop — would leave the gate
computing from its own untouched copy and still passing. A guard built as an
instance of the very divergence class it exists to prevent
(DEFECT.GENERATIVE-FIX) is worse than no guard, because it reports green.

Fixed by eliminating the duplicate rather than adding a parity assertion, the
same resolution used for the FAMILIES table in #2996. applyBudget's inline
construction is extracted to an exported buildBudgetFragments(), which both
applyBudget and the gate now call; the 1024 plan floor is exported as
PLAN_FLOOR_CHARS instead of being re-declared in the test. The extraction is pure
— verified behavior-preserving at budget=2000: hardFailed false, omitted
['context'], projectMd shrunk, plan truncation ~27.8%, all headers present. There
is no longer a second copy to diverge from.

Also fixed a vacuous assertion the same review caught: isolatePrefix was pinned
across the sweep, but no production fragment sets isolate:true, so the value is
always '' and the check could never fail. The pinning assertion stays, with an
honest comment that nothing in production sets it today, and a second test now
constructs an isolate:true fragment set and proves the prefix is non-empty and
byte-identical across a roomy and a severely tight budget — which is what makes
the first assertion capable of detecting a real change.

Refs #3065

* chore(#3065): backfill changeset pr number to 3068

---------

Co-authored-by: sim <sim@local>
2026-08-04 22:41:01 -04:00
Tom Boucher
ed360cd99f chore(#2995): extend fragment emission to agents/ and reclaim size-cap headroom (#3058)
* feat(#2995): extend fragment emission to agents/ across every read point

Epic #1671 Phase 6.4. `composeWorkflow` stripped `<!-- gsd:section -->` markers
only for `gsd-core/workflows/`, so a marked agent shipped its markers verbatim
into every runtime — and agent text is loaded into a subagent's context on every
dispatch.

The issue proposed widening the `copyWithPathReplacement` guard. That is a no-op
for agents: agents never traverse that function. Agent content is read for
emission at five independent points, and the obvious chokepoint
`stageAgentsForProfile` short-circuits on the DEFAULT `full` profile
(`skills === '*'` returns the real unstaged directory), so a hook placed there is
dead code on most installs.

Composition now happens at two call sites instead of five parallel surfaces:
`stageAgentsForRuntimeWithConverter` (with `agentsKind` and `kimiAgentsKind`
routed through it via an identity converter) and the inline agent loop in
bin/install.js. Both compose BEFORE any path rewrite, so a `.claude/` ->
`.windsurf/` regex can never reach inside a marker attribute — the ordering
#2930 established for workflows.

`installCodexConfig` was the fifth read point: Codex embeds each agent's prompt
into a per-agent `.toml` via its own readFileSync. Call-graph analysis missed it;
the exhaustive per-runtime emission sweep found it. That is why the new guard is
behavioral rather than structural — a sixth read point fails the sweep without
anyone remembering to extend a list.

tests/agent-fragments-emission.install.test.cjs spawns a real installer for every
runtime at every agent-bearing scope, derived from RUNTIME_META and the
capability registry at run time so a new runtime cannot be silently
under-covered. It asserts markers are absent AND the `when="always"` body is
retained, so marker-absence cannot be satisfied by dropping content. An
identity-composer negative control proves the assertion can fail.

Verified: 0 install failures, 0 marker leaks, body retained on 27 runtime/scope
paths; red before the wiring on claude(global+local), zcode(global+local),
kimi, codex and opencode.

Refs #2995

* chore(#2995): give the tightest agents headroom and correct the design lock

Epic #1671 Phase 6.4, second half.

`agents/gsd-verifier.md` had 12 bytes of headroom under its 49,152-byte LARGE
cap and `agents/gsd-debugger.md` had 147 under its 57,344-byte XL cap. Both now
extract reference material to `gsd-core/references/` behind an @-reference — the
documented DEFECT.AGENT-FILE-SIZE-CAP-BREACH remedy:

  gsd-verifier  49,140 -> 46,371 B   headroom    12 -> 2,781
  gsd-debugger  57,197 -> 48,851 B   headroom   147 -> 8,493

Byte accounting proves no content was lost: the combined agent+reference delta
is exactly the new files' headers plus the agents' slim replacement blocks. Each
agent keeps its routing table and a one-line summary per entry, so it degrades
gracefully on a runtime that does not inline @-references.

`agents/gsd-planner.md` is untouched and still passes both char guards
(49,130 < 49,152); it needed no change, so it took none.

The other nine LARGE/XL agents carry NO gsd:section markers, and that is
deliberate, not deferred. `when=` selection is read from
gsd-core/workflows/section-manifest.json, which gen-section-manifest.cjs derives
from gsd-core/workflows/*.md only — shape `{workflows: ...}`, no per-agent key,
no per-agent init entry point. An agent atom therefore fails admission gate (2)
("a fact the init seam demonstrably computes at a real entry point") and would
evaluate false forever while looking like working gating. Marking agents would
manufacture exactly the silent-inertness rot the frozen vocabulary exists to
prevent.

ADR-1671 gains three amendments, two of which close gaps /adr-phase-coverage
found against what actually merged:

  - The 19 -> 29 vocabulary widening shipped in #2994 with no coordinated ADR
    amendment, which that bullet's own rule forbids. Recorded now.
  - `flag:--verify-only` was one of six atoms #2992 withheld and deferred to
    "the LARGE/XL rollout phase". Five shipped; this one is permanently
    rejected, and that disposition lived only in a merged PR body.
  - Phase 6.4's own finding: emission extends to agents/, gating does not.

CONTEXT.md's glossary was stale on both seams — Workflow Fragments Module still
listed the original 4-atom vocabulary and described when= as "not yet acted on",
and Section Manifest Module still described InvocationFacts as
{waveFlag, phaseNumber, hasPriorPhases}. Both now match the shipped contract.

Inventory manifest regenerated AFTER build:lib per the documented ordering
landmine; 19 install-tree fixtures pick up the two new references.

Refs #2995

* chore(#2995): correct the compose-site count and mark the raw stager

Self-review found two comment defects in the prior commit. The agentsKind
comment claimed composition lands at TWO call sites; it is three, since
installCodexConfig's per-agent .toml writer was added after that comment was
written. And stageAgentsForProfile is now production-dead — both callers route
through the composing stager — while staying exported and unit-tested, which
makes it a trap: it does a raw copyFileSync and short-circuits to the unstaged
source directory under the default profile, so a future caller would silently
reintroduce the marker-shipping path. Its JSDoc now says so.

* test(#2995): guard the marker-documenting-doc class for agents

Widening the composer's scope to agents/ makes reachable the exact class #2930
narrowed scope to avoid: a file that DOCUMENTS the marker syntax with an
unfenced example is indistinguishable from a real marker, so the composer drops
that line from the emitted artifact.

Three rows. A fenced example must compose byte-identically. No shipped agent may
carry a marker outside a fence — asserted by parsing every real agent and
requiring zero explicit sections, which is what makes the fence protection
load-bearing rather than decorative. And a non-vacuity row asserts an UNFENCED
marker IS parsed as a real marker, so if that ever stops being true the second
row is guarding nothing.

Also applies two review findings: stageAgentsForProfile's new JSDoc claimed it
had no production caller, which is false — bin/install.js's _stageAgents still
calls it, and its consumers compose before writing. Corrected to state the
invariant instead. And a let/const nit in the emission sweep.

* fix(#2995): keep verifier status vocabulary in the agent, fix a wrong fixture

The first remote run came back red with three failures. Both root causes were
mine.

1. tests/agent-frontmatter.test.cjs requires agents/gsd-verifier.md to literally
   contain HOLLOW and DISCONNECTED. The Step 4b extraction moved that status
   vocabulary into gsd-core/references/verifier-wiring-patterns.md, so the agent
   no longer had it.

   Byte accounting said no content was lost, and byte-wise that was true — but a
   contract required those tokens to live IN THE AGENT. That is ADR-1671:66's
   flexReserve floor stated concretely: a load-bearing fragment must not be
   trimmed out of its host, and "the bytes still exist somewhere" is not the
   test. The two status tables are restored to the agent and deliberately
   mirrored in the reference with a note saying so, so the procedure there still
   reads standalone. gsd-verifier lands at 47,069 B — headroom 12 -> 2,083,
   rather than the 2,781 the first attempt claimed.

2. Row 12b of the new marker-documentation guard asserted that an unfenced
   marker example parses as a real marker, and threw instead:
   "unmatched /gsd:section close marker". The grammar is WHOLE-LINE only. The
   fixture had put the OPEN marker inline mid-sentence, so it was correctly not
   recognised as an open while the close, on its own line, was.

   That is a real refinement of the hazard this guard exists for: only a marker
   on its OWN line is mis-parsed — which is exactly how a documentation example
   is normally written. Row 12b now uses a whole-line marker, and a new row 12c
   pins the inline case as explicitly NOT a marker.

No test was weakened to accommodate the change; the change was corrected to
satisfy the tests.

Refs #2995

* chore(#2995): backfill changeset pr number to 3058

---------

Co-authored-by: sim <sim@local>
2026-08-04 18:10:31 -04:00
Dennis Kim
d3ddcaba1c fix(#2785): implement missing gate predicate evaluators (#2816)
* fix(#2785): implement missing gate predicate evaluators

* fix(#2785): gate predicate numerical coercion

* fix(#2785): address evaluator review findings

* fix(#2785): use safe frontmatter read seam

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:16:56 -04:00