Commit Graph

4910 Commits

Author SHA1 Message Date
Tom Boucher
de78f2eef2 docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate (#3010)
* docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate

security-model.md, USER-GUIDE.md, ARCHITECTURE.md, COMMANDS.md,
FEATURES.md, and gsd-planner.md's STRIDE template (+ ja-JP mirrors)
described the pre-ADR-0656 design: slopcheck as the install-or-degrade
gate, with unavailability degrading every package to [ASSUMED].
ADR-0656 inverted this months ago — registry-API verdicts (npm/PyPI/
crates.io) are the gate; slopcheck is an optional escalate-only adapter
that no shipped configuration wires. Verified every replacement claim
against src/package-legitimacy.cts (checkPackages, classifyPackage,
lookupNpm/lookupPypi/lookupCrates) via Memtrace before writing it, so
the corrected prose matches the live implementation rather than
restating the ADR from memory.

Restored docs/explanation/security-model.md:79-84 (and its ja-JP
mirror) to original wording after an orthogonal spec review caught
that an earlier draft had edited the "Why WebSearch packages are
always [ASSUMED]" paragraph — inside the range issue #2775 explicitly
named as correct and to leave alone.

The ja-JP mirror was missing the closing clause present in the
corrected English original ("its absence leaves registry-API verdicts
intact rather than downgrading everything to [ASSUMED]") — added for
parity. This completes the ja-JP mirror the issue's acceptance
criteria named explicitly.

zh-CN/ko-KR/pt-BR (not named by #2775, but carrying the same stale
design) get the mechanical portion of the same fix: command-string
swaps, table headers, ARCHITECTURE.md diagram labels, and technical-
term swaps that reuse a word already attested elsewhere in the same
file (合法性/적법성/legitimidade for "legitimacy") — surrounding prose
untouched. The remainder in those three locales — full-paragraph
rewrites of the corrected degrade-path mechanism, deleted "External
dependency" bullets, and "manually install slopcheck" code blocks —
needs prose composed by a fluent speaker of each language and is filed
as open-gsd/gsd-core#3002 with an exact file:line inventory.

* test(#2775): acknowledge gsd-planner.md byte growth from the STRIDE-row fix

agents/gsd-planner.md grew 14 bytes (49309 -> 49323) from the STRIDE
supply-chain row correction (slopcheck -> package-legitimacy gate).
Emitted agent/workflow files are byte-tracked; this fragment
acknowledges the growth per tests/emitted-attribution.test.cjs's
"differential attribution over the real tree" check.

* docs(#2775): close ja-JP FEATURES.md gap; fix a ko-KR transliterated heading

docs/ja-JP/FEATURES.md:2808 still read the katakana transliteration
"スロップチェック verdict" in REQ-PKG-GATE-01 — invisible to a literal
"slopcheck" grep, so it was missed when ja-JP parity was checked and
declared complete. Corrected to "正当性判定" (legitimacy verdict),
matching the term already established in ja-JP/explanation/
security-model.md and ja-JP/USER-GUIDE.md. This was the only
remaining ja-JP gap; a full sweep for the transliterated form across
docs/ja-JP/ now returns zero hits, and the ja-JP mirror is genuinely
at parity.

docs/ko-KR/USER-GUIDE.md:398's heading "슬롭체크 판정:" had the same
transliteration problem. Fixed inline to "적법성 판정:", reusing the
적법성/legitimacy word already attested two lines below in the same
table. A parallel sweep of zh-CN and pt-BR found no transliterated
forms of "slopcheck" in either locale. The remaining transliterated
occurrence in ko-KR (USER-GUIDE.md:406, the lead-in to the
pip-install code block) needs prose composition like the rest of that
block and is added to open-gsd/gsd-core#3002's inventory.

* chore(#2775): backfill changeset PR number to 3010

---------

Co-authored-by: sim <sim@local>
2026-08-02 20:24:42 -04:00
Tom Boucher
97f2af29da fix(#2852): isolate wave-cleanup blocks to their own entry (#3009)
* test(#2852): add failing-first regression coverage for wave-cleanup isolation

Adds the #2852 test matrix to executeWorktreeWaveCleanupPlan: per-entry
block reasons must isolate to the blocked entry instead of aborting the
rest of the wave, and a deletion must only block when another wave
member's branch still touches the deleted path. These fail against the
current implementation (RED) — the fix lands in the next commit.

* fix(#2852): isolate wave-cleanup blocks to their own entry and scope the deletions guard to real dependents

executeWorktreeWaveCleanupPlan aborted the rest of a cleanup wave on the
first blocked entry (branch_mismatch, base_mismatch, worktree_dirty,
merge_failed, etc.), dumping every remaining entry into `pending`
untouched instead of evaluating it. Every per-entry block reason now
isolates via `continue` instead of `break` + bulk pending push. The one
exception is a failed --no-ff merge, which can leave repoRoot itself
mid-merge: that path now attempts `git merge --abort` and only halts the
remaining wave if the abort itself fails (an unrecoverable repo-level
failure), matching every other block reason's isolation.

The `branch_contains_deletions` guard also blocked any deletion
unconditionally, even one nothing else in the wave depends on (the
reported repro: folding a test file into a sibling and deleting the
original). It now blocks only when another wave member's branch still
touches the deleted path — computed lazily per wave so a run with no
deletions pays no extra git call, and fails closed (still blocks) when
a sibling's diff cannot be determined.

* fix(#2852): eagerly cache each entry's own diff to fix an ordering bug in the deletions-overlap check

The deletions cross-entry overlap check (previous commit) computed
each "other" entry's touched-files set lazily, the first time some
later entry's overlap check needed it. That is wrong: once an entry
has already been merged earlier in the same loop pass, its branch
becomes an ancestor of HEAD, and `git diff --name-only HEAD...branch`
silently collapses to empty. A dependent entry that appears BEFORE the
deleting entry in the manifest (and has therefore already merged by
the time the deletion check runs) would be missed, letting a
genuinely-depended-on deletion through undetected — a live violation
of the negative-space acceptance criterion, caught by /code-review's
Spec-axis before this shipped.

Fixed by populating each entry's touched-files cache eagerly, during
that entry's own turn in the loop, immediately before its own merge
attempt (the only step that can move HEAD) — so every entry's diff is
captured before it could possibly have been merged, regardless of
manifest order. Adds a regression test reproducing the exact broken
ordering (dependent merges first, then a later entry tries to delete
the file it depends on).

* refactor(#2852): extract shared git name-only line parser

/code-review's Standards axis flagged duplicated parsing logic:
`stdout.split('\n').map((l) => l.trim()).filter(Boolean)` appeared at
both the per-entry deletion list and the cross-entry touched-files
cache added by this fix. Extracted into parseGitNameOnlyLines(), used
by both call sites, so the two can't silently drift apart.

* revert(#2852): scope the fix to wave-isolation only, restore unconditional deletions guard

#2852's own triage comment explicitly deferred the deletions-guard
policy question as a separate product decision ("Policy/enhancement
ask, not a defect ... Out of scope: deciding or implementing an
opt-in mechanism for intentional deletions"). All four of the issue's
actual acceptance criteria concern wave isolation only. The prior two
commits on this branch built a cross-entry deletion-dependency
heuristic that substituted a derived judgment for that deferred
product decision — out of scope for a confirmed-bug fix.

Reverts: getEntryChangedFiles, touchedFilesCache, the overlapUnknown
fail-closed branch, parseGitNameOnlyLines, and the eager per-turn
cache-population call. `branch_contains_deletions` now blocks
unconditionally again (byte-identical trigger condition to pre-fix);
the only change is `continue` instead of `break` + bulk `pending.push`,
same as the other 7 block reasons.

Keeps: the full wave-isolation fix (all 8 sites) and the merge_failed
/ git merge --abort recovery-and-carve-out, both squarely inside the
issue's actual acceptance criteria.

The deferred opt-in-for-intentional-deletions decision is filed as
#3003, citing #2852's triage as origin.

* refactor(#2852): extract blockEntry() helper to remove duplicated block-assembly across 8 sites

/code-review's Standards axis flagged the repeated
"result.status='blocked'; result.reason=...; result.stderr=...;
results.push(result); ok=false;" shape at every one of the 8 per-entry
block sites this fix touches. Extracted into blockEntry(), called at
each site; each call site still owns its own continue/break decision.
No behavior change.

* fix(#2852): check actual repo state instead of git merge --abort's exit code

The merge_failed recovery path decided "genuinely unrecoverable, halt
the wave" based on whether `git merge --abort` itself exited
successfully. That is not a reliable signal: git refuses many merges
(e.g. "your local changes would be overwritten by merge") WITHOUT
ever creating a MERGE_HEAD, in which case repoRoot's tree was never
touched — but `git merge --abort` still fails with "There is no merge
to abort (MERGE_HEAD missing)?" in that exact safe case. Trusting
that exit code alone misclassified an ordinary per-entry merge
failure as a repo-level one and stranded the rest of the wave — the
exact defect #2852 exists to fix, reintroduced through the recovery
path (caught in review).

Fixed by checking repoRoot's actual state directly via
`git rev-parse --verify -q MERGE_HEAD` after the abort attempt:
MERGE_HEAD present means genuinely still mid-merge (unrecoverable,
halt); absent means safe (isolate and continue), whether because no
merge state was ever entered or because abort successfully cleared
it. An unexpected git error or timeout degrades to the conservative
"still mid-merge" answer rather than throwing or guessing.

Rewrites the "unrecoverable merge_failed" test, which previously used
the safe "There is no merge to abort" string as its unrecoverable
example — that pinned the defect as correct behavior. Adds the
missing case: an ordinary merge_failed that never entered a merge
state must not abort the wave.

* test(#2852): cover repoRootStillMidMerge's fail-closed branches

/code-review flagged that the two conservative fail-closed branches of
repoRootStillMidMerge (a timeout on the post-abort MERGE_HEAD check,
and an unexpected non-0/1 exit code such as a fatal git error) had no
test coverage — exactly the branches most likely to hide a mutation
survivor (e.g. a flipped `timedOut` check or a flipped final `return
true`). Adds both cases: each must halt the wave (fail closed) rather
than assume repoRoot is safe when its state cannot be verified.

* chore(#2852): backfill changeset PR number to 3009

---------

Co-authored-by: sim <sim@local>
2026-08-02 19:51:43 -04:00
Tom Boucher
67e2ff7b25 fix(#2855): scope the phase-locator archived-milestone fallback to the active workstream (#3008)
* test(#2855): add failing-first regression test for cross-workstream archive leak

Covers findPhaseInternal/getArchivedPhaseDirs in src/phase-locator.cts
resolving a pending workstream phase to an unrelated workstream's (or
flat-mode's) archived phase because the archive fallback hardcodes the
project-root .planning/milestones/ tree. Fails against the current
implementation; the fix lands in a follow-up commit.

* fix(#2855): scope phase-locator archived-milestone fallback to the active workstream

findPhaseInternal and getArchivedPhaseDirs in src/phase-locator.cts hardcoded
the project-root .planning/milestones/ tree when falling back to search
archived phases, ignoring GSD_WORKSTREAM. A pending phase in one workstream
whose own phases/ directory didn't exist yet would silently resolve to a
same-numbered phase archived under an unrelated workstream's (or flat-mode's)
history, complete with stale plan/summary counts and an archived status.

Route the archive fallback through planningDir(cwd) instead — the same
workstream-aware helper the active-phase search (three lines above) and the
archive-write path (archivePhaseDirectories in milestone.cts) already use.
Flat/non-workstream projects are unaffected: planningDir(cwd) with no
GSD_WORKSTREAM resolves to the same root .planning path as before.

Also switch the reported relBase/basePath from a hardcoded
'.planning/milestones/...' literal to path.relative(cwd, archivePath), so the
paths returned to callers stay consistent with wherever the archive actually
resolved to (root or workstream-scoped).

* chore(#2855): add changeset for phase-locator workstream archive fix

* fix(#2855): normalize getArchivedPhaseDirs basePath to posix separators

Orthogonal code-review finding: findPhaseInternal's relBase/directory field
was explicitly toPosixPath-normalized, but getArchivedPhaseDirs's basePath
used a bare path.relative() call, leaving it native-separator on Windows —
an inconsistency between two sibling "relative path from cwd" report fields
introduced by the same #2855 fix. Wrap basePath in toPosixPath to match, and
update the two existing assertions that compared basePath against path.join
output (which would break on Windows now that the field is guaranteed posix)
to compare against forward-slash literals instead, matching how the sibling
`directory` field is already asserted elsewhere in this suite.

* refactor(#2855): share archive-directory resolution between findPhaseInternal and getArchivedPhaseDirs

Orthogonal code-review finding: the two functions carried independent copies
of the same resolve-milestonesDir-then-enumerate-archive-dirs logic — the
exact shape that let the original #2855 bug (hardcoded root path) exist in
one copy while the workstream-aware active-phase search sat three lines
above it. Extract listArchiveVersionDirs(cwd) as the single seam both
functions now consume, so a future change to how the archive tree is located
only needs to happen once. Byte-for-behaviour preserved: readSubdirectories
and searchPhaseInDir already self-contain their own try/catch and never
throw, so moving the iteration outside the old inline try block changes
nothing observable (verified via manual repro scripts covering leak
prevention, positive resolution, flat-mode parity, and multi-milestone
reverse-sort ordering).

* test(#2855): demonstrate ROADMAP.md presence does not affect the archive-leak guard

Orthogonal code-review (spec axis) finding: issue #2855's AC1 states the
guard must hold "regardless of whether workstream A's roadmap already lists
the phase and when it doesn't yet" — an explicit two-value dimension that
had no direct test coverage; it was only inferable by reading
findPhaseInternal's source and confirming it never touches ROADMAP.md.
Add a parametrized test creating the workstream's ROADMAP.md with and
without a matching Phase heading, asserting the archive-leak guard resolves
identically (null) either way.

* chore(#2855): backfill changeset PR number to 3008

---------

Co-authored-by: sim <sim@local>
2026-08-02 19:24:53 -04:00
Tom Boucher
51f32d2d40 fix(#2658): detect trae runtime and resolve its instruction file to a concrete rules file (#3006)
* test(#2658): add failing-first regression for trae runtime detection and instruction path

Covers all three collided defects reported in #2658 plus a fourth
instance of defect 1 (ingest-docs.md) found while diagnosing it:
missing trae detection in workflow runtime-detection blocks, the
CLAUDE.md path-mutilation bug in both the js/cjs and md install-time
converters, and the missing projectInstructionFile capability
declaration. Fails against current source; the next commit fixes it.

* fix(#2658): detect trae runtime and resolve its instruction file to a concrete rules file

Three defects collided to produce the reported ".claude/.trae/rules/"
path:

1. new-project.md and ingest-docs.md's runtime-detection blocks only
   recognized codex/gemini/opencode and fell through to RUNTIME=claude
   for trae. Both now recognize the /.trae/ execution-context path and
   the TRAE_CONFIG_DIR env var before the claude fallback.
2. RUNTIME_CONTENT_DISPATCH.trae.js (bin/install.js) replaced bare
   "CLAUDE.md" before the ".claude/" prefix was handled, mutilating
   ".claude/CLAUDE.md" into ".claude/.trae/rules/". Now replaces the
   full ".claude/CLAUDE.md" path first, and targets a concrete file.
3. convertClaudeToTraeMarkdown (mirrored in bin/install.js and
   src/runtime-artifact-conversion.cts per the #2094 output-parity
   test) had the same class of bug with a different wrong output
   (".trae/.trae/rules/", from its generic ".claude/" rewrite firing
   after the bare CLAUDE.md rewrite). Both mirrors now match full-path
   forms before the bare/generic patterns, converging on the same
   concrete file as the js/cjs converter.
4. capabilities/trae/capability.json didn't declare
   hostBehaviors.projectInstructionFile, so getProjectInstructionFile
   fell through to the generic AGENTS.md default even when RUNTIME=trae
   was resolved correctly. Now declares ".trae/rules/rules.md",
   regenerated into gsd-core/bin/lib/capability-registry.cjs via
   npm run gen:capability-registry.

Closes #2658

* chore(#2658): add changeset

* fix(#2658): preserve arbitrary runtime-dir prefixes in the trae path rewrite

Found by the end-to-end --trae install regression test (not by static
trace) across two verification runs:

1. copyWithPathReplacement runs a generic ~/.claude/, $HOME/.claude/,
   and ./.claude/ -> runtime-dir rewrite on every .md file BEFORE
   calling convertClaudeToTraeMarkdown. The prior fix's
   .claude/CLAUDE.md-specific patterns never fire on that
   already-rewritten text, and the bare fallback still doubled
   whatever prefix the generic pass substituted. A first attempt
   handled only the fixed "./.trae/" shape and missed the
   $HOME/.claude/ and ~/.claude/ forms gsd-core/workflows/profile-user.md
   actually uses, which post-rewrite become an arbitrary absolute
   local-install-root path, not the fixed relative shape. Fixed with a
   prefix-preserving pattern that captures whatever precedes a
   ".trae/" tail and fixes only the filename suffix, instead of
   assuming one fixed shape.

2. The fix's own explanatory comments literally spelled out the
   malformed strings and the instruction filename as contiguous text.
   Since these two files ship verbatim into local --trae installs,
   where they are themselves run through the same find/replace, the
   comments got "fixed" right along with the real code, leaking the
   malformed string into the installed tree. Rewrote every comment in
   both mirror copies to never spell either the instruction filename
   or a malformed shape as one contiguous token.

Adds an emitted-drift-ack fragment: the corrected replacement target
for every CLAUDE.md mention (bare directory -> concrete file) changes
trae-emitted output for every repo file that mentions CLAUDE.md, not
only the ones that hit the originally reported bug.

* test(#2658): extend parity test with arbitrary-prefix .trae/ inputs

The bin/install.js vs runtime-artifact-conversion.cjs parity assertion
for convertClaudeToTraeMarkdown only fed the pre-existing bare
.claude/CLAUDE.md input through both implementations. Feed the
prefix-preserving cases (relative, nested-absolute, tilde, $HOME,
backtick-wrapped) plus a property-based check through both, so a
future edit to only one copy of the .trae/-tail regex fails this
test instead of silently diverging.

* chore(#2658): backfill changeset PR number to 3006

---------

Co-authored-by: sim <sim@local>
2026-08-02 18:29:47 -04:00
Tom Boucher
d770365753 docs(#2999): document the takeover process for a capability, reviewer lane, or EoS integration (#3000)
* docs(#2999): document the capability / reviewer-lane / EoS takeover process

The capability ecosystem documented a complete forward lifecycle — develop,
publish, version, import, update, remove, turn off — but nothing covering a
change of maintainer for an entry that already exists. Adds
docs/how-to/take-over-a-capability-or-eos.md defining four takeover modes
(consensual handoff, adoption fork, first-party absorption, retirement), the
PR shape each takes, a per-surface snapshot of the inherited user-visible
contract, and an install-continuity checklist.

Also corrects .github/PULL_REQUEST_TEMPLATE/registry-entry.md, which directed
contributors to a 'Registry' Discussions category that does not exist — the
category is named 'EoS Registry' per docs/registries/README.md, and because
'discussion' is a required field the thread must exist before the PR is
opened, so the wrong name stalled contributors at the first required step.

Closes #2999

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2999): backfill changeset pr number to 3000

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:02:58 -04:00
Tom Boucher
a987cf2731 chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init

Extends the init bundle with a typed per-invocation section manifest so an
invocation loads only the branch guidance it will actually take.

The three flag/state-gated branches in execute-phase.md move into their own
step files; the parent keeps its gsd:section markers wrapping a one-line
on-demand reference, so each section's prose lives in exactly one file and
the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator
derives the shipped section manifest from those markers, and a new pure
evaluator maps invocation facts to applicable section ids.

The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser
(Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary
and the predicate map stay exhaustively in sync.

Closes #2932

* fix(#2932): fail closed on prototype-chain when values

An isolated adversarial review found WHEN_PREDICATES[section.when] was a
bracket lookup on a plain-prototype object, so inherited Object.prototype
members resolved as predicates: "constructor"/"toString"/"valueOf"/
"hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and
"__proto__" threw an untyped TypeError carrying no .reason. Both violate
the module's documented fail-closed contract, and the manifest is read from
disk at run time so it cannot be assumed trustworthy.

Builds the predicate map on a null prototype and guards the lookup with an
explicit Object.hasOwn check. Adds table-driven coverage for nine
Object.prototype-shaped keys asserting the TYPED reason (asserting only
that it throws would still pass while broken) plus a fast-check property
injecting a hostile value at an arbitrary document position.

* test(#2932): retarget execute-phase step assertions at extracted step files

* fix(#2932): emit typed reasons for generator lib-load and write failures

* fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures

* chore(#2932): backfill changeset pr number to 2987

---------

Co-authored-by: sim <sim@local>
2026-08-02 12:34:41 -04:00
kyle-the-dev
33985c11a9 test(#2429): add codex local installer regression (#2436)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-02 01:32:30 -04:00
kyle-the-dev
ce38d44811 fix(#2777): remove stale codex local home metadata (#2831)
* fix(#2777): remove stale codex local home metadata

* chore(#2777): add changeset for codex local layout metadata

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-02 01:00:52 -04:00
Adnan
137f3fbb9c fix(#2562): scope workstream progress/status to the current milestone (#2588)
* fix(#2562): scope workstream progress/status to the current milestone

`workstream progress` / `workstream status` / `workstream list` share one
derivation that could report a workstream's CURRENT milestone as
"milestone complete" / 100% while phases in that milestone were unstarted,
in progress, or failing verification. Three coupled defects:

1. The shipped signal was project-lifetime, not milestone-scoped:
   workstreamMilestoneShipped() returned true if ANY *-ROADMAP.md snapshot
   existed or "SHIPPED" appeared anywhere in ROADMAP.md. Every prior shipped
   milestone leaves a permanent collapsed <summary>✅ … SHIPPED</summary>
   block, so any post-v1.0 workstream was pinned to "milestone complete"
   forever (over-correction from #1913).
2. The denominator dropped declared-but-unscaffolded phases, and completed
   PRIOR-milestone phase directories inflated the numerator, letting
   progress_percent round to 100 while real work remained.
3. Phase completeness ignored the VERIFICATION verdict — SUMMARY >= PLAN
   count alone marked a phase complete even with a human_needed verdict.

Fix: derive both numerator and denominator from artifacts scoped to the
current milestone. The current version comes from the workstream STATE.md
`milestone:` field (ROADMAP in-progress markers can be stale); the ROADMAP
`## Progress` table maps every phase — including dirless ones — to its
milestone, and the matching set is both the denominator and the directory
membership filter. The shipped signal now requires the CURRENT version's
archived ROADMAP snapshot (REQUIREMENTS snapshots are not accepted; they can
be written at milestone start) or the current milestone's own line marked
shipped. Phases with an explicit failing verdict (gaps_found/human_needed)
count as in_progress; missing/unknown/stale are left untouched so
verifier-disabled projects do not regress to never-complete.

Greenfield roadmaps with no versioned Progress table, and projects whose
current version cannot be determined, keep the prior behaviour.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2562): add changeset for workstream milestone-scoping fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(#2562): parse the Progress table via findTableWithColumns

The ad-hoc pipe-table regex tripped the local/no-adhoc-markdown-parsing
ESLint rule. Use the canonical markdown-table helper instead: the
milestone-grouped RoadmapProgress variant is located by its required
`Phase` + `Milestone` columns and cells are addressed by column NAME,
so the parser tolerates column reordering and injected columns. The
`flat` variant (no Milestone column) yields no attribution, which is
the intended fallback to legacy counting.

Behaviour is unchanged: verified against a real multi-workstream project
(same status/percent/phase and plan counts before and after).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2562): count table-only phases in the unscoped denominator

Addresses the reporter's repro detail: a phase declared as a `## Progress`
table row with no `### Phase N` heading is missed by countRoadmapPhases
EVEN WHEN other headings exist — the heading regex counts 1 for a
"1 heading + 1 table-only" roadmap — not just in the zero-heading fallback
path. Milestone scoping did not cover this, because a flat Progress table
(no Milestone column) carries no per-phase attribution, so greenfield and
single-milestone projects kept the old heading-only denominator and the
declared phase silently vanished from it.

When milestone scoping cannot engage, the denominator is now the union of
the Progress table's declared phase numbers and the phase directories, so
neither source can shrink it. Verified against the reporter's minimal
fixture (phase 1: 1 PLAN + 1 SUMMARY + gaps_found; phase 2: table row only,
no heading, no dir), which now reports 0/2 at 0% across all four
table/STATE permutations instead of 1/1 at 100%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2562): attribute dir-only sub-phases to their parent's milestone

A sub-phase directory inserted mid-milestone (e.g. `30.1-…` under a
table-declared phase 30) usually has no ROADMAP Progress-table row, so it
had no milestone attribution and was scoped out of the rollup entirely —
its completed work was invisible and it could never hold the percentage
below 100.

It now inherits its parent phase's milestone and joins BOTH sides of the
calculation. Both sides is the load-bearing part: adding it to the
numerator alone would let completed_phases exceed a denominator that never
counted it, cap back to 100% via Math.min, and reintroduce exactly the
defect this issue reports. A regression test pins that failure mode (all
declared phases complete + an in-progress dir-only sub-phase → 75%, not
100%).

Attribution is deliberately one-directional: a sub-phase counts only when
its PARENT is in the current milestone, so a follow-up created in a later
milestone under an older parent is excluded rather than misattributed —
conservative (under-count) rather than falsely inflating.

Verified on a real project: the reported workstream moves from 2/6 (33%)
to 3/7 (43%), the 3/7 being the honest figure — a completed sub-phase that
was previously invisible now counts, and so does its plan total.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2562): describe the denominator + sub-phase fixes in the changeset

The fragment was written at the first commit and only covered the three
original defects. Bring it up to date with what actually ships: the
table-only-phase denominator union (heading-only counting dropped a
declared phase even when other headings existed) and sub-phase milestone
inheritance across both sides of the calculation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(#2562): promote the canonical phase-key surface to phase-id

state.cts kept `phaseKeyFromToken`/`phaseKeyFromDir` private, so every other
module that had to compare two independently-derived phase references — a
ROADMAP table cell against a phase directory, say — wrote its own regex. That
is the defect class #2562 reports: a padded `01` and an unpadded `1-slug` land
in different key spaces and the comparison silently yields nothing.

Move the pair to the phase-id owner module and add `phaseKeyFromProse` (for
ROADMAP/STATE prose, markdown emphasis stripped) and `parentPhaseKey` (a
sub-phase's parent). state.cts imports them; its call sites are unchanged.

* fix(#2562): own milestone-shipped detection and accept a workstream scope

Three changes to the module that owns milestone parsing, so its consumers stop
reimplementing it:

- `isMilestoneShippedInRoadmap(content, version)` answers "does the ROADMAP mark
  THIS milestone shipped" from heading and `<summary>` lines only. A bullet that
  merely names the version (`- [x] 03-01: ship the v2.0 login endpoint`) is prose
  about a phase, not a milestone verdict. The version token is boundary-matched
  with `(?![\w.-])` — `\b` does not bound it, since `.` is a non-word character,
  so a shipped `v2.0.1` heading would otherwise close `v2.0`.
- `extractCurrentMilestone` and `getMilestonePhaseFilter` take an optional
  trailing workstream name and thread it to `planningDir(cwd, ws)`. A caller
  iterating workstreams cannot set `GSD_WORKSTREAM` per iteration, which is what
  the existing resolution falls back to. Omitted, resolution is unchanged.
- `getMilestonePhaseFilter` exposes `versionScoped`, true only when the phase set
  really is one milestone's. On an unversioned roadmap `phaseCount` spans the
  project's lifetime and must not be read as a current-milestone denominator.

The closed/active milestone-marker patterns were kept in three byte-identical
copies; they are hoisted to module scope as one `isClosedMilestoneHeading`.

* fix(#2562): derive membership and denominator from one phase-key space

The milestone scoping added earlier in this PR derived the ROADMAP table key and
the phase-directory key with two different regexes, and dropped rows it could not
attribute. Each of those was another way to reproduce the symptom this issue
reports — a rollup contradicting its own `phases[]` listing:

- a padded `| 01. … |` row never matched a `1-slug` directory (and a bespoke
  `^0*(\d+…)` never matched `PROJ-05-…` at all), so phases fell out of the
  milestone entirely and the percentage collapsed or pinned;
- a blank or malformed Milestone cell deleted the phase from BOTH sides, letting
  an unstarted phase vanish and the remainder round to 100%;
- shipped detection scanned bullets, so any checkmarked line naming the version
  closed the milestone;
- the numerator counted per-directory while the denominator counted distinct
  phases, so a stale same-numbered directory (Bug #2445's scenario) pushed
  `completed_phases` past the denominator, where `Math.min` capped it to 100%
  and hid the unstarted phase.

Both sides now key off the phase-id owner module (`phaseKeyFromDir` /
`phaseKeyFromProse`), directory membership additionally consults
`getMilestonePhaseFilter` when that filter is genuinely version-scoped, and the
denominator is the union of the roadmap's declarations with the member
directories' keys — so `completed_phases <= denominator` holds by construction.
The Builder asserts it and throws; the `Math.min` cap survives only on the legacy
unscoped path, where the denominator is a heading count that cannot bound the
numerator. An unattributable row degrades over-inclusively (kept, never dropped),
matching the degrade direction roadmap-parser already commits to.

* test(#2562): boundary coverage for each milestone-scoping reproduction

One test per way the scoping could still report "milestone complete"/100% while
phases are incomplete: zero-padded rows vs padded dirs (and the mirror),
project-code-prefixed dirs, a blank/malformed Milestone cell, a checkmarked
bullet naming the version, a shipped `v2.0.1` heading against a current `v2.0`,
and a stale same-numbered directory. Plus the current milestone's own shipped
heading (the signal must survive the boundary fix), the Builder's
numerator-above-denominator throw, a parity check that every non-`passed`
verifier status blocks completeness, and a guard that scoping reads the
workstream's ROADMAP rather than the project root's.

Reverting only `src/` reddens six of them.

* docs(#2562): record the milestone-scoped semantics and its consumer impact

CONTEXT.md: the Workstream Inventory Module's completion fields now describe the
current milestone, not the workstream's lifetime; phase-id owns the canonical
phase-key surface; roadmap-parser owns milestone shipped/active classification
and takes an optional workstream scope.

Changeset: name the behaviour change explicitly — `roadmap_phase_count`,
`completed_phases` and `progress_percent` change meaning with no schema signal,
and `getOtherActiveWorkstreamInventories` filters on the derived status, so
consumers see real movement.

* fix(#2562): collapse every zero-padding spelling to one phase key

A property test over the key surface — table cell and directory decorated
INDEPENDENTLY, which is the point — found a divergence neither review named:
`padStart(2, '0')` is a no-op once the input is already ≥2 characters, so `5`
normalised to `05` while `005` stayed `005`. A `| 5. … |` row and a `005-slug`
directory therefore never compared equal, which is the same failure mode as the
padded-vs-unpadded blocker, one level down.

The strip belongs in `phaseKeyFromToken`, not in `normalizePhaseName`: applying
it to the latter regressed multi-decimal leading-zero plan IDs (`001.10-PLAN.md`
capture + wave assignment), which rely on its verbatim rendering. Confining it
to the key surface fixes the comparison and leaves rendering untouched.

Also tightens `isMilestoneShippedInRoadmap`'s patterns to anchored,
complementary character classes so an untrusted ROADMAP cannot drive
backtracking, and makes the project-code test discriminating — it previously
passed pre-fix, because an unresolvable key collapsed scoping to the whole
roadmap and happened to land on the same number. It now carries a
prior-milestone directory that a collapse would wrongly admit.

* fix(#2562): prefer the milestone-attributing Progress table; pin the seams

Three gaps the earlier self-check missed:

- Both RoadmapProgress variants carry a `Plans Complete` column, so probing it
  first picked a FLAT table appearing earlier in the document over the
  milestone-grouped one that actually carries the attribution. Every row came
  back unattributed, was treated as current-milestone, and silently re-admitted
  prior-milestone phases. The attributing shape is probed first; flipping the
  order reddens the new test.
- `lint-phase-id-drift` exempts phase-id.cts by design, so it is silent on
  `phaseKeyFromToken`'s own segment strip by construction — not evidence. Its
  interaction with `stripProjectCodePrefix` (which runs AFTER) is pinned across
  project codes and hyphenated ids, including the pre-existing `M1-46-6` vs
  `M1-46-6-rs` asymmetry, which is `extractPhaseToken`'s #2043/#2232 slug-word
  rule and not something to "fix" by accident.
- `listWorkstreamInventories` loops every workstream with no try/catch, so a
  REACHABLE Builder-invariant throw would take down `workstream list`/`status`/
  `progress` for all of them. The invariant test only exercised the pure Builder
  with hand-built inputs. A test now drives `inspectWorkstream` over every
  adversarial shape at once (prior-milestone dirs, three colliding duplicates,
  a dirless declaration, an unattributed row, a project-code prefix, a dir-only
  sub-phase) and asserts it does not throw and the invariant holds — so the
  throw stays a contract assertion for external callers, not a runtime path.

Also covers the active-marker-wins rule (`## v2.0 — 🚧 IN PROGRESS … ✅` must not
mark shipped), which nothing exercised.

* fix(#2562): scope a declared-but-empty current milestone instead of falling back to history

The review's open MAJOR. `STATE.md`'s `milestone:` field updates the moment
`/gsd-new-milestone` writes the heading, while the `## Progress` table and phase
sections land later. In that window nothing attributes a phase to the current
milestone, `scoped` went false, and the fallback counted the project's ENTIRE
phase history as both numerator and denominator — a milestone with zero work
done reported 100% off its predecessors'. That is #2562's own symptom reached by
a different precondition, and none of the 16 tests covered it.

Reproduced first, four ROADMAP shapes, at `inspectWorkstream` rather than the
Builder — the Builder takes the scoping decision as an input, so a Builder-level
test proves it honours a flag, not that the derivation sets it. Three of the
four reported 2/2 100% with no phase of the current milestone begun.

Which signal witnesses the state depends on the ROADMAP's shape, and no single
one covers all three:

- `## v3.0` exists but declares no phases. `getMilestonePhaseFilter` DOES locate
  the section and sets `versionScoped`, then the zero-phase pass-all degrade
  resets it to false — erasing the only evidence the milestone exists. Neither
  existing flag survives that path, so this adds `versionSectionFound`, set
  beside `versionScoped` and deliberately preserved through the degrade.
- No section for this version at all, in a roadmap that versions its others —
  the existing `missingExplicitVersion`, already exposed and tested.
- Unversioned headings, but a Progress table attributing every row elsewhere:
  neither filter flag fires and the table is the only witness.

A ROADMAP that attributes NO versions anywhere matches none of them, which is
the point. Its rows parse with `version: null`, land in `currentMilestoneKeys`,
and never reach the new branch. `readCurrentMilestoneVersion` returns a non-null
version for very nearly every project (`getMilestoneInfo` defaults to `v1.0`),
so keying off `currentVersion` alone would have zeroed out every free-form
legacy project — the condition looks fussy for that reason. A test pins it.

Within an empty milestone, membership inverts: a directory belongs unless
another milestone's row claims it. Excluding everything would have dropped a
phase scaffolded before the roadmap caught up from BOTH sides of the rollup, and
hiding real work is the same class of defect as inventing it — this codebase
degrades over-inclusive, never under.

Scoping is now stated by the caller (`milestoneScoped`) rather than inferred
from `currentMilestonePhaseCount > 0`. That inference was the root cause: it
cannot represent a milestone that is scoped AND legitimately zero-phase, so the
Builder read "no phases yet" as "no scoping" and reopened the whole-history
path. The count-derived value stays the default for callers that say nothing.

A regression test also pins that a zero denominator does not trip the Builder's
`completed_phases <= denominator` throw, since `listWorkstreamInventories` has
no try/catch and a crash on every freshly-declared milestone would be worse than
a wrong percentage.

The changeset and CONTEXT.md no longer claim membership is derived in "ONE" /
"a SINGLE" phase-key space. `getMilestonePhaseFilter` still runs its own
`normalizePhaseIdSegments`; the signals are OR'd so a divergence can only widen
membership, but two normalisers coexist and the docs now say so.

* fix(#2562): cross-validate the shipped marker against the milestone's artifacts

`status: "milestone complete"` was asserted from the shipped marker alone, so a
single payload could report it beside `progress_percent: 67` — this issue's own
symptom, reached through `status` rather than the percentage.

The marker is now a claim checked against the milestone's own artifacts, and the
two signals are checked at DIFFERENT strengths because one check cannot serve
both. A `heading` marker (operator-typed, live ROADMAP) is refused on a short
completion ratio, which also catches phases declared but never scaffolded. A
`snapshot` marker is NOT ratio-gated: `milestone complete` moves the milestone's
phase dirs into `milestones/<version>-phases/` (milestone.cts:755-762) while
copying — never truncating — the live ROADMAP (:671-674), so a correctly
archived milestone reads 0/N by construction and a ratio gate would strip
`milestone complete` from every archived milestone. It is refused instead when
an in-milestone phase dir is still live and unfinished, reachable because
`milestone complete` does not advance STATE's `milestone:` field
(state-transition.cts:1335 vs :1224). `legacy` stays ungated — only reachable
when scoping is off.

A refused marker does not fall through to a STATE field claiming the same thing;
against contradicting artifacts neither source may report completion. The
refusal surfaces as `milestone_shipped_unverified` rather than staying silent,
distinct from `status_conflict` (derived-vs-field only).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* test(#2562): pin both marker strengths, the archived guard, and the owner modules

workstream-inventory: four tests, all four red against the prior src and green
with it. A live-ROADMAP SHIPPED heading over an incomplete milestone is refused;
an archived snapshot SURVIVES its phase dirs being moved out (the regression the
obvious single ratio-gate would cause — swapping the snapshot branch to that
gate reddens this AND the pre-existing `CURRENT-version snapshot marks the
milestone complete` at :321); an archived snapshot is refused once a phase is
reopened under it; and a refused marker is not re-asserted by a STATE field
claiming the same.

roadmap-parser: `isMilestoneShippedInRoadmap` gets unit coverage at its owner
module rather than only through the inventory that consumes it, plus two
characterisation tests for `getMilestonePhaseFilter`'s legacy call surface —
omitting the new trailing `ws` param is indistinguishable from `undefined`/`null`,
and the `GSD_WORKSTREAM` env fallback still resolves. These characterise the
call surface; they do not stand in for coverage of its individual callers.

phase-id: the `phaseKeyFrom*` / `parentPhaseKey` one-key-space contract, incl. a
property that padding a directory number never changes its key.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* docs(#2562): record the two-strength shipped cross-check + milestone_shipped_unverified

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* fix(#2562): refuse an archived snapshot on a DIRTY archive, not just a live-unfinished dir

The snapshot arm checked `liveIncompletePhases > 0`, which misses the shape
@davesienkowski reproduced: a COMPLETE live dir beside a phase declared in the
Progress table with no directory. Nothing is live-and-unfinished, the marker
sails through, and `cmdWorkstreamProgress` returns
`{"status":"milestone complete","progress_percent":50}` — the reported symptom
verbatim, from one payload. Reproduced at 483a3ba30 before changing anything.

His diagnosis is the right one and better than mine: an in-milestone directory
outliving the archive means the archive is not CLEAN, and once that is true the
completion ratio is meaningful again. So the check is the conjunction — any live
in-milestone dir AND `completedPhases < effectivePhaseCount`. That strictly
subsumes the old predicate (an incomplete member dir is in the denominator and
not the numerator, so the ratio is always short when one exists) and leaves the
clean-archive guard green, since a clean archive has no live dirs at all.

Also corrects the module comment: the `scoped &&` prefix ungates all three
signals, not just `legacy`. That is correct behaviour — unscoped, the
denominator is the whole-roadmap count and membership is everything, so there is
no current-milestone artifact set to check a current-milestone claim against —
but the comment claimed otherwise. And the `milestone.cts` citations were ~28
lines stale after the rebase; they are now :700-702 (copy) and :783-790 (move),
re-verified against this head.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* fix(#2562): project milestone_shipped_unverified from list, status and progress

The inventory carried the field and every renderer dropped it — `workstream.cts`
was not in this PR's diff at all — so at the CLI a refused marker looked exactly
like no marker: a fallback `status` and nothing saying one was seen and rejected.
That is the silent collapse this issue is about, reintroduced one layer up, and
it made the changeset's "visible rather than silent" claim false at every
surface.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* test(#2562): pin the dirty-archive shape and the CLI projection

Five tests, all five red against the prior src and green with it.

The reviewer's repro at the builder: an archived snapshot with a COMPLETE live
dir beside a dirless declared phase must be refused, and status must not
contradict the percentage.

Four at the CLI via runGsdTools, the surface that was dropping the field rather
than the builder that already had it: `workstream progress`/`status`/`list` each
project `milestone_shipped_unverified: true` for that workstream, and a clean
archive still reports `false` with `status: "milestone complete"`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

* docs(#2562): correct the snapshot check, the scoped-only caveat and the CLI claim

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpFzuEHKTN1jaypSN44rzd

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 23:06:43 -04:00
0xdhx
f4d6747d21 fix(#2738): report graphify query budget outcome and stop between-tier over-trimming (#2819)
* fix(#2738): report budget outcome from graphify query and stop over-trimming between tiers

applyBudget retains seed nodes unconditionally, so the seed set is a floor
the edge-tier reduction cannot go below — a --budget 500 request could
return the full seed payload (~119k tokens measured) with no signal of the
miss. Add budget_met + budget_estimate to the budget result and surface
them through graphifyQuery when a budget was requested.

Secondary: the tier loop estimated against the full pre-filter node set,
so a tier removal that already satisfied the budget (once its orphaned
nodes were excluded) still triggered the next, higher-confidence tier
drop. Recompute reachability and the estimate after each tier and break
as soon as the pruned result fits.

Adjacent same-class instance from pre-submit review: the CLI forwards
--budget 0 but truthiness checks silently treated it as no budget and
returned the unbounded result. Test budget presence with != null so a
zero budget is honored and reported as an unmeetable miss.

* docs(changeset): backfill PR number for #2738 fragment

* fix(#2738): estimate the payload as emitted, not a private compact form

budget_met measured a different payload than the caller receives. The
estimator serialized a compact `{nodes, edges}`, while output() emits the
whole response pretty-printed (2-space indent, plus the term/total_*/trimmed
wrapper keys). Measured on the repo's own SAMPLE_GRAPH fixture: reported 183
tokens against 302 actually emitted — 1.65x — so `--budget 200` returned
budget_met: true while handing back 302 tokens. That is worse than the old
silent miss: an automated consumer stops checking a signal that is
confidently wrong.

Fix the basis rather than the number:

- io.cts gains serializeForOutput(), the single definition of the wire form.
  output() now calls it, so the estimator and the emitter cannot drift on
  indentation or shape. Pure extraction; output()'s behaviour is unchanged.
- graphify builds its response through one buildQueryResponse() used by both
  the emitter and the estimator, so the estimate describes exactly the bytes
  returned.
- The tier loop estimates on that same basis, so it keeps trimming until the
  real payload fits instead of stopping at a smaller internal measure. This
  makes budget_met === (budget_estimate <= budget) true by construction.
- Drop the module-private chars/4 helper for prompt-budget's estimateTokens —
  the repo's single token scale, per the rule phase-estimation.cts documents.

budget_estimate is self-referential (its own digits are part of the emitted
bytes), resolved by iterating to a fixed point; the sequence only ever grows,
so it settles in a couple of passes and errs toward over-reporting.

Tests pin estimator to emitter so this cannot silently re-diverge if output()
ever changes its indentation. Both new tests fail against the pre-fix source.

* test(#2738): property-test the budget-limit and reporting contract

RULESET.TESTS.property-based-testing names budget-limit contracts, and this
module is the literal case: #2819 turns it into a *reporting* contract, which
is what properties express well. Five invariants over arbitrary small graphs
and any budget >= 0:

- budget_met === (budget_estimate <= budget)
- budget_estimate === the tokens actually emitted
- the seed set is a floor the reduction never goes below (the changeset's
  "seeds are a floor" claim, previously asserted only for one hand-built
  fixture)
- total_nodes/total_edges match the returned arrays
- a larger budget never yields a smaller payload

Two notes on what these do and do not prove. The emitted-payload property
fails against the pre-fix source; the budget_met/budget_estimate agreement
property does NOT — pre-fix both derived from the same wrong number, so it
is a contract guard, not a regression proof.

Monotonicity is asserted over the payload (node/edge counts), not over
budget_estimate: the estimate measures emitted bytes exactly, and budget_met
renders as "false" (5 chars) or "true" (4), so an identical payload can
measure one token larger when the budget is missed. That is the estimate
being honest, not a monotonicity break.

The generator seeds on `label` — seedAndExpand matches label/description,
never id/name, and a fixture that gets this wrong expands to nothing and
passes vacuously.

* test(#2738): pin the budget boundary and label the forward-guard test

Two test-coverage gaps from review.

RULESET.TESTS.boundary-coverage wants limit-1 / limit / limit+1. The added
tests used 1, 0, 200, 100000, 50 — all far from the decision point. The
branch that matters is `estimate <= budgetTokens`, so the input that decides
it is budget === estimate exactly: that is the one value where an off-by-one
in the comparison flips budget_met, and nothing else in the suite would
catch it. Pinned at the limit and either side of it.

The "omits budget fields when no budget was requested" test passes unchanged
on next — graphifyQuery never set those keys before the fix, so both
assertions already held. It has value as a forward guard against the spread
leaking budget fields, but it is not failing-first and should not be counted
toward RULESET.TESTS.regression-must-fail-first. Said so in a comment, so a
later reader does not mistake it for the regression proof.

* fix(#2738): keep a non-finite budget out of the budget path

Switching `!budgetTokens` to `budgetTokens == null` widened the internal
contract to admit NaN, where every `estimate <= NaN` is false: the loop
strips all three tiers and returns a seeds-only payload that is
indistinguishable from a legitimate aggressive trim.

Unreachable through the CLI — graphify-command-router rejects a non-numeric
--budget with makeInvalidArgs before graphifyQuery is called — so this is
hardening, not a live defect. It is still worth guarding: graphifyQuery and
applyBudget are module-level entry points a future caller could reach
without the router's validation, and the failure mode is silent.

Number.isFinite also routes Infinity to the no-budget path, deliberately: an
unbounded budget is not a budget, and parseInt cannot produce one anyway.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:36:48 -04:00
Adnan
fc3fde05ee feat(#2646): surface unresolved deferred-items.md at milestone close (#2983)
* feat(#2646): surface unresolved deferred-items.md at milestone close

auditOpenArtifacts scanned eight categories; deferred-items.md was not
among them. #2287 made that file readable at the PHASE boundary
(audit-uat, /gsd-progress check 7), but one boundary up it stayed
invisible — and phase directories archive to milestones/vX.Y-phases/ by
default (#1871), so an out-of-scope discovery a phase agent correctly
recorded rather than fixed left the live tree at milestone close having
never reached the [R]/[A]/[C] prompt that exists to catch exactly this.

Adds deferred_items as a ninth scanner plus its count, its items entry
and its report section. The workflow needed no change: complete-milestone
branches on "any section with count > 0", so the new category flows
through the existing prompt.

The resolved/unresolved predicate is NOT reimplemented. uat.cjs already
exports parseDeferredItems, which owns the parsing rule (entries under a
`## Deferred Items` level-2 heading, else the whole file fail-safe;
RESOLVED only on an explicit case-insensitive `status: resolved` field).
The scanner requires it lazily, inside the scan, so audit-command-router's
property that a route never loads the module it does not need is
preserved. Two readers of one file sharing one predicate is the point —
duplicating the inequality is how they drift into disagreeing about what
"open" means.

Regression test proves fail-first: 9 of its 10 cases go red against the
pre-change tree. The tenth is the deliberate no-regression boundary (a
clean tree emits no section) and is green both ways.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4

* docs(#2646): document the pre-close artifact audit and its nine categories

The /gsd-complete-milestone entry did not mention the audit at all, so
the gate that can stop a close was undocumented — and this change adds a
category to it. Tabulates all nine with their source artifact and what
makes each one "open", plus the [R]/[A]/[C] outcomes.

Also disambiguates the one genuinely confusing thing: the per-phase
deferred-items.md scanned here is NOT the `## Deferred Items` section the
[A] path writes into STATE.md. Same name, different artifact, opposite
ends of the flow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4

* chore(#2646): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:27:33 -04:00
0xdhx
c61dd49d95 enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes

Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks
(decision: 'block', exit 2) a whole-file Write collapsing a curated
.planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md,
STATE.md) below 40% of its on-disk line count. Files under 40 lines
are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message)
bypasses for legitimate milestone resets.

Fix 3 of #973 — the only defense independent of per-agent tool config.
Registered on the Claude plugin surface (hooks.json), settings-json
runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi
spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and
INVENTORY regenerated; regression tests negative-controlled (16/16
RED with the hook absent, 16/16 GREEN with it present).

* chore(#2255): backfill changeset pr number to 2301

* enhance(#2255): address review — fail-closed reads, typed block output, registration, property test

Review fixes for trek-e's CHANGES_REQUESTED on PR #2301:

- Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES
  (no-shipping-drift test).
- Blocker 3: update the always-on hook enumerations in ADR-766 and
  CONTEXT.md from six to seven.
- Major 4: fail CLOSED on non-ENOENT read errors — only a missing file
  (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a
  typed readError field and the override still honored. Tested, with a
  negative control against the pre-fix hook.
- Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES
  budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the
  floor; sub-floor always exempt), boundary examples pinned.
- Major 6: block output now carries typed oldLines/newLines/
  overrideEnvVar fields; tests assert on those instead of regexing the
  free-form reason string.
- Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS
  bypass on macOS/Windows); limit+1 boundary tests added for both the
  floor and the ratio.

* enhance(#2255): engage the write guard on Kimi's native payload shape

The guard shipped with Claude-vocabulary checks (tool_name 'Write',
tool_input.file_path), which #2304 showed leaves a guard dormant on
Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli
forwards its native payload verbatim — tool_name 'WriteFile' (bare or
module-qualified) and tool_input.path per its tool schemas
(src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown
name, and exited 0.

Apply the same per-guard normalization PR #2326 gives the three
sibling guards (name + field mapping, inlined — hook scripts stage as
standalone files), and write the block reason to stderr as well as
stdout JSON: Kimi feeds stderr, not stdout, back to the model on
exit 2, so a stdout-only reason blocks without telling the model why
or naming the documented override.

Regression tests pipe Kimi-shaped payloads (engage, qualified-name,
stderr-reason) plus exemption pins (StrReplaceFile stays out of scope
by design; non-curated paths pass) — verified red against the pre-fix
guard, green after.

* enhance(#2255): rebase onto next; regenerate golden-parity fixtures

* enhance(#2255): wire the escape hatch into complete-milestone's reorganize step

Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP
reorganize — the tree's only legitimate milestone reset and the exact caller
GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the
rewrite through a shell write with the hatch set on the command (a hook
inherits the runtime env, so a bare Write cannot carry a per-step override),
and a binding test derives the env var name from the guard's typed output and
asserts (a) the workflow step sets it and (b) the guard passes the identical
catastrophic payload under it — so the next complete-milestone.md edit cannot
silently re-break the wiring.

* enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload

Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string
reconstruction were unreachable-by-effect — the guard exits 0 for any
tool_name !== 'Write', so nothing ever read the fields they set, leaving
guaranteed-surviving mutants against the Stryker bar. The map now carries
only WriteFile -> 'Write'; the StrReplaceFile exemption test message states
the fall-through it actually exercises.

* enhance(#2255): review minors — American spellings; writeSync before exit(2)

Minor 1: normalised/normalise -> American house style. Minor 2: the two
block paths wrote stdout+stderr via async pipe writes then exit(2) —
async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block
payload durable.

* enhance(#2255): assert stderr equals the typed reason, not raw prose

Minor 3: the last raw-text match in the suite pinned override-name prose on
stderr. The contract is "stderr carries the reason Kimi feeds back" — now
asserted as stderr non-empty and byte-equal to the parsed stdout.reason.

* enhance(#2255): bind the write-guard's Kimi normalization into the parity test

Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with
nothing binding it. This extends PR #2326's kimi-guard-normalization-parity
test (same path and helpers, authored as a superset so either merge order
resolves cleanly): sibling byte-parity is existence-gated zero-or-all —
trivially green until #2326 lands, full-strength after — and the write-guard
copy is bound semantically (map is the value-inverse of convertKimiToolName;
the Kimi name for Write must map, or the guard is dormant on Kimi; the
path -> file_path half must be present). Byte-parity is deliberately not
asserted for this copy: it legitimately omits the Edit-class mapping
(Major 1 — dead code in a Write-only guard).

* enhance(#2255): refresh golden-parity fixtures for revised guard + workflow

* chore(#2255): regenerate golden fixtures after rebase onto next

The committed fixture hashes were generated against a tree predating
next's latest 11 commits, which independently modified the same
install-parity surface. Rebased onto next and regenerated with
`npm run gen:golden`.

Verified: against upstream/next the regenerated fixtures differ by
exactly this PR's own entries -- hooks/gsd-write-guard.js (new),
hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and
gsd-core/workflows/complete-milestone.md. No unrelated drift.

* fix(#2255): regenerate workflow size baseline for complete-milestone

`complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2
review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step,
but tests/workflow-size-baseline.json was never regenerated. The
per-file workflow baseline test (issue #1074) failed on
ubuntu-latest/22 and both macOS shard 1/3 jobs.

The growth is justified: it is the escape-hatch binding requested in
review round 2 (the guard must not hard-block the tree's only
legitimate milestone reset), not incidental bloat.

Regenerated via `npm run size:baseline`; the diff is exactly the one
entry.

* chore(#2255): regenerate golden fixtures and size baseline after rebase onto next

* enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes

Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert
(no PreToolUse hook exists on Bash in this family; the write succeeded by
dodging the guard, not by the override firing) and the protection was prose.
The hatch is now a transport code consults: complete-milestone's reorganize
step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the
Write tool as the sanctioned path, and the guard — at the block point only —
verifies the sentinel is fresh (15 min) and names the pending target, then
CONSUMES it and allows that one write. Path-bound + single-use + freshness
keep it from becoming a standing unlock. The env var remains as the
interactive transport, where it can actually reach the hook.

Regression tests written first (negative control: 3 failed pre-fix): the
armed-sentinel Write passes and consumes; stale does not exempt; a token for
a different file neither exempts nor is consumed; the binding test now takes
the sentinel name from the guard's typed output (overrideSentinel), asserts
the step arms it, and asserts the step no longer routes the rewrite around
Write via a shell pipe.

Also in this commit, same file:
- m2: block emission is exception-safe — emitBlock() wraps both writeSync
  sites in their own try/catch that still exits 2, so an EPIPE can no longer
  convert fail-closed into the outer catch's fail-open.
- Header discloses the two reviewed design limits (cumulative sequential
  shrink; lexical match vs symlinked paths) per round-5 scoping.

* docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4)

USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the
runtime-hooks-surface registration comment, and the changeset now describe
both hatches — the single-use sentinel for workflow steps and the env var
for interactive use — instead of implying a per-step env can reach a hook.
The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)`
parenthetical the repo's fragments use (round-5 m4).

* chore(#2255): regenerate derived families on the rebased tree (full sweep)

Full generator sweep after rebasing onto next @ the body-parser-patched
lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every
regen delta verified to be either a PR-owned entry (gsd-write-guard.js,
complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's
committed value for entries our arbitrary-side conflict resolution had left
stale (all 18 runtime fixtures checked mechanically).

* test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync

The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY
retry budget; the sentinel disarm now rides it like every other teardown.

* chore(#2255): regenerate derived families after rebase onto next

Full sweep on the rebased tree (build -> gen-inventory-manifest ->
gen:golden -> size:baseline). Every delta is either a PR-owned entry
(hooks/gsd-write-guard.js, its registration surfaces
hooks/managed-hooks-registry.cjs and the two plugin buses,
gsd-core/workflows/complete-milestone.md) or exact convergence to
next's committed value across all 18 runtime fixtures.

* chore(#2255): regenerate derived families after rebase onto next @ a5180d96

Rebase onto current `next` (a5180d96) resolved 12 conflicting
golden-install-parity fixtures; all regenerated via the full generator
sweep (build, gen:golden, size:baseline) rather than a single generator.

`lint:generated-sync` reports every generated artifact in sync. All 45
differing fixture keys and the single workflow-size-baseline entry map
to files this PR actually touches; no foreign drift.

* fix(#2255): remove the stale unguarded reorganize_roadmap step (round-8 blocker)

complete-milestone.md carried a second ROADMAP-collapsing step,
`reorganize_roadmap`, distinct from the sentinel-armed
`reorganize_roadmap_and_delete_originals` this PR wired. It is a vestige
of the pre-archive-then-reorganize design: it sits BEFORE
archive_milestone, so executing it as written would collapse ROADMAP.md
before the archive snapshots the full phase detail — and its Write is
exactly the shape gsd-write-guard hard-blocks, with no hatch armed. The
file's own success criteria describe only one reorganize outcome
(Backlog-preserving, overwrite-in-place — the later step's properties),
and archive_milestone points forward to "the reorganize step".

Removed rather than wired, per the round-8 review's confirm-and-remove
option. A new binding test asserts the sentinel-armed step is the ONLY
reorganize step in the workflow, so an unguarded collapse step cannot be
silently reintroduced (negative-controlled: fails against the pre-fix
tree). Golden-parity fixtures and the size baseline regenerate for the
shrunk file; every changed fixture key is complete-milestone.md's own.

* test(#2255): document why the read-error injection is a path collision, not an fs monkeypatch

Round-8 nit: the non-ENOENT tests inject via a directory-at-target-path
collision instead of the repo's fs-method monkeypatch pattern. That is
deliberate, not drift — runHook exercises the hook as a spawnSync child
process, so an in-process fs.readFileSync patch (the pattern the cited
siblings use on require'd, in-process code) can never reach the code
under test. Record the reasoning at the injection site.

* chore(#2255): regenerate derived families after rebase onto next @ 0d08c320

Rebase onto current next (0d08c320) for the CONFLICTING/DIRTY state. All 32
conflicts were generated artifacts (19 golden-install-parity, 12 install-tree,
workflow-size-baseline); resolved arbitrarily and regenerated via a full
generator sweep (build, gen:golden, size:baseline, gen-inventory-manifest)
rather than hand-merged. No source conflicts.

Regen diff verified against the PR's changed-file set: 7 distinct differing
keys, all PR-owned (gsd-write-guard.js, managed-hooks-registry.cjs,
plugins/gsd-core.js, complete-milestone.md, and their .kimi mirrors).
lint:generated-sync clean.

* chore(#2255): regenerate derived families after rebase onto next @ 9138271b

Conflict set was 20 paths, every one a generated artifact, zero source
conflicts — resolved arbitrarily during the replay and regenerated here,
per the maintainer's round-9 recipe (never hand-merged).

Generator sweep (not just gen:golden): npm run build, gen:golden,
size:baseline, gen-inventory-manifest, gen:registry. INVENTORY-MANIFEST
came back byte-identical, so the merged value was already correct.

Regen diff verified == PR-touched entries: every differing leaf key
attributes to a file this PR changes (complete-milestone.md, the guard,
managed-hooks-registry.cjs, both plugin buses) — derived from the
changed-file set, not a feature-name filter.

* fix(#2255): port the #2595 Kimi-payload hardening the base landed for the sibling guards

Self-found via the round-10 base sweep: a8b40fa5 (#2547/PR #2595) hardened
gsd-prompt/read/worktree-path-guard's normalizeKimiPayload after this PR was
authored, and this guard carried all three pre-fix shapes:

- `path` now wins outright over a model-supplied `file_path` — kimi-cli's
  WriteFile schema carries no `file_path`, so a payload pairing a curated
  `path` with a spurious `file_path: ""` had the guard reading '' and
  exiting 0 while kimi-cli wrote to `path` (one-key bypass, no crash needed).
- null/primitive payloads pass through deliberately instead of throwing into
  the fail-open catch.
- the `file_path` read is typed (`typeof === 'string'`), closing the
  crash-to-allow on `[]`/`{}` — this is what the base-added
  kimi-guard-typed-payload-reads scan flags.

Negative-controlled: the shadowing payload exits 0 against the pre-fix guard
and 2 against this one.

* fix(#2255): round 9 Major 2 — stop printing the sentinel recipe in the agent-facing denial

The block message ended by telling the blocked agent exactly how to arm the
single-use sentinel — handing the #973 actor a mechanical self-authorization
in the moment it is blocked. The recipe is gone from the prose; the
corrective advice and the env override's name stay (the latter is a #2255
acceptance criterion, and a per-step env prefix cannot reach a hook anyway),
and the typed overrideSentinel field stays for the binding tests. The hatch
remains documented in USER-GUIDE.md and complete-milestone.md, where humans
and the workflow engine read.

* fix(#2255): round 9 Minors 1-2 — realpath-resolve the target before the curated match; disclose the /i Linux cost

Minor 1: a Write to a non-curated path that symlinks into a curated file was
not matched while writeFileSync followed the link — the target is now
realpath-resolved before the curated match (ENOENT keeps the lexical
resolution so new-file Writes still pass; any other realpath error falls
through to the read, which fails closed). Negative-controlled: the symlink
payload exits 0 against the pre-fix guard, 2 against this one. Test skips on
win32, where symlink creation needs privilege.

Minor 2: the header's design-limits block now names the unconditional /i
cost on case-sensitive Linux (a genuinely distinct .planning/roadmap.md is
also treated as curated) next to the stateless limit, and drops the closed
symlink limit.

* test(#2255): round 9 Minors 3-4 — CRLF counting pin + a passing Write leaves a fresh sentinel unburned

Minor 3: countLines' split('\n') is CRLF-safe for a count (the \r rides
along), confirmed by trace in the review — this pins it against this repo's
recurring CRLF regressions, on both sides of the compare and at the 40%
boundary.

Minor 4: consumeSentinelFor runs only after the ratio check would block, so
a within-tolerance Write never burns the workflow's token — true by
construction, previously un-asserted.

* fix(#2255): round 9 Major 3 — correct the stale env-var line in archive_milestone's summary

complete-milestone.md's "After archival" bullet still said the reorganize
happens "under GSD_ALLOW_PLANNING_SHRINK=1" — the wording from the round-2
design this PR's own history rejected in round 5 (a per-step env var cannot
reach a hook; setting it in a Bash step silently does nothing). It now points
at the sentinel mechanics the reorganize step actually documents, matching
that step and USER-GUIDE.md.

* docs(#2255): round 9 Major 1 — user-facing docs state the stateless per-Write limit

The changeset and USER-GUIDE described the guard as covering "catastrophically
shrinks" with no caveat, while the stateless design was disclosed only in the
hook header — an operator reading the shipped docs would conclude iterative
erosion is covered. Both surfaces now state the per-Write comparison and the
erosion non-goal explicitly, in line with what the guard does.

* fix(#2255): round 10 Blocker 1 — arm the hatch in the roadmapper's curated writes

gsd-roadmapper Step 7 Writes both .planning/ROADMAP.md and .planning/STATE.md
wholesale, and /gsd:new-milestone spawns it against the OUTGOING milestone's
files. new-milestone's `phases.clear` archives phase DIRECTORIES, never
ROADMAP.md, so nothing compacts it first, and no ordering rule forces
/gsd:complete-milestone to run before /gsd:new-milestone.

Measured against the shipped hook at the #973 file size (292 lines): a new
4-phase roadmap lands at 18.2% and an 8-phase one at 31.8% — both blocked; only
a 12-phase replacement (45.5%) clears. So the collapse is reachable on a
first-party flow, which is the failure mode #2255 para 4 says the guard must not
have.

Arms the single-use sentinel immediately before each Write, one per target
(the token is path-bound and single-use, so one arming cannot cover both), and
gates each arming on `[ -f ]` so the /gsd:new-project path — where the guard
already exempts via ENOENT — strands no unconsumed token.

Binding test derives the sentinel name from the guard's own typed output, so a
rename on either side fails the suite instead of silently unwiring the hatch.
Negative-controlled: 2 assertions fail against the pre-fix roadmapper.

* docs(#2255): round 10 Major 2 — scope the claim to the guarantee that ships

The sentinel is a plain file, so an agent that would reason past an advisory
can arm one with a single Bash call it is already permitted to make. The
"only defense independent of per-agent tool config" framing outran that, and
the changeset was on its way into CHANGELOG.md.

Retitles the claim on all three surfaces (changeset, guard header, USER-GUIDE)
to what the guard actually delivers: it blocks accidental and single-shot
collapse and is not a defense against a determined agent; what it converts is
"ignore a sentence" into "take one deliberate, path-bound, single-use,
auditable action".

Pinned by test on the DURABLE surfaces only — the guard header and USER-GUIDE.
The changeset fragment is deliberately not pinned: it is consumed at release,
so a test reading it would start failing the moment the release lands. The
bound-statement assertion normalizes comment markers and whitespace first, so
it pins the claim rather than the paragraph's line wrapping.

Negative-controlled: both assertions fail against the pre-fix surfaces.

* test(#2255): acknowledge the roadmapper growth from the round 10 Blocker 1 wiring

The emitted-attribution gate (#2719/#2767) flags gsd-roadmapper.md growing 1130
bytes without an acknowledgment. The growth is the Blocker 1 sentinel wiring
plus the rationale a future editor needs to keep it, so it gets an ack fragment
rather than a silencing regen — the gate's own message is explicit that there is
nothing left to regenerate.

Fragment is PR-scoped (2301-…) per the gate's naming instruction, and uses the
plain-string reason form the shipped fragments use.

Verified against the TRUE upstream tip, not the fork's origin/next: a stale
origin made this same gate report unrelated phantom drift (1 emitted path + 6
grown files + 5 stale acks) that vanishes when GSD_EMITTED_BASE is pinned.

* test(#2255): renumber the roadmapper PROSE_ALLOWLIST pin after the Step 7 wiring

CI red on shard 2/3, all four platforms. The #2751 gate keys PROSE_ALLOWLIST on
{file, line}; the Blocker 1 wiring added 18 lines above the allowlisted
parenthetical in agents/gsd-roadmapper.md, moving it 624 -> 642. Both halves of
the gate then fired: the moved line reads as a new offender, and the stale
entry no longer matches anything.

Line content at 642 is byte-identical to what the entry describes — a
descriptive "e.g." naming SDK queries a user could run — so this is a
renumber, not a re-classification.

Swept the defect class rather than the instance: agents/gsd-roadmapper.md is
the only line-pinned reference to any file this round changed.

Negative-controlled: both assertions fail against the un-renumbered allowlist.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:19:49 -04:00
0xdhx
cc3ee301a7 fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root

installSharedHooksBundle wrote `{"type":"commonjs"}` over
<configRoot>/package.json unconditionally — no existence check, no merge,
no backup — on every install and every /gsd-update re-install. On the 11
affected runtimes that file is often user-owned; on OpenCode and Kilo it is
the documented place to declare local-plugin npm dependencies, so a user's
name/type/dependencies/scripts were destroyed on each run.

The uninstall path already read the file and unlinked it only on an exact
content match. That asymmetry was the defect: the discipline existed in the
codebase, it just was not applied on the write side.

Move the marker into the directories GSD creates and fills with its own .js
files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the
nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop
writing the config root entirely. New src/commonjs-marker.cts owns the marker
string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed
on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker,
so install and uninstall cannot drift apart again.

Nothing else depended on the config-root marker: package identity is baked at
build time (#378/#498) and version resolution prefers gsd-core/VERSION and
already tolerates a missing root package.json (#1383) — Codex has installed
without one all along. A package.json in plugins/ or extensions/ is inert to
plugin discovery, which globs *.{ts,js} only (see installer-migration 006).

Uninstall retires the pre-fix config-root marker, so upgrading users are
cleaned up on removal, and still never touches a file it did not write.

* fix(#2544): point the changeset fragment at the filed PR

The fragment's `pr:` field is only knowable after `gh pr create` returns.

* fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set

bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the
linted source), so it belongs in the ADR-457 ignore list like its siblings.
Clears the lint-tests no-var failure and the repo-invariants
"linted xor ignored" migration-state test.

* fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root

The UPGRADE 1 test still asserted the pre-#2544 marker location
(~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the
directory GSD itself creates — matching the updated golden-install-parity
and install-tree fixtures. Also asserts the root marker is NOT written.

* fix(#2544): make the CommonJS marker write path non-fatal

Review round 2, Major 3 + Minor 1 + the stagedHooks nit.

ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site
caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole
install with a raw stack trace. Every other marker interaction in the module is
best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker
swallows read failures — and this was the write path, i.e. the one most likely
to fail on a locked-down config dir. It now returns a new 'failed' outcome and
both call sites warn and continue.

Sibling found while sweeping for the same defect class: fs.mkdirSync sat
OUTSIDE the try block, so an unwritable parent threw past the guard entirely.
Creating the directory is the same environmental hazard as writing into it, so
it moved inside.

Also in this file:

- The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks
  alone. stagedHooks is computed from the SOURCE listing before the copy loop,
  so it stays true when the copies land but verifyInstalled() then fails —
  marking a hooks/ GSD did not successfully populate claims an ownership the
  install did not earn.
- The uninstall rmdir of the native plugin dir is gated on GSD having actually
  removed something from it. Hoisting it out of the adapter-exists guard (so
  the marker-only case could prune) had silently widened it into deleting a
  user-created but empty plugins/ or extensions/ dir — the same "don't touch
  territory GSD didn't fill" principle this issue is about, inverted.
- Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the
  same call site that writes its replacement. That path is outside kimi's
  configDir, so installer-migration 007 structurally cannot reach it.

* fix(#2544): retire the stale config-root marker via installer-migration 007

Review round 2, Major 1 — the PR's headline claim was false for existing
installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale
{"type":"commonjs"} at the config root, so their config root stayed pinned to
CommonJS and their dependency manifest stayed gone until they uninstalled.

The migration is unusual in one way, and it is the part worth reviewing: the
config-root marker was never recorded in gsd-file-manifest.json (writeManifest
records hooks/, agents/, commands/, scripts/ and the native plugin, never a root
package.json), so classifyArtifact answers 'unknown' for it and the planner's
own guard downgrades a remove-managed on an 'unknown' classification to
preserve-user. 007 therefore supplies the "purpose-built detector for an old
GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions —
exact content match, the same predicate removeCommonJsMarker has always used —
and declares the resulting classification on the action. A package.json with any
other content is left untouched, and there is deliberately no backup-and-remove
branch: a non-matching file here is not a patched GSD artifact, it is somebody
else's file.

Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`:
validateStringArray requires the field to be non-empty WHEN PRESENT, while the
runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan
time and the migration never runs. The metadata test pins this.

Kimi is a deliberate carve-out, named in the migration's own header: its marker
lived at ~/.kimi, outside kimi's configDir, and migration relPaths are
structurally confined to configDir. It is retired by the installer instead.

Registration: shipped-migrations table, .gitignore for the emitted .cjs, the
EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not
copied from migration 006 by rote — 006 needs no entry because it imports
nothing, while 007 imports node builtins, so tsc emits its __importDefault
helper and the `var` in it trips no-var. This is the same lint gate that made
round 1 red.

* test(#2544): fault-injection and multi-runtime marker coverage

Review round 2, Major 2 + Minors 4 and 5.

Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and
the suite had no fs monkeypatching at all. Every branch now covered is one whose
doc comment claims it as the module's safety posture:

- classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule),
  with an ENOENT control alongside it so the test discriminates rather than
  just asserting one side
- classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never
  downgrades to the permissive answer) — the fixture's bytes are exactly GSD's
  marker, so the test fails if the code ever answers on content it could not read
- a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already
  covered with a real symlink, the directory case needs no injection at all)
- the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx'
- the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the
  mkdirSync that used to sit outside the guard
- removeCommonJsMarker unlink throw -> false

These save and restore fs methods in `finally` rather than using chmod 0o000,
which does not fault under root and would pass vacuously in root Docker and CI.

Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both
kimi locations now have behavioral coverage, install and uninstall, each paired
with a user-authored-file case proving GSD leaves it alone.

Minor 5 — the stagedHooks gate had no assertion behind its stated reason.
A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime
that declares skipSharedHooksInstall and asserted to stay marker-free, with its
user content intact.

Also regression-tests the uninstall rmdir gate from the previous commit: an
empty plugin dir GSD removed nothing from must survive.

* docs(#2544): correct stale marker prose, register the module, document the trade-off

Review round 2, Minors 2, 3 and 6.

Minor 2 — six files asserted the installed ROOT ships the synthetic marker.
None was load-bearing (all three walk-up consumers are VERSION-first with
try/catch and the marker never carried a `version`), but ADR-457:52 is the
rationale for keeping a generated module, so a future reader would mis-derive
the constraint from it. Each site is corrected to what is now true: the
installed tree carries no package.json with a .name at all, because the only
ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own
directories.

Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js
and the platform-gate test both described `require('../package.json').name`
resolving to undefined; post-#2544 that require does not resolve at all, so the
history is kept accurate and the present-tense claim corrected rather than just
moved. And src/runtime-artifact-conversion.cts described the no-root-package.json
case as Codex-only — it is now every runtime, which strengthens that comment's
own argument for lazy resolution. The generated .cjs sibling needs no edit: it
is gitignored build output, not a tracked file.

Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer
module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added,
including the fail-closed posture and the never-throws contract.

Minor 6 — the plugins//extensions/ marker shadows the config root for all .js
siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is
exactly what #2544's Fix section prescribed and it is disclosed in the PR body,
but the PR body is not documentation. It now lives in the OpenCode section of
docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than
a pure improvement, with the .ts mitigation and a fallback for ESM plugins.

* test(#2544): attribute the CommonJS marker in the emitted-provenance rules

The differential emitted-attribution gate (#2723, landed on `next` after this
branch was cut) went red on the macOS shards once this PR rebased onto it. Two
distinct causes, both real gaps rather than noise:

1. `plugins/package.json` and `extensions/package.json` matched NO rule — the
   `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an
   unattributed emitted family.
2. `hooks/package.json` fell through to `hooks-built`, which attributes an
   emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no
   `hooks/package.json` in the repo, so it resolved to a nonexistent path.

Cause 2 is exactly the failure already documented three lines above it for
Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix
follows that precedent rather than inventing one: `package.json` is excluded
from `hooks-built` the same way, and a dedicated `commonjs-marker` rule
attributes the family across all four roots it can appear in (both hooks roots
plus `plugins`/`extensions`) to the sources that actually emit it.

Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for
a one-off ripple and goes stale by design — the gate fails a stale ack precisely
so it cannot pre-clear the next change on that path. These markers are a
permanent part of the emitted tree from #2544 onward, so they need standing
attribution.

Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3
provenance errors + 12 unattributed paths before, 35/35 green after.

* fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker

#2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned
in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via
dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's
module on the two properties that matter:

  - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a
    DANGLING one, so a dangling package.json symlink classified as absent and
    the write went straight through it. Demonstrated: against the pre-fix copy,
    ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json
    symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory.
  - create: a plain writeFileSync leaves the classify->write window open, where
    commonjs-marker creates with flag:'wx' (O_EXCL).

Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's
own docstring already claimed was the single place these rules are enforced.
Exported signatures are unchanged (still boolean), so bin/install.js and the
#2717 tests are unaffected.

The new subtest is the only coverage that fails if the duplicate is ever
reintroduced — the two implementations agree on every non-adversarial input, so
the existing suites pass against both.

* test(#2544): pin the stagedHooks gate on zcode, not windsurf

The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall
kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the
marker was correctly absent.

#2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via
dedicated paths and get the marker beside those scripts. Measured on this tree,
windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning
behaviour that is now wrong, not the gate it was written for.

ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin
surface to spawn hooks, so GSD stages no .js there by either route (measured: 0
staged, no marker). The property under test is unchanged — a user-created hooks/
directory GSD never fills stays marker-free.

* test(#2544): use the shared cleanup helper in the migration test

Addresses the review's Major 1. The suppression's stated reason — "no helpers
import available" — was not correct: tests/helpers.cjs exports cleanup, and the
other test file added in this same PR imports it (tests/commonjs-marker.test.cjs).

The local reimplementation dropped two protections that are live on this repo's
windows-latest lane: the CWD guard (Windows cannot remove a directory that is the
current working directory) and the 20 x 250ms retry budget that absorbs the
deferred-scan handle Windows Defender holds on newly-written files.

Local function and suppression both removed; local/no-raw-rmsync-in-tests now
passes without one.

* test(#2544): expect hooks/package.json for the #2717 runtimes

The fresh-install contract table predates #2717, which stages cursor/windsurf/
codex .js hooks via dedicated paths and writes the CommonJS marker beside them.
All three therefore now receive hooks/package.json legitimately.

Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with
the marker; cline/copilot/trae/zcode stage none and get none, so their contracts
are unchanged.

* fix(#2544): gate the #2717 marker writes on having staged something

The three dedicated marker writers #2717 added ran unconditionally. Each one
mkdirs hooks/ up front and stages its scripts conditionally on the source
existing, so with an absent or empty hook source they created a directory,
filled it with nothing, and marked it as GSD's anyway.

That is the same write-into-someone-else's-territory this issue is about, and
installSharedHooksBundle already guards the identical case with `stagedHooks`.
The dedicated paths now carry the matching gate:

  - cursor / windsurf: `installedScripts.size > 0`
  - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that
    hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY
    entry landed.

Covered for cursor and windsurf by driving each writer against a src tree whose
hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its
trigger state needs a package tree where hooks/dist exists but holds none of the
allowlist, which is not constructible from a real checkout.

* test(#2544): scope the commonjs-marker sources per root

The rule declared one flat source list for every marker root, so
`extensions/package.json` was attributed to runtime-hooks-surface.cts (which
never writes there) and `.kimi/hooks/package.json` to install-engine.cts.

That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source,
so a flat list containing bin/install.js let any change anywhere in that
13k-line file authorise marker drift for every root — the blanket escape hatch
this file's own agents-verbatim comment refuses for exactly the same reason.

Sources are now derived per root from ctx.rel. Note the rule ctx is
`{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent
every path down one branch silently.

* test(#2544): state precisely what the zcode assertion pins

The comment claimed the test pinned installSharedHooksBundle's `stagedHooks`
gate. It does not, and neither did the windsurf version it replaced: zcode
declares skipSharedHooksInstall, so the outer guard skips that helper entirely
and the gate is never evaluated. The test passes on the runtime exclusion.

What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays
marker-free — is still worth having, and is what the review asked for. The two
`staging zero hook scripts` tests are the ones that pin a real staged-nothing
gate. Comment corrected rather than left implying coverage that is not there.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:00:23 -04:00
Tom Boucher
b62589b73f fix(#2840): exclude runtime from defaults.json spread into project config (#2985)
* test(#2840): add regression for runtime poisoning from defaults.json

* fix(#2840): exclude runtime from defaults.json spread into project config

runtime is host-specific (written by whichever installer ran last). On a
machine with 2+ runtimes, it poisons every new project config — e.g. a Codex
install's runtime:'codex' leaks into Claude Code projects. Now excluded from
the userDefaults spread, mirroring the resolve_model_ids guard (#2297).

* chore(#2840): add changeset fragment

* fix(#2840): add new test file to lint-test-file-count allowlist

* chore(#2840): backfill changeset PR number 2985

---------

Co-authored-by: sim <sim@local>
2026-08-01 16:33:43 -04:00
Tom Boucher
33fd203ccd test(#2966): loop QA walk — drive real scenarios across all five loop steps (#2976)
* test(#2966): loop QA walk — drive real scenarios across all five loop steps

Adds a headless walk that carries accumulating project state across
discuss -> plan -> execute -> verify -> ship against one temp project,
layered over the existing tests/helpers.cjs runGsdTools substrate.

Findings carry severity. A violation breaks a stated contract and fails
the build; a smell is legal under today's implementation but structurally
questionable, is recorded, and never reddens CI. Without that split an
oracle set derived from current behavior can only ever confirm current
behavior -- the harness could not say "this works and is still wrong".

The end-to-end test asserts the walk produces at least one smell: a QA
harness that reports nothing on a first run against a real engine is far
more likely mis-specified than the engine is perfect. It deliberately does
not pin smell ids or counts, which would re-freeze current behavior.

First run against the real engine: 0 violations, 3 smell classes --
init returns agents_dir outside the project tree; smart-entry emits prose
unconditionally so routing cannot be asserted; state-snapshot reports a
missing STATE.md through a payload key with exit 0.

Also fixes tests/fixtures/index.cjs: createFixture with git:true and
planning:false staged nothing, so the commit failed with "nothing to
commit". That combination was unreachable until greenfield needed it.

Extends RULESET.TESTS.feedback-loop-convergence from estimation to the
loop itself. Design lock: docs/adr/2966-loop-qa-walk.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): wire fault injection, make perturbations discriminating

Independent review found tests/qa/mutations.cjs entirely unwired: 462
lines exercised only by their own unit tests, with no mutation hook in
the scenario DSL and no scenario applying one, while the module header
and the ADR described fault injection in the present tense. Dead code
documented as live.

Adds a `mutate` step field, three perturbation scenarios, and a wiring
detector: a self-test scenario whose expectations are known-false and
which MUST fail. The previous anti-vacuity check asserted only that the
walk produced a smell, which passes on well-known engine behavior
regardless of whether the harness wiring works.

First perturbation attempt produced zero signal -- progress does not
structurally parse ROADMAP.md, so a corrupted roadmap sailed through. A
perturbation that cannot fail is the same defect in a new costume.
Probes now target roadmap get-phase, and each mutated step runs a clean
baseline first so `mutationObserved` records whether the corruption
changed anything at all.

Also clears four review findings: classify() returned PROSE for exit-0
with empty stdout; `warnings` was structurally unpopulatable on the
success path (execFileSync discards it) and is now documented as
error-path-only; read-only-idempotence passed vacuously when asked to
check idempotence without the data to check it; the ADR miscounted the
oracles.

Discrimination matrix across 8 mutations x 6 commands: bom,
duplicate-phase-id and escaped-pipes are absorbed silently by every
probed surface, and progress / smart-entry / roadmap validate never
reacted to any mutation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): add path-containment guard for scenario-supplied targets

Security review found scenario-supplied paths joined to the temp project
with no containment check. step.mutate.target and agent.write keys were
validated only as non-empty strings, so a target of ../../../../etc/hosts
reached fs.unlinkSync / fs.writeFileSync / fs.symlinkSync outside the
project. The symlink mutation was worst: it read the traversed file, wrote
a sibling copy, deleted the original and symlinked it back.

Not exploitable today -- all shipped scenarios target .planning/ROADMAP.md
and scenarios are repo-committed, not runtime input. Fixed anyway: it is a
live primitive any future scenario or copied helper can reach.

Adds tests/qa/paths.cjs with resolveWithin(): rejects absolute paths, NUL
bytes and empty input, normalizes separators unconditionally, and requires
containment by path segment so a sibling like <base>-evil is not treated as
inside. Non-existent targets resolve via nearest existing ancestor rather
than falling back to a lexical compare. Scenario load now rejects traversing
or absolute targets up front.

oracles.cjs previously carried its own copy of the containment logic; both
now share paths.cjs, since a duplicated containment check is exactly the
divergence class this repo calls out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): complete trajectory corpus, report emission, boundary-aware oracle

Adds the remaining trajectories and drives all 11 mutations end-to-end.
20 scenarios, 72 steps, 0 violations, 25 smells.

Adds qa-report.json with per-step verdicts and a copy-pasteable repro
command, plus --keep / GSD_QA_KEEP=1 to preserve a failing tree. A repro
line for a tree that was not preserved is marked NOT RUNNABLE rather than
emitting a command pointing at a deleted directory.

monotonic-progress is now boundary-aware. Two scenarios had been trimmed
to stop the oracle complaining at a milestone rollover, which destroys the
signal the trajectory exists to produce. Evidence: counters legitimately
reset to zero at milestone complete, but the payload milestone_version
lags until a new ROADMAP.md is written. So the oracle now scopes by
milestone plus workstream, keeps a same-scope decrease as a violation, and
records a boundary crossing as a smell. Both scenarios walk the real
boundary again.

Standards review fixes: oracle findings now carry a structured subject so
tests assert on typed fields instead of substring-matching the free-form
detail string, resolveWithin throws a typed EPATHESCAPE error, and the
absolute-path predicate scenario.cjs had re-implemented now comes from
paths.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): fix silently-vacuous fixtures and guard the class

Every fixture carried its #2371 provenance comment BEFORE the frontmatter
block, and extractFrontmatter returns {} when anything precedes the opening
---. So every scenario reading status/phase/name was operating on an empty
object and reporting green. Nine fixtures repositioned; the comment stays,
it just moves below the closing ---.

Both UAT fixtures lacked a parser-recognized result block, so
evaluateUatPassed saw checks.length===0 and could never return passed:true.
The uat-fail-then-remediate scenario could not have proven a remediation.
Its expect block only inspected blockers, which is empty before AND after,
which is why the corpus never noticed. Both fixtures now carry real result
blocks and the scenario asserts passed and no_uat_artifacts on each side of
the flip.

The actual deliverable is the guard: a fixture-integrity block asserting
every fixture with a frontmatter shape parses to a non-empty object, that
every fixture carries its provenance marker, and that the two UAT fixtures
produce opposite verdicts through the real evaluateUatPassed. The first
guard written required --- at byte 0, which would never have fired on the
regression it exists to prevent; it was rewritten and proven by deliberately
re-breaking a fixture.

No engine defect here. no_uat_artifacts means no parsed check items, not no
UAT files, and it was reporting correctly on fixtures that had none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): make the walk report — smell ratchet, baseline, CI job

The harness computed smells into a gitignored qa-report.json that nothing
read. In CI it surfaced nothing at all: violations failed the build, but the
half of the tool that says "this works and is still wrong" was inert. A QA
tool nobody hears is decoration.

Adds a ratchet on the same idiom this repo already uses three times over
(the regression-test-name allowlist, the emitted-drift acks, the size
baseline): a committed smell-baseline.json, per-PR acknowledgment fragments
under tests/qa/smell-acks/, and a ratchet script wired into CI.

The design invariant is preserved exactly. A smell still never fails a build
on its own merits. What fails is an UNACKNOWLEDGED NEW smell -- the absence
of a decision -- leaving an author two honest exits: fix it, or record a
fragment with a real reason. An empty reason is rejected. The baseline is
shrink-only, so a fixed smell must prune its entry. Violations remain
unacknowledgeable.

Fingerprints are composed only from stable fields (oracle id, scenario,
argv, subject discriminator) -- never temp paths, timestamps or counts.
Verified byte-identical across two runs in separate temp dirs; an unstable
fingerprint would have false-positived every CI run.

CI gains a qa-loop-walk job that runs the suite and the ratchet, uploads the
report with `if: always()` (it matters most when it failed), and renders a
summary a reviewer reads without downloading anything.

Also fixes the report runner invoking main() unconditionally on require, so
importing it double-ran every scenario and clobbered its own output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): every smell terminates in a defect or a fixed detector

The baseline accepted a smell with a free-text reason. That is a mechanism
for designing smells in -- an allowlist nobody revisits. The harness is
brand new, so nothing it found is inherited legacy; every finding is a
FIRST finding. Each must now terminate in exactly one of two states:

  REAL           -> an assigned defect, entry carries the issue number
  FALSE POSITIVE -> the detector is wrong and gets fixed, never baselined

There is no third "accepted with a good explanation" state, so the ratchet
now requires a positive-integer `issue` on every entry. A reason may remain
as a human note but can never substitute. `--update` refuses to invent
issue numbers: a new smell is written with `issue: null` and a TODO, and
the next plain run rejects it, forcing triage rather than accumulation.

Working the 21 existing entries through that rule found 16 were my own
detectors being wrong:

value-hygiene (10) flagged $.agents_dir, a field whose entire contract is
to point at the install tree outside any project. Fixed with a leaf-key
allowlist of contractually-external fields, verified as the only such key
in the init payload. Genuinely unexpected out-of-project paths still smell.

monotonic-progress (6) fired on legitimate boundary crossings -- milestone
v1.0 to v2.0, workstream beta to alpha -- and on one payload carrying no
scope fields at all, where a change cannot even be known. Scope changes now
reset silently and scope-less observations are skipped. The same-scope
decrease remains a violation; that is the real invariant and is regression-
guarded.

The five survivors are real and now tracked: soft-error-exit-zero (#2980),
untyped-success (#2979). Baseline 25 -> 5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): keep the ratchet out of the tarball, unpin the qa CI job

The remote matrix returned failed -- 3 unique failures, identical on
node22 and node24, both root causes in this branch's own diff.

The ratchet lives under scripts/, which ships in the npm tarball, and it
requires three modules under tests/, which does not. In a published
install it is MODULE_NOT_FOUND at load. This is exactly the class the
#2858 guard was added to catch, and it caught it. Fixed the way #2858
fixed the same shape for its own repo-only CI script: a targeted files[]
negation, so the ratchet stays in the repo for CI and out of the tarball.
Not solved by moving or inlining the required modules -- the ratchet must
keep using the same code the harness uses, or the two drift.

Verified both directions: the script is no longer in the pack list, and
build-hooks.js, fix-slash-commands.cjs and gen-capability-registry.cjs are
all still shipped. Over-negating there would have broken installs, since
bin/install.js requires them.

The qa-loop-walk job also carried CI_REBASE_BASE_SHA copied from a
neighbouring job without the paired GSD_EMITTED_BASE, which the #2854
invariant forbids by name: diverging them makes the differential compare a
tree against a baseline from a different commit. The job runs only the qa
suite and the ratchet and invokes no emitted-attribution test, so it needs
no rebase-pinned base at all -- the step was removed rather than paired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): stop monotonic-progress going blind on scope-less payloads

The full remote suite caught a false NEGATIVE I introduced while fixing a
false positive. Silencing the boundary-crossing noise had made the oracle
skip ANY observation lacking milestone fields -- so a minimal payload like
{total_summaries: n} produced no violation at all, and the oracle stopped
catching the exact defect it exists to catch. For a QA tool that is
strictly worse than the noise it replaced.

Scope is only indeterminate when the two observations DISAGREE about
having it:

  both scoped, same scope, decrease -> VIOLATION
  both scoped, different scope      -> reset silently
  NEITHER scoped, decrease          -> VIOLATION   (the regression)
  mixed                             -> skip the comparison

Implementing the mixed case surfaced a second blind spot: advancing the
reference point on a skipped pair lets a scope-less observation sitting
between two same-scope ones mask a real decrease. Mixed now leaves the
reference untouched. All four branches carry explicit coverage; only one
did before, which is why this shipped.

The self-test that failed was right and the code was wrong, so the code
moved. Corpus behavior is unchanged: still 5 smells, 0 new, 0 stale, 0
violations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:13:28 -04:00
Tom Boucher
628648d63a chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)
* fix(#2931): preserve protected regions and cap emitted per-runtime bytes

Route every runtime brand swap through applyClaudeCodeBrandSwap so
"Claude Code" survives verbatim inside <runtime_compatibility> regions
(#2284b). The fix existed only in bin/install.js's local copies; the
src/*.cts exports still used a naive replace, so binding install.js to
the single source -- as this phase does for the Windsurf family --
would have silently regressed those runtimes. A table-driven parity
guard now covers all nine brand-swapping converters.

De-duplicate the Windsurf converter family: delete the six local copies
in bin/install.js and bind the four exported ones by reference, guarded
by reference-identity assertions (the ADR-1508/#1675 pattern). The two
unexported helpers and an unused tool table go with them.

Replace the Windsurf 12,000-byte throw with description truncation,
matching the bound its sibling skill converter already applied. The
throw could only fire on an ~11.7 KB frontmatter description: the
largest emitted workflow is 311 bytes. Truncation makes the cap
unreachable by construction and leaves 12,000 in exactly one place,
eliminating the dual-surface duplication rather than testing for it.

Add the emitted-byte cap gate: buildEmittedSizes captures LF- and
<HOME>-normalized bytes from the walk buildParityManifest already
performs, and evaluateEmittedCaps asserts them against a per-runtime
cap table with dead-rule detection. buildParityManifest's return shape
is deliberately unchanged -- diffEmitted compares its values with
===, so making them objects would report all 8,529 emitted paths as
moved. A regression test pins the values as strings.

Add a deterministic trim-safety gate over composeWithinBudget's
omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity
rule, replacing the model-graded eval gate the issue described.

* docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract

* fix(#2931): bound the windsurf command name and single-source the brand swap

Review findings from the orthogonal passes, all fixed inline.

The claim that removing the 12,000-byte throw left total emission
"bounded by construction" was false. The #1615 regex constrains the
character class but not the length, and commandName is interpolated
three times into the emitted workflow: a 20,000-character name emitted
60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate,
clearly-labelled size control that THROWS -- commandName is the @-ref
path target, so truncating it would point the workflow at a file that
does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The
#1615 security regex is untouched and still runs first. 128 is generous:
the longest shipped name is gsd-plan-review-convergence at 27.

Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe
truncation helper. It still used a UTF-16 slice(0,177) -- the exact
surrogate-splitting bug the helper was written to avoid, in the very
sibling the helper's comment cites as its model. Bounds are unchanged,
so output is byte-identical for every shipped command (descriptions max
out at 99 chars).

Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting
the local copy. Adding it to the .cts left two unlinked implementations
of identical logic -- the drift class this change exists to remove.
Verified byte-identical across eight fixtures and five sequential calls
before merging, and guarded by a reference-identity assertion.

Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344),
add fast-check property coverage for the trim-safety contract, and use
fc.pre instead of a bare return in a property callback.

* test(#2931): fix three test-authoring bugs the remote matrix caught

The remote runner returned 8 unique failures on 6f15cdeb8. All three
causes were in the test files, not the modules under test -- local
harnesses exercise the modules directly, so nothing executed the test
bodies until the matrix did.

`{ __proto__: [...] }` in an object literal sets the prototype instead
of an own key, so the JSON round-trip erased it and the cap table never
saw a reserved runtime key. The production rejection was already
correct; the test could not reach it. Use a computed key.

Two cap fixtures tripped orthogonal error paths rather than the paths
they name: one declared windsurf in the cap table but omitted it from
sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern
matching nothing (a genuine dead rule). Both now include a compliant
artifact so the intended branch is what is asserted. The dead-rule and
unknown-runtime contracts are deliberate and unchanged.

`const { root } = makeSyntheticConfig({ ... `${root}` })` referenced
`root` from inside its own initializer -- a temporal dead zone error.
makeSyntheticConfig now optionally takes a (root) => files factory.

Also raise the npm pack --dry-run bound 60s -> 120s in the shipped-
scripts packaging test. That failure is NOT from this branch: the file
is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s
here, and the run recorded 60,637ms against a 60,000ms bound -- a
timeout under 28,948-test parallel contention, not a slowdown. Fixed
rather than deferred because a bound that tight is fragile regardless
of which branch trips it.

* chore(#2931): backfill changeset pr number to 2984

---------

Co-authored-by: sim <sim@local>
2026-08-01 16:00:14 -04:00
Tom Boucher
4df6d884b3 fix(#2641): treat absent capture_artifacts as enabled (schema default) (#2982)
* test(#2641): add regression for mempalace-capture gate default inversion

* fix(#2641): treat absent capture_artifacts as enabled (schema default)

The gate used `capture_artifacts !== true` which treated absent (undefined)
as disabled — inverted from the capability registry's declared default of
true. Changed to `capture_artifacts === false` (disabled only on explicit
false), matching the sibling gsd-mempalace-recall skill's correct pattern.

* chore(#2641): add changeset fragment

* chore(#2641): backfill changeset PR number 2982

---------

Co-authored-by: sim <sim@local>
2026-08-01 15:23:58 -04:00
Tom Boucher
6c96b13cfe fix(#2639): warn when local is ahead of origin before forking phase branch (#2981)
* test(#2639): add regression for local-ahead-of-origin warning in handle_branching

execute-phase.md's handle_branching forks from origin/$DEFAULT_BRANCH. When
local is ahead (unpushed commits), the phase branch silently misses them.
The test asserts the workflow checks for local-ahead-of-origin and warns.

* fix(#2639): warn when local is ahead of origin before forking phase branch

handle_branching forks the phase branch from origin/$DEFAULT_BRANCH. When
local $DEFAULT_BRANCH is ahead (unpushed commits like plan/research docs),
the fork silently misses those commits. Now a loud WARNING is printed to
stderr naming the commit count and advising the user, matching the existing
uncommitted-changes warning pattern.

* chore(#2639): add changeset fragment

* fix(#2639): condense warning under ADR-857 cap + merge emitted-drift ack

gsd-test gate caught: (1) execute-phase.md exceeded the 93600-byte Phase 6
ceiling — condensed the WARNING from 3 echo lines to 1. (2) emitted-attribution
flagged the growth without an ack — merged into the existing #2930 ack fragment
(execute-phase.md was already acked there; can't have two acks for the same path).

* fix(#2639): condense warning further to clear the 93400 comfortable-margin gate

The ADR-857 Phase 6 test has two assertions: <93600 (hard ceiling) and
<=93400 (comfortable margin). Condensed from 3 lines to 2 to fit under
93400 (now 93369).

* chore(#2639): backfill changeset PR number 2981

---------

Co-authored-by: sim <sim@local>
2026-08-01 14:24:05 -04:00
Tom Boucher
34633fa4ec fix(#2893): preserve prose below the JSON ledger on windows append/waive/fixed (#2975)
* test(#2893): add regression for append destroying prose below JSON ledger

writeLedgerAtomic overwrites the entire file with renderLedger(ledger),
dropping any prose below the JSON closing fence. The test creates a
WINDOWS.md with prose sections below the ledger, appends an entry, and
asserts the prose survives.

* fix(#2893): preserve prose below the JSON ledger on append/waive/fixed

writeLedgerAtomic was overwriting the entire WINDOWS.md with
renderLedger(ledger), which reconstructs only the frontmatter + header +
table + JSON block — silently destroying any prose a user wrote below the
JSON closing fence.

Now the writer reads the existing file before overwriting, extracts content
after the closing fence, and appends it to the rendered ledger. First-write
(no existing file) proceeds normally with no prose to preserve.

* fix(#2893): address review — correct fence search + idempotency test

BLOCKER from isolated adversarial review: indexOf(JSON_FENCE_CLOSE) matched
the OPENING fence ('json' starts with ''), duplicating the entire
JSON body as prose on every write. Now searches for the closing fence
starting AFTER the opening fence, mirroring parseJsonBlock.

Test hardened: non-empty initial ledger, second append (idempotency — prose
appears exactly once, exactly one JSON fence open), parseLedger round-trip.

* chore(#2893): add changeset fragment

* chore(#2893): backfill changeset PR number 2975

---------

Co-authored-by: sim <sim@local>
2026-08-01 12:59:09 -04:00
Tom Boucher
640eaee16e chore(#2930): fragmentize execute-phase.md and prove per-runtime composed emission (#2972)
* feat(#2930): fragmentize plan-phase.md workflow into per-runtime-composed sections

Adds src/workflow-fragments.cts (in-file <!-- gsd:section --> marker
parser/composer, ADR-1671 epic #1671 Phase 3), wires it into
bin/install.js's copyWithPathReplacement emission path, and pilots the
marker grammar on gsd-core/workflows/plan-phase.md.

Bookkeeping ripple for the new src/*.cts module: .gitignore,
eslint.config.mjs, docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json,
and a CONTEXT.md glossary entry. Amends ADR-1671 with open questions 1
and 2 resolutions and records the closed when= applicability grammar.
Adds docs/reference/workflow-fragments.md and an ARCHITECTURE.md
section documenting the marker authoring model.

* fix(#2930): put allow-test-rule issue ref on the same line as the marker

lint-allow-test-rule-refs.cjs requires the #NNN issue reference on the
same source line as `allow-test-rule:`; it was one line below and read
as an unreferenced novel exemption.

* docs(#2930): link the orphaned gate-predicates reference from the docs index

Found while adding the workflow-fragments reference doc: docs/reference/gate-predicates.md
shipped without an entry in docs/README.md, so it was unreachable from the docs index.
Fixed inline rather than deferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2930): scope composition to workflows, add typed failure reasons

Review findings from two orthogonal passes:

- Scope composeWorkflow to gsd-core/workflows/ only. It previously ran on
  every .md the installer copied, so a future agent/command/reference doc
  documenting the marker syntax with an unfenced example would have been
  mis-parsed and silently stripped — a lossy drop the phase forbids.
- Add a frozen REASON enum; failures attach a typed .reason and tests assert
  on it instead of matching free-form message text (CONTRIBUTING.md:635-694).
- Derive the property generator's when= values from WHEN_VOCABULARY instead
  of duplicating them (DEFECT.GENERATIVE-FIX).
- Add adversarial parser fixtures: Unicode headings, NUL, U+FFFD, BOM,
  fence-within-fence, tilde and indented fences, lone-CR marker line.
- Document why --mvp is structurally unmarkable: its content is interleaved,
  not sectioned, so the whole-line grammar cannot reach it.

Also fixes two stale tests on this branch, each reproduced on the unmodified
tree before correction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2930): retarget the pilot from plan-phase to execute-phase

The full remote matrix went red on both Linux lanes. Root cause was ours:
tests/phase6-capstone-conformance.test.cjs holds a PRE_PHASE6 ceiling of
94519 bytes for plan-phase.md, asserting an ADR-857 Phase-6 completion
property. That is a third size gate beyond the tier caps and the
differential ratchet, and it left plan-phase.md just 36 bytes of headroom
rather than the 3821 computed from the XL cap. The 330 marker bytes
overran it by 294.

Raising the ceiling is not an option: it is a red line certifying another
ADR's completion. plan-phase.md is reverted to byte-identical origin/next
and the pilot moves to execute-phase.md, which has 728 bytes of headroom
under its own ceiling and lands at 93147 with 3 marker pairs.

The vocabulary narrows to the atoms actually used: always, flag:--wave,
state:gap-closure-phase, state:has-prior-phases.

Recorded in the ADR: every branch the epic names lives in plan-phase.md,
which cannot be fragmentized until caps move from source to emitted bytes.
That is direct evidence for the epic's premise and may reorder phases 3-4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2930): backfill changeset PR number (#2972)

* fix(#2930): make the emission install tests portable on Windows

The windows-latest lane went red on three tests in the new install suite;
Linux was green. Both causes were in the test harness, not the module.

Root normalization: the opencode converter always embeds the install root
forward-slashed, but the tests stripped it with the native-separator string
from mkdtemp. On Windows that never matched, so the root leaked through
unstripped — and because the real and stub install roots have different
prefix lengths, that length difference landed directly in the byte-delta
assertion (344 observed vs 275 expected). Normalize both text and root to
one separator form before stripping.

@-ref resolution: the helper stripped only the @~/ and @$HOME/ forms, so a
Windows absolute ref (@C:/Users/...) fell through and was joined onto the
root, producing ...\@C:\Users\... Strip the @ first, then detect
absoluteness from the token's own shape (POSIX, drive-letter, or UNC) with
no platform branching, so every OS takes the same path.

Neither assertion was weakened; the exact-equality byte check is the point
of the test and still holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2930): document every REASON member and guard the doc/enum parity

Code review found the reference doc's 'Fails closed' list covering 10 of the
11 frozen REASON members — MALFORMED_ATTRIBUTES (parseAttrs rejects malformed
key="value" syntax) had no bullet, and it is distinct from
UNRECOGNIZED_ATTRIBUTE, which is valid syntax with an unknown key.

Two parallel surfaces sharing one constant with nothing asserting they agree is
the DEFECT.GENERATIVE-FIX class, so the same commit adds the parity assertion:
the test derives the enum side from the built module and the doc side by parsing
the reference page, keyed on the reason IDENTIFIER rather than prose so a
reworded bullet does not break it, and reports set differences in both
directions by name.

Proven non-vacuous: removing the MALFORMED_ATTRIBUTES bullet turns the suite
red naming that exact member; restoring it returns 44/44.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 12:12:20 -04:00
Tom Boucher
f0bb0787c9 fix(#2640): report truthful state_updated + keep progress frontmatter in sync after phase remove (#2974)
* test(#2640): add regression for state_updated false positive + stale progress

Three cases: (1) state_updated reflects actual content change, not just
file existence; (2) progress.total_phases resync'd even when the body lacks
'Total Phases:' (the no-op guard was skipping syncStateFrontmatter);
(3) state_updated is false when STATE.md doesn't exist.

* fix(#2640): report truthful state_updated + force frontmatter resync

Two defects in cmdPhaseRemove:

1. state_updated was fs.existsSync(statePath) — trivially true, since the file
   existed before and readModifyWriteStateMd never deletes it. Now captures
   the boolean return from readModifyWriteStateMd (changed from void to
   boolean: true when content was written, false on no-op).

2. progress.* frontmatter stayed stale when the body lacked 'Total Phases:'
   or 'of N' — readModifyWriteStateMd's no-op guard (#948) skipped
   syncStateFrontmatter when the body transform was unchanged. Now the
   transform forces a body diff when a phase was actually removed, so the
   guard passes and syncStateFrontmatter rebuilds progress.* from the
   post-deletion disk/ROADMAP state.

* fix(#2640): address review — gate forced-diff on targetDir, strengthen assertions

Two MAJOR findings from isolated adversarial review:
1. Forced-diff injected a spurious 'Total Phases:' line even when no directory
   was removed (targetDir === null). Now gated on targetDir !== null.
2. Test #2 asserted 'not 3' instead of '2' — would pass for any wrong count.
   Now asserts exact value. Test #1 strengthened to assert body Total Phases
   and frontmatter total_phases both equal 1.

* chore(#2640): add changeset fragment

* chore(#2640): backfill changeset PR number 2974

---------

Co-authored-by: sim <sim@local>
2026-08-01 11:56:34 -04:00
Tom Boucher
000a322489 fix(#2941): hint at global: prefix when bare skill name matches a global skill (#2973)
* test(#2941): add regression for bare-skill-name global: hint

When a bare skill name matches an existing global skill, the skip warning
must hint at the global: prefix. When no global skill matches, the original
'Skill not found' warning is unchanged.

* fix(#2941): hint at global: prefix when a bare skill name matches a global skill

When buildAgentSkillsBlock skips a bare skill name that doesn't exist as a
project-relative path, check if it matches an existing global skill. If so,
append a hint to the warning: 'a global skill named X exists; use global:X
to reference it'. When no global skill matches, the warning is unchanged.

getGlobalSkillDir and getGlobalSkillsBase are already imported in this module
for the global: branch. The hint is guarded on globalSkillsBase being non-null
(runtimes without a skills directory don't support the prefix).

* chore(#2941): add changeset fragment

* chore(#2941): backfill changeset PR number 2973

---------

Co-authored-by: sim <sim@local>
2026-08-01 10:58:50 -04:00
Tom Boucher
608be0e7cf fix(#2913): distinguish empty cherry-pick from genuine conflict in hotfix create (#2970)
* test(#2913): add regression for hotfix empty-cherry-pick discrimination

Two layers: (1) real-git test proving the discrimination logic (check for
unmerged paths → skip if empty, abort if conflict) is correct; (2) source-text
assertions proving the logic and summary heading are in release.yml.

Before the fix: release.yml treats any non-zero cherry-pick exit as a conflict,
so an already-applied commit (empty pick) aborts the entire hotfix create.

* fix(#2913): distinguish empty cherry-pick from genuine conflict

git cherry-pick exits non-zero for BOTH genuine conflicts AND empty picks
(the change is already present by content). The hotfix create job treated
every non-zero exit as a conflict, so an already-applied commit (notably the
structural 'chore: sync next package version' that follows every release
finalize) aborted the entire run.

Now the error handler checks for unmerged paths (git diff --diff-filter=U):
- No unmerged paths → already applied by content → skip, record, continue.
- Unmerged paths present → genuine conflict → existing abort/push/exit-1
  behavior and operator guidance, unchanged.

The job summary now has a separate 'Skipped (already applied by content)'
heading, distinct from 'Skipped (feat/refactor/etc)'.

* fix(#2913): address review — post-skip continuation test + conflict observability

Two findings from isolated adversarial review:
1. MINOR (test gap): no test proved the sequencer is clean after --skip, so a
   regression switching --skip to --quit would pass green. Added a second
   cherry-pick after the skip asserting it succeeds.
2. MINOR (observability): SKIPPED_EMPTY was dropped on the conflict-exit path
   — already-applied commits before a genuine conflict were silently lost from
   the summary. Now the conflict summary emits them under a dedicated heading.

* chore(#2913): add changeset fragment

* chore(#2913): backfill changeset PR number 2970

---------

Co-authored-by: sim <sim@local>
2026-08-01 10:21:27 -04:00
Tom Boucher
d3305fc3a5 fix(#2858): exclude gen-emitted-baseline.cjs from npm tarball + class-extinction guard (#2968)
* test(#2858): add class-extinction guard for shipped-script require boundary

Every shipped scripts/**/*.cjs must be require-able using only shipped paths.
The guard resolves the tarball file list via npm pack --dry-run --json (not a
hardcoded list) and statically checks each require() call against the shipped
set. A script requiring ../tests/** (which does not ship) is a violation.

* fix(#2858): exclude gen-emitted-baseline.cjs from the npm tarball

scripts/gen-emitted-baseline.cjs is repo-only CI tooling (CI workflows +
test fixtures spawn it from a checkout). It requires three modules from
tests/, which does not ship — so in a published install it is
MODULE_NOT_FOUND at load time.

Add a targeted files[] negation (!scripts/gen-emitted-baseline.cjs) so the
script stays in the repo for CI use but does not ship. Other scripts that
ship and are required by bin/install.js (build-hooks.js,
fix-slash-commands.cjs, gen-capability-registry.cjs) are unaffected.

The class-extinction guard test in
tests/packaging-shipped-scripts-require-only-shipped.test.cjs ensures no
shipped script can require outside the shipped tree going forward.

* test(#2858): widen guard to .js + strip block comments (review fixes)

Two findings from isolated adversarial review:
1. MAJOR: the guard only checked .cjs files, but scripts/build-hooks.js
   ships and is required by bin/install.js — a .js file with a broken
   require would bypass the guard. Widened filter to /\.(cjs|js)$/.
2. MODERATE: the static parser could false-positive on require() calls
   inside /* */ block comments or inline // comments. Now strips both
   before matching.

* chore(#2858): add changeset fragment

* chore(#2858): backfill changeset PR number 2968

* fix(#2858): add issue ref to allow-test-rule exemption (ADR-456)

CI lint-tests caught: the allow-test-rule comment needs a 'see #NNN' ref
per ADR-456. Added '(see #2858)' to the integration-test-input exemption.

---------

Co-authored-by: sim <sim@local>
2026-08-01 09:29:15 -04:00
Tom Boucher
1db7dcd9bf fix(#2849): strip trailing hyphen after 60-char slug truncation (#2967)
* test(#2849): add failing regression for trailing-hyphen-after-truncation

The strip ran before .substring(0, 60), so a cut landing on a separator
produced a slug ending in '-'. Four cases: the exact issue repro (59 a's +
space + tail), a boundary landing before a separator, leading-hyphen survival,
and a long-Cyrillic transliteration+truncation case.

* fix(#2849): strip trailing hyphen after 60-char truncation

generateSlugInternal ran the ^-+|-+$ hyphen strip BEFORE .substring(0, 60),
so a title whose 60-character cut landed on a separator yielded a slug ending
in '-' — the very thing the strip step exists to prevent.

Reorder so the strip runs after truncation. Truncation cannot introduce a
leading hyphen, so the full ^-+|-+$ pass last is equivalent for leading
hyphens and fixes the trailing-hyphen-after-truncation case.

Latin-script output for titles ≤ 60 chars is byte-identical; only titles
whose truncation boundary lands on a separator change (from broken to clean).

* test(#2849): add all-separator collapses-to-empty boundary case

Surfaced by isolated adversarial review: pin the contract that input
which is entirely separators ('!!!', '!'.repeat(70)) reduces to '' —
not null, not a stray hyphen — both short and past the 60-char truncation.

* chore(#2849): add changeset fragment

pr:0 placeholder; will backfill the real PR number after the PR exists.

* chore(#2849): backfill changeset PR number 2967

---------

Co-authored-by: sim <sim@local>
2026-08-01 08:24:05 -04:00
Tom Boucher
0bb7525a62 fix(#2943): rename get-library-docs -> query-docs; correct the ctx7 fallback rationale (#2963)
* test(#2943): parity guard against the nonexistent get-library-docs tool

Second context7 naming drift after #2017 (which guarded the plugin-marketplace
PREFIX). #2017's guard only checks tools: frontmatter lines, not prose bodies —
which is where the broken tool NAME (get-library-docs) lived. The context7 MCP
server registers only resolve-library-id and query-docs; get-library-docs is a
stale copy from upstream's own README.

Scans the shipped prose surface (agents/, gsd-core/references|workflows/,
commands/gsd/, skills/) and fails if any artifact instructs an agent to call
mcp__context7__get-library-docs. Excludes tests/ (a fixture may use the name as
a negative input) and CHANGELOG/RELEASE-NOTES-LEGACY (history).

Fails-first: 4 offenders today (gsd-executor.md:29,
research-documentation-lookup.md:5, discovery-phase.md:68 & :104).

* fix(#2943): rename get-library-docs to query-docs and correct the ctx7 fallback rationale

The context7 MCP server registers only resolve-library-id and query-docs
(verified against upstream packages/mcp/src/index.ts); get-library-docs is a
stale name copied from upstream's own README. Four shipped prose sites instructed
agents to call a tool the server does not register, so every research path that
loaded the canonical reference either errored, fell through to the ctx7 CLI
branch, or fabricated a result.

- research-documentation-lookup.md, gsd-executor.md, discovery-phase.md (x2):
  get-library-docs -> query-docs, params context7CompatibleLibraryId/topic ->
  libraryId/query (the registered contract).
- Same files' ctx7 CLI fallback rationale: the cited cause
  (anthropics/claude-code#13898 'strips MCP tools from agents with a tools:
  frontmatter restriction') was wrong on two counts — #13898 is closed and was
  never about tools: frontmatter. Rewritten to describe the real mechanism
  (custom subagents cannot see project-scoped .mcp.json; they only inherit
  user-scoped ~/.claude/mcp.json). The fallback itself is kept.
- discovery-phase.md 'mode: code/info' dropped — query-docs takes libraryId +
  query only; the code-vs-concepts intent is now expressed via the query text.

resolve-library-id is unchanged (still registered upstream). CHANGELOG and
RELEASE-NOTES-LEGACY citations are historical record, left as-is.

* chore(#2943): add changeset fragment (pr:0 placeholder)

* test(#2943): widen parity-guard scan surface to docs/ (isolated-review finding)

The isolated adversarial review flagged that SCAN_DIRS omitted docs/, which
ships docs/AGENTS.md — agent-consumed prose carrying 8 mcp__context7__* refs.
No false negative today (it uses only the wildcard), but a future banned-name
addition there would slip through, recreating the exact drift this guard exists
to prevent. Add docs/ to the scan surface, with an EXCLUDED_FILES set for
historical record (docs/RELEASE-NOTES-LEGACY.md, CHANGELOG.md) that must not be
rewritten to satisfy the guard.

* fix(#2943): update shifted PROSE_ALLOWLIST line + acknowledge gsd-executor.md growth

The gsd-test gate caught two real consequences of the rationale rewrite in
agents/gsd-executor.md (the +2-line corrected mechanism description shifted
line numbers below it):

1. tests/no-bare-gsd-tools-command-position.test.cjs: the legitimate
   'gsd-tools query commit' descriptive mention moved from line 791 -> 793.
   Update the PROSE_ALLOWLIST entry to the new line (the mention is unchanged,
   just relocated by my edit above it). Without this the gate reports both a
   stale allowlist entry (791) and a new offender (793) for the same mention.
2. tests/emitted-drift-acks/2943-context7-tool-name.json: gsd-executor.md grew
   95 bytes (the accurate mechanism rationale is longer than the wrong one-line
   #13898 attribution it replaces). Acknowledge the growth with the reason.

Both are mandated by the gate, not optional. The rename itself (get-library-docs
-> query-docs) is byte-neutral-ish; only the rationale rewrite grew the file.

* chore(#2943): backfill changeset PR number 2963

---------

Co-authored-by: sim <sim@local>
2026-08-01 01:34:05 -04:00
Tom Boucher
73418c516f fix(#2956): scope Phase extraction to ## Current Position (3rd gen of #2444/#2567) (#2961)
* test(#2956): fail-first regressions for Phase scoped to ## Current Position

Third generation of #2444 / #2567. Stopped At / Paused At were scoped to
## Session; Phase (canonically in ## Current Position per templates/state.md)
was left unscoped, so a historical Phase: / **Phase:** line in an archive
section overwrites current_phase on every write. Since current_phase is
routing input for gsd-progress / --next, the rewind routes work to the wrong
phase.

Six failing-first regressions + one round-trip:
- shape B: bold **Phase:** 19 archive BELOW the section
- shape C: plain archive Phase: 19 ABOVE the section
- bootstrap h3 ### Current Position variant
- CRLF variant
- Phase token in decisions prose (over-broad-fix guard)
- Paused At read-path parity with the write seam (## Session)
- write-then-read round trip stays at 22 (read/write agreement)

Folded into tests/state.test.cjs (lint:regression-test-names bans a
new tests/bug-NNNN-*.test.cjs file).

* fix(#2956): scope Phase extraction to ## Current Position at both seams

Third generation of #2444 / #2567. Stopped At / Paused At were scoped to
## Session by those fixes; Phase (canonically in ## Current Position per
templates/state.md) was left unscoped, so a historical Phase: / **Phase:**
line in an archive section silently overwrote current_phase on every write.
Because current_phase is routing input for gsd-progress / --next, the rewind
routes work to the wrong phase.

Fix mirrors the proven #2444 seam exactly:
- new matchCurrentPositionSection helper (collectSection-based, CRLF-tolerant,
  level-flexible for the bootstrap ### Current Position h3 variant), sited next
  to matchSessionSection.
- read path (cmdStateSnapshot): extract Phase from matchCurrentPositionSection
  ?? body. Also scope Paused At to matchSessionSection ?? body so the read seam
  agrees with the write seam (which already scoped Paused At to ## Session).
- write path (buildStateFrontmatter): extract Phase from
  matchCurrentPositionSection ?? bodyContent.

stateExtractField itself is untouched (its bold/plain precedence is load-bearing
for other fields — the #3265 test depends on it), and preferNewerLastActivity is
untouched (Last Activity has no canonical section; its date-direction guard is
deliberate). Fall back to full-body when no ## Current Position section exists
so files without the heading keep current behaviour.

* chore(#2956): add changeset fragment (pr:0 placeholder, backfill after PR)

* test(#2956): make round-trip test actually trigger the write-path resync

The write-then-read round-trip test used 'state update Status "Executing"' on a
fixture with no Status field, so the update was a no-op (updated:false) and no
frontmatter resync ran through buildStateFrontmatter — the assertion on the
written frontmatter then failed not because the fix is wrong, but because no
write happened. Add a **Status:** field so the update performs a real field
update (updated:true) and forces the resync. Verified locally: pre-fix this
writes current_phase:19 (the archive value); post-fix it writes 22.

The code fix is correct (5 of 7 RED tests passed; the 2 failures were this
defective test). This is the 'fix the bad test' half of the TDD-loop rule.

* chore(#2956): backfill changeset PR number 2961

---------

Co-authored-by: sim <sim@local>
2026-08-01 00:02:41 -04:00
Tom Boucher
9ac0dfad58 chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor

Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared
context-composer seam. Its success condition is that review-prompt output does
not change, and the only authority on "did not change" is the behavior that
shipped before the refactor. Capture that behavior now, while it is still the
live implementation.

47 characterization cases, every `expected` value computed by executing the
current implementation rather than hand-authored — the independence
CONTRIBUTING.md "Fixture provenance (#2371)" asks for.

A corpus is only worth what it can detect, so this one was validated by
mutation rather than assumed. Five deliberate defects were injected and each
must be caught by at least one case:

  - the note reserve deducted unconditionally instead of only under pressure
  - the pressure test relaxed from `>` to `>=`
  - a no-op head-shrink still setting the shrunk flag
  - the per-plan floor dropped from the proportional share
  - drop order reversed

Two of those exposed real holes in the first cut of this corpus, and the cases
that close them exist because of it:

  - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a
    floored plan group, and the 1024-char floor absorbed the entire trim, so the
    mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap,
    which makes the strict inequality observable as context kept vs omitted.

  - No case reached proportional-truncate at all — B6 and B7 both hard-failed
    the min-set pre-check first, leaving planTruncationPct at 0 across every
    case and the floor semantics entirely unexercised. Rebudgeted to 700 and
    1100 so the min-set fits and the truncate step is actually reached; they now
    record 40.20% and 48.80%.

The A4/A10 families sweep the pressure boundary from both sides, which is where
this function has regressed before: CONTEXT.md's
LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions
that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band,
because the suite paired a trivially-fitting budget with a trivially-overflowing
one and never sampled between them. A4 pins that nothing is trimmed from the cap
down to 81 tokens under it; A10 pins that pressure fires at +1. Together with
A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures.

Two facts the corpus establishes that the design notes had wrong:

  - "" and null sections are NOT distinguished. applyBudget uses truthy checks
    throughout, so an empty-string section is treated as absent: not rendered,
    not dropped, never recorded in `omitted`. B13b pins this while the ladder is
    actively trimming, where only the non-empty `research` is dropped.

  - Sizing matters. B12/B13 were first written at a budget where both hard-failed
    the min-set check and returned "", so comparing them compared two empty
    strings and proved nothing.

Committed as its own commit, ahead of the refactor, and regenerated against the
pre-refactor implementation, so the oracle is demonstrably independent of the
change it will adjudicate.

Refs #2929

* refactor(#2929): extract the context-composer seam from prompt-budget

Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer —
per-runtime artifact emission — but it is walled inside the cross-AI review
pipeline. Lift it into a shared seam so later phases can call it, without
changing what the review pipeline emits.

ADR-1671 specifies the composer as "priority + binary-search cutoff to a
per-runtime budget". Read against the code it generalizes, that contract cannot
express the thing being generalized. applyBudget is not a cutoff: it is a fixed
five-step ladder in which each section carries its own shrink strategy, and only
three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N
lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor;
instructions and roadmap are never touched at all. A cutoff composer sorts by
priority and discards the tail — it has no way to say "shrink this one",
"truncate that one but never below 1 KB each", or "these three are the only
droppables, in this order". Building to the literal contract and routing
prompt-budget through it would have silently changed review-prompt output, which
is the one outcome this phase forbids.

So shrink strategies are the core abstraction here, and cutoff becomes one
strategy among them — the right one for per-runtime emission in Phases 3-4, not
for this ladder. That is an elaboration of the ADR's intent, not a departure
from it, and ADR-1671 is updated to say so.

Three decisions worth stating:

  - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan
    of surviving fragments and never a string. assemblePrompt's rendering is
    prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and
    owning it in the composer would force emission to adopt prompt-shaped
    rendering. The split is what lets one seam serve both consumers.

  - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its
    chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires
    for emission caps. The existing code converts a token budget to a character
    budget with a hardcoded `* 4`; that assumption is now an explicit
    `charsPerUnit` inverse, which is precisely what a byte unit needs in order to
    reuse this.

  - The entry point is `composeWithinBudget`, not `applyBudget`. That name
    already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter
    being an unrelated graph-edge budget. A third would make every symbol search
    in this repo ambiguous, and it already misresolves: preflight and impact
    queries for "applyBudget" return graphify's.

Behavior is unchanged and proven so: all 47 characterization cases reproduce
byte-identically, and the corpus is mutation-validated rather than merely green
(see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and
from eighteen mutable accumulators to two, both inside a helper copied verbatim.

estimateTokens deliberately stays in prompt-budget and keeps its exact math:
src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins
plan estimates and recorded actuals to that same scale, so moving or changing it
would silently break the calibration loop.

Refs #2929

* docs(#2929): document the context-composer seam and amend ADR-1671

Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new
domain modules), and a mutation-matrix entry for the new module.

The ADR amendment is the substantive part. ADR-1671 specified the composer as
"priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2
established that a cutoff alone cannot express the function the platform
generalizes, so the ADR now records shrink strategies as the core abstraction
with cutoff as one strategy among them, reserved for per-runtime emission in
Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned
against that contract and would otherwise be planned against a mechanism that
does not work.

The mutation-matrix entry is not bookkeeping. Stryker scores per module against
a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the
extracted code unmeasured while prompt-budget's own score floated free of the
logic it used to cover. context-composer gets its own entry at the same floor.

Refs #2929

* test(#2929): pin the effectiveBudget rounding mode in the parity corpus

An isolated correctness review found a real blind spot: mutating
`Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of
the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator
happened to produce a whole number, so floor, round and ceil all agreed and the
rounding mode was entirely unpinned by a corpus whose whole job is to pin
observable behavior.

Three cases fix that by straddling the .5 boundary:

  A11  95 * 0.90  = 85.5   floor 85, round 86  -> the two disagree
  A12  97 * 0.90  = 87.3   floor and round agree; ceil (88) does not
  A13  93 * 0.85  = 79.05  same guard at a non-multiple-of-10 margin, so the
                           margin arithmetic is exercised and not just the budget

A11 alone catches the round mutation; all three catch ceil. Regenerated against
the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so
the expanded corpus keeps the independence property the original capture had.

The corpus is now mutation-validated against seven injected defects, every one
caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink
setting its flag, the truncate floor ignored, drop order reversed, and both
rounding-mode changes.

Refs #2929

* feat(#2929): flexReserve floors and the byte-stable isolate prefix

Two of issue #2929's "Done when" items were unimplemented rather than deferred,
and an isolated review flagged them alongside my own audit. Both are part of
ADR-1671's composer contract, so shipping the seam without them would have left
Phases 3-4 building against a contract that does not exist yet.

flexReserve is a per-fragment floor in measure units that every strategy must
respect, which is what makes it different from the pre-existing floorChars: that
one is a chars-denominated detail of proportional-truncate alone and is retained
unchanged. A floored fragment is never dropped, is never head-shrunk below its
floor, and raises its own proportional cap. A fragment already smaller than its
floor is untouchable outright. Metadata gains `floored`, listing the ids whose
floor actually prevented a trim — a guarantee no caller can observe is a
guarantee no test can hold you to.

isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed,
never dropped, but still counted, because a prefix excluded from accounting
would silently under-count real context. Metadata gains `isolatePrefix` so a
caller can hash or assert on the exact bytes. Declaring an isolate fragment
after a non-isolate one throws: a prefix that is not at the front is not a
prefix, and accepting it would make the cross-runtime stability claim
meaningless.

Adds tests/context-composer.test.cjs for the exact new semantics and
tests/context-composer.property.test.cjs for the five invariants, including the
budget-monotonicity property the issue names explicitly. Both are registered in
the mutation matrix, since coverage does not migrate with relocated code.

prompt-budget uses neither feature, and its output is unchanged: all 50 corpus
cases still reproduce byte-identically.

Refs #2929

* chore(#2929): allowlist the prompt-budget parity suite

The parity corpus needs its own test file and that makes prompt-budget a
three-file module against a limit of two. The lint offers consolidation or an
allowlist entry with justification; the entry is the right call here.

Consolidation would mean folding the characterization suite into
prompt-budget.test.cjs, which is the one thing that should not happen to it. The
parity suite is a distinct concern with a distinct lifecycle: it is generated
rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own
scoring target, and its failure means something categorically different from a
unit-test failure — not "this behavior is wrong" but "observable output moved".
Burying it inside a general unit file would obscure exactly that signal.

The allowlist is an identity ratchet, so this entry pins today's three exact
filenames: adding a fourth still fails, and dropping back to two requires
removing the entry.

Refs #2929

* fix(#2929): register the new module with two gates it was missing

The remote matrix caught three defects that no local check could, because the
local runner is blocked in this repo and these suites had therefore never
executed. Eight failures, identical on node22 and node24, so nothing
environment-shaped.

Two are the new-module ripple. A net-new src/*.cts lands in six places and this
change had reached four of them — .gitignore, INVENTORY, the manifest, and the
CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not
be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet
baseline (a deliberate review-visible mirror of the matrix floors, which every
COVERED module must carry). Both are now registered, the ratchet at the same
floor of 66 the matrix declares.

The third was a test asserting an outcome it had made impossible. It set
budget:1 alongside a 400-char required fragment, so the group budget came out at
-99 and the proportional-truncate step was skipped entirely — the deliberate
"non-positive group budget is skipped, never clamped" rule inherited from the
original ladder. Nothing was trimmed, and the test then asserted a truncation.
Rebudgeted so the step actually runs, with the arithmetic written out in a
comment so the next reader does not have to re-derive why 120 rather than 80.

Fixing that surfaced a genuine bug in the composer. `floored` is documented as
recording fragments whose flexReserve prevented a trim that would otherwise have
happened, but the push sat in the else-branch of "content did not change", so it
only fired when nothing was trimmed at all. A fragment truncated to a
reserve-raised cap has also had a trim prevented — 40 characters' worth in the
test above — and was silently absent from the field that exists to make the
guarantee observable. The condition was already right; it was in the wrong
branch. Now recorded on both paths: a drop prevented outright, and a truncation
capped higher than the share alone would have allowed.

Parity is unaffected — prompt-budget never sets flexReserve, so the branch is
unreachable from every corpus path, and all 50 cases still match.

Refs #2929

* chore(#2929): backfill changeset PR number (#2958)

* chore(#2929): correct the corpus case count in the changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-07-31 23:03:13 -04:00
Tom Boucher
f6257f3745 ci(#2952): shard the full test lane instead of widening its cap (#2960)
* ci(#2952): budget CI job timeouts by headroom over measured cost

`origin/next` was red. The only failing check was `Required tests`, and its
sole cause was `test (ubuntu-latest, 24)` reported `cancelled` — GitHub's
conclusion for a job that exceeds its `timeout-minutes`, not a button press.

That job's `ubuntu-latest / 24` entry is the only `scope: full` matrix entry:
it runs the whole unit suite under c8 coverage, then the scripts/ coverage
floor, integration, security, install and slow, serially on one runner. On
next@5a0a9f097 it ran 15m16s against `timeout-minutes: 15` and was axed 23s
into `npm run test:slow`. Projected to a completed slow step (27s on the last
green run) the lane costs ~15m20s.

Confirmed hypothesis: the budget, not the suite. The lane had been riding the
ceiling all day — 12m03s, 11m34s, 11m50s, 14m25s, 14m51s — and crossed on
three of the last four full-lane runs (d2d2f7c08, 07603df8f, 5a0a9f097). There
is no pathological test: the unit run is cost-first bin-packed into 12 chunks,
chunk 1 is gated by run-tests-harness.test.cjs at 169s (expensive by design —
it spawns real harness subprocesses, one of which exercises the per-chunk
timeout), and the remaining chunks are 32-94s. 769s is the honest cost of 689
files under c8. Re-running could not have helped; the work exceeded the budget.

Review of the first cut surfaced the same defect one runner away: `full test
(windows-latest, 22, shard 3/3)` reached 18m59s against its own 20-minute cap
on 05b170e44 (94%) and 18m14s on 81eeb8a53 (91%). That lane has already blown
its cap twice (#1051, #1212). Fixed here rather than deferred.

The first cut also asserted `test >= test-full`, which is unsound — those two
budgets are dominated by different platforms, so their ordering carries no
meaning. Replaced with the invariant that actually generalises: every lane is
held to a headroom FACTOR over its own measured cost. `test` 15 -> 25 (1.5x of
16m), `test-full` 20 -> 30 (1.5x of 19m), `test-inert` unchanged at 15.

tests/ci-test-job-timeout-budget.test.cjs locks that rule. No unit test can
prove a lane still FITS its budget — only a real run measures that — but a
budget can no longer be lowered back beneath what its lane is known to need,
and a lane that gets slower must be re-measured rather than excused.

Separately: the earlier `failure` at 05b170e44 was an unrelated, already-fixed
CONTEXT-INDEX.json drift (07603df8f re-synced it; lint-tests is green at HEAD).
07603df8f's own run hit this same timeout, which is why it never reported green.

Refs #869, #1051, #1212

* ci(#2952): shard the full test lane instead of widening its cap

The `scope: full` lane was the only unsharded lane in this file. It ran the
entire unit suite under c8 on one runner, grew past a 15-minute cap, and
reddened `next`. Raising the cap bought room; it did not change the shape, and
the same lane would have walked back into the ceiling. Shard it, the way #1212
answered this for the Windows lane.

Balance comes from measurement, not file counts. scripts/run-tests.cjs already
partitions by measured per-file duration using LPT (#2472); the table it reads
was 10 days stale — 638 of 695 files timed, 64 missing, including the whole
context-predicates group. Regenerated from a verified matrix run: 700 files, 0
missing. On that table the 685-file unit suite splits 19.37m / 19.37m / 19.37m
— 0.0% spread — and the split is a total, disjoint cover with 0 files dropped.
Completeness, disjointness, balance and determinism of the partition itself are
already pinned against selectShard in run-tests-harness.test.cjs, including a
fast-check property, so this change does not restate them.

Sharding a COVERAGE run is the part that needs care. A per-shard percentage is
meaningless — shard 2 never executes shard 1's files, so those read 0% — and
leaving the gate on the shards would have quietly measured a third of the tree.
Each shard now renders no report and only leaves raw V8 dumps; a new
`coverage-gate` job merges all three into one coverage/tmp and runs the gate
there. c8's default temp directory is where the download lands, so the ≥70%
lines / ≥60% branches gate and the ≥55% scripts floor run unmodified against
merged data.

Both surfaces call the same npm scripts rather than inlining c8 into YAML, so
the include/exclude globs and both thresholds stay defined once in package.json.
The workflow holding its own copy is the divergence this repo has a rule
against, and the new test cross-checks package.json so an inline reintroduction
fails rather than drifts.

tests/ci-full-lane-sharding.test.cjs covers the two ways this stays GREEN while
being wrong: an incomplete shard set (declare 1/3 and 2/3, never 3/3, and a
third of the suite silently stops running) and a coverage gate that stops being
required. required-tests now depends on coverage-gate and fails on it, while
still tolerating `skipped` so docs-only PRs are not blocked.

`timeout-minutes: 25` on the lane is deliberately left alone. The budget test
requires a real measurement before a lane's declared cost changes, and the
sharded cost is not measured until this PR's own CI run.

Refs #1212, #2472

* ci(#2952): tighten the sharded lane's budget to its measured cost

The sharding commit deliberately left `timeout-minutes: 25` alone, because
tests/ci-test-job-timeout-budget.test.cjs requires a real measurement before a
lane's declared cost changes and the sharded cost did not exist yet.

It exists now. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s, and
coverage-gate 1m20s. Shard 1 is the long pole because the unsharded aux suites
ride along on it, which is deliberate — they total ~1m35s and sharding them
would cost more than it saves.

So the lane's budget is 15 against a slowest measured shard of 8 minutes
(~1.9x), and coverage-gate joins LANE_COSTS at 2 minutes. 15 is the same number
the lane blew before sharding; the work behind it is now a third the size.

Merged coverage was checked against the pre-shard single-runner baseline rather
than assumed from a green check: 94.36 stmts / 96.3 funcs / 94.36 lines
identical, branches 84.22 vs 84.21 — one branch across two different trees,
noise rather than a regression.

---------

Co-authored-by: sim <sim@local>
2026-07-31 22:16:26 -04:00
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
JusticeWay
7b204ad2ac enhance(#2530): extend UAT checkpoint frame language pack (9 more languages) (#2564)
* feat: extend UAT checkpoint frame language pack (9 more languages)

response_language is a free-form config value, but CHECKPOINT_FRAMES only
covered 9 languages — any other configured language silently fell back to
the English frame. Add Dutch, Polish, Russian, Ukrainian, Turkish, Hindi,
Arabic, Vietnamese, and Indonesian frames plus their aliases, with a
regression test asserting each resolves instead of falling back.

Follow-up to #2402 (PR #2457).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: add changeset for #2527

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2530): list UAT checkpoint frame languages in CONFIGURATION.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2530): point changeset fragment at PR #2557

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2530): address Unicode language-pack review

* fix: address checkpoint language review

* fix: count spacing combining marks in checkpoint width

* test: verify checkpoint aliases structurally

* fix: isolate RTL checkpoint frames

* fix: isolate RTL checkpoint frames correctly

* test(#2530): assert checkpoint aliases neither collide nor go unreachable

Review Minor #1. A duplicate alias key was invisible to the existing
catalog tests: the runtime object is well-formed after JS collapses the
literal, the self-alias assertion still holds, and the losing language
just stops resolving. tsc catches the byte-equal case (TS1117), but not
the two that survive compilation — an alias whose NFC-lowercase form
already belongs to another language, and an alias not in lookup form at
all, which resolveCheckpointFrame() can never produce.

The check reads the source literal rather than the object, since the
object no longer records what was written. Both assertions are
independently load-bearing: an NFD twin of an existing alias trips the
collision check, an uppercase alias trips the unreachability check.

Review Minor #2: changeset retyped Changed -> Added. Nine wholly new
supported response_language values are an addition under Keep a
Changelog, not a modification of existing behavior.

* test(#2530): check alias collisions on the catalog, not its source

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:01:50 -04:00
Tom Boucher
f092c6da85 fix(#2649): diagnose-issues + execute-plan run worktree.base-check before worktree dispatch (#2955)
* test(#2649): failing-first — diagnose-issues + execute-plan must run base-check before worktree dispatch

* fix(#2649): diagnose-issues + execute-plan run worktree.base-check before dispatch

diagnose-issues.md spawn_agents and execute-plan.md Pattern A spawned
worktree-isolated subagents (gsd-debugger / gsd-executor) without the
pre-dispatch worktree.base-check gate that execute-phase (#683/#1369) and
quick (#1941) already run. Claude Code's isolation="worktree" forks from
origin/HEAD, not live local HEAD; without the gate, the documented GSD steady
state (commit every step locally, push only on request) hits the verify-only
worktree_branch_check guard's exit-42 halt mid-investigation with no auto-degrade.

Mirror the quick.md #1941 pattern: before dispatch, run
`gsd_run query worktree.base-check --pick shouldDegrade`; if true, print its
message + a #2649 warning to stderr and set USE_WORKTREES=false (sequential
main-tree dispatch). The verify-only guard stays as a backstop in both cases.

Per the triage and #2649 acceptance criterion 5, execute-plan.md's Pattern A
(identified as a second site with the identical gap) is fixed in the SAME change
— same bug class, same one-line gate, two workflow files — rather than filed as
a separate follow-up.

* fix(#2649): ack the diagnose-issues + execute-plan growth (per-PR fragment)

The two workflow files grew vs next (diagnose-issues.md +1381, execute-plan.md
+905) adding the #2649 base-check gate. emitted-attribution requires an ack;
this is a per-PR fragment under tests/emitted-drift-acks/ (#2914 mechanism,
replacing the legacy shared emitted-drift-ack.json).

* test(#2649): tighten base-check ordering assertion + guard backstop survival

Address code-review minors:
- the ordering assertion was a loose disjunction that passed even if the
  base-check moved AFTER the dispatch; tighten to assert base-check < Agent()
  (the real invariant).
- add a test that the verify-only <worktree_branch_check> backstop remains
  embedded in the Agent() prompt (acceptance criterion 4 — the base-check is a
  pre-dispatch degrade, the guard is a post-fork fail-closed backstop; both
  layers must survive).

* changeset(#2649): diagnose-issues + execute-plan auto-degrade on stale worktree base

* changeset(#2649): backfill PR number 2955

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 19:15:04 -04:00
Tom Boucher
388837219d fix(#2648): phase.complete refuses when non-retired plans lack summaries (fail-closed coverage gate) (#2953)
* test(#2648): failing-first — phase complete must refuse when plans lack summaries

Adds findUnsummarizedPlans to core-utils (mirrors countMatchedSummaries but
returns the unmatched plan files) and a PHASE_PLAN_COVERAGE_INCOMPLETE error
reason, plus a 3-case regression block in phase.test.cjs. The gate itself is
NOT yet wired into cmdPhaseComplete (reverted for the RED run), so the
'blocks completion' and 'superseded does not block' cases must FAIL (the
pre-fix code completes silently).

* fix(#2648): phase.complete refuses when non-retired plans lack summaries

cmdPhaseComplete gated only on a single *-VERIFICATION.md status, so a phase
could close complete while an arbitrary number of its plans had no completion
record (confirmed incident: 6/30 plans unexecuted incl. the phase's entire
final UI scope, every signal green). Add a fail-closed plan-coverage gate that
refuses completion when any plan lacks a matching *-SUMMARY.md, naming the
missing plans, UNLESS the plan is retired via machine-readable status:
superseded frontmatter (#2349) — closing the Goodhart hole (delete a SUMMARY to
raise the %) without regressing the lock/recovery pattern.

Uses scanPhasePlans (superseded-AWARE) + new findUnsummarizedPlans helper so the
gate, the count, and the named list can never disagree. Evaluated before the
verification-gate transaction so a refusal fails fast without mutating
ROADMAP/STATE. milestone.complete's parallel gap is explicitly out of scope
(separate seam, separate PR).

* fix(#2648): test fixtures — give #1752 phase plans summaries + STATE.md in coverage fixture

The plan-coverage gate (#2648) correctly blocks phase completion when a plan
lacks a SUMMARY. Two test fixtures needed updating to reflect the new contract:
- #1752 (total_phases-decrement cascade): its 8 phase dirs each had a PLAN.md
  with no SUMMARY. The test's concern is the total_phases cascade, not plan
  coverage, so add a matching SUMMARY to each to keep the phase fully-covered
  and isolate the #1752 behavior.
- the #2648 coverage-gate fixture: write STATE.md (createTempProject scaffolds
  .planning/phases but not STATE.md) so the 'ROADMAP/STATE unchanged on refusal'
  assertions have a file to read.

* fix(#2648): security — fail closed on unreadable plan dir + sanitize msg + surface superseded

Address the security-review blocker (B1) and hardening (M1/m2):
- B1 (blocker): the gate failed OPEN when scanPhasePlans could not read the
  phase dir (it swallows readdirSync errors → empty plan set → gate sees zero
  unsummarized plans → passes). A coverage gate that passes when it cannot read
  the plans re-opens the #2648 hole under any I/O failure. Now readdirSync the
  dir explicitly and fail closed (PHASE_PLAN_COVERAGE_INCOMPLETE) on a throw;
  a readable empty dir still passes (legitimately complete empty phase).
- m2: sanitize plan filenames (strip C0 controls / DEL) before interpolating
  into the error message — they come raw from readdirSync and could spoof the
  terminal in plain-error mode.
- M1: surface the count of plans excluded as status: superseded so a reviewer
  can audit which work was declared retired (the marker is a committable,
  review-time-trusted bypass; keep it visible).
- Add a 4th regression case: unreadable plan dir (ENOTDIR via a file, not chmod
  0o000 which root bypasses) must fail closed.

* test(#2648): drop unreachable B1 case — no root-safe unreadable-dir repro

The B1 fail-closed-on-unreadable-dir defense stays in src/phase.cts (cheap +
correct), but it cannot be unit-tested cross-platform: any condition that makes
the phase dir unreadable to the gate's readdirSync ALSO fails findPhaseInternal
upstream ('Phase N not found') before the gate runs, and chmod 0o000 is
forbidden (root bypasses it in root CI). Document the gap in the test file;
remove the case that asserted a reason the upstream error pre-empts.

* style(#2648): drop unnecessary type assertions flagged by lint:ci

scanPhasePlans returns typed string[] arrays, so the `as string[]` casts on
coverageScan.planFiles/summaryFiles were redundant (@typescript-eslint/
no-unnecessary-type-assertion). Compute supersededCount from typed lengths;
only the phaseInfo['plans'] cast remains (it is genuinely unknown).

* changeset(#2648): phase.complete refuses when plans lack summaries

* changeset(#2648): backfill PR number 2953

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 18:11:40 -04:00
Tom Boucher
5a0a9f0972 fix(#2944): remove the catastrophic-backtracking regex from the ADR-1671 example (#2950)
* fix(#2944): remove the catastrophic-backtracking regex from the example

The non-shipping Option-E reference example carried its own copy of the
predicate-id regex, which nested a dot-containing character class inside a
dot-prefixed repeat. A run of N consecutive dots therefore had exponentially
many partitions. Measured on next before this change: 30 dots 54ms, 35 66ms,
40 807ms — so roughly 55-60 dots hangs for hours.

Not exploitable where it sits: the example is outside tsconfig.build.json,
outside the npm package files list, outside the installer and outside tests,
so no build step or CI job parses anything with it. Fixed because the entire
point of a reference example is that people copy it forward, and ADR-1671
presents this one as the pattern for the platform.

Ports the linear per-segment validation that #2928 gave the production module,
so the two copies agree: both parse the real CONTEXT.md to 415 predicates
across 20 classes with 0 duplicates. Doubled-dot ids are now rejected here
too, matching production, and the grammar comment records it.

Also refreshes the example's committed index, which #2928 made stale when it
removed the duplicate predicate from CONTEXT.md.

Closes #2944

* test(#2944): guard predicate-index sync and example/production parity

Two regression tests for the two defects in this PR.

Index sync: asserts the committed docs/CONTEXT-INDEX.json equals a fresh parse
of CONTEXT.md, naming any diverging predicate ids. The merge race that reddened
next was invisible to both PRs involved and only surfaced on the next PR to run
lint:ci; this puts the same check inside the suite, which runs on every PR, and
a mutation test proves the assertion is not vacuous.

Example/production parity: asserts both copies of the parser report the same
count, classes and duplicates for the real CONTEXT.md, and agree verdict-for-
verdict over a table of id shapes. The divergence WAS the bug — production went
linear-time while the example kept the backtracking regex, with nothing
asserting they agreed. Also pins the example rejecting a 60-dot id, with the
clean rejection as the binding assertion and wall-clock only as a smoke check.

Notes a real tension rather than hiding it: ADR-1671 says the example sits
outside tests/, and this imports it. The ADR's intent is that the example is
not compiled, packaged or installed — not that it may silently rot. A parity
guard does not ship it. The file states this so a reviewer can object.

* fix(#2944): address both isolated review passes

Two independent reviewers (correctness and security axes, neither the author).
Security found nothing — it measured linearity to 100k chars across dots,
hyphens, underscores and mixed classes, and showed prototype pollution is
structurally unreachable because the first-segment pattern forbids
lowercase and underscore-leading ids. The correctness pass found three
blockers, all real.

Blocker: the parity test violated ADR-1671 verbatim. The ADR lists FOUR
exclusions for the reference example, the fourth being the CI test suite, and
the test imported it from tests/ while its own justification comment cited only
three -- constructing a rationale around the exclusion it broke. Moved to
scripts/lint-example-parser-parity.cjs wired into lint:ci; a lint script is not
the test suite, so the exclusion stands. The test file keeps only the
docs/CONTEXT-INDEX.json sync check.

Blocker: the mutation test leaked its temp dir. Its callback took no `t`, so a
failing assertion skipped the bare cleanup call. Now registered via t.after(),
matching the convention adr-index-gate.test.cjs documents.

Blocker: the example's own committed index carries the identical merge-race
staleness this PR fixes for the production one, and nothing guarded it.
Deliberately NOT fixed by wiring the example's --check into CI: that artifact
bakes line numbers, so it re-drifts on any unrelated CONTEXT.md line shift --
exactly ADR-1671 open question 4 -- and would make CI routinely red. The new
lint asserts the line-INDEPENDENT facts instead: count, class map, duplicate
set, and every (id, value) pair. Proven non-vacuous both ways: mutating a value
fails and names the id, mutating only a line number passes.

Major: a real divergence the parity claim would have missed. Production rejects
values containing an embedded CR, LF, U+2028 or U+2029; the example did not, so
a value with an embedded lone CR was rejected by one copy and accepted by the
other. Ported, and now covered by the parity table.

Also, found while verifying rather than reported: malformed diagnostics covered
only empty values. A doubled-dot id, a space in an id, and a lowercase-leading
id were all dropped silently. That contradicts the module's own intent -- a
typo should be diagnosable, and a space in an id is a likely one -- and
predicates are contractually cited, so a silently vanished predicate is the
failure mode that matters. Each rejection class now carries a named reason in
both copies, while ordinary inline code still yields none.

Trues up counts my own change staled: the example README and ADR-1671's
prototype figures said 416 and 393/18 against a real 415/20/0.

Closes #2944

* chore(#2944): backfill changeset PR number 2950

---------

Co-authored-by: sim <sim@local>
2026-07-31 15:44:19 -04:00
Tom Boucher
07603df8f2 fix(#2647): code-fixer worktree under .claude/worktrees/, not a hardcoded /tmp path (#2942)
* test(#2647): failing-first — fixer worktree path must be repo-relative not /tmp

* fix(#2647): place code-fixer worktree under .claude/worktrees/, not /tmp

The gsd-code-fixer agent hand-rolled its worktree at a hardcoded
`/tmp/sv-${padded_phase}-reviewfix-XXXXXX` mktemp path. On Windows/Git Bash
that landed OUTSIDE the project tree — outside the agent session's permission
allowlist, so every Read inside the worktree prompted (~25/run) — and mktemp's
MAX_PATH-avoidance substitute produced an un-removable `C:/mvwtNN` path.

Place the worktree repo-relative under `.claude/worktrees/` (the same dir the
harness-managed executor worktrees use: gitignored via `.claude/`, inside the
session's permission scope), with a $$-PID + epoch suffix for concurrency
uniqueness (replacing mktemp's XXXXXX). $main_repo is resolved the same way
the cleanup tail already resolves it.

Three sites updated: setup_worktree bash, concrete-steps prose, critical_rules.
The #2990 `-b "$reviewfix_branch"` invariant is preserved (the folded test
asserts it). Failing-first regression added to the #2990 suite in
tests/agent-frontmatter.test.cjs.

* test(#2647): update #2686 path assertion to expect .claude/worktrees/, not /tmp

The #2686 regression test encoded the worktree location as a hardcoded
`/tmp/sv-` path (matching sibling GSD agents at the time). #2647 showed that
breaks Windows/Git Bash (worktree outside the project tree → permission prompts;
mktemp MAX_PATH substitute un-removable). Update the #2686 path assertion to
require the repo-relative `.claude/worktrees/` location and forbid `/tmp/sv-`.
The #2686 isolation + cleanup assertions are unchanged.

* fix(#2647): word-boundary wt= parse + ack the fixer growth vs next

Two follow-ups to the #2647 GREEN run:
- parseWtAssignments matched `prior_wt=` (no word boundary), polluting the
  set and tripping the repo-relative + concurrency-unique assertions. Anchor
  on (?:^|\s)wt= so only the real worktree-path assignment is captured.
- emitted-attribution: gsd-code-fixer.md grew 1875 bytes vs origin/next. Update
  the emitted-drift-ack entry to attribute the #2647 worktree-path change
  (supersedes the prior #2825 attribution, whose growth is already in next).

* fix(#2647): address review — validate padded_phase at the sink + tighten test

Code-review + security-review both APPROVED with one actionable minor:
padded_phase is interpolated into a worktree PATH and a git BRANCH NAME, but
was only validated by the orchestrator (code-review-fix.md), not at the agent
sink. The agent prompt is a literal bash contract any caller can spawn, so add
a `[[ =~ ^[0-9]+(\.[0-9]+)?$ ]]` self-defense check rejecting traversal/shell
metachars (defense-in-depth; not a present vuln — the only caller validates).

Also tighten the concurrency-uniqueness test to require BOTH $$ AND $(date +%s)
(either-alone was too lax per review). Update the emitted-drift-ack reason to
cover the added validation growth.

* changeset(#2647): code-fixer worktree under .claude/worktrees not /tmp

* changeset(#2647): backfill PR number 2942

* chore(#2938): regenerate stale docs/CONTEXT-INDEX.json on next

#2938 (#2928) updated the CONTEXT.md RULESET prose for the new per-PR
emitted-drift-ack fragment mechanism (#2914) but shipped a CONTEXT-INDEX.json
generated from the OLD prose. lint:generated-sync fails on every PR that
rebases onto next after #2938 (the regen produces a 3-line diff bringing three
RULESET entries — AGENT_SIZE_BUDGET, EMITTED_ATTRIBUTION, WORKFLOW_SIZE_BUDGET
— in sync with the prose already on next). Mechanical regen via
`node scripts/gen-context-index.cjs --write`; idempotent; surfaced by the
#2647 rebase. No behavioral change.

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 14:45:51 -04:00
Tom Boucher
05b170e448 chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam

Productionizes the ADR-1671 Option-E reference example as a real module:
src/context-predicates.cts (parser + selector + index builder) compiled to
gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's
--check/--write drift-guard idiom and wired into lint:generated-sync.

Parser behavior is deliberately prototype-equivalent in this commit so the
next commit's regression matrix binds to the real defects rather than to a
missing module.

Two locked design deviations from the prototype:
- duplicates carry a count, not line numbers
- the committed index carries no line field at all, resolving ADR-1671 open
  question 4: an artifact without line numbers cannot drift on a line shift,
  so promoting --check to a CI gate does not make it routinely red

Also reconciles the one remaining duplicate predicate ID
(RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording
is removed) so the gate can land fail-closed on duplicates.

Refs #1671

* test(#2928): failing-first matrix for the predicate fact-store

Adds the regression matrix from the phase test plan: parser declaration
forms, fence and comment regions, ID/value grammar boundaries at
limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard
CLI, the selector query surface, and four document-shaped fast-check
properties.

Seven rows are RED for behavioral reasons against the ported parser:
indented-bare, star-list, plus-list and numbered-list declaration forms are
dropped; a tilde fence and a four-backtick fence containing a shorter fence
are not skipped; and a multi-line HTML comment is parsed as live. Eleven
selector rows are RED because the query surface is not wired yet.

Negative fixtures come from real repo documents that predate the grammar
(CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the
fixture-provenance rule, and the property generators are document-shaped
rather than seeded from our own serializer.

Refs #1671

* fix(#2928): consume the shared fence scanner, relocate the index, wire the selector

Drives the failing-first matrix green.

Parser: replaces the ported naive triple-backtick toggle with the shared
markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord
gain an export keyword — the only change to that module, which has 71
upstream dependents — because it already returns line-indexed spans, which
is exactly what a line-reporting parser needs. It also already documents
itself as the second copy of the fence state machine pending consolidation;
adding a third copy here would have been the generative-fix divergence this
repo warns about. A parity suite now pins predicate fence-skipping against
that scanner across eight fence shapes. HTML-comment skipping stays local
because the sectionizer has no comment scanner. Declaration forms widen to
indented-bare, star, plus and numbered list items.

Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The
remote matrix run caught the original choice — a committed .cjs there ships
~120KB of CONTEXT.md prose into a runtime module, and two content guards
fired truthfully on it (a leaked .claude install path, and four hardcoded
package-name literals). Neither guard was allowlisted; the artifact moved
instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to
require it — it is a drift-detection artifact, so the selector parses
CONTEXT.md live and is always current.

Generator: adds a frozen REASON enum and --check --json so the gate's
outcome is asserted structurally instead of by matching prose, and
--context-path/--index-path so tests drive the real CLI against a temp tree
with no filesystem monkeypatching.

Selector: gsd_run query context-predicates with --class/--prefix/--contains,
structured output carrying a matched count, own-property guards, and no
project-root resolution. Registering it exposed that the query dispatch
table and the usage string had drifted: a new parity test found 20 routed
commands missing from the usage list, all added here rather than deferred.

Refs #1671

* test(#2928): lock the newly-public scanFencedBlocks contract

Exporting scanFencedBlocks made it public API for the first time, so it
needs its own contract test independent of the consumer that motivated the
export. Memtrace's co-change analysis flagged the gap: this suite changes
together with markdown-sectionizer.cts 8 times in 90 days and was absent
from the diff.

Covers the documented rules: 0-based indices, -1 for an unterminated fence,
the same-char/>=length/no-trailing-text closer rule, a shorter fence inside
a longer one staying content, CommonMark 4.5 backtick-in-info-string, and
<=3-space indent tolerance.

Refs #1671

* fix(#2928): address both isolated review passes

Two independent reviewers (correctness axis and security axis, neither the
author) found seven findings. All are fixed here with regression tests; none
deferred.

BLOCKER — comment-blind fence scanning caused silent, permanent predicate
loss. The HTML-comment scan and the fence scan ran as two independent passes,
and the fence scanner is comment-blind, so a fence delimiter inside an HTML
comment with no later close read as an unterminated fence and skipped every
remaining line to EOF. Worse, the drift-guard could not catch it: it diffs
against a baseline produced by the same corrupted parse. The two constructs
now interleave in a single pass so each suppresses the other's boundary
detection while active, covered in both directions. The parity suite still
binds this scanner to markdown-sectionizer's for comment-free documents, so
the two cannot diverge unnoticed.

BLOCKER — the selector was not consumed anywhere, leaving the phase's
acceptance criterion unmet. Now wired into the pre-work predicate-citation
step in contributor-standards, which is the repo's actual brief-assembly
path; no code-level brief assembler exists to wire into.

MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id
regex nested a dot-containing character class inside a dot-prefixed repeat,
so N consecutive dots had exponentially many partitions: 40 dots took 565ms
and growth was exponential. CI runs this parser over a pull request's own
CONTEXT.md, so any contributor could have hung a shared runner with one
line. Replaced with linear per-segment validation. Doubled-dot ids are now
rejected; the real document contains none.

MAJOR — the duplicate-id gate had only ever been proven on synthetic
fixtures. A test now re-inserts the exact line this branch removed and
asserts the real generator names it.

MAJOR — --check together with --write silently let write win, turning the
gate into a writer; a missing path value resolved to the cwd and leaked an
EISDIR stack trace. Both are now clean usage errors.

MINOR — the hoisted skip-list was exported as a live mutable Set; replaced
with a read-only predicate. MINOR — flag-shaped selector values were
unmatchable; the inline --flag=value form now provides the escape hatch.

Refs #1671

* chore(#2928): backfill changeset PR number 2938

---------

Co-authored-by: sim <sim@local>
2026-07-31 13:17:01 -04:00
Tom Boucher
c043f2946c fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next

tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900.
Every entry is scoped to the diff that introduced it (#2789), so once merged
to next it is at the base by definition -- spent and inert. Its presence is
still load-bearing though: each PR rewrites the paths map wholesale, making a
persistent base copy a shared cell. Five of six conflicting PRs in the open
queue collided on this file and nothing else.

Deletes the stale document and adds a push-to-next guard asserting it stays
absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check
against the base is the #2768 shape #2789 exists to end.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2914): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2914): per-PR ack fragments instead of one shared mutable file

The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json
whose paths map every PR rewrote wholesale. That is a shared mutable cell: any
two PRs needing an ack edit the same lines and conflict. Five of six conflicting
PRs in the open queue collided on this file and nothing else.

Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same
shape .changeset/ already uses to solve this exact problem. Two PRs pick
different filenames, so they cannot collide, and fragments lingering on next
are harmless rather than toxic.

The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An
earlier delete-only attempt failed verification twice: the ratchet lost the
spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no
ack. Relocating preserves every acknowledgment.

The legacy single file is still READ (unioned with the fragments) because five
open PRs carry it; dropping support would break all of them. A duplicate path
key across sources is a hard error, never last-wins.

The push-to-next guard is retargeted accordingly: it now asserts only that the
legacy SHARED file never reappears on next. Fragments may persist harmlessly.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:15:29 -04:00
Tom Boucher
d2d2f7c088 fix(#2848): non-Latin titles no longer produce empty slugs (Cyrillic transliteration) (#2934)
* test(#2848): add failing-first regression for non-Latin slug transliteration

generateSlugInternal and slugify both strip non-ASCII chars with no
transliteration step, so an all-Cyrillic title reduces to an empty slug.
12-row matrix: Cyrillic regression (both impls), Latin negative control,
multi-letter mappings, soft/hard sign drops, Ukrainian extras, null
contract, CJK unaffected, mixed scripts, truncation parity, slugify's
distinct no-truncate contract.

* fix(#2848): transliterate Cyrillic titles to ASCII before slug strip

Both generateSlugInternal (src/core-utils.cts) and slugify
(src/gsd2-import.cts) stripped non-ASCII with no transliteration, so an
all-Cyrillic title reduced to an empty slug, producing unnamed phase
directories (01-) and empty milestone_slug init JSON.

Add a shared transliterateForSlug primitive (core-utils) covering Russian
+ the reported Ukrainian/Belarusian extras (і ї є ґ ў), with multi-letter
mappings (ж→zh ч→ch ш→sh щ→sch ю→yu я→ya) and dropped soft/hard signs
(ъ ь). It runs BEFORE the existing ASCII filter, so Latin-script text
hits zero map entries and is byte-for-byte unchanged (negative control).
slugify consumes the shared primitive, preserving its distinct single
hyphen-strip + no-truncation contract. CJK/unmapped scripts keep the
existing strip-to-ASCII behavior.

Also corrects two test assertions to match the chosen й→y mapping and the
б→b (not bie) transliteration.

* changeset(#2848): Fixed — non-Latin slug transliteration

* changeset(#2848): backfill PR number 2934

---------

Co-authored-by: sim <sim@local>
2026-07-31 10:56:11 -04:00
Tom Boucher
81eeb8a53a docs(#2926): refresh ADR-1671 with findings re-verified on next (#2936)
Re-measured the Option-E prototype's reported figures against CONTEXT.md on
next (2026-07-31) and recorded the delta rather than overwriting the June
numbers:

- index counts 393/18 (2026-06-24) -> 416/20 today; CONTEXT.md gained the
  PROBE (11) and PROHIB (10) classes
- of the 3 duplicate predicate IDs, only RULESET.WORKFLOW_MARKDOWN.FENCES
  remains; the two RULESET.GEMINI.* went with the Gemini runtime removal
- gen-context-index.cjs --check exits 1 on next, so Phase 0's "--check green
  in CI" criterion is unmet (invisible to CI: the example sits outside tests/)

Adds Open question 4 (index keyed on baked line numbers re-drifts on any
CONTEXT.md line shift, which matters once Phase 1 promotes --check to a CI
gate), and records the reviewer-proposed eval-gate question as resolved by
the PROBE.*/PROHIB.* predicate classes (ADR-550 D4/D7, ADR-1606).

Also names both surfaces of the Windsurf 12 KB throw in Decision 2, since it
is duplicated byte-identically in bin/install.js and
src/runtime-artifact-conversion.cts.

Docs-only. No code, no runtime-loaded text, no behavior change.

Co-authored-by: sim <sim@local>
2026-07-31 10:52:52 -04:00
Rezolv
76b7d73039 fix(#2733): route gate-passed spec-phase paths into the probe steps (#2779)
* fix(#2733): route gate-passed spec-phase paths into the probe steps

All four gate-passed transitions in spec-phase.md said "Jump to Step 6",
textually bypassing the mandatory Step 5.5 edge-completeness and Step 5.6
prohibition-completeness probes. Steps 5.5/5.6 were spliced between Step 5
and Step 6 by two later feature commits and the pre-existing jumps were
never re-pointed, so no jump instruction in the file reached Step 5.5 at
all and both probes were unreachable dead prose.

Re-point the four gate-passed jumps (lines 129, 162, 168, 170) to Step 5.5.
Control then flows 5.5 -> 5.6 -> 6 as the probes' own preconditions
prescribe. The max-rounds "write anyway" bypasses and the probes' own
"proceed to Step 6" exits are deliberately unchanged.

Add tests/spec-phase-probe-reachability.test.cjs, which derives the
mandatory probe steps from the file's own headings rather than hardcoding
5.5/5.6, so a future spliced-in probe step is covered without editing the
test. It also locks the two coupled constraints: the max-rounds bypass must
not be redirected into a probe, and each probe must keep its own onward exit.

The existing probe contract tests are untouched and still pass; both scope
from the "## Step 5.5"/"## Step 5.6" heading onward and were structurally
incapable of observing the upstream jump text.

* chore(changeset): Fixed fragment for #2779 (spec-phase probe reachability)

* fix(#2733): route Step 5.5's own soft gate into Step 5.6

Round-1 review blocker. The four upstream gate-passed jumps were re-pointed to
Step 5.5, but Step 5.5's own terminal soft gate at :305 still read "proceed to
Step 6" - so the COMMON path (all applicable edges resolved) skipped the
prohibition-completeness probe outright. Same defect class as the four this PR
already fixed, on the success path of the very step being fixed: the SPEC shipped
with an empty Prohibitions section instead of an empty Edge Coverage one.

Its sibling at :393 is byte-identical yet correct, because Step 6 genuinely
follows Step 5.6. Position, not phrasing, is the discriminator.

The guard could not see it: the transition matcher keyed only on the literal
"Jump to Step", and :305 says "proceed to Step". Widened it to a verb alternation
(jump/proceed/continue/go/return/skip + "to Step N", case-insensitive) and
renamed it TRANSITION_RE to match what it now models. This makes the file's own
docstring promise - that a future spliced-in probe is covered without editing the
test - true for a step whose exit is worded differently. Verified no false
positives: the two pre-existing "continue to Step 3/4" transitions are upstream
of both probes but target pre-probe steps, and the max-rounds bypass block
contains no step transitions at all.

Fail-first verified before fixing :305 - with the widened matcher against the
unfixed workflow the guard fails naming exactly "spec-phase.md:305 jumps to Step
6, skipping mandatory Step 5.6", 4 pass / 1 fail; after the fix, 5/5. The two
sibling probe contract tests stay 16/16.

Also from review:

- STEP_HEADING_RE gains an explicit \r? before $. Without it, on a CRLF checkout
  `.` stops before the \r and the unanchored $ fails to match, yielding ZERO
  steps and vacuously passing every assertion in the file. Not live today
  (.gitattributes forces eol=lf) but this repo has a recurring CRLF-regex bug
  class, so the guard no longer leans on it.
- allow-test-rule category corrected to source-text-is-the-product; the previous
  runtime-contract-is-the-product is not one of the six recognized categories
  (CONTRIBUTING.md:609-619).
- changeset body given the documented bold-lead-in form.
- emitted-drift ack reason updated: +8 -> +10 bytes across five transitions
  (31987 -> 31997), DEFAULT tier, cap 40960.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 10:20:17 -04:00
Tom Boucher
d49a7d0c4d fix(#2853): roadmap.update-plan-progress preserves hand-written annotations (#2916)
* test(#2853): add failing-first regression for plan-progress annotation preservation

The count-bump regex's trailing [^\n]+ swallowed the whole Plans line and
the replacement wrote back only the regenerated count, deleting any
hand-written annotation after it. 8-row matrix covers bold/plain forms,
bare template form, executed path, CRLF, and idempotency.

* fix(#2853): preserve hand-written annotations in roadmap plan-progress bump

The count-bump regex's trailing [^\n]+ swallowed the entire Plans line and
the replacement wrote back only the regenerated count, deleting any
hand-written prose after the count (e.g. a gap-closure annotation).

The verb owns the count token only. Capture the existing count token ($2)
and the trailing line text ($3), and rebuild the line as
<label><new count><surviving text>. Trailing text is preserved ONLY when a
real count token preceded it, so the fresh-template bracketed placeholder
(`[Number of plans…]`) is still replaced cleanly rather than glued after
the count (pre-#2853 behaviour on the template path preserved). CRLF \r is
preserved via [^\r\n].

Widens replaceInCurrentMilestone to accept a replacement callback (needed to
branch on whether the count group matched). The bare Plans: checklist header
is still skipped — the lazy match lands on the summary line first and a
count-less bare header yields no count to anchor preservation to.

* changeset(#2853): backfill PR number 2916

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: sim <sim@local>
2026-07-31 09:54:55 -04:00
Tom Boucher
8635cc447a chore(#2913): prune changeset fragments already promoted in the v1.9.1 CHANGELOG (#2922)
The v1.9.1 finalize consumed these 8 fragments on hotfix/1.9.1 and that
deletion reached main, but the back-merge did not propagate it to next
(c3c6566ac does not touch .changeset/). All 8 are already rendered into
the [1.9.1] section of CHANGELOG.md on next, so leaving them would make
the next release emit duplicate entries for changes already shipped.

Verified each fragment's lead phrase is present in next's CHANGELOG
before removal. No content is lost.

Co-authored-by: sim <sim@local>
2026-07-31 09:51:55 -04:00
clezcoding
9bd0dbf0dd docs(#2534): rewrite your-first-project tutorial for beginners (#2569)
* docs(#2534): rewrite your-first-project tutorial for beginners

Adds a loop mental-model primer (Mermaid), per-step "what just happened"
callouts, a prerequisites flow, a glossary and a troubleshooting table.
Same commands, same .planning artefacts, same to-do CLI example.

Closes #2534

* docs(#2534): make the tutorial runtime-agnostic (all IDEs)

Adds a "Pick your runtime" section (Cursor, Claude Code, OpenCode, Codex,
Gemini CLI, Copilot, Windsurf, Kilo, Cline, Qwen, Antigravity, ...) with the
installer flag and command syntax per runtime (/gsd-*, /gsd:* colon form, and
Cline rules). Keeps the same guaranteed worked example and .planning artefacts.

Closes #2534

* docs(#2534): address review - drop gsd-cursor aside + dead hero comment

- Remove the '(pair with the gsd-cursor EoS ...)' parenthetical from the Cursor row.
- Remove the commented-out reference to a non-existent hero asset.
(Gemini CLI references retained: --gemini is still live in bin/install.js on next.)

Closes #2534

* docs(#2534): fix review defects (keep multi-runtime)

- Replace dead Gemini CLI / --gemini with its live successor Antigravity
  (#1928); remove the invalid --gemini row/flag everywhere.
- Replace fabricated Step 1 output with realistic installer lines
  (71 skills/commands + destination suffix; exact lines vary by runtime).
- Fix 'Skip research' -> choose 'No' on the real Research prompt.
- behaviours -> behaviors (2x).

Multi-runtime 'Pick your runtime' section retained per author intent;
scope re-approval on #2534 still pending.

* docs(#2534): scope tutorial back to single-runtime (Claude Code)

Per trek-e's 2026-07-27 review, resolve the multi-runtime blockers by
returning to the approved scope:

- Remove the 'Pick your runtime' table + per-runtime notes; leave a one-
  line pointer to docs/how-to/install-on-your-runtime.md (which already
  documents all runtimes) rather than duplicate it (avoids the drift).
  This kills Blocker 1 (Antigravity is slash-hyphen, not colon) and
  Blocker 2 (Codex is $gsd-*) at the source.
- Step 1 uses --claude concretely; config-dir prose is Claude-local.
- Step 5: fix singular researcher (plan-phase spawns one gsd-phase-
  researcher), and make the research choice consistent with Step 3
  (choose 'Skip research'); drop the RESEARCH.md artifact line.
- Glossary/troubleshooting/prereqs/Step 2 de-multi-runtimed.

Returns the PR to #2534's approved 'docs-only, same commands' scope.

* docs(#2534): correct tutorial prerequisites and outputs

* docs(#2534): match tutorial research prompts to workflow

* docs(#2534): complete tutorial step guidance

---------

Co-authored-by: clezcoding <clezcoding@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 09:48:54 -04:00
Rezolv
2f66788cd3 docs(#2619): add ADR-2619 observability and shareable diagnostics (#2862)
Completes ADR-0174 §6's observability rollout and adds the outbound trust
boundary that ADR-1577's inbound boundary has no counterpart for.

D1 (wire the seam behind the existing opt-in gate) shipped via #2620 / PR
#2621. D1b records that the unconditional stderr-on-error rule at 0174:105
is the target state, deferred behind an explicit --json-errors envelope
version plus migration note -- disclosed as a partial supersede rather than
retconned. D2-D5 are Directional; D6's non-goals are binding.

Regenerates docs/adr/README.md via scripts/gen-adr-index.cjs --write.

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 09:43:04 -04:00
Tom Boucher
49793465d7 docs(#2915): how-to for listing a reviewer lane, and correct the stale listing section (#2917)
* docs(#2904): how-to for listing a reviewer lane in the registry

#2912 shipped the Reviewer Lane Registry, which lands in 1.9.1. Two docs
consequences.

New: docs/how-to/list-your-reviewer-lane.md. A Diataxis how-to for the
publish task -- which of the three catalogs applies (and why a runtime
carrying a reviewer body lists under its primary install shape instead),
opening the required discussion thread BEFORE the PR, the three fields
that reject entries most often (slug grammar differs from id, flags stay
kebab when the slug is snake, install/uninstall must be copy-pasteable),
regenerate-don't-hand-edit, and register-once-then-Releases. Links the
registry README for the field table rather than duplicating it -- the
spec is reference, this is the task flow.

Corrected: ship-a-reviewer-lane.md said "Listing your lane is not wired
yet" and pointed at #2904 as future work. #2906 merged at 11:25Z and
#2912 at 12:11Z, so that section shipped false the moment the registry
landed. Replaced with the publish pointer.

Indexed the new guide and the generated catalog in docs/README.md, and
added the guide to develop-a-capability.md's ecosystem list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2904): caveat credential-bearing configKeys in both the guide and the spec

Isolated security review found the worked entry's `configKeys:
["acme.api_key"]` modelled storing a live credential with no note on
where that value ends up.

Verified: config values are written in plaintext to
.planning/config.json (docs/CONFIGURATION.md:227 -- masking is
display-only, "that file is the security boundary"), and
planning.commit_docs defaults to true (:466). So a credential declared
that way lands in the installing user's git repository unless they have
gitignored .planning/. None of the twelve first-party lanes does this --
they own only review.models.*, host, and prompt-budget keys.

The pattern originates in docs/registries/README.md:221, shipped by
#2912, so the caveat goes on BOTH surfaces rather than only on the copy
that inherited it -- the spec's example is what future authors will read
first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2915): backfill changeset PR number

pr: 0 -> 2917.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 09:29:51 -04:00
Tom Boucher
a9aba61f89 Merge pull request #2920 from open-gsd/chore/backmerge-main-to-next-4f1cce98
chore: back-merge main → next (4f1cce98)
2026-07-31 09:15:08 -04:00
github-actions[bot]
c3c6566ac2 chore: back-merge main into next (4f1cce98) 2026-07-31 13:14:39 +00:00
Tom Boucher
932f99907c Merge pull request #2919 from open-gsd/chore/sync-next-version-1.9.1
chore: sync next package version to 1.9.1
2026-07-31 09:12:18 -04:00
github-actions[bot]
854c93533c chore: sync next package version to 1.9.1 2026-07-31 13:12:09 +00:00
Tom Boucher
4f1cce9875 Merge pull request #2918 from open-gsd/hotfix/1.9.1
chore: merge release v1.9.1 to main
2026-07-31 09:12:06 -04:00