Commit Graph

2399 Commits

Author SHA1 Message Date
Tom Boucher
f0bb0787c9 fix(#2640): report truthful state_updated + keep progress frontmatter in sync after phase remove (#2974)
* test(#2640): add regression for state_updated false positive + stale progress

Three cases: (1) state_updated reflects actual content change, not just
file existence; (2) progress.total_phases resync'd even when the body lacks
'Total Phases:' (the no-op guard was skipping syncStateFrontmatter);
(3) state_updated is false when STATE.md doesn't exist.

* fix(#2640): report truthful state_updated + force frontmatter resync

Two defects in cmdPhaseRemove:

1. state_updated was fs.existsSync(statePath) — trivially true, since the file
   existed before and readModifyWriteStateMd never deletes it. Now captures
   the boolean return from readModifyWriteStateMd (changed from void to
   boolean: true when content was written, false on no-op).

2. progress.* frontmatter stayed stale when the body lacked 'Total Phases:'
   or 'of N' — readModifyWriteStateMd's no-op guard (#948) skipped
   syncStateFrontmatter when the body transform was unchanged. Now the
   transform forces a body diff when a phase was actually removed, so the
   guard passes and syncStateFrontmatter rebuilds progress.* from the
   post-deletion disk/ROADMAP state.

* fix(#2640): address review — gate forced-diff on targetDir, strengthen assertions

Two MAJOR findings from isolated adversarial review:
1. Forced-diff injected a spurious 'Total Phases:' line even when no directory
   was removed (targetDir === null). Now gated on targetDir !== null.
2. Test #2 asserted 'not 3' instead of '2' — would pass for any wrong count.
   Now asserts exact value. Test #1 strengthened to assert body Total Phases
   and frontmatter total_phases both equal 1.

* chore(#2640): add changeset fragment

* chore(#2640): backfill changeset PR number 2974

---------

Co-authored-by: sim <sim@local>
2026-08-01 11:56:34 -04:00
Tom Boucher
000a322489 fix(#2941): hint at global: prefix when bare skill name matches a global skill (#2973)
* test(#2941): add regression for bare-skill-name global: hint

When a bare skill name matches an existing global skill, the skip warning
must hint at the global: prefix. When no global skill matches, the original
'Skill not found' warning is unchanged.

* fix(#2941): hint at global: prefix when a bare skill name matches a global skill

When buildAgentSkillsBlock skips a bare skill name that doesn't exist as a
project-relative path, check if it matches an existing global skill. If so,
append a hint to the warning: 'a global skill named X exists; use global:X
to reference it'. When no global skill matches, the warning is unchanged.

getGlobalSkillDir and getGlobalSkillsBase are already imported in this module
for the global: branch. The hint is guarded on globalSkillsBase being non-null
(runtimes without a skills directory don't support the prefix).

* chore(#2941): add changeset fragment

* chore(#2941): backfill changeset PR number 2973

---------

Co-authored-by: sim <sim@local>
2026-08-01 10:58:50 -04:00
Tom Boucher
608be0e7cf fix(#2913): distinguish empty cherry-pick from genuine conflict in hotfix create (#2970)
* test(#2913): add regression for hotfix empty-cherry-pick discrimination

Two layers: (1) real-git test proving the discrimination logic (check for
unmerged paths → skip if empty, abort if conflict) is correct; (2) source-text
assertions proving the logic and summary heading are in release.yml.

Before the fix: release.yml treats any non-zero cherry-pick exit as a conflict,
so an already-applied commit (empty pick) aborts the entire hotfix create.

* fix(#2913): distinguish empty cherry-pick from genuine conflict

git cherry-pick exits non-zero for BOTH genuine conflicts AND empty picks
(the change is already present by content). The hotfix create job treated
every non-zero exit as a conflict, so an already-applied commit (notably the
structural 'chore: sync next package version' that follows every release
finalize) aborted the entire run.

Now the error handler checks for unmerged paths (git diff --diff-filter=U):
- No unmerged paths → already applied by content → skip, record, continue.
- Unmerged paths present → genuine conflict → existing abort/push/exit-1
  behavior and operator guidance, unchanged.

The job summary now has a separate 'Skipped (already applied by content)'
heading, distinct from 'Skipped (feat/refactor/etc)'.

* fix(#2913): address review — post-skip continuation test + conflict observability

Two findings from isolated adversarial review:
1. MINOR (test gap): no test proved the sequencer is clean after --skip, so a
   regression switching --skip to --quit would pass green. Added a second
   cherry-pick after the skip asserting it succeeds.
2. MINOR (observability): SKIPPED_EMPTY was dropped on the conflict-exit path
   — already-applied commits before a genuine conflict were silently lost from
   the summary. Now the conflict summary emits them under a dedicated heading.

* chore(#2913): add changeset fragment

* chore(#2913): backfill changeset PR number 2970

---------

Co-authored-by: sim <sim@local>
2026-08-01 10:21:27 -04:00
Tom Boucher
d3305fc3a5 fix(#2858): exclude gen-emitted-baseline.cjs from npm tarball + class-extinction guard (#2968)
* test(#2858): add class-extinction guard for shipped-script require boundary

Every shipped scripts/**/*.cjs must be require-able using only shipped paths.
The guard resolves the tarball file list via npm pack --dry-run --json (not a
hardcoded list) and statically checks each require() call against the shipped
set. A script requiring ../tests/** (which does not ship) is a violation.

* fix(#2858): exclude gen-emitted-baseline.cjs from the npm tarball

scripts/gen-emitted-baseline.cjs is repo-only CI tooling (CI workflows +
test fixtures spawn it from a checkout). It requires three modules from
tests/, which does not ship — so in a published install it is
MODULE_NOT_FOUND at load time.

Add a targeted files[] negation (!scripts/gen-emitted-baseline.cjs) so the
script stays in the repo for CI use but does not ship. Other scripts that
ship and are required by bin/install.js (build-hooks.js,
fix-slash-commands.cjs, gen-capability-registry.cjs) are unaffected.

The class-extinction guard test in
tests/packaging-shipped-scripts-require-only-shipped.test.cjs ensures no
shipped script can require outside the shipped tree going forward.

* test(#2858): widen guard to .js + strip block comments (review fixes)

Two findings from isolated adversarial review:
1. MAJOR: the guard only checked .cjs files, but scripts/build-hooks.js
   ships and is required by bin/install.js — a .js file with a broken
   require would bypass the guard. Widened filter to /\.(cjs|js)$/.
2. MODERATE: the static parser could false-positive on require() calls
   inside /* */ block comments or inline // comments. Now strips both
   before matching.

* chore(#2858): add changeset fragment

* chore(#2858): backfill changeset PR number 2968

* fix(#2858): add issue ref to allow-test-rule exemption (ADR-456)

CI lint-tests caught: the allow-test-rule comment needs a 'see #NNN' ref
per ADR-456. Added '(see #2858)' to the integration-test-input exemption.

---------

Co-authored-by: sim <sim@local>
2026-08-01 09:29:15 -04:00
Tom Boucher
1db7dcd9bf fix(#2849): strip trailing hyphen after 60-char slug truncation (#2967)
* test(#2849): add failing regression for trailing-hyphen-after-truncation

The strip ran before .substring(0, 60), so a cut landing on a separator
produced a slug ending in '-'. Four cases: the exact issue repro (59 a's +
space + tail), a boundary landing before a separator, leading-hyphen survival,
and a long-Cyrillic transliteration+truncation case.

* fix(#2849): strip trailing hyphen after 60-char truncation

generateSlugInternal ran the ^-+|-+$ hyphen strip BEFORE .substring(0, 60),
so a title whose 60-character cut landed on a separator yielded a slug ending
in '-' — the very thing the strip step exists to prevent.

Reorder so the strip runs after truncation. Truncation cannot introduce a
leading hyphen, so the full ^-+|-+$ pass last is equivalent for leading
hyphens and fixes the trailing-hyphen-after-truncation case.

Latin-script output for titles ≤ 60 chars is byte-identical; only titles
whose truncation boundary lands on a separator change (from broken to clean).

* test(#2849): add all-separator collapses-to-empty boundary case

Surfaced by isolated adversarial review: pin the contract that input
which is entirely separators ('!!!', '!'.repeat(70)) reduces to '' —
not null, not a stray hyphen — both short and past the 60-char truncation.

* chore(#2849): add changeset fragment

pr:0 placeholder; will backfill the real PR number after the PR exists.

* chore(#2849): backfill changeset PR number 2967

---------

Co-authored-by: sim <sim@local>
2026-08-01 08:24:05 -04:00
Tom Boucher
0bb7525a62 fix(#2943): rename get-library-docs -> query-docs; correct the ctx7 fallback rationale (#2963)
* test(#2943): parity guard against the nonexistent get-library-docs tool

Second context7 naming drift after #2017 (which guarded the plugin-marketplace
PREFIX). #2017's guard only checks tools: frontmatter lines, not prose bodies —
which is where the broken tool NAME (get-library-docs) lived. The context7 MCP
server registers only resolve-library-id and query-docs; get-library-docs is a
stale copy from upstream's own README.

Scans the shipped prose surface (agents/, gsd-core/references|workflows/,
commands/gsd/, skills/) and fails if any artifact instructs an agent to call
mcp__context7__get-library-docs. Excludes tests/ (a fixture may use the name as
a negative input) and CHANGELOG/RELEASE-NOTES-LEGACY (history).

Fails-first: 4 offenders today (gsd-executor.md:29,
research-documentation-lookup.md:5, discovery-phase.md:68 & :104).

* fix(#2943): rename get-library-docs to query-docs and correct the ctx7 fallback rationale

The context7 MCP server registers only resolve-library-id and query-docs
(verified against upstream packages/mcp/src/index.ts); get-library-docs is a
stale name copied from upstream's own README. Four shipped prose sites instructed
agents to call a tool the server does not register, so every research path that
loaded the canonical reference either errored, fell through to the ctx7 CLI
branch, or fabricated a result.

- research-documentation-lookup.md, gsd-executor.md, discovery-phase.md (x2):
  get-library-docs -> query-docs, params context7CompatibleLibraryId/topic ->
  libraryId/query (the registered contract).
- Same files' ctx7 CLI fallback rationale: the cited cause
  (anthropics/claude-code#13898 'strips MCP tools from agents with a tools:
  frontmatter restriction') was wrong on two counts — #13898 is closed and was
  never about tools: frontmatter. Rewritten to describe the real mechanism
  (custom subagents cannot see project-scoped .mcp.json; they only inherit
  user-scoped ~/.claude/mcp.json). The fallback itself is kept.
- discovery-phase.md 'mode: code/info' dropped — query-docs takes libraryId +
  query only; the code-vs-concepts intent is now expressed via the query text.

resolve-library-id is unchanged (still registered upstream). CHANGELOG and
RELEASE-NOTES-LEGACY citations are historical record, left as-is.

* chore(#2943): add changeset fragment (pr:0 placeholder)

* test(#2943): widen parity-guard scan surface to docs/ (isolated-review finding)

The isolated adversarial review flagged that SCAN_DIRS omitted docs/, which
ships docs/AGENTS.md — agent-consumed prose carrying 8 mcp__context7__* refs.
No false negative today (it uses only the wildcard), but a future banned-name
addition there would slip through, recreating the exact drift this guard exists
to prevent. Add docs/ to the scan surface, with an EXCLUDED_FILES set for
historical record (docs/RELEASE-NOTES-LEGACY.md, CHANGELOG.md) that must not be
rewritten to satisfy the guard.

* fix(#2943): update shifted PROSE_ALLOWLIST line + acknowledge gsd-executor.md growth

The gsd-test gate caught two real consequences of the rationale rewrite in
agents/gsd-executor.md (the +2-line corrected mechanism description shifted
line numbers below it):

1. tests/no-bare-gsd-tools-command-position.test.cjs: the legitimate
   'gsd-tools query commit' descriptive mention moved from line 791 -> 793.
   Update the PROSE_ALLOWLIST entry to the new line (the mention is unchanged,
   just relocated by my edit above it). Without this the gate reports both a
   stale allowlist entry (791) and a new offender (793) for the same mention.
2. tests/emitted-drift-acks/2943-context7-tool-name.json: gsd-executor.md grew
   95 bytes (the accurate mechanism rationale is longer than the wrong one-line
   #13898 attribution it replaces). Acknowledge the growth with the reason.

Both are mandated by the gate, not optional. The rename itself (get-library-docs
-> query-docs) is byte-neutral-ish; only the rationale rewrite grew the file.

* chore(#2943): backfill changeset PR number 2963

---------

Co-authored-by: sim <sim@local>
2026-08-01 01:34:05 -04:00
Tom Boucher
73418c516f fix(#2956): scope Phase extraction to ## Current Position (3rd gen of #2444/#2567) (#2961)
* test(#2956): fail-first regressions for Phase scoped to ## Current Position

Third generation of #2444 / #2567. Stopped At / Paused At were scoped to
## Session; Phase (canonically in ## Current Position per templates/state.md)
was left unscoped, so a historical Phase: / **Phase:** line in an archive
section overwrites current_phase on every write. Since current_phase is
routing input for gsd-progress / --next, the rewind routes work to the wrong
phase.

Six failing-first regressions + one round-trip:
- shape B: bold **Phase:** 19 archive BELOW the section
- shape C: plain archive Phase: 19 ABOVE the section
- bootstrap h3 ### Current Position variant
- CRLF variant
- Phase token in decisions prose (over-broad-fix guard)
- Paused At read-path parity with the write seam (## Session)
- write-then-read round trip stays at 22 (read/write agreement)

Folded into tests/state.test.cjs (lint:regression-test-names bans a
new tests/bug-NNNN-*.test.cjs file).

* fix(#2956): scope Phase extraction to ## Current Position at both seams

Third generation of #2444 / #2567. Stopped At / Paused At were scoped to
## Session by those fixes; Phase (canonically in ## Current Position per
templates/state.md) was left unscoped, so a historical Phase: / **Phase:**
line in an archive section silently overwrote current_phase on every write.
Because current_phase is routing input for gsd-progress / --next, the rewind
routes work to the wrong phase.

Fix mirrors the proven #2444 seam exactly:
- new matchCurrentPositionSection helper (collectSection-based, CRLF-tolerant,
  level-flexible for the bootstrap ### Current Position h3 variant), sited next
  to matchSessionSection.
- read path (cmdStateSnapshot): extract Phase from matchCurrentPositionSection
  ?? body. Also scope Paused At to matchSessionSection ?? body so the read seam
  agrees with the write seam (which already scoped Paused At to ## Session).
- write path (buildStateFrontmatter): extract Phase from
  matchCurrentPositionSection ?? bodyContent.

stateExtractField itself is untouched (its bold/plain precedence is load-bearing
for other fields — the #3265 test depends on it), and preferNewerLastActivity is
untouched (Last Activity has no canonical section; its date-direction guard is
deliberate). Fall back to full-body when no ## Current Position section exists
so files without the heading keep current behaviour.

* chore(#2956): add changeset fragment (pr:0 placeholder, backfill after PR)

* test(#2956): make round-trip test actually trigger the write-path resync

The write-then-read round-trip test used 'state update Status "Executing"' on a
fixture with no Status field, so the update was a no-op (updated:false) and no
frontmatter resync ran through buildStateFrontmatter — the assertion on the
written frontmatter then failed not because the fix is wrong, but because no
write happened. Add a **Status:** field so the update performs a real field
update (updated:true) and forces the resync. Verified locally: pre-fix this
writes current_phase:19 (the archive value); post-fix it writes 22.

The code fix is correct (5 of 7 RED tests passed; the 2 failures were this
defective test). This is the 'fix the bad test' half of the TDD-loop rule.

* chore(#2956): backfill changeset PR number 2961

---------

Co-authored-by: sim <sim@local>
2026-08-01 00:02:41 -04:00
Tom Boucher
9ac0dfad58 chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor

Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared
context-composer seam. Its success condition is that review-prompt output does
not change, and the only authority on "did not change" is the behavior that
shipped before the refactor. Capture that behavior now, while it is still the
live implementation.

47 characterization cases, every `expected` value computed by executing the
current implementation rather than hand-authored — the independence
CONTRIBUTING.md "Fixture provenance (#2371)" asks for.

A corpus is only worth what it can detect, so this one was validated by
mutation rather than assumed. Five deliberate defects were injected and each
must be caught by at least one case:

  - the note reserve deducted unconditionally instead of only under pressure
  - the pressure test relaxed from `>` to `>=`
  - a no-op head-shrink still setting the shrunk flag
  - the per-plan floor dropped from the proportional share
  - drop order reversed

Two of those exposed real holes in the first cut of this corpus, and the cases
that close them exist because of it:

  - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a
    floored plan group, and the 1024-char floor absorbed the entire trim, so the
    mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap,
    which makes the strict inequality observable as context kept vs omitted.

  - No case reached proportional-truncate at all — B6 and B7 both hard-failed
    the min-set pre-check first, leaving planTruncationPct at 0 across every
    case and the floor semantics entirely unexercised. Rebudgeted to 700 and
    1100 so the min-set fits and the truncate step is actually reached; they now
    record 40.20% and 48.80%.

The A4/A10 families sweep the pressure boundary from both sides, which is where
this function has regressed before: CONTEXT.md's
LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions
that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band,
because the suite paired a trivially-fitting budget with a trivially-overflowing
one and never sampled between them. A4 pins that nothing is trimmed from the cap
down to 81 tokens under it; A10 pins that pressure fires at +1. Together with
A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures.

Two facts the corpus establishes that the design notes had wrong:

  - "" and null sections are NOT distinguished. applyBudget uses truthy checks
    throughout, so an empty-string section is treated as absent: not rendered,
    not dropped, never recorded in `omitted`. B13b pins this while the ladder is
    actively trimming, where only the non-empty `research` is dropped.

  - Sizing matters. B12/B13 were first written at a budget where both hard-failed
    the min-set check and returned "", so comparing them compared two empty
    strings and proved nothing.

Committed as its own commit, ahead of the refactor, and regenerated against the
pre-refactor implementation, so the oracle is demonstrably independent of the
change it will adjudicate.

Refs #2929

* refactor(#2929): extract the context-composer seam from prompt-budget

Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer —
per-runtime artifact emission — but it is walled inside the cross-AI review
pipeline. Lift it into a shared seam so later phases can call it, without
changing what the review pipeline emits.

ADR-1671 specifies the composer as "priority + binary-search cutoff to a
per-runtime budget". Read against the code it generalizes, that contract cannot
express the thing being generalized. applyBudget is not a cutoff: it is a fixed
five-step ladder in which each section carries its own shrink strategy, and only
three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N
lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor;
instructions and roadmap are never touched at all. A cutoff composer sorts by
priority and discards the tail — it has no way to say "shrink this one",
"truncate that one but never below 1 KB each", or "these three are the only
droppables, in this order". Building to the literal contract and routing
prompt-budget through it would have silently changed review-prompt output, which
is the one outcome this phase forbids.

So shrink strategies are the core abstraction here, and cutoff becomes one
strategy among them — the right one for per-runtime emission in Phases 3-4, not
for this ladder. That is an elaboration of the ADR's intent, not a departure
from it, and ADR-1671 is updated to say so.

Three decisions worth stating:

  - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan
    of surviving fragments and never a string. assemblePrompt's rendering is
    prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and
    owning it in the composer would force emission to adopt prompt-shaped
    rendering. The split is what lets one seam serve both consumers.

  - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its
    chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires
    for emission caps. The existing code converts a token budget to a character
    budget with a hardcoded `* 4`; that assumption is now an explicit
    `charsPerUnit` inverse, which is precisely what a byte unit needs in order to
    reuse this.

  - The entry point is `composeWithinBudget`, not `applyBudget`. That name
    already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter
    being an unrelated graph-edge budget. A third would make every symbol search
    in this repo ambiguous, and it already misresolves: preflight and impact
    queries for "applyBudget" return graphify's.

Behavior is unchanged and proven so: all 47 characterization cases reproduce
byte-identically, and the corpus is mutation-validated rather than merely green
(see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and
from eighteen mutable accumulators to two, both inside a helper copied verbatim.

estimateTokens deliberately stays in prompt-budget and keeps its exact math:
src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins
plan estimates and recorded actuals to that same scale, so moving or changing it
would silently break the calibration loop.

Refs #2929

* docs(#2929): document the context-composer seam and amend ADR-1671

Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new
domain modules), and a mutation-matrix entry for the new module.

The ADR amendment is the substantive part. ADR-1671 specified the composer as
"priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2
established that a cutoff alone cannot express the function the platform
generalizes, so the ADR now records shrink strategies as the core abstraction
with cutoff as one strategy among them, reserved for per-runtime emission in
Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned
against that contract and would otherwise be planned against a mechanism that
does not work.

The mutation-matrix entry is not bookkeeping. Stryker scores per module against
a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the
extracted code unmeasured while prompt-budget's own score floated free of the
logic it used to cover. context-composer gets its own entry at the same floor.

Refs #2929

* test(#2929): pin the effectiveBudget rounding mode in the parity corpus

An isolated correctness review found a real blind spot: mutating
`Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of
the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator
happened to produce a whole number, so floor, round and ceil all agreed and the
rounding mode was entirely unpinned by a corpus whose whole job is to pin
observable behavior.

Three cases fix that by straddling the .5 boundary:

  A11  95 * 0.90  = 85.5   floor 85, round 86  -> the two disagree
  A12  97 * 0.90  = 87.3   floor and round agree; ceil (88) does not
  A13  93 * 0.85  = 79.05  same guard at a non-multiple-of-10 margin, so the
                           margin arithmetic is exercised and not just the budget

A11 alone catches the round mutation; all three catch ceil. Regenerated against
the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so
the expanded corpus keeps the independence property the original capture had.

The corpus is now mutation-validated against seven injected defects, every one
caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink
setting its flag, the truncate floor ignored, drop order reversed, and both
rounding-mode changes.

Refs #2929

* feat(#2929): flexReserve floors and the byte-stable isolate prefix

Two of issue #2929's "Done when" items were unimplemented rather than deferred,
and an isolated review flagged them alongside my own audit. Both are part of
ADR-1671's composer contract, so shipping the seam without them would have left
Phases 3-4 building against a contract that does not exist yet.

flexReserve is a per-fragment floor in measure units that every strategy must
respect, which is what makes it different from the pre-existing floorChars: that
one is a chars-denominated detail of proportional-truncate alone and is retained
unchanged. A floored fragment is never dropped, is never head-shrunk below its
floor, and raises its own proportional cap. A fragment already smaller than its
floor is untouchable outright. Metadata gains `floored`, listing the ids whose
floor actually prevented a trim — a guarantee no caller can observe is a
guarantee no test can hold you to.

isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed,
never dropped, but still counted, because a prefix excluded from accounting
would silently under-count real context. Metadata gains `isolatePrefix` so a
caller can hash or assert on the exact bytes. Declaring an isolate fragment
after a non-isolate one throws: a prefix that is not at the front is not a
prefix, and accepting it would make the cross-runtime stability claim
meaningless.

Adds tests/context-composer.test.cjs for the exact new semantics and
tests/context-composer.property.test.cjs for the five invariants, including the
budget-monotonicity property the issue names explicitly. Both are registered in
the mutation matrix, since coverage does not migrate with relocated code.

prompt-budget uses neither feature, and its output is unchanged: all 50 corpus
cases still reproduce byte-identically.

Refs #2929

* chore(#2929): allowlist the prompt-budget parity suite

The parity corpus needs its own test file and that makes prompt-budget a
three-file module against a limit of two. The lint offers consolidation or an
allowlist entry with justification; the entry is the right call here.

Consolidation would mean folding the characterization suite into
prompt-budget.test.cjs, which is the one thing that should not happen to it. The
parity suite is a distinct concern with a distinct lifecycle: it is generated
rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own
scoring target, and its failure means something categorically different from a
unit-test failure — not "this behavior is wrong" but "observable output moved".
Burying it inside a general unit file would obscure exactly that signal.

The allowlist is an identity ratchet, so this entry pins today's three exact
filenames: adding a fourth still fails, and dropping back to two requires
removing the entry.

Refs #2929

* fix(#2929): register the new module with two gates it was missing

The remote matrix caught three defects that no local check could, because the
local runner is blocked in this repo and these suites had therefore never
executed. Eight failures, identical on node22 and node24, so nothing
environment-shaped.

Two are the new-module ripple. A net-new src/*.cts lands in six places and this
change had reached four of them — .gitignore, INVENTORY, the manifest, and the
CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not
be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet
baseline (a deliberate review-visible mirror of the matrix floors, which every
COVERED module must carry). Both are now registered, the ratchet at the same
floor of 66 the matrix declares.

The third was a test asserting an outcome it had made impossible. It set
budget:1 alongside a 400-char required fragment, so the group budget came out at
-99 and the proportional-truncate step was skipped entirely — the deliberate
"non-positive group budget is skipped, never clamped" rule inherited from the
original ladder. Nothing was trimmed, and the test then asserted a truncation.
Rebudgeted so the step actually runs, with the arithmetic written out in a
comment so the next reader does not have to re-derive why 120 rather than 80.

Fixing that surfaced a genuine bug in the composer. `floored` is documented as
recording fragments whose flexReserve prevented a trim that would otherwise have
happened, but the push sat in the else-branch of "content did not change", so it
only fired when nothing was trimmed at all. A fragment truncated to a
reserve-raised cap has also had a trim prevented — 40 characters' worth in the
test above — and was silently absent from the field that exists to make the
guarantee observable. The condition was already right; it was in the wrong
branch. Now recorded on both paths: a drop prevented outright, and a truncation
capped higher than the share alone would have allowed.

Parity is unaffected — prompt-budget never sets flexReserve, so the branch is
unreachable from every corpus path, and all 50 cases still match.

Refs #2929

* chore(#2929): backfill changeset PR number (#2958)

* chore(#2929): correct the corpus case count in the changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-07-31 23:03:13 -04:00
Tom Boucher
f6257f3745 ci(#2952): shard the full test lane instead of widening its cap (#2960)
* ci(#2952): budget CI job timeouts by headroom over measured cost

`origin/next` was red. The only failing check was `Required tests`, and its
sole cause was `test (ubuntu-latest, 24)` reported `cancelled` — GitHub's
conclusion for a job that exceeds its `timeout-minutes`, not a button press.

That job's `ubuntu-latest / 24` entry is the only `scope: full` matrix entry:
it runs the whole unit suite under c8 coverage, then the scripts/ coverage
floor, integration, security, install and slow, serially on one runner. On
next@5a0a9f097 it ran 15m16s against `timeout-minutes: 15` and was axed 23s
into `npm run test:slow`. Projected to a completed slow step (27s on the last
green run) the lane costs ~15m20s.

Confirmed hypothesis: the budget, not the suite. The lane had been riding the
ceiling all day — 12m03s, 11m34s, 11m50s, 14m25s, 14m51s — and crossed on
three of the last four full-lane runs (d2d2f7c08, 07603df8f, 5a0a9f097). There
is no pathological test: the unit run is cost-first bin-packed into 12 chunks,
chunk 1 is gated by run-tests-harness.test.cjs at 169s (expensive by design —
it spawns real harness subprocesses, one of which exercises the per-chunk
timeout), and the remaining chunks are 32-94s. 769s is the honest cost of 689
files under c8. Re-running could not have helped; the work exceeded the budget.

Review of the first cut surfaced the same defect one runner away: `full test
(windows-latest, 22, shard 3/3)` reached 18m59s against its own 20-minute cap
on 05b170e44 (94%) and 18m14s on 81eeb8a53 (91%). That lane has already blown
its cap twice (#1051, #1212). Fixed here rather than deferred.

The first cut also asserted `test >= test-full`, which is unsound — those two
budgets are dominated by different platforms, so their ordering carries no
meaning. Replaced with the invariant that actually generalises: every lane is
held to a headroom FACTOR over its own measured cost. `test` 15 -> 25 (1.5x of
16m), `test-full` 20 -> 30 (1.5x of 19m), `test-inert` unchanged at 15.

tests/ci-test-job-timeout-budget.test.cjs locks that rule. No unit test can
prove a lane still FITS its budget — only a real run measures that — but a
budget can no longer be lowered back beneath what its lane is known to need,
and a lane that gets slower must be re-measured rather than excused.

Separately: the earlier `failure` at 05b170e44 was an unrelated, already-fixed
CONTEXT-INDEX.json drift (07603df8f re-synced it; lint-tests is green at HEAD).
07603df8f's own run hit this same timeout, which is why it never reported green.

Refs #869, #1051, #1212

* ci(#2952): shard the full test lane instead of widening its cap

The `scope: full` lane was the only unsharded lane in this file. It ran the
entire unit suite under c8 on one runner, grew past a 15-minute cap, and
reddened `next`. Raising the cap bought room; it did not change the shape, and
the same lane would have walked back into the ceiling. Shard it, the way #1212
answered this for the Windows lane.

Balance comes from measurement, not file counts. scripts/run-tests.cjs already
partitions by measured per-file duration using LPT (#2472); the table it reads
was 10 days stale — 638 of 695 files timed, 64 missing, including the whole
context-predicates group. Regenerated from a verified matrix run: 700 files, 0
missing. On that table the 685-file unit suite splits 19.37m / 19.37m / 19.37m
— 0.0% spread — and the split is a total, disjoint cover with 0 files dropped.
Completeness, disjointness, balance and determinism of the partition itself are
already pinned against selectShard in run-tests-harness.test.cjs, including a
fast-check property, so this change does not restate them.

Sharding a COVERAGE run is the part that needs care. A per-shard percentage is
meaningless — shard 2 never executes shard 1's files, so those read 0% — and
leaving the gate on the shards would have quietly measured a third of the tree.
Each shard now renders no report and only leaves raw V8 dumps; a new
`coverage-gate` job merges all three into one coverage/tmp and runs the gate
there. c8's default temp directory is where the download lands, so the ≥70%
lines / ≥60% branches gate and the ≥55% scripts floor run unmodified against
merged data.

Both surfaces call the same npm scripts rather than inlining c8 into YAML, so
the include/exclude globs and both thresholds stay defined once in package.json.
The workflow holding its own copy is the divergence this repo has a rule
against, and the new test cross-checks package.json so an inline reintroduction
fails rather than drifts.

tests/ci-full-lane-sharding.test.cjs covers the two ways this stays GREEN while
being wrong: an incomplete shard set (declare 1/3 and 2/3, never 3/3, and a
third of the suite silently stops running) and a coverage gate that stops being
required. required-tests now depends on coverage-gate and fails on it, while
still tolerating `skipped` so docs-only PRs are not blocked.

`timeout-minutes: 25` on the lane is deliberately left alone. The budget test
requires a real measurement before a lane's declared cost changes, and the
sharded cost is not measured until this PR's own CI run.

Refs #1212, #2472

* ci(#2952): tighten the sharded lane's budget to its measured cost

The sharding commit deliberately left `timeout-minutes: 25` alone, because
tests/ci-test-job-timeout-budget.test.cjs requires a real measurement before a
lane's declared cost changes and the sharded cost did not exist yet.

It exists now. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s, and
coverage-gate 1m20s. Shard 1 is the long pole because the unsharded aux suites
ride along on it, which is deliberate — they total ~1m35s and sharding them
would cost more than it saves.

So the lane's budget is 15 against a slowest measured shard of 8 minutes
(~1.9x), and coverage-gate joins LANE_COSTS at 2 minutes. 15 is the same number
the lane blew before sharding; the work behind it is now a third the size.

Merged coverage was checked against the pre-shard single-runner baseline rather
than assumed from a green check: 94.36 stmts / 96.3 funcs / 94.36 lines
identical, branches 84.22 vs 84.21 — one branch across two different trees,
noise rather than a regression.

---------

Co-authored-by: sim <sim@local>
2026-07-31 22:16:26 -04:00
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
JusticeWay
7b204ad2ac enhance(#2530): extend UAT checkpoint frame language pack (9 more languages) (#2564)
* feat: extend UAT checkpoint frame language pack (9 more languages)

response_language is a free-form config value, but CHECKPOINT_FRAMES only
covered 9 languages — any other configured language silently fell back to
the English frame. Add Dutch, Polish, Russian, Ukrainian, Turkish, Hindi,
Arabic, Vietnamese, and Indonesian frames plus their aliases, with a
regression test asserting each resolves instead of falling back.

Follow-up to #2402 (PR #2457).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: add changeset for #2527

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2530): list UAT checkpoint frame languages in CONFIGURATION.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2530): point changeset fragment at PR #2557

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2530): address Unicode language-pack review

* fix: address checkpoint language review

* fix: count spacing combining marks in checkpoint width

* test: verify checkpoint aliases structurally

* fix: isolate RTL checkpoint frames

* fix: isolate RTL checkpoint frames correctly

* test(#2530): assert checkpoint aliases neither collide nor go unreachable

Review Minor #1. A duplicate alias key was invisible to the existing
catalog tests: the runtime object is well-formed after JS collapses the
literal, the self-alias assertion still holds, and the losing language
just stops resolving. tsc catches the byte-equal case (TS1117), but not
the two that survive compilation — an alias whose NFC-lowercase form
already belongs to another language, and an alias not in lookup form at
all, which resolveCheckpointFrame() can never produce.

The check reads the source literal rather than the object, since the
object no longer records what was written. Both assertions are
independently load-bearing: an NFD twin of an existing alias trips the
collision check, an uppercase alias trips the unreachability check.

Review Minor #2: changeset retyped Changed -> Added. Nine wholly new
supported response_language values are an addition under Keep a
Changelog, not a modification of existing behavior.

* test(#2530): check alias collisions on the catalog, not its source

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:01:50 -04:00
Tom Boucher
f092c6da85 fix(#2649): diagnose-issues + execute-plan run worktree.base-check before worktree dispatch (#2955)
* test(#2649): failing-first — diagnose-issues + execute-plan must run base-check before worktree dispatch

* fix(#2649): diagnose-issues + execute-plan run worktree.base-check before dispatch

diagnose-issues.md spawn_agents and execute-plan.md Pattern A spawned
worktree-isolated subagents (gsd-debugger / gsd-executor) without the
pre-dispatch worktree.base-check gate that execute-phase (#683/#1369) and
quick (#1941) already run. Claude Code's isolation="worktree" forks from
origin/HEAD, not live local HEAD; without the gate, the documented GSD steady
state (commit every step locally, push only on request) hits the verify-only
worktree_branch_check guard's exit-42 halt mid-investigation with no auto-degrade.

Mirror the quick.md #1941 pattern: before dispatch, run
`gsd_run query worktree.base-check --pick shouldDegrade`; if true, print its
message + a #2649 warning to stderr and set USE_WORKTREES=false (sequential
main-tree dispatch). The verify-only guard stays as a backstop in both cases.

Per the triage and #2649 acceptance criterion 5, execute-plan.md's Pattern A
(identified as a second site with the identical gap) is fixed in the SAME change
— same bug class, same one-line gate, two workflow files — rather than filed as
a separate follow-up.

* fix(#2649): ack the diagnose-issues + execute-plan growth (per-PR fragment)

The two workflow files grew vs next (diagnose-issues.md +1381, execute-plan.md
+905) adding the #2649 base-check gate. emitted-attribution requires an ack;
this is a per-PR fragment under tests/emitted-drift-acks/ (#2914 mechanism,
replacing the legacy shared emitted-drift-ack.json).

* test(#2649): tighten base-check ordering assertion + guard backstop survival

Address code-review minors:
- the ordering assertion was a loose disjunction that passed even if the
  base-check moved AFTER the dispatch; tighten to assert base-check < Agent()
  (the real invariant).
- add a test that the verify-only <worktree_branch_check> backstop remains
  embedded in the Agent() prompt (acceptance criterion 4 — the base-check is a
  pre-dispatch degrade, the guard is a post-fork fail-closed backstop; both
  layers must survive).

* changeset(#2649): diagnose-issues + execute-plan auto-degrade on stale worktree base

* changeset(#2649): backfill PR number 2955

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 19:15:04 -04:00
Tom Boucher
388837219d fix(#2648): phase.complete refuses when non-retired plans lack summaries (fail-closed coverage gate) (#2953)
* test(#2648): failing-first — phase complete must refuse when plans lack summaries

Adds findUnsummarizedPlans to core-utils (mirrors countMatchedSummaries but
returns the unmatched plan files) and a PHASE_PLAN_COVERAGE_INCOMPLETE error
reason, plus a 3-case regression block in phase.test.cjs. The gate itself is
NOT yet wired into cmdPhaseComplete (reverted for the RED run), so the
'blocks completion' and 'superseded does not block' cases must FAIL (the
pre-fix code completes silently).

* fix(#2648): phase.complete refuses when non-retired plans lack summaries

cmdPhaseComplete gated only on a single *-VERIFICATION.md status, so a phase
could close complete while an arbitrary number of its plans had no completion
record (confirmed incident: 6/30 plans unexecuted incl. the phase's entire
final UI scope, every signal green). Add a fail-closed plan-coverage gate that
refuses completion when any plan lacks a matching *-SUMMARY.md, naming the
missing plans, UNLESS the plan is retired via machine-readable status:
superseded frontmatter (#2349) — closing the Goodhart hole (delete a SUMMARY to
raise the %) without regressing the lock/recovery pattern.

Uses scanPhasePlans (superseded-AWARE) + new findUnsummarizedPlans helper so the
gate, the count, and the named list can never disagree. Evaluated before the
verification-gate transaction so a refusal fails fast without mutating
ROADMAP/STATE. milestone.complete's parallel gap is explicitly out of scope
(separate seam, separate PR).

* fix(#2648): test fixtures — give #1752 phase plans summaries + STATE.md in coverage fixture

The plan-coverage gate (#2648) correctly blocks phase completion when a plan
lacks a SUMMARY. Two test fixtures needed updating to reflect the new contract:
- #1752 (total_phases-decrement cascade): its 8 phase dirs each had a PLAN.md
  with no SUMMARY. The test's concern is the total_phases cascade, not plan
  coverage, so add a matching SUMMARY to each to keep the phase fully-covered
  and isolate the #1752 behavior.
- the #2648 coverage-gate fixture: write STATE.md (createTempProject scaffolds
  .planning/phases but not STATE.md) so the 'ROADMAP/STATE unchanged on refusal'
  assertions have a file to read.

* fix(#2648): security — fail closed on unreadable plan dir + sanitize msg + surface superseded

Address the security-review blocker (B1) and hardening (M1/m2):
- B1 (blocker): the gate failed OPEN when scanPhasePlans could not read the
  phase dir (it swallows readdirSync errors → empty plan set → gate sees zero
  unsummarized plans → passes). A coverage gate that passes when it cannot read
  the plans re-opens the #2648 hole under any I/O failure. Now readdirSync the
  dir explicitly and fail closed (PHASE_PLAN_COVERAGE_INCOMPLETE) on a throw;
  a readable empty dir still passes (legitimately complete empty phase).
- m2: sanitize plan filenames (strip C0 controls / DEL) before interpolating
  into the error message — they come raw from readdirSync and could spoof the
  terminal in plain-error mode.
- M1: surface the count of plans excluded as status: superseded so a reviewer
  can audit which work was declared retired (the marker is a committable,
  review-time-trusted bypass; keep it visible).
- Add a 4th regression case: unreadable plan dir (ENOTDIR via a file, not chmod
  0o000 which root bypasses) must fail closed.

* test(#2648): drop unreachable B1 case — no root-safe unreadable-dir repro

The B1 fail-closed-on-unreadable-dir defense stays in src/phase.cts (cheap +
correct), but it cannot be unit-tested cross-platform: any condition that makes
the phase dir unreadable to the gate's readdirSync ALSO fails findPhaseInternal
upstream ('Phase N not found') before the gate runs, and chmod 0o000 is
forbidden (root bypasses it in root CI). Document the gap in the test file;
remove the case that asserted a reason the upstream error pre-empts.

* style(#2648): drop unnecessary type assertions flagged by lint:ci

scanPhasePlans returns typed string[] arrays, so the `as string[]` casts on
coverageScan.planFiles/summaryFiles were redundant (@typescript-eslint/
no-unnecessary-type-assertion). Compute supersededCount from typed lengths;
only the phaseInfo['plans'] cast remains (it is genuinely unknown).

* changeset(#2648): phase.complete refuses when plans lack summaries

* changeset(#2648): backfill PR number 2953

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 18:11:40 -04:00
Tom Boucher
5a0a9f0972 fix(#2944): remove the catastrophic-backtracking regex from the ADR-1671 example (#2950)
* fix(#2944): remove the catastrophic-backtracking regex from the example

The non-shipping Option-E reference example carried its own copy of the
predicate-id regex, which nested a dot-containing character class inside a
dot-prefixed repeat. A run of N consecutive dots therefore had exponentially
many partitions. Measured on next before this change: 30 dots 54ms, 35 66ms,
40 807ms — so roughly 55-60 dots hangs for hours.

Not exploitable where it sits: the example is outside tsconfig.build.json,
outside the npm package files list, outside the installer and outside tests,
so no build step or CI job parses anything with it. Fixed because the entire
point of a reference example is that people copy it forward, and ADR-1671
presents this one as the pattern for the platform.

Ports the linear per-segment validation that #2928 gave the production module,
so the two copies agree: both parse the real CONTEXT.md to 415 predicates
across 20 classes with 0 duplicates. Doubled-dot ids are now rejected here
too, matching production, and the grammar comment records it.

Also refreshes the example's committed index, which #2928 made stale when it
removed the duplicate predicate from CONTEXT.md.

Closes #2944

* test(#2944): guard predicate-index sync and example/production parity

Two regression tests for the two defects in this PR.

Index sync: asserts the committed docs/CONTEXT-INDEX.json equals a fresh parse
of CONTEXT.md, naming any diverging predicate ids. The merge race that reddened
next was invisible to both PRs involved and only surfaced on the next PR to run
lint:ci; this puts the same check inside the suite, which runs on every PR, and
a mutation test proves the assertion is not vacuous.

Example/production parity: asserts both copies of the parser report the same
count, classes and duplicates for the real CONTEXT.md, and agree verdict-for-
verdict over a table of id shapes. The divergence WAS the bug — production went
linear-time while the example kept the backtracking regex, with nothing
asserting they agreed. Also pins the example rejecting a 60-dot id, with the
clean rejection as the binding assertion and wall-clock only as a smoke check.

Notes a real tension rather than hiding it: ADR-1671 says the example sits
outside tests/, and this imports it. The ADR's intent is that the example is
not compiled, packaged or installed — not that it may silently rot. A parity
guard does not ship it. The file states this so a reviewer can object.

* fix(#2944): address both isolated review passes

Two independent reviewers (correctness and security axes, neither the author).
Security found nothing — it measured linearity to 100k chars across dots,
hyphens, underscores and mixed classes, and showed prototype pollution is
structurally unreachable because the first-segment pattern forbids
lowercase and underscore-leading ids. The correctness pass found three
blockers, all real.

Blocker: the parity test violated ADR-1671 verbatim. The ADR lists FOUR
exclusions for the reference example, the fourth being the CI test suite, and
the test imported it from tests/ while its own justification comment cited only
three -- constructing a rationale around the exclusion it broke. Moved to
scripts/lint-example-parser-parity.cjs wired into lint:ci; a lint script is not
the test suite, so the exclusion stands. The test file keeps only the
docs/CONTEXT-INDEX.json sync check.

Blocker: the mutation test leaked its temp dir. Its callback took no `t`, so a
failing assertion skipped the bare cleanup call. Now registered via t.after(),
matching the convention adr-index-gate.test.cjs documents.

Blocker: the example's own committed index carries the identical merge-race
staleness this PR fixes for the production one, and nothing guarded it.
Deliberately NOT fixed by wiring the example's --check into CI: that artifact
bakes line numbers, so it re-drifts on any unrelated CONTEXT.md line shift --
exactly ADR-1671 open question 4 -- and would make CI routinely red. The new
lint asserts the line-INDEPENDENT facts instead: count, class map, duplicate
set, and every (id, value) pair. Proven non-vacuous both ways: mutating a value
fails and names the id, mutating only a line number passes.

Major: a real divergence the parity claim would have missed. Production rejects
values containing an embedded CR, LF, U+2028 or U+2029; the example did not, so
a value with an embedded lone CR was rejected by one copy and accepted by the
other. Ported, and now covered by the parity table.

Also, found while verifying rather than reported: malformed diagnostics covered
only empty values. A doubled-dot id, a space in an id, and a lowercase-leading
id were all dropped silently. That contradicts the module's own intent -- a
typo should be diagnosable, and a space in an id is a likely one -- and
predicates are contractually cited, so a silently vanished predicate is the
failure mode that matters. Each rejection class now carries a named reason in
both copies, while ordinary inline code still yields none.

Trues up counts my own change staled: the example README and ADR-1671's
prototype figures said 416 and 393/18 against a real 415/20/0.

Closes #2944

* chore(#2944): backfill changeset PR number 2950

---------

Co-authored-by: sim <sim@local>
2026-07-31 15:44:19 -04:00
Tom Boucher
07603df8f2 fix(#2647): code-fixer worktree under .claude/worktrees/, not a hardcoded /tmp path (#2942)
* test(#2647): failing-first — fixer worktree path must be repo-relative not /tmp

* fix(#2647): place code-fixer worktree under .claude/worktrees/, not /tmp

The gsd-code-fixer agent hand-rolled its worktree at a hardcoded
`/tmp/sv-${padded_phase}-reviewfix-XXXXXX` mktemp path. On Windows/Git Bash
that landed OUTSIDE the project tree — outside the agent session's permission
allowlist, so every Read inside the worktree prompted (~25/run) — and mktemp's
MAX_PATH-avoidance substitute produced an un-removable `C:/mvwtNN` path.

Place the worktree repo-relative under `.claude/worktrees/` (the same dir the
harness-managed executor worktrees use: gitignored via `.claude/`, inside the
session's permission scope), with a $$-PID + epoch suffix for concurrency
uniqueness (replacing mktemp's XXXXXX). $main_repo is resolved the same way
the cleanup tail already resolves it.

Three sites updated: setup_worktree bash, concrete-steps prose, critical_rules.
The #2990 `-b "$reviewfix_branch"` invariant is preserved (the folded test
asserts it). Failing-first regression added to the #2990 suite in
tests/agent-frontmatter.test.cjs.

* test(#2647): update #2686 path assertion to expect .claude/worktrees/, not /tmp

The #2686 regression test encoded the worktree location as a hardcoded
`/tmp/sv-` path (matching sibling GSD agents at the time). #2647 showed that
breaks Windows/Git Bash (worktree outside the project tree → permission prompts;
mktemp MAX_PATH substitute un-removable). Update the #2686 path assertion to
require the repo-relative `.claude/worktrees/` location and forbid `/tmp/sv-`.
The #2686 isolation + cleanup assertions are unchanged.

* fix(#2647): word-boundary wt= parse + ack the fixer growth vs next

Two follow-ups to the #2647 GREEN run:
- parseWtAssignments matched `prior_wt=` (no word boundary), polluting the
  set and tripping the repo-relative + concurrency-unique assertions. Anchor
  on (?:^|\s)wt= so only the real worktree-path assignment is captured.
- emitted-attribution: gsd-code-fixer.md grew 1875 bytes vs origin/next. Update
  the emitted-drift-ack entry to attribute the #2647 worktree-path change
  (supersedes the prior #2825 attribution, whose growth is already in next).

* fix(#2647): address review — validate padded_phase at the sink + tighten test

Code-review + security-review both APPROVED with one actionable minor:
padded_phase is interpolated into a worktree PATH and a git BRANCH NAME, but
was only validated by the orchestrator (code-review-fix.md), not at the agent
sink. The agent prompt is a literal bash contract any caller can spawn, so add
a `[[ =~ ^[0-9]+(\.[0-9]+)?$ ]]` self-defense check rejecting traversal/shell
metachars (defense-in-depth; not a present vuln — the only caller validates).

Also tighten the concurrency-uniqueness test to require BOTH $$ AND $(date +%s)
(either-alone was too lax per review). Update the emitted-drift-ack reason to
cover the added validation growth.

* changeset(#2647): code-fixer worktree under .claude/worktrees not /tmp

* changeset(#2647): backfill PR number 2942

* chore(#2938): regenerate stale docs/CONTEXT-INDEX.json on next

#2938 (#2928) updated the CONTEXT.md RULESET prose for the new per-PR
emitted-drift-ack fragment mechanism (#2914) but shipped a CONTEXT-INDEX.json
generated from the OLD prose. lint:generated-sync fails on every PR that
rebases onto next after #2938 (the regen produces a 3-line diff bringing three
RULESET entries — AGENT_SIZE_BUDGET, EMITTED_ATTRIBUTION, WORKFLOW_SIZE_BUDGET
— in sync with the prose already on next). Mechanical regen via
`node scripts/gen-context-index.cjs --write`; idempotent; surfaced by the
#2647 rebase. No behavioral change.

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 14:45:51 -04:00
Tom Boucher
05b170e448 chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam

Productionizes the ADR-1671 Option-E reference example as a real module:
src/context-predicates.cts (parser + selector + index builder) compiled to
gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's
--check/--write drift-guard idiom and wired into lint:generated-sync.

Parser behavior is deliberately prototype-equivalent in this commit so the
next commit's regression matrix binds to the real defects rather than to a
missing module.

Two locked design deviations from the prototype:
- duplicates carry a count, not line numbers
- the committed index carries no line field at all, resolving ADR-1671 open
  question 4: an artifact without line numbers cannot drift on a line shift,
  so promoting --check to a CI gate does not make it routinely red

Also reconciles the one remaining duplicate predicate ID
(RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording
is removed) so the gate can land fail-closed on duplicates.

Refs #1671

* test(#2928): failing-first matrix for the predicate fact-store

Adds the regression matrix from the phase test plan: parser declaration
forms, fence and comment regions, ID/value grammar boundaries at
limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard
CLI, the selector query surface, and four document-shaped fast-check
properties.

Seven rows are RED for behavioral reasons against the ported parser:
indented-bare, star-list, plus-list and numbered-list declaration forms are
dropped; a tilde fence and a four-backtick fence containing a shorter fence
are not skipped; and a multi-line HTML comment is parsed as live. Eleven
selector rows are RED because the query surface is not wired yet.

Negative fixtures come from real repo documents that predate the grammar
(CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the
fixture-provenance rule, and the property generators are document-shaped
rather than seeded from our own serializer.

Refs #1671

* fix(#2928): consume the shared fence scanner, relocate the index, wire the selector

Drives the failing-first matrix green.

Parser: replaces the ported naive triple-backtick toggle with the shared
markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord
gain an export keyword — the only change to that module, which has 71
upstream dependents — because it already returns line-indexed spans, which
is exactly what a line-reporting parser needs. It also already documents
itself as the second copy of the fence state machine pending consolidation;
adding a third copy here would have been the generative-fix divergence this
repo warns about. A parity suite now pins predicate fence-skipping against
that scanner across eight fence shapes. HTML-comment skipping stays local
because the sectionizer has no comment scanner. Declaration forms widen to
indented-bare, star, plus and numbered list items.

Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The
remote matrix run caught the original choice — a committed .cjs there ships
~120KB of CONTEXT.md prose into a runtime module, and two content guards
fired truthfully on it (a leaked .claude install path, and four hardcoded
package-name literals). Neither guard was allowlisted; the artifact moved
instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to
require it — it is a drift-detection artifact, so the selector parses
CONTEXT.md live and is always current.

Generator: adds a frozen REASON enum and --check --json so the gate's
outcome is asserted structurally instead of by matching prose, and
--context-path/--index-path so tests drive the real CLI against a temp tree
with no filesystem monkeypatching.

Selector: gsd_run query context-predicates with --class/--prefix/--contains,
structured output carrying a matched count, own-property guards, and no
project-root resolution. Registering it exposed that the query dispatch
table and the usage string had drifted: a new parity test found 20 routed
commands missing from the usage list, all added here rather than deferred.

Refs #1671

* test(#2928): lock the newly-public scanFencedBlocks contract

Exporting scanFencedBlocks made it public API for the first time, so it
needs its own contract test independent of the consumer that motivated the
export. Memtrace's co-change analysis flagged the gap: this suite changes
together with markdown-sectionizer.cts 8 times in 90 days and was absent
from the diff.

Covers the documented rules: 0-based indices, -1 for an unterminated fence,
the same-char/>=length/no-trailing-text closer rule, a shorter fence inside
a longer one staying content, CommonMark 4.5 backtick-in-info-string, and
<=3-space indent tolerance.

Refs #1671

* fix(#2928): address both isolated review passes

Two independent reviewers (correctness axis and security axis, neither the
author) found seven findings. All are fixed here with regression tests; none
deferred.

BLOCKER — comment-blind fence scanning caused silent, permanent predicate
loss. The HTML-comment scan and the fence scan ran as two independent passes,
and the fence scanner is comment-blind, so a fence delimiter inside an HTML
comment with no later close read as an unterminated fence and skipped every
remaining line to EOF. Worse, the drift-guard could not catch it: it diffs
against a baseline produced by the same corrupted parse. The two constructs
now interleave in a single pass so each suppresses the other's boundary
detection while active, covered in both directions. The parity suite still
binds this scanner to markdown-sectionizer's for comment-free documents, so
the two cannot diverge unnoticed.

BLOCKER — the selector was not consumed anywhere, leaving the phase's
acceptance criterion unmet. Now wired into the pre-work predicate-citation
step in contributor-standards, which is the repo's actual brief-assembly
path; no code-level brief assembler exists to wire into.

MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id
regex nested a dot-containing character class inside a dot-prefixed repeat,
so N consecutive dots had exponentially many partitions: 40 dots took 565ms
and growth was exponential. CI runs this parser over a pull request's own
CONTEXT.md, so any contributor could have hung a shared runner with one
line. Replaced with linear per-segment validation. Doubled-dot ids are now
rejected; the real document contains none.

MAJOR — the duplicate-id gate had only ever been proven on synthetic
fixtures. A test now re-inserts the exact line this branch removed and
asserts the real generator names it.

MAJOR — --check together with --write silently let write win, turning the
gate into a writer; a missing path value resolved to the cwd and leaked an
EISDIR stack trace. Both are now clean usage errors.

MINOR — the hoisted skip-list was exported as a live mutable Set; replaced
with a read-only predicate. MINOR — flag-shaped selector values were
unmatchable; the inline --flag=value form now provides the escape hatch.

Refs #1671

* chore(#2928): backfill changeset PR number 2938

---------

Co-authored-by: sim <sim@local>
2026-07-31 13:17:01 -04:00
Tom Boucher
c043f2946c fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next

tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900.
Every entry is scoped to the diff that introduced it (#2789), so once merged
to next it is at the base by definition -- spent and inert. Its presence is
still load-bearing though: each PR rewrites the paths map wholesale, making a
persistent base copy a shared cell. Five of six conflicting PRs in the open
queue collided on this file and nothing else.

Deletes the stale document and adds a push-to-next guard asserting it stays
absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check
against the base is the #2768 shape #2789 exists to end.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2914): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2914): per-PR ack fragments instead of one shared mutable file

The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json
whose paths map every PR rewrote wholesale. That is a shared mutable cell: any
two PRs needing an ack edit the same lines and conflict. Five of six conflicting
PRs in the open queue collided on this file and nothing else.

Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same
shape .changeset/ already uses to solve this exact problem. Two PRs pick
different filenames, so they cannot collide, and fragments lingering on next
are harmless rather than toxic.

The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An
earlier delete-only attempt failed verification twice: the ratchet lost the
spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no
ack. Relocating preserves every acknowledgment.

The legacy single file is still READ (unioned with the fragments) because five
open PRs carry it; dropping support would break all of them. A duplicate path
key across sources is a hard error, never last-wins.

The push-to-next guard is retargeted accordingly: it now asserts only that the
legacy SHARED file never reappears on next. Fragments may persist harmlessly.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:15:29 -04:00
Tom Boucher
d2d2f7c088 fix(#2848): non-Latin titles no longer produce empty slugs (Cyrillic transliteration) (#2934)
* test(#2848): add failing-first regression for non-Latin slug transliteration

generateSlugInternal and slugify both strip non-ASCII chars with no
transliteration step, so an all-Cyrillic title reduces to an empty slug.
12-row matrix: Cyrillic regression (both impls), Latin negative control,
multi-letter mappings, soft/hard sign drops, Ukrainian extras, null
contract, CJK unaffected, mixed scripts, truncation parity, slugify's
distinct no-truncate contract.

* fix(#2848): transliterate Cyrillic titles to ASCII before slug strip

Both generateSlugInternal (src/core-utils.cts) and slugify
(src/gsd2-import.cts) stripped non-ASCII with no transliteration, so an
all-Cyrillic title reduced to an empty slug, producing unnamed phase
directories (01-) and empty milestone_slug init JSON.

Add a shared transliterateForSlug primitive (core-utils) covering Russian
+ the reported Ukrainian/Belarusian extras (і ї є ґ ў), with multi-letter
mappings (ж→zh ч→ch ш→sh щ→sch ю→yu я→ya) and dropped soft/hard signs
(ъ ь). It runs BEFORE the existing ASCII filter, so Latin-script text
hits zero map entries and is byte-for-byte unchanged (negative control).
slugify consumes the shared primitive, preserving its distinct single
hyphen-strip + no-truncation contract. CJK/unmapped scripts keep the
existing strip-to-ASCII behavior.

Also corrects two test assertions to match the chosen й→y mapping and the
б→b (not bie) transliteration.

* changeset(#2848): Fixed — non-Latin slug transliteration

* changeset(#2848): backfill PR number 2934

---------

Co-authored-by: sim <sim@local>
2026-07-31 10:56:11 -04:00
Rezolv
76b7d73039 fix(#2733): route gate-passed spec-phase paths into the probe steps (#2779)
* fix(#2733): route gate-passed spec-phase paths into the probe steps

All four gate-passed transitions in spec-phase.md said "Jump to Step 6",
textually bypassing the mandatory Step 5.5 edge-completeness and Step 5.6
prohibition-completeness probes. Steps 5.5/5.6 were spliced between Step 5
and Step 6 by two later feature commits and the pre-existing jumps were
never re-pointed, so no jump instruction in the file reached Step 5.5 at
all and both probes were unreachable dead prose.

Re-point the four gate-passed jumps (lines 129, 162, 168, 170) to Step 5.5.
Control then flows 5.5 -> 5.6 -> 6 as the probes' own preconditions
prescribe. The max-rounds "write anyway" bypasses and the probes' own
"proceed to Step 6" exits are deliberately unchanged.

Add tests/spec-phase-probe-reachability.test.cjs, which derives the
mandatory probe steps from the file's own headings rather than hardcoding
5.5/5.6, so a future spliced-in probe step is covered without editing the
test. It also locks the two coupled constraints: the max-rounds bypass must
not be redirected into a probe, and each probe must keep its own onward exit.

The existing probe contract tests are untouched and still pass; both scope
from the "## Step 5.5"/"## Step 5.6" heading onward and were structurally
incapable of observing the upstream jump text.

* chore(changeset): Fixed fragment for #2779 (spec-phase probe reachability)

* fix(#2733): route Step 5.5's own soft gate into Step 5.6

Round-1 review blocker. The four upstream gate-passed jumps were re-pointed to
Step 5.5, but Step 5.5's own terminal soft gate at :305 still read "proceed to
Step 6" - so the COMMON path (all applicable edges resolved) skipped the
prohibition-completeness probe outright. Same defect class as the four this PR
already fixed, on the success path of the very step being fixed: the SPEC shipped
with an empty Prohibitions section instead of an empty Edge Coverage one.

Its sibling at :393 is byte-identical yet correct, because Step 6 genuinely
follows Step 5.6. Position, not phrasing, is the discriminator.

The guard could not see it: the transition matcher keyed only on the literal
"Jump to Step", and :305 says "proceed to Step". Widened it to a verb alternation
(jump/proceed/continue/go/return/skip + "to Step N", case-insensitive) and
renamed it TRANSITION_RE to match what it now models. This makes the file's own
docstring promise - that a future spliced-in probe is covered without editing the
test - true for a step whose exit is worded differently. Verified no false
positives: the two pre-existing "continue to Step 3/4" transitions are upstream
of both probes but target pre-probe steps, and the max-rounds bypass block
contains no step transitions at all.

Fail-first verified before fixing :305 - with the widened matcher against the
unfixed workflow the guard fails naming exactly "spec-phase.md:305 jumps to Step
6, skipping mandatory Step 5.6", 4 pass / 1 fail; after the fix, 5/5. The two
sibling probe contract tests stay 16/16.

Also from review:

- STEP_HEADING_RE gains an explicit \r? before $. Without it, on a CRLF checkout
  `.` stops before the \r and the unanchored $ fails to match, yielding ZERO
  steps and vacuously passing every assertion in the file. Not live today
  (.gitattributes forces eol=lf) but this repo has a recurring CRLF-regex bug
  class, so the guard no longer leans on it.
- allow-test-rule category corrected to source-text-is-the-product; the previous
  runtime-contract-is-the-product is not one of the six recognized categories
  (CONTRIBUTING.md:609-619).
- changeset body given the documented bold-lead-in form.
- emitted-drift ack reason updated: +8 -> +10 bytes across five transitions
  (31987 -> 31997), DEFAULT tier, cap 40960.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 10:20:17 -04:00
Tom Boucher
d49a7d0c4d fix(#2853): roadmap.update-plan-progress preserves hand-written annotations (#2916)
* test(#2853): add failing-first regression for plan-progress annotation preservation

The count-bump regex's trailing [^\n]+ swallowed the whole Plans line and
the replacement wrote back only the regenerated count, deleting any
hand-written annotation after it. 8-row matrix covers bold/plain forms,
bare template form, executed path, CRLF, and idempotency.

* fix(#2853): preserve hand-written annotations in roadmap plan-progress bump

The count-bump regex's trailing [^\n]+ swallowed the entire Plans line and
the replacement wrote back only the regenerated count, deleting any
hand-written prose after the count (e.g. a gap-closure annotation).

The verb owns the count token only. Capture the existing count token ($2)
and the trailing line text ($3), and rebuild the line as
<label><new count><surviving text>. Trailing text is preserved ONLY when a
real count token preceded it, so the fresh-template bracketed placeholder
(`[Number of plans…]`) is still replaced cleanly rather than glued after
the count (pre-#2853 behaviour on the template path preserved). CRLF \r is
preserved via [^\r\n].

Widens replaceInCurrentMilestone to accept a replacement callback (needed to
branch on whether the count group matched). The bare Plans: checklist header
is still skipped — the lazy match lands on the summary line first and a
count-less bare header yields no count to anchor preservation to.

* changeset(#2853): backfill PR number 2916

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: sim <sim@local>
2026-07-31 09:54:55 -04:00
Tom Boucher
90771ddf02 enh(#2904): add a reviewer entry type so third-party reviewer lanes are discoverable (#2912)
* feat(#2904): add a `reviewer` entry type so third-party reviewer lanes are discoverable

ADR-2782 made a reviewer lane installable by a third party, but neither
discoverability catalog could hold one. The Community Capability Registry
requires a non-empty `loopExtensionPoints` and forbids a lane from declaring
any hook kind, so a `role: "reviewer"` entry is unsatisfiable by construction;
the EoS Registry is for ADR-1239 host integrations, which a lane is not.

Adds a third catalog — `docs/registries/reviewers.json` →
`docs/registries/reviewer-registry.md` — whose `interactions` describes the
lane: slug, flags, transport, evidenceClass, reviewsSection, requiresBinaries,
configKeys, runtimeCompat.

The lane vocabulary is a hand-written mirror of `capability-validator.cjs`
(the same pattern as `AXES` mirroring `HOST_INTEGRATION_AXES`), with parity
enforced by tests/registry-reviewer-parity.test.cjs. `slug` deliberately uses
the runtime `LANE_SLUG_RE` grammar rather than the registry's kebab-only `id`
rule, so real lanes (`lm_studio`, `4o-mini`) are not rejected.

Two binary type branches became three-way Map dispatch. Both now fail loudly
on an unrecognized type instead of silently treating it as a capability —
`renderMarkdown` in particular writes a committed catalog file, so a silent
wrong-title render was the worst failure mode available.

Also fixed while here: `gen-registry.cjs` parsed source JSON with no error
handling, so a malformed or non-array `capabilities.json` surfaced as a raw
SyntaxError/TypeError instead of an actionable CLI error.

Closes #2904

* fix(#2904): bound and sanitize untrusted registry `interactions` strings

Review findings from the pre-PR passes.

Security (isolated pass): `interactions` string fields reached the generated,
committed Markdown catalog with no control-character check and no length
bound. A `reviewsSection` carrying ESC and a `requiresBinaries` element
carrying NUL plus 5000 characters validated clean and landed verbatim in the
rendered page — `mdInline` escapes Markdown metacharacters and collapses CRLF,
but nothing else. The identical gap already existed on the capability type's
`configKeys`/`requires`/`runtimeCompat`/`produces`/`consumes`, so it is fixed
there too rather than inherited into a third type.

`hasDisallowedControlChar` is lifted to module scope so exactly one
implementation exists, and a shared `validateStringArrayField` enforces
control-character rejection, a 200-character element cap and a 50-element
array cap for both types.

Correctness (standards pass): `renderMarkdown`'s per-entry summary builder was
still an if/else-if chain whose final `else` was the capability branch — the
one per-type dispatch point this change had not converted, and the same silent
fallthrough it removes elsewhere. It now lives in `RENDER_META` alongside the
title, so a fourth type cannot silently inherit capability's rendering. All
three types' rendered output is byte-identical to before the refactor.

Also corrects a test comment that still claimed the reviewer suites were
failing-first against an unmodified module.

* chore(#2904): backfill changeset PR number (#2912)
2026-07-31 08:11:57 -04:00
Tom Boucher
42f4f184c0 fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context

verify-summary's Pattern 1 matched any backticked path-like token with no
context check, so a prose mention of a future deliverable (`shared/types.ts`
in a 'next phase will add…' sentence) was checked for existence and its absence
failed the verdict on a healthy phase. #2685 added shape filtering but no
context check.

- src/verify.cts: both extraction patterns now require a claim label on the line
  (Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no
  longer matches; genuine labeled claims still do.
- gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION
  so relative claim paths resolve against the project root, not the raw cwd
  (subdirectory invocation no longer manufactures missing files).

Regression tests: prose mention not treated as a claim; prose-only SUMMARY
passes; absent claimed file still fails.

* chore(#2844): backfill changeset PR 2910

---------

Co-authored-by: Test <test@example.com>
2026-07-31 03:16:15 -04:00
Tom Boucher
39dbe5e0f5 fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH (#2909)
* fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH

findProjectRoot's heuristic (3) used isInsideGitRepo(parent), which only checked
'does SOME .git exist between start and the ancestor' — it never verified the
.git was co-located with / bounded the trusted .planning/. A nested child repo
(own .git, no .planning) under an ancestor GSD project satisfied the check, so
resolution silently crossed into the ancestor project (wrong identity, exit 0).

Add nearestGitRoot(from, upTo) (fs-walk, no spawn) and use it in heuristics (3)
and (4): if the caller is inside its own nested repo whose root is strictly below
the candidate ancestor, do not return that ancestor. The plain-descendant (#1414),
co-located .git+.planning, and sub_repos/multiRepo cases are unchanged.

Regression test: a nested child .git under an ancestor .planning no longer
resolves to the ancestor; the co-located single-repo case still resolves.

* chore(#2843): backfill changeset PR 2909

---------

Co-authored-by: Test <test@example.com>
2026-07-31 02:30:50 -04:00
Tom Boucher
5d0fd4dc53 fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point (#2905)
* fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point

gsd-code-fixer was the only writer that hand-rolled a git worktree inside the
agent prompt and the only one that never read workflow.use_worktrees. With the
setting explicitly false, --fix still created worktrees; the fresh worktree had
no node_modules, so runs improvised a teardown whose `rm -rf` followed a Windows
junction into the REAL node_modules (silent data loss, 3x observed).

Defect 1: gate setup_worktree + its cleanup tail on workflow.use_worktrees
(same gsd_run query config-get read the four sibling workflows use). When false:
edit/commit in the main checkout (wt='.', no temp branch, no sentinel, no cleanup).

Defect 2: forbid rm -rf on a possible reparse point in the spec — never fall
through to a destructive remove; on failure, stop and surface the error.

Defect 3: REVIEW-FIX records where verification ran (main checkout vs worktree).

The transactional worktree path (#2839/#2990/#2686) is unchanged when worktrees
are enabled. Docs-parity guards in tests/code-review.test.cjs bind the spec to
the fix.

* chore(#2825): backfill changeset PR 2905

* fix(#2825): drop stale agents/gsd-code-fixer.toml ack + spent entries (emitted-attribution)

gsd-test failed: agents/gsd-code-fixer.toml is a STALE ack here — that TOML was
changed by #2834 (now in next), not this PR. The 4 base-carried acks
(autonomous/discuss-phase-assumptions/next/plan-phase) are spent/inert.

---------

Co-authored-by: Test <test@example.com>
2026-07-31 01:57:50 -04:00
Tom Boucher
00c859fae5 fix(#2834): write defaults.json before agent TOML generation on clean Codex install (#2900)
* fix(#2834): write defaults.json before agent TOML generation on clean Codex install

Extracted writeNonClaudeDefaults(runtime) and called it BEFORE installCodexConfig
so the runtime-aware model resolver has resolve_model_ids=omit + runtime=codex in
~/.gsd/defaults.json before agent TOMLs are generated. Pre-fix, a clean first Codex
install generated TOMLs with no model fields (the resolver didn't know the runtime);
a second run fixed it. The original inline defaults-write block (which ran AFTER agent
generation) is replaced by the earlier function call (idempotent).

* chore(#2834): changeset fragment

* fix+test(#2834): acknowledge codex TOML drift (emitted-attribution) + fix test comment window

The emitted-attribution gate flags 19 codex agent TOMLs that now carry model-routing
fields (the fix's correct effect) but can't link them to a .md or src/ change (the fix
is in bin/install.js ordering). Acknowledge the drift in emitted-drift-ack.json. Fix the
test's comment-detection window (300 chars to capture the #2834 rationale).

* fix(#2834): ack remaining 15 codex TOML drift paths

* chore(#2834): backfill changeset PR number (2900)

* chore(#2834): ack code-review.md growth from concurrent merge (rebase pickup)

* fix(#2834): remove stale code-review.md ack (emitted-attribution failure)

CI failed: 'differential attribution over the real tree' — the code-review.md
ack added in 6e0b4b3b8 ('ack code-review.md growth from concurrent merge') is
STALE: this PR's diff does not touch code-review.md (only bin/install.js + tests
+ changeset), so the ack explains growth that isn't here. The base already
absorbed the concurrent code-review.md growth; the ack is inert here and the
gate flags it as stale. Remove it.

---------

Co-authored-by: Test <test@example.com>
2026-07-31 00:47:55 -04:00
Tom Boucher
9f567a1627 fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row (#2902)
* fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row

Two coupled defects in the requirement traceability state machine:

Defect 1 (terminal state): requirements revert-phase (the gaps_found response)
left a row at 'Gaps Found' with no inverse — neither mark-complete's /^pending$/i
guard nor the phase-complete reconcile's /^(?:pending|in progress)$/i accepted it,
so a single failed verification stranded every requirement permanently and blocked
the milestone. Widen both guards to accept 'gaps found' so a genuinely-satisfied
stranded row reaches Complete again.

Defect 2 (false success): mark-complete ORed checkboxHit || tableHit for 'updated',
so on a Gaps Found row it flipped the checkbox but could not move the row, yet
reported updated:true. When a traceability table has a row for an ID, gate 'updated'
on the row moving (tableHit) — a checkbox-only flip on a table-bearing file no
longer lies. The #2140 table_unmatched path (no row for the ID) is preserved.

* chore(#2788): backfill changeset PR 2902

---------

Co-authored-by: Test <test@example.com>
2026-07-31 00:04:04 -04:00
Tom Boucher
79ed181ec0 fix(#2667): run-with-timeout mediates .cmd/.bat spawns on Windows (CVE-2024-27980); fallow pre-pass names failure kind (#2897)
* fix(#2667): mediate .cmd/.bat/.exe spawns on Windows; split fallow pre-pass failure diagnostic

run-with-timeout spawned .cmd/.bat/.exe commands without shell:true on Windows,
tripping Node's CVE-2024-27980 EINVAL (April 2024 security hardening). The fallow
structural pre-pass then no-op'd silently — a hard execution failure read the same
as 'optional dependency absent'.

(A) gsd-core/bin/gsd-tools.cjs runWithTimeout: gate shell:true on
    (win32 && command ends in .cmd/.bat/.exe). Narrow by design — never fires for
    the 7 `bash -c` callers (command is `bash`, no such suffix), so the recorded
    no-shell-for-argv-array security contract (DEFECT.UNBOUNDED-SUBPROCESS) is
    preserved; cmdArgs stays an array. POSIX untouched.
(B) code-review.md fallow pre-pass: name the failure KIND (timeout / spawn failure
    / crash / not-found) so a Windows .cmd spawn failure is not mistaken for an
    absent binary.

Regression test in tests/run-with-timeout.test.cjs gated to win32 (.cmd/.bat/.exe
shims run with exit 0 + non-empty stdout; pre-fix EINVAL → exit 125/empty). POSIX
negative-space test guards the unchanged bash -c callers.

* chore(#2667): changeset fragment

* chore(#2667): backfill changeset PR 2897 + correct body (cmd.exe array, not shell:true)

* fix(#2667): exclude .exe from the win32 spawn-mediation gate; ack code-review.md growth

CI caught two failures on the first push:

1. windows-24: 'exits 124 when the wall-clock budget is exceeded' regressed. The
   gate matched .exe, so the HANG command (node.exe -e 'setTimeout(...)') was
   wrapped in 'cmd.exe /c node.exe ...' — the wrapped child escaped the timeout
   cap's process-group reap (exit 124 never fired; hit the 30s harness backstop)
   AND cmd.exe risked mis-parsing the -e script arg. .exe is INTENTIONALLY
   excluded now: real PE executables spawn fine directly; only .cmd/.bat are the
   CVE-2024-27980 EINVAL cases. The .exe test becomes a negative-space test
   (node.exe spawned directly, exit 0).

2. ubuntu-22: emitted-attribution — code-review.md grew 1177 bytes from the
   #2667 fallow pre-pass failure-KIND case statement; acknowledge it.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 23:14:17 -04:00
Tom Boucher
6e0bc50142 fix(#2666): code-review scopes root-level + extensionless build files, cross-checks against git diff (#2895)
* test(#2666): add regression + docs-parity guards for code-review file scoper

The Tier-2 SUMMARY.md extractor dropped every repository-root file (no `/`)
and every extensionless build file (Dockerfile/Makefile/etc.) via an AND-joined
predicate. Adds behavioral tests against the pure-function mirror plus
docs-parity structural guards that bind the shipped workflow .md to the fix.

RED: the docs-parity guards fail against the pre-fix shipped predicate.

* fix(#2666): accept root-level + extensionless build files in code-review scope; intersect-and-warn

Two coordinated edits to gsd-core/workflows/code-review.md compute_file_scope:

(A) Tier-2 SUMMARY extractor: replace the AND-joined predicate
`/\\//.test(raw) && /\\.[A-Za-z0-9]+$/.test(raw)` (which required BOTH a
directory separator AND a trailing extension, silently dropping every
root-level file and every extensionless build file) with a relaxed predicate
that accepts any path with a trailing extension OR a known extensionless
build basename (Dockerfile/Containerfile/Makefile/Justfile/Procfile).

(B) Tier-3: convert the eq-zero git-diff gate into an intersect-and-warn —
whenever a reliable diff base is available, cross-check the SUMMARY scope
against `git diff --name-only` and warn about (then add) any changed files
the SUMMARY extractor did not surface. Portable (bash 3.2, no associative
arrays) so a partial SUMMARY result can no longer silently ship an
incomplete review scope.

* chore(#2666): changeset fragment

* fix(#2666): use exact whole-line matching (grep -Fxq) in Tier-3 cross-check

Adversarial review found the unanchored `case "$IN_SCOPE" in *"$file"$\\n*`
substring membership test would false-match: a root-level `Dockerfile` in the
diff substring-matches an already-scoped `docker/Dockerfile`, silently skipping
it — reintroducing the exact class of silent-scope-loss bug this PR fixes.

Switch to `grep -Fxq` (exact whole-line match). Add docs-parity guard for
exact matching + the basename-collision regression.

* fix(#2666): resolve gsd-test failures — paraphrase predicate in comment, ack code-review.md growth

gsd-test caught 3 issues on 7469c3f22:
1. The docs-parity guard fired on the .md COMMENT which restated the buggy
   predicate verbatim — paraphrase the comment so it no longer contains the
   exact string the guard detects.
2. Cascade subtest failure from #1.
3. emitted-attribution: code-review.md grew 2814 bytes — acknowledge the
   deliberate #2666 growth in tests/emitted-drift-ack.json.

* chore(#2666): backfill changeset PR number 2895

---------

Co-authored-by: Test <test@example.com>
2026-07-30 22:27:32 -04:00
Tom Boucher
f093738412 fix(#2891): normalize emitted version against the measured tree, not the measuring repo (#2894)
* fix(#2891): normalize emitted version against the measured tree, not the measuring repo

buildParityManifest normalized the install-time {{GSD_VERSION}} stamp using
PKG_VERSION, bound at module load from the MEASURING repo's package.json. Since
#2767, currentManifests({repoRoot}) measures a DIFFERENT checkout, so during a
release cut the baseline worktree (origin/next, 1.8.0) was normalized with the
current tree's version (1.9.0) and its literal 1.8.0 stamp survived into the
hash. All 364 emitted hook paths diverged and the differential attribution gate
hard-failed every finalize/rc run.

Normalize against the version of the tree that PRODUCED the emitted output:
buildParityManifest takes an explicit pkgVersion, and currentManifests resolves
it from the measured tree via a new fail-closed measuredPackageVersion().

* chore(#2891): backfill changeset pr number (#2894)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 22:14:47 -04:00
Tom Boucher
cd5b8643ae fix(#2828): state sync reports correct total_phases on a flat unmilestoned roadmap (#2892)
* fix+test(#2828): total_phases uses roadmap count on flat unmilestoned roadmap

The read-path disk-scan cache fell back to phaseDirs.length (1) when milestoneBounded
was false, even though roadmapPhaseCount (6) was correct for a flat roadmap (no sibling
milestones to conflate). Use roadmapPhaseCount as the floor when > 0, matching the
write-path (cmdStateSync) which already did this. The milestoneBounded flag still flows
to milestoneUnbounded for the percent-skip (#1761 guard preserved). Regression test
asserts state-sync writes progress.total_phases:6 for a flat 6-phase roadmap + 1 phase dir.

* chore(#2828): changeset fragment

* test(#2828): add negative-space coverage (Math.max floor mutant + no-roadmap fallback) — review findings

The 6-phase test alone couldn't kill a Math.max-dropping mutant (1<6). Add: a
3-phase-dir/2-roadmap-phase case proving Math.max(dirs,count) floor; a no-roadmap
case proving phaseDirs.length fallback.

* fix(#2828): refine — distinguish flat unmilestoned from milestoned-unbounded (preserve #1761)

The first-pass fix (roadmapPhaseCount > 0 always) re-broke #1761: a milestoned-
unbounded roadmap (asserted milestone not among existing version headings) conflated
sibling milestones (8 = 4+4). Refine with a hasMilestoneSectioning discriminator:
^#{2,3}(?!Phase) detects non-Phase h2/h3 milestone section headings. A FLAT roadmap
(only ### Phase headings + a # title) has none → safe to use roadmapPhaseCount; a
SECTIONED-but-unbounded roadmap has them → fall back to phaseDirs.length (#1761).
Verified both cases locally (flat→6, sectioned-unbounded→1).

* test(#2828): remove two fragile negative-space tests (phase-dir scanner internals)

The Math.max-floor and no-roadmap tests made assumptions about the phase-dir scanner's
internals (which dirs count as 'realized') that didn't hold. The core regression test
(6-phase flat → total_phases:6) plus the existing #1761 conflation tests (which the
refined fix preserves) provide sufficient coverage.

* chore(#2828): backfill changeset PR number 2892

* fix(#2828): replace ReDoS-prone regex in regression test with line-by-line parse

CodeQL flagged the nested-quantifier regex (`(?:[ \t]+\w+:.+\r?\n?)*?`) in
tests/issue-2828-flat-roadmap-total-phases.test.cjs as a high-severity
catastrophic-backtracking risk. Rewrite the STATE.md progress.total_phases
extraction as a ReDoS-safe line-by-line block walk.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 21:19:36 -04:00
Tom Boucher
54cb4145bf fix(#2765): bump brace-expansion to patched 1.1.18/5.0.9 (high-severity devDep advisory) (#2888)
* fix(#2765): bump brace-expansion to patched 1.1.18/5.0.9 (high-severity devDep advisory)

npm audit fix (non-breaking) bumps the lockfile: brace-expansion 1.1.15→1.1.18
(eslint-nested via minimatch@3.x) and 5.0.6→5.0.9 (stryker-nested). Both 1.1.18 and
5.0.9 were published 2026-07-30 as the patch backports for GHSA-3jxr-9vmj-r5cp /
GHSA-mh99-v99m-4gvg (range <=5.0.7). No overrides needed (in-range bump), no major
bumps, no --force. Production (npm audit --omit=dev) unaffected (devDep only).
Add a structural test pinning the installed versions so the bump can't silently regress.

* chore(#2765): changeset fragment

* fix(#2765): correct changeset issue ref + parse patch version as number (review findings)

- changeset cited #2762 (typo) — fix to (#2765).
- test compared v.split('.')[2] as a string (false-pass for 1.1.9) — parse all segments as Number.

* chore(#2765): backfill changeset PR number (2888)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 20:12:28 -04:00
Tom Boucher
4bd6fb066b chore(#2880): close ADR-2143 deployment misses — table-regex fingerprint + state-document seam migration (#2889)
* chore(#2880): close ADR-2143 seam misses — widen table-regex fingerprint, migrate state-document onto the seam

The no-adhoc-markdown-parsing rule matched only a negated class whose sole
member was a pipe ([^|]), so the stricter and more common [^|\n] spelling
evaded it entirely -- src/state-document.cts hand-rolled exactly that shape
and linted clean. Widen the fingerprint to any negated class excluding a
pipe, which is the ADR-2143 section 7 prohibition as written.

With the rule fixed, state-document.cts goes red. Replace tableRowPattern
with locateFieldRow: a line scan using the markdown-table seam's
splitTableRow for cell semantics, returning the value cell's byte range, and
splice that range instead of running a whole-document content.replace. An
edit now physically cannot cross a row boundary (section 4).

Behavior is frozen -- stateReplaceField has 79 dependents across 5 command
processes. Characterization tests lock all 14 table-branch rows plus CRLF,
extract round-trip and the withFallback caller shape; a fast-check property
asserts every non-target line stays byte-identical.

Refs #2880, epic #2143

* fix(#2880): address adversarial review — lone-CR rows, field-name padding, quadratic scan, over-broad fingerprint

Isolated adversarial review found four defects in the first commit.

1. locateFieldRow split lines on \n only. JS treats a lone \r as a line
   terminator, so the regex it replaced matched rows separated by bare CR.
   "| Phase | 3 |\r| Other | 9 |" returned 3 before and null after. Now
   CR, LF and CRLF are all terminators, byte offsets unchanged.

2. The field name was normalised with trim().toLowerCase(). The old regex
   embedded it verbatim, so its whitespace had to be absorbed by the row's
   own padding -- and because the group is a literal-character match rather
   than a whitespace class, a tab-padded cell does not accept a
   space-padded name. Replaced with an offset-aligned search reproducing
   the original backtracking exactly.

3. The widened fingerprint regex had two unbounded [^\]]* around an
   optional and ran quadratically over every regex source in every linted
   file: 256000 chars took 23 seconds. Replaced with a single-pass scanner
   that never rescans; the same input is now ~1ms.

4. The fingerprint also matched non-table idioms such as [^\s|] and [^"|].
   Narrowed to a class excluding the pipe plus only \n, \r or \t.

Differential fuzz against origin/next: 20000 cases, 0 mismatches.

Refs #2880, epic #2143

* test(#2880): drop wall-clock assertion from the ReDoS regression guard

local/no-elapsed-assertion flagged the elapsed-time check, and CLAUDE.md
bans timing assertions outright as flaky. The 256000-char input stays as
the regression guard for the quadratic scan; correctness of the verdict is
what is asserted. If the quadratic path returns, the test stops completing
and surfaces as a suite timeout rather than a silent pass.

Also adds the changeset fragment for #2880.

Refs #2880

* fix(#2880): spec-correct case folding, property tests, naming

Code-review findings.

The field-name comparison used toLowerCase(). The regex it replaced used
/i WITHOUT /u, and ECMAScript Canonicalize deliberately does not fold a
non-ASCII character onto an ASCII one -- KELVIN SIGN U+212A matched ASCII
K where the old code returned null. Replaced with spec-correct
Canonicalize, including the multi-character uppercase case (eszett -> SS),
which a naive uppercase comparison also gets wrong.

Added the fast-check property tests CLAUDE.md requires for parsers: one
for the negated-class scanner, one for the field-name fold semantics, each
against an independent reference implementation. Both reference impls
failed on first run against real bugs, so neither property is vacuous.

Renamed p2/p3 to name the exactly-three-pipes invariant, and reduced a
duplicated comment to a cross-reference.

Differential fuzz vs origin/next: 20000 runs, 0 mismatches, with the
harness proven to discriminate the KELVIN case.

Refs #2880

* chore(#2880): backfill changeset PR number (#2889)

* docs(#2890): correct the local ESLint plugin path in CONTEXT.md

CONTEXT.md named the local AST-rule plugin directory as
scripts/eslint-rules/, which does not exist. The real location is
eslint-rules/ at the repo root -- what eslint.config.mjs actually
imports -- and CONTEXT.md's own later entry already says so
explicitly, so the file disagreed with itself.

Found by a line-by-line audit of all 1036 lines against the live
graph; this was the only confirmed inaccuracy.

Closes #2890

---------

Co-authored-by: Test <test@example.com>
2026-07-30 19:55:03 -04:00
Tom Boucher
7372d99a26 enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales

The reviewer lane roster was hand-enumerated across five documentation
surfaces and three workflow files that had drifted apart: --kimi-code was
missing from all four translated COMMANDS.md mirrors, --coderabbit from
every workflow forwarding list, and --antigravity from FEATURES.md.

Adds checkReviewerDocsParity, a second pure gate deliberately separate from
checkReviewerLaneParity so a stale doc cannot make the runtime checker look
red. Workflows now derive their flag lists from a new review-lane flags
query instead of hand-enumerating them, which also retires the unanchored
grep that matched --agy inside --antigravity.

Documents the previously absent reviewer body and hostBehaviors field in
the capability manifest reference.

Closes #2800
Closes #2781
Closes #2272

* fix(#2800): key the docs parity table arm on first-cell position

Review found the flag arm was file-scoped, so the forwarding row that lists
every flag in its third cell satisfied it on its own. Deleting a lane's own
reviewer-table row -- the #2781 regression this gate exists to prevent --
therefore passed undetected.

Arm 4 keys on the FIRST table cell, which separates a lane row from the
forwarding row structurally and in every locale. Regression test included.

* fix(#2800): shape-filter the flags subcommand output

All three consumers read review-lane flags through an unquoted command
substitution so the output word-splits into loop items. Phase 2 admits
third-party overlay lanes, so an overlay flag containing whitespace would
inject a second loop item and one containing a glob would expand against
the cwd. Emit only well-formed flags so neither reaches the shell.

* fix(#2800): remove the regex length ceiling and count only prose mentions

Review found two real defects in the docs parity gate.

The never-throws contract was false: building a RegExp from a declared flag
or section title throws SyntaxError past ~100k chars, and Phase 2 admits
overlay lanes whose declared strings are untrusted in length. Every one of
these matches is literal, so String.includes replaces the regex outright,
which also deletes escapeLiteral and the llama.cpp escaping it existed for.

Arm 1 was context-blind: a flag mentioned only inside a fenced example or a
commented-out row counted as documented. Both are stripped before matching.

Also advertises all 13 lane flags in the argument-hint and corrects a stale
eleven-lane count in the slug grammar note.

* test(#2800): repoint the convergence suite off deleted workflow text

The derived flag loop deleted the literal per-flag grep lines four tests
matched on. Two of those failed loudly. The behavioral and property tests
failed SILENTLY instead: their end marker no longer resolved, so the parse
block extracted empty and both passed vacuously, and the property test's
gsd_run stub had a no-op default that hid it.

All now share one extractor and execute the real deployed block through a
gsd_run shim backed by the actual binary. The whitelist assertions become an
anti-parity check: re-adding a hand-written flag list must fail.

Also repairs two vacuous cases in the docs parity suite. The unreadable-doc
test called its own mock rather than the reader, and the integration test
bounded nothing, so a doc losing its marker would have been silently skipped
and still passed green.

* fix(#2800): run the derived flag loop after the launcher preamble

The remote matrix caught a real runtime bug, not a test artifact. In
autonomous.md and plan-review-convergence.md the launcher preamble that
defines gsd_run lives in a separate, LATER bash fence than the derived loop.
Each fence is its own shell, so gsd_run was undefined where the loop ran:
the command substitution yielded nothing and zero reviewer flags would have
been forwarded. Worse than the drift this epic fixes, and silent.

The whole CONVERGENCE_ARGS construction moves as one unit, because the
--max-cycles append sits between the loop and the preamble and would
otherwise have run against an uninitialized variable and then been dropped
by the relocated initializer.

Also documents all 13 lane flags in help/modes/full.md, which the repo gates
bidirectionally against each command's argument-hint.

* test(#2800): repoint the two converge suites off deleted flag literals

Both asserted workflow.includes('--codex') against the hand-enumerated list
the derived loop removed. They now assert the derivation itself, keep --all
and --text (convergence controls, still literal), and add an anti-parity
guard so re-adding a hardcoded list fails.

The lost pass-through proof is replaced with a real one: every flag the
tests used to hardcode is asserted present in the actual roster emitted by
the binary, which is the property the old assertion was protecting.

* test(#2800): acknowledge the workflow byte growth from the derived flag loop

* chore(#2800): backfill changeset pr number to 2882

* fix(#2800): strip HTML comments to a fixed point in the parity gate

CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the
single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick
construction smuggles a commented-out row past the gate and it counts as
documented. Not an injection risk here since nothing is rendered, but it is
the exact false pass this helper exists to prevent.

Strips to a fixed point, then treats any surviving opener as unterminated so
the multi-line branch closes it on a later line. Terminates because every
pass strictly shortens the string.

* test(#2800): pin the comment-smuggling regression with a real reproducer

The obvious fixture for this class does not reproduce it: <!--<!---->-->
leaves a dangling --> rather than a live <!--, and is caught either way, so
it would have passed with and without the fix. The join-trick construction
(<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely
regresses on the single-pass strip and is what the test now uses.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 19:14:13 -04:00
Tom Boucher
b7b5c3712c fix(#2762): chunked --reviews replans instead of no-op + outline resume marker written to file (#2887)
* test(#2762): chunked --reviews must replan, not no-op (outline marker + per-plan --reviews exception)

* fix(#2762): chunked --reviews replans plans instead of skipping 100% + outline resume marker written to file

Defect A: §8.5.1 outline resume-check greped for a marker the agent only RETURNED (never
wrote to the file) → outline always re-ran (broke crash-resume). Fix: the outline agent
writes ## OUTLINE COMPLETE into the file.
Defect B: §8.5.2 per-plan resume-check skipped any plan with frontmatter, no --reviews
exception → --reviews skipped 100% of plans (contradicted §6 'go straight to replanning').
Fix: gate the skip on --reviews being ABSENT. Crash-resume (non-reviews) still skips.
Condensed adjacent §8.5 prose to keep plan-phase.md under the 94519B cap (net -33B).

* chore(#2762): changeset fragment

* chore(#2762): backfill changeset PR number (2887)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 18:52:04 -04:00
Tom Boucher
3af1941948 fix(#2772): resolve four discuss-phase text inconsistencies (dead MAX_PASSES read, gate-prompts drift, circular auto_advance, answer_validation drift) (#2886)
* test(#2772): structural guards for the four discuss-phase text inconsistencies

* fix(#2772): resolve four discuss-phase text inconsistencies

1. auto.md: remove the dead MAX_PASSES/max_discuss_passes config read (contradicted
   the mandated single-pass rule + wasted a shim invocation per auto run).
2. gate-prompts.md: context-handling options now match the actual check_existing flow
   (Update it | View it | Skip, not Overwrite|Append|Cancel); gray-area-option no longer
   mandates 'Let Claude decide' (contradicts discuss-phase.md's no-cop-out rule).
3. discuss-phase.md: auto_advance fallback ends the workflow instead of routing back to
   the already-run confirm_creation step (circular).
4. discuss-phase-assumptions.md: re-sync answer_validation to the parent canonical block
   (had drifted — lost the 'Other' empty-text branch).

* chore(#2772): changeset fragment

* fix+test(#2772): also fix the assumptions auto_advance circularity (review minor 1) + add positive test anchors (review minor 2)

The sibling discuss-phase-assumptions.md had the identical auto_advance→confirm_creation
circularity; fix it the same way (end the workflow). Add positive anchors to both
auto_advance tests so a re-phrased regression can't slip past. File #2885 for the dead
max_discuss_passes config still advertised in settings/registry/docs (review minor 3).

* fix(#2772): keep discuss-phase.md under the 32000B #717 cap + ack assumptions growth

The auto_advance fixes + the assumptions answer_validation re-sync grew both files
past the emitted-attribution gate (and discuss-phase.md past the #717 32000B cap).
Condense the auto_advance prose in both files (discuss-phase.md now net -11, under
cap; auto.md already net -4650 from the MAX_PASSES shim removal). Add
discuss-phase-assumptions.md to tests/emitted-drift-ack.json for its residual +220
(answer_validation re-sync + auto_advance fix).

* chore(#2772): backfill changeset PR number (2886)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 18:14:34 -04:00
Tom Boucher
dbb0a653be fix(#2771): advisor mode spawns registered gsd-advisor-researcher subagent instead of general-purpose (#2884)
* test(#2771): advisor mode must spawn gsd-advisor-researcher, not general-purpose

* fix(#2771): spawn registered gsd-advisor-researcher subagent instead of general-purpose in advisor mode

universal-anti-patterns rule 10 (injected into discuss-phase via <required_reading>)
says NEVER use non-GSD agent types. The advisor mode spawned general-purpose and
manually told the agent to read the def — but gsd-advisor-researcher IS registered,
so spawning by type auto-loads it. Drop the manual-read prompt line (re-specifying
the def is a drift risk) and use the registered type.

* chore(#2771): changeset fragment (mentions follow-up #2883)

* test(#2771): widen manual-read-line regex to deny phrasing variants (review minor)

/read\s+@.*gsd-advisor-researcher\.md/i (case-insensitive, any 'read @' lead-in)
so a drift variant like 'Read @' or 'Load @' can't sneak the manual-def-read back in.

* chore(#2771): backfill changeset PR number (2884)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 17:25:55 -04:00
Tom Boucher
185da024cb fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block (#2881)
* test(#2770): empty contextPath argument must fail closed, not green-skip the decision-coverage gate

The handler conflated empty-arg (caller error) with file-missing (legitimate skip),
returning passed:true/skipped on an empty argument. Add: empty arg → passed:false;
real-path-to-absent-file → legitimate green skip preserved; omitted arg → fail closed.

* fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block

Handler (check-command-router.cts): split the guard — empty/missing contextPath
argument is a caller error (fail closed, passed:false, mirrors #1365); a real path
whose file genuinely does not exist keeps the legitimate green skip.

Workflow (plan-phase.md): recompute CONTEXT_PATH inside the consuming Bash block
(it was set in the step-1 init block, which does not survive into the separately-
spawned gate block — so the gate ran with an empty arg and silently green-skipped).

* chore(#2770): changeset fragment

* fix(#2770): guard workflow empty-glob case (review blocker) + update drift-guard test

The handler now fails closed on an empty contextPath arg, so the workflow's
unguarded glob (empty when a phase genuinely has no CONTEXT.md) would invoke the
gate with an empty arg → passed:false → exit 1, hard-halting the legitimate
'Continue without context' plan-phase path. Guard the empty-glob case: only run the
gate when a CONTEXT.md actually exists. Update the F1 drift-guard test (which gave
false coverage — it only checked for the ${CONTEXT_PATH} token) to assert the
in-block recompute AND the empty-glob guard.

* fix(#2770): keep plan-phase.md under ADR-857 size cap + ack emitted drift + fix drift-guard window

The workflow fix grew plan-phase.md past the ADR-857 phase-6 size cap (94519B) and
triggered emitted-attribution. Condense adjacent §13a prose/JSON to offset (net
+89B, under cap). Add tests/emitted-drift-ack.json acknowledging the residual growth.
Widen the drift-guard test window (the gate invocation is now nested in the empty-glob
guard, so the old 400-char window missed the glob recompute).

* chore(#2770): backfill changeset PR number (2881)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 16:55:10 -04:00
Tom Boucher
7e9ce08e2c test(#2736): property-test parsePhaseFromProse name precedence (#2878)
#2821 reworked parsePhaseFromProse's NAME precedence (dash-vs-paren choice,
status-keyword tails, `Milestone:` tails, paren-stripped separator search, the
lone-ALL-CAPS-tail rule) but shipped it pinned only by hand-picked examples.

The two pre-existing property blocks cover phase-token ANCHORING (#2111) and
"N of M" phase extraction; neither touches name precedence. This adds nine
properties over the canonical parser.

P1 and P9 are the delta guards: both fail against the pre-#2821 paren-first
parser (proved with a standalone mutation harness), because each requires a
genuine em-dash name to win over a co-present parenthetical. P9 additionally
exercises the paren-stripped separator search, since the losing parenthetical
itself contains an em-dash.

P2-P4 are characterization tests for precedence contracts both parser versions
satisfy; P5-P8 pin totality and phase-token extraction, which #2821 left alone.
The section comment states which is which so a future reader does not mistake
the characterization tests for delta guards.

The generator's status-word exclusion filter is a test-local mirror of the
private, unexported STATUSY_TAIL_RE. A divergence-guard test pins that mirror
to observable parser behavior (not source text, which the no-source-grep rule
forbids), so an implementation vocabulary change fails loudly instead of
silently weakening every property that depends on the filter.

Also clears six pre-existing no-unused-vars lint warnings surfaced by the lint
run for this change (dead bindings in four unrelated test files, deleted rather
than underscore-renamed); removing the never-called openCodeBlock helper
cascaded to its now-unused fs/path/reviewPath consts in both fold scopes.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 15:27:52 -04:00
Tom Boucher
7cbfb2819f fix(#2764): extend no-path-literal-in-assert to membership/substring checks (catches the #2728 Windows CI shape) (#2879)
* test(#2764): no-path-literal-in-assert must flag membership/substring checks over path-returners

The equality-only visitor missed .includes/.indexOf/.startsWith/.endsWith over a
path-returning receiver (directly or via a .map() hop) — these pass lint and fail on
Windows. Add RuleTester invalid cases for each shape + valid cases for the
suppressions (POSIX-normalized receiver, non-path receiver, no-slash arg).

* fix(#2764): extend no-path-literal-in-assert to membership/substring checks over path-returners

Add a third CallExpression shape: .includes/.indexOf/.startsWith/.endsWith/.match
over a receiver that traces to a path-returning call (directly, or through ONE
.map(f => path-…) hop — the #2728 shape that passed lint and failed Windows CI).
Reuse isPathReturningCall/isPosixSlashStringLiteral/isPosixNormalizerCall; respect
the normalizer + Windows-excluded suppressions. The rule is scoped to tests/**/*.test.cjs.

* chore(#2764): changeset fragment

* test+docs(#2764): add .match test coverage, Windows-excluded membership case, fix doc accuracy (review findings)

- MAJOR: .match was supported but untested — add invalid (.match with slash) +
  valid (.match no-slash) RuleTester cases.
- MINOR: Windows-excluded membership case was claimed by the matrix but absent —
  add a platform-guard valid case in membership form.
- MINOR: .map() hop comment said 'ONE' but recursion allows nesting — reword to
  'recursively'.
- Add known-boundary (d) note for the membership/.match shape.
- Fix a premature comment-close from a literal */ in the glob example.

* chore(#2764): backfill changeset PR number (2879)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 15:26:43 -04:00
Tom Boucher
ef7a71fc09 fix(#2754): make parseStateMd CRLF-safe — frontmatter no longer drops on Windows-authored STATE.md (#2865)
* test(#2754): parseStateMd must parse CRLF STATE.md identically to LF

The frontmatter fence regex and downstream splits used literal \n, so a CRLF
STATE.md dropped the ENTIRE frontmatter block. Assert CRLF/LF parity (the
production contract) across full frontmatter, next_phases flow + block forms,
the progress nested block, and null handling.

* fix(#2754): make parseStateMd CRLF-safe — frontmatter fence + splits use \r?\n

The fence regex, scalar-line split, next_phases block-list regex, and progress
block regex all used literal \n, so a CRLF (Windows-authored) STATE.md dropped
the ENTIRE frontmatter block — every statusline field was silently absent. Use
\r?\n throughout, mirroring the CRLF-safe extractFrontmatter in src/frontmatter.cts.
LF behavior unchanged.

* chore(#2754): changeset fragment

* test(#2754): pin parseStateMd↔extractFrontmatter parity (Generative-Fix Divergence guard)

The statusline parseStateMd and the canonical extractFrontmatter both derive GSD
state from STATE.md frontmatter and diverged once already (the CRLF bug this PR
fixes). Add a cross-parser parity assertion (CLAUDE.md parallel-surfaces rule)
over the overlapping fields, under both LF and CRLF, so a future divergence on a
scalar shape or line ending is caught here. (isolated-review minor finding)

* chore(#2754): backfill changeset PR number (2865)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 14:34:28 -04:00
0xdhx
82ca13f5a5 fix(#2736): write current_phase_name from the transition intent, not the lossy prose round-trip (#2821)
* fix(#2736): intent-first current_phase_name on transitions; dash-first prose precedence

Primary: completePhase (adapter) and beginPhase (via readModifyWriteStateMd
options) pass the intent-held display name to syncStateFrontmatter as an
authoritative override, applied after every derive/preserve/carry-forward
step — so the lossy prose round-trip can never destroy a name the transition
just resolved. Names containing a parenthetical
(`Closer-ruling measurement (D1a)`) now land in frontmatter verbatim instead
of collapsing to the parenthetical (`D1a`).

Secondary (#1695 AC #3 residual): parsePhaseFromProse prefers the em-dash
name when it is a genuine name (not a status keyword, not a `Milestone:`
tail), else falls back to the parenthetical — satisfying both first-party
writer shapes (`N — Name (aside)` and `N (Name) — EXECUTING`). Still lossy
for paren-containing names, which is why the intent-first override is the
primary fix.

plannedPhase carries no name in its intent, so it is naturally out of scope.

Fixes #2736

* docs(changeset): backfill PR number for #2736 fragment

* fix(#2736): drop an unnecessary type assertion on result.data

StateTransitionResult.data is already `Record<string, unknown> | undefined`,
so the cast was a no-op and tripped @typescript-eslint/no-unnecessary-type-
assertion (CI lint-tests red on the first push; every test lane was green).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-30 13:53:39 -04:00
Tom Boucher
3f6b063fbb chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans

Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers
iterate declared lanes instead of hand-authored per-CLI bash.

Five additive descriptor amendments, each forced by a lane that ships today:
- LaneHandler gains 'opencode' — the lane rebuilds its review from assistant
  text parts of a --format json stream; a plain stdout copy re-breaks #1936.
- modelConfigKey — antigravity's key is review.models.agy, not .antigravity,
  so resolving by slug silently dropped a configured model.
- defaultHost/fallbackModel — Phase 4 federated every *_host with a default of
  empty string; the real fallback only existed in the bash.
- args becomes an argv template with a closed four-placeholder vocabulary.
  Positional splicing produced 'codex --model M -o F exec --ephemeral', which
  is not a valid invocation: codex injects in the middle, twice.
- kimi-code lane, with the bounded command-capability probe (needle
  --output-format) that tells Kimi Code from the legacy python kimi-cli.

Parity gate re-pointed: the workflow-text families it scanned are the text this
phase deletes, so they are replaced by descriptor-to-registry parity plus an
anti-parity check that no bespoke leg returns.

jq, curl and external timeout/gtimeout all drop out of the review path.

Refs #2782

* chore(#2799): add review-lane query surface and widen the manifest vocabulary

Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow
loops over, projects all twelve lanes into their capability manifests, and
widens capability-validator for the amendments.

opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's
own admission rule: one lane, justified by a documented upstream defect data
cannot express (#1936 — the agent can end its turn with zero output tokens and
--format default then drops the assistant text entirely).

Two bugs caught by an end-to-end stub run and fixed here:
- loadConfigResolved returns a provenance wrapper, not the config; using it
  directly resolved every key to undefined, which reads as 'nothing
  configured' and silently dropped every model override.
- hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced
  with a PATH scan that spawns nothing at all.

Refs #2782

* chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews

Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved
lanes, and renders REVIEWS.md sections from each lane's declared
reviewsSection instead of thirteen hardcoded headings. review.md drops from
1104 lines to 507 (61KB to 28.7KB).

Parity gate re-pointed, as agreed: the leg-marker and section-heading families
scanned exactly the text this phase deletes, so they are replaced by
descriptor-to-registry parity in both directions, plus an anti-parity check
that fires if a bespoke leg is ever re-added. Enum, emitting sites and the
Object.keys lock moved together.

The budget-trim helper is hoisted out of the Ollama leg: it was always
lane-agnostic, and any lane may now declare a promptBudgetKey.

Refs #2782

* feat(#2799): bind the consented egress host and re-verify it at invocation

Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3
but was not implemented: ConsentRecord had no host field and nothing in the
tree bound one, so this phase's rule-4 comparison had no baseline.

ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design:
isValidConsentRecord does not require it, so every record already on disk
stays valid and no re-consent storm fires (D4 rule 5). It is deliberately
excluded from disclosureSignature — the loader has no config resolver, so
folding a config-derived value in would make loader and lifecycle compute
different signatures for the same manifest and re-prompt forever.

Install resolves hostConfigKey (falling back to the lane's declared
defaultHost, which is what the invocation path uses) and records it.
Invocation re-resolves and blocks on mismatch rather than silently
redirecting. Absence allows: no record, or a record predating the field,
means nothing to compare — denying there would break every existing
local-model user on upgrade.

Refs #2782

* test(#2799): cover the resolver, runner and handlers; retarget the parity suites

Adds the golden invocation-plan table (one row per shipped lane, derived from
the bash legs rather than the descriptor types) plus runner coverage for the
probe, empty-output policy, the three handlers and the egress check.

Retargets the existing suites onto the new contract: descriptor-to-registry
parity, the anti-parity check, the opencode handler, and the twelfth lane.

Two corrections found by running them:
- modelConfigKey was required; that breaks D4 rule 2, since a reviewer
  manifest authored before this phase would fail validation on upgrade. It is
  optional, read as null when absent.
- the antigravity non-zero-exit test pre-seeded the transcript, which asserted
  that a STALE entry leaks through — the exact bug the watermark prevents. The
  spawn now appends, as the real tool does.

Refs #2782

* fix(#2799): restore agy --add-dir and the self-report prompt in the handler

Retargeting the three legacy reviewer suites off the deleted bash surfaced two
real regressions in the port, both #2176:

- --add-dir was dropped. Without it agy's permission context never receives the
  cwd repo, so the agent anchors on its own scratch dir and reviews the plan
  text in isolation — the exact failure the Review Instructions forbid. It is
  capability-probed, because an older agy rejects the unknown flag outright and
  a lane that fails to start is worse than one running on the prompt anchor.
- the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS
  self-report, which is what makes a blind review distinguishable from a
  grounded one. antigravity now builds its own prompt variant.

Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned
model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own
log is the only evidence that anything failed.

The three suites now assert against the plan and the handler instead of
matching fence text, so they no longer need allow-test-rule exemptions.

Refs #2782

* docs(#2799): document the declared lanes, the new flag, and dropped prerequisites

COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph,
which is now false: no lane requires jq, curl or an external timeout. Adds the
changed-egress-destination behavior, since a blocked lane is something a user
can hit.

CONFIGURATION.md records that the model config key is declared per lane rather
than derived from the flag — antigravity's is review.models.agy — and adds
review.models.kimi-code.

reviewer-instances.md now routes an instance through its lane's single
invocation seam instead of a copied per-adapter bash block, which is what lets
a cross-cutting fix reach instances for free. That required implementing the
--model/--agent/--as flags it documents; --model re-resolves through the lane's
argv template rather than splicing, so the flag lands where the lane declares
it rather than ahead of a subcommand.

CONTEXT.md glossary gains both new modules.

Refs #2782

* chore(#2799): drop the stale emitted-drift acknowledgment

The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md.
That file now shrinks by ~32KB and every emitted hash that moved is
attributable to this diff, so the ack no longer explains anything. Removing
the last entry means removing the file: its presence is the alarm, and an
empty one signals nothing.

Verified by deleting it and re-running the attribution and provenance gates
plus lint:ci — all green without it.

Refs #2782

* docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782

Five additive amendments, each forced by a lane that ships today, plus two
corrections the phase had to make rather than work around:

- D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so
  this phase's rule-4 comparison had no baseline. Recorded because an ADR
  asserting a rule was delivered is exactly what stops a later phase checking.
- The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families
  scanned the text this phase deletes.

Also records that D7's 'skip the probe where no bounding mechanism exists'
carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded
on every stock macOS host, which ships neither timeout nor gtimeout.

Refs #2782

* fix(#2799): close four defects found by adversarial review

Two confirmed bugs, both reproduced before fixing:

- resolveLanePlan was not total. An openai-http lane with a missing or
  non-object invoke dereferenced inv.hostConfigKey and threw, contradicting
  the module's own documented contract; the spawn branch guarded correctly and
  the http branch did not. The CLI seam resolves every selected lane in one
  map, so one malformed overlay manifest would have aborted the whole review
  rather than dropping its own lane. Guarded, plus a per-lane try/catch at the
  seam so a throw can never take down siblings.
- A reviewer-instance model was silently dropped for any lane declaring
  modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates
  that cli is a known slug but never that the slug accepts a model, so a user
  could configure one, get a clean run, and never learn a different model
  reviewed their plan. Now warns explicitly.

Two hardening fixes:

- The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in
  the resolver rather than inherited from a validator that does not run on this
  path — the module documents itself as the overlay-manifest trust boundary, so
  it should not depend on someone else having checked.
- normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses
  with an empty hostname, so it became 'localhost://11434' and was compared and
  requested as if real. An empty hostname now means not-a-URL.

Also documents the one gap that cannot be closed here: the antigravity
watermark is keyed by workspace, so two concurrent reviews of the same repo
share a transcript. agy exposes no per-invocation id to filter on, so the
handler now states which half of its never-stale guarantee actually holds.

Refs #2782

* test(#2799): retarget the remaining eight review.md-asserting suites

The remote runner found 37 failures the local sweep missed (it hit the shell's
two-minute cap before reaching these). All eight extract per-CLI bash from
review.md that this phase deletes; each protects a real invariant, so each is
retargeted onto the plan, the runner or the handler rather than removed.

Three real defects surfaced by doing so:

- effort args never reached ANY lane. model-resolver.cjs exports no
  resolveExecution, so effortFor silently returned [] every time. Restored by
  calling the same bounded resolve-execution query the bash legs used — and
  NOT with --raw, which prints the resolved effort rather than the picked
  field, so claude got 'low' instead of '--effort low'.
- the timeout guidance lost 'a silent empty output is a timeout kill, not a
  crash' — the operator note that exists because of the Codex 0xc0000142
  misdiagnosis. Restored.
- the opencode handler dropped EMPTY assistant text parts. The shipped jq was
  , and  only substitutes for false/null — an empty
  string is truthy in jq and contributed a blank line. Found by a property
  test shrinking to ['', ''].

The opencode property suite no longer spawns jq at all, which deletes the
#2099 hang mechanism it was architected around rather than mitigating it.

Refs #2782

* fix(#2799): register the two new generated modules, and untrack them

The remote runner caught build output committed to git. Both new modules
compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated
that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own
review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each
bin/lib/*.cjs is linted xor ignored according to migration state" failed.

Registered both in .gitignore and eslint.config.mjs alongside the Phase 1
module, and dropped them from the index. Nothing about the shipped behaviour
changes; the artifacts are rebuilt by build:lib.

This is the new-.cts-module registration ripple, and it is the one part of it I
had not completed - the CONTEXT.md glossary and the inventory manifest were
already done.

Refs #2782

* chore(#2799): backfill changeset pr number to 2861

* chore(#2799): backfill changeset pr number to 2861

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:48:06 -04:00
Tom Boucher
97e9fd50bd fix(#2752): make tool_input.path authoritative over model-controlled file_path in gsd-phase-boundary.sh (#2860)
* test(#2752): path is authoritative over model-controlled file_path in phase-boundary hook

Rewrite the #2304 precedence test (which pinned the buggy file_path-wins
behavior with an impossible no-tool_name {file_path,path} payload) to assert
the correct precedence with realistic Kimi-shaped payloads: a real .planning/
write with a decoy file_path must still fire the reminder (suppression repro),
and a write elsewhere with a decoy .planning/ file_path must NOT fabricate one
(fabrication repro). Update the parity vocabulary alarm to the new expression.

* fix(#2752): make tool_input.path authoritative over model-controlled file_path in gsd-phase-boundary.sh

The hook consulted file_path first, path second. kimi-cli executes on path and
sends path only; file_path on a Kimi payload is always model-supplied. So a
model-supplied decoy file_path could suppress the reminder for a real .planning/
write or fabricate one for a file never touched. Flip the precedence so path is
authoritative and file_path is the fallback (Claude Code emits file_path and no
path, so the fallback must remain). Mirrors the #2595 JS-guard fix.

* chore(#2752): changeset fragment

* chore(#2752): clarify comment/changeset — JS guards use upstream normalization, shell hook applies precedence directly (review minor 1)

* chore(#2752): backfill changeset PR number (2860)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 12:32:00 -04:00
Tom Boucher
aa19e3478c fix(#2854): pin the emitted gate to the base the tree was merged with (#2859)
* test(#2854): failing-first coverage for CI baseline export provenance

Extracts the export decision out of main() behind injected IO so it is
unit-testable, preserving today's export-whenever-present semantics, and
adds the matrix that proves those semantics are wrong.

The PR lane restores the emitted baseline keyed on the PR's recorded base
sha while the gate resolves the base ref live, so the two drift whenever
next advances mid-flight. The restore was published straight to
GSD_EMITTED_BASELINE, where a mismatch is fatal, turning a recoverable
cache into a hard failure on diffs that touched nothing related.

Also renames the stale-env fixture from 'from-cache-restore.json' to an
operator-pin name: that fixture asserted the exact conflation this bug
is, documenting the defect as intended behavior.

Refs #2854

* fix(#2854): validate a restored baseline before publishing it as an operator pin

GSD_EMITTED_BASELINE is an operator pin: resolveBaseline() treats a mismatch
there as a hard stop, because the operator said "use this one". CI published
its cache restore to that same variable whenever the file merely existed, so
a restore keyed on the PR's recorded base sha - while the gate resolves the
base ref live - turned a recoverable cache into a fatal error whenever next
advanced mid-flight. Required tests went red on diffs that touched nothing
related, and named a test file the contributor never opened.

The export step is the boundary, so it is the boundary that validates. It now
publishes only a baseline already valid for the sha under test, judged by
validateBaseline so the staleness rule keeps one definition. Anything refused
is left to be found via the cache path, where a mismatch degrades to the
in-job build exactly as ADR-2719 SS5 specifies. The operator hard stop is
untouched, and the fast path still hits on a current cache.

Also reports the sources actually reached rather than asserting all three
ran: the failure message claimed an in-job build it had returned before
calling, sending contributors after a rebuild that never happened.

Fixes #2854

* fix(#2854): pin the emitted gate to the base the tree was actually merged with

The differential compared a tree built on one commit against a baseline at a
different one. "Rebase check" merges pull_request.base.sha, pinned by #2472 so
all 12 matrix jobs agree on one tree, but resolveBase() fell through to
origin/next, which fetch-depth 0 leaves at the live tip. Nothing set
GSD_EMITTED_BASE, so whenever next advanced mid-flight the two disagreed.

The cached baseline, keyed on base.sha, was correct for that tree and was
rejected as STALE by a target that was not. Required tests went red on diffs
that touched nothing related, naming a test file the contributor never opened.

The near miss is the worse half: had resolution gotten past the baseline step,
a baseline at the live tip would have attributed commits merged to next in
between to the PR under test. The hard stop was shielding us from a wrong
answer, so making it fall through would have made this worse.

Pins GSD_EMITTED_BASE to the same expression as CI_REBASE_BASE_SHA in every
rebase-merged job, with a parity test asserting the two cannot diverge. That
parity check immediately caught a third lane, test-inert, that merges a pinned
base and had been missed.

Fixes #2854

* fix(#2854): keep the export step self-contained across the package boundary

scripts/ ships in the npm tarball and tests/ does not, so requiring the
validator across that boundary is MODULE_NOT_FOUND in a published install.
The export step now reads the same GSD_EMITTED_BASE pin the gate resolves
through, and applies a cheap self-contained precondition; validateBaseline
remains the sole authority and still runs downstream on whatever is
published, so there is no second opinion to drift.

Reading the pin rather than re-deriving a base is the point: a second,
divergent base lookup is exactly what caused this bug.

The same hazard pre-exists in scripts/gen-emitted-baseline.cjs, which ships
and requires three tests/ modules. Filed as #2858 rather than folded in:
fixing it means relocating the shared helpers out of tests/ and updating
every consumer, which would bury this change.

Refs #2854

* fix(#2854): stop announcing the restored cache through the operator-pin door

Two independent reviewers found the same blocker in the previous approach.
Validating before publishing to GSD_EMITTED_BASELINE only narrowed the hole:
the precondition gated on sha equality alone, so a document with a MATCHING
sha but a wrong schema version or malformed manifests was still announced as
an operator pin and still hard-stopped downstream. That reproduces this bug's
own class, triggered by malformation instead of staleness.

The step was never load-bearing. The cache restores to
.gsd-cache/emitted-baseline.json, which is resolveBaseline's DEFAULT_CACHE_PATH
and is read whether or not anything announces it. Publishing the same file to
the pin door could only ever convert recoverable into fatal, so the step and
its script are deleted rather than made cleverer. Every failure mode now
degrades to the in-job build by construction, and validateBaseline is once
again the only thing that judges a baseline.

Coverage moves to where the behaviour lives: stale sha, wrong schema version,
manifests array/absent, non-object documents, unreadable file, and the 39/40/41
hex boundary all assert degradation via the cache path. Adds the empty-pin case
a reviewer flagged as untested - the pin is job-level env, so on push events it
renders as an empty string, and only baseRefCandidates' truthy check keeps it
out of the candidate list.

Fixes #2854

---------

Co-authored-by: Test <test@example.com>
2026-07-30 11:19:45 -04:00
Tom Boucher
b36e3b7e1f fix(#2751): normalize bare gsd-tools command-position calls to gsd_run in shipped source (#2851)
* test(#2751): regression guard — no command-position bare gsd-tools calls

Agents/workflows instructed bare `gsd-tools <verb>` invocations that fail with
'command not found' on a shim-only install (#725 fixed only the Codex conversion
pipeline; the Claude-facing source shipped them verbatim). Adds a source-text
guard (allow-test-rule: source-text-is-the-product) scanning agents/*.md +
gsd-core/workflows/*.md for the operative shape `gsd-tools <verb> <arg>`,
excluding command -v probes / resolver definitions, with a documented
PROSE_ALLOWLIST for descriptive mentions that name the command without
instructing literal invocation. A stale-allowlist check ensures entries stay
real. RED first; source fix lands next commit.

* fix(#2751): normalize bare gsd-tools calls to gsd_run in Claude-facing source

The 12 command-position bare `gsd-tools <verb>` instructions across agents/ and
gsd-core/workflows/ failed with 'command not found' on a shim-only install (no
gsd-tools binary on PATH). #725 fixed this only for the Codex install-conversion
pipeline; the Claude-facing SOURCE shipped the bare calls verbatim, and new ones
kept accumulating (new-project.md:114 landed 12 days AFTER #725 closed).

Rewrite each operative site to the portable `gsd_run` resolver that the same
files already define (3-18x each) — a pure command-position token swap
preserving all arguments, flags, --files, and surrounding prose. Every runtime
now benefits from one source change instead of each needing its own converter.

Touched sites (12): gsd-intel-updater (validate/snapshot/extract-exports),
gsd-code-fixer (query commit), gsd-planner (learnings.query), gsd-project-
researcher (websearch/research-plan/classify-confidence), gsd-phase-researcher
(websearch/research-plan/classify-confidence), new-project (project-instruction-
file / commit --files), new-milestone (commit --files). Preserves command -v
gsd-tools probes, resolver-snippet definitions, and the 4 descriptive prose
mentions that NAME the command without instructing invocation. RED @ 1bb12ba1.

* chore(#2751): allow-test-rule issue ref + changeset fragment

Add the (#2751) tracking ref to the source-text-is-the-product annotation per
ADR-456, and the .changeset Fixed fragment (pr:0, backfilled post-PR).

* fix(#2751): convert remaining command-position bare gsd-tools calls (verify-summary, windows, worktree, smart-entry, quick-tasks-append)

Isolated adversarial review (Step 4) found the first pass missed genuine
command-position bare calls because the regression test's hand-maintained
6-verb list silently false-passed verify-summary (the 'verify' branch matched
the prefix then died on the hyphen) and omitted windows/worktree/smart-entry/
quick-tasks-append entirely. Convert these 8 additional operative sites across
new-project.md, new-milestone.md, ship.md, execute-phase.md, progress.md,
smart-entry.md, quick.md.

* chore(#2751): backfill changeset PR number (2851)

* fix(#2751): normalize allowlist paths to forward slashes — Windows path-separator false-flag

The PROSE_ALLOWLIST is keyed by file:line using forward-slash paths, but
path.relative() returns backslash separators on Windows, so the allowlist
lookup failed and the 6 descriptive mentions were flagged as offenders on
the windows-latest CI lane. Normalize rel to forward slashes before the
lookup so the allowlist matches identically on every OS.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 11:18:07 -04:00
Tom Boucher
0408276791 chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema

Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the
ADR's swap amendment: a federated config slice lives inside a
capabilities/<id>/capability.json, and three of the five key families had
no capability directory until 5a created them.

Four key families move to the lanes that use them; the central-schema
removal and the federated addition land in this one commit because the
exclusivity invariant fails the build on a key present in both.
review.max_prompt_tokens, review.default_reviewers and
review.reviewer_instances describe policy ACROSS lanes and stay central.

Two things the issue did not name, both found while building it:

1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys
   against manifest.validKeys only, and two of the four families
   (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>)
   were pattern-backed. That is not cosmetic: isCentralConfigKey consults
   those patterns and mergeFederatedConfig skips every key for which it
   returns true, so declaring a slice while the pattern survived would
   have shipped an INERT slice behind a green gate — the exact
   half-migrated shape the invariant exists to prevent. The gate now
   loads the patterns from the same manifest the runtime reads.

2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a
   federated key always resolves to its declared default. The three
   fallback guards in review.md checked only empty-or-"null", so a user
   who set the GLOBAL review.max_prompt_tokens would have silently lost
   trimming on the HTTP lanes. The guards now treat 0 as unset.

D9 says review.models.<slug> is owned by "the lane whose slug it names".
That is false for one lane: the shipped key is review.models.agy while
the slug is antigravity. Ownership follows the lane; the key name is
preserved, because renaming would break every config that sets it.

Existing tests updated rather than left asserting the old world:
config-get on a cleared federated key yields empty instead of
not-found (what #2046 actually protects — never persisting the literal
"null" — is unchanged and still asserted); the config-schema dynamic
pattern representative moves to reviewer_instances; the
prototype-pollution guard case moves to a surviving dynamic prefix so
alert #26 keeps its coverage, with a new case asserting the old key is
now rejected earlier; and Phase 2's harvest-widening inertness assertion
becomes an ownership assertion, since Phase 4 is what consumes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives

A federated config key always resolves to its declared default, so an
unset per-lane prompt budget needed a value the workflow could treat as
'not configured'. The first cut used 0 — which is wrong: 0 is already a
LEGITIMATE per-lane budget meaning 'do not trim this lane' (the
early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it
as unset would have silently switched a user who deliberately disabled
trimming for one lane onto the global budget.

The sentinel is now -1, which is not a valid token budget, so all three
states stay distinguishable: unset falls back to global, an explicit 0
disables trimming for that lane, and an explicit N is used. Locked by
three CLI round-trip tests.

Surfaced by the isolated security reviewer before it crashed mid-run;
verified independently against the shipped trim guard rather than taken
on trust.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): update central-registration assertions and stay under the review.md cap

The remote runner caught both; my local sweep missed the files.

1. tests/plan-review-convergence.test.cjs asserted the three local-server
   host keys are in VALID_CONFIG_KEYS. They are federated to their lane
   capabilities now, and the exclusivity invariant forbids a key living
   in both places. What #2306-local actually protects is that config-set
   ACCEPTS them, so that is what is asserted — via isValidConfigKey, the
   predicate config-set itself uses, which spans central and federated.
   A second assertion pins federated ownership, so a silent reversion
   back to the central schema fails too.

2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap
   is a red line, not a budget to raise. The three per-lane budget guard
   comments were near-identical; condensed to one terse line each. 61371
   bytes, 69 to spare. Real extraction to workflows/review/modes/ is
   Phase 5b/6 work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs

Isolated security review findings.

MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse
error and returned [], while its sibling loadCentralConfigKeys, reading
the SAME file, writes to stderr and throws ExitError(1) on that identical
failure class. Fail-open here defeats the gate this function exists to
feed: with zero patterns, validateCrossCapability's pattern-collision
check silently passes and an inert federated slice ships green. It was
masked in the one production call site only because loadCentralConfigKeys
runs first against the same path — a coincidence of ordering, not a
guarantee, and this function is exported and called standalone. The two
now share a contract: ENOENT is the legitimate absent case, anything else
throws loudly. A single unparseable PATTERN is still skipped, which
degrades to "checked less" rather than blocking every build. The branch
had zero coverage; it now has two tests (malformed JSON, EISDIR).

MINOR — docs/CONFIGURATION.md still listed review.models.qwen and
review.models.cursor as settable, ~770 lines below this PR's own new
Ownership section. Those lanes take no model flag, so they declare no
model key and config-set now rejects them. Rows removed; the missing
review.models.agy row added; the per-reviewer budget row corrected to
name only the lanes that own a budget key, and to document that a
per-lane 0 disables trimming for that lane.

Also fixes a shadowed "raw" binding introduced by the fail-closed change,
which made the generator unrequirable — caught immediately by its own
--check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2797): backfill changeset pr number to 2841

* fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env

tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next
stays green, and it is currently blocking at least three unrelated PRs
(#2841, #2832, #2827) with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

The existing loop comment attributes this to `git add .` rehashing O(n^2)
blobs "before the object write had landed" and works around it by staging
one path per iteration. That is not the cause: sequential execFileSync
calls cannot race each other's object writes, and the failure persisted
after that change — it simply moved to a lower commit index.

The cause is that the git() helper inherited the runner's environment. A
leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's
index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the
blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole
operation. In every case `git commit` then cannot resolve a blob it just
staged, which is precisely the error above.

Verified by negative control: with GIT_DIR exported, this test fails on
the unfixed helper (the git commands operate on the wrong repository
entirely); with the helper stripping GIT_* it passes. The single-path
staging is kept — it is genuinely less work — but it is no longer load
bearing.

Found while shipping #2797. Fixed in place rather than deferred: it is a
defect surfaced during the work, and it is blocking other contributors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2452): build the base-advance commits empty, removing the lost-object class

The base-ref guard has been failing in CI with:

  error: invalid object 100644 <sha> for 'base-N.txt'
  error: Error building trees

It is currently red on at least three unrelated PRs (#2841, #2832,
#2827) while next stays green.

Two theories have now been tried and neither held. #1881 blamed `git
add .` rehashing O(n^2) blobs and switched to staging one path per
iteration; the failure moved from commit 32 to commit 25 and carried on.
The preceding commit here made the git helper hermetic against leaked
GIT_* environment — that IS a real vulnerability (with GIT_DIR exported
the helper operates on the wrong repository entirely, proven by negative
control) but it produces a different error than CI reports, so it is not
demonstrably the cause either.

Neither trigger reproduces off-CI, so this stops guessing at the trigger
and removes the failure CLASS instead. The loop needs base-branch DEPTH
and nothing else: no assertion reads these commits' contents, and
base-side files cannot appear in `origin/base...HEAD` regardless.
`--allow-empty` writes no blob and no tree, so there is no object for the
index to reference and lose. It is also far less work than 60
write+hash+index cycles.

The guard still proves its mechanism: the test asserts that a --depth=1
base fetch FAILS and a full fetch resolves, so a broken topology would
surface immediately rather than passing vacuously.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Test <test@example.com>
2026-07-30 08:52:56 -04:00
Tom Boucher
4f6935e29b fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks (#2846)
* test(#2717): CommonJS marker for cursor/windsurf/codex staged .js hooks

Cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) stage .js
hook scripts via dedicated paths that bypass installSharedHooksBundle — the
only writer of the {"type":"commonjs"} marker. Under a config root declaring
{"type":"module"}, Node loaded those scripts as ESM and every require()
failed with 'require is not defined', silently disabling the runtime's hooks.

Adds regression tests (RED first, fix lands next commit):
- parametrized cursor/windsurf/codex install asserts hooks/package.json exists
  with exactly GSD's marker content;
- end-to-end: a cursor require()-using hook loads under a planted ESM-typed
  config root without the require-is-not-defined error;
- the ensureCommonJsMarker / removeCommonJsMarkerIfGsdOwned contract: GSD
  markers are removed on uninstall, user-authored package.json is never touched.

* fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks

The {"type":"commonjs"} marker lived only inside installSharedHooksBundle,
which cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) never
reach. Their .js hooks are staged by dedicated paths, so under a config root
declaring {"type":"module"} Node loaded them as ESM and every require()
failed with 'require is not defined', silently disabling those runtimes' hooks.

Decouple the marker write into a shared helper so any code path that stages
.js hooks can ensure it lands in the SAME directory as the scripts:

- src/runtime-hooks-surface.cts: add ensureCommonJsMarker(dir) +
  removeCommonJsMarkerIfGsdOwned(dir) (byte-identical content to
  installSharedHooksBundle's marker; preserves a user-authored package.json on
  both write and uninstall). Call ensureCommonJsMarker(hooksDir) from
  writeCursorHooksJson + writeWindsurfHooksJson; call
  removeCommonJsMarkerIfGsdOwned on their matching remove paths. Export both.
- bin/install.js: call hooksSurface.ensureCommonJsMarker after the codex hook
  copy; call hooksSurface.removeCommonJsMarkerIfGsdOwned in the generic
  hooks-removal loop (safe no-op where no marker exists).

No change to which runtimes receive the shared bundle, the !isCodex gate,
skipSharedHooksInstall, or kimi/kimi-code/cline/copilot/trae/zcode (all
unchanged — audit in the diagnosis). RED @ dbb7d2bb (6 failures: 3 missing
markers + the ESM require error + missing helpers); GREEN pending.

* docs(#2717): changeset fragment (pr:0, backfilled post-PR)

* chore(#2717): regen codex/cursor/windsurf install-tree fixtures + attribution ack

The fix adds hooks/package.json to those three runtimes' install trees (the
new CommonJS marker), so the golden install-tree fixtures gain one path each
(regenerated via npm run gen:install-tree). emitted-attribution (ADR-2719)
flags the 3 emitted hooks/package.json paths under the hooks-built rule;
acknowledge them. Also drops 5 spent ack entries left by now-merged PRs
(#2694 code-review.md/code-review-fix.md, #2695 worker/registry, #2794
review.md) — they are stale on this branch (base already carries them).

* fix(#2717): codex ESM-root behavioral test + hooks-built provenance for package.json

Two review-driven follow-ups on the #2717 fix:
- Adversarial review noted the ESM-root behavioral test covered only cursor;
  refactor it into a helper and add a codex case (the !isCodex-gated path most
  likely to regress, whose marker write lives in bin/install.js). gsd-check-update.js
  require()s at module load, so it surfaces the ESM failure immediately.
- emitted-provenance flagged hooks/package.json as 'attributed source does not
  exist' — the marker is code-derived (a fixed literal emitted by
  ensureCommonJsMarker at install time), not built from a tracked source. Route
  the hooks-built rule's sources/transforms for package.json to the surface
  source file, mirroring the existing .cmd-shim sub-family.

* chore(#2717): drop now-redundant hooks/package.json attribution ack

The hooks-built provenance routing (prior commit) now self-attributes the
emitted hooks/package.json to src/runtime-hooks-surface.cts, which IS in this
diff — so the attribution is self-explaining and the emitted-drift-ack entry
became stale. Delete the (now-empty) ack file per ADR-2719's empty-file rule.

* docs(changeset): backfill #2717 PR number to 2846
2026-07-29 22:18:48 -04:00
Tom Boucher
b12d4df03b fix(#2694): normalize CRLF before frontmatter-boundary match in code-review workflows (#2839)
* test(#2694): CRLF frontmatter boundary regression for code-review workflows

The code-review / code-review-fix workflows embed inline node -e one-liners
whose frontmatter boundary regex used a literal \n, silently returning null
on CRLF-saved SUMMARY.md / REVIEW.md / REVIEW-FIX.md artifacts and dropping
every file in that summary (acceptance: per-artifact, no warning when the
phase aggregate stays non-zero).

Adds:
- behavioral CRLF==LF boundary extraction tests (replica of the shipped
  one-liner's boundary step), proving the buggy literal-\n returns null on
  CRLF while the fixed normalize-then-match yields a byte-identical body;
- a structural-regression-guard (allow-test-rule: structural-regression-guard)
  that reads the two shipped workflow files and asserts every boundary site
  normalizes \r\n -> \n before matching, so a revert of the fix is caught.

* fix(#2694): normalize CRLF before frontmatter-boundary match in code-review workflows

The code-review and code-review-fix workflows embed nine inline node -e
one-liners that extract YAML frontmatter via a boundary regex
  content.match(/^---\n([\s\S]*?)\n---/)
The literal \n defeated any CRLF-saved artifact (\r between --- and the line
terminator), so SUMMARY.md / REVIEW.md / REVIEW-FIX.md saved with CRLF
endings silently contributed zero files (or 'unknown' status / 'invalid')
with no per-artifact warning. The Tier-3 git-diff fallback only fires when
the aggregate across all summaries is zero, so a single CRLF summary among
LF summaries produced no signal at all.

Normalize \r\n -> \n once before the existing boundary match at all nine
sites (code-review.md x3, code-review-fix.md x6). Byte-identical to the LF
path; mirrors the canonical src/frontmatter.cts extractFrontmatter intent
(CRLF == LF at the boundary); zero risk of \r leaking into field values
consumed by the inner JS or the shell grep/cut pipeline.

RED @ 94d0213 (3 failures, structural guard caught the shipped-text bug,
both linux-node22+24 lanes).GREEN pending.

* chore(#2694): acknowledge code-review workflow growth + changeset fragment

emitted-attribution (ADR-2719) reports the byte growth from the CRLF-normalize
insertion in code-review.md (+69) and code-review-fix.md (+138); both are the
intended #2694 fix. Adds the .changeset Fixed fragment (pr:0, backfilled post-PR).

* test(#2694): mixed CRLF/LF phase yields the union of both artifacts (criterion 2)

The spec-axis review flagged that acceptance criterion 2 (a phase with a mix
of CRLF-affected and unaffected artifacts no longer silently drops the CRLF
artifact's contribution) was only transitively satisfied. Adds an explicit
mixed-phase test replicating the full shipped Tier-2 extractor (boundary +
inner key_files parse) across one LF and one CRLF SUMMARY.md, asserting the
union of both — plus a RED proof showing the buggy boundary drops the CRLF
artifact silently (aggregate non-zero, so the Tier-3 eq-zero fallback never
fired). Locks the silent-partial-masking behavior the triage named as the more
serious half of the defect.

* docs(changeset): backfill #2694 PR number to 2839

* fix(#2694): make the CRLF regression test itself CRLF-lint-clean

CI lint-tests caught that the new test tripped local/no-crlf-fragile-split:
- the frontmatter boundary regex replicas (fixed + buggy) were RegExpLiterals
  with a bare \n; the rule flags frontmatter-shape regexes unconditionally.
  Build them via new RegExp(...) (byte-identical .source to the shipped literal)
  so the faithful replica is not a lint violation — the buggy replica MUST keep
  the literal \n, that is the bug it demonstrates.
- the structural guard's src.split('\n') on the readFileSync'd workflow file
  was genuinely CRLF-fragile; use /\r?\n/ per the rule's canonical fix.
- the allow-test-rule annotation gains its (#2694) tracking ref per ADR-456.

lint:ci now exit 0 (incl. lint-allow-test-rule-refs, lint-emitted-drift-ack,
lint-fix-has-regression-test: PASS).
2026-07-29 19:49:53 -04:00
Tom Boucher
1fc21cdee0 fix(#2716): route non-user-facing conventional types to an Internal bucket, omit from release notes (#2838)
* test(#2716): failing-first regression for non-user-facing types → Internal bucket

* fix(#2716): route non-user-facing conventional types to an Internal bucket, omit from release notes

* fix(#2716): update SAMPLE_BODY/Discord/property tests for Internal bucket; fix stale CONTRIBUTING sentence (review)

* test(#2716): relax Discord Enhancement assertion (enhance: prefix not stripped by cleanBullet)

* docs(changeset): #2716 non-user-facing types omitted from release notes

* docs(changeset): backfill #2716 PR number to 2838
2026-07-29 17:06:04 -04:00
Tom Boucher
6a9babda69 chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data

Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3.

- Five reviewers GSD never installs into become lane-only role:reviewer
  capabilities with no runtime body, no runtimeCompat and no install surface:
  gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no
  descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail,
  which is now deleted outright.
- The six hosts that are ALSO reviewers gain a reviewer body alongside their
  runtime body. Their runtime bodies are byte-identical to next -- verified per
  capability against the git blob, not asserted -- so no install behaviour moves.
- KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported
  deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived
  legacy alias for one release; where a capability carries both, the body wins
  and the slug appears once. Alias removal is Phase 7 (#2801).

THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity,
claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama,
opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it,
and the test asserts that literal list rather than a count.

kimi-code is deliberately NOT declared here. It is net-new with no
invoke_reviewers leg, so declaring it now would make it selectable but not
invocable -- present in --all, selected, emitting an empty section for the whole
5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b
alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a
reviewer at all and gains nothing.

The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it
deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field,
including probe and invoke sub-fields. All eleven are byte-identical, key order
included. The epic's premise is that the manifest and the core descriptor
describe the same lane with NO translation layer, and Phase 2's review already
caught one divergence that every other test missed.

Two ADR corrections folded in, as Phases 1-3 each did:

1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798
   claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable:
   D9 assigns review.<host>_host to lane capabilities that do not exist until
   THIS phase creates them, and a federated config slice must live inside
   capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4.
2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs
   bin/lib/*.cjs modules, not capability directories -- antigravity, opencode
   and qwen appear zero times in it -- and gen-inventory-manifest --check passes
   with the five new dirs and no edit.

Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug
pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the
shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported
LANE_SLUG_RE. The prose had not followed the code.

Closes #2798

* fix(#2798): catalogue reviewer capabilities in the generated matrix

The capability matrix rendered exactly two tables, feature and runtime, via
renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD
role, so every role:"reviewer" capability was silently dropped from the
first-party catalogue.

The drift guard did not catch it, and could not: --check compares generated
output against the committed file, and both omitted the five lanes identically,
so it reported "up to date" while five shipped capabilities were invisible in
the one document that is supposed to list what ships. A guard blind to an entire
role is not guarding.

This phase is what exposed it -- it ships the first role:"reviewer"
capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect
found while working is fixed in the current change, which overrides
one-concern-per-PR).

Verified red-before-green: with a lane row deleted from the matrix, --check now
exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so
there was nothing for the guard to compare.

Phase 6 (#2800) still owns enriching the matrix with lane-specific detail
(slug/flag/transport columns) and the locale parity gate. This is the narrower
fix: the capabilities APPEAR at all.

* fix(#2798): close two hardening gaps and record three limits durably

Isolated security review (5 targets, no blockers) reproduced two gaps in the new
deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is
generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for
reuse and carries no other validation, so it must not depend on its caller.

- A whitespace-only slug passed the length>0 test verbatim and occupied a roster
  entry it could never match. Slugs are now trimmed before the emptiness test. A
  blank body correctly falls through to the legacy alias rather than DROPPING the
  lane, which would have been worse than the blank slug.
- KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there
  breaks import for EVERY consumer rather than degrading selection. It is now
  guarded, yielding an empty roster on a malformed registry. That is a visible
  degradation, not a silent one: under D4 an explicitly requested reviewer that
  is unavailable is an ERROR, so /gsd:review --claude against an empty roster
  fails loudly. This also removes an asymmetry -- the sibling capability-trust
  module documents its collectors as TOTAL and wraps them for exactly this reason.

Also records three findings that previously existed ONLY in squash-merged PR
bodies, which is not a durable record:

- ADR-2782 D5 gains an implementation note explaining why the resolved host is
  deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds
  the resolved host; the loader has no config resolver, so folding it in would
  make the loader and lifecycle compute different signatures for one manifest and
  re-prompt forever. The binding is split: signature covers the SHA-pinned
  manifest fields, the consent record stores the resolved host, and Phase 5b
  re-resolves at invocation -- which is where rule 4 already puts the check. A
  reader comparing rule 1 to the code would otherwise conclude it is unimplemented.
- CONTEXT.md's capability-trust entry still described THREE executable surfaces.
  Phase 3 added the fourth and made that false; corrected here, since it is drift
  this epic introduced rather than Phase 6's new-glossary-term work.
- stableJson documents the NaN/Infinity/undefined -> null signature collision and
  why it is unreachable (JSON grammar has no such literal, so JSON.parse throws
  first). Reachability rests entirely on the ingest path staying JSON.parse-only,
  so the note lives where someone would break it.

* chore(#2798): backfill changeset pr number to 2837
2026-07-29 16:58:02 -04:00