Commit Graph

4881 Commits

Author SHA1 Message Date
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
JusticeWay
7b204ad2ac enhance(#2530): extend UAT checkpoint frame language pack (9 more languages) (#2564)
* feat: extend UAT checkpoint frame language pack (9 more languages)

response_language is a free-form config value, but CHECKPOINT_FRAMES only
covered 9 languages — any other configured language silently fell back to
the English frame. Add Dutch, Polish, Russian, Ukrainian, Turkish, Hindi,
Arabic, Vietnamese, and Indonesian frames plus their aliases, with a
regression test asserting each resolves instead of falling back.

Follow-up to #2402 (PR #2457).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: add changeset for #2527

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2530): list UAT checkpoint frame languages in CONFIGURATION.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2530): point changeset fragment at PR #2557

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2530): address Unicode language-pack review

* fix: address checkpoint language review

* fix: count spacing combining marks in checkpoint width

* test: verify checkpoint aliases structurally

* fix: isolate RTL checkpoint frames

* fix: isolate RTL checkpoint frames correctly

* test(#2530): assert checkpoint aliases neither collide nor go unreachable

Review Minor #1. A duplicate alias key was invisible to the existing
catalog tests: the runtime object is well-formed after JS collapses the
literal, the self-alias assertion still holds, and the losing language
just stops resolving. tsc catches the byte-equal case (TS1117), but not
the two that survive compilation — an alias whose NFC-lowercase form
already belongs to another language, and an alias not in lookup form at
all, which resolveCheckpointFrame() can never produce.

The check reads the source literal rather than the object, since the
object no longer records what was written. Both assertions are
independently load-bearing: an NFD twin of an existing alias trips the
collision check, an uppercase alias trips the unreachability check.

Review Minor #2: changeset retyped Changed -> Added. Nine wholly new
supported response_language values are an addition under Keep a
Changelog, not a modification of existing behavior.

* test(#2530): check alias collisions on the catalog, not its source

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:01:50 -04:00
Tom Boucher
f092c6da85 fix(#2649): diagnose-issues + execute-plan run worktree.base-check before worktree dispatch (#2955)
* test(#2649): failing-first — diagnose-issues + execute-plan must run base-check before worktree dispatch

* fix(#2649): diagnose-issues + execute-plan run worktree.base-check before dispatch

diagnose-issues.md spawn_agents and execute-plan.md Pattern A spawned
worktree-isolated subagents (gsd-debugger / gsd-executor) without the
pre-dispatch worktree.base-check gate that execute-phase (#683/#1369) and
quick (#1941) already run. Claude Code's isolation="worktree" forks from
origin/HEAD, not live local HEAD; without the gate, the documented GSD steady
state (commit every step locally, push only on request) hits the verify-only
worktree_branch_check guard's exit-42 halt mid-investigation with no auto-degrade.

Mirror the quick.md #1941 pattern: before dispatch, run
`gsd_run query worktree.base-check --pick shouldDegrade`; if true, print its
message + a #2649 warning to stderr and set USE_WORKTREES=false (sequential
main-tree dispatch). The verify-only guard stays as a backstop in both cases.

Per the triage and #2649 acceptance criterion 5, execute-plan.md's Pattern A
(identified as a second site with the identical gap) is fixed in the SAME change
— same bug class, same one-line gate, two workflow files — rather than filed as
a separate follow-up.

* fix(#2649): ack the diagnose-issues + execute-plan growth (per-PR fragment)

The two workflow files grew vs next (diagnose-issues.md +1381, execute-plan.md
+905) adding the #2649 base-check gate. emitted-attribution requires an ack;
this is a per-PR fragment under tests/emitted-drift-acks/ (#2914 mechanism,
replacing the legacy shared emitted-drift-ack.json).

* test(#2649): tighten base-check ordering assertion + guard backstop survival

Address code-review minors:
- the ordering assertion was a loose disjunction that passed even if the
  base-check moved AFTER the dispatch; tighten to assert base-check < Agent()
  (the real invariant).
- add a test that the verify-only <worktree_branch_check> backstop remains
  embedded in the Agent() prompt (acceptance criterion 4 — the base-check is a
  pre-dispatch degrade, the guard is a post-fork fail-closed backstop; both
  layers must survive).

* changeset(#2649): diagnose-issues + execute-plan auto-degrade on stale worktree base

* changeset(#2649): backfill PR number 2955

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 19:15:04 -04:00
Tom Boucher
388837219d fix(#2648): phase.complete refuses when non-retired plans lack summaries (fail-closed coverage gate) (#2953)
* test(#2648): failing-first — phase complete must refuse when plans lack summaries

Adds findUnsummarizedPlans to core-utils (mirrors countMatchedSummaries but
returns the unmatched plan files) and a PHASE_PLAN_COVERAGE_INCOMPLETE error
reason, plus a 3-case regression block in phase.test.cjs. The gate itself is
NOT yet wired into cmdPhaseComplete (reverted for the RED run), so the
'blocks completion' and 'superseded does not block' cases must FAIL (the
pre-fix code completes silently).

* fix(#2648): phase.complete refuses when non-retired plans lack summaries

cmdPhaseComplete gated only on a single *-VERIFICATION.md status, so a phase
could close complete while an arbitrary number of its plans had no completion
record (confirmed incident: 6/30 plans unexecuted incl. the phase's entire
final UI scope, every signal green). Add a fail-closed plan-coverage gate that
refuses completion when any plan lacks a matching *-SUMMARY.md, naming the
missing plans, UNLESS the plan is retired via machine-readable status:
superseded frontmatter (#2349) — closing the Goodhart hole (delete a SUMMARY to
raise the %) without regressing the lock/recovery pattern.

Uses scanPhasePlans (superseded-AWARE) + new findUnsummarizedPlans helper so the
gate, the count, and the named list can never disagree. Evaluated before the
verification-gate transaction so a refusal fails fast without mutating
ROADMAP/STATE. milestone.complete's parallel gap is explicitly out of scope
(separate seam, separate PR).

* fix(#2648): test fixtures — give #1752 phase plans summaries + STATE.md in coverage fixture

The plan-coverage gate (#2648) correctly blocks phase completion when a plan
lacks a SUMMARY. Two test fixtures needed updating to reflect the new contract:
- #1752 (total_phases-decrement cascade): its 8 phase dirs each had a PLAN.md
  with no SUMMARY. The test's concern is the total_phases cascade, not plan
  coverage, so add a matching SUMMARY to each to keep the phase fully-covered
  and isolate the #1752 behavior.
- the #2648 coverage-gate fixture: write STATE.md (createTempProject scaffolds
  .planning/phases but not STATE.md) so the 'ROADMAP/STATE unchanged on refusal'
  assertions have a file to read.

* fix(#2648): security — fail closed on unreadable plan dir + sanitize msg + surface superseded

Address the security-review blocker (B1) and hardening (M1/m2):
- B1 (blocker): the gate failed OPEN when scanPhasePlans could not read the
  phase dir (it swallows readdirSync errors → empty plan set → gate sees zero
  unsummarized plans → passes). A coverage gate that passes when it cannot read
  the plans re-opens the #2648 hole under any I/O failure. Now readdirSync the
  dir explicitly and fail closed (PHASE_PLAN_COVERAGE_INCOMPLETE) on a throw;
  a readable empty dir still passes (legitimately complete empty phase).
- m2: sanitize plan filenames (strip C0 controls / DEL) before interpolating
  into the error message — they come raw from readdirSync and could spoof the
  terminal in plain-error mode.
- M1: surface the count of plans excluded as status: superseded so a reviewer
  can audit which work was declared retired (the marker is a committable,
  review-time-trusted bypass; keep it visible).
- Add a 4th regression case: unreadable plan dir (ENOTDIR via a file, not chmod
  0o000 which root bypasses) must fail closed.

* test(#2648): drop unreachable B1 case — no root-safe unreadable-dir repro

The B1 fail-closed-on-unreadable-dir defense stays in src/phase.cts (cheap +
correct), but it cannot be unit-tested cross-platform: any condition that makes
the phase dir unreadable to the gate's readdirSync ALSO fails findPhaseInternal
upstream ('Phase N not found') before the gate runs, and chmod 0o000 is
forbidden (root bypasses it in root CI). Document the gap in the test file;
remove the case that asserted a reason the upstream error pre-empts.

* style(#2648): drop unnecessary type assertions flagged by lint:ci

scanPhasePlans returns typed string[] arrays, so the `as string[]` casts on
coverageScan.planFiles/summaryFiles were redundant (@typescript-eslint/
no-unnecessary-type-assertion). Compute supersededCount from typed lengths;
only the phaseInfo['plans'] cast remains (it is genuinely unknown).

* changeset(#2648): phase.complete refuses when plans lack summaries

* changeset(#2648): backfill PR number 2953

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 18:11:40 -04:00
Tom Boucher
5a0a9f0972 fix(#2944): remove the catastrophic-backtracking regex from the ADR-1671 example (#2950)
* fix(#2944): remove the catastrophic-backtracking regex from the example

The non-shipping Option-E reference example carried its own copy of the
predicate-id regex, which nested a dot-containing character class inside a
dot-prefixed repeat. A run of N consecutive dots therefore had exponentially
many partitions. Measured on next before this change: 30 dots 54ms, 35 66ms,
40 807ms — so roughly 55-60 dots hangs for hours.

Not exploitable where it sits: the example is outside tsconfig.build.json,
outside the npm package files list, outside the installer and outside tests,
so no build step or CI job parses anything with it. Fixed because the entire
point of a reference example is that people copy it forward, and ADR-1671
presents this one as the pattern for the platform.

Ports the linear per-segment validation that #2928 gave the production module,
so the two copies agree: both parse the real CONTEXT.md to 415 predicates
across 20 classes with 0 duplicates. Doubled-dot ids are now rejected here
too, matching production, and the grammar comment records it.

Also refreshes the example's committed index, which #2928 made stale when it
removed the duplicate predicate from CONTEXT.md.

Closes #2944

* test(#2944): guard predicate-index sync and example/production parity

Two regression tests for the two defects in this PR.

Index sync: asserts the committed docs/CONTEXT-INDEX.json equals a fresh parse
of CONTEXT.md, naming any diverging predicate ids. The merge race that reddened
next was invisible to both PRs involved and only surfaced on the next PR to run
lint:ci; this puts the same check inside the suite, which runs on every PR, and
a mutation test proves the assertion is not vacuous.

Example/production parity: asserts both copies of the parser report the same
count, classes and duplicates for the real CONTEXT.md, and agree verdict-for-
verdict over a table of id shapes. The divergence WAS the bug — production went
linear-time while the example kept the backtracking regex, with nothing
asserting they agreed. Also pins the example rejecting a 60-dot id, with the
clean rejection as the binding assertion and wall-clock only as a smoke check.

Notes a real tension rather than hiding it: ADR-1671 says the example sits
outside tests/, and this imports it. The ADR's intent is that the example is
not compiled, packaged or installed — not that it may silently rot. A parity
guard does not ship it. The file states this so a reviewer can object.

* fix(#2944): address both isolated review passes

Two independent reviewers (correctness and security axes, neither the author).
Security found nothing — it measured linearity to 100k chars across dots,
hyphens, underscores and mixed classes, and showed prototype pollution is
structurally unreachable because the first-segment pattern forbids
lowercase and underscore-leading ids. The correctness pass found three
blockers, all real.

Blocker: the parity test violated ADR-1671 verbatim. The ADR lists FOUR
exclusions for the reference example, the fourth being the CI test suite, and
the test imported it from tests/ while its own justification comment cited only
three -- constructing a rationale around the exclusion it broke. Moved to
scripts/lint-example-parser-parity.cjs wired into lint:ci; a lint script is not
the test suite, so the exclusion stands. The test file keeps only the
docs/CONTEXT-INDEX.json sync check.

Blocker: the mutation test leaked its temp dir. Its callback took no `t`, so a
failing assertion skipped the bare cleanup call. Now registered via t.after(),
matching the convention adr-index-gate.test.cjs documents.

Blocker: the example's own committed index carries the identical merge-race
staleness this PR fixes for the production one, and nothing guarded it.
Deliberately NOT fixed by wiring the example's --check into CI: that artifact
bakes line numbers, so it re-drifts on any unrelated CONTEXT.md line shift --
exactly ADR-1671 open question 4 -- and would make CI routinely red. The new
lint asserts the line-INDEPENDENT facts instead: count, class map, duplicate
set, and every (id, value) pair. Proven non-vacuous both ways: mutating a value
fails and names the id, mutating only a line number passes.

Major: a real divergence the parity claim would have missed. Production rejects
values containing an embedded CR, LF, U+2028 or U+2029; the example did not, so
a value with an embedded lone CR was rejected by one copy and accepted by the
other. Ported, and now covered by the parity table.

Also, found while verifying rather than reported: malformed diagnostics covered
only empty values. A doubled-dot id, a space in an id, and a lowercase-leading
id were all dropped silently. That contradicts the module's own intent -- a
typo should be diagnosable, and a space in an id is a likely one -- and
predicates are contractually cited, so a silently vanished predicate is the
failure mode that matters. Each rejection class now carries a named reason in
both copies, while ordinary inline code still yields none.

Trues up counts my own change staled: the example README and ADR-1671's
prototype figures said 416 and 393/18 against a real 415/20/0.

Closes #2944

* chore(#2944): backfill changeset PR number 2950

---------

Co-authored-by: sim <sim@local>
2026-07-31 15:44:19 -04:00
Tom Boucher
07603df8f2 fix(#2647): code-fixer worktree under .claude/worktrees/, not a hardcoded /tmp path (#2942)
* test(#2647): failing-first — fixer worktree path must be repo-relative not /tmp

* fix(#2647): place code-fixer worktree under .claude/worktrees/, not /tmp

The gsd-code-fixer agent hand-rolled its worktree at a hardcoded
`/tmp/sv-${padded_phase}-reviewfix-XXXXXX` mktemp path. On Windows/Git Bash
that landed OUTSIDE the project tree — outside the agent session's permission
allowlist, so every Read inside the worktree prompted (~25/run) — and mktemp's
MAX_PATH-avoidance substitute produced an un-removable `C:/mvwtNN` path.

Place the worktree repo-relative under `.claude/worktrees/` (the same dir the
harness-managed executor worktrees use: gitignored via `.claude/`, inside the
session's permission scope), with a $$-PID + epoch suffix for concurrency
uniqueness (replacing mktemp's XXXXXX). $main_repo is resolved the same way
the cleanup tail already resolves it.

Three sites updated: setup_worktree bash, concrete-steps prose, critical_rules.
The #2990 `-b "$reviewfix_branch"` invariant is preserved (the folded test
asserts it). Failing-first regression added to the #2990 suite in
tests/agent-frontmatter.test.cjs.

* test(#2647): update #2686 path assertion to expect .claude/worktrees/, not /tmp

The #2686 regression test encoded the worktree location as a hardcoded
`/tmp/sv-` path (matching sibling GSD agents at the time). #2647 showed that
breaks Windows/Git Bash (worktree outside the project tree → permission prompts;
mktemp MAX_PATH substitute un-removable). Update the #2686 path assertion to
require the repo-relative `.claude/worktrees/` location and forbid `/tmp/sv-`.
The #2686 isolation + cleanup assertions are unchanged.

* fix(#2647): word-boundary wt= parse + ack the fixer growth vs next

Two follow-ups to the #2647 GREEN run:
- parseWtAssignments matched `prior_wt=` (no word boundary), polluting the
  set and tripping the repo-relative + concurrency-unique assertions. Anchor
  on (?:^|\s)wt= so only the real worktree-path assignment is captured.
- emitted-attribution: gsd-code-fixer.md grew 1875 bytes vs origin/next. Update
  the emitted-drift-ack entry to attribute the #2647 worktree-path change
  (supersedes the prior #2825 attribution, whose growth is already in next).

* fix(#2647): address review — validate padded_phase at the sink + tighten test

Code-review + security-review both APPROVED with one actionable minor:
padded_phase is interpolated into a worktree PATH and a git BRANCH NAME, but
was only validated by the orchestrator (code-review-fix.md), not at the agent
sink. The agent prompt is a literal bash contract any caller can spawn, so add
a `[[ =~ ^[0-9]+(\.[0-9]+)?$ ]]` self-defense check rejecting traversal/shell
metachars (defense-in-depth; not a present vuln — the only caller validates).

Also tighten the concurrency-uniqueness test to require BOTH $$ AND $(date +%s)
(either-alone was too lax per review). Update the emitted-drift-ack reason to
cover the added validation growth.

* changeset(#2647): code-fixer worktree under .claude/worktrees not /tmp

* changeset(#2647): backfill PR number 2942

* chore(#2938): regenerate stale docs/CONTEXT-INDEX.json on next

#2938 (#2928) updated the CONTEXT.md RULESET prose for the new per-PR
emitted-drift-ack fragment mechanism (#2914) but shipped a CONTEXT-INDEX.json
generated from the OLD prose. lint:generated-sync fails on every PR that
rebases onto next after #2938 (the regen produces a 3-line diff bringing three
RULESET entries — AGENT_SIZE_BUDGET, EMITTED_ATTRIBUTION, WORKFLOW_SIZE_BUDGET
— in sync with the prose already on next). Mechanical regen via
`node scripts/gen-context-index.cjs --write`; idempotent; surfaced by the
#2647 rebase. No behavioral change.

---------

Co-authored-by: sim <sim@users.noreply.github.com>
2026-07-31 14:45:51 -04:00
Tom Boucher
05b170e448 chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam

Productionizes the ADR-1671 Option-E reference example as a real module:
src/context-predicates.cts (parser + selector + index builder) compiled to
gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's
--check/--write drift-guard idiom and wired into lint:generated-sync.

Parser behavior is deliberately prototype-equivalent in this commit so the
next commit's regression matrix binds to the real defects rather than to a
missing module.

Two locked design deviations from the prototype:
- duplicates carry a count, not line numbers
- the committed index carries no line field at all, resolving ADR-1671 open
  question 4: an artifact without line numbers cannot drift on a line shift,
  so promoting --check to a CI gate does not make it routinely red

Also reconciles the one remaining duplicate predicate ID
(RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording
is removed) so the gate can land fail-closed on duplicates.

Refs #1671

* test(#2928): failing-first matrix for the predicate fact-store

Adds the regression matrix from the phase test plan: parser declaration
forms, fence and comment regions, ID/value grammar boundaries at
limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard
CLI, the selector query surface, and four document-shaped fast-check
properties.

Seven rows are RED for behavioral reasons against the ported parser:
indented-bare, star-list, plus-list and numbered-list declaration forms are
dropped; a tilde fence and a four-backtick fence containing a shorter fence
are not skipped; and a multi-line HTML comment is parsed as live. Eleven
selector rows are RED because the query surface is not wired yet.

Negative fixtures come from real repo documents that predate the grammar
(CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the
fixture-provenance rule, and the property generators are document-shaped
rather than seeded from our own serializer.

Refs #1671

* fix(#2928): consume the shared fence scanner, relocate the index, wire the selector

Drives the failing-first matrix green.

Parser: replaces the ported naive triple-backtick toggle with the shared
markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord
gain an export keyword — the only change to that module, which has 71
upstream dependents — because it already returns line-indexed spans, which
is exactly what a line-reporting parser needs. It also already documents
itself as the second copy of the fence state machine pending consolidation;
adding a third copy here would have been the generative-fix divergence this
repo warns about. A parity suite now pins predicate fence-skipping against
that scanner across eight fence shapes. HTML-comment skipping stays local
because the sectionizer has no comment scanner. Declaration forms widen to
indented-bare, star, plus and numbered list items.

Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The
remote matrix run caught the original choice — a committed .cjs there ships
~120KB of CONTEXT.md prose into a runtime module, and two content guards
fired truthfully on it (a leaked .claude install path, and four hardcoded
package-name literals). Neither guard was allowlisted; the artifact moved
instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to
require it — it is a drift-detection artifact, so the selector parses
CONTEXT.md live and is always current.

Generator: adds a frozen REASON enum and --check --json so the gate's
outcome is asserted structurally instead of by matching prose, and
--context-path/--index-path so tests drive the real CLI against a temp tree
with no filesystem monkeypatching.

Selector: gsd_run query context-predicates with --class/--prefix/--contains,
structured output carrying a matched count, own-property guards, and no
project-root resolution. Registering it exposed that the query dispatch
table and the usage string had drifted: a new parity test found 20 routed
commands missing from the usage list, all added here rather than deferred.

Refs #1671

* test(#2928): lock the newly-public scanFencedBlocks contract

Exporting scanFencedBlocks made it public API for the first time, so it
needs its own contract test independent of the consumer that motivated the
export. Memtrace's co-change analysis flagged the gap: this suite changes
together with markdown-sectionizer.cts 8 times in 90 days and was absent
from the diff.

Covers the documented rules: 0-based indices, -1 for an unterminated fence,
the same-char/>=length/no-trailing-text closer rule, a shorter fence inside
a longer one staying content, CommonMark 4.5 backtick-in-info-string, and
<=3-space indent tolerance.

Refs #1671

* fix(#2928): address both isolated review passes

Two independent reviewers (correctness axis and security axis, neither the
author) found seven findings. All are fixed here with regression tests; none
deferred.

BLOCKER — comment-blind fence scanning caused silent, permanent predicate
loss. The HTML-comment scan and the fence scan ran as two independent passes,
and the fence scanner is comment-blind, so a fence delimiter inside an HTML
comment with no later close read as an unterminated fence and skipped every
remaining line to EOF. Worse, the drift-guard could not catch it: it diffs
against a baseline produced by the same corrupted parse. The two constructs
now interleave in a single pass so each suppresses the other's boundary
detection while active, covered in both directions. The parity suite still
binds this scanner to markdown-sectionizer's for comment-free documents, so
the two cannot diverge unnoticed.

BLOCKER — the selector was not consumed anywhere, leaving the phase's
acceptance criterion unmet. Now wired into the pre-work predicate-citation
step in contributor-standards, which is the repo's actual brief-assembly
path; no code-level brief assembler exists to wire into.

MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id
regex nested a dot-containing character class inside a dot-prefixed repeat,
so N consecutive dots had exponentially many partitions: 40 dots took 565ms
and growth was exponential. CI runs this parser over a pull request's own
CONTEXT.md, so any contributor could have hung a shared runner with one
line. Replaced with linear per-segment validation. Doubled-dot ids are now
rejected; the real document contains none.

MAJOR — the duplicate-id gate had only ever been proven on synthetic
fixtures. A test now re-inserts the exact line this branch removed and
asserts the real generator names it.

MAJOR — --check together with --write silently let write win, turning the
gate into a writer; a missing path value resolved to the cwd and leaked an
EISDIR stack trace. Both are now clean usage errors.

MINOR — the hoisted skip-list was exported as a live mutable Set; replaced
with a read-only predicate. MINOR — flag-shaped selector values were
unmatchable; the inline --flag=value form now provides the escape hatch.

Refs #1671

* chore(#2928): backfill changeset PR number 2938

---------

Co-authored-by: sim <sim@local>
2026-07-31 13:17:01 -04:00
Tom Boucher
c043f2946c fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next

tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900.
Every entry is scoped to the diff that introduced it (#2789), so once merged
to next it is at the base by definition -- spent and inert. Its presence is
still load-bearing though: each PR rewrites the paths map wholesale, making a
persistent base copy a shared cell. Five of six conflicting PRs in the open
queue collided on this file and nothing else.

Deletes the stale document and adds a push-to-next guard asserting it stays
absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check
against the base is the #2768 shape #2789 exists to end.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2914): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2914): per-PR ack fragments instead of one shared mutable file

The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json
whose paths map every PR rewrote wholesale. That is a shared mutable cell: any
two PRs needing an ack edit the same lines and conflict. Five of six conflicting
PRs in the open queue collided on this file and nothing else.

Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same
shape .changeset/ already uses to solve this exact problem. Two PRs pick
different filenames, so they cannot collide, and fragments lingering on next
are harmless rather than toxic.

The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An
earlier delete-only attempt failed verification twice: the ratchet lost the
spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no
ack. Relocating preserves every acknowledgment.

The legacy single file is still READ (unioned with the fragments) because five
open PRs carry it; dropping support would break all of them. A duplicate path
key across sources is a hard error, never last-wins.

The push-to-next guard is retargeted accordingly: it now asserts only that the
legacy SHARED file never reappears on next. Fragments may persist harmlessly.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:15:29 -04:00
Tom Boucher
d2d2f7c088 fix(#2848): non-Latin titles no longer produce empty slugs (Cyrillic transliteration) (#2934)
* test(#2848): add failing-first regression for non-Latin slug transliteration

generateSlugInternal and slugify both strip non-ASCII chars with no
transliteration step, so an all-Cyrillic title reduces to an empty slug.
12-row matrix: Cyrillic regression (both impls), Latin negative control,
multi-letter mappings, soft/hard sign drops, Ukrainian extras, null
contract, CJK unaffected, mixed scripts, truncation parity, slugify's
distinct no-truncate contract.

* fix(#2848): transliterate Cyrillic titles to ASCII before slug strip

Both generateSlugInternal (src/core-utils.cts) and slugify
(src/gsd2-import.cts) stripped non-ASCII with no transliteration, so an
all-Cyrillic title reduced to an empty slug, producing unnamed phase
directories (01-) and empty milestone_slug init JSON.

Add a shared transliterateForSlug primitive (core-utils) covering Russian
+ the reported Ukrainian/Belarusian extras (і ї є ґ ў), with multi-letter
mappings (ж→zh ч→ch ш→sh щ→sch ю→yu я→ya) and dropped soft/hard signs
(ъ ь). It runs BEFORE the existing ASCII filter, so Latin-script text
hits zero map entries and is byte-for-byte unchanged (negative control).
slugify consumes the shared primitive, preserving its distinct single
hyphen-strip + no-truncation contract. CJK/unmapped scripts keep the
existing strip-to-ASCII behavior.

Also corrects two test assertions to match the chosen й→y mapping and the
б→b (not bie) transliteration.

* changeset(#2848): Fixed — non-Latin slug transliteration

* changeset(#2848): backfill PR number 2934

---------

Co-authored-by: sim <sim@local>
2026-07-31 10:56:11 -04:00
Tom Boucher
81eeb8a53a docs(#2926): refresh ADR-1671 with findings re-verified on next (#2936)
Re-measured the Option-E prototype's reported figures against CONTEXT.md on
next (2026-07-31) and recorded the delta rather than overwriting the June
numbers:

- index counts 393/18 (2026-06-24) -> 416/20 today; CONTEXT.md gained the
  PROBE (11) and PROHIB (10) classes
- of the 3 duplicate predicate IDs, only RULESET.WORKFLOW_MARKDOWN.FENCES
  remains; the two RULESET.GEMINI.* went with the Gemini runtime removal
- gen-context-index.cjs --check exits 1 on next, so Phase 0's "--check green
  in CI" criterion is unmet (invisible to CI: the example sits outside tests/)

Adds Open question 4 (index keyed on baked line numbers re-drifts on any
CONTEXT.md line shift, which matters once Phase 1 promotes --check to a CI
gate), and records the reviewer-proposed eval-gate question as resolved by
the PROBE.*/PROHIB.* predicate classes (ADR-550 D4/D7, ADR-1606).

Also names both surfaces of the Windsurf 12 KB throw in Decision 2, since it
is duplicated byte-identically in bin/install.js and
src/runtime-artifact-conversion.cts.

Docs-only. No code, no runtime-loaded text, no behavior change.

Co-authored-by: sim <sim@local>
2026-07-31 10:52:52 -04:00
Rezolv
76b7d73039 fix(#2733): route gate-passed spec-phase paths into the probe steps (#2779)
* fix(#2733): route gate-passed spec-phase paths into the probe steps

All four gate-passed transitions in spec-phase.md said "Jump to Step 6",
textually bypassing the mandatory Step 5.5 edge-completeness and Step 5.6
prohibition-completeness probes. Steps 5.5/5.6 were spliced between Step 5
and Step 6 by two later feature commits and the pre-existing jumps were
never re-pointed, so no jump instruction in the file reached Step 5.5 at
all and both probes were unreachable dead prose.

Re-point the four gate-passed jumps (lines 129, 162, 168, 170) to Step 5.5.
Control then flows 5.5 -> 5.6 -> 6 as the probes' own preconditions
prescribe. The max-rounds "write anyway" bypasses and the probes' own
"proceed to Step 6" exits are deliberately unchanged.

Add tests/spec-phase-probe-reachability.test.cjs, which derives the
mandatory probe steps from the file's own headings rather than hardcoding
5.5/5.6, so a future spliced-in probe step is covered without editing the
test. It also locks the two coupled constraints: the max-rounds bypass must
not be redirected into a probe, and each probe must keep its own onward exit.

The existing probe contract tests are untouched and still pass; both scope
from the "## Step 5.5"/"## Step 5.6" heading onward and were structurally
incapable of observing the upstream jump text.

* chore(changeset): Fixed fragment for #2779 (spec-phase probe reachability)

* fix(#2733): route Step 5.5's own soft gate into Step 5.6

Round-1 review blocker. The four upstream gate-passed jumps were re-pointed to
Step 5.5, but Step 5.5's own terminal soft gate at :305 still read "proceed to
Step 6" - so the COMMON path (all applicable edges resolved) skipped the
prohibition-completeness probe outright. Same defect class as the four this PR
already fixed, on the success path of the very step being fixed: the SPEC shipped
with an empty Prohibitions section instead of an empty Edge Coverage one.

Its sibling at :393 is byte-identical yet correct, because Step 6 genuinely
follows Step 5.6. Position, not phrasing, is the discriminator.

The guard could not see it: the transition matcher keyed only on the literal
"Jump to Step", and :305 says "proceed to Step". Widened it to a verb alternation
(jump/proceed/continue/go/return/skip + "to Step N", case-insensitive) and
renamed it TRANSITION_RE to match what it now models. This makes the file's own
docstring promise - that a future spliced-in probe is covered without editing the
test - true for a step whose exit is worded differently. Verified no false
positives: the two pre-existing "continue to Step 3/4" transitions are upstream
of both probes but target pre-probe steps, and the max-rounds bypass block
contains no step transitions at all.

Fail-first verified before fixing :305 - with the widened matcher against the
unfixed workflow the guard fails naming exactly "spec-phase.md:305 jumps to Step
6, skipping mandatory Step 5.6", 4 pass / 1 fail; after the fix, 5/5. The two
sibling probe contract tests stay 16/16.

Also from review:

- STEP_HEADING_RE gains an explicit \r? before $. Without it, on a CRLF checkout
  `.` stops before the \r and the unanchored $ fails to match, yielding ZERO
  steps and vacuously passing every assertion in the file. Not live today
  (.gitattributes forces eol=lf) but this repo has a recurring CRLF-regex bug
  class, so the guard no longer leans on it.
- allow-test-rule category corrected to source-text-is-the-product; the previous
  runtime-contract-is-the-product is not one of the six recognized categories
  (CONTRIBUTING.md:609-619).
- changeset body given the documented bold-lead-in form.
- emitted-drift ack reason updated: +8 -> +10 bytes across five transitions
  (31987 -> 31997), DEFAULT tier, cap 40960.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 10:20:17 -04:00
Tom Boucher
d49a7d0c4d fix(#2853): roadmap.update-plan-progress preserves hand-written annotations (#2916)
* test(#2853): add failing-first regression for plan-progress annotation preservation

The count-bump regex's trailing [^\n]+ swallowed the whole Plans line and
the replacement wrote back only the regenerated count, deleting any
hand-written annotation after it. 8-row matrix covers bold/plain forms,
bare template form, executed path, CRLF, and idempotency.

* fix(#2853): preserve hand-written annotations in roadmap plan-progress bump

The count-bump regex's trailing [^\n]+ swallowed the entire Plans line and
the replacement wrote back only the regenerated count, deleting any
hand-written prose after the count (e.g. a gap-closure annotation).

The verb owns the count token only. Capture the existing count token ($2)
and the trailing line text ($3), and rebuild the line as
<label><new count><surviving text>. Trailing text is preserved ONLY when a
real count token preceded it, so the fresh-template bracketed placeholder
(`[Number of plans…]`) is still replaced cleanly rather than glued after
the count (pre-#2853 behaviour on the template path preserved). CRLF \r is
preserved via [^\r\n].

Widens replaceInCurrentMilestone to accept a replacement callback (needed to
branch on whether the count group matched). The bare Plans: checklist header
is still skipped — the lazy match lands on the summary line first and a
count-less bare header yields no count to anchor preservation to.

* changeset(#2853): backfill PR number 2916

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: sim <sim@local>
2026-07-31 09:54:55 -04:00
Tom Boucher
8635cc447a chore(#2913): prune changeset fragments already promoted in the v1.9.1 CHANGELOG (#2922)
The v1.9.1 finalize consumed these 8 fragments on hotfix/1.9.1 and that
deletion reached main, but the back-merge did not propagate it to next
(c3c6566ac does not touch .changeset/). All 8 are already rendered into
the [1.9.1] section of CHANGELOG.md on next, so leaving them would make
the next release emit duplicate entries for changes already shipped.

Verified each fragment's lead phrase is present in next's CHANGELOG
before removal. No content is lost.

Co-authored-by: sim <sim@local>
2026-07-31 09:51:55 -04:00
clezcoding
9bd0dbf0dd docs(#2534): rewrite your-first-project tutorial for beginners (#2569)
* docs(#2534): rewrite your-first-project tutorial for beginners

Adds a loop mental-model primer (Mermaid), per-step "what just happened"
callouts, a prerequisites flow, a glossary and a troubleshooting table.
Same commands, same .planning artefacts, same to-do CLI example.

Closes #2534

* docs(#2534): make the tutorial runtime-agnostic (all IDEs)

Adds a "Pick your runtime" section (Cursor, Claude Code, OpenCode, Codex,
Gemini CLI, Copilot, Windsurf, Kilo, Cline, Qwen, Antigravity, ...) with the
installer flag and command syntax per runtime (/gsd-*, /gsd:* colon form, and
Cline rules). Keeps the same guaranteed worked example and .planning artefacts.

Closes #2534

* docs(#2534): address review - drop gsd-cursor aside + dead hero comment

- Remove the '(pair with the gsd-cursor EoS ...)' parenthetical from the Cursor row.
- Remove the commented-out reference to a non-existent hero asset.
(Gemini CLI references retained: --gemini is still live in bin/install.js on next.)

Closes #2534

* docs(#2534): fix review defects (keep multi-runtime)

- Replace dead Gemini CLI / --gemini with its live successor Antigravity
  (#1928); remove the invalid --gemini row/flag everywhere.
- Replace fabricated Step 1 output with realistic installer lines
  (71 skills/commands + destination suffix; exact lines vary by runtime).
- Fix 'Skip research' -> choose 'No' on the real Research prompt.
- behaviours -> behaviors (2x).

Multi-runtime 'Pick your runtime' section retained per author intent;
scope re-approval on #2534 still pending.

* docs(#2534): scope tutorial back to single-runtime (Claude Code)

Per trek-e's 2026-07-27 review, resolve the multi-runtime blockers by
returning to the approved scope:

- Remove the 'Pick your runtime' table + per-runtime notes; leave a one-
  line pointer to docs/how-to/install-on-your-runtime.md (which already
  documents all runtimes) rather than duplicate it (avoids the drift).
  This kills Blocker 1 (Antigravity is slash-hyphen, not colon) and
  Blocker 2 (Codex is $gsd-*) at the source.
- Step 1 uses --claude concretely; config-dir prose is Claude-local.
- Step 5: fix singular researcher (plan-phase spawns one gsd-phase-
  researcher), and make the research choice consistent with Step 3
  (choose 'Skip research'); drop the RESEARCH.md artifact line.
- Glossary/troubleshooting/prereqs/Step 2 de-multi-runtimed.

Returns the PR to #2534's approved 'docs-only, same commands' scope.

* docs(#2534): correct tutorial prerequisites and outputs

* docs(#2534): match tutorial research prompts to workflow

* docs(#2534): complete tutorial step guidance

---------

Co-authored-by: clezcoding <clezcoding@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 09:48:54 -04:00
Rezolv
2f66788cd3 docs(#2619): add ADR-2619 observability and shareable diagnostics (#2862)
Completes ADR-0174 §6's observability rollout and adds the outbound trust
boundary that ADR-1577's inbound boundary has no counterpart for.

D1 (wire the seam behind the existing opt-in gate) shipped via #2620 / PR
#2621. D1b records that the unconditional stderr-on-error rule at 0174:105
is the target state, deferred behind an explicit --json-errors envelope
version plus migration note -- disclosed as a partial supersede rather than
retconned. D2-D5 are Directional; D6's non-goals are binding.

Regenerates docs/adr/README.md via scripts/gen-adr-index.cjs --write.

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-31 09:43:04 -04:00
Tom Boucher
49793465d7 docs(#2915): how-to for listing a reviewer lane, and correct the stale listing section (#2917)
* docs(#2904): how-to for listing a reviewer lane in the registry

#2912 shipped the Reviewer Lane Registry, which lands in 1.9.1. Two docs
consequences.

New: docs/how-to/list-your-reviewer-lane.md. A Diataxis how-to for the
publish task -- which of the three catalogs applies (and why a runtime
carrying a reviewer body lists under its primary install shape instead),
opening the required discussion thread BEFORE the PR, the three fields
that reject entries most often (slug grammar differs from id, flags stay
kebab when the slug is snake, install/uninstall must be copy-pasteable),
regenerate-don't-hand-edit, and register-once-then-Releases. Links the
registry README for the field table rather than duplicating it -- the
spec is reference, this is the task flow.

Corrected: ship-a-reviewer-lane.md said "Listing your lane is not wired
yet" and pointed at #2904 as future work. #2906 merged at 11:25Z and
#2912 at 12:11Z, so that section shipped false the moment the registry
landed. Replaced with the publish pointer.

Indexed the new guide and the generated catalog in docs/README.md, and
added the guide to develop-a-capability.md's ecosystem list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2904): caveat credential-bearing configKeys in both the guide and the spec

Isolated security review found the worked entry's `configKeys:
["acme.api_key"]` modelled storing a live credential with no note on
where that value ends up.

Verified: config values are written in plaintext to
.planning/config.json (docs/CONFIGURATION.md:227 -- masking is
display-only, "that file is the security boundary"), and
planning.commit_docs defaults to true (:466). So a credential declared
that way lands in the installing user's git repository unless they have
gitignored .planning/. None of the twelve first-party lanes does this --
they own only review.models.*, host, and prompt-budget keys.

The pattern originates in docs/registries/README.md:221, shipped by
#2912, so the caveat goes on BOTH surfaces rather than only on the copy
that inherited it -- the spec's example is what future authors will read
first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2915): backfill changeset PR number

pr: 0 -> 2917.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 09:29:51 -04:00
Tom Boucher
a9aba61f89 Merge pull request #2920 from open-gsd/chore/backmerge-main-to-next-4f1cce98
chore: back-merge main → next (4f1cce98)
2026-07-31 09:15:08 -04:00
github-actions[bot]
c3c6566ac2 chore: back-merge main into next (4f1cce98) 2026-07-31 13:14:39 +00:00
Tom Boucher
932f99907c Merge pull request #2919 from open-gsd/chore/sync-next-version-1.9.1
chore: sync next package version to 1.9.1
2026-07-31 09:12:18 -04:00
github-actions[bot]
854c93533c chore: sync next package version to 1.9.1 2026-07-31 13:12:09 +00:00
Tom Boucher
4f1cce9875 Merge pull request #2918 from open-gsd/hotfix/1.9.1
chore: merge release v1.9.1 to main
2026-07-31 09:12:06 -04:00
github-actions[bot]
957ebd8e6c chore: promote CHANGELOG for v1.9.1 2026-07-31 13:11:25 +00:00
Tom Boucher
538cb0fc1d enh(#2904): add a reviewer entry type so third-party reviewer lanes are discoverable (#2912)
* feat(#2904): add a `reviewer` entry type so third-party reviewer lanes are discoverable

ADR-2782 made a reviewer lane installable by a third party, but neither
discoverability catalog could hold one. The Community Capability Registry
requires a non-empty `loopExtensionPoints` and forbids a lane from declaring
any hook kind, so a `role: "reviewer"` entry is unsatisfiable by construction;
the EoS Registry is for ADR-1239 host integrations, which a lane is not.

Adds a third catalog — `docs/registries/reviewers.json` →
`docs/registries/reviewer-registry.md` — whose `interactions` describes the
lane: slug, flags, transport, evidenceClass, reviewsSection, requiresBinaries,
configKeys, runtimeCompat.

The lane vocabulary is a hand-written mirror of `capability-validator.cjs`
(the same pattern as `AXES` mirroring `HOST_INTEGRATION_AXES`), with parity
enforced by tests/registry-reviewer-parity.test.cjs. `slug` deliberately uses
the runtime `LANE_SLUG_RE` grammar rather than the registry's kebab-only `id`
rule, so real lanes (`lm_studio`, `4o-mini`) are not rejected.

Two binary type branches became three-way Map dispatch. Both now fail loudly
on an unrecognized type instead of silently treating it as a capability —
`renderMarkdown` in particular writes a committed catalog file, so a silent
wrong-title render was the worst failure mode available.

Also fixed while here: `gen-registry.cjs` parsed source JSON with no error
handling, so a malformed or non-array `capabilities.json` surfaced as a raw
SyntaxError/TypeError instead of an actionable CLI error.

Closes #2904

* fix(#2904): bound and sanitize untrusted registry `interactions` strings

Review findings from the pre-PR passes.

Security (isolated pass): `interactions` string fields reached the generated,
committed Markdown catalog with no control-character check and no length
bound. A `reviewsSection` carrying ESC and a `requiresBinaries` element
carrying NUL plus 5000 characters validated clean and landed verbatim in the
rendered page — `mdInline` escapes Markdown metacharacters and collapses CRLF,
but nothing else. The identical gap already existed on the capability type's
`configKeys`/`requires`/`runtimeCompat`/`produces`/`consumes`, so it is fixed
there too rather than inherited into a third type.

`hasDisallowedControlChar` is lifted to module scope so exactly one
implementation exists, and a shared `validateStringArrayField` enforces
control-character rejection, a 200-character element cap and a 50-element
array cap for both types.

Correctness (standards pass): `renderMarkdown`'s per-entry summary builder was
still an if/else-if chain whose final `else` was the capability branch — the
one per-type dispatch point this change had not converted, and the same silent
fallthrough it removes elsewhere. It now lives in `RENDER_META` alongside the
title, so a fourth type cannot silently inherit capability's rendering. All
three types' rendered output is byte-identical to before the refactor.

Also corrects a test comment that still claimed the reviewer suites were
failing-first against an unmodified module.

* chore(#2904): backfill changeset PR number (#2912)

(cherry picked from commit 90771ddf02)
2026-07-31 08:19:59 -04:00
Tom Boucher
f72f70ad39 docs(#2782): how-to for declaring a reviewer lane in a capability (#2906)
* docs(#2782): how-to for declaring a reviewer lane in a capability

ADR-2782 shipped the Reviewer Lane capability surface in 1.9.0, but the
only documentation was reference (capability-manifest.md) and rationale
(ADR-2782). A capability author had no task-oriented path from "I have a
review CLI" to "/gsd:review invokes it".

Adds docs/how-to/ship-a-reviewer-lane.md: role selection, the spawn and
openai-http worked examples, federated config ownership, the
review-lane query surface as the verification step, what the install
disclosure and egress-host re-verification mean for the author, and the
data-only boundary with the two named CLIs that do not fit today.

Also corrects the manifest reference's `invoke` row, which understated
three enums against the shipped validator: promptChannel omitted `argv`,
effortChannel omitted `env`, and the openai-http sub-shape plus the
required-with-file-arg `outputArg` were undocumented entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): correct four claims found by the two orthogonal reviews

Security review (1 major, 2 minor):
- "a reserved slug" implied the gsd-/anthropic- namespace rule, which
  guards the capability id, not reviewer.slug. The slug guard is
  isReservedName (__proto__/constructor/prototype), a prototype-pollution
  barrier. Both rules are now stated and kept apart.
- Added ADR-2782 D5's own caveat verbatim: disclosure and host pinning
  make the channel visible, pinned and revocable, not safe, and
  consent-at-install is a weaker gate for a standing egress channel than
  for a hook.
- Named integrity/SHA pinning and engines.gsd as the controls that make
  the disclosure tamper-evident and the version range enforceable.

Correctness review (1 major, 1 minor):
- Claimed a name collision is "a hard failure at install". It is not.
  installCapability never runs validateCrossCapability; the check runs in
  loadRegistry, and a colliding overlay is dropped from acceptedMap with a
  warning while the install reports success. Documented as the quiet
  failure mode it is, with the symptom to look for.
- The feature-only field list omitted hooks and activationKey, both of
  which FEATURE_FIELDS_FORBIDDEN_ON_REVIEWER rejects.

Both worked examples re-validated against validateCapability() -> [].

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): restore Kimi Code to the cross-AI reviewer list

set-up-cross-ai-review.md named eleven reviewers; twelve lanes ship. The
kimi-code lane (added by #2718, declared as manifest data by #2798) was
never added here — the same roster-drift class as #2781, which #2800's
parity gate covers for COMMANDS.md and FEATURES.md but not for how-to
prose.

Also points readers at the declared-lane model rather than a static list:
the roster is now generated, a capability can ship its own lane, and
`gsd-tools review-lane sections` answers "what do I actually have".

Verified against the twelve declared bodies in capabilities/*/capability.json.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): use the invocation form the runtime descriptors actually declare

The new guide used /gsd:review. Nothing in this repo produces that form.

- capability-validator VALID_COMMAND_STYLES is {slash-hyphen, shell-var};
  there is no colon/namespaced style in the vocabulary at all.
- 18 of 19 runtime capabilities declare commandStyle "slash-hyphen",
  claude included; codex is "shell-var". Every artifactLayout prefix is
  "gsd-".
- A plain-file command install never namespaces, so .claude/commands/
  gsd-review.md is typed /gsd-review.
- The Claude Code plugin surface would namespace on plugin.json "name",
  which is "gsd-core" -- so the plugin form would be /gsd-core:review.
  The commands/gsd/ subdirectory is cosmetic and contributes nothing to
  the invoked name.

So /gsd:review is neither the installed form nor the plugin form. Uses
/gsd-review, matching set-up-cross-ai-review.md and the 18 descriptors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): state that lane listing is not wired yet

The guide could describe authoring, validating, and installing a lane but
implied a publishing path that does not exist. Neither discoverability
catalog can accept one: a Community Capability Registry entry requires a
non-empty loopExtensionPoints plus hookKinds, and a role:"reviewer"
capability is forbidden from declaring steps/contributions/gates, so both
fields are unsatisfiable rather than merely unset. The EoS Registry is
ADR-1239 host integrations, which a lane is not.

The registry schema predates the reviewer role by 17 days (#2182 Jul 11,
ADR-2782 Jul 28) and registry-schema.cjs has zero occurrences of
"reviewer". Tracked for a 1.9.x point release by #2904.

Says so explicitly, and tells authors NOT to file a loop extension point
they do not use to get past validation -- a schema satisfiable only by
lying is one that will be lied to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2782): backfill PR number and fix the changeset invocation form

pr: 0 -> 2906.

Also corrects /gsd:review -> /gsd-review in the fragment body. The
fragment renders into CHANGELOG.md, which is a reader-facing docs surface
and is never passed through the install-time converter -- so the colon
form would ship the #2903 drift into a permanent release artifact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2907): propagate the invocation-form fix and restore lane order

Two findings from an isolated review of the post-review delta.

- docs/README.md and develop-a-capability.md still said /gsd:review in
  the cross-links added for the new guide. The form was corrected in the
  guide itself but not in the two entries pointing at it, leaving three
  docs making the same claim in two different forms.
- set-up-cross-ai-review.md inserted Kimi Code between Antigravity and
  Ollama. REVIEWER_LANES is ordered by write_reviews order and kimi-code
  is 12th, appended after llama_cpp; the prose list mirrored declaration
  order before this change and now does again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 557d46984e)
2026-07-31 08:19:59 -04:00
Tom Boucher
7112c6ca47 fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context

verify-summary's Pattern 1 matched any backticked path-like token with no
context check, so a prose mention of a future deliverable (`shared/types.ts`
in a 'next phase will add…' sentence) was checked for existence and its absence
failed the verdict on a healthy phase. #2685 added shape filtering but no
context check.

- src/verify.cts: both extraction patterns now require a claim label on the line
  (Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no
  longer matches; genuine labeled claims still do.
- gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION
  so relative claim paths resolve against the project root, not the raw cwd
  (subdirectory invocation no longer manufactures missing files).

Regression tests: prose mention not treated as a claim; prose-only SUMMARY
passes; absent claimed file still fails.

* chore(#2844): backfill changeset PR 2910

---------

Co-authored-by: Test <test@example.com>
(cherry picked from commit 42f4f184c0)
2026-07-31 08:19:59 -04:00
Tom Boucher
e1b275766d fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH (#2909)
* fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH

findProjectRoot's heuristic (3) used isInsideGitRepo(parent), which only checked
'does SOME .git exist between start and the ancestor' — it never verified the
.git was co-located with / bounded the trusted .planning/. A nested child repo
(own .git, no .planning) under an ancestor GSD project satisfied the check, so
resolution silently crossed into the ancestor project (wrong identity, exit 0).

Add nearestGitRoot(from, upTo) (fs-walk, no spawn) and use it in heuristics (3)
and (4): if the caller is inside its own nested repo whose root is strictly below
the candidate ancestor, do not return that ancestor. The plain-descendant (#1414),
co-located .git+.planning, and sub_repos/multiRepo cases are unchanged.

Regression test: a nested child .git under an ancestor .planning no longer
resolves to the ancestor; the co-located single-repo case still resolves.

* chore(#2843): backfill changeset PR 2909

---------

Co-authored-by: Test <test@example.com>
(cherry picked from commit 39dbe5e0f5)
2026-07-31 08:19:59 -04:00
Tom Boucher
3aabb0f441 fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point (#2905)
* fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point

gsd-code-fixer was the only writer that hand-rolled a git worktree inside the
agent prompt and the only one that never read workflow.use_worktrees. With the
setting explicitly false, --fix still created worktrees; the fresh worktree had
no node_modules, so runs improvised a teardown whose `rm -rf` followed a Windows
junction into the REAL node_modules (silent data loss, 3x observed).

Defect 1: gate setup_worktree + its cleanup tail on workflow.use_worktrees
(same gsd_run query config-get read the four sibling workflows use). When false:
edit/commit in the main checkout (wt='.', no temp branch, no sentinel, no cleanup).

Defect 2: forbid rm -rf on a possible reparse point in the spec — never fall
through to a destructive remove; on failure, stop and surface the error.

Defect 3: REVIEW-FIX records where verification ran (main checkout vs worktree).

The transactional worktree path (#2839/#2990/#2686) is unchanged when worktrees
are enabled. Docs-parity guards in tests/code-review.test.cjs bind the spec to
the fix.

* chore(#2825): backfill changeset PR 2905

* fix(#2825): drop stale agents/gsd-code-fixer.toml ack + spent entries (emitted-attribution)

gsd-test failed: agents/gsd-code-fixer.toml is a STALE ack here — that TOML was
changed by #2834 (now in next), not this PR. The 4 base-carried acks
(autonomous/discuss-phase-assumptions/next/plan-phase) are spent/inert.

---------

Co-authored-by: Test <test@example.com>
(cherry picked from commit 5d0fd4dc53)
2026-07-31 08:19:59 -04:00
Tom Boucher
dc73680532 fix(#2834): write defaults.json before agent TOML generation on clean Codex install (#2900)
* fix(#2834): write defaults.json before agent TOML generation on clean Codex install

Extracted writeNonClaudeDefaults(runtime) and called it BEFORE installCodexConfig
so the runtime-aware model resolver has resolve_model_ids=omit + runtime=codex in
~/.gsd/defaults.json before agent TOMLs are generated. Pre-fix, a clean first Codex
install generated TOMLs with no model fields (the resolver didn't know the runtime);
a second run fixed it. The original inline defaults-write block (which ran AFTER agent
generation) is replaced by the earlier function call (idempotent).

* chore(#2834): changeset fragment

* fix+test(#2834): acknowledge codex TOML drift (emitted-attribution) + fix test comment window

The emitted-attribution gate flags 19 codex agent TOMLs that now carry model-routing
fields (the fix's correct effect) but can't link them to a .md or src/ change (the fix
is in bin/install.js ordering). Acknowledge the drift in emitted-drift-ack.json. Fix the
test's comment-detection window (300 chars to capture the #2834 rationale).

* fix(#2834): ack remaining 15 codex TOML drift paths

* chore(#2834): backfill changeset PR number (2900)

* chore(#2834): ack code-review.md growth from concurrent merge (rebase pickup)

* fix(#2834): remove stale code-review.md ack (emitted-attribution failure)

CI failed: 'differential attribution over the real tree' — the code-review.md
ack added in 6e0b4b3b8 ('ack code-review.md growth from concurrent merge') is
STALE: this PR's diff does not touch code-review.md (only bin/install.js + tests
+ changeset), so the ack explains growth that isn't here. The base already
absorbed the concurrent code-review.md growth; the ack is inert here and the
gate flags it as stale. Remove it.

---------

Co-authored-by: Test <test@example.com>
(cherry picked from commit 00c859fae5)
2026-07-31 08:19:59 -04:00
Tom Boucher
a38e4d080d fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row (#2902)
* fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row

Two coupled defects in the requirement traceability state machine:

Defect 1 (terminal state): requirements revert-phase (the gaps_found response)
left a row at 'Gaps Found' with no inverse — neither mark-complete's /^pending$/i
guard nor the phase-complete reconcile's /^(?:pending|in progress)$/i accepted it,
so a single failed verification stranded every requirement permanently and blocked
the milestone. Widen both guards to accept 'gaps found' so a genuinely-satisfied
stranded row reaches Complete again.

Defect 2 (false success): mark-complete ORed checkboxHit || tableHit for 'updated',
so on a Gaps Found row it flipped the checkbox but could not move the row, yet
reported updated:true. When a traceability table has a row for an ID, gate 'updated'
on the row moving (tableHit) — a checkbox-only flip on a table-bearing file no
longer lies. The #2140 table_unmatched path (no row for the ID) is preserved.

* chore(#2788): backfill changeset PR 2902

---------

Co-authored-by: Test <test@example.com>
(cherry picked from commit 9f567a1627)
2026-07-31 08:19:59 -04:00
Tom Boucher
14cfbdaad0 fix(#2667): run-with-timeout mediates .cmd/.bat spawns on Windows (CVE-2024-27980); fallow pre-pass names failure kind (#2897)
* fix(#2667): mediate .cmd/.bat/.exe spawns on Windows; split fallow pre-pass failure diagnostic

run-with-timeout spawned .cmd/.bat/.exe commands without shell:true on Windows,
tripping Node's CVE-2024-27980 EINVAL (April 2024 security hardening). The fallow
structural pre-pass then no-op'd silently — a hard execution failure read the same
as 'optional dependency absent'.

(A) gsd-core/bin/gsd-tools.cjs runWithTimeout: gate shell:true on
    (win32 && command ends in .cmd/.bat/.exe). Narrow by design — never fires for
    the 7 `bash -c` callers (command is `bash`, no such suffix), so the recorded
    no-shell-for-argv-array security contract (DEFECT.UNBOUNDED-SUBPROCESS) is
    preserved; cmdArgs stays an array. POSIX untouched.
(B) code-review.md fallow pre-pass: name the failure KIND (timeout / spawn failure
    / crash / not-found) so a Windows .cmd spawn failure is not mistaken for an
    absent binary.

Regression test in tests/run-with-timeout.test.cjs gated to win32 (.cmd/.bat/.exe
shims run with exit 0 + non-empty stdout; pre-fix EINVAL → exit 125/empty). POSIX
negative-space test guards the unchanged bash -c callers.

* chore(#2667): changeset fragment

* chore(#2667): backfill changeset PR 2897 + correct body (cmd.exe array, not shell:true)

* fix(#2667): exclude .exe from the win32 spawn-mediation gate; ack code-review.md growth

CI caught two failures on the first push:

1. windows-24: 'exits 124 when the wall-clock budget is exceeded' regressed. The
   gate matched .exe, so the HANG command (node.exe -e 'setTimeout(...)') was
   wrapped in 'cmd.exe /c node.exe ...' — the wrapped child escaped the timeout
   cap's process-group reap (exit 124 never fired; hit the 30s harness backstop)
   AND cmd.exe risked mis-parsing the -e script arg. .exe is INTENTIONALLY
   excluded now: real PE executables spawn fine directly; only .cmd/.bat are the
   CVE-2024-27980 EINVAL cases. The .exe test becomes a negative-space test
   (node.exe spawned directly, exit 0).

2. ubuntu-22: emitted-attribution — code-review.md grew 1177 bytes from the
   #2667 fallow pre-pass failure-KIND case statement; acknowledge it.

---------

Co-authored-by: Test <test@example.com>
(cherry picked from commit 79ed181ec0)
2026-07-31 08:19:59 -04:00
github-actions[bot]
13fcbe45a6 chore: bump version to 1.9.1 for hotfix 2026-07-31 12:19:17 +00:00
Tom Boucher
90771ddf02 enh(#2904): add a reviewer entry type so third-party reviewer lanes are discoverable (#2912)
* feat(#2904): add a `reviewer` entry type so third-party reviewer lanes are discoverable

ADR-2782 made a reviewer lane installable by a third party, but neither
discoverability catalog could hold one. The Community Capability Registry
requires a non-empty `loopExtensionPoints` and forbids a lane from declaring
any hook kind, so a `role: "reviewer"` entry is unsatisfiable by construction;
the EoS Registry is for ADR-1239 host integrations, which a lane is not.

Adds a third catalog — `docs/registries/reviewers.json` →
`docs/registries/reviewer-registry.md` — whose `interactions` describes the
lane: slug, flags, transport, evidenceClass, reviewsSection, requiresBinaries,
configKeys, runtimeCompat.

The lane vocabulary is a hand-written mirror of `capability-validator.cjs`
(the same pattern as `AXES` mirroring `HOST_INTEGRATION_AXES`), with parity
enforced by tests/registry-reviewer-parity.test.cjs. `slug` deliberately uses
the runtime `LANE_SLUG_RE` grammar rather than the registry's kebab-only `id`
rule, so real lanes (`lm_studio`, `4o-mini`) are not rejected.

Two binary type branches became three-way Map dispatch. Both now fail loudly
on an unrecognized type instead of silently treating it as a capability —
`renderMarkdown` in particular writes a committed catalog file, so a silent
wrong-title render was the worst failure mode available.

Also fixed while here: `gen-registry.cjs` parsed source JSON with no error
handling, so a malformed or non-array `capabilities.json` surfaced as a raw
SyntaxError/TypeError instead of an actionable CLI error.

Closes #2904

* fix(#2904): bound and sanitize untrusted registry `interactions` strings

Review findings from the pre-PR passes.

Security (isolated pass): `interactions` string fields reached the generated,
committed Markdown catalog with no control-character check and no length
bound. A `reviewsSection` carrying ESC and a `requiresBinaries` element
carrying NUL plus 5000 characters validated clean and landed verbatim in the
rendered page — `mdInline` escapes Markdown metacharacters and collapses CRLF,
but nothing else. The identical gap already existed on the capability type's
`configKeys`/`requires`/`runtimeCompat`/`produces`/`consumes`, so it is fixed
there too rather than inherited into a third type.

`hasDisallowedControlChar` is lifted to module scope so exactly one
implementation exists, and a shared `validateStringArrayField` enforces
control-character rejection, a 200-character element cap and a 50-element
array cap for both types.

Correctness (standards pass): `renderMarkdown`'s per-entry summary builder was
still an if/else-if chain whose final `else` was the capability branch — the
one per-type dispatch point this change had not converted, and the same silent
fallthrough it removes elsewhere. It now lives in `RENDER_META` alongside the
title, so a fourth type cannot silently inherit capability's rendering. All
three types' rendered output is byte-identical to before the refactor.

Also corrects a test comment that still claimed the reviewer suites were
failing-first against an unmodified module.

* chore(#2904): backfill changeset PR number (#2912)
2026-07-31 08:11:57 -04:00
Tom Boucher
557d46984e docs(#2782): how-to for declaring a reviewer lane in a capability (#2906)
* docs(#2782): how-to for declaring a reviewer lane in a capability

ADR-2782 shipped the Reviewer Lane capability surface in 1.9.0, but the
only documentation was reference (capability-manifest.md) and rationale
(ADR-2782). A capability author had no task-oriented path from "I have a
review CLI" to "/gsd:review invokes it".

Adds docs/how-to/ship-a-reviewer-lane.md: role selection, the spawn and
openai-http worked examples, federated config ownership, the
review-lane query surface as the verification step, what the install
disclosure and egress-host re-verification mean for the author, and the
data-only boundary with the two named CLIs that do not fit today.

Also corrects the manifest reference's `invoke` row, which understated
three enums against the shipped validator: promptChannel omitted `argv`,
effortChannel omitted `env`, and the openai-http sub-shape plus the
required-with-file-arg `outputArg` were undocumented entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): correct four claims found by the two orthogonal reviews

Security review (1 major, 2 minor):
- "a reserved slug" implied the gsd-/anthropic- namespace rule, which
  guards the capability id, not reviewer.slug. The slug guard is
  isReservedName (__proto__/constructor/prototype), a prototype-pollution
  barrier. Both rules are now stated and kept apart.
- Added ADR-2782 D5's own caveat verbatim: disclosure and host pinning
  make the channel visible, pinned and revocable, not safe, and
  consent-at-install is a weaker gate for a standing egress channel than
  for a hook.
- Named integrity/SHA pinning and engines.gsd as the controls that make
  the disclosure tamper-evident and the version range enforceable.

Correctness review (1 major, 1 minor):
- Claimed a name collision is "a hard failure at install". It is not.
  installCapability never runs validateCrossCapability; the check runs in
  loadRegistry, and a colliding overlay is dropped from acceptedMap with a
  warning while the install reports success. Documented as the quiet
  failure mode it is, with the symptom to look for.
- The feature-only field list omitted hooks and activationKey, both of
  which FEATURE_FIELDS_FORBIDDEN_ON_REVIEWER rejects.

Both worked examples re-validated against validateCapability() -> [].

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): restore Kimi Code to the cross-AI reviewer list

set-up-cross-ai-review.md named eleven reviewers; twelve lanes ship. The
kimi-code lane (added by #2718, declared as manifest data by #2798) was
never added here — the same roster-drift class as #2781, which #2800's
parity gate covers for COMMANDS.md and FEATURES.md but not for how-to
prose.

Also points readers at the declared-lane model rather than a static list:
the roster is now generated, a capability can ship its own lane, and
`gsd-tools review-lane sections` answers "what do I actually have".

Verified against the twelve declared bodies in capabilities/*/capability.json.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): use the invocation form the runtime descriptors actually declare

The new guide used /gsd:review. Nothing in this repo produces that form.

- capability-validator VALID_COMMAND_STYLES is {slash-hyphen, shell-var};
  there is no colon/namespaced style in the vocabulary at all.
- 18 of 19 runtime capabilities declare commandStyle "slash-hyphen",
  claude included; codex is "shell-var". Every artifactLayout prefix is
  "gsd-".
- A plain-file command install never namespaces, so .claude/commands/
  gsd-review.md is typed /gsd-review.
- The Claude Code plugin surface would namespace on plugin.json "name",
  which is "gsd-core" -- so the plugin form would be /gsd-core:review.
  The commands/gsd/ subdirectory is cosmetic and contributes nothing to
  the invoked name.

So /gsd:review is neither the installed form nor the plugin form. Uses
/gsd-review, matching set-up-cross-ai-review.md and the 18 descriptors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2782): state that lane listing is not wired yet

The guide could describe authoring, validating, and installing a lane but
implied a publishing path that does not exist. Neither discoverability
catalog can accept one: a Community Capability Registry entry requires a
non-empty loopExtensionPoints plus hookKinds, and a role:"reviewer"
capability is forbidden from declaring steps/contributions/gates, so both
fields are unsatisfiable rather than merely unset. The EoS Registry is
ADR-1239 host integrations, which a lane is not.

The registry schema predates the reviewer role by 17 days (#2182 Jul 11,
ADR-2782 Jul 28) and registry-schema.cjs has zero occurrences of
"reviewer". Tracked for a 1.9.x point release by #2904.

Says so explicitly, and tells authors NOT to file a loop extension point
they do not use to get past validation -- a schema satisfiable only by
lying is one that will be lied to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2782): backfill PR number and fix the changeset invocation form

pr: 0 -> 2906.

Also corrects /gsd:review -> /gsd-review in the fragment body. The
fragment renders into CHANGELOG.md, which is a reader-facing docs surface
and is never passed through the install-time converter -- so the colon
form would ship the #2903 drift into a permanent release artifact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2907): propagate the invocation-form fix and restore lane order

Two findings from an isolated review of the post-review delta.

- docs/README.md and develop-a-capability.md still said /gsd:review in
  the cross-links added for the new guide. The form was corrected in the
  guide itself but not in the two entries pointing at it, leaving three
  docs making the same claim in two different forms.
- set-up-cross-ai-review.md inserted Kimi Code between Antigravity and
  Ollama. REVIEWER_LANES is ordered by write_reviews order and kimi-code
  is 12th, appended after llama_cpp; the prose list mirrored declaration
  order before this change and now does again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:25:01 -04:00
Tom Boucher
42f4f184c0 fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context

verify-summary's Pattern 1 matched any backticked path-like token with no
context check, so a prose mention of a future deliverable (`shared/types.ts`
in a 'next phase will add…' sentence) was checked for existence and its absence
failed the verdict on a healthy phase. #2685 added shape filtering but no
context check.

- src/verify.cts: both extraction patterns now require a claim label on the line
  (Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no
  longer matches; genuine labeled claims still do.
- gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION
  so relative claim paths resolve against the project root, not the raw cwd
  (subdirectory invocation no longer manufactures missing files).

Regression tests: prose mention not treated as a claim; prose-only SUMMARY
passes; absent claimed file still fails.

* chore(#2844): backfill changeset PR 2910

---------

Co-authored-by: Test <test@example.com>
2026-07-31 03:16:15 -04:00
Tom Boucher
39dbe5e0f5 fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH (#2909)
* fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH

findProjectRoot's heuristic (3) used isInsideGitRepo(parent), which only checked
'does SOME .git exist between start and the ancestor' — it never verified the
.git was co-located with / bounded the trusted .planning/. A nested child repo
(own .git, no .planning) under an ancestor GSD project satisfied the check, so
resolution silently crossed into the ancestor project (wrong identity, exit 0).

Add nearestGitRoot(from, upTo) (fs-walk, no spawn) and use it in heuristics (3)
and (4): if the caller is inside its own nested repo whose root is strictly below
the candidate ancestor, do not return that ancestor. The plain-descendant (#1414),
co-located .git+.planning, and sub_repos/multiRepo cases are unchanged.

Regression test: a nested child .git under an ancestor .planning no longer
resolves to the ancestor; the co-located single-repo case still resolves.

* chore(#2843): backfill changeset PR 2909

---------

Co-authored-by: Test <test@example.com>
2026-07-31 02:30:50 -04:00
Tom Boucher
5d0fd4dc53 fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point (#2905)
* fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point

gsd-code-fixer was the only writer that hand-rolled a git worktree inside the
agent prompt and the only one that never read workflow.use_worktrees. With the
setting explicitly false, --fix still created worktrees; the fresh worktree had
no node_modules, so runs improvised a teardown whose `rm -rf` followed a Windows
junction into the REAL node_modules (silent data loss, 3x observed).

Defect 1: gate setup_worktree + its cleanup tail on workflow.use_worktrees
(same gsd_run query config-get read the four sibling workflows use). When false:
edit/commit in the main checkout (wt='.', no temp branch, no sentinel, no cleanup).

Defect 2: forbid rm -rf on a possible reparse point in the spec — never fall
through to a destructive remove; on failure, stop and surface the error.

Defect 3: REVIEW-FIX records where verification ran (main checkout vs worktree).

The transactional worktree path (#2839/#2990/#2686) is unchanged when worktrees
are enabled. Docs-parity guards in tests/code-review.test.cjs bind the spec to
the fix.

* chore(#2825): backfill changeset PR 2905

* fix(#2825): drop stale agents/gsd-code-fixer.toml ack + spent entries (emitted-attribution)

gsd-test failed: agents/gsd-code-fixer.toml is a STALE ack here — that TOML was
changed by #2834 (now in next), not this PR. The 4 base-carried acks
(autonomous/discuss-phase-assumptions/next/plan-phase) are spent/inert.

---------

Co-authored-by: Test <test@example.com>
2026-07-31 01:57:50 -04:00
Tom Boucher
00c859fae5 fix(#2834): write defaults.json before agent TOML generation on clean Codex install (#2900)
* fix(#2834): write defaults.json before agent TOML generation on clean Codex install

Extracted writeNonClaudeDefaults(runtime) and called it BEFORE installCodexConfig
so the runtime-aware model resolver has resolve_model_ids=omit + runtime=codex in
~/.gsd/defaults.json before agent TOMLs are generated. Pre-fix, a clean first Codex
install generated TOMLs with no model fields (the resolver didn't know the runtime);
a second run fixed it. The original inline defaults-write block (which ran AFTER agent
generation) is replaced by the earlier function call (idempotent).

* chore(#2834): changeset fragment

* fix+test(#2834): acknowledge codex TOML drift (emitted-attribution) + fix test comment window

The emitted-attribution gate flags 19 codex agent TOMLs that now carry model-routing
fields (the fix's correct effect) but can't link them to a .md or src/ change (the fix
is in bin/install.js ordering). Acknowledge the drift in emitted-drift-ack.json. Fix the
test's comment-detection window (300 chars to capture the #2834 rationale).

* fix(#2834): ack remaining 15 codex TOML drift paths

* chore(#2834): backfill changeset PR number (2900)

* chore(#2834): ack code-review.md growth from concurrent merge (rebase pickup)

* fix(#2834): remove stale code-review.md ack (emitted-attribution failure)

CI failed: 'differential attribution over the real tree' — the code-review.md
ack added in 6e0b4b3b8 ('ack code-review.md growth from concurrent merge') is
STALE: this PR's diff does not touch code-review.md (only bin/install.js + tests
+ changeset), so the ack explains growth that isn't here. The base already
absorbed the concurrent code-review.md growth; the ack is inert here and the
gate flags it as stale. Remove it.

---------

Co-authored-by: Test <test@example.com>
2026-07-31 00:47:55 -04:00
Tom Boucher
9f567a1627 fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row (#2902)
* fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row

Two coupled defects in the requirement traceability state machine:

Defect 1 (terminal state): requirements revert-phase (the gaps_found response)
left a row at 'Gaps Found' with no inverse — neither mark-complete's /^pending$/i
guard nor the phase-complete reconcile's /^(?:pending|in progress)$/i accepted it,
so a single failed verification stranded every requirement permanently and blocked
the milestone. Widen both guards to accept 'gaps found' so a genuinely-satisfied
stranded row reaches Complete again.

Defect 2 (false success): mark-complete ORed checkboxHit || tableHit for 'updated',
so on a Gaps Found row it flipped the checkbox but could not move the row, yet
reported updated:true. When a traceability table has a row for an ID, gate 'updated'
on the row moving (tableHit) — a checkbox-only flip on a table-bearing file no
longer lies. The #2140 table_unmatched path (no row for the ID) is preserved.

* chore(#2788): backfill changeset PR 2902

---------

Co-authored-by: Test <test@example.com>
2026-07-31 00:04:04 -04:00
Tom Boucher
9f73703a4c Merge pull request #2901 from open-gsd/chore/backmerge-main-to-next-efbbcc35
chore: back-merge main → next (efbbcc35)
2026-07-30 23:17:14 -04:00
github-actions[bot]
ee519e883c chore: back-merge main into next (efbbcc35) 2026-07-31 03:16:47 +00:00
Tom Boucher
79ed181ec0 fix(#2667): run-with-timeout mediates .cmd/.bat spawns on Windows (CVE-2024-27980); fallow pre-pass names failure kind (#2897)
* fix(#2667): mediate .cmd/.bat/.exe spawns on Windows; split fallow pre-pass failure diagnostic

run-with-timeout spawned .cmd/.bat/.exe commands without shell:true on Windows,
tripping Node's CVE-2024-27980 EINVAL (April 2024 security hardening). The fallow
structural pre-pass then no-op'd silently — a hard execution failure read the same
as 'optional dependency absent'.

(A) gsd-core/bin/gsd-tools.cjs runWithTimeout: gate shell:true on
    (win32 && command ends in .cmd/.bat/.exe). Narrow by design — never fires for
    the 7 `bash -c` callers (command is `bash`, no such suffix), so the recorded
    no-shell-for-argv-array security contract (DEFECT.UNBOUNDED-SUBPROCESS) is
    preserved; cmdArgs stays an array. POSIX untouched.
(B) code-review.md fallow pre-pass: name the failure KIND (timeout / spawn failure
    / crash / not-found) so a Windows .cmd spawn failure is not mistaken for an
    absent binary.

Regression test in tests/run-with-timeout.test.cjs gated to win32 (.cmd/.bat/.exe
shims run with exit 0 + non-empty stdout; pre-fix EINVAL → exit 125/empty). POSIX
negative-space test guards the unchanged bash -c callers.

* chore(#2667): changeset fragment

* chore(#2667): backfill changeset PR 2897 + correct body (cmd.exe array, not shell:true)

* fix(#2667): exclude .exe from the win32 spawn-mediation gate; ack code-review.md growth

CI caught two failures on the first push:

1. windows-24: 'exits 124 when the wall-clock budget is exceeded' regressed. The
   gate matched .exe, so the HANG command (node.exe -e 'setTimeout(...)') was
   wrapped in 'cmd.exe /c node.exe ...' — the wrapped child escaped the timeout
   cap's process-group reap (exit 124 never fired; hit the 30s harness backstop)
   AND cmd.exe risked mis-parsing the -e script arg. .exe is INTENTIONALLY
   excluded now: real PE executables spawn fine directly; only .cmd/.bat are the
   CVE-2024-27980 EINVAL cases. The .exe test becomes a negative-space test
   (node.exe spawned directly, exit 0).

2. ubuntu-22: emitted-attribution — code-review.md grew 1177 bytes from the
   #2667 fallow pre-pass failure-KIND case statement; acknowledge it.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 23:14:17 -04:00
Tom Boucher
b54e023f05 Merge pull request #2899 from open-gsd/chore/sync-next-version-1.9.0
chore: sync next package version to 1.9.0
2026-07-30 23:14:09 -04:00
github-actions[bot]
4232a79396 chore: sync next package version to 1.9.0 2026-07-31 03:14:04 +00:00
Tom Boucher
efbbcc359e Merge pull request #2898 from open-gsd/release/1.9.0
chore: merge release v1.9.0 to main
2026-07-30 23:14:01 -04:00
github-actions[bot]
7d270a205c chore: promote CHANGELOG for v1.9.0 2026-07-31 03:13:23 +00:00
github-actions[bot]
63503b2fc1 chore: finalize v1.9.0 2026-07-31 03:00:43 +00:00
Tom Boucher
6e0bc50142 fix(#2666): code-review scopes root-level + extensionless build files, cross-checks against git diff (#2895)
* test(#2666): add regression + docs-parity guards for code-review file scoper

The Tier-2 SUMMARY.md extractor dropped every repository-root file (no `/`)
and every extensionless build file (Dockerfile/Makefile/etc.) via an AND-joined
predicate. Adds behavioral tests against the pure-function mirror plus
docs-parity structural guards that bind the shipped workflow .md to the fix.

RED: the docs-parity guards fail against the pre-fix shipped predicate.

* fix(#2666): accept root-level + extensionless build files in code-review scope; intersect-and-warn

Two coordinated edits to gsd-core/workflows/code-review.md compute_file_scope:

(A) Tier-2 SUMMARY extractor: replace the AND-joined predicate
`/\\//.test(raw) && /\\.[A-Za-z0-9]+$/.test(raw)` (which required BOTH a
directory separator AND a trailing extension, silently dropping every
root-level file and every extensionless build file) with a relaxed predicate
that accepts any path with a trailing extension OR a known extensionless
build basename (Dockerfile/Containerfile/Makefile/Justfile/Procfile).

(B) Tier-3: convert the eq-zero git-diff gate into an intersect-and-warn —
whenever a reliable diff base is available, cross-check the SUMMARY scope
against `git diff --name-only` and warn about (then add) any changed files
the SUMMARY extractor did not surface. Portable (bash 3.2, no associative
arrays) so a partial SUMMARY result can no longer silently ship an
incomplete review scope.

* chore(#2666): changeset fragment

* fix(#2666): use exact whole-line matching (grep -Fxq) in Tier-3 cross-check

Adversarial review found the unanchored `case "$IN_SCOPE" in *"$file"$\\n*`
substring membership test would false-match: a root-level `Dockerfile` in the
diff substring-matches an already-scoped `docker/Dockerfile`, silently skipping
it — reintroducing the exact class of silent-scope-loss bug this PR fixes.

Switch to `grep -Fxq` (exact whole-line match). Add docs-parity guard for
exact matching + the basename-collision regression.

* fix(#2666): resolve gsd-test failures — paraphrase predicate in comment, ack code-review.md growth

gsd-test caught 3 issues on 7469c3f22:
1. The docs-parity guard fired on the .md COMMENT which restated the buggy
   predicate verbatim — paraphrase the comment so it no longer contains the
   exact string the guard detects.
2. Cascade subtest failure from #1.
3. emitted-attribution: code-review.md grew 2814 bytes — acknowledge the
   deliberate #2666 growth in tests/emitted-drift-ack.json.

* chore(#2666): backfill changeset PR number 2895

---------

Co-authored-by: Test <test@example.com>
2026-07-30 22:27:32 -04:00
Tom Boucher
f093738412 fix(#2891): normalize emitted version against the measured tree, not the measuring repo (#2894)
* fix(#2891): normalize emitted version against the measured tree, not the measuring repo

buildParityManifest normalized the install-time {{GSD_VERSION}} stamp using
PKG_VERSION, bound at module load from the MEASURING repo's package.json. Since
#2767, currentManifests({repoRoot}) measures a DIFFERENT checkout, so during a
release cut the baseline worktree (origin/next, 1.8.0) was normalized with the
current tree's version (1.9.0) and its literal 1.8.0 stamp survived into the
hash. All 364 emitted hook paths diverged and the differential attribution gate
hard-failed every finalize/rc run.

Normalize against the version of the tree that PRODUCED the emitted output:
buildParityManifest takes an explicit pkgVersion, and currentManifests resolves
it from the measured tree via a new fail-closed measuredPackageVersion().

* chore(#2891): backfill changeset pr number (#2894)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 22:14:47 -04:00
Tom Boucher
cd5b8643ae fix(#2828): state sync reports correct total_phases on a flat unmilestoned roadmap (#2892)
* fix+test(#2828): total_phases uses roadmap count on flat unmilestoned roadmap

The read-path disk-scan cache fell back to phaseDirs.length (1) when milestoneBounded
was false, even though roadmapPhaseCount (6) was correct for a flat roadmap (no sibling
milestones to conflate). Use roadmapPhaseCount as the floor when > 0, matching the
write-path (cmdStateSync) which already did this. The milestoneBounded flag still flows
to milestoneUnbounded for the percent-skip (#1761 guard preserved). Regression test
asserts state-sync writes progress.total_phases:6 for a flat 6-phase roadmap + 1 phase dir.

* chore(#2828): changeset fragment

* test(#2828): add negative-space coverage (Math.max floor mutant + no-roadmap fallback) — review findings

The 6-phase test alone couldn't kill a Math.max-dropping mutant (1<6). Add: a
3-phase-dir/2-roadmap-phase case proving Math.max(dirs,count) floor; a no-roadmap
case proving phaseDirs.length fallback.

* fix(#2828): refine — distinguish flat unmilestoned from milestoned-unbounded (preserve #1761)

The first-pass fix (roadmapPhaseCount > 0 always) re-broke #1761: a milestoned-
unbounded roadmap (asserted milestone not among existing version headings) conflated
sibling milestones (8 = 4+4). Refine with a hasMilestoneSectioning discriminator:
^#{2,3}(?!Phase) detects non-Phase h2/h3 milestone section headings. A FLAT roadmap
(only ### Phase headings + a # title) has none → safe to use roadmapPhaseCount; a
SECTIONED-but-unbounded roadmap has them → fall back to phaseDirs.length (#1761).
Verified both cases locally (flat→6, sectioned-unbounded→1).

* test(#2828): remove two fragile negative-space tests (phase-dir scanner internals)

The Math.max-floor and no-roadmap tests made assumptions about the phase-dir scanner's
internals (which dirs count as 'realized') that didn't hold. The core regression test
(6-phase flat → total_phases:6) plus the existing #1761 conflation tests (which the
refined fix preserves) provide sufficient coverage.

* chore(#2828): backfill changeset PR number 2892

* fix(#2828): replace ReDoS-prone regex in regression test with line-by-line parse

CodeQL flagged the nested-quantifier regex (`(?:[ \t]+\w+:.+\r?\n?)*?`) in
tests/issue-2828-flat-roadmap-total-phases.test.cjs as a high-severity
catastrophic-backtracking risk. Rewrite the STATE.md progress.total_phases
extraction as a ReDoS-safe line-by-line block walk.

---------

Co-authored-by: Test <test@example.com>
2026-07-30 21:19:36 -04:00
Tom Boucher
54cb4145bf fix(#2765): bump brace-expansion to patched 1.1.18/5.0.9 (high-severity devDep advisory) (#2888)
* fix(#2765): bump brace-expansion to patched 1.1.18/5.0.9 (high-severity devDep advisory)

npm audit fix (non-breaking) bumps the lockfile: brace-expansion 1.1.15→1.1.18
(eslint-nested via minimatch@3.x) and 5.0.6→5.0.9 (stryker-nested). Both 1.1.18 and
5.0.9 were published 2026-07-30 as the patch backports for GHSA-3jxr-9vmj-r5cp /
GHSA-mh99-v99m-4gvg (range <=5.0.7). No overrides needed (in-range bump), no major
bumps, no --force. Production (npm audit --omit=dev) unaffected (devDep only).
Add a structural test pinning the installed versions so the bump can't silently regress.

* chore(#2765): changeset fragment

* fix(#2765): correct changeset issue ref + parse patch version as number (review findings)

- changeset cited #2762 (typo) — fix to (#2765).
- test compared v.split('.')[2] as a string (false-pass for 1.1.9) — parse all segments as Number.

* chore(#2765): backfill changeset PR number (2888)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 20:12:28 -04:00