Commit Graph

1811 Commits

Author SHA1 Message Date
Tom Boucher
cf2e66b39e feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten

Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): backgroundDispatch citations in matrix + CONTEXT note

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1708): address review findings on typed dispatch-flatten

Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): backfill backgroundDispatch in role:runtime test fixtures

Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate

fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1708): add changeset for typed dispatch-flatten

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1708): remove stray temp PR-body file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1708): add issue ref to bug-853 allow-test-rule annotations

ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 15:37:33 -04:00
Tom Boucher
481d121dd4 feat(#1711): no-posix-mode-bit-assert AST rule (Phase 2) (#1718)
Phase 2 of epic #1702. Closes #1711.
2026-06-25 15:26:23 -04:00
Tom Boucher
a72dbfa58c feat(#1707): no-path-literal-in-assert AST rule + portability foundation (#1710)
Phase 1 of epic #1702. Closes #1707.
2026-06-25 14:37:10 -04:00
Tom Boucher
fb5f89db10 feat(#1704): destSubpath write-confinement (ADR-1239 Phase B) (#1706)
* feat(#1679): confine install writes within configHome

ADR-1239 Phase B write-confinement: a pure assertDestWithinConfigHome(configDir, destSubpath) rejects a destSubpath that escapes configHome (path traversal / NUL byte) at plan-build time on BOTH the install and uninstall plan paths; surface.applySurface and installOpencodeFamilySkills route through it, and _copyStaged carries a defense-in-depth containment check. Security-load-bearing for the Phase C third-party-descriptor loader.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1704): add changeset for destSubpath write-confinement

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1704): fix windows path-portability in confinement test

The N1 'accepts a true child subpath' assertion compared against path.join (no drive resolution) while the helper uses path.resolve — on Windows that mismatches the C: drive prefix. Compute the expected via path.resolve to mirror the helper. Windows-CI-only failure (local gsd-test is Mac+Linux).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 13:20:25 -04:00
Tom Boucher
30d4b85de5 feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module

ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1684): validate and document host-integration axes (16 runtimes)

Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add host-integration capability matrix and adr amendment

New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1684): harden dispatch negotiation edge cases

Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1684): register host-integration.cjs in lint-ignore and inventory

New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add changeset fragment for host-integration interface

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1684): add how-to for sourcing a host's integration axes

Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 11:33:06 -04:00
Tom Boucher
2b215b4163 feat(#1688): warn on stale model bake for static-frontmatter runtimes (#1692)
* docs(#1650): fix stale opencode install-path claim in core settings

* feat(#1688): warn on stale model bake for static-frontmatter runtimes

* chore(#1688): backfill changeset pr field with real PR number

* test(#1688): make resolveAgentDir assertions use path.join for windows

* docs(#1688): codify windows path-literal-in-assert anti-pattern + align test
2026-06-25 09:12:40 -04:00
Tom Boucher
c438fd396a test(#1676): fast-check property tests for path-prefix collapse + rewrite idempotency (#1686)
Delivers the property coverage promised in #1511's test scope but not landed
(follow-up #1676, epic #1507 / ADR-1508). Adds
tests/enh-1676-path-prefix-collapse-idempotency.property.test.cjs covering:

  (A) $HOME-collapse invariant for _computePathPrefix — global-under-home
      projects to $HOME/<suffix>/ (exact equality, not substring, so short
      homes like /root or /a do not false-positive); opencode is the
      documented exception (absolute form, never $HOME).
  (B) backslash->posix invariance (#1615 Windows path-leak fix).
  (C) path-rewrite idempotency for _applyRuntimeRewrites across the
      path-rewriting runtimes (f(f(c)) === f(c); attribution held at
      undefined to isolate the path axis). Non-vacuous: a sanity assertion
      proves the first pass actually rewrites the seed ~/.claude/ refs.

Pure test addition — no production code changed. Uses the shared
fast-check-setup (seed=42, numRuns=200). Closes #1676
2026-06-24 21:34:56 -04:00
Tom Boucher
466a2f2866 refactor(#1675): dedup augment converter family — single-source from conversion module (#1685)
bin/install.js held byte-identical duplicate definitions of the augment
converter family (convertSlashCommandsToAugmentSkillMentions,
convertClaudeToAugmentMarkdown, getAugmentSkillAdapterHeader,
convertClaudeCommandToAugmentSkill, convertClaudeAgentToAugmentAgent) that
already exist canonically in src/runtime-artifact-conversion.cts (generated
to gsd-core/bin/lib/runtime-artifact-conversion.cjs). Deferred Phase 1->2
cleanup tracked in #1675 (epic #1507 / ADR-1508).

Deleted the five local copies; install.js now binds the three PUBLIC
converters from runtimeArtifactConversion (same pattern as getDirName /
processAttribution in #1510). The two private helpers live only in the
conversion module now. module.exports preserved (re-exported).

Behavior-preserving: four converters byte-identical; the fifth
(convertClaudeAgentToAugmentAgent) differed only by an inert let->const
(variable never reassigned). Extends the DEFECT.GENERATIVE-FIX
reference-identity parity guard in enh-1511 to assert single-sourcing.

Closes #1675
2026-06-24 21:34:53 -04:00
Behruz Nassre Esfahani
47906b052d fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) (#1550)
* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final)

BSD/macOS mktemp only substitutes the XXXXXX template when it is the final
path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md`
return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent
workflow runs collide on the same temp manifest/body file — one run can
overwrite or consume another's. Reproduced on macOS: the second call to the
suffixed template fails `mkstemp: File exists`.

Fix: use a suffixless `XXXXXX` template (so it IS the final component), then
rename to add the intended extension — portable across BSD + GNU userlands,
no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at
every site.

Affected workflow temp files:
- execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest)
- quick.md:         gsd-quick-worktree-*.json
- spec-phase.md:    edge-probe-reqs-*.json
- ship.md:          gsd-pr-body-*.md
- profile-user.md:  gsd-profile-answers-*.json, gsd-profile-analysis-*.json

The execute-phase.md edit uses a compact intermediate var + trailing comment
to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the
workflow size baseline accordingly. Validated on macOS: 20 concurrent calls
yield 20 unique randomized paths.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): add changeset fragment (Fixed)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix

Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp
template whose XXXXXX run is followed by a filename suffix (the BSD/macOS
non-randomizing form). Fails on the six pre-fix instances and passes on
the fix, and locks the copy-paste-prone idiom out of future workflows.
Mirrors the bug-637 hardcoded-$HOME workflow guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): rename regression test to fix- prefix (regression-test-names lint)

New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names
ratchet; use the fix- prefix (matches the fix-1445 precedent).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint)

lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to
carry a #NNN reference (don't allowlist). Add (#1520) to the source-text
exemption.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1520): abort touched mktemp chains on failure (|| exit 1)

Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's
suggested failure guard. If mktemp fails, $VAR is empty and the subsequent
mv/write lands on an unintended relative path. Add `|| exit 1` to all six
touched chains so a mktemp failure aborts the snippet. Regenerated the
workflow size baseline for the slightly longer lines.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): rebase onto next — regen size baseline + describe rename

Resolve the workflow-size-baseline.json conflict from next advancing by
regenerating from the current workflow sizes. Also rename the test describe
from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit)

The two profile-user.md temp sites this PR already rewrites kept a hardcoded
/tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for
consistency and macOS-correctness (some sandboxes have no writable /tmp).
Regenerated the workflow size baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1520): regen size baseline after rebase onto next

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:27:45 -04:00
Andreas Brauchli
cbf7c82841 feat(#323): fish-shell support in post-install PATH suggestion (#727)
* feat(#323): fish-shell support in post-install PATH suggestion

Two additive changes to the post-install PATH-suggestion seam, both scoped
to existing functions.

A. Projection: add a fish entry to the persist-mode shell-action list in
   projectPathActionProjection() (src/shell-command-projection.cts). fish has
   no `export`/`$PATH`-list syntax, so the existing zsh/bash `export PATH=...`
   commands are inert when pasted. The new entry emits the fish-native
   `fish_add_path '<dir>'` (fish 3.2+, persists via the universal-variable
   store, de-duplicating). The directory is single-quoted with the same POSIX
   literal escaping as the zsh/bash siblings; verified round-tripping through
   real fish 3.7.0 for paths containing quotes, spaces, `$`, `*`, backticks
   and unicode.

B. Detection: add homePathCoveredByFishConfig() in bin/install.js, called
   from maybeSuggestPathExport() alongside homePathCoveredByRc(). fish does
   not use sh-style `export PATH=` rc files, so a fish user whose
   fish_user_paths already covers the global bin would otherwise get a
   false-positive "not on your PATH" warning on every install. Two
   side-effect-free detection routes (no fish subprocess):

   1. The universal-variable store (~/.config/fish/fish_variables). fish
      serializes this with `full_escape`: every byte outside [A-Za-z0-9/_]
      becomes `\xHH` (space -> \x20, `-` -> \x2d, `.` -> \x2e, `$` -> \x24,
      unicode -> \uXXXX) and list elements are joined by the literal 4-char
      token `\x1e` (NOT a raw 0x1e byte). The detector splits on `\x1e`,
      decodes the escapes, then compares each as an absolute literal — a
      decoded `$` is part of the directory name, not an unexpanded variable.
      Verified against real fish 3.7.0 output.
   2. config.fish (`fish_add_path`, `set -gx PATH`, `set -Ux fish_user_paths`)
      — plain shell tokens: HOME forms ($HOME/${HOME}/~) are expanded and a
      token still holding `$` (e.g. `$PATH`, `$fish_user_paths`) is skipped.

   Honours $XDG_CONFIG_HOME and always also checks ~/.config/fish.

No behaviour change for bash/zsh/PowerShell/cmd/Git-Bash users: their entries
and command strings are unchanged; the fish entry is additive and the fish
detector only narrows the set of cases that warn.

Tests: update the projection length assertion (2 -> 3) and fish escaping in
bug-3441; add fish detection + suppression cases in install-path-detection
(uvar store with real fish escaping, dot/hyphen/space/$-literal decode
regressions, config.fish routes, commented-out, relative-segment guard,
unreadable-file fault injection, suppression and emission via
maybeSuggestPathExport).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(changeset): add Changed fragment for #323 fish PATH support (#727)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#323): address review — action-only fish docs, decoder property test, win32 guard

Addresses @trek-e's review on #727:

- docs (blocker): keep the how-to action-only (Diátaxis). Drop the
  `# fish — persists via …` comment and the internal-mechanism clause
  naming fish_variables/config.fish; leave one command + the exec-fish
  directive.
- tests (minor): extract decodeFishUniversalValue to a pure, exported
  module function and add fast-check round-trip properties
  (decode(fishEscape(p)) === p over arbitrary unicode, abs-path variant,
  totality). Consolidated into install-path-detection.test.cjs to respect
  the install test-file-count ratchet.
- tests (follow-up): port #721's win32 negative-projection test (no fish
  action on win32; persist projection is PowerShell/cmd.exe/Git Bash).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#323): address review — drop unused 'after' import, clarify escaping comment

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:20:52 -04:00
Alex V.
1a46109b97 enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb

Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb
(coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain;
gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt.
Non-breaking — additive only.

arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR).

* fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline

- C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives)
- C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws
- C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface)
- C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access)
- inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows)
- size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step)
- eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage)

* fix(#1579): register eval family in alias-drift gates

Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs
families and to familyArrayKeys in the manifest-coverage test, so the eval
family lands under the same drift guard as every sibling family
(state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review.

check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10.

* docs(#1579): use half-open verdict band ranges in CLI-TOOLS

overall_score is fractional and thresholds are >=80/>=60/>=40, so a score
in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel
bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit.

* fix(#1579): validate eval.score CLI inputs

Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior.
2026-06-24 17:16:13 -04:00
Alex V.
a63684c222 enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking

Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).

arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).

* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized

- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
  circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
  in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
  false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
  drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
  the noted pre-existing follow-up.

* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate

The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.

Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.

* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer

trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
 - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
 - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
   document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.

Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.

* docs(#1577): document security.injection_blocking + boundary seam

trek-e Major 2 + Minor:
 - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
   the Full Schema and a Security Settings subsection, distinguishing it from
   the workflow.security_* namespace; honest circuit-breaker-not-redactor
   framing matching ADR-1577 / security-model.
 - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.

Verified: lint:docs ok; config-field-docs + contributor-standards green.

* test(#1577): make read-injection property test git-text, not binary

trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.

Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.

* docs(#1577): align untrusted boundary docs

Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.

* docs(#1577): align ADR ingest agent count

Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 17:07:23 -04:00
Behruz Nassre Esfahani
3870fafe74 fix(#1571): resolve schema-drift phase by token, not substring (#1640)
* fix(#1571): resolve schema-drift phase by token, not substring

verify schema-drift <phase> resolved the phase directory with a naive
entry.name.includes(phaseArg) test, so a non-existent phase could
silently match a different phase whose directory name merely contained
the requested token (e.g. "1" matched "11-expansion"), running the drift
gate against the wrong phase. Use the canonical phaseTokenMatches +
normalizePhaseName, matching find-phase, verify phase-completeness, and
this file's own unstarted-phase check.

Regression coverage folded into tests/schema-drift.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1571): add changeset for schema-drift token-match fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 17:03:01 -04:00
Tom Boucher
51c5497460 fix(#1668): restore CRLF By-Phase row-upsert regression (resolved via #1655) (#1669)
#1668 (phase complete did not persist the By-Phase row on a CRLF STATE.md) no longer
reproduces on next: #1655's restructure of updatePerformanceMetricsSection (By-Phase table
upsert now runs BEFORE the velocity manipulation) resolved it as a side effect. Verified
clean on next (CRLF STATE.md -> phase complete -> row upserted, placeholder removed,
velocity derived). Restore the #1658 test's full row-upsert + placeholder assertions that
#1667 had relaxed (they now pass), locking the fix.
2026-06-24 16:50:50 -04:00
Tom Boucher
d101daff30 fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt (#1654)
* fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt

Both onboarding paths (Step 2a auto-mode + Step 5 interactive) enumerated only
4 profiles (Balanced/Quality/Budget/Inherit), omitting 'adaptive' even though the
model catalog (model-catalog.json profiles) and docs/CONFIGURATION.md register 5.
Mirrors the proven /gsd:settings two-question split (#3784): Q1 routes between
Adaptive/Standard-tier/Inherit; Q2 (conditional on Q1=Standard) picks
Quality/Balanced/Budget — keeping every AskUserQuestion within the 4-option cap.
Both config-new-project example payloads now list adaptive. Regression cases
folded into the owning tests/new-project-mvp-prompt.test.cjs (per the
lint-regression-test-names ban on new top-level bug-NNNN files): each AI Models
prompt makes adaptive reachable, all 5 profiles reachable, 4-option cap honored,
both example enums include adaptive, brace balance. Workflow size baseline bumped
(new-project.md 62324 -> 66138 bytes; still well under the XL hard cap).

* chore(#1516): backfill changeset pr ref to 1654
2026-06-24 14:46:44 -04:00
Tom Boucher
c583bcc02c fix(#1659): dedup By-Phase rows across padded/unpadded phase numbers (#1663)
* fix(#1659): dedup By-Phase rows across padded/unpadded phase numbers

phaseRowPattern matched the phase number literally (escapeRegex(String(phaseNum))),
so a seeded zero-padded row '| 05 |' was not matched by 'phase complete 5' (pattern '| 5 |'),
producing a duplicate row that double-counted the phase. Canonicalize a numeric phase to
its integer form (Number('05')===Number('5')===5) and match with a 0* prefix so 5/05/005
all collapse to the same row in either direction. Regression folded into state.test.cjs:
seeded '| 05 |' + 'phase complete 5' yields exactly one phase-5 row. Non-numeric phase IDs
retain the literal escapeRegex match.

* chore(#1659): backfill changeset pr ref to 1663

* fix(#1659): add verification fixture to padded-dedup test under #1522 gate
2026-06-24 14:41:45 -04:00
Tom Boucher
ff161f2281 fix(#1582): derive phase-complete velocity from By-Phase table (idempotent) (#1655)
* fix(#1582): derive phase-complete velocity from By-Phase table (idempotent)

updatePerformanceMetricsSection blind-added summaryCount onto the prior velocity
total on every phase complete, so re-running phase complete on an already-complete
phase incremented the total each time (the sibling of #4, which fixed the Completed
Phases counter the same way). The velocity total is now derived as the sum of the
By-Phase table's Plans column AFTER the row upsert — re-completing a phase upserts
the same row, so the sum is stable; a hand-edited inflated total self-heals downward
to the true sum on the next completion. When the By-Phase table is absent the total
is left unchanged (no crash). Strengthens the misnamed 'idempotent' test (its comment
explicitly declined to assert velocity idempotency — the latent gap) and adds a
self-heal regression; corrects the #320 behavior-lock velocity assertion which had
encoded the blind-add (3 = 1+2 double-count) — the derived value is 2.

* chore(#1582): backfill changeset pr ref to 1655

* fix(#1582): velocity sum tolerates indented By-Phase rows (codex review)

Adversarial review (codex, gpt-5.5/high) flagged that byPhaseTablePattern's
data-row capture allows leading whitespace ([ \t]*\|), but the derive sum was
anchored at ^\| and would skip indented hand-edited/legacy rows — capturing them
in the table but silently undercounting. Align the sum regex (^\s*\|) with the
table capture's tolerance. Adds an indented-row regression. Two other codex
findings are pre-existing and out of scope: padded/unpadded phase dedup
(phaseRowPattern, identical in old code — derive yields the same value as the old
blind-add) and CRLF tables (the shared byPhaseTablePattern header requires bare
\n, so the upsert was already broken on CRLF; the fix changes stale-vs-
double-count, does not worsen it).

* fix(#1582): add verification fixtures to velocity tests under #1522 gate

Post-rebase onto next+#1548, the #1582 velocity tests (self-heal, indented-row) use
phase complete, which now fail-closes under #1522's canonical verification gate without a
passed *-VERIFICATION.md. Add writePassedVerification(tmpDir,'02-next','02') to both.
2026-06-24 14:36:01 -04:00
Tom Boucher
f202d243cb fix(#1658): #1658 test — add verification fixture + assert velocity (CRLF downstream bug filed) (#1667)
Under #1522's verification gate (#1548), phase complete fail-closes without a passed
VERIFICATION.md, so the #1658 test's fixture needed one. The By-Phase *row* upsert on CRLF
has a separate downstream bug (phase complete updates velocity + status but doesn't persist
the table row on CRLF, while it does on LF; byPhaseTablePattern matches CRLF — verified
directly), tracked separately. Relax the assertion to phase-complete-succeeds + velocity-
updates, proving CRLF STATE.md is processed end-to-end.
2026-06-24 14:21:37 -04:00
Tom Boucher
80607bec93 fix(#1658): make byPhaseTablePattern CRLF-tolerant on STATE.md tables (#1662)
* fix(#1658): make byPhaseTablePattern CRLF-tolerant on STATE.md tables

byPhaseTablePattern required a bare \n after the header and separator rows, so a
STATE.md with CRLF (\r\n) line endings (Windows, or hand-edited) had its By-Phase
table treated as absent: phase complete never upserted the row (and the velocity-from-
table derivation went stale). Make the header/separator terminators and the closing
lookahead CRLF-tolerant ([ \t]*\r?\n, (?=\r?\n|$)). Backward-compatible with LF.
Regression folded into tests/state.test.cjs: phase complete on a CRLF STATE.md upserts
the row and removes the placeholder. CONTRIBUTING's QA matrix lists Mixed CRLF/LF as a
required parser case.

* chore(#1658): backfill changeset pr ref to 1662
2026-06-24 13:52:47 -04:00
Tom Boucher
752df8adb4 fix(#1657): recover malformed (non-object) ~/.gsd/defaults.json in finishInstall (#1661)
* fix(#1657): recover malformed (non-object) ~/.gsd/defaults.json in finishInstall

JSON.parse of defaults.json succeeds for valid-JSON-but-non-object values (null, [],
42, "str"), which then bypassed the parse catch: null threw a TypeError on property
access (swallowed by the outer try/catch), and array/number/string had resolve_model_ids
set on a non-object whose JSON.stringify round-trip kept the broken shape. The non-Claude
finishInstall step now resets any non-object (null, non-object, or array) parse result to
{} before reading/writing, so the file is repaired and resolve_model_ids defaults normally.
Regression folded into the owning tests/bug-410-install-defaults-test-mode-guard.test.cjs
(parameterized over null/[]/42/"str").

* chore(#1657): backfill changeset pr ref to 1661
2026-06-24 13:48:16 -04:00
Tom Boucher
e1d768dd78 fix(#1660): fail-closed frontmatter set of object-list fields instead of silent no-op (#1664)
* fix(#1660): fail-closed frontmatter set of object-list fields instead of silent no-op

cmdFrontmatterSet reported {updated:true} even when spliceFrontmatter returned the
content unchanged, which happened whenever the new value's extractFrontmatter projection
equalled the original's — notably for object-list fields like must_haves, whose
{path,provides} items flatten to scalar strings under the lossy parser. Detect a no-op
(newContent === content) for a dict-valued field and surface an error directing the user
to edit the file directly, instead of silently accepting a no-op set. Scalars and scalar
arrays round-trip faithfully, so idempotent sets of those are intentionally NOT flagged
(two precision regression tests lock this). Folded into frontmatter-cli.test.cjs.

* chore(#1660): backfill changeset pr ref to 1664

* refactor(#1660): extract noOpObjectListSetError as pure tested helper (Stryker coverage)

cmdFrontmatterSet is not in Stryker's property/unit test set, so the inline no-op
detection added survivors that dropped the frontmatter module below its 62% mutation
threshold. Extract the detection into a pure exported helper noOpObjectListSetError and
unit-test every branch directly (changed content, scalar, scalar-array, null, dict
no-op). cmdFrontmatterSet now calls the helper. Same pattern as the #1572 spliceFrontmatter
coverage fix.
2026-06-24 13:39:17 -04:00
Tom Boucher
f615eb9ef3 fix(#1572): preserve must_haves object-lists across frontmatter set/merge (#1656)
* fix(#1572): preserve must_haves object-lists across frontmatter set/merge

spliceFrontmatter round-tripped the WHOLE frontmatter through extractFrontmatter
(a scalar-only parser) then reconstructFrontmatter (a lossy serializer), so any
must_haves object-list — artifacts {path, provides}, prohibitions {statement,
status} — was flattened to scalar strings and re-emitted as a malformed inline
array whenever an UNRELATED field changed, silently dropping every provides:/
status: value. The write now preserves the original raw text for any top-level
key whose value is structurally unchanged between the original parse and the new
object (generalizing the existing whole-document no-op guard to per-key
fidelity), and regenerates only the key that actually changed. The key set is
still defined by newObj (the cmdSet/cmdMerge flow always passes the full merged
object). spliceFrontmatter's only callers are cmdFrontmatterSet/Merge — the
STATE.md read-modify-write family calls reconstructFrontmatter directly and is
unaffected. Regression cases folded into tests/frontmatter-cli.test.cjs:
artifacts/prohibitions object-lists survive set and merge; idempotent on repeat
sets. Asserted via parseMustHavesBlock (the structure-preserving parser).

* chore(#1572): backfill changeset pr ref to 1656

* fix(#1572): fail-closed when set/merge would emit [object Object] (codex review)

Adversarial review (codex, gpt-5.5/high) flagged that directly setting a must_haves
object-list (a CHANGED key) still routed through the lossy reconstructFrontmatter,
emitting literal "[object Object]" and destroying the data. The reported case
(mutating an UNRELATED field) was already fixed by per-key raw-text preservation,
but the changed-object-list path was still silently lossy. Add fail-closed: when a
regenerated key's text contains the "[object Object]" sentinel, spliceFrontmatter
throws — cmdFrontmatterSet/Merge error out WITHOUT writing, directing the user to
edit the file directly. The no-frontmatter (generate-from-scratch) path is guarded
the same way. Adds a test that a refused set leaves the file unchanged and the
original object-list intact. Codex finding #2 (a contrived flattened-projection
no-op) is a deeper limitation noted in the PR — non-destructive, and the fail-closed
message already directs users to edit object-list blocks directly.

* test(#1572): add spliceFrontmatter per-key preservation + fail-closed unit coverage

Stryker mutates gsd-core/bin/lib/frontmatter.cjs against tests/frontmatter.{property,unit}.test.cjs
(MinScore 62). The #1572 regression cases live in frontmatter-cli.test.cjs, which is NOT in
Stryker's test set, so the new functions (sliceTopLevelFrontmatterSegments, the per-key
preserve/regenerate/drop/append loop, regenerateFrontmatterKey's [object Object] fail-closed)
had surviving mutants that dropped the module below threshold. Add unit-level coverage in
frontmatter.unit.test.cjs exercising every new branch directly via spliceFrontmatter:
unchanged object-list preserved (provides survives) when a scalar sibling changes; changed
scalar regenerates only that key; orphan keys dropped; new keys appended; indented nested
block stays attached to its parent key; whole-document no-op returns input verbatim; both
fail-closed paths (changed object-list + no-frontmatter) throw.
2026-06-24 13:30:46 -04:00
Tom Boucher
b205e4c2b2 fix(#1639): parseDecisions handles the titled-colon bullet form (#1665)
* fix(#1639): parseDecisions handles titled-colon bullet form

bulletColonRe anchors on ':**' (colon immediately before close-bold) and bulletEmDashRe
requires an em-dash, so the titled-colon form '- **D-NN: Title.** body' (title between the
colon and the closing **) matched neither and was dropped by the parse-miss guard. When all
decisions used the titled convention, parseDecisions returned 0 and check.decision-coverage-
plan passed vacuously — the same false-coverage failure mode as #1343/#1364/#1365. Add a
third per-form regex bulletTitledColonRe, checked LAST (strict superset of bulletColonRe,
so it only catches bullets the other two miss — minimal blast radius); id + [tags]
trackability honored. Regression folded into decisions.test.cjs: titled-colon parses,
coexists with colon/em-dash, tags, all-titled-13 no longer vacuously 0.

* fix(#1639): tighten titled-colon title to [^:*]* so malformed pre-colon-run bullets still reject

The first cut's title run [^*]* was too permissive: it matched a genuinely-malformed
bullet with a colon in the pre-separator freeform run (e.g. 'D-07 ratio 3:1:**') by
treating the 3:1 colon as the separator, regressing the #1343 parse-miss guard tests.
Tighten the title to [^:*]* (no colon, no star) so the separator colon remains the only
colon permitted before ** — matching bulletColonRe's existing [^:*]* discipline. Valid
titled forms (colon-free titles) still parse; the malformed colon-in-freeform case still
falls through to the parse-miss guard.

* chore(#1639): backfill changeset pr ref to 1665
2026-06-24 13:26:25 -04:00
Tom Boucher
4631923982 fix(#1569): preserve explicit resolve_model_ids in non-Claude installs (#1653)
* fix(#1569): preserve explicit resolve_model_ids in non-Claude installs

The non-Claude finishInstall step keyed its resolve_model_ids:"omit" write on
!== "omit", so an explicit true opt-in (resolveModelInternal returns full model
IDs) was silently clobbered on every install/upgrade across all 14 non-Claude
runtimes, making generated agent manifests inherit the active chat model instead
of pinning the resolved model. Now only absent/falsy is defaulted to "omit"; an
explicit true (and an existing "omit") is preserved. Regression test
parameterizes across codex/opencode/gemini and covers the absent/false/idempotent/
claude/malformed boundaries.

* chore(#1569): backfill changeset pr ref to 1653

* fix(#1569): default non-canonical resolve_model_ids values to omit (codex review)

Adversarial review (codex, gpt-5.5/high) flagged that the original allowlist-by-
enumeration condition (undefined/null/false -> omit) preserved malformed values
(0, "", "yes", {}) instead of defaulting them to omit, letting them leak Claude
aliases a non-Claude runtime cannot resolve. Switch to an allowlist condition
(existing !== true && existing !== 'omit') so only an explicit canonical true
opt-in and an existing omit are preserved; everything else defaults to the safe
non-Claude omit. Adds a parameterized test over [0, "", "yes", {}].
2026-06-24 13:21:19 -04:00
Jeremy McSpadden
77c7b4fc9d fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition

* no-mistakes(review): Fix canonical verification closeout gates

* no-mistakes(review): Fix verify-work frontmatter promotion command

* no-mistakes(review): Fix stale verification gates

* no-mistakes(review): Fix canonical verification routing gates

* no-mistakes(review): Fix verification dependency and runtime routing gates

* no-mistakes(review): Block stale verification bypasses

* fix: handle large init manager outputs in verification workflows

* chore: update changeset pr number

* fix(verify-work): use fresh verification.status for stale gate

The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.

* fix(init): skip roadmap-checked phases when selecting next_phase

Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.

* fix: gaps_found not overridden by stale, transition uses canonical verification

- verification.cts: check gaps_found before stale so gap-closure routing
  is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
  already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
  to avoid false-positive blocks from body text matching

* ci: retrigger tests after rebase

* fix(transition): replace gsd_run advisory check with awk frontmatter extraction

The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.

Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.

The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.

Also update workflow-size-baseline.json for the updated transition.md size.

Fixes: runtime-launcher-parity test (B)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: re-check verification under planning lock in phase complete

Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.

* fix(transition): gate on canonical verification.status including stale

Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).

* Fix workflow verification gates for yolo transition and stale routing

Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.

* fix(transition): use verification.status query for stale-aware advisory check

The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.

Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.

Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.

Also update workflow-size-baseline.json for the updated transition.md size.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* ci: trigger test matrix for 525b946

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(transition): restore awk frontmatter extraction for pre-shim verification check

The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(#1522): clarify transition verification gate wording

* fix(#1522): update transition workflow size baseline

* fix(#1522): update workflow-size-baseline after rebase onto next

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)

Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 13:19:10 -04:00
Tom Boucher
6214039358 refactor(#1644): Hub extension — exitReason? field on InvalidArgs + adapter honestification (#1645)
Phase 1 of parent #1641. Implements the contract documented in the
Phase 0 ADR-0174 §5 amendment (#1642 / #1643).

src/command-routing-hub.cts
  * InvalidArgsResult interface gains optional exitReason?: string
    (carries an ERROR_REASON enum value, separate from reason which is
    the explanation text).
  * makeInvalidArgs(arg, reason, exitReason?) factory conditionally adds
    the field only when the third arg is truthy — preserves the strict-
    keys invariant tested at command-routing-hub.test.cjs:444.
  * _VARIANT_SCHEMA.InvalidArgs.allowed Set extended to include
    'exitReason' so the runtime validator does not coerce well-formed
    extended Results to HandlerFailure.

src/cjs-command-router-adapter.cts
  * Honestified the wrapper comment: the runtime check ('ok' in result)
    already passes any {ok:*} object through, so the historical
    {ok:true, data} return type was a lie for err Results. The lying
    cast is preserved because the Hub's export = syntax doesn't expose
    HubResult for import; the Hub's _validateErrResult runtime-validates
    the actual shape.
  * Result→error() translation branched: when InvalidArgs carries
    exitReason, the adapter calls error(result.reason, result.exitReason)
    so the JSON-error envelope (GSD_JSON_ERRORS=1) preserves the typed
    ERROR_REASON value. When exitReason is absent, error(msg) is called
    with exactly one arg — byte-identical with prior behavior.
  * RouteCjsCommandFamilyOptions.error and RouteHubCommandFamilyOptions
    .error callback types widened from (message) to (message, reason?)
    to match io.cts's actual error() signature.

CONTEXT.md
  * Command Routing Hub predicate updated to document the new field,
    factory signature, and dispatcher translation contract.

Tests (TDD red→green)
  * tests/command-routing-hub.test.cjs: 8 new tests covering 2-arg
    (strict-keys), 3-arg (key present), undefined, empty string, frozen
    result, hub.dispatch propagation, and validator acceptance.
  * tests/cjs-command-router-adapter.test.cjs: 2 new tests covering
    exitReason passed as second arg + byte-identical prior behavior when
    absent.

Verification
  * npm run test:unit: 2448 tests, 0 fail (no regressions)
  * gsd-test-summary on docker: outcome=passed, 0 failures
    (RULESET.PR-FLOW.docker-before-push)

Memtrace blast radius: LOW (get_impact makeInvalidArgs → 3 nodes; the
optional field is non-breaking for the 1 existing caller routePhaseCommand).
2026-06-23 22:47:57 -04:00
Tom Boucher
bcc5a6d1ba fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command

Capability hook install (applyCapabilitySharedEdits) wrote each settings.json
hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a
fail-closed guard could then block the whole session), and emitted a bare
single-quoted script path so a .js-family hook from a git/tarball source
without +x failed with Permission denied on every matching call.

- Pass through an optional declared `matcher` (entry-level sibling of `hooks`);
  absent => omitted (match-all), so existing shipped capabilities are unchanged.
- Validate `matcher` in the declaration (non-empty string, no control chars).
- Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party);
  .sh and others keep the bare quoted path (unchanged).

Root cause: the manifest hook schema (validator rule C4) was {event, script}
only with no matcher, and applyCapabilitySharedEdits never read or wrote one;
the command used shellSingleQuote(absScript) with no node prefix.

Regression tests fail-first on both defects (matcher dropped; bare path) and
pass after the fix; #1460 command assertions updated for the node prefix.

* chore(#1634): backfill changeset pr:1638

* fix(#1634): resolve lint and windows CI failures

- validator: replace the control-character range regex with a char-code loop.
  The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char
  codes are equally precise and lint-clean. Behavior unchanged (still rejects
  matchers containing ASCII control characters incl. DEL).
- test: gate the executable-bit precondition on POSIX. Windows fs does not
  honor POSIX write modes (a 0o644 write reads back as 0o666), so the
  precondition is meaningless there and failed the windows-latest lane. The
  node-prefix assertion — the actual fix — is platform-independent and still
  runs everywhere.

* docs(#1634): amend ADR-894 for optional lifecycle hook matcher

The `role: "feature"` `hooks[]` entry now carries an optional `matcher`
(settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document
the field in the §2 schema table and record a Grilling-amendments entry:
the install path projects a declared matcher onto the emitted settings.json
hook entry (absent = match-all, so shipped capabilities are unchanged), and
per-runtime matcher projection (ADR-857 D8) stays a separate concern. This
amendment ships with the fix that introduced the field rather than as a
follow-up.

* docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md

Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a
test that writes a file with a POSIX mode and then asserts statSync().mode
& 0o777 === <octal> passes on macOS/Linux but fails on windows-latest
(Windows fs does not honor POSIX write modes — reads back 0o666). Added as a
machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/
prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward:
gate the mode-bit precondition on process.platform !== 'win32' and keep the
platform-independent behavioral assertion running everywhere.
2026-06-23 21:49:43 -04:00
Behruz Nassre Esfahani
fc5ca178a2 test(#1178): consolidate duplicated agent-roster helper into tests/helpers (#1420)
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers

The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated
across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and
inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting
listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites
through it. Semantically-different sites (installed-dest dirs, absolute-path
returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform
multi-family inventory table) are left intact, each with a one-line comment.

Test-only; no production code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1178): note AGENTS_DIR export is for future call sites

Review nit: clarify that the currently-unused AGENTS_DIR export is intentional
— available for future tests needing the canonical source agents path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:34:08 -04:00
Joe Seymour
9d12725e4e fix(#1619): normalize pruned mise node execPath to the stable shim in normalizeNodePath (#1621)
* fix(#1619): normalize pruned mise node execPath to the stable shim

resolveNodeRunner() bakes process.execPath into managed .js hook commands.
Node realpaths execPath, so under mise it resolves to a concrete
<data>/installs/node/<ver>/bin/node that mise prunes on `mise up`, after
which every managed hook 404s — the same ephemeral-path failure #977 fixed
for fnm and #3181 for Homebrew. normalizeNodePath now rewrites a mise
versioned install path to the stable sibling shim <data>/shims/node when it
exists (deriving <data> from execPath so a custom MISE_DATA_DIR works),
falling back to the raw execPath otherwise. Tests folded into
install.test.cjs per the regression test-name lint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(changeset): set pr number to 1621

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Joe Seymour <joese@iarx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:17:03 -04:00
Tom Boucher
207d8f1697 fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).

- planner: add a Severity column to the STRIDE threat register; assign
  severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
  severity enum; redefine threats_open as the count of OPEN threats whose
  severity is at or above block_on (none => 0). Below-threshold opens are
  reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.

No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 19:04:13 -04:00
Tom Boucher
f9d9dfb4bc fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.

- New reference gsd-core/references/security-asvs-levels.md defines L1
  (opportunistic), L2 (standard), L3 (comprehensive) for both planner
  threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
  hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
  L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
  by extracting the goal-backward worked example to planner-guidance.md.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 18:48:58 -04:00
Tom Boucher
94be6d5b60 fix(#1625): resolve security config in secure-phase.md before auditor handoff (#1633) 2026-06-23 17:24:59 -04:00
Tom Boucher
01ef08ed2a Merge pull request #1632 from open-gsd/fix/1628-config-set-security-enum-validation 2026-06-23 17:24:23 -04:00
Tom Boucher
0c4d570541 fix(#1628): type-safe config-set validation — close JSON-coercion enum bypass + enforce capability schema
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently':

1. Missing guards: workflow.security_block_on (enum) and
   workflow.security_asvs_level (integer 1-3) had no store-time validation.

2. Systemic JSON-coercion bypass: every string-enum guard used
   VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed
   before validation, String(["member"]) === "member" let a JSON array
   slip through and an array was stored in a scalar key. Reproduced on
   human_verify_mode, statusline.context_position, context_guard_mode,
   fallow.scope/profile, source_grounding_authority, drift_action, context.

3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum,
   25 boolean, 2 number, 1 string) had no hardcoded guard, so any value —
   including coerced arrays/objects and out-of-enum strings like
   code_review_depth=garbage — was stored silently.

Fix: a type-safe assertEnumValue() helper (requires typeof === 'string'
before membership), routed through all nine central string-enum guards
(messages preserved byte-for-byte); plus a generic capability-registry
validation block that validates every capability key against its declared
type/values (enum via the registry's values — single source of truth —
boolean, number, string). Behavioral regression tests cover every central
enum key and representative capability keys (array + object coercion
rejected, out-of-enum rejected, valid accepted) with boundary coverage for
the security keys.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-23 16:10:34 -04:00
Tom Boucher
db2b4d326a fix(#1629): cleanup legacy .devin/skills/gsd- dirs on Windsurf reinstall
Pre-#1615 Windsurf installs wrote skills under .devin/skills/gsd-*/ (Devin Desktop preferred dir, #1085). PR #1615 moved Windsurf to .windsurf/workflows/ but never cleaned up the old layout. Users upgrading from a pre-#1615 install were left with dead .devin/skills/gsd-* directories that nothing reads anymore.

Fix: added cleanupWindsurfLegacyDevinSkills() which mirrors the Codex cleanupCodexSkillMetadataSidecars() pattern. Runs on Windsurf local install, removes GSD-managed .devin/skills/gsd-* dirs, preserves user content (non-gsd- dirs, gsd-dev-preferences per #2973, symlinks). Empty .devin/ and .devin/skills/ containers are pruned; non-empty ones are left intact.

5 regression tests: removes gsd-* dirs; preserves user content; skips symlinks (escape guard); no-op when absent; end-to-end install removes pre-staged legacy artifacts.

Refs #1629 (Finding B; Finding A addressed in #1630).
2026-06-23 15:54:07 -04:00
Tom Boucher
b65939cdff fix(#1629): copy Windsurf command bodies so workflow delegation targets exist
PR #1622 (issue #1615) shipped Windsurf /gsd-* workflow wrappers that delegate to command bodies at <targetDir>/.windsurf/gsd-core/commands/gsd/X.md via a hardcoded @~/.claude/gsd-core/commands/gsd/ path. The path-rewrite pipeline correctly substitutes ~/.claude/ to the install target. But the source gsd-core/ dir does not ship with commands/ — the canonical command source lives at the package root (commands/gsd/). Without this copy, every /gsd-* workflow in Cascade references a file that does not exist. The slash commands appear in the / menu but silently fail when invoked because the LLM is told to read a missing file.

None of the original reviews caught this: not the security review, not Codex's adversarial orthogonal review (gpt-5.5/high), not Memtrace's graph-backed review. It was surfaced by a #1629 regression test that verifies 'every workflow @-reference target exists on disk after install' — the test failed, revealing the bug.

Fix: for Windsurf local installs, copy commands/gsd/*.md into <targetDir>/gsd-core/commands/gsd/ via copyWithPathReplacement (applies the same path+brand rewrites as the rest of the install). Guarded on isWindsurf && !isGlobal since global Windsurf workflow install is an explicit no-op.

Documented as DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED in CONTEXT.md so the pattern is locked in: any new converter emitting a wrapper that delegates to another file MUST verify the delegation target is actually installed.
2026-06-23 15:28:54 -04:00
Tom Boucher
658ea33cb6 fix(#1615): applySurface rewrites commands kind, not just skills
Codex adversarial orthogonal review of PR #1622 surfaced that applySurface (src/surface.cts) only called rewriteStagedSkillBodies for kind='skills', skipping kind='commands'. The gap meant /gsd-surface profile changes on any runtime with commands kinds (windsurf, opencode, kilo, cursor, augment, codebuddy, gemini) wrote raw @~/.claude/... references into synced command/workflow bodies, which fail at invocation time on non-Claude runtimes.

For Windsurf specifically, this left workflow files containing @~/.claude/gsd-core/commands/gsd/X.md after a profile change — paths that don't exist on a Windsurf install. Verified by the new regression test which fails before the fix (workflow bodies contained @~/.claude/) and passes after (workflow bodies reference the install target).

Captures the return value of rewriteStagedCommandBodies (temp dir path — commands rewrite uses copy-then-rewrite to avoid mutating the package source), syncs from the temp dir, then cleans up. Type annotations satisfy typescript-eslint strict mode.

Findings 2 (install ordering) and 3 (legacy .devin cleanup) from the same review are tracked in #1629 — both real but out of scope for #1615.
2026-06-23 14:52:32 -04:00
Tom Boucher
4ed208e74b fix(#1615): validate commandName to prevent workflow prompt injection
Codex peer review of PR #1622 surfaced that convertClaudeCommandToWindsurfWorkflow interpolated commandName unsanitized into a markdown body that Windsurf loads as an LLM-readable workflow. A plugin author who controls a commands/gsd/*.md filename could inject newlines, markdown structure, or path components (..) to manipulate the workflow body.

Validate commandName at function entry against /^(?:gsd-)?[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/ — rejects slashes, backslashes, spaces, dots, control chars, trailing dash. Pattern requires alphanumeric ending so gsd- alone (which would slice to empty stem) is also rejected. Throws with a JSON.stringify-escaped preview (no literal newlines in the error message).

Applied to both bin/install.js (where tests import from) and src/runtime-artifact-conversion.cts (production source). 18 positive + 22 negative test cases lock in the validation.
2026-06-23 14:38:38 -04:00
Tom Boucher
527142ad2e fix(#1615): normalize Windows backslash paths in workflow content
computePathPrefix returned a Windows-style path (with backslashes from path.join) into markdown @-references. Workflow file content on Windows ended up with mixed separators, breaking substring checks in install/install-runtime-artifacts tests on windows-latest CI only.

Normalize resolvedTarget and homeDir to forward slashes inside computePathPrefix. The prefix is always substituted into markdown body text, which uses POSIX paths universally. Idempotent on POSIX.

Also normalizes the two test assertions to forward-slash form so they pass on Windows. Adds a regression test for backslash-style input.

Documents DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT + RULESET.CONTENT-PATH-NORMALIZATION in CONTEXT.md so this anti-pattern stops recurring.
2026-06-23 14:26:44 -04:00
Tom Boucher
6d782e309d test(#1615): update Windsurf workflow expectations 2026-06-23 12:51:46 -04:00
Tom Boucher
fc2a7c0555 fix(#1615): install Windsurf slash workflows 2026-06-23 12:10:21 -04:00
Tom Boucher
c28cccbf85 Merge pull request #1595 from davesienkowski/feat/1592-plan-drift-precheck
feat(#1592): add plan:pre codebase-drift pre-check before planner runs
2026-06-23 10:56:29 -04:00
Tom Boucher
1b95762661 Merge pull request #1568 from behruznassre/fix/1514-retired-phase-total-phases
fix(#1514): exclude retired/folded phases from progress.total_phases
2026-06-23 10:55:41 -04:00
Tom Boucher
ce7fffffc4 test(#1614): classify Antigravity skills as flat 2026-06-23 10:40:55 -04:00
Tom Boucher
da2de3a183 Merge branch 'next' into fix/1514-retired-phase-total-phases 2026-06-23 10:22:18 -04:00
Tom Boucher
cbd21092a9 fix(#1614): install Antigravity skills flat 2026-06-23 10:21:06 -04:00
Tom Boucher
ba96c70b14 feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.

- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
  block (extractFrontmatter can't — its `-` items are scalars-only; this is a
  focused parser, sibling of parseMustHavesBlock), validates each entry, and
  classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
  typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
  AND non-empty all-`pass` verification AND zero validation errors. Everything
  else — judgment, empty/failing verification, malformed entry — routes to the
  human (fail-safe). A malformed block falls back to legacy prose extraction and
  surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
  human_judgment:true); verify-work extract_tests consumes it; create_uat_file
  marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
  eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
  USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
  for the null-entry/comment-header/mis-indent cases found in adversarial review.

Closes #1602

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Rezolv
652142521b enhance(#1549): validate PR-title issue-ref convention at open time (#1576)
* enhance(#1549): validate PR-title issue-ref convention at open time

The release changelog is title-driven: release.yml generates "What's Changed"
from PR titles, then format-github-release-notes.cjs buckets each line by its
conventional-commit prefix and relies on a `(#<issue>)` in the title to render
the issue link. Both rules were enforced only socially, so titles like
`fix(core): ...` (no issue link) and `[security] fix(...): ...` (leading tag
defeats the `^fix` bucket anchor -> mis-filed under Enhancement) silently broke
the changelog, landing on the maintainer as release-time cleanup.

Extract the title matcher into one shared module consumed by BOTH the changelog
classifier and a new PR-title CI gate, so a title that passes the gate cannot
mis-bucket in the changelog (single source of truth).

- scripts/lib/conventional-title.cjs (new): classifyBucket + evaluatePrTitle +
  the anchored regexes. One matcher, two consumers.
- scripts/release-notes/format-github-release-notes.cjs: classifyTitle now
  delegates to classifyBucket (behavior preserved; existing tests green).
- .github/workflows/pr-title-validator.yml (new): runs evaluatePrTitle on
  pull_request opened/edited/reopened/synchronize, for ALL authors (the drift
  came from member PRs). Trusted base-ref checkout; WARN_ONLY knob for rollout.
- tests/conventional-title.test.cjs (new): bucket + gate cases incl. the
  leading-tag mis-bucket (backfills the untested classifyTitle case) and a
  cross-check that the classifier delegates to the shared matcher.
- CONTRIBUTING.md: document the `type(#<issue>):` rule and no-leading-tag.

Claude-Session: https://claude.ai/code/session_01UMV5Qr3H4oFikbuiEauGQk

* fix(#1549): check out the PR in pr-title-validator so the new matcher resolves

The workflow checked out the base branch (next) as a trusted policy source, but
the shared matcher (scripts/lib/conventional-title.cjs) is introduced by this PR
and does not exist on next yet — so require() failed and validate-title errored
on its own introducing PR. Check out the PR's merge ref instead: the matcher
under review is present, the check is self-consistent, and a fork pull_request
runs read-only with no secrets, so running the PR's own pure-string regex is
safe.

* fix(#1549): move conventional-title.cjs out of installed scripts/lib/

bin/install.js bundles every file under scripts/lib/ into the user-installed
payload (the changeset CLI's dependencies), and install.test.cjs (#935) asserts
that exact set. The new matcher is release/CI tooling that must NOT ship to
users, so placing it in scripts/lib/ both broke the install manifest test and
would have shipped dead code. Relocate it next to its consumer in
scripts/release-notes/ (which the installer does not copy) and update the three
require paths (classifier, workflow, test) + the CONTRIBUTING reference.

install.test.cjs now 125/125; conventional-title + release-notes suites green;
lint:ci clean.

* fix(#1549): load title matcher from trusted base ref, not PR code

Addresses review (Solvely-Colin + trek-e): the gate checked out the PR
merge ref and require()'d evaluatePrTitle from PR-controlled code, so any
future PR could edit conventional-title.cjs to return { valid: true } and
wave its own malformed title through — a self-bypassable required check.

Load the matcher from a base-branch checkout instead (ref:
github.event.pull_request.base.ref), the same trusted-policy-source pattern
pr-target-validator.yml already uses. The PR can change its title but not
the ruler that measures it. An existsSync bootstrap guard skips the check
when the matcher isn't on the base branch yet (the introducing PR); every
PR after merge is fully gated. This keeps the single shared matcher (#1549's
whole point) rather than forking the regex into the workflow.

Also per review:
- add tests/conventional-title.property.test.cjs (fast-check): any
  `type(#n): summary` round-trips to valid; evaluatePrTitle/classifyBucket
  are total functions (never throw).
- pin the `fix(#):` zero-digit boundary as missing-issue-ref.

Claude-Session: https://claude.ai/code/session_01VqUHNQCh71pEqjo96zkgQL

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 22:44:44 -04:00
Dave
d1f7ba82f2 feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.

Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
  CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
  form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
  registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
  via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).

Closes #1592

Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
2026-06-22 20:30:59 -04:00
Behruz Nassre Esfahani
0271231910 test(#1514): add fast-check property + all-retired boundary for retired parser
Per review (test-standard items):
- Property test (RULESET.TESTS.property-based-testing): extractRetiredPhaseNumbers
  is the parsing core, so add a fast-check property — k of n checklist phases
  struck → exactly the k canonical keys returned, across randomized phase counts
  and numeric/zero-padded/project-code ID forms. Exposed via a `_`-prefixed test
  seam (mirrors the existing _setLockProbes seams), no public API surface added.
- Boundary (RULESET.TESTS.boundary-coverage): all-retired case (k === n) →
  total_phases 0, via state json.

No production behavior change; the exclusion logic is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 17:22:29 -07:00