Commit Graph

3624 Commits

Author SHA1 Message Date
Dave
8f428bdd41 feat(1259-01): deterministic prohibition-enforcement producer + check route
- Author src/prohibition-enforcement.cts: the test-tier prohibition PRODUCER/gate (ADR-550 D5d heavy half)
  locate wired check (node-test|lint-rule) -> confirm fail-first -> run -> build enforcementEvidence -> dispositionForProhibition
  pure/deterministic with injectable runCheck; missing/failing/non-fail-first -> hard-gate (both modes); passing -> green
- Route check prohibition-enforcement in src/check-command-router.cts (same family as ui-plan-gate / tdd-review-checkpoint)
- Add tests/prohibition-enforcement.test.cjs (behavioral, typed-field, injected runner) — new module within <=2 budget
- Register the built bin/lib surface: docs/INVENTORY.md row + regenerated docs/INVENTORY-MANIFEST.json
- .gitignore: add the emitted gsd-core/bin/lib/prohibition-enforcement.cjs (build artifact, ADR-457)
- No src/probe-core.cts edit — the green/fail-closed policy seam already exists
2026-06-15 12:44:29 -04:00
Dave
6115ab216d test(1259-01): RED — test-tier enforcement (both check kinds) before producer exists
- Extend prohibition-probe.verify-tier.test.cjs with the ENFORCEMENT half (#1259, ADR-550 D5d)
- Require the not-yet-built gsd-core/bin/lib/prohibition-enforcement.cjs (RED)
- Cover both wired-check kinds (node-test + no-source-grep lint-rule) and miss/fail hard-gate
- Typed-field assertions only; injected runCheck (no real subprocess)
- Keep the 2 original fail-closed tests verbatim
2026-06-15 12:38:48 -04:00
Tom Boucher
b370787155 Merge pull request #1262 from open-gsd/chore/sync-next-version-1.5.0-rc.4
chore: sync next package version to 1.5.0-rc.4
2026-06-15 01:07:49 -04:00
github-actions[bot]
04e4d631c4 chore: sync next package version to 1.5.0-rc.4 2026-06-15 05:07:43 +00:00
Tom Boucher
cf68841220 enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents

- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#1243): document plugin-provided skills in agent_skills

Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.

Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)

- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1243): regenerate agent-size baseline for the Skill-tool grant

The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).

* chore(#1243): add Added changeset fragment

* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 00:49:12 -04:00
Tom Boucher
00c447701d fix(#1257): update pipe-table Status/Phase/Plan cells in planned-phase + begin-phase (#1260)
* fix(#1257): update pipe-table Status/Phase/Plan cells in planned-phase + begin-phase

cmdStatePlannedPhase ran its body-field replacements on the full file content,
so the case-insensitive ^Status: pattern matched the YAML frontmatter `status:`
line before the body `| Status | … |` cell — the cell never advanced to
'Ready to execute' and syncStateFrontmatter re-derived the stale 'planning'
status (the #1230 delta heuristic preserved it). cmdStateBeginPhase had
pipe-table else-branches only for Status/Last activity (#1256), so the Current
Position `| Phase |` / `| Plan |` cells were left stale while a spurious inline
`Phase: N — EXECUTING` line was prepended.

Both handlers now strip frontmatter before body-field replacement and update the
pipe-table cells in place via stateReplaceField, matching inline-format
behaviour. Systemic residual of #1255 / #1256.

Adds 4 regression tests (#1257 block in tests/state.test.cjs) covering both
findings — RED before the fix, GREEN after; full state-area suite (437) and
local unit suite stay green.

Closes #1257

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1257): backfill changeset pr number (#1260)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 00:29:18 -04:00
Tom Boucher
2668fbbeb0 fix(#1255): advance frontmatter status for pipe-table STATE.md on begin/complete-phase (#1256)
* fix(#1255): advance frontmatter status for pipe-table STATE.md on begin/complete-phase

state begin-phase/complete-phase called stateReplaceField on the FULL file content, so its case-insensitive ^Status: plain-pattern matched the YAML frontmatter status: line first (no g flag) and never updated the body pipe-table Status cell; syncStateFrontmatter then re-derived the stale status from the unchanged body, freezing the frontmatter status. Fix: strip frontmatter before the body-field replacements (operate on body only), reassemble with frontmatter preserved, so the body Status cell updates and the frontmatter derives correctly — for inline AND pipe-table body formats. Also corrects the Current Position pipe-table else-branches (Status/Phase/Last-activity) to write bare, consistent cell values. Pipe-table Status is a supported body format (not rewritten to inline).

Closes #1255

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1255): add changeset for pipe-table state status fix (#1256)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1255): fold pipe-table regression into state.test.cjs + Windows-portable frontmatter regex

Per the 2026-06 audit, new tests/bug-NNNN-*.test.cjs files are banned (lint-regression-test-names) — folded the 7 #1255 regressions into tests/state.test.cjs and removed the standalone file + its lint-test-file-count allowlist entry. Also fixed the frontmatter assertions' /^---\n/ anchors to /^---\r?\n/ (windows-test-parity-guard frontmatterAnchorLiteralNewline).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 22:42:28 -04:00
Rezolv
3556450b0d feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no
shape cue matched, no `shapes` override) as a single soft `unclassified — review
manually` candidate instead of silently dropping it — the exact blind spot the
probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays
silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the
candidate is left `unresolved`, never auto-`backstop` (a missing shape is not
evidence an edge exists).

Closes #1110
2026-06-14 21:44:21 -04:00
Rezolv
395fb519e7 feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644.
2026-06-14 21:29:11 -04:00
Tom Boucher
9a4a448c27 docs(#1235): ADR — migrate agent conversion to the descriptor-driven install path (#1253)
Records the design gate for moving agent conversion off the inline bin/install.js loop onto the descriptor path (ADR-3660): the two-(really three-)path problem, the 10 parity behaviors the descriptor agents path must gain (verified against code — the issue's 7 plus Qwen/Hermes branding, the Codex TOML sidecar, and stale-agent cleanup), an AgentConverterContext contract to fix the leaky (content)=>string converter signature, and an incremental per-runtime cutover gated on byte-for-byte golden parity (full + minimal mode). Status: Proposed. Codex-reviewed for technical accuracy.

Closes #1235

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:54:20 -04:00
Tom Boucher
b783410815 refactor(#1191): inject clock/reset testability seams + handle valid-null settings (#1233)
* refactor(#1191): inject clock/reset testability seams + handle valid-null settings

- worktree-safety reapOrphanWorktrees: injectable deps.nowMs clock for deterministic stale-lock boundary tests (mirrors snapshotWorktreeInventory's options.nowMs).

- active-workstream-store: _resetControllingTtyCacheForTests() seam clears the memoized controlling-TTY probe cache; test replaces require.cache busting.

- gen-capability-registry: export stripGeneratedComment (additive); test imports the real helper + equivalence assertion, keeping the deliberate drift oracle.

- install.js readSettings: a successfully-parsed JSON null is treated as empty settings ({}) instead of being mis-reported as malformed; genuine parse failures still warn. readSettings/stripJsonComments exported (GSD_TEST_MODE-guarded require) for real behavioral tests.

Closes #1191

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1191): add changeset for valid-null settings fix (#1233)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1191): replace Stryker-incompatible structural reset test with behavioral isTTY-spy

The seam-2 reset test read the BUILT active-workstream-store.cjs and grepped for 'didProbeControllingTtyToken = false' — Stryker instruments that file so the literal is absent, failing the mutation DRY RUN. Replaced with a behavioral test that spies on process.stdin.isTTY access count to prove a post-reset probe re-runs (kills the didProbe-reset mutant) without reading source text. Local stryker: dry run passes, score 85.21% >= 80.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:56 -04:00
Tom Boucher
55eb1fd72e refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam (#1242)
* refactor(#1190): extract ADR-22 drift-guard decision logic into a testable seam

ADR-22's severity mapping, authority auto-upgrade, and rung>=3 hard-block lived only as prose in plan-review-convergence.md — untestable. Extracted into src/plan-drift-guard.cts (pure: AUTHORITY_RUNGS, getEffectiveAuthority, classifyDriftSeverity) + a gsd-tools drift-guard CLI seam (authority/severity), and rewired the workflow to call the seam deterministically instead of reasoning the decision in prose. 47 unit/e2e/structural tests cover the full severity table, the grep->intel auto-upgrade, and the rung>=3 HIGH hard-block.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for ADR-22 drift-guard seam (#1242)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): register ADR-22 module in eslint-ignore + inventory manifest + docs-exempt changeset

Full-matrix CI surfaced new-module/command governance ripples beyond the lint-tests chain: (1) tsc-generated plan-drift-guard.cjs must be in the eslint ignore list (551-eslint-bin-lib-coverage); (2) docs/INVENTORY-MANIFEST.json must include the new module/command (regen via gen-inventory-manifest.cjs --write); (3) a type:Added changeset triggers docs-required — added a docs-exempt marker (internal seam, no user-facing surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:36 -04:00
Tom Boucher
6ef78fe417 feat(#1188): branch-coverage floors + promote no-source-grep to error on tests/** (#1250)
* feat(#1188): promote no-source-grep to error on tests/**

Apply local/no-source-grep as an error to tests/**/*.test.cjs (was warn on bin/scripts only). Only 2 real violations surfaced (repo-layout.test.cjs structural guard-placement checks on bin/install.js) — marked allow-test-rule with #1188 reason. PART A (branch-coverage floor) follows separately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1188): add c8 branch-coverage floors (60 global, 70 per-file on UNMUTATED modules)

c8 enforced line floors only. Add a --branches 60 global floor (current ~82.5%, margin mirrors the lines-70-vs-91% gap) to test:coverage + test:coverage:unit, plus a chained 'c8 check-coverage --per-file --branches 70' for the high-risk UNMUTATED modules state/phase/verify/init (current min 78.3% on verify). The per-file check reuses the coverage data the suite run just produced (the proven scripts-floor pattern) so it adds no second suite run.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1188): per-module branch check must pin lines/funcs/stmts to 0

Standalone 'c8 check-coverage' defaults unspecified metrics to 90 and enforces them, so --branches 70 alone also failed on lines (verify.cjs 85.29% < default 90). Pin --lines 0 --functions 0 --statements 0 so only the branch floor (70) is enforced on the UNMUTATED modules. Data-reuse confirmed working (it computed verify's real %).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 20:52:13 -04:00
Tom Boucher
6724581285 fix(#1230): preserve frontmatter status/stopped_at when a state write doesn't change the body source (#1252)
* fix(#1230): preserve frontmatter status/stopped_at when a state write doesn't change the body source

Adds a delta heuristic in `readModifyWriteStateMd`: snapshots the body
Status and Stopped At fields before the transform, then after
syncStateFrontmatter runs, restores the existing frontmatter values for
any field whose body source was not changed by this write. Integrates
cleanly with the existing !resync progress-restore block (computes postFm
once, applies both restorations, reconstructs frontmatter once).

Regression tests added to bug-905 test file (same theme, no new file).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1230): add Fixed changeset fragment

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 20:51:53 -04:00
Tom Boucher
8da48e22ae fix(#1229): count bullet-only phases so phase.add stops reusing an existing phase number (#1249)
* fix(#1229): count bullet-only phases + guard against number collision in phase.add

Before this fix, the phase.add number scan only checked ### Phase N: section
headers and on-disk phases/N-* directories. A phase that existed only as a
roadmap bullet (e.g. "- [ ] **Phase 11: ...**") was invisible to both scans,
causing phase.add to silently assign a duplicate number.

Fix: add a bullet-entry regex scan (all checkbox variants: [ ], [x], [~], with
or without ** bold markers) to the set-based phase-number collection in
cmdPhaseAdd. Also added a post-compute collision guard that advances the
candidate past any already-used number.

Regression tests added to tests/phase.test.cjs (bug #1229 describe block):
bullet-only collision, [x]/[~] variants, plain-bullet, and baseline preservation.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1229): add Fixed changeset fragment

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-14 17:52:15 -04:00
Tom Boucher
f52a7a5f77 feat(#1245): add capability ecosystem ADR, PRD, and developer documentation (#1248)
Phase 0 of the Capability Ecosystem epic (#1244): the design record and the
third-party-author documentation set, with no runtime or code changes.

- docs/adr/1244-capability-ecosystem.md — architecture decision record
  (amends/extends ADR-857 Decisions 7 & 8)
- docs/prd/1244-capability-ecosystem.md — product requirements
- Diataxis docs: tutorial, how-to (publish/import/version/remove),
  reference (manifest schema, /gsd:capability command, capability matrix),
  explanation (trust model); cross-links added to develop-a-capability.md

Refs #1244

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 17:51:43 -04:00
Tom Boucher
10ae85cbbf fix(#1238): keep node resolvable in bug-891 PATH isolation so home-fallback subtests don't false-fail when node co-locates with a global gsd-tools shim (#1247) 2026-06-14 17:18:17 -04:00
Tom Boucher
fbf5d5eab6 docs(#1239): ADR — GSD as an embeddable orchestration engine (#1241)
* docs(#1239): ADR for GSD as an embeddable orchestration engine

Records the design to invert GSD from a standalone installer that projects
onto a host into an embeddable orchestration engine a host loads as a plugin,
driven through a negotiated host-integration interface. Unifies ADR-1016
projection (the declarative adapter) with imperative embedding behind one
contract; grounded in a 9-host capability survey.

Closes #1239

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#1239): de-slash illustrative command placeholders (docs-parity gate)

Replace illustrative /gsd:x /gsd-x /gsd.x placeholders with namespace-prefix
wording so the docs-parity live-registry check does not parse them as
non-live commands.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 17:16:40 -04:00
Tom Boucher
1fa7bc594c refactor(#1190): extract ADR-230 PR-target branch policy into a tested, fork-safe seam (#1246)
ADR-230's branching-model gate (pr-target-validator.yml) decided allowed/blocked PR targets via inline regex in github-script — untestable. Extracted the decision into committed scripts/pr-target-policy.cjs (classifyPrTarget(base,head)->{decision}), and rewired the workflow to checkout the BASE ref (trusted; fork-tamper-safe) + require the module. Behavior-identical (Codex-verified char-by-char regex equivalence + all side-effects preserved). 70 tests incl. an equivalence oracle battery + hyphen-boundary negatives. Added contents:read for the checkout. Re-attribution: no ADR-230 test references exist (issue's '2 misattributed files' claim not borne out).

Closes #1190

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 17:16:19 -04:00
Tom Boucher
00acbc8868 fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads (#1240)
* fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads

Before this fix, bin/install.js copied scripts/changeset/ and scripts/lib/
into the runtime config dir but omitted scripts/fix-slash-commands.cjs.
gsd-core/bin/lib/command-roster.cjs requires this file at module load via
require('../../../scripts/fix-slash-commands.cjs'), so every gsd-tools
command crashed with MODULE_NOT_FOUND on every installed runtime.

Four changes:
- bin/install.js copy step: copy fix-slash-commands.cjs into <configDir>/scripts/
  with source-missing hard-fail and verifyFileInstalled smoke check
- bin/install.js writeManifest: track scripts/fix-slash-commands.cjs (not
  covered by the changeset/lib subdir loops)
- bin/install.js uninstall: best-effort unlinkSync before scripts/ rmdir
- scripts/fix-slash-commands.cjs readCmdNames(): wrap readdirSync in
  try/catch returning [] so skill-based/global installs without a local
  commands/gsd/ directory do not throw ENOENT

Tests added to tests/install.test.cjs (6 new tests):
- smoke: install() copies fix-slash-commands.cjs
- e2e: spawned gsd-tools.cjs does not crash with MODULE_NOT_FOUND
- manifest: writeManifest() tracks the file
- uninstall: uninstall() removes the file
- readCmdNames unit: export returns an array
- readCmdNames spawn: absent COMMANDS_DIR returns exit 0 (no throw)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1223): backfill changeset PR number (#1240)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 16:03:26 -04:00
Tom Boucher
0b3a2e5f9c feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15)

ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test.

Closes #1190

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1190): add changeset for progress --converge surface (#1237)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline

CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:52:22 -04:00
Tom Boucher
e97b5d6b7b fix(#1217): bound acquireStateLock retries on recoverable errno (no busy-spin) (#1236)
* fix(#1217): bound acquireStateLock retries on recoverable errno (no busy-spin)

The recoverable-errno branch (`ACQUIRE_LOCK_RETRY_ERRNOS`) previously called
`continue` directly, skipping both `clock.sleep()` and the 30 000 ms budget
check. A permanently-failing ENOENT (e.g. parent dir removed) would spin at
100% CPU forever with the event loop fully blocked, making the OS-level
`timeout` the only escape.

Fix: extract a `checkBudgetAndSleep(context)` helper and call it from BOTH
the recoverable-errno path and the EEXIST contention path so every retry is
bounded and backed off identically.

Regression tests added to tests/clock-seam.test.cjs:
- persistent ENOENT throws budget-exceeded error (clock must advance via sleep)
- transient ENOENT (2 retries then success) acquires lock normally

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1217): backfill changeset PR number (#1236)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:10:59 -04:00
Tom Boucher
cafb874c4a fix(#1224): accept --pr 0 placeholder at changeset creation (#1231)
* fix(#1224): accept --pr 0 placeholder at changeset creation

The required-field guard `!opts.pr` treated the integer 0 as falsy,
rejecting the documented `pr: 0` two-push placeholder with a usage
error (exit 2). Non-numeric `--pr abc` (NaN) was also silently
accepted before (passes `!NaN === true`... actually `!NaN` is true, so
NaN would trigger the guard already). The new explicit checks use
`opts.pr === null` for missing flag and `Number.isNaN` for non-numeric
input, accepting all finite integer values including 0.

The merge-time safety net in parse.cjs (`pr <= 0` → INVALID_PR) is
unchanged — a pr:0 fragment is still rejected at lint/render time.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1224): backfill changeset PR number (#1231)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 15:02:54 -04:00
Tom Boucher
15fbdf6879 chore(#1175): remove dead tool-name converters + fix unused-arg lint (#1234)
Removes three unreferenced functions from bin/install.js:
- convertCursorToolName (zero callers across src/, bin/, scripts/, tests/)
- convertWindsurfToolName (zero callers across src/, bin/, scripts/, tests/)
- convertAugmentToolName (zero callers across src/, bin/, scripts/, tests/)

Dead code only — behavior-preserving. All three were superseded when
the conversion logic moved to src/runtime-artifact-conversion.cts.
lint and build both clean.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 14:52:25 -04:00
Tom Boucher
2ae5fcdcf8 test(#1189): cover ADR-0006 planningPaths() consumption in init handlers (#1226)
Add a black-box regression guard (8 tests in tests/init.test.cjs) asserting the init handlers (execute-phase, plan-phase, phase-op, milestone-op) resolve workstream-scoped planning paths under GSD_WORKSTREAM via planningPaths()/planningDir(), never the flat .planning form.

Each positive case asserts both the scoped value and not-equal-to-flat (genuine guard, not coverage credit); fixtures seeded workstream-scoped; every CLI call pins GSD_WORKSTREAM and GSD_PROJECT for hermeticity. Verified by mutation (flat join fails exactly the 4 positive assertions) and re-verified under a polluted parent env. Test-only; src/ unmodified.

Closes #1189

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 14:20:47 -04:00
Tom Boucher
73b7f45140 feat(#1173): wire agent converters into descriptor-driven install path (#1227)
Extends `dispatchKindEntry` in `runtime-artifact-layout.cts` to route
agents-kind entries through a converter when the descriptor carries a
non-null `converter` field. Adds `stageAgentsForRuntimeWithConverter`
to `install-profiles.cts`, expands `VALID_CONVERTER_NAMES` with the 9
agent converter names, and adds a fail-first behavioral test suite
(9 tests) proving the new wiring end-to-end.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 14:20:21 -04:00
Tom Boucher
bf634b95c3 feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver

Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the
`gsd-tools capability set` subcommand: the write-side inverse of the capability
resolver (ADR-1213). One desired capability state projects onto the substrates —
`enabled` drives the runtime surface (canonical on/off), `gates` drive federated
config keys (hook granularity), install profile is a read-only floor — then
re-resolves and reports divergence (assert-and-report), so "off means off" holds
as a write-time invariant. Adds batched setConfigValues; routes gsd:settings
capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to,
ADR-1213, CONTEXT.md term.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1213): add changeset for Capability State Writer (#1225)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 12:51:17 -04:00
Tom Boucher
22f56f4431 ci(#1212): shard windows full-test lane to remove timeout cliff (#1222)
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.

Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.

- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
  round-robin partition (fileIndex % n === i-1) over the SORTED selected file
  list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
  pure no-op. The 28K Windows argv chunking is preserved within each shard. A
  legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
  sharding still hits the discovery hard error. Composes with --suite and is
  order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
  cross-product (explicit include rows — a base shard dim does not cross-product
  with include legs, and a nested matrix.leg.os is unresolvable by the H1
  shell-policy linter). Unit suite runs sharded; integration/security run once
  per leg (shard 1). The Required tests fan-in is unchanged: it already needs
  test-full and checks the matrix-aggregate result, so a failed/cancelled shard
  fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
  balance, determinism, boundaries, fast-check property) + parseShardArg
  validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
  guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
  denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.

Closes #1212

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 12:25:29 -04:00
Tom Boucher
7edd18fd2b feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165.
2026-06-14 12:23:07 -04:00
Tom Boucher
82a0f561e1 fix(#1182): extract agent converters + tool-name tables into conversion module (#1220)
Extracts all 9 convertClaudeAgentTo* functions together with their full
dependency closure (claudeToCopilotTools, claudeToGeminiTools,
convertCopilotToolName, convertGeminiToolName) from bin/install.js into
src/runtime-artifact-conversion.cts and adds them to the module's export=
block.

Adds 6 regression tests in tests/copilot-install.test.cjs including a
DEFECT.GENERATIVE-FIX parity guard that asserts claudeToCopilotTools is
identical in module and bin/install.js.

Inline copies in bin/install.js are retained (#1175 will remove them).
This unblocks #1173 (descriptor-driven dispatch).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:44:42 -04:00
Tom Boucher
8c6a0d65e6 fix(#1203): gate free-form ROADMAP deprecation warning on phase_id_convention (#1218)
* fix(#1203): gate free-form ROADMAP deprecation warning on phase_id_convention

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number for #1218

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1203): use canonical gsd-tools roadmap upgrade command in warning

Match the migration command string to the canonical form used in
verify.cts and roadmap-command-router.cts (gsd-tools roadmap upgrade
--convention milestone-prefixed, dry-run by default) instead of the
non-canonical 'gsd roadmap upgrade --apply'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:44:36 -04:00
Tom Boucher
2e8f4f6de1 fix(#1202): make verify key-links wave-aware for planned future files (#1219)
* fix(#1202): make verify key-links wave-aware for planned future files

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number (#1219)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 11:28:14 -04:00
Tom Boucher
1a186013a4 fix(#1205): roadmapper applies phase_id_convention to generated phase IDs (#1215)
* fix(#1205): roadmapper applies phase_id_convention to generated phase IDs

- Add Phase ID Convention section to <phase_identification> block:
  documents sequential (default) vs milestone-prefixed forms, and
  instructs the agent to read phase_id_convention from config.json
- Update <output_formats> to show both header and checklist forms for
  sequential and milestone-prefixed conventions with examples
  (e.g. ### Phase 1-01: Name, - [ ] **Phase 1-01: Name**)
- Add TDD regression test tests/bug-1205-roadmapper-convention.test.cjs
  (5 assertions, confirmed fail-first then pass after fix)
- Update tests/agent-size-baseline.json to reflect legitimate growth
- Add .changeset/brave-otters-leap.md (Fixed, pr:0 placeholder)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset pr: 1215 for fix/1205

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1205): move phase_id_convention regression into roadmapper-granularity.test.cjs

lint-regression-test-names rejects new standalone bug-NNNN-*.test.cjs files;
regression cases must live in the owning module's test file.

Move the 5 phase_id_convention assertions (#1205 regression) from the
removed tests/bug-1205-roadmapper-convention.test.cjs into
tests/roadmapper-granularity.test.cjs as a new describe block, alongside
the existing granularity calibration tests. Also update the allow-test-rule
comment to cover both #163 and #1205 surface contracts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1205): fix lint-allow-test-rule-refs for roadmapper-granularity

- Add issue ref (see #1205) to allow-test-rule comment in
  tests/roadmapper-granularity.test.cjs so lint-allow-test-rule-refs
  passes (new exemptions require #NNN per ADR-456)
- Prune stale 'source-text-is-the-product' entry from
  scripts/lint-allow-test-rule-refs.allowlist.json (ratchet-down;
  comment now compliant and no longer needs grandfathering)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:46:22 -04:00
Tom Boucher
d444864bf8 fix(#1194): correct inverted statusline auto-compact buffer math (#1211)
* fix(#1194): correct inverted statusline auto-compact buffer math

The reserved-buffer percentage was computed as (acw/totalCtx)*100 — the
usable fraction — instead of (1 - acw/totalCtx)*100 — the reserved fraction.
When acw == totalCtx this produced buffer=100%, making the usable-range
denominator zero and pinning `used` at a constant 100% regardless of real
remaining context.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number (#1211)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:46:00 -04:00
Tom Boucher
e79e4e18b3 fix(#1186): guard hotfix version regex against leading zeros (ADR-218) (#1214)
Replace `^[0-9]+\.[0-9]+\.[1-9][0-9]*$` with
`^(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.[1-9][0-9]*$` in the hotfix
branch of the validate-version step, consistent with the two sibling
patterns already using the strict (0|[1-9][0-9]*) guard. Adds
regression assertions to tests/adr-218-release-version-validation.test.cjs
that prove `01.2.3` and `1.02.3` are rejected.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:45:42 -04:00
Tom Boucher
9e5d4b266b fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs (#1207)
* fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs via SessionStart hook

Claude Code marketplace plugin installs unpack the package into the
version-pinned plugin cache and never run bin/install.js, so
~/.claude/gsd-core/ is never created. Agents, commands, and templates
markdown-@-include the canonical ~/.claude/gsd-core/... path (which
expands ~ but NOT ${CLAUDE_PLUGIN_ROOT}), so every include resolved to
nothing and agents (e.g. the executor) failed.

Add a SessionStart hook (hooks/gsd-ensure-canonical-path.js) that, on a
plugin install, symlinks the canonical path's immutable subdirs (bin,
contexts, references, templates, workflows) to the plugin's bundled
gsd-core/ tree. It changes zero @-references, is a no-op in classic
installs, preserves user-generated files (USER-PROFILE.md, STATE.md),
prunes stale links so it self-heals after `claude plugin update`, uses
Windows junctions, and rejects bundled/canonical paths that escape the
resolved plugin root (no traversal, no clobber).

Registered in HOOKS_TO_COPY (build-hooks), MANAGED_HOOKS, hooks.json
SessionStart (runs first, timeout 5), and BUNDLED_GSD_HOOK_FILES.
Behavioral regression tests folded into issue-766-plugin-manifest.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#997): backfill changeset PR number to #1207

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 09:37:41 -04:00
Tom Boucher
cd638d5a71 docs(#1209): map ready-for-agent to confirmed-bug in triage-labels (#1210)
The triage-label mapping pointed ready-for-agent at the stale `confirmed`
label. The live verified-bug gate is `confirmed-bug` (RULESET.CONTRIB.CLASSIFY.fix
requires confirmed/confirmed-bug; 200+ issues and bug-remediation workflows use
confirmed-bug). Update the row + note so /triage applies the correct gate, and
mark `confirmed` as legacy/back-compat only.

Closes #1209

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 09:32:30 -04:00
Tom Boucher
98866a0c69 feat(#1187): per-module Stryker mutation-score ratchet (ADR-456 80% floor) (#1200)
* feat(#1187): per-module mutation-score ratchet + graduate core-utils

ADR-456's 80% mutation floor was unenforceable as a single global break=50:
4 of 6 covered modules sit at 63-79% and forcing them to 80 would require
brittle exact-string assertions on equivalent string-literal mutants (a
Goodhart's-Law trap). Instead, each covered module declares a minScore floor
(locked at its measured score, TARGET 80) enforced per CI shard via
stryker --break, ratcheting up over time without brittle tests.
- mutation-matrix.cjs: minScore per module + TARGET_MUTATION_SCORE=80, emitted
  in the matrix; require.main guard + exports for testability.
- mutation.yml: per-shard --break <minScore>.
- stryker.config.mjs: global break 50->60 as a local backstop (CI uses minScore).
- Graduated core-utils (measured 77.5%, floor 75).
- context-utilization 79.5->92.3% via behavioral killers (state classification
  outputs + error-value contract, not exact-string matches) -> minScore 80 (TARGET).
- ratchet-integrity guard test (28 cases).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): pass mutation break via MUTATION_BREAK env (no stryker --break flag)

Adversarial review caught that Stryker 9.x has no --break CLI flag, so the
per-shard 'stryker run --break <minScore>' errored out every mutation shard.
Read the per-module floor from process.env.MUTATION_BREAK in stryker.config.mjs
and set it per shard via env in mutation.yml. Red-green verified: MUTATION_BREAK=99
exits 1, =80 exits 0. Also make the ratchet guard monotonic (RATCHET_BASELINE
floors; lowering a floor now fails the guard unless the baseline is edited).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): fail closed on bad MUTATION_BREAK + monotonic ratchet baseline

Code review: Number(env)||60 failed OPEN — an empty/invalid MUTATION_BREAK
(e.g. a future module missing minScore -> matrix expands to '') silently
degraded the shard to break 60, letting a high-floor module regress undetected.
resolveMutationBreak() now returns 60 only when the env is truly unset (local
backstop) and THROWS on present-but-empty/non-numeric/out-of-range (fail closed);
stryker.config.mjs imports it via createRequire. Also make RATCHET_BASELINE an
equality mirror (=== not >=) so any floor change is explicit in review and no
floor can be silently lowered. Tests: 46 (incl resolveMutationBreak cases).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): recalibrate config-schema/prompt-budget floors to CI scores

First CI mutation run failed two shards: the floors were set from local Stryker
runs whose TIMEOUTS were counted as kills (env-variable), inflating scores. CI
runs with timeout~0, so the real deterministic scores are lower:
- config-schema: local 69.7% -> CI 54.55% (5 local timeouts vanished) -> floor 52
- prompt-budget: local 99.6% -> CI 68.33% (239 local timeouts vanished) -> floor 66
Calibrate floors from CI (the documented source of truth) and record the lesson
in the comment so future floors aren't set from timeout-inflated local runs.
Baseline updated to match. The other 5 shards passed (deterministic CI scores
above their floors).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 08:25:41 -04:00
Tom Boucher
f660a9e55a Merge pull request #1208 from open-gsd/chore/sync-next-version-1.5.0-rc.3
chore: sync next package version to 1.5.0-rc.3
2026-06-14 08:25:06 -04:00
github-actions[bot]
8a02b417c4 chore: sync next package version to 1.5.0-rc.3 2026-06-14 12:25:00 +00:00
Tom Boucher
fdac556746 fix(#1160): resolve capability surface from installed skill layouts (#1206)
* fix(#1160): resolve capability surface from installed skill layouts

In a global skills-runtime install (e.g. Codex at ~/.codex), gsd-tools.cjs
runs from <configDir>/gsd-core/bin/ and the commands/gsd source tree is
absent — only <configDir>/skills/gsd-<stem>/SKILL.md files exist.
_resolveCommandsGsdDir() returned a path that does not exist there, so
loadSkillsManifest returned an empty Map. resolveSurface then materialised
the '*' (full) profile sentinel by enumerating that empty manifest → empty
surfaced Set → every capability reported surfaced=false/enabled=false
regardless of project config. As a result `loop render-hooks verify:post`
returned activeHooks:[] even with workflow.security_enforcement and
workflow.nyquist_validation enabled, silently disabling the security and
Nyquist gates.

Fix: add _loadInstalledSkillsManifest(configDir) that scans configDir/skills/
for gsd-<stem>/SKILL.md dirs and builds the same Map shape, and
_resolveManifest(commandsGsdDir, configDir) that prefers the source tree when
present (preserving repo-checkout behaviour) and falls back to the installed
skills layout otherwise. Both resolveCapabilityRuntimeState call sites use
_resolveManifest. Both helpers are exported for direct unit-testing.

Tests: capability-state.test.cjs gains a faithful installed-runtime e2e block
that copies gsd-core/bin + scripts + package.json into a temp install root
with no reachable commands/gsd, then runs the real gsd-tools.cjs against an
installed skills/ layout. It asserts capability state reports security &
nyquist enabled and verify:post includes security->secure-phase and
nyquist->validate-phase; a disabled-config negative confirms no
over-activation. This block FAILS before the fix (activeHooks:[]) and PASSES
after. Plus unit coverage for the two new helpers and the empty-surface
pre-fix scenario.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 08:22:43 -04:00
Tom Boucher
e9f9ae49c8 fix(#1146): single base-branch resolver across forking workflows (#1198)
* fix(#1146): single base-branch resolver across forking workflows

Replaces duplicated per-workflow bash detection that silently fell through
to :-main on repos where origin/HEAD is unset (git init+remote add+fetch
without set-head, most CI checkouts, many worktrees).

New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch`
with full precedence ladder: git.base_branch config override → origin/HEAD
symref → git remote show origin (authoritative) → local branch presence →
"main". All git subprocesses bounded with timeouts; degrades gracefully.

Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the
single resolver. Removes 14 lines of duplicated detection bash across the
five workflows.

Includes 7 behavioral tests covering the full precedence ladder including the
key regression case (master repo, origin/HEAD unset → must return "master",
NOT "main") and an anti-regression guard that fails if any workflow
re-introduces the :-main/:-master fallback pattern.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number #1198

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1146): drop stray PR-body file from branch

pr-1146-body.md was committed during changeset backfill but must not
be tracked in the repo. Content preserved externally for PR body use.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): add tests for flat base_branch config key and both-branch tie-break

Closes two mutation gaps identified in adversarial review:
- A2: flat {base_branch: ...} at config root (legacy key form) was covered
  by code but unguarded against mutation of lines 74-75 in resolver
- H: tier-4 tie-break when both main+master exist locally (main wins,
  per tryLocalBranch JSDoc) was documented but untested

9/9 tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1146): allowlist workflow-literal guard as runtime-contract exemption

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks

handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted
and run verbatim by behavioral tests that lack the gsd_run preamble.
Adding a || fallback ladder (git symbolic-ref then echo main) keeps the
unified resolver as primary in real workflows while letting the test harness
succeed without gsd_run defined.

Also propagates updated runtime-launcher preamble to pr-branch.md (added in
origin/next MemPalace PR) and regenerates workflow-size-baseline.json.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 08:22:40 -04:00
Tom Boucher
a375c4b354 feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)

Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).

Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).

HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#956): backfill changeset PR number (#1201)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 02:09:10 -04:00
Tom Boucher
b1e8a74708 fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks

discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and
LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no
`loop render-hooks` dispatch and was absent from the conformance gate's
HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post.

- discuss-phase.md: add minimal discuss:pre (before analyze_phase) and
  discuss:post (after write_context) render-hooks dispatch steps that
  delegate consumption to a new shared reference (kept under the 32KB
  #2551 budget; no inline subagent dispatch token).
- references/loop-hook-dispatch.md: new canonical, point-agnostic contract
  for consuming `loop render-hooks --raw` activeHooks (contribution/step/
  gate) — single source for hook consumption across host loops.
- gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and
  export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host
  file) — one source of truth for the host-loop file + wired-point set.
- phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES
  and shared scanWiredPoints (was a hand-maintained duplicate omitting
  discuss-phase.md + a duplicated regex).
- gen-capability-registry.cjs: add validateHooksWired() gen-time guard that
  rejects a capability hook declared at a valid-but-unwired loop point, with
  a clear remediation message — failure now surfaces at gen --check/--write
  time instead of deep in the full conformance suite.
- tests (capability-registry.test.cjs): regression + anti-pattern parity
  guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES;
  POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into
  the discuss-class gap again.
- docs/INVENTORY*: register the new reference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1196): backfill changeset PR number (#1199)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 01:39:54 -04:00
Tom Boucher
be132445aa fix(#1159): phase complete ignores historical/deferred requirement metadata (#1197)
* fix(#1159): phase complete ignores historical/deferred requirement metadata

Defect A: `phase complete` emitted a false "has unresolved gaps" warning
when a VERIFICATION.md file's frontmatter contained `status: passed` but
the body contained `previous_status: gaps_found`. The full-text regex
`/status: gaps_found/` matched the substring inside `previous_status:
gaps_found`, producing a spurious warning. Fixed by reading only the
frontmatter `status` key via `extractFrontmatter` in both `phase.cts`
(phase complete warning check) and `commands.cts` (determinePhaseStatus).

Defect B: requirement IDs under explicitly deferred/backlog/future/v2
section headings in REQUIREMENTS.md were flagged as missing from the
Traceability table. Fixed by splitting the body into markdown sections,
detecting deferred-intent headings via a keyword regex, and skipping
those sections when collecting IDs to check against the table.

Both fixes include boundary tests (false positive suppressed; genuine
gap/missing warnings still fire) in 4-phase-complete-cjs-regression.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1159): address adversarial-review findings (subheading + case sensitivity)

- Defect B: Replace section-split approach with line-by-line depth-tracking
  to correctly propagate deferred status to sub-headings and ignore headings
  inside fenced code blocks. Previously `## Future Backlog` / `### Sub` would
  leak sub-heading IDs (split created a new non-deferred section per heading).
- Defect A/commands: restore case-insensitive status matching via toLowerCase()
  to match the prior /status:\s*passed/i regex semantics.
- Add test #1159-B-4 covering the nested-subheading case.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number for #1159 fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 00:00:37 -04:00
Tom Boucher
5fa4dcd78c fix: recover silently-excluded test dirs + test-architecture audit hardening (#1195)
* fix: recurse test discovery so subdir test suites actually run

scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.

Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: retire 5 verified-worthless tests

Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
  covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
  and bug-782-cline-skills-emission)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: add ADR-218 release version-validation coverage

ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: redesign weak tests into behavioral, deterministic assertions

Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
  bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
  plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
  context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
  core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
  collision) and unconditional plugin.json schema validation (issue-766)

Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add no-tautological-assert lint rule, error in test suite

New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).

Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gate new allow-test-rule exemptions to require an issue ref

ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: add ADR test-audit evidence report (#1192)

Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture

feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: address adversarial-review findings

Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
  capability-registry.cjs in place (concurrency hazard) — uses in-memory
  checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
  /* */ too, matching no-source-grep) so a block comment can't bypass it;
  one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
  rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
  assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address code-review findings (subdir discovery, rule + test gaps)

xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
  Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
  silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{}  equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
  writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)

The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: reconcile allow-test-rule allowlist after rebase onto next

Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:35:08 -04:00
Tom Boucher
3ebd13c45d fix(#1145): implement query user-story.validate handler (#1193)
* fix(#1145): implement query user-story.validate handler

`query user-story.validate` was a phantom command invoked by mvp-phase.md
(line 102) and verify-work.md (line 170) but had no CJS handler. Every
call exited with "Unknown command: user-story". The dotted-form dispatcher
strips the `query` prefix, splits on `.`, yielding command='user-story'
which fell through to the `default:` case with no registered capability
handler.

Adds a `case 'user-story':` handler inline in gsd-tools.cjs (same pattern
as the #1140 fix in PR #1148). The handler:

- Validates "As a [role], I want to [capability], so that [outcome]."
- Uses \S anchors to require non-whitespace content in each slot
  (whitespace-only slots like "As a  , I want to ..." now correctly
  return valid:false — found by adversarial Codex review)
- Returns { valid: boolean, errors: string[], slots: {role, capability,
  outcome} | null }
- Supports --pick valid for bare boolean output (verify-work.md usage)
- Added 'user-story' to SKIP_ROOT_RESOLUTION (pure string validation)
- Added 'user-story' to TOP_LEVEL_USAGE command list

Regression tests added to tests/commands.test.cjs (per lint-regression-
test-names policy; new bug-NNNN standalone files are banned).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(changeset): backfill PR number for #1145 fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 23:34:44 -04:00
Tom Boucher
3827054954 fix(#1156): support table-format STATE.md and insert missing roadmap plan rows (#1172)
* fix(#1162): support table-format STATE.md in state field read/replace

Extend stateExtractField and stateReplaceField in state-document.cts to
detect and operate on pipe-table rows (| Field | value |).  The separator
row | --- | --- | is excluded from matching.  updateCurrentPositionFields
in state.cts now falls through to stateReplaceField for table-format
Current Position sections when the inline Status:/Last activity: patterns
do not match.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1163): insert missing plan checklist rows in roadmap update-plan-progress

cmdRoadmapUpdatePlanProgress now inserts `- [ ] NN-XX-PLAN.md` checkbox
rows under the phase Plans: line when no per-plan checkbox rows exist yet
(fresh template).  Rows are sorted ascending and any already-summarised
plans are immediately marked [x].  The planCountPattern is extended to
also match plain `Plans:` (in addition to bold `**Plans:**`) so plan
counts are updated in both template variants.  The existing-rows check
covers both top-level and indented checkbox forms to preserve idempotency.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1163): fill partial plan-row gaps and scope insertion to active milestone

- Finding 1 (HIGH): replace all-or-nothing rowsAlreadyPresent guard with per-file
  set-difference so missing rows are inserted even when SOME plan rows already exist
- Finding 2 (MEDIUM): extend planCountPattern to recognise **Plans**: (canonical
  template form — bold word + outer colon) alongside **Plans:** and plain Plans:;
  use two-pattern fallback for row insertion to anchor under Plans: checklist header
  rather than the **Plans**: summary line
- Finding 3 (MEDIUM): scope row insertion to the active (post-</details>) milestone
  region so duplicate phase headings in archived sections never receive new rows
- Finding 4 (LOW): rename misleading "pipe-like content" test to honestly describe
  what it tests (multi-row table isolation); add out-of-scope escaped-pipe comment
- Remove now-unused anyCheckboxMatched variable (lint clean)
- Add 5 adversarial regression tests (pre-fix failures confirmed)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1163): scope missing-plan detection to active milestone; preserve authored state fields in table format

- roadmap.cts: compute activeRegion (post-</details> slice) once and use it
  for BOTH missingPlans detection and row insertion, so archived <details>
  rows no longer suppress active-section inserts (Finding 1 code-review)
- roadmap.cts: change (Plans:) inner capture to non-capturing (?:Plans:)
  in insertRowsPatternA to prevent group-numbering shift (Finding 3)
- state.cts: mirror inline-branch preserve-authored guards onto table-format
  branches in updateCurrentPositionFields — Status table branch checks
  isInList/matchesPattern before replacing; Last Activity table branch
  checks isDateShape/inList, preserving executor-authored narrative prose
  (Finding 2 code-review)
- tests: add three failing-first regression tests confirming each finding

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1162,#1163): fold table-format + plan-row regressions into owning module tests

Move bug-1162 cases into tests/state.test.cjs and bug-1163 cases into
tests/roadmap.test.cjs under named regression describe blocks; delete the
standalone bug-NNNN files and prune their allowlist entries.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1156): add changeset fragment for table-format state + roadmap insert fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:56:18 -04:00
Tom Boucher
4db185da74 fix(#1161): make phase complete idempotent for roadmap completion dates (#1177)
* fix(#1161): preserve existing roadmap completion date on repeat phase complete

Guard the Completed cell write in cmdRoadmapUpdatePlanProgress so that a
non-empty, non-placeholder date is never overwritten on repeat invocations.
Also routes the date source through realClock.today() so GSD_NOW_MS pins
the written date deterministically in tests (clock seam parity with the
rest of the codebase). Regression tests added to
4-phase-complete-cjs-regression.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1161): preserve completion date in cmdPhaseComplete (real handler) + drive regression via phase complete CLI

- Import realClock in phase.cts; switch cmdPhaseComplete from bare `new Date()` to `realClock.today()` so the clock seam is honoured in tests
- Apply preserve-or-stamp guard to both 5-col (cells[4]) and 4-col (cells[3]) paths in cmdPhaseComplete: keep existing non-empty, non-dash date; only stamp today on first completion
- Rewrite the #1161 describe block in 4-phase-complete-cjs-regression.test.cjs to drive the real handler via runGsdTools('phase complete 1') end-to-end; cases (a)/(b)/(c)/(d) all FAIL against unfixed build and pass post-fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1161): preserve only date-shaped completion cells (self-heal garbage); correct test helper for 5-col

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1161): add changeset fragment for completion-date idempotence

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:46:10 -04:00
Tom Boucher
ae8bb707bc refactor(#1170): remove hand-maintained INVENTORY count scalars (#1179)
* refactor(#1170): remove hand-maintained INVENTORY count scalars

The `(N shipped)` heading counts in docs/INVENTORY.md were absolute
scalars that collided silently on merge: two branches each bumping the
same integer to N+1 produced a clean git merge whose value the merged
filesystem (N+2) contradicted, hard-failing inventory-counts.test.cjs on
the CI merge commit across all platforms (DEFECT.INVENTORY-MERGE-UNDERCOUNT).

- Strip the six `(N shipped)` heading counts + the two prose footnote
  counts; repoint the intro to INVENTORY-MANIFEST.json as the registry.
- Drop the decorative `generated` date from the manifest + its
  strip-before-compare branch in gen-inventory-manifest.cjs (it conflicted
  on cross-day merges and is read by nothing).
- Delete inventory-counts.test.cjs (scalar-vs-disk gate, the collision
  source); its drift protection is subsumed by the merge-safe set-membership
  test inventory-manifest-sync.test.cjs, which stays as the sole gate.
- Add inventory-headings-countfree.test.cjs guard (fails if a count is
  re-added to a heading).
- Fix already-broken count-bearing cross-doc anchors to stable count-free
  slugs in ARCHITECTURE.md + multi-agent-orchestration.md.
- Retire the now-impossible DEFECT.INVENTORY-MERGE-UNDERCOUNT + obsolete
  RULESET.DOC-CONSISTENCY in CONTEXT.md; de-count DEFECT.INVENTORY-DRIFT;
  correct stale MANIFEST-CANONICAL-KEY (all six families canonical).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1170): backfill changeset PR number (#1179)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:30:03 -04:00