Commit Graph

120 Commits

Author SHA1 Message Date
Tom Boucher
b3165930c3 docs(857): refresh stale "out of scope" comments on landed phase-6 cutovers (#1126)
The render-hooks resolver and capability-state resolver headers (and their CONTEXT.md glossary entries) still claimed "no workflow calls this yet" / "consumed by nothing yet (phase-6 wiring is out of scope)". Both are now false: loop.render-hooks is consumed live by plan-phase.md/autonomous.md (plan:pre ui-phase, verify:post ui-review), and capability-state is the `gsd-tools capability state` CLI diagnostic.

Comment/prose only — no behavior change. Surfaced by the ADR-857 phase 1-5 completeness audit.

Closes #1125

Refs #857

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 16:06:50 -04:00
Tom Boucher
827011b865 fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs,
overwriting/diluting a hand-crafted instruction file. --force was parsed but
silently dropped, and nothing guarded an existing non-GSD file.

- Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers
  (hand-crafted) is left untouched; report action:"skipped". --force (now wired
  through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check
  uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe.
- Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid
  auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned
  across the handler default, config-defaults.manifest.json, buildNewProjectConfig,
  the config template, new-project.md, and cmdGenerateClaudeProfile; advisory
  read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still
  writes AGENTS.md.

Closes #1098

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 16:03:51 -04:00
Tom Boucher
fb8e7a3a65 fix(#1114): write-profile resolves the active runtime config home for USER-PROFILE.md (#1119)
Under Codex, `write-profile` wrote ~/.claude/gsd-core/USER-PROFILE.md while Codex
discuss-phase advisor-mode (installed under ~/.codex) checked the Codex home and
never found the profile, so advisor-mode silently stayed disabled. Resolve the
default output via the runtime-aware getGlobalConfigDir (GSD_RUNTIME / config.runtime),
mirroring cmdGenerateDevPreferences. Claude unchanged; --output still wins.

Closes #1114

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:54:48 -04:00
Tom Boucher
e837bfc1de fix(#1101): record-session updates ## Session Continuity in place, no duplicate block (#1113)
The reported symptom (recorded:false yet STATE.md frontmatter still mutated) was
already resolved on next by #944/#948. This fixes the residual: the DWIM
auto-create recognised only the canonical `## Session` heading, so a bootstrap
`## Session Continuity` section fell through to the append branch and produced a
second `## Session` block. Insert only the missing canonical fields after the
`## Session Continuity` heading — preserving the heading and any prose (e.g.
"Next recommended action") — and teach the snapshot / frontmatter readers to
recognise that heading (the optional ` Continuity` group still excludes
`## Session Continuity Archive`, preserving #2444 scoping).

Closes #1101

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:54:28 -04:00
Tom Boucher
4beceb02fe fix(#1103): preserve newline before Plans: header in roadmap.annotate-dependencies (#1111)
The match regex's `(?:^|\n)` anchor consumes a leading newline on mid-string
matches (the common case where a `**Plans:** N plans` summary line precedes a
bare `Plans:` block). The replacement dropped that newline, fusing the summary
line onto the header — e.g. `**Plans:** 3 plansPlans:`. Re-emit the consumed
newline so markdown line boundaries are preserved.

Closes #1103

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 15:54:25 -04:00
Rezolv
c70842b6df enhance(spec-phase): edge-probe precision probe should name tie-breaking / rounding-mode (#1108)
* feat(spec-phase): name tie-breaking/rounding mode in the edge-probe precision probe (#1102)

The precision edge probe fired on numeric-range requirements but its text never
named the most common rounding failure mode — tie-breaking / rounding mode
(half-up vs half-to-even, ceil/floor/truncate). In an A/B experiment the weak
tier (haiku) false-passed 50% of tie-class defects against the surfaced-
unresolved spec; naming the rule in the probe text drove honest abstention
83%→100% (false-pass 17%→0%, haiku 50%→0%) with no category regression
(precision still fires only on numeric-range).

Sharpen the single TAXONOMY probe string and update every rendering together so
the doc↔fixture↔machine contract (edge-probe-docs-fixtures.test.cjs) stays green:
src/edge-probe.cts, the edge-probe.md taxonomy table + 2 worked examples, the
01-round-half-even and 04-money-rounding fixtures, and the resolve-edge-coverage
how-to. Prose-only — no consumer keys on the probe text; SHAPE_CUES firing
unchanged; fully backward compatible.

* chore(changeset): Changed fragment for edge-probe precision probe text (#1108)
2026-06-12 13:04:31 -04:00
Rezolv
3e836fef0d feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/

Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.

Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.

Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
  planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
  source-checkout-gated build fallback) instead of LLM re-derivation; the engine
  capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
  resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
  per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
  validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.

Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.

* test(#550): RED — status×verification re-cut + probe-core engine specs

Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):

  status: resolved | dismissed | unresolved   (lifecycle, shared)
  verification: explicit | backstop | null     (only when resolved)

- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
  to be extracted — validateResolution(r, validators), validateRequirement,
  analyzeCoverage(items, resolutions?, validators), byVerification rollup,
  runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
  {resolved,backstop}; coverage gains byVerification.{explicit,backstop};
  proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
  COUNT preserved on every fixture (closed set = resolved+dismissed; doc
  line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
  rewritten to the two-axis model.

Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).

* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)

Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.

probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
  verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
  items[] (core never assumes propose is deterministic — edge resolves via
  LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
  count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
  {categories, verification, requiredFieldsByVerification} (ADR-550 #5)

edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.

* chore(#550): register probe-core.cjs artifact in ledgers + inventory

New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:

- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
  never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
  is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
  edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).

probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.

* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]

trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.

Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.

* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard

Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:

- validateResolution now enforces the 'verification is null unless resolved'
  invariant for EVERY status (not just resolved): a dismissed/unresolved
  resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
  (was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
  it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
  contract; a new test locks that an all-dismissed run is NOT affirmatively covered
  (byVerification is the honest gate).

Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.

* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics

Re-review #5 (trek-e) clarity edits:

- Decision 5: annotate that only contract item (a) ships on #584 (the edge
  adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
  parenthetical — the blessed/implemented semantics are count-preserved = the
  CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
  carrying the per-tier resolved-status breakdown. The old parenthetical
  contradicted the shipped count.

* test(#550): cover runProbeCli structural-guard numeric-count branch

Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.

* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)

The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.

* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)

A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.

* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)

templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.

* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)

The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.

* docs(#550): add how-to for resolving edge-coverage findings (B1)

Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.

* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)

trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.

* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)

trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.

* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)

trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.

* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe

Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:

  - spec-phase.md  15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
  - plan-phase.md  93135 -> 94253 (+1118): covered/backstop edge lift into
    must_haves.truths (the live <downstream_consumer> block)

Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.

* chore(#550): reconcile INVENTORY headline counts after rebase onto next

Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
  - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
  - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs

Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
2026-06-12 11:05:31 -04:00
Tom Boucher
9e2ef2c94d fix(#1091): thread install scope into skill converters so local Antigravity/Copilot installs use workspace paths (#1092)
The skills layout wrapper (skillsKind) invoked every per-runtime skill
converter as realConverter(content, skillName, runtime, cmdNames). The
3rd positional arg is overloaded: claude/kimi/cline converters read
`runtime` there, but the copilot/antigravity converters read `isGlobal`
there — so they received the truthy runtime string and always took the
global path branch, leaking ~/.gemini/antigravity/ and ~/.copilot/ into
local/workspace installs instead of .agent/ and .github/.

Thread `scope` from resolveRuntimeArtifactLayout -> dispatchKindEntry ->
skillsKind, derive isGlobal = scope === 'global', and pass it as a
non-colliding 5th positional arg. Move isGlobal out of the colliding 3rd
slot in the two converter signatures (3rd/4th become ignored
_runtime/_cmdNames, matching the kimi convention). The fix flows through
the shared ArtifactKind.stage closure, so applySurface re-apply inherits
it via the same seam.

Regression test exercises the wrapper seam (installRuntimeArtifacts at
local scope) for both runtimes and asserts workspace paths, not global.

Closes #1091

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 00:20:28 -04:00
Tom Boucher
b4a7eabaae feat(#1085): migrate windsurf workspace skills to .devin/ + fix global content refs (#1093)
Fresh windsurf/devin-desktop workspace installs write skills under .devin/ (legacy .windsurf/ recognized); global ~/.codeium/windsurf/ unchanged. Also threads real isGlobal through _applyRuntimeRewrites so global skill content references the codeium path. Closes #1085.
2026-06-11 23:17:01 -04:00
Tom Boucher
77a671ec53 feat(#791): migrate antigravity workspace base dir .agent → .agents (#1090)
Fresh antigravity workspace installs write under the canonical .agents/ (plural) base; legacy .agent/ stays recognized (dual-read). Global ~/.gemini/antigravity/ path unchanged. Closes #791.
2026-06-11 22:37:55 -04:00
Tom Boucher
94872662e9 feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor

postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in
runtime-hooks-surface.cts) now select the event-name dialect from
registry.runtimes[id].runtime.hookEvents instead of the hardcoded
(runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini'
→ AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the
applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents
'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude
dialect (matches the old else).

The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact;
runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED —
hookEvents (2-value) is too coarse to drive them (the event set differs within
hookEvents='claude'); per-event-set drive tracked in #1076.

Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND
pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for
gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread).

Closes #1077

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI

hooks/dist is gitignored and absent in scoped/windows CI jobs that do not
pre-run build:hooks. Without it, install() finds no hook files and all
AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty,
failing every hook-presence assertion. Added an idempotent ensureHooksDist()
called in a top-level before() — mirrors the pattern from
bug-376-claude-js-hook-gsd-rewriter.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors

Purely additive: four new fields on all 16 runtime capability.json descriptors,
validator extended with three new closed-vocab sets, registry regenerated.
Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new
required fields so all 255 capability-registry tests continue to pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): drive per-event hook guards from extendedHookEvents descriptor

Replace hardcoded runtime-name checks (isQwen||runtime==='claude',
runtime==='claude', isGemini) in applySettingsJsonHooks with a single
descriptor-driven extendedEvents array derived from the new opts field.
Remove isQwen and isGemini derivations (no remaining uses after the three
guard blocks are migrated). Wire extendedHookEvents from the capability
registry in bin/install.js call site. Add behavioral regression test
(enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is
purely descriptor-based and runtime-name-agnostic.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY

- Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs
  and read installSurface / writesSharedSettings / permissionWriter from
  runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2).
- ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface.
- Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read
  installSurface directly from capMap descriptor bodies, breaking the require cycle
  (adapter now requires the generated registry; gen-script must not require the adapter).
- Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests)
  pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract.
- Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip

- Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in
  applySettingsJsonHooks (src/runtime-hooks-surface.cts).
- Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with
  hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations
  (ADR-857 phase 5g drive 3).
- Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks
  call site in bin/install.js using the established
  _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom.
- Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites
  proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name;
  (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously
  hardcoded to skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016

Add Decision 7a documenting the four axes added in the 5f-completion pass,
update axis counts from "six" to "twelve", note 5f-completion drives as done
in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done
and 5g (InstallPlan capstone) remains the only open phase.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level)

The test fixture makeRuntimeCapMap did not include installSurface in the runtime object,
so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw.
Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct
installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests.
The gate implementation already reads r.installSurface correctly from the descriptor level.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests

GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a
runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none,
codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors.

GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events
(BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family
events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'.

Added 10 rejection tests in suite 27 covering each gate + each new field validator.
All 16 real runtimes satisfy both gates (verified before coding).
Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES,
VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup

1. bin/install.js: add explicit literal fallback for hooksSurface when the committed
   capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json').
   The descriptor is always the source of truth in normal operation.

2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty)
   to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty.
   ensureHooksDist() in before() guarantees hook files exist.

3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime
   already fully covered by Test 1's deepStrictEqual over all four fields.

4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the
   parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface
   absent → typeof r.installSurface !== 'string' → soft-skip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete)

ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in
runtime-config-adapter-registry collects install-level descriptor axes into
the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope
section updated: 5g capstone is complete, phase 5 fully materialized.
ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves).
ADR-58 Implementation note added (2026-06-11): realized in
runtime-config-adapter-registry (co-located with adapter-selection).
CONTEXT.md Runtime Config Adapter Registry entry extended to document
resolveInstallPlan and both-halves realization.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1082): update install drift guard to the resolveInstallPlan seam (5g)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 22:36:47 -04:00
Tom Boucher
93c5ecd645 feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792.
2026-06-11 22:32:49 -04:00
Tom Boucher
ce0ec808fb docs(#793): disambiguate Trae IDE vs trae-agent in runtime notes (#1083)
Comment-only disambiguation of the trae runtime (Trae IDE vs trae-agent) + community-soft-confirmed note for ~/.trae/skills/. Closes #793.
2026-06-11 22:25:04 -04:00
Tom Boucher
2023f47c64 fix(#1058): cross-reference install manifest in validate agents to catch .md/.toml pair drift (#1079)
* fix(#1058): cross-reference install manifest in validate agents to catch pair drift

`validate agents` considered an agent installed if ANY supported file format was
present on disk. The Codex installer generates a per-agent PAIR (agents/gsd-*.md
AND agents/gsd-*.toml) and records both in gsd-file-manifest.json, so a partial
generated install — one side of the pair missing — was reported as healthy
(agents_found: true, missing: []), masking an incomplete Codex agent install.

checkAgentsInstalled now cross-references the install manifest beside the agents
dir (path.dirname(agentsDir)/gsd-file-manifest.json): for each expected agent, if
the manifest tracks files for it and any tracked file is absent on disk, the
agent is reported in a new `incomplete` list and agents_found becomes false. The
check no-ops when no manifest is present (preserves bundled/claude behavior) and
is scoped to expected agents so retired/stale manifest entries cannot false-flag.

Regression cases added to tests/agent-install-validation.test.cjs cover the
drift case, the complete-pair (no false positive), and the no-manifest no-op.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1058): add changeset for validate-agents manifest pair-drift fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 20:50:08 -04:00
Tom Boucher
9855ea3f39 fix(#1070): recognize "Complete ✓" terminal status in planned-phase transition (#1078)
* fix(#1070): recognize "Complete ✓" terminal status in planned-phase transition

LLM phase executors (e.g. OpenCode) may write `Status: Complete ✓` into
STATE.md when finishing a phase. `state planned-phase` then failed to advance
the Status field on both the frontmatter `**Status:**` line and the Current
Position `Status:` line, because `Complete ✓` matched neither
KNOWN_TEMPLATE_DEFAULTS['Status'] nor any KNOWN_STATUS_PATTERNS entry — so it
was preserved as an executor-authored value and the state machine stayed stuck
on the prior phase.

Add a narrow, fully-anchored pattern `/^Complete\s*[✓✔✅☑]?\s*$/i` to
KNOWN_STATUS_PATTERNS so a bare `Complete` / `Complete ✓` terminal marker yields
to the next phase's `Ready to execute`. Both Status writers consult this array,
so the single addition fixes both paths. Caveat-bearing statuses like
`Complete but needs manual QA` are not matched and remain preserved.

Regression cases added to tests/state.test.cjs (planned-phase block) exercising
both code paths plus the preservation guarantee.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1070): add changeset for planned-phase Complete-status fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 20:30:58 -04:00
Tom Boucher
f11462e58e refactor(#1067): phase 5f-1b — extract the settings-json hook block (applySettingsJsonHooks) — ADR-857/1016 (#1075)
* refactor(#1067): promote referencesHook to runtime-hooks-surface module scope

referencesHook was declared as a local function inside install() but also
called in finishInstall() (module scope), meaning JS hoisting was the only
thing making it work from finishInstall. Move it to src/runtime-hooks-surface.cts,
export it, and have both call sites in install.js use the module's copy.

This is the prerequisite for COMMIT 2 (applySettingsJsonHooks extraction)
per ADR-857 phase 5f-1b.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1067): extract applySettingsJsonHooks to runtime-hooks-surface module

ADR-857 phase 5f-1b: move the ~457-line settings.json hook-registration block
from install() into applySettingsJsonHooks(settings, opts) in
src/runtime-hooks-surface.cts. install() replaces the block with a single call.

Behavior-preserving: all runtime=== guards, isGemini/isQwen/isOpencode/isKilo
derivations, postToolEvent/preToolEvent dialect branches, idempotency checks,
fs.existsSync guards, and console.log/warn messages are verbatim.

Opts bag: 13 fields — runtime, isGlobal, targetDir, postToolEvent,
updateCheckCommand, contextMonitorCommand, promptGuardCommand, readGuardCommand,
readInjectionScannerCommand, configReloadCommand, hookOpts, localCmd,
localShellCmd. preToolEvent computed inside (from runtime). workflowGuardCommand
/ worktreePathGuardCommand / validateCommitCommand / graphifyUpdateCommand /
sessionStateCommand / phaseBoundaryCommand / contextMonitorFile also computed
inside. settings.hooks-only mutations confirmed.

5 source-scan tests updated to read runtime-hooks-surface.cts alongside
install.js (concatenated), so structural regression guards remain valid at
their new canonical location.

install.js: 12700 → 12254 lines (−446).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-11 19:17:06 -04:00
Tom Boucher
58bfae9d6a refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module — ADR-857/1016 (#1064)
* refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module

Extract the structurally-isolated hook-surface writer functions (cline/cursor/
copilot/codex-hooks-json + buildHookCommand + atomicWriteFileSync + node/bash runner
resolvers) out of bin/install.js into a new src/runtime-hooks-surface.cts module
(-693 LOC from install.js). Behavior-preserving: install.js requires + re-exports
the moved functions (module.exports surface preserved); no descriptor reads, no
behavior change. Prerequisite for the descriptor-drive (5f-2), mirroring ADR-3660's
artifactLayout extract→drive split.

Review caught + fixed 3 coupling issues: (HIGH) the module's atomicWriteFileSync
dropped the shared __atomicWrittenTmps temp-tracking → now ONE shared set (module
owns it, install.js aliases it, both cleanups read it); (drift) buildHookCommand
called resolveNodeRunner(opts) vs the original resolveNodeRunner() → reverted; two
source-grep tests (workflow-guard, sh-hook-paths) that scanned install.js for the
moved functions → made behavioral/non-vacuous; duplicate runner resolvers consolidated.

Settings-json hook block (~648 LOC) deferred to 5f-1b; descriptor-drive to 5f-2.
New-module checklist done. ~62 hook test files green; gsd-test 17592/0.

Closes #1059

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1059): reconcile CLI Modules count after merging next (uat-predicate)

Merging current next (which added uat-predicate.cjs via #247) alongside this
branch's runtime-hooks-surface.cjs put the filesystem at 107 bin/lib modules, but
both sides had independently bumped the INVENTORY headline 105→106 so the merge
under-counted. Set "CLI Modules (107 shipped)" + regenerate INVENTORY-MANIFEST.json.
Both module rows already present. Fixes inventory-counts.test.cjs (the only CI red).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 17:21:04 -04:00
Tom Boucher
9223f2f4c8 feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results (#1063)
* feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results

Wire the already-reserved `phase.uat-passed` alias (subcommand `uat-passed`,
mutation:false) into the phase command router with a new markdown-aware
predicate that evaluates HUMAN-UAT results and reports pass only when every
required check passes. Post-SDK-retirement (ADR-0174/#174) successor to the
SDK-framed #70, with no SDK-specific API surface.

New pure module src/uat-predicate.cts:
- stripFalsePositiveContexts: frontmatter -> HTML-comment -> CommonMark-style
  fenced-block state machine (tracks delimiter char+length) -> blockquote,
  each a small composable step, so a `result: passed` inside frontmatter, a
  fenced/~~~ block (incl. ~~~ nested in a ``` fence), a comment, or a
  blockquote is never counted.
- parseUatResultItems: heading-block parser, column-0-anchored same-line
  result; a heading with no result -> `missing` (fail-closed).
- analyzeMarkdown: unterminated fence/comment detection (malformed -> blocker).
- evaluateUatPassed: allowlist pass/verification semantics; passed = no
  blockers && >=1 check && all passing; no_uat_artifacts discriminator (no
  vacuous pass); optional requireVerification policy hook.

Thin cmdPhaseUatPassed handler in phase.cts; router closure rejects unknown
flags via makeInvalidArgs. Hardened across two Codex adversarial passes
(vacuous pass, dropped failing tests, permissive verification status,
nested-fence escape, cross-line result value, masked unterminated comment) —
all fixed fail-closed. New unit + CLI-integration suites incl. a fast-check
property test; docs, CONTEXT glossary, inventory, and changeset updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#247): backfill changeset PR number (#1063)

* fix(#247): indexOf paired-scan for unterminated-comment detection

CodeQL js/incomplete-multi-character-sanitization (high) flagged the
`raw.replace(/<!--[\s\S]*?-->/g,'')`-then-`.includes('<!--')` detection in
analyzeMarkdown as incomplete sanitization (a single regex pass can leave a
residual `<!--`). Replace it with a paired left-to-right indexOf scan that
contains no `.replace()` of the comment token — CodeQL-clean and strictly
more correct (a closed earlier comment can never mask a later unterminated
one). Behaviour unchanged; 98 predicate tests + scoped docker run green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 16:37:17 -04:00
Tom Boucher
8813ee5f95 feat(#429): HARD GATE on negative-grep literals echoed in plan <action> bodies (#1062)
Convert the planner's soft comment-text guideline into a plan-write-time
HARD GATE. When an acceptance criterion negative-greps for a literal
(`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an
`<action>` body (JSDoc samples, head-comment references, "what NOT to do"
snippets), the executor's commit-time verify gate later fails on the
comment echo rather than a real regression — wasting cycles and training
the executor to distrust the gate.

`verify.plan-structure` (the `validate_plan` step) now scans for this:
- confidently-extracted (quoted) negative-grep literal echoed in an
  <action> → error (valid:false), failing plan creation
- unquoted/ambiguous grep target → warning (fallback policy)
- `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal
- positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope

Adds the `<comment_text_discipline>` block to gsd-planner.md, the full
rules + allowlist example to planner-antipatterns.md, and regression
fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary
case proving positive-count gate 11-02 is not flagged).

Closes #429

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 15:46:49 -04:00
Tom Boucher
58ed55683e feat(#1049): phase 5d — drive artifactLayout from the runtime descriptor (retire the 128-LOC switch) (#1053)
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.

Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.

validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.

Closes #1049

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 14:08:46 -04:00
Tom Boucher
734c56ccfe feat(#1046): phase 5c — drive commandStyle from the runtime descriptor (runtime-slash codex-check → lookup) (#1048)
formatGsdSlash now reads commandStyle from registry.runtimes[id].runtime.commandStyle
(lazy require of the committed capability-registry.cjs) instead of the hardcoded
if (rt === 'codex'). Equivalence-preserving (Codex-verified): codex (shell-var) →
$gsd- + lowercased token; all 15 others (slash-hyphen) + unknown → /gsd- +
case-preserved token. canonicalizeRuntimeName + input normalization + claude default
preserved; no circular load (mirrors 5b runtime-homes pattern).

Added a registry-parity test: 16 parametrized sub-tests derive the expected prefix +
lowercasing from each runtime's commandStyle, proving the prefix is a pure function of
the registry (catches future hardcode-vs-registry divergence).

Closes #1046

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 12:37:25 -04:00
Tom Boucher
6968e04d8a feat(#1040): phase 5b — drive configHome from the runtime descriptor — ADR-857/1016 (#1043)
* feat(#1040): phase 5b — drive configHome from the runtime descriptor (runtime-homes switch → lookup)

getGlobalConfigDir now resolves configHome from registry.runtimes[id].runtime.configHome
via a single resolveConfigHomeFromDescriptor(configHome, {env, home, existsSync})
(dot-home / dot-home-nested / xdg / generic-agents-root), replacing the hardcoded
16-runtime switch. Equivalence-preserving: byte-identical config dirs for all 16
runtimes + grok + default + explicitDir (Codex-verified, no divergence).

Nuances preserved: xdg env[1] is a FILE path → path.dirname; existsSync injection
seam keeps antigravity/kimi probe tests hermetic; grok stays hardcoded
(GROK_AGENTS_HOME → ~/.agents, not in the 16); copilot two-env fallback; explicitDir
short-circuit. getGlobalSkillsBase unchanged (out of scope). Lazy require of the
committed capability-registry.cjs (no circular load).

Test env-clearing lists (install.test ENV_KEYS, bug-3126 envKeys) now derived from
the registry runtime configHome.env arrays — auto-correct, closes the missing
KIMI_CONFIG_DIR gap. New 81-case golden equivalence test.

Closes #1040

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1040): make windsurf golden path Windows-portable (path.join, not POSIX literal)

The dot-home-nested windsurf equivalence case hardcoded '/home/u/.codeium/windsurf'
but the resolver builds it via path.join(home,parent,name) → backslashes on Windows.
Use path.join for the expected value. Test-only; production resolver unchanged.
Defensive scan confirmed it was the only path.join-derived hardcoded literal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 12:08:11 -04:00
Tom Boucher
728788fdb3 fix(#1034): correct stale intel.cts comments to canonical INTEL_FILES names (#1036)
Three maintainer-facing comments in src/intel.cts referenced the retired
short/markdown intel filenames (files.json, deps.json, arch.md) and described
arch as searched "text lines", though the implementation iterates the exported
INTEL_FILES map and parses every entry — arch included — as JSON.

Update the comments to cite the canonical names (file-roles.json,
dependency-graph.json, arch-decisions.json) and the JSON-for-arch behavior.
Comment-only; no code or behavior change. Deferred from the #1000 fix.

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:55:55 -04:00
Tom Boucher
e4dfa6b9ea fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer

The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:

1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
   --changed-since/--base for changed-files scoping (no file-list input), and
   --max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
   0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
   the output — i.e. it threw away exactly the findings it exists to surface.
   Success is now decided by whether a valid fallow JSON report was produced,
   not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
   (unusedExports/duplicates/circularDependencies) fallow never shipped, and was
   dead code (the workflow embedded raw JSON; its tests asserted the fictional
   schema, one even calling a non-existent runFallowAudit and passing vacuously).

Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.

Closes #1012

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1012): backfill changeset PR number to 1044

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:50 -04:00
Tom Boucher
ab8b84286e fix(#1013): resolve worktree.baseRef from user/global settings cascade (#1038)
* fix(#1013): resolve worktree.baseRef from user/global settings cascade

cmdWorktreeBaseCheck resolved worktree.baseRef from the project checkout's
.claude/ only (settings.local.json then settings.json). A user/global
worktree.baseRef:"head" — the layer /config writes and the harness honors,
and the only sensible place for a machine-wide preference — was invisible. On a
phase lane (HEAD ahead of origin/HEAD, or no origin/HEAD symref) base-check
returned shouldDegrade:true and execute-phase forced sequential execution,
silently losing the parallel worktree execution the user configured.
CLAUDE_CONFIG_DIR (relocated user config dir) was also ignored.

resolveEffectiveBaseRef now accepts an optional user/global config dir and reads
its settings.json as a third, lowest-precedence layer (project local > project
shared > user/global). cmdWorktreeBaseCheck resolves it via
getGlobalConfigDir('claude'), which honors CLAUDE_CONFIG_DIR. The existing
project-level reads stay as higher-precedence overrides and the injectable
readFile seam is preserved, keeping the unit tests hermetic.

Closes #1013

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1013): backfill changeset PR number to 1038

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:30 -04:00
Tom Boucher
9ed8c7d574 feat(#1026): §5.6/ui-phase cutover — first gate dispatch (plan:pre step + blocking gate) (#1028)
Replace plan-phase.md §5.6 (a 6-branch inline UI gate) with a capability-driven
loop.render-hooks plan:pre dispatch — the FIRST gate dispatch in any workflow.
A step (ui-phase, when:workflow.ui_phase) + a new blocking gate
(when:workflow.ui_safety_gate). New ui.plan-gate check verb returns
{frontend, hasUiSpec, block}; the dispatch runs it unconditionally then fires
the active step (pipeline) or halts on the active blocking gate (manual). The
gate-handling (run check.query; halt if blocking+block) is the reusable
phase-6 template for blocking-gate cutovers.

Config semantics fixed per #1022 + maintainer call: ui_phase gates plan-time
UI-SPEC generation, ui_safety_gate gates the planning block. Common case + all
ui_phase=false cases are equivalence-preserving; the one intended change is
{ui_phase:true, ui_safety_gate:false} now auto-generating in pipelines.

Review found it broken twice (non-generic dispatch, phase-lookup divergence,
then the step-only check nested in a gate loop) — fixed; final Codex pass
verified all 8 (ui_phase,ui_safety_gate)x{pipeline,manual} cases correct.
gsd-ui-phase skill + autonomous §3a.5 untouched (§3a.5 deferred).
getRoadmapPhaseWithFallback mirrors cmdRoadmapGetPhase for lookup parity.

Closes #1026

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 00:58:21 -04:00
Rezolv
28ac89d810 fix(#1008): tolerate EAGAIN + short writes in io output()/error() (#1009)
* fix(#1008): tolerate EAGAIN + short writes in io output()/error()

The I/O Module wrote stdout/stderr with a bare fs.writeSync(fd, data), assuming
it blocks until the kernel accepts every byte. That is false when the fd is a
non-blocking pipe (as under the parallel node:test runner on Linux CI): a full
pipe throws EAGAIN and a partially-drained pipe returns a short count. The former
caused spurious failures (e.g. bug-974 graphify property test threw EAGAIN); the
latter risked silently truncating output.

Add writeAllSync(fd, data): loop on short counts and retry EAGAIN/EINTR with a
bounded backoff. The backoff sleep buffer is allocated lazily on the first retry
(rare) and reused — keeping it out of module load avoids perturbing the
SharedArrayBuffer-allocation accounting in perf-316 and costs nothing on the
common no-retry path. Route output() and error() through it; non-transient
errors (EPIPE) still propagate. Mirrors the transient-errno handling already
applied to STATE.md lock acquisition (ACQUIRE_LOCK_RETRY_ERRNOS / #3776).

Regression cases live in tests/io.test.cjs (the owning module's file, per the
regression-test-name placement policy) and inject fs.writeSync via mock.method:
EAGAIN/EINTR retry, short-write no-truncation, EPIPE still surfaces, and error()
retries while still exit(1). Red against the pre-fix bare-writeSync io.cjs.

* chore(#1008): add Fixed changeset for io EAGAIN/short-write fix
2026-06-10 17:05:07 -04:00
Tom Boucher
972a41a528 fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:)

Closes #967

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#967): backfill changeset pr number (990)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:10:40 -04:00
Tom Boucher
921a7cd618 fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op (#986)
* fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op

When `--budget` was the last arg or followed by a non-numeric token,
parseInt(undefined/NaN-string, 10) produced NaN. NaN is falsy so both
the router check and applyBudget gate silently skipped budget trimming.
The query ran unbounded with no warning.

Fix: guard in graphify-command-router.cts — if args[budgetIdx+1] is
absent or parses to NaN, emit ERROR_REASON.USAGE and return early.
Defensive fix in graphify.cts: tighten `if (!budgetTokens)` →
`if (budgetTokens == null)` and `if (options.budget)` →
`if (options.budget != null)` so a real 0/NaN caller is handled
predictably by both independent guards.

Regression tests: 16 cases (unit/mock, subprocess, property-based)
covering boundary inputs: missing value, non-numeric, valid integers,
and fast-check properties over the budget parse contract.

Closes #974

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#974): backfill changeset pr number (986)

* fix(#974): convert static property (b) to real fc.assert property with constrained generator + deterministic seed

The test named "property: --budget as last arg always produces usage error"
was a static test with no fc.assert — it only checked a single hardcoded
term ("someterm") and could never flake or produce a fast-check path. This
is a generator/property bug (case b): the test was mislabeled as a property
test but lacked parameterization.

Root-cause of CI path "46:0:0" / seed 42 failure: a naive parameterized
version without the !startsWith('--') filter could feed term='--budget',
causing args.indexOf('--budget') to hit index 2 (the term slot) rather than
index 3 (the flag slot), placing the router in a different code path. The
property still holds — NaN detection fires on rawBudget='--budget' — but
the assertion text referenced the wrong invariant, making the failure appear
spurious. Fix: constrain the generator to non-flag terms (filter out
strings starting with '--') and pass explicit { seed: 42, numRuns: 200 } to
fc.assert so the test is fully deterministic in CI regardless of GSD_FC_SEED.

No change to src/graphify-command-router.cts (router is correct).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): constrain non-numeric budget property generator to genuinely-NaN values and pin fast-check seeds (CI determinism)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): use a strict valid-term generator and pin seeds so budget property tests are deterministic

Replace fc.string({minLength:1}).filter(!startsWith('--')) term generators in
properties (b) and (d) with a shared validTerm = fc.stringMatching(/^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/)
that is alphanumeric-leading and contains no whitespace, flags, or sign-numerics.
This eliminates the class of CI failures where the old generator produced out-of-
contract inputs (" ", "+5", empty) that the router legitimately rejects for reasons
outside the budget-parse contract under test. Pin distinct seeds (1001–1004) on
every fc.assert for CI determinism. Stress-tested at numRuns=100000 per property
and across 10 seeds (1,2,7,13,42,43,44,99,12345,999999) at numRuns=2000 — all pass.
No src change: gsd-core/bin/lib/graphify-command-router.cjs is correct.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): replace flaky term-fuzzing properties with deterministic examples; keep value-fuzzing properties

The validTerm regex /^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/ admitted single-digit
strings like "0" which are falsy; the router's `if (!term)` guard fires before
the budget-missing-value path, producing a spurious errFn call. Properties (b)
and (d), which test the BUDGET contract (not term handling), are replaced with
deterministic example loops over fixed valid terms. Properties (a) and (c),
which fuzz the BUDGET VALUE with a fixed term, are kept unchanged (seed+numRuns
pinned). The validTerm generator is fully removed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:58:42 -04:00
Tom Boucher
898d55788e fix(#976): detect command+args (wrapped) hook registrations in installer presence checks (#994)
* fix(#976): detect command+args (wrapped) hook registrations in installer presence checks

Add referencesHook() helper that inspects both h.command (standard form) and
h.args[] (args-form / wrapped-launcher form) when checking whether a managed
hook is already registered.  Rewrite all has*Hook predicates and the
alreadyHas* guards to use it so args-form registrations suppress the duplicate
stock string-command entry that was previously appended on every install/update.

Also add an explicit args-form skip to rewriteLegacyManagedNodeHookCommands so
entries with a non-empty args[] are left untouched (they are intentional user
wrappers, not legacy bare-node commands to migrate).

Extend isManagedHookCommand() in shell-command-projection.cts with an optional
args: unknown[] parameter that checks whether any arg's basename matches the
managed hook surface set — backward compatible; existing callers are unaffected.

Regression test added to tests/install-regressions.test.cjs:
- two-pass install with an args-form SessionStart entry pre-written to
  settings.local.json asserts exactly 1 hook entry remains after reinstall
  (previously 2 — the original args-form + a new stock string-command duplicate)
- rewriteLegacyManagedNodeHookCommands test asserts args-form entries unchanged

Closes #976

* chore(#976): backfill changeset pr number (994)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:44:24 -04:00
Tom Boucher
22bb82e205 fix(#965): emit structured json error for unexpected handler throws under --json-errors (#987)
* fix(#965): emit structured json error for unexpected handler throws under --json-errors

When GSD_JSON_ERRORS=1 / --json-errors is active, an unexpected (non-ExitError) throw
in a handler now emits { ok: false, reason: "sdk_fail_fast", message } to stderr instead
of a raw stack trace. The plain-text behaviour (no json-error mode) is unchanged.

Closes #965

* chore(#965): backfill changeset pr number (987)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:40:12 -04:00
Tom Boucher
caca4d255c feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) (#988)
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)

Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).

Closes #985

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)

The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 10:00:30 -04:00
Tom Boucher
9a03539c2d feat(#981): audit-uat + audit-open command cutover — commands-only capability (ADR-857 phase 4d-impl-3) (#984)
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.

Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.

Closes #981

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 08:34:49 -04:00
Tom Boucher
77bfd943dd feat(#972): graphify command cutover — first capability owning a command family (ADR-857 phase 4d-impl-2) (#975)
Migrate graphify into a Capability that owns its `graphify` command family,
dispatched via the registry (#961 mechanism) instead of a hardcoded case.
graphify is now an enable/disable plug-in.

- src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command),
  reproduces the removed case EXACTLY (query +--budget, status, diff, build,
  hidden build snapshot, usage/unknown errors); injectable _graphify test seam.
- capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify],
  config:{graphify.enabled default false}, commands:[{family:graphify, module,
  router:routeGraphifyCommand}].
- removed case 'graphify' from gsd-tools.cjs; graphify now flows
  default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router.
- regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/
  capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op.

Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per
subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing
output shapes; existing graphify tests pass unchanged through the new path.

Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op
quirk, preserved for equivalence, filed separately.

Closes #972

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 07:12:57 -04:00
Tom Boucher
1c86368785 fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch (#955)
* test(#947): add regression tests and update stale Hermes assertions

- Add bug-947-hermes-gsd-prefix.test.cjs: 12 TDD tests covering fresh
  install canonical layout, bare-stem migration, manifest key format,
  and non-Hermes runtime isolation
- Update hermes-skills-migration.test.cjs: bare-stem → gsd-prefixed
  path and name assertions (#947 canonical layout)
- Update install-nested-layout.test.cjs: Hermes NEST matrix prefix ''
  → 'gsd-'
- Update install-regressions.test.cjs: Defect #1 now seeds bare-stem
  dirs (help/, quick/) and asserts gsd-help/ canonical output; use
  real GSD stems so readGsdCommandNames() migration finds them
- Update install-runtime-artifacts.test.cjs: Hermes nested layout and
  legacy migration assertions align with gsd- prefix
- Update install.test.cjs: Hermes install test uses gsd- prefixed paths
- Update runtime-artifact-layout.test.cjs: prefix '' → 'gsd-'

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch

Hermes skills were installing under bare-stem paths
(skills/gsd/<stem>/SKILL.md, name: <stem>) due to prefix: '' set in
ADR-3660 / #3664. This broke /gsd-<stem> dispatch and forced users to
invoke skills without the gsd- namespace prefix.

- src/runtime-artifact-layout.cts: change Hermes skillsKind prefix
  from '' to 'gsd-'; skills now land at skills/gsd/gsd-<stem>/SKILL.md
  with name: gsd-<stem>
- bin/install.js _runLegacyInstallMigrations: invert the #3664
  migration — remove stale bare-stem dirs (using readGsdCommandNames()
  to distinguish GSD-owned stems from user content), keep gsd-* dirs
  which are now canonical
- bin/install.js _runLegacyUninstallCleanup: also remove bare-stem
  dirs on uninstall for clean teardown
- bin/install.js uninstallRuntimeArtifacts: post-cleanup removes
  DESCRIPTION.md and empty skills/gsd/ category dir on Hermes
- bin/install.js: remove skillListPrefix Hermes exception (now uses
  shared 'gsd-' path)
- docs/adr/3660-runtime-artifact-layout-module.md: document #947
  reversal of the bare-stem sub-decision

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: add changeset for #947 fix (#955)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#947): remove ALL pre-migration bare-stem Hermes skills on reinstall (adversarial review)

Replace readGsdCommandNames()-based bare-stem cleanup (which missed skills
not in the commands source tree, e.g. dev-preferences) with
_removeHermesBareStemDirs(), called AFTER the install loop when the exact
set of installed gsd-<stem>/ dirs is authoritative. For every gsd-<stem>/
written this run, the corresponding bare skills/gsd/<stem>/ is removed.
User-owned bare dirs with no gsd-<stem> counterpart are preserved.

Add two adversarial-review regression tests that FAIL on old code:
- bare skills/gsd/dev-preferences/ removed when gsd-dev-preferences/ installed
- user-owned bare dir with no gsd-<stem> counterpart is preserved (no over-deletion)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 00:22:36 -04:00
Tom Boucher
46967baae8 fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944) (#952)
* test(#948): add regression tests for no-op write guard and record-session auto-create (#944)

Red before fix: 11/15 tests fail. Green after: 15/15.
Covers zero-match patch byte-identity, milestone_name preservation,
stopped_at frontmatter-wins, record-session auto-create fallback, and
adversarial fixtures (CRLF, empty body, non-canonical labels).

Also registers bug-948-state-noop-write-guard.test.cjs in the state
bucket of lint-test-file-count.allowlist.json.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944)

Shared root cause: `readModifyWriteStateMd` wrote STATE.md unconditionally
even when the transform produced no change, and `syncStateFrontmatter`
re-derived frontmatter from the possibly-stale body on every write.

Three coordinated fixes in src/state.cts:

1. readModifyWriteStateMd: add no-op guard — when transform result ===
   input content, skip the write entirely (no platformWriteSync, no
   last_updated bump, no frontmatter re-derive). Fixes #948 zero-match
   phantom write and the #944 phantom last_updated bump.

2. syncStateFrontmatter: extend existing-frontmatter preserve logic —
   fall back to existingFm['milestone_name'] / existingFm['milestone']
   when the derived value is the template placeholder 'milestone'
   (getMilestoneInfo returns this literal when it cannot match the
   version in ROADMAP.md); prefer existingFm['stopped_at'] /
   existingFm['paused_at'] over a body-derived value (the frontmatter
   value, written by the canonical record-session path, wins over stale
   historical body lines). Mirrors the fallback already in cmdStateJson.

3. cmdStateRecordSession: when --stopped-at / --resume-file are supplied
   but body labels are absent, DWIM auto-create a canonical ## Session
   section (mirroring how add-decision / add-blocker / record-metric
   auto-create their sections). Never return a silent recorded:false when
   the caller supplied values.

SDK check: no sdk/src/state.ts exists in this repo (the comment in
cmdStateSnapshot references a sibling concern in the TypeScript SDK
codebase, which is a separate repo not present here).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add changeset for PR #952 (fix #948/#944)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#948): correct stopped_at preserve rule; adjust test for sync behaviour

The "always prefer frontmatter stopped_at" rule in syncStateFrontmatter
was too aggressive — it broke phase.complete which intentionally updates
stopped_at in the body and expects syncStateFrontmatter to pick it up.

The primary fix (no-op guard in readModifyWriteStateMd) already prevents
the stale-body-overwrites-frontmatter scenario from #948: the file is not
written when the transform produces no change, so syncStateFrontmatter
never runs on a zero-match patch. The body-derived value can only win when
an actual write occurs, which means the body was legitimately updated.

Reverted to the original #905 rule for stopped_at/paused_at: fall back to
existing frontmatter only when the derived value is absent (empty/null).

Also adjusted the sync-suite test to assert what state sync actually does
(milestone_name preservation) rather than a stopped_at-wins property that
state sync does not have by design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#944): update existing session block in place (adversarial review)

HIGH finding: the DWIM auto-create in cmdStateRecordSession was appending
a second ## Session block unconditionally, even when one already existed
with non-canonical content (e.g. a markdown table). Both
buildStateFrontmatter and cmdStateSnapshot read only the FIRST ## Session
block via regex, so the newly-written Stopped at / Resume file values
landed in the second, invisible block — frontmatter stopped_at stayed
stale and state-snapshot returned nulls.

Fix: check for an existing ## Session heading. When one is present,
normalize that section in place by replacing its body with canonical
**Last session:** / **Stopped at:** / **Resume file:** bold-label lines.
Only append a brand-new section when NO ## Session heading exists.

LOW finding: the auto-create scaffold emits **Last session:** but
cmdStateSnapshot only matched **Last Date:**, so session.last_date was
null after auto-create despite a valid timestamp being written.

Fix: extend the lastDateMatch regex in cmdStateSnapshot to also accept
**Last session:** / Last session: (the form the scaffold writes).

Tests: 3 new tests added to bug-948-state-noop-write-guard.test.cjs that
confirmed failure against the previous HEAD and pass after this fix:
- exactly one ## Session block after record-session with non-canonical existing block
- state-snapshot sees correct stopped_at via first Session block (not a duplicate)
- state-snapshot session.last_date is non-null after auto-create on body-less file

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#944): improve in-place section replace to cleanly remove old body content

The previous regex `/(^## Session[ \t]*$)([\s\S]*?)(?=\n^## |\n*$)/im`
with a lazy match consumed nothing after the heading, so old non-canonical
body content (e.g. table rows) remained after the new canonical lines.

While functionally correct (parsers found the canonical lines first in the
FIRST ## Session block), it left stale content in the section. Replace with
a negative-lookahead per-line pattern that consumes all content from the
heading up to (but not including) the next ## heading, producing a clean
section with only the canonical bold-label lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 00:22:29 -04:00
Tom Boucher
dc139b38e0 feat(#949): install/surface consume derived profiles/clusters (ADR-857 phase 4c) (#954)
Make install + surface read the registry's derived profileMembership/
capabilityClusters so a capability's tier drives what installs + surfaces.
resolveProfile (when given the registry) unions capability skills for the
profiles its tier implies before the requires: closure; resolveSurface merges
capabilityClusters into the cluster map. bin/install.js, /gsd:surface, and the
capability-state resolver all thread the registry.

Shipped as a proven no-op: the UI capability is reconciled to tier:full (its
skills were full-only in the hand-authored profiles), so it contributes only to
the full profile (already the '*' sentinel) and core/standard are unchanged.
Equivalence tests prove resolveProfile/resolveSurface/listSurface/staging/
capability-state are identical with vs without the registry; the core-alias
staging path is verified equivalent (empty manifest → raw PROFILES.core).

Closes #949

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 22:35:42 -04:00
Tom Boucher
85cfa5dc13 feat(#945): unified capability-state resolver (ADR-857 phase 4b) (#946)
Add a read-side query composing the three toggle systems into one
per-capability view. resolveCapabilityState({registry, installedSkills,
surfacedSkills, config, cwd}) reports installed (skills ⊆ resolved install
profile), surfaced (skills ⊆ resolved surface), and per-hook active (no when →
active; non-empty-string when → resolved via _resolveActivationValue; empty/
non-string → inactive), with no forced composite verdict. cmdCapabilityState
does the I/O (resolveProfile + resolveSurface + loadConfig), resolves the
runtime config dir via the canonical getGlobalConfigDir (--config-dir override),
and surfaces resolution failures as warnings rather than a false installed='*'.
Routed as `gsd-tools capability state`.

Additive: install/surface/workflows untouched; consumed by nothing.

Closes #945

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 16:18:23 -04:00
Tom Boucher
74a121bb4f fix(#934): reapply verifier handles missing pristine baseline post-rename (#937)
* fix(#934): reapply verifier handles missing pristine baseline post-rename

Gap 1 (verify-reapply-patches.cjs): when backup-meta.json records a
pristine_hash for a file but gsd-pristine/ has no corresponding snapshot
on disk, the verifier fell to over-broad mode and produced false
FAIL_USER_LINES_MISSING. Fix: return advisory OK_NO_BASELINE (non-blocking,
exit 0) so the verifier does not block on files it cannot reason about.

Gap 2 (new migration 004): migration 003 removed legacy get-shit-done/
runtime files but left gsd-pristine/get-shit-done/ orphan snapshots in
place. Those stale snapshots referenced get-shit-done/... key paths that
no longer match the active gsd-core/... layout. Fix: add migration
004-prune-stale-pristine-get-shit-done (NOT editing 003, preserving its
checksum — ref #670 guard) to remove all files under
gsd-pristine/get-shit-done/ as GSD-managed pristine snapshots.

Includes tests: bug-934 OK_NO_BASELINE assertions in the verifier test,
new installer-migration-prune-stale-pristine.test.cjs, updated
installer-migrations baseline-lock checksum for 004.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#934): rename migration to satisfy legacy-name guard + mark intentional path refs

Rename src/installer-migrations/004-prune-stale-pristine-get-shit-done.cts
→ 004-prune-stale-pristine-snapshots.cts so the filename no longer contains the
forbidden token.  Update .gitignore and eslint.config.mjs to track the new built
path.  Add gsd-allow-legacy-name markers to the remaining intentional uses of the
legacy path string in the migration body (lines 3 and 100) and in tests
(installer-migration-prune-stale-pristine.test.cjs lines 202 and 226; and the
baseline-lock key in installer-migrations.test.cjs:1469).  Update the baseline
checksum for migration 2026-06-09-prune-stale-pristine-get-shit-done to reflect
the two new marker comments added to its body.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 12:51:15 -04:00
Colin Johnson
eadcf0fa12 Merge branch 'next' into kimi-runtime-support 2026-06-09 10:27:59 -04:00
Tom Boucher
8ee66ade77 fix(#929): cmdSkillManifest discovers nested concrete skills (#933)
Scans `gsd-ns-<router>/skills/<stem>/SKILL.md` in addition to the
existing flat `<stem>/SKILL.md` layout, so gsd-health and gsd-settings
report the correct concrete skill count on nested-layout runtimes
(cline, qwen, hermes, augment, trae, antigravity).

Guard: descent into a `skills/` subdir is restricted to `gsd-ns-*`
router directories — unrelated user dirs that happen to have a
`skills/` subdir are not traversed. Dual-routed concretes (same name
under two routers) are deduped within each root.

Adds a negative-case regression test: verifies that a non-`gsd-ns-*`
dir (e.g. `my-tool/`, `gsd-settings/`) with its own `skills/`
subdir does NOT contribute nested entries to the manifest.

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 10:25:55 -04:00
Colin Johnson
27aca73b6e Merge branch 'next' into kimi-runtime-support 2026-06-09 10:15:41 -04:00
Tom Boucher
03d9fcdd65 fix(#924): revert Claude to flat skill layout so concrete skills are discoverable (#928)
PR #883 nested Claude skills 3 levels deep under gsd-ns-*/skills/<stem>/SKILL.md.
Claude Code's Skill tool scans only one level under ~/.claude/skills/ — nested
concrete skills were never listed and Skill(skill="gsd-plan-phase") calls failed.

Revert to flat layout: all ~61 concrete skills at ~/.claude/skills/gsd-<name>/SKILL.md.
The 6 other runtimes confirmed as non-recursive scanners (cline, qwen, hermes, augment,
trae, antigravity) retain their nested layout — only Claude changes.

Tradeoff: ~61 top-level skill dirs return to the flat install, but they are discoverable
and invokable. Nested concretes were invisible to the Skill tool entirely.

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 08:59:20 -04:00
Viktorplus
c6c3a51277 Merge branch 'next' into kimi-runtime-support 2026-06-09 13:31:21 +02:00
Tom Boucher
5670feaef5 feat(#918): loop.render-hooks resolver — consume the Capability Registry (ADR-857 phase 3c) (#920)
Add the loop.render-hooks resolver: the first registry-consuming query.
`gsd-tools loop render-hooks <point>` validates the point against the
authoritative canonical 12, reads the registry's materialized byLoopPoint
hooks, filters them by activation, and emits a JSON envelope {point,
activeHooks, rendered} with ordered markdown.

Activation resolves each hook's `when` key by precedence: loadConfig value
(post-cutover federated) -> raw config.json workstream/root single-key lookup
(pre-cutover central override) -> registry configSchema default (so a
default:true capability hook is active out-of-the-box) -> inactive. Guarded
single-value reads only (no merged object built from untrusted keys).

Registry-only: no workflow calls the resolver yet (wiring is the phase-6
cutover). Completes the phase-3 trio.

Closes #918

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 00:24:12 -04:00
Tom Boucher
6dbd895028 feat(#910): federated config merge in config-loader (ADR-857 phase 3b) (#914)
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.

Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.

Closes #910

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 23:37:41 -04:00
Colin Johnson
bf8813b7a0 Merge branch 'next' into kimi-runtime-support 2026-06-08 23:04:37 -04:00
Tom Boucher
d12809985e fix(#905): preserve STATE.md frontmatter scalars in syncStateFrontmatter (#907)
syncStateFrontmatter was silently dropping current_phase, current_phase_name,
current_plan, and progress when body annotations were absent (e.g. after an
agent or tool rewrote the body). These scalars can only be derived from body
annotations — when absent, buildStateFrontmatter returns nothing for those
keys. Added existingFm fallbacks mirroring the same pattern already applied in
cmdStateJson, so every writeStateMd call preserves the existing values instead
of stripping them. Also extended cmdStateJson with the same fallbacks for the
three non-progress scalars.

Adds regression test (7 cases) + lint-test-file-count allowlist entry.

Closes #905

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:52:43 -04:00
Tom Boucher
808df9110c fix(#892): parse checklist-style roadmap phases in validate/verify (#908)
buildRoadmapPhaseVariants() only matched heading-style phases (## Phase N:),
silently skipping the supported checklist format (- [x] **Phase N: name**).
This caused W007 false-positives for every on-disk phase dir when the project
uses a checklist ROADMAP. Fix adds a second regex pass (mirroring the existing
buildNotStartedPhaseVariants() approach). Also refactors the duplicate
inline heading-only regex in cmdValidateConsistency() to delegate to
buildRoadmapPhaseVariants() (DRY). Regression test in
tests/bug-892-validate-checklist-roadmap-phases.test.cjs covers both paths.

Closes #892

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:52:19 -04:00
Tom Boucher
6edc39c4eb chore(#893): remove dead loadConfig from configuration.cts (#912)
`loadConfig` in configuration.cts was superseded by config-loader.cts
(ADR-857 phase 2e, #885). Exhaustive grep confirms no caller imports
loadConfig from configuration.cjs — all live callers use config-loader.cjs
or the core.cjs back-compat re-export. configuration.cts now provides only
the pure normalization and defaults primitives (normalizeLegacyKeys,
mergeDefaults, migrateOnDisk, CONFIG_DEFAULTS) that config-loader.cts
depends on. Updated CONTEXT.md and docs/INVENTORY.md to reflect the
narrowed module surface.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 22:51:43 -04:00