Commit Graph

3530 Commits

Author SHA1 Message Date
Tom Boucher
7c07fce70f fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes (#1084)
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes

On runtimes that execute each fenced bash block in a separate shell process
(e.g. Claude Code — documented behavior: each Bash command is a separate
process; inline shell functions and exported vars do not persist between
calls), the once-per-file gsd_run() function was undefined in every block
after the preamble block, and the call was swallowed by
`2>/dev/null || echo "{}"` into silent empty state.

Fix (budget-neutral session-level resolution):
- Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own
  location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm
  `bin` field (global installs) and shipped to local installs via the recursive
  gsd-core/ copy.
- The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"`
  to the file named by $CLAUDE_ENV_FILE (Claude Code's documented
  env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from
  PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline
  gsd_run() definition remains the fallback for all other runtimes. The
  single-quoted dir neutralizes shell metacharacters at source time.
- Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files.
- XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md
  to 93135; legitimate content growth, ratchet-up per #717).

Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper
delegation and end-to-end PATH persistence (sourcing the env file with a
space-bearing install path).

Known limitation: an install path containing a literal single-quote yields a
malformed env-file line and falls back to the status quo (no regression);
rare on sanitized home directories.

Closes #381

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#381): add changeset for gsd_run fresh-shell reachability fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit)

Windows Git Bash (msys2) does not honor Node's chmod exec bit for
PATH-executing extension-less scripts, so the bare `gsd_run` command lookup
failed there even though the env-file PATH persistence was correct. The
env-file content assertions (the fix's actual cross-platform logic) still run
on every platform; only the final source-and-execute sub-step is gated to
non-win32. Global installs on Windows are covered by npm's generated bin shim.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 21:42:52 -04:00
Tom Boucher
2023f47c64 fix(#1058): cross-reference install manifest in validate agents to catch .md/.toml pair drift (#1079)
* fix(#1058): cross-reference install manifest in validate agents to catch pair drift

`validate agents` considered an agent installed if ANY supported file format was
present on disk. The Codex installer generates a per-agent PAIR (agents/gsd-*.md
AND agents/gsd-*.toml) and records both in gsd-file-manifest.json, so a partial
generated install — one side of the pair missing — was reported as healthy
(agents_found: true, missing: []), masking an incomplete Codex agent install.

checkAgentsInstalled now cross-references the install manifest beside the agents
dir (path.dirname(agentsDir)/gsd-file-manifest.json): for each expected agent, if
the manifest tracks files for it and any tracked file is absent on disk, the
agent is reported in a new `incomplete` list and agents_found becomes false. The
check no-ops when no manifest is present (preserves bundled/claude behavior) and
is scoped to expected agents so retired/stale manifest entries cannot false-flag.

Regression cases added to tests/agent-install-validation.test.cjs cover the
drift case, the complete-pair (no false positive), and the no-manifest no-op.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1058): add changeset for validate-agents manifest pair-drift fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 20:50:08 -04:00
Tom Boucher
9855ea3f39 fix(#1070): recognize "Complete ✓" terminal status in planned-phase transition (#1078)
* fix(#1070): recognize "Complete ✓" terminal status in planned-phase transition

LLM phase executors (e.g. OpenCode) may write `Status: Complete ✓` into
STATE.md when finishing a phase. `state planned-phase` then failed to advance
the Status field on both the frontmatter `**Status:**` line and the Current
Position `Status:` line, because `Complete ✓` matched neither
KNOWN_TEMPLATE_DEFAULTS['Status'] nor any KNOWN_STATUS_PATTERNS entry — so it
was preserved as an executor-authored value and the state machine stayed stuck
on the prior phase.

Add a narrow, fully-anchored pattern `/^Complete\s*[✓✔✅☑]?\s*$/i` to
KNOWN_STATUS_PATTERNS so a bare `Complete` / `Complete ✓` terminal marker yields
to the next phase's `Ready to execute`. Both Status writers consult this array,
so the single addition fixes both paths. Caveat-bearing statuses like
`Complete but needs manual QA` are not matched and remain preserved.

Regression cases added to tests/state.test.cjs (planned-phase block) exercising
both code paths plus the preservation guarantee.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1070): add changeset for planned-phase Complete-status fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 20:30:58 -04:00
Tom Boucher
f11462e58e refactor(#1067): phase 5f-1b — extract the settings-json hook block (applySettingsJsonHooks) — ADR-857/1016 (#1075)
* refactor(#1067): promote referencesHook to runtime-hooks-surface module scope

referencesHook was declared as a local function inside install() but also
called in finishInstall() (module scope), meaning JS hoisting was the only
thing making it work from finishInstall. Move it to src/runtime-hooks-surface.cts,
export it, and have both call sites in install.js use the module's copy.

This is the prerequisite for COMMIT 2 (applySettingsJsonHooks extraction)
per ADR-857 phase 5f-1b.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1067): extract applySettingsJsonHooks to runtime-hooks-surface module

ADR-857 phase 5f-1b: move the ~457-line settings.json hook-registration block
from install() into applySettingsJsonHooks(settings, opts) in
src/runtime-hooks-surface.cts. install() replaces the block with a single call.

Behavior-preserving: all runtime=== guards, isGemini/isQwen/isOpencode/isKilo
derivations, postToolEvent/preToolEvent dialect branches, idempotency checks,
fs.existsSync guards, and console.log/warn messages are verbatim.

Opts bag: 13 fields — runtime, isGlobal, targetDir, postToolEvent,
updateCheckCommand, contextMonitorCommand, promptGuardCommand, readGuardCommand,
readInjectionScannerCommand, configReloadCommand, hookOpts, localCmd,
localShellCmd. preToolEvent computed inside (from runtime). workflowGuardCommand
/ worktreePathGuardCommand / validateCommitCommand / graphifyUpdateCommand /
sessionStateCommand / phaseBoundaryCommand / contextMonitorFile also computed
inside. settings.hooks-only mutations confirmed.

5 source-scan tests updated to read runtime-hooks-surface.cts alongside
install.js (concatenated), so structural regression guards remain valid at
their new canonical location.

install.js: 12700 → 12254 lines (−446).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-11 19:17:06 -04:00
Tom Boucher
7749e4f16e docs(#1065): record DEFECT.INVENTORY-MERGE-UNDERCOUNT in CONTEXT.md (#1066)
The INVENTORY.md "CLI Modules (N shipped)" headline is an ABSOLUTE filesystem
count, not a delta. Two branches each adding a bin/lib module both bump N->N+1,
so the merge under-counts and tests/inventory-counts.test.cjs hard-fails on the
CI merge-commit across all platforms — even though each branch passed local
gsd-test (which runs branch-only, not branch+next). Hit #844 and #1059.

Adds a DEFECT.INVENTORY-MERGE-UNDERCOUNT predicate (.symptom/.detect/.fix-forward/
.prevention) after DEFECT.INVENTORY-DRIFT: re-derive the count from the filesystem
+ regen the manifest BEFORE every push on module-adding branches.

Closes #1065

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 17:24:08 -04:00
Tom Boucher
58bfae9d6a refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module — ADR-857/1016 (#1064)
* refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module

Extract the structurally-isolated hook-surface writer functions (cline/cursor/
copilot/codex-hooks-json + buildHookCommand + atomicWriteFileSync + node/bash runner
resolvers) out of bin/install.js into a new src/runtime-hooks-surface.cts module
(-693 LOC from install.js). Behavior-preserving: install.js requires + re-exports
the moved functions (module.exports surface preserved); no descriptor reads, no
behavior change. Prerequisite for the descriptor-drive (5f-2), mirroring ADR-3660's
artifactLayout extract→drive split.

Review caught + fixed 3 coupling issues: (HIGH) the module's atomicWriteFileSync
dropped the shared __atomicWrittenTmps temp-tracking → now ONE shared set (module
owns it, install.js aliases it, both cleanups read it); (drift) buildHookCommand
called resolveNodeRunner(opts) vs the original resolveNodeRunner() → reverted; two
source-grep tests (workflow-guard, sh-hook-paths) that scanned install.js for the
moved functions → made behavioral/non-vacuous; duplicate runner resolvers consolidated.

Settings-json hook block (~648 LOC) deferred to 5f-1b; descriptor-drive to 5f-2.
New-module checklist done. ~62 hook test files green; gsd-test 17592/0.

Closes #1059

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1059): reconcile CLI Modules count after merging next (uat-predicate)

Merging current next (which added uat-predicate.cjs via #247) alongside this
branch's runtime-hooks-surface.cjs put the filesystem at 107 bin/lib modules, but
both sides had independently bumped the INVENTORY headline 105→106 so the merge
under-counted. Set "CLI Modules (107 shipped)" + regenerate INVENTORY-MANIFEST.json.
Both module rows already present. Fixes inventory-counts.test.cjs (the only CI red).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 17:21:04 -04:00
Tom Boucher
9223f2f4c8 feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results (#1063)
* feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results

Wire the already-reserved `phase.uat-passed` alias (subcommand `uat-passed`,
mutation:false) into the phase command router with a new markdown-aware
predicate that evaluates HUMAN-UAT results and reports pass only when every
required check passes. Post-SDK-retirement (ADR-0174/#174) successor to the
SDK-framed #70, with no SDK-specific API surface.

New pure module src/uat-predicate.cts:
- stripFalsePositiveContexts: frontmatter -> HTML-comment -> CommonMark-style
  fenced-block state machine (tracks delimiter char+length) -> blockquote,
  each a small composable step, so a `result: passed` inside frontmatter, a
  fenced/~~~ block (incl. ~~~ nested in a ``` fence), a comment, or a
  blockquote is never counted.
- parseUatResultItems: heading-block parser, column-0-anchored same-line
  result; a heading with no result -> `missing` (fail-closed).
- analyzeMarkdown: unterminated fence/comment detection (malformed -> blocker).
- evaluateUatPassed: allowlist pass/verification semantics; passed = no
  blockers && >=1 check && all passing; no_uat_artifacts discriminator (no
  vacuous pass); optional requireVerification policy hook.

Thin cmdPhaseUatPassed handler in phase.cts; router closure rejects unknown
flags via makeInvalidArgs. Hardened across two Codex adversarial passes
(vacuous pass, dropped failing tests, permissive verification status,
nested-fence escape, cross-line result value, masked unterminated comment) —
all fixed fail-closed. New unit + CLI-integration suites incl. a fast-check
property test; docs, CONTEXT glossary, inventory, and changeset updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#247): backfill changeset PR number (#1063)

* fix(#247): indexOf paired-scan for unterminated-comment detection

CodeQL js/incomplete-multi-character-sanitization (high) flagged the
`raw.replace(/<!--[\s\S]*?-->/g,'')`-then-`.includes('<!--')` detection in
analyzeMarkdown as incomplete sanitization (a single regex pass can leave a
residual `<!--`). Replace it with a paired left-to-right indexOf scan that
contains no `.replace()` of the comment token — CodeQL-clean and strictly
more correct (a closed earlier comment can never mask a later unterminated
one). Behaviour unchanged; 98 predicate tests + scoped docker run green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 16:37:17 -04:00
Tom Boucher
8813ee5f95 feat(#429): HARD GATE on negative-grep literals echoed in plan <action> bodies (#1062)
Convert the planner's soft comment-text guideline into a plan-write-time
HARD GATE. When an acceptance criterion negative-greps for a literal
(`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an
`<action>` body (JSDoc samples, head-comment references, "what NOT to do"
snippets), the executor's commit-time verify gate later fails on the
comment echo rather than a real regression — wasting cycles and training
the executor to distrust the gate.

`verify.plan-structure` (the `validate_plan` step) now scans for this:
- confidently-extracted (quoted) negative-grep literal echoed in an
  <action> → error (valid:false), failing plan creation
- unquoted/ambiguous grep target → warning (fallback policy)
- `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal
- positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope

Adds the `<comment_text_discipline>` block to gsd-planner.md, the full
rules + allowlist example to planner-antipatterns.md, and regression
fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary
case proving positive-count gate 11-02 is not flagged).

Closes #429

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 15:46:49 -04:00
Tom Boucher
0cc37a94c2 feat(#1056): phase 5e — close ConverterName enum + configFormat↔installSurface parity guard (#1057)
Two gen-time validation tightenings (validation-only; serialized registry content
unchanged; bin/install.js + adapter + descriptors untouched):

Part B: validateArtifactKindEntry now requires artifactLayout[].converter ∈
VALID_CONVERTER_NAMES (15 names, all exported by install.js) ∪ {null} — a typo'd
converter fails at gen time instead of silently → installExports[name]===undefined
at install time.

Part A: a HARD buildRegistry parity gate asserts each runtime descriptor's
configFormat agrees with the adapter registry's installSurface via a fixed mapping
(cursor-hooks-json/profile-marker-only→none, codex-toml→toml, copilot-instructions→
markdown, cline-rules→markdown-dir, settings-json→settings-json) — keeps configFormat
from drifting; prerequisite-validation for the deferred full drive (#1055).

The full config-writing drive (retire resolveRuntimeConfigIntent) is deferred to
#1055: configFormat is lossy vs installSurface (cursor vs profile-marker both → none;
opencode/kilo permissionWriter has no descriptor field) → needs schema extension.

Closes #1056

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 14:40:45 -04:00
Tom Boucher
4698b3e349 fix(#1051): force-exit + per-chunk timeout for the windows full-test lane; close leaked test handles (#1054)
The `full test (windows-latest, 22)` job intermittently got CANCELLED at its
20m wall-clock cap with no failed test step — a false-negative gate (recurrence
of #869). Root cause: a unit test leaves an open event-loop handle, so the
chunk's `node --test` child hangs ~150s on Windows after its last test prints;
two such stalls push the already-~13m job past 20m.

Fix (defense in depth):
- run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0)
  so the runner exits once all tests finish regardless of lingering handles —
  the durable backstop. Account for the flag in the argv-length ceiling.
- run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env
  RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the
  chunk's files, so a hung chunk can never silently eat the job budget.
- perf-316 test: terminate both Worker threads on all paths (afterEach +
  finally) so they cannot outlive the test.
- locking-bugs test: kill spawned children in a finally that wraps the whole
  spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier
  timeout no longer leaks live child processes.
- Refresh the stale synckit comment (synckit/SDK bridge was removed).

Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout
and fails with a clear message; force-exit lets a chunk with a leaked handle
exit cleanly.

Closes #1051

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 14:11:50 -04:00
Tom Boucher
58ed55683e feat(#1049): phase 5d — drive artifactLayout from the runtime descriptor (retire the 128-LOC switch) (#1053)
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.

Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.

validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.

Closes #1049

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 14:08:46 -04:00
Tom Boucher
cec7e704d6 feat(#435): expand workflow-policy linter to full Cartesian matrix cross-product (#1050)
`expandRunsOn` enumerated a multi-axis `strategy.matrix` one key at a time,
producing partial realization contexts. A true `os × shell` matrix therefore
left `${{ matrix.shell }}` unresolvable against any `{ os: ... }`-only context,
firing spurious `UNRESOLVABLE_MATRIX` violations and leaving shell-pinning
coverage incomplete on Cartesian jobs.

Enumerate the full GitHub Actions cross-product of all base-list matrix keys
(every `matrix.<k>` array, excluding the `include`/`exclude` control keys) via
a named `cartesianProduct` helper. Each realization's context now carries a
value for every matrix key, so `${{ matrix.<key> }}` resolves per realization.
Single-axis matrices keep byte-for-byte identical output; only multi-axis
matrices change shape. The `include` and `exclude` blocks are unchanged
(full tuple-aware exclude is a documented out-of-scope follow-up).

Tests: updated the `os × shell` test to assert post-fix behavior (4 step
realizations, 2 WRONG_SHELL_FOR_OS, 0 UNRESOLVABLE_MATRIX); added a compliant
`os × node-version` cross-product test; added a fast-check property test that
the realization count equals the product of axis lengths and that no axis key
is dropped from any context.

Closes #435

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 13:36:47 -04:00
Tom Boucher
9e3b056b15 fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779.
2026-06-11 13:36:23 -04:00
Tom Boucher
734c56ccfe feat(#1046): phase 5c — drive commandStyle from the runtime descriptor (runtime-slash codex-check → lookup) (#1048)
formatGsdSlash now reads commandStyle from registry.runtimes[id].runtime.commandStyle
(lazy require of the committed capability-registry.cjs) instead of the hardcoded
if (rt === 'codex'). Equivalence-preserving (Codex-verified): codex (shell-var) →
$gsd- + lowercased token; all 15 others (slash-hyphen) + unknown → /gsd- +
case-preserved token. canonicalizeRuntimeName + input normalization + claude default
preserved; no circular load (mirrors 5b runtime-homes pattern).

Added a registry-parity test: 16 parametrized sub-tests derive the expected prefix +
lowercasing from each runtime's commandStyle, proving the prefix is a pure function of
the registry (catches future hardcode-vs-registry divergence).

Closes #1046

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 12:37:25 -04:00
Tom Boucher
6968e04d8a feat(#1040): phase 5b — drive configHome from the runtime descriptor — ADR-857/1016 (#1043)
* feat(#1040): phase 5b — drive configHome from the runtime descriptor (runtime-homes switch → lookup)

getGlobalConfigDir now resolves configHome from registry.runtimes[id].runtime.configHome
via a single resolveConfigHomeFromDescriptor(configHome, {env, home, existsSync})
(dot-home / dot-home-nested / xdg / generic-agents-root), replacing the hardcoded
16-runtime switch. Equivalence-preserving: byte-identical config dirs for all 16
runtimes + grok + default + explicitDir (Codex-verified, no divergence).

Nuances preserved: xdg env[1] is a FILE path → path.dirname; existsSync injection
seam keeps antigravity/kimi probe tests hermetic; grok stays hardcoded
(GROK_AGENTS_HOME → ~/.agents, not in the 16); copilot two-env fallback; explicitDir
short-circuit. getGlobalSkillsBase unchanged (out of scope). Lazy require of the
committed capability-registry.cjs (no circular load).

Test env-clearing lists (install.test ENV_KEYS, bug-3126 envKeys) now derived from
the registry runtime configHome.env arrays — auto-correct, closes the missing
KIMI_CONFIG_DIR gap. New 81-case golden equivalence test.

Closes #1040

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1040): make windsurf golden path Windows-portable (path.join, not POSIX literal)

The dot-home-nested windsurf equivalence case hardcoded '/home/u/.codeium/windsurf'
but the resolver builds it via path.join(home,parent,name) → backslashes on Windows.
Use path.join for the expected value. Test-only; production resolver unchanged.
Defensive scan confirmed it was the only path.join-derived hardcoded literal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 12:08:11 -04:00
Tom Boucher
728788fdb3 fix(#1034): correct stale intel.cts comments to canonical INTEL_FILES names (#1036)
Three maintainer-facing comments in src/intel.cts referenced the retired
short/markdown intel filenames (files.json, deps.json, arch.md) and described
arch as searched "text lines", though the implementation iterates the exported
INTEL_FILES map and parses every entry — arch included — as JSON.

Update the comments to cite the canonical names (file-roles.json,
dependency-graph.json, arch-decisions.json) and the JSON-for-arch behavior.
Comment-only; no code or behavior change. Deferred from the #1000 fix.

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:55:55 -04:00
Tom Boucher
1fab2e10ba fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver (#1045)
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver

Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker,
gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks.
On a shim-only install — where gsd-tools.cjs exists under the runtime home but
gsd-tools is NOT on PATH — those calls fail with "command not found" and the
agent silently skips init/state/validate/commit ceremony, deferring to the
orchestrator or bypassing GSD bookkeeping entirely.

were never migrated, so it persisted on Claude Code and every other runtime that
consumes the source agents directly. Only gsd-phase-researcher.md carried a
resolver — and a stale, claude-only truncated one.

Fix (all runtimes):
- Inject the canonical multi-runtime gsd_run preamble (byte-equal to
  _runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/
  augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of
  the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every
  command-position bare gsd-tools to gsd_run.
- Upgrade gsd-phase-researcher.md's stale resolver to the canonical one.
- Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the
  sync caught and corrected a mis-placed preamble during development).
- Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards
  to agents/ so no runtime can silently regress.

Closes #1041

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1041): backfill changeset PR number to 1045

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:42:59 -04:00
Tom Boucher
e4dfa6b9ea fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer

The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:

1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
   --changed-since/--base for changed-files scoping (no file-list input), and
   --max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
   0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
   the output — i.e. it threw away exactly the findings it exists to surface.
   Success is now decided by whether a valid fallow JSON report was produced,
   not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
   (unusedExports/duplicates/circularDependencies) fallow never shipped, and was
   dead code (the workflow embedded raw JSON; its tests asserted the fictional
   schema, one even calling a non-existent runFallowAudit and passing vacuously).

Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.

Closes #1012

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1012): backfill changeset PR number to 1044

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:50 -04:00
Tom Boucher
1e3ce6df05 fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline (#1042)
* fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline

The gsd-research-synthesizer agent intermittently hits an LLM false-refusal:
instead of writing .planning/research/SUMMARY.md with the Write tool, it returns
the SUMMARY.md content inline and fabricates a non-existent write restriction
(e.g. "the runtime is blocking file writes"). The shipped prompt hardening
(#240) is necessary but insufficient — the false-refusal recurs under some
context loads, and a drifting subagent then leaves gsd-roadmapper to fail with
"SUMMARY.md not found".

Adds an orchestrator-level self-heal to new-project.md and new-milestone.md:
after the synthesizer returns, verify .planning/research/SUMMARY.md exists; if it
is missing but the agent returned content inline, the orchestrator persists that
content with the Write tool (logging a warning) before spawning gsd-roadmapper;
if missing with no content, surface the error and stop rather than proceed
against a missing SUMMARY.md. This absorbs the failure mode deterministically
instead of depending on the subagent never drifting.

Closes #222

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#222): backfill changeset PR number to 1042

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:39 -04:00
Tom Boucher
ab8b84286e fix(#1013): resolve worktree.baseRef from user/global settings cascade (#1038)
* fix(#1013): resolve worktree.baseRef from user/global settings cascade

cmdWorktreeBaseCheck resolved worktree.baseRef from the project checkout's
.claude/ only (settings.local.json then settings.json). A user/global
worktree.baseRef:"head" — the layer /config writes and the harness honors,
and the only sensible place for a machine-wide preference — was invisible. On a
phase lane (HEAD ahead of origin/HEAD, or no origin/HEAD symref) base-check
returned shouldDegrade:true and execute-phase forced sequential execution,
silently losing the parallel worktree execution the user configured.
CLAUDE_CONFIG_DIR (relocated user config dir) was also ignored.

resolveEffectiveBaseRef now accepts an optional user/global config dir and reads
its settings.json as a third, lowest-precedence layer (project local > project
shared > user/global). cmdWorktreeBaseCheck resolves it via
getGlobalConfigDir('claude'), which honors CLAUDE_CONFIG_DIR. The existing
project-level reads stay as higher-precedence overrides and the injectable
readFile seam is preserved, keeping the unit tests hermetic.

Closes #1013

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1013): backfill changeset PR number to 1038

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:30 -04:00
Tom Boucher
12d6285b4f fix(#1000): align gsd-intel-updater output to canonical intel filenames (#1037)
* fix(#1000): align gsd-intel-updater output to canonical intel filenames

The intel-updater agent was instructed to write short names (files.json,
apis.json, deps.json) and a markdown arch.md, but the intel library + gsd-tools
intel CLI read only the canonical long names from INTEL_FILES (file-roles.json,
api-map.json, dependency-graph.json, arch-decisions.json as JSON). After
/gsd:map-codebase --query refresh the agent output was orphaned — intel status
and validate reported the canonical files missing and intel query returned
nothing.

Renames every short reference to its INTEL_FILES canonical name and converts the
arch output from markdown to queryable arch-decisions.json. Adds a drift-proof
regression test (derived from the exported INTEL_FILES map) in the owning
module's test file tests/intel.test.cjs, reviving the maintainer-approved
approach from closed PR #608.

Closes #1000

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1000): backfill changeset PR number to 1037

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:18 -04:00
Tom Boucher
5599e5a9cf fix(#1004): detect http-route hook registrations in installer presence check (#1032)
* fix(#1004): detect http-route hook registrations in installer presence check

referencesHook only inspected h.command and h.args, so a managed hook
re-registered as a type:"http" entry (local hook-server routing) — whose
identity lives only in h.url — was invisible. The installer then appended a
stock command duplicate on every install/update, running the hook twice per
event. Adds the h.url arm, mirroring the #976 args-form fix.

Closes #1004

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1004): backfill changeset PR number to 1032

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:36:03 -04:00
Tom Boucher
fbd62cd84f feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).

Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).

Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).

Closes #1035

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 09:34:07 -04:00
github-actions[bot]
b2222fa769 docs(#860): ADR-3660 addendum — Nth-runtime-via-existing-layout is an addendum, not a new ADR
Maintainer governance decision (waiving a standalone ADR for Qoder, PR #1021):
registering a runtime that reuses the existing profile-marker-only install
surface + a single skills kind is an enhancement governed by ADR-3660 via
addendum. Records the addendum-vs-new-ADR qualifying criteria, the normative
agent-frontmatter contract (name+description only — the sibling converter
shape), and Qoder as the first runtime logged under this path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 08:39:22 -04:00
Tom Boucher
8bd5e07c58 feat(#1031): autonomous §3a.5 plan:pre ui-phase cutover — last inlined ui-phase site (step-only) (#1033)
Cut over autonomous.md §3a.5 (autonomous plan:pre ui-phase step) to the
loop.render-hooks plan:pre dispatch — completing the ui-phase migration begun
in #1026 (plan-phase.md §5.6). Step-only, non-blocking: autonomous is always
pipeline, so it fires active kind==step hooks and never runs the manual-only
plan:pre blocking gate.

Skip condition keys on "no active step hooks" (not empty activeHooks), so the
gate-only {ui_phase:false, ui_safety_gate:true} case skips silently with no
spurious warning — matching OLD §3a.5. Fires gsd-ui-phase under the identical
precondition (frontend + no UI-SPEC + workflow.ui_phase active), bare
${PHASE_NUM} args. Replaces the inline ui-safety-gate.cjs probe + config-get
with render-hooks + the ui.plan-gate check verb.

Codex caught the gate-only spurious-warning divergence on the first pass; fixed
+ re-confirmed equivalence-preserving. gsd-ui-phase skill, §5.6, §3d.5 untouched.

Closes #1031

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 08:31:35 -04:00
Tom Boucher
9ed8c7d574 feat(#1026): §5.6/ui-phase cutover — first gate dispatch (plan:pre step + blocking gate) (#1028)
Replace plan-phase.md §5.6 (a 6-branch inline UI gate) with a capability-driven
loop.render-hooks plan:pre dispatch — the FIRST gate dispatch in any workflow.
A step (ui-phase, when:workflow.ui_phase) + a new blocking gate
(when:workflow.ui_safety_gate). New ui.plan-gate check verb returns
{frontend, hasUiSpec, block}; the dispatch runs it unconditionally then fires
the active step (pipeline) or halts on the active blocking gate (manual). The
gate-handling (run check.query; halt if blocking+block) is the reusable
phase-6 template for blocking-gate cutovers.

Config semantics fixed per #1022 + maintainer call: ui_phase gates plan-time
UI-SPEC generation, ui_safety_gate gates the planning block. Common case + all
ui_phase=false cases are equivalence-preserving; the one intended change is
{ui_phase:true, ui_safety_gate:false} now auto-generating in pipelines.

Review found it broken twice (non-generic dispatch, phase-lookup divergence,
then the step-only check nested in a gate loop) — fixed; final Codex pass
verified all 8 (ui_phase,ui_safety_gate)x{pipeline,manual} cases correct.
gsd-ui-phase skill + autonomous §3a.5 untouched (§3a.5 deferred).
getRoadmapPhaseWithFallback mirrors cmdRoadmapGetPhase for lookup parity.

Closes #1026

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 00:58:21 -04:00
Tom Boucher
5f48a42514 docs(#1022): resolve step-vs-gate model question — steps additive, gates block, mode self-gates (#1025)
Record the resolution of #1022 (surfaced scoping the §5.6 ui-phase cutover):
a step is purely additive and never halts the host; host-blocking preconditions
are gates (blocking/onError:halt) — no hook-model change. Runtime/mode context
(auto/chain vs manual) self-gates in the skill, not via when (config-only).

§5.6 decomposes into the existing plan:pre step (ui-phase, self-gates on
frontend + pipeline) + a new plan:pre gate (frontend-and-no-UI-SPEC → halt,
when: workflow.ui_safety_gate) that blocks planning in manual mode — preserving
the "run /gsd:ui-phase first" UX (maintainer call: pipelines-only auto-fire).
The render-hooks dispatch template grows to handle gates, not just steps.

Recorded in ADR-894 (clarification) + CONTEXT.md
(RULESET.CAPABILITY.step-additive-gate-blocks). Unblocks the §5.6 cutover.

Closes #1022

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 22:51:31 -04:00
Tom Boucher
2ac6592096 feat(#1023): first phase-6 cutover — ui-review (verify:post) inline → loop.render-hooks dispatch (#1024)
Replace the inlined ui-review invocation in autonomous.md §3d.5 with a
loop.render-hooks verify:post dispatch — the first workflow to consume
render-hooks and fire a skill from it (closes the #1018 live-execution residual
as real wiring). Capability-driven, equivalence-preserving for the current
registry (only ui-review at verify:post, default on): fires gsd-ui-review under
the same precondition (UI-SPEC exists via consumes-gate + workflow.ui_review).

Gate findings (real pattern issues, fixed so every future cutover inherits them):
- bug-2643 static "Skill() references a real skill" check vs templated
  Skill(skill="gsd-${ref.skill}") dispatch → skip ${...}-templated names.
- Coverage moved, not lost: gen-capability-registry now validates
  steps[].ref.skill in skills + ref.agent in agents + rejects gsd- double-prefix.
- Tightened §3d.5 tests; markdown clarity (consumes rule, LLM-native JSON read,
  UI-REVIEW.md score hint).

gsd-ui-review skill + §3a.5/ui-phase untouched. §5.6/ui-phase cutover deferred
(#1022 step-can-halt-vs-gate model question).

Closes #1023

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 22:40:58 -04:00
Tom Boucher
7eea991884 test(#1018): hook-firing de-risk spike — structural off-means-off proof + cutover contract (#1020)
Spike disposition (a): prove the loop.render-hooks mechanism at the data level
and document the cutover-shape finding, without touching plan-phase.md §5.6.

- tests/loop-hook-firing-spike.test.cjs: makes "off means off" executable — a
  host-computed aggregate (hostConsume) that is a pure function of the active
  hook set; UI active-by-default → 1, off → 0 == independently-computed base
  (not tautological); synthetic 2-hook ordering; second point verify:post.
- CONTEXT.md RULESET.CAPABILITY.cutover-self-gating: the loop hook is coarse;
  phase-context detection + mode self-gate in the skill (ADR-894); §5.6 UI gate
  is the worked example of what must move into gsd-ui-phase before cutover.

Finding: the UI plan:pre hook is already inlined (§5.6) — cutover is move
host-logic-into-skill, not a 1:1 swap. Live LLM-execution of injected hook
markdown is the residual, carried into the first phase-6 cutover acceptance.

Closes #1018

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 21:40:30 -04:00
Tom Boucher
61b29d2af6 docs(#1016): ADR-1016 runtime capability descriptor (ADR-857 phase 5 design) (#1019)
Realize ADR-857 Branch 8 (host-CLI support as role:runtime Capabilities) +
materialize ADR-58's InstallPlan. Closes the 6-axis descriptor vocabulary,
absorbs the hard-case runtimes as data, and stages the install migration as a
5a→5g ladder (4/6 axes already modularized; InstallPlan is the 5g capstone,
reachable by collection, not a big-bang rewrite).

Amended after a plugin-side design grill (rubber-duck + grill-with-docs vs
#956 MemPalace / #999 Impeccable):
- New term Connected Capability (CONTEXT.md) — a Capability whose integration
  shape brings its own external process/service/state; orthogonal to authorship.
  Named, tracked gap (vehicle #956); current schema does not express it.
- Narrowed the dogfood claim: the descriptor dogfoods the runtime interface
  only, not the feature-plugin/Connected path.
- Structural "off means off" rule (CONTEXT.md RULESET.CAPABILITY.off-means-off):
  the host derives shared outputs from active hooks; a hook adds/is-counted,
  never mutates host source. Ratify in ADR-894; proven by spike #1018.
- Hand-waves resolved by code: sandboxTier real-but-thin; model-catalog
  orthogonal (not a 7th axis); converters closed into a ConverterName enum +
  added the kimi-agents artifact kind; configHome is pure read-only.
- Hook-firing path is unproven (render-hooks built, never consumed) → spike
  #1018 must prove render-hooks→live-workflow execution before phase-5 build.

Design-only; no code. Status: Proposed.

Closes #1016

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 21:00:37 -04:00
Colin Johnson
76f42ddb4b feat(#1014): add Claude Fable 5 model config (#1015) 2026-06-10 20:32:26 -04:00
Rezolv
28ac89d810 fix(#1008): tolerate EAGAIN + short writes in io output()/error() (#1009)
* fix(#1008): tolerate EAGAIN + short writes in io output()/error()

The I/O Module wrote stdout/stderr with a bare fs.writeSync(fd, data), assuming
it blocks until the kernel accepts every byte. That is false when the fd is a
non-blocking pipe (as under the parallel node:test runner on Linux CI): a full
pipe throws EAGAIN and a partially-drained pipe returns a short count. The former
caused spurious failures (e.g. bug-974 graphify property test threw EAGAIN); the
latter risked silently truncating output.

Add writeAllSync(fd, data): loop on short counts and retry EAGAIN/EINTR with a
bounded backoff. The backoff sleep buffer is allocated lazily on the first retry
(rare) and reused — keeping it out of module load avoids perturbing the
SharedArrayBuffer-allocation accounting in perf-316 and costs nothing on the
common no-retry path. Route output() and error() through it; non-transient
errors (EPIPE) still propagate. Mirrors the transient-errno handling already
applied to STATE.md lock acquisition (ACQUIRE_LOCK_RETRY_ERRNOS / #3776).

Regression cases live in tests/io.test.cjs (the owning module's file, per the
regression-test-name placement policy) and inject fs.writeSync via mock.method:
EAGAIN/EINTR retry, short-write no-truncation, EPIPE still surfaces, and error()
retries while still exit(1). Red against the pre-fix bare-writeSync io.cjs.

* chore(#1008): add Fixed changeset for io EAGAIN/short-write fix
2026-06-10 17:05:07 -04:00
Jeremy McSpadden
092340d18a fix(#711): wire autonomous convergence flag (#729)
* fix(#711): wire autonomous convergence flag

* Update wise-ibex-tumble.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 16:10:07 -04:00
Tom Boucher
19edab21da fix(#1006): rc CHANGELOG preview crash on malformed changeset fragment + validate fragment content at the gate (#1007)
* fix(#1006): harden render --preview against fragment parse failures

`render --preview` wrote `report.preview` unconditionally. When a `.changeset`
fragment fails to parse, `cmdRender` early-returns with `{exitCode:1, report:
{failures}}` and NO `preview` key, so `process.stdout.write(undefined)` threw
ERR_INVALID_ARG_TYPE and the rc release job's "Preview CHANGELOG" step died with
a cryptic TypeError that masked the real cause.

Guard the preview write on `typeof report.preview === 'string'` (ADR-227: shape,
not just type); when absent, fall through to the existing failure reporter that
names the offending fragment and exits non-zero — identical to a non-preview
render. Also backfills the stray placeholder `pr: 0` -> `pr: 939` in
.changeset/936-convergence-inline-plan-phase.md that triggered the live failure.

Regression test (red-then-green verified) added at the render --preview seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1006): validate changeset fragment content at the Changeset Required gate

The `Changeset Required` gate (scripts/changeset/lint.cjs) only checked that a
`.changeset/*.md` fragment EXISTS in the PR diff; it never validated the
fragment's contents. So a malformed fragment (e.g. an un-backfilled `pr: 0`
placeholder) silently merged to `next` and only detonated later in the rc
release job. This is the upstream prevention for #1006 — the crash hardening
turns the failure into a clear message, this stops the bad fragment ever
reaching the release path.

evaluateLint now accepts `fragmentFailures` and fails with the typed reason
`fail_invalid_fragment` (naming each offending file) before the existence/
opt-out checks — a malformed fragment beats `no-changelog`, since it will break
the render regardless. main() reads + parseFragment()s every changed fragment:
a deleted fragment (not on disk) is skipped, a present-but-unreadable one fails
closed. Tests assert on the typed LINT_REASON enum (no raw-text matching), a
precedence case over the opt-out label, and an end-to-end suite that drives the
real main() against a temp git repo (malformed -> fail, valid -> pass, deleted
-> skipped) so the wiring is regression-proof.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1006): assert the typed --json report in the preview regression test

Code review flagged the preview parse-failure regression test for positive
raw-text matching on CLI output (`combined.includes('bad-fragment.md')` /
`'invalid_pr'`), which this repo's testing standards forbid. Keep the non-json
`runRenderRaw` call for the negative crash proof (the ERR_INVALID_ARG_TYPE
crash lives only on the non-json stdout.write path), and add a `--json`
invocation that asserts the offending fragment + typed `invalid_pr` reason via
the structured `report.failures[]` surface instead of rendered prose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 15:55:31 -04:00
Jeremy McSpadden
f61b97276e fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings

* merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 14:54:15 -04:00
Jeremy McSpadden
fb37fa7dd5 fix(#725): route Codex gsd-tools calls through shim (#731)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 14:44:43 -04:00
Joe
5e8a723089 feat(templates): add optional Business Context section to PROJECT.md template (#756)
* feat(templates): add optional Business Context section to PROJECT.md template

Adds an optional `## Business Context` section (Customer, Revenue model,
Success metric, Strategy notes) between Core Value and Requirements, for
monetized or customer-facing projects. Optional by default — an HTML comment
tells non-business projects to delete it; capped at four one-line fields to
stay a constraint reference, not a business plan. The milestone evolution
review in complete-milestone.md checks it only when the section is present.

Refs #72

* chore(changeset): set pr number for #72 fragment

* test(#72): add source-text-is-the-product exemption marker

Addresses review Minor #1 on PR #756. The contract test reads the
PROJECT.md template and complete-milestone workflow .md files and
asserts on their content (the local/no-source-grep pattern). Those
.md files ARE the product surface, so this is a valid
source-text-is-the-product case. Add the explicit // allow-test-rule
marker per RULESET.TESTS.no-source-grep.exemption so intent is
audit-traceable before the rule promotes to error (#453).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-10 14:13:01 -04:00
Tom Boucher
4c10eb2253 fix(#991): inject configured agent_skills into code-review family subagents (#1005)
* fix(#991): inject configured agent_skills into code-review family subagents

code-review.md, code-review-fix.md, and eval-review.md spawned their
subagents (gsd-code-reviewer / gsd-code-fixer / gsd-eval-auditor) without
querying or injecting the project-configured agent_skills, while ~20 sibling
workflows do. Subagents don't inherit the orchestrator's auto-loaded context,
so this injection is the only channel — reviewers/fixers/auditors silently ran
without the configured rule/skill context.

Mirror the established sibling idiom: add
`VAR=$(gsd_run query agent-skills <agent-type>)` in each workflow's initialize
step and interpolate `${VAR}` into every Agent() spawn of that type. This
covers all spawn sites, including code-review-fix.md's --auto loop which
re-spawns gsd-code-reviewer in addition to the two gsd-code-fixer spawns.

Regression test reads the workflow text (source-text-is-the-product) and
asserts each file queries agent-skills for every agent type it spawns and
interpolates the result at least once per spawn.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#991): add changeset for code-review agent_skills injection fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 14:07:23 -04:00
github-actions[bot]
0976d849f2 chore: remove stray orchestration temp files (issue-996-regression.md, pr-1001-body.md) leaked via #1002 git add -A 2026-06-10 14:01:08 -04:00
Tom Boucher
adaf3e17d8 fix(#1001): make bug-969 hardening tests hermetic + move build tsbuildinfo out of shipped tree (regression from #996) (#1002)
* fix(#969): make bug-969 hardening tests hermetic and move build tsbuildinfo out of shipped tree (regression from #996)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#969): self-heal legacy bin-local tsbuildinfo and make sentinel test hermetic (adversarial-review follow-ups)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): set pr number to 1002

* docs(#1001): record DEFECT.SHARED-ARTIFACT-MUTATION-IN-CONCURRENT-TEST anti-pattern in CONTEXT.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 13:58:56 -04:00
Tom Boucher
88e30d5342 test(#969): fix stale-build flake (incremental + re-emit-on-missing) and make runGsdTools retry-once before surfacing subprocess kills (#996)
Closes #969

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 12:04:59 -04:00
Tom Boucher
1fd5c86a1e fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity) (#995)
* fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity)

Both convertClaudeToWindsurfMarkdown and convertClaudeToTraeMarkdown only
handled trailing-slash .claude/ forms; bare ~/.claude and $HOME/.claude
references (e.g. configDir = ~/.claude, RUNTIME_CONFIG_DIR=".../$HOME/.claude")
survived conversion and pointed users at the wrong config dir.

Fix: add bare-form replacements using negative lookahead (?![\w-]) to protect
.claude-plugin and .claudeignore, mirroring Cline (#782) and Codex (#570) precedent.
Also adds CLAUDE_CONFIG_DIR -> WINDSURF_CONFIG_DIR / TRAE_CONFIG_DIR rewrite.
_applyRuntimeRewrites windsurf case gets matching \b-anchored bare-form lines,
mirroring the existing trae case.

Closes #983

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#983): backfill changeset pr number (995)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:43:31 -04:00
radioflyer28
d617735bed fix(codex): avoid partial model effort pinning (#842)
Co-authored-by: Andrew Kriz <akriz@vt.edu>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-10 11:41:01 -04:00
Tom Boucher
36b68ac81d fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath (#992)
* fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath

Closes #977

* chore(#977): backfill changeset pr number (992)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 11:16:13 -04:00
Tom Boucher
972a41a528 fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:)

Closes #967

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#967): backfill changeset pr number (990)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 11:10:40 -04:00
Tom Boucher
921a7cd618 fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op (#986)
* fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op

When `--budget` was the last arg or followed by a non-numeric token,
parseInt(undefined/NaN-string, 10) produced NaN. NaN is falsy so both
the router check and applyBudget gate silently skipped budget trimming.
The query ran unbounded with no warning.

Fix: guard in graphify-command-router.cts — if args[budgetIdx+1] is
absent or parses to NaN, emit ERROR_REASON.USAGE and return early.
Defensive fix in graphify.cts: tighten `if (!budgetTokens)` →
`if (budgetTokens == null)` and `if (options.budget)` →
`if (options.budget != null)` so a real 0/NaN caller is handled
predictably by both independent guards.

Regression tests: 16 cases (unit/mock, subprocess, property-based)
covering boundary inputs: missing value, non-numeric, valid integers,
and fast-check properties over the budget parse contract.

Closes #974

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#974): backfill changeset pr number (986)

* fix(#974): convert static property (b) to real fc.assert property with constrained generator + deterministic seed

The test named "property: --budget as last arg always produces usage error"
was a static test with no fc.assert — it only checked a single hardcoded
term ("someterm") and could never flake or produce a fast-check path. This
is a generator/property bug (case b): the test was mislabeled as a property
test but lacked parameterization.

Root-cause of CI path "46:0:0" / seed 42 failure: a naive parameterized
version without the !startsWith('--') filter could feed term='--budget',
causing args.indexOf('--budget') to hit index 2 (the term slot) rather than
index 3 (the flag slot), placing the router in a different code path. The
property still holds — NaN detection fires on rawBudget='--budget' — but
the assertion text referenced the wrong invariant, making the failure appear
spurious. Fix: constrain the generator to non-flag terms (filter out
strings starting with '--') and pass explicit { seed: 42, numRuns: 200 } to
fc.assert so the test is fully deterministic in CI regardless of GSD_FC_SEED.

No change to src/graphify-command-router.cts (router is correct).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): constrain non-numeric budget property generator to genuinely-NaN values and pin fast-check seeds (CI determinism)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): use a strict valid-term generator and pin seeds so budget property tests are deterministic

Replace fc.string({minLength:1}).filter(!startsWith('--')) term generators in
properties (b) and (d) with a shared validTerm = fc.stringMatching(/^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/)
that is alphanumeric-leading and contains no whitespace, flags, or sign-numerics.
This eliminates the class of CI failures where the old generator produced out-of-
contract inputs (" ", "+5", empty) that the router legitimately rejects for reasons
outside the budget-parse contract under test. Pin distinct seeds (1001–1004) on
every fc.assert for CI determinism. Stress-tested at numRuns=100000 per property
and across 10 seeds (1,2,7,13,42,43,44,99,12345,999999) at numRuns=2000 — all pass.
No src change: gsd-core/bin/lib/graphify-command-router.cjs is correct.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#974): replace flaky term-fuzzing properties with deterministic examples; keep value-fuzzing properties

The validTerm regex /^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/ admitted single-digit
strings like "0" which are falsy; the router's `if (!term)` guard fires before
the budget-missing-value path, producing a spurious errFn call. Properties (b)
and (d), which test the BUDGET contract (not term handling), are replaced with
deterministic example loops over fixed valid terms. Properties (a) and (c),
which fuzz the BUDGET VALUE with a fixed term, are kept unchanged (seed+numRuns
pinned). The validTerm generator is fully removed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:58:42 -04:00
Tom Boucher
898d55788e fix(#976): detect command+args (wrapped) hook registrations in installer presence checks (#994)
* fix(#976): detect command+args (wrapped) hook registrations in installer presence checks

Add referencesHook() helper that inspects both h.command (standard form) and
h.args[] (args-form / wrapped-launcher form) when checking whether a managed
hook is already registered.  Rewrite all has*Hook predicates and the
alreadyHas* guards to use it so args-form registrations suppress the duplicate
stock string-command entry that was previously appended on every install/update.

Also add an explicit args-form skip to rewriteLegacyManagedNodeHookCommands so
entries with a non-empty args[] are left untouched (they are intentional user
wrappers, not legacy bare-node commands to migrate).

Extend isManagedHookCommand() in shell-command-projection.cts with an optional
args: unknown[] parameter that checks whether any arg's basename matches the
managed hook surface set — backward compatible; existing callers are unaffected.

Regression test added to tests/install-regressions.test.cjs:
- two-pass install with an args-form SessionStart entry pre-written to
  settings.local.json asserts exactly 1 hook entry remains after reinstall
  (previously 2 — the original args-form + a new stock string-command duplicate)
- rewriteLegacyManagedNodeHookCommands test asserts args-form entries unchanged

Closes #976

* chore(#976): backfill changeset pr number (994)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:44:24 -04:00
Tom Boucher
2981983bae fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md (#989)
* fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md

gsd-planner shipped Write but not Edit — the same writer-agent gap fixed for six
agents in #571/#581. Without Edit, an in-place ROADMAP update fell back to a
whole-file Write that truncated committed milestone history (292→16 lines in a
real incident).

Changes:
- agents/gsd-planner.md: add Edit to tools: frontmatter (adjacent to Write)
- agents/gsd-planner.md: update_roadmap step now directs Edit (scoped), with an
  explicit blocking prohibition on whole-file Write of ROADMAP.md or any existing
  curated .planning/ file
- agents/gsd-planner.md: Write contract section clarifies Write is authorized only
  for net-new PLAN.md creation; existing files must use Edit
- tests/agent-frontmatter.test.cjs: extend SECTION_WRITER_AGENTS list (#581 test)
  to cover gsd-planner — fails before fix, passes after
- .changeset/973-gsd-planner-edit-tool.md: Fixed changeset, pr:0

Closes #973

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#973): backfill changeset pr number (989)

* fix(#973): trim gsd-planner.md prose under agent size cap (keep Edit + scoped-Edit-for-ROADMAP rule)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:44:20 -04:00
Tom Boucher
22bb82e205 fix(#965): emit structured json error for unexpected handler throws under --json-errors (#987)
* fix(#965): emit structured json error for unexpected handler throws under --json-errors

When GSD_JSON_ERRORS=1 / --json-errors is active, an unexpected (non-ExitError) throw
in a handler now emits { ok: false, reason: "sdk_fail_fast", message } to stderr instead
of a raw stack trace. The plain-text behaviour (no json-error mode) is unchanged.

Closes #965

* chore(#965): backfill changeset pr number (987)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:40:12 -04:00
Tom Boucher
626575cbc5 fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works (#982)
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works

The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so
`options.force` was always `undefined` and the guard inside `cmdMilestoneComplete`
(which tells users to "Re-run with --force to override") could never be bypassed.

Add `const force = args.includes('--force')` and pass it into the options object.
The guard already honors `options.force` — no changes to milestone.cts needed.

Closes #978

* chore(#978): backfill changeset pr number (982)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 10:40:06 -04:00