Commit Graph

376 Commits

Author SHA1 Message Date
Tom Boucher
05db5b53e4 fix(#853): gate manager/autonomous background dispatch by runtime (#863)
* fix(#853): gate manager/autonomous bg dispatch by runtime

/gsd-manager and /gsd-autonomous --interactive dispatched Plan/Execute
via Agent(run_in_background=true). On Claude Code a backgrounded agent
has no Agent/Task tool, so it cannot spawn the nested subagents those
pipelines need — per-plan worktree-isolated executors, the plan-checker,
and the verifier. The phases reported complete but isolation and
independent verification silently never ran, even with use_worktrees /
plan_check / verifier enabled.

Both workflows now resolve the runtime (config-get runtime, default
claude) before dispatching: run plan/execute INLINE on Claude Code so
the nested pipeline runs, and background-dispatch only on runtimes where
a backgrounded agent can still nest. Mirrors execute-phase.md's existing
Codex fail-closed precedent. Reconciles the stale unconditional
background/overlap/lean-context claims elsewhere in both workflows and
in the docs. Adds a content regression test pinning the gate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#853): add changeset for runtime-gated bg dispatch

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 3697e6768f)
2026-06-09 03:34:32 +00:00
Tom Boucher
1b6bd66f2c feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) (#821)
* feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged)

Wire three new context-tracking events (SubagentStop, Stop, PreCompact) to
gsd-context-monitor so context-headroom warnings surface at model-stop and
subagent-finalisation moments — not just on PostToolUse.  Add a new
FileChanged hook (gsd-config-reload.js) that hot-reloads .planning/config.json
context mid-session when the user edits it, injecting a config summary as
hookSpecificOutput.additionalContext.  Updates plugin manifest hooks.json,
managed-hooks-registry, installer-migration-report allowlist, and
shell-command-projection cleanup tables.  Tests: 21 new assertions in
enh-770-claude-hook-events.test.cjs; enh-788 and issue-766 test suites updated.

Closes #770

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#770): document newly-registered Claude Code lifecycle hooks

Add a Hook coverage table to the Claude Code npm installer section of
docs/how-to/install-on-your-runtime.md describing SubagentStop, Stop,
PreCompact, and the new FileChanged (gsd-config-reload.js) hook that
hot-reloads .planning/config.json mid-session. Also fixes the changeset
frontmatter (adds type: Added + pr: 821) so docs-lint can consume the
fragment.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): add gsd-config-reload.js to INVENTORY.md and regenerate manifest

The feat commit added hooks/gsd-config-reload.js but did not bump the
Hooks count in docs/INVENTORY.md (14→15) or add the new row, and did not
regenerate docs/INVENTORY-MANIFEST.json. Both inventory-counts and
inventory-manifest-sync tests failed across the full CI matrix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): make lifecycle-hook tests deterministic on scoped runner

Replace the shared hooks/dist/ ensemble setup (ensureHooksDist /
teardownHooksDist) in the Claude hook tests with per-test isolation:
pre-populate each test's own tmpDir/.claude/hooks/ with stub files and
pass installerMigrations:[] to install() so the first-time-baseline
migration does not remove the stubs before the copy step can run.

Root cause: hooks/dist/ is gitignored and absent on a fresh npm ci.
ensureHooksDist() created it and teardownHooksDist() deleted it, but
with --test-concurrency=4 both test files ran concurrently as separate
Node.js worker processes sharing the same filesystem.  One file's
afterEach teardown deleted hooks/dist/ while the other file's install()
was copying from it, producing an ENOENT (reproduced 2/10 runs locally).

The additional issue: even with pre-placed stubs surviving the copy race,
the 000-first-time-baseline migration classified hooks/gsd-*.js as
bundled-gsd-hook artifacts, auto-removed them, and the copy step never
re-ran (hooks/dist/ absent) — leaving contextMonitorFile missing and all
hook registrations silently skipped (the 'got: []' symptom).

Fix: pre-populate targetDir/hooks/ per-test (isolated temp dir) AND pass
installerMigrations:[] so the baseline scan is skipped.  The Qwen suites
already used this pattern correctly; the Claude suites are aligned to it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): ship gsd-config-reload.js by adding it to build-hooks HOOKS_TO_COPY

The #770 feature added hooks/gsd-config-reload.js and registered it in
MANAGED_HOOKS, the installer, INVENTORY, and the test EXPECTED_ALL_HOOKS
list — but never added it to scripts/build-hooks.js HOOKS_TO_COPY. As a
result the hook was never copied into hooks/dist/ during the build, so:

  - the hook would never ship to users (real production bug — the
    FileChanged config-reload feature was dead-on-arrival), and
  - install-minimal-hooks.test.cjs #1755 ("all expected hooks are copied
    from hooks/dist/ to target", ".js hooks are executable after copy",
    "manifest contains .js hook entries") failed on any environment with
    a clean checkout (no pre-existing hooks/dist/): coverage, full test
    macos-22/macos-24, test ubuntu-24.

The failures were masked locally only by a stale hooks/dist/ left from a
prior build (build-hooks copies into dist without clearing it). On CI's
fresh `npm ci` there is no dist, so the omission surfaced.

Fix: add 'gsd-config-reload.js' to HOOKS_TO_COPY so build-hooks stages it
into hooks/dist/ alongside the other JS hooks. Verified by removing
hooks/dist/ and rerunning the full suite green (0 fail).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): make config prototype-pollution beforeEach deterministic on scoped runner

Root cause: the #663 and alert-#26 prototype-pollution describe blocks
seeded .planning/config.json in beforeEach via a bare
runGsdTools('config-ensure-section') whose result was discarded. That
command runs in a spawned gsd-tools child; on the scoped CI lane
(--test-concurrency=4, config.test.cjs scheduled alongside the heavy
install/tarball suites that #770 pulled into the targeted set) the child
can be transiently killed under resource pressure (non-zero exit, empty
stderr — an OS-level kill, not an app error). The swallowed failure left
config.json absent, so the first subtest's readConfig() threw ENOENT
opening <tmp>/.planning/config.json. Only 1 of 4 subtests failed,
confirming a per-invocation transient, not a deterministic miss; the full
suite schedules files differently so config.test.cjs did not collide with
those heavy neighbors → passed there.

Fix: add ensureConfigReady(tmpDir) which retries config-ensure-section on
ANY failure or missing file and throws a clear diagnostic if it still
cannot create config.json, then use it in both prototype-pollution
beforeEach blocks. Setup is now deterministic under load; the #663/alert-#26
security assertions are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:36:11 -04:00
Tom Boucher
d32b8db635 feat(#836): no-LLM duplicate-issue detection + challenge + 1-day auto-close (#843)
* feat(#836): no-LLM duplicate-issue detection + challenge + 1-day auto-close

Adds a deterministic (no-LLM) duplicate-issue governance lifecycle:

- scripts/issue-dedupe.cjs: pure, unit-tested module (tokenize, Sørensen–Dice
  title similarity, scoreCandidates, renderChallengeComment, shouldClose) with
  fail-safe destructive-action guards.
- duplicate-check.yml (issues:opened): scores new-issue title against open
  issues, posts a challenge comment + applies the pending `possible-duplicate`
  label on a clear match.
- duplicate-sweep.yml (daily cron): closes possible-duplicate issues whose
  challenge comment is >24h old with no human reply and no 👎 veto; honors
  exempt labels; re-checks the label immediately before close (TOCTOU guard);
  strips the label on close to avoid reopen loops.
- remove-duplicate-label.yml (issue_comment:created): clears the label and
  applies needs-maintainer-review when any human responds.
- bug_report.yml / docs_issue.yml: add the required "I searched existing
  issues" preflight checkbox so all five forms force a pre-search attestation.
- docs/agents/triage-labels.md: document the label + lifecycle.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#836): add changeset fragment for duplicate-issue detection

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:35:43 -04:00
Tom Boucher
40d48c0508 feat(#815): add /gsd-update --next to install the @next RC channel (#839)
Adds an opt-in --next (alias --rc) flag to /gsd-update targeting the @next RC dist-tag (ADR #660), with a {latest,next} allowlist enforced at three layers, channel-aware version check + banner, and byte-for-byte unchanged default @latest behavior.

Closes #815
2026-06-07 20:22:07 -04:00
Tom Boucher
5b4880522b docs(#832): add how-to guide for minimal install / skill profiles (#835)
Add docs/how-to/install-minimal-and-add-skills.md covering the --minimal
/ --core-only / --profile=core install, the core/standard/full profiles,
and growing the surface live via /gsd:surface or on reinstall. Register
it in the docs/README.md How-to guides index.

Closes #832

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 19:36:28 -04:00
Tom Boucher
1040fb792e feat(#777): register Cursor-native hooks (.cursor/hooks.json) for session-start/post-tool parity (#831)
* feat(#777): register Cursor-native hooks (.cursor/hooks.json) for session-start/post-tool parity

- Add gsd-cursor-session-start.js: injects STATE.md presence reminder (or
  new-project nudge) into Cursor sessions via the sessionStart hook event
- Add gsd-cursor-post-tool.js: emits an additional_context nudge when
  write-class tool calls touch .planning/ files (postToolUse hook event)
- Add 'cursor-hooks-json' installSurface to runtime-config-adapter-registry;
  writeCursorHooksJson/reconcileCursorHooksJson write the canonical
  { version: 1, hooks: { sessionStart, postToolUse } } JSON shape with
  idempotent reconciliation that preserves user-owned hook entries
- Hook scripts are copied with /gsd:→gsd- rewrite so installed files
  contain no colon-form slash-command refs (bug-376 invariant)
- 20 new tests in tests/cursor-hooks.test.cjs cover all reconciler paths,
  entry helpers, removal, runtime adapter surface, and hook script behavior
- Update CONTEXT.md, ARCHITECTURE.md, installer-migrations.md, and
  000-first-time-baseline.cts to include Cursor hooks.json surface

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#777): build hooks/dist on demand in bug-376 test for scoped/windows CI

hooks/dist is gitignored and only produced by `npm run build:hooks`.
The CI scoped (ubuntu-latest/node-22) and windows (windows-latest/node-24)
test jobs do NOT run build:hooks before executing tests, so bug-376's
prerequisite suite was failing with "hooks/dist not found" on both legs.

Add ensureHooksDist() helper (mirrors bug-3357 pattern) that builds
hooks/dist on demand in the before() hooks of prerequisite and Suite 3.
Also add ensureHooksDist() call to Suite 3's before() so the snapshot
step is also hermetic.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 19:22:48 -04:00
Tom Boucher
e04e757672 feat(#776): Gemini hook events (BeforeAgent/AfterAgent/BeforeModel) + hooksConfig.enabled check (#829)
* feat(#776): Gemini hook events (BeforeAgent/AfterAgent/BeforeModel) + hooksConfig.enabled check

Register three new Gemini-CLI hook events on install:
  - BeforeAgent: fires before agent planning; wired to gsd-context-monitor
  - AfterAgent: fires after final response generation; wired to gsd-context-monitor
  - BeforeModel: fires before each LLM call (per-turn); wired to gsd-context-monitor

All three reuse gsd-context-monitor.js (no new hook files). Uninstall cleanup
loop extended to remove the new events. Non-array guard added for robustness
against malformed settings.

Also detect hooksConfig.enabled:false in Gemini settings and emit a clear
warning — without this check, all registered hooks silently do nothing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: update changeset pr: 829

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#776): document Gemini hook events

Add hook coverage table to the Gemini CLI section of install-on-your-runtime.md,
covering the three new events (BeforeAgent/AfterAgent/BeforeModel wired to
gsd-context-monitor) plus a callout for the hooksConfig.enabled:false silent
failure mode detected by the installer.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 19:16:32 -04:00
Tom Boucher
cb284962bf feat(#789): elevate CodeBuddy — slash commands (#830)
* feat(#789): elevate CodeBuddy — emit slash commands (+ document subagent/MCP scope)

Emit a CodeBuddy slash-command surface so GSD workflows appear in the
'/' menu, reaching parity with other elevated runtimes.

- Add convertClaudeCommandToCodebuddyCommand and register a commands/
  artifact kind for the codebuddy runtime (commands/gsd-<name>.md),
  consistent with the Cursor (#785) and Augment (#790) commands surfaces.
- Mark emitted skills user-invocable:false so the commands surface is the
  sole '/' entry point (no duplicate /gsd-* entries); skills stay
  model-invocable. CodeBuddy's SKILL.md supports this field.
- Normalize $HOME/.codebuddy (bare + slash) path forms in runtime
  rewrites so --config-dir/local installs don't leak the default home.
- Report installed commands/ count on install; uninstall prunes gsd-*
  commands while preserving user-owned commands.

Scope: subagents (~/.codebuddy/agents/) are already emitted by the
generic agents block (unchanged); no mcp.json is written (gsd ships no
MCP server, and CodeBuddy's mcp.json registers only external servers).

Closes #789

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#789): set changeset pr number to 830

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:47 -04:00
Tom Boucher
28c8d524fe feat(#778): cross-runtime command enrichment (Gemini {{args}}/!{}, Qwen priority) (#825)
* feat(#778): cross-runtime command enrichment (Gemini {{args}}/!{}, Qwen priority)

Enrich the installer's per-runtime command/skill generators with native,
verified, additive fields:

- Gemini CLI: map Claude's $ARGUMENTS -> Gemini's {{args}} in generated TOML
  commands so typed arguments interpolate; inject live .planning/STATE.md into
  /gsd:progress via a fixed, injection-safe !{cat .planning/STATE.md 2>/dev/null}
  shell block (no interpolated input).
- Qwen Code: emit the optional numeric `priority` field on main-loop skills so
  the most-used workflows sort first in the /skills list (higher = earlier per
  the Qwen skills spec; the issue's inverse numbering was corrected).

OpenCode per-command model/agent/subtask/variant enrichment was evaluated and
intentionally not implemented: `model` reintroduces the #1156
ProviderModelNotFoundError regression for non-Anthropic providers (the converter
deliberately strips model:), `subtask`/`agent` change execution semantics for
GSD's interactive commands, and `variant` is not in the OpenCode command schema.

Schemas verified against primary docs (Gemini custom-commands, Qwen skills,
OpenCode commands/skills). Adds tests/enh-778-* and how-to + USER-GUIDE docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#778): set changeset PR number to 825

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:42 -04:00
Tom Boucher
67703d1586 feat(#772): adopt stable Codex hook events + commandWindows for Windows parity (#827)
* feat(#772): adopt stable Codex hook events + commandWindows for Windows parity

Register three new stable Codex hook events (SubagentStart, Stop,
PostToolUse) wired to gsd-context-monitor.js so Codex installs get
the same context-headroom tracking at subagent and session boundaries
that Claude/Qwen already have.

Add commandWindows field to the SessionStart hook entry on Windows so
Codex uses the .cmd shim directly (Git Bash/MSYS cannot POSIX-exec
node.exe). commandWindows is only emitted on win32; POSIX is unchanged.

Refactor reconcileCodexHooksJsonSessionStart into a generic
reconcileCodexHooksJsonEvent so any event name can be reconciled with
the same dedup/preserve-user-entries logic.

Add gsd-context-monitor.js and .cmd to MANAGED_HOOK_COMMAND_BASENAMES
_BY_SURFACE so idempotent re-runs de-duplicate entries correctly.

30 new tests covering: export surface, event registration for each of
the three events, commandWindows parity (POSIX vs win32), idempotency,
uninstall, and user-entry preservation.

Closes #772

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#772): windows path normalization + docs-lint

- Normalize scriptPath backslashes to forward slashes in
  ensureCodexHooksJsonEvent and ensureCodexHooksJsonSessionStart so that
  isManagedHookCommand can match stored commands against configDir on
  Windows CI runners. path.resolve returns backslash paths on Windows,
  but when platform is not 'win32' (e.g. platform:'linux' in tests),
  projectManagedHookCommand skips normalization — producing a mismatch
  that breaks idempotency deduplication (the same hook entry appended
  twice on re-register). Forward-slash paths are always valid in both
  Node.js and Codex, so the normalization is safe for all platforms.
- Fix changeset pr: 0 → 827 to resolve fail_malformed_fragment.
- Add Codex hook coverage table to docs/how-to/install-on-your-runtime.md
  documenting the SubagentStart/Stop/PostToolUse events + commandWindows
  Windows-parity field added by this enhancement.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:36 -04:00
Tom Boucher
41f91b2e88 feat(#769): adopt context:fork + effort on heavy workflow skills (#820)
* feat(#769): emit context:fork + effort: frontmatter on heavy workflow skills

Add `context: fork` and `effort: xhigh` to the three heaviest workflow
commands (plan-phase, execute-phase, autonomous) and `effort: low` to the
two quick-status commands (progress, stats).

On Claude Code, `context: fork` runs the skill in an isolated subagent
context window so the main session's context budget is protected.
`effort: xhigh` / `effort: low` signal the appropriate token-budget tier to
the runtime. Both fields are silently ignored by runtimes that do not
recognise them (Gemini, Codex, Cursor, etc.) — no behaviour change outside
Claude Code.

Update convertClaudeCommandToClaudeSkill in bin/install.js to preserve
`context:` and `effort:` when rewriting source command files to SKILL.md for
a Claude global install. Add install-suite tests to assert the fields are
present in both source commands and the installed SKILL.md output.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#769): tighten regex assertions + add execute/plan-phase effort coverage

Fix low-severity adversarial finding: tighten test regex patterns from
`\s*` to `[ \t]*` so they cannot match across newlines (CRLF parity).
Add missing effort: xhigh assertions for gsd-execute-phase and gsd-plan-phase
SKILL.md install output to complete the black-box coverage gap.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:30 -04:00
Tom Boucher
19280510fe feat(#774): emit service_tier/model_verbosity in Codex agent TOML + agents/openai.yaml skill chip (#828)
* feat(#774): emit service_tier/model_verbosity in Codex agent TOML + agents/openai.yaml skill chip

- Add service_tier = "flex" and model_verbosity = "low" to the Codex
  ConfigProfile TOML for light-tier agents (gsd-research-synthesizer,
  gsd-codebase-mapper, gsd-plan-checker, and 8 others identified via
  AGENT_DEFAULT_TIERS). Field names/values verified against Codex schema
  (profile_toml.rs / config_types.rs Verbosity enum). Non-light agents
  are unaffected.

- Add generateCodexSkillMetadataYaml() and writeCodexSkillMetadataFiles():
  after installRuntimeArtifacts, iterate every gsd-* skill directory,
  read the short-description already emitted in the SKILL.md frontmatter
  by convertClaudeCommandToCodexSkill, and write agents/openai.yaml with
  interface.display_name and interface.short_description for the Codex
  TUI skill picker chip.
  - yamlQuote (JSON.stringify) handles all YAML-unsafe chars.
  - User-owned gsd-dev-preferences dir is never overwritten.
  - Errors per-skill are swallowed so a bad SKILL.md can't abort install.
  - agents/openai.yaml is covered by the snapshot/rollback system and
    manifest hash (writeManifest hashes skill dirs recursively).
  - Uninstall symmetry: _removeGsdEntries removes whole gsd-* dirs.

- 21 new tests in codex-config.test.cjs covering service_tier/verbosity
  TOML emission, YAML generation (round-trip via js-yaml), and
  writeCodexSkillMetadataFiles including an e2e integration test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#774): correct docs-lint coverage — proper changeset format + USER-GUIDE entry

Rewrite the changeset fragment from old @opengsd/gsd-core:patch format to the
required type:/pr: schema so the docs-lint parser can consume it.  Add a new
"Codex skill picker and agent scheduling (#774)" section to docs/USER-GUIDE.md
describing the flex-tier scheduling and /skills TUI chip enrichments — both are
user-visible and belong in docs rather than behind a docs-exempt marker.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:25 -04:00
Tom Boucher
dc7f1557c6 feat(#768): pre-populate settings.json permissions.allow/deny for Claude Code (#819)
* feat(#768): pre-populate settings.json permissions.allow/deny for Claude Code

Adds mergeClaudePermissions() to bin/install.js which non-destructively
appends GSD's known-safe tool-call patterns to permissions.allow and
defense-in-depth credential-file patterns to permissions.deny during
Claude Code installs. Merge is idempotent (no duplicates on reinstall)
and additive (existing user entries preserved). Uninstall removes only
the exact GSD-owned entries.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: update changeset pr number to 819

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:20 -04:00
Tom Boucher
a3aa0ae142 feat(#775): ship a gemini-extension.json extension package (#818)
Add a Gemini CLI extension package so users can install, update, and
remove GSD through Gemini's own extension lifecycle and have it appear in
`gemini extensions list`:

  gemini extensions install https://github.com/open-gsd/gsd-core
  gemini extensions update gsd-core
  gemini extensions uninstall gsd-core
  gemini extensions link /path/to/gsd-core   # dev

This mirrors the additive Claude Code plugin manifest (#766): a thin,
version-stamped manifest enforced by an in-repo drift test. The extension
ships the context-file payload (GEMINI.md), loaded into every Gemini
session; slash-command/agent/hook TOML projection into the extension is a
documented follow-up. The manual `npx gsd-core --gemini` installer (which
provides the /gsd:* commands) is unchanged — purely additive, no breaking
change.

- gemini-extension.json: name=binName, version tracks package.json,
  description, contextFileName=GEMINI.md (minimal; no mcpServers — gsd
  ships no MCP server)
- GEMINI.md: Gemini-session context payload
- package.json: add both artifacts to files[] so they publish
- CONTEXT.md: add "Gemini Extension Package" glossary entry
- docs: USER-GUIDE + install-on-your-runtime how-to
- tests/issue-775-gemini-extension.test.cjs: manifest validity, version
  parity with package.json, contextFileName existence, files[] publication

Closes #775

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:54:08 -04:00
Tom Boucher
3aed02822d chore(#771): convert agent color: hex/magenta values to documented named colors (#823)
* chore(#771): convert agent color: hex/magenta values to documented named colors

Claude Code's sub-agent `color:` field documents only 8 named colors
(red, blue, green, yellow, purple, orange, pink, cyan). Twelve agent
files used hex values and two used the undocumented `magenta`; convert
each to the nearest documented named color so the intended per-agent
TUI color differentiation is spec-compliant.

- agents/*.md: 14 color values hex/magenta -> nearest named color
- scripts/research-profiles.cjs: update the 3 generated research-agent
  profiles (source of truth) so gen-research-agents stays in sync
- docs/AGENTS.md: update documented colors; add missing Color rows for
  gsd-nyquist-auditor, gsd-project-researcher, gsd-phase-researcher
- tests/agent-frontmatter.test.cjs: add regression guard asserting every
  agent color: is in the documented named-color set

Closes #771

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#771): add changeset

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:50:50 -04:00
Tom Boucher
ea0d8f09b1 enh(#784): emit native skills for OpenCode + Kilo runtimes (#810)
* feat(#784): emit native skills for OpenCode + Kilo runtimes

OpenCode and Kilo share a config schema and both discover on-demand
skills from skills/<name>/SKILL.md. The installer previously emitted
only flat commands (command/) and file-based agents (agents/) for these
runtimes. Add a shared OpenCode-family skill writer that stages each GSD
command as a spec-compliant SKILL.md (name matching the directory,
description 1-1024 chars), wired through the runtime artifact layout so
uninstall cleans skills/ automatically. Skills respect the active
install profile.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#784): correct skill body paths + preserve user dev-preferences

Address adversarial-review findings:
- Add opencode/kilo cases to _applyRuntimeRewrites so staged SKILL.md
  bodies are re-pointed from the converter's hardcoded default config dir
  to the actual install target (fixes --local / --config-dir installs;
  commands/agents already did this by applying pathPrefix pre-conversion).
- Preserve user-owned skills/gsd-dev-preferences across reinstall in
  installOpencodeFamilySkills (snapshot+restore around the gsd-* prune),
  matching installRuntimeArtifacts.
- Export installOpencodeFamilySkills and add regression tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#784): guarantee command/skill body parity, fix kilo-alt double-rewrite

Follow-up adversarial-review found the post-conversion path rewrite could
double-rewrite custom Kilo dirs (kilo -> kilo-alt -> kilo-alt-alt) because
the kilo pathPrefix is a $HOME (non-absolute) superset of the hardcoded
default base. Restructure so OpenCode/Kilo skills mirror copyFlattenedCommands
exactly: stage raw commands, apply pathPrefix BEFORE conversion via a new
shared applyOpencodeFamilyPathPrefix() helper (now used by both the command
and skill writers), then convert. This guarantees byte-for-byte command/
skill body parity for global, --local, and --config-dir installs and removes
the prefix-overlap hazard. Drop the fragile _applyRuntimeRewrites opencode/
kilo case. Strengthen the path regression test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(#784): derive opencode/kilo skills from the same staged command set

Pass the installer's _stageSkills() output directly to
installOpencodeFamilySkills instead of re-staging via the layout, so the
command/ and skills/ surfaces always cover the identical profile-resolved
set — including the --minimal/--core-only alias path, which stages
differently from a plain --profile=core. Verified: minimal install now
emits 8 commands and 8 skills.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#784): set changeset PR number to 810

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#784): fully escape backslashes in test helper (CodeQL js/incomplete-string-escaping)

Replace the dot-only escape `replace(/[.]/g, '\\.')` with a complete
regex-escape pattern `replace(/[\\.*+?^${}()|[\]]/g, '\\$&')` so all
regex metacharacters (including backslash itself) in `defaultBase` are
safely escaped before interpolation into `new RegExp(...)`.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 16:13:06 -04:00
Tom Boucher
10087a48f5 feat(#790): emit Augment slash commands (~/.augment/commands/) (#808)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-07 15:52:22 -04:00
Tom Boucher
5e4e7de1ff enhancement(#782): emit gsd skills to ~/.cline/skills for Cline >= v3.48 (#809)
Cline added a global skills system (~/.cline/skills/<name>/SKILL.md) in
v3.48.0, but gsd treated Cline as rules-only and emitted zero skills
(getGlobalSkillsBase('cline')=null, empty artifact kinds). This makes gsd
emit skills for Cline at global scope, alongside the existing .clinerules.

- runtime-homes: getGlobalSkillsBase('cline') -> ~/.cline/skills (was null)
- runtime-artifact-layout: cline emits a skills kind for GLOBAL scope only
  (local stays .clinerules-only), mirroring claude's scope dispatch
- install.js: convertClaudeCommandToClineSkill emits name+description-only
  SKILL.md frontmatter (Cline/agentskills.io spec; no Claude-specific
  allowed-tools/argument-hint/agent), hyphen-normalized + .cline/-rewritten
  body; global cline routed through the skills path while .clinerules is
  still written; _applyRuntimeRewrites cline case handles custom
  CLINE_CONFIG_DIR; convertClaudeToCliineMarkdown also rewrites bare
  ~/.claude and CLAUDE_CONFIG_DIR
- docs: install-on-your-runtime.md documents Cline global skills vs local rules
- tests: converter (name+description-only), global emission, skills+.clinerules
  coexistence, scope-aware layout, custom-dir paths, idempotency

Closes #782

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 15:48:40 -04:00
Tom Boucher
5618505f9a feat(#788): expand Qwen Code hook-event coverage (#807)
* feat(#788): expand Qwen Code hook-event coverage to 4 new events

Register SubagentStop, Stop, PreCompact (gsd-context-monitor.js) and
UserPromptSubmit (gsd-prompt-guard.js) in the Qwen Code installer.
Guard is isQwen-only — Claude Code and all other runtimes are unchanged.
Uninstall loop extended to include the 4 new event names.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#788): reconcile to 3 Qwen-only events — defer UserPromptSubmit

gsd-prompt-guard exits unless tool_name is Write|Edit (PreToolUse
payload shape); UserPromptSubmit carries raw user-prompt text with no
tool_name field, so wiring it would be a silent no-op.  Deferred to a
follow-on issue.

Artifacts made consistent:
- bin/install.js: drop UserPromptSubmit registration block; uninstall
  loop drops UPS from event list
- .changeset/788-qwen-hook-events.md: corrected to 3 events + rationale
- docs/how-to/install-on-your-runtime.md: remove UPS row from hook table
- tests/enh-788-qwen-hook-events.test.cjs: assert UPS NOT registered;
  fix idempotency suite to persist settings between installs; drop
  UPS-specific assertions

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#788): update changeset PR number to #807

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#788): prune stale install-bucket allowlist entry for enh-788 test

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-07 15:34:32 -04:00
Tom Boucher
3025a6846e fix(#812): honor COPILOT_HOME in Copilot global config-dir resolution (#814)
* fix(#812): honor COPILOT_HOME in Copilot global config-dir resolution

getGlobalConfigDir('copilot') resolved the global config directory using
only --config-dir > COPILOT_CONFIG_DIR > ~/.copilot, ignoring the
COPILOT_HOME env var. Per GitHub's Copilot CLI docs, COPILOT_HOME
overrides the default ~/.copilot location (and user-level hooks are read
from $COPILOT_HOME/hooks/), so a global --copilot install wrote all
artifacts (skills, agents, copilot-instructions.md, the gsd-session.json
hook) to ~/.copilot even when the user relocated their Copilot home,
making them undiscoverable by Copilot CLI.

Mirror the codex/CODEX_HOME branch: precedence is now
--config-dir > COPILOT_CONFIG_DIR > COPILOT_HOME > ~/.copilot. Uninstall
uses the same resolver, so it stays symmetric.

Also: document COPILOT_HOME in the installer --help notes, the
USER-GUIDE env-var table, and the installer-migrations Copilot row; and
clear COPILOT_HOME in the two default-path test suites so they stay
hermetic now that the resolver honors it.

Closes #812

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#812): add changeset for PR #814

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 15:27:46 -04:00
Tom Boucher
eefef2ec19 feat(#787): elevate Cline — .clinerules/ dir form, PreToolUse hook, AGENTS.md (#803)
* feat(#787): elevate Cline — .clinerules/ dir form, PreToolUse hook, AGENTS.md

Migrate the installer's Cline output from a single-file .clinerules to the
.clinerules/ directory form (.clinerules/gsd.md), which is the prerequisite for
Cline's v3.36 hooks (a path cannot be both a file and a directory). Add a
.clinerules/hooks/PreToolUse lifecycle hook implementing Cline's JSON stdin ->
{cancel,errorMessage,contextModification} protocol; it guards .planning/
artifacts and fails open. On global installs, merge GSD instructions into the
cross-tool ~/.agents/AGENTS.md target (marker-delimited, merge-safe). A legacy
single-file .clinerules is migrated in place; --uninstall removes the new
artifacts and strips the AGENTS.md GSD block.

Also fixes the uninstall targetDir for Cline local installs (it pointed at
./.cline instead of the project root) and re-runs writeManifest after the
Cline artifacts are written so they are hash-tracked.

Self-contained: implemented independently of the #782 Cline skills work.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#787): address review findings

- Scope PreToolUse hook path-walk to PATH_KEY fields only (eliminates false
  positive when doc body content mentions .planning/)
- Use lstatSync + isSymbolicLink() for migration guard so GSD never writes
  through a user's symlinked .clinerules into an external directory
- Add regression tests for both cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#787): set changeset pr: 803

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 15:20:55 -04:00
Tom Boucher
74a818308e feat(#785): write .cursor/commands/ Cursor 1.6 slash-command surface (#805)
* feat(#785): write .cursor/commands/ as Cursor 1.6 slash-command surface

Cursor 1.6 (released 2025-09-12) introduced plain-markdown slash commands
in `.cursor/commands/<name>.md` — no frontmatter, invocable via `/` in the
Agent input. GSD previously emitted only `~/.cursor/skills/` for Cursor.

This PR wires a second artifact kind for `cursor` in
`runtime-artifact-layout.cts`: `convertedCommandsKind('commands', 'gsd-',
'convertClaudeCommandToCursorCommand', configDir)`. The new kind applies the
same `convertClaudeToCursorMarkdown` transforms (tool renames, brand
substitution, slash-command normalisation) and then strips YAML frontmatter
so the output is plain prose. Skills output is unchanged.

`stageCommandsForRuntimeFlat` in `install-profiles.cts` stages each source
`.md` as a flat `<stem>.md` in a temp dir; the existing `_copyStaged` commands
path then prefixes and copies to `<configDir>/commands/`.

`.cursor/mcp.json` is explicitly OUT OF SCOPE: GSD ships no MCP server; the
`mcpServers` schema cannot be usefully populated by the installer.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#785): address review nit

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-07 15:07:14 -04:00
Tom Boucher
5e852611e3 feat(#786): elevate GitHub Copilot installer — lifecycle hook + AGENTS.md (#804)
* feat(#786): elevate Copilot installer with lifecycle hook + AGENTS.md

Emit a self-contained sessionStart hook config (.github/hooks/gsd-session.json
local, ~/.copilot/hooks/gsd-session.json global) and write AGENTS.md at the repo
root (Copilot CLI reads it as primary instructions) alongside
copilot-instructions.md. The hook is an inline `command` hook (no separate hook
script), so it cannot dangle. Uninstall removes both and preserves user content.

Verified against GitHub Copilot CLI primary docs: hooks-configuration (camelCase
events, version+hooks shape, inline bash/powershell command hooks) and
add-custom-instructions (AGENTS.md read at repo root as primary instructions).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#786): set changeset pr number to 804

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 14:53:32 -04:00
Tom Boucher
f66c4a082c feat(#766): distribute gsd-core as a native Claude Code plugin (#797)
* feat(#766): distribute gsd-core as a native Claude Code plugin

Add an additive .claude-plugin/plugin.json manifest plus hooks/hooks.json so gsd-core can be installed as a first-class Claude Code plugin (marketplace or zero-friction @skills-dir), with /gsd-core: namespaced commands and lifecycle management — alongside the unchanged npm/file-copy installer.

- .claude-plugin/plugin.json: validated with 'claude plugin validate --strict'
- hooks/hooks.json: mirrors the installer's always-on Claude hook wiring via ${CLAUDE_PLUGIN_ROOT}
- package.json: ship .claude-plugin in the npm tarball
- tests/issue-766-plugin-manifest.test.cjs: manifest + always-on-hook-contract drift guards
- docs: install-on-your-runtime.md + FEATURES.md

Closes #766

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#766): add ADR-766 + glossary entry for Claude Code Plugin Manifest Module

Record the plugin manifest as the Seam projecting gsd-core's artifact surfaces onto the Claude Code plugin contract (sibling of the Runtime Artifact Layout Module, ADR-3660), with the defined kind->field mapping and the always-on hook projection rule.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 12:02:39 -04:00
Tom Boucher
2f07443119 refactor(#60): make runtime config adapter registry explicit (#795)
* refactor(#60): make runtime config adapter registry explicit

Replace scattered inline `runtime === '...'` config-mutation branching in
bin/install.js with an explicit, typed adapter registry. The new
src/runtime-config-adapter-registry.cts maps each of the 15 supported
runtimes to a config intent { installSurface, writesSharedSettings,
finishPermissionWriter }; install()/finishInstall() dispatch by resolved
intent instead of runtime-name checks (cursor/windsurf/trae collapse to one
profile-marker-only branch).

Behavior-preserving: the same config files are written for the same runtimes
(opencode still writes both settings.json and its permissions; kilo writes
only its permissions; codex minimal-mode and opencode GSD_TEST_MODE guards
unchanged). Unknown runtimes fail loudly via TypeError, with an Object.hasOwn
barrier so prototype-chain keys (__proto__/constructor) also throw rather than
returning a bogus intent. Leads the installer-refactor chain (#58 -> #60 ->
#56), building on ADR-58.

Closes #60

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#60): add changeset for runtime config adapter registry

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#60): register Runtime Config Adapter Registry in CONTEXT.md glossary

Per docs/contributor-standards.md, every new Module/seam must get a
`### <Name>` entry under the domain glossary. Adds the entry for the
runtime-config-adapter-registry seam introduced in this PR (interface,
policy boundary, source file, ADR-58 / #60 cross-references).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 11:16:03 -04:00
Tom Boucher
5237fae537 feat(#759): non-destructive CHANGELOG preview in the rc release job (#763)
The rc action publishes a release candidate to @next for testing but
never surfaces the curated CHANGELOG section for the version under test —
render only runs destructively at finalize (#715), so there was no safe
way to preview the upcoming notes during the RC window.

Add a --preview mode to scripts/changeset/cli.cjs cmdRender: it renders
the dated release section to stdout via the existing renderChangelog/
serializeChangelog path (with priorChangelog: null, so only the new
section is emitted), reuses the shared injectEmptyPlaceholder helper for
zero-fragment releases, and returns WITHOUT writing CHANGELOG.md or
deleting any .changeset fragment. Wire a "Preview CHANGELOG" step into
the rc job that renders to a file (standalone command, so a malformed
fragment fails the step) and cats it to the job summary and log.

Closes #759

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 10:05:21 -04:00
Tom Boucher
3011b937ad docs(#58): add ADR for Runtime Install Policy Module boundary (#762)
Record the Runtime Install Policy Module decision and ownership
boundary: install policy projects a pure, typed install plan by
composing artifact placements (ADR-3660) and command text (ADR-0009)
plus per-runtime config intentions, with no filesystem IO; runtime
adapters consume the plan and execute concrete file mutations and
format-specific config rendering. Explicitly records what stays
outside the policy module (TOML/JSON/Markdown serialization, merge
semantics, filesystem effects).

Adds the ADR index row in docs/adr/README.md and a glossary entry in
CONTEXT.md. Leads the installer-refactor chain (#58 -> #60 -> #56).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 09:55:25 -04:00
Tom Boucher
4056d830bc refactor(#651): consolidate verification-status routing into one queryable seam (#755)
* refactor(#651): consolidate verification-status routing into one queryable seam

The passed/gaps_found/human_needed verification status was re-encoded as
bare strings across three prose surfaces (gsd-verifier emits, execute-phase
routes, ship gates), each independently deciding the per-status next action
with no parity coupling — the DEFECT.GENERATIVE-FIX class.

Give the enum one home: src/verification.cts (-> bin/lib/verification.cjs)
exposing `gsd_run query verification.status <phaseDir>` returning a typed
{status, next_action, next_command}. ship.md and execute-phase.md now consume
the query instead of re-deriving the routing in prose; gsd-verifier.md points
at the shared vocabulary as the single emitter (values unchanged).

Also fixes the latent broad-grep status misread (DEFECT.FRONTMATTER-SCALAR-
BROAD-GREP): execute-phase.md read `grep "^status:"` over the whole report, so
a body `status:` line could misroute a valid phase. Extraction is now
frontmatter-scoped in one place. A parity test fails if a verifier status
gains no route. Lands the two CONTEXT.md DEFECT entries captured on the issue.

Closes #651

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#651): set changeset pr to 755

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 01:26:56 -04:00
Tom Boucher
f7e902f1cf feat(#52): add agent_skills_security.trusted_global_roots allowlist for global skills (#754)
* feat(#52): add agent_skills_security.trusted_global_roots allowlist

Opt-in allowlist so a global: agent skill whose SKILL.md realpath resolves
outside the default global skills base (e.g. ~/.claude/skills) is accepted
when its real target lies under a user-declared trusted root. Default [] is
byte-identical to prior behavior; the symlink-escape guard is preserved and
simply re-applied against each declared root.

- src/security.cts: loadTrustedGlobalRoots — tilde-expand (~ and ~/), reject
  project-relative and dangerously broad roots (filesystem/UNC root, homedir),
  realpath-canonicalize each root every run and drop non-existent ones.
- src/init.cts: on base-check failure the guard consults the trusted roots
  (hoisted out of the loop); emits a stderr NOTE when a skill is accepted via
  a trusted root so the widened boundary is visible.
- src/core.cts: thread agent_skills_security through loadConfig.
- config-schema.manifest.json: allow the new key path.
- docs/CONFIGURATION.md: document the option and its security model.
- tests/agent-skills.test.cjs: unit + end-to-end CLI coverage (regression,
  feature, negative, broad-root hardening, stderr NOTE).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#52): add changeset fragment for trusted_global_roots (#754)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 01:06:41 -04:00
Tom Boucher
ad1203f9d4 feat(#25): scope gsd-verifier Step 7b to enumerate-or-single-test; forbid full-suite re-runs (#753)
* feat(#25): scope gsd-verifier Step 7b to enumerate-or-single-test; forbid full-suite re-runs

Step 7b's lone test example (`npm test -- --grep "$PHASE_TEST_PATTERN"`) is
mocha/vitest/jest-specific, where `--grep` filters which tests *execute*. Models
generalized it to `cargo test --workspace 2>&1 | grep X` (runs the whole suite,
filters only *output*) and repeated it once per must-have, adding minutes per
verification with no new evidence after the first run.

Replace the example with language-agnostic guidance: prove a test EXISTS via
enumeration (`cargo test -- --list` / `pytest --collect-only` / `npx vitest
list` / `go test -list`), and prove it PASSES via a single named test
(`cargo test <name> -- --exact` / `pytest -k` / `npx vitest run -t`). Add a
Spot-check constraint forbidding more than one full-suite run per verification
or piping a full run through grep per must-have, while still permitting one
saved run + grep when a full run is genuinely required. docs/AGENTS.md gains a
one-line Key-behaviors note, and a new test asserts the Step 7b content.

Scoped per the maintainer decision on the issue: folded into Step 7b (no new
top-level Step 7a) with no VERIFICATION.md label changes.

Closes #25

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#25): add Changed changeset fragment for PR #753

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 00:46:27 -04:00
Tom Boucher
7e76f1a736 feat(#703): add --granularity override flag to /gsd:plan-phase (#750)
* feat(#703): add --granularity override flag to /gsd:plan-phase

Add a `--granularity <coarse|standard|fine>` flag to /gsd:plan-phase that
overrides the configured planning granularity for a single invocation.

The override is a new highest-priority tier above the existing precedence
chain (granularities[phaseType] -> granularity -> planning.granularity ->
'standard') in resolveGranularityInternal; when the flag is absent, resolution
is byte-for-byte unchanged. cmdInitPlanPhase now resolves with phaseType
'planning' so granularities.planning participates, and emits the resolved
value in the init JSON, which the plan-phase workflow forwards to the planner
prompt. Invalid values are rejected at the CLI boundary via a shared
assertValidGranularityOverride helper.

Closes #703

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#703): set changeset pr to 750

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 23:56:53 -04:00
Tom Boucher
cf8bd3cd5e fix(#683): auto-degrade phase execution to sequential on worktree base mismatch (#749)
* fix(#683): auto-degrade phase execution to sequential on worktree base mismatch

Claude Code forks worktree-isolated executors off the repository default
branch (origin/HEAD), not the orchestrator's HEAD. Running /gsd-execute-phase
on a branch diverged from the default (unmerged milestone/feature branch) left
every executor without the phase's plan files and tripped the
worktree-branch-check guard with `exit 42` — 100% reproducible, all OSes.

- New module src/worktree-base-ref.cts: HEAD-vs-fork-base drift detection
  (origin/HEAD with symbolic-ref fallback) and no-clobber worktree.baseRef
  management, exposed as `worktree base-check` / `worktree set-baseref`.
- execute-phase.md: pre-dispatch, for Claude Code with worktrees enabled,
  auto-degrades the run to sequential on the main tree when a base mismatch
  is detected, recommending worktree.baseRef:"head". The exit-42 guard stays
  as a backstop.
- Installer: fresh local Claude installs set worktree.baseRef:"head" in
  .claude/settings.local.json (no-clobber, respecting an explicit shared
  settings.json value); upgrades print an opt-in notice pointing at
  `gsd-tools worktree set-baseref`.
- Docs: how-to guide, CLI/config reference, planning-config cross-ref.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): auto-apply worktree.baseRef on upgrade; gate fresh+upgrade on use_worktrees

Per maintainer direction: on a local Claude Code UPGRADE, set
worktree.baseRef:"head" automatically (no opt-in notice) when the project's
workflow.use_worktrees is enabled, instead of merely printing a remediation
notice. For consistency the FRESH path is now gated the same way: both paths
compute worktrees-enabled once (bounded walk-up read of .planning/config.json,
default enabled unless workflow.use_worktrees === false) and apply the
no-clobber baseRef only when enabled — never overwriting an explicit value in
settings.local.json or a shared settings.json. gsd-tools worktree set-baseref
remains for manual use. Docs + changeset updated; tests hardened (file-exists
assertions, fresh+disabled case, upgrade idempotency).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): measure workflow byte-budget on LF, fixing Windows-only CI failure

The workflow-size-budget test failed only on Windows: git checks out the .md
files as CRLF (no eol=lf in .gitattributes) and byteCount used
fs.statSync().size (raw on-disk bytes), counting an extra \r per line. That
inflated execute-phase.md — the XL high-water-mark file pinned near its ceiling
by the tighten-only ratchet — from 88492 LF bytes to ~90245 on Windows, over
the 90000 XL ceiling, while passing on the LF-checkout Mac/Linux runners.

The ceilings are explicitly "calibrated against raw `wc -c`" on an LF checkout,
so the measurement should be LF-based on every platform. byteCount now reads the
file and counts Buffer.byteLength after stripping CR, making the budget
platform-independent (a no-op on LF checkouts; verified statSync === normalized
for all 88 workflow files). No ceilings changed. Added a regression test
asserting CRLF and LF content of the same file count identically.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#683): make worktree-base-ref test path mocks Windows-safe (path.join)

tests/worktree-base-ref.test.cjs keyed its injected readFile/writeFile mocks
(and a few expected `file` values) with forward-slash template literals like
`${claudeDir}/settings.local.json`. The module composes those paths with
path.join(), which emits backslashes on Windows, so the mock keys never matched
the module's lookup → readFile returned null → resolveEffectiveBaseRef /
cmdWorktreeBaseCheck / cmdWorktreeSetBaseRef (and the JSONC variants) failed on
the Windows full-test runner only (they passed on Mac/Linux, and the install
tests passed because they use the real filesystem). The module is correct;
only the test fixtures hardcoded '/'.

All mock keys and path assertions now use path.join(base, ...) mirroring the
module, so they match on every platform (no-op on POSIX). 19 path references
across 16 lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 23:40:24 -04:00
Tom Boucher
1bea220d58 refactor(#720): lazy-load MVP-only reference bodies on non-MVP runs (#746)
* refactor(#720): lazy-load MVP-only reference bodies (eager @-import → gated Read)

Convert eager @-imports of MVP-only reference bodies into lazy "Read" instructions
gated on MVP_MODE / WALKING_SKELETON / MVP+TDD, so non-MVP planning/execution runs
no longer pull MVP guidance into context. Covers both the workflow files and the
planner/executor agent definitions (the dominant context-cost path):

- workflows/plan-phase.md: planner-mvp-mode.md + skeleton-template.md (L146/936/937/941)
- workflows/execute-phase.md: execute-mvp-tdd.md halt-report ref, now gated on gate-trip (L191)
- agents/gsd-planner.md: planner-mvp-mode.md, user-story-template.md, skeleton-template.md
- agents/gsd-executor.md: execute-mvp-tdd.md

The dedicated always-MVP mvp-phase workflow keeps its eager imports (intentional).
Behaviour is unchanged; non-MVP runs simply carry less loaded context. Adds a
regression guard mirroring the discuss-phase lazy-load test, and documents the
conformance in docs/ARCHITECTURE.md.

Refs #720

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#720): add changeset fragment (pr #746)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 18:23:17 -04:00
Tom Boucher
b79b767629 refactor(bin): replace process.exit() in CLI entrypoints + ratchet rule to error (#738) (#741)
Part 2 of 2 of the n/no-process-exit cleanup (completes umbrella #738; part 1
was #739/scripts). Converts the 20 flagged process.exit() calls in the three
hand-written gsd-core/bin CLI entrypoints and flips n/no-process-exit to error.

- New src/cli-exit.cts -> gsd-core/bin/lib/cli-exit.cjs (ExitError + runMain),
  the gsd-core-side equivalent of scripts/lib/cli-exit.cjs; registered in
  .gitignore, eslint ignores, and the inventory manifest like its siblings.
- gsd-tools.cjs: 13 apply-prompt-budget exits -> throw ExitError; main()->runMain.
- verify-reapply-patches.cjs: 6 exits -> throw ExitError / return verdict; runMain.
- check-latest-version.cjs: 1 exit -> return verdict; runMain.
- eslint.config.mjs: n/no-process-exit warn -> error.

Scope note: the gsd-core/bin/lib/*.cjs modules (core, state, profile-pipeline,
roadmap-command-router, adr-parser, ui-safety-gate) are tsc-generated and
eslint-ignored (ADR-457), so their process.exit calls were never flagged and are
intentionally left untouched. Only the linted hand-written entrypoints are in scope.

Exit codes verified unchanged for all three entrypoints.

Closes #738

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 16:27:05 -04:00
Tom Boucher
42b74100f1 feat(#159): auto-use existing RESEARCH.md in /gsd:plan-phase --research-phase (#718)
* feat(#159): auto-use existing RESEARCH.md in /gsd:plan-phase --research-phase

When RESEARCH.md already exists in research-only mode and neither --research
nor --view is passed, emit a one-line notice and exit cleanly instead of
prompting update/view/skip. This matches the promptless auto-use of standard
/gsd:plan-phase <N> (§5.1) and removes the §5.0/§5.1 inconsistency, making
AI-agent and CLI invocations non-interactive in the common case. The two
explicit-flag escape hatches (--research to refresh, --view to print) cover
any deviation.

Closes #159

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#159): point changeset fragment at PR #718

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#159): tighten research-phase reference register (Diataxis)

Make the 'no modifier' research-phase entries descriptive rather than
imperative and drop the trailing 'pass --research/--view' clauses, which
duplicated the adjacent --research/--view documentation. Reference docs
describe; the recovery flags are documented in their own entries. The
emitted runtime notice in the workflow keeps naming the flags (in-band
recovery), unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 11:09:54 -04:00
Tom Boucher
31edeb1b54 feat(#717): re-base workflow size budget on bytes + document quality rationale (#719)
* feat(#717): re-base workflow size budget on bytes + document quality rationale

Re-base tests/workflow-size-budget.test.cjs from line counts to byte
counts (matches Codex's 32,768-byte project_doc_max_bytes cap; deterministic,
no tokenizer). Tier ceilings: XL=90000, LARGE=54000, DEFAULT=38000, GRACE=3000;
discuss-phase target re-expressed as <30 KB. The #597 tighten-only ratchet and
per-file budget semantics are preserved unchanged — only the unit swaps.

byteCount() uses fs.statSync().size to match `wc -c` (includes trailing
newline), deliberately not lineCount()'s newline-stripping.

Document the context-rot / attention-budget QUALITY rationale (independent of
prompt caching) in the test JSDoc and docs/ARCHITECTURE.md, plus the
Goodhart caveat: the byte budget measures one file, so the real goal is
bounded *loaded* context — eager @-imports game the proxy; legitimate
extraction is lazy. Update CONTEXT.md RULESET.WORKFLOW_SIZE_BUDGET to bytes and
remove a stale duplicate ruleset entry that still said "1800 lines".

Defers the #3182 MVP-mode split (tracked separately): MVP is a cross-cutting
concern woven through plan-phase/execute-phase, not a discrete extractable mode.

Closes #717

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#717): add changeset fragment for byte-budget re-base

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 20:19:08 -04:00
Tom Boucher
11afca2968 feat(#656): Research module — content-addressed cache + provider seam + registry-API legitimacy (#664)
* feat(#656): add Research Store module (content-addressed cache, TTL staleness)

Content-addressed research cache behind a clock seam: researchKey (sha256, deterministic), putResearch/getResearch ({hit,stale}, never throws), ttlForSource (curated HIGH 30d / MED 7d / web LOW 1d), two-tier resolveStorePath (curated -> ~/.gsd/research-cache, web/synthesis -> project .planning/research/.cache). 28 behavioral + property tests; boundary coverage at ttl-1/ttl/ttl+1.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): add Research Provider module (waterfall + confidence + plan)

Single source of truth for the Balanced provider waterfall (docs Context7->Ref->Jina, web Exa+Tavily, fallback Perplexity/Brave, Firecrawl scrape-only). classifyConfidence stamps HIGH|MEDIUM|LOW by provider (never throws). providerAvailability maps config flags to usable providers. planResearch checks the Research Store (injected seam) and returns cache-hits + a per-question fetch plan, falling through the waterfall to the always-available websearch terminal. 22 behavioral + property tests.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): add Package Legitimacy module (registry-API verdicts, slopcheck optional)

Replaces the pip-install-or-degrade slopcheck prose gate with code: classifyPackage (pure, never throws) computes OK|SUS|SLOP from tunable thresholds (minAgeDays 30, minWeeklyDownloads 1000, requireRepo). checkPackages queries injectable npm/PyPI/crates registry adapters (real https with 5s timeout, degraded-not-thrown on failure); slopcheck is one optional adapter that can only escalate severity, never degrade to [ASSUMED]. 34 behavioral + property tests; boundary coverage on age and downloads (limit-1/limit/limit+1).

Known follow-up: real npm adapter must add api.npmjs.org last-week downloads fetch (currently null -> unknown-downloads). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): detect Tavily/Ref/Perplexity/Jina provider keys; complete npm downloads adapter

config: add tavily_search/ref_search/perplexity/jina availability flags (env var or ~/.gsd/<x>_api_key), mirroring brave_search/exa_search/firecrawl, so the Research Provider waterfall can gate them. package-legitimacy: real npm adapter now fetches api.npmjs.org last-week downloads (bounded, degraded-not-thrown) so weeklyDownloads is populated. +12 config tests; 34 legitimacy tests unchanged.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#656): expose Research seam via gsd-tools query (research-plan, research-store, package-legitimacy)

Routes the L2-hybrid surface so agents reach it as CLI: 'query research-store get/put' (cache, HOME-sandboxable), 'query research-plan --input' (cache-hits + fetch plan from planResearch), 'query package-legitimacy check --ecosystem' (async registry verdicts). Commands skip .planning root resolution and appear in top-level usage. 5 behavioral runGsdTools tests; command-contract unchanged (335).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#656): document Research module (CONTEXT predicates, ADR-0656, architecture, changeset)

Adds GSD-RESEARCH.* + DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT predicates to CONTEXT.md, ADR-0656 recording the L2-hybrid seam decision, a docs/ARCHITECTURE.md Research Module subsection, and an Added changeset fragment (pr:0, backfill on PR). Notes the #657 deferrals (agent collapse + install.js MCP mapping).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): sync inventory for research modules

Regenerate INVENTORY-MANIFEST.json and bump docs/INVENTORY.md CLI Modules count 82->85 with rows for research-store/research-provider/package-legitimacy (DEFECT.INVENTORY-DRIFT).

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): eslint-ignore generated research .cjs artifacts (ADR-457)

research-store/research-provider/package-legitimacy .cjs are tsc-generated from src/*.cts, so they belong in the ESLint ignore block (lint the .cts source, not the emitted .cjs). Fixes tests/551-eslint-bin-lib-coverage.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): backfill changeset pr number to #664

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#656): satisfy eslint lint-tests gate

Fix 20 eslint errors in the new research files: use helpers.cleanup() instead of raw fs.rmSync() in tests (local/no-raw-rmsync-in-tests, Windows-EBUSY retry budget); drop redundant '| string' union members and unnecessary type assertions; deterministic object normalization in researchKey (no-base-to-string). Logic unchanged; 6180 tests still green.

Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): harden package legitimacy per review (W1/W2/I3/I4)

W1: httpsGet now reads statusCode; npm/PyPI/crates map 404 -> exists:false -> SLOP (registry-existence is the #1 slopsquatting defense; previously only npm caught it). Transport made injectable (_setHttpGet) for hermetic 404 tests. W2: suspicious-postinstall is now terminal SLOP independent of the optional slopcheck adapter, and the regex drops the bare https?:// arm (over-fired on esbuild/sharp/node-gyp) for shell-exec/download-exec signatures only. I3: checkPackages now threads version to registry.lookup and adapters verify that specific version exists. I4: moreServerVerdict -> moreSevereVerdict. +11 regression tests (all RED-first); 45 total green.

Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): research-store tier coherence + freshness + version TTL (W4/I1/I2/I4)

I1: tier now derives from source (curated -> user ~/.gsd, else -> project .planning), not kind, so put-tier and get-tier can't diverge; kind is a key component only. W4: getResearch searches both tiers and returns the freshest (non-stale preferred), never letting a stale curated entry shadow a fresh web one; blank version caps TTL at 1 day (no 30d on version-blind keys). I2: atomic platformWriteSync instead of raw fs.writeFileSync on the shared global path. I4: dropped the dead ttlForSource arm. CLI get now searches both tiers. +5 RED-first regression tests; 38 green.

Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): expose classifyConfidence as a CLI route, killing dead code (W3)

Adds 'gsd-tools query classify-confidence --provider X [--verified]' so research agents get the confidence tier FROM CODE (provider waterfall + verification lever) instead of asserting it in prose. classifyConfidence previously had no runtime caller. HIGH means 'trusted provider'; --verified raises web results to MEDIUM (verification semantics documented in ADR-0656). +4 behavioral tests.

Addresses review by @davesienkowski on #664 (W3). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): close Codex adversarial-review findings (path-traversal, version-age, malformed-cache)

HIGH: research key must be 64-hex sha256 (isValidResearchKey) + resolved-path containment check in put/get + CLI validation -> blocks '../../x' arbitrary-file-write. HIGH: package legitimacy now derives publishedAt from the REQUESTED version (npm time[version], PyPI releases[version] upload_time, crates versions[].created_at) so a new malicious version of an old package can't inherit old age and evade 'too-new'. MEDIUM: getResearch validates entry shape (finite fetched_at + positive ttl + required fields) -> malformed cache entry is a miss, not fresh-forever. +regression tests (RED-first); 111 green.

Codex adversarial review (required pre-PR gate). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): close code-review correctness findings

(1) package-legitimacy CLI now rejects unknown --flags instead of silently consuming the following package as a flag value; only --ecosystem takes a value. (2) crates recent_downloads (90-day) normalized to a weekly figure before the minWeeklyDownloads threshold (was ~13x too lenient). (3) research-plan --input validates parsed JSON is an object with an Array questions before destructuring -> clean usage error instead of an uncaught TypeError on null/bad input. (4) research-store put rejects a flag value that is itself a --flag (no more storing '--source' as content). (5) planResearch skips questions whose text is not a non-empty string instead of emitting question:undefined. +13 RED-first regression tests; 143 green.

Code-review gate. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(#657): extract researcher documentation_lookup to shared @-reference

6 researcher agents carried a near-duplicate <documentation_lookup> block; consolidate into gsd-core/references/research-documentation-lookup.md (@-included). Unifies the ctx7 CLI fallback to the safer 'command -v ctx7' guard (drops silent 'npx --yes ctx7@latest' execution in 5 agents). Behavior-preserving dedup; inventory 63->64 references. Phase A of the agent collapse.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(#657): extract researcher philosophy + verification-protocol to shared @-references

philosophy and the pitfalls+pre-submission-checklist common-core were near-duplicated in project/phase researchers; consolidate into gsd-core/references/research-{philosophy,verification-protocol}.md (@-included). phase-researcher keeps its 3 extra checklist items inline. Pre-submission domains checklist made agent-agnostic so project-researcher doesn't lose features/architecture coverage. Write-contract intentionally left inline (bug-214 tests assert it verbatim). Inventory 64->66 refs. Behavior-preserving. Phase A.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): wire gsd-phase-researcher to the Research seam (Phase B / S1)

The phase researcher now CALLS the code seam instead of carrying inline mechanics: provider waterfall -> 'gsd-tools query research-plan' (+ research-store put to cache digests); confidence-tier prose -> 'gsd-tools query classify-confidence'; slopcheck pip-install protocol -> 'gsd-tools query package-legitimacy check'. This makes the Research module a real runtime consumer (validates the seam end-to-end, addresses reviewer S1) and removes the duplicated waterfall/confidence/slopcheck prose. RESEARCH.md output contract, commit step, structured returns, and Phase-A @-includes unchanged. package-legitimacy-gate.test.cjs rewritten prose-grep -> behavioral (asserts the seam invocation).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): wire gsd-project-researcher to the seam + add tavily/ref/jina MCP tools (Phase C.1)

project-researcher now calls gsd-tools query research-plan / classify-confidence (+ research-store put) instead of the inline provider waterfall + confidence-tier prose (mirrors the phase-researcher rewire; no package-legitimacy — phase-only). Output contract (STACK/FEATURES/ARCHITECTURE/PITFALLS/SUMMARY.md + sections, no-commit, structured returns, Phase-A @-includes) unchanged. Adds mcp__tavily/ref/jina__* to the project/phase/ui researcher tools frontmatter (Balanced provider set) so install.js MCP mapping (C.2) has a consumer.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#657): cover tavily/ref/jina MCP install handling + frontmatter parity guard (Phase C.2)

Investigation: exa/firecrawl have no explicit per-runtime tool-mapping — every mcp__<server>__* except context7 rides the generic passthrough (Copilot lowercases; OpenCode/Cursor/Windsurf/Augment keep as-is; Gemini auto-discovers). tavily/ref/jina are handled identically, no install path broken. Added 12 copilot-install passthrough tests + a mcp-tool-inheritance parity guard (tavily co-declared with exa, jina with firecrawl, ref present across the 3 web researchers) so the MCP set can't drift. No io.github registry ids invented (none sourceable in-repo); documented as a follow-up. 488 tests green.

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#657): profiles as source of truth for researcher agents + drift-guard (Phase C.3)

scripts/research-profiles.cjs declares each of the 7 researcher agents' identity + contract (name, description, color, tools, required @-includes, required gsd-tools seam calls, output-contract markers). scripts/gen-research-agents.cjs --check validates every committed agent against its profile; --write regenerates ONLY the frontmatter from profiles (body untouched) and is a verified no-op against the current agents (zero diff = fidelity). tests/research-agent-profiles.test.cjs is the DEFECT.GENERATIVE-FIX drift guard. Design note: profiles govern the generatable/contract surface rather than destructively regenerating the disparate operational prose bodies (those were deduped via @-includes in Phase A). scripts/ is not inventoried (no inventory change).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#657): complete agent provider-dispatch + parity guard; align legitimacy field; validate profiles

Adversarial-review findings: (HIGH) the seam-wired agents' Step-C dispatch only mapped 6 providers, so a planResearch result of jina/ref/perplexity/brave (reachable via the waterfall fallbacks) had no handling -> agent stall; completed both agents' dispatch to all 9 PROVIDER_WATERFALL ids + a catch-all, and added a parity test asserting agent dispatch stays in sync with research-provider PROVIDER_WATERFALL (DEFECT.GENERATIVE-FIX). (MEDIUM) phase-researcher package-legitimacy JSON example used 'package' but the module returns 'name' -> aligned. (LOW) gen-research-agents checkAgent now returns a clear failure for a malformed profile instead of throwing. +parity/validation tests (RED-first).

Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#656): make classifyConfidence verification-evidence-driven (W3)

Confidence conflated provider authority with claim verification — context7/ref
stamped HIGH purely by provider identity, and the only verification lever was a
self-set --verified flag. Split into two axes: provider authority (static) +
verification evidence (code-computed). HIGH now requires ground-truth
corroboration (legitimacyVerdict OK), independent of provider; authority alone
caps at MEDIUM; SLOP caps at LOW; the self-reported --verified is demoted to a
MEDIUM-only web lever. HIGH = corroborated-against-authoritative-source, not a
correctness guarantee. Adds --legitimacy-verdict to the classify-confidence CLI;
updates CONTEXT.md predicate + ADR-0656 (tier set unchanged, ADR-consistent).

Addresses davesienkowski's W3 review on #664.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#656): bind classify-confidence verdict to code, closing CLI self-grading

Adversarial review found the new --legitimacy-verdict flag was caller-supplied,
so an agent could self-assert OK->HIGH without any real legitimacy check —
reintroducing the exact self-grading hole W3 closes. Remove the free flag; the
CLI now computes the verdict via checkPackages only when --package/--ecosystem
is given (code-computed, not agent-asserted). Update the stale CLI test
(context7 alone -> MEDIUM) and extend the property test to vary legitimacyVerdict.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 17:58:48 -04:00
Tom Boucher
b44e16a47a docs: document --only and --text flags for /gsd-autonomous (#695) (#716)
Adds --only N and --text to the COMMANDS.md reference table and the
run-phases-autonomously how-to guide, revised to fit the Diataxis
framework (reference: factual/parallel rows; how-to: goal-framed sections).
2026-06-05 17:43:19 -04:00
Tom Boucher
0e259a589c fix(#705): route hardcoded $HOME gsd-tools invocations in agents/commands through gsd_run (#707)
* fix(#705): route hardcoded $HOME gsd-tools invocations in agents/commands through gsd_run

The hardcoded `node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" <cmd>` form
(fixed for workflows in #621/#637) survived in agent/command surfaces and
misresolves on global/shim-only installs. Route every agent-executed
invocation through the resolved `gsd_run` launcher in gsd-phase-researcher,
gsd-planner (load_graph_context extracted to a shared reference to stay under
the planner size budget), import, and graphify. Add a regression guard over
agents/ + commands/ + gsd-core/references/ bash blocks. User-facing display
messages and docs are intentionally left untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#705): use repo changeset fragment format (type: Fixed, pr: 707)

The hand-written fragment used the standard changesets package format
(package: bump) which lacks the type:/pr: frontmatter the repo's
docs-required lint consumes (fail_malformed_fragment / missing_type).
Regenerated via scripts/changeset/new.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 09:26:29 -04:00
Tom Boucher
e3571b2e1b fix(#676): update tests/docs referencing deleted hotfix.yml (#680)
hotfix.yml was deleted (folded into release.yml). Remove the now-broken
release-coverage-scope and policy-release-no-npm-self-upgrade assertions that
readFileSync'd hotfix.yml (release.yml equivalents retained), drop the dead
install-smoke.yml path trigger, and update VERSIONING.md / docs/branching.md
prose to describe hotfixes via the Release workflow with a patch version.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 14:57:18 -04:00
Tom Boucher
cdd78bd2aa fix(#670): self-healing recovery for installer-migration checksum drift (#675)
Editing the body of an already-released installer migration drifts its computed
checksum (it hashes plan.toString()). The integrity guard then hard-aborted
every prior install on upgrade with "applied migration checksum changed" — a
100% reproducible blocker (v1.3.0, all platforms).

Already-applied migrations are filtered out of `pending` and never re-run, so
a drifted checksum is functionally inert. ADR-0008 anticipates checksum-mismatch
state as something the install-state layer must handle gracefully (plan -> apply
-> recover/report), not abort on.

This supersedes the published-checksum allowlist merged in #674 (per-release
maintenance debt — every historical checksum hand-pinned, still throws for any
unregistered value) with a general, self-healing recovery:

- Replace the throwing guard with non-fatal `collectAppliedChecksumDrift`,
  surfaced on `plan.checksumDrift`.
- Reconcile drifted stored checksums durably on the next state write
  (`reconcileDriftedChecksums`), idempotently (no perpetual writes).
- Relocate the "shipped migration bodies are immutable" rule to a CI baseline
  test that locks every shipped migration's checksum and fails on body drift —
  where #615 should have been caught, instead of blocking users.

Removes #674's legacyChecksums field, per-migration checksum pins,
published-checksums.json fixture, and compat test.

Fixes #670

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 13:44:05 -04:00
Colin
b8e15c7a98 fix: accept published installer migration checksums 2026-06-04 12:30:13 -04:00
Tom Boucher
0afed31904 docs(#660): add ADR for release-from-next-head release model (#661)
Replaces the persistent/frozen release branch + hand-moved tag with:
release always cut from next's head, immutable tags minted once at
finalize, next on a -dev stream, and @next dist-tag as the RC surface.

Closes #660

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 23:26:18 -04:00
Tom Boucher
14d0238cf5 docs: roll up legacy release notes into a single archive (#658) (#659)
Consolidate all pre-rename release notes (the retired get-shit-done-cc /
get-shit-done-redux lineage, 1.0.0 -> 1.50.0-canary.1) into one condensed,
read-only archive so the legacy 1.x version numbers no longer collide with
the current @opengsd/gsd-core line.

- Add docs/RELEASE-NOTES-LEGACY.md: rename banner, master version-index
  table, condensed per-version sections (1.42.3 -> 1.0.0), and a separate
  pre-release & canary builds section. Stale install commands stripped.
- Trim CHANGELOG.md to the current @opengsd/gsd-core line only; replace the
  Legacy Release History block with a pointer to the archive and drop the
  orphaned numbered legacy reference-link definitions.
- Remove the 10 standalone docs/RELEASE-v*.md files.
- Repoint docs/CANARY.md and docs/FEATURES.md links to the new archive.

Closes #658

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 23:04:44 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
91f5bd7cb1 feat(#607): rebuild get-shit-done-cc → gsd-core migration (per-package cache + installer auto-cleanup + --dry-run) (#611)
* feat(#607): rebuild get-shit-done-cc → gsd-core migration

Leftover get-shit-done-cc installs poisoned the shared update cache,
causing a permanent false "update available". Rebuild the migration so a
stale old install is both harmless and actively removed.

- Per-package update cache filename (gsd-update-check-<slug>.json) in the
  shared ~/.cache/gsd dir, single-sourced via package-identity; writers
  stamp package_name and readers reject foreign/absent lineage. Multi-
  runtime visibility preserved (same shared dir + filename across runtimes).
- New get-shit-done/bin/lib/legacy-cleanup.cjs seam: detects code-file
  references to the old package + the legacy fixed-name cache across home
  runtime dirs; installer auto-cleans on every install; --dry-run previews
  and mutates nothing. User hooks and dev-preferences are never touched.
- update.md cache-clear globs gsd-update-check*.json across ALL supported
  runtimes (adds cursor/windsurf/augment/trae/qwen/hermes/codebuddy/cline).
- Fix worker MODULE_NOT_FOUND post-install (ship managed-hooks-registry.cjs
  + degrade gracefully) so the per-package cache is always written.
- Diataxis how-to: docs/cleanup-get-shit-done-cc.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#607): add changeset for PR #611

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#607): set USERPROFILE alongside HOME in dry-run install test for Windows

os.homedir() reads USERPROFILE on win32, so HOME-only isolation let the
spawned installer scan the real runner home on windows-latest. Set both.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 12:19:57 -04:00
Tom Boucher
3bb2f8f1c5 docs: rebrand to GSD Core and restructure docs with Diataxis (#605)
* chore: wire docs/agents config into AGENTS.md Agent skills section

Add the `## Agent skills` discovery block pointing the engineering
skills at the existing docs/agents/{issue-tracker,triage-labels,domain}.md
files (issue tracker, triage label mapping, single-context domain docs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: rebrand to GSD Core and restructure docs with Diataxis

Reorganise the root README and docs/ around the Diataxis framework
(tutorials, how-to guides, reference, explanation), add new how-to
guides and schema references (STATE.md / CONTEXT.md / PLAN.md /
planning artifacts), and cross-link the whole set. Update the lone
legacy gsd-build reference to open-gsd; keep internal get-shit-done/
filesystem paths unchanged (directory rename tracked separately in
open-gsd/gsd-core#604). Regenerate the ja-JP, ko-KR, pt-BR and zh-CN
localised trees to mirror the new structure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: backfill changeset PR number (#605)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 08:13:09 -04:00
Tom Boucher
a11ba2dfcb feat(#68): per-phase granularity overrides (granularities.<phaseType>) (#595)
Closes #68. Per-phase-type granularity overrides via granularities.<phaseType>, mirroring models.<phaseType>. Includes maintainer-authorized sdk-seam reference cleanup.
2026-06-01 20:53:46 -04:00
Tom Boucher
9ffe45a7c3 feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation (#591)
* feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation

Tighten the Granularity Calibration buckets in gsd-roadmapper (Coarse 3-5->2-4,
Standard 5-8->4-6, Fine 8-12->6-10) and append inline Key guidance naming the
thin-phase failure pattern (single requirement / internal-quality goal /
task-shaped success criteria) with instruction to fold into a neighbor rather
than create a standalone phase. Implements the maintainer-approved proposal
verbatim.

Update the canonical English docs that hardcoded the old phase-count numbers:
docs/CONFIGURATION.md and docs/FEATURES.md. Translated docs are
community-maintained and are not updated per-PR (CONTRIBUTING.md language
policy).

Prompt/doc text only; no code, format, or downstream-consumer changes. Agent
size-budget and skills-awareness tests pass; full suite green.

Closes #163

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#163): add Changed changeset for roadmapper granularity tightening

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#163): lock tightened gsd-roadmapper granularity buckets

source-text-is-the-product test asserting the Granularity Calibration table
holds the tightened ranges (Coarse 2-4, Standard 4-6, Fine 6-10), that no row
maps to an old bucket, and that the Key paragraph carries the thin-phase
folding guidance. Would fail if the values regress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 20:37:49 -04:00
Tom Boucher
0c1084ba88 fix(#48): verify-only worktree_branch_check + orchestrator cwd-drift guard (#590)
Closes #48. Makes the canonical worktree_branch_check fragment verify-only/fail-closed (exit 42, no git reset self-recovery), adds an orchestrator fail-closed collection rule and a cwd-drift guard at execute_waves entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 20:10:59 -04:00