* fix(#4721): give cleanup-wave's merge its own timeout, report merge_timed_out, and restore the index a killed merge leaves staged `worktree cleanup-wave` ran `git merge --no-ff` under the module-wide DEFAULT_GIT_TIMEOUT_MS (10 s) that is sized for plumbing calls. The merge is the one call in the wave that runs user hooks, so a repo whose pre-merge-commit hook is a test-suite gate lost every code-bearing executor merge. Three things went wrong at once, each fixed here: 1. Budget. The merge now passes an explicit timeout — DEFAULT_MERGE_TIMEOUT_MS (10 min), overridable via deps.mergeTimeoutMs. Every other git call in the wave keeps the module default; the shared constant is untouched, because every other caller is exactly what its 10 s comment describes. 2. Reason. A merge that does time out blocks on `merge_timed_out`, and its stderr names the budget and says the hook may still be running, instead of `merge_failed` carrying whatever the hook had printed before git was killed — which made a healthy executor branch look broken. 3. Residue. A merge killed during its hook has already staged the merged tree into the primary's index but never wrote MERGE_HEAD, so `git merge --abort` finds nothing and repoRootStillMidMerge (#2852) reads the primary as clean while the executor's whole diff sits staged against the old HEAD; a `git commit` from that state squashes the executor's history into one parent. After any failed merge the wave now reads `git diff --cached --name-only`; anything staged is the merge's own (git refuses to start a merge when the index differs from HEAD), so it runs `git reset --merge` — restores exactly those paths, keeps unrelated unstaged edits — and re-reads. Restored paths are reported as WAVE_CLEANUP_WARNING.MERGE_RESIDUE_RESTORED and the wave continues; a still-dirty or unreadable index reports MERGE_RESIDUE_LEFT_STAGED and halts the remaining entries, the same repo-level carve-out an unfinished merge takes. Tests: five mock-driven rows (budget wiring incl. the deps override, the timeout classification with restore, the no-reset control for an ordinary refused merge, an unrestorable residue halting the wave, an unverifiable index failing closed) plus a real-git row that runs a sleeping pre-merge-commit hook under a 1 s budget and asserts HEAD unmoved, index and worktree clean, the executor branch intact — with the same fixture merging cleanly under the default budget as its negative control. Two existing #2852 rows gain a handler for the new post-failure index read. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * docs(#4721): add Fixed changeset Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * test(#4721): release the real-git fixtures with t.after, not try/finally The two real-git rows cleaned up their scratch repo in a `finally` block; this file's own convention for fixture teardown is the test context's `t.after(() => cleanup(dir))`, and the house PR ruleset flags `finally` in a test body. Behaviour-neutral. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * fix(#4721): gate the residue restore on the timeout, re-apply a merge autostash, and correct the hook census Three findings from the pre-file adversarial review of the previous commit, each driven on real git before changing code: 1. A merge git REFUSED ("your local changes … would be overwritten") also leaves no MERGE_HEAD — and that refusal is exactly what a pre-existing dirty primary index earns. The residue restore read that index as the merge's own and `reset --merge`d the operator's staged work away (driven: a staged edit to an unrelated file was discarded and reported as "restored"). The restore now runs ONLY when the merge timed out; a refusal is an immediate exit, never a timeout, so on that path nothing is read or reset. 2. `merge.autoStash=true` lets a merge start on a dirty index by parking the work in MERGE_AUTOSTASH, which a killed merge never re-applies. `git reset --merge` moves that stash into the stash list; the wave now runs `git stash pop --index` afterwards (the outcome `merge --abort` gives an autostashed merge), and reports WAVE_CLEANUP_WARNING.MERGE_AUTOSTASH_UNRESTORED (path null) when the pop fails or the autostash state could not be read — the work stays in the stash, the index is clean, the wave continues. Because of this the reset runs on a timed-out merge even when the index reads clean. 3. The merge is not the only hook-running git call in the module: `worktree add` runs post-checkout and every ref update runs reference-transaction. It is the only call that runs the commit-family hooks, which is what the budget is for. Comments and docs say so now. Tests: the "ordinary merge_failed" control becomes the regression row for finding 1 (strict mock — a `diff --cached` or `reset --merge` on a refused merge throws), plus a mock row for the autostash pop (dirty and clean index, pop success and failure), and two real-git rows: a refused merge over pre-existing staged work leaves it byte-identical, and a killed merge under merge.autoStash restores the executor residue AND puts the operator's staged work back. The real-git hook now sleeps 4 s against a 1.5 s budget for margin on slow runners. The two #2852 handlers added earlier are removed — the residue read no longer fires on their path. Negative control: 4 of the 10 #4721 rows fail on the previous commit. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * fix(#4721): key the residue restore on a killed merge, and re-read the index after a failed autostash pop Two more findings from the continuation review, both driven: 1. An externally delivered SIGTERM leaves the same staged/no-MERGE_HEAD state as the timeout, and the seam reports it as exitCode null + signal with timedOut false — so the timeout-only gate skipped the restore on a state it was written for. The gate is now "killed": timedOut, or a null exit code with a signal. A refused merge still exits with a code and is still never touched. The reason stays merge_failed for a signal kill. 2. A failed `git stash pop --index` keeps the stash entry but can leave conflict entries (UU) and partially applied paths, after which the next merge fails on "you have unmerged files"; the code returned halt:false on the strength of the pre-pop recheck. The index is now re-read after a failed pop and a dirty result halts the wave as merge_residue_left_staged alongside the merge_autostash_unrestored warning. Also driven and now documented rather than changed: a kill that lands once MERGE_HEAD exists (inside commit-msg) is the ordinary #2852 abort path — `git merge --abort` restores the tree and re-applies an autostash itself, unstaged, as git does for any aborted autostashed merge. Tests: the pop-failure mock row now asserts the post-pop re-read and gains a conflict-leftover variant that halts; a signal-kill mock row; a real-git row with the sleeping hook moved to commit-msg (timed out, no residue warnings, MERGE_HEAD cleared, primary clean). 414 pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * fix(#4721): key the kill gate on the seam's signal, not on a null exit code The shell projection seam normalizes a signal death to exitCode 1 and carries the signal alongside (`_spawnResult`: `result.status ?? 1`), so the previous `exitCode === null && signal` gate could never fire in production and the unit row that covered it modelled a shape the seam does not emit (caught in the round-3 review). The gate is now `timedOut || signal`; a refused merge exits with a code and no signal. The mock row uses the real shape, and a mocked spawnSync signal death driven through the compiled seam reaches `reset --merge` and reports the residue restored. Also: three comments that still said "at its budget" / "runs user hooks" / "the index is clean", and the CLI-TOOLS sentence that reserved `merge_failed` for refusals and conflicts, now name the signal case. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TbbqrGJMuiuLftAMVLayb9 * chore(#4721): set changeset fragment pr to 4766 --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
GSD Core documentation
Documentation is organised into four quadrants: tutorials help you learn by doing, how-to guides solve specific tasks, reference states authoritative facts, and explanation explores concepts and design decisions.
Language versions: English · Português (pt-BR) · 日本語 · 简体中文
Tutorials
- Your first project — install to first shipped phase, one guaranteed path
- Onboarding an existing codebase — bring GSD Core to a brownfield repo
- Build your first capability — author a tiny declarative capability and watch it act in the loop
- Install your first capability — install a third-party capability end-to-end: consent, verify, check for updates, remove
How-to guides
- Install on your runtime — runtime-specific install steps for all 16 supported runtimes
- Install a minimal GSD and add skills later — install only the core skills, then grow the surface with profiles and
/gsd-surface - Attach a plugin-provided skill to a GSD agent — use the
global:plugin:skillentry form to load Claude Code plugin skills into agent prompts - Discuss a phase — capture implementation decisions before planning begins
- Resolve edge-coverage findings — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- Probe edges in a non-English project — get real edge coverage on a spec written in another language, and tell "no edges here" apart from "the probe could not read it"
- Resolve prohibition findings — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions
- Resolve an unreachable-workflow finding — wire or fully sweep a shipped workflow that no command, agent, or skill references
- Acknowledge emitted-artifact drift — declare a deliberate emitted-byte ripple or workflow/agent growth in a commit trailer, and migrate an older ack fragment
- Change the STATE.md schema — add, change or remove a STATE.md frontmatter key and keep the template and all five reference documents in step
- Resolve verify-command path findings — fix an
<automated>verify command whose target directory does not resolve from the executor's cwd - State a failing direction — say what output constitutes failure for an
<automated>verify command, and migrate a phase planned before the rule - Resolve a contract-drift finding — bring an agent's completion contract, read-tag gate, or deleted-file test reference back into agreement with the registry
- Resolve unreachable-guard findings — fix shell guards whose fallback arm cannot run, and tell "nothing to report" apart from "could not look"
- Declare a hook's crash policy — terminate a GSD hook with
allow/deny/crash, declare itsON_CRASHpolicy, and tell a hook's own crash apart from a check that could not run at all - Resolve a skipped capability probe — act on a coverage gate that held your phase for an unestablished scope, or a planning checkpoint that reported
skippedinstead of a verdict - Diagnose which gsd-tools is running — tell this package's tool apart from the predecessor's colliding binary and from a gsd-core too old to identify itself
- Resolve an ESLint glob-coverage finding — bring a source file that matches no lint rule under coverage, or record a reasoned exemption
- Resolve a raw-terminator finding — pick
runMain/ExitError,terminateNow, orprocess.exitCodefor alocal/require-registered-exitfinding, and know the two allowlist entries and the rule's documented evasions - Adopt the v2 exit contract — turn on
gsd-tools's versioned exit-code projection, read the code table including what80(DEGRADED) means, and migrate a CI gate that treats any non-zero exit as fatal - Read the statusline freshness marker — turn on
state ~N commits back, and tell "STATE.md is fresh" apart from "freshness could not be established" - Consume the planning snapshot — read
planning inspectfrom a dashboard or harness, and tell "nothing to report" apart from "could not look" - Read CI timeout budget signals — find the near-cap warning on a run, read the accumulated
tests/ci-timeout-budget-history.jsonltrend, and know which lever (cap, shard balance, shard-1 contents) a repeatedly-near-cap lane calls for - Consume the state contract — read
.planning/state.jsonfrom a workbench or editor extension, gate on the contract version, and tell "nothing to show" apart from "could not look" - Keep planning docs out of a shared repo — make
.planning/local-only, including untracking files git already tracks (the step.gitignorealone cannot do) - Publish PRs without planning artifacts — keep
.planning/committed locally, so worktrees and/gsd-undokeep working, whileplanning.pr_strictkeeps every planning path out of the branch you push - Plan a phase — run research, decompose work, and verify plan quality
- Verify a dependency-compatibility claim — act on a compatibility claim the researcher left
[ASSUMED], and tell "nothing declared" apart from "a constraint is declared" and "the lookup failed" - Execute a phase — run plans in parallel waves with fresh-context subagents
- Enable parallel reviewer lanes — cut a multi-reviewer
/gsd-reviewpass toward its slowest lane, and tell a rate-limited lane apart from one that was never selected - Enable concurrent per-plan planners in chunked mode — dispatch chunked
/gsd-plan-phase's per-plan Tasks together within one outline Wave instead of one at a time, and know when the setting has no effect - Verify and ship — walk through completed work, diagnose failures, and create the PR
- Catch complexity before it compounds — enable the post-execute refactor hook, read a proposal's score vs. anchor delta, and accept or decline it
- Run phases autonomously — use autonomous mode for unattended phase execution
- Handle quick and fast tasks — use
/gsd-quickand/gsd-fastfor ad-hoc work outside the phase loop - Batch quick tasks — run several
/gsd-quick-shaped tasks together with/gsd-quick-batch, understand capacity/isolation, and recover a failed or interrupted batch - Configure model profiles — switch between quality, balanced, and budget model tiers
- Control which host runtime GSD reports — read the
agent_runtimeladder, understand what host detection looks at, and pin the runtime when detection is not what you want - Set up cross-AI review — configure a second AI to review code produced by the primary agent
- Scope code review depth by path — escalate
/gsd-code-reviewtodeepfor sensitive directories while the rest of the repo stays at the default depth - Work in parallel with workstreams — run independent lines of work simultaneously using workstreams
- Isolate work with workspaces — use workspaces to sandbox experimental or risky changes
- Debug a failed execution — diagnose and recover from broken or incomplete phase execution
- Interpret scope-conformance warnings — read the advisory the worktree-wave merge emits when a plan branch commits outside its declared scope
- Interpret install-shadow warnings — read the advisory GSD Core emits when a
/gsd-*trigger is installed at both scopes and one silently wins, and tell "nothing to report" apart from "could not look" - Interpret
state validateresults — read thescopereason codes and tell "nothing to report" apart from "could not look" - Spike and sketch — use
/gsd-spikeand/gsd-sketchfor exploratory work before committing to a plan - Design a UI phase — use the UI phase loop for frontend and visual work
- Enable live-DOM verification — opt a project into browser-backed UI acceptance checks during execution, handle the browser-profile lock, and tell "nothing to report" apart from "could not look"
- Develop a Capability for GSD 1.5+ — add feature Capabilities, hook fragments, and registry entries
- Develop a task-content resolver capability — declare a
taskContentResolversoexecute-plan.mdresolves per-task content from your external issue tracker instead ofPLAN.md - Ship a reviewer lane in your capability — declare a
reviewerbody so/gsd-reviewdiscovers, invokes, and renders your external review CLI or model endpoint - List your reviewer lane in the registry — publish a lane you have built to the Reviewer Lane Registry so other people can find and install it
- Take over a capability or EoS integration — assume maintainership of an existing third-party capability, reviewer lane, or EoS host integration through a handoff, an adoption fork, first-party absorption, or a de-listing
- Add or update a host's integration — set a host's documentation-sourced
runtime.hostIntegrationaxes (ADR-1239 Phase A), with theundocumentedsentinel rule - Migrate an install test to the executed plan — convert an
fs.existsSync-probing install test group to a value assertion againstinstallRuntimeArtifacts's executed-plan return, and test against a fake fs adapter - Vendor a dependency — add a third-party package
gsd-core/bin/**needs at runtime as a verbatim vendored artifact, keep it out ofdependencies, and pick the right upstream bundle - Turn a capability off (and keep it off) — disable a capability via the surface, or gate individual hooks off without removing the capability
- Drive GSD from a tracker issue — start a phase from a GitHub, Linear, or Jira issue
- Migrate from GSD 2 — upgrade an existing GSD 2 project to GSD Core
- Update GSD — re-run the installer to pick up the latest release
- Clean up get-shit-done-cc — remove leftover old-package artifacts that cause a spurious
⬆ /gsd-updateindicator after migrating to@opengsd/gsd-core - Fix the worktree base-mismatch (exit 42) error — resolve the branch-divergence condition that halts parallel phase execution
- Recover and troubleshoot — fix common problems, rebuild context, and uninstall
Reference
- Commands — every command with flags and examples
- Configuration — full config schema, model profiles, git branching strategies
- CLI tools —
gsd-tools.cjsprogrammatic API for workflows and agents - JSON error mode —
gsd-toolsfailure channels: faults (stderr, exit 1) vs degraded results (stdout, exit 0), and the reason-code taxonomy - Features — complete feature index
- Inventory — installed skills and surface map
- STATE.md schema — field-by-field reference for
.planning/STATE.md - CONTEXT.md schema — field-by-field reference for
.planning/phases/<N>/CONTEXT.md - PLAN.md schema — field-by-field reference for
.planning/phases/<N>/PLAN.md - Planning artifacts — all
.planning/files and their roles - Review and verification capabilities — code review, security, and Nyquist capability ownership and hook contracts
- Gate predicates — canonical specification of the phase-gate predicate vocabulary
- Capability matrix — generated catalogue of every capability's role, tier, extension points, hook kinds, and
engines.gsd - Exit code reference — generated catalogue of every registered process exit code, its name, meaning, and owning module, plus the reserved bands and the v1/v2 exit contract
- Capability manifest — the full
capability.jsonschema and validation rules gsd capabilitycommand — install / update / remove / list reference for third-party capabilities- Workflow fragments — in-file
<!-- gsd:section -->marker grammar for fragmentizing workflow markdown at emission time - Partition rules for compact-content splits — the protected-content list, sentinel syntax, and the five CI checks a
workflow.compact_contentspine/detail split must obey - Reviewer Lane Registry — generated catalogue of third-party reviewer lanes, with their flags, transport, and install commands
Explanation
- Context engineering — how context rot forms and how GSD Core prevents it
- The phase loop — design rationale for the Discuss → Plan → Execute → Verify → Ship cycle
- Multi-agent orchestration — how subagents are spawned, scoped, and coordinated
- Security model — trust boundaries, permissions, and safe automation
- The capability trust model — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- How overlay capabilities compose — why first-party always wins and how the loader resolves precedence, conflicts, and fail-open load-failure warnings
- Architecture — system architecture, agent model, and data flow
- The Embeddable Orchestration System — one public, versioned contract for embedding GSD across many hosts
- Discuss modes — assumptions mode vs interview mode for
/gsd-discuss-phase - Context monitoring — context window monitoring hook architecture
- Issue-driven orchestration — recipe for driving GSD from a tracker issue using existing primitives
Related
- What's new in 1.7.0 — curated highlights of the 1.7.0 release
- Root README — landing page, quickstart, and documentation overview
- Changelog — release history