* fix(#4881): trust worktree.baseRef:"head" in harness mode and keep the #4868 observation as what restores it under a WorktreeCreate hook The pre-dispatch base check still derived its harness-mode verdict from the retired #48 premise that the harness never reads worktree.baseRef: with "head" set and HEAD diverged from origin/HEAD it degraded every wave with baseref-head-ignored-by-harness. #4868 inserted an observation of a clean prior harness worktree at HEAD ahead of that comparison, but an execute-phase run never has one at the moment it checks — the base-check runs before any dispatch, a degraded wave creates no worktrees, and a wave that did run in worktrees has them removed and HEAD moved before the next check — so the common case was unchanged (#4881 repro states 1 and 4). Re-scope of the closed #4752 onto current next, with #4868 kept: - branch a trusts "head" in both isolation modes, the way the harness is measured to behave (#4588: three settings layers, three OSes), and the spawn-time exit-42 guard stays the observation-based backstop; - a Claude Code WorktreeCreate hook in any settings file the check reads, or a file that does not parse, withholds that trust — the hook creates the worktree without applying the setting — and the inferred comparison runs, degrading with baseref-head-bypassed-by-hook; - the #4868 observation (b2) now sits behind branch a: it is reached only when "head" was not trusted outright, and on a hook host it is what restores the trust — a hook that forks from HEAD leaves exactly that evidence, one that forks elsewhere never does. Its per-HEAD cache is unchanged. It is skipped under an explicit --observed-fork-base, which outranks an inference from a prior worktree; - --observed-fork-base <sha> threads a measured fork base through the evaluation (strict full-hex, TypeError otherwise). Tests: the #4868 rows are unchanged and still reachable (they run with the setting unset); the #4752 rows re-land, with the one exit-128 row updated to the degrade #4734 pinned since; a new #4881 block pins the start-of-run state (baseref-head, git never consulted), the hook + observation composition in both directions, and that an explicit observation skips the probe. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS * docs(#4881): rewrite the eight surfaces that still state the harness ignores worktree.baseRef:"head" Every prose surface #4868 left untouched still asserted the retired #48 premise as verified fact, starting with the step file the orchestrator reads. Each now describes the measured behaviour, the WorktreeCreate-hook exception, the --observed-fork-base input, and the #4868 observation as what lifts the hook degrade; docs/CLI-TOOLS.md gains the fork-from-head-observed reason row #4868 did not document. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS * chore(#4881): set changeset fragment pr to 4921 * fix(#4881): withhold the #4868 observation under the hook interlock The hook interlock this PR added withheld the worktree.baseRef:"head" trust but still let b2's prior-worktree observation restore it, and that observation cannot be attributed to the hook. The evidence is a clean agent worktree sitting at the orchestrator HEAD; nothing on disk records which creator left it there, so one the plain harness created BEFORE a WorktreeCreate hook was configured — with HEAD unmoved since — reads as evidence for the hook. It was the one fail-open branch in a mechanism documented as fail-closed. Keying the observation cache by hook configuration does not close it. observeHarnessForkFromHead has two legs: a HEAD-keyed cache and a live probe over .claude/worktrees/agent-*. A hook-keyed cache simply misses, and the miss falls through to the probe, which re-finds the same stale worktree and re-confirms. The probe takes no hook input at all. So the observation is not consulted under the interlock rather than re-keyed: on a hook host the only admissible positive signal is an explicit --observed-fork-base measurement of the dispatch in hand, and absent one the inferred comparison runs and a mismatch degrades with baseref-head-bypassed-by-hook, leaving the spawn-time exit-42 guard as the backstop. Scoped to the case branch a. declined to trust: "head" set AND a hook (or an unparseable layer) in the harness's path. With no "head" setting b2 is #4868's own arm and is unchanged, hook or not — re-scoping that trust is a separate question this PR does not open, and a test pins the boundary. Cost, stated: a hook host with "head" set, on a branch diverged from origin/HEAD and passing no observation, now runs sequentially. It still runs parallel when HEAD matches origin/HEAD. No workflow threads an observation today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ * refactor(#4881): drop the now-unused forkRef message-builder parameter buildMsgBaserefHeadIgnored stopped reading forkRef when the message was made mode-neutral, and the parameter was retained with `void forkRef;` for symmetry with its two sibling builders, which do read it. Symmetry is not reason enough to keep a dead parameter on a module-private function with one caller, so drop it (#4921 review). Behaviour is unchanged; the message text is pinned by an existing full-string assertion, which is what covers the only real risk here — transposing the two remaining arguments at the call site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ * test(#4881): exercise both FULL_SHA_RE alternatives at their boundaries The invalid-observation list pinned 39 and 41 hex around the 40-hex SHA-1 arm but left the 64-hex SHA-256 arm's own +/-1 boundary unexercised, which the repo's boundary-coverage convention asks for (#4921 review). Adds 63, 65 and a 64-length non-hex string. The regex already rejected all three -- this is coverage of correct behaviour, not a fix -- so it carries no negative control against a pre-fix base. That it is not vacuous was shown instead by widening the arm to {63,65}, under which the row fails by name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ * docs(#4881): finish the surface sweep the hook interlock owes Self-found by this round's own pre-push adversarial review, over six passes. Eight sites, four classes. FIVE were prose still asserting the stance the interlock overturned, which reads as live canon to anyone arriving cold: docs/CLI-TOOLS.md, docs/CONFIGURATION.md and gsd-core/references/planning-config.md each still said the hook degrade is lifted "unless/until a clean prior harness worktree is observed"; observeHarnessForkFromHead's own header still said a qualifying worktree "can only exist if the harness forked from HEAD"; and a test name still called the observation "required on a hook host" when it is now inadmissible there. My own sweep had grepped for "restores"/"lifts" and missed every one -- the ordinary failure of a grep, which returns what you thought to search for. The SIXTH is the same class one step worse: gsd-core/workflows/execute-plan.md still said flatly that Claude Code's isolation="worktree" "forks from origin/HEAD, not live local HEAD" -- in a paragraph THIS PR already edits, a few sentences after the clause it corrected. A tombstone makes only its own line clean; adjoining text asserting the dead stance is the other half of the same defect. Now qualified on the setting, with the hook exception named. The SEVENTH is a proof-strength overstatement that predates this PR, with a driven counterexample: a worktree created from an older base and since `git checkout --detach`ed onto HEAD is clean, sits at HEAD, and satisfies the probe identically, so "can only exist" was false. The worktree's own reflog does retain that original checkout -- the information is not lost, the probe simply does not consult it. The EIGHTH is that the header described only one of the function's two legs. A cache hit returns the prior conclusion without reading any worktree, so "the probe reads a worktree's present state" was true of the live probe and false of the cache. The header now separates them, and names the cache's blindness as a third reason the observation is inadmissible under a hook. Gaps seven and eight are inherited from #4868 and accepted there for the no-hook case. Nothing about the mechanism changes here; only what the header claims for it. No behavioural change -- comments, prose, and one test's registered name. Emitted-Drift-Ack-Growth: execute-plan.md — the Pattern A paragraph gained a qualifying clause: it stated flatly that Claude Code forks from origin/HEAD, which is the premise this PR retires, a few sentences after the clause the PR had already corrected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
GSD Core documentation
Documentation is organised into four quadrants: tutorials help you learn by doing, how-to guides solve specific tasks, reference states authoritative facts, and explanation explores concepts and design decisions.
Language versions: English · Português (pt-BR) · 日本語 · 简体中文
Tutorials
- Your first project — install to first shipped phase, one guaranteed path
- Onboarding an existing codebase — bring GSD Core to a brownfield repo
- Build your first capability — author a tiny declarative capability and watch it act in the loop
- Install your first capability — install a third-party capability end-to-end: consent, verify, check for updates, remove
How-to guides
- Install on your runtime — runtime-specific install steps for all 16 supported runtimes
- Install a minimal GSD and add skills later — install only the core skills, then grow the surface with profiles and
/gsd-surface - Attach a plugin-provided skill to a GSD agent — use the
global:plugin:skillentry form to load Claude Code plugin skills into agent prompts - Discuss a phase — capture implementation decisions before planning begins
- Resolve edge-coverage findings — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- Probe edges in a non-English project — get real edge coverage on a spec written in another language, and tell "no edges here" apart from "the probe could not read it"
- Resolve prohibition findings — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions
- Resolve an unreachable-workflow finding — wire or fully sweep a shipped workflow that no command, agent, or skill references
- Acknowledge emitted-artifact drift — declare a deliberate emitted-byte ripple or workflow/agent growth in a commit trailer, and migrate an older ack fragment
- Change the STATE.md schema — add, change or remove a STATE.md frontmatter key and keep the template and all five reference documents in step
- Resolve verify-command path findings — fix an
<automated>verify command whose target directory does not resolve from the executor's cwd - State a failing direction — say what output constitutes failure for an
<automated>verify command, and migrate a phase planned before the rule - Resolve a contract-drift finding — bring an agent's completion contract, read-tag gate, or deleted-file test reference back into agreement with the registry
- Resolve unreachable-guard findings — fix shell guards whose fallback arm cannot run, and tell "nothing to report" apart from "could not look"
- Declare a hook's crash policy — terminate a GSD hook with
allow/deny/crash, declare itsON_CRASHpolicy, and tell a hook's own crash apart from a check that could not run at all - Resolve a skipped capability probe — act on a coverage gate that held your phase for an unestablished scope, or a planning checkpoint that reported
skippedinstead of a verdict - Diagnose which gsd-tools is running — tell this package's tool apart from the predecessor's colliding binary and from a gsd-core too old to identify itself
- Resolve an ESLint glob-coverage finding — bring a source file that matches no lint rule under coverage, or record a reasoned exemption
- Resolve a raw-terminator finding — pick
runMain/ExitError,terminateNow, orprocess.exitCodefor alocal/require-registered-exitfinding, and know the two allowlist entries and the rule's documented evasions - Adopt the v2 exit contract — turn on
gsd-tools's versioned exit-code projection, read the code table including what80(DEGRADED) means, and migrate a CI gate that treats any non-zero exit as fatal - Read the statusline freshness marker — turn on
state ~N commits back, and tell "STATE.md is fresh" apart from "freshness could not be established" - Consume the planning snapshot — read
planning inspectfrom a dashboard or harness, and tell "nothing to report" apart from "could not look" - Read CI timeout budget signals — find the near-cap warning on a run, read the accumulated
tests/ci-timeout-budget-history.jsonltrend, and know which lever (cap, shard balance, shard-1 contents) a repeatedly-near-cap lane calls for - Consume the state contract — read
.planning/state.jsonfrom a workbench or editor extension, gate on the contract version, and tell "nothing to show" apart from "could not look" - Keep planning docs out of a shared repo — make
.planning/local-only, including untracking files git already tracks (the step.gitignorealone cannot do) - Publish PRs without planning artifacts — keep
.planning/committed locally, so worktrees and/gsd-undokeep working, whileplanning.pr_strictkeeps every planning path out of the branch you push - Plan a phase — run research, decompose work, and verify plan quality
- Verify a dependency-compatibility claim — act on a compatibility claim the researcher left
[ASSUMED], and tell "nothing declared" apart from "a constraint is declared" and "the lookup failed" - Execute a phase — run plans in parallel waves with fresh-context subagents
- Enable parallel reviewer lanes — cut a multi-reviewer
/gsd-reviewpass toward its slowest lane, and tell a rate-limited lane apart from one that was never selected - Enable concurrent per-plan planners in chunked mode — dispatch chunked
/gsd-plan-phase's per-plan Tasks together within one outline Wave instead of one at a time, and know when the setting has no effect - Verify and ship — walk through completed work, diagnose failures, and create the PR
- Catch complexity before it compounds — enable the post-execute refactor hook, read a proposal's score vs. anchor delta, and accept or decline it
- Run phases autonomously — use autonomous mode for unattended phase execution
- Handle quick and fast tasks — use
/gsd-quickand/gsd-fastfor ad-hoc work outside the phase loop - Batch quick tasks — run several
/gsd-quick-shaped tasks together with/gsd-quick-batch, understand capacity/isolation, and recover a failed or interrupted batch - Configure model profiles — switch between quality, balanced, and budget model tiers
- Control which host runtime GSD reports — read the
agent_runtimeladder, understand what host detection looks at, and pin the runtime when detection is not what you want - Set up cross-AI review — configure a second AI to review code produced by the primary agent
- Scope code review depth by path — escalate
/gsd-code-reviewtodeepfor sensitive directories while the rest of the repo stays at the default depth - Work in parallel with workstreams — run independent lines of work simultaneously using workstreams
- Isolate work with workspaces — use workspaces to sandbox experimental or risky changes
- Debug a failed execution — diagnose and recover from broken or incomplete phase execution
- Interpret scope-conformance warnings — read the advisory the worktree-wave merge emits when a plan branch commits outside its declared scope
- Interpret install-shadow warnings — read the advisory GSD Core emits when a
/gsd-*trigger is installed at both scopes and one silently wins, and tell "nothing to report" apart from "could not look" - Interpret
state validateresults — read thescopereason codes and tell "nothing to report" apart from "could not look" - Spike and sketch — use
/gsd-spikeand/gsd-sketchfor exploratory work before committing to a plan - Design a UI phase — use the UI phase loop for frontend and visual work
- Enable live-DOM verification — opt a project into browser-backed UI acceptance checks during execution, handle the browser-profile lock, and tell "nothing to report" apart from "could not look"
- Enable UI interaction capture — let
/gsd-ui-review's auditor capture hover, focus, open-menu and filled-form states through thechrome-devtoolsCLI from Bash, with no MCP server - Develop a Capability for GSD 1.5+ — add feature Capabilities, hook fragments, and registry entries
- Develop a task-content resolver capability — declare a
taskContentResolversoexecute-plan.mdresolves per-task content from your external issue tracker instead ofPLAN.md - Ship a reviewer lane in your capability — declare a
reviewerbody so/gsd-reviewdiscovers, invokes, and renders your external review CLI or model endpoint - List your reviewer lane in the registry — publish a lane you have built to the Reviewer Lane Registry so other people can find and install it
- Take over a capability or EoS integration — assume maintainership of an existing third-party capability, reviewer lane, or EoS host integration through a handoff, an adoption fork, first-party absorption, or a de-listing
- Add or update a host's integration — set a host's documentation-sourced
runtime.hostIntegrationaxes (ADR-1239 Phase A), with theundocumentedsentinel rule - Migrate an install test to the executed plan — convert an
fs.existsSync-probing install test group to a value assertion againstinstallRuntimeArtifacts's executed-plan return, and test against a fake fs adapter - Vendor a dependency — add a third-party package
gsd-core/bin/**needs at runtime as a verbatim vendored artifact, keep it out ofdependencies, and pick the right upstream bundle - Turn a capability off (and keep it off) — disable a capability via the surface, or gate individual hooks off without removing the capability
- Drive GSD from a tracker issue — start a phase from a GitHub, Linear, or Jira issue
- Migrate from GSD 2 — upgrade an existing GSD 2 project to GSD Core
- Update GSD — re-run the installer to pick up the latest release
- Clean up get-shit-done-cc — remove leftover old-package artifacts that cause a spurious
⬆ /gsd-updateindicator after migrating to@opengsd/gsd-core - Fix the worktree base-mismatch (exit 42) error — resolve the branch-divergence condition that halts parallel phase execution
- Recover and troubleshoot — fix common problems, rebuild context, and uninstall
Reference
- Commands — every command with flags and examples
- Configuration — full config schema, model profiles, git branching strategies
- CLI tools —
gsd-tools.cjsprogrammatic API for workflows and agents - JSON error mode —
gsd-toolsfailure channels: faults (stderr, exit 1) vs degraded results (stdout, exit 0), and the reason-code taxonomy - Features — complete feature index
- Inventory — installed skills and surface map
- STATE.md schema — field-by-field reference for
.planning/STATE.md - CONTEXT.md schema — field-by-field reference for
.planning/phases/<N>/CONTEXT.md - PLAN.md schema — field-by-field reference for
.planning/phases/<N>/PLAN.md - Planning artifacts — all
.planning/files and their roles - Review and verification capabilities — code review, security, and Nyquist capability ownership and hook contracts
- Gate predicates — canonical specification of the phase-gate predicate vocabulary
- Capability matrix — generated catalogue of every capability's role, tier, extension points, hook kinds, and
engines.gsd - Exit code reference — generated catalogue of every registered process exit code, its name, meaning, and owning module, plus the reserved bands and the v1/v2 exit contract
- Capability manifest — the full
capability.jsonschema and validation rules gsd capabilitycommand — install / update / remove / list reference for third-party capabilities- Workflow fragments — in-file
<!-- gsd:section -->marker grammar for fragmentizing workflow markdown at emission time - Partition rules for compact-content splits — the protected-content list, sentinel syntax, and the five CI checks a
workflow.compact_contentspine/detail split must obey - Reviewer Lane Registry — generated catalogue of third-party reviewer lanes, with their flags, transport, and install commands
Explanation
- Context engineering — how context rot forms and how GSD Core prevents it
- The phase loop — design rationale for the Discuss → Plan → Execute → Verify → Ship cycle
- Multi-agent orchestration — how subagents are spawned, scoped, and coordinated
- Security model — trust boundaries, permissions, and safe automation
- The capability trust model — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- How overlay capabilities compose — why first-party always wins and how the loader resolves precedence, conflicts, and fail-open load-failure warnings
- Architecture — system architecture, agent model, and data flow
- The Embeddable Orchestration System — one public, versioned contract for embedding GSD across many hosts
- Discuss modes — assumptions mode vs interview mode for
/gsd-discuss-phase - Context monitoring — context window monitoring hook architecture
- Issue-driven orchestration — recipe for driving GSD from a tracker issue using existing primitives
Related
- What's new in 1.7.0 — curated highlights of the 1.7.0 release
- Root README — landing page, quickstart, and documentation overview
- Changelog — release history