Files
msd-core/docs/USER-GUIDE.md
Tom Boucher 241646a43a fix(#4651): classify .env names by final extension, and close the trailing-dot alias bypass — Phase 1 of #4636 (#4659)
* test(#4651): failing-first coverage for final-extension classification

Phase 1 of epic #4636, absorbing #4580. Tests only; no fix. These MUST fail.

The guard classifies a name by comparing everything after `.env.` as one
token against a set whose members are FINAL EXTENSIONS. So `.env.local.example`
yields suffix `local.example`, which is not a member, and a committed
secret-free template is refused. That is a category error, not strictness.

Two arms are covered because the same classification is hand-rolled twice in
one file: `isSecretBasename` for Read/Bash, and `globAltSelectsSecret`
(`lit.startsWith('.env.')`) for Grep globs. Fixing one alone would ship a
guard that allows `cat .env.local.example` while refusing
`Grep --glob '.env.local.example'` — the same file, the same hook, opposite
answers. A cross-arm parity loop over one shared list asserts the two cannot
drift.

Rows that exist because they are the ones nobody enumerates:

- `.env.example.local` must stay BLOCKED. Final extension is `local`; this is
  dotenv's documented local-override convention and a real secret. Any fix
  shaped as "contains example" admits it.
- `.env.local.` must stay BLOCKED — empty final extension is not a member.
- `.env.` must stay ALLOWED. Note #4580's proposed patch adds
  `if (suffix === '') return true;`, which flips it to blocked; that breaks the
  existing `allows` assertion in this suite and broadens the protected set,
  which epic #4636's non-goals forbid. Not applied.
- `.env.local.exam*` (partial glob literal) must stay BLOCKED — it can select
  `.env.local`, and a partial literal cannot be classified.
- `*.example` and `*` must stay ALLOWED — regression protection on the arm
  that already works.

Local behavioral repro of the current guard, confirming the tests fail for the
right reason rather than by construction:

  .env.local.example  rc=2 (blocked)   <- the defect
  .env.example        rc=0 (allowed)
  .env.local          rc=2 (blocked)
  .env.example.local  rc=2 (blocked)
  .env.               rc=0 (allowed)
  glob .env.local.example  rc=2        <- the second arm

Regressions are folded into the owning module's suite rather than a new
tests/fix-NNNN-*.test.cjs file, per scripts/lint-regression-test-names.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): classify by final extension so .env.<name>.example is readable

Phase 1 of epic #4636, absorbing #4580. Implements ADR-4650 decision 5.

The guard compared everything after `.env.` as ONE token against a set whose
members are FINAL EXTENSIONS. `.env.local.example` yielded `local.example`,
which is not a member, so a committed, secret-free template was refused — the
guard blocked the one file that exists so nobody has to open the real `.env`.

That is a category error, not strictness. The fix is not "add local.example to
the set"; it is to compare the right token. hooks/lib/filename-classification.js
now owns that distinction and is the only place it is expressed.

Both arms are fixed, because the same classification was hand-rolled twice in
this one file:

  - isSecretBasename (Read/Bash) now tests finalExtension(suffix).
  - globAltSelectsSecret (Grep --glob) split its first branch. With no
    wildcard the alternative IS a whole filename, so it is classified exactly
    via isSecretBasename. With a wildcard present the literal is only a
    PARTIAL prefix (`.env.local.exam*` can still select `.env.local`) and
    cannot be classified, so the original conservative rule stays.

Fixing only the first would have shipped a self-contradicting guard: `cat
.env.local.example` allowed while `Grep --glob '.env.local.example'` refused —
same file, same hook, opposite answers. A cross-arm parity loop over one shared
list now asserts the two cannot drift.

Two deliberate departures from #4580's suggested patch, both verified:

  - Its `if (suffix === '') return true;` is NOT applied. That flips `.env.`
    from allowed to blocked, breaking an existing assertion in this suite and
    broadening the protected set, which epic #4636's non-goals forbid.
  - `fullSuffix` was drafted alongside finalExtension and removed before
    commit: zero production consumers, and none planned (Phases 2-4 are
    containment, duplicate draining and the path-join ratchet, none of which
    classify filenames). A zero-caller export is dead code. The distinction is
    pinned instead by a test asserting finalExtension('local.example') is
    'example' and explicitly NOT 'local.example'.

The protected set is unchanged. `.env.example.local` stays BLOCKED — its final
extension is `local`, dotenv's local-override convention and a real secret;
any fix shaped as "contains example" admits it.

Scoped out by measurement, not assumption: src/validate.cts:395 and
src/phase.cts:1674 also hand-roll lastIndexOf('.'), but both parse phase
identifiers (`3.2` -> parent `3`), owned by the phase-id.cts seam. Folding
them in would repeat this same category error in the opposite direction.

Checkpoint 1 (prove RED) on the tests-only commit 91d3d6e1: outcome=failed,
26 failures / 45330, all 26 in the two new test files, zero pre-existing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): document the widened template exemption and cover the Bash arm

Two findings from the isolated adversarial review, both fixed in place.

1. The header's "Stated cost" passage named only the four literal template
   names, but since this change the exemption keys on the FINAL EXTENSION, so
   the trusted set is `.env.<anything>.{example,sample,template,dist}` — an
   unbounded family. The reviewer demonstrated it: `.env.prod-real-secrets.example`
   is allowed. That is the deliberate and necessary cost of fixing #4580, but
   it was materially larger than what the header disclosed, and a silent
   expansion of a security guard's trusted set is not acceptable. The passage
   now states the family, the concrete bypass, and that it applies across
   Read, Grep and Bash alike.

2. The cross-arm parity loop asserted Read and the exact-literal Grep glob but
   not Bash, whose `namesSecret` -> `isSecretBasename` path is genuinely
   distinct. The Bash arm was covered only by two one-off tests outside the
   shared table, so the table could not have caught a drift there. The loop now
   drives all three arms from the same TEMPLATES/SECRETS arrays.

No classification logic changed. The Read-arm behavioral table is byte-identical
before and after: rc=0 for .env.local.example / .env.example / .env. ; rc=2 for
.env.local / .env.example.local / .env / .secrets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): treat trailing dots and spaces as aliases of the protected file

Closes a Windows path-alias bypass surfaced by the isolated adversarial review
of this phase. Maintainer-approved as in scope.

Win32 strips trailing dots and spaces from every path component, so `.env.`,
`.env..`, `.env `, `.env. `, `.env .`, `.secrets.` and `.secrets ` all resolve
to the real `.env` / `.secrets` on Windows. The guard allowed every one of them
— a bypass of a file it already protects, reachable from Read, Grep and Bash
alike. `isSecretBasename` now normalizes the basename before classifying.

The whole class is fixed, not the reported name. `.env.` alone would have left
`.secrets.` and the trailing-space forms open, which is the same
one-cause-explains-every-failure trap this epic exists to close.

Two consequences, both measured rather than assumed:

  - `.env.example.` flips blocked -> ALLOWED. It aliases the already-trusted
    `.env.example` template, so this is correct; it was previously blocked only
    because the trailing dot broke final-extension parsing.
  - A Bash token that is exactly `.env` plus trailing whitespace flips
    allowed -> BLOCKED. Verified this is CONSISTENCY, not a new false-positive
    class: the bare `.env` token was ALREADY blocked as an operand in the same
    position before this change, so the alias now simply behaves like the thing
    it aliases.

The header's "No whitespace trimming" guarantee is preserved and now stated
precisely: leading and interior whitespace is still never trimmed, so prose
like a commit message mentioning `.env` in a sentence stays prose and stays
allowed. Only TRAILING dots and spaces are stripped. Two tests pin that.

This lands at the same behavior #4580's proposed `if (suffix === '') return
true;` would have produced for `.env.`, which this phase earlier rejected. The
rejection was correct on its stated grounds — that line broadens the protected
set, which epic #4636's non-goals forbid. The Windows framing is different:
normalizing an alias of an already-protected file is not a broadening, and the
fix is reached by normalization rather than by special-casing an empty suffix,
so it generalizes to `.secrets.` and the space forms.

Cannot be reproduced on this host — the remote matrix is Linux-only and Windows
coverage arrives from CI — so this ships on the Win32 path-normalization
contract plus the CI lane, and that limitation is stated rather than implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): one owner for path segmentation, closing a Read/Grep divergence

Four findings from the two-axis review, all fixed in place.

The real one: the guard had TWO path-segmentation rules. `lastSegment` (used
by Read and Bash via `namesSecret`) splits on both `/` and `\`, while
`classifyGrepGlob` hand-rolled its own on `/` only. Measured:

  Read  of `config\.env`      rc=2  BLOCKED
  Grep  --glob 'config\.env'  rc=0  ALLOWED

Same logical file, opposite answers — precisely the divergence this epic
exists to remove, sitting inside the file this phase was already fixing.
`lastSegment` now lives in hooks/lib/filename-classification.js and both arms
call it. All five path-bearing cases (both separators) now agree.

Note on how this was nearly missed: the first measurement of it reported
"both allow", which looked like the reviewer was wrong. That reading was a
measurement artifact — `config\.env` inside a printf'd JSON payload is an
invalid escape, so the hook fails open at rc=0 and the test was observing
JSON breakage rather than the predicate. Re-measured with correct escaping,
the divergence is real. The tests added here use properly escaped literals
and were verified by running, not by reasoning about the escaping.

Also fixed:

  - Both fast-check properties were satisfied by a degenerate
    always-return-'' implementation: "never contains a dot / is a suffix" and
    "never ends with dot-or-space / is a prefix" are both trivially true of
    the empty string. They now additionally pin content preservation — the
    removed tail must match /^[. ]*$/, and a name with nothing to strip must
    come back unchanged.
  - The cross-arm parity loop used only bare basenames, so it could not have
    caught the divergence above. It now covers path-bearing names with both
    separators.
  - That loop's description overclaimed: Read and Bash BOTH route through
    `namesSecret`, so they are not independent paths; only the Grep glob arm
    is genuinely separate. The description now says so rather than implying
    three-way independence.
  - `normalizeWindowsBasename` runs on every platform, not only Windows. Its
    doc now states that explicitly: the guard must answer identically
    everywhere, and a name is judged by what Win32 would resolve it to.

No classification logic changed; the 12-name regression sweep is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): regenerate install-tree goldens, correct the guard's user-facing docs

Three things, all consequences of the fix rather than new behavior.

1. Install-tree goldens. `hooks/lib/filename-classification.js` is a SHIPPED
   file — package.json `files` includes `hooks` — so every per-runtime install
   tree gains a path. Checkpoint 2 failed on exactly this: 11 failures, all in
   tests/golden-install-tree.test.cjs, against 45356 passing. Regenerated via
   scripts/gen-install-tree-fixtures.cjs; 11 goldens changed, matching the 11
   failures one-for-one.

   This ripple was identified at design time and then not acted on. Fleet's
   impact preview named golden-install-tree.test.cjs before any code was
   written, and 40-design.md records it under "Ripples identified". Writing a
   risk down is not the same as discharging it, and a full matrix run was spent
   discovering something already known.

2. docs/USER-GUIDE.md made a precise and now-false claim about the guard's
   protected set: it named `.env.example` / `.sample` / `.template` / `.dist`
   as the four exempt names. The exemption keys on the FINAL EXTENSION, so the
   exempt set is the unbounded family `.env.<anything>.{example,sample,template,dist}`.
   The page now states that family, the widened residual, that order matters
   and only the last segment counts (`.env.example.local` is a secret), and
   that trailing dots and spaces are stripped because Windows resolves them to
   the protected file. A wrong user-facing model of what a security guard
   protects is worth correcting even though Fixed/Security changesets are
   exempt from the required-docs rule.

   docs/ARCHITECTURE.md and docs/INVENTORY.md say "templates such as
   `.env.example` exempt" — non-exhaustive, still true, deliberately left
   alone. Same for the ja-JP / zh-CN / ko-KR / pt-BR rows, which carry the same
   hedged phrasing; hand-translating a security description unreviewed is not
   something to do silently.

3. Two changeset fragments, not one. A refusal corrected is `Fixed`; a bypass
   closed is `Security`. Folding the second into the first would under-report
   it in the release notes. Both carry `pr: 0` for backfill once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): backfill changeset PR number to 4659

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 08:42:43 -04:00

62 KiB
Raw Blame History

GSD User Guide

A narrative companion guide to GSD Core — orient yourself here, then follow the links into the dedicated docs.

GSD Core's documentation is organised by Diataxis. Browse by goal: Tutorials · How-to guides · Reference · Explanation · Docs index


Table of Contents

For driving GSD directly from a GitHub / Linear / Jira issue, see the Issue-driven orchestration guide — a recipe that maps tracker issues onto the workspace → discuss → plan → execute → verify → review → ship loop using existing GSD primitives.


Slash-command form

GSD ships the same set of skills to every supported runtime, using the hyphen slash-form spelling:

  • Hyphen form — /gsd-command-name — used by Claude Code, Copilot, OpenCode, Kilo, Cursor, Windsurf, Augment, Antigravity, and Trae.

The installer writes this form into the command directory of each runtime you target.

Namespace routing primer (gsd-ns-*, v1.40+)

Architecture

GSD ships six namespace router bundles (gsd-ns-workflow, gsd-ns-project, gsd-ns-review, gsd-ns-context, gsd-ns-ideate, gsd-ns-manage). On runtimes with non-recursive skill loaders, the installer emits these 6 routers as the only top-level skill entries; the ~61 concrete skills are nested under each router at <router>/skills/<name>/SKILL.md. This reduces the eager skill-listing overhead to ≈6 entries instead of ≈67.

Each router's body contains a routing table. When the model receives a request, it reads the router, identifies the relevant sub-skill by name, then opens skills/<name>/SKILL.md via a file-path Read. The concrete skill is fully available — it is not invocable by bare name through the Skill tool's top-level listing, but is reachable through the router.

The nested layout applies only to runtimes with confirmed non-recursive skill loaders: Cline, Qwen, Hermes, Augment, Trae. Claude's loader is also non-recursive, but #924 reverted it flat because the Skill tool hard-errors on unknown names rather than re-routing via the router. Antigravity's loader is also non-recursive, but #1614 moved it flat because agy scans only skills/<name>/SKILL.md — nested sub-skills were unreachable. Other recursive or unconfirmed loaders (Cursor, Codex, Copilot, Windsurf, CodeBuddy, OpenCode, Kilo) retain the flat layout unchanged.

Namespace Router bundle Routes to
Phase pipeline gsd-ns-workflow discuss / plan / execute / verify / phase / progress
Project lifecycle gsd-ns-project milestones, audits, summary
Quality gates gsd-ns-review code review, debug, audit, security, eval, ui
Codebase intelligence gsd-ns-context map, graphify, docs, learnings
Exploration & capture gsd-ns-ideate explore, sketch, spike, spec, capture
Management gsd-ns-manage config, workspace, workstreams, thread, update, ship, inbox

Slash commands are unaffected

On runtimes that install a commands surface (commands/gsd), slash commands such as /gsd-plan-phase continue to work directly — the nesting applies only to the Skill tool's top-level listing, not to the commands directory.

Migration note (breaking change on nesting runtimes)

On the seven nesting runtimes listed above, upgrading to v1.40 changes skill invocation behaviour:

  • Before: each of the ~67 concrete gsd-<name> skills appeared at the top level and was invocable by bare name through the Skill tool.
  • After: only the 6 gsd-ns-* router bundles appear at the top level. Concrete skills are reachable via the router's routing table and a Read skills/<name>/SKILL.md call. Direct bare-name invocation of concrete skills through the Skill tool's listing no longer works.
  • Slash commands unchanged: /gsd-plan-phase, /gsd-discuss-phase, etc. still work directly where a commands surface is installed.
  • Upgrade prune: the installer's existing prune step removes the legacy top-level gsd-<concrete>/ skill directories on upgrade — no manual cleanup is needed.

Reading GSD's output

GSD marks its sections with Markdown, not with drawn borders. There are exactly three forms, and every workflow, checkpoint and report uses them:

You see It means
### GSD ► {STAGE NAME} A major workflow transition — planning, executing a wave, verifying, completing
### CHECKPOINT: {Type} GSD is waiting on you. The bolded **→ …** line at the bottom tells you what to type
--- A break between two sections — most often before the ▶ Next Up block at the end of a completion

Panels that used to be drawn with box characters are now a heading followed by their rows as ordinary lines. A checkpoint reads:

### CHECKPOINT: Verification Required

Progress: 5/8 tasks complete
Task: Responsive dashboard layout

How to verify:
  1. Visit: http://localhost:3000/dashboard

---

**→ YOUR ACTION: Type "approved" or describe issues**

Why there are no drawn borders

Earlier releases framed stage banners between two 53-character runs of ━, and checkpoints were a 62-column box drawn with ╔, ║ and ╚. Those runs are ordinary text to whatever renders GSD's output. In a pane narrower than the run, the rule wraps and the leftover glyphs land on a second line — so the border comes apart from the heading it was framing, and the output looks broken rather than merely narrow. Reported against the Codex desktop interface in #3028.

A Markdown heading and a thematic break carry the same structure without committing to a width, so they read correctly in a narrow pane and a wide one.

This is unconditional: every runtime gets the same output. The alternative — a capability flag that kept line-art for terminal-oriented runtimes — was weighed and rejected, because it leaves two output conventions to keep in sync forever for a gain that is aesthetic rather than structural. What a terminal loses is the drawn frame; what it keeps is the title, the hierarchy and the break, none of which depended on the frame. The reasoning is recorded in gsd-core/references/ui-brand.md § Why this is unconditional, not per-runtime. If you run GSD in a host where the heading form reads worse than the old boxes did, that is worth reporting — it is the evidence that would justify the flag.

Single-cell tokens are unaffected and unchanged: the status symbols (✓ ✗ ◆ ○ ⚠), the GSD ► prefix, and the ten-cell progress gauge (Progress: ████████░░ 80%) are not runs and do not wrap.

The convention is specified in gsd-core/references/ui-brand.md and enforced against all shipped content by tests/responsive-separators.test.cjs.


Project lifecycle overview

The core GSD loop is: discuss → plan → execute → verify → ship, repeated per phase. The full step-by-step walkthrough — including example outputs, what files get created, and all the flags in play — is in the dedicated tutorial.

See Your first project.

For onboarding an existing codebase before starting a new milestone, run /gsd-onboard or see Onboarding an existing codebase.

Relevant flags at a glance:

Flag Command When to use
--auto /gsd-new-project Skip interactive questions, ingest from a PRD file
--research /gsd-quick Add a research agent to an ad-hoc task
--validate /gsd-quick Add plan-checking and post-execution verification
--chain /gsd-discuss-phase Auto-chain discuss → plan → execute without stopping
--skip-research /gsd-plan-phase Skip research agents when the domain is already familiar
--draft /gsd-ship Create a draft PR instead of a ready-for-review one

For the full command reference with all flags, see docs/COMMANDS.md. For configuration options (model profiles, workflow agents, git branching), see docs/CONFIGURATION.md.


Workflow Diagrams

Full Project Lifecycle

  ┌──────────────────────────────────────────────────┐
  │                   NEW PROJECT                    │
  │  /gsd-new-project                                │
  │  Questions -> Research -> Requirements -> Roadmap│
  └─────────────────────────┬────────────────────────┘
                            │
             ┌──────────────▼─────────────┐
             │      FOR EACH PHASE:       │
             │                            │
             │  ┌────────────────────┐    │
             │  │ /gsd-discuss-phase │    │  <- Lock in preferences
             │  └──────────┬─────────┘    │
             │             │              │
             │  ┌──────────▼─────────┐    │
             │  │ /gsd-ui-phase      │    │  <- Design contract (frontend)
             │  └──────────┬─────────┘    │
             │             │              │
             │  ┌──────────▼─────────┐    │
             │  │ /gsd-plan-phase    │    │  <- Research + Plan + Verify
             │  └──────────┬─────────┘    │
             │             │              │
             │  ┌──────────▼─────────┐    │
             │  │ /gsd-execute-phase │    │  <- Parallel execution
             │  └──────────┬─────────┘    │
             │             │              │
             │  ┌──────────▼─────────┐    │
             │  │ /gsd-verify-work   │    │  <- Manual UAT
             │  └──────────┬─────────┘    │
             │             │              │
             │  ┌──────────▼─────────┐    │
             │  │ /gsd-ship          │    │  <- Create PR (optional)
             │  └──────────┬─────────┘    │
             │             │              │
             │     Next Phase?────────────┘
             │             │ No
             └─────────────┼──────────────┘
                            │
            ┌───────────────▼──────────────┐
            │  /gsd-audit-milestone        │
            │  /gsd-complete-milestone     │
            └───────────────┬──────────────┘
                            │
                   Another milestone?
                       │          │
                      Yes         No -> Done!
                       │
               ┌───────▼──────────────┐
               │  /gsd-new-milestone  │
               └──────────────────────┘

Planning Agent Coordination

  /gsd-plan-phase N
         │
         ├── Phase Researcher (x4 parallel)
         │     ├── Stack researcher
         │     ├── Features researcher
         │     ├── Architecture researcher
         │     └── Pitfalls researcher
         │           │
         │     ┌──────▼──────┐
         │     │ RESEARCH.md │
         │     └──────┬──────┘
         │            │
         │     ┌──────▼──────┐
         │     │   Planner   │  <- Reads PROJECT.md, REQUIREMENTS.md,
         │     │             │     CONTEXT.md, RESEARCH.md
         │     └──────┬──────┘
         │            │
         │     ┌──────▼───────────┐     ┌────────┐
         │     │   Plan Checker   │────>│ PASS?  │
         │     └──────────────────┘     └───┬────┘
         │                                  │
         │                             Yes  │  No
         │                              │   │   │
         │                              │   └───┘  (loop, up to 3x)
         │                              │
         │                        ┌─────▼──────┐
         │                        │ PLAN files │
         │                        └────────────┘
         └── Done

Validation Architecture (Nyquist Layer)

During plan-phase research, GSD maps automated test coverage to each phase requirement before any code is written. The researcher detects your existing test infrastructure, maps each requirement to a specific test command, and identifies any test scaffolding that must be created before implementation begins (Wave 0 tasks). The plan-checker enforces this as an 8th verification dimension: plans where tasks lack automated verify commands will not be approved.

Output: {phase}-VALIDATION.md — the feedback contract for the phase.

Disable: Set workflow.nyquist_validation: false in /gsd-settings for rapid prototyping phases where test infrastructure isn't the focus.

Retroactive Validation (/gsd-validate-phase)

For phases executed before Nyquist validation existed, or for existing codebases with only traditional test suites, retroactively audit and fill coverage gaps:

  /gsd-validate-phase N
         |
         +-- Detect state (VALIDATION.md exists? SUMMARY.md exists?)
         |
         +-- Discover: scan implementation, map requirements to tests
         |
         +-- Analyze gaps: which requirements lack automated verification?
         |
         +-- Present gap plan for approval
         |
         +-- Spawn auditor: generate tests, run, debug (max 3 attempts)
         |
         +-- Update VALIDATION.md
               |
               +-- COMPLIANT -> all requirements have automated checks
               +-- PARTIAL -> some gaps escalated to manual-only

The auditor never modifies implementation code — only test files and VALIDATION.md. If a test reveals an implementation bug, it's flagged as an escalation for you to address.

Assumptions Discussion Mode

By default, /gsd-discuss-phase asks open-ended questions about your implementation preferences. Assumptions mode inverts this: GSD reads your codebase first, surfaces structured assumptions about how it would build the phase, and asks only for corrections.

Enable: Set workflow.discuss_mode to 'assumptions' via /gsd-settings.

See docs/workflow-discuss-mode.md for the full discuss-mode reference.

Decision Coverage Gates

The discuss-phase captures implementation decisions in CONTEXT.md under a <decisions> block as numbered bullets (- **D-01:** …). Two gates ensure those decisions survive into plans and shipped code.

Plan-phase translation gate (blocking). After planning, GSD refuses to mark the phase planned until every trackable decision appears in at least one plan's scanned surfaces: front-matter must_haves/truths/objective, a ## must_haves/truths/tasks/objective heading, or an <objective>/<tasks>/<task>/<action>/<read_first>/<behavior>/<verify>/<acceptance_criteria>/<done> tag body.

Verify-phase validation gate (non-blocking). During verification, GSD searches plans, SUMMARY.md, modified files, and recent commit messages for each trackable decision. Misses are logged to VERIFICATION.md as a warning section; verification status is unchanged.

Opting a decision out. Move it under the ### Claude's Discretion heading inside <decisions>, or tag it: - **D-08 [informational]:** …, - **D-09 [folded]:** …, - **D-10 [deferred]:** ….

Disabling the gates. Set workflow.context_coverage_gate: false in .planning/config.json (or via /gsd-settings). Default is true.

Execution Wave Coordination

  /gsd-execute-phase N
         │
         ├── Analyze plan dependencies
         │
         ├── Wave 1 (independent plans):
         │     ├── Executor A (fresh 200K context) -> commit
         │     └── Executor B (fresh 200K context) -> commit
         │
         ├── Wave 2 (depends on Wave 1):
         │     └── Executor C (fresh 200K context) -> commit
         │
         └── Verifier
               ├── Check codebase against phase goals
               ├── Test quality audit (disabled tests, circular patterns, assertion strength)
               │
               ├── PASS -> VERIFICATION.md (success)
               └── FAIL -> Issues logged for /gsd-verify-work

Isolated-run Recovery (fail-safe)

When a worktree-isolated run is rejected — the user declines to merge it, or the run over-reached the requested scope, or the orchestrator surfaces recovery guidance for a blocked plan — GSD halts safely and offers two options: (a) re-attempt in a fresh, narrowly-scoped worktree, or (b) inspect or discard the rejected worktree without merging. GSD never defaults recovery to editing the primary checkout (main). Any path that edits the primary checkout requires explicit, clearly-labeled confirmation from the user first. This behavior is unconditional and applies to both /gsd-execute-phase (worktree executor waves) and /gsd-quick (quick-mode isolated runs).


UI Design Contract

AI-generated frontends are visually inconsistent not because Claude Code is bad at UI but because no design contract existed before execution. /gsd-ui-phase locks the design contract before planning; /gsd-ui-review audits the result after execution.

For the full workflow, configuration, shadcn initialisation, and the registry safety gate, see Design a UI phase.

Quick reference:

Command Description
/gsd-ui-phase [N] Generate UI-SPEC.md design contract for a frontend phase
/gsd-ui-review [N] Retroactive 6-pillar visual audit of implemented UI
Setting Default Description
workflow.ui_phase true Generate UI design contracts for frontend phases
workflow.ui_safety_gate true plan-phase prompts to run /gsd-ui-phase for frontend phases

Spiking & Sketching

Use /gsd-spike to validate technical feasibility before planning, and /gsd-sketch to explore visual direction before designing. Both store artifacts in .planning/ and integrate with the project-skills system via their wrap-up companions.

For the full workflow and flow diagram, see Spike and sketch.

Typical flow:

/gsd-spike "SSE vs WebSocket"     # Validate the approach
/gsd-spike --wrap-up              # Package learnings

/gsd-sketch "real-time feed UI"   # Explore the design
/gsd-sketch --wrap-up             # Package decisions

/gsd-discuss-phase N              # Lock in preferences (now informed by spike + sketch)
/gsd-plan-phase N                 # Plan with confidence

Backlog & Threads

Backlog Parking Lot

Ideas that aren't ready for active planning go into the backlog using 999.x numbering, keeping them outside the active phase sequence.

/gsd-capture --backlog "GraphQL API layer"     # Creates 999.1-graphql-api-layer/
/gsd-capture --backlog "Mobile responsive"     # Creates 999.2-mobile-responsive/

Backlog items get full phase directories, so you can use /gsd-discuss-phase 999.1 to explore an idea further or /gsd-plan-phase 999.1 when it's ready. Backlog directories (and the 0-* pre-milestone directory some projects carry) are excluded from /gsd-progress, /gsd-stats, and phase listings for the current milestone — they stay out of the active phase sequence for counting purposes too, not just for planning.

Review and promote with /gsd-review-backlog — it shows all backlog items and lets you promote (move to active sequence), keep (leave in backlog), or remove (delete).

Seeds

Seeds are forward-looking ideas with trigger conditions. Unlike backlog items, seeds surface automatically when the right milestone arrives.

/gsd-capture --seed "Add real-time collab when WebSocket infra is in place"

/gsd-new-milestone scans all seeds and presents matches. Storage: .planning/seeds/SEED-NNN-slug.md

Once you've parked a few, audit them on demand instead of waiting for the next milestone to surface them:

/gsd-capture --list-seeds            # Review every parked seed
/gsd-capture --list-seeds dormant    # Narrow to one status

This is read-only — it renders an audit table (ID, status, scope, trigger, title) and a per-status summary, and never modifies a seed. Filter by dormant, active, or triggered when you only want to see seeds in one state.

Persistent Context Threads

Threads are lightweight cross-session knowledge stores for work that spans multiple sessions but doesn't belong to any specific phase.

/gsd-thread                              # List all threads
/gsd-thread fix-deploy-key-auth          # Resume existing thread
/gsd-thread "Investigate TCP timeout"    # Create new thread

Threads can be promoted to phases (/gsd-phase) or backlog items (/gsd-capture --backlog) when they mature. Storage: .planning/threads/{slug}.md


Workstreams & Workspaces

Workstreams and workspaces both provide isolation, but at different levels.

Workstreams share the same codebase and git history but isolate planning artifacts — lighter weight, good for working on multiple milestone areas concurrently. See Work in parallel with workstreams.

Workspaces create separate repo worktrees with their own .planning/ — heavier, for feature-branch or multi-repo isolation. See Isolate work with workspaces.

Command Purpose
/gsd-workstreams create <name> Create a new workstream with isolated planning state
/gsd-workstreams switch <name> Switch active context to a different workstream
/gsd-workstreams list Show all workstreams and which is active
/gsd-workstreams complete <name> Mark a workstream as done and archive its state
# Workspace example — feature branch isolation
/gsd-workspace --new --name feature-b --repos .
cd ~/gsd-workspaces/feature-b
/gsd-new-project

/gsd-workspace --list
/gsd-workspace --remove feature-b

Security

Defense-in-Depth (v1.27)

GSD generates markdown files that become LLM system prompts. This means any user-controlled text flowing into planning artifacts is a potential indirect prompt injection vector. v1.27 introduced centralised security hardening:

Path Traversal Prevention: All user-supplied file paths (--text-file, --prd) are validated to resolve within the project directory. macOS /var → /private/var symlink resolution is handled.

Prompt Injection Detection: The security.cjs module scans for known injection patterns in user-supplied text before it enters planning artifacts.

Runtime Hooks:

  • gsd-prompt-guard.js — Scans Write/Edit calls to .planning/ for injection patterns (always active, advisory-only)
  • gsd-workflow-guard.js — Warns on file edits outside GSD workflow context (opt-in via hooks.workflow_guard)
  • gsd-write-guard.js — Hard-blocks a whole-file Write that catastrophically shrinks a curated .planning/ artifact (ROADMAP.md, milestone roadmaps, STATE.md) below 40% of its on-disk line count; files under 40 lines are exempt. The check is stateless per Write, comparing each payload against the file's current on-disk size — a single-shot collapse (the #973 shape) is blocked, but a sequence of individually-tolerated shrinks that erodes the file across several Writes is not detected. For a legitimate milestone reset or large deletion, bypass once with the single-use sentinel — write the target's path into .planning/.gsd-allow-shrink (fresh within 15 minutes; consumed by the allowed write) — or, interactively, with GSD_ALLOW_PLANNING_SHRINK=1 in the runtime's environment. Scope the guarantee accordingly: this stops accidental and single-shot collapse, and is not a defense against a determined agent — the sentinel is a plain file, so anything with shell access can arm one; what it buys is that the bypass becomes a deliberate, path-bound, single-use and auditable action rather than a sentence to reason past (always active, blocking; #2255, fix 3 of #973)
  • gsd-secret-read-guard.js — Hard-blocks reads of secret files — .env, .env.<suffix> and .secrets, matched case-insensitively (.ENV, .Secrets) — through Read (file_path), Grep (an explicit path, or a glob that selects them, judged per brace alternative) and Bash (operands, input redirects, $( ) / backtick / <( ) bodies, and git show <ref>:<path> shapes). A shell interpreter (bash/sh/zsh/dash/ksh) has its script scanned however it arrives — -c '…', a <( ) file operand, a heredoc / here-string, or a pipe from a knowable echo/printf source (echo cat .env | bash) — as do eval's joined operands, a source/. process-substitution operand, and find … | xargs cat pipelines (upstream literal names become the sub-command's read operands). A name is exempt when its final extension is example, sample, template or dist, so .env.example and equally .env.local.example / .env.production.sample stay readable (they are the templates GSD's own phase prompt reads). Stated cost: the exempt set is that unbounded family, not four fixed names — a real secret is not protected if it is named to end in one of those four extensions. Order matters, and only the last segment counts: .env.example.local is a secret and stays blocked. Trailing dots and spaces are stripped before the name is judged, because Windows resolves .env., .env and .secrets. to the protected file itself. Existence checks ([ -f .env ], ls .env*, test, stat, rm, touch, echo, …) pass. Not covered, by construction: $VAR indirection (bash -c "$CMD"), shell globs (cat .e*), interpreter one-liners, a piped script from a non-echo/printf source (cat gen.sh | bash, curl … | sh), reads inside scripts the agent runs, and a Grep glob: '*' reaching a .env that is not gitignored — none are statically resolvable by a hook. This replaces the Read(.env) / Read(.env.*) / Read(.secrets) permission deny rules the installer used to write: on Claude Code ≥ 2.1.259 any Read() deny rule makes every cd DIR && grep … compound prompt for approval even in auto mode, while a hook denial is not a permission rule and applies in auto and bypassPermissions alike (always active, blocking; #4221)

CI Scanner: prompt-injection-scan.security.test.cjs scans all agent, workflow, and command files for embedded injection vectors.


Package Legitimacy Gate (v1.42.1)

AI coding tools hallucinate package names. Attackers pre-register those names on npm, PyPI, and crates.io with malicious post-install scripts — a technique called slopsquatting. v1.42.1 adds a three-layer gate that stops this before it reaches your shell.

In RESEARCH.md — every phase that recommends external packages includes a ## Package Legitimacy Audit table:

## Package Legitimacy Audit

| Package | Registry | Age | Downloads | Source Repo | Verdict | Disposition |
|---------|----------|-----|-----------|-------------|---------|-------------|
| express | npm | 13 yrs | 100M+/wk | github.com/expressjs/express | [OK] | Approved |
| some-new-util | npm | 3 days | 47 | none | [SLOP] | REMOVED |
| api-bridge | npm | 6 mo | 1.2k/wk | github.com/user/api-bridge | [SUS] | Flagged |

[SLOP] packages are removed from RESEARCH.md entirely and never reach the planner.

In PLAN.md — [SUS] or [ASSUMED] packages trigger a checkpoint:human-verify task before the install.

During execution — if an install fails, the executor surfaces a checkpoint and stops rather than silently trying an alternative.

Legitimacy verdicts:

Verdict Meaning GSD action
[OK] Passes all legitimacy checks Proceeds — no checkpoint added
[SUS] Suspicious signals Flagged; planner adds checkpoint:human-verify
[SLOP] High-confidence hallucination Removed from RESEARCH.md; never reaches planner

Verdicts are computed from live registry APIs (npm, PyPI, crates.io) — there is no separate tool to install. slopcheck is an optional escalate-only adapter (it can raise a verdict but never lower one); no shipped configuration wires it, and its absence does not change the gate's behavior.


Code Review Workflow

After executing a phase, run a structured code review before UAT. See Set up cross-AI review for the full workflow.

/gsd-code-review 3               # Review all changed files in phase 3
/gsd-code-review 3 --depth=deep  # Deep cross-file review
/gsd-code-review 3 --fix         # Fix Critical + Warning findings atomically
/gsd-code-review 3 --fix --auto  # Fix and re-review until clean (max 3 iterations)
/gsd-audit-fix                   # Audit + classify + fix (medium+ severity, max 5)

The review step slots in after execution and before UAT:

/gsd-execute-phase N  ->  /gsd-code-review N  ->  /gsd-code-review N --fix  ->  /gsd-verify-work N

Optional external source-review lanes (#4209): /gsd-code-review accepts the same reviewer-lane flags as /gsd-review (run gsd_run review-lane flags to list the flags your installation's roster declares, e.g. --codex, --agy). Adding one asks that lane to independently review the same file scope alongside the internal gsd-code-reviewer agent; its findings are unverified corroborating evidence that gsd-code-reviewer re-checks against the actual source before writing anything to REVIEW.md — there is still exactly one REVIEW.md. No reviewer-lane flag is the default and reviews with only the internal agent, unchanged from before #4209. This is separate from /gsd-review, which reviews PLAN.md files before execution, not source code — see Set up cross-AI review.

/gsd-code-review 3 --codex       # Corroborate the internal review with the codex reviewer lane

Coverage-Aware UAT Routing

Historically, /gsd-verify-work turned every ## Accomplishments bullet in a SUMMARY into a manual checkpoint — even deliverables already covered one-to-one by a passing unit test. With a green test suite you were still asked to re-confirm things the tests had already proven, every phase.

GSD now lets the executor record, at authoring time, how each deliverable was verified. When a SUMMARY.md carries a coverage: frontmatter block (see the coverage: block reference), /gsd-verify-work routes deterministically:

  • Auto-passed — a deliverable marked human_judgment: false whose verification list is non-empty and entirely pass is recorded as passed (source: automated) and never prompted.
  • Presented — everything else is shown to you for sign-off: anything flagged human_judgment: true (visual adequacy, multi-device behaviour, subjective quality), anything with no verification, anything not fully passing, and any malformed entry.

The asymmetry is deliberate. The worst outcome is auto-passing something broken that UAT existed to catch, so auto-pass is the narrow, fully-proven case and uncertainty always routes back to you. Flipping the flag alone cannot skip a prompt — a passing test reference is also required. SUMMARYs without a coverage: block behave exactly as before (prose-based checkpoints), so nothing changes for existing or un-migrated phases.


Command And Configuration Reference

  • Command Reference: see docs/COMMANDS.md for every stable command's flags, subcommands, and examples.
  • Configuration Reference: see docs/CONFIGURATION.md for the full config.json schema, model-profile table, git branching strategies, and security settings.
  • Discuss Mode: see docs/workflow-discuss-mode.md for interview vs assumptions mode.

Graphify capability gate (tri-state, v1.43+)

Graphify commands (graphify status, graphify build, graphify query, graphify diff) now respect the full tri-state capability gate:

  1. Installed — the gsd-graphify-* skills are present in the active install profile.
  2. Surfaced — those skills appear on the current runtime surface (e.g., in ~/.claude/commands/gsd/).
  3. Config-enabled — graphify.enabled: true is set in .planning/config.json.

All three conditions must be true. Setting graphify.enabled: true alone is no longer sufficient if graphify has not been installed and surfaced. If graphify commands return { disabled: true } after upgrading, verify that the install profile includes graphify skills (gsd-tools capability state) and re-run the installer to surface them.

Intel capability gate (tri-state, v1.44+)

Intel commands (intel status, intel query, intel diff, intel snapshot, intel validate, intel api-surface) now respect the full tri-state capability gate (same resolver as graphify above):

  1. Installed — the intel capability is present in the active install profile (intel has no skill files, so this is vacuously true for all profiles).
  2. Surfaced — the intel capability is on the current runtime surface (vacuously true for all surfaces since intel registers no skill stems).
  3. Config-enabled — intel.enabled: true is set in .planning/config.json.

For intel, conditions 1 and 2 are always satisfied (intel has no skill files). The effective gate is intel.enabled in config — the same behaviour as before, but now enforced through the shared isCapabilityActive('intel', cwd) resolver rather than a direct config read. This means intel honours the full capability-state pipeline, including any future install-profile or surface restrictions. If intel commands return { disabled: true }, ensure intel.enabled: true is set in .planning/config.json and verify gsd-tools capability state shows intel as active.

Compact content mode (workflow.compact_content, v4139+)

GSD's own workflow instructions, planning-artifact templates, and (for non-Claude runtimes) agent personas are prose — and a large eagerly-loaded instruction body is context the model spends on GSD's own orchestration rather than on your code. workflow.compact_content (default false) opts a project into token-minimized variants of that content wherever one has been authored, without changing what GSD actually does.

Why turn it on: finite attention, not price. The point of this key is not a cheaper invocation — with prompt caching, the per-request cost of re-sending a large instruction file is already small. The point is what a large always-loaded instruction body costs in attention: every byte of GSD's own prose sitting in context is a byte not spent reasoning about your codebase. That cost is paid whether or not the tokens were cheap to transport. Turn it on when you're running long sessions, working in a large codebase that already competes for context, or on a runtime with a small context window; leave it off (the default) if you'd rather have every elaboration and worked example available up front, or you're evaluating GSD for the first time and want full detail while you learn how it thinks.

What actually changes. Nothing is compressed at runtime. Every compact variant is a hand-authored, reviewed file sitting beside its canonical sibling — the key only chooses which one GSD reads. Guardrails, output-format contracts, few-shot examples, and security language are never shortened or dropped in a compact variant; only rarely-needed elaboration and restatement are. With the key off, behavior is unchanged from before this feature existed.

How to turn it on: answer "Yes" to the Compact Content question during /gsd-new-project, or run /gsd-settings (or /gsd-config with no flag) on an existing project and toggle Compact Content. See docs/CONFIGURATION.md for the mechanics and ADR-4139 for the full design rationale.


Usage Examples

New Project (Full Cycle)

claude --dangerously-skip-permissions
/gsd-new-project            # Answer questions, configure, approve roadmap
/clear
/gsd-discuss-phase 1        # Lock in your preferences
/gsd-ui-phase 1             # Design contract (frontend phases)
/gsd-plan-phase 1           # Research + plan + verify
/gsd-execute-phase 1        # Parallel execution
/gsd-verify-work 1          # Manual UAT
/gsd-ship 1                 # Create PR from verified work
/gsd-ui-review 1            # Visual audit (frontend phases)
/clear
/gsd-progress --next                   # Auto-detect and run next step
...
/gsd-audit-milestone        # Check everything shipped
/gsd-complete-milestone     # Archive, tag, done
/gsd-pause-work --report         # Generate session summary

Caution

The permissions flag is optional. It skips per-file confirmation while GSD's sub-agents read and write files. Use it only in low-stakes or throwaway contexts. To keep confirmations enabled, start with claude instead. For real work, read the security model first.

New Project from Existing Document

/gsd-new-project --auto @prd.md   # Auto-runs research/requirements/roadmap from your doc
/clear
/gsd-discuss-phase 1               # Normal flow from here

Existing Codebase

/gsd-onboard                # Safely map, ingest docs, and initialize planning
# Follow the printed top-level handoff commands, then rerun /gsd-onboard
# (normal phase workflow from here)

/gsd-onboard routes through /gsd-map-codebase, /gsd-ingest-docs, and /gsd-new-project without nesting interactive workflows or overwriting existing planning files silently.

Post-execute drift detection (#2003). After every /gsd-execute-phase, GSD checks whether the phase introduced enough structural change to make .planning/codebase/STRUCTURE.md stale. Flip the behavior with:

/gsd-settings workflow.drift_action auto-remap       # remap automatically
/gsd-settings workflow.drift_threshold 5             # tune sensitivity

Plan Drift Guard

Default-on. The plan drift guard (plan_review.source_grounding: true) runs during plan review and verifies that every symbol your plans cite — decorators, classes, functions, CLI flags — actually exists in your source tree at review time. This catches hallucinated names before any execution agent runs.

Two axes, one switch. The same guard also runs a cross-artifact fact-drift pass: when ROADMAP.md, PLAN.md, STATE.md and CONTEXT.md state the same fact in contradictory ways — a phase marked complete in one and in progress in the other, a success criterion the plan restates with a different outcome, a term used against its CONTEXT.md definition — you get an advisory finding in REVIEWS.md naming both locations and which one is authoritative. It keys on contradicting knowledge, not on similar-looking text, so a plan that simply restates a criterion in its own words is not flagged. The findings never block convergence.

What it catches:

  • Functions referenced in a PLAN.md step that don't exist in source
  • Class or decorator names that were renamed or removed since the plan was written
  • CLI flags documented in a plan that are not defined in the argument parser
  • Module paths cited in implementation steps that resolve to no files

Needs-acknowledgement behavior. When the guard finds a missing symbol, it emits a needs-acknowledgement notice in the plan review output rather than hard-blocking. You can acknowledge and proceed (the symbol may be intentionally new) or request a plan revision. The guard does not auto-reject plans — it surfaces signal for human decision.

Works without intel. By default the guard uses grep/ripgrep to search source files — no pre-indexing required. If you have run /gsd-map-codebase with intel.enabled: true, set plan_review.source_grounding_authority: intel to use the faster pre-built api-map.json index instead.

# Enable/disable (default: on)
/gsd-settings plan_review.source_grounding true
/gsd-settings plan_review.source_grounding false

# Switch resolver authority
/gsd-settings plan_review.source_grounding_authority grep   # live grep (default)
/gsd-settings plan_review.source_grounding_authority intel  # pre-indexed api-map.json

Toggle at project setup (/gsd-new-project asks during workflow preferences) or any time via /gsd-settings (Planning section → Drift Guard).

Quick Bug Fix

/gsd-quick
> "Fix the login button not responding on mobile Safari"

Resuming After a Break

/gsd-progress               # See where you left off and what's next
# or
/gsd-resume-work            # Full context restoration from last session

Preparing for Release

/gsd-audit-milestone        # Check requirements coverage, detect stubs
/gsd-complete-milestone     # Archive, tag, done

Speed vs Quality Presets

Scenario Mode Granularity Profile Research Plan Check Verifier
Prototyping yolo coarse budget off off off
Normal dev interactive standard balanced on on on
Production interactive fine quality on on on

Skipping discuss-phase in autonomous mode: When running in yolo mode, set workflow.skip_discuss: true via /gsd-settings.

Mid-Milestone Scope Changes

/gsd-phase                  # Append a new phase to the roadmap (default mode)
/gsd-phase --insert 3       # Insert urgent work between phases 3 and 4
/gsd-phase --remove 7       # Descope phase 7 and renumber
/gsd-phase --edit 4         # Edit any field of phase 4 in place

Troubleshooting

For a comprehensive troubleshooting guide, see Recover and troubleshoot. The most common issues are summarised below.

Programmatic CLI (gsd-tools query vs gsd-tools.cjs)

For automation, prefer gsd-tools query with a registered subcommand (see CLI-TOOLS.md — SDK and programmatic access and QUERY-HANDLERS.md). The legacy node $HOME/.claude/gsd-core/bin/gsd-tools.cjs CLI remains supported.

STATE.md Out of Sync

node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" state validate          # Detect drift
node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" state sync --verify     # Preview changes
node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" state sync              # Reconstruct STATE.md

state validate's report carries a scope field alongside valid — valid:true means no drift was found, but scope says whether the check could actually run at all (a phase that could not be resolved, or an unreadable frontmatter/phases directory, reports valid:true too, because there was nothing to flag). See Interpret state validate results before treating a passing state validate as "clean."

A Command Looks Frozen After "Spawning..."

GSD subagents run in a separate context window — their work is invisible to the parent session while in progress. Do not interrupt the session. Wait for the result; research and planning agents routinely take 1–5 minutes.

Context Degradation During Long Sessions

Clear your context window between major commands: /clear in Claude Code. GSD is designed around fresh contexts — every subagent gets a clean 200K window. Use /gsd-resume-work or /gsd-progress to restore state after clearing.

Plans Seem Wrong or Misaligned

Run /gsd-discuss-phase [N] before planning. Most plan quality issues come from Claude making assumptions that CONTEXT.md would have prevented.

Execution Fails or Produces Stubs

Check that the plan was not too ambitious. Plans should have 2–3 tasks maximum. Re-plan with smaller scope.

Lost Track of Where You Are

Run /gsd-progress. It reads all state files and tells you exactly where you are and what to do next.

Model Costs Too High

Switch to budget profile: /gsd-config --profile budget. Disable research and plan-check agents via /gsd-settings if the domain is familiar.

Tuning model cost by phase (models) — added in v1.40

Add a models block to .planning/config.json:

{
  "model_profile": "balanced",
  "models": {
    "planning": "opus",
    "discuss": "opus",
    "research": "sonnet",
    "execution": "opus",
    "verification": "sonnet",
    "completion": "sonnet"
  }
}

Need a per-agent exception? Add model_overrides alongside — it wins over models:

{
  "models": { "research": "sonnet" },
  "model_overrides": {
    "gsd-codebase-mapper": "haiku"
  }
}

For the full mapping table and resolution-precedence rules, see Per-Phase-Type Models.

Cheap-by-default with dynamic_routing — added in v1.40

{
  "dynamic_routing": {
    "enabled": true,
    "tier_models": {
      "light":    "haiku",
      "standard": "sonnet",
      "heavy":    "opus"
    },
    "escalate_on_failure": true,
    "max_escalations": 1
  }
}

For the full agent → tier mapping, see Dynamic Routing.

Trim MCP servers to reduce per-turn cost

Before tuning model_profile or models.<phase_type>, audit which MCP servers your harness has enabled. Every enabled MCP server injects its tool schema into every turn — heavyweight servers can cost 20k+ tokens each.

This is a harness setting, not a GSD setting. The toggle lives in .claude/settings.json:

{
  "enabledMcpjsonServers": ["context7"],
  "disabledMcpjsonServers": ["playwright", "mac-tools"]
}

Quick audit before a long phase:

  • Are any browser / playwright tools enabled when this phase has no UI work?
  • Are any platform-specific tools enabled when not needed?
  • Are any project-specific MCPs from a different project still enabled here?

Each disabled server removes its schema from every subsequent turn. Trimming MCPs compounds with model_profile tuning — both levers are additive, and MCP savings show up immediately across every subagent the orchestrator spawns.

For the full audit, harness reference, and the composition note with model_profile, see MCP Tool Schema Cost in the bundled context-budget.md reference.

Using Non-Claude Runtimes (Codex, OpenCode, Antigravity CLI, Kilo)

Codex CLI minimum supported version: 0.130.0 (issue #3562).

If you installed GSD for a non-Claude runtime, the installer already configured model resolution. No manual setup is needed — resolve_model_ids: "omit" is set automatically, which tells GSD to skip Anthropic model ID resolution and let the runtime choose its own default model.

To assign different models on a non-Claude runtime:

{
  "resolve_model_ids": "omit",
  "model_overrides": {
    "gsd-planner": "o3",
    "gsd-executor": "o4-mini",
    "gsd-debugger": "o3"
  }
}

Codex skill picker and agent scheduling (#774)

GSD enriches each Codex install with an additional artifact:

  • Flex-tier scheduling — light-tier agents (haiku-equivalent) emit service_tier = "flex" and model_verbosity = "low" in their agent TOML. The Codex scheduler routes these agents to the flex tier (lower cost, background processing) and suppresses verbose token output.

GSD skills appear in the Codex /skills picker via their SKILL.md file, which Codex discovers automatically. No agents/openai.yaml sidecar is emitted — doing so caused duplicate autocomplete entries (#1326).

This enrichment is written automatically at install time and requires no manual configuration. Requires Codex CLI ≥ 0.130.0.

Switching from Claude to Codex with one config change (#2517)

{
  "runtime": "codex",
  "model_profile": "balanced"
}

See Runtime-Aware Profiles.

Per-runtime command enrichment

When generating artifacts, the installer adapts GSD commands to each runtime's native command schema:

  • Qwen Code — main-loop skills carry Qwen's numeric priority field so the most-used workflows (e.g. new-project, plan-phase, execute-phase) sort first in the /skills list; utility skills are left unset. Higher values sort earlier; the field affects only the /skills list order.

See How to install GSD Core on your runtime for the full per-runtime details.

Manual install / no-Node.js setup

If you cannot run the GSD installer, you cannot use the source files in agents/ directly — they are in Claude Code's native frontmatter format. For OpenCode, two transformations are required:

Field GSD source format OpenCode-valid format Action
tools: Read, Bash, Grep (comma-string) Not a frontmatter field Remove the tools: line entirely
color: Plain CSS color name Hex or OpenCode semantic name Convert to hex or remove

Alternative: run the installer on any machine with Node.js:

npx @opengsd/gsd-core@latest --opencode --global

Installing for Cline

npx @opengsd/gsd-core --cline --global   # applies to all projects
npx @opengsd/gsd-core --cline --local    # this project only

Installing for CodeBuddy

npx @opengsd/gsd-core --codebuddy --global

GSD installs four surfaces for CodeBuddy: /gsd-* slash commands in ~/.codebuddy/commands/, subagents in ~/.codebuddy/agents/, model-invocable skills in ~/.codebuddy/skills/, and settings.json hooks. The skills are emitted with user-invocable: false so the slash commands are the single / menu surface (no duplicate entries).

Installing for Qwen Code

npx @opengsd/gsd-core --qwen --global

Installing for Prerelease Editions

Set the runtime's *_CONFIG_DIR env var to the prerelease directory before running the installer:

WINDSURF_CONFIG_DIR=~/.codeium/windsurf-next npx @opengsd/gsd-core@latest --windsurf --global

Env-var reference for supported runtimes:

Runtime Stable default Override env var
Claude Code ~/.claude CLAUDE_CONFIG_DIR
OpenCode XDG_CONFIG_HOME/opencode OPENCODE_CONFIG_DIR
Codex (per Codex CLI) --config-dir flag
Copilot ~/.copilot COPILOT_CONFIG_DIR (or COPILOT_HOME)
Cursor ~/.cursor CURSOR_CONFIG_DIR
Windsurf / Devin Desktop ~/.codeium/windsurf WINDSURF_CONFIG_DIR
Antigravity auto-detected ANTIGRAVITY_CONFIG_DIR
Augment ~/.augment AUGMENT_CONFIG_DIR
Trae ~/.trae TRAE_CONFIG_DIR
Qwen Code ~/.qwen QWEN_CONFIG_DIR
Kilo ~/.config/kilo KILO_CONFIG_DIR
CodeBuddy ~/.codebuddy CODEBUDDY_CONFIG_DIR
Cline ~/.cline CLINE_CONFIG_DIR

Using Claude Code with Non-Anthropic Providers

Switch to the inherit profile: /gsd-config --profile inherit. This makes all agents use your current session model.

Working on a Sensitive/Private Project

Set commit_docs: false during /gsd-new-project or via /gsd-settings. Add .planning/ to your .gitignore.

GSD Update Overwrote My Local Changes

Which recovery you need depends on whether you modified a GSD file or added your own:

  • You edited a file GSD ships (an agent prompt, a workflow). Since v1.17 the installer backs it up to gsd-local-patches/. Run /gsd-update --reapply to merge your changes back.

  • You added your own file inside a GSD-managed directory (a custom skill under skills/, an extra file in commands/gsd/). The installer saves it to gsd-user-files-backup/, and the update offers to restore it once the new version is installed. If you declined, or the backup is left over from an older update, restore it any time:

    node <config-dir>/gsd-core/bin/gsd-tools.cjs restore-custom-files --config-dir <config-dir> --apply
    

    Run it without --apply first to see what would be restored. The backup is never deleted, and the restore skips any file that would overwrite something the new release ships.

Install or Refresh a Release Candidate

To install or refresh GSD from the @next RC dist-tag (the pre-release channel established by ADR #660), run:

/gsd-update --next
# or equivalently:
/gsd-update --rc

The same scope/runtime detection, changelog preview, custom-file backup, and cache clearing apply. Omitting --next/--rc keeps targeting @latest (stable channel, no change). Only the @latest and @next channels are supported — no arbitrary dist-tag can be passed.

Cannot Update via npm

See docs/manual-update.md for a step-by-step manual update procedure.

Workflow Diagnostics (/gsd-forensics)

When a workflow fails in a non-obvious way, run /gsd-forensics to generate a diagnostic report covering git history anomalies, artifact integrity, and state inconsistencies. Output goes to .planning/forensics/.

Pre-populated Permissions (Claude Code)

Since v1.3.1, the installer pre-populates ~/.claude/settings.json (or settings.local.json for local installs) with the core permissions GSD needs:

{
  "permissions": {
    "allow": [
      "Bash(npx gsd-core *)",
      "Read(.planning/*)",
      "Edit(.planning/*)",
      "Read(STATE.md)",
      "Edit(STATE.md)"
    ]
  }
}

These entries eliminate first-run approval prompts for GSD's own tool calls. The merge is non-destructive — your existing permissions are preserved and GSD entries are only appended. Uninstalling GSD removes exactly these entries and preserves any others.

Secret-file protection moved from deny rules to a hook (#4221). Earlier versions also wrote three permissions.deny rules — Read(.env), Read(.env.*) and Read(.secrets). Claude Code 2.1.259 hardened the Bash-side enforcement of Read() deny rules so that any such rule makes every cd DIR && grep … / cd DIR && cat … compound prompt for approval, even in auto mode — and GSD's subagents emit hundreds of those per session. The same protection now ships as the always-on gsd-secret-read-guard.js PreToolUse hook (Read, Grep and Bash; see Runtime Hooks above for what it covers and its documented gaps). A hook denial is not a permission rule, so it never arms that check, and it applies in auto and bypassPermissions modes alike. On install and uninstall the three retired strings are removed from permissions.deny (and an emptied deny array is dropped). Note the removal is byte-exact: a rule you wrote by hand that is identical to one of the three is indistinguishable from the installer's and is removed as well — re-add it if you want both layers.

Executor Subagent Gets "Permission denied" on Bash Commands

Add the required patterns to ~/.claude/settings.json. Core patterns needed for all stacks:

"Bash(git add:*)",
"Bash(git commit:*)",
"Bash(git merge:*)",
"Bash(git worktree:*)",
"Bash(git rebase:*)",
"Bash(git reset:*)",
"Bash(git checkout:*)",
"Bash(git switch:*)",
"Bash(git restore:*)",
"Bash(git stash:*)",
"Bash(git rm:*)",
"Bash(git mv:*)",
"Bash(git fetch:*)",
"Bash(git cherry-pick:*)",
"Bash(git apply:*)",
"Bash(gh:*)"

Per-project permissions: add the same permissions.allow block to .claude/settings.local.json in your project root instead of ~/.claude/settings.json.

Parallel Execution Causes Build Lock Errors

GSD handles this automatically since v1.26. If you're on an older version, add to your project's CLAUDE.md:

## Git Commit Rules for Agents
All subagent/executor commits MUST use `--no-verify`.

To disable parallel execution entirely: /gsd-settings → set parallelization.enabled to false.


Recovery Quick Reference

Problem Solution
Lost context / new session /gsd-resume-work or /gsd-progress
Phase went wrong git revert the phase commits, then re-plan
Need to change scope /gsd-phase (default), /gsd-phase --insert, or /gsd-phase --remove
Something broke /gsd-debug "description" (add --diagnose for analysis without fixes)
STATE.md out of sync state validate then state sync — check the report's scope field, not just valid (interpret results)
Workflow state seems corrupted /gsd-forensics
Quick targeted fix /gsd-quick
Plan doesn't match your vision /gsd-discuss-phase [N] then re-plan
Costs running high /gsd-config --profile budget and /gsd-settings to toggle agents off
Update broke local changes /gsd-update --reapply
Custom file gone after an update gsd-tools restore-custom-files --config-dir <dir> --apply
Want session summary for stakeholder /gsd-pause-work --report
Don't know what step is next /gsd-progress --next
Parallel execution build errors Update GSD or set parallelization.enabled: false

Project File Structure

.planning/
  PROJECT.md              # Project vision and context (always loaded)
  REQUIREMENTS.md         # Scoped v1/v2 requirements with IDs
  ROADMAP.md              # Phase breakdown with status tracking
  STATE.md                # Decisions, blockers, session memory
  config.json             # Workflow configuration
  MILESTONES.md           # Completed milestone archive
  HANDOFF.json            # Structured session handoff (from /gsd-pause-work)
  research/               # Domain research from /gsd-new-project
  reports/                # Session reports (from /gsd-pause-work --report)
  todos/
    pending/              # Captured ideas awaiting work
    completed/             # Completed todos
  debug/                  # Active debug sessions
    resolved/             # Archived debug sessions
  spikes/                 # Feasibility experiments (from /gsd-spike)
    NNN-name/             # Experiment code + README with verdict
    MANIFEST.md           # Index of all spikes
  sketches/               # HTML mockups (from /gsd-sketch)
    NNN-name/             # index.html (2-3 variants) + README
    themes/
      default.css         # Shared CSS variables for all sketches
    MANIFEST.md           # Index of all sketches with winners
  codebase/               # Brownfield codebase mapping (from /gsd-map-codebase or /gsd-onboard)
  onboarding/             # Brownfield onboarding summary (from /gsd-onboard)
  phases/
    XX-phase-name/
      XX-YY-PLAN.md       # Atomic execution plans
      XX-YY-SUMMARY.md    # Execution outcomes and decisions
      CONTEXT.md          # Your implementation preferences
      RESEARCH.md         # Ecosystem research findings
      VERIFICATION.md     # Post-execution verification results
      XX-UI-SPEC.md       # UI design contract (from /gsd-ui-phase)
      XX-UI-REVIEW.md     # Visual audit scores (from /gsd-ui-review)
  ui-reviews/             # Screenshots from /gsd-ui-review (gitignored)