* test(01-01): define reviewer-support trait contract Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02): validator rejects non-boolean values with an exact field path, accepts missing/true/false, and the real code-review capability.json steps must declare supportsReviewerLanes: true. Add loop-resolver projection coverage proving the trait reaches activeHooks verbatim for a provider-neutral synthetic step (not code-review-specific), and that omitted/false values stay inert (no key on the active hook). All 8 new assertions fail today: the validator has no such field, and loop-resolver has nothing to project. RED before GREEN. * feat(01-01): declare reviewer-capable steps Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional boolean opt-in trait, step-scoped (not capability-wide). Only a literal true validates and projects; false/omitted stay inert (no key on the projected active hook), and every non-boolean type fails capability-validator.cjs with an exact field-path error. Opt both existing code-review steps (execute:post, execute:wave:post) into the trait in capabilities/code-review/capability.json. Project the validated field through src/loop-resolver.cts into activeHooks so a provider-neutral generic interpreter can read it without any code-review-specific knowledge. Document the field in docs/reference/capability-manifest.md and regenerate gsd-core/bin/lib/capability-registry.cjs via the generator (never hand-edited). Makes all 8 RED assertions from the prior commit pass. * test(01-02): define shared reviewer dispatch - Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes: inert when the supportsReviewerLanes trait is off or nothing is selected, exactly-once plan/invoke per selected lane, duplicate-alias dedup, the bounded metadata-only source-review prompt (repo root, paths+baseSha, depth, four fixed prohibitions), and capability-neutral reuse via a second synthetic step context. - RED: module under test (src/reviewer-step-dispatch.cts) does not exist yet, so require() fails and every assertion is unreached. * feat(01-02): dispatch reviewers for opted-in steps - Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps), ONE interpreter for a step's supportsReviewerLanes trait. Reuses resolveReviewerSelection for selection and resolveLanePlan for planning (both already-existing, pure building blocks); invocation is the one required, caller-injected seam (deps.invoke) since runLane needs OS-aware spawn plumbing this module does not own. - trait !== true, or a selection resolving to zero lanes, dispatches nothing (zero plan/invoke calls). Each selected lane is planned and invoked exactly once, in the selector's deduped/sorted order. - buildSourceReviewPrompt assembles a metadata-only bounded prompt (repo root, canonical paths + base SHA, depth, four fixed prohibitions) — never file contents — written once per dispatch and shared across every invoked lane. - GREEN: tests/reviewer-step-dispatch.test.cjs now passes. * test(01-02): define reviewer dispatch failures - Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed matrix: an explicitly requested lane the selector could not resolve still lets the OTHER resolved lane run, but the aggregate result must never read as a clean success (and 'every explicit lane unavailable' must be distinguishable from the plain no-flags-passed inert case); request-level validation (path traversal, absolute paths outside repoRoot, empty/non-string paths, missing depth/base SHA) halts the whole dispatch before any lane is planned or invoked; a per-lane prompt-budget overflow hard-fails only that lane before invoke while its sibling still runs. - RED: src/reviewer-step-dispatch.cts does not yet implement any of these guards, so 9 of the new assertions fail against the current (Task 1) implementation. * fix(01-02): fail closed in reviewer dispatch - src/reviewer-step-dispatch.cts: add the fail-closed guards the prior commit deliberately left out. An explicitly requested lane the selector could not resolve no longer lets the aggregate read as a clean success — lanes that DID resolve still run and keep their results (never narrow the requested set), but selection.errors now flips the aggregate ok to false, and 'every explicit lane unavailable' is now distinguishable (SELECTION_FAILED) from the plain no-flags-passed inert case (NO_LANES_SELECTED). - Add request-level validation (validatePaths, depth/baseSha presence) that halts the WHOLE dispatch before any lane is planned or invoked: path traversal, absolute paths outside repoRoot, empty/non-string paths, and missing provenance are all rejected up front. - Add per-lane prompt-budget enforcement (resolveBudget, mirroring gsd-tools.cjs's budgetFor convention including budget 0 = unbounded): a lane whose resolved budget the prompt exceeds hard-fails before invoke runs for it, without cancelling a sibling lane already planned. - Document the supportsReviewerLanes trait and its dispatch-step interpreter in gsd-core/references/loop-hook-dispatch.md. - GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass; no regressions in the review-lane/reviewer-selection/prompt-budget suites (356 passing). * test(01-03): define optional source reviewer flow RED: assert code-review.md dispatches roster-derived reviewer-lane flags through a single review-lane dispatch-step call (DISP-01..05), that the no-flag path stays byte-for-behavior unchanged (COMP-01), and that external evidence reaching the internal reviewer prompt is marked unverified (CONS-02). Also covers the CLI contract directly: no-op with no explicit selection, and fail-closed on an explicit unknown lane (SAFE-07) via real gsd-tools.cjs subprocess calls. * feat(01-03): route optional source reviewers GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches canonical reviewer-lane flags against the merged first-party + installed roster (never a hand-maintained list) and, only when at least one is present, calls the shared reviewer-step interpreter exactly once with the already-resolved repo root, file scope, depth, and base SHA. Its evidence paths are appended to the internal reviewer prompt via ${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane flag leaves the internal-only dispatch byte-for-behavior unchanged (COMP-01). Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI route `dispatchReviewerLanes` wires through, but never implemented the gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add it to the existing review-lane router, reusing the same effort-aware plan building and runner deps `plan`/`invoke` already use (factored into buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard the CLI's own `detected` set on whether an explicit flag was passed: resolveReviewerSelection's no-explicit-selection fallback is "select every detected reviewer" (the correct default for /gsd:review), and passing it an unconditionally non-empty detected set would silently invoke the whole roster on every no-flag code review, violating COMP-01. * test(01-03): define external finding consolidation RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as untrusted input — independently re-verifies every claim against the actual current source, resists a prompt-injection attempt embedded in evidence text, and folds a verified claim into the existing Narrative Findings section with no second REVIEW.md schema (CONS-01..03). Also assert code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * feat(01-03): consolidate external review evidence GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence> as untrusted data, independently re-verifies every cited claim against the actual current source before it can appear in REVIEW.md, and explicitly resists prompt injection embedded in evidence text (never a command, no matter what it claims to be). A verified claim folds into the existing Narrative Findings section with (external: {slug}) provenance — one REVIEW.md schema only, no separate external-findings section. code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * fix(01-02): gitignore the reviewer-step-dispatch build artifact 01-02 added src/reviewer-step-dispatch.cts but never added its npm run build:lib output to .gitignore, unlike every sibling gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked noise in git status. * docs(01-04): publish user and command contract for reviewer-lane source review - Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on failure, findings independently consolidated into the single REVIEW.md - Add the same contract to the docs/features/code-review-pipeline.md fragment and regenerate docs/FEATURES.md from it - Preserve /gsd-review as the plan-review command; cross-reference it rather than duplicating the reviewer roster - Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md drift owned by source already shipped in Plans 01-01/01-03 but never regenerated (npm run regen:derived had not been run in this worktree) * docs(01-04): align architecture and agent ownership docs for reviewer-lane trait - ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes) through the shared dispatchReviewerLanes interpreter to the existing review-lane plan/invoke machinery, ending at gsd-code-reviewer as the sole REVIEW.md consolidator - AGENTS.md: document gsd-code-reviewer's full-context verification scope and its treatment of external reviewer evidence as unverified input - No new diagram, abstraction, or config key; docs/CONFIGURATION.md is unchanged since the feature adds no setting or default * fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact Same gap as the earlier .gitignore fix: 01-02 added src/reviewer-step-dispatch.cts but never added its generated gsd-core/bin/lib/reviewer-step-dispatch.cjs output to eslint.config.mjs's ignore list like every sibling generated file, so tsc's emitted __importDefault CommonJS-interop var tripped no-var. * fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md 01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster row in docs/INVENTORY.md — required by design, since a role sentence cannot be generated — was never added. * fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added supportsReviewerLanes: true to that step and this fixture was not updated. * chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md Both files grew as a direct, intended consequence of wiring optional reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes step and the untrusted-evidence consolidation contract) — not incidental drift. Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209) Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209) * test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes From internal code review: dispatched must be false when zero lanes actually reached plan(), and a throwing plan()/invoke() for one lane must not discard results already collected for a sibling lane — matching the fail-closed pattern gsd-tools.cjs already uses for the same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794). Refs: gsd-core-dks.16, gsd-core-dks.17 * fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review - WR-01: dispatched now tracks whether any lane actually reached plan(), not results.length — an unresolvable selected slug no longer reports dispatched:true. - WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a throw for one lane can never discard results already collected for a sibling lane, matching the same guard gsd-tools.cjs already has around the identical resolveLanePlan call. - IN-01: documents the intentional budget===0-is-unbounded convention (#2797) the caller already relies on. - IN-02: review-lane dispatch-step no longer blocks indefinitely on an un-piped interactive TTY; fails closed to empty paths instead. Refs: gsd-core-dks.16, gsd-core-dks.17 * docs(01-05): add changeset fragment for PR #17 * fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract agents/gsd-code-reviewer.md's untrusted-evidence section and its pinning regression test both quote injection phrases as the exact attack they defend against/detect — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing allowlist entries, not an actual injection vector. * test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip From CodeRabbit review: WR-02's earlier fix only wrapped plan() — writePromptFile()/deps.invoke() still ran unguarded, so a throw there still aborted every later selected lane. Also covers the dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit lanes silently not running when no prior review and no phase-start commit exist). * fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved Previously an explicit reviewer-lane request with no prior review and no resolvable phase-start commit reached dispatch-step with an empty --base-sha, which fails closed via missing_provenance — correct, but silent about why explicitly requested lanes didn't run. Now skip dispatch entirely in that case with a stderr warning naming the actual cause. * fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan() WR-02's original fix only guarded plan() — a throw from writePromptFile() or deps.invoke() still aborted the whole dispatch, discarding results already collected for lanes processed earlier in the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it. * fix(01-05): WR-02b mock must throw only on the first writePromptFile() call The committed mock threw unconditionally, so codex's retry also threw and failed for the same reason as claude's — the test could not distinguish 'sibling still runs' from 'sibling also breaks'. Gate the throw to the first call, matching WR-02/WR-02c's single-failure intent. * fix(#4209): close review findings from adversarial + critical-code-reviewer pass Two independent reviews (agy adversarial review, Opus critical-code-reviewer + ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and fixed here: - dispatch-step's reducer silently swallowed whole-dispatch rejections (invalid paths, missing provenance, etc); it now checks parsed.ok/reason. - spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review; now shares the single compute_file_scope derivation. - the external reviewer prompt had no actual review request or citation requirement, only prohibitions; added both. - removed the supportsReviewerLanes trait plumbing (capability registry, validator, loop-resolver, docs, tests) — it was never consulted by the real dispatch path, which gates on explicit CLI flags instead. - flag-resolution require() was a fragile cwd-relative literal that failed silently on non-vendored installs; now resolves via GSD_TOOLS's own directory and warns instead of swallowing failure. - reducer didn't unwrap the @file: overflow protocol for large payloads. - deduplicated resolveBudget/budgetFor into one resolveLaneBudget. - lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a second dispatch can't overwrite prior evidence. - validatePaths rejects control characters, closing a markdown-injection vector into the external prompt via crafted filenames. - reworded the one line that tripped prompt-injection-scan.sh instead of allowlisting the whole production prompt file. - fixed a stale docstring range and a dispatched-field ordering bug. - added 3 integration tests executing the actual reducer against synthetic dispatch-step JSON, replacing markdown-substring-only assertions. 771/771 tests pass across every touched suite; tsc --noEmit clean. * fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait The maintainer's approval on issue #4209 explicitly redirected implementation shape: reviewer-lane dispatch must be a reusable capability/step-dispatch trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call itself. My previous commit (e2558326) deleted that trait entirely after finding it declared-but-never-consulted, which was backwards — the fix was to wire it, not remove it. Restores the trait (capability.json, generated registry, validator, loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes now resolves its own active hook via `gsd_run loop render-hooks` for the configured workflow.code_review_point and only proceeds to CLI-flag matching when supportsReviewerLanes reads true. Explicit flags no longer bypass the trait; a matching flag with the trait false resolves zero slugs (proven by a new integration test executing the real fence with both trait states). Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer redirect requires the capability layer, not the workflow, own the opt-in decision). * fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point Both an agy adversarial review and an Opus critical-code-reviewer pass independently found the same gap in my previous commit (9b2c3773d): the trait check I wired into code-review.md only protected code-review's OWN invocation — gsd-tools.cjs's dispatch-step handler still hardcoded `trait: true` unconditionally, so a second capability declaring supportsReviewerLanes would get zero enforcement from the shared CLI unless it correctly re-implemented the ~15-line render-hooks scrape itself. That is exactly the "each workflow.md hand-wiring the call" the maintainer's redirect said to eliminate. Moves the trait check into dispatch-step itself: given --cap-id/--point, it self-invokes `loop render-hooks <point>` (relocating the one subprocess code-review.md used to spawn for this, not adding a new one) and derives the real trait from that capId's active hook, rather than trusting a caller-passed boolean. code-review.md now only passes --cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or gates on the trait itself — the ~20-line scrape it previously carried is gone. Any other capability opts into the identical enforcement by declaring the trait and passing the same two flags. Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input variable (they proved a bash branch honors a variable, not that the variable reflects the real capability manifest) with three integration tests that invoke the real dispatch-step CLI against the real first-party capability registry: the real code-review trait resolves true, an unknown --cap-id resolves false (trait_not_enabled, fail-closed), and omitting --cap-id/--point entirely resolves false (no context means no opt-in). Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check (agy-F1 was incomplete), and delete the promptWritten per-lane coupling flag — the prompt write is idempotent, so writing it once per lane instead of gating on "did any lane write it yet" removes a latent bug where a deps.plan override that ever varies promptPath per lane would silently skip writing for a later lane. Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count drops (the trait scrape moved into dispatch-step), but the file still grew this session across multiple commits; acknowledging per the growth-tracking convention. * fix(#4209): remove per-run token waste from the shipped prompts Runtime prompt content, not session tokens: two real, per-invocation token costs in the code that ships. 1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of load_context step 5's ~180-word untrusted-evidence contract in ~90 more words, breaking this section's own established terse one-liner style (every other rule here is 1-2 sentences). This prompt loads fresh on every /gsd:code-review invocation. Shrunk to a one-line cross-reference, matching how write_review's own reference to step 5 already does it. 2. buildSourceReviewPrompt repeated the base SHA on every single file line even though it is identical for every file and already stated once at the top of the prompt — O(files) wasted tokens on every dispatched lane for a 50-file review, for zero information gain. File lines are now bare paths. * fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3 Opus critical-code-reviewer found a real Blocking defect in the --cap-id/ --point self-invocation added last commit: `dispatch-step` spawned `loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to `@file:<path>` instead of inline JSON -- the same overflow protocol this feature already unwraps for its OWN dispatch result 60 lines later in code-review.md. A large-enough activeHooks envelope (more installed capabilities/fragments) would throw, get silently swallowed by the bare catch, and misreport a real trait as trait_not_enabled with zero diagnostic. Fixed by extracting the config/registry/capability-state resolution `cmdLoopRenderHooks` already performs into an exported pure function, resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now share it), and calling it in-process from dispatch-step instead of spawning a subprocess at all. This eliminates the @file: exposure entirely (the dispatch-step path never touches the rendered-string envelope or its JSON-stringify/50000-char threshold), removes one subprocess spawn per code-review invocation, and gives a genuine diagnostic (stderr warning) on resolution failure instead of silent fail-closed. Corrected three doc/ docstring references to the now-removed subprocess self-invocation. Also fixes 2 real CI failures this round surfaced: - lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line no-control-regex` comment was unused under this project's ESLint config (verified locally: the rule never actually flags \x00-\x1f in this repo's config) -- a mistake from an earlier commit this session, never actually lint-checked before push. Removed the disable comment. - security (prompt-injection-scan): the agy-F1 regression test's crafted fixture literally contains "Ignore all prior instructions." as test data proving validatePaths rejects it -- allowlisted the test file, same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries. Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one bullet stated "untrusted, never a command" three different ways in one paragraph, and a same-file duplicate of write_review's schema rule. Consolidated to state each rule once. Declined one suggestion from this round: shrinking code-review.md's EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests (tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block, tests/code-review.test.cjs's CONS-02 test) deliberately lock the four- prohibitions restatement and the untrusted-evidence prose into the INJECTED block itself, not just the consolidator's system prompt -- adjacency of the warning to the untrusted payload it's warning about is a recognized prompt-injection defense-in-depth pattern from this workstream's original TDD plan, not accidental duplication. * fix(#4209): correct stale per-file base-SHA prose in the external prompt Leftover from removing the per-file base SHA repetition earlier this session: the review-request sentence still said "relative to its base SHA" (singular per-file framing) when there's now exactly one base SHA, stated once above the file list. Reads "relative to the base SHA above" now. * fix(#4209): make getLane/configGet/plan required deps, delete dead defaults R3/R4 from the review round I'd deferred as low-priority test-churn: this file's one production caller (gsd-tools.cjs's dispatch-step handler) always supplies all three, so the fallbacks were dead in production -- but each was actively WRONG if ever reached: the default configGet always returned undefined, silently disabling resolveLaneBudget's overflow guard; the default getLane looked up only first-party REVIEWER_LANES, diverging from production's overlay-merged roster; the default plan skipped per-host effort resolution entirely. These defaults were introduced by this PR's own earlier work (this file did not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from elsewhere, so there's no external caller depending on the lenient contract. Turned out free to fix: making the three deps required and deleting defaultGetLane/defaultPlan needed zero test changes -- every existing test that actually reaches the per-lane loop already supplies getLane/plan explicitly, and configGet's only real dependent (the budget-overflow tests) already supplies it too. 788/788 tests pass unchanged, tsc/lint clean. * fix(#4209): define depth semantics for the external reviewer lane Verified this was a real bug, not a match to existing convention as I'd claimed when declining the suggestion earlier this session: the internal gsd-code-reviewer agent's own system prompt carries a full <depth_levels> block defining what quick/standard/deep mean and do (agents/gsd-code- reviewer.md:68-99). The external reviewer lane has no access to that persona at all -- it only ever sees buildSourceReviewPrompt's bounded text, which sent the bare depth label with zero definition to a third-party CLI with no other source of truth for what "standard" means. Added depthMeaning(), condensed from the internal reviewer's own <depth_levels> definitions so the two stay consistent, and interpolated it into the review-request sentence. 150/150 tests pass, tsc/lint clean. * fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide whether to dispatch at all. This file's own documented rule (its depth-resolution guard, stated explicitly a few hundred lines earlier) is that a guard and the extraction it protects must run as one shell control-flow decision, because markdown-fenced blocks do not share shell state -- this step violated its own file's rule for the entire feature's gating condition. Merged the roster-resolution fence and the dispatch-decision fence into one continuous bash block, removing the intervening prose that split them. Fixed the stderr-based failure detection in the same edit (RQ-01: checking whether stderr is non-empty misfires on any benign Node warning; now checks the actual exit status of the roster-resolution command). Verified by extracting the merged fence and executing it standalone, driving both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty, SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests pass, tsc/lint clean. * fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write Batch of Required/Suggestion fixes from the Opus critical-code-reviewer + writing-for-agents pass: - CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch blocks, commented-out code) and deep (error propagation, state mutation consistency, circular dependencies) relative to the real <depth_levels> block, and had zero test coverage. Restored full accuracy and added tests that read the real agents/gsd-code-reviewer.md file directly, so drift between the two can't recur silently. Unrecognised depth now normalizes to standard's definition, matching that agent's own documented rule, instead of rendering an undefined bare label. - RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt `paths` does, but weren't checked for control characters like paths were (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and applied it to all four fields at the same provenance-check boundary. runDir previously had zero validation at all. - S1: deleted the dead `identity` parameter on `invoke` -- the one production caller already ignores it, no test read it by name. - S2: hoisted the shared prompt write above the per-lane loop -- promptPath is derived from runDir alone (constant across lanes by construction), so writing it once is both correct and cheaper than the per-lane write R1 introduced earlier this session. Discovered and fixed a real regression from the naive version of this hoist: an unguarded throw would have escaped dispatchReviewerLanes as an uncaught exception instead of a clean per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason, matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with a dedicated regression test. - S3: moved `planned = true` past the budget-overflow gate, so `dispatched` only reports true once a lane has cleared BOTH plan and budget checks. - S5: relayed gsd-code-reviewer.md's own "performance issues are out of scope unless also correctness issues" policy into the external-lane prompt, which previously had no such guidance and could return findings the internal reviewer's own contract excludes. - RQ-05 (partial): shrunk this file's own header docstring's restatement of the trait-reuse architecture to a pointer at gsd-core/references/loop-hook-dispatch.md, the canonical home. 234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke` already share. code-review.md's ~18-line inline `node -e` reimplementing `loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in gsd-tools.cjs) is now a single call to this subcommand -- the exact violation code-review-flags.cjs's own header warns against ("this is the canonical flag-parsing surface -- do not replicate inline bash parsing"). RQ-03: an empty --cap-id XOR --point now warns distinctly from the legitimate no-context opt-out (both absent) -- a caller that named a capability without its point was silently indistinguishable from a correct opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only ever fires when the config-get COMMAND ITSELF fails (config-get already resolves the manifest's own schema default in the normal case), but that failure was previously silent. RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait resolved inside dispatch-step" explanation was restated in full in 5 places across this session's own review cycles. Consolidated to ONE canonical statement in gsd-core/references/loop-hook-dispatch.md; the other 4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md, code-review.md's step-opening comment) now point at it instead. W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two inert cases when capability-validator.cjs already rejects non-boolean at load -- restated as the two cases that actually reach this code. Removed a "do not hand-roll trait resolution" prohibition whose target no longer exists once the positive description precedes it. W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing block means proceed as normal") -- an absent optional block already means proceed as normal without being told. W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the made-up compound "byte-for-behavior [un]changed" with the token this session's own docs already coined for this concept (inert) and the word that means what byte-for-behavior was reaching for (unchanged). W-10: dispatch_reviewer_lanes had no completion criterion -- added one sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set, either populated or empty). This exact sentence would have caught the cross-fence bug fixed two commits ago at authoring time. Declined from this round, with reasoning: W-02/W-03 (trim the untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) -- two tests deliberately lock this as intentional adjacency-based prompt-injection defense-in-depth, not accidental duplication (see this branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site `trap ... EXIT`) -- would fire at the end of the CREATING fence, before spawn_reviewer's agent ever reads the evidence files, given this file's own documented fenced-block execution model; the existing named cross-reference between creation and cleanup already satisfies the co-location concern without introducing that regression. 853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get fallback lived in an earlier, separate fence from the fence that consumes it via --point, split only by prose (not a guard, per this step's own documented rule). Merged into the single continuous fence and added a structural test asserting exactly one bash fence in the step. The new end-to-end regression test for this used --codex, which drives the fence's real `review-lane dispatch-step` call and, with the codex binary present on PATH, spawns the real external CLI — which then blocks on interactive auth with no stdin (BL-01). Stubbed gsd_run for `review-lane dispatch-step` only (captures argv instead of executing), keeping the real config-get/explicit-from-argv calls the test is actually about. * fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment Round-5 review (Opus) warning-tier findings: - WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but a control-character injection attempt" — a caller distinguishing a config problem from a security event couldn't tell them apart. Split into MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid). - WR-05: validatePaths' containment check was lexical only (path.resolve), so a symlink whose own path sits inside repoRoot could still point outside it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can legitimately name a file already deleted in a stale worktree), realpathing repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't false-positive-reject its own real children. - WR-08: a comment in the per-lane loop still said a throwing writePromptFile() was caught there — stale since the prompt write was hoisted above the loop in an earlier round. WR-03 (validate depth against the quick/standard/deep enum) was considered and declined: this dispatcher is deliberately capability-neutral (see the existing "synthetic step context" test, which passes a non-code-review depth label on purpose to prove no code-review-specific special-casing exists). WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and WR-07 (reason omitted on the aggregate return) were verified against source and are not bugs — see review notes. * docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap Round-5 review (Opus, BL-03) flagged that an early exit between dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A trap-based cleanup was considered and rejected: if a step genuinely runs as a separate process, a trap set at creation time would fire at the end of that SAME fence, deleting the directory before spawn_reviewer/commit_review ever read it — worse than the leak it would fix. review.md's own gather_context/cleanup pair for the identical resource class (a run-scoped reviewer temp dir) already makes and documents this exact trade-off: cleanup runs only on a documented success path, and a leftover $TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording that precedent here so this isn't re-raised as a live gap in a future review. * fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes paths: ['docs/spec.md'] as a synthetic, never-read path proving the dispatcher has no code-review-specific special-casing. lint-docs-guard- registration correctly flagged this as an unregistered docs/ path reference — add the docs-guard-exempt marker and its pinned baseline entry, the same pattern every other synthetic docs/ literal in this test suite already uses. * fix(#4209): backfill changeset pr: field with the real upstream PR number changeset-lint's fail_pr_field_drift caught the fragment still pointing at the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this branch is now also open against. * docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every decision in ADR-2782 (D1-D9) and every prior dated amendment governs the `role: "reviewer"` capability body and its one consumer, /gsd:review. This PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary feature capability's `steps[]` entry, projected through loop-resolver.cts and resolved in-process via resolveActiveHooksForPoint - is a different capability axis (steps/gates/contributions) that the ADR's own scope note explicitly places out of reach. Per docs/contributor-standards.md's "Amending an accepted ADR", an in-place dated section is the established, lighter-weight path for an addition that stays within the ADR's existing decisions - used twice already in this same file - so this appends a third dated entry documenting the new seam, its consumer, and why it reuses the existing D1-D9-governed plan/invoke machinery rather than adding a second one. No decision is reversed; no new Amends/Amended-by pair is needed since the steps/gates/contributions axis already carries reciprocal links to ADR-857 and ADR-894. * fix(#4209): close two test-quality gaps trek-e's review found Minor 1: validatePaths (a path-shape parser guarding the prompt- injection/path-traversal trust boundary) had only example-based coverage, violating ADR-456's rule that parsers/budget limits carry at least one fast-check property test. Adds three: safe-segment paths are never rejected, a single leading "../" always escapes the one-segment repoRoot, and a control character anywhere is always rejected - one property per rejection reason validatePaths owns. Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only ever exercised far below budget or at budget:0 (unbounded), never at the exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds three exact-boundary tests using the real estimateTokens/ buildSourceReviewPrompt the module calls internally, so the resolved token count is exact rather than approximated: budget == estimate (must pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must pass). Also extracts okPlan()'s fixture timeoutMs into a named constant - local/no-adhoc-timeout-literal (#4446) landed on next after this branch was authored and flagged the pre-existing literal on rebase; it is fixture data for a synthetic plan object dispatchReviewerLanes never waits on, a distinct class from tests/helpers/timeouts.cjs's real subprocess norms. * fix(#4209): update docs-guard-registration baseline for the new ADR citation reviewer-step-dispatch.test.cjs's new fast-check property tests cite docs/adr/456-test-rigor-architecture.md in a justifying comment (never a real read). lint-docs-guard-registration fingerprints every docs/ path string an exempted test file mentions and fails on drift so a human re-confirms the exemption still holds - re-confirmed, and the baseline is updated to match. * fix(#4209): point changeset pr: field at the fork PR for CI validation changeset-lint's fail_pr_field_drift check compares the fragment's pr: field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH), not a fixed target. Rehearsing this branch on fork PR davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but fails here. Backfill to 4323 happens again, as the last commit, immediately before the approved push to open-gsd#4323 - never leaving pr: 17 on the branch that ships upstream. * fix(#4209): reject promptChannel:none lanes from source-review dispatch CodeRabbit found a real scope mismatch: coderabbit's lane declares promptChannel: 'none' and reviews the working tree on its own terms, fed nothing (review.md:367). Silently dispatching it through dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope buildSourceReviewPrompt promises and let the lane review whatever it independently sees fit, violating this interpreter's own scoped, metadata-only contract. Reject before plan()/invoke(), same as an unresolved slug. * fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file CodeRabbit found the whole-file match on workflowContent would still pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts of this 1000+-line workflow, proving nothing about the actual evidence block's contract. Line-filtered via splitLines (not a bare-\n regex spanning readFileSync content) so this stays CRLF-portable and passes local/no-unbounded-quantifier and local/no-crlf-fragile-split. * fix(#4209): guard DISPATCH_JSON substitution and capture its stderr CodeRabbit found the dispatch-step command substitution unguarded: a non-zero exit could leave DISPATCH_JSON empty (or halt the step under errexit with no warning), and the downstream reducer would only ever report the generic unparseable_dispatch_output reason, discarding the command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/ EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface it in a warning on failure, and fall back to a parseable dispatch_ command_failed JSON stub so the reducer's existing reason-reporting path still fires. * docs(#4209): fix byte-for-behavior wording and missing colon, regenerate CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the established repo term for output-identical unchanged behavior) and a missing colon after the bold "Optional external reviewer lanes (#4209)" lead-in in docs/features/code-review-pipeline.md. Fixed in the two hand-authored sources (commands/gsd/code-review.md, docs/features/ code-review-pipeline.md) and regenerated the two derived projections (skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/ FEATURES.md via gen-features.cjs) so they stay in sync. * fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI) The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds double-quoted JSON keys inside a single-quoted shell literal. That extra quote density, inside an already quote-heavy ~8KB driver string, passed bash -n and the full local suite on Linux but broke Windows Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end to end (#4209 round 5)` failed on two Windows CI shards with `bash -c: unexpected EOF while looking for matching '''` — a Windows argv-to- command-line re-quoting edge case, reproducible on rerun, not a flake. Root-caused via gh api job logs plus a byte-identical local reconstruction of the test's own driver script. Fix: drop the fabricated stub. The downstream node -e reducer already falls back to reason `unparseable_dispatch_output` on any JSON.parse failure, so an empty/partial DISPATCH_JSON on command failure is still handled correctly, with zero new quoting risk. * revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI) Two materially different mechanisms for the same CodeRabbit Nitpick ("Trivial | Quick win") both broke Windows Git-Bash reproducibly: a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ... matching '''") and, after removing that, a plain `head -1 "$VAR"` inside a nested command substitution ("unexpected EOF ... matching '"'"). Both passed bash -n and the full local suite on Linux every time; both failed the SAME test deterministically on Windows CI. Two attempts at the same class of fix (nested-quote construction near this exact step) is the retry limit - reverting to the original, already-shipped, Windows-verified unguarded form rather than continuing to guess at a third quoting mechanism for a Trivial- severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for anyone attempting this again: the fix belongs outside this specific markdown-fence-driver test harness (e.g., a real .sh helper script) if it's worth doing at all. * fix(#4209): backfill changeset pr: field to the real upstream PR before push Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy changeset-lint's PR-number check while rehearsing there; this is the last commit before the approved push to the real upstream PR (open-gsd/gsd-core#4323), so the field points at that PR number again. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
60 KiB
GSD User Guide
A narrative companion guide to GSD Core — orient yourself here, then follow the links into the dedicated docs.
GSD Core's documentation is organised by Diataxis. Browse by goal: Tutorials · How-to guides · Reference · Explanation · Docs index
Table of Contents
- Slash-command forms
- Namespace routing primer
- Reading GSD's output
- Project lifecycle overview
- Workflow Diagrams
- UI Design Contract
- Spiking & Sketching
- Backlog & Threads
- Workstreams & Workspaces
- Security
- Usage Examples
- Troubleshooting
- Recovery Quick Reference
- Project File Structure
- Related
For driving GSD directly from a GitHub / Linear / Jira issue, see the Issue-driven orchestration guide — a recipe that maps tracker issues onto the workspace → discuss → plan → execute → verify → review → ship loop using existing GSD primitives.
Slash-command form
GSD ships the same set of skills to every supported runtime, using the hyphen slash-form spelling:
- Hyphen form —
/gsd-command-name— used by Claude Code, Copilot, OpenCode, Kilo, Cursor, Windsurf, Augment, Antigravity, and Trae.
The installer writes this form into the command directory of each runtime you target.
Namespace routing primer (gsd-ns-*, v1.40+)
Architecture
GSD ships six namespace router bundles (gsd-ns-workflow, gsd-ns-project, gsd-ns-review, gsd-ns-context, gsd-ns-ideate, gsd-ns-manage). On runtimes with non-recursive skill loaders, the installer emits these 6 routers as the only top-level skill entries; the ~61 concrete skills are nested under each router at <router>/skills/<name>/SKILL.md. This reduces the eager skill-listing overhead to ≈6 entries instead of ≈67.
Each router's body contains a routing table. When the model receives a request, it reads the router, identifies the relevant sub-skill by name, then opens skills/<name>/SKILL.md via a file-path Read. The concrete skill is fully available — it is not invocable by bare name through the Skill tool's top-level listing, but is reachable through the router.
The nested layout applies only to runtimes with confirmed non-recursive skill loaders: Cline, Qwen, Hermes, Augment, Trae. Claude's loader is also non-recursive, but #924 reverted it flat because the Skill tool hard-errors on unknown names rather than re-routing via the router. Antigravity's loader is also non-recursive, but #1614 moved it flat because agy scans only skills/<name>/SKILL.md — nested sub-skills were unreachable. Other recursive or unconfirmed loaders (Cursor, Codex, Copilot, Windsurf, CodeBuddy, OpenCode, Kilo) retain the flat layout unchanged.
| Namespace | Router bundle | Routes to |
|---|---|---|
| Phase pipeline | gsd-ns-workflow |
discuss / plan / execute / verify / phase / progress |
| Project lifecycle | gsd-ns-project |
milestones, audits, summary |
| Quality gates | gsd-ns-review |
code review, debug, audit, security, eval, ui |
| Codebase intelligence | gsd-ns-context |
map, graphify, docs, learnings |
| Exploration & capture | gsd-ns-ideate |
explore, sketch, spike, spec, capture |
| Management | gsd-ns-manage |
config, workspace, workstreams, thread, update, ship, inbox |
Slash commands are unaffected
On runtimes that install a commands surface (commands/gsd), slash commands such as /gsd-plan-phase continue to work directly — the nesting applies only to the Skill tool's top-level listing, not to the commands directory.
Migration note (breaking change on nesting runtimes)
On the seven nesting runtimes listed above, upgrading to v1.40 changes skill invocation behaviour:
- Before: each of the ~67 concrete
gsd-<name>skills appeared at the top level and was invocable by bare name through the Skill tool. - After: only the 6
gsd-ns-*router bundles appear at the top level. Concrete skills are reachable via the router's routing table and aRead skills/<name>/SKILL.mdcall. Direct bare-name invocation of concrete skills through the Skill tool's listing no longer works. - Slash commands unchanged:
/gsd-plan-phase,/gsd-discuss-phase, etc. still work directly where a commands surface is installed. - Upgrade prune: the installer's existing prune step removes the legacy top-level
gsd-<concrete>/skill directories on upgrade — no manual cleanup is needed.
Reading GSD's output
GSD marks its sections with Markdown, not with drawn borders. There are exactly three forms, and every workflow, checkpoint and report uses them:
| You see | It means |
|---|---|
### GSD ► {STAGE NAME} |
A major workflow transition — planning, executing a wave, verifying, completing |
### CHECKPOINT: {Type} |
GSD is waiting on you. The bolded **→ …** line at the bottom tells you what to type |
--- |
A break between two sections — most often before the ▶ Next Up block at the end of a completion |
Panels that used to be drawn with box characters are now a heading followed by their rows as ordinary lines. A checkpoint reads:
### CHECKPOINT: Verification Required
Progress: 5/8 tasks complete
Task: Responsive dashboard layout
How to verify:
1. Visit: http://localhost:3000/dashboard
---
**→ YOUR ACTION: Type "approved" or describe issues**
Why there are no drawn borders
Earlier releases framed stage banners between two 53-character runs
of ━, and checkpoints were a 62-column box drawn with ╔, ║ and ╚. Those
runs are ordinary text to whatever renders GSD's output. In a pane narrower than
the run, the rule wraps and the leftover glyphs land on a second line — so the
border comes apart from the heading it was framing, and the output looks broken
rather than merely narrow. Reported against the Codex desktop interface in
#3028.
A Markdown heading and a thematic break carry the same structure without committing to a width, so they read correctly in a narrow pane and a wide one.
This is unconditional: every runtime gets the same output. The alternative — a
capability flag that kept line-art for terminal-oriented runtimes — was weighed
and rejected, because it leaves two output conventions to keep in sync forever
for a gain that is aesthetic rather than structural. What a terminal loses is the
drawn frame; what it keeps is the title, the hierarchy and the break, none of
which depended on the frame. The reasoning is recorded in
gsd-core/references/ui-brand.md § Why this is unconditional, not per-runtime.
If you run GSD in a host where the heading form reads worse than the old boxes
did, that is worth reporting — it is the evidence that would justify the flag.
Single-cell tokens are unaffected and unchanged: the status symbols
(✓ ✗ ◆ ○ ⚠), the GSD ► prefix, and the ten-cell progress gauge
(Progress: ████████░░ 80%) are not runs and do not wrap.
The convention is specified in gsd-core/references/ui-brand.md and enforced
against all shipped content by tests/responsive-separators.test.cjs.
Project lifecycle overview
The core GSD loop is: discuss → plan → execute → verify → ship, repeated per phase. The full step-by-step walkthrough — including example outputs, what files get created, and all the flags in play — is in the dedicated tutorial.
See Your first project.
For onboarding an existing codebase before starting a new milestone, run /gsd-onboard or see Onboarding an existing codebase.
Relevant flags at a glance:
| Flag | Command | When to use |
|---|---|---|
--auto |
/gsd-new-project |
Skip interactive questions, ingest from a PRD file |
--research |
/gsd-quick |
Add a research agent to an ad-hoc task |
--validate |
/gsd-quick |
Add plan-checking and post-execution verification |
--chain |
/gsd-discuss-phase |
Auto-chain discuss → plan → execute without stopping |
--skip-research |
/gsd-plan-phase |
Skip research agents when the domain is already familiar |
--draft |
/gsd-ship |
Create a draft PR instead of a ready-for-review one |
For the full command reference with all flags, see docs/COMMANDS.md. For configuration options (model profiles, workflow agents, git branching), see docs/CONFIGURATION.md.
Workflow Diagrams
Full Project Lifecycle
┌──────────────────────────────────────────────────┐
│ NEW PROJECT │
│ /gsd-new-project │
│ Questions -> Research -> Requirements -> Roadmap│
└─────────────────────────┬────────────────────────┘
│
┌──────────────▼─────────────┐
│ FOR EACH PHASE: │
│ │
│ ┌────────────────────┐ │
│ │ /gsd-discuss-phase │ │ <- Lock in preferences
│ └──────────┬─────────┘ │
│ │ │
│ ┌──────────▼─────────┐ │
│ │ /gsd-ui-phase │ │ <- Design contract (frontend)
│ └──────────┬─────────┘ │
│ │ │
│ ┌──────────▼─────────┐ │
│ │ /gsd-plan-phase │ │ <- Research + Plan + Verify
│ └──────────┬─────────┘ │
│ │ │
│ ┌──────────▼─────────┐ │
│ │ /gsd-execute-phase │ │ <- Parallel execution
│ └──────────┬─────────┘ │
│ │ │
│ ┌──────────▼─────────┐ │
│ │ /gsd-verify-work │ │ <- Manual UAT
│ └──────────┬─────────┘ │
│ │ │
│ ┌──────────▼─────────┐ │
│ │ /gsd-ship │ │ <- Create PR (optional)
│ └──────────┬─────────┘ │
│ │ │
│ Next Phase?────────────┘
│ │ No
└─────────────┼──────────────┘
│
┌───────────────▼──────────────┐
│ /gsd-audit-milestone │
│ /gsd-complete-milestone │
└───────────────┬──────────────┘
│
Another milestone?
│ │
Yes No -> Done!
│
┌───────▼──────────────┐
│ /gsd-new-milestone │
└──────────────────────┘
Planning Agent Coordination
/gsd-plan-phase N
│
├── Phase Researcher (x4 parallel)
│ ├── Stack researcher
│ ├── Features researcher
│ ├── Architecture researcher
│ └── Pitfalls researcher
│ │
│ ┌──────▼──────┐
│ │ RESEARCH.md │
│ └──────┬──────┘
│ │
│ ┌──────▼──────┐
│ │ Planner │ <- Reads PROJECT.md, REQUIREMENTS.md,
│ │ │ CONTEXT.md, RESEARCH.md
│ └──────┬──────┘
│ │
│ ┌──────▼───────────┐ ┌────────┐
│ │ Plan Checker │────>│ PASS? │
│ └──────────────────┘ └───┬────┘
│ │
│ Yes │ No
│ │ │ │
│ │ └───┘ (loop, up to 3x)
│ │
│ ┌─────▼──────┐
│ │ PLAN files │
│ └────────────┘
└── Done
Validation Architecture (Nyquist Layer)
During plan-phase research, GSD maps automated test coverage to each phase requirement before any code is written. The researcher detects your existing test infrastructure, maps each requirement to a specific test command, and identifies any test scaffolding that must be created before implementation begins (Wave 0 tasks). The plan-checker enforces this as an 8th verification dimension: plans where tasks lack automated verify commands will not be approved.
Output: {phase}-VALIDATION.md — the feedback contract for the phase.
Disable: Set workflow.nyquist_validation: false in /gsd-settings for rapid prototyping phases where test infrastructure isn't the focus.
Retroactive Validation (/gsd-validate-phase)
For phases executed before Nyquist validation existed, or for existing codebases with only traditional test suites, retroactively audit and fill coverage gaps:
/gsd-validate-phase N
|
+-- Detect state (VALIDATION.md exists? SUMMARY.md exists?)
|
+-- Discover: scan implementation, map requirements to tests
|
+-- Analyze gaps: which requirements lack automated verification?
|
+-- Present gap plan for approval
|
+-- Spawn auditor: generate tests, run, debug (max 3 attempts)
|
+-- Update VALIDATION.md
|
+-- COMPLIANT -> all requirements have automated checks
+-- PARTIAL -> some gaps escalated to manual-only
The auditor never modifies implementation code — only test files and VALIDATION.md. If a test reveals an implementation bug, it's flagged as an escalation for you to address.
Assumptions Discussion Mode
By default, /gsd-discuss-phase asks open-ended questions about your implementation preferences. Assumptions mode inverts this: GSD reads your codebase first, surfaces structured assumptions about how it would build the phase, and asks only for corrections.
Enable: Set workflow.discuss_mode to 'assumptions' via /gsd-settings.
See docs/workflow-discuss-mode.md for the full discuss-mode reference.
Decision Coverage Gates
The discuss-phase captures implementation decisions in CONTEXT.md under a <decisions> block as numbered bullets (- **D-01:** …). Two gates ensure those decisions survive into plans and shipped code.
Plan-phase translation gate (blocking). After planning, GSD refuses to mark the phase planned until every trackable decision appears in at least one plan's scanned surfaces: front-matter must_haves/truths/objective, a ## must_haves/truths/tasks/objective heading, or an <objective>/<tasks>/<task>/<action>/<read_first>/<behavior>/<verify>/<acceptance_criteria>/<done> tag body.
Verify-phase validation gate (non-blocking). During verification, GSD searches plans, SUMMARY.md, modified files, and recent commit messages for each trackable decision. Misses are logged to VERIFICATION.md as a warning section; verification status is unchanged.
Opting a decision out. Move it under the ### Claude's Discretion heading inside <decisions>, or tag it: - **D-08 [informational]:** …, - **D-09 [folded]:** …, - **D-10 [deferred]:** ….
Disabling the gates. Set workflow.context_coverage_gate: false in .planning/config.json (or via /gsd-settings). Default is true.
Execution Wave Coordination
/gsd-execute-phase N
│
├── Analyze plan dependencies
│
├── Wave 1 (independent plans):
│ ├── Executor A (fresh 200K context) -> commit
│ └── Executor B (fresh 200K context) -> commit
│
├── Wave 2 (depends on Wave 1):
│ └── Executor C (fresh 200K context) -> commit
│
└── Verifier
├── Check codebase against phase goals
├── Test quality audit (disabled tests, circular patterns, assertion strength)
│
├── PASS -> VERIFICATION.md (success)
└── FAIL -> Issues logged for /gsd-verify-work
Isolated-run Recovery (fail-safe)
When a worktree-isolated run is rejected — the user declines to merge it, or the run over-reached the requested scope, or the orchestrator surfaces recovery guidance for a blocked plan — GSD halts safely and offers two options: (a) re-attempt in a fresh, narrowly-scoped worktree, or (b) inspect or discard the rejected worktree without merging. GSD never defaults recovery to editing the primary checkout (main). Any path that edits the primary checkout requires explicit, clearly-labeled confirmation from the user first. This behavior is unconditional and applies to both /gsd-execute-phase (worktree executor waves) and /gsd-quick (quick-mode isolated runs).
UI Design Contract
AI-generated frontends are visually inconsistent not because Claude Code is bad at UI but because no design contract existed before execution. /gsd-ui-phase locks the design contract before planning; /gsd-ui-review audits the result after execution.
For the full workflow, configuration, shadcn initialisation, and the registry safety gate, see Design a UI phase.
Quick reference:
| Command | Description |
|---|---|
/gsd-ui-phase [N] |
Generate UI-SPEC.md design contract for a frontend phase |
/gsd-ui-review [N] |
Retroactive 6-pillar visual audit of implemented UI |
| Setting | Default | Description |
|---|---|---|
workflow.ui_phase |
true |
Generate UI design contracts for frontend phases |
workflow.ui_safety_gate |
true |
plan-phase prompts to run /gsd-ui-phase for frontend phases |
Spiking & Sketching
Use /gsd-spike to validate technical feasibility before planning, and /gsd-sketch to explore visual direction before designing. Both store artifacts in .planning/ and integrate with the project-skills system via their wrap-up companions.
For the full workflow and flow diagram, see Spike and sketch.
Typical flow:
/gsd-spike "SSE vs WebSocket" # Validate the approach
/gsd-spike --wrap-up # Package learnings
/gsd-sketch "real-time feed UI" # Explore the design
/gsd-sketch --wrap-up # Package decisions
/gsd-discuss-phase N # Lock in preferences (now informed by spike + sketch)
/gsd-plan-phase N # Plan with confidence
Backlog & Threads
Backlog Parking Lot
Ideas that aren't ready for active planning go into the backlog using 999.x numbering, keeping them outside the active phase sequence.
/gsd-capture --backlog "GraphQL API layer" # Creates 999.1-graphql-api-layer/
/gsd-capture --backlog "Mobile responsive" # Creates 999.2-mobile-responsive/
Backlog items get full phase directories, so you can use /gsd-discuss-phase 999.1 to explore an idea further or /gsd-plan-phase 999.1 when it's ready. Backlog directories (and the 0-* pre-milestone directory some projects carry) are excluded from /gsd-progress, /gsd-stats, and phase listings for the current milestone — they stay out of the active phase sequence for counting purposes too, not just for planning.
Review and promote with /gsd-review-backlog — it shows all backlog items and lets you promote (move to active sequence), keep (leave in backlog), or remove (delete).
Seeds
Seeds are forward-looking ideas with trigger conditions. Unlike backlog items, seeds surface automatically when the right milestone arrives.
/gsd-capture --seed "Add real-time collab when WebSocket infra is in place"
/gsd-new-milestone scans all seeds and presents matches. Storage: .planning/seeds/SEED-NNN-slug.md
Once you've parked a few, audit them on demand instead of waiting for the next milestone to surface them:
/gsd-capture --list-seeds # Review every parked seed
/gsd-capture --list-seeds dormant # Narrow to one status
This is read-only — it renders an audit table (ID, status, scope, trigger, title) and a per-status summary, and never modifies a seed. Filter by dormant, active, or triggered when you only want to see seeds in one state.
Persistent Context Threads
Threads are lightweight cross-session knowledge stores for work that spans multiple sessions but doesn't belong to any specific phase.
/gsd-thread # List all threads
/gsd-thread fix-deploy-key-auth # Resume existing thread
/gsd-thread "Investigate TCP timeout" # Create new thread
Threads can be promoted to phases (/gsd-phase) or backlog items (/gsd-capture --backlog) when they mature. Storage: .planning/threads/{slug}.md
Workstreams & Workspaces
Workstreams and workspaces both provide isolation, but at different levels.
Workstreams share the same codebase and git history but isolate planning artifacts — lighter weight, good for working on multiple milestone areas concurrently. See Work in parallel with workstreams.
Workspaces create separate repo worktrees with their own .planning/ — heavier, for feature-branch or multi-repo isolation. See Isolate work with workspaces.
| Command | Purpose |
|---|---|
/gsd-workstreams create <name> |
Create a new workstream with isolated planning state |
/gsd-workstreams switch <name> |
Switch active context to a different workstream |
/gsd-workstreams list |
Show all workstreams and which is active |
/gsd-workstreams complete <name> |
Mark a workstream as done and archive its state |
# Workspace example — feature branch isolation
/gsd-workspace --new --name feature-b --repos .
cd ~/gsd-workspaces/feature-b
/gsd-new-project
/gsd-workspace --list
/gsd-workspace --remove feature-b
Security
Defense-in-Depth (v1.27)
GSD generates markdown files that become LLM system prompts. This means any user-controlled text flowing into planning artifacts is a potential indirect prompt injection vector. v1.27 introduced centralised security hardening:
Path Traversal Prevention: All user-supplied file paths (--text-file, --prd) are validated to resolve within the project directory. macOS /var → /private/var symlink resolution is handled.
Prompt Injection Detection: The security.cjs module scans for known injection patterns in user-supplied text before it enters planning artifacts.
Runtime Hooks:
gsd-prompt-guard.js— Scans Write/Edit calls to.planning/for injection patterns (always active, advisory-only)gsd-workflow-guard.js— Warns on file edits outside GSD workflow context (opt-in viahooks.workflow_guard)gsd-write-guard.js— Hard-blocks a whole-fileWritethat catastrophically shrinks a curated.planning/artifact (ROADMAP.md, milestone roadmaps,STATE.md) below 40% of its on-disk line count; files under 40 lines are exempt. The check is stateless per Write, comparing each payload against the file's current on-disk size — a single-shot collapse (the #973 shape) is blocked, but a sequence of individually-tolerated shrinks that erodes the file across several Writes is not detected. For a legitimate milestone reset or large deletion, bypass once with the single-use sentinel — write the target's path into.planning/.gsd-allow-shrink(fresh within 15 minutes; consumed by the allowed write) — or, interactively, withGSD_ALLOW_PLANNING_SHRINK=1in the runtime's environment. Scope the guarantee accordingly: this stops accidental and single-shot collapse, and is not a defense against a determined agent — the sentinel is a plain file, so anything with shell access can arm one; what it buys is that the bypass becomes a deliberate, path-bound, single-use and auditable action rather than a sentence to reason past (always active, blocking; #2255, fix 3 of #973)gsd-secret-read-guard.js— Hard-blocks reads of secret files —.env,.env.<suffix>and.secrets, matched case-insensitively (.ENV,.Secrets) — through Read (file_path), Grep (an explicitpath, or aglobthat selects them, judged per brace alternative) and Bash (operands, input redirects,$( )/ backtick /<( )bodies, andgit show <ref>:<path>shapes). A shell interpreter (bash/sh/zsh/dash/ksh) has its script scanned however it arrives —-c '…', a<( )file operand, a heredoc / here-string, or a pipe from a knowableecho/printfsource (echo cat .env | bash) — as doeval's joined operands, asource/.process-substitution operand, andfind … | xargs catpipelines (upstream literal names become the sub-command's read operands)..env.example/.env.sample/.env.template/.env.diststay readable (they are the templates GSD's own phase prompt reads — a real secret stored under one of those names is not protected), and existence checks ([ -f .env ],ls .env*,test,stat,rm,touch,echo, …) pass. Not covered, by construction:$VARindirection (bash -c "$CMD"), shell globs (cat .e*), interpreter one-liners, a piped script from a non-echo/printfsource (cat gen.sh | bash,curl … | sh), reads inside scripts the agent runs, and a Grepglob: '*'reaching a.envthat is not gitignored — none are statically resolvable by a hook. This replaces theRead(.env)/Read(.env.*)/Read(.secrets)permission deny rules the installer used to write: on Claude Code ≥ 2.1.259 anyRead()deny rule makes everycd DIR && grep …compound prompt for approval even inautomode, while a hook denial is not a permission rule and applies inautoandbypassPermissionsalike (always active, blocking; #4221)
CI Scanner: prompt-injection-scan.security.test.cjs scans all agent, workflow, and command files for embedded injection vectors.
Package Legitimacy Gate (v1.42.1)
AI coding tools hallucinate package names. Attackers pre-register those names on npm, PyPI, and crates.io with malicious post-install scripts — a technique called slopsquatting. v1.42.1 adds a three-layer gate that stops this before it reaches your shell.
In RESEARCH.md — every phase that recommends external packages includes a ## Package Legitimacy Audit table:
## Package Legitimacy Audit
| Package | Registry | Age | Downloads | Source Repo | Verdict | Disposition |
|---------|----------|-----|-----------|-------------|---------|-------------|
| express | npm | 13 yrs | 100M+/wk | github.com/expressjs/express | [OK] | Approved |
| some-new-util | npm | 3 days | 47 | none | [SLOP] | REMOVED |
| api-bridge | npm | 6 mo | 1.2k/wk | github.com/user/api-bridge | [SUS] | Flagged |
[SLOP] packages are removed from RESEARCH.md entirely and never reach the planner.
In PLAN.md — [SUS] or [ASSUMED] packages trigger a checkpoint:human-verify task before the install.
During execution — if an install fails, the executor surfaces a checkpoint and stops rather than silently trying an alternative.
Legitimacy verdicts:
| Verdict | Meaning | GSD action |
|---|---|---|
[OK] |
Passes all legitimacy checks | Proceeds — no checkpoint added |
[SUS] |
Suspicious signals | Flagged; planner adds checkpoint:human-verify |
[SLOP] |
High-confidence hallucination | Removed from RESEARCH.md; never reaches planner |
Verdicts are computed from live registry APIs (npm, PyPI, crates.io) — there
is no separate tool to install. slopcheck is an optional escalate-only
adapter (it can raise a verdict but never lower one); no shipped
configuration wires it, and its absence does not change the gate's behavior.
Code Review Workflow
After executing a phase, run a structured code review before UAT. See Set up cross-AI review for the full workflow.
/gsd-code-review 3 # Review all changed files in phase 3
/gsd-code-review 3 --depth=deep # Deep cross-file review
/gsd-code-review 3 --fix # Fix Critical + Warning findings atomically
/gsd-code-review 3 --fix --auto # Fix and re-review until clean (max 3 iterations)
/gsd-audit-fix # Audit + classify + fix (medium+ severity, max 5)
The review step slots in after execution and before UAT:
/gsd-execute-phase N -> /gsd-code-review N -> /gsd-code-review N --fix -> /gsd-verify-work N
Optional external source-review lanes (#4209): /gsd-code-review accepts the same reviewer-lane flags as /gsd-review (run gsd_run review-lane flags to list the flags your installation's roster declares, e.g. --codex, --agy). Adding one asks that lane to independently review the same file scope alongside the internal gsd-code-reviewer agent; its findings are unverified corroborating evidence that gsd-code-reviewer re-checks against the actual source before writing anything to REVIEW.md — there is still exactly one REVIEW.md. No reviewer-lane flag is the default and reviews with only the internal agent, unchanged from before #4209. This is separate from /gsd-review, which reviews PLAN.md files before execution, not source code — see Set up cross-AI review.
/gsd-code-review 3 --codex # Corroborate the internal review with the codex reviewer lane
Coverage-Aware UAT Routing
Historically, /gsd-verify-work turned every ## Accomplishments bullet in a SUMMARY into a manual checkpoint — even deliverables already covered one-to-one by a passing unit test. With a green test suite you were still asked to re-confirm things the tests had already proven, every phase.
GSD now lets the executor record, at authoring time, how each deliverable was verified. When a SUMMARY.md carries a coverage: frontmatter block (see the coverage: block reference), /gsd-verify-work routes deterministically:
- Auto-passed — a deliverable marked
human_judgment: falsewhoseverificationlist is non-empty and entirelypassis recorded as passed (source: automated) and never prompted. - Presented — everything else is shown to you for sign-off: anything flagged
human_judgment: true(visual adequacy, multi-device behaviour, subjective quality), anything with no verification, anything not fully passing, and any malformed entry.
The asymmetry is deliberate. The worst outcome is auto-passing something broken that UAT existed to catch, so auto-pass is the narrow, fully-proven case and uncertainty always routes back to you. Flipping the flag alone cannot skip a prompt — a passing test reference is also required. SUMMARYs without a coverage: block behave exactly as before (prose-based checkpoints), so nothing changes for existing or un-migrated phases.
Command And Configuration Reference
- Command Reference: see
docs/COMMANDS.mdfor every stable command's flags, subcommands, and examples. - Configuration Reference: see
docs/CONFIGURATION.mdfor the fullconfig.jsonschema, model-profile table, git branching strategies, and security settings. - Discuss Mode: see
docs/workflow-discuss-mode.mdfor interview vs assumptions mode.
Graphify capability gate (tri-state, v1.43+)
Graphify commands (graphify status, graphify build, graphify query, graphify diff) now respect the full tri-state capability gate:
- Installed — the
gsd-graphify-*skills are present in the active install profile. - Surfaced — those skills appear on the current runtime surface (e.g., in
~/.claude/commands/gsd/). - Config-enabled —
graphify.enabled: trueis set in.planning/config.json.
All three conditions must be true. Setting graphify.enabled: true alone is no longer sufficient if graphify has not been installed and surfaced. If graphify commands return { disabled: true } after upgrading, verify that the install profile includes graphify skills (gsd-tools capability state) and re-run the installer to surface them.
Intel capability gate (tri-state, v1.44+)
Intel commands (intel status, intel query, intel diff, intel snapshot, intel validate, intel api-surface) now respect the full tri-state capability gate (same resolver as graphify above):
- Installed — the intel capability is present in the active install profile (intel has no skill files, so this is vacuously true for all profiles).
- Surfaced — the intel capability is on the current runtime surface (vacuously true for all surfaces since intel registers no skill stems).
- Config-enabled —
intel.enabled: trueis set in.planning/config.json.
For intel, conditions 1 and 2 are always satisfied (intel has no skill files). The effective gate is intel.enabled in config — the same behaviour as before, but now enforced through the shared isCapabilityActive('intel', cwd) resolver rather than a direct config read. This means intel honours the full capability-state pipeline, including any future install-profile or surface restrictions. If intel commands return { disabled: true }, ensure intel.enabled: true is set in .planning/config.json and verify gsd-tools capability state shows intel as active.
Usage Examples
New Project (Full Cycle)
claude --dangerously-skip-permissions
/gsd-new-project # Answer questions, configure, approve roadmap
/clear
/gsd-discuss-phase 1 # Lock in your preferences
/gsd-ui-phase 1 # Design contract (frontend phases)
/gsd-plan-phase 1 # Research + plan + verify
/gsd-execute-phase 1 # Parallel execution
/gsd-verify-work 1 # Manual UAT
/gsd-ship 1 # Create PR from verified work
/gsd-ui-review 1 # Visual audit (frontend phases)
/clear
/gsd-progress --next # Auto-detect and run next step
...
/gsd-audit-milestone # Check everything shipped
/gsd-complete-milestone # Archive, tag, done
/gsd-pause-work --report # Generate session summary
Caution
The permissions flag is optional. It skips per-file confirmation while GSD's sub-agents read and write files. Use it only in low-stakes or throwaway contexts. To keep confirmations enabled, start with
claudeinstead. For real work, read the security model first.
New Project from Existing Document
/gsd-new-project --auto @prd.md # Auto-runs research/requirements/roadmap from your doc
/clear
/gsd-discuss-phase 1 # Normal flow from here
Existing Codebase
/gsd-onboard # Safely map, ingest docs, and initialize planning
# Follow the printed top-level handoff commands, then rerun /gsd-onboard
# (normal phase workflow from here)
/gsd-onboard routes through /gsd-map-codebase, /gsd-ingest-docs, and /gsd-new-project without nesting interactive workflows or overwriting existing planning files silently.
Post-execute drift detection (#2003). After every /gsd-execute-phase, GSD checks whether the phase introduced enough structural change to make .planning/codebase/STRUCTURE.md stale. Flip the behavior with:
/gsd-settings workflow.drift_action auto-remap # remap automatically
/gsd-settings workflow.drift_threshold 5 # tune sensitivity
Plan Drift Guard
Default-on. The plan drift guard (plan_review.source_grounding: true) runs during plan review and verifies that every symbol your plans cite — decorators, classes, functions, CLI flags — actually exists in your source tree at review time. This catches hallucinated names before any execution agent runs.
Two axes, one switch. The same guard also runs a cross-artifact fact-drift pass: when ROADMAP.md, PLAN.md, STATE.md and CONTEXT.md state the same fact in contradictory ways — a phase marked complete in one and in progress in the other, a success criterion the plan restates with a different outcome, a term used against its CONTEXT.md definition — you get an advisory finding in REVIEWS.md naming both locations and which one is authoritative. It keys on contradicting knowledge, not on similar-looking text, so a plan that simply restates a criterion in its own words is not flagged. The findings never block convergence.
What it catches:
- Functions referenced in a PLAN.md step that don't exist in source
- Class or decorator names that were renamed or removed since the plan was written
- CLI flags documented in a plan that are not defined in the argument parser
- Module paths cited in implementation steps that resolve to no files
Needs-acknowledgement behavior. When the guard finds a missing symbol, it emits a needs-acknowledgement notice in the plan review output rather than hard-blocking. You can acknowledge and proceed (the symbol may be intentionally new) or request a plan revision. The guard does not auto-reject plans — it surfaces signal for human decision.
Works without intel. By default the guard uses grep/ripgrep to search source files — no pre-indexing required. If you have run /gsd-map-codebase with intel.enabled: true, set plan_review.source_grounding_authority: intel to use the faster pre-built api-map.json index instead.
# Enable/disable (default: on)
/gsd-settings plan_review.source_grounding true
/gsd-settings plan_review.source_grounding false
# Switch resolver authority
/gsd-settings plan_review.source_grounding_authority grep # live grep (default)
/gsd-settings plan_review.source_grounding_authority intel # pre-indexed api-map.json
Toggle at project setup (/gsd-new-project asks during workflow preferences) or any time via /gsd-settings (Planning section → Drift Guard).
Quick Bug Fix
/gsd-quick
> "Fix the login button not responding on mobile Safari"
Resuming After a Break
/gsd-progress # See where you left off and what's next
# or
/gsd-resume-work # Full context restoration from last session
Preparing for Release
/gsd-audit-milestone # Check requirements coverage, detect stubs
/gsd-complete-milestone # Archive, tag, done
Speed vs Quality Presets
| Scenario | Mode | Granularity | Profile | Research | Plan Check | Verifier |
|---|---|---|---|---|---|---|
| Prototyping | yolo |
coarse |
budget |
off | off | off |
| Normal dev | interactive |
standard |
balanced |
on | on | on |
| Production | interactive |
fine |
quality |
on | on | on |
Skipping discuss-phase in autonomous mode: When running in yolo mode, set workflow.skip_discuss: true via /gsd-settings.
Mid-Milestone Scope Changes
/gsd-phase # Append a new phase to the roadmap (default mode)
/gsd-phase --insert 3 # Insert urgent work between phases 3 and 4
/gsd-phase --remove 7 # Descope phase 7 and renumber
/gsd-phase --edit 4 # Edit any field of phase 4 in place
Troubleshooting
For a comprehensive troubleshooting guide, see Recover and troubleshoot. The most common issues are summarised below.
Programmatic CLI (gsd-tools query vs gsd-tools.cjs)
For automation, prefer gsd-tools query with a registered subcommand (see CLI-TOOLS.md — SDK and programmatic access and QUERY-HANDLERS.md). The legacy node $HOME/.claude/gsd-core/bin/gsd-tools.cjs CLI remains supported.
STATE.md Out of Sync
node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" state validate # Detect drift
node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" state sync --verify # Preview changes
node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" state sync # Reconstruct STATE.md
state validate's report carries a scope field alongside valid — valid:true means no drift was found, but scope says whether the check could actually run at all (a phase that could not be resolved, or an unreadable frontmatter/phases directory, reports valid:true too, because there was nothing to flag). See Interpret state validate results before treating a passing state validate as "clean."
A Command Looks Frozen After "Spawning..."
GSD subagents run in a separate context window — their work is invisible to the parent session while in progress. Do not interrupt the session. Wait for the result; research and planning agents routinely take 1–5 minutes.
Context Degradation During Long Sessions
Clear your context window between major commands: /clear in Claude Code. GSD is designed around fresh contexts — every subagent gets a clean 200K window. Use /gsd-resume-work or /gsd-progress to restore state after clearing.
Plans Seem Wrong or Misaligned
Run /gsd-discuss-phase [N] before planning. Most plan quality issues come from Claude making assumptions that CONTEXT.md would have prevented.
Execution Fails or Produces Stubs
Check that the plan was not too ambitious. Plans should have 2–3 tasks maximum. Re-plan with smaller scope.
Lost Track of Where You Are
Run /gsd-progress. It reads all state files and tells you exactly where you are and what to do next.
Model Costs Too High
Switch to budget profile: /gsd-config --profile budget. Disable research and plan-check agents via /gsd-settings if the domain is familiar.
Tuning model cost by phase (models) — added in v1.40
Add a models block to .planning/config.json:
{
"model_profile": "balanced",
"models": {
"planning": "opus",
"discuss": "opus",
"research": "sonnet",
"execution": "opus",
"verification": "sonnet",
"completion": "sonnet"
}
}
Need a per-agent exception? Add model_overrides alongside — it wins over models:
{
"models": { "research": "sonnet" },
"model_overrides": {
"gsd-codebase-mapper": "haiku"
}
}
For the full mapping table and resolution-precedence rules, see Per-Phase-Type Models.
Cheap-by-default with dynamic_routing — added in v1.40
{
"dynamic_routing": {
"enabled": true,
"tier_models": {
"light": "haiku",
"standard": "sonnet",
"heavy": "opus"
},
"escalate_on_failure": true,
"max_escalations": 1
}
}
For the full agent → tier mapping, see Dynamic Routing.
Trim MCP servers to reduce per-turn cost
Before tuning model_profile or models.<phase_type>, audit which MCP servers your harness has enabled. Every enabled MCP server injects its tool schema into every turn — heavyweight servers can cost 20k+ tokens each.
This is a harness setting, not a GSD setting. The toggle lives in .claude/settings.json:
{
"enabledMcpjsonServers": ["context7"],
"disabledMcpjsonServers": ["playwright", "mac-tools"]
}
Quick audit before a long phase:
- Are any browser / playwright tools enabled when this phase has no UI work?
- Are any platform-specific tools enabled when not needed?
- Are any project-specific MCPs from a different project still enabled here?
Each disabled server removes its schema from every subsequent turn. Trimming MCPs compounds with model_profile tuning — both levers are additive, and MCP savings show up immediately across every subagent the orchestrator spawns.
For the full audit, harness reference, and the composition note with model_profile, see MCP Tool Schema Cost in the bundled context-budget.md reference.
Using Non-Claude Runtimes (Codex, OpenCode, Antigravity CLI, Kilo)
Codex CLI minimum supported version:
0.130.0(issue #3562).
If you installed GSD for a non-Claude runtime, the installer already configured model resolution. No manual setup is needed — resolve_model_ids: "omit" is set automatically, which tells GSD to skip Anthropic model ID resolution and let the runtime choose its own default model.
To assign different models on a non-Claude runtime:
{
"resolve_model_ids": "omit",
"model_overrides": {
"gsd-planner": "o3",
"gsd-executor": "o4-mini",
"gsd-debugger": "o3"
}
}
Codex skill picker and agent scheduling (#774)
GSD enriches each Codex install with an additional artifact:
- Flex-tier scheduling — light-tier agents (haiku-equivalent) emit
service_tier = "flex"andmodel_verbosity = "low"in their agent TOML. The Codex scheduler routes these agents to the flex tier (lower cost, background processing) and suppresses verbose token output.
GSD skills appear in the Codex /skills picker via their SKILL.md file, which Codex discovers automatically. No agents/openai.yaml sidecar is emitted — doing so caused duplicate autocomplete entries (#1326).
This enrichment is written automatically at install time and requires no manual configuration. Requires Codex CLI ≥ 0.130.0.
Switching from Claude to Codex with one config change (#2517)
{
"runtime": "codex",
"model_profile": "balanced"
}
Per-runtime command enrichment
When generating artifacts, the installer adapts GSD commands to each runtime's native command schema:
- Qwen Code — main-loop skills carry Qwen's numeric
priorityfield so the most-used workflows (e.g.new-project,plan-phase,execute-phase) sort first in the/skillslist; utility skills are left unset. Higher values sort earlier; the field affects only the/skillslist order.
See How to install GSD Core on your runtime for the full per-runtime details.
Manual install / no-Node.js setup
If you cannot run the GSD installer, you cannot use the source files in agents/ directly — they are in Claude Code's native frontmatter format. For OpenCode, two transformations are required:
| Field | GSD source format | OpenCode-valid format | Action |
|---|---|---|---|
tools: |
Read, Bash, Grep (comma-string) |
Not a frontmatter field | Remove the tools: line entirely |
color: |
Plain CSS color name | Hex or OpenCode semantic name | Convert to hex or remove |
Alternative: run the installer on any machine with Node.js:
npx @opengsd/gsd-core@latest --opencode --global
Installing for Cline
npx @opengsd/gsd-core --cline --global # applies to all projects
npx @opengsd/gsd-core --cline --local # this project only
Installing for CodeBuddy
npx @opengsd/gsd-core --codebuddy --global
GSD installs four surfaces for CodeBuddy: /gsd-* slash commands in ~/.codebuddy/commands/, subagents in ~/.codebuddy/agents/, model-invocable skills in ~/.codebuddy/skills/, and settings.json hooks. The skills are emitted with user-invocable: false so the slash commands are the single / menu surface (no duplicate entries).
Installing for Qwen Code
npx @opengsd/gsd-core --qwen --global
Installing for Prerelease Editions
Set the runtime's *_CONFIG_DIR env var to the prerelease directory before running the installer:
WINDSURF_CONFIG_DIR=~/.codeium/windsurf-next npx @opengsd/gsd-core@latest --windsurf --global
Env-var reference for supported runtimes:
| Runtime | Stable default | Override env var |
|---|---|---|
| Claude Code | ~/.claude |
CLAUDE_CONFIG_DIR |
| OpenCode | XDG_CONFIG_HOME/opencode |
OPENCODE_CONFIG_DIR |
| Codex | (per Codex CLI) | --config-dir flag |
| Copilot | ~/.copilot |
COPILOT_CONFIG_DIR (or COPILOT_HOME) |
| Cursor | ~/.cursor |
CURSOR_CONFIG_DIR |
| Windsurf / Devin Desktop | ~/.codeium/windsurf |
WINDSURF_CONFIG_DIR |
| Antigravity | auto-detected | ANTIGRAVITY_CONFIG_DIR |
| Augment | ~/.augment |
AUGMENT_CONFIG_DIR |
| Trae | ~/.trae |
TRAE_CONFIG_DIR |
| Qwen Code | ~/.qwen |
QWEN_CONFIG_DIR |
| Kilo | ~/.config/kilo |
KILO_CONFIG_DIR |
| CodeBuddy | ~/.codebuddy |
CODEBUDDY_CONFIG_DIR |
| Cline | ~/.cline |
CLINE_CONFIG_DIR |
Using Claude Code with Non-Anthropic Providers
Switch to the inherit profile: /gsd-config --profile inherit. This makes all agents use your current session model.
Working on a Sensitive/Private Project
Set commit_docs: false during /gsd-new-project or via /gsd-settings. Add .planning/ to your .gitignore.
GSD Update Overwrote My Local Changes
Which recovery you need depends on whether you modified a GSD file or added your own:
-
You edited a file GSD ships (an agent prompt, a workflow). Since v1.17 the installer backs it up to
gsd-local-patches/. Run/gsd-update --reapplyto merge your changes back. -
You added your own file inside a GSD-managed directory (a custom skill under
skills/, an extra file incommands/gsd/). The installer saves it togsd-user-files-backup/, and the update offers to restore it once the new version is installed. If you declined, or the backup is left over from an older update, restore it any time:node <config-dir>/gsd-core/bin/gsd-tools.cjs restore-custom-files --config-dir <config-dir> --applyRun it without
--applyfirst to see what would be restored. The backup is never deleted, and the restore skips any file that would overwrite something the new release ships.
Install or Refresh a Release Candidate
To install or refresh GSD from the @next RC dist-tag (the pre-release channel established by ADR #660), run:
/gsd-update --next
# or equivalently:
/gsd-update --rc
The same scope/runtime detection, changelog preview, custom-file backup, and cache clearing apply. Omitting --next/--rc keeps targeting @latest (stable channel, no change). Only the @latest and @next channels are supported — no arbitrary dist-tag can be passed.
Cannot Update via npm
See docs/manual-update.md for a step-by-step manual update procedure.
Workflow Diagnostics (/gsd-forensics)
When a workflow fails in a non-obvious way, run /gsd-forensics to generate a diagnostic report covering git history anomalies, artifact integrity, and state inconsistencies. Output goes to .planning/forensics/.
Pre-populated Permissions (Claude Code)
Since v1.3.1, the installer pre-populates ~/.claude/settings.json (or
settings.local.json for local installs) with the core permissions GSD needs:
{
"permissions": {
"allow": [
"Bash(npx gsd-core *)",
"Read(.planning/*)",
"Edit(.planning/*)",
"Read(STATE.md)",
"Edit(STATE.md)"
]
}
}
These entries eliminate first-run approval prompts for GSD's own tool calls. The merge is non-destructive — your existing permissions are preserved and GSD entries are only appended. Uninstalling GSD removes exactly these entries and preserves any others.
Secret-file protection moved from deny rules to a hook (#4221). Earlier
versions also wrote three permissions.deny rules — Read(.env),
Read(.env.*) and Read(.secrets). Claude Code 2.1.259 hardened the
Bash-side enforcement of Read() deny rules so that any such rule makes every
cd DIR && grep … / cd DIR && cat … compound prompt for approval, even in
auto mode — and GSD's subagents emit hundreds of those per session. The same
protection now ships as the always-on gsd-secret-read-guard.js PreToolUse hook
(Read, Grep and Bash; see Runtime Hooks above for what it covers and its
documented gaps). A hook denial is not a permission rule, so it never arms that
check, and it applies in auto and bypassPermissions modes alike. On install
and uninstall the three retired strings are removed from permissions.deny
(and an emptied deny array is dropped). Note the removal is byte-exact: a
rule you wrote by hand that is identical to one of the three is indistinguishable
from the installer's and is removed as well — re-add it if you want both layers.
Executor Subagent Gets "Permission denied" on Bash Commands
Add the required patterns to ~/.claude/settings.json. Core patterns needed for all stacks:
"Bash(git add:*)",
"Bash(git commit:*)",
"Bash(git merge:*)",
"Bash(git worktree:*)",
"Bash(git rebase:*)",
"Bash(git reset:*)",
"Bash(git checkout:*)",
"Bash(git switch:*)",
"Bash(git restore:*)",
"Bash(git stash:*)",
"Bash(git rm:*)",
"Bash(git mv:*)",
"Bash(git fetch:*)",
"Bash(git cherry-pick:*)",
"Bash(git apply:*)",
"Bash(gh:*)"
Per-project permissions: add the same permissions.allow block to .claude/settings.local.json in your project root instead of ~/.claude/settings.json.
Parallel Execution Causes Build Lock Errors
GSD handles this automatically since v1.26. If you're on an older version, add to your project's CLAUDE.md:
## Git Commit Rules for Agents
All subagent/executor commits MUST use `--no-verify`.
To disable parallel execution entirely: /gsd-settings → set parallelization.enabled to false.
Recovery Quick Reference
| Problem | Solution |
|---|---|
| Lost context / new session | /gsd-resume-work or /gsd-progress |
| Phase went wrong | git revert the phase commits, then re-plan |
| Need to change scope | /gsd-phase (default), /gsd-phase --insert, or /gsd-phase --remove |
| Something broke | /gsd-debug "description" (add --diagnose for analysis without fixes) |
| STATE.md out of sync | state validate then state sync — check the report's scope field, not just valid (interpret results) |
| Workflow state seems corrupted | /gsd-forensics |
| Quick targeted fix | /gsd-quick |
| Plan doesn't match your vision | /gsd-discuss-phase [N] then re-plan |
| Costs running high | /gsd-config --profile budget and /gsd-settings to toggle agents off |
| Update broke local changes | /gsd-update --reapply |
| Custom file gone after an update | gsd-tools restore-custom-files --config-dir <dir> --apply |
| Want session summary for stakeholder | /gsd-pause-work --report |
| Don't know what step is next | /gsd-progress --next |
| Parallel execution build errors | Update GSD or set parallelization.enabled: false |
Project File Structure
.planning/
PROJECT.md # Project vision and context (always loaded)
REQUIREMENTS.md # Scoped v1/v2 requirements with IDs
ROADMAP.md # Phase breakdown with status tracking
STATE.md # Decisions, blockers, session memory
config.json # Workflow configuration
MILESTONES.md # Completed milestone archive
HANDOFF.json # Structured session handoff (from /gsd-pause-work)
research/ # Domain research from /gsd-new-project
reports/ # Session reports (from /gsd-pause-work --report)
todos/
pending/ # Captured ideas awaiting work
completed/ # Completed todos
debug/ # Active debug sessions
resolved/ # Archived debug sessions
spikes/ # Feasibility experiments (from /gsd-spike)
NNN-name/ # Experiment code + README with verdict
MANIFEST.md # Index of all spikes
sketches/ # HTML mockups (from /gsd-sketch)
NNN-name/ # index.html (2-3 variants) + README
themes/
default.css # Shared CSS variables for all sketches
MANIFEST.md # Index of all sketches with winners
codebase/ # Brownfield codebase mapping (from /gsd-map-codebase or /gsd-onboard)
onboarding/ # Brownfield onboarding summary (from /gsd-onboard)
phases/
XX-phase-name/
XX-YY-PLAN.md # Atomic execution plans
XX-YY-SUMMARY.md # Execution outcomes and decisions
CONTEXT.md # Your implementation preferences
RESEARCH.md # Ecosystem research findings
VERIFICATION.md # Post-execution verification results
XX-UI-SPEC.md # UI design contract (from /gsd-ui-phase)
XX-UI-REVIEW.md # Visual audit scores (from /gsd-ui-review)
ui-reviews/ # Screenshots from /gsd-ui-review (gitignored)
Related
- Docs index
- Commands
- Configuration
- The phase loop
- Community Capability Registry & EoS Registry — discover third-party Capabilities and EoS host integrations