* enhance(#3911): give hooks an exit seam that needs no build ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the epic's 128 terminators and cannot reach `terminateNow` today. The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as gsd-agent-isolation-guard.js already does for two other modules — is rejected. That precedent carries its own warning (#3582): those files are tsc output, gitignored and absent on a raw plugin-marketplace or git-clone install, so the hook must first call ensureRuntimeBuild() to self-heal. Making the module a hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and its failure mode is precisely the fail-open this phase exists to remove: a guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam` already encodes that concern, and Design B would have had to add an ensureRuntimeBuild() call to all 19 hooks to satisfy it. So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the registry, preserving the invariant `src/cli-exit.cts`'s own header states: it imports nothing but node:fs and its sibling registry, and the generator dual-emits that sibling alongside each copy so a relative require resolves next to whichever copy loaded it. Shipping needed no change — build-hooks.js already declares HOOKS_SUBDIRS_TO_COPY = ['lib']. Proven, not asserted: the two files are copied into an otherwise-empty tmpdir and a child process requires them and terminates — PASS exits 0, HOOK_DENY exits 2 with the payload on both stdout and stderr. That test fails the moment the hooks copy gains a require reaching outside hooks/lib/. Also fixed inline: the registry's fifth target let any `--write` test overwrite the real committed hooks/lib/exit-code-registry.js, because the test helper derived only three of the other output paths. It now redirects all five, and a regression test asserts every committed artifact is byte-identical after a redirected write. Install-tree goldens pick up the two new shipped paths across 11 runtimes — insertions only, no removals. lint:ci was green while they were stale, so this was found by regenerating rather than by a gate. Verification runs on the remote runner. Refs #3911 * enhance(#3911): declare a crash policy, and migrate the write guard Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`, hand-written because the cli-exit copy beside it is generated: allow(payload) exit 0 deny(payload, stderr?) exit 2 crash(onCrash, payload) whichever the hook DECLARED `crash()` takes the policy as a required argument with no default, which is the whole mechanism: fail-open by accident stops being expressible. A hook must name ALLOW or DENY at the call site, and an unrecognized value terminates INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission does not. `gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by citing this hook's `emitBlock` — but modeled it as sending the same bytes to both streams, when `emitBlock` actually sends full JSON to stdout and only the bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim back to the model. Migrating as written would have turned a readable sentence into a JSON blob for Kimi-backed agents. #3911 requires both "all 19 hooks terminate through terminateNow" and "no hook's effective default changes". Those are jointly satisfiable only by teaching the seam to carry a distinct stderr payload, so `terminateNow` gains an optional third argument: omitted, behavior is byte-for-byte what it was; a string is written raw, which is exactly the Kimi case. The doc comment's inaccurate claim about emitBlock is corrected in place. Proven rather than asserted: the pre-migration file is reconstructed from HEAD and driven with the same catastrophic-shrink payload as the migrated one — exit code, stdout and stderr all byte-identical. Verification runs on the remote runner. Refs #3911 * enhance(#3911): all 19 hooks terminate through the seam Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk now reports zero `process.exit(` call sites across every `hooks/*.js` — down from the 91 the census measured. Each hook with an outer catch declares its policy once, at module top, with the reason that policy is right for that specific guard: a read guard that cannot scan must not block the read; a statusline that renders every prompt must degrade rather than crash; an injection scanner must not retroactively block a result already returned. Those sentences are the deliverable — they are what turns fail-open-by-accident into fail-open-on-purpose. No hook's effective default changed. Wiring exposed two defects, both fixed here rather than noted. A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s `emitForceAddBlock`, matching the pattern already known from the write guard — full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the `stderrPayload` argument added in the previous commit, which is now carrying its second real caller rather than one special case. More seriously, `terminateNow` emitted both streams inside ONE try, so a payload that failed to serialize aborted before the stderr write ever ran. The two windsurf guards write nothing to stdout on a block and only a reason string to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny that silently loses its reason, which is the exact "fails with success" class this epic exists to close. The streams are now emitted independently, each with its own guard, and `undefined` means "nothing to write for this stream" rather than an error. Regression tests inject a throwing write on one fd and assert the other still receives its payload; they fail against the single-try version. Byte-identity was proven per hook, not assumed: each pre-change file is reconstructed from HEAD and driven side by side with the migrated one across its normal path, its deny path, malformed stdin and empty stdin — exit code, stdout and stderr compared. Verification runs on the remote runner. Refs #3911 * enhance(#3911): harden the three shell hooks, and pin every hook's policy `gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh` gain `set -euo pipefail`. The expected hazard did not materialize, and that is worth recording: every intentionally-non-zero command in all three is already the condition of an `if`/`elif`, which `set -e` never fires on, and none of them reads a possibly-unset variable or pipes through a grep that may legitimately match nothing. No `|| true` guards were needed. Each hook was still checked command-by-command before the flags went in rather than after. Twenty-one before/after cases across the three hooks — disabled and enabled, planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload shape, quoted and unquoted `-m`, valid and over-long Conventional Commits — all match on exit code, stdout and stderr. The hardening is shown to actually fire, not merely added: with a stubbed `node` that fails at the JSON-emit step, phase-boundary and session-state go from silently exiting 0 with empty stdout to failing visibly with the error surfaced. No such case could be constructed for `gsd-validate-commit.sh`, whose every statement already sits inside an if-condition — recorded as unproven rather than claimed. `tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks for, table-driven over all 19 hooks rather than 76 hand-written cases: normal allow, deny where a deny path exists, crash-honors-the-declared-policy, and an unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The deny assertions encode each hook's ACTUAL stream split rather than a uniform shape, since four of the six deliberately differ. A drift guard enumerates `hooks/*.js` and fails if a terminating hook is ever added without a row. Writing those tests surfaced two hooks that emit a block decision in their JSON body and exit 0. Both were checked rather than assumed, and neither is a fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js` follows Cursor's JSON-body protocol. They are deliberately left alone — a mechanical sweep to `deny()` would have broken exactly these two. Verification runs on the remote runner. Refs #3911 * fix(#3838): the commit validator says when it could not validate #3911 claims to subsume #3838. Measurement said otherwise, so this closes it for real rather than by assertion. `set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash exempts a command used as an `if` condition from `set -e`, and all three of the hook's swallow-and-pass sites are exactly that shape. Verified against the hardened hook with a node shim that fails only the classifier call — a non-conforming commit still exited 0 with empty stdout AND empty stderr, indistinguishable from "your commit conforms". That is the defect verbatim. All three sites named in #3838 now capture the real exit status instead of consuming it as a condition, and each distinguishes its genuine negative from "could not run": - the classifier: 0 = is a git commit, 1 = genuinely not one, anything else = could not classify. Its `node -e` now wraps the require and the call in try/catch and exits 3 on a throw, so a broken require chain can never be mistaken for `isGitSubcommand` legitimately returning false — which is the arm that matters, since `token-scanner.cjs` is a gitignored build artifact and a fresh checkout lands there. - the opt-in config read and the JSON command extraction get the same treatment. On "could not run" the hook emits a diagnostic to stderr naming which check failed and why, then exits 0. The issue confirms this is safe — it is a PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it the smallest sufficient fix. The gate still fails open, but it can no longer do so silently, which is the whole complaint: a validator that disables itself quietly costs more than one that is absent, because it is trusted. Both controls are unchanged and pinned by tests: a conforming commit still passes silently, a non-conforming one still exits 2 with its existing block payload. The defect test asserts stderr is non-empty and names the failure; it fails against the pre-fix hook. Verification runs on the remote runner. Refs #3911, #3838 * docs(#3911): document the hook crash-policy contract Reference and Explanation via a new docs/features fragment (FEATURES.md is generated from it), INVENTORY rows for the three new hooks/lib files, and an ARCHITECTURE note on the hooks section. How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md — a hook author now has to choose and declare a crash policy, which is more than one step and crosses into which harness protocol their hook speaks. It covers allow/deny/crash, writing an ON_CRASH reason that is actually useful, when a deny needs a distinct stderr payload, the two hooks whose harness reads a JSON-body decision and must NOT use deny(), and what to do when a check cannot run at all — with #3838 as the worked example. Refs #3911 * test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup Two review findings. The acceptance criterion 'hooks/dist/** stays in parity via the build seam (lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks something else — that a hook requiring a compiled gsd-core/bin/lib module also calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a recorded ship-blocking bug where a new hook never shipped because a copy list missed it. The suite now builds dist through the repo's own ensureBuiltHooks(), byte-compares each shipped copy against its source, and spawns a child that requires the SHIPPED dist copy and denies — which is what catches a copy that exists but cannot resolve its sibling registry. gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent trap on EXIT replaces them, guarded so cleanup cannot alter the exit status. Behavior-neutral across five cases, with temp-file counts taken before and after each run. Refs #3911 * fix(#3911): stage transitive hook lib requires, not just one level The remote run returned 7 failures across 3 real causes. The important one is a PRODUCTION bug this phase exposed rather than caused. `writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly one level deep and never re-scanned the lib files it staged for their own sibling requires. Nothing had a transitive lib dependency before, so the gap was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js made real Cursor installs ship a bundle that dies at require time with MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor hook runs to completion. The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so the injection scanner crashed at require time and its exit-1 was being read as a policy decision. Migrated to copyScriptWithDeps, which walks the require graph — the repo's recorded rule for this class, since adding another copyFileSync keeps it alive for the next person. The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib file it expected to be named in the abort message; the same throw now fires for a different file first. Its assertion is unchanged in substance — staging still must abort rather than ship a broken hook — only the name is no longer pinned. The last one was my own test asserting an uppercase reason code. Measured against origin/next: the pre-change hook emits the same lowercase 'config_unreadable', so the test was wrong, not the migration. Corrected to the real value rather than making the code match the test. Verification runs on the remote runner. Refs #3911 * chore(#3911): regenerate the cursor install-tree golden The staging fix means a Cursor install now correctly carries the two transitive lib files it was silently missing. Additive only — no path was removed. The golden diff is the evidence the packaging defect was real. Refs #3911 * chore(#3911): backfill the changeset PR number Refs #3911 * fix(#3911): a git probe that timed out is not a negative A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just past the 2000ms budget these hooks give their git probes. The three that passed took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns, the hook reads the non-zero result as "not a git repo", and allows with exit 0 and empty stdout AND empty stderr. Under load, the guards silently stop guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks this phase is about. The repo had already recognized the class in one place — gsd-cursor-subagent-start.js fail-closed-denies on `git_timed_out` (#3045) — but nowhere else. `hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than folding all four into `status !== 0`. Three guards route their eight git probes through it. The resolution is the same shape #3838 took, and the same one that issue endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes on any path** — a developer on a loaded machine is still not blocked, which keeps #3911's declaration-pass contract intact for exit codes. What changes is that the hook now says on stderr which probe could not answer, instead of presenting silence as a clean verdict. Scope was checked across every hooks/*.js, not just the three that failed: gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only a cosmetic display segment, not an allow/deny decision, and are left alone. The C2 deny assertion was a real-race test — it demanded exit 2 while a slow git legitimately yields 0. It now requires the hook to either deny, or allow with a diagnostic naming the probe that could not run; a silent allow still fails, so the assertion is not vacuous. A deterministic regression stubs git on PATH to sleep past the budget rather than waiting for load to reproduce it. Verification runs on the remote runner. Refs #3911 * test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows The deterministic timeout regression stubbed git on PATH and asserted the guard reports rather than silently allows. It passes on Linux and macOS and failed on Windows in 83ms and 176ms — the stub was never invoked at all. Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The git.cmd branch could not have worked and is removed rather than left implying a Windows path that does. Adding shell:true to the hooks to serve a test would change product behavior and widen an injection surface, so the case is skipped on win32 only, with the mechanism written into the skip reason so a future reader does not 'fix' it that way. Linux and macOS keep the coverage, and macOS is where the underlying fail-open was actually caught. Refs #3911 --------- Co-authored-by: sim <sim@local>
77 KiB
GSD Core Architecture
System architecture for contributors and advanced users. For user-facing documentation, see Feature Reference or User Guide.
Table of Contents
- System Overview
- Design Principles
- Component Architecture
- Agent Model
- Data Flow
- File System Layout
- Installer Architecture
- Hook System
- CLI Tools Layer
- Runtime Abstraction
System Overview
GSD Core is a meta-prompting framework that sits between the user and AI coding agents (Claude Code, Kimi CLI, OpenCode, Kilo, Codex, Copilot, Antigravity, Trae, Cline, Augment Code). It provides:
- Context engineering — Structured artifacts that give the AI everything it needs per task (see Context engineering)
- Multi-agent orchestration — Thin orchestrators that spawn specialized agents with fresh context windows (see Multi-agent orchestration)
- Spec-driven development — Requirements → research → plans → execution → verification pipeline
- State management — Persistent project memory across sessions and context resets
┌──────────────────────────────────────────────────────┐
│ USER │
│ /gsd-command [args] │
└─────────────────────┬────────────────────────────────┘
│
┌─────────────────────▼────────────────────────────────┐
│ COMMAND LAYER │
│ commands/gsd/*.md — Prompt-based command files │
│ (Claude Code custom commands / Codex skills) │
└─────────────────────┬────────────────────────────────┘
│
┌─────────────────────▼────────────────────────────────┐
│ WORKFLOW LAYER │
│ gsd-core/workflows/*.md — Orchestration logic │
│ (Reads references, spawns agents, manages state) │
└──────┬──────────────┬─────────────────┬──────────────┘
│ │ │
┌──────▼──────┐ ┌─────▼─────┐ ┌────────▼───────┐
│ AGENT │ │ AGENT │ │ AGENT │
│ (fresh │ │ (fresh │ │ (fresh │
│ context) │ │ context)│ │ context) │
└──────┬──────┘ └─────┬─────┘ └────────┬───────┘
│ │ │
┌──────▼──────────────▼─────────────────▼──────────────┐
│ CLI TOOLS LAYER │
│ gsd-tools.cjs command families + domain modules │
│ command-routing-hub + observability seams │
└──────────────────────┬───────────────────────────────┘
│
┌──────────────────────▼───────────────────────────────┐
│ FILE SYSTEM (.planning/) │
│ PROJECT.md | REQUIREMENTS.md | ROADMAP.md │
│ STATE.md | config.json | phases/ | research/ │
└──────────────────────────────────────────────────────┘
Design Principles
1. Fresh Context Per Agent
Every agent spawned by an orchestrator gets a clean context window (up to 200K tokens). This eliminates context rot — the quality degradation that happens as an AI fills its context window with accumulated conversation.
2. Thin Orchestrators
Workflow files (gsd-core/workflows/*.md) never do heavy lifting. They:
- Load context via
gsd-tools.cjs init <workflow> - Spawn specialized agents with focused prompts
- Collect results and route to the next step
- Update state between steps
3. File-Based State
All state lives in .planning/ as human-readable Markdown and JSON. No database, no server, no external dependencies. This means:
- State survives context resets (
/clear) - State is inspectable by both humans and agents
- State can be committed to git for team visibility
4. Absent = Enabled
Workflow feature flags follow the absent = enabled pattern. If a key is missing from config.json, it defaults to true. Users explicitly disable features; they don't need to enable defaults.
5. Defense in Depth
Multiple layers prevent common failure modes:
- Plans are verified before execution (plan-checker agent)
- Execution produces atomic commits per task
- Post-execution verification checks against phase goals
- UAT provides human verification as final gate
Component Architecture
Commands (commands/gsd/*.md)
User-facing entry points. Each file contains YAML frontmatter (name, description, allowed-tools) and a prompt body that bootstraps the workflow. Commands are installed as:
- Claude Code: Custom slash commands (hyphen form,
/gsd-command-name) - OpenCode / Kilo: Slash commands (hyphen form,
/gsd-command-name) - Codex: Skills (
$gsd-command-name) - Copilot: Slash commands (hyphen form,
/gsd-command-name) - Kimi CLI: Agent Skills (
/skill:gsd-command-name) plus an explicit custom agent launch withkimi --agent-file - Antigravity: Skills
Total commands: see docs/INVENTORY.md for the authoritative count and full roster.
Two-stage hierarchical routing (v1.40, #2792)
To keep the eager skill-listing token cost low, v1.40 introduces six namespace meta-skills (gsd-workflow, gsd-project, gsd-quality, gsd-context, gsd-manage, gsd-ideate — sourced from commands/gsd/ns-*.md, but the invocable name: is the bare form shown here) layered above the concrete sub-skills. On runtimes with non-recursive skill loaders (cline, qwen, hermes, augment, trae) the installer now realizes this fully: it emits only the 6 namespace router bundles as top-level skills and nests the ~61 concrete skills under <router>/skills/<name>/SKILL.md, so the eager listing is ≈6 entries instead of ≈67. The model selects a namespace router, which instructs it to read the nested concrete skill file via a routing table embedded in the router body. On these runtimes concrete skills are not directly invocable by bare name via the Skill tool; they are reachable through the router. Slash commands (/gsd-*, via the separate commands surface) are unaffected where the runtime has one. On runtimes with recursive or unconfirmed skill loaders (claude global, cursor, codex, copilot, windsurf, codebuddy, opencode, kilo, antigravity) the layout remains flat — all skills emitted at the top level as before. Antigravity moved from nested to flat in #1614: agy scans only skills/<name>/SKILL.md, so nested sub-skills were unreachable. Claude was reverted to flat in #924: the Skill tool hard-errors on unknown names rather than re-routing via the router, so nested concrete skills were uninvokable.
The router descriptions use pipe-separated keyword tags (≤ 60 chars) per the Tool Attention research showing keyword-dense tags outperform prose for routing at ~40 % the token cost.
MCP token-budget interaction
The eager skill listing is one of two recurring per-turn token costs. The other is the MCP tool schema injected by every enabled MCP server in .claude/settings.json. Heavyweight MCP servers (browser/playwright, Mac-tools, Windows-tools) can each cost 20 k+ tokens per turn — often dwarfing what model_profile tuning saves. The toggle lives in the Claude Code harness (enabledMcpjsonServers / disabledMcpjsonServers in .claude/settings.json) and is not a GSD concern. Together, the two-stage routing layer (#2792) and disciplined MCP enablement are the largest cost levers per turn. See docs/USER-GUIDE.md and references/context-budget.md for the audit checklist.
Workflows (gsd-core/workflows/*.md)
Orchestration logic that commands reference. Contains the step-by-step process including:
- Context loading via
gsd-tools.cjs inithandlers - Agent spawn instructions with model resolution
- Gate/checkpoint definitions
- State update patterns
- Error handling and recovery
Total workflows: see docs/INVENTORY.md for the authoritative count and full roster.
Progressive disclosure for workflows
Workflow files are loaded verbatim into Claude's context every time the
corresponding /gsd-* command is invoked. The workflow size budget enforced by
tests/workflow-size-budget.test.cjs keeps each file bounded, mirroring the
the agent size-budget convention. The budget is measured in bytes (#717), not lines:
line count over-penalizes prose and under-catches token-dense tables and code
blocks, whereas bytes are deterministic and match the unit our vendors bound on
— Codex truncates instruction docs past 32,768 bytes (project_doc_max_bytes).
We adopt that unit, not that exact number: the XL/LARGE ceilings below sit above
32,768 because these are grandfathered top-level orchestrators loaded by Claude,
not Codex AGENTS.md docs.
| Tier | Per-file byte limit |
|---|---|
XL |
90,000 — top-level orchestrators (execute-phase, plan-phase, new-project) |
LARGE |
54,000 — multi-step planners and large feature workflows |
DEFAULT |
38,000 — focused single-purpose workflows (the target tier) |
Ceilings are not fixed forever: under the tighten-only ratchet (#597) each one tracks its tier's current high-water mark within a small grace band, so budgets may only decrease over time.
Why the budget exists. With prompt caching the per-invocation cost of a large workflow is modest (cache reads run ~10% of input). The stronger, caching-independent reason is quality: as context grows, recall and reasoning degrade ("context rot" / attention budget), so leaner, higher-signal instructions produce better plans. The ceiling protects the agent's attention, not just the token bill.
Because the budget measures one file, it is a proxy for the real goal —
bounded loaded context. Extraction only helps when the extracted content is
loaded lazily (Read at the step that needs it). Moving prose into a file
that is still eagerly @-imported shrinks the measured file without shrinking
loaded context, which games the proxy rather than serving the goal.
workflows/discuss-phase.md is held to a stricter <30,000-byte ceiling per
the discuss-phase byte budget (#717; the discuss-phase/modes split keeps it ≈32000 bytes). When a workflow grows
beyond its tier, extract per-mode bodies into
workflows/<workflow>/modes/<mode>.md, templates into
workflows/<workflow>/templates/, and shared knowledge into
gsd-core/references/. The parent file becomes a thin dispatcher that
Reads only the mode and template files needed for the current invocation.
workflows/discuss-phase/ is the canonical example of this pattern —
parent dispatches, modes/ holds per-flag behavior (power.md, all.md,
auto.md, chain.md, text.md, batch.md, analyze.md, default.md,
advisor.md), and templates/ holds CONTEXT.md, DISCUSSION-LOG.md, and
checkpoint.json schemas that are read only when the corresponding output
file is being written.
workflows/plan-phase.md, workflows/execute-phase.md, and the
gsd-planner / gsd-executor agent definitions apply the same discipline
to their MVP-only reference bodies — planner-mvp-mode.md,
user-story-template.md, skeleton-template.md, and execute-mvp-tdd.md
are referenced for the planner/executor to Read only on MVP,
Walking-Skeleton, or MVP+TDD paths, rather than eagerly @-imported, so
non-MVP runs do not pay their context cost (guards against the "@-import
behind a conditional still loads eagerly" leak; see #720). The dedicated
mvp-phase workflow keeps its eager imports, since it is always MVP.
Agents (agents/*.md)
Specialized agent definitions with frontmatter specifying:
name— Agent identifierdescription— Role and purposetools— Allowed tool access (Read, Write, Edit, Bash, Grep, Glob, WebSearch, etc.)color— Terminal output color for visual distinction
Total agents: 33
References (gsd-core/references/*.md)
Shared knowledge documents that workflows and agents @-reference (see docs/INVENTORY.md for the authoritative full roster):
Core references:
checkpoints.md— Checkpoint type definitions and interaction patternsgates.md— 4 canonical gate types (Confirm, Quality, Safety, Transition) wired into plan-checker and verifiermodel-profiles.md— Per-agent model tier assignmentsmodel-profile-resolution.md— Model resolution algorithm documentationverification-patterns.md— How to verify different artifact typesverification-overrides.md— Per-artifact verification override rulesplanning-config.md— Full config schema and behaviorgit-integration.md— Git commit, branching, and history patternsgit-planning-commit.md— Planning directory commit conventionsquestioning.md— Dream extraction philosophy for project initializationtdd.md— Test-driven development integration patternsui-brand.md— Visual output formatting patternscommon-bug-patterns.md— Common bug patterns for code review and verification
Workflow references:
agent-contracts.md— Formal interface between orchestrators and agentscontext-budget.md— Context window budget allocation rulescontinuation-format.md— Session continuation/resume formatdomain-probes.md— Domain-specific probing questions for discuss-phasegate-prompts.md— Gate/checkpoint prompt templatesrevision-loop.md— Plan revision iteration patternsuniversal-anti-patterns.md— Common anti-patterns to detect and avoidartifact-types.md— Planning artifact type definitionsphase-argument-parsing.md— Phase argument parsing conventionsdecimal-phase-calculation.md— Decimal sub-phase numbering rulesworkstream-flag.md— Workstream active pointer conventionsuser-profiling.md— User behavioral profiling methodologythinking-partner.md— Conditional thinking partner activation at decision points
Thinking model references:
References for integrating thinking-class models (o3, o4-mini, Gemini 2.5 Pro) into GSD workflows:
thinking-models-debug.md— Thinking model patterns for debugging workflowsthinking-models-execution.md— Thinking model patterns for execution agentsthinking-models-planning.md— Thinking model patterns for planning agentsthinking-models-research.md— Thinking model patterns for research agentsthinking-models-verification.md— Thinking model patterns for verification agents
Modular planner decomposition:
The planner agent (agents/gsd-planner.md) was decomposed from a single monolithic file into a core agent plus reference modules to stay under the 50K character limit imposed by some runtimes:
planner-gap-closure.md— Gap closure mode behavior (reads VERIFICATION.md, targeted replanning)planner-reviews.md— Cross-AI review integration (reads REVIEWS.md from/gsd-review)planner-revision.md— Plan revision patterns for iterative refinement
Templates (gsd-core/templates/)
Markdown templates for all planning artifacts. Used by gsd-tools.cjs template fill / phase.scaffold (and top-level scaffold) to create pre-structured files:
project.md,requirements.md,roadmap.md,state.md— Core project filesphase-prompt.md— Phase execution prompt templatesummary.md(+summary-minimal.md,summary-standard.md,summary-complex.md) — Granularity-aware summary templatesDEBUG.md— Debug session tracking templateUI-SPEC.md,UAT.md,VALIDATION.md— Specialized verification templatesdiscussion-log.md— Discussion audit trail templatecodebase/— Brownfield mapping templates (stack, architecture, conventions, concerns, structure, testing, integrations)research-project/— Research output templates (SUMMARY, STACK, FEATURES, ARCHITECTURE, PITFALLS)
Hooks (hooks/)
Runtime hooks that integrate with the host AI agent:
| Hook | Event | Purpose |
|---|---|---|
gsd-statusline.js |
statusLine |
Displays model (long-context suffixes like (1M context) collapse to a compact (1M) badge), task, directory, and context usage bar |
gsd-context-monitor.js |
PostToolUse / AfterTool |
Injects agent-facing context warnings at 35%/25% remaining |
gsd-check-update.js |
SessionStart |
Foreground trigger for the background update check |
gsd-ensure-canonical-path.js |
SessionStart |
For Claude Code plugin installs, symlinks ~/.claude/gsd-core/{bin,contexts,references,templates,workflows} to the plugin's bundled tree so @~/.claude/gsd-core/... includes resolve; runs first in SessionStart, no-op in classic installs, self-heals after claude plugin update (#997) |
gsd-check-update-worker.js |
(helper) | Background worker spawned by gsd-check-update.js; no direct event registration |
gsd-prompt-guard.js |
PreToolUse |
Scans .planning/ writes for prompt injection patterns (advisory) |
gsd-read-injection-scanner.js |
PostToolUse |
Scans Read tool output for injected instructions in untrusted content |
gsd-workflow-guard.js |
PreToolUse |
Detects file edits outside GSD workflow context (advisory, opt-in via hooks.workflow_guard) |
gsd-read-guard.js |
PreToolUse |
Advisory guard preventing Edit/Write on files not yet read in the session |
gsd-session-state.sh |
PostToolUse |
Session state tracking for shell-based runtimes |
gsd-validate-commit.sh |
PostToolUse |
Commit validation for conventional commit enforcement |
gsd-phase-boundary.sh |
PostToolUse |
Phase boundary detection for workflow transitions |
See docs/INVENTORY.md for the authoritative hook roster.
Crash policy (ADR-3889 Phase 7, #3911). Every enforcement hook terminates
through hooks/lib/hook-exit.js's allow(payload) (exit 0), deny(payload, stderrPayload?) (exit 2), or crash(onCrash, payload) — the last dispatching
per a HOOK_ON_CRASH policy the hook must declare explicitly (ALLOW or
DENY, no default), so a hook's fail-open/fail-closed stance is a visible
declaration rather than an inference from a bare process.exit(N). Two hooks
are deliberate exceptions — gsd-read-injection-scanner.js (PostToolUse) and
gsd-cursor-subagent-start.js (Cursor) — whose harnesses read the block
decision from the JSON response body at exit 0, not from the exit code, so
they never call deny(). See
Declare a hook's crash policy.
Command Routing Hub (gsd-core/bin/lib/command-routing-hub.cjs)
CJS command family routers dispatch through CommandRoutingHub. The hub owns the no-throw pure-result contract (hub.dispatch() catches internal exceptions and returns { ok: false, kind, ...typedPayload }) and the closed runtime error taxonomy (UnknownCommand, InvalidArgs, HandlerRefusal, HandlerFailure). Router adapters remain thin CLI translators — they build the hub, call dispatch, then map the Result to output()/error() calls. The runtime is single-path (no dual-runtime mode selection). See docs/adr/0174-retire-gsd-sdk-package-boundary.md.
Planned (ADR-2346 / epic #2345): the
runCommand73-case switch is being dissolved into a two-layer dispatch — families via thecommandFamiliesregistry (ADR-959 mechanism, completed) and single-purpose leaf verbs via a table filling the prepared_dispatchNonFamilyseam — collapsingrunCommandto a ~15-line dispatcher. Behavior-preserving; tracked phase-by-phase under epic #2345. The current-state description above holds until each phase lands.
Capability Command Dispatch (gsd-core/bin/gsd-tools.cjs, ADR-1244 D7)
Command families declared by capabilities (commands: [{ family, module, router }]) are dispatched from the registry rather than a hardcoded switch. The runCommand default arm tries, in order:
- First-party —
dispatchCapabilityCommandagainst the frozencapability-registry.cjscommandFamilies, loading the router frombin/lib/. The in-tree families (graphify,intel,audit) reach their routers this way (the legacy hardcoded switch is retired). - Third-party (installed overlay) —
dispatchOverlayCapabilityCommandcallsloadRegistry({ includeInstalled })and dispatches a family only when itscapIdappears in_overlay.commandRoots. The loader lists a command root only for an accepted overlay capability with a committed ledger entry (consent gate), and the router module isrequire()'d from that capability's install root, confined by basename validation +realpathcontainment (rejecting..traversal and symlink escape). This is the one point where third-party capability code executes; see the capability trust model for the consent + confinement + project-scope trust boundary.
Both paths share the same guards: prototype-pollution-safe command keys, an own-property router check, and synchronous-only routers (an async router is a fail-fast error).
Research Module (src/research-{store,provider}.cts, src/package-legitimacy.cts)
The Research Module implements an L2-hybrid seam: code owns the cache, provider policy, and package legitimacy verdicts; MCP owns the actual network fetch.
Three compiled modules (generated to gsd-core/bin/lib/*.cjs per ADR-457) are reachable via gsd-tools query research-plan | research-store | package-legitimacy:
- Research Store — content-addressed cache (
sha256(ecosystem+library+version+query+kind)) with per-source TTL (curated-doc: 30 d, medium: 7 d, web/synthesis: 1 d) and two storage tiers:~/.gsd/research-cachefor cross-project curated-doc hits,.planning/research/.cachefor project-local web/synthesis results. - Research Provider — single
PROVIDER_WATERFALL(Context7→Ref→Jina→websearchfor docs;Exa→Tavily→Perplexity→Brave→websearchfor web;Firecrawl→Jinafor scrape-only).planResearch()returns cache hits plus a fetch plan;classifyConfidence()stampsHIGH|MEDIUM|LOWby provider tier. - Package Legitimacy — registry-API verdicts (npm/PyPI/crates.io injectable adapters) producing
OK|SUS|SLOPper package.slopcheckis an optional escalate-only adapter; absence leaves registry verdicts intact rather than downgrading everything to[ASSUMED].
Data flow:
agent
│
▼
gsd-tools query research-plan ← Research Provider: check cache, build fetch plan
│
├── [cache hits] ──────────────────► RESEARCH.md (digest only, no raw content)
│
└── [fetch plan] ──────────────────► MCP fetch (agent calls MCP tools with the plan)
│
▼
gsd-tools query research-store (put)
│
▼
RESEARCH.md path returned to orchestrator
Agents always return a RESEARCH.md path, never raw fetched content. Context discipline is enforced through subagent isolation, compact provider output, and fetch-to-disk. See ADR-0656.
Context Predicate Fact-Store (src/context-predicates.cts, ADR-1671)
The CONTEXT.md predicate fact-store — every backtick-wrapped CLASS.subkey=value declaration in the repo-root CONTEXT.md — has a compiled parser/selector seam (generated to gsd-core/bin/lib/context-predicates.cjs per ADR-457) reachable live via gsd-tools query context-predicates --class|--prefix|--contains. Fence-aware line skipping mirrors markdown-sectionizer.cts's exported scanFencedBlocks delimiter-matching rule exactly (proven by a fence-skip parity test suite), but is scanned by a LOCAL, interleaved single pass rather than a call into that seam directly: fences and HTML comments must mutually suppress each other's open/close detection while either is active (a fence delimiter inside a real comment, or a comment token inside a real fence, must not falsely toggle the other construct), and that precedence cannot be resolved by two independent passes over scanFencedBlocks's comment-blind output — see src/context-predicates.cts's module doc comment.
scripts/gen-context-index.cjs --check is the CI drift-guard for the committed docs/CONTEXT-INDEX.json artifact: it fails on staleness between a fresh parse of CONTEXT.md and the committed file, and on any duplicate predicate ID. It is wired into lint:generated-sync (so lint:ci, so CI). docs/CONTEXT-INDEX.json is generated — never hand-edit it; regenerate with gen-context-index.cjs --write (also wired into build, after build:lib, and into regen:derived). The generator require()s the compiled context-predicates.cjs, so it must run after build:lib in any pipeline; .github/workflows/test.yml does this.
The committed index intentionally carries no line field for any predicate (ADR-1671 open question 4, resolved by #2928) — committed-but-uncompared metadata goes silently stale, the same defect class the drift-guard exists to catch, with the alarm removed. The live gsd-tools query context-predicates parse still returns line/section for callers that want to cite a source location. See ADR-1671 and CLI Tools Reference.
Workflow Fragmentization and Emission (src/workflow-fragments.cts, ADR-1671)
Workflow markdown under gsd-core/workflows/*.md can mark one or more sections with an
in-file <!-- gsd:section id="<id>" when="<when>" --> / <!-- /gsd:section --> pair. A
compiled parser/composer seam (generated to gsd-core/bin/lib/workflow-fragments.cjs per
ADR-457) partitions a marked document into fragments and recomposes them through the shared
context-composer.cjs budget seam (ADR-1671, #2929) before any per-runtime converter sees the
text — so a marker attribute can never be corrupted by a .claude/ → .windsurf/-style
path-rewrite regex. bin/install.js's copyWithPathReplacement calls composeWorkflow on
every workflow file at emit time; an unmarked file (88 of the 89 shipped workflows today)
parses to a single implicit fragment and round-trips byte-identical, so this is a no-op for
every workflow that hasn't opted in yet.
Every fragment in this phase carries the verbatim strategy, so composition is structurally
non-lossy — nothing is trimmed regardless of budget. Fence and HTML-comment interleaving
reuses the same LOCAL, single-pass, mutually-suppressing scan discipline as
context-predicates.cts (see above), so a marker-shaped line inside a fenced code block or an
unrelated comment is never misread as structural. Markers are stripped at emit — the
installed artifact carries no build metadata and is smaller than the source by exactly the
stripped marker bytes.
See Reference: Workflow fragments for the full marker
grammar, the frozen when= vocabulary, and fail-closed authoring rules, and
ADR-1671 (open questions 1 and 2) for why
in-file markers were chosen over separate fragment files or a sidecar manifest.
Section Manifest (src/section-manifest.cts, ADR-1671 Phases 5 and 6.1)
Two seams turn a workflow's gsd:section markers into per-invocation applicability data.
scripts/gen-section-manifest.cjs --write (wired into build after build:lib, and into
lint:generated-sync) scans gsd-core/workflows/*.md and writes the committed
gsd-core/workflows/section-manifest.json, keyed per workflow —
{workflows: {"<name>": [{id, when, read}]}} — where read is the path of the step file the
section's body was extracted to. A workflow with no marked sections contributes no key at
all: an absent key means degraded/unknown (the caller reads every section, the safe superset),
while a key present with an empty array means "computed, nothing applies". The generator reuses
parseWorkflowSections unchanged rather than re-implementing marker parsing, and fails closed
(--check) on a marker naming a step file that does not exist, a step file no marker
references, or a committed artifact still carrying the pre-6.1 flat {sections: [...]} shape.
A separate pure evaluator, src/section-manifest.cts (compiled to
gsd-core/bin/lib/section-manifest.cjs per ADR-457), maps one invocation's facts —
{flags, phaseNumber, hasPriorPhases} plus the optional needsCodebaseMap, phaseMvpMode and
worktreesEnabled booleans — to an included/excluded partition of section ids via
selectSections. flags is a ReadonlySet<string> of flag tokens; because parseNamedArgs
always materializes a boolean flag key (false when the token was absent, never undefined),
presence is truthiness, and the init router folds a boolean flag's own false into the
absent sentinel before the facts are built. Per Greenspun's Tenth Rule, the evaluator is a total
lookup over the frozen 14-atom when= vocabulary, never a parser: WHEN_PREDICATES is a
hand-written literal map that never derives a predicate from its atom string, and an
unrecognized value fails closed rather than being silently excluded. An atom is admitted only
when it has both a real consuming section and a fact the init seam actually computes — an atom
without the latter would evaluate false forever and silently disable its own section.
execute-phase.md's partial-wave and gap-closure-artifacts sections — previously inlined
directly per #2930's pilot — now delegate to dedicated step files under
gsd-core/workflows/execute-phase/steps/, the same pattern the pre-existing regression-gate
section already used.
CLI Tools (gsd-core/bin/)
Node.js CLI utility (gsd-tools.cjs) with domain modules split across gsd-core/bin/lib/ (see docs/INVENTORY.md for the authoritative roster):
| Module | Responsibility |
|---|---|
config-loader.cjs |
Project config loading — defaults merge, legacy-key migration, workstream overlay, unknown-key/profile-override validation, and federated config overlay (ADR-857 phase 3b) (extracted from core.cjs, ADR-857) |
federated-config.cjs |
Defensive merge of capability-declared config slices (ADR-857 phase 3b); exports mergeFederatedConfig; live for migrated Capability keys that are absent from the central config schema |
core-utils.cjs |
Shared low-level utility primitives — POSIX path normalization, sub-repo/subdirectory scanning, phase file stats, slug/one-liner/plan-id helpers, time-ago (extracted from core.cjs, ADR-857) |
core.cjs |
Shared utilities; compatibility re-exports for planning, I/O (io.cjs), and phase-id helpers |
io.cjs |
CLI I/O primitives — output/error emission, JSON-error mode, large-payload temp-file spillover |
phase-id.cjs |
Pure phase-id parsing/matching helpers — normalize, token match, regex builders (extracted from core.cjs, ADR-857) |
phase-locator.cjs |
Phase-directory search and location — active-phase discovery (searchPhaseInDir, findPhaseInternal) and archived-phase-dir enumeration (getArchivedPhaseDirs), matching phase ids/tokens against the filesystem (extracted from core.cjs, ADR-857) |
roadmap-parser.cjs |
ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from core.cjs, ADR-857) |
planning-workspace.cjs |
Planning seam (planningDir, planningPaths, active workstream routing, .planning/.lock) |
state.cjs |
STATE.md parsing, updating, progression, metrics |
phase.cjs |
Phase directory operations, decimal numbering, plan indexing |
roadmap.cjs |
ROADMAP.md parsing, phase extraction, plan progress |
config.cjs |
config.json read/write, section initialization |
verify.cjs |
Plan structure, phase completeness, reference, commit validation |
template.cjs |
Template selection and filling with variable substitution |
frontmatter.cjs |
YAML frontmatter CRUD operations |
init.cjs |
Compound context loading for each workflow type |
milestone.cjs |
Milestone archival, requirements marking |
commands.cjs |
Misc commands (slug, timestamp, todos, scaffolding, stats) |
model-profiles.cjs |
Model profile resolution table |
model-resolver.cjs |
Model and effort resolution policy — resolves model, tier, granularity, effort, and fast-mode for a given agent from project config and model profiles/catalog (extracted from core.cjs, ADR-857) |
security.cjs |
Path traversal prevention, prompt injection detection, safe JSON parsing, shell argument validation |
uat.cjs |
UAT file parsing, verification debt tracking, audit-uat support |
docs.cjs |
Docs-update workflow init, Markdown scanning, monorepo detection |
workstream.cjs |
Workstream CRUD, migration, session-scoped active pointer |
schema-detect.cjs |
Schema-drift detection for ORM patterns (Prisma, Drizzle, etc.) |
profile-pipeline.cjs |
User behavioral profiling data pipeline, session file scanning |
profile-output.cjs |
Profile rendering, USER-PROFILE.md and dev-preferences.md generation |
context-predicates.cjs |
CONTEXT.md predicate fact-store parser/selector (ADR-1671, #2928); backs query context-predicates and scripts/gen-context-index.cjs's docs/CONTEXT-INDEX.json drift guard; compiled from src/context-predicates.cts |
loop-host-contract.cjs |
Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts; emitted by scripts/gen-loop-host-contract.cjs from workflow markers (ADR-894 §3); consumed by gen-capability-registry.cjs |
capability-loader.cjs |
Runtime registry overlay loader (ADR-1244 D2) — loadRegistry({ includeInstalled }) composes the frozen first-party registry with a validated installed overlay of third-party capability manifests read from global $GSD_HOME/.gsd/capabilities/ and project <projectRoot>/.gsd/capabilities/; first-party always wins; load-time engines.gsd re-gate skips incompatible overlays with a warning; gate-kind hooks on skipped capabilities fail OPEN — no gate is injected; a loud warning (stderr + envelope warnings) names the load failure and the gsd capability remove <id> remediation (#2009) |
capability-registry.cjs |
Generated central Capability Registry — role-partitioned index of all co-located capability declarations; emitted by scripts/gen-capability-registry.cjs (ADR-894 §5) |
loop-resolver.cjs |
Loop Extension Point resolver — ADR-857 phase 3c registry-consuming query; consumes resolved Capability State, filters byLoopPoint by capability enablement plus config activation, renders active hooks as markdown, emits { point, activeHooks, rendered } envelope; gsd-tools loop render-hooks <point> [--config-dir <path>] |
capability-state.cjs |
Unified capability-state resolver — ADR-857 phase 4b/6; composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; pure resolveCapabilityState, reusable resolveCapabilityRuntimeState, I/O cmdCapabilityState, and convenience predicate isCapabilityActive(capId, cwd); gsd-tools capability state [--config-dir <path>] emits { runtimeConfigDir, capabilities[] } where each entry carries enabled (installed && surfaced) and active (enabled && configActivation via the capability's activationKey; absent key → active===enabled) |
capability-validator.cjs |
Shared capability conformance validator (ADR-1244 D2) — extracted from scripts/gen-capability-registry.cjs so the build-time generator and the runtime overlay loader share one validateCapability(manifest) implementation; generative-parity is CI-guarded |
graphify-command-router.cjs |
ADR-959 capability command router — first real capability command cutover (phase 4d-impl-2); extracted from the case 'graphify': arm in gsd-tools.cjs; dispatches build/query/status/diff subcommands; discovered via commandFamilies in the capability registry |
audit-command-router.cjs |
ADR-959 capability command router (phase 4d-impl-3); extracted from the case 'audit-uat': and case 'audit-open': arms in gsd-tools.cjs; routeAuditUat → uat.cjs:cmdAuditUat, routeAuditOpen → audit.cjs:{auditOpenArtifacts,formatAuditReport}; discovered via commandFamilies in the capability registry |
intel-command-router.cjs |
ADR-959 capability command router (phase 4d-impl-4, last first-party cutover); extracted from the case 'intel': arm in gsd-tools.cjs; routeIntelCommand → all 9 intel subcommands via lazy require('./intel.cjs'); preserves non-raw timeAgo transform on status.files[*].updated_at; discovered via commandFamilies in the capability registry |
runtime-hooks-surface.cjs |
Standalone hook-surface writer module (ADR-857 phase 5f-1); owns Cline rules/agents-md/pre-tool-use hook generation, Cursor hooks.json reconciliation, Copilot session-hook config, and Codex hook-block management; extracted verbatim from bin/install.js with no logic change. |
Agent Model
Orchestrator → Agent Pattern
Orchestrator (workflow .md)
│
├── Load context: gsd-tools.cjs init <workflow> <phase>
│ Returns JSON with: project info, config, state, phase details
│
├── Resolve model: gsd-tools.cjs resolve-model <agent-name>
│ Returns: opus | sonnet | haiku | inherit
│
├── Spawn Agent (Task/SubAgent call)
│ ├── Agent prompt (agents/*.md)
│ ├── Context payload (init JSON)
│ ├── Model assignment
│ └── Tool permissions
│
├── Collect result
│
└── Update state: gsd-tools.cjs state update / state patch / state advance-plan
Primary Agent Spawn Categories
Conceptual spawn-pattern taxonomy for the primary agents. For the authoritative agent roster (including the advanced/specialized agents such as gsd-pattern-mapper, gsd-code-reviewer, gsd-code-fixer, gsd-ai-researcher, gsd-domain-researcher, gsd-eval-planner, gsd-eval-auditor, gsd-framework-selector, gsd-debug-session-manager, gsd-intel-updater), see docs/INVENTORY.md.
| Category | Agents | Parallelism |
|---|---|---|
| Researchers | gsd-project-researcher, gsd-phase-researcher, gsd-ui-researcher, gsd-advisor-researcher | 4 parallel (stack, features, architecture, pitfalls); advisor spawns during discuss-phase |
| Synthesizers | gsd-research-synthesizer | Sequential (after researchers complete) |
| Planners | gsd-planner, gsd-roadmapper | Sequential |
| Checkers | gsd-plan-checker, gsd-integration-checker, gsd-ui-checker, gsd-nyquist-auditor | Sequential (verification loop, max 3 iterations) |
| Executors | gsd-executor | Parallel within waves, sequential across waves |
| Verifiers | gsd-verifier | Sequential (after all executors complete) |
| Mappers | gsd-codebase-mapper | 4 parallel (tech, arch, quality, concerns) |
| Debuggers | gsd-debugger | Sequential (interactive) |
| Auditors | gsd-ui-auditor, gsd-security-auditor | Sequential |
| Doc Writers | gsd-doc-writer, gsd-doc-verifier | Sequential (writer then verifier) |
| Profilers | gsd-user-profiler | Sequential |
| Analyzers | gsd-assumptions-analyzer | Sequential (during discuss-phase) |
Wave Execution Model
During execute-phase, plans are grouped into dependency waves:
Wave Analysis:
Plan 01 (no deps) ─┐
Plan 02 (no deps) ─┤── Wave 1 (parallel)
Plan 03 (depends: 01) ─┤── Wave 2 (waits for Wave 1)
Plan 04 (depends: 02) ─┘
Plan 05 (depends: 03,04) ── Wave 3 (waits for Wave 2)
Each executor gets:
- Fresh 200K context window (or up to 1M for models that support it)
- The specific PLAN.md to execute
- Project context (PROJECT.md, STATE.md)
- Phase context (CONTEXT.md, RESEARCH.md if available)
Adaptive Context Enrichment (1M Models)
When the context window is 500K+ tokens (1M-class models like Opus 4.6, Sonnet 4.6), subagent prompts are automatically enriched with additional context that would not fit in standard 200K windows:
- Executor agents receive prior wave SUMMARY.md files and the phase CONTEXT.md/RESEARCH.md, enabling cross-plan awareness within a phase
- Verifier agents receive all PLAN.md, SUMMARY.md, CONTEXT.md files plus REQUIREMENTS.md, enabling history-aware verification
The orchestrator reads context_window from config (gsd-tools.cjs config-get context_window) and conditionally includes richer context when the value is >= 500,000. For standard 200K windows, prompts use truncated versions with cache-friendly ordering to maximize context efficiency.
Parallel Commit Safety
When multiple executors run within the same wave, two mechanisms prevent conflicts:
--no-verifycommits — Parallel agents skip pre-commit hooks (which can cause build lock contention, e.g., cargo lock fights in Rust projects). The orchestrator runsgit hook run pre-commitonce after each wave completes.- STATE.md file locking — All
writeStateMd()calls use lockfile-based mutual exclusion (STATE.md.lockwithO_EXCLatomic creation). This prevents the read-modify-write race condition where two agents read STATE.md, modify different fields, and the last writer overwrites the other's changes. Includes stale lock detection (10s timeout) and spin-wait with jitter.
The STATE.md Write Path
Locking decides who writes. A separate contract decides what survives the write.
STATE.md carries the same fact in two places — YAML frontmatter and the document body — and the body is authoritative. Every write therefore re-derives frontmatter from the body, which raises the question the write path exists to answer: when a re-derived value disagrees with the one already in frontmatter, which wins?
FIELD_CLASSIFICATION (src/state-transition.cts) answers it per field, declaring a preservation policy — preserve-when-unchanged, preserve-always, preserve-if-placeholder, derive — that applyStatePreservation executes after syncStateFrontmatter re-derives. (A fifth policy, clear, was listed here until ADR-3408 §8.6's amendment removed it: no row used it and no executor existed for it.)
The pipeline's precondition is a type, not a convention (ADR-3473 §8.6). A policy row can only be honored if the pre-write frontmatter snapshot it compares against is actually present. That snapshot now travels as a StateTransaction, built by openStateTransaction() — preservation applies — or rebuildStateTransaction() — it does not. Both carry the snapshot, and a transaction cannot be constructed without one: an absent snapshot is a construction failure, not a runtime skip. That distinction is the whole point. Previously the snapshot was nulled to signal "re-derive from disk", so a declared preserve-always row and a silently-skipped one were indistinguishable at runtime, which is how a curated progress: block was erased by verbs that had nothing to do with progress.
rebuildStateTransaction() is the typed form of ADR-3408 §8.3's closed exception list: state sync, which exists to let the body win, and /gsd-health --repair's factory reset. Both are deliberate and permanent, not debt — and because the type names them, the write-path drift guard no longer has to track them as strings in a ratcheted baseline.
What a command reports it wrote is the same snapshot, read back (ADR-3473 §8.7). Every state.* command returns an updated array. That array is now derived by comparing what was actually persisted against the transaction's pre-write snapshot: a field appears if, and only if, its persisted value changed. One comparison answers both of the questions that used to need separate machinery — a field the caller asked for that the pipeline then discarded is persisted-equals-snapshot and drops out, and a field nobody asked for that the write moved anyway is different and appears. Nothing is filtered by its preservation policy.
Reporting is at leaf granularity: when a single counter moves you are told progress.total_plans, not progress. The leaves are the ones the field-classification table already declares, so the report is bounded by a schema rather than by walking the document.
One field is excluded, and it is excluded for its provenance rather than its policy: last_updated is stamped on every save regardless of what you changed, so admitting it would make state patch's success signal — which is simply whether updated is non-empty — permanently true, and a patch in which every field failed would report success. state_head is deliberately not excluded: it is recomputed on every save but only changes when the commit it records actually moved, so reporting it tells you something true.
A practical consequence worth knowing: these arrays are now longer than they used to be, because they used to under-report. If you compare one exactly, expect more entries — and expect them to be the ones that really changed.
ADR-3408 is the normative contract for that path: one executor per declared policy, one write seam, and reports computed from what was actually persisted rather than from what the caller intended to write. Where the contract and the code disagree, the code is the defect. It is the write-side counterpart of ADR-3180, which gave each read-side derivation a single owner.
ADR-3473 owns the invariants that sit outside that contract. ADR-3408 governs what survives a write; it does not govern the pipeline's precondition, what a command reports it wrote, or where the set of STATE.md keys, types and enums is declared. ADR-3473 owns those, alongside document parsing, enumeration, and the return contract of every routine that can fail. It is the third application of ADR-3180's mechanism and the first whose success metric requires the guard surface to shrink as each seam lands.
Data Flow
New Project Flow
User input (idea description)
│
▼
Questions (questioning.md philosophy)
│
▼
4x Project Researchers (parallel)
├── Stack → STACK.md
├── Features → FEATURES.md
├── Architecture → ARCHITECTURE.md
└── Pitfalls → PITFALLS.md
│
▼
Research Synthesizer → SUMMARY.md
│
▼
Requirements extraction → REQUIREMENTS.md
│
▼
Roadmapper → ROADMAP.md
│
▼
User approval → STATE.md initialized
Phase Execution Flow
discuss-phase → CONTEXT.md (user preferences)
│
▼
ui-phase → UI-SPEC.md (design contract, optional)
│
▼
plan-phase
├── Research gate (blocks if RESEARCH.md has unresolved open questions)
├── Phase Researcher → RESEARCH.md
│ └── Package Legitimacy Gate: registry-API verdict on every package; [SLOP] removed,
│ [SUS]/[ASSUMED] flagged; Audit table written to RESEARCH.md
├── Planner (with reachability check) → PLAN.md files
│ └── checkpoint:human-verify injected before [ASSUMED]/[SUS] installs;
│ T-{phase}-SC STRIDE row added for install-bearing plans
├── Plan Checker → Verify loop (max 3x)
├── Requirements coverage gate (REQ-IDs → plans)
└── Decision coverage gate (CONTEXT.md `<decisions>` → plans, BLOCKING — #2492)
│
▼
state planned-phase → STATE.md (Planned/Ready to execute)
│
▼
execute-phase (context reduction: truncated prompts, cache-friendly ordering)
├── Wave analysis (dependency grouping)
├── Executor per plan → code + atomic commits
├── SUMMARY.md per plan
└── Verifier → VERIFICATION.md
└── Decision coverage gate (CONTEXT.md decisions → shipped artifacts, NON-BLOCKING — #2492)
│
▼
verify-work → UAT.md (user acceptance testing)
│
▼
ui-review → UI-REVIEW.md (visual audit, optional)
Context Propagation
Each workflow stage produces artifacts that feed into subsequent stages:
PROJECT.md ────────────────────────────────────────────► All agents
REQUIREMENTS.md ───────────────────────────────────────► Planner, Verifier, Auditor
ROADMAP.md ────────────────────────────────────────────► Orchestrators
STATE.md ──────────────────────────────────────────────► All agents (decisions, blockers)
CONTEXT.md (per phase) ────────────────────────────────► Researcher, Planner, Executor
RESEARCH.md (per phase) ───────────────────────────────► Planner, Plan Checker
PLAN.md (per plan) ────────────────────────────────────► Executor, Plan Checker
SUMMARY.md (per plan) ─────────────────────────────────► Verifier, State tracking
UI-SPEC.md (per phase) ────────────────────────────────► Executor, UI Auditor
File System Layout
Installation Files
~/.claude/ # Claude Code (global install)
├── skills/gsd-ns-*/SKILL.md # Global skills — nesting runtimes: 6 namespace routers (authoritative roster: docs/INVENTORY.md)
│ └── skills/<name>/SKILL.md # concrete skills nested under each router
│ (flat runtimes: skills/gsd-*/SKILL.md — all ~67 skills at top level)
├── commands/gsd/*.md # Local Claude installs use slash commands instead of global skills
├── gsd-core/
│ ├── bin/gsd-tools.cjs # CLI utility
│ ├── bin/lib/*.cjs # Domain modules (authoritative roster: docs/INVENTORY.md)
│ ├── workflows/*.md # Workflow definitions (authoritative roster: docs/INVENTORY.md)
│ ├── references/*.md # Shared reference docs (authoritative roster: docs/INVENTORY.md)
│ └── templates/ # Planning artifact templates
├── agents/*.md # Agent definitions (authoritative roster: docs/INVENTORY.md)
├── hooks/*.js # Node.js hooks (statusline, guards, monitors, update check)
├── hooks/*.sh # Shell hooks (session state, commit validation, phase boundary)
├── settings.json # Hook registrations
└── VERSION # Installed version number
Equivalent paths for other runtimes:
- OpenCode:
~/.config/opencode/global or./.opencode/local - Kilo:
~/.config/kilo/global or./.kilo/local - Kimi CLI: first-existing generic global root (
~/.config/agents/recommended, then~/.agents/if itsskills/directory already exists); local install is deferred and guarded - Codex:
~/.codex/global or./.codex/local - Copilot:
~/.copilot/global or./.github/local - Antigravity: auto-detected global root (
~/.gemini/antigravity/,~/.gemini/antigravity-ide/, or~/.gemini/antigravity-cli/) for settings and runtime files; global skills/agents under~/.gemini/config/(the machine-local discovery dir, #3738) or./.agent/local - Cursor:
~/.cursor/global or./.cursor/local - Windsurf/Devin Desktop:
~/.codeium/windsurf/global config or./.windsurf/local workflows - Augment Code:
~/.augment/global or./.augment/local - Trae:
~/.trae/global or./.trae/local - Qwen Code:
~/.qwen/global or./.qwen/local - Hermes Agent:
~/.hermes/global or./.hermes/local - CodeBuddy:
~/.codebuddy/global or./.codebuddy/local - Cline:
~/.cline/global or project-root.clineruleslocal
Project Files (.planning/)
.planning/
├── PROJECT.md # Project vision, constraints, decisions, evolution rules
├── REQUIREMENTS.md # Scoped requirements (v1/v2/out-of-scope)
├── ROADMAP.md # Phase breakdown with status tracking
├── STATE.md # Living memory: position, decisions, blockers, metrics
├── config.json # Workflow configuration
├── MILESTONES.md # Completed milestone archive
├── research/ # Domain research from /gsd-new-project
│ ├── SUMMARY.md
│ ├── STACK.md
│ ├── FEATURES.md
│ ├── ARCHITECTURE.md
│ └── PITFALLS.md
├── codebase/ # Brownfield mapping (from /gsd-map-codebase or /gsd-onboard)
├── onboarding/ # Brownfield onboarding summary (from /gsd-onboard)
│ ├── STACK.md # YAML frontmatter carries `last_mapped_commit`
│ ├── ARCHITECTURE.md # for the post-execute drift gate (#2003)
│ ├── CONVENTIONS.md
│ ├── CONCERNS.md
│ ├── STRUCTURE.md
│ ├── TESTING.md
│ └── INTEGRATIONS.md
├── phases/
│ └── XX-phase-name/
│ ├── XX-CONTEXT.md # User preferences (from discuss-phase)
│ ├── XX-RESEARCH.md # Ecosystem research (from plan-phase)
│ ├── XX-YY-PLAN.md # Execution plans
│ ├── XX-YY-SUMMARY.md # Execution outcomes
│ ├── XX-VERIFICATION.md # Post-execution verification
│ ├── XX-VALIDATION.md # Nyquist test coverage mapping
│ ├── XX-UI-SPEC.md # UI design contract (from ui-phase)
│ ├── XX-UI-REVIEW.md # Visual audit scores (from ui-review)
│ └── XX-UAT.md # User acceptance test results
├── quick/ # Quick task tracking
│ └── YYMMDD-xxx-slug/
│ ├── PLAN.md
│ └── SUMMARY.md
├── todos/
│ ├── pending/ # Captured ideas
│ └── completed/ # Completed todos
├── threads/ # Persistent context threads (from /gsd-thread)
├── seeds/ # Forward-looking ideas (from /gsd-capture --seed)
├── debug/ # Active debug sessions
│ ├── *.md # Active sessions
│ ├── resolved/ # Archived sessions
│ └── knowledge-base.md # Persistent debug learnings
├── ui-reviews/ # Screenshots from /gsd-ui-review (gitignored)
└── continue-here.md # Context handoff (from pause-work)
Post-Execute Codebase Drift Gate (#2003)
After the last wave of /gsd-execute-phase commits, the workflow runs a
non-blocking codebase_drift_gate step (between schema_drift_gate and
verify_phase_goal). It compares the diff last_mapped_commit..HEAD
against .planning/codebase/STRUCTURE.md and counts four kinds of
structural elements:
- New directories outside mapped paths
- New barrel exports at
(packages|apps)/<name>/src/index.* - New migration files
- New route modules under
routes/orapi/
If the count meets workflow.drift_threshold (default 3), the gate either
warns (default) with the suggested /gsd-map-codebase --paths … command,
or auto-remaps (workflow.drift_action = auto-remap) by spawning
gsd-codebase-mapper scoped to the affected paths. Any error in detection
or remap is logged and the phase continues — drift detection cannot fail
verification.
last_mapped_commit lives in YAML frontmatter at the top of each
.planning/codebase/*.md file; bin/lib/drift.cjs provides
readMappedCommit and writeMappedCommit round-trip helpers.
Installer Architecture
The installer (bin/install.js, ~10,700 lines) handles:
- Runtime detection — Interactive prompt or CLI flags (
--claude,--opencode,--kimi,--kilo,--codex,--copilot,--antigravity,--cursor,--windsurf,--augment,--trae,--qwen,--hermes,--codebuddy,--cline,--all) - Location selection — Global (
--global) or local (--local) - File deployment — Copies commands, skills, workflows, references, templates, agents, and hooks
- Runtime adaptation — Transforms file content per runtime:
- Claude Code: Uses as-is
- OpenCode: Converts commands/agents to OpenCode-compatible flat command + subagent format
- Kilo: Reuses the OpenCode conversion pipeline with Kilo config paths
- Codex: Generates TOML config + skills from commands
- Kimi CLI: Generates Agent Skills under
skills/gsd-*/SKILL.md, custom agent YAML/prompt files, and explicitkimi_cli.tools.*module paths - Copilot: Maps tool names (Read→read, Bash→execute, etc.)
- Antigravity: Skills-first with Google model equivalents; adjusts hook event names (
AfterToolinstead ofPostToolUse) - Cursor: Skills-first with Cursor rule references
- Windsurf: Skills-first with Windsurf rule references
- Trae: Skills-first install to
~/.trae/./.traewith nosettings.jsonor hook integration - Qwen Code: Skills-first with Qwen-branded path and prompt rewrites
- Hermes Agent: Category-based skills under
skills/gsd/ - CodeBuddy: Skills-first with CodeBuddy path and prompt rewrites
- Cline: Writes
.clinerulesfor rule-based integration - Augment Code: Skills-first with full skill conversion and config management
- Path normalization — Replaces
~/.claude/paths with runtime-specific paths - Settings integration — Registers hooks in runtime's
settings.json - Patch backup — Since v1.17, backs up locally modified files to
gsd-local-patches/for/gsd-update --reapply - Manifest tracking — Writes
gsd-file-manifest.jsonfor clean uninstall. The manifest also records whichruntimeand whichscope(global/local) wrote it, under amanifestVersionschema field, so a reader can answer "which surfaces are installed, at which scopes" without inferring it from the directory the file sits in (ADR 2866, #2872). Manifests written before that carry no such fields and are read without error — no reinstall is required. See Installer Migrations → File Manifest - Uninstall mode —
--uninstallremoves all GSD files, hooks, and settings
installRuntimeArtifacts (install-engine.cjs) returns the executed plan it ran — per kind, per
scope, including on the combined OpenCode/Kilo family path, which previously early-returned void —
rather than being observable only by re-reading disk afterward. Its destination-writing IO (copies,
removals, snapshot/restore, best-effort cleanup) now routes through an injectable fs seam,
install-fs-adapter.cjs, so a full install can be exercised against a fake adapter with zero real
destination IO; locating this package's own source tree remains real by design (a destination-fake
is never seeded with the repo's own paths). Writes stay byte-identical and existing void-ignoring
callers are unaffected. This completes ADR 58's
registry → adapter → helpers → cleanup rollout — the cleanup step had not previously landed
(#2874, epic #2866 Phase 5).
Install-time file moves, stale-artifact cleanup, config rewrites, and user-data preservation are governed by the Installer Migration Module. See Installer Migrations and ADR 0008. The migration module also owns the gated first-time baseline scan for legacy installs, classifying known runtime install surfaces before later migrations remove or rewrite anything.
The plan drift guard (plan_review.source_grounding) — which verifies symbol references in generated plans against live source before execution — is specified in ADR 22.
The same switch gates a second, cross-artifact axis: a fact-drift pass that compares the same fact as stated in ROADMAP.md, PLAN.md, STATE.md and CONTEXT.md and reports contradictions (a phase status, a success criterion, a requirement ID, a glossary term) with both locations and the authoritative side named. Where the source-grounding axis grounds a plan against code, this one grounds the planning artifacts against each other. It keys on contradicting knowledge rather than similar-looking text, and is advisory only — it never sets hardBlock and never contributes to the convergence counts.
Platform Handling
- Windows:
windowsHideon child processes, EPERM/EACCES protection on protected directories, path separator normalization - WSL: Detects Windows Node.js running on WSL and warns about path mismatches
- Docker/CI: Supports
CLAUDE_CONFIG_DIRenv var for custom config directory locations
Hook System
Architecture
Runtime Engine (Claude Code / Antigravity CLI)
│
├── statusLine event ──► gsd-statusline.js
│ Reads: stdin (session JSON)
│ Writes: stdout (formatted status), /tmp/claude-ctx-{session}.json (bridge)
│
├── PostToolUse/AfterTool event ──► gsd-context-monitor.js
│ Reads: stdin (tool event JSON), /tmp/claude-ctx-{session}.json (bridge)
│ Writes: stdout (hookSpecificOutput with additionalContext warning)
│
└── SessionStart event
├──► gsd-ensure-canonical-path.js (runs first)
│ Reads: ${CLAUDE_PLUGIN_ROOT}/gsd-core/ (plugin installs only)
│ Writes: ~/.claude/gsd-core/{bin,contexts,references,templates,workflows} symlinks
│ (no-op in classic installs; preserves user files; self-heals)
└──► gsd-check-update.js
Reads: VERSION file
Writes: ~/.claude/cache/gsd-update-check.json (spawns background process)
Context Monitor Thresholds
| Remaining Context | Level | Agent Behavior |
|---|---|---|
| > 35% | Normal | No warning injected |
| ≤ 35% | WARNING | "Avoid starting new complex work" |
| ≤ 25% | CRITICAL | "Context nearly exhausted, inform user" |
Debounce: 5 tool uses between repeated warnings. Severity escalation (WARNING→CRITICAL) bypasses debounce.
Safety Properties
- All hooks wrap in try/catch, exit silently on error
- stdin timeout guard (3s) prevents hanging on pipe issues
- Stale metrics (>60s old) are ignored
- Missing bridge files handled gracefully (subagents, fresh sessions)
- Context monitor is advisory — never issues imperative commands that override user preferences
Package Legitimacy Gate (v1.42.1)
The researcher → planner → executor pipeline includes a supply-chain gate against slopsquatting (AI-hallucinated package names pre-registered with malicious post-install scripts).
Threat model: GSD automates the full path from "researcher names a package" to "executor runs npm install". A hallucinated name that passes npm view (proving only registration, not legitimacy) would previously flow through undetected. ~20% of AI-generated package references are hallucinated; ~43% of those names recur consistently across prompts, making pre-registration economically viable for attackers.
Gate layers:
| Layer | Component | Action |
|---|---|---|
| Research | gsd-phase-researcher |
Runs gsd-tools query package-legitimacy check --ecosystem <npm|pypi|crates> <pkgs>; writes ## Package Legitimacy Audit table to RESEARCH.md; strips [SLOP] packages before RESEARCH.md is written |
| Planning | gsd-planner |
Reads Audit table; inserts checkpoint:human-verify before any [ASSUMED] or [SUS] install task; adds T-{phase}-SC STRIDE supply-chain row to <threat_model> |
| Execution | gsd-executor |
RULE 3 excludes package installation from auto-fix scope; failed installs surface as checkpoints, never silent substitutions |
Claim provenance integration: Package names discovered via WebSearch are tagged [ASSUMED] (not [VERIFIED]) regardless of the registry-API verdict. This extends the existing [ASSUMED] / [VERIFIED] / [CITED] provenance system by enforcing the provenance tag as a hard gate at the install boundary — [ASSUMED] always generates a checkpoint:human-verify in PLAN.md.
Ecosystem coverage: The gate resolves signals directly from each ecosystem's registry API rather than a single generic check — registry.npmjs.org + api.npmjs.org/downloads (Node), pypi.org/pypi/<pkg>/json (Python), the crates.io API (Rust). This catches cross-ecosystem hallucination (~9% rate documented in 2025 USENIX research).
Graceful degradation: Each registry adapter degrades to null signals (never throws) on a failed lookup; missing signals push a package to [SUS], which is gated behind the same checkpoint:human-verify checkpoint as [ASSUMED]. Research and planning proceed; the system never hard-fails on a network or tool outage. slopcheck is an optional escalate-only adapter — it can only raise a verdict, never lower it, and is not the install-or-degrade gate. No shipped configuration wires it.
Security Hooks (v1.27)
For a conceptual overview of how the hook and guard layers fit into the broader security approach, see Security model.
Prompt Guard (gsd-prompt-guard.js):
- Triggers on Write/Edit to
.planning/files - Scans content for prompt injection patterns (role override, instruction bypass, system tag injection)
- Advisory-only — logs detection, does not block
- Patterns are inlined (subset of
security.cjs) for hook independence
Read Injection Scanner (gsd-read-injection-scanner.js):
- Triggers on
Read/WebFetch/WebSearchPostToolUse events - Advisory by default; blocks only
HIGHseverity, and only whensecurity.injection_blockingistrue - Severity is
LOWfor 1-2 matched patterns,HIGHfor 3 or more - Skips content shorter than 20 characters, and skips excluded paths (
.planning/,REVIEW.md,CHECKPOINT*, security/injection docs, and GSD's own staged hook bundle) - Rule ids: the
MD-LINK-*markdown-link rules mirrored fromsecurity.cjs'sMARKDOWN_LINK_PATTERNS, plusINJECTION-PATTERN,INVISIBLE-UNICODE, andUNICODE-TAG-BLOCK - Patterns are shared with
gsd-prompt-guard.jsviahooks/lib/injection-patterns.js(#3504); the markdown-link list is inlined for hook independence - Output contract:
hookSpecificOutputcarries bothadditionalContext(the human-readable advisory sentence) andfindings— an array of{ ruleId, match }records naming each rule that fired.findingsis the structured surface; the advisory is rendered from it, so the two cannot disagree.matchisnullfor rules with no captured text (INVISIBLE-UNICODE,UNICODE-TAG-BLOCK). Consumers should readfindingsrather than parsing the advisory text.
Workflow Guard (gsd-workflow-guard.js):
- Triggers on Write/Edit to non-
.planning/files - Detects edits outside GSD workflow context (no active
/gsd-command or Task subagent) - Advises using
/gsd-quickor/gsd-fastfor state-tracked changes - Opt-in via
hooks.workflow_guard: true(default: false)
Runtime Abstraction
GSD supports multiple AI coding runtimes through a unified command/workflow architecture:
Runtime Install Contract Matrix
This matrix describes the runtime surfaces the installer materializes today. The migration-specific ownership and source snapshots live in Installer Migrations.
| Runtime | Global root | Local root | Invocation surface | Agent surface | Config and hooks |
|---|---|---|---|---|---|
| Claude Code | ~/.claude |
./.claude |
Global skills/gsd-*/SKILL.md (flat, #924); local commands/gsd/*.md |
agents/gsd-*.md |
settings.json hook and statusLine entries |
| OpenCode | ~/.config/opencode |
./.opencode |
commands/gsd-*.md |
agents/gsd-*.md |
opencode.json or opencode.jsonc; no GSD hooks |
| Kilo | ~/.config/kilo |
./.kilo |
command/gsd-*.md |
agents/gsd-*.md |
kilo.json or kilo.jsonc; no GSD hooks |
| Kimi CLI | First-existing generic root: ~/.config/agents recommended, then ~/.agents when ~/.agents/skills exists and ~/.config/agents/skills does not |
Deferred and guarded | skills/gsd-*/SKILL.md (flat) invoked as /skill:gsd-* |
agents/gsd.yaml, agents/gsd.md, and agents/subagents/gsd-* YAML/prompt pairs |
Explicit kimi --agent-file <configRoot>/agents/gsd.yaml; no GSD hooks or statusline |
| Codex | ~/.codex |
./.codex |
skills/gsd-*/SKILL.md (flat) |
agents/ source markdown plus per-agent TOML (Codex auto-discovers each agents/gsd-*.toml; this is the sole canonical role registration, #2406) |
config.toml bare [agents] dispatch-tuning scalar (max_depth, no per-role [agents.gsd-*] tables), [features].hooks (canonical; legacy alias codex_hooks is recognized and migrated forward on reinstall, #3566), and hook tables |
| GitHub Copilot | ~/.copilot |
./.github |
skills/gsd-*/SKILL.md (flat), copilot-instructions.md, and AGENTS.md (repo root, local) |
.agent.md files |
Self-contained sessionStart hook (hooks/gsd-session.json, inline command type); no statusline |
| Antigravity | auto-detected: ~/.gemini/antigravity, ~/.gemini/antigravity-ide, or ~/.gemini/antigravity-cli |
./.agent |
~/.gemini/config/skills/gsd-*/SKILL.md (flat, #1614; global home override #3738) |
~/.gemini/config/agents/gsd-*.md (#3738) |
Gemini-style settings.json hook entries when installed by GSD |
| Cursor | ~/.cursor |
./.cursor |
skills/gsd-*/SKILL.md (flat) |
agents/gsd-*.md |
Rule references under rules/; hooks.json with sessionStart context injection and postToolUse STATE.md monitor (#777) |
| Windsurf | ~/.codeium/windsurf config |
./.windsurf |
workflows/gsd-*.md slash-command workflows |
No custom-agent artifact surface | No GSD hooks |
| Augment Code | ~/.augment |
./.augment |
skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
No GSD hooks or statusline |
| Trae | ~/.trae |
./.trae |
skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
Rule references under rules/; no GSD hooks |
| Qwen Code | ~/.qwen |
./.qwen |
skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
Common GSD settings and hook entries where supported |
| Hermes Agent | ~/.hermes |
./.hermes |
skills/gsd/ns-*/SKILL.md (6 routers, prefix='') + skills/gsd/ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
Common GSD settings and hook entries where supported |
| CodeBuddy | ~/.codebuddy |
./.codebuddy |
skills/gsd-*/SKILL.md (flat, user-invocable: false) |
agents/gsd-*.md |
/gsd-* slash commands under commands/; common GSD settings and hook entries where supported |
| Cline | ~/.cline |
project root | skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) + .clinerules |
Rules only | No GSD hooks or statusline |
Upstream Contract Sources
Runtime install expectations are checked against primary documentation where available. The current source snapshot is 2026-05-11, with Kimi CLI rechecked on 2026-06-07:
- Claude Code: Anthropic slash commands, settings, hooks, and subagents docs.
- OpenCode and Kilo: OpenCode config docs and Kilo custom subagent docs.
- Qwen Code: command/config docs; Qwen command docs were last updated 2026-05-06.
- Kimi CLI: Agent Skills docs for user-level brand roots and first-existing
generic roots (
~/.config/agents/skills/recommended, then~/.agents/skills/), plus Agents docs for YAML files,system_prompt_path,kimi_cli.tools.*module paths, and explicitkimi --agent-filelaunch. - Codex: OpenAI Codex docs and
config-schema.json; the installer also carries Codex 0.124.0 compatibility for agent table shape. - Copilot, Cursor, Cline, Augment, Hermes, and CodeBuddy: vendor docs for custom instructions, rules, skills, or config.
- Antigravity, Windsurf, and Trae: source-limited rows. The installer documents current compatibility shims, and migrations must refresh those sources before rewriting their config.
Abstraction Points
- Tool name mapping — Each runtime has its own tool names (e.g., Claude's
Bash→ Copilot'sexecute) - Hook event names — Claude uses
PostToolUse, Antigravity usesAfterTool - Agent frontmatter — Each runtime has its own agent definition format
- Path conventions — Each runtime stores config in different directories
- Model references —
inheritprofile lets GSD defer to runtime's model selection
The installer handles all translation at install time. Workflows and agents are written in Claude Code's native format and transformed during deployment.