* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI gsd-ui-auditor is chartered to audit interaction and handed a capture driver with no interaction verb: `npx playwright screenshot` cannot click, fill, hover, press or snapshot, so a hover state, an open menu, a focus ring or a form's validation state never appears in its evidence and every Experience Design finding degrades to code reading. Implements the shape approved at triage, not a new capability: - capabilities/ui/capability.json declares `workflow.ui_interaction_capture` (boolean, default false) on the capability that already owns the auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated. - gsd-core/workflows/ui-review.md reads the key through gsd_run and hands it to the auditor as `interaction_capture:` in the spawn <config> block — the auditor carries no gsd_run resolver, so the key travels by value. - agents/gsd-ui-auditor.md gains an anchored interaction-capture section AFTER the static block. With the key on and a Chrome binary resolved it starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on an --isolated profile, opens the dev URL the static block reached, takes the a11y snapshot for element uids, captures the baseline and a Tab focus-ring state, drives the UI-SPEC's interactive components, saves console output, and stops the daemon unconditionally. Key off, no dev server, or no Chrome: one status line, and the Playwright-only static path runs exactly as before — the static fence is untouched. Needs only Bash: no MCP server, no tools: change. Chromium-only by nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so readiness is polled through evaluate_script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): bind the interaction-capture shape and containment - manifest, generated registry, config schema and config-set/loadConfig all know workflow.ui_interaction_capture as a default-off boolean, and hand-written non-booleans fall to the slice default - the orchestrator reads the key and hands it down; the auditor never grows a gsd_run dependency - the static fence stays Playwright-only and the interaction fence chrome-devtools-only, so key-off is today's path - the interaction fence runs under bash with a stub driver on PATH: key off / absent / no dev server / no Chrome invoke nothing; the happy path starts first and stops last on the [selected] pageId with the documented flags; a failed capture is removed and not counted; new_page and start failures still honour the stop-only-if-started rule; CHROME_BIN and CHROME_DEVTOOLS_MCP_VERSION overrides flow through - docs/CONFIGURATION.md row shape; registered in the docs-guard lane Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * docs(#4223): document workflow.ui_interaction_capture and its how-to - docs/CONFIGURATION.md: one row in the workflow.* table, default-off - docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the interaction-capture section adds, skips and never claims - docs/how-to/enable-ui-interaction-capture.md: turn it on, read the `**Interaction captures:**` outcomes, what it does not do, turn it off - docs/README.md: index the how-to beside live-DOM verification Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): add changeset Added-type fragment; pr: carries the issue number until the PR exists. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace; the hyphen form is retired there and the slash-command-namespace guard rejects it. docs/ keep the hyphen form by convention. Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered. Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): per-run daemon session, bounded navigation, and step failures that count Three findings from the pre-file adversarial review of the interaction fence, folded in: - `--sessionId <epoch>-<pid>` on every driver call. `start` restarts whatever daemon shares its session and --isolated isolates only the browser profile, so two concurrent audits — or an audit beside the operator's own CLI daemon — would otherwise stop each other. The CLI accepts hex and dashes only; the id is validated by the test stub. - `new_page --timeout 30000`: the one verb that takes a bound, placed before every verb that does not, so a hung page is caught first. - a failed take_snapshot or press_key now increments the failure count and is named on stdout; two clean screenshots can no longer read as `0 failed` after the step that gives the interactions their uids failed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal Second review round, both reviewers: - session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a subshell, so two audits forked from one parent in the same second shared an id and could stop each other's daemon (driven by the reviewer) - `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver under Git Bash still matches the `$` anchor, and `|| true` on the assignment so a failed new_page cannot abort the block under `set -e -o pipefail` before the unconditional stop - a failed take_snapshot removes any snapshot.txt it left or inherited from a reused directory, so stale uids never drive the interactions - `<config>` placeholder is `{interaction_capture}`, lowercase like its `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt template, not a bash heredoc - how-to: the `not captured` row no longer claims the daemon started Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges Third review round: - a new_page that prints a page line and then exits non-zero is a failed navigation, not a page id: the exit status is checked in an `if` before the output is parsed (driven by the reviewer against the previous `|| true`, which masked exactly that) - regression cases for what the last two rounds added: CRLF driver output, a stale snapshot removed on failure, partial-output new_page failure, and the whole fence under `set -e -o pipefail` (both the failed-navigation path and the happy path) - the harness whitelist gains `date`; the session-id assertion now requires all three parts, so a silently empty epoch cannot hide again Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder - the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes, over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces; the interaction section's comments are tightened to the same content in fewer bytes (23559 now). No bash changed — the fence's own tests and the real-browser run are unchanged. - .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which the post-create backfill rewrites to the PR's own number. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): set changeset fragment pr to 4477 * test(#4223): compare the fence's status path with the separator the fence uses On the windows-latest lane the happy-path case failed on `\interaction` vs `/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a literal slash, and the assertion built its expectation with path.join. Every other case in the file passed on that lane, including the CRLF and errexit/pipefail ones. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): drop the inert file-header allow-test-rule marker Review round 1 on #4477: the `source-text-is-the-product` marker sat at line 2, outside no-source-grep's 8-line lookahead of every readFileSync site (the first is ~60 lines down), so it suppressed nothing. It was also unnecessary: every read in this file is a .md/.json path, which the rule does not trigger on. Deleted rather than relocated — there is no site to relocate it to. Negative control: `eslint` on the file is clean without it. * chore(#4223): regenerate the platform-conformance-tier lists for the new test Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier gate after this branch opened; its two committed lists must name every file under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never been in them. Once the branch was updated against next the lists were stale and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier --check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's fragment-single-edit-propagation, which sees the same staleness as regen:derived touching files beyond the fragment edit under test. Regenerated with the repo's own generators, no hand-editing. The general tier goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry; both --check arms are clean. Verified the red is this PR's own file and not base drift: at upstream/next both generators report "list matches" (546 / 196), and our committed copies were byte-identical to next's before this commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ * fix(#4223): bound, confine and trap the chrome-devtools driver fence Round 4 — three findings in one fence, interleaved on the same lines, so one commit: - Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a background job in its own process group (`set -m`) under a watchdog that kills the whole group at the ceiling — TERM, then KILL two seconds later. One pid is not enough: npm forwards SIGTERM only to its direct child, so killing `npx` alone leaves the client holding the fence's stdout and a `$(cdt … new_page)` capture blocked past the ceiling (driven against a real npx tree by the round's adversarial review; the pid-only first cut of this commit had exactly that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )` subshell: a subshell inherits bash's saved copies of the caller's stdio (the fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a watchdog outlived its kill — measured as the intermittent 30 s test run the review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 -- -pgid`, every 0.1 s) and stands down by itself once the group is empty; nothing ever signals it. The group, not the leader pid: a child can outlive the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll stood down at once and left the substitution open-ended (driven by the round's adversarial review at 6× the ceiling; a pgid cannot be reused while any member lives, which a bare pid can). The daemon `start` launches is spawned detached (its own session) and never in that group. Two platforms forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM; sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git Bash a signal to a watchdog still starting up hung the fence's `wait` for it: 18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs, while a fence slowed by xtrace, or three tests run alone, never hit it (a startup race; the mechanism is not pinned further). Polling: 0/300 slow calls and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3 full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU, msys and busybox sleep; BSD sleep documents it). A clock that cannot launch (`sleep … || exit 0`) stands the watchdog down rather than firing at once and killing a healthy call — by design that leaves a hung call unbounded, the pre-round-4 behaviour, instead of failing a healthy one. A hung call returns once its group is gone: at the ceiling, plus up to the 2 s TERM-to-KILL grace. The KILL after the grace is sent only to a group that is still alive: a pgid freed during the grace can be reused, and an unconditional KILL could hit an unrelated group (the round's review). `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog rather than either. - --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may write under the run's interaction/ directory and nowhere else. Relative, like every --filePath (unchanged from rounds 1-3): the daemon resolves both against one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and path.resolve()s both), and a relative path needs no dialect translation — an absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native daemon cannot resolve (CI's windows conformance shard caught the first cut). --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified), so the documented floor moves from ^1.8.0 to ^1.9.0, where --allowUnrestrictedPaths is deprecated. - `stop` is owed by an EXIT trap after a successful `start`, not by position (it replaces any earlier EXIT trap — none exists in this file); the explicit call keeps it in order, a flag makes the trap a no-op afterwards, and only the shell that installed the trap may act: a subshell copy of the fence state carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced locally at 3/40 under load: the second `stop` came from a subshell pid, never main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist and CI's macos conformance job showed the guard comparing empty to empty. The fence was driven under bash 3.2.57 for the injected-subshell, errexit failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy paths. A failed resize_page is a counted failed step now, not the one bare command an errexit runner could abort on. Prose in the section is tightened to pay for the mechanism: 23559 -> 24517 bytes against the 24576 DEFAULT-tier cap. Tests: the stub driver hangs as a real child tree (sh waiting on a child that holds stdout — never an exec), so a pid-only kill fails the new aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative- controlled: it blocks for the harness's whole cap on the old wrapper). A hung start and a hung capture are cut off within ceiling + grace + slack and still reach stop; an injected bare failure under errexit reaches stop through the trap, exactly once; an injected subshell call of cdt_stop issues nothing; the happy path issues exactly one stop; every driver call site names a ceiling and the only bare $CDT is the wrapper's own spawn; the start line carries --workspace with the capture directory, every --filePath lies under it, and no code line carries --allowUnrestrictedPaths. A driver whose leader exits at once while a child keeps holding the capture pipe is still cut off at the ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone (negative-controlled: the trap form kills `start` in under 20 ms). The harness EXPORTS its stub-only PATH — unexported, the exec'd watchdog fell through to bash's compiled-in default PATH and never saw the stub dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash, pid-preserving). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the accessibility tree, with entered form values) and console.txt (which can carry tokens) were committable by `git add .`. The gate now ignores `interaction/` as a directory — the next artifact type is covered by construction — and it appends whatever an existing .gitignore lacks instead of writing once. The write-once form was the same defect one step later: every project that had already run an audit would never have received the new pattern at all. Tests run the gate fence under bash: a fresh file carries every pattern; an image-only file from an earlier audit gains interaction/ and keeps its own header without duplicating present lines; a second run appends nothing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * test(#4223): declare the interaction-capture anchor as a comment marker The #4324 colon-token gate (slash-command-namespace) landed on next after this branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an unconvertible /gsd: command token. It is a section anchor of the same family as gsd:live-dom-families and gsd:write-continue, so it is declared in COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added gates against the merged tree; CI at ca8d2508 predates the gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: CI Rebase Check <ci@gsd-redux>
48 KiB
GSD Agent Reference
Full role cards for 22 primary agents plus concise stubs for 12 advanced/specialized agents (34 shipped agents total). The
agents/directory anddocs/INVENTORY.mdare the authoritative roster; see Architecture for context.
Overview
GSD uses a multi-agent architecture where thin orchestrators (workflow files) spawn specialized agents with fresh context windows. Each agent has a focused role, limited tool access, and produces specific artifacts.
Required reading (#3423): the canonical spawn-block tag is <required_reading> on BOTH sides — orchestrators emit it, and gating agents enforce it ("you MUST use the Read tool to load every file listed there before performing any other actions"). The legacy <files_to_read> emit-tag is retired and banned repo-wide by tests/agent-required-reading-consistency.test.cjs, because a mismatched pair silently disarms the enforcement clause.
Agent Categories
The table below covers the 22 primary agents detailed in this section. Thirteen additional shipped agents (pattern-mapper, debug-session-manager, code-reviewer, code-fixer, ai-researcher, domain-researcher, eval-planner, eval-auditor, framework-selector, intel-updater, doc-classifier, doc-synthesizer, mempalace-curator) have concise stubs in the Advanced and Specialized Agents section below. For the authoritative 35-agent roster, see
docs/INVENTORY.mdand theagents/directory.
| Category | Count | Agents |
|---|---|---|
| Researchers | 3 | project-researcher, phase-researcher, ui-researcher |
| Analyzers | 2 | assumptions-analyzer, advisor-researcher |
| Synthesizers | 1 | research-synthesizer |
| Planners | 1 | planner |
| Roadmappers | 1 | roadmapper |
| Executors | 1 | executor |
| Checkers | 3 | plan-checker, integration-checker, ui-checker |
| Verifiers | 2 | verifier, dom-verifier |
| Auditors | 3 | nyquist-auditor, ui-auditor, security-auditor |
| Mappers | 1 | codebase-mapper |
| Debuggers | 1 | debugger |
| Doc Writers | 2 | doc-writer, doc-verifier |
| Profilers | 1 | user-profiler |
Agent Details
gsd-project-researcher
Role: Researches domain ecosystem before roadmap creation.
| Property | Value |
|---|---|
| Spawned by | /gsd-new-project, /gsd-new-milestone |
| Parallelism | 4 instances (stack, features, architecture, pitfalls) |
| Tools | Read, Write, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__, mcp__firecrawl__, mcp__exa__, mcp__tavily__, mcp__ref__, mcp__jina__, mcp__perplexity__ |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | .planning/research/STACK.md, FEATURES.md, ARCHITECTURE.md, PITFALLS.md |
Capabilities:
- Web search for current ecosystem information
- Context7 MCP integration for library documentation
- Writes research documents directly to disk (reduces orchestrator context load)
gsd-phase-researcher
Role: Researches how to implement a specific phase before planning.
| Property | Value |
|---|---|
| Spawned by | /gsd-plan-phase |
| Parallelism | 4 instances (same focus areas as project researcher) |
| Tools | Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__, mcp__firecrawl__, mcp__exa__, mcp__tavily__, mcp__ref__, mcp__jina__, mcp__perplexity__ |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | {phase}-RESEARCH.md |
Capabilities:
- Reads CONTEXT.md to focus research on user's decisions
- Investigates implementation patterns for the specific phase domain
- Detects test infrastructure for Nyquist validation mapping
- Tags in-repo discrete values (enums, schema unions, error codes, status constants, paths)
[VERIFIED]only after reading the source-of-truth file that run, citing path and line range, and quoting the values verbatim - Refuses
[VERIFIED]for a compatibility claim resting on missing metadata (nopython_requires, noenginesfield, no per-version classifier, no changelog entry, no matching support-matrix row) — an absence constrains no version, and an enumerated allow-list that stops short of the target is still an absence, so only a positive falsification attempt with its failing output pasted earns the tag; anything less stays[ASSUMED]
gsd-ui-researcher
Role: Produces UI design contracts for frontend phases.
| Property | Value |
|---|---|
| Spawned by | /gsd-ui-phase |
| Parallelism | Single instance |
| Tools | Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__, mcp__firecrawl__, mcp__exa__, mcp__tavily__, mcp__ref__, mcp__jina__* |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | {phase}-UI-SPEC.md |
Capabilities:
- Detects design system state (shadcn components.json, Tailwind config, existing tokens)
- Offers shadcn initialization for React/Next.js/Vite projects
- Asks only unanswered design contract questions
- Enforces registry safety gate for third-party components
- Enumerates the component inventory rather than recalling it (#2845): the UI-SPEC's
## Component Inventorycarries a provenance line — the command that enumerated it, the count it returned, the resolved<package>@<version>, and the date — or, when nothing can enumerate it, aCould not enumerate: <reason>record in the same slot
gsd-assumptions-analyzer
Role: Deeply analyzes codebase for a phase and returns structured assumptions with evidence, confidence levels, and consequences if wrong.
| Property | Value |
|---|---|
| Spawned by | discuss-phase-assumptions workflow (when workflow.discuss_mode = 'assumptions') |
| Parallelism | Single instance |
| Tools | Read, Bash, Grep, Glob, Skill |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | Structured assumptions with decision statements, evidence file paths, confidence levels |
Key behaviors:
- Reads ROADMAP.md phase description and prior CONTEXT.md files
- Searches codebase for files related to the phase (components, patterns, similar features)
- Reads 5-15 most relevant source files to form evidence-based assumptions
- Classifies confidence: Confident (clear from code), Likely (reasonable inference), Unclear (could go multiple ways)
- Flags topics that need external research (library compatibility, ecosystem best practices)
- Output calibrated by tier: full_maturity (3-5 areas), standard (3-4), minimal_decisive (2-3)
gsd-advisor-researcher
Role: Researches a single gray area decision during discuss-phase advisor mode and returns a structured comparison table.
| Property | Value |
|---|---|
| Spawned by | discuss-phase workflow (when ADVISOR_MODE = true) |
| Parallelism | Multiple instances (one per gray area) |
| Tools | Read, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__ |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | 5-column comparison table (Option / Pros / Cons / Complexity / Recommendation) with rationale paragraph |
Key behaviors:
- Researches a single assigned gray area using Claude's knowledge, Context7, and web search
- Produces genuinely viable options — no padding with filler alternatives
- Complexity column uses impact surface + risk (never time estimates)
- Recommendations are conditional ("Rec if X", "Rec if Y") — never single-winner ranking
- Output calibrated by tier: full_maturity (3-5 options with maturity signals), standard (2-4), minimal_decisive (2 options, decisive recommendation)
gsd-research-synthesizer
Role: Combines outputs from parallel researchers into a unified summary.
| Property | Value |
|---|---|
| Spawned by | /gsd-new-project (after 4 researchers complete) |
| Parallelism | Single instance (sequential after researchers) |
| Tools | Read, Write, Bash, Skill |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | .planning/research/SUMMARY.md |
gsd-planner
Role: Creates executable phase plans with task breakdown, dependency analysis, and goal-backward verification.
| Property | Value |
|---|---|
| Spawned by | /gsd-plan-phase, /gsd-quick |
| Parallelism | Single instance |
| Tools | Read, Write, Edit, Bash, Glob, Grep, Skill, WebFetch, mcp__context7__, mcp__plugin_context7_context7__ |
| Model (balanced) | Opus |
| Color | Green |
| Produces | {phase}-{N}-PLAN.md files |
Key behaviors:
- Reads PROJECT.md, REQUIREMENTS.md, CONTEXT.md, RESEARCH.md
- Creates 2-3 atomic task plans sized for single context windows
- Uses XML structure with
<task>elements - Emits a
<fails_when>sibling for every runnable<automated>verify command, naming what output constitutes failure (#3172) - Includes
read_firstandacceptance_criteriasections - Groups plans into dependency waves
- Applies an ordered minimum-solution check after preserving locked decisions and requirement coverage, preferring existing project behavior, standard-library or native-platform capability, and already-installed dependencies before new implementation (#4089)
- Performs reachability check to validate plan steps reference accessible files and APIs (v1.32)
- Enforces a comment-text discipline HARD GATE at plan-write time (
verify.plan-structure): a literal that an acceptance criterion negative-greps for (grep -c 'LIT' file == 0) must not appear verbatim in an<action>body; violations fail plan creation. Use<!-- planner-discipline-allow: LIT -->to allowlist a legitimate occurrence. (#429)
gsd-roadmapper
Role: Creates project roadmaps with phase breakdown and requirement mapping.
| Property | Value |
|---|---|
| Spawned by | /gsd-new-project |
| Parallelism | Single instance |
| Tools | Read, Write, Bash, Glob, Grep, Skill |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | ROADMAP.md |
Key behaviors:
- Maps requirements to phases (traceability)
- Derives success criteria from requirements
- Respects granularity setting for phase count
- Validates coverage (every v1 requirement mapped to a phase)
gsd-executor
Role: Executes GSD plans with atomic commits, deviation handling, and checkpoint protocols.
| Property | Value |
|---|---|
| Spawned by | /gsd-execute-phase, /gsd-quick |
| Parallelism | Multiple (parallel within waves, sequential across waves) |
| Tools | Read, Write, Edit, Bash, Grep, Glob, Skill, mcp__context7__, mcp__plugin_context7_context7__ |
| Model (balanced) | Sonnet |
| Color | Yellow |
| Produces | Code changes, git commits, {phase}-{N}-SUMMARY.md |
Key behaviors:
- Fresh 200K context window per plan
- Follows XML task instructions precisely
- Atomic git commit per completed task
- Handles task types: auto, tracer, checkpoint (human-verify, decision, human-action)
- Tracer feedback gate: after a
tracerslice, verifies it end-to-end before expansion tasks — autonomous runs halt on failure; interactive runs honorworkflow.human_verify_mode(under theend-of-phasedefault an automated-only<verify>continues with no checkpoint; otherwise a human-verify checkpoint is emitted, #3299) - Reports deviations from plan in SUMMARY.md
- Invokes node repair on verification failure
gsd-plan-checker
Role: Verifies plans will achieve phase goals before execution.
| Property | Value |
|---|---|
| Spawned by | /gsd-plan-phase (verification loop, max 3 iterations) |
| Parallelism | Single instance (iterative) |
| Tools | Read, Bash, Glob, Grep, Skill |
| Disallowed Tools | Write, Edit, MultiEdit |
| Model (balanced) | Sonnet |
| Color | Green |
| Produces | PASS/FAIL verdict with specific feedback |
Verification Dimensions — labels match the agent's own ## Dimension <N> headings:
| # | Dimension |
|---|---|
| 1 | Requirement coverage |
| 2 | Task completeness |
| 3 | Dependency correctness |
| 3b | Undeclared / temporal coupling — advisory; flags same-wave plan pairs coupled through shared mutable state or execution order with no depends_on between them |
| 4 | Key links planned |
| 5 | Scope sanity |
| 6 | Verification derivation |
| 7 | Context compliance (when CONTEXT.md exists) |
| 7b | Scope reduction detection |
| 7c | Architectural tier compliance (when RESEARCH.md defines a responsibility map) |
| 8 | Nyquist compliance (when enabled) — checks 8a-8e cover automated-verify presence, feedback latency, sampling continuity and Wave 0 completeness; check 8f blocks a runnable <automated> command with no stated <fails_when> failing direction (#3172) |
| 9 | Cross-plan data contracts |
| 10 | CLAUDE.md compliance |
| 11 | Research resolution |
| 12 | Pattern compliance |
Three further dimensions carry no number: Verify Command Format Sanity, Verify Command Path Resolvability, and Numeric/Factual Claim Authority.
gsd-integration-checker
Role: Verifies cross-phase integration and end-to-end flows.
| Property | Value |
|---|---|
| Spawned by | /gsd-audit-milestone |
| Parallelism | Single instance |
| Tools | Read, Bash, Grep, Glob, Skill |
| Disallowed Tools | Write, Edit, MultiEdit |
| Model (balanced) | Sonnet |
| Color | Blue |
| Produces | Integration verification report |
gsd-ui-checker
Role: Validates UI-SPEC.md design contracts against quality dimensions.
| Property | Value |
|---|---|
| Spawned by | /gsd-ui-phase (validation loop, max 2 iterations) |
| Parallelism | Single instance |
| Tools | Read, Bash, Glob, Grep, Skill |
| Disallowed Tools | Write, Edit, MultiEdit |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | BLOCK/FLAG/PASS verdict |
Verification Dimensions — labels match the agent's own ## Dimension <N> headings:
| # | Dimension |
|---|---|
| 1 | Copywriting |
| 2 | Visuals |
| 3 | Color |
| 4 | Typography |
| 5 | Spacing |
| 6 | Registry Safety |
| 7 | Inventory Provenance |
Key behaviors:
- Inventory provenance (#2845): a UI-SPEC whose component inventory carries no provenance line is reported as a defect, and the inventory is downgraded from a closed allowlist to a non-exhaustive list of known-good components — so an executor is never blocked from something the spec merely failed to mention. A spec with no inventory at all PASSes, which is what keeps every UI-SPEC written before the dimension existed validating unchanged. The checker never executes the recorded command; it reads the spec as a document. Limits, because the dimension is narrower than it reads: the line makes an inventory's origin falsifiable, not verified — nothing re-runs the command or compares the count, so a fabricated line passes; the rule is agent-applied like the other six, not a schema check; and "never executes the recorded command" is an instruction rather than a capability boundary, since the checker holds a
Bashgrant it needs for the agent-skills bootstrap. See Security model → Trade-offs and limits and How to design a UI phase. - Adversarial stance / "The Auditor" (#1578): applies explicit BLOCK/FLAG/PASS tiers and an anti-capitulation rule that resists author-framing pressure while still allowing self-correction when the prior dimension application was mistaken. Persona effects are strongest on Sonnet-class reasoning and unvalidated on budget/Haiku-class routing; the criteria and evidence remain authoritative.
gsd-verifier
Role: Verifies phase goal achievement through goal-backward analysis.
| Property | Value |
|---|---|
| Spawned by | /gsd-execute-phase (after all executors complete) |
| Parallelism | Single instance |
| Tools | Read, Write, Bash, Grep, Glob, Skill |
| Disallowed Tools | Edit, MultiEdit |
| Model (balanced) | Sonnet |
| Color | Green |
| Produces | {phase}-VERIFICATION.md |
Key behaviors:
-
Checks codebase against phase goals, not just task completion
-
PASS/FAIL with specific evidence
-
Logs issues for
/gsd-verify-workto address -
Milestone scope filtering: gaps addressed in later phases are marked as "deferred", not reported as failures (v1.32)
-
Test quality audit (v1.32): verifies that tests prove what they claim by checking for disabled/skipped tests on requirements, circular test patterns (system generating its own expected values), assertion strength (existence vs. value vs. behavioral), and expected value provenance. Blockers from test quality audit override an otherwise passing verification
-
Runs the full workspace test suite at most once per verification — proves a test exists by enumeration and that it passes via a single named test, never re-running the whole suite per must-have.
-
Behavior-dependent calibration (#966): a must-have that asserts a state transition or a cancellation/cleanup/ordering invariant is marked
⚠️ PRESENT_BEHAVIOR_UNVERIFIED(notVERIFIED) when no test exercises it — excluded from theverified_truthsscore, counted in thebehavior_unverifiedfrontmatter field, and routed to human verification, so a cleanN/Ncertifies behavioral evidence rather than mere symbol presence. -
Coincidental-reliance advisory (#1955): a truth that reaches
✓ VERIFIEDis additionally asked why it holds. When the recorded evidence shows the truth holding for an incidental reason —undeclared-precondition,incidental-ordering, orfixture-only— the verdict is qualified as✓ VERIFIED (coincidental-reliance)and the truth is listed in thecoincidental_reliance_itemsfrontmatter field with what to harden. This is advisory: the base✓ VERIFIEDtoken is unchanged, the truth still counts towardverified_truths, the overallstatusis unaffected, and no human-verification item is emitted — a passing phase still passes. It classifies evidence the verifier already gathered rather than asking it to rate its own confidence — but it is honestly an endogenous check, andgsd-core/references/honest-verifier.mdrecords that endogenous gates are measurably weaker than the exogenousbackstoptag it routes on. Advisory status is the consequence, not a coincidence: a miss costs exactly today's behaviour (a plain✓ VERIFIED) and a false positive costs one line of prose, never a failed phase, so a weaker mechanism is affordable here in a way it would not be on a pass/fail axis. Its precision is unmeasured. It complements the two existing axes:PRESENT_BEHAVIOR_UNVERIFIEDis no behavioral evidence,insufficient_specis an under-specified truth, and this is evidence that exists and passes for the wrong reason.The advisory is carried by two surfaces.
agents/gsd-verifier.md(Step 3, sub-step 5c) holds the detection rule, and the verifier's eagerly-importedgsd-core/references/verifier-phase-gates.mdpoints at the canonical report template@~/.claude/gsd-core/templates/verification-report.md, whose## Guidelinescarry the same instruction. (The former third surface, the retiredverify-phaseworkflow, was deleted as an orphan in #1892 — every verification path is subagent-shaped today.) -
Convergence evidence gate (#3304): during re-verification (after a
/gsd-plan-phase --gapscycle), Step 7's anti-pattern scan re-runs at full, unbounded scope — by design — but a 🛑 Blocker it finds no longer auto-reverts a completed gap-closure round on its own judgment alone. A blocker other than the self-evidencing debt-marker check (TBD/FIXME/XXX) blocks unconditionally only if it is a carried-forward gap from the priorVERIFICATION.mdor its file was git-modified since the prior pass (a regression, fail-closed toward blocking when history is unresolvable); otherwise it predates the gap-closure round unflagged and needs deterministic evidence — a named test run red, or another concrete reproducible artifact — to stay blocking. Unevidenced, it downgrades to theadvisory:frontmatter list and the report's "Advisory (New Scope, Unevidenced)" section instead of settingstatus: gaps_found. Full algorithm ingsd-core/references/verifier-evidence-gate.md.
gsd-nyquist-auditor
Role: Fills Nyquist validation gaps by generating tests.
| Property | Value |
|---|---|
| Spawned by | /gsd-validate-phase |
| Parallelism | Single instance |
| Tools | Read, Write, Edit, Bash, Glob, Grep, Skill |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | Test files, updated VALIDATION.md |
Key behaviors:
- Never modifies implementation code — only test files
- Max 3 attempts per gap
- Flags implementation bugs as escalations for user
gsd-ui-auditor
Role: Retroactive 6-pillar visual audit of implemented frontend code.
| Property | Value |
|---|---|
| Spawned by | /gsd-ui-review |
| Parallelism | Single instance |
| Tools | Read, Write, Bash, Grep, Glob, Skill |
| Disallowed Tools | Edit, MultiEdit |
| Model (balanced) | Sonnet |
| Color | Pink |
| Produces | {phase}-UI-REVIEW.md with scores |
| Interaction capture | workflow.ui_interaction_capture (default false) |
6 Audit Pillars (scored 1-4):
- Copywriting
- Visuals
- Color
- Typography
- Spacing
- Experience Design
Interaction capture (default-off). The static screenshots are three viewport captures of
the first paint, taken with npx playwright screenshot, which has no interaction verb — so a
hover state, an open menu, a focus ring or a form's validation state never appears in them.
With workflow.ui_interaction_capture on, /gsd-ui-review passes interaction_capture: true
in the auditor's <config> block and the auditor adds post-interaction captures through the
chrome-devtools CLI (the second binary in the chrome-devtools-mcp package), driven from
Bash against a throwaway --isolated profile: no MCP server, no tools: change. It needs an
installed Chrome (CHROME_BIN overrides discovery). With the key off, or no Chrome resolved,
the section prints one line and the Playwright-only path runs exactly as before. The report's
**Interaction captures:** field carries the outcome — off, skipped with its reason, or the
number of states captured — and an interaction state that was not captured is never reported as
observed. See Enable UI interaction capture.
gsd-dom-verifier
Role: Observes a live DOM and reports which of a wave's stated UI acceptance criteria hold. Additive — never blocks.
| Property | Value |
|---|---|
| Spawned by | live-dom-uat capability step at execute:wave:post |
| Parallelism | One per wave |
| Tools | Read, Write, Glob, Grep, mcp__chrome-devtools__, mcp__claude-in-chrome__ |
| Disallowed Tools | Edit, Bash, the Playwright MCP family |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | {phase}-DOM-VERIFY.md |
| Gated by | workflow.live_dom_uat (default false) |
This is the only GSD agent carrying browser MCP tools. gsd-executor is deliberately not widened — for a first-party agent the static tools: list is the only control that exists (ADR-1244 D2, ADR-857 D4). It carries no Bash: it does not start dev servers or shell out.
Outcome codes (nothing_to_report and could_not_look are never conflated):
outcome |
reason |
Meaning |
|---|---|---|
verified |
ok |
Criteria existed and were observed |
nothing_to_report |
no_criteria |
The wave stated no UI acceptance criteria |
could_not_look |
no_browser_mcp |
No browser MCP answered |
could_not_look |
profile_locked |
Another instance holds the browser profile |
could_not_look |
target_unreachable |
Nothing serving the target |
Reference: Enable live-DOM verification · Explanation
gsd-codebase-mapper
Role: Explores codebase and writes structured analysis documents.
| Property | Value |
|---|---|
| Spawned by | /gsd-map-codebase, post-execute drift gate in /gsd-execute-phase |
| Parallelism | 4 instances (tech, architecture, quality, concerns) |
| Tools | Read, Bash, Grep, Glob, Write, Skill |
| Model (balanced) | Haiku |
| Color | Cyan |
| Produces | .planning/codebase/*.md (7 documents, with last_mapped_commit frontmatter) |
Key behaviors:
- Read-only exploration + structured output
- Writes documents directly to disk
- No reasoning required — pattern extraction from file contents
--paths <p1,p2,...> scope hint (#2003):
Accepts an optional --paths directive in its prompt. When present, the
mapper restricts Glob/Grep/Bash exploration to the listed repo-relative path
prefixes — this is the incremental-remap path used by the post-execute
codebase-drift gate. Path values that contain .., start with /, or
include shell metacharacters are rejected. Without the hint, the mapper
runs its default whole-repo scan.
gsd-debugger
Role: Investigates bugs using scientific method with persistent state.
| Property | Value |
|---|---|
| Spawned by | /gsd-debug, /gsd-verify-work (for failures) |
| Parallelism | Single instance (interactive) |
| Tools | Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch |
| Model (balanced) | Sonnet |
| Color | Orange |
| Produces | .planning/debug/*.md, knowledge-base updates |
Debug Session Lifecycle:
gathering → investigating → fixing → verifying → awaiting_human_verify → resolved
Key behaviors:
- Tracks hypotheses, evidence, and eliminated theories
- State persists across context resets
- Requires human verification before marking resolved
- Runs a multi-signal fix-acceptance guardrail (mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm) before accepting a fix; degrades gracefully when Stryker or a test suite is absent
- Ranks suspect code by Ochiai suspiciousness from test pass/fail coverage (spectrum-based fault localization) before forming hypotheses; skips cleanly when no coverage exists
- Branches root-cause analysis across ≥2 Ishikawa categories and applies an AND-gate check before committing root_cause (guards against 5-Whys single-cause bias); root_cause may hold a set when the AND-gate fires
- Classifies each failure as Bohrbug / Heisenbug-Mandelbug / Concurrency at Phase 1.75 and routes the investigation technique accordingly (routes Bohrbugs to SBFL+bisect, Heisenbugs to record-replay/stability with SBFL skipped, Concurrency to the atomicity/order/deadlock checklist)
- Hardens regression tests via PBT shrinking (minimized counterexample as the seed), explicit oracle classification (specified/derived/metamorphic/implicit), and boundary neighbors around the fixed equivalence class
- Emits a blameless-postmortem Prevention block at resolution (branching 5-Whys, why-wasn't-this-caught, a concrete recurrence guard) and records
why_not_caught+recurrence_guardin the knowledge base so the same bug class is prevented, not just fixed - Recalls prior resolved sessions semantically via MemPalace at Phase 0 (top-k meaning-similar), catching same-root-cause/different-wording cases keyword overlap misses; falls back to keyword matching when MemPalace is absent
- Appends to persistent knowledge base on resolution
- Consults knowledge base on new sessions
gsd-user-profiler
Role: Analyzes session messages across 8 behavioral dimensions to produce a scored developer profile.
| Property | Value |
|---|---|
| Spawned by | /gsd-profile-user |
| Parallelism | Single instance |
| Tools | Read |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | USER-PROFILE.md, CLAUDE.md profile section |
Behavioral Dimensions: Communication style, decision patterns, debugging approach, UX preferences, vendor choices, frustration triggers, learning style, explanation depth.
Key behaviors:
- Read-only agent — analyzes extracted session data, does not modify files
- Produces scored dimensions with confidence levels and evidence citations
- Questionnaire fallback when session history is unavailable
gsd-doc-writer
Role: Writes and updates project documentation. Spawned with a doc_assignment block specifying doc type, mode, and project context.
| Property | Value |
|---|---|
| Spawned by | /gsd-docs-update |
| Parallelism | Multiple instances (one per doc type) |
| Tools | Read, Bash, Grep, Glob, Write, Edit, Skill |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | Project documentation files (README, architecture, API docs, etc.) |
Key behaviors:
- Supports modes: create, update, supplement, fix
- Handles doc types: readme, architecture, getting_started, development, testing, api, configuration, deployment, contributing, custom
- Monorepo-aware: can generate per-package READMEs
- Fix mode accepts failure objects from gsd-doc-verifier for targeted corrections
- Writes directly to disk — does not return content to orchestrator
gsd-doc-verifier
Role: Verifies factual claims in generated documentation against the live codebase.
| Property | Value |
|---|---|
| Spawned by | /gsd-docs-update (after doc-writer completes) |
| Parallelism | Multiple instances (one per doc file) |
| Tools | Read, Write, Bash, Grep, Glob |
| Disallowed Tools | Edit, MultiEdit |
| Model (balanced) | Sonnet |
| Color | Orange |
| Produces | Structured JSON verification results per doc |
Key behaviors:
- Extracts checkable claims (file paths, function names, CLI commands, config keys)
- Verifies each claim against filesystem using tools only — no assumptions
- Writes structured JSON result file for orchestrator to process
- Failed claims feed back to doc-writer in fix mode
gsd-security-auditor
Role: Verifies threat mitigations from PLAN.md threat model exist in implemented code.
| Property | Value |
|---|---|
| Spawned by | /gsd-secure-phase |
| Parallelism | Single instance |
| Tools | Read, Bash, Glob, Grep, Skill |
| Model (balanced) | Sonnet |
| Color | Red |
| Produces | Structured verdict (SECURED / OPEN_THREATS / ESCALATE) — orchestrator writes {phase}-SECURITY.md (#2119) |
Key behaviors:
- Verifies each threat by its declared disposition (mitigate / accept / transfer)
- Does NOT scan blindly for new vulnerabilities — verifies declared mitigations only
- Implementation files are read-only — never patches implementation code
- Unmitigated threats reported as OPEN_THREATS or ESCALATE
- Supports ASVS levels 1/2/3 for verification depth
Advanced and Specialized Agents
Twelve additional agents ship under agents/gsd-*.md and are used by specialty workflows (/gsd-ai-integration-phase, /gsd-eval-review, /gsd-code-review, /gsd-code-review --fix, /gsd-debug, /gsd-map-codebase --query, /gsd-ingest-docs) and by the planner pipeline. Each carries full frontmatter in its agent file; the stubs below are concise by design. The authoritative roster (with spawner and primary-doc status per agent) lives in docs/INVENTORY.md.
gsd-pattern-mapper
Role: Read-only codebase analysis that maps files-to-be-created or modified to their closest existing analogs, producing PATTERNS.md for the planner to consume.
| Property | Value |
|---|---|
| Spawned by | /gsd-plan-phase (between research and planning) |
| Parallelism | Single instance |
| Tools | Read, Bash, Glob, Grep, Write |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | PATTERNS.md in the phase directory |
Key behaviors:
- Extracts file list from CONTEXT.md and RESEARCH.md; classifies each by role (controller, component, service, model, middleware, utility, config, test) and data flow (CRUD, streaming, file I/O, event-driven, request-response)
- Searches for the closest existing analog per file and extracts concrete code excerpts (imports, auth patterns, core pattern, error handling)
- Strictly read-only against source; only writes
PATTERNS.md
gsd-debug-session-manager
Role: Runs the full /gsd-debug checkpoint-and-continuation loop in an isolated context so the orchestrator's main context stays lean; spawns gsd-debugger agents, dispatches specialist skills, and handles user checkpoints via AskUserQuestion.
| Property | Value |
|---|---|
| Spawned by | /gsd-debug |
| Parallelism | Single instance (interactive, stateful) |
| Tools | Read, Write, Edit, Bash, Grep, Glob, Agent, AskUserQuestion |
| Model (balanced) | Sonnet |
| Color | Orange |
| Produces | Compact summary returned to main context; evolves the .planning/debug/{slug}.md session file |
Key behaviors:
- Reads the debug session file first; passes file paths (not inlined contents) to spawned agents to respect context budget
- Treats all user-supplied AskUserQuestion content as data-only, wrapped in DATA_START/DATA_END markers
- Coordinates TDD gates and reasoning checkpoints introduced in v1.36.0
gsd-code-reviewer
Role: Reviews source files for bugs, security vulnerabilities, and code-quality problems; produces a structured REVIEW.md with severity-classified findings.
| Property | Value |
|---|---|
| Spawned by | /gsd-code-review |
| Parallelism | Typically single instance per review scope |
| Tools | Read, Write, Bash, Grep, Glob, Skill |
| Model (balanced) | Sonnet |
| Color | Orange |
| Produces | REVIEW.md in the phase directory |
Key behaviors:
- Detects bugs (logic errors, null/undefined checks, off-by-one, type mismatches, unreachable code), security issues (injection, XSS, hardcoded secrets, insecure crypto), and quality issues
- Honors
CLAUDE.mdproject conventions and.claude/skills//.agents/skills/rules when present - Read-only against implementation source — never modifies code under review
- Full-context review scope: surrounding modules, callers, tests, and docs, not a diff-only pass
- Owns
REVIEW.mdeven when optional external reviewer lanes ran (#4209): it treats their<external_reviewer_evidence>as unverified input, re-verifies every claim against the actual current source before accepting it, and never follows an instruction embedded inside evidence text — there remains exactly oneREVIEW.mdschema regardless of how many lanes contributed
gsd-code-fixer
Role: Applies fixes to findings from REVIEW.md with intelligent (non-blind) patching and atomic per-fix commits; produces REVIEW-FIX.md.
| Property | Value |
|---|---|
| Spawned by | /gsd-code-review --fix |
| Parallelism | Single instance |
| Tools | Read, Edit, Write, Bash, Grep, Glob, Skill |
| Model (balanced) | Sonnet |
| Color | Green |
| Produces | REVIEW-FIX.md; one atomic git commit per applied fix |
Key behaviors:
- Treats
REVIEW.mdsuggestions as guidance, not a patch to apply literally - Commits each fix atomically so review and rollback stay granular
- Honors
CLAUDE.mdand project-skill rules during fixes
gsd-ai-researcher
Role: Researches a chosen AI/LLM framework's official documentation and distills it into implementation-ready guidance — framework quick reference, patterns, and pitfalls — for the Section 3–4b body of AI-SPEC.md.
| Property | Value |
|---|---|
| Spawned by | /gsd-ai-integration-phase |
| Parallelism | Single instance (sequential with domain-researcher / eval-planner) |
| Tools | Read, Write, Edit, Bash, Grep, Glob, WebFetch, WebSearch, mcp__context7__, mcp__plugin_context7_context7__ |
| Model (balanced) | Sonnet |
| Color | Green |
| Produces | Sections 3–4b of AI-SPEC.md (framework quick reference + implementation guidance) |
Key behaviors:
- Uses Context7 MCP when available; falls back to the
ctx7CLI via Bash when MCP tools are stripped from the agent - Anchors guidance to the specific use case, not generic framework overviews
gsd-domain-researcher
Role: Surfaces the business-domain and real-world evaluation context for an AI system — expert rubric ingredients, failure modes, regulatory context — before the eval-planner turns it into measurable rubrics. Writes Section 1b of AI-SPEC.md.
| Property | Value |
|---|---|
| Spawned by | /gsd-ai-integration-phase |
| Parallelism | Single instance |
| Tools | Read, Write, Edit, Bash, Grep, Glob, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__ |
| Model (balanced) | Sonnet |
| Color | Purple |
| Produces | Section 1b of AI-SPEC.md |
Key behaviors:
- Researches the domain, not the technical framework — its output feeds the eval-planner downstream
- Produces rubric ingredients that downstream evaluators can turn into measurable criteria
gsd-eval-planner
Role: Designs the structured evaluation strategy for an AI phase — failure modes, eval dimensions with rubrics, tooling, reference dataset, guardrails, production monitoring. Writes Sections 5–7 of AI-SPEC.md.
| Property | Value |
|---|---|
| Spawned by | /gsd-ai-integration-phase |
| Parallelism | Single instance (sequential after domain-researcher) |
| Tools | Read, Write, Edit, Bash, Grep, Glob, AskUserQuestion |
| Model (balanced) | Sonnet |
| Color | Orange |
| Produces | Sections 5–7 of AI-SPEC.md (Evaluation Strategy, Guardrails, Production Monitoring) |
Required reading: gsd-core/references/ai-evals.md (evaluation framework).
Key behaviors:
- Turns domain-researcher rubric ingredients into measurable, tooled evaluation criteria
- Does not re-derive domain context — reads Section 1 and 1b of
AI-SPEC.mdas established input
gsd-eval-auditor
Role: Retroactive audit of an implemented AI phase's evaluation coverage against its planned AI-SPEC.md eval strategy. Scores each eval dimension COVERED / PARTIAL / MISSING and produces EVAL-REVIEW.md.
| Property | Value |
|---|---|
| Spawned by | /gsd-eval-review |
| Parallelism | Single instance |
| Tools | Read, Write, Bash, Grep, Glob, Skill |
| Disallowed Tools | Edit, MultiEdit |
| Model (balanced) | Sonnet |
| Color | Red |
| Produces | EVAL-REVIEW.md with dimension scores, findings, and remediation guidance |
Required reading: gsd-core/references/ai-evals.md.
Key behaviors:
- Compares the implemented codebase against the planned eval strategy — never re-plans
- Reads implementation files incrementally to respect context budget
gsd-framework-selector
Role: Interactive decision-matrix agent that runs a ≤6-question interview, scores candidate AI/LLM frameworks, and returns a ranked recommendation with rationale.
| Property | Value |
|---|---|
| Spawned by | /gsd-ai-integration-phase |
| Parallelism | Single instance (interactive) |
| Tools | Read, Bash, Grep, Glob, WebSearch, AskUserQuestion |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | Scored ranked recommendation (structured return to orchestrator) |
Required reading: gsd-core/references/ai-frameworks.md (decision matrix).
Key behaviors:
- Scans
package.json,pyproject.toml,requirements*.txtfor existing AI libraries before the interview to avoid recommending a rejected framework - Asks only what the codebase scan and CONTEXT.md have not already answered
gsd-intel-updater
Role: Reads project source and writes structured intel (JSON + Markdown) into .planning/intel/, building a queryable codebase knowledge base that other agents use instead of performing expensive fresh exploration.
| Property | Value |
|---|---|
| Spawned by | /gsd-map-codebase --query (refresh / update flows) |
| Parallelism | Single instance |
| Tools | Read, Write, Bash, Glob, Grep |
| Model (balanced) | Sonnet |
| Color | Cyan |
| Produces | .planning/intel/*.json (and companion Markdown) consumed by gsd-tools query intel |
Key behaviors:
- Writes current state only — no temporal language, every claim references an actual file path
- Uses Glob / Read / Grep for cross-platform correctness; Bash is reserved for
gsd-tools query intelCLI calls
gsd-doc-classifier
Role: Classifies a single planning document as ADR, PRD, SPEC, DOC, or UNKNOWN. Extracts title, scope summary, and cross-references. Writes a JSON classification file used by gsd-doc-synthesizer to build a consolidated context.
| Property | Value |
|---|---|
| Spawned by | /gsd-ingest-docs (parallel fan-out over the doc corpus) |
| Parallelism | One instance per input document |
| Tools | Read, Write, Grep, Glob |
| Model (balanced) | Haiku |
| Color | Yellow |
| Produces | One JSON classification file per input doc (type, title, scope, refs) |
Key behaviors:
- Single-doc scope — never synthesizes or resolves conflicts (that is the synthesizer's job)
- Heuristic-first classification; returns UNKNOWN when the doc lacks type signals rather than guessing
- Extraction discipline (#1578): few-shot input→output exemplars plus a terminal schema restatement; marks a field absent rather than fabricating a value when the doc lacks the signal.
gsd-doc-synthesizer
Role: Synthesizes classified planning docs into a single consolidated context. Applies precedence rules, detects cross-reference cycles, enforces LOCKED-vs-LOCKED hard-blocks, and writes INGEST-CONFLICTS.md with three buckets (auto-resolved, competing-variants, unresolved-blockers).
| Property | Value |
|---|---|
| Spawned by | /gsd-ingest-docs (after classifier fan-in) |
| Parallelism | Single instance |
| Tools | Read, Write, Grep, Glob, Bash |
| Model (balanced) | Sonnet |
| Color | Orange |
| Produces | Consolidated context for .planning/ plus INGEST-CONFLICTS.md report |
Key behaviors:
- Hard-blocks on LOCKED-vs-LOCKED ADR contradictions instead of silently picking a winner
- Follows the
references/doc-conflict-engine.mdcontract so/gsd-importand/gsd-ingest-docsproduce consistent conflict reports - Extraction discipline (#1578): few-shot exemplars plus a terminal schema restatement and a mark-absent (no-fabrication) rule for missing fields.
gsd-mempalace-curator
Role: Ship-time memory curation — writes per-agent diary entries, proposes and creates cross-project tunnels, runs wing-scoped sync pruning, and mirrors extract-learnings output into MemPalace's temporal knowledge graph with provenance.
| Property | Value |
|---|---|
| Spawned by | MemPalace capability at ship:post (when mempalace.enabled = true); diary/tunnels/KG-mirror are then refined by their own toggles |
| Parallelism | Single instance |
| Tools | Read, Bash, Grep, Glob |
| Model (balanced) | Sonnet |
| Produces | Diary entry in MemPalace, wing tunnel proposals, KG provenance records |
Key behaviors:
- Best-effort only — every operation is
onError: skip; a MemPalace failure never halts the loop - Wing-scoped sync pruning (
mempalace sync --wing <wing> --apply) — never runs a global prune - Cross-project tunnel proposals when
mempalace.cross_project_tunnels = true - Mirrors
extract-learningsdecisions, lessons, patterns, and surprises into the KG withsource_drawer_idprovenance - Requires MemPalace MCP server or CLI to be reachable; writes a skip-notice stub when unavailable
Agent Tool Permissions Summary
Scope: this table covers the 22 primary agents only. The 13 advanced/specialized agents listed above carry their own tool surfaces in their
agents/gsd-*.mdfrontmatter (summarized in the per-agent stubs above and indocs/INVENTORY.md).
| Agent | Read | Write | Edit | Bash | Grep | Glob | WebSearch | WebFetch | MCP |
|---|---|---|---|---|---|---|---|---|---|
| project-researcher | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| phase-researcher | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| ui-researcher | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| assumptions-analyzer | ✓ | ✓ | ✓ | ✓ | |||||
| advisor-researcher | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ||
| research-synthesizer | ✓ | ✓ | ✓ | ||||||
| planner | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ||
| roadmapper | ✓ | ✓ | ✓ | ✓ | ✓ | ||||
| executor | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |||
| plan-checker | ✓ | ✓ | ✓ | ✓ | |||||
| integration-checker | ✓ | ✓ | ✓ | ✓ | |||||
| ui-checker | ✓ | ✓ | ✓ | ✓ | |||||
| verifier | ✓ | ✓ | ✓ | ✓ | ✓ | ||||
| nyquist-auditor | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |||
| ui-auditor | ✓ | ✓ | ✓ | ✓ | ✓ | ||||
| codebase-mapper | ✓ | ✓ | ✓ | ✓ | ✓ | ||||
| debugger | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ||
| user-profiler | ✓ | ||||||||
| doc-writer | ✓ | ✓ | ✓ | ✓ | ✓ | ||||
| doc-verifier | ✓ | ✓ | ✓ | ✓ | ✓ | ||||
| security-auditor | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
Principle of Least Privilege:
- Checkers are read-only (no Write/Edit) — they evaluate, never modify
- Researchers have web access — they need current ecosystem information
- Executors have Edit — they modify code but not web access
- Mappers have Write — they write analysis documents but not Edit (no code changes)
Completion Contracts (machine-enforced)
Every agent's return contract is declared in gsd-core/references/agent-contracts.md's Agent Registry table — (Agent, Completion Markers, Consumed by, Kind) — and enforced by npm run check:contract-drift (part of lint:ci).
The Kind column records how a caller actually detects the agent's completion:
| Kind | Detection mechanism |
|---|---|
sentinel-match |
Exact-case string match against a declared marker (by a workflow, command, or another agent) |
artifact+query |
The agent writes a file; the caller reads or queries that artifact |
structured-return |
The agent returns parseable sections/JSON inline; the caller reads the return text |
When you add an agent or change what it returns, update its registry row in the same change — a stale row is a build failure, not a documentation cleanup for later. Markers are extracted fence-aware (a heading inside a fenced block is the emitted template; the same words outside a fence are prose documentation), producer scope includes @-included gsd-core/references/** files, and consumers are matched exact-case (a case-insensitive hit is reported as a collision, never accepted). A marker that is deliberately emitted but matched by nothing carries an (unconsumed: <reason>) annotation — an auditable exemption that waives only the consumer requirement.
The same check also enforces the read-tag pairing: whenever a declared consumer emits <required_reading>, the producing agent's instructions must reference the gate (directly or via an @-included reference) — and the retired <files_to_read> vocabulary may not reappear under workflows/, commands/, or agents/.
For acting on a specific finding, see How to resolve a contract-drift finding.