Files
msd-core/docs/AGENTS.md
Behruz Nassre Esfahani 622f43353c fix(#3299): tracer feedback gate honors workflow.human_verify_mode (#3390)
* fix(#3299): tracer feedback gate honors workflow.human_verify_mode

The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.

Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.

The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.

`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.

Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.

Fixes #3299

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): reconcile the canonical schema table and the stale acceptance test

Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.

1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
   schema reference for the tracer task-type contract, and its Task-types row
   still claimed interactive runs unconditionally present a
   checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
   first round; this one was missed, so the authoritative reference was the
   wrong answer. The row now carries the human_verify_mode-conditional
   behavior and points at the canonical precedence chain.

2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
   tracer ROW EXISTS, never its content, which is why CI could not see the
   drift. It now asserts the row's actual claims and rejects the pre-#3299
   wording. Separately, the #1945 acceptance test named 'interactive run emits
   checkpoint:human-verify after the tracer' kept passing only because its
   substrings still occur in the fallback clause, while its name asserted the
   opposite of shipped behavior. Renamed and narrowed to what #1945 still
   guarantees, plus a new interactiveIsConditional pin so the unconditional
   prose cannot be restored under a passing substring check.

3. plan-md.md's <verify> row now documents that the legacy bare-text form
   (valid, and still shown at :179) does not reach the #3299 auto-continue —
   only a <verify> carrying <automated> does — so the benefit is silently
   unreachable for tracers using that format.

Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions

Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.

MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.

MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:

- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
  'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
  same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
  appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
  presence. Deliberately brittle: CONTEXT.md names that table the canonical
  schema reference, so a wording change must be a conscious edit in both
  places.

Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs

Peer review round 4. Blacklisting did not hold, twice over:

- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
  forms in the auto-continue clause. Round 4 defeated that by appending
  'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
  before expansion' — none of the banned tokens, same restored interruption
  after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
  '<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
  legacy bare form operative. 107/107 passed across tracer, planner and the
  three size-cap suites.

Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.

These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.

Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): strip comments, require uniqueness, pin whole regions

Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):

- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
  wrong copy: every extractor selected the commented decoy. Worked against the
  planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
  checkpoint_protocol before expansion, regardless of the mode-specific rules
  below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
  approval' below the canonical table. The pinned text was untouched, so
  equality held while the shipped meaning inverted.

The shape that holds, applied to every operative surface:
  1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
  2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
     cannot hide behind a correct first one;
  3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
     suffix can override what the pin proves.

Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.

Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.

Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): drop the superseded exact-placeholder planner assertion

Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.

Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.

Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner

Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:

- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
  internally uses the repo's interleaved fence/comment scanner. Instrumenting
  candidate lines as throwaway predicate declarations borrows that scanner with
  no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
  edits silently turned the guards into decoy checks:
    * a forgotten '-->' comments the live rule through to EOF, and the
      balanced-only stripper still saw and accepted the commented rule;
    * a normal fenced documentation example of the rule, plus a whitespace-only
      reformat of the live list item, made the selector choose the example.
  Neither needs intent. A dangling comment is a typo; a fenced example is good
  documentation. Together they reproduce exactly the accidental drift #3299 came
  from — with CI green.

The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.

Verified both ordinary-edit scenarios now fail the suite (each was green before).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): close the operative-selection gaps the maintainer blocked on

trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.

1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
   replaced a matched candidate with an UNINDENTED marker regardless of the
   original line's indentation. A 4-space-indented CommonMark code block is not
   skipped by parsePredicates (it accepts indented declarations by design), so
   stripping the indent PROMOTED an indented decoy to operative — the exact
   inversion of the guard's purpose. The marker now preserves the original
   indent, and a candidate that is itself indented 4+ spaces is never injected.

2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
   pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
   pollute the count. Now filters on a Set of the indexes actually injected on
   this call.

3. RAW FENCE SELECTION (planner). The template test matched the first raw
   ```xml fence after the marker with no fence/comment awareness — the one
   selection in the suite that was not operative-aware — so a commented-out
   decoy template between the marker and the real one would be selected while
   the live template regressed. The opener must now be operative AND the first
   non-blank line after the marker.

4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
   a fenced example containing a ### / <type line truncated the pinned region
   early — a false FAILURE on a legitimate doc edit. End anchors now go through
   the same operative filter as start anchors.

Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.

Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): allow-list operative indentation; pin marker provenance

Review round 9.

BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", "  \t" and "   \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.

MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.

MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.

Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.

KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge

The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in 362d0434b (#3370) once #2728's entries were spent.
#3370's own fragment now owns execute-plan.md at the base, so a new
3299-*.json naming that path would collide — mergeAckSources rejects a
duplicate key across fragments rather than silently last-winning.

Re-arms #3370's entry instead, the mechanism the gate is built for (a spent
ack whose reason changes in the diff is live again), carrying #3370's own
reason forward verbatim so the base growth keeps its account.

Verified: emitted-attribution 175/175 against origin/next@be9329b10.

* fix(#3299): honor golden rule 6 in the tracer gate, extract the chain

Addresses the review on #3390 (B1-B3, M1-M4, minors).

B3 — checkpoints.md asserted two incompatible rules about the same gate.
Golden rule 6 says gate="blocking-human" stops for a human in every mode;
the precedence table scoped row 1 to interactive runs, so a first-match
chain let an auto-mode tracer carrying that gate fall to row 2 and
auto-continue. Rule 6 wins: row 1 is now "Any run, any mode", the
justification sentence it falsified is gone, and the STOP is evaluated
before the auto-mode branch at all three dispatch sites — gsd-executor.md,
execute-plan.md and the plan-md.md schema row. Unreachable by our planner
is not unreachable: src/verify.cts parses only `type` and never consults
`gate` on non-checkpoint tasks, so an imported PLAN.md can carry it.

B1 — the LARGE-tier cap. gsd-executor.md is 49150 on next against a 49152
cap, so this PR could not add a byte. Extracted rather than trimmed: the
precedence chain now lives only in checkpoints.md (already @-imported by
<checkpoint_protocol>, so no new load), and the duplicate summary inside
that protocol section is a pointer. The rationale the earlier trim
deleted is restored — "production-quality, never a throwaway" and
"Pouring more layers onto a broken foundation...". Result 49097: 55 bytes
under the cap and a net 53-byte REDUCTION against next, so the PR returns
headroom instead of consuming it.

B2 — merged upstream/next and resolved all three drift-ack conflicts.
2775 changed shape upstream (string -> {reason}); adopted the new form.

M1 — the 2775 ack claimed the Nyquist Rule sat "twelve lines earlier"; it
is ~75 lines. Corrected to "earlier in the file".
M2 — ack arithmetic restated from measurement, not from a stale base. The
2943 #3299 append is DELETED: with gsd-executor.md now shrinking there is
no ripple to acknowledge, and emitted-attribution correctly flagged the
entry as stale.
M3 — changeset rewritten to the documented bold-lead + em-dash one-liner.
M4 — the two self-defeated shapes are gone. The planner-human-verify-mode
presence checks now go through operativeLineIndexes. The config-get check
does NOT: all three reads live inside ```bash fences, which is their
correct executable form, and that selector excludes fenced lines by
design. It instead pins exactly one live, uncommented, fenced read per
file — mutation-tested against both a commented-out read and a duplicate.

Minors — dangling colon lead-in dropped, a "below" pointer that pointed
above corrected, and the `(default)` asymmetry between the two dispatch
copies aligned.

Two defects the merge surfaced, both caught only by the full suite:
the new #3576 gate rejected this PR's own bare `references/checkpoints.md`
cite in planner-human-verify-mode.md (rewritten to the canonical
gsd-core/ form), and the line-keyed PROSE_ALLOWLIST entry for
gsd-executor.md needed 794 -> 795 after this change shifted the line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): correct the size record the 08-22 merge falsified

Review round: one Major, four Minors.

Major — the #3299 arm's arithmetic was measured before the merge and is
now wrong in a document whose whole purpose is to be an accurate size
record. Re-measured at head: execute-plan.md is 39315 B on next and
40111 B here, so the 796-byte delta was right but the endpoints and the
headroom were not (849 bytes against DEFAULT_CAP 40960, not 1003). The
superseded figures are named rather than silently replaced. Confirmed
the workflow cap counts LF BYTES while the agent cap counts CHARACTERS —
two caps in two units, one per file.

Minor 1 — 2943-context7-tool-name.json reverted to next. JSON.parse of
both sides was already identical; the diff was an em-dash/times-sign
re-serialization left over from adding and then removing the #3299 arm.
No business in this PR.

Minor 2 — the duplicated `tracer row Autonomy cell` test is gone. Both
copies were new here and carried the same ~8-line canonical string; the
one removed selected its row with a raw startsWith find, the shape this
suite records at :477 as defeated in round 1. Its rationale — why the
cell is pinned EXACTLY, and the append-a-contradiction attack that
defeated keyword matching — is carried onto the surviving fence-aware
copy rather than deleted with it.

Minor 3 — the executor's condensed interactive clause said only "re-run,
continue", which does not distinguish pass from fail; read in isolation
it invites expansion onto a broken slice, the outcome the gate exists to
prevent. Now "re-run; fails → HALT as above, passes → continue, no
checkpoint". The pinned expected string moved with it. Executor at
48,905 chars, 247 under the cap.

Minor 4 — 2775 asserted two different current sizes for gsd-planner.md.
The stale half is next's own text taken wholesale, so the contradiction
was inherited; it now reads as a before-figure rather than a current one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): cite the plan-md example by section, not by a drifting line

Review round 7, Nit N-1. The 2775 ack fragment justified its one-line
formatting with "matching docs/reference/plan-md.md:207's own example
style". At head, :207 is prose; the one-line <verify><automated>
example it means is at :222. The citation was accurate when written
(77c2fda, f23205c) and drifted with a later merge of next.

Re-pointed by section rather than by line — it has already drifted
once, and the fragment's whole purpose is to be an accurate record —
and the drift itself is recorded inline so the correction does not
quietly overwrite what the earlier number said.

Also narrows the changeset's "any task with gate=blocking-human" to
"any tracer carrying gate=blocking-human" (found by Codex in the
whole-PR pass). Golden rule 6 and the #3299 decision table both scope
that gate to checkpoints and to the tracer feedback gate; the normal
type="auto" branch never inspects `gate`, so the wider claim promised
behavior the implementation does not have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): answer fence-delimiter liveness by insertion, not replacement

Review round 9. The round-8 fence-awareness fix was itself unsound, in the same
class it was added to close.

`operativeLineIndexes` detects operative lines by REPLACING each candidate with
a throwaway predicate declaration and asking `parsePredicates` which survived.
Sound for ordinary content lines. Not sound for a fence DELIMITER, which is
exactly what the tracer-template selection passed it: deleting every ```xml
OPENER leaves each matching closer to become an opener, and since
`computeSkippedLineFlags` is a strict FORWARD state machine, fence parity
inverts for the whole remainder of the document.

Measured against the real file rather than argued:

  agents/gsd-planner.md has 3 live top-level ```xml openers — 0-based 180, 232,
  262. operativeLineIndexes reported 180 and 262. Line 232, the "Task-level TDD"
  example, read NON-OPERATIVE — a wrong answer from a helper whose only job is
  that question.

It passed only by parity coincidence, and one extra live example anywhere
earlier flipped it to a false FAILURE blaming a decoy that does not exist:

  HEAD as-is                  | anchor 260 | openIdx 262 | ASSERTION PASSES
  +1 unrelated ```xml example | anchor 265 | openIdx 267 | ASSERTION *** FAILS ***

Fixed by asking the question a way that perturbs nothing. `isOperativePosition`
INSERTS a marker on its own line immediately before the candidate instead of
replacing it. Insertion preserves every delimiter, and because the skip-state
machine runs strictly forward, a line inserted at `idx` observes exactly the
fence/comment state the candidate observes, with nothing but the marker between
them — so marker-operative IS the candidate's position-liveness.

The review's suggested direction (substitute a same-shaped opener that still
opens a fence) cannot work here: the marker would then be inside the fence and
would never parse as a predicate at all.

Position-liveness is not content-liveness, so the helper also rejects a line
that is entirely comment (`<!-- ```xml -->`), rather than leaving that to each
caller's own shape test to happen to exclude.

`operativeLineIndexes` now THROWS when its candidate regex matches a fence
delimiter, so the unsound route cannot be reached again by a future caller
rather than only being fixed at the one site that got it wrong.

Verified with the same extra-example scenario above: with the fix, all 35 rows
stay green. Teeth: reverting the call site to `operativeLineSet` turns the
tracer-template row red on the new guard. The regression row pins both live
openers (the second is the one the deletion route lost), the block-commented
and same-line-commented openers, a line inside a fence, and re-checks both
openers after unrelated lines shift above them.

Only tests/tracer-bullet.test.cjs changes — no agent file is touched, so the
5-char gsd-planner.md and 19-byte gsd-executor.md headroom are unaffected.

Verified: `npm run lint:ci` exit 0; full `npm test` 31307 tests / 31292 pass /
0 fail / 14 skipped, TMPDIR unset, against a freshly synced origin/next.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): guard the delimiter class, match the scanner, pin the assignment

Codex full-PR review of #3390, run against the round-9 head. Three defects,
two of them in the code that round added.

1. The mode-read pin survived the regression it exists to catch.
   `READ` matched the config-get substring only, so rewriting the shipped line
   as `IGNORED_MODE=$(gsd_run query config-get ...)` kept the row green while
   nothing defined HUMAN_VERIFY_MODE — the gate falls through to STOP and #3299
   is back with the suite passing. The regex now requires the assignment. A
   lookahead after `end-of-phase` closes the other half: the bare prefix also
   accepted `--default end-of-phase-wrong`. Proven by mutation: renaming the
   variable in agents/gsd-executor.md now turns that row red, and did not before.

2. The round-9 fence-delimiter guard was a SAMPLE of the class, not the class.
   It probed a fixed list of five delimiter strings. `~~~xml`, ```json, `~~~~`
   and arbitrary info strings all walk past any list short enough to write down
   — the guard was added precisely because one such regex had already slipped
   through. Now matched against the lines the regex actually selects in the
   document, which cannot go stale and cannot miss a spelling nobody thought of.
   Four such spellings pinned as rows.

3. `isOperativePosition` disagreed with the scanner it delegates to.
   For `<!-- closed --> real content` it stripped the span, found surviving
   content, and answered "live". `computeSkippedLineFlags` skips an ENTIRE line
   whose trimmed text starts with `<!--`, balanced or not, before it considers
   fences at all. Verified directly against parsePredicates. It now applies the
   scanner's own rule instead of out-reasoning it. Latent for the present caller
   (its anchored ```xml shape cannot match a comment-prefixed line), real in
   general.

Disclosed rather than fixed, and raised with the maintainer: the exact executor
region pin ends before the second operative tracer-gate paragraph at
agents/gsd-executor.md:327, which is only heading-checked — so contradictory
later instructions could ship. How much of that file to pin is a call for its
owner.

Independently probed isOperativePosition across 19 edge cases before the review
(line 0, CRLF, tab / 4-space / mixed " \t" indentation, 0-3 space fences, nested
fences, ~~~ fences, info strings, bounds); all correct. That probe is what
surfaced finding 2, which the review then confirmed from the other direction.

Verified: `npm run lint:ci` exit 0; full `npm test` 31296 tests / 31281 pass /
0 fail / 14 skipped, TMPDIR unset. One caveat stated rather than smoothed over:
in that run tests/planning-snapshot.test.cjs was truncated by concurrency after
row A5 — 11 tests did not execute, which a 0-fail aggregate cannot show. Re-run
in isolation it is 87 tests / 87 pass / 0 fail, and it is untouched by this
change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-23 18:43:53 -04:00

44 KiB
Raw Blame History

GSD Agent Reference

Full role cards for 22 primary agents plus concise stubs for 12 advanced/specialized agents (34 shipped agents total). The agents/ directory and docs/INVENTORY.md are the authoritative roster; see Architecture for context.


Overview

GSD uses a multi-agent architecture where thin orchestrators (workflow files) spawn specialized agents with fresh context windows. Each agent has a focused role, limited tool access, and produces specific artifacts.

Required reading (#3423): the canonical spawn-block tag is <required_reading> on BOTH sides — orchestrators emit it, and gating agents enforce it ("you MUST use the Read tool to load every file listed there before performing any other actions"). The legacy <files_to_read> emit-tag is retired and banned repo-wide by tests/agent-required-reading-consistency.test.cjs, because a mismatched pair silently disarms the enforcement clause.

Agent Categories

The table below covers the 22 primary agents detailed in this section. Thirteen additional shipped agents (pattern-mapper, debug-session-manager, code-reviewer, code-fixer, ai-researcher, domain-researcher, eval-planner, eval-auditor, framework-selector, intel-updater, doc-classifier, doc-synthesizer, mempalace-curator) have concise stubs in the Advanced and Specialized Agents section below. For the authoritative 35-agent roster, see docs/INVENTORY.md and the agents/ directory.

Category Count Agents
Researchers 3 project-researcher, phase-researcher, ui-researcher
Analyzers 2 assumptions-analyzer, advisor-researcher
Synthesizers 1 research-synthesizer
Planners 1 planner
Roadmappers 1 roadmapper
Executors 1 executor
Checkers 3 plan-checker, integration-checker, ui-checker
Verifiers 2 verifier, dom-verifier
Auditors 3 nyquist-auditor, ui-auditor, security-auditor
Mappers 1 codebase-mapper
Debuggers 1 debugger
Doc Writers 2 doc-writer, doc-verifier
Profilers 1 user-profiler

Agent Details

gsd-project-researcher

Role: Researches domain ecosystem before roadmap creation.

Property Value
Spawned by /gsd-new-project, /gsd-new-milestone
Parallelism 4 instances (stack, features, architecture, pitfalls)
Tools Read, Write, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__, mcp__firecrawl__, mcp__exa__, mcp__tavily__, mcp__ref__, mcp__jina__, mcp__perplexity__
Model (balanced) Sonnet
Color Cyan
Produces .planning/research/STACK.md, FEATURES.md, ARCHITECTURE.md, PITFALLS.md

Capabilities:

  • Web search for current ecosystem information
  • Context7 MCP integration for library documentation
  • Writes research documents directly to disk (reduces orchestrator context load)

gsd-phase-researcher

Role: Researches how to implement a specific phase before planning.

Property Value
Spawned by /gsd-plan-phase
Parallelism 4 instances (same focus areas as project researcher)
Tools Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__, mcp__firecrawl__, mcp__exa__, mcp__tavily__, mcp__ref__, mcp__jina__, mcp__perplexity__
Model (balanced) Sonnet
Color Cyan
Produces {phase}-RESEARCH.md

Capabilities:

  • Reads CONTEXT.md to focus research on user's decisions
  • Investigates implementation patterns for the specific phase domain
  • Detects test infrastructure for Nyquist validation mapping
  • Tags in-repo discrete values (enums, schema unions, error codes, status constants, paths) [VERIFIED] only after reading the source-of-truth file that run, citing path and line range, and quoting the values verbatim
  • Refuses [VERIFIED] for a compatibility claim resting on missing metadata (no python_requires, no engines field, no per-version classifier, no changelog entry, no matching support-matrix row) — an absence constrains no version, and an enumerated allow-list that stops short of the target is still an absence, so only a positive falsification attempt with its failing output pasted earns the tag; anything less stays [ASSUMED]

gsd-ui-researcher

Role: Produces UI design contracts for frontend phases.

Property Value
Spawned by /gsd-ui-phase
Parallelism Single instance
Tools Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__, mcp__firecrawl__, mcp__exa__, mcp__tavily__, mcp__ref__, mcp__jina__*
Model (balanced) Sonnet
Color Purple
Produces {phase}-UI-SPEC.md

Capabilities:

  • Detects design system state (shadcn components.json, Tailwind config, existing tokens)
  • Offers shadcn initialization for React/Next.js/Vite projects
  • Asks only unanswered design contract questions
  • Enforces registry safety gate for third-party components
  • Enumerates the component inventory rather than recalling it (#2845): the UI-SPEC's ## Component Inventory carries a provenance line — the command that enumerated it, the count it returned, the resolved <package>@<version>, and the date — or, when nothing can enumerate it, a Could not enumerate: <reason> record in the same slot

gsd-assumptions-analyzer

Role: Deeply analyzes codebase for a phase and returns structured assumptions with evidence, confidence levels, and consequences if wrong.

Property Value
Spawned by discuss-phase-assumptions workflow (when workflow.discuss_mode = 'assumptions')
Parallelism Single instance
Tools Read, Bash, Grep, Glob, Skill
Model (balanced) Sonnet
Color Cyan
Produces Structured assumptions with decision statements, evidence file paths, confidence levels

Key behaviors:

  • Reads ROADMAP.md phase description and prior CONTEXT.md files
  • Searches codebase for files related to the phase (components, patterns, similar features)
  • Reads 5-15 most relevant source files to form evidence-based assumptions
  • Classifies confidence: Confident (clear from code), Likely (reasonable inference), Unclear (could go multiple ways)
  • Flags topics that need external research (library compatibility, ecosystem best practices)
  • Output calibrated by tier: full_maturity (3-5 areas), standard (3-4), minimal_decisive (2-3)

gsd-advisor-researcher

Role: Researches a single gray area decision during discuss-phase advisor mode and returns a structured comparison table.

Property Value
Spawned by discuss-phase workflow (when ADVISOR_MODE = true)
Parallelism Multiple instances (one per gray area)
Tools Read, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__
Model (balanced) Sonnet
Color Cyan
Produces 5-column comparison table (Option / Pros / Cons / Complexity / Recommendation) with rationale paragraph

Key behaviors:

  • Researches a single assigned gray area using Claude's knowledge, Context7, and web search
  • Produces genuinely viable options — no padding with filler alternatives
  • Complexity column uses impact surface + risk (never time estimates)
  • Recommendations are conditional ("Rec if X", "Rec if Y") — never single-winner ranking
  • Output calibrated by tier: full_maturity (3-5 options with maturity signals), standard (2-4), minimal_decisive (2 options, decisive recommendation)

gsd-research-synthesizer

Role: Combines outputs from parallel researchers into a unified summary.

Property Value
Spawned by /gsd-new-project (after 4 researchers complete)
Parallelism Single instance (sequential after researchers)
Tools Read, Write, Bash, Skill
Model (balanced) Sonnet
Color Purple
Produces .planning/research/SUMMARY.md

gsd-planner

Role: Creates executable phase plans with task breakdown, dependency analysis, and goal-backward verification.

Property Value
Spawned by /gsd-plan-phase, /gsd-quick
Parallelism Single instance
Tools Read, Write, Edit, Bash, Glob, Grep, Skill, WebFetch, mcp__context7__, mcp__plugin_context7_context7__
Model (balanced) Opus
Color Green
Produces {phase}-{N}-PLAN.md files

Key behaviors:

  • Reads PROJECT.md, REQUIREMENTS.md, CONTEXT.md, RESEARCH.md
  • Creates 2-3 atomic task plans sized for single context windows
  • Uses XML structure with <task> elements
  • Includes read_first and acceptance_criteria sections
  • Groups plans into dependency waves
  • Performs reachability check to validate plan steps reference accessible files and APIs (v1.32)
  • Enforces a comment-text discipline HARD GATE at plan-write time (verify.plan-structure): a literal that an acceptance criterion negative-greps for (grep -c 'LIT' file == 0) must not appear verbatim in an <action> body; violations fail plan creation. Use <!-- planner-discipline-allow: LIT --> to allowlist a legitimate occurrence. (#429)

gsd-roadmapper

Role: Creates project roadmaps with phase breakdown and requirement mapping.

Property Value
Spawned by /gsd-new-project
Parallelism Single instance
Tools Read, Write, Bash, Glob, Grep, Skill
Model (balanced) Sonnet
Color Purple
Produces ROADMAP.md

Key behaviors:

  • Maps requirements to phases (traceability)
  • Derives success criteria from requirements
  • Respects granularity setting for phase count
  • Validates coverage (every v1 requirement mapped to a phase)

gsd-executor

Role: Executes GSD plans with atomic commits, deviation handling, and checkpoint protocols.

Property Value
Spawned by /gsd-execute-phase, /gsd-quick
Parallelism Multiple (parallel within waves, sequential across waves)
Tools Read, Write, Edit, Bash, Grep, Glob, Skill, mcp__context7__, mcp__plugin_context7_context7__
Model (balanced) Sonnet
Color Yellow
Produces Code changes, git commits, {phase}-{N}-SUMMARY.md

Key behaviors:

  • Fresh 200K context window per plan
  • Follows XML task instructions precisely
  • Atomic git commit per completed task
  • Handles task types: auto, tracer, checkpoint (human-verify, decision, human-action)
  • Tracer feedback gate: after a tracer slice, verifies it end-to-end before expansion tasks — autonomous runs halt on failure; interactive runs honor workflow.human_verify_mode (under the end-of-phase default an automated-only <verify> continues with no checkpoint; otherwise a human-verify checkpoint is emitted, #3299)
  • Reports deviations from plan in SUMMARY.md
  • Invokes node repair on verification failure

gsd-plan-checker

Role: Verifies plans will achieve phase goals before execution.

Property Value
Spawned by /gsd-plan-phase (verification loop, max 3 iterations)
Parallelism Single instance (iterative)
Tools Read, Bash, Glob, Grep, Skill
Disallowed Tools Write, Edit, MultiEdit
Model (balanced) Sonnet
Color Green
Produces PASS/FAIL verdict with specific feedback

Verification Dimensions — labels match the agent's own ## Dimension <N> headings:

# Dimension
1 Requirement coverage
2 Task completeness
3 Dependency correctness
3b Undeclared / temporal coupling — advisory; flags same-wave plan pairs coupled through shared mutable state or execution order with no depends_on between them
4 Key links planned
5 Scope sanity
6 Verification derivation
7 Context compliance (when CONTEXT.md exists)
7b Scope reduction detection
7c Architectural tier compliance (when RESEARCH.md defines a responsibility map)
8 Nyquist compliance (when enabled)
9 Cross-plan data contracts
10 CLAUDE.md compliance
11 Research resolution
12 Pattern compliance

Two further dimensions carry no number: Verify Command Format Sanity and Numeric/Factual Claim Authority.


gsd-integration-checker

Role: Verifies cross-phase integration and end-to-end flows.

Property Value
Spawned by /gsd-audit-milestone
Parallelism Single instance
Tools Read, Bash, Grep, Glob, Skill
Disallowed Tools Write, Edit, MultiEdit
Model (balanced) Sonnet
Color Blue
Produces Integration verification report

gsd-ui-checker

Role: Validates UI-SPEC.md design contracts against quality dimensions.

Property Value
Spawned by /gsd-ui-phase (validation loop, max 2 iterations)
Parallelism Single instance
Tools Read, Bash, Glob, Grep, Skill
Disallowed Tools Write, Edit, MultiEdit
Model (balanced) Sonnet
Color Cyan
Produces BLOCK/FLAG/PASS verdict

Verification Dimensions — labels match the agent's own ## Dimension <N> headings:

# Dimension
1 Copywriting
2 Visuals
3 Color
4 Typography
5 Spacing
6 Registry Safety
7 Inventory Provenance

Key behaviors:

  • Inventory provenance (#2845): a UI-SPEC whose component inventory carries no provenance line is reported as a defect, and the inventory is downgraded from a closed allowlist to a non-exhaustive list of known-good components — so an executor is never blocked from something the spec merely failed to mention. A spec with no inventory at all PASSes, which is what keeps every UI-SPEC written before the dimension existed validating unchanged. The checker never executes the recorded command; it reads the spec as a document. Limits, because the dimension is narrower than it reads: the line makes an inventory's origin falsifiable, not verified — nothing re-runs the command or compares the count, so a fabricated line passes; the rule is agent-applied like the other six, not a schema check; and "never executes the recorded command" is an instruction rather than a capability boundary, since the checker holds a Bash grant it needs for the agent-skills bootstrap. See Security model → Trade-offs and limits and How to design a UI phase.
  • Adversarial stance / "The Auditor" (#1578): applies explicit BLOCK/FLAG/PASS tiers and an anti-capitulation rule that resists author-framing pressure while still allowing self-correction when the prior dimension application was mistaken. Persona effects are strongest on Sonnet-class reasoning and unvalidated on budget/Haiku-class routing; the criteria and evidence remain authoritative.

gsd-verifier

Role: Verifies phase goal achievement through goal-backward analysis.

Property Value
Spawned by /gsd-execute-phase (after all executors complete)
Parallelism Single instance
Tools Read, Write, Bash, Grep, Glob, Skill
Disallowed Tools Edit, MultiEdit
Model (balanced) Sonnet
Color Green
Produces {phase}-VERIFICATION.md

Key behaviors:

  • Checks codebase against phase goals, not just task completion

  • PASS/FAIL with specific evidence

  • Logs issues for /gsd-verify-work to address

  • Milestone scope filtering: gaps addressed in later phases are marked as "deferred", not reported as failures (v1.32)

  • Test quality audit (v1.32): verifies that tests prove what they claim by checking for disabled/skipped tests on requirements, circular test patterns (system generating its own expected values), assertion strength (existence vs. value vs. behavioral), and expected value provenance. Blockers from test quality audit override an otherwise passing verification

  • Runs the full workspace test suite at most once per verification — proves a test exists by enumeration and that it passes via a single named test, never re-running the whole suite per must-have.

  • Behavior-dependent calibration (#966): a must-have that asserts a state transition or a cancellation/cleanup/ordering invariant is marked ⚠️ PRESENT_BEHAVIOR_UNVERIFIED (not VERIFIED) when no test exercises it — excluded from the verified_truths score, counted in the behavior_unverified frontmatter field, and routed to human verification, so a clean N/N certifies behavioral evidence rather than mere symbol presence.

  • Coincidental-reliance advisory (#1955): a truth that reaches ✓ VERIFIED is additionally asked why it holds. When the recorded evidence shows the truth holding for an incidental reason — undeclared-precondition, incidental-ordering, or fixture-only — the verdict is qualified as ✓ VERIFIED (coincidental-reliance) and the truth is listed in the coincidental_reliance_items frontmatter field with what to harden. This is advisory: the base ✓ VERIFIED token is unchanged, the truth still counts toward verified_truths, the overall status is unaffected, and no human-verification item is emitted — a passing phase still passes. It classifies evidence the verifier already gathered rather than asking it to rate its own confidence — but it is honestly an endogenous check, and gsd-core/references/honest-verifier.md records that endogenous gates are measurably weaker than the exogenous backstop tag it routes on. Advisory status is the consequence, not a coincidence: a miss costs exactly today's behaviour (a plain ✓ VERIFIED) and a false positive costs one line of prose, never a failed phase, so a weaker mechanism is affordable here in a way it would not be on a pass/fail axis. Its precision is unmeasured. It complements the two existing axes: PRESENT_BEHAVIOR_UNVERIFIED is no behavioral evidence, insufficient_spec is an under-specified truth, and this is evidence that exists and passes for the wrong reason.

    The advisory is carried by two surfaces. agents/gsd-verifier.md (Step 3, sub-step 5c) holds the detection rule, and the verifier's eagerly-imported gsd-core/references/verifier-phase-gates.md points at the canonical report template @~/.claude/gsd-core/templates/verification-report.md, whose ## Guidelines carry the same instruction. (The former third surface, the retired verify-phase workflow, was deleted as an orphan in #1892 — every verification path is subagent-shaped today.)


gsd-nyquist-auditor

Role: Fills Nyquist validation gaps by generating tests.

Property Value
Spawned by /gsd-validate-phase
Parallelism Single instance
Tools Read, Write, Edit, Bash, Glob, Grep, Skill
Model (balanced) Sonnet
Color Purple
Produces Test files, updated VALIDATION.md

Key behaviors:

  • Never modifies implementation code — only test files
  • Max 3 attempts per gap
  • Flags implementation bugs as escalations for user

gsd-ui-auditor

Role: Retroactive 6-pillar visual audit of implemented frontend code.

Property Value
Spawned by /gsd-ui-review
Parallelism Single instance
Tools Read, Write, Bash, Grep, Glob, Skill
Disallowed Tools Edit, MultiEdit
Model (balanced) Sonnet
Color Pink
Produces {phase}-UI-REVIEW.md with scores

6 Audit Pillars (scored 1-4):

  1. Copywriting
  2. Visuals
  3. Color
  4. Typography
  5. Spacing
  6. Experience Design

gsd-dom-verifier

Role: Observes a live DOM and reports which of a wave's stated UI acceptance criteria hold. Additive — never blocks.

Property Value
Spawned by live-dom-uat capability step at execute:wave:post
Parallelism One per wave
Tools Read, Write, Glob, Grep, mcp__chrome-devtools__, mcp__claude-in-chrome__
Disallowed Tools Edit, Bash, the Playwright MCP family
Model (balanced) Sonnet
Color Cyan
Produces {phase}-DOM-VERIFY.md
Gated by workflow.live_dom_uat (default false)

This is the only GSD agent carrying browser MCP tools. gsd-executor is deliberately not widened — for a first-party agent the static tools: list is the only control that exists (ADR-1244 D2, ADR-857 D4). It carries no Bash: it does not start dev servers or shell out.

Outcome codes (nothing_to_report and could_not_look are never conflated):

outcome reason Meaning
verified ok Criteria existed and were observed
nothing_to_report no_criteria The wave stated no UI acceptance criteria
could_not_look no_browser_mcp No browser MCP answered
could_not_look profile_locked Another instance holds the browser profile
could_not_look target_unreachable Nothing serving the target

Reference: Enable live-DOM verification · Explanation


gsd-codebase-mapper

Role: Explores codebase and writes structured analysis documents.

Property Value
Spawned by /gsd-map-codebase, post-execute drift gate in /gsd-execute-phase
Parallelism 4 instances (tech, architecture, quality, concerns)
Tools Read, Bash, Grep, Glob, Write, Skill
Model (balanced) Haiku
Color Cyan
Produces .planning/codebase/*.md (7 documents, with last_mapped_commit frontmatter)

Key behaviors:

  • Read-only exploration + structured output
  • Writes documents directly to disk
  • No reasoning required — pattern extraction from file contents

--paths <p1,p2,...> scope hint (#2003): Accepts an optional --paths directive in its prompt. When present, the mapper restricts Glob/Grep/Bash exploration to the listed repo-relative path prefixes — this is the incremental-remap path used by the post-execute codebase-drift gate. Path values that contain .., start with /, or include shell metacharacters are rejected. Without the hint, the mapper runs its default whole-repo scan.


gsd-debugger

Role: Investigates bugs using scientific method with persistent state.

Property Value
Spawned by /gsd-debug, /gsd-verify-work (for failures)
Parallelism Single instance (interactive)
Tools Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch
Model (balanced) Sonnet
Color Orange
Produces .planning/debug/*.md, knowledge-base updates

Debug Session Lifecycle: gathering → investigating → fixing → verifying → awaiting_human_verify → resolved

Key behaviors:

  • Tracks hypotheses, evidence, and eliminated theories
  • State persists across context resets
  • Requires human verification before marking resolved
  • Runs a multi-signal fix-acceptance guardrail (mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm) before accepting a fix; degrades gracefully when Stryker or a test suite is absent
  • Ranks suspect code by Ochiai suspiciousness from test pass/fail coverage (spectrum-based fault localization) before forming hypotheses; skips cleanly when no coverage exists
  • Branches root-cause analysis across ≥2 Ishikawa categories and applies an AND-gate check before committing root_cause (guards against 5-Whys single-cause bias); root_cause may hold a set when the AND-gate fires
  • Classifies each failure as Bohrbug / Heisenbug-Mandelbug / Concurrency at Phase 1.75 and routes the investigation technique accordingly (routes Bohrbugs to SBFL+bisect, Heisenbugs to record-replay/stability with SBFL skipped, Concurrency to the atomicity/order/deadlock checklist)
  • Hardens regression tests via PBT shrinking (minimized counterexample as the seed), explicit oracle classification (specified/derived/metamorphic/implicit), and boundary neighbors around the fixed equivalence class
  • Emits a blameless-postmortem Prevention block at resolution (branching 5-Whys, why-wasn't-this-caught, a concrete recurrence guard) and records why_not_caught + recurrence_guard in the knowledge base so the same bug class is prevented, not just fixed
  • Recalls prior resolved sessions semantically via MemPalace at Phase 0 (top-k meaning-similar), catching same-root-cause/different-wording cases keyword overlap misses; falls back to keyword matching when MemPalace is absent
  • Appends to persistent knowledge base on resolution
  • Consults knowledge base on new sessions

gsd-user-profiler

Role: Analyzes session messages across 8 behavioral dimensions to produce a scored developer profile.

Property Value
Spawned by /gsd-profile-user
Parallelism Single instance
Tools Read
Model (balanced) Sonnet
Color Purple
Produces USER-PROFILE.md, CLAUDE.md profile section

Behavioral Dimensions: Communication style, decision patterns, debugging approach, UX preferences, vendor choices, frustration triggers, learning style, explanation depth.

Key behaviors:

  • Read-only agent — analyzes extracted session data, does not modify files
  • Produces scored dimensions with confidence levels and evidence citations
  • Questionnaire fallback when session history is unavailable

gsd-doc-writer

Role: Writes and updates project documentation. Spawned with a doc_assignment block specifying doc type, mode, and project context.

Property Value
Spawned by /gsd-docs-update
Parallelism Multiple instances (one per doc type)
Tools Read, Bash, Grep, Glob, Write, Edit, Skill
Model (balanced) Sonnet
Color Purple
Produces Project documentation files (README, architecture, API docs, etc.)

Key behaviors:

  • Supports modes: create, update, supplement, fix
  • Handles doc types: readme, architecture, getting_started, development, testing, api, configuration, deployment, contributing, custom
  • Monorepo-aware: can generate per-package READMEs
  • Fix mode accepts failure objects from gsd-doc-verifier for targeted corrections
  • Writes directly to disk — does not return content to orchestrator

gsd-doc-verifier

Role: Verifies factual claims in generated documentation against the live codebase.

Property Value
Spawned by /gsd-docs-update (after doc-writer completes)
Parallelism Multiple instances (one per doc file)
Tools Read, Write, Bash, Grep, Glob
Disallowed Tools Edit, MultiEdit
Model (balanced) Sonnet
Color Orange
Produces Structured JSON verification results per doc

Key behaviors:

  • Extracts checkable claims (file paths, function names, CLI commands, config keys)
  • Verifies each claim against filesystem using tools only — no assumptions
  • Writes structured JSON result file for orchestrator to process
  • Failed claims feed back to doc-writer in fix mode

gsd-security-auditor

Role: Verifies threat mitigations from PLAN.md threat model exist in implemented code.

Property Value
Spawned by /gsd-secure-phase
Parallelism Single instance
Tools Read, Bash, Glob, Grep, Skill
Model (balanced) Sonnet
Color Red
Produces Structured verdict (SECURED / OPEN_THREATS / ESCALATE) — orchestrator writes {phase}-SECURITY.md (#2119)

Key behaviors:

  • Verifies each threat by its declared disposition (mitigate / accept / transfer)
  • Does NOT scan blindly for new vulnerabilities — verifies declared mitigations only
  • Implementation files are read-only — never patches implementation code
  • Unmitigated threats reported as OPEN_THREATS or ESCALATE
  • Supports ASVS levels 1/2/3 for verification depth

Advanced and Specialized Agents

Twelve additional agents ship under agents/gsd-*.md and are used by specialty workflows (/gsd-ai-integration-phase, /gsd-eval-review, /gsd-code-review, /gsd-code-review --fix, /gsd-debug, /gsd-map-codebase --query, /gsd-ingest-docs) and by the planner pipeline. Each carries full frontmatter in its agent file; the stubs below are concise by design. The authoritative roster (with spawner and primary-doc status per agent) lives in docs/INVENTORY.md.

gsd-pattern-mapper

Role: Read-only codebase analysis that maps files-to-be-created or modified to their closest existing analogs, producing PATTERNS.md for the planner to consume.

Property Value
Spawned by /gsd-plan-phase (between research and planning)
Parallelism Single instance
Tools Read, Bash, Glob, Grep, Write
Model (balanced) Sonnet
Color Purple
Produces PATTERNS.md in the phase directory

Key behaviors:

  • Extracts file list from CONTEXT.md and RESEARCH.md; classifies each by role (controller, component, service, model, middleware, utility, config, test) and data flow (CRUD, streaming, file I/O, event-driven, request-response)
  • Searches for the closest existing analog per file and extracts concrete code excerpts (imports, auth patterns, core pattern, error handling)
  • Strictly read-only against source; only writes PATTERNS.md

gsd-debug-session-manager

Role: Runs the full /gsd-debug checkpoint-and-continuation loop in an isolated context so the orchestrator's main context stays lean; spawns gsd-debugger agents, dispatches specialist skills, and handles user checkpoints via AskUserQuestion.

Property Value
Spawned by /gsd-debug
Parallelism Single instance (interactive, stateful)
Tools Read, Write, Edit, Bash, Grep, Glob, Agent, AskUserQuestion
Model (balanced) Sonnet
Color Orange
Produces Compact summary returned to main context; evolves the .planning/debug/{slug}.md session file

Key behaviors:

  • Reads the debug session file first; passes file paths (not inlined contents) to spawned agents to respect context budget
  • Treats all user-supplied AskUserQuestion content as data-only, wrapped in DATA_START/DATA_END markers
  • Coordinates TDD gates and reasoning checkpoints introduced in v1.36.0

gsd-code-reviewer

Role: Reviews source files for bugs, security vulnerabilities, and code-quality problems; produces a structured REVIEW.md with severity-classified findings.

Property Value
Spawned by /gsd-code-review
Parallelism Typically single instance per review scope
Tools Read, Write, Bash, Grep, Glob, Skill
Model (balanced) Sonnet
Color Orange
Produces REVIEW.md in the phase directory

Key behaviors:

  • Detects bugs (logic errors, null/undefined checks, off-by-one, type mismatches, unreachable code), security issues (injection, XSS, hardcoded secrets, insecure crypto), and quality issues
  • Honors CLAUDE.md project conventions and .claude/skills/ / .agents/skills/ rules when present
  • Read-only against implementation source — never modifies code under review

gsd-code-fixer

Role: Applies fixes to findings from REVIEW.md with intelligent (non-blind) patching and atomic per-fix commits; produces REVIEW-FIX.md.

Property Value
Spawned by /gsd-code-review --fix
Parallelism Single instance
Tools Read, Edit, Write, Bash, Grep, Glob, Skill
Model (balanced) Sonnet
Color Green
Produces REVIEW-FIX.md; one atomic git commit per applied fix

Key behaviors:

  • Treats REVIEW.md suggestions as guidance, not a patch to apply literally
  • Commits each fix atomically so review and rollback stay granular
  • Honors CLAUDE.md and project-skill rules during fixes

gsd-ai-researcher

Role: Researches a chosen AI/LLM framework's official documentation and distills it into implementation-ready guidance — framework quick reference, patterns, and pitfalls — for the Section 3–4b body of AI-SPEC.md.

Property Value
Spawned by /gsd-ai-integration-phase
Parallelism Single instance (sequential with domain-researcher / eval-planner)
Tools Read, Write, Edit, Bash, Grep, Glob, WebFetch, WebSearch, mcp__context7__, mcp__plugin_context7_context7__
Model (balanced) Sonnet
Color Green
Produces Sections 3–4b of AI-SPEC.md (framework quick reference + implementation guidance)

Key behaviors:

  • Uses Context7 MCP when available; falls back to the ctx7 CLI via Bash when MCP tools are stripped from the agent
  • Anchors guidance to the specific use case, not generic framework overviews

gsd-domain-researcher

Role: Surfaces the business-domain and real-world evaluation context for an AI system — expert rubric ingredients, failure modes, regulatory context — before the eval-planner turns it into measurable rubrics. Writes Section 1b of AI-SPEC.md.

Property Value
Spawned by /gsd-ai-integration-phase
Parallelism Single instance
Tools Read, Write, Edit, Bash, Grep, Glob, WebSearch, WebFetch, mcp__context7__, mcp__plugin_context7_context7__
Model (balanced) Sonnet
Color Purple
Produces Section 1b of AI-SPEC.md

Key behaviors:

  • Researches the domain, not the technical framework — its output feeds the eval-planner downstream
  • Produces rubric ingredients that downstream evaluators can turn into measurable criteria

gsd-eval-planner

Role: Designs the structured evaluation strategy for an AI phase — failure modes, eval dimensions with rubrics, tooling, reference dataset, guardrails, production monitoring. Writes Sections 5–7 of AI-SPEC.md.

Property Value
Spawned by /gsd-ai-integration-phase
Parallelism Single instance (sequential after domain-researcher)
Tools Read, Write, Edit, Bash, Grep, Glob, AskUserQuestion
Model (balanced) Sonnet
Color Orange
Produces Sections 5–7 of AI-SPEC.md (Evaluation Strategy, Guardrails, Production Monitoring)

Required reading: gsd-core/references/ai-evals.md (evaluation framework).

Key behaviors:

  • Turns domain-researcher rubric ingredients into measurable, tooled evaluation criteria
  • Does not re-derive domain context — reads Section 1 and 1b of AI-SPEC.md as established input

gsd-eval-auditor

Role: Retroactive audit of an implemented AI phase's evaluation coverage against its planned AI-SPEC.md eval strategy. Scores each eval dimension COVERED / PARTIAL / MISSING and produces EVAL-REVIEW.md.

Property Value
Spawned by /gsd-eval-review
Parallelism Single instance
Tools Read, Write, Bash, Grep, Glob, Skill
Disallowed Tools Edit, MultiEdit
Model (balanced) Sonnet
Color Red
Produces EVAL-REVIEW.md with dimension scores, findings, and remediation guidance

Required reading: gsd-core/references/ai-evals.md.

Key behaviors:

  • Compares the implemented codebase against the planned eval strategy — never re-plans
  • Reads implementation files incrementally to respect context budget

gsd-framework-selector

Role: Interactive decision-matrix agent that runs a ≤6-question interview, scores candidate AI/LLM frameworks, and returns a ranked recommendation with rationale.

Property Value
Spawned by /gsd-ai-integration-phase
Parallelism Single instance (interactive)
Tools Read, Bash, Grep, Glob, WebSearch, AskUserQuestion
Model (balanced) Sonnet
Color Cyan
Produces Scored ranked recommendation (structured return to orchestrator)

Required reading: gsd-core/references/ai-frameworks.md (decision matrix).

Key behaviors:

  • Scans package.json, pyproject.toml, requirements*.txt for existing AI libraries before the interview to avoid recommending a rejected framework
  • Asks only what the codebase scan and CONTEXT.md have not already answered

gsd-intel-updater

Role: Reads project source and writes structured intel (JSON + Markdown) into .planning/intel/, building a queryable codebase knowledge base that other agents use instead of performing expensive fresh exploration.

Property Value
Spawned by /gsd-map-codebase --query (refresh / update flows)
Parallelism Single instance
Tools Read, Write, Bash, Glob, Grep
Model (balanced) Sonnet
Color Cyan
Produces .planning/intel/*.json (and companion Markdown) consumed by gsd-tools query intel

Key behaviors:

  • Writes current state only — no temporal language, every claim references an actual file path
  • Uses Glob / Read / Grep for cross-platform correctness; Bash is reserved for gsd-tools query intel CLI calls

gsd-doc-classifier

Role: Classifies a single planning document as ADR, PRD, SPEC, DOC, or UNKNOWN. Extracts title, scope summary, and cross-references. Writes a JSON classification file used by gsd-doc-synthesizer to build a consolidated context.

Property Value
Spawned by /gsd-ingest-docs (parallel fan-out over the doc corpus)
Parallelism One instance per input document
Tools Read, Write, Grep, Glob
Model (balanced) Haiku
Color Yellow
Produces One JSON classification file per input doc (type, title, scope, refs)

Key behaviors:

  • Single-doc scope — never synthesizes or resolves conflicts (that is the synthesizer's job)
  • Heuristic-first classification; returns UNKNOWN when the doc lacks type signals rather than guessing
  • Extraction discipline (#1578): few-shot input→output exemplars plus a terminal schema restatement; marks a field absent rather than fabricating a value when the doc lacks the signal.

gsd-doc-synthesizer

Role: Synthesizes classified planning docs into a single consolidated context. Applies precedence rules, detects cross-reference cycles, enforces LOCKED-vs-LOCKED hard-blocks, and writes INGEST-CONFLICTS.md with three buckets (auto-resolved, competing-variants, unresolved-blockers).

Property Value
Spawned by /gsd-ingest-docs (after classifier fan-in)
Parallelism Single instance
Tools Read, Write, Grep, Glob, Bash
Model (balanced) Sonnet
Color Orange
Produces Consolidated context for .planning/ plus INGEST-CONFLICTS.md report

Key behaviors:

  • Hard-blocks on LOCKED-vs-LOCKED ADR contradictions instead of silently picking a winner
  • Follows the references/doc-conflict-engine.md contract so /gsd-import and /gsd-ingest-docs produce consistent conflict reports
  • Extraction discipline (#1578): few-shot exemplars plus a terminal schema restatement and a mark-absent (no-fabrication) rule for missing fields.

gsd-mempalace-curator

Role: Ship-time memory curation — writes per-agent diary entries, proposes and creates cross-project tunnels, runs wing-scoped sync pruning, and mirrors extract-learnings output into MemPalace's temporal knowledge graph with provenance.

Property Value
Spawned by MemPalace capability at ship:post (when mempalace.enabled = true); diary/tunnels/KG-mirror are then refined by their own toggles
Parallelism Single instance
Tools Read, Bash, Grep, Glob
Model (balanced) Sonnet
Produces Diary entry in MemPalace, wing tunnel proposals, KG provenance records

Key behaviors:

  • Best-effort only — every operation is onError: skip; a MemPalace failure never halts the loop
  • Wing-scoped sync pruning (mempalace sync --wing <wing> --apply) — never runs a global prune
  • Cross-project tunnel proposals when mempalace.cross_project_tunnels = true
  • Mirrors extract-learnings decisions, lessons, patterns, and surprises into the KG with source_drawer_id provenance
  • Requires MemPalace MCP server or CLI to be reachable; writes a skip-notice stub when unavailable

Agent Tool Permissions Summary

Scope: this table covers the 22 primary agents only. The 13 advanced/specialized agents listed above carry their own tool surfaces in their agents/gsd-*.md frontmatter (summarized in the per-agent stubs above and in docs/INVENTORY.md).

Agent Read Write Edit Bash Grep Glob WebSearch WebFetch MCP
project-researcher ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
phase-researcher ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
ui-researcher ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
assumptions-analyzer ✓ ✓ ✓ ✓
advisor-researcher ✓ ✓ ✓ ✓ ✓ ✓ ✓
research-synthesizer ✓ ✓ ✓
planner ✓ ✓ ✓ ✓ ✓ ✓ ✓
roadmapper ✓ ✓ ✓ ✓ ✓
executor ✓ ✓ ✓ ✓ ✓ ✓
plan-checker ✓ ✓ ✓ ✓
integration-checker ✓ ✓ ✓ ✓
ui-checker ✓ ✓ ✓ ✓
verifier ✓ ✓ ✓ ✓ ✓
nyquist-auditor ✓ ✓ ✓ ✓ ✓ ✓
ui-auditor ✓ ✓ ✓ ✓ ✓
codebase-mapper ✓ ✓ ✓ ✓ ✓
debugger ✓ ✓ ✓ ✓ ✓ ✓ ✓
user-profiler ✓
doc-writer ✓ ✓ ✓ ✓ ✓
doc-verifier ✓ ✓ ✓ ✓ ✓
security-auditor ✓ ✓ ✓ ✓ ✓ ✓

Principle of Least Privilege:

  • Checkers are read-only (no Write/Edit) — they evaluate, never modify
  • Researchers have web access — they need current ecosystem information
  • Executors have Edit — they modify code but not web access
  • Mappers have Write — they write analysis documents but not Edit (no code changes)

Completion Contracts (machine-enforced)

Every agent's return contract is declared in gsd-core/references/agent-contracts.md's Agent Registry table — (Agent, Completion Markers, Consumed by, Kind) — and enforced by npm run check:contract-drift (part of lint:ci).

The Kind column records how a caller actually detects the agent's completion:

Kind Detection mechanism
sentinel-match Exact-case string match against a declared marker (by a workflow, command, or another agent)
artifact+query The agent writes a file; the caller reads or queries that artifact
structured-return The agent returns parseable sections/JSON inline; the caller reads the return text

When you add an agent or change what it returns, update its registry row in the same change — a stale row is a build failure, not a documentation cleanup for later. Markers are extracted fence-aware (a heading inside a fenced block is the emitted template; the same words outside a fence are prose documentation), producer scope includes @-included gsd-core/references/** files, and consumers are matched exact-case (a case-insensitive hit is reported as a collision, never accepted). A marker that is deliberately emitted but matched by nothing carries an (unconsumed: <reason>) annotation — an auditable exemption that waives only the consumer requirement.

The same check also enforces the read-tag pairing: whenever a declared consumer emits <required_reading>, the producing agent's instructions must reference the gate (directly or via an @-included reference) — and the retired <files_to_read> vocabulary may not reappear under workflows/, commands/, or agents/.

For acting on a specific finding, see How to resolve a contract-drift finding.