Files
msd-core/docs/ARCHITECTURE.md
Dennis Alexis Valin Dittrich 18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00

79 KiB

GSD Core Architecture

System architecture for contributors and advanced users. For user-facing documentation, see Feature Reference or User Guide.


Table of Contents


System Overview

GSD Core is a meta-prompting framework that sits between the user and AI coding agents (Claude Code, Kimi CLI, OpenCode, Kilo, Codex, Copilot, Antigravity, Trae, Cline, Augment Code). It provides:

  1. Context engineering — Structured artifacts that give the AI everything it needs per task (see Context engineering)
  2. Multi-agent orchestration — Thin orchestrators that spawn specialized agents with fresh context windows (see Multi-agent orchestration)
  3. Spec-driven development — Requirements → research → plans → execution → verification pipeline
  4. State management — Persistent project memory across sessions and context resets
┌──────────────────────────────────────────────────────┐
│                      USER                            │
│            /gsd-command [args]                        │
└─────────────────────┬────────────────────────────────┘
                      │
┌─────────────────────▼────────────────────────────────┐
│              COMMAND LAYER                            │
│   commands/gsd/*.md — Prompt-based command files      │
│   (Claude Code custom commands / Codex skills)        │
└─────────────────────┬────────────────────────────────┘
                      │
┌─────────────────────▼────────────────────────────────┐
│              WORKFLOW LAYER                           │
│   gsd-core/workflows/*.md — Orchestration logic  │
│   (Reads references, spawns agents, manages state)    │
└──────┬──────────────┬─────────────────┬──────────────┘
       │              │                 │
┌──────▼──────┐ ┌─────▼─────┐ ┌────────▼───────┐
│  AGENT      │ │  AGENT    │ │  AGENT         │
│  (fresh     │ │  (fresh   │ │  (fresh        │
│   context)  │ │   context)│ │   context)     │
└──────┬──────┘ └─────┬─────┘ └────────┬───────┘
       │              │                 │
┌──────▼──────────────▼─────────────────▼──────────────┐
│              CLI TOOLS LAYER                          │
│   gsd-tools.cjs command families + domain modules      │
│   command-routing-hub + observability seams            │
└──────────────────────┬───────────────────────────────┘
                       │
┌──────────────────────▼───────────────────────────────┐
│              FILE SYSTEM (.planning/)                 │
│   PROJECT.md | REQUIREMENTS.md | ROADMAP.md          │
│   STATE.md | config.json | phases/ | research/       │
└──────────────────────────────────────────────────────┘

Design Principles

1. Fresh Context Per Agent

Every agent spawned by an orchestrator gets a clean context window (up to 200K tokens). This eliminates context rot — the quality degradation that happens as an AI fills its context window with accumulated conversation.

2. Thin Orchestrators

Workflow files (gsd-core/workflows/*.md) never do heavy lifting. They:

  • Load context via gsd-tools.cjs init <workflow>
  • Spawn specialized agents with focused prompts
  • Collect results and route to the next step
  • Update state between steps

3. File-Based State

All state lives in .planning/ as human-readable Markdown and JSON. No database, no server, no external dependencies. This means:

  • State survives context resets (/clear)
  • State is inspectable by both humans and agents
  • State can be committed to git for team visibility

4. Absent = Enabled

Workflow feature flags follow the absent = enabled pattern. If a key is missing from config.json, it defaults to true. Users explicitly disable features; they don't need to enable defaults.

5. Defense in Depth

Multiple layers prevent common failure modes:

  • Plans are verified before execution (plan-checker agent)
  • Execution produces atomic commits per task
  • Post-execution verification checks against phase goals
  • UAT provides human verification as final gate

Component Architecture

Commands (commands/gsd/*.md)

User-facing entry points. Each file contains YAML frontmatter (name, description, allowed-tools) and a prompt body that bootstraps the workflow. Commands are installed as:

  • Claude Code: Custom slash commands (hyphen form, /gsd-command-name)
  • OpenCode / Kilo: Slash commands (hyphen form, /gsd-command-name)
  • Codex: Skills ($gsd-command-name)
  • Copilot: Slash commands (hyphen form, /gsd-command-name)
  • Kimi CLI: Agent Skills (/skill:gsd-command-name) plus an explicit custom agent launch with kimi --agent-file
  • Antigravity: Skills

Total commands: see docs/INVENTORY.md for the authoritative count and full roster.

Two-stage hierarchical routing (v1.40, #2792)

To keep the eager skill-listing token cost low, v1.40 introduces six namespace meta-skills (gsd-workflow, gsd-project, gsd-quality, gsd-context, gsd-manage, gsd-ideate — sourced from commands/gsd/ns-*.md, but the invocable name: is the bare form shown here) layered above the concrete sub-skills. On runtimes with non-recursive skill loaders (cline, qwen, hermes, augment, trae) the installer now realizes this fully: it emits only the 6 namespace router bundles as top-level skills and nests the ~61 concrete skills under <router>/skills/<name>/SKILL.md, so the eager listing is ≈6 entries instead of ≈67. The model selects a namespace router, which instructs it to read the nested concrete skill file via a routing table embedded in the router body. On these runtimes concrete skills are not directly invocable by bare name via the Skill tool; they are reachable through the router. Slash commands (/gsd-*, via the separate commands surface) are unaffected where the runtime has one. On runtimes with recursive or unconfirmed skill loaders (claude global, cursor, codex, copilot, windsurf, codebuddy, opencode, kilo, antigravity) the layout remains flat — all skills emitted at the top level as before. Antigravity moved from nested to flat in #1614: agy scans only skills/<name>/SKILL.md, so nested sub-skills were unreachable. Claude was reverted to flat in #924: the Skill tool hard-errors on unknown names rather than re-routing via the router, so nested concrete skills were uninvokable.

The router descriptions use pipe-separated keyword tags (≤ 60 chars) per the Tool Attention research showing keyword-dense tags outperform prose for routing at ~40 % the token cost.

MCP token-budget interaction

The eager skill listing is one of two recurring per-turn token costs. The other is the MCP tool schema injected by every enabled MCP server in .claude/settings.json. Heavyweight MCP servers (browser/playwright, Mac-tools, Windows-tools) can each cost 20 k+ tokens per turn — often dwarfing what model_profile tuning saves. The toggle lives in the Claude Code harness (enabledMcpjsonServers / disabledMcpjsonServers in .claude/settings.json) and is not a GSD concern. Together, the two-stage routing layer (#2792) and disciplined MCP enablement are the largest cost levers per turn. See docs/USER-GUIDE.md and references/context-budget.md for the audit checklist.

Workflows (gsd-core/workflows/*.md)

Orchestration logic that commands reference. Contains the step-by-step process including:

  • Context loading via gsd-tools.cjs init handlers
  • Agent spawn instructions with model resolution
  • Gate/checkpoint definitions
  • State update patterns
  • Error handling and recovery

Total workflows: see docs/INVENTORY.md for the authoritative count and full roster.

Progressive disclosure for workflows

Workflow files are loaded verbatim into Claude's context every time the corresponding /gsd-* command is invoked. The workflow size budget enforced by tests/workflow-size-budget.test.cjs keeps each file bounded, mirroring the the agent size-budget convention. The budget is measured in bytes (#717), not lines: line count over-penalizes prose and under-catches token-dense tables and code blocks, whereas bytes are deterministic and match the unit our vendors bound on — Codex truncates instruction docs past 32,768 bytes (project_doc_max_bytes). We adopt that unit, not that exact number: the XL/LARGE ceilings below sit above 32,768 because these are grandfathered top-level orchestrators loaded by Claude, not Codex AGENTS.md docs.

Tier Per-file byte limit
XL 90,000 — top-level orchestrators (execute-phase, plan-phase, new-project)
LARGE 54,000 — multi-step planners and large feature workflows
DEFAULT 38,000 — focused single-purpose workflows (the target tier)

Ceilings are not fixed forever: under the tighten-only ratchet (#597) each one tracks its tier's current high-water mark within a small grace band, so budgets may only decrease over time.

Why the budget exists. With prompt caching the per-invocation cost of a large workflow is modest (cache reads run ~10% of input). The stronger, caching-independent reason is quality: as context grows, recall and reasoning degrade ("context rot" / attention budget), so leaner, higher-signal instructions produce better plans. The ceiling protects the agent's attention, not just the token bill.

Because the budget measures one file, it is a proxy for the real goal — bounded loaded context. Extraction only helps when the extracted content is loaded lazily (Read at the step that needs it). Moving prose into a file that is still eagerly @-imported shrinks the measured file without shrinking loaded context, which games the proxy rather than serving the goal.

workflows/discuss-phase.md is held to a stricter <30,000-byte ceiling per the discuss-phase byte budget (#717; the discuss-phase/modes split keeps it ≈32000 bytes). When a workflow grows beyond its tier, extract per-mode bodies into workflows/<workflow>/modes/<mode>.md, templates into workflows/<workflow>/templates/, and shared knowledge into gsd-core/references/. The parent file becomes a thin dispatcher that Reads only the mode and template files needed for the current invocation.

workflows/discuss-phase/ is the canonical example of this pattern — parent dispatches, modes/ holds per-flag behavior (power.md, all.md, auto.md, chain.md, text.md, batch.md, analyze.md, default.md, advisor.md), and templates/ holds CONTEXT.md, DISCUSSION-LOG.md, and checkpoint.json schemas that are read only when the corresponding output file is being written.

workflows/plan-phase.md, workflows/execute-phase.md, and the gsd-planner / gsd-executor agent definitions apply the same discipline to their MVP-only reference bodies — planner-mvp-mode.md, user-story-template.md, skeleton-template.md, and execute-mvp-tdd.md are referenced for the planner/executor to Read only on MVP, Walking-Skeleton, or MVP+TDD paths, rather than eagerly @-imported, so non-MVP runs do not pay their context cost (guards against the "@-import behind a conditional still loads eagerly" leak; see #720). The dedicated mvp-phase workflow keeps its eager imports, since it is always MVP.

Agents (agents/*.md)

Specialized agent definitions with frontmatter specifying:

  • name — Agent identifier
  • description — Role and purpose
  • tools — Allowed tool access (Read, Write, Edit, Bash, Grep, Glob, WebSearch, etc.)
  • color — Terminal output color for visual distinction

Total agents: 33

References (gsd-core/references/*.md)

Shared knowledge documents that workflows and agents @-reference (see docs/INVENTORY.md for the authoritative full roster):

Core references:

  • checkpoints.md — Checkpoint type definitions and interaction patterns
  • gates.md — 4 canonical gate types (Confirm, Quality, Safety, Transition) wired into plan-checker and verifier
  • model-profiles.md — Per-agent model tier assignments
  • model-profile-resolution.md — Model resolution algorithm documentation
  • verification-patterns.md — How to verify different artifact types
  • verification-overrides.md — Per-artifact verification override rules
  • planning-config.md — Full config schema and behavior
  • git-integration.md — Git commit, branching, and history patterns
  • git-planning-commit.md — Planning directory commit conventions
  • questioning.md — Dream extraction philosophy for project initialization
  • tdd.md — Test-driven development integration patterns
  • ui-brand.md — Visual output formatting patterns
  • common-bug-patterns.md — Common bug patterns for code review and verification

Workflow references:

  • agent-contracts.md — Formal interface between orchestrators and agents
  • context-budget.md — Context window budget allocation rules
  • continuation-format.md — Session continuation/resume format
  • domain-probes.md — Domain-specific probing questions for discuss-phase
  • gate-prompts.md — Gate/checkpoint prompt templates
  • revision-loop.md — Plan revision iteration patterns
  • universal-anti-patterns.md — Common anti-patterns to detect and avoid
  • artifact-types.md — Planning artifact type definitions
  • phase-argument-parsing.md — Phase argument parsing conventions
  • decimal-phase-calculation.md — Decimal sub-phase numbering rules
  • workstream-flag.md — Workstream active pointer conventions
  • user-profiling.md — User behavioral profiling methodology
  • thinking-partner.md — Conditional thinking partner activation at decision points

Thinking model references:

References for integrating thinking-class models (o3, o4-mini, Gemini 2.5 Pro) into GSD workflows:

  • thinking-models-debug.md — Thinking model patterns for debugging workflows
  • thinking-models-execution.md — Thinking model patterns for execution agents
  • thinking-models-planning.md — Thinking model patterns for planning agents
  • thinking-models-research.md — Thinking model patterns for research agents
  • thinking-models-verification.md — Thinking model patterns for verification agents

Modular planner decomposition:

The planner agent (agents/gsd-planner.md) was decomposed from a single monolithic file into a core agent plus reference modules to stay under the 50K character limit imposed by some runtimes:

  • planner-gap-closure.md — Gap closure mode behavior (reads VERIFICATION.md, targeted replanning)
  • planner-reviews.md — Cross-AI review integration (reads REVIEWS.md from /gsd-review)
  • planner-revision.md — Plan revision patterns for iterative refinement

Templates (gsd-core/templates/)

Markdown templates for all planning artifacts. Used by gsd-tools.cjs template fill / phase.scaffold (and top-level scaffold) to create pre-structured files:

  • project.md, requirements.md, roadmap.md, state.md — Core project files
  • phase-prompt.md — Phase execution prompt template
  • summary.md (+ summary-minimal.md, summary-standard.md, summary-complex.md) — Granularity-aware summary templates
  • DEBUG.md — Debug session tracking template
  • UI-SPEC.md, UAT.md, VALIDATION.md — Specialized verification templates
  • discussion-log.md — Discussion audit trail template
  • codebase/ — Brownfield mapping templates (stack, architecture, conventions, concerns, structure, testing, integrations)
  • research-project/ — Research output templates (SUMMARY, STACK, FEATURES, ARCHITECTURE, PITFALLS)

Hooks (hooks/)

Runtime hooks that integrate with the host AI agent:

Hook Event Purpose
gsd-statusline.js statusLine Displays model (long-context suffixes like (1M context) collapse to a compact (1M) badge), task, directory, and context usage bar
gsd-context-monitor.js PostToolUse / AfterTool Injects agent-facing context warnings at 35%/25% remaining
gsd-check-update.js SessionStart Foreground trigger for the background update check
gsd-ensure-canonical-path.js SessionStart For Claude Code plugin installs, symlinks ~/.claude/gsd-core/{bin,contexts,references,templates,workflows} to the plugin's bundled tree so @~/.claude/gsd-core/... includes resolve; runs first in SessionStart, no-op in classic installs, self-heals after claude plugin update (#997)
gsd-check-update-worker.js (helper) Background worker spawned by gsd-check-update.js; no direct event registration
gsd-prompt-guard.js PreToolUse Scans .planning/ writes for prompt injection patterns (advisory)
gsd-read-injection-scanner.js PostToolUse Scans Read tool output for injected instructions in untrusted content
gsd-workflow-guard.js PreToolUse Detects file edits outside GSD workflow context (advisory, opt-in via hooks.workflow_guard)
gsd-secret-read-guard.js PreToolUse Hard-blocks Read / Grep / Bash reads of .env, .env.<suffix> (templates such as .env.example exempt) and .secrets; replaces the installer-written Read(.env*) permission deny rules, which made every cd DIR && grep … compound prompt for approval on Claude Code ≥ 2.1.259 (#4221)
gsd-read-guard.js PreToolUse Advisory guard preventing Edit/Write on files not yet read in the session
gsd-session-state.sh SessionStart Session state tracking for shell-based runtimes
gsd-validate-commit.sh PreToolUse Commit validation for conventional commit enforcement
gsd-phase-boundary.sh PostToolUse Phase boundary detection for workflow transitions

See docs/INVENTORY.md for the authoritative hook roster.

Crash policy (ADR-3889 Phase 7, #3911). Every enforcement hook terminates through hooks/lib/hook-exit.js's allow(payload) (exit 0), deny(payload, stderrPayload?) (exit 2), or crash(onCrash, payload) — the last dispatching per a HOOK_ON_CRASH policy the hook must declare explicitly (ALLOW or DENY, no default), so a hook's fail-open/fail-closed stance is a visible declaration rather than an inference from a bare process.exit(N). Two hooks are deliberate exceptions — gsd-read-injection-scanner.js (PostToolUse) and gsd-cursor-subagent-start.js (Cursor) — whose harnesses read the block decision from the JSON response body at exit 0, not from the exit code, so they never call deny(). See Declare a hook's crash policy.

Command Routing Hub (gsd-core/bin/lib/command-routing-hub.cjs)

CJS command family routers dispatch through CommandRoutingHub. The hub owns the no-throw pure-result contract (hub.dispatch() catches internal exceptions and returns { ok: false, kind, ...typedPayload }) and the closed runtime error taxonomy (UnknownCommand, InvalidArgs, HandlerRefusal, HandlerFailure). Router adapters remain thin CLI translators — they build the hub, call dispatch, then map the Result to output()/error() calls. The runtime is single-path (no dual-runtime mode selection). See docs/adr/0174-retire-gsd-sdk-package-boundary.md.

Planned (ADR-2346 / epic #2345): the runCommand 73-case switch is being dissolved into a two-layer dispatch — families via the commandFamilies registry (ADR-959 mechanism, completed) and single-purpose leaf verbs via a table filling the prepared _dispatchNonFamily seam — collapsing runCommand to a ~15-line dispatcher. Behavior-preserving; tracked phase-by-phase under epic #2345. The current-state description above holds until each phase lands.

Capability Command Dispatch (gsd-core/bin/gsd-tools.cjs, ADR-1244 D7)

Command families declared by capabilities (commands: [{ family, module, router }]) are dispatched from the registry rather than a hardcoded switch. The runCommand default arm tries, in order:

  1. First-party — dispatchCapabilityCommand against the frozen capability-registry.cjs commandFamilies, loading the router from bin/lib/. The in-tree families (graphify, intel, audit) reach their routers this way (the legacy hardcoded switch is retired).
  2. Third-party (installed overlay) — dispatchOverlayCapabilityCommand calls loadRegistry({ includeInstalled }) and dispatches a family only when its capId appears in _overlay.commandRoots. The loader lists a command root only for an accepted overlay capability with a committed ledger entry (consent gate), and the router module is require()'d from that capability's install root, confined by basename validation + realpath containment (rejecting .. traversal and symlink escape). This is the one point where third-party capability code executes; see the capability trust model for the consent + confinement + project-scope trust boundary.

Both paths share the same guards: prototype-pollution-safe command keys, an own-property router check, and synchronous-only routers (an async router is a fail-fast error).

Reviewer-Lane Capability Trait (#4209, ADR-2782)

/gsd-code-review optionally corroborates its internal review with external reviewer lanes (--codex, --agy, ...), gated by the reusable supportsReviewerLanes capability-step trait and dispatched through the single dispatchReviewerLanes interpreter — see gsd-core/references/loop-hook-dispatch.md for the trait and src/reviewer-step-dispatch.cts for the interpreter's fail-closed contract. gsd-code-reviewer is the sole consolidator: it independently re-verifies every external claim against the actual source before writing anything to REVIEW.md, so a lane's evidence is corroborating input, never a second output schema.

Research Module (src/research-{store,provider}.cts, src/package-legitimacy.cts)

The Research Module implements an L2-hybrid seam: code owns the cache, provider policy, and package legitimacy verdicts; MCP owns the actual network fetch.

Three compiled modules (generated to gsd-core/bin/lib/*.cjs per ADR-457) are reachable via gsd-tools query research-plan | research-store | package-legitimacy:

  • Research Store — content-addressed cache (sha256(ecosystem+library+version+query+kind)) with per-source TTL (curated-doc: 30 d, medium: 7 d, web/synthesis: 1 d) and two storage tiers: ~/.gsd/research-cache for cross-project curated-doc hits, .planning/research/.cache for project-local web/synthesis results.
  • Research Provider — single PROVIDER_WATERFALL (Context7→Ref→Jina→websearch for docs; Exa→Tavily→Perplexity→Brave→websearch for web; Firecrawl→Jina for scrape-only). planResearch() returns cache hits plus a fetch plan; classifyConfidence() stamps HIGH|MEDIUM|LOW by provider tier.
  • Package Legitimacy — registry-API verdicts (npm/PyPI/crates.io injectable adapters) producing OK|SUS|SLOP per package. slopcheck is an optional escalate-only adapter; absence leaves registry verdicts intact rather than downgrading everything to [ASSUMED].

Data flow:

agent
  │
  ▼
gsd-tools query research-plan          ← Research Provider: check cache, build fetch plan
  │
  ├── [cache hits] ──────────────────► RESEARCH.md (digest only, no raw content)
  │
  └── [fetch plan] ──────────────────► MCP fetch (agent calls MCP tools with the plan)
                                          │
                                          ▼
                                    gsd-tools query research-store (put)
                                          │
                                          ▼
                                    RESEARCH.md path returned to orchestrator

Agents always return a RESEARCH.md path, never raw fetched content. Context discipline is enforced through subagent isolation, compact provider output, and fetch-to-disk. See ADR-0656.

Context Predicate Fact-Store (src/context-predicates.cts, ADR-1671)

The CONTEXT.md predicate fact-store — every backtick-wrapped CLASS.subkey=value declaration in the repo-root CONTEXT.md — has a compiled parser/selector seam (generated to gsd-core/bin/lib/context-predicates.cjs per ADR-457) reachable live via gsd-tools query context-predicates --class|--prefix|--contains. Fence-aware line skipping mirrors markdown-sectionizer.cts's exported scanFencedBlocks delimiter-matching rule exactly (proven by a fence-skip parity test suite), but is scanned by a LOCAL, interleaved single pass rather than a call into that seam directly: fences and HTML comments must mutually suppress each other's open/close detection while either is active (a fence delimiter inside a real comment, or a comment token inside a real fence, must not falsely toggle the other construct), and that precedence cannot be resolved by two independent passes over scanFencedBlocks's comment-blind output — see src/context-predicates.cts's module doc comment.

scripts/gen-context-index.cjs --check is the CI drift-guard for the committed docs/CONTEXT-INDEX.json artifact: it fails on staleness between a fresh parse of CONTEXT.md and the committed file, and on any duplicate predicate ID. It is wired into lint:generated-sync (so lint:ci, so CI). docs/CONTEXT-INDEX.json is generated — never hand-edit it; regenerate with gen-context-index.cjs --write (also wired into build, after build:lib, and into regen:derived). The generator require()s the compiled context-predicates.cjs, so it must run after build:lib in any pipeline; .github/workflows/test.yml does this.

The committed index intentionally carries no line field for any predicate (ADR-1671 open question 4, resolved by #2928) — committed-but-uncompared metadata goes silently stale, the same defect class the drift-guard exists to catch, with the alarm removed. The live gsd-tools query context-predicates parse still returns line/section for callers that want to cite a source location. See ADR-1671 and CLI Tools Reference.

Workflow Fragmentization and Emission (src/workflow-fragments.cts, ADR-1671)

Workflow markdown under gsd-core/workflows/*.md can mark one or more sections with an in-file <!-- gsd:section id="<id>" when="<when>" --> / <!-- /gsd:section --> pair. A compiled parser/composer seam (generated to gsd-core/bin/lib/workflow-fragments.cjs per ADR-457) partitions a marked document into fragments and recomposes them through the shared context-composer.cjs budget seam (ADR-1671, #2929) before any per-runtime converter sees the text — so a marker attribute can never be corrupted by a .claude/ → .windsurf/-style path-rewrite regex. bin/install.js's copyWithPathReplacement calls composeWorkflow on every workflow file at emit time; an unmarked file (88 of the 89 shipped workflows today) parses to a single implicit fragment and round-trips byte-identical, so this is a no-op for every workflow that hasn't opted in yet.

Every fragment in this phase carries the verbatim strategy, so composition is structurally non-lossy — nothing is trimmed regardless of budget. Fence and HTML-comment interleaving reuses the same LOCAL, single-pass, mutually-suppressing scan discipline as context-predicates.cts (see above), so a marker-shaped line inside a fenced code block or an unrelated comment is never misread as structural. Markers are stripped at emit — the installed artifact carries no build metadata and is smaller than the source by exactly the stripped marker bytes.

See Reference: Workflow fragments for the full marker grammar, the frozen when= vocabulary, and fail-closed authoring rules, and ADR-1671 (open questions 1 and 2) for why in-file markers were chosen over separate fragment files or a sidecar manifest.

Section Manifest (src/section-manifest.cts, ADR-1671 Phases 5 and 6.1)

Two seams turn a workflow's gsd:section markers into per-invocation applicability data. scripts/gen-section-manifest.cjs --write (wired into build after build:lib, and into lint:generated-sync) scans gsd-core/workflows/*.md and writes the committed gsd-core/workflows/section-manifest.json, keyed per workflow — {workflows: {"<name>": [{id, when, read}]}} — where read is the path of the step file the section's body was extracted to. A workflow with no marked sections contributes no key at all: an absent key means degraded/unknown (the caller reads every section, the safe superset), while a key present with an empty array means "computed, nothing applies". The generator reuses parseWorkflowSections unchanged rather than re-implementing marker parsing, and fails closed (--check) on a marker naming a step file that does not exist, a step file no marker references, or a committed artifact still carrying the pre-6.1 flat {sections: [...]} shape.

A separate pure evaluator, src/section-manifest.cts (compiled to gsd-core/bin/lib/section-manifest.cjs per ADR-457), maps one invocation's facts — {flags, phaseNumber, hasPriorPhases} plus the optional needsCodebaseMap, phaseMvpMode and worktreesEnabled booleans — to an included/excluded partition of section ids via selectSections. flags is a ReadonlySet<string> of flag tokens; because parseNamedArgs always materializes a boolean flag key (false when the token was absent, never undefined), presence is truthiness, and the init router folds a boolean flag's own false into the absent sentinel before the facts are built. Per Greenspun's Tenth Rule, the evaluator is a total lookup over the frozen 14-atom when= vocabulary, never a parser: WHEN_PREDICATES is a hand-written literal map that never derives a predicate from its atom string, and an unrecognized value fails closed rather than being silently excluded. An atom is admitted only when it has both a real consuming section and a fact the init seam actually computes — an atom without the latter would evaluate false forever and silently disable its own section.

execute-phase.md's partial-wave and gap-closure-artifacts sections — previously inlined directly per #2930's pilot — now delegate to dedicated step files under gsd-core/workflows/execute-phase/steps/, the same pattern the pre-existing regression-gate section already used.

CLI Tools (gsd-core/bin/)

Node.js CLI utility (gsd-tools.cjs) with domain modules split across gsd-core/bin/lib/ (see docs/INVENTORY.md for the authoritative roster):

Module Responsibility
config-loader.cjs Project config loading — defaults merge, legacy-key migration, workstream overlay, unknown-key/profile-override validation, and federated config overlay (ADR-857 phase 3b) (extracted from core.cjs, ADR-857)
federated-config.cjs Defensive merge of capability-declared config slices (ADR-857 phase 3b); exports mergeFederatedConfig; live for migrated Capability keys that are absent from the central config schema
core-utils.cjs Shared low-level utility primitives — POSIX path normalization, sub-repo/subdirectory scanning, phase file stats, slug/one-liner/plan-id helpers, time-ago (extracted from core.cjs, ADR-857)
core.cjs Shared utilities; compatibility re-exports for planning, I/O (io.cjs), and phase-id helpers
io.cjs CLI I/O primitives — output/error emission, JSON-error mode, large-payload temp-file spillover
phase-id.cjs Pure phase-id parsing/matching helpers — normalize, token match, regex builders (extracted from core.cjs, ADR-857)
phase-locator.cjs Phase-directory search and location — active-phase discovery (searchPhaseInDir, findPhaseInternal) and archived-phase-dir enumeration (getArchivedPhaseDirs), matching phase ids/tokens against the filesystem (extracted from core.cjs, ADR-857)
roadmap-parser.cjs ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from core.cjs, ADR-857)
planning-workspace.cjs Planning seam (planningDir, planningPaths, active workstream routing, .planning/.lock)
state.cjs STATE.md parsing, updating, progression, metrics
phase.cjs Phase directory operations, decimal numbering, plan indexing
roadmap.cjs ROADMAP.md parsing, phase extraction, plan progress
config.cjs config.json read/write, section initialization
verify.cjs Plan structure, phase completeness, reference, commit validation
template.cjs Template selection and filling with variable substitution
frontmatter.cjs YAML frontmatter CRUD operations
init.cjs Compound context loading for each workflow type
milestone.cjs Milestone archival, requirements marking
commands.cjs Misc commands (slug, timestamp, todos, scaffolding, stats)
model-profiles.cjs Model profile resolution table
model-resolver.cjs Model and effort resolution policy — resolves model, tier, granularity, effort, and fast-mode for a given agent from project config and model profiles/catalog (extracted from core.cjs, ADR-857)
security.cjs Path traversal prevention, prompt injection detection, safe JSON parsing, shell argument validation
uat.cjs UAT file parsing, verification debt tracking, audit-uat support
docs.cjs Docs-update workflow init, Markdown scanning, monorepo detection
workstream.cjs Workstream CRUD, migration, session-scoped active pointer
schema-detect.cjs Schema-drift detection for ORM patterns (Prisma, Drizzle, etc.)
profile-pipeline.cjs User behavioral profiling data pipeline, session file scanning
profile-output.cjs Profile rendering, USER-PROFILE.md and dev-preferences.md generation
context-predicates.cjs CONTEXT.md predicate fact-store parser/selector (ADR-1671, #2928); backs query context-predicates and scripts/gen-context-index.cjs's docs/CONTEXT-INDEX.json drift guard; compiled from src/context-predicates.cts
loop-host-contract.cjs Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts; emitted by scripts/gen-loop-host-contract.cjs from workflow markers (ADR-894 §3); consumed by gen-capability-registry.cjs
capability-loader.cjs Runtime registry overlay loader (ADR-1244 D2) — loadRegistry({ includeInstalled }) composes the frozen first-party registry with a validated installed overlay of third-party capability manifests read from global $GSD_HOME/.gsd/capabilities/ and project <projectRoot>/.gsd/capabilities/; first-party always wins; load-time engines.gsd re-gate skips incompatible overlays with a warning; gate-kind hooks on skipped capabilities fail OPEN — no gate is injected; a loud warning (stderr + envelope warnings) names the load failure and the gsd capability remove <id> remediation (#2009)
capability-registry.cjs Generated central Capability Registry — role-partitioned index of all co-located capability declarations; emitted by scripts/gen-capability-registry.cjs (ADR-894 §5)
loop-resolver.cjs Loop Extension Point resolver — ADR-857 phase 3c registry-consuming query; consumes resolved Capability State, filters byLoopPoint by capability enablement plus config activation, renders active hooks as markdown, emits { point, activeHooks, rendered } envelope; gsd-tools loop render-hooks <point> [--config-dir <path>]
capability-state.cjs Unified capability-state resolver — ADR-857 phase 4b/6; composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; pure resolveCapabilityState, reusable resolveCapabilityRuntimeState, I/O cmdCapabilityState, and convenience predicate isCapabilityActive(capId, cwd); gsd-tools capability state [--config-dir <path>] emits { runtimeConfigDir, capabilities[] } where each entry carries enabled (installed && surfaced) and active (enabled && configActivation via the capability's activationKey; absent key → active===enabled)
capability-validator.cjs Shared capability conformance validator (ADR-1244 D2) — extracted from scripts/gen-capability-registry.cjs so the build-time generator and the runtime overlay loader share one validateCapability(manifest) implementation; generative-parity is CI-guarded
graphify-command-router.cjs ADR-959 capability command router — first real capability command cutover (phase 4d-impl-2); extracted from the case 'graphify': arm in gsd-tools.cjs; dispatches build/query/status/diff subcommands; discovered via commandFamilies in the capability registry
audit-command-router.cjs ADR-959 capability command router (phase 4d-impl-3); extracted from the case 'audit-uat': and case 'audit-open': arms in gsd-tools.cjs; routeAuditUat → uat.cjs:cmdAuditUat, routeAuditOpen → audit.cjs:{auditOpenArtifacts,formatAuditReport}; discovered via commandFamilies in the capability registry
intel-command-router.cjs ADR-959 capability command router (phase 4d-impl-4, last first-party cutover); extracted from the case 'intel': arm in gsd-tools.cjs; routeIntelCommand → all 9 intel subcommands via lazy require('./intel.cjs'); preserves non-raw timeAgo transform on status.files[*].updated_at; discovered via commandFamilies in the capability registry
runtime-hooks-surface.cjs Standalone hook-surface writer module (ADR-857 phase 5f-1); owns Cline rules/agents-md/pre-tool-use hook generation, Cursor hooks.json reconciliation, Copilot session-hook config, and Codex hook-block management; extracted verbatim from bin/install.js with no logic change.

Agent Model

Orchestrator → Agent Pattern

Orchestrator (workflow .md)
    │
    ├── Load context: gsd-tools.cjs init <workflow> <phase>
    │   Returns JSON with: project info, config, state, phase details
    │
    ├── Resolve model: gsd-tools.cjs resolve-model <agent-name>
    │   Returns: opus | sonnet | haiku | inherit
    │
    ├── Spawn Agent (Task/SubAgent call)
    │   ├── Agent prompt (agents/*.md)
    │   ├── Context payload (init JSON)
    │   ├── Model assignment
    │   └── Tool permissions
    │
    ├── Collect result
    │
    └── Update state: gsd-tools.cjs state update / state patch / state advance-plan

Primary Agent Spawn Categories

Conceptual spawn-pattern taxonomy for the primary agents. For the authoritative agent roster (including the advanced/specialized agents such as gsd-pattern-mapper, gsd-code-reviewer, gsd-code-fixer, gsd-ai-researcher, gsd-domain-researcher, gsd-eval-planner, gsd-eval-auditor, gsd-framework-selector, gsd-debug-session-manager, gsd-intel-updater), see docs/INVENTORY.md.

Category Agents Parallelism
Researchers gsd-project-researcher, gsd-phase-researcher, gsd-ui-researcher, gsd-advisor-researcher 4 parallel (stack, features, architecture, pitfalls); advisor spawns during discuss-phase
Synthesizers gsd-research-synthesizer Sequential (after researchers complete)
Planners gsd-planner, gsd-roadmapper Sequential
Checkers gsd-plan-checker, gsd-integration-checker, gsd-ui-checker, gsd-nyquist-auditor Sequential (verification loop, max 3 iterations)
Executors gsd-executor Parallel within waves, sequential across waves
Verifiers gsd-verifier Sequential (after all executors complete)
Mappers gsd-codebase-mapper 4 parallel (tech, arch, quality, concerns)
Debuggers gsd-debugger Sequential (interactive)
Auditors gsd-ui-auditor, gsd-security-auditor Sequential
Doc Writers gsd-doc-writer, gsd-doc-verifier Sequential (writer then verifier)
Profilers gsd-user-profiler Sequential
Analyzers gsd-assumptions-analyzer Sequential (during discuss-phase)

Wave Execution Model

During execute-phase, plans are grouped into dependency waves:

Wave Analysis:
  Plan 01 (no deps)      ─┐
  Plan 02 (no deps)      ─┤── Wave 1 (parallel)
  Plan 03 (depends: 01)  ─┤── Wave 2 (waits for Wave 1)
  Plan 04 (depends: 02)  ─┘
  Plan 05 (depends: 03,04) ── Wave 3 (waits for Wave 2)

Each executor gets:

  • Fresh 200K context window (or up to 1M for models that support it)
  • The specific PLAN.md to execute
  • Project context (PROJECT.md, STATE.md)
  • Phase context (CONTEXT.md, RESEARCH.md if available)

Adaptive Context Enrichment (1M Models)

When the context window is 500K+ tokens (1M-class models like Opus 4.6, Sonnet 4.6), subagent prompts are automatically enriched with additional context that would not fit in standard 200K windows:

  • Executor agents receive prior wave SUMMARY.md files and the phase CONTEXT.md/RESEARCH.md, enabling cross-plan awareness within a phase
  • Verifier agents receive all PLAN.md, SUMMARY.md, CONTEXT.md files plus REQUIREMENTS.md, enabling history-aware verification

The orchestrator reads context_window from config (gsd-tools.cjs config-get context_window) and conditionally includes richer context when the value is >= 500,000. For standard 200K windows, prompts use truncated versions with cache-friendly ordering to maximize context efficiency.

Parallel Commit Safety

When multiple executors run within the same wave, two mechanisms prevent conflicts:

  1. --no-verify commits — Parallel agents skip pre-commit hooks (which can cause build lock contention, e.g., cargo lock fights in Rust projects). The orchestrator runs git hook run pre-commit once after each wave completes.
  2. STATE.md file locking — All writeStateMd() calls use lockfile-based mutual exclusion (STATE.md.lock with O_EXCL atomic creation). This prevents the read-modify-write race condition where two agents read STATE.md, modify different fields, and the last writer overwrites the other's changes. Includes stale lock detection (10s timeout) and spin-wait with jitter.

The STATE.md Write Path

Locking decides who writes. A separate contract decides what survives the write.

STATE.md carries the same fact in two places — YAML frontmatter and the document body — and the body is authoritative. Every write therefore re-derives frontmatter from the body, which raises the question the write path exists to answer: when a re-derived value disagrees with the one already in frontmatter, which wins?

FIELD_CLASSIFICATION (src/state-transition.cts) answers it per field, declaring a preservation policy — preserve-when-unchanged, preserve-always, preserve-if-placeholder, derive — that applyStatePreservation executes after syncStateFrontmatter re-derives. (A fifth policy, clear, was listed here until ADR-3408 §8.6's amendment removed it: no row used it and no executor existed for it.)

The pipeline's precondition is a type, not a convention (ADR-3473 §8.6). A policy row can only be honored if the pre-write frontmatter snapshot it compares against is actually present. That snapshot now travels as a StateTransaction, built by openStateTransaction() — preservation applies — or rebuildStateTransaction() — it does not. Both carry the snapshot, and a transaction cannot be constructed without one: an absent snapshot is a construction failure, not a runtime skip. That distinction is the whole point. Previously the snapshot was nulled to signal "re-derive from disk", so a declared preserve-always row and a silently-skipped one were indistinguishable at runtime, which is how a curated progress: block was erased by verbs that had nothing to do with progress.

rebuildStateTransaction() is the typed form of ADR-3408 §8.3's closed exception list: state sync, which exists to let the body win, and /gsd-health --repair's factory reset. Both are deliberate and permanent, not debt — and because the type names them, the write-path drift guard no longer has to track them as strings in a ratcheted baseline.

What a command reports it wrote is the same snapshot, read back (ADR-3473 §8.7). Every state.* command returns an updated array. That array is now derived by comparing what was actually persisted against the transaction's pre-write snapshot: a field appears if, and only if, its persisted value changed. One comparison answers both of the questions that used to need separate machinery — a field the caller asked for that the pipeline then discarded is persisted-equals-snapshot and drops out, and a field nobody asked for that the write moved anyway is different and appears. Nothing is filtered by its preservation policy.

Reporting is at leaf granularity: when a single counter moves you are told progress.total_plans, not progress. The leaves are the ones the field-classification table already declares, so the report is bounded by a schema rather than by walking the document.

One field is excluded, and it is excluded for its provenance rather than its policy: last_updated is stamped on every save regardless of what you changed, so admitting it would make state patch's success signal — which is simply whether updated is non-empty — permanently true, and a patch in which every field failed would report success. state_head is deliberately not excluded: it is recomputed on every save but only changes when the commit it records actually moved, so reporting it tells you something true.

A practical consequence worth knowing: these arrays are now longer than they used to be, because they used to under-report. If you compare one exactly, expect more entries — and expect them to be the ones that really changed.

ADR-3408 is the normative contract for that path: one executor per declared policy, one write seam, and reports computed from what was actually persisted rather than from what the caller intended to write. Where the contract and the code disagree, the code is the defect. It is the write-side counterpart of ADR-3180, which gave each read-side derivation a single owner.

ADR-3473 owns the invariants that sit outside that contract. ADR-3408 governs what survives a write; it does not govern the pipeline's precondition, what a command reports it wrote, or where the set of STATE.md keys, types and enums is declared. ADR-3473 owns those, alongside document parsing, enumeration, and the return contract of every routine that can fail. It is the third application of ADR-3180's mechanism and the first whose success metric requires the guard surface to shrink as each seam lands.


Data Flow

New Project Flow

User input (idea description)
    │
    ▼
Questions (questioning.md philosophy)
    │
    ▼
4x Project Researchers (parallel)
    ├── Stack → STACK.md
    ├── Features → FEATURES.md
    ├── Architecture → ARCHITECTURE.md
    └── Pitfalls → PITFALLS.md
    │
    ▼
Research Synthesizer → SUMMARY.md
    │
    ▼
Requirements extraction → REQUIREMENTS.md
    │
    ▼
Roadmapper → ROADMAP.md
    │
    ▼
User approval → STATE.md initialized

Phase Execution Flow

discuss-phase → CONTEXT.md (user preferences)
    │
    ▼
ui-phase → UI-SPEC.md (design contract, optional)
    │
    ▼
plan-phase
    ├── Research gate (blocks if RESEARCH.md has unresolved open questions)
    ├── Phase Researcher → RESEARCH.md
    │       └── Package Legitimacy Gate: registry-API verdict on every package; [SLOP] removed,
    │           [SUS]/[ASSUMED] flagged; Audit table written to RESEARCH.md
    ├── Planner (with reachability check) → PLAN.md files
    │       └── checkpoint:human-verify injected before [ASSUMED]/[SUS] installs;
    │           T-{phase}-SC STRIDE row added for install-bearing plans
    ├── Plan Checker → Verify loop (max 3x)
    ├── Requirements coverage gate (REQ-IDs → plans)
    └── Decision coverage gate (CONTEXT.md `<decisions>` → plans, BLOCKING — #2492)
    │
    ▼
state planned-phase → STATE.md (Planned/Ready to execute)
    │
    ▼
execute-phase (context reduction: truncated prompts, cache-friendly ordering)
    ├── Wave analysis (dependency grouping)
    ├── Executor per plan → code + atomic commits
    ├── SUMMARY.md per plan
    └── Verifier → VERIFICATION.md
        └── Decision coverage gate (CONTEXT.md decisions → shipped artifacts, NON-BLOCKING — #2492)
    │
    ▼
verify-work → UAT.md (user acceptance testing)
    │
    ▼
ui-review → UI-REVIEW.md (visual audit, optional)

Context Propagation

Each workflow stage produces artifacts that feed into subsequent stages:

PROJECT.md ────────────────────────────────────────────► All agents
REQUIREMENTS.md ───────────────────────────────────────► Planner, Verifier, Auditor
ROADMAP.md ────────────────────────────────────────────► Orchestrators
STATE.md ──────────────────────────────────────────────► All agents (decisions, blockers)
CONTEXT.md (per phase) ────────────────────────────────► Researcher, Planner, Executor
RESEARCH.md (per phase) ───────────────────────────────► Planner, Plan Checker
PLAN.md (per plan) ────────────────────────────────────► Executor, Plan Checker
SUMMARY.md (per plan) ─────────────────────────────────► Verifier, State tracking
UI-SPEC.md (per phase) ────────────────────────────────► Executor, UI Auditor

File System Layout

Installation Files

~/.claude/                          # Claude Code (global install)
├── skills/gsd-ns-*/SKILL.md        # Global skills — nesting runtimes: 6 namespace routers (authoritative roster: docs/INVENTORY.md)
│   └── skills/<name>/SKILL.md     #   concrete skills nested under each router
│   (flat runtimes: skills/gsd-*/SKILL.md — all ~67 skills at top level)
├── commands/gsd/*.md               # Local Claude installs use slash commands instead of global skills
├── gsd-core/
│   ├── bin/gsd-tools.cjs           # CLI utility
│   ├── bin/lib/*.cjs               # Domain modules (authoritative roster: docs/INVENTORY.md)
│   ├── workflows/*.md              # Workflow definitions (authoritative roster: docs/INVENTORY.md)
│   ├── references/*.md             # Shared reference docs (authoritative roster: docs/INVENTORY.md)
│   └── templates/                  # Planning artifact templates
├── agents/*.md                     # Agent definitions (authoritative roster: docs/INVENTORY.md)
├── hooks/*.js                      # Node.js hooks (statusline, guards, monitors, update check)
├── hooks/*.sh                      # Shell hooks (session state, commit validation, phase boundary)
├── settings.json                   # Hook registrations
└── VERSION                         # Installed version number

Equivalent paths for other runtimes:

  • OpenCode: ~/.config/opencode/ global or ./.opencode/ local
  • Kilo: ~/.config/kilo/ global or ./.kilo/ local
  • Kimi CLI: first-existing generic global root (~/.config/agents/ recommended, then ~/.agents/ if its skills/ directory already exists); local install is deferred and guarded
  • Codex: ~/.codex/ global or ./.codex/ local
  • Copilot: ~/.copilot/ global or ./.github/ local
  • Antigravity: auto-detected global root (~/.gemini/antigravity/, ~/.gemini/antigravity-ide/, or ~/.gemini/antigravity-cli/) for settings and runtime files; global skills/agents under ~/.gemini/config/ (the machine-local discovery dir, #3738) or ./.agent/ local
  • Cursor: ~/.cursor/ global or ./.cursor/ local
  • Windsurf/Devin Desktop: ~/.codeium/windsurf/ global config or ./.windsurf/ local workflows
  • Augment Code: ~/.augment/ global or ./.augment/ local
  • Trae: ~/.trae/ global or ./.trae/ local
  • Qwen Code: ~/.qwen/ global or ./.qwen/ local
  • Hermes Agent: ~/.hermes/ global or ./.hermes/ local
  • CodeBuddy: ~/.codebuddy/ global or ./.codebuddy/ local
  • Cline: ~/.cline/ global or project-root .clinerules local

Project Files (.planning/)

.planning/
├── PROJECT.md              # Project vision, constraints, decisions, evolution rules
├── REQUIREMENTS.md         # Scoped requirements (v1/v2/out-of-scope)
├── ROADMAP.md              # Phase breakdown with status tracking
├── STATE.md                # Living memory: position, decisions, blockers, metrics
├── config.json             # Workflow configuration
├── MILESTONES.md           # Completed milestone archive
├── research/               # Domain research from /gsd-new-project
│   ├── SUMMARY.md
│   ├── STACK.md
│   ├── FEATURES.md
│   ├── ARCHITECTURE.md
│   └── PITFALLS.md
├── codebase/               # Brownfield mapping (from /gsd-map-codebase or /gsd-onboard)
├── onboarding/             # Brownfield onboarding summary (from /gsd-onboard)
│   ├── STACK.md            # YAML frontmatter carries `last_mapped_commit`
│   ├── ARCHITECTURE.md     # for the post-execute drift gate (#2003)
│   ├── CONVENTIONS.md
│   ├── CONCERNS.md
│   ├── STRUCTURE.md
│   ├── TESTING.md
│   └── INTEGRATIONS.md
├── phases/
│   └── XX-phase-name/
│       ├── XX-CONTEXT.md       # User preferences (from discuss-phase)
│       ├── XX-RESEARCH.md      # Ecosystem research (from plan-phase)
│       ├── XX-YY-PLAN.md       # Execution plans
│       ├── XX-YY-SUMMARY.md    # Execution outcomes
│       ├── XX-VERIFICATION.md  # Post-execution verification
│       ├── XX-VALIDATION.md    # Nyquist test coverage mapping
│       ├── XX-UI-SPEC.md       # UI design contract (from ui-phase)
│       ├── XX-UI-REVIEW.md     # Visual audit scores (from ui-review)
│       └── XX-UAT.md           # User acceptance test results
├── quick/                  # Quick task tracking
│   └── YYMMDD-xxx-slug/
│       ├── PLAN.md
│       └── SUMMARY.md
├── todos/
│   ├── pending/            # Captured ideas
│   └── completed/          # Completed todos
├── threads/               # Persistent context threads (from /gsd-thread)
├── seeds/                 # Forward-looking ideas (from /gsd-capture --seed)
├── debug/                  # Active debug sessions
│   ├── *.md                # Active sessions
│   ├── resolved/           # Archived sessions
│   └── knowledge-base.md   # Persistent debug learnings
├── ui-reviews/             # Screenshots from /gsd-ui-review (gitignored)
└── continue-here.md        # Context handoff (from pause-work)

Post-Execute Codebase Drift Gate (#2003)

After the last wave of /gsd-execute-phase commits, the workflow runs a non-blocking codebase_drift_gate step (between schema_drift_gate and verify_phase_goal). It compares the diff last_mapped_commit..HEAD against .planning/codebase/STRUCTURE.md and counts four kinds of structural elements:

  1. New directories outside mapped paths
  2. New barrel exports at (packages|apps)/<name>/src/index.*
  3. New migration files
  4. New route modules under routes/ or api/

If the count meets workflow.drift_threshold (default 3), the gate either warns (default) with the suggested /gsd-map-codebase --paths … command, or auto-remaps (workflow.drift_action = auto-remap) by spawning gsd-codebase-mapper scoped to the affected paths. Any error in detection or remap is logged and the phase continues — drift detection cannot fail verification.

last_mapped_commit lives in YAML frontmatter at the top of each .planning/codebase/*.md file; bin/lib/drift.cjs provides readMappedCommit and writeMappedCommit round-trip helpers.


Installer Architecture

The installer (bin/install.js, ~10,700 lines) handles:

  1. Runtime detection — Interactive prompt or CLI flags (--claude, --opencode, --kimi, --kilo, --codex, --copilot, --antigravity, --cursor, --windsurf, --augment, --trae, --qwen, --hermes, --codebuddy, --cline, --all)
  2. Location selection — Global (--global) or local (--local)
  3. File deployment — Copies commands, skills, workflows, references, templates, agents, and hooks
  4. Runtime adaptation — Transforms file content per runtime:
  • Claude Code: Uses as-is
  • OpenCode: Converts commands/agents to OpenCode-compatible flat command + subagent format
  • Kilo: Reuses the OpenCode conversion pipeline with Kilo config paths
  • Codex: Generates TOML config + skills from commands
  • Kimi CLI: Generates Agent Skills under skills/gsd-*/SKILL.md, custom agent YAML/prompt files, and explicit kimi_cli.tools.* module paths
  • Copilot: Maps tool names (Read→read, Bash→execute, etc.)
  • Antigravity: Skills-first with Google model equivalents; adjusts hook event names (AfterTool instead of PostToolUse)
  • Cursor: Skills-first with Cursor rule references
  • Windsurf: Skills-first with Windsurf rule references
  • Trae: Skills-first install to ~/.trae / ./.trae with no settings.json or hook integration
  • Qwen Code: Skills-first with Qwen-branded path and prompt rewrites
  • Hermes Agent: Category-based skills under skills/gsd/
  • CodeBuddy: Skills-first with CodeBuddy path and prompt rewrites
  • Cline: Writes .clinerules for rule-based integration
  • Augment Code: Skills-first with full skill conversion and config management
  1. Path normalization — Replaces ~/.claude/ paths with runtime-specific paths
  2. Settings integration — Registers hooks in runtime's settings.json
  3. Patch backup — Since v1.17, backs up locally modified files to gsd-local-patches/ for /gsd-update --reapply
  4. Manifest tracking — Writes gsd-file-manifest.json for clean uninstall. The manifest also records which runtime and which scope (global/local) wrote it, under a manifestVersion schema field, so a reader can answer "which surfaces are installed, at which scopes" without inferring it from the directory the file sits in (ADR 2866, #2872). Manifests written before that carry no such fields and are read without error — no reinstall is required. See Installer Migrations → File Manifest
  5. Uninstall mode — --uninstall removes all GSD files, hooks, and settings

installRuntimeArtifacts (install-engine.cjs) returns the executed plan it ran — per kind, per scope, including on the combined OpenCode/Kilo family path, which previously early-returned void — rather than being observable only by re-reading disk afterward. Its destination-writing IO (copies, removals, snapshot/restore, best-effort cleanup) now routes through an injectable fs seam, install-fs-adapter.cjs, so a full install can be exercised against a fake adapter with zero real destination IO; locating this package's own source tree remains real by design (a destination-fake is never seeded with the repo's own paths). Writes stay byte-identical and existing void-ignoring callers are unaffected. This completes ADR 58's registry → adapter → helpers → cleanup rollout — the cleanup step had not previously landed (#2874, epic #2866 Phase 5).

Install-time file moves, stale-artifact cleanup, config rewrites, and user-data preservation are governed by the Installer Migration Module. See Installer Migrations and ADR 0008. The migration module also owns the gated first-time baseline scan for legacy installs, classifying known runtime install surfaces before later migrations remove or rewrite anything.

The plan drift guard (plan_review.source_grounding) — which verifies symbol references in generated plans against live source before execution — is specified in ADR 22.

The same switch gates a second, cross-artifact axis: a fact-drift pass that compares the same fact as stated in ROADMAP.md, PLAN.md, STATE.md and CONTEXT.md and reports contradictions (a phase status, a success criterion, a requirement ID, a glossary term) with both locations and the authoritative side named. Where the source-grounding axis grounds a plan against code, this one grounds the planning artifacts against each other. It keys on contradicting knowledge rather than similar-looking text, and is advisory only — it never sets hardBlock and never contributes to the convergence counts.

Platform Handling

  • Windows: windowsHide on child processes, EPERM/EACCES protection on protected directories, path separator normalization
  • WSL: Detects Windows Node.js running on WSL and warns about path mismatches
  • Docker/CI: Supports CLAUDE_CONFIG_DIR env var for custom config directory locations

Hook System

Architecture

Runtime Engine (Claude Code / Antigravity CLI)
    │
    ├── statusLine event ──► gsd-statusline.js
    │   Reads: stdin (session JSON)
    │   Writes: stdout (formatted status), /tmp/claude-ctx-{session}.json (bridge)
    │
    ├── PostToolUse/AfterTool event ──► gsd-context-monitor.js
    │   Reads: stdin (tool event JSON), /tmp/claude-ctx-{session}.json (bridge)
    │   Writes: stdout (hookSpecificOutput with additionalContext warning)
    │
    └── SessionStart event
        ├──► gsd-ensure-canonical-path.js   (runs first)
        │    Reads:  ${CLAUDE_PLUGIN_ROOT}/gsd-core/ (plugin installs only)
        │    Writes: ~/.claude/gsd-core/{bin,contexts,references,templates,workflows} symlinks
        │            (no-op in classic installs; preserves user files; self-heals)
        └──► gsd-check-update.js
             Reads:  VERSION file
             Writes: ~/.claude/cache/gsd-update-check.json (spawns background process)

Context Monitor Thresholds

Remaining Context Level Agent Behavior
> 35% Normal No warning injected
≤ 35% WARNING "Avoid starting new complex work"
≤ 25% CRITICAL "Context nearly exhausted, inform user"

Debounce: 5 tool uses between repeated warnings. Severity escalation (WARNING→CRITICAL) bypasses debounce.

Safety Properties

  • All hooks wrap in try/catch, exit silently on error
  • stdin timeout guard (3s) prevents hanging on pipe issues
  • Stale metrics (>60s old) are ignored
  • Missing bridge files handled gracefully (subagents, fresh sessions)
  • Context monitor is advisory — never issues imperative commands that override user preferences

Package Legitimacy Gate (v1.42.1)

The researcher → planner → executor pipeline includes a supply-chain gate against slopsquatting (AI-hallucinated package names pre-registered with malicious post-install scripts).

Threat model: GSD automates the full path from "researcher names a package" to "executor runs npm install". A hallucinated name that passes npm view (proving only registration, not legitimacy) would previously flow through undetected. ~20% of AI-generated package references are hallucinated; ~43% of those names recur consistently across prompts, making pre-registration economically viable for attackers.

Gate layers:

Layer Component Action
Research gsd-phase-researcher Runs gsd-tools query package-legitimacy check --ecosystem <npm|pypi|crates> <pkgs>; writes ## Package Legitimacy Audit table to RESEARCH.md; strips [SLOP] packages before RESEARCH.md is written
Planning gsd-planner Reads Audit table; inserts checkpoint:human-verify before any [ASSUMED] or [SUS] install task; adds T-{phase}-SC STRIDE supply-chain row to <threat_model>
Execution gsd-executor RULE 3 excludes package installation from auto-fix scope; failed installs surface as checkpoints, never silent substitutions

Claim provenance integration: Package names discovered via WebSearch are tagged [ASSUMED] (not [VERIFIED]) regardless of the registry-API verdict. This extends the existing [ASSUMED] / [VERIFIED] / [CITED] provenance system by enforcing the provenance tag as a hard gate at the install boundary — [ASSUMED] always generates a checkpoint:human-verify in PLAN.md.

Ecosystem coverage: The gate resolves signals directly from each ecosystem's registry API rather than a single generic check — registry.npmjs.org + api.npmjs.org/downloads (Node), pypi.org/pypi/<pkg>/json (Python), the crates.io API (Rust). This catches cross-ecosystem hallucination (~9% rate documented in 2025 USENIX research).

Graceful degradation: Each registry adapter degrades to null signals (never throws) on a failed lookup; missing signals push a package to [SUS], which is gated behind the same checkpoint:human-verify checkpoint as [ASSUMED]. Research and planning proceed; the system never hard-fails on a network or tool outage. slopcheck is an optional escalate-only adapter — it can only raise a verdict, never lower it, and is not the install-or-degrade gate. No shipped configuration wires it.


Security Hooks (v1.27)

For a conceptual overview of how the hook and guard layers fit into the broader security approach, see Security model.

Prompt Guard (gsd-prompt-guard.js):

  • Triggers on Write/Edit to .planning/ files
  • Scans content for prompt injection patterns (role override, instruction bypass, system tag injection)
  • Advisory-only — logs detection, does not block
  • Patterns are inlined (subset of security.cjs) for hook independence
  • Output contract: hookSpecificOutput carries both additionalContext and findings — an array of { ruleId, match } records (INJECTION-PATTERN or INVISIBLE-UNICODE), module-local to this hook (not shared with gsd-read-injection-scanner.js's own RULE_IDS). The advisory is rendered from findings via a single mapper, so the two cannot disagree. Consumers should read findings rather than parsing the advisory text.

Read Injection Scanner (gsd-read-injection-scanner.js):

  • Triggers on Read / WebFetch / WebSearch PostToolUse events
  • Advisory by default; blocks only HIGH severity, and only when security.injection_blocking is true
  • Severity is LOW for 1-2 matched patterns, HIGH for 3 or more
  • Skips content shorter than 20 characters, and skips excluded paths (.planning/, REVIEW.md, CHECKPOINT*, security/injection docs, and GSD's own staged hook bundle)
  • Rule ids: the MD-LINK-* markdown-link rules mirrored from security.cjs's MARKDOWN_LINK_PATTERNS, plus INJECTION-PATTERN, INVISIBLE-UNICODE, and UNICODE-TAG-BLOCK
  • Patterns are shared with gsd-prompt-guard.js via hooks/lib/injection-patterns.js (#3504); the markdown-link list is inlined for hook independence
  • Output contract: hookSpecificOutput carries additionalContext (the human-readable advisory sentence), findings — an array of { ruleId, match } records naming each rule that fired — plus severity (LOW for 1-2 matches, HIGH for 3+) and source (the scanned file path, URL, or search: <query> string). findings and severity are the structured surface; the advisory is rendered from them, so the three cannot disagree. match is null for rules with no captured text (INVISIBLE-UNICODE, UNICODE-TAG-BLOCK). Consumers should read findings/severity/source rather than parsing the advisory text.

Workflow Guard (gsd-workflow-guard.js):

  • Triggers on Write/Edit to non-.planning/ files
  • Detects edits outside GSD workflow context (no active /gsd- command or Task subagent)
  • Advises using /gsd-quick or /gsd-fast for state-tracked changes
  • Opt-in via hooks.workflow_guard: true (default: false)
  • Output contract: the advisory leg's hookSpecificOutput carries code: 'WORKFLOW_ADVISORY' alongside additionalContext. This is distinct from the hook's separate force-add block leg (code: 'WORKTREE_AGENT_FORCE_ADD_FORBIDDEN', decision: 'block') — the two are disambiguated by code, never by presence.

Runtime Abstraction

GSD supports multiple AI coding runtimes through a unified command/workflow architecture:

Runtime Install Contract Matrix

This matrix describes the runtime surfaces the installer materializes today. The migration-specific ownership and source snapshots live in Installer Migrations.

Runtime Global root Local root Invocation surface Agent surface Config and hooks
Claude Code ~/.claude ./.claude Global skills/gsd-*/SKILL.md (flat, #924); local commands/gsd/*.md agents/gsd-*.md settings.json hook and statusLine entries
OpenCode ~/.config/opencode ./.opencode commands/gsd-*.md agents/gsd-*.md opencode.json or opencode.jsonc; no GSD hooks
Kilo ~/.config/kilo ./.kilo command/gsd-*.md agents/gsd-*.md kilo.json or kilo.jsonc; no GSD hooks
Kimi CLI First-existing generic root: ~/.config/agents recommended, then ~/.agents when ~/.agents/skills exists and ~/.config/agents/skills does not Deferred and guarded skills/gsd-*/SKILL.md (flat) invoked as /skill:gsd-* agents/gsd.yaml, agents/gsd.md, and agents/subagents/gsd-* YAML/prompt pairs Explicit kimi --agent-file <configRoot>/agents/gsd.yaml; no GSD hooks or statusline
Codex ~/.codex ./.codex skills/gsd-*/SKILL.md (flat) agents/ source markdown plus per-agent TOML (Codex auto-discovers each agents/gsd-*.toml; this is the sole canonical role registration, #2406) config.toml bare [agents] dispatch-tuning scalar (max_depth, no per-role [agents.gsd-*] tables), [features].hooks (canonical; legacy alias codex_hooks is recognized and migrated forward on reinstall, #3566), and hook tables
GitHub Copilot ~/.copilot ./.github skills/gsd-*/SKILL.md (flat), copilot-instructions.md, and AGENTS.md (repo root, local) .agent.md files Self-contained sessionStart hook (hooks/gsd-session.json, inline command type); no statusline
Antigravity auto-detected: ~/.gemini/antigravity, ~/.gemini/antigravity-ide, or ~/.gemini/antigravity-cli ./.agent ~/.gemini/config/skills/gsd-*/SKILL.md (flat, #1614; global home override #3738) ~/.gemini/config/agents/gsd-*.md (#3738) Gemini-style settings.json hook entries when installed by GSD
Cursor ~/.cursor ./.cursor skills/gsd-*/SKILL.md (flat) agents/gsd-*.md Rule references under rules/; hooks.json with sessionStart context injection and postToolUse STATE.md monitor (#777)
Windsurf ~/.codeium/windsurf config ./.windsurf workflows/gsd-*.md slash-command workflows No custom-agent artifact surface No GSD hooks
Augment Code ~/.augment ./.augment skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) agents/gsd-*.md No GSD hooks or statusline
Trae ~/.trae ./.trae skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) agents/gsd-*.md Rule references under rules/; no GSD hooks
Qwen Code ~/.qwen ./.qwen skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) agents/gsd-*.md Common GSD settings and hook entries where supported
Hermes Agent ~/.hermes ./.hermes skills/gsd/ns-*/SKILL.md (6 routers, prefix='') + skills/gsd/ns-*/skills/<name>/SKILL.md (nested concretes) agents/gsd-*.md Common GSD settings and hook entries where supported
CodeBuddy ~/.codebuddy ./.codebuddy skills/gsd-*/SKILL.md (flat, user-invocable: false) agents/gsd-*.md /gsd-* slash commands under commands/; common GSD settings and hook entries where supported
Cline ~/.cline project root skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) + .clinerules Rules only No GSD hooks or statusline

Upstream Contract Sources

Runtime install expectations are checked against primary documentation where available. The current source snapshot is 2026-05-11, with Kimi CLI rechecked on 2026-06-07:

  • Claude Code: Anthropic slash commands, settings, hooks, and subagents docs.
  • OpenCode and Kilo: OpenCode config docs and Kilo custom subagent docs.
  • Qwen Code: command/config docs; Qwen command docs were last updated 2026-05-06.
  • Kimi CLI: Agent Skills docs for user-level brand roots and first-existing generic roots (~/.config/agents/skills/ recommended, then ~/.agents/skills/), plus Agents docs for YAML files, system_prompt_path, kimi_cli.tools.* module paths, and explicit kimi --agent-file launch.
  • Codex: OpenAI Codex docs and config-schema.json; the installer also carries Codex 0.124.0 compatibility for agent table shape.
  • Copilot, Cursor, Cline, Augment, Hermes, and CodeBuddy: vendor docs for custom instructions, rules, skills, or config.
  • Antigravity, Windsurf, and Trae: source-limited rows. The installer documents current compatibility shims, and migrations must refresh those sources before rewriting their config.

Abstraction Points

  1. Tool name mapping — Each runtime has its own tool names (e.g., Claude's Bash → Copilot's execute)
  2. Hook event names — Claude uses PostToolUse, Antigravity uses AfterTool
  3. Agent frontmatter — Each runtime has its own agent definition format
  4. Path conventions — Each runtime stores config in different directories
  5. Model references — inherit profile lets GSD defer to runtime's model selection

The installer handles all translation at install time. Workflows and agents are written in Claude Code's native format and transformed during deployment.