* test(01-01): define reviewer-support trait contract Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02): validator rejects non-boolean values with an exact field path, accepts missing/true/false, and the real code-review capability.json steps must declare supportsReviewerLanes: true. Add loop-resolver projection coverage proving the trait reaches activeHooks verbatim for a provider-neutral synthetic step (not code-review-specific), and that omitted/false values stay inert (no key on the active hook). All 8 new assertions fail today: the validator has no such field, and loop-resolver has nothing to project. RED before GREEN. * feat(01-01): declare reviewer-capable steps Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional boolean opt-in trait, step-scoped (not capability-wide). Only a literal true validates and projects; false/omitted stay inert (no key on the projected active hook), and every non-boolean type fails capability-validator.cjs with an exact field-path error. Opt both existing code-review steps (execute:post, execute:wave:post) into the trait in capabilities/code-review/capability.json. Project the validated field through src/loop-resolver.cts into activeHooks so a provider-neutral generic interpreter can read it without any code-review-specific knowledge. Document the field in docs/reference/capability-manifest.md and regenerate gsd-core/bin/lib/capability-registry.cjs via the generator (never hand-edited). Makes all 8 RED assertions from the prior commit pass. * test(01-02): define shared reviewer dispatch - Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes: inert when the supportsReviewerLanes trait is off or nothing is selected, exactly-once plan/invoke per selected lane, duplicate-alias dedup, the bounded metadata-only source-review prompt (repo root, paths+baseSha, depth, four fixed prohibitions), and capability-neutral reuse via a second synthetic step context. - RED: module under test (src/reviewer-step-dispatch.cts) does not exist yet, so require() fails and every assertion is unreached. * feat(01-02): dispatch reviewers for opted-in steps - Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps), ONE interpreter for a step's supportsReviewerLanes trait. Reuses resolveReviewerSelection for selection and resolveLanePlan for planning (both already-existing, pure building blocks); invocation is the one required, caller-injected seam (deps.invoke) since runLane needs OS-aware spawn plumbing this module does not own. - trait !== true, or a selection resolving to zero lanes, dispatches nothing (zero plan/invoke calls). Each selected lane is planned and invoked exactly once, in the selector's deduped/sorted order. - buildSourceReviewPrompt assembles a metadata-only bounded prompt (repo root, canonical paths + base SHA, depth, four fixed prohibitions) — never file contents — written once per dispatch and shared across every invoked lane. - GREEN: tests/reviewer-step-dispatch.test.cjs now passes. * test(01-02): define reviewer dispatch failures - Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed matrix: an explicitly requested lane the selector could not resolve still lets the OTHER resolved lane run, but the aggregate result must never read as a clean success (and 'every explicit lane unavailable' must be distinguishable from the plain no-flags-passed inert case); request-level validation (path traversal, absolute paths outside repoRoot, empty/non-string paths, missing depth/base SHA) halts the whole dispatch before any lane is planned or invoked; a per-lane prompt-budget overflow hard-fails only that lane before invoke while its sibling still runs. - RED: src/reviewer-step-dispatch.cts does not yet implement any of these guards, so 9 of the new assertions fail against the current (Task 1) implementation. * fix(01-02): fail closed in reviewer dispatch - src/reviewer-step-dispatch.cts: add the fail-closed guards the prior commit deliberately left out. An explicitly requested lane the selector could not resolve no longer lets the aggregate read as a clean success — lanes that DID resolve still run and keep their results (never narrow the requested set), but selection.errors now flips the aggregate ok to false, and 'every explicit lane unavailable' is now distinguishable (SELECTION_FAILED) from the plain no-flags-passed inert case (NO_LANES_SELECTED). - Add request-level validation (validatePaths, depth/baseSha presence) that halts the WHOLE dispatch before any lane is planned or invoked: path traversal, absolute paths outside repoRoot, empty/non-string paths, and missing provenance are all rejected up front. - Add per-lane prompt-budget enforcement (resolveBudget, mirroring gsd-tools.cjs's budgetFor convention including budget 0 = unbounded): a lane whose resolved budget the prompt exceeds hard-fails before invoke runs for it, without cancelling a sibling lane already planned. - Document the supportsReviewerLanes trait and its dispatch-step interpreter in gsd-core/references/loop-hook-dispatch.md. - GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass; no regressions in the review-lane/reviewer-selection/prompt-budget suites (356 passing). * test(01-03): define optional source reviewer flow RED: assert code-review.md dispatches roster-derived reviewer-lane flags through a single review-lane dispatch-step call (DISP-01..05), that the no-flag path stays byte-for-behavior unchanged (COMP-01), and that external evidence reaching the internal reviewer prompt is marked unverified (CONS-02). Also covers the CLI contract directly: no-op with no explicit selection, and fail-closed on an explicit unknown lane (SAFE-07) via real gsd-tools.cjs subprocess calls. * feat(01-03): route optional source reviewers GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches canonical reviewer-lane flags against the merged first-party + installed roster (never a hand-maintained list) and, only when at least one is present, calls the shared reviewer-step interpreter exactly once with the already-resolved repo root, file scope, depth, and base SHA. Its evidence paths are appended to the internal reviewer prompt via ${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane flag leaves the internal-only dispatch byte-for-behavior unchanged (COMP-01). Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI route `dispatchReviewerLanes` wires through, but never implemented the gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add it to the existing review-lane router, reusing the same effort-aware plan building and runner deps `plan`/`invoke` already use (factored into buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard the CLI's own `detected` set on whether an explicit flag was passed: resolveReviewerSelection's no-explicit-selection fallback is "select every detected reviewer" (the correct default for /gsd:review), and passing it an unconditionally non-empty detected set would silently invoke the whole roster on every no-flag code review, violating COMP-01. * test(01-03): define external finding consolidation RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as untrusted input — independently re-verifies every claim against the actual current source, resists a prompt-injection attempt embedded in evidence text, and folds a verified claim into the existing Narrative Findings section with no second REVIEW.md schema (CONS-01..03). Also assert code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * feat(01-03): consolidate external review evidence GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence> as untrusted data, independently re-verifies every cited claim against the actual current source before it can appear in REVIEW.md, and explicitly resists prompt injection embedded in evidence text (never a command, no matter what it claims to be). A verified claim folds into the existing Narrative Findings section with (external: {slug}) provenance — one REVIEW.md schema only, no separate external-findings section. code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff. * fix(01-02): gitignore the reviewer-step-dispatch build artifact 01-02 added src/reviewer-step-dispatch.cts but never added its npm run build:lib output to .gitignore, unlike every sibling gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked noise in git status. * docs(01-04): publish user and command contract for reviewer-lane source review - Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on failure, findings independently consolidated into the single REVIEW.md - Add the same contract to the docs/features/code-review-pipeline.md fragment and regenerate docs/FEATURES.md from it - Preserve /gsd-review as the plan-review command; cross-reference it rather than duplicating the reviewer roster - Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md drift owned by source already shipped in Plans 01-01/01-03 but never regenerated (npm run regen:derived had not been run in this worktree) * docs(01-04): align architecture and agent ownership docs for reviewer-lane trait - ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes) through the shared dispatchReviewerLanes interpreter to the existing review-lane plan/invoke machinery, ending at gsd-code-reviewer as the sole REVIEW.md consolidator - AGENTS.md: document gsd-code-reviewer's full-context verification scope and its treatment of external reviewer evidence as unverified input - No new diagram, abstraction, or config key; docs/CONFIGURATION.md is unchanged since the feature adds no setting or default * fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact Same gap as the earlier .gitignore fix: 01-02 added src/reviewer-step-dispatch.cts but never added its generated gsd-core/bin/lib/reviewer-step-dispatch.cjs output to eslint.config.mjs's ignore list like every sibling generated file, so tsc's emitted __importDefault CommonJS-interop var tripped no-var. * fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md 01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster row in docs/INVENTORY.md — required by design, since a role sentence cannot be generated — was never added. * fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added supportsReviewerLanes: true to that step and this fixture was not updated. * chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md Both files grew as a direct, intended consequence of wiring optional reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes step and the untrusted-evidence consolidation contract) — not incidental drift. Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209) Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209) * test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes From internal code review: dispatched must be false when zero lanes actually reached plan(), and a throwing plan()/invoke() for one lane must not discard results already collected for a sibling lane — matching the fail-closed pattern gsd-tools.cjs already uses for the same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794). Refs: gsd-core-dks.16, gsd-core-dks.17 * fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review - WR-01: dispatched now tracks whether any lane actually reached plan(), not results.length — an unresolvable selected slug no longer reports dispatched:true. - WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a throw for one lane can never discard results already collected for a sibling lane, matching the same guard gsd-tools.cjs already has around the identical resolveLanePlan call. - IN-01: documents the intentional budget===0-is-unbounded convention (#2797) the caller already relies on. - IN-02: review-lane dispatch-step no longer blocks indefinitely on an un-piped interactive TTY; fails closed to empty paths instead. Refs: gsd-core-dks.16, gsd-core-dks.17 * docs(01-05): add changeset fragment for PR #17 * fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract agents/gsd-code-reviewer.md's untrusted-evidence section and its pinning regression test both quote injection phrases as the exact attack they defend against/detect — same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing allowlist entries, not an actual injection vector. * test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip From CodeRabbit review: WR-02's earlier fix only wrapped plan() — writePromptFile()/deps.invoke() still ran unguarded, so a throw there still aborted every later selected lane. Also covers the dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit lanes silently not running when no prior review and no phase-start commit exist). * fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved Previously an explicit reviewer-lane request with no prior review and no resolvable phase-start commit reached dispatch-step with an empty --base-sha, which fails closed via missing_provenance — correct, but silent about why explicitly requested lanes didn't run. Now skip dispatch entirely in that case with a stderr warning naming the actual cause. * fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan() WR-02's original fix only guarded plan() — a throw from writePromptFile() or deps.invoke() still aborted the whole dispatch, discarding results already collected for lanes processed earlier in the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it. * fix(01-05): WR-02b mock must throw only on the first writePromptFile() call The committed mock threw unconditionally, so codex's retry also threw and failed for the same reason as claude's — the test could not distinguish 'sibling still runs' from 'sibling also breaks'. Gate the throw to the first call, matching WR-02/WR-02c's single-failure intent. * fix(#4209): close review findings from adversarial + critical-code-reviewer pass Two independent reviews (agy adversarial review, Opus critical-code-reviewer + ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and fixed here: - dispatch-step's reducer silently swallowed whole-dispatch rejections (invalid paths, missing provenance, etc); it now checks parsed.ok/reason. - spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review; now shares the single compute_file_scope derivation. - the external reviewer prompt had no actual review request or citation requirement, only prohibitions; added both. - removed the supportsReviewerLanes trait plumbing (capability registry, validator, loop-resolver, docs, tests) — it was never consulted by the real dispatch path, which gates on explicit CLI flags instead. - flag-resolution require() was a fragile cwd-relative literal that failed silently on non-vendored installs; now resolves via GSD_TOOLS's own directory and warns instead of swallowing failure. - reducer didn't unwrap the @file: overflow protocol for large payloads. - deduplicated resolveBudget/budgetFor into one resolveLaneBudget. - lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a second dispatch can't overwrite prior evidence. - validatePaths rejects control characters, closing a markdown-injection vector into the external prompt via crafted filenames. - reworded the one line that tripped prompt-injection-scan.sh instead of allowlisting the whole production prompt file. - fixed a stale docstring range and a dispatched-field ordering bug. - added 3 integration tests executing the actual reducer against synthetic dispatch-step JSON, replacing markdown-substring-only assertions. 771/771 tests pass across every touched suite; tsc --noEmit clean. * fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait The maintainer's approval on issue #4209 explicitly redirected implementation shape: reviewer-lane dispatch must be a reusable capability/step-dispatch trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call itself. My previous commit (e2558326) deleted that trait entirely after finding it declared-but-never-consulted, which was backwards — the fix was to wire it, not remove it. Restores the trait (capability.json, generated registry, validator, loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes now resolves its own active hook via `gsd_run loop render-hooks` for the configured workflow.code_review_point and only proceeds to CLI-flag matching when supportsReviewerLanes reads true. Explicit flags no longer bypass the trait; a matching flag with the trait false resolves zero slugs (proven by a new integration test executing the real fence with both trait states). Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer redirect requires the capability layer, not the workflow, own the opt-in decision). * fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point Both an agy adversarial review and an Opus critical-code-reviewer pass independently found the same gap in my previous commit (9b2c3773d): the trait check I wired into code-review.md only protected code-review's OWN invocation — gsd-tools.cjs's dispatch-step handler still hardcoded `trait: true` unconditionally, so a second capability declaring supportsReviewerLanes would get zero enforcement from the shared CLI unless it correctly re-implemented the ~15-line render-hooks scrape itself. That is exactly the "each workflow.md hand-wiring the call" the maintainer's redirect said to eliminate. Moves the trait check into dispatch-step itself: given --cap-id/--point, it self-invokes `loop render-hooks <point>` (relocating the one subprocess code-review.md used to spawn for this, not adding a new one) and derives the real trait from that capId's active hook, rather than trusting a caller-passed boolean. code-review.md now only passes --cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or gates on the trait itself — the ~20-line scrape it previously carried is gone. Any other capability opts into the identical enforcement by declaring the trait and passing the same two flags. Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input variable (they proved a bash branch honors a variable, not that the variable reflects the real capability manifest) with three integration tests that invoke the real dispatch-step CLI against the real first-party capability registry: the real code-review trait resolves true, an unknown --cap-id resolves false (trait_not_enabled, fail-closed), and omitting --cap-id/--point entirely resolves false (no context means no opt-in). Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check (agy-F1 was incomplete), and delete the promptWritten per-lane coupling flag — the prompt write is idempotent, so writing it once per lane instead of gating on "did any lane write it yet" removes a latent bug where a deps.plan override that ever varies promptPath per lane would silently skip writing for a later lane. Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count drops (the trait scrape moved into dispatch-step), but the file still grew this session across multiple commits; acknowledging per the growth-tracking convention. * fix(#4209): remove per-run token waste from the shipped prompts Runtime prompt content, not session tokens: two real, per-invocation token costs in the code that ships. 1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of load_context step 5's ~180-word untrusted-evidence contract in ~90 more words, breaking this section's own established terse one-liner style (every other rule here is 1-2 sentences). This prompt loads fresh on every /gsd:code-review invocation. Shrunk to a one-line cross-reference, matching how write_review's own reference to step 5 already does it. 2. buildSourceReviewPrompt repeated the base SHA on every single file line even though it is identical for every file and already stated once at the top of the prompt — O(files) wasted tokens on every dispatched lane for a 50-file review, for zero information gain. File lines are now bare paths. * fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3 Opus critical-code-reviewer found a real Blocking defect in the --cap-id/ --point self-invocation added last commit: `dispatch-step` spawned `loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to `@file:<path>` instead of inline JSON -- the same overflow protocol this feature already unwraps for its OWN dispatch result 60 lines later in code-review.md. A large-enough activeHooks envelope (more installed capabilities/fragments) would throw, get silently swallowed by the bare catch, and misreport a real trait as trait_not_enabled with zero diagnostic. Fixed by extracting the config/registry/capability-state resolution `cmdLoopRenderHooks` already performs into an exported pure function, resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now share it), and calling it in-process from dispatch-step instead of spawning a subprocess at all. This eliminates the @file: exposure entirely (the dispatch-step path never touches the rendered-string envelope or its JSON-stringify/50000-char threshold), removes one subprocess spawn per code-review invocation, and gives a genuine diagnostic (stderr warning) on resolution failure instead of silent fail-closed. Corrected three doc/ docstring references to the now-removed subprocess self-invocation. Also fixes 2 real CI failures this round surfaced: - lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line no-control-regex` comment was unused under this project's ESLint config (verified locally: the rule never actually flags \x00-\x1f in this repo's config) -- a mistake from an earlier commit this session, never actually lint-checked before push. Removed the disable comment. - security (prompt-injection-scan): the agy-F1 regression test's crafted fixture literally contains "Ignore all prior instructions." as test data proving validatePaths rejects it -- allowlisted the test file, same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries. Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one bullet stated "untrusted, never a command" three different ways in one paragraph, and a same-file duplicate of write_review's schema rule. Consolidated to state each rule once. Declined one suggestion from this round: shrinking code-review.md's EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests (tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block, tests/code-review.test.cjs's CONS-02 test) deliberately lock the four- prohibitions restatement and the untrusted-evidence prose into the INJECTED block itself, not just the consolidator's system prompt -- adjacency of the warning to the untrusted payload it's warning about is a recognized prompt-injection defense-in-depth pattern from this workstream's original TDD plan, not accidental duplication. * fix(#4209): correct stale per-file base-SHA prose in the external prompt Leftover from removing the per-file base SHA repetition earlier this session: the review-request sentence still said "relative to its base SHA" (singular per-file framing) when there's now exactly one base SHA, stated once above the file list. Reads "relative to the base SHA above" now. * fix(#4209): make getLane/configGet/plan required deps, delete dead defaults R3/R4 from the review round I'd deferred as low-priority test-churn: this file's one production caller (gsd-tools.cjs's dispatch-step handler) always supplies all three, so the fallbacks were dead in production -- but each was actively WRONG if ever reached: the default configGet always returned undefined, silently disabling resolveLaneBudget's overflow guard; the default getLane looked up only first-party REVIEWER_LANES, diverging from production's overlay-merged roster; the default plan skipped per-host effort resolution entirely. These defaults were introduced by this PR's own earlier work (this file did not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from elsewhere, so there's no external caller depending on the lenient contract. Turned out free to fix: making the three deps required and deleting defaultGetLane/defaultPlan needed zero test changes -- every existing test that actually reaches the per-lane loop already supplies getLane/plan explicitly, and configGet's only real dependent (the budget-overflow tests) already supplies it too. 788/788 tests pass unchanged, tsc/lint clean. * fix(#4209): define depth semantics for the external reviewer lane Verified this was a real bug, not a match to existing convention as I'd claimed when declining the suggestion earlier this session: the internal gsd-code-reviewer agent's own system prompt carries a full <depth_levels> block defining what quick/standard/deep mean and do (agents/gsd-code- reviewer.md:68-99). The external reviewer lane has no access to that persona at all -- it only ever sees buildSourceReviewPrompt's bounded text, which sent the bare depth label with zero definition to a third-party CLI with no other source of truth for what "standard" means. Added depthMeaning(), condensed from the internal reviewer's own <depth_levels> definitions so the two stay consistent, and interpolated it into the review-request sentence. 150/150 tests pass, tsc/lint clean. * fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide whether to dispatch at all. This file's own documented rule (its depth-resolution guard, stated explicitly a few hundred lines earlier) is that a guard and the extraction it protects must run as one shell control-flow decision, because markdown-fenced blocks do not share shell state -- this step violated its own file's rule for the entire feature's gating condition. Merged the roster-resolution fence and the dispatch-decision fence into one continuous bash block, removing the intervening prose that split them. Fixed the stderr-based failure detection in the same edit (RQ-01: checking whether stderr is non-empty misfires on any benign Node warning; now checks the actual exit status of the roster-resolution command). Verified by extracting the merged fence and executing it standalone, driving both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty, SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests pass, tsc/lint clean. * fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write Batch of Required/Suggestion fixes from the Opus critical-code-reviewer + writing-for-agents pass: - CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch blocks, commented-out code) and deep (error propagation, state mutation consistency, circular dependencies) relative to the real <depth_levels> block, and had zero test coverage. Restored full accuracy and added tests that read the real agents/gsd-code-reviewer.md file directly, so drift between the two can't recur silently. Unrecognised depth now normalizes to standard's definition, matching that agent's own documented rule, instead of rendering an undefined bare label. - RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt `paths` does, but weren't checked for control characters like paths were (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and applied it to all four fields at the same provenance-check boundary. runDir previously had zero validation at all. - S1: deleted the dead `identity` parameter on `invoke` -- the one production caller already ignores it, no test read it by name. - S2: hoisted the shared prompt write above the per-lane loop -- promptPath is derived from runDir alone (constant across lanes by construction), so writing it once is both correct and cheaper than the per-lane write R1 introduced earlier this session. Discovered and fixed a real regression from the naive version of this hoist: an unguarded throw would have escaped dispatchReviewerLanes as an uncaught exception instead of a clean per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason, matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with a dedicated regression test. - S3: moved `planned = true` past the budget-overflow gate, so `dispatched` only reports true once a lane has cleared BOTH plan and budget checks. - S5: relayed gsd-code-reviewer.md's own "performance issues are out of scope unless also correctness issues" policy into the external-lane prompt, which previously had no such guidance and could return findings the internal reviewer's own contract excludes. - RQ-05 (partial): shrunk this file's own header docstring's restatement of the trait-reuse architecture to a pointer at gsd-core/references/loop-hook-dispatch.md, the canonical home. 234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke` already share. code-review.md's ~18-line inline `node -e` reimplementing `loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in gsd-tools.cjs) is now a single call to this subcommand -- the exact violation code-review-flags.cjs's own header warns against ("this is the canonical flag-parsing surface -- do not replicate inline bash parsing"). RQ-03: an empty --cap-id XOR --point now warns distinctly from the legitimate no-context opt-out (both absent) -- a caller that named a capability without its point was silently indistinguishable from a correct opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only ever fires when the config-get COMMAND ITSELF fails (config-get already resolves the manifest's own schema default in the normal case), but that failure was previously silent. RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait resolved inside dispatch-step" explanation was restated in full in 5 places across this session's own review cycles. Consolidated to ONE canonical statement in gsd-core/references/loop-hook-dispatch.md; the other 4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md, code-review.md's step-opening comment) now point at it instead. W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two inert cases when capability-validator.cjs already rejects non-boolean at load -- restated as the two cases that actually reach this code. Removed a "do not hand-roll trait resolution" prohibition whose target no longer exists once the positive description precedes it. W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing block means proceed as normal") -- an absent optional block already means proceed as normal without being told. W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the made-up compound "byte-for-behavior [un]changed" with the token this session's own docs already coined for this concept (inert) and the word that means what byte-for-behavior was reaching for (unchanged). W-10: dispatch_reviewer_lanes had no completion criterion -- added one sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set, either populated or empty). This exact sentence would have caught the cross-fence bug fixed two commits ago at authoring time. Declined from this round, with reasoning: W-02/W-03 (trim the untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) -- two tests deliberately lock this as intentional adjacency-based prompt-injection defense-in-depth, not accidental duplication (see this branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site `trap ... EXIT`) -- would fire at the end of the CREATING fence, before spawn_reviewer's agent ever reads the evidence files, given this file's own documented fenced-block execution model; the existing named cross-reference between creation and cleanup already satisfies the co-location concern without introducing that regression. 853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean. * fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get fallback lived in an earlier, separate fence from the fence that consumes it via --point, split only by prose (not a guard, per this step's own documented rule). Merged into the single continuous fence and added a structural test asserting exactly one bash fence in the step. The new end-to-end regression test for this used --codex, which drives the fence's real `review-lane dispatch-step` call and, with the codex binary present on PATH, spawns the real external CLI — which then blocks on interactive auth with no stdin (BL-01). Stubbed gsd_run for `review-lane dispatch-step` only (captures argv instead of executing), keeping the real config-get/explicit-from-argv calls the test is actually about. * fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment Round-5 review (Opus) warning-tier findings: - WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but a control-character injection attempt" — a caller distinguishing a config problem from a security event couldn't tell them apart. Split into MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid). - WR-05: validatePaths' containment check was lexical only (path.resolve), so a symlink whose own path sits inside repoRoot could still point outside it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can legitimately name a file already deleted in a stale worktree), realpathing repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't false-positive-reject its own real children. - WR-08: a comment in the per-lane loop still said a throwing writePromptFile() was caught there — stale since the prompt write was hoisted above the loop in an earlier round. WR-03 (validate depth against the quick/standard/deep enum) was considered and declined: this dispatcher is deliberately capability-neutral (see the existing "synthetic step context" test, which passes a non-code-review depth label on purpose to prove no code-review-specific special-casing exists). WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and WR-07 (reason omitted on the aggregate return) were verified against source and are not bugs — see review notes. * docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap Round-5 review (Opus, BL-03) flagged that an early exit between dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A trap-based cleanup was considered and rejected: if a step genuinely runs as a separate process, a trap set at creation time would fire at the end of that SAME fence, deleting the directory before spawn_reviewer/commit_review ever read it — worse than the leak it would fix. review.md's own gather_context/cleanup pair for the identical resource class (a run-scoped reviewer temp dir) already makes and documents this exact trade-off: cleanup runs only on a documented success path, and a leftover $TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording that precedent here so this isn't re-raised as a live gap in a future review. * fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes paths: ['docs/spec.md'] as a synthetic, never-read path proving the dispatcher has no code-review-specific special-casing. lint-docs-guard- registration correctly flagged this as an unregistered docs/ path reference — add the docs-guard-exempt marker and its pinned baseline entry, the same pattern every other synthetic docs/ literal in this test suite already uses. * fix(#4209): backfill changeset pr: field with the real upstream PR number changeset-lint's fail_pr_field_drift caught the fragment still pointing at the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this branch is now also open against. * docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every decision in ADR-2782 (D1-D9) and every prior dated amendment governs the `role: "reviewer"` capability body and its one consumer, /gsd:review. This PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary feature capability's `steps[]` entry, projected through loop-resolver.cts and resolved in-process via resolveActiveHooksForPoint - is a different capability axis (steps/gates/contributions) that the ADR's own scope note explicitly places out of reach. Per docs/contributor-standards.md's "Amending an accepted ADR", an in-place dated section is the established, lighter-weight path for an addition that stays within the ADR's existing decisions - used twice already in this same file - so this appends a third dated entry documenting the new seam, its consumer, and why it reuses the existing D1-D9-governed plan/invoke machinery rather than adding a second one. No decision is reversed; no new Amends/Amended-by pair is needed since the steps/gates/contributions axis already carries reciprocal links to ADR-857 and ADR-894. * fix(#4209): close two test-quality gaps trek-e's review found Minor 1: validatePaths (a path-shape parser guarding the prompt- injection/path-traversal trust boundary) had only example-based coverage, violating ADR-456's rule that parsers/budget limits carry at least one fast-check property test. Adds three: safe-segment paths are never rejected, a single leading "../" always escapes the one-segment repoRoot, and a control character anywhere is always rejected - one property per rejection reason validatePaths owns. Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only ever exercised far below budget or at budget:0 (unbounded), never at the exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds three exact-boundary tests using the real estimateTokens/ buildSourceReviewPrompt the module calls internally, so the resolved token count is exact rather than approximated: budget == estimate (must pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must pass). Also extracts okPlan()'s fixture timeoutMs into a named constant - local/no-adhoc-timeout-literal (#4446) landed on next after this branch was authored and flagged the pre-existing literal on rebase; it is fixture data for a synthetic plan object dispatchReviewerLanes never waits on, a distinct class from tests/helpers/timeouts.cjs's real subprocess norms. * fix(#4209): update docs-guard-registration baseline for the new ADR citation reviewer-step-dispatch.test.cjs's new fast-check property tests cite docs/adr/456-test-rigor-architecture.md in a justifying comment (never a real read). lint-docs-guard-registration fingerprints every docs/ path string an exempted test file mentions and fails on drift so a human re-confirms the exemption still holds - re-confirmed, and the baseline is updated to match. * fix(#4209): point changeset pr: field at the fork PR for CI validation changeset-lint's fail_pr_field_drift check compares the fragment's pr: field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH), not a fixed target. Rehearsing this branch on fork PR davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but fails here. Backfill to 4323 happens again, as the last commit, immediately before the approved push to open-gsd#4323 - never leaving pr: 17 on the branch that ships upstream. * fix(#4209): reject promptChannel:none lanes from source-review dispatch CodeRabbit found a real scope mismatch: coderabbit's lane declares promptChannel: 'none' and reviews the working tree on its own terms, fed nothing (review.md:367). Silently dispatching it through dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope buildSourceReviewPrompt promises and let the lane review whatever it independently sees fit, violating this interpreter's own scoped, metadata-only contract. Reject before plan()/invoke(), same as an unresolved slug. * fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file CodeRabbit found the whole-file match on workflowContent would still pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts of this 1000+-line workflow, proving nothing about the actual evidence block's contract. Line-filtered via splitLines (not a bare-\n regex spanning readFileSync content) so this stays CRLF-portable and passes local/no-unbounded-quantifier and local/no-crlf-fragile-split. * fix(#4209): guard DISPATCH_JSON substitution and capture its stderr CodeRabbit found the dispatch-step command substitution unguarded: a non-zero exit could leave DISPATCH_JSON empty (or halt the step under errexit with no warning), and the downstream reducer would only ever report the generic unparseable_dispatch_output reason, discarding the command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/ EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface it in a warning on failure, and fall back to a parseable dispatch_ command_failed JSON stub so the reducer's existing reason-reporting path still fires. * docs(#4209): fix byte-for-behavior wording and missing colon, regenerate CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the established repo term for output-identical unchanged behavior) and a missing colon after the bold "Optional external reviewer lanes (#4209)" lead-in in docs/features/code-review-pipeline.md. Fixed in the two hand-authored sources (commands/gsd/code-review.md, docs/features/ code-review-pipeline.md) and regenerated the two derived projections (skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/ FEATURES.md via gen-features.cjs) so they stay in sync. * fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI) The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds double-quoted JSON keys inside a single-quoted shell literal. That extra quote density, inside an already quote-heavy ~8KB driver string, passed bash -n and the full local suite on Linux but broke Windows Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end to end (#4209 round 5)` failed on two Windows CI shards with `bash -c: unexpected EOF while looking for matching '''` — a Windows argv-to- command-line re-quoting edge case, reproducible on rerun, not a flake. Root-caused via gh api job logs plus a byte-identical local reconstruction of the test's own driver script. Fix: drop the fabricated stub. The downstream node -e reducer already falls back to reason `unparseable_dispatch_output` on any JSON.parse failure, so an empty/partial DISPATCH_JSON on command failure is still handled correctly, with zero new quoting risk. * revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI) Two materially different mechanisms for the same CodeRabbit Nitpick ("Trivial | Quick win") both broke Windows Git-Bash reproducibly: a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ... matching '''") and, after removing that, a plain `head -1 "$VAR"` inside a nested command substitution ("unexpected EOF ... matching '"'"). Both passed bash -n and the full local suite on Linux every time; both failed the SAME test deterministically on Windows CI. Two attempts at the same class of fix (nested-quote construction near this exact step) is the retry limit - reverting to the original, already-shipped, Windows-verified unguarded form rather than continuing to guess at a third quoting mechanism for a Trivial- severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for anyone attempting this again: the fix belongs outside this specific markdown-fence-driver test harness (e.g., a real .sh helper script) if it's worth doing at all. * fix(#4209): backfill changeset pr: field to the real upstream PR before push Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy changeset-lint's PR-number check while rehearsing there; this is the last commit before the approved push to the real upstream PR (open-gsd/gsd-core#4323), so the field points at that PR number again. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
66 KiB
ADR-2782: Reviewer Lane — the cross-AI reviewer handoff becomes a declared capability surface
- Status: Accepted
- Date: 2026-07-28
- Amended: 2026-07-29 by Phase 1 (#2794) — D1's
flagbecomesflags[]; D2'spromptChannelgainsnone,outputChannelgainsfile-argwith a companionoutputArg; D8's uniqueness invariant restated over the flattened flag set. All four are additive widenings of closed enums, each forced by a shipped lane the original survey did not cover. See Amendments at the end. - Issue: #2782 (epic); Phase 0 tracked by #2793
- Amends: ADR-857 (extension points as data — extends D7/D8 in the same "amend, not reverse" sense ADR-1244 D8 established) · ADR-894 (adds a role-typed body and a third role) · ADR-1016 (the runtime body is no longer the only body a
role: "runtime"capability may carry; its closed-vocabulary principle is upheld, not relaxed — see D6) · ADR-1244 (D5 gains a fourth executable-surface disclosure class; D9's matrix gains a lane column) - Unchanged and explicitly out of scope: ADR-0011 (reviewer selection precedence) · ADR-1517 (the
REVIEWS.mdcontract and reviewer instances) - Subsumes: #2690 (core single-sourcing — lands as Phase 1 under this ADR rather than as its own design)
Context
A cross-AI reviewer lane — one external CLI or model endpoint that /gsd:review hands a plan to
for independent review — is declared today in three unrelated places, none of which is the
capability system, and none of which a third party can extend.
1. The roster is half registry-derived, half hardcoded. src/review-reviewer-selection.cts
derives slugs from runtime.hostBehaviors.reviewerCli === true (:40-49), then concatenates a
hardcoded NON_RUNTIME_REVIEWER_SLUGS tail (:32-38) for five reviewers that have no
capabilities/<id>/ directory at all. The module's own comment says exactly this. Six capabilities
carry the flag; five reviewers have no descriptor of any kind.
2. The invocation contract is prose. gsd-core/workflows/review.md is 1070 lines;
invoke_reviewers spans roughly 60% of it as hand-authored per-CLI bash. Each leg re-implements
probe, argv shape, model lookup, effort channel, timeout, stderr capture, and empty-output policy.
3. The output contract is prose. write_reviews hardcodes per-reviewer section headings,
including two literal instance names.
Three structural consequences follow, and they are why this is a capability question rather than only a refactor question.
(a) reviewerCli is a bare boolean in an undocumented, unvalidated bag. hostBehaviors
appears zero times in docs/reference/capability-manifest.md — not in the envelope table, not
in the runtime-body axis table — and scripts/gen-capability-registry.cjs does not validate its
keys. The one field that decides reviewer membership is unspecified, unvalidated, and carries no
invocation data. A capability author can discover it only by reading
src/review-reviewer-selection.cts:47.
(b) Reviewer-ness is welded to runtime-ness, and the runtime body structurally cannot hold a lane
contract. capability-manifest.md:141 states the runtime body is "a closed 8-axis (plus 4
install-surface) vocabulary; no feature-only fields (skills, agents, steps, contributions,
gates, hooks) are permitted," and gen-capability-registry.cjs:505 enforces the consequence — a
role: "runtime" capability is stored whole into runtimes[] and its config/steps/
contributions/gates are never harvested. A reviewer lane therefore cannot own its own federated
config keys. That is why review.models.*, review.ollama_host, review.lm_studio_host,
review.llama_cpp_host, and review.max_prompt_tokens_per_reviewer.* all live in the central
schema instead of with the lane that uses them — the exact half-migrated shape the config-key
exclusivity invariant exists to prevent.
(c) A reviewer that is not a GSD install target has nowhere to live. gemini, coderabbit,
ollama, lm_studio, and llama_cpp are review or model CLIs GSD never installs into. There is no
capabilities/<id>/ for them, so they are a hardcoded tail by necessity, not by choice.
Net: adding a reviewer lane is a core patch. It means editing the roster module or a runtime
descriptor, hand-authoring a bash leg, hand-adding a write_reviews heading, adding central config
keys, and updating five prose surfaces. #2718 was that patch in flight (PR #2776, closed in favor
of this design); #2781 is the documentation drift it produced. Cross-cutting fixes land per-leg:
#2494 and #2605 were the same empty-output defect filed twice; #2475 (effort channel), #2589 (model
lookup), #2295 (resolved-model recording), and #2272 (flag parity) are the same shape.
What a survey of the twelve lanes actually shows
The design was drafted assuming one lane shape. Reading all twelve legs disproved that, and the correction is the most consequential decision in this ADR (D2).
| Family | Lanes | Shape |
|---|---|---|
| Spawned CLI | gemini, claude, codex, coderabbit, opencode, qwen, cursor, antigravity, kimi-code |
Binary + argv; prompt via stdin or argv; stdout captured, stderr to a .err sidecar |
| OpenAI-compatible HTTP | ollama, lm_studio, llama_cpp |
No binary. curl to /v1/chat/completions on a user-configured host; model discovered via GET /v1/models piped through jq |
Three of twelve lanes are not spawned binaries at all. Timeout floors genuinely diverge — a measured
~570 s for Codex at xhigh effort and ~525 s for headless Claude drive a 900 000 ms floor with
1 200 000 ms for those two, while the Antigravity leg runs a 600 s external cap over a 540 s native
--print-timeout, and the HTTP lanes use 120 s. Five lanes require jq on PATH. The Antigravity
leg carries a deliberate three-layer fallback for an upstream stdout bug.
Divergence between lanes is real and frequently correct. The value of a descriptor is therefore one place where divergence is declared, not one behavior imposed on every lane.
Decisions
D1 — A reviewer body on the capability manifest, admissible on two roles
A reviewer lane is declared as data in a reviewer body:
{
"id": "acme-reviewer",
"role": "reviewer",
"version": "1.0.0",
"title": "Acme Review CLI",
"description": "Cross-AI plan review lane backed by the Acme CLI.",
"tier": "full",
"requires": [],
"engines": { "gsd": ">=1.9.0" },
"reviewer": {
"slug": "acme",
"flags": ["--acme"],
"transport": "spawn",
"probe": { "kind": "command-exists", "binary": "acme" },
"invoke": {
"binary": "acme",
"args": ["review", "--format", "text"],
"promptChannel": "stdin",
"outputChannel": "stdout",
"modelArg": "--model",
"effortChannel": "argv"
},
"timeoutFloorMs": 900000,
"emptyOutput": "stub-with-stderr",
"reviewsSection": "Acme Review",
"evidenceClass": "source-grounded",
"requiresBinaries": [],
"promptBudgetKey": null,
"handler": null
},
"config": {
"review.models.acme": {
"type": "string",
"default": "",
"description": "Model passed to the Acme reviewer lane."
}
}
}
The body is admissible on role: "runtime" — so the six capabilities that are both install
targets and reviewers (claude, codex, cursor, opencode, qwen, antigravity) keep exactly
one manifest — and on a new role: "reviewer" (D3) for lanes that are not install targets.
This is the amendment to ADR-1016: a role: "runtime" capability may now carry a reviewer body
alongside its runtime body. The runtime body itself remains closed and unchanged; no feature-only
field becomes permissible on it. A lane body is a third thing, not a relaxation of the second.
Because a lane may own a federated config slice, gen-capability-registry.cjs must harvest
config from a lane-bearing capability of either role — the specific limitation at :505 that
context (b) describes.
D2 — transport is a closed discriminator, and it selects the invoke sub-shape
reviewer.transport is a closed enum: spawn | openai-http.
spawn |
openai-http |
|
|---|---|---|
invoke.binary |
required | forbidden |
invoke.args |
required (array) | forbidden |
invoke.promptChannel |
stdin | argv | argv-file-ref | none |
forbidden |
invoke.outputChannel |
stdout | file-arg |
forbidden |
invoke.outputArg |
required iff outputChannel: "file-arg", else forbidden |
forbidden |
invoke.hostConfigKey |
forbidden | required (dotted config key holding the base URL) |
invoke.path |
forbidden | required (e.g. /v1/chat/completions) |
invoke.modelDiscovery |
forbidden | closed enum: none | first-from-models-endpoint |
invoke.modelArg |
optional | forbidden (model travels in the JSON body) |
invoke.effortChannel |
closed enum: none | argv | env |
none |
invoke.env |
optional: object of environment name/value pairs, string values only | forbidden (no child process to carry an environment) |
A manifest declaring fields from both sub-shapes, or neither, fails validation. The discriminator is explicit rather than inferred from field presence: inference leaves a manifest with both — or with neither — carrying undefined meaning, which is precisely what a closed vocabulary exists to prevent.
promptChannel: "argv-file-ref" exists because two lanes (cursor, kimi-code) take the prompt as
an argv argument, and passing a full plan set inline would approach the 32 767-character Windows
execFileSync ceiling. The file-reference form passes a short instruction naming a prompt file in
the run directory. That instruction must also carry the absolute repository root, because an
argv-fed CLI does not reliably inherit the review's working directory — the existing cursor and
kimi-code legs already do this by hand (review.md:447-448, :550-552).
outputChannel is a required, named, closed-enum field rather than an implicit assumption, because
the alternative — a lane that writes its review to a file and prints nothing — is a shape a real CLI
can take, and an unnamed assumption is the thing a later contributor silently violates.
Amended 2026-07-29 (#2794): this ADR originally recorded outputChannel as having "exactly one
member (stdout) today" and described the file-writing lane as a shape a real CLI could take. It
already does. codex captures its review through its own -o/--output-last-message <FILE> and
discards stdout, because on Windows it writes process-teardown output to stdout after the final
message, and a stdout redirect would append that noise to a non-empty file — slipping past the
empty-output guard as a silently polluted review (#1698). The enum therefore ships with two members,
and file-arg carries a companion outputArg naming the argument that takes the path: knowing the
review lands in a file is useless without it.
promptChannel likewise gains none. coderabbit is fed no prompt at all — it reviews the
working-tree diff and accepts neither a prompt nor a model flag (review.md:367). The original
three-member enum had no way to say "this lane receives nothing", which would have forced Phase 2 to
either invent a sentinel or mis-declare the lane.
Both were found by building Phase 1's descriptor table against all eleven shipped legs. That is the
same evidence path that produced openai-http in the first place, and it is the process working:
the vocabulary widens on a lane that exists, never on speculation.
Three further declared fields carry per-lane divergence that would otherwise live only in prose:
| Field | Values | Why it exists |
|---|---|---|
evidenceClass |
source-grounded | diff-only |
CodeRabbit reviews a diff, not the source tree, and its findings are deliberately down-weighted in synthesis (review.md:367). Today that caveat is a prose annotation a reader may miss; declaring it lets write_reviews render the caveat from data |
requiresBinaries |
string[] | External tools the lane needs on PATH — jq for five lanes. A missing prerequisite reports the lane unavailable with an install hint rather than running it into an empty review |
promptBudgetKey |
dotted config key | null |
Per-lane prompt trimming (prepare_trimmed_prompt_for_reviewer, review.md:646-704) is keyed per slug today; the key becomes the lane's own federated config (D9) |
Naming note: the field is requiresBinaries, not requires. The envelope already carries a
requires (capability-id dependencies, ADR-1244). Two fields named requires at different nesting
depths with unrelated semantics is a defect waiting to happen; the collision was caught in review of
this ADR and renamed here rather than left for a downstream phase to trip over.
D3 — A third role, role: "reviewer", for lanes that are not install targets
gemini, coderabbit, ollama, lm_studio, and llama_cpp become first-party capabilities with
a reviewer body, no runtime body, and no install surface — which is the honest description of
what they are. runtimeCompat is not required for this role (it declares which host runtimes a
feature surfaces through; a lane surfaces through none).
tier remains required, because it is the source of truth for install-profile membership. A
role: "reviewer" capability therefore receives profile membership from deriveProfileMembership
(gen-capability-registry.cjs:201-213) like any other. That membership is inert: the capability
contributes no artifacts, so there is nothing to install. This is stated explicitly because a reader
encountering a lane in an install profile would otherwise reasonably assume it installs something.
Rejected: one role for every lane, splitting codex into codex + codex-reviewer. It is the
cleaner discriminator and was rejected for churn — six manifests would each fragment into two
capabilities and two ids, complicating roster derivation for no gain.
D4 — The reviewer body is optional and absent-safe at every layer
This is a normative MUST, and it governs every downstream phase.
- A capability with no
reviewerbody is simply not a lane. This is never a validation error. Most runtime capabilities are install targets only; a validator that errors on an absent body would break the majority of the registry. - An overlay declaring a
roleor a field this GSD version does not know is skipped with a warning via the existingengines.gsdhard gate (ADR-1244 D6) — never a crash. This is the forward half: a capability built for a newer GSD degrades to discovered-but-inactive. - An unknown field inside a
reviewerbody is ignored with a warning rather than failing validation. - A lane naming an unknown
handlerfails closed — the lane is unavailable; the registry does not crash. - A capability with no
reviewerbody must not perturb its disclosure signature (D5). An absent body that changed the signature would force spurious re-consent across every installed capability.
The asymmetry is deliberate and is Postel's Law applied with a boundary: liberal in what a manifest may omit, strict in what it asserts. Permissiveness about absence is forward compatibility; permissiveness about assertions would be an untyped escape hatch.
Absent-safe governs discovery, never explicit selection
Rules 1–5 describe what happens when the system is looking for lanes. They do not apply once a
user has named one. If a user runs /gsd:review --acme and the acme lane is unavailable — because
its capability was skipped under rule 2, because its handler failed closed under rule 4, because a
prerequisite binary is missing, or because its egress destination changed (D5) — that is an
error, surfaced and non-silent. It is not an informational note, and the run does not quietly
proceed with a thinner reviewer set.
This is called out because the current implementation does the opposite: an unavailable
explicitly-requested reviewer is recorded as an info (review-reviewer-selection.cts:246-248).
The workflow's own guidance already names why that is wrong — "a cross-AI review that silently drops
a lane is blind in one eye" (review.md:304) — and a design whose whole premise is more lanes from
less trusted sources must not inherit a silently-degrading selector. Correcting this is Phase 1's
responsibility, because Phase 1 is where the selector is single-sourced.
In one line: not finding a lane nobody asked for is normal; failing to run a lane somebody asked for is an error.
Where warnings surface
"Skipped with a warning" means nothing unless a human sees it. Warnings arising at build time
(registry generation over first-party capabilities) are emitted by gen-capability-registry.cjs on
stderr, and fail the build only where D8's uniqueness invariants are breached. Warnings arising at
load time — an overlay skipped by engines.gsd, an unknown field, a handler that failed
closed — surface on the /gsd:review run that would have used the lane, and in
gsd capability list, which is where a user goes to ask why a capability is inactive. A warning
written only to a build log nobody reads is not a warning.
D5 — A fourth executable-surface disclosure class: the reviewer lane
ADR-1244 D5 rule 2 requires that executable surfaces be disclosed and consented at install, and
names three classes: hooks, command modules, and mcpServers. A reviewer lane is a fourth, and it
is materially different from the other three: it receives data. A lane is piped the plan text,
the requirements, the research findings, and the CONTEXT.md decisions, and its output is read back
into REVIEWS.md. That is an egress channel for the most sensitive artifacts GSD produces.
Making lanes pluggable without a disclosure class would open a data-exfiltration path behind a manifest field. The trust work is therefore the gating requirement of this design, not polish.
discloseExecutableSurfaces gains a reviewer-lane surface that discloses, by transport:
spawn— the binary and its full declaredargs, in both rendered and raw form, exactly as MCP servers already discloseargv/rawArgs(capability-trust.cts:688-690).openai-http— the destination host URL resolved fromhostConfigKey, plus thehostConfigKeyitself. Disclosingcurlwould be technically true and practically meaningless; the destination is the disclosure that matters. Alocalhostdestination is still disclosed, and is distinguished from a remote one.
Both forms additionally disclose the egress payload classes — plan text, requirements, research
findings, CONTEXT.md decisions — rather than an unhelpful "sends data to the tool".
Disclosing the binary without its args is insufficient, and this is not hypothetical. A lane
declaring binary: "python3" with innocuous args could, in a later version, change args to
["-c", "<arbitrary program>"] without the binary changing at all. That is precisely the bug class
#1459 already fixed for MCP servers, and a binary-only disclosure would reopen it. args is
therefore disclosed and signature-bound.
The lane folds into disclosureSignature / signatureForManifest as stable sorted JSON, exactly as
env/cwd do for MCP servers (#1459). executableSetChanged treats any of the following as an
executable-set change for the auto-update re-consent trigger (ADR-1244 D5 rule 4): adding or removing
a lane, changing its slug, transport, binary, args, hostConfigKey, promptChannel or
handler, or changing any other field of its declared invoke object — the residual added by
#2483, which is what stops this list going stale again. See the 2026-08-05 amendment: an enumeration
of "the fields that matter" had already fallen eight fields behind by the time env arrived, so the
signature no longer relies on one.
The egress destination is re-verified at invocation, not only at install
hostConfigKey is the one consent-bound value that does not live in the SHA-pinned bundle. It
names a key in .planning/config.json, which is user- and CI-editable at any time with no
re-install and no integrity check — unlike every existing consent-bound field (command, args,
env, cwd, url), all of which come from the manifest itself (capability-trust.cts:74-125).
Left unaddressed, this is a real hole: a lane consented against http://localhost:8080 could be
silently redirected to a remote host by a later config edit — including one arriving through an
ordinary pull request touching .planning/config.json — and every subsequent review run would
egress plans, requirements, research, and decisions to the new destination with no re-prompt.
Therefore, normatively:
- The consent record binds the resolved host, not merely the config key.
- Before invoking an
openai-httplane the runtime re-resolveshostConfigKeyand compares the result against the consented host. - On mismatch the lane is blocked, not silently redirected; the user is told the destination changed and must re-consent. A blocked lane reports like any other unavailable lane — it never degrades to running against the new host.
- This check lives on the invocation path (Phase 5b), not only in the install path.
A host change is a change of who receives the user's plans. It is the most security-relevant mutation in this design, and it must not be reachable by editing a JSON file.
Implementation note added by Phase 3 (#2796) — how rule 1 is actually satisfied.
The resolved host is deliberately excluded from the disclosure signature, and a reader comparing rule 1 to
capability-trust.ctsmust not mistake that for the rule being unimplemented.
signatureForManifest(manifest, stagedDir?)is the single consent key that both the loader and the lifecycle compute, explicitly so the two "can never drift". The loader has no config resolver —hostConfigKeynames a key in.planning/config.json, which is outside the SHA-pinned bundle. Folding the resolved host into the signature would therefore make the loader and the lifecycle compute different signatures for the same manifest, producing a permanent false-mismatch loop that re-prompts forever.So the binding is split, and rule 1 still holds end to end:
- the signature binds the manifest-derived lane fields —
slug,transport,binary,args,hostConfigKey,promptChannel,handler, plus every other declaredinvokefield via the #2483 residual (2026-08-05: the enumeration alone was eight fields short, includingdefaultHost, which is itself an egress destination);- the consent record additionally stores the resolved host, which is what rule 1 requires;
- Phase 5b re-resolves and compares at invocation and blocks on mismatch, which is where rule 4 already places the check.
reviewsSectionandtimeoutFloorMsare excluded for a different reason: a cosmetic change must not force re-consent, because a prompt carrying no security information is how users learn to click through — the same failure this decision cites when rejecting a per-run egress prompt.(Recorded here rather than in the PR that made the decision. A squash-merged PR body is not a durable record: it is invisible to anyone reading the ADR later, which is exactly who needs this.)
Stated honestly, and consistent with ADR-1244 D5's own acknowledgment that there is no sandbox: even with the above, consent-at-install remains a weaker gate for a standing egress channel than for a hook. A user consents once; the lane thereafter receives every plan on every review run. Disclosure plus destination re-verification makes the channel visible, pinned, and revocable — it does not make it safe. A per-run egress prompt was considered and rejected as consent fatigue that trains users to approve blindly.
D6 — handler is a closed enum of first-party names; third-party lanes are data-only
Lane divergence is real (context above), so the descriptor must not promise uniformity. Where a lane
needs genuinely imperative behavior — the Antigravity three-layer fallback is the canary —
reviewer.handler names an imperative module by closed first-party name, rather than growing
conditionals inside data.
The enum ships with exactly these members. null is the default and covers eight of the twelve
lanes:
handler |
Lanes | What it owns that data cannot express |
|---|---|---|
null |
the other eight | Nothing — the declared vocabulary suffices |
"antigravity" |
antigravity |
The three-layer fallback for an upstream stdout bug; the two-level timeout (a 600 s external timeout/gtimeout cap wrapping a 540 s native --print-timeout, review.md:560); and the stale-response watermark guard that rejects a cached conversation from a prior run |
"openai-compatible" |
ollama, lm_studio, llama_cpp |
Model discovery against /v1/models, the JSON request/response shape, and the served-model mismatch warning raised when the responding model differs from the one requested (review.md:794-797) |
Enumerating the members here is deliberate. A "closed enum" whose membership is left to the implementing phase is not closed, and three separate phases would each have invented a different list.
On timeoutFloorMs and the Antigravity two-level timeout. The descriptor carries one scalar,
timeoutFloorMs — the outer wall-clock bound every lane gets. A lane whose tool has its own
internal timeout (Antigravity's --print-timeout is the only current case) expresses that inner
bound in its handler, not in the descriptor. Adding a second declarative timeout field to serve
one first-party lane would be speculative generality; the handler seam exists for exactly this. The
delegation is stated here so a reader does not wonder where the measured 540 s went.
This upholds rather than relaxes ADR-1016. That ADR's core principle is that "a runtime that needs a
shape no existing primitive expresses is supported by adding a first-party primitive … never by
embedding arbitrary code or an open escape hatch in the descriptor," and its §Alternatives #2
explicitly rejected an open escape hatch. handler is the same construction as ADR-1016
Decision 3's closed ConverterName: the descriptor references a first-party function by name and
never embeds it.
The consequence must be stated plainly, because it caps this epic's headline claim. Third-party lanes are data-only. A third-party CLI needing a shape the closed vocabulary lacks is blocked on a first-party PR. The honest claim is most lanes, declaratively — not any lane.
The escalation path is the ADR-1016 model, and it is documented rather than implied: file an issue
naming the primitive the vocabulary lacks; it is reviewed and added first-party. D2 is the worked
example of that path already functioning — the openai-http transport exists precisely because a
survey produced evidence that three real lanes did not fit, and the vocabulary widened on evidence
rather than on speculation.
Two real CLIs that would NOT fit today, named so the boundary is a known quantity rather than a surprise for the first third-party author who hits it:
- Aider mutates the repository by default — it edits files and commits. The vocabulary has no way
to declare "this tool must be invoked read-only", and the existing lanes achieve that only by
asking politely inside the prompt text (
review.md:448: "Do not edit any files"), which a coding-agent CLI is under no obligation to honor. A repo-mutating reviewer is a materially different safety posture from a read-only one, and the descriptor does not currently express it. - Plandex requires a stateful session (
plandex new) before a review turn. D2's singlebinary+args+ prompt-channel shape describes one invocation; it cannot express a two-phase setup-then-invoke sequence.
Neither is a reason to reject this design — both are exactly the "file an issue naming the primitive"
case, and both would likely be served by two future primitives: a declared invocation-safety posture,
and a setup phase on the descriptor. They are recorded because an ADR claiming a closed vocabulary
is sufficient, without naming what it excludes, is claiming more than it verified.
Revisiting this to permit a third-party handler module confined to the capability install root
(the ADR-1244 D7 model, which does allow third-party command modules) would genuinely deliver "any
plugin can ship a lane." It is rejected here because it reverses an ADR-1016 rejection rather
than amending it, and because D7 itself calls third-party code execution the highest-risk surface
and sequences it last. It should be revisited only with its own ADR and its own evidence.
D7 — probe.kind is a closed enum wider than existence, and every probe is bounded
probe.kind is a closed enum:
| Kind | Fields | Semantics |
|---|---|---|
command-exists |
binary |
command -v <binary> |
command-capability |
binary, needle, timeoutMs |
<binary> --help bounded, matched against needle |
http-reachable |
hostConfigKey, path, timeoutMs |
Bounded GET; reachable ⇒ available |
command-exists alone is structurally insufficient, and the evidence is concrete: kimi is
claimed by both Kimi Code CLI (Node) and the legacy Python kimi-cli — which is a separate,
first-party, non-reviewer runtime capability in this repo. An existence-only probe registers the
wrong tool. This was found in review of PR #2776 and is the reason the vocabulary ships wider than
one member.
Every probe that starts a process or a connection MUST be bounded. This repo carries a named
Unbounded Subprocesses defect class, and the original Kimi probe was a live instance of it: an
unbounded kimi --help | grep that ran on every /gsd:review invocation regardless of which
flags were passed, so a user whose Kimi binary waited on a first-run consent or auth prompt would
hang every future review — including reviews that never asked for that lane.
command-capability bounds via external timeout, falling back to gtimeout (the precedent
already set by the Antigravity block at review.md:560). Stock macOS ships neither; where no
bounding mechanism is available the probe is skipped and the lane reported unavailable, which
degrades a lane rather than hanging a command.
D8 — Uniqueness is a build-time conformance invariant
Across the merged first-party ∪ overlay set, reviewer.slug, reviewer.flags, and
reviewer.reviewsSection are each unique — for flags, over the flattened set of every lane's
flags, since one lane may declare several. A collision fails the build gate.
Amended 2026-07-29 (#2794): D1 originally declared a singular flag, and this invariant was
stated over it. antigravity is selected by both --antigravity and --agy
(review.md, docs/COMMANDS.md), which a single-valued field cannot express, so the field is
flags: string[] and uniqueness flattens across lanes. reviewsSection
uniqueness is not cosmetic: two lanes sharing a heading would silently merge their output in
REVIEWS.md, producing a review that appears to have consensus it does not have.
An overlay lane colliding with a first-party lane is rejected, first-party winning — the existing
id-uniqueness precedent (capability-manifest.md:167).
Reviewer instances are not lanes. review.reviewer_instances.<name> = {cli, model?, agent?}
(ADR-1517) lets one model-capable adapter run as several reviewer identities. Instances resolve
through a lane and continue to; they do not participate in the roster, the flag set, or this
uniqueness check.
D9 — Reviewer config keys become federated, and the roster derives from declared lanes
review.models.*, review.<host>_host, and review.max_prompt_tokens_per_reviewer.* move from the
central schema to federated config slices owned by their lane capabilities. Key names and
existing .planning/config.json files are unchanged; only validation provenance moves, so no user
migration is required. Per the config-key exclusivity invariant (capability-manifest.md:173), the
central-schema removal and the federated addition must land in the same commit or the build gate
fails on a key present in both.
Ownership is per-key and per-lane, so that no key is owned twice — Phase 4 implements this table rather than re-deriving it:
| Key | Owner | Notes |
|---|---|---|
review.models.<slug> |
the lane whose slug it names |
One key per lane; a lane with no model override declares none |
review.ollama_host |
ollama |
The hostConfigKey its own descriptor points at (D2) |
review.lm_studio_host |
lm_studio |
ditto |
review.llama_cpp_host |
llama_cpp |
ditto |
review.max_prompt_tokens_per_reviewer.<slug> |
the lane whose slug it names |
The lane's promptBudgetKey (D2) resolves to this |
review.max_prompt_tokens |
stays central | A global default across all lanes; owned by no single lane, so federating it would be wrong |
review.default_reviewers |
stays central | Selection policy over lanes (ADR-0011), not a property of any lane |
review.reviewer_instances |
stays central | Instance→lane mapping (ADR-1517); an instance is not a lane (D8) |
The last three rows matter as much as the first five: a key that describes policy across lanes must not be federated into one, and the exclusivity invariant would not catch that error — it only catches a key owned twice, not a key owned by the wrong side.
KNOWN_REVIEWER_SLUGS derives from declared reviewer bodies. hostBehaviors.reviewerCli survives
as a derived legacy alias for one release and is then removed. Where both a body and the alias
are present, the body wins. The field is undocumented, so external users are unlikely — but
"undocumented" is not "unused", which is why it gets a deprecation window and a changeset note
rather than a silent removal. The removal is owned by a named phase (#2801), not left implicit.
Consequences
Positive.
- Adding a reviewer becomes one manifest installed through
gsd capability install <url>— no core patch, no workflow edit, no release cycle — for any lane the vocabulary expresses. - A cross-cutting fix (empty output, effort channel, model lookup) becomes a single-site change covering every lane, retiring the #2494 → #2605 cadence.
- The roster gets one generated source, which makes the
DEFECT.GENERATIVE-FIXparity assertion for #2781 mechanical rather than per-lane. - Third-party lanes arrive behind the existing trust gate — disclosure, consent, SHA pin,
engines.gsd, reserved namespaces — instead of as an unreviewable prose block. - A lane owns its own configuration, closing a half-migrated config surface.
Negative, and accepted.
- The closed vocabulary must grow, under review, when a genuinely new lane shape appears. This is intentional friction and it is the trust boundary. D2 shows the cost is real: the first survey already forced one widening.
- Third-party lanes are data-only (D6). "Any plugin can ship a reviewer lane" overstates what this delivers; the ADR and the epic should both say most.
- Consent-at-install is a weaker gate for a standing egress channel than for a hook (D5). There is no sandbox.
discloseExecutableSurfacesis already cyclomatic 51 / cognitive 99 with five dependents. Adding a fourth class lands in an existing hotspot; the implementing phase should extract per-class helpers rather than grow the switch, and should expect the mutation gate to bite.- Two declaration mechanisms coexist for one release (D9).
- Normalizing empty-output handling is observable on lanes that previously returned nothing silently. That is a bug fix that breaks a workaround, and it needs a changeset note rather than a silent correction.
Explicitly unchanged: reviewer selection precedence (ADR-0011), the REVIEWS.md contract
(ADR-1517), and every existing lane's observable command shape.
Implementation phases (dependency-ordered)
Verified with /adr-phase-coverage: every deliverable is claimed by exactly one phase, every
hand-off lands, and every user-facing capability has a phase that wires its entry point.
Every decision is mapped to the phase that delivers it, so no decision is left to be "handled somewhere".
| Phase | Issue | Delivers | Deliverable |
|---|---|---|---|
| 0 | #2793 | — | This ADR |
| 1 | #2794 | D4 (explicit-selection carve-out only) | Core single-sourced invocation descriptor + DEFECT.GENERATIVE-FIX parity assertion; corrects the silently-degrading selector — closes #2690 |
| 2 | #2795 | D1, D2, D3, D4, D7, D8 | Manifest reviewer body and the third role; transport and probe.kind closed enums; registry harvest, validation, uniqueness; the absent-safe invariant and its warning channel |
| 3 | #2796 | D5 | The fourth trust-disclosure class: binary + args / host + hostConfigKey, egress payload classes, signature binding |
| 4 | #2797 | D9 (config half) | Federated config migration per the ownership table, same-commit |
| 5a | #2798 | D9 (roster half) | The 11 existing lanes declare reviewer bodies; roster derives; hardcoded tail deleted |
| 5b | #2799 | D6, D5 (invocation-time host re-verification) | invoke_reviewers / write_reviews iterate lanes; the antigravity and openai-compatible handler modules ported from the existing bash legs; the kimi-code lane — closes #2718 |
| 6 | #2800 | — | Docs, hostBehaviors documentation gap, capability matrix, locale parity gate — closes #2781 |
| 7 | #2801 | D9 (alias removal) | Remove the hostBehaviors.reviewerCli alias, the release after 5a |
Two mappings are worth calling out because a reader would otherwise assume the wrong phase. D6's handler modules are code, not data — porting Antigravity's ~100-line three-layer fallback and the three OpenAI-compatible lanes into named first-party modules is Phase 5b's work, delivered alongside the iteration that calls them. And D5 splits across two phases: the disclosure itself is Phase 3, but the invocation-time destination re-verification necessarily lands in Phase 5b, because that is where the invocation path is built.
Why kimi-code lands in 5b and not 5a. 5a makes the roster derive from declared bodies, but 5b
is what makes invoke_reviewers iterate them. The eleven existing lanes already have hand-authored
legs, so declaring them in 5a changes nothing observable. kimi-code is net-new with no leg —
declaring it in 5a would make it selectable but not invocable: present in --all, selected, and
producing an empty section for the whole 5a → 5b window. Landing it with the iteration keeps the
Phase 1 parity assertion green across the entire migration.
Alternatives considered
- A single unified
invokeshape. The design this ADR started from. Rejected on evidence: a read of all twelve legs found three that are HTTP endpoints with no binary (see Context). Had it shipped, Phase 2 would have bolted on an implicit second shape or stranded three lanes in the hardcoded tail this epic exists to delete. - Transport inferred from field presence (
binary⇒ spawn,hostConfigKey⇒ http). Fewer fields; rejected because a manifest with both or neither has undefined meaning. - A spawn-only body, leaving the three HTTP lanes in core. Smaller and sooner; rejected because it preserves a hardcoded tail and permanently bars a third party from shipping a local-model lane — the epic's own problem statement in miniature.
- Core descriptor table only (#2690 as filed). Single-sources invocation inside
review-reviewer-selection.ctsand collapses the eleven blocks. Cheaper and lands sooner, and it does fix the cross-cutting-defect cadence — but it does not make lanes installable: still a core patch, still no trust gate, still no federated config. Not discarded — adopted as Phase 1, so the descriptor shape is designed once under this ADR rather than twice. - Keep
hostBehaviors.reviewerCli, just document and validate it. Cheapest, and it does close the documentation gap. Rejected because it leaves problems (b) and (c) intact: a lane still cannot own its config, and the five non-installable reviewers still have nowhere to live. - A third-party
handlermodule confined to the install root. See D6 — the only option that genuinely delivers "any plugin"; rejected here as reversing rather than amending ADR-1016, and as the surface ADR-1244 D7 sequences last. Revisit with its own ADR. - Route lanes through MCP. Rejected: reviewers are batch, single-shot, ten-to-twenty-minute
invocations. An MCP server lifecycle adds nothing, and
mcpServersdisclosure already covers the cases that genuinely are servers. - One
role: "reviewer"for every lane, splitting the six dual-purpose runtimes. Cleaner discriminator; rejected for churn (D3).
Amendments
2026-07-29 — vocabulary widened by Phase 1 (#2794)
Phase 1 built the core descriptor table against all eleven shipped legs, which is the first time every lane's contract was written down in one place. That surfaced four cases the original survey did not cover. All four are additive widenings of closed enums, none reverses a decision, and each is forced by a lane that exists today rather than by a hypothetical:
| # | Decision | Was | Is | Forced by |
|---|---|---|---|---|
| 1 | D2 | promptChannel: stdin | argv | argv-file-ref |
adds none |
coderabbit is fed no prompt — it reviews the working-tree diff |
| 2 | D2 | outputChannel: stdout ("exactly one member today") |
adds file-arg |
codex already writes via -o/--output-last-message and discards stdout (#1698) |
| 3 | D2 | — | adds outputArg, required iff file-arg |
knowing the review lands in a file is useless without the argument naming it |
| 4 | D1, D8 | flag: string |
flags: string[], uniqueness flattened |
antigravity is selected by both --antigravity and --agy |
Why this is the process working, not a design failure. D2 already records that the original
draft assumed one lane shape and that reading all twelve legs disproved it — openai-http exists
because a survey produced evidence, not because anyone predicted it. These four are the same
mechanism at the next level of detail: the vocabulary widens when a real lane does not fit, under
review, and never on speculation. D6's escalation path ("file an issue naming the primitive the
vocabulary lacks") is for third parties; a first-party phase that finds the gap while implementing
amends the ADR directly, which is what happened here.
What this does not change. No decision is reversed. transport remains a closed two-member
discriminator selecting the invoke sub-shape; probe.kind and handler are untouched; the
absent-safe invariant (D4), the disclosure class (D5), and the config-ownership table (D9) are
unaffected. Phase 2 (#2795) implements the manifest validator against the vocabulary as amended
here, which is the point of amending rather than leaving it for Phase 2 to rediscover.
2026-07-29 — three factual corrections from Phase 2 (#2795)
Implementing the validator required reading the code each claim rests on. Three statements above did not survive that reading. None changes a decision; each would have misdirected a later phase, which is precisely why they are corrected here rather than worked around in code.
1. The cause of the stranded config keys was misattributed (Context (b), Scope of changes, D9).
The ADR attributes reviewer config keys living centrally to the runtime body forbidding feature-only
fields. That is not the mechanism. FEATURE_FIELDS_FORBIDDEN_ON_RUNTIME is
['skills','agents','steps','contributions','gates','hooks','activationKey'] — config is not in
it, and a role: "runtime" capability carrying a config slice passes validation today. The real
cause is two harvest sites that never read it:
gen-capability-registry.cjsnested its config-harvest loop inside therole === 'feature'branch, so a non-feature capability'sconfigwas silently dropped fromconfigKeys/configSchema.validateCrossCapabilityopened its config-key ownership loop withif (cap.role !== 'feature') … continue, so a non-feature capability was exempt from both single-ownership and the central-schema collision check.
Phase 2 fixes both by filtering on the presence of a config slice rather than on the role.
This matters for Phase 4 (#2797), which would otherwise have been designed against a constraint that
does not exist — and it means the exclusivity invariant was, until now, unenforced for every
non-feature capability rather than merely unused.
2. D3's profile-membership claim is inverted.
D3 states that a role: "reviewer" capability "receives profile membership from
deriveProfileMembership (gen-capability-registry.cjs:201-213) like any other" and that the
membership is inert. It receives no membership: that function skips any capability without a
non-empty skills array, and a lane-only capability has none. The intended outcome — a reviewer
capability installs nothing — holds exactly as D3 wanted, and tier remains required as the source
of truth for the requires-closure. Only the stated mechanism was wrong, and a Phase 5a author
following D3 would have gone looking for membership that is not there.
3. The specified capability folder names for two lanes would fail the build.
The Scope-of-changes section and #2798 both name capabilities/lm_studio/ and
capabilities/llama_cpp/. Both would be rejected: id must equal the folder name and match
KEBAB_RE (/^[a-z][a-z0-9-]*$/), which does not admit _. Three namespaces are in play for one
lane and they are deliberately not the same string:
| value | casing | fixed by | |
|---|---|---|---|
capability id / folder |
lm-studio |
kebab | the id conformance invariant |
reviewer.slug |
lm_studio |
snake | the shipped roster and review.lm_studio_host, which D9 leaves unchanged |
reviewer.flags |
--lm-studio |
kebab | the shipped flag |
Phase 2 therefore validates reviewer.slug against its own pattern rather than reusing KEBAB_RE,
which would have rejected two shipped lanes. Phase 5a must create capabilities/lm-studio/ and
capabilities/llama-cpp/, each declaring the snake-case slug.
(Corrected 2026-07-29: this paragraph first recorded the pattern as /^[a-z][a-z0-9_-]*$/. Phase 2's
own security review caught that as a divergence from Phase 1's exported LANE_SLUG_RE, which permits
a leading digit — a model-named lane such as 4o-mini would have been accepted by the core descriptor
and rejected by the manifest validator, reintroducing exactly the translation layer this epic deletes.
The shipped pattern is /^[a-z0-9][a-z0-9_-]*$/, and a parity assertion now fails if the two ever
drift again.)
2026-07-29 — Phase 4 and Phase 5a are swapped (ordering correction from Phase 5a)
The phase table above runs Phase 4 (federated config) before Phase 5a (lane declarations), and #2798 states "Depends on Phases 2 and 4". That ordering is inverted, and it makes Phase 4 unsatisfiable.
D9's ownership table assigns review.ollama_host to the ollama lane, review.lm_studio_host to
lm_studio, and review.llama_cpp_host to llama_cpp. A federated config slice lives inside a
capabilities/<id>/capability.json — and those capability directories do not exist until Phase 5a
creates them. Verified before the swap: capabilities/{ollama,lm_studio,llama_cpp,gemini,coderabbit}
were all absent, and only the six reviewerCli-flagged runtime capabilities existed.
So in the stated order Phase 4 has nowhere to put three of its five key families, and its own "Done
when" — "review.<host>_host owned by lane capabilities" — cannot be met. Shipping it unmet would
be a failed deployment under CI.GATE.acceptance-criteria-required.
5a's stated dependency on Phase 4 is likewise unfounded: declaring a reviewer body requires only the
manifest vocabulary from Phase 2. The real dependency graph is Phase 2 → 5a → 4, with
5a → 5b → 6 unchanged. Nothing about either phase's content changes — only their order.
A second correction, to #2798's acceptance list. It requires "docs/INVENTORY.md updated +
node scripts/gen-inventory-manifest.cjs --write run after build:lib". That rests on a false
premise: the inventory catalogs bin/lib/*.cjs modules, not capability directories —
antigravity, opencode and qwen appear zero times in it, and INVENTORY-MANIFEST.json's six
families contain no capabilities/ entry at all. gen-inventory-manifest.cjs --check passes with the
five new capability directories added and no inventory edit. The item is vacuous for this phase, and
inventing an edit to satisfy it would introduce drift rather than prevent it.
2026-07-30 — vocabulary widened by Phase 5b (#2799)
Phase 5b is the cutover: it deletes the ~640 lines of hand-authored bash and runs every lane from the declaration. Building the resolver against all twelve legs — the first time each leg's runtime contract, not just its shape, had to be reproduced — surfaced five gaps. All five are additive, each is forced by a lane that ships today, and none reverses a decision. This is the same mechanism D2 and the Phase 1 amendment record, at the next level of detail.
| # | Decision | Was | Is | Forced by |
|---|---|---|---|---|
| 1 | D6 | handler: null | antigravity | openai-compatible |
adds opencode |
opencode's review is RECONSTRUCTED from assistant text parts of a --format json stream; a plain stdout copy writes the raw JSON envelope into REVIEWS.md (#1936). Admitted under the enum's own second arm — a documented upstream defect data cannot express — exactly as antigravity was |
| 2 | D1 | model key implicit as review.models.<slug> |
adds reviewer.modelConfigKey |
antigravity's slug is antigravity but its shipped key is review.models.agy. Resolving by slug misses it and silently ignores a configured model, disabling the pinned-model escape hatch #2073 added |
| 3 | D2 | invoke.args a fixed array |
an argv template over a closed four-member placeholder set ({{model}}, {{effort}}, {{output}}, {{prompt}}) |
The injected pieces do not all go in the same place: codex injects the model after its exec subcommand and the output file later still, while five lanes end with a bare - that must stay last. Positional splicing produced codex --model M -o F exec --ephemeral …, which is not a valid invocation |
| 4 | D2 | openai-http invoke had no default |
adds invoke.defaultHost and invoke.fallbackModel |
Phase 4 federated every review.*_host with a default of "", so the real fallback (http://localhost:11434, llama3, …) existed only inside the bash leg. A data-driven lane would POST to a garbage URL |
| 5 | D7 | — | kimi-code lands with a command-capability probe |
Net-new lane, per the phase table. kimi is claimed by both Kimi Code CLI and the legacy Python kimi-cli (analysis from closed PR #2776, credit @drungrin) |
modelConfigKey is OPTIONAL, and that is D4 rule 2 rather than a convenience. It did not exist
before this phase, so requiring it would fail validation on every reviewer manifest authored against
an earlier GSD. Absent reads as null.
D5 rule 1 was recorded as delivered by Phase 3 and was not implemented. The implementation note
added to D5 on 2026-07-29 states that "the consent record additionally stores the resolved host".
It did not: ConsentRecord carried no host field, recordProjectConsent accepted none, and nothing
in the tree bound one — so this phase's rule-4 comparison had no baseline to compare against. Phase
5b implements it, as an optional reviewerHost that isValidConsentRecord does not require, so
no record already on disk is invalidated and no re-consent storm fires (D4 rule 5). It stays out of
disclosureSignature for the reason that note gives. Recorded here because the ADR asserting a rule
was delivered is precisely what would stop a later phase from checking.
Three runtime dependencies leave the review path, and two of them were platform holes rather than
mere overhead: jq (absent on stock Windows/Git-Bash, #2589 — it gated five lanes), curl, and the
external timeout/gtimeout the Antigravity leg probed for. Stock macOS ships neither killer, so
D7's "where no bounding mechanism is available the probe is skipped" carve-out was, in practice, that
lane running unbounded on every stock Mac. spawnSync's native timeout is always available, so the
bound is now unconditional and that carve-out is obsolete.
The DEFECT.GENERATIVE-FIX parity gate is re-pointed. Phase 1's assertion required a literal
<!-- reviewer-lane: <slug> --> per lane inside invoke_reviewers and a literal
## <Section> Review per lane inside write_reviews — the exact text this phase deletes. Those two
families could not be kept without keeping the hand-maintained per-lane blocks the epic exists to
remove, so they are replaced by descriptor ↔ registry parity in both directions (the registry is
what the runtime iterates once lanes are data) plus an anti-parity assertion that fires if a
bespoke leg is ever re-added. That also gives #2781/Phase 6 the mechanical single source its docs and
locale gate needs, which per-leg text could never provide.
2026-08-03 — D2 spawn invoke vocabulary widened by #2483 (invoke.env)
The claude lane was the only reviewer additionally inheriting the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory — a context asymmetry against the
workflow's own independent-review premise, since gemini sees only the assembled prompt and codex
runs --ephemeral. Closing it needs two environment variables set for that one spawn. Additive, and
forced by a lane that ships today.
| # | Decision | Was | Is | Forced by |
|---|---|---|---|---|
| 1 | D2 | spawn invoke carried no way to shape the child's environment |
adds invoke.env — optional, an object of environment name/value pairs with string values only; forbidden on openai-http |
The claude lane must spawn with CLAUDE_CODE_DISABLE_CLAUDE_MDS=1 CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 (#2483). The pairs are static per lane, so this is declared data — not a handler, which D6 reserves for behavior data cannot express |
Declared data rather than a handler, and D6 is the wrong authority for it. An earlier revision of
this change cited D6 in the source comment. D6 governs the closed handler enum — imperative
behavior admitted first-party — and says nothing about the invoke field vocabulary, which is D2's
territory. The citation did not cover the widening, which is why this entry exists rather than a code
comment pointing at the wrong decision.
env is OPTIONAL, per D4 rule 2, exactly as modelConfigKey was in the Phase 5b entry above: it
did not exist before this change, so requiring it would fail validation on every reviewer manifest
authored against an earlier GSD. Absent means the lane inherits the environment unchanged.
Forbidden on openai-http, and registered in the discriminator to make that enforceable. An
openai-http lane spawns no child, so an environment pair there has no referent. The first
implementation validated env's shape but left it out of SPAWN_ONLY_INVOKE_FIELDS — the list the
openai-http arm rejects against — so it was accepted on that transport in silence, alone among the
spawn fields. Both registrations are required; neither implies the other.
What this does not change. No decision is reversed. transport remains a closed two-member
discriminator; effortChannel stays in neither field list because D2 defines it for both
transports, so it is shared rather than spawn-only.
env IS added to the D5 disclosure signature, and to the human consent prompt. An installed
overlay reviewer body reaches resolveLanePlan and is executed: routeReviewLane builds its lane
map from mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true})) (D8, #2927 /
#3062), and that merge is field-identical per D1 — it admits the overlay body without deep-validating
invoke, precisely because the invocation seam is where a lane is re-validated before it runs. So a
third-party manifest can declare env on a reviewer lane and have those pairs applied to the spawned
child. Undisclosed, that is arbitrary code execution behind a consent prompt that never mentioned it
(NODE_OPTIONS=--require ./evil.js; LD_PRELOAD on POSIX). D5 already folds env into MCP-server
disclosure and names that exact shape as the reason; reviewer lanes now carry the identical treatment.
2026-08-05: the earlier reading — that
envneeded no disclosure because manifestinvokefields never reachresolveLanePlan— is withdrawn. It was true when written and #3062 retired it. See git history for the superseded text.
The enumeration was the defect, not the missing name. env was the ninth invoke field that
reaches resolveLanePlan without being bound by the D5 lane signature; the other eight were
defaultHost, path, outputChannel, outputArg, modelArg, effortChannel, modelDiscovery and
fallbackModel. Two of those are egress-relevant on their own — defaultHost is the destination the
manifest itself declares, used whenever hostConfigKey resolves to nothing (configured ?? declaredDefault), and path completes the URL — so a lane with an unresolved config key disclosed
"(unresolved …)" while shipping the D5 egress payload classes to an address of the manifest's
choosing. Adding a ninth name would have left a tenth open, so the lane element instead carries a
residual of every other declared invoke key, mirroring the rawConfig completeness backstop the
MCP surface has carried since #1459 finding 5. env and defaultHost are additionally named
explicitly, mirroring that same line's deliberate explicit-then-backstop overlap.
D4.5 is preserved one level down, and the cost it guards against does not arise here anyway. The residual element is appended to the lane tuple only when the lane declares something beyond the eight already-bound fields, so a lane declaring none keeps a byte-identical signature. State the scope of that property honestly, because it is easy to oversell in both directions:
- It is vacuous for any VALID lane, and that is the honest statement. Once the residual covers the
lane body's outer fields too, a lane that produces no residual is one declaring no
flags, noprobe, noemptyOutput, noevidenceClass, norequiresBinariesand nopromptBudgetKey— i.e. a body the validator rejects. Measured across the twelve first-party reviewer capabilities: zero are in the byte-identical class. The conditional append is still correct — it keeps the signature minimal and means the residual element carries information when present — but it is a property of the encoding, not a claim that anyone's signature is unchanged. - And no capability is re-prompted regardless. A code change to
disclosureSignaturecannot invalidate an existing consent:hasProjectConsentmatches on the recomputed bundlecontentHash— the signature is explicitly "no longer the security binding" (#1459 CB-1/CB-2) — and the upgrade path'sexecutableSetChanged(old, new)compares two disclosures both computed by the current code, so widening the signature shifts both sides equally. - What the widening actually buys is therefore forward-looking and is the whole point: an upgrade
whose manifest edits
env,defaultHost, or any other declaredinvokefield now registers as an executable-surface change and re-consents, where previously it could change what the lane runs in silence. First-party capabilities never enter this path at all — the install flow blocks a first-party id before trust evaluation.
Validation is defence in depth; consent is the boundary. invoke.env is validated for object
shape, POSIX name grammar and string values, and a denylist refuses execution-primitive names
outright on a reviewer lane — PATH, NODE_OPTIONS, LD_PRELOAD, DYLD_INSERT_LIBRARIES,
BASH_ENV, PYTHONPATH, PERL5OPT, RUBYOPT, GIT_SSH_COMMAND, JAVA_TOOL_OPTIONS and
siblings, matched case-insensitively because Windows environment lookup is. PATH is included
deliberately: it is the most complete primitive of the set, and a lane needing a specific executable
declares an absolute invoke.binary rather than reshaping the child's PATH.
State the limit plainly, because the list invites being mistaken for the control: it cannot be complete against an arbitrary third-party child, and disclosure runs before validation — on manifests validation would reject. So the boundary remains install-time consent, which shows every declared pair and binds it to the signature; execution-primitive names additionally carry a warning line in the prompt. A name missing from both lists costs a quieter line on a value the user is still shown.
One inconsistency this entry closes, and it was real. Consequences above states that adding a
reviewer is "one manifest … no core patch", and CONTEXT.md, gsd-core/workflows/review.md and
resolveLanePlan's own header all describe overlay manifests reaching the resolver — while the
runtime, until #3062, built laneBySlug solely from the first-party table and rejected every slug
absent from it. Four documents on one side, the runtime on the other. #3062 resolved it in the
documents' favour, which is what makes the disclosure above mandatory rather than defensive.
2026-09-07 — a new consumer axis: the supportsReviewerLanes capability-step trait (#4209)
Every decision above (D1–D9, and every dated entry so far) governs the role: "reviewer" capability
body — the shape of a lane declaration — and its one consumer, /gsd:review. #4209 adds a second,
unrelated consumer: /gsd:code-review, which does not want a second internal review pipeline, only
the existing resolveLanePlan/runner.runLane machinery reused to corroborate its own single
internal reviewer with external source-review evidence. That consumer is not a role: "reviewer"
capability at all — it is an ordinary role: "feature" capability's steps[] entry (the
capability-manifest.md axis this ADR's own scope note, Context (b), names as structurally distinct
from the runtime/reviewer body and explicitly out of this ADR's reach). This entry records that
extension, since it is additive to the reviewer-lane surface without being expressible inside D1–D9.
What was added. A step entry in a role: "feature" capability's steps[] array (per
capability-manifest.md's existing steps table) may carry an optional supportsReviewerLanes: true
field alongside its required point/ref/produces/consumes/onError. Validated in
capability-validator.cjs (must be the literal boolean true; any other type fails validation;
false/omitted are inert — no key on the projected ActiveHook). Projected through
loop-resolver.cts's resolveLoopHooks/resolveActiveHooksForPoint onto the step's ActiveHook as
supportsReviewerLanes: true. A workflow step whose ActiveHook carries the trait may call the new,
capability-neutral interpreter dispatchReviewerLanes (src/reviewer-step-dispatch.cts), which reuses
resolveReviewerSelection and resolveLanePlan/runner.runLane — the SAME D1–D9-governed
plan/invoke machinery this ADR already specifies — rather than reimplementing dispatch. No new
invocation mechanism was created; only a new, generic activation seam for the existing one.
Why this is additive, not a reversal. No D1–D9 decision changes. The reviewer capability body,
its ten decisions, resolveLanePlan, and runner.runLane are consumed exactly as specified;
/gsd:review itself is untouched. What is new is a second caller of that machinery, reached through
a different capability axis than this ADR covers, and a trust boundary this ADR never needed: a
role: "feature" step's dispatch target now receives lane output as evidence to independently
re-verify, not as a second output schema — enforced by dispatchReviewerLanes's fail-closed request
validation (path traversal, absolute paths, symlink escape, missing/invalid provenance, budget
overflow) and by gsd-code-reviewer's untrusted-evidence consolidation contract, neither of which
/gsd:review's existing consumer needed since it already fully owns its own output contract.
Scope this does not touch. steps/gates/contributions as a capability axis are governed by
ADR-857 (Loop Extension Points) and ADR-894
(declaration format), both already Amended by this ADR for the reviewer axis — this entry does not
add a new Amends relationship to either, since supportsReviewerLanes is one optional field on an
already-Amends-covered steps[] entry, not a new axis of its own.