Files
msd-core/gsd-core/workflows/review.md
Tom Boucher 929e02cb2c enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict

ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking
answer. Three families do exactly that today; this commit pins each one RED.

Measured on this tree, 2026-08-27:

  intel query, .planning/intel/file-roles.json nested 12000 deep
    -> exit 1, "Error: Maximum call stack size exceeded"
       searchJsonEntries / matchesInValue carry no depth parameter at all.
       The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage
       (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage
       never received it.

  same fixture nested 48 and 49 deep
    -> both return total=1 at exit 0, truncated=undefined
       Nothing distinguishes "searched to the bottom" from "stopped looking".

  query phase-plan-index, a plan whose depends_on names an unresolvable token
    -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it
                  in wave 1"]
       The token is never mentioned. computeDependencyLevels drops the edge
       with `if (!resolvedDep) continue;`, every plan becomes a root, and the
       tool then reports the author's correct wave: as the thing that is wrong.

  countPhasePlansAndSummaries with fs.readdirSync throwing EACCES
    -> hasContext:false, indistinguishable from a phase that simply has no
       CONTEXT.md. context_read_error is undefined.

The shapes these tests assert against, chosen here so the implementation has a
target rather than inventing one later: `truncated: boolean` on the intel query
result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and
`context_read_error: string | null` per analyzed phase.

Deliberately green, and they must stay that way — each stops the fix from
over-firing:

  depth 48 is found and NOT flagged truncated (the ceiling is inclusive)
  a shallow miss reports no truncation           (noise control, N1)
  10,000 siblings at depth 2 are unaffected      (the bound is DEPTH, N2)
  a genuine wave: mismatch on a fully-resolved DAG still warns (N3)
  a genuinely missing directory is absent, not an error
  the emitted depends_on display mapping still passes an unresolved token
    through verbatim — already pinned by the existing #3785 test, so no
    duplicate was added

T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the
real CLI and reads the emitted JSON, because a unit assertion on
computeDependencyLevels would have passed throughout #3427's life.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3885): no silent swallow, and no verdict manufactured from dropped data

Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed
into an output that reads as authoritative.

The recursion bound, restored but NOT verbatim (src/intel.cts)

  MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and
  matchesInValue, which carried no depth parameter at all. The bound existed in
  the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the
  surviving .cts lineage never received it — §8.3's "a consolidation may not
  delete an invariant along with the surface that held it", demonstrated.

  Measured before: a .planning/intel file nested 12000 deep exits 1 with
  "Error: Maximum call stack size exceeded". Reachable from a project document.

  The original returned a bare `false` at the ceiling. Restoring that verbatim
  would trade a crash for a silent "no match" when the truth is "I stopped
  looking" — the same class this epic exists to close, and ADR-3473 Decision 4
  forbids it. So the bound carries a truncation signal:

    nesting 47 -> found,     truncated false
    nesting 48 -> found,     truncated false      (the ceiling is inclusive)
    nesting 49 -> not found, truncated TRUE
    nesting 12000 -> exit 0, truncated TRUE, no RangeError

  A shallow document that simply has no match reports truncated FALSE — the
  flag means "I stopped early", never "I found nothing", or it would be noise.
  The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected.

The dropped edge is named, and stops being blamed on the author (src/phase.cts)

  computeDependencyLevels dropped every unresolvable depends_on token with a
  bare `continue`. Each drop makes a plan a root, so the whole phase collapses
  to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as
  the thing that was wrong.

  Before:
    warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in
                wave 1"]
  After:
    warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not
                resolve to any plan in this phase — edge dropped, wave placement
                for this plan may be unreliable"]

  The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG
  and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId
  stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is
  deliberately not built here. The emitted depends_on display mapping still
  passes an unresolved token through verbatim (#3785).

No artifact from failed inputs (gsd-core/workflows/review.md, #3352)

  A failed lane leaves no result file, so "every lane failed" is exactly "the
  aggregate JSONL has zero lines" — the gate condition already existed as a
  byproduct. REVIEWS.md is no longer written in that case, and the commit step
  is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT
  counted as a failure. Per-lane output and non-empty .err are preserved to
  .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that
  the lanes failed at all; the commit step names one file, never a glob, so the
  diagnostics are not swept in.

Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2)

  Four callers collapsed an EACCES on a phase directory into [] and reported
  hasContext:false — byte-identical to a phase that simply has no CONTEXT.md.
  Each now names the directory it could not read. A genuinely missing directory
  stays absent rather than becoming an error, which is what keeps the fix from
  over-firing.

Fatal errno folded into a retry set: audited, no defect found

  Reported as a verified negative rather than padded with a change.
  withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776;
  atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by
  construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they
  return or rethrow the final error rather than swallowing it. estimate-cli's
  sole caller surfaces that rethrow as write_error in its JSON output.
  Manufacturing a diff to make the checkbox look worked-on is the Goodhart
  outcome Decision 6 exists to prevent.

Disclosed: R46 (the commit step names one file, never a glob) is a real
regression guard but is NOT independently failing-first — the commit fence is
byte-identical pre- and post-fix, so it only fails pre-fix through its shared
extraction dependency. Recorded rather than claimed as fail-first.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence

Two review findings, both real, both in my own change.

An isolated adversarial review found the evidence-preservation block never
checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in
a SEPARATE fenced block. A disk-full or unwritable phase directory therefore
still destroyed the only copy of the failed lanes' output — reintroducing the
exact #3352 data loss this item exists to stop, inside the fix for it.

Preservation and cleanup are now one block, because each fenced block is a
separate execution and a shell variable cannot carry between them. mkdir -p and
each cp are exit-checked; cleanup runs only when preservation succeeded, and a
failure warns naming the intact run directory. "Nothing to preserve" is not a
failure and still cleans up. Driven three ways: success removes run_dir, failure
leaves it intact with the warning, nothing-to-preserve removes it. The failure is
induced by a file-vs-directory conflict rather than chmod 0o000, which root
bypasses.

The new unresolved-depends_on warning embedded a user-authored token verbatim:

  warnings: ["Plan 03-02: depends_on token \"evil
  Plan 03-01: FORGED WARNING\" does not resolve ..."]

The JSON wire form is safe, and the security reviewer judged it non-exploitable
for that reason. It is escaped anyway through formatDiagnosticToken — the helper
#3884 added one phase earlier for exactly this class. warnings[] is an array a
consumer naturally prints line by line, and not reusing the sibling fix is the
generative-fix-divergence shape this epic exists to close. The same treatment is
applied to context_read_error / phase_dir_read_error, which embed a phase
directory path a repository can choose, and to the fs error message, which
echoes the raw path itself.

Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow
array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per
negative space N2, disclosed rather than left to be discovered.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot"

Blocker from the round-2 isolated review, and it is my own inconsistency:
this phase applied "unreadable is not absent" to phase directories and left it
broken in the file it was already editing.

  chmod 000 .planning/intel/file-roles.json
  gsd-tools intel query <term>
  -> {"matches":[],"total":0,"truncated":false}  exit 0

safeReadJson swallowed every read failure and returned null, so an EACCES was
byte-indistinguishable from an absent file AND from a genuine no-match. Now it
separates three states: ENOENT stays silently absent, because not every project
has every intel file and intelQuery loops over all of them expecting misses;
EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel
file previously read as "no matches" too — same defect, same fix.

Threading that outcome through the other three callers found something worse
than the reported case. intelDiff returned no_baseline:true for a corrupt or
unreadable snapshot — not a silent failure but an actively FALSE verdict, telling
the caller they never took a snapshot when they did. That is §8.5's headline
case, so it is fixed and tested rather than noted. intelStatus and
intelApiSurface collapsed the same way; intelApiSurface additionally printed a
"not yet populated" banner that was simply untrue.

Every row is failing-first, including the absent-file ones — the field is new,
so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin
that the fix does not OVER-fire on the ordinary absent case, which is what would
turn this into noise on every project lacking an intel file. IO failure is
injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root
bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as
a test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): build the pathological intel fixture as text, not by stringifying a nested object

The remote runner came back red on Linux with two failures, both
T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on
macOS. The product was never at fault.

writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then
JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed
the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned.
Linux's container stack is smaller than macOS's, which is the whole of the
platform difference.

Measured, with the same document built as JSON TEXT so nothing in the building
process recurses:

  depth=100    rc=0 truncated=true
  depth=5000   rc=0 truncated=true
  depth=12000  rc=0 truncated=true
  depth=60000  rc=0 truncated=true

V8 parses this shape iteratively; only stringify recurses. The bound works at
every depth tried.

The fixture is now built by string concatenation. That is also the more faithful
input — a real deeply nested JSON document on disk is exactly what the bound
guards, where a stringified object was only ever a way to produce one.

The depth stays 12000. Lowering it would have made the test pass by weakening it
to accommodate a fixture bug, and 12000 is a legitimate pathological input the
product handles. T4 remains a genuine fail-first: rebuilt against the parent of
the commit that added the bound, the string-built depth-12000 fixture still
drives the CLI to rc=1 with "Error: Maximum call stack size exceeded".

A comment records why the fixture is text, so it is not "simplified" back into a
macOS-green / Linux-red test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3885): backfill the changeset PR number

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): normalize path separators before splicing into the workflow's bash

CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the
remote runner were all green.

  AssertionError: commit must name the single REVIEWS.md file; got:
    --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md

Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The
harness spliced an OS-native temp path into the extracted bash, and bash consumes
\U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke
RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory
survived — which is the other two assertions.

This is a fixture defect, not a product one, and that was checked rather than
assumed. In production the phase directory is toPosixPath-normalized at every
call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and
the run directory is created by "mktemp -d" running inside the bash block itself
(gsd-core/workflows/review.md:163), which emits POSIX-style output even under
Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it.

The file's pre-existing #3034 harness splices raw native paths too, but only ever
inside double-quoted assignments, so it never tripped this — my new harness
followed that convention faithfully into the one place where it does not hold.
Both now splice through toPosixPath from shell-command-projection, the
established seam, which is a no-op on POSIX and mirrors what production does.

No assertion was weakened. "commit must name the single REVIEWS.md file" and
"the run dir must still be destroyed" still assert exactly that; only how the
fixture supplies its path changed. Nothing is skipped on Windows — a t.skip()
here would have hidden the question of whether the exposure was real, which is
the question that mattered.

Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI
string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input
produces a byte-identical shape, proving the normalization is idempotent.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): stop the harness making the deleted run dir its own cwd

Windows shard 3/3 stayed red after the separator fix, on two assertions the
separator fix never touched:

  AssertionError: the run dir must still be destroyed
  AssertionError: nothing to preserve is not a failure — run dir must still be removed

The separators were a real bug and fixing them fixed the --files assertion. They
were not this bug, and two CI cycles went into the wrong axis before I stopped
converting path forms and looked at what the harness actually does.

runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's
working directory WAS the directory the block under test then removes with
rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified
locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live
process's working directory cannot be removed. So on Windows the directory
survived and both assertions failed, on macOS and Linux it vanished and they
passed. Nothing to do with slashes.

Harness-only. Production never cd's into the run directory; every reference is by
absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself
(gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged.

Fix: the child now runs with its cwd in an unrelated temp directory that the
block under test never deletes. Neither assertion was weakened, and nothing is
skipped on Windows — the tests in this file carry no platform guard and run
there unconditionally, which is how this surfaced at all.

Honest limit: the Windows failure mode cannot be reproduced on macOS, because
POSIX permits the very thing Windows refuses. The diagnosis is grounded in that
documented divergence and in the fact that only the Windows lane failed, but the
green outcome on windows-latest is unverified until CI runs it.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 04:12:47 -04:00

38 KiB

Cross-AI peer review — invoke external AI CLIs to independently review phase plans. Each CLI gets the same prompt (PROJECT.md context, phase plans, requirements) and produces structured feedback. Results are combined into REVIEWS.md for the planner to incorporate via --reviews flag.

This implements adversarial review: different AI models catch different blind spots. A plan that survives review from 2-3 independent AI systems is more robust.

Check which AI CLIs are available on the system:
_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; _gsd_at() { for _p; do if [ -f "$_p" ]; then GSD_TOOLS="$_p"; return 0; fi; done; return 1; }; if _gsd_at "${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif unset -f gsd_run; _G="$(command -v gsd_run)"; then GSD_TOOLS="$_G"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif _gsd_at "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd_run is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; GSD_IDENTITY_STATUS=unverified; case "$(gsd_run runtime-identity --raw 2>/dev/null || true)" in '{"packageName":"@opengsd/gsd-core"'*'}') GSD_IDENTITY_STATUS=ok;; esac; export GSD_IDENTITY_STATUS; [ "$GSD_IDENTITY_STATUS" = ok ] || echo "WARNING: \"$GSD_TOOLS\" did not prove it is @opengsd/gsd-core - it is either a different package or an @opengsd/gsd-core older than the runtime-identity verb. See docs/how-to/diagnose-a-foreign-gsd-tools.md" >&2; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
# Check each CLI
command -v gemini >/dev/null 2>&1 && echo "gemini:available" || echo "gemini:missing"
command -v claude >/dev/null 2>&1 && echo "claude:available" || echo "claude:missing"
command -v codex >/dev/null 2>&1 && echo "codex:available" || echo "codex:missing"
command -v coderabbit >/dev/null 2>&1 && echo "coderabbit:available" || echo "coderabbit:missing"
command -v opencode >/dev/null 2>&1 && echo "opencode:available" || echo "opencode:missing"
command -v qwen >/dev/null 2>&1 && echo "qwen:available" || echo "qwen:missing"
command -v cursor-agent >/dev/null 2>&1 && echo "cursor:available" || echo "cursor:missing"
command -v agy >/dev/null 2>&1 && echo "antigravity:available" || echo "antigravity:missing"
command -v kimi >/dev/null 2>&1 && echo "kimi-code:available" || echo "kimi-code:missing"

# Check local model servers (OpenAI-compatible HTTP API — no CLI binary required)
OLLAMA_HOST=$(gsd_run query config-get review.ollama_host --raw 2>/dev/null || echo "")
if [ -z "$OLLAMA_HOST" ] || [ "$OLLAMA_HOST" = "null" ]; then OLLAMA_HOST="http://localhost:11434"; fi
curl -s --max-time 2 "${OLLAMA_HOST}/v1/models" >/dev/null 2>&1 && echo "ollama:available" || echo "ollama:missing"

LM_STUDIO_HOST=$(gsd_run query config-get review.lm_studio_host --raw 2>/dev/null || echo "")
if [ -z "$LM_STUDIO_HOST" ] || [ "$LM_STUDIO_HOST" = "null" ]; then LM_STUDIO_HOST="http://localhost:1234"; fi
curl -s --max-time 2 "${LM_STUDIO_HOST}/v1/models" >/dev/null 2>&1 && echo "lm_studio:available" || echo "lm_studio:missing"

LLAMA_CPP_HOST=$(gsd_run query config-get review.llama_cpp_host --raw 2>/dev/null || echo "")
if [ -z "$LLAMA_CPP_HOST" ] || [ "$LLAMA_CPP_HOST" = "null" ]; then LLAMA_CPP_HOST="http://localhost:8080"; fi
curl -s --max-time 2 "${LLAMA_CPP_HOST}/v1/models" >/dev/null 2>&1 && echo "llama_cpp:available" || echo "llama_cpp:missing"

# jq prerequisite (#2589). The config/model/budget lookups in this workflow no
# longer need jq — they use the native --raw/--pick flags. But the lanes listed
# under "jq-dependent reviewer lanes" below parse structured JSON that gsd-tools
# does not emit (OpenAI-compatible /v1/chat/completions responses, opencode's
# JSONL event stream, agy's conversation cache), so they cannot run without jq.
# Probe it here rather than letting each lane swallow exit 127 into empty output.
command -v jq >/dev/null 2>&1 && echo "jq:available" || echo "jq:missing"

jq-dependent reviewer lanes. jq is a production prerequisite for the ollama, lm_studio, llama_cpp, opencode, and antigravity lanes only. If detect_clis reports jq:missing, treat those five as undetected — they follow the same "known-but-undetected" path as a missing CLI. Which path that is depends on how the lane was selected (see the precedence rules below): reached through review.default_reviewers or --all it is an info note and the lane is ignored; named by an explicit flag it is an error, because the user asserted that lane. Tell the user to install jq:

NOTE: jq is not on PATH — the ollama, lm_studio, llama_cpp, opencode, and
antigravity reviewer lanes are unavailable. Install jq (https://jqlang.org/download/)
or select a lane that does not require it (--gemini, --claude, --codex,
--coderabbit, --qwen, --cursor).

The remaining lanes (gemini, claude, codex, coderabbit, qwen, cursor) do not require jq and must stay selectable on a jq-less host.

Parse flags from $ARGUMENTS:

  • --gemini → include Gemini
  • --claude → include Claude
  • --codex → include Codex
  • --coderabbit → include CodeRabbit
  • --opencode → include OpenCode
  • --qwen → include Qwen Code
  • --cursor → include Cursor
  • --agy or --antigravity → include Antigravity CLI
  • --kimi-code → include Kimi CLI
  • --ollama → include Ollama (local server, OpenAI-compatible)
  • --lm-studio → include LM Studio (local server, OpenAI-compatible)
  • --llama-cpp → include llama.cpp (local server, OpenAI-compatible)
  • --all → include all available (CLIs + running local servers)
  • No flags → if review.default_reviewers is set, include only configured reviewers that are detected; otherwise include all available

Reviewer-selection precedence:

  1. Individual reviewer flags (--gemini, --codex, etc.)
  2. --all
  3. review.default_reviewers
  4. No key + no flags → all detected reviewers

Explicit reviewer flags are an assertion, not a preference (ADR-2782 D4). A lane the user named on the command line and that cannot run is an error, surfaced and non-silent — even when other named lanes did run. Do not proceed with a thinner reviewer set and report success: --gemini --qwen on a host without qwen fails, it does not quietly become a Gemini-only review. This applies however the lane became unavailable — binary missing, prerequisite jq absent, or a local server not reachable.

The asymmetry is deliberate: not finding a lane nobody asked for is normal; failing to run a lane somebody asked for is an error. A user who wants "whatever is available" has --all; a user who wants a preferred set has review.default_reviewers. Both stay lenient below.

review.default_reviewers behavior:

  • Value must be a non-empty array of slug strings (configured via gsd config-set review.default_reviewers '["gemini","codex"]')
  • Unknown slugs warn and are ignored
  • Known-but-undetected slugs emit an info note and are ignored — a configured default is a preference evaluated across many hosts, so a subset being present is expected, not an error
  • If all configured reviewers are unavailable, fail with an actionable message

If section_manifest is null or "reviewer-instances-note-1" is in its included list: read and execute gsd-core/workflows/review/steps/reviewer-instances-note-1.md. Otherwise skip — do not read the file.

If no CLIs are available:

No external AI CLIs found. Install at least one:
- gemini: https://github.com/google-gemini/gemini-cli
- codex: https://github.com/openai/codex
- claude: https://github.com/anthropics/claude-code
- opencode: https://opencode.ai (leverages GitHub Copilot subscription models)
- qwen: https://github.com/nicepkg/qwen-code (Alibaba Qwen models)
- cursor: https://cursor.com (Cursor IDE agent mode)
- agy: curl -fsSL https://antigravity.google/cli/install.sh | bash (Antigravity CLI — free with Google credentials)

Then run /gsd:review again.

Exit.

Determine which CLI to skip based on the current runtime environment:

# Environment-based runtime detection (priority order)
if [ "$ANTIGRAVITY_AGENT" = "1" ]; then
  # Antigravity is a separate client — all CLIs are external, skip none
  SELF_CLI="none"
elif [ -n "$CURSOR_SESSION_ID" ]; then
  # Running inside Cursor agent — skip cursor for independence
  SELF_CLI="cursor"
elif [ -n "$CLAUDE_CODE_ENTRYPOINT" ]; then
  # Running inside Claude Code CLI — skip claude for independence
  SELF_CLI="claude"
else
  # Other environments (Gemini CLI, Codex CLI, etc.)
  # Fall back to AI self-identification to decide which CLI to skip
  SELF_CLI="auto"
fi

Rules:

  • If SELF_CLI="none" → invoke ALL available CLIs (no skip)
  • If SELF_CLI="claude" → skip claude, use gemini/codex
  • If SELF_CLI="auto" → the executing AI identifies itself and skips its own CLI
  • At least one DIFFERENT CLI must be available for the review to proceed.
Collect phase artifacts for the review prompt:
INIT=$(gsd_run query init.review "${PHASE_ARG}")
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi

# #2358: ONE run-scoped temp dir (portable via ${TMPDIR:-/tmp}) so overlapping
# runs never collide or read each other's stale files.
RUN_DIR=$(mktemp -d "${TMPDIR:-/tmp}/gsd-review-XXXXXX")
echo "RUN_DIR=$RUN_DIR"

Read from init: phase_dir, phase_number, padded_phase.

Capture RUN_DIR above (created ONCE) and thread it into every {run_dir} placeholder and $RUN_DIR/${RUN_DIR} reference within a bash block. Do NOT re-run mktemp -d later — every block must resolve to this same directory, or build_prompt's writes and invoke_reviewers' reads split.

Then read:

  1. .planning/PROJECT.md (first 80 lines — project context)
  2. Phase section from .planning/ROADMAP.md
  3. All *-PLAN.md files in the phase directory
  4. *-CONTEXT.md if present (user decisions)
  5. *-RESEARCH.md if present (domain research)
  6. .planning/REQUIREMENTS.md (requirements this phase addresses)
Build a structured review prompt:
# Cross-AI Plan Review Request

You are reviewing implementation plans for a software project phase.
Provide structured feedback on plan quality, completeness, and risks.

## Project Context
{first 80 lines of PROJECT.md}

## Phase {N}: {phase name}
### Roadmap Section
{roadmap phase section}

### Requirements Addressed
{requirements for this phase}

### User Decisions (CONTEXT.md)
{context if present}

### Research Findings
{research if present}

### Plans to Review
{all PLAN.md contents}

## Review Instructions

**Verify against source — do not review the plan text in isolation.** The plans reference real files, migrations, routes, and tests in this repo.
1. Open the referenced files and check each claim against the actual code.
2. For every strength or concern, cite concrete `path/to/file:line` evidence plus the mechanism.
3. When a plan asserts a mechanism works (a guard, a query filter, a test that exercises a path), trace whether it actually does what is claimed — do not take the plan's word for it.
4. If you cannot read the repo (no file access), say so and downgrade that finding to an open question rather than asserting it.

Findings citing `file:line` evidence are weighted far more heavily than impressionistic ones; a review that only restates the plan's own claims has low value.

Analyze each plan and provide:

1. **Summary** — One-paragraph assessment
2. **Strengths** — What's well-designed (bullet points)
3. **Concerns** — Potential issues, gaps, risks (bullet points with severity: HIGH/MEDIUM/LOW)
4. **Suggestions** — Specific improvements (bullet points)
5. **Risk Assessment** — Overall risk level (LOW/MEDIUM/HIGH) with justification

Focus on:
- Missing edge cases or error handling
- Dependency ordering issues
- Scope creep or over-engineering
- Security considerations
- Performance implications
- Whether the plans actually achieve the phase goals

Output your review in markdown format.

Write to a temp file: {run_dir}/gsd-review-prompt.md

Also write individual section files so the budget tool can re-trim per reviewer:

# #2962: zsh aborts the block on an unmatched for-list glob (nomatch); bash passes it through. nullglob both.
shopt -s nullglob 2>/dev/null; setopt NULL_GLOB 2>/dev/null

RUN_DIR="{run_dir}"   # from gather_context

# Write individual section files for per-reviewer budget trimming
# These are always written so reviewers with a budget can invoke prompt-budget
cp "$INSTRUCTIONS_BLOCK_FILE" "${RUN_DIR}/gsd-review-instructions.md"
cp "$ROADMAP_SECTION_FILE" "${RUN_DIR}/gsd-review-roadmap.md"

# Plan files: copy each PLAN.md to a predictable numbered path
PLAN_INDEX=0
for PLAN_FILE in "${PHASE_DIR}"/*-PLAN.md; do
  PADDED_IDX=$(printf '%02d' "$PLAN_INDEX")
  cp "$PLAN_FILE" "${RUN_DIR}/gsd-review-plan-${PADDED_IDX}.md"
  PLAN_INDEX=$((PLAN_INDEX + 1))
done

# Optional section files (only if content was included in the combined prompt)
if [ -f ".planning/PROJECT.md" ]; then
  cp .planning/PROJECT.md "${RUN_DIR}/gsd-review-project.md"
fi
_CTX=( "${PHASE_DIR}"/*-CONTEXT.md )
if [ ${#_CTX[@]} -gt 0 ]; then
  cat "${_CTX[@]}" > "${RUN_DIR}/gsd-review-context.md"
fi
_RESEARCH=( "${PHASE_DIR}"/*-RESEARCH.md )
if [ ${#_RESEARCH[@]} -gt 0 ]; then
  cat "${_RESEARCH[@]}" > "${RUN_DIR}/gsd-review-research.md"
fi
if [ -f ".planning/REQUIREMENTS.md" ]; then
  cp .planning/REQUIREMENTS.md "${RUN_DIR}/gsd-review-requirements.md"
fi

Note: INSTRUCTIONS_BLOCK_FILE, ROADMAP_SECTION_FILE, and PHASE_DIR come from prompt assembly; RUN_DIR is the run-scoped dir from gather_context (#2358) re-assigned from {run_dir} above. Copy the temp files written during prompt assembly to these section paths (or write each section here if the prompt was built inline).

Every reviewer lane is **declared data** (ADR-2782). This step iterates the lanes the selection resolved; it does not enumerate them. Adding a reviewer is a capability manifest, not an edit here.

Do not re-add a per-CLI block. A <!-- reviewer-lane: … --> marker anywhere in this step now FAILS the parity gate (checkReviewerLaneParity → bespoke_leg_present). Lane divergence is declared in the manifest — timeout floor, probe, prompt/output channel, empty-output policy — and behaviour that data genuinely cannot express is a named first-party handler (ADR-2782 D6), never a bespoke block here.

Timeout guidance (#2194): prompt-fed source-grounded reviews are slow — measured ~570 s for Codex at xhigh effort and ~525 s for headless Claude on a large plan set. Each lane declares its own timeoutFloorMs and the runner enforces it internally, but the Bash tool call wrapping the loop below must still be given a high timeout: — at least 900000, and 1200000 when Codex or headless Claude are in the selection — or the host kills the whole loop mid-lane. On Claude Code, raise the host cap via BASH_MAX_TIMEOUT_MS if a review can exceed it.

A silent empty output after a long run is a timeout kill, not a crash — the Codex 0xc0000142 misdiagnosis persisted for exactly this reason, because an empty result cannot distinguish the two on its own. Treat an empty result on a slow lane as a dropped lane and re-run with more time rather than diagnosing a CLI or sandbox failure. A cross-AI review that silently drops a lane is blind in one eye.

No hook-trust bypass (#2479): no lane passes a hook-trust bypass flag and none runs a capability probe for one. That flag only bypasses persisted hook trust (a first-run condition) and flagless invocations work in steady state, while host-harness safety classifiers deny commands carrying it. An environment that genuinely hits an untrusted-hook prompt surfaces through the .err capture and the empty-output stub as a dropped lane with diagnosable stderr, not silent attrition. Do not reintroduce the flag (even spelled out in prose — a regression test bans the literal file-wide).

If section_manifest is null or "reviewer-instances-note-2" is in its included list: read and execute gsd-core/workflows/review/steps/reviewer-instances-note-2.md. Otherwise skip — do not read the file.

Lanes run sequentially by default — concurrent invocation trips provider rate limits, and a lane lost to one is a cross-AI review that quietly went blind in one eye. A project whose providers can accept the concurrency opts in with review.parallel_lanes: true (#3034): the selected lanes are dispatched together and all joined before aggregation. The default is unchanged, and convergence cycles stay sequential either way — only the lanes within one pass overlap.

# #2962: zsh aborts the block on an unmatched for-list glob (nomatch); bash passes it through. nullglob both.
shopt -s nullglob 2>/dev/null; setopt NULL_GLOB 2>/dev/null

RUN_DIR="{run_dir}"
REPO_ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
# SELECTED_REVIEWERS is the comma-separated result of reviewer selection (ADR-0011 precedence:
# explicit flags > --all > review.default_reviewers > all detected). Unchanged by this phase.

# #3034: opt-in concurrent lane dispatch. STRICT equality on "true" is deliberate — "1", "yes" and
# "TRUE" must NOT opt in, so a mistyped config gets the conservative behaviour rather than firing
# concurrent requests at a rate-limited provider. Note the `|| echo "false"` fallback is the
# OPPOSITE polarity from the commit_docs guard, which fails OPEN: there, failing open preserves the
# user's intent; here it would fire exactly the requests the default exists to prevent.
PARALLEL_LANES=$(gsd_run query config-get review.parallel_lanes --raw 2>/dev/null || echo "false")

# Shared budget-trim helper. Was defined inside the Ollama leg; it is lane-agnostic, so it is
# hoisted here now that any lane may declare a promptBudgetKey. Returns non-zero when the budget
# is too small for the minimum review set (prompt-budget exit 2 / 11).
prepare_trimmed_prompt_for_reviewer() {
  REVIEWER_KEY="$1"; REVIEWER_BUDGET="$2"; OUTPUT_PROMPT="$3"; OUTPUT_META="$4"

  PLAN_FILE_ARGS=""
  for p in "$RUN_DIR"/gsd-review-plan-*.md; do
    [ -f "$p" ] && PLAN_FILE_ARGS="$PLAN_FILE_ARGS --plan-file $p"
  done
  PROJECT_ARG=""
  [ -f "$RUN_DIR/gsd-review-project.md" ] && PROJECT_ARG="--project-file $RUN_DIR/gsd-review-project.md"
  CONTEXT_ARG=""
  [ -f "$RUN_DIR/gsd-review-context.md" ] && CONTEXT_ARG="--context-file $RUN_DIR/gsd-review-context.md"
  RESEARCH_ARG=""
  [ -f "$RUN_DIR/gsd-review-research.md" ] && RESEARCH_ARG="--research-file $RUN_DIR/gsd-review-research.md"
  REQUIREMENTS_ARG=""
  [ -f "$RUN_DIR/gsd-review-requirements.md" ] && REQUIREMENTS_ARG="--requirements-file $RUN_DIR/gsd-review-requirements.md"

  gsd_run query prompt-budget \
    --budget "$REVIEWER_BUDGET" \
    --instructions-file "$RUN_DIR/gsd-review-instructions.md" \
    --roadmap-file "$RUN_DIR/gsd-review-roadmap.md" \
    $PLAN_FILE_ARGS $PROJECT_ARG $CONTEXT_ARG $RESEARCH_ARG $REQUIREMENTS_ARG \
    --output-prompt "$OUTPUT_PROMPT" \
    --output-metadata "$OUTPUT_META"
  return $?
}

gsd_run query review-lane plan \
  --selected "$SELECTED_REVIEWERS" --run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" --json \
  > "$RUN_DIR/gsd-review-lanes.json"

# One lane, start to finish. Hoisted into a function so the sequential and concurrent paths share
# ONE body: two dispatch bodies kept in sync by hand is the generative-fix divergence ADR-2782 spent
# a phase deleting, and it is what let #2494/#2605 be filed twice as the same defect.
#
# The result goes to a SLUG-SCOPED file, never a shared append. Concurrent O_APPEND is atomic only
# below PIPE_BUF (4096 on Linux, 512 on some platforms), so a lane result above that bound could
# interleave — and write_reviews parses this JSONL to render the models:/model_sources: frontmatter,
# so a torn line is a broken REVIEWS.md, not a cosmetic log defect.
run_review_lane() {
  # `local` is hygiene, not a live fix: each `&`-dispatched call already forks its own subshell, so
  # concurrent lanes cannot share these today. Scoped anyway so the isolation is a property of this
  # function rather than of the dispatch mechanism happening to fork.
  local SLUG LANE_BUDGET PROMPT_ARG TRIMMED
  SLUG="$1"
  # Per-lane prompt budget. The lane declares its own `promptBudgetKey`; `plan` resolved it,
  # applying #2797's sentinel rule (-1 = unset → fall back to the global budget; 0 legitimately
  # means "do not trim this lane"). Trimming itself stays in prompt-budget, which owns it.
  LANE_BUDGET=$(gsd_run query review-lane plan --selected "$SLUG" --run-dir "$RUN_DIR" \
                  --repo-root "$REPO_ROOT" --json 2>/dev/null \
                | sed -n 's/.*"promptBudget": *\([0-9-]*\).*/\1/p' | head -1)
  PROMPT_ARG=""
  if [ -n "$LANE_BUDGET" ] && [ "$LANE_BUDGET" != "null" ] && [ "$LANE_BUDGET" -gt 0 ] 2>/dev/null; then
    TRIMMED="$RUN_DIR/gsd-review-prompt-$SLUG.md"
    if prepare_trimmed_prompt_for_reviewer "$SLUG" "$LANE_BUDGET" "$TRIMMED" \
         "$RUN_DIR/gsd-review-prompt-$SLUG.metadata.json"; then
      PROMPT_ARG="--prompt-file $TRIMMED"
    else
      # A budget too small for the minimum review set drops the lane just as silently as an empty
      # response used to (#2605), so leave the skip visible in the review output, not only on stderr.
      echo "$SLUG review skipped: prompt budget (${LANE_BUDGET} tokens) too small for the minimum review set." \
        > "$RUN_DIR/gsd-review-$SLUG.md"
      # Was `continue` when this was a loop body. Inside a function that keyword is not the loop
      # control it looks like — `return 0` is what skips this lane and leaves it with no result line.
      return 0
    fi
  fi

  # One invocation, whatever the lane's transport, prompt channel, output channel or handler.
  # `--explicit` marks a lane the user NAMED: ADR-2782 D4 — not finding a lane nobody asked for is
  # normal, failing to run one somebody asked for is an error.
  gsd_run query review-lane invoke --slug "$SLUG" \
    --run-dir "$RUN_DIR" --repo-root "$REPO_ROOT" $PROMPT_ARG $EXPLICIT_FLAG --json \
    > "$RUN_DIR/gsd-review-lane-result-$SLUG.json"
}

# Split ONCE, de-duplicated, and reuse for both loops below. Two reasons, and the second is
# load-bearing: a slug repeated in SELECTED_REVIEWERS would put TWO concurrent background jobs on
# `> "$RUN_DIR/gsd-review-lane-result-$SLUG.json"` — the same file, both truncating. The shared-append
# form this replaced could not corrupt itself that way, so de-duping is what keeps the concurrent
# path no worse than the sequential one. Selection de-dupes today (the roster is a Set;
# review.default_reviewers normalizes lowercase-unique), but reachability analysis is not a contract
# and the next caller should not have to redo it.
#
# A plain string accumulator, not an array: zsh and bash disagree on array indexing and this block
# runs under both (see the nullglob/NULL_GLOB pairing above).
DISPATCH_SLUGS=""
for SLUG in $(echo "$SELECTED_REVIEWERS" | tr ',' ' '); do
  case " $DISPATCH_SLUGS " in
    *" $SLUG "*) continue ;;
  esac
  DISPATCH_SLUGS="$DISPATCH_SLUGS $SLUG"
done

for SLUG in $DISPATCH_SLUGS; do
  if [ "$PARALLEL_LANES" = "true" ]; then
    run_review_lane "$SLUG" &
  else
    run_review_lane "$SLUG"
  fi
done

# Join every dispatched lane. A bare `wait` with no background jobs returns 0, so the sequential
# path needs no guard around it. NOTHING below this line may run before every lane has finished —
# write_reviews renders REVIEWS.md and the consensus summary from the aggregate below, and a review
# assembled from a partial set looks complete while silently missing a reviewer.
wait

# Aggregate in SELECTED_REVIEWERS order, NOT completion order, so the JSONL a concurrent run
# produces is byte-identical to the one a sequential run produces. This is post-join and therefore
# single-threaded, so `>>` here is safe. A lane that was budget-skipped, or that never started,
# leaves no result file and correctly contributes no line.
for SLUG in $DISPATCH_SLUGS; do
  LANE_RESULT="$RUN_DIR/gsd-review-lane-result-$SLUG.json"
  if [ -f "$LANE_RESULT" ]; then
    cat "$LANE_RESULT" >> "$RUN_DIR/gsd-review-lane-results.jsonl"
  fi
done

Each lane leaves {run_dir}/gsd-review-<slug>.md — its review, or a diagnostic stub carrying the captured stderr (and, for an OpenAI-compatible lane, the raw response body, where such a server puts its error JSON on an HTTP 4xx/5xx while still exiting 0). A stub is never mistaken for a clean review: it keeps its "failed or returned empty output" header (#2494/#2605/#2794).

A lane that will not run reports a typed reason rather than an empty file — missing_binary, probe_failed, probe_timeout, missing_required_binary, host_unreachable, egress_host_changed, unknown_handler, budget_too_small. egress_host_changed means the lane was consented to send plans to one destination and .planning/config.json now names another; it is blocked, not silently redirected (ADR-2782 D5).

Display progress:

### GSD ► CROSS-AI REVIEW — Phase {N}

◆ Reviewing with {CLI}... done ✓
◆ Reviewing with {CLI}... done ✓
**#3352 (ADR-3473 §8.5): no artifact from failed inputs.** Before rendering anything, gate on whether any lane actually produced a result — "every lane failed" is exactly "the aggregate JSONL has zero lines" (§`invoke_reviewers`'s aggregation loop already builds this file as a byproduct; a lane that never started or was budget-skipped contributes no line either way).
RUN_DIR="{run_dir}"
JSONL="$RUN_DIR/gsd-review-lane-results.jsonl"
LANE_LINES=0
[ -f "$JSONL" ] && LANE_LINES=$(wc -l < "$JSONL" | tr -d ' ')

TOTAL_LANE_FAILURE="false"
ALL_LANES_SKIPPED="false"
if [ "${LANE_LINES:-0}" -eq 0 ]; then
  # Zero lines means every dispatched lane left no result JSON — either every one
  # was budget-skipped (N5: a skip is not a failure) or every one actually failed
  # to run. Re-derive the dispatched-slug set the same way invoke_reviewers did
  # (SELECTED_REVIEWERS is a shell block boundary — recompute, do not assume the
  # earlier step's local DISPATCH_SLUGS variable survived into this block).
  DISPATCH_SLUGS=""
  for SLUG in $(echo "$SELECTED_REVIEWERS" | tr ',' ' '); do
    case " $DISPATCH_SLUGS " in
      *" $SLUG "*) continue ;;
    esac
    DISPATCH_SLUGS="$DISPATCH_SLUGS $SLUG"
  done
  # Distinguish by whether every dispatched slug's stub markdown says "skipped":
  # a skip stub always does (see run_review_lane's budget branch, which writes
  # this exact text before returning without ever invoking the lane); a real
  # failure stub does not. If a slug has no stub at all, it is not a skip.
  DISPATCHED_COUNT=0
  SKIPPED_COUNT=0
  for SLUG in $DISPATCH_SLUGS; do
    DISPATCHED_COUNT=$((DISPATCHED_COUNT + 1))
    STUB="$RUN_DIR/gsd-review-$SLUG.md"
    if [ -f "$STUB" ] && grep -q "review skipped: prompt budget" "$STUB" 2>/dev/null; then
      SKIPPED_COUNT=$((SKIPPED_COUNT + 1))
    fi
  done
  if [ "$DISPATCHED_COUNT" -gt 0 ] && [ "$SKIPPED_COUNT" -eq "$DISPATCHED_COUNT" ]; then
    ALL_LANES_SKIPPED="true"
  else
    TOTAL_LANE_FAILURE="true"
  fi
fi
  • If ALL_LANES_SKIPPED=true: do NOT write REVIEWS.md and do NOT run the commit below — there is nothing to review. Report to the user that every selected lane was budget-skipped (not a failure) and stop; do not proceed to present_results' summary claiming a review ran.
  • If TOTAL_LANE_FAILURE=true: do NOT write REVIEWS.md and do NOT run the commit below. Report the total lane failure to the user (name the lanes that were dispatched and point at their .err/stub files preserved under .review-diagnostics/ by present_results) and stop.
  • Otherwise (at least one lane produced a result — R1, unchanged): proceed exactly as below.

Combine all review responses into {phase_dir}/{padded_phase}-REVIEWS.md:

After all reviewers complete, collect trim metadata files written during the run. For each reviewer that was trimmed (i.e. a .metadata.json file exists and hardFailed or omitted is non-empty, or projectMdShrunk is true, or planTruncationPct > 0), include a trimmed_reviewers block in the frontmatter. Omit the key entirely if no reviewer was trimmed.

Reviewer instances (#1517, optional): when instances ran, frontmatter records their names, each gets its own ## <Adapter> Review (<instance>) section, and ≥2 same-cli instances print a one-line shared-adapter caveat. Format in gsd-core/references/reviewer-instances.md.

Resolved model (#2295): each lane's review-lane invoke --json line in {run_dir}/gsd-review-lane-results.jsonl carries a model object; render models: from its value and model_sources: from its source. Both maps carry exactly one entry per reviewer that appears in reviewers: — the two key sets always match. Write the literal unknown rather than omitting a key: an omitted key is indistinguishable from the feature not having run, and a reader must be able to tell no model recorded from nothing to look at. Emit every models:/model_sources: value as a DOUBLE-QUOTED YAML scalar — a legitimate model id can contain : (llama3:70b, qwen2.5:7b), which is unquotable as a bare scalar; a control character is already refused at the recording seam, so quoting is what closes the remaining :/#/leading-- cases. When GSD applied a reasoning effort to a lane, its value already carries a (reasoning=<level>) suffix (e.g. gpt-5.6-sol (reasoning=high)) — render it as-is, without re-deriving or re-formatting it.

---
phase: {N}
reviewers: [gemini, claude, codex, coderabbit, opencode, qwen, cursor, antigravity, ollama, lm_studio, llama_cpp]  # populate at runtime with only the reviewers actually invoked
reviewed_at: {ISO timestamp}
plans_reviewed: [{list of PLAN.md files}]
models:                   # resolved model per reviewer; `unknown` when not recoverable
  codex: "gpt-5.6-sol (reasoning=low)"
  antigravity: "unknown"
model_sources:            # how each value above was determined
  codex: "banner"
  antigravity: "unknown"
trimmed_reviewers:        # only present if at least one reviewer was trimmed
  ollama:
    budget: 6000
    effective_budget: 5400
    estimated_tokens: 5380
    omitted: [context, research]
    project_md_shrunk: true
    plan_truncation_pct: 22
    hard_failed: false
    note_injected: true
---

# Cross-AI Plan Review — Phase {N}

<!-- Sections are RENDERED from each lane's declared `reviewsSection`, in descriptor order.
     There is deliberately no hardcoded per-reviewer heading list here any more: a hand-maintained
     list is exactly the drift #2781 was filed about, and it silently disagreed with the roster.
     `gsd_run query review-lane sections --selected "$SELECTED_REVIEWERS"` emits
     `<slug><TAB><reviewsSection>` in order; for each row, emit:

         ## <reviewsSection> Review

         {contents of {run_dir}/gsd-review-<slug>.md}

         ---

     Two headings must NOT be generated from this list, because they are not lanes:
       * `## <Adapter> Review (<instance>)` — an ADR-1517 reviewer INSTANCE resolves THROUGH a lane
         and is rendered from the instance list, not the lane list (ADR-2782 D8).
       * `## Consensus Summary` — not a review section at all.

     A lane whose `evidenceClass` is `diff-only` (CodeRabbit) carries its caveat from data: it never
     received the source-grounding prompt, so its verdict is folded in as a diff observation and is
     not weighted as a grounded plan review. -->

## Consensus Summary

{synthesize common concerns across all reviewers. CodeRabbit is a diff-only reviewer (it never received the source-grounding prompt), so do not weight its verdict as a grounded plan review — fold in its diff findings, but base plan-level consensus on the prompt-fed reviewers. A reviewer output carrying the `[reviewed-without-repo-access]` marker (or beginning with `REVIEWED-WITHOUT-REPO-ACCESS`) ran without repo access (#2176) — treat it the same way: note its concerns, but do not count its verdict at full consensus weight. A reviewer output carrying the `[reviewed-without-source-citations]` marker (#3194) declared source-grounded evidence but cited no `file:line` evidence, so it reviewed the plan text only — treat it the same way: note its concerns, but do not count its verdict at full consensus weight.}

### Agreed Strengths
{strengths mentioned by 2+ reviewers}

### Agreed Concerns
{concerns raised by 2+ reviewers — highest priority}

### Divergent Views
{where reviewers disagreed — worth investigating}

Commit (only reached when TOTAL_LANE_FAILURE and ALL_LANES_SKIPPED are both false — the gate above):

gsd_run query commit "docs: cross-AI review for phase {N}" --files {phase_dir}/{padded_phase}-REVIEWS.md
**If `write_reviews` set `TOTAL_LANE_FAILURE=true` or `ALL_LANES_SKIPPED=true`, skip the success summary below entirely** — no `REVIEWS.md` was written or committed, so there is nothing to present as complete. Report instead:
### GSD ► REVIEW FAILED

Phase {N}: every selected reviewer lane {failed to produce a result|was budget-skipped} — no
REVIEWS.md was written.

{If the preserve+cleanup block below reports success: "Diagnostics preserved:
{phase_dir}/.review-diagnostics/". If it reports failure: relay its own warning verbatim —
it names the intact run directory holding the un-preserved evidence instead.}

Otherwise (at least one lane succeeded), display summary:

### GSD ► REVIEW COMPLETE

Phase {N} reviewed by {count} AI systems.

Consensus concerns:
{top 3 shared concerns}

Full review: {padded_phase}-REVIEWS.md

To incorporate feedback into planning:
  /gsd:plan-phase {N} --reviews

#3352 (ADR-3473 §8.5, R3): preserve per-lane evidence before destroying it. Regardless of which branch above ran, the run's temp directory is the only record that a lane failed at all — copy it beside the phase's artifacts BEFORE cleanup. A lane that produced no output at all (L4) leaves nothing to preserve; that is a smaller diagnostics folder, not a fabricated one, and is NOT a preservation failure. This copy is deliberately NOT part of the commit above (N6) — that step names only {padded_phase}-REVIEWS.md explicitly, never a directory glob, so .review-diagnostics/ is never swept into it.

Preservation and cleanup MUST run in the same fenced block below (a shell variable cannot survive across separate fences — each is its own process). mkdir -p and every cp are exit-status checked; rm -rf "$RUN_DIR" runs ONLY if nothing was preserved (nothing to preserve is not a failure) or everything that needed preserving was copied successfully. If preservation fails partway, $RUN_DIR is left intact and a message names it as the location of the un-preserved evidence — a leftover temp directory is far cheaper than destroyed evidence:

shopt -s nullglob 2>/dev/null; setopt NULL_GLOB 2>/dev/null

RUN_DIR="{run_dir}"
DIAG_DIR="{phase_dir}/.review-diagnostics"

_DIAG_MD=( "$RUN_DIR"/gsd-review-*.md )
_DIAG_ERR=()
for f in "$RUN_DIR"/gsd-review-*.err; do
  [ -s "$f" ] && _DIAG_ERR+=("$f")
done

_PRESERVE_OK=true
if [ ${#_DIAG_MD[@]} -gt 0 ] || [ ${#_DIAG_ERR[@]} -gt 0 ]; then
  if mkdir -p "$DIAG_DIR"; then
    if [ ${#_DIAG_MD[@]} -gt 0 ] && ! cp "${_DIAG_MD[@]}" "$DIAG_DIR/"; then
      _PRESERVE_OK=false
    fi
    if [ ${#_DIAG_ERR[@]} -gt 0 ] && ! cp "${_DIAG_ERR[@]}" "$DIAG_DIR/"; then
      _PRESERVE_OK=false
    fi
  else
    _PRESERVE_OK=false
  fi
fi

if [ "$_PRESERVE_OK" = "true" ]; then
  rm -rf "$RUN_DIR"
else
  echo "WARNING: evidence preservation to $DIAG_DIR failed — leaving the un-preserved run directory intact at: $RUN_DIR" >&2
fi

<success_criteria>

  • At least one external CLI invoked successfully
  • REVIEWS.md written with structured feedback
  • Consensus summary synthesized from multiple reviewers
  • Temp files cleaned up
  • User knows how to use feedback (/gsd:plan-phase --reviews) </success_criteria>