Files
msd-core/docs/how-to/set-up-cross-ai-review.md
Tom Boucher b54c1c5848 fix(#4709): retire the Gemini CLI reviewer lane (#4716)
* fix(#4709): retire the Gemini CLI reviewer lane

Google stopped serving Gemini CLI for the free/Pro/Ultra tiers on 2026-06-18 —
the same sunset that removed the gemini RUNTIME in #1928 (shipped 1.8.0). GSD
targets solo developers, so those tiers ARE the user path: the lane spawned
`gemini {{model}} -p -`, a binary that no longer answers for the majority of
users, and five locales documented it as a supported choice.

The lane was re-created after #1928 by the reviewer-lane-as-manifest-data work
(6a9babda69, #2798/#2837, ADR-2782). Per the maintainer that re-creation was an
error in that buildout rather than a considered decision, so this corrects a
mistake and needs no ADR-2782 amendment.

Reviewer roster: 12 lanes / 13 flags -> 11 lanes / 12 flags.

TWO sources of truth had to be removed, not one. Deleting
capabilities/gemini/capability.json left the capability registry at 11 lanes
while src/review-lane-descriptor.cts's hand-maintained REVIEWER_LANES array
still carried its own complete gemini entry at 12 — precisely the disagreement
checkReviewerLaneParity exists to catch. Both are gone; both parity checkers
now run clean against the real tree (lane parity ok/0 violations, docs parity
0 violations).

Surfaces stripped of the dead flag:
- capabilities/gemini/ deleted; registry and capability-matrix regenerated
- src/review-lane-descriptor.cts: REVIEWER_LANES entry, docblock count, and the
  three doc comments that used --gemini as a live example
- commands/gsd/{review,plan-review-convergence,autonomous,progress}.md and the
  four matching skills/*/SKILL.md: argument-hint frontmatter and flag bullets
- gsd-core/workflows/help/modes/{full,full.compact}.md: /gsd-help signatures,
  the detected-CLI list, and the reviewer-title list
- gsd-core/workflows/settings-integrations.md: the integrations wizard no longer
  offers "Gemini" as a model option, and the settable-keys list drops it
- gsd-core/workflows/review.md: the `command -v gemini` probe, the --gemini
  flag, the roster frontmatter, the install pointer to the sunset repo, and the
  jq-less / precedence / self-skip lane lists
- gsd-core/workflows/sync-skills.md: "two runtimes (grok, gemini) resolve to
  ANOTHER runtime's skills root" is now one runtime; gemini never aliased
  anything, it fell through canonicalizeRuntimeName to a fail-closed default
- docs/{CONFIGURATION,COMMANDS,CLI-TOOLS}.md, docs/reference/capability-matrix.md,
  docs/how-to/set-up-cross-ai-review.md — including its `npm install -g
  @google/gemini-cli` instruction and the two rows recommending --gemini
- docs/features/{cross-ai-peer-review,opt-in-parallel-reviewer-lanes}.md as the
  generator inputs behind docs/FEATURES.md, plus the three locale FEATURES.md
  signature lines the docs-parity gate covers (the #2781 class: a flag change
  that never reaches the mirrors)

Counts reconciled against measurement rather than arithmetic: 8 timeout keys of
11 lanes, 11 budget keys, 9 model keys, and four hardcoded literals in
tests/reviewer-lane-declarations.test.cjs (NEW_LANE_ONLY_IDS 5->4, LITERAL_ROSTER
12->11, two roster counts 12->11).

BEHAVIOR CHANGE, accepted deliberately: `gsd config-set review.models.gemini`
now errors with "Unknown config key". An existing key already in
.planning/config.json still parses and is simply never read, so no project fails
to load. This is the repo's own documented policy for exactly this case
(docs/CONFIGURATION.md:327 — "a key left over from a removed reviewer validated
silently and was never read. Such a key is now rejected by config-set"), so no
installer migration ships. Note my first measurement of this was WRONG: I tested
config-get, which reads undeclared keys fine, and generalised. Read and write are
different surfaces and gave different answers.

Antigravity is untouched throughout — its --antigravity/--agy flags,
review.models.agy, ~/.gemini/antigravity configHome, ~/.gemini/config global
skills root (#3738), hookEvents "gemini", GEMINI.md instruction file, and every
gemini-* model id it actually runs on.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): changeset for the reviewer-lane retirement

Type Removed: the --gemini flag and its three config keys are user-visible
surface that no longer exists.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): close the 24 test failures and the locale-doc gap the gates found

An adversarial review and a full matrix run between them found substantially
more fallout than inspection had. All of it is this PR's own, and all of it is
fixed rather than waved off.

THE MATRIX RUN FOUND 24 FAILURES ACROSS 6 FILES. Inspection had predicted two.
The dominant class was a test helper that looks up a lane by slug and throws
`no declared lane 'gemini'`:

- tests/feat-2483-review-claude-mds-guard.test.cjs (6) — used gemini as the
  "other declared first-party lane" to contrast against claude's env
  suppression. Now qwen, verified from source as a lane that declares no `env`
  (only claude does), so the contrast still holds.
- tests/review-lane-descriptor.test.cjs (6) — the duplicate-flag and
  duplicate-section fixtures deliberately COLLIDED with a real declared lane to
  prove the parity checker reports a duplicate. `--gemini`/`Gemini` no longer
  collide with anything, so the checker reported
  `descriptor_lane_not_in_registry:acme` instead and the tests proved nothing.
  Now collide with `--codex`/`Codex`, reproduced against the real checker.
- tests/review-reviewer-selection.test.cjs (3) — these distinguish KNOWN-but-
  undetected from UNKNOWN. gemini flipped categories, inverting what they
  proved. The known case now uses qwen; `__nope__` stays the unknown fixture.
- tests/review-default-reviewers-resolution.test.cjs (2), and
  tests/settings-integrations.test.cjs (3) — the wizard now offers three
  reviewer CLIs, not four, so the test and its name say three.
- Two count assertions the earlier sweep missed outright:
  reviewer-lane-declarations.test.cjs:359 (`length, 12`) and
  reviewer-docs-parity.test.cjs:681 (`>= 12`).

THE LOCALE-DOC GAP, and why the parity gate stayed green over it. All four
locale mirrors still documented `--gemini` as a live reviewer flag. The
docs-parity checker asserts the PRESENCE of every current flag and never the
ABSENCE of a retired one, so "0 violations" was never evidence those files were
clean — my earlier reading of it as such was wrong. This is the #2781
locale-drift class in the opposite direction. Fixed across 12 locale files:
COMMANDS.md flag lists and table rows, CONFIGURATION.md `review.models.gemini`
rows and reviewer prose, CLI-TOOLS.md config examples, and
set-up-cross-ai-review.md including its install block and its
which-reviewer-to-choose row, which now recommends Antigravity.

ALSO FOUND, and instructive about my own method: docs/CONFIGURATION.md:297 still
carried a `review.models.gemini` row. My sweep had missed it because my grep
excluded lines matching `gemini-[0-9]` to spare Google's model ids — and that
row's example value is `"gemini-2.5-pro"` on the same line. The exclusion built
to avoid false positives created a false negative.

Remaining comment/example sites: src/review-reviewer-selection.cts:309 and
src/config.cts:598 named the dead flag and key as examples;
gsd-core/references/planning-config.md:269 likewise; and
review-reviewer-selection.cts:22 claimed in the PRESENT tense that gemini is a
lane-only reviewer capability. Line 38 of that same docblock says "Before this
phase the five non-runtime reviewers (gemini, ...)" and is left exactly as is —
that is past-tense history, and rewriting it would falsify the record.

Deliberately still deferred to Phase 4, because it is the RUNTIME axis rather
than the reviewer lane: the locale install-on-your-runtime.md `--gemini --global`
instructions, the USER-GUIDE colon-form notes, and the ARCHITECTURE
runtime-detection flag lists.

Both parity checkers green against the real tree; lint:ci exit 0.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): backfill the changeset PR number

pr: 0 -> 4716, now that the PR exists. Never guessed ahead of the number.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 03:03:44 -04:00

7.2 KiB
Raw Blame History

How to set up cross-AI review

Goal: Configure which AI reviewers participate in plan review, run a review of a planned phase, and use the feedback to converge on a plan with no HIGH-severity concerns.

Prerequisites: The phase has been planned ({phase}-PLAN.md files exist in .planning/phases/). At least one external AI CLI is installed and authenticated.


Decide which reviewers to use

GSD Core can route review requests to any combination of: Claude (separate session), Codex CLI, CodeRabbit, OpenCode, Qwen Code, Cursor, Antigravity CLI, Ollama, LM Studio, llama.cpp, and Kimi Code.

That list is not fixed. Each of those is a declared reviewer lane, and a capability can ship its own — see Ship a reviewer lane in your capability. To see exactly which lanes your installation has, run gsd-tools review-lane sections.

Each reviewer runs the same structured prompt against your PLAN.md files independently. Because different models have different blind spots, multi-reviewer consensus catches more issues than any single reviewer.

If you have no external CLIs installed yet, install at least one:

# Antigravity CLI (free with Google credentials)
curl -fsSL https://antigravity.google/cli/install.sh | bash

# Codex CLI
npm install -g @openai/codex

Set default reviewers (optional)

By default, /gsd-review runs all detected CLIs. To pin a subset as project defaults:

/gsd-config --integrations

The integrations wizard covers API keys, code-review CLI routing, and the review.default_reviewers list. Set the list to the reviewers you want as the no-flag default — for example ["codex","claude"].

Alternatively, set it directly with gsd-tools:

gsd config-set review.default_reviewers '["codex","claude"]'

For the full integration settings schema (API keys, model overrides per reviewer, local server host addresses), see Configuration.

If you need multiple independent reviewer voices from the same adapter, configure review.reviewer_instances and add those instance names to review.default_reviewers. Instance names run only through review.default_reviewers; they are not valid /gsd-review flags. See Reviewer instances for the schema.


Run a review

Standard review (uses your configured defaults or all detected CLIs)

/gsd-review --phase 3

GSD invokes each reviewer in sequence, collects structured feedback (Summary, Strengths, Concerns at HIGH/MEDIUM/LOW, Suggestions, Risk Assessment), and writes the combined output to .planning/phases/03-.../03-REVIEWS.md.

Select a single reviewer for a one-off run

/gsd-review --phase 3 --agy
/gsd-review --phase 3 --codex
/gsd-review --phase 3 --cursor

Any explicit flag overrides both the --all default and review.default_reviewers for that run.

Run every available reviewer in parallel

/gsd-review --phase 3 --all

--all always overrides config and runs the full detected set, including any configured local model servers (Ollama, LM Studio, llama.cpp).

Local model server reviewers

If you run Ollama or LM Studio locally, they are included automatically with --all when the server is reachable. You can also target them explicitly:

/gsd-review --phase 3 --ollama
/gsd-review --phase 3 --lm-studio

Configure the host addresses and model selection under review.* keys via /gsd-config --integrations if the defaults (localhost:11434 / localhost:1234) do not apply.


Read the review output

The {padded_phase}-REVIEWS.md file contains:

  • Individual reviews from each reviewer with severity-classified concerns
  • A Consensus Summary section that synthesises concerns raised by two or more reviewers — start here for the highest-priority signal
  • A Divergent Views section for areas where reviewers disagreed
  • models: and model_sources: frontmatter maps — the resolved model each reviewer actually ran under, and how that value was determined

Which model produced a review

Compare two reviewers' verdicts only after checking what actually produced each one — models: in the frontmatter gives the model per reviewer, and model_sources: gives the mechanism that recovered it.

If a reviewer's entry reads unknown, pin it: set review.models.<slug> for that lane (the key suffix is not always the lane's slug — Antigravity's is review.models.agy) so the next run records pinned. See Code-review CLI routing for the full key table.

Some unknown values are expected, not a bug to chase: lanes that accept no model at all (cursor, qwen, coderabbit), and any lane whose CLI didn't disclose one on this run. pinned is a certain value; banner and transcript are recovered from third-party CLI output and can degrade to unknown after an upstream release changes that output.

A models: entry like gpt-5.6-sol (reasoning=high) is not a formatting quirk: the (reasoning=<level>) suffix reflects a reasoning effort GSD itself applied to that lane, driven by your effort.* config — not the CLI's own default.


Incorporate feedback into the plan

Once you have reviewed the output, replan incorporating the feedback:

/gsd-plan-phase 3 --reviews

The planner reads REVIEWS.md and adjusts the plans to address the concerns before saving.


Automate the plan–review–replan loop

For phases where you want to iterate until all HIGH-severity concerns are resolved, use the convergence loop:

/gsd-plan-review-convergence 3

This runs plan-phase → review → replan → re-review up to three cycles (default). The loop exits when the HIGH-concern count reaches zero.

Convergence with a specific reviewer

/gsd-plan-review-convergence 3 --codex
/gsd-plan-review-convergence 3 --agy

Convergence with all reviewers and a higher cycle cap

/gsd-plan-review-convergence 3 --all --max-cycles 5

Stall detection: if the HIGH-concern count is not decreasing across cycles, GSD warns you. When the cycle cap is reached with open HIGH concerns, an escalation gate asks whether to proceed or review manually.


Conditionals: which reviewers to choose

Situation Recommended approach
You have Antigravity already installed --agy is always a good starting reviewer
You want free multi-reviewer coverage --agy (Google credentials) + --claude
Your project is OpenAI-heavy add --codex for an OpenAI-model perspective
You want GitHub Copilot's model add --opencode
You want to avoid API costs entirely configure Ollama with a local model and use --ollama
You need maximum coverage before a release /gsd-plan-review-convergence N --all
You're iterating quickly and want fast feedback pick one CLI: /gsd-review --phase N --agy