* fix(#4709): retire the Gemini CLI reviewer lane
Google stopped serving Gemini CLI for the free/Pro/Ultra tiers on 2026-06-18 —
the same sunset that removed the gemini RUNTIME in #1928 (shipped 1.8.0). GSD
targets solo developers, so those tiers ARE the user path: the lane spawned
`gemini {{model}} -p -`, a binary that no longer answers for the majority of
users, and five locales documented it as a supported choice.
The lane was re-created after #1928 by the reviewer-lane-as-manifest-data work
(6a9babda69, #2798/#2837, ADR-2782). Per the maintainer that re-creation was an
error in that buildout rather than a considered decision, so this corrects a
mistake and needs no ADR-2782 amendment.
Reviewer roster: 12 lanes / 13 flags -> 11 lanes / 12 flags.
TWO sources of truth had to be removed, not one. Deleting
capabilities/gemini/capability.json left the capability registry at 11 lanes
while src/review-lane-descriptor.cts's hand-maintained REVIEWER_LANES array
still carried its own complete gemini entry at 12 — precisely the disagreement
checkReviewerLaneParity exists to catch. Both are gone; both parity checkers
now run clean against the real tree (lane parity ok/0 violations, docs parity
0 violations).
Surfaces stripped of the dead flag:
- capabilities/gemini/ deleted; registry and capability-matrix regenerated
- src/review-lane-descriptor.cts: REVIEWER_LANES entry, docblock count, and the
three doc comments that used --gemini as a live example
- commands/gsd/{review,plan-review-convergence,autonomous,progress}.md and the
four matching skills/*/SKILL.md: argument-hint frontmatter and flag bullets
- gsd-core/workflows/help/modes/{full,full.compact}.md: /gsd-help signatures,
the detected-CLI list, and the reviewer-title list
- gsd-core/workflows/settings-integrations.md: the integrations wizard no longer
offers "Gemini" as a model option, and the settable-keys list drops it
- gsd-core/workflows/review.md: the `command -v gemini` probe, the --gemini
flag, the roster frontmatter, the install pointer to the sunset repo, and the
jq-less / precedence / self-skip lane lists
- gsd-core/workflows/sync-skills.md: "two runtimes (grok, gemini) resolve to
ANOTHER runtime's skills root" is now one runtime; gemini never aliased
anything, it fell through canonicalizeRuntimeName to a fail-closed default
- docs/{CONFIGURATION,COMMANDS,CLI-TOOLS}.md, docs/reference/capability-matrix.md,
docs/how-to/set-up-cross-ai-review.md — including its `npm install -g
@google/gemini-cli` instruction and the two rows recommending --gemini
- docs/features/{cross-ai-peer-review,opt-in-parallel-reviewer-lanes}.md as the
generator inputs behind docs/FEATURES.md, plus the three locale FEATURES.md
signature lines the docs-parity gate covers (the #2781 class: a flag change
that never reaches the mirrors)
Counts reconciled against measurement rather than arithmetic: 8 timeout keys of
11 lanes, 11 budget keys, 9 model keys, and four hardcoded literals in
tests/reviewer-lane-declarations.test.cjs (NEW_LANE_ONLY_IDS 5->4, LITERAL_ROSTER
12->11, two roster counts 12->11).
BEHAVIOR CHANGE, accepted deliberately: `gsd config-set review.models.gemini`
now errors with "Unknown config key". An existing key already in
.planning/config.json still parses and is simply never read, so no project fails
to load. This is the repo's own documented policy for exactly this case
(docs/CONFIGURATION.md:327 — "a key left over from a removed reviewer validated
silently and was never read. Such a key is now rejected by config-set"), so no
installer migration ships. Note my first measurement of this was WRONG: I tested
config-get, which reads undeclared keys fine, and generalised. Read and write are
different surfaces and gave different answers.
Antigravity is untouched throughout — its --antigravity/--agy flags,
review.models.agy, ~/.gemini/antigravity configHome, ~/.gemini/config global
skills root (#3738), hookEvents "gemini", GEMINI.md instruction file, and every
gemini-* model id it actually runs on.
Refs #4709
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(#4709): changeset for the reviewer-lane retirement
Type Removed: the --gemini flag and its three config keys are user-visible
surface that no longer exists.
Refs #4709
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#4709): close the 24 test failures and the locale-doc gap the gates found
An adversarial review and a full matrix run between them found substantially
more fallout than inspection had. All of it is this PR's own, and all of it is
fixed rather than waved off.
THE MATRIX RUN FOUND 24 FAILURES ACROSS 6 FILES. Inspection had predicted two.
The dominant class was a test helper that looks up a lane by slug and throws
`no declared lane 'gemini'`:
- tests/feat-2483-review-claude-mds-guard.test.cjs (6) — used gemini as the
"other declared first-party lane" to contrast against claude's env
suppression. Now qwen, verified from source as a lane that declares no `env`
(only claude does), so the contrast still holds.
- tests/review-lane-descriptor.test.cjs (6) — the duplicate-flag and
duplicate-section fixtures deliberately COLLIDED with a real declared lane to
prove the parity checker reports a duplicate. `--gemini`/`Gemini` no longer
collide with anything, so the checker reported
`descriptor_lane_not_in_registry:acme` instead and the tests proved nothing.
Now collide with `--codex`/`Codex`, reproduced against the real checker.
- tests/review-reviewer-selection.test.cjs (3) — these distinguish KNOWN-but-
undetected from UNKNOWN. gemini flipped categories, inverting what they
proved. The known case now uses qwen; `__nope__` stays the unknown fixture.
- tests/review-default-reviewers-resolution.test.cjs (2), and
tests/settings-integrations.test.cjs (3) — the wizard now offers three
reviewer CLIs, not four, so the test and its name say three.
- Two count assertions the earlier sweep missed outright:
reviewer-lane-declarations.test.cjs:359 (`length, 12`) and
reviewer-docs-parity.test.cjs:681 (`>= 12`).
THE LOCALE-DOC GAP, and why the parity gate stayed green over it. All four
locale mirrors still documented `--gemini` as a live reviewer flag. The
docs-parity checker asserts the PRESENCE of every current flag and never the
ABSENCE of a retired one, so "0 violations" was never evidence those files were
clean — my earlier reading of it as such was wrong. This is the #2781
locale-drift class in the opposite direction. Fixed across 12 locale files:
COMMANDS.md flag lists and table rows, CONFIGURATION.md `review.models.gemini`
rows and reviewer prose, CLI-TOOLS.md config examples, and
set-up-cross-ai-review.md including its install block and its
which-reviewer-to-choose row, which now recommends Antigravity.
ALSO FOUND, and instructive about my own method: docs/CONFIGURATION.md:297 still
carried a `review.models.gemini` row. My sweep had missed it because my grep
excluded lines matching `gemini-[0-9]` to spare Google's model ids — and that
row's example value is `"gemini-2.5-pro"` on the same line. The exclusion built
to avoid false positives created a false negative.
Remaining comment/example sites: src/review-reviewer-selection.cts:309 and
src/config.cts:598 named the dead flag and key as examples;
gsd-core/references/planning-config.md:269 likewise; and
review-reviewer-selection.cts:22 claimed in the PRESENT tense that gemini is a
lane-only reviewer capability. Line 38 of that same docblock says "Before this
phase the five non-runtime reviewers (gemini, ...)" and is left exactly as is —
that is past-tense history, and rewriting it would falsify the record.
Deliberately still deferred to Phase 4, because it is the RUNTIME axis rather
than the reviewer lane: the locale install-on-your-runtime.md `--gemini --global`
instructions, the USER-GUIDE colon-form notes, and the ARCHITECTURE
runtime-detection flag lists.
Both parity checkers green against the real tree; lint:ci exit 0.
Refs #4709
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(#4709): backfill the changeset PR number
pr: 0 -> 4716, now that the PR exists. Never guessed ahead of the number.
Refs #4709
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
318 lines
12 KiB
JavaScript
318 lines
12 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* Characterization tests for the reviewer selection module.
|
|
* Locks the normalizeConfiguredDefaultReviewers and resolveReviewerSelection
|
|
* export shapes and key policy decisions.
|
|
*/
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
|
|
const {
|
|
KNOWN_REVIEWER_SLUGS,
|
|
normalizeConfiguredDefaultReviewers,
|
|
resolveReviewerSelection,
|
|
} = require('../gsd-core/bin/lib/review-reviewer-selection.cjs');
|
|
|
|
describe('KNOWN_REVIEWER_SLUGS', () => {
|
|
test('known slug appears in selected with no warning; unknown slug produces a warning and is dropped', () => {
|
|
const knownSlug = KNOWN_REVIEWER_SLUGS[0];
|
|
const unknownSlug = '__not_a_real_reviewer__';
|
|
|
|
const knownResult = resolveReviewerSelection({
|
|
detected: [knownSlug],
|
|
explicitFlags: [],
|
|
allFlag: false,
|
|
configuredDefaultReviewers: [knownSlug],
|
|
});
|
|
assert.ok(
|
|
knownResult.selected.includes(knownSlug),
|
|
`expected known slug "${knownSlug}" to appear in selected`,
|
|
);
|
|
assert.ok(
|
|
knownResult.warnings.length === 0,
|
|
`expected no warnings for known slug "${knownSlug}", got: ${JSON.stringify(knownResult.warnings)}`,
|
|
);
|
|
|
|
const unknownResult = resolveReviewerSelection({
|
|
detected: [unknownSlug],
|
|
explicitFlags: [],
|
|
allFlag: false,
|
|
configuredDefaultReviewers: [unknownSlug],
|
|
});
|
|
assert.ok(
|
|
!unknownResult.selected.includes(unknownSlug),
|
|
`expected unknown slug "${unknownSlug}" to be dropped from selected`,
|
|
);
|
|
assert.ok(
|
|
unknownResult.warnings.some((w) => w.includes(unknownSlug)),
|
|
`expected a warning mentioning "${unknownSlug}", got: ${JSON.stringify(unknownResult.warnings)}`,
|
|
);
|
|
});
|
|
});
|
|
|
|
describe('normalizeConfiguredDefaultReviewers', () => {
|
|
test('returns absent=true for undefined', () => {
|
|
const r = normalizeConfiguredDefaultReviewers(undefined);
|
|
assert.ok(r.absent);
|
|
assert.deepStrictEqual(r.values, []);
|
|
assert.deepStrictEqual(r.errors, []);
|
|
});
|
|
|
|
test('returns absent=true for null', () => {
|
|
const r = normalizeConfiguredDefaultReviewers(null);
|
|
assert.ok(r.absent);
|
|
});
|
|
|
|
test('returns error for non-array', () => {
|
|
const r = normalizeConfiguredDefaultReviewers('gemini');
|
|
assert.ok(!r.absent);
|
|
assert.ok(r.errors.length > 0);
|
|
});
|
|
|
|
test('returns error for empty array', () => {
|
|
const r = normalizeConfiguredDefaultReviewers([]);
|
|
assert.ok(!r.absent);
|
|
assert.ok(r.errors.length > 0);
|
|
});
|
|
|
|
test('normalizes slugs to lowercase', () => {
|
|
const r = normalizeConfiguredDefaultReviewers(['Gemini', 'CLAUDE']);
|
|
assert.ok(!r.absent);
|
|
assert.ok(r.values.includes('gemini'));
|
|
assert.ok(r.values.includes('claude'));
|
|
});
|
|
|
|
test('deduplicates slugs case-insensitively', () => {
|
|
const r = normalizeConfiguredDefaultReviewers(['gemini', 'GEMINI']);
|
|
assert.ok(!r.absent);
|
|
assert.strictEqual(r.values.filter((s) => s === 'gemini').length, 1);
|
|
});
|
|
|
|
test('records error for invalid slug format', () => {
|
|
const r = normalizeConfiguredDefaultReviewers(['gem@ini']);
|
|
assert.ok(r.errors.some((e) => e.includes('invalid reviewer slug')));
|
|
});
|
|
});
|
|
|
|
describe('resolveReviewerSelection', () => {
|
|
test('explicit_flags source — returns intersection of flags and detected', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini', 'claude'],
|
|
explicitFlags: ['gemini'],
|
|
allFlag: false,
|
|
});
|
|
assert.equal(r.source, 'explicit_flags');
|
|
assert.deepStrictEqual(r.selected, ['gemini']);
|
|
});
|
|
|
|
test('all_flag source — returns all detected', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini', 'claude'],
|
|
explicitFlags: [],
|
|
allFlag: true,
|
|
});
|
|
assert.equal(r.source, 'all_flag');
|
|
assert.ok(r.selected.includes('gemini'));
|
|
assert.ok(r.selected.includes('claude'));
|
|
});
|
|
|
|
test('no_config_all_detected source — returns all detected when no config', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: [],
|
|
allFlag: false,
|
|
});
|
|
assert.equal(r.source, 'no_config_all_detected');
|
|
assert.deepStrictEqual(r.selected, ['gemini']);
|
|
});
|
|
|
|
test('selected is sorted alphabetically', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['claude', 'gemini'],
|
|
explicitFlags: [],
|
|
allFlag: true,
|
|
});
|
|
assert.deepStrictEqual(r.selected, [...r.selected].sort());
|
|
});
|
|
|
|
test('empty detected with no flags/config falls back to no_config_all_detected with empty selected, warnings, and errors', () => {
|
|
const r = resolveReviewerSelection({ detected: [] });
|
|
assert.equal(r.source, 'no_config_all_detected');
|
|
assert.deepStrictEqual(r.selected, []);
|
|
assert.deepStrictEqual(r.warnings, []);
|
|
assert.deepStrictEqual(r.errors, []);
|
|
});
|
|
});
|
|
|
|
/**
|
|
* ADR-2782 D4 (#2794) — absent-safe governs DISCOVERY, never explicit selection.
|
|
*
|
|
* "Not finding a lane nobody asked for is normal; failing to run a lane somebody
|
|
* asked for is an error."
|
|
*
|
|
* Before this change every explicit miss was an `info`. A TOTAL miss still
|
|
* errored, but only as a side effect of the selected set coming out empty — so
|
|
* the PARTIAL miss (`--gemini --qwen` on a host without qwen) had no signal at
|
|
* all: the review ran with a thinner reviewer set and present_results reported
|
|
* success. The workflow's own guidance names why that is wrong — "a cross-AI
|
|
* review that silently drops a lane is blind in one eye".
|
|
*/
|
|
describe('resolveReviewerSelection — explicit flags are an assertion (ADR-2782 D4)', () => {
|
|
const errorsMentioning = (r, slug) => r.errors.filter((e) => e.includes(slug));
|
|
|
|
test('an explicit flag for a detected reviewer selects it with no message', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['gemini'],
|
|
});
|
|
assert.deepStrictEqual(r.selected, ['gemini']);
|
|
assert.deepStrictEqual(r.errors, []);
|
|
assert.deepStrictEqual(r.infos, []);
|
|
});
|
|
|
|
test('a partial explicit miss errors instead of degrading silently', () => {
|
|
// THE regression row. Pre-fix this produced errors: [] and an info note,
|
|
// and the run proceeded one-eyed.
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['gemini', 'qwen'],
|
|
});
|
|
assert.deepStrictEqual(r.selected, ['gemini'], 'the detected lane is still selected');
|
|
assert.strictEqual(
|
|
errorsMentioning(r, 'qwen').length,
|
|
1,
|
|
`expected exactly one error naming qwen, got: ${JSON.stringify(r.errors)}`,
|
|
);
|
|
assert.deepStrictEqual(r.infos, [], 'the miss must not be downgraded to an info');
|
|
});
|
|
|
|
test('a sole explicit flag that is undetected errors per-slug and in aggregate', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['qwen'],
|
|
});
|
|
assert.deepStrictEqual(r.selected, []);
|
|
assert.strictEqual(errorsMentioning(r, 'qwen').length, 1);
|
|
// The pre-existing aggregate message is preserved, not replaced — the
|
|
// per-slug errors must not suppress it.
|
|
assert.ok(
|
|
r.errors.some((e) => e.includes('no selected reviewers are available')),
|
|
`expected the aggregate error to survive, got: ${JSON.stringify(r.errors)}`,
|
|
);
|
|
});
|
|
|
|
test('every missing explicit flag produces its own error, in a stable order', () => {
|
|
const forward = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['gemini', 'qwen', 'codex'],
|
|
});
|
|
const reversed = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['gemini', 'codex', 'qwen'],
|
|
});
|
|
assert.strictEqual(errorsMentioning(forward, 'qwen').length, 1);
|
|
assert.strictEqual(errorsMentioning(forward, 'codex').length, 1);
|
|
// Order must not depend on the order flags appeared on the command line.
|
|
assert.deepStrictEqual(forward.errors, reversed.errors);
|
|
});
|
|
|
|
test('a duplicate explicit flag produces exactly one error', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['qwen', 'qwen'],
|
|
});
|
|
assert.strictEqual(errorsMentioning(r, 'qwen').length, 1);
|
|
});
|
|
|
|
test('explicit flag matching is case-insensitive', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['QWEN'],
|
|
});
|
|
assert.strictEqual(errorsMentioning(r, 'qwen').length, 1);
|
|
});
|
|
|
|
test('an explicit flag with nothing detected at all errors', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: [],
|
|
explicitFlags: ['gemini'],
|
|
});
|
|
assert.deepStrictEqual(r.selected, []);
|
|
assert.strictEqual(errorsMentioning(r, 'gemini').length, 1);
|
|
});
|
|
|
|
test('pre-existing config errors do not suppress per-slug explicit errors', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: ['qwen'],
|
|
configuredDefaultReviewers: 'not-an-array',
|
|
});
|
|
assert.strictEqual(errorsMentioning(r, 'qwen').length, 1);
|
|
assert.ok(r.errors.some((e) => e.includes('must be a JSON array')));
|
|
// Guarded on the PRE-branch error count, so the aggregate fires exactly when
|
|
// it did before this change — i.e. not here, because a config error already
|
|
// existed.
|
|
assert.ok(!r.errors.some((e) => e.includes('no selected reviewers are available')));
|
|
});
|
|
|
|
test('non-string explicit flags are coerced, never thrown on', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
explicitFlags: [null, 0, { a: 1 }],
|
|
});
|
|
assert.strictEqual(r.source, 'explicit_flags');
|
|
assert.strictEqual(r.errors.length, 3 + 1, 'three unknown flags plus the aggregate');
|
|
});
|
|
});
|
|
|
|
/**
|
|
* The other half of D4, and the reason the carve-out is scoped to explicit
|
|
* flags only: discovery paths stay lenient. `--all` is a quantifier over what
|
|
* exists; `review.default_reviewers` is a preference evaluated across many
|
|
* hosts. Neither is an assertion about a specific lane, so neither errors.
|
|
*/
|
|
describe('resolveReviewerSelection — discovery paths stay lenient (ADR-2782 D4)', () => {
|
|
test('--all does not error on undetected lanes', () => {
|
|
const r = resolveReviewerSelection({
|
|
detected: ['gemini'],
|
|
allFlag: true,
|
|
});
|
|
assert.equal(r.source, 'all_flag');
|
|
assert.deepStrictEqual(r.selected, ['gemini']);
|
|
assert.deepStrictEqual(r.errors, []);
|
|
assert.deepStrictEqual(r.infos, []);
|
|
});
|
|
|
|
test('a configured default that is undetected stays an info, not an error', () => {
|
|
// `qwen` is a genuinely KNOWN lane (REVIEWER_LANES) that is simply absent from `detected` here —
|
|
// the shape this test needs. `gemini` no longer works as this fixture: it was retired from
|
|
// REVIEWER_LANES by #4709, so it now falls into the UNKNOWN-slug branch (a warning) instead of
|
|
// the known-but-undetected branch (an info) this test exists to prove.
|
|
const r = resolveReviewerSelection({
|
|
detected: ['qwen'],
|
|
configuredDefaultReviewers: ['qwen', 'codex'],
|
|
});
|
|
assert.equal(r.source, 'config_default');
|
|
assert.deepStrictEqual(r.selected, ['qwen']);
|
|
assert.deepStrictEqual(r.errors, [], 'a preference miss must not become an error');
|
|
assert.ok(
|
|
r.infos.some((i) => i.includes('codex')),
|
|
`expected an info naming codex, got: ${JSON.stringify(r.infos)}`,
|
|
);
|
|
});
|
|
|
|
test('an unknown configured slug stays a warning', () => {
|
|
// `qwen` here plays the genuinely KNOWN, detected lane so `selected`/`errors` behave as expected;
|
|
// `__nope__` remains the genuinely UNKNOWN slug under test — `gemini` would ALSO now be a valid
|
|
// fixture for the unknown branch (retired from REVIEWER_LANES by #4709), but `__nope__` already
|
|
// names that branch unambiguously without relying on retirement history.
|
|
const r = resolveReviewerSelection({
|
|
detected: ['qwen'],
|
|
configuredDefaultReviewers: ['qwen', '__nope__'],
|
|
});
|
|
assert.deepStrictEqual(r.errors, []);
|
|
assert.ok(r.warnings.some((w) => w.includes('__nope__')));
|
|
});
|
|
});
|