Files
msd-core/docs/how-to/run-phases-autonomously.md
Tom Boucher b54c1c5848 fix(#4709): retire the Gemini CLI reviewer lane (#4716)
* fix(#4709): retire the Gemini CLI reviewer lane

Google stopped serving Gemini CLI for the free/Pro/Ultra tiers on 2026-06-18 —
the same sunset that removed the gemini RUNTIME in #1928 (shipped 1.8.0). GSD
targets solo developers, so those tiers ARE the user path: the lane spawned
`gemini {{model}} -p -`, a binary that no longer answers for the majority of
users, and five locales documented it as a supported choice.

The lane was re-created after #1928 by the reviewer-lane-as-manifest-data work
(6a9babda69, #2798/#2837, ADR-2782). Per the maintainer that re-creation was an
error in that buildout rather than a considered decision, so this corrects a
mistake and needs no ADR-2782 amendment.

Reviewer roster: 12 lanes / 13 flags -> 11 lanes / 12 flags.

TWO sources of truth had to be removed, not one. Deleting
capabilities/gemini/capability.json left the capability registry at 11 lanes
while src/review-lane-descriptor.cts's hand-maintained REVIEWER_LANES array
still carried its own complete gemini entry at 12 — precisely the disagreement
checkReviewerLaneParity exists to catch. Both are gone; both parity checkers
now run clean against the real tree (lane parity ok/0 violations, docs parity
0 violations).

Surfaces stripped of the dead flag:
- capabilities/gemini/ deleted; registry and capability-matrix regenerated
- src/review-lane-descriptor.cts: REVIEWER_LANES entry, docblock count, and the
  three doc comments that used --gemini as a live example
- commands/gsd/{review,plan-review-convergence,autonomous,progress}.md and the
  four matching skills/*/SKILL.md: argument-hint frontmatter and flag bullets
- gsd-core/workflows/help/modes/{full,full.compact}.md: /gsd-help signatures,
  the detected-CLI list, and the reviewer-title list
- gsd-core/workflows/settings-integrations.md: the integrations wizard no longer
  offers "Gemini" as a model option, and the settable-keys list drops it
- gsd-core/workflows/review.md: the `command -v gemini` probe, the --gemini
  flag, the roster frontmatter, the install pointer to the sunset repo, and the
  jq-less / precedence / self-skip lane lists
- gsd-core/workflows/sync-skills.md: "two runtimes (grok, gemini) resolve to
  ANOTHER runtime's skills root" is now one runtime; gemini never aliased
  anything, it fell through canonicalizeRuntimeName to a fail-closed default
- docs/{CONFIGURATION,COMMANDS,CLI-TOOLS}.md, docs/reference/capability-matrix.md,
  docs/how-to/set-up-cross-ai-review.md — including its `npm install -g
  @google/gemini-cli` instruction and the two rows recommending --gemini
- docs/features/{cross-ai-peer-review,opt-in-parallel-reviewer-lanes}.md as the
  generator inputs behind docs/FEATURES.md, plus the three locale FEATURES.md
  signature lines the docs-parity gate covers (the #2781 class: a flag change
  that never reaches the mirrors)

Counts reconciled against measurement rather than arithmetic: 8 timeout keys of
11 lanes, 11 budget keys, 9 model keys, and four hardcoded literals in
tests/reviewer-lane-declarations.test.cjs (NEW_LANE_ONLY_IDS 5->4, LITERAL_ROSTER
12->11, two roster counts 12->11).

BEHAVIOR CHANGE, accepted deliberately: `gsd config-set review.models.gemini`
now errors with "Unknown config key". An existing key already in
.planning/config.json still parses and is simply never read, so no project fails
to load. This is the repo's own documented policy for exactly this case
(docs/CONFIGURATION.md:327 — "a key left over from a removed reviewer validated
silently and was never read. Such a key is now rejected by config-set"), so no
installer migration ships. Note my first measurement of this was WRONG: I tested
config-get, which reads undeclared keys fine, and generalised. Read and write are
different surfaces and gave different answers.

Antigravity is untouched throughout — its --antigravity/--agy flags,
review.models.agy, ~/.gemini/antigravity configHome, ~/.gemini/config global
skills root (#3738), hookEvents "gemini", GEMINI.md instruction file, and every
gemini-* model id it actually runs on.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): changeset for the reviewer-lane retirement

Type Removed: the --gemini flag and its three config keys are user-visible
surface that no longer exists.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): close the 24 test failures and the locale-doc gap the gates found

An adversarial review and a full matrix run between them found substantially
more fallout than inspection had. All of it is this PR's own, and all of it is
fixed rather than waved off.

THE MATRIX RUN FOUND 24 FAILURES ACROSS 6 FILES. Inspection had predicted two.
The dominant class was a test helper that looks up a lane by slug and throws
`no declared lane 'gemini'`:

- tests/feat-2483-review-claude-mds-guard.test.cjs (6) — used gemini as the
  "other declared first-party lane" to contrast against claude's env
  suppression. Now qwen, verified from source as a lane that declares no `env`
  (only claude does), so the contrast still holds.
- tests/review-lane-descriptor.test.cjs (6) — the duplicate-flag and
  duplicate-section fixtures deliberately COLLIDED with a real declared lane to
  prove the parity checker reports a duplicate. `--gemini`/`Gemini` no longer
  collide with anything, so the checker reported
  `descriptor_lane_not_in_registry:acme` instead and the tests proved nothing.
  Now collide with `--codex`/`Codex`, reproduced against the real checker.
- tests/review-reviewer-selection.test.cjs (3) — these distinguish KNOWN-but-
  undetected from UNKNOWN. gemini flipped categories, inverting what they
  proved. The known case now uses qwen; `__nope__` stays the unknown fixture.
- tests/review-default-reviewers-resolution.test.cjs (2), and
  tests/settings-integrations.test.cjs (3) — the wizard now offers three
  reviewer CLIs, not four, so the test and its name say three.
- Two count assertions the earlier sweep missed outright:
  reviewer-lane-declarations.test.cjs:359 (`length, 12`) and
  reviewer-docs-parity.test.cjs:681 (`>= 12`).

THE LOCALE-DOC GAP, and why the parity gate stayed green over it. All four
locale mirrors still documented `--gemini` as a live reviewer flag. The
docs-parity checker asserts the PRESENCE of every current flag and never the
ABSENCE of a retired one, so "0 violations" was never evidence those files were
clean — my earlier reading of it as such was wrong. This is the #2781
locale-drift class in the opposite direction. Fixed across 12 locale files:
COMMANDS.md flag lists and table rows, CONFIGURATION.md `review.models.gemini`
rows and reviewer prose, CLI-TOOLS.md config examples, and
set-up-cross-ai-review.md including its install block and its
which-reviewer-to-choose row, which now recommends Antigravity.

ALSO FOUND, and instructive about my own method: docs/CONFIGURATION.md:297 still
carried a `review.models.gemini` row. My sweep had missed it because my grep
excluded lines matching `gemini-[0-9]` to spare Google's model ids — and that
row's example value is `"gemini-2.5-pro"` on the same line. The exclusion built
to avoid false positives created a false negative.

Remaining comment/example sites: src/review-reviewer-selection.cts:309 and
src/config.cts:598 named the dead flag and key as examples;
gsd-core/references/planning-config.md:269 likewise; and
review-reviewer-selection.cts:22 claimed in the PRESENT tense that gemini is a
lane-only reviewer capability. Line 38 of that same docblock says "Before this
phase the five non-runtime reviewers (gemini, ...)" and is left exactly as is —
that is past-tense history, and rewriting it would falsify the record.

Deliberately still deferred to Phase 4, because it is the RUNTIME axis rather
than the reviewer lane: the locale install-on-your-runtime.md `--gemini --global`
instructions, the USER-GUIDE colon-form notes, and the ARCHITECTURE
runtime-detection flag lists.

Both parity checkers green against the real tree; lint:ci exit 0.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): backfill the changeset PR number

pr: 0 -> 4716, now that the PR exists. Never guessed ahead of the number.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 03:03:44 -04:00

7.3 KiB

How to run phases autonomously

Run all remaining phases — or a bounded range of them — unattended, so GSD moves through discuss → plan → execute for each phase without you driving every step.

For background on what the phase loop is doing during an autonomous run, see The phase loop.


Prerequisites

  • An active project with .planning/ROADMAP.md and .planning/STATE.md
  • All phases you want to run must be in a state that autonomous mode can drive (pending or in-progress; not already complete)
  • Any design decisions you care about should already be in PROJECT.md or captured via a prior /gsd-discuss-phase — autonomous mode can surface grey areas interactively only when you use --interactive

Run all remaining phases

/gsd-autonomous

GSD reads ROADMAP.md, discovers every incomplete phase in numeric order, and runs discuss → plan → execute on each one. After all phases complete it automatically runs the milestone lifecycle: audit → complete → cleanup.


Run a specific range of phases

Use --from and --to to bound the run. Both flags accept decimal phase numbers (e.g. 3.1).

/gsd-autonomous --from 3          # phases 3, 4, 5 … (skip already-done phases 1 and 2)
/gsd-autonomous --to 5            # phases up to and including 5
/gsd-autonomous --from 3 --to 5   # exactly phases 3, 4, and 5

When --to is reached the lifecycle step is skipped, because not all milestone phases are done. The completion banner tells you how to resume:

Resume with: /gsd-autonomous --from 6

Run a single phase

To run exactly one phase without triggering the milestone lifecycle, use --only N:

/gsd-autonomous --only 4

If the phase is already complete, autonomous mode exits immediately with a message rather than re-running it.


Run with plan convergence

Use --converge when you want each phase to run the plan-review convergence loop before execution. Both /gsd-autonomous and /gsd-progress --next --auto support this flag.

gsd config-set workflow.plan_review_convergence true

# Via autonomous (multi-phase or single-phase):
/gsd-autonomous --only 4 --converge
/gsd-autonomous --from 3 --to 5 --converge --all --max-cycles 5

# Via progress --next --auto (step-chaining with convergence):
/gsd-progress --next --auto --converge
/gsd-progress --next --auto --converge --codex --max-cycles 4

--cross-ai is accepted as an alias for --converge. Reviewer flags supported by /gsd-plan-review-convergence pass through unchanged, including --codex, --claude, --opencode, --ollama, --lm-studio, --llama-cpp, --all, and --max-cycles N.

If workflow.plan_review_convergence is not enabled, the command stops before planning and prints the enable command instead of silently falling back to regular planning.


Run with interactive discuss

By default, autonomous mode answers discuss questions automatically using smart discuss (batch table proposals). If you want to answer design questions yourself while keeping plan and execute out of the main context:

/gsd-autonomous --interactive

In interactive mode:

  • /gsd-discuss-phase runs inline and waits for your answers
  • On runtimes that support nested background dispatch, planning and execution are dispatched as background agents so you can discuss the next phase while the current one builds; on Claude Code, planning and execution run inline (the next phase's discuss does not overlap)
  • The main context stays lean — only discuss conversations accumulate (on runtimes with background dispatch; on Claude Code, inline plan/execute also accumulate)

Run on a non-Claude runtime

To run autonomously on a runtime that does not support the AskUserQuestion tool (for example Codex CLI or Gemini CLI), add --text:

/gsd-autonomous --text
/gsd-autonomous --from 3 --text

All interactive prompts become plain numbered lists; type the choice number to respond. When combined with --converge, --text is also forwarded to the convergence loop via CONVERGENCE_ARGS so reviewer prompts inside plan-review convergence use the same plain-text mode.


What safety gates still apply

Autonomous mode does not bypass GSD's quality pipeline. Each phase still:

  • Runs the plan-checker before execution
  • Reads VERIFICATION.md after execution and routes on the result
  • Pauses and asks you what to do when verification status is human_needed or gaps_found
  • Stops and presents options (fix and retry, skip phase, or stop) if any step fails

The only difference from manual execution is that passed verification advances automatically — you are not prompted between phases unless a decision is required.

The package legitimacy gate also remains active. If a plan includes a checkpoint:human-verify task for a suspicious package, the executor will stop and surface the checkpoint. Autonomous mode will not silently install flagged packages.


When not to use autonomous mode

Do not use /gsd-autonomous when:

  • Phases have unsettled design decisions. If you have not run /gsd-discuss-phase and your PROJECT.md does not capture your preferences, smart discuss will make autonomous choices you may not agree with. Run discuss interactively first, or use --interactive.

  • You need fine-grained control over a single phase. For one phase, /gsd-execute-phase N gives you step-by-step output and lets you react before continuing. Use --only N if you want the autonomous quality pipeline on a single phase but do not need step-by-step interaction.

  • The phase has novel or high-risk work. Autonomous mode skips pauses unless it hits a blocker. On a phase where you expect surprises, stay in the loop with manual execution.

  • You are mid-phase with partial execution. Autonomous mode picks up incomplete phases but it does not resume a partially-executed wave. Use /gsd-execute-phase N to finish a phase that is already in progress.

If a run stops partway through, see Debug a failed execution for how to diagnose what went wrong.


Checking progress during a run

Autonomous mode prints a progress banner before each phase:

 GSD ► AUTONOMOUS ▸ Phase 3/7: Auth Middleware [████░░░░] 28%

If you need to check where the run stands mid-session, open another terminal and run:

/gsd-progress

Resuming after a stop

If autonomous mode stops — whether you chose "Stop autonomous mode" from the blocker prompt, or the session was interrupted — resume from where it left off:

/gsd-autonomous --from 4     # replace 4 with the first incomplete phase number

GSD skips already-complete phases automatically, so it is safe to re-run from an earlier phase number if you are not sure where the run stopped.

If a prior run recorded a Deferred Verification entry in STATE.md, later /gsd-autonomous reruns skip that phase instead of re-entering the same deferral prompt. Resume deferred work with the exact command shown in the table:

/gsd-verify-work 4          # for verification_deferred_human
/gsd-plan-phase 6 --gaps   # for verification_deferred_gaps