* chore(#4729): guard the retired-runtime name, and finish the locale residue Phase 5 of 5 on epic #4709, and the phase that closes it. Two parts, one concern: make the tree clean, and keep it clean. The guard is inert until the tree is clean, and shipping the cleanup without the guard is the one-bug-at-a-time pattern this epic exists to end. WHY A GUARD, AND WHY LAST Nothing in CI answered "does any shipped surface still present a retired runtime as live?", and the two gates that look like they should cannot. checkReviewerDocsParity is one-directional: it asserts the PRESENCE of every declared reviewer flag and never the ABSENCE of a retired one, so in #4716 it reported 0 violations while all four locale mirrors still documented --gemini as a live reviewer flag, with usage examples. And tests/gemini-runtime-removed.test.cjs is scoped by construction - its own docblock limits it to the installer CLI contract and the runtime-name-policy exports; it never reads docs/**, gsd-core/workflows/**, commands/** or agents/**. Every extension to it during this epic was a hand-added assertion for a surface somebody had already noticed. A guard written earlier would have red-flagged the very references phases 1b-4b were removing, which is why it lands last. PART A - THE RESIDUE, INCLUDING WORK I SHIPPED INCOMPLETE Each site was judged against its ENGLISH counterpart, not on its own: README.{ja-JP,ko-KR,pt-BR,zh-CN}.md :9 :24 :46 English README.md has ZERO occurrences -> substituted "Antigravity CLI, Kimi CLI" how-to/execute-a-phase.md:88 x4 locales fixed in #4728 -> substitute how-to/verify-and-ship.md:89 x4 locales fixed in #4728 -> substitute FEATURES.md cross-AI CLI list :1419 no Gemini -> DELETE FEATURES.md REQ-MULTI-RT-01 :1709 -> substitute FEATURES.md REQ-SKILLS-03 :1952 -> rewrite FEATURES.md REQ-QUOTA-02 :3256 deleted upstream -> delete VERSIONING.md:133 stale manifest -> see below The twelve README occurrences were an adversarial reviewer's BLOCKER, and the reason they survived my own sweep is structural: root-level *.md was outside the guard's scan set, so the repo's most-read runtime-advertising surface was invisible to the guard meant to police it. :46 is a live installer-runtime claim - it tells the reader the installer will offer a runtime that no longer exists. Checked for the duplicate-name trap before substituting: neither Antigravity nor Kimi appears anywhere in those four files. Two of these are mine to own: I fixed the ENGLISH execute-a-phase.md and verify-and-ship.md in #4728 and left all four mirrors behind. Unfinished work, not a deferral. Two more show why "substitute Gemini -> Antigravity" is the wrong default: in the cross-AI list and REQ-QUOTA-02 English DELETES the name, because Antigravity was already in the list or the classifier had dropped it. Substituting would have duplicated a name - the identical trap ARCHITECTURE.md:24 set in #4728, where English holds Kimi CLI in that slot. VERSIONING.md:133 is a different and worse defect than translation lag. Under "Manifest Version Sync" it listed gemini-extension.json as a version-synced manifest. That file is ABSENT from the repo, and scripts/sync-manifest-versions.cjs says so in its own comment - "#1928: gemini-extension.json was removed with the gemini runtime ... it is no longer a registered manifest" - while VERSIONED_MANIFESTS holds plugin.json, marketplace.json and vscode/package.json. So the doc named a manifest that does not exist AND omitted the one that replaced it. Both fixed, verified against the owning code rather than inferred from the name. The replacement bullet cites #1942, the issue that actually registered vscode/package.json, matching the convention of its neighbours. pt-BR/FEATURES.md is a 77-line stub genuinely lacking two sites, and ko-KR has no REQ-QUOTA-02 line. Skipped and recorded, never invented. PART B - THE GUARD scripts/lint-retired-runtime-name.cjs, modelled on scripts/lint-legacy-dir-name.cjs - the repo's own precedent for this problem shape (forbid a retired token, allowlist frozen content, self-exempt via a split literal, a REPO_ROOT test seam, lib/cli-exit.cjs, exit 0/1). Case sensitivity IS the mechanism, not an accident. The naive guard - "the string gemini must not appear" - is WRONG, not merely noisy: that string is load-bearing across Antigravity's real on-disk contract. A case-sensitive, standalone, capitalised name works because every legitimate reference is spelled differently and therefore cannot match: lowercase config homes (~/.gemini/antigravity, ~/.gemini/config, #3738), lowercase hyphenated model ids (gemini-2.5-flash-lite), uppercase env vars (GEMINI_API_KEY), and GEMINI.md. Table-driven, so the next retired runtime costs one row. THE ALLOWLIST IS THE ENTIRE RISK SURFACE, so it is three tiers, not one. Two rounds of isolated adversarial review reshaped it; both are recorded in .gsd/bug/chore-4729-gemini-drift-guard/60-review.json. ROUND 2 FOUND ONE ROOT CAUSE BEHIND TWO SEPARATE HOLES, and it was mine: both Tier-1 rules treated the ABSENCE of a runtime word as a GRANT. A veto list can never be complete, so "no runtime word found" silently exempted every phrasing nobody had enumerated. Demonstrated: `The installer now offers Gemini 3.`, `Supported agents include Gemini 3, Kimi, and Cursor.` and three more exited 0, as did `Suportamos Gemini, no estilo padrao, como runtime de instalacao.` and `Gemini 兼容,并且是受支持的运行时之一。`, both of which literally contain `runtime` or `运行时`. The fix was to stop enumerating exceptions and invert the evidence direction: Tier 1(a) - the hook DIALECT Antigravity inherits. Position is language-dependent and MEASURED: en Gemini-style/-compatible, ja Gemini スタイル, ko Gemini 스타일/호환, zh Gemini 风格 / 与 Gemini 兼容的, pt "no estilo Gemini" / "compatível com Gemini" where the qualifier PRECEDES the name. The marker must now form an ADJACENT COMPOUND with the name, not merely sit in a +/-24-character window - that window let `| Antigravity | Gemini-style hooks | Gemini support is live |` exit 0, one legitimate reference licensing a fresh live claim 21 characters later. The runtime-word veto is now LINE-GLOBAL. Ten real lines legitimately pair a dialect compound with a runtime word (`~/.gemini/antigravity-cli` in a table cell, "runtime files" in the same sentence); each is an explicit pin rather than a reason to loosen the veto for everyone. Measured: widening it surfaced exactly those ten and no others. Tier 1(b) - the provider/model axis. A version optionally followed by a qualifier, including full-width digits and CJK punctuation, AND positive model-axis evidence on the line, AND no runtime word. The positive requirement is the part that matters: all eight real model-axis lines in the repo name a model explicitly, so requiring it costs nothing on the real tree while flagging every laundering attempt. It is also the honest resolution of the agent/target tension below - rather than guess at an exhaustive veto list, stop treating an empty veto as evidence. Tier 2 - PINNED OCCURRENCES, now SPAN-SCOPED. A pin excuses only a match falling INSIDE an occurrence of its own snippet. Line-level containment let `Known provider menu update: Gemini CLI is once again a selectable GSD runtime.` and `Install target: Google (Gemini) - choose Gemini CLI as your GSD runtime.` both exit 0, because a short snippet elsewhere on the line pre-approved a brand-new claim. Span scoping makes short snippets safe: `Google (Gemini)` can only ever excuse the match inside those 15 characters. A LOAD-TIME validator now requires every pin to contain a retired name, and it immediately caught five of MY OWN pins whose snippets sat BESIDE the name rather than covering it - each would have shipped permanently inert and permanently reported stale. All pins were then reconciled in one pass. A pin is also marked used by PRESENCE on the line now, rather than only on the Tier-2 branch. Previously a pinned line that a general rule also matched never marked its pin used, producing a provably FALSE "no line matches pinned snippet" whose printed remedy told the maintainer to delete a pin that was still needed. Tier 3 - blanket trust, and a new occurrence inside it IS invisible. CHANGELOG.md and `.changeset/` - the rendered changelog and its source, one surface - plus six append-only directories. All 21 `.changeset/` hits were measured to be fragments DESCRIBING the retirement or a fix to it, 464 of them under archived/; a fragment can only describe what already shipped and is deleted at release, so pinning them would be friction with no signal. The cost is stated in the guard's own header rather than hidden. THE SCAN SET IS NOW EVERY TRACKED *.md FILE (1165 read). The original prefix list left `.github/`, `.changeset/`, `capabilities/`, `playbooks/` and `references/` invisible - and `.changeset/*.md` renders into CHANGELOG.md, so a live claim introduced there was invisible at BOTH ends. The escape hatch must now carry a justification (`gsd-allow-retired-runtime-name: <reason>`). A bare marker is rejected: it is checked first, excuses the whole line, and the failure message advertises it, so an unexplained one is indistinguishable from a silenced defect. Plus an anti-vacuity floor counting files actually READ, not files listed - a candidate count stays healthy-looking even if every read failed. A FALSE NEGATIVE I INTRODUCED, AND CLOSED The model-display escape began as a blanket /^ \d/ - "space then a digit" - which also matched "Install for Gemini 2.5 CLI as a supported runtime.", laundering a genuine stale-runtime claim through an attached version number. That was the THIRD appearance of one failure shape in this epic: an exclusion added to suppress false positives creating a false negative. #4716's sweep excluded lines matching gemini-[0-9] to spare Google's model ids, and thereby hid a stale review.models.gemini row whose example value was "gemini-2.5-pro" ON THE SAME LINE. Round 2 then produced the FOURTH and FIFTH instances, which is why the fix this time was to invert the rule's evidence direction rather than to enumerate more exceptions. The veto is word-anchored for Latin terms - unanchored, case-insensitive "CLI" matched inside "client" and would have vetoed legitimate model lists - and raw for CJK terms, where \b is ASCII-word-based and would never fire beside an ideograph, so anchoring them would silently disable the veto in ja/ko/zh. "agent" and "target" were deliberately left OUT: both occur throughout ordinary prose ("AI coding agents (Claude Code, Codex, Gemini 2.5 Pro)"), so vetoing on them would red correct content instead of catching runtime claims. The reasoning is in the guard's comment, not just the omission - and Tier 1(b)'s positive-evidence requirement is what makes that omission safe, since the rule no longer depends on the veto list being complete. COVERAGE tests/lint-retired-runtime-name.test.cjs drives the guard through its GSD_LINT_RETIRED_RUNTIME_REPO_ROOT seam against fixture repos, mirroring tests/lint-legacy-dir-name.test.cjs. A guard never observed failing is not a guard, and this epic already shipped one that was vacuous for 2 of its 5 files, so properties are paired against BOTH failure modes - too broad silently absorbs a future defect, too narrow reds on legitimate content. Floor boundaries are covered at 149/150/151. The round-2 reviewer's sharpest point was about that claim, and it was right: the first matrix's pairing was "true of the properties chosen, not of the predicate's actual surface" - not one of its twenty properties could see the dialect adjacency hole, a non-adjacent runtime word, pin shadowing, or an over-broad pin colliding with a new line. Every one of those is now a committed regression using the reviewer's own attack line verbatim, and the local fixture harness went from 14 cases to 35 (PASS=35 FAIL=0). That harness earned a finding of its own. Its first run reported PASS=2 FAIL=12 with BOTH passes VACUOUS: `git add` has no -q flag on this build, so nothing staged, every fixture hit the empty-walk error path, and the two checks that assert an ABSENCE passed off that error path rather than off real guard logic. A staging failure is now fatal and every absence-asserting check first proves the walk ran and the expected violation was flagged. Later, one case failed because its fixture supplied only one of a pinned file's two approved lines, so the stale-pin check fired correctly - the expectation was wrong, not the guard. Telling those two apart is the whole value of running a matrix rather than reasoning about one. On the two orthogonal reviews: the isolated adversarial pass executed a great deal of code, across two rounds, against its own fixture repos. The security pass did NOT - it self-discloses that it verified by reading only, because node --test is hard-blocked here. Saying so plainly, because "two orthogonal reviews" without that caveat overstates what the second one established. It also raised, and I cleared by measurement, a concern that importing escapeRegex from a gitignored build artifact would break lint:ci on an unbuilt clone: six other tracked scripts already require that exact path, three of them already in lint:ci, and .github/workflows/test.yml:192-193 runs `npm run build:lib` immediately before it for exactly this reason. Part A has no new test deliberately - those edits are covered by the guard itself inside lint:ci, and a separate per-locale assertion would duplicate it and then drift from it. The one exception is the root README case, which IS pinned: that residue was invisible to the guard rather than merely unasserted, so the fix is a scan-set change and needs its own regression test. No mode-bit read-failure fixture was added on purpose: the benches run as root, where chmod-based IO injection is vacuous, so such a test would assert nothing. The test's fixture helpers write throwaway docs/ paths, which trips lint-docs-guard-registration's reader-name heuristic. Resolved the way that lint documents - a header `// docs-guard-exempt:` marker plus a baseline entry - because the fixtures only WRITE scratch data and never read shipped docs; the baseline was re-confirmed, not merely extended, each time locale and adversarial fixtures were added. scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own --write path, since a new test file changes the count lint:generated-sync reads. Fixes #4729 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4729): backfill changeset PR number (#4753) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
GSD Core documentation
Documentation is organised into four quadrants: tutorials help you learn by doing, how-to guides solve specific tasks, reference states authoritative facts, and explanation explores concepts and design decisions.
Language versions: English · Português (pt-BR) · 日本語 · 简体中文
Tutorials
- Your first project — install to first shipped phase, one guaranteed path
- Onboarding an existing codebase — bring GSD Core to a brownfield repo
- Build your first capability — author a tiny declarative capability and watch it act in the loop
- Install your first capability — install a third-party capability end-to-end: consent, verify, check for updates, remove
How-to guides
- Install on your runtime — runtime-specific install steps for all 16 supported runtimes
- Install a minimal GSD and add skills later — install only the core skills, then grow the surface with profiles and
/gsd-surface - Attach a plugin-provided skill to a GSD agent — use the
global:plugin:skillentry form to load Claude Code plugin skills into agent prompts - Discuss a phase — capture implementation decisions before planning begins
- Resolve edge-coverage findings — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- Probe edges in a non-English project — get real edge coverage on a spec written in another language, and tell "no edges here" apart from "the probe could not read it"
- Resolve prohibition findings — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions
- Resolve an unreachable-workflow finding — wire or fully sweep a shipped workflow that no command, agent, or skill references
- Acknowledge emitted-artifact drift — declare a deliberate emitted-byte ripple or workflow/agent growth in a commit trailer, and migrate an older ack fragment
- Change the STATE.md schema — add, change or remove a STATE.md frontmatter key and keep the template and all five reference documents in step
- Resolve verify-command path findings — fix an
<automated>verify command whose target directory does not resolve from the executor's cwd - State a failing direction — say what output constitutes failure for an
<automated>verify command, and migrate a phase planned before the rule - Resolve a contract-drift finding — bring an agent's completion contract, read-tag gate, or deleted-file test reference back into agreement with the registry
- Resolve unreachable-guard findings — fix shell guards whose fallback arm cannot run, and tell "nothing to report" apart from "could not look"
- Declare a hook's crash policy — terminate a GSD hook with
allow/deny/crash, declare itsON_CRASHpolicy, and tell a hook's own crash apart from a check that could not run at all - Resolve a skipped capability probe — act on a coverage gate that held your phase for an unestablished scope, or a planning checkpoint that reported
skippedinstead of a verdict - Diagnose which gsd-tools is running — tell this package's tool apart from the predecessor's colliding binary and from a gsd-core too old to identify itself
- Resolve an ESLint glob-coverage finding — bring a source file that matches no lint rule under coverage, or record a reasoned exemption
- Resolve a raw-terminator finding — pick
runMain/ExitError,terminateNow, orprocess.exitCodefor alocal/require-registered-exitfinding, and know the two allowlist entries and the rule's documented evasions - Adopt the v2 exit contract — turn on
gsd-tools's versioned exit-code projection, read the code table including what80(DEGRADED) means, and migrate a CI gate that treats any non-zero exit as fatal - Read the statusline freshness marker — turn on
state ~N commits back, and tell "STATE.md is fresh" apart from "freshness could not be established" - Consume the planning snapshot — read
planning inspectfrom a dashboard or harness, and tell "nothing to report" apart from "could not look" - Read CI timeout budget signals — find the near-cap warning on a run, read the accumulated
tests/ci-timeout-budget-history.jsonltrend, and know which lever (cap, shard balance, shard-1 contents) a repeatedly-near-cap lane calls for - Consume the state contract — read
.planning/state.jsonfrom a workbench or editor extension, gate on the contract version, and tell "nothing to show" apart from "could not look" - Keep planning docs out of a shared repo — make
.planning/local-only, including untracking files git already tracks (the step.gitignorealone cannot do) - Publish PRs without planning artifacts — keep
.planning/committed locally, so worktrees and/gsd-undokeep working, whileplanning.pr_strictkeeps every planning path out of the branch you push - Plan a phase — run research, decompose work, and verify plan quality
- Verify a dependency-compatibility claim — act on a compatibility claim the researcher left
[ASSUMED], and tell "nothing declared" apart from "a constraint is declared" and "the lookup failed" - Execute a phase — run plans in parallel waves with fresh-context subagents
- Enable parallel reviewer lanes — cut a multi-reviewer
/gsd-reviewpass toward its slowest lane, and tell a rate-limited lane apart from one that was never selected - Enable concurrent per-plan planners in chunked mode — dispatch chunked
/gsd-plan-phase's per-plan Tasks together within one outline Wave instead of one at a time, and know when the setting has no effect - Verify and ship — walk through completed work, diagnose failures, and create the PR
- Catch complexity before it compounds — enable the post-execute refactor hook, read a proposal's score vs. anchor delta, and accept or decline it
- Run phases autonomously — use autonomous mode for unattended phase execution
- Handle quick and fast tasks — use
/gsd-quickand/gsd-fastfor ad-hoc work outside the phase loop - Batch quick tasks — run several
/gsd-quick-shaped tasks together with/gsd-quick-batch, understand capacity/isolation, and recover a failed or interrupted batch - Configure model profiles — switch between quality, balanced, and budget model tiers
- Control which host runtime GSD reports — read the
agent_runtimeladder, understand what host detection looks at, and pin the runtime when detection is not what you want - Set up cross-AI review — configure a second AI to review code produced by the primary agent
- Scope code review depth by path — escalate
/gsd-code-reviewtodeepfor sensitive directories while the rest of the repo stays at the default depth - Work in parallel with workstreams — run independent lines of work simultaneously using workstreams
- Isolate work with workspaces — use workspaces to sandbox experimental or risky changes
- Debug a failed execution — diagnose and recover from broken or incomplete phase execution
- Interpret scope-conformance warnings — read the advisory the worktree-wave merge emits when a plan branch commits outside its declared scope
- Interpret install-shadow warnings — read the advisory GSD Core emits when a
/gsd-*trigger is installed at both scopes and one silently wins, and tell "nothing to report" apart from "could not look" - Interpret
state validateresults — read thescopereason codes and tell "nothing to report" apart from "could not look" - Spike and sketch — use
/gsd-spikeand/gsd-sketchfor exploratory work before committing to a plan - Design a UI phase — use the UI phase loop for frontend and visual work
- Enable live-DOM verification — opt a project into browser-backed UI acceptance checks during execution, handle the browser-profile lock, and tell "nothing to report" apart from "could not look"
- Develop a Capability for GSD 1.5+ — add feature Capabilities, hook fragments, and registry entries
- Develop a task-content resolver capability — declare a
taskContentResolversoexecute-plan.mdresolves per-task content from your external issue tracker instead ofPLAN.md - Ship a reviewer lane in your capability — declare a
reviewerbody so/gsd-reviewdiscovers, invokes, and renders your external review CLI or model endpoint - List your reviewer lane in the registry — publish a lane you have built to the Reviewer Lane Registry so other people can find and install it
- Take over a capability or EoS integration — assume maintainership of an existing third-party capability, reviewer lane, or EoS host integration through a handoff, an adoption fork, first-party absorption, or a de-listing
- Add or update a host's integration — set a host's documentation-sourced
runtime.hostIntegrationaxes (ADR-1239 Phase A), with theundocumentedsentinel rule - Migrate an install test to the executed plan — convert an
fs.existsSync-probing install test group to a value assertion againstinstallRuntimeArtifacts's executed-plan return, and test against a fake fs adapter - Vendor a dependency — add a third-party package
gsd-core/bin/**needs at runtime as a verbatim vendored artifact, keep it out ofdependencies, and pick the right upstream bundle - Turn a capability off (and keep it off) — disable a capability via the surface, or gate individual hooks off without removing the capability
- Drive GSD from a tracker issue — start a phase from a GitHub, Linear, or Jira issue
- Migrate from GSD 2 — upgrade an existing GSD 2 project to GSD Core
- Update GSD — re-run the installer to pick up the latest release
- Clean up get-shit-done-cc — remove leftover old-package artifacts that cause a spurious
⬆ /gsd-updateindicator after migrating to@opengsd/gsd-core - Fix the worktree base-mismatch (exit 42) error — resolve the branch-divergence condition that halts parallel phase execution
- Recover and troubleshoot — fix common problems, rebuild context, and uninstall
Reference
- Commands — every command with flags and examples
- Configuration — full config schema, model profiles, git branching strategies
- CLI tools —
gsd-tools.cjsprogrammatic API for workflows and agents - JSON error mode —
gsd-toolsfailure channels: faults (stderr, exit 1) vs degraded results (stdout, exit 0), and the reason-code taxonomy - Features — complete feature index
- Inventory — installed skills and surface map
- STATE.md schema — field-by-field reference for
.planning/STATE.md - CONTEXT.md schema — field-by-field reference for
.planning/phases/<N>/CONTEXT.md - PLAN.md schema — field-by-field reference for
.planning/phases/<N>/PLAN.md - Planning artifacts — all
.planning/files and their roles - Review and verification capabilities — code review, security, and Nyquist capability ownership and hook contracts
- Gate predicates — canonical specification of the phase-gate predicate vocabulary
- Capability matrix — generated catalogue of every capability's role, tier, extension points, hook kinds, and
engines.gsd - Exit code reference — generated catalogue of every registered process exit code, its name, meaning, and owning module, plus the reserved bands and the v1/v2 exit contract
- Capability manifest — the full
capability.jsonschema and validation rules gsd capabilitycommand — install / update / remove / list reference for third-party capabilities- Workflow fragments — in-file
<!-- gsd:section -->marker grammar for fragmentizing workflow markdown at emission time - Partition rules for compact-content splits — the protected-content list, sentinel syntax, and the five CI checks a
workflow.compact_contentspine/detail split must obey - Reviewer Lane Registry — generated catalogue of third-party reviewer lanes, with their flags, transport, and install commands
Explanation
- Context engineering — how context rot forms and how GSD Core prevents it
- The phase loop — design rationale for the Discuss → Plan → Execute → Verify → Ship cycle
- Multi-agent orchestration — how subagents are spawned, scoped, and coordinated
- Security model — trust boundaries, permissions, and safe automation
- The capability trust model — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- How overlay capabilities compose — why first-party always wins and how the loader resolves precedence, conflicts, and fail-open load-failure warnings
- Architecture — system architecture, agent model, and data flow
- The Embeddable Orchestration System — one public, versioned contract for embedding GSD across many hosts
- Discuss modes — assumptions mode vs interview mode for
/gsd-discuss-phase - Context monitoring — context window monitoring hook architecture
- Issue-driven orchestration — recipe for driving GSD from a tracker issue using existing primitives
Related
- What's new in 1.7.0 — curated highlights of the 1.7.0 release
- Root README — landing page, quickstart, and documentation overview
- Changelog — release history