Tom Boucher e2bfc06558 fix(#4709): a retired runtime id must not resolve to Claude Code (#4756)
* fix(#4709): a retired runtime id must not resolve to Claude Code

AC#1 of epic #4709 — the last unmet acceptance criterion. Every other phase
(#4711, #4716, #4732, #4743, #4753) is merged; the epic does not close until
this lands.

THE DEFECT, MEASURED

Five runtime-resolution accessors resolved a RETIRED id to a plausible-looking
value, indistinguishable from the same call with a canonical id. Measured on
5d4c98cde7 by executing the built modules:

  getRuntimeLabel('gemini')              -> 'Claude Code'
  getProjectInstructionFile('gemini')    -> 'AGENTS.md'
  getGlobalConfigHomeFragment('gemini')  -> "'.claude'"
  getGlobalConfigDir('gemini')           -> ~/.claude   (byte-identical to 'claude')
  getDirName('gemini')                   -> '.claude'

So asking for a runtime Google sunset on 2026-06-18 wrote into Claude Code's
global config home and labelled the install "Claude Code". Nothing errored and
nothing warned.

AC#1 names four accessors. getDirName is the fifth, found by a reviewer: same
module, same silent-wrong-answer class, and it feeds capability-state's
runtimeConfigDir. Fixing only the four the criterion happened to list would
have left the defect reachable, so it is guarded too.

WHY THE CHECK CANNOT LIVE IN CANONICALIZATION

canonicalizeRuntimeName returns null for 'gemini', 'gemini-cli', 'Gemini',
'GEMINI' AND for ''. After canonicalization a retired id, an unknown id and an
empty string are the same value, so anything keyed off the canonical form
cannot tell them apart — it would have to treat all three alike, which is the
behaviour being fixed. The check therefore runs on the RAW input.

WHAT THIS DELIBERATELY DOES NOT DO

The criterion reads "reject a non-canonical runtime id". Taken literally that
overturns three recorded decisions, so the narrower reading was put to the
maintainer as a blocking question and this implements the answer: RETIRED ids
throw, unknown and future ids keep falling back.

Preserved:

  - The #1529 contract, written into getProjectInstructionFile's own docblock
    as a mapping table ending "unknown / future runtimes -> AGENTS.md (safe
    cross-agent default)". That default exists so a runtime GSD has never heard
    of still gets a working instruction file.
  - ADR-1239 Phase B / #1679, which preserved GLOBAL_CONFIG_HOME_FRAGMENTS
    BYTE-FOR-BYTE when it collapsed a 14-branch chain, with golden install
    parity asserting generated hook output is unchanged across every runtime.
  - The explicit `if (!runtime) return <default>` branch. Empty string is a
    supported input, not a non-canonical id.

The distinction the code encodes: ABSENCE OF KNOWLEDGE IS NOT THE SAME AS
RECORDED RETIREMENT. Unknown means "no information, degrade safely". Retired
means "we know it is gone and we know what replaced it" — and silently
substituting a different product for it is the defect.

ONE INACCURACY IN THE CRITERION, RECORDED RATHER THAN REPEATED

AC#1 says the accessors return "a Claude Code value". True for getRuntimeLabel,
getGlobalConfigHomeFragment, getGlobalConfigDir and getDirName — but
getProjectInstructionFile returns 'AGENTS.md', which is not a Claude value at
all. The defect it points at is real for all of them, so the fix covers all of
them, but the wording is wrong for one.

MATCHING

RETIRED_RUNTIME_DETAILS is a Map keyed by canonical retired id, and
RETIRED_RUNTIME_SPELLINGS maps every spelling to that id. Both are Maps, not
object literals: a literal indexed by a computed key resolves INHERITED
properties, so '__proto__' and 'constructor' were truthy and threw with every
field `undefined`, while isRetiredRuntimeId — which already went through a Set
— correctly answered false for the same input. Two guards disagreeing about one
id is worse than either answer. A Map has no prototype keys, so that hazard is
structural rather than patched. The predicate and the assertion now share one
normaliser and one table and cannot diverge.

Candidates are normalised NFKC + lowercase + strip non-alphanumerics. Folding
the separators makes 'gemini-cli', 'gemini_cli', 'gemini.cli' and 'geminicli'
one key instead of four near-misses found one at a time, and NFKC folds the
full-width 'gemini' a CJK keyboard produces. It stays MEMBERSHIP matching,
never prefix or substring: 'gemini-2.5-pro' folds to 'gemini25pro' and
'gemini-3.1-pro-preview' to 'gemini31propreview', neither a member, so Google's
live model ids — part of Antigravity's real on-disk contract — are untouched.

Homoglyph folding is deliberately not attempted, and a Cyrillic 'і' would slip
through. These values arrive from argv and env, trusted inputs here, and a
mapping broad enough to catch deliberate homoglyphs would start catching
legitimate ids. Stated rather than left for the next reader to discover.

This over-broad-match trap is the recurring shape of the whole epic: an
exclusion or match written wider than its subject. Four occurrences, each
cited: #4716's `gemini-[0-9]` sweep exclusion hid a stale review.models.gemini
row whose value was "gemini-2.5-pro" on the same line; #4753's first
model-display escape was a blanket /^ \d/ that laundered "Gemini 2.5 CLI as a
supported runtime."; its dialect rule then used a +/-24-character window that
let one legitimate reference license a live claim 21 characters away; and its
model rule treated the ABSENCE of a runtime word as a grant, passing five
unqualified live-runtime claims. Earlier drafts of this message and its
artifacts said "five" in one place and "three" in another with nothing cited;
it is four, listed here, and the artifacts now agree.

THE THROW

RetiredRuntimeError carries `code: 'GSD_RETIRED_RUNTIME'` so a caller can
handle this case without string-matching a message that may be reworded, and
the message names the id, the successor and the retiring issue.
assertNotRetiredRuntime runs as the FIRST statement of each accessor, including
before getGlobalConfigDir's explicitDir branch, so an explicit directory cannot
mask a runtime that is gone.

`gsd-tools query project-instruction-file --runtime gemini` answered the new
throw with a raw stack trace — a user-facing regression this change introduced.
Its sibling routeSkillsRoot already emitted a clean single-line error for an
unknown runtime, so that route now maps GSD_RETIRED_RUNTIME through the same
`error()` helper, and a test asserts the contract directly: non-zero exit,
stderr naming Antigravity and #1928, and no stack frame. It was the only
unwrapped call site in that CLI; I checked the rest rather than assuming.

getRuntimeNewProjectCommand is deliberately NOT guarded: its value does not
vary by runtime in a way that makes a retired id a wrong answer, so throwing
would cost callers a crash without correcting anything. Verified by observing
it return the same value across claude, codex, opencode, kimi, antigravity,
copilot and an unknown id.

RECONCILING THE TESTS THAT PINNED THE DEFECT

The full remote matrix went red with 14 failures, and every one was a
pre-existing test asserting the fallback this criterion calls a defect. One had
already been caught locally by review; the matrix found the other thirteen
across four files. They were reconciled by intent, not blanket-inverted:

  - Tests whose SUBJECT is the retired runtime — "gemini falls back on label /
    config-fragment / new-project surfaces", "gemini no longer maps to
    GEMINI.md (defaults to AGENTS.md)", "gemini is no longer a known runtime —
    falls back to AGENTS.md" — had pinned the defect, titles and all. Their
    assertions are INVERTED rather than deleted, so the history of what the
    behaviour used to be stays attached to the test that pinned it.
  - Tests whose SUBJECT is "an unregistered id falls back generically", with
    gemini merely the SAMPLE, still assert a TRUE property that this change
    deliberately preserved. Those keep their assertion and switch the sample to
    a genuinely unknown id, with a retired-id refusal pinned alongside so both
    halves of the distinction sit together.
  - The project-instruction-file parity loop dropped gemini from its
    parametrised runtimes — both sides now refuse, so there is no value to
    agree on — and gained a dedicated refusal-parity test.

A FIFTEENTH was then found by executing the touched suites locally, in process,
one file at a time — `tests/runtime-name-policy.test.cjs:135` asserted
`getProjectInstructionFile('gemini-cli') === 'AGENTS.md'`, and its own comment
read "gemini-cli was an alias for gemini", which is exactly why that spelling
is now a retired one rather than a merely-unrecognised one. Inverted like the
rest.

Two remote runs on this change were avoidable: the first by reconciling the
tests that pinned the old behaviour before shipping, the second by executing
the touched suites locally first. The matrix is the authority; it is not the
discovery mechanism. Local per-file execution is bounded and cheap and is not
the banned `node --test` fan-out.

All five touched suites now pass in process: runtime-name-policy 47/47,
gemini-runtime-removed 32/32, project-instruction-file-parity 12/12,
runtime-homes-legacy-ids-drift-guard 2/2, install 452/452.

COVERAGE

Failing-first, one per accessor as the criterion demands, each proven RED
against 5d4c98cde7 before the fix existed — the table at the top of this
message IS that baseline, and the exports the tests import did not exist yet
either.

Asserting only the throw would pass if every id threw, which would break every
install, so each property is paired with its opposite: every canonical id still
resolves on all five accessors with byte-identical values; '' keeps its
documented branch; a genuinely unknown id keeps 'Claude Code' / 'AGENTS.md' /
'.claude' / ~/.claude. That last one is the load-bearing negative — it is the
decision the maintainer chose to preserve, so a later patch that "tightens" the
guard to reject all non-canonical ids turns it red with the reason attached.

Boundary coverage maps limit-1/limit/limit+1 onto set membership: 'gemin',
'geminix', 'gemini-2.5-pro' and 'gemini-3.1-pro-preview' must NOT throw, the
retired id and its folded spellings must. '__proto__', 'constructor' and
'  CONSTRUCTOR  ' are pinned as must-not-throw, and predicate/assertion
agreement is asserted directly. Several assert.throws calls initially passed a
string as the second argument, which node treats as the MESSAGE rather than a
matcher, so they asserted nothing about the error; they now use a real
predicate checking the code.

The tests live in the owning modules' suites rather than a new issue-named
file: lint-regression-test-names rejects new bug-NNNN/fix-NNNN/issue-NNNN test
files outright and directs the regression to the owning module's suite.
scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own
--write path, since the tracked test-file count moved.

Fixes #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): backfill changeset PR number (#4756)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 22:22:16 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%