Commit Graph

2 Commits

Author SHA1 Message Date
Tom Boucher
e2bfc06558 fix(#4709): a retired runtime id must not resolve to Claude Code (#4756)
* fix(#4709): a retired runtime id must not resolve to Claude Code

AC#1 of epic #4709 — the last unmet acceptance criterion. Every other phase
(#4711, #4716, #4732, #4743, #4753) is merged; the epic does not close until
this lands.

THE DEFECT, MEASURED

Five runtime-resolution accessors resolved a RETIRED id to a plausible-looking
value, indistinguishable from the same call with a canonical id. Measured on
5d4c98cde7 by executing the built modules:

  getRuntimeLabel('gemini')              -> 'Claude Code'
  getProjectInstructionFile('gemini')    -> 'AGENTS.md'
  getGlobalConfigHomeFragment('gemini')  -> "'.claude'"
  getGlobalConfigDir('gemini')           -> ~/.claude   (byte-identical to 'claude')
  getDirName('gemini')                   -> '.claude'

So asking for a runtime Google sunset on 2026-06-18 wrote into Claude Code's
global config home and labelled the install "Claude Code". Nothing errored and
nothing warned.

AC#1 names four accessors. getDirName is the fifth, found by a reviewer: same
module, same silent-wrong-answer class, and it feeds capability-state's
runtimeConfigDir. Fixing only the four the criterion happened to list would
have left the defect reachable, so it is guarded too.

WHY THE CHECK CANNOT LIVE IN CANONICALIZATION

canonicalizeRuntimeName returns null for 'gemini', 'gemini-cli', 'Gemini',
'GEMINI' AND for ''. After canonicalization a retired id, an unknown id and an
empty string are the same value, so anything keyed off the canonical form
cannot tell them apart — it would have to treat all three alike, which is the
behaviour being fixed. The check therefore runs on the RAW input.

WHAT THIS DELIBERATELY DOES NOT DO

The criterion reads "reject a non-canonical runtime id". Taken literally that
overturns three recorded decisions, so the narrower reading was put to the
maintainer as a blocking question and this implements the answer: RETIRED ids
throw, unknown and future ids keep falling back.

Preserved:

  - The #1529 contract, written into getProjectInstructionFile's own docblock
    as a mapping table ending "unknown / future runtimes -> AGENTS.md (safe
    cross-agent default)". That default exists so a runtime GSD has never heard
    of still gets a working instruction file.
  - ADR-1239 Phase B / #1679, which preserved GLOBAL_CONFIG_HOME_FRAGMENTS
    BYTE-FOR-BYTE when it collapsed a 14-branch chain, with golden install
    parity asserting generated hook output is unchanged across every runtime.
  - The explicit `if (!runtime) return <default>` branch. Empty string is a
    supported input, not a non-canonical id.

The distinction the code encodes: ABSENCE OF KNOWLEDGE IS NOT THE SAME AS
RECORDED RETIREMENT. Unknown means "no information, degrade safely". Retired
means "we know it is gone and we know what replaced it" — and silently
substituting a different product for it is the defect.

ONE INACCURACY IN THE CRITERION, RECORDED RATHER THAN REPEATED

AC#1 says the accessors return "a Claude Code value". True for getRuntimeLabel,
getGlobalConfigHomeFragment, getGlobalConfigDir and getDirName — but
getProjectInstructionFile returns 'AGENTS.md', which is not a Claude value at
all. The defect it points at is real for all of them, so the fix covers all of
them, but the wording is wrong for one.

MATCHING

RETIRED_RUNTIME_DETAILS is a Map keyed by canonical retired id, and
RETIRED_RUNTIME_SPELLINGS maps every spelling to that id. Both are Maps, not
object literals: a literal indexed by a computed key resolves INHERITED
properties, so '__proto__' and 'constructor' were truthy and threw with every
field `undefined`, while isRetiredRuntimeId — which already went through a Set
— correctly answered false for the same input. Two guards disagreeing about one
id is worse than either answer. A Map has no prototype keys, so that hazard is
structural rather than patched. The predicate and the assertion now share one
normaliser and one table and cannot diverge.

Candidates are normalised NFKC + lowercase + strip non-alphanumerics. Folding
the separators makes 'gemini-cli', 'gemini_cli', 'gemini.cli' and 'geminicli'
one key instead of four near-misses found one at a time, and NFKC folds the
full-width 'gemini' a CJK keyboard produces. It stays MEMBERSHIP matching,
never prefix or substring: 'gemini-2.5-pro' folds to 'gemini25pro' and
'gemini-3.1-pro-preview' to 'gemini31propreview', neither a member, so Google's
live model ids — part of Antigravity's real on-disk contract — are untouched.

Homoglyph folding is deliberately not attempted, and a Cyrillic 'і' would slip
through. These values arrive from argv and env, trusted inputs here, and a
mapping broad enough to catch deliberate homoglyphs would start catching
legitimate ids. Stated rather than left for the next reader to discover.

This over-broad-match trap is the recurring shape of the whole epic: an
exclusion or match written wider than its subject. Four occurrences, each
cited: #4716's `gemini-[0-9]` sweep exclusion hid a stale review.models.gemini
row whose value was "gemini-2.5-pro" on the same line; #4753's first
model-display escape was a blanket /^ \d/ that laundered "Gemini 2.5 CLI as a
supported runtime."; its dialect rule then used a +/-24-character window that
let one legitimate reference license a live claim 21 characters away; and its
model rule treated the ABSENCE of a runtime word as a grant, passing five
unqualified live-runtime claims. Earlier drafts of this message and its
artifacts said "five" in one place and "three" in another with nothing cited;
it is four, listed here, and the artifacts now agree.

THE THROW

RetiredRuntimeError carries `code: 'GSD_RETIRED_RUNTIME'` so a caller can
handle this case without string-matching a message that may be reworded, and
the message names the id, the successor and the retiring issue.
assertNotRetiredRuntime runs as the FIRST statement of each accessor, including
before getGlobalConfigDir's explicitDir branch, so an explicit directory cannot
mask a runtime that is gone.

`gsd-tools query project-instruction-file --runtime gemini` answered the new
throw with a raw stack trace — a user-facing regression this change introduced.
Its sibling routeSkillsRoot already emitted a clean single-line error for an
unknown runtime, so that route now maps GSD_RETIRED_RUNTIME through the same
`error()` helper, and a test asserts the contract directly: non-zero exit,
stderr naming Antigravity and #1928, and no stack frame. It was the only
unwrapped call site in that CLI; I checked the rest rather than assuming.

getRuntimeNewProjectCommand is deliberately NOT guarded: its value does not
vary by runtime in a way that makes a retired id a wrong answer, so throwing
would cost callers a crash without correcting anything. Verified by observing
it return the same value across claude, codex, opencode, kimi, antigravity,
copilot and an unknown id.

RECONCILING THE TESTS THAT PINNED THE DEFECT

The full remote matrix went red with 14 failures, and every one was a
pre-existing test asserting the fallback this criterion calls a defect. One had
already been caught locally by review; the matrix found the other thirteen
across four files. They were reconciled by intent, not blanket-inverted:

  - Tests whose SUBJECT is the retired runtime — "gemini falls back on label /
    config-fragment / new-project surfaces", "gemini no longer maps to
    GEMINI.md (defaults to AGENTS.md)", "gemini is no longer a known runtime —
    falls back to AGENTS.md" — had pinned the defect, titles and all. Their
    assertions are INVERTED rather than deleted, so the history of what the
    behaviour used to be stays attached to the test that pinned it.
  - Tests whose SUBJECT is "an unregistered id falls back generically", with
    gemini merely the SAMPLE, still assert a TRUE property that this change
    deliberately preserved. Those keep their assertion and switch the sample to
    a genuinely unknown id, with a retired-id refusal pinned alongside so both
    halves of the distinction sit together.
  - The project-instruction-file parity loop dropped gemini from its
    parametrised runtimes — both sides now refuse, so there is no value to
    agree on — and gained a dedicated refusal-parity test.

A FIFTEENTH was then found by executing the touched suites locally, in process,
one file at a time — `tests/runtime-name-policy.test.cjs:135` asserted
`getProjectInstructionFile('gemini-cli') === 'AGENTS.md'`, and its own comment
read "gemini-cli was an alias for gemini", which is exactly why that spelling
is now a retired one rather than a merely-unrecognised one. Inverted like the
rest.

Two remote runs on this change were avoidable: the first by reconciling the
tests that pinned the old behaviour before shipping, the second by executing
the touched suites locally first. The matrix is the authority; it is not the
discovery mechanism. Local per-file execution is bounded and cheap and is not
the banned `node --test` fan-out.

All five touched suites now pass in process: runtime-name-policy 47/47,
gemini-runtime-removed 32/32, project-instruction-file-parity 12/12,
runtime-homes-legacy-ids-drift-guard 2/2, install 452/452.

COVERAGE

Failing-first, one per accessor as the criterion demands, each proven RED
against 5d4c98cde7 before the fix existed — the table at the top of this
message IS that baseline, and the exports the tests import did not exist yet
either.

Asserting only the throw would pass if every id threw, which would break every
install, so each property is paired with its opposite: every canonical id still
resolves on all five accessors with byte-identical values; '' keeps its
documented branch; a genuinely unknown id keeps 'Claude Code' / 'AGENTS.md' /
'.claude' / ~/.claude. That last one is the load-bearing negative — it is the
decision the maintainer chose to preserve, so a later patch that "tightens" the
guard to reject all non-canonical ids turns it red with the reason attached.

Boundary coverage maps limit-1/limit/limit+1 onto set membership: 'gemin',
'geminix', 'gemini-2.5-pro' and 'gemini-3.1-pro-preview' must NOT throw, the
retired id and its folded spellings must. '__proto__', 'constructor' and
'  CONSTRUCTOR  ' are pinned as must-not-throw, and predicate/assertion
agreement is asserted directly. Several assert.throws calls initially passed a
string as the second argument, which node treats as the MESSAGE rather than a
matcher, so they asserted nothing about the error; they now use a real
predicate checking the code.

The tests live in the owning modules' suites rather than a new issue-named
file: lint-regression-test-names rejects new bug-NNNN/fix-NNNN/issue-NNNN test
files outright and directs the regression to the owning module's suite.
scripts/lib/macos-conformance-tier.generated.cjs regenerated through its own
--write path, since the tracked test-file count moved.

Fixes #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): backfill changeset PR number (#4756)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 22:22:16 -04:00
Tom Boucher
5d4c98cde7 chore(#4729): guard the retired-runtime name, and finish the locale residue (#4753)
* chore(#4729): guard the retired-runtime name, and finish the locale residue

Phase 5 of 5 on epic #4709, and the phase that closes it. Two parts, one
concern: make the tree clean, and keep it clean. The guard is inert until the
tree is clean, and shipping the cleanup without the guard is the
one-bug-at-a-time pattern this epic exists to end.

WHY A GUARD, AND WHY LAST

Nothing in CI answered "does any shipped surface still present a retired
runtime as live?", and the two gates that look like they should cannot.
checkReviewerDocsParity is one-directional: it asserts the PRESENCE of every
declared reviewer flag and never the ABSENCE of a retired one, so in #4716 it
reported 0 violations while all four locale mirrors still documented --gemini
as a live reviewer flag, with usage examples. And
tests/gemini-runtime-removed.test.cjs is scoped by construction - its own
docblock limits it to the installer CLI contract and the runtime-name-policy
exports; it never reads docs/**, gsd-core/workflows/**, commands/** or
agents/**. Every extension to it during this epic was a hand-added assertion
for a surface somebody had already noticed.

A guard written earlier would have red-flagged the very references phases
1b-4b were removing, which is why it lands last.

PART A - THE RESIDUE, INCLUDING WORK I SHIPPED INCOMPLETE

Each site was judged against its ENGLISH counterpart, not on its own:

  README.{ja-JP,ko-KR,pt-BR,zh-CN}.md :9 :24 :46  English README.md has ZERO
                                                  occurrences -> substituted
                                                  "Antigravity CLI, Kimi CLI"
  how-to/execute-a-phase.md:88  x4 locales        fixed in #4728 -> substitute
  how-to/verify-and-ship.md:89  x4 locales        fixed in #4728 -> substitute
  FEATURES.md cross-AI CLI list                   :1419 no Gemini -> DELETE
  FEATURES.md REQ-MULTI-RT-01                     :1709 -> substitute
  FEATURES.md REQ-SKILLS-03                       :1952 -> rewrite
  FEATURES.md REQ-QUOTA-02                        :3256 deleted upstream -> delete
  VERSIONING.md:133                               stale manifest -> see below

The twelve README occurrences were an adversarial reviewer's BLOCKER, and the
reason they survived my own sweep is structural: root-level *.md was outside
the guard's scan set, so the repo's most-read runtime-advertising surface was
invisible to the guard meant to police it. :46 is a live installer-runtime
claim - it tells the reader the installer will offer a runtime that no longer
exists. Checked for the duplicate-name trap before substituting: neither
Antigravity nor Kimi appears anywhere in those four files.

Two of these are mine to own: I fixed the ENGLISH execute-a-phase.md and
verify-and-ship.md in #4728 and left all four mirrors behind. Unfinished work,
not a deferral.

Two more show why "substitute Gemini -> Antigravity" is the wrong default: in
the cross-AI list and REQ-QUOTA-02 English DELETES the name, because
Antigravity was already in the list or the classifier had dropped it.
Substituting would have duplicated a name - the identical trap
ARCHITECTURE.md:24 set in #4728, where English holds Kimi CLI in that slot.

VERSIONING.md:133 is a different and worse defect than translation lag. Under
"Manifest Version Sync" it listed gemini-extension.json as a version-synced
manifest. That file is ABSENT from the repo, and
scripts/sync-manifest-versions.cjs says so in its own comment - "#1928:
gemini-extension.json was removed with the gemini runtime ... it is no longer
a registered manifest" - while VERSIONED_MANIFESTS holds plugin.json,
marketplace.json and vscode/package.json. So the doc named a manifest that
does not exist AND omitted the one that replaced it. Both fixed, verified
against the owning code rather than inferred from the name. The replacement
bullet cites #1942, the issue that actually registered vscode/package.json,
matching the convention of its neighbours.

pt-BR/FEATURES.md is a 77-line stub genuinely lacking two sites, and ko-KR has
no REQ-QUOTA-02 line. Skipped and recorded, never invented.

PART B - THE GUARD

scripts/lint-retired-runtime-name.cjs, modelled on
scripts/lint-legacy-dir-name.cjs - the repo's own precedent for this problem
shape (forbid a retired token, allowlist frozen content, self-exempt via a
split literal, a REPO_ROOT test seam, lib/cli-exit.cjs, exit 0/1).

Case sensitivity IS the mechanism, not an accident. The naive guard - "the
string gemini must not appear" - is WRONG, not merely noisy: that string is
load-bearing across Antigravity's real on-disk contract. A case-sensitive,
standalone, capitalised name works because every legitimate reference is
spelled differently and therefore cannot match: lowercase config homes
(~/.gemini/antigravity, ~/.gemini/config, #3738), lowercase hyphenated model
ids (gemini-2.5-flash-lite), uppercase env vars (GEMINI_API_KEY), and
GEMINI.md. Table-driven, so the next retired runtime costs one row.

THE ALLOWLIST IS THE ENTIRE RISK SURFACE, so it is three tiers, not one. Two
rounds of isolated adversarial review reshaped it; both are recorded in
.gsd/bug/chore-4729-gemini-drift-guard/60-review.json.

ROUND 2 FOUND ONE ROOT CAUSE BEHIND TWO SEPARATE HOLES, and it was mine: both
Tier-1 rules treated the ABSENCE of a runtime word as a GRANT. A veto list can
never be complete, so "no runtime word found" silently exempted every phrasing
nobody had enumerated. Demonstrated: `The installer now offers Gemini 3.`,
`Supported agents include Gemini 3, Kimi, and Cursor.` and three more exited 0,
as did `Suportamos Gemini, no estilo padrao, como runtime de instalacao.` and
`Gemini 兼容,并且是受支持的运行时之一。`, both of which literally contain `runtime`
or `运行时`. The fix was to stop enumerating exceptions and invert the evidence
direction:

  Tier 1(a) - the hook DIALECT Antigravity inherits. Position is
  language-dependent and MEASURED: en Gemini-style/-compatible, ja Gemini
  スタイル, ko Gemini 스타일/호환, zh Gemini 风格 / 与 Gemini 兼容的, pt "no estilo
  Gemini" / "compatível com Gemini" where the qualifier PRECEDES the name. The
  marker must now form an ADJACENT COMPOUND with the name, not merely sit in a
  +/-24-character window - that window let `| Antigravity | Gemini-style hooks
  | Gemini support is live |` exit 0, one legitimate reference licensing a
  fresh live claim 21 characters later. The runtime-word veto is now
  LINE-GLOBAL. Ten real lines legitimately pair a dialect compound with a
  runtime word (`~/.gemini/antigravity-cli` in a table cell, "runtime files"
  in the same sentence); each is an explicit pin rather than a reason to
  loosen the veto for everyone. Measured: widening it surfaced exactly those
  ten and no others.

  Tier 1(b) - the provider/model axis. A version optionally followed by a
  qualifier, including full-width digits and CJK punctuation, AND positive
  model-axis evidence on the line, AND no runtime word. The positive
  requirement is the part that matters: all eight real model-axis lines in the
  repo name a model explicitly, so requiring it costs nothing on the real tree
  while flagging every laundering attempt. It is also the honest resolution of
  the agent/target tension below - rather than guess at an exhaustive veto
  list, stop treating an empty veto as evidence.

  Tier 2 - PINNED OCCURRENCES, now SPAN-SCOPED. A pin excuses only a match
  falling INSIDE an occurrence of its own snippet. Line-level containment let
  `Known provider menu update: Gemini CLI is once again a selectable GSD
  runtime.` and `Install target: Google (Gemini) - choose Gemini CLI as your
  GSD runtime.` both exit 0, because a short snippet elsewhere on the line
  pre-approved a brand-new claim. Span scoping makes short snippets safe:
  `Google (Gemini)` can only ever excuse the match inside those 15 characters.
  A LOAD-TIME validator now requires every pin to contain a retired name, and
  it immediately caught five of MY OWN pins whose snippets sat BESIDE the name
  rather than covering it - each would have shipped permanently inert and
  permanently reported stale. All pins were then reconciled in one pass.

  A pin is also marked used by PRESENCE on the line now, rather than only on
  the Tier-2 branch. Previously a pinned line that a general rule also matched
  never marked its pin used, producing a provably FALSE "no line matches
  pinned snippet" whose printed remedy told the maintainer to delete a pin
  that was still needed.

  Tier 3 - blanket trust, and a new occurrence inside it IS invisible.
  CHANGELOG.md and `.changeset/` - the rendered changelog and its source, one
  surface - plus six append-only directories. All 21 `.changeset/` hits were
  measured to be fragments DESCRIBING the retirement or a fix to it, 464 of
  them under archived/; a fragment can only describe what already shipped and
  is deleted at release, so pinning them would be friction with no signal. The
  cost is stated in the guard's own header rather than hidden.

THE SCAN SET IS NOW EVERY TRACKED *.md FILE (1165 read). The original prefix
list left `.github/`, `.changeset/`, `capabilities/`, `playbooks/` and
`references/` invisible - and `.changeset/*.md` renders into CHANGELOG.md, so a
live claim introduced there was invisible at BOTH ends.

The escape hatch must now carry a justification
(`gsd-allow-retired-runtime-name: <reason>`). A bare marker is rejected: it is
checked first, excuses the whole line, and the failure message advertises it,
so an unexplained one is indistinguishable from a silenced defect.

Plus an anti-vacuity floor counting files actually READ, not files listed - a
candidate count stays healthy-looking even if every read failed.

A FALSE NEGATIVE I INTRODUCED, AND CLOSED

The model-display escape began as a blanket /^ \d/ - "space then a digit" -
which also matched "Install for Gemini 2.5 CLI as a supported runtime.",
laundering a genuine stale-runtime claim through an attached version number.

That was the THIRD appearance of one failure shape in this epic: an exclusion
added to suppress false positives creating a false negative. #4716's sweep
excluded lines matching gemini-[0-9] to spare Google's model ids, and thereby
hid a stale review.models.gemini row whose example value was "gemini-2.5-pro"
ON THE SAME LINE. Round 2 then produced the FOURTH and FIFTH instances, which
is why the fix this time was to invert the rule's evidence direction rather
than to enumerate more exceptions.

The veto is word-anchored for Latin terms - unanchored, case-insensitive "CLI"
matched inside "client" and would have vetoed legitimate model lists - and raw
for CJK terms, where \b is ASCII-word-based and would never fire beside an
ideograph, so anchoring them would silently disable the veto in ja/ko/zh.
"agent" and "target" were deliberately left OUT: both occur throughout
ordinary prose ("AI coding agents (Claude Code, Codex, Gemini 2.5 Pro)"), so
vetoing on them would red correct content instead of catching runtime claims.
The reasoning is in the guard's comment, not just the omission - and Tier
1(b)'s positive-evidence requirement is what makes that omission safe, since
the rule no longer depends on the veto list being complete.

COVERAGE

tests/lint-retired-runtime-name.test.cjs drives the guard through its
GSD_LINT_RETIRED_RUNTIME_REPO_ROOT seam against fixture repos, mirroring
tests/lint-legacy-dir-name.test.cjs. A guard never observed failing is not a
guard, and this epic already shipped one that was vacuous for 2 of its 5
files, so properties are paired against BOTH failure modes - too broad
silently absorbs a future defect, too narrow reds on legitimate content. Floor
boundaries are covered at 149/150/151.

The round-2 reviewer's sharpest point was about that claim, and it was right:
the first matrix's pairing was "true of the properties chosen, not of the
predicate's actual surface" - not one of its twenty properties could see the
dialect adjacency hole, a non-adjacent runtime word, pin shadowing, or an
over-broad pin colliding with a new line. Every one of those is now a
committed regression using the reviewer's own attack line verbatim, and the
local fixture harness went from 14 cases to 35 (PASS=35 FAIL=0).

That harness earned a finding of its own. Its first run reported PASS=2
FAIL=12 with BOTH passes VACUOUS: `git add` has no -q flag on this build, so
nothing staged, every fixture hit the empty-walk error path, and the two
checks that assert an ABSENCE passed off that error path rather than off real
guard logic. A staging failure is now fatal and every absence-asserting check
first proves the walk ran and the expected violation was flagged. Later, one
case failed because its fixture supplied only one of a pinned file's two
approved lines, so the stale-pin check fired correctly - the expectation was
wrong, not the guard. Telling those two apart is the whole value of running a
matrix rather than reasoning about one.

On the two orthogonal reviews: the isolated adversarial pass executed a great
deal of code, across two rounds, against its own fixture repos. The security
pass did NOT - it self-discloses that it verified by reading only, because
node --test is hard-blocked here. Saying so plainly, because "two orthogonal
reviews" without that caveat overstates what the second one established. It
also raised, and I cleared by measurement, a concern that importing
escapeRegex from a gitignored build artifact would break lint:ci on an unbuilt
clone: six other tracked scripts already require that exact path, three of
them already in lint:ci, and .github/workflows/test.yml:192-193 runs
`npm run build:lib` immediately before it for exactly this reason.

Part A has no new test deliberately - those edits are covered by the guard
itself inside lint:ci, and a separate per-locale assertion would duplicate it
and then drift from it. The one exception is the root README case, which IS
pinned: that residue was invisible to the guard rather than merely unasserted,
so the fix is a scan-set change and needs its own regression test.

No mode-bit read-failure fixture was added on purpose: the benches run as
root, where chmod-based IO injection is vacuous, so such a test would assert
nothing.

The test's fixture helpers write throwaway docs/ paths, which trips
lint-docs-guard-registration's reader-name heuristic. Resolved the way that
lint documents - a header `// docs-guard-exempt:` marker plus a baseline entry
- because the fixtures only WRITE scratch data and never read shipped docs;
the baseline was re-confirmed, not merely extended, each time locale and
adversarial fixtures were added. scripts/lib/macos-conformance-tier.generated.cjs
regenerated through its own --write path, since a new test file changes the
count lint:generated-sync reads.

Fixes #4729

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4729): backfill changeset PR number (#4753)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 20:50:26 -04:00