Files
msd-core/src/model-resolver.cts
Tom Boucher 9410f7e6e6 enhance(#3897): ADR-3473 §8.3 rungs 2-4 — runtime marker, derived Codex sandbox, short-form depends_on (#3941)
* test(#3897): failing-first coverage for §8.3 rungs 2-4

ADR-3473 §8.3 has four rungs; #3883/PR #3896 shipped the first. This pins the
other three RED before any fix.

Rung 2 — the install marker has four readers and resolveRuntime is not one.

  resolveRuntime resolves GSD_RUNTIME > config.runtime > 'claude' and reads no
  marker at all, while bin/install.js writes one (#2297) and FOUR hand-rolled
  readInstallRuntimeMarker copies exist: src/model-resolver.cts:65 (cached, with
  test seams), hooks/gsd-agent-isolation-guard.js:112, and TWICE in
  hooks/gsd-cursor-subagent-start.js at :346 and :355. Four copies of one rule.

  Fixtures and seam names mined from PR #3382 rather than re-derived; it
  implemented this rung and was closed "not on the merits".

Rung 3 — the sandbox map, and the fallback that was the real defect.

  Measured across all 35 files in agents/, deriving workspace-write iff tools:
  declares Write or Edit:

    - all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly,
      zero disagreements — the map carries nothing the contract does not
    - 24 roles fall through `|| 'read-only'`, of which 16 declare Write or Edit

  So the map is redundant and the silent fallback is the defect. The maintainer
  chose to derive but hold those 16 at read-only pending the question of whether
  Codex enforces sandbox_mode or merely advises; HALT.md records it.

  T20 asserts the emitted sandbox_mode PER ROLE against a captured baseline, not
  in aggregate — an aggregate passes while one role silently widens, which is
  the proxy-instead-of-identity shape this repo names. T24 and T25 fail on a
  stale hold, so the hold list cannot rot into the subset map being deleted.

Rung 4 — shortFormToId, recovered rather than invented.

  I nearly reported this as another wrong §8.3 claim: `git log -S shortFormToId`
  returns only documentation commits. That was the wrong instrument. Direct
  inspection of sdk/src/query/phase.ts at 11918dcc3^ shows five occurrences, and
  the tests match that code rather than a guess at its semantics — including
  first-write-wins on a duplicate short form.

  T43 asserts at the consumer's output: the emitted `waves` map from the real
  CLI, which pre-fix collapses to {"1":[...]} because every short-form edge is
  dropped. A unit assertion on resolveDependencyId would have passed throughout
  this defect's life.

Observed RED, this tree:
  rung 2   11/11 fail — no marker rung, no seams
  rung 3   T23,T24,T25,T26,T30 fail; T28 fails (validate agents passes a TOML
           whose sandbox_mode disagrees — it checks presence only)
  rung 4   T42,T44 fail; T43,T49 fail with waves collapsed to a single wave 1

Green and staying green: T20/T21/T22/T27 as captured baselines, #3885's
unresolvable-token warning and wave-verdict suppression, and #3785's
display-mapping passthrough. If the third tier over-reaches, those go red — that
is their job.

Disclosed weakness: T45 (a canonical id with no dash is not short-form indexed)
cannot be isolated behaviorally, because planMap always masks it. It is a
non-crash boundary pin, weaker than the other rows, and is recorded as such
rather than presented as equivalent.

Design:      .gsd/phase/feat-3897-adr3473-83-rungs/40-design.md
Test matrix: .gsd/phase/feat-3897-adr3473-83-rungs/50-test-matrix.md
Decision:    .gsd/phase/feat-3897-adr3473-83-rungs/HALT.md

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3897): §8.3 rungs 2-4 — one marker reader, a derived sandbox, the third depends_on tier

ADR-3473 §8.3 has four rungs. #3883/PR #3896 shipped the first. These are the
other three.

Rung 2 — the install marker had four readers, and resolveRuntime was not one.

  resolveRuntime resolved GSD_RUNTIME > config.runtime > 'claude' and read no
  marker, while bin/install.js writes one (#2297) and four hand-rolled
  readInstallRuntimeMarker copies existed: src/model-resolver.cts (cached, with
  seams), hooks/gsd-agent-isolation-guard.js, and twice in
  hooks/gsd-cursor-subagent-start.js.

  model-resolver's was already the house idiom, so it was promoted rather than
  replaced: src/runtime-slash.cts now owns it, and model-resolver plus both
  hooks delegate. The hooks reach it through ensureRuntimeBuild(), the seam
  lint-hooks-runtime-build-seam enforces. No import cycle existed - checked
  both directions before moving anything.

  The marker is the THIRD rung: env > project config > marker > 'claude'.

  N1 was checked rather than assumed, and my first reading of it was wrong. A
  marker holding an unknown name comes back essentially verbatim, which looked
  like a validation gap. Measured against the env rung with the same inputs -
  including "../../etc/passwd" and "claude;rm -rf /" - the two are identical,
  because they share resolveRuntimeNameFromCandidates. N1 asks for exactly that,
  and it is met. The residual (the shared normalizer normalizes shape, it does
  not validate against the known-runtime set) is pre-existing on the env rung
  and plausibly deliberate, since a new runtime should not need a code change.
  The marker also does not widen the trust boundary in any real sense: it lives
  inside the install tree beside the code, so anyone who can write it can write
  runtime-slash.cjs itself.

Rung 3 — the map was redundant; the silent fallback was the defect.

  Measured across all 35 files in agents/, deriving workspace-write iff tools:
  declares Write or Edit: all 11 CODEX_AGENT_SANDBOX entries derive to their
  mapped value exactly, zero disagreements. The map carried nothing the contract
  did not already have, so it is DELETED rather than reconciled. What was
  actually broken is `|| 'read-only'`, which silently under-granted 24 of 35
  roles.

  16 of those 24 declare Write or Edit and would widen under derivation. Per the
  maintainer's decision (HALT.md), they are held at read-only pending the
  question of whether Codex enforces sandbox_mode or merely advises. Emitted
  TOML is therefore byte-identical for all 35 roles - asserted per role, not in
  aggregate, because an aggregate passes while one role silently widens.

  The hold list self-invalidates. A hold whose role no longer derives broader
  fails, and so does a hold naming a role with no agents/<name>.md. Without
  that it would rot into exactly the hand-maintained subset map being deleted,
  and this commit's own ledger claim would become false over time. Both cases
  were proved by injecting them and watching them throw.

  Two committed tests asserted the deleted map's existence and contents. They
  were pinning the thing being removed, so the tests moved rather than the
  production code: the 11 role-value pairs survive as a test-local
  PRE_3897_CODEX_AGENT_SANDBOX baseline, and the assertions now drive the real
  derivation against real agents/*.md. The coverage is preserved; only its
  source moved out of production code.

  validate agents gains checkCodexSandboxPosture, mirroring the existing
  checkCodexModelPosture: each installed TOML's sandbox_mode must equal the
  role's expected value, failing with role, expected and found. It previously
  checked file presence and manifest completeness only, so a TOML whose
  sandbox_mode disagreed passed.

Rung 4 — shortFormToId, recovered rather than invented.

  I nearly reported this as another wrong §8.3 claim: git log -S returns only
  documentation commits. Wrong instrument. sdk/src/query/phase.ts at 11918dcc3^
  carries five occurrences, and the implementation here matches that code rather
  than a guess at its semantics - including first-write-wins on a duplicate
  short form, deterministic from the sorted plan order.

  It resolves the bare plan number: depends_on: ["01"] now reaches
  26-01-auth-hardening. That is a control-flow change, not a diagnostic one -
  plans that silently collapsed into a single wave 1 now execute in their
  declared waves, and execute-phase.md consumes those wave values.

  In-phase only, by construction: the map is built from this phase's rawPlans,
  so a same-named short form in another phase does not resolve.

  #3785's display-mapping passthrough and #3885's unresolvable-token warning and
  wave-verdict suppression are untouched and stay green. If the third tier had
  over-reached, those are what would have caught it.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): close a fail-open I introduced, and wire the posture check to its command

Two blockers from review. Both are mine, and one is a security regression my own
change created.

1. A held role could escape its hold by editing its own frontmatter.

  The Codex install loop set the sandbox identity from the agent's frontmatter
  `name:` field rather than from its filename, so the hold lookup keyed off a
  value the file itself declares:

    deriveCodexSandboxMode('gsd-doc-writer',   <real file>)          -> read-only
    deriveCodexSandboxMode('gsd-doc-writer-x', <same file, name: edited>) -> workspace-write
    deriveCodexSandboxMode('GSD-Doc-Writer',   <same file, name: recased>) -> workspace-write

  What makes this a blocker rather than a nit is the DIRECTION. The deleted
  CODEX_AGENT_SANDBOX map had the identical lookup-key quirk, but it was an
  allowlist: an unmatched key fell back to read-only, which is safe. The new
  scheme derives workspace-write from the tool contract and uses the hold as a
  subtraction, so the same mismatch fails OPEN. I converted a fail-closed quirk
  into a fail-open one and did not notice; the isolated reviewer proved it by
  execution.

  Neither safety net caught it. validateCodexSandboxHolds only checks that
  <key>.md exists, never that a file's derived identity matches its key.
  checkCodexSandboxPosture looks the canonical source up by the installed TOML's
  filename, finds nothing for a renamed agent, and treats it as a custom
  non-roster agent — silently no violation.

  The identity is now the FILENAME STEM, which is what validateCodexSandboxHolds
  already validates and what an attacker editing frontmatter cannot change
  without renaming the file — at which point the existing validator catches it.
  The lookup is case-insensitive so a recase does not slip past either. The
  frontmatter name still drives the TOML body and filename, unchanged; only the
  sandbox identity moved.

  All 35 roster files were checked: name matches filename stem everywhere, so a
  stricter "they must agree or throw" invariant would have been safe against real
  content. It is deliberately NOT added — it would abort an install on a tampered
  file where emitting a correctly-derived read-only TOML is the safer outcome.
  Recorded as a fork rather than decided silently.

2. checkCodexSandboxPosture was exported and never called.

  cmdValidateAgents (src/verify.cts) called checkAgentsInstalled and
  checkCodexModelPosture only; grep for the sandbox check in that file returned
  nothing. So criterion 3 — "validate agents fails on semantic drift, not only on
  missing files" — was unmet, and `validate agents` behaved exactly as before.
  That is ADR-3473 Decision 2's named shape: a declared policy with no executor.

  It also meant the T28 test asserted at the helper's return value while the
  COMMAND stayed broken — the ADR-3180 Decision 4(b) failure this epic exists to
  close, committed by me while enforcing it elsewhere in the same epic.

  Now wired as an additive `sandbox_posture` field beside `codex_posture`,
  following the sibling precedent exactly. Drift is report-only, not a non-zero
  exit, because that is what checkCodexModelPosture does — two sibling posture
  checks disagreeing about whether a violation is fatal would be its own defect.
  The choice is recorded in a comment rather than left implicit. A consumer-output
  test now drives the real CLI and asserts on the emitted JSON, and was shown
  failing before the wiring and passing after.

Also corrected a stale artifact: the design's Known limit L1 still claimed rung 3
was not in this deliverable, written while it was halted and false once the
maintainer unblocked it.

Verified after both fixes: the three bypass probes all return read-only, the
per-role table is 35/35 byte-identical, and both hold self-invalidation cases
still throw.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): the marker rung, the derived sandbox, and the bare plan-number depends_on

Reference: the runtime precedence ladder in docs/CLI-TOOLS.md gains the install
marker rung; docs/COMMANDS.md documents validate agents' new sandbox_posture
field; docs/reference/plan-md.md documents that depends_on accepts the bare plan
number.

Explanation: a docs/features fragment keyed id 3897, so it cannot collide with a
concurrent PR hand-allocating a section number, regenerated into FEATURES.md.

ADR-3473 §8.3 gains an ANSWER blockquote in the document's own correction style,
recording what was measured and built against the section's 2026-08-26 correction
- including the qualification that checkAgentsInstalled itself still checks
presence only, and the semantic assertion lives in a sibling wired into validate
agents rather than folded into it.

No how-to. Both user-visible changes are zero-step: a non-Claude install resolving
its own runtime, and plans executing in their declared waves, both happen without
the user doing anything. docs/how-to/control-the-reported-host-runtime.md covers a
DIFFERENT ladder (resolveReportedRuntime / agent_runtime) that this change does
not touch, and was deliberately left alone rather than edited by association.

No tutorial - nothing multi-step to walk through. docs/AGENTS.md unchanged: it
documents Claude-side tools frontmatter, never Codex sandbox_mode, and the
emitted tools contract did not change.

The prompt layer documents depends_on only by example, not by schema, so nothing
there needed editing - and few-shot-examples/plan-checker.md already showed
depends_on: ['01'], which now actually resolves.

Translated copies of plan-md.md are untouched; the project treats translations as
community-maintained.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): move the sandbox derivation out of the installer, off the install path, and off a third parser

The full suite came back with 26 failures across four files. Three distinct
causes, mapped individually rather than assuming the first explained the rest.

A. Requiring bin/install.js printed the GSD banner to stdout and corrupted
   `validate agents` JSON.

     Unexpected token '', "[36m   ██"... is not valid JSON

   checkCodexSandboxPosture reached deriveCodexSandboxMode by lazily requiring
   bin/install.js, whose module load prints the ASCII banner. So the command
   emitted banner bytes before its JSON and every JSON consumer broke, including
   ten tests that predate this branch. src/ reaching into bin/ was backwards
   layering that happened to also be loud.

   The derivation now lives in src/codex-agent-toml.cts - the existing Codex TOML
   domain module, no new module and no six-gate ripple - and both bin/install.js
   and src/agent-install-check.cts import it. One owner, which is §8.3's rule
   applied to the fix for §8.3.

B. The stale-hold throw fired on a legitimate partial source dir, and masked a
   security assertion.

   validateCodexSandboxHolds treated "this hold's .md is absent from the install
   SOURCE dir" as a stale hold and threw. A test fixture, or any partial install
   source, legitimately contains a couple of agents. Worse, it threw BEFORE the
   path-escape check, so a test asserting that a `../../evil` frontmatter name is
   rejected got my unrelated error instead of the traversal rejection it was
   written for. A fail-closed check of mine was hiding a real security check.

   The "no stale holds, shrink-only" invariant is a property of the repo's
   canonical agents/ roster, not of whatever directory an install happens to read.
   It is off the runtime path and enforced where it belongs, in the tests that
   already existed for it. A partial source dir now installs cleanly, and the
   evil-name case throws with its own escapes-configHome message again.

C. T8 depended on ambient process.env state.

   The marker/env parity assertion round-tripped through live process.env. It now
   compares against resolveExplicitRuntime's already-exported dependency-injection
   parameter - deterministic and hermetic, same claim. Proven still falsifiable
   rather than assumed: with the marker rung's normalization temporarily bypassed
   the two rungs diverge ("codex\n../../etc/passwd" vs "codex-../../etc/passwd")
   and the assertion fails, then passes again once reverted.

One correction folded in along the way. The first version of the move added
private _extractFrontmatterAndBody/_extractFrontmatterField helpers to
codex-agent-toml.cts - a THIRD copy of frontmatter extraction, where the graph
already shows two (bin/install.js:2348, runtime-artifact-conversion.cts:893).
Adding a third inside the epic whose thesis is one implementation per rule is not
defensible. deriveCodexSandboxMode no longer parses anything: it takes
(identity, toolsValue) and each caller supplies the tools value using the
extractor it already has. Both helpers are deleted. The identity argument is
still the filename stem, so the fail-open fix is untouched.

Verified after all three: `validate agents --raw` emits parseable JSON with no
banner and both posture fields; the four hold-bypass probes still return
read-only; the per-role table is 35/35 byte-identical at 26 read-only / 9
workspace-write; the hold list is still 16.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): drop a dev-only transitive dep, make the derivation total, retire a stale fallback test

Suite down to 7 failures from 26. Three more causes, mapped individually.

A. My extractor import dragged in a script that does not exist in an installed
   tree.

     Cannot find module '../../../scripts/fix-slash-commands.cjs'

   Chain: src/agent-install-check.cts imported runtime-artifact-conversion.cjs,
   which requires command-roster.cjs, whose line 36 requires
   ../../../scripts/fix-slash-commands.cjs. That path exists in the repo and not
   in an install, so every test exercising a synthetic install dir died at module
   load. I picked that extractor for convenience without checking what it pulls
   in - the same mistake that produced the banner bug, one layer further out.

   agent-install-check now uses a single-purpose extractToolsLine on
   codex-agent-toml.cts. That is deliberately NOT a general frontmatter parser:
   we deleted those helpers a commit ago for good reason, and this reads one
   line. Verified from outside the repo root that requiring either module prints
   nothing and does not throw.

B. A test pinned the deleted name-based fallback.

   'defaults unknown agents to read-only' called generateCodexAgentToml with a
   fixture declaring tools: Read, Write, Edit. Under derivation an unknown agent
   with a writing contract correctly derives workspace-write - design row S6, a
   new writing role gets the contract, not the pin. The behavior it asserted was
   the silent fallback this rung deleted; identity no longer decides the sandbox.

   Replaced with two rows rather than a flipped string: no tools declared ->
   read-only (absence is not a grant), and Write/Edit declared -> workspace-write.
   Strictly more coverage than the row it replaces.

C. The stale-hold check still threw per derivation call.

   Last commit took the roster-existence check off the install path, but
   deriveCodexSandboxMode itself still threw when a hold's role did not derive
   broader FOR THE CONTENT IT WAS HANDED - so it fired on any synthetic fixture
   for a held role.

   The throw is gone, and it cost nothing: if a held role's content does not
   derive broader, the hold pins read-only and derivation returns read-only
   anyway, so the hold is a no-op and there is nothing to fail about. The
   staleness invariant is a property of the real agents/ roster, and
   validateCodexSandboxHolds still enforces it there - confirmed against the real
   roster after the change, not assumed.

   deriveCodexSandboxMode is now total: every (identity, toolsValue) including
   undefined and null returns read-only or workspace-write, never throws.

Verified: validate agents emits parseable JSON; the four hold-bypass probes
return read-only; the per-role table is 35/35 at 26 read-only / 9
workspace-write.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): put the rung-3 decision in the shipped docs instead of pointing at an ignored path

The ADR entry and the feature fragment both ended their rung-3 explanation with
"see .gsd/phase/feat-3897-adr3473-83-rungs/45-decision-rung3-sandbox.md". That
directory is gitignored (.gitignore:55), so the rationale for holding 16 roles at
read-only was reachable only from the machine that produced it. A reader of the
ADR got a pointer to nothing.

Both now carry the reasoning inline: the criterion asks both that the sandbox
derive from the declared tool contract and that no role gain a broader sandbox,
and those cannot both hold, because a faithful derivation widens 16 roles the
deleted map never listed and that fell through its silent read-only default. The
resolution is derive-and-hold - the derivation owns the rule now, each hold is
released as its enforcement question is answered, and a hold is reversible where
a widened sandbox that turns out to be enforced is not.

Checked before assuming this was a defect class: CONTEXT.md cites
.gsd/phase/<slug>/40-design.md as its standard Design: provenance line in eight
module entries, and four other shipped docs do the same. Citing a phase artifact
is an established convention here, so those are left alone. What was wrong was
specific to these two: they put load-bearing rationale behind the pointer instead
of provenance.

docs/FEATURES.md regenerated from the fragment via scripts/gen-features.cjs
rather than hand-edited.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3897): close a fail-open, stop a silent mis-resolution, and read a declaration as a declaration

Two orthogonal reviews on the shipped sha. Three of the findings are the same
failure class this epic exists to close, committed inside it.

1. BLOCKER - the sandbox was decided for one identity and applied to another.

   bin/install.js derived sandbox_mode for the filename stem and then wrote the
   result to `${name}.toml`, where name comes from the file's own frontmatter.
   Make the two disagree and a HELD role's artifact goes wide:

     rename gsd-doc-writer.md -> gsd-doc-writer-v2.md, keep name: gsd-doc-writer
       -> stem is unheld, derives workspace-write, lands on gsd-doc-writer.toml
     add any gsd-*.md whose frontmatter name: is a held role
       -> clobbers that role's toml with workspace-write

   Both emit read-only on origin/next, because the deleted map was an allowlist
   and a miss fell back safe. This is a regression my change introduced. The
   previous review round moved the HOLD KEY off frontmatter to the filename stem
   and left the OUTPUT PATH on frontmatter; my own comment at install.js:6985
   calls that value attacker-editable, four lines above the line that uses it as
   the filename.

   The decision is now made over BOTH candidate identities, most-restrictive
   wins: if either the stem or the emitted name is held, the mode is read-only.

2. MAJOR - hold matching was toLowerCase() only, so confusables escaped.

   Turkish dotted/dotless i, fullwidth, NFD, trailing space/NBSP/dot/newline,
   ./ and ../agents/ all slipped the hold and emitted workspace-write.
   Identities are now basenamed, trimmed of NBSP/zero-width/control characters,
   NFKC-normalized and lowercased - and anything still carrying a character
   outside [a-z0-9._-] is treated as suspicious and derives read-only. We do not
   enumerate confusables; every shipped roster file is ASCII, so refusing to
   widen on an identity we cannot recognize is fail-closed with no false
   positives on real content.

3. MAJOR - the short-form depends_on tier mis-resolved SILENTLY.

   shortFormToId keyed on the last dash-segment of any canonical id with no
   constraint that it is a plan number, so a phase holding 09-FIX-auth-PLAN.md
   made depends_on: ["auth"] bind at wave 2 with zero warnings. This is the
   worst shape in the epic: the unresolvable-token warning fires on a DROPPED
   token, so a MIS-RESOLVED one is invisible and the tool reports a confident
   wave assignment built from a wrong edge. A wrong edge is worse than a missing
   one.

   The segment must now match /^\d+$/, which is exactly the contract
   docs/reference/plan-md.md already documents. This tier was recovered verbatim
   from the retired SDK lineage, which carried the same defect; we are
   deliberately NOT preserving it bug-for-bug, and the comment says so, so the
   next reader does not "restore" it.

4. MAJOR - the derivation was reading a declaration as an absence.

   extractToolsLine read one line, so a YAML list-form tools: block returned only
   its first item. Two roster files use list form, and gsd-nyquist-auditor
   declares Write and Edit there - parsed as "- Read", found no write tool, and
   emitted read-only. Rung 3's headline claim is that sandbox_mode derives from
   the declared tool contract; that claim was false for 2 of 35 roles and
   materially wrong for 1. Reading a declaration as an absence is the silent-drop
   class this epic exists to close.

   Renamed extractToolsValue and taught it both shapes. gsd-nyquist-auditor now
   derives workspace-write and joins CODEX_SANDBOX_HOLDS as its 17th entry, per
   the standing derive-and-hold decision - so emitted TOML stays byte-identical
   at 26 read-only / 9 workspace-write while the hold list finally records every
   role that would widen. A previous pass declined this fix because it moved the
   count; that inverts the priority. Byte-identity is preserved THROUGH the hold,
   not by leaving a parser broken.

   Divergence check, because this is where that bug hides: both paths feeding
   sandbox derivation - install.js's emitter and checkCodexSandboxPosture - now
   route through the one extractor. The tools readers in
   runtime-artifact-conversion and install.js's other frontmatter call sites
   serve Claude-side emission and do not feed sandbox derivation.

Also fixed, each real: the posture check's `found` used a naive whole-file regex
where its own sibling uses the block-aware scanner, so prose inside
developer_instructions produced a false violation; `found` skipped
truncatePostureValue and leaked a 300-char value into validate agents output;
deriveCodexSandboxMode's absolute never-throws claim was false for an object with
a throwing toString; T49 could not falsify cross-phase leakage (its target phase
had its own 01, so a globally-scoped map passed too); T20/N6 iterated a hardcoded
table and pinned the FIXTURE size, so a 36th agent would be silently unchecked;
three tests reimplemented the code they were testing instead of importing it; and
T2-T4 deleted GSD_RUNTIME without restoring it.

Verified: hold list 17, gsd-nyquist-auditor derives workspace-write unheld and
emits read-only held, roster 35/35 at 26/9, depends_on ["auth"] no longer
resolves while ["01"] still does, both identity-bypass cases and every confusable
vector emit read-only.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3897): the hold list is 17, and the reason the 17th was missing

The count read 16 because the derivation could not read the declaration it
claimed to derive from: the tools reader was single-line, so a YAML list-form
tools: block returned only its first item and gsd-nyquist-auditor's declared
Write and Edit were read as an absence.

Both the ADR entry and the feature fragment now carry the corrected count and the
reason for it, rather than a silently updated number. Deriving from a declaration
you cannot parse is not deriving, and a flattering count is worse than a wrong
one because it looks settled.

docs/FEATURES.md regenerated from the fragment.

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3897): backfill changeset pr number

Refs #3897

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 15:19:01 -04:00

935 lines
41 KiB
TypeScript

/**
* Model Resolver — Model and effort resolution policy
*
* ADR-857 rollout phase 2f: extracted from core.cts (issue #888).
* Owns model and effort resolution policy: resolves the model, runtime tier,
* planning granularity, reasoning effort, and fast-mode for a given agent by
* reading project config and resolving against the model profiles and catalog.
* Behaviour is preserved byte-for-behaviour from the prior location; only
* the module boundary moved. The core.cjs re-export spine was retired in
* epic #1267; callers import resolvers from model-resolver.cjs directly.
*
* Dependencies (leaf modules only):
* - node:fs / node:path (read the per-install .gsd-runtime marker + project config for the #2297 omit gate)
* - ./runtime-name-policy.cjs (resolveRuntimeNameFromCandidates — canonicalize the active runtime)
* - ./planning-workspace.cjs (planningDir — workstream/project-aware project-config path)
* - ./config-loader.cjs (loadConfig)
* - ./configuration.cjs (CONFIG_DEFAULTS as CANONICAL_CONFIG_DEFAULTS)
* - ./model-profiles.cjs (MODEL_PROFILES, AGENT_TO_PHASE_TYPE, AGENT_DEFAULT_TIERS, VALID_AGENT_TIERS, nextTier)
* - ./model-catalog.cjs (MODEL_ALIAS_MAP, RUNTIME_PROFILE_MAP, PROVIDER_PRESETS, VALID_TIERS,
* CLAUDE_AGENT_ALIASES — re-exported below for back-compat, #3241)
*/
// eslint-disable-next-line @typescript-eslint/no-require-imports
import configLoaderModule = require('./config-loader.cjs');
const { loadConfig } = configLoaderModule;
// ─── Configuration Module (for CANONICAL_CONFIG_DEFAULTS used by effort/fast_mode resolvers) ─
import { CONFIG_DEFAULTS as CANONICAL_CONFIG_DEFAULTS } from './configuration.cjs';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import modelProfiles = require('./model-profiles.cjs');
const { MODEL_PROFILES, AGENT_TO_PHASE_TYPE, AGENT_DEFAULT_TIERS, VALID_AGENT_TIERS, nextTier } = modelProfiles;
import { MODEL_ALIAS_MAP, RUNTIME_PROFILE_MAP, PROVIDER_PRESETS, VALID_TIERS, CLAUDE_AGENT_ALIASES, mergeEffortTierDefaults } from './model-catalog.cjs';
import fs from 'node:fs';
import path from 'node:path';
import { resolveRuntimeNameFromCandidates } from './runtime-name-policy.cjs';
import {
readInstallRuntimeMarker,
_setInstallRuntimeMarkerForTests,
_resetInstallRuntimeMarkerCacheForTests,
} from './runtime-slash.cjs';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import planningWorkspaceMod = require('./planning-workspace.cjs');
const { planningDir } = planningWorkspaceMod;
// ─── #2297: per-install runtime identity for the resolve_model_ids:"omit" gate ─
//
// The installer writes `resolve_model_ids:"omit"` into the SHARED
// ~/.gsd/defaults.json for every runtime that lacks native model aliases (#1156).
// Because that file is machine-wide, a non-Claude install would otherwise poison
// a Claude no-project resolution into returning '' — silently defeating Claude's
// adaptive tier aliases. The "omit" must therefore apply only when a runtime that
// genuinely lacks native aliases is the one resolving.
//
// In a no-project session there is no `.planning/config.json` (so config.runtime
// is null) and GSD_RUNTIME is not exported by gsd-core, so the only reliable
// current-runtime signal is the per-install marker the installer co-locates next
// to VERSION at <install>/gsd-core/.gsd-runtime (this file's dir is
// <install>/gsd-core/bin/lib). Precedence for the gate: project config.runtime →
// GSD_RUNTIME env (manual/CI override + test seam) → install marker → 'claude'.
//
// `claude` is currently the ONLY runtime with nativeModelAliases:true; a
// registry-parity test guards this set so a future alias-capable runtime fails
// loudly here instead of silently omitting.
const RUNTIMES_WITH_NATIVE_ALIASES: ReadonlySet<string> = new Set(['claude']);
// #3897 rung 2: the marker reader + its cache and test seams were promoted to
// the canonical owner, `runtime-slash.cts` (imported above) — this module now
// consumes that single implementation instead of holding its own copy. N5:
// behaviour and the seam contract are unchanged by the move; the re-exports
// below (`export =` at the bottom of this file) preserve every existing
// caller's `require('./model-resolver.cjs')` surface byte-for-behaviour.
// The runtime whose install is actually resolving, canonicalized so an alias or
// case variant (e.g. "claude-code"/"Claude") cannot defeat the native-alias
// check below (#2297 review). Precedence mirrors resolveRuntime()
// (runtime-slash.cts): GSD_RUNTIME env → project config.runtime → per-install
// .gsd-runtime marker → 'claude'.
function resolveActiveRuntime(config: Record<string, unknown>): string {
return resolveRuntimeNameFromCandidates(
process.env['GSD_RUNTIME'],
config['runtime'],
readInstallRuntimeMarker(),
) || 'claude';
}
// Did the PROJECT's own config (root `.planning/config.json` or the active
// workstream/project override) explicitly set resolve_model_ids to "omit"?
// Project config takes precedence over the shared ~/.gsd/defaults.json (#2297
// out-of-scope guard + #2517 finding #4): an explicit project "omit" is honored
// regardless of runtime, whereas an "omit" that came only from the global
// defaults is ignored by native-alias runtimes. Workstream/project-scope aware
// via planningDir (mirrors loadConfig's precedence: workstream value wins over
// root); a plain read avoids loadConfig's normalization side effects.
function projectExplicitlySetsOmit(cwd: string): boolean {
const wsDir = planningDir(cwd);
const rootDir = path.join(cwd, '.planning');
const layers = wsDir === rootDir ? [rootDir] : [wsDir, rootDir]; // workstream > root
for (const dir of layers) {
try {
const parsed = JSON.parse(fs.readFileSync(path.join(dir, 'config.json'), 'utf8')) as Record<string, unknown>;
const value = parsed?.['resolve_model_ids'];
// First layer that sets the key wins (matches loadConfig's deep-merge
// precedence). A layer that omits the key falls through to the next.
if (value !== undefined) return value === 'omit';
} catch {
// Absent/unreadable layer — try the next.
}
}
return false;
}
// ─── Model alias resolution ───────────────────────────────────────────────────
interface TierEntryResolved {
model: string;
reasoning_effort?: string;
[key: string]: unknown;
}
interface ResolveTierEntryOpts {
runtime: string | null | undefined;
tier: string | null | undefined;
overrides: Record<string, unknown> | null | undefined;
}
/**
* #2517 — Resolve the runtime-aware tier entry for (runtime, tier).
*/
function resolveTierEntry({ runtime, tier, overrides }: ResolveTierEntryOpts): TierEntryResolved | null {
if (!runtime || !tier) return null;
const runtimeMap = RUNTIME_PROFILE_MAP as unknown as Record<string, Record<string, Record<string, unknown>>>;
const builtin = runtimeMap[runtime]?.[tier] || null;
const overridesMap = overrides as Record<string, Record<string, unknown>> | null | undefined;
const userRaw = overridesMap?.[runtime]?.[tier];
let userEntry: Record<string, unknown> | null = null;
if (userRaw) {
userEntry = typeof userRaw === 'string' ? { model: userRaw } : (userRaw as Record<string, unknown>);
}
if (!builtin && !userEntry) return null;
return { ...(builtin || {}), ...(userEntry || {}) } as TierEntryResolved;
}
/**
* Convenience wrapper used by resolveModelInternal.
*/
function _resolveRuntimeTier(config: Record<string, unknown>, tier: string): TierEntryResolved | null {
return resolveTierEntry({
runtime: config['runtime'] as string | null | undefined,
tier,
overrides: config['model_profile_overrides'] as Record<string, unknown> | null | undefined,
});
}
// Reverse of the Claude tier-default IDs, plus the Fable alias which Claude
// Code's Agent tool accepts but which is not a GSD model-profile tier (#1133).
const CLAUDE_POLICY_ID_TO_ALIAS: Record<string, string> = {
...Object.fromEntries(
Object.entries(MODEL_ALIAS_MAP)
.filter((e): e is [string, string] => typeof e[1] === 'string')
.map(([aliasName, id]) => [id, aliasName]),
),
'claude-fable-5': 'fable',
};
// CLAUDE_AGENT_ALIASES moved to ./model-catalog.cts (#3241) — imported above
// and re-exported below for back-compat (bin/install.js:474,
// tests/codex-config.test.cjs:24 depend on the name being on this module).
// Dedupe stderr warnings so repeated agent resolutions don't spam (#1133).
const _modelPolicyUnmappableWarned = new Set<string>();
function warnModelPolicyUnmappable(agentType: string, policyModel: string, tier: string): void {
const key = `${agentType}::${policyModel}::${tier}`;
if (_modelPolicyUnmappableWarned.has(key)) return;
_modelPolicyUnmappableWarned.add(key);
// MUST go to stderr — resolve-model's JSON result is parsed from stdout.
process.stderr.write(
`gsd: warning — model_policy resolved "${policyModel}" for ${agentType}, ` +
`but it has no Claude agent alias; using "${tier}" instead.\n`,
);
}
// Test-only: reset the model_policy warn-dedupe cache between cases (#1133).
function _resetModelPolicyWarningCacheForTests(): void {
_modelPolicyUnmappableWarned.clear();
}
// Dedupe stderr warnings for unmappable model_overrides Claude IDs (#2041).
const _modelOverrideUnmappableWarned = new Set<string>();
function warnModelOverrideUnmappable(agentType: string, overrideValue: string): void {
const key = `${agentType}::${overrideValue}`;
if (_modelOverrideUnmappableWarned.has(key)) return;
_modelOverrideUnmappableWarned.add(key);
// Cap emission length so an oversized or secret-shaped value cannot leak in
// full to stderr/logs (#2041 security review). MUST go to stderr — resolve-
// model's JSON result is parsed from stdout.
const safe = overrideValue.length > 64 ? overrideValue.slice(0, 64) + '…' : overrideValue;
process.stderr.write(
`gsd: warning — model_overrides value "${safe}" for ${agentType} ` +
`has no Claude agent alias; falling through to tier resolution.\n`,
);
}
// Test-only: reset the model_overrides warn-dedupe cache between cases (#2041).
function _resetModelOverrideWarningCacheForTests(): void {
_modelOverrideUnmappableWarned.clear();
}
/**
* #2041 — Map a `model_overrides` value to its Claude Agent-tool alias on the
* claude runtime, mirroring the `model_policy` path (#1144). Claude Code's
* Agent tool `model` parameter documents only tier aliases (opus/sonnet/haiku/
* fable); a full Claude model ID returned verbatim is silently dropped by the
* spawner. Returns the value to return verbatim, or null to signal "fall
* through to normal tier/dynamic-routing resolution" (used when a Claude full
* ID has no alias — matches model_policy's warn-and-fall-through). Non-Claude
* runtimes and non-Claude values always pass through verbatim.
*
* Hardening (code+security review): a `typeof` guard preserves the pre-fix
* no-crash behavior if a malformed config surfaces a non-string value, and an
* `Object.hasOwn` lookup defeats `__proto__`/`constructor` lookups on the plain
* object literal so those reserved keys cannot return a truthy non-string.
*/
function mapClaudeOverrideForRuntime(
override: string,
configRuntime: string | null | undefined,
agentType: string,
): string | null {
// Defensive: model_overrides is typed Record<string,string> but a malformed
// config could surface a non-string; pass through verbatim (preserving the
// pre-fix no-crash behaviour) and let the downstream Agent tool reject it.
if (typeof override !== 'string') return override;
const onClaude = !configRuntime || configRuntime === 'claude';
if (!onClaude) return override;
// Object.hasOwn guards against __proto__/constructor returning a truthy
// non-string from the plain object literal (#2041 security review).
if (Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, override)) {
return CLAUDE_POLICY_ID_TO_ALIAS[override];
}
if (CLAUDE_AGENT_ALIASES.has(override)) return override;
if (override.startsWith('claude-')) {
warnModelOverrideUnmappable(agentType, override);
return null;
}
return override;
}
/**
* #49 — Provider-neutral model policy preset resolution.
*/
function resolveModelPolicy(policy: Record<string, unknown> | null | undefined, tier: string | null | undefined): string | null {
if (!policy || typeof policy !== 'object') return null;
if (!tier) return null;
const runtime = policy['runtime'];
const rtOverrides = policy['runtime_tiers'];
if (runtime && typeof runtime === 'string' && rtOverrides && typeof rtOverrides === 'object') {
const rtOverridesMap = rtOverrides as Record<string, unknown>;
if (Object.hasOwn(rtOverridesMap, runtime)) {
const runtimeEntry = rtOverridesMap[runtime];
if (runtimeEntry && typeof runtimeEntry === 'object' && Object.hasOwn(runtimeEntry, tier)) {
const raw = (runtimeEntry as Record<string, unknown>)[tier];
if (raw != null) {
const entry = typeof raw === 'string' ? { model: raw } : (raw as Record<string, unknown>);
if (entry && entry['model']) return entry['model'] as string;
}
}
}
}
const provider = policy['provider'];
if (!provider || typeof provider !== 'string') return null;
if (provider === 'generic' || provider === 'custom') {
const TIER_TO_POLICY_KEY: Record<string, string> = { opus: 'high', sonnet: 'medium', haiku: 'low' };
const policyKey = TIER_TO_POLICY_KEY[tier];
if (!policyKey) return null;
const v = policy[policyKey];
return (v && typeof v === 'string') ? v : null;
}
const presetsMap = PROVIDER_PRESETS as Record<string, Record<string, Record<string, { model: string } | null>>>;
if (!Object.hasOwn(presetsMap, provider)) return null;
const presetForProvider = presetsMap[provider];
if (!presetForProvider || typeof presetForProvider !== 'object') return null;
if (!Object.hasOwn(presetForProvider, tier)) return null;
const tierPresets = presetForProvider[tier];
if (!tierPresets || typeof tierPresets !== 'object') return null;
const budget = (policy['budget'] && typeof policy['budget'] === 'string') ? policy['budget'] : 'medium';
if (!Object.hasOwn(tierPresets, budget)) return null;
const budgetEntry = tierPresets[budget];
if (!budgetEntry || !budgetEntry.model) return null;
return budgetEntry.model;
}
/**
* #2229 — the profile/phase-type tier for (config, agentType).
*
* Extracted verbatim from resolveModelInternal's step 2 so the same expression can
* answer "which tier did GSD resolve?" without also resolving a model id. The
* extraction is behaviour-preserving by construction: resolveModelInternal calls
* straight back into it.
*
* Returns null when the agent has no catalog entry and the profile is not `inherit`.
*/
function computeProfileTier(config: Record<string, unknown>, agentType: string): string | null {
// eslint-disable-next-line @typescript-eslint/no-base-to-string
const profile = String(config['model_profile'] || 'balanced').toLowerCase();
// Own-property guard: agentType is an unvalidated CLI positional (the
// `resolve-model <agent-type>` argument is never checked against a known
// agent list), so a prototype-chain agentType ("toString", "constructor")
// would otherwise return an inherited member from this plain object
// instead of undefined — verified reachable purely via the CLI.
const modelProfilesMap = MODEL_PROFILES as unknown as Record<string, Record<string, string>>;
const agentModels = Object.hasOwn(modelProfilesMap, agentType) ? modelProfilesMap[agentType] : undefined;
const phaseType = (AGENT_TO_PHASE_TYPE)[agentType];
const configModels = config['models'] as Record<string, string> | null | undefined;
const phaseTypeTier = (phaseType && configModels && typeof configModels === 'object')
? configModels[phaseType]
: undefined;
return (phaseTypeTier && VALID_TIERS.has(phaseTypeTier))
? phaseTypeTier
: (profile === 'inherit'
? 'inherit'
: (agentModels
// Own-property guard: `profile` is a config-supplied string
// (config['model_profile'], lower-cased); an already-lowercase
// prototype-chain key ("constructor", "__proto__") would otherwise
// return an inherited non-string member instead of falling back to
// 'balanced' (verified: profile:"constructor"/"__proto__" leaked a
// function/object through both the tier and model resolution paths).
? ((Object.hasOwn(agentModels, profile) ? agentModels[profile] : undefined) || agentModels['balanced'])
: null));
}
/**
* #2229 — the effective model TIER for (config, agentType), as a signal a workflow can
* read: `gsd_run query resolve-model <agent> --pick tier`.
*
* Why this is not just "look at the resolved model": on every runtime the installer
* configures with `resolve_model_ids: "omit"` — which is every non-Claude runtime, see
* docs/CONFIGURATION.md — resolveModelInternal deliberately returns '' below. A guard
* keyed on the model id therefore cannot tell a budget-tier run from a top-tier one
* there, while the tier itself is computed ABOVE that early-return and stays knowable.
*
* Honesty contract — this never guesses, because a guard that reports a wrong tier is
* worse than one that reports none:
* - a per-agent `model_overrides` pin naming a known alias (or a full Claude id that
* maps to one) reports that alias;
* - a pin that maps to nothing reports 'unknown' — a raw model id carries no tier;
* - `model_profile: inherit` reports 'inherit' — the session model is not ours to name;
* - an agent with no catalog entry reports 'unknown'.
*
* Callers must treat 'unknown' and 'inherit' as "cannot tell", never as "adequate".
*/
function resolveTierFromConfig(config: Record<string, unknown>, agentType: string): string {
const rawOverrides = config['model_overrides'];
const modelOverrides = (rawOverrides && typeof rawOverrides === 'object' && !Array.isArray(rawOverrides))
? rawOverrides as Record<string, string>
: null;
// Own-property guard: agentType is a caller-supplied string (the
// `resolve-model <agent-type>` CLI positional is not validated against a
// known agent list); a prototype-chain agentType ("toString",
// "constructor") against ANY model_overrides object — even `{}` — would
// otherwise return an inherited member instead of undefined.
const override = (modelOverrides && Object.hasOwn(modelOverrides, agentType))
? modelOverrides[agentType]
: undefined;
if (override && typeof override === 'string') {
if (CLAUDE_AGENT_ALIASES.has(override)) return override;
// Own-property guard: this indexes a plain object with a config-supplied
// string, so a prototype-chain key ("toString", "constructor", "valueOf")
// would otherwise return an inherited member instead of undefined — and a
// function-valued tier is dropped entirely by JSON.stringify, silently
// removing the key a guard depends on.
const alias = Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, override)
? CLAUDE_POLICY_ID_TO_ALIAS[override]
: undefined;
if (typeof alias === 'string' && alias) return alias;
return 'unknown';
}
const profileTier = computeProfileTier(config, agentType);
// #3282 — mirror resolveModelInternal's step 2.5 (model_policy preset). The
// profile tier alone under-reports: model_policy can dispatch a DIFFERENT
// tier than the profile implies (e.g. a `balanced` profile's "sonnet" tier
// combined with `model_policy: {budget: 'low'}` actually spawns "haiku"),
// and reporting the profile tier there is exactly the under-report this
// fixes — a haiku-tier run must never be reported as "sonnet". Skipped
// under the same condition resolveModelInternal skips it (no tier, or
// "inherit" — the session model is not ours to name).
if (profileTier && profileTier !== 'inherit') {
const mergedPolicy = config['model_policy']
? { ...(config['model_policy'] as Record<string, unknown>), runtime: (config['runtime'] as string | null | undefined) || 'claude' }
: null;
const policyModel = resolveModelPolicy(mergedPolicy, profileTier);
if (policyModel) {
// Map the policy-resolved id back to a tier alias with the same
// own-property-guarded lookups used above. If it maps, that alias IS
// the tier that actually runs — report it (the fix). If it does not
// map — including every non-Claude runtime, where resolveModelInternal
// returns the policy model verbatim with no tier meaning — the model
// carries no tier we can name; report 'unknown' rather than falling
// back to the profile tier, which would silently reintroduce the
// under-report this block exists to close.
const aliasForId = Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, policyModel)
? CLAUDE_POLICY_ID_TO_ALIAS[policyModel]
: undefined;
if (typeof aliasForId === 'string' && aliasForId) return aliasForId;
if (CLAUDE_AGENT_ALIASES.has(policyModel)) return policyModel;
return 'unknown';
}
}
return profileTier || 'unknown';
}
function resolveTierInternal(cwd: string, agentType: string): string {
return resolveTierFromConfig(loadConfig(cwd), agentType);
}
function resolveModelInternal(cwd: string, agentType: string): string {
const config = loadConfig(cwd);
// 1. Per-agent override (#2041: map Claude full IDs → Agent-tool aliases on
// the claude runtime, mirroring the model_policy path #1144; non-Claude
// runtimes and non-Claude values pass through verbatim).
const modelOverrides = config['model_overrides'] as Record<string, string> | null | undefined;
// Own-property guard (see resolveTierFromConfig above): without it, an
// agentType of "toString" against `model_overrides: {}` returned the
// inherited Function.prototype.toString as the resolved "model" — verified
// reachable purely via the CLI, no override value needed.
const override = (modelOverrides && Object.hasOwn(modelOverrides, agentType))
? modelOverrides[agentType]
: undefined;
if (override) {
const mapped = mapClaudeOverrideForRuntime(override, config['runtime'] as string | null | undefined, agentType);
if (mapped !== null) return mapped;
// Unmappable Claude ID — fall through to tier resolution (matches model_policy).
}
// 2. Compute the tier (#2229: shared with resolveTierFromConfig so the tier a
// workflow reads and the tier a model is resolved from can never diverge).
// eslint-disable-next-line @typescript-eslint/no-base-to-string
const profile = String(config['model_profile'] || 'balanced').toLowerCase();
// Own-property guard (see computeProfileTier above): without it, agentType
// "toString" returned Function.prototype.toString as `agentModels`
// (truthy), which skipped the "unknown agent" fallback below and made
// resolveModelInternal return undefined instead of a tier-derived string.
const modelProfilesMapForModel = MODEL_PROFILES as unknown as Record<string, Record<string, string>>;
const agentModels = Object.hasOwn(modelProfilesMapForModel, agentType) ? modelProfilesMapForModel[agentType] : undefined;
const tier = computeProfileTier(config, agentType);
// 2.5. model_policy preset (#49, #1133)
const configRuntime = config['runtime'] as string | null | undefined;
if (tier && tier !== 'inherit') {
const onClaude = !configRuntime || configRuntime === 'claude';
const effectiveRuntime = configRuntime || 'claude';
const mergedPolicy = config['model_policy']
? { ...(config['model_policy'] as Record<string, unknown>), runtime: effectiveRuntime }
: null;
const policyModel = resolveModelPolicy(mergedPolicy, tier);
if (policyModel) {
// Non-Claude runtimes take full model IDs verbatim (unchanged behavior).
if (!onClaude) return policyModel;
// Claude Code's Agent tool takes tier aliases (opus/sonnet/haiku/fable),
// not full model IDs — map the policy-resolved ID back to an alias (#1133).
const aliasForId = Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, policyModel)
? CLAUDE_POLICY_ID_TO_ALIAS[policyModel]
: undefined;
if (typeof aliasForId === 'string' && aliasForId) return aliasForId;
// The policy value may already be a bare Claude agent alias (e.g. "fable").
if (CLAUDE_AGENT_ALIASES.has(policyModel)) return policyModel;
// No Claude alias for this ID (e.g. a pinned minor version like
// claude-opus-4-5). Warn once and fall through to the tier alias rather
// than returning an ID Claude Code cannot spawn.
warnModelPolicyUnmappable(agentType, policyModel, tier);
}
}
// 3. Runtime-aware resolution (#2517)
if (configRuntime && configRuntime !== 'claude' && tier && tier !== 'inherit') {
const entry = _resolveRuntimeTier(config, tier);
if (entry?.model) return entry.model;
}
// 4. resolve_model_ids: "omit" — runtime-aware (#2297). Honor "omit" when the
// PROJECT explicitly set it (user intent — project config wins, #2517 finding
// #4) OR when the active runtime genuinely lacks native model aliases. Only a
// native-alias runtime (Claude) ignores an "omit" that came solely from the
// SHARED ~/.gsd/defaults.json — the #2297 poisoning fix — and falls through to
// its tier aliases below. Active runtime: GSD_RUNTIME → config.runtime → the
// per-install .gsd-runtime marker → 'claude' (canonicalized).
// NOTE: a non-Claude runtime that HAS a populated runtime-tier map already
// returned its own model id at step 3 above, before this gate — for those the
// explicit-project-omit honoring here is moot (step 3 wins, by #2517 design).
if (config['resolve_model_ids'] === 'omit'
&& (projectExplicitlySetsOmit(cwd) || !RUNTIMES_WITH_NATIVE_ALIASES.has(resolveActiveRuntime(config)))) {
return '';
}
// 5. Profile lookup (Claude-native default).
if (!agentModels) {
return profile === 'quality' ? 'opus'
: profile === 'budget' ? 'haiku'
: profile === 'inherit' ? 'inherit'
: 'sonnet';
}
if (tier === 'inherit') return 'inherit';
const alias = tier;
// Only the explicit `true` opt-in materializes full model IDs (#1569). Guard
// against the loose-truthy check catching a "omit" that a native-alias runtime
// ignored above (#2297): "omit" must fall through to the tier ALIAS here, not
// be materialized into a full ID Claude's Agent tool cannot spawn.
if (config['resolve_model_ids'] === true) {
return (MODEL_ALIAS_MAP as Record<string, string>)[alias!] || alias!;
}
return alias!;
}
const VALID_GRANULARITIES = new Set(['coarse', 'standard', 'fine']);
/**
* Resolve the planning granularity for a phase type (#68).
*/
function resolveGranularityInternal(cwd: string, phaseType: string | null | undefined, override?: string | null): string {
if (override !== undefined && override !== null && override !== '') {
if (VALID_GRANULARITIES.has(override)) {
return override;
}
}
const config = loadConfig(cwd);
const configGranularities = config['granularities'] as Record<string, string> | null | undefined;
const perPhase = (phaseType && configGranularities && typeof configGranularities === 'object')
? configGranularities[phaseType]
: undefined;
if (perPhase && VALID_GRANULARITIES.has(perPhase)) {
return perPhase;
}
if (config['granularity'] !== undefined && config['granularity'] !== null && config['granularity'] !== '') {
return config['granularity'] as string;
}
const planning = config['planning'] as Record<string, unknown> | null | undefined;
const planningGran = planning && planning['granularity'];
if (planningGran !== undefined && planningGran !== null && planningGran !== '') {
return planningGran as string;
}
return 'standard';
}
/**
* Validate a CLI granularity override at the command boundary. Empty/null/undefined
* are treated as "no override" (no-op). An invalid non-empty value calls `fail`.
*/
function assertValidGranularityOverride(
override: string | null | undefined,
fail: (msg: string) => never,
): void {
if (override !== undefined && override !== null && override !== '' && !VALID_GRANULARITIES.has(override)) {
fail(`invalid granularity '${override}' (valid: ${[...VALID_GRANULARITIES].join(', ')})`);
}
}
/**
* #3024 — Resolve a model for a specific dynamic-routing attempt.
*/
function resolveModelForTier(cwd: string, agentType: string, attempt?: number): string {
const config = loadConfig(cwd);
const attemptN = Number.isInteger(attempt) && (attempt as number) > 0 ? (attempt as number) : 0;
const modelOverrides = config['model_overrides'] as Record<string, string> | null | undefined;
// Own-property guard (see resolveTierFromConfig above): without it, an
// agentType of "toString" against `model_overrides: {}` returned the
// inherited Function.prototype.toString as the resolved "model" — verified
// reachable purely via the CLI, no override value needed.
const override = (modelOverrides && Object.hasOwn(modelOverrides, agentType))
? modelOverrides[agentType]
: undefined;
if (override) {
const mapped = mapClaudeOverrideForRuntime(override, config['runtime'] as string | null | undefined, agentType);
if (mapped !== null) return mapped;
// Unmappable Claude ID — fall through to dynamic_routing / model_policy resolution.
}
if (config['model_policy'] && config['runtime'] && config['runtime'] !== 'claude') {
return resolveModelInternal(cwd, agentType);
}
const dr = config['dynamic_routing'] as Record<string, unknown> | null | undefined;
if (!dr || typeof dr !== 'object' || dr['enabled'] !== true) {
return resolveModelInternal(cwd, agentType);
}
const tierModels = dr['tier_models'] as Record<string, string> | null | undefined;
if (!tierModels || typeof tierModels !== 'object') {
return resolveModelInternal(cwd, agentType);
}
const defaultTier = (AGENT_DEFAULT_TIERS)[agentType];
if (!defaultTier || !(VALID_AGENT_TIERS).has(defaultTier)) {
return resolveModelInternal(cwd, agentType);
}
const maxEscalations = Number.isInteger(dr['max_escalations']) && (dr['max_escalations'] as number) >= 0
? (dr['max_escalations'] as number)
: 1;
const escalationEnabled = dr['escalate_on_failure'] !== false;
const effectiveAttempt = escalationEnabled
? Math.min(attemptN, maxEscalations)
: 0;
let tier = defaultTier;
for (let i = 0; i < effectiveAttempt; i += 1) {
const next = (nextTier)(tier);
if (!next || next === tier) break;
tier = next;
}
const alias = tierModels[tier];
if (typeof alias !== 'string' || alias.length === 0) {
return resolveModelInternal(cwd, agentType);
}
return alias;
}
/**
* #2296 — Outcome of consulting the provider-escalation ladder for one attempt.
*
* `from`/`to` describe the PROVIDER ladder only; they are equal whenever the
* ladder was not consulted or had nothing to offer.
*/
interface ProviderEscalationResult {
from: string;
to: string;
escalated: boolean;
exhausted: boolean;
attempted: string[];
index: number;
}
/**
* Keep only usable model ids: non-empty strings. A malformed config can put
* anything in here (nulls, numbers, blank strings), and a blank model id would
* resolve to an unusable agent invocation rather than failing visibly. Invalid
* entries are dropped and the surviving order is preserved, so the ladder stays
* predictable (ADR 227 — validate shape, not just type).
*/
function sanitizeProviderEscalation(raw: unknown): string[] {
if (!Array.isArray(raw)) return [];
return raw.filter((entry): entry is string => typeof entry === 'string' && entry.trim().length > 0);
}
/**
* #2296 — Resolve the model for one attempt of the PROVIDER escalation ladder.
*
* The tier ladder (`resolveModelForTier`) escalates within one provider's
* `tier_models`, which does not help when that provider is the thing that is
* throttled. This walks `dynamic_routing.provider_escalation` instead: an
* ordered list of alternative model ids, capped by
* `min(max_escalations, list length)`.
*
* `applicable` is the caller's policy decision (only a quota-exceeded
* classification should consult this ladder). It is a parameter rather than a
* class check here so this module keeps depending only on leaf modules, per
* CONTEXT.md's Model Resolution module contract.
*
* Attempt 0 — and every non-applicable call — stays on the source model.
* `exhausted` reports that the ladder is spent so the caller can fail loudly
* naming every model it tried.
*/
function resolveProviderEscalation(
cwd: string,
agentType: string,
attempt: number | undefined,
applicable: boolean,
): ProviderEscalationResult {
// The model that would be used with no provider escalation at all.
const from = resolveModelForTier(cwd, agentType, 0);
const stay = (exhausted = false): ProviderEscalationResult => ({
from,
to: from,
escalated: false,
exhausted,
attempted: [from],
index: 0,
});
if (!applicable) return stay();
const config = loadConfig(cwd);
const dr = config['dynamic_routing'] as Record<string, unknown> | null | undefined;
if (!dr || typeof dr !== 'object' || dr['enabled'] !== true) return stay();
if (dr['escalate_on_failure'] === false) return stay();
const list = sanitizeProviderEscalation(dr['provider_escalation']);
if (list.length === 0) return stay();
// Same default and same validity rule as the tier ladder above — a negative or
// non-integer max_escalations is invalid config, not a request for zero.
const maxEscalations = Number.isInteger(dr['max_escalations']) && (dr['max_escalations'] as number) >= 0
? (dr['max_escalations'] as number)
: 1;
const cap = Math.min(maxEscalations, list.length);
// An explicit cap of 0 means the ladder exists but is spent before it starts.
if (cap === 0) return stay(true);
const attemptN = Number.isInteger(attempt) && (attempt as number) > 0 ? (attempt as number) : 0;
if (attemptN === 0) return stay();
const index = Math.min(attemptN, cap);
return {
from,
to: list[index - 1],
escalated: true,
exhausted: attemptN > cap,
attempted: [from, ...list.slice(0, index)],
index,
};
}
// ─── #443 — Unified effort + fast_mode resolvers ─────────────────────────────
const VALID_EFFORTS = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max'];
// #3533 (10d): the VOCABULARY carries one more member than the LADDER —
// 'inherit' is a declarable effort choice ("follow the session", expressed by
// OMITTING the effort key at the writer) but not a level nextEffort may step
// into. Keeping it out of VALID_EFFORTS means escalation (resolveEffortForTier)
// never walks past an explicit inherit: nextEffort('inherit') is null.
const EFFORT_SET = new Set([...VALID_EFFORTS, 'inherit']);
/**
* Walk one step up the effort ladder from `e`.
*/
function nextEffort(e: string): string | null {
const i = VALID_EFFORTS.indexOf(e);
if (i < 0) return null;
return VALID_EFFORTS[Math.min(i + 1, VALID_EFFORTS.length - 1)];
}
interface EffortOpts {
override?: string;
}
interface FastModeOpts {
override?: boolean;
}
/**
* #443 — Resolve a universal effort string for (cwd, agentType).
*/
function resolveEffortInternal(cwd: string, agentType: string, opts?: EffortOpts): string {
// Step 1: invocation override
if (opts && typeof opts.override === 'string' && EFFORT_SET.has(opts.override)) {
return opts.override;
}
const config = loadConfig(cwd);
const effortCfg = (config['effort'] && typeof config['effort'] === 'object' && !Array.isArray(config['effort']))
? (config['effort'] as Record<string, unknown>)
: null;
// Step 2: agent_overrides
if (effortCfg) {
const ao = effortCfg['agent_overrides'];
if (ao && typeof ao === 'object' && !Array.isArray(ao)) {
const v = (ao as Record<string, unknown>)[agentType];
if (typeof v === 'string' && EFFORT_SET.has(v)) return v;
}
} else {
const canonicalEffort = (CANONICAL_CONFIG_DEFAULTS)['effort'];
const mao = canonicalEffort && typeof canonicalEffort === 'object'
? (canonicalEffort as Record<string, unknown>)['agent_overrides']
: undefined;
if (mao && typeof mao === 'object' && !Array.isArray(mao)) {
const v = (mao as Record<string, unknown>)[agentType];
if (typeof v === 'string' && EFFORT_SET.has(v)) return v;
}
}
// Step 3: routing_tier_defaults by agent's default tier.
// #3531 (10c): the config block merges OVER the manifest tier defaults
// rather than replacing them — an effort block without
// routing_tier_defaults (or missing this agent's tier) falls back to the
// manifest built-in for that tier instead of skipping to effort.default.
// Invalid config values are dropped by the merge, so the manifest value for
// the tier surfaces (the same "invalid falls through" rule every layer has).
const agentTier = (AGENT_DEFAULT_TIERS)[agentType];
if (agentTier) {
const canonicalEffort = (CANONICAL_CONFIG_DEFAULTS)['effort'];
const manifestDefaults = canonicalEffort && typeof canonicalEffort === 'object'
? (canonicalEffort as Record<string, unknown>)['routing_tier_defaults'] as Record<string, string> | undefined
: undefined;
const isValidEffort = (v: unknown): v is string => typeof v === 'string' && EFFORT_SET.has(v);
const merged = mergeEffortTierDefaults(
manifestDefaults,
effortCfg ? effortCfg['routing_tier_defaults'] : undefined,
isValidEffort,
);
const v = merged[agentTier];
if (isValidEffort(v)) return v;
}
// Step 4: effort.default
if (effortCfg) {
const d = effortCfg['default'];
if (typeof d === 'string' && EFFORT_SET.has(d)) return d;
} else {
const canonicalEffort = (CANONICAL_CONFIG_DEFAULTS)['effort'];
const d = canonicalEffort && typeof canonicalEffort === 'object'
? (canonicalEffort as Record<string, unknown>)['default']
: undefined;
if (typeof d === 'string' && EFFORT_SET.has(d)) return d;
}
// Step 5: hardcoded default
return 'high';
}
/**
* #443 — Resolve fast_mode boolean for (cwd, agentType).
*/
function resolveFastModeInternal(cwd: string, agentType: string, opts?: FastModeOpts): boolean {
// Step 1: invocation override
if (opts && typeof opts.override === 'boolean') {
return opts.override;
}
const config = loadConfig(cwd);
const fmCfg = (config['fast_mode'] && typeof config['fast_mode'] === 'object' && !Array.isArray(config['fast_mode']))
? (config['fast_mode'] as Record<string, unknown>)
: null;
// Step 2: agent_overrides
if (fmCfg) {
const ao = fmCfg['agent_overrides'];
if (ao && typeof ao === 'object' && !Array.isArray(ao)) {
const v = (ao as Record<string, unknown>)[agentType];
if (typeof v === 'boolean') return v;
}
}
// Step 3: routing_tier_defaults by agent's default tier.
const agentTier = (AGENT_DEFAULT_TIERS)[agentType];
if (agentTier) {
if (fmCfg && fmCfg['routing_tier_defaults'] &&
typeof fmCfg['routing_tier_defaults'] === 'object' &&
!Array.isArray(fmCfg['routing_tier_defaults'])) {
const v = (fmCfg['routing_tier_defaults'] as Record<string, unknown>)[agentTier];
if (typeof v === 'boolean') return v;
} else if (!fmCfg) {
const canonicalFm = (CANONICAL_CONFIG_DEFAULTS)['fast_mode'];
const manifestDefaults = canonicalFm && typeof canonicalFm === 'object'
? (canonicalFm as Record<string, unknown>)['routing_tier_defaults']
: undefined;
if (manifestDefaults && typeof manifestDefaults === 'object') {
const v = (manifestDefaults as Record<string, unknown>)[agentTier];
if (typeof v === 'boolean') return v;
}
}
}
// Step 4: fast_mode.enabled
if (fmCfg && typeof fmCfg['enabled'] === 'boolean') {
return fmCfg['enabled'];
}
// Step 5: hardcoded default
return false;
}
/**
* #443 — Resolve effort for a dynamic-routing attempt (with escalation).
*/
function resolveEffortForTier(cwd: string, agentType: string, attempt?: number): string {
const base = resolveEffortInternal(cwd, agentType);
const config = loadConfig(cwd);
const dr = config['dynamic_routing'] as Record<string, unknown> | null | undefined;
if (!dr || typeof dr !== 'object' || dr['enabled'] !== true) {
return base;
}
if (dr['escalate_on_failure'] === false) {
return base;
}
const maxEscalations = Number.isInteger(dr['max_escalations']) && (dr['max_escalations'] as number) >= 0
? (dr['max_escalations'] as number)
: 1;
const attemptN = Number.isInteger(attempt) && (attempt as number) > 0 ? (attempt as number) : 0;
const effectiveAttempt = Math.min(attemptN, maxEscalations);
let current = base;
for (let i = 0; i < effectiveAttempt; i++) {
const next = nextEffort(current);
if (!next || next === current) break;
current = next;
}
return current;
}
export = {
resolveTierEntry,
CLAUDE_AGENT_ALIASES,
resolveModelPolicy,
resolveModelInternal,
resolveTierInternal,
resolveTierFromConfig,
_resetModelPolicyWarningCacheForTests,
_resetModelOverrideWarningCacheForTests,
_setInstallRuntimeMarkerForTests,
_resetInstallRuntimeMarkerCacheForTests,
VALID_GRANULARITIES,
resolveGranularityInternal,
assertValidGranularityOverride,
resolveModelForTier,
resolveProviderEscalation,
VALID_EFFORTS,
EFFORT_SET,
nextEffort,
resolveEffortInternal,
resolveFastModeInternal,
resolveEffortForTier,
};