* test(#3897): failing-first coverage for §8.3 rungs 2-4 ADR-3473 §8.3 has four rungs; #3883/PR #3896 shipped the first. This pins the other three RED before any fix. Rung 2 — the install marker has four readers and resolveRuntime is not one. resolveRuntime resolves GSD_RUNTIME > config.runtime > 'claude' and reads no marker at all, while bin/install.js writes one (#2297) and FOUR hand-rolled readInstallRuntimeMarker copies exist: src/model-resolver.cts:65 (cached, with test seams), hooks/gsd-agent-isolation-guard.js:112, and TWICE in hooks/gsd-cursor-subagent-start.js at :346 and :355. Four copies of one rule. Fixtures and seam names mined from PR #3382 rather than re-derived; it implemented this rung and was closed "not on the merits". Rung 3 — the sandbox map, and the fallback that was the real defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: - all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements — the map carries nothing the contract does not - 24 roles fall through `|| 'read-only'`, of which 16 declare Write or Edit So the map is redundant and the silent fallback is the defect. The maintainer chose to derive but hold those 16 at read-only pending the question of whether Codex enforces sandbox_mode or merely advises; HALT.md records it. T20 asserts the emitted sandbox_mode PER ROLE against a captured baseline, not in aggregate — an aggregate passes while one role silently widens, which is the proxy-instead-of-identity shape this repo names. T24 and T25 fail on a stale hold, so the hold list cannot rot into the subset map being deleted. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: `git log -S shortFormToId` returns only documentation commits. That was the wrong instrument. Direct inspection of sdk/src/query/phase.ts at 11918dcc3^ shows five occurrences, and the tests match that code rather than a guess at its semantics — including first-write-wins on a duplicate short form. T43 asserts at the consumer's output: the emitted `waves` map from the real CLI, which pre-fix collapses to {"1":[...]} because every short-form edge is dropped. A unit assertion on resolveDependencyId would have passed throughout this defect's life. Observed RED, this tree: rung 2 11/11 fail — no marker rung, no seams rung 3 T23,T24,T25,T26,T30 fail; T28 fails (validate agents passes a TOML whose sandbox_mode disagrees — it checks presence only) rung 4 T42,T44 fail; T43,T49 fail with waves collapsed to a single wave 1 Green and staying green: T20/T21/T22/T27 as captured baselines, #3885's unresolvable-token warning and wave-verdict suppression, and #3785's display-mapping passthrough. If the third tier over-reaches, those go red — that is their job. Disclosed weakness: T45 (a canonical id with no dash is not short-form indexed) cannot be isolated behaviorally, because planMap always masks it. It is a non-crash boundary pin, weaker than the other rows, and is recorded as such rather than presented as equivalent. Design: .gsd/phase/feat-3897-adr3473-83-rungs/40-design.md Test matrix: .gsd/phase/feat-3897-adr3473-83-rungs/50-test-matrix.md Decision: .gsd/phase/feat-3897-adr3473-83-rungs/HALT.md Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3897): §8.3 rungs 2-4 — one marker reader, a derived sandbox, the third depends_on tier ADR-3473 §8.3 has four rungs. #3883/PR #3896 shipped the first. These are the other three. Rung 2 — the install marker had four readers, and resolveRuntime was not one. resolveRuntime resolved GSD_RUNTIME > config.runtime > 'claude' and read no marker, while bin/install.js writes one (#2297) and four hand-rolled readInstallRuntimeMarker copies existed: src/model-resolver.cts (cached, with seams), hooks/gsd-agent-isolation-guard.js, and twice in hooks/gsd-cursor-subagent-start.js. model-resolver's was already the house idiom, so it was promoted rather than replaced: src/runtime-slash.cts now owns it, and model-resolver plus both hooks delegate. The hooks reach it through ensureRuntimeBuild(), the seam lint-hooks-runtime-build-seam enforces. No import cycle existed - checked both directions before moving anything. The marker is the THIRD rung: env > project config > marker > 'claude'. N1 was checked rather than assumed, and my first reading of it was wrong. A marker holding an unknown name comes back essentially verbatim, which looked like a validation gap. Measured against the env rung with the same inputs - including "../../etc/passwd" and "claude;rm -rf /" - the two are identical, because they share resolveRuntimeNameFromCandidates. N1 asks for exactly that, and it is met. The residual (the shared normalizer normalizes shape, it does not validate against the known-runtime set) is pre-existing on the env rung and plausibly deliberate, since a new runtime should not need a code change. The marker also does not widen the trust boundary in any real sense: it lives inside the install tree beside the code, so anyone who can write it can write runtime-slash.cjs itself. Rung 3 — the map was redundant; the silent fallback was the defect. Measured across all 35 files in agents/, deriving workspace-write iff tools: declares Write or Edit: all 11 CODEX_AGENT_SANDBOX entries derive to their mapped value exactly, zero disagreements. The map carried nothing the contract did not already have, so it is DELETED rather than reconciled. What was actually broken is `|| 'read-only'`, which silently under-granted 24 of 35 roles. 16 of those 24 declare Write or Edit and would widen under derivation. Per the maintainer's decision (HALT.md), they are held at read-only pending the question of whether Codex enforces sandbox_mode or merely advises. Emitted TOML is therefore byte-identical for all 35 roles - asserted per role, not in aggregate, because an aggregate passes while one role silently widens. The hold list self-invalidates. A hold whose role no longer derives broader fails, and so does a hold naming a role with no agents/<name>.md. Without that it would rot into exactly the hand-maintained subset map being deleted, and this commit's own ledger claim would become false over time. Both cases were proved by injecting them and watching them throw. Two committed tests asserted the deleted map's existence and contents. They were pinning the thing being removed, so the tests moved rather than the production code: the 11 role-value pairs survive as a test-local PRE_3897_CODEX_AGENT_SANDBOX baseline, and the assertions now drive the real derivation against real agents/*.md. The coverage is preserved; only its source moved out of production code. validate agents gains checkCodexSandboxPosture, mirroring the existing checkCodexModelPosture: each installed TOML's sandbox_mode must equal the role's expected value, failing with role, expected and found. It previously checked file presence and manifest completeness only, so a TOML whose sandbox_mode disagreed passed. Rung 4 — shortFormToId, recovered rather than invented. I nearly reported this as another wrong §8.3 claim: git log -S returns only documentation commits. Wrong instrument. sdk/src/query/phase.ts at 11918dcc3^ carries five occurrences, and the implementation here matches that code rather than a guess at its semantics - including first-write-wins on a duplicate short form, deterministic from the sorted plan order. It resolves the bare plan number: depends_on: ["01"] now reaches 26-01-auth-hardening. That is a control-flow change, not a diagnostic one - plans that silently collapsed into a single wave 1 now execute in their declared waves, and execute-phase.md consumes those wave values. In-phase only, by construction: the map is built from this phase's rawPlans, so a same-named short form in another phase does not resolve. #3785's display-mapping passthrough and #3885's unresolvable-token warning and wave-verdict suppression are untouched and stay green. If the third tier had over-reached, those are what would have caught it. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open I introduced, and wire the posture check to its command Two blockers from review. Both are mine, and one is a security regression my own change created. 1. A held role could escape its hold by editing its own frontmatter. The Codex install loop set the sandbox identity from the agent's frontmatter `name:` field rather than from its filename, so the hold lookup keyed off a value the file itself declares: deriveCodexSandboxMode('gsd-doc-writer', <real file>) -> read-only deriveCodexSandboxMode('gsd-doc-writer-x', <same file, name: edited>) -> workspace-write deriveCodexSandboxMode('GSD-Doc-Writer', <same file, name: recased>) -> workspace-write What makes this a blocker rather than a nit is the DIRECTION. The deleted CODEX_AGENT_SANDBOX map had the identical lookup-key quirk, but it was an allowlist: an unmatched key fell back to read-only, which is safe. The new scheme derives workspace-write from the tool contract and uses the hold as a subtraction, so the same mismatch fails OPEN. I converted a fail-closed quirk into a fail-open one and did not notice; the isolated reviewer proved it by execution. Neither safety net caught it. validateCodexSandboxHolds only checks that <key>.md exists, never that a file's derived identity matches its key. checkCodexSandboxPosture looks the canonical source up by the installed TOML's filename, finds nothing for a renamed agent, and treats it as a custom non-roster agent — silently no violation. The identity is now the FILENAME STEM, which is what validateCodexSandboxHolds already validates and what an attacker editing frontmatter cannot change without renaming the file — at which point the existing validator catches it. The lookup is case-insensitive so a recase does not slip past either. The frontmatter name still drives the TOML body and filename, unchanged; only the sandbox identity moved. All 35 roster files were checked: name matches filename stem everywhere, so a stricter "they must agree or throw" invariant would have been safe against real content. It is deliberately NOT added — it would abort an install on a tampered file where emitting a correctly-derived read-only TOML is the safer outcome. Recorded as a fork rather than decided silently. 2. checkCodexSandboxPosture was exported and never called. cmdValidateAgents (src/verify.cts) called checkAgentsInstalled and checkCodexModelPosture only; grep for the sandbox check in that file returned nothing. So criterion 3 — "validate agents fails on semantic drift, not only on missing files" — was unmet, and `validate agents` behaved exactly as before. That is ADR-3473 Decision 2's named shape: a declared policy with no executor. It also meant the T28 test asserted at the helper's return value while the COMMAND stayed broken — the ADR-3180 Decision 4(b) failure this epic exists to close, committed by me while enforcing it elsewhere in the same epic. Now wired as an additive `sandbox_posture` field beside `codex_posture`, following the sibling precedent exactly. Drift is report-only, not a non-zero exit, because that is what checkCodexModelPosture does — two sibling posture checks disagreeing about whether a violation is fatal would be its own defect. The choice is recorded in a comment rather than left implicit. A consumer-output test now drives the real CLI and asserts on the emitted JSON, and was shown failing before the wiring and passing after. Also corrected a stale artifact: the design's Known limit L1 still claimed rung 3 was not in this deliverable, written while it was halted and false once the maintainer unblocked it. Verified after both fixes: the three bypass probes all return read-only, the per-role table is 35/35 byte-identical, and both hold self-invalidation cases still throw. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the marker rung, the derived sandbox, and the bare plan-number depends_on Reference: the runtime precedence ladder in docs/CLI-TOOLS.md gains the install marker rung; docs/COMMANDS.md documents validate agents' new sandbox_posture field; docs/reference/plan-md.md documents that depends_on accepts the bare plan number. Explanation: a docs/features fragment keyed id 3897, so it cannot collide with a concurrent PR hand-allocating a section number, regenerated into FEATURES.md. ADR-3473 §8.3 gains an ANSWER blockquote in the document's own correction style, recording what was measured and built against the section's 2026-08-26 correction - including the qualification that checkAgentsInstalled itself still checks presence only, and the semantic assertion lives in a sibling wired into validate agents rather than folded into it. No how-to. Both user-visible changes are zero-step: a non-Claude install resolving its own runtime, and plans executing in their declared waves, both happen without the user doing anything. docs/how-to/control-the-reported-host-runtime.md covers a DIFFERENT ladder (resolveReportedRuntime / agent_runtime) that this change does not touch, and was deliberately left alone rather than edited by association. No tutorial - nothing multi-step to walk through. docs/AGENTS.md unchanged: it documents Claude-side tools frontmatter, never Codex sandbox_mode, and the emitted tools contract did not change. The prompt layer documents depends_on only by example, not by schema, so nothing there needed editing - and few-shot-examples/plan-checker.md already showed depends_on: ['01'], which now actually resolves. Translated copies of plan-md.md are untouched; the project treats translations as community-maintained. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): move the sandbox derivation out of the installer, off the install path, and off a third parser The full suite came back with 26 failures across four files. Three distinct causes, mapped individually rather than assuming the first explained the rest. A. Requiring bin/install.js printed the GSD banner to stdout and corrupted `validate agents` JSON. Unexpected token '', "[36m ██"... is not valid JSON checkCodexSandboxPosture reached deriveCodexSandboxMode by lazily requiring bin/install.js, whose module load prints the ASCII banner. So the command emitted banner bytes before its JSON and every JSON consumer broke, including ten tests that predate this branch. src/ reaching into bin/ was backwards layering that happened to also be loud. The derivation now lives in src/codex-agent-toml.cts - the existing Codex TOML domain module, no new module and no six-gate ripple - and both bin/install.js and src/agent-install-check.cts import it. One owner, which is §8.3's rule applied to the fix for §8.3. B. The stale-hold throw fired on a legitimate partial source dir, and masked a security assertion. validateCodexSandboxHolds treated "this hold's .md is absent from the install SOURCE dir" as a stale hold and threw. A test fixture, or any partial install source, legitimately contains a couple of agents. Worse, it threw BEFORE the path-escape check, so a test asserting that a `../../evil` frontmatter name is rejected got my unrelated error instead of the traversal rejection it was written for. A fail-closed check of mine was hiding a real security check. The "no stale holds, shrink-only" invariant is a property of the repo's canonical agents/ roster, not of whatever directory an install happens to read. It is off the runtime path and enforced where it belongs, in the tests that already existed for it. A partial source dir now installs cleanly, and the evil-name case throws with its own escapes-configHome message again. C. T8 depended on ambient process.env state. The marker/env parity assertion round-tripped through live process.env. It now compares against resolveExplicitRuntime's already-exported dependency-injection parameter - deterministic and hermetic, same claim. Proven still falsifiable rather than assumed: with the marker rung's normalization temporarily bypassed the two rungs diverge ("codex\n../../etc/passwd" vs "codex-../../etc/passwd") and the assertion fails, then passes again once reverted. One correction folded in along the way. The first version of the move added private _extractFrontmatterAndBody/_extractFrontmatterField helpers to codex-agent-toml.cts - a THIRD copy of frontmatter extraction, where the graph already shows two (bin/install.js:2348, runtime-artifact-conversion.cts:893). Adding a third inside the epic whose thesis is one implementation per rule is not defensible. deriveCodexSandboxMode no longer parses anything: it takes (identity, toolsValue) and each caller supplies the tools value using the extractor it already has. Both helpers are deleted. The identity argument is still the filename stem, so the fail-open fix is untouched. Verified after all three: `validate agents --raw` emits parseable JSON with no banner and both posture fields; the four hold-bypass probes still return read-only; the per-role table is 35/35 byte-identical at 26 read-only / 9 workspace-write; the hold list is still 16. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): drop a dev-only transitive dep, make the derivation total, retire a stale fallback test Suite down to 7 failures from 26. Three more causes, mapped individually. A. My extractor import dragged in a script that does not exist in an installed tree. Cannot find module '../../../scripts/fix-slash-commands.cjs' Chain: src/agent-install-check.cts imported runtime-artifact-conversion.cjs, which requires command-roster.cjs, whose line 36 requires ../../../scripts/fix-slash-commands.cjs. That path exists in the repo and not in an install, so every test exercising a synthetic install dir died at module load. I picked that extractor for convenience without checking what it pulls in - the same mistake that produced the banner bug, one layer further out. agent-install-check now uses a single-purpose extractToolsLine on codex-agent-toml.cts. That is deliberately NOT a general frontmatter parser: we deleted those helpers a commit ago for good reason, and this reads one line. Verified from outside the repo root that requiring either module prints nothing and does not throw. B. A test pinned the deleted name-based fallback. 'defaults unknown agents to read-only' called generateCodexAgentToml with a fixture declaring tools: Read, Write, Edit. Under derivation an unknown agent with a writing contract correctly derives workspace-write - design row S6, a new writing role gets the contract, not the pin. The behavior it asserted was the silent fallback this rung deleted; identity no longer decides the sandbox. Replaced with two rows rather than a flipped string: no tools declared -> read-only (absence is not a grant), and Write/Edit declared -> workspace-write. Strictly more coverage than the row it replaces. C. The stale-hold check still threw per derivation call. Last commit took the roster-existence check off the install path, but deriveCodexSandboxMode itself still threw when a hold's role did not derive broader FOR THE CONTENT IT WAS HANDED - so it fired on any synthetic fixture for a held role. The throw is gone, and it cost nothing: if a held role's content does not derive broader, the hold pins read-only and derivation returns read-only anyway, so the hold is a no-op and there is nothing to fail about. The staleness invariant is a property of the real agents/ roster, and validateCodexSandboxHolds still enforces it there - confirmed against the real roster after the change, not assumed. deriveCodexSandboxMode is now total: every (identity, toolsValue) including undefined and null returns read-only or workspace-write, never throws. Verified: validate agents emits parseable JSON; the four hold-bypass probes return read-only; the per-role table is 35/35 at 26 read-only / 9 workspace-write. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): put the rung-3 decision in the shipped docs instead of pointing at an ignored path The ADR entry and the feature fragment both ended their rung-3 explanation with "see .gsd/phase/feat-3897-adr3473-83-rungs/45-decision-rung3-sandbox.md". That directory is gitignored (.gitignore:55), so the rationale for holding 16 roles at read-only was reachable only from the machine that produced it. A reader of the ADR got a pointer to nothing. Both now carry the reasoning inline: the criterion asks both that the sandbox derive from the declared tool contract and that no role gain a broader sandbox, and those cannot both hold, because a faithful derivation widens 16 roles the deleted map never listed and that fell through its silent read-only default. The resolution is derive-and-hold - the derivation owns the rule now, each hold is released as its enforcement question is answered, and a hold is reversible where a widened sandbox that turns out to be enforced is not. Checked before assuming this was a defect class: CONTEXT.md cites .gsd/phase/<slug>/40-design.md as its standard Design: provenance line in eight module entries, and four other shipped docs do the same. Citing a phase artifact is an established convention here, so those are left alone. What was wrong was specific to these two: they put load-bearing rationale behind the pointer instead of provenance. docs/FEATURES.md regenerated from the fragment via scripts/gen-features.cjs rather than hand-edited. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3897): close a fail-open, stop a silent mis-resolution, and read a declaration as a declaration Two orthogonal reviews on the shipped sha. Three of the findings are the same failure class this epic exists to close, committed inside it. 1. BLOCKER - the sandbox was decided for one identity and applied to another. bin/install.js derived sandbox_mode for the filename stem and then wrote the result to `${name}.toml`, where name comes from the file's own frontmatter. Make the two disagree and a HELD role's artifact goes wide: rename gsd-doc-writer.md -> gsd-doc-writer-v2.md, keep name: gsd-doc-writer -> stem is unheld, derives workspace-write, lands on gsd-doc-writer.toml add any gsd-*.md whose frontmatter name: is a held role -> clobbers that role's toml with workspace-write Both emit read-only on origin/next, because the deleted map was an allowlist and a miss fell back safe. This is a regression my change introduced. The previous review round moved the HOLD KEY off frontmatter to the filename stem and left the OUTPUT PATH on frontmatter; my own comment at install.js:6985 calls that value attacker-editable, four lines above the line that uses it as the filename. The decision is now made over BOTH candidate identities, most-restrictive wins: if either the stem or the emitted name is held, the mode is read-only. 2. MAJOR - hold matching was toLowerCase() only, so confusables escaped. Turkish dotted/dotless i, fullwidth, NFD, trailing space/NBSP/dot/newline, ./ and ../agents/ all slipped the hold and emitted workspace-write. Identities are now basenamed, trimmed of NBSP/zero-width/control characters, NFKC-normalized and lowercased - and anything still carrying a character outside [a-z0-9._-] is treated as suspicious and derives read-only. We do not enumerate confusables; every shipped roster file is ASCII, so refusing to widen on an identity we cannot recognize is fail-closed with no false positives on real content. 3. MAJOR - the short-form depends_on tier mis-resolved SILENTLY. shortFormToId keyed on the last dash-segment of any canonical id with no constraint that it is a plan number, so a phase holding 09-FIX-auth-PLAN.md made depends_on: ["auth"] bind at wave 2 with zero warnings. This is the worst shape in the epic: the unresolvable-token warning fires on a DROPPED token, so a MIS-RESOLVED one is invisible and the tool reports a confident wave assignment built from a wrong edge. A wrong edge is worse than a missing one. The segment must now match /^\d+$/, which is exactly the contract docs/reference/plan-md.md already documents. This tier was recovered verbatim from the retired SDK lineage, which carried the same defect; we are deliberately NOT preserving it bug-for-bug, and the comment says so, so the next reader does not "restore" it. 4. MAJOR - the derivation was reading a declaration as an absence. extractToolsLine read one line, so a YAML list-form tools: block returned only its first item. Two roster files use list form, and gsd-nyquist-auditor declares Write and Edit there - parsed as "- Read", found no write tool, and emitted read-only. Rung 3's headline claim is that sandbox_mode derives from the declared tool contract; that claim was false for 2 of 35 roles and materially wrong for 1. Reading a declaration as an absence is the silent-drop class this epic exists to close. Renamed extractToolsValue and taught it both shapes. gsd-nyquist-auditor now derives workspace-write and joins CODEX_SANDBOX_HOLDS as its 17th entry, per the standing derive-and-hold decision - so emitted TOML stays byte-identical at 26 read-only / 9 workspace-write while the hold list finally records every role that would widen. A previous pass declined this fix because it moved the count; that inverts the priority. Byte-identity is preserved THROUGH the hold, not by leaving a parser broken. Divergence check, because this is where that bug hides: both paths feeding sandbox derivation - install.js's emitter and checkCodexSandboxPosture - now route through the one extractor. The tools readers in runtime-artifact-conversion and install.js's other frontmatter call sites serve Claude-side emission and do not feed sandbox derivation. Also fixed, each real: the posture check's `found` used a naive whole-file regex where its own sibling uses the block-aware scanner, so prose inside developer_instructions produced a false violation; `found` skipped truncatePostureValue and leaked a 300-char value into validate agents output; deriveCodexSandboxMode's absolute never-throws claim was false for an object with a throwing toString; T49 could not falsify cross-phase leakage (its target phase had its own 01, so a globally-scoped map passed too); T20/N6 iterated a hardcoded table and pinned the FIXTURE size, so a 36th agent would be silently unchecked; three tests reimplemented the code they were testing instead of importing it; and T2-T4 deleted GSD_RUNTIME without restoring it. Verified: hold list 17, gsd-nyquist-auditor derives workspace-write unheld and emits read-only held, roster 35/35 at 26/9, depends_on ["auth"] no longer resolves while ["01"] still does, both identity-bypass cases and every confusable vector emit read-only. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3897): the hold list is 17, and the reason the 17th was missing The count read 16 because the derivation could not read the declaration it claimed to derive from: the tools reader was single-line, so a YAML list-form tools: block returned only its first item and gsd-nyquist-auditor's declared Write and Edit were read as an absence. Both the ADR entry and the feature fragment now carry the corrected count and the reason for it, rather than a silently updated number. Deriving from a declaration you cannot parse is not deriving, and a flattering count is worse than a wrong one because it looks settled. docs/FEATURES.md regenerated from the fragment. Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3897): backfill changeset pr number Refs #3897 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
935 lines
41 KiB
TypeScript
935 lines
41 KiB
TypeScript
/**
|
|
* Model Resolver — Model and effort resolution policy
|
|
*
|
|
* ADR-857 rollout phase 2f: extracted from core.cts (issue #888).
|
|
* Owns model and effort resolution policy: resolves the model, runtime tier,
|
|
* planning granularity, reasoning effort, and fast-mode for a given agent by
|
|
* reading project config and resolving against the model profiles and catalog.
|
|
* Behaviour is preserved byte-for-behaviour from the prior location; only
|
|
* the module boundary moved. The core.cjs re-export spine was retired in
|
|
* epic #1267; callers import resolvers from model-resolver.cjs directly.
|
|
*
|
|
* Dependencies (leaf modules only):
|
|
* - node:fs / node:path (read the per-install .gsd-runtime marker + project config for the #2297 omit gate)
|
|
* - ./runtime-name-policy.cjs (resolveRuntimeNameFromCandidates — canonicalize the active runtime)
|
|
* - ./planning-workspace.cjs (planningDir — workstream/project-aware project-config path)
|
|
* - ./config-loader.cjs (loadConfig)
|
|
* - ./configuration.cjs (CONFIG_DEFAULTS as CANONICAL_CONFIG_DEFAULTS)
|
|
* - ./model-profiles.cjs (MODEL_PROFILES, AGENT_TO_PHASE_TYPE, AGENT_DEFAULT_TIERS, VALID_AGENT_TIERS, nextTier)
|
|
* - ./model-catalog.cjs (MODEL_ALIAS_MAP, RUNTIME_PROFILE_MAP, PROVIDER_PRESETS, VALID_TIERS,
|
|
* CLAUDE_AGENT_ALIASES — re-exported below for back-compat, #3241)
|
|
*/
|
|
|
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
import configLoaderModule = require('./config-loader.cjs');
|
|
const { loadConfig } = configLoaderModule;
|
|
|
|
// ─── Configuration Module (for CANONICAL_CONFIG_DEFAULTS used by effort/fast_mode resolvers) ─
|
|
import { CONFIG_DEFAULTS as CANONICAL_CONFIG_DEFAULTS } from './configuration.cjs';
|
|
|
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
import modelProfiles = require('./model-profiles.cjs');
|
|
const { MODEL_PROFILES, AGENT_TO_PHASE_TYPE, AGENT_DEFAULT_TIERS, VALID_AGENT_TIERS, nextTier } = modelProfiles;
|
|
|
|
import { MODEL_ALIAS_MAP, RUNTIME_PROFILE_MAP, PROVIDER_PRESETS, VALID_TIERS, CLAUDE_AGENT_ALIASES, mergeEffortTierDefaults } from './model-catalog.cjs';
|
|
|
|
import fs from 'node:fs';
|
|
import path from 'node:path';
|
|
import { resolveRuntimeNameFromCandidates } from './runtime-name-policy.cjs';
|
|
import {
|
|
readInstallRuntimeMarker,
|
|
_setInstallRuntimeMarkerForTests,
|
|
_resetInstallRuntimeMarkerCacheForTests,
|
|
} from './runtime-slash.cjs';
|
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
import planningWorkspaceMod = require('./planning-workspace.cjs');
|
|
const { planningDir } = planningWorkspaceMod;
|
|
|
|
// ─── #2297: per-install runtime identity for the resolve_model_ids:"omit" gate ─
|
|
//
|
|
// The installer writes `resolve_model_ids:"omit"` into the SHARED
|
|
// ~/.gsd/defaults.json for every runtime that lacks native model aliases (#1156).
|
|
// Because that file is machine-wide, a non-Claude install would otherwise poison
|
|
// a Claude no-project resolution into returning '' — silently defeating Claude's
|
|
// adaptive tier aliases. The "omit" must therefore apply only when a runtime that
|
|
// genuinely lacks native aliases is the one resolving.
|
|
//
|
|
// In a no-project session there is no `.planning/config.json` (so config.runtime
|
|
// is null) and GSD_RUNTIME is not exported by gsd-core, so the only reliable
|
|
// current-runtime signal is the per-install marker the installer co-locates next
|
|
// to VERSION at <install>/gsd-core/.gsd-runtime (this file's dir is
|
|
// <install>/gsd-core/bin/lib). Precedence for the gate: project config.runtime →
|
|
// GSD_RUNTIME env (manual/CI override + test seam) → install marker → 'claude'.
|
|
//
|
|
// `claude` is currently the ONLY runtime with nativeModelAliases:true; a
|
|
// registry-parity test guards this set so a future alias-capable runtime fails
|
|
// loudly here instead of silently omitting.
|
|
const RUNTIMES_WITH_NATIVE_ALIASES: ReadonlySet<string> = new Set(['claude']);
|
|
|
|
// #3897 rung 2: the marker reader + its cache and test seams were promoted to
|
|
// the canonical owner, `runtime-slash.cts` (imported above) — this module now
|
|
// consumes that single implementation instead of holding its own copy. N5:
|
|
// behaviour and the seam contract are unchanged by the move; the re-exports
|
|
// below (`export =` at the bottom of this file) preserve every existing
|
|
// caller's `require('./model-resolver.cjs')` surface byte-for-behaviour.
|
|
|
|
// The runtime whose install is actually resolving, canonicalized so an alias or
|
|
// case variant (e.g. "claude-code"/"Claude") cannot defeat the native-alias
|
|
// check below (#2297 review). Precedence mirrors resolveRuntime()
|
|
// (runtime-slash.cts): GSD_RUNTIME env → project config.runtime → per-install
|
|
// .gsd-runtime marker → 'claude'.
|
|
function resolveActiveRuntime(config: Record<string, unknown>): string {
|
|
return resolveRuntimeNameFromCandidates(
|
|
process.env['GSD_RUNTIME'],
|
|
config['runtime'],
|
|
readInstallRuntimeMarker(),
|
|
) || 'claude';
|
|
}
|
|
|
|
// Did the PROJECT's own config (root `.planning/config.json` or the active
|
|
// workstream/project override) explicitly set resolve_model_ids to "omit"?
|
|
// Project config takes precedence over the shared ~/.gsd/defaults.json (#2297
|
|
// out-of-scope guard + #2517 finding #4): an explicit project "omit" is honored
|
|
// regardless of runtime, whereas an "omit" that came only from the global
|
|
// defaults is ignored by native-alias runtimes. Workstream/project-scope aware
|
|
// via planningDir (mirrors loadConfig's precedence: workstream value wins over
|
|
// root); a plain read avoids loadConfig's normalization side effects.
|
|
function projectExplicitlySetsOmit(cwd: string): boolean {
|
|
const wsDir = planningDir(cwd);
|
|
const rootDir = path.join(cwd, '.planning');
|
|
const layers = wsDir === rootDir ? [rootDir] : [wsDir, rootDir]; // workstream > root
|
|
for (const dir of layers) {
|
|
try {
|
|
const parsed = JSON.parse(fs.readFileSync(path.join(dir, 'config.json'), 'utf8')) as Record<string, unknown>;
|
|
const value = parsed?.['resolve_model_ids'];
|
|
// First layer that sets the key wins (matches loadConfig's deep-merge
|
|
// precedence). A layer that omits the key falls through to the next.
|
|
if (value !== undefined) return value === 'omit';
|
|
} catch {
|
|
// Absent/unreadable layer — try the next.
|
|
}
|
|
}
|
|
return false;
|
|
}
|
|
|
|
// ─── Model alias resolution ───────────────────────────────────────────────────
|
|
|
|
interface TierEntryResolved {
|
|
model: string;
|
|
reasoning_effort?: string;
|
|
[key: string]: unknown;
|
|
}
|
|
|
|
interface ResolveTierEntryOpts {
|
|
runtime: string | null | undefined;
|
|
tier: string | null | undefined;
|
|
overrides: Record<string, unknown> | null | undefined;
|
|
}
|
|
|
|
/**
|
|
* #2517 — Resolve the runtime-aware tier entry for (runtime, tier).
|
|
*/
|
|
function resolveTierEntry({ runtime, tier, overrides }: ResolveTierEntryOpts): TierEntryResolved | null {
|
|
if (!runtime || !tier) return null;
|
|
|
|
const runtimeMap = RUNTIME_PROFILE_MAP as unknown as Record<string, Record<string, Record<string, unknown>>>;
|
|
const builtin = runtimeMap[runtime]?.[tier] || null;
|
|
const overridesMap = overrides as Record<string, Record<string, unknown>> | null | undefined;
|
|
const userRaw = overridesMap?.[runtime]?.[tier];
|
|
|
|
let userEntry: Record<string, unknown> | null = null;
|
|
if (userRaw) {
|
|
userEntry = typeof userRaw === 'string' ? { model: userRaw } : (userRaw as Record<string, unknown>);
|
|
}
|
|
|
|
if (!builtin && !userEntry) return null;
|
|
return { ...(builtin || {}), ...(userEntry || {}) } as TierEntryResolved;
|
|
}
|
|
|
|
/**
|
|
* Convenience wrapper used by resolveModelInternal.
|
|
*/
|
|
function _resolveRuntimeTier(config: Record<string, unknown>, tier: string): TierEntryResolved | null {
|
|
return resolveTierEntry({
|
|
runtime: config['runtime'] as string | null | undefined,
|
|
tier,
|
|
overrides: config['model_profile_overrides'] as Record<string, unknown> | null | undefined,
|
|
});
|
|
}
|
|
|
|
// Reverse of the Claude tier-default IDs, plus the Fable alias which Claude
|
|
// Code's Agent tool accepts but which is not a GSD model-profile tier (#1133).
|
|
const CLAUDE_POLICY_ID_TO_ALIAS: Record<string, string> = {
|
|
...Object.fromEntries(
|
|
Object.entries(MODEL_ALIAS_MAP)
|
|
.filter((e): e is [string, string] => typeof e[1] === 'string')
|
|
.map(([aliasName, id]) => [id, aliasName]),
|
|
),
|
|
'claude-fable-5': 'fable',
|
|
};
|
|
// CLAUDE_AGENT_ALIASES moved to ./model-catalog.cts (#3241) — imported above
|
|
// and re-exported below for back-compat (bin/install.js:474,
|
|
// tests/codex-config.test.cjs:24 depend on the name being on this module).
|
|
|
|
// Dedupe stderr warnings so repeated agent resolutions don't spam (#1133).
|
|
const _modelPolicyUnmappableWarned = new Set<string>();
|
|
function warnModelPolicyUnmappable(agentType: string, policyModel: string, tier: string): void {
|
|
const key = `${agentType}::${policyModel}::${tier}`;
|
|
if (_modelPolicyUnmappableWarned.has(key)) return;
|
|
_modelPolicyUnmappableWarned.add(key);
|
|
// MUST go to stderr — resolve-model's JSON result is parsed from stdout.
|
|
process.stderr.write(
|
|
`gsd: warning — model_policy resolved "${policyModel}" for ${agentType}, ` +
|
|
`but it has no Claude agent alias; using "${tier}" instead.\n`,
|
|
);
|
|
}
|
|
|
|
// Test-only: reset the model_policy warn-dedupe cache between cases (#1133).
|
|
function _resetModelPolicyWarningCacheForTests(): void {
|
|
_modelPolicyUnmappableWarned.clear();
|
|
}
|
|
|
|
// Dedupe stderr warnings for unmappable model_overrides Claude IDs (#2041).
|
|
const _modelOverrideUnmappableWarned = new Set<string>();
|
|
function warnModelOverrideUnmappable(agentType: string, overrideValue: string): void {
|
|
const key = `${agentType}::${overrideValue}`;
|
|
if (_modelOverrideUnmappableWarned.has(key)) return;
|
|
_modelOverrideUnmappableWarned.add(key);
|
|
// Cap emission length so an oversized or secret-shaped value cannot leak in
|
|
// full to stderr/logs (#2041 security review). MUST go to stderr — resolve-
|
|
// model's JSON result is parsed from stdout.
|
|
const safe = overrideValue.length > 64 ? overrideValue.slice(0, 64) + '…' : overrideValue;
|
|
process.stderr.write(
|
|
`gsd: warning — model_overrides value "${safe}" for ${agentType} ` +
|
|
`has no Claude agent alias; falling through to tier resolution.\n`,
|
|
);
|
|
}
|
|
|
|
// Test-only: reset the model_overrides warn-dedupe cache between cases (#2041).
|
|
function _resetModelOverrideWarningCacheForTests(): void {
|
|
_modelOverrideUnmappableWarned.clear();
|
|
}
|
|
|
|
/**
|
|
* #2041 — Map a `model_overrides` value to its Claude Agent-tool alias on the
|
|
* claude runtime, mirroring the `model_policy` path (#1144). Claude Code's
|
|
* Agent tool `model` parameter documents only tier aliases (opus/sonnet/haiku/
|
|
* fable); a full Claude model ID returned verbatim is silently dropped by the
|
|
* spawner. Returns the value to return verbatim, or null to signal "fall
|
|
* through to normal tier/dynamic-routing resolution" (used when a Claude full
|
|
* ID has no alias — matches model_policy's warn-and-fall-through). Non-Claude
|
|
* runtimes and non-Claude values always pass through verbatim.
|
|
*
|
|
* Hardening (code+security review): a `typeof` guard preserves the pre-fix
|
|
* no-crash behavior if a malformed config surfaces a non-string value, and an
|
|
* `Object.hasOwn` lookup defeats `__proto__`/`constructor` lookups on the plain
|
|
* object literal so those reserved keys cannot return a truthy non-string.
|
|
*/
|
|
function mapClaudeOverrideForRuntime(
|
|
override: string,
|
|
configRuntime: string | null | undefined,
|
|
agentType: string,
|
|
): string | null {
|
|
// Defensive: model_overrides is typed Record<string,string> but a malformed
|
|
// config could surface a non-string; pass through verbatim (preserving the
|
|
// pre-fix no-crash behaviour) and let the downstream Agent tool reject it.
|
|
if (typeof override !== 'string') return override;
|
|
const onClaude = !configRuntime || configRuntime === 'claude';
|
|
if (!onClaude) return override;
|
|
// Object.hasOwn guards against __proto__/constructor returning a truthy
|
|
// non-string from the plain object literal (#2041 security review).
|
|
if (Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, override)) {
|
|
return CLAUDE_POLICY_ID_TO_ALIAS[override];
|
|
}
|
|
if (CLAUDE_AGENT_ALIASES.has(override)) return override;
|
|
if (override.startsWith('claude-')) {
|
|
warnModelOverrideUnmappable(agentType, override);
|
|
return null;
|
|
}
|
|
return override;
|
|
}
|
|
|
|
/**
|
|
* #49 — Provider-neutral model policy preset resolution.
|
|
*/
|
|
function resolveModelPolicy(policy: Record<string, unknown> | null | undefined, tier: string | null | undefined): string | null {
|
|
if (!policy || typeof policy !== 'object') return null;
|
|
if (!tier) return null;
|
|
|
|
const runtime = policy['runtime'];
|
|
const rtOverrides = policy['runtime_tiers'];
|
|
if (runtime && typeof runtime === 'string' && rtOverrides && typeof rtOverrides === 'object') {
|
|
const rtOverridesMap = rtOverrides as Record<string, unknown>;
|
|
if (Object.hasOwn(rtOverridesMap, runtime)) {
|
|
const runtimeEntry = rtOverridesMap[runtime];
|
|
if (runtimeEntry && typeof runtimeEntry === 'object' && Object.hasOwn(runtimeEntry, tier)) {
|
|
const raw = (runtimeEntry as Record<string, unknown>)[tier];
|
|
if (raw != null) {
|
|
const entry = typeof raw === 'string' ? { model: raw } : (raw as Record<string, unknown>);
|
|
if (entry && entry['model']) return entry['model'] as string;
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
const provider = policy['provider'];
|
|
if (!provider || typeof provider !== 'string') return null;
|
|
|
|
if (provider === 'generic' || provider === 'custom') {
|
|
const TIER_TO_POLICY_KEY: Record<string, string> = { opus: 'high', sonnet: 'medium', haiku: 'low' };
|
|
const policyKey = TIER_TO_POLICY_KEY[tier];
|
|
if (!policyKey) return null;
|
|
const v = policy[policyKey];
|
|
return (v && typeof v === 'string') ? v : null;
|
|
}
|
|
|
|
const presetsMap = PROVIDER_PRESETS as Record<string, Record<string, Record<string, { model: string } | null>>>;
|
|
if (!Object.hasOwn(presetsMap, provider)) return null;
|
|
const presetForProvider = presetsMap[provider];
|
|
if (!presetForProvider || typeof presetForProvider !== 'object') return null;
|
|
|
|
if (!Object.hasOwn(presetForProvider, tier)) return null;
|
|
const tierPresets = presetForProvider[tier];
|
|
if (!tierPresets || typeof tierPresets !== 'object') return null;
|
|
|
|
const budget = (policy['budget'] && typeof policy['budget'] === 'string') ? policy['budget'] : 'medium';
|
|
if (!Object.hasOwn(tierPresets, budget)) return null;
|
|
const budgetEntry = tierPresets[budget];
|
|
if (!budgetEntry || !budgetEntry.model) return null;
|
|
|
|
return budgetEntry.model;
|
|
}
|
|
|
|
/**
|
|
* #2229 — the profile/phase-type tier for (config, agentType).
|
|
*
|
|
* Extracted verbatim from resolveModelInternal's step 2 so the same expression can
|
|
* answer "which tier did GSD resolve?" without also resolving a model id. The
|
|
* extraction is behaviour-preserving by construction: resolveModelInternal calls
|
|
* straight back into it.
|
|
*
|
|
* Returns null when the agent has no catalog entry and the profile is not `inherit`.
|
|
*/
|
|
function computeProfileTier(config: Record<string, unknown>, agentType: string): string | null {
|
|
// eslint-disable-next-line @typescript-eslint/no-base-to-string
|
|
const profile = String(config['model_profile'] || 'balanced').toLowerCase();
|
|
// Own-property guard: agentType is an unvalidated CLI positional (the
|
|
// `resolve-model <agent-type>` argument is never checked against a known
|
|
// agent list), so a prototype-chain agentType ("toString", "constructor")
|
|
// would otherwise return an inherited member from this plain object
|
|
// instead of undefined — verified reachable purely via the CLI.
|
|
const modelProfilesMap = MODEL_PROFILES as unknown as Record<string, Record<string, string>>;
|
|
const agentModels = Object.hasOwn(modelProfilesMap, agentType) ? modelProfilesMap[agentType] : undefined;
|
|
const phaseType = (AGENT_TO_PHASE_TYPE)[agentType];
|
|
const configModels = config['models'] as Record<string, string> | null | undefined;
|
|
const phaseTypeTier = (phaseType && configModels && typeof configModels === 'object')
|
|
? configModels[phaseType]
|
|
: undefined;
|
|
return (phaseTypeTier && VALID_TIERS.has(phaseTypeTier))
|
|
? phaseTypeTier
|
|
: (profile === 'inherit'
|
|
? 'inherit'
|
|
: (agentModels
|
|
// Own-property guard: `profile` is a config-supplied string
|
|
// (config['model_profile'], lower-cased); an already-lowercase
|
|
// prototype-chain key ("constructor", "__proto__") would otherwise
|
|
// return an inherited non-string member instead of falling back to
|
|
// 'balanced' (verified: profile:"constructor"/"__proto__" leaked a
|
|
// function/object through both the tier and model resolution paths).
|
|
? ((Object.hasOwn(agentModels, profile) ? agentModels[profile] : undefined) || agentModels['balanced'])
|
|
: null));
|
|
}
|
|
|
|
/**
|
|
* #2229 — the effective model TIER for (config, agentType), as a signal a workflow can
|
|
* read: `gsd_run query resolve-model <agent> --pick tier`.
|
|
*
|
|
* Why this is not just "look at the resolved model": on every runtime the installer
|
|
* configures with `resolve_model_ids: "omit"` — which is every non-Claude runtime, see
|
|
* docs/CONFIGURATION.md — resolveModelInternal deliberately returns '' below. A guard
|
|
* keyed on the model id therefore cannot tell a budget-tier run from a top-tier one
|
|
* there, while the tier itself is computed ABOVE that early-return and stays knowable.
|
|
*
|
|
* Honesty contract — this never guesses, because a guard that reports a wrong tier is
|
|
* worse than one that reports none:
|
|
* - a per-agent `model_overrides` pin naming a known alias (or a full Claude id that
|
|
* maps to one) reports that alias;
|
|
* - a pin that maps to nothing reports 'unknown' — a raw model id carries no tier;
|
|
* - `model_profile: inherit` reports 'inherit' — the session model is not ours to name;
|
|
* - an agent with no catalog entry reports 'unknown'.
|
|
*
|
|
* Callers must treat 'unknown' and 'inherit' as "cannot tell", never as "adequate".
|
|
*/
|
|
function resolveTierFromConfig(config: Record<string, unknown>, agentType: string): string {
|
|
const rawOverrides = config['model_overrides'];
|
|
const modelOverrides = (rawOverrides && typeof rawOverrides === 'object' && !Array.isArray(rawOverrides))
|
|
? rawOverrides as Record<string, string>
|
|
: null;
|
|
// Own-property guard: agentType is a caller-supplied string (the
|
|
// `resolve-model <agent-type>` CLI positional is not validated against a
|
|
// known agent list); a prototype-chain agentType ("toString",
|
|
// "constructor") against ANY model_overrides object — even `{}` — would
|
|
// otherwise return an inherited member instead of undefined.
|
|
const override = (modelOverrides && Object.hasOwn(modelOverrides, agentType))
|
|
? modelOverrides[agentType]
|
|
: undefined;
|
|
if (override && typeof override === 'string') {
|
|
if (CLAUDE_AGENT_ALIASES.has(override)) return override;
|
|
// Own-property guard: this indexes a plain object with a config-supplied
|
|
// string, so a prototype-chain key ("toString", "constructor", "valueOf")
|
|
// would otherwise return an inherited member instead of undefined — and a
|
|
// function-valued tier is dropped entirely by JSON.stringify, silently
|
|
// removing the key a guard depends on.
|
|
const alias = Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, override)
|
|
? CLAUDE_POLICY_ID_TO_ALIAS[override]
|
|
: undefined;
|
|
if (typeof alias === 'string' && alias) return alias;
|
|
return 'unknown';
|
|
}
|
|
|
|
const profileTier = computeProfileTier(config, agentType);
|
|
|
|
// #3282 — mirror resolveModelInternal's step 2.5 (model_policy preset). The
|
|
// profile tier alone under-reports: model_policy can dispatch a DIFFERENT
|
|
// tier than the profile implies (e.g. a `balanced` profile's "sonnet" tier
|
|
// combined with `model_policy: {budget: 'low'}` actually spawns "haiku"),
|
|
// and reporting the profile tier there is exactly the under-report this
|
|
// fixes — a haiku-tier run must never be reported as "sonnet". Skipped
|
|
// under the same condition resolveModelInternal skips it (no tier, or
|
|
// "inherit" — the session model is not ours to name).
|
|
if (profileTier && profileTier !== 'inherit') {
|
|
const mergedPolicy = config['model_policy']
|
|
? { ...(config['model_policy'] as Record<string, unknown>), runtime: (config['runtime'] as string | null | undefined) || 'claude' }
|
|
: null;
|
|
const policyModel = resolveModelPolicy(mergedPolicy, profileTier);
|
|
if (policyModel) {
|
|
// Map the policy-resolved id back to a tier alias with the same
|
|
// own-property-guarded lookups used above. If it maps, that alias IS
|
|
// the tier that actually runs — report it (the fix). If it does not
|
|
// map — including every non-Claude runtime, where resolveModelInternal
|
|
// returns the policy model verbatim with no tier meaning — the model
|
|
// carries no tier we can name; report 'unknown' rather than falling
|
|
// back to the profile tier, which would silently reintroduce the
|
|
// under-report this block exists to close.
|
|
const aliasForId = Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, policyModel)
|
|
? CLAUDE_POLICY_ID_TO_ALIAS[policyModel]
|
|
: undefined;
|
|
if (typeof aliasForId === 'string' && aliasForId) return aliasForId;
|
|
if (CLAUDE_AGENT_ALIASES.has(policyModel)) return policyModel;
|
|
return 'unknown';
|
|
}
|
|
}
|
|
|
|
return profileTier || 'unknown';
|
|
}
|
|
|
|
function resolveTierInternal(cwd: string, agentType: string): string {
|
|
return resolveTierFromConfig(loadConfig(cwd), agentType);
|
|
}
|
|
|
|
function resolveModelInternal(cwd: string, agentType: string): string {
|
|
const config = loadConfig(cwd);
|
|
|
|
// 1. Per-agent override (#2041: map Claude full IDs → Agent-tool aliases on
|
|
// the claude runtime, mirroring the model_policy path #1144; non-Claude
|
|
// runtimes and non-Claude values pass through verbatim).
|
|
const modelOverrides = config['model_overrides'] as Record<string, string> | null | undefined;
|
|
// Own-property guard (see resolveTierFromConfig above): without it, an
|
|
// agentType of "toString" against `model_overrides: {}` returned the
|
|
// inherited Function.prototype.toString as the resolved "model" — verified
|
|
// reachable purely via the CLI, no override value needed.
|
|
const override = (modelOverrides && Object.hasOwn(modelOverrides, agentType))
|
|
? modelOverrides[agentType]
|
|
: undefined;
|
|
if (override) {
|
|
const mapped = mapClaudeOverrideForRuntime(override, config['runtime'] as string | null | undefined, agentType);
|
|
if (mapped !== null) return mapped;
|
|
// Unmappable Claude ID — fall through to tier resolution (matches model_policy).
|
|
}
|
|
|
|
// 2. Compute the tier (#2229: shared with resolveTierFromConfig so the tier a
|
|
// workflow reads and the tier a model is resolved from can never diverge).
|
|
// eslint-disable-next-line @typescript-eslint/no-base-to-string
|
|
const profile = String(config['model_profile'] || 'balanced').toLowerCase();
|
|
// Own-property guard (see computeProfileTier above): without it, agentType
|
|
// "toString" returned Function.prototype.toString as `agentModels`
|
|
// (truthy), which skipped the "unknown agent" fallback below and made
|
|
// resolveModelInternal return undefined instead of a tier-derived string.
|
|
const modelProfilesMapForModel = MODEL_PROFILES as unknown as Record<string, Record<string, string>>;
|
|
const agentModels = Object.hasOwn(modelProfilesMapForModel, agentType) ? modelProfilesMapForModel[agentType] : undefined;
|
|
const tier = computeProfileTier(config, agentType);
|
|
|
|
// 2.5. model_policy preset (#49, #1133)
|
|
const configRuntime = config['runtime'] as string | null | undefined;
|
|
if (tier && tier !== 'inherit') {
|
|
const onClaude = !configRuntime || configRuntime === 'claude';
|
|
const effectiveRuntime = configRuntime || 'claude';
|
|
const mergedPolicy = config['model_policy']
|
|
? { ...(config['model_policy'] as Record<string, unknown>), runtime: effectiveRuntime }
|
|
: null;
|
|
const policyModel = resolveModelPolicy(mergedPolicy, tier);
|
|
if (policyModel) {
|
|
// Non-Claude runtimes take full model IDs verbatim (unchanged behavior).
|
|
if (!onClaude) return policyModel;
|
|
// Claude Code's Agent tool takes tier aliases (opus/sonnet/haiku/fable),
|
|
// not full model IDs — map the policy-resolved ID back to an alias (#1133).
|
|
const aliasForId = Object.hasOwn(CLAUDE_POLICY_ID_TO_ALIAS, policyModel)
|
|
? CLAUDE_POLICY_ID_TO_ALIAS[policyModel]
|
|
: undefined;
|
|
if (typeof aliasForId === 'string' && aliasForId) return aliasForId;
|
|
// The policy value may already be a bare Claude agent alias (e.g. "fable").
|
|
if (CLAUDE_AGENT_ALIASES.has(policyModel)) return policyModel;
|
|
// No Claude alias for this ID (e.g. a pinned minor version like
|
|
// claude-opus-4-5). Warn once and fall through to the tier alias rather
|
|
// than returning an ID Claude Code cannot spawn.
|
|
warnModelPolicyUnmappable(agentType, policyModel, tier);
|
|
}
|
|
}
|
|
|
|
// 3. Runtime-aware resolution (#2517)
|
|
if (configRuntime && configRuntime !== 'claude' && tier && tier !== 'inherit') {
|
|
const entry = _resolveRuntimeTier(config, tier);
|
|
if (entry?.model) return entry.model;
|
|
}
|
|
|
|
// 4. resolve_model_ids: "omit" — runtime-aware (#2297). Honor "omit" when the
|
|
// PROJECT explicitly set it (user intent — project config wins, #2517 finding
|
|
// #4) OR when the active runtime genuinely lacks native model aliases. Only a
|
|
// native-alias runtime (Claude) ignores an "omit" that came solely from the
|
|
// SHARED ~/.gsd/defaults.json — the #2297 poisoning fix — and falls through to
|
|
// its tier aliases below. Active runtime: GSD_RUNTIME → config.runtime → the
|
|
// per-install .gsd-runtime marker → 'claude' (canonicalized).
|
|
// NOTE: a non-Claude runtime that HAS a populated runtime-tier map already
|
|
// returned its own model id at step 3 above, before this gate — for those the
|
|
// explicit-project-omit honoring here is moot (step 3 wins, by #2517 design).
|
|
if (config['resolve_model_ids'] === 'omit'
|
|
&& (projectExplicitlySetsOmit(cwd) || !RUNTIMES_WITH_NATIVE_ALIASES.has(resolveActiveRuntime(config)))) {
|
|
return '';
|
|
}
|
|
|
|
// 5. Profile lookup (Claude-native default).
|
|
if (!agentModels) {
|
|
return profile === 'quality' ? 'opus'
|
|
: profile === 'budget' ? 'haiku'
|
|
: profile === 'inherit' ? 'inherit'
|
|
: 'sonnet';
|
|
}
|
|
if (tier === 'inherit') return 'inherit';
|
|
const alias = tier;
|
|
|
|
// Only the explicit `true` opt-in materializes full model IDs (#1569). Guard
|
|
// against the loose-truthy check catching a "omit" that a native-alias runtime
|
|
// ignored above (#2297): "omit" must fall through to the tier ALIAS here, not
|
|
// be materialized into a full ID Claude's Agent tool cannot spawn.
|
|
if (config['resolve_model_ids'] === true) {
|
|
return (MODEL_ALIAS_MAP as Record<string, string>)[alias!] || alias!;
|
|
}
|
|
|
|
return alias!;
|
|
}
|
|
|
|
const VALID_GRANULARITIES = new Set(['coarse', 'standard', 'fine']);
|
|
|
|
/**
|
|
* Resolve the planning granularity for a phase type (#68).
|
|
*/
|
|
function resolveGranularityInternal(cwd: string, phaseType: string | null | undefined, override?: string | null): string {
|
|
if (override !== undefined && override !== null && override !== '') {
|
|
if (VALID_GRANULARITIES.has(override)) {
|
|
return override;
|
|
}
|
|
}
|
|
const config = loadConfig(cwd);
|
|
const configGranularities = config['granularities'] as Record<string, string> | null | undefined;
|
|
const perPhase = (phaseType && configGranularities && typeof configGranularities === 'object')
|
|
? configGranularities[phaseType]
|
|
: undefined;
|
|
if (perPhase && VALID_GRANULARITIES.has(perPhase)) {
|
|
return perPhase;
|
|
}
|
|
if (config['granularity'] !== undefined && config['granularity'] !== null && config['granularity'] !== '') {
|
|
return config['granularity'] as string;
|
|
}
|
|
const planning = config['planning'] as Record<string, unknown> | null | undefined;
|
|
const planningGran = planning && planning['granularity'];
|
|
if (planningGran !== undefined && planningGran !== null && planningGran !== '') {
|
|
return planningGran as string;
|
|
}
|
|
return 'standard';
|
|
}
|
|
|
|
/**
|
|
* Validate a CLI granularity override at the command boundary. Empty/null/undefined
|
|
* are treated as "no override" (no-op). An invalid non-empty value calls `fail`.
|
|
*/
|
|
function assertValidGranularityOverride(
|
|
override: string | null | undefined,
|
|
fail: (msg: string) => never,
|
|
): void {
|
|
if (override !== undefined && override !== null && override !== '' && !VALID_GRANULARITIES.has(override)) {
|
|
fail(`invalid granularity '${override}' (valid: ${[...VALID_GRANULARITIES].join(', ')})`);
|
|
}
|
|
}
|
|
|
|
/**
|
|
* #3024 — Resolve a model for a specific dynamic-routing attempt.
|
|
*/
|
|
function resolveModelForTier(cwd: string, agentType: string, attempt?: number): string {
|
|
const config = loadConfig(cwd);
|
|
const attemptN = Number.isInteger(attempt) && (attempt as number) > 0 ? (attempt as number) : 0;
|
|
|
|
const modelOverrides = config['model_overrides'] as Record<string, string> | null | undefined;
|
|
// Own-property guard (see resolveTierFromConfig above): without it, an
|
|
// agentType of "toString" against `model_overrides: {}` returned the
|
|
// inherited Function.prototype.toString as the resolved "model" — verified
|
|
// reachable purely via the CLI, no override value needed.
|
|
const override = (modelOverrides && Object.hasOwn(modelOverrides, agentType))
|
|
? modelOverrides[agentType]
|
|
: undefined;
|
|
if (override) {
|
|
const mapped = mapClaudeOverrideForRuntime(override, config['runtime'] as string | null | undefined, agentType);
|
|
if (mapped !== null) return mapped;
|
|
// Unmappable Claude ID — fall through to dynamic_routing / model_policy resolution.
|
|
}
|
|
|
|
if (config['model_policy'] && config['runtime'] && config['runtime'] !== 'claude') {
|
|
return resolveModelInternal(cwd, agentType);
|
|
}
|
|
|
|
const dr = config['dynamic_routing'] as Record<string, unknown> | null | undefined;
|
|
if (!dr || typeof dr !== 'object' || dr['enabled'] !== true) {
|
|
return resolveModelInternal(cwd, agentType);
|
|
}
|
|
|
|
const tierModels = dr['tier_models'] as Record<string, string> | null | undefined;
|
|
if (!tierModels || typeof tierModels !== 'object') {
|
|
return resolveModelInternal(cwd, agentType);
|
|
}
|
|
|
|
const defaultTier = (AGENT_DEFAULT_TIERS)[agentType];
|
|
if (!defaultTier || !(VALID_AGENT_TIERS).has(defaultTier)) {
|
|
return resolveModelInternal(cwd, agentType);
|
|
}
|
|
|
|
const maxEscalations = Number.isInteger(dr['max_escalations']) && (dr['max_escalations'] as number) >= 0
|
|
? (dr['max_escalations'] as number)
|
|
: 1;
|
|
const escalationEnabled = dr['escalate_on_failure'] !== false;
|
|
const effectiveAttempt = escalationEnabled
|
|
? Math.min(attemptN, maxEscalations)
|
|
: 0;
|
|
|
|
let tier = defaultTier;
|
|
for (let i = 0; i < effectiveAttempt; i += 1) {
|
|
const next = (nextTier)(tier);
|
|
if (!next || next === tier) break;
|
|
tier = next;
|
|
}
|
|
|
|
const alias = tierModels[tier];
|
|
if (typeof alias !== 'string' || alias.length === 0) {
|
|
return resolveModelInternal(cwd, agentType);
|
|
}
|
|
return alias;
|
|
}
|
|
|
|
/**
|
|
* #2296 — Outcome of consulting the provider-escalation ladder for one attempt.
|
|
*
|
|
* `from`/`to` describe the PROVIDER ladder only; they are equal whenever the
|
|
* ladder was not consulted or had nothing to offer.
|
|
*/
|
|
interface ProviderEscalationResult {
|
|
from: string;
|
|
to: string;
|
|
escalated: boolean;
|
|
exhausted: boolean;
|
|
attempted: string[];
|
|
index: number;
|
|
}
|
|
|
|
/**
|
|
* Keep only usable model ids: non-empty strings. A malformed config can put
|
|
* anything in here (nulls, numbers, blank strings), and a blank model id would
|
|
* resolve to an unusable agent invocation rather than failing visibly. Invalid
|
|
* entries are dropped and the surviving order is preserved, so the ladder stays
|
|
* predictable (ADR 227 — validate shape, not just type).
|
|
*/
|
|
function sanitizeProviderEscalation(raw: unknown): string[] {
|
|
if (!Array.isArray(raw)) return [];
|
|
return raw.filter((entry): entry is string => typeof entry === 'string' && entry.trim().length > 0);
|
|
}
|
|
|
|
/**
|
|
* #2296 — Resolve the model for one attempt of the PROVIDER escalation ladder.
|
|
*
|
|
* The tier ladder (`resolveModelForTier`) escalates within one provider's
|
|
* `tier_models`, which does not help when that provider is the thing that is
|
|
* throttled. This walks `dynamic_routing.provider_escalation` instead: an
|
|
* ordered list of alternative model ids, capped by
|
|
* `min(max_escalations, list length)`.
|
|
*
|
|
* `applicable` is the caller's policy decision (only a quota-exceeded
|
|
* classification should consult this ladder). It is a parameter rather than a
|
|
* class check here so this module keeps depending only on leaf modules, per
|
|
* CONTEXT.md's Model Resolution module contract.
|
|
*
|
|
* Attempt 0 — and every non-applicable call — stays on the source model.
|
|
* `exhausted` reports that the ladder is spent so the caller can fail loudly
|
|
* naming every model it tried.
|
|
*/
|
|
function resolveProviderEscalation(
|
|
cwd: string,
|
|
agentType: string,
|
|
attempt: number | undefined,
|
|
applicable: boolean,
|
|
): ProviderEscalationResult {
|
|
// The model that would be used with no provider escalation at all.
|
|
const from = resolveModelForTier(cwd, agentType, 0);
|
|
const stay = (exhausted = false): ProviderEscalationResult => ({
|
|
from,
|
|
to: from,
|
|
escalated: false,
|
|
exhausted,
|
|
attempted: [from],
|
|
index: 0,
|
|
});
|
|
|
|
if (!applicable) return stay();
|
|
|
|
const config = loadConfig(cwd);
|
|
const dr = config['dynamic_routing'] as Record<string, unknown> | null | undefined;
|
|
if (!dr || typeof dr !== 'object' || dr['enabled'] !== true) return stay();
|
|
if (dr['escalate_on_failure'] === false) return stay();
|
|
|
|
const list = sanitizeProviderEscalation(dr['provider_escalation']);
|
|
if (list.length === 0) return stay();
|
|
|
|
// Same default and same validity rule as the tier ladder above — a negative or
|
|
// non-integer max_escalations is invalid config, not a request for zero.
|
|
const maxEscalations = Number.isInteger(dr['max_escalations']) && (dr['max_escalations'] as number) >= 0
|
|
? (dr['max_escalations'] as number)
|
|
: 1;
|
|
const cap = Math.min(maxEscalations, list.length);
|
|
// An explicit cap of 0 means the ladder exists but is spent before it starts.
|
|
if (cap === 0) return stay(true);
|
|
|
|
const attemptN = Number.isInteger(attempt) && (attempt as number) > 0 ? (attempt as number) : 0;
|
|
if (attemptN === 0) return stay();
|
|
|
|
const index = Math.min(attemptN, cap);
|
|
return {
|
|
from,
|
|
to: list[index - 1],
|
|
escalated: true,
|
|
exhausted: attemptN > cap,
|
|
attempted: [from, ...list.slice(0, index)],
|
|
index,
|
|
};
|
|
}
|
|
|
|
// ─── #443 — Unified effort + fast_mode resolvers ─────────────────────────────
|
|
|
|
const VALID_EFFORTS = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max'];
|
|
// #3533 (10d): the VOCABULARY carries one more member than the LADDER —
|
|
// 'inherit' is a declarable effort choice ("follow the session", expressed by
|
|
// OMITTING the effort key at the writer) but not a level nextEffort may step
|
|
// into. Keeping it out of VALID_EFFORTS means escalation (resolveEffortForTier)
|
|
// never walks past an explicit inherit: nextEffort('inherit') is null.
|
|
const EFFORT_SET = new Set([...VALID_EFFORTS, 'inherit']);
|
|
|
|
/**
|
|
* Walk one step up the effort ladder from `e`.
|
|
*/
|
|
function nextEffort(e: string): string | null {
|
|
const i = VALID_EFFORTS.indexOf(e);
|
|
if (i < 0) return null;
|
|
return VALID_EFFORTS[Math.min(i + 1, VALID_EFFORTS.length - 1)];
|
|
}
|
|
|
|
interface EffortOpts {
|
|
override?: string;
|
|
}
|
|
|
|
interface FastModeOpts {
|
|
override?: boolean;
|
|
}
|
|
|
|
/**
|
|
* #443 — Resolve a universal effort string for (cwd, agentType).
|
|
*/
|
|
function resolveEffortInternal(cwd: string, agentType: string, opts?: EffortOpts): string {
|
|
// Step 1: invocation override
|
|
if (opts && typeof opts.override === 'string' && EFFORT_SET.has(opts.override)) {
|
|
return opts.override;
|
|
}
|
|
|
|
const config = loadConfig(cwd);
|
|
const effortCfg = (config['effort'] && typeof config['effort'] === 'object' && !Array.isArray(config['effort']))
|
|
? (config['effort'] as Record<string, unknown>)
|
|
: null;
|
|
|
|
// Step 2: agent_overrides
|
|
if (effortCfg) {
|
|
const ao = effortCfg['agent_overrides'];
|
|
if (ao && typeof ao === 'object' && !Array.isArray(ao)) {
|
|
const v = (ao as Record<string, unknown>)[agentType];
|
|
if (typeof v === 'string' && EFFORT_SET.has(v)) return v;
|
|
}
|
|
} else {
|
|
const canonicalEffort = (CANONICAL_CONFIG_DEFAULTS)['effort'];
|
|
const mao = canonicalEffort && typeof canonicalEffort === 'object'
|
|
? (canonicalEffort as Record<string, unknown>)['agent_overrides']
|
|
: undefined;
|
|
if (mao && typeof mao === 'object' && !Array.isArray(mao)) {
|
|
const v = (mao as Record<string, unknown>)[agentType];
|
|
if (typeof v === 'string' && EFFORT_SET.has(v)) return v;
|
|
}
|
|
}
|
|
|
|
// Step 3: routing_tier_defaults by agent's default tier.
|
|
// #3531 (10c): the config block merges OVER the manifest tier defaults
|
|
// rather than replacing them — an effort block without
|
|
// routing_tier_defaults (or missing this agent's tier) falls back to the
|
|
// manifest built-in for that tier instead of skipping to effort.default.
|
|
// Invalid config values are dropped by the merge, so the manifest value for
|
|
// the tier surfaces (the same "invalid falls through" rule every layer has).
|
|
const agentTier = (AGENT_DEFAULT_TIERS)[agentType];
|
|
if (agentTier) {
|
|
const canonicalEffort = (CANONICAL_CONFIG_DEFAULTS)['effort'];
|
|
const manifestDefaults = canonicalEffort && typeof canonicalEffort === 'object'
|
|
? (canonicalEffort as Record<string, unknown>)['routing_tier_defaults'] as Record<string, string> | undefined
|
|
: undefined;
|
|
const isValidEffort = (v: unknown): v is string => typeof v === 'string' && EFFORT_SET.has(v);
|
|
const merged = mergeEffortTierDefaults(
|
|
manifestDefaults,
|
|
effortCfg ? effortCfg['routing_tier_defaults'] : undefined,
|
|
isValidEffort,
|
|
);
|
|
const v = merged[agentTier];
|
|
if (isValidEffort(v)) return v;
|
|
}
|
|
|
|
// Step 4: effort.default
|
|
if (effortCfg) {
|
|
const d = effortCfg['default'];
|
|
if (typeof d === 'string' && EFFORT_SET.has(d)) return d;
|
|
} else {
|
|
const canonicalEffort = (CANONICAL_CONFIG_DEFAULTS)['effort'];
|
|
const d = canonicalEffort && typeof canonicalEffort === 'object'
|
|
? (canonicalEffort as Record<string, unknown>)['default']
|
|
: undefined;
|
|
if (typeof d === 'string' && EFFORT_SET.has(d)) return d;
|
|
}
|
|
|
|
// Step 5: hardcoded default
|
|
return 'high';
|
|
}
|
|
|
|
/**
|
|
* #443 — Resolve fast_mode boolean for (cwd, agentType).
|
|
*/
|
|
function resolveFastModeInternal(cwd: string, agentType: string, opts?: FastModeOpts): boolean {
|
|
// Step 1: invocation override
|
|
if (opts && typeof opts.override === 'boolean') {
|
|
return opts.override;
|
|
}
|
|
|
|
const config = loadConfig(cwd);
|
|
const fmCfg = (config['fast_mode'] && typeof config['fast_mode'] === 'object' && !Array.isArray(config['fast_mode']))
|
|
? (config['fast_mode'] as Record<string, unknown>)
|
|
: null;
|
|
|
|
// Step 2: agent_overrides
|
|
if (fmCfg) {
|
|
const ao = fmCfg['agent_overrides'];
|
|
if (ao && typeof ao === 'object' && !Array.isArray(ao)) {
|
|
const v = (ao as Record<string, unknown>)[agentType];
|
|
if (typeof v === 'boolean') return v;
|
|
}
|
|
}
|
|
|
|
// Step 3: routing_tier_defaults by agent's default tier.
|
|
const agentTier = (AGENT_DEFAULT_TIERS)[agentType];
|
|
if (agentTier) {
|
|
if (fmCfg && fmCfg['routing_tier_defaults'] &&
|
|
typeof fmCfg['routing_tier_defaults'] === 'object' &&
|
|
!Array.isArray(fmCfg['routing_tier_defaults'])) {
|
|
const v = (fmCfg['routing_tier_defaults'] as Record<string, unknown>)[agentTier];
|
|
if (typeof v === 'boolean') return v;
|
|
} else if (!fmCfg) {
|
|
const canonicalFm = (CANONICAL_CONFIG_DEFAULTS)['fast_mode'];
|
|
const manifestDefaults = canonicalFm && typeof canonicalFm === 'object'
|
|
? (canonicalFm as Record<string, unknown>)['routing_tier_defaults']
|
|
: undefined;
|
|
if (manifestDefaults && typeof manifestDefaults === 'object') {
|
|
const v = (manifestDefaults as Record<string, unknown>)[agentTier];
|
|
if (typeof v === 'boolean') return v;
|
|
}
|
|
}
|
|
}
|
|
|
|
// Step 4: fast_mode.enabled
|
|
if (fmCfg && typeof fmCfg['enabled'] === 'boolean') {
|
|
return fmCfg['enabled'];
|
|
}
|
|
|
|
// Step 5: hardcoded default
|
|
return false;
|
|
}
|
|
|
|
/**
|
|
* #443 — Resolve effort for a dynamic-routing attempt (with escalation).
|
|
*/
|
|
function resolveEffortForTier(cwd: string, agentType: string, attempt?: number): string {
|
|
const base = resolveEffortInternal(cwd, agentType);
|
|
|
|
const config = loadConfig(cwd);
|
|
const dr = config['dynamic_routing'] as Record<string, unknown> | null | undefined;
|
|
if (!dr || typeof dr !== 'object' || dr['enabled'] !== true) {
|
|
return base;
|
|
}
|
|
if (dr['escalate_on_failure'] === false) {
|
|
return base;
|
|
}
|
|
|
|
const maxEscalations = Number.isInteger(dr['max_escalations']) && (dr['max_escalations'] as number) >= 0
|
|
? (dr['max_escalations'] as number)
|
|
: 1;
|
|
|
|
const attemptN = Number.isInteger(attempt) && (attempt as number) > 0 ? (attempt as number) : 0;
|
|
const effectiveAttempt = Math.min(attemptN, maxEscalations);
|
|
|
|
let current = base;
|
|
for (let i = 0; i < effectiveAttempt; i++) {
|
|
const next = nextEffort(current);
|
|
if (!next || next === current) break;
|
|
current = next;
|
|
}
|
|
return current;
|
|
}
|
|
|
|
export = {
|
|
resolveTierEntry,
|
|
CLAUDE_AGENT_ALIASES,
|
|
resolveModelPolicy,
|
|
resolveModelInternal,
|
|
resolveTierInternal,
|
|
resolveTierFromConfig,
|
|
_resetModelPolicyWarningCacheForTests,
|
|
_resetModelOverrideWarningCacheForTests,
|
|
_setInstallRuntimeMarkerForTests,
|
|
_resetInstallRuntimeMarkerCacheForTests,
|
|
VALID_GRANULARITIES,
|
|
resolveGranularityInternal,
|
|
assertValidGranularityOverride,
|
|
resolveModelForTier,
|
|
resolveProviderEscalation,
|
|
VALID_EFFORTS,
|
|
EFFORT_SET,
|
|
nextEffort,
|
|
resolveEffortInternal,
|
|
resolveFastModeInternal,
|
|
resolveEffortForTier,
|
|
};
|