From 004e9dd7411f33dd5ca7ee3881e4a8a38c75396c Mon Sep 17 00:00:00 2001 From: Tom Boucher Date: Sat, 22 Aug 2026 20:51:55 -0400 Subject: [PATCH] fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * test(#3007): failing-first suite for per-model Codex effort capability RED by construction. Binds to behavior renderEffortForRuntime does not yet have: an optional third `model` argument, a per-model advertised-level table, `max` passing through instead of clamping to `xhigh`, `minimal` clamping to `low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/ `reason`) so a downgrade is legible from resolver output rather than silent. Two of these pin defects that exist on next today: - `max` is discarded. Both Codex models whose catalog entries are retrievable (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing. - `minimal` is emitted to a model that refuses it. providerPresets.openai. haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's advertised floor is `low`. GSD is sending a value into a document Codex itself validates. The parity test is what pins that fixed, and it names the offending path/model/effort when it trips. Also corrects tests/model-resolver.test.cjs:351, which asserted renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as though it were a contract. ADR-443 recorded "Codex has no max" as fact and it was true when written; Codex has since added both `max` and `ultra`. That is a stale premise, so the assertion is corrected here rather than worked around. The property test asserts the invariant the whole change exists for: a rendered effort is always a level the target model actually advertises, or an explicit rejection. There is no third outcome. * fix(#3007): resolve Codex effort per model, and make every clamp visible Codex declares supported_reasoning_levels per MODEL and validates against it, so a single per-runtime capability set cannot be right for all of them. GSD's was wrong in both directions at once. `max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped max -> xhigh on that basis; it was accurate when written, and Codex has since added both `max` and `ultra`. Every Codex model whose catalog entry is retrievable advertises `max`, so the clamp was discarding a level the provider supports, silently, on the most-used path. `minimal` stops reaching Codex. No Codex model advertises it -- both retrievable entries floor at `low` -- yet providerPresets.openai.haiku.low paired gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the receiver validates and refuses into a file the receiver reads. Being unconservative in what you send is the half of Postel's rule with no defensible reading, so that preset is corrected and a parity test pins it. `ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum reasoning with automatic task delegation": at ultra, effective_multi_agent_mode returns Proactive and Codex spawns sub-agents on its own initiative, underneath GSD's orchestration rather than inside it (#2167). It is a mode switch, not a reasoning depth, so it is not added to the universal ladder -- which stays provider-agnostic by ADR-443's design -- and it is rejected even for gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered and rejected: that silently discards what the user actually asked for. Clamping is now visible. RenderedEffort carries requested/clamped/reason and resolve-execution surfaces them. The previous table clamped correctly but invisibly, so a user asking for `max` on Codex had no way to find out they were getting `xhigh` -- exactly the failure mode the robustness principle's modern critique warns about, and why "be liberal" has to mean "liberal and loud". Also closes a latent trap found while reviewing the implementation: the clamp-up loop walks the ladder upward, and for a future model advertising `ultra` but not `max` it would have selected `ultra` as the clamp target -- re-entering by the back door the mode the rejection above exists to keep out. A clamp may never produce a value that a direct request for that value would refuse. Unreachable with today's catalog, which is why no test caught it; a test now asserts the invariant directly. Signature stability is preserved: the third `model` argument is optional and the two-argument form still resolves, against the family baseline. That form's BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old answer would have fixed the defect only where a model happened to be threaded through and left it live everywhere else. tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract and is corrected here rather than worked around. * fix(#3007): close every review finding on the Codex effort alignment Two isolated reviewers, correctness and security. Both found the same two blockers, and the per-model work was inert on every surface that matters until this commit. BLOCKER — resolve-execution never passed the model and discarded the clamp. cmdResolveExecution called the two-argument form and emitted only effort_rendered/effort_param/effort_propagation, so the per-model table was unreachable from production code (tests were its only caller) and requested/ clamped/reason were computed and thrown away. Requested outcome 3 names "the effective rendered effort in resolver output" specifically, so the feature was unmet on the exact surface the issue asks for. Now passes the resolved model and emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the existing key convention rather than introducing a nested object. BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed a nested {"effort": ...} sample; the real result is flat and those keys were absent entirely. A reference doc asserting a JSON path a reader can copy is worse than no doc. Corrected against the actual emitted key set. MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex kept minimal in its supported set and still clamped max down to xhigh, so the invocation-time and install-time channels disagreed about the same runtime's capability: --host codex with max emitted xhigh while the generated TOML said max. This is the repo's documented generative-fix-divergence class, so both tables now cross-reference each other and a parity test fails if they ever diverge again. MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null _baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback never fired and every effort rendered as null. And a non-array value made the Set constructor throw at module load — model-catalog.cjs is required across the whole CLI, so one bad JSON value killed every command, not just codex effort. Guarded on size and filtered to array values; both degrade to the hardcoded baseline. MAJOR — value widened to a nullable string with two consumers left behind. runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a null effort key in generated frontmatter); install-effort-resolver still declared a non-nullable return, a structural lie that silently defeated null checking. Both corrected, both omitting the key on null — the same posture as 'inherit', where omission means "follow the host default". MAJOR — the per-model table is inert today, and the docs now say so. All three shipped models advertise the same usable range and ultra (sol's only differentiator) is rejected for every model, so no observable output differs by model. The table stays because Codex declares capability per model and the sets are free to diverge — a single per-runtime assumption is precisely what went stale and produced this issue — but overselling it as a visible per-model feature would have been the same class of error as the doc blocker above. Tests: three passed under a full revert and are strengthened rather than deleted, since each guards a real contract (#3533's inherit rule, the undeclared-host rule, off-ladder handling) — they now also assert the clamp-visibility fields, which only exist after this change. The fast-check property is kept for its shrinking, and a deterministic nested loop over the full cross-product now sits beside it so coverage is exhaustive rather than sampled. Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg form and would have written a literal null reasoning effort on the ultra path; CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT. The installer defect was found by the co-change gate, not by a reviewer — install.js is a historical co-change partner of model-catalog.cts that this diff had not touched. * test(#3007): correct assertions that pinned Codex's stale effort premise Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the shipped commit. Every one is a stale pin, not a defect: each was probed against the built module before its expectation was changed, and none failed for a reason other than this premise correction. Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale by a production change must not ride inside another commit, because the release-sdk hotfix cherry-pick filter routes by subject prefix and a correction buried under the wrong prefix ships a half-state (v1.42.3, #3621). The most valuable one was tests/model-resolver.test.cjs's cross-provider validity invariant, which hardcoded the Codex enum as `minimal|low|medium|high|xhigh` and failed with "real API would 400". That message is now false in both directions: Codex accepts `max`, and rejects `minimal`, which no model advertises. The enum is corrected to `low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the "would the real API refuse this" check worth having, and it was right to fail here. It simply carried the stale fact in its own fixture. Test NAMES were corrected alongside their assertions wherever the name asserted the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal passthrough". A renamed test that still claims the old thing is worse than a failing one, and a green test whose name states a falsehood is how the next reader inherits the wrong premise. Both channels are covered: install-time (renderEffortForRuntime, and the generated .toml in install-runtime-artifacts) and invocation-time argv (effort-surface-axis). They were deliberately brought into agreement in this change, so their assertions had to move together. Each site carries a #3007 comment recording that Codex gained max/ultra and that capability is declared per model, so a future reader can tell this was a deliberate premise correction rather than a test bent to fit an implementation. * test(#3007): separate the effort-precedence case from the clamp case The previous stale-assertion pass over-corrected one test. It saw `effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`, assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to "max passes through" and changed the expectation to `max`. The remote runner disagreed. Reproduced against the real CLI: with that config and `gsd-planner`, the resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`, `effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE — gsd-planner is heavy/opus tier and its routing-tier default outranks `effort.default` — so `max` never reaches the renderer at all. The test says nothing about clamping and never did; it only looked like a clamp pin because both mechanisms happened to yield the same string. Restored to `xhigh` and renamed to say what it actually tests. It now also asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is what makes it impossible to mistake for a clamp pin again: those two fields prove the value is what the resolver produced rather than something the renderer downgraded. Before #3007 there was no way to tell the two apart from the output — which is precisely why the previous pass could not tell them apart either. Added the test that was actually missing: `effort.agent_overrides`, which outranks the tier default, so the requested level genuinely reaches the renderer and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified by probe before asserting. One test now pins the precedence rule and the other pins the #3007 behavior, and neither can be read as the other. That the clamp-visibility fields are what resolved this is a small argument for having added them. * chore(#3007): backfill changeset pr number to 3765 * test(#3007): put model-catalog under the mutation gate The Stryker shard showed as `skipping` on this PR despite the diff rewriting model-catalog's effort logic. That was legitimate, not a detection bug: `model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the whole module — including everything #3007 touches — sat entirely outside mutation scoring with has_work "false". Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no temp dirs. That shape is not stylistic — it is the #2790 precedent this file already documents. Stryker's command runner treats a whole `node --test ` invocation as ONE test costing whatever its slowest case costs, and re-runs it per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses runGsdTools throughout) would reproduce exactly the 15-minute shard-cap cancellation #2790 hit. The integration file is unaffected and keeps running in full in the normal test job. Coverage spans the module rather than only the diff, because the score is measured over the whole file: effort rendering across every model and ladder level in both channels, the prototype-chain host guard, the exported enums and maps, isAnthropicFlavoredModel's provider namespacings, the profile projections, nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are worth naming — every uncovered exported function is score given away, and mergeEffortTierDefaults turned out to have a genuinely interesting contract (#3531: a partial override merges over the built-ins rather than replacing them, and isValid gates the VALUE, not the tier name, so an unknown tier key is still merged in). Every expectation was probed against the built module before being asserted. minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment. Floors in this repo are measured, not chosen — the existing entries sit at 94, 75 and 56 — and they can only be measured in CI, because mutation shards run `node --test`, which is hard-blocked locally. The first CI run on this branch reports the real number and the floor gets ratcheted to it before merge. A placeholder of 1 reaching `next` would make the gate decorative: it would pass whether or not a single mutant is ever killed. Note the target is "never regress from measured", not a fixed 80 — planning-inspect sits at 56 and is documented as an accepted ratchet candidate. * test(#3007): bootstrap model-catalog's mutation floor legally The placeholder floor was structurally illegal and the remote run said so. tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module must carry a matching RATCHET_BASELINE entry in the same diff, minScore must EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three. That is the ratchet working exactly as intended — a floor nobody can satisfy accidentally is the point of it. Bootstrapped at 50 in both places. Fifty is not a measured score and the comment says so plainly: it is the minimum the guard permits, and it coincides with Stryker's own configured `break` threshold, so it is the lowest legal starting point for a module that has never been measured. It still must be ratcheted to floor(measured) - 1 before this PR merges. Also corrected a real defect in the file's own instructions. "HOW TO UPDATE" step 1 read "Run the per-module Stryker shard locally" — which cannot be done here, and which the same file contradicts eighty lines further down, where the #2790 scores are recorded as "not a local run; mutation shards run `node --test`, hard-blocked in this repo's local environment". stryker.config.mjs confirms the command runner invokes `node --test` once per mutant, and .claude/hooks/block-local-node-test.sh denies exactly that. So the documented first step sends the next contributor at a wall. Rewritten to describe the path that works — push, read the measured score off the CI shard, then set the floor and its baseline together in one diff — and to say why local measurement is not available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched. The two-step is inherent to the environment rather than a shortcut: a floor cannot be measured before the first CI run exists, and the guard rightly refuses to accept an unmeasured one below its minimum. * test(#3007): ratchet model-catalog's mutation floor to its measured score The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this file's own rule, minScore = floor(measured) - 1, which is the same arithmetic every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94. Both halves moved together, because the ratchet guard asserts minScore equals its RATCHET_BASELINE entry and would reject them drifting apart. The spawn-free unit surface is vindicated by the clock: 57 seconds, against a 15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing the shard at tests/model-resolver.test.cjs — #2790 recorded shards being CANCELLED at that cap when they targeted a runGsdTools-heavy integration file. The registry comment is rewritten rather than deleted. It previously warned that the floor was provisional and must not ship that way; leaving that text next to a measured floor would make the file lie in the other direction. It now records the measurement the way the sibling entries do, including that 59.62 sits below TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 — comfortably clear of its own floor with real room to grow. Raise it as the tests improve; never lower it. Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is an honest floor for a module that had NO mutation coverage at all an hour ago, and it is now pinned so it cannot silently regress. --------- Co-authored-by: sim --- .changeset/wise-otters-greet.md | 5 + CONTEXT.md | 2 +- bin/install.js | 14 +- docs/CONFIGURATION.md | 76 +++- ...48-unified-effort-and-fast-mode-routing.md | 63 +++ docs/how-to/configure-model-profiles.md | 45 ++ gsd-core/bin/shared/model-catalog.json | 8 +- scripts/mutation-matrix.cjs | 45 +- src/commands.cts | 14 +- src/install-effort-resolver.cts | 12 +- src/model-catalog.cts | 162 ++++++- src/runtime-artifact-conversion.cts | 9 +- tests/effort-surface-axis.test.cjs | 8 +- tests/install-runtime-artifacts.test.cjs | 15 +- tests/model-catalog.unit.test.cjs | 381 ++++++++++++++++ tests/model-resolver.test.cjs | 410 +++++++++++++++++- tests/mutation-matrix-ratchet.test.cjs | 1 + 17 files changed, 1219 insertions(+), 51 deletions(-) create mode 100644 .changeset/wise-otters-greet.md create mode 100644 tests/model-catalog.unit.test.cjs diff --git a/.changeset/wise-otters-greet.md b/.changeset/wise-otters-greet.md new file mode 100644 index 000000000..e17afc127 --- /dev/null +++ b/.changeset/wise-otters-greet.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 3765 +--- +**Codex reasoning effort is now resolved per model, and every clamp is visible** — `max` reaches Codex instead of being silently downgraded to `xhigh`, `minimal` clamps up to `low` instead of being sent to models that reject it, and `resolve-execution` reports the level you asked for alongside the one actually rendered. `ultra` is refused outright because it switches Codex into proactive task delegation underneath GSD's own orchestration. (#3007) diff --git a/CONTEXT.md b/CONTEXT.md index 808cf1812..a8de3bbd8 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -242,7 +242,7 @@ The seam deciding whether `.planning/` artifacts reach git. `commit_docs` resolv **Per-phase override (#3587, epic #2292 Phase 3).** A NEW tier resolves ABOVE the chain above, entirely inside `cmdCommit` (`src/commands.cts`'s `resolveCommitDocsPolicy`/`resolvePhaseCommitDocsOverride`) — deliberately NOT inside `loadConfigResolved`, which has no phase context and is called by nearly every command. The dynamic config key `phase_commit_docs.` (a `{ "": boolean }` map, registered in `config-schema.manifest.json`'s `dynamicKeyPatterns` and threaded through `config-loader.cts`'s `_baseConfig` projection the same way `agent_skills` is — a dynamic key absent from that hand-maintained allowlist is silently dropped on read, the exact failure mode `features.` demonstrates today) lets a tech lead commit one phase's artifacts while the project-wide `commit_docs` stays `false` (or the reverse). The phase being committed is resolved via the PRE-EXISTING `detectPhaseNumberFromFiles(files)` (the same #2539-hardened, project-code-aware derivation `branching_strategy` already used one branch below) and normalized through `normalizePhaseName` on both sides of the comparison, so `3`/`03`/`PROJ-03` hit one entry and a value scoped to a DIFFERENT phase never leaks. A non-boolean stored value (`"true"`, `1`, `null`) is never coerced — it falls through to the pre-existing chain untouched. When this tier suppresses a commit, the envelope's reason is `skipped_commit_docs_phase_false` — deliberately distinct from `skipped_commit_docs_false`, so a per-phase suppression is never reported as "your project setting is false" when it is actually `true`. **AC4 (byte-identical when unset):** with no `phase_commit_docs` key, this tier is a no-op and the three-tier chain above resolves exactly as before — pinned by the `folded:phase-commit-docs` block's C1-C5 in `tests/commit-docs-bypass.test.cjs`. The `phase_commit_docs.` grammar is a hand-copy of the canonical `PHASE_NUMBER_TOKEN_SOURCE` (`src/phase-id.cts`, #2128) into the hand-maintained schema manifest — pinned against drift by that same block's describe 'E' in `tests/commit-docs-bypass.test.cjs` (behavioral, over a shared shape list, per CLAUDE.md's Generative Fix Divergence class). **The two reasons are ordered, not peers, and `skipped_gitignored` is near-unreachable in a real project** (measured #3585): `cmdCommit` tests resolved `commit_docs` FIRST, and whenever `.planning/config.json` exists — which it does in every initialized project — the loader's gitignore auto-detect has already resolved that value to `false`, so the first branch returns `skipped_commit_docs_false` and the `isGitIgnored` branch below it is never reached. `skipped_gitignored` fires only when `config.json` is absent entirely, so the loader falls back to the `true` default and `cmdCommit`'s own check is what fires. A gitignore-driven skip therefore reports the config-driven reason; both members are behaviorally pinned by `tests/commit-docs-bypass.test.cjs` (B1-B3 and G1) so a rename fails loudly, but the reason a user sees does not distinguish *which* input suppressed the commit. **The gate is bypassable only from OUTSIDE the code**: a workflow step that types `git add` into its own shell reaches the index without passing through `cmdCommit`, and no code change can intercept that. Two guards therefore enforce it as text rather than at runtime: `tests/commit-files-pathspec.test.cjs` (#2269) requires every shipped `commit` invocation to declare `--files`, and `tests/commit-docs-bypass.test.cjs` (#1783, made repo-wide by #3585) requires every shipped `git add` able to reach `.planning/` to sit inside an **executable** `commit_docs` check — a markdown prose conditional ("**If `commit_docs` is true:**") is not a guard, because the bash block below it runs regardless. Both consume one shell tokenizer (`tests/helpers/shipped-command-scan.cjs`) and one exemption marker (`# gsd-scan-ignore: #NNN`, reason must cite a tracking ref per ADR-456). Guard state does not cross a fenced-block boundary: each fenced block is its own shell, so a guard opened in one block does not protect a `git add` in the next. Known limit: `.gitignore` has no effect on files git already TRACKS, so a project that committed `.planning/` before ignoring it keeps staging those paths — see `gsd-core/references/planning-config.md` and #3586. **`cmdCheckCommit`** (Command Module, verb `check-commit`) is a SEPARATE reader of the same `commit_docs` value, for callers outside GSD's own commit path: it inspects the staged set directly and refuses (non-zero exit) when `commit_docs` is `false` and any staged path is under `.planning/`; otherwise it allows. `gsd-tools commit-docs-guard enable`/`disable` (#3588) is the opt-in installer for a `.git/hooks/pre-commit` hook that shells out to exactly this verb, closing the one bypass the text-scan guards above cannot reach — a human or script running a bare `git add -A && git commit` in their OWN shell, outside any GSD-shipped workflow. The hook is identified by a `# gsd-core:commit-docs-guard` marker line (presence-checked, not byte-equality), is written only on explicit request (no install path wires it by default — locked by `tests/commands.test.cjs`'s E2 row), refuses rather than overwrite or delete a foreign `pre-commit`, resolves the real hooks dir via `git rev-parse --git-path hooks` so a linked worktree or submodule whose `.git` is a FILE works, and refuses outright when `core.hooksPath` is already set, because a written-but-ignored hook is worse than a refusal. ### Model Catalog Module -Leaf module owning the **static** model tables and the closed vocabularies derived from them — the tier/runtime/provider enums (`VALID_TIERS`, `VALID_AGENT_TIERS`, `KNOWN_RUNTIMES`, `KNOWN_PROVIDERS`, `RUNTIMES_WITH_REASONING_EFFORT`, `RUNTIMES_WITH_FAST_MODE`, `ADAPTIVE_TIER_VALUES`), the alias and profile maps (`MODEL_ALIAS_MAP`, `RUNTIME_PROFILE_MAP`, `PROVIDER_PRESETS`), effort rendering (`renderEffortForRuntime`), and the agent→model projections (`getAgentToModelMapForProfile`, `formatAgentToModelMapAsTable`). A **genuine leaf**: it imports `node:path` and its own `model-catalog.json` and nothing else, which is what makes it the correct home for anything several unrelated surfaces must agree on. Model *ids* live in `model-catalog.json`, never inline — changing one means regenerating goldens (`UPDATE_GOLDEN`). Also owns the **Anthropic-flavored-model rule** (#3241, ADR-2313): `CLAUDE_AGENT_ALIASES` (the frozen four-alias set `opus`/`sonnet`/`haiku`/`fable`) and `isAnthropicFlavoredModel(model)`, which is true for a bare tier alias or for any `claude-*` id in any provider namespacing (`anthropic/claude-*`, `us.anthropic.claude-*`); no OpenAI/Codex model id contains "claude", so the case-insensitive substring test is safe and exhaustive. It was **moved down here from the Model Resolver Module** rather than shared from there, because the Codex posture check (`agent-install-check`, epic #2313 Phase 2) and the Codex `.toml` sync (`commands`, Phase 3) both need the rule and neither may take the `config-loader` dependency `model-resolver` would have brought; `model-resolver` re-exports it for back-compat and a parity test fails if the two ever fork. _Avoid_: "the model list" (ambiguous between the catalog JSON and the derived enums). See Model Resolver Module, ADR-2313, and ADR-0003. +Leaf module owning the **static** model tables and the closed vocabularies derived from them — the tier/runtime/provider enums (`VALID_TIERS`, `VALID_AGENT_TIERS`, `KNOWN_RUNTIMES`, `KNOWN_PROVIDERS`, `RUNTIMES_WITH_REASONING_EFFORT`, `RUNTIMES_WITH_FAST_MODE`, `ADAPTIVE_TIER_VALUES`), the alias and profile maps (`MODEL_ALIAS_MAP`, `RUNTIME_PROFILE_MAP`, `PROVIDER_PRESETS`), effort rendering (`renderEffortForRuntime`), and the agent→model projections (`getAgentToModelMapForProfile`, `formatAgentToModelMapAsTable`). Also owns the per-model Codex effort capability table (`CODEX_MODEL_EFFORT`, sourced from `model-catalog.json`'s `codexModelEffort` with a `_baseline` fallback for unknown model ids) because Codex declares `supported_reasoning_levels` per model rather than per runtime, so `renderEffortForRuntime('codex', level, model)` resolves against that table. A **genuine leaf**: it imports `node:path` and its own `model-catalog.json` and nothing else, which is what makes it the correct home for anything several unrelated surfaces must agree on. Model *ids* live in `model-catalog.json`, never inline — changing one means regenerating goldens (`UPDATE_GOLDEN`). Also owns the **Anthropic-flavored-model rule** (#3241, ADR-2313): `CLAUDE_AGENT_ALIASES` (the frozen four-alias set `opus`/`sonnet`/`haiku`/`fable`) and `isAnthropicFlavoredModel(model)`, which is true for a bare tier alias or for any `claude-*` id in any provider namespacing (`anthropic/claude-*`, `us.anthropic.claude-*`); no OpenAI/Codex model id contains "claude", so the case-insensitive substring test is safe and exhaustive. It was **moved down here from the Model Resolver Module** rather than shared from there, because the Codex posture check (`agent-install-check`, epic #2313 Phase 2) and the Codex `.toml` sync (`commands`, Phase 3) both need the rule and neither may take the `config-loader` dependency `model-resolver` would have brought; `model-resolver` re-exports it for back-compat and a parity test fails if the two ever fork. _Avoid_: "the model list" (ambiguous between the catalog JSON and the derived enums). See Model Resolver Module, ADR-2313, and ADR-0003. ### Model Resolver Module Module owning model and effort resolution policy: resolves the model, runtime tier, planning granularity, reasoning effort, and fast-mode for a given agent by reading project config and resolving against the model profiles and catalog (`resolveModelInternal`, `resolveModelPolicy`, `resolveTierEntry`, `resolveModelForTier`, `resolveGranularityInternal`, `resolveEffortInternal`, `resolveFastModeInternal`, `resolveEffortForTier`, `nextEffort`, `assertValidGranularityOverride`). Depends only on leaf modules (`config-loader` for `loadConfig`, `configuration` for defaults, `model-profiles` and `model-catalog` for the static tables) — no other core dependency. Extracted from the Core module per ADR-857 rollout phase 2f (#888) — the final core.cts decomposition step; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. **`CLAUDE_AGENT_ALIASES` no longer lives here** — it moved down to the Model Catalog Module (#3241, ADR-2313 Phase 1) so the Agent Install Check and Codex-sync surfaces can consume the alias rule without taking a `config-loader` dependency this module would have dragged with it; it is still **re-exported** from here, so existing importers (`bin/install.js`, `tests/codex-config.test.cjs`) are unaffected and a parity test asserts both modules expose the same set. Source of truth: `gsd-core/bin/lib/model-resolver.cjs` (generated from `src/model-resolver.cts`). diff --git a/bin/install.js b/bin/install.js index f96fe9829..1c2b027b5 100755 --- a/bin/install.js +++ b/bin/install.js @@ -4012,8 +4012,10 @@ function generateCodexAgentToml(agentName, agentContent, modelOverrides = null, // #443 — Unified effort for Codex .toml. Uses the same config-driven precedence chain // as the Claude .md effort injection (resolveInstallTimeEffort), so both runtimes read // from the same effort.agent_overrides / effort.routing_tier_defaults / effort.default - // config source. Codex does not support 'max' → clamped to 'xhigh' by - // gsdRenderEffortForRuntime('codex', ...). + // config source. #3007 — Codex advertises supported_reasoning_levels per model, so the + // pinned model id is passed through and the value is resolved against that model's own + // set: 'max' now passes, 'minimal' clamps up to 'low', and 'ultra' is refused (no key + // emitted) rather than clamped to a fabricated level. // #838 — Do not pin effort when Codex is intentionally inheriting the parent // chat model. A TOML with no `model` but a static `model_reasoning_effort` // creates confusing partial routing: model follows the Codex UI while effort @@ -4023,8 +4025,12 @@ function generateCodexAgentToml(agentName, agentContent, modelOverrides = null, // #3533 (10d): 'inherit' means OMIT the pin — the agent follows the host's // own effort default. Never write the literal. if (_universalEffortCodex !== 'inherit') { - const _renderedEffortCodex = _getGsdEffortCatalog().renderEffortForRuntime('codex', _universalEffortCodex).value; - lines.push(`model_reasoning_effort = ${JSON.stringify(_renderedEffortCodex)}`); + const _renderedEffortCodex = _getGsdEffortCatalog().renderEffortForRuntime('codex', _universalEffortCodex, pinnedModel).value; + // #3007 — 'ultra' is rejected by the model's supported_reasoning_levels and + // renders as null. Omit the key entirely rather than write a literal `null`. + if (_renderedEffortCodex !== null) { + lines.push(`model_reasoning_effort = ${JSON.stringify(_renderedEffortCodex)}`); + } } } diff --git a/docs/CONFIGURATION.md b/docs/CONFIGURATION.md index 146cfce9e..9127e425a 100644 --- a/docs/CONFIGURATION.md +++ b/docs/CONFIGURATION.md @@ -1521,7 +1521,60 @@ minimal < low < medium < high < xhigh < max Effort is rendered per-runtime: `output_config.effort` for Claude (Claude Code subagent `effort` frontmatter / `CLAUDE_CODE_EFFORT_LEVEL` env), `model_reasoning_effort` for Codex (Responses API `reasoning.effort`). -**Cross-provider clamping:** `max` is Anthropic-only — it clamps to `xhigh` on Codex. `minimal` is Codex-only — it clamps to `low` on Claude. +**Cross-provider clamping:** `minimal` is Anthropic-unsupported — it clamps to `low` on Claude. + +**Codex effort is resolved per model, not per runtime (#3007).** Codex advertises a +`supported_reasoning_levels` set on each model and validates against it, so the same universal level +can pass cleanly on one model and clamp on another. GSD therefore renders against the model's own +advertised set: + +| Model | Advertised levels | +|---|---| +| `gpt-5.6-sol` | `low`, `medium`, `high`, `xhigh`, `max`, `ultra` | +| `gpt-5.6-terra` | `low`, `medium`, `high`, `xhigh`, `max` | +| `gpt-5.6-luna` | `low`, `medium`, `high`, `xhigh`, `max` | +| any other / unknown id | `low`, `medium`, `high`, `xhigh`, `max` (family baseline) | + +Today every shipped Codex model advertises the same usable range, so the same effort resolves +identically across `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` — `ultra` is sol's only +differentiator, and GSD rejects it for every model regardless (see below), so no observable output +currently differs by model. The table is per-model, not per-runtime, because Codex declares +capability per model and the sets are free to diverge — the previous single per-runtime assumption +is exactly what went stale and produced this change. + +Three consequences: + +- **`max` reaches Codex.** It is no longer clamped to `xhigh`. Earlier GSD releases described `max` + as Anthropic-only; that was accurate when written and Codex has since added it. If you set `max` + for a Codex agent, your generated `model_reasoning_effort` now says `max` where it previously said + `xhigh`. +- **`minimal` no longer reaches Codex.** No Codex model advertises it, so it clamps up to `low` — + the floor every model does advertise. GSD previously emitted `minimal` verbatim, which Codex + rejects. +- **`ultra` is refused outright**, and is not part of GSD's ladder. See below. + +**Every clamp is now visible.** `resolve-execution` reports the level you asked for alongside the +level actually rendered, so a downgrade is legible instead of silent. These are flat keys in the +same result object as `effort_rendered` — there is no nested `effort` object: + +```json +{ + "effort_rendered": "low", + "effort_requested": "minimal", + "effort_clamped": true, + "effort_clamp_reason": "requested 'minimal' is not in gpt-5.6-luna's advertised reasoning levels; clamped up to its floor, 'low'." +} +``` + +**Why `ultra` is rejected rather than clamped.** Codex's own catalog describes `ultra` as *"Maximum +reasoning with automatic task delegation"* — it is a mode switch, not a louder `max`. At `ultra` +Codex enters proactive multi-agent mode and spawns sub-agents on its own initiative, which would run +underneath GSD's orchestration rather than inside it ([#2167](https://github.com/open-gsd/gsd-core/issues/2167)). +GSD refuses it for every model, including `gpt-5.6-sol`, which does advertise it. This is +deliberately stricter than Codex requires: Codex only applies proactive mode to V2 sessions and +never to spawned sub-agents, but GSD writes effort into generated agent files at install time and +cannot know the session source of a future invocation. Clamping `ultra` down to `max` was rejected +as an option — it would silently discard what you actually asked for. The model-catalog's `reasoning_effort` per-tier hint is a legacy field kept for reference; effort is now config-driven. @@ -1663,18 +1716,23 @@ Use `node gsd-tools.cjs resolve-execution [--effort ] [--fas ```json { - "model": "opus", - "profile": "balanced", - "effort": "xhigh", - "effort_rendered": "xhigh", - "effort_param": "output_config.effort", - "effort_propagation": "frontmatter", - "fast_mode": false, + "model": "opus", + "profile": "balanced", + "effort": "xhigh", + "effort_rendered": "xhigh", + "effort_param": "output_config.effort", + "effort_propagation": "frontmatter", + "effort_requested": "xhigh", + "effort_clamped": false, + "effort_clamp_reason": null, + "fast_mode": false, "fast_mode_supported": false } ``` -`effort_param` tells you which runtime parameter to set. `fast_mode_supported` tells you whether the configured runtime supports per-agent fast_mode propagation. +`effort_param` tells you which runtime parameter to set. `effort_requested` is the level you asked +for (before any clamp); `effort_rendered` is what actually shipped. `effort_clamped` is `true` only +when the two differ, and `effort_clamp_reason` explains why (`null` when unclamped). `fast_mode_supported` tells you whether the configured runtime supports per-agent fast_mode propagation. --- diff --git a/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md b/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md index 3263e878f..a277e5892 100644 --- a/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md +++ b/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md @@ -59,6 +59,69 @@ Recorded as a dated section rather than by editing the amendment above or Decisi **Boundary — unchanged.** [ADR-2313](2313-codex-passive-model-posture.md) still owns the static/install-time channel; [ADR-1239](1239-gsd-embeddable-orchestration-engine.md)'s `effortSurface` amendment still owns the invocation-time argv channel. This amendment changes neither, and changes no runtime behavior at all. +## Amendment (#3007, 2026-08-22) — the Codex capability premise went stale + +**What changed underneath this ADR.** The Context section below records, as fact, that Codex's +ladder is `minimal, low, medium, high, xhigh` and that it *"has `minimal`; no `max`"*, and Decision +item 2 clamps `max → xhigh` on that basis. That was accurate when written. It is no longer: +Codex's `ReasoningEffort` now accepts `none, minimal, low, medium, high, xhigh, max, ultra`, and +capability is declared **per model** via `supported_reasoning_levels`, which Codex validates against +(`validate_spawn_agent_reasoning_effort`) and exposes through `model/list`. + +Verified against Codex's own `codex-rs/models-manager/models.json`: + +| Model | `supported_reasoning_levels` | `default_reasoning_level` | +|---|---|---| +| `gpt-5.6-sol` | low, medium, high, xhigh, max, **ultra** | `low` | +| `gpt-5.6-luna` | low, medium, high, xhigh, max | `medium` | + +**No Codex model advertises `minimal`** — the level this ADR called Codex-only. + +**What this amendment changes.** Decision item 2's per-runtime clamp table is superseded for Codex +by a per-**model** advertised set, with the family baseline `low, medium, high, xhigh, max` for any +model GSD does not know: + +| Universal level | Claude rendering | Codex rendering (was → now) | +|---|---|---| +| `minimal` | `low` (clamped) | `minimal` → **`low` (clamped)** | +| `low`–`xhigh` | unchanged | unchanged | +| `max` | `max` | `xhigh` (clamped) → **`max` (passes)** | +| `ultra` | *not on the ladder* | **rejected, never clamped** | + +Two defects this corrects, both live on `next` before it: + +1. `max` was silently discarded for every Codex model, all of which advertise it. +2. `providerPresets.openai.haiku.low` paired `gpt-5.6-luna` with `reasoning_effort: "minimal"` — a + level luna does not advertise, written into a document Codex itself validates. + +**Clamping becomes visible rather than silent.** `RenderedEffort` gains `requested` / `clamped` / +`reason`, and `resolve-execution` surfaces them as the flat result keys `effort_requested`, +`effort_clamped`, and `effort_clamp_reason` — siblings of the existing `effort_rendered`, not a +nested `effort` object. The old table clamped correctly-but-invisibly, so a user asking for `max` on +Codex had no way to learn they were getting `xhigh` — the failure mode Postel's robustness critique +warns about, and the reason "be liberal" here has to mean "liberal and loud". + +**No model divergence is observable today.** All three shipped Codex models advertise the same +usable set (`low`…`max`), and `ultra` — sol's only differentiator — is rejected for every model +regardless. So today, the same requested level renders identically across `gpt-5.6-sol`, +`gpt-5.6-terra`, and `gpt-5.6-luna`; the per-model table exists because Codex declares capability +per model and the sets are free to diverge, not because a user can currently observe a difference. + +**`ultra` is refused, not laddered.** Codex's catalog describes it as *"Maximum reasoning with +automatic task delegation"*, and at `ultra` Codex enters proactive multi-agent mode +(`effective_multi_agent_mode` → `Proactive`), spawning sub-agents on its own initiative underneath +GSD's orchestration ([#2167](https://github.com/open-gsd/gsd-core/issues/2167)). It is a mode +switch, not a reasoning depth, so it is not added to the universal ladder — which stays +provider-agnostic per Decision item 2 — and it is rejected for every model including `gpt-5.6-sol`, +which advertises it. GSD is deliberately stricter than Codex here: Codex applies proactive mode only +to V2 sessions and never to spawned sub-agents, but GSD writes effort at install time and cannot +know the session source of a future invocation. + +**What this amendment does NOT change.** The universal ladder itself, the cascade, the resolver +precedence, the two channel boundaries above, and Claude's rendering are all untouched. The +per-model table is static and will go stale exactly as this premise did; runtime discovery via +`model/list` is the known escape hatch and was deliberately deferred as the larger step. + ## Context ### Effort control and fast mode in Claude Opus 4.8 diff --git a/docs/how-to/configure-model-profiles.md b/docs/how-to/configure-model-profiles.md index 7c4c0c380..cef73192f 100644 --- a/docs/how-to/configure-model-profiles.md +++ b/docs/how-to/configure-model-profiles.md @@ -228,6 +228,51 @@ Codex UI drives both rather than one following GSD and the other following your > The installer prints a one-time notice when it drops a pin. If you were on a ChatGPT account, this > is the change that stops the 400s — nothing to do. +### Allocating for execution-heavy workflows on Codex + +Execution and verification account for most of the model calls in a long GSD run — planning happens +once per phase, execution happens per plan, and verification runs over everything produced. On +2026-07-30 OpenAI cut GPT-5.6 Luna API pricing by 80% and Terra by 20%, and reduced how many credits +both consume against Codex paid-plan quotas while leaving subscription prices and quota budgets +unchanged. Sol was unchanged. That makes the cheaper models materially cheaper for exactly the +high-volume half of a workflow. + +GSD does not add a routing surface for this — the levers below already express it, and +[#2935](https://github.com/open-gsd/gsd-core/issues/2935) was closed as already-implemented on +precisely that basis. Keep Sol where the reasoning is worth the spend, and put the volume on Terra +or Luna: + +```json +{ + "runtime": "codex", + "model_overrides": { + "gsd-planner": "gpt-5.6-sol", + "gsd-debugger": "gpt-5.6-sol", + "gsd-executor": "gpt-5.6-terra", + "gsd-verifier": "gpt-5.6-luna" + } +} +``` + +Prefer `models` when you want the split by *phase type* rather than by agent — it maps the six phase +types at once and every agent carries a `phaseType`, so it survives the roster changing under you: + +```json +{ + "models": { "planning": "opus", "execution": "sonnet", "verification": "haiku" } +} +``` + +Two things worth knowing before you tune this: + +- **Effort is a separate lever from model, and it is now per-model.** Dropping to Luna does not force + you to drop effort — Luna advertises everything up to `max`. See + [Configuration reference — effort](../CONFIGURATION.md#model-profiles) for the per-model table and + which levels clamp. +- **These are cost/limit tradeoffs, not quality claims.** The 2026-07-30 change was a pricing and + credit-accounting change; it did not alter model quality. Sol remains the strongest model for + planning and hard debugging, which is why it stays there above. + **If you want per-agent model IDs on any non-Claude runtime:** ```json diff --git a/gsd-core/bin/shared/model-catalog.json b/gsd-core/bin/shared/model-catalog.json index 6b987062c..1ecff022b 100644 --- a/gsd-core/bin/shared/model-catalog.json +++ b/gsd-core/bin/shared/model-catalog.json @@ -112,7 +112,7 @@ "openai": { "opus": { "low": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" }, "medium": { "model": "gpt-5.6-sol", "reasoning_effort": "high" }, "high": { "model": "gpt-5.6-sol", "reasoning_effort": "xhigh" } }, "sonnet": { "low": { "model": "gpt-5.6-luna", "reasoning_effort": "low" }, "medium": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" }, "high": { "model": "gpt-5.6-sol", "reasoning_effort": "medium" } }, - "haiku": { "low": { "model": "gpt-5.6-luna", "reasoning_effort": "minimal" }, "medium": { "model": "gpt-5.6-luna", "reasoning_effort": "medium" }, "high": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" } } + "haiku": { "low": { "model": "gpt-5.6-luna", "reasoning_effort": "low" }, "medium": { "model": "gpt-5.6-luna", "reasoning_effort": "medium" }, "high": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" } } }, "google": { "opus": { "low": { "model": "gemini-2.5-flash-lite" }, "medium": { "model": "gemini-3-flash" }, "high": { "model": "gemini-3.1-pro-preview" } }, @@ -130,6 +130,12 @@ "haiku": { "low": null, "medium": null, "high": null } } }, + "codexModelEffort": { + "_baseline": ["low", "medium", "high", "xhigh", "max"], + "gpt-5.6-sol": ["low", "medium", "high", "xhigh", "max", "ultra"], + "gpt-5.6-terra": ["low", "medium", "high", "xhigh", "max"], + "gpt-5.6-luna": ["low", "medium", "high", "xhigh", "max"] + }, "agents": { "gsd-planner": { "golden": "opus", "balanced": "opus", "budget": "sonnet", "phaseType": "planning", "routingTier": "heavy" }, "gsd-roadmapper": { "golden": "opus", "balanced": "sonnet", "budget": "sonnet", "phaseType": "planning", "routingTier": "heavy" }, diff --git a/scripts/mutation-matrix.cjs b/scripts/mutation-matrix.cjs index 86be866fb..acd0cda1b 100644 --- a/scripts/mutation-matrix.cjs +++ b/scripts/mutation-matrix.cjs @@ -92,10 +92,16 @@ function readStdinSync() { // confirmed equivalent mutant is acceptable. // // HOW TO UPDATE: -// 1. Run the per-module Stryker shard locally. -// 2. Note the reported score. -// 3. Set minScore = floor(score) - 1 (never lower than current value). -// 4. Open a PR — the CI gate will enforce the new floor on every future run. +// 1. The per-module Stryker shard CANNOT be run locally: Stryker's command +// runner invokes `node --test` once per mutant (see stryker.config.mjs), +// and this repo hard-blocks local `node --test` via +// .claude/hooks/block-local-node-test.sh. Push the branch instead and +// let CI run the shard for the changed module. +// 2. Read the measured score from the CI shard's output. +// 3. Set minScore = floor(measured) - 1 (never lower than current value) +// and update the matching RATCHET_BASELINE entry in the same diff. +// 4. Open/update the PR — the CI gate will enforce the new floor on every +// future run. /** Long-run target for all modules (ADR-456). */ const TARGET_MUTATION_SCORE = 80; @@ -259,6 +265,37 @@ const COVERED = { ], minScore: 94, }, + // model-catalog: net-new registration by #3007. The module was entirely + // outside mutation scoring (has_work: "false") before this entry, so the + // #3007 per-model Codex effort rewrite (renderEffortForRuntime's + // CODEX_MODEL_EFFORT lookup, the 'ultra' policy rejection, the ladder + // walk-up clamp) had zero mutation coverage. + // + // Same #2790 precedent as planning-inspect above: this shard points at a + // dedicated tests/model-catalog.unit.test.cjs, NOT tests/model-resolver.test.cjs + // — that integration file uses runGsdTools heavily and would hit the same + // 15-minute shard-cap cancellation #2790 documented (a `node --test ` + // invocation is ONE test costing whatever its slowest case costs, re-run + // per mutant). tests/model-catalog.unit.test.cjs is spawn-free, in-process, + // and runs in well under a second. + // + // Measured CI score (GitHub Actions run 32605073352, job 97108869486): + // model-catalog 59.62% → floor 58 (248 killed, 168 survived, 0 timeouts, + // 0 errors; below TARGET_MUTATION_SCORE (80) — ratchet candidate like + // planning-inspect (56): comfortably clears its own floor but has real + // room to grow. Raise as its tests improve, never lower it.) + // Floor follows this file's documented rule, minScore = floor(measured) - 1, + // matching the sibling precedent exactly (57.03 → 56, 76.58 → 75, 95.65 → 94). + // + // The shard completed in 57 seconds — concrete evidence the spawn-free + // unit-file design above worked: the #2790 precedent's 15-minute shard-cap + // cancellations do not apply here, and for comparison the `frontmatter` + // shard in the same run took 9m46s. + 'model-catalog': { + cjs: 'gsd-core/bin/lib/model-catalog.cjs', + tests: ['tests/model-catalog.unit.test.cjs'], + minScore: 58, + }, }; // ── Files that, when changed, invalidate ALL modules ───────────────────────── diff --git a/src/commands.cts b/src/commands.cts index 477f2f3a6..b2055ca7a 100644 --- a/src/commands.cts +++ b/src/commands.cts @@ -619,7 +619,12 @@ function cmdResolveExecution(cwd: string, agentType: string | undefined, raw: bo const fastMode = resolveFastModeInternal(cwd, agentType!, fastModeOpts); const runtime = (config['runtime'] as string) || 'claude'; - const rendered = renderEffortForRuntime(runtime, effort); + // #3007: pass the resolved model so the per-model advertised-effort ceiling + // (CODEX_MODEL_EFFORT) is reachable from this production seam. `model` may + // be a tier alias or a non-Codex id for other runtimes — that's fine and + // must not be special-cased here: advertisedCodexEffort() falls back to the + // family baseline for any id it doesn't recognize. + const rendered = renderEffortForRuntime(runtime, effort, model); const fastModeSupported = RUNTIMES_WITH_FAST_MODE.has(runtime); @@ -675,6 +680,9 @@ function cmdResolveExecution(cwd: string, agentType: string | undefined, raw: bo effort_rendered: rendered.value, effort_param: rendered.param, effort_propagation: rendered.channel, + effort_requested: rendered.requested, + effort_clamped: rendered.clamped, + effort_clamp_reason: rendered.reason, effort_effective: effortEffective, effort_effective_source: effortEffectiveSource, fast_mode: fastMode, @@ -855,8 +863,10 @@ function cmdEffortSync(cwd: string, raw: boolean, opts?: { dryRun?: boolean; con continue; } + // `runtime` is guaranteed 'claude' by the guard above (#3007: only + // codex's 'ultra' rejection can produce a null value). const rendered = renderEffortForRuntime(runtime, universalEffort); - const newEffortValue = rendered.value; + const newEffortValue = rendered.value as string; // eslint-disable-next-line local/no-unbounded-quantifier -- lazy `*?` bounded by the `^---$/m` closing anchor, no nested quantifier, measured linear to 5MB (no-closing-marker adversarial input) const fmMatch = /^---\r?\n([\s\S]*?)^---\r?$/m.exec(content); diff --git a/src/install-effort-resolver.cts b/src/install-effort-resolver.cts index 736362fa4..1e0cb86e7 100644 --- a/src/install-effort-resolver.cts +++ b/src/install-effort-resolver.cts @@ -76,7 +76,11 @@ function _readGsdConfigFile(absPath: string, label: string): Record; - renderEffortForRuntime: (runtime: string, effort: string) => { value: string }; + // #3007: `value` is `string | null` — declaring it `string` here was a + // structural lie that silently defeated TS null-checking for anything + // routed through this seam (a rejected/unrenderable effort level renders + // null, e.g. 'ultra' or an exhausted catalog clamp). + renderEffortForRuntime: (runtime: string, effort: string) => { value: string | null }; EFFORT_MANIFEST_TIER_DEFAULTS: Record; EFFORT_MANIFEST_DEFAULT: string; } @@ -94,7 +98,11 @@ function _getGsdEffortCatalog(): EffortCatalog { // eslint-disable-next-line @typescript-eslint/no-require-imports -- model-catalog.cjs is an export= CommonJS module const { AGENT_DEFAULT_TIERS, renderEffortForRuntime } = require('./model-catalog.cjs') as { AGENT_DEFAULT_TIERS: Record; - renderEffortForRuntime: (runtime: string, effort: string) => { value: string }; + // #3007: `value` is `string | null` — declaring it `string` here was a + // structural lie that silently defeated TS null-checking for anything + // routed through this seam (a rejected/unrenderable effort level renders + // null, e.g. 'ultra' or an exhausted catalog clamp). + renderEffortForRuntime: (runtime: string, effort: string) => { value: string | null }; }; // This module lives in gsd-core/bin/lib/, so the shared manifest is one level diff --git a/src/model-catalog.cts b/src/model-catalog.cts index eb8eb2151..a6c598002 100644 --- a/src/model-catalog.cts +++ b/src/model-catalog.cts @@ -55,6 +55,7 @@ export interface ModelCatalog { runtimeTierDefaults: Record>; providerPresets: Record>>; agents: Record; + codexModelEffort?: Record; } let catalog: ModelCatalog | null = null; @@ -144,6 +145,43 @@ export const RUNTIMES_WITH_REASONING_EFFORT: Set = new Set( export const PROVIDER_PRESETS: Record>> = _catalog.providerPresets ?? {}; +// ─── #3007 — Codex per-model effort capability ─────────────────────────────── +// +// Codex's own `models.json` publishes `supported_reasoning_levels` per model and +// rejects an unsupported level at request time, so GSD must be conservative +// about what it sends: read the ceiling as DATA from the catalog (never +// branch on model id in code) and fall back to the family baseline for any +// model the catalog doesn't know about. +// (b) A malformed catalog entry (e.g. a non-array value like `"gpt-x": 5`) must +// degrade to "ignore that entry", never throw — model-catalog.cjs is required +// across the whole CLI, so one bad JSON value must not kill every command. +// `new Set(5)` would throw at module load; filter to array values first. +export const CODEX_MODEL_EFFORT: Record> = Object.fromEntries( + Object.entries(_catalog.codexModelEffort ?? {}) + .filter(([, levels]) => Array.isArray(levels)) + .map(([model, levels]) => [model, new Set(levels)]) +); +// (a) `??` only catches null/undefined. A malformed `"_baseline": null` still +// produces `new Set(null)` above — an empty Set, which is truthy — so a bare +// `??` fallback would never fire and every level would silently lose its +// advertised set (every effort would render as `value: null`). Guard on +// `.size > 0` so an empty/missing/malformed baseline always falls back to the +// hardcoded floor instead of failing open. +const CODEX_EFFORT_BASELINE: Set = + CODEX_MODEL_EFFORT['_baseline'] && CODEX_MODEL_EFFORT['_baseline'].size > 0 + ? CODEX_MODEL_EFFORT['_baseline'] + : new Set(['low', 'medium', 'high', 'xhigh', 'max']); + +function advertisedCodexEffort(model: string | null | undefined): Set { + if (typeof model !== 'string' || model.length === 0) return CODEX_EFFORT_BASELINE; + return Object.prototype.hasOwnProperty.call(CODEX_MODEL_EFFORT, model) ? CODEX_MODEL_EFFORT[model] : CODEX_EFFORT_BASELINE; +} + +// The full universal effort ladder, low-to-high. Used only to find "the +// nearest advertised level below" when a requested level isn't supported — +// never to invent behaviour for a level that isn't on it at all. +const EFFORT_LADDER: string[] = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + // KNOWN_PROVIDERS excludes 'generic' — it is a sentinel (all null entries) that // forces users to supply model IDs via model_profile_overrides. It is not a // real catalog-backed provider (#49). @@ -229,20 +267,35 @@ export const EFFORT_RENDERING: Record = { }, }, codex: { + // #3007: 'max' and 'minimal' are stale here — Codex's per-model table + // (CODEX_MODEL_EFFORT above) is now the source of truth for what a given + // model actually advertises, and every model in the family baseline DOES + // advertise 'max' (no model advertises 'minimal'). This runtime-level + // spec is kept in sync with the family baseline so the two tables can + // never disagree; renderEffortForRuntime layers the per-model ceiling + // (and the 'ultra' policy rejection) on top of it. + // KEEP IN SYNC with EFFORT_ARGV.codex below — same family baseline, two + // channels (install-time vs invocation-time); they must never diverge. param: 'model_reasoning_effort', channel: 'api', - supported: new Set(['minimal', 'low', 'medium', 'high', 'xhigh']), + supported: new Set(['low', 'medium', 'high', 'xhigh', 'max']), clamp(level: string): string { - if (level === 'max') return 'xhigh'; + if (level === 'minimal') return 'low'; return level; }, }, }; export interface RenderedEffort { - value: string; + value: string | null; param: string | null; channel: string | null; + /** The level as originally asked for. */ + requested?: string; + /** true ONLY when value !== requested. */ + clamped?: boolean; + /** Why, when clamped or rejected; null otherwise. */ + reason?: string | null; } // ─── Invocation-time (argv) effort rendering ───────────────────────────────── @@ -281,10 +334,14 @@ export const EFFORT_ARGV: Record = { }, // First-party Codex docs: `model_reasoning_effort` is a config-only key with no // dedicated flag, so the generic `-c key=value` override is the only argv route. + // #3007: KEEP IN SYNC with EFFORT_RENDERING.codex above — this table must match + // the family baseline exactly (no 'minimal', 'max' passes through unclamped), + // otherwise the argv channel and the install-time channel disagree about the + // same runtime's capability. codex: { render: (level: string): string[] => ['-c', `model_reasoning_effort=${level}`], - supported: new Set(['minimal', 'low', 'medium', 'high', 'xhigh']), - clamp: (level: string): string => (level === 'max' ? 'xhigh' : level), + supported: new Set(['low', 'medium', 'high', 'xhigh', 'max']), + clamp: (level: string): string => (level === 'minimal' ? 'low' : level), }, }; @@ -323,23 +380,108 @@ export function renderEffortArgv( /** * Render a universal effort string for a specific runtime. + * + * `model` (#3007) is consulted ONLY for codex: Codex's per-model + * `supported_reasoning_levels` means the same universal level can be a clean + * pass-through on one model and a clamp (or, for 'ultra', an outright + * rejection) on another. Every other runtime ignores the third argument + * entirely — passing a model id to claude changes nothing. */ -export function renderEffortForRuntime(runtime: string, universalEffort: string): RenderedEffort { +export function renderEffortForRuntime(runtime: string, universalEffort: string, model?: string | null): RenderedEffort { // #3533 (10d): 'inherit' is not a wire level on ANY runtime — it means // "omit the key / pass no argument and follow the session/host default". // Renderers must never emit it as a literal; null param/channel tells - // resolve-execution consumers there is no propagation. + // resolve-execution consumers there is no propagation. Never measured + // against any supported set. if (universalEffort === 'inherit') { - return { value: 'inherit', param: null, channel: null }; + return { value: 'inherit', param: null, channel: null, requested: 'inherit', clamped: false, reason: null }; } const spec = EFFORT_RENDERING[runtime]; if (!spec) { - return { value: universalEffort, param: null, channel: null }; + return { value: universalEffort, param: null, channel: null, requested: universalEffort, clamped: false, reason: null }; } + + if (runtime === 'codex') { + // #2167 — 'ultra' turns on Codex's automatic task delegation, which would + // let Codex spawn agents underneath GSD's own orchestration. GSD rejects + // it unconditionally as a POLICY call, never as a capability clamp — this + // holds even for a model (e.g. gpt-5.6-sol) that DOES advertise 'ultra', + // so it is never softened down to 'max'. + if (universalEffort === 'ultra') { + return { + value: null, + param: null, + channel: null, + requested: 'ultra', + clamped: false, + reason: "'ultra' turns on Codex's automatic task delegation, which would let Codex spawn agents underneath GSD's own orchestration (#2167); GSD rejects it regardless of what the model advertises.", + }; + } + + const allowed = advertisedCodexEffort(model); + if (allowed.has(universalEffort)) { + return { value: universalEffort, param: spec.param, channel: spec.channel, requested: universalEffort, clamped: false, reason: null }; + } + + const idx = EFFORT_LADDER.indexOf(universalEffort); + if (idx === -1) { + // Not on the ladder at all (e.g. 'MAX') — preserve prior behaviour: + // fall through to the runtime-level clamp rather than inventing new + // handling for input the ladder doesn't recognize. + return { value: spec.clamp(universalEffort), param: spec.param, channel: spec.channel, requested: universalEffort, clamped: false, reason: null }; + } + // Every model's advertised set is a contiguous run up to 'max' (or 'ultra' + // for sol, already handled above), so the only unsupported level in + // practice is 'minimal' — below every model's floor. Walk UP the ladder + // to the nearest level the model actually advertises (its floor): there + // is nothing below 'minimal' to fall back to. + // Walking UP is safe today only because every advertised set floors at + // 'low' — a future model whose floor is, say, 'high' would silently turn + // a requested 'low' into 'high': MORE reasoning and MORE cost than asked + // for, with no error. `clamped`/`reason` below is what makes that + // escalation visible to a caller instead of a silent cost surprise, which + // is why those fields are not optional decoration. + for (let i = idx + 1; i < EFFORT_LADDER.length; i++) { + const candidate = EFFORT_LADDER[i]; + // 'ultra' is never a valid clamp target: it would re-enter, by the back + // door, the delegation mode the #2167 rejection above exists to keep + // out. A clamp may never produce a value that a direct request for + // that same value would have refused. + if (candidate === 'ultra') { + continue; + } + if (allowed.has(candidate)) { + return { + value: candidate, + param: spec.param, + channel: spec.channel, + requested: universalEffort, + clamped: true, + reason: `requested '${universalEffort}' is not in ${model ? `${model}'s` : "the codex family baseline's"} advertised reasoning levels; clamped up to its floor, '${candidate}'.`, + }; + } + } + // No advertised level at or above the request either (shouldn't happen + // given today's catalog data, but never throw): reject rather than emit + // an unsupported level. + return { + value: null, + param: null, + channel: null, + requested: universalEffort, + clamped: false, + reason: `requested '${universalEffort}' is not in ${model ? `${model}'s` : "the codex family baseline's"} advertised reasoning levels, and no advertised level is available either.`, + }; + } + + const value = spec.clamp(universalEffort); return { - value: spec.clamp(universalEffort), + value, param: spec.param, channel: spec.channel, + requested: universalEffort, + clamped: value !== universalEffort, + reason: value !== universalEffort ? `requested '${universalEffort}' clamped to '${value}' for ${runtime}.` : null, }; } diff --git a/src/runtime-artifact-conversion.cts b/src/runtime-artifact-conversion.cts index 3161b65b1..f3e412583 100644 --- a/src/runtime-artifact-conversion.cts +++ b/src/runtime-artifact-conversion.cts @@ -3503,7 +3503,14 @@ function applyAgentFrontmatterExtensions( // effort key, so skipping injection is the whole job. if (universalEffort !== 'inherit') { const renderedEffort = _getGsdEffortCatalog().renderEffortForRuntime(runtime, universalEffort).value; - result = injectEffortFrontmatter(result, renderedEffort); + // #3007: `value` is `string | null` — a rejected/unrenderable level (e.g. + // 'ultra', or a catalog with no advertised level at or above the request) + // renders null. Same posture as the 'inherit' case above: omit the key + // entirely rather than writing a literal `effort: null`, so the host + // falls back to its own default instead of failing to parse. + if (renderedEffort !== null) { + result = injectEffortFrontmatter(result, renderedEffort); + } } const disallowedTools = READONLY_AGENT_DISALLOWED_TOOLS[agentName]; if (disallowedTools) result = injectDisallowedToolsFrontmatter(result, disallowedTools); diff --git a/tests/effort-surface-axis.test.cjs b/tests/effort-surface-axis.test.cjs index bfa51b087..ce19f5cba 100644 --- a/tests/effort-surface-axis.test.cjs +++ b/tests/effort-surface-axis.test.cjs @@ -297,10 +297,14 @@ describe('#2481 renderEffortArgv — per-host syntax and clamping', () => { ); }); + // #3007: corrected — Codex gained 'max' (declared per-model), and no Codex + // model advertises 'minimal', so 'max' now passes through and 'minimal' + // clamps to 'low' instead. test('clamps the provider-unique tail levels', () => { - // claude has no `minimal`; codex has no `max`. + // claude has no `minimal`; codex has no `minimal` either (clamps to 'low'). assert.deepEqual(renderEffortArgv('claude', 'minimal', 'argv').argv, ['--effort', 'low']); - assert.deepEqual(renderEffortArgv('codex', 'max', 'argv').argv, ['-c', 'model_reasoning_effort=xhigh']); + assert.deepEqual(renderEffortArgv('codex', 'max', 'argv').argv, ['-c', 'model_reasoning_effort=max']); + assert.deepEqual(renderEffortArgv('codex', 'minimal', 'argv').argv, ['-c', 'model_reasoning_effort=low']); }); test('emits nothing when the surface is not argv', () => { diff --git a/tests/install-runtime-artifacts.test.cjs b/tests/install-runtime-artifacts.test.cjs index 01f3f4bb6..4ba9b6039 100644 --- a/tests/install-runtime-artifacts.test.cjs +++ b/tests/install-runtime-artifacts.test.cjs @@ -3834,7 +3834,10 @@ describe('#443 Config-driven: effort.agent_overrides drives install-time effort' `gsd-planner.toml should have model_reasoning_effort = "low" from config override\nActual:\n${tomlContent.slice(0, 500)}`); }); - test('Codex .toml clamps effort max → xhigh when agent_overrides.gsd-planner=max', () => { + // #3007: corrected — Codex's gpt-5.6-sol advertises 'max' in its own + // supported_reasoning_levels, so install-time rendering now passes 'max' + // through instead of clamping it to 'xhigh'. + test('Codex .toml renders effort max → max when agent_overrides.gsd-planner=max', () => { const projectDir = path.dirname(codexHome); // Overwrite config with max override. #3241: include an explicit // model_overrides pin (D1 removed the resolver-only auto-embed). @@ -3860,11 +3863,11 @@ describe('#443 Config-driven: effort.agent_overrides drives install-time effort' ); assert.match(tomlContent, /^model\s*=\s*"gpt-5.6-sol"$/m, `gsd-planner.toml should pin Codex model when runtime:"codex" is configured\nActual:\n${tomlContent.slice(0, 500)}`); - // Codex does not support 'max' → clamped to 'xhigh' - assert.match(tomlContent, /^model_reasoning_effort\s*=\s*"xhigh"$/m, - `gsd-planner.toml should clamp max → xhigh for Codex\nActual:\n${tomlContent.slice(0, 500)}`); - assert.doesNotMatch(tomlContent, /model_reasoning_effort\s*=\s*"max"/, - 'Codex .toml must never contain model_reasoning_effort = "max"'); + // gpt-5.6-sol advertises 'max' -> renders through unchanged, no clamp. + assert.match(tomlContent, /^model_reasoning_effort\s*=\s*"max"$/m, + `gsd-planner.toml should render max for Codex on a model that advertises it\nActual:\n${tomlContent.slice(0, 500)}`); + assert.doesNotMatch(tomlContent, /model_reasoning_effort\s*=\s*"xhigh"/, + 'Codex .toml must not clamp "max" down to "xhigh" for a model that advertises max'); }); }); diff --git a/tests/model-catalog.unit.test.cjs b/tests/model-catalog.unit.test.cjs new file mode 100644 index 000000000..f02aa49f5 --- /dev/null +++ b/tests/model-catalog.unit.test.cjs @@ -0,0 +1,381 @@ +'use strict'; + +/** + * FAST, IN-PROCESS mutation-testing surface for `model-catalog.cjs` (#3007). + * + * Root cause this file exists to fix: `model-catalog.cjs` was entirely + * outside Stryker's covered-module list, so the #3007 per-model Codex effort + * rewrite (`renderEffortForRuntime`, `CODEX_MODEL_EFFORT`, the 'ultra' + * rejection, the ladder walk-up) had zero mutation coverage. Following the + * #2790 precedent (see `tests/planning-inspect.unit.test.cjs`), this is a + * dedicated, spawn-free, in-process unit file rather than pointing the shard + * at an integration test — `tests/model-resolver.test.cjs` uses + * `runGsdTools` heavily and would hit the same 15-minute shard-cap cancel + * that #2790 documented (one `node --test ` invocation costs whatever + * its slowest case costs, per mutant). + * + * NEVER spawn a child process here — no `runGsdTools`, `spawnSync`, + * `execFileSync`, or CLI invocation of any kind, and no filesystem writes. + * Every case below requires the BUILT `.cjs` artifact directly and calls its + * exports in-process. + * + * Every value asserted below was verified by requiring the built lib + * directly and inspecting the real returned object — never guessed from + * reading the source alone (CLAUDE.md "verify assertions by executing, not + * retyping"). + */ + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); + +const catalog = require('../gsd-core/bin/lib/model-catalog.cjs'); + +const { + VALID_TIERS, + VALID_AGENT_TIERS, + KNOWN_RUNTIMES, + KNOWN_PROVIDERS, + RUNTIMES_WITH_REASONING_EFFORT, + RUNTIMES_WITH_FAST_MODE, + ADAPTIVE_TIER_VALUES, + CODEX_MODEL_EFFORT, + MODEL_ALIAS_MAP, + PROVIDER_PRESETS, + isAnthropicFlavoredModel, + getAgentToModelMapForProfile, + formatAgentToModelMapAsTable, + renderEffortArgv, + renderEffortForRuntime, + nextTier, + mergeEffortTierDefaults, +} = catalog; + +describe('model-catalog: exported enums/maps', () => { + test('VALID_TIERS is opus/sonnet/haiku plus inherit', () => { + assert.deepEqual(new Set(VALID_TIERS), new Set(['opus', 'sonnet', 'haiku', 'inherit'])); + }); + + test('VALID_AGENT_TIERS is light/standard/heavy', () => { + assert.deepEqual(new Set(VALID_AGENT_TIERS), new Set(['light', 'standard', 'heavy'])); + }); + + test('ADAPTIVE_TIER_VALUES excludes inherit', () => { + assert.deepEqual(new Set(ADAPTIVE_TIER_VALUES), new Set(['opus', 'sonnet', 'haiku'])); + assert.equal(ADAPTIVE_TIER_VALUES.has('inherit'), false); + }); + + test('KNOWN_RUNTIMES includes codex and claude', () => { + assert.equal(KNOWN_RUNTIMES.has('codex'), true); + assert.equal(KNOWN_RUNTIMES.has('claude'), true); + }); + + test('KNOWN_PROVIDERS excludes generic sentinel', () => { + assert.equal(KNOWN_PROVIDERS.has('generic'), false); + assert.equal(KNOWN_PROVIDERS.has('anthropic'), true); + assert.equal(KNOWN_PROVIDERS.has('openai'), true); + }); + + test('RUNTIMES_WITH_REASONING_EFFORT is codex-only', () => { + assert.deepEqual(new Set(RUNTIMES_WITH_REASONING_EFFORT), new Set(['codex'])); + }); + + test('RUNTIMES_WITH_FAST_MODE is api-only', () => { + assert.deepEqual(new Set(RUNTIMES_WITH_FAST_MODE), new Set(['api'])); + }); + + test('CODEX_MODEL_EFFORT has a baseline and per-model sets, sol includes ultra', () => { + assert.ok(CODEX_MODEL_EFFORT['_baseline'] instanceof Set); + assert.equal(CODEX_MODEL_EFFORT['_baseline'].has('max'), true); + assert.equal(CODEX_MODEL_EFFORT['gpt-5.6-sol'].has('ultra'), true); + assert.equal(CODEX_MODEL_EFFORT['gpt-5.6-terra'].has('ultra'), false); + assert.equal(CODEX_MODEL_EFFORT['gpt-5.6-luna'].has('ultra'), false); + }); + + test('MODEL_ALIAS_MAP maps opus/sonnet/haiku to claude model ids', () => { + assert.equal(MODEL_ALIAS_MAP.opus, 'claude-opus-4-8'); + assert.equal(MODEL_ALIAS_MAP.sonnet, 'claude-sonnet-5'); + assert.equal(MODEL_ALIAS_MAP.haiku, 'claude-haiku-4-5'); + }); + + test('PROVIDER_PRESETS is a non-empty object keyed by provider name', () => { + assert.equal(typeof PROVIDER_PRESETS, 'object'); + assert.ok(Object.keys(PROVIDER_PRESETS).length > 0); + assert.ok('anthropic' in PROVIDER_PRESETS); + }); +}); + +describe('model-catalog: isAnthropicFlavoredModel', () => { + test('bare tier aliases are flavored', () => { + assert.equal(isAnthropicFlavoredModel('opus'), true); + assert.equal(isAnthropicFlavoredModel('sonnet'), true); + assert.equal(isAnthropicFlavoredModel('haiku'), true); + assert.equal(isAnthropicFlavoredModel('fable'), true); + }); + + test('claude-* ids are flavored in every provider namespacing', () => { + assert.equal(isAnthropicFlavoredModel('claude-opus-4-8'), true); + assert.equal(isAnthropicFlavoredModel('anthropic/claude-opus-4-8'), true); + assert.equal(isAnthropicFlavoredModel('us.anthropic.claude-opus-4-8'), true); + }); + + test('negative case: gpt-* id is not flavored', () => { + assert.equal(isAnthropicFlavoredModel('gpt-5.6-sol'), false); + }); + + test('non-string input is not flavored', () => { + assert.equal(isAnthropicFlavoredModel(123), false); + assert.equal(isAnthropicFlavoredModel(null), false); + assert.equal(isAnthropicFlavoredModel(undefined), false); + }); +}); + +describe('model-catalog: getAgentToModelMapForProfile / formatAgentToModelMapAsTable', () => { + const EXPECTED_PLANNER_TIER = { quality: 'opus', balanced: 'opus', budget: 'sonnet', adaptive: 'opus' }; + const EXPECTED_MAPPER_TIER = { quality: 'sonnet', balanced: 'haiku', budget: 'haiku', adaptive: 'haiku' }; + + for (const profile of ['quality', 'balanced', 'budget', 'adaptive']) { + test(`profile '${profile}' returns the expected tier for known agents`, () => { + const map = getAgentToModelMapForProfile(profile); + assert.equal(map['gsd-planner'], EXPECTED_PLANNER_TIER[profile]); + assert.equal(map['gsd-codebase-mapper'], EXPECTED_MAPPER_TIER[profile]); + assert.ok(Object.keys(map).length > 0); + for (const value of Object.values(map)) { + assert.equal(typeof value, 'string'); + assert.ok(value.length > 0); + } + }); + } + + test("profile 'inherit' maps every agent to the literal string 'inherit'", () => { + const map = getAgentToModelMapForProfile('inherit'); + for (const value of Object.values(map)) { + assert.equal(value, 'inherit'); + } + }); + + test('invalid profile falls back to balanced', () => { + const balanced = getAgentToModelMapForProfile('balanced'); + const invalid = getAgentToModelMapForProfile('totally-bogus-profile'); + assert.deepEqual(invalid, balanced); + }); + + test('formatAgentToModelMapAsTable pads columns and renders a header/separator', () => { + const out = formatAgentToModelMapAsTable({ agentA: 'model-x', b: 'model-y-longer' }); + const lines = out.split('\n'); + assert.equal(lines[0].includes('Agent'), true); + assert.equal(lines[0].includes('Model'), true); + assert.equal(lines[1].includes('┼'), true); + assert.equal(lines[2].includes('agentA'), true); + assert.equal(lines[2].includes('model-x'), true); + assert.equal(lines[3].includes('model-y-longer'), true); + }); +}); + +describe('model-catalog: renderEffortForRuntime — codex per-model', () => { + test("sol's 'max' passes through unclamped", () => { + const r = renderEffortForRuntime('codex', 'max', 'gpt-5.6-sol'); + assert.equal(r.value, 'max'); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + assert.equal(r.param, 'model_reasoning_effort'); + assert.equal(r.channel, 'api'); + }); + + test("'minimal' clamps up to 'low' with clamped:true and a non-empty reason (sol/terra/luna)", () => { + for (const model of ['gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.6-luna']) { + const r = renderEffortForRuntime('codex', 'minimal', model); + assert.equal(r.value, 'low'); + assert.equal(r.clamped, true); + assert.equal(typeof r.reason, 'string'); + assert.ok(r.reason.length > 0); + assert.ok(r.reason.includes(model)); + } + }); + + test("'ultra' is rejected outright even for sol which advertises it (value:null)", () => { + const r = renderEffortForRuntime('codex', 'ultra', 'gpt-5.6-sol'); + assert.equal(r.value, null); + assert.equal(r.param, null); + assert.equal(r.channel, null); + assert.equal(r.clamped, false); + assert.ok(r.reason.includes('#2167')); + }); + + test('low/medium/high/xhigh pass through unclamped with reason:null', () => { + for (const level of ['low', 'medium', 'high', 'xhigh']) { + const r = renderEffortForRuntime('codex', level, 'gpt-5.6-terra'); + assert.equal(r.value, level); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + } + }); + + test('unknown model falls back to the family baseline', () => { + const r = renderEffortForRuntime('codex', 'max', 'gpt-9-never-heard-of-it'); + assert.equal(r.value, 'max'); + assert.equal(r.clamped, false); + }); + + test('omitted / null / empty-string model all fall back to the family baseline', () => { + const omitted = renderEffortForRuntime('codex', 'minimal'); + const nullModel = renderEffortForRuntime('codex', 'minimal', null); + const emptyModel = renderEffortForRuntime('codex', 'minimal', ''); + for (const r of [omitted, nullModel, emptyModel]) { + assert.equal(r.value, 'low'); + assert.equal(r.clamped, true); + assert.ok(r.reason.includes('codex family baseline')); + } + }); + + test("off-ladder input ('MAX') falls through to the runtime-level clamp unchanged", () => { + const r = renderEffortForRuntime('codex', 'MAX', 'gpt-5.6-terra'); + assert.equal(r.value, 'MAX'); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + }); +}); + +describe('model-catalog: renderEffortForRuntime — cross-runtime', () => { + test("'inherit' passes through on every known runtime", () => { + for (const runtime of [...KNOWN_RUNTIMES, 'totally-unknown-runtime']) { + const r = renderEffortForRuntime(runtime, 'inherit'); + assert.equal(r.value, 'inherit'); + assert.equal(r.param, null); + assert.equal(r.channel, null); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + } + }); + + test('unknown runtime passes the requested value through unchanged', () => { + const r = renderEffortForRuntime('totally-unknown-runtime', 'high'); + assert.equal(r.value, 'high'); + assert.equal(r.param, null); + assert.equal(r.channel, null); + assert.equal(r.clamped, false); + }); + + test('claude passes high through unchanged and clamps minimal to low', () => { + const high = renderEffortForRuntime('claude', 'high'); + assert.equal(high.value, 'high'); + assert.equal(high.clamped, false); + assert.equal(high.reason, null); + + const minimal = renderEffortForRuntime('claude', 'minimal'); + assert.equal(minimal.value, 'low'); + assert.equal(minimal.clamped, true); + assert.ok(minimal.reason.includes('claude')); + }); +}); + +describe('model-catalog: renderEffortArgv', () => { + test("effortSurface gate: 'none'/undefined produce empty, 'argv' produces an argument", () => { + assert.deepEqual(renderEffortArgv('codex', 'max', 'none'), { argv: [], value: null, host: 'codex' }); + assert.deepEqual(renderEffortArgv('codex', 'max', undefined), { argv: [], value: null, host: 'codex' }); + const r = renderEffortArgv('codex', 'max', 'argv'); + assert.deepEqual(r.argv, ['-c', 'model_reasoning_effort=max']); + assert.equal(r.value, 'max'); + }); + + test("codex 'max' passes through, 'minimal' clamps to 'low'", () => { + const max = renderEffortArgv('codex', 'max', 'argv'); + assert.deepEqual(max.argv, ['-c', 'model_reasoning_effort=max']); + assert.equal(max.value, 'max'); + + const minimal = renderEffortArgv('codex', 'minimal', 'argv'); + assert.deepEqual(minimal.argv, ['-c', 'model_reasoning_effort=low']); + assert.equal(minimal.value, 'low'); + }); + + test('unknown host produces empty result, never throws', () => { + const r = renderEffortArgv('totally-bogus-host', 'high', 'argv'); + assert.deepEqual(r, { argv: [], value: null, host: 'totally-bogus-host' }); + }); + + test('prototype-chain host names (__proto__, constructor, toString) return empty, never throw', () => { + for (const host of ['__proto__', 'constructor', 'toString']) { + const r = renderEffortArgv(host, 'high', 'argv'); + assert.deepEqual(r.argv, []); + assert.equal(r.value, null); + assert.equal(r.host, host); + } + }); +}); + +describe('model-catalog: nextTier', () => { + // Probed directly against the built module: light -> standard -> heavy, + // and heavy SATURATES at 'heavy' rather than wrapping back to 'light'. + test("advances light -> standard and standard -> heavy", () => { + assert.equal(nextTier('light'), 'standard'); + assert.equal(nextTier('standard'), 'heavy'); + }); + + test("saturates at the top tier: 'heavy' stays 'heavy'", () => { + assert.equal(nextTier('heavy'), 'heavy'); + }); + + test('unknown/garbage tier returns null', () => { + assert.equal(nextTier('bogus'), null); + assert.equal(nextTier('LIGHT'), null); // case-sensitive, not in the order array + }); + + test('empty string, null, undefined, and non-string input all return null', () => { + assert.equal(nextTier(''), null); + assert.equal(nextTier(null), null); + assert.equal(nextTier(undefined), null); + assert.equal(nextTier(123), null); + }); +}); + +describe('model-catalog: mergeEffortTierDefaults (#3531)', () => { + const manifest = { opus: 'sonnet', sonnet: 'haiku', haiku: 'opus' }; + const isValid = (v) => typeof v === 'string' && ['opus', 'sonnet', 'haiku'].includes(v); + + test('absent/empty override leaves the manifest values untouched', () => { + assert.deepEqual(mergeEffortTierDefaults(manifest, undefined, isValid), manifest); + assert.deepEqual(mergeEffortTierDefaults(manifest, {}, isValid), manifest); + }); + + test('partial override changes only the named tier; other tiers keep their built-in values', () => { + const merged = mergeEffortTierDefaults(manifest, { sonnet: 'opus' }, isValid); + assert.equal(merged.sonnet, 'opus'); + assert.equal(merged.opus, manifest.opus); + assert.equal(merged.haiku, manifest.haiku); + }); + + test('an invalid override value for a tier is ignored; the built-in for that tier is retained', () => { + const merged = mergeEffortTierDefaults(manifest, { sonnet: 'bogus' }, isValid); + assert.equal(merged.sonnet, manifest.sonnet); + }); + + test('an override key not present in the manifest is still merged in (isValid gates values, not tier names)', () => { + const merged = mergeEffortTierDefaults(manifest, { newtier: 'opus' }, isValid); + assert.equal(merged.newtier, 'opus'); + assert.equal(merged.opus, manifest.opus); + assert.equal(merged.sonnet, manifest.sonnet); + assert.equal(merged.haiku, manifest.haiku); + }); + + test('__proto__/constructor/prototype override keys are skipped (house pollution guard)', () => { + const merged = mergeEffortTierDefaults(manifest, { __proto__: 'opus', constructor: 'sonnet', prototype: 'haiku' }, isValid); + assert.deepEqual(merged, manifest); + assert.equal(Object.prototype.hasOwnProperty.call(merged, '__proto__'), false); + assert.equal(Object.prototype.hasOwnProperty.call(merged, 'constructor'), false); + }); + + test('null, non-object, and array config all leave the manifest untouched', () => { + assert.deepEqual(mergeEffortTierDefaults(manifest, null, isValid), manifest); + assert.deepEqual(mergeEffortTierDefaults(manifest, 'bogus', isValid), manifest); + assert.deepEqual(mergeEffortTierDefaults(manifest, ['x'], isValid), manifest); + }); + + test('a falsy/missing manifest merges over an empty base object', () => { + assert.deepEqual(mergeEffortTierDefaults(null, { sonnet: 'opus' }, isValid), { sonnet: 'opus' }); + }); + + test('the function is pure: it never mutates the manifest it is given', () => { + const original = { ...manifest }; + mergeEffortTierDefaults(manifest, { sonnet: 'opus' }, isValid); + assert.deepEqual(manifest, original); + }); +}); diff --git a/tests/model-resolver.test.cjs b/tests/model-resolver.test.cjs index 23096f411..b6e184232 100644 --- a/tests/model-resolver.test.cjs +++ b/tests/model-resolver.test.cjs @@ -349,7 +349,10 @@ describe('#3533 effort inherit: expressible at every layer, never a wire level', } // Concrete levels unchanged. assert.strictEqual(renderEffortForRuntime('claude', 'minimal').value, 'low'); - assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'xhigh'); + // #3007: corrected — Codex DOES advertise 'max' (per-model table), so + // ADR-443's "Codex has no max" premise went stale and this pinned the + // defect (clamping 'max' down to 'xhigh') instead of the fix. + assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'max'); assert.strictEqual(renderEffortForRuntime('claude', 'xhigh').value, 'xhigh'); }); }); @@ -1641,9 +1644,13 @@ const { // Anthropic: output_config.effort — https://docs.anthropic.com (Claude API) // OpenAI: model_reasoning_effort — https://platform.openai.com/docs (Codex) // ───────────────────────────────────────────────────────────────────────────── +// #3007: corrected — Codex's own models.json now advertises 'max' (and 'ultra', +// which is policy-rejected separately, #2167) but no Codex model advertises +// 'minimal'. The old enum here ('minimal'..'xhigh', no 'max') encoded ADR-443's +// stale premise, not the real API. const PROVIDER_EFFORT_ENUMS = { claude: new Set(['low', 'medium', 'high', 'xhigh', 'max']), - codex: new Set(['minimal', 'low', 'medium', 'high', 'xhigh']), + codex: new Set(['low', 'medium', 'high', 'xhigh', 'max']), }; // Helper: write config.json into a temp project @@ -1672,8 +1679,10 @@ describe('#443 integration (a): cross-provider validity invariant', () => { }); // Documented clamps must hold exactly - test("render('codex','max').value === 'xhigh' (max is Anthropic-only)", () => { - assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'xhigh'); + // #3007: corrected — Codex gained 'max' (declared per-model via + // supported_reasoning_levels); 'max' is no longer Anthropic-only. + test("render('codex','max').value === 'max' (Codex now advertises max)", () => { + assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'max'); }); test("render('claude','minimal').value === 'low' (minimal is Codex-only)", () => { @@ -2653,9 +2662,11 @@ describe('#443 resolveEffortForTier escalation', () => { // ─── Rendering / clamping ────────────────────────────────────────────────────── describe('#443 renderEffortForRuntime', () => { - test('codex: "max" clamps to "xhigh"', () => { + // #3007: corrected — Codex gained 'max' (per-model supported_reasoning_levels); + // it no longer clamps to 'xhigh'. + test('codex: "max" passes through as "max"', () => { const r = renderEffortForRuntime('codex', 'max'); - assert.strictEqual(r.value, 'xhigh'); + assert.strictEqual(r.value, 'max'); assert.strictEqual(r.param, 'model_reasoning_effort'); }); @@ -2666,8 +2677,10 @@ describe('#443 renderEffortForRuntime', () => { assert.strictEqual(renderEffortForRuntime('codex', 'xhigh').value, 'xhigh'); }); - test('codex: "minimal" passthrough', () => { - assert.strictEqual(renderEffortForRuntime('codex', 'minimal').value, 'minimal'); + // #3007: corrected — no Codex model advertises 'minimal'; it now clamps up + // to the family floor, 'low', instead of passing through. + test('codex: "minimal" clamps to "low"', () => { + assert.strictEqual(renderEffortForRuntime('codex', 'minimal').value, 'low'); }); test('claude: "minimal" clamps to "low"', () => { @@ -2728,7 +2741,13 @@ describe('#443 resolve-execution CLI command', () => { assert.ok('profile' in output, 'should have profile field'); }); - test('codex runtime -> effort_param=model_reasoning_effort, max clamps to xhigh, fast_mode_supported=false', () => { + // NOTE: effort.default: 'max' never reaches the renderer for gsd-planner here — + // gsd-planner is a heavy/opus-tier agent, and its routing-tier default outranks + // effort.default in resolution precedence, so the resolved level is 'xhigh' before + // the renderer ever sees 'max'. effort_clamped=false and effort_requested='xhigh' + // prove this is precedence, not the #3007 clamp — do not "correct" this back to + // expecting 'max'. + test('codex runtime -> effort_param=model_reasoning_effort, tier default outranks effort.default, fast_mode_supported=false', () => { writeConfig(tmpDir, { runtime: 'codex', effort: { default: 'max' }, @@ -2738,6 +2757,29 @@ describe('#443 resolve-execution CLI command', () => { const output = JSON.parse(result.output); assert.strictEqual(output.effort_param, 'model_reasoning_effort'); assert.strictEqual(output.effort_rendered, 'xhigh'); + assert.strictEqual(output.effort_clamped, false); + assert.strictEqual(output.effort_requested, 'xhigh'); + // fast_mode_supported: codex does not support fast mode via subagent + assert.strictEqual(output.fast_mode_supported, false); + }); + + // #3007: Codex gained 'max' (per-model supported_reasoning_levels), so 'max' now + // renders through unchanged instead of clamping to 'xhigh'. agent_overrides is used + // here (not effort.default) because it outranks the routing-tier default, which is + // what actually lets 'max' reach the renderer end-to-end. + test('codex runtime -> max survives to the wire via agent_overrides, fast_mode_supported=false', () => { + writeConfig(tmpDir, { + runtime: 'codex', + effort: { default: 'max', agent_overrides: { 'gsd-planner': 'max' } }, + }); + const result = runGsdTools(['resolve-execution', 'gsd-planner'], tmpDir, { HOME: tmpDir }); + assert.ok(result.success, `Command failed: ${result.error}`); + const output = JSON.parse(result.output); + assert.strictEqual(output.effort_param, 'model_reasoning_effort'); + assert.strictEqual(output.effort, 'max'); + assert.strictEqual(output.effort_rendered, 'max'); + assert.strictEqual(output.effort_requested, 'max'); + assert.strictEqual(output.effort_clamped, false); // fast_mode_supported: codex does not support fast mode via subagent assert.strictEqual(output.fast_mode_supported, false); }); @@ -5928,3 +5970,353 @@ describe('issue #2612: partial override merge for new Group A runtimes', () => { }); }); } + +// ─── #3007: Codex effort capability is per-model ───────────────────────────── +// +// Ground truth (Codex's models.json): gpt-5.6-sol advertises +// low/medium/high/xhigh/max/ultra; gpt-5.6-luna and gpt-5.6-terra advertise +// low/medium/high/xhigh/max (no ultra); no Codex model advertises 'minimal'. +// An unknown/omitted model id falls back to the family baseline +// (low/medium/high/xhigh/max). + +const CODEX_MODEL_EFFORT_SETS = { + 'gpt-5.6-sol': new Set(['low', 'medium', 'high', 'xhigh', 'max', 'ultra']), + 'gpt-5.6-luna': new Set(['low', 'medium', 'high', 'xhigh', 'max']), + 'gpt-5.6-terra': new Set(['low', 'medium', 'high', 'xhigh', 'max']), +}; +const CODEX_FAMILY_BASELINE = new Set(['low', 'medium', 'high', 'xhigh', 'max']); +function advertisedCodexEfforts(model) { + return CODEX_MODEL_EFFORT_SETS[model] || CODEX_FAMILY_BASELINE; +} + +describe('#3007 — Codex effort capability is per-model, and every clamp is visible', () => { + test('max survives for a model that advertises it', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max', 'gpt-5.6-sol'); + assert.strictEqual(r.value, 'max'); + assert.strictEqual(r.clamped, false); + assert.strictEqual(r.reason, null); + }); + + test('max survives on luna, not only sol', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max', 'gpt-5.6-luna'); + assert.strictEqual(r.value, 'max'); + }); + + test('a level below the ceiling is not reported as clamped', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'xhigh', 'gpt-5.6-sol'); + assert.strictEqual(r.value, 'xhigh'); + assert.strictEqual(r.clamped, false); + }); + + test('ultra is rejected, never clamped to max', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'ultra', 'gpt-5.6-sol'); + assert.strictEqual(r.value, null); + assert.ok(typeof r.reason === 'string' && /deleg/i.test(r.reason), `reason should mention delegation: ${r.reason}`); + }); + + test('minimal clamps to low on a model that floors at low', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal', 'gpt-5.6-luna'); + assert.strictEqual(r.value, 'low'); + assert.strictEqual(r.clamped, true); + assert.ok(typeof r.reason === 'string' && r.reason.length > 0); + }); + + test('minimal clamps on sol too', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal', 'gpt-5.6-sol'); + assert.strictEqual(r.value, 'low'); + assert.strictEqual(r.clamped, true); + }); + + test('an unknown model id gets the family baseline, not a clamp', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max', 'gpt-9.9-unreleased'); + assert.strictEqual(r.value, 'max'); + assert.strictEqual(r.clamped, false); + }); + + test('an unknown model id still floors at low', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal', 'gpt-9.9-unreleased'); + assert.strictEqual(r.value, 'low'); + assert.strictEqual(r.clamped, true); + }); + + test('the model-less form resolves against the family baseline', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max'); + assert.strictEqual(r.value, 'max'); + }); + + test('the model-less form still floors at low', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal'); + assert.strictEqual(r.value, 'low'); + }); + + test('the two-argument signature keeps working for every level', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const UNCHANGED = new Set(['low', 'medium', 'high', 'xhigh']); + for (const level of ['minimal', 'low', 'medium', 'high', 'xhigh', 'max']) { + let r; + assert.doesNotThrow(() => { r = renderEffortForRuntime('codex', level); }); + if (UNCHANGED.has(level)) { + assert.strictEqual(r.value, level, `level ${level} should pass through unchanged`); + } + } + }); + + test('claude rendering is untouched by the codex table', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + assert.strictEqual(renderEffortForRuntime('claude', 'max').value, 'max'); + assert.strictEqual(renderEffortForRuntime('claude', 'minimal').value, 'low'); + // passing a codex model id as the 3rd arg to claude must change nothing + const withModel = renderEffortForRuntime('claude', 'max', 'gpt-5.6-sol'); + assert.strictEqual(withModel.value, 'max'); + }); + + test('inherit is not measured against any supported set', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + for (const runtime of ['claude', 'codex', 'something-unknown']) { + for (const model of [undefined, 'gpt-5.6-sol']) { + const r = renderEffortForRuntime(runtime, 'inherit', model); + assert.deepStrictEqual( + { value: r.value, param: r.param, channel: r.channel }, + { value: 'inherit', param: null, channel: null }, + `runtime=${runtime} model=${JSON.stringify(model)}`, + ); + // #3007's new requested/clamped/reason fields must be honest on this + // path too: 'inherit' is never a clamp target, so it can never be + // reported as clamped, and it echoes itself back as `requested`. + assert.deepStrictEqual( + { requested: r.requested, clamped: r.clamped, reason: r.reason }, + { requested: 'inherit', clamped: false, reason: null }, + `runtime=${runtime} model=${JSON.stringify(model)}`, + ); + } + } + }); + + test('an undeclared runtime still renders nothing', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('something-unknown', 'high'); + assert.strictEqual(r.param, null); + assert.strictEqual(r.channel, null); + // #3007's new fields must also be honest for a host with no spec at all: + // the value passes straight through, unclamped, and there is nothing to + // explain about it. + assert.strictEqual(r.value, 'high'); + assert.strictEqual(r.requested, 'high'); + assert.strictEqual(r.clamped, false); + assert.strictEqual(r.reason, null); + }); + + test('a bare tier alias is not treated as a model id', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + let r; + assert.doesNotThrow(() => { r = renderEffortForRuntime('codex', 'max', 'opus'); }); + assert.strictEqual(r.value, 'max'); + }); + + test('an anthropic-flavored id is not an effort-capability error', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + let r; + assert.doesNotThrow(() => { r = renderEffortForRuntime('codex', 'max', 'claude-opus-4-8'); }); + assert.strictEqual(r.value, 'max'); + }); + + test('an empty model argument behaves exactly like omitting it', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + for (const level of ['max', 'minimal']) { + const omitted = renderEffortForRuntime('codex', level); + for (const emptyish of ['', null, undefined]) { + assert.deepStrictEqual( + renderEffortForRuntime('codex', level, emptyish), + omitted, + `level=${level} emptyish=${JSON.stringify(emptyish)}`, + ); + } + } + }); + + test('effort matching is case-sensitive, as the ladder always was', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + // 'MAX' is not 'max' — whatever the renderer does with an unrecognized + // level, it must not silently treat it as a clean pass-through of 'max'. + // 'MAX' is off the EFFORT_LADDER entirely, so #3007's per-model path + // must fall through to the runtime-level clamp verbatim rather than + // inventing new handling — and it must report that verbatim pass-through + // as NOT clamped, with `requested` echoing the exact (unrecognized) input. + // Under a fully reverted #3007, `clamped`/`requested` do not exist on the + // returned object at all, so this fails there too. + const r = renderEffortForRuntime('codex', 'MAX', 'gpt-5.6-sol'); + assert.notStrictEqual(r.value, 'max'); + assert.strictEqual(r.value, 'MAX'); + assert.strictEqual(r.clamped, false); + assert.strictEqual(r.requested, 'MAX'); + }); + + test('a clamp never lands on ultra, even for a model that only advertises it', () => { + // The clamp-up loop in renderEffortForRuntime walks EFFORT_LADDER + // upward from the requested level looking for the model's floor. Today's + // catalog can't actually exercise the 'ultra'-as-clamp-target path — every + // model advertises 'max', so the allowed.has() fast path always returns + // first. This test guards a latent path, not a currently-reachable one: + // do not delete it as redundant just because it never fails today. The + // invariant it protects is general — for EVERY model and EVERY ladder + // level, a clamp must never produce 'ultra', because that would re-enter + // by the back door the delegation mode the #2167 rejection exists to + // keep out. + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const models = [undefined, 'gpt-5.6-sol', 'gpt-5.6-luna', 'gpt-5.6-terra', 'gpt-9.9-unreleased']; + const levels = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + for (const model of models) { + for (const level of levels) { + const r = renderEffortForRuntime('codex', level, model); + assert.notStrictEqual(r.value, 'ultra', `model=${JSON.stringify(model)} level=${level}`); + } + } + }); +}); + +// ─── #3007 PARITY: known-defect gauntlet ────────────────────────────────────── + +test('every catalog preset ships an effort its own model supports', () => { + const catalogPath = path.join(__dirname, '..', 'gsd-core', 'bin', 'shared', 'model-catalog.json'); + const catalog = JSON.parse(fs.readFileSync(catalogPath, 'utf8')); + const offenders = []; + + const checkEntry = (entryPath, entry) => { + if (!entry || typeof entry !== 'object') return; + const model = entry.model; + const effort = entry.reasoning_effort; + if (typeof model !== 'string' || !model.startsWith('gpt-') || typeof effort !== 'string') return; + const allowed = advertisedCodexEfforts(model); + if (!allowed.has(effort)) { + offenders.push(`${entryPath}: model=${model} reasoning_effort=${effort} (not in {${[...allowed].join(', ')}})`); + } + }; + + const codexDefaults = catalog.runtimeTierDefaults && catalog.runtimeTierDefaults.codex; + if (codexDefaults) { + for (const [tier, entry] of Object.entries(codexDefaults)) { + checkEntry(`runtimeTierDefaults.codex.${tier}`, entry); + } + } + + const openaiPresets = catalog.providerPresets && catalog.providerPresets.openai; + if (openaiPresets) { + for (const [tier, profiles] of Object.entries(openaiPresets)) { + if (!profiles || typeof profiles !== 'object') continue; + for (const [profile, entry] of Object.entries(profiles)) { + checkEntry(`providerPresets.openai.${tier}.${profile}`, entry); + } + } + } + + assert.deepStrictEqual(offenders, [], `offending presets (model does not advertise the assigned effort):\n${offenders.join('\n')}`); +}); + +// ─── #3007 PARITY: argv channel must agree with the render-for-runtime channel ─ +// +// `renderEffortArgv('codex', ...)` (invocation-time, `-c model_reasoning_effort=`) +// and `renderEffortForRuntime('codex', ...)` (install-time / api channel) each +// read their own EFFORT_ARGV.codex / EFFORT_RENDERING.codex tables. Those two +// tables must describe the SAME capability, or a user gets a different answer +// depending on which code path asked — the repo's documented "generative fix +// divergence" class (two surfaces reading one fact that can drift apart). + +describe('#3007 PARITY: argv channel agrees with renderEffortForRuntime for codex', () => { + test('renderEffortArgv and renderEffortForRuntime never disagree across the ladder', () => { + const { renderEffortArgv, renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const LADDER = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + for (const level of LADDER) { + const argvResult = renderEffortArgv('codex', level, 'argv'); + const runtimeResult = renderEffortForRuntime('codex', level); + if (level === 'ultra') { + // 'ultra' is Codex's automatic-delegation switch (#2167): the + // install-time channel rejects it outright (value: null, policy + // reason). The argv channel has no concept of that policy rejection + // — it simply isn't in EFFORT_ARGV.codex's supported set, so it also + // degrades to a `null` value. Both channels landing on `null` here + // is the explicit agreement contract for this level; it is not a + // case the parity check can skip. + assert.strictEqual(argvResult.value, null, `argv channel must also refuse ultra: got ${argvResult.value}`); + assert.strictEqual(runtimeResult.value, null, `runtime channel must refuse ultra: got ${runtimeResult.value}`); + continue; + } + assert.strictEqual( + argvResult.value, + runtimeResult.value, + `argv/runtime channels disagree for level=${level}: argv=${argvResult.value} runtime=${runtimeResult.value}`, + ); + } + }); +}); + +// ─── #3007 PROPERTY: rendered codex effort is always within the model's ceiling ─ + +describe('#3007 PROPERTY: renderEffortForRuntime never renders a level the model does not advertise', () => { + test('every (model, level) pair is exhaustively checked, not sampled', () => { + // fc.constantFrom over MODELS x LADDER with numRuns: 200 is very likely to + // hit all 4 x 7 = 28 pairs but is not GUARANTEED to. This deterministic + // nested loop covers the full cross-product with certainty; it is kept + // alongside the fast-check property below (not instead of it) because the + // repo requires a property test for a closed-vocabulary contract, and the + // property still adds shrinking value on failure. + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const MODELS = ['gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.6-luna', 'gpt-9.9-unreleased-model']; + const LADDER = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + for (const model of MODELS) { + for (const level of LADDER) { + const allowed = advertisedCodexEfforts(model); + const r = renderEffortForRuntime('codex', level, model); + if (r.value === null) { + assert.strictEqual(level, 'ultra', `only ultra may be rejected: model=${model} level=${level}`); + continue; + } + assert.ok(allowed.has(r.value), `rendered value not in model's advertised set: model=${model} level=${level} value=${r.value}`); + if (r.value === level) { + assert.strictEqual(r.clamped, false, `pass-through reported as clamped: model=${model} level=${level}`); + } else { + assert.strictEqual(r.clamped, true, `changed value not reported as clamped: model=${model} level=${level} value=${r.value}`); + } + } + } + }); + + test('for every (model, level) pair, the outcome is pass-through, clamp, or reject', () => { + const fc = require('./helpers/fast-check-setup.cjs'); + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const MODELS = ['gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.6-luna', 'gpt-9.9-unreleased-model']; + const LADDER = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + fc.assert( + fc.property(fc.constantFrom(...MODELS), fc.constantFrom(...LADDER), (model, level) => { + const allowed = advertisedCodexEfforts(model); + const r = renderEffortForRuntime('codex', level, model); + if (r.value === null) { + // Rejection is a POLICY outcome, not a capability one, so it is not + // predicted by the advertised set. `ultra` is refused even on sol, + // which does advertise it: GSD is deliberately stricter than Codex + // because ultra turns on automatic task delegation (#2167). + // Reserving null for exactly `ultra` is what keeps that a decision + // rather than a side effect — any OTHER null is a bug. + assert.strictEqual(level, 'ultra', `only ultra may be rejected: model=${model} level=${level}`); + return; + } + assert.ok(allowed.has(r.value), `rendered value not in model's advertised set: model=${model} level=${level} value=${r.value}`); + if (r.value === level) { + assert.strictEqual(r.clamped, false, `pass-through reported as clamped: model=${model} level=${level}`); + } else { + assert.strictEqual(r.clamped, true, `changed value not reported as clamped: model=${model} level=${level} value=${r.value}`); + } + }), + { seed: 3007, numRuns: 200 }, + ); + }); +}); diff --git a/tests/mutation-matrix-ratchet.test.cjs b/tests/mutation-matrix-ratchet.test.cjs index 497cb3041..525ce5e2a 100644 --- a/tests/mutation-matrix-ratchet.test.cjs +++ b/tests/mutation-matrix-ratchet.test.cjs @@ -202,6 +202,7 @@ const RATCHET_BASELINE = { 'planning-inspect': 56, // CI run 32392791843: 57.03% (unit shard); ratchet candidate vs TARGET 80 'plan-document': 75, // CI run 32392791843: 76.58% (unit shard) 'planning-command-router': 94, // CI run 32392791843: 95.65% (unit shard); already exceeds TARGET 80 + 'model-catalog': 58, // #3007: measured 59.62% in CI (248 killed / 168 survived); floor(59.62)-1 }; describe('mutation-matrix ratchet: floor equality enforcement', () => {