diff --git a/.changeset/wise-otters-greet.md b/.changeset/wise-otters-greet.md new file mode 100644 index 000000000..e17afc127 --- /dev/null +++ b/.changeset/wise-otters-greet.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 3765 +--- +**Codex reasoning effort is now resolved per model, and every clamp is visible** — `max` reaches Codex instead of being silently downgraded to `xhigh`, `minimal` clamps up to `low` instead of being sent to models that reject it, and `resolve-execution` reports the level you asked for alongside the one actually rendered. `ultra` is refused outright because it switches Codex into proactive task delegation underneath GSD's own orchestration. (#3007) diff --git a/CONTEXT.md b/CONTEXT.md index 808cf1812..a8de3bbd8 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -242,7 +242,7 @@ The seam deciding whether `.planning/` artifacts reach git. `commit_docs` resolv **Per-phase override (#3587, epic #2292 Phase 3).** A NEW tier resolves ABOVE the chain above, entirely inside `cmdCommit` (`src/commands.cts`'s `resolveCommitDocsPolicy`/`resolvePhaseCommitDocsOverride`) — deliberately NOT inside `loadConfigResolved`, which has no phase context and is called by nearly every command. The dynamic config key `phase_commit_docs.` (a `{ "": boolean }` map, registered in `config-schema.manifest.json`'s `dynamicKeyPatterns` and threaded through `config-loader.cts`'s `_baseConfig` projection the same way `agent_skills` is — a dynamic key absent from that hand-maintained allowlist is silently dropped on read, the exact failure mode `features.` demonstrates today) lets a tech lead commit one phase's artifacts while the project-wide `commit_docs` stays `false` (or the reverse). The phase being committed is resolved via the PRE-EXISTING `detectPhaseNumberFromFiles(files)` (the same #2539-hardened, project-code-aware derivation `branching_strategy` already used one branch below) and normalized through `normalizePhaseName` on both sides of the comparison, so `3`/`03`/`PROJ-03` hit one entry and a value scoped to a DIFFERENT phase never leaks. A non-boolean stored value (`"true"`, `1`, `null`) is never coerced — it falls through to the pre-existing chain untouched. When this tier suppresses a commit, the envelope's reason is `skipped_commit_docs_phase_false` — deliberately distinct from `skipped_commit_docs_false`, so a per-phase suppression is never reported as "your project setting is false" when it is actually `true`. **AC4 (byte-identical when unset):** with no `phase_commit_docs` key, this tier is a no-op and the three-tier chain above resolves exactly as before — pinned by the `folded:phase-commit-docs` block's C1-C5 in `tests/commit-docs-bypass.test.cjs`. The `phase_commit_docs.` grammar is a hand-copy of the canonical `PHASE_NUMBER_TOKEN_SOURCE` (`src/phase-id.cts`, #2128) into the hand-maintained schema manifest — pinned against drift by that same block's describe 'E' in `tests/commit-docs-bypass.test.cjs` (behavioral, over a shared shape list, per CLAUDE.md's Generative Fix Divergence class). **The two reasons are ordered, not peers, and `skipped_gitignored` is near-unreachable in a real project** (measured #3585): `cmdCommit` tests resolved `commit_docs` FIRST, and whenever `.planning/config.json` exists — which it does in every initialized project — the loader's gitignore auto-detect has already resolved that value to `false`, so the first branch returns `skipped_commit_docs_false` and the `isGitIgnored` branch below it is never reached. `skipped_gitignored` fires only when `config.json` is absent entirely, so the loader falls back to the `true` default and `cmdCommit`'s own check is what fires. A gitignore-driven skip therefore reports the config-driven reason; both members are behaviorally pinned by `tests/commit-docs-bypass.test.cjs` (B1-B3 and G1) so a rename fails loudly, but the reason a user sees does not distinguish *which* input suppressed the commit. **The gate is bypassable only from OUTSIDE the code**: a workflow step that types `git add` into its own shell reaches the index without passing through `cmdCommit`, and no code change can intercept that. Two guards therefore enforce it as text rather than at runtime: `tests/commit-files-pathspec.test.cjs` (#2269) requires every shipped `commit` invocation to declare `--files`, and `tests/commit-docs-bypass.test.cjs` (#1783, made repo-wide by #3585) requires every shipped `git add` able to reach `.planning/` to sit inside an **executable** `commit_docs` check — a markdown prose conditional ("**If `commit_docs` is true:**") is not a guard, because the bash block below it runs regardless. Both consume one shell tokenizer (`tests/helpers/shipped-command-scan.cjs`) and one exemption marker (`# gsd-scan-ignore: #NNN`, reason must cite a tracking ref per ADR-456). Guard state does not cross a fenced-block boundary: each fenced block is its own shell, so a guard opened in one block does not protect a `git add` in the next. Known limit: `.gitignore` has no effect on files git already TRACKS, so a project that committed `.planning/` before ignoring it keeps staging those paths — see `gsd-core/references/planning-config.md` and #3586. **`cmdCheckCommit`** (Command Module, verb `check-commit`) is a SEPARATE reader of the same `commit_docs` value, for callers outside GSD's own commit path: it inspects the staged set directly and refuses (non-zero exit) when `commit_docs` is `false` and any staged path is under `.planning/`; otherwise it allows. `gsd-tools commit-docs-guard enable`/`disable` (#3588) is the opt-in installer for a `.git/hooks/pre-commit` hook that shells out to exactly this verb, closing the one bypass the text-scan guards above cannot reach — a human or script running a bare `git add -A && git commit` in their OWN shell, outside any GSD-shipped workflow. The hook is identified by a `# gsd-core:commit-docs-guard` marker line (presence-checked, not byte-equality), is written only on explicit request (no install path wires it by default — locked by `tests/commands.test.cjs`'s E2 row), refuses rather than overwrite or delete a foreign `pre-commit`, resolves the real hooks dir via `git rev-parse --git-path hooks` so a linked worktree or submodule whose `.git` is a FILE works, and refuses outright when `core.hooksPath` is already set, because a written-but-ignored hook is worse than a refusal. ### Model Catalog Module -Leaf module owning the **static** model tables and the closed vocabularies derived from them — the tier/runtime/provider enums (`VALID_TIERS`, `VALID_AGENT_TIERS`, `KNOWN_RUNTIMES`, `KNOWN_PROVIDERS`, `RUNTIMES_WITH_REASONING_EFFORT`, `RUNTIMES_WITH_FAST_MODE`, `ADAPTIVE_TIER_VALUES`), the alias and profile maps (`MODEL_ALIAS_MAP`, `RUNTIME_PROFILE_MAP`, `PROVIDER_PRESETS`), effort rendering (`renderEffortForRuntime`), and the agent→model projections (`getAgentToModelMapForProfile`, `formatAgentToModelMapAsTable`). A **genuine leaf**: it imports `node:path` and its own `model-catalog.json` and nothing else, which is what makes it the correct home for anything several unrelated surfaces must agree on. Model *ids* live in `model-catalog.json`, never inline — changing one means regenerating goldens (`UPDATE_GOLDEN`). Also owns the **Anthropic-flavored-model rule** (#3241, ADR-2313): `CLAUDE_AGENT_ALIASES` (the frozen four-alias set `opus`/`sonnet`/`haiku`/`fable`) and `isAnthropicFlavoredModel(model)`, which is true for a bare tier alias or for any `claude-*` id in any provider namespacing (`anthropic/claude-*`, `us.anthropic.claude-*`); no OpenAI/Codex model id contains "claude", so the case-insensitive substring test is safe and exhaustive. It was **moved down here from the Model Resolver Module** rather than shared from there, because the Codex posture check (`agent-install-check`, epic #2313 Phase 2) and the Codex `.toml` sync (`commands`, Phase 3) both need the rule and neither may take the `config-loader` dependency `model-resolver` would have brought; `model-resolver` re-exports it for back-compat and a parity test fails if the two ever fork. _Avoid_: "the model list" (ambiguous between the catalog JSON and the derived enums). See Model Resolver Module, ADR-2313, and ADR-0003. +Leaf module owning the **static** model tables and the closed vocabularies derived from them — the tier/runtime/provider enums (`VALID_TIERS`, `VALID_AGENT_TIERS`, `KNOWN_RUNTIMES`, `KNOWN_PROVIDERS`, `RUNTIMES_WITH_REASONING_EFFORT`, `RUNTIMES_WITH_FAST_MODE`, `ADAPTIVE_TIER_VALUES`), the alias and profile maps (`MODEL_ALIAS_MAP`, `RUNTIME_PROFILE_MAP`, `PROVIDER_PRESETS`), effort rendering (`renderEffortForRuntime`), and the agent→model projections (`getAgentToModelMapForProfile`, `formatAgentToModelMapAsTable`). Also owns the per-model Codex effort capability table (`CODEX_MODEL_EFFORT`, sourced from `model-catalog.json`'s `codexModelEffort` with a `_baseline` fallback for unknown model ids) because Codex declares `supported_reasoning_levels` per model rather than per runtime, so `renderEffortForRuntime('codex', level, model)` resolves against that table. A **genuine leaf**: it imports `node:path` and its own `model-catalog.json` and nothing else, which is what makes it the correct home for anything several unrelated surfaces must agree on. Model *ids* live in `model-catalog.json`, never inline — changing one means regenerating goldens (`UPDATE_GOLDEN`). Also owns the **Anthropic-flavored-model rule** (#3241, ADR-2313): `CLAUDE_AGENT_ALIASES` (the frozen four-alias set `opus`/`sonnet`/`haiku`/`fable`) and `isAnthropicFlavoredModel(model)`, which is true for a bare tier alias or for any `claude-*` id in any provider namespacing (`anthropic/claude-*`, `us.anthropic.claude-*`); no OpenAI/Codex model id contains "claude", so the case-insensitive substring test is safe and exhaustive. It was **moved down here from the Model Resolver Module** rather than shared from there, because the Codex posture check (`agent-install-check`, epic #2313 Phase 2) and the Codex `.toml` sync (`commands`, Phase 3) both need the rule and neither may take the `config-loader` dependency `model-resolver` would have brought; `model-resolver` re-exports it for back-compat and a parity test fails if the two ever fork. _Avoid_: "the model list" (ambiguous between the catalog JSON and the derived enums). See Model Resolver Module, ADR-2313, and ADR-0003. ### Model Resolver Module Module owning model and effort resolution policy: resolves the model, runtime tier, planning granularity, reasoning effort, and fast-mode for a given agent by reading project config and resolving against the model profiles and catalog (`resolveModelInternal`, `resolveModelPolicy`, `resolveTierEntry`, `resolveModelForTier`, `resolveGranularityInternal`, `resolveEffortInternal`, `resolveFastModeInternal`, `resolveEffortForTier`, `nextEffort`, `assertValidGranularityOverride`). Depends only on leaf modules (`config-loader` for `loadConfig`, `configuration` for defaults, `model-profiles` and `model-catalog` for the static tables) — no other core dependency. Extracted from the Core module per ADR-857 rollout phase 2f (#888) — the final core.cts decomposition step; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. **`CLAUDE_AGENT_ALIASES` no longer lives here** — it moved down to the Model Catalog Module (#3241, ADR-2313 Phase 1) so the Agent Install Check and Codex-sync surfaces can consume the alias rule without taking a `config-loader` dependency this module would have dragged with it; it is still **re-exported** from here, so existing importers (`bin/install.js`, `tests/codex-config.test.cjs`) are unaffected and a parity test asserts both modules expose the same set. Source of truth: `gsd-core/bin/lib/model-resolver.cjs` (generated from `src/model-resolver.cts`). diff --git a/bin/install.js b/bin/install.js index f96fe9829..1c2b027b5 100755 --- a/bin/install.js +++ b/bin/install.js @@ -4012,8 +4012,10 @@ function generateCodexAgentToml(agentName, agentContent, modelOverrides = null, // #443 — Unified effort for Codex .toml. Uses the same config-driven precedence chain // as the Claude .md effort injection (resolveInstallTimeEffort), so both runtimes read // from the same effort.agent_overrides / effort.routing_tier_defaults / effort.default - // config source. Codex does not support 'max' → clamped to 'xhigh' by - // gsdRenderEffortForRuntime('codex', ...). + // config source. #3007 — Codex advertises supported_reasoning_levels per model, so the + // pinned model id is passed through and the value is resolved against that model's own + // set: 'max' now passes, 'minimal' clamps up to 'low', and 'ultra' is refused (no key + // emitted) rather than clamped to a fabricated level. // #838 — Do not pin effort when Codex is intentionally inheriting the parent // chat model. A TOML with no `model` but a static `model_reasoning_effort` // creates confusing partial routing: model follows the Codex UI while effort @@ -4023,8 +4025,12 @@ function generateCodexAgentToml(agentName, agentContent, modelOverrides = null, // #3533 (10d): 'inherit' means OMIT the pin — the agent follows the host's // own effort default. Never write the literal. if (_universalEffortCodex !== 'inherit') { - const _renderedEffortCodex = _getGsdEffortCatalog().renderEffortForRuntime('codex', _universalEffortCodex).value; - lines.push(`model_reasoning_effort = ${JSON.stringify(_renderedEffortCodex)}`); + const _renderedEffortCodex = _getGsdEffortCatalog().renderEffortForRuntime('codex', _universalEffortCodex, pinnedModel).value; + // #3007 — 'ultra' is rejected by the model's supported_reasoning_levels and + // renders as null. Omit the key entirely rather than write a literal `null`. + if (_renderedEffortCodex !== null) { + lines.push(`model_reasoning_effort = ${JSON.stringify(_renderedEffortCodex)}`); + } } } diff --git a/docs/CONFIGURATION.md b/docs/CONFIGURATION.md index 146cfce9e..9127e425a 100644 --- a/docs/CONFIGURATION.md +++ b/docs/CONFIGURATION.md @@ -1521,7 +1521,60 @@ minimal < low < medium < high < xhigh < max Effort is rendered per-runtime: `output_config.effort` for Claude (Claude Code subagent `effort` frontmatter / `CLAUDE_CODE_EFFORT_LEVEL` env), `model_reasoning_effort` for Codex (Responses API `reasoning.effort`). -**Cross-provider clamping:** `max` is Anthropic-only — it clamps to `xhigh` on Codex. `minimal` is Codex-only — it clamps to `low` on Claude. +**Cross-provider clamping:** `minimal` is Anthropic-unsupported — it clamps to `low` on Claude. + +**Codex effort is resolved per model, not per runtime (#3007).** Codex advertises a +`supported_reasoning_levels` set on each model and validates against it, so the same universal level +can pass cleanly on one model and clamp on another. GSD therefore renders against the model's own +advertised set: + +| Model | Advertised levels | +|---|---| +| `gpt-5.6-sol` | `low`, `medium`, `high`, `xhigh`, `max`, `ultra` | +| `gpt-5.6-terra` | `low`, `medium`, `high`, `xhigh`, `max` | +| `gpt-5.6-luna` | `low`, `medium`, `high`, `xhigh`, `max` | +| any other / unknown id | `low`, `medium`, `high`, `xhigh`, `max` (family baseline) | + +Today every shipped Codex model advertises the same usable range, so the same effort resolves +identically across `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` — `ultra` is sol's only +differentiator, and GSD rejects it for every model regardless (see below), so no observable output +currently differs by model. The table is per-model, not per-runtime, because Codex declares +capability per model and the sets are free to diverge — the previous single per-runtime assumption +is exactly what went stale and produced this change. + +Three consequences: + +- **`max` reaches Codex.** It is no longer clamped to `xhigh`. Earlier GSD releases described `max` + as Anthropic-only; that was accurate when written and Codex has since added it. If you set `max` + for a Codex agent, your generated `model_reasoning_effort` now says `max` where it previously said + `xhigh`. +- **`minimal` no longer reaches Codex.** No Codex model advertises it, so it clamps up to `low` — + the floor every model does advertise. GSD previously emitted `minimal` verbatim, which Codex + rejects. +- **`ultra` is refused outright**, and is not part of GSD's ladder. See below. + +**Every clamp is now visible.** `resolve-execution` reports the level you asked for alongside the +level actually rendered, so a downgrade is legible instead of silent. These are flat keys in the +same result object as `effort_rendered` — there is no nested `effort` object: + +```json +{ + "effort_rendered": "low", + "effort_requested": "minimal", + "effort_clamped": true, + "effort_clamp_reason": "requested 'minimal' is not in gpt-5.6-luna's advertised reasoning levels; clamped up to its floor, 'low'." +} +``` + +**Why `ultra` is rejected rather than clamped.** Codex's own catalog describes `ultra` as *"Maximum +reasoning with automatic task delegation"* — it is a mode switch, not a louder `max`. At `ultra` +Codex enters proactive multi-agent mode and spawns sub-agents on its own initiative, which would run +underneath GSD's orchestration rather than inside it ([#2167](https://github.com/open-gsd/gsd-core/issues/2167)). +GSD refuses it for every model, including `gpt-5.6-sol`, which does advertise it. This is +deliberately stricter than Codex requires: Codex only applies proactive mode to V2 sessions and +never to spawned sub-agents, but GSD writes effort into generated agent files at install time and +cannot know the session source of a future invocation. Clamping `ultra` down to `max` was rejected +as an option — it would silently discard what you actually asked for. The model-catalog's `reasoning_effort` per-tier hint is a legacy field kept for reference; effort is now config-driven. @@ -1663,18 +1716,23 @@ Use `node gsd-tools.cjs resolve-execution [--effort ] [--fas ```json { - "model": "opus", - "profile": "balanced", - "effort": "xhigh", - "effort_rendered": "xhigh", - "effort_param": "output_config.effort", - "effort_propagation": "frontmatter", - "fast_mode": false, + "model": "opus", + "profile": "balanced", + "effort": "xhigh", + "effort_rendered": "xhigh", + "effort_param": "output_config.effort", + "effort_propagation": "frontmatter", + "effort_requested": "xhigh", + "effort_clamped": false, + "effort_clamp_reason": null, + "fast_mode": false, "fast_mode_supported": false } ``` -`effort_param` tells you which runtime parameter to set. `fast_mode_supported` tells you whether the configured runtime supports per-agent fast_mode propagation. +`effort_param` tells you which runtime parameter to set. `effort_requested` is the level you asked +for (before any clamp); `effort_rendered` is what actually shipped. `effort_clamped` is `true` only +when the two differ, and `effort_clamp_reason` explains why (`null` when unclamped). `fast_mode_supported` tells you whether the configured runtime supports per-agent fast_mode propagation. --- diff --git a/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md b/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md index 3263e878f..a277e5892 100644 --- a/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md +++ b/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md @@ -59,6 +59,69 @@ Recorded as a dated section rather than by editing the amendment above or Decisi **Boundary — unchanged.** [ADR-2313](2313-codex-passive-model-posture.md) still owns the static/install-time channel; [ADR-1239](1239-gsd-embeddable-orchestration-engine.md)'s `effortSurface` amendment still owns the invocation-time argv channel. This amendment changes neither, and changes no runtime behavior at all. +## Amendment (#3007, 2026-08-22) — the Codex capability premise went stale + +**What changed underneath this ADR.** The Context section below records, as fact, that Codex's +ladder is `minimal, low, medium, high, xhigh` and that it *"has `minimal`; no `max`"*, and Decision +item 2 clamps `max → xhigh` on that basis. That was accurate when written. It is no longer: +Codex's `ReasoningEffort` now accepts `none, minimal, low, medium, high, xhigh, max, ultra`, and +capability is declared **per model** via `supported_reasoning_levels`, which Codex validates against +(`validate_spawn_agent_reasoning_effort`) and exposes through `model/list`. + +Verified against Codex's own `codex-rs/models-manager/models.json`: + +| Model | `supported_reasoning_levels` | `default_reasoning_level` | +|---|---|---| +| `gpt-5.6-sol` | low, medium, high, xhigh, max, **ultra** | `low` | +| `gpt-5.6-luna` | low, medium, high, xhigh, max | `medium` | + +**No Codex model advertises `minimal`** — the level this ADR called Codex-only. + +**What this amendment changes.** Decision item 2's per-runtime clamp table is superseded for Codex +by a per-**model** advertised set, with the family baseline `low, medium, high, xhigh, max` for any +model GSD does not know: + +| Universal level | Claude rendering | Codex rendering (was → now) | +|---|---|---| +| `minimal` | `low` (clamped) | `minimal` → **`low` (clamped)** | +| `low`–`xhigh` | unchanged | unchanged | +| `max` | `max` | `xhigh` (clamped) → **`max` (passes)** | +| `ultra` | *not on the ladder* | **rejected, never clamped** | + +Two defects this corrects, both live on `next` before it: + +1. `max` was silently discarded for every Codex model, all of which advertise it. +2. `providerPresets.openai.haiku.low` paired `gpt-5.6-luna` with `reasoning_effort: "minimal"` — a + level luna does not advertise, written into a document Codex itself validates. + +**Clamping becomes visible rather than silent.** `RenderedEffort` gains `requested` / `clamped` / +`reason`, and `resolve-execution` surfaces them as the flat result keys `effort_requested`, +`effort_clamped`, and `effort_clamp_reason` — siblings of the existing `effort_rendered`, not a +nested `effort` object. The old table clamped correctly-but-invisibly, so a user asking for `max` on +Codex had no way to learn they were getting `xhigh` — the failure mode Postel's robustness critique +warns about, and the reason "be liberal" here has to mean "liberal and loud". + +**No model divergence is observable today.** All three shipped Codex models advertise the same +usable set (`low`…`max`), and `ultra` — sol's only differentiator — is rejected for every model +regardless. So today, the same requested level renders identically across `gpt-5.6-sol`, +`gpt-5.6-terra`, and `gpt-5.6-luna`; the per-model table exists because Codex declares capability +per model and the sets are free to diverge, not because a user can currently observe a difference. + +**`ultra` is refused, not laddered.** Codex's catalog describes it as *"Maximum reasoning with +automatic task delegation"*, and at `ultra` Codex enters proactive multi-agent mode +(`effective_multi_agent_mode` → `Proactive`), spawning sub-agents on its own initiative underneath +GSD's orchestration ([#2167](https://github.com/open-gsd/gsd-core/issues/2167)). It is a mode +switch, not a reasoning depth, so it is not added to the universal ladder — which stays +provider-agnostic per Decision item 2 — and it is rejected for every model including `gpt-5.6-sol`, +which advertises it. GSD is deliberately stricter than Codex here: Codex applies proactive mode only +to V2 sessions and never to spawned sub-agents, but GSD writes effort at install time and cannot +know the session source of a future invocation. + +**What this amendment does NOT change.** The universal ladder itself, the cascade, the resolver +precedence, the two channel boundaries above, and Claude's rendering are all untouched. The +per-model table is static and will go stale exactly as this premise did; runtime discovery via +`model/list` is the known escape hatch and was deliberately deferred as the larger step. + ## Context ### Effort control and fast mode in Claude Opus 4.8 diff --git a/docs/how-to/configure-model-profiles.md b/docs/how-to/configure-model-profiles.md index 7c4c0c380..cef73192f 100644 --- a/docs/how-to/configure-model-profiles.md +++ b/docs/how-to/configure-model-profiles.md @@ -228,6 +228,51 @@ Codex UI drives both rather than one following GSD and the other following your > The installer prints a one-time notice when it drops a pin. If you were on a ChatGPT account, this > is the change that stops the 400s — nothing to do. +### Allocating for execution-heavy workflows on Codex + +Execution and verification account for most of the model calls in a long GSD run — planning happens +once per phase, execution happens per plan, and verification runs over everything produced. On +2026-07-30 OpenAI cut GPT-5.6 Luna API pricing by 80% and Terra by 20%, and reduced how many credits +both consume against Codex paid-plan quotas while leaving subscription prices and quota budgets +unchanged. Sol was unchanged. That makes the cheaper models materially cheaper for exactly the +high-volume half of a workflow. + +GSD does not add a routing surface for this — the levers below already express it, and +[#2935](https://github.com/open-gsd/gsd-core/issues/2935) was closed as already-implemented on +precisely that basis. Keep Sol where the reasoning is worth the spend, and put the volume on Terra +or Luna: + +```json +{ + "runtime": "codex", + "model_overrides": { + "gsd-planner": "gpt-5.6-sol", + "gsd-debugger": "gpt-5.6-sol", + "gsd-executor": "gpt-5.6-terra", + "gsd-verifier": "gpt-5.6-luna" + } +} +``` + +Prefer `models` when you want the split by *phase type* rather than by agent — it maps the six phase +types at once and every agent carries a `phaseType`, so it survives the roster changing under you: + +```json +{ + "models": { "planning": "opus", "execution": "sonnet", "verification": "haiku" } +} +``` + +Two things worth knowing before you tune this: + +- **Effort is a separate lever from model, and it is now per-model.** Dropping to Luna does not force + you to drop effort — Luna advertises everything up to `max`. See + [Configuration reference — effort](../CONFIGURATION.md#model-profiles) for the per-model table and + which levels clamp. +- **These are cost/limit tradeoffs, not quality claims.** The 2026-07-30 change was a pricing and + credit-accounting change; it did not alter model quality. Sol remains the strongest model for + planning and hard debugging, which is why it stays there above. + **If you want per-agent model IDs on any non-Claude runtime:** ```json diff --git a/gsd-core/bin/shared/model-catalog.json b/gsd-core/bin/shared/model-catalog.json index 6b987062c..1ecff022b 100644 --- a/gsd-core/bin/shared/model-catalog.json +++ b/gsd-core/bin/shared/model-catalog.json @@ -112,7 +112,7 @@ "openai": { "opus": { "low": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" }, "medium": { "model": "gpt-5.6-sol", "reasoning_effort": "high" }, "high": { "model": "gpt-5.6-sol", "reasoning_effort": "xhigh" } }, "sonnet": { "low": { "model": "gpt-5.6-luna", "reasoning_effort": "low" }, "medium": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" }, "high": { "model": "gpt-5.6-sol", "reasoning_effort": "medium" } }, - "haiku": { "low": { "model": "gpt-5.6-luna", "reasoning_effort": "minimal" }, "medium": { "model": "gpt-5.6-luna", "reasoning_effort": "medium" }, "high": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" } } + "haiku": { "low": { "model": "gpt-5.6-luna", "reasoning_effort": "low" }, "medium": { "model": "gpt-5.6-luna", "reasoning_effort": "medium" }, "high": { "model": "gpt-5.6-terra", "reasoning_effort": "medium" } } }, "google": { "opus": { "low": { "model": "gemini-2.5-flash-lite" }, "medium": { "model": "gemini-3-flash" }, "high": { "model": "gemini-3.1-pro-preview" } }, @@ -130,6 +130,12 @@ "haiku": { "low": null, "medium": null, "high": null } } }, + "codexModelEffort": { + "_baseline": ["low", "medium", "high", "xhigh", "max"], + "gpt-5.6-sol": ["low", "medium", "high", "xhigh", "max", "ultra"], + "gpt-5.6-terra": ["low", "medium", "high", "xhigh", "max"], + "gpt-5.6-luna": ["low", "medium", "high", "xhigh", "max"] + }, "agents": { "gsd-planner": { "golden": "opus", "balanced": "opus", "budget": "sonnet", "phaseType": "planning", "routingTier": "heavy" }, "gsd-roadmapper": { "golden": "opus", "balanced": "sonnet", "budget": "sonnet", "phaseType": "planning", "routingTier": "heavy" }, diff --git a/scripts/mutation-matrix.cjs b/scripts/mutation-matrix.cjs index 86be866fb..acd0cda1b 100644 --- a/scripts/mutation-matrix.cjs +++ b/scripts/mutation-matrix.cjs @@ -92,10 +92,16 @@ function readStdinSync() { // confirmed equivalent mutant is acceptable. // // HOW TO UPDATE: -// 1. Run the per-module Stryker shard locally. -// 2. Note the reported score. -// 3. Set minScore = floor(score) - 1 (never lower than current value). -// 4. Open a PR — the CI gate will enforce the new floor on every future run. +// 1. The per-module Stryker shard CANNOT be run locally: Stryker's command +// runner invokes `node --test` once per mutant (see stryker.config.mjs), +// and this repo hard-blocks local `node --test` via +// .claude/hooks/block-local-node-test.sh. Push the branch instead and +// let CI run the shard for the changed module. +// 2. Read the measured score from the CI shard's output. +// 3. Set minScore = floor(measured) - 1 (never lower than current value) +// and update the matching RATCHET_BASELINE entry in the same diff. +// 4. Open/update the PR — the CI gate will enforce the new floor on every +// future run. /** Long-run target for all modules (ADR-456). */ const TARGET_MUTATION_SCORE = 80; @@ -259,6 +265,37 @@ const COVERED = { ], minScore: 94, }, + // model-catalog: net-new registration by #3007. The module was entirely + // outside mutation scoring (has_work: "false") before this entry, so the + // #3007 per-model Codex effort rewrite (renderEffortForRuntime's + // CODEX_MODEL_EFFORT lookup, the 'ultra' policy rejection, the ladder + // walk-up clamp) had zero mutation coverage. + // + // Same #2790 precedent as planning-inspect above: this shard points at a + // dedicated tests/model-catalog.unit.test.cjs, NOT tests/model-resolver.test.cjs + // — that integration file uses runGsdTools heavily and would hit the same + // 15-minute shard-cap cancellation #2790 documented (a `node --test ` + // invocation is ONE test costing whatever its slowest case costs, re-run + // per mutant). tests/model-catalog.unit.test.cjs is spawn-free, in-process, + // and runs in well under a second. + // + // Measured CI score (GitHub Actions run 32605073352, job 97108869486): + // model-catalog 59.62% → floor 58 (248 killed, 168 survived, 0 timeouts, + // 0 errors; below TARGET_MUTATION_SCORE (80) — ratchet candidate like + // planning-inspect (56): comfortably clears its own floor but has real + // room to grow. Raise as its tests improve, never lower it.) + // Floor follows this file's documented rule, minScore = floor(measured) - 1, + // matching the sibling precedent exactly (57.03 → 56, 76.58 → 75, 95.65 → 94). + // + // The shard completed in 57 seconds — concrete evidence the spawn-free + // unit-file design above worked: the #2790 precedent's 15-minute shard-cap + // cancellations do not apply here, and for comparison the `frontmatter` + // shard in the same run took 9m46s. + 'model-catalog': { + cjs: 'gsd-core/bin/lib/model-catalog.cjs', + tests: ['tests/model-catalog.unit.test.cjs'], + minScore: 58, + }, }; // ── Files that, when changed, invalidate ALL modules ───────────────────────── diff --git a/src/commands.cts b/src/commands.cts index 477f2f3a6..b2055ca7a 100644 --- a/src/commands.cts +++ b/src/commands.cts @@ -619,7 +619,12 @@ function cmdResolveExecution(cwd: string, agentType: string | undefined, raw: bo const fastMode = resolveFastModeInternal(cwd, agentType!, fastModeOpts); const runtime = (config['runtime'] as string) || 'claude'; - const rendered = renderEffortForRuntime(runtime, effort); + // #3007: pass the resolved model so the per-model advertised-effort ceiling + // (CODEX_MODEL_EFFORT) is reachable from this production seam. `model` may + // be a tier alias or a non-Codex id for other runtimes — that's fine and + // must not be special-cased here: advertisedCodexEffort() falls back to the + // family baseline for any id it doesn't recognize. + const rendered = renderEffortForRuntime(runtime, effort, model); const fastModeSupported = RUNTIMES_WITH_FAST_MODE.has(runtime); @@ -675,6 +680,9 @@ function cmdResolveExecution(cwd: string, agentType: string | undefined, raw: bo effort_rendered: rendered.value, effort_param: rendered.param, effort_propagation: rendered.channel, + effort_requested: rendered.requested, + effort_clamped: rendered.clamped, + effort_clamp_reason: rendered.reason, effort_effective: effortEffective, effort_effective_source: effortEffectiveSource, fast_mode: fastMode, @@ -855,8 +863,10 @@ function cmdEffortSync(cwd: string, raw: boolean, opts?: { dryRun?: boolean; con continue; } + // `runtime` is guaranteed 'claude' by the guard above (#3007: only + // codex's 'ultra' rejection can produce a null value). const rendered = renderEffortForRuntime(runtime, universalEffort); - const newEffortValue = rendered.value; + const newEffortValue = rendered.value as string; // eslint-disable-next-line local/no-unbounded-quantifier -- lazy `*?` bounded by the `^---$/m` closing anchor, no nested quantifier, measured linear to 5MB (no-closing-marker adversarial input) const fmMatch = /^---\r?\n([\s\S]*?)^---\r?$/m.exec(content); diff --git a/src/install-effort-resolver.cts b/src/install-effort-resolver.cts index 736362fa4..1e0cb86e7 100644 --- a/src/install-effort-resolver.cts +++ b/src/install-effort-resolver.cts @@ -76,7 +76,11 @@ function _readGsdConfigFile(absPath: string, label: string): Record; - renderEffortForRuntime: (runtime: string, effort: string) => { value: string }; + // #3007: `value` is `string | null` — declaring it `string` here was a + // structural lie that silently defeated TS null-checking for anything + // routed through this seam (a rejected/unrenderable effort level renders + // null, e.g. 'ultra' or an exhausted catalog clamp). + renderEffortForRuntime: (runtime: string, effort: string) => { value: string | null }; EFFORT_MANIFEST_TIER_DEFAULTS: Record; EFFORT_MANIFEST_DEFAULT: string; } @@ -94,7 +98,11 @@ function _getGsdEffortCatalog(): EffortCatalog { // eslint-disable-next-line @typescript-eslint/no-require-imports -- model-catalog.cjs is an export= CommonJS module const { AGENT_DEFAULT_TIERS, renderEffortForRuntime } = require('./model-catalog.cjs') as { AGENT_DEFAULT_TIERS: Record; - renderEffortForRuntime: (runtime: string, effort: string) => { value: string }; + // #3007: `value` is `string | null` — declaring it `string` here was a + // structural lie that silently defeated TS null-checking for anything + // routed through this seam (a rejected/unrenderable effort level renders + // null, e.g. 'ultra' or an exhausted catalog clamp). + renderEffortForRuntime: (runtime: string, effort: string) => { value: string | null }; }; // This module lives in gsd-core/bin/lib/, so the shared manifest is one level diff --git a/src/model-catalog.cts b/src/model-catalog.cts index eb8eb2151..a6c598002 100644 --- a/src/model-catalog.cts +++ b/src/model-catalog.cts @@ -55,6 +55,7 @@ export interface ModelCatalog { runtimeTierDefaults: Record>; providerPresets: Record>>; agents: Record; + codexModelEffort?: Record; } let catalog: ModelCatalog | null = null; @@ -144,6 +145,43 @@ export const RUNTIMES_WITH_REASONING_EFFORT: Set = new Set( export const PROVIDER_PRESETS: Record>> = _catalog.providerPresets ?? {}; +// ─── #3007 — Codex per-model effort capability ─────────────────────────────── +// +// Codex's own `models.json` publishes `supported_reasoning_levels` per model and +// rejects an unsupported level at request time, so GSD must be conservative +// about what it sends: read the ceiling as DATA from the catalog (never +// branch on model id in code) and fall back to the family baseline for any +// model the catalog doesn't know about. +// (b) A malformed catalog entry (e.g. a non-array value like `"gpt-x": 5`) must +// degrade to "ignore that entry", never throw — model-catalog.cjs is required +// across the whole CLI, so one bad JSON value must not kill every command. +// `new Set(5)` would throw at module load; filter to array values first. +export const CODEX_MODEL_EFFORT: Record> = Object.fromEntries( + Object.entries(_catalog.codexModelEffort ?? {}) + .filter(([, levels]) => Array.isArray(levels)) + .map(([model, levels]) => [model, new Set(levels)]) +); +// (a) `??` only catches null/undefined. A malformed `"_baseline": null` still +// produces `new Set(null)` above — an empty Set, which is truthy — so a bare +// `??` fallback would never fire and every level would silently lose its +// advertised set (every effort would render as `value: null`). Guard on +// `.size > 0` so an empty/missing/malformed baseline always falls back to the +// hardcoded floor instead of failing open. +const CODEX_EFFORT_BASELINE: Set = + CODEX_MODEL_EFFORT['_baseline'] && CODEX_MODEL_EFFORT['_baseline'].size > 0 + ? CODEX_MODEL_EFFORT['_baseline'] + : new Set(['low', 'medium', 'high', 'xhigh', 'max']); + +function advertisedCodexEffort(model: string | null | undefined): Set { + if (typeof model !== 'string' || model.length === 0) return CODEX_EFFORT_BASELINE; + return Object.prototype.hasOwnProperty.call(CODEX_MODEL_EFFORT, model) ? CODEX_MODEL_EFFORT[model] : CODEX_EFFORT_BASELINE; +} + +// The full universal effort ladder, low-to-high. Used only to find "the +// nearest advertised level below" when a requested level isn't supported — +// never to invent behaviour for a level that isn't on it at all. +const EFFORT_LADDER: string[] = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + // KNOWN_PROVIDERS excludes 'generic' — it is a sentinel (all null entries) that // forces users to supply model IDs via model_profile_overrides. It is not a // real catalog-backed provider (#49). @@ -229,20 +267,35 @@ export const EFFORT_RENDERING: Record = { }, }, codex: { + // #3007: 'max' and 'minimal' are stale here — Codex's per-model table + // (CODEX_MODEL_EFFORT above) is now the source of truth for what a given + // model actually advertises, and every model in the family baseline DOES + // advertise 'max' (no model advertises 'minimal'). This runtime-level + // spec is kept in sync with the family baseline so the two tables can + // never disagree; renderEffortForRuntime layers the per-model ceiling + // (and the 'ultra' policy rejection) on top of it. + // KEEP IN SYNC with EFFORT_ARGV.codex below — same family baseline, two + // channels (install-time vs invocation-time); they must never diverge. param: 'model_reasoning_effort', channel: 'api', - supported: new Set(['minimal', 'low', 'medium', 'high', 'xhigh']), + supported: new Set(['low', 'medium', 'high', 'xhigh', 'max']), clamp(level: string): string { - if (level === 'max') return 'xhigh'; + if (level === 'minimal') return 'low'; return level; }, }, }; export interface RenderedEffort { - value: string; + value: string | null; param: string | null; channel: string | null; + /** The level as originally asked for. */ + requested?: string; + /** true ONLY when value !== requested. */ + clamped?: boolean; + /** Why, when clamped or rejected; null otherwise. */ + reason?: string | null; } // ─── Invocation-time (argv) effort rendering ───────────────────────────────── @@ -281,10 +334,14 @@ export const EFFORT_ARGV: Record = { }, // First-party Codex docs: `model_reasoning_effort` is a config-only key with no // dedicated flag, so the generic `-c key=value` override is the only argv route. + // #3007: KEEP IN SYNC with EFFORT_RENDERING.codex above — this table must match + // the family baseline exactly (no 'minimal', 'max' passes through unclamped), + // otherwise the argv channel and the install-time channel disagree about the + // same runtime's capability. codex: { render: (level: string): string[] => ['-c', `model_reasoning_effort=${level}`], - supported: new Set(['minimal', 'low', 'medium', 'high', 'xhigh']), - clamp: (level: string): string => (level === 'max' ? 'xhigh' : level), + supported: new Set(['low', 'medium', 'high', 'xhigh', 'max']), + clamp: (level: string): string => (level === 'minimal' ? 'low' : level), }, }; @@ -323,23 +380,108 @@ export function renderEffortArgv( /** * Render a universal effort string for a specific runtime. + * + * `model` (#3007) is consulted ONLY for codex: Codex's per-model + * `supported_reasoning_levels` means the same universal level can be a clean + * pass-through on one model and a clamp (or, for 'ultra', an outright + * rejection) on another. Every other runtime ignores the third argument + * entirely — passing a model id to claude changes nothing. */ -export function renderEffortForRuntime(runtime: string, universalEffort: string): RenderedEffort { +export function renderEffortForRuntime(runtime: string, universalEffort: string, model?: string | null): RenderedEffort { // #3533 (10d): 'inherit' is not a wire level on ANY runtime — it means // "omit the key / pass no argument and follow the session/host default". // Renderers must never emit it as a literal; null param/channel tells - // resolve-execution consumers there is no propagation. + // resolve-execution consumers there is no propagation. Never measured + // against any supported set. if (universalEffort === 'inherit') { - return { value: 'inherit', param: null, channel: null }; + return { value: 'inherit', param: null, channel: null, requested: 'inherit', clamped: false, reason: null }; } const spec = EFFORT_RENDERING[runtime]; if (!spec) { - return { value: universalEffort, param: null, channel: null }; + return { value: universalEffort, param: null, channel: null, requested: universalEffort, clamped: false, reason: null }; } + + if (runtime === 'codex') { + // #2167 — 'ultra' turns on Codex's automatic task delegation, which would + // let Codex spawn agents underneath GSD's own orchestration. GSD rejects + // it unconditionally as a POLICY call, never as a capability clamp — this + // holds even for a model (e.g. gpt-5.6-sol) that DOES advertise 'ultra', + // so it is never softened down to 'max'. + if (universalEffort === 'ultra') { + return { + value: null, + param: null, + channel: null, + requested: 'ultra', + clamped: false, + reason: "'ultra' turns on Codex's automatic task delegation, which would let Codex spawn agents underneath GSD's own orchestration (#2167); GSD rejects it regardless of what the model advertises.", + }; + } + + const allowed = advertisedCodexEffort(model); + if (allowed.has(universalEffort)) { + return { value: universalEffort, param: spec.param, channel: spec.channel, requested: universalEffort, clamped: false, reason: null }; + } + + const idx = EFFORT_LADDER.indexOf(universalEffort); + if (idx === -1) { + // Not on the ladder at all (e.g. 'MAX') — preserve prior behaviour: + // fall through to the runtime-level clamp rather than inventing new + // handling for input the ladder doesn't recognize. + return { value: spec.clamp(universalEffort), param: spec.param, channel: spec.channel, requested: universalEffort, clamped: false, reason: null }; + } + // Every model's advertised set is a contiguous run up to 'max' (or 'ultra' + // for sol, already handled above), so the only unsupported level in + // practice is 'minimal' — below every model's floor. Walk UP the ladder + // to the nearest level the model actually advertises (its floor): there + // is nothing below 'minimal' to fall back to. + // Walking UP is safe today only because every advertised set floors at + // 'low' — a future model whose floor is, say, 'high' would silently turn + // a requested 'low' into 'high': MORE reasoning and MORE cost than asked + // for, with no error. `clamped`/`reason` below is what makes that + // escalation visible to a caller instead of a silent cost surprise, which + // is why those fields are not optional decoration. + for (let i = idx + 1; i < EFFORT_LADDER.length; i++) { + const candidate = EFFORT_LADDER[i]; + // 'ultra' is never a valid clamp target: it would re-enter, by the back + // door, the delegation mode the #2167 rejection above exists to keep + // out. A clamp may never produce a value that a direct request for + // that same value would have refused. + if (candidate === 'ultra') { + continue; + } + if (allowed.has(candidate)) { + return { + value: candidate, + param: spec.param, + channel: spec.channel, + requested: universalEffort, + clamped: true, + reason: `requested '${universalEffort}' is not in ${model ? `${model}'s` : "the codex family baseline's"} advertised reasoning levels; clamped up to its floor, '${candidate}'.`, + }; + } + } + // No advertised level at or above the request either (shouldn't happen + // given today's catalog data, but never throw): reject rather than emit + // an unsupported level. + return { + value: null, + param: null, + channel: null, + requested: universalEffort, + clamped: false, + reason: `requested '${universalEffort}' is not in ${model ? `${model}'s` : "the codex family baseline's"} advertised reasoning levels, and no advertised level is available either.`, + }; + } + + const value = spec.clamp(universalEffort); return { - value: spec.clamp(universalEffort), + value, param: spec.param, channel: spec.channel, + requested: universalEffort, + clamped: value !== universalEffort, + reason: value !== universalEffort ? `requested '${universalEffort}' clamped to '${value}' for ${runtime}.` : null, }; } diff --git a/src/runtime-artifact-conversion.cts b/src/runtime-artifact-conversion.cts index 3161b65b1..f3e412583 100644 --- a/src/runtime-artifact-conversion.cts +++ b/src/runtime-artifact-conversion.cts @@ -3503,7 +3503,14 @@ function applyAgentFrontmatterExtensions( // effort key, so skipping injection is the whole job. if (universalEffort !== 'inherit') { const renderedEffort = _getGsdEffortCatalog().renderEffortForRuntime(runtime, universalEffort).value; - result = injectEffortFrontmatter(result, renderedEffort); + // #3007: `value` is `string | null` — a rejected/unrenderable level (e.g. + // 'ultra', or a catalog with no advertised level at or above the request) + // renders null. Same posture as the 'inherit' case above: omit the key + // entirely rather than writing a literal `effort: null`, so the host + // falls back to its own default instead of failing to parse. + if (renderedEffort !== null) { + result = injectEffortFrontmatter(result, renderedEffort); + } } const disallowedTools = READONLY_AGENT_DISALLOWED_TOOLS[agentName]; if (disallowedTools) result = injectDisallowedToolsFrontmatter(result, disallowedTools); diff --git a/tests/effort-surface-axis.test.cjs b/tests/effort-surface-axis.test.cjs index bfa51b087..ce19f5cba 100644 --- a/tests/effort-surface-axis.test.cjs +++ b/tests/effort-surface-axis.test.cjs @@ -297,10 +297,14 @@ describe('#2481 renderEffortArgv — per-host syntax and clamping', () => { ); }); + // #3007: corrected — Codex gained 'max' (declared per-model), and no Codex + // model advertises 'minimal', so 'max' now passes through and 'minimal' + // clamps to 'low' instead. test('clamps the provider-unique tail levels', () => { - // claude has no `minimal`; codex has no `max`. + // claude has no `minimal`; codex has no `minimal` either (clamps to 'low'). assert.deepEqual(renderEffortArgv('claude', 'minimal', 'argv').argv, ['--effort', 'low']); - assert.deepEqual(renderEffortArgv('codex', 'max', 'argv').argv, ['-c', 'model_reasoning_effort=xhigh']); + assert.deepEqual(renderEffortArgv('codex', 'max', 'argv').argv, ['-c', 'model_reasoning_effort=max']); + assert.deepEqual(renderEffortArgv('codex', 'minimal', 'argv').argv, ['-c', 'model_reasoning_effort=low']); }); test('emits nothing when the surface is not argv', () => { diff --git a/tests/install-runtime-artifacts.test.cjs b/tests/install-runtime-artifacts.test.cjs index 01f3f4bb6..4ba9b6039 100644 --- a/tests/install-runtime-artifacts.test.cjs +++ b/tests/install-runtime-artifacts.test.cjs @@ -3834,7 +3834,10 @@ describe('#443 Config-driven: effort.agent_overrides drives install-time effort' `gsd-planner.toml should have model_reasoning_effort = "low" from config override\nActual:\n${tomlContent.slice(0, 500)}`); }); - test('Codex .toml clamps effort max → xhigh when agent_overrides.gsd-planner=max', () => { + // #3007: corrected — Codex's gpt-5.6-sol advertises 'max' in its own + // supported_reasoning_levels, so install-time rendering now passes 'max' + // through instead of clamping it to 'xhigh'. + test('Codex .toml renders effort max → max when agent_overrides.gsd-planner=max', () => { const projectDir = path.dirname(codexHome); // Overwrite config with max override. #3241: include an explicit // model_overrides pin (D1 removed the resolver-only auto-embed). @@ -3860,11 +3863,11 @@ describe('#443 Config-driven: effort.agent_overrides drives install-time effort' ); assert.match(tomlContent, /^model\s*=\s*"gpt-5.6-sol"$/m, `gsd-planner.toml should pin Codex model when runtime:"codex" is configured\nActual:\n${tomlContent.slice(0, 500)}`); - // Codex does not support 'max' → clamped to 'xhigh' - assert.match(tomlContent, /^model_reasoning_effort\s*=\s*"xhigh"$/m, - `gsd-planner.toml should clamp max → xhigh for Codex\nActual:\n${tomlContent.slice(0, 500)}`); - assert.doesNotMatch(tomlContent, /model_reasoning_effort\s*=\s*"max"/, - 'Codex .toml must never contain model_reasoning_effort = "max"'); + // gpt-5.6-sol advertises 'max' -> renders through unchanged, no clamp. + assert.match(tomlContent, /^model_reasoning_effort\s*=\s*"max"$/m, + `gsd-planner.toml should render max for Codex on a model that advertises it\nActual:\n${tomlContent.slice(0, 500)}`); + assert.doesNotMatch(tomlContent, /model_reasoning_effort\s*=\s*"xhigh"/, + 'Codex .toml must not clamp "max" down to "xhigh" for a model that advertises max'); }); }); diff --git a/tests/model-catalog.unit.test.cjs b/tests/model-catalog.unit.test.cjs new file mode 100644 index 000000000..f02aa49f5 --- /dev/null +++ b/tests/model-catalog.unit.test.cjs @@ -0,0 +1,381 @@ +'use strict'; + +/** + * FAST, IN-PROCESS mutation-testing surface for `model-catalog.cjs` (#3007). + * + * Root cause this file exists to fix: `model-catalog.cjs` was entirely + * outside Stryker's covered-module list, so the #3007 per-model Codex effort + * rewrite (`renderEffortForRuntime`, `CODEX_MODEL_EFFORT`, the 'ultra' + * rejection, the ladder walk-up) had zero mutation coverage. Following the + * #2790 precedent (see `tests/planning-inspect.unit.test.cjs`), this is a + * dedicated, spawn-free, in-process unit file rather than pointing the shard + * at an integration test — `tests/model-resolver.test.cjs` uses + * `runGsdTools` heavily and would hit the same 15-minute shard-cap cancel + * that #2790 documented (one `node --test ` invocation costs whatever + * its slowest case costs, per mutant). + * + * NEVER spawn a child process here — no `runGsdTools`, `spawnSync`, + * `execFileSync`, or CLI invocation of any kind, and no filesystem writes. + * Every case below requires the BUILT `.cjs` artifact directly and calls its + * exports in-process. + * + * Every value asserted below was verified by requiring the built lib + * directly and inspecting the real returned object — never guessed from + * reading the source alone (CLAUDE.md "verify assertions by executing, not + * retyping"). + */ + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); + +const catalog = require('../gsd-core/bin/lib/model-catalog.cjs'); + +const { + VALID_TIERS, + VALID_AGENT_TIERS, + KNOWN_RUNTIMES, + KNOWN_PROVIDERS, + RUNTIMES_WITH_REASONING_EFFORT, + RUNTIMES_WITH_FAST_MODE, + ADAPTIVE_TIER_VALUES, + CODEX_MODEL_EFFORT, + MODEL_ALIAS_MAP, + PROVIDER_PRESETS, + isAnthropicFlavoredModel, + getAgentToModelMapForProfile, + formatAgentToModelMapAsTable, + renderEffortArgv, + renderEffortForRuntime, + nextTier, + mergeEffortTierDefaults, +} = catalog; + +describe('model-catalog: exported enums/maps', () => { + test('VALID_TIERS is opus/sonnet/haiku plus inherit', () => { + assert.deepEqual(new Set(VALID_TIERS), new Set(['opus', 'sonnet', 'haiku', 'inherit'])); + }); + + test('VALID_AGENT_TIERS is light/standard/heavy', () => { + assert.deepEqual(new Set(VALID_AGENT_TIERS), new Set(['light', 'standard', 'heavy'])); + }); + + test('ADAPTIVE_TIER_VALUES excludes inherit', () => { + assert.deepEqual(new Set(ADAPTIVE_TIER_VALUES), new Set(['opus', 'sonnet', 'haiku'])); + assert.equal(ADAPTIVE_TIER_VALUES.has('inherit'), false); + }); + + test('KNOWN_RUNTIMES includes codex and claude', () => { + assert.equal(KNOWN_RUNTIMES.has('codex'), true); + assert.equal(KNOWN_RUNTIMES.has('claude'), true); + }); + + test('KNOWN_PROVIDERS excludes generic sentinel', () => { + assert.equal(KNOWN_PROVIDERS.has('generic'), false); + assert.equal(KNOWN_PROVIDERS.has('anthropic'), true); + assert.equal(KNOWN_PROVIDERS.has('openai'), true); + }); + + test('RUNTIMES_WITH_REASONING_EFFORT is codex-only', () => { + assert.deepEqual(new Set(RUNTIMES_WITH_REASONING_EFFORT), new Set(['codex'])); + }); + + test('RUNTIMES_WITH_FAST_MODE is api-only', () => { + assert.deepEqual(new Set(RUNTIMES_WITH_FAST_MODE), new Set(['api'])); + }); + + test('CODEX_MODEL_EFFORT has a baseline and per-model sets, sol includes ultra', () => { + assert.ok(CODEX_MODEL_EFFORT['_baseline'] instanceof Set); + assert.equal(CODEX_MODEL_EFFORT['_baseline'].has('max'), true); + assert.equal(CODEX_MODEL_EFFORT['gpt-5.6-sol'].has('ultra'), true); + assert.equal(CODEX_MODEL_EFFORT['gpt-5.6-terra'].has('ultra'), false); + assert.equal(CODEX_MODEL_EFFORT['gpt-5.6-luna'].has('ultra'), false); + }); + + test('MODEL_ALIAS_MAP maps opus/sonnet/haiku to claude model ids', () => { + assert.equal(MODEL_ALIAS_MAP.opus, 'claude-opus-4-8'); + assert.equal(MODEL_ALIAS_MAP.sonnet, 'claude-sonnet-5'); + assert.equal(MODEL_ALIAS_MAP.haiku, 'claude-haiku-4-5'); + }); + + test('PROVIDER_PRESETS is a non-empty object keyed by provider name', () => { + assert.equal(typeof PROVIDER_PRESETS, 'object'); + assert.ok(Object.keys(PROVIDER_PRESETS).length > 0); + assert.ok('anthropic' in PROVIDER_PRESETS); + }); +}); + +describe('model-catalog: isAnthropicFlavoredModel', () => { + test('bare tier aliases are flavored', () => { + assert.equal(isAnthropicFlavoredModel('opus'), true); + assert.equal(isAnthropicFlavoredModel('sonnet'), true); + assert.equal(isAnthropicFlavoredModel('haiku'), true); + assert.equal(isAnthropicFlavoredModel('fable'), true); + }); + + test('claude-* ids are flavored in every provider namespacing', () => { + assert.equal(isAnthropicFlavoredModel('claude-opus-4-8'), true); + assert.equal(isAnthropicFlavoredModel('anthropic/claude-opus-4-8'), true); + assert.equal(isAnthropicFlavoredModel('us.anthropic.claude-opus-4-8'), true); + }); + + test('negative case: gpt-* id is not flavored', () => { + assert.equal(isAnthropicFlavoredModel('gpt-5.6-sol'), false); + }); + + test('non-string input is not flavored', () => { + assert.equal(isAnthropicFlavoredModel(123), false); + assert.equal(isAnthropicFlavoredModel(null), false); + assert.equal(isAnthropicFlavoredModel(undefined), false); + }); +}); + +describe('model-catalog: getAgentToModelMapForProfile / formatAgentToModelMapAsTable', () => { + const EXPECTED_PLANNER_TIER = { quality: 'opus', balanced: 'opus', budget: 'sonnet', adaptive: 'opus' }; + const EXPECTED_MAPPER_TIER = { quality: 'sonnet', balanced: 'haiku', budget: 'haiku', adaptive: 'haiku' }; + + for (const profile of ['quality', 'balanced', 'budget', 'adaptive']) { + test(`profile '${profile}' returns the expected tier for known agents`, () => { + const map = getAgentToModelMapForProfile(profile); + assert.equal(map['gsd-planner'], EXPECTED_PLANNER_TIER[profile]); + assert.equal(map['gsd-codebase-mapper'], EXPECTED_MAPPER_TIER[profile]); + assert.ok(Object.keys(map).length > 0); + for (const value of Object.values(map)) { + assert.equal(typeof value, 'string'); + assert.ok(value.length > 0); + } + }); + } + + test("profile 'inherit' maps every agent to the literal string 'inherit'", () => { + const map = getAgentToModelMapForProfile('inherit'); + for (const value of Object.values(map)) { + assert.equal(value, 'inherit'); + } + }); + + test('invalid profile falls back to balanced', () => { + const balanced = getAgentToModelMapForProfile('balanced'); + const invalid = getAgentToModelMapForProfile('totally-bogus-profile'); + assert.deepEqual(invalid, balanced); + }); + + test('formatAgentToModelMapAsTable pads columns and renders a header/separator', () => { + const out = formatAgentToModelMapAsTable({ agentA: 'model-x', b: 'model-y-longer' }); + const lines = out.split('\n'); + assert.equal(lines[0].includes('Agent'), true); + assert.equal(lines[0].includes('Model'), true); + assert.equal(lines[1].includes('┼'), true); + assert.equal(lines[2].includes('agentA'), true); + assert.equal(lines[2].includes('model-x'), true); + assert.equal(lines[3].includes('model-y-longer'), true); + }); +}); + +describe('model-catalog: renderEffortForRuntime — codex per-model', () => { + test("sol's 'max' passes through unclamped", () => { + const r = renderEffortForRuntime('codex', 'max', 'gpt-5.6-sol'); + assert.equal(r.value, 'max'); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + assert.equal(r.param, 'model_reasoning_effort'); + assert.equal(r.channel, 'api'); + }); + + test("'minimal' clamps up to 'low' with clamped:true and a non-empty reason (sol/terra/luna)", () => { + for (const model of ['gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.6-luna']) { + const r = renderEffortForRuntime('codex', 'minimal', model); + assert.equal(r.value, 'low'); + assert.equal(r.clamped, true); + assert.equal(typeof r.reason, 'string'); + assert.ok(r.reason.length > 0); + assert.ok(r.reason.includes(model)); + } + }); + + test("'ultra' is rejected outright even for sol which advertises it (value:null)", () => { + const r = renderEffortForRuntime('codex', 'ultra', 'gpt-5.6-sol'); + assert.equal(r.value, null); + assert.equal(r.param, null); + assert.equal(r.channel, null); + assert.equal(r.clamped, false); + assert.ok(r.reason.includes('#2167')); + }); + + test('low/medium/high/xhigh pass through unclamped with reason:null', () => { + for (const level of ['low', 'medium', 'high', 'xhigh']) { + const r = renderEffortForRuntime('codex', level, 'gpt-5.6-terra'); + assert.equal(r.value, level); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + } + }); + + test('unknown model falls back to the family baseline', () => { + const r = renderEffortForRuntime('codex', 'max', 'gpt-9-never-heard-of-it'); + assert.equal(r.value, 'max'); + assert.equal(r.clamped, false); + }); + + test('omitted / null / empty-string model all fall back to the family baseline', () => { + const omitted = renderEffortForRuntime('codex', 'minimal'); + const nullModel = renderEffortForRuntime('codex', 'minimal', null); + const emptyModel = renderEffortForRuntime('codex', 'minimal', ''); + for (const r of [omitted, nullModel, emptyModel]) { + assert.equal(r.value, 'low'); + assert.equal(r.clamped, true); + assert.ok(r.reason.includes('codex family baseline')); + } + }); + + test("off-ladder input ('MAX') falls through to the runtime-level clamp unchanged", () => { + const r = renderEffortForRuntime('codex', 'MAX', 'gpt-5.6-terra'); + assert.equal(r.value, 'MAX'); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + }); +}); + +describe('model-catalog: renderEffortForRuntime — cross-runtime', () => { + test("'inherit' passes through on every known runtime", () => { + for (const runtime of [...KNOWN_RUNTIMES, 'totally-unknown-runtime']) { + const r = renderEffortForRuntime(runtime, 'inherit'); + assert.equal(r.value, 'inherit'); + assert.equal(r.param, null); + assert.equal(r.channel, null); + assert.equal(r.clamped, false); + assert.equal(r.reason, null); + } + }); + + test('unknown runtime passes the requested value through unchanged', () => { + const r = renderEffortForRuntime('totally-unknown-runtime', 'high'); + assert.equal(r.value, 'high'); + assert.equal(r.param, null); + assert.equal(r.channel, null); + assert.equal(r.clamped, false); + }); + + test('claude passes high through unchanged and clamps minimal to low', () => { + const high = renderEffortForRuntime('claude', 'high'); + assert.equal(high.value, 'high'); + assert.equal(high.clamped, false); + assert.equal(high.reason, null); + + const minimal = renderEffortForRuntime('claude', 'minimal'); + assert.equal(minimal.value, 'low'); + assert.equal(minimal.clamped, true); + assert.ok(minimal.reason.includes('claude')); + }); +}); + +describe('model-catalog: renderEffortArgv', () => { + test("effortSurface gate: 'none'/undefined produce empty, 'argv' produces an argument", () => { + assert.deepEqual(renderEffortArgv('codex', 'max', 'none'), { argv: [], value: null, host: 'codex' }); + assert.deepEqual(renderEffortArgv('codex', 'max', undefined), { argv: [], value: null, host: 'codex' }); + const r = renderEffortArgv('codex', 'max', 'argv'); + assert.deepEqual(r.argv, ['-c', 'model_reasoning_effort=max']); + assert.equal(r.value, 'max'); + }); + + test("codex 'max' passes through, 'minimal' clamps to 'low'", () => { + const max = renderEffortArgv('codex', 'max', 'argv'); + assert.deepEqual(max.argv, ['-c', 'model_reasoning_effort=max']); + assert.equal(max.value, 'max'); + + const minimal = renderEffortArgv('codex', 'minimal', 'argv'); + assert.deepEqual(minimal.argv, ['-c', 'model_reasoning_effort=low']); + assert.equal(minimal.value, 'low'); + }); + + test('unknown host produces empty result, never throws', () => { + const r = renderEffortArgv('totally-bogus-host', 'high', 'argv'); + assert.deepEqual(r, { argv: [], value: null, host: 'totally-bogus-host' }); + }); + + test('prototype-chain host names (__proto__, constructor, toString) return empty, never throw', () => { + for (const host of ['__proto__', 'constructor', 'toString']) { + const r = renderEffortArgv(host, 'high', 'argv'); + assert.deepEqual(r.argv, []); + assert.equal(r.value, null); + assert.equal(r.host, host); + } + }); +}); + +describe('model-catalog: nextTier', () => { + // Probed directly against the built module: light -> standard -> heavy, + // and heavy SATURATES at 'heavy' rather than wrapping back to 'light'. + test("advances light -> standard and standard -> heavy", () => { + assert.equal(nextTier('light'), 'standard'); + assert.equal(nextTier('standard'), 'heavy'); + }); + + test("saturates at the top tier: 'heavy' stays 'heavy'", () => { + assert.equal(nextTier('heavy'), 'heavy'); + }); + + test('unknown/garbage tier returns null', () => { + assert.equal(nextTier('bogus'), null); + assert.equal(nextTier('LIGHT'), null); // case-sensitive, not in the order array + }); + + test('empty string, null, undefined, and non-string input all return null', () => { + assert.equal(nextTier(''), null); + assert.equal(nextTier(null), null); + assert.equal(nextTier(undefined), null); + assert.equal(nextTier(123), null); + }); +}); + +describe('model-catalog: mergeEffortTierDefaults (#3531)', () => { + const manifest = { opus: 'sonnet', sonnet: 'haiku', haiku: 'opus' }; + const isValid = (v) => typeof v === 'string' && ['opus', 'sonnet', 'haiku'].includes(v); + + test('absent/empty override leaves the manifest values untouched', () => { + assert.deepEqual(mergeEffortTierDefaults(manifest, undefined, isValid), manifest); + assert.deepEqual(mergeEffortTierDefaults(manifest, {}, isValid), manifest); + }); + + test('partial override changes only the named tier; other tiers keep their built-in values', () => { + const merged = mergeEffortTierDefaults(manifest, { sonnet: 'opus' }, isValid); + assert.equal(merged.sonnet, 'opus'); + assert.equal(merged.opus, manifest.opus); + assert.equal(merged.haiku, manifest.haiku); + }); + + test('an invalid override value for a tier is ignored; the built-in for that tier is retained', () => { + const merged = mergeEffortTierDefaults(manifest, { sonnet: 'bogus' }, isValid); + assert.equal(merged.sonnet, manifest.sonnet); + }); + + test('an override key not present in the manifest is still merged in (isValid gates values, not tier names)', () => { + const merged = mergeEffortTierDefaults(manifest, { newtier: 'opus' }, isValid); + assert.equal(merged.newtier, 'opus'); + assert.equal(merged.opus, manifest.opus); + assert.equal(merged.sonnet, manifest.sonnet); + assert.equal(merged.haiku, manifest.haiku); + }); + + test('__proto__/constructor/prototype override keys are skipped (house pollution guard)', () => { + const merged = mergeEffortTierDefaults(manifest, { __proto__: 'opus', constructor: 'sonnet', prototype: 'haiku' }, isValid); + assert.deepEqual(merged, manifest); + assert.equal(Object.prototype.hasOwnProperty.call(merged, '__proto__'), false); + assert.equal(Object.prototype.hasOwnProperty.call(merged, 'constructor'), false); + }); + + test('null, non-object, and array config all leave the manifest untouched', () => { + assert.deepEqual(mergeEffortTierDefaults(manifest, null, isValid), manifest); + assert.deepEqual(mergeEffortTierDefaults(manifest, 'bogus', isValid), manifest); + assert.deepEqual(mergeEffortTierDefaults(manifest, ['x'], isValid), manifest); + }); + + test('a falsy/missing manifest merges over an empty base object', () => { + assert.deepEqual(mergeEffortTierDefaults(null, { sonnet: 'opus' }, isValid), { sonnet: 'opus' }); + }); + + test('the function is pure: it never mutates the manifest it is given', () => { + const original = { ...manifest }; + mergeEffortTierDefaults(manifest, { sonnet: 'opus' }, isValid); + assert.deepEqual(manifest, original); + }); +}); diff --git a/tests/model-resolver.test.cjs b/tests/model-resolver.test.cjs index 23096f411..b6e184232 100644 --- a/tests/model-resolver.test.cjs +++ b/tests/model-resolver.test.cjs @@ -349,7 +349,10 @@ describe('#3533 effort inherit: expressible at every layer, never a wire level', } // Concrete levels unchanged. assert.strictEqual(renderEffortForRuntime('claude', 'minimal').value, 'low'); - assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'xhigh'); + // #3007: corrected — Codex DOES advertise 'max' (per-model table), so + // ADR-443's "Codex has no max" premise went stale and this pinned the + // defect (clamping 'max' down to 'xhigh') instead of the fix. + assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'max'); assert.strictEqual(renderEffortForRuntime('claude', 'xhigh').value, 'xhigh'); }); }); @@ -1641,9 +1644,13 @@ const { // Anthropic: output_config.effort — https://docs.anthropic.com (Claude API) // OpenAI: model_reasoning_effort — https://platform.openai.com/docs (Codex) // ───────────────────────────────────────────────────────────────────────────── +// #3007: corrected — Codex's own models.json now advertises 'max' (and 'ultra', +// which is policy-rejected separately, #2167) but no Codex model advertises +// 'minimal'. The old enum here ('minimal'..'xhigh', no 'max') encoded ADR-443's +// stale premise, not the real API. const PROVIDER_EFFORT_ENUMS = { claude: new Set(['low', 'medium', 'high', 'xhigh', 'max']), - codex: new Set(['minimal', 'low', 'medium', 'high', 'xhigh']), + codex: new Set(['low', 'medium', 'high', 'xhigh', 'max']), }; // Helper: write config.json into a temp project @@ -1672,8 +1679,10 @@ describe('#443 integration (a): cross-provider validity invariant', () => { }); // Documented clamps must hold exactly - test("render('codex','max').value === 'xhigh' (max is Anthropic-only)", () => { - assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'xhigh'); + // #3007: corrected — Codex gained 'max' (declared per-model via + // supported_reasoning_levels); 'max' is no longer Anthropic-only. + test("render('codex','max').value === 'max' (Codex now advertises max)", () => { + assert.strictEqual(renderEffortForRuntime('codex', 'max').value, 'max'); }); test("render('claude','minimal').value === 'low' (minimal is Codex-only)", () => { @@ -2653,9 +2662,11 @@ describe('#443 resolveEffortForTier escalation', () => { // ─── Rendering / clamping ────────────────────────────────────────────────────── describe('#443 renderEffortForRuntime', () => { - test('codex: "max" clamps to "xhigh"', () => { + // #3007: corrected — Codex gained 'max' (per-model supported_reasoning_levels); + // it no longer clamps to 'xhigh'. + test('codex: "max" passes through as "max"', () => { const r = renderEffortForRuntime('codex', 'max'); - assert.strictEqual(r.value, 'xhigh'); + assert.strictEqual(r.value, 'max'); assert.strictEqual(r.param, 'model_reasoning_effort'); }); @@ -2666,8 +2677,10 @@ describe('#443 renderEffortForRuntime', () => { assert.strictEqual(renderEffortForRuntime('codex', 'xhigh').value, 'xhigh'); }); - test('codex: "minimal" passthrough', () => { - assert.strictEqual(renderEffortForRuntime('codex', 'minimal').value, 'minimal'); + // #3007: corrected — no Codex model advertises 'minimal'; it now clamps up + // to the family floor, 'low', instead of passing through. + test('codex: "minimal" clamps to "low"', () => { + assert.strictEqual(renderEffortForRuntime('codex', 'minimal').value, 'low'); }); test('claude: "minimal" clamps to "low"', () => { @@ -2728,7 +2741,13 @@ describe('#443 resolve-execution CLI command', () => { assert.ok('profile' in output, 'should have profile field'); }); - test('codex runtime -> effort_param=model_reasoning_effort, max clamps to xhigh, fast_mode_supported=false', () => { + // NOTE: effort.default: 'max' never reaches the renderer for gsd-planner here — + // gsd-planner is a heavy/opus-tier agent, and its routing-tier default outranks + // effort.default in resolution precedence, so the resolved level is 'xhigh' before + // the renderer ever sees 'max'. effort_clamped=false and effort_requested='xhigh' + // prove this is precedence, not the #3007 clamp — do not "correct" this back to + // expecting 'max'. + test('codex runtime -> effort_param=model_reasoning_effort, tier default outranks effort.default, fast_mode_supported=false', () => { writeConfig(tmpDir, { runtime: 'codex', effort: { default: 'max' }, @@ -2738,6 +2757,29 @@ describe('#443 resolve-execution CLI command', () => { const output = JSON.parse(result.output); assert.strictEqual(output.effort_param, 'model_reasoning_effort'); assert.strictEqual(output.effort_rendered, 'xhigh'); + assert.strictEqual(output.effort_clamped, false); + assert.strictEqual(output.effort_requested, 'xhigh'); + // fast_mode_supported: codex does not support fast mode via subagent + assert.strictEqual(output.fast_mode_supported, false); + }); + + // #3007: Codex gained 'max' (per-model supported_reasoning_levels), so 'max' now + // renders through unchanged instead of clamping to 'xhigh'. agent_overrides is used + // here (not effort.default) because it outranks the routing-tier default, which is + // what actually lets 'max' reach the renderer end-to-end. + test('codex runtime -> max survives to the wire via agent_overrides, fast_mode_supported=false', () => { + writeConfig(tmpDir, { + runtime: 'codex', + effort: { default: 'max', agent_overrides: { 'gsd-planner': 'max' } }, + }); + const result = runGsdTools(['resolve-execution', 'gsd-planner'], tmpDir, { HOME: tmpDir }); + assert.ok(result.success, `Command failed: ${result.error}`); + const output = JSON.parse(result.output); + assert.strictEqual(output.effort_param, 'model_reasoning_effort'); + assert.strictEqual(output.effort, 'max'); + assert.strictEqual(output.effort_rendered, 'max'); + assert.strictEqual(output.effort_requested, 'max'); + assert.strictEqual(output.effort_clamped, false); // fast_mode_supported: codex does not support fast mode via subagent assert.strictEqual(output.fast_mode_supported, false); }); @@ -5928,3 +5970,353 @@ describe('issue #2612: partial override merge for new Group A runtimes', () => { }); }); } + +// ─── #3007: Codex effort capability is per-model ───────────────────────────── +// +// Ground truth (Codex's models.json): gpt-5.6-sol advertises +// low/medium/high/xhigh/max/ultra; gpt-5.6-luna and gpt-5.6-terra advertise +// low/medium/high/xhigh/max (no ultra); no Codex model advertises 'minimal'. +// An unknown/omitted model id falls back to the family baseline +// (low/medium/high/xhigh/max). + +const CODEX_MODEL_EFFORT_SETS = { + 'gpt-5.6-sol': new Set(['low', 'medium', 'high', 'xhigh', 'max', 'ultra']), + 'gpt-5.6-luna': new Set(['low', 'medium', 'high', 'xhigh', 'max']), + 'gpt-5.6-terra': new Set(['low', 'medium', 'high', 'xhigh', 'max']), +}; +const CODEX_FAMILY_BASELINE = new Set(['low', 'medium', 'high', 'xhigh', 'max']); +function advertisedCodexEfforts(model) { + return CODEX_MODEL_EFFORT_SETS[model] || CODEX_FAMILY_BASELINE; +} + +describe('#3007 — Codex effort capability is per-model, and every clamp is visible', () => { + test('max survives for a model that advertises it', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max', 'gpt-5.6-sol'); + assert.strictEqual(r.value, 'max'); + assert.strictEqual(r.clamped, false); + assert.strictEqual(r.reason, null); + }); + + test('max survives on luna, not only sol', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max', 'gpt-5.6-luna'); + assert.strictEqual(r.value, 'max'); + }); + + test('a level below the ceiling is not reported as clamped', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'xhigh', 'gpt-5.6-sol'); + assert.strictEqual(r.value, 'xhigh'); + assert.strictEqual(r.clamped, false); + }); + + test('ultra is rejected, never clamped to max', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'ultra', 'gpt-5.6-sol'); + assert.strictEqual(r.value, null); + assert.ok(typeof r.reason === 'string' && /deleg/i.test(r.reason), `reason should mention delegation: ${r.reason}`); + }); + + test('minimal clamps to low on a model that floors at low', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal', 'gpt-5.6-luna'); + assert.strictEqual(r.value, 'low'); + assert.strictEqual(r.clamped, true); + assert.ok(typeof r.reason === 'string' && r.reason.length > 0); + }); + + test('minimal clamps on sol too', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal', 'gpt-5.6-sol'); + assert.strictEqual(r.value, 'low'); + assert.strictEqual(r.clamped, true); + }); + + test('an unknown model id gets the family baseline, not a clamp', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max', 'gpt-9.9-unreleased'); + assert.strictEqual(r.value, 'max'); + assert.strictEqual(r.clamped, false); + }); + + test('an unknown model id still floors at low', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal', 'gpt-9.9-unreleased'); + assert.strictEqual(r.value, 'low'); + assert.strictEqual(r.clamped, true); + }); + + test('the model-less form resolves against the family baseline', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'max'); + assert.strictEqual(r.value, 'max'); + }); + + test('the model-less form still floors at low', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('codex', 'minimal'); + assert.strictEqual(r.value, 'low'); + }); + + test('the two-argument signature keeps working for every level', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const UNCHANGED = new Set(['low', 'medium', 'high', 'xhigh']); + for (const level of ['minimal', 'low', 'medium', 'high', 'xhigh', 'max']) { + let r; + assert.doesNotThrow(() => { r = renderEffortForRuntime('codex', level); }); + if (UNCHANGED.has(level)) { + assert.strictEqual(r.value, level, `level ${level} should pass through unchanged`); + } + } + }); + + test('claude rendering is untouched by the codex table', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + assert.strictEqual(renderEffortForRuntime('claude', 'max').value, 'max'); + assert.strictEqual(renderEffortForRuntime('claude', 'minimal').value, 'low'); + // passing a codex model id as the 3rd arg to claude must change nothing + const withModel = renderEffortForRuntime('claude', 'max', 'gpt-5.6-sol'); + assert.strictEqual(withModel.value, 'max'); + }); + + test('inherit is not measured against any supported set', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + for (const runtime of ['claude', 'codex', 'something-unknown']) { + for (const model of [undefined, 'gpt-5.6-sol']) { + const r = renderEffortForRuntime(runtime, 'inherit', model); + assert.deepStrictEqual( + { value: r.value, param: r.param, channel: r.channel }, + { value: 'inherit', param: null, channel: null }, + `runtime=${runtime} model=${JSON.stringify(model)}`, + ); + // #3007's new requested/clamped/reason fields must be honest on this + // path too: 'inherit' is never a clamp target, so it can never be + // reported as clamped, and it echoes itself back as `requested`. + assert.deepStrictEqual( + { requested: r.requested, clamped: r.clamped, reason: r.reason }, + { requested: 'inherit', clamped: false, reason: null }, + `runtime=${runtime} model=${JSON.stringify(model)}`, + ); + } + } + }); + + test('an undeclared runtime still renders nothing', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const r = renderEffortForRuntime('something-unknown', 'high'); + assert.strictEqual(r.param, null); + assert.strictEqual(r.channel, null); + // #3007's new fields must also be honest for a host with no spec at all: + // the value passes straight through, unclamped, and there is nothing to + // explain about it. + assert.strictEqual(r.value, 'high'); + assert.strictEqual(r.requested, 'high'); + assert.strictEqual(r.clamped, false); + assert.strictEqual(r.reason, null); + }); + + test('a bare tier alias is not treated as a model id', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + let r; + assert.doesNotThrow(() => { r = renderEffortForRuntime('codex', 'max', 'opus'); }); + assert.strictEqual(r.value, 'max'); + }); + + test('an anthropic-flavored id is not an effort-capability error', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + let r; + assert.doesNotThrow(() => { r = renderEffortForRuntime('codex', 'max', 'claude-opus-4-8'); }); + assert.strictEqual(r.value, 'max'); + }); + + test('an empty model argument behaves exactly like omitting it', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + for (const level of ['max', 'minimal']) { + const omitted = renderEffortForRuntime('codex', level); + for (const emptyish of ['', null, undefined]) { + assert.deepStrictEqual( + renderEffortForRuntime('codex', level, emptyish), + omitted, + `level=${level} emptyish=${JSON.stringify(emptyish)}`, + ); + } + } + }); + + test('effort matching is case-sensitive, as the ladder always was', () => { + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + // 'MAX' is not 'max' — whatever the renderer does with an unrecognized + // level, it must not silently treat it as a clean pass-through of 'max'. + // 'MAX' is off the EFFORT_LADDER entirely, so #3007's per-model path + // must fall through to the runtime-level clamp verbatim rather than + // inventing new handling — and it must report that verbatim pass-through + // as NOT clamped, with `requested` echoing the exact (unrecognized) input. + // Under a fully reverted #3007, `clamped`/`requested` do not exist on the + // returned object at all, so this fails there too. + const r = renderEffortForRuntime('codex', 'MAX', 'gpt-5.6-sol'); + assert.notStrictEqual(r.value, 'max'); + assert.strictEqual(r.value, 'MAX'); + assert.strictEqual(r.clamped, false); + assert.strictEqual(r.requested, 'MAX'); + }); + + test('a clamp never lands on ultra, even for a model that only advertises it', () => { + // The clamp-up loop in renderEffortForRuntime walks EFFORT_LADDER + // upward from the requested level looking for the model's floor. Today's + // catalog can't actually exercise the 'ultra'-as-clamp-target path — every + // model advertises 'max', so the allowed.has() fast path always returns + // first. This test guards a latent path, not a currently-reachable one: + // do not delete it as redundant just because it never fails today. The + // invariant it protects is general — for EVERY model and EVERY ladder + // level, a clamp must never produce 'ultra', because that would re-enter + // by the back door the delegation mode the #2167 rejection exists to + // keep out. + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const models = [undefined, 'gpt-5.6-sol', 'gpt-5.6-luna', 'gpt-5.6-terra', 'gpt-9.9-unreleased']; + const levels = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + for (const model of models) { + for (const level of levels) { + const r = renderEffortForRuntime('codex', level, model); + assert.notStrictEqual(r.value, 'ultra', `model=${JSON.stringify(model)} level=${level}`); + } + } + }); +}); + +// ─── #3007 PARITY: known-defect gauntlet ────────────────────────────────────── + +test('every catalog preset ships an effort its own model supports', () => { + const catalogPath = path.join(__dirname, '..', 'gsd-core', 'bin', 'shared', 'model-catalog.json'); + const catalog = JSON.parse(fs.readFileSync(catalogPath, 'utf8')); + const offenders = []; + + const checkEntry = (entryPath, entry) => { + if (!entry || typeof entry !== 'object') return; + const model = entry.model; + const effort = entry.reasoning_effort; + if (typeof model !== 'string' || !model.startsWith('gpt-') || typeof effort !== 'string') return; + const allowed = advertisedCodexEfforts(model); + if (!allowed.has(effort)) { + offenders.push(`${entryPath}: model=${model} reasoning_effort=${effort} (not in {${[...allowed].join(', ')}})`); + } + }; + + const codexDefaults = catalog.runtimeTierDefaults && catalog.runtimeTierDefaults.codex; + if (codexDefaults) { + for (const [tier, entry] of Object.entries(codexDefaults)) { + checkEntry(`runtimeTierDefaults.codex.${tier}`, entry); + } + } + + const openaiPresets = catalog.providerPresets && catalog.providerPresets.openai; + if (openaiPresets) { + for (const [tier, profiles] of Object.entries(openaiPresets)) { + if (!profiles || typeof profiles !== 'object') continue; + for (const [profile, entry] of Object.entries(profiles)) { + checkEntry(`providerPresets.openai.${tier}.${profile}`, entry); + } + } + } + + assert.deepStrictEqual(offenders, [], `offending presets (model does not advertise the assigned effort):\n${offenders.join('\n')}`); +}); + +// ─── #3007 PARITY: argv channel must agree with the render-for-runtime channel ─ +// +// `renderEffortArgv('codex', ...)` (invocation-time, `-c model_reasoning_effort=`) +// and `renderEffortForRuntime('codex', ...)` (install-time / api channel) each +// read their own EFFORT_ARGV.codex / EFFORT_RENDERING.codex tables. Those two +// tables must describe the SAME capability, or a user gets a different answer +// depending on which code path asked — the repo's documented "generative fix +// divergence" class (two surfaces reading one fact that can drift apart). + +describe('#3007 PARITY: argv channel agrees with renderEffortForRuntime for codex', () => { + test('renderEffortArgv and renderEffortForRuntime never disagree across the ladder', () => { + const { renderEffortArgv, renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const LADDER = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + for (const level of LADDER) { + const argvResult = renderEffortArgv('codex', level, 'argv'); + const runtimeResult = renderEffortForRuntime('codex', level); + if (level === 'ultra') { + // 'ultra' is Codex's automatic-delegation switch (#2167): the + // install-time channel rejects it outright (value: null, policy + // reason). The argv channel has no concept of that policy rejection + // — it simply isn't in EFFORT_ARGV.codex's supported set, so it also + // degrades to a `null` value. Both channels landing on `null` here + // is the explicit agreement contract for this level; it is not a + // case the parity check can skip. + assert.strictEqual(argvResult.value, null, `argv channel must also refuse ultra: got ${argvResult.value}`); + assert.strictEqual(runtimeResult.value, null, `runtime channel must refuse ultra: got ${runtimeResult.value}`); + continue; + } + assert.strictEqual( + argvResult.value, + runtimeResult.value, + `argv/runtime channels disagree for level=${level}: argv=${argvResult.value} runtime=${runtimeResult.value}`, + ); + } + }); +}); + +// ─── #3007 PROPERTY: rendered codex effort is always within the model's ceiling ─ + +describe('#3007 PROPERTY: renderEffortForRuntime never renders a level the model does not advertise', () => { + test('every (model, level) pair is exhaustively checked, not sampled', () => { + // fc.constantFrom over MODELS x LADDER with numRuns: 200 is very likely to + // hit all 4 x 7 = 28 pairs but is not GUARANTEED to. This deterministic + // nested loop covers the full cross-product with certainty; it is kept + // alongside the fast-check property below (not instead of it) because the + // repo requires a property test for a closed-vocabulary contract, and the + // property still adds shrinking value on failure. + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const MODELS = ['gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.6-luna', 'gpt-9.9-unreleased-model']; + const LADDER = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + for (const model of MODELS) { + for (const level of LADDER) { + const allowed = advertisedCodexEfforts(model); + const r = renderEffortForRuntime('codex', level, model); + if (r.value === null) { + assert.strictEqual(level, 'ultra', `only ultra may be rejected: model=${model} level=${level}`); + continue; + } + assert.ok(allowed.has(r.value), `rendered value not in model's advertised set: model=${model} level=${level} value=${r.value}`); + if (r.value === level) { + assert.strictEqual(r.clamped, false, `pass-through reported as clamped: model=${model} level=${level}`); + } else { + assert.strictEqual(r.clamped, true, `changed value not reported as clamped: model=${model} level=${level} value=${r.value}`); + } + } + } + }); + + test('for every (model, level) pair, the outcome is pass-through, clamp, or reject', () => { + const fc = require('./helpers/fast-check-setup.cjs'); + const { renderEffortForRuntime } = require('../gsd-core/bin/lib/model-catalog.cjs'); + const MODELS = ['gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.6-luna', 'gpt-9.9-unreleased-model']; + const LADDER = ['minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra']; + fc.assert( + fc.property(fc.constantFrom(...MODELS), fc.constantFrom(...LADDER), (model, level) => { + const allowed = advertisedCodexEfforts(model); + const r = renderEffortForRuntime('codex', level, model); + if (r.value === null) { + // Rejection is a POLICY outcome, not a capability one, so it is not + // predicted by the advertised set. `ultra` is refused even on sol, + // which does advertise it: GSD is deliberately stricter than Codex + // because ultra turns on automatic task delegation (#2167). + // Reserving null for exactly `ultra` is what keeps that a decision + // rather than a side effect — any OTHER null is a bug. + assert.strictEqual(level, 'ultra', `only ultra may be rejected: model=${model} level=${level}`); + return; + } + assert.ok(allowed.has(r.value), `rendered value not in model's advertised set: model=${model} level=${level} value=${r.value}`); + if (r.value === level) { + assert.strictEqual(r.clamped, false, `pass-through reported as clamped: model=${model} level=${level}`); + } else { + assert.strictEqual(r.clamped, true, `changed value not reported as clamped: model=${model} level=${level} value=${r.value}`); + } + }), + { seed: 3007, numRuns: 200 }, + ); + }); +}); diff --git a/tests/mutation-matrix-ratchet.test.cjs b/tests/mutation-matrix-ratchet.test.cjs index 497cb3041..525ce5e2a 100644 --- a/tests/mutation-matrix-ratchet.test.cjs +++ b/tests/mutation-matrix-ratchet.test.cjs @@ -202,6 +202,7 @@ const RATCHET_BASELINE = { 'planning-inspect': 56, // CI run 32392791843: 57.03% (unit shard); ratchet candidate vs TARGET 80 'plan-document': 75, // CI run 32392791843: 76.58% (unit shard) 'planning-command-router': 94, // CI run 32392791843: 95.65% (unit shard); already exceeds TARGET 80 + 'model-catalog': 58, // #3007: measured 59.62% in CI (248 killed / 168 survived); floor(59.62)-1 }; describe('mutation-matrix ratchet: floor equality enforcement', () => {