Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted.
26 KiB
ADR 443: Unified cross-provider effort controls and fast-mode-aware routing
- Status: Accepted (2026-08-19)
- Date: 2026-05-28
- Tracking issue: #443
Why this is still Proposed (audited 2026-07-17)
The audit confirmed the cross-provider resolver/renderer/CLI machinery genuinely shipped: resolveEffortInternal, resolveEffortForTier, renderEffortForRuntime, RUNTIMES_WITH_FAST_MODE, and cmdResolveExecution (src/model-resolver.cts:534,654; src/commands.cts) implement the cascade and clamping exactly as Decision items 1–3, 5, and 6 describe, and static install-time propagation is real and end-to-end tested — tests/install-runtime-artifacts.test.cjs's describe('#443 Claude install: effort: injected into frontmatter') runs the actual install() function and reads the resulting agent .md files off disk, confirming msd-planner gets effort: xhigh, msd-codebase-mapper gets effort: low, and msd-executor gets effort: high. That test predates the QA audit below (landed 2026-05-29 in the original #443 PR, commit 5ca646f01), so the "resolver-only, nothing reaches the runtime" framing of the original flavor-text problem this ADR set out to fix is fixed for the static path.
The blocker. Decision item 1's cascade names an "(1) orchestrator invocation override" as the highest-precedence layer, and Decision item 6 adds a dynamic escalation path ("effort steps up the ladder on a failed attempt"). Both exist only as CLI-callable resolver code — resolveEffortInternal's invocation-override step (src/model-resolver.cts:535) and resolveEffortForTier's attempt-based escalation (src/model-resolver.cts:654) — exercised solely by unit/CLI tests. Nothing in the shipped orchestration actually calls them: a search across every file in msd-core/workflows/*.md and agents/*.md for resolve-execution or CLAUDE_CODE_EFFORT_LEVEL returns zero hits; the only workflow-level mentions of "effort" are documentation of the config keys in settings-advanced.md's confirmation table. The only propagation channel actually wired into a real MSD flow is the static one (config → install() → frontmatter, baked once at install time) — the ADR's own decided design promises more than that, and the more-than-static-baking part has no consumer. Separately, the repo's own dated QA test-architecture audit (docs/issueevidence/1192-adr-test-audit-2026-06-13.md, produced under issue #1192, closed COMPLETED) rated ADR-443 "partial ... end-to-end effort propagation untested" and named it in its action plan ("Strengthen ... ADR-443 end-to-end effort propagation," line 220); that action item was never converted into a tracked follow-up issue, and no commit since 2026-06-13 addresses it. That audit's blanket "untested" framing overstates the gap — the static path is tested — but the underlying signal (a decided mechanism with no live caller) is real and independently confirmed here.
Unblock condition. Either (a) wire the orchestrator-invocation-override and attempt-based-escalation paths into an actual MSD workflow or agent dispatch (so resolveEffortForTier's escalation and resolveEffortInternal's invocation-override step have a real caller outside src/commands.cts's CLI surface and tests), and add a test exercising that live path the way tests/install-runtime-artifacts.test.cjs exercises the static one; or (b) if the ADR's intended scope is in fact limited to static install-time propagation, amend Decision items 1 and 6 to say so explicitly and close out audit issue #1192's action-plan item 18 with a note pointing at the shipped install-wiring tests. Either is a maintainer call this file records but does not make.
Amendment (2026-07-21): path (a) chosen; audit corrected (#2481)
The maintainer call above has been made: path (a). Raised by #2475 — reviewer CLIs invoked as subprocesses by the review workflow silently inherit whatever reasoning effort sits in the user's own global CLI config, because no shipped orchestration resolves effort at invocation time. That is this ADR's blocker surfacing as a user-visible defect, not a new problem.
Two deferrals recorded above are closed by this change — resolved, not re-tracked:
- Audit issue #1192's action-plan item 18 ("Strengthen … ADR-443 end-to-end effort propagation") — which the blocker text notes "was never converted into a tracked follow-up issue" — is satisfied by this change, which supplies the end-to-end propagation and the live-path tests it asked for. It is closed out, not converted into another follow-up.
- The choice between (a) and (b) — which this file previously recorded without making — is resolved as (a). Scope is not limited to static install-time propagation. Choosing the path is not the same as completing it; see the status table below for what remains.
How path (a) is being satisfied. The consumer is defined through the Host-Integration Interface rather than by hard-coding per-CLI effort syntax into a workflow: ADR-1239 gains an effortSurface axis declaring how each host accepts reasoning effort (argv | none), so a universal effort value resolved by this ADR's cascade is rendered per host through the negotiated descriptor. EFFORT_RENDERING (src/model-catalog.cts) — whose channel vocabulary is frontmatter | api, both install-time — collapses into that descriptor data rather than growing a parallel per-runtime table. Its callers today are exactly the two channels this ADR already ships: the static install-time renderer (bin/install.js, via src/install-effort-resolver.cts) and the manual query resolve-execution / effort-sync CLI surface (src/commands.cts). No workflow or agent dispatch calls it — which is the blocker restated in terms of the renderer rather than the resolver.
What this change actually delivers — and what it does not. Path (a) names two mechanisms needing a live caller. This change delivers neither of them; it delivers a third thing the blocker did not anticipate, and the audit of the other two turns out to have been stale.
| Path (a) mechanism | Status |
|---|---|
Decision item 1 — resolveEffortInternal's invocation-override step (--effort) |
Still no live caller. No workflow, reference, or agent passes --effort to resolve-execution. This change does not add one. |
Decision item 6 — resolveEffortForTier's attempt-based escalation |
Already satisfied — by #2296, not by this change. msd-core/references/execute-phase-quota-recovery.md calls resolve-execution msd-executor --attempt "${QUOTA_ATTEMPT:-1}" --failure-class quota-exceeded, and that reference is @-included into msd-core/workflows/execute-phase.md, so it executes as part of the live workflow. |
| New here: the resolved effort cascade reaches a spawned host as an invocation argument | Delivered. msd-core/workflows/review.md calls resolve-execution … --host <id> per reviewer and appends the rendered argument, gated by the host's negotiated effortSurface. |
The blocker's grep was stale in two ways. It searched only msd-core/workflows/*.md and agents/*.md; references/*.md is @-included into workflows and is therefore just as live — that is where #2296's escalation caller sits. And the blocker text was written 2026-07-17, three days before #2296 landed (455ad49ae, 2026-07-20), so its "zero hits" finding was correct on the day and has since been overtaken.
This ADR therefore remains Proposed. Decision item 6's condition is met (by #2296); Decision item 1's is not. The corpus rule for ratifying a stale Proposed requires the decided mechanism to demonstrably exist in the tree, and the invocation-override step still has no caller outside the CLI surface and tests. Status flips when item 1 gains a live caller and a test exercises that path through a workflow rather than through msd-tools directly.
Boundary. #2313 owns the static/install-time effort channel for Codex (model_reasoning_effort in generated ~/.codex/agents/<agent>.toml, plus a sync path) and explicitly places orchestrator effort-override drift outside its scope. That is the static channel this ADR already ships; the work above is the invocation-time channel it does not.
Amendment (2026-08-19): Decision item 1 scoped to the operator surface; this ADR is ratified (#2475)
Recorded as a dated section rather than by editing the amendment above or Decision item 1 itself, since ADRs here are append-only.
The remaining maintainer call has been made: unblock path (b), applied to Decision item 1 only. The 2026-07-21 amendment above left this ADR Proposed on a single condition — that item 1's orchestrator invocation override (precedence step 1) gain a caller in shipped orchestration. It does not gain one. The scope is corrected instead.
Why the original (b) wording does not fit, and what replaces it. The unblock condition offered (b) as "if the ADR's intended scope is in fact limited to static install-time propagation, amend Decision items 1 and 6 to say so explicitly." That sentence is now false on both counts and must not be adopted verbatim: item 6 has a live orchestration caller (#2296, msd-core/references/execute-phase-quota-recovery.md, @-included into execute-phase.md), and #2481 delivered a live invocation-time channel in which the resolved cascade reaches a spawned host as an argv argument, gated by ADR-1239's effortSurface. This ADR's scope is emphatically not limited to install-time propagation. What is narrowed is one precedence step, for a reason (b) did not anticipate.
What Decision item 1's invocation override is. An operator-facing surface, not an orchestration channel. msd_run query resolve-execution <agent> --effort <level> exists so a human — or an agent diagnosing routing — can ask "what would this agent run at if effort were X?" without mutating project or home configuration. That is its whole job, and it does it today. Shipped orchestration deliberately does not pass --effort: workflows resolve effort through precedence steps 2–5 (effort.agent_overrides → effort.routing_tier_defaults → effort.default → built-in), which is the configured, reviewable, per-project path. An override baked into a workflow would be an unconfigurable constant overriding the user's own configuration at the top of the cascade — the inverse of what step 1 is for.
Why no consumer was invented to satisfy the gate. The unblock condition is a proxy for design maturity; wiring a caller purely to flip a status is the textbook case of a measure becoming a target, producing a number that looks better while the design gets worse. No user has asked for a per-invocation effort override, and #2475's actual complaint — reviewer CLIs silently inheriting a global effort default — is closed by the cascade→argv path, not by step 1. Adding an unrequested knob to ratify an ADR would also invert the very property ADR-1239's effortSurface amendment claims for itself ("a wired axis, not a declared-but-unconsumed one"): a consumed-but-unrequested one is no better.
--effort is supported and is NOT deprecated. Scoping it out of orchestration says nothing about its standing as a CLI flag. It ships, it is documented, it is covered by tests, and callers may rely on it. This paragraph exists so that a future reader — or a dead-code sweep — does not read "no orchestration caller" as "unused, remove it".
What is deliberately not built. There is no per-run effort override for review lanes ("make this review cheap"). A user wanting that edits effort.*. Should a real need appear, it is a new enhancement to be judged on its merits — not a deferred obligation of this ADR, and not a request that was refused.
Status: Proposed → Accepted. The corpus rule for ratifying a stale Proposed requires the decided mechanism to demonstrably exist in the tree. With item 1's decided scope narrowed to the operator surface, every Decision item now clears that bar: items 1–3 and 5 in the resolver, renderer, and resolve-execution output; item 4 in the static install-time propagation the 2026-07-17 audit confirmed end-to-end; item 6 in #2296's escalation caller; and the cascade's delivery to a spawned host in #2481's review-lane wiring.
The invariant this ratification rests on, and its guard. Ratification is conditional on item 1 staying operator-only. tests/effort-surface-axis.test.cjs walks msd-core/workflows, msd-core/references, agents, and commands and fails if any of them invokes resolve-execution … --effort in either accepted shape (--effort <level> or --effort=<level>). That guard was widened in this change: it previously matched only the space-separated form, so --effort=low would have passed it silently. Wiring an orchestration caller therefore requires amending this ADR first — the guard is the enforcement of this decision, not a snapshot of a gap awaiting closure.
Boundary — unchanged. ADR-2313 still owns the static/install-time channel; ADR-1239's effortSurface amendment still owns the invocation-time argv channel. This amendment changes neither, and changes no runtime behavior at all.
Amendment (#3007, 2026-08-22) — the Codex capability premise went stale
What changed underneath this ADR. The Context section below records, as fact, that Codex's
ladder is minimal, low, medium, high, xhigh and that it "has minimal; no max", and Decision
item 2 clamps max → xhigh on that basis. That was accurate when written. It is no longer:
Codex's ReasoningEffort now accepts none, minimal, low, medium, high, xhigh, max, ultra, and
capability is declared per model via supported_reasoning_levels, which Codex validates against
(validate_spawn_agent_reasoning_effort) and exposes through model/list.
Verified against Codex's own codex-rs/models-manager/models.json:
| Model | supported_reasoning_levels |
default_reasoning_level |
|---|---|---|
gpt-5.6-sol |
low, medium, high, xhigh, max, ultra | low |
gpt-5.6-luna |
low, medium, high, xhigh, max | medium |
No Codex model advertises minimal — the level this ADR called Codex-only.
What this amendment changes. Decision item 2's per-runtime clamp table is superseded for Codex
by a per-model advertised set, with the family baseline low, medium, high, xhigh, max for any
model MSD does not know:
| Universal level | Claude rendering | Codex rendering (was → now) |
|---|---|---|
minimal |
low (clamped) |
minimal → low (clamped) |
low–xhigh |
unchanged | unchanged |
max |
max |
xhigh (clamped) → max (passes) |
ultra |
not on the ladder | rejected, never clamped |
Two defects this corrects, both live on next before it:
maxwas silently discarded for every Codex model, all of which advertise it.providerPresets.openai.haiku.lowpairedgpt-5.6-lunawithreasoning_effort: "minimal"— a level luna does not advertise, written into a document Codex itself validates.
Clamping becomes visible rather than silent. RenderedEffort gains requested / clamped /
reason, and resolve-execution surfaces them as the flat result keys effort_requested,
effort_clamped, and effort_clamp_reason — siblings of the existing effort_rendered, not a
nested effort object. The old table clamped correctly-but-invisibly, so a user asking for max on
Codex had no way to learn they were getting xhigh — the failure mode Postel's robustness critique
warns about, and the reason "be liberal" here has to mean "liberal and loud".
No model divergence is observable today. All three shipped Codex models advertise the same
usable set (low…max), and ultra — sol's only differentiator — is rejected for every model
regardless. So today, the same requested level renders identically across gpt-5.6-sol,
gpt-5.6-terra, and gpt-5.6-luna; the per-model table exists because Codex declares capability
per model and the sets are free to diverge, not because a user can currently observe a difference.
ultra is refused, not laddered. Codex's catalog describes it as "Maximum reasoning with
automatic task delegation", and at ultra Codex enters proactive multi-agent mode
(effective_multi_agent_mode → Proactive), spawning sub-agents on its own initiative underneath
MSD's orchestration (#2167). It is a mode
switch, not a reasoning depth, so it is not added to the universal ladder — which stays
provider-agnostic per Decision item 2 — and it is rejected for every model including gpt-5.6-sol,
which advertises it. MSD is deliberately stricter than Codex here: Codex applies proactive mode only
to V2 sessions and never to spawned sub-agents, but MSD writes effort at install time and cannot
know the session source of a future invocation.
What this amendment does NOT change. The universal ladder itself, the cascade, the resolver
precedence, the two channel boundaries above, and Claude's rendering are all untouched. The
per-model table is static and will go stale exactly as this premise did; runtime discovery via
model/list is the known escape hatch and was deliberately deferred as the larger step.
Context
Effort control and fast mode in Claude Opus 4.8
Claude Opus 4.8 introduced two orthogonal execution controls relevant to MSD's agent orchestration:
-
Effort control — API request field
output_config.effort(string enum). Anthropic levels:low,medium,high,xhigh,max; Opus 4.8 defaults tohigh. In Claude Code it is exposed as/effort, the--effortCLI flag, theCLAUDE_CODE_EFFORT_LEVELenv var, theeffortLevelsettings.json key (acceptslow/medium/high/xhigh;maxis session-only), and — critically for orchestration — a per-subagenteffortfrontmatter key (shipped per anthropics/claude-code issue #31536, CLOSED/COMPLETED). -
Fast mode — API request field
speed(standard|fast);fastenables high output-tokens-per-second inference. Pricing for Opus 4.8 fast mode is $10/$50 per MTok in/out vs $5/$25 standard. In Claude Code it is the interactive/fasttoggle ONLY — there is no settings.json key, env var, or subagent-frontmatter mechanism to enable fast mode for a spawned subagent.
MSD already routes WHICH model runs a task (routingTier heavy/standard/light, model_profile quality/balanced/budget/adaptive/inherit, model_overrides, and dynamic_routing escalation). It had no way to control HOW HARD the model reasons or WHICH speed tier it uses.
The "flavor text" problem: issue #2517
Issue #2517 added resolveReasoningEffortInternal and made query resolve-model emit a reasoning_effort field derived from the Codex runtime's per-tier catalog values (model-catalog.json runtimeTierDefaults.codex.*.reasoning_effort). However, a codebase audit found that NO orchestrator, workflow, or agent ever consumes that emitted field — it is never passed to an actual Codex invocation. The resolver computed a value and a test asserted the computed JSON, but the value reached no runtime. The feature was inert ("flavor text, no code"): asserting a resolver's return value is not the same as asserting the control reaches the model.
Cross-provider effort enum mismatch
The two providers' effort enums are NOT identical:
- Anthropic/Claude (Opus 4.8):
low,medium,high,xhigh,max(hasmax; nominimal) - OpenAI/Codex (
model_reasoning_effort/ Responses APIreasoning.effort; SDKReasoningEffortranksnone=0,minimal=1,low=2,medium=3,high=4,xhigh=5):minimal,low,medium,high,xhigh(hasminimal; nomax)
Common core: low, medium, high, xhigh.
Decision
-
Introduce a single universal
effortconfig knob (and an orthogonalfast_modeknob) that compose with model selection rather than replace it. Resolution precedence mirrors the existing model cascade: (1) orchestrator invocation override, (2)effort.agent_overrides[agent], (3)effort.routing_tier_defaults[routingTier], (4)effort.default, (5) built-in defaulthigh. Same cascade forfast_modewith built-in defaultfalse. Invalid enum values at any level are ignored and fall through (mirrors theVALID_TIERSgate inresolveModelInternal) so a typo never silently breaks resolution. -
The universal effort value is provider-agnostic; a per-runtime renderer maps it to each runtime's wire parameter, clamping the genuinely-unique tail levels:
- Claude / API: param
output_config.effort(Claude Code: subagenteffortfrontmatter /CLAUDE_CODE_EFFORT_LEVELenv).minimalclamps tolow(Claude has nominimal);low/medium/high/xhigh/maxpass through. - Codex: param
model_reasoning_effort(Responses APIreasoning.effort).maxclamps toxhigh(Codex has nomax);minimal/low/medium/high/xhighpass through.
Universal level Claude rendering Codex rendering minimallow(clamped)minimallowlowlowmediummediummediumhigh(default)highhighxhighxhighxhighmaxmaxxhigh(clamped) - Claude / API: param
-
Fold the inert
reasoning_effortoutput into this unified model.query resolve-modelis preserved for back-compat; a NEWquery resolve-executionis the superset that emits:model,effort(universal), the per-runtime rendered effort, the wire param name, the propagation channel,fast_mode, andfast_mode_supported. Each config key ships help text naming exactly which runtime field/invocation it drives. -
Make effort actually reach the runtime (close the flavor-text gap). Claude is first-class: the resolved effort propagates to spawned subagents via the
effortfrontmatter /CLAUDE_CODE_EFFORT_LEVELenv. Tests assert end-to-end propagation, not just resolver return values. -
Fast mode honesty: because Claude Code has no per-subagent fast-mode mechanism,
fast_modeis resolved and surfaced (with afast_mode_supportedflag,falsefor the claude runtime's subagents) but is NEVER emitted as a fake frontmatter key — doing so would be a silent no-op. It propagates only where the runtime supports it (APIspeed:"fast"). -
Dynamic-routing integration is additive: a new effort-escalation path (effort steps up the ladder on a failed attempt BEFORE model-tier escalation) is gated on the same
dynamic_routing.enabled/escalate_on_failureswitches and does NOT modifyresolveModelForTier(so existing feat-3024 behavior is unchanged).
Consequences
Positive
- One coherent effort policy across all runtimes; Claude effort is first-class and actually wired.
- The dead
reasoning_effortfield becomes meaningful; finer-grained cost/quality control (a light-tier scanning agent can runloweffort; a heavy planning agentxhigh) without changing model class. - Effort-first escalation reduces unnecessary model upgrades.
- Cross-provider clamping is explicit and documented.
Negative
- The universal enum is the union of two providers' ladders, so two levels (
max,minimal) are runtime-specific and clamp when rendered to the other provider — users must understand the mapping (mitigated by help text and the table above). - Fast mode remains asymmetric: it cannot be forced per-subagent on Claude Code, only at session level or on API-direct runtimes.
- Updating issue-2517's tests to assert real wiring is a deliberate behavior/contract change (the old "null on claude" assertion encoded the now-false premise that Claude has no effort control).
Alternatives Considered
(a) Global effort env override (e.g. a single CLAUDE_CODE_EFFORT_LEVEL for the whole session) — rejected: caps cost but starves heavy agents that legitimately need deep reasoning; static global breaks the per-tier design.
(b) Model selection alone (status quo) — rejected: choosing Haiku for light tasks reduces cost, but within one model class there is no way to tune reasoning depth; a quality profile pays full reasoning cost even for scanning.
(c) Static per-agent effort only — rejected: loses context sensitivity; the same agent doing trivial vs complex work should not always get the same effort.
(d) A separate effort field kept fully parallel to Codex's existing reasoning_effort (two independent lanes) — rejected: produces two overlapping fields that can diverge and confuse; Codex's reasoning_effort is better modeled as one rendering of the single universal effort.
(e) Overloading the existing reasoning_effort field to also carry Claude effort — rejected: it would conflate a Codex-specific wire name with the universal concept and break the clean per-runtime rendering.
References
- Tracking issue: #443
- Prior art (inert reasoning_effort): #2517;
tests/model-resolver.test.cjs(folds formerissue-2517-runtime-aware-profiles) - dynamic_routing escalation: #3024;
tests/model-profiles.test.cjs(folds formerfeat-3024-dynamic-routing, consolidation epic #1969) - phase-type tiers: #3023
- Anthropic effort API:
output_config.effort(low/medium/high/xhigh/max); fast mode:speed(standard/fast) - Claude Code effort:
/effort,--effort,CLAUDE_CODE_EFFORT_LEVEL,effortLevelsetting, subagenteffortfrontmatter (anthropics/claude-code #31536, completed); fast mode:/fast(interactive only) - OpenAI Codex effort:
model_reasoning_effortconfig key; Responses APIreasoning.effort;ReasoningEffortenumnone<minimal<low<medium<high<xhigh