Files
msd-core/docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md
Tom Boucher 67a9243cf1 chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants

The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.

Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:

- scripts/gen-adr-index.cjs generates the index between markers and validates
  the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
  successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.

Correct the lifecycle metadata the gate surfaced, without flipping any status:

- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
  Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
  the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
  capability system shipped and epic #857 is closed. Ratification is a
  maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
  the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: capture stderr via spawnSync; record ADR-0010 draft supersession

Two fixes surfaced by the first gsd-test run and by regenerating the index:

- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
  through the thrown error on non-zero exit. The `--write` path exits 0 while
  reporting outstanding violations on stderr, so the helper always saw ''.
  spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
  "earlier draft superseded by ADR-0011" while the file itself still said
  Proposed. Deriving the index from the files would have dropped that
  assertion and resurrected a superseded draft as a live decision, so it is
  recorded at its source, with the reciprocal Supersedes on ADR-0011.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174

src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").

ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (11918dcc3), so the candidate can no longer resolve in any layout: a source
repo has no sdk/, and an install layout points it at ~/.claude/sdk/shared/, which
the installer never writes -- the original #3288 bug. It was dead weight implying
a package boundary this repo no longer has.

No test depends on it: the #3288 regression tests in tests/install.test.cjs write
their own synthetic old-path fixture and assert it throws.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: ratify nine shipped ADRs; record why ten others stay Proposed

The corpus carried 19 Proposed ADRs, most describing architecture that had
already shipped. A Proposed label on live architecture tells contributors and
agents the decision is an unbuilt idea -- the capability system and EoS were
both being misread that way.

Audited all 19 against the shipped tree and GitHub. Each candidate flip then had
to survive two independent reviewers instructed to refute it.

Ratified Proposed -> Accepted, each with a dated Ratification section carrying
the verified evidence (file:line, symbols, tests, issue state):

  857  capability system      894  declaration format   1244 capability ecosystem
  1577 injection boundary     1610 size-budget ratchet  1990 existing-code onboarding
  15   cross-AI convergence   22   plan-drift guard     0011 default reviewers

Held ten, each now carrying a "Why this is still Proposed" section naming the
blocker and its unblock condition, so the audit is not repeated:

  2264 its own headline acceptance criterion is unmet in the tree
  230  live branch protection contradicts the decided spec (1 approval, not 2)
  660  the namesake release/<version> re-cut is manual, not automated
  959  issue #2346 is approved and plans its graduation as its own ADR
  1213 the shipped writer's return shape differs from the decided interface
  443  the orchestrator override path has no live caller
  1143 / 1606 each states its own bar for acceptance; neither is met
  612 / 1671 legitimately open

Shipped code proved necessary but not sufficient: eight ADRs had every named
module, symbol, and test present with their epics closed, and still failed the
bar. That lesson is written into README.md's ratification procedure.

Also corrected ADR-857's "Supersedes (generalizes)" to "Subsumes": taken
literally it would have marked two live seams dead -- ADR-0011 (surface.cts:348)
and ADR-58 (runtime-artifact-install-plan.cts:82). Both keep Accepted status and
gain Subsumed-by pointers.

Index: Active 39->48, Proposed 19->10, Superseded/Legacy 7. 65 total.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: harden gen-adr-index against hostile titles and non-ADR filenames (#2356)

Three findings from the pre-PR orthogonal security review, all confirmed:

- An ADR title containing the literal ADR-INDEX:END marker was emitted verbatim
  into its table cell, relocating the splice boundary so the NEXT --write
  spliced against the wrong marker and truncated README.md. Titles now render
  through cellText(), which escapes pipes and angle brackets -- making an HTML
  comment (and any other HTML) unformable from ADR-authored text.
- A docs/adr/*.md without a numeric prefix crashed on match(...)[1] of null.
  Such a file is also invisible to the index -- the very failure this gate
  exists to prevent -- so it is now reported as a naming-convention violation
  naming the file and the fix.
- Tests leaked their mkdtemp dirs. They now use helpers.createTempDir/cleanup
  via t.after(); helpers.cleanup carries the Windows-EBUSY retry budget that a
  raw fs.rmSync lacks (caught by local/no-raw-rmsync-in-tests).

Adds five regression tests: marker hijack, HTML injection, pipe cell-break,
non-conforming filename, and splice stability across repeated writes.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: close two gate false-passes; read ## Supersedes sections (#2356)

Second round of confirmed findings from the pre-PR orthogonal code review. Both
false-passes matter more than a false-fail: a gate that silently misses a
violation is worse than no gate, because it is trusted.

- A relation field mixing a link with a bare id silently dropped the bare claim:
  the check tested `rel.links.length` (does this field have ANY link?) instead
  of whether THAT id was linked. `Supersedes: [ADR-0001](...), ADR-0011` passed
  clean -- accepting exactly the ambiguous bare reference the rule forbids. Now
  each bare id is checked against the ids actually linked in the same field, so
  a repeat in trailing prose stays quiet while an unlinked claim is flagged.
- The ratification guard (`statusToken !== 'Accepted'`) skipped BOTH relation
  directions, which killed the IN check entirely: `supersedes.in` is only ever
  populated on an ADR whose status IS `Superseded`, so a dangling `Superseded by
  X` where X never claims it always passed. The guard now applies to OUT only --
  a prospective claim must not obligate its target, but an ADR's statement about
  ITSELF is always owed a reciprocal.
- Fixing that surfaced a parser gap: ADR-0174 declares its supersessions in a
  `## Supersedes` table SECTION, not a header field, and headerBlock() stops at
  the first `##`. The repo's best-documented supersession was invisible. Section
  form is now parsed for both relations.
- Replaced a vacuous test: the em-dash negation case passed whether or not
  NEGATED_RELATION_RE matched (a mutation to /$^/ survived). It now carries a
  link that would create a failing asymmetric relation if negation did not fire.

Also removes docs/adr/9401-test-target.md -- a synthetic fixture a reviewer
created in the worktree while reproducing a finding, swept in by `git add -A`.

Adds regression tests for each: mixed link+bare, linked-and-repeated-in-prose,
dangling superseded-by from a non-Accepted ADR, and the ADR-0174 section shape.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: escape backslashes before pipes in the ADR index cell renderer (#2356)

CodeQL js/incomplete-sanitization (high) on scripts/gen-adr-index.cjs: cellText()
escaped `|` -> `\|` without first escaping the backslash. Markdown's escape
character is the backslash, so the input `\|` became `\\|`, which renders as a
literal backslash followed by an UNESCAPED pipe -- re-opening the cell break the
pipe escape exists to prevent. Order is load-bearing: escape the escape
character first, then everything that emits one.

Same class as the index-marker hijack fixed earlier: ADR-authored text breaking
out of the cell it is rendered into.

Adds a regression test asserting a `\|`-bearing title leaves exactly the row's
own 5 unescaped delimiters and cannot forge a Status cell. Uses split(/\r?\n/)
per local/no-crlf-fragile-split -- a literal "\n" split is CRLF-fragile on the
Windows CI leg.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 10:51:58 -04:00

12 KiB
Raw Blame History

ADR 443: Unified cross-provider effort controls and fast-mode-aware routing

  • Status: Proposed (2026-05-28)
  • Date: 2026-05-28
  • Tracking issue: #443

Why this is still Proposed (audited 2026-07-17)

The audit confirmed the cross-provider resolver/renderer/CLI machinery genuinely shipped: resolveEffortInternal, resolveEffortForTier, renderEffortForRuntime, RUNTIMES_WITH_FAST_MODE, and cmdResolveExecution (src/model-resolver.cts:534,654; src/commands.cts) implement the cascade and clamping exactly as Decision items 1–3, 5, and 6 describe, and static install-time propagation is real and end-to-end tested — tests/install-runtime-artifacts.test.cjs's describe('#443 Claude install: effort: injected into frontmatter') runs the actual install() function and reads the resulting agent .md files off disk, confirming gsd-planner gets effort: xhigh, gsd-codebase-mapper gets effort: low, and gsd-executor gets effort: high. That test predates the QA audit below (landed 2026-05-29 in the original #443 PR, commit 5ca646f01), so the "resolver-only, nothing reaches the runtime" framing of the original flavor-text problem this ADR set out to fix is fixed for the static path.

The blocker. Decision item 1's cascade names an "(1) orchestrator invocation override" as the highest-precedence layer, and Decision item 6 adds a dynamic escalation path ("effort steps up the ladder on a failed attempt"). Both exist only as CLI-callable resolver code — resolveEffortInternal's invocation-override step (src/model-resolver.cts:535) and resolveEffortForTier's attempt-based escalation (src/model-resolver.cts:654) — exercised solely by unit/CLI tests. Nothing in the shipped orchestration actually calls them: a search across every file in gsd-core/workflows/*.md and agents/*.md for resolve-execution or CLAUDE_CODE_EFFORT_LEVEL returns zero hits; the only workflow-level mentions of "effort" are documentation of the config keys in settings-advanced.md's confirmation table. The only propagation channel actually wired into a real GSD flow is the static one (config → install() → frontmatter, baked once at install time) — the ADR's own decided design promises more than that, and the more-than-static-baking part has no consumer. Separately, the repo's own dated QA test-architecture audit (docs/issueevidence/1192-adr-test-audit-2026-06-13.md, produced under issue #1192, closed COMPLETED) rated ADR-443 "partial ... end-to-end effort propagation untested" and named it in its action plan ("Strengthen ... ADR-443 end-to-end effort propagation," line 220); that action item was never converted into a tracked follow-up issue, and no commit since 2026-06-13 addresses it. That audit's blanket "untested" framing overstates the gap — the static path is tested — but the underlying signal (a decided mechanism with no live caller) is real and independently confirmed here.

Unblock condition. Either (a) wire the orchestrator-invocation-override and attempt-based-escalation paths into an actual GSD workflow or agent dispatch (so resolveEffortForTier's escalation and resolveEffortInternal's invocation-override step have a real caller outside src/commands.cts's CLI surface and tests), and add a test exercising that live path the way tests/install-runtime-artifacts.test.cjs exercises the static one; or (b) if the ADR's intended scope is in fact limited to static install-time propagation, amend Decision items 1 and 6 to say so explicitly and close out audit issue #1192's action-plan item 18 with a note pointing at the shipped install-wiring tests. Either is a maintainer call this file records but does not make.

Context

Effort control and fast mode in Claude Opus 4.8

Claude Opus 4.8 introduced two orthogonal execution controls relevant to GSD's agent orchestration:

  1. Effort control — API request field output_config.effort (string enum). Anthropic levels: low, medium, high, xhigh, max; Opus 4.8 defaults to high. In Claude Code it is exposed as /effort, the --effort CLI flag, the CLAUDE_CODE_EFFORT_LEVEL env var, the effortLevel settings.json key (accepts low/medium/high/xhigh; max is session-only), and — critically for orchestration — a per-subagent effort frontmatter key (shipped per anthropics/claude-code issue #31536, CLOSED/COMPLETED).

  2. Fast mode — API request field speed (standard|fast); fast enables high output-tokens-per-second inference. Pricing for Opus 4.8 fast mode is $10/$50 per MTok in/out vs $5/$25 standard. In Claude Code it is the interactive /fast toggle ONLY — there is no settings.json key, env var, or subagent-frontmatter mechanism to enable fast mode for a spawned subagent.

GSD already routes WHICH model runs a task (routingTier heavy/standard/light, model_profile quality/balanced/budget/adaptive/inherit, model_overrides, and dynamic_routing escalation). It had no way to control HOW HARD the model reasons or WHICH speed tier it uses.

The "flavor text" problem: issue #2517

Issue #2517 added resolveReasoningEffortInternal and made query resolve-model emit a reasoning_effort field derived from the Codex runtime's per-tier catalog values (model-catalog.json runtimeTierDefaults.codex.*.reasoning_effort). However, a codebase audit found that NO orchestrator, workflow, or agent ever consumes that emitted field — it is never passed to an actual Codex invocation. The resolver computed a value and a test asserted the computed JSON, but the value reached no runtime. The feature was inert ("flavor text, no code"): asserting a resolver's return value is not the same as asserting the control reaches the model.

Cross-provider effort enum mismatch

The two providers' effort enums are NOT identical:

  • Anthropic/Claude (Opus 4.8): low, medium, high, xhigh, max (has max; no minimal)
  • OpenAI/Codex (model_reasoning_effort / Responses API reasoning.effort; SDK ReasoningEffort ranks none=0, minimal=1, low=2, medium=3, high=4, xhigh=5): minimal, low, medium, high, xhigh (has minimal; no max)

Common core: low, medium, high, xhigh.

Decision

  1. Introduce a single universal effort config knob (and an orthogonal fast_mode knob) that compose with model selection rather than replace it. Resolution precedence mirrors the existing model cascade: (1) orchestrator invocation override, (2) effort.agent_overrides[agent], (3) effort.routing_tier_defaults[routingTier], (4) effort.default, (5) built-in default high. Same cascade for fast_mode with built-in default false. Invalid enum values at any level are ignored and fall through (mirrors the VALID_TIERS gate in resolveModelInternal) so a typo never silently breaks resolution.

  2. The universal effort value is provider-agnostic; a per-runtime renderer maps it to each runtime's wire parameter, clamping the genuinely-unique tail levels:

    • Claude / API: param output_config.effort (Claude Code: subagent effort frontmatter / CLAUDE_CODE_EFFORT_LEVEL env). minimal clamps to low (Claude has no minimal); low/medium/high/xhigh/max pass through.
    • Codex: param model_reasoning_effort (Responses API reasoning.effort). max clamps to xhigh (Codex has no max); minimal/low/medium/high/xhigh pass through.
    Universal level Claude rendering Codex rendering
    minimal low (clamped) minimal
    low low low
    medium medium medium
    high (default) high high
    xhigh xhigh xhigh
    max max xhigh (clamped)
  3. Fold the inert reasoning_effort output into this unified model. query resolve-model is preserved for back-compat; a NEW query resolve-execution is the superset that emits: model, effort (universal), the per-runtime rendered effort, the wire param name, the propagation channel, fast_mode, and fast_mode_supported. Each config key ships help text naming exactly which runtime field/invocation it drives.

  4. Make effort actually reach the runtime (close the flavor-text gap). Claude is first-class: the resolved effort propagates to spawned subagents via the effort frontmatter / CLAUDE_CODE_EFFORT_LEVEL env. Tests assert end-to-end propagation, not just resolver return values.

  5. Fast mode honesty: because Claude Code has no per-subagent fast-mode mechanism, fast_mode is resolved and surfaced (with a fast_mode_supported flag, false for the claude runtime's subagents) but is NEVER emitted as a fake frontmatter key — doing so would be a silent no-op. It propagates only where the runtime supports it (API speed:"fast").

  6. Dynamic-routing integration is additive: a new effort-escalation path (effort steps up the ladder on a failed attempt BEFORE model-tier escalation) is gated on the same dynamic_routing.enabled / escalate_on_failure switches and does NOT modify resolveModelForTier (so existing feat-3024 behavior is unchanged).

Consequences

Positive

  • One coherent effort policy across all runtimes; Claude effort is first-class and actually wired.
  • The dead reasoning_effort field becomes meaningful; finer-grained cost/quality control (a light-tier scanning agent can run low effort; a heavy planning agent xhigh) without changing model class.
  • Effort-first escalation reduces unnecessary model upgrades.
  • Cross-provider clamping is explicit and documented.

Negative

  • The universal enum is the union of two providers' ladders, so two levels (max, minimal) are runtime-specific and clamp when rendered to the other provider — users must understand the mapping (mitigated by help text and the table above).
  • Fast mode remains asymmetric: it cannot be forced per-subagent on Claude Code, only at session level or on API-direct runtimes.
  • Updating issue-2517's tests to assert real wiring is a deliberate behavior/contract change (the old "null on claude" assertion encoded the now-false premise that Claude has no effort control).

Alternatives Considered

(a) Global effort env override (e.g. a single CLAUDE_CODE_EFFORT_LEVEL for the whole session) — rejected: caps cost but starves heavy agents that legitimately need deep reasoning; static global breaks the per-tier design.

(b) Model selection alone (status quo) — rejected: choosing Haiku for light tasks reduces cost, but within one model class there is no way to tune reasoning depth; a quality profile pays full reasoning cost even for scanning.

(c) Static per-agent effort only — rejected: loses context sensitivity; the same agent doing trivial vs complex work should not always get the same effort.

(d) A separate effort field kept fully parallel to Codex's existing reasoning_effort (two independent lanes) — rejected: produces two overlapping fields that can diverge and confuse; Codex's reasoning_effort is better modeled as one rendering of the single universal effort.

(e) Overloading the existing reasoning_effort field to also carry Claude effort — rejected: it would conflate a Codex-specific wire name with the universal concept and break the clean per-runtime rendering.

References

  • Tracking issue: #443
  • Prior art (inert reasoning_effort): #2517; tests/issue-2517-runtime-aware-profiles.test.cjs
  • dynamic_routing escalation: #3024; tests/model-profiles.test.cjs (folds former feat-3024-dynamic-routing, consolidation epic #1969)
  • phase-type tiers: #3023
  • Anthropic effort API: output_config.effort (low/medium/high/xhigh/max); fast mode: speed (standard/fast)
  • Claude Code effort: /effort, --effort, CLAUDE_CODE_EFFORT_LEVEL, effortLevel setting, subagent effort frontmatter (anthropics/claude-code #31536, completed); fast mode: /fast (interactive only)
  • OpenAI Codex effort: model_reasoning_effort config key; Responses API reasoning.effort; ReasoningEffort enum none<minimal<low<medium<high<xhigh