* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants
The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.
Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:
- scripts/gen-adr-index.cjs generates the index between markers and validates
the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.
Correct the lifecycle metadata the gate surfaced, without flipping any status:
- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
capability system shipped and epic #857 is closed. Ratification is a
maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: capture stderr via spawnSync; record ADR-0010 draft supersession
Two fixes surfaced by the first gsd-test run and by regenerating the index:
- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
through the thrown error on non-zero exit. The `--write` path exits 0 while
reporting outstanding violations on stderr, so the helper always saw ''.
spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
"earlier draft superseded by ADR-0011" while the file itself still said
Proposed. Deriving the index from the files would have dropped that
assertion and resurrected a superseded draft as a live decision, so it is
recorded at its source, with the reciprocal Supersedes on ADR-0011.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174
src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").
ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (11918dcc3), so the candidate can no longer resolve in any layout: a source
repo has no sdk/, and an install layout points it at ~/.claude/sdk/shared/, which
the installer never writes -- the original #3288 bug. It was dead weight implying
a package boundary this repo no longer has.
No test depends on it: the #3288 regression tests in tests/install.test.cjs write
their own synthetic old-path fixture and assert it throws.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: ratify nine shipped ADRs; record why ten others stay Proposed
The corpus carried 19 Proposed ADRs, most describing architecture that had
already shipped. A Proposed label on live architecture tells contributors and
agents the decision is an unbuilt idea -- the capability system and EoS were
both being misread that way.
Audited all 19 against the shipped tree and GitHub. Each candidate flip then had
to survive two independent reviewers instructed to refute it.
Ratified Proposed -> Accepted, each with a dated Ratification section carrying
the verified evidence (file:line, symbols, tests, issue state):
857 capability system 894 declaration format 1244 capability ecosystem
1577 injection boundary 1610 size-budget ratchet 1990 existing-code onboarding
15 cross-AI convergence 22 plan-drift guard 0011 default reviewers
Held ten, each now carrying a "Why this is still Proposed" section naming the
blocker and its unblock condition, so the audit is not repeated:
2264 its own headline acceptance criterion is unmet in the tree
230 live branch protection contradicts the decided spec (1 approval, not 2)
660 the namesake release/<version> re-cut is manual, not automated
959 issue #2346 is approved and plans its graduation as its own ADR
1213 the shipped writer's return shape differs from the decided interface
443 the orchestrator override path has no live caller
1143 / 1606 each states its own bar for acceptance; neither is met
612 / 1671 legitimately open
Shipped code proved necessary but not sufficient: eight ADRs had every named
module, symbol, and test present with their epics closed, and still failed the
bar. That lesson is written into README.md's ratification procedure.
Also corrected ADR-857's "Supersedes (generalizes)" to "Subsumes": taken
literally it would have marked two live seams dead -- ADR-0011 (surface.cts:348)
and ADR-58 (runtime-artifact-install-plan.cts:82). Both keep Accepted status and
gain Subsumed-by pointers.
Index: Active 39->48, Proposed 19->10, Superseded/Legacy 7. 65 total.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: harden gen-adr-index against hostile titles and non-ADR filenames (#2356)
Three findings from the pre-PR orthogonal security review, all confirmed:
- An ADR title containing the literal ADR-INDEX:END marker was emitted verbatim
into its table cell, relocating the splice boundary so the NEXT --write
spliced against the wrong marker and truncated README.md. Titles now render
through cellText(), which escapes pipes and angle brackets -- making an HTML
comment (and any other HTML) unformable from ADR-authored text.
- A docs/adr/*.md without a numeric prefix crashed on match(...)[1] of null.
Such a file is also invisible to the index -- the very failure this gate
exists to prevent -- so it is now reported as a naming-convention violation
naming the file and the fix.
- Tests leaked their mkdtemp dirs. They now use helpers.createTempDir/cleanup
via t.after(); helpers.cleanup carries the Windows-EBUSY retry budget that a
raw fs.rmSync lacks (caught by local/no-raw-rmsync-in-tests).
Adds five regression tests: marker hijack, HTML injection, pipe cell-break,
non-conforming filename, and splice stability across repeated writes.
Refs #2356
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: close two gate false-passes; read ## Supersedes sections (#2356)
Second round of confirmed findings from the pre-PR orthogonal code review. Both
false-passes matter more than a false-fail: a gate that silently misses a
violation is worse than no gate, because it is trusted.
- A relation field mixing a link with a bare id silently dropped the bare claim:
the check tested `rel.links.length` (does this field have ANY link?) instead
of whether THAT id was linked. `Supersedes: [ADR-0001](...), ADR-0011` passed
clean -- accepting exactly the ambiguous bare reference the rule forbids. Now
each bare id is checked against the ids actually linked in the same field, so
a repeat in trailing prose stays quiet while an unlinked claim is flagged.
- The ratification guard (`statusToken !== 'Accepted'`) skipped BOTH relation
directions, which killed the IN check entirely: `supersedes.in` is only ever
populated on an ADR whose status IS `Superseded`, so a dangling `Superseded by
X` where X never claims it always passed. The guard now applies to OUT only --
a prospective claim must not obligate its target, but an ADR's statement about
ITSELF is always owed a reciprocal.
- Fixing that surfaced a parser gap: ADR-0174 declares its supersessions in a
`## Supersedes` table SECTION, not a header field, and headerBlock() stops at
the first `##`. The repo's best-documented supersession was invisible. Section
form is now parsed for both relations.
- Replaced a vacuous test: the em-dash negation case passed whether or not
NEGATED_RELATION_RE matched (a mutation to /$^/ survived). It now carries a
link that would create a failing asymmetric relation if negation did not fire.
Also removes docs/adr/9401-test-target.md -- a synthetic fixture a reviewer
created in the worktree while reproducing a finding, swept in by `git add -A`.
Adds regression tests for each: mixed link+bare, linked-and-repeated-in-prose,
dangling superseded-by from a non-Accepted ADR, and the ADR-0174 section shape.
Refs #2356
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: escape backslashes before pipes in the ADR index cell renderer (#2356)
CodeQL js/incomplete-sanitization (high) on scripts/gen-adr-index.cjs: cellText()
escaped `|` -> `\|` without first escaping the backslash. Markdown's escape
character is the backslash, so the input `\|` became `\\|`, which renders as a
literal backslash followed by an UNESCAPED pipe -- re-opening the cell break the
pipe escape exists to prevent. Order is load-bearing: escape the escape
character first, then everything that emits one.
Same class as the index-marker hijack fixed earlier: ADR-authored text breaking
out of the cell it is rendered into.
Adds a regression test asserting a `\|`-bearing title leaves exactly the row's
own 5 unescaped delimiters and cannot forge a Status cell. Uses split(/\r?\n/)
per local/no-crlf-fragile-split -- a literal "\n" split is CRLF-fragile on the
Windows CI leg.
Refs #2356
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
12 KiB
ADR 443: Unified cross-provider effort controls and fast-mode-aware routing
- Status: Proposed (2026-05-28)
- Date: 2026-05-28
- Tracking issue: #443
Why this is still Proposed (audited 2026-07-17)
The audit confirmed the cross-provider resolver/renderer/CLI machinery genuinely shipped: resolveEffortInternal, resolveEffortForTier, renderEffortForRuntime, RUNTIMES_WITH_FAST_MODE, and cmdResolveExecution (src/model-resolver.cts:534,654; src/commands.cts) implement the cascade and clamping exactly as Decision items 1–3, 5, and 6 describe, and static install-time propagation is real and end-to-end tested — tests/install-runtime-artifacts.test.cjs's describe('#443 Claude install: effort: injected into frontmatter') runs the actual install() function and reads the resulting agent .md files off disk, confirming gsd-planner gets effort: xhigh, gsd-codebase-mapper gets effort: low, and gsd-executor gets effort: high. That test predates the QA audit below (landed 2026-05-29 in the original #443 PR, commit 5ca646f01), so the "resolver-only, nothing reaches the runtime" framing of the original flavor-text problem this ADR set out to fix is fixed for the static path.
The blocker. Decision item 1's cascade names an "(1) orchestrator invocation override" as the highest-precedence layer, and Decision item 6 adds a dynamic escalation path ("effort steps up the ladder on a failed attempt"). Both exist only as CLI-callable resolver code — resolveEffortInternal's invocation-override step (src/model-resolver.cts:535) and resolveEffortForTier's attempt-based escalation (src/model-resolver.cts:654) — exercised solely by unit/CLI tests. Nothing in the shipped orchestration actually calls them: a search across every file in gsd-core/workflows/*.md and agents/*.md for resolve-execution or CLAUDE_CODE_EFFORT_LEVEL returns zero hits; the only workflow-level mentions of "effort" are documentation of the config keys in settings-advanced.md's confirmation table. The only propagation channel actually wired into a real GSD flow is the static one (config → install() → frontmatter, baked once at install time) — the ADR's own decided design promises more than that, and the more-than-static-baking part has no consumer. Separately, the repo's own dated QA test-architecture audit (docs/issueevidence/1192-adr-test-audit-2026-06-13.md, produced under issue #1192, closed COMPLETED) rated ADR-443 "partial ... end-to-end effort propagation untested" and named it in its action plan ("Strengthen ... ADR-443 end-to-end effort propagation," line 220); that action item was never converted into a tracked follow-up issue, and no commit since 2026-06-13 addresses it. That audit's blanket "untested" framing overstates the gap — the static path is tested — but the underlying signal (a decided mechanism with no live caller) is real and independently confirmed here.
Unblock condition. Either (a) wire the orchestrator-invocation-override and attempt-based-escalation paths into an actual GSD workflow or agent dispatch (so resolveEffortForTier's escalation and resolveEffortInternal's invocation-override step have a real caller outside src/commands.cts's CLI surface and tests), and add a test exercising that live path the way tests/install-runtime-artifacts.test.cjs exercises the static one; or (b) if the ADR's intended scope is in fact limited to static install-time propagation, amend Decision items 1 and 6 to say so explicitly and close out audit issue #1192's action-plan item 18 with a note pointing at the shipped install-wiring tests. Either is a maintainer call this file records but does not make.
Context
Effort control and fast mode in Claude Opus 4.8
Claude Opus 4.8 introduced two orthogonal execution controls relevant to GSD's agent orchestration:
-
Effort control — API request field
output_config.effort(string enum). Anthropic levels:low,medium,high,xhigh,max; Opus 4.8 defaults tohigh. In Claude Code it is exposed as/effort, the--effortCLI flag, theCLAUDE_CODE_EFFORT_LEVELenv var, theeffortLevelsettings.json key (acceptslow/medium/high/xhigh;maxis session-only), and — critically for orchestration — a per-subagenteffortfrontmatter key (shipped per anthropics/claude-code issue #31536, CLOSED/COMPLETED). -
Fast mode — API request field
speed(standard|fast);fastenables high output-tokens-per-second inference. Pricing for Opus 4.8 fast mode is $10/$50 per MTok in/out vs $5/$25 standard. In Claude Code it is the interactive/fasttoggle ONLY — there is no settings.json key, env var, or subagent-frontmatter mechanism to enable fast mode for a spawned subagent.
GSD already routes WHICH model runs a task (routingTier heavy/standard/light, model_profile quality/balanced/budget/adaptive/inherit, model_overrides, and dynamic_routing escalation). It had no way to control HOW HARD the model reasons or WHICH speed tier it uses.
The "flavor text" problem: issue #2517
Issue #2517 added resolveReasoningEffortInternal and made query resolve-model emit a reasoning_effort field derived from the Codex runtime's per-tier catalog values (model-catalog.json runtimeTierDefaults.codex.*.reasoning_effort). However, a codebase audit found that NO orchestrator, workflow, or agent ever consumes that emitted field — it is never passed to an actual Codex invocation. The resolver computed a value and a test asserted the computed JSON, but the value reached no runtime. The feature was inert ("flavor text, no code"): asserting a resolver's return value is not the same as asserting the control reaches the model.
Cross-provider effort enum mismatch
The two providers' effort enums are NOT identical:
- Anthropic/Claude (Opus 4.8):
low,medium,high,xhigh,max(hasmax; nominimal) - OpenAI/Codex (
model_reasoning_effort/ Responses APIreasoning.effort; SDKReasoningEffortranksnone=0,minimal=1,low=2,medium=3,high=4,xhigh=5):minimal,low,medium,high,xhigh(hasminimal; nomax)
Common core: low, medium, high, xhigh.
Decision
-
Introduce a single universal
effortconfig knob (and an orthogonalfast_modeknob) that compose with model selection rather than replace it. Resolution precedence mirrors the existing model cascade: (1) orchestrator invocation override, (2)effort.agent_overrides[agent], (3)effort.routing_tier_defaults[routingTier], (4)effort.default, (5) built-in defaulthigh. Same cascade forfast_modewith built-in defaultfalse. Invalid enum values at any level are ignored and fall through (mirrors theVALID_TIERSgate inresolveModelInternal) so a typo never silently breaks resolution. -
The universal effort value is provider-agnostic; a per-runtime renderer maps it to each runtime's wire parameter, clamping the genuinely-unique tail levels:
- Claude / API: param
output_config.effort(Claude Code: subagenteffortfrontmatter /CLAUDE_CODE_EFFORT_LEVELenv).minimalclamps tolow(Claude has nominimal);low/medium/high/xhigh/maxpass through. - Codex: param
model_reasoning_effort(Responses APIreasoning.effort).maxclamps toxhigh(Codex has nomax);minimal/low/medium/high/xhighpass through.
Universal level Claude rendering Codex rendering minimallow(clamped)minimallowlowlowmediummediummediumhigh(default)highhighxhighxhighxhighmaxmaxxhigh(clamped) - Claude / API: param
-
Fold the inert
reasoning_effortoutput into this unified model.query resolve-modelis preserved for back-compat; a NEWquery resolve-executionis the superset that emits:model,effort(universal), the per-runtime rendered effort, the wire param name, the propagation channel,fast_mode, andfast_mode_supported. Each config key ships help text naming exactly which runtime field/invocation it drives. -
Make effort actually reach the runtime (close the flavor-text gap). Claude is first-class: the resolved effort propagates to spawned subagents via the
effortfrontmatter /CLAUDE_CODE_EFFORT_LEVELenv. Tests assert end-to-end propagation, not just resolver return values. -
Fast mode honesty: because Claude Code has no per-subagent fast-mode mechanism,
fast_modeis resolved and surfaced (with afast_mode_supportedflag,falsefor the claude runtime's subagents) but is NEVER emitted as a fake frontmatter key — doing so would be a silent no-op. It propagates only where the runtime supports it (APIspeed:"fast"). -
Dynamic-routing integration is additive: a new effort-escalation path (effort steps up the ladder on a failed attempt BEFORE model-tier escalation) is gated on the same
dynamic_routing.enabled/escalate_on_failureswitches and does NOT modifyresolveModelForTier(so existing feat-3024 behavior is unchanged).
Consequences
Positive
- One coherent effort policy across all runtimes; Claude effort is first-class and actually wired.
- The dead
reasoning_effortfield becomes meaningful; finer-grained cost/quality control (a light-tier scanning agent can runloweffort; a heavy planning agentxhigh) without changing model class. - Effort-first escalation reduces unnecessary model upgrades.
- Cross-provider clamping is explicit and documented.
Negative
- The universal enum is the union of two providers' ladders, so two levels (
max,minimal) are runtime-specific and clamp when rendered to the other provider — users must understand the mapping (mitigated by help text and the table above). - Fast mode remains asymmetric: it cannot be forced per-subagent on Claude Code, only at session level or on API-direct runtimes.
- Updating issue-2517's tests to assert real wiring is a deliberate behavior/contract change (the old "null on claude" assertion encoded the now-false premise that Claude has no effort control).
Alternatives Considered
(a) Global effort env override (e.g. a single CLAUDE_CODE_EFFORT_LEVEL for the whole session) — rejected: caps cost but starves heavy agents that legitimately need deep reasoning; static global breaks the per-tier design.
(b) Model selection alone (status quo) — rejected: choosing Haiku for light tasks reduces cost, but within one model class there is no way to tune reasoning depth; a quality profile pays full reasoning cost even for scanning.
(c) Static per-agent effort only — rejected: loses context sensitivity; the same agent doing trivial vs complex work should not always get the same effort.
(d) A separate effort field kept fully parallel to Codex's existing reasoning_effort (two independent lanes) — rejected: produces two overlapping fields that can diverge and confuse; Codex's reasoning_effort is better modeled as one rendering of the single universal effort.
(e) Overloading the existing reasoning_effort field to also carry Claude effort — rejected: it would conflate a Codex-specific wire name with the universal concept and break the clean per-runtime rendering.
References
- Tracking issue: #443
- Prior art (inert reasoning_effort): #2517;
tests/issue-2517-runtime-aware-profiles.test.cjs - dynamic_routing escalation: #3024;
tests/model-profiles.test.cjs(folds formerfeat-3024-dynamic-routing, consolidation epic #1969) - phase-type tiers: #3023
- Anthropic effort API:
output_config.effort(low/medium/high/xhigh/max); fast mode:speed(standard/fast) - Claude Code effort:
/effort,--effort,CLAUDE_CODE_EFFORT_LEVEL,effortLevelsetting, subagenteffortfrontmatter (anthropics/claude-code #31536, completed); fast mode:/fast(interactive only) - OpenAI Codex effort:
model_reasoning_effortconfig key; Responses APIreasoning.effort;ReasoningEffortenumnone<minimal<low<medium<high<xhigh