* test(#443): RED unified effort + fast_mode + resolve-execution All 68 tests failing as expected — no implementation yet. Covers: effort cascade (tier defaults, overrides, invalid fallthrough), fast_mode cascade (boolean-only, tier defaults), resolveEffortForTier escalation, renderEffortForRuntime clamping, resolve-execution CLI, config schema new keys, QA hostile-input matrix. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(#443): unified cross-provider effort + fast_mode knobs and resolve-execution query Adds config-driven effort control (universal ladder: minimal<low<medium<high<xhigh<max) and fast_mode propagation knobs, with per-runtime rendering that clamps the unique tail values (max=Anthropic-only clamps to xhigh on Codex; minimal=Codex-only clamps to low on Claude). Key changes: - config-schema.manifest.json: add effort.default, fast_mode.enabled as validKeys; add 4 dynamicKeyPatterns for effort.routing_tier_defaults, effort.agent_overrides, fast_mode.routing_tier_defaults, fast_mode.agent_overrides; fix stale _comment - config-defaults.manifest.json: add effort and fast_mode blocks with tier defaults - model-catalog.cjs: add EFFORT_RENDERING map, renderEffortForRuntime(), RUNTIMES_WITH_FAST_MODE - model-profiles.cjs: re-export new catalog exports - core.cjs: add resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, VALID_EFFORTS, EFFORT_SET, nextEffort; pass effort/fast_mode through loadConfig - commands.cjs: replace reasoning_effort in cmdResolveModel with unified effort; add cmdResolveExecution (superset command with effort_rendered, effort_param, effort_propagation, fast_mode, fast_mode_supported) - gsd-tools.cjs: add resolve-execution case with --effort/--fast-mode/--attempt flags - tests/feat-443: 69 tests covering cascade, rendering, escalation, CLI, schema, QA matrix - tests/commands.test.cjs: convert 3 reasoning_effort assertions to unified effort - docs/CONFIGURATION.md: document effort + fast_mode + resolve-execution sections - settings-advanced.md: list new effort/fast_mode keys in confirmation table Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#443): remove dead catalog effort lane; unify codex effort through renderEffortForRuntime - Remove resolveReasoningEffortInternal (catalog-driven effort function) from core.cjs and its export; remove from commands.cjs destructure import - Convert tests/issue-2517-runtime-aware-profiles.test.cjs: all 11 effort assertions now use resolveEffortInternal + renderEffortForRuntime; Claude effort is first-class (output_config.effort); unknown runtimes assert param===null - Convert tests/feat-3023-model-phase-types.test.cjs: replace the entire resolveReasoningEffortInternal describe with unified effort assertions; effort derives from AGENT_DEFAULT_TIERS routing tier, not phase-type tier Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(#443): ADR for unified cross-provider effort + fast-mode routing Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(#443): architecture-level QA invariants + test-strategy doc Add 48-test integration suite (feat-443-effort-fast-mode.integration.test.cjs) covering 8 architectural invariants: cross-provider validity (never emit a value the real API would 400 on), param/channel contract stability, resolve-execution JSON contract (all 8 keys + correct types), totality across the full 33-agent registry, fast-mode honesty (claude always fast_mode_supported=false), precedence first-valid-wins matrix for both effort and fast_mode cascades, dynamic-routing composition (effort escalation independent of model tier), and config-set round-trip for all new effort/* and fast_mode/* key namespaces. Append test-strategy section with invariant rationale and E2E gap documentation to docs/TESTING-SUITES.md. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#443): add failing install-wiring tests for effort per-runtime injection (RED) TDD RED: 10 failing tests covering: - Claude .md gets effort: injected per tier (planner=xhigh, mapper=low, executor=high) - Gemini .md does NOT get effort: (already passing — Gemini-safe) - Codex .toml gets model_reasoning_effort via unified resolver - Config-driven: effort.agent_overrides drives both Claude .md and Codex .toml - Source purity: agents/*.md have no effort: key (already passing) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(#443): wire effort per-runtime at install (Claude .md frontmatter + Codex .toml unified) - Import AGENT_DEFAULT_TIERS and renderEffortForRuntime from model-catalog.cjs - Add readGsdEffectiveEffortConfig(targetDir): reads merged effort config from .planning/config.json (per-project wins) + ~/.gsd/defaults.json (global fallback), same probe pattern as readGsdRuntimeProfileResolver - Add resolveInstallTimeEffort(effortCfg, agentName): pure function matching resolveEffortInternal() precedence (agent_overrides > routing_tier_defaults > default > 'high') without loadConfig side-effects (no sub-repo detection, no migration writes) - Claude agent copy loop: inject `effort: <value>` into frontmatter ONLY for runtime === 'claude'; all other .md runtimes (Gemini, Qwen, Hermes, etc.) stay effort-free (Gemini-safe source contract preserved in agents/*.md) - generateCodexAgentToml: add effortCfg param; emit model_reasoning_effort from unified resolver (replaces old catalog entry.reasoning_effort); Codex clamps max → xhigh via renderEffortForRuntime('codex', ...) - installCodexConfig: pass readGsdEffectiveEffortConfig(targetDir) to generateCodexAgentToml so per-project config wins for Codex .toml too - Update failing tests to GREEN: 12/12 pass; all 17 install tests pass; 2847/2848 unit tests pass (1 pre-existing failure: policy-shell-pinning) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#443): source install effort defaults from manifest (kill drift) + guard test Replace hardcoded _GSD_EFFORT_MANIFEST_TIER_DEFAULTS and the 'high' fallback in resolveInstallTimeEffort with values read from config-defaults.manifest.json at module init, using the same __dirname-relative path install.js already uses for all shared manifests. Add feat-443-effort-defaults-drift.test.cjs to assert equality between install.js's runtime constants and the manifest on every CI run. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): reconcile Codex TOML tests with unified effort design The #443 unified effort resolver makes generateCodexAgentToml always emit model_reasoning_effort (driven by resolveInstallTimeEffort, not model_profile_overrides). The test 'generated TOML omits reasoning_effort when runtime has none' had an obsolete premise — model_profile_overrides.reasoning_effort:'' no longer suppresses unified effort. Convert it to assert the new invariant: Codex TOML always carries a valid model_reasoning_effort from the agent's routing tier (xhigh for gsd-planner, a heavy-tier agent), while model_profile_overrides model override is still respected. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): make install.js effort resolution lazy (no load-time side effects breaking launcher-parity) Replace module-load-time IIFE + hard throw (config-defaults.manifest.json read) and top-level require of model-catalog.cjs with a lazy _getGsdEffortCatalog() getter that initialises on first call from resolveInstallTimeEffort / generateCodexAgentToml / Claude .md effort injection. Requiring install.js in unrelated test contexts (e.g. runtime-launcher-parity) no longer triggers manifest IO or throws, eliminating the load-time side effect that changed subprocess exit codes / stderr on the bench. Drift-guard exports (_GSD_EFFORT_MANIFEST_TIER_DEFAULTS / _GSD_EFFORT_MANIFEST_DEFAULT) preserved as lazy getter properties on module.exports so feat-443-effort-defaults-drift still validates them without forcing eager load. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): isolate install-wiring test HOME to stop \$HOME/.claude pollution breaking launcher-parity runGlobalInstall() now redirects HOME to a per-call isolated tmpdir in addition to the existing runtime-specific env-var redirects (CLAUDE_CONFIG_DIR, GEMINI_CONFIG_DIR, CODEX_HOME). This ensures install.js code that uses os.homedir() directly — including the ~/.cache/gsd update-check deletion, ~/.gsd/defaults.json reads, and any HOME-relative npm subprocess writes — never touches the real \$HOME during the test. Without the HOME isolation the install test (which is new to this branch and is now picked up by Docker's raw \`tests/*.test.cjs\` glob) could write or delete files under the real \$HOME, causing runtime-launcher-parity test (D) to fail: (D) asserts a loud non-zero exit when \$RUNTIME_DIR/gsd-tools.cjs is absent and gsd-tools is not on PATH, but the launcher's \$HOME/.claude fallback arm succeeds if \$HOME/.claude/get-shit-done/bin/gsd-tools.cjs exists. Also sets GSD_SKIP_STALE_SDK_CHECK=1 to suppress the \`npm ls -g\` subprocess that the global installer spawns — irrelevant to effort-wiring assertions, slow, and potentially writes to ~/.npm cache. All 12 feat-443 install-wiring assertions preserved. Drift-guard 5/5. Unit suite 2848/2850 (pre-existing policy-shell-pinning.test.cjs failure on next). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#443): add changeset fragment for effort + fast-mode routing Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(#443): set GSD_TEST_MODE before requiring install.js in drift-guard test to prevent HOME leak Without GSD_TEST_MODE=1, require('bin/install.js') runs the module's main install block (guarded by !GSD_TEST_MODE), performing a real global Claude install into $HOME/.claude/. On CI ubuntu where node is on standard PATH, the launcher's $HOME/.claude fallback arm then finds gsd-tools.cjs, causing runtime-launcher-parity test (D) to exit zero when it must exit non-zero. Root cause: feat-443-effort-defaults-drift.test.cjs (unit suite) runs alphabetically before runtime-launcher-parity.test.cjs in the same node --test invocation. Each runs in a separate worker process but shares the same HOME. The drift test's install leaks gsd-tools.cjs into that HOME, then the launcher test's bash subprocess finds it via the $HOME/.claude arm. Fix: add process.env.GSD_TEST_MODE = '1' at the top of the drift-guard test, before the require(installPath) call. This matches the pattern used by feat-443-effort-fast-mode.test.cjs and feat-443-effort-install-wiring .install.test.cjs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): deterministic resolve-execution arg parsing + validate install-time effort (Codex adversarial findings) Finding 1: resolve-execution --effort low gsd-planner misrouted 'low' as the agent. Replace find(non-dash) with a proper flag-consuming loop that collects a single positional; validate missing/extra positionals and malformed --attempt values. Finding 2: resolveInstallTimeEffort returned unvalidated effort strings (e.g. "ultra") verbatim. Each precedence layer now checks GSD_EFFORT_SET (imported once from core.cjs) before accepting a value, mirroring resolveEffortInternal exactly. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): newline-agnostic effort frontmatter injection (Windows CRLF) + CRLF-safe assertions Extracts injectEffortFrontmatter(content, effortValue) pure helper that detects EOL (LF vs CRLF) from the opening '---' line and inserts 'effort: <value>' before the closing '---' delimiter using the same EOL as the surrounding frontmatter. Regex now uses /^---\r?\n([\s\S]*?)^---\r?$/m instead of the LF-only /^(---\n[\s\S]*?)(---)(\n|$)/ that silently skipped CRLF files on Windows (git core.autocrlf=true checkout). Also adds 7 unit tests covering LF, CRLF, idempotency, no-frontmatter, and complex frontmatter cases. Exports injectEffortFrontmatter from module.exports. Fixes 6 CI failures in tests/feat-443-effort-install-wiring.install.test.cjs on windows-latest runners (lines 138, 145, 152, 261, 345, 356). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -972,6 +972,123 @@ The `dynamic_routing` block is **disabled by default** — `enabled: false` (or
|
||||
|
||||
`dynamic_routing` is structurally a *cost lever*: you pay Opus rates only for the hard cases that warrant Opus. Compose with `model_overrides` for per-agent exceptions (override always wins).
|
||||
|
||||
---
|
||||
|
||||
### Effort Control (`effort`) — added in v1.42
|
||||
|
||||
> Unified cross-provider effort knob. Added in [#443](https://github.com/open-gsd/get-shit-done-redux/issues/443).
|
||||
|
||||
Control the reasoning effort of agent invocations with a single config. The universal ladder is:
|
||||
|
||||
```
|
||||
minimal < low < medium < high < xhigh < max
|
||||
```
|
||||
|
||||
Effort is rendered per-runtime: `output_config.effort` for Claude (Claude Code subagent `effort` frontmatter / `CLAUDE_CODE_EFFORT_LEVEL` env), `model_reasoning_effort` for Codex (Responses API `reasoning.effort`).
|
||||
|
||||
**Cross-provider clamping:** `max` is Anthropic-only — it clamps to `xhigh` on Codex. `minimal` is Codex-only — it clamps to `low` on Claude.
|
||||
|
||||
The model-catalog's `reasoning_effort` per-tier hint is a legacy field kept for reference; effort is now config-driven.
|
||||
|
||||
**Precedence (highest → lowest):**
|
||||
1. Invocation override (e.g. `--effort` flag on `resolve-execution`)
|
||||
2. `effort.agent_overrides[<agent-id>]`
|
||||
3. `effort.routing_tier_defaults[<light|standard|heavy>]`
|
||||
4. `effort.default`
|
||||
5. `"high"` (Anthropic Opus 4.8 universal default)
|
||||
|
||||
```json
|
||||
{
|
||||
"effort": {
|
||||
"default": "high",
|
||||
"routing_tier_defaults": {
|
||||
"light": "low",
|
||||
"standard": "high",
|
||||
"heavy": "xhigh"
|
||||
},
|
||||
"agent_overrides": {
|
||||
"gsd-planner": "max"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Settings
|
||||
|
||||
| Key | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `effort.default` | enum | `"high"` | Global fallback effort level. Applies when no tier or agent override matches. |
|
||||
| `effort.routing_tier_defaults.light` | enum | `"low"` | Effort for light-tier agents (fast mappers/scanners). |
|
||||
| `effort.routing_tier_defaults.standard` | enum | `"high"` | Effort for standard-tier agents (workhorse agents). |
|
||||
| `effort.routing_tier_defaults.heavy` | enum | `"xhigh"` | Effort for heavy-tier agents (deep reasoning). |
|
||||
| `effort.agent_overrides.<agent-id>` | enum | (none) | Per-agent effort override. Beats tier defaults. |
|
||||
|
||||
Valid effort values: `minimal`, `low`, `medium`, `high`, `xhigh`, `max`.
|
||||
|
||||
---
|
||||
|
||||
### Fast Mode (`fast_mode`) — added in v1.42
|
||||
|
||||
> Per-agent fast_mode propagation knob. Added in [#443](https://github.com/open-gsd/get-shit-done-redux/issues/443).
|
||||
|
||||
Control whether fast_mode is propagated to agent invocations. Only accepts real booleans — string `"true"` is rejected.
|
||||
|
||||
**Note:** `fast_mode` is only propagatable via API runtimes (`api` speed:"fast"). Claude Code has no per-subagent fast-mode mechanism — `/fast` is session-level only, so emitting a `fast_mode` frontmatter key on a Claude subagent is a silent no-op. `fast_mode_supported` in `resolve-execution` output tells you if the configured runtime supports it.
|
||||
|
||||
**Precedence (highest → lowest):**
|
||||
1. Invocation override (e.g. `--fast-mode` flag on `resolve-execution`)
|
||||
2. `fast_mode.agent_overrides[<agent-id>]` (boolean)
|
||||
3. `fast_mode.routing_tier_defaults[<light|standard|heavy>]` (boolean)
|
||||
4. `fast_mode.enabled` (boolean)
|
||||
5. `false`
|
||||
|
||||
```json
|
||||
{
|
||||
"fast_mode": {
|
||||
"enabled": false,
|
||||
"routing_tier_defaults": {
|
||||
"light": true,
|
||||
"standard": false,
|
||||
"heavy": false
|
||||
},
|
||||
"agent_overrides": {}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Settings
|
||||
|
||||
| Key | Type | Default | Description |
|
||||
|---|---|---|---|
|
||||
| `fast_mode.enabled` | boolean | `false` | Global fast_mode flag. Only honored when no tier/agent override matches. |
|
||||
| `fast_mode.routing_tier_defaults.light` | boolean | `true` | Fast mode for light-tier agents. |
|
||||
| `fast_mode.routing_tier_defaults.standard` | boolean | `false` | Fast mode for standard-tier agents. |
|
||||
| `fast_mode.routing_tier_defaults.heavy` | boolean | `false` | Fast mode for heavy-tier agents. |
|
||||
| `fast_mode.agent_overrides.<agent-id>` | boolean | (none) | Per-agent fast_mode override. |
|
||||
|
||||
---
|
||||
|
||||
### Execution Query (`resolve-execution`)
|
||||
|
||||
Use `node gsd-tools.cjs resolve-execution <agent-type> [--effort <level>] [--fast-mode <true|false>] [--attempt <n>]` to get the full resolved execution context for an agent:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "opus",
|
||||
"profile": "balanced",
|
||||
"effort": "xhigh",
|
||||
"effort_rendered": "xhigh",
|
||||
"effort_param": "output_config.effort",
|
||||
"effort_propagation": "frontmatter",
|
||||
"fast_mode": false,
|
||||
"fast_mode_supported": false
|
||||
}
|
||||
```
|
||||
|
||||
`effort_param` tells you which runtime parameter to set. `fast_mode_supported` tells you whether the configured runtime supports per-agent fast_mode propagation.
|
||||
|
||||
---
|
||||
|
||||
### Non-Claude Runtimes (Codex, OpenCode, Gemini CLI, Kilo)
|
||||
|
||||
> **Codex CLI minimum supported version: `0.130.0`** (issue [#3562](https://github.com/open-gsd/get-shit-done-redux/issues/3562)).
|
||||
|
||||
@@ -94,3 +94,139 @@ node scripts/ci-test-scope.cjs --base origin/next --head HEAD
|
||||
- Avoid stack-trace or error-message prose assertions. Assert `err.code`, structured JSON fields, or enums — Node minor releases routinely tweak error wording.
|
||||
- Prefer `node:test`, `node:assert/strict`, and `node:test` mocks. No external test frameworks.
|
||||
- Coverage uses `c8` and propagates `NODE_V8_COVERAGE` through the harness's child process.
|
||||
|
||||
---
|
||||
|
||||
## Test strategy: #443 effort + fast_mode engine
|
||||
|
||||
> Feature: unified cross-provider effort and fast_mode knobs (issue #443).
|
||||
> Test files: `tests/feat-443-effort-fast-mode.test.cjs` (unit),
|
||||
> `tests/feat-443-effort-fast-mode.integration.test.cjs` (integration).
|
||||
|
||||
### Testing pyramid
|
||||
|
||||
| Layer | File | What it covers |
|
||||
|---|---|---|
|
||||
| **Unit** | `feat-443-effort-fast-mode.test.cjs` | Pure logic: cascade rules, clamping, escalation math, malformed config handling, schema key validation. No CLI subprocess. |
|
||||
| **Integration** | `feat-443-effort-fast-mode.integration.test.cjs` | Architecture-level invariants: cross-provider validity, totality across the 33-agent registry, CLI JSON contract, config round-trip, fast-mode honesty. Real subprocesses via `runGsdTools`. |
|
||||
| **E2E** *(pending)* | *(not yet wired)* | Propagation layer: effort frontmatter / `CLAUDE_CODE_EFFORT_LEVEL` env actually reaching a spawned Claude Code subagent. See "Gaps" below. |
|
||||
|
||||
### Architectural invariants
|
||||
|
||||
Each invariant exists to prevent a specific class of production failure.
|
||||
|
||||
#### (a) Cross-provider validity
|
||||
|
||||
**What:** `renderEffortForRuntime(runtime, universalEffort).value` must always
|
||||
be a member of the runtime's real provider enum. Ground-truth enums are defined
|
||||
as local constants in the test — not sourced from the implementation.
|
||||
|
||||
```
|
||||
PROVIDER_EFFORT_ENUMS = {
|
||||
claude: Set { 'low', 'medium', 'high', 'xhigh', 'max' } // Anthropic output_config.effort
|
||||
codex: Set { 'minimal', 'low', 'medium', 'high', 'xhigh' } // OpenAI model_reasoning_effort
|
||||
}
|
||||
```
|
||||
|
||||
**Why:** Passing a value outside these sets results in a 400 from the real API.
|
||||
The clamping logic (`max -> xhigh` for codex; `minimal -> low` for claude) must
|
||||
hold for every cell of the VALID_EFFORTS × runtimes matrix.
|
||||
|
||||
#### (b) Param/channel contract
|
||||
|
||||
**What:** Each runtime exposes a stable `param` string (the native API field
|
||||
name) and `channel` (how the value is propagated). Unknown runtimes return
|
||||
`param: null, channel: null` and pass the effort value through unchanged.
|
||||
|
||||
**Why:** Callers read `.param` to construct the dispatch payload. A regression
|
||||
here would silently drop effort from subagent invocations.
|
||||
|
||||
#### (c) Resolve-execution JSON contract
|
||||
|
||||
**What:** The `gsd-tools resolve-execution <agent>` command emits a JSON object
|
||||
with all eight keys present and typed correctly: `model` (string), `profile`
|
||||
(string), `effort` (VALID_EFFORTS member), `effort_rendered` (string),
|
||||
`effort_param` (string|null), `effort_propagation` (string|null), `fast_mode`
|
||||
(boolean), `fast_mode_supported` (boolean).
|
||||
|
||||
**Why:** Orchestrators and workflow dispatchers parse this JSON. A missing or
|
||||
mistyped field silently breaks downstream consumers.
|
||||
|
||||
#### (d) Totality across the real registry
|
||||
|
||||
**What:** For every agent in the 33-agent registry, `resolveEffortInternal`
|
||||
returns a VALID_EFFORTS member (never undefined/null), `resolveFastModeInternal`
|
||||
returns a strict boolean, and `renderEffortForRuntime('claude', effort)` stays
|
||||
within the claude provider enum.
|
||||
|
||||
**Why:** A catalog addition that introduces a missing `routingTier` mapping
|
||||
would otherwise produce `undefined` and propagate silently.
|
||||
|
||||
#### (e) Fast-mode honesty invariant
|
||||
|
||||
**What:** When the runtime is `claude`, `fast_mode_supported` in
|
||||
resolve-execution output is always `false`, regardless of the fast_mode config.
|
||||
`RUNTIMES_WITH_FAST_MODE` contains only `'api'`.
|
||||
|
||||
**Why:** Claude Code's `/fast` toggle is session-level only. Emitting
|
||||
`fast_mode: true` as frontmatter on a Claude subagent is a silent no-op.
|
||||
Advertising `fast_mode_supported: true` for claude would cause orchestrators to
|
||||
believe the knob was wired when it is not.
|
||||
|
||||
#### (f) Precedence first-valid-wins
|
||||
|
||||
**What:** Both effort and fast_mode use a layered cascade. The test table covers
|
||||
all four effort layers (invocation override → agent_overrides →
|
||||
routing_tier_defaults → default) and all five fast_mode layers, including the
|
||||
case where an invalid value at a higher layer correctly falls through.
|
||||
|
||||
**Why:** Silent precedence bugs (e.g., a numeric value in agent_overrides not
|
||||
being rejected) would override intentional user config.
|
||||
|
||||
#### (g) Dynamic-routing composition
|
||||
|
||||
**What:** `resolveEffortForTier` escalates effort by attempt number
|
||||
independently of the model tier mapping. The test verifies the effort ladder
|
||||
(`low -> medium -> high -> xhigh -> max`), the `max` clamp, the
|
||||
`max_escalations` cap, and that `escalate_on_failure: false` suppresses
|
||||
escalation entirely.
|
||||
|
||||
**Why:** Effort escalation and model escalation share configuration
|
||||
(`dynamic_routing`) but must operate independently; coupling them would cause
|
||||
over-escalation or under-escalation.
|
||||
|
||||
#### (h) Config-tooling round-trip
|
||||
|
||||
**What:** `gsd-tools config-set` accepts all new key namespaces
|
||||
(`effort.default`, `effort.routing_tier_defaults.<tier>`,
|
||||
`effort.agent_overrides.<agent>`, `fast_mode.enabled`,
|
||||
`fast_mode.routing_tier_defaults.<tier>`, `fast_mode.agent_overrides.<agent>`)
|
||||
without an "Unknown config key" error, and values set via `config-set` are
|
||||
reflected in `resolve-execution` output.
|
||||
|
||||
**Why:** The schema validation gate (`VALID_CONFIG_KEYS` + `DYNAMIC_KEY_PATTERNS`)
|
||||
is separate from the resolver logic. A key missing from the schema would produce
|
||||
a silent write failure and appear as a bug only at runtime.
|
||||
|
||||
### Coverage targets
|
||||
|
||||
| Suite | Target |
|
||||
|---|---|
|
||||
| Unit | Every cascade rule, every fallthrough, every clamp. All function branches in `resolveEffortInternal`, `resolveFastModeInternal`, `resolveEffortForTier`, `renderEffortForRuntime`. |
|
||||
| Integration | All 8 architectural invariants. All 33 registered agents. All 6 provider × effort combinations for the valid-enum check. Full config-set key namespace. |
|
||||
|
||||
### Gaps / not yet covered
|
||||
|
||||
**E2E orchestrator-spawn-propagation layer (pending follow-up wiring):**
|
||||
The integration tests verify that GSD resolves and renders effort values
|
||||
correctly. They do NOT verify that the rendered values actually reach a spawned
|
||||
Claude Code or Codex subagent at runtime. Specifically uncovered:
|
||||
|
||||
- `CLAUDE_CODE_EFFORT_LEVEL` env var being set and read by a spawned claude subprocess
|
||||
- `output_config.effort` frontmatter key surviving the AGENTS.md template substitution
|
||||
- `model_reasoning_effort` field surviving serialization into a Codex API request body
|
||||
- Fast-mode `speed: "fast"` field reaching an `api`-runtime request when `fast_mode_supported: true`
|
||||
|
||||
These require spawning real subagents (or stubs thereof) and asserting on the
|
||||
process environment / request payload — a scope that belongs in a future E2E
|
||||
suite under `*.slow.test.cjs` or dedicated fixture-driven integration work.
|
||||
|
||||
93
docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md
Normal file
93
docs/adr/443-opus48-unified-effort-and-fast-mode-routing.md
Normal file
@@ -0,0 +1,93 @@
|
||||
# ADR 443: Unified cross-provider effort controls and fast-mode-aware routing
|
||||
|
||||
- **Status:** Proposed (2026-05-28)
|
||||
- **Date:** 2026-05-28
|
||||
- **Tracking issue:** [#443](https://github.com/open-gsd/get-shit-done-redux/issues/443)
|
||||
|
||||
## Context
|
||||
|
||||
### Effort control and fast mode in Claude Opus 4.8
|
||||
|
||||
Claude Opus 4.8 introduced two orthogonal execution controls relevant to GSD's agent orchestration:
|
||||
|
||||
1. **Effort control** — API request field `output_config.effort` (string enum). Anthropic levels: `low`, `medium`, `high`, `xhigh`, `max`; Opus 4.8 defaults to `high`. In Claude Code it is exposed as `/effort`, the `--effort` CLI flag, the `CLAUDE_CODE_EFFORT_LEVEL` env var, the `effortLevel` settings.json key (accepts `low`/`medium`/`high`/`xhigh`; `max` is session-only), and — critically for orchestration — a per-subagent `effort` frontmatter key (shipped per anthropics/claude-code issue #31536, CLOSED/COMPLETED).
|
||||
|
||||
2. **Fast mode** — API request field `speed` (`standard`|`fast`); `fast` enables high output-tokens-per-second inference. Pricing for Opus 4.8 fast mode is $10/$50 per MTok in/out vs $5/$25 standard. In Claude Code it is the interactive `/fast` toggle ONLY — there is no settings.json key, env var, or subagent-frontmatter mechanism to enable fast mode for a spawned subagent.
|
||||
|
||||
GSD already routes WHICH model runs a task (routingTier `heavy`/`standard`/`light`, model_profile `quality`/`balanced`/`budget`/`adaptive`/`inherit`, `model_overrides`, and dynamic_routing escalation). It had no way to control HOW HARD the model reasons or WHICH speed tier it uses.
|
||||
|
||||
### The "flavor text" problem: issue #2517
|
||||
|
||||
Issue #2517 added `resolveReasoningEffortInternal` and made `query resolve-model` emit a `reasoning_effort` field derived from the Codex runtime's per-tier catalog values (`model-catalog.json` `runtimeTierDefaults.codex.*.reasoning_effort`). However, a codebase audit found that NO orchestrator, workflow, or agent ever consumes that emitted field — it is never passed to an actual Codex invocation. The resolver computed a value and a test asserted the computed JSON, but the value reached no runtime. The feature was inert ("flavor text, no code"): asserting a resolver's return value is not the same as asserting the control reaches the model.
|
||||
|
||||
### Cross-provider effort enum mismatch
|
||||
|
||||
The two providers' effort enums are NOT identical:
|
||||
|
||||
- **Anthropic/Claude (Opus 4.8):** `low`, `medium`, `high`, `xhigh`, `max` (has `max`; no `minimal`)
|
||||
- **OpenAI/Codex** (`model_reasoning_effort` / Responses API `reasoning.effort`; SDK `ReasoningEffort` ranks `none=0`, `minimal=1`, `low=2`, `medium=3`, `high=4`, `xhigh=5`): `minimal`, `low`, `medium`, `high`, `xhigh` (has `minimal`; no `max`)
|
||||
|
||||
Common core: `low`, `medium`, `high`, `xhigh`.
|
||||
|
||||
## Decision
|
||||
|
||||
1. **Introduce a single universal `effort` config knob** (and an orthogonal `fast_mode` knob) that compose with model selection rather than replace it. Resolution precedence mirrors the existing model cascade: (1) orchestrator invocation override, (2) `effort.agent_overrides[agent]`, (3) `effort.routing_tier_defaults[routingTier]`, (4) `effort.default`, (5) built-in default `high`. Same cascade for `fast_mode` with built-in default `false`. Invalid enum values at any level are ignored and fall through (mirrors the `VALID_TIERS` gate in `resolveModelInternal`) so a typo never silently breaks resolution.
|
||||
|
||||
2. **The universal effort value is provider-agnostic; a per-runtime renderer maps it to each runtime's wire parameter**, clamping the genuinely-unique tail levels:
|
||||
|
||||
- **Claude / API:** param `output_config.effort` (Claude Code: subagent `effort` frontmatter / `CLAUDE_CODE_EFFORT_LEVEL` env). `minimal` clamps to `low` (Claude has no `minimal`); `low`/`medium`/`high`/`xhigh`/`max` pass through.
|
||||
- **Codex:** param `model_reasoning_effort` (Responses API `reasoning.effort`). `max` clamps to `xhigh` (Codex has no `max`); `minimal`/`low`/`medium`/`high`/`xhigh` pass through.
|
||||
|
||||
| Universal level | Claude rendering | Codex rendering |
|
||||
| --- | --- | --- |
|
||||
| `minimal` | `low` (clamped) | `minimal` |
|
||||
| `low` | `low` | `low` |
|
||||
| `medium` | `medium` | `medium` |
|
||||
| `high` (default) | `high` | `high` |
|
||||
| `xhigh` | `xhigh` | `xhigh` |
|
||||
| `max` | `max` | `xhigh` (clamped) |
|
||||
|
||||
3. **Fold the inert `reasoning_effort` output into this unified model.** `query resolve-model` is preserved for back-compat; a NEW `query resolve-execution` is the superset that emits: `model`, `effort` (universal), the per-runtime rendered effort, the wire param name, the propagation channel, `fast_mode`, and `fast_mode_supported`. Each config key ships help text naming exactly which runtime field/invocation it drives.
|
||||
|
||||
4. **Make effort actually reach the runtime (close the flavor-text gap).** Claude is first-class: the resolved effort propagates to spawned subagents via the `effort` frontmatter / `CLAUDE_CODE_EFFORT_LEVEL` env. Tests assert end-to-end propagation, not just resolver return values.
|
||||
|
||||
5. **Fast mode honesty:** because Claude Code has no per-subagent fast-mode mechanism, `fast_mode` is resolved and surfaced (with a `fast_mode_supported` flag, `false` for the claude runtime's subagents) but is NEVER emitted as a fake frontmatter key — doing so would be a silent no-op. It propagates only where the runtime supports it (API `speed:"fast"`).
|
||||
|
||||
6. **Dynamic-routing integration is additive:** a new effort-escalation path (effort steps up the ladder on a failed attempt BEFORE model-tier escalation) is gated on the same `dynamic_routing.enabled` / `escalate_on_failure` switches and does NOT modify `resolveModelForTier` (so existing feat-3024 behavior is unchanged).
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- One coherent effort policy across all runtimes; Claude effort is first-class and actually wired.
|
||||
- The dead `reasoning_effort` field becomes meaningful; finer-grained cost/quality control (a light-tier scanning agent can run `low` effort; a heavy planning agent `xhigh`) without changing model class.
|
||||
- Effort-first escalation reduces unnecessary model upgrades.
|
||||
- Cross-provider clamping is explicit and documented.
|
||||
|
||||
### Negative
|
||||
|
||||
- The universal enum is the union of two providers' ladders, so two levels (`max`, `minimal`) are runtime-specific and clamp when rendered to the other provider — users must understand the mapping (mitigated by help text and the table above).
|
||||
- Fast mode remains asymmetric: it cannot be forced per-subagent on Claude Code, only at session level or on API-direct runtimes.
|
||||
- Updating issue-2517's tests to assert real wiring is a deliberate behavior/contract change (the old "null on claude" assertion encoded the now-false premise that Claude has no effort control).
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
**(a) Global effort env override** (e.g. a single `CLAUDE_CODE_EFFORT_LEVEL` for the whole session) — rejected: caps cost but starves heavy agents that legitimately need deep reasoning; static global breaks the per-tier design.
|
||||
|
||||
**(b) Model selection alone (status quo)** — rejected: choosing Haiku for light tasks reduces cost, but within one model class there is no way to tune reasoning depth; a quality profile pays full reasoning cost even for scanning.
|
||||
|
||||
**(c) Static per-agent effort only** — rejected: loses context sensitivity; the same agent doing trivial vs complex work should not always get the same effort.
|
||||
|
||||
**(d) A separate `effort` field kept fully parallel to Codex's existing `reasoning_effort` (two independent lanes)** — rejected: produces two overlapping fields that can diverge and confuse; Codex's `reasoning_effort` is better modeled as one rendering of the single universal effort.
|
||||
|
||||
**(e) Overloading the existing `reasoning_effort` field to also carry Claude effort** — rejected: it would conflate a Codex-specific wire name with the universal concept and break the clean per-runtime rendering.
|
||||
|
||||
## References
|
||||
|
||||
- Tracking issue: #443
|
||||
- Prior art (inert reasoning_effort): #2517; `tests/issue-2517-runtime-aware-profiles.test.cjs`
|
||||
- dynamic_routing escalation: #3024; `tests/feat-3024-dynamic-routing.test.cjs`
|
||||
- phase-type tiers: #3023
|
||||
- Anthropic effort API: `output_config.effort` (`low`/`medium`/`high`/`xhigh`/`max`); fast mode: `speed` (`standard`/`fast`)
|
||||
- Claude Code effort: `/effort`, `--effort`, `CLAUDE_CODE_EFFORT_LEVEL`, `effortLevel` setting, subagent `effort` frontmatter (anthropics/claude-code #31536, completed); fast mode: `/fast` (interactive only)
|
||||
- OpenAI Codex effort: `model_reasoning_effort` config key; Responses API `reasoning.effort`; `ReasoningEffort` enum `none<minimal<low<medium<high<xhigh`
|
||||
Reference in New Issue
Block a user