* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (b6e6a22fc) regenerated the same
artifacts. Conflict resolution picked a side to unblock the rebase; a true
regeneration on the combined tree then produced further drift, confirming the
resolved content was stale and would have dropped #2402's fixture changes.
Regenerated goldens, install-tree, size baseline, and INVENTORY-MANIFEST from
the merged tree. docs/INVENTORY.md keeps BOTH new reference rows.
execute-phase.md is 92782 LF bytes with both #2402's and this PR's extractions
applied — under the frozen ceiling (hard <93600, margin <=93400).
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
9.3 KiB
How to configure model profiles
Choose the right model tier strategy for your project, then tune individual agents or entire phase types without writing a large override block. This guide starts with the simplest lever and works up to dynamic routing.
The four profiles (plus adaptive and inherit)
Set model_profile in .planning/config.json or via /gsd-config --profile <name>:
| Profile | Planner | Executor | Researchers | Verifier | Use when |
|---|---|---|---|---|---|
quality |
Opus | Opus | Opus | Sonnet | Production-quality work where cost is secondary |
balanced |
Opus | Sonnet | Sonnet | Sonnet | Normal development — the default |
budget |
Sonnet | Sonnet | Haiku | Haiku | Rapid prototyping, cost-sensitive contexts |
adaptive |
Opus | Sonnet | Sonnet | Sonnet | Resolves the same way as the other tiers under runtime-aware profiles; use when switching between runtimes frequently |
inherit |
(session model) | (session model) | (session model) | (session model) | Non-Anthropic providers (OpenRouter, local models) — all agents follow your current session model |
The table above shows a representative subset. All 33 shipped agents have explicit per-profile tier assignments in sdk/shared/model-catalog.json. For the full table see Model Profiles in the configuration reference.
Quick switch via command:
/gsd-config --profile balanced # Normal development
/gsd-config --profile budget # Prototyping or high-cost phases
/gsd-config --profile quality # Production release
/gsd-config --profile inherit # OpenRouter, local models
Or edit .planning/config.json directly:
{
"model_profile": "balanced"
}
Per-agent overrides (model_overrides)
If a single agent needs a different tier without changing the whole profile, use model_overrides:
{
"model_profile": "balanced",
"model_overrides": {
"gsd-executor": "opus",
"gsd-codebase-mapper": "haiku"
}
}
Valid values: opus, sonnet, haiku, inherit, or any fully-qualified model ID (e.g. "openai/o3", "google/gemini-2.5-pro").
model_overrides can be set per-project in .planning/config.json or globally in ~/.gsd/defaults.json. Per-project entries win on conflict; non-conflicting global entries are preserved.
Important for Codex and OpenCode: Those runtimes embed the resolved model into each agent's static config at install time. After editing model_overrides, re-run the installer for the change to take effect:
npx @opengsd/gsd-core@latest --codex --global # or --opencode, --kilo, etc.
GSD will also warn you if you forget: workflow entry commands (gsd init plan-phase, gsd init execute-phase, etc.) detect when .planning/config.json or ~/.gsd/defaults.json is newer than your installed agent files and print a one-line stderr reminder naming the changed file and the re-install command. The check is read-only and runs only on codex and opencode; Claude Code resolves models at spawn time and is unaffected. (#1688)
Per-phase-type models (models)
If you want to say "Opus for planning, Sonnet for everything else" without learning all 33 agent names, use the models block. It maps six phase types to tier aliases:
{
"model_profile": "balanced",
"models": {
"planning": "opus",
"discuss": "opus",
"research": "sonnet",
"execution": "opus",
"verification": "sonnet",
"completion": "sonnet"
}
}
Phase types and their agents:
| Phase type | Agents covered |
|---|---|
planning |
gsd-planner, gsd-roadmapper, gsd-pattern-mapper |
research |
gsd-phase-researcher, gsd-project-researcher, gsd-research-synthesizer, gsd-codebase-mapper, gsd-ui-researcher |
execution |
gsd-executor, gsd-debugger, gsd-doc-writer |
verification |
gsd-verifier, gsd-plan-checker, gsd-integration-checker, gsd-nyquist-auditor, gsd-ui-checker, gsd-ui-auditor, gsd-doc-verifier, gsd-code-reviewer |
discuss |
gsd-assumptions-analyzer |
completion |
Reserved — no subagent today; accepted by schema for forward compatibility |
The models block accepts tier aliases only (opus, sonnet, haiku, inherit). For a fully-qualified model ID, use model_overrides per agent instead.
Combining models with a per-agent exception:
{
"model_profile": "balanced",
"models": {
"research": "sonnet"
},
"model_overrides": {
"gsd-codebase-mapper": "haiku"
}
}
All five research agents resolve to sonnet except gsd-codebase-mapper, which is pinned to haiku.
Dynamic routing — start cheap, escalate on failure
If you want to pay for cheaper tiers by default and only escalate when an agent fails a quality gate, enable dynamic_routing:
{
"dynamic_routing": {
"enabled": true,
"tier_models": {
"light": "haiku",
"standard": "sonnet",
"heavy": "opus"
},
"escalate_on_failure": true,
"max_escalations": 1
}
}
Each agent has a default tier (light, standard, or heavy). On the first attempt, GSD picks tier_models[default_tier]. If the orchestrator detects a soft failure (verification inconclusive, plan-check flagged, etc.), it re-spawns the agent one tier up. max_escalations caps the total retries.
Agents that already sit at heavy cannot escalate further.
Turning off escalation while keeping dynamic resolution:
{
"dynamic_routing": {
"enabled": true,
"escalate_on_failure": false
}
}
Every attempt uses tier_models[default_tier] regardless of outcome — useful when you want explicit tier-to-model mapping without the escalation behaviour.
dynamic_routing is disabled by default. Omitting the block or setting enabled: false preserves static resolution.
Keep going when a provider throttles you
The tier ladder above escalates within one provider. When the provider itself is the thing
that ran out of quota, a heavier tier on the same account is still throttled. Add
provider_escalation — an ordered list of fallback model IDs — to keep the phase moving
instead of stopping for a manual restart:
{
"dynamic_routing": {
"enabled": true,
"tier_models": { "light": "haiku", "standard": "sonnet", "heavy": "opus" },
"provider_escalation": ["gpt-5", "nvidia/llama-3.3"],
"max_escalations": 2
}
}
When an executor dies on a rate limit, GSD classifies the error body, switches to the next
model in the list, logs the swap (sonnet → gpt-5), and waits out any Retry-After the
provider sent. The walk is capped at min(max_escalations, provider_escalation.length).
Once the list is spent, GSD names every model it tried and hands you the normal recovery
prompt — it never silently retries the exhausted one.
This is most useful on providers without a guaranteed SLA (Nvidia NIM, OpenRouter, and
other third-party OpenCode models), where a throttle mid-phase is routine. It only fires on
quota / rate-limit failures; other failures keep the tier ladder. Leaving
provider_escalation unset preserves the manual wait-for-reset behaviour exactly.
Using GSD on non-Anthropic runtimes
If you installed GSD for Codex, OpenCode, Antigravity CLI, or Kilo, the installer already set resolve_model_ids: "omit" in your config. This tells GSD to skip Anthropic model ID resolution and let the runtime choose its own default model. No manual setup is needed for the basic case.
If you want tiered models on Codex:
{
"runtime": "codex",
"model_profile": "balanced"
}
GSD resolves each tier alias to the Codex-native model and reasoning effort defined in the runtime tier map.
If you want per-agent model IDs on any non-Claude runtime:
{
"resolve_model_ids": "omit",
"model_overrides": {
"gsd-planner": "o3",
"gsd-executor": "o4-mini",
"gsd-debugger": "o3"
}
}
For the full runtime-aware profiles reference and the model_policy surface (provider-neutral presets added in v1.42), see Configuration reference — Model Profiles.
Resolution precedence (highest to lowest)
When multiple layers apply, the resolver picks the highest-priority entry:
1. model_overrides[<agent>] — per-agent; full IDs; targeted exception
2. dynamic_routing.tier_models[<tier>] — when enabled; escalates on soft failure
3. models[<phase_type>] — coarse phase-level tier
4. model_profile (per-agent column) — global tier strategy
5. Runtime default — when nothing else applies
Choosing the right lever
| You want | Use |
|---|---|
| One tier strategy for all agents | model_profile |
| Coarse phase-level tuning ("Opus for planning") | models.<phase_type> |
| Per-agent precision ("force Haiku on the codebase mapper") | model_overrides[<agent>] |
| A fully-qualified model ID for a specific agent | model_overrides[<agent>]: "openai/gpt-5" |
| Start cheap, escalate only on failure | dynamic_routing |
| All agents follow the session model (non-Anthropic provider) | model_profile: "inherit" |