Files
msd-core/docs/how-to/configure-model-profiles.md
Tom Boucher 455ad49ae3 feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded

Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.

Red until the resolver, CLI flag, and manifest key land.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#2296): config-gated provider escalation on quota-exceeded

The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.

- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
  capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
  Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
  the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
  validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
  passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
  every model tried once the ladder is spent.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2296): extract quota recovery to a reference fragment; regen goldens

The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.

That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.

Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2351): make the C1 orphan-reaping test load-independent

tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.

The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.

Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md

* chore(#2296): regenerate fixtures after rebase onto #2402

The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (b6e6a22fc) regenerated the same
artifacts. Conflict resolution picked a side to unblock the rebase; a true
regeneration on the combined tree then produced further drift, confirming the
resolved content was stale and would have dropped #2402's fixture changes.

Regenerated goldens, install-tree, size baseline, and INVENTORY-MANIFEST from
the merged tree. docs/INVENTORY.md keeps BOTH new reference rows.

execute-phase.md is 92782 LF bytes with both #2402's and this PR's extractions
applied — under the frozen ceiling (hard <93600, margin <=93400).

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 14:59:30 -04:00

9.3 KiB

How to configure model profiles

Choose the right model tier strategy for your project, then tune individual agents or entire phase types without writing a large override block. This guide starts with the simplest lever and works up to dynamic routing.


The four profiles (plus adaptive and inherit)

Set model_profile in .planning/config.json or via /gsd-config --profile <name>:

Profile Planner Executor Researchers Verifier Use when
quality Opus Opus Opus Sonnet Production-quality work where cost is secondary
balanced Opus Sonnet Sonnet Sonnet Normal development — the default
budget Sonnet Sonnet Haiku Haiku Rapid prototyping, cost-sensitive contexts
adaptive Opus Sonnet Sonnet Sonnet Resolves the same way as the other tiers under runtime-aware profiles; use when switching between runtimes frequently
inherit (session model) (session model) (session model) (session model) Non-Anthropic providers (OpenRouter, local models) — all agents follow your current session model

The table above shows a representative subset. All 33 shipped agents have explicit per-profile tier assignments in sdk/shared/model-catalog.json. For the full table see Model Profiles in the configuration reference.

Quick switch via command:

/gsd-config --profile balanced   # Normal development
/gsd-config --profile budget     # Prototyping or high-cost phases
/gsd-config --profile quality    # Production release
/gsd-config --profile inherit    # OpenRouter, local models

Or edit .planning/config.json directly:

{
  "model_profile": "balanced"
}

Per-agent overrides (model_overrides)

If a single agent needs a different tier without changing the whole profile, use model_overrides:

{
  "model_profile": "balanced",
  "model_overrides": {
    "gsd-executor": "opus",
    "gsd-codebase-mapper": "haiku"
  }
}

Valid values: opus, sonnet, haiku, inherit, or any fully-qualified model ID (e.g. "openai/o3", "google/gemini-2.5-pro").

model_overrides can be set per-project in .planning/config.json or globally in ~/.gsd/defaults.json. Per-project entries win on conflict; non-conflicting global entries are preserved.

Important for Codex and OpenCode: Those runtimes embed the resolved model into each agent's static config at install time. After editing model_overrides, re-run the installer for the change to take effect:

npx @opengsd/gsd-core@latest --codex --global   # or --opencode, --kilo, etc.

GSD will also warn you if you forget: workflow entry commands (gsd init plan-phase, gsd init execute-phase, etc.) detect when .planning/config.json or ~/.gsd/defaults.json is newer than your installed agent files and print a one-line stderr reminder naming the changed file and the re-install command. The check is read-only and runs only on codex and opencode; Claude Code resolves models at spawn time and is unaffected. (#1688)


Per-phase-type models (models)

If you want to say "Opus for planning, Sonnet for everything else" without learning all 33 agent names, use the models block. It maps six phase types to tier aliases:

{
  "model_profile": "balanced",
  "models": {
    "planning":      "opus",
    "discuss":       "opus",
    "research":      "sonnet",
    "execution":     "opus",
    "verification":  "sonnet",
    "completion":    "sonnet"
  }
}

Phase types and their agents:

Phase type Agents covered
planning gsd-planner, gsd-roadmapper, gsd-pattern-mapper
research gsd-phase-researcher, gsd-project-researcher, gsd-research-synthesizer, gsd-codebase-mapper, gsd-ui-researcher
execution gsd-executor, gsd-debugger, gsd-doc-writer
verification gsd-verifier, gsd-plan-checker, gsd-integration-checker, gsd-nyquist-auditor, gsd-ui-checker, gsd-ui-auditor, gsd-doc-verifier, gsd-code-reviewer
discuss gsd-assumptions-analyzer
completion Reserved — no subagent today; accepted by schema for forward compatibility

The models block accepts tier aliases only (opus, sonnet, haiku, inherit). For a fully-qualified model ID, use model_overrides per agent instead.

Combining models with a per-agent exception:

{
  "model_profile": "balanced",
  "models": {
    "research": "sonnet"
  },
  "model_overrides": {
    "gsd-codebase-mapper": "haiku"
  }
}

All five research agents resolve to sonnet except gsd-codebase-mapper, which is pinned to haiku.


Dynamic routing — start cheap, escalate on failure

If you want to pay for cheaper tiers by default and only escalate when an agent fails a quality gate, enable dynamic_routing:

{
  "dynamic_routing": {
    "enabled": true,
    "tier_models": {
      "light":    "haiku",
      "standard": "sonnet",
      "heavy":    "opus"
    },
    "escalate_on_failure": true,
    "max_escalations": 1
  }
}

Each agent has a default tier (light, standard, or heavy). On the first attempt, GSD picks tier_models[default_tier]. If the orchestrator detects a soft failure (verification inconclusive, plan-check flagged, etc.), it re-spawns the agent one tier up. max_escalations caps the total retries.

Agents that already sit at heavy cannot escalate further.

Turning off escalation while keeping dynamic resolution:

{
  "dynamic_routing": {
    "enabled": true,
    "escalate_on_failure": false
  }
}

Every attempt uses tier_models[default_tier] regardless of outcome — useful when you want explicit tier-to-model mapping without the escalation behaviour.

dynamic_routing is disabled by default. Omitting the block or setting enabled: false preserves static resolution.

Keep going when a provider throttles you

The tier ladder above escalates within one provider. When the provider itself is the thing that ran out of quota, a heavier tier on the same account is still throttled. Add provider_escalation — an ordered list of fallback model IDs — to keep the phase moving instead of stopping for a manual restart:

{
  "dynamic_routing": {
    "enabled": true,
    "tier_models": { "light": "haiku", "standard": "sonnet", "heavy": "opus" },
    "provider_escalation": ["gpt-5", "nvidia/llama-3.3"],
    "max_escalations": 2
  }
}

When an executor dies on a rate limit, GSD classifies the error body, switches to the next model in the list, logs the swap (sonnet → gpt-5), and waits out any Retry-After the provider sent. The walk is capped at min(max_escalations, provider_escalation.length). Once the list is spent, GSD names every model it tried and hands you the normal recovery prompt — it never silently retries the exhausted one.

This is most useful on providers without a guaranteed SLA (Nvidia NIM, OpenRouter, and other third-party OpenCode models), where a throttle mid-phase is routine. It only fires on quota / rate-limit failures; other failures keep the tier ladder. Leaving provider_escalation unset preserves the manual wait-for-reset behaviour exactly.


Using GSD on non-Anthropic runtimes

If you installed GSD for Codex, OpenCode, Antigravity CLI, or Kilo, the installer already set resolve_model_ids: "omit" in your config. This tells GSD to skip Anthropic model ID resolution and let the runtime choose its own default model. No manual setup is needed for the basic case.

If you want tiered models on Codex:

{
  "runtime": "codex",
  "model_profile": "balanced"
}

GSD resolves each tier alias to the Codex-native model and reasoning effort defined in the runtime tier map.

If you want per-agent model IDs on any non-Claude runtime:

{
  "resolve_model_ids": "omit",
  "model_overrides": {
    "gsd-planner":   "o3",
    "gsd-executor":  "o4-mini",
    "gsd-debugger":  "o3"
  }
}

For the full runtime-aware profiles reference and the model_policy surface (provider-neutral presets added in v1.42), see Configuration reference — Model Profiles.


Resolution precedence (highest to lowest)

When multiple layers apply, the resolver picks the highest-priority entry:

1. model_overrides[<agent>]           — per-agent; full IDs; targeted exception
2. dynamic_routing.tier_models[<tier>] — when enabled; escalates on soft failure
3. models[<phase_type>]               — coarse phase-level tier
4. model_profile (per-agent column)   — global tier strategy
5. Runtime default                    — when nothing else applies

Choosing the right lever

You want Use
One tier strategy for all agents model_profile
Coarse phase-level tuning ("Opus for planning") models.<phase_type>
Per-agent precision ("force Haiku on the codebase mapper") model_overrides[<agent>]
A fully-qualified model ID for a specific agent model_overrides[<agent>]: "openai/gpt-5"
Start cheap, escalate only on failure dynamic_routing
All agents follow the session model (non-Anthropic provider) model_profile: "inherit"