2a73f53cb3d9f5515e5145aecd4aea19f47211f8
9 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
b2961c3f69 |
fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) (#2336)
* test(#2070): fail-first tests for adaptive model_profile and models tier validation Encodes the three acceptance criteria from #2070 plus the boundary cases the resolver silently ignores today (non-string values, empty string, mistyped phase-type key), and pins VALID_TIERS to a catalog-derived set. Red phase: these fail against current src/ by design. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) W004 sourced its profile list from a hand-maintained literal that predated the adaptive profile, so `"model_profile": "adaptive"` was false-flagged. It now reads VALID_PROFILES, which model-catalog.cts derives from model-catalog.json. models.<phase_type> was validated nowhere: the resolver's tier gate silently drops unknown values, so a typo like `"planning": "opuss"` was an undiagnosable no-op. A new W022 flags unknown phase-type keys and invalid tier values (including non-string values, which the same gate also drops). VALID_TIERS moves from a function-local literal in model-resolver.cts to a catalog-derived export, so health and the resolver cannot disagree by construction rather than by parity test. Object.values(adaptiveTierMap) is ['opus','sonnet','haiku'] plus 'inherit' — identical to the previous literal, so resolution behavior is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): changeset for validate health adaptive profile + W022 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): close review findings — malformed models, tier-list duplication, changeset gate Review of the initial fix surfaced three real defects, folded in per the no-defer rule: 1. verify.cts: the W022 guard skipped a top-level `models` that is present but not a plain object (`[]`, `"opus"`, `5`, `true`). The resolver ignores those identically, so they were the same undiagnosable no-op #2070 targets — just one level up. They now warn; absent/null/{} stay silent. 2. config-loader.cts: RUNTIME_OVERRIDE_TIERS was a second hardcoded copy of the tier vocabulary this change had just de-hardcoded elsewhere. It now derives from the catalog via ADAPTIVE_TIER_VALUES (no 'inherit' — runtime overrides resolve to a concrete tier). Byte-equivalent to the old literal. 3. scripts/changeset/lint.cjs: USER_FACING_PREFIXES omitted `src/`. Post-ADR-457 the product source is src/*.cts compiled to a gitignored gsd-core/bin/lib, so the `gsd-core/` prefix is dead coverage for library code and a src/-only PR could merge with no release note — including this one. Adding `src/` closes the gate; tests/ stays non-user-facing. Also corrects a false docstring in the VALID_TIERS test: value-equality cannot detect a re-hardcoded literal, so the test no longer claims it does. Two review findings were rejected with evidence rather than actioned: - W021 double-allocation is governed by ADR-612 ("W021 renumber -> void ... kept, message-disambiguated"), not a defect. - Global-defaults validation would be a false-positive generator: config-loader reads ~/.gsd/defaults.json only on the "no .planning/" branch, and health early-returns E001 without .planning/, so those values provably never affect resolution in any context health can run. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2070): regenerate install goldens for the changeset-lint change scripts/ ships as an installed artifact, so scripts/changeset/lint.cjs's content hash is pinned in all 18 runtime golden fixtures. Adding 'src/' to USER_FACING_PREFIXES changed that hash and tripped every golden parity check. Regenerated via `npm run gen:golden`; the only delta is the lint.cjs hash. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): backfill PR number 2336 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9ad2bab4be |
fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime (#2332)
* fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime The installer writes resolve_model_ids:"omit" for non-alias runtimes into the machine-wide ~/.gsd/defaults.json (#1156); any runtime read it back, so install order silently flipped Claude's adaptive tier aliases (executor->sonnet, planner->opus) to '' in no-project sessions. Resolution is now scoped to the runtime actually resolving, identified by a new per-install <install>/gsd-core/.gsd-runtime marker (installer writes it beside VERSION). The "omit" branch returns '' only when the PROJECT explicitly set omit (honored for all runtimes, #2517 finding #4) OR the active runtime lacks native aliases. Claude ignores a global-defaults-only omit and keeps its aliases; the active runtime is canonicalized (GSD_RUNTIME -> config.runtime -> marker -> claude) so alias/case spellings can't defeat the check; explicit project omit is workstream/project-scope aware; explicit true still materializes IDs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2297): backfill PR number 2332 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a22333034b |
fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml (#2312)
* fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml generateCodexAgentToml embedded a per-agent `model_overrides` value verbatim as the Codex `.toml` `model`, leaking GSD/Claude tier aliases (opus/sonnet/haiku/fable) and `claude-*` ids. Codex/ChatGPT rejects those (400 "The 'sonnet' model is not supported when using Codex with a ChatGPT account"), and since spawn_agent has no inline model param, the model is baked into the .toml at install time — so the orchestrator could not recover and fell back to the non-equivalent generic-agent workaround. Translate a GSD tier alias through the Codex tier map (sonnet -> gpt-5.6-terra); drop with a deduped warning any Anthropic-flavored value with no Codex mapping (fable) or a `claude-*` id, so emission falls through to the runtime-aware resolver or Codex's default. A final safety gate blocks an Anthropic-flavored model from the runtime- resolver path too (runtime/target mismatch). Mirrors the Claude-side override guard (#2041). Real Codex/OpenAI model ids in model_overrides still pass through verbatim (#2256 preserved); runtime:"codex" tier resolution unchanged (#2517). Adds regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2310): backfill changeset PR number to #2312 * fix(#2310): Codex passive-model posture — omit Anthropic-flavored model (all namespacings) Adopt ADR-1239's passive/session-only posture for Codex model handling: a Codex agent .toml `model` is embedded ONLY for an explicit real-Codex model_overrides pin; any Anthropic-flavored value is omitted so the agent inherits the always- available session model (never a 400). - model_overrides tier alias (opus/sonnet/haiku/fable) or a Claude model id → omit (was: translate to gpt-*); an explicit real-Codex model id → embed verbatim (#2256 preserved). - Detect ALL Anthropic namespacings, not just `claude-*`: single-source the canonical CLAUDE_AGENT_ALIASES from model-resolver.cts and treat any id whose value contains "claude" (case-insensitive) as Anthropic-flavored — catching `anthropic/claude-*` and `us.anthropic.claude-*` (the forms the catalog assigns to opencode/hermes/kilo), which reach a Codex .toml via the runtime-resolver path on a mixed-runtime + Codex install. - The final safety gate applies to the runtime-resolver path too. The full passive posture (removing #2517's runtime-resolver per-tier embedding + a correctness health-check + a Codex TOML sync path) is tracked as the ADR-2310 epic #2313. Regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e95af39a8c |
fix(#2041): address code+security review findings
- add typeof guard so a non-string override passes through verbatim instead of crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1] - use Object.hasOwn() for the alias lookup so __proto__/constructor cannot return a truthy non-string from the plain object literal [LOW-D3] - cap the unmappable-override stderr warning at 64 chars so an oversized or secret-shaped value cannot leak in full to stderr/logs [LOW-D4] - remove the unused mapClaudeOverrideForRuntime export (helpers are covered behaviourally via resolveModelInternal/resolveModelForTier) [NIT] - add resolveModelForTier unmappable-override fall-through test (closes the mutation-score gap) [MEDIUM-1] - add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2] Both orthogonal reviews returned APPROVE with no Critical/High findings. |
||
|
|
f214f1320d |
fix(#2041): map model_overrides full claude IDs to agent-tool aliases
model_overrides values that are full Claude model IDs (claude-sonnet-5, claude-opus-4-8, claude-haiku-4-5, claude-fable-5) were returned verbatim on the claude runtime and handed to the Claude Agent tool, whose typed model parameter documents only tier aliases (opus/sonnet/haiku/fable). The model_policy path already mapped full IDs -> aliases via CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides skipped that mapping, so the two resolver paths produced different shapes for the same underlying Claude model. The fix mirrors #1144 on the override path via a shared mapClaudeOverrideForRuntime helper used by both resolveModelInternal and resolveModelForTier. Bare aliases pass through verbatim; non-Claude runtimes and non-Claude custom/vendor values keep full IDs verbatim (parity). An unmappable Claude ID (e.g. claude-opus-4-5) warns once to stderr and falls through to tier resolution, exactly as the model_policy path already does. Alias mapping is also the documented best practice (prevents staleness when new model versions ship). |
||
|
|
8c3d934a90 |
refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) (#1295)
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) After T0–T6 nothing imports core, so retire the spine and its scaffolding: - delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact; remove its .gitignore + eslint-ignore entries) - delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the package.json lint:ci chain - regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface) - sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired, callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner, and false present-tense core.cjs claims in leaf-module docstrings The ADR-857 decomposition is complete: the former Core god-module is fully dissolved into its leaf modules; no re-export spine remains. No behaviour change. Closes #1294 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1294): migrate the computed-path core.cjs importers the literal grep missed bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed path, and bin/install.js was never in the convergence lint's scan roots), and ~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG forms the literal-string migration grep missed. Route install.js's symbols to their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET-> model-resolver) and repoint/adjust the test references to the leaves. Recovers the 161 'Cannot find module core.cjs' failures from the spine deletion. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
44024aa535 |
fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto next. The hotfix was authored against src/core.cts (v1.4.4); on next the resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888), so the patch is re-applied there rather than cherry-picked. resolveModelInternal step 2.5 now honors model_policy on the claude runtime: the policy-resolved full model ID is mapped back to a Claude Code agent alias via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 -> fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no Claude alias warns once to stderr (deduped by agentType::policyModel::tier) and falls back to the configured tier alias. Non-claude runtimes return full IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged. The warn-dedupe cache lives in model-resolver.cts; core.cts composes the exported _resetRuntimeWarningCacheForTests to clear both that cache and the config-loader warning cache (config-loader cannot import model-resolver -- circular dependency). Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path. Forward-port of #1133 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
185935379a |
refactor(#888): extract model+effort resolution into model-resolver.cts (#890)
ADR-857 rollout phase 2f — the FINAL core.cts decomposition. Move the model and effort resolution cluster (resolveModelInternal, resolveModelPolicy, resolveTierEntry, _resolveRuntimeTier, resolveModelForTier, resolveGranularityInternal, assertValidGranularityOverride, resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, nextEffort + VALID_GRANULARITIES/ VALID_EFFORTS/EFFORT_SET + interfaces) out of core.cts into a new leaf module src/model-resolver.cts. core.cts re-exports the 13 public symbols (callers in init/docs/commands unchanged; export= set byte-identical). Cycle-free: model-resolver imports only leaves (config-loader for loadConfig, configuration for defaults, model-profiles + model-catalog for the static tables). Removed 6 now-unused imports from core (verified zero remaining references, none re-exported). This completes the god-module decomposition: core.cts 2271 -> 389 lines (~83%), now a thin re-export spine over seven clean leaves (io, phase-id, roadmap-parser, core-utils, phase-locator, config-loader, model-resolver). New-CLI-module checklist done (.gitignore, eslint, INVENTORY 96->97 + row, manifest, ARCHITECTURE, CONTEXT.md "Model Resolver Module"). Adds tests/model-resolver.test.cjs (81 tests: behavioral + shim-identity + adversarial). Gates: lint, code-review (export set byte-identical; import-removal verified), security-review, codex adversarial-review (all 0 findings; verbatim move). Mac 4303 pass; clean-build docker 13117 pass, 0 fail. Closes #888 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |