682eaae3f047280d69ce3da75fc35afafeca02ae
13 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b7cca0363f |
fix(#3531): merge routing_tier_defaults over manifest tier defaults (#3539)
* test(#3531): failing-first suite for routing_tier_defaults manifest merge * fix(#3531): merge routing_tier_defaults over manifest tier defaults * docs(#3531): document routing_tier_defaults merge-over-built-ins semantics * fix(#3531): correct test helper scope, update folded #443 expectations, guard merge keys * test(#3531): pin tiers in effort-sync and surface-axis fixtures post-merge * chore(#3531): backfill changeset pr number * fix(#3531): correct rebase resolution — keep both 3531 and 3533 test blocks intact * test(#3531): pin inherit/effort fixtures to the layer that reaches tiered agents --------- Co-authored-by: sim <sim@local> |
||
|
|
50d5368add | fix(#3533): effort inherit — expressible, omitted at writers, never re-added (#3541) | ||
|
|
2dbee3ebdd |
enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229. Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone. Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined. To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged. Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change). Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes. |
||
|
|
58d73dd220 |
enhance(#3241): omit the codex per-agent model by default (#3276)
* test(#3241): failing-first suite for the codex passive model posture Locks ADR-2313's D1-D5 before any production code exists, so the tests bind to the behavior rather than to whatever the implementation happens to do. Red-first (fail against the current tree): - the resolver path emits no `model` and no `model_reasoning_effort` - a whitespace-only model_overrides value yields no pin - isAnthropicFlavoredModel / CLAUDE_AGENT_ALIASES on model-catalog - the one-time install notice, and its once-per-install dedupe Regression guards (pass today, must keep passing): resolver-null via `inherit` and via absent runtime; a resolver that resolves to nothing; empty-string and non-string overrides; and the light-tier service_tier/model_verbosity fields, which are NOT coupled to the model pin and would silently regress if the implementation coupled them. Classifying each test as red-first or regression guard is deliberate. A test that passes on both sides of the change proves nothing, and this epic has already shipped two such rows before catching them. The whitespace case is a live defect, not a quirk: `' '` is truthy, survives the type guard, is not Anthropic-flavored, and is embedded verbatim as `model = " "` — the same class the #2310 guard exists to stop. Same function, same path, fixed in this phase per CLAUDE.md §3. Two matrix rows were dropped as vacuous rather than shipped green: a 64-char truncation case (the pinned notice interpolates no user-controlled value, so it cannot exhibit truncation) and a newline hazard that the input surface cannot reach. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3241): omit the codex per-agent model by default Implements ADR-2313 D1-D5. generateCodexAgentToml no longer embeds the runtime resolver's per-tier Codex model, so an agent inherits the always-available session model instead of a pin a ChatGPT-account Codex may not expose. model_reasoning_effort disappears with it via the existing hasPinnedModel coupling (#838) — no logic change needed there. Supersedes #2517's embedding on the default path only. An explicit real-Codex model_overrides pin is still embedded verbatim, and the #2310 Anthropic-flavored guard is retained: the model_overrides route to it is still live even though the resolver route is now unreachable. The shared rule moves down a layer. CLAUDE_AGENT_ALIASES leaves model-resolver for model-catalog — a genuine leaf importing only node:path and its own JSON — with isAnthropicFlavoredModel defined beside it, and is re-exported from model-resolver so every existing importer is untouched. This is what lets Phase 2's install-check and Phase 3's sync consume the rule without taking the config-loader dependency model-resolver would have dragged into a module documented as pure read/verify with 33 dependents. A parity test fails if the two ever fork. Also fixes a live defect surfaced while writing the tests: a whitespace-only model_overrides value was truthy, survived the type guard, was not Anthropic-flavored, and so was embedded verbatim as `model = " "` — the same class the #2310 guard exists to stop, reached by a different route. Trimmed before the truthiness test. It is deliberately not routed through _warnCodexModelOverrideDropped, whose text would misdescribe a blank field as a mis-typed model. Adds the one-time install notice (maintainer direction, recorded as an ADR-2313 amendment): one stderr line naming model_overrides and the session model, deduped per install rather than per agent, and emitted only for the population that actually loses a pin. service_tier and model_verbosity stay decoupled from the model (#774); a regression guard asserts they still emit with nothing pinned. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3241): amend ADR-2313, add the model-catalog glossary entry ADR-2313 gains two dated amendments rather than edits to its merged text, since ADRs here are append-only. The first records that a deprecation notice IS offered, reversing the Migration section's "no deprecation window" position, and states why that position was wrong rather than just superseding it: the ADR identified the API-key population as losing something real and then declined to warn it, in the same document. Hyrum's guidance was applied to the recourse and not to the notice. The second records the whitespace-only model_overrides defect and notes that D2 always implied the fix — the implementation simply never enforced it and no test covered the case. CONTEXT.md gains a Model Catalog Module entry. The module had none, which is why the glossary gate passed without one: check-glossary-refs verifies that references resolve, not that modules are documented. The entry records why the Anthropic-flavored rule lives there rather than in model-resolver, so a later reader does not "helpfully" move it back. The Model Resolver entry is updated to point at its new home and note the back-compat re-export. docs/CONFIGURATION.md carried a claim that is now false: that the resolved tier ID is embedded into agent frontmatter at install time on codex and opencode. Corrected to name codex as the exception, with the 400 symptom and the model_overrides recourse. Changeset leads with the user-visible change and the migration line rather than the implementation, per the ADR's Hyrum's-Law analysis. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): only notice a lost pin when one was actually embeddable Review finding from an isolated reviewer. The deprecation notice gated on whether the runtime resolver would have returned *any* model, but the question that matters is whether that model would have been *embedded*. Those differ. The #2310 safety gate already rejected an Anthropic- flavored model arriving from the resolver path before Phase 1 — so for a mixed-runtime config resolving to a claude-* id against a Codex install target, the user never had that pin. The notice told them they lost something they never got, and pointed them at model_overrides for no reason. The existing #2310 test drives exactly that path but asserts only the emitted `model` line, never stderr, which is why it slipped through. Now covered. Deliberately unchanged: an Anthropic-flavored model_overrides value plus a legal resolver model fires BOTH the override warning and the notice. That is correct — pre-Phase-1 the guard dropped the override, execution fell through to the resolver, and the resolver's model was embedded, so that user did lose a pin. Two messages, two distinct true facts, and the prefixes differ (`gsd: warning — ` vs `gsd: notice — `) so the one-notice-per-install contract holds. A regression test now pins that behavior so it does not get "simplified" away. Of the three tests added, only the first is red-first; the other two pass on both sides by design and are labelled as guards — one against over-correcting the fix into silence, one against removing the intentional double message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): reset the notice dedupe via a seam, not a require.cache bust The remote runner caught a regression I introduced: the #2760 post-write-validation test began failing with the validator override no longer intercepting. Cause, confirmed by trace rather than guessed: the new #3241 review tests deleted require.cache for bin/install.js and re-required it mid suite, to clear the notice's module-level dedupe flag. But runCodexInstall destructures `install` at file load, closing over the ORIGINAL module's exports. After the cache bust a second instance existed, so the test's `installModule.__codexSchemaValidator = ...` mutated the new object while the code under test still called the old one. The override silently stopped intercepting, the real validator ran and passed on GSD-emitted output, and the abort-and-restore path was never exercised. Cache-busting a module mid-suite breaks every later test that assumes a single instance, which every other test in the file is entitled to. So the fix is a seam, not a workaround: bin/install.js exports _resetCodexNoticeDedupeForTests(), and the three tests call it directly instead of reloading the module. The flag is module-level by design — the dedupe is per-install and install() already resets it — so a unit test driving generateCodexAgentToml directly needs an explicit way to reset it. That is now what it has. Swept the rest of the #3241 diff for the same hazard; this flag was the only shared module-level state introduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): reset both codex dedupe stores, not just the notice flag Second incomplete fix, same class one layer down. bin/install.js keeps TWO module-level dedupe stores and the require.cache bust I removed had been papering over both; my replacement seam cleared only one. _codexModelOverrideDroppedWarned is a Set keyed `${agent}::${value}`. tests/codex-config.test.cjs:558 already emits for `gsd-executor::sonnet`, so by the time the review test using the same agent and value ran, _warnCodexModelOverrideDropped was a silent no-op and the expected warning never appeared. The seam now clears both stores and is renamed to say so. Its comment records that per-install dedupe lives in module scope deliberately and that this is the single sanctioned way for a unit test to clear it. Swept bin/install.js for every other module-scope mutable a test could latch. Two are inert (capability registries assigned once at require time; selectedRuntimes computed once from argv). One is a genuine latent hazard and is deliberately NOT folded in: attributionCache (:1654) memoizes getCommitAttribution by runtime name for process lifetime, so two in-process installs of one runtime with differing attribution config would collide. It is unreachable from any current test and is a different concern from Codex warning dedupe, so it stays out of this PR rather than widening it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3241): correct the codex tier-routing how-to The docs gate forced the task-oriented quadrant and found the worst defect in this change's documentation surface. docs/how-to/configure-model-profiles.md carried a section titled "If you want tiered models on Codex" telling users to set runtime:codex + model_profile:balanced, promising "GSD resolves each tier alias to the Codex-native model and reasoning effort defined in the runtime tier map." That is exactly the behavior this PR removes — a how-to page confidently instructing users to do something that no longer works, which is worse than a missing page because it fails at the moment of use. Rewritten to state that Codex does no tier routing, give the model_overrides pin as the supported alternative, and name the two constraints on what may be pinned: it must be a real Codex model id, and the account must actually expose it — GSD cannot verify the second, so the honest advice when unsure is to omit the pin. Carries an upgrade note for both account types, since the change is a no-op for ChatGPT accounts and a real loss for API-key ones. Also tightened the same page's claim that Codex "embeds the resolved model" at install time — now true only of an explicit override. The re-install instruction it supports is still correct and still needed, so only the premise moved. Both the required-docs set (COMMANDS.md + FEATURES.md) and lint-docs-required.cjs would have passed before this commit, since CONFIGURATION.md and the ADR had already moved. Neither checks the quadrant a user in trouble actually opens. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3241): backfill changeset pr number (#3276) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
b2961c3f69 |
fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) (#2336)
* test(#2070): fail-first tests for adaptive model_profile and models tier validation Encodes the three acceptance criteria from #2070 plus the boundary cases the resolver silently ignores today (non-string values, empty string, mistyped phase-type key), and pins VALID_TIERS to a catalog-derived set. Red phase: these fail against current src/ by design. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) W004 sourced its profile list from a hand-maintained literal that predated the adaptive profile, so `"model_profile": "adaptive"` was false-flagged. It now reads VALID_PROFILES, which model-catalog.cts derives from model-catalog.json. models.<phase_type> was validated nowhere: the resolver's tier gate silently drops unknown values, so a typo like `"planning": "opuss"` was an undiagnosable no-op. A new W022 flags unknown phase-type keys and invalid tier values (including non-string values, which the same gate also drops). VALID_TIERS moves from a function-local literal in model-resolver.cts to a catalog-derived export, so health and the resolver cannot disagree by construction rather than by parity test. Object.values(adaptiveTierMap) is ['opus','sonnet','haiku'] plus 'inherit' — identical to the previous literal, so resolution behavior is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): changeset for validate health adaptive profile + W022 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): close review findings — malformed models, tier-list duplication, changeset gate Review of the initial fix surfaced three real defects, folded in per the no-defer rule: 1. verify.cts: the W022 guard skipped a top-level `models` that is present but not a plain object (`[]`, `"opus"`, `5`, `true`). The resolver ignores those identically, so they were the same undiagnosable no-op #2070 targets — just one level up. They now warn; absent/null/{} stay silent. 2. config-loader.cts: RUNTIME_OVERRIDE_TIERS was a second hardcoded copy of the tier vocabulary this change had just de-hardcoded elsewhere. It now derives from the catalog via ADAPTIVE_TIER_VALUES (no 'inherit' — runtime overrides resolve to a concrete tier). Byte-equivalent to the old literal. 3. scripts/changeset/lint.cjs: USER_FACING_PREFIXES omitted `src/`. Post-ADR-457 the product source is src/*.cts compiled to a gitignored gsd-core/bin/lib, so the `gsd-core/` prefix is dead coverage for library code and a src/-only PR could merge with no release note — including this one. Adding `src/` closes the gate; tests/ stays non-user-facing. Also corrects a false docstring in the VALID_TIERS test: value-equality cannot detect a re-hardcoded literal, so the test no longer claims it does. Two review findings were rejected with evidence rather than actioned: - W021 double-allocation is governed by ADR-612 ("W021 renumber -> void ... kept, message-disambiguated"), not a defect. - Global-defaults validation would be a false-positive generator: config-loader reads ~/.gsd/defaults.json only on the "no .planning/" branch, and health early-returns E001 without .planning/, so those values provably never affect resolution in any context health can run. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2070): regenerate install goldens for the changeset-lint change scripts/ ships as an installed artifact, so scripts/changeset/lint.cjs's content hash is pinned in all 18 runtime golden fixtures. Adding 'src/' to USER_FACING_PREFIXES changed that hash and tripped every golden parity check. Regenerated via `npm run gen:golden`; the only delta is the lint.cjs hash. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): backfill PR number 2336 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9ad2bab4be |
fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime (#2332)
* fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime The installer writes resolve_model_ids:"omit" for non-alias runtimes into the machine-wide ~/.gsd/defaults.json (#1156); any runtime read it back, so install order silently flipped Claude's adaptive tier aliases (executor->sonnet, planner->opus) to '' in no-project sessions. Resolution is now scoped to the runtime actually resolving, identified by a new per-install <install>/gsd-core/.gsd-runtime marker (installer writes it beside VERSION). The "omit" branch returns '' only when the PROJECT explicitly set omit (honored for all runtimes, #2517 finding #4) OR the active runtime lacks native aliases. Claude ignores a global-defaults-only omit and keeps its aliases; the active runtime is canonicalized (GSD_RUNTIME -> config.runtime -> marker -> claude) so alias/case spellings can't defeat the check; explicit project omit is workstream/project-scope aware; explicit true still materializes IDs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2297): backfill PR number 2332 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a22333034b |
fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml (#2312)
* fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml generateCodexAgentToml embedded a per-agent `model_overrides` value verbatim as the Codex `.toml` `model`, leaking GSD/Claude tier aliases (opus/sonnet/haiku/fable) and `claude-*` ids. Codex/ChatGPT rejects those (400 "The 'sonnet' model is not supported when using Codex with a ChatGPT account"), and since spawn_agent has no inline model param, the model is baked into the .toml at install time — so the orchestrator could not recover and fell back to the non-equivalent generic-agent workaround. Translate a GSD tier alias through the Codex tier map (sonnet -> gpt-5.6-terra); drop with a deduped warning any Anthropic-flavored value with no Codex mapping (fable) or a `claude-*` id, so emission falls through to the runtime-aware resolver or Codex's default. A final safety gate blocks an Anthropic-flavored model from the runtime- resolver path too (runtime/target mismatch). Mirrors the Claude-side override guard (#2041). Real Codex/OpenAI model ids in model_overrides still pass through verbatim (#2256 preserved); runtime:"codex" tier resolution unchanged (#2517). Adds regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2310): backfill changeset PR number to #2312 * fix(#2310): Codex passive-model posture — omit Anthropic-flavored model (all namespacings) Adopt ADR-1239's passive/session-only posture for Codex model handling: a Codex agent .toml `model` is embedded ONLY for an explicit real-Codex model_overrides pin; any Anthropic-flavored value is omitted so the agent inherits the always- available session model (never a 400). - model_overrides tier alias (opus/sonnet/haiku/fable) or a Claude model id → omit (was: translate to gpt-*); an explicit real-Codex model id → embed verbatim (#2256 preserved). - Detect ALL Anthropic namespacings, not just `claude-*`: single-source the canonical CLAUDE_AGENT_ALIASES from model-resolver.cts and treat any id whose value contains "claude" (case-insensitive) as Anthropic-flavored — catching `anthropic/claude-*` and `us.anthropic.claude-*` (the forms the catalog assigns to opencode/hermes/kilo), which reach a Codex .toml via the runtime-resolver path on a mixed-runtime + Codex install. - The final safety gate applies to the runtime-resolver path too. The full passive posture (removing #2517's runtime-resolver per-tier embedding + a correctness health-check + a Codex TOML sync path) is tracked as the ADR-2310 epic #2313. Regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e95af39a8c |
fix(#2041): address code+security review findings
- add typeof guard so a non-string override passes through verbatim instead of crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1] - use Object.hasOwn() for the alias lookup so __proto__/constructor cannot return a truthy non-string from the plain object literal [LOW-D3] - cap the unmappable-override stderr warning at 64 chars so an oversized or secret-shaped value cannot leak in full to stderr/logs [LOW-D4] - remove the unused mapClaudeOverrideForRuntime export (helpers are covered behaviourally via resolveModelInternal/resolveModelForTier) [NIT] - add resolveModelForTier unmappable-override fall-through test (closes the mutation-score gap) [MEDIUM-1] - add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2] Both orthogonal reviews returned APPROVE with no Critical/High findings. |
||
|
|
f214f1320d |
fix(#2041): map model_overrides full claude IDs to agent-tool aliases
model_overrides values that are full Claude model IDs (claude-sonnet-5, claude-opus-4-8, claude-haiku-4-5, claude-fable-5) were returned verbatim on the claude runtime and handed to the Claude Agent tool, whose typed model parameter documents only tier aliases (opus/sonnet/haiku/fable). The model_policy path already mapped full IDs -> aliases via CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides skipped that mapping, so the two resolver paths produced different shapes for the same underlying Claude model. The fix mirrors #1144 on the override path via a shared mapClaudeOverrideForRuntime helper used by both resolveModelInternal and resolveModelForTier. Bare aliases pass through verbatim; non-Claude runtimes and non-Claude custom/vendor values keep full IDs verbatim (parity). An unmappable Claude ID (e.g. claude-opus-4-5) warns once to stderr and falls through to tier resolution, exactly as the model_policy path already does. Alias mapping is also the documented best practice (prevents staleness when new model versions ship). |
||
|
|
8c3d934a90 |
refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) (#1295)
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) After T0–T6 nothing imports core, so retire the spine and its scaffolding: - delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact; remove its .gitignore + eslint-ignore entries) - delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the package.json lint:ci chain - regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface) - sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired, callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner, and false present-tense core.cjs claims in leaf-module docstrings The ADR-857 decomposition is complete: the former Core god-module is fully dissolved into its leaf modules; no re-export spine remains. No behaviour change. Closes #1294 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1294): migrate the computed-path core.cjs importers the literal grep missed bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed path, and bin/install.js was never in the convergence lint's scan roots), and ~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG forms the literal-string migration grep missed. Route install.js's symbols to their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET-> model-resolver) and repoint/adjust the test references to the leaves. Recovers the 161 'Cannot find module core.cjs' failures from the spine deletion. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
44024aa535 |
fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto next. The hotfix was authored against src/core.cts (v1.4.4); on next the resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888), so the patch is re-applied there rather than cherry-picked. resolveModelInternal step 2.5 now honors model_policy on the claude runtime: the policy-resolved full model ID is mapped back to a Claude Code agent alias via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 -> fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no Claude alias warns once to stderr (deduped by agentType::policyModel::tier) and falls back to the configured tier alias. Non-claude runtimes return full IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged. The warn-dedupe cache lives in model-resolver.cts; core.cts composes the exported _resetRuntimeWarningCacheForTests to clear both that cache and the config-loader warning cache (config-loader cannot import model-resolver -- circular dependency). Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path. Forward-port of #1133 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
185935379a |
refactor(#888): extract model+effort resolution into model-resolver.cts (#890)
ADR-857 rollout phase 2f — the FINAL core.cts decomposition. Move the model and effort resolution cluster (resolveModelInternal, resolveModelPolicy, resolveTierEntry, _resolveRuntimeTier, resolveModelForTier, resolveGranularityInternal, assertValidGranularityOverride, resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, nextEffort + VALID_GRANULARITIES/ VALID_EFFORTS/EFFORT_SET + interfaces) out of core.cts into a new leaf module src/model-resolver.cts. core.cts re-exports the 13 public symbols (callers in init/docs/commands unchanged; export= set byte-identical). Cycle-free: model-resolver imports only leaves (config-loader for loadConfig, configuration for defaults, model-profiles + model-catalog for the static tables). Removed 6 now-unused imports from core (verified zero remaining references, none re-exported). This completes the god-module decomposition: core.cts 2271 -> 389 lines (~83%), now a thin re-export spine over seven clean leaves (io, phase-id, roadmap-parser, core-utils, phase-locator, config-loader, model-resolver). New-CLI-module checklist done (.gitignore, eslint, INVENTORY 96->97 + row, manifest, ARCHITECTURE, CONTEXT.md "Model Resolver Module"). Adds tests/model-resolver.test.cjs (81 tests: behavioral + shim-identity + adversarial). Gates: lint, code-review (export set byte-identical; import-removal verified), security-review, codex adversarial-review (all 0 findings; verbatim move). Mac 4303 pass; clean-build docker 13117 pass, 0 fail. Closes #888 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |