Commit Graph

4295 Commits

Author SHA1 Message Date
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
cce405c9b7 Merge pull request #2058 from open-gsd/fix/2046-config-unset-null
fix(#2046): config-set <key> null unsets (clears) the key instead of persisting "null"
2026-07-07 13:47:59 -04:00
Tom Boucher
3742270c87 Merge branch 'next' into fix/2046-config-unset-null 2026-07-07 13:26:16 -04:00
Tom Boucher
bc56e26ca3 Merge pull request #2059 from open-gsd/fix/2043-phase-token-single-digit-slug
fix(#2043): reject single-digit slug word in phase-token extraction (all sites)
2026-07-07 13:25:54 -04:00
Tom Boucher
9accce2b72 Merge branch 'next' into fix/2043-phase-token-single-digit-slug 2026-07-07 13:13:31 -04:00
Tom Boucher
9a9f304690 Merge pull request #2060 from open-gsd/fix/1857-test-gate-watch-mode-timeout
fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
2026-07-07 13:13:12 -04:00
Tom Boucher
12cc1955b2 Merge branch 'next' into fix/1857-test-gate-watch-mode-timeout 2026-07-07 12:58:57 -04:00
Tom Boucher
a6922c6d52 Merge pull request #2057 from open-gsd/fix/1821-kilo-dead-hook-copy
fix(#1821): stop copying dead hook scripts for Kilo and ZCode
2026-07-07 12:58:37 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
6efd2b4016 fix(#2043): reject single-digit slug word in phase-token extraction (all sites)
A phase whose slug's first word is a bare single digit (dir
`46-6-rs-pipeline-orchestrator`, roadmap phase "6 Rs Pipeline Orchestrator" →
slug `6-rs-…`) had its phase token over-collected as `46-6` instead of `46`, so
`gsd-tools` phase-by-number lookups (init.plan-phase, init.phase-op, and
downstream execute/verify/ship) resolved phase_dir=null / has_context=false.

Root cause: an over-broad "looks numeric" test (`/^\d/`, `\d+(?:-\d+)*`)
classified token segments and could not distinguish a legitimate zero-padded
sub-phase segment (`01`, `02`) from a single-digit slug word (`6`). Zero-padded
phase/sub-phase segments are always ≥2 digits, so requiring ≥2 digits is the
structural distinguisher. Applied consistently across every same-class
implementation the triage identified (fixing extractPhaseToken alone leaves the
health-check and milestone-filter subsystems exposed):

- src/phase-id.cts extractPhaseToken — a pure-numeric leading segment continues
  only with ≥2-digit segments; a letter-prefixed milestone id (`M1`) still
  admits its single-digit sub-phase (`M1-2`).
- src/validate.cts PHASE_TOKEN_FROM_DIR_RE (W005/W006/W007 health checks) and
  canonicalPlanStem (I001 plan/summary pairing).
- src/roadmap-parser.cts isDirInMilestone numericRe (getMilestonePhaseFilter —
  highest blast radius: a false negative silently excludes a phase from the
  milestone).
- src/core-utils.cts + its verbatim duplicate src/phase.cts extractCanonicalPlanId.

Regression tests added across tests/{phase-id,health-validation,core-utils,
roadmap-parser}.test.cjs using the shared `46-6-rs-…` fixture, asserting the
token resolves to `46` (not `46-6`) while legit multi-segment tokens
(`01-02`, `02-03-04`, `M1-2`, `68-01`, `02-01`) are unchanged. The
roadmap-parser test exercises the real getMilestonePhaseFilter end-to-end.

The residual ≥2-digit-leading-slug case (e.g. `NN-2024-roadmap`) stays out of
scope — subsumed by the #612 / #565 phase-ID convention work, per the issue.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 10:15:37 -04:00
Tom Boucher
dd21bb7451 fix(#2046): config-set <key> null unsets (removes) the key instead of writing "null"
`gsd-tools config-set <key> null` — the "Clear" action documented in
settings-integrations.md / settings-advanced.md — previously fell through the
value parser as the literal STRING "null" and persisted it. Consequences:
"cleared" keys stayed set (config-get returned truthy "null"), and for secret
keys (brave_search/firecrawl/exa_search) a masked success line hid a truthy
4-char value on disk that integrations could pass along as a real credential.
There was also no unset/delete verb at all.

Fix: parse a bare `null` to JS null and short-circuit to a real UNSET that
DELETES the key from config.json — the semantic the docs already describe
("Remove the stored key" / "remove the key by setting it to null"). Deleting
(not persisting JSON null) is the correct clear: a persisted null is still a
present value consumers must special-case.

- src/config.cts:
  - parse block: `else if (val === 'null') parsedValue = null;`
  - cmdConfigSet: when parsedValue === null, short-circuit BEFORE the typed
    per-key validator gauntlet (so clearing an enum/boolean/number key removes
    it rather than being rejected) and before the project_code special-case;
    mask the previous value for secret keys in the output.
  - new `unsetConfigValue()` + `_unsetNestedValue()` mirroring setConfigValue/
    _setNestedValue: same prototype-pollution guard, but never creates missing
    intermediates and never prunes empty parents; returns { previousValue,
    existed }. Unsetting a never-set key is an idempotent no-op success.
- tests/config.test.cjs: new suite covering non-secret routing key, secret key,
  typed-enum-key bypass (context), idempotent unset, literal-"null"-on-disk
  guard, the unset-path prototype-pollution guard (alert #26 parity), and a
  4-segment deep-nested unset.
- tests/review-model-config.test.cjs: update the stale round-trip test that
  codified the bug (asserted config-set null → config-get returns "null") to the
  fixed contract — the model key is removed; the review workflow's
  `[ -n "$VAR" ] && [ "$VAR" != "null" ]` guard handles the empty read as
  "no override → reviewer default", same as the old "null" sentinel.

The 4 documented "Clear" flows (settings-integrations.md, settings-advanced.md)
were verified — their prose already describes removal, so the fix makes them
accurate rather than aspirational; no doc wording change required.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 09:41:14 -04:00
Tom Boucher
2b806904e2 fix(#1821): stop copying dead hook scripts for Kilo and ZCode (hooksSurface:none)
Kilo and ZCode both declare `hooksSurface: 'none'` and have no plugin surface,
so the GSD installer staged lifecycle hook scripts (hooks/*.js, hooks/*.sh,
hooks/lib/) plus a `{"type":"commonjs"}` package.json marker into their config
dirs where nothing ever invokes them — dead weight (#1821).

The installer's two hook-copy guards at bin/install.js were still on the legacy
hardcoded runtime-name list and never excluded Kilo or ZCode. Add
`&& !isKilo && !isZcode` to both (isZcode added to the install() runtimeFlags
destructure).

OpenCode — which #1821 also named — is deliberately NOT excluded: since the
issue was filed, #1914 shipped a native OpenCode plugin (plugins/gsd-core.js)
that spawns those exact staged hooks via OpenCode's event bus and requires both
the hook scripts and the CommonJS package.json marker. Excluding OpenCode would
regress #1914, so its hooks stay live. The genuinely-dead cases are Kilo & ZCode.

- bin/install.js: add `&& !isKilo && !isZcode` to the hooks/dist copy guard and
  the hooks/lib copy guard; document the OpenCode-vs-Kilo/ZCode split.
- tests/install-minimal-hooks.test.cjs: regression test asserting Kilo and ZCode
  receive no gsd-*.js/.sh hooks or hooks/lib, while OpenCode keeps its hooks +
  #1914 plugin and Claude keeps its hooks (over-exclusion guard).
- tests/fixtures/golden-install-parity/{kilo,zcode}.json: drop the 21 hooks/*
  entries and the package.json marker they no longer receive (opencode unchanged).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 09:40:01 -04:00
Tom Boucher
a3b9cbaeb8 Merge pull request #2054 from open-gsd/fix/2045-third-party-skills-surface
fix(#2045): third-party capability skills surface correctly
2026-07-07 08:38:23 -04:00
Tom Boucher
847de596b8 fix: third-party capability skills surface correctly (#2045)
A skills-only role:feature third-party capability installed 'active' but its skills never reached the runtime surface, capability enable/set rejected it as 'unknown capability', and capability list disagreed with capability state. Three defects, fix shape 1b (teach resolveSurface, no on-disk linking):

D1 (materialization): resolveSurface built the surfaced skills Set only from the on-disk manifest; third-party cap skills live at ~/.gsd/capabilities/<id>/skills/ and never entered the Set -> surfaced:false. Fix: union registry.capabilityClusters values into the Set in the full-profile branch (idempotent for first-party, additive for third-party; prototype-pollution guarded).

D2 (enable/set unknown): setCapabilityState validated the capId against the static first-party registry instead of the composed overlay-aware registry. Fix: validate against loadRegistry({includeInstalled}) like capability-state.cts does.

D3 (list vs state): capability list derived status purely from ledger existence, never surface composition. Fix: add a 'surfaced' field to list rows sourced from the same resolver capability state uses, so list and state agree.

Regression coverage: tests/issue-2045-third-party-skills-surface.test.cjs asserts all five acceptance criteria (resolveSurface union, state surfaced:true, enable/set not-unknown, list/state agreement, first-party + unknown-id regression). gsd-test: 23916/23916 green (linux-node22 + linux-node24).
2026-07-07 08:25:33 -04:00
Tom Boucher
2372a221eb Merge pull request #1835 from open-gsd/feat/1820-specless-predicate-rail
feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them
2026-07-07 08:05:30 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
Tom Boucher
d2ae7fe391 Merge pull request #2053 from open-gsd/chore/sync-next-version-1.7.0-rc.4
chore: sync next package version to 1.7.0-rc.4
2026-07-07 02:08:34 -04:00
github-actions[bot]
a6263bba8a chore: sync next package version to 1.7.0-rc.4 2026-07-07 06:08:27 +00:00
Tom Boucher
ed2afd6204 Merge pull request #2051 from open-gsd/fix/2003-capability-state-runtime-flag
fix(#2003): add --runtime override to capability state + loop render-hooks
2026-07-07 01:40:36 -04:00
Tom Boucher
0d7c15badc Merge branch 'next' into fix/2003-capability-state-runtime-flag 2026-07-07 00:24:27 -04:00
Tom Boucher
29c3ed102c Merge pull request #1994 from open-gsd/codex/gsd-onboard
feat(#1990): add brownfield onboarding workflow
2026-07-07 00:23:55 -04:00
Tom Boucher
8a935a08a3 Merge branch 'next' into fix/2003-capability-state-runtime-flag 2026-07-07 00:19:08 -04:00
Tom Boucher
bc751a64ec docs(#1990): point adr index at renamed file 2026-07-07 00:01:57 -04:00
Tom Boucher
1171499f38 docs(#1990): remove pre-rename ADR filename 2026-07-07 00:01:56 -04:00
Tom Boucher
d0b8eacd3c docs(#1990): rename ADR to Existing Code Onboarding 2026-07-07 00:01:55 -04:00
Tom Boucher
e8fb05e965 docs(#1990): index ADR-1990 in adr README 2026-07-06 23:58:50 -04:00
Tom Boucher
3c7d722ed9 docs(#1990): add ADR-1990 onboard projection module 2026-07-06 23:58:27 -04:00
Tom Boucher
d7129222c0 Merge branch 'next' into codex/gsd-onboard 2026-07-06 23:49:49 -04:00
Tom Boucher
cc0372007a Merge pull request #1991 from jslitzkerttcu/fix/1941-quick-worktree-stale-base
fix(#1941): degrade /gsd-quick worktree dispatch when fork base is stale
2026-07-06 23:49:25 -04:00
Tom Boucher
579ad30eae docs(#2003): backfill changeset pr number to 2051 2026-07-06 23:44:25 -04:00
Tom Boucher
9b34fd5c08 Merge branch 'next' into fix/1941-quick-worktree-stale-base 2026-07-06 23:37:23 -04:00
Tom Boucher
327b6409e8 fix(#2003): address code+security review findings
- warn (don't silently ignore) when --runtime is an unknown runtime that
  canonicalizeRuntimeName rejects; the warning surfaces via warnings[] so a
  typo like --runtime cluade or a runtime known to runtime-homes but not the
  alias manifest (e.g. grok) no longer silently resolves to the persisted
  runtime's config dir on this diagnostic command [M-1]
- add end-to-end CLI test for loop render-hooks --runtime (the exact command
  the bug report calls out as silently no-op'ing) [L-2]
- add closed-vocabulary rejection test: crafted --runtime values
  (../../etc/passwd, __proto__, --config-dir, garbage) are rejected, warn,
  and fall through to the persisted runtime — pins the security-load-bearing
  contract [NIT-01]
- add boundary tests: --config-dir wins over --runtime (precedence); missing
  --runtime value errors with USAGE [N-1]

Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed --runtime cannot coerce getGlobalConfigDir into an
arbitrary path (closed-vocabulary Map lookup + registry hash-key gate) and
does not expand the trust surface beyond the existing operator-controlled
--config-dir flag.
2026-07-06 23:32:06 -04:00
Tom Boucher
def745fa6b test(#2003): regenerate golden-install-parity fixtures for gsd-tools.cjs change
The --runtime parsing + help-text edit to gsd-core/bin/gsd-tools.cjs changes
the installed file's content (gsd-tools.cjs is installed and compared by the
golden snapshot, unlike gsd-core/bin/lib/ which is excluded). Regenerated via
UPDATE_GOLDEN=1; every runtime's manifest updates exactly one line (the
gsd-tools.cjs hash).
2026-07-06 23:09:54 -04:00
Tom Boucher
49552b3485 docs(#2003): add changeset fragment for --runtime override 2026-07-06 22:57:39 -04:00
Tom Boucher
ab82e73af3 fix(#2003): add --runtime override to capability state + loop render-hooks
resolveCapabilityRuntimeState derived the config dir from resolveRuntime(cwd)
(GSD_RUNTIME -> config.runtime -> 'claude') when no --config-dir was passed,
so a repo with persisted runtime:'codex' resolved the config dir to ~/.codex
where the Claude skill isn't installed -> surfaced:false / hooks silently
no-op when the operator drove from Claude Code. capability state and loop
render-hooks parsed only --config-dir, never --runtime, so there was no way
to assert the actually-active runtime.

Add a runtimeOverride param to resolveCapabilityRuntimeState (canonicalized
via runtime-name-policy so aliases like codex-app work); when present it
short-circuits the persisted-runtime fallback and resolves getGlobalConfigDir
for the explicit runtime. Thread --runtime through cmdCapabilityState and
cmdLoopRenderHooks, and parse it in gsd-tools.cjs for both commands (dual
--runtime X / --runtime=X form, mirroring --config-dir and the existing
capability-set --runtime precedent). Help text updated.

Without the override, behavior is byte-identical to today (regression-guarded).
2026-07-06 22:57:05 -04:00
Tom Boucher
6a15ab9345 test(#2003): add regression tests for --runtime override on capability state/loop render-hooks
Mirrors the #1160 installed-layout block for the runtime auto-detection gap.
Covers: runtimeOverride='claude' bypasses persisted config.runtime:'codex';
no override still honours persisted runtime (regression guard); alias
canonicalization (codex-app -> codex); and an end-to-end CLI test proving
'capability state --runtime claude' resolves the Claude config dir despite a
persisted runtime:'codex'.

Expected RED against unfixed resolveCapabilityRuntimeState (no runtimeOverride
param) and unfixed gsd-tools.cjs (no --runtime parsing for capability state /
loop render-hooks).
2026-07-06 22:57:05 -04:00
Codesmith
192764f0c3 chore(#1990): recapture zcode golden install fixture with onboard artifacts
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-07 02:55:18 +00:00
Tom Boucher
080bacdb4b Merge branch 'next' into codex/gsd-onboard 2026-07-06 22:43:32 -04:00
Tom Boucher
46d18c0800 Merge pull request #2049 from open-gsd/fix/1858-flat-layout-skill-manifest
fix(#1858): detect flat commands/gsd-*.md layout in _resolveManifest
2026-07-06 22:42:29 -04:00
Tom Boucher
12b4c625ad Merge branch 'next' into fix/1858-flat-layout-skill-manifest 2026-07-06 22:33:07 -04:00
Tom Boucher
4424a366d7 docs(#1858): backfill changeset pr number to 2049 2026-07-06 22:18:40 -04:00
Tom Boucher
2c853822a1 fix(#1858): address code+security review findings
- restructure _loadFlatCommandsGsdManifest try/catch to wrap read+parse+set
  together (mirrors loadSkillsManifest exactly), so a thrown parser degrades
  both keys to [] — closes the latent catch-scope parity drift [Nit-1]
- add boundary test: gsd-.md (empty stem) skipped, gsd-x.md single-char stem
  kept (slice(4,-3) boundary) [Low-1]
- add unreadable-file test (POSIX-gated): both keys degrade to [] [Low-2]
- strengthen parity test to also compare requires + _calls_agents_ VALUES,
  not just the stem set [Nit-2]

Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed no new trust-boundary crossing, prototype-pollution
immune (Map/Set throughout), and symlink/path-traversal surface identical to
the pre-existing nested loader (not a regression).
2026-07-06 22:06:30 -04:00
Tom Boucher
d99e1f9c82 Merge pull request #2048 from open-gsd/fix/2041-model-overrides-claude-alias
fix(#2041): map model_overrides Claude IDs to Agent-tool aliases
2026-07-06 21:36:12 -04:00
Tom Boucher
e3d053619c docs(#1858): add changeset fragment for flat-layout manifest fix 2026-07-06 21:14:44 -04:00
Tom Boucher
9dcd370332 fix(#1858): detect flat commands/gsd-<stem>.md layout in _resolveManifest
_resolveManifest only recognized the nested source layout (commands/gsd/*.md)
and the installed-runtime skills layout (skills/gsd-<stem>/SKILL.md). A flat
source install (Claude local project shape: commands/gsd-<stem>.md, no
commands/gsd/ subdir) matched neither branch, so the manifest came back empty
and resolveSurface materialized the full profile to an empty Set — silently
reporting every skill-bearing capability as surfaced:false / enabled:false /
active:false. The nyquist/code-review/security/ui verify:post and execute:post
hooks never fired even with their workflow.* toggles on.

Add a third branch: when commandsGsdDir is absent, scan dirname(commandsGsdDir)
for gsd-<stem>.md files, strip the gsd- prefix, and build the same Map shape
the nested loader produces (requires via shared parseRequires, companion
_calls_agents_<stem> via shared parseCallsAgents). Falls through to the
installed-skills branch when the flat dir has no gsd-*.md files (precedence:
nested > flat-source > installed).

Also export parseCallsAgents from install-profiles so capability-state reuses
the SAME parser the nested loader uses (no drift; mirrors the existing
parseRequires export+reuse pattern).
2026-07-06 21:13:42 -04:00
Tom Boucher
4b1825e5f8 test(#1858): add regression tests for flat commands/gsd-*.md layout
Mirrors the #1160 installed-layout tests for the flat source layout
(<repo>/commands/gsd-<stem>.md, no commands/gsd/ subdir). Covers stem
extraction (strip gsd- prefix), requires parsing via shared parseRequires,
companion _calls_agents_ key parity, _resolveManifest flat-branch detection,
precedence (flat-empty falls through to installed), and a generative-parity
assertion that the flat loader and nested loader produce identical stem sets
for the real command tree.

Expected RED against unfixed capability-state.cts (_resolveManifest has no
flat branch; _loadFlatCommandsGsdManifest not exported).
2026-07-06 21:10:27 -04:00
Tom Boucher
6136aa20de docs(#2041): backfill changeset pr number to 2048 2026-07-06 20:59:38 -04:00
Tom Boucher
e95af39a8c fix(#2041): address code+security review findings
- add typeof guard so a non-string override passes through verbatim instead of
  crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1]
- use Object.hasOwn() for the alias lookup so __proto__/constructor cannot
  return a truthy non-string from the plain object literal [LOW-D3]
- cap the unmappable-override stderr warning at 64 chars so an oversized or
  secret-shaped value cannot leak in full to stderr/logs [LOW-D4]
- remove the unused mapClaudeOverrideForRuntime export (helpers are covered
  behaviourally via resolveModelInternal/resolveModelForTier) [NIT]
- add resolveModelForTier unmappable-override fall-through test (closes the
  mutation-score gap) [MEDIUM-1]
- add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2]

Both orthogonal reviews returned APPROVE with no Critical/High findings.
2026-07-06 20:45:40 -04:00
Tom Boucher
13e40fcbf7 docs(#2041): add changeset fragment for model_overrides alias fix 2026-07-06 20:45:40 -04:00
Tom Boucher
f214f1320d fix(#2041): map model_overrides full claude IDs to agent-tool aliases
model_overrides values that are full Claude model IDs (claude-sonnet-5,
claude-opus-4-8, claude-haiku-4-5, claude-fable-5) were returned verbatim on
the claude runtime and handed to the Claude Agent tool, whose typed model
parameter documents only tier aliases (opus/sonnet/haiku/fable). The
model_policy path already mapped full IDs -> aliases via
CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides skipped that mapping, so
the two resolver paths produced different shapes for the same underlying
Claude model. The fix mirrors #1144 on the override path via a shared
mapClaudeOverrideForRuntime helper used by both resolveModelInternal and
resolveModelForTier. Bare aliases pass through verbatim; non-Claude runtimes
and non-Claude custom/vendor values keep full IDs verbatim (parity). An
unmappable Claude ID (e.g. claude-opus-4-5) warns once to stderr and falls
through to tier resolution, exactly as the model_policy path already does.
Alias mapping is also the documented best practice (prevents staleness when
new model versions ship).
2026-07-06 20:06:30 -04:00