Commit Graph

20 Commits

Author SHA1 Message Date
Tom Boucher
9a41a95212 fix(#4717): consult the per-install runtime marker at both identity seams (#4861)
* test(#4717): add failing-first coverage for the two runtime-identity marker seams

* fix(#4717): consult the per-install runtime marker at both identity seams

resolveReportedRuntime (agent_runtime) and loadConfigResolved
(config.runtime) both ignored the per-install .gsd-runtime marker that
resolveRuntime and the model-resolver gate already read. On a
multi-runtime machine (e.g. a globally exported CODEX_HOME), host sniffing
misreported every Claude Code session as codex, and a shared
defaults.json stamped by the first non-Claude install leaked its runtime
to every other one.

Seam 1: the reported-runtime ladder becomes explicit > install marker >
host detection > claude. Seam 2: loadConfigResolved fills an empty
config.runtime from GSD_RUNTIME then the marker, copy-on-write (the
builtin-defaults branch returns a shared object). Explicit runtimes and
marker-less trees are unchanged.

* fix(#4717): a marker-detected runtime opts into its tier map (decision a)

* fix(#4717): stamped-defaults leg, marker fail-safe, docs, review fold-ins

* chore(#4717): backfill changeset PR number (4861)

---------

Co-authored-by: sim <sim@local>
2026-09-18 13:50:01 -04:00
Tom Boucher
9b750dc00a fix(#4505): resolve models through the active runtime and the tier table (#4726)
* test(#4505): cover runtime-aware overrides and routing precedence

Failing-first for both halves of the issue, plus the precedence layers a naive
fix silently defeats.

Every row drives the REAL CLI in a subprocess. That is load-bearing: the defect
is WHICH function the shipped call sites reach, so a row calling the resolver
in-process would pass while every real spawn stayed broken. It also makes
GSD_RUNTIME hermetic -- it is ambient, and an in-process row would leak it into
its neighbours.

Two fixture mechanics are documented in the helper because each silently
invalidates a row when got wrong, and both were found by measuring rather than
by reading the loader:

  - the loader reads `process.env.GSD_HOME || os.homedir()`, so redirecting only
    HOME leaves a developer's real ~/.gsd/defaults.json in play;
  - the mere EXISTENCE of a .planning/ directory disables the shared-defaults
    layer, so a fixture that creates one stops exercising the "poisoned global"
    path #2297 acceptance #4 is about. Measured: .planning/ with config ->
    gpt-5.6-terra; .planning/ present but empty -> gpt-5.6-terra; no .planning/
    at all -> "".

Rows cover: both reported repros; the init payload a real spawn reads; the
omit gate, the runtime tier map and model_overrides each outranking the tier
table; model/tier coherence under dynamic routing; resolve-execution with the
attempt absent; max_escalations at limit-1/limit/limit+1 plus a cap of 0; and
four fail-safe rows pinning that only a value canonicalizing to a recognised
non-Claude runtime may outrank an omit.

Registers the docs-guard exemption path: the rows quote the documented
first-spawn contract in comments. The file still never READS a docs/ path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4505): resolve models through the active runtime and the tier table

Consolidates #4495 and #4493. One gap: the function every agent spawn goes
through consulted neither mechanism that was supposed to make resolution
runtime- and tier-aware.

Half A -- the runtime was READ from config, not resolved. resolveActiveRuntime
(GSD_RUNTIME -> config.runtime -> per-install marker -> claude) existed and
worked, but was called exactly once in the file. Every other site read
config['runtime'] raw, and that key is normally absent, so
model_profile_overrides.<runtime>.<tier> was inert for any install that
identifies its runtime through the environment or the marker. Same override
both times, differing only in WHERE the runtime is declared:

  GSD_RUNTIME=opencode, no runtime key   ->  sonnet          (ignored)
  runtime:"opencode" in the config       ->  TEST-OPENCODE   (works)

Half B -- nothing consulted dynamic_routing on the first spawn, though
docs/features/dynamic-routing-with-failure-tier-escalation.md documents
"the resolver picks tier_models[default_tier] for the FIRST spawn".

The tier-table lookup is extracted into ONE helper both entry points call, so
the first-spawn value and the escalated value cannot drift; resolveModelInternal
calls it at attempt 0 and resolveModelForTier at the real attempt.

Placement is the documented composition, not a convenience. The same doc says
"model_overrides always wins; dynamic_routing.tier_models[<tier>] resolves above
models.<phase_type> and model_profile" -- so the step sits BELOW model_overrides,
the model_policy preset, the runtime tier map, the resolve_model_ids:"omit" gate
and the claude tier override, and ABOVE the profile lookup. An earlier cut routed
every call site through resolveModelForTier instead, which returns the tier model
directly and therefore skipped three of those layers: with an omit and a
non-Claude runtime it handed out a model id where the gate had returned "".

Criterion 1 is applied in full, including the two value-policy reads #4192 had
recorded as "NOT via resolveActiveRuntime". The tests decided it: switching them
breaks nothing, so that reading was never enforced -- and the old behaviour
defeated #4192's own principle that an explicit pin must not be silently
unpinned (claude-opus-4-8 under GSD_RUNTIME=opencode collapsed to the Claude-only
alias opus). #4192's comment is updated in place rather than left stale.

The step-3 opt-in signal is CANONICALIZED. Comparing the raw config field against
the literal 'claude' made runtime:"Claude", "claude-code" and even 5 count as
non-Claude opt-ins and outrank an explicit omit -- failing OPEN in exactly the
#2297 case the guard exists to protect. null now covers both "not a string" and
"not a runtime we recognise", and both read as NOT an opt-in.

Verified cell by cell against a pristine origin/next worktree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4505): backfill changeset PR number (#4726)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:40:17 -04:00
Tom Boucher
fd4aac5670 fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime

Two documented model-configuration contracts did not hold on the claude
runtime (confirmed-bug scope from the issue triage):

Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of
resolveModelInternal gated runtime-aware tier resolution on
configRuntime !== 'claude', so the key's only reader was never consulted,
while workflows/settings-advanced.md writes it for claude-runtime users.
A new step 4.5 resolves ONLY the user's override entry (never the builtin
claude tier map, so unpinned installs keep resolving aliases). An
override value that maps to a current tier alias collapses to that alias
(byte-equivalent, the #2041 protection); anything else — a pinned older
generation, a bare alias repoint, a non-Anthropic id — resolves verbatim.
It sits after the resolve_model_ids:'omit' gate so an explicit project
omit still wins (#2297) and before the alias return so
resolve_model_ids:true cannot re-materialize the pin to the latest id.

Finding 2 — fully-qualified claude-* ids in model_overrides were
warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable
branch, #2041), while the docs promise any fully-qualified model id is
valid. The unmappable branch now passes the pin through verbatim with a
warn-once breadcrumb (text describes the pass-through). Dropping it
silently unpinned the operator's explicit choice — the exact 'profile
can misrepresent what actually runs' defect of #4192. Mappable ids and
non-claude values behave exactly as before; resolveModelForTier shares
the mapping; the tier honesty signal is unchanged (raw ids still report
'unknown'); the model_policy path is untouched.

Docs updated to the agreed contract (CONFIGURATION.md false 'Claude
example' corrected; how-to + shipped reference document the pin
semantics, the fable alias, and the tier-override composition).

* test(#4192): pin explicit model pin resolution on the claude runtime

28 failing-first rows across the resolver seam and the resolve-model CLI:
pinned-generation fidelity (tier override + per-agent verbatim pins,
object form, explicit runtime), unpinned controls byte-stable (no
override, other runtime/tier, inherit, project omit, precedence),
adversarial rows (prototype-chain keys, malformed values, warn-once
dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through
runGsdTools. The stale #2041 fall-through assertions now pin the
pass-through contract; mappable-id collapse assertions unchanged.

* chore(#4192): add changeset fragment

* chore(#4192): backfill PR number in changeset fragment

---------

Co-authored-by: ZCode <zcode@localhost>
2026-09-06 10:17:50 -04:00
Tom Boucher
107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00
Tom Boucher
004e9dd741 fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability

RED by construction. Binds to behavior renderEffortForRuntime does not yet
have: an optional third `model` argument, a per-model advertised-level table,
`max` passing through instead of clamping to `xhigh`, `minimal` clamping to
`low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/
`reason`) so a downgrade is legible from resolver output rather than silent.

Two of these pin defects that exist on next today:

- `max` is discarded. Both Codex models whose catalog entries are retrievable
  (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing.
- `minimal` is emitted to a model that refuses it. providerPresets.openai.
  haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's
  advertised floor is `low`. GSD is sending a value into a document Codex
  itself validates. The parity test is what pins that fixed, and it names the
  offending path/model/effort when it trips.

Also corrects tests/model-resolver.test.cjs:351, which asserted
renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as
though it were a contract. ADR-443 recorded "Codex has no max" as fact and it
was true when written; Codex has since added both `max` and `ultra`. That is a
stale premise, so the assertion is corrected here rather than worked around.

The property test asserts the invariant the whole change exists for: a rendered
effort is always a level the target model actually advertises, or an explicit
rejection. There is no third outcome.

* fix(#3007): resolve Codex effort per model, and make every clamp visible

Codex declares supported_reasoning_levels per MODEL and validates against it,
so a single per-runtime capability set cannot be right for all of them. GSD's
was wrong in both directions at once.

`max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped
max -> xhigh on that basis; it was accurate when written, and Codex has since
added both `max` and `ultra`. Every Codex model whose catalog entry is
retrievable advertises `max`, so the clamp was discarding a level the provider
supports, silently, on the most-used path.

`minimal` stops reaching Codex. No Codex model advertises it -- both retrievable
entries floor at `low` -- yet providerPresets.openai.haiku.low paired
gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the
receiver validates and refuses into a file the receiver reads. Being
unconservative in what you send is the half of Postel's rule with no defensible
reading, so that preset is corrected and a parity test pins it.

`ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum
reasoning with automatic task delegation": at ultra, effective_multi_agent_mode
returns Proactive and Codex spawns sub-agents on its own initiative, underneath
GSD's orchestration rather than inside it (#2167). It is a mode switch, not a
reasoning depth, so it is not added to the universal ladder -- which stays
provider-agnostic by ADR-443's design -- and it is rejected even for
gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered
and rejected: that silently discards what the user actually asked for.

Clamping is now visible. RenderedEffort carries requested/clamped/reason and
resolve-execution surfaces them. The previous table clamped correctly but
invisibly, so a user asking for `max` on Codex had no way to find out they were
getting `xhigh` -- exactly the failure mode the robustness principle's modern
critique warns about, and why "be liberal" has to mean "liberal and loud".

Also closes a latent trap found while reviewing the implementation: the clamp-up
loop walks the ladder upward, and for a future model advertising `ultra` but not
`max` it would have selected `ultra` as the clamp target -- re-entering by the
back door the mode the rejection above exists to keep out. A clamp may never
produce a value that a direct request for that value would refuse. Unreachable
with today's catalog, which is why no test caught it; a test now asserts the
invariant directly.

Signature stability is preserved: the third `model` argument is optional and the
two-argument form still resolves, against the family baseline. That form's
BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old
answer would have fixed the defect only where a model happened to be threaded
through and left it live everywhere else.

tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract
and is corrected here rather than worked around.

* fix(#3007): close every review finding on the Codex effort alignment

Two isolated reviewers, correctness and security. Both found the same two
blockers, and the per-model work was inert on every surface that matters until
this commit.

BLOCKER — resolve-execution never passed the model and discarded the clamp.
cmdResolveExecution called the two-argument form and emitted only
effort_rendered/effort_param/effort_propagation, so the per-model table was
unreachable from production code (tests were its only caller) and requested/
clamped/reason were computed and thrown away. Requested outcome 3 names "the
effective rendered effort in resolver output" specifically, so the feature was
unmet on the exact surface the issue asks for. Now passes the resolved model and
emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the
existing key convention rather than introducing a nested object.

BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed
a nested {"effort": ...} sample; the real result is flat and those keys were
absent entirely. A reference doc asserting a JSON path a reader can copy is worse
than no doc. Corrected against the actual emitted key set.

MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex
kept minimal in its supported set and still clamped max down to xhigh, so the
invocation-time and install-time channels disagreed about the same runtime's
capability: --host codex with max emitted xhigh while the generated TOML said
max. This is the repo's documented generative-fix-divergence class, so both
tables now cross-reference each other and a parity test fails if they ever
diverge again.

MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null
_baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback
never fired and every effort rendered as null. And a non-array value made the Set
constructor throw at module load — model-catalog.cjs is required across the whole
CLI, so one bad JSON value killed every command, not just codex effort. Guarded
on size and filtered to array values; both degrade to the hardcoded baseline.

MAJOR — value widened to a nullable string with two consumers left behind.
runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a
null effort key in generated frontmatter); install-effort-resolver still declared
a non-nullable return, a structural lie that silently defeated null checking.
Both corrected, both omitting the key on null — the same posture as 'inherit',
where omission means "follow the host default".

MAJOR — the per-model table is inert today, and the docs now say so. All three
shipped models advertise the same usable range and ultra (sol's only
differentiator) is rejected for every model, so no observable output differs by
model. The table stays because Codex declares capability per model and the sets
are free to diverge — a single per-runtime assumption is precisely what went
stale and produced this issue — but overselling it as a visible per-model feature
would have been the same class of error as the doc blocker above.

Tests: three passed under a full revert and are strengthened rather than deleted,
since each guards a real contract (#3533's inherit rule, the undeclared-host
rule, off-ladder handling) — they now also assert the clamp-visibility fields,
which only exist after this change. The fast-check property is kept for its
shrinking, and a deterministic nested loop over the full cross-product now sits
beside it so coverage is exhaustive rather than sampled.

Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg
form and would have written a literal null reasoning effort on the ultra path;
CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT.
The installer defect was found by the co-change gate, not by a reviewer —
install.js is a historical co-change partner of model-catalog.cts that this diff
had not touched.

* test(#3007): correct assertions that pinned Codex's stale effort premise

Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the
shipped commit. Every one is a stale pin, not a defect: each was probed against
the built module before its expectation was changed, and none failed for a
reason other than this premise correction.

Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale
by a production change must not ride inside another commit, because the
release-sdk hotfix cherry-pick filter routes by subject prefix and a correction
buried under the wrong prefix ships a half-state (v1.42.3, #3621).

The most valuable one was tests/model-resolver.test.cjs's cross-provider
validity invariant, which hardcoded the Codex enum as
`minimal|low|medium|high|xhigh` and failed with "real API would 400". That
message is now false in both directions: Codex accepts `max`, and rejects
`minimal`, which no model advertises. The enum is corrected to
`low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the
"would the real API refuse this" check worth having, and it was right to fail
here. It simply carried the stale fact in its own fixture.

Test NAMES were corrected alongside their assertions wherever the name asserted
the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal
passthrough". A renamed test that still claims the old thing is worse than a
failing one, and a green test whose name states a falsehood is how the next
reader inherits the wrong premise.

Both channels are covered: install-time (renderEffortForRuntime, and the
generated .toml in install-runtime-artifacts) and invocation-time argv
(effort-surface-axis). They were deliberately brought into agreement in this
change, so their assertions had to move together.

Each site carries a #3007 comment recording that Codex gained max/ultra and that
capability is declared per model, so a future reader can tell this was a
deliberate premise correction rather than a test bent to fit an implementation.

* test(#3007): separate the effort-precedence case from the clamp case

The previous stale-assertion pass over-corrected one test. It saw
`effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`,
assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to
"max passes through" and changed the expectation to `max`. The remote runner
disagreed.

Reproduced against the real CLI: with that config and `gsd-planner`, the
resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`,
`effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE —
gsd-planner is heavy/opus tier and its routing-tier default outranks
`effort.default` — so `max` never reaches the renderer at all. The test says
nothing about clamping and never did; it only looked like a clamp pin because
both mechanisms happened to yield the same string.

Restored to `xhigh` and renamed to say what it actually tests. It now also
asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is
what makes it impossible to mistake for a clamp pin again: those two fields prove
the value is what the resolver produced rather than something the renderer
downgraded. Before #3007 there was no way to tell the two apart from the output —
which is precisely why the previous pass could not tell them apart either.

Added the test that was actually missing: `effort.agent_overrides`, which
outranks the tier default, so the requested level genuinely reaches the renderer
and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified
by probe before asserting.

One test now pins the precedence rule and the other pins the #3007 behavior, and
neither can be read as the other. That the clamp-visibility fields are what
resolved this is a small argument for having added them.

* chore(#3007): backfill changeset pr number to 3765

* test(#3007): put model-catalog under the mutation gate

The Stryker shard showed as `skipping` on this PR despite the diff rewriting
model-catalog's effort logic. That was legitimate, not a detection bug:
`model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the
whole module — including everything #3007 touches — sat entirely outside
mutation scoring with has_work "false".

Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs
is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no
temp dirs. That shape is not stylistic — it is the #2790 precedent this file
already documents. Stryker's command runner treats a whole `node --test <file>`
invocation as ONE test costing whatever its slowest case costs, and re-runs it
per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses
runGsdTools throughout) would reproduce exactly the 15-minute shard-cap
cancellation #2790 hit. The integration file is unaffected and keeps running in
full in the normal test job.

Coverage spans the module rather than only the diff, because the score is
measured over the whole file: effort rendering across every model and ladder
level in both channels, the prototype-chain host guard, the exported enums and
maps, isAnthropicFlavoredModel's provider namespacings, the profile projections,
nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are
worth naming — every uncovered exported function is score given away, and
mergeEffortTierDefaults turned out to have a genuinely interesting contract
(#3531: a partial override merges over the built-ins rather than replacing them,
and isValid gates the VALUE, not the tier name, so an unknown tier key is still
merged in). Every expectation was probed against the built module before being
asserted.

minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment.
Floors in this repo are measured, not chosen — the existing entries sit at 94, 75
and 56 — and they can only be measured in CI, because mutation shards run
`node --test`, which is hard-blocked locally. The first CI run on this branch
reports the real number and the floor gets ratcheted to it before merge. A
placeholder of 1 reaching `next` would make the gate decorative: it would pass
whether or not a single mutant is ever killed.

Note the target is "never regress from measured", not a fixed 80 — planning-inspect
sits at 56 and is documented as an accepted ratchet candidate.

* test(#3007): bootstrap model-catalog's mutation floor legally

The placeholder floor was structurally illegal and the remote run said so.
tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module
must carry a matching RATCHET_BASELINE entry in the same diff, minScore must
EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three.
That is the ratchet working exactly as intended — a floor nobody can satisfy
accidentally is the point of it.

Bootstrapped at 50 in both places. Fifty is not a measured score and the comment
says so plainly: it is the minimum the guard permits, and it coincides with
Stryker's own configured `break` threshold, so it is the lowest legal starting
point for a module that has never been measured. It still must be ratcheted to
floor(measured) - 1 before this PR merges.

Also corrected a real defect in the file's own instructions. "HOW TO UPDATE"
step 1 read "Run the per-module Stryker shard locally" — which cannot be done
here, and which the same file contradicts eighty lines further down, where the
#2790 scores are recorded as "not a local run; mutation shards run `node --test`,
hard-blocked in this repo's local environment". stryker.config.mjs confirms the
command runner invokes `node --test` once per mutant, and
.claude/hooks/block-local-node-test.sh denies exactly that. So the documented
first step sends the next contributor at a wall. Rewritten to describe the path
that works — push, read the measured score off the CI shard, then set the floor
and its baseline together in one diff — and to say why local measurement is not
available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched.

The two-step is inherent to the environment rather than a shortcut: a floor
cannot be measured before the first CI run exists, and the guard rightly refuses
to accept an unmeasured one below its minimum.

* test(#3007): ratchet model-catalog's mutation floor to its measured score

The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no
timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this
file's own rule, minScore = floor(measured) - 1, which is the same arithmetic
every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94.

Both halves moved together, because the ratchet guard asserts minScore equals its
RATCHET_BASELINE entry and would reject them drifting apart.

The spawn-free unit surface is vindicated by the clock: 57 seconds, against a
15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the
whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing
the shard at tests/model-resolver.test.cjs — #2790 recorded shards being
CANCELLED at that cap when they targeted a runGsdTools-heavy integration file.

The registry comment is rewritten rather than deleted. It previously warned that
the floor was provisional and must not ship that way; leaving that text next to a
measured floor would make the file lie in the other direction. It now records the
measurement the way the sibling entries do, including that 59.62 sits below
TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 —
comfortably clear of its own floor with real room to grow. Raise it as the tests
improve; never lower it.

Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is
an honest floor for a module that had NO mutation coverage at all an hour ago,
and it is now pinned so it cannot silently regress.

---------

Co-authored-by: sim <sim@local>
2026-08-22 20:51:55 -04:00
Tom Boucher
682eaae3f0 enh(#2876): retire the dead and pass-through exports from bin/install.js (#3615)
* enh(#2876): retire the dead and pass-through exports from bin/install.js

The installer exported 197 names and had zero production consumers - every
non-test require of it repo-wide sits inside a comment. Its interface was
shaped by test access, not by callers.

Removes 9 dead exports and 61 pass-throughs, repointing their tests onto the
extracted modules' own interfaces. 197 down to 127.

Every count in the issue was wrong: 197 exports not 188, 9 dead not 12, 61
pass-throughs not 49, 44 test files not 42 - and the audit itself then missed
7 more consumer files. restoreUserArtifacts was on the dead list but ceased to
exist in phase 6, and two _GSD_EFFORT_MANIFEST_* names listed as dead are now
genuinely asserted, so acting on that list would have deleted live exports.

7 of the 9 dead names collide with an independent declaration that install.js
delegates TO. Each removal was justified by which declaration a reference
resolves to, never by whether the name appears somewhere.

Coverage parity was the gate rather than test greenness: per-file counts were
captured before any edit and diffed after. 44 of 45 files are byte-identical;
the single delta is one added assertion, not a loss.

The sweep for scattered require sites found two forms static grep misses -
require(VARIABLE) and multi-line require() - plus tests asserting that
install.js re-exports the SAME object, which now assert retirement instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2876): close review findings — restore the duplicate-body guard, sweep orphaned code

Both review engines found real defects in the first cut.

The DEFECT.GENERATIVE-FIX single-owner guard from #1511 had been repointed
from a reference-identity check to install.X === undefined. Those are not
equivalent: the guard exists to catch a duplicate function body reintroduced
into install.js, and the replacement passes cleanly if that duplicate is used
internally and never exported. It now walks bin/install.js's real top-level
bindings, so it catches a duplicate under either shape, exported or not -
strictly stronger than the check it replaced. Proved by injecting a duplicate
and watching it go red.

That weakening survived the coverage-parity gate because the assertion count
never moved. The gate compares counts, so an assertion that changes meaning
rather than number is invisible to it.

Removing the exports had orphaned their wrapper bodies: 14 dead wrappers, 9
consts and 9 destructure entries, several pre-existing and found by the same
sweep. Dead code left in the file this phase exists to shrink.

Three more comments claimed re-exports this phase removed, and tests were
reading Cursor and Windsurf hook constants from install.js's local copy while
calling functions from the hooks surface - equal today, with nothing holding
them equal. The local consts now reference the owning module.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2876): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 09:23:44 -04:00
Tom Boucher
3ab0007164 enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19)

preserveUserArtifacts held user files only in an in-memory Map across the
wipe, so any process death between preserve and restore lost them outright.

Seven call sites, not the four the issue records. Three of them never called
the helper at all - they open-coded the same read/wipe/write - so searching
for callers under-counted by construction; the extra sites were found by
sweeping for the pattern instead.

The worst is the mainline install path, where the crash window spans the
entire gsd-core tree copy rather than a single rmSync.

Adds src/user-artifact-staging.cts: durable on-disk staging with a record
written after the copies land as the commit point, plus recovery of orphaned
batches on the next run - without recovery the staged bytes survive but the
user's file is still gone, which would pass its own test while delivering
nothing.

Routes copyPreservingSymlink through installFs() so staging cannot bypass the
install fs seam, and reunites its symlink-safety docblock with the function it
documents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-3574 with four claims disproved by implementation

Implementing Phase 6 disproved four statements the ADR rests on. The central
decision - no single materializer - is unaffected and stands.

Corrected: decision 3 was already satisfied, so nothing was extracted; the
agents-bypass runtime set omitted claude, kilo and opencode, and closing it
needed three new pieces of descriptor contract rather than proceeding on its
own terms; three of the four blockers the layout comment names were already
stale; and F19 is seven call sites, not four.

Records the generalizable lesson: the defect is the pattern of holding user
data in memory across a wipe, not the helper, so searching for callers of the
helper under-counts by construction.

Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and
notes that copyPreservingSymlink needed routing through the install fs seam
before it could be reused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close dangling-symlink blind spot and harden staging recovery

An adversarial review found the F19 staging work shipped red and unsafe.

Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween
missed dangling symlinks in both its root check and its per-segment walk,
because it probed with existsSync, which is false for a link whose target does
not exist. Fixing only the new module would have reused a guard that was
itself blind. This guard protects the whole install tree.

Recovery no longer throws: it degrades per entry and per file, so one bad
batch cannot block the others. Previously an unrecoverable entry propagated
out of the first statement of install and uninstall, before the cleanup that
would have removed it - wedging the installer permanently.

Partial fs adapters now throw on any omitted method instead of silently
reaching the real filesystem, closing the trap that let a test poison list
pass while real IO happened.

Staged names must be flat, recovery refuses a dangling destination symlink,
and a batch whose recovery genuinely failed is no longer swept - it was
discarding the only durable copy of the file it had just failed to restore.

Replaces three tests that could not fail, including the one labelled negative
proof.

Known limitation, documented not closed: concurrent installs sharing a staging
key can still lose a batch. A real fix needs a cross-process lock.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enh(#2875): make the descriptor authoritative for the agents kind

Deletes the inline agent-staging loop in bin/install.js and the
_DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from
its capability descriptor instead of an inline hostBehaviors dispatch.

Closing it needed three pieces of contract the descriptor pipeline never had,
all reducible to one missing input - per-agent resolution context: a
frontmatter-extensions step for claude's effort and disallowedTools, per-agent
model-override resolution for kilo and opencode, and a named branding
converter for hermes, whose rewrite data was already declared.

Seven runtimes were on the loop, not the six the design recorded - kimi-code
was found by a golden fixture, not by analysis. claude-local and kimi-code
both silently lost their agents mid-change; the fixtures caught both and the
cause was fixed rather than the fixtures regenerated.

A parity harness gates the migration: both pipelines over identical inputs,
byte-identical output including filenames, per runtime. It is demonstrated
red before being trusted. Surface and install paths converge for all seven,
which also fixes surface previously writing no agents for these runtimes.

Codex's config.toml strip stays put - it mutates host config, which no
descriptor kind models.

Also routes install-model-override-resolver and install-effort-resolver
through the install fs seam. Both leaked real filesystem IO from the install
call tree; the stricter adapter is what exposed them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): record the agents-descriptor migration and correct the ADR count

The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host
integration guide told readers to join a set that is gone. Replaces that with
what is now true - declare an agents entry and it installs, on the surface
path as well as install - and points anyone needing a per-agent transform at
the three extension points rather than at a new inline branch.

Corrects the ADR amendment: seven runtimes were on the inline loop, not six.
kimi-code was found by a golden fixture going red, not by reading. That is the
third short count this phase, all from enumerating by symbol or set membership
when the thing that matters is a behavior.

Adds the Changed changeset for the surface-path convergence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-2866 - claude global always wrote agents on disk

The claude row's global=[skills] described what capability.json declared, not
what the installer wrote. bin/install.js's inline agent-staging loop was never
scope-gated and never consulted the descriptor, so a claude --global install
has always written agents/gsd-*.md.

Phase 6 closes the gap by deleting that loop and declaring agents on claude's
descriptor at global scope. On-disk bytes are unchanged - the golden fixtures
did not move, which is the evidence that the descriptor, not the installer,
was incomplete.

#2218 is unaffected: agents are not trigger-bearing, so the wider row does not
introduce a new shadowing case.

Records the warning that an incomplete descriptor is invisible while a second
code path silently does its work, and only surfaces when the two are forced
into agreement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close review findings across staging, agents and the parity harness

Two independent reviews of this branch found defects the local gates missed.

Security: a dangling symlink at a migration destination allowed writing
outside configDir - the same class this change claimed to close, missed at the
terminal write of the flow being added. The staging-root resolver threw as the
first statement of install and uninstall, so a hostile symlink bricked both,
and symlinked-configDir users lost uninstall as well as install; it now
degrades instead of aborting. Recovery gained a source-side symlink check and
now refuses a relative destDir, which resolved against cwd. Converter dispatch
gained a runtime allowlist - lint-time validation stopped mattering once this
branch promoted that dispatch from the surface path to real installs.

Correctness: claude --local --minimal exited 1 because the minimal profile
legitimately yields zero agents and the new path treated that as a failure.
cline --local silently lost its agents - its descriptor declared none while
the deleted loop wrote them unconditionally. The agents prune was widened to
any gsd-* entry and destroyed user files it never owned.

The parity harness, on which the migration's safety argument rested, drove a
synthetic registry and never byte-compared the shipped descriptors; two of its
trap rows could not fail. It now drives the real registry across 13
runtime-scope rows including kimi-code and cline-local, and its red-proof is
demonstrated by corrupting a live capability.json. Three goldens that had
encoded the cline regression as expected behavior were corrected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close findings from both mandated review engines

/security-review found the staging source-side walk honouring
GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write
destination. A symlinked files/ component dereferenced because
copyPreservingSymlink lstats the leaf only, so an intermediate link is
followed. The source walk no longer honours the opt-in; the destination check
still does.

/code-review spec axis found this branch had reintroduced its own bug:
migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after
the legacy dir was wiped and before the staged batch was restored, so a
planted symlink bricked uninstall permanently and orphaned the batch. Refusal
kept, abort removed.

kimi-code local silently lost its agents, the same class as the cline bug, and
the parity harness recorded that exclusion as intentional - the third test in
this branch to pin a regression as correct.

--minimal now creates an empty agents/ dir that never existed. Behaviour
restored rather than softening the changeset, so its byte-identical claim
stays true.

Standards axis: try/finally removed from twelve test bodies, fast-check
properties added for parseOwnerPid, boundary coverage at the grace window and
the ancestor-probe depth, a parity assertion for the staging-root helper
duplicated across two files, and the 8-deep config walk deduplicated.

Records 60-review.json with every finding and disposition from five passes,
including the smells left unfixed and why.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): prune stale agents unconditionally in minimal mode

The previous round stopped an empty agents/ directory being created when the
resolved profile yields no agents. That was implemented by skipping the agents
kind entirely, which also skipped its stale-agent prune - so a full to minimal
downgrade left stale gsd-* agents behind.

The deleted inline loop pruned unconditionally and only skipped writing. Those
are three separate conditions, not one: prune always, write only when there is
something to write, create the directory only when writing.

Both call sites now run _removeGsdEntries before the empty-staged early exit.
The symlink-escape guard moved with it, since the prune also touches dest.
Codex .toml agents and the config.toml stanzas are cleaned again, and
user-owned agents are still preserved.

The agents/ directory is left in place after a prune empties it, matching
every sibling kind - none of them remove the destination directory itself.

Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh
install, so fixture generation is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): document interrupted-install recovery for user-owned files

The durable-staging fix is invisible to the user it protects. Someone whose
install died mid-flight has no way to know USER-PROFILE.md was staged before
the delete, that the next run restores it, or that recovery happens at the
start of that run rather than in the background.

Written as the task the user has - finish the interrupted command - rather
than as a description of the mechanism, and states what it will not do:
overwrite a file already present, or touch staging belonging to another
install still running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2875): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2875): assert the J8 model override without building a regex

CodeQL flagged incomplete string escaping: the assertion interpolated the
override value into a RegExp while escaping only forward slashes, which is
meaningless in a constructor, leaving real metacharacters unescaped.

The failure direction was the dangerous one - a metacharacter would have made
the match more permissive, so the row would pass when it should fail. That
matters here because J8 exists precisely because an earlier revision was a
tautology; the rewrite reintroduced a different way for the same assertion to
stop discriminating.

Replaced with a line-wise exact match, so no regex is constructed at all.
Swept the other test files this branch adds; no sibling instances.

lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full
metachar-escape copy, so a single slash replace slipped under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 17:25:53 -04:00
Tom Boucher
b7cca0363f fix(#3531): merge routing_tier_defaults over manifest tier defaults (#3539)
* test(#3531): failing-first suite for routing_tier_defaults manifest merge

* fix(#3531): merge routing_tier_defaults over manifest tier defaults

* docs(#3531): document routing_tier_defaults merge-over-built-ins semantics

* fix(#3531): correct test helper scope, update folded #443 expectations, guard merge keys

* test(#3531): pin tiers in effort-sync and surface-axis fixtures post-merge

* chore(#3531): backfill changeset pr number

* fix(#3531): correct rebase resolution — keep both 3531 and 3533 test blocks intact

* test(#3531): pin inherit/effort fixtures to the layer that reaches tiered agents

---------

Co-authored-by: sim <sim@local>
2026-08-15 08:19:35 -04:00
Tom Boucher
50d5368add fix(#3533): effort inherit — expressible, omitted at writers, never re-added (#3541) 2026-08-15 07:00:44 -04:00
sim
c444051bef test(#3339): fix orthogonal-review findings — Wave 7 fold
Standards/Spec-axis review + Memtrace graph pass found real issues in
the just-folded suites, all fixed here:

- tests/phase.test.cjs: the issue explicitly asked to dedupe overlapping
  fixtures between issue-2945/issue-2949 — the fold preserved both test
  sets correctly (genuinely distinct code paths, 0 tests dropped, already
  verified correct) but left each fold block with its own near-identical
  copy of a runVerifiedPhaseComplete(args, tmpDir) helper, the fixture-
  level dedup the issue actually asked for. Consolidated into one shared
  definition without disturbing the file's other, unrelated same-named
  helper at module scope (would have collided if hoisted directly).
  Removed two now-unused local runGsdTools destructurings left behind by
  the consolidation (both blocks already close over the module-scope
  import at line 26).
- tests/review-lane-descriptor.test.cjs: two .find() results dereferenced
  without a presence guard (same defect class Wave 3/#3335 already found
  and fixed once in this epic) — added assert.ok() guards matching the
  repo's established style.
- tests/model-resolver.test.cjs: documented the 1-of-80 dropped duplicate
  test with an inline comment, matching this same wave's host-integration
  fold's convention of citing drops by exact reference instead of leaving
  a reviewer to reconstruct the justification via git archaeology.

No test() count changed in any file. No production code touched.
2026-08-12 08:25:14 -04:00
sim
1bc7f7e6b0 test(#3339): fold the state/phase/dispatch & model-profile issue-* cluster — Wave 7
Folds 9 legacy issue-*.test.cjs regression files (140 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
LAST of 4 issue-* waves — closes out the 74-file fix-*/issue-* backlog
(pending BUG_FILE_RE extension, held for a follow-up commit until Wave 6
is confirmed merged, per the epic's own zero-backlog precondition).

- issue-2828-flat-roadmap-total-phases.test.cjs (1) + issue-3204-state-
  writer-phase-count.test.cjs (21): both target state-document.cjs
  buildStateFrontmatter via different CLI entrypoints — merged jointly
  into state-document.test.cjs, 0 dropped.
- issue-2945-phase-complete-checkbox-rollback.test.cjs (4) + issue-2949-
  phase-complete-stage3-sentinel.test.cjs (4): both target phase.cts
  cmdPhaseComplete; issue explicitly warned of overlap — verified
  disjoint fixtures/assertions, 0 dropped, merged into phase.test.cjs.
- issue-2927-reviewer-lane-overlay-invocation.test.cjs (10) merged into
  review-lane-descriptor.test.cjs.
- issue-2939-dispatch-flatten-maxdepth.test.cjs (9) merged into
  host-integration.test.cjs, 2 dropped as verified exact duplicates.
- issue-2977-frontmatter-bom.test.cjs (5) merged into frontmatter.test.cjs.
- issue-2045-third-party-skills-surface.test.cjs (6) merged into
  capability-loader.test.cjs.
- issue-2517-runtime-aware-profiles.test.cjs (80, the largest single
  fold in the epic) merged into model-resolver.test.cjs, 1 dropped as a
  verified true duplicate (checked against src/model-resolver.cts logic,
  not just title similarity).

Fixed a genuine eslint irregular-whitespace finding: a literal BOM
character embedded in a doc comment (pre-existing content from the
original #2977 source, illustrating what a BOM looks like) — replaced
with a readable U+FEFF notation.

3 stale doc references found and fixed (docs/adr/2313, 3180, 443).

Zero net test-coverage loss. No production code changed.
2026-08-12 08:25:14 -04:00
Tom Boucher
a875372f18 test(#3335): fold the workflow-content & phase-lifecycle fix-* cluster — Wave 3 (#3373)
* test(#3335): fold the workflow-content & phase-lifecycle fix-* cluster — Wave 3

Folds 13 legacy fix-*.test.cjs regression files (131 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053):

- 6 files with no prior target coverage: renamed (git mv) into new suites
  (spike-manifest-scoping, ship-note, add-todo, workflow-jq-dependency,
  resolve-execution-dynamic-routing, clock)
- 7 files merged into 5 pre-existing suites (worktree-base-ref x2,
  model-resolver, phase-locator x2, frontmatter, verification-status),
  deduplicated against existing coverage

Zero net test-coverage loss: every source assertion preserved or verified
as a genuine pre-existing duplicate. No production code changed.

Last fix-* wave (Wave 1 #3341, Wave 2 #3342 already merged); 4 issue-*
waves remain in #3315.

* test(#3335): fix orthogonal-review findings — Wave 3 fold

Standards-axis review + Memtrace graph pass found real defects in the
just-folded suites, all fixed here:

- phase-locator.test.cjs: pinned an unseeded fast-check property test
  (CONTRIBUTING.md determinism requirement), matching the sibling test's
  seed:7 convention.
- phase-locator.test.cjs: added assert.ok() presence guards after 9
  data.plans.find() calls that were dereferenced unguarded, inconsistent
  with 5 sibling tests in the same file that already guard correctly.
  Latent robustness gap — an omitted plan would throw an opaque TypeError
  instead of a clear assertion failure.
- Standardized the fold-wrapper convention (block-scoped __foldDescribe)
  across worktree-base-ref.test.cjs, verification-status.test.cjs, and
  phase-locator.test.cjs to match the pattern already established in
  frontmatter.test.cjs and model-resolver.test.cjs from earlier folds.
- worktree-base-ref.test.cjs: moved a mid-file require to the top-of-file
  require block.
- Eliminated duplicated env-isolation helpers: model-resolver.test.cjs and
  phase-locator.test.cjs each reimplemented GSD_WORKSTREAM/GSD_PROJECT
  save-restore independently; factored a shared isolateWorkstreamEnv()/
  restoreWorkstreamEnv() into tests/helpers.cjs and pointed both call
  sites at it.

No test() count changed in any file. No production code touched.

---------

Co-authored-by: sim <sim@local>
2026-08-11 22:23:55 -04:00
Rezolv
2dbee3ebdd enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229.

Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone.

Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined.

To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged.

Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change).

Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes.
2026-08-09 23:05:41 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
e95af39a8c fix(#2041): address code+security review findings
- add typeof guard so a non-string override passes through verbatim instead of
  crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1]
- use Object.hasOwn() for the alias lookup so __proto__/constructor cannot
  return a truthy non-string from the plain object literal [LOW-D3]
- cap the unmappable-override stderr warning at 64 chars so an oversized or
  secret-shaped value cannot leak in full to stderr/logs [LOW-D4]
- remove the unused mapClaudeOverrideForRuntime export (helpers are covered
  behaviourally via resolveModelInternal/resolveModelForTier) [NIT]
- add resolveModelForTier unmappable-override fall-through test (closes the
  mutation-score gap) [MEDIUM-1]
- add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2]

Both orthogonal reviews returned APPROVE with no Critical/High findings.
2026-07-06 20:45:40 -04:00
Tom Boucher
9ea5519bc0 test(#2041): add regression test for model_overrides claude alias mapping
Mirrors the #1133 model_policy alias-mapping tests for the model_overrides
path. Covers AC1-AC6: mappable Claude full IDs (claude-sonnet-5/opus-4-8/
haiku-4-5/fable-5) resolve to aliases on runtime:claude; bare aliases pass
through; non-claude runtimes keep full IDs verbatim; unmappable Claude IDs
warn-once + fall through; resolveModelForTier escalation path also maps;
non-Claude custom/vendor values pass through verbatim (regression guards).

Expected RED against unfixed model-resolver.cts (override short-circuit at
lines 162-167 / 288-290 returns override verbatim with no alias mapping).
2026-07-06 19:33:43 -04:00
Tom Boucher
697cbb1f05 test(#1977): consolidate 22 misc + repo-invariant regression tests
Final epic-#1969 batch. Fold 22 issue-named files: the 4 genuine repo-wide invariant
scans (551-eslint-bin-lib-coverage, bug-3054 stale /gsd-next, bug-3810 no-gsd-sdk-runtime-refs,
feat-3593 cli-negative-universal) into a NEW shared repo-invariants.test.cjs; the other 18 as
singletons into their nearest module suite (model-resolver, codex-config, runtime-converters,
security, state-transition, worktree-safety, roadmap-parser, etc.). Verbatim block-scoped
describe wrappers; 334 subtests conserved 1:1.

Host-env pre-check (B2+B6): the 6 CLI folds into GSD_TEST_MODE-setting hosts (model-resolver/
codex-config/runtime-converters) are benign — each origin independently sets GSD_TEST_MODE=1
itself (idempotent), unlike the B6 real-install case.

Regenerates regression-name allowlist (222->213), ratchets file-count allowlist (state 17->16),
makes 7 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids).
Repoints 2 tests/ refs in docs/TESTING-SUITES.md. lint:ci green.

Part of epic #1969. Closes #1977.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:06:11 -04:00
Tom Boucher
85ed50cc4f test(#1972): consolidate 94 command/module regression tests into subject suites
Fold 94 issue-named command/module regression files into the canonical test file
that owns each subject-under-test, across 52 existing suites (state, config, frontmatter,
roadmap-parser, capability-registry, shell-command-projection-dispatch, plan-phase-drift-guard,
health-validation, runtime-converters, commands, etc.). Verbatim block-scoped describe
wrappers; 881 subtests conserved 1:1. No new test files.

Host-env pre-check (per B2): the only GSD_WORKSTREAM/GSD_PROJECT-touching destinations
(intel, planning-workspace) clear those vars hermetically, so folded CLI tests are safe.

Regenerates regression-name allowlist (222->162), ratchets file-count allowlist across
8 buckets (validate entry removed after dropping <=2), makes 34 relocated allow-test-rule
exemptions issue-ref-compliant (ADR-456; prunes 34 stale ids). Repoints CONTEXT.md +
ADR-0002/443/1235/3524 test-file references. lint:ci green.

Part of epic #1969. Closes #1972.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 08:59:23 -04:00
Tom Boucher
c76827afbc refactor(#1291): T6 — migrate test files off the core spine ahead of deletion (#1293)
The convergence lint only scanned src/ + gsd-core/bin, so ~35 test files
still imported core.cjs. Repoint all 33 behaviour importers to the leaf
modules directly (same symbol->leaf map as the src migration; leaves are the
objects core re-exported by reference), delete the now-meaningless
shim-identity describe blocks in the 8 leaf tests, and delete tests/core.test.cjs
(forwarded-behaviour coverage now lives at the leaves; resolveWorktreeRoot
test relocated to worktree-safety in T0) and tests/lint-core-spine-imports.test.cjs
(the lint is removed in T-final). Dropped the stale core.test.cjs entries from
the allow-test-rule-refs allowlist; eslint-rules RuleTester fixture path
pointed at io.cjs.

After T6: ZERO test imports core.cjs. core.cts still builds (now fully unused);
T-final deletes it. No behaviour change.

Closes #1291

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:11:20 -04:00
Tom Boucher
185935379a refactor(#888): extract model+effort resolution into model-resolver.cts (#890)
ADR-857 rollout phase 2f — the FINAL core.cts decomposition. Move the model and
effort resolution cluster (resolveModelInternal, resolveModelPolicy,
resolveTierEntry, _resolveRuntimeTier, resolveModelForTier,
resolveGranularityInternal, assertValidGranularityOverride, resolveEffortInternal,
resolveFastModeInternal, resolveEffortForTier, nextEffort + VALID_GRANULARITIES/
VALID_EFFORTS/EFFORT_SET + interfaces) out of core.cts into a new leaf module
src/model-resolver.cts. core.cts re-exports the 13 public symbols (callers in
init/docs/commands unchanged; export= set byte-identical).

Cycle-free: model-resolver imports only leaves (config-loader for loadConfig,
configuration for defaults, model-profiles + model-catalog for the static
tables). Removed 6 now-unused imports from core (verified zero remaining
references, none re-exported).

This completes the god-module decomposition: core.cts 2271 -> 389 lines (~83%),
now a thin re-export spine over seven clean leaves (io, phase-id, roadmap-parser,
core-utils, phase-locator, config-loader, model-resolver).

New-CLI-module checklist done (.gitignore, eslint, INVENTORY 96->97 + row,
manifest, ARCHITECTURE, CONTEXT.md "Model Resolver Module"). Adds
tests/model-resolver.test.cjs (81 tests: behavioral + shim-identity + adversarial).

Gates: lint, code-review (export set byte-identical; import-removal verified),
security-review, codex adversarial-review (all 0 findings; verbatim move). Mac
4303 pass; clean-build docker 13117 pass, 0 fail.

Closes #888

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 17:26:40 -04:00