Files
msd-core/scripts/mutation-matrix.cjs
Tom Boucher 004e9dd741 fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability

RED by construction. Binds to behavior renderEffortForRuntime does not yet
have: an optional third `model` argument, a per-model advertised-level table,
`max` passing through instead of clamping to `xhigh`, `minimal` clamping to
`low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/
`reason`) so a downgrade is legible from resolver output rather than silent.

Two of these pin defects that exist on next today:

- `max` is discarded. Both Codex models whose catalog entries are retrievable
  (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing.
- `minimal` is emitted to a model that refuses it. providerPresets.openai.
  haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's
  advertised floor is `low`. GSD is sending a value into a document Codex
  itself validates. The parity test is what pins that fixed, and it names the
  offending path/model/effort when it trips.

Also corrects tests/model-resolver.test.cjs:351, which asserted
renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as
though it were a contract. ADR-443 recorded "Codex has no max" as fact and it
was true when written; Codex has since added both `max` and `ultra`. That is a
stale premise, so the assertion is corrected here rather than worked around.

The property test asserts the invariant the whole change exists for: a rendered
effort is always a level the target model actually advertises, or an explicit
rejection. There is no third outcome.

* fix(#3007): resolve Codex effort per model, and make every clamp visible

Codex declares supported_reasoning_levels per MODEL and validates against it,
so a single per-runtime capability set cannot be right for all of them. GSD's
was wrong in both directions at once.

`max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped
max -> xhigh on that basis; it was accurate when written, and Codex has since
added both `max` and `ultra`. Every Codex model whose catalog entry is
retrievable advertises `max`, so the clamp was discarding a level the provider
supports, silently, on the most-used path.

`minimal` stops reaching Codex. No Codex model advertises it -- both retrievable
entries floor at `low` -- yet providerPresets.openai.haiku.low paired
gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the
receiver validates and refuses into a file the receiver reads. Being
unconservative in what you send is the half of Postel's rule with no defensible
reading, so that preset is corrected and a parity test pins it.

`ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum
reasoning with automatic task delegation": at ultra, effective_multi_agent_mode
returns Proactive and Codex spawns sub-agents on its own initiative, underneath
GSD's orchestration rather than inside it (#2167). It is a mode switch, not a
reasoning depth, so it is not added to the universal ladder -- which stays
provider-agnostic by ADR-443's design -- and it is rejected even for
gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered
and rejected: that silently discards what the user actually asked for.

Clamping is now visible. RenderedEffort carries requested/clamped/reason and
resolve-execution surfaces them. The previous table clamped correctly but
invisibly, so a user asking for `max` on Codex had no way to find out they were
getting `xhigh` -- exactly the failure mode the robustness principle's modern
critique warns about, and why "be liberal" has to mean "liberal and loud".

Also closes a latent trap found while reviewing the implementation: the clamp-up
loop walks the ladder upward, and for a future model advertising `ultra` but not
`max` it would have selected `ultra` as the clamp target -- re-entering by the
back door the mode the rejection above exists to keep out. A clamp may never
produce a value that a direct request for that value would refuse. Unreachable
with today's catalog, which is why no test caught it; a test now asserts the
invariant directly.

Signature stability is preserved: the third `model` argument is optional and the
two-argument form still resolves, against the family baseline. That form's
BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old
answer would have fixed the defect only where a model happened to be threaded
through and left it live everywhere else.

tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract
and is corrected here rather than worked around.

* fix(#3007): close every review finding on the Codex effort alignment

Two isolated reviewers, correctness and security. Both found the same two
blockers, and the per-model work was inert on every surface that matters until
this commit.

BLOCKER — resolve-execution never passed the model and discarded the clamp.
cmdResolveExecution called the two-argument form and emitted only
effort_rendered/effort_param/effort_propagation, so the per-model table was
unreachable from production code (tests were its only caller) and requested/
clamped/reason were computed and thrown away. Requested outcome 3 names "the
effective rendered effort in resolver output" specifically, so the feature was
unmet on the exact surface the issue asks for. Now passes the resolved model and
emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the
existing key convention rather than introducing a nested object.

BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed
a nested {"effort": ...} sample; the real result is flat and those keys were
absent entirely. A reference doc asserting a JSON path a reader can copy is worse
than no doc. Corrected against the actual emitted key set.

MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex
kept minimal in its supported set and still clamped max down to xhigh, so the
invocation-time and install-time channels disagreed about the same runtime's
capability: --host codex with max emitted xhigh while the generated TOML said
max. This is the repo's documented generative-fix-divergence class, so both
tables now cross-reference each other and a parity test fails if they ever
diverge again.

MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null
_baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback
never fired and every effort rendered as null. And a non-array value made the Set
constructor throw at module load — model-catalog.cjs is required across the whole
CLI, so one bad JSON value killed every command, not just codex effort. Guarded
on size and filtered to array values; both degrade to the hardcoded baseline.

MAJOR — value widened to a nullable string with two consumers left behind.
runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a
null effort key in generated frontmatter); install-effort-resolver still declared
a non-nullable return, a structural lie that silently defeated null checking.
Both corrected, both omitting the key on null — the same posture as 'inherit',
where omission means "follow the host default".

MAJOR — the per-model table is inert today, and the docs now say so. All three
shipped models advertise the same usable range and ultra (sol's only
differentiator) is rejected for every model, so no observable output differs by
model. The table stays because Codex declares capability per model and the sets
are free to diverge — a single per-runtime assumption is precisely what went
stale and produced this issue — but overselling it as a visible per-model feature
would have been the same class of error as the doc blocker above.

Tests: three passed under a full revert and are strengthened rather than deleted,
since each guards a real contract (#3533's inherit rule, the undeclared-host
rule, off-ladder handling) — they now also assert the clamp-visibility fields,
which only exist after this change. The fast-check property is kept for its
shrinking, and a deterministic nested loop over the full cross-product now sits
beside it so coverage is exhaustive rather than sampled.

Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg
form and would have written a literal null reasoning effort on the ultra path;
CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT.
The installer defect was found by the co-change gate, not by a reviewer —
install.js is a historical co-change partner of model-catalog.cts that this diff
had not touched.

* test(#3007): correct assertions that pinned Codex's stale effort premise

Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the
shipped commit. Every one is a stale pin, not a defect: each was probed against
the built module before its expectation was changed, and none failed for a
reason other than this premise correction.

Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale
by a production change must not ride inside another commit, because the
release-sdk hotfix cherry-pick filter routes by subject prefix and a correction
buried under the wrong prefix ships a half-state (v1.42.3, #3621).

The most valuable one was tests/model-resolver.test.cjs's cross-provider
validity invariant, which hardcoded the Codex enum as
`minimal|low|medium|high|xhigh` and failed with "real API would 400". That
message is now false in both directions: Codex accepts `max`, and rejects
`minimal`, which no model advertises. The enum is corrected to
`low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the
"would the real API refuse this" check worth having, and it was right to fail
here. It simply carried the stale fact in its own fixture.

Test NAMES were corrected alongside their assertions wherever the name asserted
the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal
passthrough". A renamed test that still claims the old thing is worse than a
failing one, and a green test whose name states a falsehood is how the next
reader inherits the wrong premise.

Both channels are covered: install-time (renderEffortForRuntime, and the
generated .toml in install-runtime-artifacts) and invocation-time argv
(effort-surface-axis). They were deliberately brought into agreement in this
change, so their assertions had to move together.

Each site carries a #3007 comment recording that Codex gained max/ultra and that
capability is declared per model, so a future reader can tell this was a
deliberate premise correction rather than a test bent to fit an implementation.

* test(#3007): separate the effort-precedence case from the clamp case

The previous stale-assertion pass over-corrected one test. It saw
`effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`,
assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to
"max passes through" and changed the expectation to `max`. The remote runner
disagreed.

Reproduced against the real CLI: with that config and `gsd-planner`, the
resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`,
`effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE —
gsd-planner is heavy/opus tier and its routing-tier default outranks
`effort.default` — so `max` never reaches the renderer at all. The test says
nothing about clamping and never did; it only looked like a clamp pin because
both mechanisms happened to yield the same string.

Restored to `xhigh` and renamed to say what it actually tests. It now also
asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is
what makes it impossible to mistake for a clamp pin again: those two fields prove
the value is what the resolver produced rather than something the renderer
downgraded. Before #3007 there was no way to tell the two apart from the output —
which is precisely why the previous pass could not tell them apart either.

Added the test that was actually missing: `effort.agent_overrides`, which
outranks the tier default, so the requested level genuinely reaches the renderer
and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified
by probe before asserting.

One test now pins the precedence rule and the other pins the #3007 behavior, and
neither can be read as the other. That the clamp-visibility fields are what
resolved this is a small argument for having added them.

* chore(#3007): backfill changeset pr number to 3765

* test(#3007): put model-catalog under the mutation gate

The Stryker shard showed as `skipping` on this PR despite the diff rewriting
model-catalog's effort logic. That was legitimate, not a detection bug:
`model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the
whole module — including everything #3007 touches — sat entirely outside
mutation scoring with has_work "false".

Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs
is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no
temp dirs. That shape is not stylistic — it is the #2790 precedent this file
already documents. Stryker's command runner treats a whole `node --test <file>`
invocation as ONE test costing whatever its slowest case costs, and re-runs it
per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses
runGsdTools throughout) would reproduce exactly the 15-minute shard-cap
cancellation #2790 hit. The integration file is unaffected and keeps running in
full in the normal test job.

Coverage spans the module rather than only the diff, because the score is
measured over the whole file: effort rendering across every model and ladder
level in both channels, the prototype-chain host guard, the exported enums and
maps, isAnthropicFlavoredModel's provider namespacings, the profile projections,
nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are
worth naming — every uncovered exported function is score given away, and
mergeEffortTierDefaults turned out to have a genuinely interesting contract
(#3531: a partial override merges over the built-ins rather than replacing them,
and isValid gates the VALUE, not the tier name, so an unknown tier key is still
merged in). Every expectation was probed against the built module before being
asserted.

minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment.
Floors in this repo are measured, not chosen — the existing entries sit at 94, 75
and 56 — and they can only be measured in CI, because mutation shards run
`node --test`, which is hard-blocked locally. The first CI run on this branch
reports the real number and the floor gets ratcheted to it before merge. A
placeholder of 1 reaching `next` would make the gate decorative: it would pass
whether or not a single mutant is ever killed.

Note the target is "never regress from measured", not a fixed 80 — planning-inspect
sits at 56 and is documented as an accepted ratchet candidate.

* test(#3007): bootstrap model-catalog's mutation floor legally

The placeholder floor was structurally illegal and the remote run said so.
tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module
must carry a matching RATCHET_BASELINE entry in the same diff, minScore must
EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three.
That is the ratchet working exactly as intended — a floor nobody can satisfy
accidentally is the point of it.

Bootstrapped at 50 in both places. Fifty is not a measured score and the comment
says so plainly: it is the minimum the guard permits, and it coincides with
Stryker's own configured `break` threshold, so it is the lowest legal starting
point for a module that has never been measured. It still must be ratcheted to
floor(measured) - 1 before this PR merges.

Also corrected a real defect in the file's own instructions. "HOW TO UPDATE"
step 1 read "Run the per-module Stryker shard locally" — which cannot be done
here, and which the same file contradicts eighty lines further down, where the
#2790 scores are recorded as "not a local run; mutation shards run `node --test`,
hard-blocked in this repo's local environment". stryker.config.mjs confirms the
command runner invokes `node --test` once per mutant, and
.claude/hooks/block-local-node-test.sh denies exactly that. So the documented
first step sends the next contributor at a wall. Rewritten to describe the path
that works — push, read the measured score off the CI shard, then set the floor
and its baseline together in one diff — and to say why local measurement is not
available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched.

The two-step is inherent to the environment rather than a shortcut: a floor
cannot be measured before the first CI run exists, and the guard rightly refuses
to accept an unmeasured one below its minimum.

* test(#3007): ratchet model-catalog's mutation floor to its measured score

The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no
timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this
file's own rule, minScore = floor(measured) - 1, which is the same arithmetic
every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94.

Both halves moved together, because the ratchet guard asserts minScore equals its
RATCHET_BASELINE entry and would reject them drifting apart.

The spawn-free unit surface is vindicated by the clock: 57 seconds, against a
15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the
whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing
the shard at tests/model-resolver.test.cjs — #2790 recorded shards being
CANCELLED at that cap when they targeted a runGsdTools-heavy integration file.

The registry comment is rewritten rather than deleted. It previously warned that
the floor was provisional and must not ship that way; leaving that text next to a
measured floor would make the file lie in the other direction. It now records the
measurement the way the sibling entries do, including that 59.62 sits below
TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 —
comfortably clear of its own floor with real room to grow. Raise it as the tests
improve; never lower it.

Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is
an honest floor for a module that had NO mutation coverage at all an hour ago,
and it is now pinned so it cannot silently regress.

---------

Co-authored-by: sim <sim@local>
2026-08-22 20:51:55 -04:00

481 lines
20 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env node
'use strict';
/**
* scripts/mutation-matrix.cjs
*
* Single source of truth for the ADR-457 Stryker mutation gate dynamic matrix.
*
* Computes which covered modules changed vs a base ref and emits a GitHub
* Actions matrix JSON so CI can run one Stryker shard per changed module in
* parallel rather than a single serial run over all modules.
*
* Usage:
* node scripts/mutation-matrix.cjs --base origin/next
* printf 'src/config-schema.cts\n' | node scripts/mutation-matrix.cjs
* node scripts/mutation-matrix.cjs --base origin/next --print
*
* Output (stdout, default): JSON object
* {
* "has_work": "true"|"false",
* "matrix": {
* "include": [
* { "name": "<module>", "mutate": "gsd-core/bin/lib/<module>.cjs", "tests": "<space-joined test files>" },
* ...
* ]
* }
* }
*
* Exit codes: 0 always (empty matrix is not an error, has_work "false").
*/
const { execFileSync } = require('child_process');
const fs = require('fs');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
// ── Resilient stdin reader ────────────────────────────────────────────────────
// On macOS, libuv sets the stdin pipe fd to non-blocking mode. A synchronous
// readFileSync(process.stdin.fd) can therefore throw EAGAIN ("resource
// temporarily unavailable") when the writer hasn't yet filled the pipe — this
// is intermittent under heavy CI shard load and causes a spurious status 2
// exit. We work around it by calling fs.readSync in a loop and retrying on
// EAGAIN with a 1 ms synchronous pause (Atomics.wait on a fresh SharedArrayBuffer
// — no hot spin, no real-clock dependency, works under --experimental-vm-modules).
/**
* Read all of stdin synchronously, retrying on EAGAIN.
*
* @returns {string} UTF-8 decoded full stdin content.
*/
function readStdinSync() {
const BUF_SIZE = 64 * 1024; // 64 KB chunks
const buf = Buffer.allocUnsafe(BUF_SIZE);
const chunks = [];
for (;;) {
let bytesRead;
try {
bytesRead = fs.readSync(process.stdin.fd, buf, 0, BUF_SIZE, null);
} catch (err) {
if (err.code === 'EAGAIN') {
// Non-blocking pipe not yet ready — yield for ~1 ms then retry.
Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), 0, 0, 1);
continue;
}
if (err.code === 'EOF') {
break;
}
throw err;
}
if (bytesRead === 0) {
break; // Clean EOF
}
chunks.push(Buffer.from(buf.slice(0, bytesRead)));
}
return Buffer.concat(chunks).toString('utf8');
}
// ── Per-module mutation score ratchet ─────────────────────────────────────────
// ADR-456 / issue #1187: every covered module declares a minScore floor.
//
// HOW THE RATCHET WORKS:
// • minScore locks in the current measured mutation score (minus a 1–2 pt
// margin for run-to-run timeout variance).
// • CI fails a shard if the module's live score drops below its minScore.
// • Raise minScore (never lower) as a module's tests improve.
// • The goal is every module reaching TARGET_MUTATION_SCORE (80).
//
// GOODHART SAFETY: scores are improved by writing genuine behavioural
// assertions that kill real mutants — never by adding brittle exact-string
// matches on incidental output. A justified `// Stryker disable` on a
// confirmed equivalent mutant is acceptable.
//
// HOW TO UPDATE:
// 1. The per-module Stryker shard CANNOT be run locally: Stryker's command
// runner invokes `node --test` once per mutant (see stryker.config.mjs),
// and this repo hard-blocks local `node --test` via
// .claude/hooks/block-local-node-test.sh. Push the branch instead and
// let CI run the shard for the changed module.
// 2. Read the measured score from the CI shard's output.
// 3. Set minScore = floor(measured) - 1 (never lower than current value)
// and update the matching RATCHET_BASELINE entry in the same diff.
// 4. Open/update the PR — the CI gate will enforce the new floor on every
// future run.
/** Long-run target for all modules (ADR-456). */
const TARGET_MUTATION_SCORE = 80;
// ── Single source of truth: covered modules ───────────────────────────────────
// Each entry: { cjs: '<built artifact>', tests: ['tests/...', ...], minScore: N }
//
// minScore is the CI break threshold for this module's shard.
// Floors are measured scores minus 1–2 pts for run-to-run variance.
// Measured CI scores 2026-06-14 (issue #1187, timeout-free — source of truth):
// context-utilization 92.31% → floor 80 (target already met)
// prompt-budget 68.33% → floor 66 (local was 99.6% — TIMEOUT INFLATION; CI is the truth)
// frontmatter 63.35% → floor 62
// adr-parser 69.30% → floor 68
// config-schema 54.55% → floor 52 (local was 69.7% — TIMEOUT INFLATION; CI is the truth)
// active-workstream-store 81.91% → floor 80
// core-utils 77.52% → floor 75
//
// LESSON: floors MUST be calibrated from CI mutation runs (CI runs with
// timeout≈0, deterministic). Local runs count timeouts as kills and
// inflate scores significantly (prompt-budget: 99.6% local vs 68.3% CI;
// config-schema: 69.7% local vs 54.55% CI). Never set a floor from a
// local run without CI cross-check.
const COVERED = {
'context-utilization': {
cjs: 'gsd-core/bin/lib/context-utilization.cjs',
tests: [
'tests/context-utilization.property.test.cjs',
],
// After mutation-killer assertions added in #1187: measured 92.31% (2026-06-14).
// 3 survivors are __esModule boilerplate (genuinely equivalent CJS interop mutants).
// minScore raised to TARGET (80) — module now meets ADR-456 goal.
minScore: 80,
},
// context-composer: extracted from prompt-budget by #2929. Needs its own entry because
// mutation coverage does not migrate with relocated code — scoring only prompt-budget.cjs
// would leave the extracted ladder unmeasured.
'context-composer': {
cjs: 'gsd-core/bin/lib/context-composer.cjs',
tests: [
'tests/prompt-budget-parity.test.cjs',
'tests/prompt-budget.unit.test.cjs',
'tests/context-composer.test.cjs',
'tests/context-composer.property.test.cjs',
],
minScore: 66,
},
'prompt-budget': {
cjs: 'gsd-core/bin/lib/prompt-budget.cjs',
tests: [
'tests/prompt-budget.property.test.cjs',
'tests/prompt-budget.unit.test.cjs',
],
// CI 68.33% timeout-free (164 killed / 1 timeout / 240 total) 2026-06-14;
// local was 99.6% — timeout inflation. Floor = 68 - 2 margin.
minScore: 66,
},
frontmatter: {
cjs: 'gsd-core/bin/lib/frontmatter.cjs',
tests: [
'tests/frontmatter.property.test.cjs',
'tests/frontmatter.unit.test.cjs',
// #1882 added the unterminated-fence detection to frontmatter.cjs, and the tests that
// constrain it live here. Without this entry the mutants in that branch are covered by
// no test in the shard, so the module's score drops even though the behaviour is tested.
'tests/unusable-input.test.cjs',
],
minScore: 62,
},
'adr-parser': {
cjs: 'gsd-core/bin/lib/adr-parser.cjs',
tests: [
'tests/adr-parser.property.test.cjs',
'tests/adr-parser.test.cjs',
'tests/adr-parser.unit.test.cjs',
],
minScore: 68,
},
'config-schema': {
cjs: 'gsd-core/bin/lib/config-schema.cjs',
tests: [
'tests/config-schema.property.test.cjs',
],
// CI 54.55% timeout-free (18 killed / 0 timeout / 33 total) 2026-06-14;
// local was 69.7% — timeout inflation. Floor = 54 - 2 margin.
minScore: 52,
},
'active-workstream-store': {
cjs: 'gsd-core/bin/lib/active-workstream-store.cjs',
tests: [
'tests/active-workstream-store.test.cjs',
'tests/active-workstream-store.unit.test.cjs',
],
minScore: 80,
},
'core-utils': {
cjs: 'gsd-core/bin/lib/core-utils.cjs',
tests: [
'tests/core-utils.test.cjs',
],
minScore: 75, // measured 77.52% (2026-06-14, issue #1187); floor = 77 - 2
},
// planning-inspect / plan-document / planning-command-router: net-new modules
// added by #2790. Registered here so the Stryker gate stops SKIPPING them
// (previously has_work: "false" — ~1000 LOC entirely outside mutation scoring).
//
// WHY THESE SHARDS POINT AT tests/planning-inspect.unit.test.cjs, NOT
// tests/planning-inspect.test.cjs. CI evidence: two shards pointed at the
// integration file were CANCELLED at the workflow's 15-minute cap —
// "Mutation testing 4% (elapsed: ~3m, remaining: ~1h 19m) 27/640 tested".
// tests/planning-inspect.test.cjs is INTEGRATION-shaped (91 cases, most
// spawning a `gsd-tools` child process via `runGsdTools`); Stryker's command
// runner treats the whole `node --test <file>` invocation as ONE test costing
// whatever the slowest case costs (measured ~20s), and re-runs that entire
// file once per mutant — 640 mutants x 20s cannot finish in 15 minutes.
// tests/planning-inspect.unit.test.cjs is the dedicated, spawn-free,
// in-process mutation surface for exactly these three modules (measured
// locally: the whole file runs in well under a second) — the same shape
// every other entry in this registry already uses (*.property.test.cjs /
// *.unit.test.cjs). The integration suite is UNAFFECTED by this change: it
// keeps running in full in the normal (non-mutation) test job, and remains
// the source of truth for spawn-boundary/CLI-dispatch/read-only-proof
// behaviour that an in-process unit file cannot exercise.
//
// Measured CI scores (GitHub Actions run 32392791843, all three shards
// PASSED — not a local run; mutation shards run `node --test`, hard-blocked
// in this repo's local environment):
// planning-command-router 95.65% → floor 94 (already exceeds TARGET_MUTATION_SCORE (80))
// plan-document 76.58% → floor 75
// planning-inspect 57.03% → floor 56 (well below TARGET (80) — ratchet
// candidate; comfortably clears its own floor but has real room to grow.
// Raise as its tests improve, never lower it.)
//
// All three shards point at tests/planning-inspect.unit.test.cjs (in-process,
// spawn-free, ~0.3s dry run), not tests/planning-inspect.test.cjs — that is
// what made measurement possible at all. The integration file spawns a
// subprocess per case via runGsdTools; Stryker's command runner treats the
// whole `node --test <file>` invocation as one test costing whatever the
// slowest case costs (measured ~20s), and re-runs that entire file once per
// mutant, so 640 mutants x 20s could not finish inside the 15-minute shard
// cap. The integration suite is unaffected by this change: it keeps running
// in full in the normal (non-mutation) test job.
'planning-inspect': {
cjs: 'gsd-core/bin/lib/planning-inspect.cjs',
tests: [
'tests/planning-inspect.unit.test.cjs',
],
minScore: 56,
},
'plan-document': {
cjs: 'gsd-core/bin/lib/plan-document.cjs',
tests: [
'tests/planning-inspect.unit.test.cjs',
],
minScore: 75,
},
'planning-command-router': {
cjs: 'gsd-core/bin/lib/planning-command-router.cjs',
tests: [
'tests/planning-inspect.unit.test.cjs',
],
minScore: 94,
},
// model-catalog: net-new registration by #3007. The module was entirely
// outside mutation scoring (has_work: "false") before this entry, so the
// #3007 per-model Codex effort rewrite (renderEffortForRuntime's
// CODEX_MODEL_EFFORT lookup, the 'ultra' policy rejection, the ladder
// walk-up clamp) had zero mutation coverage.
//
// Same #2790 precedent as planning-inspect above: this shard points at a
// dedicated tests/model-catalog.unit.test.cjs, NOT tests/model-resolver.test.cjs
// — that integration file uses runGsdTools heavily and would hit the same
// 15-minute shard-cap cancellation #2790 documented (a `node --test <file>`
// invocation is ONE test costing whatever its slowest case costs, re-run
// per mutant). tests/model-catalog.unit.test.cjs is spawn-free, in-process,
// and runs in well under a second.
//
// Measured CI score (GitHub Actions run 32605073352, job 97108869486):
// model-catalog 59.62% → floor 58 (248 killed, 168 survived, 0 timeouts,
// 0 errors; below TARGET_MUTATION_SCORE (80) — ratchet candidate like
// planning-inspect (56): comfortably clears its own floor but has real
// room to grow. Raise as its tests improve, never lower it.)
// Floor follows this file's documented rule, minScore = floor(measured) - 1,
// matching the sibling precedent exactly (57.03 → 56, 76.58 → 75, 95.65 → 94).
//
// The shard completed in 57 seconds — concrete evidence the spawn-free
// unit-file design above worked: the #2790 precedent's 15-minute shard-cap
// cancellations do not apply here, and for comparison the `frontmatter`
// shard in the same run took 9m46s.
'model-catalog': {
cjs: 'gsd-core/bin/lib/model-catalog.cjs',
tests: ['tests/model-catalog.unit.test.cjs'],
minScore: 58,
},
};
// ── Files that, when changed, invalidate ALL modules ─────────────────────────
// Changes to the Stryker config, this script itself, or any covered test file
// affect all mutation scores and must force a full re-run.
const GLOBAL_TRIGGERS = new Set([
'stryker.config.mjs',
'scripts/mutation-matrix.cjs',
]);
// Also flag all test files that belong to any covered module as global triggers.
for (const mod of Object.values(COVERED)) {
for (const t of mod.tests) {
GLOBAL_TRIGGERS.add(t);
}
}
// ── Argument parsing ──────────────────────────────────────────────────────────
function parseArgs(argv) {
const out = { base: null, print: false };
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
if (arg === '--base') {
out.base = argv[++i];
if (!out.base || out.base.startsWith('--')) {
throw new Error('--base requires a value');
}
} else if (arg.startsWith('--base=')) {
out.base = arg.slice('--base='.length);
if (!out.base) throw new Error('--base requires a value');
} else if (arg === '--print') {
out.print = true;
} else if (arg === '--help' || arg === '-h') {
console.log([
'Usage:',
' node scripts/mutation-matrix.cjs --base <ref> [--print]',
' printf "src/foo.cts\\n" | node scripts/mutation-matrix.cjs [--print]',
'',
'Options:',
' --base <ref> Git ref to diff against (default: origin/${GITHUB_BASE_REF:-next})',
' --print Human-readable output instead of JSON',
].join('\n'));
throw new ExitError(0);
} else {
throw new Error(`unknown argument: ${arg}`);
}
}
return out;
}
// ── Changed-file resolution ───────────────────────────────────────────────────
function resolveChangedFiles(args) {
// When --base is provided, always use git diff (regardless of stdin).
// When --base is absent AND stdin is not a TTY (isTTY is falsy / undefined),
// read a newline-delimited file list from stdin.
if (!args.base && process.stdin.isTTY !== true) {
const raw = readStdinSync();
return raw.split('\n').map(l => l.trim()).filter(Boolean);
}
// Otherwise (--base given, or stdin is a real TTY), diff against the base ref.
const defaultBase = `origin/${process.env.GITHUB_BASE_REF || 'next'}`;
const base = args.base || defaultBase;
const stdout = execFileSync('git', ['diff', '--name-only', `${base}...HEAD`], {
encoding: 'utf8',
});
return stdout.split('\n').map(l => l.trim()).filter(Boolean);
}
// ── Module classification ─────────────────────────────────────────────────────
function computeMatrix(changedFiles) {
// Check for global triggers first — if any hit, include every covered module.
const allModuleNames = Object.keys(COVERED);
for (const f of changedFiles) {
if (GLOBAL_TRIGGERS.has(f)) {
return allModuleNames;
}
}
// Otherwise find which modules have their src/*.cts changed.
const changed = new Set();
for (const f of changedFiles) {
// Match src/<module>.cts (top-level src/, not nested)
const m = f.match(/^src\/([^/]+)\.cts$/);
if (m && COVERED[m[1]]) {
changed.add(m[1]);
}
}
return [...changed];
}
// ── Output formatting ─────────────────────────────────────────────────────────
function buildResult(moduleNames) {
const include = moduleNames.map(name => ({
name,
mutate: COVERED[name].cjs,
tests: COVERED[name].tests.join(' '),
minScore: COVERED[name].minScore,
}));
return {
has_work: include.length > 0 ? 'true' : 'false',
matrix: { include },
};
}
function printHuman(result, changedFiles) {
console.log(`Changed files (${changedFiles.length}):`);
for (const f of changedFiles) console.log(` ${f}`);
console.log('');
console.log(`has_work: ${result.has_work}`);
console.log(`Shards (${result.matrix.include.length}):`);
for (const shard of result.matrix.include) {
console.log(` [${shard.name}]`);
console.log(` mutate: ${shard.mutate}`);
console.log(` tests: ${shard.tests}`);
console.log(` minScore: ${shard.minScore}`);
}
}
// ── Main ──────────────────────────────────────────────────────────────────────
function main() {
try {
const args = parseArgs(process.argv.slice(2));
const changedFiles = resolveChangedFiles(args);
const moduleNames = computeMatrix(changedFiles);
const result = buildResult(moduleNames);
if (args.print) {
printHuman(result, changedFiles);
} else {
console.log(JSON.stringify(result, null, 2));
}
} catch (err) {
if (err instanceof ExitError) throw err;
console.error(`mutation-matrix: ${err.message}`);
throw new ExitError(2);
}
}
// ── MUTATION_BREAK resolver ───────────────────────────────────────────────────
/**
* Resolves the per-shard mutation break threshold from the MUTATION_BREAK env var.
*
* Fail-closed contract:
* - undefined → 60 (local run: no env set, documented backstop)
* - set but empty (e.g. CI matrix.minScore missing) → throws (wiring error)
* - non-numeric or out-of-range [1, 100] → throws (invalid config)
* - valid integer string → returns that number
*
* This function is the single call site for reading MUTATION_BREAK.
* stryker.config.mjs imports and calls it so CI shards with a bad
* MUTATION_BREAK fail immediately rather than silently falling back to 60
* and bypassing a per-module floor above 60 (e.g. prompt-budget: 90).
*
* @param {string|undefined} raw - value of process.env.MUTATION_BREAK
* @returns {number}
*/
function resolveMutationBreak(raw) {
if (raw === undefined) {
// Local run with no MUTATION_BREAK set — use documented backstop.
return 60;
}
if (typeof raw !== 'string' || raw.trim() === '') {
throw new Error(
'MUTATION_BREAK is set but empty — CI shard wiring is broken (matrix.minScore missing?)'
);
}
const n = Number(raw);
if (!Number.isFinite(n) || n < 1 || n > 100) {
throw new Error(
`MUTATION_BREAK invalid: "${raw}" (expected a per-module minScore 1-100)`
);
}
return n;
}
// Export internals for programmatic use (tests/mutation-matrix-ratchet.test.cjs).
// The require.main guard prevents main() from running when this file is require()d.
module.exports = { COVERED, TARGET_MUTATION_SCORE, resolveMutationBreak, readStdinSync };
if (require.main === module) runMain(main);