Files
msd-core/scripts/mutation-matrix.cjs
Tom Boucher 382bf7c423 fix(#3706): deliver the resolved reasoning effort to OpenCode subagents (#3867)
* test(#3706): failing-first coverage for OpenCode variant emission and frontmatter escaping

* fix(#3706): emit the resolved reasoning effort as OpenCode's variant key

`query resolve-execution` resolved an effort level for every agent, but the
OpenCode bake wrote only `model:` — the effort never reached the generated
agent, so subagents ran at whatever the runtime defaulted the model to. This
is the effort-side twin of the model-side defect fixed in #3705.

The key is written only when an `effort` block is actually configured.
`resolveInstallTimeEffort` always returns a level (the catalog default is
`high`), so gating on its return value would stamp `variant: high` into every
existing OpenCode install — and OpenCode resolves a variant name against a
`variants` map in the user's `opencode.jsonc`, so a value nobody declared is
not a safe default. Gating on `readGsdEffectiveEffortConfig` keeps installs
that never asked for effort routing byte-identical.

Kilo does not receive the key: `EFFORT_ARGV` declares surfaces for claude,
opencode and codex and has no kilo entry. This is deliberately asymmetric with
the model side, where #2794 J8 requires the two runtimes to resolve alike.

Both frontmatter sinks now route through `frontmatterScalar`, which quotes and
escapes any value that is not a plain scalar. The raw interpolation predates
this change, but it was already shown by execution during the #3705 security
review to let a config value containing a newline inject additional top-level
keys (`tools:`, `permission:`) into a generated agent file. This change adds a
second write to that sink, so it is closed here rather than doubled.

* fix(#3706): quote frontmatter values YAML would not read back verbatim

Self-review of the predicate added in the previous commit. Treating
/^[A-Za-z0-9._:/@+-]+$/ as 'safe to emit bare' answers the wrong question:
a value can match it and still not round-trip.

  - A leading '@' is a YAML *reserved* indicator and may not open a plain
    scalar at all, so a scoped ID like '@org/model' emitted bare is a parse
    error, not an ambiguity — the whole agent file becomes unreadable.
  - 'no' / 'y' / 'off' / 'null' resolve to booleans and null, so a variant
    with one of those names would match no entry in the user's variants map.
  - '12:30' resolves to 750 under YAML 1.1 sexagesimal, and ':' is legal
    mid-identifier here, so the form is reachable rather than contrived.

Real model IDs pass every clause and stay bare, so already-generated files
remain byte-identical.

* fix(#3706): route variant through the declared effort seam and cover the live path

Addresses six findings from the isolated review, all confirmed by execution.

The tests were the serious one: they required `../bin/install.js` while the fix
landed in src/, which compiles to gsd-core/bin/lib/. They exercised a different
copy of the converter than the one the bake actually uses, so the whole suite
was green-by-construction against unchanged code and the remote run failed all
13. Every case now runs against BOTH copies from one table, which doubles as the
parity assertion the generative-fix note in runtime-artifact-conversion.cts asks
for, and bin/install.js carries the mirrored change.

Emission no longer hand-rolls the value. It goes through `renderEffortArgv`,
the declared OpenCode effort seam (EFFORT_ARGV.opencode: its own supported set
and clamp). That is what rejects a level that is not a wire value — above all
`inherit`, which per #3533 (10d) means "omit the key and follow the host
default" and was previously written literally, naming a variant that cannot
resolve. Reachable two ways, both now pinned: an agent_overrides entry and a
routing_tier_defaults entry. A bare effort.default does NOT reach a tiered
agent (the #3531 tier ladder answers first), so a test written against
`default` alone asserts nothing — that is pinned too.

The plain-scalar decision moved into frontmatter.cts beside
`scalarNeedsDoubleQuoting` rather than sitting next to it as a second, weaker
predicate. `agentScalarNeedsDoubleQuoting` is a documented superset: it adds a
trailing `:` (read as a nested mapping key, which fails the whole frontmatter),
boolean/null words, and numeric-looking values including YAML 1.1 sexagesimal.

Docs now state the cascade plainly: the gate is on effort being configured at
all, not on the individual agent being named, so every generated OpenCode agent
gets a variant line once any effort block exists.

* test(#3706): assert the two frontmatterScalar copies cannot diverge

A hand-picked adversarial corpus plus a fast-check property over
YAML-significant strings, both run against bin/install.js and the live
src copy. Verified the property can actually fail: mutating one copy's
quoting rule is killed well inside the run budget.

* fix(#3706): close the review findings — predicate, seam, and dead mirror

Third review round; every item below was confirmed by execution.

The scalar predicate was wrong in two families, both found by a round-trip
property test rather than by reading. Basing it on scalarNeedsDoubleQuoting
dropped the "first character must be alphanumeric" clause, so `~`, `.inf`,
`.nan`, `+1`, `-0` and `.5` went out bare and came back as null/floats/ints;
and that base predicate only inspects the FIRST character, so an embedded `: `
(a nested mapping, i.e. a parse error) or ` #` (a comment, i.e. silent
truncation) also passed. Dates round out the set: `2026-08-25` opens
alphanumeric, survives every other clause, and YAML resolves it to a Date.
The property now asserts the contract directly over generated values instead
of trusting an enumerated character list.

The bin/install.js mirror is gone. Its premise was false — install.js already
requires bin/lib at :65 — and it was unreachable besides: install.js's
convertClaudeToOpencodeFrontmatter has no `isAgent: true` call site, because
its agents path resolves converters from the compiled module. It was a third
copy of the YAML rules serving a test rather than a caller, so the file is
back to origin/next and the tests target the live copy only.

Effort clamping moved to `clampEffortForHost`, which renderEffortArgv now
delegates to. The layout was calling renderEffortArgv with a hardcoded 'argv'
to borrow its clamp, which read as if the frontmatter key were gated on the
invocation-time axis. It is not: claude declares effortSurface "argv" and
independently bakes an effort: key. One capability table, one clamp, two
channels that no longer pretend to be each other.

Also corrects an earlier claim of mine: adding EFFORT_RENDERING.opencode would
NOT have made `effort sync` write the wrong key, because it guards on the
runtime name before it ever renders. The seam choice stands on other grounds.
`effort sync` still skips OpenCode, but its stated reason claimed OpenCode
"does not use effort: frontmatter", which this change makes false — so the
message now says what is actually true.

* docs(#3706): restate the changeset around the round-trip contract

* fix(#3706): restore the changeset fragment belonging to #3809

An earlier commit in this branch picked the first file in .changeset/ by
glob order instead of the fragment created for this issue, and overwrote
agile-geese-squeak.md (PR 3815 / #3809) with this change's body. Restored
verbatim from origin/next; this change's text now lives in its own
patient-cranes-parade.md, where it was created.

* feat(#3706): maintain the OpenCode variant key from effort sync

Install bakes the resolved effort into OpenCode agent frontmatter as
`variant:`, so `effort sync` has to maintain it or a config change only takes
effect on reinstall — and its skip message claimed OpenCode does not use
frontmatter effort at all, which this issue made false.

cmdEffortSyncOpencode mirrors the codex branch: resolve per agent, clamp
through the declared OpenCode capability, then write, strip, or skip. A null
target means the key must not exist, which covers both "no effort configured"
and "resolved to inherit or to an unsupported level" — the same states under
which install writes nothing, so sync and install agree by construction.

The frontmatter line-editors are key-parameterised rather than copied:
setEffortFrontmatter / removeEffortFrontmatter are now thin wrappers over the
same internals the variant path uses, and a test pins that the claude `effort:`
behavior did not move. The child-process test harness fixes both HOME and
USERPROFILE, so the hermetic-config assertions cannot pass vacuously on Windows.

* fix(#3706): scope the frontmatter line editors to the matched block

Found by the security review of the sync path, reported as correctness rather
than vulnerability, and reproduced against pre-fix code before being fixed.

Both editors matched the frontmatter with a regex that can match a block after
a preamble, then derived the EOL and the opening-fence length from the START OF
THE FILE. On a CRLF document with a preamble those disagree, the offsets shift
by one byte, and the reassembled document comes back with a mangled fence
(`---\rname: x`). Both now take the EOL from the matched block.

`setFrontmatterKeyLine` additionally did a whole-file `/m` replace when the key
already existed, gated only on the key being present in the frontmatter body —
so a preamble line starting with the same key was rewritten instead of the
frontmatter one. It now replaces inside the frontmatter span only, which is the
hazard `removeFrontmatterKeyLine` already documented and guarded against.

Neither is reachable from an install-written `gsd-*.md` (those begin at byte 0
with `---`), and both predate this change — but the editors are in this diff
because #3706 key-parameterised them, so they are fixed here rather than left
for the next caller to trip over. Three regression tests, each confirmed to
fail against the pre-fix build.

* fix(#3706): treat a present-but-empty key as present, and pin the real seam

Fourth review round.

The MAJOR one: both sync branches read the current value with `(.+?)`, which
needs at least one character, so a key present with an EMPTY value read as
"key absent". When the target was also null the code concluded "already
correct" and skipped — leaving the key in the file, where it reads back as
YAML `null`: exactly the unresolvable-variant state this change exists to
prevent. Whitespace decided whether it fired, since `variant:   ` matched and
`variant:` did not. Presence and value are now separate questions at both the
opencode and the claude branch.

The OpenCode writer now follows the codex branch rather than the claude one:
tmp file plus retryRenameSync with orphan cleanup, and a write failure skips
that agent and is reported instead of aborting the sweep. Same granularity,
same transient-Windows-lock exposure, so the hardened sibling was the right
precedent.

Also: the generic line-editors escape their interpolated key, the JSDoc
stranded by the clampEffortForHost extraction is back on renderEffortArgv, and
a cast that declared a nullable function as non-nullable is corrected.

Tests close the gaps the review listed — empty value (both spellings), CRLF
round-trip through write and strip, the symlink guard, a body line starting
`variant:`, a file with no frontmatter, and the YAML classes that actually
broke the predicate. The new layout-seam test drives the real stage() path and
was verified to FAIL when `variant` is removed from the converter call; a seam
test that survives cutting the seam is worse than none.

* fix(#3706): clear the round-five review findings

No blockers or majors this round; the repo's review gate is zero-tolerance, so
the minors are cleared too.

A duplicated key was only half-stripped: the strip regex had no `g` flag, so a
frontmatter carrying the key twice lost one occurrence, reported success, and
left the "a null target means the key must not exist" invariant false on disk —
converging only on a second run. Such a document is already invalid YAML, so
this is robustness rather than a live corruption path, but a successful sync
has to leave the invariant true.

A run in which every write failed still summarised as `ok`, so a caller could
not tell "nothing to do" from "everything failed". The OpenCode branch now
reports `failed` when any write failed. The write-failure path was also the
newest code in the change with no coverage at all; it now has a test that
injects the failure by monkeypatching the write, per CLAUDE.md §4, rather than
by chmod — mode bits do not bite under root in CI.

`CodexEffortSyncWriteFailure` is renamed `EffortSyncWriteFailure` now that two
branches share it. Removed a guard on the claude concrete path that was
provably unreachable — no member of EFFORT_SET renders null there, so it read
as protection that did not exist. The claude inherit path's presence check is
load-bearing and untouched.

Three stale statements corrected: the OpenCode result shape matches codex's,
not claude's, now that it emits write_failures; the `thread()` test helper now
calls `clampEffortForHost` so it genuinely mirrors the layout instead of
merely claiming to; and a test helper restored `USERPROFILE` by assignment,
writing the literal string "undefined" into the environment on POSIX — it
deletes now.

* fix(#3706): converge the set path, degrade on unreadable files, preserve mode

Rounds five and six of review. No blockers or majors; the review gate is
zero-tolerance, so the minors are cleared too.

`setFrontmatterKeyLine` was the mirror of a defect already fixed in its
sibling: `remove` was made global, `set` was not, so on a frontmatter carrying
the key twice it rewrote the first and left a stale second. Last-wins YAML
readers honour the stale value while the sync's own first-occurrence read
reports "in sync" — permanently non-converging. It now collapses to exactly one
occurrence, in the position of the first, so ordinary single-occurrence
documents stay byte-identical (verified across seven shapes before and after).

An unreadable agent file used to throw and abort the entire sweep, while a
failed WRITE in the same loop degraded into a report. The OpenCode branch now
reports read failures alongside write failures; the claude branch degrades to a
skip without a new result field, because its shape is long-standing and widely
consumed and one bad file aborting the sweep is the actual defect.

The tmp+rename publish dropped the original file's mode — a plain writeFileSync
preserves it, a rename does not — so a 0600 agent came back 0644. Both the
OpenCode and the codex branch now carry the original's permission bits across
the publish, masked with 0o7777: the raw stat mode includes the file-type bits,
and POSIX leaves those unspecified for chmod. Linux is the only OS the remote
matrix runs, so relying on Darwin's tolerance would have been untestable here.

Also documents the `from` contract on EffortSyncChange (null means the key was
absent, '' means present with an empty value — a distinction earlier rounds
introduced and then collapsed in the output), adds OpenCode to the docs
paragraph enumerating where the key is omitted under inherit, and records in a
comment that the 'failed' summary reaches only raw mode and does not change the
exit code, which is a CLI-contract change affecting all three branches and is
deliberately not made here.

* fix(#3706): guard the codex read, close the tmp permission window, rename the failure type

Round seven, plus one thing I found myself.

`cmdEffortSyncCodex` still had an unguarded `fs.readFileSync` — a read fault on
one agent exited 1 and aborted the whole sweep. The claude and opencode
branches were both guarded earlier this round and codex was missed, with the
unguarded read sitting ten lines above the chmod block the previous commit did
edit. It now reports read failures the way the OpenCode branch does, and a read
failure flips its summary to `failed` — which write failures did not do there
either, so both are corrected for consistency.

The tmp file was created at the default mode and only tightened afterwards, so
a 0600 agent's contents sat in a 0644 file for the length of the publish. I
measured the window rather than assuming it, then closed it by passing the
mode at creation. The chmod after the write is deliberately RETAINED and
commented: the `mode` option only applies when the file is actually created, so
a leftover tmp from an earlier crashed run would be truncated and reused at its
old mode, and the chmod is what corrects that.

`EffortSyncWriteFailure` is renamed `EffortSyncFileFailure` — it was typing a
`read_failures` array, the same naming-lie the `Codex…` prefix had last round.

Also pins the codex mode preservation with a test. It only writes on a path
that genuinely rewrites the file, so the fixture is an Anthropic-flavoured
model pin the sync strips, and the test asserts the content changed before
checking the mode — otherwise it would pass on a sync that did nothing.

* fix(#3706): guard the claude writes and share one escaping rule

The security sign-off caught a comment of mine that was factually wrong: the
new claude read guard said the failure is folded in "like the write path in
this same loop does", and there was no write guard in that loop. Rather than
correct the sentence, both claude write sites are now guarded the way the read
is — a failed file is skipped, the sweep continues, and the raw summary token
flips to `failed`. The JSON shape stays frozen deliberately, because it is
long-standing and widely consumed; the token is the channel that can carry the
signal without a compatibility risk, which is the reviewer's own suggestion.

That makes all three branches consistent: reads and writes guarded everywhere,
per-file failures degrade instead of aborting, and every branch reports
`failed` rather than `ok` when something did not sync.

`setFrontmatterKeyLine` interpolated its value raw while the install-side
writer quoted through the shared helpers — two writers of the same frontmatter
key disagreeing on escaping, the divergence class this repo requires closed.
They now share one rule. Verified no churn: all six effort levels are plain
scalars and emit byte-identically, with claude's documented minimal-to-low
clamp the only difference in the table, exactly as before.

* fix(#3706): publish claude agent writes atomically too

Both reviewers found this independently, and it is data loss rather than a
reporting gap. The claude branch wrote in place, so `fs.writeFileSync`'s
O_TRUNC meant a post-open fault left the agent file truncated or half-written:
an injected ENOSPC produced an empty file, and under `ulimit -f` a 60000-byte
agent came back as 512 bytes of wrong content. The guard added earlier this
round then counted that destroyed file as `skipped`, which in JSON mode is
indistinguishable from "already in sync" — so a caller would have read the
sweep as clean while an agent on disk was corrupt.

It now publishes the way the codex and opencode branches already do: write to
a tmp file created at the original's masked mode, chmod, then retryRenameSync,
with the tmp unlinked and the agent skipped on any failure. The corrupting case
is gone rather than merely reported, which matters because this branch
deliberately takes no new result key.

I had claimed all three branches were consistent after the previous commit.
That was true for degradation and reporting and not for atomicity; the reviewer
caught the overclaim. It is true now.

Also sorts the claude file list, which the other two branches already did —
readdir order is platform-dependent, so leaving it unsorted made the reported
`changes` ordering differ across machines for identical inputs.

* chore(#3706): backfill the changeset PR number

pr:0 placeholder replaced with the real PR now that gh api returned it.

* test(#3706): kill the frontmatter mutants this change introduced

CI's Stryker frontmatter shard scored 60.58 against a break floor of 62.
The cause is documented in the lane's own config, from #1882: this PR added a
multi-clause predicate to frontmatter.cts and exported the escaper, but the
tests constraining them live in tests/runtime-converters.test.cjs, which that
shard does not run — so every mutant in the new code was uncovered there even
though the behaviour is tested elsewhere.

The fix is assertions that kill real mutants, per the repo's own instruction,
not a lowered floor and not a Stryker disable: scripts/mutation-matrix.cjs is
untouched. Each clause of agentScalarNeedsDoubleQuoting now has a true case AND
a near-miss that must answer the opposite way, so flipping the clause fails a
specific named test — alnum-first against `a-b`, trailing `:` against `foo:bar`,
embedded `: ` against `a:b`, embedded ` #` against `a#b`, the word list against
`yes1`/`nullish`, the numeric forms against `1a`/`0xzz`, the timestamp against
`2026-08-25x`, plus the case-insensitive spellings that pin the `i` flag.
escapeDoubleQuoted is pinned on exact output, including a case constructed so
that escaping in the wrong ORDER yields a different string.

Two of my expectations were wrong and are asserted as the code actually
behaves: `12:99` is NOT quoted, because the sexagesimal alternative never
range-checks minutes and so does not match — which is right, since YAML would
not read it as sexagesimal either; and `20260825` is quoted by the numeric
clause rather than the timestamp one, being a bare integer.

* chore(#3706): ratchet the frontmatter mutation floor to 65

The lane measured 66.67 on PR 3867 after the mutant-killing unit tests landed —
above its pre-change 63.35 baseline, not merely recovered. Step 3 of this
file's own HOW TO UPDATE procedure says to set minScore = floor(measured) - 1
in the same diff, so 62 becomes 65 and the improvement is locked in rather than
left free to slide back.

The ledger of measured scores now records the new measurement, why the shard
broke in the first place (logic added to frontmatter.cts whose only tests lived
in a file this lane does not run — the same trap the #1882 note describes), and
one discrepancy: step 3 also says to update "the matching RATCHET_BASELINE
entry", but no such declaration exists in this file. The name appears only in
that comment, so minScore and the ledger are all there is to update.

* fix(#3706): update RATCHET_BASELINE alongside the raised floor

The ratchet test caught the previous commit: it raised COVERED['frontmatter']
.minScore to 65 without updating the baseline that mirrors it, which is exactly
the mismatch that guard exists to make visible in review.

I had claimed RATCHET_BASELINE did not exist. It does — in
tests/mutation-matrix-ratchet.test.cjs, not in scripts/mutation-matrix.cjs,
which is the only file I searched before concluding it was a stale reference.
The ledger comment is corrected to say where it lives and to record that the
guard caught the error rather than leaving my wrong claim on the record.

* docs(#3706): put the mutation ledger entries back under their own dates

The 2026-08-25 measurement was spliced into the middle of the 2026-06-14 list,
so adr-parser, config-schema, active-workstream-store and core-utils ended up
sitting under the wrong heading and misattributing their measurement dates.
That ledger is what a future change reads to calibrate a floor, so a wrong date
there is not cosmetic. Each measurement is now under the date it was taken.

Also drops the first-person account of my own mistake from the entry — the
factual half (where RATCHET_BASELINE lives, and that it is updated in the same
diff) is what a reader needs; the confession is not.

---------

Co-authored-by: sim <sim@local>
2026-08-25 19:54:30 -04:00

526 lines
23 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env node
'use strict';
/**
* scripts/mutation-matrix.cjs
*
* Single source of truth for the ADR-457 Stryker mutation gate dynamic matrix.
*
* Computes which covered modules changed vs a base ref and emits a GitHub
* Actions matrix JSON so CI can run one Stryker shard per changed module in
* parallel rather than a single serial run over all modules.
*
* Usage:
* node scripts/mutation-matrix.cjs --base origin/next
* printf 'src/config-schema.cts\n' | node scripts/mutation-matrix.cjs
* node scripts/mutation-matrix.cjs --base origin/next --print
*
* Output (stdout, default): JSON object
* {
* "has_work": "true"|"false",
* "matrix": {
* "include": [
* { "name": "<module>", "mutate": "gsd-core/bin/lib/<module>.cjs", "tests": "<space-joined test files>" },
* ...
* ]
* }
* }
*
* Exit codes: 0 always (empty matrix is not an error, has_work "false").
*/
const { execFileSync } = require('child_process');
const fs = require('fs');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
// ── Resilient stdin reader ────────────────────────────────────────────────────
// On macOS, libuv sets the stdin pipe fd to non-blocking mode. A synchronous
// readFileSync(process.stdin.fd) can therefore throw EAGAIN ("resource
// temporarily unavailable") when the writer hasn't yet filled the pipe — this
// is intermittent under heavy CI shard load and causes a spurious status 2
// exit. We work around it by calling fs.readSync in a loop and retrying on
// EAGAIN with a 1 ms synchronous pause (Atomics.wait on a fresh SharedArrayBuffer
// — no hot spin, no real-clock dependency, works under --experimental-vm-modules).
/**
* Read all of stdin synchronously, retrying on EAGAIN.
*
* @returns {string} UTF-8 decoded full stdin content.
*/
function readStdinSync() {
const BUF_SIZE = 64 * 1024; // 64 KB chunks
const buf = Buffer.allocUnsafe(BUF_SIZE);
const chunks = [];
for (;;) {
let bytesRead;
try {
bytesRead = fs.readSync(process.stdin.fd, buf, 0, BUF_SIZE, null);
} catch (err) {
if (err.code === 'EAGAIN') {
// Non-blocking pipe not yet ready — yield for ~1 ms then retry.
Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), 0, 0, 1);
continue;
}
if (err.code === 'EOF') {
break;
}
throw err;
}
if (bytesRead === 0) {
break; // Clean EOF
}
chunks.push(Buffer.from(buf.slice(0, bytesRead)));
}
return Buffer.concat(chunks).toString('utf8');
}
// ── Per-module mutation score ratchet ─────────────────────────────────────────
// ADR-456 / issue #1187: every covered module declares a minScore floor.
//
// HOW THE RATCHET WORKS:
// • minScore locks in the current measured mutation score (minus a 1–2 pt
// margin for run-to-run timeout variance).
// • CI fails a shard if the module's live score drops below its minScore.
// • Raise minScore (never lower) as a module's tests improve.
// • The goal is every module reaching TARGET_MUTATION_SCORE (80).
//
// GOODHART SAFETY: scores are improved by writing genuine behavioural
// assertions that kill real mutants — never by adding brittle exact-string
// matches on incidental output. A justified `// Stryker disable` on a
// confirmed equivalent mutant is acceptable.
//
// HOW TO UPDATE:
// 1. The per-module Stryker shard CANNOT be run locally: Stryker's command
// runner invokes `node --test` once per mutant (see stryker.config.mjs),
// and this repo hard-blocks local `node --test` via
// .claude/hooks/block-local-node-test.sh. Push the branch instead and
// let CI run the shard for the changed module.
// 2. Read the measured score from the CI shard's output.
// 3. Set minScore = floor(measured) - 1 (never lower than current value)
// and update the matching RATCHET_BASELINE entry in the same diff.
// 4. Open/update the PR — the CI gate will enforce the new floor on every
// future run.
/** Long-run target for all modules (ADR-456). */
const TARGET_MUTATION_SCORE = 80;
// ── Single source of truth: covered modules ───────────────────────────────────
// Each entry: { cjs: '<built artifact>', tests: ['tests/...', ...], minScore: N }
//
// minScore is the CI break threshold for this module's shard.
// Floors are measured scores minus 1–2 pts for run-to-run variance.
// Measured CI scores 2026-06-14 (issue #1187, timeout-free — source of truth):
// context-utilization 92.31% → floor 80 (target already met)
// prompt-budget 68.33% → floor 66 (local was 99.6% — TIMEOUT INFLATION; CI is the truth)
// frontmatter 63.35% → floor 62 (SUPERSEDED — see 2026-08-25 below)
// adr-parser 69.30% → floor 68
// config-schema 54.55% → floor 52 (local was 69.7% — TIMEOUT INFLATION; CI is the truth)
// active-workstream-store 81.91% → floor 80
// core-utils 77.52% → floor 75
//
// Measured CI score 2026-08-25 (#3706, PR 3867):
// frontmatter 66.67% → floor 65
// #3706 added agentScalarNeedsDoubleQuoting to frontmatter.cts and exported
// escapeDoubleQuoted, but the tests constraining them lived in
// tests/runtime-converters.test.cjs, which this lane does NOT run — the same trap the
// #1882 note on the frontmatter entry describes. The shard fell to 60.58 and broke the
// floor. Direct unit tests for both were added to tests/frontmatter.unit.test.cjs, each
// clause paired with a near-miss that must answer the opposite way, which took the module
// above its pre-change score. Floor ratcheted per the HOW TO UPDATE formula above, and
// RATCHET_BASELINE — which lives in tests/mutation-matrix-ratchet.test.cjs, not here — is
// updated in the same diff as that procedure requires.
//
// LESSON: floors MUST be calibrated from CI mutation runs (CI runs with
// timeout≈0, deterministic). Local runs count timeouts as kills and
// inflate scores significantly (prompt-budget: 99.6% local vs 68.3% CI;
// config-schema: 69.7% local vs 54.55% CI). Never set a floor from a
// local run without CI cross-check.
const COVERED = {
'context-utilization': {
cjs: 'gsd-core/bin/lib/context-utilization.cjs',
tests: [
'tests/context-utilization.property.test.cjs',
],
// After mutation-killer assertions added in #1187: measured 92.31% (2026-06-14).
// 3 survivors are __esModule boilerplate (genuinely equivalent CJS interop mutants).
// minScore raised to TARGET (80) — module now meets ADR-456 goal.
minScore: 80,
},
// context-composer: extracted from prompt-budget by #2929. Needs its own entry because
// mutation coverage does not migrate with relocated code — scoring only prompt-budget.cjs
// would leave the extracted ladder unmeasured.
'context-composer': {
cjs: 'gsd-core/bin/lib/context-composer.cjs',
tests: [
'tests/prompt-budget-parity.test.cjs',
'tests/prompt-budget.unit.test.cjs',
'tests/context-composer.test.cjs',
'tests/context-composer.property.test.cjs',
],
minScore: 66,
},
'prompt-budget': {
cjs: 'gsd-core/bin/lib/prompt-budget.cjs',
tests: [
'tests/prompt-budget.property.test.cjs',
'tests/prompt-budget.unit.test.cjs',
],
// CI 68.33% timeout-free (164 killed / 1 timeout / 240 total) 2026-06-14;
// local was 99.6% — timeout inflation. Floor = 68 - 2 margin.
minScore: 66,
},
frontmatter: {
cjs: 'gsd-core/bin/lib/frontmatter.cjs',
tests: [
'tests/frontmatter.property.test.cjs',
'tests/frontmatter.unit.test.cjs',
// #1882 added the unterminated-fence detection to frontmatter.cjs, and the tests that
// constrain it live here. Without this entry the mutants in that branch are covered by
// no test in the shard, so the module's score drops even though the behaviour is tested.
'tests/unusable-input.test.cjs',
],
minScore: 65,
},
'adr-parser': {
cjs: 'gsd-core/bin/lib/adr-parser.cjs',
tests: [
'tests/adr-parser.property.test.cjs',
'tests/adr-parser.test.cjs',
'tests/adr-parser.unit.test.cjs',
],
minScore: 68,
},
'config-schema': {
cjs: 'gsd-core/bin/lib/config-schema.cjs',
tests: [
'tests/config-schema.property.test.cjs',
],
// CI 54.55% timeout-free (18 killed / 0 timeout / 33 total) 2026-06-14;
// local was 69.7% — timeout inflation. Floor = 54 - 2 margin.
minScore: 52,
},
'active-workstream-store': {
cjs: 'gsd-core/bin/lib/active-workstream-store.cjs',
tests: [
'tests/active-workstream-store.test.cjs',
'tests/active-workstream-store.unit.test.cjs',
],
minScore: 80,
},
'core-utils': {
cjs: 'gsd-core/bin/lib/core-utils.cjs',
tests: [
'tests/core-utils.test.cjs',
],
minScore: 75, // measured 77.52% (2026-06-14, issue #1187); floor = 77 - 2
},
// planning-inspect / plan-document / planning-command-router: net-new modules
// added by #2790. Registered here so the Stryker gate stops SKIPPING them
// (previously has_work: "false" — ~1000 LOC entirely outside mutation scoring).
//
// WHY THESE SHARDS POINT AT tests/planning-inspect.unit.test.cjs, NOT
// tests/planning-inspect.test.cjs. CI evidence: two shards pointed at the
// integration file were CANCELLED at the workflow's 15-minute cap —
// "Mutation testing 4% (elapsed: ~3m, remaining: ~1h 19m) 27/640 tested".
// tests/planning-inspect.test.cjs is INTEGRATION-shaped (91 cases, most
// spawning a `gsd-tools` child process via `runGsdTools`); Stryker's command
// runner treats the whole `node --test <file>` invocation as ONE test costing
// whatever the slowest case costs (measured ~20s), and re-runs that entire
// file once per mutant — 640 mutants x 20s cannot finish in 15 minutes.
// tests/planning-inspect.unit.test.cjs is the dedicated, spawn-free,
// in-process mutation surface for exactly these three modules (measured
// locally: the whole file runs in well under a second) — the same shape
// every other entry in this registry already uses (*.property.test.cjs /
// *.unit.test.cjs). The integration suite is UNAFFECTED by this change: it
// keeps running in full in the normal (non-mutation) test job, and remains
// the source of truth for spawn-boundary/CLI-dispatch/read-only-proof
// behaviour that an in-process unit file cannot exercise.
//
// Measured CI scores (GitHub Actions run 32392791843, all three shards
// PASSED — not a local run; mutation shards run `node --test`, hard-blocked
// in this repo's local environment):
// planning-command-router 95.65% → floor 94 (already exceeds TARGET_MUTATION_SCORE (80))
// plan-document 76.58% → floor 75
// planning-inspect 57.03% → floor 56 (well below TARGET (80) — ratchet
// candidate; comfortably clears its own floor but has real room to grow.
// Raise as its tests improve, never lower it.)
//
// All three shards point at tests/planning-inspect.unit.test.cjs (in-process,
// spawn-free, ~0.3s dry run), not tests/planning-inspect.test.cjs — that is
// what made measurement possible at all. The integration file spawns a
// subprocess per case via runGsdTools; Stryker's command runner treats the
// whole `node --test <file>` invocation as one test costing whatever the
// slowest case costs (measured ~20s), and re-runs that entire file once per
// mutant, so 640 mutants x 20s could not finish inside the 15-minute shard
// cap. The integration suite is unaffected by this change: it keeps running
// in full in the normal (non-mutation) test job.
'planning-inspect': {
cjs: 'gsd-core/bin/lib/planning-inspect.cjs',
tests: [
'tests/planning-inspect.unit.test.cjs',
],
minScore: 56,
},
'plan-document': {
cjs: 'gsd-core/bin/lib/plan-document.cjs',
tests: [
'tests/planning-inspect.unit.test.cjs',
],
minScore: 75,
},
'planning-command-router': {
cjs: 'gsd-core/bin/lib/planning-command-router.cjs',
tests: [
'tests/planning-inspect.unit.test.cjs',
],
minScore: 94,
},
// model-catalog: net-new registration by #3007. The module was entirely
// outside mutation scoring (has_work: "false") before this entry, so the
// #3007 per-model Codex effort rewrite (renderEffortForRuntime's
// CODEX_MODEL_EFFORT lookup, the 'ultra' policy rejection, the ladder
// walk-up clamp) had zero mutation coverage.
//
// Same #2790 precedent as planning-inspect above: this shard points at a
// dedicated tests/model-catalog.unit.test.cjs, NOT tests/model-resolver.test.cjs
// — that integration file uses runGsdTools heavily and would hit the same
// 15-minute shard-cap cancellation #2790 documented (a `node --test <file>`
// invocation is ONE test costing whatever its slowest case costs, re-run
// per mutant). tests/model-catalog.unit.test.cjs is spawn-free, in-process,
// and runs in well under a second.
//
// Measured CI score (GitHub Actions run 32605073352, job 97108869486):
// model-catalog 59.62% → floor 58 (248 killed, 168 survived, 0 timeouts,
// 0 errors; below TARGET_MUTATION_SCORE (80) — ratchet candidate like
// planning-inspect (56): comfortably clears its own floor but has real
// room to grow. Raise as its tests improve, never lower it.)
// Floor follows this file's documented rule, minScore = floor(measured) - 1,
// matching the sibling precedent exactly (57.03 → 56, 76.58 → 75, 95.65 → 94).
//
// The shard completed in 57 seconds — concrete evidence the spawn-free
// unit-file design above worked: the #2790 precedent's 15-minute shard-cap
// cancellations do not apply here, and for comparison the `frontmatter`
// shard in the same run took 9m46s.
'model-catalog': {
cjs: 'gsd-core/bin/lib/model-catalog.cjs',
tests: ['tests/model-catalog.unit.test.cjs'],
minScore: 58,
},
// state-contract: net-new module from #3227. Without this entry the
// Stryker gate reports has_work: "false" and SKIPS it entirely — the
// exact gap #2790 (planning-inspect / plan-document / planning-command-router)
// and #3007 (model-catalog) each had to fix after the fact.
//
// Same #2790 precedent as planning-inspect / model-catalog above: this
// shard points at tests/state-contract.unit.test.cjs, NOT
// tests/state-contract.test.cjs — the latter spawns a `gsd-tools` child
// process per case via runGsdTools, and Stryker's command runner treats
// the whole `node --test <file>` invocation as ONE test costing whatever
// its slowest case costs, re-run once per mutant, so it cannot finish
// inside the 15-minute shard cap. tests/state-contract.unit.test.cjs is
// spawn-free and in-process.
//
// Measured CI score (GitHub Actions run 32769289750, job 97565813640,
// `Stryker (state-contract)`, PASSED in 2m23s):
// state-contract 66.25% → floor 65 (below TARGET_MUTATION_SCORE (80) —
// ratchet candidate like planning-inspect (56) and model-catalog (58):
// comfortably clears its own floor but has real room to grow. Raise as
// its tests improve, never lower it.)
// Floor follows this file's documented rule, minScore = floor(measured) - 1,
// matching the sibling precedent exactly (57.03 → 56, 76.58 → 75,
// 95.65 → 94, 59.62 → 58, 66.25 → 65).
//
// The floor MUST come from a CI shard, never a local run: local runs count
// timeouts as kills and inflate scores badly (this file already records
// prompt-budget 99.6% local vs 68.33% CI, and config-schema 69.7% local vs
// 54.55% CI).
'state-contract': {
cjs: 'gsd-core/bin/lib/state-contract.cjs',
tests: ['tests/state-contract.unit.test.cjs'],
minScore: 65,
},
};
// ── Files that, when changed, invalidate ALL modules ─────────────────────────
// Changes to the Stryker config, this script itself, or any covered test file
// affect all mutation scores and must force a full re-run.
const GLOBAL_TRIGGERS = new Set([
'stryker.config.mjs',
'scripts/mutation-matrix.cjs',
]);
// Also flag all test files that belong to any covered module as global triggers.
for (const mod of Object.values(COVERED)) {
for (const t of mod.tests) {
GLOBAL_TRIGGERS.add(t);
}
}
// ── Argument parsing ──────────────────────────────────────────────────────────
function parseArgs(argv) {
const out = { base: null, print: false };
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
if (arg === '--base') {
out.base = argv[++i];
if (!out.base || out.base.startsWith('--')) {
throw new Error('--base requires a value');
}
} else if (arg.startsWith('--base=')) {
out.base = arg.slice('--base='.length);
if (!out.base) throw new Error('--base requires a value');
} else if (arg === '--print') {
out.print = true;
} else if (arg === '--help' || arg === '-h') {
console.log([
'Usage:',
' node scripts/mutation-matrix.cjs --base <ref> [--print]',
' printf "src/foo.cts\\n" | node scripts/mutation-matrix.cjs [--print]',
'',
'Options:',
' --base <ref> Git ref to diff against (default: origin/${GITHUB_BASE_REF:-next})',
' --print Human-readable output instead of JSON',
].join('\n'));
throw new ExitError(0);
} else {
throw new Error(`unknown argument: ${arg}`);
}
}
return out;
}
// ── Changed-file resolution ───────────────────────────────────────────────────
function resolveChangedFiles(args) {
// When --base is provided, always use git diff (regardless of stdin).
// When --base is absent AND stdin is not a TTY (isTTY is falsy / undefined),
// read a newline-delimited file list from stdin.
if (!args.base && process.stdin.isTTY !== true) {
const raw = readStdinSync();
return raw.split('\n').map(l => l.trim()).filter(Boolean);
}
// Otherwise (--base given, or stdin is a real TTY), diff against the base ref.
const defaultBase = `origin/${process.env.GITHUB_BASE_REF || 'next'}`;
const base = args.base || defaultBase;
const stdout = execFileSync('git', ['diff', '--name-only', `${base}...HEAD`], {
encoding: 'utf8',
});
return stdout.split('\n').map(l => l.trim()).filter(Boolean);
}
// ── Module classification ─────────────────────────────────────────────────────
function computeMatrix(changedFiles) {
// Check for global triggers first — if any hit, include every covered module.
const allModuleNames = Object.keys(COVERED);
for (const f of changedFiles) {
if (GLOBAL_TRIGGERS.has(f)) {
return allModuleNames;
}
}
// Otherwise find which modules have their src/*.cts changed.
const changed = new Set();
for (const f of changedFiles) {
// Match src/<module>.cts (top-level src/, not nested)
const m = f.match(/^src\/([^/]+)\.cts$/);
if (m && COVERED[m[1]]) {
changed.add(m[1]);
}
}
return [...changed];
}
// ── Output formatting ─────────────────────────────────────────────────────────
function buildResult(moduleNames) {
const include = moduleNames.map(name => ({
name,
mutate: COVERED[name].cjs,
tests: COVERED[name].tests.join(' '),
minScore: COVERED[name].minScore,
}));
return {
has_work: include.length > 0 ? 'true' : 'false',
matrix: { include },
};
}
function printHuman(result, changedFiles) {
console.log(`Changed files (${changedFiles.length}):`);
for (const f of changedFiles) console.log(` ${f}`);
console.log('');
console.log(`has_work: ${result.has_work}`);
console.log(`Shards (${result.matrix.include.length}):`);
for (const shard of result.matrix.include) {
console.log(` [${shard.name}]`);
console.log(` mutate: ${shard.mutate}`);
console.log(` tests: ${shard.tests}`);
console.log(` minScore: ${shard.minScore}`);
}
}
// ── Main ──────────────────────────────────────────────────────────────────────
function main() {
try {
const args = parseArgs(process.argv.slice(2));
const changedFiles = resolveChangedFiles(args);
const moduleNames = computeMatrix(changedFiles);
const result = buildResult(moduleNames);
if (args.print) {
printHuman(result, changedFiles);
} else {
console.log(JSON.stringify(result, null, 2));
}
} catch (err) {
if (err instanceof ExitError) throw err;
console.error(`mutation-matrix: ${err.message}`);
throw new ExitError(2);
}
}
// ── MUTATION_BREAK resolver ───────────────────────────────────────────────────
/**
* Resolves the per-shard mutation break threshold from the MUTATION_BREAK env var.
*
* Fail-closed contract:
* - undefined → 60 (local run: no env set, documented backstop)
* - set but empty (e.g. CI matrix.minScore missing) → throws (wiring error)
* - non-numeric or out-of-range [1, 100] → throws (invalid config)
* - valid integer string → returns that number
*
* This function is the single call site for reading MUTATION_BREAK.
* stryker.config.mjs imports and calls it so CI shards with a bad
* MUTATION_BREAK fail immediately rather than silently falling back to 60
* and bypassing a per-module floor above 60 (e.g. prompt-budget: 90).
*
* @param {string|undefined} raw - value of process.env.MUTATION_BREAK
* @returns {number}
*/
function resolveMutationBreak(raw) {
if (raw === undefined) {
// Local run with no MUTATION_BREAK set — use documented backstop.
return 60;
}
if (typeof raw !== 'string' || raw.trim() === '') {
throw new Error(
'MUTATION_BREAK is set but empty — CI shard wiring is broken (matrix.minScore missing?)'
);
}
const n = Number(raw);
if (!Number.isFinite(n) || n < 1 || n > 100) {
throw new Error(
`MUTATION_BREAK invalid: "${raw}" (expected a per-module minScore 1-100)`
);
}
return n;
}
// Export internals for programmatic use (tests/mutation-matrix-ratchet.test.cjs).
// The require.main guard prevents main() from running when this file is require()d.
module.exports = { COVERED, TARGET_MUTATION_SCORE, resolveMutationBreak, readStdinSync };
if (require.main === module) runMain(main);