Files
msd-core/scripts/mutation-matrix.cjs
Tom Boucher 9ac0dfad58 chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor

Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared
context-composer seam. Its success condition is that review-prompt output does
not change, and the only authority on "did not change" is the behavior that
shipped before the refactor. Capture that behavior now, while it is still the
live implementation.

47 characterization cases, every `expected` value computed by executing the
current implementation rather than hand-authored — the independence
CONTRIBUTING.md "Fixture provenance (#2371)" asks for.

A corpus is only worth what it can detect, so this one was validated by
mutation rather than assumed. Five deliberate defects were injected and each
must be caught by at least one case:

  - the note reserve deducted unconditionally instead of only under pressure
  - the pressure test relaxed from `>` to `>=`
  - a no-op head-shrink still setting the shrunk flag
  - the per-plan floor dropped from the proportional share
  - drop order reversed

Two of those exposed real holes in the first cut of this corpus, and the cases
that close them exist because of it:

  - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a
    floored plan group, and the 1024-char floor absorbed the entire trim, so the
    mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap,
    which makes the strict inequality observable as context kept vs omitted.

  - No case reached proportional-truncate at all — B6 and B7 both hard-failed
    the min-set pre-check first, leaving planTruncationPct at 0 across every
    case and the floor semantics entirely unexercised. Rebudgeted to 700 and
    1100 so the min-set fits and the truncate step is actually reached; they now
    record 40.20% and 48.80%.

The A4/A10 families sweep the pressure boundary from both sides, which is where
this function has regressed before: CONTEXT.md's
LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions
that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band,
because the suite paired a trivially-fitting budget with a trivially-overflowing
one and never sampled between them. A4 pins that nothing is trimmed from the cap
down to 81 tokens under it; A10 pins that pressure fires at +1. Together with
A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures.

Two facts the corpus establishes that the design notes had wrong:

  - "" and null sections are NOT distinguished. applyBudget uses truthy checks
    throughout, so an empty-string section is treated as absent: not rendered,
    not dropped, never recorded in `omitted`. B13b pins this while the ladder is
    actively trimming, where only the non-empty `research` is dropped.

  - Sizing matters. B12/B13 were first written at a budget where both hard-failed
    the min-set check and returned "", so comparing them compared two empty
    strings and proved nothing.

Committed as its own commit, ahead of the refactor, and regenerated against the
pre-refactor implementation, so the oracle is demonstrably independent of the
change it will adjudicate.

Refs #2929

* refactor(#2929): extract the context-composer seam from prompt-budget

Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer —
per-runtime artifact emission — but it is walled inside the cross-AI review
pipeline. Lift it into a shared seam so later phases can call it, without
changing what the review pipeline emits.

ADR-1671 specifies the composer as "priority + binary-search cutoff to a
per-runtime budget". Read against the code it generalizes, that contract cannot
express the thing being generalized. applyBudget is not a cutoff: it is a fixed
five-step ladder in which each section carries its own shrink strategy, and only
three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N
lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor;
instructions and roadmap are never touched at all. A cutoff composer sorts by
priority and discards the tail — it has no way to say "shrink this one",
"truncate that one but never below 1 KB each", or "these three are the only
droppables, in this order". Building to the literal contract and routing
prompt-budget through it would have silently changed review-prompt output, which
is the one outcome this phase forbids.

So shrink strategies are the core abstraction here, and cutoff becomes one
strategy among them — the right one for per-runtime emission in Phases 3-4, not
for this ladder. That is an elaboration of the ADR's intent, not a departure
from it, and ADR-1671 is updated to say so.

Three decisions worth stating:

  - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan
    of surviving fragments and never a string. assemblePrompt's rendering is
    prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and
    owning it in the composer would force emission to adopt prompt-shaped
    rendering. The split is what lets one seam serve both consumers.

  - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its
    chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires
    for emission caps. The existing code converts a token budget to a character
    budget with a hardcoded `* 4`; that assumption is now an explicit
    `charsPerUnit` inverse, which is precisely what a byte unit needs in order to
    reuse this.

  - The entry point is `composeWithinBudget`, not `applyBudget`. That name
    already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter
    being an unrelated graph-edge budget. A third would make every symbol search
    in this repo ambiguous, and it already misresolves: preflight and impact
    queries for "applyBudget" return graphify's.

Behavior is unchanged and proven so: all 47 characterization cases reproduce
byte-identically, and the corpus is mutation-validated rather than merely green
(see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and
from eighteen mutable accumulators to two, both inside a helper copied verbatim.

estimateTokens deliberately stays in prompt-budget and keeps its exact math:
src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins
plan estimates and recorded actuals to that same scale, so moving or changing it
would silently break the calibration loop.

Refs #2929

* docs(#2929): document the context-composer seam and amend ADR-1671

Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new
domain modules), and a mutation-matrix entry for the new module.

The ADR amendment is the substantive part. ADR-1671 specified the composer as
"priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2
established that a cutoff alone cannot express the function the platform
generalizes, so the ADR now records shrink strategies as the core abstraction
with cutoff as one strategy among them, reserved for per-runtime emission in
Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned
against that contract and would otherwise be planned against a mechanism that
does not work.

The mutation-matrix entry is not bookkeeping. Stryker scores per module against
a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the
extracted code unmeasured while prompt-budget's own score floated free of the
logic it used to cover. context-composer gets its own entry at the same floor.

Refs #2929

* test(#2929): pin the effectiveBudget rounding mode in the parity corpus

An isolated correctness review found a real blind spot: mutating
`Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of
the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator
happened to produce a whole number, so floor, round and ceil all agreed and the
rounding mode was entirely unpinned by a corpus whose whole job is to pin
observable behavior.

Three cases fix that by straddling the .5 boundary:

  A11  95 * 0.90  = 85.5   floor 85, round 86  -> the two disagree
  A12  97 * 0.90  = 87.3   floor and round agree; ceil (88) does not
  A13  93 * 0.85  = 79.05  same guard at a non-multiple-of-10 margin, so the
                           margin arithmetic is exercised and not just the budget

A11 alone catches the round mutation; all three catch ceil. Regenerated against
the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so
the expanded corpus keeps the independence property the original capture had.

The corpus is now mutation-validated against seven injected defects, every one
caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink
setting its flag, the truncate floor ignored, drop order reversed, and both
rounding-mode changes.

Refs #2929

* feat(#2929): flexReserve floors and the byte-stable isolate prefix

Two of issue #2929's "Done when" items were unimplemented rather than deferred,
and an isolated review flagged them alongside my own audit. Both are part of
ADR-1671's composer contract, so shipping the seam without them would have left
Phases 3-4 building against a contract that does not exist yet.

flexReserve is a per-fragment floor in measure units that every strategy must
respect, which is what makes it different from the pre-existing floorChars: that
one is a chars-denominated detail of proportional-truncate alone and is retained
unchanged. A floored fragment is never dropped, is never head-shrunk below its
floor, and raises its own proportional cap. A fragment already smaller than its
floor is untouchable outright. Metadata gains `floored`, listing the ids whose
floor actually prevented a trim — a guarantee no caller can observe is a
guarantee no test can hold you to.

isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed,
never dropped, but still counted, because a prefix excluded from accounting
would silently under-count real context. Metadata gains `isolatePrefix` so a
caller can hash or assert on the exact bytes. Declaring an isolate fragment
after a non-isolate one throws: a prefix that is not at the front is not a
prefix, and accepting it would make the cross-runtime stability claim
meaningless.

Adds tests/context-composer.test.cjs for the exact new semantics and
tests/context-composer.property.test.cjs for the five invariants, including the
budget-monotonicity property the issue names explicitly. Both are registered in
the mutation matrix, since coverage does not migrate with relocated code.

prompt-budget uses neither feature, and its output is unchanged: all 50 corpus
cases still reproduce byte-identically.

Refs #2929

* chore(#2929): allowlist the prompt-budget parity suite

The parity corpus needs its own test file and that makes prompt-budget a
three-file module against a limit of two. The lint offers consolidation or an
allowlist entry with justification; the entry is the right call here.

Consolidation would mean folding the characterization suite into
prompt-budget.test.cjs, which is the one thing that should not happen to it. The
parity suite is a distinct concern with a distinct lifecycle: it is generated
rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own
scoring target, and its failure means something categorically different from a
unit-test failure — not "this behavior is wrong" but "observable output moved".
Burying it inside a general unit file would obscure exactly that signal.

The allowlist is an identity ratchet, so this entry pins today's three exact
filenames: adding a fourth still fails, and dropping back to two requires
removing the entry.

Refs #2929

* fix(#2929): register the new module with two gates it was missing

The remote matrix caught three defects that no local check could, because the
local runner is blocked in this repo and these suites had therefore never
executed. Eight failures, identical on node22 and node24, so nothing
environment-shaped.

Two are the new-module ripple. A net-new src/*.cts lands in six places and this
change had reached four of them — .gitignore, INVENTORY, the manifest, and the
CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not
be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet
baseline (a deliberate review-visible mirror of the matrix floors, which every
COVERED module must carry). Both are now registered, the ratchet at the same
floor of 66 the matrix declares.

The third was a test asserting an outcome it had made impossible. It set
budget:1 alongside a 400-char required fragment, so the group budget came out at
-99 and the proportional-truncate step was skipped entirely — the deliberate
"non-positive group budget is skipped, never clamped" rule inherited from the
original ladder. Nothing was trimmed, and the test then asserted a truncation.
Rebudgeted so the step actually runs, with the arithmetic written out in a
comment so the next reader does not have to re-derive why 120 rather than 80.

Fixing that surfaced a genuine bug in the composer. `floored` is documented as
recording fragments whose flexReserve prevented a trim that would otherwise have
happened, but the push sat in the else-branch of "content did not change", so it
only fired when nothing was trimmed at all. A fragment truncated to a
reserve-raised cap has also had a trim prevented — 40 characters' worth in the
test above — and was silently absent from the field that exists to make the
guarantee observable. The condition was already right; it was in the wrong
branch. Now recorded on both paths: a drop prevented outright, and a truncation
capped higher than the share alone would have allowed.

Parity is unaffected — prompt-budget never sets flexReserve, so the branch is
unreachable from every corpus path, and all 50 cases still match.

Refs #2929

* chore(#2929): backfill changeset PR number (#2958)

* chore(#2929): correct the corpus case count in the changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-07-31 23:03:13 -04:00

383 lines
15 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env node
'use strict';
/**
* scripts/mutation-matrix.cjs
*
* Single source of truth for the ADR-457 Stryker mutation gate dynamic matrix.
*
* Computes which covered modules changed vs a base ref and emits a GitHub
* Actions matrix JSON so CI can run one Stryker shard per changed module in
* parallel rather than a single serial run over all modules.
*
* Usage:
* node scripts/mutation-matrix.cjs --base origin/next
* printf 'src/config-schema.cts\n' | node scripts/mutation-matrix.cjs
* node scripts/mutation-matrix.cjs --base origin/next --print
*
* Output (stdout, default): JSON object
* {
* "has_work": "true"|"false",
* "matrix": {
* "include": [
* { "name": "<module>", "mutate": "gsd-core/bin/lib/<module>.cjs", "tests": "<space-joined test files>" },
* ...
* ]
* }
* }
*
* Exit codes: 0 always (empty matrix is not an error, has_work "false").
*/
const { execFileSync } = require('child_process');
const fs = require('fs');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
// ── Resilient stdin reader ────────────────────────────────────────────────────
// On macOS, libuv sets the stdin pipe fd to non-blocking mode. A synchronous
// readFileSync(process.stdin.fd) can therefore throw EAGAIN ("resource
// temporarily unavailable") when the writer hasn't yet filled the pipe — this
// is intermittent under heavy CI shard load and causes a spurious status 2
// exit. We work around it by calling fs.readSync in a loop and retrying on
// EAGAIN with a 1 ms synchronous pause (Atomics.wait on a fresh SharedArrayBuffer
// — no hot spin, no real-clock dependency, works under --experimental-vm-modules).
/**
* Read all of stdin synchronously, retrying on EAGAIN.
*
* @returns {string} UTF-8 decoded full stdin content.
*/
function readStdinSync() {
const BUF_SIZE = 64 * 1024; // 64 KB chunks
const buf = Buffer.allocUnsafe(BUF_SIZE);
const chunks = [];
for (;;) {
let bytesRead;
try {
bytesRead = fs.readSync(process.stdin.fd, buf, 0, BUF_SIZE, null);
} catch (err) {
if (err.code === 'EAGAIN') {
// Non-blocking pipe not yet ready — yield for ~1 ms then retry.
Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), 0, 0, 1);
continue;
}
if (err.code === 'EOF') {
break;
}
throw err;
}
if (bytesRead === 0) {
break; // Clean EOF
}
chunks.push(Buffer.from(buf.slice(0, bytesRead)));
}
return Buffer.concat(chunks).toString('utf8');
}
// ── Per-module mutation score ratchet ─────────────────────────────────────────
// ADR-456 / issue #1187: every covered module declares a minScore floor.
//
// HOW THE RATCHET WORKS:
// • minScore locks in the current measured mutation score (minus a 1–2 pt
// margin for run-to-run timeout variance).
// • CI fails a shard if the module's live score drops below its minScore.
// • Raise minScore (never lower) as a module's tests improve.
// • The goal is every module reaching TARGET_MUTATION_SCORE (80).
//
// GOODHART SAFETY: scores are improved by writing genuine behavioural
// assertions that kill real mutants — never by adding brittle exact-string
// matches on incidental output. A justified `// Stryker disable` on a
// confirmed equivalent mutant is acceptable.
//
// HOW TO UPDATE:
// 1. Run the per-module Stryker shard locally.
// 2. Note the reported score.
// 3. Set minScore = floor(score) - 1 (never lower than current value).
// 4. Open a PR — the CI gate will enforce the new floor on every future run.
/** Long-run target for all modules (ADR-456). */
const TARGET_MUTATION_SCORE = 80;
// ── Single source of truth: covered modules ───────────────────────────────────
// Each entry: { cjs: '<built artifact>', tests: ['tests/...', ...], minScore: N }
//
// minScore is the CI break threshold for this module's shard.
// Floors are measured scores minus 1–2 pts for run-to-run variance.
// Measured CI scores 2026-06-14 (issue #1187, timeout-free — source of truth):
// context-utilization 92.31% → floor 80 (target already met)
// prompt-budget 68.33% → floor 66 (local was 99.6% — TIMEOUT INFLATION; CI is the truth)
// frontmatter 63.35% → floor 62
// adr-parser 69.30% → floor 68
// config-schema 54.55% → floor 52 (local was 69.7% — TIMEOUT INFLATION; CI is the truth)
// active-workstream-store 81.91% → floor 80
// core-utils 77.52% → floor 75
//
// LESSON: floors MUST be calibrated from CI mutation runs (CI runs with
// timeout≈0, deterministic). Local runs count timeouts as kills and
// inflate scores significantly (prompt-budget: 99.6% local vs 68.3% CI;
// config-schema: 69.7% local vs 54.55% CI). Never set a floor from a
// local run without CI cross-check.
const COVERED = {
'context-utilization': {
cjs: 'gsd-core/bin/lib/context-utilization.cjs',
tests: [
'tests/context-utilization.property.test.cjs',
],
// After mutation-killer assertions added in #1187: measured 92.31% (2026-06-14).
// 3 survivors are __esModule boilerplate (genuinely equivalent CJS interop mutants).
// minScore raised to TARGET (80) — module now meets ADR-456 goal.
minScore: 80,
},
// context-composer: extracted from prompt-budget by #2929. Needs its own entry because
// mutation coverage does not migrate with relocated code — scoring only prompt-budget.cjs
// would leave the extracted ladder unmeasured.
'context-composer': {
cjs: 'gsd-core/bin/lib/context-composer.cjs',
tests: [
'tests/prompt-budget-parity.test.cjs',
'tests/prompt-budget.unit.test.cjs',
'tests/context-composer.test.cjs',
'tests/context-composer.property.test.cjs',
],
minScore: 66,
},
'prompt-budget': {
cjs: 'gsd-core/bin/lib/prompt-budget.cjs',
tests: [
'tests/prompt-budget.property.test.cjs',
'tests/prompt-budget.unit.test.cjs',
],
// CI 68.33% timeout-free (164 killed / 1 timeout / 240 total) 2026-06-14;
// local was 99.6% — timeout inflation. Floor = 68 - 2 margin.
minScore: 66,
},
frontmatter: {
cjs: 'gsd-core/bin/lib/frontmatter.cjs',
tests: [
'tests/frontmatter.property.test.cjs',
'tests/frontmatter.unit.test.cjs',
// #1882 added the unterminated-fence detection to frontmatter.cjs, and the tests that
// constrain it live here. Without this entry the mutants in that branch are covered by
// no test in the shard, so the module's score drops even though the behaviour is tested.
'tests/unusable-input.test.cjs',
],
minScore: 62,
},
'adr-parser': {
cjs: 'gsd-core/bin/lib/adr-parser.cjs',
tests: [
'tests/adr-parser.property.test.cjs',
'tests/adr-parser.test.cjs',
'tests/adr-parser.unit.test.cjs',
],
minScore: 68,
},
'config-schema': {
cjs: 'gsd-core/bin/lib/config-schema.cjs',
tests: [
'tests/config-schema.property.test.cjs',
],
// CI 54.55% timeout-free (18 killed / 0 timeout / 33 total) 2026-06-14;
// local was 69.7% — timeout inflation. Floor = 54 - 2 margin.
minScore: 52,
},
'active-workstream-store': {
cjs: 'gsd-core/bin/lib/active-workstream-store.cjs',
tests: [
'tests/active-workstream-store.test.cjs',
'tests/active-workstream-store.unit.test.cjs',
],
minScore: 80,
},
'core-utils': {
cjs: 'gsd-core/bin/lib/core-utils.cjs',
tests: [
'tests/core-utils.test.cjs',
],
minScore: 75, // measured 77.52% (2026-06-14, issue #1187); floor = 77 - 2
},
};
// ── Files that, when changed, invalidate ALL modules ─────────────────────────
// Changes to the Stryker config, this script itself, or any covered test file
// affect all mutation scores and must force a full re-run.
const GLOBAL_TRIGGERS = new Set([
'stryker.config.mjs',
'scripts/mutation-matrix.cjs',
]);
// Also flag all test files that belong to any covered module as global triggers.
for (const mod of Object.values(COVERED)) {
for (const t of mod.tests) {
GLOBAL_TRIGGERS.add(t);
}
}
// ── Argument parsing ──────────────────────────────────────────────────────────
function parseArgs(argv) {
const out = { base: null, print: false };
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
if (arg === '--base') {
out.base = argv[++i];
if (!out.base || out.base.startsWith('--')) {
throw new Error('--base requires a value');
}
} else if (arg.startsWith('--base=')) {
out.base = arg.slice('--base='.length);
if (!out.base) throw new Error('--base requires a value');
} else if (arg === '--print') {
out.print = true;
} else if (arg === '--help' || arg === '-h') {
console.log([
'Usage:',
' node scripts/mutation-matrix.cjs --base <ref> [--print]',
' printf "src/foo.cts\\n" | node scripts/mutation-matrix.cjs [--print]',
'',
'Options:',
' --base <ref> Git ref to diff against (default: origin/${GITHUB_BASE_REF:-next})',
' --print Human-readable output instead of JSON',
].join('\n'));
throw new ExitError(0);
} else {
throw new Error(`unknown argument: ${arg}`);
}
}
return out;
}
// ── Changed-file resolution ───────────────────────────────────────────────────
function resolveChangedFiles(args) {
// When --base is provided, always use git diff (regardless of stdin).
// When --base is absent AND stdin is not a TTY (isTTY is falsy / undefined),
// read a newline-delimited file list from stdin.
if (!args.base && process.stdin.isTTY !== true) {
const raw = readStdinSync();
return raw.split('\n').map(l => l.trim()).filter(Boolean);
}
// Otherwise (--base given, or stdin is a real TTY), diff against the base ref.
const defaultBase = `origin/${process.env.GITHUB_BASE_REF || 'next'}`;
const base = args.base || defaultBase;
const stdout = execFileSync('git', ['diff', '--name-only', `${base}...HEAD`], {
encoding: 'utf8',
});
return stdout.split('\n').map(l => l.trim()).filter(Boolean);
}
// ── Module classification ─────────────────────────────────────────────────────
function computeMatrix(changedFiles) {
// Check for global triggers first — if any hit, include every covered module.
const allModuleNames = Object.keys(COVERED);
for (const f of changedFiles) {
if (GLOBAL_TRIGGERS.has(f)) {
return allModuleNames;
}
}
// Otherwise find which modules have their src/*.cts changed.
const changed = new Set();
for (const f of changedFiles) {
// Match src/<module>.cts (top-level src/, not nested)
const m = f.match(/^src\/([^/]+)\.cts$/);
if (m && COVERED[m[1]]) {
changed.add(m[1]);
}
}
return [...changed];
}
// ── Output formatting ─────────────────────────────────────────────────────────
function buildResult(moduleNames) {
const include = moduleNames.map(name => ({
name,
mutate: COVERED[name].cjs,
tests: COVERED[name].tests.join(' '),
minScore: COVERED[name].minScore,
}));
return {
has_work: include.length > 0 ? 'true' : 'false',
matrix: { include },
};
}
function printHuman(result, changedFiles) {
console.log(`Changed files (${changedFiles.length}):`);
for (const f of changedFiles) console.log(` ${f}`);
console.log('');
console.log(`has_work: ${result.has_work}`);
console.log(`Shards (${result.matrix.include.length}):`);
for (const shard of result.matrix.include) {
console.log(` [${shard.name}]`);
console.log(` mutate: ${shard.mutate}`);
console.log(` tests: ${shard.tests}`);
console.log(` minScore: ${shard.minScore}`);
}
}
// ── Main ──────────────────────────────────────────────────────────────────────
function main() {
try {
const args = parseArgs(process.argv.slice(2));
const changedFiles = resolveChangedFiles(args);
const moduleNames = computeMatrix(changedFiles);
const result = buildResult(moduleNames);
if (args.print) {
printHuman(result, changedFiles);
} else {
console.log(JSON.stringify(result, null, 2));
}
} catch (err) {
if (err instanceof ExitError) throw err;
console.error(`mutation-matrix: ${err.message}`);
throw new ExitError(2);
}
}
// ── MUTATION_BREAK resolver ───────────────────────────────────────────────────
/**
* Resolves the per-shard mutation break threshold from the MUTATION_BREAK env var.
*
* Fail-closed contract:
* - undefined → 60 (local run: no env set, documented backstop)
* - set but empty (e.g. CI matrix.minScore missing) → throws (wiring error)
* - non-numeric or out-of-range [1, 100] → throws (invalid config)
* - valid integer string → returns that number
*
* This function is the single call site for reading MUTATION_BREAK.
* stryker.config.mjs imports and calls it so CI shards with a bad
* MUTATION_BREAK fail immediately rather than silently falling back to 60
* and bypassing a per-module floor above 60 (e.g. prompt-budget: 90).
*
* @param {string|undefined} raw - value of process.env.MUTATION_BREAK
* @returns {number}
*/
function resolveMutationBreak(raw) {
if (raw === undefined) {
// Local run with no MUTATION_BREAK set — use documented backstop.
return 60;
}
if (typeof raw !== 'string' || raw.trim() === '') {
throw new Error(
'MUTATION_BREAK is set but empty — CI shard wiring is broken (matrix.minScore missing?)'
);
}
const n = Number(raw);
if (!Number.isFinite(n) || n < 1 || n > 100) {
throw new Error(
`MUTATION_BREAK invalid: "${raw}" (expected a per-module minScore 1-100)`
);
}
return n;
}
// Export internals for programmatic use (tests/mutation-matrix-ratchet.test.cjs).
// The require.main guard prevents main() from running when this file is require()d.
module.exports = { COVERED, TARGET_MUTATION_SCORE, resolveMutationBreak, readStdinSync };
if (require.main === module) runMain(main);