Files
msd-core/tests/ci-full-lane-sharding.test.cjs
Tom Boucher f6257f3745 ci(#2952): shard the full test lane instead of widening its cap (#2960)
* ci(#2952): budget CI job timeouts by headroom over measured cost

`origin/next` was red. The only failing check was `Required tests`, and its
sole cause was `test (ubuntu-latest, 24)` reported `cancelled` — GitHub's
conclusion for a job that exceeds its `timeout-minutes`, not a button press.

That job's `ubuntu-latest / 24` entry is the only `scope: full` matrix entry:
it runs the whole unit suite under c8 coverage, then the scripts/ coverage
floor, integration, security, install and slow, serially on one runner. On
next@5a0a9f097 it ran 15m16s against `timeout-minutes: 15` and was axed 23s
into `npm run test:slow`. Projected to a completed slow step (27s on the last
green run) the lane costs ~15m20s.

Confirmed hypothesis: the budget, not the suite. The lane had been riding the
ceiling all day — 12m03s, 11m34s, 11m50s, 14m25s, 14m51s — and crossed on
three of the last four full-lane runs (d2d2f7c08, 07603df8f, 5a0a9f097). There
is no pathological test: the unit run is cost-first bin-packed into 12 chunks,
chunk 1 is gated by run-tests-harness.test.cjs at 169s (expensive by design —
it spawns real harness subprocesses, one of which exercises the per-chunk
timeout), and the remaining chunks are 32-94s. 769s is the honest cost of 689
files under c8. Re-running could not have helped; the work exceeded the budget.

Review of the first cut surfaced the same defect one runner away: `full test
(windows-latest, 22, shard 3/3)` reached 18m59s against its own 20-minute cap
on 05b170e44 (94%) and 18m14s on 81eeb8a53 (91%). That lane has already blown
its cap twice (#1051, #1212). Fixed here rather than deferred.

The first cut also asserted `test >= test-full`, which is unsound — those two
budgets are dominated by different platforms, so their ordering carries no
meaning. Replaced with the invariant that actually generalises: every lane is
held to a headroom FACTOR over its own measured cost. `test` 15 -> 25 (1.5x of
16m), `test-full` 20 -> 30 (1.5x of 19m), `test-inert` unchanged at 15.

tests/ci-test-job-timeout-budget.test.cjs locks that rule. No unit test can
prove a lane still FITS its budget — only a real run measures that — but a
budget can no longer be lowered back beneath what its lane is known to need,
and a lane that gets slower must be re-measured rather than excused.

Separately: the earlier `failure` at 05b170e44 was an unrelated, already-fixed
CONTEXT-INDEX.json drift (07603df8f re-synced it; lint-tests is green at HEAD).
07603df8f's own run hit this same timeout, which is why it never reported green.

Refs #869, #1051, #1212

* ci(#2952): shard the full test lane instead of widening its cap

The `scope: full` lane was the only unsharded lane in this file. It ran the
entire unit suite under c8 on one runner, grew past a 15-minute cap, and
reddened `next`. Raising the cap bought room; it did not change the shape, and
the same lane would have walked back into the ceiling. Shard it, the way #1212
answered this for the Windows lane.

Balance comes from measurement, not file counts. scripts/run-tests.cjs already
partitions by measured per-file duration using LPT (#2472); the table it reads
was 10 days stale — 638 of 695 files timed, 64 missing, including the whole
context-predicates group. Regenerated from a verified matrix run: 700 files, 0
missing. On that table the 685-file unit suite splits 19.37m / 19.37m / 19.37m
— 0.0% spread — and the split is a total, disjoint cover with 0 files dropped.
Completeness, disjointness, balance and determinism of the partition itself are
already pinned against selectShard in run-tests-harness.test.cjs, including a
fast-check property, so this change does not restate them.

Sharding a COVERAGE run is the part that needs care. A per-shard percentage is
meaningless — shard 2 never executes shard 1's files, so those read 0% — and
leaving the gate on the shards would have quietly measured a third of the tree.
Each shard now renders no report and only leaves raw V8 dumps; a new
`coverage-gate` job merges all three into one coverage/tmp and runs the gate
there. c8's default temp directory is where the download lands, so the ≥70%
lines / ≥60% branches gate and the ≥55% scripts floor run unmodified against
merged data.

Both surfaces call the same npm scripts rather than inlining c8 into YAML, so
the include/exclude globs and both thresholds stay defined once in package.json.
The workflow holding its own copy is the divergence this repo has a rule
against, and the new test cross-checks package.json so an inline reintroduction
fails rather than drifts.

tests/ci-full-lane-sharding.test.cjs covers the two ways this stays GREEN while
being wrong: an incomplete shard set (declare 1/3 and 2/3, never 3/3, and a
third of the suite silently stops running) and a coverage gate that stops being
required. required-tests now depends on coverage-gate and fails on it, while
still tolerating `skipped` so docs-only PRs are not blocked.

`timeout-minutes: 25` on the lane is deliberately left alone. The budget test
requires a real measurement before a lane's declared cost changes, and the
sharded cost is not measured until this PR's own CI run.

Refs #1212, #2472

* ci(#2952): tighten the sharded lane's budget to its measured cost

The sharding commit deliberately left `timeout-minutes: 25` alone, because
tests/ci-test-job-timeout-budget.test.cjs requires a real measurement before a
lane's declared cost changes and the sharded cost did not exist yet.

It exists now. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s, and
coverage-gate 1m20s. Shard 1 is the long pole because the unsharded aux suites
ride along on it, which is deliberate — they total ~1m35s and sharding them
would cost more than it saves.

So the lane's budget is 15 against a slowest measured shard of 8 minutes
(~1.9x), and coverage-gate joins LANE_COSTS at 2 minutes. 15 is the same number
the lane blew before sharding; the work behind it is now a third the size.

Merged coverage was checked against the pre-shard single-runner baseline rather
than assumed from a green check: 94.36 stmts / 96.3 funcs / 94.36 lines
identical, branches 84.22 vs 84.21 — one branch across two different trees,
noise rather than a regression.

---------

Co-authored-by: sim <sim@local>
2026-07-31 22:16:26 -04:00

278 lines
11 KiB
JavaScript

'use strict';
/**
* The full test lane is sharded, and the coverage gate that sharding displaced
* is still wired in — .github/workflows/test.yml (#2952).
*
* The `scope: full` lane was the only unsharded lane in this file. It ran the
* entire unit suite under c8 on a single runner, grew past a 15-minute cap, and
* reddened `next` (#2952). Raising the cap treated the symptom; sharding is the
* shape fix, and it is the same answer #1212 reached for the Windows lane.
*
* Sharding introduces two failure modes that stay GREEN while being wrong, so
* both are pinned here:
*
* 1. An incomplete shard set. If the matrix declares shards 1/3 and 2/3 but
* never 3/3, a third of the unit suite simply stops running and every check
* still passes. The partition MATH is already covered — completeness,
* disjointness, balance and determinism are asserted against selectShard in
* run-tests-harness.test.cjs (#1212), including a fast-check property. What
* is NOT covered there, and is asserted here, is that the WORKFLOW asks for
* a complete set: same denominator everywhere, numerators exactly 1..N.
*
* 2. A dropped coverage gate. A sharded run leaves each runner with a partial
* picture — shard 2 never executes shard 1's files, so those read 0%. The
* ≥70% gate therefore cannot live on the shards; it moved to `coverage-gate`,
* which merges every shard's raw V8 dumps. If that job silently stopped
* being required, coverage enforcement would vanish without any red check.
*/
const test = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const yaml = require('js-yaml');
const WORKFLOWS_DIR = path.join(__dirname, '..', '.github', 'workflows');
function loadWorkflow(name) {
return yaml.load(fs.readFileSync(path.join(WORKFLOWS_DIR, name), 'utf8'));
}
/**
* Parse an `i/n` shard spec into { index, total }, or null if malformed.
* Mirrors the grammar scripts/run-tests.cjs accepts for --shard.
*/
function parseShardSpec(spec) {
const m = /^(\d+)\/(\d+)$/.exec(String(spec));
if (!m) return null;
const index = Number(m[1]);
const total = Number(m[2]);
if (!Number.isInteger(index) || !Number.isInteger(total)) return null;
if (total < 1 || index < 1 || index > total) return null;
return { index, total };
}
/** True iff `specs` is exactly one complete shard set: same N, numerators 1..N. */
function isCompleteShardSet(specs) {
if (specs.length === 0) return false;
const parsed = specs.map(parseShardSpec);
if (parsed.some((p) => p === null)) return false;
const total = parsed[0].total;
if (parsed.some((p) => p.total !== total)) return false;
if (parsed.length !== total) return false;
const seen = new Set(parsed.map((p) => p.index));
return seen.size === total && [...seen].every((i) => i >= 1 && i <= total);
}
test('the full test lane is sharded and complete (#2952)', async (t) => {
const workflow = loadWorkflow('test.yml');
const include = workflow.jobs.test.strategy.matrix.include;
const fullLanes = include.filter((e) => e.scope === 'full');
await t.test('the full lane is actually sharded, not a single runner', () => {
assert.ok(fullLanes.length > 0, 'expected at least one `scope: full` matrix entry');
assert.ok(
fullLanes.length > 1,
'the `scope: full` lane is back to a single unsharded entry. That is the '
+ '#2952 regression: the whole unit suite under c8 on one runner grew past '
+ 'its cap and reddened `next`.',
);
for (const lane of fullLanes) {
assert.ok(
lane.shard !== undefined,
`a \`scope: full\` matrix entry declares no shard: ${JSON.stringify(lane)}`,
);
}
});
await t.test('the declared shards form one complete set', () => {
const specs = fullLanes.map((e) => e.shard);
assert.ok(
isCompleteShardSet(specs),
`the full lane's shards ${JSON.stringify(specs)} are not a complete set. `
+ 'Every entry must share one denominator N and the numerators must be '
+ 'exactly 1..N — a missing numerator silently stops running that slice of '
+ 'the unit suite while every check stays green.',
);
});
await t.test('only the full lane is sharded', () => {
for (const lane of include.filter((e) => e.scope !== 'full')) {
assert.equal(
lane.shard, undefined,
`non-full lane ${JSON.stringify(lane)} declares a shard; the scoped and `
+ 'Windows lanes run a selected file list, not a partition.',
);
}
});
await t.test('each shard runs its own slice, not the whole suite', () => {
const unitStep = workflow.jobs.test.steps.find(
(s) => typeof s.run === 'string' && s.run.includes('test:coverage:unit:raw'),
);
assert.ok(unitStep, 'no step in the test job runs the raw unit coverage script');
assert.match(
unitStep.run, /--shard \$\{\{ matrix\.shard \}\}/,
'the unit step does not pass matrix.shard through to run-tests.cjs, so '
+ 'every shard would run the ENTIRE suite — N times the cost, no speedup.',
);
});
await t.test('the aux suites are pinned to one shard that actually exists', () => {
// Pinning to a LIVE shard is the whole assertion. A pin to a shard the
// matrix no longer declares — say the count moves to 4 and these `if:`
// conditions keep naming 1/3 — means the aux suites stop running entirely
// while every check stays green. Checking only that some shard literal is
// present would not catch that, so the literal is resolved against the
// shards the matrix actually declares.
const declared = fullLanes.map((e) => String(e.shard));
const auxScripts = ['test:integration', 'test:security', 'test:install', 'test:slow'];
const pins = new Set();
for (const script of auxScripts) {
const step = workflow.jobs.test.steps.find(
(s) => typeof s.run === 'string' && s.run.includes(script),
);
assert.ok(step, `no step runs \`npm run ${script}\``);
const pin = /matrix\.shard == '([^']+)'/.exec(String(step.if));
assert.ok(
pin,
`the \`${script}\` step is not pinned to a single shard; it would run `
+ 'once per shard and multiply its cost for no extra signal.',
);
assert.ok(
declared.includes(pin[1]),
`the \`${script}\` step is pinned to shard '${pin[1]}', which the matrix `
+ `does not declare (${JSON.stringify(declared)}). That condition can `
+ 'never be true, so this suite would silently never run.',
);
pins.add(pin[1]);
}
assert.equal(
pins.size, 1,
`the aux suites are split across shards ${JSON.stringify([...pins])}; they `
+ 'are meant to run together on exactly one.',
);
});
});
test('the merged coverage gate survives sharding (#2952)', async (t) => {
const workflow = loadWorkflow('test.yml');
await t.test('a dedicated coverage-gate job exists and consumes the shards', () => {
const gate = workflow.jobs['coverage-gate'];
assert.ok(gate, '.github/workflows/test.yml declares no `coverage-gate` job');
assert.ok(
(gate.needs || []).includes('test'),
'coverage-gate must depend on `test` — it merges that job\'s shard artifacts',
);
const merges = gate.steps.some(
(s) => String(s.uses || '').includes('download-artifact')
&& s.with && s.with['merge-multiple'] === true,
);
assert.ok(
merges,
'coverage-gate does not download the shard artifacts with merge-multiple. '
+ 'Without every shard merged into one coverage/tmp, the gate scores a '
+ 'partial run: files no shard in hand executed read 0%.',
);
});
await t.test('both coverage thresholds are still enforced', () => {
const runs = workflow.jobs['coverage-gate'].steps
.map((s) => s.run).filter((r) => typeof r === 'string').join('\n');
const pkg = JSON.parse(
fs.readFileSync(path.join(__dirname, '..', 'package.json'), 'utf8'),
);
assert.match(
runs, /test:coverage:report/,
'coverage-gate never runs the coverage report+gate script — the >=70% '
+ 'lines / >=60% branches gate on gsd-core/bin/lib would be gone',
);
assert.match(
runs, /test:coverage:scripts-floor/,
'coverage-gate never enforces the >=55% scripts/ floor',
);
// The workflow calls npm scripts precisely so the thresholds are defined
// once. If a future edit inlines c8 into the YAML, the two surfaces can
// drift silently — the gate would still be green while measuring something
// other than what package.json says.
assert.match(
pkg.scripts['test:coverage:report'], /check-coverage-gate\.cjs/,
'test:coverage:report no longer runs check-coverage-gate.cjs',
);
assert.match(
pkg.scripts['test:coverage:scripts-floor'], /--lines 55/,
'test:coverage:scripts-floor no longer enforces 55%',
);
assert.match(
pkg.scripts['test:coverage:unit:raw'], /--suite unit/,
'test:coverage:unit:raw no longer runs the unit suite',
);
assert.doesNotMatch(
pkg.scripts['test:coverage:unit:raw'], /--shard/,
'the shard must come from the workflow matrix, not be baked into the script',
);
});
await t.test('required-tests fails when the coverage gate fails', () => {
const required = workflow.jobs['required-tests'];
assert.ok(
(required.needs || []).includes('coverage-gate'),
'required-tests does not depend on coverage-gate, so a red gate could not '
+ 'block a merge',
);
const summarize = required.steps[0];
assert.ok(
'COVERAGE_GATE_RESULT' in (summarize.env || {}),
'required-tests does not read needs.coverage-gate.result',
);
assert.match(
summarize.run, /coverage-gate did not pass/,
'required-tests never fails on a red coverage-gate — depending on a job '
+ 'without checking its result makes the dependency decorative',
);
});
await t.test('a skipped coverage gate is not treated as a failure', () => {
// coverage-gate is conditioned on product_changed, exactly like the test
// lane. A docs-only PR skips both; treating `skipped` as red would block
// every one of them.
assert.match(
workflow.jobs['required-tests'].steps[0].run,
/COVERAGE_GATE_RESULT" != "skipped"/,
'required-tests treats a skipped coverage-gate as a failure',
);
});
});
test('shard-spec parsing is exact at its boundaries (#2952)', () => {
assert.deepEqual(parseShardSpec('1/3'), { index: 1, total: 3 });
assert.deepEqual(parseShardSpec('3/3'), { index: 3, total: 3 });
assert.deepEqual(parseShardSpec('1/1'), { index: 1, total: 1 });
// index 0 is below the range, index total+1 above it.
assert.equal(parseShardSpec('0/3'), null);
assert.equal(parseShardSpec('4/3'), null);
assert.equal(parseShardSpec('1/0'), null);
assert.equal(parseShardSpec('1'), null);
assert.equal(parseShardSpec('a/3'), null);
assert.equal(parseShardSpec(''), null);
assert.equal(parseShardSpec(undefined), null);
// A complete set, and the three ways one stops being complete.
assert.equal(isCompleteShardSet(['1/3', '2/3', '3/3']), true);
assert.equal(isCompleteShardSet(['1/3', '2/3']), false, 'missing numerator');
assert.equal(isCompleteShardSet(['1/3', '2/3', '2/3']), false, 'duplicate numerator');
assert.equal(isCompleteShardSet(['1/3', '2/3', '3/4']), false, 'mixed denominator');
assert.equal(isCompleteShardSet([]), false, 'empty set');
});