* ci(#2952): budget CI job timeouts by headroom over measured cost `origin/next` was red. The only failing check was `Required tests`, and its sole cause was `test (ubuntu-latest, 24)` reported `cancelled` — GitHub's conclusion for a job that exceeds its `timeout-minutes`, not a button press. That job's `ubuntu-latest / 24` entry is the only `scope: full` matrix entry: it runs the whole unit suite under c8 coverage, then the scripts/ coverage floor, integration, security, install and slow, serially on one runner. On next@5a0a9f097 it ran 15m16s against `timeout-minutes: 15` and was axed 23s into `npm run test:slow`. Projected to a completed slow step (27s on the last green run) the lane costs ~15m20s. Confirmed hypothesis: the budget, not the suite. The lane had been riding the ceiling all day — 12m03s, 11m34s, 11m50s, 14m25s, 14m51s — and crossed on three of the last four full-lane runs (d2d2f7c08,07603df8f,5a0a9f097). There is no pathological test: the unit run is cost-first bin-packed into 12 chunks, chunk 1 is gated by run-tests-harness.test.cjs at 169s (expensive by design — it spawns real harness subprocesses, one of which exercises the per-chunk timeout), and the remaining chunks are 32-94s. 769s is the honest cost of 689 files under c8. Re-running could not have helped; the work exceeded the budget. Review of the first cut surfaced the same defect one runner away: `full test (windows-latest, 22, shard 3/3)` reached 18m59s against its own 20-minute cap on05b170e44(94%) and 18m14s on81eeb8a53(91%). That lane has already blown its cap twice (#1051, #1212). Fixed here rather than deferred. The first cut also asserted `test >= test-full`, which is unsound — those two budgets are dominated by different platforms, so their ordering carries no meaning. Replaced with the invariant that actually generalises: every lane is held to a headroom FACTOR over its own measured cost. `test` 15 -> 25 (1.5x of 16m), `test-full` 20 -> 30 (1.5x of 19m), `test-inert` unchanged at 15. tests/ci-test-job-timeout-budget.test.cjs locks that rule. No unit test can prove a lane still FITS its budget — only a real run measures that — but a budget can no longer be lowered back beneath what its lane is known to need, and a lane that gets slower must be re-measured rather than excused. Separately: the earlier `failure` at05b170e44was an unrelated, already-fixed CONTEXT-INDEX.json drift (07603df8fre-synced it; lint-tests is green at HEAD). 07603df8f's own run hit this same timeout, which is why it never reported green. Refs #869, #1051, #1212 * ci(#2952): shard the full test lane instead of widening its cap The `scope: full` lane was the only unsharded lane in this file. It ran the entire unit suite under c8 on one runner, grew past a 15-minute cap, and reddened `next`. Raising the cap bought room; it did not change the shape, and the same lane would have walked back into the ceiling. Shard it, the way #1212 answered this for the Windows lane. Balance comes from measurement, not file counts. scripts/run-tests.cjs already partitions by measured per-file duration using LPT (#2472); the table it reads was 10 days stale — 638 of 695 files timed, 64 missing, including the whole context-predicates group. Regenerated from a verified matrix run: 700 files, 0 missing. On that table the 685-file unit suite splits 19.37m / 19.37m / 19.37m — 0.0% spread — and the split is a total, disjoint cover with 0 files dropped. Completeness, disjointness, balance and determinism of the partition itself are already pinned against selectShard in run-tests-harness.test.cjs, including a fast-check property, so this change does not restate them. Sharding a COVERAGE run is the part that needs care. A per-shard percentage is meaningless — shard 2 never executes shard 1's files, so those read 0% — and leaving the gate on the shards would have quietly measured a third of the tree. Each shard now renders no report and only leaves raw V8 dumps; a new `coverage-gate` job merges all three into one coverage/tmp and runs the gate there. c8's default temp directory is where the download lands, so the ≥70% lines / ≥60% branches gate and the ≥55% scripts floor run unmodified against merged data. Both surfaces call the same npm scripts rather than inlining c8 into YAML, so the include/exclude globs and both thresholds stay defined once in package.json. The workflow holding its own copy is the divergence this repo has a rule against, and the new test cross-checks package.json so an inline reintroduction fails rather than drifts. tests/ci-full-lane-sharding.test.cjs covers the two ways this stays GREEN while being wrong: an incomplete shard set (declare 1/3 and 2/3, never 3/3, and a third of the suite silently stops running) and a coverage gate that stops being required. required-tests now depends on coverage-gate and fails on it, while still tolerating `skipped` so docs-only PRs are not blocked. `timeout-minutes: 25` on the lane is deliberately left alone. The budget test requires a real measurement before a lane's declared cost changes, and the sharded cost is not measured until this PR's own CI run. Refs #1212, #2472 * ci(#2952): tighten the sharded lane's budget to its measured cost The sharding commit deliberately left `timeout-minutes: 25` alone, because tests/ci-test-job-timeout-budget.test.cjs requires a real measurement before a lane's declared cost changes and the sharded cost did not exist yet. It exists now. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s, and coverage-gate 1m20s. Shard 1 is the long pole because the unsharded aux suites ride along on it, which is deliberate — they total ~1m35s and sharding them would cost more than it saves. So the lane's budget is 15 against a slowest measured shard of 8 minutes (~1.9x), and coverage-gate joins LANE_COSTS at 2 minutes. 15 is the same number the lane blew before sharding; the work behind it is now a third the size. Merged coverage was checked against the pre-shard single-runner baseline rather than assumed from a green check: 94.36 stmts / 96.3 funcs / 94.36 lines identical, branches 84.22 vs 84.21 — one branch across two different trees, noise rather than a regression. --------- Co-authored-by: sim <sim@local>
158 lines
6.8 KiB
JavaScript
158 lines
6.8 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* CI job timeout budgets — .github/workflows/test.yml (#2952).
|
|
*
|
|
* A GitHub Actions job that exceeds its `timeout-minutes` is reported
|
|
* `cancelled`, not `failed`. That conclusion propagates into `Required tests`
|
|
* and reddens the branch, while reading like someone hit the cancel button —
|
|
* which is what makes this failure mode expensive to diagnose and worth a gate.
|
|
*
|
|
* It has now happened three times in this repo: #1051 and #1212 on the Windows
|
|
* full-test lane, and #2952 on the unsharded `test` lane, whose
|
|
* `ubuntu-latest / 24` entry is the only `scope: full` matrix entry — it runs
|
|
* the whole unit suite under c8 coverage, then the scripts/ coverage floor,
|
|
* integration, security, install and slow, serially on one runner. Every one of
|
|
* those was the same root cause: a budget sized to what the lane cost that
|
|
* week, with no headroom for the suite to grow into.
|
|
*
|
|
* So the rule enforced here is not a fixed number per lane — it is a HEADROOM
|
|
* FACTOR over each lane's measured cost. A lane may be slow; what it may not be
|
|
* is budgeted to finish with seconds to spare.
|
|
*
|
|
* This is a budget assertion, not a duration assertion. No unit test can prove
|
|
* a lane still FITS its budget — only a real CI run measures that. What this
|
|
* file guarantees is that a budget cannot be quietly lowered back beneath what
|
|
* its lane is already known to need.
|
|
*/
|
|
|
|
const test = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const yaml = require('js-yaml');
|
|
|
|
const WORKFLOWS_DIR = path.join(__dirname, '..', '.github', 'workflows');
|
|
|
|
function loadWorkflow(name) {
|
|
return yaml.load(fs.readFileSync(path.join(WORKFLOWS_DIR, name), 'utf8'));
|
|
}
|
|
|
|
/**
|
|
* Multiplier applied to a lane's measured cost to get its required budget.
|
|
* 1.5x is enough slack for ordinary suite growth across a release cycle without
|
|
* letting a genuinely runaway lane hide behind a large number.
|
|
*/
|
|
const HEADROOM_FACTOR = 1.5;
|
|
|
|
/**
|
|
* Measured wall-clock cost per lane, in whole minutes rounded UP, each from a
|
|
* named run. Raise an entry only alongside a fresh measurement — never to make
|
|
* a red gate green. Raising a measurement raises the required budget with it,
|
|
* which is the point: a lane that got slower must be re-budgeted, not excused.
|
|
*/
|
|
const LANE_COSTS = [
|
|
{
|
|
job: 'test',
|
|
measuredMinutes: 8,
|
|
// Sharded three ways as of #2952, so this is ONE shard's cost, not the
|
|
// whole unit suite. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s.
|
|
// Shard 1 is the long pole because the unsharded aux suites ride on it.
|
|
// Before sharding the same lane cost 15m20s and blew a 15-minute cap.
|
|
evidence: 'run 30677442953 — 7m12s slowest shard',
|
|
},
|
|
{
|
|
job: 'test-full',
|
|
measuredMinutes: 19,
|
|
// Worst observed shard is `full test (windows-latest, 22, shard 3/3)`:
|
|
// 18m59s on 05b170e44 and 18m14s on 81eeb8a53. The Windows shards are slow
|
|
// for platform reasons, not extra work.
|
|
evidence: 'run 30650559192 — 18m59s, windows-22 shard 3/3',
|
|
},
|
|
{
|
|
job: 'coverage-gate',
|
|
measuredMinutes: 2,
|
|
// Downloads three shards' raw V8 dumps, renders one merged report and runs
|
|
// both thresholds. Run 30677442953: 1m20s end to end, most of it npm ci.
|
|
evidence: 'run 30677442953 — 1m20s',
|
|
},
|
|
{
|
|
job: 'test-inert',
|
|
measuredMinutes: 2,
|
|
// Runs only the targeted-test step when no product code changed; observed
|
|
// around a minute. Listed so its budget cannot be dropped to nothing.
|
|
evidence: 'targeted-only lane, ~1m observed',
|
|
},
|
|
];
|
|
|
|
function requiredBudgetMinutes(measuredMinutes, headroomFactor = HEADROOM_FACTOR) {
|
|
return Math.ceil(measuredMinutes * headroomFactor);
|
|
}
|
|
|
|
function hasSufficientBudget(budgetMinutes, measuredMinutes, headroomFactor = HEADROOM_FACTOR) {
|
|
return Number.isInteger(budgetMinutes)
|
|
&& budgetMinutes >= requiredBudgetMinutes(measuredMinutes, headroomFactor);
|
|
}
|
|
|
|
test('CI job timeout budgets carry headroom over measured cost (#2952)', async (t) => {
|
|
const workflow = loadWorkflow('test.yml');
|
|
|
|
for (const lane of LANE_COSTS) {
|
|
await t.test(`${lane.job} is budgeted above its measured cost`, () => {
|
|
assert.ok(
|
|
workflow.jobs && Object.prototype.hasOwnProperty.call(workflow.jobs, lane.job),
|
|
`.github/workflows/test.yml declares no job \`${lane.job}\`. If it was `
|
|
+ 'renamed or removed, update LANE_COSTS in this file to match — do not '
|
|
+ 'delete the entry to make this pass.',
|
|
);
|
|
|
|
const budget = workflow.jobs[lane.job]['timeout-minutes'];
|
|
const required = requiredBudgetMinutes(lane.measuredMinutes);
|
|
|
|
assert.equal(
|
|
typeof budget, 'number',
|
|
`.github/workflows/test.yml jobs.${lane.job} must declare timeout-minutes`,
|
|
);
|
|
assert.ok(
|
|
hasSufficientBudget(budget, lane.measuredMinutes),
|
|
`jobs.${lane.job}.timeout-minutes is ${budget}, but the lane measured `
|
|
+ `${lane.measuredMinutes}m (${lane.evidence}) and needs at least `
|
|
+ `${required} — ${HEADROOM_FACTOR}x — so suite growth does not breach `
|
|
+ 'the cap. A job that exceeds timeout-minutes is reported `cancelled` '
|
|
+ 'and reddens `Required tests`.',
|
|
);
|
|
});
|
|
}
|
|
|
|
// Boundary coverage on the predicate that decides every lane above:
|
|
// required-1 must be rejected, required and required+1 accepted.
|
|
await t.test('budget sufficiency is exact at the boundary', () => {
|
|
for (const lane of LANE_COSTS) {
|
|
const required = requiredBudgetMinutes(lane.measuredMinutes);
|
|
|
|
assert.equal(hasSufficientBudget(required - 1, lane.measuredMinutes), false,
|
|
`${lane.job}: a budget one minute under the requirement must be rejected`);
|
|
assert.equal(hasSufficientBudget(required, lane.measuredMinutes), true,
|
|
`${lane.job}: a budget exactly at the requirement must be accepted`);
|
|
assert.equal(hasSufficientBudget(required + 1, lane.measuredMinutes), true,
|
|
`${lane.job}: a budget over the requirement must be accepted`);
|
|
}
|
|
});
|
|
|
|
await t.test('a non-integer budget is not a sufficient budget', () => {
|
|
// `timeout-minutes: 24.5` is not something GitHub accepts; treating it as
|
|
// sufficient would let a malformed workflow through this gate.
|
|
assert.equal(hasSufficientBudget(24.5, 16), false);
|
|
assert.equal(hasSufficientBudget(Number.NaN, 16), false);
|
|
assert.equal(hasSufficientBudget(undefined, 16), false);
|
|
});
|
|
|
|
await t.test('requiredBudgetMinutes rounds up rather than truncating', () => {
|
|
// 15 * 1.5 = 22.5 — truncation would hand back 22 and under-budget the lane.
|
|
assert.equal(requiredBudgetMinutes(15, 1.5), 23);
|
|
assert.equal(requiredBudgetMinutes(16, 1.5), 24);
|
|
assert.equal(requiredBudgetMinutes(19, 1.5), 29);
|
|
assert.equal(requiredBudgetMinutes(10, 1.5), 15);
|
|
});
|
|
});
|