* ci(#2952): budget CI job timeouts by headroom over measured cost `origin/next` was red. The only failing check was `Required tests`, and its sole cause was `test (ubuntu-latest, 24)` reported `cancelled` — GitHub's conclusion for a job that exceeds its `timeout-minutes`, not a button press. That job's `ubuntu-latest / 24` entry is the only `scope: full` matrix entry: it runs the whole unit suite under c8 coverage, then the scripts/ coverage floor, integration, security, install and slow, serially on one runner. On next@5a0a9f097 it ran 15m16s against `timeout-minutes: 15` and was axed 23s into `npm run test:slow`. Projected to a completed slow step (27s on the last green run) the lane costs ~15m20s. Confirmed hypothesis: the budget, not the suite. The lane had been riding the ceiling all day — 12m03s, 11m34s, 11m50s, 14m25s, 14m51s — and crossed on three of the last four full-lane runs (d2d2f7c08,07603df8f,5a0a9f097). There is no pathological test: the unit run is cost-first bin-packed into 12 chunks, chunk 1 is gated by run-tests-harness.test.cjs at 169s (expensive by design — it spawns real harness subprocesses, one of which exercises the per-chunk timeout), and the remaining chunks are 32-94s. 769s is the honest cost of 689 files under c8. Re-running could not have helped; the work exceeded the budget. Review of the first cut surfaced the same defect one runner away: `full test (windows-latest, 22, shard 3/3)` reached 18m59s against its own 20-minute cap on05b170e44(94%) and 18m14s on81eeb8a53(91%). That lane has already blown its cap twice (#1051, #1212). Fixed here rather than deferred. The first cut also asserted `test >= test-full`, which is unsound — those two budgets are dominated by different platforms, so their ordering carries no meaning. Replaced with the invariant that actually generalises: every lane is held to a headroom FACTOR over its own measured cost. `test` 15 -> 25 (1.5x of 16m), `test-full` 20 -> 30 (1.5x of 19m), `test-inert` unchanged at 15. tests/ci-test-job-timeout-budget.test.cjs locks that rule. No unit test can prove a lane still FITS its budget — only a real run measures that — but a budget can no longer be lowered back beneath what its lane is known to need, and a lane that gets slower must be re-measured rather than excused. Separately: the earlier `failure` at05b170e44was an unrelated, already-fixed CONTEXT-INDEX.json drift (07603df8fre-synced it; lint-tests is green at HEAD). 07603df8f's own run hit this same timeout, which is why it never reported green. Refs #869, #1051, #1212 * ci(#2952): shard the full test lane instead of widening its cap The `scope: full` lane was the only unsharded lane in this file. It ran the entire unit suite under c8 on one runner, grew past a 15-minute cap, and reddened `next`. Raising the cap bought room; it did not change the shape, and the same lane would have walked back into the ceiling. Shard it, the way #1212 answered this for the Windows lane. Balance comes from measurement, not file counts. scripts/run-tests.cjs already partitions by measured per-file duration using LPT (#2472); the table it reads was 10 days stale — 638 of 695 files timed, 64 missing, including the whole context-predicates group. Regenerated from a verified matrix run: 700 files, 0 missing. On that table the 685-file unit suite splits 19.37m / 19.37m / 19.37m — 0.0% spread — and the split is a total, disjoint cover with 0 files dropped. Completeness, disjointness, balance and determinism of the partition itself are already pinned against selectShard in run-tests-harness.test.cjs, including a fast-check property, so this change does not restate them. Sharding a COVERAGE run is the part that needs care. A per-shard percentage is meaningless — shard 2 never executes shard 1's files, so those read 0% — and leaving the gate on the shards would have quietly measured a third of the tree. Each shard now renders no report and only leaves raw V8 dumps; a new `coverage-gate` job merges all three into one coverage/tmp and runs the gate there. c8's default temp directory is where the download lands, so the ≥70% lines / ≥60% branches gate and the ≥55% scripts floor run unmodified against merged data. Both surfaces call the same npm scripts rather than inlining c8 into YAML, so the include/exclude globs and both thresholds stay defined once in package.json. The workflow holding its own copy is the divergence this repo has a rule against, and the new test cross-checks package.json so an inline reintroduction fails rather than drifts. tests/ci-full-lane-sharding.test.cjs covers the two ways this stays GREEN while being wrong: an incomplete shard set (declare 1/3 and 2/3, never 3/3, and a third of the suite silently stops running) and a coverage gate that stops being required. required-tests now depends on coverage-gate and fails on it, while still tolerating `skipped` so docs-only PRs are not blocked. `timeout-minutes: 25` on the lane is deliberately left alone. The budget test requires a real measurement before a lane's declared cost changes, and the sharded cost is not measured until this PR's own CI run. Refs #1212, #2472 * ci(#2952): tighten the sharded lane's budget to its measured cost The sharding commit deliberately left `timeout-minutes: 25` alone, because tests/ci-test-job-timeout-budget.test.cjs requires a real measurement before a lane's declared cost changes and the sharded cost did not exist yet. It exists now. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s, and coverage-gate 1m20s. Shard 1 is the long pole because the unsharded aux suites ride along on it, which is deliberate — they total ~1m35s and sharding them would cost more than it saves. So the lane's budget is 15 against a slowest measured shard of 8 minutes (~1.9x), and coverage-gate joins LANE_COSTS at 2 minutes. 15 is the same number the lane blew before sharding; the work behind it is now a third the size. Merged coverage was checked against the pre-shard single-runner baseline rather than assumed from a green check: 94.36 stmts / 96.3 funcs / 94.36 lines identical, branches 84.22 vs 84.21 — one branch across two different trees, noise rather than a regression. --------- Co-authored-by: sim <sim@local>
This commit is contained in:
174
.github/workflows/test.yml
vendored
174
.github/workflows/test.yml
vendored
@@ -112,10 +112,22 @@ jobs:
|
||||
run: npm run lint:ci
|
||||
|
||||
test:
|
||||
name: test (${{ matrix.os }}, ${{ matrix.node-version }})
|
||||
name: test (${{ matrix.os }}, ${{ matrix.node-version }}${{ matrix.shard && format(', shard {0}', matrix.shard) || '' }})
|
||||
needs: changes
|
||||
if: needs.changes.outputs.product_changed == 'true'
|
||||
runs-on: ${{ matrix.os }}
|
||||
# #2952: this lane is sharded three ways (see the matrix below), so the
|
||||
# budget covers ONE shard, not the whole unit suite. Measured on run
|
||||
# 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s. Shard 1 is the long
|
||||
# pole because the unsharded aux suites (integration/security/install/slow)
|
||||
# ride along on it — that is deliberate; they total ~1m35s and sharding them
|
||||
# would cost more than it saves.
|
||||
#
|
||||
# 15 is ~1.9x the slowest measured shard. Before sharding this same lane ran
|
||||
# 15m20s against a 15-minute cap and was killed mid-run, which is the whole
|
||||
# of #2952 — the number is unchanged, the work behind it is a third the size.
|
||||
# tests/ci-test-job-timeout-budget.test.cjs holds every lane here to a
|
||||
# headroom factor over its own measured cost.
|
||||
timeout-minutes: 15
|
||||
env:
|
||||
GSD_PLUGIN_ROOT: .ci-gsd-plugin-root-disabled
|
||||
@@ -136,12 +148,31 @@ jobs:
|
||||
# and one Windows shell/path lane. The non-primary OS/runtime lanes run
|
||||
# scoped tests from scripts/ci-test-scope.cjs; Ubuntu/Node 24 runs the
|
||||
# broader default suite.
|
||||
#
|
||||
# #2952: the `scope: full` lane is SHARDED three ways. It was the only
|
||||
# unsharded lane in this file, and the whole unit suite under c8 on one
|
||||
# runner grew until it blew a 15-minute cap and reddened `next`. Raising
|
||||
# the cap treated the symptom; sharding changes the shape. Shards are
|
||||
# partitioned by MEASURED per-file duration (tests/test-timings.json)
|
||||
# using LPT in scripts/run-tests.cjs — the same cost-aware packer #2472
|
||||
# gave the `test-full` lane, which measures 0.0% spread across 3 bins on
|
||||
# the current table. The aux suites (integration/security/install/slow)
|
||||
# total ~1m35s and are not worth sharding; they run on shard 1 only.
|
||||
- os: ubuntu-latest
|
||||
node-version: 22
|
||||
scope: targeted
|
||||
- os: ubuntu-latest
|
||||
node-version: 24
|
||||
scope: full
|
||||
shard: 1/3
|
||||
- os: ubuntu-latest
|
||||
node-version: 24
|
||||
scope: full
|
||||
shard: 2/3
|
||||
- os: ubuntu-latest
|
||||
node-version: 24
|
||||
scope: full
|
||||
shard: 3/3
|
||||
- os: windows-latest
|
||||
node-version: 24
|
||||
scope: windows
|
||||
@@ -226,55 +257,47 @@ jobs:
|
||||
if: matrix.scope != 'full'
|
||||
run: node scripts/run-tests.cjs --files-from .ci-selected-tests.txt
|
||||
|
||||
# The unit suite runs ONCE here, under c8 with the coverage gate — the
|
||||
# former standalone `coverage` job duplicated this lane's entire unit
|
||||
# run (~4 min of runner time per PR) just to collect the same numbers.
|
||||
- name: Run unit tests (coverage gate ≥70% on gsd-core/bin/lib)
|
||||
# #2952: each shard runs its slice of the unit suite under c8 but renders
|
||||
# NO report and enforces NO gate — it only leaves raw V8 dumps in
|
||||
# coverage/tmp. A per-shard percentage is meaningless (shard 2 never
|
||||
# executes shard 1's files, so every file outside its slice reads 0%), so
|
||||
# the gate has to see all three slices merged. The `coverage-gate` job
|
||||
# below does that. Both call the SAME npm scripts, so the c8 globs and the
|
||||
# ≥70% lib / ≥55% scripts floor thresholds stay defined once, in package.json.
|
||||
- name: Run unit tests (shard ${{ matrix.shard }}, raw coverage only)
|
||||
if: matrix.scope == 'full'
|
||||
env:
|
||||
NODE_OPTIONS: --max-old-space-size=6144
|
||||
run: npm run test:coverage:unit
|
||||
run: npm run test:coverage:unit:raw -- --shard ${{ matrix.shard }}
|
||||
|
||||
# Second-tier floor over the CI/release/lint tooling itself. Re-slices
|
||||
# the SAME V8 coverage data left in coverage/tmp by the run above — no
|
||||
# extra suite execution. Audit 2026-06: scripts/ measured 65.95%; the
|
||||
# 55% floor prevents a collapse to zero-coverage tooling while leaving
|
||||
# headroom for variance. The threshold lives in package.json next to
|
||||
# the 70% lib gate — raise deliberately, never lower.
|
||||
- name: Coverage floor — scripts/ tooling (≥55%)
|
||||
# Raw V8 dumps, not rendered reports — coverage-gate merges these. They
|
||||
# compress hard in the artifact zip, but they are still bulk: keep the
|
||||
# retention at one day so they cannot accumulate against the org quota.
|
||||
- name: Upload raw coverage for merge
|
||||
if: matrix.scope == 'full'
|
||||
run: npm run test:coverage:scripts-floor
|
||||
|
||||
- name: Upload coverage artifact
|
||||
if: always() && matrix.scope == 'full'
|
||||
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||
with:
|
||||
name: coverage-unit
|
||||
# coverage/tmp holds raw per-process V8 dumps (>1 GB for the full
|
||||
# unit suite) — exclude it; the rendered reports are the artifact.
|
||||
path: |
|
||||
coverage/
|
||||
!coverage/tmp
|
||||
.nyc_output/
|
||||
if-no-files-found: ignore
|
||||
# ~440 MB per full run; same-run diagnostic never consumed by other
|
||||
# jobs — keep short so it cannot accumulate against the org quota.
|
||||
retention-days: 3
|
||||
name: coverage-tmp-shard-${{ strategy.job-index }}
|
||||
path: coverage/tmp
|
||||
if-no-files-found: error
|
||||
retention-days: 1
|
||||
|
||||
# The aux suites are small (~1m35s combined) and unsharded — running them
|
||||
# on every shard would triple their cost for no signal.
|
||||
- name: Run integration tests
|
||||
if: matrix.scope == 'full'
|
||||
if: matrix.scope == 'full' && matrix.shard == '1/3'
|
||||
run: npm run test:integration
|
||||
|
||||
- name: Run security tests
|
||||
if: matrix.scope == 'full'
|
||||
if: matrix.scope == 'full' && matrix.shard == '1/3'
|
||||
run: npm run test:security
|
||||
|
||||
- name: Run install tests
|
||||
if: matrix.scope == 'full' && needs.changes.outputs.full_matrix == 'true'
|
||||
if: matrix.scope == 'full' && matrix.shard == '1/3' && needs.changes.outputs.full_matrix == 'true'
|
||||
run: npm run test:install
|
||||
|
||||
- name: Run slow tests
|
||||
if: matrix.scope == 'full' && needs.changes.outputs.full_matrix == 'true'
|
||||
if: matrix.scope == 'full' && matrix.shard == '1/3' && needs.changes.outputs.full_matrix == 'true'
|
||||
run: npm run test:slow
|
||||
|
||||
test-inert:
|
||||
@@ -353,7 +376,13 @@ jobs:
|
||||
# and a NESTED `leg.os` key is not resolvable by the H1 shell-policy linter
|
||||
# in scripts/workflow-policy.cjs, which reads `matrix.os`/`matrix.shell`
|
||||
# directly. Explicit rows keep both the cross-product and the linter happy.)
|
||||
timeout-minutes: 20
|
||||
# #2952: `full test (windows-latest, 22, shard 3/3)` reached 18m59s (94% of
|
||||
# a 20-minute cap) on 05b170e44 and 18m14s (91%) on 81eeb8a53. The Windows
|
||||
# shards are slow for platform reasons — process spawn and filesystem cost,
|
||||
# not extra work — and this lane has already blown its cap twice before
|
||||
# (#1051, #1212). 30 is ~1.5x the worst observed shard, restoring the
|
||||
# headroom that #1212's sharding bought and this suite has since eaten.
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
GSD_PLUGIN_ROOT: .ci-gsd-plugin-root-disabled
|
||||
# #2854: pin the emitted gate's baseline to the SAME commit the tree was merged
|
||||
@@ -482,6 +511,67 @@ jobs:
|
||||
if: matrix.shard == 1
|
||||
run: npm run test:security
|
||||
|
||||
# #2952: the coverage gate that the `test` lane used to run inline. Sharding
|
||||
# the unit suite means no single runner sees the whole picture, so the gate
|
||||
# moves here, downloads every shard's raw V8 dumps into one coverage/tmp, and
|
||||
# runs the SAME commands as before against the merged data. Splitting the
|
||||
# gate out is what makes sharding safe: nothing about the thresholds or the
|
||||
# include/exclude globs changes, only where they are evaluated.
|
||||
coverage-gate:
|
||||
name: Coverage gate (merged shards)
|
||||
needs: [changes, test]
|
||||
if: needs.changes.outputs.product_changed == 'true' && needs.test.result == 'success'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
||||
with:
|
||||
persist-credentials: true
|
||||
token: ${{ github.token }}
|
||||
|
||||
- name: Set up Node.js 24
|
||||
uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
|
||||
with:
|
||||
node-version: 24
|
||||
cache: 'npm'
|
||||
|
||||
- name: Install dependencies
|
||||
run: npm ci
|
||||
|
||||
# merge-multiple flattens every shard's dumps into ONE coverage/tmp, which
|
||||
# is exactly the layout c8 expects from a single run. The dumps are
|
||||
# per-process files with distinct names, so there is nothing to collide.
|
||||
- name: Download every shard's raw coverage
|
||||
uses: actions/download-artifact@018cc2cf5baa6db3ef3c5f8a56943fffe632ef53 # v6.0.0
|
||||
with:
|
||||
pattern: coverage-tmp-shard-*
|
||||
path: coverage/tmp
|
||||
merge-multiple: true
|
||||
|
||||
# c8's default --temp-directory is ./coverage/tmp, which is exactly where
|
||||
# the download above landed every shard's dumps — so these run unmodified
|
||||
# against merged data and the thresholds stay defined in package.json.
|
||||
- name: Report merged coverage + gate gsd-core/bin/lib (≥70% lines, ≥60% branches)
|
||||
run: npm run test:coverage:report
|
||||
|
||||
# Second-tier floor over the CI/release/lint tooling itself. Re-slices the
|
||||
# SAME merged V8 data — no extra suite execution. Audit 2026-06: scripts/
|
||||
# measured 65.95%; the 55% floor prevents a collapse to zero-coverage
|
||||
# tooling while leaving headroom for variance. Raise deliberately, never lower.
|
||||
- name: Coverage floor — scripts/ tooling (≥55%)
|
||||
run: npm run test:coverage:scripts-floor
|
||||
|
||||
- name: Upload merged coverage report
|
||||
if: always()
|
||||
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||
with:
|
||||
name: coverage-unit-merged
|
||||
path: |
|
||||
coverage/
|
||||
!coverage/tmp
|
||||
if-no-files-found: ignore
|
||||
retention-days: 3
|
||||
|
||||
required-tests:
|
||||
name: Required tests
|
||||
needs:
|
||||
@@ -490,6 +580,7 @@ jobs:
|
||||
- test
|
||||
- test-inert
|
||||
- test-full
|
||||
- coverage-gate
|
||||
if: always()
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 1
|
||||
@@ -503,6 +594,7 @@ jobs:
|
||||
TEST_RESULT: ${{ needs.test.result }}
|
||||
INERT_RESULT: ${{ needs.test-inert.result }}
|
||||
FULL_TEST_RESULT: ${{ needs.test-full.result }}
|
||||
COVERAGE_GATE_RESULT: ${{ needs.coverage-gate.result }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
echo "code_changed=$CODE_CHANGED"
|
||||
@@ -512,6 +604,7 @@ jobs:
|
||||
echo "test=$TEST_RESULT"
|
||||
echo "test-inert=$INERT_RESULT"
|
||||
echo "test-full=$FULL_TEST_RESULT"
|
||||
echo "coverage-gate=$COVERAGE_GATE_RESULT"
|
||||
|
||||
if [ "$CHANGES_RESULT" != "success" ]; then
|
||||
echo "::error::test scope detection did not pass"
|
||||
@@ -533,12 +626,23 @@ jobs:
|
||||
echo "::error::test matrix did not pass"
|
||||
exit 1
|
||||
fi
|
||||
# The coverage gates (lib 70% + scripts/ 55% floor) run inside the
|
||||
# ubuntu/24 full lane of the `test` matrix, so TEST_RESULT covers them.
|
||||
# #2952: the coverage gates no longer run inside the `test` matrix —
|
||||
# sharding moved them to the `coverage-gate` job, which merges every
|
||||
# shard's dumps. TEST_RESULT therefore does NOT cover them any more;
|
||||
# COVERAGE_GATE_RESULT below is what gates coverage.
|
||||
if [ "$FULL_TEST_RESULT" != "success" ] && [ "$FULL_TEST_RESULT" != "skipped" ]; then
|
||||
echo "::error::full parity matrix did not pass"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# #2952: the coverage gate is skipped when product code did not change
|
||||
# (same condition as the test lane). Only a non-success, non-skipped
|
||||
# result is a failure — treating `skipped` as red would block every
|
||||
# docs-only PR.
|
||||
if [ "$COVERAGE_GATE_RESULT" != "success" ] && [ "$COVERAGE_GATE_RESULT" != "skipped" ]; then
|
||||
echo "::error::coverage-gate did not pass"
|
||||
exit 1
|
||||
fi
|
||||
else
|
||||
if [ "$INERT_RESULT" != "success" ]; then
|
||||
echo "::error::inert CI lane did not pass"
|
||||
|
||||
@@ -129,6 +129,8 @@
|
||||
"test:coverage": "c8 --check-coverage --lines 70 --branches 60 --reporter text --include 'gsd-core/bin/lib/*.cjs' --exclude 'tests/**' --all node scripts/run-tests.cjs",
|
||||
"test:coverage:scripts-floor": "c8 check-coverage --lines 55 --include 'scripts/**/*.cjs' --exclude 'tests/**' --all",
|
||||
"test:coverage:unit": "c8 --reporter text --reporter json-summary --include 'gsd-core/bin/lib/*.cjs' --exclude 'tests/**' --all node scripts/run-tests.cjs --suite unit && node scripts/check-coverage-gate.cjs",
|
||||
"test:coverage:unit:raw": "c8 --reporter none node scripts/run-tests.cjs --suite unit",
|
||||
"test:coverage:report": "c8 report --reporter text --reporter json-summary --include 'gsd-core/bin/lib/*.cjs' --exclude 'tests/**' --all && node scripts/check-coverage-gate.cjs",
|
||||
"test:coverage:all": "npm run test:coverage",
|
||||
"test:mutation": "stryker run",
|
||||
"test:mutation:since": "stryker run --incremental --since origin/next"
|
||||
|
||||
277
tests/ci-full-lane-sharding.test.cjs
Normal file
277
tests/ci-full-lane-sharding.test.cjs
Normal file
@@ -0,0 +1,277 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* The full test lane is sharded, and the coverage gate that sharding displaced
|
||||
* is still wired in — .github/workflows/test.yml (#2952).
|
||||
*
|
||||
* The `scope: full` lane was the only unsharded lane in this file. It ran the
|
||||
* entire unit suite under c8 on a single runner, grew past a 15-minute cap, and
|
||||
* reddened `next` (#2952). Raising the cap treated the symptom; sharding is the
|
||||
* shape fix, and it is the same answer #1212 reached for the Windows lane.
|
||||
*
|
||||
* Sharding introduces two failure modes that stay GREEN while being wrong, so
|
||||
* both are pinned here:
|
||||
*
|
||||
* 1. An incomplete shard set. If the matrix declares shards 1/3 and 2/3 but
|
||||
* never 3/3, a third of the unit suite simply stops running and every check
|
||||
* still passes. The partition MATH is already covered — completeness,
|
||||
* disjointness, balance and determinism are asserted against selectShard in
|
||||
* run-tests-harness.test.cjs (#1212), including a fast-check property. What
|
||||
* is NOT covered there, and is asserted here, is that the WORKFLOW asks for
|
||||
* a complete set: same denominator everywhere, numerators exactly 1..N.
|
||||
*
|
||||
* 2. A dropped coverage gate. A sharded run leaves each runner with a partial
|
||||
* picture — shard 2 never executes shard 1's files, so those read 0%. The
|
||||
* ≥70% gate therefore cannot live on the shards; it moved to `coverage-gate`,
|
||||
* which merges every shard's raw V8 dumps. If that job silently stopped
|
||||
* being required, coverage enforcement would vanish without any red check.
|
||||
*/
|
||||
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const yaml = require('js-yaml');
|
||||
|
||||
const WORKFLOWS_DIR = path.join(__dirname, '..', '.github', 'workflows');
|
||||
|
||||
function loadWorkflow(name) {
|
||||
return yaml.load(fs.readFileSync(path.join(WORKFLOWS_DIR, name), 'utf8'));
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse an `i/n` shard spec into { index, total }, or null if malformed.
|
||||
* Mirrors the grammar scripts/run-tests.cjs accepts for --shard.
|
||||
*/
|
||||
function parseShardSpec(spec) {
|
||||
const m = /^(\d+)\/(\d+)$/.exec(String(spec));
|
||||
if (!m) return null;
|
||||
const index = Number(m[1]);
|
||||
const total = Number(m[2]);
|
||||
if (!Number.isInteger(index) || !Number.isInteger(total)) return null;
|
||||
if (total < 1 || index < 1 || index > total) return null;
|
||||
return { index, total };
|
||||
}
|
||||
|
||||
/** True iff `specs` is exactly one complete shard set: same N, numerators 1..N. */
|
||||
function isCompleteShardSet(specs) {
|
||||
if (specs.length === 0) return false;
|
||||
const parsed = specs.map(parseShardSpec);
|
||||
if (parsed.some((p) => p === null)) return false;
|
||||
const total = parsed[0].total;
|
||||
if (parsed.some((p) => p.total !== total)) return false;
|
||||
if (parsed.length !== total) return false;
|
||||
const seen = new Set(parsed.map((p) => p.index));
|
||||
return seen.size === total && [...seen].every((i) => i >= 1 && i <= total);
|
||||
}
|
||||
|
||||
test('the full test lane is sharded and complete (#2952)', async (t) => {
|
||||
const workflow = loadWorkflow('test.yml');
|
||||
const include = workflow.jobs.test.strategy.matrix.include;
|
||||
const fullLanes = include.filter((e) => e.scope === 'full');
|
||||
|
||||
await t.test('the full lane is actually sharded, not a single runner', () => {
|
||||
assert.ok(fullLanes.length > 0, 'expected at least one `scope: full` matrix entry');
|
||||
assert.ok(
|
||||
fullLanes.length > 1,
|
||||
'the `scope: full` lane is back to a single unsharded entry. That is the '
|
||||
+ '#2952 regression: the whole unit suite under c8 on one runner grew past '
|
||||
+ 'its cap and reddened `next`.',
|
||||
);
|
||||
for (const lane of fullLanes) {
|
||||
assert.ok(
|
||||
lane.shard !== undefined,
|
||||
`a \`scope: full\` matrix entry declares no shard: ${JSON.stringify(lane)}`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
await t.test('the declared shards form one complete set', () => {
|
||||
const specs = fullLanes.map((e) => e.shard);
|
||||
assert.ok(
|
||||
isCompleteShardSet(specs),
|
||||
`the full lane's shards ${JSON.stringify(specs)} are not a complete set. `
|
||||
+ 'Every entry must share one denominator N and the numerators must be '
|
||||
+ 'exactly 1..N — a missing numerator silently stops running that slice of '
|
||||
+ 'the unit suite while every check stays green.',
|
||||
);
|
||||
});
|
||||
|
||||
await t.test('only the full lane is sharded', () => {
|
||||
for (const lane of include.filter((e) => e.scope !== 'full')) {
|
||||
assert.equal(
|
||||
lane.shard, undefined,
|
||||
`non-full lane ${JSON.stringify(lane)} declares a shard; the scoped and `
|
||||
+ 'Windows lanes run a selected file list, not a partition.',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
await t.test('each shard runs its own slice, not the whole suite', () => {
|
||||
const unitStep = workflow.jobs.test.steps.find(
|
||||
(s) => typeof s.run === 'string' && s.run.includes('test:coverage:unit:raw'),
|
||||
);
|
||||
assert.ok(unitStep, 'no step in the test job runs the raw unit coverage script');
|
||||
assert.match(
|
||||
unitStep.run, /--shard \$\{\{ matrix\.shard \}\}/,
|
||||
'the unit step does not pass matrix.shard through to run-tests.cjs, so '
|
||||
+ 'every shard would run the ENTIRE suite — N times the cost, no speedup.',
|
||||
);
|
||||
});
|
||||
|
||||
await t.test('the aux suites are pinned to one shard that actually exists', () => {
|
||||
// Pinning to a LIVE shard is the whole assertion. A pin to a shard the
|
||||
// matrix no longer declares — say the count moves to 4 and these `if:`
|
||||
// conditions keep naming 1/3 — means the aux suites stop running entirely
|
||||
// while every check stays green. Checking only that some shard literal is
|
||||
// present would not catch that, so the literal is resolved against the
|
||||
// shards the matrix actually declares.
|
||||
const declared = fullLanes.map((e) => String(e.shard));
|
||||
const auxScripts = ['test:integration', 'test:security', 'test:install', 'test:slow'];
|
||||
const pins = new Set();
|
||||
|
||||
for (const script of auxScripts) {
|
||||
const step = workflow.jobs.test.steps.find(
|
||||
(s) => typeof s.run === 'string' && s.run.includes(script),
|
||||
);
|
||||
assert.ok(step, `no step runs \`npm run ${script}\``);
|
||||
|
||||
const pin = /matrix\.shard == '([^']+)'/.exec(String(step.if));
|
||||
assert.ok(
|
||||
pin,
|
||||
`the \`${script}\` step is not pinned to a single shard; it would run `
|
||||
+ 'once per shard and multiply its cost for no extra signal.',
|
||||
);
|
||||
assert.ok(
|
||||
declared.includes(pin[1]),
|
||||
`the \`${script}\` step is pinned to shard '${pin[1]}', which the matrix `
|
||||
+ `does not declare (${JSON.stringify(declared)}). That condition can `
|
||||
+ 'never be true, so this suite would silently never run.',
|
||||
);
|
||||
pins.add(pin[1]);
|
||||
}
|
||||
|
||||
assert.equal(
|
||||
pins.size, 1,
|
||||
`the aux suites are split across shards ${JSON.stringify([...pins])}; they `
|
||||
+ 'are meant to run together on exactly one.',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
test('the merged coverage gate survives sharding (#2952)', async (t) => {
|
||||
const workflow = loadWorkflow('test.yml');
|
||||
|
||||
await t.test('a dedicated coverage-gate job exists and consumes the shards', () => {
|
||||
const gate = workflow.jobs['coverage-gate'];
|
||||
assert.ok(gate, '.github/workflows/test.yml declares no `coverage-gate` job');
|
||||
assert.ok(
|
||||
(gate.needs || []).includes('test'),
|
||||
'coverage-gate must depend on `test` — it merges that job\'s shard artifacts',
|
||||
);
|
||||
|
||||
const merges = gate.steps.some(
|
||||
(s) => String(s.uses || '').includes('download-artifact')
|
||||
&& s.with && s.with['merge-multiple'] === true,
|
||||
);
|
||||
assert.ok(
|
||||
merges,
|
||||
'coverage-gate does not download the shard artifacts with merge-multiple. '
|
||||
+ 'Without every shard merged into one coverage/tmp, the gate scores a '
|
||||
+ 'partial run: files no shard in hand executed read 0%.',
|
||||
);
|
||||
});
|
||||
|
||||
await t.test('both coverage thresholds are still enforced', () => {
|
||||
const runs = workflow.jobs['coverage-gate'].steps
|
||||
.map((s) => s.run).filter((r) => typeof r === 'string').join('\n');
|
||||
const pkg = JSON.parse(
|
||||
fs.readFileSync(path.join(__dirname, '..', 'package.json'), 'utf8'),
|
||||
);
|
||||
|
||||
assert.match(
|
||||
runs, /test:coverage:report/,
|
||||
'coverage-gate never runs the coverage report+gate script — the >=70% '
|
||||
+ 'lines / >=60% branches gate on gsd-core/bin/lib would be gone',
|
||||
);
|
||||
assert.match(
|
||||
runs, /test:coverage:scripts-floor/,
|
||||
'coverage-gate never enforces the >=55% scripts/ floor',
|
||||
);
|
||||
|
||||
// The workflow calls npm scripts precisely so the thresholds are defined
|
||||
// once. If a future edit inlines c8 into the YAML, the two surfaces can
|
||||
// drift silently — the gate would still be green while measuring something
|
||||
// other than what package.json says.
|
||||
assert.match(
|
||||
pkg.scripts['test:coverage:report'], /check-coverage-gate\.cjs/,
|
||||
'test:coverage:report no longer runs check-coverage-gate.cjs',
|
||||
);
|
||||
assert.match(
|
||||
pkg.scripts['test:coverage:scripts-floor'], /--lines 55/,
|
||||
'test:coverage:scripts-floor no longer enforces 55%',
|
||||
);
|
||||
assert.match(
|
||||
pkg.scripts['test:coverage:unit:raw'], /--suite unit/,
|
||||
'test:coverage:unit:raw no longer runs the unit suite',
|
||||
);
|
||||
assert.doesNotMatch(
|
||||
pkg.scripts['test:coverage:unit:raw'], /--shard/,
|
||||
'the shard must come from the workflow matrix, not be baked into the script',
|
||||
);
|
||||
});
|
||||
|
||||
await t.test('required-tests fails when the coverage gate fails', () => {
|
||||
const required = workflow.jobs['required-tests'];
|
||||
assert.ok(
|
||||
(required.needs || []).includes('coverage-gate'),
|
||||
'required-tests does not depend on coverage-gate, so a red gate could not '
|
||||
+ 'block a merge',
|
||||
);
|
||||
|
||||
const summarize = required.steps[0];
|
||||
assert.ok(
|
||||
'COVERAGE_GATE_RESULT' in (summarize.env || {}),
|
||||
'required-tests does not read needs.coverage-gate.result',
|
||||
);
|
||||
assert.match(
|
||||
summarize.run, /coverage-gate did not pass/,
|
||||
'required-tests never fails on a red coverage-gate — depending on a job '
|
||||
+ 'without checking its result makes the dependency decorative',
|
||||
);
|
||||
});
|
||||
|
||||
await t.test('a skipped coverage gate is not treated as a failure', () => {
|
||||
// coverage-gate is conditioned on product_changed, exactly like the test
|
||||
// lane. A docs-only PR skips both; treating `skipped` as red would block
|
||||
// every one of them.
|
||||
assert.match(
|
||||
workflow.jobs['required-tests'].steps[0].run,
|
||||
/COVERAGE_GATE_RESULT" != "skipped"/,
|
||||
'required-tests treats a skipped coverage-gate as a failure',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
test('shard-spec parsing is exact at its boundaries (#2952)', () => {
|
||||
assert.deepEqual(parseShardSpec('1/3'), { index: 1, total: 3 });
|
||||
assert.deepEqual(parseShardSpec('3/3'), { index: 3, total: 3 });
|
||||
assert.deepEqual(parseShardSpec('1/1'), { index: 1, total: 1 });
|
||||
|
||||
// index 0 is below the range, index total+1 above it.
|
||||
assert.equal(parseShardSpec('0/3'), null);
|
||||
assert.equal(parseShardSpec('4/3'), null);
|
||||
assert.equal(parseShardSpec('1/0'), null);
|
||||
|
||||
assert.equal(parseShardSpec('1'), null);
|
||||
assert.equal(parseShardSpec('a/3'), null);
|
||||
assert.equal(parseShardSpec(''), null);
|
||||
assert.equal(parseShardSpec(undefined), null);
|
||||
|
||||
// A complete set, and the three ways one stops being complete.
|
||||
assert.equal(isCompleteShardSet(['1/3', '2/3', '3/3']), true);
|
||||
assert.equal(isCompleteShardSet(['1/3', '2/3']), false, 'missing numerator');
|
||||
assert.equal(isCompleteShardSet(['1/3', '2/3', '2/3']), false, 'duplicate numerator');
|
||||
assert.equal(isCompleteShardSet(['1/3', '2/3', '3/4']), false, 'mixed denominator');
|
||||
assert.equal(isCompleteShardSet([]), false, 'empty set');
|
||||
});
|
||||
157
tests/ci-test-job-timeout-budget.test.cjs
Normal file
157
tests/ci-test-job-timeout-budget.test.cjs
Normal file
@@ -0,0 +1,157 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* CI job timeout budgets — .github/workflows/test.yml (#2952).
|
||||
*
|
||||
* A GitHub Actions job that exceeds its `timeout-minutes` is reported
|
||||
* `cancelled`, not `failed`. That conclusion propagates into `Required tests`
|
||||
* and reddens the branch, while reading like someone hit the cancel button —
|
||||
* which is what makes this failure mode expensive to diagnose and worth a gate.
|
||||
*
|
||||
* It has now happened three times in this repo: #1051 and #1212 on the Windows
|
||||
* full-test lane, and #2952 on the unsharded `test` lane, whose
|
||||
* `ubuntu-latest / 24` entry is the only `scope: full` matrix entry — it runs
|
||||
* the whole unit suite under c8 coverage, then the scripts/ coverage floor,
|
||||
* integration, security, install and slow, serially on one runner. Every one of
|
||||
* those was the same root cause: a budget sized to what the lane cost that
|
||||
* week, with no headroom for the suite to grow into.
|
||||
*
|
||||
* So the rule enforced here is not a fixed number per lane — it is a HEADROOM
|
||||
* FACTOR over each lane's measured cost. A lane may be slow; what it may not be
|
||||
* is budgeted to finish with seconds to spare.
|
||||
*
|
||||
* This is a budget assertion, not a duration assertion. No unit test can prove
|
||||
* a lane still FITS its budget — only a real CI run measures that. What this
|
||||
* file guarantees is that a budget cannot be quietly lowered back beneath what
|
||||
* its lane is already known to need.
|
||||
*/
|
||||
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const yaml = require('js-yaml');
|
||||
|
||||
const WORKFLOWS_DIR = path.join(__dirname, '..', '.github', 'workflows');
|
||||
|
||||
function loadWorkflow(name) {
|
||||
return yaml.load(fs.readFileSync(path.join(WORKFLOWS_DIR, name), 'utf8'));
|
||||
}
|
||||
|
||||
/**
|
||||
* Multiplier applied to a lane's measured cost to get its required budget.
|
||||
* 1.5x is enough slack for ordinary suite growth across a release cycle without
|
||||
* letting a genuinely runaway lane hide behind a large number.
|
||||
*/
|
||||
const HEADROOM_FACTOR = 1.5;
|
||||
|
||||
/**
|
||||
* Measured wall-clock cost per lane, in whole minutes rounded UP, each from a
|
||||
* named run. Raise an entry only alongside a fresh measurement — never to make
|
||||
* a red gate green. Raising a measurement raises the required budget with it,
|
||||
* which is the point: a lane that got slower must be re-budgeted, not excused.
|
||||
*/
|
||||
const LANE_COSTS = [
|
||||
{
|
||||
job: 'test',
|
||||
measuredMinutes: 8,
|
||||
// Sharded three ways as of #2952, so this is ONE shard's cost, not the
|
||||
// whole unit suite. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s.
|
||||
// Shard 1 is the long pole because the unsharded aux suites ride on it.
|
||||
// Before sharding the same lane cost 15m20s and blew a 15-minute cap.
|
||||
evidence: 'run 30677442953 — 7m12s slowest shard',
|
||||
},
|
||||
{
|
||||
job: 'test-full',
|
||||
measuredMinutes: 19,
|
||||
// Worst observed shard is `full test (windows-latest, 22, shard 3/3)`:
|
||||
// 18m59s on 05b170e44 and 18m14s on 81eeb8a53. The Windows shards are slow
|
||||
// for platform reasons, not extra work.
|
||||
evidence: 'run 30650559192 — 18m59s, windows-22 shard 3/3',
|
||||
},
|
||||
{
|
||||
job: 'coverage-gate',
|
||||
measuredMinutes: 2,
|
||||
// Downloads three shards' raw V8 dumps, renders one merged report and runs
|
||||
// both thresholds. Run 30677442953: 1m20s end to end, most of it npm ci.
|
||||
evidence: 'run 30677442953 — 1m20s',
|
||||
},
|
||||
{
|
||||
job: 'test-inert',
|
||||
measuredMinutes: 2,
|
||||
// Runs only the targeted-test step when no product code changed; observed
|
||||
// around a minute. Listed so its budget cannot be dropped to nothing.
|
||||
evidence: 'targeted-only lane, ~1m observed',
|
||||
},
|
||||
];
|
||||
|
||||
function requiredBudgetMinutes(measuredMinutes, headroomFactor = HEADROOM_FACTOR) {
|
||||
return Math.ceil(measuredMinutes * headroomFactor);
|
||||
}
|
||||
|
||||
function hasSufficientBudget(budgetMinutes, measuredMinutes, headroomFactor = HEADROOM_FACTOR) {
|
||||
return Number.isInteger(budgetMinutes)
|
||||
&& budgetMinutes >= requiredBudgetMinutes(measuredMinutes, headroomFactor);
|
||||
}
|
||||
|
||||
test('CI job timeout budgets carry headroom over measured cost (#2952)', async (t) => {
|
||||
const workflow = loadWorkflow('test.yml');
|
||||
|
||||
for (const lane of LANE_COSTS) {
|
||||
await t.test(`${lane.job} is budgeted above its measured cost`, () => {
|
||||
assert.ok(
|
||||
workflow.jobs && Object.prototype.hasOwnProperty.call(workflow.jobs, lane.job),
|
||||
`.github/workflows/test.yml declares no job \`${lane.job}\`. If it was `
|
||||
+ 'renamed or removed, update LANE_COSTS in this file to match — do not '
|
||||
+ 'delete the entry to make this pass.',
|
||||
);
|
||||
|
||||
const budget = workflow.jobs[lane.job]['timeout-minutes'];
|
||||
const required = requiredBudgetMinutes(lane.measuredMinutes);
|
||||
|
||||
assert.equal(
|
||||
typeof budget, 'number',
|
||||
`.github/workflows/test.yml jobs.${lane.job} must declare timeout-minutes`,
|
||||
);
|
||||
assert.ok(
|
||||
hasSufficientBudget(budget, lane.measuredMinutes),
|
||||
`jobs.${lane.job}.timeout-minutes is ${budget}, but the lane measured `
|
||||
+ `${lane.measuredMinutes}m (${lane.evidence}) and needs at least `
|
||||
+ `${required} — ${HEADROOM_FACTOR}x — so suite growth does not breach `
|
||||
+ 'the cap. A job that exceeds timeout-minutes is reported `cancelled` '
|
||||
+ 'and reddens `Required tests`.',
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
// Boundary coverage on the predicate that decides every lane above:
|
||||
// required-1 must be rejected, required and required+1 accepted.
|
||||
await t.test('budget sufficiency is exact at the boundary', () => {
|
||||
for (const lane of LANE_COSTS) {
|
||||
const required = requiredBudgetMinutes(lane.measuredMinutes);
|
||||
|
||||
assert.equal(hasSufficientBudget(required - 1, lane.measuredMinutes), false,
|
||||
`${lane.job}: a budget one minute under the requirement must be rejected`);
|
||||
assert.equal(hasSufficientBudget(required, lane.measuredMinutes), true,
|
||||
`${lane.job}: a budget exactly at the requirement must be accepted`);
|
||||
assert.equal(hasSufficientBudget(required + 1, lane.measuredMinutes), true,
|
||||
`${lane.job}: a budget over the requirement must be accepted`);
|
||||
}
|
||||
});
|
||||
|
||||
await t.test('a non-integer budget is not a sufficient budget', () => {
|
||||
// `timeout-minutes: 24.5` is not something GitHub accepts; treating it as
|
||||
// sufficient would let a malformed workflow through this gate.
|
||||
assert.equal(hasSufficientBudget(24.5, 16), false);
|
||||
assert.equal(hasSufficientBudget(Number.NaN, 16), false);
|
||||
assert.equal(hasSufficientBudget(undefined, 16), false);
|
||||
});
|
||||
|
||||
await t.test('requiredBudgetMinutes rounds up rather than truncating', () => {
|
||||
// 15 * 1.5 = 22.5 — truncation would hand back 22 and under-budget the lane.
|
||||
assert.equal(requiredBudgetMinutes(15, 1.5), 23);
|
||||
assert.equal(requiredBudgetMinutes(16, 1.5), 24);
|
||||
assert.equal(requiredBudgetMinutes(19, 1.5), 29);
|
||||
assert.equal(requiredBudgetMinutes(10, 1.5), 15);
|
||||
});
|
||||
});
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user