From 5039d499248af08519c248d44fcdfb0dfa59aeb9 Mon Sep 17 00:00:00 2001 From: sim Date: Wed, 5 Aug 2026 23:22:59 -0400 Subject: [PATCH] ci(#3057): shard the scoped Windows lane, the last unsharded one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The scoped Windows lane reached exactly 15m05s and was cancelled on four consecutive shas of PR #3094. A job that exceeds timeout-minutes reports as CANCELLED rather than FAILURE, which is why it first read as infrastructure noise; the giveaway is that the duration equals the cap. The Required tests rollup fans that job in, so it red-blocked merge while every other lane — including all three sharded full-Windows shards — was green. The trigger was a change to the shared test helper, which scopes the install-heavy suites into the selected list. The lane normally runs about eight minutes; with that list it does not fit. It was the only unsharded lane left in this job, so it was the only one without headroom to absorb a large scoped list. Issue #869 hit this exact cliff on the sibling lane and named the durable answer in its own follow-up: a timeout bump moves the cliff, sharding removes it. #2952 then sharded the full lane. This finishes that work. The runner already supports it — the shard partition is applied after scope selection, so it composes with a selected file list rather than only with a suite, and the partition is cost-weighted from the measured timings table. The job name template already renders a shard suffix when one is present, so the three entries name themselves. No individual matrix job is a required status check; the rollup is, and it is name-independent, so renaming these jobs does not touch branch protection. timeout-minutes stays at 15. Each shard now does roughly a third of the work, so the cap goes from binding to backstop without being raised. The lane-shape tests were generalized rather than relaxed: the complete-shard-set invariant now runs per sharded scope instead of only over the full lane, and "only the full lane is sharded" became "targeted is the only unsharded lane". A new assertion pins the shard through to the runner — without it the three shards would each run the entire selected list, triple the cost and no speedup, and every check would stay green. No LANE_COSTS entry is added for the new shards. The only recorded cost for that lane is the pre-sharding run that hit the cap, and inventing a post-sharding number would be exactly the kind of unmeasured claim the rest of that table avoids. The estimate and the reason are written down instead, to be replaced by a real measurement. Refs #3057 Co-Authored-By: Claude Opus 5 --- .github/workflows/test.yml | 39 ++++++++-- tests/ci-full-lane-sharding.test.cjs | 87 +++++++++++++++-------- tests/ci-test-job-timeout-budget.test.cjs | 12 ++++ 3 files changed, 101 insertions(+), 37 deletions(-) diff --git a/.github/workflows/test.yml b/.github/workflows/test.yml index f716b8ebb..24d8e8912 100644 --- a/.github/workflows/test.yml +++ b/.github/workflows/test.yml @@ -116,16 +116,25 @@ jobs: needs: changes if: needs.changes.outputs.product_changed == 'true' runs-on: ${{ matrix.os }} - # #2952: this lane is sharded three ways (see the matrix below), so the - # budget covers ONE shard, not the whole unit suite. Measured on run - # 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s. Shard 1 is the long - # pole because the unsharded aux suites (integration/security/install/slow) - # ride along on it — that is deliberate; they total ~1m35s and sharding them - # would cost more than it saves. + # #2952: the `scope: full` lane is sharded three ways (see the matrix + # below), so the budget covers ONE shard, not the whole unit suite. + # Measured on run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s. Shard + # 1 is the long pole because the unsharded aux suites + # (integration/security/install/slow) ride along on it — that is + # deliberate; they total ~1m35s and sharding them would cost more than it + # saves. # # 15 is ~1.9x the slowest measured shard. Before sharding this same lane ran # 15m20s against a 15-minute cap and was killed mid-run, which is the whole # of #2952 — the number is unchanged, the work behind it is a third the size. + # + # The `scope: windows` lane is sharded three ways for the same reason + # (see #3057): on PR #3094 it reached 15m05s against this same cap and was + # CANCELLED, four shas in a row — a change to tests/helpers.cjs scoped in + # the install-heavy suites and pushed the single Windows lane over the top. + # Per #869, a timeout bump only moves that cliff; sharding removes it. The + # cap stays at 15 unchanged: each shard now does roughly a third of the + # work, so headroom improves rather than needing a raise. # tests/ci-test-job-timeout-budget.test.cjs holds every lane here to a # headroom factor over its own measured cost. timeout-minutes: 15 @@ -158,6 +167,13 @@ jobs: # gave the `test-full` lane, which measures 0.0% spread across 3 bins on # the current table. The aux suites (integration/security/install/slow) # total ~1m35s and are not worth sharding; they run on shard 1 only. + # + # #3057: the `scope: windows` lane is now sharded three ways too, the + # same fix applied to the same cliff (#869's stated durable follow-up). + # It hit exactly 15m05s and was CANCELLED on PR #3094, four shas + # straight, after a change to tests/helpers.cjs scoped in the + # install-heavy suites. Its shards use the same `--shard i/n` + # flag on scripts/run-tests.cjs, applied AFTER scope selection. - os: ubuntu-latest node-version: 22 scope: targeted @@ -176,6 +192,15 @@ jobs: - os: windows-latest node-version: 24 scope: windows + shard: 1/3 + - os: windows-latest + node-version: 24 + scope: windows + shard: 2/3 + - os: windows-latest + node-version: 24 + scope: windows + shard: 3/3 steps: # Windows lane on checkout v5.0.1 (drops includeIf; no auth flake, uses Node 24 natively). @@ -255,7 +280,7 @@ jobs: - name: Run scoped tests if: matrix.scope != 'full' - run: node scripts/run-tests.cjs --files-from .ci-selected-tests.txt + run: node scripts/run-tests.cjs --files-from .ci-selected-tests.txt${{ matrix.shard && format(' --shard {0}', matrix.shard) || '' }} # #2952: each shard runs its slice of the unit suite under c8 but renders # NO report and enforces NO gate — it only leaves raw V8 dumps in diff --git a/tests/ci-full-lane-sharding.test.cjs b/tests/ci-full-lane-sharding.test.cjs index 2721d82b6..7b3a425ce 100644 --- a/tests/ci-full-lane-sharding.test.cjs +++ b/tests/ci-full-lane-sharding.test.cjs @@ -1,13 +1,20 @@ 'use strict'; /** - * The full test lane is sharded, and the coverage gate that sharding displaced - * is still wired in — .github/workflows/test.yml (#2952). + * The full test lane and the scoped Windows lane are both sharded, and the + * coverage gate that sharding displaced is still wired in — + * .github/workflows/test.yml (#2952, #3057). * * The `scope: full` lane was the only unsharded lane in this file. It ran the * entire unit suite under c8 on a single runner, grew past a 15-minute cap, and * reddened `next` (#2952). Raising the cap treated the symptom; sharding is the - * shape fix, and it is the same answer #1212 reached for the Windows lane. + * shape fix, and it is the same answer #1212 reached for the Windows lane at + * the time. + * + * The `scope: windows` lane then hit the identical cliff itself: it reached + * exactly 15m05s and was CANCELLED on PR #3094, four shas in a row. Per #869's + * stated durable follow-up, it is now sharded three ways too (#3057), leaving + * `scope: targeted` as the only lane in this job with no shard. * * Sharding introduces two failure modes that stay GREEN while being wrong, so * both are pinned here: @@ -69,44 +76,64 @@ test('the full test lane is sharded and complete (#2952)', async (t) => { const workflow = loadWorkflow('test.yml'); const include = workflow.jobs.test.strategy.matrix.include; const fullLanes = include.filter((e) => e.scope === 'full'); + const windowsLanes = include.filter((e) => e.scope === 'windows'); + // The only lane in this job with no shard is `scope: targeted` — the fast, + // single-runner default lane. Both `full` and `windows` are sharded. + const shardedScopes = { full: fullLanes, windows: windowsLanes }; - await t.test('the full lane is actually sharded, not a single runner', () => { - assert.ok(fullLanes.length > 0, 'expected at least one `scope: full` matrix entry'); - assert.ok( - fullLanes.length > 1, - 'the `scope: full` lane is back to a single unsharded entry. That is the ' - + '#2952 regression: the whole unit suite under c8 on one runner grew past ' - + 'its cap and reddened `next`.', - ); - for (const lane of fullLanes) { + for (const [scope, lanes] of Object.entries(shardedScopes)) { + await t.test(`the \`scope: ${scope}\` lane is actually sharded, not a single runner`, () => { + assert.ok(lanes.length > 0, `expected at least one \`scope: ${scope}\` matrix entry`); assert.ok( - lane.shard !== undefined, - `a \`scope: full\` matrix entry declares no shard: ${JSON.stringify(lane)}`, + lanes.length > 1, + `the \`scope: ${scope}\` lane is back to a single unsharded entry. That is ` + + 'the #2952/#3057 regression: the whole suite on one runner grows past ' + + 'its cap and reddens `next`.', ); - } - }); + for (const lane of lanes) { + assert.ok( + lane.shard !== undefined, + `a \`scope: ${scope}\` matrix entry declares no shard: ${JSON.stringify(lane)}`, + ); + } + }); - await t.test('the declared shards form one complete set', () => { - const specs = fullLanes.map((e) => e.shard); - assert.ok( - isCompleteShardSet(specs), - `the full lane's shards ${JSON.stringify(specs)} are not a complete set. ` - + 'Every entry must share one denominator N and the numerators must be ' - + 'exactly 1..N — a missing numerator silently stops running that slice of ' - + 'the unit suite while every check stays green.', - ); - }); + await t.test(`the \`scope: ${scope}\` lane's declared shards form one complete set`, () => { + const specs = lanes.map((e) => e.shard); + assert.ok( + isCompleteShardSet(specs), + `the \`scope: ${scope}\` lane's shards ${JSON.stringify(specs)} are not a ` + + 'complete set. Every entry must share one denominator N and the ' + + 'numerators must be exactly 1..N — a missing numerator silently stops ' + + 'running that slice of the suite while every check stays green.', + ); + }); + } - await t.test('only the full lane is sharded', () => { - for (const lane of include.filter((e) => e.scope !== 'full')) { + await t.test('no other lane is sharded', () => { + for (const lane of include.filter((e) => e.scope !== 'full' && e.scope !== 'windows')) { assert.equal( lane.shard, undefined, - `non-full lane ${JSON.stringify(lane)} declares a shard; the scoped and ` - + 'Windows lanes run a selected file list, not a partition.', + `lane ${JSON.stringify(lane)} declares a shard but is neither \`scope: full\` ` + + 'nor `scope: windows` — the targeted lane runs a selected file list, not ' + + 'a partition.', ); } }); + await t.test('the scoped Windows shards are passed through to run-tests.cjs', () => { + const scopedStep = workflow.jobs.test.steps.find( + (s) => s.name === 'Run scoped tests', + ); + assert.ok(scopedStep, 'no "Run scoped tests" step in the test job'); + assert.match( + scopedStep.run, /matrix\.shard/, + 'the "Run scoped tests" step does not reference matrix.shard, so the ' + + 'windows lane\'s three shards would each run the entire selected file ' + + 'list — N times the cost, no speedup.', + ); + }); + await t.test('each shard runs its own slice, not the whole suite', () => { const unitStep = workflow.jobs.test.steps.find( (s) => typeof s.run === 'string' && s.run.includes('test:coverage:unit:raw'), diff --git a/tests/ci-test-job-timeout-budget.test.cjs b/tests/ci-test-job-timeout-budget.test.cjs index 4c52de419..cedce1415 100644 --- a/tests/ci-test-job-timeout-budget.test.cjs +++ b/tests/ci-test-job-timeout-budget.test.cjs @@ -59,6 +59,18 @@ const LANE_COSTS = [ // whole unit suite. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s. // Shard 1 is the long pole because the unsharded aux suites ride on it. // Before sharding the same lane cost 15m20s and blew a 15-minute cap. + // + // This one `timeout-minutes` also covers the `scope: windows` matrix + // entries — GitHub applies a single job-level budget across every matrix + // combination, not one per entry. That lane is now sharded three ways too + // (#3057), but no post-sharding per-shard measurement exists yet: its only + // recorded cost is the PRE-sharding whole-suite run that hit 15m05s and was + // CANCELLED on PR #3094. Each of its three shards should now cost roughly a + // third of that (~5m), which is already comfortably under the 8m/12m this + // entry requires — so no separate LANE_COSTS entry is added on a number + // that has not actually been measured. Replace this estimate with a real + // measured shard cost once one exists, the same discipline every other + // entry here follows. evidence: 'run 30677442953 — 7m12s slowest shard', }, {