chore(#4603): retire the test-full CI job (#4604)

* chore(#4603): retire the test-full CI job

Phase 2 (#4591) added test-conformance but left test-full (the pre-existing
full-suite Windows/macOS replay) running unchanged, gated on the same
full_matrix flag, downgraded only from a hard gate to a non-blocking
::warning:: -- framed as "a non-gating safety net for one release cycle."
No phase or issue ever retired it. Result: every full_matrix=true PR ran
10 OS-specific jobs (test-full's 6 + test-conformance's 4, purely
additive) instead of the original 6 -- the epic's own goal (reduce
runner-minutes) was measurably regressing, not improving, for the
majority of PRs.

This phase was missing from the original 4-phase epic decomposition; the
epic (#4589) has been amended to add it as Phase 5 (see its comment
thread), and this issue was filed as the tracked sub-issue.

Deletes the test-full job from .github/workflows/test.yml entirely, along
with every reference to it: required-tests' needs/FULL_TEST_RESULT
warning branch, ci-timeout-report.cjs's JOB_RULES entry,
ci-test-job-timeout-budget.test.cjs's LANE_COSTS/staticLanes/testFullRule
entries, ci-test-scope.test.cjs's test-full-specific tests (preserving
three unrelated tests that were nested in the same describe block, moved
under a renamed describe rather than deleted), and docs mentions.
test-conformance is now the sole gating signal for real-OS coverage.

Two separate defects found and fixed while auditing every test-full
reference:
- tests/ci-pr-mergeability.test.cjs's GATED['test.yml'] safety-critical
  array (jobs that must needs: the mergeability preflight) had test-full
  but was missing test-conformance entirely -- Phase 2 never added it.
  Verified the real workflow wiring was already correct (test-conformance
  does have needs: [changes, preflight]); this was a test-coverage gap,
  not a live defect. Fixed by swapping the array entry.
- docs/TESTING-SUITES.md's "## CI matrix" section was substantially stale
  independent of this phase (predating even #2952's coverage-gate split).
  Rewritten against the real, current job topology, verified directly
  against test.yml rather than trusted from memory.

An isolated code-review pass found and fixed two minor inaccuracies in the
rewritten docs table (two jobs' "Gated on" column didn't match their real
if: condition exactly). An isolated security-review pass found no
qualifying findings -- every compute-provisioning job already carries
needs: preflight directly, unaffected by this deletion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(ci): isolate 7 more heavy test files from chunk-weight packing

`next`'s own push-triggered Tests run failed: `conformance test
(windows-latest, 24, shard 2/3)` chunk 3/6 was killed after 600019ms.
Root cause: state.test.cjs (weight 21.35, measured) was packed alongside
companions by run-tests.cjs's LPT chunk packer, the same failure mode
that previously hit codex-config.test.cjs (weight 17.87) twice and got a
dedicated fix (ISOLATED_HEAVY_FILES, #4497) -- but state.test.cjs was
never added to that set.

This is a direct, unintended consequence of epic #4589 Phase 2: the new
platform-conformance-tier job packs only ~546 files per shard (vs. the
~950-file full suite the packer used to balance against), so the same
absolute-weight outlier now represents a larger share of a smaller, more
homogeneous pool -- the LPT packer has fewer light files to pad around
it with. This was a real, foreseeable side effect of shrinking the
packing pool that nobody checked for when Phase 2 shipped.

A first attempt at this fix hand-picked 4 candidates by eyeballing a
truncated weight list and missed 3 heavier ones -- caught by an isolated
code-review pass (blocker: emitted-attribution.test.cjs at 66.2% of the
Windows chunk budget, install-minimal-hooks.test.cjs at 61.1%,
install.test.cjs at 47.1%, all above codex-config.test.cjs's own
44.7% -- the ratio that already proved dangerous twice). Corrected by
systematically computing weight/budget for every unit-suite file and
isolating everything at or above that same ratio: 7 files total, plus
the pre-existing codex-config.test.cjs (8 total).

Added a durable regression test (tests/run-tests-harness.test.cjs) that
re-derives this exact computation from the live tests/test-timings.json
on every run, so a future heavy file crossing this threshold fails the
test instead of silently reintroducing this failure -- not just a
one-time manual sweep.

Verified end-to-end: simulated the real 3-way windows shard split of the
actual conformance-tier file list with the real packing functions. Max
packable-chunk weight across all 3 shards is now 27.04 / 24.10 / 23.91
(shard 2 is the exact shard that failed on next), comfortably under the
40 budget -- versus 40+ and a 600s kill before this fix.

A second isolated code-review + security-review pass on the corrected
diff found nothing further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-09-10 12:46:04 -04:00
committed by GitHub
parent bcd99696d3
commit 181c4c8659
10 changed files with 204 additions and 380 deletions

View File

@@ -324,8 +324,8 @@ jobs:
# runner grew until it blew a 15-minute cap and reddened `next`. Raising
# the cap treated the symptom; sharding changes the shape. Shards are
# partitioned by MEASURED per-file duration (tests/test-timings.json)
# using LPT in scripts/run-tests.cjs — the same cost-aware packer #2472
# gave the `test-full` lane, which measures 0.0% spread across 3 bins on
# using LPT in scripts/run-tests.cjs — the #2472 cost-aware packer,
# which measures 0.0% spread across 3 bins on
# the current table. The aux suites (integration/security/install/slow)
# run on shard 1 only — sharding them too was evaluated and rejected for
# #4070 (see tests/run-tests-harness.test.cjs's "selectShard
@@ -595,210 +595,20 @@ jobs:
- name: Run scoped tests
run: node scripts/run-tests.cjs --files-from .ci-selected-tests.txt
test-full:
name: full test (${{ matrix.os }}, ${{ matrix.node-version }}, shard ${{ matrix.shard }}/3)
needs: [changes, preflight]
if: needs.changes.outputs.code_changed == 'true' && needs.changes.outputs.full_matrix == 'true'
runs-on: ${{ matrix.os }}
defaults:
run:
shell: ${{ matrix.shell }}
# The unit suite is sharded across 3 parallel runners per OS/node leg
# (#1212, cost-weighted in #2472). Each shard runs a deterministic
# cost-balanced third of the sorted unit-file list via
# `run-tests.cjs --suite unit --shard i/3`, so per-job
# wall-clock scales as O(total/3) and stays well under the cap as the suite
# grows — replacing the #869 timeout bump (15→20m) which only deferred the
# cliff. The cap stays at 20m as a generous backstop; a healthy shard now
# finishes in roughly a third of the old single-lane wall-clock.
#
# The matrix is the cross-product of 2 OS/node legs × 3 shards = 6 jobs
# (windows-latest/24, macos-latest/24 — the Node floor is 24, so there is
# no separate node-22 leg to cross-product against), enumerated explicitly
# as `include:` rows. (A base `shard: [1,2,3]`
# dimension would NOT cross-product against `include` legs — include rows
# sharing no key with the base matrix are appended as standalone combos —
# and a NESTED `leg.os` key is not resolvable by the H1 shell-policy linter
# in scripts/workflow-policy.cjs, which reads `matrix.os`/`matrix.shell`
# directly. Explicit rows keep both the cross-product and the linter happy.)
# #2952: `full test (windows-latest, 22, shard 3/3)` reached 18m59s (94% of
# a 20-minute cap) on 05b170e44 and 18m14s (91%) on 81eeb8a53. The Windows
# shards are slow for platform reasons — process spawn and filesystem cost,
# not extra work — and this lane has already blown its cap twice before
# (#1051, #1212).
# #3787: fresh measurement on windows-latest/24 shard 3/3 hit 26m18s (run
# 32614439702), so the 18m59s figure above is stale and the 30-minute cap
# only had ~1.14x headroom, in violation of this repo's own 1.5x rule
# (tests/ci-test-job-timeout-budget.test.cjs). 1.5x of 27m requires 41m
# minimum; 45 is used instead of the bare minimum because shard
# composition is unstable — adding one test file reshuffled 115 of 268
# unit files between shards — so the per-shard worst case moves run to
# run and a budget pinned to the exact minimum would be re-breached by
# the next file anyone adds.
timeout-minutes: 45
env:
GSD_PLUGIN_ROOT: .ci-gsd-plugin-root-disabled
# #2665: strict on Linux/macOS, report-only on Windows (see the `test` job note).
GSD_STRICT_LIVE_CONFIG_GUARD: ${{ matrix.os != 'windows-latest' && '1' || '' }}
# #2854: pin the emitted gate's baseline to the SAME commit the tree was merged
# with. "Rebase check" merges `pull_request.base.sha` (pinned by #2472 so all 12
# matrix jobs agree on one tree), but `resolveBase()` otherwise falls through to
# `origin/next`, which `fetch-depth: 0` leaves at the LIVE tip. Whenever `next`
# advanced mid-flight the gate compared a tree built on base.sha against a
# baseline at a newer commit — so the correctly-keyed cache was rejected as
# "stale" and the run hard-failed on diffs that touched nothing related.
# This must stay equal to CI_REBASE_BASE_SHA; a test asserts that parity.
GSD_EMITTED_BASE: ${{ github.event.pull_request.base.sha }}
# #4196: pin the npm-audit baseline the SAME way GSD_EMITTED_BASE pins
# its own baseline (see the comment above) -- origin/next is live under
# fetch-depth: 0 and can advance mid-run; base.sha is fixed for the life
# of the run. For a push event, github.event.before is git's own record
# of the ref's state immediately before this push landed -- correct
# even when a rebase-merged PR lands as multiple discrete commits in
# one push (HEAD~1 would be wrong there: it could already contain an
# earlier commit's newly-introduced vulnerable package, masking it).
# #4241: a merge_group event carries no pull_request/push context, so
# without this arm it silently fell through to the '' branch --
# resolveBaselineRef()'s documented-unreachable origin/next live-tip
# fallback (npm-audit-baseline.cjs), reopening the exact race #4196
# fixed, but only for merge-queue runs. github.event.merge_group.base_sha
# is "the SHA of the merge group's parent commit" (GitHub's merge_group
# webhook payload) -- the base tip the temporary merge-group commit was
# built against, pinned for the life of the run same as the other two arms.
AUDIT_BASELINE_REF: ${{ github.event_name == 'pull_request' && github.event.pull_request.base.sha || (github.event_name == 'push' && github.event.before) || (github.event_name == 'merge_group' && github.event.merge_group.base_sha) || '' }}
strategy:
fail-fast: false
matrix:
include:
- os: windows-latest
node-version: 24
shell: pwsh
shard: 1
- os: windows-latest
node-version: 24
shell: pwsh
shard: 2
- os: windows-latest
node-version: 24
shell: pwsh
shard: 3
- os: macos-latest
node-version: 24
shell: 'zsh {0}'
shard: 1
- os: macos-latest
node-version: 24
shell: 'zsh {0}'
shard: 2
- os: macos-latest
node-version: 24
shell: 'zsh {0}'
shard: 3
steps:
- uses: actions/checkout@93cb6efe18208431cddfb8368fd83d5badbf9bfd # v5.0.1 (Windows)
if: runner.os == 'Windows'
with:
fetch-depth: 0
persist-credentials: true
token: ${{ github.token }}
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2 (Linux/macOS)
if: runner.os != 'Windows'
with:
fetch-depth: 0
persist-credentials: true
token: ${{ github.token }}
- name: Record job start time (near-cap check)
run: node -e 'require("fs").appendFileSync(process.env.GITHUB_ENV, "CI_JOB_START_EPOCH_MS=" + Date.now() + "\n")'
- name: Guard — require GitHub-hosted runner
run: node scripts/ci-guard-runner.cjs
- name: Rebase check — merge PR base branch into PR head
if: github.event_name == 'pull_request'
env:
GITHUB_TOKEN: ${{ github.token }}
# Pin every job of this run to ONE base commit (#2472). Each job runs
# this step independently, minutes apart across a 12-job matrix, so
# merging the moving branch ref lets jobs see different trees when the
# base advances mid-run. The sharded lane needs all jobs to agree on a
# partition, and disagreement there drops a test file silently while
# CI stays green. base.sha is fixed for the life of the run.
CI_REBASE_BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: node scripts/ci-rebase-check.cjs
- name: Set up Node.js ${{ matrix.node-version }}
uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
with:
node-version: ${{ matrix.node-version }}
cache: 'npm'
- name: Environment check
run: npm run check:env
- name: Install dependencies
run: npm ci
- name: Dependency integrity gate
run: node scripts/check-npm-integrity.cjs
# #2724 (ADR-2719 §5): restore the differential attribution check's cached
# baseline, keyed on the PR's base sha. A miss degrades to an in-job build
# rather than failing (resolveBaseline()'s documented precedence).
- name: Restore emitted-baseline cache
if: github.event_name == 'pull_request'
uses: actions/cache/restore@0057852bfaa89a56745cba8c7296529d2fc39830 # v4.3.0
with:
path: .gsd-cache/emitted-baseline.json
key: emitted-baseline-${{ github.event.pull_request.base.sha }}
# The heavy unit suite is split across the 3 shards — each runs a
# deterministic cost-balanced third of the sorted unit-file list (#2472). The
# union of shards 1/3 + 2/3 + 3/3 is the full unit suite, so coverage is
# unchanged; only wall-clock per job drops to ~total/3.
- name: Run unit tests (shard ${{ matrix.shard }}/3)
run: node scripts/run-tests.cjs --suite unit --shard ${{ matrix.shard }}/3
# Integration and security suites are small; run them once per OS/node
# leg (on shard 1 only) instead of redundantly on all three shards. They
# still run on every leg (3 times total, once per platform), so each
# platform's integration/security coverage is unchanged — only the
# 3x-per-leg duplication is removed.
- name: Run integration tests
if: matrix.shard == 1
run: npm run test:integration
- name: Run security tests
if: matrix.shard == 1
run: npm run test:security
- name: Check job budget (near-cap advisory)
if: always()
continue-on-error: true
env:
CI_JOB_LABEL: "full test (${{ matrix.os }}, ${{ matrix.node-version }}, shard ${{ matrix.shard }}/3)"
CI_JOB_TIMEOUT_MINUTES: '45'
run: node scripts/ci-check-job-near-cap.cjs
# #4591 (epic #4589 Phase 2): runs ONLY the platform-conformance-tier file
# list (scripts/lib/platform-conformance-tier.generated.cjs) on real
# Windows/macOS — the OS-agnostic bulk of the suite already ran once on
# ubuntu-latest in the `test` job above. Gated the same way `test-full` is
# (product code changed AND full_matrix). Windows is sharded 3 ways
# (matching test-full's own precedent) after the first real CI run measured
# it: unsharded, windows-latest hit its 45-minute timeout and was CANCELLED
# (started 03:41:19Z, cancelled 04:26:25Z, run 34434252144) while
# macos-latest finished the identical file set in 26m58s — this conformance
# tier is dominated by subprocess-spawning tests (343/565 files), and
# Windows process-spawn overhead is the documented reason test-full's own
# windows lane needed the same fix (#3057). macos-latest stays unsharded; it
# has real headroom (27m against the 45m cap).
# `test-full` (above) keeps running the WHOLE suite on the
# same OSes as a non-gating safety net for one release cycle (see
# `required-tests` below) — this job is the new GATING signal for
# windows/macos coverage.
# ubuntu-latest in the `test` job above. Gated on product code changed AND
# full_matrix (#4591); this is the sole gating signal for windows/macos
# coverage (#4603 retired the parallel legacy full-matrix job). Windows is
# sharded 3 ways, for the same reason a prior full-suite windows lane
# needed 3-way sharding (#3057, since retired — #4603), after the first
# real CI run measured it: unsharded, windows-latest hit its 45-minute
# timeout and was CANCELLED (started 03:41:19Z, cancelled 04:26:25Z, run
# 34434252144) while macos-latest finished the identical file set in
# 26m58s — this conformance tier is dominated by subprocess-spawning tests
# (343/565 files). macos-latest stays unsharded; it has real headroom (27m
# against the 45m cap).
test-conformance:
name: conformance test (${{ matrix.os }}, ${{ matrix.node-version }}${{ matrix.shard && format(', shard {0}', matrix.shard) || '' }})
needs: [changes, preflight]
@@ -810,9 +620,8 @@ jobs:
# No LANE_COSTS entry exists yet for this brand-new job — see
# tests/ci-test-job-timeout-budget.test.cjs's own header: "no unit test
# can prove a lane fits its budget — only a real CI run measures that."
# Generous and unchanged from test-full's own budget until a real
# measurement exists; the smaller file count should comfortably undercut
# this ceiling.
# Generous until a real measurement exists; the smaller file count
# should comfortably undercut this ceiling.
timeout-minutes: 45
env:
GSD_PLUGIN_ROOT: .ci-gsd-plugin-root-disabled
@@ -1055,7 +864,6 @@ jobs:
- lint-tests
- test
- test-inert
- test-full
- test-conformance
- coverage-gate
- qa-loop-walk
@@ -1074,7 +882,6 @@ jobs:
TEST_RESULT: ${{ needs.test.result }}
TEST_CONFORMANCE_RESULT: ${{ needs.test-conformance.result }}
INERT_RESULT: ${{ needs.test-inert.result }}
FULL_TEST_RESULT: ${{ needs.test-full.result }}
COVERAGE_GATE_RESULT: ${{ needs.coverage-gate.result }}
QA_LOOP_WALK_RESULT: ${{ needs.qa-loop-walk.result }}
run: |
@@ -1088,7 +895,6 @@ jobs:
echo "test=$TEST_RESULT"
echo "test-conformance=$TEST_CONFORMANCE_RESULT"
echo "test-inert=$INERT_RESULT"
echo "test-full=$FULL_TEST_RESULT"
echo "coverage-gate=$COVERAGE_GATE_RESULT"
echo "qa-loop-walk=$QA_LOOP_WALK_RESULT"
@@ -1143,17 +949,6 @@ jobs:
# sharding moved them to the `coverage-gate` job, which merges every
# shard's dumps. TEST_RESULT therefore does NOT cover them any more;
# COVERAGE_GATE_RESULT below is what gates coverage.
# #4591 (epic #4589 Phase 2): downgraded from a hard gate to a
# non-blocking safety net for one release cycle, per the epic's
# phased-rollout design — test-conformance (above) is now the real
# gate for windows/macos coverage. A red full-suite run here is
# still surfaced (it means the conformance-tier classifier missed
# something) but does not block merge; retire this job entirely
# in a follow-up once the classifier has proven itself over one
# release cycle with no such miss.
if [ "$FULL_TEST_RESULT" != "success" ] && [ "$FULL_TEST_RESULT" != "skipped" ]; then
echo "::warning::full legacy parity matrix did not pass (non-gating safety net — see #4591)"
fi
# #2952: the coverage gate is skipped when product code did not change
# (same condition as the test lane). Only a non-success, non-skipped
@@ -1180,7 +975,7 @@ jobs:
# #2724 (ADR-2719 §5): publishes the differential attribution check's baseline
# artifact after `next` advances, keyed on the merge sha. PR lanes restore it
# (see the `test` and `test-full` jobs' "Restore emitted-baseline cache" steps),
# (see the `test` and `test-conformance` jobs' "Restore emitted-baseline cache" steps),
# keyed on `pull_request.base.sha` — "the next sha the PR was merged with", the
# ADR's own phrasing. Not required by `required-tests`: a miss here degrades PR
# lanes to an in-job build (resolveBaseline()'s documented precedence) rather

View File

@@ -403,7 +403,7 @@ Two properties are load-bearing and easy to break:
running* — gating on it self-deadlocks the pipeline. It is read for the
failure message only.
Gated: `test.yml` (`lint-tests`, `test`, `test-inert`, `test-full`,
Gated: `test.yml` (`lint-tests`, `test`, `test-inert`,
`test-conformance`, `coverage-gate`, `qa-loop-walk`, `required-tests`), `install-smoke.yml`,
`mutation.yml`, `security-scan.yml`, `docs-required.yml`,
`changeset-required.yml`, `default-flip-documentation.yml`, `branch-naming.yml`.
@@ -457,43 +457,33 @@ sha, which invalidates any prior remote-runner verification. Rebase *last*.
## CI matrix
The `Tests` workflow runs every PR through a scoped gate generated by
`scripts/ci-test-scope.cjs`.
The `Tests` workflow (`.github/workflows/test.yml`) runs every PR through a scope
computed by `scripts/ci-test-scope.cjs`'s `classify()`, which sets two flags —
`product_changed` and `full_matrix` — from the changed-file list.
All lanes run on **Node 24** — the `engines.node` floor (`>=24.0.0`) and the
only supported runtime.
| Lane | Scope |
|---|---|
| `ubuntu-latest` (scoped) | scoped tests — fast PR signal |
| `ubuntu-latest` (full, sharded) | unit + integration + security |
| `windows-latest` | scoped Windows/path/shell tests |
| `macos-latest` | full parity when required |
| Job | Lanes | Gated on | Purpose |
|---|---|---|---|
| `test` | `ubuntu-latest` (1 targeted + 3-shard full) + `windows-latest` (3-shard, Windows/path/shell-scoped) | `product_changed == 'true'` | The default, always-scoped PR signal — the full `unit`/`integration`/`security` suites run once, sharded, on Linux; Windows runs the Windows-sensitive subset plus every changed test file |
| `test-inert` | `ubuntu-latest` | `code_changed == 'true' && product_changed != 'true'` | A lightweight lane for PRs that touch only administrative/policy workflow files (code changed, but nothing that needs the real matrix) |
| `test-conformance` | `windows-latest` (3-shard) + `macos-latest` (unsharded) | `code_changed == 'true' && full_matrix == 'true'` | Runs only the **platform-conformance-tier** file list (`scripts/lib/platform-conformance-tier.generated.cjs`, epic #4589 Phase 2/#4591) on real Windows/macOS — the sole gating signal for real-OS coverage. Retired the parallel legacy full-suite matrix in #4603. |
| `coverage-gate` | `ubuntu-latest` | `product_changed == 'true' && test.result == 'success'` | Merges every `test` shard's coverage dumps and evaluates the threshold once (sharding moved this out of the `test` job itself — #2952) |
| `qa-loop-walk` | `ubuntu-latest` | `product_changed == 'true'` | The QA smell-ratchet scenario walk (see "The QA smell ratchet" below) |
| `required-tests` | `ubuntu-latest` | `always()` | Aggregates every job above into the one branch-protection-required check |
- **Scoped tests** are selected from the changed paths, plus a small CLI/package
smoke set. They are for confidence on the affected surface, not for counting
tests.
`full_matrix` fires on any changed `tests/**/*.test.cjs` file unconditionally (restored
by #4421 after #962's narrowing let a real regression through undetected), plus a
handful of curated `RULES` (workflow/installer/hooks/env-gate changes) — see
`scripts/ci-test-scope.cjs`'s own `classify()` for the exact, current rule set; this
doc intentionally does not restate it in full, to avoid drifting out of sync with it.
The default PR gate runs the broad `unit` (under the c8 coverage gate),
`integration`, and `security` suites once on Ubuntu / Node 24, scoped tests on
a second Ubuntu / Node 24 lane, and scoped tests on Windows / Node 24.
"Scoped" means the diff-selected list from the rule table — not the full suite
and not a fixed smoke set (the fixed smoke list is only the empty-selection
fallback). The Windows lane's list is the Windows-sensitive subset of the
selection, plus **every changed test file, unconditionally** (the #494
invariant, narrowed): a modified test is exercised on the divergent OS before
merge at per-file cost, without paying for the three full parity lanes.
PRs touching workflow, package, test-runner, install, release, or
Windows-sensitive surfaces also run the full parity matrix on macOS and the
older Windows runtime, plus `install` and `slow` on the primary Ubuntu lane.
Everything (including the full parity matrix) runs on every push to `next`,
which covers the residual macOS / Windows cross-product for scoped PRs.
Coverage runs inside the Ubuntu / Node 24 full lane (not a separate job — that
duplicated the entire unit run) and stays single-lane because multiplying
coverage across OS/runtime lanes adds cost without improving the threshold
signal. Note the gate's deliberate blind spot: it measures
Coverage is evaluated by the dedicated `coverage-gate` job (moved out of the `test`
job by #2952, once sharding meant no single `test` runner saw the whole picture) and
stays single-lane (Ubuntu / Node 24 only) because multiplying coverage across
OS/runtime lanes adds cost without improving the threshold signal. Note the gate's
deliberate blind spot: it measures
`gsd-core/bin/lib/*.cjs` only — `scripts/`, `hooks/`, and `bin/` are
unenforced, and `stryker.config.mjs` additionally excludes ~48% of lib lines
from mutation testing (see the UNMUTATED list there). Widening either gate is
@@ -540,18 +530,18 @@ would, so a badly stale table cannot collapse the suite into a few fat chunks.
### CI job timeout budgets: report + near-cap warning (#4036)
Every matrixed job — `test` and `test-full` in `test.yml`, `mutate` in
Every matrixed job — `test` in `test.yml`, `mutate` in
`mutation.yml`, `smoke` in `install-smoke.yml` — declares a `timeout-minutes`
cap. `tests/ci-test-job-timeout-budget.test.cjs` enforces that each checked-in
cap stays at least the **headroom factor** (1.5x) above a documented,
hand-measured cost for that job — now all four of the jobs above, not just
`test`/`test-full`/`coverage-gate`/`test-inert` as before.
hand-measured cost for that job — now all three of the jobs above, not just
`test`/`coverage-gate`/`test-inert` as before.
`test-conformance` in `test.yml` (#4591, epic #4589 Phase 2) runs the
`scripts/lib/platform-conformance-tier.generated.cjs` file list on
`windows-latest` (sharded three ways) and `macos-latest` (unsharded) — the
gating signal for real-OS coverage, replacing `test-full`'s old gating role
for one release cycle while `test-full` runs as a non-gating safety net. It
sole gating signal for real-OS coverage (#4603 retired the parallel
legacy full-matrix safety-net job). It
also declares a `timeout-minutes` cap and runs the same in-job near-cap check
described below, but it has no `LANE_COSTS` entry in
`tests/ci-test-job-timeout-budget.test.cjs` yet — no real, completed
@@ -561,7 +551,7 @@ does not cover it until one lands.
Two runtime mechanisms sit on top of that static gate, both new in #4036:
- **In-job near-cap check** (`scripts/ci-check-job-near-cap.cjs`) — the last
step of each of the four jobs computes elapsed-vs-cap from a start-time
step of each of the three jobs computes elapsed-vs-cap from a start-time
marker recorded as that job's first step. At >=90% of budget it emits a
`::warning::` annotation (visible in the PR Checks UI) and a
`$GITHUB_STEP_SUMMARY` block. Advisory only — it never fails the job. Known
@@ -572,7 +562,7 @@ Two runtime mechanisms sit on top of that static gate, both new in #4036:
`scripts/ci-timeout-report.cjs`) — runs daily and on `workflow_dispatch`. It
polls GitHub's Actions REST API for recently completed jobs across
`test.yml`, `mutation.yml`, and `install-smoke.yml`, resolves each job's
declared cap (a literal `timeout-minutes` for `test`/`test-full`/`test-conformance`/`smoke`, or
declared cap (a literal `timeout-minutes` for `test`/`test-conformance`/`smoke`, or
`scripts/mutation-matrix.cjs`'s `COVERED[<module>].timeoutMinutes` for
`mutate`'s per-module shards), and appends any new `(runId, jobName)`
records to `tests/ci-timeout-budget-history.jsonl`. Unlike the in-job check,

View File

@@ -1,12 +1,11 @@
# How to read CI timeout budget signals
Every matrixed CI job (`test`, `test-full`, `test-conformance` in `.github/workflows/test.yml`;
Every matrixed CI job (`test`, `test-conformance` in `.github/workflows/test.yml`;
`mutate` in `mutation.yml`; `smoke` in `install-smoke.yml`) now reports how close it ran to its
`timeout-minutes` cap. This page is for a maintainer trying to answer: *is a lane drifting
toward its cap, and where do I look?* (`test-conformance` runs the platform-conformance-tier file
list — `scripts/lib/platform-conformance-tier.generated.cjs` — on `windows-latest`, sharded three
ways, and `macos-latest`, unsharded; it is the gating signal for real-OS coverage, replacing
`test-full`'s old gating role for one release cycle.)
ways, and `macos-latest`, unsharded; it is the sole gating signal for real-OS coverage.)
## 1. A single run crossed 90% of its budget

View File

@@ -33,7 +33,6 @@ const WORKFLOWS_DIR = path.join(__dirname, '..', '.github', 'workflows');
const JOB_RULES = [
{ workflowFile: 'test.yml', jobKey: 'test-inert', test: (name) => name === 'test (inert CI)' },
{ workflowFile: 'test.yml', jobKey: 'test', test: (name) => name.startsWith('test (') && name !== 'test (inert CI)' },
{ workflowFile: 'test.yml', jobKey: 'test-full', test: (name) => name.startsWith('full test (') },
{ workflowFile: 'test.yml', jobKey: 'coverage-gate', test: (name) => name === 'Coverage gate (merged shards)' },
{ workflowFile: 'test.yml', jobKey: 'test-conformance', test: (name) => name.startsWith('conformance test (') },
{ workflowFile: 'install-smoke.yml', jobKey: 'smoke', test: (name) => name.startsWith('smoke (') },

View File

@@ -16,8 +16,8 @@
* (`scripts/affected-tests-lib.cjs`'s `PR_EXCLUDED_SUITES`; "PRs must never
* select or run these"), and this generator's output feeds a `pull_request`-
* triggered job; (2) `integration`/`security` already run via their own
* separate, dedicated, unsharded, shard-1-only steps in the `test`/
* `test-full` jobs (.github/workflows/test.yml) — folding any of them into
* separate, dedicated, unsharded, shard-1-only steps in the `test` job
* (.github/workflows/test.yml) — folding any of them into
* this job's generic `--files-from` + `--shard` invocation is unproven and,
* per the incident below, unsafe. (3) `qa` (loop-walk-suite files) already
* runs via its own separate, dedicated `qa-loop-walk` job
@@ -46,10 +46,10 @@
* What stands in for it: (1) the most recent push-triggered run on `next`
* (unconditionally full-matrix) is green on every OS for every file in this
* classification, confirmed before this classifier was built; (2) the
* existing `test-full` job keeps running the WHOLE suite on real Windows/
* macOS as a non-gating safety net for one release cycle (.github/workflows/
* test.yml) — a classifier miss surfaces as a visible warning there, not a
* silent gap, before the safety net is retired. This is the same
* legacy full-matrix job ran the WHOLE suite on real Windows/macOS as a
* non-gating safety net for one release cycle (.github/workflows/test.yml)
* before it was retired (#4603) — a classifier miss during that cycle would
* have surfaced as a visible warning there, not a silent gap. This is the same
* static-analysis-substitutes-for-real-OS-execution stance ADR-1703's whole
* rule catalog already takes; it is a real, disclosed limit, not a silent
* substitution.

View File

@@ -671,16 +671,56 @@ function packChunks(files, { weightOf, maxWeight, maxChars, fixedOverhead }) {
// which files end up as that gamble's companions is not something a future
// PR can predict or control.
//
// Isolating it into its own chunk, unconditionally, on every platform,
// removes the gamble at its source rather than tuning the shared budget a
// third time around a moving target: no other file's packing changes (this
// file simply never enters the shared pool `packChunks` balances), and no
// future single-file addition can silently reintroduce this exact failure by
// landing in its chunk. If a future profiling pass genuinely speeds up
// codex-config.test.cjs itself, this isolation can be revisited — this is a
// packing-side mitigation for a KNOWN file's cost, not a statement that the
// cost is irreducible.
const ISOLATED_HEAVY_FILES = new Set(['codex-config.test.cjs']);
// 2026-09-10 (epic #4589 Phase 2/#4591, #4603): the SAME failure hit
// state.test.cjs (weight 21.35, heavier than codex-config.test.cjs) on
// `next`'s own push-triggered Tests run, `conformance test (windows-latest,
// 24, shard 2/3)` chunk 3/6 — 600019ms, killed. Root cause is not a new
// outlier: state.test.cjs was already this heavy before Phase 2 existed.
// What changed is the POOL it gets packed against. Phase 2's
// platform-conformance-tier job packs only the ~546 conformance-tier files
// per shard (vs. the ~950-file full suite `packChunks` used to balance
// against), so the same absolute-weight outlier now represents a much
// larger share of a much smaller, more homogeneous pool — the LPT packer has
// fewer light files available to pad around it. This is a structural risk of
// the smaller conformance-tier pool, not a one-off.
//
// A first attempt at this fix hand-picked a handful of candidates by eye and
// missed three heavier files — caught by an isolated code-review pass, which
// is the reason this comment says "systematically", not "we looked at the
// obvious ones". The corrected method: codex-config.test.cjs's own weight
// (17.87) is 44.7% of the Windows MAX_FILES_PER_CHUNK budget (40) — that
// ratio, not a round "~45%", is the actual established threshold, since it's
// the exact file two prior documented incidents already proved dangerous.
// Computing weight/budget for EVERY unit-suite file in the timings table
// (`suiteOf(f) === null` — the same eligibility test that decides
// conformance-tier membership) and keeping everything at or above that ratio
// found SEVEN files, not four: run-tests-harness.test.cjs (31.23, 78.1%),
// emitted-attribution.test.cjs (26.47, 66.2%), install-minimal-hooks.test.cjs
// (24.45, 61.1%), phase.test.cjs (23.31, 58.3%), state.test.cjs (21.35,
// 53.4%), config.test.cjs (19.76, 49.4%), install.test.cjs (18.84, 47.1%) —
// each at or above codex-config.test.cjs's own proven-dangerous ratio.
//
// Isolating all eight (these seven plus codex-config.test.cjs) into their own
// chunk, unconditionally, on every platform, removes the gamble at its
// source rather than tuning the shared budget again around a moving target:
// no other file's packing changes (these files simply never enter the shared
// pool `packChunks` balances), and no future single-file addition can
// silently reintroduce this exact failure by landing in one of their chunks.
// If a future profiling pass genuinely speeds any of them up, this isolation
// can be revisited — this is a packing-side mitigation for KNOWN files' cost,
// not a statement that the cost is irreducible. If a FUTURE file's measured
// weight ever crosses this same ratio, it needs the same treatment; nothing
// currently re-runs this sweep automatically when the timings table changes.
const ISOLATED_HEAVY_FILES = new Set([
'codex-config.test.cjs',
'run-tests-harness.test.cjs',
'emitted-attribution.test.cjs',
'install-minimal-hooks.test.cjs',
'phase.test.cjs',
'state.test.cjs',
'config.test.cjs',
'install.test.cjs',
]);
/**
* Split `files` (absolute or repo-relative paths) into `{isolated, packable}`

View File

@@ -655,7 +655,7 @@ describe('ci-pr-mergeability: CLI', () => {
/** workflow file -> job ids that must be gated on the preflight. */
const GATED = Object.freeze({
'test.yml': ['lint-tests', 'test', 'test-inert', 'test-full', 'coverage-gate', 'qa-loop-walk', 'required-tests'],
'test.yml': ['lint-tests', 'test', 'test-inert', 'test-conformance', 'coverage-gate', 'qa-loop-walk', 'required-tests'],
'install-smoke.yml': ['smoke', 'smoke-unpacked'],
'mutation.yml': ['detect', 'mutation-gate'],
'security-scan.yml': ['security'],

View File

@@ -96,17 +96,6 @@ const LANE_COSTS = [
// once one exists, the same discipline every other entry here follows.
evidence: 'run 33278340189 — 13m48s completed; run 33285384930 — CANCELLED at ~14m51s (#4070)',
},
{
job: 'test-full',
measuredMinutes: 27,
// Lane moved from windows-22 to windows-latest/24 and is now sharded three
// ways. Worst observed shard is `full test (windows-latest, 24, shard
// 3/3)`: 26m18s on run 32614439702 (shard 2/3 23m36s, shard 1/3 19m22s),
// and 23m17s for shard 2/3 on run 32603886007. The previous 18m59s /
// windows-22 figure recorded here predated this cost and is stale — the
// lane is measurably slower now, not merely relabeled.
evidence: 'run 32614439702 — 26m18s, windows-latest/24 shard 3/3',
},
{
job: 'coverage-gate',
measuredMinutes: 2,
@@ -252,7 +241,6 @@ test('mutation.yml mutate job timeout budgets (#4036)', async (t) => {
test('near-cap check CI_JOB_TIMEOUT_MINUTES literals match each job\'s own timeout-minutes (#4036)', async (t) => {
const staticLanes = [
{ workflowFile: 'test.yml', jobKey: 'test', envLiteral: '32' },
{ workflowFile: 'test.yml', jobKey: 'test-full', envLiteral: '45' },
{ workflowFile: 'test.yml', jobKey: 'test-conformance', envLiteral: '45' },
{ workflowFile: 'install-smoke.yml', jobKey: 'smoke', envLiteral: '12' },
];
@@ -281,11 +269,6 @@ test('near-cap check CI_JOB_TIMEOUT_MINUTES literals match each job\'s own timeo
'test.yml jobs.test.name no longer starts with "test (" — update JOB_RULES in scripts/ci-timeout-report.cjs to match');
assert.equal(testRule.test('test (ubuntu-latest, 24, shard 1/3)'), true);
const testFullRule = JOB_RULES.find((r) => r.workflowFile === 'test.yml' && r.jobKey === 'test-full');
assert.ok(testWorkflow.jobs['test-full'].name.startsWith('full test ('),
'test.yml jobs.test-full.name no longer starts with "full test (" — update JOB_RULES to match');
assert.equal(testFullRule.test('full test (windows-latest, 24, shard 1/3)'), true);
const testInertRule = JOB_RULES.find((r) => r.workflowFile === 'test.yml' && r.jobKey === 'test-inert');
assert.equal(testWorkflow.jobs['test-inert'].name, 'test (inert CI)',
'test.yml jobs.test-inert.name changed — update JOB_RULES to match');

View File

@@ -461,80 +461,9 @@ describe('test.yml changes job contract (#837)', () => {
});
});
describe('test-full shard matrix parity (#1212)', () => {
// DEFECT.GENERATIVE-FIX: the sharded windows full-test lane has TWO surfaces
// that must agree — the `shard:` matrix array (how many parallel jobs run)
// and the `/N` denominator in `run-tests.cjs --suite unit --shard i/N` (how
// many slices the runner partitions the suite into). If they diverge (e.g.
// someone grows `shard: [1,2,3,4]` but leaves `--shard ${{ matrix.shard }}/3`),
// shards silently overlap and one shard errors out. This parity assertion
// fails the moment the two drift.
describe('test.yml gating internals (#2472 / #4241)', () => {
const yaml = require('js-yaml');
function loadTestFull() {
const text = fs.readFileSync(path.join(WORKFLOWS_DIR, 'test.yml'), 'utf8');
const doc = yaml.load(text);
return { text, job: doc.jobs['test-full'] };
}
test('distinct shard values are 1..N matching the --shard /N denominator, on every leg', () => {
const { job } = loadTestFull();
const include = job.strategy.matrix.include;
assert.ok(Array.isArray(include), 'test-full matrix must enumerate `include:` rows');
assert.ok(
include.every(r => Number.isInteger(r.shard)),
'every include row must carry an integer `shard:` key',
);
const distinctShards = [...new Set(include.map(r => r.shard))].sort((a, b) => a - b);
const n = distinctShards.length;
// Distinct shard values must be exactly 1..n (1-based, contiguous) so the
// runner's cost-balanced shard selection covers every file with no gaps/overlaps.
assert.deepStrictEqual(
distinctShards,
Array.from({ length: n }, (_, i) => i + 1),
`distinct shard values must be 1..${n} (1-based, contiguous), got ${JSON.stringify(distinctShards)}`,
);
// Every OS/node leg must appear once per shard (full cross-product) — no
// leg may silently skip a shard, which would drop a third of its coverage.
const legs = [...new Set(include.map(r => `${r.os}|${r['node-version']}`))];
for (const leg of legs) {
const [os, node] = leg.split('|');
const shardsForLeg = include
.filter(r => r.os === os && String(r['node-version']) === node)
.map(r => r.shard)
.sort((a, b) => a - b);
assert.deepStrictEqual(
shardsForLeg,
distinctShards,
`leg ${leg} must run all shards ${JSON.stringify(distinctShards)}, got ${JSON.stringify(shardsForLeg)}`,
);
}
// Full cross-product: every (leg, shard) pair is present exactly once, so
// the row count equals legs × shards with no duplicate/missing combination.
const pairKey = r => `${r.os}|${r['node-version']}|${r.shard}`;
assert.strictEqual(new Set(include.map(pairKey)).size, legs.length * n);
assert.strictEqual(include.length, legs.length * n);
// Find the `--shard ${{ matrix.shard }}/<N>` denominator in the unit step.
const unitStep = job.steps.find(
s => typeof s.run === 'string' && s.run.includes('run-tests.cjs') && s.run.includes('--shard'),
);
assert.ok(unitStep, 'test-full must have a step running run-tests.cjs --shard');
const m = /--shard\s+\$\{\{\s*matrix\.shard\s*\}\}\/(\d+)/.exec(unitStep.run);
assert.ok(m, `could not parse --shard i/N denominator from: ${unitStep.run}`);
const denominator = Number(m[1]);
assert.strictEqual(
denominator,
n,
`shard count (${n}) and --shard /N denominator (${denominator}) must match — ` +
`update both the per-row \`shard:\` values and the \`/N\` in the run command together.`,
);
});
// #2472: every job of a run must merge ONE base commit. Each job runs the
// rebase-check step independently, minutes apart across the matrix, so
// merging the moving branch ref lets jobs see different trees when the base
@@ -591,19 +520,14 @@ describe('test-full shard matrix parity (#1212)', () => {
}
});
test('required-tests fan-in still needs test-full and keeps the protected name', () => {
test('required-tests fan-in keeps the protected name', () => {
// Hyrum's Law: branch protection requires a status check literally named
// "Required tests". Renaming it (or dropping test-full from its needs)
// would silently break the gate. Pin both.
// "Required tests". Renaming it would silently break the gate.
const text = fs.readFileSync(path.join(WORKFLOWS_DIR, 'test.yml'), 'utf8');
const doc = yaml.load(text);
const fanIn = doc.jobs['required-tests'];
assert.ok(fanIn, 'required-tests job must exist');
assert.strictEqual(fanIn.name, 'Required tests', 'the branch-protection check name must stay "Required tests"');
assert.ok(
Array.isArray(fanIn.needs) && fanIn.needs.includes('test-full'),
'required-tests must `needs: test-full` so all shard legs aggregate into the gate',
);
});
test('workflow triggers on merge_group (#4241)', () => {

View File

@@ -2875,9 +2875,20 @@ describe('analyzeChunkEvents (#3889)', () => {
// its chunk and retrigger the per-chunk timeout two prior incidents already
// hit. These tests pin partitionIsolatedFiles directly — the pure split, not
// the chunk-execution loop around it.
//
// 2026-09-10 (#4603): epic #4589 Phase 2's platform-conformance-tier job packs
// a much smaller file pool per shard than the full suite did, which exposed
// the SAME failure on state.test.cjs (weight 21.35, heavier than
// codex-config.test.cjs) on `next` itself. Re-running the same weight-table
// analysis found two more unisolated files at or above the same ~45%-of-budget
// threshold: run-tests-harness.test.cjs (31.23) and phase.test.cjs (23.31),
// plus config.test.cjs (19.76) just under codex-config.test.cjs's own 45% but
// still heavier than several already-risky files. All four added to
// ISOLATED_HEAVY_FILES; see scripts/run-tests.cjs's own comment for the full
// weight/budget accounting.
const { ISOLATED_HEAVY_FILES, partitionIsolatedFiles } = require('../scripts/run-tests.cjs');
describe('partitionIsolatedFiles (#4497 codex-config.test.cjs chunk isolation)', () => {
describe('partitionIsolatedFiles (#4497 codex-config.test.cjs chunk isolation, extended #4603)', () => {
test('an isolated-heavy file is split out, in its own bucket, everything else stays packable', () => {
const files = [
'/repo/tests/a.test.cjs',
@@ -2889,6 +2900,30 @@ describe('partitionIsolatedFiles (#4497 codex-config.test.cjs chunk isolation)',
assert.deepStrictEqual(packable, ['/repo/tests/a.test.cjs', '/repo/tests/b.test.cjs']);
});
test('#4603: every newly-isolated heavy file is split out individually, in original order', () => {
const files = [
'/repo/tests/a.test.cjs',
'/repo/tests/state.test.cjs',
'/repo/tests/b.test.cjs',
'/repo/tests/phase.test.cjs',
'/repo/tests/run-tests-harness.test.cjs',
'/repo/tests/config.test.cjs',
'/repo/tests/c.test.cjs',
];
const { isolated, packable } = partitionIsolatedFiles(files);
assert.deepStrictEqual(isolated, [
'/repo/tests/state.test.cjs',
'/repo/tests/phase.test.cjs',
'/repo/tests/run-tests-harness.test.cjs',
'/repo/tests/config.test.cjs',
]);
assert.deepStrictEqual(packable, [
'/repo/tests/a.test.cjs',
'/repo/tests/b.test.cjs',
'/repo/tests/c.test.cjs',
]);
});
test('matches by BASENAME, so it isolates regardless of platform path separator or directory prefix', () => {
const files = [
'C:\\repo\\tests\\codex-config.test.cjs',
@@ -2918,7 +2953,66 @@ describe('partitionIsolatedFiles (#4497 codex-config.test.cjs chunk isolation)',
assert.deepStrictEqual(partitionIsolatedFiles([]), { isolated: [], packable: [] });
});
test('ISOLATED_HEAVY_FILES currently names exactly codex-config.test.cjs (documents the set the fix scoped to)', () => {
assert.deepStrictEqual([...ISOLATED_HEAVY_FILES], ['codex-config.test.cjs']);
test('ISOLATED_HEAVY_FILES currently names exactly the eight known-heavy files (documents the set the fix scoped to)', () => {
assert.deepStrictEqual(
[...ISOLATED_HEAVY_FILES].sort(),
[
'codex-config.test.cjs',
'config.test.cjs',
'emitted-attribution.test.cjs',
'install-minimal-hooks.test.cjs',
'install.test.cjs',
'phase.test.cjs',
'run-tests-harness.test.cjs',
'state.test.cjs',
].sort(),
);
});
// #4603: a durable guard, not a one-time snapshot. A first attempt at this
// fix hand-picked candidates by eye and missed three heavier files (caught
// by an isolated code-review pass) — this test closes that gap by
// RE-DERIVING the same weight/budget computation from the live timings
// table on every run, so a future test file crossing the same threshold
// fails this test instead of silently reintroducing the per-chunk-timeout
// failure this whole mechanism exists to prevent.
test('#4603: no unisolated unit-suite file exceeds ISOLATED_HEAVY_FILES\' own established threshold', () => {
const { suiteOf } = require('../scripts/lib/suite-detection.cjs');
const table = require('../tests/test-timings.json');
const timings = table.timings;
const values = Object.values(timings).filter(
(v) => typeof v === 'number' && Number.isFinite(v) && v >= 0,
);
const mean = values.reduce((sum, v) => sum + v, 0) / values.length;
const WINDOWS_BUDGET = 40;
// The threshold is codex-config.test.cjs's OWN ratio — the exact file two
// prior documented incidents proved dangerous — not an arbitrarily chosen
// round number. This makes the test self-consistent even if
// WINDOWS_BUDGET or the timings table changes: it always asks "is this
// file at least as dangerous as the file we already know is dangerous?"
assert.ok(
Object.hasOwn(timings, 'codex-config.test.cjs'),
'codex-config.test.cjs must remain in the timings table to anchor this threshold',
);
const codexRatio = timings['codex-config.test.cjs'] / mean / WINDOWS_BUDGET;
const exceedsThreshold = [];
for (const [file, ms] of Object.entries(timings)) {
if (typeof ms !== 'number' || !Number.isFinite(ms) || ms < 0) continue;
if (suiteOf(file) !== null) continue; // suite-tagged files never enter this pool
const ratio = ms / mean / WINDOWS_BUDGET;
if (ratio >= codexRatio && !ISOLATED_HEAVY_FILES.has(file)) {
exceedsThreshold.push(`${file} (${(ratio * 100).toFixed(1)}% of budget)`);
}
}
assert.deepStrictEqual(
exceedsThreshold,
[],
`file(s) at/above codex-config.test.cjs's own danger ratio (${(codexRatio * 100).toFixed(1)}%) ` +
`are not in ISOLATED_HEAVY_FILES: ${exceedsThreshold.join(', ')} — add them, following ` +
`scripts/run-tests.cjs's ISOLATED_HEAVY_FILES comment for the pattern`,
);
});
});