* chore(#4603): retire the test-full CI job Phase 2 (#4591) added test-conformance but left test-full (the pre-existing full-suite Windows/macOS replay) running unchanged, gated on the same full_matrix flag, downgraded only from a hard gate to a non-blocking ::warning:: -- framed as "a non-gating safety net for one release cycle." No phase or issue ever retired it. Result: every full_matrix=true PR ran 10 OS-specific jobs (test-full's 6 + test-conformance's 4, purely additive) instead of the original 6 -- the epic's own goal (reduce runner-minutes) was measurably regressing, not improving, for the majority of PRs. This phase was missing from the original 4-phase epic decomposition; the epic (#4589) has been amended to add it as Phase 5 (see its comment thread), and this issue was filed as the tracked sub-issue. Deletes the test-full job from .github/workflows/test.yml entirely, along with every reference to it: required-tests' needs/FULL_TEST_RESULT warning branch, ci-timeout-report.cjs's JOB_RULES entry, ci-test-job-timeout-budget.test.cjs's LANE_COSTS/staticLanes/testFullRule entries, ci-test-scope.test.cjs's test-full-specific tests (preserving three unrelated tests that were nested in the same describe block, moved under a renamed describe rather than deleted), and docs mentions. test-conformance is now the sole gating signal for real-OS coverage. Two separate defects found and fixed while auditing every test-full reference: - tests/ci-pr-mergeability.test.cjs's GATED['test.yml'] safety-critical array (jobs that must needs: the mergeability preflight) had test-full but was missing test-conformance entirely -- Phase 2 never added it. Verified the real workflow wiring was already correct (test-conformance does have needs: [changes, preflight]); this was a test-coverage gap, not a live defect. Fixed by swapping the array entry. - docs/TESTING-SUITES.md's "## CI matrix" section was substantially stale independent of this phase (predating even #2952's coverage-gate split). Rewritten against the real, current job topology, verified directly against test.yml rather than trusted from memory. An isolated code-review pass found and fixed two minor inaccuracies in the rewritten docs table (two jobs' "Gated on" column didn't match their real if: condition exactly). An isolated security-review pass found no qualifying findings -- every compute-provisioning job already carries needs: preflight directly, unaffected by this deletion. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(ci): isolate 7 more heavy test files from chunk-weight packing `next`'s own push-triggered Tests run failed: `conformance test (windows-latest, 24, shard 2/3)` chunk 3/6 was killed after 600019ms. Root cause: state.test.cjs (weight 21.35, measured) was packed alongside companions by run-tests.cjs's LPT chunk packer, the same failure mode that previously hit codex-config.test.cjs (weight 17.87) twice and got a dedicated fix (ISOLATED_HEAVY_FILES, #4497) -- but state.test.cjs was never added to that set. This is a direct, unintended consequence of epic #4589 Phase 2: the new platform-conformance-tier job packs only ~546 files per shard (vs. the ~950-file full suite the packer used to balance against), so the same absolute-weight outlier now represents a larger share of a smaller, more homogeneous pool -- the LPT packer has fewer light files to pad around it with. This was a real, foreseeable side effect of shrinking the packing pool that nobody checked for when Phase 2 shipped. A first attempt at this fix hand-picked 4 candidates by eyeballing a truncated weight list and missed 3 heavier ones -- caught by an isolated code-review pass (blocker: emitted-attribution.test.cjs at 66.2% of the Windows chunk budget, install-minimal-hooks.test.cjs at 61.1%, install.test.cjs at 47.1%, all above codex-config.test.cjs's own 44.7% -- the ratio that already proved dangerous twice). Corrected by systematically computing weight/budget for every unit-suite file and isolating everything at or above that same ratio: 7 files total, plus the pre-existing codex-config.test.cjs (8 total). Added a durable regression test (tests/run-tests-harness.test.cjs) that re-derives this exact computation from the live tests/test-timings.json on every run, so a future heavy file crossing this threshold fails the test instead of silently reintroducing this failure -- not just a one-time manual sweep. Verified end-to-end: simulated the real 3-way windows shard split of the actual conformance-tier file list with the real packing functions. Max packable-chunk weight across all 3 shards is now 27.04 / 24.10 / 23.91 (shard 2 is the exact shard that failed on next), comfortably under the 40 budget -- versus 40+ and a 600s kill before this fix. A second isolated code-review + security-review pass on the corrected diff found nothing further. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
67 lines
3.9 KiB
Markdown
67 lines
3.9 KiB
Markdown
# How to read CI timeout budget signals
|
|
|
|
Every matrixed CI job (`test`, `test-conformance` in `.github/workflows/test.yml`;
|
|
`mutate` in `mutation.yml`; `smoke` in `install-smoke.yml`) now reports how close it ran to its
|
|
`timeout-minutes` cap. This page is for a maintainer trying to answer: *is a lane drifting
|
|
toward its cap, and where do I look?* (`test-conformance` runs the platform-conformance-tier file
|
|
list — `scripts/lib/platform-conformance-tier.generated.cjs` — on `windows-latest`, sharded three
|
|
ways, and `macos-latest`, unsharded; it is the sole gating signal for real-OS coverage.)
|
|
|
|
## 1. A single run crossed 90% of its budget
|
|
|
|
Open the job's page in the Actions run — two places show it, both populated by the same
|
|
computation (`scripts/lib/ci-job-timing.cjs`):
|
|
|
|
- **The Checks tab annotation.** A `::warning::` line renders as an expandable warning banner
|
|
on the PR's Checks summary, naming the job, its elapsed time, its cap, and the percentage —
|
|
visible without opening the job's logs.
|
|
- **The job's step summary.** The same information, as a Markdown line, appended to the job's
|
|
own summary page (`$GITHUB_STEP_SUMMARY`) by that job's own "Check job budget (near-cap
|
|
advisory)" step — always the job's last step.
|
|
|
|
Neither signal fails the job. A near-cap warning on an otherwise-green run means exactly what
|
|
it says: this run finished, but with less margin than the headroom-factor gate assumes it has.
|
|
|
|
**If the warning never appears even on a job that was actually cancelled at its cap**, that is
|
|
expected — a killed job never reaches its last step, so the in-job check never runs. See §2.
|
|
|
|
## 2. Checking the accumulated trend
|
|
|
|
`.github/workflows/ci-timeout-report.yml` runs daily (and on-demand via `workflow_dispatch`). It
|
|
polls GitHub's Actions REST API directly — independent of whether any individual job's own
|
|
near-cap step ran — so it also catches jobs that were actually cancelled by a timeout breach
|
|
(GitHub's Jobs API still reports `started_at`/`completed_at` for a cancelled job).
|
|
|
|
Each run's new rows land in `tests/ci-timeout-budget-history.jsonl`, one JSON object per line:
|
|
|
|
```json
|
|
{"runId":123456,"jobName":"test (ubuntu-latest, 24, shard 1/3)","workflowFile":"test.yml","sha":"...","completedAt":"...","elapsedMs":432000,"timeoutMinutes":15,"pct":0.8,"runEvent":"push"}
|
|
```
|
|
|
|
`runEvent` matters for `install-smoke.yml`'s `smoke` job specifically — its `pull_request` runs
|
|
use a smaller matrix (no `macos-latest` `full_only` row) than its `push` runs, so a `pct` figure
|
|
only means the same thing across rows sharing the same `runEvent`.
|
|
|
|
Because `next` is a protected branch, the report never pushes directly to it — each scheduled
|
|
run opens (or the prior run's already merged, in which case a fresh one opens) a small,
|
|
data-only PR carrying just that run's new rows, titled `chore: CI timeout budget report — run
|
|
<id>`. Merge these like any other PR; there is nothing to review beyond "did the numbers land."
|
|
|
|
## 3. A lane is repeatedly near-cap — what to do
|
|
|
|
Neither mechanism here decides what to do about a lane that's genuinely trending toward its
|
|
cap. That is a maintainer call among three options, each with real tradeoffs:
|
|
|
|
- **Raise the `timeout-minutes` cap** for that job.
|
|
- **Rebalance the shard split** so no single shard carries a disproportionate share of the
|
|
suite (see `scripts/run-tests.cjs`'s `selectShard`, which packs shards by measured cost from
|
|
`tests/test-timings.json`).
|
|
- **Trim what runs on the long-pole shard** — for the `test` job, shard 1 also carries the
|
|
unsharded aux suites (integration/security/install/slow); moving one elsewhere changes what
|
|
shard 1 costs.
|
|
|
|
`tests/ci-test-job-timeout-budget.test.cjs` will keep failing to accept a lowered
|
|
`timeout-minutes` beneath 1.5x whatever `LANE_COSTS`/`COVERED[*].timeoutMinutes` records as that
|
|
job's last measured cost — raising the cap back down is not something either mechanism will
|
|
silently allow.
|