* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% Adds two new mechanisms plus an audit-coverage extension: - scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic - scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check, wired into test/test-full/mutate/smoke as each job's last step - scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml: scheduled REST-API poll that appends new records to tests/ci-timeout-budget-history.jsonl and opens a small data-only PR - tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate (mutation.yml) and smoke (install-smoke.yml), which previously had no headroom-factor gate coverage at all Does not change any timeout-minutes value, shard composition, or shard-1 contents — those stay maintainer policy calls per the issue's own scope. * fix(#4036): address two-orthogonal-review findings - Parity tests guarding the two hand-duplicated literals this design cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES vs each job's own timeout-minutes, and ci-timeout-report.cjs's JOB_RULES name-prefixes vs each job's actual name: template. - Thread run.event through as runEvent on every persisted record, so PR-context and push-context install-smoke timings (genuinely different matrix shape) are distinguishable in the history rather than silently conflated under one job name. - Replace the Windows near-cap start-time step's ambiguous PowerShell +/>> precedence with GitHub's documented string-interpolation form. - Move github.run_id out of direct ${{ }} shell interpolation into an env: var in the new scheduled workflow, per this repo's own expression-injection-safe convention. * test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs npm run gen:install-tree — scripts/ ships wholesale into the installed package (per ADR/known-defect precedent from #4012's own PR history: a new scripts/lib/*.cjs file needs its golden entry regenerated or every runtime's install-tree test fails). Confirmed via gsd-test: this was the sole cause of the first real verification run's 25 failures (all in tests/golden-install-tree.test.cjs, one per runtime). Top-level scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs) are not individually tracked in these fixtures — consistent with every other existing top-level scripts/*.cjs file, so no entry was expected or added for those two. * fix(#4036): register new lib file with installer, fix H1 shell policy - bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a hand-maintained registry, not generated — tests/install.test.cjs asserts every scripts/lib/ file is enumerated here) - test.yml: replace the two OS-conditional "Record job start time" step pairs (test + test-full jobs) with a single unconditional `node -e` step. The prior pair's Windows variant declared an explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker statically flags against every OS a job's matrix can realize, independent of the step's own if: gate. A single Node one-liner needs no shell override at all — it's syntactically valid and behaves identically under bash, zsh, and pwsh — which is both H1 compliant and removes the last OS-specific shell syntax from this change entirely. Both defects were found by a real gsd-test run, not local gates — lint:ci and build:lib were clean throughout because neither the scripts/lib/ install-manifest parity check nor the H1 shell-policy baseline runs as part of lint:ci; both are gsd-test-only suites. * docs(#4036): how-to for reading CI timeout budget signals The phase-gate docs check correctly flagged the enablement sequence as 3 real steps (read the near-cap warning, find the accumulated trend file, pick the right maintainer lever) — a reference table can't carry a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from docs/README.md. * chore(#4036): backfill changeset PR number (4043) --------- Co-authored-by: sim <sim@local>
3.6 KiB
How to read CI timeout budget signals
Every matrixed CI job (test, test-full in .github/workflows/test.yml; mutate in
mutation.yml; smoke in install-smoke.yml) now reports how close it ran to its
timeout-minutes cap. This page is for a maintainer trying to answer: is a lane drifting
toward its cap, and where do I look?
1. A single run crossed 90% of its budget
Open the job's page in the Actions run — two places show it, both populated by the same
computation (scripts/lib/ci-job-timing.cjs):
- The Checks tab annotation. A
::warning::line renders as an expandable warning banner on the PR's Checks summary, naming the job, its elapsed time, its cap, and the percentage — visible without opening the job's logs. - The job's step summary. The same information, as a Markdown line, appended to the job's
own summary page (
$GITHUB_STEP_SUMMARY) by that job's own "Check job budget (near-cap advisory)" step — always the job's last step.
Neither signal fails the job. A near-cap warning on an otherwise-green run means exactly what it says: this run finished, but with less margin than the headroom-factor gate assumes it has.
If the warning never appears even on a job that was actually cancelled at its cap, that is expected — a killed job never reaches its last step, so the in-job check never runs. See §2.
2. Checking the accumulated trend
.github/workflows/ci-timeout-report.yml runs daily (and on-demand via workflow_dispatch). It
polls GitHub's Actions REST API directly — independent of whether any individual job's own
near-cap step ran — so it also catches jobs that were actually cancelled by a timeout breach
(GitHub's Jobs API still reports started_at/completed_at for a cancelled job).
Each run's new rows land in tests/ci-timeout-budget-history.jsonl, one JSON object per line:
{"runId":123456,"jobName":"test (ubuntu-latest, 24, shard 1/3)","workflowFile":"test.yml","sha":"...","completedAt":"...","elapsedMs":432000,"timeoutMinutes":15,"pct":0.8,"runEvent":"push"}
runEvent matters for install-smoke.yml's smoke job specifically — its pull_request runs
use a smaller matrix (no macos-latest full_only row) than its push runs, so a pct figure
only means the same thing across rows sharing the same runEvent.
Because next is a protected branch, the report never pushes directly to it — each scheduled
run opens (or the prior run's already merged, in which case a fresh one opens) a small,
data-only PR carrying just that run's new rows, titled chore: CI timeout budget report — run <id>. Merge these like any other PR; there is nothing to review beyond "did the numbers land."
3. A lane is repeatedly near-cap — what to do
Neither mechanism here decides what to do about a lane that's genuinely trending toward its cap. That is a maintainer call among three options, each with real tradeoffs:
- Raise the
timeout-minutescap for that job. - Rebalance the shard split so no single shard carries a disproportionate share of the
suite (see
scripts/run-tests.cjs'sselectShard, which packs shards by measured cost fromtests/test-timings.json). - Trim what runs on the long-pole shard — for the
testjob, shard 1 also carries the unsharded aux suites (integration/security/install/slow); moving one elsewhere changes what shard 1 costs.
tests/ci-test-job-timeout-budget.test.cjs will keep failing to accept a lowered
timeout-minutes beneath 1.5x whatever LANE_COSTS/COVERED[*].timeoutMinutes records as that
job's last measured cost — raising the cap back down is not something either mechanism will
silently allow.