Files
msd-core/docs/how-to/read-ci-timeout-signals.md
Tom Boucher 370cfc6680 enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90%

Adds two new mechanisms plus an audit-coverage extension:

- scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic
- scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check,
  wired into test/test-full/mutate/smoke as each job's last step
- scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml:
  scheduled REST-API poll that appends new records to
  tests/ci-timeout-budget-history.jsonl and opens a small data-only PR
- tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate
  (mutation.yml) and smoke (install-smoke.yml), which previously had no
  headroom-factor gate coverage at all

Does not change any timeout-minutes value, shard composition, or shard-1
contents — those stay maintainer policy calls per the issue's own scope.

* fix(#4036): address two-orthogonal-review findings

- Parity tests guarding the two hand-duplicated literals this design
  cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES
  vs each job's own timeout-minutes, and ci-timeout-report.cjs's
  JOB_RULES name-prefixes vs each job's actual name: template.
- Thread run.event through as runEvent on every persisted record, so
  PR-context and push-context install-smoke timings (genuinely
  different matrix shape) are distinguishable in the history rather
  than silently conflated under one job name.
- Replace the Windows near-cap start-time step's ambiguous PowerShell
  +/>> precedence with GitHub's documented string-interpolation form.
- Move github.run_id out of direct ${{ }} shell interpolation into an
  env: var in the new scheduled workflow, per this repo's own
  expression-injection-safe convention.

* test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs

npm run gen:install-tree — scripts/ ships wholesale into the installed
package (per ADR/known-defect precedent from #4012's own PR history: a
new scripts/lib/*.cjs file needs its golden entry regenerated or every
runtime's install-tree test fails). Confirmed via gsd-test: this was the
sole cause of the first real verification run's 25 failures (all in
tests/golden-install-tree.test.cjs, one per runtime). Top-level
scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs)
are not individually tracked in these fixtures — consistent with every
other existing top-level scripts/*.cjs file, so no entry was expected
or added for those two.

* fix(#4036): register new lib file with installer, fix H1 shell policy

- bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a
  hand-maintained registry, not generated — tests/install.test.cjs
  asserts every scripts/lib/ file is enumerated here)
- test.yml: replace the two OS-conditional "Record job start time"
  step pairs (test + test-full jobs) with a single unconditional
  `node -e` step. The prior pair's Windows variant declared an
  explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker
  statically flags against every OS a job's matrix can realize,
  independent of the step's own if: gate. A single Node one-liner
  needs no shell override at all — it's syntactically valid and
  behaves identically under bash, zsh, and pwsh — which is both H1
  compliant and removes the last OS-specific shell syntax from this
  change entirely.

Both defects were found by a real gsd-test run, not local gates —
lint:ci and build:lib were clean throughout because neither the
scripts/lib/ install-manifest parity check nor the H1 shell-policy
baseline runs as part of lint:ci; both are gsd-test-only suites.

* docs(#4036): how-to for reading CI timeout budget signals

The phase-gate docs check correctly flagged the enablement sequence as
3 real steps (read the near-cap warning, find the accumulated trend
file, pick the right maintainer lever) — a reference table can't carry
a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from
docs/README.md.

* chore(#4036): backfill changeset PR number (4043)

---------

Co-authored-by: sim <sim@local>
2026-08-29 16:13:15 -04:00

3.6 KiB

How to read CI timeout budget signals

Every matrixed CI job (test, test-full in .github/workflows/test.yml; mutate in mutation.yml; smoke in install-smoke.yml) now reports how close it ran to its timeout-minutes cap. This page is for a maintainer trying to answer: is a lane drifting toward its cap, and where do I look?

1. A single run crossed 90% of its budget

Open the job's page in the Actions run — two places show it, both populated by the same computation (scripts/lib/ci-job-timing.cjs):

  • The Checks tab annotation. A ::warning:: line renders as an expandable warning banner on the PR's Checks summary, naming the job, its elapsed time, its cap, and the percentage — visible without opening the job's logs.
  • The job's step summary. The same information, as a Markdown line, appended to the job's own summary page ($GITHUB_STEP_SUMMARY) by that job's own "Check job budget (near-cap advisory)" step — always the job's last step.

Neither signal fails the job. A near-cap warning on an otherwise-green run means exactly what it says: this run finished, but with less margin than the headroom-factor gate assumes it has.

If the warning never appears even on a job that was actually cancelled at its cap, that is expected — a killed job never reaches its last step, so the in-job check never runs. See §2.

2. Checking the accumulated trend

.github/workflows/ci-timeout-report.yml runs daily (and on-demand via workflow_dispatch). It polls GitHub's Actions REST API directly — independent of whether any individual job's own near-cap step ran — so it also catches jobs that were actually cancelled by a timeout breach (GitHub's Jobs API still reports started_at/completed_at for a cancelled job).

Each run's new rows land in tests/ci-timeout-budget-history.jsonl, one JSON object per line:

{"runId":123456,"jobName":"test (ubuntu-latest, 24, shard 1/3)","workflowFile":"test.yml","sha":"...","completedAt":"...","elapsedMs":432000,"timeoutMinutes":15,"pct":0.8,"runEvent":"push"}

runEvent matters for install-smoke.yml's smoke job specifically — its pull_request runs use a smaller matrix (no macos-latest full_only row) than its push runs, so a pct figure only means the same thing across rows sharing the same runEvent.

Because next is a protected branch, the report never pushes directly to it — each scheduled run opens (or the prior run's already merged, in which case a fresh one opens) a small, data-only PR carrying just that run's new rows, titled chore: CI timeout budget report — run <id>. Merge these like any other PR; there is nothing to review beyond "did the numbers land."

3. A lane is repeatedly near-cap — what to do

Neither mechanism here decides what to do about a lane that's genuinely trending toward its cap. That is a maintainer call among three options, each with real tradeoffs:

  • Raise the timeout-minutes cap for that job.
  • Rebalance the shard split so no single shard carries a disproportionate share of the suite (see scripts/run-tests.cjs's selectShard, which packs shards by measured cost from tests/test-timings.json).
  • Trim what runs on the long-pole shard — for the test job, shard 1 also carries the unsharded aux suites (integration/security/install/slow); moving one elsewhere changes what shard 1 costs.

tests/ci-test-job-timeout-budget.test.cjs will keep failing to accept a lowered timeout-minutes beneath 1.5x whatever LANE_COSTS/COVERED[*].timeoutMinutes records as that job's last measured cost — raising the cap back down is not something either mechanism will silently allow.