Files
msd-core/tests/ci-test-job-timeout-budget.test.cjs
Tom Boucher 107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00

173 lines
7.8 KiB
JavaScript

'use strict';
/**
* CI job timeout budgets — .github/workflows/test.yml (#2952).
*
* A GitHub Actions job that exceeds its `timeout-minutes` is reported
* `cancelled`, not `failed`. That conclusion propagates into `Required tests`
* and reddens the branch, while reading like someone hit the cancel button —
* which is what makes this failure mode expensive to diagnose and worth a gate.
*
* It has now happened three times in this repo: #1051 and #1212 on the Windows
* full-test lane, and #2952 on the unsharded `test` lane, whose
* `ubuntu-latest / 24` entry is the only `scope: full` matrix entry — it runs
* the whole unit suite under c8 coverage, then the scripts/ coverage floor,
* integration, security, install and slow, serially on one runner. Every one of
* those was the same root cause: a budget sized to what the lane cost that
* week, with no headroom for the suite to grow into.
*
* So the rule enforced here is not a fixed number per lane — it is a HEADROOM
* FACTOR over each lane's measured cost. A lane may be slow; what it may not be
* is budgeted to finish with seconds to spare.
*
* This is a budget assertion, not a duration assertion. No unit test can prove
* a lane still FITS its budget — only a real CI run measures that. What this
* file guarantees is that a budget cannot be quietly lowered back beneath what
* its lane is already known to need.
*/
const test = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const yaml = require('js-yaml');
const WORKFLOWS_DIR = path.join(__dirname, '..', '.github', 'workflows');
function loadWorkflow(name) {
return yaml.load(fs.readFileSync(path.join(WORKFLOWS_DIR, name), 'utf8'));
}
/**
* Multiplier applied to a lane's measured cost to get its required budget.
* 1.5x is enough slack for ordinary suite growth across a release cycle without
* letting a genuinely runaway lane hide behind a large number.
*/
const HEADROOM_FACTOR = 1.5;
/**
* Measured wall-clock cost per lane, in whole minutes rounded UP, each from a
* named run. Raise an entry only alongside a fresh measurement — never to make
* a red gate green. Raising a measurement raises the required budget with it,
* which is the point: a lane that got slower must be re-budgeted, not excused.
*/
const LANE_COSTS = [
{
job: 'test',
measuredMinutes: 8,
// Sharded three ways as of #2952, so this is ONE shard's cost, not the
// whole unit suite. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s.
// Shard 1 is the long pole because the unsharded aux suites ride on it.
// Before sharding the same lane cost 15m20s and blew a 15-minute cap.
//
// This one `timeout-minutes` also covers the `scope: windows` matrix
// entries — GitHub applies a single job-level budget across every matrix
// combination, not one per entry. That lane is now sharded three ways too
// (#3057), but no post-sharding per-shard measurement exists yet: its only
// recorded cost is the PRE-sharding whole-suite run that hit 15m05s and was
// CANCELLED on PR #3094. Each of its three shards should now cost roughly a
// third of that (~5m), which is already comfortably under the 8m/12m this
// entry requires — so no separate LANE_COSTS entry is added on a number
// that has not actually been measured. Replace this estimate with a real
// measured shard cost once one exists, the same discipline every other
// entry here follows.
evidence: 'run 30677442953 — 7m12s slowest shard',
},
{
job: 'test-full',
measuredMinutes: 27,
// Lane moved from windows-22 to windows-latest/24 and is now sharded three
// ways. Worst observed shard is `full test (windows-latest, 24, shard
// 3/3)`: 26m18s on run 32614439702 (shard 2/3 23m36s, shard 1/3 19m22s),
// and 23m17s for shard 2/3 on run 32603886007. The previous 18m59s /
// windows-22 figure recorded here predated this cost and is stale — the
// lane is measurably slower now, not merely relabeled.
evidence: 'run 32614439702 — 26m18s, windows-latest/24 shard 3/3',
},
{
job: 'coverage-gate',
measuredMinutes: 2,
// Downloads three shards' raw V8 dumps, renders one merged report and runs
// both thresholds. Run 30677442953: 1m20s end to end, most of it npm ci.
evidence: 'run 30677442953 — 1m20s',
},
{
job: 'test-inert',
measuredMinutes: 2,
// Runs only the targeted-test step when no product code changed; observed
// around a minute. Listed so its budget cannot be dropped to nothing.
evidence: 'targeted-only lane, ~1m observed',
},
];
function requiredBudgetMinutes(measuredMinutes, headroomFactor = HEADROOM_FACTOR) {
return Math.ceil(measuredMinutes * headroomFactor);
}
function hasSufficientBudget(budgetMinutes, measuredMinutes, headroomFactor = HEADROOM_FACTOR) {
return Number.isInteger(budgetMinutes)
&& budgetMinutes >= requiredBudgetMinutes(measuredMinutes, headroomFactor);
}
test('CI job timeout budgets carry headroom over measured cost (#2952)', async (t) => {
const workflow = loadWorkflow('test.yml');
for (const lane of LANE_COSTS) {
await t.test(`${lane.job} is budgeted above its measured cost`, () => {
assert.ok(
workflow.jobs && Object.prototype.hasOwnProperty.call(workflow.jobs, lane.job),
`.github/workflows/test.yml declares no job \`${lane.job}\`. If it was `
+ 'renamed or removed, update LANE_COSTS in this file to match — do not '
+ 'delete the entry to make this pass.',
);
const budget = workflow.jobs[lane.job]['timeout-minutes'];
const required = requiredBudgetMinutes(lane.measuredMinutes);
assert.equal(
typeof budget, 'number',
`.github/workflows/test.yml jobs.${lane.job} must declare timeout-minutes`,
);
assert.ok(
hasSufficientBudget(budget, lane.measuredMinutes),
`jobs.${lane.job}.timeout-minutes is ${budget}, but the lane measured `
+ `${lane.measuredMinutes}m (${lane.evidence}) and needs at least `
+ `${required} — ${HEADROOM_FACTOR}x — so suite growth does not breach `
+ 'the cap. A job that exceeds timeout-minutes is reported `cancelled` '
+ 'and reddens `Required tests`.',
);
});
}
// Boundary coverage on the predicate that decides every lane above:
// required-1 must be rejected, required and required+1 accepted.
await t.test('budget sufficiency is exact at the boundary', () => {
for (const lane of LANE_COSTS) {
const required = requiredBudgetMinutes(lane.measuredMinutes);
assert.equal(hasSufficientBudget(required - 1, lane.measuredMinutes), false,
`${lane.job}: a budget one minute under the requirement must be rejected`);
assert.equal(hasSufficientBudget(required, lane.measuredMinutes), true,
`${lane.job}: a budget exactly at the requirement must be accepted`);
assert.equal(hasSufficientBudget(required + 1, lane.measuredMinutes), true,
`${lane.job}: a budget over the requirement must be accepted`);
}
});
await t.test('a non-integer budget is not a sufficient budget', () => {
// `timeout-minutes: 24.5` is not something GitHub accepts; treating it as
// sufficient would let a malformed workflow through this gate.
assert.equal(hasSufficientBudget(24.5, 16), false);
assert.equal(hasSufficientBudget(Number.NaN, 16), false);
assert.equal(hasSufficientBudget(undefined, 16), false);
});
await t.test('requiredBudgetMinutes rounds up rather than truncating', () => {
// 15 * 1.5 = 22.5 — truncation would hand back 22 and under-budget the lane.
assert.equal(requiredBudgetMinutes(15, 1.5), 23);
assert.equal(requiredBudgetMinutes(16, 1.5), 24);
assert.equal(requiredBudgetMinutes(19, 1.5), 29);
assert.equal(requiredBudgetMinutes(10, 1.5), 15);
});
});