Files
msd-core/.changeset/fierce-geese-hop.md
Tom Boucher 86452da7cb fix(#4070): reserve shard 1's aux-suite cost out of the LPT unit-test packer (#4072)
* test(#4070): failing-first regression for shard1 aux-suite budget imbalance

- selectShard has no way to reserve virtual weight on a bin, so the LPT
  unit-test packer cannot account for shard 1's fixed aux-suite cost
  (integration/security/install/slow all pinned to shard 1/3).
- test.yml wires no such reserve into the workflow.
- ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test`
  was stale (7m12s from run 30677442953, predating the aux-suite growth);
  corrected to the real evidence cited in #4070 (13m48s / cancelled at
  ~14m51s), which now honestly fails the file's own 1.5x headroom policy
  against the current 15-minute cap.

All three are expected RED on this commit; see
.gsd/bug/fix-4070-shard1-aux-suite-budget/50-test-matrix.md.

* fix(#4070): reserve shard 1's aux-suite cost out of the LPT unit-test packer

selectShard now accepts an optional initialWeights array giving one or more
bins a virtual head start before any file is placed, so LPT converges each
bin's FINAL total (assigned weight + head start) toward equal instead of
raw assigned weight alone. test.yml wires RUN_TESTS_SHARD_RESERVE=1:77 into
the full-scope unit-test step (gated on matrix.scope == 'full', so the
unrelated windows lane is unaffected) -- 77 weight units is the empirical
conversion of the aux suites' ~220s measured fixed cost, derived against
the real tests/test-timings.json (see the diagnosis artifact for the full
computation).

Also corrects two pieces of now-stale bookkeeping this issue exposed:
- ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test`
  carried a 7m12s figure that predated the aux-suite growth; corrected to
  the real pre-fix evidence (13m48s / cancelled at ~14m51s, issue #4070),
  which requires raising timeout-minutes from 15 to 21 (1.5x headroom over
  the real worst-case measurement) to satisfy the file's own policy.
- test.yml's job-header and matrix comments, which still claimed the aux
  suites cost "~1m35s combined" (they now measure ~216-224s).

Closes the gap the previous commit's failing-first tests proved: selectShard
had no reserve-capacity mechanism and test.yml wired none in.

* fix(#4070): correct the reserved-weight property bound; cover main()'s reserve bounds check

Isolated code review found a genuine gap and gsd-test's real run confirmed a real
test bug it exposed:

- The fast-check property "no shard exceeds average(+reserve) + heaviest file" was
  falsified by gsd-test itself (weights=[1,1,1], total=2, reserve=6 on bin 0):
  selectShard is correct, the BOUND was wrong. A reserve large enough that its bin
  never receives a real item stays at exactly that reserve forever -- no amount of
  routing real items elsewhere can dilute a fixed head start below itself -- so the
  true bound is max(reserve, the classic Graham term), not the Graham term alone.
  Verified the corrected bound against the exact counterexample plus 20,000
  additional random trials (zero violations) before re-running gsd-test.

- Isolated review (MAJOR): the shard-total bounds check on RUN_TESTS_SHARD_RESERVE
  and its console.error fallback in main() were untested end-to-end --
  parseShardReserve itself has no concept of the shard total, so only main()
  enforces that guard, and nothing exercised it through the subprocess seam. Added
  an E2E harness test that sets RUN_TESTS_SHARD_RESERVE to an out-of-range index
  via the real CLI, asserts the fallback warning fires, AND asserts the resulting
  file selection is byte-identical to a no-reserve control run against the same
  injected timings table -- proving the reserve was actually ignored, not just
  that a warning printed.

* chore(#4070): backfill changeset PR number

pr:0 -> pr:4072

* fix(#4070): strip leaked RUN_TESTS_SHARD_RESERVE from the harness test's child env

Real GH Actions CI on this PR (run 33288554040, ubuntu shard 2/3) failed 7
tests in the shard-partitioning describe block, all with the same symptom:
`run-tests: no tests in suite "all"` where a real file count was expected.
gsd-test's own dockerized bench run never showed this, and ubuntu shards 1/3
and 3/3 (which run the same test.yml step) passed clean -- the discrepancy is
the tell: only shard 2/3 happened to schedule this specific test FILE for
that run, and the outer CI job's own environment is where the leak lives.

Root cause: test.yml's "Run unit tests" step now sets
RUN_TESTS_SHARD_RESERVE=1:77 (this issue's own reserve mechanism) on the
OUTER job that runs `npm run test:coverage:unit:raw -- --shard N`. The
harness test file's runHarness() helper spawns run-tests.cjs as a CHILD of
that same job via `{...process.env, ...extraEnv}`, so every pre-existing
--shard test in this describe block silently inherited the ambient reserve
-- even though none of them know it exists. A reserve of 77 weight units
utterly dwarfs the ~0.3 total weight of the 9-file synthetic fixtures these
tests use (none are in the real timings table, so all fall back to the same
tiny median weight), so shard index 1 is routed zero files every time --
exactly the observed "no tests" failures, and exactly the skewed 5/4 split
observed on the shard-2 test that expected a plain 3/3/3 round-robin.

Reproduced locally end to end (set RUN_TESTS_SHARD_RESERVE=1:77, spawn the
old runHarness against a synthetic 9-file fixture, --shard 1/3 -- reproduces
the exact "no tests in suite \"all\"" stderr) and confirmed the fix (env
stripped unless a test opts in via extraEnv, as the #4070 E2E bounds-check
test already does) resolves it, before re-running gsd-test.

This is a genuine bug this PR introduced -- a new ambient env var that a
pre-existing subprocess-spawning test helper didn't know to isolate against
-- not a pre-existing flake and not resource contention.

---------

Co-authored-by: sim <sim@local>
2026-08-30 00:43:02 -04:00

339 B

type, pr
type pr
Fixed 4072

CI shard 1 no longer runs at 92-99% of its timeout cap. The full-scope unit-test shard balancer now reserves shard 1's fixed aux-suite cost (integration/security/install/slow) before packing unit-test files onto it, instead of leaving shard 1 to carry that cost on top of an equal unit-test share. (#4070)