* test(#4070): failing-first regression for shard1 aux-suite budget imbalance - selectShard has no way to reserve virtual weight on a bin, so the LPT unit-test packer cannot account for shard 1's fixed aux-suite cost (integration/security/install/slow all pinned to shard 1/3). - test.yml wires no such reserve into the workflow. - ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test` was stale (7m12s from run 30677442953, predating the aux-suite growth); corrected to the real evidence cited in #4070 (13m48s / cancelled at ~14m51s), which now honestly fails the file's own 1.5x headroom policy against the current 15-minute cap. All three are expected RED on this commit; see .gsd/bug/fix-4070-shard1-aux-suite-budget/50-test-matrix.md. * fix(#4070): reserve shard 1's aux-suite cost out of the LPT unit-test packer selectShard now accepts an optional initialWeights array giving one or more bins a virtual head start before any file is placed, so LPT converges each bin's FINAL total (assigned weight + head start) toward equal instead of raw assigned weight alone. test.yml wires RUN_TESTS_SHARD_RESERVE=1:77 into the full-scope unit-test step (gated on matrix.scope == 'full', so the unrelated windows lane is unaffected) -- 77 weight units is the empirical conversion of the aux suites' ~220s measured fixed cost, derived against the real tests/test-timings.json (see the diagnosis artifact for the full computation). Also corrects two pieces of now-stale bookkeeping this issue exposed: - ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test` carried a 7m12s figure that predated the aux-suite growth; corrected to the real pre-fix evidence (13m48s / cancelled at ~14m51s, issue #4070), which requires raising timeout-minutes from 15 to 21 (1.5x headroom over the real worst-case measurement) to satisfy the file's own policy. - test.yml's job-header and matrix comments, which still claimed the aux suites cost "~1m35s combined" (they now measure ~216-224s). Closes the gap the previous commit's failing-first tests proved: selectShard had no reserve-capacity mechanism and test.yml wired none in. * fix(#4070): correct the reserved-weight property bound; cover main()'s reserve bounds check Isolated code review found a genuine gap and gsd-test's real run confirmed a real test bug it exposed: - The fast-check property "no shard exceeds average(+reserve) + heaviest file" was falsified by gsd-test itself (weights=[1,1,1], total=2, reserve=6 on bin 0): selectShard is correct, the BOUND was wrong. A reserve large enough that its bin never receives a real item stays at exactly that reserve forever -- no amount of routing real items elsewhere can dilute a fixed head start below itself -- so the true bound is max(reserve, the classic Graham term), not the Graham term alone. Verified the corrected bound against the exact counterexample plus 20,000 additional random trials (zero violations) before re-running gsd-test. - Isolated review (MAJOR): the shard-total bounds check on RUN_TESTS_SHARD_RESERVE and its console.error fallback in main() were untested end-to-end -- parseShardReserve itself has no concept of the shard total, so only main() enforces that guard, and nothing exercised it through the subprocess seam. Added an E2E harness test that sets RUN_TESTS_SHARD_RESERVE to an out-of-range index via the real CLI, asserts the fallback warning fires, AND asserts the resulting file selection is byte-identical to a no-reserve control run against the same injected timings table -- proving the reserve was actually ignored, not just that a warning printed. * chore(#4070): backfill changeset PR number pr:0 -> pr:4072 * fix(#4070): strip leaked RUN_TESTS_SHARD_RESERVE from the harness test's child env Real GH Actions CI on this PR (run 33288554040, ubuntu shard 2/3) failed 7 tests in the shard-partitioning describe block, all with the same symptom: `run-tests: no tests in suite "all"` where a real file count was expected. gsd-test's own dockerized bench run never showed this, and ubuntu shards 1/3 and 3/3 (which run the same test.yml step) passed clean -- the discrepancy is the tell: only shard 2/3 happened to schedule this specific test FILE for that run, and the outer CI job's own environment is where the leak lives. Root cause: test.yml's "Run unit tests" step now sets RUN_TESTS_SHARD_RESERVE=1:77 (this issue's own reserve mechanism) on the OUTER job that runs `npm run test:coverage:unit:raw -- --shard N`. The harness test file's runHarness() helper spawns run-tests.cjs as a CHILD of that same job via `{...process.env, ...extraEnv}`, so every pre-existing --shard test in this describe block silently inherited the ambient reserve -- even though none of them know it exists. A reserve of 77 weight units utterly dwarfs the ~0.3 total weight of the 9-file synthetic fixtures these tests use (none are in the real timings table, so all fall back to the same tiny median weight), so shard index 1 is routed zero files every time -- exactly the observed "no tests" failures, and exactly the skewed 5/4 split observed on the shard-2 test that expected a plain 3/3/3 round-robin. Reproduced locally end to end (set RUN_TESTS_SHARD_RESERVE=1:77, spawn the old runHarness against a synthetic 9-file fixture, --shard 1/3 -- reproduces the exact "no tests in suite \"all\"" stderr) and confirmed the fix (env stripped unless a test opts in via extraEnv, as the #4070 E2E bounds-check test already does) resolves it, before re-running gsd-test. This is a genuine bug this PR introduced -- a new ambient env var that a pre-existing subprocess-spawning test helper didn't know to isolate against -- not a pre-existing flake and not resource contention. --------- Co-authored-by: sim <sim@local>
339 B
339 B
type, pr
| type | pr |
|---|---|
| Fixed | 4072 |
CI shard 1 no longer runs at 92-99% of its timeout cap. The full-scope unit-test shard balancer now reserves shard 1's fixed aux-suite cost (integration/security/install/slow) before packing unit-test files onto it, instead of leaving shard 1 to carry that cost on top of an equal unit-test share. (#4070)