* test(#4070): failing-first regression for shard1 aux-suite budget imbalance
- selectShard has no way to reserve virtual weight on a bin, so the LPT
unit-test packer cannot account for shard 1's fixed aux-suite cost
(integration/security/install/slow all pinned to shard 1/3).
- test.yml wires no such reserve into the workflow.
- ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test`
was stale (7m12s from run 30677442953, predating the aux-suite growth);
corrected to the real evidence cited in #4070 (13m48s / cancelled at
~14m51s), which now honestly fails the file's own 1.5x headroom policy
against the current 15-minute cap.
All three are expected RED on this commit; see
.gsd/bug/fix-4070-shard1-aux-suite-budget/50-test-matrix.md.
* fix(#4070): reserve shard 1's aux-suite cost out of the LPT unit-test packer
selectShard now accepts an optional initialWeights array giving one or more
bins a virtual head start before any file is placed, so LPT converges each
bin's FINAL total (assigned weight + head start) toward equal instead of
raw assigned weight alone. test.yml wires RUN_TESTS_SHARD_RESERVE=1:77 into
the full-scope unit-test step (gated on matrix.scope == 'full', so the
unrelated windows lane is unaffected) -- 77 weight units is the empirical
conversion of the aux suites' ~220s measured fixed cost, derived against
the real tests/test-timings.json (see the diagnosis artifact for the full
computation).
Also corrects two pieces of now-stale bookkeeping this issue exposed:
- ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test`
carried a 7m12s figure that predated the aux-suite growth; corrected to
the real pre-fix evidence (13m48s / cancelled at ~14m51s, issue #4070),
which requires raising timeout-minutes from 15 to 21 (1.5x headroom over
the real worst-case measurement) to satisfy the file's own policy.
- test.yml's job-header and matrix comments, which still claimed the aux
suites cost "~1m35s combined" (they now measure ~216-224s).
Closes the gap the previous commit's failing-first tests proved: selectShard
had no reserve-capacity mechanism and test.yml wired none in.
* fix(#4070): correct the reserved-weight property bound; cover main()'s reserve bounds check
Isolated code review found a genuine gap and gsd-test's real run confirmed a real
test bug it exposed:
- The fast-check property "no shard exceeds average(+reserve) + heaviest file" was
falsified by gsd-test itself (weights=[1,1,1], total=2, reserve=6 on bin 0):
selectShard is correct, the BOUND was wrong. A reserve large enough that its bin
never receives a real item stays at exactly that reserve forever -- no amount of
routing real items elsewhere can dilute a fixed head start below itself -- so the
true bound is max(reserve, the classic Graham term), not the Graham term alone.
Verified the corrected bound against the exact counterexample plus 20,000
additional random trials (zero violations) before re-running gsd-test.
- Isolated review (MAJOR): the shard-total bounds check on RUN_TESTS_SHARD_RESERVE
and its console.error fallback in main() were untested end-to-end --
parseShardReserve itself has no concept of the shard total, so only main()
enforces that guard, and nothing exercised it through the subprocess seam. Added
an E2E harness test that sets RUN_TESTS_SHARD_RESERVE to an out-of-range index
via the real CLI, asserts the fallback warning fires, AND asserts the resulting
file selection is byte-identical to a no-reserve control run against the same
injected timings table -- proving the reserve was actually ignored, not just
that a warning printed.
* chore(#4070): backfill changeset PR number
pr:0 -> pr:4072
* fix(#4070): strip leaked RUN_TESTS_SHARD_RESERVE from the harness test's child env
Real GH Actions CI on this PR (run 33288554040, ubuntu shard 2/3) failed 7
tests in the shard-partitioning describe block, all with the same symptom:
`run-tests: no tests in suite "all"` where a real file count was expected.
gsd-test's own dockerized bench run never showed this, and ubuntu shards 1/3
and 3/3 (which run the same test.yml step) passed clean -- the discrepancy is
the tell: only shard 2/3 happened to schedule this specific test FILE for
that run, and the outer CI job's own environment is where the leak lives.
Root cause: test.yml's "Run unit tests" step now sets
RUN_TESTS_SHARD_RESERVE=1:77 (this issue's own reserve mechanism) on the
OUTER job that runs `npm run test:coverage:unit:raw -- --shard N`. The
harness test file's runHarness() helper spawns run-tests.cjs as a CHILD of
that same job via `{...process.env, ...extraEnv}`, so every pre-existing
--shard test in this describe block silently inherited the ambient reserve
-- even though none of them know it exists. A reserve of 77 weight units
utterly dwarfs the ~0.3 total weight of the 9-file synthetic fixtures these
tests use (none are in the real timings table, so all fall back to the same
tiny median weight), so shard index 1 is routed zero files every time --
exactly the observed "no tests" failures, and exactly the skewed 5/4 split
observed on the shard-2 test that expected a plain 3/3/3 round-robin.
Reproduced locally end to end (set RUN_TESTS_SHARD_RESERVE=1:77, spawn the
old runHarness against a synthetic 9-file fixture, --shard 1/3 -- reproduces
the exact "no tests in suite \"all\"" stderr) and confirmed the fix (env
stripped unless a test opts in via extraEnv, as the #4070 E2E bounds-check
test already does) resolves it, before re-running gsd-test.
This is a genuine bug this PR introduced -- a new ambient env var that a
pre-existing subprocess-spawning test helper didn't know to isolate against
-- not a pre-existing flake and not resource contention.
---------
Co-authored-by: sim <sim@local>