* fix(#2472): weight-aware shard partition Windows shard 1/3 hit the 20-minute job cap with no failing assertion. Root cause is the shard layer, not the chunk layer: selectShard partitioned by sorted ARRAY INDEX (k % n, #1212), which balances file COUNTS and ignores file COST. On the real unit suite that produced 12.4m / 19.2m / 15.2m — a 1.23x max/ideal ratio leaving the heaviest shard 5% under the cap. Because assignment keyed off position, inserting one test file re-indexed every file after it and could tip that shard over; deterministic, so a re-run reproduced it exactly. This is NOT the chunk packer (#2456/#2463). That fix works and applies one level down, WITHIN a shard. The across-shard partition predated it and never consumed the cost table. Both layers now share one cost model. selectShard takes an optional weightOf and, when given one, partitions by LPT (longest-processing-time-first) — the same algorithm packChunks uses. Omitting it keeps the legacy round-robin byte-identical, so every existing test above still exercises that path unchanged and callers without timing data lose nothing. A missing timings table yields uniform weight 1, under which LPT degenerates to the equal-count split. Projected on the real suite: 16.4/17.3/13.0 -> 15.6/15.6/15.6 (worst shard 17.3m -> 15.6m). Tests: a skewed-cost regression (round-robin clusters all four heavy files onto one shard at 2.98x ideal; LPT does not), back-compat equivalence, determinism, tie-breaking, order preservation, and two fast-check properties — the partition is exhaustive and disjoint (getting this wrong silently DROPS tests from CI, the worst failure mode for a harness), and no shard exceeds average + heaviest file. Two assertions were corrected during authoring rather than shipped wrong: - an initial "LPT within 4/3 of ideal" bound was false. The 4/3 figure is relative to the OPTIMAL makespan, not the average, and the two differ when item sizes force a pairing. Replaced with Graham's average+max bound, which is what is actually provable. - "weighted is never worse than round-robin" is also false; fast-check falsified it with [19316,10190,1,9128,29353,20227] over 2 shards (rr 48670, lpt 48671). Round-robin can win by luck on a specific input. Dropped, with the counterexample recorded in place so it is not re-asserted later. Closes #2472 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): rotate tied bins; restore #1212 test block; lazy cost table Isolated-review findings, all fixed. HIGH — zero weights collapsed the whole partition onto shard 1. The lightest-bin scan compared weight only, and adding a zero-weight file leaves its bin's weight unchanged, so bin 0 stayed tied-minimum forever and every such file landed on it. Verified: all-zero weights gave shard1=[a..f], shard2=[], shard3=[] — two of three CI runners idle while one ran everything. Reachable through safeWeight's own clamp (a NaN/negative/Infinity entry in a corrupted or hand-edited timings table) and through any genuine 0ms measurement, so the clamp reproduced the exact failure its comment claimed to prevent. Ties now break on file COUNT after weight, which rotates. Pinned by two regression tests (all-zero, and clamped NaN/negative/Infinity) plus a property over list size x shard count. The live table has no 0ms entries (min 19ms), so production was not affected — but nothing prevented it. MEDIUM — the new describe block had swallowed #1212's pre-existing property test, which is why a test under a "weight-aware" heading never passed a weigher. That was a bad block boundary in the previous commit, not a bad test: the #2472 describe was opened before #1212's last test instead of after. Moved back where it belongs; #1212 is 762-879 and #2472 is 894-1082. LOW — that relocated property test ran unseeded. Seeded (12120) per the repo's property-test convention so a failure reproduces. Verified passing under the new seed. LOW — hoisting the timings load above the shard block charged a readFileSync + JSON.parse to invocations that exit before needing it (empty selection, --files matching nothing). Now lazily memoized, so neither consumer reads the table unless it is used and it is still read at most once. Real-suite projection unchanged at 15.6m / 15.6m / 15.6m. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#2472): correct stale round-robin sharding descriptions The partition is now cost-balanced, so the header block in run-tests.cjs and the two comments in test.yml describing '--shard' as a round-robin over sorted file index were actively wrong. Updated to describe LPT over measured duration, and to state the degenerate case explicitly: with no timing data every file weighs the same and the partition collapses back to k % n, which is why the pre-existing #1212 CLI tests still pass unchanged (their nine synthetic files are absent from the timings table, so all take the identical median weight). Remaining 'round-robin' mentions are correct — they describe the unweighted fallback path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): shard diagnostics, cost-routing E2E test, table validation Second orthogonal review (operational lens) findings, all fixed. HIGH — cross-runner partition divergence. Each of the up-to-12 CI jobs runs its own 'merge base into head' and computes its own partition, so if the inputs differ between jobs (the file list, or the timings table) two jobs can place the same file in different shards or in none. Every job stays internally exhaustive and disjoint, so nothing errors: a test simply never runs and CI stays green. The risk class is pre-existing — round-robin diverges identically when the file set differs between jobs, which is literally this issue's insertion instability — but weighting adds tests/test-timings.json as a second input that must match, so it widens the hole. Properly closing it means pinning the partition inputs per run, a workflow change beyond this fix. What IS closed here is the silence. Each shard now prints an input fingerprint over the FULL pre-partition list and the weight assigned to each file — deliberately not this shard's slice, which would differ by design and be useless for comparison. All shard jobs of one run must print an identical sig; a mismatch is direct proof the runners disagreed about the input. Verified: three independent computations agree, and the sig changes when the input drifts by one file. MEDIUM — nothing proved main() actually threads fileWeightOf() into selectShard. Every pre-existing --shard E2E test uses synthetic filenames absent from the real table, so all collapse to a uniform median weight, under which LPT is mathematically identical to k % n — a typo on that one wiring line would have passed the whole suite. Added an E2E test that injects a table via RUN_TESTS_TIMINGS_FILE with differing costs, placing the heavy files at exactly the indices round-robin hands to shard 1, and asserts shard 1 does NOT receive all three. Plus a test that all three shards emit the same sig. MEDIUM/LOW — no observability. The diagnostic line now reports files, weighed count, aggregate weight, and whether the table loaded, so a table that silently failed to parse shows table=absent/weighed=0 instead of being indistinguishable from a healthy load. (The reviewer confirmed the advisory fallback is already live on next: feat-2296-provider-escalation.test.cjs is missing from the table.) LOW — typeof [] === 'object', so a hand-edit turning the map into a list was accepted as a valid table. Now rejected via Array.isArray, falling back to uniform weight like any other malformed table. LOW — stale round-robin wording in ci-test-scope.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): pin every CI job to one base commit Closes the cross-runner divergence at its source instead of only making it visible. Each job of a run executes the rebase-check step independently, minutes apart across a 12-job matrix, and merged the MOVING origin/<branch> ref. If the base advanced mid-run, different jobs merged different trees. That was survivable when jobs only had to agree on pass/fail; it is not once they must agree on a PARTITION. Each shard job computes the whole split and keeps its own slice, so jobs working from different trees can place a file in two shards or in none — and every job still looks internally consistent, so nothing errors. A test silently never runs and CI stays green. ci-rebase-check.cjs now accepts CI_REBASE_BASE_SHA and pins BOTH the fetch and the merge to that one commit, so the two can never disagree. test.yml passes github.event.pull_request.base.sha on all three rebase-check steps; that value is fixed for the life of a run, so all jobs merge the identical base. This also closes the PRE-EXISTING half of the divergence. Round-robin had the same exposure whenever the test-file set differed between jobs — that is this issue's insertion instability — so the pin fixes the older hole too, not just the timings-table input weighting added. Only a full 40-hex sha is accepted; empty (push/workflow_dispatch), malformed, or injected values fall back to the branch ref rather than handing an arbitrary string to git fetch as a refspec. resolveBaseRefs is extracted pure and exported, and runMain is guarded behind require.main === module, so the pin contract is testable without spawning git. Tests (tests/ci-test-scope.test.cjs): every rebase-check step must carry the pin; a valid sha pins both refs; absence falls back correctly; and five hostile values — short sha, uppercase, --upload-pack= injection, ref expression, empty — are each rejected. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
131 lines
5.0 KiB
JavaScript
131 lines
5.0 KiB
JavaScript
'use strict';
|
|
// ci-rebase-check.cjs — Merge the PR base branch into the current PR head.
|
|
// Replaces the inline bash "Rebase check — merge PR base branch into PR head" step.
|
|
// Shell-agnostic: invoked as `node scripts/ci-rebase-check.cjs` from any shell.
|
|
//
|
|
// Required environment variables (set by the workflow step's `env:` block):
|
|
// GITHUB_TOKEN — access token for remote set-url
|
|
// GITHUB_BASE_REF — PR base branch name (set by GitHub Actions on pull_request events)
|
|
// GITHUB_REPOSITORY — owner/repo (set by GitHub Actions)
|
|
//
|
|
// Optional:
|
|
// CI_REBASE_BASE_SHA — pin the merge to one exact base commit (#2472).
|
|
//
|
|
// Why the pin matters. Every job of a run executes this step independently, at
|
|
// whatever wall-clock moment it gets there — and Windows/macOS installs skew
|
|
// that by minutes across a 12-job matrix. Merging the moving `origin/<branch>`
|
|
// ref means that if the base advances mid-run, different jobs merge different
|
|
// trees. That was survivable when jobs only had to agree on pass/fail, but the
|
|
// sharded lane makes them agree on a PARTITION: each shard job computes the
|
|
// whole split and keeps its own slice, so jobs working from different trees can
|
|
// place a file in two shards or in none. Each job still looks internally
|
|
// consistent, so nothing errors — a test silently never runs and CI stays
|
|
// green. Pinning every job to `github.event.pull_request.base.sha`, which is
|
|
// fixed for the life of the run, removes the divergence at its source rather
|
|
// than detecting it after the fact.
|
|
//
|
|
// Exit 0 = merged cleanly (or merge was a no-op).
|
|
// Exit 1 = merge conflict or fetch failure.
|
|
|
|
const { execFileSync } = require('child_process');
|
|
|
|
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
|
|
|
|
function run(cmd, args, opts) {
|
|
try {
|
|
execFileSync(cmd, args, { stdio: 'inherit', ...opts });
|
|
return true; // success sentinel; execFileSync returns null with stdio:'inherit'
|
|
} catch (e) {
|
|
return false;
|
|
}
|
|
}
|
|
|
|
function runOrThrow(cmd, args, label) {
|
|
try {
|
|
execFileSync(cmd, args, { stdio: 'inherit' });
|
|
} catch (e) {
|
|
throw new ExitError(1, `::error::${label} failed`);
|
|
}
|
|
}
|
|
|
|
const token = process.env.GITHUB_TOKEN || '';
|
|
const baseBranch = process.env.GITHUB_BASE_REF || 'main';
|
|
const repo = process.env.GITHUB_REPOSITORY || '';
|
|
// Resolve what to fetch and what to merge, pinned together so they can never
|
|
// disagree. Pure and exported so the pin contract is testable without spawning
|
|
// git: env in, refs out.
|
|
//
|
|
// Only a full 40-hex sha is accepted. Anything else — empty on push/dispatch
|
|
// events, or a malformed/injected value — falls back to the branch ref,
|
|
// preserving the pre-#2472 behavior rather than handing an arbitrary string to
|
|
// `git fetch` as a refspec.
|
|
function resolveBaseRefs(env = process.env, fallbackBranch = 'main') {
|
|
const branch = env.GITHUB_BASE_REF || fallbackBranch;
|
|
const raw = env.CI_REBASE_BASE_SHA || '';
|
|
const sha = /^[0-9a-f]{40}$/.test(raw) ? raw : null;
|
|
return {
|
|
branch,
|
|
sha,
|
|
pinned: sha !== null,
|
|
fetchRef: sha || branch,
|
|
mergeRef: sha || `origin/${branch}`,
|
|
};
|
|
}
|
|
|
|
const { fetchRef, mergeRef } = resolveBaseRefs(process.env, 'main');
|
|
|
|
function main() {
|
|
// Configure git identity (needed for merge commit).
|
|
runOrThrow('git', ['config', 'user.email', 'ci@gsd-redux'], 'git config user.email');
|
|
runOrThrow('git', ['config', 'user.name', 'CI Rebase Check'], 'git config user.name');
|
|
|
|
// Set authenticated remote URL.
|
|
if (token && repo) {
|
|
runOrThrow(
|
|
'git',
|
|
['remote', 'set-url', 'origin', `https://x-access-token:${token}@github.com/${repo}.git`],
|
|
'git remote set-url'
|
|
);
|
|
}
|
|
|
|
// Fetch base branch with retry.
|
|
for (let attempt = 1; attempt <= 3; attempt++) {
|
|
const result = run('git', ['fetch', 'origin', fetchRef]);
|
|
if (result) {
|
|
break;
|
|
}
|
|
if (attempt === 3) {
|
|
throw new ExitError(1, `::error::git fetch origin ${fetchRef} failed after 3 attempts.`);
|
|
}
|
|
// Wait before retry: attempt * 4 seconds.
|
|
const waitMs = attempt * 4000;
|
|
const deadline = Date.now() + waitMs;
|
|
while (Date.now() < deadline) { /* busy wait, acceptable in CI */ }
|
|
}
|
|
|
|
// Attempt merge.
|
|
try {
|
|
execFileSync('git', ['merge', '--no-edit', '--no-ff', mergeRef], { stdio: 'inherit' });
|
|
} catch (e) {
|
|
process.stderr.write(
|
|
`::error::This PR cannot cleanly merge origin/${baseBranch}. Rebase your branch onto current ${baseBranch} and push again.\n`
|
|
);
|
|
process.stderr.write('::error::Conflicting files:\n');
|
|
try {
|
|
execFileSync('git', ['diff', '--name-only', '--diff-filter=U'], { stdio: 'inherit' });
|
|
} catch (_) { /* ignore */ }
|
|
try {
|
|
execFileSync('git', ['merge', '--abort'], { stdio: 'inherit' });
|
|
} catch (_) { /* ignore */ }
|
|
throw new ExitError(1);
|
|
}
|
|
}
|
|
|
|
// Only run when invoked as the CI step. Guarded so a test can require this
|
|
// module for resolveBaseRefs without firing git fetch/merge as a side effect.
|
|
if (require.main === module) {
|
|
runMain(main);
|
|
}
|
|
|
|
module.exports = { resolveBaseRefs };
|