f7df920681f233ae0fe064ee659550bdf41ff708
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bf0d715733 |
fix(#2472): cost-balanced test sharding and pinned CI base commit (#2480)
* fix(#2472): weight-aware shard partition Windows shard 1/3 hit the 20-minute job cap with no failing assertion. Root cause is the shard layer, not the chunk layer: selectShard partitioned by sorted ARRAY INDEX (k % n, #1212), which balances file COUNTS and ignores file COST. On the real unit suite that produced 12.4m / 19.2m / 15.2m — a 1.23x max/ideal ratio leaving the heaviest shard 5% under the cap. Because assignment keyed off position, inserting one test file re-indexed every file after it and could tip that shard over; deterministic, so a re-run reproduced it exactly. This is NOT the chunk packer (#2456/#2463). That fix works and applies one level down, WITHIN a shard. The across-shard partition predated it and never consumed the cost table. Both layers now share one cost model. selectShard takes an optional weightOf and, when given one, partitions by LPT (longest-processing-time-first) — the same algorithm packChunks uses. Omitting it keeps the legacy round-robin byte-identical, so every existing test above still exercises that path unchanged and callers without timing data lose nothing. A missing timings table yields uniform weight 1, under which LPT degenerates to the equal-count split. Projected on the real suite: 16.4/17.3/13.0 -> 15.6/15.6/15.6 (worst shard 17.3m -> 15.6m). Tests: a skewed-cost regression (round-robin clusters all four heavy files onto one shard at 2.98x ideal; LPT does not), back-compat equivalence, determinism, tie-breaking, order preservation, and two fast-check properties — the partition is exhaustive and disjoint (getting this wrong silently DROPS tests from CI, the worst failure mode for a harness), and no shard exceeds average + heaviest file. Two assertions were corrected during authoring rather than shipped wrong: - an initial "LPT within 4/3 of ideal" bound was false. The 4/3 figure is relative to the OPTIMAL makespan, not the average, and the two differ when item sizes force a pairing. Replaced with Graham's average+max bound, which is what is actually provable. - "weighted is never worse than round-robin" is also false; fast-check falsified it with [19316,10190,1,9128,29353,20227] over 2 shards (rr 48670, lpt 48671). Round-robin can win by luck on a specific input. Dropped, with the counterexample recorded in place so it is not re-asserted later. Closes #2472 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): rotate tied bins; restore #1212 test block; lazy cost table Isolated-review findings, all fixed. HIGH — zero weights collapsed the whole partition onto shard 1. The lightest-bin scan compared weight only, and adding a zero-weight file leaves its bin's weight unchanged, so bin 0 stayed tied-minimum forever and every such file landed on it. Verified: all-zero weights gave shard1=[a..f], shard2=[], shard3=[] — two of three CI runners idle while one ran everything. Reachable through safeWeight's own clamp (a NaN/negative/Infinity entry in a corrupted or hand-edited timings table) and through any genuine 0ms measurement, so the clamp reproduced the exact failure its comment claimed to prevent. Ties now break on file COUNT after weight, which rotates. Pinned by two regression tests (all-zero, and clamped NaN/negative/Infinity) plus a property over list size x shard count. The live table has no 0ms entries (min 19ms), so production was not affected — but nothing prevented it. MEDIUM — the new describe block had swallowed #1212's pre-existing property test, which is why a test under a "weight-aware" heading never passed a weigher. That was a bad block boundary in the previous commit, not a bad test: the #2472 describe was opened before #1212's last test instead of after. Moved back where it belongs; #1212 is 762-879 and #2472 is 894-1082. LOW — that relocated property test ran unseeded. Seeded (12120) per the repo's property-test convention so a failure reproduces. Verified passing under the new seed. LOW — hoisting the timings load above the shard block charged a readFileSync + JSON.parse to invocations that exit before needing it (empty selection, --files matching nothing). Now lazily memoized, so neither consumer reads the table unless it is used and it is still read at most once. Real-suite projection unchanged at 15.6m / 15.6m / 15.6m. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#2472): correct stale round-robin sharding descriptions The partition is now cost-balanced, so the header block in run-tests.cjs and the two comments in test.yml describing '--shard' as a round-robin over sorted file index were actively wrong. Updated to describe LPT over measured duration, and to state the degenerate case explicitly: with no timing data every file weighs the same and the partition collapses back to k % n, which is why the pre-existing #1212 CLI tests still pass unchanged (their nine synthetic files are absent from the timings table, so all take the identical median weight). Remaining 'round-robin' mentions are correct — they describe the unweighted fallback path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): shard diagnostics, cost-routing E2E test, table validation Second orthogonal review (operational lens) findings, all fixed. HIGH — cross-runner partition divergence. Each of the up-to-12 CI jobs runs its own 'merge base into head' and computes its own partition, so if the inputs differ between jobs (the file list, or the timings table) two jobs can place the same file in different shards or in none. Every job stays internally exhaustive and disjoint, so nothing errors: a test simply never runs and CI stays green. The risk class is pre-existing — round-robin diverges identically when the file set differs between jobs, which is literally this issue's insertion instability — but weighting adds tests/test-timings.json as a second input that must match, so it widens the hole. Properly closing it means pinning the partition inputs per run, a workflow change beyond this fix. What IS closed here is the silence. Each shard now prints an input fingerprint over the FULL pre-partition list and the weight assigned to each file — deliberately not this shard's slice, which would differ by design and be useless for comparison. All shard jobs of one run must print an identical sig; a mismatch is direct proof the runners disagreed about the input. Verified: three independent computations agree, and the sig changes when the input drifts by one file. MEDIUM — nothing proved main() actually threads fileWeightOf() into selectShard. Every pre-existing --shard E2E test uses synthetic filenames absent from the real table, so all collapse to a uniform median weight, under which LPT is mathematically identical to k % n — a typo on that one wiring line would have passed the whole suite. Added an E2E test that injects a table via RUN_TESTS_TIMINGS_FILE with differing costs, placing the heavy files at exactly the indices round-robin hands to shard 1, and asserts shard 1 does NOT receive all three. Plus a test that all three shards emit the same sig. MEDIUM/LOW — no observability. The diagnostic line now reports files, weighed count, aggregate weight, and whether the table loaded, so a table that silently failed to parse shows table=absent/weighed=0 instead of being indistinguishable from a healthy load. (The reviewer confirmed the advisory fallback is already live on next: feat-2296-provider-escalation.test.cjs is missing from the table.) LOW — typeof [] === 'object', so a hand-edit turning the map into a list was accepted as a valid table. Now rejected via Array.isArray, falling back to uniform weight like any other malformed table. LOW — stale round-robin wording in ci-test-scope.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): pin every CI job to one base commit Closes the cross-runner divergence at its source instead of only making it visible. Each job of a run executes the rebase-check step independently, minutes apart across a 12-job matrix, and merged the MOVING origin/<branch> ref. If the base advanced mid-run, different jobs merged different trees. That was survivable when jobs only had to agree on pass/fail; it is not once they must agree on a PARTITION. Each shard job computes the whole split and keeps its own slice, so jobs working from different trees can place a file in two shards or in none — and every job still looks internally consistent, so nothing errors. A test silently never runs and CI stays green. ci-rebase-check.cjs now accepts CI_REBASE_BASE_SHA and pins BOTH the fetch and the merge to that one commit, so the two can never disagree. test.yml passes github.event.pull_request.base.sha on all three rebase-check steps; that value is fixed for the life of a run, so all jobs merge the identical base. This also closes the PRE-EXISTING half of the divergence. Round-robin had the same exposure whenever the test-file set differed between jobs — that is this issue's insertion instability — so the pin fixes the older hole too, not just the timings-table input weighting added. Only a full 40-hex sha is accepted; empty (push/workflow_dispatch), malformed, or injected values fall back to the branch ref rather than handing an arbitrary string to git fetch as a refspec. resolveBaseRefs is extracted pure and exported, and runMain is guarded behind require.main === module, so the pin contract is testable without spawning git. Tests (tests/ci-test-scope.test.cjs): every rebase-check step must carry the pin; a valid sha pins both refs; absence falls back correctly; and five hostile values — short sha, uppercase, --upload-pack= injection, ref expression, empty — are each rejected. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f729101eec |
refactor(scripts): replace process.exit() with ExitError + runMain handler (#739) (#740)
Part 1 of 2 of the n/no-process-exit cleanup (umbrella #738): convert every process.exit() call in standalone scripts/** CLIs to the rule-compliant pattern. - New shared helper scripts/lib/cli-exit.cjs: ExitError(code,message) + runMain() which translates a thrown ExitError / returned number into process.exitCode (never process.exit()), flushing output and still firing process.on('exit'). - main()-based entrypoints: throw new ExitError(code) for errors, return <code> for verdicts; invoked via runMain(main). Child exit codes preserved via return. - top-level-only scripts: imperative body extracted into main() so mid-flow aborts (throw ExitError) actually halt; pure consts/helpers stay at module scope. - diff-touches-shipped-paths.cjs: stdin event handling restructured to an async read so the whole flow runs under runMain; uncaughtException/unhandledRejection nets replaced by an in-band catch that preserves EXIT_ERROR=2. Exit codes verified unchanged for every converted script (success/error/help and the 0/1/2 semantic codes in diff-touches). Rule stays warn here; flipped to error in part 2 (#738) once gsd-core/bin/** is also clean. Refs #739 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ba231ecbfc |
chore: clean up clear-cut ESLint warnings (#732) (#734)
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts). No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving. Closes #732 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
48b1e35187 |
fix(#431): enforce H1 shell policy (linux=bash, macOS=zsh, windows=pwsh) across PR + release gates (#434)
* test(#431): policy-shell-pinning linter — RED baseline (37 violations on origin/next) Adds scripts/workflow-policy.cjs: H1 shell-policy linter with POLICY map, VIOLATION enum, matrix expansion, effective-shell resolution order, and runPolicyLint({ workflowsDir }) entry point. Adds tests/policy-shell-pinning.test.cjs: 8 tests (baseline + 6 synthetic counter-tests). Synthetic tests 2–7 pass; baseline test is intentionally RED (37 violations: 28 in test.yml, 9 in install-smoke.yml — all macos/windows lanes using shell: bash instead of native zsh/pwsh). Adds js-yaml@4.1.1 as devDependency for YAML parsing. * fix(#431): switch ubuntu/windows lanes to native shells; extract bash-isms to Node Remove all explicit shell: bash pins from ubuntu-only jobs (changes, lint-tests, coverage, required-tests, smoke-unpacked) — ubuntu runner default is bash, which is both H1-compliant and the runner default, making the pin redundant. For the test and test-full mixed-OS jobs (ubuntu+windows, windows+macos): - Move bash-ism steps to shell-agnostic Node scripts: scripts/ci-guard-runner.cjs — RUNNER_ENVIRONMENT check scripts/ci-rebase-check.cjs — git fetch+merge PR base branch scripts/check-npm-integrity.cjs — Node port of check-npm-integrity.sh scripts/ci-prepare-test-scope.cjs — write .ci-selected-tests.txt scripts/ci-smoke-skip.cjs — set skip= output for full-only matrix entries - Remove shell: bash from simple npm/node command steps (runner default applies) This brings Windows violations from 19 to 0. Remaining 17 violations are all MACOS_MISSING_EXPLICIT_ZSH in mixed-OS matrix jobs (test-full: windows+macos, install-smoke smoke: ubuntu+macos) — these require job splitting to fix; see BLOCKER in PR description. * fix(#431): update workflow-shell-pinning test for H1 policy The old test required all Windows-targeting npm steps to pin shell: bash (to prevent pwsh stderr-swallow). Under H1, Windows runners must use pwsh (native, no pin needed) — shell: bash on Windows is now the violation, not the fix. Update findViolations() to flag npm steps with effectiveShell === 'bash' (rather than effectiveShell === null). Update synthetic tests to verify the H1-inverted semantics: defaults.run.shell: bash on Windows is now 2 violations, not 0. Update test name and assertion messages to describe the H1 constraint rather than the old missing-pin constraint. * fix(#431): extend policy linter to resolve matrix.shell expressions - expandRunsOn now captures all matrix.include row keys as realization context (os, node-version, shell, full_only, etc.) instead of only os - effectiveShell now accepts a realizationContext and resolves ${{ matrix.<key> }} expressions against it before checking policy - Unresolvable matrix key in shell expression emits UNRESOLVABLE_MATRIX - Add 3 new tests: positive (zsh+pwsh per row → 0 violations), counter (bash in macOS row → WRONG_SHELL_FOR_OS), counter (missing shell key → UNRESOLVABLE_MATRIX) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): apply matrix.shell pattern to test-full and smoke jobs (clears BLOCKER) test-full job (test.yml): - Add shell: pwsh/zsh per matrix.include row (windows-latest→pwsh, macos-latest→zsh) - Add job-level defaults.run.shell: ${{ matrix.shell }} - No step-level shell pins existed to remove smoke job (install-smoke.yml): - Add shell: bash/zsh per matrix.include row (ubuntu→bash, macos→zsh) - Add job-level defaults.run.shell: ${{ matrix.shell }} - No step-level shell pins existed to remove Policy linter now reports 0 violations across all workflow files. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#431): migrate .sh check scripts to .cjs; remove .sh originals - Add scripts/check-env.cjs: Node.js port of check-env.sh with identical exit codes (0/1/2), human-readable and --json output, --help flag, and all 5 checks (node-version, npm-version, lockfile-present, lockfile-sync, version-manager-pin) - Migrate all callers: - package.json check:env → node scripts/check-env.cjs - package.json check:integrity → node scripts/check-npm-integrity.cjs - scripts/ci-test-scope.cjs path strings → .cjs equivalents - .github/workflows/release.yml rc+finalize jobs → node .cjs (drop chmod+x) - .github/workflows/security-scan.yml → node .cjs (drop chmod+x) - tests/check-env.test.cjs → spawn node process.execPath [.cjs] - tests/npm-integrity-gate.test.cjs → spawn node process.execPath [.cjs] - Delete scripts/check-env.sh and scripts/check-npm-integrity.sh Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#431): update doc references from .sh to .cjs Update SECURITY.md and docs/contributing/bootstrap.md to reference the canonical Node invocation instead of the removed bash scripts. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): use per-step shell:matrix.shell instead of defaults.run.shell (GHA compat) GHA does not reliably resolve matrix expressions inside defaults.run.shell. Per-step shell: always resolves correctly. Removed the defaults.run.shell block from the test-full job (test.yml) and the smoke job (install-smoke.yml), and added shell: \${{ matrix.shell }} directly on every run: step in both jobs. Codex finding: defaults.run.shell with matrix expressions is not a GHA-supported pattern; per-step shell: is the safe form. * fix(#431): policy linter validates every matrix.include row independently Removed runner-label-only dedup from expandRunsOn() in workflow-policy.cjs. The prior guard (if !realizations.find(r => r.runner === runner)) collapsed two macos-latest rows with different node-version/shell contexts into one, hiding the second row's policy violation. Each matrix.include row is a distinct CI realization with its own context; validating it twice is harmless but skipping it causes false negatives. Added counter-test (Test 8) in tests/policy-shell-pinning.test.cjs: two macos-latest rows (shell:zsh compliant + shell:bash violation) must produce exactly one WRONG_SHELL_FOR_OS violation on the second row. * fix(#431): remove dedup-by-runner in Cartesian matrix.<key> expansion (Codex round 3) The base-list path in expandRunsOn (matrix.<key> arrays, e.g. matrix.os) previously guarded each push with `if (!realizations.find(r => r.runner === runner))`, collapsing duplicate runner values into a single realization and hiding policy violations on later rows of a Cartesian matrix. Remove the guard unconditionally; each entry in the base-list array now produces its own realization, matching the same fix already applied to the matrix.include path. Add counter-test "Cartesian matrix os × shell — dedup must not collapse rows by runner alone": matrix.os: [macos-latest, macos-latest] + shell: ${{ matrix.shell }} now yields 2 realizations (not 1). Documents that Cartesian cross-product expansion (carrying all keys into realization context) is a separate follow-up; current violations are UNRESOLVABLE_MATRIX pending that work. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): remove 60s timeout regression on npm ci --dry-run (parity with check-env.sh) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): ci-rebase-check.cjs — return truthy sentinel on success (Codex round 4) run() used execFileSync with stdio:'inherit', which returns null on success. Caller checked `result !== null`, always false → every successful fetch fell through to "failed after 3 attempts" exit-1 path. Fix: run() now returns true on success, false on failure. Update caller from `result !== null` to `if (result)`. Adds tests/ci-rebase-check.test.cjs (5 tests) covering the sentinel contract and a local-bare-remote integration smoke that verifies the full fetch+merge path exits 0 when fetch succeeds. --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> |