Commit Graph

4 Commits

Author SHA1 Message Date
Tom Boucher
bf0d715733 fix(#2472): cost-balanced test sharding and pinned CI base commit (#2480)
* fix(#2472): weight-aware shard partition

Windows shard 1/3 hit the 20-minute job cap with no failing assertion. Root
cause is the shard layer, not the chunk layer: selectShard partitioned by
sorted ARRAY INDEX (k % n, #1212), which balances file COUNTS and ignores
file COST. On the real unit suite that produced 12.4m / 19.2m / 15.2m — a
1.23x max/ideal ratio leaving the heaviest shard 5% under the cap. Because
assignment keyed off position, inserting one test file re-indexed every file
after it and could tip that shard over; deterministic, so a re-run reproduced
it exactly.

This is NOT the chunk packer (#2456/#2463). That fix works and applies one
level down, WITHIN a shard. The across-shard partition predated it and never
consumed the cost table. Both layers now share one cost model.

selectShard takes an optional weightOf and, when given one, partitions by LPT
(longest-processing-time-first) — the same algorithm packChunks uses. Omitting
it keeps the legacy round-robin byte-identical, so every existing test above
still exercises that path unchanged and callers without timing data lose
nothing. A missing timings table yields uniform weight 1, under which LPT
degenerates to the equal-count split.

Projected on the real suite: 16.4/17.3/13.0 -> 15.6/15.6/15.6 (worst shard
17.3m -> 15.6m).

Tests: a skewed-cost regression (round-robin clusters all four heavy files
onto one shard at 2.98x ideal; LPT does not), back-compat equivalence,
determinism, tie-breaking, order preservation, and two fast-check properties
— the partition is exhaustive and disjoint (getting this wrong silently DROPS
tests from CI, the worst failure mode for a harness), and no shard exceeds
average + heaviest file.

Two assertions were corrected during authoring rather than shipped wrong:
- an initial "LPT within 4/3 of ideal" bound was false. The 4/3 figure is
  relative to the OPTIMAL makespan, not the average, and the two differ when
  item sizes force a pairing. Replaced with Graham's average+max bound, which
  is what is actually provable.
- "weighted is never worse than round-robin" is also false; fast-check
  falsified it with [19316,10190,1,9128,29353,20227] over 2 shards (rr 48670,
  lpt 48671). Round-robin can win by luck on a specific input. Dropped, with
  the counterexample recorded in place so it is not re-asserted later.

Closes #2472

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): rotate tied bins; restore #1212 test block; lazy cost table

Isolated-review findings, all fixed.

HIGH — zero weights collapsed the whole partition onto shard 1. The
lightest-bin scan compared weight only, and adding a zero-weight file leaves
its bin's weight unchanged, so bin 0 stayed tied-minimum forever and every
such file landed on it. Verified: all-zero weights gave shard1=[a..f],
shard2=[], shard3=[] — two of three CI runners idle while one ran everything.
Reachable through safeWeight's own clamp (a NaN/negative/Infinity entry in a
corrupted or hand-edited timings table) and through any genuine 0ms
measurement, so the clamp reproduced the exact failure its comment claimed to
prevent. Ties now break on file COUNT after weight, which rotates. Pinned by
two regression tests (all-zero, and clamped NaN/negative/Infinity) plus a
property over list size x shard count. The live table has no 0ms entries
(min 19ms), so production was not affected — but nothing prevented it.

MEDIUM — the new describe block had swallowed #1212's pre-existing property
test, which is why a test under a "weight-aware" heading never passed a
weigher. That was a bad block boundary in the previous commit, not a bad
test: the #2472 describe was opened before #1212's last test instead of
after. Moved back where it belongs; #1212 is 762-879 and #2472 is 894-1082.

LOW — that relocated property test ran unseeded. Seeded (12120) per the
repo's property-test convention so a failure reproduces. Verified passing
under the new seed.

LOW — hoisting the timings load above the shard block charged a readFileSync
+ JSON.parse to invocations that exit before needing it (empty selection,
--files matching nothing). Now lazily memoized, so neither consumer reads the
table unless it is used and it is still read at most once.

Real-suite projection unchanged at 15.6m / 15.6m / 15.6m.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2472): correct stale round-robin sharding descriptions

The partition is now cost-balanced, so the header block in run-tests.cjs
and the two comments in test.yml describing '--shard' as a round-robin over
sorted file index were actively wrong. Updated to describe LPT over measured
duration, and to state the degenerate case explicitly: with no timing data
every file weighs the same and the partition collapses back to k % n, which
is why the pre-existing #1212 CLI tests still pass unchanged (their nine
synthetic files are absent from the timings table, so all take the identical
median weight).

Remaining 'round-robin' mentions are correct — they describe the unweighted
fallback path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): shard diagnostics, cost-routing E2E test, table validation

Second orthogonal review (operational lens) findings, all fixed.

HIGH — cross-runner partition divergence. Each of the up-to-12 CI jobs runs
its own 'merge base into head' and computes its own partition, so if the
inputs differ between jobs (the file list, or the timings table) two jobs can
place the same file in different shards or in none. Every job stays
internally exhaustive and disjoint, so nothing errors: a test simply never
runs and CI stays green.

The risk class is pre-existing — round-robin diverges identically when the
file set differs between jobs, which is literally this issue's insertion
instability — but weighting adds tests/test-timings.json as a second input
that must match, so it widens the hole. Properly closing it means pinning the
partition inputs per run, a workflow change beyond this fix.

What IS closed here is the silence. Each shard now prints an input
fingerprint over the FULL pre-partition list and the weight assigned to each
file — deliberately not this shard's slice, which would differ by design and
be useless for comparison. All shard jobs of one run must print an identical
sig; a mismatch is direct proof the runners disagreed about the input.
Verified: three independent computations agree, and the sig changes when the
input drifts by one file.

MEDIUM — nothing proved main() actually threads fileWeightOf() into
selectShard. Every pre-existing --shard E2E test uses synthetic filenames
absent from the real table, so all collapse to a uniform median weight, under
which LPT is mathematically identical to k % n — a typo on that one wiring
line would have passed the whole suite. Added an E2E test that injects a
table via RUN_TESTS_TIMINGS_FILE with differing costs, placing the heavy
files at exactly the indices round-robin hands to shard 1, and asserts shard 1
does NOT receive all three. Plus a test that all three shards emit the same
sig.

MEDIUM/LOW — no observability. The diagnostic line now reports files,
weighed count, aggregate weight, and whether the table loaded, so a table
that silently failed to parse shows table=absent/weighed=0 instead of being
indistinguishable from a healthy load. (The reviewer confirmed the advisory
fallback is already live on next: feat-2296-provider-escalation.test.cjs is
missing from the table.)

LOW — typeof [] === 'object', so a hand-edit turning the map into a list was
accepted as a valid table. Now rejected via Array.isArray, falling back to
uniform weight like any other malformed table.

LOW — stale round-robin wording in ci-test-scope.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): pin every CI job to one base commit

Closes the cross-runner divergence at its source instead of only making it
visible.

Each job of a run executes the rebase-check step independently, minutes apart
across a 12-job matrix, and merged the MOVING origin/<branch> ref. If the base
advanced mid-run, different jobs merged different trees. That was survivable
when jobs only had to agree on pass/fail; it is not once they must agree on a
PARTITION. Each shard job computes the whole split and keeps its own slice, so
jobs working from different trees can place a file in two shards or in none —
and every job still looks internally consistent, so nothing errors. A test
silently never runs and CI stays green.

ci-rebase-check.cjs now accepts CI_REBASE_BASE_SHA and pins BOTH the fetch and
the merge to that one commit, so the two can never disagree. test.yml passes
github.event.pull_request.base.sha on all three rebase-check steps; that value
is fixed for the life of a run, so all jobs merge the identical base.

This also closes the PRE-EXISTING half of the divergence. Round-robin had the
same exposure whenever the test-file set differed between jobs — that is this
issue's insertion instability — so the pin fixes the older hole too, not just
the timings-table input weighting added.

Only a full 40-hex sha is accepted; empty (push/workflow_dispatch), malformed,
or injected values fall back to the branch ref rather than handing an arbitrary
string to git fetch as a refspec. resolveBaseRefs is extracted pure and
exported, and runMain is guarded behind require.main === module, so the pin
contract is testable without spawning git.

Tests (tests/ci-test-scope.test.cjs): every rebase-check step must carry the
pin; a valid sha pins both refs; absence falls back correctly; and five hostile
values — short sha, uppercase, --upload-pack= injection, ref expression, empty
— are each rejected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 09:15:59 -04:00
Tom Boucher
f729101eec refactor(scripts): replace process.exit() with ExitError + runMain handler (#739) (#740)
Part 1 of 2 of the n/no-process-exit cleanup (umbrella #738): convert every
process.exit() call in standalone scripts/** CLIs to the rule-compliant pattern.

- New shared helper scripts/lib/cli-exit.cjs: ExitError(code,message) + runMain()
  which translates a thrown ExitError / returned number into process.exitCode
  (never process.exit()), flushing output and still firing process.on('exit').
- main()-based entrypoints: throw new ExitError(code) for errors, return <code>
  for verdicts; invoked via runMain(main). Child exit codes preserved via return.
- top-level-only scripts: imperative body extracted into main() so mid-flow
  aborts (throw ExitError) actually halt; pure consts/helpers stay at module scope.
- diff-touches-shipped-paths.cjs: stdin event handling restructured to an async
  read so the whole flow runs under runMain; uncaughtException/unhandledRejection
  nets replaced by an in-band catch that preserves EXIT_ERROR=2.

Exit codes verified unchanged for every converted script (success/error/help and
the 0/1/2 semantic codes in diff-touches). Rule stays warn here; flipped to error
in part 2 (#738) once gsd-core/bin/** is also clean.

Refs #739

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 16:13:13 -04:00
Tom Boucher
ba231ecbfc chore: clean up clear-cut ESLint warnings (#732) (#734)
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts).

No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving.

Closes #732

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 11:24:48 -04:00
Tom Boucher
48b1e35187 fix(#431): enforce H1 shell policy (linux=bash, macOS=zsh, windows=pwsh) across PR + release gates (#434)
* test(#431): policy-shell-pinning linter — RED baseline (37 violations on origin/next)

Adds scripts/workflow-policy.cjs: H1 shell-policy linter with POLICY map,
VIOLATION enum, matrix expansion, effective-shell resolution order, and
runPolicyLint({ workflowsDir }) entry point.

Adds tests/policy-shell-pinning.test.cjs: 8 tests (baseline + 6 synthetic
counter-tests). Synthetic tests 2–7 pass; baseline test is intentionally RED
(37 violations: 28 in test.yml, 9 in install-smoke.yml — all macos/windows
lanes using shell: bash instead of native zsh/pwsh).

Adds js-yaml@4.1.1 as devDependency for YAML parsing.

* fix(#431): switch ubuntu/windows lanes to native shells; extract bash-isms to Node

Remove all explicit shell: bash pins from ubuntu-only jobs (changes, lint-tests,
coverage, required-tests, smoke-unpacked) — ubuntu runner default is bash, which
is both H1-compliant and the runner default, making the pin redundant.

For the test and test-full mixed-OS jobs (ubuntu+windows, windows+macos):
- Move bash-ism steps to shell-agnostic Node scripts:
    scripts/ci-guard-runner.cjs       — RUNNER_ENVIRONMENT check
    scripts/ci-rebase-check.cjs       — git fetch+merge PR base branch
    scripts/check-npm-integrity.cjs   — Node port of check-npm-integrity.sh
    scripts/ci-prepare-test-scope.cjs — write .ci-selected-tests.txt
    scripts/ci-smoke-skip.cjs         — set skip= output for full-only matrix entries
- Remove shell: bash from simple npm/node command steps (runner default applies)

This brings Windows violations from 19 to 0. Remaining 17 violations are all
MACOS_MISSING_EXPLICIT_ZSH in mixed-OS matrix jobs (test-full: windows+macos,
install-smoke smoke: ubuntu+macos) — these require job splitting to fix; see
BLOCKER in PR description.

* fix(#431): update workflow-shell-pinning test for H1 policy

The old test required all Windows-targeting npm steps to pin shell: bash
(to prevent pwsh stderr-swallow). Under H1, Windows runners must use
pwsh (native, no pin needed) — shell: bash on Windows is now the
violation, not the fix.

Update findViolations() to flag npm steps with effectiveShell === 'bash'
(rather than effectiveShell === null). Update synthetic tests to verify
the H1-inverted semantics: defaults.run.shell: bash on Windows is now 2
violations, not 0. Update test name and assertion messages to describe
the H1 constraint rather than the old missing-pin constraint.

* fix(#431): extend policy linter to resolve matrix.shell expressions

- expandRunsOn now captures all matrix.include row keys as realization
  context (os, node-version, shell, full_only, etc.) instead of only os
- effectiveShell now accepts a realizationContext and resolves
  ${{ matrix.<key> }} expressions against it before checking policy
- Unresolvable matrix key in shell expression emits UNRESOLVABLE_MATRIX
- Add 3 new tests: positive (zsh+pwsh per row → 0 violations),
  counter (bash in macOS row → WRONG_SHELL_FOR_OS), counter (missing
  shell key → UNRESOLVABLE_MATRIX)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): apply matrix.shell pattern to test-full and smoke jobs (clears BLOCKER)

test-full job (test.yml):
- Add shell: pwsh/zsh per matrix.include row (windows-latest→pwsh,
  macos-latest→zsh)
- Add job-level defaults.run.shell: ${{ matrix.shell }}
- No step-level shell pins existed to remove

smoke job (install-smoke.yml):
- Add shell: bash/zsh per matrix.include row (ubuntu→bash, macos→zsh)
- Add job-level defaults.run.shell: ${{ matrix.shell }}
- No step-level shell pins existed to remove

Policy linter now reports 0 violations across all workflow files.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#431): migrate .sh check scripts to .cjs; remove .sh originals

- Add scripts/check-env.cjs: Node.js port of check-env.sh with
  identical exit codes (0/1/2), human-readable and --json output,
  --help flag, and all 5 checks (node-version, npm-version,
  lockfile-present, lockfile-sync, version-manager-pin)
- Migrate all callers:
  - package.json check:env → node scripts/check-env.cjs
  - package.json check:integrity → node scripts/check-npm-integrity.cjs
  - scripts/ci-test-scope.cjs path strings → .cjs equivalents
  - .github/workflows/release.yml rc+finalize jobs → node .cjs (drop chmod+x)
  - .github/workflows/security-scan.yml → node .cjs (drop chmod+x)
  - tests/check-env.test.cjs → spawn node process.execPath [.cjs]
  - tests/npm-integrity-gate.test.cjs → spawn node process.execPath [.cjs]
- Delete scripts/check-env.sh and scripts/check-npm-integrity.sh

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#431): update doc references from .sh to .cjs

Update SECURITY.md and docs/contributing/bootstrap.md to reference the
canonical Node invocation instead of the removed bash scripts.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): use per-step shell:matrix.shell instead of defaults.run.shell (GHA compat)

GHA does not reliably resolve matrix expressions inside defaults.run.shell.
Per-step shell: always resolves correctly. Removed the defaults.run.shell block
from the test-full job (test.yml) and the smoke job (install-smoke.yml), and
added shell: \${{ matrix.shell }} directly on every run: step in both jobs.

Codex finding: defaults.run.shell with matrix expressions is not a
GHA-supported pattern; per-step shell: is the safe form.

* fix(#431): policy linter validates every matrix.include row independently

Removed runner-label-only dedup from expandRunsOn() in workflow-policy.cjs.
The prior guard (if !realizations.find(r => r.runner === runner)) collapsed
two macos-latest rows with different node-version/shell contexts into one,
hiding the second row's policy violation.

Each matrix.include row is a distinct CI realization with its own context;
validating it twice is harmless but skipping it causes false negatives.

Added counter-test (Test 8) in tests/policy-shell-pinning.test.cjs:
two macos-latest rows (shell:zsh compliant + shell:bash violation) must
produce exactly one WRONG_SHELL_FOR_OS violation on the second row.

* fix(#431): remove dedup-by-runner in Cartesian matrix.<key> expansion (Codex round 3)

The base-list path in expandRunsOn (matrix.<key> arrays, e.g. matrix.os)
previously guarded each push with `if (!realizations.find(r => r.runner === runner))`,
collapsing duplicate runner values into a single realization and hiding policy
violations on later rows of a Cartesian matrix.

Remove the guard unconditionally; each entry in the base-list array now produces
its own realization, matching the same fix already applied to the matrix.include path.

Add counter-test "Cartesian matrix os × shell — dedup must not collapse rows by
runner alone": matrix.os: [macos-latest, macos-latest] + shell: ${{ matrix.shell }}
now yields 2 realizations (not 1). Documents that Cartesian cross-product expansion
(carrying all keys into realization context) is a separate follow-up; current violations
are UNRESOLVABLE_MATRIX pending that work.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): remove 60s timeout regression on npm ci --dry-run (parity with check-env.sh)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#431): ci-rebase-check.cjs — return truthy sentinel on success (Codex round 4)

run() used execFileSync with stdio:'inherit', which returns null on success.
Caller checked `result !== null`, always false → every successful fetch fell
through to "failed after 3 attempts" exit-1 path.

Fix: run() now returns true on success, false on failure.
Update caller from `result !== null` to `if (result)`.

Adds tests/ci-rebase-check.test.cjs (5 tests) covering the sentinel contract
and a local-bare-remote integration smoke that verifies the full fetch+merge
path exits 0 when fetch succeeds.

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-05-28 09:23:59 -04:00