f7df920681f233ae0fe064ee659550bdf41ff708
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3fac6e629f |
test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73. Timeouts are sized from evidence already in the tree rather than a house default, because this wave spawns installers rather than git plumbing and an undersized bound does not catch a hang -- it manufactures CI flake, which is worse, since a flake gets re-run instead of investigated. install.test.cjs records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while another lane passed the same commit in 12.7s, so full installs are bound at 120000ms against that recorded incident. Also adds an auditable escape to the guard's timeout ceiling. The 600000ms cap was set in #3143 from partial evidence, but fragment-single-edit- propagation carries a documented, load-tested 900000ms bound on a run that chains a full build plus eight generators -- the guard would have rejected a correct timeout the moment that file left the allowlist. A value above the ceiling is now permitted only with an inline allow-spawn-timeout-ceiling marker carrying a non-empty reason. It raises the ceiling; it never waives the requirement for a bound, which is asserted directly. install-shared.cjs keeps its hand-rolled assert rather than routing through throwIfFailed: its message embeds both streams, and throwIfFailed carries only a trimmed stderr. The message now also names the outcome, so a bounded timeout reads as such across its 38 importers instead of as expected null to equal 0. * test(#3145): extract class-norm timeouts and correct the build-hooks sizing A pre-PR review found 52 copies of four class-norm timeout constants across this wave. These are not per-suite fixture bindings -- they are shared facts about how long a class of subprocess takes, derived from a recorded bench incident. That norm already moved once (60000 to 120000 after a real ETIMEDOUT), and 52 copies would have drifted the next time it moved. Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and converts the copies. A site that genuinely differs -- a real tsc compile, or regen:derived -- keeps its own local constant with its own justification. Also corrects a misclassification: scripts/build-hooks.js was sized as a build at 120000 in twelve places and 60000 in another, but it compiles and bundles nothing. Its own header says no bundling needed; it copies pre-built files and syntax-checks them with vm. Three different values bounded one script; now there is one. * test(#3145): fix red CI — lint self-match and a Windows chunk overrun Two failures on PR 3176. lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The fixture exists to prove an unrelated marker does NOT suppress the rule, so it carries that marker's literal text as test data. Split via concatenation, the same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own self-match problem. The explanatory comment needed the same treatment. The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped seven minutes before the kill, so this was an overrun rather than a slow chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs regen:derived bounded at 900000ms, which is larger than the whole chunk budget, so the chunk killer always fires first and it can never complete there. Both the test and that bound predate this change; modifying the file pulled it into the Windows targeted set and exposed it. Skipped on Windows with the reason recorded; the Linux lanes cover it. The 900000 bound and its ceiling marker are unchanged -- they are correct. * test(#3145): refresh the stale test-timings cost table The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs packs chunks by measured duration from tests/test-timings.json, and an unknown file falls back to the table's median weight -- advisory by design, but it silently underweights exactly the files that matter. Four of the failing chunk's 22 files were absent from the table, including the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s (it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s. Both were weighted as average, so the chunk's total weight read 53.68 against a budget of 60 and the packer produced a single chunk. Regenerated from a passing full-suite run, per the remedy the script itself documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since gen-test-timings.cjs replaces the table wholesale rather than merging. Proven against the real packer: the same 22 files now weigh 103.91 and split into two chunks. No logic, budget, or timeout was changed; raising a budget to make a red gate pass is not a fix. --------- Co-authored-by: sim <sim@local> |
||
|
|
1d208e5af6 |
test(#3144): bound the git/worktree cluster onto the process seam (#3152)
* test(#3144): bound the git/worktree cluster onto the process seam Migrates 180 unbounded sync spawn sites across 19 files. Every previously unbounded call now carries an explicit timeout with a comment giving the number and why. The migration is not a callee swap. execSync and execFileSync throw on a non-zero exit and the seam never does, so each site was classified first: sites that rely on the throw route to gitOrThrow, and sites that already read .status to detect an EXPECTED non-zero -- an intended cherry-pick conflict, a rev-parse outside a repo driving a skip -- route to the never-throwing runGit instead, which would otherwise throw on exactly the exit being probed for. Two same-named git() helpers in worktree-cleanup.test.cjs have different return contracts, one trimmed and one raw; both are preserved rather than unified. Collapses five hand-rolled throw wrappers onto one throwIfFailed in git-fixture.cjs, which gitOrThrow now also uses so the shape cannot drift. Allowlist drops 139 to 120; BASELINE lowered to match. * test(#3144): fix pre-PR review findings Documents throwIfFailed in the CONTEXT.md glossary and CONTRIBUTING.md -- it became the shared throw mechanism without either doc naming it. Routes the sixth and seventh hand-rolled copies of the throw shape through throwIfFailed (worktree-baseref-install, worktree-safety-reap); the first consolidation missed both. Converts ci-rebase-check's 8 fixture-setup calls from unchecked runGit to gitOrThrow so a failed setup step aborts where it fails rather than surfacing later as a confusing failure against the wrong subject. Adds 12 direct unit tests for throwIfFailed, which until now was only exercised transitively. Splits verify.test.cjs's non-git grep/sed bound off GIT_TIMEOUT_MS. --------- Co-authored-by: sim <sim@local> |
||
|
|
a28dcec981 |
chore(#597): replace count-based ratchet guards with AST lint + named-set allowlists (#603)
The windows-test-parity ratchet greps test source for fs.rmSync-without-
maxRetries (and six other Windows-portability anti-patterns), failing when an
integer offender COUNT exceeds a frozen baseline (rmSync: 95). A count ratchet
is a Goodhart metric: fixing one offender and adding another keeps the count
constant, so a new defect slips through green. Replace it — and every other
count ratchet in the repo — with a layered, masking-proof design.
Behavioral seam test
- tests/helpers-cleanup.test.cjs proves helpers.cleanup() carries the Windows
EBUSY retry budget. cleanup() delegates retries to Node's fs.rmSync via
maxRetries (it owns no loop), so the test asserts the option contract
(recursive/force/maxRetries>0/retryDelay>0) + real-FS removal + the cwd-guard,
rather than a loop that does not exist. The EBUSY risk is now tested ONCE at
the helper, not approximated textually at every call site.
Write-time ESLint rule (AST-accurate, replaces the grep)
- eslint-rules/no-raw-rmsync-in-tests.cjs (error in tests/**/*.test.cjs) bans
raw fs.rmSync, steering to cleanup(). Catches member, computed (fs['rmSync']),
destructured and aliased forms; escape hatch is inline
`// eslint-disable-next-line local/no-raw-rmsync-in-tests -- <reason>` only.
- Migrated 336 raw fs.rmSync teardown calls across ~116 test files to cleanup().
~18 genuinely load-bearing sites (mid-test SUT/fault-injection removals,
error-swallowing or name-colliding local teardown helpers) keep the raw call
with an inline eslint-disable + reason.
Shared anti-ratchet primitive
- scripts/lib/allowlist-ratchet.cjs:
- assertWithinAllowlist: fails on NOVEL ids (new offender introduced) AND on
STALE ids (a known offender was fixed but not pruned) — identity, not count,
and a ratchet DOWN toward zero.
- assertTightCeiling: a size/length budget whose ceiling must stay within a
grace band of the high-water mark, so budgets may only tighten, never creep.
Ratchets converted onto the primitive
- windows-test-parity-guard.test.cjs: rmSync rule deleted (now ESLint-enforced);
the remaining six patterns moved from integer baselines to named-set
allowlists with ratchet-down.
- scripts/lint-test-file-count.{cjs,allowlist.json}: per-module integer counts →
named filename sets (closes the swap-a-file-keep-the-count blind spot); a
module dropping under cap now FAILS to force pruning its allowlist entry.
- enh-2790 skill-count `<= 63` → named skill allowlist (ratchets toward ~58).
Size budgets hardened (tighten-only)
- agent-size / workflow-size / feat-3039 help-tiered: ceilings lowered to the
current high-water mark and an assertTightCeiling anti-creep check added per
tier. Fixed external-contract limits (description ≤100 chars, agent ≤100 KB)
are intentionally left as-is — they are not grandfathered creeping budgets.
No user-facing behavior change (tests + tooling only); no USER_FACING_PREFIXES
touched, so no changeset fragment is required.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
48b1e35187 |
fix(#431): enforce H1 shell policy (linux=bash, macOS=zsh, windows=pwsh) across PR + release gates (#434)
* test(#431): policy-shell-pinning linter — RED baseline (37 violations on origin/next) Adds scripts/workflow-policy.cjs: H1 shell-policy linter with POLICY map, VIOLATION enum, matrix expansion, effective-shell resolution order, and runPolicyLint({ workflowsDir }) entry point. Adds tests/policy-shell-pinning.test.cjs: 8 tests (baseline + 6 synthetic counter-tests). Synthetic tests 2–7 pass; baseline test is intentionally RED (37 violations: 28 in test.yml, 9 in install-smoke.yml — all macos/windows lanes using shell: bash instead of native zsh/pwsh). Adds js-yaml@4.1.1 as devDependency for YAML parsing. * fix(#431): switch ubuntu/windows lanes to native shells; extract bash-isms to Node Remove all explicit shell: bash pins from ubuntu-only jobs (changes, lint-tests, coverage, required-tests, smoke-unpacked) — ubuntu runner default is bash, which is both H1-compliant and the runner default, making the pin redundant. For the test and test-full mixed-OS jobs (ubuntu+windows, windows+macos): - Move bash-ism steps to shell-agnostic Node scripts: scripts/ci-guard-runner.cjs — RUNNER_ENVIRONMENT check scripts/ci-rebase-check.cjs — git fetch+merge PR base branch scripts/check-npm-integrity.cjs — Node port of check-npm-integrity.sh scripts/ci-prepare-test-scope.cjs — write .ci-selected-tests.txt scripts/ci-smoke-skip.cjs — set skip= output for full-only matrix entries - Remove shell: bash from simple npm/node command steps (runner default applies) This brings Windows violations from 19 to 0. Remaining 17 violations are all MACOS_MISSING_EXPLICIT_ZSH in mixed-OS matrix jobs (test-full: windows+macos, install-smoke smoke: ubuntu+macos) — these require job splitting to fix; see BLOCKER in PR description. * fix(#431): update workflow-shell-pinning test for H1 policy The old test required all Windows-targeting npm steps to pin shell: bash (to prevent pwsh stderr-swallow). Under H1, Windows runners must use pwsh (native, no pin needed) — shell: bash on Windows is now the violation, not the fix. Update findViolations() to flag npm steps with effectiveShell === 'bash' (rather than effectiveShell === null). Update synthetic tests to verify the H1-inverted semantics: defaults.run.shell: bash on Windows is now 2 violations, not 0. Update test name and assertion messages to describe the H1 constraint rather than the old missing-pin constraint. * fix(#431): extend policy linter to resolve matrix.shell expressions - expandRunsOn now captures all matrix.include row keys as realization context (os, node-version, shell, full_only, etc.) instead of only os - effectiveShell now accepts a realizationContext and resolves ${{ matrix.<key> }} expressions against it before checking policy - Unresolvable matrix key in shell expression emits UNRESOLVABLE_MATRIX - Add 3 new tests: positive (zsh+pwsh per row → 0 violations), counter (bash in macOS row → WRONG_SHELL_FOR_OS), counter (missing shell key → UNRESOLVABLE_MATRIX) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): apply matrix.shell pattern to test-full and smoke jobs (clears BLOCKER) test-full job (test.yml): - Add shell: pwsh/zsh per matrix.include row (windows-latest→pwsh, macos-latest→zsh) - Add job-level defaults.run.shell: ${{ matrix.shell }} - No step-level shell pins existed to remove smoke job (install-smoke.yml): - Add shell: bash/zsh per matrix.include row (ubuntu→bash, macos→zsh) - Add job-level defaults.run.shell: ${{ matrix.shell }} - No step-level shell pins existed to remove Policy linter now reports 0 violations across all workflow files. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#431): migrate .sh check scripts to .cjs; remove .sh originals - Add scripts/check-env.cjs: Node.js port of check-env.sh with identical exit codes (0/1/2), human-readable and --json output, --help flag, and all 5 checks (node-version, npm-version, lockfile-present, lockfile-sync, version-manager-pin) - Migrate all callers: - package.json check:env → node scripts/check-env.cjs - package.json check:integrity → node scripts/check-npm-integrity.cjs - scripts/ci-test-scope.cjs path strings → .cjs equivalents - .github/workflows/release.yml rc+finalize jobs → node .cjs (drop chmod+x) - .github/workflows/security-scan.yml → node .cjs (drop chmod+x) - tests/check-env.test.cjs → spawn node process.execPath [.cjs] - tests/npm-integrity-gate.test.cjs → spawn node process.execPath [.cjs] - Delete scripts/check-env.sh and scripts/check-npm-integrity.sh Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#431): update doc references from .sh to .cjs Update SECURITY.md and docs/contributing/bootstrap.md to reference the canonical Node invocation instead of the removed bash scripts. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): use per-step shell:matrix.shell instead of defaults.run.shell (GHA compat) GHA does not reliably resolve matrix expressions inside defaults.run.shell. Per-step shell: always resolves correctly. Removed the defaults.run.shell block from the test-full job (test.yml) and the smoke job (install-smoke.yml), and added shell: \${{ matrix.shell }} directly on every run: step in both jobs. Codex finding: defaults.run.shell with matrix expressions is not a GHA-supported pattern; per-step shell: is the safe form. * fix(#431): policy linter validates every matrix.include row independently Removed runner-label-only dedup from expandRunsOn() in workflow-policy.cjs. The prior guard (if !realizations.find(r => r.runner === runner)) collapsed two macos-latest rows with different node-version/shell contexts into one, hiding the second row's policy violation. Each matrix.include row is a distinct CI realization with its own context; validating it twice is harmless but skipping it causes false negatives. Added counter-test (Test 8) in tests/policy-shell-pinning.test.cjs: two macos-latest rows (shell:zsh compliant + shell:bash violation) must produce exactly one WRONG_SHELL_FOR_OS violation on the second row. * fix(#431): remove dedup-by-runner in Cartesian matrix.<key> expansion (Codex round 3) The base-list path in expandRunsOn (matrix.<key> arrays, e.g. matrix.os) previously guarded each push with `if (!realizations.find(r => r.runner === runner))`, collapsing duplicate runner values into a single realization and hiding policy violations on later rows of a Cartesian matrix. Remove the guard unconditionally; each entry in the base-list array now produces its own realization, matching the same fix already applied to the matrix.include path. Add counter-test "Cartesian matrix os × shell — dedup must not collapse rows by runner alone": matrix.os: [macos-latest, macos-latest] + shell: ${{ matrix.shell }} now yields 2 realizations (not 1). Documents that Cartesian cross-product expansion (carrying all keys into realization context) is a separate follow-up; current violations are UNRESOLVABLE_MATRIX pending that work. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): remove 60s timeout regression on npm ci --dry-run (parity with check-env.sh) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#431): ci-rebase-check.cjs — return truthy sentinel on success (Codex round 4) run() used execFileSync with stdio:'inherit', which returns null on success. Caller checked `result !== null`, always false → every successful fetch fell through to "failed after 3 attempts" exit-1 path. Fix: run() now returns true on success, false on failure. Update caller from `result !== null` to `if (result)`. Adds tests/ci-rebase-check.test.cjs (5 tests) covering the sentinel contract and a local-bare-remote integration smoke that verifies the full fetch+merge path exits 0 when fetch succeeds. --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> |