Commit Graph

4 Commits

Author SHA1 Message Date
Tom Boucher
c28134ab39 fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host

Adds local/no-duplicate-fold-marker, an AST rule that reports the second and
every subsequent `folded:<name>` marker in a host file, plus RuleTester cases
and a tree-wide regression assertion.

Failing-first on purpose: the rule is registered at error and the 25 duplicated
regions are still present, so eslint and the new tree-wide test are RED. The
deletions land in the next commit.

The marker key is the whitespace-delimited token after `folded:` — not the
issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on
tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and
feat-443-effort-fast-mode are two distinct folded suites.

Refs #3271

* fix(#3271): delete 25 duplicated folded suites from three install hosts

Three consolidated install suites each carried a verbatim second copy of a
contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and
green — each duplicated block registered and ran twice on every lane.

  tests/install.test.cjs                   5981-9937  (3957 lines, 18 blocks)
  tests/install-minimal-hooks.test.cjs     2734-4015  (1282 lines,  5 blocks)
  tests/install-write-confinement.test.cjs 1754-2321  ( 568 lines,  2 blocks)

Introduced by 6d072435d (#1975 re-applying #1970's hunks on a tree that already
had them, 2026-07-03) — one stale-base re-application, three files, one commit.
Verified by marker-count bisect: 1 at 4f779eda4 and 0cc7a1a42, 2 from 6d072435d
onward.

The later copy is deleted in each case, so every file returns to what its
authoring batch produced and blame on the surviving lines stays accurate.
local/no-duplicate-fold-marker, red on the previous commit, is now green.

tests/model-resolver.test.cjs is untouched: the issue lists it, but its two
blocks are folded from two different files and are not identical. It is a false
positive of the issue's own grep, whose `[a-z0-9-]*` key truncates at `.`.

Fixes #3271

* test(#3271): property-test marker identity and pin the alias non-goal

Three review findings, all fixed inline:

1. foldMarkerOf is a parser and carried no fast-check property test. Raised
   independently by the /code-review standards axis and the isolated adversarial
   pass; the file already establishes the fc.property-driving-ruleTester idiom
   for a sibling rule. Added, two arms over markers generated from [a-z0-9-._]:
   the same marker twice always reports exactly once against firstLine 1, and
   two distinct markers never collide. The alphabet includes `.` on purpose —
   an implementation keyed on the issue's [a-z0-9-]* slice passes arm 1 and
   fails arm 2, which is exactly the model-resolver false positive.

2. meta.docs.category was the novel value 'Test hygiene'; all 16 sibling local
   rules use 'Best Practices', 'Portability' or 'Reliability'. Now
   'Best Practices'.

3. A call through a further alias (const d = __foldDescribe) was unreported and
   undocumented — accidental rather than deliberate. It is now the fourth entry
   in the rule's documented non-goals, with the reason, and pinned by a valid
   RuleTester case so it cannot drift silently.

Refs #3271

* test(#3271): name the step and elapsed time when a baseline build fails

buildBaselineAtRef runs four bounded steps and, when one exceeded its bound,
threw a bare "spawnSync ETIMEDOUT" naming neither the step nor how long
anything took. Diagnosing one real failure took four separate experiments to
recover information the throw already had.

Each step is now timed, and any throw carries the breakdown: which step failed,
its elapsed time, the timings of every step that completed before it, all three
bounds, and the tail of the child's captured stdout/stderr.

The failure message is deliberately the carrier. On the remote runner the
captured output field comes back empty in failures.json while error and stack
survive verbatim, so the message is the only channel that reaches a reader of a
remote verdict.

Refs #3271

* fix(#3271): size the baseline generator bound for the machine it runs on

Instrumentation from a real remote-runner failure gave the breakdown:

  git-worktree-add=15.1s  npm-run-build-lib=19.8s  gen-emitted-baseline=FAILED@300.1s

Steps 1 and 2 are comfortable. Only the generator exceeds its bound, and it is
not hung — it needs more than 300s there.

Measured ladder for that step: ~22s idle in a container, ~39s end-to-end in a
clean container, ~142s with 8 CPU burners on 8 cores, and >300s under the real
suite. Its cost is 19 sequential installer spawns, and spawn latency is exactly
where a container degrades worst (3.9x slower than host, against 1.1x for file
IO) — which is why a CPU-only load test did not reproduce it and why four
earlier hypotheses (container slowness, network, shallow clone, CPU contention)
all measured clean.

The 300s bound was sized on an idle machine for a step that never runs on one.
Under the remote runner the on-disk baseline cache is structurally absent — CI
restores it via actions/cache keyed on github.event.pull_request.base.sha, a key
that exists only inside GitHub Actions — so this slow path runs on every remote
verification. The result: this gate has passed 0 times in 754 runs, failing 80
times and never once executing successfully.

Raised to the 600000ms ceiling that local/no-unbounded-spawn treats as the
largest meaningful bound; the other two bounds are untouched. This makes the
gate RUN, which is the point: the alternative considered and rejected was
degrading the timeout to a skip, and that was measured to turn the suite green
with the gate silently not running at all.

The real remedy is making the cache reachable from the remote runner so the
in-job build returns to being the rare fallback ADR-2719 §5 describes. That is a
gsd-test-runner change, not one this repo can make.

Refs #3271

* fix(#3271): tolerate an overlay source that vanishes mid-walk

Observed on the remote runner, three runs across three different branches:

  ENOENT: no such file or directory, link '/work/hooks/dist/gsd-config-reload.js'
    -> '/tmp/gsd-2930-overlay-6nOZay/hooks/dist/gsd-config-reload.js'

buildOverlayRepo enumerates names with readdirSync and then acts on each one, so
statSync, copyFileSync and linkSync all sit in a TOCTOU window. hooks/dist is
regenerated by an ATOMIC REPLACE (scripts/build-hooks.js unlinks and renames), so
any concurrently running test that rebuilds hooks retires a just-listed name
mid-walk and the overlay dies on it. linkOrCopyFile already tolerated EXDEV and
EPERM; ENOENT went straight through.

On ENOENT the source is now re-examined ONCE rather than slept on. An atomic
rename is a single syscall, so by the time the failure surfaces the successor is
either already in place (the retry succeeds) or the path has genuinely left the
tree, in which case there is nothing to mirror and the leaf is skipped. No sleep
and no spin: a timing-based wait here would be the very flake being fixed. Every
other errno still propagates untouched, so a real permission or IO fault stays a
hard failure.

Five tests hold the boundary: gone-for-good skips without retrying, mid-replace
retries exactly once and places the file, EACCES still throws, a real linkSync
ENOENT is injected by monkeypatching fs and restoring it in a finally (never a
mode-bit trick, which root bypasses), and isMissingPath accepts only ENOENT.

Refs #3271

* fix(#3271): order the timeout ladder inward-out and lock it

Two review blockers, both real.

The generator bound had been raised to 600000ms — exactly the whole-chunk timeout
in scripts/run-tests.cjs:973. A step bound equal to the chunk ceiling loses the
race: the chunk is killed first and the failure arrives as an opaque "no failed
step" kill, so the per-step diagnostic added a commit earlier was built and then
made unreachable in the same change.

Separately the #2767 test declared a per-test timeout of 300000ms, BELOW the
inner bound it was meant to permit, so it could still die at the exact 300s
ceiling this was supposed to lift — via node:test's timeout rather than
spawnSync's. Its sibling declared 900000ms, above the chunk ceiling, which is the
same opaque-kill hazard from the other direction.

The three bounds only produce a useful failure if they fire inward-out, so they
now do: step 360s, per-test 480s, chunk 600s. 360s is ~3x the passing observation
(91.6s / 115.8s) and 20% above the censored 300.1s timeout, while leaving 240s of
chunk headroom for every other file sharing it. Four tests lock the ordering,
including a drift guard on the exported values — without it, editing a call
site's literal timeout would leave the ordering assertions passing while the real
ladder inverted.

Also from review:

- err.gsdBaselineStep and err.gsdBaselineTimings were written and never read
  anywhere in the tree; only the rewritten message is consumed. Removed rather
  than kept as speculative surface.
- buildOverlayRepo discarded placeVanishableLeaf's boolean at both call sites, so
  a vanished leaf left the overlay with no accounting at all. It now collects the
  skipped paths and warns once. Not thrown: a source that left the tree really is
  not part of the snapshot, and throwing would reintroduce the crash the
  tolerance removes — but silence would let a dropped leaf resurface later as an
  unrelated missing-file assertion.
- The instrumentation commit shipped no test. One now drives a real failure and
  asserts the message names the step, its elapsed time, and the bounds.

Refs #3271

* chore(#3271): backfill the changeset PR number

* fix(#3271): bound a hook fan-out as its own class, not as a bare probe

CI failure on PR #3285, job full test (windows-latest, 22, shard 2/3) — every
other lane green, including windows-latest node 24 across all three shards:

  not ok 1 - blocks push when any to-be-pushed commit matches local blocked regex
    error: bash .githooks\pre-push failed — outcome=timed_out exitCode=null stderr=
    duration_ms: 15040.2168

A bound, not a hang: the test supplies stdin via input:, so the hook is not
blocked reading its ref list, and the duration lands exactly on the 15000ms
bound.

The site used PROBE_TIMEOUT_MS, which tests/helpers/timeouts.cjs documents as "a
single short CLI query or node -e probe against a temp fixture". This is not
that. It spawns bash running .githooks/pre-push, and the hook then invokes a MOCK
git that is itself a bash script, so one runHook is roughly four Git Bash spawns.
On Windows each is Defender-scanned and the first hook test in a file pays cold
start on top. That module's own docstring warns against precisely this: a call
site that differs from its class must not be forced onto a shared value that does
not describe it.

HOOK_FANOUT_TIMEOUT_MS is that missing class — 60000ms, 4x the bound that failed
and half INSTALL_TIMEOUT_MS, which is the right order: a hook fan-out is much
lighter than a full installer run and far heavier than reading back a version
string. Two tests lock the ordering against both neighbours, including one
asserting real margin over the censored 15040ms observation, since a bound that
merely matched what was measured would be the same defect again.

Scoped deliberately: the other ~360 runHook sites keep their current bounds. This
adds the norm and applies it where a real failure demonstrated the need, rather
than sweeping a value across sites with no evidence for any of them.

Refs #3271

---------

Co-authored-by: sim <sim@local>
2026-08-09 23:41:44 -04:00
Tom Boucher
3fac6e629f test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam

Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73.

Timeouts are sized from evidence already in the tree rather than a house
default, because this wave spawns installers rather than git plumbing and an
undersized bound does not catch a hang -- it manufactures CI flake, which is
worse, since a flake gets re-run instead of investigated. install.test.cjs
records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while
another lane passed the same commit in 12.7s, so full installs are bound at
120000ms against that recorded incident.

Also adds an auditable escape to the guard's timeout ceiling. The 600000ms
cap was set in #3143 from partial evidence, but fragment-single-edit-
propagation carries a documented, load-tested 900000ms bound on a run that
chains a full build plus eight generators -- the guard would have rejected a
correct timeout the moment that file left the allowlist. A value above the
ceiling is now permitted only with an inline allow-spawn-timeout-ceiling
marker carrying a non-empty reason. It raises the ceiling; it never waives
the requirement for a bound, which is asserted directly.

install-shared.cjs keeps its hand-rolled assert rather than routing through
throwIfFailed: its message embeds both streams, and throwIfFailed carries
only a trimmed stderr. The message now also names the outcome, so a bounded
timeout reads as such across its 38 importers instead of as
expected null to equal 0.

* test(#3145): extract class-norm timeouts and correct the build-hooks sizing

A pre-PR review found 52 copies of four class-norm timeout constants across
this wave. These are not per-suite fixture bindings -- they are shared facts
about how long a class of subprocess takes, derived from a recorded bench
incident. That norm already moved once (60000 to 120000 after a real
ETIMEDOUT), and 52 copies would have drifted the next time it moved.

Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and
converts the copies. A site that genuinely differs -- a real tsc compile, or
regen:derived -- keeps its own local constant with its own justification.

Also corrects a misclassification: scripts/build-hooks.js was sized as a
build at 120000 in twelve places and 60000 in another, but it compiles and
bundles nothing. Its own header says no bundling needed; it copies pre-built
files and syntax-checks them with vm. Three different values bounded one
script; now there is one.

* test(#3145): fix red CI — lint self-match and a Windows chunk overrun

Two failures on PR 3176.

lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The
fixture exists to prove an unrelated marker does NOT suppress the rule, so it
carries that marker's literal text as test data. Split via concatenation, the
same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own
self-match problem. The explanatory comment needed the same treatment.

The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped
seven minutes before the kill, so this was an overrun rather than a slow
chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs
regen:derived bounded at 900000ms, which is larger than the whole chunk
budget, so the chunk killer always fires first and it can never complete
there. Both the test and that bound predate this change; modifying the file
pulled it into the Windows targeted set and exposed it. Skipped on Windows
with the reason recorded; the Linux lanes cover it. The 900000 bound and its
ceiling marker are unchanged -- they are correct.

* test(#3145): refresh the stale test-timings cost table

The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs
packs chunks by measured duration from tests/test-timings.json, and an
unknown file falls back to the table's median weight -- advisory by design,
but it silently underweights exactly the files that matter.

Four of the failing chunk's 22 files were absent from the table, including
the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s
(it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s.
Both were weighted as average, so the chunk's total weight read 53.68 against
a budget of 60 and the packer produced a single chunk.

Regenerated from a passing full-suite run, per the remedy the script itself
documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since
gen-test-timings.cjs replaces the table wholesale rather than merging.

Proven against the real packer: the same 22 files now weigh 103.91 and split
into two chunks. No logic, budget, or timeout was changed; raising a budget
to make a red gate pass is not a fix.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:18:18 -04:00
Tom Boucher
d83e58eea0 fix(#437,#439,#440): restore defaults.run.shell + 'zsh {0}' format + Windows .cmd shell:true (PR #434 fallout) (#438)
* fix(#437): restore defaults.run.shell at job level (step-level matrix expr rejected by GHA)

Per actions/runner workflow-v1.0.json schema, `jobs.<job_id>.defaults.run.shell`
allows `matrix` context (job-defaults-run has context:[matrix,...]); step-level
`shell:` does not (run-step's shell field is plain string with no context array).
PR #434 used step-level shell:${{matrix.shell}}, which GHA's parser rejects with
"Unrecognized named-value: 'matrix'" — blocking every push to next and every
release.yml dispatch.

This commit:
- Removes step-level `shell: ${{ matrix.shell }}` from test-full (test.yml)
  and smoke (install-smoke.yml) jobs (17 directives).
- Adds `defaults.run.shell: ${{ matrix.shell }}` at job level in those two jobs.
- Fixes pre-existing shellcheck SC2129 in test.yml (individual >> redirects →
  grouped brace form) and SC2010 in install-smoke.yml (ls|grep → glob loop).

Verified locally with actionlint 1.7.12 (exit 0). Policy linter still 0 violations
(matrix.shell now resolves via job.defaults.run.shell which the linter already
handles per workflow-policy.cjs:effectiveShell).

Refs: actions/runner#444 (open since 2020), GHA contexts page section "Context availability".

* fix(#439): inline ci-smoke-skip back to shell (Node port required pre-checkout file resolution)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#440): use platform-correct npm.cmd on Windows for spawn (and surface-check other Node ports)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): use 'zsh {0}' format string in matrix.shell for macOS (zsh not in GHA built-ins)

Per https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions
(jobs.<job_id>.defaults.run.shell section):

  "You can use built-in shell keywords like bash, pwsh, python, sh, cmd, and
  powershell, or define a custom set of shell options."

zsh is not in the built-ins list. GHA accepts custom shells via a format string
containing '{0}', which it replaces with the temporary script file path at
runtime (same pattern as the perl {0} example in the docs).

Bare `shell: zsh` triggers: "Invalid shell option. Shell must be a valid
built-in or a format string containing '{0}'".

Precursor: 514cb429 introduced the matrix shell-pinning pattern; this completes
it by switching the macOS rows from the bare value to the required format string.

Also updates scripts/workflow-policy.cjs to normalise 'zsh {0}' to 'zsh' before
the policy comparison, so the repo-baseline test continues to pass (the linter
was correctly treating 'zsh {0}' as a distinct value from the policy 'zsh').

Affects:
- .github/workflows/test.yml: test-full matrix (node 22 + node 24 macOS rows)
- .github/workflows/install-smoke.yml: smoke matrix (macOS node 24 row)
- scripts/workflow-policy.cjs: detectViolation strips ' {0}' format suffix

* fix(#440): add shell:true to spawnSync on Windows for .cmd files (Node docs requirement)

Per https://nodejs.org/docs/latest-v22.x/api/child_process.html:

  ".bat and .cmd files require a terminal to run and cannot be launched
  directly with execFile(). To run these scripts on Windows, use
  child_process.spawn() with the shell option, child_process.exec(), or
  spawn cmd.exe with the script as an argument."

  "On Windows, .bat and .cmd files require a shell to execute. Use
  child_process.exec() or child_process.spawn() with the shell: true option."

On Windows, npm is installed as npm.cmd (a batch wrapper). Without
shell: true, spawnSync resolves the binary directly and fails with
ENOENT / "npm binary not found on PATH" because the OS cannot execute
a .cmd file without cmd.exe as the intermediary.

The fix uses `shell: process.platform === 'win32'` so the shell spawning
is only activated on Windows; macOS/Linux continue to resolve the plain
npm binary directly with shell: false, preserving the existing behaviour
on non-Windows platforms.

Updated both spawnSync(npmCmd, ...) call sites:
- npm --version check (line 182)
- npm ci --dry-run lockfile-sync check (line 215)

* fix(#437): bug-410 defaults test — set USERPROFILE for Windows os.homedir() redirect

On Windows, os.homedir() reads USERPROFILE (not HOME), so the test's
process.env.HOME = FAKE_HOME redirect was silently ignored. finishInstall's
path.join(os.homedir(), '.gsd') resolved to the real user home and the
defaults.json write either failed (permissions) or landed outside the temp
dir, causing the existsSync assertion to return false.

Fix: also set process.env.USERPROFILE = FAKE_HOME so os.homedir() returns
the sandboxed directory on Windows. Node.js docs (os.homedir):
https://nodejs.org/docs/latest-v22.x/api/os.html#oshomedir

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): precommit-alias-drift hook test — use path.delimiter for PATH

Hardcoded ':' PATH separator breaks Windows where process.env.PATH uses ';'.
The malformed PATH passed to bash caused the mock git/npm stubs in binDir
to be invisible to the hook script; npm was never called and the marker
file never written.

Fix: replace ':' with path.delimiter in both PATH constructions so the
env var is well-formed on Windows (';') and POSIX (':') alike.

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): prepush-enterprise-email hook test — use path.delimiter for PATH

Same root cause as precommit-alias-drift: hardcoded ':' PATH separator is
invalid on Windows (';' required). The malformed PATH meant bash ran the
real git binary instead of the mock stub, which rejected the placeholder
SHAs 'refs-local-sha' / 'refs-remote-sha' with a fatal ambiguous-argument
error rather than returning the fixture commit list.

Fix: replace ':' with path.delimiter in both execFileSync PATH env values.

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): set MSYS2_PATH_TYPE=inherit so mock stubs take precedence in Git Bash PATH

Root cause: Git Bash (MSYS2) on Windows prepends its own system directories
(/mingw64/bin, /usr/bin, /bin) to the PATH at process startup before the
user-supplied Windows PATH entries. This placed the real git/npm binaries
ahead of the mock stubs in binDir even though binDir was first in the Windows
PATH passed to execFileSync. The path.delimiter fix (0042fe0d) made the PATH
syntactically correct for Windows (semicolons) but did not change the MSYS2
system-dir prepend order.

The real git rejected placeholder SHAs (refs-local-sha, refs-remote-sha) with
"fatal: ambiguous argument", producing the observed Windows CI failure. For the
pre-commit test, the real git output nothing (no staged files on a fresh
checkout), so the grep match failed and npm was never called.

Fix: set MSYS2_PATH_TYPE=inherit in the env passed to both bash spawns.
With inherit, MSYS2 uses only the converted Windows PATH without prepending
system directories, so binDir (converted from Windows to POSIX) is first in
the search path and the mock stubs are found.

grep/tr/printf remain available: the GHA Windows runner PATH includes
C:\Program Files\Git\usr\bin which contains these utilities; MSYS2 converts
that Windows entry to a POSIX path on startup. The /usr/bin/env shebang in
mock stubs resolves through MSYS2's virtual filesystem mount (not via PATH)
and is always accessible regardless of MSYS2_PATH_TYPE.

On macOS/Linux this variable is ignored; no behaviour change on those platforms.

Source: https://www.msys2.org/wiki/MSYS2-introduction/#path
(MSYS2_PATH_TYPE controls whether system dirs are prepended to converted PATH)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): hook test mocks — use cmd-shim pattern for Windows bin resolution

On Windows, bash (Git Bash / MSYS2) resolves PATH commands by scanning for
extensionless files, but cmd.exe and Win32 process creation resolve via
PATHEXT (.CMD, .BAT, .EXE). When execFileSync('bash', [hookPath]) runs a
hook that calls `git` or `npm`, both resolution paths may fire. The previous
approach set MSYS2_PATH_TYPE=inherit in the child env, but that variable is
only read in /etc/profile (login-shell path) — bash launched without --login
never sources /etc/profile, so the variable had no effect:
https://github.com/msys2/MSYS2-packages/blob/master/filesystem/profile

Fix: adopt the cmd-shim three-file pattern used by npm itself:
https://github.com/npm/cmd-shim
For each mock binary, write:
  <name>          extensionless bash script (bash PATH scan)
  <name>.cmd      batch wrapper delegating to bash (PATHEXT / cmd.exe)
  <name>.ps1      PowerShell wrapper (completeness)

This is the same approach used by stevemao/mock-bin for test mocking with
Windows CI green on AppVeyor:
https://github.com/stevemao/mock-bin

The .cmd and .ps1 files are only written on process.platform === 'win32'.
MSYS2_PATH_TYPE is removed from the child env — it was ineffective and is
no longer needed with the shim files in place.

* fix(#437): tarball-smoke — raise CHILD_TIMEOUT_MS on Windows to 600 s

The CI failure showed a test duration of 120003.1812 ms — matching the
previous CHILD_TIMEOUT_MS = 120_000 exactly. When spawnSync hits its
timeout, it sends SIGTERM and returns { status: null, stdout: '', stderr: '' }
per the Node.js docs:
https://nodejs.org/docs/latest-v22.x/api/child_process.html
  "status: <number> | <null> — The exit code of the subprocess, or null if
   the subprocess terminated due to a signal."

The installResult check is `status !== 0`; null !== 0 is true, so the
timeout fired the INSTALL_FAILED path with empty stdout/stderr, which made
the root cause invisible in CI logs.

GitHub-hosted Windows runners are slower than Linux/macOS for
filesystem-heavy operations (npm install -g of a 1499-file tarball):
https://docs.github.com/en/actions/using-github-hosted-runners/about-github-hosted-runners/about-github-hosted-runners#standard-github-hosted-runners-for-public-repositories

Fix: use 600_000 ms (10 min) on Windows, keeping 120_000 ms on POSIX.
600 s matches the SLOW_HOST_TIMEOUT already used in the test before() helper
for the pack + install fixture step.

Also expose `signal` and `installError` in the INSTALL_FAILED details object
so a future timeout (status=null, signal='SIGTERM', stdout='') is immediately
diagnosable in CI logs without guesswork.

* fix(#437): chmod +x via bash on Windows for hook test mocks (root cause: fs.writeFileSync mode=0o755 no-op on NTFS)

Root cause: Node's fs.writeFileSync mode=0o755 is a no-op for the execute
bit on Windows NTFS. Per https://nodejs.org/docs/latest-v22.x/api/fs.html:
"on Windows only the write permission can be changed." Bash's access(X_OK)
therefore skips the mock file; the real git/npm binary is found later in PATH
and the hook runs against real state instead of the test double.

Fix: after writeFileSync, invoke Git Bash's chmod via the POSIX emulation
layer (Cygwin/MSYS2), which sets the NTFS execute ACL that Node cannot reach:

    const posixPath = filePath.replace(/\\/g, '/');
    execFileSync('bash', ['-c', `chmod +x "${posixPath}"`], { stdio: 'pipe' });

execFileSync('bash', ...) works because Git for Windows ships bash on PATH in
all GHA Windows runners. Forward-slash conversion is required because MSYS2
bash auto-converts /c/foo paths but not mixed-separator paths.

Why prior approaches didn't take effect:
- MSYS2_PATH_TYPE=inherit: only read in /etc/profile (login-shell path);
  execFileSync('bash', ...) launches non-interactively without --login, so
  /etc/profile is never sourced.
  Ref: https://github.com/msys2/MSYS2-packages/blob/master/filesystem/profile
- .cmd/.ps1 cmd-shim wrappers: bash does POSIX command resolution and does
  not honor PATHEXT, so wrappers are not found by bash's own PATH scan.
  They are not wrong (kept for non-bash callers) but do not fix bash's X_OK.

Files changed: tests/precommit-alias-drift-hook.test.cjs,
               tests/prepush-enterprise-email-hook.test.cjs

* refactor(#437): hooks use GIT_OVERRIDE/NPM_OVERRIDE env-var DI; tests drop PATH-mocking

Four prior rounds (path.delimiter join, MSYS2_PATH_TYPE=inherit, cmd-shim
.cmd/.ps1 wrappers, chmod-via-bash post-write) all failed to make MSYS2
bash's PATH-lookup find the mock executables. The root cause is that none
of those approaches can reliably override bash's own command-resolution
on NTFS without fighting NTFS execute-ACLs or login-shell profile sourcing.

The simplest robust solution is to bypass PATH entirely:

Hooks: each hook now binds GIT_CMD="${GIT_OVERRIDE:-git}" (and NPM_CMD for
pre-commit) at the top. When env vars are unset the hooks invoke bare
`git`/`npm` exactly as before — zero behavior change for users.

Tests: writeMockBin/binDir/PATH manipulation replaced by writeMock(), which
writes a .sh mock to a tmpDir and passes its absolute path via GIT_OVERRIDE
/ NPM_OVERRIDE in the execFileSync env. Bash inside the hook executes the
path directly via the seam — no PATH scan, no NTFS ACL check, no MSYS2
profile dependency.

Test-rigor principle: the new seam (env-var injection) is platform-
independent and doesn't rely on bash's command-resolution mechanism on
the host OS.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 18:06:19 -04:00
Tom Boucher
006cdafe8f ci(drift): enforce alias freshness checks in CI and contributor flow (#2910)
Merging alias-drift guardrails and local hook hardening.
2026-04-30 14:19:46 -04:00