Commit Graph

5 Commits

Author SHA1 Message Date
Tom Boucher
107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00
Tom Boucher
4b69dc346b fix(#2725): repoint the pre-commit alias guard at sources git can actually stage, drop nine dead ones (#3273)
* test(#2725): failing-first coverage for the inert pre-commit alias guard

Replaces two stale tests that asserted on `sdk/src/query/command-manifest.phase.ts`
— a path retired with the SDK boundary (ADR-0174), so both passed forever while
guarding nothing.

The new matrix drives .githooks/pre-commit through its GIT_OVERRIDE/NPM_OVERRIDE
seams and asserts on the real tracked sources the drift checker reads. Red until
the guard is repointed off the gitignored build outputs it currently watches.

* fix(#2725): repoint the pre-commit alias guard at sources git can actually stage

`.githooks/pre-commit` carried ten staged-path guards and not one of them could
do its job.

Nine anchored on `^sdk/…`, a tree retired by ADR-0174, and invoked npm scripts
that no longer exist (`check:state-document-fresh` and eight siblings). Their
only reachable behavior was to abort the commit with `Missing script` — which
required matching a path that cannot exist, so they were dead twice over.

The tenth was the real defect. Its npm target does exist, but it matched
`^gsd-core/bin/lib/command-aliases\.cjs$` — a gitignored build output
(.gitignore:172). An ignored path never appears in `git diff --cached
--name-only`, so the guard was not stale, it was structurally unmatchable: it
watched the derived layer instead of the source layer, from the day it was
written.

Repointed at the nine tracked `src/*.cts` sources `scripts/check-alias-drift.cjs`
actually reads. The family table moves to `scripts/lib/alias-drift-families.cjs`
so the checker and the hook derive their surface from one list, and a new parity
test fails if a family is added to the checker without the hook learning to watch
its source — the rot mechanism, not just this instance of it.

Matching is now `grep -Fxqf` (fixed strings, whole line): exact path equality,
no regex anchors to get wrong as the list grows. Staged paths are collected into
a variable before matching, because `grep -q` exits on first match and would
SIGPIPE its upstream, which under `set -o pipefail` turns a successful match into
a non-zero pipeline status.

The two CONTRIBUTING.md recipes pasted copies of the hook bodies inline — a third
parallel surface, and one that had already drifted: the pre-commit copy carried
the same dead `sdk/` patterns, and the pre-push copy would have overwritten the
committed hook with a paraphrase that drops the GIT_OVERRIDE seam its test drives.
Both now just point git at the committed files. The pre-push recipe's
`'@example-corp\\.com$'` also never matched anything — inside single quotes bash
keeps both backslashes.

`.githooks/pre-push` was audited for the same rot and has none: it keys on no
paths, no-ops unless GSD_BLOCKED_AUTHOR_REGEX is set, and is covered. Unchanged.

Hooks stay opt-in. Nothing registers `core.hooksPath` for you, per
CONTRIBUTING.md's documented one-time setup.

* test(#2725): make the hook/checker parity assertion bidirectional

Review finding from the isolated adversarial pass: the parity row only caught
the hook UNDER-watching relative to scripts/lib/alias-drift-families.cjs. Drop a
family from the module and the hook would keep watching its source with nothing
to notice — the same divergence class this change exists to close, just pointed
the other way.

The new row takes every `src/*-command-router.cts` on disk as the universe and
asserts the hook stays silent for the eight routers the drift check does not
read. Both directions are now covered by running the real hook, not by comparing
two lists in the test.

Also corrects a CONTRIBUTING.md overclaim caught by the standards pass: 9 of the
11 watched paths derive from the module, not all 11 — bash cannot require a CJS
module, so the hook carries literals and the test is what binds them.

* fix(#2725): ship the shared family table and fix the mock that hid its own rows

Three defects the remote runner caught that local probing did not.

The mock `git` in the regression test emitted its staged-path payload as
`printf '%s' "src/command-aliases.cts\n"`. Bash does not expand `\n` inside a
double-quoted string and printf does not expand escapes in a `%s` argument, so
the mock produced one unterminated line containing a literal backslash-n. No
whole-line match could ever succeed, and every row that expects the hook to FIRE
failed while every row that expects silence passed — which is exactly the
signature the run reported: 9 failures, all of them fire-expecting rows. The
payload now goes through a file the mock `cat`s, which is byte-exact and is what
makes the CR-terminated and empty-staged-list rows mean what they say.

`scripts/lib/alias-drift-families.cjs` was not enumerated in
`GSD_SCRIPTS_LIB_FILES` (bin/install.js:377), so the installer never copied it.
That is not cosmetic: `scripts/check-alias-drift.cjs` ships, and it now requires
this module — an installed tree would have failed with MODULE_NOT_FOUND the
first time a consumer ran `check:alias-drift`. Added to the manifest, which is
what the #3184 install/uninstall parity tests assert against `readdirSync`.

Regenerated the 19 committed install-tree fixtures via `npm run gen:install-tree`
to record the new emitted path. The diff is +1 line per fixture and nothing else.

`npm run lint:ci` exits 0. The earlier claim that `scripts/` is outside the
emitted surface was wrong: `scripts/lib/` is copied into every runtime's install
tree, which is why 19 golden-install-tree cases moved.

* chore(#2725): backfill changeset pr: 3273

---------

Co-authored-by: sim <sim@local>
2026-08-09 18:29:16 -04:00
Tom Boucher
3fac6e629f test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam

Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73.

Timeouts are sized from evidence already in the tree rather than a house
default, because this wave spawns installers rather than git plumbing and an
undersized bound does not catch a hang -- it manufactures CI flake, which is
worse, since a flake gets re-run instead of investigated. install.test.cjs
records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while
another lane passed the same commit in 12.7s, so full installs are bound at
120000ms against that recorded incident.

Also adds an auditable escape to the guard's timeout ceiling. The 600000ms
cap was set in #3143 from partial evidence, but fragment-single-edit-
propagation carries a documented, load-tested 900000ms bound on a run that
chains a full build plus eight generators -- the guard would have rejected a
correct timeout the moment that file left the allowlist. A value above the
ceiling is now permitted only with an inline allow-spawn-timeout-ceiling
marker carrying a non-empty reason. It raises the ceiling; it never waives
the requirement for a bound, which is asserted directly.

install-shared.cjs keeps its hand-rolled assert rather than routing through
throwIfFailed: its message embeds both streams, and throwIfFailed carries
only a trimmed stderr. The message now also names the outcome, so a bounded
timeout reads as such across its 38 importers instead of as
expected null to equal 0.

* test(#3145): extract class-norm timeouts and correct the build-hooks sizing

A pre-PR review found 52 copies of four class-norm timeout constants across
this wave. These are not per-suite fixture bindings -- they are shared facts
about how long a class of subprocess takes, derived from a recorded bench
incident. That norm already moved once (60000 to 120000 after a real
ETIMEDOUT), and 52 copies would have drifted the next time it moved.

Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and
converts the copies. A site that genuinely differs -- a real tsc compile, or
regen:derived -- keeps its own local constant with its own justification.

Also corrects a misclassification: scripts/build-hooks.js was sized as a
build at 120000 in twelve places and 60000 in another, but it compiles and
bundles nothing. Its own header says no bundling needed; it copies pre-built
files and syntax-checks them with vm. Three different values bounded one
script; now there is one.

* test(#3145): fix red CI — lint self-match and a Windows chunk overrun

Two failures on PR 3176.

lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The
fixture exists to prove an unrelated marker does NOT suppress the rule, so it
carries that marker's literal text as test data. Split via concatenation, the
same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own
self-match problem. The explanatory comment needed the same treatment.

The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped
seven minutes before the kill, so this was an overrun rather than a slow
chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs
regen:derived bounded at 900000ms, which is larger than the whole chunk
budget, so the chunk killer always fires first and it can never complete
there. Both the test and that bound predate this change; modifying the file
pulled it into the Windows targeted set and exposed it. Skipped on Windows
with the reason recorded; the Linux lanes cover it. The 900000 bound and its
ceiling marker are unchanged -- they are correct.

* test(#3145): refresh the stale test-timings cost table

The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs
packs chunks by measured duration from tests/test-timings.json, and an
unknown file falls back to the table's median weight -- advisory by design,
but it silently underweights exactly the files that matter.

Four of the failing chunk's 22 files were absent from the table, including
the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s
(it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s.
Both were weighted as average, so the chunk's total weight read 53.68 against
a budget of 60 and the packer produced a single chunk.

Regenerated from a passing full-suite run, per the remedy the script itself
documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since
gen-test-timings.cjs replaces the table wholesale rather than merging.

Proven against the real packer: the same 22 files now weigh 103.91 and split
into two chunks. No logic, budget, or timeout was changed; raising a budget
to make a red gate pass is not a fix.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:18:18 -04:00
Tom Boucher
d83e58eea0 fix(#437,#439,#440): restore defaults.run.shell + 'zsh {0}' format + Windows .cmd shell:true (PR #434 fallout) (#438)
* fix(#437): restore defaults.run.shell at job level (step-level matrix expr rejected by GHA)

Per actions/runner workflow-v1.0.json schema, `jobs.<job_id>.defaults.run.shell`
allows `matrix` context (job-defaults-run has context:[matrix,...]); step-level
`shell:` does not (run-step's shell field is plain string with no context array).
PR #434 used step-level shell:${{matrix.shell}}, which GHA's parser rejects with
"Unrecognized named-value: 'matrix'" — blocking every push to next and every
release.yml dispatch.

This commit:
- Removes step-level `shell: ${{ matrix.shell }}` from test-full (test.yml)
  and smoke (install-smoke.yml) jobs (17 directives).
- Adds `defaults.run.shell: ${{ matrix.shell }}` at job level in those two jobs.
- Fixes pre-existing shellcheck SC2129 in test.yml (individual >> redirects →
  grouped brace form) and SC2010 in install-smoke.yml (ls|grep → glob loop).

Verified locally with actionlint 1.7.12 (exit 0). Policy linter still 0 violations
(matrix.shell now resolves via job.defaults.run.shell which the linter already
handles per workflow-policy.cjs:effectiveShell).

Refs: actions/runner#444 (open since 2020), GHA contexts page section "Context availability".

* fix(#439): inline ci-smoke-skip back to shell (Node port required pre-checkout file resolution)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#440): use platform-correct npm.cmd on Windows for spawn (and surface-check other Node ports)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): use 'zsh {0}' format string in matrix.shell for macOS (zsh not in GHA built-ins)

Per https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions
(jobs.<job_id>.defaults.run.shell section):

  "You can use built-in shell keywords like bash, pwsh, python, sh, cmd, and
  powershell, or define a custom set of shell options."

zsh is not in the built-ins list. GHA accepts custom shells via a format string
containing '{0}', which it replaces with the temporary script file path at
runtime (same pattern as the perl {0} example in the docs).

Bare `shell: zsh` triggers: "Invalid shell option. Shell must be a valid
built-in or a format string containing '{0}'".

Precursor: 514cb429 introduced the matrix shell-pinning pattern; this completes
it by switching the macOS rows from the bare value to the required format string.

Also updates scripts/workflow-policy.cjs to normalise 'zsh {0}' to 'zsh' before
the policy comparison, so the repo-baseline test continues to pass (the linter
was correctly treating 'zsh {0}' as a distinct value from the policy 'zsh').

Affects:
- .github/workflows/test.yml: test-full matrix (node 22 + node 24 macOS rows)
- .github/workflows/install-smoke.yml: smoke matrix (macOS node 24 row)
- scripts/workflow-policy.cjs: detectViolation strips ' {0}' format suffix

* fix(#440): add shell:true to spawnSync on Windows for .cmd files (Node docs requirement)

Per https://nodejs.org/docs/latest-v22.x/api/child_process.html:

  ".bat and .cmd files require a terminal to run and cannot be launched
  directly with execFile(). To run these scripts on Windows, use
  child_process.spawn() with the shell option, child_process.exec(), or
  spawn cmd.exe with the script as an argument."

  "On Windows, .bat and .cmd files require a shell to execute. Use
  child_process.exec() or child_process.spawn() with the shell: true option."

On Windows, npm is installed as npm.cmd (a batch wrapper). Without
shell: true, spawnSync resolves the binary directly and fails with
ENOENT / "npm binary not found on PATH" because the OS cannot execute
a .cmd file without cmd.exe as the intermediary.

The fix uses `shell: process.platform === 'win32'` so the shell spawning
is only activated on Windows; macOS/Linux continue to resolve the plain
npm binary directly with shell: false, preserving the existing behaviour
on non-Windows platforms.

Updated both spawnSync(npmCmd, ...) call sites:
- npm --version check (line 182)
- npm ci --dry-run lockfile-sync check (line 215)

* fix(#437): bug-410 defaults test — set USERPROFILE for Windows os.homedir() redirect

On Windows, os.homedir() reads USERPROFILE (not HOME), so the test's
process.env.HOME = FAKE_HOME redirect was silently ignored. finishInstall's
path.join(os.homedir(), '.gsd') resolved to the real user home and the
defaults.json write either failed (permissions) or landed outside the temp
dir, causing the existsSync assertion to return false.

Fix: also set process.env.USERPROFILE = FAKE_HOME so os.homedir() returns
the sandboxed directory on Windows. Node.js docs (os.homedir):
https://nodejs.org/docs/latest-v22.x/api/os.html#oshomedir

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): precommit-alias-drift hook test — use path.delimiter for PATH

Hardcoded ':' PATH separator breaks Windows where process.env.PATH uses ';'.
The malformed PATH passed to bash caused the mock git/npm stubs in binDir
to be invisible to the hook script; npm was never called and the marker
file never written.

Fix: replace ':' with path.delimiter in both PATH constructions so the
env var is well-formed on Windows (';') and POSIX (':') alike.

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): prepush-enterprise-email hook test — use path.delimiter for PATH

Same root cause as precommit-alias-drift: hardcoded ':' PATH separator is
invalid on Windows (';' required). The malformed PATH meant bash ran the
real git binary instead of the mock stub, which rejected the placeholder
SHAs 'refs-local-sha' / 'refs-remote-sha' with a fatal ambiguous-argument
error rather than returning the fixture commit list.

Fix: replace ':' with path.delimiter in both execFileSync PATH env values.

Refs: #437 (fix/437-restore-defaults-run-shell), Windows pwsh compat

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): set MSYS2_PATH_TYPE=inherit so mock stubs take precedence in Git Bash PATH

Root cause: Git Bash (MSYS2) on Windows prepends its own system directories
(/mingw64/bin, /usr/bin, /bin) to the PATH at process startup before the
user-supplied Windows PATH entries. This placed the real git/npm binaries
ahead of the mock stubs in binDir even though binDir was first in the Windows
PATH passed to execFileSync. The path.delimiter fix (0042fe0d) made the PATH
syntactically correct for Windows (semicolons) but did not change the MSYS2
system-dir prepend order.

The real git rejected placeholder SHAs (refs-local-sha, refs-remote-sha) with
"fatal: ambiguous argument", producing the observed Windows CI failure. For the
pre-commit test, the real git output nothing (no staged files on a fresh
checkout), so the grep match failed and npm was never called.

Fix: set MSYS2_PATH_TYPE=inherit in the env passed to both bash spawns.
With inherit, MSYS2 uses only the converted Windows PATH without prepending
system directories, so binDir (converted from Windows to POSIX) is first in
the search path and the mock stubs are found.

grep/tr/printf remain available: the GHA Windows runner PATH includes
C:\Program Files\Git\usr\bin which contains these utilities; MSYS2 converts
that Windows entry to a POSIX path on startup. The /usr/bin/env shebang in
mock stubs resolves through MSYS2's virtual filesystem mount (not via PATH)
and is always accessible regardless of MSYS2_PATH_TYPE.

On macOS/Linux this variable is ignored; no behaviour change on those platforms.

Source: https://www.msys2.org/wiki/MSYS2-introduction/#path
(MSYS2_PATH_TYPE controls whether system dirs are prepended to converted PATH)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#437): hook test mocks — use cmd-shim pattern for Windows bin resolution

On Windows, bash (Git Bash / MSYS2) resolves PATH commands by scanning for
extensionless files, but cmd.exe and Win32 process creation resolve via
PATHEXT (.CMD, .BAT, .EXE). When execFileSync('bash', [hookPath]) runs a
hook that calls `git` or `npm`, both resolution paths may fire. The previous
approach set MSYS2_PATH_TYPE=inherit in the child env, but that variable is
only read in /etc/profile (login-shell path) — bash launched without --login
never sources /etc/profile, so the variable had no effect:
https://github.com/msys2/MSYS2-packages/blob/master/filesystem/profile

Fix: adopt the cmd-shim three-file pattern used by npm itself:
https://github.com/npm/cmd-shim
For each mock binary, write:
  <name>          extensionless bash script (bash PATH scan)
  <name>.cmd      batch wrapper delegating to bash (PATHEXT / cmd.exe)
  <name>.ps1      PowerShell wrapper (completeness)

This is the same approach used by stevemao/mock-bin for test mocking with
Windows CI green on AppVeyor:
https://github.com/stevemao/mock-bin

The .cmd and .ps1 files are only written on process.platform === 'win32'.
MSYS2_PATH_TYPE is removed from the child env — it was ineffective and is
no longer needed with the shim files in place.

* fix(#437): tarball-smoke — raise CHILD_TIMEOUT_MS on Windows to 600 s

The CI failure showed a test duration of 120003.1812 ms — matching the
previous CHILD_TIMEOUT_MS = 120_000 exactly. When spawnSync hits its
timeout, it sends SIGTERM and returns { status: null, stdout: '', stderr: '' }
per the Node.js docs:
https://nodejs.org/docs/latest-v22.x/api/child_process.html
  "status: <number> | <null> — The exit code of the subprocess, or null if
   the subprocess terminated due to a signal."

The installResult check is `status !== 0`; null !== 0 is true, so the
timeout fired the INSTALL_FAILED path with empty stdout/stderr, which made
the root cause invisible in CI logs.

GitHub-hosted Windows runners are slower than Linux/macOS for
filesystem-heavy operations (npm install -g of a 1499-file tarball):
https://docs.github.com/en/actions/using-github-hosted-runners/about-github-hosted-runners/about-github-hosted-runners#standard-github-hosted-runners-for-public-repositories

Fix: use 600_000 ms (10 min) on Windows, keeping 120_000 ms on POSIX.
600 s matches the SLOW_HOST_TIMEOUT already used in the test before() helper
for the pack + install fixture step.

Also expose `signal` and `installError` in the INSTALL_FAILED details object
so a future timeout (status=null, signal='SIGTERM', stdout='') is immediately
diagnosable in CI logs without guesswork.

* fix(#437): chmod +x via bash on Windows for hook test mocks (root cause: fs.writeFileSync mode=0o755 no-op on NTFS)

Root cause: Node's fs.writeFileSync mode=0o755 is a no-op for the execute
bit on Windows NTFS. Per https://nodejs.org/docs/latest-v22.x/api/fs.html:
"on Windows only the write permission can be changed." Bash's access(X_OK)
therefore skips the mock file; the real git/npm binary is found later in PATH
and the hook runs against real state instead of the test double.

Fix: after writeFileSync, invoke Git Bash's chmod via the POSIX emulation
layer (Cygwin/MSYS2), which sets the NTFS execute ACL that Node cannot reach:

    const posixPath = filePath.replace(/\\/g, '/');
    execFileSync('bash', ['-c', `chmod +x "${posixPath}"`], { stdio: 'pipe' });

execFileSync('bash', ...) works because Git for Windows ships bash on PATH in
all GHA Windows runners. Forward-slash conversion is required because MSYS2
bash auto-converts /c/foo paths but not mixed-separator paths.

Why prior approaches didn't take effect:
- MSYS2_PATH_TYPE=inherit: only read in /etc/profile (login-shell path);
  execFileSync('bash', ...) launches non-interactively without --login, so
  /etc/profile is never sourced.
  Ref: https://github.com/msys2/MSYS2-packages/blob/master/filesystem/profile
- .cmd/.ps1 cmd-shim wrappers: bash does POSIX command resolution and does
  not honor PATHEXT, so wrappers are not found by bash's own PATH scan.
  They are not wrong (kept for non-bash callers) but do not fix bash's X_OK.

Files changed: tests/precommit-alias-drift-hook.test.cjs,
               tests/prepush-enterprise-email-hook.test.cjs

* refactor(#437): hooks use GIT_OVERRIDE/NPM_OVERRIDE env-var DI; tests drop PATH-mocking

Four prior rounds (path.delimiter join, MSYS2_PATH_TYPE=inherit, cmd-shim
.cmd/.ps1 wrappers, chmod-via-bash post-write) all failed to make MSYS2
bash's PATH-lookup find the mock executables. The root cause is that none
of those approaches can reliably override bash's own command-resolution
on NTFS without fighting NTFS execute-ACLs or login-shell profile sourcing.

The simplest robust solution is to bypass PATH entirely:

Hooks: each hook now binds GIT_CMD="${GIT_OVERRIDE:-git}" (and NPM_CMD for
pre-commit) at the top. When env vars are unset the hooks invoke bare
`git`/`npm` exactly as before — zero behavior change for users.

Tests: writeMockBin/binDir/PATH manipulation replaced by writeMock(), which
writes a .sh mock to a tmpDir and passes its absolute path via GIT_OVERRIDE
/ NPM_OVERRIDE in the execFileSync env. Bash inside the hook executes the
path directly via the seam — no PATH scan, no NTFS ACL check, no MSYS2
profile dependency.

Test-rigor principle: the new seam (env-var injection) is platform-
independent and doesn't rely on bash's command-resolution mechanism on
the host OS.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 18:06:19 -04:00
Tom Boucher
006cdafe8f ci(drift): enforce alias freshness checks in CI and contributor flow (#2910)
Merging alias-drift guardrails and local hook hardening.
2026-04-30 14:19:46 -04:00