Files
msd-core/docs/adr/4641-windows-selector-consolidation.md
Tom Boucher 4d65c248e5 fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector

Tests only, committed ahead of the implementation so the RED run is real.

- tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a
  ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative
  cases proving seam calls and path-call-plus-slash-literal are not platform
  signals; positive pins that genuine platform content, seam-bypassing spawns,
  chmod and symlink still classify in; macOS signal set and generated list
  unchanged.
- tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest
  rows and test-conformance still has 3 windows + 1 macOS.
- tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a
  non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint
  test does, and resolveSelection rejects the retired windows scope.

Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): delete the second Windows selector and narrow the conformance tier

Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped
conformance tier on real Windows/macOS — was not met. Measured on PR #4640
(run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7
changed test files running on a real Windows runner twice.

Two selectors, only one in the epic's scope. The test job's three scope:windows
shards predate the epic (#494, sharded #3057) and gate on product_changed, not
full_matrix, so they fire on every product PR whatever Phase 3's classifier
decides. They are deleted; test-conformance becomes the sole Windows selector,
as it already was for macOS. Non-Linux jobs 7 -> 4.

Gating the lane instead was rejected as provably redundant: for a test file
reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file),
and that same predicate sets full_matrix, which turns test-conformance on. Every
file a gated lane would run is already covered in the same run. The lane's one
non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename
heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix
instead of feeding a parallel lane.

Two detectors matched the repo's own test idiom rather than any platform signal
and carried 226 of the tier's sole-signal membership against 41 for the other
eight: process-seam-subprocess (335 files, 118 unique) matches the
tests/helpers.cjs entry points nearly every CLI test uses, and going through the
seam is the opposite of a platform signal since shell-command-projection takes
platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs
only a path call anywhere plus a slash literal anywhere, and that class is
already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed.
Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured.

Adds the size gate Phase 2 never had, as a ratio against a live denominator so
it cannot stop binding as the suite grows.

292 files leave real-OS Windows execution. The drop-out set was audited: 14 have
a platform-suggestive filename and all 14 are static source-text analyses or
seam-mediated CLI tests. raw-child-process was investigated as a suspected false
negative and left unchanged -- relaxing it adds 13 files, all false positives.

macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated
macos-conformance-tier.generated.cjs is byte-identical at 196 files.

Fixes #4641
Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): register the new ADR path in the docs-guard exempt baseline

tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md
in a comment justifying the retired windows scope; lint-docs-guard-registration
tracks that reference set, so the baseline needs the new path. Verified the
exemption still holds: the path is prose, not a filesystem read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): make the escalation tier-backed and drop every hardcoded count

Three follow-ups from measuring the first pass rather than trusting it.

The windows-hint escalation now requires tier membership as well as the
filename hint. Setting full_matrix runs test-conformance, which runs only the
tier; escalating on a test that is NOT in the tier costs four jobs and still
never runs that test on Windows. Measured over the 16 RULES entries the
narrowed predicate fires on exactly the same rules today, so this is
correct-by-construction rather than a behavior change. The broader variant --
escalate on any tier member a rule pulls in, ignoring the hint -- was measured
at 14/16 rules and rejected as over-broad.

Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as
soon as the rebase pulled in one new test file from #4253, which is the whole
argument against them: the ceilings are ratios against a live denominator, the
committed lists are pinned by comparison against a fresh classification of the
live tree, and the three named probe files now assert on their SIGNAL rather
than on membership in a literal list -- asserting by filename is the exact
error this PR fixes in the classifier.

Regenerates both lists against the rebased tree. Same-tree figures are now
547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none
added; macOS is unchanged at 197 with a zero-line diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke

An isolated adversarial review found a real false negative. Removing the
blanket process-seam-subprocess detector also removed the only coverage for
tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook
spawns options.interpreter via real spawnSync, so
runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary
executing a shell script extracted from workflow markdown. The seam argument
holds for src/shell-command-projection.cts, which takes platform as an
injected parameter; it does NOT hold for the test helpers, which spawn real
binaries. Conflating the two is what made the blanket detector look purely
noisy -- it was 99% noise wrapping a real signal.

Adds a narrow shell-interpreter-spawn category keyed on a real interpreter
option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are
added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33%
ceiling. All 9 confirmed by reading the matching source line, zero comment or
fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured
and rejected -- each adds 9 files but misses the counterexample entirely.

Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its
scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD
strictly rather than carrying the Windows report-only carve-out. The carve-out
now lives solely on test-conformance, whose matrix does include windows.

Repairs seven pre-existing assertions the category removal invalidated,
preserving each case's purpose rather than deleting coverage, and converts the
last hardcoded tier bounds to live-derived ratios -- including the macOS
sanity range that was still a magic [100, 350].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): keep the confinement test on a real OS via a documented allowlist

A security review found tests/external-descriptor-confinement.test.cjs had
dropped out of the Windows tier. It must stay in, and no content signal can
express why: it exercises isPathConfined (src/external-descriptor-trust.cts),
which uses the AMBIENT path module -- path.resolve(root, target) and path.sep
-- with no injection. Its win32 semantics (drive letters, UNC, separator) are
only reachable by actually running on Windows, and it is a security-relevant
write-confinement gate. A content classifier cannot see 'this module reads the
ambient path module', so no regex belongs here.

Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows
tier only. A Map rather than a list so an entry without a reason is impossible
by construction, and tests assert every entry names a file that exists on disk
so a stale entry fails loudly instead of rotting. This is the centrally-
enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's
portability-vocab.cjs already models -- deliberately not a heuristic.

Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is
untouched and byte-identical: the win32 concern does not apply to a POSIX
runner, and a test asserts the allowlist does not leak into that tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): inject the path impl into isPathConfined and correct the ADR count

Two review findings, both fixed rather than dispositioned.

A security review found tests/external-descriptor-confinement.test.cjs had left
real-OS execution. The allowlist pinned it back, but that only restored
INCIDENTAL coverage: isPathConfined used the ambient path module, and its test
carried POSIX-only literals, so a win32 confinement escape was unverified on
every platform including Windows. isPathConfined now takes an optional third
parameter carrying the path implementation, defaulting to the ambient module.
Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change
is purely additive and every existing two-argument caller is byte-identical.

Tests now inject path.win32 and path.posix, covering a different drive letter,
a cross-drive absolute, backslash and forward-slash traversal, UNC, and the
startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators.
Proved load-bearing: dropping the + p.sep from the prefix check fails exactly
the two boundary cases and nothing else. Callers' suites 149/149.

The spec review caught an off-by-one: the ADR narrated a 264-file tier while the
committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264
-> 265 (28.5%).

Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated
classify() as zeroing windows_tests, a key this change removes -- kept as
historical narration but labelled as such.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#131): make the unwritable-HOME test actually test something

Found by sweeping for the root-bypass class after fixing commit-files-deletion.
This one is the silent variant, and it was broken twice over.

First, the condition: the test made a fake HOME unwritable with chmod 0o500.
The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed
writable and the hostile condition never existed. Replaced with a HOME whose
PARENT is a regular file, so every write under it fails ENOTDIR at the VFS
layer for every uid -- no permission check is involved at all.

Second, and more fundamental: the probe was npm --version, which on npm 11.19.0
performs zero filesystem I/O against HOME. Proven rather than assumed --
neutralizing runNpm()'s isolation turned the sibling test red while this one
stayed green, so its assertion could never detect the regression it guards, on
any uid, with or without the condition fix. npm config get cache was tried next
and proved vacuous the same way (it only string-resolves the path). The probe is
now npm cache verify, which really does mkdir _cacache under HOME.

Re-proved load-bearing after the change: with isolation neutralized the test now
fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was
restored and verified diff-clean; suite 13/13.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the net drop-out figure in ADR-4641

The Consequences section still said 292 files leave real-OS Windows execution.
That was the count before the narrow shell-interpreter-spawn replacement
restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary
categories rather than only raw-child-process, and clarifies that the 14-file
filename audit was against the 292 initially dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the rejected concentration ceiling and its measurement

Applying Goodhart's own question to the new ceiling -- how would you make this
metric look good without improving what it represents -- surfaces a real
weakness: a ratio can be satisfied by inflating the denominator, so adding
OS-agnostic tests loosens it without narrowing the tier.

The obvious companion gate was a sole-signal concentration ceiling, since the
original defect was one detector carrying half the tier. Measured and rejected:
peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the
historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the
original defect; any threshold below it fails on a legitimate category. The
discriminator is whether a signal is platform-meaningful, which no threshold
encodes. Weakness disclosed rather than covered by a gate that does not bind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4641): add the changeset fragment for the confinement-check change

changeset-lint failed on PR #4643: the PR touches user-facing paths and carried
no fragment. The earlier no-changeset call matched #4604's CI-only precedent and
was correct then; it was not revisited once the PR grew a src/ change, which is
my miss.

The fragment describes the real user-visible improvement: the external-descriptor
write-confinement check's Windows semantics are now verified deterministically
rather than only when the suite happened to run on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the tier count in TESTING-SUITES.md

Said the tier narrowed from 546 to 254. The final committed list is 265 of 931
eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored
9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in
the ADR, in a live reference page rather than a dated record, so it states the
current truth rather than carrying an amendment note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the measured aggregate from real CI job lists

Epic #4589's closeout asserted its reduction from a static count; #4641's
acceptance criterion asks for a figure read off a real run. Recorded here:
test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run
against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4.

Also states the caveat that a PR's total CHECK count is not a clean before/after
comparison, since many gates are path-scoped and this change touches a broader
path set -- the like-for-like figure is the test.yml job count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): compare job totals the same way on both sides

The measured-aggregate table put #4640's COMPLETED run total (21) against this
run's count at matrix-expansion time (15). Those are not the same measurement:
the completed total includes the post-test Coverage gate and baseline-publisher
jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux
jobs 7 -> 4, was correct and is unchanged.

Called out in the table rather than silently corrected -- comparing two
differently-derived numbers is exactly the error class this ADR is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record measured conformance wall-clock and date the stale counterfactual

Adds the per-job durations from both runs. The honest read is that this is a
correctness win more than a speed one: file count fell 52% but wall-clock only
9-29%, because what was removed were the cheap static tests and what remains is
concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects
a future narrowing to buy time proportional to file count.

The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on
the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about --
pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its
tier was UNCHANGED at 197 files, which fixes that as runner variance and is
noted as a caution against reading a single duration as signal.

Also dates the symlink-keyword counterfactual, which cited a 254-file tier from
before the replacement category and allowlist took it to its final 265.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero

next gained #4644 mid-flight, so every absolute count shifted. Re-measured on
the tree this actually ships against (932 eligible): 548 -> 257 by detector
removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed,
9 restored. macOS 198, unchanged by this PR.

The percentages did not move across three rebases (58.8% -> 28.5%), which is
the whole argument for expressing the ceilings as ratios rather than counts --
noted in the ADR since it is now evidence rather than assertion.

Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own
win32 test cases introduced the literal win32 into the pinned file, so it
classifies in on content via win32-darwin-literal. The entry stays and the
reason is written down, because the file's real-OS need is a property of the
code under test (isPathConfined reads the ambient path module), not of the
test's text -- the text that currently saves it is incidental and could be
refactored away silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 17:00:11 -04:00

23 KiB

ADR-4641: One Windows test selector, and a proportional ceiling on the conformance tier

  • Status: Accepted
  • Date: 2026-09-11
  • Issue: #4641
  • Twin of: ADR-4593 (macOS conformance tier) — that ADR applied evidence-first sizing to the macOS signal set; this one applies the same discipline to the Windows signal set and to the number of selectors, which ADR-4593 did not cover.
  • Closes a gap in: epic #4589, whose goal — "the OS-agnostic majority of the suite runs once, on Linux, while a small and explicitly-scoped platform-conformance tier covers genuinely OS-specific behavior" — was not achieved by its five merged phases.

Context

Epic #4589 closed 2026-09-10 with every acceptance box checked. Measured on PR #4640 (run 34618834118), a routine tests/**-touching PR, two things were true that the goal text forbids.

Two independent Windows selectors

selector gate what it runs
test job, three scope: windows rows product_changed windows_tests = every changed tests/*.test.cjs, unconditionally
test-conformance, three windows shards code_changed && full_matrix the generated conformance tier

The epic replaced test-full with test-conformance and never touched the first selector, which predates it (#494, sharded in #3057). On PR #4640, 5 of 7 changed test files ran on a real Windows runner twice. Non-Linux job count was 7, not the 4 the epic's closeout reported — that figure compared test-full (6) against test-conformance (4) and omitted the three always-on scope: windows shards from both sides. Counting every non-Linux job, the epic moved 9 → 7, not 6 → 4.

The tier was 59% of the suite

node scripts/gen-platform-conformance-tier.cjs reported 546 of 930 eligible unit-suite files (58.7%) when this was diagnosed. Measured on the tree this PR actually ships against (932 eligible, after #4253 and #4644 landed on next mid-flight) the same comparison is 548 → 257 by detector removal alone. Absolute counts drift every time next gains a test file; the percentages did not move at all across three rebases (58.8% → 28.5%), which is the whole reason the ceilings are ratios. Measured per-category contribution, where UNIQUE is the count of files for which that category is the sole signal — i.e. the marginal cost of keeping it:

category total UNIQUE
process-seam-subprocess 335 118
hardcoded-path-vs-path-call 328 108
raw-child-process 96 19
symlink-keyword 86 6
win32-darwin-literal 77 1
process-platform 73 1
chmod-mode-bit 73 4
windows-env-var 69 6
windows-shell-token 23 4
os-platform 2 0

Two categories carried 226 of the tier's sole-signal membership; the other eight carried 41 combined.

  • process-seam-subprocess matches runNode( / runGit( / runHook( / runGsdTools( / gitOrThrow( — the repo's own tests/helpers.cjs entry points, used by nearly every CLI test. For the overwhelming majority of them, going through the seam is the opposite of a platform signal: src/shell-command-projection.cts takes platform as an injected parameter, and tests/shell-command-projection-dispatch.test.cjs already exercises PowerShell/cmd.exe/PATHEXT in-process on Linux by passing platform: 'win32' as data. That is epic #4589's own argument for why the cutover was safe, applied against itself. But see the Consequences section: this detector was 99% noise wrapping a real signal — the ~9 tests that spawn a real shell via runHook's interpreter option — and that signal is preserved by a narrow replacement category rather than lost with the blanket one.
  • hardcoded-path-vs-path-call requires a path.join|resolve|…( call anywhere in the file AND a quoted '/…' literal anywhere in the file, with no proximity. In a Node test suite both are universal. The defect class it gestures at is already enforced by Linux-runnable ESLint rules under ADR-1703 (no-hardcoded-tmp, no-path-literal-in-assert).

The repo had already reached this conclusion for one consumer and not the other. gen-platform-conformance-tier.cjs carried:

// Two CATEGORIES entries precise enough for TEST-file classification (this
// module's own purpose) but far too broad for SOURCE-file reachability
const NOISY_FOR_SOURCE_REACHABILITY = new Set(['hardcoded-path-vs-path-call', 'symlink-keyword']);

Measured there: classifyContent over src/ flagged 100 of 235 files; excluding these two narrowed it to 28, "all verified to carry a genuine platform-conditional branch." The same over-breadth verdict was reached, recorded, and then not applied to the tier itself.

Why no gate caught it

Phase 2's acceptance criterion was "conformance-tier file list exists as a single source of truth." It checked that the list exists. No phase asserted it was small, and no test would have failed if the classifier had put all 930 files in. The #4591 per-file parity-baseline diff that would have caught it was explicitly never performed — disclosed in the generator's own header as a KNOWN LIMIT — and Phase 5 then retired test-full, the safety net that disclosure named as its compensating control.

Decision

1. Delete the scope: windows lane; port its residue to reachability

The alternative — gating windows.add(file) on reachesConformanceTierOrSeam — was evaluated and rejected as provably redundant. For a test file that predicate is literally:

if (file.startsWith('tests/') && file.endsWith('.test.cjs')) {
  const { CONFORMANCE_TIER_FILES } = loadConformanceTier();
  return CONFORMANCE_TIER_FILES.includes(file);
}

and the same predicate is what sets full_matrix, which is what turns test-conformance on. So every file a gated lane would run is (a) already in the tier and (b) has already caused the conformance lane to run the whole tier in the same workflow run. Gating does not reduce the duplication; it makes it total.

The lane's one non-redundant contribution is the isWindowsHint arm — tests pulled in by a path RULE whose filename contains windows/win32/shell/path. That is a filename substring heuristic, precisely the kind of unproven heuristic #4592 replaced with reachability. It is therefore ported into reachesConformanceTierOrSeam: such a RULE-pulled test now sets full_matrix = true, and the conformance lane covers it. The signal is preserved; the parallel lane is not.

The escalation is tier-backed, and that condition is load-bearing. Three variants were measured over the 16 entries of RULES:

variant predicate rules firing verdict
A isWindowsHint(t) 6/16 Can fire on a test that is not in the tier — full_matrix goes true, the conformance lane runs, and the hinted test still never runs on Windows. Cost without coverage.
B (shipped) isWindowsHint(t) && reachesConformanceTierOrSeam(t) 6/16 Identical firing set to A today, so no behavior change — but correct by construction: it can only escalate when the conformance lane will actually run the file.
C reachesConformanceTierOrSeam(t) alone 14/16 Rejected as over-broad. Would newly escalate most ordinary product-code PRs (src/, agents/, commands/, hooks/, skills/, config paths) — tier membership alone is too weak a trigger.

A and B coincide only because every windows-hint test currently pulled in by a rule happens to be in the tier except one (tests/normalize-path-in-content.rule.test.cjs, which has zero signals). B is shipped because that coincidence is not an invariant.

Live effect is deliberately small: of the six rules that fire, four already set fullMatrix: true (no-op), inert CI's escalation is overridden downstream by the inert-CI reset, and exactly one — portability lint rules (ADR-1703) — genuinely changes behavior.

test-conformance becomes the sole Windows selector, matching how it already is the sole macOS selector.

2. Remove the two house-idiom detectors, and add one narrow replacement

process-seam-subprocess and hardcoded-path-vs-path-call are deleted from CATEGORIES. hardcoded-path-vs-path-call leaves NOISY_FOR_SOURCE_REACHABILITY with it (the set now holds symlink-keyword alone). Tier, all measured on one tree (932 eligible): 548 → 257 (58.8% → 27.6%) by removing the two detectors, then 257 → 266 (28.5%) once the narrow shell-interpreter-spawn replacement added 9 genuinely shell-spawning tests back. Net: 282 files removed, 9 restored. ALWAYS_REAL_OS currently adds 0 — see Consequences for why it is still there.

Measured, the change is surgical: src/ reachability is 28 → 28, zero files change status, because hardcoded-path-vs-path-call was already excluded there and no src/ file matches the test-helper regexes. Phase 3's classifier behavior for src/ diffs is provably unchanged.

3. A proportional ceiling, asserted failing-first

The missing Phase 2 gate is added as a test, and it is expressed as a ratio against a live denominator, not a count:

tier measured ceiling
Windows (CONFORMANCE_TIER_FILES) 28.5% 33%
macOS (MACOS_CONFORMANCE_TIER_FILES) 21.2% 25%

An absolute count goes stale as the suite grows and silently stops binding; the property that matters — "a tier, not the suite" — is inherently proportional. The ceiling is deliberately not today's emitted value, which #4641 rules out explicitly as a non-bound.

Consequences

  • Non-Linux jobs on a full_matrix PR: 7 → 4. Against the true pre-epic baseline of 9, epic #4589 plus this ADR deliver 9 → 4 (-56%), versus the -33% its closeout claimed against a denominator that excluded this lane.

    Measured, not computed — read off real job lists rather than derived from the workflow file, which is the verification epic #4589's own closeout skipped:

    PR #4640 (the trigger) PR #4643 (this change)
    jobs in the completed test.yml run 21 17
    non-Linux jobs 7 4
    test job 4 ubuntu + 3 windows 4 ubuntu, 0 windows
    conformance tier size 548 files 266 files

    Both job totals are counted the same way — every job in the completed run, which includes the post-test Coverage gate and baseline-publisher jobs. An earlier draft of this table compared #4640's completed total against this run's count at matrix-expansion time, before those trailing jobs exist; that is an apples-to-oranges comparison and the kind of error this ADR is otherwise about, so it is called out rather than quietly corrected.

    One caveat stated rather than glossed: a PR's total check count is not a clean before/after, because many gates are path-scoped and this change touches a broader path set than #4640. The like-for-like figure is the test.yml job count and its non-Linux portion, which is what the epic's goal was about.

    Wall-clock, measured on both runs — and the honest read is that this is a correctness win more than a speed one:

    conformance job #4640 (548-file tier) #4643 (266-file tier)
    windows shard 1/3 29m47s 21m12s -29%
    windows shard 2/3 29m00s 26m21s -9%
    windows shard 3/3 40m24s 31m27s -22%
    macOS 17m48s 21m02s +18%

    File count fell 52% but wall-clock only 9-29%, because the files removed were the cheap static ones — the tier that remains is concentrated in genuinely expensive spawn-heavy work, which is exactly what it should contain. Do not expect a future narrowing to buy time proportional to file count. The macOS figure moved the wrong way while its tier was unchanged by this PR (198 files; it tracks next's test count, not this change), which fixes it as runner variance rather than an effect of this change, and is a caution against reading any single duration as signal.

    The load-bearing number is shard 3/3: it ran at 40m24s against a 45-minute cap, 90% of the cliff that #869 and #3057 were both filed about. Pulling it to 31m27s restores real headroom.

  • 282 test files leave real-OS Windows execution — 291 dropped when the two detectors were removed, 9 restored by the narrow shell-interpreter-spawn replacement. This is a real coverage change, not a refactor. It is defensible because every file that stays out does so by losing a signal that was never a platform signal — each remains covered by the Linux run, and the files that genuinely spawn a real binary are untouched or restored (raw-child-process, 96 files; shell-interpreter-spawn, 33).

    The drop-out set was audited rather than assumed. Of those initially dropped, 14 had a filename suggesting platform relevance (/windows|win32|shell|path|platform|posix|crlf|symlink|exec|spawn|subprocess/i), and each was inspected. Six carry an explicit allow-test-rule: source-text-is-the-product or structural-regression-guard marker; the rest were read individually.

    That audit initially reached the wrong conclusion, and the correction is the most important thing in this ADR. Its first pass concluded all 14 were static analyses or seam-mediated CLI tests. An adversarial review found a counterexample by reading call semantics rather than filenames: tests/execute-phase-worktree-guard.test.cjs calls

    runHook('-c', [guardScript()], { interpreter: 'bash', cwd: dir, … })
    

    and tests/helpers/process-seam.cjs's runHook spawns options.interpreter through a real spawnSync. With interpreter: 'bash' that is a real bash binary executing a shell script extracted from workflow markdown, doing real git plumbing — bash availability, quoting, and git output parsing all differ on Windows. No injected-platform unit test stands in for that.

    The seam argument therefore needs a boundary it did not originally state. "Going through the seam is not a platform signal" is true of src/shell-command-projection.cts, which takes platform as an injected parameter. It is not true of tests/helpers/process-seam.cjs, whose runHook/runGit spawn real binaries. Conflating the two is what made the original process-seam-subprocess detector look purely noisy: it was 99% noise wrapping a real signal.

    On ALWAYS_REAL_OS adding zero today — disclosed, not hidden. The allowlist holds one entry, tests/external-descriptor-confinement.test.cjs, and it currently contributes 0 files, because this PR's own win32 test cases introduced the literal win32 into that file and it now classifies in on content via win32-darwin-literal. A future reader measuring the allowlist's marginal contribution will get zero and may conclude the mechanism is dead. It is not, and the entry stays: the file's real-OS need is a property of the CODE UNDER TEST — isPathConfined reads the ambient path module — not of the test's text, and the text that currently saves it is incidental. Rewrite those cases to use a helper without the literal and the file drops out silently. The pin exists precisely for that, and the tests assert every entry names a file that exists so a stale entry fails loudly rather than rotting.

    The fix is a narrow replacement category rather than restoring the blanket one:

    { name: 'shell-interpreter-spawn',
      test: (c) => /interpreter:\s*['"`](bash|sh|zsh|dash|pwsh|powershell|cmd)['"`]/.test(c) }
    

    Measured 2026-09-11: 33 eligible files match, 9 of them were outside the tier and are added back, taking it from 257 to 266 of 932 (27.6% → 28.5%), which is the committed total. Still under the 33% ceiling. Every one of the 9 was confirmed by reading the matching source line — all are live interpreter: options on real runHook/runHookSeam calls, zero comment or fixture matches. Two narrower alternatives (runGit( alone; non-node spawnSeam() were measured and rejected: each adds 9 files but misses the counterexample entirely, because it spawns through runHook's interpreter option rather than through runGit.

    The lesson is recorded deliberately: an audit that selects candidates by filename inherits exactly the defect this ADR is fixing in the classifier. The 14-file filename sweep was the right first cut and the wrong last word.

    The worked example is tests/windows-robustness.test.cjs, which was on this ADR's own first-draft "must remain in the tier" list because of its filename. It does not spawn anything: it reads other files' source text and asserts on it (assert.match(region, /windowsHide:\s*true/)), and its apparent spawnSync( / execFileSync( occurrences are string literals used as search anchors into those other files. It is fully Linux-runnable and correctly drops out. Selecting it by name would have been the same error the classifier makes — and a test now pins that it drops out, with the reason, so nobody "fixes" it back in.

    The clearest statement of this ADR's thesis is one the repo already wrote. tests/hardcoded-paths.test.cjs, itself a drop-out, opens: "Statically scans source files to catch hardcoded platform-specific paths… Catches issues that previously required a real Windows runner to detect."

  • classify() no longer returns a windows_tests key; ci-prepare-test-scope.cjs no longer accepts a windows scope. Both are removed rather than left inert, so a future reader cannot mistake a dead output for a live one.

  • ADR-4593's decision is unaffected: MACOS_CATEGORIES is a separate array, chmod-mode-bit and symlink-keyword keep their recorded rationale and their definitions, and the macOS tier is unchanged — the regenerated macos-conformance-tier.generated.cjs is byte-identical to the one on next (git diff reports zero changed lines). Its five prose citations of the 546 figure are left as written: they were accurate on 2026-09-10 and an ADR is a dated record, not a live reference page. ADR-4593 instead carries a short amendment note pointing here, so a reader who arrives at the 546 figure learns it has since moved without the original reasoning being rewritten underneath them.

    Neither tier's size is asserted as a literal count anywhere in the test suite: the ceilings are ratios against a live denominator, and the macOS list is pinned by comparing the committed file to a fresh classification of the live tree. A count hardcoded in a test is a failure scheduled for whenever the suite next grows — which is exactly how the first draft of this work broke.

Risk accepted, and why it is not a rerun of #962

#962 narrowed Windows coverage and was rescinded (#4421) after a macOS-only failure merged green. That failure was root-caused to a rendered-text-length assertion sensitive to tmpdir path length — a violation of ADR-456's typed-surface mandate that nothing enforced. Epic #4589 Phase 1 shipped that enforcement (local/no-rendered-text-length-assert, error from the moment it landed). The specific defect class that made the last narrowing unsafe is now statically prevented, which is the condition #4589 itself named as the precondition for narrowing. That is the difference, and it is why this narrowing rests on an enforced invariant rather than on optimism.

Rejected alternatives

  • Gate the lane instead of deleting it. Rejected: provably redundant, shown above. Gating would have produced a lane whose main arm duplicates the conformance lane exactly and whose residual arm is a filename heuristic.
  • Keep the lane ungated as deliberate belt-and-braces. Rejected: it does not function as a safety net for the files it duplicates, and for the files it does not duplicate it selects them by filename substring. Paying three Windows runners per PR for that is not a trade-off, it is an accident preserved.
  • Narrow hardcoded-path-vs-path-call to same-line proximity rather than deleting it. Rejected: ADR-1703's Linux-runnable rules already enforce the class, so real-OS execution buys nothing.
  • Also drop symlink-keyword (measured at the time as 228 rather than 254, before the shell-interpreter-spawn replacement took the tier to its final 266). Rejected: worth 6 unique files, and ADR-4593 reuses it in MACOS_CATEGORIES with recorded rationale.
  • Narrow chmod-mode-bit's bare-octal arm. #4641's text named this as a co-driver. Measurement says otherwise: 51 files match only via the bare-octal arm, but for 4 is chmod-mode-bit the sole signal. Changing it would invalidate ADR-4593's measured macOS table for a 4-file benefit. Rejected on evidence; the issue's claim is corrected here.
  • An absolute file-count ceiling. Rejected: goes stale under suite growth and stops binding without anyone noticing — the same failure shape as Phase 2's "the list exists" criterion.
  • A companion "sole-signal concentration" ceiling — no single category may be the sole signal for more than N% of the tier. Proposed because the ratio ceiling has a real Goodhart weakness: a ratio can be satisfied by inflating the denominator, so adding OS-agnostic tests loosens it without narrowing the tier. Concentration looked like the harder-to-fake companion, since the original defect was precisely one detector carrying half the tier. Measured, and rejected on the numbers. Post-fix the peak sole-signal share is raw-child-process at ~53/266 = ~20%, against the two historic offenders at 21.6% (process-seam-subprocess) and 19.8% (hardcoded-path-vs-path-call). Any threshold above 20% would have missed the original defect; any threshold below it fails today on a category that is entirely legitimate — a test that spawns a real subprocess genuinely needs a real OS. Concentration cannot separate "a big honest category" from "a big dishonest one"; the discriminator is whether the signal is platform-meaningful, which is a judgement no threshold encodes. The ratio ceiling stands alone, with its denominator-inflation weakness disclosed rather than papered over by a second gate that does not actually bind.
  • Relaxing raw-child-process to drop its content.includes('child_process') precondition. Investigated and rejected on measurement. The narrowing appeared to unmask a false negative: tests/windows-robustness.test.cjs contains spawnSync( and execFileSync( yet does not match raw-child-process, which looked like the precondition being over-tight (its stated rationale — stopping a local identifier such as spawnResult from matching — is already served by the trailing paren in \bspawnSync\(). Relaxing it was measured to add 13 files, and reading the matching line in each showed all 13 are false positives: comment text, jsdoc prose describing a return shape, template-literal code fixtures fed to an ESLint rule under test, and the search-anchor string literals described above. There is no false negative. The precondition stays exactly as written, and this paragraph exists so the same apparent bug is not "fixed" next time.