Files
msd-core/docs/adr/4593-macos-conformance-tier-architecture.md
Tom Boucher 4d65c248e5 fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector

Tests only, committed ahead of the implementation so the RED run is real.

- tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a
  ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative
  cases proving seam calls and path-call-plus-slash-literal are not platform
  signals; positive pins that genuine platform content, seam-bypassing spawns,
  chmod and symlink still classify in; macOS signal set and generated list
  unchanged.
- tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest
  rows and test-conformance still has 3 windows + 1 macOS.
- tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a
  non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint
  test does, and resolveSelection rejects the retired windows scope.

Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): delete the second Windows selector and narrow the conformance tier

Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped
conformance tier on real Windows/macOS — was not met. Measured on PR #4640
(run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7
changed test files running on a real Windows runner twice.

Two selectors, only one in the epic's scope. The test job's three scope:windows
shards predate the epic (#494, sharded #3057) and gate on product_changed, not
full_matrix, so they fire on every product PR whatever Phase 3's classifier
decides. They are deleted; test-conformance becomes the sole Windows selector,
as it already was for macOS. Non-Linux jobs 7 -> 4.

Gating the lane instead was rejected as provably redundant: for a test file
reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file),
and that same predicate sets full_matrix, which turns test-conformance on. Every
file a gated lane would run is already covered in the same run. The lane's one
non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename
heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix
instead of feeding a parallel lane.

Two detectors matched the repo's own test idiom rather than any platform signal
and carried 226 of the tier's sole-signal membership against 41 for the other
eight: process-seam-subprocess (335 files, 118 unique) matches the
tests/helpers.cjs entry points nearly every CLI test uses, and going through the
seam is the opposite of a platform signal since shell-command-projection takes
platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs
only a path call anywhere plus a slash literal anywhere, and that class is
already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed.
Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured.

Adds the size gate Phase 2 never had, as a ratio against a live denominator so
it cannot stop binding as the suite grows.

292 files leave real-OS Windows execution. The drop-out set was audited: 14 have
a platform-suggestive filename and all 14 are static source-text analyses or
seam-mediated CLI tests. raw-child-process was investigated as a suspected false
negative and left unchanged -- relaxing it adds 13 files, all false positives.

macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated
macos-conformance-tier.generated.cjs is byte-identical at 196 files.

Fixes #4641
Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): register the new ADR path in the docs-guard exempt baseline

tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md
in a comment justifying the retired windows scope; lint-docs-guard-registration
tracks that reference set, so the baseline needs the new path. Verified the
exemption still holds: the path is prose, not a filesystem read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): make the escalation tier-backed and drop every hardcoded count

Three follow-ups from measuring the first pass rather than trusting it.

The windows-hint escalation now requires tier membership as well as the
filename hint. Setting full_matrix runs test-conformance, which runs only the
tier; escalating on a test that is NOT in the tier costs four jobs and still
never runs that test on Windows. Measured over the 16 RULES entries the
narrowed predicate fires on exactly the same rules today, so this is
correct-by-construction rather than a behavior change. The broader variant --
escalate on any tier member a rule pulls in, ignoring the hint -- was measured
at 14/16 rules and rejected as over-broad.

Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as
soon as the rebase pulled in one new test file from #4253, which is the whole
argument against them: the ceilings are ratios against a live denominator, the
committed lists are pinned by comparison against a fresh classification of the
live tree, and the three named probe files now assert on their SIGNAL rather
than on membership in a literal list -- asserting by filename is the exact
error this PR fixes in the classifier.

Regenerates both lists against the rebased tree. Same-tree figures are now
547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none
added; macOS is unchanged at 197 with a zero-line diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke

An isolated adversarial review found a real false negative. Removing the
blanket process-seam-subprocess detector also removed the only coverage for
tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook
spawns options.interpreter via real spawnSync, so
runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary
executing a shell script extracted from workflow markdown. The seam argument
holds for src/shell-command-projection.cts, which takes platform as an
injected parameter; it does NOT hold for the test helpers, which spawn real
binaries. Conflating the two is what made the blanket detector look purely
noisy -- it was 99% noise wrapping a real signal.

Adds a narrow shell-interpreter-spawn category keyed on a real interpreter
option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are
added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33%
ceiling. All 9 confirmed by reading the matching source line, zero comment or
fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured
and rejected -- each adds 9 files but misses the counterexample entirely.

Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its
scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD
strictly rather than carrying the Windows report-only carve-out. The carve-out
now lives solely on test-conformance, whose matrix does include windows.

Repairs seven pre-existing assertions the category removal invalidated,
preserving each case's purpose rather than deleting coverage, and converts the
last hardcoded tier bounds to live-derived ratios -- including the macOS
sanity range that was still a magic [100, 350].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): keep the confinement test on a real OS via a documented allowlist

A security review found tests/external-descriptor-confinement.test.cjs had
dropped out of the Windows tier. It must stay in, and no content signal can
express why: it exercises isPathConfined (src/external-descriptor-trust.cts),
which uses the AMBIENT path module -- path.resolve(root, target) and path.sep
-- with no injection. Its win32 semantics (drive letters, UNC, separator) are
only reachable by actually running on Windows, and it is a security-relevant
write-confinement gate. A content classifier cannot see 'this module reads the
ambient path module', so no regex belongs here.

Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows
tier only. A Map rather than a list so an entry without a reason is impossible
by construction, and tests assert every entry names a file that exists on disk
so a stale entry fails loudly instead of rotting. This is the centrally-
enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's
portability-vocab.cjs already models -- deliberately not a heuristic.

Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is
untouched and byte-identical: the win32 concern does not apply to a POSIX
runner, and a test asserts the allowlist does not leak into that tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): inject the path impl into isPathConfined and correct the ADR count

Two review findings, both fixed rather than dispositioned.

A security review found tests/external-descriptor-confinement.test.cjs had left
real-OS execution. The allowlist pinned it back, but that only restored
INCIDENTAL coverage: isPathConfined used the ambient path module, and its test
carried POSIX-only literals, so a win32 confinement escape was unverified on
every platform including Windows. isPathConfined now takes an optional third
parameter carrying the path implementation, defaulting to the ambient module.
Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change
is purely additive and every existing two-argument caller is byte-identical.

Tests now inject path.win32 and path.posix, covering a different drive letter,
a cross-drive absolute, backslash and forward-slash traversal, UNC, and the
startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators.
Proved load-bearing: dropping the + p.sep from the prefix check fails exactly
the two boundary cases and nothing else. Callers' suites 149/149.

The spec review caught an off-by-one: the ADR narrated a 264-file tier while the
committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264
-> 265 (28.5%).

Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated
classify() as zeroing windows_tests, a key this change removes -- kept as
historical narration but labelled as such.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#131): make the unwritable-HOME test actually test something

Found by sweeping for the root-bypass class after fixing commit-files-deletion.
This one is the silent variant, and it was broken twice over.

First, the condition: the test made a fake HOME unwritable with chmod 0o500.
The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed
writable and the hostile condition never existed. Replaced with a HOME whose
PARENT is a regular file, so every write under it fails ENOTDIR at the VFS
layer for every uid -- no permission check is involved at all.

Second, and more fundamental: the probe was npm --version, which on npm 11.19.0
performs zero filesystem I/O against HOME. Proven rather than assumed --
neutralizing runNpm()'s isolation turned the sibling test red while this one
stayed green, so its assertion could never detect the regression it guards, on
any uid, with or without the condition fix. npm config get cache was tried next
and proved vacuous the same way (it only string-resolves the path). The probe is
now npm cache verify, which really does mkdir _cacache under HOME.

Re-proved load-bearing after the change: with isolation neutralized the test now
fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was
restored and verified diff-clean; suite 13/13.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the net drop-out figure in ADR-4641

The Consequences section still said 292 files leave real-OS Windows execution.
That was the count before the narrow shell-interpreter-spawn replacement
restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary
categories rather than only raw-child-process, and clarifies that the 14-file
filename audit was against the 292 initially dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the rejected concentration ceiling and its measurement

Applying Goodhart's own question to the new ceiling -- how would you make this
metric look good without improving what it represents -- surfaces a real
weakness: a ratio can be satisfied by inflating the denominator, so adding
OS-agnostic tests loosens it without narrowing the tier.

The obvious companion gate was a sole-signal concentration ceiling, since the
original defect was one detector carrying half the tier. Measured and rejected:
peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the
historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the
original defect; any threshold below it fails on a legitimate category. The
discriminator is whether a signal is platform-meaningful, which no threshold
encodes. Weakness disclosed rather than covered by a gate that does not bind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4641): add the changeset fragment for the confinement-check change

changeset-lint failed on PR #4643: the PR touches user-facing paths and carried
no fragment. The earlier no-changeset call matched #4604's CI-only precedent and
was correct then; it was not revisited once the PR grew a src/ change, which is
my miss.

The fragment describes the real user-visible improvement: the external-descriptor
write-confinement check's Windows semantics are now verified deterministically
rather than only when the suite happened to run on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the tier count in TESTING-SUITES.md

Said the tier narrowed from 546 to 254. The final committed list is 265 of 931
eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored
9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in
the ADR, in a live reference page rather than a dated record, so it states the
current truth rather than carrying an amendment note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the measured aggregate from real CI job lists

Epic #4589's closeout asserted its reduction from a static count; #4641's
acceptance criterion asks for a figure read off a real run. Recorded here:
test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run
against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4.

Also states the caveat that a PR's total CHECK count is not a clean before/after
comparison, since many gates are path-scoped and this change touches a broader
path set -- the like-for-like figure is the test.yml job count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): compare job totals the same way on both sides

The measured-aggregate table put #4640's COMPLETED run total (21) against this
run's count at matrix-expansion time (15). Those are not the same measurement:
the completed total includes the post-test Coverage gate and baseline-publisher
jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux
jobs 7 -> 4, was correct and is unchanged.

Called out in the table rather than silently corrected -- comparing two
differently-derived numbers is exactly the error class this ADR is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record measured conformance wall-clock and date the stale counterfactual

Adds the per-job durations from both runs. The honest read is that this is a
correctness win more than a speed one: file count fell 52% but wall-clock only
9-29%, because what was removed were the cheap static tests and what remains is
concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects
a future narrowing to buy time proportional to file count.

The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on
the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about --
pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its
tier was UNCHANGED at 197 files, which fixes that as runner variance and is
noted as a caution against reading a single duration as signal.

Also dates the symlink-keyword counterfactual, which cited a 254-file tier from
before the replacement category and allowlist took it to its final 265.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero

next gained #4644 mid-flight, so every absolute count shifted. Re-measured on
the tree this actually ships against (932 eligible): 548 -> 257 by detector
removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed,
9 restored. macOS 198, unchanged by this PR.

The percentages did not move across three rebases (58.8% -> 28.5%), which is
the whole argument for expressing the ceilings as ratios rather than counts --
noted in the ADR since it is now evidence rather than assertion.

Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own
win32 test cases introduced the literal win32 into the pinned file, so it
classifies in on content via win32-darwin-literal. The entry stays and the
reason is written down, because the file's real-OS need is a property of the
code under test (isPathConfined reads the ambient path module), not of the
test's text -- the text that currently saves it is incidental and could be
refactored away silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 17:00:11 -04:00

11 KiB

ADR-4593: A macOS-specific conformance-tier classifier, separate from the Windows-oriented one

  • Status: Accepted
  • Date: 2026-09-10
  • Issue: #4593 — Phase 4 (final) of epic #4589
  • Twin of: ADR-1703 (Windows portability enforcement architecture) — this ADR applies ADR-1703's evidence-first, static-classifier discipline to a second, macOS-specific surface, rather than reusing ADR-1703's Windows-oriented signal set unmodified.

Amendment (2026-09-11, ADR-4641): the general / Windows-oriented tier this ADR compares against was 546 files when this ADR was written, and every figure below is accurate as of that date. ADR-4641 has since removed two over-broad detectors from CATEGORIES, taking it to 254. Nothing in this ADR's decision changes: MACOS_CATEGORIES is a separate array, chmod-mode-bit and symlink-keyword keep the definitions and rationale recorded here, and the macOS tier remains 196 files. The historical figures below are deliberately left unedited.

Context

test-conformance's macos-latest CI leg (#4591, epic #4589 Phase 2) runs only the platform-conformance-tier file list — a content classifier over tests/**/*.test.cjs that flags files needing real-OS coverage, replacing a full-suite replay. That list's signal set (scripts/gen-platform-conformance-tier.cjs's CATEGORIES) was built for Windows: windows-shell-token, windows-env-var, hardcoded-path-vs-path-call, etc. macos-latest today runs the identical 546-file list, which is Windows-oriented, not evidence-backed for macOS specifically.

#4593 was originally filed against test-full's macOS legs, a job Phase 5 (#4603) has since deleted; the issue was corrected before design work started to target test-conformance's macos-latest leg instead — the "shrink from full replay" half of the original ask was already done by Phase 2.

Naively narrowing the general tier to macOS by simply dropping its three Windows-specific categories (windows-shell-token, windows-env-var, hardcoded-path-vs-path-call) barely narrows anything: measured, 546 -> 424 files (78% retained), because most files match multiple signals simultaneously and only need one surviving signal to stay in the tier. That is nowhere near the narrow, evidence-backed surface the issue asks for.

Decision

Build a second, macOS-specific classifier (classifyMacosContent / MACOS_CATEGORIES) in the same generator file, matching the issue's own named surface (zsh dispatch, case-sensitivity, darwin-specific behavior) plus two categories reused verbatim from the general tier that are genuinely Unix-relevant rather than Windows-motivated:

Category Test Rationale Files matched (of 930 eligible)
darwin-literal /\bdarwin\b/ (darwin alone, not the general tier's win32-darwin-literal OR) The general tier's combined win32|darwin category can't distinguish which branch fired; a file mentioning only win32 says nothing about macOS. 8
zsh-dispatch /\bzsh\b/i Directly names the issue's own "zsh shell dispatch" surface. 14
case-sensitivity /case.?insensitiv|case.?sensitiv/i Directly names the issue's own "case-sensitivity" surface (macOS's default case-insensitive-but-preserving filesystem). 74
chmod-mode-bit reused verbatim from the general tier's CATEGORIES Unix permission-bit semantics — genuinely macOS/Linux-relevant, not a Windows category (Windows has no chmod). 73
symlink-keyword reused verbatim from the general tier's CATEGORIES Symlink handling differs materially on macOS (case-insensitive-but-preserving FS, different default symlink permissions) — not Windows-motivated the way windows-shell-token etc. are. 86

Applied to the real tests/ unit-suite tree (930 files eligible under the same suiteOf(f) === null gate the general tier uses), the union is 196 files (21% of eligible files, vs. the general tier's 546/930, 59%) — a materially narrower, macOS-evidence-backed tier. The per-category counts above include tests/platform-conformance-tier.test.cjs's own new fixture strings for zsh-dispatch/case-sensitivity (e.g. "shell: 'zsh {0}'") — self-referential by one file each, since this test file lives in the same tree it classifies. Caught by an isolated code-review pass comparing this table against a fresh live count; the union total was unaffected (that file was already in the tier via an unrelated signal).

Rejected: a standalone CRLF/autocrlf signal

The issue's own surface description names "CRLF-checkout behavior" as part of macOS's risk. This was measured and rejected as a standalone macOS signal:

  • The obvious first attempt, a literal /\r\n/ match, hit 197/952 files (21% of all test files) — far too broad, because it matches routine defensive newline-normalization code (.replace(/\r\n/g, '\n')) present throughout the suite for unrelated reasons, not files specifically at CRLF-checkout risk.
  • Narrowing to the actual named terms, /\bCRLF\b|autocrlf/i, still hit 143 files — still too broad to be a precise macOS-only signal.
  • Root cause: in this codebase, CRLF is primarily a Windows checkout concern — ADR-1703 files its no-crlf-fragile-split rule under DEFECT.WINDOWS-TEST-PORTABILITY, not a macOS defect class. A CRLF-keyword signal was therefore pulling in files already covered by the general, Windows-oriented tier, not narrowing macOS coverage specifically.

The issue's "CRLF-checkout" framing is treated as "this is one of the risk categories macOS needs coverage for" — already satisfied, for files that actually need it, by the general tier — not "this is a macOS-exclusive signal to build."

Implementation

  • scripts/gen-platform-conformance-tier.cjs gains MACOS_CATEGORIES, classifyMacosContent, classifyMacosTree, and renderMacosGeneratedFile, mirroring the general tier's CATEGORIES/classifyContent/classifyTree/renderGeneratedFile exactly (same suiteOf(f) === null eligibility gate — see ADR-1703's sibling doc-comment in that file for why suite-tagged files are excluded entirely rather than merely deprioritized). A new --target windows (default, unchanged) / --target macos CLI flag selects which of the two independent generated outputs a given --check/--write/print invocation targets, so package.json's lint:generated-sync/regen:derived chains invoke the same script twice rather than needing a second script file.
  • New committed output scripts/lib/macos-conformance-tier.generated.cjs, exporting MACOS_CONFORMANCE_TIER_FILES — same generated-file convention (one array entry per line, sorted, GENERATED FILE header banner) as platform-conformance-tier.generated.cjs. Gated by the same lint:generated-sync --check fail-safe as the general tier — no new fail-safe mechanism needed; this is a second instance of an already-proven pattern.
  • .github/workflows/test.yml's test-conformance job: the macos-latest leg's "prepare conformance-tier test list" step reads MACOS_CONFORMANCE_TIER_FILES instead of the shared CONFORMANCE_TIER_FILES, writing into the same conventional .ci-conformance-tests.txt filename so the downstream run-tests.cjs --files-from step is unchanged. The windows-latest legs, their shard count, and everything else about the job are untouched — they keep using the general, Windows-inclusive tier.

Explicit requirement for future widening proposals

Per issue #4593's "Done when": before any future proposal to widen macOS coverage beyond this 196-file tier, check whether the motivating regression class is already covered by Phase 1's no-rendered-text-length-assert ESLint rule (#4590, cataloged in ADR-1703) — an author-time, zero-escape-hatch rule that catches exactly the shape behind #4421 (a rendered-text-length assertion embedding an OS-derived path, e.g. os.tmpdir(), whose length differs between macOS's /private/var/folders/… prefix and Linux's shorter one). #4421's root cause was that static-assertion shape, not a real macOS behavioral divergence requiring dynamic real-OS execution to catch. A future explorer proposing full macOS/Linux parity coverage should confirm new evidence of a behavioral divergence this static classifier and no-rendered-text-length-assert both miss, rather than re-proposing parity on the strength of #4421 alone — #4421 is already covered.

Consequences

Positive: macos-latest now runs a narrower, macOS-evidence-backed 196-file tier (down from the shared 546-file Windows-oriented list) — a real reduction (~64%) in real-OS macOS CI minutes for files with no macOS-specific signal, without dropping coverage for files that do carry one. Same generator, same fail-safe convention, same install/build ripple discipline as the general tier — no new mechanism class.

Cost / risk: a second static classifier is still subject to the same disclosed limit as the general tier (ADR-1703 / gen-platform-conformance-tier.cjs's own header doc-comment): it is content-based, not a real per-file behavioral diff. Mitigated the same way — the most recent full-matrix next run was green on every file in this classification before it was built. The legacy full-matrix macOS job's actual safety-net window was much shorter than originally planned: Phase 2 (#4591) added it as a non-gating safety net intended for "one release cycle," but Phase 5 (#4603) retired it roughly 4 hours later, the same day, after discovering it was purely additive to the new conformance job (10 OS-specific CI jobs per PR instead of the intended reduction) rather than a genuine transition period — see #4603 for the full accounting. macos-latest's coverage here has not yet had an extended real-world safety-net window of its own; that is an accepted, disclosed risk of shipping Phase 4 promptly rather than waiting, consistent with this epic's general preference for fast, evidence-driven iteration over a long unmonitored parallel-running period.

Alternatives considered

  1. Apply the general tier's signals minus the three Windows-specific categories. Rejected — measured at 424/546 files (78% retained), nowhere near the narrow, macOS-evidence-backed surface the issue asks for; see "Context" above.
  2. A standalone CRLF/autocrlf signal. Rejected — see "Rejected: a standalone CRLF/autocrlf signal" above; both attempted regexes (197 and 143 files) were too broad and were really re-selecting Windows-relevant files already covered by the general tier.
  3. A single, unified classifier covering both Windows and macOS signals, with per-OS filtering at CI-invocation time. Rejected — would still need two independent signal sets internally to produce two independently-narrow lists, so it does not actually simplify anything over two sibling classifier functions in the same file; a single combined list reproduces the 78%-retained problem from alternative 1.