Commit Graph

30 Commits

Author SHA1 Message Date
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
Tom Boucher
4d65c248e5 fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector

Tests only, committed ahead of the implementation so the RED run is real.

- tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a
  ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative
  cases proving seam calls and path-call-plus-slash-literal are not platform
  signals; positive pins that genuine platform content, seam-bypassing spawns,
  chmod and symlink still classify in; macOS signal set and generated list
  unchanged.
- tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest
  rows and test-conformance still has 3 windows + 1 macOS.
- tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a
  non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint
  test does, and resolveSelection rejects the retired windows scope.

Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): delete the second Windows selector and narrow the conformance tier

Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped
conformance tier on real Windows/macOS — was not met. Measured on PR #4640
(run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7
changed test files running on a real Windows runner twice.

Two selectors, only one in the epic's scope. The test job's three scope:windows
shards predate the epic (#494, sharded #3057) and gate on product_changed, not
full_matrix, so they fire on every product PR whatever Phase 3's classifier
decides. They are deleted; test-conformance becomes the sole Windows selector,
as it already was for macOS. Non-Linux jobs 7 -> 4.

Gating the lane instead was rejected as provably redundant: for a test file
reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file),
and that same predicate sets full_matrix, which turns test-conformance on. Every
file a gated lane would run is already covered in the same run. The lane's one
non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename
heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix
instead of feeding a parallel lane.

Two detectors matched the repo's own test idiom rather than any platform signal
and carried 226 of the tier's sole-signal membership against 41 for the other
eight: process-seam-subprocess (335 files, 118 unique) matches the
tests/helpers.cjs entry points nearly every CLI test uses, and going through the
seam is the opposite of a platform signal since shell-command-projection takes
platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs
only a path call anywhere plus a slash literal anywhere, and that class is
already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed.
Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured.

Adds the size gate Phase 2 never had, as a ratio against a live denominator so
it cannot stop binding as the suite grows.

292 files leave real-OS Windows execution. The drop-out set was audited: 14 have
a platform-suggestive filename and all 14 are static source-text analyses or
seam-mediated CLI tests. raw-child-process was investigated as a suspected false
negative and left unchanged -- relaxing it adds 13 files, all false positives.

macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated
macos-conformance-tier.generated.cjs is byte-identical at 196 files.

Fixes #4641
Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): register the new ADR path in the docs-guard exempt baseline

tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md
in a comment justifying the retired windows scope; lint-docs-guard-registration
tracks that reference set, so the baseline needs the new path. Verified the
exemption still holds: the path is prose, not a filesystem read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): make the escalation tier-backed and drop every hardcoded count

Three follow-ups from measuring the first pass rather than trusting it.

The windows-hint escalation now requires tier membership as well as the
filename hint. Setting full_matrix runs test-conformance, which runs only the
tier; escalating on a test that is NOT in the tier costs four jobs and still
never runs that test on Windows. Measured over the 16 RULES entries the
narrowed predicate fires on exactly the same rules today, so this is
correct-by-construction rather than a behavior change. The broader variant --
escalate on any tier member a rule pulls in, ignoring the hint -- was measured
at 14/16 rules and rejected as over-broad.

Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as
soon as the rebase pulled in one new test file from #4253, which is the whole
argument against them: the ceilings are ratios against a live denominator, the
committed lists are pinned by comparison against a fresh classification of the
live tree, and the three named probe files now assert on their SIGNAL rather
than on membership in a literal list -- asserting by filename is the exact
error this PR fixes in the classifier.

Regenerates both lists against the rebased tree. Same-tree figures are now
547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none
added; macOS is unchanged at 197 with a zero-line diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke

An isolated adversarial review found a real false negative. Removing the
blanket process-seam-subprocess detector also removed the only coverage for
tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook
spawns options.interpreter via real spawnSync, so
runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary
executing a shell script extracted from workflow markdown. The seam argument
holds for src/shell-command-projection.cts, which takes platform as an
injected parameter; it does NOT hold for the test helpers, which spawn real
binaries. Conflating the two is what made the blanket detector look purely
noisy -- it was 99% noise wrapping a real signal.

Adds a narrow shell-interpreter-spawn category keyed on a real interpreter
option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are
added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33%
ceiling. All 9 confirmed by reading the matching source line, zero comment or
fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured
and rejected -- each adds 9 files but misses the counterexample entirely.

Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its
scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD
strictly rather than carrying the Windows report-only carve-out. The carve-out
now lives solely on test-conformance, whose matrix does include windows.

Repairs seven pre-existing assertions the category removal invalidated,
preserving each case's purpose rather than deleting coverage, and converts the
last hardcoded tier bounds to live-derived ratios -- including the macOS
sanity range that was still a magic [100, 350].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): keep the confinement test on a real OS via a documented allowlist

A security review found tests/external-descriptor-confinement.test.cjs had
dropped out of the Windows tier. It must stay in, and no content signal can
express why: it exercises isPathConfined (src/external-descriptor-trust.cts),
which uses the AMBIENT path module -- path.resolve(root, target) and path.sep
-- with no injection. Its win32 semantics (drive letters, UNC, separator) are
only reachable by actually running on Windows, and it is a security-relevant
write-confinement gate. A content classifier cannot see 'this module reads the
ambient path module', so no regex belongs here.

Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows
tier only. A Map rather than a list so an entry without a reason is impossible
by construction, and tests assert every entry names a file that exists on disk
so a stale entry fails loudly instead of rotting. This is the centrally-
enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's
portability-vocab.cjs already models -- deliberately not a heuristic.

Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is
untouched and byte-identical: the win32 concern does not apply to a POSIX
runner, and a test asserts the allowlist does not leak into that tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): inject the path impl into isPathConfined and correct the ADR count

Two review findings, both fixed rather than dispositioned.

A security review found tests/external-descriptor-confinement.test.cjs had left
real-OS execution. The allowlist pinned it back, but that only restored
INCIDENTAL coverage: isPathConfined used the ambient path module, and its test
carried POSIX-only literals, so a win32 confinement escape was unverified on
every platform including Windows. isPathConfined now takes an optional third
parameter carrying the path implementation, defaulting to the ambient module.
Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change
is purely additive and every existing two-argument caller is byte-identical.

Tests now inject path.win32 and path.posix, covering a different drive letter,
a cross-drive absolute, backslash and forward-slash traversal, UNC, and the
startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators.
Proved load-bearing: dropping the + p.sep from the prefix check fails exactly
the two boundary cases and nothing else. Callers' suites 149/149.

The spec review caught an off-by-one: the ADR narrated a 264-file tier while the
committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264
-> 265 (28.5%).

Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated
classify() as zeroing windows_tests, a key this change removes -- kept as
historical narration but labelled as such.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#131): make the unwritable-HOME test actually test something

Found by sweeping for the root-bypass class after fixing commit-files-deletion.
This one is the silent variant, and it was broken twice over.

First, the condition: the test made a fake HOME unwritable with chmod 0o500.
The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed
writable and the hostile condition never existed. Replaced with a HOME whose
PARENT is a regular file, so every write under it fails ENOTDIR at the VFS
layer for every uid -- no permission check is involved at all.

Second, and more fundamental: the probe was npm --version, which on npm 11.19.0
performs zero filesystem I/O against HOME. Proven rather than assumed --
neutralizing runNpm()'s isolation turned the sibling test red while this one
stayed green, so its assertion could never detect the regression it guards, on
any uid, with or without the condition fix. npm config get cache was tried next
and proved vacuous the same way (it only string-resolves the path). The probe is
now npm cache verify, which really does mkdir _cacache under HOME.

Re-proved load-bearing after the change: with isolation neutralized the test now
fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was
restored and verified diff-clean; suite 13/13.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the net drop-out figure in ADR-4641

The Consequences section still said 292 files leave real-OS Windows execution.
That was the count before the narrow shell-interpreter-spawn replacement
restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary
categories rather than only raw-child-process, and clarifies that the 14-file
filename audit was against the 292 initially dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the rejected concentration ceiling and its measurement

Applying Goodhart's own question to the new ceiling -- how would you make this
metric look good without improving what it represents -- surfaces a real
weakness: a ratio can be satisfied by inflating the denominator, so adding
OS-agnostic tests loosens it without narrowing the tier.

The obvious companion gate was a sole-signal concentration ceiling, since the
original defect was one detector carrying half the tier. Measured and rejected:
peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the
historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the
original defect; any threshold below it fails on a legitimate category. The
discriminator is whether a signal is platform-meaningful, which no threshold
encodes. Weakness disclosed rather than covered by a gate that does not bind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4641): add the changeset fragment for the confinement-check change

changeset-lint failed on PR #4643: the PR touches user-facing paths and carried
no fragment. The earlier no-changeset call matched #4604's CI-only precedent and
was correct then; it was not revisited once the PR grew a src/ change, which is
my miss.

The fragment describes the real user-visible improvement: the external-descriptor
write-confinement check's Windows semantics are now verified deterministically
rather than only when the suite happened to run on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the tier count in TESTING-SUITES.md

Said the tier narrowed from 546 to 254. The final committed list is 265 of 931
eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored
9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in
the ADR, in a live reference page rather than a dated record, so it states the
current truth rather than carrying an amendment note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the measured aggregate from real CI job lists

Epic #4589's closeout asserted its reduction from a static count; #4641's
acceptance criterion asks for a figure read off a real run. Recorded here:
test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run
against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4.

Also states the caveat that a PR's total CHECK count is not a clean before/after
comparison, since many gates are path-scoped and this change touches a broader
path set -- the like-for-like figure is the test.yml job count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): compare job totals the same way on both sides

The measured-aggregate table put #4640's COMPLETED run total (21) against this
run's count at matrix-expansion time (15). Those are not the same measurement:
the completed total includes the post-test Coverage gate and baseline-publisher
jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux
jobs 7 -> 4, was correct and is unchanged.

Called out in the table rather than silently corrected -- comparing two
differently-derived numbers is exactly the error class this ADR is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record measured conformance wall-clock and date the stale counterfactual

Adds the per-job durations from both runs. The honest read is that this is a
correctness win more than a speed one: file count fell 52% but wall-clock only
9-29%, because what was removed were the cheap static tests and what remains is
concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects
a future narrowing to buy time proportional to file count.

The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on
the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about --
pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its
tier was UNCHANGED at 197 files, which fixes that as runner variance and is
noted as a caution against reading a single duration as signal.

Also dates the symlink-keyword counterfactual, which cited a 254-file tier from
before the replacement category and allowlist took it to its final 265.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero

next gained #4644 mid-flight, so every absolute count shifted. Re-measured on
the tree this actually ships against (932 eligible): 548 -> 257 by detector
removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed,
9 restored. macOS 198, unchanged by this PR.

The percentages did not move across three rebases (58.8% -> 28.5%), which is
the whole argument for expressing the ceilings as ratios rather than counts --
noted in the ADR since it is now evidence rather than assertion.

Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own
win32 test cases introduced the literal win32 into the pinned file, so it
classifies in on content via win32-darwin-literal. The entry stays and the
reason is written down, because the file's real-OS need is a property of the
code under test (isPathConfined reads the ambient path module), not of the
test's text -- the text that currently saves it is incidental and could be
refactored away silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 17:00:11 -04:00
Tom Boucher
0b928fe28c feat(#4592): replace the blanket test-file full_matrix rule with reachability (#4602)
scripts/ci-test-scope.cjs's classify() previously set full_matrix=true for
ANY changed tests/**/*.test.cjs file, unconditionally (restored by #4421
after #962's narrowing let a real macOS-only regression, PR #4384, land
undetected). This replaces that blanket rule with a reachability check
against real data instead of a path prefix:

- A changed test file forces full_matrix only when it is present in Phase
  2's committed CONFORMANCE_TIER_FILES list (scripts/lib/platform-
  conformance-tier.generated.cjs) -- direct membership, not a graph walk.
- A changed src/ file forces full_matrix when its own content carries a
  genuine platform-conditional signal, reusing gen-platform-conformance-
  tier.cjs's classifyContent with a narrowed, source-code-safe signal
  subset (excludes two categories -- hardcoded-path-vs-path-call and
  symlink-keyword -- empirically found to flag 100/235 src/ files when
  applied verbatim, versus 28/235 with the narrow subset, all verified to
  carry genuine platform branches). New export: NOISY_FOR_SOURCE_REACHABILITY.
- A change to the classification mechanism's own definition files
  (gen-platform-conformance-tier.cjs, the generated tier list, or
  suite-detection.cjs) always forces full_matrix -- the mechanism being
  changed cannot presume its own new output is safe.
- Any computation error (a require/read failure, a malformed module) fails
  safe to full_matrix=true, per the issue's explicit requirement.

The existing RULES array entries with their own fullMatrix:true (workflow
automation, installer/package layout, hooks, environment/dependency gates,
test harness) are deliberately left untouched -- they are curated,
narrowly-scoped triggers for "this diff changes the CI/installer/hooks
mechanism itself," a different and still-valid reason than "product code
might reach a platform branch." Disclosed in .gsd/phase/.../40-design.md
as a scope decision, since the issue's "Done when" wording read broader
than its "Proposed work" bullets.

Two design assumptions were caught and corrected before any code was
written (rubber-duck pass, documented in 40-design.md): (1) reusing Phase
2's classifyContent verbatim against src/ was far too noisy; (2) a single
hardcoded seam file (src/shell-command-projection.cts only, per CLAUDE.md's
"single platform seam" framing) would have silently missed genuine,
independent platform branches in src/runtime-hooks-surface.cts,
src/capability-lock.cts, src/capability-ledger.cts, and src/surface.cts --
reintroducing the #4421 failure shape inside src/ instead of tests/.

An isolated code-review pass found and fixed one real defect (a dead,
untested branch that would have survived Stryker mutation testing) and one
design-doc completeness gap (2 of 8 "narrow" signal categories were left
implicitly rather than explicitly audited). An isolated security-review
pass found no qualifying findings.

tests/ci-test-scope.test.cjs gains the full #4592 boundary-case matrix
(.gsd/phase/.../50-test-matrix.md), including a named #4421 regression case
proving tests/state-todos-render.test.cjs still forces full_matrix, now for
the documented reason instead of the removed blanket rule. Two pre-existing
tests were corrected: one used a nonexistent fixture path (src/semver.cts
-> src/semver-compare.cts, a real file); one (A3) asserted the exact old
blanket-rule behavior this issue removes, updated to the new, verified-
correct expectation.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:21 -04:00
Tom Boucher
181c4c8659 chore(#4603): retire the test-full CI job (#4604)
* chore(#4603): retire the test-full CI job

Phase 2 (#4591) added test-conformance but left test-full (the pre-existing
full-suite Windows/macOS replay) running unchanged, gated on the same
full_matrix flag, downgraded only from a hard gate to a non-blocking
::warning:: -- framed as "a non-gating safety net for one release cycle."
No phase or issue ever retired it. Result: every full_matrix=true PR ran
10 OS-specific jobs (test-full's 6 + test-conformance's 4, purely
additive) instead of the original 6 -- the epic's own goal (reduce
runner-minutes) was measurably regressing, not improving, for the
majority of PRs.

This phase was missing from the original 4-phase epic decomposition; the
epic (#4589) has been amended to add it as Phase 5 (see its comment
thread), and this issue was filed as the tracked sub-issue.

Deletes the test-full job from .github/workflows/test.yml entirely, along
with every reference to it: required-tests' needs/FULL_TEST_RESULT
warning branch, ci-timeout-report.cjs's JOB_RULES entry,
ci-test-job-timeout-budget.test.cjs's LANE_COSTS/staticLanes/testFullRule
entries, ci-test-scope.test.cjs's test-full-specific tests (preserving
three unrelated tests that were nested in the same describe block, moved
under a renamed describe rather than deleted), and docs mentions.
test-conformance is now the sole gating signal for real-OS coverage.

Two separate defects found and fixed while auditing every test-full
reference:
- tests/ci-pr-mergeability.test.cjs's GATED['test.yml'] safety-critical
  array (jobs that must needs: the mergeability preflight) had test-full
  but was missing test-conformance entirely -- Phase 2 never added it.
  Verified the real workflow wiring was already correct (test-conformance
  does have needs: [changes, preflight]); this was a test-coverage gap,
  not a live defect. Fixed by swapping the array entry.
- docs/TESTING-SUITES.md's "## CI matrix" section was substantially stale
  independent of this phase (predating even #2952's coverage-gate split).
  Rewritten against the real, current job topology, verified directly
  against test.yml rather than trusted from memory.

An isolated code-review pass found and fixed two minor inaccuracies in the
rewritten docs table (two jobs' "Gated on" column didn't match their real
if: condition exactly). An isolated security-review pass found no
qualifying findings -- every compute-provisioning job already carries
needs: preflight directly, unaffected by this deletion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(ci): isolate 7 more heavy test files from chunk-weight packing

`next`'s own push-triggered Tests run failed: `conformance test
(windows-latest, 24, shard 2/3)` chunk 3/6 was killed after 600019ms.
Root cause: state.test.cjs (weight 21.35, measured) was packed alongside
companions by run-tests.cjs's LPT chunk packer, the same failure mode
that previously hit codex-config.test.cjs (weight 17.87) twice and got a
dedicated fix (ISOLATED_HEAVY_FILES, #4497) -- but state.test.cjs was
never added to that set.

This is a direct, unintended consequence of epic #4589 Phase 2: the new
platform-conformance-tier job packs only ~546 files per shard (vs. the
~950-file full suite the packer used to balance against), so the same
absolute-weight outlier now represents a larger share of a smaller, more
homogeneous pool -- the LPT packer has fewer light files to pad around
it with. This was a real, foreseeable side effect of shrinking the
packing pool that nobody checked for when Phase 2 shipped.

A first attempt at this fix hand-picked 4 candidates by eyeballing a
truncated weight list and missed 3 heavier ones -- caught by an isolated
code-review pass (blocker: emitted-attribution.test.cjs at 66.2% of the
Windows chunk budget, install-minimal-hooks.test.cjs at 61.1%,
install.test.cjs at 47.1%, all above codex-config.test.cjs's own
44.7% -- the ratio that already proved dangerous twice). Corrected by
systematically computing weight/budget for every unit-suite file and
isolating everything at or above that same ratio: 7 files total, plus
the pre-existing codex-config.test.cjs (8 total).

Added a durable regression test (tests/run-tests-harness.test.cjs) that
re-derives this exact computation from the live tests/test-timings.json
on every run, so a future heavy file crossing this threshold fails the
test instead of silently reintroducing this failure -- not just a
one-time manual sweep.

Verified end-to-end: simulated the real 3-way windows shard split of the
actual conformance-tier file list with the real packing functions. Max
packable-chunk weight across all 3 shards is now 27.04 / 24.10 / 23.91
(shard 2 is the exact shard that failed on next), comfortably under the
40 budget -- versus 40+ and a 600s kill before this fix.

A second isolated code-review + security-review pass on the corrected
diff found nothing further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:04 -04:00
Tom Boucher
2920bbc022 fix(#4421): rescind #494's macOS full-matrix skip on changed test files (#4427) 2026-09-06 15:42:23 -04:00
Tom Boucher
02c6955162 chore(#4241): add merge_group trigger to test.yml (#4242)
* chore(ci): add merge_group trigger to test.yml (#4241)

GitHub's merge queue fires the `merge_group` event for the temporary
merge-group commit it creates when a PR is added to the queue - not
`pull_request` or `push`. Without this trigger, `required-tests` (the
registered "Required tests" branch-protection check) never schedules
for a queued PR, permanently stalling the queue on a check that never
runs. This is workflow-side prerequisite wiring only; enabling the
merge queue itself is a separate manual branch-protection step.

* fix(#4241): pin AUDIT_BASELINE_REF for merge_group events too

Code review on this branch caught that AUDIT_BASELINE_REF's ternary
only branched on pull_request/push, so a merge_group run silently fell
through to '' -- scripts/npm-audit-baseline.cjs's resolveBaselineRef()
documents its origin/next live-tip fallback as unreachable from CI
specifically because AUDIT_BASELINE_REF is "always set by test.yml".
Reopens the exact race #4196 fixed, but only for merge-queue runs.

Extends all three AUDIT_BASELINE_REF pins (test, test-inert, test-full)
to also branch on merge_group, using github.event.merge_group.base_sha
(confirmed against GitHub's own webhook payload schema: "the SHA of the
merge group's parent commit") -- the base tip the temporary
merge-group commit was built against.

Adds a regression test asserting every AUDIT_BASELINE_REF pin branches
on merge_group with the correct field.

---------

Co-authored-by: sim <sim@local>
2026-09-03 15:00:04 -04:00
Tom Boucher
107eb8c1d9 feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on dacae9273 while the PR that caused it (#3746) was green on every
check.

The docs-lint job in .github/workflows/docs-required.yml -- an ALREADY-REQUIRED
context -- now selects and runs the docs guards that read the specific docs files
the PR changed.

  scripts/docs-guard-registry.cjs    test file -> the docs paths it reads (63)
  scripts/select-docs-guards.cjs     pure (changedPaths, registry) -> test files
  scripts/lint-docs-guard-registration.cjs   drift guard, wired into lint:ci

scripts/ci-test-scope.cjs is NOT touched -- `git diff origin/next --` on it is
empty -- so #764's saving stands and its 21 pinning tests are untouched.

Selection: exact path; trailing-slash directory prefix (boundary-checked --
docs/adrenaline.md does NOT match docs/adr/, which a naive startsWith gets
wrong); and '*' for the 6 entries that walk docs/ generally or read a computed
path. Unknown maps to '*' -- guessing narrow is how a guard silently stops
running. Measured: a typo fix selects 6 of 63; docs/AGENTS.md selects 12;
docs/COMMANDS.md selects 18.

Four things this got wrong first, each found by an independent reviewer or by
probe, and each having been asserted safe in a comment:

1. The registry started as a RULE in ci-test-scope.cjs's RULES, on the theory
   that classify()'s !codeChanged normalization made it inert. True for
   docs-ONLY diffs; false for MIXED docs+code diffs, where codeChanged is true
   and the normalization never runs:

     node scripts/ci-test-scope.cjs --files "docs/a.md src/semver.cts"
       with the RULE:  25 targeted_tests
       origin/next:     3 targeted_tests

   Category error: RULES is the scoped lane's input; a docs-guard registry is a
   lane manifest for a consumer that never calls classify(). Extracted; pinned
   by value.

2. The second attempt was a dedicated workflow with paths: [docs/**]. Such a
   workflow never reports on a non-docs PR, so it can never be a required
   context without hanging every non-docs PR -- and a non-required check does not
   block a merge, so the guard would have been advisory and #3753 unfixed.
   docs-required.yml already has no paths: filter, already supplies the required
   docs-lint context, already computes docs_changed, and already ran one docs
   guard gated on it. Generalizing that step needs no ruleset edit at all.

3. The registry and the drift lint were built from ONE path-segment heuristic, so
   both were blind identically -- and blind at the guard that motivated the issue.
   The reader-call regex required a character BEFORE its keyword, so a callee
   named exactly read( / load( / parse( / doc( / file( / content( could never
   match; and only an INLINE path.join(ROOT,'docs','X.md') argument was caught,
   missing the two-step-via-variable form -- the MAJORITY spelling -- plus
   template literals and concatenation. Detector 1 fired on 14 of ~450 files, so
   35 genuine guards sat unregistered while the lint reported 0 violations,
   including cursor-reviewer (reads docs/COMMANDS.md, asserts
   .includes('--cursor')) and inventory-headings-countfree. The "accepted blind
   spot" this shipped with was the common case, not a fringe.

4. With detection fixed the true population is 115 files: 63 genuine guards, 52
   incidental. Running all 63 in a REQUIRED check on a one-line typo fix is the
   cost #764 exists to avoid -- install.test.cjs is 7840 lines and reads exactly
   one docs file, docs/AGENTS.md, for its frontmatter. Dropping it reproduces the
   bug; running it for a typo elsewhere is waste. Hence the map.

Then a second review round found six more, all fixed here:

- fragment-single-edit-propagation.install.test.cjs was EXEMPTED as
  "overlay fixture only". False: it reads the real docs/registries/eos.json and
  asserts on a registry entry name, and reads the real ADR-0001 and asserts its
  H1. A docs-only PR touching either would have gone green and red next -- #3753
  shipping again, from inside the fix for it. Now registered against both paths,
  and all 52 remaining exemptions were re-audited one by one.
- The SUITES-collision guard compared RAW registry keys, but run-tests.cjs strips
  a leading `tests/` BEFORE its suite check. So it caught 'all' and missed
  'tests/all' -- the only spelling that can actually occur, since every key
  carries the prefix. One typo would have run all 824 test files inside the
  required job. Now normalized the same way run-tests.cjs normalizes.
- The lint failed OPEN on an unreadable tests dir or candidate file: 0 violations,
  ok:true. A guard that cannot read its input must never report success.
- The exemption ratchet gated identity only, so a baselined file that later
  STARTED asserting on shipped docs stayed exempt silently -- 52 permanently blind
  files. The baseline now fingerprints the docs paths each exempted file
  references and fails when that set changes, naming what changed.
- The exemption marker was still honored inside a multi-line template literal in
  the header window. The scanner now tracks template-literal and block-comment
  state.
- `git diff --name-only | grep '^docs/'` silently dropped C-quoted non-ASCII docs
  paths, making docs_changed=false a green zero-guard check. Both call sites now
  pass -c core.quotepath=false.
- The run step was gated on hashFiles(), which a force-committed
  .docs-guard-tests.txt would satisfy. The step now rm -f's both scratch files
  first and gates on an output it sets itself.

Three empty states, deliberately distinct, because conflating them rebuilds
#3753: an empty or malformed registry HARD-FAILS; docs changed with no guard
covering them logs and skips; no docs change is already gated. The middle state
must never be expressed as an empty --files-from, which prints `no tests in suite
"all"` and exits 0 -- a green check that guarded nothing. With the current
registry that state is unreachable, because the six '*' entries always match;
the branch is kept as defensive handling for a future registry and says so.

timeout-minutes: 15 bounds the required job against a hanging fork-supplied test;
it had none. npm ci was added because the job never installed dependencies -- the
previous single-file step got away without it, the registry does not.

docs/contributing/docs-guard-registration.md documents the rule, following its
sibling cross-platform-portability-rules.md, and CONTRIBUTING.md's CI Test
Quality Checks table links to it. It is also load-bearing: without a docs/ file
in the diff this PR would not have triggered its own lane, shipping an
unexercised change to a required check.

One unrelated fix, included because this PR surfaced it and CLAUDE.md forbids
deferring a defect found while working. On this branch's first CI run,
`full test (windows-latest, 24, shard 3/3)` was CANCELLED at exactly 30 minutes;
tests were still passing 0.8s before the cancel, so it is a wall-clock timeout,
not a hang, and a cancelled job reddens `Required tests`.

The cause is not this PR's test file, which costs ~60ms. Shard composition is
unstable: adding ONE file to the unit suite reshuffled 115 of 268 files between
shards, and shard 3 drew a heavier mix. Underneath that is a real pre-existing
defect. tests/ci-test-job-timeout-budget.test.cjs requires every lane's budget to
be >= 1.5x its MEASURED cost -- "a lane that got slower must be re-budgeted, not
excused" -- and its test-full entry recorded 19m from a windows-22 shard. That is
stale. Measured on `next` with none of this PR's changes present: 26m18s (run
32614439702, windows-latest/24 shard 3/3), 23m36s and 23m17s on shard 2/3. So the
lane costs ~26m and the 30-minute cap carried 1.14x headroom, not 1.5x. The gate
had been out of compliance with its own rule; this PR was merely the file
addition that reshuffled shard 3 past the cliff.

Fixed as that file prescribes: measuredMinutes 19 -> 27 with fresh evidence, and
test-full timeout-minutes 30 -> 45. The rule's minimum for 27m is 41; 45 is
deliberately above it because the reshuffle means per-shard worst case moves run
to run, and a budget pinned to the exact minimum would be re-breached by the next
test file anyone adds. Only that one job's timeout changed; test.yml's scope,
matrix and steps are untouched, so #764's saving is unaffected.

Raising that cap let the Windows shard finish (28m45s, inside 45) and uncovered
a real failure the 30-minute cancel had been masking:
`new quick-task branch branches off origin/main (#2916)` died with
`outcome=timed_out exitCode=null`, SIGTERM, at the 15000ms bound.

tests/quick-branching.test.cjs:149 `runStep` runs a `#!/usr/bin/env bash` script
executing MULTIPLE git commands, but was bound to GIT_TIMEOUT_MS (15000) -- the
norm for a SINGLE git plumbing call. tests/helpers/timeouts.cjs already documents
this exact failure and exists to fix it: HOOK_FANOUT_TIMEOUT_MS was created after
PR #3285 recorded "outcome=timed_out exitCode=null at exactly the 15000ms probe
bound while every other lane passed the same commit", and calls that "a bound
sized for the wrong class, not a slow machine". Our failure is that case
verbatim, so both sites move to the class norm rather than to a bigger number.

The same class also failed on `next` itself 21 hours earlier -- run 32608945654,
windows-latest/24 shard 1/3, `plan touching only src/ in a submodule project
keeps worktree isolation ENABLED` -- where tests/worktree-safety.test.cjs:5845
`runGate` fans out to `git config --file .gitmodules` under a hardcoded 30000.
Fixed too, since it is a defect in the tree regardless of which branch surfaced
it.

A survey of the whole tests/ tree found the same class-mismatch at further
bash fan-out sites bound under 60000ms, and the maintainer approved sweeping
them rather than leaving them latent to surface the same way one at a time. 16
fan-out sites across 16 files now use the class norm.

The sweep is class-correctness, not raising numbers until things pass. Sites
were moved ONLY where the bash body demonstrably spawns something (git, node,
npm, a CLI); self-contained shell snippets were left where they are, and are
listed as deliberately unchanged: pure if/printf bodies (copilot-install), pure
array/case builtins (code-review-pipeline-regression:638), a documented
pure-shell gsd_run stub (host-integration), single-process hook calls
(workflow-guard:222/271/302), and a deliberately tight 5000ms fast-check hook
(gsd-write-guard.property). Nothing was lowered. process-seam.test.cjs:513
(literal 300) is untouched on purpose -- it tests timeout BEHAVIOR, so raising
it would destroy what it asserts.

Shared file-level constants were the trap here, and were handled per file rather
than by redefinition: GIT_TIMEOUT_MS has ~15 users in git-base-branch and only 1
is a fan-out; WORKTREE_TIMEOUT_MS has 16 users in worktree.test.cjs and 3 are;
PROBE_TIMEOUT_MS has several in three more files. In each the CALL SITE was
changed and the constant left alone, so no single-plumbing-call site silently
inherited a 60s bound. The one exception is hooks-opt-in.test.cjs, where
HOOK_TIMEOUT_MS has exactly one consumer -- spawnHook, the fan-out itself -- so
redefining it is identical in effect and reads better.

Only two of these sites have actually been observed failing. The rest cite that
shared class and those two run ids rather than inventing evidence of their own.

Co-authored-by: sim <sim@local>
2026-08-23 21:21:21 -04:00
Tom Boucher
b9f51836e6 refactor(#3180): ADR-3180 behavior contract + cross-surface drift guardrails (#3223)
* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract

The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a
lower bound for the third consecutive time, and that two derivation families
had never been named at all.

ADR-3180 gains Decision 7 — a normative behavior contract that says what the
right answer IS for each derivation, not merely who owns it. A reviewer with
no written rule can only ask "does this look like the others", which is how a
fifth copy passes review. Decision 4 gains (d) scan surface is every authored
surface and an owner FILE is never exempt, only its named functions; and (e)
a surface that cannot be consolidated today ships ratcheted, never unguarded.

Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined
copies of its own body across five modules. All six now route through it;
`clampPercentFromFraction` is added for the one caller that already held a
fraction. Every migration is behaviour-identical — clampPercent's first line IS
the `total > 0 ? … : 0` ternary each copy carried. Guarded by
lint-completion-ratio-drift.cjs, which reports zero re-derivations with no
file-level exemption.

Prompt layer: workflow markdown re-derives live-plan counting in raw shell
(#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs
scans it with a shrink-only baseline of the 7 sites that exist today — new
sites fail, and a baseline entry that stops firing fails too, so an
acknowledgment can never outlive the thing it describes.

lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only
the four named canonical functions are exempt now. The blanket exemption was
pointed at the one file most likely to grow the next copy, and it had.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage

Five findings from the two orthogonal review passes, all fixed.

Decision 4(c) breach: the completion-ratio identity test asserted at the
OWNER, which is exactly the bypass that decision exists to close — a consumer
can call clampPercent and then post-process locally, leaving both the lint and
an owner-level test green. It now drives `roadmap analyze`, `query progress`
and `stats` and asserts on their own output, over a fixture containing a
`status: superseded` plan so a consumer that re-counted raw files would report
60 where the owner reports 75.

Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the
issue that removes them. They name Phase 8 (#3218) now.

The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical
sites were one indistinguishable key and migrating either would have left the
guard green with the other alive. Entries carry an occurrence count; fewer than
acknowledged fails as a partial migration, more fails as a new copy.

Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's
test already had, and the fast-check property tests CONTRIBUTING requires for
clamp/budget-limit functions.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes)

`tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs`
under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`.
A fixed wall-clock budget around a double spawn, running inside a container
that is concurrently executing the full ~31k-test suite, fails by construction
under load.

Confirmed against three full matrix runs. Every failure was shaped
`null !== 0` — the child was KILLED, never an assertion about the thing under
test. One captured probe had already printed the correct resolution
(`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It
reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The
victim subset varies by run and by lane.

What these tests are actually about is suite-token RESOLUTION — `unit` as a
bare token in --files/--files-from. Executing the seeded trivial files is
incidental and is the entire timeout surface, so the assertions move
in-process against the same functions `main()` calls, in the same order.
`parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are
exported for that; no behavior, signature or logic changed.

No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the
harness for real and asserts exit codes end to end, on a 120s budget.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: delete the three elapsed-time assertions

CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all
three are load-sensitive: on a saturated bench each can fail while the code
under test is correct. In every case the load-bearing assertion sits on the
line above and the timing line adds no discrimination.

run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s
harness backstop?" — is already answered by the assertion above it. A backstop
kills by signal, which surfaces as status null, never 124. Observed directly
this session: three matrix runs produced exactly that null shape from killed
children.

normalize-test-command and context-predicates: both bounded a ReDoS check.
A threshold only ever separates "fast" from "slightly slow", which is bench
load, not correctness — catastrophic backtracking on 800 KB of input does not
take 251ms, it does not finish at all. A real regression therefore shows up as
the suite being killed on that test, which is louder and more reliable than a
number. The structural assertions (returned unchanged; cleanly rejected) are
what actually carry those tests, and they stay.

The sweep now reports zero elapsed-time assertions in tests/. The remaining
Date.now() uses are unique-path suffixes, barrier deadlines, fixture
timestamps and fake mtimes — none of them assertions.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3180): backfill changeset PR number (#3223)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows

The baseline keys on (file, trimmed text). `file` came from scanTree's
`path.relative()`, which uses NATIVE separators, while the committed baseline
stores POSIX. On Windows every violation was therefore unmatched — reported as
FRESH — and every baseline entry matched nothing — reported as STALE. The guard
failed 100% of the time there, on both CI shards:

  ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale
    + { file: 'gsd-core\\workflows\\execute-plan.md', ... }

The remote runner this repo gates on is Linux-only and cannot see this class at
all; the GitHub Actions Windows lane is what caught it.

Normalization is unconditional — never gated on process.platform. A
platform-conditional normalizer makes the POSIX path the special case and
leaves the Windows branch unexercised on every other OS, which is the same
blind spot in a different place. It is applied at one seam inside
findPromptDrift, which builds `file` on every returned violation, so the
baseline key, the --update writer, the stderr report and the tests all consume
one normalized value.

The regression tests drive a Windows-shaped relPath directly and run on every
OS rather than skipping off-Windows — a test that only runs on the platform
where the bug lives is why this escaped. They include a sanity check that
un-normalized input does NOT match, so the assertion cannot pass vacuously.

Audited the three sibling guards: none keys against a committed cross-platform
baseline, and their exemption keys are path.join-built, so producer and
consumer share the native convention. Left correct code alone rather than
making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing
there would break those three on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 16:05:17 -04:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
sim
7dd9e59f6b test(#3090): stop exempting violations under categories that do not fit
An allow-test-rule annotation citing a category that does not apply is worse
than no annotation, because it reads as reviewed. Eight were confirmed by
reading the assertions each one covered, and auditing the rest found five more
plus one refutation — a converter test whose wording described the wrong
mechanism while the covered assertion genuinely was deployed-text.

The instructive one used the CANONICAL string for the same mistake: STATE.md
command output labelled as a deployed artifact. A canonical string is not
evidence the category fits, which is why normalising strings alone would have
laundered the problem rather than fixed it. Every mapping the audit had inferred
rather than code-verified was spot-checked before rewriting, and the ones that
turned out not to fit were re-annotated rather than relabelled.

Fourteen STATE.md assertions had a typed extractor available all along and now
use it; their annotations came out because nothing needs exempting. Eight
assertions genuinely need a production change first — CLI stdout and stderr with
no structured mode — and are tagged pending-migration-to-typed-ir citing #3090,
which is what that category is for. It had zero real uses before this, while one
file carried a real citation to migration issue #2974 under a non-canonical tag.

Six annotations covered assertions that do no text matching at all. An exemption
for a violation that does not exist is noise that makes the real ones harder to
audit; those are removed.

atomic-write-coverage gains the annotation it always warranted — its own
docstring describes a structural-regression-guard while the file carried none.

Fifty-nine non-canonical strings across roughly thirty files are normalised, and
the allow-test-rule allowlist is regenerated to match. 472 annotations became
463: every one now uses a canonical category, and the two remaining
non-canonical strings are ESLint RuleTester fixtures, not annotations.

Refs #3057

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 17:20:56 -04:00
Tom Boucher
1c1af70a4b refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator

Removes the 19 committed path->hash manifests, the two per-file size
baselines, tests/golden-install-parity.test.cjs, and
scripts/gen-golden-install-parity-zcode.cjs. These were pure functions
of the source tree (ADR-2719); the differential attribution check
(tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs)
is now the sole gate for emitted-artifact propagation.

tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs
are unchanged (ADR-2719 section 7 exception).

Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's
existence guard, .gitattributes, package.json scripts, the emitted-provenance
totality guard's IO, the differential check's baseline acquisition, CI
wiring to publish/restore the baseline artifact, and docs.

* refactor(#2724): make the differential attribution check self-sufficient

Three fixes required to delete the golden fixtures without breaking CI:

- scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs
  from the three rules that named it. #2759's missingRuleTestFiles guard
  hard-throws at module load if a rule names a test file absent from
  disk, which would break the changes job on every PR the moment the
  fixture-deletion commit landed.

- tests/helpers/emitted-provenance.cjs: loadManifests() read the
  committed golden fixture directory. With that directory deleted at
  every future ref, this would throw at module load forever, taking
  the Phase 2 totality guard down with it. Rebuilt from real installer
  spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest),
  the same shape emitted-runtime.cjs's currentManifests() already uses.

- tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs:
  the real-tree test's baseline acquisition swaps from
  baselineManifestsAtRef(base) (git show at a ref that no longer carries
  fixtures) to resolveBaseline()'s documented precedence: env, then the
  on-disk cache, then an in-job build. The build fallback
  (buildBaselineAtRef, new) checks out base into a throwaway git
  worktree and runs the new scripts/gen-emitted-baseline.cjs there --
  no npm ci needed, since bin/install.js and the test helper shells are
  Node-builtins-only. That script also publishes the baseline artifact
  from CI's push-to-next job (wired in a follow-up commit).

* refactor(#2724): retire the merge-driver bridge and per-file size baselines

The Phase 1 bridge (#2721) is retired now that the artifacts it guarded
are deleted: scripts/git-merge-regen-driver.cjs, its test, and the
'setup:merge-driver' npm script are removed, and the .gitattributes
merge=gsd-regen/linguist-generated block for the three deleted-path
globs is dropped. tests/fixtures/install-tree/*.json keeps its normal
merge behavior, unchanged (ADR-2719 section 7).

scripts/update-size-baseline.cjs and its test are removed: their sole
purpose was regenerating tests/workflow-size-baseline.json and
tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm
script and its step in 'regen:derived' go with it. The per-file
baseline describe blocks in tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs are removed for the same reason; the
independent loose-tier hard caps are untouched. The differential
attribution check's size ratchet (tests/emitted-diff.cjs, already
shipped in #2723) is the replacement anti-creep mechanism.

'npm run gen:golden' is replaced by 'npm run gen:install-tree', which
keeps regenerating tests/fixtures/install-tree/*.json (the one artifact
family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's
error messages point at the new command name.

tests/golden-parity-single-source.test.cjs's anti-divergence guard
(#2266) is retargeted from the two deleted golden-parity consumers to
their two replacements (tests/helpers/emitted-runtime.cjs and
tests/helpers/emitted-provenance.cjs), which import buildParityManifest
the same way — the divergence risk the guard exists for is unchanged.

Also wires CI: a new publish-emitted-baseline job runs
scripts/gen-emitted-baseline.cjs after a push to next and caches the
result keyed on the sha; the test and test-full jobs restore that cache
on pull_request events, keyed on the PR's base sha, and export
GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree
test to pick up.

* docs(#2724): flip ADR-2719 to Accepted and update contributor docs

Status: Proposed -> Accepted. Regenerated docs/adr/README.md index.

CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET.
EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET.
AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry)
no longer point at the deleted golden-install-parity fixtures, size
baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver /
git-merge-regen-driver.cjs bridge. Editing shipped content now
requires zero manual fixture regeneration, documented against the
differential attribution check instead of the deleted commands.

* docs(#2724): add changeset for removed golden-parity commands

* fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref

check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs
token in CONTEXT.md resolves to a real file. The RULESET.
EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks,
which the checker reads as a live reference, not historical prose.

* test(#2724): retarget ci-test-scope tests off the deleted golden test

tests/ci-test-scope.test.cjs asserted specific RULES entries select
tests/golden-install-parity.test.cjs, and that every rule selecting it
also selects both emitted gates. Both premises broke when the golden
test was deleted (#2724): the deleted filename never re-appears in
targeted_tests, and there was no longer a third file for the gates to
travel alongside. Retargeted the two selection describe blocks to
assert tests/emitted-provenance.test.cjs directly (the drift guard the
golden gate's rules were retargeted to), and simplified the third block
to assert the two emitted gates always travel together, without
reference to the golden filename.

* docs(#2724): repoint two contributor how-to guides at the differential check

Both guides told contributors to regenerate a baseline against
tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed
at the differential attribution check (tests/emitted-attribution.test.cjs,
ADR-2719), which needs no manual regeneration step.

* fix(#2724): repair phase6-capstone-conformance's deleted-baseline read

An independent orthogonal review caught a real regression this branch
introduced into a test file the branch's diff never touched:
tests/phase6-capstone-conformance.test.cjs read
tests/workflow-size-baseline.json (deleted earlier in this branch) with
no fallback, so the whole suite would throw ENOENT the moment this
branch landed. The test's actual intent — prove the host-loop workflow
files are real, tracked, non-empty docs — is preserved by asserting the
live byte count via the same shared counter (scripts/workflow-size.cjs)
the size guards already use, instead of a committed snapshot.

Also, from the same review: a stale doc comment in
scripts/workflow-size.cjs still named the deleted
scripts/update-size-baseline.cjs as a consumer, and
buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left
two fs.rmSync calls unguarded against masking the primary result/error,
inconsistent with the try/catch already wrapping the git cleanup beside
them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef
explaining why it (and its siblings) are kept despite having no
production caller post-cutover — they still answer real questions
about refs that predate the cutover.

* fix(#2724): repair three real regressions found by remote verification

1. tests/emitted-provenance.test.cjs's two hostile-input tests
   (non-object manifest, unreadable fixture) drove loadManifests(tmp)
   and monkeypatched fs.readFileSync, both premised on the deleted
   fixture-directory read this branch already replaced with real
   installer spawns -- the negative assertions silently stopped firing.
   loadManifests() now accepts injected {families, install, build,
   clean} (defaulting to production values), giving the tests a real
   seam to drive a bad build result and a build failure through the
   ACTUAL loader instead of a reimplementation, and added coverage that
   clean() still runs on both paths.

2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE'
   steps hardcoded shell: bash, which is wrong on windows-latest (native
   pwsh) and on test-full's macos-latest legs (native zsh per that job's
   own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning
   .test.cjs) caught it. Replaced the inline bash script with
   scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a
   bare 'node <path>' command line has no shell-specific syntax, so it
   runs correctly under bash, zsh, and pwsh without a shell override.

tests/phase6-capstone-conformance.test.cjs's deleted-baseline read
(caught by the same remote run, at a commit prior to this one) was
already fixed in d0c3b1242 and is not touched here; verified still
passing after these changes.

* fix(#2724): revive ADR-1610's new-file size cap inside the differential

An isolated review caught a real regression: deleting
tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP
(ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor)
with no successor. tests/helpers/emitted-diff.cjs's size ratchet
already 'continue's past any file absent from sizeBaseline -- exactly
the files this cap exists to bound -- so a brand-new workflow file
sized 32,769-40,960 bytes passed CI clean and shipped, then risked
silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted
and never referenced anywhere in this branch.

Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own
size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline,
name) signal the growth check already computes -- 'new' is exactly
'present in sizeCurrent, absent from sizeBaseline'. Not ack-able,
matching the tier hard caps it sits beside: the fix is extraction, not
an acknowledgment entry. Documented, disclosed narrowing: the pure
differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering
(tests/workflow-size-budget.test.cjs's classification), so a
legitimately large new file must extract rather than tier in, one
release earlier than an existing file would need to. ADR-1610 itself is
left unamended -- this restores its decision rather than re-litigating
it.

Also fixes a stale comment plus a redundant real 19-installer-spawn
assertion left over from the pre-injection-seam version of
tests/emitted-provenance.test.cjs's build-failure test, and annotates
3 of 4 stale golden-fixture citations in
docs/reference/host-integration-capability-matrix.md as superseded
(the 4th is an accurate historical PR narrative, left alone).

* fix(#2724): repair three red CI defects on the golden-fixture cutover

Windows-only provenance false attribution (defect A): the `hooks-built`
provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are
Windows-only installer output (ensureCodexHooksJsonSessionStart /
ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping
the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so
the self-attribution resolved to a path that exists on no platform. Only
windows-latest ever emits the key, so this only failed there. Fixed by
special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated
rule) — a dedicated rule would match zero paths, and therefore report as a
dead rule, on every non-Windows lane of the same totality guard. `sources`
already supported per-match functions; `transforms` is extended to support
the same shape so the attribution can vary by match within one rule.

Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef`
ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but
that script is new in this PR and therefore absent at any base ref that
predates it — every call failed closed with "Cannot find module". Fixed by
running the PR checkout's own generator against the worktree via a new `--dir`
parameter, decoupling "which copy of the script runs" from "which tree it
measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override,
threaded down to `runMinimalInstall`'s new `installScript` override). This is
not just a bootstrap fix: a differential needs ONE measurement schema applied
to both sides, or the two stop being comparable the moment that schema
evolves — running each side's own copy would silently reintroduce that risk.
Verified locally end-to-end against real origin/next: resolves a valid
{version, sha, manifests, sizes} artifact with the correct sha and no leaked
worktree.

Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let
docs-lint evaluate the fragment for the first time; it already passes
(docs/TESTING-SUITES.md and friends already document the removed scripts).

Also fixed while in this file: an eslint no-unused-vars warning surfaced by
the changed lint run (unused `cleanup` import in
tests/emitted-provenance.test.cjs).

Added regression coverage for both A and B: a cross-platform spot-check that
drives the real hooks-built rule against `.cmd` keys directly (not through a
real Windows install), and a real-tree test that drives buildBaselineAtRef
against a base ref verified (via git cat-file) to lack the generator, both
skipping honestly rather than false-passing when their precondition does not
hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test

Two isolated-review findings on PR #2767:

- `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the
  wrapped `hooks/<name>.js` script, asserting a byte-provenance link that
  does not exist — traced against buildCodexHookWindowsShimIR
  (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that
  same file) flows into the .cmd bytes, never its content. Point `sources`
  at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention
  used elsewhere in the table. Since `sources` is checked before
  `transforms` in the differential, the wrong mapping silently excused any
  .cmd byte movement caused by editing the wrapped .js file.

- The `buildBaselineAtRef` regression test skipped unless a resolvable base
  ref still lacked scripts/gen-emitted-baseline.cjs — true only until this
  PR merges, after which every base ref carries the file and the test skips
  forever with zero ongoing coverage. Rebuilt hermetically: synthesize the
  missing-generator condition in-place via git plumbing (a throwaway commit,
  child of HEAD, with just that one file removed from a scratch index),
  never touching the real working tree, HEAD, or index, and never depending
  on ambient history or remotes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path

The runner container mounts the repo at a path owned by a different uid than
the process running the suite, so git's dubious-ownership protection refuses
every git operation there. GitHub Actions never hits this because
actions/checkout registers the workspace as safe automatically; this
runner's container does not.

buildBaselineAtRef is the production build-fallback the sole remaining
emitted gate depends on (resolveBaseline's in-job-build leg), not just a
test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs)
that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's
worktree add/remove/prune, and the hermetic regression test added in the
prior commit — funnels through, plus gen-emitted-baseline.cjs's own
rev-parse (now reusing that same wrapper instead of a second execFileSync,
so the fix has one source of truth). Each call declares -c
safe.directory=<the exact directory it already operates on>, never the *
wildcard.

Audited every other helper on this surface (emitted-diff.cjs,
emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them
shell out to git at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:41:43 -04:00
Tom Boucher
d04592de58 fix(#2758): select the emitted differential wherever golden-parity runs (#2759)
* fix(#2758): select the emitted differential wherever golden-parity runs

Add tests/emitted-provenance.test.cjs and tests/emitted-attribution.test.cjs
to every scripts/ci-test-scope.cjs rule that already selects
tests/golden-install-parity.test.cjs, so the ADR-2719 dual-run differential
travels with the golden on the targeted CI lane instead of being selected by
no rule at all.

Add an independent module-load totality guard (missingRuleTestFiles) that
throws when any rule names a test file absent from disk -- the guard that
would have caught the post-Phase-4-cutover hole. It immediately surfaced
three pre-existing phantom entries left behind by consolidation epic #1969
(bug-3588/bug-10/bug-3683 filenames folded into other suites months ago but
never removed from the rule table); fixed in the same change rather than
deferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2758): trim unused exports and normalize the path require style

Code-review (Standards axis) flagged two judgement-call smells: exporting
classify/isInertCi with no caller (Speculative Generality), and requiring
path with a node: prefix while the file's other core requires do not
(inconsistent style within one file). Both addressed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 10:48:09 -04:00
Tom Boucher
bf0d715733 fix(#2472): cost-balanced test sharding and pinned CI base commit (#2480)
* fix(#2472): weight-aware shard partition

Windows shard 1/3 hit the 20-minute job cap with no failing assertion. Root
cause is the shard layer, not the chunk layer: selectShard partitioned by
sorted ARRAY INDEX (k % n, #1212), which balances file COUNTS and ignores
file COST. On the real unit suite that produced 12.4m / 19.2m / 15.2m — a
1.23x max/ideal ratio leaving the heaviest shard 5% under the cap. Because
assignment keyed off position, inserting one test file re-indexed every file
after it and could tip that shard over; deterministic, so a re-run reproduced
it exactly.

This is NOT the chunk packer (#2456/#2463). That fix works and applies one
level down, WITHIN a shard. The across-shard partition predated it and never
consumed the cost table. Both layers now share one cost model.

selectShard takes an optional weightOf and, when given one, partitions by LPT
(longest-processing-time-first) — the same algorithm packChunks uses. Omitting
it keeps the legacy round-robin byte-identical, so every existing test above
still exercises that path unchanged and callers without timing data lose
nothing. A missing timings table yields uniform weight 1, under which LPT
degenerates to the equal-count split.

Projected on the real suite: 16.4/17.3/13.0 -> 15.6/15.6/15.6 (worst shard
17.3m -> 15.6m).

Tests: a skewed-cost regression (round-robin clusters all four heavy files
onto one shard at 2.98x ideal; LPT does not), back-compat equivalence,
determinism, tie-breaking, order preservation, and two fast-check properties
— the partition is exhaustive and disjoint (getting this wrong silently DROPS
tests from CI, the worst failure mode for a harness), and no shard exceeds
average + heaviest file.

Two assertions were corrected during authoring rather than shipped wrong:
- an initial "LPT within 4/3 of ideal" bound was false. The 4/3 figure is
  relative to the OPTIMAL makespan, not the average, and the two differ when
  item sizes force a pairing. Replaced with Graham's average+max bound, which
  is what is actually provable.
- "weighted is never worse than round-robin" is also false; fast-check
  falsified it with [19316,10190,1,9128,29353,20227] over 2 shards (rr 48670,
  lpt 48671). Round-robin can win by luck on a specific input. Dropped, with
  the counterexample recorded in place so it is not re-asserted later.

Closes #2472

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): rotate tied bins; restore #1212 test block; lazy cost table

Isolated-review findings, all fixed.

HIGH — zero weights collapsed the whole partition onto shard 1. The
lightest-bin scan compared weight only, and adding a zero-weight file leaves
its bin's weight unchanged, so bin 0 stayed tied-minimum forever and every
such file landed on it. Verified: all-zero weights gave shard1=[a..f],
shard2=[], shard3=[] — two of three CI runners idle while one ran everything.
Reachable through safeWeight's own clamp (a NaN/negative/Infinity entry in a
corrupted or hand-edited timings table) and through any genuine 0ms
measurement, so the clamp reproduced the exact failure its comment claimed to
prevent. Ties now break on file COUNT after weight, which rotates. Pinned by
two regression tests (all-zero, and clamped NaN/negative/Infinity) plus a
property over list size x shard count. The live table has no 0ms entries
(min 19ms), so production was not affected — but nothing prevented it.

MEDIUM — the new describe block had swallowed #1212's pre-existing property
test, which is why a test under a "weight-aware" heading never passed a
weigher. That was a bad block boundary in the previous commit, not a bad
test: the #2472 describe was opened before #1212's last test instead of
after. Moved back where it belongs; #1212 is 762-879 and #2472 is 894-1082.

LOW — that relocated property test ran unseeded. Seeded (12120) per the
repo's property-test convention so a failure reproduces. Verified passing
under the new seed.

LOW — hoisting the timings load above the shard block charged a readFileSync
+ JSON.parse to invocations that exit before needing it (empty selection,
--files matching nothing). Now lazily memoized, so neither consumer reads the
table unless it is used and it is still read at most once.

Real-suite projection unchanged at 15.6m / 15.6m / 15.6m.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2472): correct stale round-robin sharding descriptions

The partition is now cost-balanced, so the header block in run-tests.cjs
and the two comments in test.yml describing '--shard' as a round-robin over
sorted file index were actively wrong. Updated to describe LPT over measured
duration, and to state the degenerate case explicitly: with no timing data
every file weighs the same and the partition collapses back to k % n, which
is why the pre-existing #1212 CLI tests still pass unchanged (their nine
synthetic files are absent from the timings table, so all take the identical
median weight).

Remaining 'round-robin' mentions are correct — they describe the unweighted
fallback path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): shard diagnostics, cost-routing E2E test, table validation

Second orthogonal review (operational lens) findings, all fixed.

HIGH — cross-runner partition divergence. Each of the up-to-12 CI jobs runs
its own 'merge base into head' and computes its own partition, so if the
inputs differ between jobs (the file list, or the timings table) two jobs can
place the same file in different shards or in none. Every job stays
internally exhaustive and disjoint, so nothing errors: a test simply never
runs and CI stays green.

The risk class is pre-existing — round-robin diverges identically when the
file set differs between jobs, which is literally this issue's insertion
instability — but weighting adds tests/test-timings.json as a second input
that must match, so it widens the hole. Properly closing it means pinning the
partition inputs per run, a workflow change beyond this fix.

What IS closed here is the silence. Each shard now prints an input
fingerprint over the FULL pre-partition list and the weight assigned to each
file — deliberately not this shard's slice, which would differ by design and
be useless for comparison. All shard jobs of one run must print an identical
sig; a mismatch is direct proof the runners disagreed about the input.
Verified: three independent computations agree, and the sig changes when the
input drifts by one file.

MEDIUM — nothing proved main() actually threads fileWeightOf() into
selectShard. Every pre-existing --shard E2E test uses synthetic filenames
absent from the real table, so all collapse to a uniform median weight, under
which LPT is mathematically identical to k % n — a typo on that one wiring
line would have passed the whole suite. Added an E2E test that injects a
table via RUN_TESTS_TIMINGS_FILE with differing costs, placing the heavy
files at exactly the indices round-robin hands to shard 1, and asserts shard 1
does NOT receive all three. Plus a test that all three shards emit the same
sig.

MEDIUM/LOW — no observability. The diagnostic line now reports files,
weighed count, aggregate weight, and whether the table loaded, so a table
that silently failed to parse shows table=absent/weighed=0 instead of being
indistinguishable from a healthy load. (The reviewer confirmed the advisory
fallback is already live on next: feat-2296-provider-escalation.test.cjs is
missing from the table.)

LOW — typeof [] === 'object', so a hand-edit turning the map into a list was
accepted as a valid table. Now rejected via Array.isArray, falling back to
uniform weight like any other malformed table.

LOW — stale round-robin wording in ci-test-scope.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): pin every CI job to one base commit

Closes the cross-runner divergence at its source instead of only making it
visible.

Each job of a run executes the rebase-check step independently, minutes apart
across a 12-job matrix, and merged the MOVING origin/<branch> ref. If the base
advanced mid-run, different jobs merged different trees. That was survivable
when jobs only had to agree on pass/fail; it is not once they must agree on a
PARTITION. Each shard job computes the whole split and keeps its own slice, so
jobs working from different trees can place a file in two shards or in none —
and every job still looks internally consistent, so nothing errors. A test
silently never runs and CI stays green.

ci-rebase-check.cjs now accepts CI_REBASE_BASE_SHA and pins BOTH the fetch and
the merge to that one commit, so the two can never disagree. test.yml passes
github.event.pull_request.base.sha on all three rebase-check steps; that value
is fixed for the life of a run, so all jobs merge the identical base.

This also closes the PRE-EXISTING half of the divergence. Round-robin had the
same exposure whenever the test-file set differed between jobs — that is this
issue's insertion instability — so the pin fixes the older hole too, not just
the timings-table input weighting added.

Only a full 40-hex sha is accepted; empty (push/workflow_dispatch), malformed,
or injected values fall back to the branch ref rather than handing an arbitrary
string to git fetch as a refspec. resolveBaseRefs is extracted pure and
exported, and runMain is guarded behind require.main === module, so the pin
contract is testable without spawning git.

Tests (tests/ci-test-scope.test.cjs): every rebase-check step must carry the
pin; a valid sha pins both refs; absence falls back correctly; and five hostile
values — short sha, uppercase, --upload-pack= injection, ref expression, empty
— are each rejected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 09:15:59 -04:00
Tom Boucher
89b1bef881 refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267.
2026-07-14 19:22:13 -04:00
Tom Boucher
241b08fa18 ci(#1975): exclude release-tarball-smoke from scoped lane (fixes Windows chunk timeout)
This consolidation PR's breadth (28 changed test files) exposed a scoped-test-lane
capacity limit: ci-test-scope pulls the 3–6 min release-tarball-smoke.install.test.cjs
(npm pack + npm install -g, 10MB/1499 files) into the targeted+windows lane whenever
install files change AND when it is itself a changed file — bundling it with the other
27 files overran the 600s per-chunk timeout on windows-latest-24 (deterministic).

release-tarball-smoke has its OWN dedicated workflow (.github/workflows/install-smoke.yml,
triggered on the production install paths), so its scoped-lane run is redundant. Add a
SCOPED_LANE_EXCLUDE guard that drops it from both targeted_tests and windows_tests however
it entered (matched rule OR changed-file), and remove it from the install rule's tests list.
Update the ci-test-scope.test.cjs assertion accordingly. No coverage lost.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
b2ed7940c9 test(#1975): scope GSD_TEST_MODE off in real-install folds; fix ci-test-scope fixture
gsd-test surfaced 21 failures:
- 19: folded real-install suites (bug-1834 .sh hooks, enh-2380 --skills-root, fix-1521
  install stamping, bug-2136 .sh hook version) spawn install.js and assert side effects,
  but their host suites (install-minimal-hooks/install.test/managed-hooks) set
  GSD_TEST_MODE=1 at collection time — the install child inherited it and suppressed
  the writes. Clear GSD_TEST_MODE in each of those blocks (before/after; standalone had
  it unset), so the child performs a real install.
- 2: ci-test-scope A1 used deleted tests/bug-1974-context-exhaustion-record.test.cjs as a
  fixture; scopeFor filters nonexistent paths, so it fell back to ['unit']. Repointed to
  its consolidation destination tests/perf-317-context-monitor-fs.test.cjs (an existing test).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
6d072435d0 test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their
canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks,
read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping
the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim
block-scoped describe wrappers; 427 subtests conserved 1:1.

Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT
value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools
stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec,
so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard.

Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6,
docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456;
prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md
(EN + ja/ko/pt/zh) and ADR-0002. lint:ci green.

Part of epic #1969. Closes #1975.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
9d52043f50 feat(#1733): normalize-path-in-content production AST rule + fix Windows agent-skills content leak (Phase 5) (#1736)
* feat(#1733): normalize-path-in-content production AST rule (Phase 5)

ADR-1703 Phase 5 — the first production-code rule. local/normalize-path-in-content
(src/**/*.cts, @typescript-eslint/parser): flags a path-returning fn result
(path.basename excluded — returns a separator-less filename) interpolated into an
@-reference / config-dir markdown body without .replace(/\\/g,'/') normalization,
per RULESET.CONTENT-PATH-NORMALIZATION / DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT.

Build-and-assess found the canonical defect site (computePathPrefix) already
compliant and only 1 src/ hit — a false positive (path.basename in a status
message) — eliminated by narrowing (exclude basename; require a real @-ref/
config-dir marker, not bare .md). 0 src/ violations: clean forward-prevention.

The out-of-band disable-ban now scans src/**/*.cts too (typescript-estree) so the
production rule also cannot be eslint-disabled. Registered (error) + PROTECTED_RULES;
CONTEXT.md predicates + how-to doc updated.

- RuleTester suite (26 cases)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1733): add changeset for Windows agent-skills path-leak fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: harden mutation-matrix.cjs stdin read against EAGAIN on non-blocking pipe

scripts/mutation-matrix.cjs read piped stdin via readFileSync(process.stdin.fd).
On macOS libuv marks the stdin pipe fd non-blocking, so a synchronous read can
throw EAGAIN before the writer fills the pipe — intermittently, under heavy CI
shard load — aborting the script (status 2) and flaking mutation-matrix-ratchet.
Replace with readStdinSync(): an fs.readSync loop that retries on EAGAIN (1ms
synchronous Atomics.wait yield), stops on 0-byte/EOF, and rethrows other errors.
Deterministic regression test injects EAGAIN via an fs.readSync monkeypatch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* ci: re-run golden-install-parity on src/lib + installer changes (close drift guard)

golden-install-parity hashes every installed bin/lib/*.cjs per runtime, so it
must re-run whenever the built lib could change. ci-test-scope selected it for
neither src/** nor installer changes, so a source-only edit (e.g. #1691's
milestone.cts/roadmap.cts) recompiled bin/lib and silently drifted the golden
fixtures past the scoped lane. Add golden-install-parity.test.cjs to both the
'TS runtime sources' and 'installer and package layout' selection rules, with
behavioral regression tests for each.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: review-bot <review-bot@gsd>
2026-06-25 21:49:22 -04:00
Tom Boucher
f08b177215 feat(#1726): G1-G6 portability AST rules; fix all offenders; delete the ratchet (Phase 4) (#1731)
Phase 4 of epic #1702. Closes #1726.
2026-06-25 18:24:39 -04:00
Tom Boucher
22f56f4431 ci(#1212): shard windows full-test lane to remove timeout cliff (#1222)
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.

Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.

- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
  round-robin partition (fileIndex % n === i-1) over the SORTED selected file
  list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
  pure no-op. The 28K Windows argv chunking is preserved within each shard. A
  legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
  sharding still hits the discovery hard error. Composes with --suite and is
  order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
  cross-product (explicit include rows — a base shard dim does not cross-product
  with include legs, and a nested matrix.leg.os is unresolvable by the H1
  shell-policy linter). Unit suite runs sharded; integration/security run once
  per leg (shard 1). The Required tests fan-in is unchanged: it already needs
  test-full and checks the matrix-aggregate result, so a failed/cancelled shard
  fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
  balance, determinism, boundaries, fast-check property) + parseShardArg
  validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
  guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
  denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.

Closes #1212

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 12:25:29 -04:00
Colin
a647053dcf ci(scope): narrow #494 invariant — changed tests join windows lane, not full matrix
full_matrix fired on 15/15 sampled PRs because any tests/** change forced it,
costing ~25 runner-minutes each. A changed test file now always joins the
scoped windows lane (covering the #482 OS-specific failure class per-file)
and still runs on ubuntu 22/24 via targeted_tests; the residual macOS /
windows-node-22 cross-product is covered on every push to next.

Also narrows WINDOWS_HINTS from 6 substrings (102/633 files, a ~10-minute
scoped lane) to windows/win32/shell/path — the dropped hints (workflow,
install, hook) are either platform-independent lint tests or already covered
by fullMatrix rules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
606363c416 chore(#846): remove unused PR-size labeler (size/S–XL) workflow (#848)
The PR Gate workflow's only job, size-check, labeled every PR with
size/S–size/XL based on lines changed. Those labels aren't used in any
review, triage, or automation flow, so the workflow was pure noise.

- Delete .github/workflows/pr-gate.yml
- Drop size-check from required status checks in both rulesets so PRs
  don't block forever on a check that never reports
- Remove pr-gate.yml from INERT_WORKFLOWS (ci-test-scope.cjs) and the
  knownInert list (ci-test-scope.test.cjs)
- Remove "PR Gate / size-check" from setup-branch-protection.sh

Closes #846

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 22:03:48 -04:00
Tom Boucher
988024c1a3 fix(#837): three-dot diff in ci-test-scope so docs-only PRs skip the heavy matrix (#841)
CI test-scope detection diffed changed files with a two-dot
`git diff --name-only base head`, where base is the moving tip of `next`.
A PR branch cut from a slightly older `next` surfaced every product file
`next` had gained since the merge-base, flipping product_changed/full_matrix
and running the full Windows/macOS matrix + coverage on docs-only PRs.

Switch to a three-dot `git diff --name-only base...head` (vs the merge-base),
matching GitHub's PR "Files changed" semantics. Add a regression test that
builds a stale-base topology, plus a guard test pinning `fetch-depth: 0` on
the `changes` job (required for the merge-base to be locally available).

Closes #837

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:28:45 -04:00
Tom Boucher
571d7b5a1c feat(#764): skip cross-platform test matrix for docs-only and inert-CI PRs (#798)
test.yml had no paths filter and the ci-test-scope classifier treated docs/
and every .github/workflows/* as code_changed, so documentation edits and
product-irrelevant automation tweaks still spun up the full Linux/Windows/macOS
matrix. Narrow the heavy matrix to changes that can actually affect the product
or the test pipeline.

- ci-test-scope.cjs: drop docs/ from code_changed (docs-only -> full skip; the
  required-tests fan-in still reports green). Add src/ to code_changed (it was
  missing -> a source-only PR previously skipped all tests). Add INERT_WORKFLOWS
  allowlist + isInertCi() + an "inert CI" rule, and a product_changed output that
  gates the heavy test/coverage jobs. Fail-safe: any workflow not on the inert
  allowlist defaults to the full matrix. A module-load assertion throws if a
  PROTECTED_WORKFLOWS entry (test/install-smoke/mutation/security-scan/release)
  is ever added to the inert set, so a weakening edit fails CI loudly.
- test.yml: keep the static 3-lane matrix (so the H1 shell-policy linter can
  still statically verify the Windows lane), gate test/coverage on
  product_changed, add a lightweight ubuntu-only test-inert job, and branch the
  required-tests fan-in on product_changed.
- docs-required.yml: run docs-parity-live-registry (gated on docs/ changes) so
  pure-docs PRs still catch live-registry drift without the matrix.
- tests: cover docs-only, inert-only, src/, pipeline, unknown-workflow fail-safe,
  mixed escalation, the code_changed=false -> no-lanes invariant, and protected-
  workflow tamper-evidence.

Closes #764

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 11:46:48 -04:00
Tom Boucher
cdd78bd2aa fix(#670): self-healing recovery for installer-migration checksum drift (#675)
Editing the body of an already-released installer migration drifts its computed
checksum (it hashes plan.toString()). The integrity guard then hard-aborted
every prior install on upgrade with "applied migration checksum changed" — a
100% reproducible blocker (v1.3.0, all platforms).

Already-applied migrations are filtered out of `pending` and never re-run, so
a drifted checksum is functionally inert. ADR-0008 anticipates checksum-mismatch
state as something the install-state layer must handle gracefully (plan -> apply
-> recover/report), not abort on.

This supersedes the published-checksum allowlist merged in #674 (per-release
maintenance debt — every historical checksum hand-pinned, still throws for any
unregistered value) with a general, self-healing recovery:

- Replace the throwing guard with non-fatal `collectAppliedChecksumDrift`,
  surfaced on `plan.checksumDrift`.
- Reconcile drifted stored checksums durably on the next state write
  (`reconcileDriftedChecksums`), idempotently (no perpetual writes).
- Relocate the "shipped migration bodies are immutable" rule to a CI baseline
  test that locks every shipped migration's checksum and fails on body drift —
  where #615 should have been caught, instead of blocking users.

Removes #674's legacyChecksums field, per-migration checksum pins,
published-checksums.json fixture, and compat test.

Fixes #670

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 13:44:05 -04:00
Colin
b8e15c7a98 fix: accept published installer migration checksums 2026-06-04 12:30:13 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
a7ed001b27 fix(#494): ci-test-scope selects checks a diff can break (tests->full matrix, docs->docs-parity) (#495)
classify() under-approximated breakable checks, so scoped PRs skipped the
check their diff would break and regressions reached next (#484 docs-parity,
#482 windows-22 EBUSY). Fail-safe widen: any tests/** change forces
full_matrix (OS-specific test failures); any docs/**, commands/**, agents/**
change marks code_changed and selects docs-parity-live-registry (its runtime
inputs). Updated the docs-only test that asserted the old buggy contract.

Fixes #494

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-05-29 19:03:22 -04:00
Tom Boucher
f5f51b5b49 fix(#408): align ci-test-scope smoke handling with #395 changeset (drop unconditional injection; unit fallback) (#420)
- Remove `DEFAULT_SMOKE_TESTS` and `WINDOWS_SMOKE_TESTS` constants (now dead after the unconditional injection block is dropped)
- Drop the `addAll(targeted, DEFAULT_SMOKE_TESTS)` / `addAll(windows, WINDOWS_SMOKE_TESTS)` block from the `codeChanged` branch
- When `codeChanged && targetedTests.length === 0`, push `'unit'` as the fallback suite token
- Two new regression tests in `tests/ci-test-scope.test.cjs` covering the no-injection and unit-fallback contracts (bug #408)
2026-05-27 22:59:26 -04:00
Tom Boucher
9dd2c87c6c ci: supersede #367 with combined tiered PR pipeline redesign (#369)
* ci: streamline PR pipeline gates

* test: annotate ci-test-scope stderr assertions for lint rule

---------

Co-authored-by: Colin <colin@solvely.net>
2026-05-26 19:11:58 -04:00