Commit Graph

20 Commits

Author SHA1 Message Date
Tom Boucher
1c1af70a4b refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator

Removes the 19 committed path->hash manifests, the two per-file size
baselines, tests/golden-install-parity.test.cjs, and
scripts/gen-golden-install-parity-zcode.cjs. These were pure functions
of the source tree (ADR-2719); the differential attribution check
(tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs)
is now the sole gate for emitted-artifact propagation.

tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs
are unchanged (ADR-2719 section 7 exception).

Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's
existence guard, .gitattributes, package.json scripts, the emitted-provenance
totality guard's IO, the differential check's baseline acquisition, CI
wiring to publish/restore the baseline artifact, and docs.

* refactor(#2724): make the differential attribution check self-sufficient

Three fixes required to delete the golden fixtures without breaking CI:

- scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs
  from the three rules that named it. #2759's missingRuleTestFiles guard
  hard-throws at module load if a rule names a test file absent from
  disk, which would break the changes job on every PR the moment the
  fixture-deletion commit landed.

- tests/helpers/emitted-provenance.cjs: loadManifests() read the
  committed golden fixture directory. With that directory deleted at
  every future ref, this would throw at module load forever, taking
  the Phase 2 totality guard down with it. Rebuilt from real installer
  spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest),
  the same shape emitted-runtime.cjs's currentManifests() already uses.

- tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs:
  the real-tree test's baseline acquisition swaps from
  baselineManifestsAtRef(base) (git show at a ref that no longer carries
  fixtures) to resolveBaseline()'s documented precedence: env, then the
  on-disk cache, then an in-job build. The build fallback
  (buildBaselineAtRef, new) checks out base into a throwaway git
  worktree and runs the new scripts/gen-emitted-baseline.cjs there --
  no npm ci needed, since bin/install.js and the test helper shells are
  Node-builtins-only. That script also publishes the baseline artifact
  from CI's push-to-next job (wired in a follow-up commit).

* refactor(#2724): retire the merge-driver bridge and per-file size baselines

The Phase 1 bridge (#2721) is retired now that the artifacts it guarded
are deleted: scripts/git-merge-regen-driver.cjs, its test, and the
'setup:merge-driver' npm script are removed, and the .gitattributes
merge=gsd-regen/linguist-generated block for the three deleted-path
globs is dropped. tests/fixtures/install-tree/*.json keeps its normal
merge behavior, unchanged (ADR-2719 section 7).

scripts/update-size-baseline.cjs and its test are removed: their sole
purpose was regenerating tests/workflow-size-baseline.json and
tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm
script and its step in 'regen:derived' go with it. The per-file
baseline describe blocks in tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs are removed for the same reason; the
independent loose-tier hard caps are untouched. The differential
attribution check's size ratchet (tests/emitted-diff.cjs, already
shipped in #2723) is the replacement anti-creep mechanism.

'npm run gen:golden' is replaced by 'npm run gen:install-tree', which
keeps regenerating tests/fixtures/install-tree/*.json (the one artifact
family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's
error messages point at the new command name.

tests/golden-parity-single-source.test.cjs's anti-divergence guard
(#2266) is retargeted from the two deleted golden-parity consumers to
their two replacements (tests/helpers/emitted-runtime.cjs and
tests/helpers/emitted-provenance.cjs), which import buildParityManifest
the same way — the divergence risk the guard exists for is unchanged.

Also wires CI: a new publish-emitted-baseline job runs
scripts/gen-emitted-baseline.cjs after a push to next and caches the
result keyed on the sha; the test and test-full jobs restore that cache
on pull_request events, keyed on the PR's base sha, and export
GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree
test to pick up.

* docs(#2724): flip ADR-2719 to Accepted and update contributor docs

Status: Proposed -> Accepted. Regenerated docs/adr/README.md index.

CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET.
EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET.
AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry)
no longer point at the deleted golden-install-parity fixtures, size
baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver /
git-merge-regen-driver.cjs bridge. Editing shipped content now
requires zero manual fixture regeneration, documented against the
differential attribution check instead of the deleted commands.

* docs(#2724): add changeset for removed golden-parity commands

* fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref

check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs
token in CONTEXT.md resolves to a real file. The RULESET.
EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks,
which the checker reads as a live reference, not historical prose.

* test(#2724): retarget ci-test-scope tests off the deleted golden test

tests/ci-test-scope.test.cjs asserted specific RULES entries select
tests/golden-install-parity.test.cjs, and that every rule selecting it
also selects both emitted gates. Both premises broke when the golden
test was deleted (#2724): the deleted filename never re-appears in
targeted_tests, and there was no longer a third file for the gates to
travel alongside. Retargeted the two selection describe blocks to
assert tests/emitted-provenance.test.cjs directly (the drift guard the
golden gate's rules were retargeted to), and simplified the third block
to assert the two emitted gates always travel together, without
reference to the golden filename.

* docs(#2724): repoint two contributor how-to guides at the differential check

Both guides told contributors to regenerate a baseline against
tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed
at the differential attribution check (tests/emitted-attribution.test.cjs,
ADR-2719), which needs no manual regeneration step.

* fix(#2724): repair phase6-capstone-conformance's deleted-baseline read

An independent orthogonal review caught a real regression this branch
introduced into a test file the branch's diff never touched:
tests/phase6-capstone-conformance.test.cjs read
tests/workflow-size-baseline.json (deleted earlier in this branch) with
no fallback, so the whole suite would throw ENOENT the moment this
branch landed. The test's actual intent — prove the host-loop workflow
files are real, tracked, non-empty docs — is preserved by asserting the
live byte count via the same shared counter (scripts/workflow-size.cjs)
the size guards already use, instead of a committed snapshot.

Also, from the same review: a stale doc comment in
scripts/workflow-size.cjs still named the deleted
scripts/update-size-baseline.cjs as a consumer, and
buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left
two fs.rmSync calls unguarded against masking the primary result/error,
inconsistent with the try/catch already wrapping the git cleanup beside
them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef
explaining why it (and its siblings) are kept despite having no
production caller post-cutover — they still answer real questions
about refs that predate the cutover.

* fix(#2724): repair three real regressions found by remote verification

1. tests/emitted-provenance.test.cjs's two hostile-input tests
   (non-object manifest, unreadable fixture) drove loadManifests(tmp)
   and monkeypatched fs.readFileSync, both premised on the deleted
   fixture-directory read this branch already replaced with real
   installer spawns -- the negative assertions silently stopped firing.
   loadManifests() now accepts injected {families, install, build,
   clean} (defaulting to production values), giving the tests a real
   seam to drive a bad build result and a build failure through the
   ACTUAL loader instead of a reimplementation, and added coverage that
   clean() still runs on both paths.

2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE'
   steps hardcoded shell: bash, which is wrong on windows-latest (native
   pwsh) and on test-full's macos-latest legs (native zsh per that job's
   own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning
   .test.cjs) caught it. Replaced the inline bash script with
   scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a
   bare 'node <path>' command line has no shell-specific syntax, so it
   runs correctly under bash, zsh, and pwsh without a shell override.

tests/phase6-capstone-conformance.test.cjs's deleted-baseline read
(caught by the same remote run, at a commit prior to this one) was
already fixed in d0c3b1242 and is not touched here; verified still
passing after these changes.

* fix(#2724): revive ADR-1610's new-file size cap inside the differential

An isolated review caught a real regression: deleting
tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP
(ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor)
with no successor. tests/helpers/emitted-diff.cjs's size ratchet
already 'continue's past any file absent from sizeBaseline -- exactly
the files this cap exists to bound -- so a brand-new workflow file
sized 32,769-40,960 bytes passed CI clean and shipped, then risked
silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted
and never referenced anywhere in this branch.

Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own
size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline,
name) signal the growth check already computes -- 'new' is exactly
'present in sizeCurrent, absent from sizeBaseline'. Not ack-able,
matching the tier hard caps it sits beside: the fix is extraction, not
an acknowledgment entry. Documented, disclosed narrowing: the pure
differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering
(tests/workflow-size-budget.test.cjs's classification), so a
legitimately large new file must extract rather than tier in, one
release earlier than an existing file would need to. ADR-1610 itself is
left unamended -- this restores its decision rather than re-litigating
it.

Also fixes a stale comment plus a redundant real 19-installer-spawn
assertion left over from the pre-injection-seam version of
tests/emitted-provenance.test.cjs's build-failure test, and annotates
3 of 4 stale golden-fixture citations in
docs/reference/host-integration-capability-matrix.md as superseded
(the 4th is an accurate historical PR narrative, left alone).

* fix(#2724): repair three red CI defects on the golden-fixture cutover

Windows-only provenance false attribution (defect A): the `hooks-built`
provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are
Windows-only installer output (ensureCodexHooksJsonSessionStart /
ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping
the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so
the self-attribution resolved to a path that exists on no platform. Only
windows-latest ever emits the key, so this only failed there. Fixed by
special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated
rule) — a dedicated rule would match zero paths, and therefore report as a
dead rule, on every non-Windows lane of the same totality guard. `sources`
already supported per-match functions; `transforms` is extended to support
the same shape so the attribution can vary by match within one rule.

Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef`
ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but
that script is new in this PR and therefore absent at any base ref that
predates it — every call failed closed with "Cannot find module". Fixed by
running the PR checkout's own generator against the worktree via a new `--dir`
parameter, decoupling "which copy of the script runs" from "which tree it
measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override,
threaded down to `runMinimalInstall`'s new `installScript` override). This is
not just a bootstrap fix: a differential needs ONE measurement schema applied
to both sides, or the two stop being comparable the moment that schema
evolves — running each side's own copy would silently reintroduce that risk.
Verified locally end-to-end against real origin/next: resolves a valid
{version, sha, manifests, sizes} artifact with the correct sha and no leaked
worktree.

Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let
docs-lint evaluate the fragment for the first time; it already passes
(docs/TESTING-SUITES.md and friends already document the removed scripts).

Also fixed while in this file: an eslint no-unused-vars warning surfaced by
the changed lint run (unused `cleanup` import in
tests/emitted-provenance.test.cjs).

Added regression coverage for both A and B: a cross-platform spot-check that
drives the real hooks-built rule against `.cmd` keys directly (not through a
real Windows install), and a real-tree test that drives buildBaselineAtRef
against a base ref verified (via git cat-file) to lack the generator, both
skipping honestly rather than false-passing when their precondition does not
hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test

Two isolated-review findings on PR #2767:

- `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the
  wrapped `hooks/<name>.js` script, asserting a byte-provenance link that
  does not exist — traced against buildCodexHookWindowsShimIR
  (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that
  same file) flows into the .cmd bytes, never its content. Point `sources`
  at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention
  used elsewhere in the table. Since `sources` is checked before
  `transforms` in the differential, the wrong mapping silently excused any
  .cmd byte movement caused by editing the wrapped .js file.

- The `buildBaselineAtRef` regression test skipped unless a resolvable base
  ref still lacked scripts/gen-emitted-baseline.cjs — true only until this
  PR merges, after which every base ref carries the file and the test skips
  forever with zero ongoing coverage. Rebuilt hermetically: synthesize the
  missing-generator condition in-place via git plumbing (a throwaway commit,
  child of HEAD, with just that one file removed from a scratch index),
  never touching the real working tree, HEAD, or index, and never depending
  on ambient history or remotes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path

The runner container mounts the repo at a path owned by a different uid than
the process running the suite, so git's dubious-ownership protection refuses
every git operation there. GitHub Actions never hits this because
actions/checkout registers the workspace as safe automatically; this
runner's container does not.

buildBaselineAtRef is the production build-fallback the sole remaining
emitted gate depends on (resolveBaseline's in-job-build leg), not just a
test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs)
that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's
worktree add/remove/prune, and the hermetic regression test added in the
prior commit — funnels through, plus gen-emitted-baseline.cjs's own
rev-parse (now reusing that same wrapper instead of a second execFileSync,
so the fix has one source of truth). Each call declares -c
safe.directory=<the exact directory it already operates on>, never the *
wildcard.

Audited every other helper on this surface (emitted-diff.cjs,
emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them
shell out to git at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:41:43 -04:00
Tom Boucher
d04592de58 fix(#2758): select the emitted differential wherever golden-parity runs (#2759)
* fix(#2758): select the emitted differential wherever golden-parity runs

Add tests/emitted-provenance.test.cjs and tests/emitted-attribution.test.cjs
to every scripts/ci-test-scope.cjs rule that already selects
tests/golden-install-parity.test.cjs, so the ADR-2719 dual-run differential
travels with the golden on the targeted CI lane instead of being selected by
no rule at all.

Add an independent module-load totality guard (missingRuleTestFiles) that
throws when any rule names a test file absent from disk -- the guard that
would have caught the post-Phase-4-cutover hole. It immediately surfaced
three pre-existing phantom entries left behind by consolidation epic #1969
(bug-3588/bug-10/bug-3683 filenames folded into other suites months ago but
never removed from the rule table); fixed in the same change rather than
deferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2758): trim unused exports and normalize the path require style

Code-review (Standards axis) flagged two judgement-call smells: exporting
classify/isInertCi with no caller (Speculative Generality), and requiring
path with a node: prefix while the file's other core requires do not
(inconsistent style within one file). Both addressed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 10:48:09 -04:00
Tom Boucher
bf0d715733 fix(#2472): cost-balanced test sharding and pinned CI base commit (#2480)
* fix(#2472): weight-aware shard partition

Windows shard 1/3 hit the 20-minute job cap with no failing assertion. Root
cause is the shard layer, not the chunk layer: selectShard partitioned by
sorted ARRAY INDEX (k % n, #1212), which balances file COUNTS and ignores
file COST. On the real unit suite that produced 12.4m / 19.2m / 15.2m — a
1.23x max/ideal ratio leaving the heaviest shard 5% under the cap. Because
assignment keyed off position, inserting one test file re-indexed every file
after it and could tip that shard over; deterministic, so a re-run reproduced
it exactly.

This is NOT the chunk packer (#2456/#2463). That fix works and applies one
level down, WITHIN a shard. The across-shard partition predated it and never
consumed the cost table. Both layers now share one cost model.

selectShard takes an optional weightOf and, when given one, partitions by LPT
(longest-processing-time-first) — the same algorithm packChunks uses. Omitting
it keeps the legacy round-robin byte-identical, so every existing test above
still exercises that path unchanged and callers without timing data lose
nothing. A missing timings table yields uniform weight 1, under which LPT
degenerates to the equal-count split.

Projected on the real suite: 16.4/17.3/13.0 -> 15.6/15.6/15.6 (worst shard
17.3m -> 15.6m).

Tests: a skewed-cost regression (round-robin clusters all four heavy files
onto one shard at 2.98x ideal; LPT does not), back-compat equivalence,
determinism, tie-breaking, order preservation, and two fast-check properties
— the partition is exhaustive and disjoint (getting this wrong silently DROPS
tests from CI, the worst failure mode for a harness), and no shard exceeds
average + heaviest file.

Two assertions were corrected during authoring rather than shipped wrong:
- an initial "LPT within 4/3 of ideal" bound was false. The 4/3 figure is
  relative to the OPTIMAL makespan, not the average, and the two differ when
  item sizes force a pairing. Replaced with Graham's average+max bound, which
  is what is actually provable.
- "weighted is never worse than round-robin" is also false; fast-check
  falsified it with [19316,10190,1,9128,29353,20227] over 2 shards (rr 48670,
  lpt 48671). Round-robin can win by luck on a specific input. Dropped, with
  the counterexample recorded in place so it is not re-asserted later.

Closes #2472

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): rotate tied bins; restore #1212 test block; lazy cost table

Isolated-review findings, all fixed.

HIGH — zero weights collapsed the whole partition onto shard 1. The
lightest-bin scan compared weight only, and adding a zero-weight file leaves
its bin's weight unchanged, so bin 0 stayed tied-minimum forever and every
such file landed on it. Verified: all-zero weights gave shard1=[a..f],
shard2=[], shard3=[] — two of three CI runners idle while one ran everything.
Reachable through safeWeight's own clamp (a NaN/negative/Infinity entry in a
corrupted or hand-edited timings table) and through any genuine 0ms
measurement, so the clamp reproduced the exact failure its comment claimed to
prevent. Ties now break on file COUNT after weight, which rotates. Pinned by
two regression tests (all-zero, and clamped NaN/negative/Infinity) plus a
property over list size x shard count. The live table has no 0ms entries
(min 19ms), so production was not affected — but nothing prevented it.

MEDIUM — the new describe block had swallowed #1212's pre-existing property
test, which is why a test under a "weight-aware" heading never passed a
weigher. That was a bad block boundary in the previous commit, not a bad
test: the #2472 describe was opened before #1212's last test instead of
after. Moved back where it belongs; #1212 is 762-879 and #2472 is 894-1082.

LOW — that relocated property test ran unseeded. Seeded (12120) per the
repo's property-test convention so a failure reproduces. Verified passing
under the new seed.

LOW — hoisting the timings load above the shard block charged a readFileSync
+ JSON.parse to invocations that exit before needing it (empty selection,
--files matching nothing). Now lazily memoized, so neither consumer reads the
table unless it is used and it is still read at most once.

Real-suite projection unchanged at 15.6m / 15.6m / 15.6m.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2472): correct stale round-robin sharding descriptions

The partition is now cost-balanced, so the header block in run-tests.cjs
and the two comments in test.yml describing '--shard' as a round-robin over
sorted file index were actively wrong. Updated to describe LPT over measured
duration, and to state the degenerate case explicitly: with no timing data
every file weighs the same and the partition collapses back to k % n, which
is why the pre-existing #1212 CLI tests still pass unchanged (their nine
synthetic files are absent from the timings table, so all take the identical
median weight).

Remaining 'round-robin' mentions are correct — they describe the unweighted
fallback path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): shard diagnostics, cost-routing E2E test, table validation

Second orthogonal review (operational lens) findings, all fixed.

HIGH — cross-runner partition divergence. Each of the up-to-12 CI jobs runs
its own 'merge base into head' and computes its own partition, so if the
inputs differ between jobs (the file list, or the timings table) two jobs can
place the same file in different shards or in none. Every job stays
internally exhaustive and disjoint, so nothing errors: a test simply never
runs and CI stays green.

The risk class is pre-existing — round-robin diverges identically when the
file set differs between jobs, which is literally this issue's insertion
instability — but weighting adds tests/test-timings.json as a second input
that must match, so it widens the hole. Properly closing it means pinning the
partition inputs per run, a workflow change beyond this fix.

What IS closed here is the silence. Each shard now prints an input
fingerprint over the FULL pre-partition list and the weight assigned to each
file — deliberately not this shard's slice, which would differ by design and
be useless for comparison. All shard jobs of one run must print an identical
sig; a mismatch is direct proof the runners disagreed about the input.
Verified: three independent computations agree, and the sig changes when the
input drifts by one file.

MEDIUM — nothing proved main() actually threads fileWeightOf() into
selectShard. Every pre-existing --shard E2E test uses synthetic filenames
absent from the real table, so all collapse to a uniform median weight, under
which LPT is mathematically identical to k % n — a typo on that one wiring
line would have passed the whole suite. Added an E2E test that injects a
table via RUN_TESTS_TIMINGS_FILE with differing costs, placing the heavy
files at exactly the indices round-robin hands to shard 1, and asserts shard 1
does NOT receive all three. Plus a test that all three shards emit the same
sig.

MEDIUM/LOW — no observability. The diagnostic line now reports files,
weighed count, aggregate weight, and whether the table loaded, so a table
that silently failed to parse shows table=absent/weighed=0 instead of being
indistinguishable from a healthy load. (The reviewer confirmed the advisory
fallback is already live on next: feat-2296-provider-escalation.test.cjs is
missing from the table.)

LOW — typeof [] === 'object', so a hand-edit turning the map into a list was
accepted as a valid table. Now rejected via Array.isArray, falling back to
uniform weight like any other malformed table.

LOW — stale round-robin wording in ci-test-scope.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): pin every CI job to one base commit

Closes the cross-runner divergence at its source instead of only making it
visible.

Each job of a run executes the rebase-check step independently, minutes apart
across a 12-job matrix, and merged the MOVING origin/<branch> ref. If the base
advanced mid-run, different jobs merged different trees. That was survivable
when jobs only had to agree on pass/fail; it is not once they must agree on a
PARTITION. Each shard job computes the whole split and keeps its own slice, so
jobs working from different trees can place a file in two shards or in none —
and every job still looks internally consistent, so nothing errors. A test
silently never runs and CI stays green.

ci-rebase-check.cjs now accepts CI_REBASE_BASE_SHA and pins BOTH the fetch and
the merge to that one commit, so the two can never disagree. test.yml passes
github.event.pull_request.base.sha on all three rebase-check steps; that value
is fixed for the life of a run, so all jobs merge the identical base.

This also closes the PRE-EXISTING half of the divergence. Round-robin had the
same exposure whenever the test-file set differed between jobs — that is this
issue's insertion instability — so the pin fixes the older hole too, not just
the timings-table input weighting added.

Only a full 40-hex sha is accepted; empty (push/workflow_dispatch), malformed,
or injected values fall back to the branch ref rather than handing an arbitrary
string to git fetch as a refspec. resolveBaseRefs is extracted pure and
exported, and runMain is guarded behind require.main === module, so the pin
contract is testable without spawning git.

Tests (tests/ci-test-scope.test.cjs): every rebase-check step must carry the
pin; a valid sha pins both refs; absence falls back correctly; and five hostile
values — short sha, uppercase, --upload-pack= injection, ref expression, empty
— are each rejected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 09:15:59 -04:00
Tom Boucher
89b1bef881 refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267.
2026-07-14 19:22:13 -04:00
Tom Boucher
241b08fa18 ci(#1975): exclude release-tarball-smoke from scoped lane (fixes Windows chunk timeout)
This consolidation PR's breadth (28 changed test files) exposed a scoped-test-lane
capacity limit: ci-test-scope pulls the 3–6 min release-tarball-smoke.install.test.cjs
(npm pack + npm install -g, 10MB/1499 files) into the targeted+windows lane whenever
install files change AND when it is itself a changed file — bundling it with the other
27 files overran the 600s per-chunk timeout on windows-latest-24 (deterministic).

release-tarball-smoke has its OWN dedicated workflow (.github/workflows/install-smoke.yml,
triggered on the production install paths), so its scoped-lane run is redundant. Add a
SCOPED_LANE_EXCLUDE guard that drops it from both targeted_tests and windows_tests however
it entered (matched rule OR changed-file), and remove it from the install rule's tests list.
Update the ci-test-scope.test.cjs assertion accordingly. No coverage lost.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
b2ed7940c9 test(#1975): scope GSD_TEST_MODE off in real-install folds; fix ci-test-scope fixture
gsd-test surfaced 21 failures:
- 19: folded real-install suites (bug-1834 .sh hooks, enh-2380 --skills-root, fix-1521
  install stamping, bug-2136 .sh hook version) spawn install.js and assert side effects,
  but their host suites (install-minimal-hooks/install.test/managed-hooks) set
  GSD_TEST_MODE=1 at collection time — the install child inherited it and suppressed
  the writes. Clear GSD_TEST_MODE in each of those blocks (before/after; standalone had
  it unset), so the child performs a real install.
- 2: ci-test-scope A1 used deleted tests/bug-1974-context-exhaustion-record.test.cjs as a
  fixture; scopeFor filters nonexistent paths, so it fell back to ['unit']. Repointed to
  its consolidation destination tests/perf-317-context-monitor-fs.test.cjs (an existing test).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
6d072435d0 test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their
canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks,
read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping
the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim
block-scoped describe wrappers; 427 subtests conserved 1:1.

Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT
value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools
stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec,
so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard.

Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6,
docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456;
prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md
(EN + ja/ko/pt/zh) and ADR-0002. lint:ci green.

Part of epic #1969. Closes #1975.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:22:11 -04:00
Tom Boucher
9d52043f50 feat(#1733): normalize-path-in-content production AST rule + fix Windows agent-skills content leak (Phase 5) (#1736)
* feat(#1733): normalize-path-in-content production AST rule (Phase 5)

ADR-1703 Phase 5 — the first production-code rule. local/normalize-path-in-content
(src/**/*.cts, @typescript-eslint/parser): flags a path-returning fn result
(path.basename excluded — returns a separator-less filename) interpolated into an
@-reference / config-dir markdown body without .replace(/\\/g,'/') normalization,
per RULESET.CONTENT-PATH-NORMALIZATION / DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT.

Build-and-assess found the canonical defect site (computePathPrefix) already
compliant and only 1 src/ hit — a false positive (path.basename in a status
message) — eliminated by narrowing (exclude basename; require a real @-ref/
config-dir marker, not bare .md). 0 src/ violations: clean forward-prevention.

The out-of-band disable-ban now scans src/**/*.cts too (typescript-estree) so the
production rule also cannot be eslint-disabled. Registered (error) + PROTECTED_RULES;
CONTEXT.md predicates + how-to doc updated.

- RuleTester suite (26 cases)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1733): add changeset for Windows agent-skills path-leak fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: harden mutation-matrix.cjs stdin read against EAGAIN on non-blocking pipe

scripts/mutation-matrix.cjs read piped stdin via readFileSync(process.stdin.fd).
On macOS libuv marks the stdin pipe fd non-blocking, so a synchronous read can
throw EAGAIN before the writer fills the pipe — intermittently, under heavy CI
shard load — aborting the script (status 2) and flaking mutation-matrix-ratchet.
Replace with readStdinSync(): an fs.readSync loop that retries on EAGAIN (1ms
synchronous Atomics.wait yield), stops on 0-byte/EOF, and rethrows other errors.
Deterministic regression test injects EAGAIN via an fs.readSync monkeypatch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* ci: re-run golden-install-parity on src/lib + installer changes (close drift guard)

golden-install-parity hashes every installed bin/lib/*.cjs per runtime, so it
must re-run whenever the built lib could change. ci-test-scope selected it for
neither src/** nor installer changes, so a source-only edit (e.g. #1691's
milestone.cts/roadmap.cts) recompiled bin/lib and silently drifted the golden
fixtures past the scoped lane. Add golden-install-parity.test.cjs to both the
'TS runtime sources' and 'installer and package layout' selection rules, with
behavioral regression tests for each.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: review-bot <review-bot@gsd>
2026-06-25 21:49:22 -04:00
Tom Boucher
f08b177215 feat(#1726): G1-G6 portability AST rules; fix all offenders; delete the ratchet (Phase 4) (#1731)
Phase 4 of epic #1702. Closes #1726.
2026-06-25 18:24:39 -04:00
Tom Boucher
22f56f4431 ci(#1212): shard windows full-test lane to remove timeout cliff (#1222)
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.

Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.

- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
  round-robin partition (fileIndex % n === i-1) over the SORTED selected file
  list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
  pure no-op. The 28K Windows argv chunking is preserved within each shard. A
  legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
  sharding still hits the discovery hard error. Composes with --suite and is
  order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
  cross-product (explicit include rows — a base shard dim does not cross-product
  with include legs, and a nested matrix.leg.os is unresolvable by the H1
  shell-policy linter). Unit suite runs sharded; integration/security run once
  per leg (shard 1). The Required tests fan-in is unchanged: it already needs
  test-full and checks the matrix-aggregate result, so a failed/cancelled shard
  fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
  balance, determinism, boundaries, fast-check property) + parseShardArg
  validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
  guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
  denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.

Closes #1212

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 12:25:29 -04:00
Colin
a647053dcf ci(scope): narrow #494 invariant — changed tests join windows lane, not full matrix
full_matrix fired on 15/15 sampled PRs because any tests/** change forced it,
costing ~25 runner-minutes each. A changed test file now always joins the
scoped windows lane (covering the #482 OS-specific failure class per-file)
and still runs on ubuntu 22/24 via targeted_tests; the residual macOS /
windows-node-22 cross-product is covered on every push to next.

Also narrows WINDOWS_HINTS from 6 substrings (102/633 files, a ~10-minute
scoped lane) to windows/win32/shell/path — the dropped hints (workflow,
install, hook) are either platform-independent lint tests or already covered
by fullMatrix rules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
606363c416 chore(#846): remove unused PR-size labeler (size/S–XL) workflow (#848)
The PR Gate workflow's only job, size-check, labeled every PR with
size/S–size/XL based on lines changed. Those labels aren't used in any
review, triage, or automation flow, so the workflow was pure noise.

- Delete .github/workflows/pr-gate.yml
- Drop size-check from required status checks in both rulesets so PRs
  don't block forever on a check that never reports
- Remove pr-gate.yml from INERT_WORKFLOWS (ci-test-scope.cjs) and the
  knownInert list (ci-test-scope.test.cjs)
- Remove "PR Gate / size-check" from setup-branch-protection.sh

Closes #846

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 22:03:48 -04:00
Tom Boucher
988024c1a3 fix(#837): three-dot diff in ci-test-scope so docs-only PRs skip the heavy matrix (#841)
CI test-scope detection diffed changed files with a two-dot
`git diff --name-only base head`, where base is the moving tip of `next`.
A PR branch cut from a slightly older `next` surfaced every product file
`next` had gained since the merge-base, flipping product_changed/full_matrix
and running the full Windows/macOS matrix + coverage on docs-only PRs.

Switch to a three-dot `git diff --name-only base...head` (vs the merge-base),
matching GitHub's PR "Files changed" semantics. Add a regression test that
builds a stale-base topology, plus a guard test pinning `fetch-depth: 0` on
the `changes` job (required for the merge-base to be locally available).

Closes #837

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:28:45 -04:00
Tom Boucher
571d7b5a1c feat(#764): skip cross-platform test matrix for docs-only and inert-CI PRs (#798)
test.yml had no paths filter and the ci-test-scope classifier treated docs/
and every .github/workflows/* as code_changed, so documentation edits and
product-irrelevant automation tweaks still spun up the full Linux/Windows/macOS
matrix. Narrow the heavy matrix to changes that can actually affect the product
or the test pipeline.

- ci-test-scope.cjs: drop docs/ from code_changed (docs-only -> full skip; the
  required-tests fan-in still reports green). Add src/ to code_changed (it was
  missing -> a source-only PR previously skipped all tests). Add INERT_WORKFLOWS
  allowlist + isInertCi() + an "inert CI" rule, and a product_changed output that
  gates the heavy test/coverage jobs. Fail-safe: any workflow not on the inert
  allowlist defaults to the full matrix. A module-load assertion throws if a
  PROTECTED_WORKFLOWS entry (test/install-smoke/mutation/security-scan/release)
  is ever added to the inert set, so a weakening edit fails CI loudly.
- test.yml: keep the static 3-lane matrix (so the H1 shell-policy linter can
  still statically verify the Windows lane), gate test/coverage on
  product_changed, add a lightweight ubuntu-only test-inert job, and branch the
  required-tests fan-in on product_changed.
- docs-required.yml: run docs-parity-live-registry (gated on docs/ changes) so
  pure-docs PRs still catch live-registry drift without the matrix.
- tests: cover docs-only, inert-only, src/, pipeline, unknown-workflow fail-safe,
  mixed escalation, the code_changed=false -> no-lanes invariant, and protected-
  workflow tamper-evidence.

Closes #764

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 11:46:48 -04:00
Tom Boucher
cdd78bd2aa fix(#670): self-healing recovery for installer-migration checksum drift (#675)
Editing the body of an already-released installer migration drifts its computed
checksum (it hashes plan.toString()). The integrity guard then hard-aborted
every prior install on upgrade with "applied migration checksum changed" — a
100% reproducible blocker (v1.3.0, all platforms).

Already-applied migrations are filtered out of `pending` and never re-run, so
a drifted checksum is functionally inert. ADR-0008 anticipates checksum-mismatch
state as something the install-state layer must handle gracefully (plan -> apply
-> recover/report), not abort on.

This supersedes the published-checksum allowlist merged in #674 (per-release
maintenance debt — every historical checksum hand-pinned, still throws for any
unregistered value) with a general, self-healing recovery:

- Replace the throwing guard with non-fatal `collectAppliedChecksumDrift`,
  surfaced on `plan.checksumDrift`.
- Reconcile drifted stored checksums durably on the next state write
  (`reconcileDriftedChecksums`), idempotently (no perpetual writes).
- Relocate the "shipped migration bodies are immutable" rule to a CI baseline
  test that locks every shipped migration's checksum and fails on body drift —
  where #615 should have been caught, instead of blocking users.

Removes #674's legacyChecksums field, per-migration checksum pins,
published-checksums.json fixture, and compat test.

Fixes #670

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 13:44:05 -04:00
Colin
b8e15c7a98 fix: accept published installer migration checksums 2026-06-04 12:30:13 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
a7ed001b27 fix(#494): ci-test-scope selects checks a diff can break (tests->full matrix, docs->docs-parity) (#495)
classify() under-approximated breakable checks, so scoped PRs skipped the
check their diff would break and regressions reached next (#484 docs-parity,
#482 windows-22 EBUSY). Fail-safe widen: any tests/** change forces
full_matrix (OS-specific test failures); any docs/**, commands/**, agents/**
change marks code_changed and selects docs-parity-live-registry (its runtime
inputs). Updated the docs-only test that asserted the old buggy contract.

Fixes #494

Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-05-29 19:03:22 -04:00
Tom Boucher
f5f51b5b49 fix(#408): align ci-test-scope smoke handling with #395 changeset (drop unconditional injection; unit fallback) (#420)
- Remove `DEFAULT_SMOKE_TESTS` and `WINDOWS_SMOKE_TESTS` constants (now dead after the unconditional injection block is dropped)
- Drop the `addAll(targeted, DEFAULT_SMOKE_TESTS)` / `addAll(windows, WINDOWS_SMOKE_TESTS)` block from the `codeChanged` branch
- When `codeChanged && targetedTests.length === 0`, push `'unit'` as the fallback suite token
- Two new regression tests in `tests/ci-test-scope.test.cjs` covering the no-injection and unit-fallback contracts (bug #408)
2026-05-27 22:59:26 -04:00
Tom Boucher
9dd2c87c6c ci: supersede #367 with combined tiered PR pipeline redesign (#369)
* ci: streamline PR pipeline gates

* test: annotate ci-test-scope stderr assertions for lint rule

---------

Co-authored-by: Colin <colin@solvely.net>
2026-05-26 19:11:58 -04:00