Fold 51 issue-named CLI black-box + scripts-tooling regression files into their
canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks,
read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping
the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim
block-scoped describe wrappers; 427 subtests conserved 1:1.
Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT
value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools
stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec,
so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard.
Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6,
docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456;
prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md
(EN + ja/ko/pt/zh) and ADR-0002. lint:ci green.
Part of epic #1969. Closes#1975.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.
Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.
- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
round-robin partition (fileIndex % n === i-1) over the SORTED selected file
list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
pure no-op. The 28K Windows argv chunking is preserved within each shard. A
legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
sharding still hits the discovery hard error. Composes with --suite and is
order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
cross-product (explicit include rows — a base shard dim does not cross-product
with include legs, and a nested matrix.leg.os is unresolvable by the H1
shell-policy linter). Unit suite runs sharded; integration/security run once
per leg (shard 1). The Required tests fan-in is unchanged: it already needs
test-full and checks the matrix-aggregate result, so a failed/cancelled shard
fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
balance, determinism, boundaries, fast-check property) + parseShardArg
validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.
Closes#1212
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: address adversarial-review findings
Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
capability-registry.cjs in place (concurrency hazard) — uses in-memory
checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
/* */ too, matching no-source-grep) so a block comment can't bypass it;
one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: address code-review findings (subdir discovery, rule + test gaps)
xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{} equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)
The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: reconcile allow-test-rule allowlist after rebase onto next
Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The `full test (windows-latest, 22)` job intermittently got CANCELLED at its
20m wall-clock cap with no failed test step — a false-negative gate (recurrence
of #869). Root cause: a unit test leaves an open event-loop handle, so the
chunk's `node --test` child hangs ~150s on Windows after its last test prints;
two such stalls push the already-~13m job past 20m.
Fix (defense in depth):
- run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0)
so the runner exits once all tests finish regardless of lingering handles —
the durable backstop. Account for the flag in the argv-length ceiling.
- run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env
RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the
chunk's files, so a hung chunk can never silently eat the job budget.
- perf-316 test: terminate both Worker threads on all paths (afterEach +
finally) so they cannot outlive the test.
- locking-bugs test: kill spawned children in a finally that wraps the whole
spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier
timeout no longer leaks live child processes.
- Refresh the stale synckit comment (synckit/SDK bridge was removed).
Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout
and fails with a clear message; force-exit lets a chunk with a leaked handle
exit cleanly.
Closes#1051
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
extractCurrentMilestone reads STATE.md via planningDir(cwd), which is
workstream-aware (honours GSD_PROJECT/GSD_WORKSTREAM). The fixtures write
STATE.md to the plain <tmp>/.planning/STATE.md, so a developer shell inside a
GSD workstream (GSD_WORKSTREAM exported) redirected the read to a non-existent
workstream subdir -> version=null -> closed milestone sections leaked into the
slice and assertions failed. Clean CI/Docker env never hit it. Not a Node-26
regex bug; reproduces identically on any Node with GSD_WORKSTREAM set.
- scripts/run-tests.cjs: strip GSD_PROJECT/GSD_WORKSTREAM before spawning test
children so the local runner env matches clean CI/Docker.
- tests/roadmap-phase-fallback.test.cjs: file-level beforeEach/afterEach
save/delete/restore of both vars; new regression test pinning workstream-aware
STATE.md resolution.
- tests/run-tests-harness.test.cjs: guard asserting the runner strips both vars
(so removing the deletion fails clean CI).
Closes#872
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Windows CreateProcess caps lpCommandLine at 32,767 chars. The original
`execFileSync(node, ['--test', ...546 paths])` exceeded that on every
Windows runner and exited within ~70ms with no test output. Linux/macOS
allow ~2 MB ARG_MAX so the same call worked there.
`scripts/run-tests.cjs` now splits selected files into chunks that keep
each spawn's argv under 28,000 chars (operator-overridable via
RUN_TESTS_MAX_CMDLINE_CHARS), runs them sequentially, and reports the
first non-zero exit. Cross-platform regression test forces chunking with
a low ceiling and asserts the `run-tests: chunk N/M …` stderr marker.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
scripts/run-tests.cjs gains `--suite <name>` filtering using a filename
suffix convention (`*.security.test.cjs`, `*.integration.test.cjs`, …).
Files with no marker are `unit` (the default fast lane); files with a
marker land in the matching suite. No `--suite` flag preserves the prior
behavior of running every test (backcompat for `npm test` and
`npm run test:coverage`).
New package scripts wire the suites to stable entrypoints:
test:unit, test:integration, test:install, test:security, test:slow,
test:coverage:unit, test:coverage:all. Unknown suite → exit 2 with the
list of valid suites; empty suite → exit 0 with a stderr notice so empty
lanes (e.g. `security` before adversarial tests land) don't gate CI.
CI matrix grows from `ubuntu × {22,24}` + a single macOS lane to
`{ubuntu, macos, windows} × {22, 24, 26}`. `fail-fast: false` so one
lane failure doesn't cancel siblings. Node 26 is `continue-on-error`
until actions/setup-node stabilises that image. PR CI runs unit +
integration + security on every cell; `install` and `slow` only on
`main` push. A dedicated `coverage` job runs `test:coverage:unit` on
ubuntu/Node 24 and uploads the report.
Grouping policy lives in docs/TESTING-SUITES.md with a pointer from
CONTRIBUTING.md. New harness test covers arg parsing, filter selection,
empty-suite behavior, and failure propagation.
Closes#3597.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>