Files
msd-core/docs/TESTING-SUITES.md
Tom Boucher 33fd203ccd test(#2966): loop QA walk — drive real scenarios across all five loop steps (#2976)
* test(#2966): loop QA walk — drive real scenarios across all five loop steps

Adds a headless walk that carries accumulating project state across
discuss -> plan -> execute -> verify -> ship against one temp project,
layered over the existing tests/helpers.cjs runGsdTools substrate.

Findings carry severity. A violation breaks a stated contract and fails
the build; a smell is legal under today's implementation but structurally
questionable, is recorded, and never reddens CI. Without that split an
oracle set derived from current behavior can only ever confirm current
behavior -- the harness could not say "this works and is still wrong".

The end-to-end test asserts the walk produces at least one smell: a QA
harness that reports nothing on a first run against a real engine is far
more likely mis-specified than the engine is perfect. It deliberately does
not pin smell ids or counts, which would re-freeze current behavior.

First run against the real engine: 0 violations, 3 smell classes --
init returns agents_dir outside the project tree; smart-entry emits prose
unconditionally so routing cannot be asserted; state-snapshot reports a
missing STATE.md through a payload key with exit 0.

Also fixes tests/fixtures/index.cjs: createFixture with git:true and
planning:false staged nothing, so the commit failed with "nothing to
commit". That combination was unreachable until greenfield needed it.

Extends RULESET.TESTS.feedback-loop-convergence from estimation to the
loop itself. Design lock: docs/adr/2966-loop-qa-walk.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): wire fault injection, make perturbations discriminating

Independent review found tests/qa/mutations.cjs entirely unwired: 462
lines exercised only by their own unit tests, with no mutation hook in
the scenario DSL and no scenario applying one, while the module header
and the ADR described fault injection in the present tense. Dead code
documented as live.

Adds a `mutate` step field, three perturbation scenarios, and a wiring
detector: a self-test scenario whose expectations are known-false and
which MUST fail. The previous anti-vacuity check asserted only that the
walk produced a smell, which passes on well-known engine behavior
regardless of whether the harness wiring works.

First perturbation attempt produced zero signal -- progress does not
structurally parse ROADMAP.md, so a corrupted roadmap sailed through. A
perturbation that cannot fail is the same defect in a new costume.
Probes now target roadmap get-phase, and each mutated step runs a clean
baseline first so `mutationObserved` records whether the corruption
changed anything at all.

Also clears four review findings: classify() returned PROSE for exit-0
with empty stdout; `warnings` was structurally unpopulatable on the
success path (execFileSync discards it) and is now documented as
error-path-only; read-only-idempotence passed vacuously when asked to
check idempotence without the data to check it; the ADR miscounted the
oracles.

Discrimination matrix across 8 mutations x 6 commands: bom,
duplicate-phase-id and escaped-pipes are absorbed silently by every
probed surface, and progress / smart-entry / roadmap validate never
reacted to any mutation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): add path-containment guard for scenario-supplied targets

Security review found scenario-supplied paths joined to the temp project
with no containment check. step.mutate.target and agent.write keys were
validated only as non-empty strings, so a target of ../../../../etc/hosts
reached fs.unlinkSync / fs.writeFileSync / fs.symlinkSync outside the
project. The symlink mutation was worst: it read the traversed file, wrote
a sibling copy, deleted the original and symlinked it back.

Not exploitable today -- all shipped scenarios target .planning/ROADMAP.md
and scenarios are repo-committed, not runtime input. Fixed anyway: it is a
live primitive any future scenario or copied helper can reach.

Adds tests/qa/paths.cjs with resolveWithin(): rejects absolute paths, NUL
bytes and empty input, normalizes separators unconditionally, and requires
containment by path segment so a sibling like <base>-evil is not treated as
inside. Non-existent targets resolve via nearest existing ancestor rather
than falling back to a lexical compare. Scenario load now rejects traversing
or absolute targets up front.

oracles.cjs previously carried its own copy of the containment logic; both
now share paths.cjs, since a duplicated containment check is exactly the
divergence class this repo calls out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): complete trajectory corpus, report emission, boundary-aware oracle

Adds the remaining trajectories and drives all 11 mutations end-to-end.
20 scenarios, 72 steps, 0 violations, 25 smells.

Adds qa-report.json with per-step verdicts and a copy-pasteable repro
command, plus --keep / GSD_QA_KEEP=1 to preserve a failing tree. A repro
line for a tree that was not preserved is marked NOT RUNNABLE rather than
emitting a command pointing at a deleted directory.

monotonic-progress is now boundary-aware. Two scenarios had been trimmed
to stop the oracle complaining at a milestone rollover, which destroys the
signal the trajectory exists to produce. Evidence: counters legitimately
reset to zero at milestone complete, but the payload milestone_version
lags until a new ROADMAP.md is written. So the oracle now scopes by
milestone plus workstream, keeps a same-scope decrease as a violation, and
records a boundary crossing as a smell. Both scenarios walk the real
boundary again.

Standards review fixes: oracle findings now carry a structured subject so
tests assert on typed fields instead of substring-matching the free-form
detail string, resolveWithin throws a typed EPATHESCAPE error, and the
absolute-path predicate scenario.cjs had re-implemented now comes from
paths.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): fix silently-vacuous fixtures and guard the class

Every fixture carried its #2371 provenance comment BEFORE the frontmatter
block, and extractFrontmatter returns {} when anything precedes the opening
---. So every scenario reading status/phase/name was operating on an empty
object and reporting green. Nine fixtures repositioned; the comment stays,
it just moves below the closing ---.

Both UAT fixtures lacked a parser-recognized result block, so
evaluateUatPassed saw checks.length===0 and could never return passed:true.
The uat-fail-then-remediate scenario could not have proven a remediation.
Its expect block only inspected blockers, which is empty before AND after,
which is why the corpus never noticed. Both fixtures now carry real result
blocks and the scenario asserts passed and no_uat_artifacts on each side of
the flip.

The actual deliverable is the guard: a fixture-integrity block asserting
every fixture with a frontmatter shape parses to a non-empty object, that
every fixture carries its provenance marker, and that the two UAT fixtures
produce opposite verdicts through the real evaluateUatPassed. The first
guard written required --- at byte 0, which would never have fired on the
regression it exists to prevent; it was rewritten and proven by deliberately
re-breaking a fixture.

No engine defect here. no_uat_artifacts means no parsed check items, not no
UAT files, and it was reporting correctly on fixtures that had none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): make the walk report — smell ratchet, baseline, CI job

The harness computed smells into a gitignored qa-report.json that nothing
read. In CI it surfaced nothing at all: violations failed the build, but the
half of the tool that says "this works and is still wrong" was inert. A QA
tool nobody hears is decoration.

Adds a ratchet on the same idiom this repo already uses three times over
(the regression-test-name allowlist, the emitted-drift acks, the size
baseline): a committed smell-baseline.json, per-PR acknowledgment fragments
under tests/qa/smell-acks/, and a ratchet script wired into CI.

The design invariant is preserved exactly. A smell still never fails a build
on its own merits. What fails is an UNACKNOWLEDGED NEW smell -- the absence
of a decision -- leaving an author two honest exits: fix it, or record a
fragment with a real reason. An empty reason is rejected. The baseline is
shrink-only, so a fixed smell must prune its entry. Violations remain
unacknowledgeable.

Fingerprints are composed only from stable fields (oracle id, scenario,
argv, subject discriminator) -- never temp paths, timestamps or counts.
Verified byte-identical across two runs in separate temp dirs; an unstable
fingerprint would have false-positived every CI run.

CI gains a qa-loop-walk job that runs the suite and the ratchet, uploads the
report with `if: always()` (it matters most when it failed), and renders a
summary a reviewer reads without downloading anything.

Also fixes the report runner invoking main() unconditionally on require, so
importing it double-ran every scenario and clobbered its own output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): every smell terminates in a defect or a fixed detector

The baseline accepted a smell with a free-text reason. That is a mechanism
for designing smells in -- an allowlist nobody revisits. The harness is
brand new, so nothing it found is inherited legacy; every finding is a
FIRST finding. Each must now terminate in exactly one of two states:

  REAL           -> an assigned defect, entry carries the issue number
  FALSE POSITIVE -> the detector is wrong and gets fixed, never baselined

There is no third "accepted with a good explanation" state, so the ratchet
now requires a positive-integer `issue` on every entry. A reason may remain
as a human note but can never substitute. `--update` refuses to invent
issue numbers: a new smell is written with `issue: null` and a TODO, and
the next plain run rejects it, forcing triage rather than accumulation.

Working the 21 existing entries through that rule found 16 were my own
detectors being wrong:

value-hygiene (10) flagged $.agents_dir, a field whose entire contract is
to point at the install tree outside any project. Fixed with a leaf-key
allowlist of contractually-external fields, verified as the only such key
in the init payload. Genuinely unexpected out-of-project paths still smell.

monotonic-progress (6) fired on legitimate boundary crossings -- milestone
v1.0 to v2.0, workstream beta to alpha -- and on one payload carrying no
scope fields at all, where a change cannot even be known. Scope changes now
reset silently and scope-less observations are skipped. The same-scope
decrease remains a violation; that is the real invariant and is regression-
guarded.

The five survivors are real and now tracked: soft-error-exit-zero (#2980),
untyped-success (#2979). Baseline 25 -> 5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): keep the ratchet out of the tarball, unpin the qa CI job

The remote matrix returned failed -- 3 unique failures, identical on
node22 and node24, both root causes in this branch's own diff.

The ratchet lives under scripts/, which ships in the npm tarball, and it
requires three modules under tests/, which does not. In a published
install it is MODULE_NOT_FOUND at load. This is exactly the class the
#2858 guard was added to catch, and it caught it. Fixed the way #2858
fixed the same shape for its own repo-only CI script: a targeted files[]
negation, so the ratchet stays in the repo for CI and out of the tarball.
Not solved by moving or inlining the required modules -- the ratchet must
keep using the same code the harness uses, or the two drift.

Verified both directions: the script is no longer in the pack list, and
build-hooks.js, fix-slash-commands.cjs and gen-capability-registry.cjs are
all still shipped. Over-negating there would have broken installs, since
bin/install.js requires them.

The qa-loop-walk job also carried CI_REBASE_BASE_SHA copied from a
neighbouring job without the paired GSD_EMITTED_BASE, which the #2854
invariant forbids by name: diverging them makes the differential compare a
tree against a baseline from a different commit. The job runs only the qa
suite and the ratchet and invokes no emitted-attribution test, so it needs
no rebase-pinned base at all -- the step was removed rather than paired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): stop monotonic-progress going blind on scope-less payloads

The full remote suite caught a false NEGATIVE I introduced while fixing a
false positive. Silencing the boundary-crossing noise had made the oracle
skip ANY observation lacking milestone fields -- so a minimal payload like
{total_summaries: n} produced no violation at all, and the oracle stopped
catching the exact defect it exists to catch. For a QA tool that is
strictly worse than the noise it replaced.

Scope is only indeterminate when the two observations DISAGREE about
having it:

  both scoped, same scope, decrease -> VIOLATION
  both scoped, different scope      -> reset silently
  NEITHER scoped, decrease          -> VIOLATION   (the regression)
  mixed                             -> skip the comparison

Implementing the mixed case surfaced a second blind spot: advancing the
reference point on a skipped pair lets a scope-less observation sitting
between two same-scope ones mask a real decrease. Mixed now leaves the
reference untouched. All four branches carry explicit coverage; only one
did before, which is why this shipped.

The self-test that failed was right and the code was wrong, so the code
moved. Corpus behavior is unchanged: still 5 smells, 0 new, 0 stale, 0
violations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:13:28 -04:00

26 KiB
Raw Blame History

Testing Suites

This project's tests/ directory uses filename suffix markers to group tests into named suites. The harness scripts/run-tests.cjs filters by suite when given --suite <name>. Without a flag it runs every *.test.cjs file (the historical default — unchanged).

Tracked by issue #3597.

Suites

Suite Filename pattern What goes here
unit *.test.cjs (no other marker) Default fast lane. Pure logic, no network, no external processes beyond gsd-tools. Most tests live here.
integration *.integration.test.cjs Cross-module flows: full installer end-to-end, multi-tool orchestration, anything that crosses two or more bin entry points.
install *.install.test.cjs Tests that perform a real install/uninstall against a sandbox project. Slower; PR CI skips these on PRs and runs them on main push only.
security *.security.test.cjs Adversarial input, prompt-injection guards, fixture-driven hostile-payload sweeps.
slow *.slow.test.cjs Anything that routinely takes >5s wall-clock or holds significant memory.
qa *.qa.test.cjs End-to-end walks that drive the real gsd-tools binary across multiple loop steps against one accumulating temp project, with invariant oracles after every step. Slower than unit; excluded from the fast lane.
all (any) Explicit alias for "no filter". Equivalent to running with no --suite flag.

How to place a new test

  1. Pick the most specific bucket above.
  2. Name the file with the matching suffix: tests/<feature>.<suite>.test.cjs.
  3. If unsure, leave the suffix off — the file lands in unit, the default fast lane.

Examples:

  • tests/agent-frontmatter.test.cjs — unit
  • tests/prompt-injection-guards.security.test.cjs — security
  • tests/installer-end-to-end.install.test.cjs — install
  • tests/sdk-mutation-stress.slow.test.cjs — slow
  • tests/loop-walk.qa.test.cjs — qa

The suite-suffix convention was chosen over a directory layout (tests/security/) so the 545+ existing test files don't need to move. Existing files all classify as unit until someone explicitly retags them.

Regression tests

Do not create new top-level tests/bug-NNNN-*.test.cjs files. Add the regression case to the owning module's main test file instead (e.g. a describe('regressions') block in tests/<module>.test.cjs).

node --test spawns one child process per FILE, so file count — not test count — is the unit of CI overhead, and it is worst on Windows lanes where every spawn is Defender-scanned. The 2026-06 CI audit found 244 one-off bug-* files (~38% of the suite). That population is grandfathered in scripts/lint-regression-test-names.allowlist.json and enforced by an identity ratchet (npm run lint:regression-names, part of npm run lint:ci):

  • A new bug-* file fails CI — fold it into the owning module's file.
  • Deleting/consolidating a grandfathered file requires pruning its allowlist entry, so the baseline only ever shrinks.
  • Inherited drift (the failure names files your PR didn't add — e.g. the base branch merged bug-* files without feeding the allowlist, or you rebased and carried a pre-rebase allowlist): run node scripts/lint-regression-test-names.cjs --update and commit the regenerated allowlist. Snapshot artifacts like this allowlist (and docs/INVENTORY.md) must be regenerated after rebasing, never carried through a rebase.

The ratchet deliberately covers only bug-*. Files named feat-NNNN-* / enh-NNNN-* are feature test files — one (or one per suite) per feature is the sanctioned layout (see the #443 strategy below), not a one-off regression pattern. If issue-*/perf-* one-offs start accumulating the same way bug-* did, extend the ratchet's regex and regenerate the allowlist.

Workflow & agent size budget

Tracked by issue #1074. Bytes (not lines) per #717; LF-normalized per #683.

Workflow files (gsd-core/workflows/*.md) and agent files (agents/gsd-*.md) both ship in the installed runtime and are loaded into context — workflows on every command, agents on every subagent dispatch — so their byte size is a real cost. Two sibling guards (tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs) keep that cost from creeping up invisibly, sharing one byte-counter (measureMdFiles). Growth is caught by two independent layers:

Layer What it does Where
Differential attribution size ratchet (primary, #2724 / ADR-2719 §4) The same computed-attribution check that replaced the golden-install-parity fixtures also reports growth in any gsd-core/workflows/*.md or agents/gsd-*.md file, with the exact byte delta, comparing PR HEAD against next. Unacknowledged growth is a hard failure; shrinkage needs no acknowledgment. No committed snapshot — nothing to regenerate by hand. tests/emitted-attribution.test.cjs (real-tree test) via tests/helpers/emitted-diff.cjs
Loose tier hard caps (backstop) Absolute outer red lines per tier — workflows: XL ≤ 98304, LARGE ≤ 61440, DEFAULT ≤ 40960 bytes; agents: XL ≤ 57344, LARGE ≤ 49152, DEFAULT ≤ 24576 bytes. A cap is never raised when a file approaches it: crossing it means extract, not bump. Independent of the ratchet above — unaffected by #2724. XL/LARGE/DEFAULT_CAP in each guard file

discuss-phase.md additionally has a thin-dispatcher target of < 32000 bytes (the discuss-phase progressive-disclosure split, #717). A net-new agent is DEFAULT-tier and already bounded by the DEFAULT cap — no separate new-agent cap is needed. (This tier-cap machinery is distinct from the separate 45 KB-char extraction-evidence threshold on gsd-planner enforced by tests/planner-decomposition.test.cjs — that one proves mode sections were extracted; this one bounds total agent bytes.)

How-to: a workflow or agent grew and CI is red

The differential attribution check reports the file and the byte delta. To resolve:

  1. Justify the growth in your PR (a sentence in the description is enough) — the acknowledgment entry (below) is the review record that the larger size was a deliberate, seen decision, not silent drift.
  2. Add an acknowledgment entry in tests/emitted-drift-ack.json naming the file and the reason, per CONTEXT.md's ### Emitted Artifact Provenance entry. This is deliberately a committed file, not a flag — the entry appears in your PR diff, so touching it is the visible signal.
  3. Or shrink it instead of acknowledging. Prefer extraction when the growth is incidental: for a workflow, move per-mode bodies to workflows/<name>/modes/, templates to workflows/<name>/templates/, and shared prose to gsd-core/references/; for an agent, lift shared boilerplate into gsd-core/references/ and @-reference it — then load it LAZILY. Do not convert them to eager @-required_reading includes: that shrinks the file's bytes without shrinking loaded context, so it games the guard while making the real cost worse. See workflows/discuss-phase/ for the progressive-disclosure pattern.

If a hard cap (not the ratchet) is what failed, an acknowledgment will not help — that is the signal to extract, per step 3.

Reference

Artifact Role
scripts/workflow-size.cjs Single source of truth — LF-normalized byte counter (lfByteCount) + generic measureMdFiles(dir, predicate) (backs both workflows and agents) + workflow enumeration (listWorkflowStems, measureWorkflows). Imported by both guards and by tests/helpers/emitted-runtime.cjs's currentSizes() so they can never measure differently.
tests/emitted-attribution.test.cjs + tests/helpers/emitted-diff.cjs The differential attribution check and its size ratchet (ADR-2719). The sole mechanism for both emitted-content propagation AND per-file size growth as of #2724.
tests/emitted-drift-ack.json Committed acknowledgment file for unattributable emitted-content ripples and for size growth. Absent = no acks; its presence is the alarm.
npm run regen:derived Runs every remaining generator in dependency order (build → registry → ADR index → capability matrix → inventory manifest → manifest versions → tests/fixtures/install-tree/*.json).
tests/workflow-size-budget.test.cjs The workflow tier hard-cap guards, plus the discuss-phase progressive-disclosure checks.
tests/agent-size-budget.test.cjs The agent tier hard-cap guards (the agent analog).

tests/workflow-size-baseline.json, tests/agent-size-baseline.json, tests/fixtures/golden-install-parity/*.json, scripts/update-size-baseline.cjs (npm run size:baseline), and scripts/git-merge-regen-driver.cjs (npm run setup:merge-driver) are all removed by #2724: they were pure functions of the source tree, conflicted on every merge that touched them, and their functions are now served by the differential attribution check above. tests/fixtures/install-tree/*.json is the one artifact family that stays committed and normally-merged (ADR-2719 §7) — it conflicts on 0 of 7, its diffs are readable, and it preserves "the installer stopped shipping X" as a hard absolute failure with no attribution reasoning involved. Regenerate it with npm run gen:install-tree (folded into npm run regen:derived).

The QA smell ratchet

tests/loop-walk.qa.test.cjs (the qa suite) is the QA-walk harness's own self-test. Separately, scripts/qa-smell-ratchet.cjs drives that same harness end to end against the real gsd-tools binary and turns its findings into a CI gate — run it with npm run lint:qa-smells.

The harness's oracles (tests/qa/oracles.cjs) distinguish two severities:

  • A violation is the engine breaking a documented contract. It always fails the build — baseline or no baseline, acknowledged or not.
  • A smell is legal-but-questionable behavior. A smell never fails a build on its own merits. What fails is the absence of a decision about it: an unacknowledged NEW smell, or a STALE entry in tests/qa/smell-baseline.json (one that stopped firing — the baseline is shrink-only, so a fixed or changed scenario must be pruned, not left behind).

Every smell must terminate in exactly one of TWO states — there is no third "accepted with a good explanation" state:

  1. REAL — an assigned defect. File it, then acknowledge the smell with an entry (baseline entry or tests/qa/smell-acks/ fragment) carrying that issue number.
  2. FALSE POSITIVE — the oracle itself is wrong. Fix the oracle (tests/qa/oracles.cjs) so it stops firing. It is NEVER baselined.

When the ratchet reports a NEW smell, there are exactly two legitimate responses — fix the detector, or file a defect and cite its issue number:

  1. Fix the underlying behavior (or the oracle, if it's a false positive) so the smell stops firing.
  2. File a defect and acknowledge it by adding a fragment under tests/qa/smell-acks/ — the ratchet's failure output prints a paste-ready skeleton naming the required key, id, scenario, and issue fields. issue MUST be a positive integer naming the tracking issue; a free-text reason may accompany it as an optional human note but can NEVER substitute for issue — "write an explanation" is not a way to acknowledge a smell. See tests/qa/smell-acks/README.md for the full shape and lifecycle.

Run node scripts/qa-smell-ratchet.cjs --update to regenerate tests/qa/smell-baseline.json from the current run, folding in any acked fragments and pruning stale entries. --update never invents an issue number: a genuinely new smell is written with issue: null and a TODO reason, and the very next plain (non---update) run REJECTS that entry — forcing a human to triage it before it can ship. The baseline only ever shrinks: growth happens by adding an acknowledgment carrying a real issue number (a reviewable diff), never by widening the generator's tolerance and never by prose alone.

Running suites locally

npm test                    # everything (backcompat — same as before)
npm run test:unit           # only unit
npm run test:integration    # only integration
npm run test:install        # only install
npm run test:security       # only security
npm run test:slow           # only slow

npm run test:coverage       # backcompat — coverage over EVERY test
npm run test:coverage:unit  # fast coverage signal — only unit suite
npm run test:coverage:all   # alias for test:coverage

Direct harness invocation also works:

node scripts/run-tests.cjs --suite security
node scripts/run-tests.cjs --suite=security
node scripts/run-tests.cjs --files "tests/command-contract.test.cjs tests/core.test.cjs"
node scripts/run-tests.cjs --files-from .ci-selected-tests.txt

npm run test:affected (scripts/run-affected-tests.cjs) is a local-only convenience that selects tests via the require() dependency graph of your working-tree diff. CI does not use it — CI selection is the rule table in scripts/ci-test-scope.cjs, which is the authoritative mapping. If the two disagree, trust (and fix) the rule table.

Unknown suites exit non-zero with the list of valid suites. Empty suites (e.g. --suite security before any security-tagged file exists) exit 0 with a no tests in suite "..." notice on stderr so CI lanes don't go red while a suite is being populated.

CI matrix

The Tests workflow runs every PR through a scoped gate generated by scripts/ci-test-scope.cjs.

Lane Node 22 Node 24
ubuntu-latest scoped tests unit + integration + security
windows-latest — scoped Windows/path/shell tests
macos-latest full parity when required full parity when required
  • Node 22 is the engines.node floor (>=22.0.0) — must stay green.
  • Node 24 is the default development lane.
  • Scoped tests are selected from the changed paths, plus a small CLI/package smoke set. They are for confidence on the affected surface, not for counting tests.

The default PR gate runs the broad unit (under the c8 coverage gate), integration, and security suites once on Ubuntu / Node 24, scoped tests on Ubuntu / Node 22, and scoped tests on Windows / Node 24. "Scoped" means the diff-selected list from the rule table — not the full suite and not a fixed smoke set (the fixed smoke list is only the empty-selection fallback). The Windows lane's list is the Windows-sensitive subset of the selection, plus every changed test file, unconditionally (the #494 invariant, narrowed): a modified test is exercised on the divergent OS before merge at per-file cost, without paying for the three full parity lanes.

PRs touching workflow, package, test-runner, install, release, or Windows-sensitive surfaces also run the full parity matrix on macOS and the older Windows runtime, plus install and slow on the primary Ubuntu lane. Everything (including the full parity matrix) runs on every push to next, which covers the residual macOS / Windows-Node-22 cross-product for scoped PRs.

Coverage runs inside the Ubuntu / Node 24 full lane (not a separate job — that duplicated the entire unit run) and stays single-lane because multiplying coverage across OS/runtime lanes adds cost without improving the threshold signal. Note the gate's deliberate blind spot: it measures gsd-core/bin/lib/*.cjs only — scripts/, hooks/, and bin/ are unenforced, and stryker.config.mjs additionally excludes ~48% of lib lines from mutation testing (see the UNMUTATED list there). Widening either gate is tracked work, not an accident to "fix" silently by raising thresholds.

To inspect the scope locally:

npm run ci:test-scope -- --files "commands/gsd/plan-phase.md"
node scripts/ci-test-scope.cjs --base origin/next --head HEAD

Chunk packing and the test timing table

scripts/run-tests.cjs does not hand the whole selected file list to one node --test process. It packs the files into chunks, each spawned separately, because Windows caps a command line at 32,767 characters and because each chunk gets its own 600s timeout (RUN_TESTS_CHUNK_TIMEOUT_MS) and a fresh process, which bounds memory pressure.

How files are distributed across those chunks decides whether the slowest chunk sits near that timeout while the others idle. The packer weights each file by its measured duration, read from tests/test-timings.json, and places files with LPT (longest-processing-time-first: heaviest file first, each into the currently lightest chunk). Before #2456 the weight was guessed from the filename, which mis-ranked files badly enough that the slowest chunk ran ~3.9x the lightest.

Reference

Knob Default Meaning
RUN_TESTS_MAX_FILES_PER_CHUNK 60 Per-chunk weight budget. Weights are normalized so an average-cost file weighs 1, so this still reads as "about 60 average files".
RUN_TESTS_MAX_CMDLINE_CHARS 28000 argv ceiling per chunk, with headroom under the Windows 32,767 limit.
RUN_TESTS_TIMINGS_FILE tests/test-timings.json Path to the timing table. Tests override it to inject a synthetic cost profile.
RUN_TESTS_CHUNK_TIMEOUT_MS 600000 Per-chunk timeout.

The timing table is advisory and deliberately un-gated. There is no --check mode and no CI lint that fails on staleness, because timing data legitimately varies run to run. A file missing from the table falls back to the table's median weight, and a missing or unparseable table falls back to uniform weight — so drift costs chunk balance, never a red build. A count-based floor additionally guarantees the packer never produces fewer chunks than plain count-based packing would, so a badly stale table cannot collapse the suite into a few fat chunks.

How-to: regenerate the timing table

Regenerate when the suite's cost profile has visibly drifted — after adding or removing expensive tests, not on a schedule. The input is a node:test reporter event stream from a gsd-test run:

node scripts/gen-test-timings.cjs \
  ~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node22.jsonl \
  ~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node24.jsonl

Pass every lane you have. A file's recorded time is the max across the supplied streams, not the mean: the packer exists to keep the slowest lane's slowest chunk away from the timeout, so the conservative bound is the right one. Keys are sorted so a regeneration diff shows only the files whose cost moved.

Best practices for forward-compat (Node 24/26)

  • Use process.execPath when spawning Node in tests so each matrix lane exercises the lane's Node version.
  • Avoid stack-trace or error-message prose assertions. Assert err.code, structured JSON fields, or enums — Node minor releases routinely tweak error wording.
  • Prefer node:test, node:assert/strict, and node:test mocks. No external test frameworks.
  • Coverage uses c8 and propagates NODE_V8_COVERAGE through the harness's child process.

Test strategy: #443 effort + fast_mode engine

Feature: unified cross-provider effort and fast_mode knobs (issue #443). Test files: tests/model-resolver.test.cjs (unit), tests/model-resolver.test.cjs (integration).

Testing pyramid

Layer File What it covers
Unit feat-443-effort-fast-mode.test.cjs Pure logic: cascade rules, clamping, escalation math, malformed config handling, schema key validation. No CLI subprocess.
Integration feat-443-effort-fast-mode.integration.test.cjs Architecture-level invariants: cross-provider validity, totality across the 33-agent registry, CLI JSON contract, config round-trip, fast-mode honesty. Real subprocesses via runGsdTools.
E2E (pending) (not yet wired) Propagation layer: effort frontmatter / CLAUDE_CODE_EFFORT_LEVEL env actually reaching a spawned Claude Code subagent. See "Gaps" below.

Architectural invariants

Each invariant exists to prevent a specific class of production failure.

(a) Cross-provider validity

What: renderEffortForRuntime(runtime, universalEffort).value must always be a member of the runtime's real provider enum. Ground-truth enums are defined as local constants in the test — not sourced from the implementation.

PROVIDER_EFFORT_ENUMS = {
  claude: Set { 'low', 'medium', 'high', 'xhigh', 'max' }   // Anthropic output_config.effort
  codex:  Set { 'minimal', 'low', 'medium', 'high', 'xhigh' } // OpenAI model_reasoning_effort
}

Why: Passing a value outside these sets results in a 400 from the real API. The clamping logic (max -> xhigh for codex; minimal -> low for claude) must hold for every cell of the VALID_EFFORTS × runtimes matrix.

(b) Param/channel contract

What: Each runtime exposes a stable param string (the native API field name) and channel (how the value is propagated). Unknown runtimes return param: null, channel: null and pass the effort value through unchanged.

Why: Callers read .param to construct the dispatch payload. A regression here would silently drop effort from subagent invocations.

(c) Resolve-execution JSON contract

What: The gsd-tools resolve-execution <agent> command emits a JSON object with all eight keys present and typed correctly: model (string), profile (string), effort (VALID_EFFORTS member), effort_rendered (string), effort_param (string|null), effort_propagation (string|null), fast_mode (boolean), fast_mode_supported (boolean).

Why: Orchestrators and workflow dispatchers parse this JSON. A missing or mistyped field silently breaks downstream consumers.

(d) Totality across the real registry

What: For every agent in the 33-agent registry, resolveEffortInternal returns a VALID_EFFORTS member (never undefined/null), resolveFastModeInternal returns a strict boolean, and renderEffortForRuntime('claude', effort) stays within the claude provider enum.

Why: A catalog addition that introduces a missing routingTier mapping would otherwise produce undefined and propagate silently.

(e) Fast-mode honesty invariant

What: When the runtime is claude, fast_mode_supported in resolve-execution output is always false, regardless of the fast_mode config. RUNTIMES_WITH_FAST_MODE contains only 'api'.

Why: Claude Code's /fast toggle is session-level only. Emitting fast_mode: true as frontmatter on a Claude subagent is a silent no-op. Advertising fast_mode_supported: true for claude would cause orchestrators to believe the knob was wired when it is not.

(f) Precedence first-valid-wins

What: Both effort and fast_mode use a layered cascade. The test table covers all four effort layers (invocation override → agent_overrides → routing_tier_defaults → default) and all five fast_mode layers, including the case where an invalid value at a higher layer correctly falls through.

Why: Silent precedence bugs (e.g., a numeric value in agent_overrides not being rejected) would override intentional user config.

(g) Dynamic-routing composition

What: resolveEffortForTier escalates effort by attempt number independently of the model tier mapping. The test verifies the effort ladder (low -> medium -> high -> xhigh -> max), the max clamp, the max_escalations cap, and that escalate_on_failure: false suppresses escalation entirely.

Why: Effort escalation and model escalation share configuration (dynamic_routing) but must operate independently; coupling them would cause over-escalation or under-escalation.

(h) Config-tooling round-trip

What: gsd-tools config-set accepts all new key namespaces (effort.default, effort.routing_tier_defaults.<tier>, effort.agent_overrides.<agent>, fast_mode.enabled, fast_mode.routing_tier_defaults.<tier>, fast_mode.agent_overrides.<agent>) without an "Unknown config key" error, and values set via config-set are reflected in resolve-execution output.

Why: The schema validation gate (VALID_CONFIG_KEYS + DYNAMIC_KEY_PATTERNS) is separate from the resolver logic. A key missing from the schema would produce a silent write failure and appear as a bug only at runtime.

Coverage targets

Suite Target
Unit Every cascade rule, every fallthrough, every clamp. All function branches in resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, renderEffortForRuntime.
Integration All 8 architectural invariants. All 33 registered agents. All 6 provider × effort combinations for the valid-enum check. Full config-set key namespace.

Gaps / not yet covered

E2E orchestrator-spawn-propagation layer (pending follow-up wiring): The integration tests verify that GSD resolves and renders effort values correctly. They do NOT verify that the rendered values actually reach a spawned Claude Code or Codex subagent at runtime. Specifically uncovered:

  • CLAUDE_CODE_EFFORT_LEVEL env var being set and read by a spawned claude subprocess
  • output_config.effort frontmatter key surviving the AGENTS.md template substitution
  • model_reasoning_effort field surviving serialization into a Codex API request body
  • Fast-mode speed: "fast" field reaching an api-runtime request when fast_mode_supported: true

These require spawning real subagents (or stubs thereof) and asserting on the process environment / request payload — a scope that belongs in a future E2E suite under *.slow.test.cjs or dedicated fixture-driven integration work.