Files
msd-core/docs/TESTING-SUITES.md
Tom Boucher c28134ab39 fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host

Adds local/no-duplicate-fold-marker, an AST rule that reports the second and
every subsequent `folded:<name>` marker in a host file, plus RuleTester cases
and a tree-wide regression assertion.

Failing-first on purpose: the rule is registered at error and the 25 duplicated
regions are still present, so eslint and the new tree-wide test are RED. The
deletions land in the next commit.

The marker key is the whitespace-delimited token after `folded:` — not the
issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on
tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and
feat-443-effort-fast-mode are two distinct folded suites.

Refs #3271

* fix(#3271): delete 25 duplicated folded suites from three install hosts

Three consolidated install suites each carried a verbatim second copy of a
contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and
green — each duplicated block registered and ran twice on every lane.

  tests/install.test.cjs                   5981-9937  (3957 lines, 18 blocks)
  tests/install-minimal-hooks.test.cjs     2734-4015  (1282 lines,  5 blocks)
  tests/install-write-confinement.test.cjs 1754-2321  ( 568 lines,  2 blocks)

Introduced by 6d072435d (#1975 re-applying #1970's hunks on a tree that already
had them, 2026-07-03) — one stale-base re-application, three files, one commit.
Verified by marker-count bisect: 1 at 4f779eda4 and 0cc7a1a42, 2 from 6d072435d
onward.

The later copy is deleted in each case, so every file returns to what its
authoring batch produced and blame on the surviving lines stays accurate.
local/no-duplicate-fold-marker, red on the previous commit, is now green.

tests/model-resolver.test.cjs is untouched: the issue lists it, but its two
blocks are folded from two different files and are not identical. It is a false
positive of the issue's own grep, whose `[a-z0-9-]*` key truncates at `.`.

Fixes #3271

* test(#3271): property-test marker identity and pin the alias non-goal

Three review findings, all fixed inline:

1. foldMarkerOf is a parser and carried no fast-check property test. Raised
   independently by the /code-review standards axis and the isolated adversarial
   pass; the file already establishes the fc.property-driving-ruleTester idiom
   for a sibling rule. Added, two arms over markers generated from [a-z0-9-._]:
   the same marker twice always reports exactly once against firstLine 1, and
   two distinct markers never collide. The alphabet includes `.` on purpose —
   an implementation keyed on the issue's [a-z0-9-]* slice passes arm 1 and
   fails arm 2, which is exactly the model-resolver false positive.

2. meta.docs.category was the novel value 'Test hygiene'; all 16 sibling local
   rules use 'Best Practices', 'Portability' or 'Reliability'. Now
   'Best Practices'.

3. A call through a further alias (const d = __foldDescribe) was unreported and
   undocumented — accidental rather than deliberate. It is now the fourth entry
   in the rule's documented non-goals, with the reason, and pinned by a valid
   RuleTester case so it cannot drift silently.

Refs #3271

* test(#3271): name the step and elapsed time when a baseline build fails

buildBaselineAtRef runs four bounded steps and, when one exceeded its bound,
threw a bare "spawnSync ETIMEDOUT" naming neither the step nor how long
anything took. Diagnosing one real failure took four separate experiments to
recover information the throw already had.

Each step is now timed, and any throw carries the breakdown: which step failed,
its elapsed time, the timings of every step that completed before it, all three
bounds, and the tail of the child's captured stdout/stderr.

The failure message is deliberately the carrier. On the remote runner the
captured output field comes back empty in failures.json while error and stack
survive verbatim, so the message is the only channel that reaches a reader of a
remote verdict.

Refs #3271

* fix(#3271): size the baseline generator bound for the machine it runs on

Instrumentation from a real remote-runner failure gave the breakdown:

  git-worktree-add=15.1s  npm-run-build-lib=19.8s  gen-emitted-baseline=FAILED@300.1s

Steps 1 and 2 are comfortable. Only the generator exceeds its bound, and it is
not hung — it needs more than 300s there.

Measured ladder for that step: ~22s idle in a container, ~39s end-to-end in a
clean container, ~142s with 8 CPU burners on 8 cores, and >300s under the real
suite. Its cost is 19 sequential installer spawns, and spawn latency is exactly
where a container degrades worst (3.9x slower than host, against 1.1x for file
IO) — which is why a CPU-only load test did not reproduce it and why four
earlier hypotheses (container slowness, network, shallow clone, CPU contention)
all measured clean.

The 300s bound was sized on an idle machine for a step that never runs on one.
Under the remote runner the on-disk baseline cache is structurally absent — CI
restores it via actions/cache keyed on github.event.pull_request.base.sha, a key
that exists only inside GitHub Actions — so this slow path runs on every remote
verification. The result: this gate has passed 0 times in 754 runs, failing 80
times and never once executing successfully.

Raised to the 600000ms ceiling that local/no-unbounded-spawn treats as the
largest meaningful bound; the other two bounds are untouched. This makes the
gate RUN, which is the point: the alternative considered and rejected was
degrading the timeout to a skip, and that was measured to turn the suite green
with the gate silently not running at all.

The real remedy is making the cache reachable from the remote runner so the
in-job build returns to being the rare fallback ADR-2719 §5 describes. That is a
gsd-test-runner change, not one this repo can make.

Refs #3271

* fix(#3271): tolerate an overlay source that vanishes mid-walk

Observed on the remote runner, three runs across three different branches:

  ENOENT: no such file or directory, link '/work/hooks/dist/gsd-config-reload.js'
    -> '/tmp/gsd-2930-overlay-6nOZay/hooks/dist/gsd-config-reload.js'

buildOverlayRepo enumerates names with readdirSync and then acts on each one, so
statSync, copyFileSync and linkSync all sit in a TOCTOU window. hooks/dist is
regenerated by an ATOMIC REPLACE (scripts/build-hooks.js unlinks and renames), so
any concurrently running test that rebuilds hooks retires a just-listed name
mid-walk and the overlay dies on it. linkOrCopyFile already tolerated EXDEV and
EPERM; ENOENT went straight through.

On ENOENT the source is now re-examined ONCE rather than slept on. An atomic
rename is a single syscall, so by the time the failure surfaces the successor is
either already in place (the retry succeeds) or the path has genuinely left the
tree, in which case there is nothing to mirror and the leaf is skipped. No sleep
and no spin: a timing-based wait here would be the very flake being fixed. Every
other errno still propagates untouched, so a real permission or IO fault stays a
hard failure.

Five tests hold the boundary: gone-for-good skips without retrying, mid-replace
retries exactly once and places the file, EACCES still throws, a real linkSync
ENOENT is injected by monkeypatching fs and restoring it in a finally (never a
mode-bit trick, which root bypasses), and isMissingPath accepts only ENOENT.

Refs #3271

* fix(#3271): order the timeout ladder inward-out and lock it

Two review blockers, both real.

The generator bound had been raised to 600000ms — exactly the whole-chunk timeout
in scripts/run-tests.cjs:973. A step bound equal to the chunk ceiling loses the
race: the chunk is killed first and the failure arrives as an opaque "no failed
step" kill, so the per-step diagnostic added a commit earlier was built and then
made unreachable in the same change.

Separately the #2767 test declared a per-test timeout of 300000ms, BELOW the
inner bound it was meant to permit, so it could still die at the exact 300s
ceiling this was supposed to lift — via node:test's timeout rather than
spawnSync's. Its sibling declared 900000ms, above the chunk ceiling, which is the
same opaque-kill hazard from the other direction.

The three bounds only produce a useful failure if they fire inward-out, so they
now do: step 360s, per-test 480s, chunk 600s. 360s is ~3x the passing observation
(91.6s / 115.8s) and 20% above the censored 300.1s timeout, while leaving 240s of
chunk headroom for every other file sharing it. Four tests lock the ordering,
including a drift guard on the exported values — without it, editing a call
site's literal timeout would leave the ordering assertions passing while the real
ladder inverted.

Also from review:

- err.gsdBaselineStep and err.gsdBaselineTimings were written and never read
  anywhere in the tree; only the rewritten message is consumed. Removed rather
  than kept as speculative surface.
- buildOverlayRepo discarded placeVanishableLeaf's boolean at both call sites, so
  a vanished leaf left the overlay with no accounting at all. It now collects the
  skipped paths and warns once. Not thrown: a source that left the tree really is
  not part of the snapshot, and throwing would reintroduce the crash the
  tolerance removes — but silence would let a dropped leaf resurface later as an
  unrelated missing-file assertion.
- The instrumentation commit shipped no test. One now drives a real failure and
  asserts the message names the step, its elapsed time, and the bounds.

Refs #3271

* chore(#3271): backfill the changeset PR number

* fix(#3271): bound a hook fan-out as its own class, not as a bare probe

CI failure on PR #3285, job full test (windows-latest, 22, shard 2/3) — every
other lane green, including windows-latest node 24 across all three shards:

  not ok 1 - blocks push when any to-be-pushed commit matches local blocked regex
    error: bash .githooks\pre-push failed — outcome=timed_out exitCode=null stderr=
    duration_ms: 15040.2168

A bound, not a hang: the test supplies stdin via input:, so the hook is not
blocked reading its ref list, and the duration lands exactly on the 15000ms
bound.

The site used PROBE_TIMEOUT_MS, which tests/helpers/timeouts.cjs documents as "a
single short CLI query or node -e probe against a temp fixture". This is not
that. It spawns bash running .githooks/pre-push, and the hook then invokes a MOCK
git that is itself a bash script, so one runHook is roughly four Git Bash spawns.
On Windows each is Defender-scanned and the first hook test in a file pays cold
start on top. That module's own docstring warns against precisely this: a call
site that differs from its class must not be forced onto a shared value that does
not describe it.

HOOK_FANOUT_TIMEOUT_MS is that missing class — 60000ms, 4x the bound that failed
and half INSTALL_TIMEOUT_MS, which is the right order: a hook fan-out is much
lighter than a full installer run and far heavier than reading back a version
string. Two tests lock the ordering against both neighbours, including one
asserting real margin over the censored 15040ms observation, since a bound that
merely matched what was measured would be the same defect again.

Scoped deliberately: the other ~360 runHook sites keep their current bounds. This
adds the norm and applies it where a real failure demonstrated the need, rather
than sweeping a value across sites with no evidence for any of them.

Refs #3271

---------

Co-authored-by: sim <sim@local>
2026-08-09 23:41:44 -04:00

30 KiB
Raw Blame History

Testing Suites

This project's tests/ directory uses filename suffix markers to group tests into named suites. The harness scripts/run-tests.cjs filters by suite when given --suite <name>. Without a flag it runs every *.test.cjs file (the historical default — unchanged).

Tracked by issue #3597.

Suites

Suite Filename pattern What goes here
unit *.test.cjs (no other marker) Default fast lane. Pure logic, no network, no external processes beyond gsd-tools. Most tests live here.
integration *.integration.test.cjs Cross-module flows: full installer end-to-end, multi-tool orchestration, anything that crosses two or more bin entry points.
install *.install.test.cjs Tests that perform a real install/uninstall against a sandbox project. Slower; PR CI skips these on PRs and runs them on main push only.
security *.security.test.cjs Adversarial input, prompt-injection guards, fixture-driven hostile-payload sweeps.
slow *.slow.test.cjs Anything that routinely takes >5s wall-clock or holds significant memory.
qa *.qa.test.cjs End-to-end walks that drive the real gsd-tools binary across multiple loop steps against one accumulating temp project, with invariant oracles after every step. Slower than unit; excluded from the fast lane.
all (any) Explicit alias for "no filter". Equivalent to running with no --suite flag.

How to place a new test

  1. Pick the most specific bucket above.
  2. Name the file with the matching suffix: tests/<feature>.<suite>.test.cjs.
  3. If unsure, leave the suffix off — the file lands in unit, the default fast lane.

Examples:

  • tests/agent-frontmatter.test.cjs — unit
  • tests/prompt-injection-guards.security.test.cjs — security
  • tests/installer-end-to-end.install.test.cjs — install
  • tests/sdk-mutation-stress.slow.test.cjs — slow
  • tests/loop-walk.qa.test.cjs — qa

The suite-suffix convention was chosen over a directory layout (tests/security/) so the 545+ existing test files don't need to move. Existing files all classify as unit until someone explicitly retags them.

Regression tests

Do not create new top-level tests/bug-NNNN-*.test.cjs files. Add the regression case to the owning module's main test file instead (e.g. a describe('regressions') block in tests/<module>.test.cjs).

node --test spawns one child process per FILE, so file count — not test count — is the unit of CI overhead, and it is worst on Windows lanes where every spawn is Defender-scanned. The 2026-06 CI audit found 244 one-off bug-* files (~38% of the suite). That population is grandfathered in scripts/lint-regression-test-names.allowlist.json and enforced by an identity ratchet (npm run lint:regression-names, part of npm run lint:ci):

  • A new bug-* file fails CI — fold it into the owning module's file.
  • Deleting/consolidating a grandfathered file requires pruning its allowlist entry, so the baseline only ever shrinks.
  • Inherited drift (the failure names files your PR didn't add — e.g. the base branch merged bug-* files without feeding the allowlist, or you rebased and carried a pre-rebase allowlist): run node scripts/lint-regression-test-names.cjs --update and commit the regenerated allowlist. Snapshot artifacts like this allowlist (and docs/INVENTORY.md) must be regenerated after rebasing, never carried through a rebase.

A folded suite may appear only once per host

When a standalone file is folded into its owning module's test file, the moved suite is wrapped in a self-contained block carrying a marker:

// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-376-claude-js-hook-gsd-rewriter.test.cjs — …
// ────────────────────────────────────────────────────────────────────────
{
  const { describe: __foldDescribe } = require('node:test');
  __foldDescribe("folded:bug-376-claude-js-hook-gsd-rewriter (…)", () => { … });
}

Because the block is self-contained, a second verbatim copy in the same host parses, registers, and passes — twice. Nothing in a green suite reports it. #3271 found 25 such copies (~5,800 lines) across three install suites, all from a single stale-base re-application during the consolidation epic. The cost is not only wasted CI on every lane: it is a DEFECT.GENERATIVE-FIX trap, because a contributor fixing one of those regressions edits the copy they found and leaves the other asserting the old behavior, with the suite still green.

local/no-duplicate-fold-marker (eslint-rules/no-duplicate-fold-marker.cjs, error under tests/**/*.cjs) reports the second and every later occurrence of a folded:<marker> title in one file, naming the line the first occurrence sits on. When it fires, delete the copy it points at — the two blocks are the same suite, so the fix is removal, never an eslint-disable.

It keys on the whitespace-delimited token after folded:, which matters in both directions. A narrower key that stops at . would collide feat-443-effort-fast-mode.integration with feat-443-effort-fast-mode — two genuinely distinct suites that coexist in tests/model-resolver.test.cjs. Keying on the whole title instead would let a re-fold under a different batch label slip through, which is exactly the shape #3271 took. Titles without a folded: prefix, non-literal titles, and the same marker appearing in two different host files are all left alone.

The ratchet deliberately covers only bug-*. Files named feat-NNNN-* / enh-NNNN-* are feature test files — one (or one per suite) per feature is the sanctioned layout (see the #443 strategy below), not a one-off regression pattern. If issue-*/perf-* one-offs start accumulating the same way bug-* did, extend the ratchet's regex and regenerate the allowlist.

Workflow & agent size budget

Tracked by issue #1074. Bytes (not lines) per #717; LF-normalized per #683.

Workflow files (gsd-core/workflows/*.md) and agent files (agents/gsd-*.md) both ship in the installed runtime and are loaded into context — workflows on every command, agents on every subagent dispatch — so their byte size is a real cost. Two sibling guards (tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs) keep that cost from creeping up invisibly, sharing one byte-counter (measureMdFiles). Growth is caught by two independent layers:

Layer What it does Where
Differential attribution size ratchet (primary, #2724 / ADR-2719 §4) The same computed-attribution check that replaced the golden-install-parity fixtures also reports growth in any gsd-core/workflows/*.md or agents/gsd-*.md file, with the exact byte delta, comparing PR HEAD against next. Unacknowledged growth is a hard failure; shrinkage needs no acknowledgment. No committed snapshot — nothing to regenerate by hand. tests/emitted-attribution.test.cjs (real-tree test) via tests/helpers/emitted-diff.cjs
Loose tier hard caps (backstop) Absolute outer red lines per tier — workflows: XL ≤ 98304, LARGE ≤ 61440, DEFAULT ≤ 40960 bytes; agents: XL ≤ 57344, LARGE ≤ 49152, DEFAULT ≤ 24576 bytes. A cap is never raised when a file approaches it: crossing it means extract, not bump. Independent of the ratchet above — unaffected by #2724. XL/LARGE/DEFAULT_CAP in each guard file

discuss-phase.md additionally has a thin-dispatcher target of < 32000 bytes (the discuss-phase progressive-disclosure split, #717). A net-new agent is DEFAULT-tier and already bounded by the DEFAULT cap — no separate new-agent cap is needed. (This tier-cap machinery is distinct from the separate 45 KB-char extraction-evidence threshold on gsd-planner enforced by tests/planner-decomposition.test.cjs — that one proves mode sections were extracted; this one bounds total agent bytes.)

How-to: a workflow or agent grew and CI is red

The differential attribution check reports the file and the byte delta. To resolve:

  1. Justify the growth in your PR (a sentence in the description is enough) — the acknowledgment entry (below) is the review record that the larger size was a deliberate, seen decision, not silent drift.
  2. Add an acknowledgment entry in tests/emitted-drift-ack.json naming the file and the reason, per CONTEXT.md's ### Emitted Artifact Provenance entry. This is deliberately a committed file, not a flag — the entry appears in your PR diff, so touching it is the visible signal.
  3. Or shrink it instead of acknowledging. Prefer extraction when the growth is incidental: for a workflow, move per-mode bodies to workflows/<name>/modes/, templates to workflows/<name>/templates/, and shared prose to gsd-core/references/; for an agent, lift shared boilerplate into gsd-core/references/ and @-reference it — then load it LAZILY. Do not convert them to eager @-required_reading includes: that shrinks the file's bytes without shrinking loaded context, so it games the guard while making the real cost worse. See workflows/discuss-phase/ for the progressive-disclosure pattern.

If a hard cap (not the ratchet) is what failed, an acknowledgment will not help — that is the signal to extract, per step 3.

Reference

Artifact Role
scripts/workflow-size.cjs Single source of truth — LF-normalized byte counter (lfByteCount) + generic measureMdFiles(dir, predicate) (backs both workflows and agents) + workflow enumeration (listWorkflowStems, measureWorkflows). Imported by both guards and by tests/helpers/emitted-runtime.cjs's currentSizes() so they can never measure differently.
tests/emitted-attribution.test.cjs + tests/helpers/emitted-diff.cjs The differential attribution check and its size ratchet (ADR-2719). The sole mechanism for both emitted-content propagation AND per-file size growth as of #2724.
tests/emitted-drift-ack.json Committed acknowledgment file for unattributable emitted-content ripples and for size growth. Absent = no acks; its presence is the alarm.
npm run regen:derived Runs every remaining generator in dependency order (build → registry → ADR index → capability matrix → inventory manifest → manifest versions → tests/fixtures/install-tree/*.json).
tests/workflow-size-budget.test.cjs The workflow tier hard-cap guards, plus the discuss-phase progressive-disclosure checks.
tests/agent-size-budget.test.cjs The agent tier hard-cap guards (the agent analog).

tests/workflow-size-baseline.json, tests/agent-size-baseline.json, tests/fixtures/golden-install-parity/*.json, scripts/update-size-baseline.cjs (npm run size:baseline), and scripts/git-merge-regen-driver.cjs (npm run setup:merge-driver) are all removed by #2724: they were pure functions of the source tree, conflicted on every merge that touched them, and their functions are now served by the differential attribution check above. tests/fixtures/install-tree/*.json is the one artifact family that stays committed and normally-merged (ADR-2719 §7) — it conflicts on 0 of 7, its diffs are readable, and it preserves "the installer stopped shipping X" as a hard absolute failure with no attribution reasoning involved. Regenerate it with npm run gen:install-tree (folded into npm run regen:derived).

The QA smell ratchet

tests/loop-walk.qa.test.cjs (the qa suite) is the QA-walk harness's own self-test. Separately, scripts/qa-smell-ratchet.cjs drives that same harness end to end against the real gsd-tools binary and turns its findings into a CI gate — run it with npm run lint:qa-smells.

The harness's oracles (tests/qa/oracles.cjs) distinguish two severities:

  • A violation is the engine breaking a documented contract. It always fails the build — baseline or no baseline, acknowledged or not.
  • A smell is legal-but-questionable behavior. A smell never fails a build on its own merits. What fails is the absence of a decision about it: an unacknowledged NEW smell, or a STALE entry in tests/qa/smell-baseline.json (one that stopped firing — the baseline is shrink-only, so a fixed or changed scenario must be pruned, not left behind).

Every smell must terminate in exactly one of TWO states — there is no third "accepted with a good explanation" state:

  1. REAL — an assigned defect. File it, then acknowledge the smell with an entry (baseline entry or tests/qa/smell-acks/ fragment) carrying that issue number.
  2. FALSE POSITIVE — the oracle itself is wrong. Fix the oracle (tests/qa/oracles.cjs) so it stops firing. It is NEVER baselined.

When the ratchet reports a NEW smell, there are exactly two legitimate responses — fix the detector, or file a defect and cite its issue number:

  1. Fix the underlying behavior (or the oracle, if it's a false positive) so the smell stops firing.
  2. File a defect and acknowledge it by adding a fragment under tests/qa/smell-acks/ — the ratchet's failure output prints a paste-ready skeleton naming the required key, id, scenario, and issue fields. issue MUST be a positive integer naming the tracking issue; a free-text reason may accompany it as an optional human note but can NEVER substitute for issue — "write an explanation" is not a way to acknowledge a smell. See tests/qa/smell-acks/README.md for the full shape and lifecycle.

Run node scripts/qa-smell-ratchet.cjs --update to regenerate tests/qa/smell-baseline.json from the current run, folding in any acked fragments and pruning stale entries. --update never invents an issue number: a genuinely new smell is written with issue: null and a TODO reason, and the very next plain (non---update) run REJECTS that entry — forcing a human to triage it before it can ship. The baseline only ever shrinks: growth happens by adding an acknowledgment carrying a real issue number (a reviewable diff), never by widening the generator's tolerance and never by prose alone.

Running suites locally

npm test                    # everything (backcompat — same as before)
npm run test:unit           # only unit
npm run test:integration    # only integration
npm run test:install        # only install
npm run test:security       # only security
npm run test:slow           # only slow

npm run test:coverage       # backcompat — coverage over EVERY test
npm run test:coverage:unit  # fast coverage signal — only unit suite
npm run test:coverage:all   # alias for test:coverage

Direct harness invocation also works:

node scripts/run-tests.cjs --suite security
node scripts/run-tests.cjs --suite=security
node scripts/run-tests.cjs --files "tests/command-contract.test.cjs tests/core.test.cjs"
node scripts/run-tests.cjs --files-from .ci-selected-tests.txt

npm run test:affected (scripts/run-affected-tests.cjs) is a local-only convenience that selects tests via the require() dependency graph of your working-tree diff. CI does not use it — CI selection is the rule table in scripts/ci-test-scope.cjs, which is the authoritative mapping. If the two disagree, trust (and fix) the rule table.

Unknown suites exit non-zero with the list of valid suites. Empty suites (e.g. --suite security before any security-tagged file exists) exit 0 with a no tests in suite "..." notice on stderr so CI lanes don't go red while a suite is being populated.

The live-config hermeticity guard

Every run-tests.cjs invocation snapshots GSD's own install footprint in each live runtime config directory before the suite and re-checks it afterwards. It exists because the failure it catches is silent by construction: a test that resolves a config directory from the ambient environment instead of a sandbox writes into your real ~/.claude (or $GSD_HOME/.gsd, or a Kimi config.toml), and nothing reports it. CI cannot catch this class at all — CI never has CLAUDE_CONFIG_DIR and friends set.

The guard watches only what GSD unambiguously owns — its top-level install footprint plus gsd--prefixed children of directories shared with the host agent — never whole config roots, because a host agent legitimately writing history.jsonl mid-run would make the guard cry wolf, and a guard that cries wolf gets switched off.

Two environment variables control it:

Variable Effect
GSD_STRICT_LIVE_CONFIG_GUARD=1 A detected write fails the run. Set on the Linux/macOS lanes of every CI job that runs the suite.
GSD_SKIP_LIVE_CONFIG_GUARD=1 Skips the check entirely.

Unset, the guard reports and does not fail — deliberately, not timidly. On its first CI run it surfaced pre-existing leaks on the Windows lane, where os.homedir() reads USERPROFILE and ~190 test sites sandbox HOME alone. Those are real and worth fixing, but they are a different defect class, and a brand-new gate that instantly reds an unrelated lane gets reverted rather than obeyed. Windows lanes therefore stay report-only until that sweep lands; this repo has the pattern already, in the local/no-source-grep ESLint rule that shipped at warn and was promoted to error after its cleanup (ADR 452).

GSD_SKIP_LIVE_CONFIG_GUARD is a bypass on a safety check, so it is documented here rather than left to be discovered in the source: an undocumented bypass is one people eventually set without knowing what they turned off. If you need it routinely, that is a bug report, not a workflow.

Reported paths are labelled CREATED, MODIFIED, DELETED, or UNVERIFIED. UNVERIFIED means a scan bound was hit and the path could not be attested either way — it is never the same as clean.

CI matrix

The Tests workflow runs every PR through a scoped gate generated by scripts/ci-test-scope.cjs.

Lane Node 22 Node 24
ubuntu-latest scoped tests unit + integration + security
windows-latest — scoped Windows/path/shell tests
macos-latest full parity when required full parity when required
  • Node 22 is the engines.node floor (>=22.0.0) — must stay green.
  • Node 24 is the default development lane.
  • Scoped tests are selected from the changed paths, plus a small CLI/package smoke set. They are for confidence on the affected surface, not for counting tests.

The default PR gate runs the broad unit (under the c8 coverage gate), integration, and security suites once on Ubuntu / Node 24, scoped tests on Ubuntu / Node 22, and scoped tests on Windows / Node 24. "Scoped" means the diff-selected list from the rule table — not the full suite and not a fixed smoke set (the fixed smoke list is only the empty-selection fallback). The Windows lane's list is the Windows-sensitive subset of the selection, plus every changed test file, unconditionally (the #494 invariant, narrowed): a modified test is exercised on the divergent OS before merge at per-file cost, without paying for the three full parity lanes.

PRs touching workflow, package, test-runner, install, release, or Windows-sensitive surfaces also run the full parity matrix on macOS and the older Windows runtime, plus install and slow on the primary Ubuntu lane. Everything (including the full parity matrix) runs on every push to next, which covers the residual macOS / Windows-Node-22 cross-product for scoped PRs.

Coverage runs inside the Ubuntu / Node 24 full lane (not a separate job — that duplicated the entire unit run) and stays single-lane because multiplying coverage across OS/runtime lanes adds cost without improving the threshold signal. Note the gate's deliberate blind spot: it measures gsd-core/bin/lib/*.cjs only — scripts/, hooks/, and bin/ are unenforced, and stryker.config.mjs additionally excludes ~48% of lib lines from mutation testing (see the UNMUTATED list there). Widening either gate is tracked work, not an accident to "fix" silently by raising thresholds.

To inspect the scope locally:

npm run ci:test-scope -- --files "commands/gsd/plan-phase.md"
node scripts/ci-test-scope.cjs --base origin/next --head HEAD

Chunk packing and the test timing table

scripts/run-tests.cjs does not hand the whole selected file list to one node --test process. It packs the files into chunks, each spawned separately, because Windows caps a command line at 32,767 characters and because each chunk gets its own 600s timeout (RUN_TESTS_CHUNK_TIMEOUT_MS) and a fresh process, which bounds memory pressure.

How files are distributed across those chunks decides whether the slowest chunk sits near that timeout while the others idle. The packer weights each file by its measured duration, read from tests/test-timings.json, and places files with LPT (longest-processing-time-first: heaviest file first, each into the currently lightest chunk). Before #2456 the weight was guessed from the filename, which mis-ranked files badly enough that the slowest chunk ran ~3.9x the lightest.

Reference

Knob Default Meaning
RUN_TESTS_MAX_FILES_PER_CHUNK 60 Per-chunk weight budget. Weights are normalized so an average-cost file weighs 1, so this still reads as "about 60 average files".
RUN_TESTS_MAX_CMDLINE_CHARS 28000 argv ceiling per chunk, with headroom under the Windows 32,767 limit.
RUN_TESTS_TIMINGS_FILE tests/test-timings.json Path to the timing table. Tests override it to inject a synthetic cost profile.
RUN_TESTS_CHUNK_TIMEOUT_MS 600000 Per-chunk timeout.

The timing table is advisory and deliberately un-gated. There is no --check mode and no CI lint that fails on staleness, because timing data legitimately varies run to run. A file missing from the table falls back to the table's median weight, and a missing or unparseable table falls back to uniform weight — so drift costs chunk balance, never a red build. A count-based floor additionally guarantees the packer never produces fewer chunks than plain count-based packing would, so a badly stale table cannot collapse the suite into a few fat chunks.

How-to: regenerate the timing table

Regenerate when the suite's cost profile has visibly drifted — after adding or removing expensive tests, not on a schedule. The input is a node:test reporter event stream from a gsd-test run:

node scripts/gen-test-timings.cjs \
  ~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node22.jsonl \
  ~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node24.jsonl

Pass every lane you have. A file's recorded time is the max across the supplied streams, not the mean: the packer exists to keep the slowest lane's slowest chunk away from the timeout, so the conservative bound is the right one. Keys are sorted so a regeneration diff shows only the files whose cost moved.

Best practices for forward-compat (Node 24/26)

  • Use process.execPath when spawning Node in tests so each matrix lane exercises the lane's Node version.
  • Avoid stack-trace or error-message prose assertions. Assert err.code, structured JSON fields, or enums — Node minor releases routinely tweak error wording.
  • Prefer node:test, node:assert/strict, and node:test mocks. No external test frameworks.
  • Coverage uses c8 and propagates NODE_V8_COVERAGE through the harness's child process.

Test strategy: #443 effort + fast_mode engine

Feature: unified cross-provider effort and fast_mode knobs (issue #443). Test files: tests/model-resolver.test.cjs (unit), tests/model-resolver.test.cjs (integration).

Testing pyramid

Layer File What it covers
Unit feat-443-effort-fast-mode.test.cjs Pure logic: cascade rules, clamping, escalation math, malformed config handling, schema key validation. No CLI subprocess.
Integration feat-443-effort-fast-mode.integration.test.cjs Architecture-level invariants: cross-provider validity, totality across the 33-agent registry, CLI JSON contract, config round-trip, fast-mode honesty. Real subprocesses via runGsdTools.
E2E (pending) (not yet wired) Propagation layer: effort frontmatter / CLAUDE_CODE_EFFORT_LEVEL env actually reaching a spawned Claude Code subagent. See "Gaps" below.

Architectural invariants

Each invariant exists to prevent a specific class of production failure.

(a) Cross-provider validity

What: renderEffortForRuntime(runtime, universalEffort).value must always be a member of the runtime's real provider enum. Ground-truth enums are defined as local constants in the test — not sourced from the implementation.

PROVIDER_EFFORT_ENUMS = {
  claude: Set { 'low', 'medium', 'high', 'xhigh', 'max' }   // Anthropic output_config.effort
  codex:  Set { 'minimal', 'low', 'medium', 'high', 'xhigh' } // OpenAI model_reasoning_effort
}

Why: Passing a value outside these sets results in a 400 from the real API. The clamping logic (max -> xhigh for codex; minimal -> low for claude) must hold for every cell of the VALID_EFFORTS × runtimes matrix.

(b) Param/channel contract

What: Each runtime exposes a stable param string (the native API field name) and channel (how the value is propagated). Unknown runtimes return param: null, channel: null and pass the effort value through unchanged.

Why: Callers read .param to construct the dispatch payload. A regression here would silently drop effort from subagent invocations.

(c) Resolve-execution JSON contract

What: The gsd-tools resolve-execution <agent> command emits a JSON object with all eight keys present and typed correctly: model (string), profile (string), effort (VALID_EFFORTS member), effort_rendered (string), effort_param (string|null), effort_propagation (string|null), fast_mode (boolean), fast_mode_supported (boolean).

Why: Orchestrators and workflow dispatchers parse this JSON. A missing or mistyped field silently breaks downstream consumers.

(d) Totality across the real registry

What: For every agent in the 33-agent registry, resolveEffortInternal returns a VALID_EFFORTS member (never undefined/null), resolveFastModeInternal returns a strict boolean, and renderEffortForRuntime('claude', effort) stays within the claude provider enum.

Why: A catalog addition that introduces a missing routingTier mapping would otherwise produce undefined and propagate silently.

(e) Fast-mode honesty invariant

What: When the runtime is claude, fast_mode_supported in resolve-execution output is always false, regardless of the fast_mode config. RUNTIMES_WITH_FAST_MODE contains only 'api'.

Why: Claude Code's /fast toggle is session-level only. Emitting fast_mode: true as frontmatter on a Claude subagent is a silent no-op. Advertising fast_mode_supported: true for claude would cause orchestrators to believe the knob was wired when it is not.

(f) Precedence first-valid-wins

What: Both effort and fast_mode use a layered cascade. The test table covers all four effort layers (invocation override → agent_overrides → routing_tier_defaults → default) and all five fast_mode layers, including the case where an invalid value at a higher layer correctly falls through.

Why: Silent precedence bugs (e.g., a numeric value in agent_overrides not being rejected) would override intentional user config.

(g) Dynamic-routing composition

What: resolveEffortForTier escalates effort by attempt number independently of the model tier mapping. The test verifies the effort ladder (low -> medium -> high -> xhigh -> max), the max clamp, the max_escalations cap, and that escalate_on_failure: false suppresses escalation entirely.

Why: Effort escalation and model escalation share configuration (dynamic_routing) but must operate independently; coupling them would cause over-escalation or under-escalation.

(h) Config-tooling round-trip

What: gsd-tools config-set accepts all new key namespaces (effort.default, effort.routing_tier_defaults.<tier>, effort.agent_overrides.<agent>, fast_mode.enabled, fast_mode.routing_tier_defaults.<tier>, fast_mode.agent_overrides.<agent>) without an "Unknown config key" error, and values set via config-set are reflected in resolve-execution output.

Why: The schema validation gate (VALID_CONFIG_KEYS + DYNAMIC_KEY_PATTERNS) is separate from the resolver logic. A key missing from the schema would produce a silent write failure and appear as a bug only at runtime.

Coverage targets

Suite Target
Unit Every cascade rule, every fallthrough, every clamp. All function branches in resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, renderEffortForRuntime.
Integration All 8 architectural invariants. All 33 registered agents. All 6 provider × effort combinations for the valid-enum check. Full config-set key namespace.

Gaps / not yet covered

E2E orchestrator-spawn-propagation layer (pending follow-up wiring): The integration tests verify that GSD resolves and renders effort values correctly. They do NOT verify that the rendered values actually reach a spawned Claude Code or Codex subagent at runtime. Specifically uncovered:

  • CLAUDE_CODE_EFFORT_LEVEL env var being set and read by a spawned claude subprocess
  • output_config.effort frontmatter key surviving the AGENTS.md template substitution
  • model_reasoning_effort field surviving serialization into a Codex API request body
  • Fast-mode speed: "fast" field reaching an api-runtime request when fast_mode_supported: true

These require spawning real subagents (or stubs thereof) and asserting on the process environment / request payload — a scope that belongs in a future E2E suite under *.slow.test.cjs or dedicated fixture-driven integration work.