Files
msd-core/docs/TESTING-SUITES.md
Tom Boucher 4d65c248e5 fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector

Tests only, committed ahead of the implementation so the RED run is real.

- tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a
  ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative
  cases proving seam calls and path-call-plus-slash-literal are not platform
  signals; positive pins that genuine platform content, seam-bypassing spawns,
  chmod and symlink still classify in; macOS signal set and generated list
  unchanged.
- tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest
  rows and test-conformance still has 3 windows + 1 macOS.
- tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a
  non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint
  test does, and resolveSelection rejects the retired windows scope.

Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): delete the second Windows selector and narrow the conformance tier

Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped
conformance tier on real Windows/macOS — was not met. Measured on PR #4640
(run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7
changed test files running on a real Windows runner twice.

Two selectors, only one in the epic's scope. The test job's three scope:windows
shards predate the epic (#494, sharded #3057) and gate on product_changed, not
full_matrix, so they fire on every product PR whatever Phase 3's classifier
decides. They are deleted; test-conformance becomes the sole Windows selector,
as it already was for macOS. Non-Linux jobs 7 -> 4.

Gating the lane instead was rejected as provably redundant: for a test file
reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file),
and that same predicate sets full_matrix, which turns test-conformance on. Every
file a gated lane would run is already covered in the same run. The lane's one
non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename
heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix
instead of feeding a parallel lane.

Two detectors matched the repo's own test idiom rather than any platform signal
and carried 226 of the tier's sole-signal membership against 41 for the other
eight: process-seam-subprocess (335 files, 118 unique) matches the
tests/helpers.cjs entry points nearly every CLI test uses, and going through the
seam is the opposite of a platform signal since shell-command-projection takes
platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs
only a path call anywhere plus a slash literal anywhere, and that class is
already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed.
Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured.

Adds the size gate Phase 2 never had, as a ratio against a live denominator so
it cannot stop binding as the suite grows.

292 files leave real-OS Windows execution. The drop-out set was audited: 14 have
a platform-suggestive filename and all 14 are static source-text analyses or
seam-mediated CLI tests. raw-child-process was investigated as a suspected false
negative and left unchanged -- relaxing it adds 13 files, all false positives.

macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated
macos-conformance-tier.generated.cjs is byte-identical at 196 files.

Fixes #4641
Refs #4589, #4591, #4592, #4593, #4603

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): register the new ADR path in the docs-guard exempt baseline

tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md
in a comment justifying the retired windows scope; lint-docs-guard-registration
tracks that reference set, so the baseline needs the new path. Verified the
exemption still holds: the path is prose, not a filesystem read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): make the escalation tier-backed and drop every hardcoded count

Three follow-ups from measuring the first pass rather than trusting it.

The windows-hint escalation now requires tier membership as well as the
filename hint. Setting full_matrix runs test-conformance, which runs only the
tier; escalating on a test that is NOT in the tier costs four jobs and still
never runs that test on Windows. Measured over the 16 RULES entries the
narrowed predicate fires on exactly the same rules today, so this is
correct-by-construction rather than a behavior change. The broader variant --
escalate on any tier member a rule pulls in, ignoring the hint -- was measured
at 14/16 rules and rejected as over-broad.

Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as
soon as the rebase pulled in one new test file from #4253, which is the whole
argument against them: the ceilings are ratios against a live denominator, the
committed lists are pinned by comparison against a fresh classification of the
live tree, and the three named probe files now assert on their SIGNAL rather
than on membership in a literal list -- asserting by filename is the exact
error this PR fixes in the classifier.

Regenerates both lists against the rebased tree. Same-tree figures are now
547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none
added; macOS is unchanged at 197 with a zero-line diff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke

An isolated adversarial review found a real false negative. Removing the
blanket process-seam-subprocess detector also removed the only coverage for
tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook
spawns options.interpreter via real spawnSync, so
runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary
executing a shell script extracted from workflow markdown. The seam argument
holds for src/shell-command-projection.cts, which takes platform as an
injected parameter; it does NOT hold for the test helpers, which spawn real
binaries. Conflating the two is what made the blanket detector look purely
noisy -- it was 99% noise wrapping a real signal.

Adds a narrow shell-interpreter-spawn category keyed on a real interpreter
option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are
added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33%
ceiling. All 9 confirmed by reading the matching source line, zero comment or
fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured
and rejected -- each adds 9 files but misses the counterexample entirely.

Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its
scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD
strictly rather than carrying the Windows report-only carve-out. The carve-out
now lives solely on test-conformance, whose matrix does include windows.

Repairs seven pre-existing assertions the category removal invalidated,
preserving each case's purpose rather than deleting coverage, and converts the
last hardcoded tier bounds to live-derived ratios -- including the macOS
sanity range that was still a magic [100, 350].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): keep the confinement test on a real OS via a documented allowlist

A security review found tests/external-descriptor-confinement.test.cjs had
dropped out of the Windows tier. It must stay in, and no content signal can
express why: it exercises isPathConfined (src/external-descriptor-trust.cts),
which uses the AMBIENT path module -- path.resolve(root, target) and path.sep
-- with no injection. Its win32 semantics (drive letters, UNC, separator) are
only reachable by actually running on Windows, and it is a security-relevant
write-confinement gate. A content classifier cannot see 'this module reads the
ambient path module', so no regex belongs here.

Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows
tier only. A Map rather than a list so an entry without a reason is impossible
by construction, and tests assert every entry names a file that exists on disk
so a stale entry fails loudly instead of rotting. This is the centrally-
enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's
portability-vocab.cjs already models -- deliberately not a heuristic.

Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is
untouched and byte-identical: the win32 concern does not apply to a POSIX
runner, and a test asserts the allowlist does not leak into that tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4641): inject the path impl into isPathConfined and correct the ADR count

Two review findings, both fixed rather than dispositioned.

A security review found tests/external-descriptor-confinement.test.cjs had left
real-OS execution. The allowlist pinned it back, but that only restored
INCIDENTAL coverage: isPathConfined used the ambient path module, and its test
carried POSIX-only literals, so a win32 confinement escape was unverified on
every platform including Windows. isPathConfined now takes an optional third
parameter carrying the path implementation, defaulting to the ambient module.
Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change
is purely additive and every existing two-argument caller is byte-identical.

Tests now inject path.win32 and path.posix, covering a different drive letter,
a cross-drive absolute, backslash and forward-slash traversal, UNC, and the
startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators.
Proved load-bearing: dropping the + p.sep from the prefix check fails exactly
the two boundary cases and nothing else. Callers' suites 149/149.

The spec review caught an off-by-one: the ADR narrated a 264-file tier while the
committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264
-> 265 (28.5%).

Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated
classify() as zeroing windows_tests, a key this change removes -- kept as
historical narration but labelled as such.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#131): make the unwritable-HOME test actually test something

Found by sweeping for the root-bypass class after fixing commit-files-deletion.
This one is the silent variant, and it was broken twice over.

First, the condition: the test made a fake HOME unwritable with chmod 0o500.
The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed
writable and the hostile condition never existed. Replaced with a HOME whose
PARENT is a regular file, so every write under it fails ENOTDIR at the VFS
layer for every uid -- no permission check is involved at all.

Second, and more fundamental: the probe was npm --version, which on npm 11.19.0
performs zero filesystem I/O against HOME. Proven rather than assumed --
neutralizing runNpm()'s isolation turned the sibling test red while this one
stayed green, so its assertion could never detect the regression it guards, on
any uid, with or without the condition fix. npm config get cache was tried next
and proved vacuous the same way (it only string-resolves the path). The probe is
now npm cache verify, which really does mkdir _cacache under HOME.

Re-proved load-bearing after the change: with isolation neutralized the test now
fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was
restored and verified diff-clean; suite 13/13.

Refs #4641

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the net drop-out figure in ADR-4641

The Consequences section still said 292 files leave real-OS Windows execution.
That was the count before the narrow shell-interpreter-spawn replacement
restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary
categories rather than only raw-child-process, and clarifies that the 14-file
filename audit was against the 292 initially dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the rejected concentration ceiling and its measurement

Applying Goodhart's own question to the new ceiling -- how would you make this
metric look good without improving what it represents -- surfaces a real
weakness: a ratio can be satisfied by inflating the denominator, so adding
OS-agnostic tests loosens it without narrowing the tier.

The obvious companion gate was a sole-signal concentration ceiling, since the
original defect was one detector carrying half the tier. Measured and rejected:
peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the
historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the
original defect; any threshold below it fails on a legitimate category. The
discriminator is whether a signal is platform-meaningful, which no threshold
encodes. Weakness disclosed rather than covered by a gate that does not bind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4641): add the changeset fragment for the confinement-check change

changeset-lint failed on PR #4643: the PR touches user-facing paths and carried
no fragment. The earlier no-changeset call matched #4604's CI-only precedent and
was correct then; it was not revisited once the PR grew a src/ change, which is
my miss.

The fragment describes the real user-visible improvement: the external-descriptor
write-confinement check's Windows semantics are now verified deterministically
rather than only when the suite happened to run on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): correct the tier count in TESTING-SUITES.md

Said the tier narrowed from 546 to 254. The final committed list is 265 of 931
eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored
9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in
the ADR, in a live reference page rather than a dated record, so it states the
current truth rather than carrying an amendment note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record the measured aggregate from real CI job lists

Epic #4589's closeout asserted its reduction from a static count; #4641's
acceptance criterion asks for a figure read off a real run. Recorded here:
test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run
against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4.

Also states the caveat that a PR's total CHECK count is not a clean before/after
comparison, since many gates are path-scoped and this change touches a broader
path set -- the like-for-like figure is the test.yml job count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): compare job totals the same way on both sides

The measured-aggregate table put #4640's COMPLETED run total (21) against this
run's count at matrix-expansion time (15). Those are not the same measurement:
the completed total includes the post-test Coverage gate and baseline-publisher
jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux
jobs 7 -> 4, was correct and is unchanged.

Called out in the table rather than silently corrected -- comparing two
differently-derived numbers is exactly the error class this ADR is about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): record measured conformance wall-clock and date the stale counterfactual

Adds the per-job durations from both runs. The honest read is that this is a
correctness win more than a speed one: file count fell 52% but wall-clock only
9-29%, because what was removed were the cheap static tests and what remains is
concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects
a future narrowing to buy time proportional to file count.

The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on
the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about --
pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its
tier was UNCHANGED at 197 files, which fixes that as runner variance and is
noted as a caution against reading a single duration as signal.

Also dates the symlink-keyword counterfactual, which cited a 254-file tier from
before the replacement category and allowlist took it to its final 265.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero

next gained #4644 mid-flight, so every absolute count shifted. Re-measured on
the tree this actually ships against (932 eligible): 548 -> 257 by detector
removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed,
9 restored. macOS 198, unchanged by this PR.

The percentages did not move across three rebases (58.8% -> 28.5%), which is
the whole argument for expressing the ceilings as ratios rather than counts --
noted in the ADR since it is now evidence rather than assertion.

Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own
win32 test cases introduced the literal win32 into the pinned file, so it
classifies in on content via win32-darwin-literal. The entry stays and the
reason is written down, because the file's real-OS need is a property of the
code under test (isPathConfined reads the ambient path module), not of the
test's text -- the text that currently saves it is incidental and could be
refactored away silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 17:00:11 -04:00

43 KiB
Raw Blame History

Testing Suites

This project's tests/ directory uses filename suffix markers to group tests into named suites. The harness scripts/run-tests.cjs filters by suite when given --suite <name>. Without a flag it runs every *.test.cjs file (the historical default — unchanged).

Tracked by issue #3597.

Suites

Suite Filename pattern What goes here
unit *.test.cjs (no other marker) Default fast lane. Pure logic, no network, no external processes beyond gsd-tools. Most tests live here.
integration *.integration.test.cjs Cross-module flows: full installer end-to-end, multi-tool orchestration, anything that crosses two or more bin entry points.
install *.install.test.cjs Tests that perform a real install/uninstall against a sandbox project. Slower; PR CI skips these on PRs and runs them on main push only.
security *.security.test.cjs Adversarial input, prompt-injection guards, fixture-driven hostile-payload sweeps.
slow *.slow.test.cjs Anything that routinely takes >5s wall-clock or holds significant memory.
qa *.qa.test.cjs End-to-end walks that drive the real gsd-tools binary across multiple loop steps against one accumulating temp project, with invariant oracles after every step. Slower than unit; excluded from the fast lane.
all (any) Explicit alias for "no filter". Equivalent to running with no --suite flag.

How to place a new test

  1. Pick the most specific bucket above.
  2. Name the file with the matching suffix: tests/<feature>.<suite>.test.cjs.
  3. If unsure, leave the suffix off — the file lands in unit, the default fast lane.

Examples:

  • tests/agent-frontmatter.test.cjs — unit
  • tests/prompt-injection-guards.security.test.cjs — security
  • tests/installer-end-to-end.install.test.cjs — install
  • tests/sdk-mutation-stress.slow.test.cjs — slow
  • tests/loop-walk.qa.test.cjs — qa

The suite-suffix convention was chosen over a directory layout (tests/security/) so the 545+ existing test files don't need to move. Existing files all classify as unit until someone explicitly retags them.

Regression tests

Do not create new top-level tests/bug-NNNN-*.test.cjs, tests/fix-NNNN-*.test.cjs, or tests/issue-NNNN-*.test.cjs files. Add the regression case to the owning module's main test file instead (e.g. a describe('regressions') block in tests/<module>.test.cjs).

node --test spawns one child process per FILE, so file count — not test count — is the unit of CI overhead, and it is worst on Windows lanes where every spawn is Defender-scanned. The 2026-06 CI audit found 244 one-off bug-* files (~38% of the suite). That population is grandfathered in scripts/lint-regression-test-names.allowlist.json and enforced by an identity ratchet (npm run lint:regression-names, part of npm run lint:ci), which also bans fix-* and issue-* NNNN-prefixed files the same way:

  • A new bug-*, fix-*, or issue-* NNNN-prefixed file fails CI — fold it into the owning module's file.
  • Deleting/consolidating a grandfathered file requires pruning its allowlist entry, so the baseline only ever shrinks.
  • Inherited drift (the failure names files your PR didn't add — e.g. the base branch merged bug-*/fix-*/issue-* files without feeding the allowlist, or you rebased and carried a pre-rebase allowlist): run node scripts/lint-regression-test-names.cjs --update and commit the regenerated allowlist. Snapshot artifacts like this allowlist (and docs/INVENTORY.md) must be regenerated after rebasing, never carried through a rebase.

A folded suite may appear only once per host

When a standalone file is folded into its owning module's test file, the moved suite is wrapped in a self-contained block carrying a marker:

// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-376-claude-js-hook-gsd-rewriter.test.cjs — …
// ────────────────────────────────────────────────────────────────────────
{
  const { describe: __foldDescribe } = require('node:test');
  __foldDescribe("folded:bug-376-claude-js-hook-gsd-rewriter (…)", () => { … });
}

Because the block is self-contained, a second verbatim copy in the same host parses, registers, and passes — twice. Nothing in a green suite reports it. #3271 found 25 such copies (~5,800 lines) across three install suites, all from a single stale-base re-application during the consolidation epic. The cost is not only wasted CI on every lane: it is a DEFECT.GENERATIVE-FIX trap, because a contributor fixing one of those regressions edits the copy they found and leaves the other asserting the old behavior, with the suite still green.

local/no-duplicate-fold-marker (eslint-rules/no-duplicate-fold-marker.cjs, error under tests/**/*.cjs) reports the second and every later occurrence of a folded:<marker> title in one file, naming the line the first occurrence sits on. When it fires, delete the copy it points at — the two blocks are the same suite, so the fix is removal, never an eslint-disable.

It keys on the whitespace-delimited token after folded:, which matters in both directions. A narrower key that stops at . would collide feat-443-effort-fast-mode.integration with feat-443-effort-fast-mode — two genuinely distinct suites that coexist in tests/model-resolver.test.cjs. Keying on the whole title instead would let a re-fold under a different batch label slip through, which is exactly the shape #3271 took. Titles without a folded: prefix, non-literal titles, and the same marker appearing in two different host files are all left alone.

The ratchet deliberately covers only bug-*. Files named feat-NNNN-* / enh-NNNN-* are feature test files — one (or one per suite) per feature is the sanctioned layout (see the #443 strategy below), not a one-off regression pattern. If issue-*/perf-* one-offs start accumulating the same way bug-* did, extend the ratchet's regex and regenerate the allowlist.

Workflow & agent size budget

Tracked by issue #1074. Bytes (not lines) per #717; LF-normalized per #683.

Workflow files (gsd-core/workflows/*.md) and agent files (agents/gsd-*.md) both ship in the installed runtime and are loaded into context — workflows on every command, agents on every subagent dispatch — so their byte size is a real cost. Two sibling guards (tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs) keep that cost from creeping up invisibly, sharing one byte-counter (measureMdFiles). Growth is caught by two independent layers:

Layer What it does Where
Differential attribution size ratchet (primary, #2724 / ADR-2719 §4) The same computed-attribution check that replaced the golden-install-parity fixtures also reports growth in any gsd-core/workflows/*.md or agents/gsd-*.md file, with the exact byte delta, comparing PR HEAD against next. Unacknowledged growth is a hard failure; shrinkage needs no acknowledgment. No committed snapshot — nothing to regenerate by hand. tests/emitted-attribution.test.cjs (real-tree test) via tests/helpers/emitted-diff.cjs
Loose tier hard caps (backstop) Absolute outer red lines per tier — workflows: XL ≤ 98304, LARGE ≤ 61440, DEFAULT ≤ 40960 bytes; agents: XL ≤ 57344, LARGE ≤ 49152, DEFAULT ≤ 24576 bytes. A cap is never raised when a file approaches it: crossing it means extract, not bump. Independent of the ratchet above — unaffected by #2724. XL/LARGE/DEFAULT_CAP in each guard file
Headroom census + reserved margin (visibility, #4261) Every run prints each capped file's remaining bytes and percentage used — green runs included — sorted least-headroom-first, and appends a table of the files past a 95% reserved margin to the GitHub job summary. The margin reports, it does not fail: a file at 96% is not broken, it is a file whose next contributor should extract before adding. Nothing here raises or relaxes a cap. buildHeadroomRows / marginFor in scripts/workflow-size.cjs

Why the census exists: each PR's CI measures only its own base plus its own diff, so two PRs that are individually under a cap can be jointly over it, and no run either of them produces can show that. The census does not solve that directly — measuring on the merge result would, and was deliberately left out of #4261's approved scope — but it makes the density that causes it legible before the collision, which a passing run previously did not. It also replaces the hand-written per-tier high-water comments in both guard files, which had gone stale by several kilobytes and were themselves the reason the shrinking margin went unnoticed.

discuss-phase.md additionally has a thin-dispatcher target of < 32000 bytes (the discuss-phase progressive-disclosure split, #717). A net-new agent is DEFAULT-tier and already bounded by the DEFAULT cap — no separate new-agent cap is needed. (This tier-cap machinery is distinct from the separate 45 KB-char extraction-evidence threshold on gsd-planner enforced by tests/planner-decomposition.test.cjs — that one proves mode sections were extracted; this one bounds total agent bytes.)

How-to: a workflow or agent grew and CI is red

The differential attribution check reports the file and the byte delta. To resolve:

  1. Justify the growth in your PR (a sentence in the description is enough) — the acknowledgment entry (below) is the review record that the larger size was a deliberate, seen decision, not silent drift.

  2. Add an acknowledgment trailer to one of your own commits (ADR-3942), naming the file and the reason, per CONTRIBUTING.md's "Editing shipped content" section and CONTEXT.md's ### Emitted Artifact Provenance entry:

    Emitted-Drift-Ack-Growth: explore.md — new dispatch section, reasoning ships with the block
    

    Growth keys on the bare filename as it appears under gsd-core/workflows/ or agents/; an unattributable hash ripple uses Emitted-Drift-Ack-Hash: and keys on the emitted path (which always contains a /). The two are separate namespaces — a growth trailer will not excuse a hash ripple, and the failure output says which one applies. The trailer is the review record that the larger size was a deliberate, seen decision.

    If you need to change an acknowledgment, amend the commit carrying it. That is deliberate: the trailer cannot drift out of sync with the diff it explains, because changing either changes the sha and re-runs the gate.

  3. There is nothing to clean up afterwards. The trailer is read from git log $(git merge-base <base> HEAD)..HEAD — your commits and no others — so once your PR merges it is out of range by construction. It never becomes "spent", it owns no shared key space, it cannot conflict with anyone else's, and no sweeper has to delete it. That is the whole reason ADR-3942 moved the acknowledgment off the working tree.

  4. Or shrink it instead of acknowledging. Prefer extraction when the growth is incidental: for a workflow, move per-mode bodies to workflows/<name>/modes/, templates to workflows/<name>/templates/, and shared prose to gsd-core/references/; for an agent, lift shared boilerplate into gsd-core/references/ and @-reference it — then load it LAZILY. Do not convert them to eager @-required_reading includes: that shrinks the file's bytes without shrinking loaded context, so it games the guard while making the real cost worse. See workflows/discuss-phase/ for the progressive-disclosure pattern.

If a hard cap (not the ratchet) is what failed, an acknowledgment will not help — that is the signal to extract, per step 3.

Reference

Artifact Role
scripts/workflow-size.cjs Single source of truth — LF-normalized byte counter (lfByteCount) + generic measureMdFiles(dir, predicate) (backs both workflows and agents) + workflow enumeration (listWorkflowStems, measureWorkflows). Imported by both guards and by tests/helpers/emitted-runtime.cjs's currentSizes() so they can never measure differently.
tests/emitted-attribution.test.cjs + tests/helpers/emitted-diff.cjs The differential attribution check and its size ratchet (ADR-2719). The sole mechanism for both emitted-content propagation AND per-file size growth as of #2724.
Emitted-Drift-Ack-Hash: / Emitted-Drift-Ack-Growth: commit trailers (ADR-3942) The acknowledgment mechanism for unattributable emitted-content ripples and for size growth. Read from git log $(git merge-base <base> HEAD)..HEAD — no committed file, nothing to sweep; a merged trailer is out of range by construction.
npm run regen:derived Runs every remaining generator in dependency order (build → registry → ADR index → capability matrix → inventory manifest → manifest versions → tests/fixtures/install-tree/*.json).
tests/workflow-size-budget.test.cjs The workflow tier hard-cap guards, plus the discuss-phase progressive-disclosure checks.
tests/agent-size-budget.test.cjs The agent tier hard-cap guards (the agent analog).

tests/workflow-size-baseline.json, tests/agent-size-baseline.json, tests/fixtures/golden-install-parity/*.json, scripts/update-size-baseline.cjs (npm run size:baseline), and scripts/git-merge-regen-driver.cjs (npm run setup:merge-driver) are all removed by #2724: they were pure functions of the source tree, conflicted on every merge that touched them, and their functions are now served by the differential attribution check above. tests/fixtures/install-tree/*.json is the one artifact family that stays committed and normally-merged (ADR-2719 §7) — it conflicts on 0 of 7, its diffs are readable, and it preserves "the installer stopped shipping X" as a hard absolute failure with no attribution reasoning involved. Regenerate it with npm run gen:install-tree (folded into npm run regen:derived).

The QA smell ratchet

tests/loop-walk.qa.test.cjs (the qa suite) is the QA-walk harness's own self-test. Separately, scripts/qa-smell-ratchet.cjs drives that same harness end to end against the real gsd-tools binary and turns its findings into a CI gate — run it with npm run lint:qa-smells.

The gate has three independent inputs. The harness's oracles (tests/qa/oracles.cjs) supply the first two:

  • A violation is the engine breaking a documented contract. It always fails the build — baseline or no baseline, acknowledged or not.
  • A smell is legal-but-questionable behavior. A smell never fails a build on its own merits. What fails is the absence of a decision about it: an unacknowledged NEW smell, or a STALE entry in tests/qa/smell-baseline.json (one that stopped firing — the baseline is shrink-only, so a fixed or changed scenario must be pruned, not left behind).

The third comes from the scenarios themselves, not from an oracle:

  • A scenario expectation failure is a step's declared expect not holding — the scenario asserted percent: 100 and the engine returned something else. Like a violation, it is never acknowledgeable: it carries no fingerprint, so there is no key to put in a baseline entry or an ack fragment. Fix the engine, or correct the expectation.

Expectation failures were invisible to the gate until #3597: buildReport counted them in totals.violations while collectFindings read only oracle violations, so a failing scenario printed 0 violations and exited 0. The multi-workstream scenario failed on every CI run for three weeks without reddening a build. The invariant that keeps the two honest — asserted in tests/loop-walk.qa.test.cjs — is:

collectFindings().violations.length
  + collectFindings().expectationFailures.length
  === report.totals.violations

Note also that scripts/qa-smell-ratchet.cjs only invokes its own main() under require.main === module. That guard is what lets the QA suite require() the script to test collectFindings without kicking off a real 20-scenario walk as an import side effect — the reason the gate's own logic had no test before #3597.

Every smell must terminate in exactly one of TWO states — there is no third "accepted with a good explanation" state:

  1. REAL — an assigned defect. File it, then acknowledge the smell with an entry (baseline entry or tests/qa/smell-acks/ fragment) carrying that issue number.
  2. FALSE POSITIVE — the oracle itself is wrong. Fix the oracle (tests/qa/oracles.cjs) so it stops firing. It is NEVER baselined.

When the ratchet reports a NEW smell, there are exactly two legitimate responses — fix the detector, or file a defect and cite its issue number:

  1. Fix the underlying behavior (or the oracle, if it's a false positive) so the smell stops firing.
  2. File a defect and acknowledge it by adding a fragment under tests/qa/smell-acks/ — the ratchet's failure output prints a paste-ready skeleton naming the required key, id, scenario, and issue fields. issue MUST be a positive integer naming the tracking issue; a free-text reason may accompany it as an optional human note but can NEVER substitute for issue — "write an explanation" is not a way to acknowledge a smell. See tests/qa/smell-acks/README.md for the full shape and lifecycle.

Run node scripts/qa-smell-ratchet.cjs --update to regenerate tests/qa/smell-baseline.json from the current run, folding in any acked fragments and pruning stale entries. --update never invents an issue number: a genuinely new smell is written with issue: null and a TODO reason, and the very next plain (non---update) run REJECTS that entry — forcing a human to triage it before it can ship. The baseline only ever shrinks: growth happens by adding an acknowledgment carrying a real issue number (a reviewable diff), never by widening the generator's tolerance and never by prose alone.

Running suites locally

npm test                    # everything (backcompat — same as before)
npm run test:unit           # only unit
npm run test:integration    # only integration
npm run test:install        # only install
npm run test:security       # only security
npm run test:slow           # only slow

npm run test:coverage       # backcompat — coverage over EVERY test
npm run test:coverage:unit  # fast coverage signal — only unit suite
npm run test:coverage:all   # alias for test:coverage

Direct harness invocation also works:

node scripts/run-tests.cjs --suite security
node scripts/run-tests.cjs --suite=security
node scripts/run-tests.cjs --files "tests/command-contract.test.cjs tests/core.test.cjs"
node scripts/run-tests.cjs --files-from .ci-selected-tests.txt

npm run test:affected (scripts/run-affected-tests.cjs) is a local-only convenience that selects tests via the require() dependency graph of your working-tree diff. CI does not use it — CI selection is the rule table in scripts/ci-test-scope.cjs, which is the authoritative mapping. If the two disagree, trust (and fix) the rule table.

Unknown suites exit non-zero with the list of valid suites. Empty suites (e.g. --suite security before any security-tagged file exists) exit 0 with a no tests in suite "..." notice on stderr so CI lanes don't go red while a suite is being populated.

The live-config hermeticity guard

Every run-tests.cjs invocation snapshots GSD's own install footprint in each live runtime config directory before the suite and re-checks it afterwards. It exists because the failure it catches is silent by construction: a test that resolves a config directory from the ambient environment instead of a sandbox writes into your real ~/.claude (or $GSD_HOME/.gsd, or a Kimi config.toml), and nothing reports it. CI cannot catch this class at all — CI never has CLAUDE_CONFIG_DIR and friends set.

The guard watches only what GSD unambiguously owns — its top-level install footprint plus gsd--prefixed children of directories shared with the host agent — never whole config roots, because a host agent legitimately writing history.jsonl mid-run would make the guard cry wolf, and a guard that cries wolf gets switched off.

Two environment variables control it:

Variable Effect
GSD_STRICT_LIVE_CONFIG_GUARD=1 A detected write fails the run. Set on the Linux/macOS lanes of every CI job that runs the suite.
GSD_SKIP_LIVE_CONFIG_GUARD=1 Skips the check entirely.

Unset, the guard reports and does not fail — deliberately, not timidly. On its first CI run it surfaced pre-existing leaks on the Windows lane, where os.homedir() reads USERPROFILE and ~190 test sites sandbox HOME alone. Those are real and worth fixing, but they are a different defect class, and a brand-new gate that instantly reds an unrelated lane gets reverted rather than obeyed. Windows lanes therefore stay report-only until that sweep lands; this repo has the pattern already, in the local/no-source-grep ESLint rule that shipped at warn and was promoted to error after its cleanup (ADR 452).

GSD_SKIP_LIVE_CONFIG_GUARD is a bypass on a safety check, so it is documented here rather than left to be discovered in the source: an undocumented bypass is one people eventually set without knowing what they turned off. If you need it routinely, that is a bug report, not a workflow.

Reported paths are labelled CREATED, MODIFIED, DELETED, or UNVERIFIED. UNVERIFIED means a scan bound was hit and the path could not be attested either way — it is never the same as clean.

The mergeability preflight

Before any of the matrix below is provisioned, every pull_request compute lane waits on one shared gate: PR mergeability, the reusable workflow .github/workflows/pr-mergeable-preflight.yml. A pull request with a merge conflict runs no CI at all until the conflict is resolved (#3833).

Reference

Verdict When Job result Effect on the pipeline
MERGEABLE GitHub reported mergeable: true success everything runs as normal
CONFLICTED GitHub reported mergeable: false failure every gated job is skipped; Required tests and Stryker mutation score report red
INDETERMINATE mergeability still unknown after the retry budget, or the API read failed success (fails open) everything runs as normal; a ::warning:: is emitted
SKIPPED_NOT_A_PR the event is not pull_request (push, workflow_dispatch, a release.yml call into install-smoke.yml) success everything runs as normal; zero API calls
INDETERMINATE (bootstrap) scripts/ci-pr-mergeability.cjs is absent at the base sha success (fails open) everything runs as normal; a ::warning:: names the cause

The bootstrap row is a consequence of the checkout being pinned to the base sha: the preflight runs the script as it exists on the base branch, so the script is absent on the pull request that introduces it, and on any branch whose base predates it. Absent is not "conflicted" — it is one more thing the gate cannot determine, so it takes the same fail-open path. The arm is self-healing and never fires again once the script is on the base branch; a test pins it in place so a future reader does not mistake it for dead code.

The verdict comes from scripts/ci-pr-mergeability.cjs, which polls GET /repos/{owner}/{repo}/pulls/{number} until mergeable is non-null. GitHub computes mergeability in an asynchronous background job and returns null while it is in progress, so a single cold read is never authoritative — the poll is required, not defensive.

Two properties are load-bearing and easy to break:

  • Only mergeable === true / === false decides the verdict. null is falsy, so a truthiness test (if (!mergeable)) would classify every cold read as a conflict and red every PR in the repo.
  • mergeable_state never decides the verdict. Its blocked value means "required checks have not passed", which is true while this very job is running — gating on it self-deadlocks the pipeline. It is read for the failure message only.

Gated: test.yml (lint-tests, test, test-inert, test-conformance, coverage-gate, qa-loop-walk, required-tests), install-smoke.yml, mutation.yml, security-scan.yml, docs-required.yml, changeset-required.yml, default-flip-documentation.yml, branch-naming.yml.

Deliberately not gated: the pull_request_target policy and security lanes — auto-close-unsolicited-PRs, close-draft-PRs, the target/title/template validators, issue-link, and unauthorized-approval dismissal. Gating those would let a conflicted drive-by PR evade auto-close, so they keep running on a conflicted PR and a test asserts they are never wired to the preflight.

What it deliberately does not do

It applies no label. The gate does not add or remove needs-review: merge-conflict. Eight caller workflows each invoke the preflight, so eight jobs would race to add-or-remove one label on every PR event, and it would force pull-requests: write into eight lanes that today hold contents: read. needs-review: merge-conflict stays a maintainer triage label applied during PR sweeps; the red check and its annotation are the machine signal.

It changes no branch protection. .github/rulesets/main-protection.json is untouched and needs no new required context — GitHub natively refuses to merge a pull request with conflicts, so the ruleset already blocks it. (Adding the check as required before the workflow exists on the base branch would deadlock every open PR until it landed.)

What this does not replace

scripts/ci-rebase-check.cjs is unchanged and still runs inside each matrix job, including its #2472 base-sha pin. The preflight is an early-exit optimization on GitHub's asynchronously-computed view; the in-job merge remains authoritative and covers the case where the base advances mid-run. The gate is never a safety property — every path except an explicit mergeable: false fails open, so an API outage simply restores the pre-#3833 behavior.

If your PR shows a red PR mergeability check

The annotation names the base branch and the remedy. There is one step:

git fetch origin && git rebase origin/next && git push --force-with-lease

Resolve the conflicts the rebase reports, then push. The next synchronize event re-runs the preflight and the full pipeline comes back. (Because the sequence is a single command, this has no separate how-to page — see CONTRIBUTING.md → Where Do I Open My PR? for the branching model that makes the rebase necessary in the first place.)

Note the interaction with the sha-bound pass marker: a rebase changes your HEAD sha, which invalidates any prior remote-runner verification. Rebase last.

CI matrix

The Tests workflow (.github/workflows/test.yml) runs every PR through a scope computed by scripts/ci-test-scope.cjs's classify(), which sets two flags — product_changed and full_matrix — from the changed-file list.

All lanes run on Node 24 — the engines.node floor (>=24.0.0) and the only supported runtime.

Job Lanes Gated on Purpose
test ubuntu-latest (1 targeted + 3-shard full) product_changed == 'true' The default, always-scoped PR signal — the full unit/integration/security suites run once, sharded, on Linux. Linux only: its three scope: windows shards were deleted in #4641 (ADR-4641), which found them a second, redundant Windows selector alongside test-conformance
test-inert ubuntu-latest code_changed == 'true' && product_changed != 'true' A lightweight lane for PRs that touch only administrative/policy workflow files (code changed, but nothing that needs the real matrix)
test-conformance windows-latest (3-shard) + macos-latest (unsharded) code_changed == 'true' && full_matrix == 'true' Runs only the platform-conformance-tier file list (scripts/lib/platform-conformance-tier.generated.cjs, epic #4589 Phase 2/#4591) on real Windows/macOS — since #4641 the sole Windows and macOS selector in CI, not merely the sole gating one. Retired the parallel legacy full-suite matrix in #4603; #4641 removed the second Windows selector in test and narrowed the tier from 548 to 266 of 932 eligible unit-suite files (58.8% → 28.5%, measured 2026-09-11; the absolute counts track next's test count, the percentages are what the ceiling test binds on).
coverage-gate ubuntu-latest product_changed == 'true' && test.result == 'success' Merges every test shard's coverage dumps and evaluates the threshold once (sharding moved this out of the test job itself — #2952)
qa-loop-walk ubuntu-latest product_changed == 'true' The QA smell-ratchet scenario walk (see "The QA smell ratchet" below)
required-tests ubuntu-latest always() Aggregates every job above into the one branch-protection-required check

full_matrix fires on any changed tests/**/*.test.cjs file unconditionally (restored by #4421 after #962's narrowing let a real regression through undetected), plus a handful of curated RULES (workflow/installer/hooks/env-gate changes) — see scripts/ci-test-scope.cjs's own classify() for the exact, current rule set; this doc intentionally does not restate it in full, to avoid drifting out of sync with it.

Coverage is evaluated by the dedicated coverage-gate job (moved out of the test job by #2952, once sharding meant no single test runner saw the whole picture) and stays single-lane (Ubuntu / Node 24 only) because multiplying coverage across OS/runtime lanes adds cost without improving the threshold signal. Note the gate's deliberate blind spot: it measures gsd-core/bin/lib/*.cjs only — scripts/, hooks/, and bin/ are unenforced, and stryker.config.mjs additionally excludes ~48% of lib lines from mutation testing (see the UNMUTATED list there). Widening either gate is tracked work, not an accident to "fix" silently by raising thresholds.

To inspect the scope locally:

npm run ci:test-scope -- --files "commands/gsd/plan-phase.md"
node scripts/ci-test-scope.cjs --base origin/next --head HEAD

Chunk packing and the test timing table

scripts/run-tests.cjs does not hand the whole selected file list to one node --test process. It packs the files into chunks, each spawned separately, because Windows caps a command line at 32,767 characters and because each chunk gets its own 600s timeout (RUN_TESTS_CHUNK_TIMEOUT_MS) and a fresh process, which bounds memory pressure.

How files are distributed across those chunks decides whether the slowest chunk sits near that timeout while the others idle. The packer weights each file by its measured duration, read from tests/test-timings.json, and places files with LPT (longest-processing-time-first: heaviest file first, each into the currently lightest chunk). Before #2456 the weight was guessed from the filename, which mis-ranked files badly enough that the slowest chunk ran ~3.9x the lightest.

Reference

Knob Default Meaning
RUN_TESTS_MAX_FILES_PER_CHUNK 60 Per-chunk weight budget. Weights are normalized so an average-cost file weighs 1, so this still reads as "about 60 average files".
RUN_TESTS_MAX_CMDLINE_CHARS 28000 argv ceiling per chunk, with headroom under the Windows 32,767 limit.
RUN_TESTS_TIMINGS_FILE tests/test-timings.json Path to the timing table. Tests override it to inject a synthetic cost profile.
RUN_TESTS_CHUNK_TIMEOUT_MS 600000 Per-chunk timeout.

The timing table is advisory and deliberately un-gated. There is no --check mode and no CI lint that fails on staleness, because timing data legitimately varies run to run. A file missing from the table falls back to the table's median weight, and a missing or unparseable table falls back to uniform weight — so drift costs chunk balance, never a red build. A count-based floor additionally guarantees the packer never produces fewer chunks than plain count-based packing would, so a badly stale table cannot collapse the suite into a few fat chunks.

CI job timeout budgets: report + near-cap warning (#4036)

Every matrixed job — test in test.yml, mutate in mutation.yml, smoke in install-smoke.yml — declares a timeout-minutes cap. tests/ci-test-job-timeout-budget.test.cjs enforces that each checked-in cap stays at least the headroom factor (1.5x) above a documented, hand-measured cost for that job — now all three of the jobs above, not just test/coverage-gate/test-inert as before.

test-conformance in test.yml (#4591, epic #4589 Phase 2) runs the scripts/lib/platform-conformance-tier.generated.cjs file list on windows-latest (sharded three ways) and macos-latest (unsharded) — the sole gating signal for real-OS coverage (#4603 retired the parallel legacy full-matrix safety-net job). It also declares a timeout-minutes cap and runs the same in-job near-cap check described below, but it has no LANE_COSTS entry in tests/ci-test-job-timeout-budget.test.cjs yet — no real, completed (non-cancelled) per-shard measurement exists — so the headroom-factor gate does not cover it until one lands.

Two runtime mechanisms sit on top of that static gate, both new in #4036:

  • In-job near-cap check (scripts/ci-check-job-near-cap.cjs) — the last step of each of the three jobs computes elapsed-vs-cap from a start-time marker recorded as that job's first step. At >=90% of budget it emits a ::warning:: annotation (visible in the PR Checks UI) and a $GITHUB_STEP_SUMMARY block. Advisory only — it never fails the job. Known limit: it cannot fire for a job actually killed by the timeout, since a killed job never reaches its last step. That case is caught by the second mechanism instead.
  • Scheduled trending report (.github/workflows/ci-timeout-report.yml, scripts/ci-timeout-report.cjs) — runs daily and on workflow_dispatch. It polls GitHub's Actions REST API for recently completed jobs across test.yml, mutation.yml, and install-smoke.yml, resolves each job's declared cap (a literal timeout-minutes for test/test-conformance/smoke, or scripts/mutation-matrix.cjs's COVERED[<module>].timeoutMinutes for mutate's per-module shards), and appends any new (runId, jobName) records to tests/ci-timeout-budget-history.jsonl. Unlike the in-job check, this also catches jobs killed by an actual timeout breach — GitHub's Jobs API still reports started_at/completed_at for a cancelled job. Each run opens a small, data-only PR carrying that run's new rows, since next is a protected branch and nothing pushes to it directly — the same constraint auto-backmerge.yml already works within.

This does not retune any timeout-minutes value, rebalance shard composition, or trim what runs in shard 1 — those stay maintainer policy calls made from the accumulated history, not something either mechanism decides on its own.

How-to: regenerate the timing table

Regenerate when the suite's cost profile has visibly drifted — after adding or removing expensive tests, not on a schedule. The input is a node:test reporter event stream from a gsd-test run:

node scripts/gen-test-timings.cjs \
  ~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node24.jsonl \
  ~/.local/state/gsd-test/runs/<run-id>/test-events-macos-node24.jsonl

Pass every lane you have. A file's recorded time is the max across the supplied streams, not the mean: the packer exists to keep the slowest lane's slowest chunk away from the timeout, so the conservative bound is the right one. Keys are sorted so a regeneration diff shows only the files whose cost moved.

Best practices for forward-compat (Node 24/26)

  • Use process.execPath when spawning Node in tests so each matrix lane exercises the lane's Node version.
  • Avoid stack-trace or error-message prose assertions. Assert err.code, structured JSON fields, or enums — Node minor releases routinely tweak error wording.
  • Prefer node:test, node:assert/strict, and node:test mocks. No external test frameworks.
  • Coverage uses c8 and propagates NODE_V8_COVERAGE through the harness's child process.

Test strategy: #443 effort + fast_mode engine

Feature: unified cross-provider effort and fast_mode knobs (issue #443). Test files: tests/model-resolver.test.cjs (unit), tests/model-resolver.test.cjs (integration).

Testing pyramid

Layer File What it covers
Unit feat-443-effort-fast-mode.test.cjs Pure logic: cascade rules, clamping, escalation math, malformed config handling, schema key validation. No CLI subprocess.
Integration feat-443-effort-fast-mode.integration.test.cjs Architecture-level invariants: cross-provider validity, totality across the 33-agent registry, CLI JSON contract, config round-trip, fast-mode honesty. Real subprocesses via runGsdTools.
E2E (pending) (not yet wired) Propagation layer: effort frontmatter / CLAUDE_CODE_EFFORT_LEVEL env actually reaching a spawned Claude Code subagent. See "Gaps" below.

Architectural invariants

Each invariant exists to prevent a specific class of production failure.

(a) Cross-provider validity

What: renderEffortForRuntime(runtime, universalEffort).value must always be a member of the runtime's real provider enum. Ground-truth enums are defined as local constants in the test — not sourced from the implementation.

PROVIDER_EFFORT_ENUMS = {
  claude: Set { 'low', 'medium', 'high', 'xhigh', 'max' }   // Anthropic output_config.effort
  codex:  Set { 'minimal', 'low', 'medium', 'high', 'xhigh' } // OpenAI model_reasoning_effort
}

Why: Passing a value outside these sets results in a 400 from the real API. The clamping logic (max -> xhigh for codex; minimal -> low for claude) must hold for every cell of the VALID_EFFORTS × runtimes matrix.

(b) Param/channel contract

What: Each runtime exposes a stable param string (the native API field name) and channel (how the value is propagated). Unknown runtimes return param: null, channel: null and pass the effort value through unchanged.

Why: Callers read .param to construct the dispatch payload. A regression here would silently drop effort from subagent invocations.

(c) Resolve-execution JSON contract

What: The gsd-tools resolve-execution <agent> command emits a JSON object with all eight keys present and typed correctly: model (string), profile (string), effort (VALID_EFFORTS member), effort_rendered (string), effort_param (string|null), effort_propagation (string|null), fast_mode (boolean), fast_mode_supported (boolean).

Why: Orchestrators and workflow dispatchers parse this JSON. A missing or mistyped field silently breaks downstream consumers.

(d) Totality across the real registry

What: For every agent in the 33-agent registry, resolveEffortInternal returns a VALID_EFFORTS member (never undefined/null), resolveFastModeInternal returns a strict boolean, and renderEffortForRuntime('claude', effort) stays within the claude provider enum.

Why: A catalog addition that introduces a missing routingTier mapping would otherwise produce undefined and propagate silently.

(e) Fast-mode honesty invariant

What: When the runtime is claude, fast_mode_supported in resolve-execution output is always false, regardless of the fast_mode config. RUNTIMES_WITH_FAST_MODE contains only 'api'.

Why: Claude Code's /fast toggle is session-level only. Emitting fast_mode: true as frontmatter on a Claude subagent is a silent no-op. Advertising fast_mode_supported: true for claude would cause orchestrators to believe the knob was wired when it is not.

(f) Precedence first-valid-wins

What: Both effort and fast_mode use a layered cascade. The test table covers all four effort layers (invocation override → agent_overrides → routing_tier_defaults → default) and all five fast_mode layers, including the case where an invalid value at a higher layer correctly falls through.

Why: Silent precedence bugs (e.g., a numeric value in agent_overrides not being rejected) would override intentional user config.

(g) Dynamic-routing composition

What: resolveEffortForTier escalates effort by attempt number independently of the model tier mapping. The test verifies the effort ladder (low -> medium -> high -> xhigh -> max), the max clamp, the max_escalations cap, and that escalate_on_failure: false suppresses escalation entirely.

Why: Effort escalation and model escalation share configuration (dynamic_routing) but must operate independently; coupling them would cause over-escalation or under-escalation.

(h) Config-tooling round-trip

What: gsd-tools config-set accepts all new key namespaces (effort.default, effort.routing_tier_defaults.<tier>, effort.agent_overrides.<agent>, fast_mode.enabled, fast_mode.routing_tier_defaults.<tier>, fast_mode.agent_overrides.<agent>) without an "Unknown config key" error, and values set via config-set are reflected in resolve-execution output.

Why: The schema validation gate (VALID_CONFIG_KEYS + DYNAMIC_KEY_PATTERNS) is separate from the resolver logic. A key missing from the schema would produce a silent write failure and appear as a bug only at runtime.

Coverage targets

Suite Target
Unit Every cascade rule, every fallthrough, every clamp. All function branches in resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, renderEffortForRuntime.
Integration All 8 architectural invariants. All 33 registered agents. All 6 provider × effort combinations for the valid-enum check. Full config-set key namespace.

Gaps / not yet covered

E2E orchestrator-spawn-propagation layer (pending follow-up wiring): The integration tests verify that GSD resolves and renders effort values correctly. They do NOT verify that the rendered values actually reach a spawned Claude Code or Codex subagent at runtime. Specifically uncovered:

  • CLAUDE_CODE_EFFORT_LEVEL env var being set and read by a spawned claude subprocess
  • output_config.effort frontmatter key surviving the AGENTS.md template substitution
  • model_reasoning_effort field surviving serialization into a Codex API request body
  • Fast-mode speed: "fast" field reaching an api-runtime request when fast_mode_supported: true

These require spawning real subagents (or stubs thereof) and asserting on the process environment / request payload — a scope that belongs in a future E2E suite under *.slow.test.cjs or dedicated fixture-driven integration work.