* test(#4733): pin the cap, unknown-file weight, and isolation rules Failing-first coverage for the three defects that let a Windows conformance chunk be killed at the 600s per-chunk backstop with zero failing tests. The previous boundary rows were VACUOUS: they asserted literal arithmetic (21 * 18122 <= 400000) that cannot fail, and in doing so masked a shipped win32 cap of 23 -- a value that violates the very inequality they claimed to pin. These rows constrain defaultMaxFilesPerChunk itself, from both sides, so the shipped value is a derived maximum rather than a magic number. A second vacuous row was caught by review and removed: it recomputed the isolated set from the function under test using the identical predicate, so it was empty by construction. It is replaced by an exact deepEqual against the expected basenames, a cross-platform identity row, dynamism rows in both directions, an inclusive boundary triplet, and invalid-threshold throw rows. The cross-platform identity row is the regression guard for a threshold that was briefly anchored to the per-platform file-COUNT cap; it fails if isolation ever becomes platform-dependent again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4733): derive the win32 cap, isolation bar, and unknown weight A Windows conformance chunk was killed at the 600000ms per-chunk backstop with no test having failed, taking next red. Three compounding defects. The win32 cap of 40 permitted 40 * 18122 = 724880ms against a 600000ms backstop -- 121% of it -- so two rounds of budget-tuning could not hold. The cap is now derived: 22 is the largest value satisfying cap * 18122 <= 400000. The budget is 400000, not the raw backstop, because the chunk that died summed to only ~348328ms of per-file time -- a per-chunk overhead gap of at least 1.72x that no per-file table models. A file absent from the timings table was priced at medianWeight. The table is skewed 18.8x, so an unknown weighed 0.0533 -- 19x cheaper than average, and measured 17.5x under its real cost. Unknowns are now priced at the mean. ISOLATED_HEAVY_FILES was a static Set, stale by construction. Isolation is now derived from an absolute ms bar (0.3 * 400000 = 120000ms) converted to weight units via the live table's mean, so a file that gets heavy is isolated automatically instead of waiting for someone to edit a list. Review caught that an earlier cut anchored that bar to the per-platform file-COUNT cap -- a category error, count vs weight, which silently returned seven of the historical eight files to the shared pool on linux/darwin. Since macOS runs the full matrix only after merge, that would have planted a red next no PR could catch. The bar is absolute and platform-independent. Also from review: isolation no longer requires unit-suite membership, so fragment-single-edit-propagation.install.test.cjs -- 575000ms, 96% of the backstop in one file -- is eligible; partitionIsolatedFiles throws on a non-finite or non-positive threshold instead of silently isolating nothing; and stale per-shard figures no test pinned are removed rather than recomputed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4733): backfill changeset pr number --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
44 KiB
Testing Suites
This project's tests/ directory uses filename suffix markers to group tests into named suites. The harness scripts/run-tests.cjs filters by suite when given --suite <name>. Without a flag it runs every *.test.cjs file (the historical default — unchanged).
Tracked by issue #3597.
Suites
| Suite | Filename pattern | What goes here |
|---|---|---|
unit |
*.test.cjs (no other marker) |
Default fast lane. Pure logic, no network, no external processes beyond gsd-tools. Most tests live here. |
integration |
*.integration.test.cjs |
Cross-module flows: full installer end-to-end, multi-tool orchestration, anything that crosses two or more bin entry points. |
install |
*.install.test.cjs |
Tests that perform a real install/uninstall against a sandbox project. Slower; PR CI skips these on PRs and runs them on main push only. |
security |
*.security.test.cjs |
Adversarial input, prompt-injection guards, fixture-driven hostile-payload sweeps. |
slow |
*.slow.test.cjs |
Anything that routinely takes >5s wall-clock or holds significant memory. |
qa |
*.qa.test.cjs |
End-to-end walks that drive the real gsd-tools binary across multiple loop steps against one accumulating temp project, with invariant oracles after every step. Slower than unit; excluded from the fast lane. |
all |
(any) | Explicit alias for "no filter". Equivalent to running with no --suite flag. |
How to place a new test
- Pick the most specific bucket above.
- Name the file with the matching suffix:
tests/<feature>.<suite>.test.cjs. - If unsure, leave the suffix off — the file lands in
unit, the default fast lane.
Examples:
tests/agent-frontmatter.test.cjs—unittests/prompt-injection-guards.security.test.cjs—securitytests/installer-end-to-end.install.test.cjs—installtests/sdk-mutation-stress.slow.test.cjs—slowtests/loop-walk.qa.test.cjs—qa
The suite-suffix convention was chosen over a directory layout (tests/security/) so the 545+ existing test files don't need to move. Existing files all classify as unit until someone explicitly retags them.
Regression tests
Do not create new top-level tests/bug-NNNN-*.test.cjs,
tests/fix-NNNN-*.test.cjs, or tests/issue-NNNN-*.test.cjs files. Add the
regression case to the owning module's main test file instead (e.g. a
describe('regressions') block in tests/<module>.test.cjs).
node --test spawns one child process per FILE, so file count — not test
count — is the unit of CI overhead, and it is worst on Windows lanes where
every spawn is Defender-scanned. The 2026-06 CI audit found 244 one-off
bug-* files (~38% of the suite). That population is grandfathered in
scripts/lint-regression-test-names.allowlist.json and enforced by an
identity ratchet (npm run lint:regression-names, part of npm run lint:ci),
which also bans fix-* and issue-* NNNN-prefixed files the same way:
- A new
bug-*,fix-*, orissue-*NNNN-prefixed file fails CI — fold it into the owning module's file. - Deleting/consolidating a grandfathered file requires pruning its allowlist entry, so the baseline only ever shrinks.
- Inherited drift (the failure names files your PR didn't add — e.g. the
base branch merged
bug-*/fix-*/issue-*files without feeding the allowlist, or you rebased and carried a pre-rebase allowlist): runnode scripts/lint-regression-test-names.cjs --updateand commit the regenerated allowlist. Snapshot artifacts like this allowlist (anddocs/INVENTORY.md) must be regenerated after rebasing, never carried through a rebase.
A folded suite may appear only once per host
When a standalone file is folded into its owning module's test file, the moved suite is wrapped in a self-contained block carrying a marker:
// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-376-claude-js-hook-gsd-rewriter.test.cjs — …
// ────────────────────────────────────────────────────────────────────────
{
const { describe: __foldDescribe } = require('node:test');
__foldDescribe("folded:bug-376-claude-js-hook-gsd-rewriter (…)", () => { … });
}
Because the block is self-contained, a second verbatim copy in the same host
parses, registers, and passes — twice. Nothing in a green suite reports it.
#3271 found 25 such copies
(~5,800 lines) across three install suites, all from a single stale-base
re-application during the consolidation epic. The cost is not only wasted CI on
every lane: it is a DEFECT.GENERATIVE-FIX trap, because a contributor fixing
one of those regressions edits the copy they found and leaves the other
asserting the old behavior, with the suite still green.
local/no-duplicate-fold-marker (eslint-rules/no-duplicate-fold-marker.cjs,
error under tests/**/*.cjs) reports the second and every later occurrence of a
folded:<marker> title in one file, naming the line the first occurrence sits
on. When it fires, delete the copy it points at — the two blocks are the
same suite, so the fix is removal, never an eslint-disable.
It keys on the whitespace-delimited token after folded:, which matters in both
directions. A narrower key that stops at . would collide
feat-443-effort-fast-mode.integration with feat-443-effort-fast-mode — two
genuinely distinct suites that coexist in tests/model-resolver.test.cjs. Keying
on the whole title instead would let a re-fold under a different batch label
slip through, which is exactly the shape #3271 took. Titles without a folded:
prefix, non-literal titles, and the same marker appearing in two different host
files are all left alone.
The ratchet deliberately covers only bug-*. Files named feat-NNNN-* /
enh-NNNN-* are feature test files — one (or one per suite) per feature is
the sanctioned layout (see the #443 strategy below), not a one-off regression
pattern. If issue-*/perf-* one-offs start accumulating the same way
bug-* did, extend the ratchet's regex and regenerate the allowlist.
Workflow & agent size budget
Tracked by issue #1074. Bytes (not lines) per #717; LF-normalized per #683.
Workflow files (gsd-core/workflows/*.md) and agent files (agents/gsd-*.md)
both ship in the installed runtime and are loaded into context — workflows on
every command, agents on every subagent dispatch — so their byte size is a real
cost. Two sibling guards (tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs) keep that cost from creeping up invisibly,
sharing one byte-counter (measureMdFiles). Growth is caught by two independent
layers:
| Layer | What it does | Where |
|---|---|---|
| Differential attribution size ratchet (primary, #2724 / ADR-2719 §4) | The same computed-attribution check that replaced the golden-install-parity fixtures also reports growth in any gsd-core/workflows/*.md or agents/gsd-*.md file, with the exact byte delta, comparing PR HEAD against next. Unacknowledged growth is a hard failure; shrinkage needs no acknowledgment. No committed snapshot — nothing to regenerate by hand. |
tests/emitted-attribution.test.cjs (real-tree test) via tests/helpers/emitted-diff.cjs |
| Loose tier hard caps (backstop) | Absolute outer red lines per tier — workflows: XL ≤ 98304, LARGE ≤ 61440, DEFAULT ≤ 40960 bytes; agents: XL ≤ 57344, LARGE ≤ 49152, DEFAULT ≤ 24576 bytes. A cap is never raised when a file approaches it: crossing it means extract, not bump. Independent of the ratchet above — unaffected by #2724. |
XL/LARGE/DEFAULT_CAP in each guard file |
| Headroom census + reserved margin (visibility, #4261) | Every run prints each capped file's remaining bytes and percentage used — green runs included — sorted least-headroom-first, and appends a table of the files past a 95% reserved margin to the GitHub job summary. The margin reports, it does not fail: a file at 96% is not broken, it is a file whose next contributor should extract before adding. Nothing here raises or relaxes a cap. | buildHeadroomRows / marginFor in scripts/workflow-size.cjs |
Why the census exists: each PR's CI measures only its own base plus its own diff, so two PRs that are individually under a cap can be jointly over it, and no run either of them produces can show that. The census does not solve that directly — measuring on the merge result would, and was deliberately left out of #4261's approved scope — but it makes the density that causes it legible before the collision, which a passing run previously did not. It also replaces the hand-written per-tier high-water comments in both guard files, which had gone stale by several kilobytes and were themselves the reason the shrinking margin went unnoticed.
discuss-phase.md additionally has a thin-dispatcher target of < 32000 bytes
(the discuss-phase progressive-disclosure split, #717). A net-new agent is
DEFAULT-tier and already bounded by the DEFAULT cap — no separate new-agent cap
is needed. (This tier-cap machinery is distinct from the separate 45 KB-char
extraction-evidence threshold on gsd-planner enforced by
tests/planner-decomposition.test.cjs — that one proves mode sections were
extracted; this one bounds total agent bytes.)
How-to: a workflow or agent grew and CI is red
The differential attribution check reports the file and the byte delta. To resolve:
-
Justify the growth in your PR (a sentence in the description is enough) — the acknowledgment entry (below) is the review record that the larger size was a deliberate, seen decision, not silent drift.
-
Add an acknowledgment trailer to one of your own commits (ADR-3942), naming the file and the reason, per
CONTRIBUTING.md's "Editing shipped content" section andCONTEXT.md's### Emitted Artifact Provenanceentry:Emitted-Drift-Ack-Growth: explore.md — new dispatch section, reasoning ships with the blockGrowth keys on the bare filename as it appears under
gsd-core/workflows/oragents/; an unattributable hash ripple usesEmitted-Drift-Ack-Hash:and keys on the emitted path (which always contains a/). The two are separate namespaces — a growth trailer will not excuse a hash ripple, and the failure output says which one applies. The trailer is the review record that the larger size was a deliberate, seen decision.If you need to change an acknowledgment, amend the commit carrying it. That is deliberate: the trailer cannot drift out of sync with the diff it explains, because changing either changes the sha and re-runs the gate.
-
There is nothing to clean up afterwards. The trailer is read from
git log $(git merge-base <base> HEAD)..HEAD— your commits and no others — so once your PR merges it is out of range by construction. It never becomes "spent", it owns no shared key space, it cannot conflict with anyone else's, and no sweeper has to delete it. That is the whole reason ADR-3942 moved the acknowledgment off the working tree. -
Or shrink it instead of acknowledging. Prefer extraction when the growth is incidental: for a workflow, move per-mode bodies to
workflows/<name>/modes/, templates toworkflows/<name>/templates/, and shared prose togsd-core/references/; for an agent, lift shared boilerplate intogsd-core/references/and@-reference it — then load it LAZILY. Do not convert them to eager@-required_readingincludes: that shrinks the file's bytes without shrinking loaded context, so it games the guard while making the real cost worse. Seeworkflows/discuss-phase/for the progressive-disclosure pattern.
If a hard cap (not the ratchet) is what failed, an acknowledgment will not help — that is the signal to extract, per step 3.
Reference
| Artifact | Role |
|---|---|
scripts/workflow-size.cjs |
Single source of truth — LF-normalized byte counter (lfByteCount) + generic measureMdFiles(dir, predicate) (backs both workflows and agents) + workflow enumeration (listWorkflowStems, measureWorkflows). Imported by both guards and by tests/helpers/emitted-runtime.cjs's currentSizes() so they can never measure differently. |
tests/emitted-attribution.test.cjs + tests/helpers/emitted-diff.cjs |
The differential attribution check and its size ratchet (ADR-2719). The sole mechanism for both emitted-content propagation AND per-file size growth as of #2724. |
Emitted-Drift-Ack-Hash: / Emitted-Drift-Ack-Growth: commit trailers (ADR-3942) |
The acknowledgment mechanism for unattributable emitted-content ripples and for size growth. Read from git log $(git merge-base <base> HEAD)..HEAD — no committed file, nothing to sweep; a merged trailer is out of range by construction. |
npm run regen:derived |
Runs every remaining generator in dependency order (build → registry → ADR index → capability matrix → inventory manifest → manifest versions → tests/fixtures/install-tree/*.json). |
tests/workflow-size-budget.test.cjs |
The workflow tier hard-cap guards, plus the discuss-phase progressive-disclosure checks. |
tests/agent-size-budget.test.cjs |
The agent tier hard-cap guards (the agent analog). |
tests/workflow-size-baseline.json, tests/agent-size-baseline.json,
tests/fixtures/golden-install-parity/*.json, scripts/update-size-baseline.cjs
(npm run size:baseline), and scripts/git-merge-regen-driver.cjs
(npm run setup:merge-driver) are all removed by
#2724: they were pure
functions of the source tree, conflicted on every merge that touched them, and
their functions are now served by the differential attribution check above.
tests/fixtures/install-tree/*.json is the one artifact family that stays
committed and normally-merged (ADR-2719 §7) — it conflicts on 0 of 7, its diffs
are readable, and it preserves "the installer stopped shipping X" as a hard
absolute failure with no attribution reasoning involved. Regenerate it with
npm run gen:install-tree (folded into npm run regen:derived).
The QA smell ratchet
tests/loop-walk.qa.test.cjs (the qa suite) is the QA-walk harness's own
self-test. Separately, scripts/qa-smell-ratchet.cjs drives that same harness
end to end against the real gsd-tools binary and turns its findings into a
CI gate — run it with npm run lint:qa-smells.
The gate has three independent inputs. The harness's oracles
(tests/qa/oracles.cjs) supply the first two:
- A violation is the engine breaking a documented contract. It always fails the build — baseline or no baseline, acknowledged or not.
- A smell is legal-but-questionable behavior. A smell never fails a
build on its own merits. What fails is the absence of a decision about
it: an unacknowledged NEW smell, or a STALE entry in
tests/qa/smell-baseline.json(one that stopped firing — the baseline is shrink-only, so a fixed or changed scenario must be pruned, not left behind).
The third comes from the scenarios themselves, not from an oracle:
- A scenario expectation failure is a step's declared
expectnot holding — the scenario assertedpercent: 100and the engine returned something else. Like a violation, it is never acknowledgeable: it carries no fingerprint, so there is nokeyto put in a baseline entry or an ack fragment. Fix the engine, or correct the expectation.
Expectation failures were invisible to the gate until
#3597: buildReport
counted them in totals.violations while collectFindings read only oracle
violations, so a failing scenario printed 0 violations and exited 0. The
multi-workstream scenario failed on every CI run for three weeks without
reddening a build. The invariant that keeps the two honest — asserted in
tests/loop-walk.qa.test.cjs — is:
collectFindings().violations.length
+ collectFindings().expectationFailures.length
=== report.totals.violations
Note also that scripts/qa-smell-ratchet.cjs only invokes its own main()
under require.main === module. That guard is what lets the QA suite
require() the script to test collectFindings without kicking off a real
20-scenario walk as an import side effect — the reason the gate's own logic
had no test before #3597.
Every smell must terminate in exactly one of TWO states — there is no third "accepted with a good explanation" state:
- REAL — an assigned defect. File it, then acknowledge the smell with an
entry (baseline entry or
tests/qa/smell-acks/fragment) carrying thatissuenumber. - FALSE POSITIVE — the oracle itself is wrong. Fix the oracle
(
tests/qa/oracles.cjs) so it stops firing. It is NEVER baselined.
When the ratchet reports a NEW smell, there are exactly two legitimate responses — fix the detector, or file a defect and cite its issue number:
- Fix the underlying behavior (or the oracle, if it's a false positive) so the smell stops firing.
- File a defect and acknowledge it by adding a fragment under
tests/qa/smell-acks/— the ratchet's failure output prints a paste-ready skeleton naming the requiredkey,id,scenario, andissuefields.issueMUST be a positive integer naming the tracking issue; a free-textreasonmay accompany it as an optional human note but can NEVER substitute forissue— "write an explanation" is not a way to acknowledge a smell. Seetests/qa/smell-acks/README.mdfor the full shape and lifecycle.
Run node scripts/qa-smell-ratchet.cjs --update to regenerate
tests/qa/smell-baseline.json from the current run, folding in any acked
fragments and pruning stale entries. --update never invents an issue
number: a genuinely new smell is written with issue: null and a TODO
reason, and the very next plain (non---update) run REJECTS that entry —
forcing a human to triage it before it can ship. The baseline only ever
shrinks: growth happens by adding an acknowledgment carrying a real issue
number (a reviewable diff), never by widening the generator's tolerance and
never by prose alone.
Running suites locally
npm test # everything (backcompat — same as before)
npm run test:unit # only unit
npm run test:integration # only integration
npm run test:install # only install
npm run test:security # only security
npm run test:slow # only slow
npm run test:coverage # backcompat — coverage over EVERY test
npm run test:coverage:unit # fast coverage signal — only unit suite
npm run test:coverage:all # alias for test:coverage
Direct harness invocation also works:
node scripts/run-tests.cjs --suite security
node scripts/run-tests.cjs --suite=security
node scripts/run-tests.cjs --files "tests/command-contract.test.cjs tests/core.test.cjs"
node scripts/run-tests.cjs --files-from .ci-selected-tests.txt
npm run test:affected (scripts/run-affected-tests.cjs) is a local-only
convenience that selects tests via the require() dependency graph of your
working-tree diff. CI does not use it — CI selection is the rule table in
scripts/ci-test-scope.cjs, which is the authoritative mapping. If the two
disagree, trust (and fix) the rule table.
Unknown suites exit non-zero with the list of valid suites. Empty suites (e.g. --suite security before any security-tagged file exists) exit 0 with a no tests in suite "..." notice on stderr so CI lanes don't go red while a suite is being populated.
The live-config hermeticity guard
Every run-tests.cjs invocation snapshots GSD's own install footprint in each
live runtime config directory before the suite and re-checks it afterwards. It
exists because the failure it catches is silent by construction: a test that
resolves a config directory from the ambient environment instead of a sandbox
writes into your real ~/.claude (or $GSD_HOME/.gsd, or a Kimi
config.toml), and nothing reports it. CI cannot catch this class at all —
CI never has CLAUDE_CONFIG_DIR and friends set.
The guard watches only what GSD unambiguously owns — its top-level install
footprint plus gsd--prefixed children of directories shared with the host
agent — never whole config roots, because a host agent legitimately writing
history.jsonl mid-run would make the guard cry wolf, and a guard that cries
wolf gets switched off.
Two environment variables control it:
| Variable | Effect |
|---|---|
GSD_STRICT_LIVE_CONFIG_GUARD=1 |
A detected write fails the run. Set on the Linux/macOS lanes of every CI job that runs the suite. |
GSD_SKIP_LIVE_CONFIG_GUARD=1 |
Skips the check entirely. |
Unset, the guard reports and does not fail — deliberately, not timidly. On
its first CI run it surfaced pre-existing leaks on the Windows lane, where
os.homedir() reads USERPROFILE and ~190 test sites sandbox HOME alone.
Those are real and worth fixing, but they are a different defect class, and a
brand-new gate that instantly reds an unrelated lane gets reverted rather than
obeyed. Windows lanes therefore stay report-only until that sweep lands; this
repo has the pattern already, in the local/no-source-grep ESLint rule that
shipped at warn and was promoted to error after its cleanup (ADR 452).
GSD_SKIP_LIVE_CONFIG_GUARD is a bypass on a safety check, so it is documented
here rather than left to be discovered in the source: an undocumented bypass is
one people eventually set without knowing what they turned off. If you need it
routinely, that is a bug report, not a workflow.
Reported paths are labelled CREATED, MODIFIED, DELETED, or UNVERIFIED.
UNVERIFIED means a scan bound was hit and the path could not be attested
either way — it is never the same as clean.
The mergeability preflight
Before any of the matrix below is provisioned, every pull_request compute lane
waits on one shared gate: PR mergeability, the reusable workflow
.github/workflows/pr-mergeable-preflight.yml. A pull request with a merge
conflict runs no CI at all until the conflict is resolved (#3833).
Reference
| Verdict | When | Job result | Effect on the pipeline |
|---|---|---|---|
MERGEABLE |
GitHub reported mergeable: true |
success | everything runs as normal |
CONFLICTED |
GitHub reported mergeable: false |
failure | every gated job is skipped; Required tests and Stryker mutation score report red |
INDETERMINATE |
mergeability still unknown after the retry budget, or the API read failed | success (fails open) | everything runs as normal; a ::warning:: is emitted |
SKIPPED_NOT_A_PR |
the event is not pull_request (push, workflow_dispatch, a release.yml call into install-smoke.yml) |
success | everything runs as normal; zero API calls |
INDETERMINATE (bootstrap) |
scripts/ci-pr-mergeability.cjs is absent at the base sha |
success (fails open) | everything runs as normal; a ::warning:: names the cause |
The bootstrap row is a consequence of the checkout being pinned to the base sha: the preflight runs the script as it exists on the base branch, so the script is absent on the pull request that introduces it, and on any branch whose base predates it. Absent is not "conflicted" — it is one more thing the gate cannot determine, so it takes the same fail-open path. The arm is self-healing and never fires again once the script is on the base branch; a test pins it in place so a future reader does not mistake it for dead code.
The verdict comes from scripts/ci-pr-mergeability.cjs, which polls
GET /repos/{owner}/{repo}/pulls/{number} until mergeable is non-null.
GitHub computes mergeability in an asynchronous background job and returns
null while it is in progress, so a single cold read is never authoritative —
the poll is required, not defensive.
Two properties are load-bearing and easy to break:
- Only
mergeable === true/=== falsedecides the verdict.nullis falsy, so a truthiness test (if (!mergeable)) would classify every cold read as a conflict and red every PR in the repo. mergeable_statenever decides the verdict. Itsblockedvalue means "required checks have not passed", which is true while this very job is running — gating on it self-deadlocks the pipeline. It is read for the failure message only.
Gated: test.yml (lint-tests, test, test-inert,
test-conformance, coverage-gate, qa-loop-walk, required-tests), install-smoke.yml,
mutation.yml, security-scan.yml, docs-required.yml,
changeset-required.yml, default-flip-documentation.yml, branch-naming.yml.
Deliberately not gated: the pull_request_target policy and security lanes —
auto-close-unsolicited-PRs, close-draft-PRs, the target/title/template
validators, issue-link, and unauthorized-approval dismissal. Gating those would
let a conflicted drive-by PR evade auto-close, so they keep running on a
conflicted PR and a test asserts they are never wired to the preflight.
What it deliberately does not do
It applies no label. The gate does not add or remove needs-review: merge-conflict.
Eight caller workflows each invoke the preflight, so eight jobs would race to
add-or-remove one label on every PR event, and it would force
pull-requests: write into eight lanes that today hold contents: read.
needs-review: merge-conflict stays a maintainer triage label applied during PR
sweeps; the red check and its annotation are the machine signal.
It changes no branch protection. .github/rulesets/main-protection.json is
untouched and needs no new required context — GitHub natively refuses to merge a
pull request with conflicts, so the ruleset already blocks it. (Adding the check
as required before the workflow exists on the base branch would deadlock every
open PR until it landed.)
What this does not replace
scripts/ci-rebase-check.cjs is unchanged and still runs inside each matrix job,
including its #2472 base-sha pin. The preflight is an early-exit optimization on
GitHub's asynchronously-computed view; the in-job merge remains authoritative and
covers the case where the base advances mid-run. The gate is never a safety
property — every path except an explicit mergeable: false fails open, so an
API outage simply restores the pre-#3833 behavior.
If your PR shows a red PR mergeability check
The annotation names the base branch and the remedy. There is one step:
git fetch origin && git rebase origin/next && git push --force-with-lease
Resolve the conflicts the rebase reports, then push. The next synchronize
event re-runs the preflight and the full pipeline comes back. (Because the
sequence is a single command, this has no separate how-to page — see
CONTRIBUTING.md → Where Do I Open My PR?
for the branching model that makes the rebase necessary in the first place.)
Note the interaction with the sha-bound pass marker: a rebase changes your HEAD sha, which invalidates any prior remote-runner verification. Rebase last.
CI matrix
The Tests workflow (.github/workflows/test.yml) runs every PR through a scope
computed by scripts/ci-test-scope.cjs's classify(), which sets two flags —
product_changed and full_matrix — from the changed-file list.
All lanes run on Node 24 — the engines.node floor (>=24.0.0) and the
only supported runtime.
| Job | Lanes | Gated on | Purpose |
|---|---|---|---|
test |
ubuntu-latest (1 targeted + 3-shard full) |
product_changed == 'true' |
The default, always-scoped PR signal — the full unit/integration/security suites run once, sharded, on Linux. Linux only: its three scope: windows shards were deleted in #4641 (ADR-4641), which found them a second, redundant Windows selector alongside test-conformance |
test-inert |
ubuntu-latest |
code_changed == 'true' && product_changed != 'true' |
A lightweight lane for PRs that touch only administrative/policy workflow files (code changed, but nothing that needs the real matrix) |
test-conformance |
windows-latest (3-shard) + macos-latest (unsharded) |
code_changed == 'true' && full_matrix == 'true' |
Runs only the platform-conformance-tier file list (scripts/lib/platform-conformance-tier.generated.cjs, epic #4589 Phase 2/#4591) on real Windows/macOS — since #4641 the sole Windows and macOS selector in CI, not merely the sole gating one. Retired the parallel legacy full-suite matrix in #4603; #4641 removed the second Windows selector in test and narrowed the tier from 548 to 266 of 932 eligible unit-suite files (58.8% → 28.5%, measured 2026-09-11; the absolute counts track next's test count, the percentages are what the ceiling test binds on). |
coverage-gate |
ubuntu-latest |
product_changed == 'true' && test.result == 'success' |
Merges every test shard's coverage dumps and evaluates the threshold once (sharding moved this out of the test job itself — #2952) |
qa-loop-walk |
ubuntu-latest |
product_changed == 'true' |
The QA smell-ratchet scenario walk (see "The QA smell ratchet" below) |
required-tests |
ubuntu-latest |
always() |
Aggregates every job above into the one branch-protection-required check |
full_matrix fires on any changed tests/**/*.test.cjs file unconditionally (restored
by #4421 after #962's narrowing let a real regression through undetected), plus a
handful of curated RULES (workflow/installer/hooks/env-gate changes) — see
scripts/ci-test-scope.cjs's own classify() for the exact, current rule set; this
doc intentionally does not restate it in full, to avoid drifting out of sync with it.
Coverage is evaluated by the dedicated coverage-gate job (moved out of the test
job by #2952, once sharding meant no single test runner saw the whole picture) and
stays single-lane (Ubuntu / Node 24 only) because multiplying coverage across
OS/runtime lanes adds cost without improving the threshold signal. Note the gate's
deliberate blind spot: it measures
gsd-core/bin/lib/*.cjs only — scripts/, hooks/, and bin/ are
unenforced, and stryker.config.mjs additionally excludes ~48% of lib lines
from mutation testing (see the UNMUTATED list there). Widening either gate is
tracked work, not an accident to "fix" silently by raising thresholds.
To inspect the scope locally:
npm run ci:test-scope -- --files "commands/gsd/plan-phase.md"
node scripts/ci-test-scope.cjs --base origin/next --head HEAD
Chunk packing and the test timing table
scripts/run-tests.cjs does not hand the whole selected file list to one
node --test process. It packs the files into chunks, each spawned
separately, because Windows caps a command line at 32,767 characters and because
each chunk gets its own 600s timeout (RUN_TESTS_CHUNK_TIMEOUT_MS) and a fresh
process, which bounds memory pressure.
How files are distributed across those chunks decides whether the slowest chunk
sits near that timeout while the others idle. The packer weights each file by its
measured duration, read from tests/test-timings.json, and places files with
LPT (longest-processing-time-first: heaviest file first, each into the currently
lightest chunk). Before #2456 the weight was guessed from the filename, which
mis-ranked files badly enough that the slowest chunk ran ~3.9x the lightest.
Reference
| Knob | Default | Meaning |
|---|---|---|
RUN_TESTS_MAX_FILES_PER_CHUNK |
60 (22 on win32) |
Per-chunk weight budget. Weights are normalized so an average-cost file weighs 1, so this still reads as "about 60 average files" (about 22 on win32). Windows gets a lower cap than Linux/macOS because the weight table's calibration does not transfer 1:1 to the Windows runner for install/subprocess-heavy work — see the derivation comment above DEFAULT_MAX_FILES_PER_CHUNK in scripts/run-tests.cjs. |
RUN_TESTS_MAX_CMDLINE_CHARS |
28000 |
argv ceiling per chunk, with headroom under the Windows 32,767 limit. |
RUN_TESTS_TIMINGS_FILE |
tests/test-timings.json |
Path to the timing table. Tests override it to inject a synthetic cost profile. |
RUN_TESTS_CHUNK_TIMEOUT_MS |
600000 |
Per-chunk timeout. |
The timing table is advisory and deliberately un-gated. There is no --check
mode and no CI lint that fails on staleness, because timing data legitimately
varies run to run. A file missing from the table falls back to the table's mean
weight (1), and a missing or unparseable table falls back to uniform weight — so
drift costs chunk balance, never a red build. A count-based floor additionally
guarantees the packer never produces fewer chunks than plain count-based packing
would, so a badly stale table cannot collapse the suite into a few fat chunks.
CI job timeout budgets: report + near-cap warning (#4036)
Every matrixed job — test in test.yml, mutate in
mutation.yml, smoke in install-smoke.yml — declares a timeout-minutes
cap. tests/ci-test-job-timeout-budget.test.cjs enforces that each checked-in
cap stays at least the headroom factor (1.5x) above a documented,
hand-measured cost for that job — now all three of the jobs above, not just
test/coverage-gate/test-inert as before.
test-conformance in test.yml (#4591, epic #4589 Phase 2) runs the
scripts/lib/platform-conformance-tier.generated.cjs file list on
windows-latest (sharded three ways) and macos-latest (unsharded) — the
sole gating signal for real-OS coverage (#4603 retired the parallel
legacy full-matrix safety-net job). It
also declares a timeout-minutes cap and runs the same in-job near-cap check
described below, but it has no LANE_COSTS entry in
tests/ci-test-job-timeout-budget.test.cjs yet — no real, completed
(non-cancelled) per-shard measurement exists — so the headroom-factor gate
does not cover it until one lands.
Two runtime mechanisms sit on top of that static gate, both new in #4036:
- In-job near-cap check (
scripts/ci-check-job-near-cap.cjs) — the last step of each of the three jobs computes elapsed-vs-cap from a start-time marker recorded as that job's first step. At >=90% of budget it emits a::warning::annotation (visible in the PR Checks UI) and a$GITHUB_STEP_SUMMARYblock. Advisory only — it never fails the job. Known limit: it cannot fire for a job actually killed by the timeout, since a killed job never reaches its last step. That case is caught by the second mechanism instead. - Scheduled trending report (
.github/workflows/ci-timeout-report.yml,scripts/ci-timeout-report.cjs) — runs daily and onworkflow_dispatch. It polls GitHub's Actions REST API for recently completed jobs acrosstest.yml,mutation.yml, andinstall-smoke.yml, resolves each job's declared cap (a literaltimeout-minutesfortest/test-conformance/smoke, orscripts/mutation-matrix.cjs'sCOVERED[<module>].timeoutMinutesformutate's per-module shards), and appends any new(runId, jobName)records totests/ci-timeout-budget-history.jsonl. Unlike the in-job check, this also catches jobs killed by an actual timeout breach — GitHub's Jobs API still reportsstarted_at/completed_atfor a cancelled job. Each run opens a small, data-only PR carrying that run's new rows, sincenextis a protected branch and nothing pushes to it directly — the same constraintauto-backmerge.ymlalready works within.
This does not retune any timeout-minutes value, rebalance shard composition,
or trim what runs in shard 1 — those stay maintainer policy calls made from
the accumulated history, not something either mechanism decides on its own.
How-to: regenerate the timing table
Regenerate when the suite's cost profile has visibly drifted — after adding or
removing expensive tests, not on a schedule. The input is a node:test reporter
event stream from a gsd-test run:
node scripts/gen-test-timings.cjs \
~/.local/state/gsd-test/runs/<run-id>/test-events-linux-node24.jsonl \
~/.local/state/gsd-test/runs/<run-id>/test-events-macos-node24.jsonl
Pass every lane you have. A file's recorded time is the max across the supplied streams, not the mean: the packer exists to keep the slowest lane's slowest chunk away from the timeout, so the conservative bound is the right one. Keys are sorted so a regeneration diff shows only the files whose cost moved.
Best practices for forward-compat (Node 24/26)
- Use
process.execPathwhen spawning Node in tests so each matrix lane exercises the lane's Node version. - Avoid stack-trace or error-message prose assertions. Assert
err.code, structured JSON fields, or enums — Node minor releases routinely tweak error wording. - Prefer
node:test,node:assert/strict, andnode:testmocks. No external test frameworks. - Coverage uses
c8and propagatesNODE_V8_COVERAGEthrough the harness's child process.
Test strategy: #443 effort + fast_mode engine
Feature: unified cross-provider effort and fast_mode knobs (issue #443). Test files:
tests/model-resolver.test.cjs(unit),tests/model-resolver.test.cjs(integration).
Testing pyramid
| Layer | File | What it covers |
|---|---|---|
| Unit | feat-443-effort-fast-mode.test.cjs |
Pure logic: cascade rules, clamping, escalation math, malformed config handling, schema key validation. No CLI subprocess. |
| Integration | feat-443-effort-fast-mode.integration.test.cjs |
Architecture-level invariants: cross-provider validity, totality across the 33-agent registry, CLI JSON contract, config round-trip, fast-mode honesty. Real subprocesses via runGsdTools. |
| E2E (pending) | (not yet wired) | Propagation layer: effort frontmatter / CLAUDE_CODE_EFFORT_LEVEL env actually reaching a spawned Claude Code subagent. See "Gaps" below. |
Architectural invariants
Each invariant exists to prevent a specific class of production failure.
(a) Cross-provider validity
What: renderEffortForRuntime(runtime, universalEffort).value must always
be a member of the runtime's real provider enum. Ground-truth enums are defined
as local constants in the test — not sourced from the implementation.
PROVIDER_EFFORT_ENUMS = {
claude: Set { 'low', 'medium', 'high', 'xhigh', 'max' } // Anthropic output_config.effort
codex: Set { 'minimal', 'low', 'medium', 'high', 'xhigh' } // OpenAI model_reasoning_effort
}
Why: Passing a value outside these sets results in a 400 from the real API.
The clamping logic (max -> xhigh for codex; minimal -> low for claude) must
hold for every cell of the VALID_EFFORTS × runtimes matrix.
(b) Param/channel contract
What: Each runtime exposes a stable param string (the native API field
name) and channel (how the value is propagated). Unknown runtimes return
param: null, channel: null and pass the effort value through unchanged.
Why: Callers read .param to construct the dispatch payload. A regression
here would silently drop effort from subagent invocations.
(c) Resolve-execution JSON contract
What: The gsd-tools resolve-execution <agent> command emits a JSON object
with all eight keys present and typed correctly: model (string), profile
(string), effort (VALID_EFFORTS member), effort_rendered (string),
effort_param (string|null), effort_propagation (string|null), fast_mode
(boolean), fast_mode_supported (boolean).
Why: Orchestrators and workflow dispatchers parse this JSON. A missing or mistyped field silently breaks downstream consumers.
(d) Totality across the real registry
What: For every agent in the 33-agent registry, resolveEffortInternal
returns a VALID_EFFORTS member (never undefined/null), resolveFastModeInternal
returns a strict boolean, and renderEffortForRuntime('claude', effort) stays
within the claude provider enum.
Why: A catalog addition that introduces a missing routingTier mapping
would otherwise produce undefined and propagate silently.
(e) Fast-mode honesty invariant
What: When the runtime is claude, fast_mode_supported in
resolve-execution output is always false, regardless of the fast_mode config.
RUNTIMES_WITH_FAST_MODE contains only 'api'.
Why: Claude Code's /fast toggle is session-level only. Emitting
fast_mode: true as frontmatter on a Claude subagent is a silent no-op.
Advertising fast_mode_supported: true for claude would cause orchestrators to
believe the knob was wired when it is not.
(f) Precedence first-valid-wins
What: Both effort and fast_mode use a layered cascade. The test table covers all four effort layers (invocation override → agent_overrides → routing_tier_defaults → default) and all five fast_mode layers, including the case where an invalid value at a higher layer correctly falls through.
Why: Silent precedence bugs (e.g., a numeric value in agent_overrides not being rejected) would override intentional user config.
(g) Dynamic-routing composition
What: resolveEffortForTier escalates effort by attempt number
independently of the model tier mapping. The test verifies the effort ladder
(low -> medium -> high -> xhigh -> max), the max clamp, the
max_escalations cap, and that escalate_on_failure: false suppresses
escalation entirely.
Why: Effort escalation and model escalation share configuration
(dynamic_routing) but must operate independently; coupling them would cause
over-escalation or under-escalation.
(h) Config-tooling round-trip
What: gsd-tools config-set accepts all new key namespaces
(effort.default, effort.routing_tier_defaults.<tier>,
effort.agent_overrides.<agent>, fast_mode.enabled,
fast_mode.routing_tier_defaults.<tier>, fast_mode.agent_overrides.<agent>)
without an "Unknown config key" error, and values set via config-set are
reflected in resolve-execution output.
Why: The schema validation gate (VALID_CONFIG_KEYS + DYNAMIC_KEY_PATTERNS)
is separate from the resolver logic. A key missing from the schema would produce
a silent write failure and appear as a bug only at runtime.
Coverage targets
| Suite | Target |
|---|---|
| Unit | Every cascade rule, every fallthrough, every clamp. All function branches in resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, renderEffortForRuntime. |
| Integration | All 8 architectural invariants. All 33 registered agents. All 6 provider × effort combinations for the valid-enum check. Full config-set key namespace. |
Gaps / not yet covered
E2E orchestrator-spawn-propagation layer (pending follow-up wiring): The integration tests verify that GSD resolves and renders effort values correctly. They do NOT verify that the rendered values actually reach a spawned Claude Code or Codex subagent at runtime. Specifically uncovered:
CLAUDE_CODE_EFFORT_LEVELenv var being set and read by a spawned claude subprocessoutput_config.effortfrontmatter key surviving the AGENTS.md template substitutionmodel_reasoning_effortfield surviving serialization into a Codex API request body- Fast-mode
speed: "fast"field reaching anapi-runtime request whenfast_mode_supported: true
These require spawning real subagents (or stubs thereof) and asserting on the
process environment / request payload — a scope that belongs in a future E2E
suite under *.slow.test.cjs or dedicated fixture-driven integration work.