Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes, cline, codebuddy and pi end to end: capability descriptors, installer branches and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters, hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi migrations, Kimi payload normalization in the hook guards, dead hostBehaviors vocabulary, launcher home probes, fixtures, runtime-specific tests and the prose that presented them as supported. Installer output for the six kept runtimes is byte-identical to before the prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and read-injection-scanner are left in place pending a decision.
738 lines
44 KiB
Markdown
738 lines
44 KiB
Markdown
# Testing Suites
|
||
|
||
This project's `tests/` directory uses **filename suffix markers** to group tests into named suites. The harness `scripts/run-tests.cjs` filters by suite when given `--suite <name>`. Without a flag it runs every `*.test.cjs` file (the historical default — unchanged).
|
||
|
||
> Tracked by issue [#3597](https://github.com/open-gsd/gsd-core/issues/3597).
|
||
|
||
## Suites
|
||
|
||
| Suite | Filename pattern | What goes here |
|
||
|---|---|---|
|
||
| `unit` | `*.test.cjs` (no other marker) | Default fast lane. Pure logic, no network, no external processes beyond `msd-tools`. Most tests live here. |
|
||
| `integration` | `*.integration.test.cjs` | Cross-module flows: full installer end-to-end, multi-tool orchestration, anything that crosses two or more bin entry points. |
|
||
| `install` | `*.install.test.cjs` | Tests that perform a real install/uninstall against a sandbox project. Slower; PR CI skips these on PRs and runs them on `main` push only. |
|
||
| `security` | `*.security.test.cjs` | Adversarial input, prompt-injection guards, fixture-driven hostile-payload sweeps. |
|
||
| `slow` | `*.slow.test.cjs` | Anything that routinely takes >5s wall-clock or holds significant memory. |
|
||
| `qa` | `*.qa.test.cjs` | End-to-end walks that drive the real `msd-tools` binary across multiple loop steps against one accumulating temp project, with invariant oracles after every step. Slower than `unit`; excluded from the fast lane. |
|
||
| `all` | (any) | Explicit alias for "no filter". Equivalent to running with no `--suite` flag. |
|
||
|
||
## How to place a new test
|
||
|
||
1. Pick the most specific bucket above.
|
||
2. Name the file with the matching suffix: `tests/<feature>.<suite>.test.cjs`.
|
||
3. If unsure, leave the suffix off — the file lands in `unit`, the default fast lane.
|
||
|
||
Examples:
|
||
- `tests/agent-frontmatter.test.cjs` — `unit`
|
||
- `tests/prompt-injection-guards.security.test.cjs` — `security`
|
||
- `tests/installer-end-to-end.install.test.cjs` — `install`
|
||
- `tests/sdk-mutation-stress.slow.test.cjs` — `slow`
|
||
- `tests/loop-walk.qa.test.cjs` — `qa`
|
||
|
||
The suite-suffix convention was chosen over a directory layout (`tests/security/`) so the 545+ existing test files don't need to move. Existing files all classify as `unit` until someone explicitly retags them.
|
||
|
||
## Regression tests
|
||
|
||
**Do not create new top-level `tests/bug-NNNN-*.test.cjs`,
|
||
`tests/fix-NNNN-*.test.cjs`, or `tests/issue-NNNN-*.test.cjs` files.** Add the
|
||
regression case to the owning module's main test file instead (e.g. a
|
||
`describe('regressions')` block in `tests/<module>.test.cjs`).
|
||
|
||
`node --test` spawns one child process per FILE, so file count — not test
|
||
count — is the unit of CI overhead, and it is worst on Windows lanes where
|
||
every spawn is Defender-scanned. The 2026-06 CI audit found 244 one-off
|
||
`bug-*` files (~38% of the suite). That population is grandfathered in
|
||
`scripts/lint-regression-test-names.allowlist.json` and enforced by an
|
||
identity ratchet (`npm run lint:regression-names`, part of `npm run lint:ci`),
|
||
which also bans `fix-*` and `issue-*` NNNN-prefixed files the same way:
|
||
|
||
- A **new** `bug-*`, `fix-*`, or `issue-*` NNNN-prefixed file fails CI — fold
|
||
it into the owning module's file.
|
||
- **Deleting/consolidating** a grandfathered file requires pruning its
|
||
allowlist entry, so the baseline only ever shrinks.
|
||
- **Inherited drift** (the failure names files your PR didn't add — e.g. the
|
||
base branch merged `bug-*`/`fix-*`/`issue-*` files without feeding the
|
||
allowlist, or you rebased and carried a pre-rebase allowlist): run
|
||
`node scripts/lint-regression-test-names.cjs --update` and commit the
|
||
regenerated allowlist. Snapshot artifacts like this allowlist (and
|
||
`docs/INVENTORY.md`) must be regenerated **after** rebasing, never carried
|
||
through a rebase.
|
||
|
||
### A folded suite may appear only once per host
|
||
|
||
When a standalone file is folded into its owning module's test file, the moved
|
||
suite is wrapped in a self-contained block carrying a marker:
|
||
|
||
```javascript
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
// Folded from tests/bug-376-claude-js-hook-msd-rewriter.test.cjs — …
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
{
|
||
const { describe: __foldDescribe } = require('node:test');
|
||
__foldDescribe("folded:bug-376-claude-js-hook-msd-rewriter (…)", () => { … });
|
||
}
|
||
```
|
||
|
||
Because the block is self-contained, a **second verbatim copy in the same host
|
||
parses, registers, and passes — twice.** Nothing in a green suite reports it.
|
||
[#3271](https://github.com/open-gsd/gsd-core/issues/3271) found 25 such copies
|
||
(~5,800 lines) across three install suites, all from a single stale-base
|
||
re-application during the consolidation epic. The cost is not only wasted CI on
|
||
every lane: it is a `DEFECT.GENERATIVE-FIX` trap, because a contributor fixing
|
||
one of those regressions edits the copy they found and leaves the other
|
||
asserting the old behavior, with the suite still green.
|
||
|
||
`local/no-duplicate-fold-marker` (`eslint-rules/no-duplicate-fold-marker.cjs`,
|
||
error under `tests/**/*.cjs`) reports the second and every later occurrence of a
|
||
`folded:<marker>` title in one file, naming the line the first occurrence sits
|
||
on. **When it fires, delete the copy it points at** — the two blocks are the
|
||
same suite, so the fix is removal, never an `eslint-disable`.
|
||
|
||
It keys on the whitespace-delimited token after `folded:`, which matters in both
|
||
directions. A narrower key that stops at `.` would collide
|
||
`feat-443-effort-fast-mode.integration` with `feat-443-effort-fast-mode` — two
|
||
genuinely distinct suites that coexist in `tests/model-resolver.test.cjs`. Keying
|
||
on the *whole* title instead would let a re-fold under a different batch label
|
||
slip through, which is exactly the shape #3271 took. Titles without a `folded:`
|
||
prefix, non-literal titles, and the same marker appearing in two *different* host
|
||
files are all left alone.
|
||
|
||
The ratchet deliberately covers only `bug-*`. Files named `feat-NNNN-*` /
|
||
`enh-NNNN-*` are *feature* test files — one (or one per suite) per feature is
|
||
the sanctioned layout (see the #443 strategy below), not a one-off regression
|
||
pattern. If `issue-*`/`perf-*` one-offs start accumulating the same way
|
||
`bug-*` did, extend the ratchet's regex and regenerate the allowlist.
|
||
|
||
## Workflow & agent size budget
|
||
|
||
> Tracked by issue [#1074](https://github.com/open-gsd/gsd-core/issues/1074).
|
||
> Bytes (not lines) per [#717](https://github.com/open-gsd/gsd-core/issues/717);
|
||
> LF-normalized per [#683](https://github.com/open-gsd/gsd-core/issues/683).
|
||
|
||
Workflow files (`msd-core/workflows/*.md`) and agent files (`agents/msd-*.md`)
|
||
both ship in the installed runtime and are loaded into context — workflows on
|
||
every command, agents on every subagent dispatch — so their byte size is a real
|
||
cost. Two sibling guards (`tests/workflow-size-budget.test.cjs` and
|
||
`tests/agent-size-budget.test.cjs`) keep that cost from creeping up invisibly,
|
||
sharing one byte-counter (`measureMdFiles`). Growth is caught by two independent
|
||
layers:
|
||
|
||
| Layer | What it does | Where |
|
||
|---|---|---|
|
||
| **Differential attribution size ratchet** (primary, #2724 / ADR-2719 §4) | The same computed-attribution check that replaced the golden-install-parity fixtures also reports growth in any `msd-core/workflows/*.md` or `agents/msd-*.md` file, with the exact byte delta, comparing PR HEAD against `next`. Unacknowledged growth is a hard failure; shrinkage needs no acknowledgment. No committed snapshot — nothing to regenerate by hand. | `tests/emitted-attribution.test.cjs` (real-tree test) via `tests/helpers/emitted-diff.cjs` |
|
||
| **Loose tier hard caps** (backstop) | Absolute outer red lines per tier — workflows: `XL ≤ 98304`, `LARGE ≤ 61440`, `DEFAULT ≤ 40960` bytes; agents: `XL ≤ 57344`, `LARGE ≤ 49152`, `DEFAULT ≤ 24576` bytes. A cap is **never raised** when a file approaches it: crossing it means *extract*, not bump. Independent of the ratchet above — unaffected by #2724. | `XL/LARGE/DEFAULT_CAP` in each guard file |
|
||
| **Headroom census + reserved margin** (visibility, [#4261](https://github.com/open-gsd/gsd-core/issues/4261)) | Every run prints each capped file's remaining bytes and percentage used — green runs included — sorted least-headroom-first, and appends a table of the files past a **95% reserved margin** to the GitHub job summary. The margin **reports, it does not fail**: a file at 96% is not broken, it is a file whose next contributor should extract before adding. Nothing here raises or relaxes a cap. | `buildHeadroomRows` / `marginFor` in `scripts/workflow-size.cjs` |
|
||
|
||
Why the census exists: each PR's CI measures only its own base plus its own
|
||
diff, so two PRs that are individually under a cap can be jointly over it, and
|
||
no run either of them produces can show that. The census does not solve that
|
||
directly — measuring on the merge result would, and was deliberately left out
|
||
of #4261's approved scope — but it makes the density that causes it legible
|
||
before the collision, which a passing run previously did not. It also replaces
|
||
the hand-written per-tier high-water comments in both guard files, which had
|
||
gone stale by several kilobytes and were themselves the reason the shrinking
|
||
margin went unnoticed.
|
||
|
||
`discuss-phase.md` additionally has a thin-dispatcher target of `< 32000` bytes
|
||
(the discuss-phase progressive-disclosure split, #717). A net-new agent is
|
||
DEFAULT-tier and already bounded by the DEFAULT cap — no separate new-agent cap
|
||
is needed. (This tier-cap machinery is distinct from the separate 45 KB-*char*
|
||
extraction-evidence threshold on `msd-planner` enforced by
|
||
`tests/planner-decomposition.test.cjs` — that one proves mode sections were
|
||
extracted; this one bounds total agent bytes.)
|
||
|
||
### How-to: a workflow or agent grew and CI is red
|
||
|
||
The differential attribution check reports the file and the byte delta. To resolve:
|
||
|
||
1. **Justify the growth in your PR** (a sentence in the description is enough) —
|
||
the acknowledgment entry (below) is the review record that the larger size
|
||
was a deliberate, seen decision, not silent drift.
|
||
2. **Add an acknowledgment trailer** to one of your own commits (ADR-3942),
|
||
naming the file and the reason, per `CONTRIBUTING.md`'s "Editing shipped
|
||
content" section and `CONTEXT.md`'s `### Emitted Artifact Provenance` entry:
|
||
|
||
```
|
||
Emitted-Drift-Ack-Growth: explore.md — new dispatch section, reasoning ships with the block
|
||
```
|
||
|
||
Growth keys on the **bare filename** as it appears under `msd-core/workflows/`
|
||
or `agents/`; an unattributable **hash** ripple uses
|
||
`Emitted-Drift-Ack-Hash:` and keys on the emitted path (which always contains
|
||
a `/`). The two are separate namespaces — a growth trailer will not excuse a
|
||
hash ripple, and the failure output says which one applies. The trailer is the
|
||
review record that the larger size was a deliberate, seen decision.
|
||
|
||
If you need to change an acknowledgment, amend the commit carrying it. That is
|
||
deliberate: the trailer cannot drift out of sync with the diff it explains,
|
||
because changing either changes the sha and re-runs the gate.
|
||
3. **There is nothing to clean up afterwards.** The trailer is read from
|
||
`git log $(git merge-base <base> HEAD)..HEAD` — your commits and no others —
|
||
so once your PR merges it is out of range by construction. It never becomes
|
||
"spent", it owns no shared key space, it cannot conflict with anyone else's,
|
||
and no sweeper has to delete it. That is the whole reason ADR-3942 moved the
|
||
acknowledgment off the working tree.
|
||
4. **Or shrink it instead of acknowledging.** Prefer extraction when the growth
|
||
is incidental: for a workflow, move per-mode bodies to
|
||
`workflows/<name>/modes/`, templates to `workflows/<name>/templates/`, and
|
||
shared prose to `msd-core/references/`; for an agent, lift shared boilerplate
|
||
into `msd-core/references/` and `@`-reference it — then load it **LAZILY**. Do
|
||
*not* convert them to eager `@-required_reading` includes: that shrinks the
|
||
file's bytes without shrinking loaded context, so it games the guard while
|
||
making the real cost worse. See `workflows/discuss-phase/` for the
|
||
progressive-disclosure pattern.
|
||
|
||
If a hard cap (not the ratchet) is what failed, an acknowledgment will **not**
|
||
help — that is the signal to extract, per step 3.
|
||
|
||
### Reference
|
||
|
||
| Artifact | Role |
|
||
|---|---|
|
||
| `scripts/workflow-size.cjs` | Single source of truth — LF-normalized byte counter (`lfByteCount`) + generic `measureMdFiles(dir, predicate)` (backs both workflows and agents) + workflow enumeration (`listWorkflowStems`, `measureWorkflows`). Imported by both guards and by `tests/helpers/emitted-runtime.cjs`'s `currentSizes()` so they can never measure differently. |
|
||
| `tests/emitted-attribution.test.cjs` + `tests/helpers/emitted-diff.cjs` | The differential attribution check and its size ratchet (ADR-2719). The sole mechanism for both emitted-content propagation AND per-file size growth as of #2724. |
|
||
| `Emitted-Drift-Ack-Hash:` / `Emitted-Drift-Ack-Growth:` commit trailers (ADR-3942) | The acknowledgment mechanism for unattributable emitted-content ripples and for size growth. Read from `git log $(git merge-base <base> HEAD)..HEAD` — no committed file, nothing to sweep; a merged trailer is out of range by construction. |
|
||
| `npm run regen:derived` | Runs every remaining generator in dependency order (build → registry → ADR index → capability matrix → inventory manifest → manifest versions → `tests/fixtures/install-tree/*.json`). |
|
||
| `tests/workflow-size-budget.test.cjs` | The workflow tier hard-cap guards, plus the `discuss-phase` progressive-disclosure checks. |
|
||
| `tests/agent-size-budget.test.cjs` | The agent tier hard-cap guards (the agent analog). |
|
||
|
||
`tests/workflow-size-baseline.json`, `tests/agent-size-baseline.json`,
|
||
`tests/fixtures/golden-install-parity/*.json`, `scripts/update-size-baseline.cjs`
|
||
(`npm run size:baseline`), and `scripts/git-merge-regen-driver.cjs`
|
||
(`npm run setup:merge-driver`) are all removed by
|
||
[#2724](https://github.com/open-gsd/gsd-core/issues/2724): they were pure
|
||
functions of the source tree, conflicted on every merge that touched them, and
|
||
their functions are now served by the differential attribution check above.
|
||
`tests/fixtures/install-tree/*.json` is the one artifact family that stays
|
||
committed and normally-merged (ADR-2719 §7) — it conflicts on 0 of 7, its diffs
|
||
are readable, and it preserves "the installer stopped shipping X" as a hard
|
||
absolute failure with no attribution reasoning involved. Regenerate it with
|
||
`npm run gen:install-tree` (folded into `npm run regen:derived`).
|
||
|
||
## The QA smell ratchet
|
||
|
||
`tests/loop-walk.qa.test.cjs` (the `qa` suite) is the QA-walk harness's own
|
||
self-test. Separately, `scripts/qa-smell-ratchet.cjs` drives that same harness
|
||
end to end against the real `msd-tools` binary and turns its findings into a
|
||
CI gate — run it with `npm run lint:qa-smells`.
|
||
|
||
The gate has three independent inputs. The harness's oracles
|
||
(`tests/qa/oracles.cjs`) supply the first two:
|
||
|
||
- A **violation** is the engine breaking a documented contract. It always
|
||
fails the build — baseline or no baseline, acknowledged or not.
|
||
- A **smell** is legal-but-questionable behavior. A smell **never fails a
|
||
build on its own merits**. What fails is the absence of a decision about
|
||
it: an **unacknowledged NEW smell**, or a **STALE** entry in
|
||
`tests/qa/smell-baseline.json` (one that stopped firing — the baseline is
|
||
shrink-only, so a fixed or changed scenario must be pruned, not left
|
||
behind).
|
||
|
||
The third comes from the scenarios themselves, not from an oracle:
|
||
|
||
- A **scenario expectation failure** is a step's declared `expect` not
|
||
holding — the scenario asserted `percent: 100` and the engine returned
|
||
something else. Like a violation, it is **never acknowledgeable**: it
|
||
carries no fingerprint, so there is no `key` to put in a baseline entry or
|
||
an ack fragment. Fix the engine, or correct the expectation.
|
||
|
||
Expectation failures were invisible to the gate until
|
||
[#3597](https://github.com/open-gsd/gsd-core/issues/3597): `buildReport`
|
||
counted them in `totals.violations` while `collectFindings` read only oracle
|
||
violations, so a failing scenario printed `0 violations` and exited 0. The
|
||
`multi-workstream` scenario failed on every CI run for three weeks without
|
||
reddening a build. The invariant that keeps the two honest — asserted in
|
||
`tests/loop-walk.qa.test.cjs` — is:
|
||
|
||
```
|
||
collectFindings().violations.length
|
||
+ collectFindings().expectationFailures.length
|
||
=== report.totals.violations
|
||
```
|
||
|
||
Note also that `scripts/qa-smell-ratchet.cjs` only invokes its own `main()`
|
||
under `require.main === module`. That guard is what lets the QA suite
|
||
`require()` the script to test `collectFindings` without kicking off a real
|
||
20-scenario walk as an import side effect — the reason the gate's own logic
|
||
had no test before #3597.
|
||
|
||
Every smell must terminate in exactly one of TWO states — there is no third
|
||
"accepted with a good explanation" state:
|
||
|
||
1. **REAL** — an assigned defect. File it, then acknowledge the smell with an
|
||
entry (baseline entry or `tests/qa/smell-acks/` fragment) carrying that
|
||
`issue` number.
|
||
2. **FALSE POSITIVE** — the oracle itself is wrong. Fix the oracle
|
||
(`tests/qa/oracles.cjs`) so it stops firing. It is NEVER baselined.
|
||
|
||
When the ratchet reports a NEW smell, there are exactly two legitimate
|
||
responses — fix the detector, or file a defect and cite its issue number:
|
||
|
||
1. **Fix the underlying behavior (or the oracle, if it's a false positive)**
|
||
so the smell stops firing.
|
||
2. **File a defect and acknowledge it** by adding a fragment under
|
||
`tests/qa/smell-acks/` — the ratchet's failure output prints a paste-ready
|
||
skeleton naming the required `key`, `id`, `scenario`, and `issue` fields.
|
||
`issue` MUST be a positive integer naming the tracking issue; a free-text
|
||
`reason` may accompany it as an optional human note but can NEVER
|
||
substitute for `issue` — "write an explanation" is not a way to acknowledge
|
||
a smell. See `tests/qa/smell-acks/README.md` for the full shape and
|
||
lifecycle.
|
||
|
||
Run `node scripts/qa-smell-ratchet.cjs --update` to regenerate
|
||
`tests/qa/smell-baseline.json` from the current run, folding in any acked
|
||
fragments and pruning stale entries. `--update` never invents an issue
|
||
number: a genuinely new smell is written with `issue: null` and a TODO
|
||
`reason`, and the very next plain (non-`--update`) run REJECTS that entry —
|
||
forcing a human to triage it before it can ship. The baseline only ever
|
||
shrinks: growth happens by adding an acknowledgment carrying a real issue
|
||
number (a reviewable diff), never by widening the generator's tolerance and
|
||
never by prose alone.
|
||
|
||
## Running suites locally
|
||
|
||
```bash
|
||
npm test # everything (backcompat — same as before)
|
||
npm run test:unit # only unit
|
||
npm run test:integration # only integration
|
||
npm run test:install # only install
|
||
npm run test:security # only security
|
||
npm run test:slow # only slow
|
||
|
||
npm run test:coverage # backcompat — coverage over EVERY test
|
||
npm run test:coverage:unit # fast coverage signal — only unit suite
|
||
npm run test:coverage:all # alias for test:coverage
|
||
```
|
||
|
||
Direct harness invocation also works:
|
||
|
||
```bash
|
||
node scripts/run-tests.cjs --suite security
|
||
node scripts/run-tests.cjs --suite=security
|
||
node scripts/run-tests.cjs --files "tests/command-contract.test.cjs tests/core.test.cjs"
|
||
node scripts/run-tests.cjs --files-from .ci-selected-tests.txt
|
||
```
|
||
|
||
`npm run test:affected` (scripts/run-affected-tests.cjs) is a **local-only**
|
||
convenience that selects tests via the `require()` dependency graph of your
|
||
working-tree diff. CI does not use it — CI selection is the rule table in
|
||
`scripts/ci-test-scope.cjs`, which is the authoritative mapping. If the two
|
||
disagree, trust (and fix) the rule table.
|
||
|
||
Unknown suites exit non-zero with the list of valid suites. Empty suites (e.g. `--suite security` before any security-tagged file exists) exit `0` with a `no tests in suite "..."` notice on stderr so CI lanes don't go red while a suite is being populated.
|
||
|
||
## The live-config hermeticity guard
|
||
|
||
Every `run-tests.cjs` invocation snapshots MSD's own install footprint in each
|
||
live runtime config directory before the suite and re-checks it afterwards. It
|
||
exists because the failure it catches is silent by construction: a test that
|
||
resolves a config directory from the ambient environment instead of a sandbox
|
||
writes into *your real* `~/.claude` (or `$MSD_HOME/.msd`, or a runtime's
|
||
config file), and nothing reports it. CI cannot catch this class at all —
|
||
CI never has `CLAUDE_CONFIG_DIR` and friends set.
|
||
|
||
The guard watches only what MSD unambiguously owns — its top-level install
|
||
footprint plus `msd-`-prefixed children of directories shared with the host
|
||
agent — never whole config roots, because a host agent legitimately writing
|
||
`history.jsonl` mid-run would make the guard cry wolf, and a guard that cries
|
||
wolf gets switched off.
|
||
|
||
Two environment variables control it:
|
||
|
||
| Variable | Effect |
|
||
|---|---|
|
||
| `MSD_STRICT_LIVE_CONFIG_GUARD=1` | A detected write **fails the run**. Set on the Linux/macOS lanes of every CI job that runs the suite. |
|
||
| `MSD_SKIP_LIVE_CONFIG_GUARD=1` | Skips the check entirely. |
|
||
|
||
Unset, the guard **reports and does not fail** — deliberately, not timidly. On
|
||
its first CI run it surfaced pre-existing leaks on the Windows lane, where
|
||
`os.homedir()` reads `USERPROFILE` and ~190 test sites sandbox `HOME` alone.
|
||
Those are real and worth fixing, but they are a different defect class, and a
|
||
brand-new gate that instantly reds an unrelated lane gets reverted rather than
|
||
obeyed. Windows lanes therefore stay report-only until that sweep lands; this
|
||
repo has the pattern already, in the `local/no-source-grep` ESLint rule that
|
||
shipped at `warn` and was promoted to `error` after its cleanup (ADR 452).
|
||
|
||
`MSD_SKIP_LIVE_CONFIG_GUARD` is a bypass on a safety check, so it is documented
|
||
here rather than left to be discovered in the source: an undocumented bypass is
|
||
one people eventually set without knowing what they turned off. If you need it
|
||
routinely, that is a bug report, not a workflow.
|
||
|
||
Reported paths are labelled `CREATED`, `MODIFIED`, `DELETED`, or `UNVERIFIED`.
|
||
`UNVERIFIED` means a scan bound was hit and the path could not be attested
|
||
either way — it is never the same as clean.
|
||
|
||
## The mergeability preflight
|
||
|
||
Before any of the matrix below is provisioned, every `pull_request` compute lane
|
||
waits on one shared gate: **`PR mergeability`**, the reusable workflow
|
||
`.github/workflows/pr-mergeable-preflight.yml`. **A pull request with a merge
|
||
conflict runs no CI at all until the conflict is resolved** (#3833).
|
||
|
||
### Reference
|
||
|
||
| Verdict | When | Job result | Effect on the pipeline |
|
||
|---|---|---|---|
|
||
| `MERGEABLE` | GitHub reported `mergeable: true` | success | everything runs as normal |
|
||
| `CONFLICTED` | GitHub reported `mergeable: false` | **failure** | every gated job is skipped; `Required tests` and `Stryker mutation score` report **red** |
|
||
| `INDETERMINATE` | mergeability still unknown after the retry budget, or the API read failed | success **(fails open)** | everything runs as normal; a `::warning::` is emitted |
|
||
| `SKIPPED_NOT_A_PR` | the event is not `pull_request` (push, `workflow_dispatch`, a `release.yml` call into `install-smoke.yml`) | success | everything runs as normal; **zero** API calls |
|
||
| `INDETERMINATE` (bootstrap) | `scripts/ci-pr-mergeability.cjs` is absent at the base sha | success **(fails open)** | everything runs as normal; a `::warning::` names the cause |
|
||
|
||
The bootstrap row is a consequence of the checkout being pinned to the base sha:
|
||
the preflight runs the script **as it exists on the base branch**, so the script
|
||
is absent on the pull request that introduces it, and on any branch whose base
|
||
predates it. Absent is not "conflicted" — it is one more thing the gate cannot
|
||
determine, so it takes the same fail-open path. The arm is self-healing and
|
||
never fires again once the script is on the base branch; a test pins it in place
|
||
so a future reader does not mistake it for dead code.
|
||
|
||
The verdict comes from `scripts/ci-pr-mergeability.cjs`, which polls
|
||
`GET /repos/{owner}/{repo}/pulls/{number}` until `mergeable` is non-null.
|
||
GitHub computes mergeability in an **asynchronous background job** and returns
|
||
`null` while it is in progress, so a single cold read is never authoritative —
|
||
the poll is required, not defensive.
|
||
|
||
Two properties are load-bearing and easy to break:
|
||
|
||
- **Only `mergeable === true` / `=== false` decides the verdict.** `null` is
|
||
falsy, so a truthiness test (`if (!mergeable)`) would classify every cold read
|
||
as a conflict and red every PR in the repo.
|
||
- **`mergeable_state` never decides the verdict.** Its `blocked` value means
|
||
"required checks have not passed", which is true *while this very job is
|
||
running* — gating on it self-deadlocks the pipeline. It is read for the
|
||
failure message only.
|
||
|
||
Gated: `test.yml` (`lint-tests`, `test`, `test-inert`,
|
||
`test-conformance`, `coverage-gate`, `qa-loop-walk`, `required-tests`), `install-smoke.yml`,
|
||
`mutation.yml`, `security-scan.yml`, `docs-required.yml`,
|
||
`changeset-required.yml`, `default-flip-documentation.yml`, `branch-naming.yml`.
|
||
|
||
**Deliberately not gated:** the `pull_request_target` policy and security lanes —
|
||
auto-close-unsolicited-PRs, close-draft-PRs, the target/title/template
|
||
validators, issue-link, and unauthorized-approval dismissal. Gating those would
|
||
let a conflicted drive-by PR evade auto-close, so they keep running on a
|
||
conflicted PR and a test asserts they are never wired to the preflight.
|
||
|
||
### What it deliberately does not do
|
||
|
||
**It applies no label.** The gate does not add or remove `needs-review: merge-conflict`.
|
||
Eight caller workflows each invoke the preflight, so eight jobs would race to
|
||
add-or-remove one label on every PR event, and it would force
|
||
`pull-requests: write` into eight lanes that today hold `contents: read`.
|
||
`needs-review: merge-conflict` stays a maintainer triage label applied during PR
|
||
sweeps; the red check and its annotation are the machine signal.
|
||
|
||
**It changes no branch protection.** `.github/rulesets/main-protection.json` is
|
||
untouched and needs no new required context — GitHub natively refuses to merge a
|
||
pull request with conflicts, so the ruleset already blocks it. (Adding the check
|
||
as *required* before the workflow exists on the base branch would deadlock every
|
||
open PR until it landed.)
|
||
|
||
### What this does not replace
|
||
|
||
`scripts/ci-rebase-check.cjs` is unchanged and still runs inside each matrix job,
|
||
including its #2472 base-sha pin. The preflight is an early-exit optimization on
|
||
GitHub's asynchronously-computed view; the in-job merge remains authoritative and
|
||
covers the case where the base advances mid-run. **The gate is never a safety
|
||
property** — every path except an explicit `mergeable: false` fails open, so an
|
||
API outage simply restores the pre-#3833 behavior.
|
||
|
||
### If your PR shows a red `PR mergeability` check
|
||
|
||
The annotation names the base branch and the remedy. There is one step:
|
||
|
||
```bash
|
||
git fetch origin && git rebase origin/next && git push --force-with-lease
|
||
```
|
||
|
||
Resolve the conflicts the rebase reports, then push. The next `synchronize`
|
||
event re-runs the preflight and the full pipeline comes back. (Because the
|
||
sequence is a single command, this has no separate how-to page — see
|
||
[CONTRIBUTING.md → Where Do I Open My PR?](../CONTRIBUTING.md#where-do-i-open-my-pr-branching-model)
|
||
for the branching model that makes the rebase necessary in the first place.)
|
||
|
||
Note the interaction with the sha-bound pass marker: a rebase changes your HEAD
|
||
sha, which invalidates any prior remote-runner verification. Rebase *last*.
|
||
|
||
## CI matrix
|
||
|
||
The `Tests` workflow (`.github/workflows/test.yml`) runs every PR through a scope
|
||
computed by `scripts/ci-test-scope.cjs`'s `classify()`, which sets two flags —
|
||
`product_changed` and `full_matrix` — from the changed-file list.
|
||
|
||
All lanes run on **Node 24** — the `engines.node` floor (`>=24.0.0`) and the
|
||
only supported runtime.
|
||
|
||
| Job | Lanes | Gated on | Purpose |
|
||
|---|---|---|---|
|
||
| `test` | `ubuntu-latest` (1 targeted + 3-shard full) | `product_changed == 'true'` | The default, always-scoped PR signal — the full `unit`/`integration`/`security` suites run once, sharded, on Linux. **Linux only**: its three `scope: windows` shards were deleted in #4641 (ADR-4641), which found them a second, redundant Windows selector alongside `test-conformance` |
|
||
| `test-inert` | `ubuntu-latest` | `code_changed == 'true' && product_changed != 'true'` | A lightweight lane for PRs that touch only administrative/policy workflow files (code changed, but nothing that needs the real matrix) |
|
||
| `test-conformance` | `windows-latest` (3-shard) + `macos-latest` (unsharded) | `code_changed == 'true' && full_matrix == 'true'` | Runs only the **platform-conformance-tier** file list (`scripts/lib/platform-conformance-tier.generated.cjs`, epic #4589 Phase 2/#4591) on real Windows/macOS — since #4641 the **sole** Windows and macOS selector in CI, not merely the sole gating one. Retired the parallel legacy full-suite matrix in #4603; #4641 removed the second Windows selector in `test` and narrowed the tier from 548 to 266 of 932 eligible unit-suite files (58.8% → 28.5%, measured 2026-09-11; the absolute counts track `next`'s test count, the percentages are what the ceiling test binds on). |
|
||
| `coverage-gate` | `ubuntu-latest` | `product_changed == 'true' && test.result == 'success'` | Merges every `test` shard's coverage dumps and evaluates the threshold once (sharding moved this out of the `test` job itself — #2952) |
|
||
| `qa-loop-walk` | `ubuntu-latest` | `product_changed == 'true'` | The QA smell-ratchet scenario walk (see "The QA smell ratchet" below) |
|
||
| `required-tests` | `ubuntu-latest` | `always()` | Aggregates every job above into the one branch-protection-required check |
|
||
|
||
`full_matrix` fires on any changed `tests/**/*.test.cjs` file unconditionally (restored
|
||
by #4421 after #962's narrowing let a real regression through undetected), plus a
|
||
handful of curated `RULES` (workflow/installer/hooks/env-gate changes) — see
|
||
`scripts/ci-test-scope.cjs`'s own `classify()` for the exact, current rule set; this
|
||
doc intentionally does not restate it in full, to avoid drifting out of sync with it.
|
||
|
||
Coverage is evaluated by the dedicated `coverage-gate` job (moved out of the `test`
|
||
job by #2952, once sharding meant no single `test` runner saw the whole picture) and
|
||
stays single-lane (Ubuntu / Node 24 only) because multiplying coverage across
|
||
OS/runtime lanes adds cost without improving the threshold signal. Note the gate's
|
||
deliberate blind spot: it measures
|
||
`msd-core/bin/lib/*.cjs` only — `scripts/`, `hooks/`, and `bin/` are
|
||
unenforced, and `stryker.config.mjs` additionally excludes ~48% of lib lines
|
||
from mutation testing (see the UNMUTATED list there). Widening either gate is
|
||
tracked work, not an accident to "fix" silently by raising thresholds.
|
||
|
||
To inspect the scope locally:
|
||
|
||
```bash
|
||
npm run ci:test-scope -- --files "commands/msd/plan-phase.md"
|
||
node scripts/ci-test-scope.cjs --base origin/next --head HEAD
|
||
```
|
||
|
||
## Chunk packing and the test timing table
|
||
|
||
`scripts/run-tests.cjs` does not hand the whole selected file list to one
|
||
`node --test` process. It packs the files into **chunks**, each spawned
|
||
separately, because Windows caps a command line at 32,767 characters and because
|
||
each chunk gets its own 600s timeout (`RUN_TESTS_CHUNK_TIMEOUT_MS`) and a fresh
|
||
process, which bounds memory pressure.
|
||
|
||
How files are distributed across those chunks decides whether the slowest chunk
|
||
sits near that timeout while the others idle. The packer weights each file by its
|
||
**measured duration**, read from `tests/test-timings.json`, and places files with
|
||
LPT (longest-processing-time-first: heaviest file first, each into the currently
|
||
lightest chunk). Before #2456 the weight was guessed from the filename, which
|
||
mis-ranked files badly enough that the slowest chunk ran ~3.9x the lightest.
|
||
|
||
### Reference
|
||
|
||
| Knob | Default | Meaning |
|
||
|---|---|---|
|
||
| `RUN_TESTS_MAX_FILES_PER_CHUNK` | `60` (`22` on win32) | Per-chunk weight budget. Weights are normalized so an **average-cost** file weighs 1, so this still reads as "about 60 average files" (about 22 on win32). Windows gets a lower cap than Linux/macOS because the weight table's calibration does not transfer 1:1 to the Windows runner for install/subprocess-heavy work — see the derivation comment above `DEFAULT_MAX_FILES_PER_CHUNK` in `scripts/run-tests.cjs`. |
|
||
| `RUN_TESTS_MAX_CMDLINE_CHARS` | `28000` | argv ceiling per chunk, with headroom under the Windows 32,767 limit. |
|
||
| `RUN_TESTS_TIMINGS_FILE` | `tests/test-timings.json` | Path to the timing table. Tests override it to inject a synthetic cost profile. |
|
||
| `RUN_TESTS_CHUNK_TIMEOUT_MS` | `600000` | Per-chunk timeout. |
|
||
|
||
The timing table is **advisory and deliberately un-gated**. There is no `--check`
|
||
mode and no CI lint that fails on staleness, because timing data legitimately
|
||
varies run to run. A file missing from the table falls back to the table's mean
|
||
weight (1), and a missing or unparseable table falls back to uniform weight — so
|
||
drift costs chunk *balance*, never a red build. A count-based floor additionally
|
||
guarantees the packer never produces fewer chunks than plain count-based packing
|
||
would, so a badly stale table cannot collapse the suite into a few fat chunks.
|
||
|
||
### CI job timeout budgets: report + near-cap warning (#4036)
|
||
|
||
Every matrixed job — `test` in `test.yml`, `mutate` in
|
||
`mutation.yml`, `smoke` in `install-smoke.yml` — declares a `timeout-minutes`
|
||
cap. `tests/ci-test-job-timeout-budget.test.cjs` enforces that each checked-in
|
||
cap stays at least the **headroom factor** (1.5x) above a documented,
|
||
hand-measured cost for that job — now all three of the jobs above, not just
|
||
`test`/`coverage-gate`/`test-inert` as before.
|
||
|
||
`test-conformance` in `test.yml` (#4591, epic #4589 Phase 2) runs the
|
||
`scripts/lib/platform-conformance-tier.generated.cjs` file list on
|
||
`windows-latest` (sharded three ways) and `macos-latest` (unsharded) — the
|
||
sole gating signal for real-OS coverage (#4603 retired the parallel
|
||
legacy full-matrix safety-net job). It
|
||
also declares a `timeout-minutes` cap and runs the same in-job near-cap check
|
||
described below, but it has no `LANE_COSTS` entry in
|
||
`tests/ci-test-job-timeout-budget.test.cjs` yet — no real, completed
|
||
(non-cancelled) per-shard measurement exists — so the headroom-factor gate
|
||
does not cover it until one lands.
|
||
|
||
Two runtime mechanisms sit on top of that static gate, both new in #4036:
|
||
|
||
- **In-job near-cap check** (`scripts/ci-check-job-near-cap.cjs`) — the last
|
||
step of each of the three jobs computes elapsed-vs-cap from a start-time
|
||
marker recorded as that job's first step. At >=90% of budget it emits a
|
||
`::warning::` annotation (visible in the PR Checks UI) and a
|
||
`$GITHUB_STEP_SUMMARY` block. Advisory only — it never fails the job. Known
|
||
limit: it cannot fire for a job actually killed by the timeout, since a
|
||
killed job never reaches its last step. That case is caught by the second
|
||
mechanism instead.
|
||
- **Scheduled trending report** (`.github/workflows/ci-timeout-report.yml`,
|
||
`scripts/ci-timeout-report.cjs`) — runs daily and on `workflow_dispatch`. It
|
||
polls GitHub's Actions REST API for recently completed jobs across
|
||
`test.yml`, `mutation.yml`, and `install-smoke.yml`, resolves each job's
|
||
declared cap (a literal `timeout-minutes` for `test`/`test-conformance`/`smoke`, or
|
||
`scripts/mutation-matrix.cjs`'s `COVERED[<module>].timeoutMinutes` for
|
||
`mutate`'s per-module shards), and appends any new `(runId, jobName)`
|
||
records to `tests/ci-timeout-budget-history.jsonl`. Unlike the in-job check,
|
||
this also catches jobs killed by an actual timeout breach — GitHub's Jobs
|
||
API still reports `started_at`/`completed_at` for a cancelled job. Each run
|
||
opens a small, data-only PR carrying that run's new rows, since `next` is a
|
||
protected branch and nothing pushes to it directly — the same constraint
|
||
`auto-backmerge.yml` already works within.
|
||
|
||
This does not retune any `timeout-minutes` value, rebalance shard composition,
|
||
or trim what runs in shard 1 — those stay maintainer policy calls made from
|
||
the accumulated history, not something either mechanism decides on its own.
|
||
|
||
### How-to: regenerate the timing table
|
||
|
||
Regenerate when the suite's cost profile has visibly drifted — after adding or
|
||
removing expensive tests, not on a schedule. The input is a `node:test` reporter
|
||
event stream from a `msd-test` run:
|
||
|
||
```bash
|
||
node scripts/gen-test-timings.cjs \
|
||
~/.local/state/msd-test/runs/<run-id>/test-events-linux-node24.jsonl \
|
||
~/.local/state/msd-test/runs/<run-id>/test-events-macos-node24.jsonl
|
||
```
|
||
|
||
Pass every lane you have. A file's recorded time is the **max** across the
|
||
supplied streams, not the mean: the packer exists to keep the *slowest* lane's
|
||
slowest chunk away from the timeout, so the conservative bound is the right one.
|
||
Keys are sorted so a regeneration diff shows only the files whose cost moved.
|
||
|
||
## Best practices for forward-compat (Node 24/26)
|
||
|
||
- Use `process.execPath` when spawning Node in tests so each matrix lane exercises the lane's Node version.
|
||
- Avoid stack-trace or error-message prose assertions. Assert `err.code`, structured JSON fields, or enums — Node minor releases routinely tweak error wording.
|
||
- Prefer `node:test`, `node:assert/strict`, and `node:test` mocks. No external test frameworks.
|
||
- Coverage uses `c8` and propagates `NODE_V8_COVERAGE` through the harness's child process.
|
||
|
||
---
|
||
|
||
## Test strategy: #443 effort + fast_mode engine
|
||
|
||
> Feature: unified cross-provider effort and fast_mode knobs (issue #443).
|
||
> Test files: `tests/model-resolver.test.cjs` (unit),
|
||
> `tests/model-resolver.test.cjs` (integration).
|
||
|
||
### Testing pyramid
|
||
|
||
| Layer | File | What it covers |
|
||
|---|---|---|
|
||
| **Unit** | `feat-443-effort-fast-mode.test.cjs` | Pure logic: cascade rules, clamping, escalation math, malformed config handling, schema key validation. No CLI subprocess. |
|
||
| **Integration** | `feat-443-effort-fast-mode.integration.test.cjs` | Architecture-level invariants: cross-provider validity, totality across the 33-agent registry, CLI JSON contract, config round-trip, fast-mode honesty. Real subprocesses via `runMsdTools`. |
|
||
| **E2E** *(pending)* | *(not yet wired)* | Propagation layer: effort frontmatter / `CLAUDE_CODE_EFFORT_LEVEL` env actually reaching a spawned Claude Code subagent. See "Gaps" below. |
|
||
|
||
### Architectural invariants
|
||
|
||
Each invariant exists to prevent a specific class of production failure.
|
||
|
||
#### (a) Cross-provider validity
|
||
|
||
**What:** `renderEffortForRuntime(runtime, universalEffort).value` must always
|
||
be a member of the runtime's real provider enum. Ground-truth enums are defined
|
||
as local constants in the test — not sourced from the implementation.
|
||
|
||
```
|
||
PROVIDER_EFFORT_ENUMS = {
|
||
claude: Set { 'low', 'medium', 'high', 'xhigh', 'max' } // Anthropic output_config.effort
|
||
codex: Set { 'minimal', 'low', 'medium', 'high', 'xhigh' } // OpenAI model_reasoning_effort
|
||
}
|
||
```
|
||
|
||
**Why:** Passing a value outside these sets results in a 400 from the real API.
|
||
The clamping logic (`max -> xhigh` for codex; `minimal -> low` for claude) must
|
||
hold for every cell of the VALID_EFFORTS × runtimes matrix.
|
||
|
||
#### (b) Param/channel contract
|
||
|
||
**What:** Each runtime exposes a stable `param` string (the native API field
|
||
name) and `channel` (how the value is propagated). Unknown runtimes return
|
||
`param: null, channel: null` and pass the effort value through unchanged.
|
||
|
||
**Why:** Callers read `.param` to construct the dispatch payload. A regression
|
||
here would silently drop effort from subagent invocations.
|
||
|
||
#### (c) Resolve-execution JSON contract
|
||
|
||
**What:** The `msd-tools resolve-execution <agent>` command emits a JSON object
|
||
with all eight keys present and typed correctly: `model` (string), `profile`
|
||
(string), `effort` (VALID_EFFORTS member), `effort_rendered` (string),
|
||
`effort_param` (string|null), `effort_propagation` (string|null), `fast_mode`
|
||
(boolean), `fast_mode_supported` (boolean).
|
||
|
||
**Why:** Orchestrators and workflow dispatchers parse this JSON. A missing or
|
||
mistyped field silently breaks downstream consumers.
|
||
|
||
#### (d) Totality across the real registry
|
||
|
||
**What:** For every agent in the 33-agent registry, `resolveEffortInternal`
|
||
returns a VALID_EFFORTS member (never undefined/null), `resolveFastModeInternal`
|
||
returns a strict boolean, and `renderEffortForRuntime('claude', effort)` stays
|
||
within the claude provider enum.
|
||
|
||
**Why:** A catalog addition that introduces a missing `routingTier` mapping
|
||
would otherwise produce `undefined` and propagate silently.
|
||
|
||
#### (e) Fast-mode honesty invariant
|
||
|
||
**What:** When the runtime is `claude`, `fast_mode_supported` in
|
||
resolve-execution output is always `false`, regardless of the fast_mode config.
|
||
`RUNTIMES_WITH_FAST_MODE` contains only `'api'`.
|
||
|
||
**Why:** Claude Code's `/fast` toggle is session-level only. Emitting
|
||
`fast_mode: true` as frontmatter on a Claude subagent is a silent no-op.
|
||
Advertising `fast_mode_supported: true` for claude would cause orchestrators to
|
||
believe the knob was wired when it is not.
|
||
|
||
#### (f) Precedence first-valid-wins
|
||
|
||
**What:** Both effort and fast_mode use a layered cascade. The test table covers
|
||
all four effort layers (invocation override → agent_overrides →
|
||
routing_tier_defaults → default) and all five fast_mode layers, including the
|
||
case where an invalid value at a higher layer correctly falls through.
|
||
|
||
**Why:** Silent precedence bugs (e.g., a numeric value in agent_overrides not
|
||
being rejected) would override intentional user config.
|
||
|
||
#### (g) Dynamic-routing composition
|
||
|
||
**What:** `resolveEffortForTier` escalates effort by attempt number
|
||
independently of the model tier mapping. The test verifies the effort ladder
|
||
(`low -> medium -> high -> xhigh -> max`), the `max` clamp, the
|
||
`max_escalations` cap, and that `escalate_on_failure: false` suppresses
|
||
escalation entirely.
|
||
|
||
**Why:** Effort escalation and model escalation share configuration
|
||
(`dynamic_routing`) but must operate independently; coupling them would cause
|
||
over-escalation or under-escalation.
|
||
|
||
#### (h) Config-tooling round-trip
|
||
|
||
**What:** `msd-tools config-set` accepts all new key namespaces
|
||
(`effort.default`, `effort.routing_tier_defaults.<tier>`,
|
||
`effort.agent_overrides.<agent>`, `fast_mode.enabled`,
|
||
`fast_mode.routing_tier_defaults.<tier>`, `fast_mode.agent_overrides.<agent>`)
|
||
without an "Unknown config key" error, and values set via `config-set` are
|
||
reflected in `resolve-execution` output.
|
||
|
||
**Why:** The schema validation gate (`VALID_CONFIG_KEYS` + `DYNAMIC_KEY_PATTERNS`)
|
||
is separate from the resolver logic. A key missing from the schema would produce
|
||
a silent write failure and appear as a bug only at runtime.
|
||
|
||
### Coverage targets
|
||
|
||
| Suite | Target |
|
||
|---|---|
|
||
| Unit | Every cascade rule, every fallthrough, every clamp. All function branches in `resolveEffortInternal`, `resolveFastModeInternal`, `resolveEffortForTier`, `renderEffortForRuntime`. |
|
||
| Integration | All 8 architectural invariants. All 33 registered agents. All 6 provider × effort combinations for the valid-enum check. Full config-set key namespace. |
|
||
|
||
### Gaps / not yet covered
|
||
|
||
**E2E orchestrator-spawn-propagation layer (pending follow-up wiring):**
|
||
The integration tests verify that MSD resolves and renders effort values
|
||
correctly. They do NOT verify that the rendered values actually reach a spawned
|
||
Claude Code or Codex subagent at runtime. Specifically uncovered:
|
||
|
||
- `CLAUDE_CODE_EFFORT_LEVEL` env var being set and read by a spawned claude subprocess
|
||
- `output_config.effort` frontmatter key surviving the AGENTS.md template substitution
|
||
- `model_reasoning_effort` field surviving serialization into a Codex API request body
|
||
- Fast-mode `speed: "fast"` field reaching an `api`-runtime request when `fast_mode_supported: true`
|
||
|
||
These require spawning real subagents (or stubs thereof) and asserting on the
|
||
process environment / request payload — a scope that belongs in a future E2E
|
||
suite under `*.slow.test.cjs` or dedicated fixture-driven integration work.
|