Files
msd-core/tests/run-tests-harness.test.cjs
Dennis Alexis Valin Dittrich ad1477d659 enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap

* enhance(260903-m7p): validate configured entrypoints before success

* test(260903-m7p): require pre-success entrypoint validation

* enhance(260903-m7p): gate install success on entrypoints

* test(260903-m7p): cover configured entrypoints across runtimes

* enhance(260903-m7p): cover emitted runtime entrypoints

* fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number

- finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process
  before the new configured-entrypoint assertion throws. Without a HOME +
  config-location-env sandbox that write resolved through the ambient
  environment and landed in the developer's live ~/.gsd (confirmed absent on
  origin/next baseline, present only on this branch — full-suite HERMETICITY
  WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the
  duration of the test, matching the existing in-process finishInstall/
  install() pattern in tests/install.test.cjs (#2665).
- .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder
  (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the
  fork PR number until the upstream PR number is known.

* fix(260903-m7p): repair cross-platform and pre-existing shape fallout

- tests/configured-entrypoint-validation.test.cjs: the win32 branch of
  ensureCodexHooksJsonSessionStart writes a .cmd shim under
  <codexRoot>/hooks/; create that dir in the test (the real installer only
  calls this once hooks/gsd-check-update.js already exists) and assert the
  platform-common entrypoint shape instead of a fixed non-Windows array,
  since win32 legitimately emits two entries (cmd shim + script).
- tests/install.test.cjs: finishInstall's shared settings-json return now
  carries configuredEntrypoints/rollbackInstallerMigrations for every
  runtime on that path (trae included, not just Claude/Cursor/Windsurf);
  update the trae install() exact-shape assertion to match.

* fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash

configuredEntrypointsForHook's shell branch dropped interpreterCandidates
entirely when resolveBashExecutable returned null, unlike the sibling
portableHooks runner entry a few lines below (which correctly falls back to
the literal 'bash' token). Found via agy adversarial review; verified
unreachable through the current call graph (buildHookCommand's own
resolveBashRunner==null gate already short-circuits before
recordConfiguredHookCommand runs), so this is a defensive consistency fix,
not a live-bug patch — kept for the next caller that does not share that
gate.

* chore(260903-m7p): backfill changeset pr number to the opened upstream PR

.changeset/quick-wasps-sing.md carried the fork PR number (16) as a
placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now
open, so record its real number per CONTRIBUTING.md's changeset pr-field
convention.

* fix(#4154): track already-registered hooks for entrypoint validation on update

applySettingsJsonHooks registers each guard hook only if absent, so a hook
already present from a prior install keeps its stale on-disk command. The
new entrypoint tracker always records the freshly-computed command for it,
which never matches what is actually persisted, so the exact-string filter
in finishInstall silently dropped it from validation — the Blocker case
this feature exists to catch (an already-installed entrypoint going stale
between installs) was exactly the case it never validated.

Match on the managed script's basename instead, which the persisted
command carries either way, so an already-registered hook stays in the
validated set. Regression test forces this path by mutating a
freshly-installed hook's persisted command before a second install.

* fix(#4154): distinguish an unreadable script from a missing one

validateConfiguredEntrypoints folded an EACCES statSync failure into the
same 'missing' reason as ENOENT, misreporting a real permission problem as
an absent file. Check the error code and report 'unreadable' instead.

* docs(#4154): document entrypoint validation's rollback and PATH scope

CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had
no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite
bin/install.js x CONTEXT.md being this repo's strongest co-change pairing.

The update-gsd.md how-to overstated what a validation failure undoes: for
Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/
config.toml inside install() before the aggregate validation call runs, so
there is no rollback path for that write regardless of "where available"
phrasing. Also note that interpreter resolution checks the installer's own
PATH, not necessarily the PATH a hook fires under later (#2979 launchers).

* chore(#4154): point changeset pr field at the fork PR while CI runs there

Mirrors the branch's own prior backfill commit: pr: matches whichever PR
number changeset-lint is currently validating against (fork PR #16 during
the fork-first CI/review loop), flipped back to the upstream PR number
right before the final push to open-gsd/gsd-core.

* fix(#4249): address adversarial-review findings in entrypoint validation

An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR
found several real gaps beyond the human reviewer's Blocker, verified
against source before fixing:

- Codex's install() result bound rollbackInstallerMigrations to the narrow
  installer-migrations-only rollback instead of restoreCodexSnapshot (#3245),
  the full pre-install snapshot/restore Codex already owns for exactly this
  case — a validation failure discovered outside install() reverted nothing
  of the config.toml/hooks.json that call had already written.
- The register-only-if-absent basename match from the prior fix used a bare
  substring, which an unrelated user command mentioning the same filename
  could false-positive into GSD's validated set — anchored on the
  `/hooks/<basename>` path segment instead.
- nodeCandidates checked raw process.execPath (always true — we're running
  in that process) instead of normalizeNodePath's stable version-manager
  alias, the same one buildNodeRunnerChainToken bakes as its first choice —
  a false green regardless of whether that alias itself still resolves.
- An entry with no interpreterCandidates (Cline's PreToolUse hook, or a
  Windows-Claude .sh hook invoked without a bash runner) runs via its own
  shebang; validateConfiguredEntrypoints checked only file-type, never the
  execute bit. Cline's writer also never reported an entrypoint at all.
- Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor
  hook registered across several events) were validated once per duplicate.

Each fix is covered by a new or extended test; the Codex one required
inlining runCodexInstall's env sandboxing so the rollback closure — which
re-resolves the $HOME-relative skills root live — runs before the sandbox
is torn down, matching how installAllRuntimes' real aggregate gate calls it.

* docs(#4249): document the round-2 entrypoint-validation fixes

Runtime Hooks Surface Module and Installer Module entries now name
ConfiguredEntrypoint's not-executable reason, the normalizeNodePath
alignment, Cline's tracked hook, and which install() result the
finishInstall/installAllRuntimes rollback path actually reverts per
runtime (Codex's full snapshot vs. the others' narrow migrations-only
rollback).

* chore(#4249): point changeset pr field at the upstream PR now that fork CI is green

* fix(#4249): address agy adversarial-review findings

- validateConfiguredEntrypoints: statSync alone never detects a
  chmod-000 script (it only needs parent-dir search permission), so an
  interpreter-invoked entry with an unreadable script passed validation.
  Add an explicit R_OK check for the interpreterCandidates branch only —
  the candidate-less/shebang branch already has its own X_OK gate.
- docs/how-to/update-gsd.md: the blanket "does not revert" claim was
  false for Codex, which reverts config.toml/hooks.json via its full
  pre-install snapshot; qualify it per runtime.
- tests/codex-config.test.cjs: the #4249 rollback regression test
  asserted skills/ and VERSION were reverted but never asserted
  config.toml/hooks.json were too, despite the test's own stated intent.
- CONTEXT.md: qualify which interpreterCandidates entries get
  normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's
  Windows .cmd shim as a third candidate-less case that relies on
  extension dispatch, not a shebang.

* fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit

Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node',
so it needs the execute bit (like any shebang-invoked entry), but its
interpreter is looked up on PATH by 'env' at hook-fire time (unlike
every other GSD JS hook, which bakes an absolute node path specifically
to avoid that dependency). The candidate-less/interpreterCandidates
fork treated these as mutually exclusive, so Cline's entry silently
skipped interpreter resolution entirely — a completely missing 'node'
on PATH would still validate successfully.

Add an orthogonal selfExecutable flag so both checks run for entries
that need them. (CodeRabbit finding on the fork rehearsal PR.)

* fix(#4249): address second-round adversarial review findings (opus + agy)

- validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry,
  not just interpreterCandidates ones — a self-executable shebang script
  is still opened and read by its kernel-invoked interpreter, so X_OK
  alone never proved it was readable.
- selfExecutable is now the sole, explicit source of truth for the
  execute-bit check (every producer that needs it sets the flag) instead
  of being partly inferred from an absent interpreterCandidates, which
  Cline's hybrid entry also carries.
- The execute-bit check now skips explicitly on win32 (matching
  resolveExecutableBinary's own carve-out) instead of relying on Node's
  accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows
  machine and not a test that simulates win32 on a POSIX runner.
- bin/install.js: fixed a stale comment claiming no runtime's
  install()-time writes have a rollback path — Codex's does
  (restoreCodexSnapshot) — and added the omitted Cline to both that
  comment and CONTEXT.md's equivalent lists.
- CONTEXT.md: fixed the Cline description left stale by the previous
  commit's selfExecutable addition, and rewrote the validation-mechanism
  paragraph for clarity (writing-for-agents pass).
- docs/how-to/update-gsd.md: split an overloaded 4-clause sentence.
- Removed a fault-injection integration test that could not reliably
  exercise the real installAllRuntimes -> finalize -> rollback wiring
  without fighting the installer's own pre-registration existence
  guards; the constituent pieces remain covered individually.

* fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner

X_OK is a POSIX-only concept, skipped entirely when an entry's platform
is win32 (matching production). Two test entries omitted platform,
defaulting to process.platform — on an actual windows-latest CI runner
that silently skipped the very check they were meant to exercise,
turning 'not-executable' into a false pass. Pin platform: 'linux' so
these are deterministic regardless of which OS runs the suite.

* fix(#4249): classify EPERM the same as EACCES in statSync error handling

Windows raises EPERM (not EACCES) for a parent directory that couldn't
be traversed into — was falling through to 'missing', misreporting a
genuine permission problem as a nonexistent path.

* docs(#4249): address final CodeRabbit doc-completeness findings

- CONTEXT.md: install()'s documented result shape omitted
  configuredEntrypoints; the ConfiguredEntrypoint shape omitted
  selfExecutable.
- docs/how-to/update-gsd.md: the failure-mode sentence omitted
  unreadable and lacks-execute-permission, which the installer also
  rejects.

* fix(#4249): stop double-validating every configured entrypoint on install/update

installAllRuntimes' finalize() already runs assertConfiguredEntrypoints
once over the aggregate set; finishInstall then re-ran the identical
check per runtime in the printSummaries loop right after, so every
entrypoint paid its statSync/accessSync/interpreter-resolution cost
twice on every install and update. Add entrypointsAlreadyValidated to
skip the redundant pass specifically on that path, while leaving the
check intact for any caller that invokes finishInstall directly.

* chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there

* perf(#4249): memoize interpreter candidate resolution across entrypoints

resolveExecutableBinary walked PATH once per (entry, candidate) pair; a
typical install has a dozen-plus entries sharing the same few candidate
lists (process.execPath for JS hooks, bash for shell hooks). Cache by
(platform, candidate) so each distinct pair resolves once per validation
call instead of once per entry.

* chore(#4249): point changeset pr field at the rebased rehearsal fork PR

* fix(#4249): drop entrypoint tracking from the now-dead Codex event writer

#2586 (landed on next after this branch forked) removed install.js's
CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent
no longer runs during install or update. The ConfiguredEntrypoint records
this branch added inside it were therefore unreachable and untested. Restore
the function to its upstream shape; the entrypoints it used to report were
never collected by any caller.

* refactor(#4249): drop the revalidation bypass flag and the candidate cache

Both were this PR's own micro-optimisations over a set of roughly a dozen
entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall
gate off to save one statSync/accessSync pass; `resolvedCandidateCache`
memoised resolveExecutableBinary across entries that are already deduped by
(configPath, scriptPath). Neither is measurable, and the flag was the only
way to reach finishInstall with validation disabled. finishInstall now
always validates what it is given.

* chore(#4249): point the changeset pr field back at the upstream PR

* refactor(#4249): track settings.json entrypoints without the hooksSurface gate

The install-surface writer only tracked configured entrypoints when the
runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing
asserts that axis agrees with `installSurface`, so a descriptor that broke
the coupling would silently pass `configuredEntrypoints: undefined` and drop
that runtime out of the validation this PR adds — reintroducing the exact
'reports Done! over a broken entrypoint' failure #4154 exists to close.

Remove the dependence rather than test it: everything recorded on this path
lands in settings.json by construction, and the registered-command filter
already discards entries no persisted hook references.

* chore(#4249): put the changeset body in the documented two-part format

CONTRIBUTING.md and .changeset/README.md both show
`**<bold change>** — <symptom-led explanation>.`; the fragment was a single
unbolded sentence.

* chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there

* fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback

#3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-*
and gsd-core/VERSION. The install overwrites every other GSD-owned file too —
hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest
itself — before the entrypoint-validation gate runs, so a validation failure
left the new payload sitting on top of the restored old config.

Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims,
before runInstallerMigrations so the bytes are the true pre-install state, and
restore it from both Codex rollback closures ahead of the per-surface restores.
Files only the failed install introduced are removed, read from the manifest
now on disk. The manifest is already the authoritative record of what GSD owns,
so no second hand-written list can drift out of sync, and user-owned files are
never snapshotted or removed. Every path is confined through
resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into
an arbitrary-path write.

Non-Codex runtimes are unaffected: the snapshot is gated on the same
tomlConfigInstall + non-minimal condition as #3245's.

* fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest

Two follow-on defects in the previous commit's snapshot:

- The capture was gated on `!isMinimalMode`, copied from #3245. A core/
  --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the
  manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the
  snapshot came back empty while the rollback still ran — and its removal pass
  would have deleted every file the new manifest lists. Gate on
  tomlConfigInstall alone, matching where the rollback actually reaches.

- An unreadable or unparseable prior manifest was caught alongside ENOENT and
  treated as a fresh install. That is the same empty-snapshot state, so a failed
  update over a real install with a corrupt manifest could delete its prior
  payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means
  known-empty; any other read error or a parse failure means unknown, and the
  restore closure returns without touching anything, degrading to #3245's
  narrower rollback. Deliberately not fatal — a corrupt manifest has to stay
  repairable by reinstalling over it.

Both paths are covered by red-checked regression tests.

* fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too

commit removed from the manifest snapshot. restoreCodexSnapshot is reachable
for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-*
skill dir and gsd-* agent file the snapshot does not claim — so with an empty
minimal-mode snapshot a rollback deleted the whole skills/agents surface with
nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via
the ADR-1239 skills-kind home override, so this is also the reason manifest
`skills/` keys do not resolve under configDir: that surface belongs to this
snapshot, not to the manifest-driven one.

Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal
mode — doing nothing on an early failure is the non-destructive side.

Covered by a red-checked regression test that plants bytes in an alternate-home
skill file, reinstalls under the core profile marker, and asserts the rollback
restores it.

* fix(#4249): never remove on rollback unless a prior manifest proves what predates the install

Three defects in the manifest-driven Codex rollback, all in its removal half:

- ENOENT marked the snapshot usable, arming the removal pass on a FIRST
  install. GSD may have overwritten a user's file at a manifest-tracked path
  there, and no prior manifest records the difference — so rollback deleted it
  where before it merely left it overwritten. Absent, unreadable and malformed
  manifests now all leave the prior set UNKNOWN and skip removal entirely.

- Membership was tested against the map of files whose pre-install read
  SUCCEEDED, so a tracked file that existed but was unreadable read as
  introduced-by-this-install and was removed. Track the prior manifest's paths
  in their own Set and test against that.

- The unreachable "delete the manifest when there was no prior one" branch is
  gone: usable now implies a parsed prior manifest.

Also adds the end-to-end test the aggregate gate was missing — the four Codex
rollback tests drove the closure directly, proving the restore but not the
wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes
Cline's `env node` entry fail validation for real, and asserts Codex's payload
comes back.

Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper
instead of six copies. Both new tests are red-checked.

* test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test

lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its
Windows-EBUSY retry budget. That budget is for directory trees; this removes a
single file, which unlinkSync says more precisely and the rule does not flag.

* chore(#4249): point the changeset pr field back at the upstream PR

* fix(#4249): use an unambiguous dedup key and surface partial-restore failures

trek-e's 2026-09-08 adversarial pass flagged two findings in the new
entrypoint-validation/rollback code:
- assertConfiguredEntrypoints' dedup key already used a raw NUL
  separator (introduced in ceebb65f2d), but git/Read render NUL as a
  space, so the key looked like a plain-space join to every reviewer
  that read the diff. Replace it with JSON.stringify([configPath,
  scriptPath]) so the separator is visible and unambiguous.
- restoreManagedFileSnapshot's per-file restore catch block claimed to
  'surface the original error' but only swallowed it, matching (and
  widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add
  an actual console.warn using the existing best-effort-warning
  convention, scoped to just this PR's new function.

* fix(#4249): treat a files-less prior manifest as unknown, not known-empty

agy's gemini-3.8-flash-high adversarial pass (round 5) found and I
reproduced empirically: a structurally-valid manifest missing the
files key (e.g. {"version":1}) parses without throwing, so
Object.keys(undefined || {}) silently read as 'zero files predate
this install' instead of the UNKNOWN state the malformed-manifest
guard exists to produce. Rollback's removal pass then deleted every
GSD-owned file the failed install's own manifest listed, including
ones that predated it — the exact data loss the #4249 CodeRabbit
malformed-manifest fix was supposed to prevent, reachable through a
JSON.parse success instead of a failure. Route the shapeless case
into the same catch-all UNKNOWN path via an explicit shape check.
Regression test reproduces the deletion before the fix and confirms
the file survives after it.

Also extend restoreManagedFileSnapshot's removal-pass rmSync and
final manifest-rewrite catches with the same real console.warn
trek-e's round-4 review asked for on the per-file restore catch —
same rollback function, same operator-facing-signal gap.

* docs(#4249): correct which runtimes actually leave a written config on rollback

agy's completeness audit (round 5, holistic pass) caught this new
paragraph claiming 'for every other runtime, the configuration file(s)
already written during that update are left in place' — false for
Claude Code and other settings.json-based runtimes, whose write never
happens on failure (assertConfiguredEntrypoints runs before
finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline
actually match that description, since they persist their config file
inside install() ahead of the gate. Split the one sentence into the
three actual outcomes; matches the PR body's own accurate Before/After
wording, which this doc addition had drifted from.

* fix(#4249): clean up doc/comment mismatches and dead fields from opus review

Opus critical-code-reviewer + ponytail-review pass on the final diff:
- assertConfiguredEntrypoints carried finishInstall's old docblock
  ("Apply statusline config, then print completion message") from
  before this function was inserted between comment and callee.
  finishInstall already has its own accurate #4249 comment, so the
  stale docblock is removed rather than moved.
- checked: number on ConfiguredEntrypointValidationResult and
  error.configuredEntrypointValidation on the thrown error: the first
  had zero consumers anywhere in the repo, including its own defining
  file, and is removed. The second matches an existing repo
  convention (bin/install.js's installerMigrationRollbackFailures,
  #4249 predates this PR) of attaching structured diagnostic context
  to a re-thrown Error even before a consumer exists, so it's kept.
- finishInstall's own assertConfiguredEntrypoints call is a redundant
  backstop on the real production path (installAllRuntimes's aggregate
  call already validates the superset first), but its comment read as
  though this call alone provided the before-the-write guarantee.
  Clarified rather than removed — it's the only gate for a caller that
  invokes finishInstall directly.

* chore(#4249): split the manifest-driven rollback engine out into #4544

Issue #4154 asked the installer to consume a validation failure "through
the existing rollback mechanism, without a second transaction mechanism".
The manifest-driven rollback widening added during review (capture every
path the prior gsd-file-manifest.json claims, restore those bytes, remove
what only the failed install introduced) is that second mechanism on a
plain reading. It is a real fix for a #3245-era gap, but an independent
one, so it moves to its own bug report and PR.

Removed here:
- bin/install.js: the pre-install managed-file capture block and
  restoreManagedFileSnapshot, plus its call sites in
  _codexPreConfigRollback and restoreCodexSnapshot (99 lines).
- tests/configured-entrypoint-validation.test.cjs: the five tests that
  exercise the manifest engine.
- CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the
  widened restore. update-gsd.md again documents the #3245 surfaces only.

Kept, because it is #4154's own scope:
- the entrypoint-validation gate itself;
- Codex's install() result binding rollbackInstallerMigrations to
  restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*,
  agents/gsd-*, gsd-core/VERSION);
- the !isMinimalMode gate removal on that snapshot. Binding the closure
  to the result made it reachable for a core/--minimal install, where its
  pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot
  does not claim; an empty minimal-mode snapshot therefore deleted the
  whole surface with nothing to restore.

The surviving aggregate-failure test now asserts on config.toml, a
surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which
only the manifest engine restored.

Refs #4544

* test(#4249): cover configured entrypoints through the packed install path

#4154's scope lists install smoke coverage alongside the installer gate —
"assert representative configured entrypoints resolve for supported runtime
profiles". The gate itself (assertConfiguredEntrypoints /
validateConfiguredEntrypoints) is unit-covered by in-process install() calls;
nothing proved the property survives npm pack -> npm install -g -> install.js.

Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct
config surfaces GSD writes launch paths into (settings.json, and hooks.json +
config.toml) — run the tarball-installed installer into a throwaway HOME, then
re-read that runtime's own written config and return the new
ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a
file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check
becomes a release gate on every matrix host without workflow changes.

The scan re-derives paths from the written config instead of reusing the
installer's own entrypoint list, and test I shows why that matters: a
registration the installer never touched during a run is invisible to the
in-process gate, so the install exits 0 and only reading the config back off
disk catches the dangling launch path.

* ci(#4249): pack a publish-shaped tarball in the install smoke lane

`npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs
build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs
after `npm ci` carries no hook scripts at all — the lane has been smoking a
package that differs from the published one in exactly the artifacts the
lifecycle smoke is supposed to launch.

That went unnoticed because the lane's init runs `--local`, which registers no
statusline and therefore registers no hook whose target is missing. A
`--global` install on the same tarball exits 1 on #4249's own gate
(`gsd-statusline.js (missing)`), which is what the new configured-entrypoint
cycle performs, so without this step the cycle would report INIT_FAILED
instead of checking anything.

Build hooks before packing so the smoked tarball matches prepublishOnly. The
CLI now reports 16 configured entrypoints for claude and 1 for codex instead
of zero.

* fix(#4249): scope Codex's full snapshot restore to entrypoint failures

Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage
exception un-install a Codex install that had already succeeded and already
printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps
the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints`
call.

Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration`
scopes finalize-stage rollback to installer *migrations* ("the executor uses the
journal to restore modified paths"), and this PR's own operator-facing paragraph
in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/
agents/VERSION revert to entrypoint-validation failures specifically ("If a
script is missing, unreadable, ... For Codex, this reverts ..."). The wide
behaviour is also incoherent as a transaction abort: the same doc says Cursor,
Windsurf, Kimi and Cline keep the config they wrote inside install().

Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall
hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install
bytes — on an update, silently downgrading a working Codex install to the
previous version while the user had just been told it was Done.

Select the rollback by error kind instead. `assertConfiguredEntrypoints` already
tags its error with `configuredEntrypointValidation`, so the full snapshot
restore runs for that error (and anything downstream of it, including
finishInstall's per-runtime backstop) and the installer-migrations-only closure
runs for everything else. The codex result now also exposes that narrow closure
as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning
the full restore, so the direct-call contract asserted by
tests/codex-config.test.cjs is unchanged.

Adds a regression test that installs codex+kilo together, injects EACCES on the
Kilo permission write by monkeypatching node:fs (restored in a finally — never
chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps
the bytes the successful install wrote. Verified red against the pre-fix
unconditional path.

Cline cannot host this test: its plan is writesSharedSettings:false +
finishPermissionWriter:null, so its finishInstall performs no write and has no
non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally
(unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded
fs.writeFileSync.

* docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing

CONTEXT.md still described Codex's rollback as an unconditional bind
to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint-
validation failures via rollbackInstallerMigrationsOnly and the
configuredEntrypointValidation error tag. Caught during the round-6
PR body pass.

* fix(#4249): stop rollbackInstallerMigrations meaning its own opposite

Codex's install() result bound `rollbackInstallerMigrations` to
restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the
actual installer-migrations-only closure behind
`rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name
meant the opposite of what it says, and CONTEXT.md had to concede as much
in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure
for every runtime, matching both its name and the meaning it already has
on next, and the snapshot restore gets its own Codex-only field,
`rollbackPreInstallSnapshot`. The selection in
rollbackFinalizedInstallerMigrations collapses to one line and no longer
needs a fallback chain.

Also in this commit, all against the same rollback path:

- Correct the rollbackFinalizedInstallerMigrations comment. It read as if
  the round-6 narrowing prevented any sibling-triggered revert of a Codex
  install the user has already seen "Done!" for. It does not, and is not
  meant to: `wide` is true for ANY entrypoint-validation error from ANY
  runtime, because the aggregate gate is all-or-nothing — an invalid Cline
  entrypoint reverts Codex's snapshot, which
  tests/configured-entrypoint-validation.test.cjs's 'an aggregate
  entrypoint validation failure rolls the Codex install back (#4249)'
  asserts directly. The discriminator is the error's KIND, not which
  runtime owns the failing path. Comment and CONTEXT.md now say that.

- Name the runtime in the "Configured entrypoint validation failed" error.
  ConfiguredEntrypointInvalid already carries `runtime`; the message threw
  it away, leaving an operator of a multi-runtime install unable to tell
  whose entrypoint broke — which matters precisely because the failure can
  revert a runtime that was itself fine.

- Set `configuredEntrypoints: []` explicitly on the copilot-instructions
  early return. Every other branch states the key; this one relied on
  installAllRuntimes' `(result.configuredEntrypoints || [])` defence.
  `[]` is correct, not a workaround: every Copilot hook is an inline
  printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no
  GSD-managed script or interpreter to resolve.

No behaviour change beyond the error-message text.

* docs(#4249): narrow the smoke scan's config-surface claim to what it checks

RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is
told to launch is registered in one of settings.json / hooks.json /
config.toml, and that nothing else in a config dir is runtime
configuration. Both halves are false as stated. Cline registers its hook
at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those
names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native
[[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi),
a directory separate from Kimi's own GSD configDir — the same gap
installer-migration 007 already documents as structurally unreachable.

The scan is in fact correct for what it runs against: entrypointRuntimes
defaults to claude + codex, whose launch paths do all live in those three
top-level files. Restate the docstring at that scope, name the two known
out-of-scope surfaces, and warn that adding either runtime to
entrypointRuntimes without teaching scanConfiguredEntrypoints about its
surface yields a scan that finds zero entrypoints and proves nothing.
The entrypointRuntimes default comment carried the same overgeneralization
("every other runtime reuses one of them") and is corrected with it.

Documentation only; no code change.

* fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json

An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found
that install()'s settings-json early return for an unparseable
settings.local.json omitted configuredEntrypoints and
rollbackInstallerMigrations from its result, unlike every other branch.
rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations
unconditionally, so this branch silently dropped its own installer-migration
rollback on a later finalize-stage failure.

Completed the return shape: configuredEntrypoints: [] (matching Copilot's
equally-early no-entrypoints-yet return) and rollbackInstallerMigrations
(already in closure scope). Red-then-green regression test added.

* test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs

Same internal adversarial review (round 8): this suite's before() packed the
tarball directly, without the ensureHooksDist() guard every sibling
install-test suite (install.test.cjs, install-minimal-hooks.test.cjs,
mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run
in isolation ahead of a suite that builds hooks/dist itself, this suite's
pack would ship a tarball with no hook scripts and fail closed on
SMOKE.INIT_FAILED instead of testing anything.

* fix(#4249): refresh stale test-timings weight for the codex-config split

next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's
heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the
CI shard packer's weight table (tests/test-timings.json) was never updated:
codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the
suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block
this PR extends with its own #4249 install()-pipeline test — had no entry at
all, so the packer would silently underestimate it at the table's median
weight (roughly a 9x underestimate against its real cost).

trek-e's most recent review flagged a Windows shard timeout in-flight on
codex-config.test.cjs, plausibly aggravated by this PR's own addition to that
file before the rebase moved it. Re-measured both files locally (node --test
--test-reporter=tap, max of 3 runs, matching the table's own max-across-streams
methodology) and patched just these two entries — not a full regeneration,
which would need real multi-lane CI data this session doesn't have access to.

* fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists

next's platform-conformance-tier classifier (#4591/#4598) landed after this
branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and
tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's
gen-platform-conformance-tier --check and both the Linux and macOS conformance
suites.

* fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation

next's #4733 (landed after this branch's last rebase) replaced the static
ISOLATED_HEAVY_FILES set with a threshold derived live from
tests/test-timings.json, and pins the current derived result in
EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still
named codex-config.test.cjs, whose own weight this PR already dropped from
127783ms to 189ms (after splitting its heavy install()-pipeline blocks into
codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The
live-computed set correctly no longer includes it; the pinned expectation is
updated to match.

* fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it

trek-e's review flagged two Major gaps: the thrown error read identically
regardless of which of three real outcomes a runtime hit (nothing
persisted / snapshot reverted / config left broken on disk), and no test
proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/
Kimi/Cline — only Codex's revert path was ever asserted.

assertConfiguredEntrypoints now tags each invalid entry with its actual
consequence, mirrored from docs/how-to/update-gsd.md's existing
rollback-matrix disclosure. A new test drives the same aggregate failure
through Cline (whose own entrypoint is the one that fails) and asserts
its hook file is still on disk afterward.

* fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR

One review pass (gemini-3.8-flash-high via the antigravity review lane)
against this PR's full diff against next, findings independently verified
against source before fixing:

- Copilot's install() return object was the only one of 6 runtime branches
  missing rollbackInstallerMigrations — reachable now that this PR's own
  aggregate gate runs rollback across every result on any runtime's
  entrypoint failure, not just Copilot's own.
- buildHookCommand's unresolved-bash early return skipped track() entirely,
  so a win32 install with no Git Bash silently produced an unregistered
  .sh hook instead of the 'unresolved-interpreter' validation failure
  configuredEntrypointsForHook's own comment said it would.
- release-tarball-smoke.cjs reported a Cycle 4 install failure under
  SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing
  SMOKE.INSTALL_FAILED.
- SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's
  trailing args, which also truncated any configDir containing a space
  (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan.
  Anchored the match on the already-known configDir prefix instead of a
  generic absolute-path guess: removes the ambiguity outright rather than
  patching the character class, and stays a raw-text scan on purpose (it
  catches a writer that emits a path without registering it — a
  JSON.parse of the expected schema would miss exactly that case).

One suggested finding (test-timings.json "missing" the new test file) was
verified false — that table only holds measured CI timings, populated
after a file's first real run — and one Ponytail suggestion (a
JSON.stringify dedup key) was rejected as it would reintroduce a real, if
narrow, key-collision risk for no benefit.

* fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment

changeset-lint requires pr: to match the PR it runs on (16 on the fork,
not the eventual upstream number) — rehearsal-branch convention already
established earlier in this PR's history.

lint-docs-guard-registration's quote-pairing heuristic doesn't require the
docs/ path itself to be quoted — it flags a file once ANY quote-delimited
span containing "docs/" appears anywhere in it, alongside any real fs read
call. A comment ending "...update-gsd.md's rollback-matrix paragraph"
supplied the closing quote character (the possessive apostrophe) the
heuristic paired with an unrelated single-quoted string earlier in the
file. Reworded to avoid the unquoted apostrophe next to the path.

* chore(#4249): point the changeset pr field back at the upstream PR

Fork rehearsal (PR #16) is green; the real target for this changeset is
upstream PR #4249.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 04:23:53 -04:00

3371 lines
158 KiB
JavaScript

// docs-guard-exempt: docs/TESTING-SUITES.md is cited only in a header comment; never read.
// allow-test-rule: pending-migration-to-typed-ir [#3090]
// run-tests.cjs is a CLI test harness with no --json/structured output mode;
// these tests regex/substring-match its human-readable stderr (usage errors,
// the `run-tests: suite="X" files=N: name1 name2 ...` selection line) instead
// of a frozen typed IR. Adding a structured output mode is a production
// change out of scope here. See docs/TESTING-SUITES.md and issue #3597.
// Tracked under #3090.
//
// Tests for scripts/run-tests.cjs --suite filtering (issue #3597).
//
// Drives the harness through its subprocess seam — the same seam CI uses —
// rather than importing internals. Each test seeds a temporary directory
// with mock `.test.cjs` files (each one a trivial node:test no-op) and
// runs the harness against it via GSD_TEST_DIR.
'use strict';
const { describe, test, before, after, beforeEach, afterEach } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const { spawn } = require('node:child_process');
const { runNode } = require('./helpers/process-seam.cjs');
const { toLegacyResult } = require('./helpers/git-fixture.cjs');
const { createTempDir, cleanup, CONFIG_LOCATION_ENV_KEYS } = require('./helpers.cjs');
const { splitLines } = require('../gsd-core/bin/lib/text-lines.cjs');
const HARNESS = path.join(__dirname, '..', 'scripts', 'run-tests.cjs');
// The harness under test enforces its OWN per-chunk timeout internally
// (RUN_TESTS_CHUNK_TIMEOUT_MS, default 600000ms; this file's slowest explicit
// override below is 30000ms). This outer bound must stay comfortably above
// whatever the harness itself is configured to wait for a hung chunk, plus
// `node --test` child-process startup overhead — otherwise this seam would
// kill the harness before its own timeout diagnostic fires.
const HARNESS_TIMEOUT_MS = 120000;
// Minimal valid node:test file. Each fixture file passes when executed.
const PASS_BODY = `'use strict';
const { test } = require('node:test');
test('noop', () => {});
`;
function seed(dir, names) {
for (const name of names) {
const full = path.join(dir, name);
fs.mkdirSync(path.dirname(full), { recursive: true });
fs.writeFileSync(full, PASS_BODY, 'utf8');
}
}
function runHarness(testDir, args = [], extraEnv = {}) {
// Clear node:test parent-context env so the harness's child `node --test`
// doesn't refuse to run with "recursive run() skipping running files".
const env = { ...process.env, GSD_TEST_DIR: testDir, ...extraEnv };
delete env.NODE_TEST_CONTEXT;
// #4070: strip RUN_TESTS_SHARD_RESERVE inherited from the OUTER job's own
// environment. test.yml sets it on the "Run unit tests" step for the real
// production shard 1 of the full-scope lane — and since these tests spawn
// run-tests.cjs as a CHILD of that same step, they inherit it via
// `...process.env` above like any other ambient var. Left unstripped, a
// reserve of 77 weight units utterly dwarfs these synthetic 9-file
// fixtures' combined weight (~0.3, since none of them are in the real
// timings table), so shard index 1 gets EVERY file routed away from it —
// a real, reproducible corruption of every test in this describe block,
// not a flake (confirmed live: CI run 33288554040, shard 2/3, 7 of these
// tests failed with exactly this signature). Deleted before `extraEnv` is
// applied above would be too late (spread order), so it is deleted here,
// AFTER composition, then only reinstated if a specific test opted in via
// extraEnv — preserving this file's one legitimate use (the #4070 E2E
// bounds-check test below, which sets it deliberately).
if (!Object.prototype.hasOwnProperty.call(extraEnv, 'RUN_TESTS_SHARD_RESERVE')) {
delete env.RUN_TESTS_SHARD_RESERVE;
}
const r = runNode([HARNESS, ...args], {
cwd: path.join(__dirname, '..'),
env,
timeoutMs: HARNESS_TIMEOUT_MS,
});
// toLegacyResult() alone drops `signal` (several assertions below embed it
// in their failure message) — compose it back on top, per git-fixture.cjs's
// documented "extra field" composition pattern.
return { ...toLegacyResult(r), signal: r.signal };
}
describe('run-tests.cjs harness (issue #3597)', () => {
let tmpDir;
beforeEach(() => {
tmpDir = createTempDir('gsd-3597-harness-');
});
afterEach(() => {
cleanup(tmpDir);
});
describe('argument parsing', () => {
test('unknown suite name exits non-zero with valid-suites hint', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'bogus']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /unknown suite/i);
assert.match(r.stderr, /unit/);
assert.match(r.stderr, /security/);
});
test('missing --suite value exits non-zero', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--suite']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /requires a value/i);
});
test('duplicate --suite flag is rejected', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'unit', '--suite', 'security']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /duplicate/i);
});
test('unknown positional argument is rejected', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['unit']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /unknown argument/i);
});
test('--suite=value syntax is accepted', () => {
seed(tmpDir, ['a.test.cjs', 'b.security.test.cjs']);
const r = runHarness(tmpDir, ['--suite=security']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}\nstdout: ${r.stdout}`);
});
test('missing --files value exits non-zero', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--files']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /--files requires a value/i);
});
test('duplicate --files flag is rejected', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--files', 'a.test.cjs', '--files', 'a.test.cjs']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /duplicate --files/i);
});
test('--files and --files-from cannot be combined', () => {
seed(tmpDir, ['a.test.cjs']);
const listPath = path.join(tmpDir, 'selected-tests.txt');
fs.writeFileSync(listPath, 'a.test.cjs\n', 'utf8');
const r = runHarness(tmpDir, ['--files', 'a.test.cjs', '--files-from', listPath]);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /cannot be combined/i);
});
});
describe('suite filtering', () => {
test('no flag runs ALL test files (backcompat)', () => {
seed(tmpDir, [
'a.test.cjs',
'b.security.test.cjs',
'c.integration.test.cjs',
]);
const r = runHarness(tmpDir);
assert.strictEqual(r.status, 0);
// node:test TAP output mentions each file path.
assert.ok(r.stderr.includes('a.test.cjs'), 'expected a.test.cjs in output');
assert.ok(
r.stderr.includes('b.security.test.cjs'),
'expected b.security.test.cjs in output',
);
assert.ok(
r.stderr.includes('c.integration.test.cjs'),
'expected c.integration.test.cjs in output',
);
});
test('--suite all is equivalent to no flag', () => {
seed(tmpDir, ['a.test.cjs', 'b.security.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'all']);
assert.strictEqual(r.status, 0);
assert.ok(r.stderr.includes('a.test.cjs'));
assert.ok(r.stderr.includes('b.security.test.cjs'));
});
test('--suite unit excludes marked suites', () => {
seed(tmpDir, [
'a.test.cjs',
'b.security.test.cjs',
'c.integration.test.cjs',
'd.install.test.cjs',
'e.slow.test.cjs',
]);
const r = runHarness(tmpDir, ['--suite', 'unit']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.ok(r.stderr.includes('a.test.cjs'));
assert.ok(!r.stderr.includes('b.security.test.cjs'));
assert.ok(!r.stderr.includes('c.integration.test.cjs'));
assert.ok(!r.stderr.includes('d.install.test.cjs'));
assert.ok(!r.stderr.includes('e.slow.test.cjs'));
});
test('--suite security selects only *.security.test.cjs', () => {
seed(tmpDir, [
'a.test.cjs',
'b.security.test.cjs',
'c.integration.test.cjs',
]);
const r = runHarness(tmpDir, ['--suite', 'security']);
assert.strictEqual(r.status, 0);
assert.ok(r.stderr.includes('b.security.test.cjs'));
assert.ok(!r.stderr.includes('a.test.cjs'));
assert.ok(!r.stderr.includes('c.integration.test.cjs'));
});
test('--suite integration selects only *.integration.test.cjs', () => {
seed(tmpDir, ['a.test.cjs', 'b.integration.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'integration']);
assert.strictEqual(r.status, 0);
assert.ok(r.stderr.includes('b.integration.test.cjs'));
assert.ok(!r.stderr.includes('a.test.cjs'));
});
test('--suite install selects only *.install.test.cjs', () => {
seed(tmpDir, ['a.test.cjs', 'b.install.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'install']);
assert.strictEqual(r.status, 0);
assert.ok(r.stderr.includes('b.install.test.cjs'));
});
test('--suite slow selects only *.slow.test.cjs', () => {
seed(tmpDir, ['a.test.cjs', 'b.slow.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'slow']);
assert.strictEqual(r.status, 0);
assert.ok(r.stderr.includes('b.slow.test.cjs'));
});
});
describe('empty-suite behavior', () => {
test('--suite security with zero matching files exits non-zero with an error', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'security']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /0 test files selected/i);
});
test('GSD_ALLOW_EMPTY_SUITE=1 downgrades empty suite to a warning and exits 0', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--suite', 'security'], { GSD_ALLOW_EMPTY_SUITE: '1' });
assert.strictEqual(r.status, 0);
assert.match(r.stderr, /WARNING.*0 test files selected/i);
});
test('completely empty test dir still exits non-zero (preserves prior behavior)', () => {
const r = runHarness(tmpDir);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /no test files/i);
});
});
describe('explicit file selection', () => {
test('--files runs only the named tests', () => {
seed(tmpDir, ['a.test.cjs', 'b.security.test.cjs', 'c.test.cjs']);
const r = runHarness(tmpDir, ['--files', 'a.test.cjs tests/c.test.cjs']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.ok(r.stderr.includes('a.test.cjs'));
assert.ok(r.stderr.includes('c.test.cjs'));
assert.ok(!r.stderr.includes('b.security.test.cjs'));
});
test('--files-from runs tests listed in a file', () => {
seed(tmpDir, ['a.test.cjs', 'b.security.test.cjs', 'c.test.cjs']);
const listPath = path.join(tmpDir, 'selected-tests.txt');
fs.writeFileSync(listPath, 'a.test.cjs\nb.security.test.cjs\n', 'utf8');
const r = runHarness(tmpDir, ['--files-from', listPath]);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.ok(r.stderr.includes('a.test.cjs'));
assert.ok(r.stderr.includes('b.security.test.cjs'));
assert.ok(!r.stderr.includes('c.test.cjs'));
});
test('missing explicit test file exits non-zero', () => {
seed(tmpDir, ['a.test.cjs']);
const r = runHarness(tmpDir, ['--files', 'a.test.cjs missing.test.cjs']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /requested test file\(s\) not found: missing\.test\.cjs/i);
});
});
describe('subdir file matching (findings #1 and #9)', () => {
test('bare basename resolves to its single subdir file', () => {
seed(tmpDir, ['sub/001-foo.test.cjs', 'b.test.cjs']);
const r = runHarness(tmpDir, ['--files', '001-foo.test.cjs']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.ok(r.stderr.includes('001-foo.test.cjs'));
assert.ok(!r.stderr.includes('b.test.cjs'));
});
test('full subdir relpath matches exactly', () => {
seed(tmpDir, ['sub/001-foo.test.cjs', 'b.test.cjs']);
const r = runHarness(tmpDir, ['--files', 'sub/001-foo.test.cjs']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.ok(r.stderr.includes('001-foo.test.cjs'));
assert.ok(!r.stderr.includes('b.test.cjs'));
});
test('backslash-separated subdir path resolves on all platforms', () => {
seed(tmpDir, ['sub/001-foo.test.cjs', 'b.test.cjs']);
// Simulate a Windows caller passing backslash path
const r = runHarness(tmpDir, ['--files', 'sub\\001-foo.test.cjs']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.ok(r.stderr.includes('001-foo.test.cjs'));
});
test('tests/ prefix is stripped before subdir matching', () => {
seed(tmpDir, ['sub/001-foo.test.cjs']);
const r = runHarness(tmpDir, ['--files', 'tests/sub/001-foo.test.cjs']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.ok(r.stderr.includes('001-foo.test.cjs'));
});
test('ambiguous bare basename exits non-zero with clear error', () => {
seed(tmpDir, ['sub1/dup.test.cjs', 'sub2/dup.test.cjs']);
const r = runHarness(tmpDir, ['--files', 'dup.test.cjs']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /ambiguous basename/i);
assert.match(r.stderr, /dup\.test\.cjs/);
assert.match(r.stderr, /subdir path/i);
});
});
describe('failure propagation', () => {
test('non-zero from node:test propagates through harness', () => {
const FAIL = `'use strict';
const { test } = require('node:test');
test('boom', () => { throw new Error('intentional'); });
`;
fs.writeFileSync(path.join(tmpDir, 'a.test.cjs'), FAIL, 'utf8');
const r = runHarness(tmpDir);
assert.notStrictEqual(
r.status,
0,
`expected non-zero exit; got status=${r.status} signal=${r.signal}\nSTDOUT:\n${r.stdout}\nSTDERR:\n${r.stderr}`,
);
});
});
describe('env hermeticity', () => {
// Regression guard for the two `delete process.env.GSD_PROJECT/GSD_WORKSTREAM`
// lines added in scripts/run-tests.cjs main() right after ensureBuiltArtifacts().
// If those deletions are removed, the fixture's assertions fail inside the child
// node:test process → non-zero harness exit → this test fails → CI catches it.
test('harness strips GSD_PROJECT and GSD_WORKSTREAM before running child tests', () => {
// Write a fixture that asserts both vars are absent in the child process env.
const FIXTURE = `'use strict';
const { test } = require('node:test');
const assert = require('node:assert/strict');
test('ambient GSD workstream vars are stripped by the runner', () => {
assert.strictEqual(process.env.GSD_PROJECT, undefined);
assert.strictEqual(process.env.GSD_WORKSTREAM, undefined);
});
`;
fs.writeFileSync(path.join(tmpDir, 'env-hermeticity.test.cjs'), FIXTURE, 'utf8');
// Pass both vars in the ambient env given to the harness process.
// The harness must delete them before spawning the child node:test process.
const r = runHarness(tmpDir, [], {
GSD_PROJECT: 'ambient-proj',
GSD_WORKSTREAM: 'ambient-ws',
});
assert.strictEqual(r.status, 0, r.stderr);
});
});
describe('Windows argv-overflow chunking (issue #3597)', () => {
// Windows CreateProcess caps lpCommandLine at 32,767 chars. With ~550
// tests the unchunked spawn fails instantly on Windows with no test
// output. Linux/macOS allow ~2 MB so the same path works there. The
// harness chunks selected files so each spawn stays under the ceiling,
// and chunking is observable via the `run-tests: chunk N/M …` stderr
// line. Long filenames force chunking even with a modest file count so
// the test stays fast on every platform.
test('chunks when total argv would exceed configured ceiling', () => {
// Use a deliberately low MAX_CMDLINE_CHARS so the test is independent
// of tmp-path length (varies by OS). With a 2000-char ceiling and 30
// tests at ≥100 char paths, chunking must engage and at least one
// `chunk N/M …` marker must appear in stderr.
const longPrefix = 'a-deliberately-long-test-filename-to-force-chunking-behavior-cross-platform-';
const names = Array.from({ length: 30 }, (_, i) => `${longPrefix}${String(i).padStart(4, '0')}.test.cjs`);
seed(tmpDir, names);
const r = runHarness(tmpDir, [], { RUN_TESTS_MAX_CMDLINE_CHARS: '2000' });
assert.strictEqual(
r.status,
0,
`expected zero exit; got status=${r.status} signal=${r.signal}\nSTDERR (tail):\n${r.stderr.split('\n').slice(-20).join('\n')}`,
);
assert.match(
r.stderr,
/run-tests: chunk \d+\/\d+ — \d+ files/,
`expected chunking marker in stderr; STDERR (tail):\n${r.stderr.split('\n').slice(-20).join('\n')}`,
);
});
test('chunks by file count even when argv length is below the ceiling', () => {
const names = Array.from({ length: 7 }, (_, i) => `tiny-${String(i).padStart(2, '0')}.test.cjs`);
seed(tmpDir, names);
const r = runHarness(tmpDir, [], {
RUN_TESTS_MAX_CMDLINE_CHARS: '100000',
RUN_TESTS_MAX_FILES_PER_CHUNK: '3',
});
assert.strictEqual(
r.status,
0,
`expected zero exit; got status=${r.status} signal=${r.signal}\nSTDERR:\n${r.stderr}`,
);
assert.match(
r.stderr,
/run-tests: chunk 1\/3 — 3 files/,
`expected file-count chunking marker in stderr; STDERR:\n${r.stderr}`,
);
// #2456 changed the packer from sequential first-fit to LPT, so 7 equal-cost
// files across 3 chunks now balance {3,2,2} instead of filling greedily to
// {3,3,1}. The chunk COUNT is unchanged; only the tail is no longer starved.
assert.match(
r.stderr,
/run-tests: chunk 3\/3 — 2 files/,
`expected final file-count chunking marker in stderr; STDERR:\n${r.stderr}`,
);
});
// #2088: expensive files must never all land in one chunk — otherwise the
// unsharded targeted lane packs the whole install surface into a single chunk
// that blows the 600s per-chunk backstop on the slow Windows runner. #2088
// approximated "expensive" from the filename prefix; #2456 replaced that with
// measured durations. This test carries the #2088 GUARANTEE forward onto the
// new mechanism: the cost signal is now the timings table, injected via
// RUN_TESTS_TIMINGS_FILE so the assertion does not depend on the real suite's
// (regenerable, drifting) cost profile.
function writeTimings(dir, timings) {
const p = path.join(dir, 'timings.json');
fs.writeFileSync(p, JSON.stringify({ schema_version: 1, unit: 'ms', timings }), 'utf8');
return p;
}
test('expensive files SPREAD across chunks instead of clustering (#2088, #2456, #4733)', () => {
// Discriminating by construction: the three EXPENSIVE files carry names the
// old prefix heuristic scored 1, and the three TRIVIAL ones carry the
// `install-` prefix it scored 12 — i.e. exactly inverted from their real
// cost. Under the old packer this packs {2,2,1,1} (4 chunks, with two
// expensive files sharing chunk 1).
//
// #4733: isolation is now an ABSOLUTE ms bar (ISOLATION_BUDGET_FRACTION
// * CHUNK_WORKING_BUDGET_MS = 0.3 * 400000 = 120000ms), converted to the
// packer's weight units via the LIVE table's own mean — NOT scaled by
// RUN_TESTS_MAX_FILES_PER_CHUNK (2 here) the way the old ratio-of-budget
// rule was. Each heavy file is measured at 200000ms, well above the
// 120000ms bar, so it is pulled out by partitionIsolatedFiles into its
// own dedicated chunk, BEFORE packChunks ever sees it — an even stronger
// guarantee than LPT spread: the three heavy files can never land in the
// same chunk as each other or as a trivial file. That yields 5 chunks
// total: 3 isolated singles (the heavy files, one per chunk) plus the 3
// trivial files packed by count floor (ceil(3/2)=2 chunks: {2,1}), where
// 2 is RUN_TESTS_MAX_FILES_PER_CHUNK below (packing budget only — it no
// longer influences isolation). The assertion below therefore cannot
// pass on the old algorithm, nor with the timings file removed.
const heavy = ['heavy-0.test.cjs', 'heavy-1.test.cjs', 'heavy-2.test.cjs'];
const trivial = ['install-cheap-0.test.cjs', 'install-cheap-1.test.cjs', 'install-cheap-2.test.cjs'];
seed(tmpDir, [...heavy, ...trivial]);
const timingsFile = writeTimings(tmpDir, {
...Object.fromEntries(heavy.map((f) => [f, 200000])),
...Object.fromEntries(trivial.map((f) => [f, 10])),
});
const rh = runHarness(tmpDir, [], {
RUN_TESTS_MAX_CMDLINE_CHARS: '100000',
RUN_TESTS_MAX_FILES_PER_CHUNK: '2',
RUN_TESTS_TIMINGS_FILE: timingsFile,
});
assert.strictEqual(rh.status, 0, `heavy: expected zero exit; STDERR:\n${rh.stderr}`);
assert.match(rh.stderr, /run-tests: chunk 1\/5 — 1 files/, `expected isolated heavy chunk 1/5; STDERR:\n${rh.stderr}`);
assert.match(rh.stderr, /run-tests: chunk 2\/5 — 1 files/, `expected isolated heavy chunk 2/5; STDERR:\n${rh.stderr}`);
assert.match(rh.stderr, /run-tests: chunk 3\/5 — 1 files/, `expected isolated heavy chunk 3/5; STDERR:\n${rh.stderr}`);
assert.match(rh.stderr, /run-tests: chunk 4\/5 — 2 files/, `expected packed trivial chunk 4/5; STDERR:\n${rh.stderr}`);
assert.match(rh.stderr, /run-tests: chunk 5\/5 — 1 files/, `expected packed trivial chunk 5/5; STDERR:\n${rh.stderr}`);
});
test('trivial files the old heuristic over-weighted now stay in ONE chunk (#2088, #2456)', () => {
// The other direction of the miscalibration. These four files carry the
// `install-` prefix the old heuristic scored 12, so it split them into
// FOUR single-file chunks against a 6-weight budget. Measured, they cost
// 10ms each and belong together. The absence of a split marker cannot be
// produced by the old algorithm.
const trivial = Array.from({ length: 4 }, (_, i) => `install-triv-${i}.test.cjs`);
seed(tmpDir, trivial);
const timingsFile = writeTimings(tmpDir, Object.fromEntries(trivial.map((f) => [f, 10])));
const rl = runHarness(tmpDir, [], {
RUN_TESTS_MAX_CMDLINE_CHARS: '100000',
RUN_TESTS_MAX_FILES_PER_CHUNK: '6',
RUN_TESTS_TIMINGS_FILE: timingsFile,
});
assert.strictEqual(rl.status, 0, `light: expected zero exit; STDERR:\n${rl.stderr}`);
// A single chunk emits NO `chunk N/M` split marker (it only prints when
// chunks.length > 1), so its absence proves the four stayed together.
assert.doesNotMatch(rl.stderr, /run-tests: chunk \d+\/\d+ — /, `4 trivial install-prefixed files must stay in one chunk; STDERR:\n${rl.stderr}`);
});
});
describe('shard partitioning CLI (#1212)', () => {
// The windows full-test lane is sharded across N parallel runners via a
// GitHub Actions matrix shard dimension; each shard runs
// `run-tests.cjs --suite unit --shard i/n`. Selection is a deterministic
// round-robin over the SORTED file list so duration variance spreads
// across shards. The CLI surface is validated here; the pure partition
// contract (completeness/disjointness/balance/determinism) is validated
// against the exported selectShard() in the separate describe block below.
// 9 files, sorted: shard-00..shard-08. With n=3, round-robin gives
// shard 1 -> {00,03,06}, shard 2 -> {01,04,07}, shard 3 -> {02,05,08}.
const SHARD_NAMES = Array.from(
{ length: 9 },
(_, i) => `shard-${String(i).padStart(2, '0')}.test.cjs`,
);
test('--shard 1/3 runs a balanced round-robin slice of the sorted files', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '1/3']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
// The harness echoes the selected basenames on its `files=N:` line.
assert.match(r.stderr, /files=3:/);
assert.match(r.stderr, /shard-00\.test\.cjs/);
assert.match(r.stderr, /shard-03\.test\.cjs/);
assert.match(r.stderr, /shard-06\.test\.cjs/);
// Files belonging to other shards must NOT appear in this shard's run.
assert.doesNotMatch(r.stderr, /shard-01\.test\.cjs/);
assert.doesNotMatch(r.stderr, /shard-02\.test\.cjs/);
});
// #2472 E2E: every other --shard test here uses synthetic filenames that
// are absent from the real tests/test-timings.json, so they all collapse to
// one uniform median weight — under which LPT is mathematically identical
// to k % n. That means none of them can tell whether main() actually
// threads fileWeightOf() into selectShard: a typo on that single line would
// pass the entire existing suite. This test injects a table via
// RUN_TESTS_TIMINGS_FILE with DIFFERING costs, so the weighted partition is
// provably distinguishable from round-robin.
test('--shard routes by measured cost end-to-end, not by index', (t) => {
seed(tmpDir, SHARD_NAMES);
// Heavy files sit at indices 0/3/6 — exactly the slice round-robin hands
// to shard 1. Cost-weighting must NOT put all three on one shard.
const timings = {
schema_version: 1, unit: 'ms', timings: Object.fromEntries(
SHARD_NAMES.map((n, i) => [n, i % 3 === 0 ? 30000 : 100]),
),
};
const tablePath = path.join(tmpDir, 'injected-timings.json');
fs.writeFileSync(tablePath, JSON.stringify(timings));
t.after(() => { try { fs.unlinkSync(tablePath); } catch { /* best effort */ } });
const r = runHarness(tmpDir, ['--shard', '1/3'], { RUN_TESTS_TIMINGS_FILE: tablePath });
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
// Round-robin would give shard 1 exactly {00,03,06} — all three heavies.
const heavies = ['shard-00', 'shard-03', 'shard-06']
.filter(n => new RegExp(`${n}\\.test\\.cjs`).test(r.stderr)).length;
assert.ok(
heavies < 3,
`cost-weighted shard 1 must not receive all three heavy files (that is the `
+ `round-robin split, proving the weigher was not threaded); stderr: ${r.stderr}`,
);
// And the diagnostic must show the table was actually consumed.
assert.match(r.stderr, /table=loaded/, 'shard diagnostics must report the injected table as loaded');
assert.match(r.stderr, /sig=[0-9a-f]+/, 'shard diagnostics must emit an input fingerprint');
});
// #4070 E2E: main()'s own bounds check on RUN_TESTS_SHARD_RESERVE — an
// index that does not exist for the shard total in play — is invisible to
// every pure in-memory selectShard/parseShardReserve test, because that
// check lives in main() itself (scripts/run-tests.cjs, the
// `reserve.index <= parsed.shard.total` guard and its console.error
// fallback), which only runs through the CLI subprocess seam. Proves both
// halves: the warning fires, AND the selection is provably unaffected
// (byte-identical to a control run with no RUN_TESTS_SHARD_RESERVE at
// all, both against the SAME injected timings table so the comparison
// isn't muddied by table drift).
test('RUN_TESTS_SHARD_RESERVE with an out-of-range index warns and is ignored (#4070)', () => {
seed(tmpDir, SHARD_NAMES);
const timings = {
schema_version: 1, unit: 'ms', timings: Object.fromEntries(
SHARD_NAMES.map((n, i) => [n, i % 3 === 0 ? 30000 : 100]),
),
};
const tablePath = path.join(tmpDir, 'injected-timings-oob.json');
fs.writeFileSync(tablePath, JSON.stringify(timings));
try {
// --shard 1/3 → valid indices are 1..3. "5:999" is out of range.
const withBadReserve = runHarness(
tmpDir, ['--shard', '1/3'],
{ RUN_TESTS_TIMINGS_FILE: tablePath, RUN_TESTS_SHARD_RESERVE: '5:999' },
);
assert.strictEqual(withBadReserve.status, 0, `stderr: ${withBadReserve.stderr}`);
assert.match(
withBadReserve.stderr,
/RUN_TESTS_SHARD_RESERVE="5:999" is not a valid .* for --shard total 3 — ignoring/,
`expected the out-of-range-index fallback warning; got stderr: ${withBadReserve.stderr}`,
);
const control = runHarness(
tmpDir, ['--shard', '1/3'],
{ RUN_TESTS_TIMINGS_FILE: tablePath },
);
assert.strictEqual(control.status, 0, `stderr: ${control.stderr}`);
assert.doesNotMatch(control.stderr, /RUN_TESTS_SHARD_RESERVE/, 'control run must not warn — it sets no reserve at all');
const filesLine = (s) => (s.match(/files=\d+: (.*)$/m) || [])[1] || '';
assert.strictEqual(
filesLine(withBadReserve.stderr), filesLine(control.stderr),
'an out-of-range reserve index must select EXACTLY the same files as no reserve at all — '
+ 'the fallback warning alone is not proof the reserve was actually ignored',
);
} finally {
try { fs.unlinkSync(tablePath); } catch { /* best effort */ }
}
});
test('shard diagnostics report an identical input fingerprint across shards', () => {
seed(tmpDir, SHARD_NAMES);
// The cross-runner divergence guard: every shard of one run computes the
// partition independently, so all of them must agree on the INPUT. A
// differing sig between shard jobs is the only observable signal that
// they disagreed — which is how the union could silently drop a file.
const sigOf = (out) => (out.match(/sig=([0-9a-f]+)/) || [])[1];
const sigs = [1, 2, 3].map((i) => sigOf(runHarness(tmpDir, ['--shard', `${i}/3`]).stderr));
assert.ok(sigs.every(Boolean), `every shard must emit a sig; got ${JSON.stringify(sigs)}`);
assert.strictEqual(new Set(sigs).size, 1, `all shards must fingerprint the same input; got ${sigs}`);
});
test('--shard 2/3 selects the second round-robin slice', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '2/3']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.match(r.stderr, /files=3:/);
assert.match(r.stderr, /shard-01\.test\.cjs/);
assert.match(r.stderr, /shard-04\.test\.cjs/);
assert.match(r.stderr, /shard-07\.test\.cjs/);
assert.doesNotMatch(r.stderr, /shard-00\.test\.cjs/);
});
test('--shard composes with --suite (shards the post-filter selection)', () => {
// 6 unit files + 2 security files. `--suite unit --shard 1/2` must shard
// only the unit selection, never pulling in the security files.
seed(tmpDir, [
'u0.test.cjs',
'u1.test.cjs',
'u2.test.cjs',
'u3.test.cjs',
's0.security.test.cjs',
's1.security.test.cjs',
]);
const r = runHarness(tmpDir, ['--suite', 'unit', '--shard', '1/2']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.doesNotMatch(r.stderr, /\.security\.test\.cjs/);
});
test('--shard 1/1 is a pure no-op (runs every file)', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '1/1']);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
assert.match(r.stderr, /files=9:/);
});
test('an empty shard (n > file count) exits 0 without crashing', () => {
// n=5 with only 2 files: shards 3,4,5 are legitimately empty. An empty
// shard must NOT take the "discovery is broken" hard-error path.
seed(tmpDir, ['only-a.test.cjs', 'only-b.test.cjs']);
const r = runHarness(tmpDir, ['--shard', '5/5']);
assert.strictEqual(
r.status,
0,
`empty shard must exit 0; got status=${r.status} signal=${r.signal}\nSTDERR:\n${r.stderr}`,
);
});
test('an empty suite BEFORE sharding still hits the discovery hard error', () => {
// Regression (Codex #1212 review): a genuinely empty selection
// (e.g. --suite security with zero security files) must NOT be masked
// by the empty-shard escape hatch. Only a shard that emptied a NON-empty
// list is a legitimate no-op; a pre-empty selection is a broken filter.
seed(tmpDir, ['a.test.cjs', 'b.test.cjs']); // unit files only, no security
const r = runHarness(tmpDir, ['--suite', 'security', '--shard', '1/3']);
assert.notStrictEqual(
r.status,
0,
`empty-before-shard must fail; got status=${r.status}\nSTDERR:\n${r.stderr}`,
);
assert.match(r.stderr, /0 test files selected|discovery/i);
});
test('--shard over --files is order-independent (sorted before partition)', () => {
// Regression (Codex #1212 review): the partition keys off array index,
// so the same file set passed in different --files order must produce
// the same per-shard assignment. The runner sorts the selection before
// sharding to guarantee this.
seed(tmpDir, ['x0.test.cjs', 'x1.test.cjs', 'x2.test.cjs', 'x3.test.cjs']);
const forward = runHarness(tmpDir, [
'--files', 'x0.test.cjs x1.test.cjs x2.test.cjs x3.test.cjs',
'--shard', '1/2',
]);
const reversed = runHarness(tmpDir, [
'--files', 'x3.test.cjs x2.test.cjs x1.test.cjs x0.test.cjs',
'--shard', '1/2',
]);
assert.strictEqual(forward.status, 0, `stderr: ${forward.stderr}`);
assert.strictEqual(reversed.status, 0, `stderr: ${reversed.stderr}`);
// Both runs select the SAME files (sorted shard 1/2 of x0..x3 = x0,x2).
const filesLine = (s) => (s.match(/files=\d+: (.*)$/m) || [])[1] || '';
const a = filesLine(forward.stderr).split(' ').sort().join(' ');
const b = filesLine(reversed.stderr).split(' ').sort().join(' ');
assert.strictEqual(a, b, `order-dependent shard assignment:\nforward=${a}\nreversed=${b}`);
assert.match(a, /x0\.test\.cjs/);
assert.match(a, /x2\.test\.cjs/);
});
test('--shard rejects i outside 1..n', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '0/3']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /shard/i);
});
test('--shard rejects i greater than n', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '4/3']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /shard/i);
});
test('--shard rejects n < 1', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '1/0']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /shard/i);
});
test('--shard rejects malformed (no slash) value', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '2']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /shard/i);
});
test('--shard rejects non-integer parts', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '1.5/3']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /shard/i);
});
test('duplicate --shard flag is rejected', () => {
seed(tmpDir, SHARD_NAMES);
const r = runHarness(tmpDir, ['--shard', '1/3', '--shard', '2/3']);
assert.notStrictEqual(r.status, 0);
assert.match(r.stderr, /duplicate/i);
});
test('--shard chunking engages within a shard (argv-overflow preserved)', () => {
// Each shard must still chunk its own slice so a large shard cannot
// overflow the Windows 32,767-char command-line ceiling (#3597).
const longPrefix = 'a-deliberately-long-test-filename-to-force-chunking-within-a-shard-';
const names = Array.from(
{ length: 30 },
(_, i) => `${longPrefix}${String(i).padStart(4, '0')}.test.cjs`,
);
seed(tmpDir, names);
// n=2 -> each shard gets 15 files; a 1500-char ceiling forces chunking.
const r = runHarness(tmpDir, ['--shard', '1/2'], { RUN_TESTS_MAX_CMDLINE_CHARS: '1500' });
assert.strictEqual(r.status, 0, `stderr (tail):\n${r.stderr.split('\n').slice(-20).join('\n')}`);
assert.match(r.stderr, /run-tests: chunk \d+\/\d+ — \d+ files/);
});
});
describe('per-chunk timeout + force-exit (windows hang guard, #1051)', () => {
// A unit test that leaks an open handle (un-terminated Worker, un-killed
// child_process, ref'd timer) causes node --test to hang ~150s after its
// last test prints. Two such stalls push the windows full lane past its
// 20m CI cap and the job is CANCELLED — a false-negative gate. The harness
// now adds --test-force-exit (exits once all tests finish) and a per-chunk
// timeout (kills a hung child loudly instead of silently burning the budget).
// Leaky fixture: the test passes immediately, then a ref'd setInterval keeps
// the event loop alive so `node --test` hangs unless --test-force-exit is on.
const LEAKY_BODY = `const { test } = require('node:test');
test('passes but leaks a ref-d timer', () => {});
setInterval(() => {}, 1 << 30);
`;
test('a hung chunk hits the per-chunk timeout and fails with a clear message', () => {
// Regression proof: pre-fix (no timeout guard) this hung until the OS/CI
// killed it; now it fails fast with a diagnostic message.
fs.writeFileSync(path.join(tmpDir, 'leaky.test.cjs'), LEAKY_BODY, 'utf8');
const r = runHarness(tmpDir, [], {
RUN_TESTS_NO_FORCE_EXIT: '1',
RUN_TESTS_CHUNK_TIMEOUT_MS: '2000',
});
assert.notStrictEqual(
r.status,
0,
`expected non-zero exit from timed-out chunk; got status=${r.status}\nSTDERR:\n${r.stderr}`,
);
assert.match(
r.stderr,
/exceeded the per-chunk timeout/,
`expected timeout diagnostic in stderr; STDERR:\n${r.stderr}`,
);
});
test('a timed-out chunk aborts the remaining chunks instead of burning the job budget', () => {
// Guards against a real CI failure mode (observed live on CI run
// 29749380190, windows-latest full-lane shard 2/3): the per-chunk
// timeout is HALF the 20m CI job cap (600000ms default vs a 1,200,000ms
// job), so if the loop pressed on to the next chunk after a timeout
// instead of aborting, a healthy full pass (~11m42s on that lane) could
// never finish and the CI runner cancels the whole job. That
// cancellation surfaces as an opaque "The operation was canceled."
// and buries the actual timeout diagnostic thousands of lines back in
// the log (in the real incident, ~38,000 lines from the end — and
// `gh run view --log-failed` returned nothing). Proving chunk 2 never
// runs pins the fix: abort the remaining chunks as soon as one times
// out, so the loud diagnostic above survives as the visible failure.
//
// Force exactly one file per chunk so chunk 1 = the hanging file and
// chunk 2 = a marker-writing file; sorted order (a- before b-, and
// walkTestFiles(...).sort() sorts selection) fixes which chunk runs
// first, deterministically — the margin between a 2s timeout and a
// file that never terminates is unbounded, so there is no race.
fs.writeFileSync(path.join(tmpDir, 'a-leaky.test.cjs'), LEAKY_BODY, 'utf8');
const markerPath = path.join(tmpDir, 'chunk-2-ran.marker');
const secondBody = `'use strict';
const fs = require('fs');
const { test } = require('node:test');
fs.writeFileSync(${JSON.stringify(markerPath)}, 'ran');
test('noop', () => {});
`;
fs.writeFileSync(path.join(tmpDir, 'b-marker.test.cjs'), secondBody, 'utf8');
const r = runHarness(tmpDir, [], {
RUN_TESTS_NO_FORCE_EXIT: '1',
RUN_TESTS_CHUNK_TIMEOUT_MS: '2000',
RUN_TESTS_MAX_FILES_PER_CHUNK: '1',
});
assert.notStrictEqual(
r.status,
0,
`expected non-zero exit from an aborted run; got status=${r.status}\nSTDERR:\n${r.stderr}`,
);
assert.match(
r.stderr,
/exceeded the per-chunk timeout/,
`expected timeout diagnostic in stderr; STDERR:\n${r.stderr}`,
);
assert.match(
r.stderr,
/aborting — skipping the remaining 1 chunk/,
`expected the new abort message in stderr; STDERR:\n${r.stderr}`,
);
assert.strictEqual(
fs.existsSync(markerPath),
false,
`chunk 2's marker file must NOT exist — proves the second chunk's test file never executed; STDERR:\n${r.stderr}`,
);
assert.doesNotMatch(
r.stderr,
/chunk 2\/2/,
`chunk 2's progress line must never print — proves the loop broke before starting chunk 2; STDERR:\n${r.stderr}`,
);
});
test('force-exit lets a chunk with a leaked handle exit cleanly', () => {
const nodeMajor = Number(process.versions.node.split('.')[0]);
// --test-force-exit was added in Node 22; skip on older engines.
if (nodeMajor < 22) {
return; // skip — harness test options object not available here; just return
}
fs.writeFileSync(path.join(tmpDir, 'leaky.test.cjs'), LEAKY_BODY, 'utf8');
// force-exit is ON by default (RUN_TESTS_NO_FORCE_EXIT not set).
// 30s timeout: if force-exit works the child exits promptly after the test
// passes; if force-exit failed, the 30s timeout would fire and status ≠ 0.
const r = runHarness(tmpDir, [], {
RUN_TESTS_CHUNK_TIMEOUT_MS: '30000',
});
assert.strictEqual(
r.status,
0,
`expected zero exit with force-exit enabled; got status=${r.status} signal=${r.signal}\nSTDERR:\n${r.stderr}`,
);
});
});
describe('chunk-timeout instrumentation (#3889)', () => {
// Consolidation (CI cost, #4015): this describe block used to spawn the
// harness once PER assertion group (3 success-path runs + 2 timeout-path
// runs = 5 subprocess boots). Each boot is expensive — run-tests.cjs
// starts, globs the suite, then spawns `node --test` children — so on a
// CI shard already within ~39s of its 15-minute cap, paying for 5 boots
// to check facts that all hold against the SAME run is wasted spend.
// Below, exactly ONE successful run and ONE timed-out run are captured
// once (via `before`) and every assertion group below reads from those
// captured results instead of spawning its own. This is the whole
// savings — no assertion is weakened or removed.
// The fixture served to the timeout-path run below (#4105). Parks on a
// SETTLING timer — the #4104 idiom — so the hang is a property of the
// FIXTURE on every Node line, not of the runtime: the retired
// `new Promise(() => {})` shape held NO libuv handle, so whether the
// spawned child hung was decided by the Node line's test-runner shutdown
// behavior (measured: v24/v26 happen to hold the loop open; v22.22.0
// exits on its own in ~60ms with `# cancelled 1`, so the chunk never
// reached the timeout path and T1/T4 below asserted nothing). The park
// must outlast the chunk bound below with wide margin (asserted
// structurally in the #4105 regression test), stays ~0% CPU while parked,
// and the settle guarantees the fixture SELF-TERMINATES if a kill orphans
// it — unlike `setInterval`-forever (never exits) or `while (true) {}`
// (100% CPU forever), the two shapes #4104's pairing note rules out.
const HANG_PARK_MS = 10_000;
const HANGS_BODY = `'use strict';
const { test } = require('node:test');
test('hangs forever', () => new Promise((resolve) => { setTimeout(resolve, ${HANG_PARK_MS}); }));
`;
// The per-chunk timeout the timeout-path run below arms. Hoisted (#4105)
// so the #4105 fixture guard below asserts against the SAME bound the
// harness run uses, not a re-typed literal that can drift from it.
const CHUNK_TIMEOUT_MS_FOR_HANG_RUN = 2000;
let successDir;
let successRun;
let eventsDirBefore;
let eventsDirAfter;
let timeoutDir;
let timeoutRun;
before(() => {
// Single SUCCESS-path run, reused by T2/T3/T5 below.
successDir = createTempDir('gsd-3889-success-');
seed(successDir, ['a.test.cjs']);
eventsDirBefore = fs.readdirSync(require('os').tmpdir())
.filter((n) => n.startsWith('gsd-run-tests-events-'));
successRun = runHarness(successDir, []);
eventsDirAfter = fs.readdirSync(require('os').tmpdir())
.filter((n) => n.startsWith('gsd-run-tests-events-'));
// Single TIMEOUT-path run, reused by T1/T4 below. 2000ms is kept —
// it is already the smallest value this suite used anywhere for the
// per-chunk timeout, and going lower risks flaking on a loaded CI
// box that has to boot node --test, register the hang, and observe
// the kill inside the window.
timeoutDir = createTempDir('gsd-3889-timeout-');
fs.writeFileSync(path.join(timeoutDir, 'hangs.test.cjs'), HANGS_BODY, 'utf8');
timeoutRun = runHarness(timeoutDir, [], {
RUN_TESTS_NO_FORCE_EXIT: '1',
RUN_TESTS_CHUNK_TIMEOUT_MS: String(CHUNK_TIMEOUT_MS_FOR_HANG_RUN),
});
});
after(() => {
cleanup(successDir);
cleanup(timeoutDir);
});
// T2: per-chunk elapsed timing appears on the normal SUCCESS path, not
// only when something goes wrong — this is what makes "which chunk is
// drifting toward the cap" readable across ordinary green runs.
test('a successful chunk prints its elapsed time', () => {
assert.strictEqual(successRun.status, 0, `expected a clean pass; STDERR:\n${successRun.stderr}`);
assert.match(
successRun.stderr,
/run-tests: chunk 1\/1 completed in \d+ms/,
`expected a per-chunk completion timing line; STDERR:\n${successRun.stderr}`,
);
});
// T3: the ndjson companion reporter's destination file is a temp
// artifact of the instrumentation, not a product output — it must not
// survive a successful run. Assert against the OS temp root's own
// "gsd-run-tests-events-*" prefix (scripts/run-tests.cjs's mkdtemp
// prefix) rather than any run-tests-owned directory, since that IS the
// leak surface being guarded.
test('the ndjson reporter temp dir is cleaned up after a successful run', () => {
assert.strictEqual(successRun.status, 0, `expected a clean pass; STDERR:\n${successRun.stderr}`);
assert.deepStrictEqual(
eventsDirAfter,
eventsDirBefore,
`expected no leaked gsd-run-tests-events-* temp dir after a successful run; ` +
`before=${JSON.stringify(eventsDirBefore)} after=${JSON.stringify(eventsDirAfter)}`,
);
});
// T5 (regression): the human reporter's --test-reporter-destination
// pairing must be a regular file, not os.devNull. devNull is a character
// device; Node opens the reporter destination as an fs.WriteStream and
// fsyncs it on close, and fsync on a character device fails with EINVAL
// — surfaced as "Emitted 'error' event on WriteStream instance" /
// "EINVAL: invalid argument, fsync", which crashed EVERY chunk on the
// real remote run this regresses (43/43 failures), not only the timeout
// path. The argv construction lives entirely inside main() with no
// exported seam to unit-test directly (see NOTES), so this asserts the
// closest real, externally-observable consequence: a normal successful
// run must not surface that error text, and must still complete and
// exit 0 — both of which a reintroduced devNull destination would break
// on any platform where fsync(devNull) actually returns EINVAL (this
// suite's own bench platform, historically).
test('a successful run never surfaces the devNull fsync/EINVAL reporter crash', () => {
assert.strictEqual(successRun.status, 0, `expected a clean pass; STDERR:\n${successRun.stderr}`);
assert.doesNotMatch(
successRun.stderr,
/EINVAL|invalid argument, fsync|WriteStream instance/i,
`expected no reporter-destination fsync crash; STDERR:\n${successRun.stderr}`,
);
});
// T1: on a chunk timeout, the diagnostic must NAME the file that was
// still executing — not merely list every file the chunk contained (the
// pre-instrumentation behavior). A test that hangs INSIDE its own body
// (parks on a settling timer that outlasts the chunk bound, so it never
// resolves inside the window) keeps its test:start event unmatched by any
// test:pass/test:fail in the ndjson companion reporter's output, which
// is exactly the signal the diagnostic reads back on timeout.
test('a chunk timeout names the file that was in flight when killed', () => {
assert.notStrictEqual(
timeoutRun.status,
0,
`expected non-zero exit from a timed-out chunk; got status=${timeoutRun.status}\nSTDERR:\n${timeoutRun.stderr}`,
);
assert.match(
timeoutRun.stderr,
/In flight when killed.*hangs\.test\.cjs/s,
`expected the diagnostic to NAME the in-flight file, not just list the chunk; STDERR:\n${timeoutRun.stderr}`,
);
});
// Regression (#3889 recurrence): the reporter module itself, called
// directly with no subprocess, must return nully — this is the exact
// contract violation (`return []`) that crashed every chunk on the real
// remote run this file regresses ("Expected nully to be returned from
// the 'body' function but got an instance of Array", thrown by
// node:stream's `compose` when its async-function body returns an
// iterable instead of undefined/null). Also pins the NDJSON side effect:
// only the two handled event types are appended, verbatim, one per line.
test('the reporter returns nully and appends only the handled event types as NDJSON', async () => {
const reporter = require('../scripts/lib/ndjson-reporter.cjs');
const eventsFile = path.join(tmpDir, 'ndjson-reporter-events.ndjson');
const savedEventsFile = process.env.GSD_RUN_TESTS_EVENTS_FILE;
process.env.GSD_RUN_TESTS_EVENTS_FILE = eventsFile;
try {
async function* fakeEvents() {
yield { type: 'test:start', data: { file: 'a.test.cjs', name: 't', nesting: 0, testNumber: 1 } };
yield { type: 'test:diagnostic', data: { message: 'ignored' } };
yield { type: 'test:pass', data: { file: 'a.test.cjs', name: 't', nesting: 0, testNumber: 1 } };
}
const result = await reporter(fakeEvents());
assert.strictEqual(
result ?? null,
null,
`expected the reporter to return nully (undefined/null) per stream.compose's ` +
`async-function body contract; got ${JSON.stringify(result)}`,
);
const rawContent = fs.readFileSync(eventsFile, 'utf8');
const lines = splitLines(rawContent.trim()).filter((l) => l.length > 0);
assert.strictEqual(lines.length, 3, `expected exactly 3 NDJSON lines (init marker + 2 handled events); got:\n${lines.join('\n')}`);
const [init, start, pass] = lines.map((l) => JSON.parse(l));
assert.strictEqual(init.type, 'reporter:init');
assert.strictEqual(start.type, 'test:start');
assert.strictEqual(start.file, 'a.test.cjs');
assert.strictEqual(pass.type, 'test:pass');
assert.strictEqual(pass.file, 'a.test.cjs');
} finally {
if (savedEventsFile === undefined) {
delete process.env.GSD_RUN_TESTS_EVENTS_FILE;
} else {
process.env.GSD_RUN_TESTS_EVENTS_FILE = savedEventsFile;
}
cleanup(eventsFile);
}
});
// Regression (#3889 root cause): a hang inside a test body NEVER produces
// a `test:start`/`test:pass`/`test:fail` for that subtest (node:test only
// surfaces those to the parent once the child reports completion), so
// those three event types alone can never see a hang. `test:enqueue` and
// `test:dequeue` are emitted by the RUNNER as it queues/begins a file,
// independent of completion — this pins that the reporter now records
// both, verbatim, for exactly the "enqueue then dequeue, then nothing"
// shape a real hang produces.
test('the reporter records test:enqueue and test:dequeue for the hang shape (enqueue, dequeue, nothing else)', async () => {
const reporter = require('../scripts/lib/ndjson-reporter.cjs');
const eventsFile = path.join(tmpDir, 'ndjson-reporter-hang-shape.ndjson');
const savedEventsFile = process.env.GSD_RUN_TESTS_EVENTS_FILE;
process.env.GSD_RUN_TESTS_EVENTS_FILE = eventsFile;
try {
async function* hangShapeEvents() {
yield { type: 'test:enqueue', data: { file: 'hangs.test.cjs', name: 'hangs.test.cjs', nesting: 0 } };
yield { type: 'test:dequeue', data: { file: 'hangs.test.cjs', name: 'hangs.test.cjs', nesting: 0 } };
// Never yields test:start/test:pass/test:fail — this IS the hang.
}
const result = await reporter(hangShapeEvents());
assert.strictEqual(result ?? null, null);
const rawContent = fs.readFileSync(eventsFile, 'utf8');
const lines = splitLines(rawContent.trim()).filter((l) => l.length > 0);
assert.strictEqual(
lines.length,
3,
`expected exactly 3 NDJSON lines (init marker + enqueue + dequeue); got:\n${lines.join('\n')}`,
);
const [init, enqueue, dequeue] = lines.map((l) => JSON.parse(l));
assert.strictEqual(init.type, 'reporter:init');
assert.strictEqual(enqueue.type, 'test:enqueue');
assert.strictEqual(enqueue.file, 'hangs.test.cjs');
assert.strictEqual(dequeue.type, 'test:dequeue');
assert.strictEqual(dequeue.file, 'hangs.test.cjs');
} finally {
if (savedEventsFile === undefined) {
delete process.env.GSD_RUN_TESTS_EVENTS_FILE;
} else {
process.env.GSD_RUN_TESTS_EVENTS_FILE = savedEventsFile;
}
cleanup(eventsFile);
}
});
// #3889: the init marker is the reporter's FIRST action, written before
// the `for await` loop even begins — so it must land even when the
// source event stream yields ZERO events (e.g. the child is killed
// before node:test emits anything). This pins the marker's whole
// purpose: its presence alone proves the reporter module loaded and was
// invoked, independent of whether any test ever started.
test('the reporter writes only the init marker when the source yields zero events', async () => {
const reporter = require('../scripts/lib/ndjson-reporter.cjs');
const eventsFile = path.join(tmpDir, 'ndjson-reporter-init-only.ndjson');
const savedEventsFile = process.env.GSD_RUN_TESTS_EVENTS_FILE;
process.env.GSD_RUN_TESTS_EVENTS_FILE = eventsFile;
try {
async function* emptyEvents() {}
const result = await reporter(emptyEvents());
assert.strictEqual(
result ?? null,
null,
`expected the reporter to return nully even with zero source events; got ${JSON.stringify(result)}`,
);
const rawContent = fs.readFileSync(eventsFile, 'utf8');
const lines = splitLines(rawContent.trim()).filter((l) => l.length > 0);
assert.strictEqual(
lines.length,
1,
`expected exactly 1 NDJSON line (the init marker only); got:\n${lines.join('\n')}`,
);
const [init] = lines.map((l) => JSON.parse(l));
assert.strictEqual(init.type, 'reporter:init');
assert.strictEqual(typeof init.ts, 'number');
} finally {
if (savedEventsFile === undefined) {
delete process.env.GSD_RUN_TESTS_EVENTS_FILE;
} else {
process.env.GSD_RUN_TESTS_EVENTS_FILE = savedEventsFile;
}
cleanup(eventsFile);
}
});
// T4: the pre-existing timeout / abort / force-exit behavior (#1051)
// still holds with the reporter instrumentation wired in — the new
// --test-reporter flags must not change detection, the abort-on-timeout
// control flow, or the exit code.
test('existing timeout diagnostic and abort behavior are unchanged', () => {
assert.notStrictEqual(timeoutRun.status, 0, `expected non-zero exit; STDERR:\n${timeoutRun.stderr}`);
assert.match(
timeoutRun.stderr,
/exceeded the per-chunk timeout/,
`expected the original timeout diagnostic wording to survive; STDERR:\n${timeoutRun.stderr}`,
);
assert.match(
timeoutRun.stderr,
/run-tests: chunk 1\/1 was killed after \d+ms/,
`expected the new killed/elapsed line; STDERR:\n${timeoutRun.stderr}`,
);
});
// Regression (#4105): T1/T4 above are only meaningful if the fixture
// above GENUINELY hangs, and the hang must be a property of the FIXTURE,
// not of the runtime's test-runner shutdown behavior. A never-settling
// `new Promise(() => {})` holds NO libuv handle, so whether the spawned
// child hangs is decided by the Node line: on v24/v26 the runner happens
// to hold the loop open, but on v22 the child exits on its own in ~60ms
// (`# cancelled 1`, rc=1) — the chunk then never reaches the timeout path
// and T1/T4 assert nothing about the timeout diagnostic (measured:
// v22.22.0 `node --test` exits after 61ms). RED at this test's introducing
// sha against the then-current body: on Node 24 it fails the
// self-termination arm (an unheld never-settling promise also never
// self-terminates on lines where the runner DOES hold it open), and on
// off-24 lines it fails the still-hanging arm. Deterministic guard, the
// #4104 self-exit regression's event-driven shape: spawn the EXACT served
// body directly — no harness, no runner, no polling — and observe that its
// natural 'exit' event (a) arrives no earlier than past the chunk bound
// the harness run above arms (it hangs by itself, without any runtime
// holding it up), and (b) carries a natural exit — exit code 0, no
// signal. The `{ timeout: 2 * HANG_PARK_MS }` backstop (the
// health-validation #663 pattern) fails an immortal body — one that never
// delivers 'exit' — instead of hanging the suite; a `{ timeout }` option
// is the no-elapsed-assertion-compliant bound. Liveness of a SUBPROCESS
// cannot be driven through the clock seam (mock.timers cannot reach
// inside a separately spawned node), so the guard asserts observable
// exit/signal/liveness behavior only — never a measured duration value
// (RULESET.TESTS.no-timing-assertion).
test('the #3889 hang fixture genuinely hangs by itself and self-terminates (#4105)', { timeout: 2 * HANG_PARK_MS }, (t, done) => {
// (a) Structural margin, single-sourced with the served body: the park
// must outlast the chunk bound by >= 4x, so a loaded box's scheduling
// jitter can never let the timer settle before the harness kill fires.
assert.ok(
HANG_PARK_MS >= 4 * CHUNK_TIMEOUT_MS_FOR_HANG_RUN,
`fixture park (${HANG_PARK_MS}ms) must exceed the chunk bound (${CHUNK_TIMEOUT_MS_FOR_HANG_RUN}ms) with margin`,
);
const dir = createTempDir('gsd-3889-hang-fixture-');
t.after(() => cleanup(dir));
const tf = path.join(dir, 'hangs.test.cjs');
fs.writeFileSync(tf, HANGS_BODY, 'utf8');
const child = spawn(process.execPath, [tf], { stdio: 'ignore' });
// The runner timeout above fails an immortal body; this hook guarantees
// the child is reaped on every exit path, including that one.
t.after(() => { try { child.kill('SIGKILL'); } catch { /* already gone */ } });
const spawnedAt = Date.now();
// Past the chunk bound the harness kills at: the fixture's natural exit
// must arrive no earlier than this — that is the vacuity guard. +400ms
// covers child spawn startup so the checkpoint measures the fixture,
// not boot time.
const stillHangingAt = spawnedAt + CHUNK_TIMEOUT_MS_FOR_HANG_RUN + 400;
// Fast-fail on a spawn error (execPath unspawnable): 'exit' never fires
// for a failed spawn, so without this the runner timeout would expire
// with a generic timeout message instead of the real cause (same
// hardening as #4104).
child.on('error', (err) => {
assert.fail(`could not spawn the fixture directly: ${err.message}`);
});
// `exitCode !== null` alone misses a SIGNALED exit (exitCode stays null
// when a signal ends the child); the assertions name signalCode as the
// failure when that happens.
child.on('exit', () => {
assert.ok(
Date.now() >= stillHangingAt,
`fixture must still be hanging past the ${CHUNK_TIMEOUT_MS_FOR_HANG_RUN}ms chunk bound — ` +
`it exited after only ${Date.now() - spawnedAt}ms ` +
'(#4105 regression: the hang is the runtime\'s, not the fixture\'s)',
);
assert.strictEqual(child.signalCode, null,
'fixture must SELF-terminate (natural exit) — a signal means we had to kill it (#4105/#4104 regression)');
assert.strictEqual(child.exitCode, 0, 'the parked fixture settles and its test passes cleanly');
done();
});
});
});
});
// Pure partition contract for the shard selector (#1212). Imported directly
// (no subprocess) because these are deterministic in-memory assertions about
// the round-robin partition — the cheapest, most precise way to pin
// completeness / disjointness / balance / determinism, including a fast-check
// property test (RULESET.TESTS.property-based-testing: partition is a
// bijective transformation contract).
const { parseShardArg, selectShard } = require('../scripts/run-tests.cjs');
describe('selectShard round-robin partition (#1212)', () => {
// A deterministic sorted file list; selectShard MUST NOT re-sort — the caller
// sorts once and the partition keys off array index so ordering is identical
// across Windows/macOS/Linux.
const files = Array.from({ length: 25 }, (_, i) => `f${String(i).padStart(3, '0')}.test.cjs`);
test('n=1 returns the full list unchanged (pure no-op)', () => {
assert.deepStrictEqual(selectShard(files, { index: 1, total: 1 }), files);
});
test('completeness: the union of all shards equals the full list', () => {
const n = 4;
const union = [];
for (let i = 1; i <= n; i++) union.push(...selectShard(files, { index: i, total: n }));
assert.deepStrictEqual([...union].sort(), [...files].sort());
});
test('disjointness: no file appears in two shards', () => {
const n = 4;
const seen = new Set();
for (let i = 1; i <= n; i++) {
for (const f of selectShard(files, { index: i, total: n })) {
assert.ok(!seen.has(f), `file ${f} appeared in more than one shard`);
seen.add(f);
}
}
assert.strictEqual(seen.size, files.length);
});
test('balance: shard sizes differ by at most 1', () => {
const n = 4;
const sizes = [];
for (let i = 1; i <= n; i++) sizes.push(selectShard(files, { index: i, total: n }).length);
assert.ok(Math.max(...sizes) - Math.min(...sizes) <= 1, `sizes=${sizes}`);
});
test('determinism: same input yields the same partition', () => {
const a = selectShard(files, { index: 2, total: 3 });
const b = selectShard(files, { index: 2, total: 3 });
assert.deepStrictEqual(a, b);
});
test('round-robin: shard i gets indices i-1, i-1+n, i-1+2n, …', () => {
const n = 3;
assert.deepStrictEqual(
selectShard(files, { index: 1, total: n }),
files.filter((_, k) => k % n === 0),
);
assert.deepStrictEqual(
selectShard(files, { index: 2, total: n }),
files.filter((_, k) => k % n === 1),
);
assert.deepStrictEqual(
selectShard(files, { index: 3, total: n }),
files.filter((_, k) => k % n === 2),
);
});
test('preserves relative order within a shard', () => {
const slice = selectShard(files, { index: 1, total: 3 });
const sorted = [...slice].sort();
assert.deepStrictEqual(slice, sorted);
});
test('empty shard when total > file count returns []', () => {
const two = ['a.test.cjs', 'b.test.cjs'];
assert.deepStrictEqual(selectShard(two, { index: 5, total: 5 }), []);
assert.deepStrictEqual(selectShard(two, { index: 3, total: 5 }), []);
// shards 1 and 2 still get the two files
assert.deepStrictEqual(selectShard(two, { index: 1, total: 5 }), ['a.test.cjs']);
assert.deepStrictEqual(selectShard(two, { index: 2, total: 5 }), ['b.test.cjs']);
});
test('boundary: total exactly equals file count → one file per shard', () => {
const three = ['a.test.cjs', 'b.test.cjs', 'c.test.cjs'];
for (let i = 1; i <= 3; i++) {
assert.strictEqual(selectShard(three, { index: i, total: 3 }).length, 1);
}
});
test('boundary: total = count-1 and count+1', () => {
const four = ['a.test.cjs', 'b.test.cjs', 'c.test.cjs', 'd.test.cjs'];
// count-1 = 3 shards over 4 files → sizes {2,1,1}
const sizes3 = [1, 2, 3].map(i => selectShard(four, { index: i, total: 3 }).length).sort();
assert.deepStrictEqual(sizes3, [1, 1, 2]);
// count+1 = 5 shards over 4 files → one shard empty
const sizes5 = [1, 2, 3, 4, 5].map(i => selectShard(four, { index: i, total: 5 }).length).sort();
assert.deepStrictEqual(sizes5, [0, 1, 1, 1, 1]);
});
test('property: partition is complete, disjoint, and balanced for any n,N', () => {
const fc = require('fast-check');
fc.assert(
fc.property(
fc.array(fc.string(), { minLength: 0, maxLength: 50 }),
fc.integer({ min: 1, max: 12 }),
(rawFiles, n) => {
// Caller contract: deduped + sorted list. Mirror it so the property
// exercises the same shape the runner feeds selectShard.
const list = [...new Set(rawFiles)].sort();
const shards = [];
for (let i = 1; i <= n; i++) shards.push(selectShard(list, { index: i, total: n }));
// completeness + disjointness
const flat = shards.flat();
assert.deepStrictEqual([...flat].sort(), [...list].sort());
assert.strictEqual(new Set(flat).size, list.length);
// balance — shard sizes differ by at most 1 (sizes is never empty
// because n >= 1, so there is always at least one shard).
const sizes = shards.map(s => s.length);
assert.ok(Math.max(...sizes) - Math.min(...sizes) <= 1, `sizes=${sizes}`);
},
),
// Seed pinned so a failure is reproducible (repo convention for property
// tests); previously unseeded, flagged by the #2472 isolated review.
{ numRuns: 200, seed: 12120 },
);
});
});
// ─── #2472: weight-aware shard partition ────────────────────────────────────
//
// Equal file COUNTS is not equal file COST. Round-robin by array index balances
// the former and ignores the latter, which on the real suite produced Windows
// shard weights of 12.4m / 19.2m / 15.2m against a 20-minute job cap — a 1.23x
// max/ideal ratio and a 5%-of-cap margin on the heaviest shard. Adding any test
// file re-indexed the partition and tipped it over (#2472).
//
// Passing a weigher switches the partition to LPT (longest-processing-time
// first), the same algorithm packChunks already uses one level down (#2456 /
// #2463), so both layers share one cost model. Omitting the weigher keeps the
// legacy round-robin exactly — every test above still exercises that path.
describe('selectShard weight-aware partition (#2472)', () => {
const fc = require('fast-check');
const shardWeights = (files, total, weightOf) => {
const out = [];
for (let i = 1; i <= total; i++) {
out.push(selectShard(files, { index: i, total }, weightOf).reduce((a, f) => a + weightOf(f), 0));
}
return out;
};
// The regression: a right-skewed cost distribution (the real suite's shape —
// a few dominant files among many cheap ones) laid out so that filename order
// clusters the heavy files onto one shard.
const skewed = {
'a01.test.cjs': 30000, 'a02.test.cjs': 100, 'a03.test.cjs': 100,
'a04.test.cjs': 28000, 'a05.test.cjs': 100, 'a06.test.cjs': 100,
'a07.test.cjs': 26000, 'a08.test.cjs': 100, 'a09.test.cjs': 100,
'a10.test.cjs': 24000, 'a11.test.cjs': 100, 'a12.test.cjs': 100,
};
const skewedFiles = Object.keys(skewed).sort();
const skewedWeight = (f) => skewed[f];
test('REGRESSION: round-robin clusters heavy files; LPT does not', () => {
const ideal = Object.values(skewed).reduce((a, b) => a + b, 0) / 3;
// Legacy partition (no weigher) scored with the same cost function, so the
// two strategies are compared on identical inputs.
const rr = [1, 2, 3].map((i) =>
selectShard(skewedFiles, { index: i, total: 3 }).reduce((a, f) => a + skewedWeight(f), 0));
const lpt = shardWeights(skewedFiles, 3, skewedWeight);
// Every heavy file sits at an index ≡ 0 (mod 3), so round-robin hands them
// all to shard 1 — the exact clustering that produced the 19-minute shard.
assert.ok(
Math.max(...rr) / ideal > 1.5,
`round-robin should be badly imbalanced on skewed costs; got ${rr} (ideal ${ideal})`,
);
// Graham's list-scheduling bound: makespan <= average + heaviest item. This
// is the bound that is actually provable. The 4/3 figure quoted for LPT is
// relative to the OPTIMAL makespan, not to the average — and the two differ
// whenever item sizes force a pairing, as they do here: four items above
// 24000 into three bins means some bin holds two of them, so 50000 is
// optimal even though the average is 36267.
const heaviest = Math.max(...Object.values(skewed));
assert.ok(
Math.max(...lpt) <= ideal + heaviest,
`LPT must hold the average+max bound; got ${lpt} (bound ${ideal + heaviest})`,
);
assert.ok(
Math.max(...lpt) < Math.max(...rr),
`LPT must beat round-robin on skewed costs; lpt=${lpt} rr=${rr}`,
);
});
test('back-compat: omitting the weigher reproduces round-robin exactly', () => {
for (let i = 1; i <= 3; i++) {
assert.deepStrictEqual(
selectShard(skewedFiles, { index: i, total: 3 }),
skewedFiles.filter((_, k) => k % 3 === i - 1),
'the unweighted path must stay byte-identical to the legacy partition',
);
}
});
test('a uniform weigher degenerates to equal counts', () => {
const sizes = [1, 2, 3].map(
(i) => selectShard(skewedFiles, { index: i, total: 3 }, () => 1).length,
);
assert.ok(
Math.max(...sizes) - Math.min(...sizes) <= 1,
`equal weights must give equal counts; got ${sizes}`,
);
});
test('n=1 is still a pure no-op with a weigher', () => {
assert.deepStrictEqual(selectShard(skewedFiles, { index: 1, total: 1 }, skewedWeight), skewedFiles);
});
test('preserves the caller sort order within a weighted shard', () => {
const slice = selectShard(skewedFiles, { index: 1, total: 3 }, skewedWeight);
assert.deepStrictEqual(slice, [...slice].sort(), 'downstream chunking assumes caller order');
});
test('determinism: the weighted partition is stable across calls', () => {
const a = selectShard(skewedFiles, { index: 2, total: 3 }, skewedWeight);
const b = selectShard(skewedFiles, { index: 2, total: 3 }, skewedWeight);
assert.deepStrictEqual(a, b);
});
test('ties break deterministically, not by object iteration order', () => {
const flat = () => 5;
const a = selectShard(skewedFiles, { index: 1, total: 3 }, flat);
const b = selectShard([...skewedFiles], { index: 1, total: 3 }, flat);
assert.deepStrictEqual(a, b, 'equal weights must still partition identically');
});
// A partition is a covering contract: every file lands in exactly one shard.
// Getting this wrong silently DROPS tests from CI — the worst possible failure
// mode for a test harness — so it is property-tested rather than sampled.
test('property: the weighted partition is exhaustive and disjoint', () => {
fc.assert(
fc.property(
fc.array(fc.integer({ min: 0, max: 60000 }), { minLength: 1, maxLength: 60 }),
fc.integer({ min: 1, max: 8 }),
(weights, total) => {
const files = weights.map((_, i) => `p${String(i).padStart(3, '0')}.test.cjs`);
const w = (f) => weights[Number(f.slice(1, 4))];
const seen = [];
for (let i = 1; i <= total; i++) seen.push(...selectShard(files, { index: i, total }, w));
assert.strictEqual(new Set(seen).size, seen.length, 'a file appeared in two shards');
assert.deepStrictEqual([...seen].sort(), [...files].sort(), 'a file was dropped or invented');
},
),
{ numRuns: 200, seed: 24720 },
);
});
// Graham's bound for any greedy-into-lightest schedule, and the reason the
// shard cap stops being reachable: the heaviest shard cannot exceed the
// average by more than one file's cost, however the names happen to sort.
test('property: no shard exceeds average + heaviest file', () => {
fc.assert(
fc.property(
fc.array(fc.integer({ min: 1, max: 60000 }), { minLength: 3, maxLength: 60 }),
fc.integer({ min: 2, max: 6 }),
(weights, total) => {
const files = weights.map((_, i) => `p${String(i).padStart(3, '0')}.test.cjs`);
const w = (f) => weights[Number(f.slice(1, 4))];
const sums = shardWeights(files, total, w);
const bound = weights.reduce((a, b) => a + b, 0) / total + Math.max(...weights);
assert.ok(
Math.max(...sums) <= bound + 1e-9,
`bound violated: max=${Math.max(...sums)} bound=${bound}`,
);
},
),
{ numRuns: 200, seed: 24721 },
);
});
// Deliberately NOT asserted: "weighted is never worse than round-robin".
// fast-check falsifies it — weights [19316,10190,1,9128,29353,20227] over 2
// shards give round-robin 48670 and LPT 48671. Round-robin can win by luck on
// a specific input; LPT's guarantee is the worst-case bound above, not
// universal dominance. The value for #2472 is that the bound holds for EVERY
// distribution, so no arrangement of filenames can produce the 1.23x cluster
// that round-robin allowed — which the skewed regression above pins directly.
// Zero-weight regression (isolated review, HIGH). Adding a zero-weight file
// leaves its bin's weight unchanged, so a weight-only tiebreak kept bin 0
// tied-minimum forever and every such file landed there: all-zero weights put
// the entire suite on shard 1 and left the other runners idle. Reachable via
// safeWeight's clamp of a NaN/negative/Infinity entry, or any genuine 0ms
// measurement. The file-count tiebreak is what makes ties rotate.
test('REGRESSION: zero weights still split evenly across shards', () => {
const files = Array.from({ length: 9 }, (_, i) => `z${i}.test.cjs`);
const sizes = [1, 2, 3].map((i) => selectShard(files, { index: i, total: 3 }, () => 0).length);
assert.deepStrictEqual(sizes, [3, 3, 3], `all-zero weights must not collapse onto one shard; got ${sizes}`);
});
test('REGRESSION: clamped NaN/negative/Infinity weights do not collapse', () => {
const files = ['a', 'b', 'c', 'd', 'e', 'f'].map((x) => `${x}.test.cjs`);
const hostile = { 'a.test.cjs': NaN, 'b.test.cjs': -5, 'c.test.cjs': Infinity };
const sizes = [1, 2, 3].map(
(i) => selectShard(files, { index: i, total: 3 }, (f) => (f in hostile ? hostile[f] : 1)).length,
);
assert.ok(
Math.max(...sizes) - Math.min(...sizes) <= 1,
`weights that clamp to 0 must still spread; got ${sizes}`,
);
});
test('property: zero weights spread evenly for any list size and shard count', () => {
fc.assert(
fc.property(
fc.integer({ min: 1, max: 40 }),
fc.integer({ min: 2, max: 6 }),
(n, total) => {
const files = Array.from({ length: n }, (_, i) => `p${String(i).padStart(3, '0')}.test.cjs`);
const sizes = Array.from({ length: total }, (_, i) =>
selectShard(files, { index: i + 1, total }, () => 0).length);
assert.ok(Math.max(...sizes) - Math.min(...sizes) <= 1, `sizes=${sizes}`);
},
),
{ numRuns: 200, seed: 24723 },
);
});
});
// ─── #4070: reserved-weight shard partition ─────────────────────────────────
//
// `test.yml`'s `scope: full` lane tacks four unsharded aux suites
// (integration/security/install/slow) onto shard 1 only, outside this
// packer's model entirely — it balances the UNIT-TEST slice as if all three
// shards carried equal fixed cost, when shard 1 actually carries a fixed
// aux-suite overhead the other two do not. `initialWeights` lets a caller
// give one (or more) bins a virtual head start before LPT places any real
// file, so the algorithm converges on equalizing FINAL total cost (reserve +
// assigned files) instead of raw assigned-file weight alone — the same
// "greedy into the lightest bin" placement rule, just with non-zero starting
// points.
describe('selectShard reserved-weight partition (#4070)', () => {
const fc = require('fast-check');
const uniform = Array.from({ length: 30 }, (_, i) => `u${String(i).padStart(3, '0')}.test.cjs`);
const uniformWeight = () => 10;
const finalTotals = (files, total, weightOf, initialWeights) => {
const out = [];
for (let i = 1; i <= total; i++) {
const assigned = selectShard(files, { index: i, total }, weightOf, initialWeights)
.reduce((a, f) => a + weightOf(f), 0);
out.push(assigned + ((initialWeights && initialWeights[i - 1]) || 0));
}
return out;
};
// The regression: without reserve support, bin 0 gets an EQUAL share of
// files despite already carrying a head start, so its true final total
// (assigned + reserve) sits well above the other bins' — exactly the shard
// 1 overload this issue reports. With reserve support, LPT starts bin 0
// "already heavier" and hands it fewer files so all three converge.
test('REGRESSION: an initial reserve on one bin rebalances the rest (#4070)', () => {
const reserve = 80; // 8 average-cost files' worth, on a 30-file/300-weight suite
const totals = finalTotals(uniform, 3, uniformWeight, [reserve, 0, 0]);
const spread = Math.max(...totals) - Math.min(...totals);
assert.ok(
spread <= uniformWeight(),
`a reserved bin must converge toward the others' final totals (within one file's `
+ `weight), not just add the reserve on top of an equal share; got totals=${totals} `
+ `(spread=${spread})`,
);
// The reserved bin must have been handed FEWER files than an unreserved bin —
// otherwise "rebalancing" did nothing and the reserve is purely additive.
const reservedBinFiles = selectShard(uniform, { index: 1, total: 3 }, uniformWeight, [reserve, 0, 0]).length;
const unreservedBinFiles = selectShard(uniform, { index: 2, total: 3 }, uniformWeight, [reserve, 0, 0]).length;
assert.ok(
reservedBinFiles < unreservedBinFiles,
`the reserved bin (${reservedBinFiles} files) must receive fewer files than an `
+ `unreserved bin (${unreservedBinFiles}) — otherwise the reserve had no effect on `
+ 'placement',
);
});
test('back-compat: omitting initialWeights reproduces the unreserved partition exactly', () => {
for (let i = 1; i <= 3; i++) {
assert.deepStrictEqual(
selectShard(uniform, { index: i, total: 3 }, uniformWeight),
selectShard(uniform, { index: i, total: 3 }, uniformWeight, undefined),
);
}
});
test('a zero reserve is a no-op', () => {
for (let i = 1; i <= 3; i++) {
assert.deepStrictEqual(
selectShard(uniform, { index: i, total: 3 }, uniformWeight),
selectShard(uniform, { index: i, total: 3 }, uniformWeight, [0, 0, 0]),
);
}
});
test('a reserve applies to any bin index, not only the first', () => {
const reserve = 80;
const totals = finalTotals(uniform, 3, uniformWeight, [0, reserve, 0]);
const spread = Math.max(...totals) - Math.min(...totals);
assert.ok(spread <= uniformWeight(), `expected convergence around index 2; got ${totals}`);
const reservedBinFiles = selectShard(uniform, { index: 2, total: 3 }, uniformWeight, [0, reserve, 0]).length;
const otherBinFiles = selectShard(uniform, { index: 1, total: 3 }, uniformWeight, [0, reserve, 0]).length;
assert.ok(reservedBinFiles < otherBinFiles, `reserved bin 2 should get fewer files; got ${reservedBinFiles} vs ${otherBinFiles}`);
});
test('REGRESSION: a reserve larger than the whole suite still terminates and assigns every file', () => {
const reserve = 1e9;
const shards = [];
for (let i = 1; i <= 3; i++) shards.push(selectShard(uniform, { index: i, total: 3 }, uniformWeight, [reserve, 0, 0]));
const flat = shards.flat();
assert.deepStrictEqual([...flat].sort(), [...uniform].sort(), 'every file must still be placed exactly once');
// The massively-reserved bin should get the fewest (possibly zero) files.
assert.ok(shards[0].length <= shards[1].length && shards[0].length <= shards[2].length);
});
test('REGRESSION: a hostile reserve value clamps to zero instead of poisoning placement', () => {
for (const hostile of [NaN, -5, Infinity]) {
const shards = [];
for (let i = 1; i <= 3; i++) shards.push(selectShard(uniform, { index: i, total: 3 }, uniformWeight, [hostile, 0, 0]));
const sizes = shards.map((s) => s.length);
assert.ok(
Math.max(...sizes) - Math.min(...sizes) <= 1,
`a hostile reserve (${hostile}) must clamp to 0, not collapse/starve a bin; sizes=${sizes}`,
);
}
});
// Generalizes the existing "no shard exceeds average + heaviest file" bound
// (#2472) to include a single reserved bin — but the Graham-style proof
// (the max-load bin was the argmin, hence <= average, at the moment its
// LAST item was placed) only applies to a bin that actually received at
// least one item. A reserve large enough that its bin never receives any
// real item stays at EXACTLY its initial reserve forever — no amount of
// routing real items elsewhere can dilute a fixed head start below itself
// — so the true bound is the LARGER of the classic Graham term and the
// single biggest reserve. (Counterexample that falsified the original,
// reserve-blind-to-domination version of this bound: weights=[1,1,1],
// total=2, reserve=6 on bin 0 — bin 0 receives zero items and stays at 6,
// while (sum+reserve)/total+max = 4.5+1 = 5.5 < 6.)
test('property: no shard exceeds max(reserve, average(+reserve) + heaviest file)', () => {
fc.assert(
fc.property(
fc.array(fc.integer({ min: 1, max: 60000 }), { minLength: 3, maxLength: 60 }),
fc.integer({ min: 2, max: 6 }),
fc.integer({ min: 0, max: 200000 }),
fc.integer({ min: 0, max: 5 }),
(weights, total, reserve, reserveIdxRaw) => {
const files = weights.map((_, i) => `p${String(i).padStart(3, '0')}.test.cjs`);
const w = (f) => weights[Number(f.slice(1, 4))];
const reserveIdx = reserveIdxRaw % total;
const initialWeights = Array.from({ length: total }, (_, i) => (i === reserveIdx ? reserve : 0));
const sums = finalTotals(files, total, w, initialWeights);
const grahamBound = (weights.reduce((a, b) => a + b, 0) + reserve) / total + Math.max(...weights);
const bound = Math.max(reserve, grahamBound);
assert.ok(
Math.max(...sums) <= bound + 1e-9,
`bound violated: max=${Math.max(...sums)} bound=${bound} sums=${sums}`,
);
},
),
{ numRuns: 200, seed: 24724 },
);
});
});
describe('parseShardArg (#1212)', () => {
test('parses i/n into { index, total }', () => {
assert.deepStrictEqual(parseShardArg('2/3'), { index: 2, total: 3 });
assert.deepStrictEqual(parseShardArg('1/1'), { index: 1, total: 1 });
});
const bad = ['', '2', '0/3', '4/3', '1/0', '-1/3', '1.5/3', 'a/b', '1/3/2', ' 1/3', '1 / 3'];
for (const v of bad) {
test(`rejects malformed/out-of-range value ${JSON.stringify(v)}`, () => {
const r = parseShardArg(v);
assert.ok(r && r.error, `expected an error result for ${JSON.stringify(v)}, got ${JSON.stringify(r)}`);
});
}
});
// #4070: RUN_TESTS_SHARD_RESERVE env-var grammar — "<index>:<weight>", the
// operator knob test.yml uses to tell the full-scope lane's unit-test shard
// selection that shard 1 already carries a fixed aux-suite cost. Fail-open
// on anything malformed (mirrors positiveNumberEnv's precedent elsewhere in
// this file): a typo must degrade to "no reserve", never poison placement or
// throw and take the whole CI job down with it.
describe('parseShardReserve (#4070)', () => {
const { parseShardReserve } = require('../scripts/run-tests.cjs');
test('parses "<index>:<weight>" into { index, weight }', () => {
assert.deepEqual(parseShardReserve('1:77'), { index: 1, weight: 77 });
assert.deepEqual(parseShardReserve('2:0'), { index: 2, weight: 0 });
assert.deepEqual(parseShardReserve('3:12.5'), { index: 3, weight: 12.5 });
});
for (const v of [undefined, null, '', ' ', 'x', '1', '1:', ':77', '0:77', '-1:77', '1:-5', '1.5:77', '1:abc', 'a:b', '1:2:3']) {
test(`rejects malformed value ${JSON.stringify(v)}`, () => {
assert.equal(parseShardReserve(v), null, `expected null for ${JSON.stringify(v)}`);
});
}
test('whitespace around a valid value is tolerated', () => {
assert.deepEqual(parseShardReserve(' 1:77 '), { index: 1, weight: 77 });
});
});
// ────────────────────────────────────────────────────────────────────────
// Folded from tests/bug-969-test-infra-flake-hardening.test.cjs — consolidation epic #1969 (B6 #1975)
// ────────────────────────────────────────────────────────────────────────
{
const { describe: __foldDescribe } = require('node:test');
__foldDescribe("folded:bug-969-test-infra-flake-hardening (consolidation epic #1969 B6 #1975)", () => {
'use strict';
/**
* Regression tests for bug #969 — test-infra flake hardening.
*
* Two root causes addressed:
*
* A. SIGNATURE A: "X is not a function"
* ensureBuiltArtifacts() previously short-circuited on a single sentinel
* (semver-compare.cjs). If any other migrated .cjs was stale or absent,
* it would be silently loaded in that broken state. This test proves the
* unconditional-build fix: deleting a non-sentinel artifact and invoking
* ensureBuiltArtifacts() regenerates it even when the sentinel is present.
*
* B. SIGNATURE B: misleading assertion failures from killed subprocesses
* runGsdTools() previously had no timeout, so an OOM/SIGKILL'd subprocess
* returned { success: false } and looked like a product error. This test
* proves the kill-discrimination fix: a killed/timed-out invocation now
* throws a labeled resource-starvation error, while a clean non-zero exit
* still returns { success: false, exitCode: N }.
*
* C. SIGNATURE C: "Failed to install hooks: directory is empty" in scoped CI
* hooks/dist is gitignored and NOT built by `prepare` (build:lib only), so
* the scoped test lane starts with it absent. The first install test's
* before() hook triggers build-hooks.js, which creates DIST_DIR empty then
* fills it file-by-file — a window where a concurrently-spawned install
* reader sees zero hooks and hard-fails. ensureBuiltHooks() builds hooks/dist
* ONCE upfront (same chokepoint as ensureBuiltArtifacts) so the empty window
* never exists during concurrent test execution. These tests prove it
* rebuilds when dist is absent/empty/incomplete and no-ops when complete.
*
* RULESET.TESTS.regression-must-fail-first: each test section documents what
* the old behavior would have been (fail-before) and asserts the new behavior
* (pass-after), using only behavioral invocations — no source-grep.
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const path = require('node:path');
const fs = require('node:fs');
const os = require('node:os');
const { execFileSync } = require('node:child_process');
const { ensureBuiltArtifacts, ensureBuiltHooks } = require('../scripts/run-tests.cjs');
const { cleanup } = require('./helpers.cjs');
// ---------------------------------------------------------------------------
// Part A — ensureBuiltArtifacts: unconditional rebuild
// ---------------------------------------------------------------------------
describe('bug #969 A — ensureBuiltArtifacts rebuilds stale artifacts', () => {
/**
* Helper: create a self-contained temp TypeScript project with two source files
* (sentinelmod.cts and targetmod.cts) and a tsconfig that emits to <tmp>/out.
* Returns { tmp, overrides, sentinelOut, targetOut, tsBuildInfoPath }.
*
* HERMETIC: all destructive tests use this helper. They NEVER touch the real
* gsd-core/bin/lib/*.cjs or the real tsbuildinfo. (Regression from #996 fixed here.)
*/
function makeTempProject() {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-bug969-'));
const srcDir = path.join(tmp, 'src');
const outDir = path.join(tmp, 'out');
const tsBuildInfoPath = path.join(outDir, '.tsbuildinfo');
const tsconfigPath = path.join(tmp, 'tsconfig.build.json');
fs.mkdirSync(srcDir, { recursive: true });
fs.mkdirSync(outDir, { recursive: true });
fs.writeFileSync(path.join(srcDir, 'sentinelmod.cts'), 'export const sentinelValue = 1;\n');
fs.writeFileSync(path.join(srcDir, 'targetmod.cts'), 'export const targetValue = 2;\n');
fs.writeFileSync(tsconfigPath, JSON.stringify({
compilerOptions: {
rootDir: 'src',
outDir: 'out',
module: 'commonjs',
target: 'es2022',
esModuleInterop: true,
noEmitOnError: true,
incremental: true,
tsBuildInfoFile: 'out/.tsbuildinfo',
},
include: ['src/**/*.cts'],
}, null, 2));
const overrides = { root: tmp, srcDir, outDir, tsBuildInfoPath, tsconfigPath };
const sentinelOut = path.join(outDir, 'sentinelmod.cjs');
const targetOut = path.join(outDir, 'targetmod.cjs');
return { tmp, overrides, sentinelOut, targetOut, tsBuildInfoPath };
}
/**
* FAIL-BEFORE (origin/next behavior):
* The old code contained `if (existsSync(sentinel)) return;`. When the
* sentinel (semver-compare.cjs) was present, the function returned early
* without touching any other .cjs. This test confirms the new code always
* invokes tsc — it would have returned immediately on origin/next.
*
* Specifically: on origin/next, after deleting a non-sentinel artifact +
* its tsbuildinfo and calling ensureBuiltArtifacts() with sentinel present,
* the artifact would remain absent. On the fix, tsc runs unconditionally
* and recreates it.
*
* PASS-AFTER (fix):
* The sentinel guard is removed. ensureBuiltArtifacts() always invokes tsc.
* With no tsbuildinfo present (clean state), tsc performs a full emit and
* recreates all .cjs outputs including the deleted non-sentinel artifact.
*
* HERMETIC: this test operates on a self-contained temp project. It NEVER
* touches gsd-core/bin/lib/core.cjs or the real tsbuildinfo. (Fixed from #996.)
*/
test('rebuilds a non-sentinel artifact (with no tsbuildinfo) even when sentinel exists', () => {
const { tmp, overrides, sentinelOut, targetOut, tsBuildInfoPath } = makeTempProject();
try {
// Initial build — both outputs must appear.
ensureBuiltArtifacts(overrides);
assert.ok(fs.existsSync(sentinelOut), 'initial build: sentinelmod.cjs must exist');
assert.ok(fs.existsSync(targetOut), 'initial build: targetmod.cjs must exist');
// Simulate: fresh CI checkout — target artifact missing, no tsbuildinfo.
fs.unlinkSync(targetOut);
if (fs.existsSync(tsBuildInfoPath)) fs.unlinkSync(tsBuildInfoPath);
assert.ok(!fs.existsSync(targetOut), 'pre-condition: targetmod.cjs must be absent');
assert.ok(fs.existsSync(sentinelOut), 'pre-condition: sentinelmod.cjs must still be present');
// Under the OLD code this returned immediately (sentinel present → return).
// Under the NEW code this calls tsc unconditionally → full emit → recreated.
ensureBuiltArtifacts(overrides);
assert.ok(
fs.existsSync(targetOut),
'ensureBuiltArtifacts must recreate targetmod.cjs even when sentinelmod.cjs ' +
'exists (sentinel-short-circuit was removed in fix #969)'
);
} finally {
cleanup(tmp);
}
});
/**
* PASS-AFTER: the unconditional build emits the expected output (sentinelmod.cjs).
* Uses the temp project helper so this test is fully hermetic — it never touches
* the real gsd-core/bin/lib tree.
*/
test('sentinel (semver-compare.cjs) still exists after unconditional build', () => {
const { tmp, overrides, sentinelOut } = makeTempProject();
try {
ensureBuiltArtifacts(overrides);
assert.ok(fs.existsSync(sentinelOut), 'sentinel output (sentinelmod.cjs) must exist after ensureBuiltArtifacts');
} finally {
cleanup(tmp);
}
});
/**
* PERSISTENT-MIRROR CASE — the residual hole found by adversarial review.
*
* FAIL-BEFORE (incremental: true — the old behavior on this branch):
* With "incremental": true in tsconfig.build.json, tsc reads the .tsbuildinfo
* on disk. If sources are unchanged since the last build, tsc skips re-emitting
* any outputs — including outputs that were deleted or overwritten by an rsync
* from a different branch. This is the persistent-docker-mirror scenario:
* 1. A prior branch rsync'd a stale core.cjs into bin/lib/
* 2. A stale tsbuildinfo is present (from that same branch)
* 3. ensureBuiltArtifacts() calls tsc (incremental)
* 4. tsc sees "sources unchanged vs tsbuildinfo" → no-ops → stale .cjs served
* With "incremental": true this test would FAIL because targetmod.cjs remains absent.
*
* PASS-AFTER (step-3 unlink+clean-reemit logic):
* When a missing/zero-bytes output is detected after the incremental pass,
* ensureBuiltArtifacts() unlinks the tsbuildinfo and runs tsc a second time
* (clean re-emit). The stale/missing output is always regenerated.
*
* HERMETIC: this test operates on a self-contained temp project. It NEVER
* touches gsd-core/bin/lib/core.cjs or the real tsbuildinfo. (Fixed from #996.)
*/
test('PERSISTENT-MIRROR: rebuilds stale output even when tsbuildinfo is present (non-incremental is authoritative)', () => {
const { tmp, overrides, targetOut, tsBuildInfoPath } = makeTempProject();
const STALE_TSBUILDINFO = JSON.stringify({
program: { fileNames: [], options: { incremental: true } },
version: '5.0.0',
_gsd_test_marker: 'stale-persistent-mirror',
});
try {
// Initial build to populate outputs.
ensureBuiltArtifacts(overrides);
assert.ok(fs.existsSync(targetOut), 'initial build: targetmod.cjs must exist');
// Inject a stale tsbuildinfo (mirrors: old branch rsync'd state onto workspace).
fs.writeFileSync(tsBuildInfoPath, STALE_TSBUILDINFO);
// Delete the output .cjs (mirrors: stale/missing output on the persistent mirror).
fs.unlinkSync(targetOut);
assert.ok(!fs.existsSync(targetOut), 'pre-condition: targetmod.cjs must be absent');
assert.ok(fs.existsSync(tsBuildInfoPath), 'pre-condition: tsbuildinfo must be present');
// FAIL-BEFORE (incremental: true, no step-3): tsc would read the stale
// tsbuildinfo, see "sources unchanged", and skip re-emitting targetmod.cjs
// → it would remain absent.
//
// PASS-AFTER (step-3 unlink+clean-reemit): missing output detected after
// incremental pass → tsbuildinfo unlinked → tsc runs again → targetmod.cjs
// is regenerated unconditionally.
ensureBuiltArtifacts(overrides);
assert.ok(
fs.existsSync(targetOut),
'ensureBuiltArtifacts must regenerate targetmod.cjs even when a stale ' +
'tsbuildinfo is present on disk (persistent-mirror scenario — ' +
'incremental:true alone would have no-op\'d here)'
);
// Verify the regenerated file is valid JS.
const regenerated = fs.readFileSync(targetOut, 'utf-8');
assert.ok(regenerated.length > 0, 'regenerated targetmod.cjs must be non-empty');
assert.ok(
regenerated.includes('exports.') || regenerated.includes('"use strict"'),
'regenerated targetmod.cjs must look like a valid CommonJS module'
);
} finally {
cleanup(tmp);
}
});
});
// ---------------------------------------------------------------------------
// Part B — runGsdTools: timeout + kill-signal discrimination
// ---------------------------------------------------------------------------
describe('bug #969 B — runGsdTools kill-signal discrimination', () => {
const TOOLS_PATH = path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs');
/**
* Shared helper that mirrors the production runGsdTools implementation
* (from tests/helpers.cjs) but accepts an explicit timeout so we can
* trigger the kill path in tests without waiting 60 seconds.
*
* IMPORTANT: this helper is intentionally self-contained so that the test
* proves the CONTRACT of the implementation, not just calls the real
* runGsdTools (which would need a real 60s+ hang to trigger in tests).
* We test the identical logic paths using a tiny timeout.
*/
function runGsdToolsWithTimeout(args, cwd, env, timeoutMs) {
// The session-identity subset stays a local literal on purpose: this helper
// mirrors the production one to prove its CONTRACT, so it must not simply
// re-import what it is testing. The config-LOCATION keys are the exception —
// they are a safety scrub rather than part of the contract under test, and a
// hand-copied list of them is the #2665 drift this change exists to end. So
// spread the canonical derived set (tests/helpers.cjs) and keep the rest local.
const TEST_ENV_BASE = {
GSD_SESSION_KEY: '',
CODEX_THREAD_ID: '',
CLAUDE_SESSION_ID: '',
...Object.fromEntries(CONFIG_LOCATION_ENV_KEYS.map((k) => [k, ''])),
};
try {
let result;
const childEnv = { ...process.env, ...TEST_ENV_BASE, ...(env || {}) };
const argv = Array.isArray(args)
? args
: (args.match(/(?:[^\s"']+|"[^"]*"|'[^']*')+/g) || [])
.map(t => t.replace(/"([^"]*)"/g, '$1').replace(/'([^']*)'/g, '$1'));
result = execFileSync(process.execPath, [TOOLS_PATH, ...argv], {
cwd: cwd || process.cwd(),
encoding: 'utf-8',
stdio: ['pipe', 'pipe', 'pipe'],
env: childEnv,
timeout: timeoutMs,
});
return { success: true, output: result.trim(), exitCode: 0 };
} catch (err) {
// Production kill-discrimination logic (verbatim from helpers.cjs fix).
if (err.killed || err.signal != null || err.code === 'ETIMEDOUT') {
throw new Error(
`[runGsdTools: resource-starvation / subprocess-kill] ` +
`gsd-tools was killed before completion ` +
`(signal=${err.signal}, code=${err.code}, killed=${err.killed}). ` +
`This indicates host OOM or scheduler contention, not a product bug. ` +
`stdout=${err.stdout?.toString().trim() || ''} ` +
`stderr=${err.stderr?.toString().trim() || ''}`
);
}
const stderrRaw = err.stderr?.toString().trim() || '';
const error = stderrRaw || `${err.message} [stderr: (empty) exit:${err.status ?? 1}]`;
return {
success: false,
output: err.stdout?.toString().trim() || '',
error,
exitCode: err.status ?? 1,
};
}
}
/**
* FAIL-BEFORE (origin/next behavior):
* Without a timeout, an OOM-killed subprocess threw with err.killed=true
* but the catch block fell through to `return { success: false, ... }`.
* The test consumer saw a normal {success:false} result and tried to parse
* gsd-tools output from it, causing a confusing downstream assertion fail.
*
* PASS-AFTER (fix):
* The kill-discrimination guard rethrows immediately with a labeled error
* message containing "resource-starvation / subprocess-kill". The test
* asserts on that throw rather than getting a silent {success:false}.
*
* Mechanism: we use a tiny timeout (1ms) to guarantee a timeout-kill on a
* real gsd-tools invocation (even `--help` takes >1ms to start node).
*/
test('throws a resource-starvation error when subprocess is killed/times out', () => {
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-969-'));
try {
// 1ms timeout guarantees ETIMEDOUT / killed before gsd-tools can respond.
assert.throws(
() => runGsdToolsWithTimeout(['--help'], tmpDir, {}, 1),
(err) => {
assert.ok(
err.message.includes('resource-starvation / subprocess-kill'),
`Expected labeled resource-starvation error, got: ${err.message}`
);
return true;
}
);
} finally {
cleanup(tmpDir);
}
});
/**
* Verify that a normal fast command still returns { success: true } and does
* NOT throw — i.e., the timeout addition does not break the happy path.
*/
test('returns { success: true } for a normal fast command with generous timeout', () => {
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-969-'));
try {
// 30s timeout; gsd-tools --help completes in well under 1s.
const result = runGsdToolsWithTimeout(['--help'], tmpDir, {}, 30000);
assert.ok(result.success === true, `Expected success:true, got ${JSON.stringify(result)}`);
assert.ok(typeof result.output === 'string', 'output must be a string');
} finally {
cleanup(tmpDir);
}
});
/**
* Verify that a clean non-zero exit (a real gsd-tools application error, not
* a kill) still returns { success: false } WITHOUT throwing. This preserves
* existing test behavior that asserts on error shape.
*
* We trigger a clean non-zero by invoking a command that is known to fail
* cleanly (no project directory set up).
*/
test('returns { success: false } for a clean non-zero exit (no throw)', () => {
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-969-'));
try {
// 'phase list' on a directory with no .planning/ produces a clean error exit.
const result = runGsdToolsWithTimeout(['phase', 'list'], tmpDir, {}, 30000);
assert.ok(result.success === false, `Expected success:false for clean error, got ${JSON.stringify(result)}`);
assert.ok(result.exitCode !== 0, 'exitCode must be non-zero');
// Must NOT have thrown — the clean-error path returns normally.
} finally {
cleanup(tmpDir);
}
});
});
// ---------------------------------------------------------------------------
// Part C — ensureBuiltHooks: build hooks/dist once, closing the scoped-CI
// first-build empty-dir race.
// ---------------------------------------------------------------------------
describe('bug #969 C — ensureBuiltHooks populates hooks/dist before concurrent tests', () => {
/**
* Helper: a hermetic temp dist dir + a runBuild spy. The spy records how many
* times a build was requested and, when invoked, writes the given hook files
* (simulating build-hooks.js populating DIST_DIR) so idempotency is testable.
*
* HERMETIC: never touches the real hooks/dist. Uses dependency-injected
* overrides (distDir, hookNames, runBuild) — no fs monkeypatching, so the test
* is deterministic and root/OS-independent.
*/
function makeHooksFixture(hookNames) {
const tmp = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-969-hooks-'));
const distDir = path.join(tmp, 'hooks', 'dist');
let buildCalls = 0;
const runBuild = () => {
buildCalls += 1;
fs.mkdirSync(distDir, { recursive: true });
for (const h of hookNames) {
fs.writeFileSync(path.join(distDir, h), `// ${h}\nmodule.exports = {};\n`);
}
};
const overrides = () => ({ distDir, hookNames, runBuild });
return { tmp, distDir, hookNames, overrides, calls: () => buildCalls };
}
const HOOKS = ['a-hook.js', 'b-hook.js', 'c-hook.sh'];
/**
* FAIL-BEFORE (origin/next): ensureBuiltHooks did not exist, so the export was
* undefined and there was no upfront hooks build — the first concurrent
* install test raced build-hooks.js's empty-then-fill window. The import at the
* top of this file (ensureBuiltHooks) is itself the fail-first anchor: on
* origin/next it is undefined and every test below throws "not a function".
*
* PASS-AFTER: ensureBuiltHooks() builds when hooks/dist is entirely absent.
*/
test('builds when hooks/dist is absent (fresh checkout / scoped CI)', () => {
const fx = makeHooksFixture(HOOKS);
try {
assert.equal(typeof ensureBuiltHooks, 'function',
'ensureBuiltHooks must be exported (absent on origin/next — the fail-first anchor)');
assert.ok(!fs.existsSync(fx.distDir), 'pre-condition: hooks/dist must be absent');
ensureBuiltHooks(fx.overrides());
assert.equal(fx.calls(), 1, 'a build must be triggered when dist is absent');
for (const h of HOOKS) {
assert.ok(fs.existsSync(path.join(fx.distDir, h)), `${h} must exist after build`);
}
} finally {
cleanup(fx.tmp);
}
});
/**
* BOUNDARY: dist exists but is empty (0 of N hooks) — the exact transient state
* build-hooks.js exposes between mkdir(DIST_DIR) and the first file rename.
*/
test('builds when hooks/dist exists but is empty (0 of N)', () => {
const fx = makeHooksFixture(HOOKS);
try {
fs.mkdirSync(fx.distDir, { recursive: true }); // empty dir — the race window
ensureBuiltHooks(fx.overrides());
assert.equal(fx.calls(), 1, 'an empty dist must trigger a build');
} finally {
cleanup(fx.tmp);
}
});
/**
* BOUNDARY: dist has all-but-one hook (N-1 of N) — a partially-filled dir mid
* first-build. Must still be treated as incomplete and rebuilt.
*/
test('builds when hooks/dist is partial (N-1 of N)', () => {
const fx = makeHooksFixture(HOOKS);
try {
fs.mkdirSync(fx.distDir, { recursive: true });
for (const h of HOOKS.slice(0, HOOKS.length - 1)) {
fs.writeFileSync(path.join(fx.distDir, h), 'x');
}
ensureBuiltHooks(fx.overrides());
assert.equal(fx.calls(), 1, 'a partial dist (missing one hook) must trigger a build');
} finally {
cleanup(fx.tmp);
}
});
/**
* BOUNDARY: a zero-byte hook (N of N present, but one is 0 bytes) — a truncated
* mid-write file. statSync().size === 0 must count as incomplete → rebuild.
*/
test('builds when a hook file is present but zero-byte', () => {
const fx = makeHooksFixture(HOOKS);
try {
fs.mkdirSync(fx.distDir, { recursive: true });
HOOKS.forEach((h, i) => {
fs.writeFileSync(path.join(fx.distDir, h), i === 0 ? '' : 'ok'); // first is 0 bytes
});
ensureBuiltHooks(fx.overrides());
assert.equal(fx.calls(), 1, 'a zero-byte hook must be treated as incomplete → rebuild');
} finally {
cleanup(fx.tmp);
}
});
/**
* BOUNDARY + idempotency: dist is complete (all N present, non-empty). No build
* must fire — this keeps nested run-tests spawns and repeat invocations cheap
* and avoids a redundant concurrent build against an already-populated dist.
*/
test('no-op when hooks/dist is complete (N of N non-empty)', () => {
const fx = makeHooksFixture(HOOKS);
try {
fs.mkdirSync(fx.distDir, { recursive: true });
for (const h of HOOKS) fs.writeFileSync(path.join(fx.distDir, h), 'ok');
ensureBuiltHooks(fx.overrides());
assert.equal(fx.calls(), 0, 'a complete dist must NOT trigger a build (idempotent no-op)');
} finally {
cleanup(fx.tmp);
}
});
/**
* Integration guard: the REAL default hook set (from build-hooks.js) is what
* ensureBuiltHooks checks when no override is given. Prove the real export is a
* non-empty list so the completeness predicate can never vacuously pass.
*/
test('default hook set (build-hooks.js HOOKS_TO_COPY) is a non-empty list', () => {
const { HOOKS_TO_COPY } = require('../scripts/build-hooks.js');
assert.ok(Array.isArray(HOOKS_TO_COPY) && HOOKS_TO_COPY.length > 0,
'HOOKS_TO_COPY must be a non-empty array or the completeness check is vacuous');
});
});
});
}
// ---------------------------------------------------------------------------
// #2456 — chunk weights must reflect MEASURED cost, not a filename guess.
//
// These tests drive the packer's pure IR (`packChunks` / `makeFileWeigher` /
// `loadTestTimings`) rather than the `run-tests: chunk N/M` stderr line, so they
// assert on typed values (chunk composition, weights) instead of rendered text.
// ---------------------------------------------------------------------------
const {
packChunks,
makeFileWeigher,
loadTestTimings,
positiveNumberEnv,
DEFAULT_TIMINGS_PATH,
} = require('../scripts/run-tests.cjs');
describe('chunk packing weights measured cost (#2456)', () => {
// A cost profile modelled on the real measurements in the issue. It is the
// MISCALIBRATION that matters: the pre-#2456 heuristic scored basenames
// matching /^(?:install|codex-)/ at 12 and everything else at 1, which is
// wrong in BOTH directions here —
// * run-tests-harness / release-tarball-smoke.install are the two most
// expensive files yet scored 1 (the regex is anchored to the START of the
// basename, so a mid-name "install" never matches), and
// * installer-migration-authoring / codex-declarative-reference scored 12
// while costing almost nothing.
const MEASURED_MS = {
'run-tests-harness.test.cjs': 150180,
'release-tarball-smoke.install.test.cjs': 144420,
'phase.test.cjs': 136770,
'install-minimal-hooks.test.cjs': 136330,
'config.test.cjs': 114420,
'state.test.cjs': 94290,
'commands.test.cjs': 70580,
'init.test.cjs': 64100,
'installer-migration-authoring.test.cjs': 90,
'installer-migration-report.test.cjs': 120,
'install-update-marker.test.cjs': 95,
'codex-declarative-reference.test.cjs': 110,
};
const FILES = Object.keys(MEASURED_MS);
const FIXED_OVERHEAD = 120;
const ROOMY_CHARS = 100000;
function tableFrom(timings) {
const tmp = createTempDir('gsd-2456-timings-');
const p = path.join(tmp, 'timings.json');
fs.writeFileSync(p, JSON.stringify({ schema_version: 1, unit: 'ms', timings }), 'utf8');
return { path: p, dir: tmp };
}
function packMeasured(files, maxWeight, extra = {}) {
const t = tableFrom(MEASURED_MS);
try {
return packChunks(files, {
weightOf: makeFileWeigher(loadTestTimings(t.path)),
maxWeight,
maxChars: ROOMY_CHARS,
fixedOverhead: FIXED_OVERHEAD,
...extra,
});
} finally {
cleanup(t.dir);
}
}
const costOf = (chunk) => chunk.reduce((sum, f) => sum + MEASURED_MS[f], 0);
/**
* THE REGRESSION (#2456). Under the pre-fix packer the two most expensive
* files both scored weight 1 and, being adjacent in selection order, packed
* into the SAME chunk — that chunk ran ~3.9x the lightest and sat near the
* 600s per-chunk timeout. Weighting by measured cost and packing with LPT must
* separate them.
*/
test('the two most expensive files never land in the same chunk', () => {
const chunks = packMeasured(FILES, 4);
const chunkOf = (f) => chunks.findIndex((c) => c.includes(f));
const a = chunkOf('run-tests-harness.test.cjs');
const b = chunkOf('release-tarball-smoke.install.test.cjs');
assert.ok(a !== -1 && b !== -1, 'both heavy files must be packed');
assert.notStrictEqual(
a,
b,
`the two most expensive files must be split across chunks; got both in chunk ${a + 1}. ` +
`Chunk costs (s): ${chunks.map((c) => (costOf(c) / 1000).toFixed(1)).join(', ')}`,
);
});
test('chunk costs are balanced — slowest chunk stays well under 2x the fastest', () => {
const chunks = packMeasured(FILES, 4);
const costs = chunks.map(costOf);
const ratio = Math.max(...costs) / Math.min(...costs);
assert.ok(
ratio < 2,
`LPT must balance measured cost; imbalance was ${ratio.toFixed(2)}x ` +
`(costs in s: ${costs.map((c) => (c / 1000).toFixed(1)).join(', ')})`,
);
});
test('weights follow measured duration, correcting the old prefix heuristic both ways', () => {
const t = tableFrom(MEASURED_MS);
try {
const weigh = makeFileWeigher(loadTestTimings(t.path));
// Old heuristic: 1. Actual: the most expensive file in the suite.
assert.ok(
weigh('run-tests-harness.test.cjs') > 1,
'the heaviest file must weigh above the table average',
);
// Old heuristic: 12 (install prefix). Actual: ~0.1s.
assert.ok(
weigh('installer-migration-authoring.test.cjs') < 1,
'a trivially cheap install-prefixed file must weigh below the table average',
);
// The old regex is anchored to the START of the basename, so a mid-name
// "install" scored 1. Measured, it is the second most expensive file.
assert.ok(
weigh('release-tarball-smoke.install.test.cjs')
> weigh('installer-migration-authoring.test.cjs') * 100,
'mid-name "install" must be ranked by cost, not by prefix position',
);
} finally {
cleanup(t.dir);
}
});
test('an average-cost file weighs exactly 1, preserving the MAX_FILES_PER_CHUNK scale', () => {
// Three files at 10s, 20s, 30s → mean 20s. The 20s file is the average.
const t = tableFrom({ 'a.test.cjs': 10000, 'b.test.cjs': 20000, 'c.test.cjs': 30000 });
try {
const weigh = makeFileWeigher(loadTestTimings(t.path));
assert.strictEqual(weigh('b.test.cjs'), 1, 'the mean-cost file must weigh 1');
assert.strictEqual(weigh('a.test.cjs'), 0.5);
assert.strictEqual(weigh('c.test.cjs'), 1.5);
} finally {
cleanup(t.dir);
}
});
describe('timings are advisory, never gated', () => {
test('a file missing from the table falls back to the mean weight (1), not the median', () => {
// 10s, 20s, 60s → mean 30s, median 20s. Absent files weigh 1 (the mean),
// not 20/30 (the median) — see #2456 follow-up, red-next 2026-09-14.
const t = tableFrom({ 'a.test.cjs': 10000, 'b.test.cjs': 20000, 'c.test.cjs': 60000 });
try {
const weigh = makeFileWeigher(loadTestTimings(t.path));
assert.strictEqual(weigh('brand-new-test.test.cjs'), 1);
} finally {
cleanup(t.dir);
}
});
test('an unknown file packs without error rather than failing the run', () => {
const chunks = packMeasured([...FILES, 'never-measured.test.cjs'], 4);
assert.ok(
chunks.flat().includes('never-measured.test.cjs'),
'a file absent from the timings table must still be packed',
);
});
test('a missing timings file degrades to uniform weight, not an error', () => {
// No temp dir needed — the point is a path that does NOT exist. Creating
// one here would leak it (createTempDir has no registry).
const missing = path.join(__dirname, 'no-such-dir-2456', 'does-not-exist.json');
assert.strictEqual(loadTestTimings(missing), null, 'a missing table must load as null');
const weigh = makeFileWeigher(null);
assert.strictEqual(weigh('anything.test.cjs'), 1, 'a null table must weigh every file 1');
});
test('a corrupt timings file degrades to uniform weight, not an error', () => {
const tmp = createTempDir('gsd-2456-corrupt-');
try {
const p = path.join(tmp, 'timings.json');
fs.writeFileSync(p, '{ this is not json', 'utf8');
assert.strictEqual(loadTestTimings(p), null, 'unparseable JSON must load as null');
} finally {
cleanup(tmp);
}
});
test('a structurally valid but empty table degrades to uniform weight', () => {
const t = tableFrom({});
try {
assert.strictEqual(loadTestTimings(t.path), null, 'an empty table must load as null');
} finally {
cleanup(t.dir);
}
});
test('chunk count never drops below what count-based packing would produce', () => {
const files = Array.from({ length: 30 }, (_, i) => `unmeasured-${i}.test.cjs`);
assert.ok(
packMeasured(files, 6).length >= Math.ceil(files.length / 6),
'the count floor must hold for a fully-unknown file set',
);
});
test('a right-skewed table cannot collapse unknown files into fat chunks', () => {
// The safety floor in its worst case. In a realistically right-skewed suite
// (many trivial files, a few very expensive ones) the median weight is far
// below 1, so weighting ALONE would pack 30 unknown files into a single
// chunk — exactly the failure mode this fix exists to prevent. The count
// floor pins the result at what the pre-#2456 packer produced.
const skewed = {
...Object.fromEntries(Array.from({ length: 10 }, (_, i) => [`t-${i}.test.cjs`, 100])),
'one-huge.test.cjs': 100000,
};
const t = tableFrom(skewed);
try {
const table = loadTestTimings(t.path);
assert.ok(table.medianWeight < 0.05, 'fixture must actually be right-skewed');
const files = Array.from({ length: 30 }, (_, i) => `unmeasured-${i}.test.cjs`);
const chunks = packChunks(files, {
weightOf: makeFileWeigher(table),
maxWeight: 6,
maxChars: ROOMY_CHARS,
fixedOverhead: FIXED_OVERHEAD,
});
assert.strictEqual(
chunks.length,
Math.ceil(files.length / 6),
'a right-skewed table must still chunk exactly as count-based packing did',
);
} finally {
cleanup(t.dir);
}
});
});
describe('an unmeasured file is charged the mean, not the median (red next, 2026-09-14)', () => {
// MEASURED_MS (above) is NOT right-skewed the way the real table is: its
// median (~82435) sits ABOVE its mean (~75959), so medianWeight ≈ 1.09 —
// the bug this block guards ("median under-declares an unknown file in a
// right-skewed table") is not even expressible against that fixture. Do
// NOT reuse MEASURED_MS here or fold this fixture into it; other tests
// above depend on MEASURED_MS's exact shape.
//
// SKEWED_MS models the real suite's shape instead: mostly trivial files
// with a long right tail of a few very expensive ones (real table: mean
// 7152ms, median 381ms, ratio 0.0533).
const SKEWED_MS = {
...Object.fromEntries(Array.from({ length: 16 }, (_, i) => [`small-${i}.test.cjs`, 350])),
'huge-a.test.cjs': 100000,
'huge-b.test.cjs': 200000,
'huge-c.test.cjs': 400000,
'huge-d.test.cjs': 575000,
};
test('a file absent from the table weighs 1 (the mean), not the median', () => {
const t = tableFrom(SKEWED_MS);
try {
const table = loadTestTimings(t.path);
// Prove the fixture actually models the skew this test exists to
// guard against — without this, a passing assertion below would be
// meaningless.
assert.ok(
table.medianWeight < 0.2,
`fixture must be right-skewed; got medianWeight=${table.medianWeight}`,
);
const weigh = makeFileWeigher(table);
assert.strictEqual(weigh('never-measured.test.cjs'), 1);
} finally {
cleanup(t.dir);
}
});
test('that matches what a MISSING table already does — both mean "unknown"', () => {
const t = tableFrom(SKEWED_MS);
try {
const weigh = makeFileWeigher(loadTestTimings(t.path));
const weighNull = makeFileWeigher(null);
assert.strictEqual(weighNull('anything.test.cjs'), 1);
assert.strictEqual(weighNull('anything.test.cjs'), weigh('never-measured.test.cjs'));
} finally {
cleanup(t.dir);
}
});
test('a chunk of unmeasured files declares their real share of the budget', () => {
const t = tableFrom(SKEWED_MS);
try {
const table = loadTestTimings(t.path);
const weigh = makeFileWeigher(table);
const unmeasured = Array.from({ length: 10 }, (_, i) => `new-${i}.test.cjs`);
const total = unmeasured.reduce((sum, f) => sum + weigh(f), 0);
// Under the old median rule this would have summed to about
// 10 * table.medianWeight (~0.05 against this fixture), letting ten
// unknown files hide inside a chunk that looked essentially empty.
assert.ok(total >= 10, `expected sum >= 10, got ${total}`);
} finally {
cleanup(t.dir);
}
});
});
describe('boundary coverage', () => {
// 12 files, each weighing exactly 1 (uniform) → total weight 12.
const uniform = Array.from({ length: 12 }, (_, i) => `u-${String(i).padStart(2, '0')}.test.cjs`);
const packUniform = (maxWeight) =>
packChunks(uniform, {
weightOf: () => 1,
maxWeight,
maxChars: ROOMY_CHARS,
fixedOverhead: FIXED_OVERHEAD,
});
test('weight budget at limit-1 forces a second chunk', () => {
assert.strictEqual(packUniform(11).length, 2, 'total weight 12 over budget 11 → 2 chunks');
});
test('weight budget exactly at the limit stays in one chunk', () => {
assert.strictEqual(packUniform(12).length, 1, 'total weight 12 at budget 12 → 1 chunk');
});
test('weight budget at limit+1 stays in one chunk', () => {
assert.strictEqual(packUniform(13).length, 1, 'total weight 12 under budget 13 → 1 chunk');
});
// argv ceiling: two files of identical length in one chunk occupies
// fixedOverhead + 2*(len+1) chars.
const two = ['argv-boundary-aaa.test.cjs', 'argv-boundary-bbb.test.cjs'];
const exactChars = FIXED_OVERHEAD + two.reduce((sum, f) => sum + f.length + 1, 0);
const packChars = (maxChars) =>
packChunks(two, {
weightOf: () => 1,
maxWeight: 1000, // never the binding constraint here
maxChars,
fixedOverhead: FIXED_OVERHEAD,
});
test('argv ceiling at limit-1 splits the chunk', () => {
assert.strictEqual(packChars(exactChars - 1).length, 2, 'one char short → must split');
});
test('argv ceiling exactly at the limit keeps one chunk', () => {
assert.strictEqual(packChars(exactChars).length, 1, 'exactly at the ceiling → fits');
});
test('argv ceiling at limit+1 keeps one chunk', () => {
assert.strictEqual(packChars(exactChars + 1).length, 1, 'one char spare → fits');
});
test('a single file wider than the argv ceiling is packed alone, not dropped or looped', () => {
const chunks = packChunks(['x'.repeat(500) + '.test.cjs', 'small.test.cjs'], {
weightOf: () => 1,
maxWeight: 1000,
maxChars: FIXED_OVERHEAD + 50, // narrower than the long file alone
fixedOverhead: FIXED_OVERHEAD,
});
assert.strictEqual(chunks.flat().length, 2, 'both files must survive packing');
assert.strictEqual(chunks.length, 2, 'the over-long file must occupy its own chunk');
});
test('an empty selection packs to no chunks', () => {
assert.deepStrictEqual(
packChunks([], {
weightOf: () => 1,
maxWeight: 60,
maxChars: ROOMY_CHARS,
fixedOverhead: FIXED_OVERHEAD,
}),
[],
);
});
});
describe('packing invariants', () => {
test('every selected file is packed exactly once', () => {
const chunks = packMeasured(FILES, 3);
const flat = chunks.flat();
assert.strictEqual(flat.length, FILES.length, 'no file may be dropped or duplicated');
assert.deepStrictEqual([...flat].sort(), [...FILES].sort(), 'the packed set must equal the selection');
});
test('no chunk is empty', () => {
const chunks = packMeasured(FILES, 3);
for (const [i, c] of chunks.entries()) {
assert.ok(c.length > 0, `chunk ${i + 1} must not be empty`);
}
});
test('packing is deterministic — identical input yields byte-identical chunks', () => {
assert.deepStrictEqual(packMeasured(FILES, 4), packMeasured(FILES, 4));
});
test('files keep their original selection order within a chunk', () => {
const chunks = packMeasured(FILES, 3);
for (const chunk of chunks) {
const positions = chunk.map((f) => FILES.indexOf(f));
assert.deepStrictEqual(
positions,
[...positions].sort((a, b) => a - b),
'within a chunk, files must stay in selection order',
);
}
});
});
describe('degenerate knobs degrade safely rather than hanging or crashing', () => {
// These knobs come from the environment via Number(), so a typo yields NaN
// and an explicit 0 yields 0. Both reach the chunk-count arithmetic. Before
// hardening, NaN spun packChunks' retry loop forever (a hung CI job with no
// output) and 0 threw `RangeError: Invalid array length` from Array.from.
// Every case below must return a valid packing instead.
const files = ['a.test.cjs', 'b.test.cjs', 'c.test.cjs'];
const packWith = (opts) =>
packChunks(files, {
weightOf: () => 1,
maxWeight: 60,
maxChars: ROOMY_CHARS,
fixedOverhead: FIXED_OVERHEAD,
...opts,
});
const conserves = (chunks) =>
JSON.stringify(chunks.flat().sort()) === JSON.stringify([...files].sort());
for (const [label, opts] of [
['a NaN weight budget', { maxWeight: NaN }],
['a zero weight budget', { maxWeight: 0 }],
['a negative weight budget', { maxWeight: -5 }],
['a NaN argv ceiling', { maxChars: NaN }],
['a zero argv ceiling', { maxChars: 0 }],
['a NaN fixed overhead', { fixedOverhead: NaN }],
['a weight function returning NaN', { weightOf: () => NaN }],
['a weight function returning Infinity', { weightOf: () => Infinity }],
['a weight function returning a negative', { weightOf: () => -1 }],
]) {
test(`${label} still packs every file exactly once`, () => {
const chunks = packWith(opts);
assert.ok(conserves(chunks), `${label} must still pack all files; got ${JSON.stringify(chunks)}`);
});
}
test('a legitimate but tiny weight budget does not explode the chunk count', () => {
// positiveNumberEnv accepts any positive finite number, so 1e-9 is a valid
// budget. Without an upper clamp the packer asks for files.length / 1e-9
// bins and throws `RangeError: Invalid array length`. More chunks than
// files is never useful, so the count clamps at one file per chunk.
const many = Array.from({ length: 200 }, (_, i) => `m-${i}.test.cjs`);
const chunks = packChunks(many, {
weightOf: () => 1,
maxWeight: 1e-9,
maxChars: ROOMY_CHARS,
fixedOverhead: FIXED_OVERHEAD,
});
assert.strictEqual(chunks.length, many.length, 'must clamp to one file per chunk');
assert.strictEqual(chunks.flat().length, many.length, 'no file may be dropped');
});
test('a timings table from a future schema is ignored rather than mis-read', () => {
// A v2 table could change the unit or key format; consuming it under v1
// semantics would silently mis-weight every file. Falling back to null
// (uniform weight) is the same graceful degradation as a missing table.
const tmp = createTempDir('gsd-2456-schema-');
try {
const p2 = path.join(tmp, 'timings.json');
fs.writeFileSync(p2, JSON.stringify({ schema_version: 2, timings: { 'a.test.cjs': 100 } }), 'utf8');
assert.strictEqual(loadTestTimings(p2), null, 'an unknown schema_version must load as null');
fs.writeFileSync(p2, JSON.stringify({ schema_version: 1, timings: { 'a.test.cjs': 100 } }), 'utf8');
assert.ok(loadTestTimings(p2) !== null, 'the supported schema_version must load');
} finally {
cleanup(tmp);
}
});
test('positiveNumberEnv falls back for every non-positive-finite input', () => {
for (const bad of [undefined, null, '', ' ', 'abc', '0', '-1', 'NaN', 'Infinity', '1e999']) {
assert.strictEqual(
positiveNumberEnv(bad, 60),
60,
`${JSON.stringify(bad)} must fall back to the default`,
);
}
});
test('positiveNumberEnv accepts a legitimate override', () => {
assert.strictEqual(positiveNumberEnv('3', 60), 3);
assert.strictEqual(positiveNumberEnv('0.5', 60), 0.5);
});
test('a key that resolves on Object.prototype still weighs a number, not a function', () => {
// The table is JSON-parsed, so a bare index would walk the prototype
// chain. Real selections are `*.test.cjs` basenames, which can never equal
// an Object.prototype member — so these BARE names are the only inputs
// that actually reach that path, and makeFileWeigher is exported, so it
// does not control its caller's strings. The contract is that any key not
// present in the table weighs 1 (the mean), whatever it resolves to.
const t = tableFrom({ 'a.test.cjs': 10000, 'b.test.cjs': 20000, 'c.test.cjs': 60000 });
try {
const weigh = makeFileWeigher(loadTestTimings(t.path));
for (const name of ['constructor', 'toString', 'valueOf', 'hasOwnProperty', '__proto__']) {
const w = weigh(name);
assert.strictEqual(typeof w, 'number', `${name} must weigh a number, not a function`);
assert.strictEqual(w, 1, `${name} must fall back to the mean weight`);
}
} finally {
cleanup(t.dir);
}
});
test('a __proto__ key in the table does not pollute Object.prototype', () => {
const t = tableFrom(JSON.parse('{"__proto__":{"polluted":true},"a.test.cjs":1000}'));
try {
makeFileWeigher(loadTestTimings(t.path))('a.test.cjs');
assert.strictEqual({}.polluted, undefined, 'Object.prototype must not be polluted');
} finally {
cleanup(t.dir);
}
});
});
describe('the checked-in timings table', () => {
test('loads, is non-empty, and yields usable weights', () => {
const table = loadTestTimings(DEFAULT_TIMINGS_PATH);
assert.ok(table !== null, `${DEFAULT_TIMINGS_PATH} must parse into a usable timing table`);
assert.ok(Object.keys(table.timings).length > 0, 'the table must not be empty');
assert.ok(table.mean > 0, 'the table mean must be positive');
assert.ok(table.medianWeight > 0, 'the median fallback weight must be positive');
});
// Schema guard on a static checked-in data file — NOT a timing assertion.
// Reported as a list so a corrupt regeneration names every bad entry at once
// rather than failing on the first.
test('every recorded cost is a finite non-negative number', () => {
const table = loadTestTimings(DEFAULT_TIMINGS_PATH);
const invalid = Object.entries(table.timings)
.filter(([, value]) => typeof value !== 'number' || !Number.isFinite(value) || value < 0)
.map(([file, value]) => `${file}=${value}`);
assert.deepStrictEqual(invalid, [], 'every timing-table entry must be a finite non-negative number');
});
});
});
// ─── analyzeChunkEvents (#3889 durability fix) ──────────────────────────────
//
// scripts/lib/ndjson-reporter.cjs now writes durably (fs.appendFileSync to a
// GSD_RUN_TESTS_EVENTS_FILE path) instead of yielding strings for Node to
// pipe through a buffered --test-reporter-destination WriteStream that a
// SIGKILL can wipe out before it flushes. These tests exercise the READER
// side (analyzeChunkEvents) directly against a hand-written events file, so
// they pin the reader's contract independently of whether the writer side
// managed to flush anything in a given subprocess run.
const { analyzeChunkEvents } = require('../scripts/run-tests.cjs');
describe('analyzeChunkEvents (#3889)', () => {
let tmpDir;
beforeEach(() => {
tmpDir = createTempDir('gsd-3889-events-');
});
afterEach(() => {
cleanup(tmpDir);
});
test('a missing events file is reported as an explicit read error, not silently as "no events"', () => {
const missingPath = path.join(tmpDir, 'does-not-exist.ndjson');
const result = analyzeChunkEvents(missingPath);
assert.strictEqual(result.readError, true, 'a missing file must set readError=true');
assert.strictEqual(result.sawAnyEvent, false);
assert.deepStrictEqual(result.files, []);
});
test('an existing-but-empty events file is distinguished from a missing one (readError=false)', () => {
const emptyPath = path.join(tmpDir, 'empty.ndjson');
fs.writeFileSync(emptyPath, '', 'utf8');
const result = analyzeChunkEvents(emptyPath);
assert.strictEqual(result.readError, false, 'an existing empty file must NOT be reported as a read error');
assert.strictEqual(result.sawAnyEvent, false);
assert.deepStrictEqual(result.files, []);
});
test('a truncated final line does not crash the reader and earlier complete lines are still reported', () => {
const truncatedPath = path.join(tmpDir, 'truncated.ndjson');
const complete = [
JSON.stringify({ type: 'test:start', file: 'a.test.cjs', name: 'first', nesting: 0, testNumber: 1, ts: 1000 }),
JSON.stringify({ type: 'test:pass', file: 'a.test.cjs', name: 'first', nesting: 0, testNumber: 1, ts: 1010 }),
JSON.stringify({ type: 'test:start', file: 'b.test.cjs', name: 'hangs', nesting: 0, testNumber: 1, ts: 1020 }),
].join('\n');
// Simulate a SIGKILL mid-appendFileSync: the trailing line is cut off
// partway through a JSON object, exactly as an unbuffered but non-atomic
// write can be interrupted.
const truncatedTrailer = '\n{"type":"test:start","file":"c.test.cjs","name":"cut off mid-writ';
fs.writeFileSync(truncatedPath, complete + truncatedTrailer, 'utf8');
const result = analyzeChunkEvents(truncatedPath);
assert.strictEqual(result.readError, false, 'a readable-but-truncated file must not be a read error');
// The complete test:start (b.test.cjs) with no matching pass/fail is
// still identified as in-flight despite the unparsable trailing line.
assert.deepStrictEqual(result.files, ['b.test.cjs']);
// a.test.cjs completed (start+pass), so it must NOT show as in-flight.
assert.ok(!result.files.includes('a.test.cjs'));
// c.test.cjs never parsed (truncated line), so it cannot appear either.
assert.ok(!result.files.includes('c.test.cjs'));
assert.strictEqual(result.sawAnyEvent, true, 'the complete lines before the truncation must still count as events');
});
// Regression (#3889 root cause): test:start/test:pass/test:fail are the
// exact three event types a genuine hang guarantees are never emitted —
// node:test only surfaces a subtest event to the parent once the child
// reports it, which happens on completion. test:dequeue is the RUNNER's
// own "began this file" signal and fires independent of completion; these
// four cases pin analyzeChunkEvents' dequeue-based in-flight rule directly
// against the synthetic event shapes a real hang, and a real finish, produce.
test('a dequeued file with no terminal event is reported as in flight', () => {
const eventsPath = path.join(tmpDir, 'dequeue-only.ndjson');
const lines = [
JSON.stringify({ type: 'reporter:init', ts: 900 }),
JSON.stringify({ type: 'test:enqueue', file: 'a.test.cjs', ts: 1000 }),
JSON.stringify({ type: 'test:dequeue', file: 'a.test.cjs', ts: 1010 }),
].join('\n');
fs.writeFileSync(eventsPath, lines, 'utf8');
const result = analyzeChunkEvents(eventsPath);
assert.deepStrictEqual(result.files, ['a.test.cjs']);
assert.strictEqual(result.anyDequeued, true);
assert.strictEqual(result.sawInitMarker, true);
});
test('a dequeued file that also terminates reports nothing in flight (all files finished)', () => {
const eventsPath = path.join(tmpDir, 'dequeue-then-pass.ndjson');
const lines = [
JSON.stringify({ type: 'reporter:init', ts: 900 }),
JSON.stringify({ type: 'test:enqueue', file: 'a.test.cjs', ts: 1000 }),
JSON.stringify({ type: 'test:dequeue', file: 'a.test.cjs', ts: 1010 }),
JSON.stringify({ type: 'test:pass', file: 'a.test.cjs', ts: 1020 }),
].join('\n');
fs.writeFileSync(eventsPath, lines, 'utf8');
const result = analyzeChunkEvents(eventsPath);
assert.deepStrictEqual(result.files, [], 'a terminated file must not show as in flight');
assert.strictEqual(result.anyDequeued, true, 'the file WAS dequeued — "all files finished" is a distinct state from "nothing ran"');
});
test('one terminated file followed by a second dequeued-but-unterminated file reports only the second', () => {
const eventsPath = path.join(tmpDir, 'two-files.ndjson');
const lines = [
JSON.stringify({ type: 'reporter:init', ts: 900 }),
JSON.stringify({ type: 'test:dequeue', file: 'a.test.cjs', ts: 1000 }),
JSON.stringify({ type: 'test:pass', file: 'a.test.cjs', ts: 1010 }),
JSON.stringify({ type: 'test:dequeue', file: 'b.test.cjs', ts: 1020 }),
].join('\n');
fs.writeFileSync(eventsPath, lines, 'utf8');
const result = analyzeChunkEvents(eventsPath);
assert.deepStrictEqual(result.files, ['b.test.cjs']);
assert.ok(!result.files.includes('a.test.cjs'));
});
test('an init marker with no dequeue at all is distinguished (anyDequeued=false) from "all finished"', () => {
const eventsPath = path.join(tmpDir, 'init-only.ndjson');
fs.writeFileSync(eventsPath, JSON.stringify({ type: 'reporter:init', ts: 900 }), 'utf8');
const result = analyzeChunkEvents(eventsPath);
assert.deepStrictEqual(result.files, []);
assert.strictEqual(result.anyDequeued, false);
assert.strictEqual(result.sawInitMarker, true);
assert.strictEqual(result.sawAnyEvent, false);
});
});
// ─── partitionIsolatedFiles (#4497 codex-config.test.cjs chunk isolation) ───
//
// 2026-09-07: codex-config.test.cjs (weight 17.87, genuinely measured) is
// pulled out of the weight-balanced packing pool and given its own dedicated
// chunk, on every platform, so no future single-file addition can reshuffle a
// companion into its chunk and retrigger the per-chunk timeout two prior
// incidents already hit.
//
// 2026-09-14 (#4733): the isolated set used to be a hand-maintained Set, then
// briefly a ratio (ISOLATION_RATIO) of the per-platform file-COUNT cap
// (MAX_FILES_PER_CHUNK) — a category error (count vs. weight) that also made
// the isolated set platform-dependent, silently dropping seven of the eight
// files #4497/#4603 proved dangerous back into the shared pool on
// linux/darwin (only run-tests-harness.test.cjs, at 31.23, still cleared the
// 0.447 * 60 = 26.82 threshold there). It is now an ABSOLUTE ms bar
// (ISOLATION_BUDGET_FRACTION * CHUNK_WORKING_BUDGET_MS), converted to the packer's weight units via the
// LIVE table's own mean — platform-independent by construction, since the
// timings table is not sharded by OS. These tests pin partitionIsolatedFiles
// directly — the pure split, not the chunk-execution loop around it.
const {
CHUNK_WORKING_BUDGET_MS,
ISOLATION_BUDGET_FRACTION,
partitionIsolatedFiles,
} = require('../scripts/run-tests.cjs');
// Synthetic weigher/measured-predicate builders, keyed by basename, so these
// tests do not depend on the live tests/test-timings.json table.
function weigherFromMap(weights) {
return (f) => weights[f.split(/[\\/]/).pop()] ?? 1;
}
function measuredFromMap(weights) {
return (f) => Object.hasOwn(weights, f.split(/[\\/]/).pop());
}
describe('partitionIsolatedFiles (#4497 codex-config.test.cjs chunk isolation, derived #4733)', () => {
test('a file whose weight crosses thresholdWeight is isolated', () => {
const thresholdWeight = 9.834;
const weights = { 'a.test.cjs': 1, 'heavy.test.cjs': 9.834, 'b.test.cjs': 2 };
const files = ['/repo/tests/a.test.cjs', '/repo/tests/heavy.test.cjs', '/repo/tests/b.test.cjs'];
const { isolated, packable } = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight,
});
assert.deepStrictEqual(isolated, ['/repo/tests/heavy.test.cjs']);
assert.deepStrictEqual(packable, ['/repo/tests/a.test.cjs', '/repo/tests/b.test.cjs']);
});
test('a file just under the threshold is NOT isolated', () => {
const thresholdWeight = 9.834;
const weights = { 'a.test.cjs': 1, 'almost-heavy.test.cjs': 9.8, 'b.test.cjs': 2 };
const files = ['/repo/tests/a.test.cjs', '/repo/tests/almost-heavy.test.cjs', '/repo/tests/b.test.cjs'];
const { isolated, packable } = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight,
});
assert.deepStrictEqual(isolated, []);
assert.deepStrictEqual(packable, files);
});
// #4733 regression: changing thresholdWeight alone (no change to the
// file's own weight) must move a file across the isolation line — proves
// the split tracks the threshold parameter rather than a frozen boundary.
test('changing thresholdWeight alone moves a file across the isolation line', () => {
const weight = { 'borderline.test.cjs': 9.9 };
const files = ['/repo/tests/borderline.test.cjs'];
const atOldThreshold = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weight),
isMeasured: measuredFromMap(weight),
thresholdWeight: 17.88, // 9.9 stays packable
});
assert.deepStrictEqual(atOldThreshold, { isolated: [], packable: files });
const atNewThreshold = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weight),
isMeasured: measuredFromMap(weight),
thresholdWeight: 9.834, // 9.9 now crosses it
});
assert.deepStrictEqual(atNewThreshold, { isolated: files, packable: [] });
});
// #4733 (finding 6): eligibility for isolation is "is this file heavy?",
// not "which suite does it belong to" — the suite/unit restriction that
// used to live inside partitionIsolatedFiles is gone. Suite scoping still
// happens upstream in selectFiles before this function ever sees the list
// (verified: main() calls selectFiles(allFiles, suite) to build
// `selected`, then feeds `selected` into partitionIsolatedFiles — a
// suite='unit' run therefore never presents an install-suite file here at
// all), so a heavy install-suite file IS isolated when it reaches this
// function, e.g. on an 'all'/'install'-suite run.
test('a non-unit-suite file (foo.install.test.cjs) IS isolated when heavy, since suite scoping happens upstream', () => {
const weights = { 'foo.install.test.cjs': 1000 };
const files = ['/repo/tests/foo.install.test.cjs'];
const { isolated, packable } = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight: 1,
});
assert.deepStrictEqual(isolated, files);
assert.deepStrictEqual(packable, []);
});
test('an unmeasured file is not isolated on the strength of the fallback weight alone', () => {
// weightOf returns a large fallback (as makeFileWeigher's unknown-file
// fallback of 1 normalized weight would for a tiny thresholdWeight), but
// isMeasured reports false — isolation must refuse it regardless of what
// weightOf returns.
const files = ['/repo/tests/unknown.test.cjs'];
const { isolated, packable } = partitionIsolatedFiles(files, {
weightOf: () => 1000,
isMeasured: () => false,
thresholdWeight: 1,
});
assert.deepStrictEqual(isolated, []);
assert.deepStrictEqual(packable, files);
});
test('matches by BASENAME, so it isolates regardless of platform path separator or directory prefix', () => {
const weights = { 'heavy.test.cjs': 100 };
const files = [
'C:\\repo\\tests\\heavy.test.cjs',
'/repo/tests/subdir/heavy.test.cjs',
'heavy.test.cjs',
];
const { isolated, packable } = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight: 1,
});
assert.deepStrictEqual(isolated, files, 'every path ending in the heavy basename must be isolated, regardless of prefix/separator');
assert.deepStrictEqual(packable, []);
});
test('no heavy files present: everything is packable, order preserved', () => {
const weights = { 'z.test.cjs': 1, 'a.test.cjs': 1 };
const files = ['/repo/tests/z.test.cjs', '/repo/tests/a.test.cjs'];
const { isolated, packable } = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight: 22,
});
assert.deepStrictEqual(isolated, []);
assert.deepStrictEqual(packable, files);
});
test('an empty file list produces two empty buckets', () => {
assert.deepStrictEqual(
partitionIsolatedFiles([], { weightOf: () => 1, isMeasured: () => false, thresholdWeight: 22 }),
{ isolated: [], packable: [] },
);
});
// #4733 (finding 5): a non-finite/non-positive thresholdWeight would make
// every `>=` comparison false, silently disabling isolation with no error
// — partitionIsolatedFiles must throw instead of failing open.
for (const [label, bad] of [
['zero', 0],
['negative', -1],
['NaN', NaN],
['undefined', undefined],
]) {
test(`thresholdWeight=${label} throws instead of failing open`, () => {
assert.throws(
() => partitionIsolatedFiles(['/repo/tests/a.test.cjs'], {
weightOf: () => 100,
isMeasured: () => true,
thresholdWeight: bad,
}),
/thresholdWeight must be a finite, positive number/,
);
});
}
// ── Live-table rows (#4733) ─────────────────────────────────────────────
//
// These rows pin partitionIsolatedFiles against the REAL
// tests/test-timings.json table, computing thresholdWeight the same way
// main() does (CHUNK_WORKING_BUDGET_MS/ISOLATION_BUDGET_FRACTION over the
// table's own mean).
//
// Deliberate-mutation check performed while authoring this test (not
// committed, reverted after observing the failure): temporarily changing
// ISOLATION_BUDGET_FRACTION from 0.3 to 0.1 (thresholdWeight ~5.59 instead
// of ~16.78) made row 1 below fail with an extra file
// (workflow-fragments-emission.install.test.cjs, weight 15.93) present in
// the actual isolated set but absent from EXPECTED_ISOLATED_UNIT_FILES,
// and row 2 below fail because win32/linux/darwin no longer computed the
// same set as each other under the OLD (pre-fix) platform-scaled rule this
// row is guarding against — confirming both rows are load-bearing, not
// vacuous.
//
// #4249 split codex-config.test.cjs's heavy install()-pipeline blocks into
// codex-config-hooks.test.cjs, dropping codex-config.test.cjs's own
// measured weight from 127783ms to 189ms — well under this derived bar
// (0.3 * CHUNK_WORKING_BUDGET_MS ~= 120000ms). It correctly no longer
// appears in the live-computed isolated set, so it is dropped from this
// pinned expectation too, in place of being isolated by name.
const EXPECTED_ISOLATED_UNIT_FILES = [
'config.test.cjs',
'emitted-attribution.test.cjs',
'install-minimal-hooks.test.cjs',
'install.test.cjs',
'phase.test.cjs',
'run-tests-harness.test.cjs',
'state.test.cjs',
].sort();
function liveTableFixtures() {
const { suiteOf } = require('../scripts/lib/suite-detection.cjs');
const { makeFileWeigher, makeMeasuredPredicate } = require('../scripts/run-tests.cjs');
const table = require('../tests/test-timings.json');
const values = Object.values(table.timings).filter(
(v) => typeof v === 'number' && Number.isFinite(v) && v >= 0,
);
const mean = values.reduce((sum, v) => sum + v, 0) / values.length;
const timings = { timings: table.timings, mean };
const weightOf = makeFileWeigher(timings);
const isMeasured = makeMeasuredPredicate(timings);
const thresholdWeight = (ISOLATION_BUDGET_FRACTION * CHUNK_WORKING_BUDGET_MS) / mean;
// Production scoping: suite='unit' runs feed selectFiles-filtered
// (suiteOf(f) === null) files into partitionIsolatedFiles (see
// selectFiles in scripts/run-tests.cjs and its call site in main()).
const unitFiles = Object.keys(table.timings).filter((f) => suiteOf(f) === null);
return { weightOf, isMeasured, thresholdWeight, unitFiles };
}
test('#4603/#4733: the live-table unit-suite isolated set equals the historical 8-file set exactly', () => {
const { weightOf, isMeasured, thresholdWeight, unitFiles } = liveTableFixtures();
const { isolated } = partitionIsolatedFiles(unitFiles, { weightOf, isMeasured, thresholdWeight });
const basenames = isolated.map((f) => f.split(/[\\/]/).pop()).sort();
assert.deepStrictEqual(basenames, EXPECTED_ISOLATED_UNIT_FILES);
});
test('#4733: the live-table unit-suite isolated set is identical across win32, linux, darwin', () => {
// The threshold is computed once from the table mean and does not read
// process.platform anywhere in this derivation — this row is the
// regression guard for that: it would fail the instant the threshold (or
// the set it produces) becomes platform-dependent again, the way the
// MAX_FILES_PER_CHUNK-scaled ratio was.
const { weightOf, isMeasured, thresholdWeight, unitFiles } = liveTableFixtures();
const sets = ['win32', 'linux', 'darwin'].map(() => {
const { isolated } = partitionIsolatedFiles(unitFiles, { weightOf, isMeasured, thresholdWeight });
return isolated.map((f) => f.split(/[\\/]/).pop()).sort();
});
assert.deepStrictEqual(sets[0], EXPECTED_ISOLATED_UNIT_FILES);
assert.deepStrictEqual(sets[1], sets[0]);
assert.deepStrictEqual(sets[2], sets[0]);
});
test('#4733: dynamism — a file that crosses 120000ms in a synthetic table becomes isolated', () => {
// Synthetic table: 'was-light.test.cjs' previously measured well under
// the bar, now measured at 200000ms (above ISOLATION_BUDGET_FRACTION *
// CHUNK_WORKING_BUDGET_MS = 120000ms) — proving the threshold tracks the
// table, not a frozen list.
const timings = { 'was-light.test.cjs': 200000, 'other.test.cjs': 10000 };
const values = Object.values(timings);
const mean = values.reduce((a, b) => a + b, 0) / values.length;
const table = { timings, mean };
const { makeFileWeigher, makeMeasuredPredicate } = require('../scripts/run-tests.cjs');
const weightOf = makeFileWeigher(table);
const isMeasured = makeMeasuredPredicate(table);
const thresholdWeight = (ISOLATION_BUDGET_FRACTION * CHUNK_WORKING_BUDGET_MS) / mean;
const { isolated } = partitionIsolatedFiles(['was-light.test.cjs', 'other.test.cjs'], {
weightOf,
isMeasured,
thresholdWeight,
});
assert.deepStrictEqual(isolated, ['was-light.test.cjs']);
});
test('#4733: dynamism inverse — a file dropping below 120000ms is no longer isolated', () => {
const timings = { 'now-light.test.cjs': 50000, 'other.test.cjs': 10000 };
const values = Object.values(timings);
const mean = values.reduce((a, b) => a + b, 0) / values.length;
const table = { timings, mean };
const { makeFileWeigher, makeMeasuredPredicate } = require('../scripts/run-tests.cjs');
const weightOf = makeFileWeigher(table);
const isMeasured = makeMeasuredPredicate(table);
const thresholdWeight = (ISOLATION_BUDGET_FRACTION * CHUNK_WORKING_BUDGET_MS) / mean;
const { isolated } = partitionIsolatedFiles(['now-light.test.cjs', 'other.test.cjs'], {
weightOf,
isMeasured,
thresholdWeight,
});
assert.deepStrictEqual(isolated, []);
});
// Boundary triplet on the threshold itself. partitionIsolatedFiles uses
// `weightOf(f) >= thresholdWeight` (an inclusive `>=`), so `limit` (a file
// whose weight equals thresholdWeight exactly) IS isolated, not excluded —
// this triplet documents and pins that convention.
describe('boundary triplet on thresholdWeight (inclusive >=)', () => {
const thresholdWeight = 10;
const weights = { 'below.test.cjs': 9.999999, 'at.test.cjs': 10, 'above.test.cjs': 10.000001 };
const files = ['/repo/tests/below.test.cjs', '/repo/tests/at.test.cjs', '/repo/tests/above.test.cjs'];
test('limit-1 (just under thresholdWeight) is NOT isolated', () => {
const { isolated } = partitionIsolatedFiles(['/repo/tests/below.test.cjs'], {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight,
});
assert.deepStrictEqual(isolated, []);
});
test('limit (exactly thresholdWeight) IS isolated (inclusive >=)', () => {
const { isolated } = partitionIsolatedFiles(['/repo/tests/at.test.cjs'], {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight,
});
assert.deepStrictEqual(isolated, ['/repo/tests/at.test.cjs']);
});
test('limit+1 (just over thresholdWeight) IS isolated', () => {
const { isolated } = partitionIsolatedFiles(['/repo/tests/above.test.cjs'], {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight,
});
assert.deepStrictEqual(isolated, ['/repo/tests/above.test.cjs']);
});
test('all three together, in one call, preserve packable order', () => {
const { isolated, packable } = partitionIsolatedFiles(files, {
weightOf: weigherFromMap(weights),
isMeasured: measuredFromMap(weights),
thresholdWeight,
});
assert.deepStrictEqual(isolated, ['/repo/tests/at.test.cjs', '/repo/tests/above.test.cjs']);
assert.deepStrictEqual(packable, ['/repo/tests/below.test.cjs']);
});
});
});
// ---------------------------------------------------------------------------
// 2026-09-14 — the win32 per-chunk cap must not permit a chunk that exceeds
// the 600s backstop.
//
// Measured on `next` @ca8d9d4459: a Windows conformance chunk (shard 2/3,
// chunk 4/6, 29 files) was KILLED at 600018ms. On PR #4726 (green, larger
// pool, same shard/chunk position) the equivalent chunk (29 files) measured
// 525548ms — 87.6% of the then-current cap of 40, passing by only 74s.
// Worst packed-chunk rate: 525548 / 29 = 18122 ms/file.
// ---------------------------------------------------------------------------
const { defaultMaxFilesPerChunk } = require('../scripts/run-tests.cjs');
describe('the win32 per-chunk cap must not permit a chunk that exceeds the 600s backstop', () => {
const CHUNK_TIMEOUT_MS = 600000;
const MEASURED_WORST_MS_PER_FILE = 18122; // 525548ms / 29 files, windows conformance shard 2/3
const TARGET_MS = 400000; // 67% of the backstop
// NOT load-bearing on its own: this passes for ANY shipped cap up to 33
// (600000/18122 = 33.1), so it would NOT have caught the previously-shipped
// cap of 40. It is kept only as a coarse sanity check ("we are nowhere near
// the raw backstop"); the TARGET_MS (400000ms) rows below it are what
// actually pin the shipped value.
test('the win32 cap x the measured worst per-file rate stays under the raw 600000ms backstop (coarse, non-load-bearing sanity check)', () => {
const cap = defaultMaxFilesPerChunk('win32');
const product = cap * MEASURED_WORST_MS_PER_FILE;
assert.ok(
product < CHUNK_TIMEOUT_MS,
`win32 cap=${cap} x ${MEASURED_WORST_MS_PER_FILE}ms/file = ${product}ms, ` +
`must be < the ${CHUNK_TIMEOUT_MS}ms per-chunk backstop`,
);
});
test('the win32 cap leaves the intended headroom', () => {
const cap = defaultMaxFilesPerChunk('win32');
const product = cap * MEASURED_WORST_MS_PER_FILE;
assert.ok(
product <= TARGET_MS,
`win32 cap=${cap} x ${MEASURED_WORST_MS_PER_FILE}ms/file = ${product}ms, ` +
`must be <= the ${TARGET_MS}ms target (67% of the backstop)`,
);
});
// Boundary coverage (limit / limit+1) that actually constrains the
// EXPORTED function, not just arithmetic on literals: the invariant is
// that the shipped win32 cap must be the LARGEST value satisfying
// `cap * MEASURED_WORST_MS_PER_FILE <= TARGET_MS`. Both rows call
// defaultMaxFilesPerChunk('win32') so a change to the shipped constant
// moves both assertions with it — lowering the cap unnecessarily fails
// the maximality check below, raising it fails the safety check.
//
// A limit-1 (cap-1) row is intentionally omitted: `(cap-1) * RATE <=
// TARGET_MS` is implied by the maximality check already passing at `cap`
// (if cap clears the target, cap-1 trivially does too), so it cannot add
// coverage the other two rows don't already provide.
describe('the derived cap is the largest value that still clears the target', () => {
test('the shipped cap stays under the target (safety)', () => {
const cap = defaultMaxFilesPerChunk('win32');
const product = cap * MEASURED_WORST_MS_PER_FILE;
assert.ok(
product <= TARGET_MS,
`win32 cap=${cap} x ${MEASURED_WORST_MS_PER_FILE}ms/file = ${product}ms must be <= ${TARGET_MS}ms`,
);
});
test('one file over the shipped cap WOULD exceed the target (maximality)', () => {
const cap = defaultMaxFilesPerChunk('win32');
const product = (cap + 1) * MEASURED_WORST_MS_PER_FILE;
assert.ok(
product > TARGET_MS,
`win32 cap+1=${cap + 1} x ${MEASURED_WORST_MS_PER_FILE}ms/file = ${product}ms must be > ${TARGET_MS}ms ` +
`(if this fails, the shipped cap has slack and could safely be raised)`,
);
});
});
test('non-win32 platforms are unchanged', () => {
assert.strictEqual(defaultMaxFilesPerChunk('linux'), 60);
assert.strictEqual(defaultMaxFilesPerChunk('darwin'), 60);
});
test('the cap is still overridable by RUN_TESTS_MAX_FILES_PER_CHUNK', () => {
const cap = defaultMaxFilesPerChunk('win32');
const override = positiveNumberEnv('99', cap);
assert.strictEqual(override, 99);
assert.notStrictEqual(override, cap);
});
});