diff --git a/.github/workflows/mutation.yml b/.github/workflows/mutation.yml index e5103e34f..d5703e88f 100644 --- a/.github/workflows/mutation.yml +++ b/.github/workflows/mutation.yml @@ -110,7 +110,7 @@ jobs: runs-on: ubuntu-latest # Per-shard; lower than the old 30-min serial budget. Matrix-driven so only shards that # document a measured need (mutation-matrix.cjs COVERED[].timeoutMinutes, e.g. - # frontmatter's 180) get more; every other shard keeps the 15-minute default emitted by + # frontmatter's 60) get more; every other shard keeps the 15-minute default emitted by # buildResult() in scripts/mutation-matrix.cjs. timeout-minutes: ${{ matrix.timeoutMinutes }} strategy: @@ -133,7 +133,10 @@ jobs: run: npm ci - name: Run Stryker — ${{ matrix.name }} - # MUTATION_TEST_CMD scopes the command runner to only this module's tests. + # MUTATION_TEST_FILES supplies the per-shard tap.testFiles list to + # stryker.config.mjs (resolveMutationTestFiles). With coverageAnalysis: + # perTest, Stryker re-runs only the files that actually cover each mutant + # rather than the whole list per mutant (#3915: tap-runner swap). # --mutate scopes mutation to only the changed module's built artifact. # --incremental reuses cached results for unchanged mutants. # MUTATION_BREAK passes the per-module minScore ratchet floor to @@ -147,15 +150,14 @@ jobs: # repo's no-defer rule. env: NODE_OPTIONS: '--max-old-space-size=4096' - MUTATION_TEST_CMD: node --test --test-isolation=${{ matrix.isolation }} ${{ matrix.tests }} + MUTATION_TEST_FILES: ${{ matrix.tests }} MUTATION_BREAK: ${{ matrix.minScore }} MODULE_NAME: ${{ matrix.name }} MUTATE_GLOB: ${{ matrix.mutate }} - MODULE_TESTS: ${{ matrix.tests }} run: | echo "Module: ${MODULE_NAME}" echo "Mutate: ${MUTATE_GLOB}" - echo "Tests: ${MODULE_TESTS}" + echo "Tests: ${MUTATION_TEST_FILES}" echo "MinScore: ${MUTATION_BREAK}" npx stryker run --incremental --mutate "${MUTATE_GLOB}" diff --git a/CONTEXT.md b/CONTEXT.md index cbdddb2f6..1656943c8 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -609,6 +609,8 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr `RULESET.TESTS.clock-seam=concurrency logic must accept an optional {clock=Date} parameter; tests control time via t.mock.timers.enable(['Date']) + t.mock.timers.setTime(0) + t.mock.timers.tick(N); real OS scheduler races are not a permitted test pattern after ADR 456 (2026-05-28); real-race tests are deleted once deterministic seam tests cover the same logical path; clock.cjs realClock adds nowIso() (→ new Date(this.now()).toISOString()) and today() (→ nowIso().split('T')[0]) so all date-stamping in state.cjs routes through the seam; subprocess time-pin adapter: set GSD_TEST_MODE=1 + GSD_NOW_MS= in runGsdTools env to pin the date written by the SUT without touching real wall-clock (issue #474)` `RULESET.TESTS.property-based-testing=modules implementing parsing / transformation / budget-limit / bijective contracts must include at least one fast-check (fc) property test asserting a domain invariant; invariant categories: round-trip, monotonicity, boundary-containment, idempotency; property tests live in *.test.cjs alongside unit tests; CI signal: Stryker mutation score below 80% blocks merge` `RULESET.TESTS.mutation-score=Stryker runs incremental (--since origin/next) on ubuntu-latest/Node24 CI leg; default threshold 80% killed/total; surviving mutants in scope block merge unless path is listed in stryker.config.mjs with documented reason; treat surviving mutant as a failing test specification` +`RULESET.TESTS.mutation-runner=Stryker executes every shard through the OFFICIAL @stryker-mutator/tap-runner (testRunner:'tap'), never the built-in 'command' runner (#3915); 'command' is the one runner Stryker excludes from coverage analysis, which forced coverageAnalysis:'off' and made cost strictly linear in (mutants x whole-shard test time) — the frontmatter shard measured 1751s on run 33021042847 vs 212s for the next slowest. tap.testFiles is injected per shard via MUTATION_TEST_FILES (mutation.yml env <- matrix.tests <- scripts/mutation-matrix.cjs buildResult); resolveMutationTestFiles is the SINGLE fail-closed reader and existence-checks every entry, because the tap runner's findTestyLookingFiles resolves the list with glob() and a non-matching pattern yields an EMPTY list SILENTLY (a fast, confident, meaningless run). tap.forceBail is FALSE by measurement, not preference: 3 of 26 shard test files spawn subprocesses (config-schema.property, core-utils, feat-3881-yaml-parser-consequences) and bail fires on every KILLED mutant, so leaving it on kills processes mid-spawnSync and orphans their children; Stryker's separate disableBail still skips remaining FILES, which is most of the win. tap.nodeArgs and top-level buildCommand stay UNSET so no rebuild lands between mutation and test (ADR-457). Coverage granularity is per FILE, not per test ("a test is always a test file"), so the #2790 excludeTests bans on spawn-heavy integration files remain necessary and unchanged` +`RULESET.TESTS.mutation-score-denominator=the gated number is mutation-testing-metrics' mutationScore = totalDetected/totalValid, which counts NoCoverage in the denominator EXACTLY as Survived; both Stryker's own thresholds.break (core dist/src/reporters/mutation-test-report-helper.js) and scripts/check-mutation-score-ratchet.cjs read THAT field, which is what makes the #3915 coverageAnalysis 'off'->'perTest' switch score-neutral. NEVER gate on mutationScoreBasedOnCoveredCode — it EXCLUDES NoCoverage and inflates sharply under perTest (measured on a synthetic report: 8 killed/2 survived = 80 and 80; 8 killed/2 noCoverage = 80 and 100), so swapping to the better-sounding field would make every minScore floor trivially satisfiable and the gate decorative. Under the pre-#3915 coverageAnalysis:'off' the two fields were ALWAYS identical (noCoverage was structurally 0), which is why nothing had ever pinned the choice; tests/mutation-score-ratchet.test.cjs now pins it with a non-vacuity assertion that the two numbers genuinely diverge` `RULESET.TESTS.delete-bad-tests=pass-always / vacuous-truth / source-grep / elapsed-time / real-race / permanent-allow-test-rule tests are DELETED and replaced with compliant tests in the same PR; not skipped, not commented out, not permanently exempted; replacement must cover the same logical path via typed-surface assertion or clock-seam pattern` `RULESET.TESTS.eslint-harness=ADR 452 (2026-05-28): ESLint flat config + typescript-eslint + eslint-plugin-n + eslint-plugin-no-only-tests + local plugin at eslint-rules/ (repo root, NOT scripts/eslint-rules/); replaces scripts/lint-*.cjs regex scanners (fully removed in #632); all three test-rigor rules now ship at error in tests/**/*.test.cjs scope: local/no-source-grep and local/no-magic-sleep-in-tests promoted by #3313, local/no-elapsed-assertion promoted by #3331 once #3314 delivered its ADR-456 §(a) precondition (epic #1885 was subsumed into epic #3053 and closed stale before this promotion landed)` diff --git a/docs/CONTEXT-INDEX.json b/docs/CONTEXT-INDEX.json index 20492f1c2..312dc2383 100644 --- a/docs/CONTEXT-INDEX.json +++ b/docs/CONTEXT-INDEX.json @@ -1,6 +1,6 @@ { "schemaVersion": 1, - "count": 261, + "count": 263, "classes": { "ARCH": 1, "CI": 2, @@ -17,7 +17,7 @@ "PROC": 14, "PROHIB": 10, "RELEASE-NOTES": 31, - "RULESET": 58, + "RULESET": 60, "SESSION": 9, "WAVE": 5, "WORKSTREAM": 5, @@ -1094,11 +1094,21 @@ "klass": "RULESET", "value": "module-level const src = readFileSync(...) throws before any test() registers — wrap in try/catch in test() or use lazy load" }, + { + "id": "RULESET.TESTS.mutation-runner", + "klass": "RULESET", + "value": "Stryker executes every shard through the OFFICIAL @stryker-mutator/tap-runner (testRunner:'tap'), never the built-in 'command' runner (#3915); 'command' is the one runner Stryker excludes from coverage analysis, which forced coverageAnalysis:'off' and made cost strictly linear in (mutants x whole-shard test time) — the frontmatter shard measured 1751s on run 33021042847 vs 212s for the next slowest. tap.testFiles is injected per shard via MUTATION_TEST_FILES (mutation.yml env <- matrix.tests <- scripts/mutation-matrix.cjs buildResult); resolveMutationTestFiles is the SINGLE fail-closed reader and existence-checks every entry, because the tap runner's findTestyLookingFiles resolves the list with glob() and a non-matching pattern yields an EMPTY list SILENTLY (a fast, confident, meaningless run). tap.forceBail is FALSE by measurement, not preference: 3 of 26 shard test files spawn subprocesses (config-schema.property, core-utils, feat-3881-yaml-parser-consequences) and bail fires on every KILLED mutant, so leaving it on kills processes mid-spawnSync and orphans their children; Stryker's separate disableBail still skips remaining FILES, which is most of the win. tap.nodeArgs and top-level buildCommand stay UNSET so no rebuild lands between mutation and test (ADR-457). Coverage granularity is per FILE, not per test (\"a test is always a test file\"), so the #2790 excludeTests bans on spawn-heavy integration files remain necessary and unchanged" + }, { "id": "RULESET.TESTS.mutation-score", "klass": "RULESET", "value": "Stryker runs incremental (--since origin/next) on ubuntu-latest/Node24 CI leg; default threshold 80% killed/total; surviving mutants in scope block merge unless path is listed in stryker.config.mjs with documented reason; treat surviving mutant as a failing test specification" }, + { + "id": "RULESET.TESTS.mutation-score-denominator", + "klass": "RULESET", + "value": "the gated number is mutation-testing-metrics' mutationScore = totalDetected/totalValid, which counts NoCoverage in the denominator EXACTLY as Survived; both Stryker's own thresholds.break (core dist/src/reporters/mutation-test-report-helper.js) and scripts/check-mutation-score-ratchet.cjs read THAT field, which is what makes the #3915 coverageAnalysis 'off'->'perTest' switch score-neutral. NEVER gate on mutationScoreBasedOnCoveredCode — it EXCLUDES NoCoverage and inflates sharply under perTest (measured on a synthetic report: 8 killed/2 survived = 80 and 80; 8 killed/2 noCoverage = 80 and 100), so swapping to the better-sounding field would make every minScore floor trivially satisfiable and the gate decorative. Under the pre-#3915 coverageAnalysis:'off' the two fields were ALWAYS identical (noCoverage was structurally 0), which is why nothing had ever pinned the choice; tests/mutation-score-ratchet.test.cjs now pins it with a non-vacuity assertion that the two numbers genuinely diverge" + }, { "id": "RULESET.TESTS.no-dead-regex-in-includes", "klass": "RULESET", diff --git a/examples/dynamic-context-management/CONTEXT-INDEX.json b/examples/dynamic-context-management/CONTEXT-INDEX.json index c2350e929..16c299898 100644 --- a/examples/dynamic-context-management/CONTEXT-INDEX.json +++ b/examples/dynamic-context-management/CONTEXT-INDEX.json @@ -1,6 +1,6 @@ { "schemaVersion": 1, - "count": 261, + "count": 263, "classes": { "ARCH": 1, "CI": 2, @@ -17,7 +17,7 @@ "PROC": 14, "PROHIB": 10, "RELEASE-NOTES": 31, - "RULESET": 58, + "RULESET": 60, "SESSION": 9, "WAVE": 5, "WORKSTREAM": 5, @@ -28,1567 +28,1579 @@ "id": "ARCH.SKILL.improve-codebase.next-candidates", "klass": "ARCH", "value": "[Workstream Progress Projection Module]", - "line": 667 + "line": 677 }, { "id": "CI.GATE.changeset-lint", "klass": "CI", "value": "hard-fail for user-facing code diffs unless .changeset/* or PR has no-changelog label", - "line": 651 + "line": 661 }, { "id": "CI.GATE.issue-link-required", "klass": "CI", "value": "hard-fail if PR body lacks closes/fixes/resolves #", - "line": 650 + "line": 660 }, { "id": "CONFIG.LOCATION.SEAM.in-process-scrub", "klass": "CONFIG", "value": "TEST_ENV_BASE reaches CHILD env only; a test calling install() IN-PROCESS must additionally use helpers.scrubConfigLocationEnv() in beforeEach + its restorer in afterEach — HOME/USERPROFILE sandboxing is NOT sufficient because getGlobalConfigDir is env-FIRST", - "line": 685 + "line": 695 }, { "id": "CONFIG.LOCATION.SEAM.kimi-two-homes", "klass": "CONFIG", "value": "kimi declares TWO config-location vars: KIMI_CONFIG_DIR (registry, generic Agent-Skills root via resolveKimiGlobalDir) and KIMI_SHARE_DIR (KIMI_HOOKS_TOML_DESCRIPTOR, kimi's OWN native config.toml carrying GSD's [[hooks]] block via resolveKimiHooksTomlDir); a registry-only derivation covers the first and silently misses the second", - "line": 684 + "line": 694 }, { "id": "CONFIG.LOCATION.SEAM.scrub-set", "klass": "CONFIG", "value": "tests/helpers.cjs CONFIG_LOCATION_ENV_KEYS is DERIVED from five sources rather than maintained as one hand-written list (source 4 IS a literal residue list, for vars that fit no other rung — what is never hand-listed is the SET): capability-registry runtimes[].runtime.configHome.env AND [].configHome.skillsHome.env + runtime-homes NON_REGISTRY_CONFIG_HOME_DESCRIPTORS[].env AND [].skillsHome.env (a descriptor is a descriptor — BOTH descriptor rungs walk skillsHome, which resolves independently via resolveSkillsBaseFromDescriptor) + runtime-homes GSD_LOCATION_ENV_KEYS + a residue list (GROK_AGENTS_HOME, GSD_RUNTIME, GSD_PROJECT, GSD_WORKSTREAM) + WRITE_ESCAPE_PERMISSION_ENV_KEYS (GSD_ALLOW_SYMLINKED_DEST — a permission, not a location: it names no path but disarms the symlink-escape guard, so blanking it makes the guard STRICTER, never looser); adding a config-location var means making it ENUMERABLE at one of those sources, not appending a literal", - "line": 682 + "line": 692 }, { "id": "CONFIG.LOCATION.SEAM.two-families", "klass": "CONFIG", "value": "runtime configHomes (where a third-party runtime keeps config, registry- or descriptor-declared) and GSD's OWN location vars (GSD_HOME -> $GSD_HOME/.gsd store, GSD_AGENTS_DIR -> getAgentsDir priority 1) are DISTINCT families; no registry derivation reaches the second, and treating a miss there as a registry gap is what produced review round 2", - "line": 683 + "line": 693 }, { "id": "CONFIG.SEAM.loadConfig-context", "klass": "CONFIG", "value": "loadConfig(cwd,{workstream}) replaces env-mutation fallback; no temporary process.env GSD_WORKSTREAM rewrites", - "line": 681 + "line": 691 }, { "id": "EXEC.CLASSIFY.classes", "klass": "EXEC", "value": "{class:'quota-exceeded'|'classify-handoff-bug'|'unknown-failure', sentinel?, retryAfterSeconds?}", - "line": 900 + "line": 910 }, { "id": "EXEC.CLASSIFY.cross-runtime", "klass": "EXEC", "value": "Anthropic/CC: usage limit|rate limit|quota|429|retry-after; Copilot CLI: rate_limit (stem); Codex CLI: 429|usage_limit_reached|too many requests", - "line": 902 + "line": 912 }, { "id": "EXEC.CLASSIFY.handler", "klass": "EXEC", "value": "gsd-core/bin/lib/agent-command-router.cjs:classifyAgentFailure (registered via command-aliases.cjs; mutation:false outputMode:json)", - "line": 898 + "line": 908 }, { "id": "EXEC.CLASSIFY.precedence", "klass": "EXEC", "value": "quota sentinel wins over classifyHandoffIfNeeded bug when both appear", - "line": 903 + "line": 913 }, { "id": "EXEC.CLASSIFY.proactive-signal-not-usable", "klass": "EXEC", "value": "Anthropic exposes anthropic-ratelimit-* headers + Agent SDK RateLimitEvent; Claude Code subprocess does NOT forward to hooks/statusline today (upstream #33820, #22407, #32796)", - "line": 905 + "line": 915 }, { "id": "EXEC.CLASSIFY.retry-after-parser", "klass": "EXEC", "value": "\\bretry[-_ ]after[:\\s]+(\\d+)\\b avoids embedded-word false matches like noretry-after", - "line": 904 + "line": 914 }, { "id": "EXEC.CLASSIFY.sentinel-order", "klass": "EXEC", "value": "most specific first: 429 beats too-many-requests; resource_exhausted beats quota (array order in src/agent-command-router.cts QUOTA_SENTINELS checks resource_exhausted before quota); case-insensitive; canonical sentinel value is lower-cased form", - "line": 901 + "line": 911 }, { "id": "EXEC.CLASSIFY.workflow", "klass": "EXEC", "value": "gsd-core/workflows/execute-phase.md step 7; class-distinct prompts (quota-to-wait-for-reset; classify-handoff-bug-to-spot-check; unknown-to-continue/stop)", - "line": 899 + "line": 909 }, { "id": "GSD-RESEARCH.CONTEXT-DISCIPLINE", "klass": "GSD-RESEARCH", "value": "less-context levers: subagent isolation + compact provider output + fetches-to-disk + cache-returns-digest; API clear_tool_uses/memory tool are the conceptual model, not a Claude Code harness knob", - "line": 438 + "line": 446 }, { "id": "GSD-RESEARCH.INTEGRATION.L2-hybrid", "klass": "GSD-RESEARCH", "value": "code owns cache+legitimacy+confidence+provider-pick (gsd-tools query research-plan/research-store/package-legitimacy); MCP owns the fetch; agent returns RESEARCH.md path, never raw fetches", - "line": 436 + "line": 444 }, { "id": "GSD-RESEARCH.MODULE.package-legitimacy", "klass": "GSD-RESEARCH", "value": "registry-API verdicts (npm/PyPI/crates.io injectable adapters) computed from thresholds {minAgeDays:30,minWeeklyDownloads:1000,requireRepo:true}; verdict OK|SUS|SLOP per package; slopcheck=optional adapter that can only escalate, never the install-or-degrade gate", - "line": 435 + "line": 443 }, { "id": "GSD-RESEARCH.MODULE.research-provider", "klass": "GSD-RESEARCH", "value": "single source of truth PROVIDER_WATERFALL (docs Context7->Ref->Jina->websearch; web Exa->Tavily->Perplexity->Brave->websearch; scrape Firecrawl->Jina); planResearch returns cache-hits+fetch-plan; classifyConfidence stamps HIGH|MEDIUM|LOW by provider AUTHORITY + verification EVIDENCE (HIGH requires code-computed ground-truth corroboration e.g. legitimacyVerdict OK; provider authority alone caps at MEDIUM; SLOP caps at LOW); Firecrawl is scrape-only (not in the docs or web legs)", - "line": 434 + "line": 442 }, { "id": "GSD-RESEARCH.MODULE.research-store", "klass": "GSD-RESEARCH", "value": "content-addressed cache; key=sha256(ecosystem+library+version+query+kind); getResearch->{hit,stale} never throws (mirrors graphify staleness); ttlForSource curated HIGH 30d|MED 7d|web LOW 1d; tiers: curated-doc kinds -> ~/.gsd/research-cache (cross-project), web/synthesis -> project .planning/research/.cache", - "line": 433 + "line": 441 }, { "id": "GSD-RESEARCH.PROVIDER.availability", "klass": "GSD-RESEARCH", "value": "config flags brave_search/exa_search/firecrawl/tavily_search/ref_search/perplexity/jina (env _API_KEY or ~/.gsd/_api_key); context7/jina/websearch always available; planResearch falls through waterfall to websearch terminal", - "line": 437 + "line": 445 }, { "id": "LEARNING.prompt-budget.boundary-gap", "klass": "LEARNING", "value": "PR #3708 commit 2df566ed reserved NOTE_RESERVE_TOKENS in pressure-threshold AND in minSet pre-check; both buggy paths only fire when baseTokens ∈ (effectiveBudget - NOTE_RESERVE_TOKENS, effectiveBudget]; original test suite used budgets far from that band so neither path was exercised; fix bde1ae8f confines NOTE_RESERVE accounting to post-trim assembly path only; future budget/limit code MUST add boundary fixtures per RULESET.TESTS.boundary-coverage.fixtures", - "line": 598 + "line": 606 }, { "id": "LIVE-CONFIG.GUARD.SEAM.ci-blind", "klass": "LIVE-CONFIG", "value": "the AMBIENT-ENV half stays CI-blind — CI never has these vars set, so green CI is not evidence for it; what strict mode catches in CI is the suite's own default-root leaks (HOME/USERPROFILE-derived), the guard remains the only loud signal for ambient-var escapes", - "line": 691 + "line": 701 }, { "id": "LIVE-CONFIG.GUARD.SEAM.module", "klass": "LIVE-CONFIG", "value": "scripts/live-config-guard.cjs (deliberately NOT scripts/lib/, which the installer copies to users wholesale while uninstall removes only an allowlist; excluded from the npm tarball via package.json files[] together with its whole require chain run-tests.cjs/affected-tests-lib.cjs/run-affected-tests.cjs — a partial exclusion trips the #2858 shipped-requires-only-shipped gate); exports [resolveLiveConfigRoots, resolveExtraWatchTargets, snapshotLiveConfig, diffLiveConfig, formatViolations, newestMtime]; driven by scripts/run-tests.cjs pre/post suite", - "line": 686 + "line": 696 }, { "id": "LIVE-CONFIG.GUARD.SEAM.non-root-targets", "klass": "LIVE-CONFIG", "value": "resolveExtraWatchTargets covers THREE live write surfaces that are not runtime config ROOTS (skills bases are a DELIBERATE non-target — the config-root layout misfires beneath them, so they need their own layout): $GSD_HOME/.gsd watched WHOLESALE (exclusively GSD-owned, so the shared-root trap does not apply) plus ONE config.toml per NON_REGISTRY_CONFIG_HOME_DESCRIPTORS entry, each watched as a SINGLE FILE (those roots belong to their products) — today three targets, since #2755 split Kimi CLI (~/.kimi, KIMI_SHARE_DIR) from Kimi Code (~/.kimi-code, KIMI_CODE_HOME); the targets are DERIVED by iterating that array, never by calling a named resolver, so a further descriptor is picked up without editing the guard PROVIDED it owns the same NON_REGISTRY_OWNED_FILE ('config.toml') — one that owns a different filename needs a per-descriptor mapping, the named residual the guard states at its own definition. SECOND RESIDUAL: config.toml is not all GSD writes into those roots — installSharedHooksBundle also populates /hooks/, which is UNWATCHED; closing it is a layout decision, like skills bases; passed to snapshotLiveConfig explicitly so a fixture-root caller cannot pull the real ~/.gsd into its snapshot", - "line": 688 + "line": 698 }, { "id": "LIVE-CONFIG.GUARD.SEAM.scope", "klass": "LIVE-CONFIG", "value": "ownership-based, never whole-root: GSD_OWNED_ENTRIES top-level footprint + children whose name startsWith GSD_ARTIFACT_PREFIX ('gsd-') under GSD_PREFIXED_PARENTS (dirs shared with the host agent); watching a shared root wholesale false-positives on the host's own writes and a guard that cries wolf gets disabled", - "line": 687 + "line": 697 }, { "id": "LIVE-CONFIG.GUARD.SEAM.severity", "klass": "LIVE-CONFIG", "value": "reports by default locally; CI wires GSD_STRICT_LIVE_CONFIG_GUARD=1 on Linux/macOS lanes (test.yml, all three test jobs) so a suite-produced leak FAILS those runs; Windows lanes stay report-only pending the documented pre-existing USERPROFILE sweep (~190 test sites sandbox HOME alone) — promote once that lands; skipped by GSD_SKIP_LIVE_CONFIG_GUARD=1", - "line": 690 + "line": 700 }, { "id": "LIVE-CONFIG.GUARD.SEAM.truncation", "klass": "LIVE-CONFIG", "value": "MAX_ENTRIES/MAX_DEPTH bound the walk; a bound hit sets truncated and diffLiveConfig emits kind:'unverified' — a truncated scan MUST NOT read as clean; boundary covered at {limit-1,limit,limit+1} via newestMtime's injected budget plus fast-check monotonicity, per RULESET.TESTS.boundary-coverage + RULESET.TESTS.property-based-testing", - "line": 689 + "line": 699 }, { "id": "META.RULE.brief-must-cite-doc", "klass": "META", "value": "agent prompts MUST quote the canonical doc line being applied; paraphrasing from predicate memory drifts and produces violations", - "line": 742 + "line": 752 }, { "id": "META.RULE.brief-no-paraphrase", "klass": "META", "value": "writing \"k040 — never leave changelog box unchecked\" caused 5 of 8 agents to edit CHANGELOG.md in violation of CONTRIBUTING.md L110", - "line": 743 + "line": 753 }, { "id": "META.RULE.canonical-source-precedence", "klass": "META", "value": "CONTRIBUTING.md > docs/adr/* > CONTEXT.md > agent memory", - "line": 740 + "line": 750 }, { "id": "META.RULE.read-contributing-first", "klass": "META", "value": "read CONTRIBUTING.md sections \"Pull Request Guidelines\" + \"CHANGELOG Entries\" before EVERY agent dispatch", - "line": 741 + "line": 751 }, { "id": "PLANNING.PATH.PARITY.project-scope", "klass": "PLANNING", "value": ".planning/ (never .planning/projects/); mirror planning-workspace.cjs planningDir()", - "line": 676 + "line": 686 }, { "id": "PLANNING.PATH.SEAM.helpers", "klass": "PLANNING", "value": "helpers.planningPaths delegates to workspacePlanningPaths + resolveWorkspaceContext; precedence explicit-ws > env-ws > env-project > root", - "line": 677 + "line": 687 }, { "id": "PLANNING.PATH.SEAM.init-handlers", "klass": "PLANNING", "value": "[initExecutePhase, initPlanPhase, initPhaseOp, initMilestoneOp] consume helpers.planningPaths().planning (no direct relPlanningPath join)", - "line": 678 + "line": 688 }, { "id": "PR.3267.POSTMORTEM.recovery", "klass": "PR", "value": "[issue#3270 created, label approved-enhancement applied, PR reopened, body includes \"Closes #3270\", label no-changelog applied]", - "line": 655 + "line": 665 }, { "id": "PR.3267.POSTMORTEM.root-cause", "klass": "PR", "value": "[missing issue link, missing changeset/no-changelog]", - "line": 654 + "line": 664 }, { "id": "PRED.k320.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L193-211", - "line": 746 + "line": 756 }, { "id": "PRED.k320.ci-enforcement", "klass": "PRED", "value": "scripts/changeset/lint.cjs", - "line": 752 + "line": 762 }, { "id": "PRED.k320.ci-paths-monitored", "klass": "PRED", "value": "bin/ gsd-core/ src/ agents/ commands/ hooks/ sdk/src/ sdk/prompts/", - "line": 753 + "line": 763 }, { "id": "PRED.k320.cure", "klass": "PRED", "value": "drop .changeset/--.md fragment ONLY", - "line": 748 + "line": 758 }, { "id": "PRED.k320.evidence", "klass": "PRED", "value": "PR #3302 merge-conflict against #3308 CHANGELOG.md row 2026-05-09", - "line": 755 + "line": 765 }, { "id": "PRED.k320.opt-out-label", "klass": "PRED", "value": "no-changelog", - "line": 751 + "line": 761 }, { "id": "PRED.k320.recovery", "klass": "PRED", "value": "open Removed-typed cleanup PR deleting only the redundant row", - "line": 754 + "line": 764 }, { "id": "PRED.k320.rule", "klass": "PRED", "value": "do not edit CHANGELOG.md in feature/fix/enhancement PRs", - "line": 747 + "line": 757 }, { "id": "PRED.k320.signal", "klass": "PRED", "value": "changelog-direct-edit-forbidden", - "line": 745 + "line": 755 }, { "id": "PRED.k320.tool", "klass": "PRED", "value": "npm run changeset -- --type --pr --body \"...\"", - "line": 749 + "line": 759 }, { "id": "PRED.k320.types", "klass": "PRED", "value": "Added|Changed|Deprecated|Removed|Fixed|Security", - "line": 750 + "line": 760 }, { "id": "PRED.k321.evidence", "klass": "PRED", "value": "PRs #3304/#3305 (2026-05-09): real Minor/Major findings in body, 0 threads", - "line": 761 + "line": 771 }, { "id": "PRED.k321.poll-shape", "klass": "PRED", "value": "parse pulls//reviews body AND graphql reviewThreads", - "line": 759 + "line": 769 }, { "id": "PRED.k321.resolution", "klass": "PRED", "value": "address in code; no GraphQL resolveReviewThread needed for body-only findings", - "line": 760 + "line": 770 }, { "id": "PRED.k321.shape", "klass": "PRED", "value": "CR posts \"[!CAUTION] outside the diff\" findings in review BODY, not in reviewThreads", - "line": 758 + "line": 768 }, { "id": "PRED.k321.signal", "klass": "PRED", "value": "cr-outside-diff-range-finding", - "line": 757 + "line": 767 }, { "id": "PRED.k322.cure-1", "klass": "PRED", "value": "2nd retrigger ~10min after first ack", - "line": 766 + "line": 776 }, { "id": "PRED.k322.cure-2", "klass": "PRED", "value": "if silent at 50min, treat as silent-pass with maintainer flag in merge-commit body", - "line": 767 + "line": 777 }, { "id": "PRED.k322.distinct-from", "klass": "PRED", "value": "k080", - "line": 764 + "line": 774 }, { "id": "PRED.k322.evidence", "klass": "PRED", "value": "PR #3306 (2026-05-09): 0 reviews after 50min + 2 retriggers", - "line": 769 + "line": 779 }, { "id": "PRED.k322.merge-gate-impact", "klass": "PRED", "value": "k070 real_coderabbit_review_present unsatisfied; requires maintainer judgment", - "line": 768 + "line": 778 }, { "id": "PRED.k322.shape", "klass": "PRED", "value": "ack posted, real review never lands within [5s, 410s] cooldown after burst of N PRs <15min", - "line": 765 + "line": 775 }, { "id": "PRED.k322.signal", "klass": "PRED", "value": "cr-sustained-throttle", - "line": 763 + "line": 773 }, { "id": "PRED.k323.cure-alt", "klass": "PRED", "value": "consolidate into single PR when 2+ issues share root cause", - "line": 774 + "line": 784 }, { "id": "PRED.k323.cure-pre-dispatch", "klass": "PRED", "value": "brief one agent canonical-owner; brief others to EXCLUDE shared site", - "line": 773 + "line": 783 }, { "id": "PRED.k323.evidence", "klass": "PRED", "value": "#3300 (#3297) overlapped #3306 (#3298) on add-backlog.md hunks 2026-05-09", - "line": 776 + "line": 786 }, { "id": "PRED.k323.recovery", "klass": "PRED", "value": "close smaller PR as \"subsumed by #N\" or rebase second to drop overlap hunk", - "line": 775 + "line": 785 }, { "id": "PRED.k323.shape", "klass": "PRED", "value": "2+ open issues touch same canonical bug site; each fix's sibling-audit produces overlapping diff", - "line": 772 + "line": 782 }, { "id": "PRED.k323.signal", "klass": "PRED", "value": "sibling-audit-cross-pr-overlap", - "line": 771 + "line": 781 }, { "id": "PRED.k324.cure", "klass": "PRED", "value": "verify via gh api on every agent-completion notification; never trust narrative", - "line": 780 + "line": 790 }, { "id": "PRED.k324.evidence", "klass": "PRED", "value": "2026-05-09 session: 5+ mid-monitor terminations across PRs #3232/#3271/#3251/#3255/#3262", - "line": 782 + "line": 792 }, { "id": "PRED.k324.k095-restatement", "klass": "PRED", "value": "k095 confirmed shape: agent reports \"waiting for monitor\" / \"tests still running\" then terminates", - "line": 779 + "line": 789 }, { "id": "PRED.k324.poll-shape", "klass": "PRED", "value": "gh pr view --json mergeStateStatus,statusCheckRollup + pulls//reviews + graphql reviewThreads + issues//comments tail", - "line": 781 + "line": 791 }, { "id": "PRED.k324.signal", "klass": "PRED", "value": "agent-terminates-mid-monitor", - "line": 778 + "line": 788 }, { "id": "PRED.k325.cleanup", "klass": "PRED", "value": "git worktree remove --force for aged agent worktrees", - "line": 787 + "line": 797 }, { "id": "PRED.k325.cure", "klass": "PRED", "value": "detached-HEAD: git checkout --detach $(git ls-remote origin ); modify; commit; git push --force-with-lease=: origin HEAD:refs/heads/", - "line": 786 + "line": 796 }, { "id": "PRED.k325.evidence", "klass": "PRED", "value": "2026-05-09 CHANGELOG.md strip on PRs #3300/#3302/#3304/#3305 required detached-HEAD", - "line": 788 + "line": 798 }, { "id": "PRED.k325.shape", "klass": "PRED", "value": "git checkout errors \"already used by worktree at \"", - "line": 785 + "line": 795 }, { "id": "PRED.k325.signal", "klass": "PRED", "value": "worktree-branch-lock-on-force-push", - "line": 784 + "line": 794 }, { "id": "PRED.k326.cure", "klass": "PRED", "value": "quote canonical doc verbatim in brief; mentally simulate \"if all N agents follow this brief literally, do they violate any rule?\"", - "line": 792 + "line": 802 }, { "id": "PRED.k326.evidence", "klass": "PRED", "value": "2026-05-09 brief \"k040 — update CHANGELOG.md\" → 5 of 8 agents violated CONTRIBUTING.md L110", - "line": 793 + "line": 803 }, { "id": "PRED.k326.shape", "klass": "PRED", "value": "N parallel agents amplify a single brief-vs-doc contradiction into N violations", - "line": 791 + "line": 801 }, { "id": "PRED.k326.signal", "klass": "PRED", "value": "brief-contradicts-canonical-doc", - "line": 790 + "line": 800 }, { "id": "PRED.k327.ack-shape", "klass": "PRED", "value": "body \"✅ Actions performed - Full review triggered\"", - "line": 796 + "line": 806 }, { "id": "PRED.k327.cooldown-normal", "klass": "PRED", "value": "[5s, 410s]", - "line": 799 + "line": 809 }, { "id": "PRED.k327.cooldown-throttled", "klass": "PRED", "value": "k322", - "line": 800 + "line": 810 }, { "id": "PRED.k327.distinguish-key", "klass": "PRED", "value": "len(pulls//reviews) — ack=0, real=≥1", - "line": 798 + "line": 808 }, { "id": "PRED.k327.real-review-shape", "klass": "PRED", "value": "body starts \"Actionable comments posted: N\" OR \"[!CAUTION] Some comments are outside the diff\"", - "line": 797 + "line": 807 }, { "id": "PRED.k327.signal", "klass": "PRED", "value": "cr-ack-vs-real-review", - "line": 795 + "line": 805 }, { "id": "PRED.k328.audit-list", "klass": "PRED", "value": "[heading-matches-class, closing-keyword-present, changeset-fragment-or-no-changelog-label]", - "line": 805 + "line": 815 }, { "id": "PRED.k328.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L48,L64,L81 (template links) + .github/PULL_REQUEST_TEMPLATE/{fix,enhancement,feature}.md L1 (heading text)", - "line": 803 + "line": 813 }, { "id": "PRED.k328.k100-restatement", "klass": "PRED", "value": "heading must match issue class: bug→## Fix PR, enhancement→## Enhancement PR, feature→## Feature PR", - "line": 804 + "line": 814 }, { "id": "PRED.k328.signal", "klass": "PRED", "value": "pr-template-typed-heading-required", - "line": 802 + "line": 812 }, { "id": "PRED.k329.body", "klass": "PRED", "value": "**** — . (#)", - "line": 811 + "line": 821 }, { "id": "PRED.k329.canonical-source", "klass": "PRED", "value": "CONTRIBUTING.md L196-202 + .changeset/README.md", - "line": 808 + "line": 818 }, { "id": "PRED.k329.filename", "klass": "PRED", "value": ".changeset/--.md", - "line": 809 + "line": 819 }, { "id": "PRED.k329.frontmatter", "klass": "PRED", "value": "---\\\\ntype: \\\\npr: \\\\n---", - "line": 810 + "line": 820 }, { "id": "PRED.k329.observed-clean", "klass": "PRED", "value": "#3299 sunny-ibex-wave, #3301 sturdy-rams-caper, #3306 3298-phase-dir-prefix-drift-workflows", - "line": 812 + "line": 822 }, { "id": "PRED.k329.signal", "klass": "PRED", "value": "changeset-fragment-canonical-shape", - "line": 807 + "line": 817 }, { "id": "PRED.k330.fallback", "klass": "PRED", "value": "append predicate-format findings directly to CONTEXT.md", - "line": 816 + "line": 826 }, { "id": "PRED.k330.shape", "klass": "PRED", "value": "mempalace MCP tools require explicit user call; AI cannot trigger", - "line": 815 + "line": 825 }, { "id": "PRED.k330.signal", "klass": "PRED", "value": "mempalace-diary-not-callable-by-ai", - "line": 814 + "line": 824 }, { "id": "PRED.k331.cure", "klass": "PRED", "value": "gh pr close with NO --comment flag", - "line": 821 + "line": 831 }, { "id": "PRED.k331.evidence", "klass": "PRED", "value": "2026-05-09 wave-3: violation on #3300 close, deleted within 30s", - "line": 823 + "line": 833 }, { "id": "PRED.k331.k101-restatement", "klass": "PRED", "value": "k101 includes close-time --comment flag; rationale belongs in subsuming PR's squash-merge body", - "line": 820 + "line": 830 }, { "id": "PRED.k331.recovery", "klass": "PRED", "value": "if violation lands, gh api -X DELETE repos///issues/comments/", - "line": 822 + "line": 832 }, { "id": "PRED.k331.shape", "klass": "PRED", "value": "instruction \"close with no comment (rationale)\" — parenthetical is rationale, NOT comment body", - "line": 819 + "line": 829 }, { "id": "PRED.k331.signal", "klass": "PRED", "value": "close-with-no-comment-is-literal", - "line": 818 + "line": 828 }, { "id": "PROBE.ci.surface", "klass": "PROBE", "value": "the contract (parse/validate, projection round-trip, fail-closed guards), NEVER the LLM judgment (ADR-550 D5)", - "line": 567 + "line": 575 }, { "id": "PROBE.core.seam", "klass": "PROBE", "value": "analyzeCoverage(items,resolutions?,validators) ingests ALREADY-proposed items; does NOT assume deterministic propose (ADR-550 D7b)", - "line": 560 + "line": 568 }, { "id": "PROBE.edge.verification", "klass": "PROBE", "value": "explicit|backstop", - "line": 562 + "line": 570 }, { "id": "PROBE.family", "klass": "PROBE", "value": "edge-probe(shape-axis)+prohibition-probe(must-NOT-axis)+ui-consideration-probe(UI-state-axis), shared probe-core, run as spec-phase/ui-phase soft gates (ADR-550 D7; #1867)", - "line": 558 + "line": 566 }, { "id": "PROBE.item.axes", "klass": "PROBE", "value": "status{resolved|dismissed|unresolved} x verification{|null} — orthogonal; the lifecycle enum carries no verification fact (ADR-550 D7a)", - "line": 561 + "line": 569 }, { "id": "PROBE.principle", "klass": "PROBE", "value": "verifier-reach-equals-spec-reach (a goal-backward verifier only checks assertions that exist; probes make omitted assertions exist before code) — ADR-857 verification-substrate boundary; docs/design/verifier-reach.md", - "line": 557 + "line": 565 }, { "id": "PROBE.prohib.verification", "klass": "PROBE", "value": "test|judgment", - "line": 563 + "line": 571 }, { "id": "PROBE.protocol", "klass": "PROBE", "value": "recall(adversarial over-generate)->precision(drop routine-engineering); dismissals require a non-empty reason", - "line": 559 + "line": 567 }, { "id": "PROBE.ui.axis", "klass": "PROBE", "value": "MIXED — closed compiled shape-rooted 8 (empty/loading/error/populated/partial/overflow/zero-one-many/long-text) via ui-consideration-probe adapter; open UX (real-time/a11y/i18n-RTL) prose-owned in references/domain-probes.md, NOT compiled (#1867)", - "line": 565 + "line": 573 }, { "id": "PROBE.ui.seam", "klass": "PROBE", "value": "ui-phase Step 9.5 post-verification: element-cue classify -> propose-then-confirm (partial-cue mitigation, Goodhart) -> autoResolve --auto floor (never dismiss; unclassified stays unresolved #1110) -> ## UI Considerations write-back -> plan-phase `## UI Considerations` lift rule (#1867)", - "line": 566 + "line": 574 }, { "id": "PROBE.ui.verification", "klass": "PROBE", "value": "explicit|backstop", - "line": 564 + "line": 572 }, { "id": "PROC.AGENT-DISPATCH.completion-verify", "klass": "PROC", "value": "run k324.poll-shape on every agent-completion notification", - "line": 827 + "line": 837 }, { "id": "PROC.AGENT-DISPATCH.parallel-overlap-audit", "klass": "PROC", "value": "before dispatching N sibling-audit fixers, compute file-set union and assign canonical owners", - "line": 826 + "line": 836 }, { "id": "PROC.AGENT-DISPATCH.preflight", "klass": "PROC", "value": "[read-CONTRIBUTING.md-fresh, read-relevant-ADRs, cite-specific-line-in-brief, require-closing-keyword, require-changeset-fragment, forbid-CHANGELOG.md-edit, require-isolation-worktree, forbid-self-PR-comment, mandate-trust-but-verify]", - "line": 825 + "line": 835 }, { "id": "PROC.MERGE-WAVE.changelog-strip-pattern", "klass": "PROC", "value": "detached-HEAD per k325 + git checkout main -- CHANGELOG.md + commit + force-with-lease", - "line": 831 + "line": 841 }, { "id": "PROC.MERGE-WAVE.merge-tool", "klass": "PROC", "value": "gh pr merge --squash --delete-branch", - "line": 832 + "line": 842 }, { "id": "PROC.MERGE-WAVE.merge-tool-warning", "klass": "PROC", "value": "delete-branch may fail with \"used by worktree at\" — harmless; remote branch still deleted", - "line": 833 + "line": 843 }, { "id": "PROC.MERGE-WAVE.ordering", "klass": "PROC", "value": "[wave1: isolated-files, wave2: CHANGELOG-only-overlap (better: strip per k320), wave3: same-file-overlap with explicit decision]", - "line": 829 + "line": 839 }, { "id": "PROC.MERGE-WAVE.preflight", "klass": "PROC", "value": "gh pr view --json files for every PR; identify overlap pairs; surface to maintainer", - "line": 830 + "line": 840 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.observed", "klass": "PROC", "value": "#3541 + #3542 dispatched simultaneously this session; PRs #3546 #3547 opened green; one syntax slip caught by AGENT-RETIRED-SLASH-SYNTAX-DRIFT and fixed before second PR opened", - "line": 909 + "line": 919 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.pattern", "klass": "PROC", "value": "bot triage brief → worktree per branch → parallel sub-agents do rubber-duck/RCA/TDD implementation only → top-level orchestrator owns commit + gsd-test + push + PR + changeset-pr-backfill", - "line": 907 + "line": 917 }, { "id": "PROC.PARALLEL-FIX-DISPATCH.rationale", "klass": "PROC", "value": "long-running test runs need cross-turn notifications (orchestrator-only); CONTRIBUTING.md gh-templates-first hook requires session-scoped Read calls sub-agents wouldn't otherwise make; sequencing test runs avoids GSD-TEST-CONCURRENT-OUTPUT-COLLISION", - "line": 908 + "line": 918 }, { "id": "PROC.TRIAGE.comment-shape", "klass": "PROC", "value": "lead with \"duplicate of #NNNN, fixed by PR #MMMM, in v1.X.Y\"; show current code snippet proving bug-surface gone; give @latest and @next upgrade commands; close", - "line": 912 + "line": 922 }, { "id": "PROC.TRIAGE.no-duplicate-label", "klass": "PROC", "value": "this repo has no duplicate label; framing lives in comment text + closing the issue", - "line": 913 + "line": 923 }, { "id": "PROC.TRIAGE.routing-incoming", "klass": "PROC", "value": "stale-bug-already-fixed to close as duplicate of originating issue + cite fix PR + first stable tag; release-publish-or-backport to ready-for-human; reporter-can-self-test to awaiting-retest", - "line": 911 + "line": 921 }, { "id": "PROHIB.canon-referral", "klass": "PROHIB", "value": "OWASP/GDPR/fairness-canon are REFERRED to /gsd:secure-phase+eslint, never minted as prohibitions (ADR-550 D6)", - "line": 569 + "line": 577 }, { "id": "PROHIB.descriptor.shape", "klass": "PROHIB", "value": "5 FLAT scalars (check_kind,check_target,check_rule,check_violation_fixture,check_clean_fixture) — NEVER a nested check:{} (parseMustHavesBlock is a flat parser, src/frontmatter.cts)", - "line": 574 + "line": 582 }, { "id": "PROHIB.enforce.adr", "klass": "PROHIB", "value": "docs/adr/1606-prohibition-enforcement-verify-seam.md (verify-time enforcement seam) + docs/adr/550-spec-phase-probe-contract.md (spec-phase contract)", - "line": 577 + "line": 585 }, { "id": "PROHIB.enforce.causation", "klass": "PROHIB", "value": "clean-fixture control proves the red is content-caused not env-var-set; MANDATORY for node-test (#1906 supersedes #1346 opt-in) — absent clean-fixture ⇒ node-test un-provable/fail-closed; lint-rule needs none (its subject IS the linted file)", - "line": 573 + "line": 581 }, { "id": "PROHIB.enforce.failfirst", "klass": "PROHIB", "value": "MACHINE-PROVEN against an author-supplied violation fixture (#1279); caller failFirst attestation DEMOTED to a non-authoritative hint (FF-08)", - "line": 572 + "line": 580 }, { "id": "PROHIB.enforce.green-rule", "klass": "PROHIB", "value": "passed iff provenFailFirst===true && run.passed===true (runProhibitionEnforcement); every miss/fail/un-provable HARD-GATES both modes via dispositionForProhibition's fail-closed default", - "line": 570 + "line": 578 }, { "id": "PROHIB.enforce.kinds", "klass": "PROHIB", "value": "node-test (non-vacuous red via isNonVacuousNodeTestRed; pass-side vacuity via isNonVacuousNodeTestPass) | lint-rule (eslint --format json filtered by ruleId)", - "line": 571 + "line": 579 }, { "id": "PROHIB.judgment-tier", "klass": "PROHIB", "value": "never-silent / never-hard-halt soft gate; autonomous emits \"unverified-prohibition — human review recommended\" (exogenous grading, ADR-550 D4)", - "line": 576 + "line": 584 }, { "id": "PROHIB.rail", "klass": "PROHIB", "value": "core verify rail, non-toggleable (ADR-857 verification-substrate boundary / decision #6); the verifier<->predicate contract is NOT an off-by-default capability", - "line": 575 + "line": 583 }, { "id": "PROHIB.recall", "klass": "PROHIB", "value": "LLM-prose; no compiled prohibition-probe recall engine (only the schema/projection layer is code, ADR-550 D7b)", - "line": 568 + "line": 576 }, { "id": "RELEASE-NOTES.ANTI-PATTERN", "klass": "RELEASE-NOTES", "value": "raw \"What's Changed\" PR list as final body for hotfix or feature release; \"Full Changelog only\" body for tagged release with >0 user-facing fixes", - "line": 722 + "line": 732 }, { "id": "RELEASE-NOTES.ANTI-PATTERN.implementation-first", "klass": "RELEASE-NOTES", "value": "do not lead bullet with file path or function name; lead with symptom/user-visible behavior", - "line": 723 + "line": 733 }, { "id": "RELEASE-NOTES.ANTI-PATTERN.risk-commentary", "klass": "RELEASE-NOTES", "value": "do not include \"may break\", \"be careful\", \"test thoroughly\" - release notes state what changed, not hedges about what might go wrong", - "line": 724 + "line": 734 }, { "id": "RELEASE-NOTES.DEFAULT-STATE", "klass": "RELEASE-NOTES", "value": "auto-generated body is \"What's Changed\" PR list + Full Changelog link; treat as draft, not final", - "line": 698 + "line": 708 }, { "id": "RELEASE-NOTES.EXAMPLE.hotfix", "klass": "RELEASE-NOTES", "value": "v1.41.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.41.1) - 14 fixes grouped by 6 subgroups", - "line": 726 + "line": 736 }, { "id": "RELEASE-NOTES.EXAMPLE.minor-auto-acceptable", "klass": "RELEASE-NOTES", "value": "v1.41.0 - kept auto-generated body; many small fixes with clean conventional-commit titles", - "line": 728 + "line": 738 }, { "id": "RELEASE-NOTES.EXAMPLE.rc", "klass": "RELEASE-NOTES", "value": "v1.7.0-rc.1 (https://github.com/open-gsd/gsd-core/releases/tag/v1.7.0-rc.1) - intro + Added/Changed/Fixed/Documentation taxonomy", - "line": 727 + "line": 737 }, { "id": "RELEASE-NOTES.GATE.hotfix", "klass": "RELEASE-NOTES", "value": "manual edit required; auto-generated body for vX.Y.{Z>0} is \"Full Changelog only\" and must be replaced with structured body", - "line": 699 + "line": 709 }, { "id": "RELEASE-NOTES.GATE.minor", "klass": "RELEASE-NOTES", "value": "auto-generated body acceptable when PR titles are clean; promote to structured body when >20 PRs or contains feature+refactor+fix mix", - "line": 701 + "line": 711 }, { "id": "RELEASE-NOTES.GATE.rc", "klass": "RELEASE-NOTES", "value": "manual edit recommended; auto-generated PR list is acceptable for early RCs but final RC before vX.Y.0 should match standard", - "line": 700 + "line": 710 }, { "id": "RELEASE-NOTES.RELEASE-STREAM.main-branch", "klass": "RELEASE-NOTES", "value": "next (RCs) + latest (stable); install via @next or @latest", - "line": 733 + "line": 743 }, { "id": "RELEASE-NOTES.RELEASE-STREAM.rule", "klass": "RELEASE-NOTES", "value": "streams do not mix; do not document @next in hotfix/stable notes", - "line": 734 + "line": 744 }, { "id": "RELEASE-NOTES.SCOPE", "klass": "RELEASE-NOTES", "value": "GitHub Releases body for tags vX.Y.Z, vX.Y.Z-rc.N; not CHANGELOG.md (changeset workflow owns that)", - "line": 697 + "line": 707 }, { "id": "RELEASE-NOTES.SOURCE.changesets", "klass": "RELEASE-NOTES", "value": ".changeset/*.md (frontmatter pr: + body bullets)", - "line": 713 + "line": 723 }, { "id": "RELEASE-NOTES.SOURCE.commits", "klass": "RELEASE-NOTES", "value": "git log .. --pretty=format:'%s%n%n%b' --no-merges", - "line": 712 + "line": 722 }, { "id": "RELEASE-NOTES.SOURCE.pr-bodies", "klass": "RELEASE-NOTES", "value": "gh pr view --json title,body for fixes lacking a changeset", - "line": 714 + "line": 724 }, { "id": "RELEASE-NOTES.SOURCE.precedence", "klass": "RELEASE-NOTES", "value": "changeset body > commit body > PR body > commit subject (prefer authored content over auto-generated)", - "line": 715 + "line": 725 }, { "id": "RELEASE-NOTES.STANDARD.bullet-shape", "klass": "RELEASE-NOTES", "value": "**Bold user-visible change** — explanation of what was broken or what's new, leading with symptom not implementation. Trailing (#NNN) PR ref.", - "line": 705 + "line": 715 }, { "id": "RELEASE-NOTES.STANDARD.footer.full-changelog", "klass": "RELEASE-NOTES", "value": "**Full Changelog**: https://github.com/open-gsd/gsd-core/compare/...", - "line": 709 + "line": 719 }, { "id": "RELEASE-NOTES.STANDARD.footer.hotfix", "klass": "RELEASE-NOTES", "value": "Install/upgrade: \\`npx @opengsd/gsd-core@latest\\`", - "line": 707 + "line": 717 }, { "id": "RELEASE-NOTES.STANDARD.footer.rc", "klass": "RELEASE-NOTES", "value": "Install for testing: \\`npx @opengsd/gsd-core@next\\` (per branch->dist-tag policy)", - "line": 708 + "line": 718 }, { "id": "RELEASE-NOTES.STANDARD.heading-level", "klass": "RELEASE-NOTES", "value": "## for category, ### for subgroup (area), - for bullet", - "line": 704 + "line": 714 }, { "id": "RELEASE-NOTES.STANDARD.intro", "klass": "RELEASE-NOTES", "value": "optional one-paragraph framing for RC/feature releases; omit for pure-fix hotfixes", - "line": 710 + "line": 720 }, { "id": "RELEASE-NOTES.STANDARD.subgroups", "klass": "RELEASE-NOTES", "value": "phase-planning-state | workstream | query-dispatch-cli | code-review | install | capture | docs | architecture | security", - "line": 706 + "line": 716 }, { "id": "RELEASE-NOTES.STANDARD.taxonomy", "klass": "RELEASE-NOTES", "value": "Keep-a-Changelog 1.1.0: Added | Changed | Deprecated | Removed | Fixed | Security | Documentation", - "line": 703 + "line": 713 }, { "id": "RELEASE-NOTES.TEMPLATE.hotfix", "klass": "RELEASE-NOTES", "value": "## Fixed\\n\\n### \\n- **** — . (#)\\n\\n---\\n\\nInstall/upgrade: \\`npx @opengsd/gsd-core@latest\\`\\n\\n**Full Changelog**: ", - "line": 730 + "line": 740 }, { "id": "RELEASE-NOTES.TEMPLATE.rc", "klass": "RELEASE-NOTES", "value": "\\n\\n## Added\\n### \\n- **** — . (#)\\n\\n## Changed\\n### Architecture\\n- **** — . (#)\\n\\n## Fixed\\n### \\n- **** — . (#)\\n\\n## Documentation\\n- **** — . (#)\\n\\n---\\n\\nThis is a release candidate. Install for testing:\\n\\`\\`\\`bash\\nnpx @opengsd/gsd-core@next\\n\\`\\`\\`\\n\\n**Full Changelog**: ", - "line": 731 + "line": 741 }, { "id": "RELEASE-NOTES.WORKFLOW.edit", "klass": "RELEASE-NOTES", "value": "gh release edit --notes-file ", - "line": 717 + "line": 727 }, { "id": "RELEASE-NOTES.WORKFLOW.idempotency", "klass": "RELEASE-NOTES", "value": "gh release edit overwrites body wholesale; safe to re-run after refining", - "line": 720 + "line": 730 }, { "id": "RELEASE-NOTES.WORKFLOW.token", "klass": "RELEASE-NOTES", "value": "must use .envrc GITHUB_TOKEN per RULESET.GH.AUTH.DEFAULT (this doc); never ambient gh auth", - "line": 719 + "line": 729 }, { "id": "RELEASE-NOTES.WORKFLOW.view", "klass": "RELEASE-NOTES", "value": "gh release view --json body --jq .body", - "line": 718 + "line": 728 }, { "id": "RULESET.ADR-HEADER", "klass": "RULESET", "value": "every docs/adr/NNNN-*.md must open with - **Status:** Accepted|Proposed|Superseded (by [ADR-NNNN](file.md))|Legacy + - **Date:** YYYY-MM-DD immediately after title", - "line": 622 + "line": 632 }, { "id": "RULESET.AGENT_SIZE_BUDGET", "klass": "RULESET", "value": "agent-size-budget (#1074; sibling of WORKFLOW_SIZE_BUDGET; BYTES not lines per #717/#683, rebased from lines in PR 3/3) = differential attribution size ratchet (PRIMARY anti-creep since #2724/ADR-2719 §4, same mechanism and same ack fragments (tests/emitted-drift-acks/, #2914; legacy tests/emitted-drift-ack.json still honored) as WORKFLOW_SIZE_BUDGET, scoped to agents/gsd-*.md) + loose tier hard caps (red lines, never raised on approach: XL<=57344 / LARGE<=49152 / DEFAULT<=24576); net-new agents are DEFAULT-tier (no separate new-file cap). Sizes are measured via the shared scripts/workflow-size.cjs measureMdFiles(dir,predicate) counter (tests/helpers/emitted-runtime.cjs's currentSizes() and the guard's own tier-cap checks both import it). A grown agent fails the differential guard — ack + justify, or extract LAZILY to gsd-core/references/. DISTINCT from DEFECT.AGENT-FILE-SIZE-CAP-BREACH (a separate 45K-CHAR extraction-evidence threshold on gsd-planner via planner-decomposition/reachability tests): that guard proves mode-sections were extracted; this one bounds total agent bytes. Two guards, two units (chars vs bytes), two purposes. The prior per-file baseline (tests/agent-size-baseline.json, `npm run size:baseline`) is REMOVED by #2724", - "line": 611 + "line": 621 }, { "id": "RULESET.ALLOWED-TOOLS-FRONTMATTER", "klass": "RULESET", "value": "command's allowed-tools must cover every tool the workflow calls (including Write for file creation); thin-wrapper pattern makes this easy to miss", - "line": 618 + "line": 628 }, { "id": "RULESET.ARGUMENTS-SANITIZE", "klass": "RULESET", "value": "any workflow step constructing .planning/.../{SLUG}.md path from user input ($ARGUMENTS, parsed remainder) must sanitize inline ([a-z0-9-] only, reject ..//\\\\, max-length) — \"(already sanitized)\" must trace back to explicit guard; RESUME/fallback modes need own guards", - "line": 619 + "line": 629 }, { "id": "RULESET.AUDIT.search-source-not-generated", "klass": "RULESET", "value": "verify an invariant/validation EXISTS by searching the AUTHORED source (src/*.cts OR the scripts/gen-*.cjs generator), never the generated bin/lib/*.cjs (gitignored, ADR-457); gen-time checks live in gen-*.cjs not the .cts it consumes → search BOTH before declaring absent; read generated .cjs only for output drift. Repro: grep src/*.cts for VALID_CONVERTER_NAMES → false \"5e ConverterName unenforced\"; actually enforced in gen-capability-registry.cjs. cf RULESET.TESTS.no-source-grep", - "line": 607 + "line": 617 }, { "id": "RULESET.CAPABILITY.cutover-self-gating", "klass": "RULESET", "value": "a phase-6 per-feature cutover moves the host's phase-context detection + mode/flag logic INTO the skill (self-gating, per ADR-894); the loop hook is intentionally COARSE — \"invoke skill X at point Y when config Z\" — and carries no detection/mode. WORKED EXAMPLE: plan-phase.md §5.6 UI gate (frontend-detection via ui-safety-gate.cjs + --auto/manual branch + --skip-ui bypass) must move into gsd-ui-phase before its plan:pre hook can replace the inline call without behavior loss. Spike #1018 finding.", - "line": 388 + "line": 396 }, { "id": "RULESET.CAPABILITY.off-means-off", "klass": "RULESET", "value": "the host derives shared outputs from the ACTIVE hook set (via loop.render-hooks); a hook may ADD a labeled block or be COUNTED into a host-computed aggregate (e.g. a score denominator), but NEVER mutates host source — so a disabled capability yields the base output by construction, not by authoring discipline. Ratify in ADR-894; proven by spike #1018.", - "line": 386 + "line": 394 }, { "id": "RULESET.CAPABILITY.precedence-engine-single-owner", "klass": "RULESET", "value": "the config-key four-level precedence walk (loadConfig result → workstream config.json → root config.json → registry.configSchema default → absent) is owned solely by src/capability-activation.cts: raw-value primitive resolveConfigKey(dotKey, {config,cwd,registry}) and boolean wrapper _resolveActivationValue(dotKey,config,cwd,registry); loop-resolver.cts imports the engine (no duplicate); resolveConfigValues in loop-resolver.cts delegates to resolveConfigKey; resolveCapabilityRuntimeState does NOT return registry/config — callers import capability-registry.cjs and call loadConfig(cwd) directly.", - "line": 392 + "line": 400 }, { "id": "RULESET.CAPABILITY.step-additive-gate-blocks", "klass": "RULESET", "value": "a `step` hook is purely additive (invoke skill + produce artifacts, NEVER halts the host); host-blocking preconditions are `gate`s (blocking:true, onError:halt); runtime/mode context (auto/chain vs manual) self-gates IN THE SKILL, not via `when` (config-only). §5.6 = plan:pre step (ui-phase; skill self-gates on frontend+pipeline, auto-fires only in pipelines) + a NEW plan:pre gate (frontend-and-no-UI-SPEC → halt, when:workflow.ui_safety_gate); the loop.render-hooks dispatch template handles steps AND gates. Resolves #1022.", - "line": 390 + "line": 398 }, { "id": "RULESET.CODERABBIT.GUARD.COMPLETE", "klass": "RULESET", "value": "required_checks_green && coderabbit_check_pass && graphQL(reviewThreads.unresolved_count)==0", - "line": 644 + "line": 654 }, { "id": "RULESET.CODERABBIT.GUARD.GRAPHQL", "klass": "RULESET", "value": "reviewThreads(first:100){nodes{id isResolved comments{nodes{author body path line originalLine url}}}}; use unresolved threads as authoritative, not badge text alone", - "line": 645 + "line": 655 }, { "id": "RULESET.CODERABBIT.GUARD.OPEN_PRS", "klass": "RULESET", "value": "gh pr list --repo open-gsd/gsd-core --author @me --state open; repeat near end because open PR set can change mid-run", - "line": 643 + "line": 653 }, { "id": "RULESET.CODERABBIT.GUARD.RERUN", "klass": "RULESET", "value": "after every push wait for CodeRabbit completion, then re-query unresolved threads; CodeRabbit can add new findings after earlier threads were resolved", - "line": 646 + "line": 656 }, { "id": "RULESET.CODERABBIT.GUARD.RESOLVE", "klass": "RULESET", "value": "fix validated finding -> focused tests -> commit/push -> resolveReviewThread(threadId) -> wait CI/CodeRabbit -> final unresolved_count query", - "line": 647 + "line": 657 }, { "id": "RULESET.CODERABBIT.GUARD.SCOPE", "klass": "RULESET", "value": "if a new @me open PR appears during final list, include it in the same guard pass before declaring all-open-PRs complete", - "line": 648 + "line": 658 }, { "id": "RULESET.CONTENT-PATH-NORMALIZATION", "klass": "RULESET", "value": "filesystem paths substituted into markdown body text (@-references, workflow .md, agent .md, generated docs, command bodies) MUST be normalized to POSIX forward slashes via .replace(/\\\\/g,'/') at the production source BEFORE substitution; never push normalization to tests; cross-platform content is POSIX-only; applies to: computePathPrefix output, install-path rewrites, generated shim paths emitted into .md bodies; idempotent on POSIX so unconditional; mechanically enforced by local/normalize-path-in-content (eslint, src/**/*.cts; #1733)", - "line": 849 + "line": 859 }, { "id": "RULESET.CONTRIB.CLASSIFY.enhancement", "klass": "RULESET", "value": "requires approved-enhancement before implementation", - "line": 637 + "line": 647 }, { "id": "RULESET.CONTRIB.CLASSIFY.feature", "klass": "RULESET", "value": "requires approved-feature before implementation", - "line": 638 + "line": 648 }, { "id": "RULESET.CONTRIB.CLASSIFY.fix", "klass": "RULESET", "value": "requires confirmed-bug before implementation (legacy 'confirmed' label is back-compat only for duplicate-sweep exemption, not a valid implementation gate)", - "line": 636 + "line": 646 }, { "id": "RULESET.CONTRIB.GATE.ORDER", "klass": "RULESET", "value": "issue-first -> approval-label -> code -> PR-link -> changeset/no-changelog", - "line": 635 + "line": 645 }, { "id": "RULESET.CR-THREAD-RESOLVE", "klass": "RULESET", "value": "after adding // allow-test-rule: to silence lint, resolve existing inline CR threads via graphql resolveReviewThread mutation before merge — open threads mislead future reviewers; pattern: gh api graphql -f query='mutation { resolveReviewThread(input:{threadId:\"PRRT_...\"}) { thread { isResolved } } }'", - "line": 629 + "line": 639 }, { "id": "RULESET.EMITTED_ATTRIBUTION", "klass": "RULESET", "value": "the emitted-artifact family (ADR-2719, epic #2719) — POST-CUTOVER (#2724, Phase 4). Historically tests/fixtures/golden-install-parity/*.json (19 path→hash manifests) + tests/workflow-size-baseline.json + tests/agent-size-baseline.json were all committed, PURE FUNCTIONS of the source tree whose correct merge was ALWAYS \"recompute\" — 140 of 143 conflicted-file instances across the open PR queue were these files. #2724 DELETES all three, the golden test (tests/golden-install-parity.test.cjs), the generator (scripts/gen-golden-install-parity-zcode.cjs), `npm run gen:golden`, `UPDATE_GOLDEN`, the merge-driver bridge (scripts/git-merge-regen-driver.cjs, `npm run setup:merge-driver`, the .gitattributes merge=gsd-regen block), and scripts/update-size-baseline.cjs (`npm run size:baseline`). The differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the SOLE gate for emitted-artifact propagation AND size growth — no committed artifact, nothing to hand-merge, nothing to regenerate. `npm run regen:derived` still exists for what remains committed and derived: build, registry, ADR index, capability matrix, inventory manifest, manifest versions, and `tests/fixtures/install-tree/*.json` (now `npm run gen:install-tree`, folded into `regen:derived`). tests/fixtures/install-tree/*.json is DELIBERATELY EXCLUDED from the cutover (ADR-2719 §7): it conflicts on 0 of 7, its diffs are readable, and it preserves \"the installer stopped shipping X\" as a hard absolute failure — capturing it would convert that absolute into an attribution-free auto-resolve. The baseline the differential compares against is now published by `scripts/gen-emitted-baseline.cjs` on every push to `next` (cached, keyed on sha) and restored in PR lanes via `GSD_EMITTED_BASELINE`/`resolveBaseline()` (tests/helpers/emitted-baseline.cjs); a cache miss falls back to an in-job build via a throwaway `git worktree` (tests/helpers/emitted-runtime.cjs's `buildBaselineAtRef`). REMEDIATION IS PART OF THE GATE (#2778): the failure output names its own remedy, because a gate that states a requirement and withholds the means of satisfying it is a maintainer round-trip, not a gate — ADR-2719 §3's \"conspicuous declaration\" only works if the contributor can discover how to make it. Both failing branches name a NEW fragment to create under `tests/emitted-drift-acks/` (#2914; pick a name nobody else is using), say it may not exist yet (absence is the healthy steady state), print a minimal valid document, and repeat \"do NOT regenerate anything\" — post-#2724 there is nothing left to regenerate, and hunting for a deleted baseline is the predictable wrong guess. The two branches key on DIFFERENT spaces and each says which: the hash pass keys on the EMITTED PATH (always contains a `/`), the size ratchet keys on the BARE FILENAME (`currentSizes` writes `sizes[entry.name]` from readdirSync over `gsd-core/workflows/` + `agents/`). A stale-ack failure additionally says to delete the FILE when removing its last entry, since an empty-but-present ack parses fine yet signals nothing; post-#2789 it also offers CORRECTING the entry to name the ripple actually made, which is the other honest resolution and the one a contributor usually wants. NOT ack-able and deliberately given no ack text: the `NEW_FILE_CAP` branch, whose remedy is extraction. Text is sourced from one frozen `REMEDIATION` export in tests/helpers/emitted-diff.cjs whose example document is rendered from `ACK_VERSION` via `JSON.stringify`, so the taught schema cannot drift from the accepted one (a round-trip test feeds the printed document back through `parseAck`); the message teaches ONE canonical shape even though `parseAck` also accepts a bare-string reason and a missing `version` — liberal in what it accepts, conservative in what it sends. Note the ADR's Consequences originally called the #2724 migration \"terminal\"; #2778 corrected that — it is terminal only for a PR that grows no shipped file. #2914 replaced the single shared ack file with per-PR fragments under `tests/emitted-drift-acks/` — exactly the shape `.changeset/` already uses for the identical \"every PR rewrites one shared document\" conflict problem — so two PRs needing an ack can no longer collide with each other on the FILE; the legacy file is still read and unioned in for branches that predate the split, and a duplicate path key across two sources is a hard, loudly-reported error, never silent last-wins — #3078 made that error name its two resolutions (git rm an already-merged, spent owner; APPEND prose to a still-live one, which re-arms it), because the guard runs post-merge and cannot stop the colliding PR. `tests/emitted-drift-ack.json` (the LEGACY file specifically) must NEVER persist on `next` (#2914): every entry is scoped to the diff that introduced it, so once merged it is by definition already at the base — spent and inert regardless of shape — and a persistent copy makes that ONE file a shared merge-conflict cell across every open PR that also carries an ack, exactly the \"140 of 143\" cost this whole cutover exists to remove; #2914 asserted a persisting FRAGMENT was harmless by construction and deliberately exempted the directory; #3078 REVERSED that — fragments do not share a FILE but they DO share a PATH KEY SPACE, so a fully-spent fragment on `next` owns keys it can no longer gate and the next PR growing one of those paths can declare it neither there (spent) nor in its own (duplicate), which is the #2914 wall one level down (measured at the sweep: 45 fragments owning 403 paths, up from 13/272 at triage 19 days earlier). A fragment is judged on INERTNESS, not presence: swept once EVERY entry is spent, left alone while PARTIALLY spent — the asymmetry is what keeps the re-arm-by-appending route (#2639, #2993) working, and the `0000` legacy-migration bucket #2923 created for the old shared file's 35 entries was NOT permanent (the issue's own open question resolved to NO) and went with the rest. This is enforced on `next` itself only, never as a PR-lane check: the `guard-no-ack-on-next` job in `.github/workflows/test.yml` (push-to-`next` trigger) runs `scripts/lint-emitted-drift-ack.cjs --guard-next`, which is now BOTH halves — `assertAbsentOnNext` (legacy file, fails on PRESENCE alone, valid or not) and `assertNoAllSpentFragments` (fragments, fails on all-entries-spent vs the copy at the PRE-PUSH TIP of next — CI passes `github.event.before` via `--base-ref`, because the default branch allows REBASE merges so one push can carry N commits and a bare `HEAD^` would flag a fragment the same push introduced; `HEAD^` remains only the local/manual fallback, using the SAME zero-width/whitespace-stripping prose comparison as `isSpent` so an invisible reword cannot fake a re-arm; duplicated across the scripts-ship/tests-do-not line and held by a parity test). The job's checkout REQUIRES `fetch-depth: 2` plus an explicit `git fetch --depth=1 origin $BEFORE` — at depth 1 no base commit exists locally, every fragment reads as brand-new, and the guard passes vacuously, which is exactly how the legacy half went blind after #2914 removed the file it was watching. The gate's `INVISIBLE`/`normalizeAckReason` are EXPORTED from tests/helpers/emitted-diff.cjs for the sole purpose of letting the parity test compare them against the script's duplicate; before #3078 neither was exported, so the \"parity test\" the comments promised was a tautology checking the script against itself. A PR-lane \"base ack must be absent\" check would red every open PR the instant a spent ack merged, which is the #2768 shape #2789 already ended — so this alerts AFTER the merge by design and never stops the offending PR. #3875 automated the REMEDY that alert asks for, because detection without an executable remedy is what actually failed: #3823 shipped the guard together with a static 45-fragment sweep computed at its own branch point, #3809's fragment merged to `next` while it was in flight, and the guard reddened on its own merge commit and stayed red for 24 consecutive pushes over two days — the sweep condition is computed DYNAMICALLY at merge time while a hand-authored `git rm` is fixed at BRANCH time, so on a moving branch the second can never reliably satisfy the first. `runGuardNext` therefore returns the set it reasoned about (`sweepable`, already narrowed by the #3842 hold, plus `legacyPresent` for the legacy document, which is a fixed path rather than a fragment basename and would otherwise be invisible to any sweeper), `--sweep-plan` emits that set as a work list on stdout with the prose diverted to stderr and exit 0 (a non-empty plan is the NORMAL case, and a non-zero exit would fail the step that asked for the list), and `.github/workflows/ack-fragment-sweep.yml` runs it on a timer and opens a reviewable PR rather than pushing to protected `next`. The plan is re-validated against a literal allowlist before any deletion and each path is removed under a `:(literal)` pathspec — `git rm` reads its arguments as PATHSPECS with wildmatch semantics, so a fragment named `*.json` (a legal filename that `listFragmentFiles` admits, since it filters only on the suffix) would otherwise expand to every fragment in the directory, including ones the #3842 hold deliberately withheld. An empty plan is NOT reported as success on its own: the guard is re-run without the hold to separate \"next is clean\" from \"everything is held\", the commonest holder being the sweep PR from the previous run, which touches precisely the fragments it proposed to delete and would otherwise make the automation go silently inert. cf `RULESET.WORKFLOW_SIZE_BUDGET`, `RULESET.AGENT_SIZE_BUDGET`; see `### Emitted Artifact Provenance`", - "line": 612 + "line": 622 }, { "id": "RULESET.GENERATIVE-FIX", "klass": "RULESET", "value": "parallel implementations diverge silently when no parity test enforces equality at the test layer; for any new constant/array/parser shared between two parallel surfaces (two workflow surfaces, or a generated artifact and its hand-authored source), the same commit MUST add a parity assertion that fails when the two diverge; exemplar: tests/runtime-launcher-parity.test.cjs (asserts every workflow bash block uses the canonical gsd_run launcher)", - "line": 847 + "line": 857 }, { "id": "RULESET.GH.AUTH.DEFAULT", "klass": "RULESET", "value": "source .envrc GITHUB_TOKEN before gh; exception=ambient allowed only when user explicitly says machine-only fallback", - "line": 642 + "line": 652 }, { "id": "RULESET.HARNESS.test-memory-guard", "klass": "RULESET", "value": "~/.claude/hooks/test-memory-guard.sh fires on every Bash PreToolUse; if argv[0]∈{node|vitest|jest|mocha|tsx|ts-node|tap|ava|playwright|cypress} OR matches (npm|pnpm|yarn|bun) (run )?(t|test|tests|vitest|jest); blocks via hookSpecificOutput.permissionDecision=deny when sum(RSS of running matching procs, excluding tsserver|*-mcp|claude|Electron|...) ≥ 4 GiB OR when argv[0] basename matches a running process's argv[0]. Exception: node --version|-v|--help|-h|-p|-e are trivial probes and skip the check. Designed for a 24 GB Mac where prior accidental fan-out exhausted RAM", - "line": 888 + "line": 898 }, { "id": "RULESET.MANIFEST-CANONICAL-KEY", "klass": "RULESET", "value": "docs/INVENTORY-MANIFEST.json has a single top-level key: families; ALL EIGHT families.* arrays (agents/commands/workflows/references/cli_modules/hooks flat, plus workflow_modes/workflow_steps nested — #2996, epic #1671 Phase 6.5) are canonical, consumed by test suites — tests/inventory-manifest-sync.test.cjs reads all eight, edit-phase/enh-2380/enh-2430 tests read commands+workflows; the six flat families are keyed by BARE BASENAME while the two nested families are keyed by // path, deliberately, because two workflows may each own a same-named step file and a basename key would silently drop one under a JSON-equality comparison; recursion is bounded at exactly one named subdirectory, never a general walk; the family tables live ONCE in scripts/gen-inventory-manifest.cjs and are IMPORTED by the test (the test formerly redeclared them, a DEFECT.GENERATIVE-FIX divergence that let a new family be verified by nobody while still reporting green); the old generated date field and the stale top-level workflows key are both gone; regen via node scripts/gen-inventory-manifest.cjs --write, AFTER build:lib; #3762 added the ROSTER half — tests/inventory-manifest-sync.test.cjs now also asserts every manifest entry has a hand-written row in docs/INVENTORY.md, via the pure matcher in tests/helpers/inventory-roster.cjs. Scope is the SIX FLAT families only, each searched inside its own `## ` section; workflow_steps/workflow_modes are DELIBERATELY exempt because docs/INVENTORY.md §\"Workflow Sub-Files\" is a shipped decision that they carry no hand-written per-file rows. Matching is whole-CELL-exact (never substring — the rostered host-integration-adapters/imperative-hook-bus.cjs must not satisfy the separate top-level hook-bus.cjs) and section-scoped (smart-entry.md and smart-entry.cjs are different families), EXCEPT commands, which match on the row's Source-column link to ../commands/gsd/.md because the six ns-* namespace routers deliberately RENDER a name that is not their file stem (/gsd-workflow ← ns-workflow.md) — DEFECT.DISPLAY-VALUE-AS-IDENTITY. Landing the gate required backfilling 32 pre-existing unrostered surfaces on next", - "line": 623 + "line": 633 }, { "id": "RULESET.PR-FLOW.docker-before-push", "klass": "RULESET", "value": "before ANY git push of any fix to any PR, run gsd-test (docker on the remote, mirrors ubuntu CI) and confirm exit 0. macOS-local node --test is NOT a substitute — many failures are platform-specific (path separators, case sensitivity, locale, fs semantics). Watchdog with Monitor on the output log; never set a sleep/timer and walk away. Source: user feedback 2026-05-16 — \"we don't set a timer we actively watch and record results in real time as possible\". SUPERSEDED 2026-07-17: 'confirm exit 0' is a false-green trap — piping/backgrounding can report exit 0 on a failed suite; gate on the verdict-line outcome:\"passed\" for the exact HEAD sha instead. See CLAUDE.md's gsd-test rule and the gsd-test-is-ref-based-commit-first predicate for the current, correct gating contract.", - "line": 890 + "line": 900 }, { "id": "RULESET.PR-FLOW.templates-mandatory", "klass": "RULESET", "value": "every gh pr create|edit|gh issue create|edit MUST first invoke the gh-templates-first skill and Read (Read tool, not Bash cat — k321 read-tracking) the matching template in .github/. Apply ALL required sections; never write freeform bodies. Repo enforces this via gsd-pr-template-policy GitHub Action which flags any non-templated body — the bot allows the PR to stay open only because authors are contributors-or-higher, but the warning is a real complaint that must be cured. Source: user feedback 2026-05-16 (multi-message escalation) — \"the whole reason i have that github action is because you fucking blow through and ignore using the templates\"", - "line": 892 + "line": 902 }, { "id": "RULESET.PR-SCOPE.one-concern-per-pr", "klass": "RULESET", "value": "split unrelated changes into separate PRs; cherry-pick doc changes to dedicated docs/ branch immediately, then force-push original to remove the commit", - "line": 625 + "line": 635 }, { "id": "RULESET.SHARED-HELPERS-LINT-VS-TEST", "klass": "RULESET", "value": "when a lint script and test suite both implement same constant (CANONICAL_TOOLS) or parser (parseFrontmatter, executionContextRefs), extract to scripts/*-helpers.cjs required by both — silent divergence otherwise", - "line": 620 + "line": 630 }, { "id": "RULESET.TESTS.CODERABBIT_FIX", "klass": "RULESET", "value": "prefer exported-function behavioral tests over source-grep; lint-no-source-grep rejects readFileSync source assertions without allow-test-rule", - "line": 649 + "line": 659 }, { "id": "RULESET.TESTS.boundary-coverage", "klass": "RULESET", "value": "tests MUST exercise inputs at and near the threshold/limit, not only trivial-fit and trivial-overflow; pick inputs where N ∈ {limit-1, limit, limit+1} and where pre-trim/pre-check accumulators ≈ effective limit; \"very small\" and \"very large\" inputs alone do not constitute edge-case coverage and routinely miss off-by-one + reservation-accounting bugs", - "line": 594 + "line": 602 }, { "id": "RULESET.TESTS.boundary-coverage.anti-pattern", "klass": "RULESET", "value": "test suites that pair budget:1_000_000 (trivially fits) with budget:1 (trivially overflows) and skip the boundary region; failure mode that shipped PR #3708 UNNEEDED_TRIM + FALSE_HARDFAIL regressions (commit 2df566ed, fixed bde1ae8f)", - "line": 597 + "line": 605 }, { "id": "RULESET.TESTS.boundary-coverage.fixtures", "klass": "RULESET", "value": "for any code with budget/limit/quota/threshold parameter, test suite MUST include: (a) input where SUT estimate == limit exactly, (b) input where estimate == limit - 1, (c) input where estimate == limit + 1, (d) input where any internal reserve/safety constant pushes baseline within reserve-distance of limit (catches early-pressure firing)", - "line": 596 + "line": 604 }, { "id": "RULESET.TESTS.clock-seam", "klass": "RULESET", "value": "concurrency logic must accept an optional {clock=Date} parameter; tests control time via t.mock.timers.enable(['Date']) + t.mock.timers.setTime(0) + t.mock.timers.tick(N); real OS scheduler races are not a permitted test pattern after ADR 456 (2026-05-28); real-race tests are deleted once deterministic seam tests cover the same logical path; clock.cjs realClock adds nowIso() (→ new Date(this.now()).toISOString()) and today() (→ nowIso().split('T')[0]) so all date-stamping in state.cjs routes through the seam; subprocess time-pin adapter: set GSD_TEST_MODE=1 + GSD_NOW_MS= in runGsdTools env to pin the date written by the SUT without touching real wall-clock (issue #474)", - "line": 601 + "line": 609 }, { "id": "RULESET.TESTS.coderabbit-fix-prefer", "klass": "RULESET", "value": "behavioral tests (call exported fn, capture JSON, assert typed fields) over source-grep", - "line": 592 + "line": 600 }, { "id": "RULESET.TESTS.delete-bad-tests", "klass": "RULESET", "value": "pass-always / vacuous-truth / source-grep / elapsed-time / real-race / permanent-allow-test-rule tests are DELETED and replaced with compliant tests in the same PR; not skipped, not commented out, not permanently exempted; replacement must cover the same logical path via typed-surface assertion or clock-seam pattern", - "line": 604 + "line": 614 }, { "id": "RULESET.TESTS.diagnostics", "klass": "RULESET", "value": "after JSON.parse, assert output shape (Array.isArray(output.phases)) with raw-output-prefix diagnostics before .map() — prevents opaque TypeErrors when CLI output shape changes", - "line": 593 + "line": 601 }, { "id": "RULESET.TESTS.escape-regex", "klass": "RULESET", "value": "new RegExp(\"prefix${var}\") must escapeRegex(var); phase-id.cjs exports escapeRegex (core.cjs re-export spine retired in epic #1267); phase IDs like 5.1 contain . which is metacharacter", - "line": 589 + "line": 597 }, { "id": "RULESET.TESTS.eslint-harness", "klass": "RULESET", "value": "ADR 452 (2026-05-28): ESLint flat config + typescript-eslint + eslint-plugin-n + eslint-plugin-no-only-tests + local plugin at eslint-rules/ (repo root, NOT scripts/eslint-rules/); replaces scripts/lint-*.cjs regex scanners (fully removed in #632); all three test-rigor rules now ship at error in tests/**/*.test.cjs scope: local/no-source-grep and local/no-magic-sleep-in-tests promoted by #3313, local/no-elapsed-assertion promoted by #3331 once #3314 delivered its ADR-456 §(a) precondition (epic #1885 was subsumed into epic #3053 and closed stale before this promotion landed)", - "line": 605 + "line": 615 }, { "id": "RULESET.TESTS.feedback-loop-convergence", "klass": "RULESET", "value": "when a feature's OUTPUT feeds back into its own INPUT (calibration, retry backoff, adaptive budgets, ratchets, any self-correcting signal), step-wise tests are NOT sufficient evidence of correctness: they assert `given X return Y` while the defect lives in the TRAJECTORY across iterations. Required: a closed-loop test that (a) drives the REAL end-to-end surface — not the pure core alone, since composition bugs live between surfaces — for N >= 2x the loop's window, (b) asserts convergence on the known-true value, (c) asserts the fixed point (an already-correct history must produce NO correction), and (d) asserts boundedness under an adversarial/oscillating history. Two defects shipped past a green ~26,800-test suite in epic #1952 for want of exactly this: calibration applied twice across two surfaces (factor^2, #2631) and calibration measured against its own corrected output so it oscillated to ~1.41 instead of converging on 2.0 (#2632). Every unit, boundary, property and round-trip test passed for both. HOW TO SPOT ONE (the detection tell, not a judgment call): the feature's own acceptance criterion carries a TEMPORAL QUANTIFIER — \"after N phases\", \"subsequent\", \"over time\", \"improves\", \"learns\", \"adapts\". That phrasing means the claim is about a TRAJECTORY, so a step-wise `given X return Y` test does not test the claim that was made. #1952's AC4 read \"After N phases, the error is computed and applied as a correction to SUBSEQUENT estimates\" — the tell was in plain sight and was still tested as a point. Survey of this repo (2026-07): estimation calibration is the ONLY true instance; size/mutation ratchets are exempt because they fail on both growth AND shrinkage (cannot self-satisfy), and retry ladders (node_repair_budget, plan_bounce_passes, provider_escalation) terminate rather than feed back. Test anchor: tests/estimate-loop-convergence.test.cjs", - "line": 595 + "line": 603 }, { "id": "RULESET.TESTS.guard-toplevel-readFileSync", "klass": "RULESET", "value": "module-level const src = readFileSync(...) throws before any test() registers — wrap in try/catch in test() or use lazy load", - "line": 591 + "line": 599 + }, + { + "id": "RULESET.TESTS.mutation-runner", + "klass": "RULESET", + "value": "Stryker executes every shard through the OFFICIAL @stryker-mutator/tap-runner (testRunner:'tap'), never the built-in 'command' runner (#3915); 'command' is the one runner Stryker excludes from coverage analysis, which forced coverageAnalysis:'off' and made cost strictly linear in (mutants x whole-shard test time) — the frontmatter shard measured 1751s on run 33021042847 vs 212s for the next slowest. tap.testFiles is injected per shard via MUTATION_TEST_FILES (mutation.yml env <- matrix.tests <- scripts/mutation-matrix.cjs buildResult); resolveMutationTestFiles is the SINGLE fail-closed reader and existence-checks every entry, because the tap runner's findTestyLookingFiles resolves the list with glob() and a non-matching pattern yields an EMPTY list SILENTLY (a fast, confident, meaningless run). tap.forceBail is FALSE by measurement, not preference: 3 of 26 shard test files spawn subprocesses (config-schema.property, core-utils, feat-3881-yaml-parser-consequences) and bail fires on every KILLED mutant, so leaving it on kills processes mid-spawnSync and orphans their children; Stryker's separate disableBail still skips remaining FILES, which is most of the win. tap.nodeArgs and top-level buildCommand stay UNSET so no rebuild lands between mutation and test (ADR-457). Coverage granularity is per FILE, not per test (\"a test is always a test file\"), so the #2790 excludeTests bans on spawn-heavy integration files remain necessary and unchanged", + "line": 612 }, { "id": "RULESET.TESTS.mutation-score", "klass": "RULESET", "value": "Stryker runs incremental (--since origin/next) on ubuntu-latest/Node24 CI leg; default threshold 80% killed/total; surviving mutants in scope block merge unless path is listed in stryker.config.mjs with documented reason; treat surviving mutant as a failing test specification", - "line": 603 + "line": 611 + }, + { + "id": "RULESET.TESTS.mutation-score-denominator", + "klass": "RULESET", + "value": "the gated number is mutation-testing-metrics' mutationScore = totalDetected/totalValid, which counts NoCoverage in the denominator EXACTLY as Survived; both Stryker's own thresholds.break (core dist/src/reporters/mutation-test-report-helper.js) and scripts/check-mutation-score-ratchet.cjs read THAT field, which is what makes the #3915 coverageAnalysis 'off'->'perTest' switch score-neutral. NEVER gate on mutationScoreBasedOnCoveredCode — it EXCLUDES NoCoverage and inflates sharply under perTest (measured on a synthetic report: 8 killed/2 survived = 80 and 80; 8 killed/2 noCoverage = 80 and 100), so swapping to the better-sounding field would make every minScore floor trivially satisfiable and the gate decorative. Under the pre-#3915 coverageAnalysis:'off' the two fields were ALWAYS identical (noCoverage was structurally 0), which is why nothing had ever pinned the choice; tests/mutation-score-ratchet.test.cjs now pins it with a non-vacuity assertion that the two numbers genuinely diverge", + "line": 613 }, { "id": "RULESET.TESTS.no-dead-regex-in-includes", "klass": "RULESET", "value": "src.includes(\"foo.*bar\") is always false — .* is regex metacharacter not wildcard; use new RegExp(...).test(src) or delete", - "line": 590 + "line": 598 }, { "id": "RULESET.TESTS.no-duplicate-fold-marker", "klass": "RULESET", "value": "local/no-duplicate-fold-marker ESLint AST rule (eslint-rules/no-duplicate-fold-marker.cjs, #3271) reports the 2nd and every later __foldDescribe(\"folded: ...\") call carrying a marker already seen in the SAME file, naming the first occurrence's line; error in tests/**/*.cjs. The key is the WHITESPACE-delimited token after folded:, NOT a [a-z0-9-]* slice — a slice truncates at \".\" and collides feat-443-effort-fast-mode.integration with feat-443-effort-fast-mode (two distinct suites coexisting in tests/model-resolver.test.cjs), and NOT the whole title, so a re-fold under a different batch label (\"B1 #1970\" vs \"B5 #1975\") is still caught. Deliberately silent on: a __foldDescribe title with no folded: prefix (the alias is reused for one ordinary describe in tests/review-default-reviewers-workflow.test.cjs), a plain describe(), a non-literal title, and the same marker in two DIFFERENT files (the defect class is intra-file).", - "line": 586 + "line": 594 }, { "id": "RULESET.TESTS.no-duplicate-fold-marker.why", "klass": "RULESET", "value": "consolidation epic #1969 folds are self-contained blocks, so a second verbatim copy parses, registers and PASSES twice — nothing reports it; #3271 found 25 such copies (~5,800 lines) in tests/install.test.cjs (18), tests/install-minimal-hooks.test.cjs (5) and tests/install-write-confinement.test.cjs (2), all from one stale-base re-application in 6d072435d (#1975 re-applying #1970's hunks, 2026-07-03). Ref DEFECT.GENERATIVE-FIX: the two copies drift apart silently when a contributor fixes one and leaves the other asserting the old behavior, with the suite still green.", - "line": 587 + "line": 595 }, { "id": "RULESET.TESTS.no-source-grep", "klass": "RULESET", "value": "local/no-source-grep ESLint AST rule (eslint-rules/no-source-grep.cjs) rejects readFileSync of a source .cjs/.js/.ts path bound to a var later hit with .includes()/.match()/.startsWith()/.endsWith()/.indexOf()/.search(); error in tests/**/*.test.cjs, warn in gsd-core/bin/**/*.cjs + scripts/**/*.cjs (ADR 452 retired the old regex script, removed for good in #632)", - "line": 583 + "line": 591 }, { "id": "RULESET.TESTS.no-source-grep.exemption", "klass": "RULESET", "value": "// allow-test-rule: with one-line justification; reserved for tests where the file content IS the product surface (STATE.md, config.toml, hooks.json, agent .md). Migration to typed-IR parser tracked in #2974.", - "line": 584 + "line": 592 }, { "id": "RULESET.TESTS.no-source-grep.tmp-file-traps", "klass": "RULESET", "value": "reading tmp files written by the SUT in tests still trips lint; round-trip through CLI (e.g. frontmatter get) instead of readFileSync+.includes()", - "line": 585 + "line": 593 }, { "id": "RULESET.TESTS.no-timing-assertion", "klass": "RULESET", "value": "do not assert on wall-clock elapsed time (Date.now() delta, performance.now(), process.hrtime() comparison); such assertions test the host machine not the SUT and flake on loaded CI runners; enforcement: local/no-elapsed-assertion ESLint rule, error (promoted by #3331 once #3314 delivered the ADR-456 §(a) reachability rule + deterministic backfill precondition); canonical replacement: clock-seam pattern with node:test mock.timers", - "line": 600 + "line": 608 }, { "id": "RULESET.TESTS.property-based-testing", "klass": "RULESET", "value": "modules implementing parsing / transformation / budget-limit / bijective contracts must include at least one fast-check (fc) property test asserting a domain invariant; invariant categories: round-trip, monotonicity, boundary-containment, idempotency; property tests live in *.test.cjs alongside unit tests; CI signal: Stryker mutation score below 80% blocks merge", - "line": 602 + "line": 610 }, { "id": "RULESET.TRIAGE-EXISTING-WORK", "klass": "RULESET", "value": "before writing agent brief for confirmed bug, check (1) local branches git branch -a | grep , (2) untracked/modified files on that branch, (3) stash, (4) open PRs with matching head branch — recover existing work rather than re-implement", - "line": 627 + "line": 637 }, { "id": "RULESET.WORKFLOW.COVERAGE-METADATA", "klass": "RULESET", "value": "#1602 SUMMARY frontmatter `coverage:` block (list of {id,description,requirement?,verification:[{kind∈unit|integration|e2e|automated_ui|manual_procedural|other, ref, status∈pass|fail|unknown}],human_judgment:bool,rationale?}) is the per-deliverable RTM consumed DETERMINISTICALLY by verify-work extract_tests via `gsd-tools uat classify-coverage --summary ` (src/coverage.cts → bin/lib/coverage.cjs). AUTHORING: execute-plan create_summary populates it from task results; every deliverable MUST be classified; fail-safe default = human_judgment:true + rationale. CLASSIFY CONTRACT: auto-pass (skip human) ONLY when human_judgment===false (strict boolean) AND verification non-empty AND every status==='pass' AND zero validation errors — else PRESENT to human. mode:legacy (no block) ⇒ byte-identical prose `## Accomplishments` fall-through; `coverage: []` ⇒ mode:coverage, zero entries (single-confirmation). Frozen IR: MODE/PRESENT_REASON/ERROR_CODE enums locked by tests/coverage-metadata-parser.test.cjs. extractFrontmatter CANNOT parse it (scalars-only `-` items) → dedicated parser, sibling of parseMustHavesBlock. Asymmetry by design: false-negative=redundant prompt (status quo); false-positive=shipped bug UAT existed to catch", - "line": 616 + "line": 626 }, { "id": "RULESET.WORKFLOW_EXECUTE_END_TO_END", "klass": "RULESET", "value": "standard for single-workflow commands is \"Execute end-to-end.\" (no bolded **Follow the X workflow** fragments); flag-dispatch routing uses \"execute the X workflow end-to-end.\" in routing bullets — convention verified live across ~20 commands/gsd/*.md files; no ADR currently documents this specific phrasing rule (ADR-0002 covers the adjacent but distinct command-contract/@-ref-resolution seam, not this convention)", - "line": 615 + "line": 625 }, { "id": "RULESET.WORKFLOW_EXECUTION_CONTEXT", "klass": "RULESET", "value": "@-ref in commands/gsd/*.md must resolve to an existing file on disk; regression test in tests/docs-update.test.cjs (folds former \\`bug-3135-capture-backlog-workflow\\`, consolidation epic #1969); INVENTORY.md row + INVENTORY-MANIFEST.json families.workflows must stay in sync; \"Invoked by\" attribution must move when a flag absorbs a micro-skill", - "line": 614 + "line": 624 }, { "id": "RULESET.WORKFLOW_FILE_NAMES", "klass": "RULESET", "value": "workflow files use hyphens; XML attributes must match (extract-learnings not extract_learnings); tests should pin exact hyphenated name", - "line": 613 + "line": 623 }, { "id": "RULESET.WORKFLOW_MARKDOWN.FENCES", "klass": "RULESET", "value": "preserve opening language fence when editing shell snippets in workflow markdown; malformed fence creates fresh CR threads (MD040)", - "line": 609 + "line": 619 }, { "id": "RULESET.WORKFLOW_SIZE_BUDGET", "klass": "RULESET", "value": "workflow size enforcement (#1074; BYTES not lines per #717; LF-normalized per #683) = differential attribution size ratchet (PRIMARY anti-creep since #2724/ADR-2719 §4: tests/emitted-attribution.test.cjs's real-tree test reports growth in any gsd-core/workflows/*.md with its exact byte delta vs `next`, no committed snapshot, requires an ack entry — a fragment under tests/emitted-drift-acks/, #2914; the legacy tests/emitted-drift-ack.json is still honored and unioned in) + loose tier hard caps (outer red lines, NEVER raised on approach: XL<=98304 / LARGE<=61440 / DEFAULT<=40960) + discuss-phase<32000; a file that grew fails the differential guard — add an ack entry naming the file and reason, justify the growth in the PR (or extract LAZILY-loaded content; eager @-imports don't reduce loaded context); crossing a hard cap means EXTRACT, not bump. The prior per-file baseline (tests/workflow-size-baseline.json, `npm run size:baseline`) is REMOVED by #2724. Its new-file cap (ADR-1610 Decision point 3, un-baselined files <=32768, the Codex anchor) is REVIVED inside the differential's size ratchet itself (`NEW_FILE_CAP` in tests/helpers/emitted-diff.cjs) rather than lost: \"not yet baselined\" is exactly \"present in sizeCurrent, absent from sizeBaseline\", a signal the ratchet already computes for its own reasons. NOT ack-able — same as the tier hard caps, the fix is extraction. Narrower than the original: this check cannot see XL/LARGE tiering (tests/workflow-size-budget.test.cjs's classification, invisible to the pure differential module), so a legitimately large NEW file must extract rather than tier in, one release earlier than an existing file would need to — a disclosed, deliberate simplification", - "line": 610 + "line": 620 }, { "id": "SESSION.2026-05-05", "klass": "SESSION", "value": "[PRED.k320..k331 introduced; DEFECT.SOURCE-GREP-IN-NEW-TESTS, DEFECT.CHANGESET-PR-FIELD-DRIFT, DEFECT.PHASE-DIR-PREFIX-DRIFT, DEFECT.PROMPT-INJECTION-SCAN-COLLISION; ADR-0002 thin-wrapper pattern findings folded into RULESET.WORKFLOW_*]", - "line": 878 + "line": 888 }, { "id": "SESSION.2026-05-05.sdk-bridge", "klass": "SESSION", "value": "PR #3158 SDK Runtime Bridge — observability isolation rule; strict-mode dispatchMode reporting invariant; transport decision ordering (guard before event emission); folded into Dispatch Policy Module glossary", - "line": 879 + "line": 889 }, { "id": "SESSION.2026-05-09", "klass": "SESSION", "value": "[8-PR triage wave, 7 merged + 1 subsumed; META.RULE.* introduced; WAVE.LESSON.* captured; k320/k322/k323/k326/k331 evidence; AI Ops Memory predicate format established]", - "line": 880 + "line": 890 }, { "id": "SESSION.2026-05-10", "klass": "SESSION", "value": "[ai-ops memory consolidation; release-notes standard taxonomy + templates; RELEASE-NOTES.* predicates introduced]", - "line": 881 + "line": 891 }, { "id": "SESSION.2026-05-13", "klass": "SESSION", "value": "[Shell Command Projection Module expansion (#3465-#3468); ADR-0009 superseded; new exports for subprocess dispatch and platform file I/O; phase-gated migration plan; PR #3464 three-gate invariant CI+CR+unresolved=0; PR #3470 stash-include-untracked rebase pattern]", - "line": 882 + "line": 892 }, { "id": "SESSION.2026-05-14", "klass": "SESSION", "value": "[#3095/PR #3490 EXEC.CLASSIFY.* introduced (Anthropic/Copilot/Codex/Gemini [runtime removed #1928] cross-runtime rate-limit sentinel coverage); #3489/PR #3499 DEFECT.STATE-TRAMPLE.idempotency-oracle (STATE.md current_phase field is oracle for state.complete-phase); #3488/PR #3501 DAG resolver same-phase short-form depends_on (shortFormToId index added to sdk/src/query/phase.ts); #3491/PR #3502 DEFECT.NESTED-GIT-INIT (gitWorktreeInfoInternal helper); #3493/PR #3500 extractCurrentMilestone generic Phase Details continuation past planned-milestone siblings; #3503/PR #3504 DEFECT.PATH-SUBSTRING-CHECK (trailing-slash anchor for homedir checks); #3346/PR #3505 codex AoT TOML leaf-key via extractFlatHookEventName; #3506/PR #3507 label-scoped stale-bot sub-job pattern; multi-PR triage operational lessons folded into PROC.TRIAGE.*; #3508 DEFECT.AGENT-ISOLATION-SILENT-FAIL; gsd-test image-missing auto-build (locally-built image via embedded heredoc Dockerfile); refined PRED.k322 threshold to 3 PRs/<10min]", - "line": 883 + "line": 893 }, { "id": "SESSION.2026-05-15", "klass": "SESSION", "value": "[#3537/PR #3538 DEFECT.PHASE-REGEX-FANOUT — phaseMarkdownRegexSource promoted to core.cjs and wired to 7 sites; parity-style regression test established as DEFECT.GENERATIVE-FIX exemplar; trek-e/gsd-test-runner#1 filed for DEFECT.GSD-TEST-MIRROR-POISONED — chown-back-before-exec legacy gap (poisoned holodeck mirror unstuck via authorized docker chown to remote 1000:1000); RULESET.PR-FLOW.* codified from project CLAUDE.md load-bearing rule; first dispatch under run-tests-before-create held cleanly (PR #3520 worker stopped on Docker exit 12 infra failure, orchestrator opened PR after unblock); CONTEXT.md refactored from 882 lines of mixed prose+predicates into ~500 lines of pure-predicate format with chronological session log]", - "line": 884 + "line": 894 }, { "id": "SESSION.2026-05-15.parallel-fix-dispatch", "klass": "SESSION", "value": "[#3542/PR #3546 prohibit git stash family in executor agents (shared refs/stash across worktrees); #3541/PR #3547 non-TTY resolution for installer prompt-user actions (default remove for SDK build artifacts, keep for skills/gsd-*/SKILL.md); #3545 filed for gsd-test-summary concurrent /tmp output collision; new predicates DEFECT.HOOK-OVER-ENFORCEMENT.read-tool-tracking, DEFECT.GSD-TEST-CONCURRENT-OUTPUT-COLLISION, DEFECT.SUBAGENT-LONG-RUNNING-BG-STALL, DEFECT.AGENT-RETIRED-SLASH-SYNTAX-DRIFT, PROC.PARALLEL-FIX-DISPATCH; agent-trust-but-verify caught /gsd-update retired-syntax comment slip in #3541 implementation before PR open]", - "line": 885 + "line": 895 }, { "id": "SESSION.2026-05-16", "klass": "SESSION", "value": "[multi-PR triage wave (#3577/3581/3640/3641/3642/3648/3649/3637/3639). Established global PreToolUse hook ~/.claude/hooks/test-memory-guard.sh denying new node/test spawns when sum(RSS of node|vitest|jest|...) >= 4 GiB on the 24 GB Mac OR when a same-runner process is already in argv[0] — hard deny via hookSpecificOutput.permissionDecision=deny. PR #3577 fix: revert config-ensure-section dispatch to CJS cmdConfigEnsureSection (SDK author wrote single-section semantics under a name whose legacy callers expect full-default config init); plus 3 SDK parity carve-outs (configNewProject defaults align with sdk/shared/config-defaults.manifest.json, return relative .planning/config.json path, drop quotes from Unknown config key, lead malformed-JSON error with \"Failed to read config.json:\"). PR #3649 fix: chunk node --test spawn at 28K argv ceiling (Windows CreateProcess lpCommandLine cap 32,767 was instantly aborting unchunked spawn of 546 paths). Chunking fix surfaced 14 pre-existing Windows-only test bugs (4010 pass / 14 fail; vs 0/0 before — entire suite was un-runnable on Windows). PRs #3639 + #3637 confirmed unable to stand alone (legitimately depend on Phase 6 scaffolding only present on feat/3575-enforcement-hardening) — user decision: cherry-pick into #3577 and close. Five other PRs each had ≤1 unresolved CR thread of the changeset-pr-number / null-vs-throw / implicit-Claude-runtime / docs-stale-guidance / hardcoded-tests-path family — all quick wins. New predicates: DEFECT.SDK-PORT-NAME-COLLISION, DEFECT.WINDOWS-ARGV-OVERFLOW, DEFECT.STACKED-PR-CANNOT-STAND-ALONE, DEFECT.CANARY-VERSION-LEAK, DEFECT.GSD-TEST-HOST-MID-RUN-DEATH, RULESET.HARNESS.test-memory-guard, RULESET.PR-FLOW.docker-before-push, RULESET.PR-FLOW.templates-mandatory]", - "line": 886 + "line": 896 }, { "id": "WAVE.LESSON.agent-narrative-unreliable", "klass": "WAVE", "value": "k095/k324 confirmed at scale: 5 of 8 agents terminated mid-monitor with stale claims requiring direct verification", - "line": 840 + "line": 850 }, { "id": "WAVE.LESSON.changelog-policy-violation-multiplier", "klass": "WAVE", "value": "brief contradicting CONTRIBUTING.md's changelog-fragment policy (\"CHANGELOG Entries — Drop a Fragment\" section) produced violations on 5 of 8 PRs (#3300, #3302, #3304, #3305, #3308); k326 + k320 capture", - "line": 837 + "line": 847 }, { "id": "WAVE.LESSON.cr-throttle-burst-correlation", "klass": "WAVE", "value": "8 PRs in <15min triggered k322 sustained-throttle on multiple PRs (#3306 worst case)", - "line": 838 + "line": 848 }, { "id": "WAVE.LESSON.k101-still-trips", "klass": "WAVE", "value": "even after CONTEXT.md k101 reinforcement, agent of record posted self-PR comment on close; k331 adds explicit close-time literal-instruction guard", - "line": 841 + "line": 851 }, { "id": "WAVE.LESSON.sibling-audit-overlap", "klass": "WAVE", "value": "k015-family parallel dispatch on #3297 + #3298 produced k323 add-backlog.md cross-PR overlap", - "line": 839 + "line": 849 }, { "id": "WORKSTREAM.INVARIANT.migrate-name", "klass": "WORKSTREAM", "value": "must normalize through canonical slug policy", - "line": 663 + "line": 673 }, { "id": "WORKSTREAM.INVARIANT.slug-contract", "klass": "WORKSTREAM", "value": "all .planning/workstreams/ must be addressable by set/get/status/complete", - "line": 664 + "line": 674 }, { "id": "WORKSTREAM.NAME.POLICY.cjs-module", "klass": "WORKSTREAM", "value": "gsd-core/bin/lib/workstream-name-policy.cjs owns toWorkstreamSlug + active-name/path-segment validation", - "line": 679 + "line": 689 }, { "id": "WORKSTREAM.POINTER.SEAM.cjs-module", "klass": "WORKSTREAM", "value": "gsd-core/bin/lib/active-workstream-store.cjs owns read/write self-heal for .planning/active-workstream", - "line": 680 + "line": 690 }, { "id": "WORKSTREAM.REGRESSION.test-anchor", "klass": "WORKSTREAM", "value": "tests/workstream.test.cjs::normalizes --migrate-name to a valid workstream slug", - "line": 665 + "line": 675 }, { "id": "WORKTREE.SEAM.caller-rule", "klass": "WORKTREE", "value": "verify.cjs must consume inspectWorktreeHealth for W017 classification; no ad-hoc porcelain parsing in callers", - "line": 673 + "line": 683 }, { "id": "WORKTREE.SEAM.current", "klass": "WORKTREE", "value": "Worktree Safety Policy Module", - "line": 657 + "line": 667 }, { "id": "WORKTREE.SEAM.decision-1", "klass": "WORKTREE", "value": "retain non-destructive default; destructive path only as explicit future opt-in scaffold", - "line": 661 + "line": 671 }, { "id": "WORKTREE.SEAM.default-prune-policy", "klass": "WORKTREE", "value": "metadata_prune_only (non-destructive)", - "line": 660 + "line": 670 }, { "id": "WORKTREE.SEAM.files", "klass": "WORKTREE", "value": "[gsd-core/bin/lib/worktree-safety.cjs]", - "line": 658 + "line": 668 }, { "id": "WORKTREE.SEAM.interface", "klass": "WORKTREE", "value": "[resolveWorktreeContext, parseWorktreePorcelain, planWorktreePrune, executeWorktreePrunePlan, planWorktreeRecordAgent, cmdWorktreeRecordAgent]", - "line": 659 + "line": 669 }, { "id": "WORKTREE.SEAM.invariant", "klass": "WORKTREE", "value": "parser failure must degrade to metadata_prune_only and never escalate to destructive removal", - "line": 671 + "line": 681 }, { "id": "WORKTREE.SEAM.inventory-interface", "klass": "WORKTREE", "value": "[listLinkedWorktreePaths, inspectWorktreeHealth]", - "line": 672 + "line": 682 }, { "id": "WORKTREE.SEAM.inventory-snapshot", "klass": "WORKTREE", "value": "snapshotWorktreeInventory(repoRoot,{staleAfterMs,nowMs}) is canonical linked-worktree health snapshot for callers", - "line": 675 + "line": 685 }, { "id": "WORKTREE.SEAM.test-anchor-w017", "klass": "WORKTREE", "value": "tests/orphan-worktree-detection.test.cjs + tests/worktree-safety.test.cjs", - "line": 674 + "line": 684 }, { "id": "WORKTREE.SEAM.test-anchors", "klass": "WORKTREE", "value": "[resolveWorktreeContext:has_local_planning|linked_worktree|not_git_repo|main_worktree, planWorktreePrune:git_list_failed|worktrees_present|no_worktrees|parser_throw_fallback, executeWorktreePrunePlan:missing_plan|skip_passthrough|unsupported_action|metadata_prune_only]", - "line": 670 + "line": 680 }, { "id": "WORKTREE.SEAM.test-policy", "klass": "WORKTREE", "value": "cover all decision branches in policy module before changing prune behavior", - "line": 669 + "line": 679 } ], "duplicates": [] diff --git a/package-lock.json b/package-lock.json index 3af2815f0..3873506fb 100644 --- a/package-lock.json +++ b/package-lock.json @@ -21,6 +21,7 @@ "devDependencies": { "@eslint/js": "^9.39.4", "@stryker-mutator/core": "^9.6.1", + "@stryker-mutator/tap-runner": "^9.6.1", "@types/node": "^22.19.19", "c8": "^11.0.0", "eslint": "^9.39.4", @@ -1747,6 +1748,26 @@ "node": ">=20.0.0" } }, + "node_modules/@stryker-mutator/tap-runner": { + "version": "9.6.1", + "resolved": "https://registry.npmjs.org/@stryker-mutator/tap-runner/-/tap-runner-9.6.1.tgz", + "integrity": "sha512-b5ryfiRQHH5VoWP++VEA9KYiU6lhVbE9znooFaWRr7umaAtqKmlWrFinKA6fghwiFpdxerRpctLDT8uVKzeQEw==", + "dev": true, + "license": "Apache-2.0", + "dependencies": { + "@stryker-mutator/api": "9.6.1", + "@stryker-mutator/util": "9.6.1", + "glob": "~13.0.0", + "tap-parser": "~17.0.0", + "tslib": "~2.8.0" + }, + "engines": { + "node": ">=14.18.0" + }, + "peerDependencies": { + "@stryker-mutator/core": "9.6.1" + } + }, "node_modules/@stryker-mutator/util": { "version": "9.6.1", "resolved": "https://registry.npmjs.org/@stryker-mutator/util/-/util-9.6.1.tgz", @@ -3010,6 +3031,16 @@ "node": ">= 0.6" } }, + "node_modules/events-to-array": { + "version": "2.0.3", + "resolved": "https://registry.npmjs.org/events-to-array/-/events-to-array-2.0.3.tgz", + "integrity": "sha512-f/qE2gImHRa4Cp2y1stEOSgw8wTFyUdVJX7G//bMwbaV9JqISFxg99NbmVQeP7YLnDUZ2un851jlaDrlpmGehQ==", + "dev": true, + "license": "ISC", + "engines": { + "node": ">=12" + } + }, "node_modules/eventsource": { "version": "3.0.7", "resolved": "https://registry.npmjs.org/eventsource/-/eventsource-3.0.7.tgz", @@ -4870,6 +4901,37 @@ "node": ">=8" } }, + "node_modules/tap-parser": { + "version": "17.0.0", + "resolved": "https://registry.npmjs.org/tap-parser/-/tap-parser-17.0.0.tgz", + "integrity": "sha512-Na7kB4ML7T77abJtYIlXh/aJcz54Azv0iAtOaDnLqsL4uWjU40uNFIFnZ5IvnGTuCIk5M6vjx7ZsceNGc1mcag==", + "dev": true, + "license": "BlueOak-1.0.0", + "dependencies": { + "events-to-array": "^2.0.3", + "tap-yaml": "3.0.0" + }, + "bin": { + "tap-parser": "bin/cmd.cjs" + }, + "engines": { + "node": ">= 18.6.0" + } + }, + "node_modules/tap-yaml": { + "version": "3.0.0", + "resolved": "https://registry.npmjs.org/tap-yaml/-/tap-yaml-3.0.0.tgz", + "integrity": "sha512-qtbgXJqE9xdWqlE520y+vG4c1lgqWrDHN7Y2YcrV1XudLuc2Y5aMXhAyPBGl57h8MNoprvL/mAJiISUIadvS9w==", + "dev": true, + "license": "BlueOak-1.0.0", + "dependencies": { + "yaml": "^2.4.1", + "yaml-types": "^0.3.0" + }, + "engines": { + "node": ">= 18.6.0" + } + }, "node_modules/tapable": { "version": "2.3.3", "resolved": "https://registry.npmjs.org/tapable/-/tapable-2.3.3.tgz", @@ -5298,6 +5360,36 @@ "dev": true, "license": "ISC" }, + "node_modules/yaml": { + "version": "2.9.0", + "resolved": "https://registry.npmjs.org/yaml/-/yaml-2.9.0.tgz", + "integrity": "sha512-2AvhNX3mb8zd6Zy7INTtSpl1F15HW6Wnqj0srWlkKLcpYl/gMIMJiyuGq2KeI2YFxUPjdlB+3Lc10seMLtL4cA==", + "dev": true, + "license": "ISC", + "bin": { + "yaml": "bin.mjs" + }, + "engines": { + "node": ">= 14.6" + }, + "funding": { + "url": "https://github.com/sponsors/eemeli" + } + }, + "node_modules/yaml-types": { + "version": "0.3.0", + "resolved": "https://registry.npmjs.org/yaml-types/-/yaml-types-0.3.0.tgz", + "integrity": "sha512-i9RxAO/LZBiE0NJUy9pbN5jFz5EasYDImzRkj8Y81kkInTi1laia3P3K/wlMKzOxFQutZip8TejvQP/DwgbU7A==", + "dev": true, + "license": "ISC", + "engines": { + "node": ">= 16", + "npm": ">= 7" + }, + "peerDependencies": { + "yaml": "^2.3.0" + } + }, "node_modules/yargs": { "version": "17.7.2", "resolved": "https://registry.npmjs.org/yargs/-/yargs-17.7.2.tgz", diff --git a/package.json b/package.json index e193981a8..67f41bc96 100644 --- a/package.json +++ b/package.json @@ -66,6 +66,7 @@ "devDependencies": { "@eslint/js": "^9.39.4", "@stryker-mutator/core": "^9.6.1", + "@stryker-mutator/tap-runner": "^9.6.1", "@types/node": "^22.19.19", "c8": "^11.0.0", "eslint": "^9.39.4", diff --git a/scripts/mutation-matrix.cjs b/scripts/mutation-matrix.cjs index a572a2a74..e04567c91 100644 --- a/scripts/mutation-matrix.cjs +++ b/scripts/mutation-matrix.cjs @@ -31,6 +31,7 @@ const { execFileSync } = require('child_process'); const fs = require('fs'); +const path = require('node:path'); const { ExitError, runMain } = require('./lib/cli-exit.cjs'); @@ -92,11 +93,14 @@ function readStdinSync() { // confirmed equivalent mutant is acceptable. // // HOW TO UPDATE: -// 1. The per-module Stryker shard CANNOT be run locally: Stryker's command -// runner invokes `node --test` once per mutant (see stryker.config.mjs), -// and this repo hard-blocks local `node --test` via -// .claude/hooks/block-local-node-test.sh. Push the branch instead and -// let CI run the shard for the changed module. +// 1. The per-module Stryker shard CANNOT be run locally: Stryker's tap +// runner (see stryker.config.mjs) spawns +// `node --test-reporter=tap -r ` once per covering test +// file per mutant, and .claude/hooks/block-local-node-test.sh's matcher +// still denies that form — its pattern `node(\s+-\S+)*\s+--test([\s=-]|$)` +// matches `--test-reporter` because `-` is in the trailing character +// class. Push the branch instead and let CI run the shard for the +// changed module. // 2. Read the measured score from the CI shard's output. // 3. Set minScore = floor(measured) - 1 (never lower than current value) // and update the matching RATCHET_BASELINE entry in the same diff. @@ -168,7 +172,7 @@ const TARGET_MUTATION_SCORE = 80; // named in extraTests, or named in excludeTests — so a file that newly starts // matching the naming rule (like #3888's four files would have, had they been // named `frontmatter*`) cannot silently fall through the cracks again. -const TESTS_DIR = require('node:path').join(__dirname, '..', 'tests'); +const TESTS_DIR = path.join(__dirname, '..', 'tests'); let _testRequireCache = null; /** @@ -191,7 +195,7 @@ function scanTestRequires() { entries = []; } for (const file of entries) { - const text = fs.readFileSync(require('node:path').join(TESTS_DIR, file), 'utf8'); + const text = fs.readFileSync(path.join(TESTS_DIR, file), 'utf8'); let m; REQUIRE_RE.lastIndex = 0; while ((m = REQUIRE_RE.exec(text))) { @@ -403,27 +407,26 @@ const COVERED = { 'frontmatter.test.cjs', ], minScore: 65, - // Wall-time projection, re-derived after piece 1 (dropping frontmatter.test.cjs) using - // this file's own documented method. Mutant-count factor is unchanged: source grew 1.8x - // for #3881 (1030 -> ~1850 mutants; see the #3888-era note this superseded for that - // derivation). Per-run test-command cost is re-measured on the CURRENT 6-file derived - // set (frontmatter.property/.unit/-golden-parity/-roundtrip.property + unusable-input + - // feat-3881-yaml-parser-consequences — the shard minus frontmatter.test.cjs and minus - // frontmatter-cli.test.cjs, neither of which was ever in a measured baseline): 1520ms, - // vs the documented OLD 3-file baseline of 593ms — a 2.56x per-run cost increase (down - // from the pre-piece-1 8x, since the file responsible for 3132ms of the old 4669ms - // 7-file run is gone). Applying both factors the same way the prior note did: 586s - // (documented 3-file/1030-mutant CI baseline) * 1.8 (mutants) * 2.56 (test cost) ~= - // 2700s (~45 minutes). Set to 60 minutes for margin above that projection (the same - // ~1.3x margin ratio the prior 180-minute budget used over its own 140-minute - // projection), well under GitHub Actions' 360-minute job ceiling and a 3x cut from the - // previous 180. Scoped to this shard only via timeoutMinutes below — every other shard - // keeps the 15-minute default. - timeoutMinutes: 60, - // isolation: intentionally NOT set (defaults to 'process' below) — unchanged from the - // prior audit: 'none' showed no reliable win once the test set grew past 3 files - // (overlapping distributions), and dropping frontmatter.test.cjs only shrinks the set - // further, so there is no new basis to revisit that call. + // MEASUREMENT, not a projection. Under the tap runner with coverageAnalysis: 'perTest' + // (#3915), Stryker now re-runs only the test files that cover each mutated line instead + // of all six files for every one of ~1900 mutants. Measured result: the frontmatter shard + // completed in 713s (11m53s) — GitHub Actions run 33026833181, job wall time including + // checkout and `npm ci` — against 1751s (29m11s) on the command runner in run + // 33021042847. A 59% reduction. + // 20 minutes is 1.68x the measured 713s. The override is not simply deleted because the + // shared default is 15 minutes, which 11m53s would fit inside — but only at 79% of + // budget — and this module's mutant count grew 1.8x in a single change (#3881), so a + // shard sitting at 79% of the shared default is one growth spurt from a red lane. 20 + // keeps a real margin while still cutting the previous 60-minute budget by 3x. + // Scoped to this shard only via timeoutMinutes below — every other shard keeps the + // 15-minute default emitted by buildResult(), well under GitHub Actions' 360-minute job + // ceiling. + timeoutMinutes: 20, + // isolation: no knob left to tune (#3915). Per-file process isolation is now INHERENT + // to @stryker-mutator/tap-runner — it drives Node's own `--test-reporter=tap` once per + // covering test FILE, so every file already runs in its own process by construction. + // The prior audit's 'none' vs 'process' comparison (recorded here before this change) + // is moot: there is nothing left to opt in or out of. }, // adr-parser / config-schema / active-workstream-store / core-utils: derivation reproduces // their prior hand lists exactly (every constraining file's own name already matched the @@ -534,13 +537,9 @@ const COVERED = { // per mutant). tests/model-catalog.unit.test.cjs is spawn-free, in-process, // and runs in well under a second. // - // Measured CI score (GitHub Actions run 32605073352, job 97108869486): + // Prior context (#3007, GitHub Actions run 32605073352, job 97108869486): // model-catalog 59.62% → floor 58 (248 killed, 168 survived, 0 timeouts, - // 0 errors; below TARGET_MUTATION_SCORE (80) — ratchet candidate like - // planning-inspect (56): comfortably clears its own floor but has real - // room to grow. Raise as its tests improve, never lower it.) - // Floor follows this file's documented rule, minScore = floor(measured) - 1, - // matching the sibling precedent exactly (57.03 → 56, 76.58 → 75, 95.65 → 94). + // 0 errors). SUPERSEDED by the 2026-08-27 measurement below. // // The shard completed in 57 seconds — concrete evidence the spawn-free // unit-file design above worked: the #2790 precedent's 15-minute shard-cap @@ -551,9 +550,31 @@ const COVERED = { // directly require model-catalog.cjs and match the "model-catalog*" naming rule. Measured // cost: 50ms (1-file) -> 196ms (3-file), in-process, 0 subprocess spawns; still far under // the 57s the shard already measured for the single-file set. + // + // CI run 33029755081 (2026-08-27, #3915): measured 75.26% (295 killed / 49 survived / + // 48 no-coverage / 24 runtime-error, totalValid 392). Floor = floor(75.26) - 1 = 74, + // following this file's documented convention. + // + // Why it moved so far: under #3915's tap-runner swap this shard first came back at + // 57.91% against the old floor of 58. Diagnosis from the mutation report's own JSON: + // all 24 of its RuntimeError mutants are in the module's load-time catalog bootstrap, + // so mutating them makes model-catalog.cjs throw at `require`. Under `node --test` + // that is a failed test file and the mutant counts as Killed; under the tap runner + // the process dies before emitting TAP, which Stryker classifies RuntimeError and + // EXCLUDES from the denominator. Adding those 24 back as killed reproduces + // 248/416 = 59.62 exactly — the pre-swap #3007 number — so detection never + // regressed, only its classification changed. + // + // The floor was NOT lowered to absorb that. 11 new behavioural tests in + // tests/model-catalog.unit.test.cjs killed 68 previously-surviving mutants (227 -> + // 295 killed), taking the module from 57.91 to 75.26 — now within striking distance + // of TARGET_MUTATION_SCORE (80) instead of the 59.62 it sat at before this change. + // + // The floor MUST come from a CI shard, never a local run — same rule as every other + // entry in this file. 'model-catalog': { cjs: 'gsd-core/bin/lib/model-catalog.cjs', - minScore: 58, + minScore: 74, }, // state-contract: net-new module from #3227. Without this entry the // Stryker gate reports has_work: "false" and SKIPS it entirely — the @@ -696,14 +717,10 @@ function buildResult(moduleNames) { mutate: COVERED[name].cjs, tests: COVERED[name].tests.join(' '), minScore: COVERED[name].minScore, - // node:test's default per-file process isolation; only modules that document a - // measured, audited need for 'none' opt out. - isolation: COVERED[name].isolation || 'process', // Per-shard CI job timeout in minutes. Defaults to 15 (the shared per-shard budget); // only a module that documents a measured need for more (see the frontmatter entry // above) sets a higher value. Threaded through mutation.yml's job-level - // `timeout-minutes: ${{ matrix.timeoutMinutes }}` the same way `isolation` is threaded - // through the test-runner env. + // `timeout-minutes: ${{ matrix.timeoutMinutes }}`. timeoutMinutes: COVERED[name].timeoutMinutes || 15, })); @@ -724,7 +741,6 @@ function printHuman(result, changedFiles) { console.log(` mutate: ${shard.mutate}`); console.log(` tests: ${shard.tests}`); console.log(` minScore: ${shard.minScore}`); - console.log(` isolation:${shard.isolation}`); console.log(` timeoutMinutes:${shard.timeoutMinutes}`); } } @@ -786,12 +802,123 @@ function resolveMutationBreak(raw) { return n; } +/** + * Sorted, de-duplicated union of every COVERED module's `tests` array. This + * is the tap-runner's local/full-run default (see resolveMutationTestFiles + * below) — the same union stryker.config.mjs's since-removed DEFAULT_TEST_CMD + * string used to build for the command runner. + * + * @returns {string[]} + */ +function allCoveredTests() { + return [...new Set(Object.values(COVERED).flatMap((mod) => mod.tests))].sort(); +} + +// ── MUTATION_TEST_FILES resolver ────────────────────────────────────────────── +/** + * Resolves the per-shard tap-runner test-file list from the MUTATION_TEST_FILES + * env var. Fail-closed twin of resolveMutationBreak above, for #3915's swap from + * Stryker's `command` runner to `@stryker-mutator/tap-runner`: the tap runner + * takes an explicit `tap.testFiles` array rather than a shell command, so there + * is no single string to inject a per-shard test list into — this function is + * that injection point instead. + * + * Fail-closed contract: + * - undefined → allCoveredTests() (local run: no env set, documented backstop) + * - non-string → throws (the realistic caller mistake: COVERED[*].tests is an + * array, but the env var this reads is the SPACE-JOINED STRING form of it — + * passing the array itself, or any other non-string, is a wiring bug) + * - set but empty/whitespace-only → throws (CI shard wiring is broken: + * matrix.tests missing) + * - otherwise → trim, split on whitespace, de-duplicate, sort, and, per + * entry: (1) resolve it against the repo root and reject any entry whose + * resolved path escapes the repo root (e.g. via `../` segments); (2) + * reject any entry that does not exist on disk, or that exists but is + * not a regular file (e.g. names a directory) — each failure throws + * naming the offending entry(ies) + * + * This function is the single call site for reading MUTATION_TEST_FILES. + * stryker.config.mjs imports and calls it, so a bad value must fail + * immediately rather than silently degrade: the tap runner's own + * `findTestyLookingFiles` resolves `tap.testFiles` via `glob()`, and a + * non-matching glob pattern yields an EMPTY list SILENTLY — which would + * produce a fast, confident, meaningless mutation run (every mutant reported + * killed or survived against zero tests) instead of a loud error. + * + * The `undefined` branch also runs the same on-disk existence check as every + * other branch, so a stale `extraTests`/`excludeTests` entry in COVERED fails + * loudly here rather than silently producing a shard pointed at a phantom file. + * + * @param {string|undefined} raw - value of process.env.MUTATION_TEST_FILES + * @returns {string[]} sorted, de-duplicated, existence-checked test file paths + */ +function resolveMutationTestFiles(raw) { + let entries; + if (raw === undefined) { + // Local run with no MUTATION_TEST_FILES set — use the derived full-run default. + entries = allCoveredTests(); + } else if (typeof raw !== 'string') { + throw new Error( + `MUTATION_TEST_FILES must be a string (space-joined test file paths), got ${typeof raw} — ` + + "COVERED[*].tests is an array internally, but the env var this reads is always the " + + 'SPACE-JOINED STRING form of it; passing the array (or any other non-string) directly is a wiring bug' + ); + } else if (raw.trim() === '') { + throw new Error( + 'MUTATION_TEST_FILES is set but empty — CI shard wiring is broken (matrix.tests missing?)' + ); + } else { + entries = [...new Set(raw.trim().split(/\s+/))].sort(); + } + + const repoRoot = path.join(__dirname, '..'); + + const escaped = []; + const missing = []; + const notFile = []; + for (const entry of entries) { + const resolved = path.resolve(repoRoot, entry); + const rel = path.relative(repoRoot, resolved); + if (rel === '' || rel.startsWith('..') || path.isAbsolute(rel)) { + escaped.push(entry); + continue; + } + if (!fs.existsSync(resolved)) { + missing.push(entry); + continue; + } + if (!fs.statSync(resolved).isFile()) { + notFile.push(entry); + } + } + + if (escaped.length > 0) { + throw new Error( + `MUTATION_TEST_FILES names ${escaped.length} entry(ies) that escape the repo root: ${escaped.join(', ')}` + ); + } + if (missing.length > 0) { + throw new Error( + `MUTATION_TEST_FILES names ${missing.length} file(s) that do not exist on disk: ${missing.join(', ')}` + ); + } + if (notFile.length > 0) { + throw new Error( + `MUTATION_TEST_FILES names ${notFile.length} entry(ies) that are not a regular file: ${notFile.join(', ')}` + ); + } + + return entries; +} + // Export internals for programmatic use (tests/mutation-matrix-ratchet.test.cjs). // The require.main guard prevents main() from running when this file is require()d. module.exports = { COVERED, TARGET_MUTATION_SCORE, resolveMutationBreak, + allCoveredTests, + resolveMutationTestFiles, readStdinSync, // Derivation-engine internals — exported for tests/mutation-test-derivation-drift.test.cjs // and scripts/lint-mutation-test-derivation-drift.cjs. diff --git a/stryker.config.mjs b/stryker.config.mjs index fdaa7363a..55dfde085 100644 --- a/stryker.config.mjs +++ b/stryker.config.mjs @@ -3,13 +3,15 @@ * * Mutation testing configuration for gsd-core. * - * Test runner: 'command' (built into @stryker-mutator/core) - * Runs: node --test over the lib test files via the repo's run-tests invocation. + * Test runner: 'tap' (@stryker-mutator/tap-runner) + * Runs: node --test-reporter=tap over each per-shard test file (tap.testFiles, + * resolved by scripts/mutation-matrix.cjs's resolveMutationTestFiles), one + * process per covering test file per mutant. * * Mutate scope: bin/lib/**\/*.cjs, excluding generated files and test files. * - * coverageAnalysis: 'off' — command runner does not support per-mutant coverage - * thresholds: high=80, low=60, break=50 + * coverageAnalysis: 'perTest' — the tap runner supports per-mutant coverage + * thresholds: high=80, low=60, break=per-shard MUTATION_BREAK (local fallback 60) * incremental: true — caches results; PR-scoped runs pass --mutate * * Reports: @@ -25,12 +27,11 @@ import { createRequire } from 'node:module'; const _require = createRequire(import.meta.url); // resolveMutationBreak: fail-closed resolver for MUTATION_BREAK env var. // undefined → 60 (local backstop); set-but-empty or non-numeric → throws. -// COVERED: the same derived-tests registry CI's per-shard MUTATION_TEST_CMD is built from -// (mutation.yml's `matrix.tests`) — DEFAULT_TEST_CMD below derives from it too, rather than -// hand-duplicating the union in a second literal (#3881 follow-up, mutation-matrix piece 2: -// this exact split — six modules' tests missing from the old hand-written DEFAULT_TEST_CMD, -// plus two entries belonging to no module — is what this derivation removes). -const { resolveMutationBreak, COVERED } = _require('./scripts/mutation-matrix.cjs'); +// resolveMutationTestFiles: fail-closed resolver for MUTATION_TEST_FILES env var (#3915). +// undefined → the derived union of every COVERED module's tests (local backstop); +// set-but-empty, non-string, or naming a nonexistent file → throws. See that +// function's doc-comment in scripts/mutation-matrix.cjs for the full contract. +const { resolveMutationBreak, resolveMutationTestFiles } = _require('./scripts/mutation-matrix.cjs'); // ADR-457: bin/lib/*.cjs are gitignored build artifacts (compiled from // src/*.cts by `npm run build:lib`, which the mutation CI job runs via `npm ci` @@ -64,29 +65,28 @@ const UNMUTATED = [ '!gsd-core/bin/lib/gsd2-import.cjs', ]; -// Full test command used by local runs and as the fallback when CI does not inject a -// per-shard command via MUTATION_TEST_CMD. DERIVED — never hand-edit this list; it is the -// sorted union of every COVERED module's `tests` array (itself derived by -// scripts/mutation-matrix.cjs's computeModuleTests from direct `require()`s of each -// module's built artifact — see that file's derivation-engine header). A previous -// hand-maintained literal here drifted independently of COVERED's own hand-written `tests` -// arrays: six modules' tests were missing from it, plus two entries (broken-windows.test.cjs, -// complexity-trigger.test.cjs) belonging to no COVERED module at all. Deriving both from the -// same COVERED object makes that class of drift structurally impossible. -const DEFAULT_TEST_CMD = `node --test ${ - [...new Set(Object.values(COVERED).flatMap((mod) => mod.tests))].sort().join(' ') -}`; - /** @type {import('@stryker-mutator/core').PartialStrykerOptions} */ export default { // ── Test runner ────────────────────────────────────────────────────────────── - testRunner: 'command', - commandRunner: { - // Run property + unit tests over lib only (avoids the slow integration - // suite). NO build step here: Stryker mutates the already-built .cjs and the - // tests load it directly — adding a build would rebuild over the mutation. - // In CI each matrix shard injects MUTATION_TEST_CMD with only its own tests. - command: process.env.MUTATION_TEST_CMD || DEFAULT_TEST_CMD, + testRunner: 'tap', + tap: { + testFiles: resolveMutationTestFiles(process.env.MUTATION_TEST_FILES), + // forceBail is OFF (MEASURED, #3915): a structural AST audit of all 26 shard test + // files found 3 that spawn subprocesses — tests/config-schema.property.test.cjs + // (6 runGsdTools calls), tests/core-utils.test.cjs (1), and + // tests/feat-3881-yaml-parser-consequences.test.cjs (2 runGsdTools + 2 runNode). + // @stryker-mutator/tap-runner's docs warn that with forceBail on, a runner that + // spawns child processes can be terminated prematurely — bail fires on every KILLED + // mutant (i.e. most of them), so leaving it on would kill hundreds of processes + // mid-spawnSync per shard and orphan their children. The cost of leaving it off is + // only INTRA-file early exit: Stryker's separate `disableBail` (unset, default false) + // still skips the remaining FILES in tap.testFiles after a failure, so this is no + // worse than the command runner's behaviour was. + forceBail: false, + // No nodeArgs, and no top-level buildCommand (ADR-457): the plugin's default argv, + // ["--test-reporter=tap", "-r", "{{hookFile}}", "{{testFile}}"], contains no build + // step, and ADR-457 requires Stryker to test the ALREADY-BUILT gsd-core/bin/lib/*.cjs + // artifacts with no rebuild between mutation and test. }, // ── Files to mutate ────────────────────────────────────────────────────────── @@ -99,8 +99,18 @@ export default { ], // ── Coverage ───────────────────────────────────────────────────────────────── - // 'off' is required for the command test runner — it cannot instrument per-mutant. - coverageAnalysis: 'off', + // 'perTest' (#3915): 'off' was required only because the command runner could not + // instrument per-mutant coverage. @stryker-mutator/tap-runner is an official plugin + // that does support it, so Stryker now re-runs only the test files that cover each + // mutated line rather than the full per-shard test list for every mutant. + // + // Arithmetic note: 'perTest' makes Stryker able to report a mutant as NoCoverage, and + // NoCoverage counts in the SAME denominator as Survived for `mutationScore` — so this + // reclassification is score-neutral for both `thresholds.break` and + // scripts/check-mutation-score-ratchet.cjs. Never gate on + // `mutationScoreBasedOnCoveredCode`, which EXCLUDES NoCoverage and inflates sharply + // under this setting. + coverageAnalysis: 'perTest', // ── Thresholds ─────────────────────────────────────────────────────────────── // ADR-456 / issue #1187: CI passes the per-module minScore (from diff --git a/tests/model-catalog.unit.test.cjs b/tests/model-catalog.unit.test.cjs index f02aa49f5..c7de8d8ac 100644 --- a/tests/model-catalog.unit.test.cjs +++ b/tests/model-catalog.unit.test.cjs @@ -41,11 +41,17 @@ const { CODEX_MODEL_EFFORT, MODEL_ALIAS_MAP, PROVIDER_PRESETS, + AGENT_TO_PHASE_TYPE, + AGENT_DEFAULT_TIERS, + RUNTIME_PROFILE_MAP, + EFFORT_RENDERING, + EFFORT_ARGV, isAnthropicFlavoredModel, getAgentToModelMapForProfile, formatAgentToModelMapAsTable, renderEffortArgv, renderEffortForRuntime, + clampEffortForHost, nextTier, mergeEffortTierDefaults, } = catalog; @@ -378,4 +384,128 @@ describe('model-catalog: mergeEffortTierDefaults (#3531)', () => { mergeEffortTierDefaults(manifest, { sonnet: 'opus' }, isValid); assert.deepEqual(manifest, original); }); + + // #3915 — a plain object literal's `__proto__: value` key is a prototype + // setter, not an own-enumerable property, so `Object.entries()` never + // yields it and the house-pollution guard's `tier === '__proto__'` branch + // is unreachable via a normal test fixture. `Object.fromEntries` uses + // CreateDataPropertyOrThrow internally and DOES create a genuine own + // property literally named "__proto__", which `Object.entries()` then + // yields — the only way to actually drive `tier` to that value and + // exercise the guard's first disjunct. A permissive `isValid` (accepts + // any value) is required here specifically so a disabled guard would let + // the object-valued override reach `merged['__proto__'] = value`, which + // (unlike a string value) really does repoint the prototype. + test('__proto__ pollution guard fires even for a genuine own-enumerable "__proto__" key', () => { + const permissive = () => true; + const overrideProto = Object.fromEntries([['__proto__', { polluted: true }]]); + const merged = mergeEffortTierDefaults({}, overrideProto, permissive); + assert.strictEqual(Object.getPrototypeOf(merged), Object.prototype); + assert.strictEqual(merged.polluted, undefined); + }); +}); + +// #3915 — mutation-score restoration (Stryker survivors in model-catalog.cjs). +// Each test below targets specific surviving mutants identified from the +// mutation report; see the PR/issue for the full mutant-to-test mapping. +describe('model-catalog: module load shape (#3915)', () => { + test('the compiled module is flagged __esModule (defineProperty descriptor, not a loose object literal)', () => { + assert.strictEqual(catalog.__esModule, true); + }); +}); + +describe('model-catalog: AGENT_TO_PHASE_TYPE / AGENT_DEFAULT_TIERS (#3915)', () => { + test('AGENT_TO_PHASE_TYPE maps a known agent to its exact catalog phaseType', () => { + assert.equal(AGENT_TO_PHASE_TYPE['gsd-planner'], 'planning'); + assert.equal(AGENT_TO_PHASE_TYPE['gsd-verifier'], 'verification'); + }); + + test('AGENT_DEFAULT_TIERS maps a known agent to its exact catalog routingTier', () => { + assert.equal(AGENT_DEFAULT_TIERS['gsd-planner'], 'heavy'); + assert.equal(AGENT_DEFAULT_TIERS['gsd-codebase-mapper'], 'light'); + }); +}); + +describe('model-catalog: RUNTIME_PROFILE_MAP filtering (#3915)', () => { + // The catalog's runtimeTierDefaults has runtimes whose opus/sonnet/haiku + // entries are ALL null (e.g. 'cline') as a deliberate "no defaults yet" + // sentinel, and runtimes fully populated (e.g. 'claude'). This pair is + // the exact boundary the filter's `Object.keys(filtered).length > 0` + // check exists for. + test('a runtime whose tier entries are all null is dropped entirely', () => { + assert.equal('cline' in RUNTIME_PROFILE_MAP, false); + assert.equal('kimi' in RUNTIME_PROFILE_MAP, false); + }); + + test('a fully-populated runtime keeps every tier entry, unmodified', () => { + assert.deepStrictEqual(RUNTIME_PROFILE_MAP.claude, { + opus: { model: 'claude-opus-4-8' }, + sonnet: { model: 'claude-sonnet-5' }, + haiku: { model: 'claude-haiku-4-5' }, + }); + }); +}); + +describe('model-catalog: EFFORT_RENDERING / EFFORT_ARGV supported sets (#3915)', () => { + test('EFFORT_RENDERING.claude.supported is exactly low/medium/high/xhigh/max', () => { + assert.deepStrictEqual(EFFORT_RENDERING.claude.supported, new Set(['low', 'medium', 'high', 'xhigh', 'max'])); + }); + + test('EFFORT_RENDERING.codex.supported is exactly low/medium/high/xhigh/max', () => { + assert.deepStrictEqual(EFFORT_RENDERING.codex.supported, new Set(['low', 'medium', 'high', 'xhigh', 'max'])); + }); + + test("EFFORT_RENDERING.codex.clamp maps 'minimal' to 'low' and leaves other levels unchanged", () => { + assert.equal(EFFORT_RENDERING.codex.clamp('minimal'), 'low'); + assert.equal(EFFORT_RENDERING.codex.clamp('high'), 'high'); + }); + + test('EFFORT_ARGV.claude.supported is exactly low/medium/high/xhigh/max', () => { + assert.deepStrictEqual(EFFORT_ARGV.claude.supported, new Set(['low', 'medium', 'high', 'xhigh', 'max'])); + }); + + test('EFFORT_ARGV.opencode.supported additionally includes minimal', () => { + assert.deepStrictEqual(EFFORT_ARGV.opencode.supported, new Set(['minimal', 'low', 'medium', 'high', 'xhigh', 'max'])); + }); + + test('EFFORT_ARGV.codex.supported is exactly low/medium/high/xhigh/max', () => { + assert.deepStrictEqual(EFFORT_ARGV.codex.supported, new Set(['low', 'medium', 'high', 'xhigh', 'max'])); + }); +}); + +describe('model-catalog: formatAgentToModelMapAsTable — exact column widths (#3915)', () => { + // Existing coverage only used `.includes()` on longer-than-header + // agent/model names, which never exercises `Math.max('Agent'.length, ...)` + // vs `Math.min` (padEnd never truncates, so a too-small width is + // invisible to a substring check). Short entries plus a full-string + // comparison make the width computation and the +2/-2 separator padding + // observable. + test('short agent/model names still pad to the header width, and the separator is exactly width+2 wide', () => { + const out = formatAgentToModelMapAsTable({ ab: 'xy' }); + const expected = ` Agent │ Model\n${'─'.repeat(7)}┼${'─'.repeat(7)}\n ab │ xy \n`; + assert.equal(out, expected); + }); +}); + +describe('model-catalog: clampEffortForHost (#3915)', () => { + test('non-string host, unknown host, non-string effort, and empty effort all return null; valid input clamps', () => { + assert.equal(clampEffortForHost(123, 'high'), null); + assert.equal(clampEffortForHost('totally-bogus-host', 'high'), null); + assert.equal(clampEffortForHost('claude', 123), null); + assert.equal(clampEffortForHost('claude', ''), null); + assert.equal(clampEffortForHost('claude', 'minimal'), 'low'); + assert.equal(clampEffortForHost('claude', 'high'), 'high'); + }); + + // The own-property guard must reject a host BEFORE any property lookup + // that could be coerced into a real key. `typeof host !== 'string'` short- + // circuits the `||` for a non-string host, so `EFFORT_ARGV[host]` is never + // reached even when the host's `toString()` would resolve to a real key — + // proving the type check (not just the hasOwnProperty check) is load- + // bearing, and that neither the `||` nor either disjunct can be disabled. + test('a non-string host is rejected even when it stringifies to a known key', () => { + const spoofedHost = { toString: () => 'claude' }; + assert.equal(clampEffortForHost(spoofedHost, 'high'), null); + assert.equal(clampEffortForHost('claude', 'high'), 'high'); + }); }); diff --git a/tests/mutation-matrix-ratchet.test.cjs b/tests/mutation-matrix-ratchet.test.cjs index 94dab8fba..bea6e5911 100644 --- a/tests/mutation-matrix-ratchet.test.cjs +++ b/tests/mutation-matrix-ratchet.test.cjs @@ -23,13 +23,24 @@ const { test, describe } = require('node:test'); const assert = require('node:assert/strict'); const path = require('node:path'); +const fs = require('node:fs'); +const fc = require('fast-check'); const { runNode } = require('./helpers/process-seam.cjs'); const { throwIfFailed } = require('./helpers/git-fixture.cjs'); const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs'); +const REPO_ROOT = path.resolve(__dirname, '..'); const MATRIX_SCRIPT = path.resolve(__dirname, '../scripts/mutation-matrix.cjs'); const matrix = require(MATRIX_SCRIPT); +// Deliberately RE-DERIVED here from matrix.COVERED, rather than calling the +// exported matrix.allCoveredTests() — resolveMutationTestFiles(undefined)'s +// production implementation itself calls allCoveredTests(), so comparing its +// output against that same function would make the assertion a tautology +// that passes no matter what either function does. This independent +// derivation is what makes the comparison test the production code. +const derived = [...new Set(Object.values(matrix.COVERED).flatMap((mod) => mod.tests))].sort(); + // ── (a) TARGET_MUTATION_SCORE ───────────────────────────────────────────────── describe('mutation-matrix ratchet: TARGET_MUTATION_SCORE export', () => { test('exports TARGET_MUTATION_SCORE', () => { @@ -202,7 +213,10 @@ const RATCHET_BASELINE = { 'planning-inspect': 56, // CI run 32392791843: 57.03% (unit shard); ratchet candidate vs TARGET 80 'plan-document': 75, // CI run 32392791843: 76.58% (unit shard) 'planning-command-router': 94, // CI run 32392791843: 95.65% (unit shard); already exceeds TARGET 80 - 'model-catalog': 58, // #3007: measured 59.62% in CI (248 killed / 168 survived); floor(59.62)-1 + 'model-catalog': 74, // CI run 33029755081 (2026-08-27, #3915): measured 75.26%; + // floor(75.26)-1. Supersedes the #3007 59.62% measurement + // (248 killed / 168 survived) — see scripts/mutation-matrix.cjs + // for the RuntimeError classification-change diagnosis. 'state-contract': 65, // #3227: CI run 32769289750, job 97565813640, // `Stryker (state-contract)`: measured 66.25%; floor(66.25)-1. // Below TARGET_MUTATION_SCORE (80) — ratchet candidate like @@ -305,3 +319,204 @@ describe('resolveMutationBreak: fail-closed env-var resolver', () => { assert.strictEqual(resolveMutationBreak('62'), 62); }); }); + +// ── (g) resolveMutationTestFiles behaviour (#3915: tap-runner test-file resolver) ──── +// stryker's tap-runner has no single `command` to inject a per-shard test list into — +// it takes an explicit `tap.testFiles` array — so the derived-union default that used to +// live only in stryker.config.mjs's DEFAULT_TEST_CMD string needs a resolver twin to +// resolveMutationBreak: same fail-closed shape, same single call site, this time +// producing a file-list array instead of a numeric threshold. +describe('resolveMutationTestFiles: fail-closed test-file-list resolver', () => { + const { resolveMutationTestFiles } = matrix; + + test('exports resolveMutationTestFiles as a function', () => { + assert.strictEqual(typeof resolveMutationTestFiles, 'function'); + }); + + test('undefined → sorted, de-duplicated union of every COVERED[*].tests', () => { + assert.deepStrictEqual(resolveMutationTestFiles(undefined), derived); + }); + + test('every entry of the undefined-default result exists on disk', () => { + for (const entry of resolveMutationTestFiles(undefined)) { + assert.ok( + fs.existsSync(path.resolve(REPO_ROOT, entry)), + `derived default test file does not exist on disk: ${entry}` + ); + } + }); + + test('single test file (boundary 1) → one-element array', () => { + assert.deepStrictEqual( + resolveMutationTestFiles('tests/frontmatter.unit.test.cjs'), + ['tests/frontmatter.unit.test.cjs'] + ); + }); + + test('two space-separated test files → two-element array', () => { + assert.deepStrictEqual( + resolveMutationTestFiles('tests/frontmatter.unit.test.cjs tests/unusable-input.test.cjs'), + ['tests/frontmatter.unit.test.cjs', 'tests/unusable-input.test.cjs'] + ); + }); + + test("'' (empty string) → throws (CI shard wiring error)", () => { + assert.throws(() => resolveMutationTestFiles(''), /set but empty/); + }); + + test("' ' (whitespace-only) → throws (CI shard wiring error)", () => { + assert.throws(() => resolveMutationTestFiles(' '), /set but empty/); + }); + + test("'\\t\\n ' (mixed whitespace-only) → throws (CI shard wiring error)", () => { + assert.throws(() => resolveMutationTestFiles('\t\n '), /set but empty/); + }); + + test('null → throws', () => { + assert.throws(() => resolveMutationTestFiles(null)); + }); + + test('123 (non-string) → throws', () => { + assert.throws(() => resolveMutationTestFiles(123)); + }); + + test('an array (not a string) → throws', () => { + assert.throws(() => resolveMutationTestFiles(['tests/frontmatter.unit.test.cjs'])); + }); + + test('leading/trailing whitespace is trimmed to exactly one entry with no surrounding whitespace', () => { + const result = resolveMutationTestFiles(' tests/frontmatter.unit.test.cjs '); + assert.strictEqual(result.length, 1); + assert.strictEqual(result[0], result[0].trim()); + assert.strictEqual(result[0], 'tests/frontmatter.unit.test.cjs'); + }); + + test('mixed tab/newline/space separators → exactly 3 non-empty entries', () => { + const result = resolveMutationTestFiles( + 'tests/frontmatter.unit.test.cjs\t\ttests/unusable-input.test.cjs\n tests/frontmatter.property.test.cjs' + ); + assert.strictEqual(result.length, 3); + for (const entry of result) assert.notStrictEqual(entry, ''); + }); + + test('CRLF-separated entries → exactly 2 entries, none carrying a \\r', () => { + const result = resolveMutationTestFiles( + 'tests/frontmatter.unit.test.cjs\r\ntests/unusable-input.test.cjs' + ); + assert.strictEqual(result.length, 2); + for (const entry of result) assert.ok(!entry.includes('\r'), `entry retained a \\r: ${JSON.stringify(entry)}`); + }); + + test('a repeated entry is de-duplicated to exactly 1', () => { + const result = resolveMutationTestFiles( + 'tests/frontmatter.unit.test.cjs tests/frontmatter.unit.test.cjs' + ); + assert.deepStrictEqual(result, ['tests/frontmatter.unit.test.cjs']); + }); + + test('a nonexistent test file throws, naming the bad path', () => { + assert.throws( + () => resolveMutationTestFiles('tests/does-not-exist-3915.test.cjs'), + /does-not-exist-3915/ + ); + }); + + test('one bad entry among otherwise-valid entries still throws', () => { + assert.throws( + () => resolveMutationTestFiles('tests/frontmatter.unit.test.cjs tests/does-not-exist-3915.test.cjs'), + /does-not-exist-3915/ + ); + }); + + test('an entry naming an existing directory throws, indicating it is not a regular file', () => { + assert.throws( + () => resolveMutationTestFiles('tests'), + /not a regular file/ + ); + }); + + test('an entry escaping the repo root via ../../../etc/passwd throws, indicating the escape', () => { + assert.throws( + () => resolveMutationTestFiles('../../../etc/passwd'), + /escape the repo root/ + ); + }); + + test('an entry escaping via a valid-looking prefix (tests/../../outside-3915.cjs) throws for escaping the repo root', () => { + assert.throws( + () => resolveMutationTestFiles('tests/../../outside-3915.cjs'), + /escape the repo root/ + ); + }); + + test('the full derived default, space-joined, round-trips through the resolver', () => { + assert.deepStrictEqual(resolveMutationTestFiles(derived.join(' ')), derived); + }); +}); + +describe('resolveMutationTestFiles: round-trip property', () => { + const { resolveMutationTestFiles } = matrix; + + test('any non-empty subset of the derived default, space-joined, resolves back to that de-duplicated subset', () => { + fc.assert( + fc.property( + fc.subarray(derived, { minLength: 1 }), + (subset) => { + const expected = [...new Set(subset)].sort(); + const actual = resolveMutationTestFiles(subset.join(' ')); + assert.deepStrictEqual(actual, expected); + } + ), + { seed: 3915, numRuns: 100 } + ); + }); +}); + +// ── (h) mutation matrix: isolation field removed (#3915) ───────────────────── +// stryker's tap-runner has no equivalent of `node --test --test-isolation=` +// (it drives Node's own TAP test-reporter file-by-file), so the per-shard +// `isolation` field this matrix used to emit for mutation.yml's +// MUTATION_TEST_CMD env line has nothing left to wire into and must be removed +// from buildResult()'s output entirely, not merely left unused. +describe('mutation matrix: isolation field removed (#3915)', () => { + const { resolveMutationTestFiles } = matrix; + + test('every emitted matrix entry has no isolation field but keeps the rest of the contract', () => { + const covered = matrix.COVERED || {}; + const moduleNames = Object.keys(covered); + const stdinLines = moduleNames.map((name) => `src/${name}.cts`).join('\n'); + + const spawnResult = runNode( + [MATRIX_SCRIPT], + { + input: stdinLines + '\n', + cwd: REPO_ROOT, + timeoutMs: PROBE_TIMEOUT_MS, + } + ); + throwIfFailed(spawnResult, `node ${MATRIX_SCRIPT}`); + const result = JSON.parse(spawnResult.stdout); + + assert.ok(result.matrix.include.length > 0, 'matrix.include must not be empty'); + + for (const entry of result.matrix.include) { + assert.ok( + !Object.prototype.hasOwnProperty.call(entry, 'isolation'), + `matrix entry for '${entry.name}' must NOT include an isolation field (#3915: tap-runner has no test-isolation flag)` + ); + for (const key of ['name', 'mutate', 'tests', 'minScore', 'timeoutMinutes']) { + assert.ok( + Object.prototype.hasOwnProperty.call(entry, key), + `matrix entry for '${entry.name}' must still include '${key}'` + ); + } + + // Round-trip between the two halves of the wiring: the tests string this + // matrix emits must resolve back to exactly the COVERED entry's own tests. + assert.deepStrictEqual( + resolveMutationTestFiles(entry.tests), + matrix.COVERED[entry.name].tests + ); + } + }); +}); diff --git a/tests/mutation-score-ratchet.test.cjs b/tests/mutation-score-ratchet.test.cjs index 92bd72634..57015fa2e 100644 --- a/tests/mutation-score-ratchet.test.cjs +++ b/tests/mutation-score-ratchet.test.cjs @@ -81,6 +81,62 @@ describe('extractAchievedScore: Stryker json-reporter document', () => { test('throws when the report has no scoreable mutants (empty files map — a wiring bug, never a 0% score)', () => { assert.throws(() => extractAchievedScore({ schemaVersion: '1.0', thresholds: {}, files: {} }), /no scoreable mutants/); }); + + // #3915: coverageAnalysis flips 'off' -> 'perTest' for the tap-runner swap, which is + // what makes NoCoverage a status Stryker can now actually emit for these shards (the + // command runner's 'off' analysis never produced it). extractAchievedScore must already + // treat NoCoverage exactly like Survived in the denominator, or every shard's floor + // silently becomes easier to clear the moment mutants start reporting NoCoverage instead + // of Survived. + function buildCountReport({ Killed = 0, Survived = 0, NoCoverage = 0 } = {}) { + const mutants = []; + let id = 0; + const push = (status, n) => { + for (let i = 0; i < n; i += 1) { + mutants.push({ + id: String(id++), + mutatorName: 'a', + status, + location: { start: { line: 1, column: 1 }, end: { line: 1, column: 2 } }, + }); + } + }; + push('Killed', Killed); + push('Survived', Survived); + push('NoCoverage', NoCoverage); + return { + schemaVersion: '1.0', + thresholds: { high: 80, low: 60 }, + files: { 'foo.js': { language: 'javascript', source: 'x', mutants } }, + }; + } + + test('8 Killed + 2 Survived → 80', () => { + assert.strictEqual(extractAchievedScore(buildCountReport({ Killed: 8, Survived: 2 })), 80); + }); + + // NoCoverage counts in the mutation-score denominator exactly as Survived does, which is + // what makes the #3915 coverageAnalysis 'off'->'perTest' switch score-neutral; Stryker's + // own break threshold reads this same field (mutation-test-report-helper.js). + test('8 Killed + 2 NoCoverage → 80, NOT 100', () => { + assert.strictEqual(extractAchievedScore(buildCountReport({ Killed: 8, NoCoverage: 2 })), 80); + }); + + // Non-vacuity guard (Goodhart guard): proves the previous test is not a tautology by + // showing mutationScore and mutationScoreBasedOnCoveredCode genuinely diverge on the same + // report. This fails the moment someone "improves" extractAchievedScore to read the + // better-sounding covered-code field, which would make every floor trivially satisfiable. + test('non-vacuity: mutationScoreBasedOnCoveredCode (100) diverges from extractAchievedScore (80) on the same report', () => { + const { calculateMutationTestMetrics } = require('mutation-testing-metrics'); + const report = buildCountReport({ Killed: 8, NoCoverage: 2 }); + const metrics = calculateMutationTestMetrics(report); + assert.strictEqual(metrics.systemUnderTestMetrics.metrics.mutationScoreBasedOnCoveredCode, 100); + assert.strictEqual(extractAchievedScore(report), 80); + }); + + test('0 Killed + 10 NoCoverage → 0 (a wholly-uncovered shard scores zero, not 100)', () => { + assert.strictEqual(extractAchievedScore(buildCountReport({ NoCoverage: 10 })), 0); + }); }); // ── CLI end-to-end: fail on a planted over-floor score, pass when the floor is raised ─── diff --git a/tests/mutation-tap-runner-wiring.test.cjs b/tests/mutation-tap-runner-wiring.test.cjs new file mode 100644 index 000000000..e6ba2a176 --- /dev/null +++ b/tests/mutation-tap-runner-wiring.test.cjs @@ -0,0 +1,189 @@ +'use strict'; + +/** + * tests/mutation-tap-runner-wiring.test.cjs + * + * Pins the #3915 tap-runner wiring — the config contract AND the + * workflow<->config parity, because the injected env token lives on two + * surfaces (`.github/workflows/mutation.yml`'s per-shard `env:` block and + * `stryker.config.mjs`'s reader of it) and silent drift between them would + * make every shard silently fall back to running the FULL default test list + * instead of its own module's tests — slow, wrong, and (because Stryker + * would still find SOME test constraining each mutant) green. + * + * FAILING-FIRST (#3915): stryker.config.mjs still declares `testRunner: + * 'command'` with a `commandRunner.command` built from `MUTATION_TEST_CMD`, + * and mutation.yml still injects `MUTATION_TEST_CMD` / `matrix.isolation`. + * Every test below targets the tap-runner shape those files do not have yet. + */ + +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const path = require('node:path'); +const fs = require('node:fs'); +const { pathToFileURL } = require('node:url'); +const yaml = require('js-yaml'); + +const REPO_ROOT = path.resolve(__dirname, '..'); +const CONFIG_PATH = path.join(REPO_ROOT, 'stryker.config.mjs'); +const WORKFLOW_PATH = path.join(REPO_ROOT, '.github', 'workflows', 'mutation.yml'); + +// Env keys this config is known to read. Every key is saved/restored so a +// test's env override can never leak into a sibling test. +const RELEVANT_ENV_KEYS = ['MUTATION_TEST_FILES', 'MUTATION_TEST_CMD', 'MUTATION_BREAK']; + +// Cache-busting counter: importing the SAME file:// URL twice returns the +// SAME cached ES module record, so an env-dependent config test that reused +// one URL across calls would silently observe the FIRST call's env forever — +// every subsequent env-dependent assertion would pass or fail for the wrong +// reason. Incrementing this per call forces a fresh module evaluation. +let _importCounter = 0; + +/** + * Load stryker.config.mjs with `env` applied on top of process.env for the + * duration of the import, then restore process.env exactly. `env` values of + * `undefined` delete the corresponding key rather than setting it. + * + * @param {Record} env + * @returns {Promise} the config module's default export + */ +async function loadConfig(env) { + // Save/restore the UNION of RELEVANT_ENV_KEYS and Object.keys(env), not just the fixed + // list: the PARITY test below discovers its env key NAME from the parsed workflow file + // rather than hardcoding it, so a key outside RELEVANT_ENV_KEYS can reach here — saving + // only the fixed list would let that discovered key leak into sibling tests. + const keysToRestore = new Set([...RELEVANT_ENV_KEYS, ...Object.keys(env)]); + const saved = {}; + for (const key of keysToRestore) saved[key] = process.env[key]; + for (const [key, value] of Object.entries(env)) { + if (value === undefined) delete process.env[key]; + else process.env[key] = value; + } + try { + const mod = await import(`${pathToFileURL(CONFIG_PATH).href}?v=${_importCounter++}`); + return mod.default; + } finally { + for (const key of keysToRestore) { + if (saved[key] === undefined) delete process.env[key]; + else process.env[key] = saved[key]; + } + } +} + +describe('stryker.config.mjs: tap runner contract (#3915)', () => { + test("testRunner === 'tap'", async () => { + const config = await loadConfig({}); + assert.strictEqual(config.testRunner, 'tap'); + }); + + test("coverageAnalysis === 'perTest'", async () => { + const config = await loadConfig({}); + assert.strictEqual(config.coverageAnalysis, 'perTest'); + }); + + test("coverageAnalysis !== 'off' (the literal condition #3915 requires to stop being true)", async () => { + const config = await loadConfig({}); + assert.notStrictEqual(config.coverageAnalysis, 'off'); + }); + + test('no own-property commandRunner (the command runner is gone)', async () => { + const config = await loadConfig({}); + assert.ok(!Object.prototype.hasOwnProperty.call(config, 'commandRunner')); + }); + + test('tap.forceBail === false', async () => { + const config = await loadConfig({}); + assert.strictEqual(config.tap.forceBail, false); + }); + + test('no own-property buildCommand, and tap has no own-property nodeArgs (no rebuild step reintroduced, ADR-457)', async () => { + const config = await loadConfig({}); + assert.ok(!Object.prototype.hasOwnProperty.call(config, 'buildCommand')); + assert.ok(!Object.prototype.hasOwnProperty.call(config.tap, 'nodeArgs')); + }); + + test('MUTATION_TEST_FILES set → tap.testFiles deep-equals the parsed entries', async () => { + const config = await loadConfig({ + MUTATION_TEST_FILES: 'tests/frontmatter.unit.test.cjs tests/unusable-input.test.cjs', + }); + assert.deepStrictEqual(config.tap.testFiles, [ + 'tests/frontmatter.unit.test.cjs', + 'tests/unusable-input.test.cjs', + ]); + }); + + test('MUTATION_BREAK set → thresholds.break reflects it unchanged (re-proved after the runner swap)', async () => { + const config = await loadConfig({ MUTATION_BREAK: '72' }); + assert.strictEqual(config.thresholds.break, 72); + }); + + test('MUTATION_TEST_FILES set but empty → fail-closed survives all the way to config load', async () => { + await assert.rejects(() => loadConfig({ MUTATION_TEST_FILES: '' })); + }); + + test('mutate scope unchanged by the runner swap', async () => { + const config = await loadConfig({}); + assert.ok(Array.isArray(config.mutate)); + assert.ok(config.mutate.includes('gsd-core/bin/lib/**/*.cjs')); + assert.ok( + config.mutate.some((entry) => entry.startsWith('!gsd-core/bin/lib/')), + 'mutate array must still carry at least one !gsd-core/bin/lib/... exclusion' + ); + }); +}); + +describe('mutation.yml <-> stryker.config.mjs: injected env parity (#3915)', () => { + const workflowDoc = yaml.load(fs.readFileSync(WORKFLOW_PATH, 'utf8')); + const mutateJob = workflowDoc.jobs.mutate; + const runStrykerStep = mutateJob.steps.find( + (step) => typeof step.name === 'string' && step.name.startsWith('Run Stryker') + ); + + test("the 'Run Stryker' step exists", () => { + assert.ok(runStrykerStep, "no step in the mutate job's steps has a name starting with 'Run Stryker'"); + }); + + test('step.env has key MUTATION_TEST_FILES', () => { + assert.ok(Object.prototype.hasOwnProperty.call(runStrykerStep.env, 'MUTATION_TEST_FILES')); + }); + + test('step.env does NOT have key MUTATION_TEST_CMD', () => { + assert.ok(!Object.prototype.hasOwnProperty.call(runStrykerStep.env, 'MUTATION_TEST_CMD')); + }); + + test('MUTATION_TEST_FILES value is exactly the ${{ matrix.tests }} expression (derived from mutation-matrix.cjs)', () => { + assert.strictEqual(String(runStrykerStep.env.MUTATION_TEST_FILES).trim(), '${{ matrix.tests }}'); + }); + + test('MUTATION_BREAK still references matrix.minScore', () => { + assert.ok(Object.prototype.hasOwnProperty.call(runStrykerStep.env, 'MUTATION_BREAK')); + assert.ok(String(runStrykerStep.env.MUTATION_BREAK).includes('matrix.minScore')); + }); + + test('no value anywhere in the step env mentions --test-isolation', () => { + for (const value of Object.values(runStrykerStep.env)) { + assert.ok(!String(value).includes('--test-isolation'), `unexpected --test-isolation reference: ${value}`); + } + }); + + test('no value anywhere in the whole mutate job mentions matrix.isolation', () => { + const serialized = JSON.stringify(mutateJob); + assert.ok(!serialized.includes('matrix.isolation'), 'mutate job still references matrix.isolation'); + }); + + test("the step's run string does not contain '${{' (no direct interpolation inside run:, CONTRIBUTING.md)", () => { + assert.ok(!runStrykerStep.run.includes('${{'), 'run: block contains a direct ${{ }} interpolation'); + }); + + test('PARITY: the env key name discovered from the workflow drives the config test directly, so the two surfaces cannot drift', async () => { + // Find the key in step.env that starts with MUTATION_TEST_ — this is read + // from the PARSED workflow document, not hardcoded, so a rename on either + // surface (the workflow's env key, or stryker.config.mjs's reader of it) + // breaks this test instead of the two silently drifting apart. + const discoveredKey = Object.keys(runStrykerStep.env).find((k) => k.startsWith('MUTATION_TEST_')); + assert.ok(discoveredKey, 'no MUTATION_TEST_* key found in the Run Stryker step env'); + + const config = await loadConfig({ [discoveredKey]: 'tests/frontmatter.unit.test.cjs' }); + assert.deepStrictEqual(config.tap.testFiles, ['tests/frontmatter.unit.test.cjs']); + }); +});