refactor(#3180): ADR-3180 behavior contract + cross-surface drift guardrails (#3223)

* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract

The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a
lower bound for the third consecutive time, and that two derivation families
had never been named at all.

ADR-3180 gains Decision 7 — a normative behavior contract that says what the
right answer IS for each derivation, not merely who owns it. A reviewer with
no written rule can only ask "does this look like the others", which is how a
fifth copy passes review. Decision 4 gains (d) scan surface is every authored
surface and an owner FILE is never exempt, only its named functions; and (e)
a surface that cannot be consolidated today ships ratcheted, never unguarded.

Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined
copies of its own body across five modules. All six now route through it;
`clampPercentFromFraction` is added for the one caller that already held a
fraction. Every migration is behaviour-identical — clampPercent's first line IS
the `total > 0 ? … : 0` ternary each copy carried. Guarded by
lint-completion-ratio-drift.cjs, which reports zero re-derivations with no
file-level exemption.

Prompt layer: workflow markdown re-derives live-plan counting in raw shell
(#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs
scans it with a shrink-only baseline of the 7 sites that exist today — new
sites fail, and a baseline entry that stops firing fails too, so an
acknowledgment can never outlive the thing it describes.

lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only
the four named canonical functions are exempt now. The blanket exemption was
pointed at the one file most likely to grow the next copy, and it had.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage

Five findings from the two orthogonal review passes, all fixed.

Decision 4(c) breach: the completion-ratio identity test asserted at the
OWNER, which is exactly the bypass that decision exists to close — a consumer
can call clampPercent and then post-process locally, leaving both the lint and
an owner-level test green. It now drives `roadmap analyze`, `query progress`
and `stats` and asserts on their own output, over a fixture containing a
`status: superseded` plan so a consumer that re-counted raw files would report
60 where the owner reports 75.

Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the
issue that removes them. They name Phase 8 (#3218) now.

The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical
sites were one indistinguishable key and migrating either would have left the
guard green with the other alive. Entries carry an occurrence count; fewer than
acknowledged fails as a partial migration, more fails as a new copy.

Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's
test already had, and the fast-check property tests CONTRIBUTING requires for
clamp/budget-limit functions.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes)

`tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs`
under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`.
A fixed wall-clock budget around a double spawn, running inside a container
that is concurrently executing the full ~31k-test suite, fails by construction
under load.

Confirmed against three full matrix runs. Every failure was shaped
`null !== 0` — the child was KILLED, never an assertion about the thing under
test. One captured probe had already printed the correct resolution
(`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It
reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The
victim subset varies by run and by lane.

What these tests are actually about is suite-token RESOLUTION — `unit` as a
bare token in --files/--files-from. Executing the seeded trivial files is
incidental and is the entire timeout surface, so the assertions move
in-process against the same functions `main()` calls, in the same order.
`parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are
exported for that; no behavior, signature or logic changed.

No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the
harness for real and asserts exit codes end to end, on a 120s budget.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: delete the three elapsed-time assertions

CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all
three are load-sensitive: on a saturated bench each can fail while the code
under test is correct. In every case the load-bearing assertion sits on the
line above and the timing line adds no discrimination.

run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s
harness backstop?" — is already answered by the assertion above it. A backstop
kills by signal, which surfaces as status null, never 124. Observed directly
this session: three matrix runs produced exactly that null shape from killed
children.

normalize-test-command and context-predicates: both bounded a ReDoS check.
A threshold only ever separates "fast" from "slightly slow", which is bench
load, not correctness — catastrophic backtracking on 800 KB of input does not
take 251ms, it does not finish at all. A real regression therefore shows up as
the suite being killed on that test, which is louder and more reliable than a
number. The structural assertions (returned unchanged; cleanly rejected) are
what actually carry those tests, and they stay.

The sweep now reports zero elapsed-time assertions in tests/. The remaining
Date.now() uses are unique-path suffixes, barrier deadlines, fixture
timestamps and fake mtimes — none of them assertions.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3180): backfill changeset PR number (#3223)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows

The baseline keys on (file, trimmed text). `file` came from scanTree's
`path.relative()`, which uses NATIVE separators, while the committed baseline
stores POSIX. On Windows every violation was therefore unmatched — reported as
FRESH — and every baseline entry matched nothing — reported as STALE. The guard
failed 100% of the time there, on both CI shards:

  ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale
    + { file: 'gsd-core\\workflows\\execute-plan.md', ... }

The remote runner this repo gates on is Linux-only and cannot see this class at
all; the GitHub Actions Windows lane is what caught it.

Normalization is unconditional — never gated on process.platform. A
platform-conditional normalizer makes the POSIX path the special case and
leaves the Windows branch unexercised on every other OS, which is the same
blind spot in a different place. It is applied at one seam inside
findPromptDrift, which builds `file` on every returned violation, so the
baseline key, the --update writer, the stderr report and the tests all consume
one normalized value.

The regression tests drive a Windows-shaped relPath directly and run on every
OS rather than skipping off-Windows — a test that only runs on the platform
where the bug lives is why this escaped. They include a sanity check that
un-normalized input does NOT match, so the assertion cannot pass vacuously.

Audited the three sibling guards: none keys against a committed cross-platform
baseline, and their exemption keys are path.join-built, so producer and
consumer share the native convention. Left correct code alone rather than
making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing
there would break those three on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-08-08 16:05:17 -04:00
committed by GitHub
parent 636ec92107
commit b9f51836e6
22 changed files with 2023 additions and 101 deletions

View File

@@ -891,10 +891,16 @@ const path = require('path');
const { createTempDir, cleanup } = require('./helpers.cjs');
const { runNode } = require('./helpers/process-seam.cjs');
const { toLegacyResult } = require('./helpers/git-fixture.cjs');
const { PROBE_TIMEOUT_MS } = require('./helpers/timeouts.cjs');
const HARNESS = path.join(__dirname, '..', 'scripts', 'run-tests.cjs');
// Exported so the suite-token resolution contract below can be asserted
// in-process rather than through a timed subprocess spawn (see evidence
// comment above describe('bug #641 ...')).
const {
walkTestFiles,
selectExplicitFiles,
selectFiles,
parseArgs,
} = require('../scripts/run-tests.cjs');
const PASS_BODY = `'use strict';
const { test } = require('node:test');
@@ -907,17 +913,42 @@ function seed(dir, names) {
}
}
function runHarness(testDir, args = [], extraEnv = {}) {
const env = { ...process.env, GSD_TEST_DIR: testDir, ...extraEnv };
delete env.NODE_TEST_CONTEXT;
const result = runNode([HARNESS, ...args], {
cwd: path.join(__dirname, '..'),
env,
timeoutMs: PROBE_TIMEOUT_MS,
});
return toLegacyResult(result);
// Mirrors scripts/run-tests.cjs main()'s selection path (parseArgs ->
// walkTestFiles -> selectExplicitFiles/selectFiles) exactly, called
// in-process instead of through a spawned child that then spawns a nested
// `node --test`. See the evidence comment above describe('bug #641 ...').
function selectInProcess(testDir, args) {
const parsed = parseArgs(args);
if (parsed.error) return parsed;
const allFiles = walkTestFiles(testDir, '').sort();
const usingExplicitFiles = parsed.files !== null || parsed.filesFrom !== null;
if (usingExplicitFiles) {
return selectExplicitFiles(allFiles, parsed.files, parsed.filesFrom);
}
return { files: selectFiles(allFiles, parsed.suite) };
}
// Redesign evidence (2026-08-08): these probes originally spawned
// scripts/run-tests.cjs as a real child — which itself spawns a NESTED
// `node --test` — under a fixed PROBE_TIMEOUT_MS=15000 wall-clock budget,
// concurrently with the ~31k-test full suite running in the same container.
// On the remote runner this produced intermittent failures shaped
// `null !== 0` (r.status === null: the child was KILLED at the timeout),
// never a failed assertion about suite-token resolution. Reproduced on
// `next` alone (sha e705652ba): 5 failures on linux-node22, 0 failures on
// linux-node24. The victim subset varied by run and by lane, and the
// failure count went DOWN (7 -> 4 unique failures) as an unrelated diff got
// heavier — a resource collision, not a flake. The subject under test is
// suite-TOKEN RESOLUTION (parseArgs / selectExplicitFiles / selectFiles),
// not test execution; running the seeded trivial fixture files was
// incidental and was the entire timeout surface. These probes now call the
// exported selection functions in-process — removing the wall-clock budget
// around a nested spawn, not raising its number. Real end-to-end coverage of
// run-tests.cjs spawning and running to completion (exit 0 from a real
// harness run) already exists in tests/run-tests-harness.test.cjs (e.g. 'no
// flag runs ALL test files', '--suite unit excludes marked suites',
// 'non-zero from node:test propagates through harness'), so no execution
// coverage is lost by converting the "(tests run successfully)" probe below.
describe('bug #641 — --files-from with bare suite token', () => {
let tmpDir;
@@ -935,57 +966,45 @@ describe('bug #641 — --files-from with bare suite token', () => {
const listPath = path.join(tmpDir, 'ci-selected-tests.txt');
fs.writeFileSync(listPath, 'unit\n', 'utf8');
const r = runHarness(tmpDir, ['--files-from', listPath]);
const r = selectInProcess(tmpDir, ['--files-from', listPath]);
// Must NOT exit 2 with the "not found" error.
assert.notStrictEqual(
r.status,
2,
`Expected exit 0 or 1, got 2.\nstderr: ${r.stderr}\nstdout: ${r.stdout}`,
);
assert.doesNotMatch(
r.stderr,
/requested test file\(s\) not found: unit/,
`Must not emit "not found: unit".\nstderr: ${r.stderr}`,
);
// The unit suite file (a.test.cjs) must appear in the run.
// Must NOT be the "not found" error shape (the old exit-2 crash).
assert.strictEqual(r.error, undefined, `Expected a resolved file list, got error: ${r.error}`);
// The unit suite file (a.test.cjs) must appear in the resolution.
assert.ok(
r.stderr.includes('a.test.cjs'),
`Expected a.test.cjs (unit suite) to be selected.\nstderr: ${r.stderr}`,
r.files.includes('a.test.cjs'),
`Expected a.test.cjs (unit suite) to be selected. Resolved: ${r.files}`,
);
// The security suite file must NOT be included (unit token = unit only).
assert.ok(
!r.stderr.includes('b.security.test.cjs'),
`Expected b.security.test.cjs (security suite) to be excluded.\nstderr: ${r.stderr}`,
!r.files.includes('b.security.test.cjs'),
`Expected b.security.test.cjs (security suite) to be excluded. Resolved: ${r.files}`,
);
});
test('--files-from with bare "unit" token exits 0 (tests run successfully)', () => {
// "Exits 0" is determined entirely by selection succeeding with a
// non-empty list — the seeded fixture is a trivial no-op, so executing
// it contributes nothing this in-process call doesn't already prove.
// End-to-end execution coverage of run-tests.cjs lives in
// tests/run-tests-harness.test.cjs (see the evidence comment above).
seed(tmpDir, ['a.test.cjs']);
const listPath = path.join(tmpDir, 'ci-selected-tests.txt');
fs.writeFileSync(listPath, 'unit\n', 'utf8');
const r = runHarness(tmpDir, ['--files-from', listPath]);
const r = selectInProcess(tmpDir, ['--files-from', listPath]);
assert.strictEqual(
r.status,
0,
`Expected exit 0.\nstderr: ${r.stderr}\nstdout: ${r.stdout}`,
);
assert.strictEqual(r.error, undefined, `Expected a resolved file list, got error: ${r.error}`);
assert.deepStrictEqual(r.files, ['a.test.cjs']);
});
test('--files with bare "unit" token also resolves correctly', () => {
seed(tmpDir, ['a.test.cjs', 'b.security.test.cjs']);
const r = runHarness(tmpDir, ['--files', 'unit']);
const r = selectInProcess(tmpDir, ['--files', 'unit']);
assert.notStrictEqual(
r.status,
2,
`Expected exit 0, got 2.\nstderr: ${r.stderr}`,
);
assert.doesNotMatch(r.stderr, /requested test file\(s\) not found: unit/);
assert.ok(r.stderr.includes('a.test.cjs'), `a.test.cjs must be selected.\nstderr: ${r.stderr}`);
assert.ok(!r.stderr.includes('b.security.test.cjs'), `security file must not be selected.\nstderr: ${r.stderr}`);
assert.strictEqual(r.error, undefined, `Expected exit 0, got error: ${r.error}`);
assert.ok(r.files.includes('a.test.cjs'), `a.test.cjs must be selected. Resolved: ${r.files}`);
assert.ok(!r.files.includes('b.security.test.cjs'), `security file must not be selected. Resolved: ${r.files}`);
});
test('mixed: suite token "unit" alongside an explicit file resolves both', () => {
@@ -994,13 +1013,14 @@ describe('bug #641 — --files-from with bare suite token', () => {
// 'unit' expands to [a.test.cjs, b.test.cjs]; b.test.cjs is explicit too.
fs.writeFileSync(listPath, 'unit\nb.test.cjs\n', 'utf8');
const r = runHarness(tmpDir, ['--files-from', listPath]);
const r = selectInProcess(tmpDir, ['--files-from', listPath]);
assert.strictEqual(r.status, 0, `stderr: ${r.stderr}`);
// Both unit files present; security not.
assert.ok(r.stderr.includes('a.test.cjs'), `a.test.cjs must be selected.\nstderr: ${r.stderr}`);
assert.ok(r.stderr.includes('b.test.cjs'), `b.test.cjs must be selected.\nstderr: ${r.stderr}`);
assert.ok(!r.stderr.includes('c.security.test.cjs'), `c.security.test.cjs must be excluded.\nstderr: ${r.stderr}`);
assert.strictEqual(r.error, undefined, `Expected a resolved file list, got error: ${r.error}`);
// Both unit files present exactly once (the union dedupes them); security not.
assert.ok(r.files.includes('a.test.cjs'), `a.test.cjs must be selected. Resolved: ${r.files}`);
assert.ok(r.files.includes('b.test.cjs'), `b.test.cjs must be selected. Resolved: ${r.files}`);
assert.ok(!r.files.includes('c.security.test.cjs'), `c.security.test.cjs must be excluded. Resolved: ${r.files}`);
assert.strictEqual(r.files.length, 2, `expected no duplicate b.test.cjs. Resolved: ${r.files}`);
});
test('#408 fallback: ci-test-scope "unit" sentinel does not crash run-tests', () => {
@@ -1013,15 +1033,10 @@ describe('bug #641 — --files-from with bare suite token', () => {
const listPath = path.join(tmpDir, '.ci-selected-tests.txt');
fs.writeFileSync(listPath, 'unit\n', 'utf8');
const r = runHarness(tmpDir, ['--files-from', listPath]);
const r = selectInProcess(tmpDir, ['--files-from', listPath]);
assert.strictEqual(
r.status,
0,
`#408 fallback: expected exit 0 but got ${r.status}.\nstderr: ${r.stderr}`,
);
assert.doesNotMatch(r.stderr, /not found: unit/);
assert.ok(r.stderr.includes('a.test.cjs'), `unit test must run.\nstderr: ${r.stderr}`);
assert.strictEqual(r.error, undefined, `#408 fallback: expected a resolved file list, got error: ${r.error}`);
assert.ok(r.files.includes('a.test.cjs'), `unit test must resolve. Resolved: ${r.files}`);
});
});

View File

@@ -0,0 +1,576 @@
/**
* Tests for the completion-RATIO single-owner drift guard (epic #3180,
* ADR-3180) — `scripts/lint-completion-ratio-drift.cjs`.
*
* Covers:
* - `findCompletionRatioDrift` — the per-line detection shape (Math.round-
* family + `* 100` scale + an EARLIER division on the same line), and
* its documented near-miss exclusions.
* - Function-scoped owner exemption: only `clampPercent` /
* `clampPercentFromFraction` inside `src/phase-lifecycle.cts` are
* exempt — an unrelated top-level function in that SAME file is not.
* - `scanRepo` against the real repo tree: zero unsanctioned
* re-derivations (the guard's actual contract).
* - The canonical owner itself (`gsd-core/bin/lib/phase-lifecycle.cjs`'s
* `clampPercent`/`clampPercentFromFraction`) at its numeric boundaries.
*
* Uses fs.mkdtempSync directly (matching plan-count-single-owner.test.cjs /
* milestone-window-single-owner.test.cjs's own drift-guard sections, which
* build ad hoc fixture trees rather than routing through
* tests/helpers.cjs's createTempDir/cleanup for this particular shape) —
* cleaned up in `t.after()`, never a fixed path.
*/
'use strict';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const drift = require('../scripts/lint-completion-ratio-drift.cjs');
const { clampPercent, clampPercentFromFraction } = require('../gsd-core/bin/lib/phase-lifecycle.cjs');
const { scanPhasePlans } = require('../gsd-core/bin/lib/plan-scan.cjs');
const { createTempDir, cleanup, runGsdTools } = require('./helpers.cjs');
const fc = require('./helpers/fast-check-setup.cjs');
const REPO_ROOT = path.join(__dirname, '..');
const OWNER_RELPATH = path.join('src', 'phase-lifecycle.cts');
// ─── Fixture helpers (mirrors milestone-window-single-owner.test.cjs) ─────
function planningDirOf(cwd) {
return path.join(cwd, '.planning');
}
function writeRoadmap(cwd, content) {
fs.mkdirSync(planningDirOf(cwd), { recursive: true });
fs.writeFileSync(path.join(planningDirOf(cwd), 'ROADMAP.md'), content);
}
function writeState(cwd, fields) {
fs.mkdirSync(planningDirOf(cwd), { recursive: true });
const lines = ['---'];
for (const [k, v] of Object.entries(fields)) lines.push(`${k}: ${v}`);
lines.push('---', '');
fs.writeFileSync(path.join(planningDirOf(cwd), 'STATE.md'), lines.join('\n'));
}
function writeFile(cwd, relPath, content) {
const full = path.join(cwd, relPath);
fs.mkdirSync(path.dirname(full), { recursive: true });
fs.writeFileSync(full, content);
}
// ─── POSITIVE: the Math.round family, each with a genuine completed/total ─
// division whose result is scaled by 100 AFTER the divide.
describe('findCompletionRatioDrift — positive detection across the rounding family', () => {
test('Math.round with a Math.min(100, ...) ceiling is detected', () => {
const line = 'const p = total > 0 ? Math.min(100, Math.round((done / total) * 100)) : 0;';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.strictEqual(out.length, 1);
assert.strictEqual(out[0].line, 1);
});
test('Math.floor variant is detected', () => {
const line = 'const p = Math.floor((done / total) * 100);';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.strictEqual(out.length, 1);
});
test('Math.ceil variant is detected', () => {
const line = 'const p = Math.ceil((done / total) * 100);';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.strictEqual(out.length, 1);
});
test('Math.trunc variant is detected', () => {
const line = 'const p = Math.trunc((done / total) * 100);';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.strictEqual(out.length, 1);
});
test('a version with no Math.min(100, ...) ceiling is still detected', () => {
const line = 'const p = total > 0 ? Math.round((done / total) * 100) : 0;';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.strictEqual(out.length, 1);
});
});
// ─── NEGATIVE: each near-miss, with a comment saying WHY it must not fire ─
describe('findCompletionRatioDrift — negative: documented near-misses', () => {
test('Math.round(n * 100) / 100 (2-decimal rounding) is NOT detected', () => {
// Scales FIRST, divides SECOND — clause (c)'s ordering requirement
// (divIdx < scaleIdx) is the guard's whole precision, and this idiom is
// the exact shape it exists to let through: it rounds an
// already-fractional value to 2 decimal places, unrelated to a
// completed/total percentage derivation.
const line = 'const rounded = Math.round(n * 100) / 100;';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.deepStrictEqual(out, []);
});
test('Math.floor(Math.random() * 100) is NOT detected', () => {
// Carries the rounding call and the *100 scale but no division anywhere
// on the line (DIVISION_RE finds nothing, divIdx === -1) — clause (c)
// alone excludes it.
const line = 'const p = Math.floor(Math.random() * 100);';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.deepStrictEqual(out, []);
});
test('Math.min(Math.round(ratio * 100), 100) is NOT detected', () => {
// Scales an ALREADY-COMPUTED fraction (`ratio`) — there is no division
// anywhere on this line either, so it is out of scope by domain (no
// completed/total pair is being re-derived here), the same reason the
// Math.random() case above is excluded.
const line = 'const p = Math.min(Math.round(ratio * 100), 100);';
const out = drift.findCompletionRatioDrift(line, 'src/somewhere.cts');
assert.deepStrictEqual(out, []);
});
test('a division and a * 100 scale on two DIFFERENT lines is NOT detected', () => {
// Documented per-line limit: this guard's detection window is ONE
// source line. The division happens on line 1; line 2 carries the
// Math.round-family call and the *100 scale but no division of its
// own, so DIVISION_RE finds nothing on line 2 and clause (c) excludes
// it, even though the two lines together form the exact re-derivation
// shape the guard exists to catch.
const text = [
'const frac = done / total;',
'const p = Math.round(frac * 100);',
].join('\n');
const out = drift.findCompletionRatioDrift(text, 'src/somewhere.cts');
assert.deepStrictEqual(out, []);
});
});
// ─── OWNER SCOPING: function-scoped, not file-scoped ──────────────────────
describe('findCompletionRatioDrift — owner exemption is function-scoped, not file-scoped', () => {
test('the re-derivation line inside clampPercent in src/phase-lifecycle.cts is exempt', () => {
const text = [
'function clampPercent(completed, total) {',
' return total > 0 ? Math.min(100, Math.round((completed / total) * 100)) : 0;',
'}',
].join('\n');
const out = drift.findCompletionRatioDrift(text, OWNER_RELPATH);
assert.deepStrictEqual(out, []);
});
test('the re-derivation line inside clampPercentFromFraction in src/phase-lifecycle.cts is exempt', () => {
const text = [
'function clampPercentFromFraction(fraction) {',
' return Math.min(100, Math.round((fraction * total) / 100 * 100));',
'}',
].join('\n');
// Note: the real clampPercentFromFraction body carries no division at
// all (Math.min(100, Math.round(fraction * 100))) and so never matches
// regardless of exemption — this fixture synthesizes a line that WOULD
// match the detection shape, specifically to prove the exemption itself
// (not merely the absence of a division) is what suppresses it.
const out = drift.findCompletionRatioDrift(text, OWNER_RELPATH);
assert.deepStrictEqual(out, []);
});
test('the SAME line inside a differently-named top-level function in the SAME file IS reported', () => {
// This is the point of function-scoped rather than file-scoped
// exemption: src/phase-lifecycle.cts is scanned like every other file,
// and only the two named canonical functions are exempt.
const text = [
'function someOtherFunction(completed, total) {',
' return total > 0 ? Math.min(100, Math.round((completed / total) * 100)) : 0;',
'}',
].join('\n');
const out = drift.findCompletionRatioDrift(text, OWNER_RELPATH);
assert.strictEqual(out.length, 1);
assert.strictEqual(out[0].line, 2);
});
test('exempt and non-exempt functions in ONE file: only the non-exempt line is reported', () => {
const text = [
'function clampPercent(completed, total) {',
' return total > 0 ? Math.min(100, Math.round((completed / total) * 100)) : 0;',
'}',
'',
'function someOtherFunction(completed, total) {',
' return total > 0 ? Math.min(100, Math.round((completed / total) * 100)) : 0;',
'}',
].join('\n');
const out = drift.findCompletionRatioDrift(text, OWNER_RELPATH);
assert.strictEqual(out.length, 1);
assert.strictEqual(out[0].line, 6);
});
test('the same line in a DIFFERENT, non-owner file is reported (no exemption applies)', () => {
const line = 'const p = total > 0 ? Math.min(100, Math.round((completed / total) * 100)) : 0;';
const out = drift.findCompletionRatioDrift(line, path.join('src', 'unrelated.cts'));
assert.strictEqual(out.length, 1);
});
});
// ─── MAX_REGEX_LITERAL_LEN boundary on the reported fragment ──────────────
// Sibling precedent: tests/plan-count-single-owner.test.cjs's
// `readRegexLiteralAt` MAX_REGEX_LITERAL_LEN boundary test (limit-1 reads,
// limit reads, limit+1 returns/truncates). This guard has no regex-literal
// tokenizer of its own (see module header) — it truncates its REPORTED
// fragment with `.trim().slice(0, MAX_REGEX_LITERAL_LEN)`, so the boundary
// here is on `found.length`, not on a returned `null`.
describe('findCompletionRatioDrift — MAX_REGEX_LITERAL_LEN boundary on the reported fragment', () => {
test('limit-1 reports the fragment in full, limit reports it in full, limit+1 truncates', () => {
const MAX = drift.MAX_REGEX_LITERAL_LEN;
assert.ok(Number.isInteger(MAX) && MAX > 10, `MAX_REGEX_LITERAL_LEN must be exported as an integer > 10, got ${MAX}`);
// A genuine re-derivation shape (Math.round + earlier division + *100),
// padded with trailing filler characters (never `/`, `*`, digits, or
// whitespace-adjacent tokens that could themselves alter detection) so
// the TOTAL trimmed-line length lands exactly on each boundary.
const base = 'const p = Math.round((done / total) * 100);';
const build = (totalLen) => {
if (totalLen < base.length) throw new Error(`totalLen ${totalLen} shorter than base fixture ${base.length}`);
return base + 'z'.repeat(totalLen - base.length);
};
const underBound = build(MAX - 1); // limit-1
const atBound = build(MAX); // limit
const overBound = build(MAX + 1); // limit+1
const underOut = drift.findCompletionRatioDrift(underBound, 'src/somewhere.cts');
assert.strictEqual(underOut.length, 1);
assert.strictEqual(underOut[0].found.length, MAX - 1);
assert.strictEqual(underOut[0].found, underBound);
const atOut = drift.findCompletionRatioDrift(atBound, 'src/somewhere.cts');
assert.strictEqual(atOut.length, 1);
assert.strictEqual(atOut[0].found.length, MAX);
assert.strictEqual(atOut[0].found, atBound);
const overOut = drift.findCompletionRatioDrift(overBound, 'src/somewhere.cts');
assert.strictEqual(overOut.length, 1);
assert.strictEqual(overOut[0].found.length, MAX);
// Truncated to exactly the limit-length fragment — the first MAX
// characters of overBound are byte-identical to atBound (same base,
// same padding character).
assert.strictEqual(overOut[0].found, atBound);
});
});
// ─── scanRepo — tree-walk mechanics on a synthetic tree ───────────────────
describe('scanRepo — synthetic tree', () => {
test('a violation in a fresh temp tree is reported with its file and line', (t) => {
const root = createTempDir('gsd-completion-ratio-drift-');
t.after(() => cleanup(root));
fs.mkdirSync(path.join(root, 'src'), { recursive: true });
fs.writeFileSync(
path.join(root, 'src', 'fake.cts'),
'const p = Math.round((done / total) * 100);\n',
);
const violations = drift.scanRepo(root);
assert.strictEqual(violations.length, 1);
assert.strictEqual(violations[0].file, path.join('src', 'fake.cts'));
assert.strictEqual(violations[0].line, 1);
});
test('a clean temp tree with no re-derivations reports zero violations', (t) => {
const root = createTempDir('gsd-completion-ratio-drift-');
t.after(() => cleanup(root));
fs.mkdirSync(path.join(root, 'src'), { recursive: true });
fs.writeFileSync(path.join(root, 'src', 'clean.cts'), 'const x = 1;\n');
const violations = drift.scanRepo(root);
assert.deepStrictEqual(violations, []);
});
});
// ─── WHOLE-REPO contract: the guard's actual promise ──────────────────────
test('scanRepo(repoRoot) against the real repo returns EMPTY — zero independent re-derivations', () => {
// This is the guard's actual contract ("0 independent re-derivations") —
// the test that fails the day copy number seven lands.
const violations = drift.scanRepo(REPO_ROOT);
assert.deepStrictEqual(violations, []);
});
// ─── Canonical owner: clampPercent / clampPercentFromFraction boundaries ──
describe('clampPercent — canonical owner boundaries', () => {
test('total 0 yields 0 (nothing to complete is 0%, never 100%)', () => {
assert.strictEqual(clampPercent(0, 0), 0);
});
test('total 1, completed 0 yields 0', () => {
assert.strictEqual(clampPercent(0, 1), 0);
});
test('completed === total yields 100', () => {
assert.strictEqual(clampPercent(5, 5), 100);
});
test('completed > total: ceiling holds at 100', () => {
assert.strictEqual(clampPercent(7, 5), 100);
});
test('a negative total yields 0', () => {
assert.strictEqual(clampPercent(3, -5), 0);
});
});
describe('clampPercentFromFraction — canonical owner boundaries', () => {
test('fraction 0 yields 0', () => {
assert.strictEqual(clampPercentFromFraction(0), 0);
});
test('fraction 1 (completed === total, expressed as a fraction) yields 100', () => {
assert.strictEqual(clampPercentFromFraction(1), 100);
});
test('a fraction greater than 1 (completed > total): ceiling holds at 100', () => {
assert.strictEqual(clampPercentFromFraction(1.4), 100);
});
});
// ═════════════════════════════════════════════════════════════════════════
// CONSUMER identity (ADR-3180 Decision 4c): the completion-ratio derivation's
// canonical owner is `clampPercent`/`clampPercentFromFraction`
// (src/phase-lifecycle.cts). Decision 4(c) requires the identity test to
// compare each CONSUMER's own observable output against the canonical
// owner's result for the SAME input — a consumer can call the owner and
// then post-process the counts it feeds it locally, which satisfies both
// the structural lint (no re-implementation of the rounding arithmetic) AND
// an owner-level identity test (the owner itself is untouched) while
// silently restoring divergence. tests/milestone-window-single-owner.test.cjs
// (~lines 819-857) is the correct precedent in this repo: it drives real CLI
// verbs and asserts on THEIR output, not the owner's return value alone.
//
// The fixture below deliberately includes a `status: superseded` plan
// (03-baz/03-02-PLAN.md) — the canonical live-plan-counting owner
// (`scanPhasePlans`, #3183) excludes it, but a consumer that locally
// re-filtered or hand-counted files instead of calling the owner would
// include it, producing a DIFFERENT numerator/denominator and therefore a
// DIFFERENT percentage. That divergence is what gives this identity test
// power per Decision 4(c) — the "power proof" test below shows the two
// counting strategies land on different percentages (75% vs 60%) over the
// SAME fixture.
// ═════════════════════════════════════════════════════════════════════════
describe('CONSUMER identity (ADR-3180 Decision 4c): percent fields match the canonical owner', () => {
function buildFixture(cwd) {
writeState(cwd, { milestone: 'v1.0' });
writeRoadmap(cwd, [
'## v1.0 Current 🚧',
'',
'### Phase 1: Foo',
'',
'### Phase 2: Bar',
'',
'### Phase 3: Baz',
].join('\n'));
// Phase 1 (01-foo): 2 live plans, 2 summaries, verified -> Complete.
writeFile(cwd, '.planning/phases/01-foo/01-01-PLAN.md', '# Plan\n');
writeFile(cwd, '.planning/phases/01-foo/01-02-PLAN.md', '# Plan\n');
writeFile(cwd, '.planning/phases/01-foo/01-01-SUMMARY.md', '# Summary\n');
writeFile(cwd, '.planning/phases/01-foo/01-02-SUMMARY.md', '# Summary\n');
writeFile(cwd, '.planning/phases/01-foo/VERIFICATION.md', '---\nstatus: passed\n---\n# Verification\n');
// Phase 2 (02-bar): 1 live plan, 0 summaries -> Planned, not complete.
writeFile(cwd, '.planning/phases/02-bar/02-01-PLAN.md', '# Plan\n');
// Phase 3 (03-baz): 1 live plan + 1 SUPERSEDED plan (excluded by the
// canonical owner scanPhasePlans, #2349), 1 summary, verified -> Complete.
writeFile(cwd, '.planning/phases/03-baz/03-01-PLAN.md', '# Plan\n');
writeFile(cwd, '.planning/phases/03-baz/03-01-SUMMARY.md', '# Summary\n');
writeFile(cwd, '.planning/phases/03-baz/03-02-PLAN.md', '---\nstatus: superseded\n---\n# Plan\n');
writeFile(cwd, '.planning/phases/03-baz/VERIFICATION.md', '---\nstatus: passed\n---\n# Verification\n');
return ['01-foo', '02-bar', '03-baz'];
}
// Independently compute the canonical totals by calling the OWNER
// (scanPhasePlans) directly per phase directory — never re-deriving the
// superseded-exclusion filter locally in this test.
function ownerTotals(cwd, dirs) {
let totalPlans = 0;
let totalSummaries = 0;
for (const dir of dirs) {
const scan = scanPhasePlans(path.join(cwd, '.planning', 'phases', dir));
totalPlans += scan.planCount;
totalSummaries += scan.summaryCount;
}
return { totalPlans, totalSummaries };
}
test('power proof: a naive raw-file count over this fixture computes a DIFFERENT percentage than the canonical owner', () => {
// Not itself a guard assertion — documents why the fixture below has the
// power Decision 4(c) requires. If this ever fails, the fixture has
// stopped exercising the superseded-exclusion divergence and must be
// redesigned.
const canonical = clampPercent(3, 4); // scanPhasePlans-derived: 3 summaries / 4 live plans
const naive = clampPercent(3, 5); // raw file count: 03-02-PLAN.md counted despite being superseded
assert.strictEqual(canonical, 75);
assert.strictEqual(naive, 60);
assert.notStrictEqual(canonical, naive);
});
test('roadmap analyze: progress_percent matches clampPercent(owner totals)', (t) => {
const cwd = createTempDir('gsd-completion-ratio-consumer-');
t.after(() => cleanup(cwd));
const dirs = buildFixture(cwd);
const { totalPlans, totalSummaries } = ownerTotals(cwd, dirs);
const expected = clampPercent(totalSummaries, totalPlans);
assert.strictEqual(expected, 75, 'fixture sanity: must match the power-proof test above');
const result = runGsdTools(['roadmap', 'analyze', '--cwd', cwd, '--raw'], cwd);
assert.strictEqual(result.success, true, result.error);
const analyzed = JSON.parse(result.output);
assert.strictEqual(analyzed.progress_percent, expected);
});
test('query progress: the percent field matches clampPercent(owner totals)', (t) => {
const cwd = createTempDir('gsd-completion-ratio-consumer-');
t.after(() => cleanup(cwd));
const dirs = buildFixture(cwd);
const { totalPlans, totalSummaries } = ownerTotals(cwd, dirs);
const expected = clampPercent(totalSummaries, totalPlans);
assert.strictEqual(expected, 75, 'fixture sanity: must match the power-proof test above');
const result = runGsdTools(['query', 'progress', '--cwd', cwd, '--raw'], cwd);
assert.strictEqual(result.success, true, result.error);
const rendered = JSON.parse(result.output);
assert.strictEqual(rendered.percent, expected);
});
test('stats: plan_percent matches clampPercent(owner totals); percent (phase-level) matches clampPercent of its own reported completed/total', (t) => {
const cwd = createTempDir('gsd-completion-ratio-consumer-');
t.after(() => cleanup(cwd));
const dirs = buildFixture(cwd);
const { totalPlans, totalSummaries } = ownerTotals(cwd, dirs);
const expectedPlanPercent = clampPercent(totalSummaries, totalPlans);
assert.strictEqual(expectedPlanPercent, 75, 'fixture sanity: must match the power-proof test above');
const result = runGsdTools(['stats', '--cwd', cwd, '--raw'], cwd);
assert.strictEqual(result.success, true, result.error);
const stats = JSON.parse(result.output);
// plan_percent: the completion-ratio derivation this ADR section (§7.6)
// owns — checked against the INDEPENDENTLY owner-derived totals, so a
// consumer post-filtering scanPhasePlans's output locally would fail
// this exact assertion (see the power-proof test above).
assert.strictEqual(stats.plan_percent, expectedPlanPercent);
// percent (phase-level completion): the completed/total PAIR itself is
// phase enumeration + phase completion's derivation (§7.3/§7.4 — a
// SEPARATE ADR-3180 phase, not owned by this guard; §7.6's own status
// note says numerator/denominator SAME-SCOPE-SET enforcement is Phase 7,
// not yet shipped). Re-deriving `determinePhaseStatus` locally in this
// test to independently compute completedPhases/phasesTotal would
// duplicate PRODUCTION status logic rather than exercise this
// derivation's owner. What THIS guard (§7.6, completion-ratio
// arithmetic) owns is that whatever completed/total pair `stats`
// reports is turned into a percent through the canonical clampPercent
// arithmetic, never a locally re-derived ternary/rounding — asserted
// directly against the command's own reported pair.
assert.strictEqual(stats.percent, clampPercent(stats.phases_completed, stats.phases_total));
// Fixture sanity: phases_completed/phases_total must be the values this
// fixture was designed to produce (2 of 3 phases Complete), so the
// assertion above is not vacuously true against a degenerate 0/0 pair.
assert.strictEqual(stats.phases_completed, 2);
assert.strictEqual(stats.phases_total, 3);
assert.strictEqual(stats.percent, 67);
});
});
// ═════════════════════════════════════════════════════════════════════════
// PROPERTY tests (CONTRIBUTING.md: parsers, budget limits, and bijective
// contracts require at least one fast-check property test). clampPercent /
// clampPercentFromFraction are clamp/budget-limit functions.
// ═════════════════════════════════════════════════════════════════════════
describe('clampPercent / clampPercentFromFraction — property tests (fast-check)', () => {
// Bounded integer domains (matching this repo's fast-check convention —
// see tests/derive-progress.property.test.cjs, tests/eval.property.test.cjs
// — rather than unconstrained fc.float()/fc.double(), which would exercise
// float-precision edge cases orthogonal to the property under test).
test('property: non-negative completed + any finite total -> integer result in [0, 100]', () => {
// Restricted to completed >= 0 -- every real caller in this codebase
// passes a COUNT (never a negative numerator), and clampPercent does not
// claim to clamp a negative numerator's result into [0, 100] (verified
// against the implementation: clampPercent(-5, 10) returns -50, since
// Math.min(100, x) has no LOWER bound). The property as stated is true
// only within the domain the function is actually used in.
fc.assert(
fc.property(
fc.integer({ min: 0, max: 1_000_000 }),
fc.integer({ min: -1_000_000, max: 1_000_000 }),
(completed, total) => {
const result = clampPercent(completed, total);
assert.ok(Number.isInteger(result), `result must be an integer, got ${result}`);
assert.ok(result >= 0 && result <= 100, `result must be in [0, 100], got ${result} for clampPercent(${completed}, ${total})`);
},
),
);
});
test('property: a non-positive or non-finite total always yields exactly 0 (for a non-negative completed)', () => {
// completed restricted to >= 0 for the same reason as the property
// above (a real numerator is always a count) — AND to sidestep a signed
// -0 vs +0 distinction: with total === Infinity the division short-
// circuit is skipped (Infinity > 0), so completed/Infinity is computed
// directly, and a NEGATIVE completed produces -0 (verified against the
// implementation: clampPercent(-5, Infinity) returns -0, which
// Object.is-based assert.strictEqual treats as distinct from +0 even
// though both represent 0%).
fc.assert(
fc.property(
fc.integer({ min: 0, max: 1_000_000 }),
fc.oneof(
fc.integer({ min: -1_000_000, max: 0 }),
fc.constant(NaN),
fc.constant(Infinity),
fc.constant(-Infinity),
),
(completed, total) => {
assert.strictEqual(clampPercent(completed, total), 0);
},
),
);
});
test('property: clampPercent(c, t) === clampPercentFromFraction(c / t) for every t > 0', () => {
fc.assert(
fc.property(
fc.integer({ min: -1_000_000, max: 1_000_000 }),
fc.integer({ min: 1, max: 1_000_000 }),
(completed, total) => {
assert.strictEqual(clampPercent(completed, total), clampPercentFromFraction(completed / total));
},
),
);
});
test('property: clampPercent is monotonic non-decreasing in completed for a fixed total > 0', () => {
fc.assert(
fc.property(
fc.integer({ min: 1, max: 1_000_000 }),
fc.integer({ min: -1_000_000, max: 1_000_000 }),
fc.integer({ min: -1_000_000, max: 1_000_000 }),
(total, a, b) => {
const [lo, hi] = a <= b ? [a, b] : [b, a];
assert.ok(
clampPercent(lo, total) <= clampPercent(hi, total),
`clampPercent(${lo}, ${total})=${clampPercent(lo, total)} must be <= clampPercent(${hi}, ${total})=${clampPercent(hi, total)}`,
);
},
),
);
});
});

View File

@@ -399,21 +399,20 @@ describe('parsePredicates: ID/value grammar boundaries (C)', () => {
assert.equal(r.predicates.length, 0, 'a doubled dot (empty segment) must be rejected, not silently accepted');
});
test('idGrammarValidationIsLinearTimeAgainstManyConsecutiveDots', () => {
test('idGrammarValidationRejectsManyConsecutiveDots', () => {
// DEFECT.CONTEXT-PREDICATES-ID-REDOS (MAJOR review finding): the old
// `ID_RE`'s `(?:\.[A-Za-z0-9_.-]+)*` group was exponential in the number
// of consecutive dots (measured: ~565ms for 40 dots). The structural
// per-segment validator is linear. A generous wall-clock bound is used
// only as a smoke check; the load-bearing assertion is that the result
// is a clean rejection (an id-shaped line with 60 consecutive dots has
// an empty segment at every step and must not parse).
// per-segment validator is linear. No elapsed-time bound is asserted:
// the load-bearing assertion is that the result is a clean rejection
// (an id-shaped line with 60 consecutive dots has an empty segment at
// every step and must not parse). An exponential regression would not
// finish at all, not merely exceed a threshold — so a clean rejection
// is itself the discriminator.
const dots = '.'.repeat(60);
const md = `\`A${dots}x=value\``;
const start = Date.now();
const r = parsePredicates(md);
const elapsedMs = Date.now() - start;
assert.equal(r.predicates.length, 0, 'an id with 60 consecutive dots has empty segments and must cleanly reject');
assert.ok(elapsedMs < 1000, `expected well under 1s (linear time), got ${elapsedMs}ms — possible ReDoS regression`);
});
});

View File

@@ -1119,6 +1119,42 @@ test('root confinement holds', { skip: process.platform === 'win32' ? 'symlink c
assert.strictEqual(violations.length, 0);
});
// The guard's owner file (src/roadmap-parser.cts) is no longer exempt as a
// whole file — only its named canonical functions are (FUNCTION_SCOPED_EXEMPTIONS).
// These synthesize a relPath of 'src/roadmap-parser.cts' WITHOUT touching the
// real source file, exercising findMilestoneWindowDrift directly.
test('a non-exempt top-level function in src/roadmap-parser.cts IS reported', () => {
const relPath = path.join('src', 'roadmap-parser.cts');
const text = [
'function someOtherFunction() {',
` ${violatingLine().trim()}`,
'}',
].join('\n');
const violations = driftGuard.findMilestoneWindowDrift(text, relPath);
assert.strictEqual(violations.length, 1);
assert.strictEqual(violations[0].line, 2);
});
const OWNER_FILE_EXEMPT_FUNCTIONS = [
'isMilestoneShippedInRoadmap',
'locateMilestoneHeadings',
'hasMilestoneSectioning',
'extractCurrentMilestoneScoped',
];
for (const fnName of OWNER_FILE_EXEMPT_FUNCTIONS) {
test(`owner-file exempt function ${fnName} in src/roadmap-parser.cts is NOT reported`, () => {
const relPath = path.join('src', 'roadmap-parser.cts');
const text = [
`function ${fnName}() {`,
` ${violatingLine().trim()}`,
'}',
].join('\n');
const violations = driftGuard.findMilestoneWindowDrift(text, relPath);
assert.deepStrictEqual(violations, []);
});
}
test('report output is sanitized', () => {
// findMilestoneWindowDrift returns the RAW fragment; main() sanitizes both
// `file` and `found` at the reporting boundary via the shared

View File

@@ -134,14 +134,17 @@ describe('normalizeTestCommand: security hardening (#1857 review)', () => {
});
}
test('an oversized command is returned unchanged and in linear time (no ReDoS)', () => {
test('an oversized command is returned unchanged (no ReDoS)', () => {
// The blow-up input from the review: a long "npm " run with no `test` token.
const huge = 'npm '.repeat(200000); // ~800 KB
const start = Date.now();
const out = normalizeTestCommand(huge, '/tmp');
const elapsedMs = Date.now() - start;
// No elapsed-time bound: catastrophic backtracking on an 800 KB input
// does not take 251ms, it does not finish at all. A real ReDoS
// regression manifests as the suite being killed on this test, which is
// a louder and more reliable signal than a threshold — the threshold
// only ever distinguished "fast" from "slightly slow" (bench load), not
// correctness.
assert.strictEqual(out, huge, 'oversized input must be returned unchanged');
assert.ok(elapsedMs < 250, `normalization must be fast even on adversarial input (took ${elapsedMs}ms)`);
});
test('a package.json that is not a regular file is ignored (no FIFO hang)', () => {

View File

@@ -0,0 +1,326 @@
/**
* Tests for the prompt-layer plan/summary-COUNTING drift guard (epic #3180,
* ADR-3180 Decision 4(e)) — `scripts/lint-planning-prompt-drift.cjs`.
*
* Covers:
* - `findPromptDrift` — the per-line detection shape (a `*...PLAN.md` /
* `*...SUMMARY.md` set glob AND a counting operator on the same line),
* and its documented near-miss exclusions.
* - `diffAgainstBaseline` — the three ratchet invariants (known / fresh /
* stale), keyed on TEXT not line number, exercised on synthetic input.
* - `loadBaseline` / `scanRepo` against the real, committed repo state —
* the guard's actual contract.
*
* Uses fs.mkdtempSync directly for the one synthetic-tree fixture, matching
* the sibling drift-guard test suites' own drift-guard sections — cleaned
* up in `t.after()`, never a fixed path.
*/
'use strict';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const drift = require('../scripts/lint-planning-prompt-drift.cjs');
const { findPromptDrift, scanRepo, loadBaseline, diffAgainstBaseline, toPosixRel, writeBaseline } = drift;
const { createTempDir, cleanup } = require('./helpers.cjs');
const REPO_ROOT = path.join(__dirname, '..');
// ─── POSITIVE ───────────────────────────────────────────────────────────
describe('findPromptDrift — positive detection', () => {
test('X=$(ls dir/*-PLAN.md 2>/dev/null | wc -l) is detected', () => {
const line = 'X=$(ls dir/*-PLAN.md 2>/dev/null | wc -l)';
const out = findPromptDrift(line, 'gsd-core/workflows/fake.md');
assert.strictEqual(out.length, 1);
assert.strictEqual(out[0].found, '*-PLAN.md');
assert.strictEqual(out[0].text, line);
});
test('the *-SUMMARY.md variant is detected', () => {
const line = 'X=$(ls dir/*-SUMMARY.md 2>/dev/null | wc -l)';
const out = findPromptDrift(line, 'gsd-core/workflows/fake.md');
assert.strictEqual(out.length, 1);
assert.strictEqual(out[0].found, '*-SUMMARY.md');
});
test('a grep -c variant is detected', () => {
const line = "Y=$(grep -cE '^' dir/*-PLAN.md)";
const out = findPromptDrift(line, 'gsd-core/workflows/fake.md');
assert.strictEqual(out.length, 1);
assert.strictEqual(out[0].found, '*-PLAN.md');
});
});
// ─── NEGATIVE — each with a comment saying WHY it must not fire ──────────
describe('findPromptDrift — negative: documented near-misses', () => {
test('grep -cE task-heading count inside ONE NAMED plan (no glob) is NOT detected', () => {
// A real line in gsd-core/workflows/execute-plan.md: it counts <task>
// elements INSIDE one already-named plan file — no `*` glob token
// anywhere near PLAN.md — so it is not a plan-COUNT re-derivation. A
// false positive here would redden lint:ci on an untouched file.
const line = "grep -cE '^\\s*<task[[:space:]>]' .planning/phases/[current-phase-dir]/{phase}-{plan}-PLAN.md";
const out = findPromptDrift(line, 'gsd-core/workflows/execute-plan.md');
assert.deepStrictEqual(out, []);
});
test('a *-UAT.md count is NOT detected', () => {
// UAT artifacts are a different derivation this guard does not own —
// PLAN_SUMMARY_GLOB_RE requires the literal PLAN.md or SUMMARY.md
// suffix, which "UAT.md" never satisfies.
const line = 'X=$(ls dir/*-UAT.md 2>/dev/null | wc -l)';
const out = findPromptDrift(line, 'gsd-core/workflows/fake.md');
assert.deepStrictEqual(out, []);
});
test('a line that globs plan files but does not count them is NOT detected', () => {
// Reading/iterating (cat, backup, cross-reference) over a *-PLAN.md
// glob without a counting operator is not this derivation — every
// non-counting *-PLAN.md/*-SUMMARY.md glob in plan-phase.md is exactly
// this shape and is deliberately left alone.
const line = 'cat dir/*-PLAN.md';
const out = findPromptDrift(line, 'gsd-core/workflows/fake.md');
assert.deepStrictEqual(out, []);
});
});
// ─── RATCHET MECHANICS — diffAgainstBaseline on synthetic inputs ─────────
describe('diffAgainstBaseline — ratchet invariants (synthetic)', () => {
test('a violation whose (file, text) pair is in the baseline is KNOWN: neither fresh nor stale', () => {
const baseline = [{ file: 'a.md', text: 'X=$(ls *-PLAN.md 2>/dev/null | wc -l)' }];
const violations = [
{ file: 'a.md', line: 10, found: '*-PLAN.md', text: 'X=$(ls *-PLAN.md 2>/dev/null | wc -l)' },
];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(fresh, []);
assert.deepStrictEqual(stale, []);
});
test('a violation absent from the baseline is FRESH: fails', () => {
const baseline = [];
const violations = [
{ file: 'a.md', line: 1, found: '*-PLAN.md', text: 'Y=$(grep -c dir/*-PLAN.md)' },
];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.strictEqual(fresh.length, 1);
assert.strictEqual(fresh[0].text, 'Y=$(grep -c dir/*-PLAN.md)');
assert.deepStrictEqual(stale, []);
});
test('a baseline entry matching nothing this run is STALE: fails', () => {
const baseline = [{ file: 'a.md', text: 'X=$(ls *-PLAN.md 2>/dev/null | wc -l)' }];
const violations = [];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(fresh, []);
assert.strictEqual(stale.length, 1);
assert.strictEqual(stale[0].text, 'X=$(ls *-PLAN.md 2>/dev/null | wc -l)');
});
test('keying is on TEXT not line number: the same trimmed text at a different line is still KNOWN', () => {
// This is what stops the baseline rotting on an unrelated edit that
// merely shifts line numbers (a new paragraph, a reworded step).
const baseline = [{ file: 'a.md', text: 'X=$(ls *-PLAN.md 2>/dev/null | wc -l)' }];
const violations = [
{ file: 'a.md', line: 999, found: '*-PLAN.md', text: 'X=$(ls *-PLAN.md 2>/dev/null | wc -l)' },
];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(fresh, []);
assert.deepStrictEqual(stale, []);
});
// ─── count-aware ratchet (Finding-3 fix): duplicate (file, text) pairs no
// longer make a partial migration invisible ───────────────────────────
test('a pair with count:2 fully matched by TWO occurrences is KNOWN: neither fresh nor stale', () => {
const baseline = [{ file: 'a.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)', count: 2 }];
const violations = [
{ file: 'a.md', line: 10, found: '*-PLAN.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' },
{ file: 'a.md', line: 40, found: '*-PLAN.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' },
];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(fresh, []);
assert.deepStrictEqual(stale, []);
});
test('a pair with count:2 but only ONE occurrence this run is a PARTIAL-migration STALE, naming both numbers', () => {
// This is the exact defect Finding 3 closes: migrating only ONE of two
// byte-identical sites must not be invisible to the ratchet just because
// the OTHER site still matches the (file, text) pair.
const baseline = [{ file: 'a.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)', count: 2 }];
const violations = [
{ file: 'a.md', line: 10, found: '*-PLAN.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' },
];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(fresh, []);
assert.strictEqual(stale.length, 1);
assert.strictEqual(stale[0].count, 2);
assert.strictEqual(stale[0].actualCount, 1);
});
test('a pair with count:2 and ZERO occurrences this run is fully STALE (both sites migrated)', () => {
const baseline = [{ file: 'a.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)', count: 2 }];
const { fresh, stale } = diffAgainstBaseline([], baseline);
assert.deepStrictEqual(fresh, []);
assert.strictEqual(stale.length, 1);
assert.strictEqual(stale[0].actualCount, 0);
assert.strictEqual(stale[0].count, 2);
});
test('a pair with count:1 but a THIRD occurrence appears this run: the excess occurrence is FRESH (new copy)', () => {
const baseline = [{ file: 'a.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)', count: 1 }];
const violations = [
{ file: 'a.md', line: 10, found: '*-PLAN.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' },
{ file: 'a.md', line: 55, found: '*-PLAN.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' },
];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(stale, []);
assert.strictEqual(fresh.length, 1);
assert.strictEqual(fresh[0].line, 55);
});
test('an entry with no `count` field defaults to acknowledging exactly ONE occurrence', () => {
const baseline = [{ file: 'a.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' }];
const violations = [
{ file: 'a.md', line: 10, found: '*-PLAN.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' },
{ file: 'a.md', line: 55, found: '*-PLAN.md', text: 'DISK_PLANS=$(ls *-PLAN.md | wc -l)' },
];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(stale, []);
assert.strictEqual(fresh.length, 1);
assert.strictEqual(fresh[0].line, 55);
});
});
// ─── scanRepo — tree-walk mechanics on a synthetic tree ───────────────────
describe('scanRepo — synthetic tree', () => {
test('a violation in a fresh temp tree is reported with its file, line, and text', (t) => {
const root = createTempDir('gsd-planning-prompt-drift-');
t.after(() => cleanup(root));
fs.mkdirSync(path.join(root, 'gsd-core', 'workflows'), { recursive: true });
fs.writeFileSync(
path.join(root, 'gsd-core', 'workflows', 'fake.md'),
'X=$(ls dir/*-PLAN.md 2>/dev/null | wc -l)\n',
);
const violations = scanRepo(root);
assert.strictEqual(violations.length, 1);
// Always POSIX-separated regardless of the host OS's native separator
// (`path.join` would build native separators here, which is exactly the
// Windows-vs-POSIX mismatch this guard's baseline keying must not have —
// see the Windows-shaped-path coverage below).
assert.strictEqual(violations[0].file, 'gsd-core/workflows/fake.md');
assert.strictEqual(violations[0].line, 1);
assert.strictEqual(violations[0].found, '*-PLAN.md');
});
test('a clean temp tree with no re-derivations reports zero violations', (t) => {
const root = createTempDir('gsd-planning-prompt-drift-');
t.after(() => cleanup(root));
fs.mkdirSync(path.join(root, 'gsd-core', 'workflows'), { recursive: true });
fs.writeFileSync(path.join(root, 'gsd-core', 'workflows', 'clean.md'), 'no globs or counts here\n');
const violations = scanRepo(root);
assert.deepStrictEqual(violations, []);
});
});
// ─── WINDOWS PATH-SEPARATOR NORMALIZATION — the #3223 regression ─────────
//
// `scanTree` (scripts/lib/drift-scan.cjs) builds its repo-relative path via
// `path.relative()`, which uses NATIVE separators. On Windows that is
// `gsd-core\workflows\progress.md`, while the committed baseline
// (`scripts/baselines/planning-prompt-drift-baseline.json`) stores POSIX
// paths — an un-normalized Windows path silently fails to match ANY
// baseline entry, so every violation reports FRESH and every baseline entry
// reports STALE (a 100% guard failure on Windows, caught by GitHub Actions'
// Windows CI lane on PR #3223; the Linux-only remote runner this repo
// otherwise gates on cannot see this class at all).
//
// This coverage drives the pure functions with a Windows-shaped path
// directly — no mocking of the filesystem and NOT gated on
// `process.platform` — so it fails identically on every OS pre-fix and
// passes identically on every OS post-fix. Skipping it on non-Windows would
// recreate the exact blind spot that let this ship.
describe('Windows-shaped repo-relative paths are normalized to POSIX', () => {
const WINDOWS_REL = 'gsd-core\\workflows\\progress.md';
const POSIX_REL = 'gsd-core/workflows/progress.md';
const WINDOWS_LINE = 'X=$(ls dir/*-PLAN.md 2>/dev/null | wc -l)';
test('toPosixRel converts a Windows-shaped separator run to POSIX, and is a no-op on an already-POSIX path', () => {
assert.strictEqual(toPosixRel(WINDOWS_REL), POSIX_REL);
assert.strictEqual(toPosixRel(POSIX_REL), POSIX_REL);
});
test('findPromptDrift on a Windows-shaped relPath reports a POSIX `file`, regardless of input separator', () => {
const out = findPromptDrift(WINDOWS_LINE, WINDOWS_REL);
assert.strictEqual(out.length, 1);
assert.strictEqual(out[0].file, POSIX_REL);
assert.ok(!out[0].file.includes('\\'), 'reported file must carry no backslashes');
});
test('a violation produced from a Windows-shaped path matches a POSIX baseline entry: classified KNOWN, not fresh and not stale', () => {
const baseline = [{ file: POSIX_REL, text: WINDOWS_LINE }];
const violations = findPromptDrift(WINDOWS_LINE, WINDOWS_REL);
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(fresh, []);
assert.deepStrictEqual(stale, []);
});
test('a violation produced from a Windows-shaped path does NOT match if left un-normalized (sanity check the assertion above is meaningful)', () => {
// Same inputs as the previous test, but bypassing toPosixRel to prove the
// KNOWN classification above is actually exercising normalization, not a
// coincidence of the fixture.
const baseline = [{ file: POSIX_REL, text: WINDOWS_LINE }];
const violations = [{ file: WINDOWS_REL, line: 1, found: '*-PLAN.md', text: WINDOWS_LINE }];
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.strictEqual(fresh.length, 1);
assert.strictEqual(stale.length, 1);
});
test('--update (writeBaseline) serializes a POSIX `file` for a Windows-shaped input', (t) => {
const root = createTempDir('gsd-planning-prompt-drift-update-');
t.after(() => cleanup(root));
const violations = findPromptDrift(WINDOWS_LINE, WINDOWS_REL);
writeBaseline(root, violations);
const written = JSON.parse(fs.readFileSync(path.join(root, 'scripts', 'baselines', 'planning-prompt-drift-baseline.json'), 'utf8'));
assert.strictEqual(written.entries.length, 1);
assert.strictEqual(written.entries[0].file, POSIX_REL);
assert.ok(!written.entries[0].file.includes('\\'), 'written baseline entry must carry no backslashes');
});
});
// ─── BASELINE INTEGRITY — both directions, against the real repo ─────────
test('loadBaseline on the committed baseline returns exactly 6 entries (one row per distinct (file, text) pair)', () => {
// 7 total ACKNOWLEDGED occurrences across 6 distinct pairs: plan-phase.md's
// byte-identical DISK_PLANS site fires at two different lines and is
// recorded as ONE row carrying `count: 2` (the Finding-3 fix — a
// duplicated-row baseline made migrating only one of the two sites
// invisible to the ratchet).
const { entries, errors } = loadBaseline(REPO_ROOT);
assert.deepStrictEqual(errors, []);
assert.strictEqual(entries.length, 6);
const totalAcknowledgedOccurrences = entries.reduce((sum, e) => sum + (e.count ?? 1), 0);
assert.strictEqual(totalAcknowledgedOccurrences, 7);
const planPhaseEntry = entries.find((e) => e.file === 'gsd-core/workflows/plan-phase.md');
assert.strictEqual(planPhaseEntry.count, 2);
});
test('scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale', () => {
// The guard's actual contract: every re-derivation this run finds is
// already acknowledged in the baseline, and every baseline entry still
// fires — no fresh, no stale, in either direction.
const violations = scanRepo(REPO_ROOT);
const { entries: baseline, errors } = loadBaseline(REPO_ROOT);
assert.deepStrictEqual(errors, []);
const { fresh, stale } = diffAgainstBaseline(violations, baseline);
assert.deepStrictEqual(fresh, []);
assert.deepStrictEqual(stale, []);
});

View File

@@ -57,11 +57,11 @@ describe('#2351 run-with-timeout — exit-code contract', () => {
});
test('exits 124 when the wall-clock budget is exceeded (matches GNU timeout)', () => {
const start = Date.now();
const r = runVerb(['1', '--', ...HANG]);
// exit 124 is itself the discriminator: a harness backstop kill surfaces
// as status === null (signal), never as 124 — so no elapsed-time
// assertion is needed or allowed here.
assert.equal(r.status, 124, 'a timed-out command must exit 124');
// Sanity: the cap actually fired promptly, not the 30s harness backstop.
assert.ok(Date.now() - start < 15000, 'timeout should fire near the 1s budget');
});
test('exits 127 when the command is not found (matches GNU timeout)', () => {