* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a lower bound for the third consecutive time, and that two derivation families had never been named at all. ADR-3180 gains Decision 7 — a normative behavior contract that says what the right answer IS for each derivation, not merely who owns it. A reviewer with no written rule can only ask "does this look like the others", which is how a fifth copy passes review. Decision 4 gains (d) scan surface is every authored surface and an owner FILE is never exempt, only its named functions; and (e) a surface that cannot be consolidated today ships ratcheted, never unguarded. Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined copies of its own body across five modules. All six now route through it; `clampPercentFromFraction` is added for the one caller that already held a fraction. Every migration is behaviour-identical — clampPercent's first line IS the `total > 0 ? … : 0` ternary each copy carried. Guarded by lint-completion-ratio-drift.cjs, which reports zero re-derivations with no file-level exemption. Prompt layer: workflow markdown re-derives live-plan counting in raw shell (#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs scans it with a shrink-only baseline of the 7 sites that exist today — new sites fail, and a baseline entry that stops firing fails too, so an acknowledgment can never outlive the thing it describes. lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only the four named canonical functions are exempt now. The blanket exemption was pointed at the one file most likely to grow the next copy, and it had. Refs #3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage Five findings from the two orthogonal review passes, all fixed. Decision 4(c) breach: the completion-ratio identity test asserted at the OWNER, which is exactly the bypass that decision exists to close — a consumer can call clampPercent and then post-process locally, leaving both the lint and an owner-level test green. It now drives `roadmap analyze`, `query progress` and `stats` and asserts on their own output, over a fixture containing a `status: superseded` plan so a consumer that re-counted raw files would report 60 where the owner reports 75. Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the issue that removes them. They name Phase 8 (#3218) now. The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical sites were one indistinguishable key and migrating either would have left the guard green with the other alive. Entries carry an occurrence count; fewer than acknowledged fails as a partial migration, more fails as a new copy. Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's test already had, and the fast-check property tests CONTRIBUTING requires for clamp/budget-limit functions. Refs #3180 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes) `tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs` under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`. A fixed wall-clock budget around a double spawn, running inside a container that is concurrently executing the full ~31k-test suite, fails by construction under load. Confirmed against three full matrix runs. Every failure was shaped `null !== 0` — the child was KILLED, never an assertion about the thing under test. One captured probe had already printed the correct resolution (`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The victim subset varies by run and by lane. What these tests are actually about is suite-token RESOLUTION — `unit` as a bare token in --files/--files-from. Executing the seeded trivial files is incidental and is the entire timeout surface, so the assertions move in-process against the same functions `main()` calls, in the same order. `parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are exported for that; no behavior, signature or logic changed. No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the harness for real and asserts exit codes end to end, on a 120s budget. Pre-existing on `next`, fixed here rather than deferred. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: delete the three elapsed-time assertions CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all three are load-sensitive: on a saturated bench each can fail while the code under test is correct. In every case the load-bearing assertion sits on the line above and the timing line adds no discrimination. run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s harness backstop?" — is already answered by the assertion above it. A backstop kills by signal, which surfaces as status null, never 124. Observed directly this session: three matrix runs produced exactly that null shape from killed children. normalize-test-command and context-predicates: both bounded a ReDoS check. A threshold only ever separates "fast" from "slightly slow", which is bench load, not correctness — catastrophic backtracking on 800 KB of input does not take 251ms, it does not finish at all. A real regression therefore shows up as the suite being killed on that test, which is louder and more reliable than a number. The structural assertions (returned unchanged; cleanly rejected) are what actually carry those tests, and they stay. The sweep now reports zero elapsed-time assertions in tests/. The remaining Date.now() uses are unique-path suffixes, barrier deadlines, fixture timestamps and fake mtimes — none of them assertions. Pre-existing on `next`, fixed here rather than deferred. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3180): backfill changeset PR number (#3223) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows The baseline keys on (file, trimmed text). `file` came from scanTree's `path.relative()`, which uses NATIVE separators, while the committed baseline stores POSIX. On Windows every violation was therefore unmatched — reported as FRESH — and every baseline entry matched nothing — reported as STALE. The guard failed 100% of the time there, on both CI shards: ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale + { file: 'gsd-core\\workflows\\execute-plan.md', ... } The remote runner this repo gates on is Linux-only and cannot see this class at all; the GitHub Actions Windows lane is what caught it. Normalization is unconditional — never gated on process.platform. A platform-conditional normalizer makes the POSIX path the special case and leaves the Windows branch unexercised on every other OS, which is the same blind spot in a different place. It is applied at one seam inside findPromptDrift, which builds `file` on every returned violation, so the baseline key, the --update writer, the stderr report and the tests all consume one normalized value. The regression tests drive a Windows-shaped relPath directly and run on every OS rather than skipping off-Windows — a test that only runs on the platform where the bug lives is why this escaped. They include a sanity check that un-normalized input does NOT match, so the assertion cannot pass vacuously. Audited the three sibling guards: none keys against a committed cross-platform baseline, and their exemption keys are path.join-built, so producer and consumer share the native convention. Left correct code alone rather than making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing there would break those three on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
435 lines
20 KiB
JavaScript
435 lines
20 KiB
JavaScript
#!/usr/bin/env node
|
||
'use strict';
|
||
|
||
/**
|
||
* Anti-divergence drift guard for the PROMPT-LAYER plan/summary-COUNTING seam
|
||
* (epic #3180, ADR-3180 "Planning Semantic Model Single Owner", Decision 4(e)).
|
||
*
|
||
* `scripts/lint-plan-count-drift.cjs` and `scripts/lint-milestone-window-drift.cjs`
|
||
* scan `src/` only — but the `.planning/` semantic derivations they own are ALSO
|
||
* re-derived a second time, in the PROMPT layer: the workflow markdown that
|
||
* ships to every runtime, authored as raw shell rather than TypeScript. Issue
|
||
* #1762's second reproduction traced a wrong `30 plans, 24 summaries` figure to
|
||
* a `ls -1 ... *-PLAN.md | wc -l` snippet in `gsd-core/workflows/progress.md` —
|
||
* a re-derivation no `.cts`-scoped guard can see, because it is markdown, not
|
||
* source. ADR-3180 Decision 4(a) requires whole-repo discovery; this guard
|
||
* extends that requirement from "the whole `src/` tree" to "every authored
|
||
* surface that can carry a derivation", covering the prompt layer the two
|
||
* sibling guards structurally cannot reach.
|
||
*
|
||
* Detection is intentionally NARROW, mirroring the sibling guards' precedent:
|
||
* a line is a re-derivation when it carries BOTH, in ONE source line:
|
||
* (a) a plan/summary SET GLOB — a `*` followed by a run of
|
||
* `[-A-Za-z0-9_.{}$]` characters and then the literal `PLAN.md` or
|
||
* `SUMMARY.md`. The leading `*` is load-bearing: it is what makes the
|
||
* line enumerate a SET of files rather than name one specific plan.
|
||
* `gsd-core/workflows/execute-plan.md`'s
|
||
* `grep -cE '^\s*<task[[:space:]>]' .../{phase}-{plan}-PLAN.md` counts
|
||
* TASKS *inside* one already-named plan file — it has no glob token
|
||
* (no `*` anywhere near `PLAN.md`), so it is not a plan-count
|
||
* re-derivation and correctly never matches (a).
|
||
* (b) a COUNTING operation on that same line — `wc -l`, or `grep -c`
|
||
* (optionally with bundled short flags, e.g. `grep -cE`). Reading,
|
||
* globbing, or merely LISTING plan/summary files (`ls *-PLAN.md`,
|
||
* `cat *-PLAN.md`, `--files ".../*-PLAN.md"`) without counting them is
|
||
* not this derivation and must not be flagged — every non-counting
|
||
* `*-PLAN.md`/`*-SUMMARY.md` glob in `gsd-core/workflows/plan-phase.md`
|
||
* (backup, `--files`, `cat`, cross-reference prose) is exactly this
|
||
* shape and is deliberately left alone.
|
||
* `*-UAT.md` never matches (a) — UAT artifacts are a different derivation
|
||
* this guard does not own — so `gsd-core/workflows/progress.md`'s
|
||
* `... *-UAT.md ... | wc -l` line correctly never fires even though it sits
|
||
* one line below two lines that DO.
|
||
*
|
||
* Both regexes are small, bounded, and non-backtracking by construction (a
|
||
* single fixed character class with no nested quantifiers) — `npm run
|
||
* lint:ci` runs CodeQL js/redos over this repo, the same discipline the
|
||
* sibling guards document in their own headers.
|
||
*
|
||
* Surfaces scanned (SCAN_DIRS): `gsd-core/workflows`, `commands`, `agents`,
|
||
* `skills` — the prompt-layer markdown that ships to runtimes. SCAN_EXT:
|
||
* `.md` only. The tree-walk / root-confinement / symlink / sanitizer
|
||
* machinery is SHARED with the two sibling guards via `scripts/lib/drift-scan.cjs`
|
||
* (ADR-3180 Decision 4's own "Rejected: let the new drift guard copy Phase 1's
|
||
* tree-walk / root-confinement / sanitizer") — see that module for the
|
||
* `isInsideRoot` case-sensitivity note, the `walk` symlink-confinement
|
||
* rationale, and the ReDoS-avoidance rationale for its regex-literal reader
|
||
* (unused by this guard's own regexes, which need no literal tokenizer, but
|
||
* shared for the tree walk and report sanitization).
|
||
*
|
||
* RATCHET, not an allowlist. Per ADR-3180 Decision 4(e) this guard's baseline
|
||
* (`scripts/baselines/planning-prompt-drift-baseline.json`) mirrors
|
||
* `scripts/qa-smell-ratchet.cjs`'s precedent exactly: a violation whose
|
||
* `(file, text)` pair is already RECORDED in the baseline is KNOWN and never
|
||
* fails; a violation whose pair is NOT recorded is NEW and fails, telling the
|
||
* author to route the count through the `gsd-core` CLI instead of re-deriving
|
||
* it in shell; a recorded pair that no longer fires in this run is STALE and
|
||
* ALSO fails, forcing `--update` (run by a maintainer after a migration) to
|
||
* prune it — this is what makes the baseline SHRINK-ONLY as call sites
|
||
* migrate off the shell re-derivation, rather than a list that only ever
|
||
* grows. Matching is keyed on the pair (`file`, TRIMMED source `text`), never
|
||
* the line number: a workflow markdown file's line numbers churn on every
|
||
* unrelated edit (a new paragraph, a reworded step) and a number-keyed
|
||
* baseline would need hand-maintenance on changes that have nothing to do
|
||
* with this derivation at all.
|
||
*
|
||
* COUNT, not duplicate rows. Two DIFFERENT source lines can carry the exact
|
||
* same (file, TRIMMED text) pair — `gsd-core/workflows/plan-phase.md` has two
|
||
* byte-identical `DISK_PLANS=$(ls "${PHASE_DIR}"/*-PLAN.md 2>/dev/null | wc -l
|
||
* | tr -d ' ')` sites. Keying on (file, text) alone with one baseline row per
|
||
* OCCURRENCE made a partial migration invisible: migrating ONE of the two
|
||
* sites still leaves a violation matching the row, so nothing goes fresh and
|
||
* nothing goes stale — the remaining, unmigrated copy is silently covered by
|
||
* the row meant to acknowledge the pair NO LONGER MIGRATING. Each baseline
|
||
* entry therefore carries a `count` — the number of byte-identical
|
||
* occurrences of that (file, text) pair acknowledged at this site, not a
|
||
* duplicated row per occurrence:
|
||
* - actual occurrences this run < entry.count -> STALE as a PARTIAL
|
||
* migration: some but not all acknowledged copies are gone, so the entry
|
||
* no longer describes reality and must be re-recorded via `--update`;
|
||
* - actual occurrences this run > entry.count -> the occurrences beyond
|
||
* the acknowledged count are FRESH: a new copy landed next to one that
|
||
* was already acknowledged;
|
||
* - actual occurrences this run === 0 -> fully STALE, the
|
||
* existing "site was migrated, delete the row" case;
|
||
* - actual occurrences this run === entry.count -> fully acknowledged, no
|
||
* failure.
|
||
* Line numbers stay OUT of the key even with counting — that is still what
|
||
* keeps the baseline immune to unrelated churn; `count` answers "how many",
|
||
* never "which lines".
|
||
*
|
||
* KNOWN, ACCEPTED limits of a per-line textual scan (same tradeoff the
|
||
* sibling guards document): a re-derivation whose glob and counting operator
|
||
* are split across two DIFFERENT lines (e.g. a variable holding the glob,
|
||
* counted via `wc -l` on the next line) is not caught by this narrow shape.
|
||
* That is left to code review, not this regex.
|
||
*/
|
||
|
||
const fs = require('node:fs');
|
||
const path = require('node:path');
|
||
const driftScan = require('./lib/drift-scan.cjs');
|
||
const { sanitizeForReport, scanTree } = driftScan;
|
||
|
||
// (a) A plan/summary SET GLOB: a `*` followed by a bounded run of path/brace/
|
||
// var-interpolation characters and then the literal `PLAN.md` or
|
||
// `SUMMARY.md`. The character class is fixed and the quantifier is a single
|
||
// `*` (regex "zero or more", not the shell glob character being matched) over
|
||
// that one class — no nesting, no alternation inside a repeated group, so
|
||
// there is nothing here for a backtracking engine to explore more than once.
|
||
const PLAN_SUMMARY_GLOB_RE = /\*[-A-Za-z0-9_.{}$]*(?:PLAN|SUMMARY)\.md/;
|
||
|
||
// `scanTree` (scripts/lib/drift-scan.cjs) builds its repo-relative path via
|
||
// `path.relative()`, which uses NATIVE separators: on Windows that is
|
||
// `gsd-core\workflows\execute-plan.md`, while the committed baseline
|
||
// (`scripts/baselines/planning-prompt-drift-baseline.json`) stores POSIX
|
||
// paths (`gsd-core/workflows/execute-plan.md`). Every baseline lookup in this
|
||
// guard is keyed on that path, so an un-normalized Windows path silently
|
||
// fails to match ANY baseline entry — every real violation reports as FRESH
|
||
// and every baseline entry reports as STALE (100% failure rate on Windows,
|
||
// caught by GitHub Actions' Windows CI lane on PR #3223; the remote runner
|
||
// this repo otherwise gates on is Linux-only and cannot see this class).
|
||
// Normalized UNCONDITIONALLY — never gated on `process.platform` — because a
|
||
// platform-conditional normalizer is itself the bug: it makes the POSIX path
|
||
// the tested case and leaves the Windows branch exercised only on Windows.
|
||
// Applied at the single seam `findPromptDrift` owns (the only place a
|
||
// repo-relative path enters this guard's violation objects), so ONE
|
||
// normalized value flows into all four consumers: the baseline key
|
||
// (`diffAgainstBaseline`), the `--update` writer (`writeBaseline` via
|
||
// `dedupeViolationsForBaseline`), the violation report (`main`), and the
|
||
// tests.
|
||
function toPosixRel(relPath) {
|
||
return relPath.replace(/\\/g, '/');
|
||
}
|
||
|
||
// (b) A counting operation: `wc -l`, or `grep -c` optionally followed by
|
||
// bundled short flags before the next space (e.g. `grep -cE`, `grep -cE`).
|
||
// `[A-Za-z]{0,4}` bounds the bundled-flag run so the alternative branch is
|
||
// exactly as fixed-width-bounded as `wc -l` — no unbounded quantifier chained
|
||
// to another, so nothing to backtrack.
|
||
const COUNTING_OP_RE = /wc -l|grep -c[A-Za-z]{0,4}\b/;
|
||
|
||
// Prompt-layer markdown that ships to every runtime.
|
||
const SCAN_DIRS = ['gsd-core/workflows', 'commands', 'agents', 'skills'];
|
||
const SCAN_EXT = new Set(['.md']);
|
||
|
||
const BASELINE_REL_PATH = path.join('scripts', 'baselines', 'planning-prompt-drift-baseline.json');
|
||
|
||
// ADR-3180 Decision 4(e): a baseline entry is "acknowledged, in writing, with
|
||
// the issue that owns its removal" — that is Phase 8 (#3218, "the prompt
|
||
// layer": give the workflow layer a CLI surface to ask for plan and phase
|
||
// counts, and burn this ratchet baseline to zero), NOT the epic (#3180)
|
||
// itself. #3180 is the scope authority for the whole consolidation; #3218 is
|
||
// the phase that actually deletes these shell re-derivations.
|
||
const RATCHET_OWNER_ISSUE = '#3218';
|
||
|
||
/**
|
||
* Pure: find every plan/summary-count re-derivation line in `text`.
|
||
* `relPath` is the repo-relative path (native separators or POSIX, either
|
||
* is accepted) — normalized via `toPosixRel` and attached as `file` on every
|
||
* result; this function applies no per-file exemption, so `relPath` is not
|
||
* otherwise consulted for detection.
|
||
* Returns [{ file, line, found, text }] — `file` is always POSIX-separated,
|
||
* `text` is the TRIMMED source line, the same value the baseline keys on.
|
||
*/
|
||
function findPromptDrift(text, relPath) {
|
||
const file = toPosixRel(relPath);
|
||
const out = [];
|
||
const lines = text.split('\n');
|
||
for (let i = 0; i < lines.length; i++) {
|
||
const line = lines[i];
|
||
const globMatch = PLAN_SUMMARY_GLOB_RE.exec(line);
|
||
if (!globMatch) continue;
|
||
if (!COUNTING_OP_RE.test(line)) continue;
|
||
out.push({ file, line: i + 1, found: globMatch[0], text: line.trim() });
|
||
}
|
||
return out;
|
||
}
|
||
|
||
/**
|
||
* Scan the prompt-layer markdown tree and return every re-derivation, each
|
||
* annotated with the repo-relative file path (POSIX-normalized — see
|
||
* `toPosixRel`).
|
||
*/
|
||
function scanRepo(root) {
|
||
return scanTree({
|
||
root,
|
||
scanDirs: SCAN_DIRS,
|
||
scanExt: SCAN_EXT,
|
||
onFile(rel, text) {
|
||
return findPromptDrift(text, rel);
|
||
},
|
||
});
|
||
}
|
||
|
||
/**
|
||
* Read and parse the ratchet baseline. Returns `{ entries, errors }` —
|
||
* `entries` is `[]` and `errors` names the problem when the file is missing,
|
||
* empty, invalid JSON, or malformed; callers in check mode treat a non-empty
|
||
* `errors` as a hard failure (mirrors `qa-smell-ratchet.cjs`'s `readBaseline`).
|
||
*/
|
||
function loadBaseline(root) {
|
||
const baselinePath = path.join(root, BASELINE_REL_PATH);
|
||
if (!fs.existsSync(baselinePath)) {
|
||
return { entries: [], errors: [`${BASELINE_REL_PATH} is missing — run \`node scripts/lint-planning-prompt-drift.cjs --update\` to generate it`] };
|
||
}
|
||
const raw = fs.readFileSync(baselinePath, 'utf8');
|
||
if (raw.trim() === '') {
|
||
return { entries: [], errors: [`${BASELINE_REL_PATH} is present but empty`] };
|
||
}
|
||
let doc;
|
||
try {
|
||
doc = JSON.parse(raw);
|
||
} catch (err) {
|
||
return { entries: [], errors: [`${BASELINE_REL_PATH} is not valid JSON: ${err.message}`] };
|
||
}
|
||
if (doc === null || typeof doc !== 'object' || Array.isArray(doc)) {
|
||
return { entries: [], errors: [`${BASELINE_REL_PATH} must be a JSON object, got ${Array.isArray(doc) ? 'array' : typeof doc}`] };
|
||
}
|
||
if (!Array.isArray(doc.entries)) {
|
||
return { entries: [], errors: [`${BASELINE_REL_PATH}: "entries" must be an array, got ${JSON.stringify(doc.entries)}`] };
|
||
}
|
||
const errors = [];
|
||
const entries = [];
|
||
doc.entries.forEach((entry, i) => {
|
||
const where = `${BASELINE_REL_PATH}.entries[${i}]`;
|
||
if (entry === null || typeof entry !== 'object' || Array.isArray(entry)) {
|
||
errors.push(`${where} must be an object, got ${JSON.stringify(entry)}`);
|
||
return;
|
||
}
|
||
if (typeof entry.file !== 'string' || entry.file === '') {
|
||
errors.push(`${where}.file must be a non-empty string, got ${JSON.stringify(entry.file)}`);
|
||
return;
|
||
}
|
||
if (typeof entry.text !== 'string' || entry.text === '') {
|
||
errors.push(`${where}.text must be a non-empty string, got ${JSON.stringify(entry.text)}`);
|
||
return;
|
||
}
|
||
// `count` is optional on read (diffAgainstBaseline defaults an absent
|
||
// count to 1) but when present must be a positive integer — the number
|
||
// of byte-identical (file, text) occurrences this entry acknowledges.
|
||
if (entry.count !== undefined && !(Number.isInteger(entry.count) && entry.count >= 1)) {
|
||
errors.push(`${where}.count must be a positive integer when present, got ${JSON.stringify(entry.count)}`);
|
||
return;
|
||
}
|
||
entries.push(entry);
|
||
});
|
||
return { entries, errors };
|
||
}
|
||
|
||
/**
|
||
* Diff scanned `violations` (from `scanRepo`) against baseline `entries`,
|
||
* matched by the pair (`file`, TRIMMED `text`) — never the line
|
||
* number — and COUNT-aware: an entry acknowledges `entry.count`
|
||
* (default 1 when absent) byte-identical occurrences of that pair, not
|
||
* merely its presence. Returns `{ fresh, stale }`:
|
||
* - `fresh`: violations whose (file, text) pair is NOT in the baseline at
|
||
* all (a brand new site), PLUS any occurrences of a KNOWN pair beyond
|
||
* its acknowledged `count` (a new copy landed next to an
|
||
* already-acknowledged one) — both fail the build as NEW.
|
||
* - `stale`: baseline entries whose actual occurrence count this run is
|
||
* LESS than their acknowledged `count` — zero actual
|
||
* occurrences is the fully-migrated case ("site was migrated, delete
|
||
* the row"); a positive but short count is a PARTIAL migration (some
|
||
* but not all acknowledged copies are gone). Both fail the build,
|
||
* forcing `--update` to re-record the pair (this is what keeps the
|
||
* baseline shrink-only and what makes a partial migration visible
|
||
* instead of silently covered by the still-present sibling
|
||
* occurrence).
|
||
*/
|
||
function diffAgainstBaseline(violations, baseline) {
|
||
const key = (file, text) => `${file} |