* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking answer. Three families do exactly that today; this commit pins each one RED. Measured on this tree, 2026-08-27: intel query, .planning/intel/file-roles.json nested 12000 deep -> exit 1, "Error: Maximum call stack size exceeded" searchJsonEntries / matchesInValue carry no depth parameter at all. The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it. same fixture nested 48 and 49 deep -> both return total=1 at exit 0, truncated=undefined Nothing distinguishes "searched to the bottom" from "stopped looking". query phase-plan-index, a plan whose depends_on names an unresolvable token -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] The token is never mentioned. computeDependencyLevels drops the edge with `if (!resolvedDep) continue;`, every plan becomes a root, and the tool then reports the author's correct wave: as the thing that is wrong. countPhasePlansAndSummaries with fs.readdirSync throwing EACCES -> hasContext:false, indistinguishable from a phase that simply has no CONTEXT.md. context_read_error is undefined. The shapes these tests assert against, chosen here so the implementation has a target rather than inventing one later: `truncated: boolean` on the intel query result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and `context_read_error: string | null` per analyzed phase. Deliberately green, and they must stay that way — each stops the fix from over-firing: depth 48 is found and NOT flagged truncated (the ceiling is inclusive) a shallow miss reports no truncation (noise control, N1) 10,000 siblings at depth 2 are unaffected (the bound is DEPTH, N2) a genuine wave: mismatch on a fully-resolved DAG still warns (N3) a genuinely missing directory is absent, not an error the emitted depends_on display mapping still passes an unresolved token through verbatim — already pinned by the existing #3785 test, so no duplicate was added T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the real CLI and reads the emitted JSON, because a unit assertion on computeDependencyLevels would have passed throughout #3427's life. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3885): no silent swallow, and no verdict manufactured from dropped data Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed into an output that reads as authoritative. The recursion bound, restored but NOT verbatim (src/intel.cts) MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and matchesInValue, which carried no depth parameter at all. The bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it — §8.3's "a consolidation may not delete an invariant along with the surface that held it", demonstrated. Measured before: a .planning/intel file nested 12000 deep exits 1 with "Error: Maximum call stack size exceeded". Reachable from a project document. The original returned a bare `false` at the ceiling. Restoring that verbatim would trade a crash for a silent "no match" when the truth is "I stopped looking" — the same class this epic exists to close, and ADR-3473 Decision 4 forbids it. So the bound carries a truncation signal: nesting 47 -> found, truncated false nesting 48 -> found, truncated false (the ceiling is inclusive) nesting 49 -> not found, truncated TRUE nesting 12000 -> exit 0, truncated TRUE, no RangeError A shallow document that simply has no match reports truncated FALSE — the flag means "I stopped early", never "I found nothing", or it would be noise. The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected. The dropped edge is named, and stops being blamed on the author (src/phase.cts) computeDependencyLevels dropped every unresolvable depends_on token with a bare `continue`. Each drop makes a plan a root, so the whole phase collapses to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as the thing that was wrong. Before: warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] After: warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not resolve to any plan in this phase — edge dropped, wave placement for this plan may be unreliable"] The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is deliberately not built here. The emitted depends_on display mapping still passes an unresolved token through verbatim (#3785). No artifact from failed inputs (gsd-core/workflows/review.md, #3352) A failed lane leaves no result file, so "every lane failed" is exactly "the aggregate JSONL has zero lines" — the gate condition already existed as a byproduct. REVIEWS.md is no longer written in that case, and the commit step is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT counted as a failure. Per-lane output and non-empty .err are preserved to .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that the lanes failed at all; the commit step names one file, never a glob, so the diagnostics are not swept in. Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2) Four callers collapsed an EACCES on a phase directory into [] and reported hasContext:false — byte-identical to a phase that simply has no CONTEXT.md. Each now names the directory it could not read. A genuinely missing directory stays absent rather than becoming an error, which is what keeps the fix from over-firing. Fatal errno folded into a retry set: audited, no defect found Reported as a verified negative rather than padded with a change. withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776; atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they return or rethrow the final error rather than swallowing it. estimate-cli's sole caller surfaces that rethrow as write_error in its JSON output. Manufacturing a diff to make the checkbox look worked-on is the Goodhart outcome Decision 6 exists to prevent. Disclosed: R46 (the commit step names one file, never a glob) is a real regression guard but is NOT independently failing-first — the commit fence is byte-identical pre- and post-fix, so it only fails pre-fix through its shared extraction dependency. Recorded rather than claimed as fail-first. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence Two review findings, both real, both in my own change. An isolated adversarial review found the evidence-preservation block never checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in a SEPARATE fenced block. A disk-full or unwritable phase directory therefore still destroyed the only copy of the failed lanes' output — reintroducing the exact #3352 data loss this item exists to stop, inside the fix for it. Preservation and cleanup are now one block, because each fenced block is a separate execution and a shell variable cannot carry between them. mkdir -p and each cp are exit-checked; cleanup runs only when preservation succeeded, and a failure warns naming the intact run directory. "Nothing to preserve" is not a failure and still cleans up. Driven three ways: success removes run_dir, failure leaves it intact with the warning, nothing-to-preserve removes it. The failure is induced by a file-vs-directory conflict rather than chmod 0o000, which root bypasses. The new unresolved-depends_on warning embedded a user-authored token verbatim: warnings: ["Plan 03-02: depends_on token \"evil Plan 03-01: FORGED WARNING\" does not resolve ..."] The JSON wire form is safe, and the security reviewer judged it non-exploitable for that reason. It is escaped anyway through formatDiagnosticToken — the helper #3884 added one phase earlier for exactly this class. warnings[] is an array a consumer naturally prints line by line, and not reusing the sibling fix is the generative-fix-divergence shape this epic exists to close. The same treatment is applied to context_read_error / phase_dir_read_error, which embed a phase directory path a repository can choose, and to the fs error message, which echoes the raw path itself. Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per negative space N2, disclosed rather than left to be discovered. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot" Blocker from the round-2 isolated review, and it is my own inconsistency: this phase applied "unreadable is not absent" to phase directories and left it broken in the file it was already editing. chmod 000 .planning/intel/file-roles.json gsd-tools intel query <term> -> {"matches":[],"total":0,"truncated":false} exit 0 safeReadJson swallowed every read failure and returned null, so an EACCES was byte-indistinguishable from an absent file AND from a genuine no-match. Now it separates three states: ENOENT stays silently absent, because not every project has every intel file and intelQuery loops over all of them expecting misses; EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel file previously read as "no matches" too — same defect, same fix. Threading that outcome through the other three callers found something worse than the reported case. intelDiff returned no_baseline:true for a corrupt or unreadable snapshot — not a silent failure but an actively FALSE verdict, telling the caller they never took a snapshot when they did. That is §8.5's headline case, so it is fixed and tested rather than noted. intelStatus and intelApiSurface collapsed the same way; intelApiSurface additionally printed a "not yet populated" banner that was simply untrue. Every row is failing-first, including the absent-file ones — the field is new, so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin that the fix does not OVER-fire on the ordinary absent case, which is what would turn this into noise on every project lacking an intel file. IO failure is injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as a test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): build the pathological intel fixture as text, not by stringifying a nested object The remote runner came back red on Linux with two failures, both T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on macOS. The product was never at fault. writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned. Linux's container stack is smaller than macOS's, which is the whole of the platform difference. Measured, with the same document built as JSON TEXT so nothing in the building process recurses: depth=100 rc=0 truncated=true depth=5000 rc=0 truncated=true depth=12000 rc=0 truncated=true depth=60000 rc=0 truncated=true V8 parses this shape iteratively; only stringify recurses. The bound works at every depth tried. The fixture is now built by string concatenation. That is also the more faithful input — a real deeply nested JSON document on disk is exactly what the bound guards, where a stringified object was only ever a way to produce one. The depth stays 12000. Lowering it would have made the test pass by weakening it to accommodate a fixture bug, and 12000 is a legitimate pathological input the product handles. T4 remains a genuine fail-first: rebuilt against the parent of the commit that added the bound, the string-built depth-12000 fixture still drives the CLI to rc=1 with "Error: Maximum call stack size exceeded". A comment records why the fixture is text, so it is not "simplified" back into a macOS-green / Linux-red test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3885): backfill the changeset PR number Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): normalize path separators before splicing into the workflow's bash CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the remote runner were all green. AssertionError: commit must name the single REVIEWS.md file; got: --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The harness spliced an OS-native temp path into the extracted bash, and bash consumes \U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory survived — which is the other two assertions. This is a fixture defect, not a product one, and that was checked rather than assumed. In production the phase directory is toPosixPath-normalized at every call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and the run directory is created by "mktemp -d" running inside the bash block itself (gsd-core/workflows/review.md:163), which emits POSIX-style output even under Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it. The file's pre-existing #3034 harness splices raw native paths too, but only ever inside double-quoted assignments, so it never tripped this — my new harness followed that convention faithfully into the one place where it does not hold. Both now splice through toPosixPath from shell-command-projection, the established seam, which is a no-op on POSIX and mirrors what production does. No assertion was weakened. "commit must name the single REVIEWS.md file" and "the run dir must still be destroyed" still assert exactly that; only how the fixture supplies its path changed. Nothing is skipped on Windows — a t.skip() here would have hidden the question of whether the exposure was real, which is the question that mattered. Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input produces a byte-identical shape, proving the normalization is idempotent. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): stop the harness making the deleted run dir its own cwd Windows shard 3/3 stayed red after the separator fix, on two assertions the separator fix never touched: AssertionError: the run dir must still be destroyed AssertionError: nothing to preserve is not a failure — run dir must still be removed The separators were a real bug and fixing them fixed the --files assertion. They were not this bug, and two CI cycles went into the wrong axis before I stopped converting path forms and looked at what the harness actually does. runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's working directory WAS the directory the block under test then removes with rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live process's working directory cannot be removed. So on Windows the directory survived and both assertions failed, on macOS and Linux it vanished and they passed. Nothing to do with slashes. Harness-only. Production never cd's into the run directory; every reference is by absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself (gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged. Fix: the child now runs with its cwd in an unrelated temp directory that the block under test never deletes. Neither assertion was weakened, and nothing is skipped on Windows — the tests in this file carry no platform guard and run there unconditionally, which is how this surfaced at all. Honest limit: the Windows failure mode cannot be reproduced on macOS, because POSIX permits the very thing Windows refuses. The diagnosis is grounded in that documented divergence and in the fact that only the Windows lane failed, but the green outcome on windows-latest is unverified until CI runs it. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
529 lines
23 KiB
TypeScript
529 lines
23 KiB
TypeScript
/**
|
|
* Post-planning gap analysis (#2493).
|
|
*
|
|
* Reads REQUIREMENTS.md (planning-root) and CONTEXT.md (per-phase) and compares
|
|
* each REQ-ID and D-ID against the concatenated text of all PLAN.md files in
|
|
* the phase directory. Emits a unified `Source | Item | Status` report.
|
|
*
|
|
* Gated on workflow.post_planning_gaps (default true). When false, returns
|
|
* { enabled: false } and does not scan.
|
|
*
|
|
* Coverage detection uses word-boundary regex matching to avoid false positives
|
|
* (REQ-1 must not match REQ-10).
|
|
*
|
|
* ADR-457 build-at-publish: the hand-written bin/lib/gap-checker.cjs collapsed
|
|
* to a TypeScript source of truth. Behaviour is preserved byte-for-behaviour
|
|
* from the prior hand-written .cjs; only strict types are added.
|
|
*/
|
|
|
|
import fs from 'node:fs';
|
|
import path from 'node:path';
|
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
import io = require('./io.cjs');
|
|
const { output, error, formatDiagnosticToken } = io;
|
|
import { escapeRegex } from './pattern.cjs';
|
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
import planningWorkspace = require('./planning-workspace.cjs');
|
|
const { planningPaths, planningDir, findContextMdIn } = planningWorkspace;
|
|
import { parseDecisions, extractDecisions } from './decisions.cjs';
|
|
import { iterateBullets } from './markdown-sectionizer.cjs';
|
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
import planScanMod = require('./plan-scan.cjs');
|
|
const { scanPhasePlans } = planScanMod;
|
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
import phaseIdMod = require('./phase-id.cjs');
|
|
const { scopeToPhase } = phaseIdMod;
|
|
|
|
// ─── Types ────────────────────────────────────────────────────────────────────
|
|
|
|
interface ReqItem {
|
|
id: string;
|
|
text: string;
|
|
}
|
|
|
|
interface RequirementItem extends ReqItem {
|
|
source: string;
|
|
}
|
|
|
|
type DecisionItem = ReturnType<typeof parseDecisions>[number] & { source: string };
|
|
|
|
type Item = RequirementItem | DecisionItem;
|
|
|
|
interface CoverageRow {
|
|
source: string;
|
|
item: string;
|
|
status: string;
|
|
}
|
|
|
|
interface GapCounts {
|
|
total: number;
|
|
covered: number;
|
|
uncovered: number;
|
|
}
|
|
|
|
interface GapResult {
|
|
enabled: boolean;
|
|
rows: CoverageRow[];
|
|
table: string;
|
|
summary: string;
|
|
counts: GapCounts;
|
|
/**
|
|
* #3885 (ADR-3473 §8.5): null when the phase directory is genuinely absent
|
|
* (guarded by `fs.existsSync` before the read, so readdirSync is never even
|
|
* attempted) or was read successfully. A message naming the phase directory
|
|
* when it EXISTS but `readdirSync` failed (EACCES/EIO/...) — never
|
|
* collapsed to the same `[]` an absent directory produces.
|
|
*/
|
|
phase_dir_read_error: string | null;
|
|
}
|
|
|
|
interface RunGapAnalysisOptions {
|
|
phaseReqIds?: string | null | undefined;
|
|
}
|
|
|
|
/**
|
|
* Parse REQ-IDs from REQUIREMENTS.md content.
|
|
*
|
|
* Supports both checkbox (`- [ ] **REQ-NN** ...`) and traceability table
|
|
* (`| REQ-NN | ... |`) formats.
|
|
*/
|
|
function parseRequirements(reqMd: unknown): ReqItem[] {
|
|
if (!reqMd || typeof reqMd !== 'string') return [];
|
|
const out: ReqItem[] = [];
|
|
const seen = new Set<string>();
|
|
|
|
// Prefix-agnostic ID format: REQ-01, TST-01, BACK-07, INSP-04, etc.
|
|
const ID_PATTERN = '[A-Z][A-Z0-9]*-[A-Za-z0-9_-]+';
|
|
const idRe = new RegExp(`^(${ID_PATTERN})$`);
|
|
|
|
// Checkbox-bullet path: migrate to seam's iterateBullets (checkbox markers).
|
|
// The **ID** is extracted from the bullet text caller-side — the seam provides
|
|
// the raw text; we parse the bold-ID prefix from it here.
|
|
const boldIdRe = new RegExp(`^\\*\\*(${ID_PATTERN})\\*\\*\\s*(.*)$`);
|
|
// The shipped template (gsd-core/templates/requirements.md) writes
|
|
// `- [ ] **AUTH-01**: User can sign up` — a single separator delimiter
|
|
// between the bold ID and the description. Strip AT MOST ONE leading
|
|
// delimiter (plus its surrounding whitespace) before the final `.trim()`,
|
|
// mirroring roadmap-parser.cts's `stripLeadingDelimiter` delimiter set
|
|
// (em dash, en dash, colon, hyphen) — that helper is not exported, and its
|
|
// own `+`-quantified strip removes an entire delimiter RUN, which would
|
|
// also eat a second, meaningful marker (`**X-01**: -- weird` must keep the
|
|
// `--`), so the set is mirrored here with a single-occurrence match instead
|
|
// of reused verbatim.
|
|
const ONE_LEADING_DELIMITER_RE = /^\s*[—–:-]\s*/;
|
|
for (const bullet of iterateBullets(reqMd)) {
|
|
if (bullet.marker !== 'checkbox-unchecked' && bullet.marker !== 'checkbox-checked') continue;
|
|
const m = boldIdRe.exec(bullet.text);
|
|
if (!m) continue;
|
|
const id = m[1];
|
|
if (!idRe.test(id)) continue;
|
|
if (!seen.has(id)) {
|
|
seen.add(id);
|
|
const rawText = m[2] || '';
|
|
out.push({ id, text: rawText.replace(ONE_LEADING_DELIMITER_RE, '').trim() });
|
|
}
|
|
}
|
|
|
|
// Pipe-table-row path and separator-row skip stay caller-side
|
|
// (table parsing is out of seam scope per ADR-1372 T3 spec).
|
|
const tableFirstCellRe = new RegExp(`^\\s*\\|\\s*(${ID_PATTERN})\\s*\\|`);
|
|
const separatorRowRe = /^\s*\|[\s:|-]+\|\s*$/;
|
|
const lines = reqMd.split(/\r?\n/);
|
|
|
|
for (let i = 0; i < lines.length; i += 1) {
|
|
const line = lines[i];
|
|
if (!line.includes('|')) continue;
|
|
|
|
// Skip markdown table separator rows and header rows immediately preceding them.
|
|
if (separatorRowRe.test(line)) continue;
|
|
if (i + 1 < lines.length && separatorRowRe.test(lines[i + 1])) continue;
|
|
|
|
const tm = tableFirstCellRe.exec(line);
|
|
if (!tm) continue;
|
|
const id = tm[1];
|
|
if (!seen.has(id)) {
|
|
seen.add(id);
|
|
out.push({ id, text: '' });
|
|
}
|
|
}
|
|
|
|
return out;
|
|
}
|
|
|
|
function detectCoverage(items: Item[], planText: string): CoverageRow[] {
|
|
return items.map(it => {
|
|
const re = new RegExp('\\b' + escapeRegex(it.id) + '\\b');
|
|
return {
|
|
source: it.source,
|
|
item: it.id,
|
|
status: re.test(planText) ? 'Covered' : 'Not covered',
|
|
};
|
|
});
|
|
}
|
|
|
|
function naturalKey(s: unknown): string {
|
|
return String(s).replace(/(\d+)/g, (_, n: string) => n.padStart(8, '0'));
|
|
}
|
|
|
|
function sortRows(rows: CoverageRow[]): CoverageRow[] {
|
|
const sourceOrder: Record<string, number> = { 'REQUIREMENTS.md': 0, 'CONTEXT.md': 1 };
|
|
return rows.slice().sort((a, b) => {
|
|
const so = (sourceOrder[a.source] ?? 99) - (sourceOrder[b.source] ?? 99);
|
|
if (so !== 0) return so;
|
|
return naturalKey(a.item).localeCompare(naturalKey(b.item));
|
|
});
|
|
}
|
|
|
|
function formatGapTable(rows: CoverageRow[]): string {
|
|
if (rows.length === 0) {
|
|
return '## Post-Planning Gap Analysis\n\nNo requirements or decisions to check.\n';
|
|
}
|
|
const header = '| Source | Item | Status |\n|--------|------|--------|';
|
|
const body = rows.map(r => {
|
|
const tick = r.status === 'Covered' ? '✓ Covered'
|
|
: r.status === 'Missing from REQUIREMENTS.md' ? '⚠ Missing from REQUIREMENTS.md'
|
|
: '✗ Not covered';
|
|
return `| ${r.source} | ${r.item} | ${tick} |`;
|
|
}).join('\n');
|
|
return `## Post-Planning Gap Analysis\n\n${header}\n${body}\n`;
|
|
}
|
|
|
|
function readGate(cwd: string): boolean {
|
|
const cfgPath = path.join(planningDir(cwd), 'config.json');
|
|
try {
|
|
const raw = JSON.parse(fs.readFileSync(cfgPath, 'utf-8')) as unknown;
|
|
if (raw && typeof raw === 'object' && 'workflow' in raw) {
|
|
const wf = (raw as Record<string, unknown>)['workflow'];
|
|
if (wf && typeof wf === 'object' && 'post_planning_gaps' in wf) {
|
|
const val = (wf as Record<string, unknown>)['post_planning_gaps'];
|
|
if (typeof val === 'boolean') return val;
|
|
}
|
|
}
|
|
} catch { /* fall through */ }
|
|
return true;
|
|
}
|
|
|
|
/**
|
|
* Same-prefix ascending numeric range, e.g. `SEL-01..SEL-03`. Both sides must
|
|
* share an identical prefix and a numeric suffix. Captures are:
|
|
* 1 low prefix, 2 low digits, 3 high prefix (compared to group 1 for equality), 4 high digits.
|
|
*/
|
|
const PHASE_REQ_RANGE_RE = /^(.+-)(\d+)\.\.(.+-)(\d+)$/;
|
|
|
|
/**
|
|
* Maximum number of IDs a single range token may expand to. A range whose span
|
|
* exceeds this cap stays literal (fail-closed) rather than expanding, guarding
|
|
* against pathological input like `X-1..X-100000` ballooning the comparison set.
|
|
*/
|
|
const MAX_PHASE_REQ_RANGE = 1000;
|
|
|
|
/**
|
|
* Shape filter applied to `--phase-req-ids` tokens AFTER range expansion
|
|
* (#3189). ROADMAP `**Requirements:**` lines routinely carry prose trailing
|
|
* the real ID list — locked-decision annotations, ambiguity scores,
|
|
* prohibitions, dates — and the function's own contract says callers may pass
|
|
* that roadmap value through verbatim. Without a shape filter every prose word
|
|
* survives the whitespace split and is reported as an individually-missing
|
|
* requirement, drowning the real coverage signal.
|
|
*
|
|
* The filter is intentionally wider than `ID_PATTERN` (which mandates a hyphen
|
|
* and so would reject the hyphen-less `R1`..`R8` family real roadmaps use).
|
|
* Three load-bearing details (#3189):
|
|
*
|
|
* 1. The digit lookahead `(?=.*\d)` matters. Without it ALL-CAPS prose tokens
|
|
* that appear in annotations (`LOCKED`, `TBD`-as-prose, `NONE`-as-prose)
|
|
* would pass and still reach the report as fake IDs. Every real
|
|
* requirement ID carries a digit (`R1`, `SEL-01`, `P1-P3`).
|
|
* 2. It MUST run after `expandPhaseReqIdToken`. Filtering before would
|
|
* discard the range tokens themselves (`SEL-01..SEL-03` matches no
|
|
* single-ID shape), silently undoing the #1269 / #1419 range expansion.
|
|
* 3. It is NOT reusable as `ID_PATTERN`. `ID_PATTERN` mandates a hyphen,
|
|
* so it rejects `R1`..`R8` and would leave only range-shaped tokens
|
|
* standing — a silently-empty report, which is worse.
|
|
*
|
|
* Accepts both hyphen-less digit-bearing IDs (`R1`, `R8`) and prefix-hyphen
|
|
* IDs (`REQ-01`, `SEL-01`, `BACK-07`), including multi-segment prefixes
|
|
* (`REQ2-01`). Rejects prose (`LOCKED`), punctuation (`—`, `##`, `+`),
|
|
* dates, version-likes (`0.12;`), backtick fragments, and any token still
|
|
* containing `.` (so invalid range tokens returned literal by
|
|
* `expandPhaseReqIdToken` are dropped rather than surfaced as fake IDs).
|
|
*/
|
|
const PHASE_REQ_ID_SHAPE_RE = /^(?=.*\d)[A-Z][A-Z0-9]*(?:[-_][A-Za-z0-9]+)*$/;
|
|
|
|
/**
|
|
* Expand a single `--phase-req-ids` token in place. If it is a valid ascending
|
|
* same-prefix numeric range (`<PREFIX>-NN..<PREFIX>-MM`, identical prefix both
|
|
* sides, numeric NN ≤ MM), return the individual IDs `<PREFIX>-NN … <PREFIX>-MM`
|
|
* preserving the bounds' zero-pad width. Anything that does NOT cleanly match a
|
|
* valid range stays literal (fail-closed) — returned as a single-element array.
|
|
*
|
|
* The two numeric bounds must share the same digit width; a range with
|
|
* differing widths (e.g. `SEL-9..SEL-11`) is ambiguous (padding to the wider
|
|
* width could invent IDs like `SEL-09` that never appear unpadded in
|
|
* REQUIREMENTS) and is left literal. A range spanning more than
|
|
* MAX_PHASE_REQ_RANGE IDs also stays literal.
|
|
*/
|
|
function expandPhaseReqIdToken(token: string): string[] {
|
|
const m = PHASE_REQ_RANGE_RE.exec(token);
|
|
if (!m) return [token];
|
|
const [, prefixLow, lowDigits, prefixHigh, highDigits] = m;
|
|
// Fail closed unless the prefixes are identical.
|
|
if (prefixLow !== prefixHigh) return [token];
|
|
// Fail closed unless the bounds share an identical digit width. Differing
|
|
// widths are ambiguous: padding to the wider width could invent IDs that
|
|
// never appear unpadded in REQUIREMENTS.
|
|
if (lowDigits.length !== highDigits.length) return [token];
|
|
const low = Number(lowDigits);
|
|
const high = Number(highDigits);
|
|
// Fail closed on descending ranges (NN > MM). NN == MM is a valid single-element range.
|
|
if (!Number.isFinite(low) || !Number.isFinite(high) || low > high) return [token];
|
|
// Fail closed (DoS guard) on ranges spanning more than the cap.
|
|
if (high - low + 1 > MAX_PHASE_REQ_RANGE) return [token];
|
|
// Preserve the bounds' (shared) zero-pad width.
|
|
const width = lowDigits.length;
|
|
const out: string[] = [];
|
|
for (let n = low; n <= high; n++) {
|
|
out.push(`${prefixLow}${String(n).padStart(width, '0')}`);
|
|
}
|
|
return out;
|
|
}
|
|
|
|
/**
|
|
* Normalize a raw `--phase-req-ids` argument into the scoping signal used by
|
|
* runGapAnalysis (#447). Mirrors §13's null/TBD skip semantics.
|
|
*
|
|
* undefined → flag absent: compare the whole REQUIREMENTS.md (back-compat)
|
|
* null | '' | TBD → no requirements mapped to this phase: skip the comparison
|
|
* "REQ-01,REQ-02" → restrict the comparison to these IDs
|
|
* "SEL-01..SEL-03" → range form: expands in place to SEL-01, SEL-02, SEL-03 (#1269)
|
|
*
|
|
* Range form (#1269): a list element of the shape `<PREFIX>-NN..<PREFIX>-MM`
|
|
* (identical prefix both sides, identical bound digit width, ascending numeric
|
|
* NN ≤ MM) is expanded in place to the individual IDs, preserving the bounds'
|
|
* zero-pad width; mixed lists expand in input order. Any element that does not
|
|
* cleanly match a valid ascending same-prefix numeric range (mismatched
|
|
* prefix, differing bound width, descending, non-numeric, missing bound, or
|
|
* spanning more than MAX_PHASE_REQ_RANGE IDs) stays literal — no partial
|
|
* expansion, no guessing.
|
|
*
|
|
* ID-shape filter (#3189): STRICTLY AFTER range expansion, every token is
|
|
* filtered through `PHASE_REQ_ID_SHAPE_RE`. ROADMAP `**Requirements:**` lines
|
|
* routinely carry prose trailing the real ID list (locked-decision annotations,
|
|
* ambiguity scores, prohibitions, dates); without the filter every prose word
|
|
* would be reported as an individually-missing requirement. Tokens that cannot
|
|
* be requirement IDs (prose, punctuation, dates, invalid range tokens left
|
|
* literal by `expandPhaseReqIdToken`) are dropped; an input whose every token
|
|
* is dropped collapses to `null` (skip semantics).
|
|
*
|
|
* Tolerates JSON-array-ish input (`["REQ-01","REQ-02"]`) since callers may pass
|
|
* the roadmap value through verbatim.
|
|
*/
|
|
function normalizePhaseReqIds(rawVal: unknown): string[] | null | undefined {
|
|
if (rawVal === undefined) return undefined;
|
|
if (rawVal === null) return null;
|
|
// eslint-disable-next-line @typescript-eslint/no-base-to-string
|
|
const v = String(rawVal).replace(/["'[\]()]/g, '').trim();
|
|
if (v === '' || /^(null|tbd|none)$/i.test(v)) return null;
|
|
// Tolerate comma-, space-, or newline-separated lists (callers may pass the
|
|
// roadmap value verbatim, whose serialization is not guaranteed).
|
|
const ids = v.split(/[\s,]+/).map(s => s.trim()).filter(Boolean);
|
|
// Expand range tokens (#1269) per-token AFTER the split, preserving input order.
|
|
const expanded = ids.flatMap(expandPhaseReqIdToken);
|
|
// #3189: drop tokens that cannot be requirement IDs (prose, punctuation,
|
|
// dates, invalid range tokens left literal by expandPhaseReqIdToken). MUST
|
|
// run after expand so valid ranges still expand (filtering before would
|
|
// discard `SEL-01..SEL-03` itself, undoing #1269 / #1419).
|
|
const filtered = expanded.filter(t => PHASE_REQ_ID_SHAPE_RE.test(t));
|
|
return filtered.length === 0 ? null : filtered;
|
|
}
|
|
|
|
function runGapAnalysis(cwd: string, phaseDir: string, options: RunGapAnalysisOptions = {}): GapResult {
|
|
const phaseReqIds = normalizePhaseReqIds(options.phaseReqIds);
|
|
if (!readGate(cwd)) {
|
|
return {
|
|
enabled: false,
|
|
rows: [],
|
|
table: '',
|
|
summary: 'workflow.post_planning_gaps disabled — skipping post-planning gap analysis',
|
|
counts: { total: 0, covered: 0, uncovered: 0 },
|
|
phase_dir_read_error: null,
|
|
};
|
|
}
|
|
|
|
const absPhaseDir = path.isAbsolute(phaseDir) ? phaseDir : path.join(cwd, phaseDir);
|
|
|
|
const reqPath = planningPaths(cwd).requirements;
|
|
const reqMd = fs.existsSync(reqPath) ? fs.readFileSync(reqPath, 'utf-8') : '';
|
|
let reqItems: RequirementItem[] = parseRequirements(reqMd).map(r => ({ ...r, source: 'REQUIREMENTS.md' }));
|
|
|
|
// Scope the requirements comparison to the phase's mapped REQ-IDs (#447).
|
|
// A phase that maps no requirements (phase_req_ids null/TBD) must not report
|
|
// every unrelated project REQ-ID as a gap — mirror §13's skip behavior.
|
|
// CONTEXT.md decisions (below) are always in scope regardless.
|
|
let ghostReqIds: string[] = [];
|
|
if (phaseReqIds === null) {
|
|
reqItems = [];
|
|
} else if (Array.isArray(phaseReqIds)) {
|
|
const wanted = new Set(phaseReqIds);
|
|
const foundIds = new Set(reqItems.map(r => r.id));
|
|
reqItems = reqItems.filter(r => wanted.has(r.id));
|
|
ghostReqIds = phaseReqIds.filter(id => !foundIds.has(id));
|
|
}
|
|
|
|
// Read the phase directory once; reuse the listing for both context detection
|
|
// and plan-file enumeration (avoids redundant readdirSync calls).
|
|
let phaseDirFiles: string[] = [];
|
|
// #3885 (ADR-3473 §8.5): the existsSync guard above already means a catch
|
|
// here is NEVER "genuinely absent" (ENOENT) — this directory exists, so any
|
|
// failure to list it is a real read error (EACCES/EIO/...) and must be
|
|
// named, not folded into the same `[]` an absent directory produces.
|
|
let phaseDirReadError: string | null = null;
|
|
try {
|
|
if (fs.existsSync(absPhaseDir)) phaseDirFiles = fs.readdirSync(absPhaseDir);
|
|
} catch (err) {
|
|
phaseDirReadError = `Could not read phase directory ${formatDiagnosticToken(absPhaseDir)}: ${formatDiagnosticToken((err as Error)?.message ?? String(err))}`;
|
|
}
|
|
|
|
// #3511-class: scope the raw listing to this phase dir before the
|
|
// phase-numbered -CONTEXT.md predicate. `phaseDirFiles` itself stays raw —
|
|
// it is also reused below only as a `.length > 0` guard ahead of
|
|
// scanPhasePlans's own plan-sequence-numbered enumeration, a different
|
|
// grammar that must not be scoped by phase number.
|
|
const scopedPhaseDirFiles = scopeToPhase(phaseDirFiles, path.basename(absPhaseDir));
|
|
const ctxFile = findContextMdIn(scopedPhaseDirFiles);
|
|
const ctxPath = ctxFile ? path.join(absPhaseDir, ctxFile) : null;
|
|
const ctxMd = ctxPath ? fs.readFileSync(ctxPath, 'utf-8') : '';
|
|
|
|
// Use extractDecisions so gap-checker can distinguish could-not-parse from none-present.
|
|
const ctxExtraction = extractDecisions(ctxMd);
|
|
const dItems: DecisionItem[] = ctxExtraction.decisions.map(d => ({ ...d, source: 'CONTEXT.md' }));
|
|
|
|
const items: Item[] = [...reqItems, ...dItems];
|
|
|
|
let planText = '';
|
|
try {
|
|
if (phaseDirFiles.length > 0) {
|
|
// #3183 (lint-plan-count-drift): source the live plan-file list from
|
|
// the single owner (scanPhasePlans) instead of a local `-PLAN\.md$`
|
|
// filter on the already-read listing — picks up bare PLAN.md, nested
|
|
// plans/, and excludes superseded plans, none of which the prior
|
|
// root-only exact-suffix filter did.
|
|
const files = scanPhasePlans(absPhaseDir).planFiles;
|
|
planText = files.map(f => {
|
|
try { return fs.readFileSync(path.join(absPhaseDir, f), 'utf-8'); }
|
|
catch { return ''; }
|
|
}).join('\n');
|
|
}
|
|
} catch { /* unreadable */ }
|
|
|
|
// FIX D (#1365): surface decision could-not-parse independently of whether
|
|
// requirements items exist. Without this, a could-not-parse on decisions is
|
|
// silently masked whenever REQUIREMENTS.md has ≥1 item — the mismatch must
|
|
// appear in the report regardless of the requirements row count.
|
|
if (ctxExtraction.outcome === 'could-not-parse') {
|
|
const mismatchMsg = '## Post-Planning Gap Analysis\n\nextracted 0 of N — possible format mismatch in CONTEXT.md decisions block.\n';
|
|
// If there are also requirement items, include them in the return with the
|
|
// mismatch summary appended, so the caller still sees requirement coverage.
|
|
// #2334 HIGH 1: gate on `ghostReqIds.length > 0` too — identical defect to
|
|
// the one fixed at #2316-6b (~34 lines below, at the `items.length === 0`
|
|
// early return): a phase whose EVERY cited REQ-ID is unregistered has
|
|
// `items.length === 0` (all its requirement items were filtered out at
|
|
// ~line 297) but still has real ghost rows to report. Without this guard,
|
|
// a single malformed `<decisions>` line in CONTEXT.md made an all-ghost
|
|
// phase's ghost rows silently vanish (this could-not-parse branch fell
|
|
// through to the bare `mismatchMsg`-only return below, dropping ghost
|
|
// rows that the general path further down correctly surfaces).
|
|
if (items.length > 0 || ghostReqIds.length > 0) {
|
|
const rows = sortRows([
|
|
...detectCoverage(items, planText),
|
|
...ghostReqIds.map(id => ({ source: 'REQUIREMENTS.md', item: id, status: 'Missing from REQUIREMENTS.md' })),
|
|
]);
|
|
const covered = rows.filter(r => r.status === 'Covered').length;
|
|
const uncovered = rows.length - covered;
|
|
const coverageSummary = uncovered === 0
|
|
? `✓ All ${rows.length} items covered by plans`
|
|
: `⚠ ${uncovered} of ${rows.length} items not covered by any plan`;
|
|
return {
|
|
enabled: true,
|
|
rows,
|
|
table: formatGapTable(rows) + '\n' + coverageSummary + '\n\n' + mismatchMsg,
|
|
summary: coverageSummary + '; extracted 0 of N — possible format mismatch',
|
|
counts: { total: rows.length, covered, uncovered },
|
|
phase_dir_read_error: phaseDirReadError,
|
|
};
|
|
}
|
|
return {
|
|
enabled: true,
|
|
rows: [],
|
|
table: mismatchMsg,
|
|
summary: 'extracted 0 of N — possible format mismatch',
|
|
counts: { total: 0, covered: 0, uncovered: 0 },
|
|
phase_dir_read_error: phaseDirReadError,
|
|
};
|
|
}
|
|
|
|
// #1365: if no items at all, surface a clean no-check message.
|
|
// #2316-6b: this must NOT fire when `ghostReqIds` is non-empty — a phase
|
|
// whose EVERY cited REQ-ID is unregistered has `items.length === 0` (all
|
|
// its requirement items were filtered out at ~line 297) but still has real
|
|
// ghost rows to report below. Without this guard, an all-orphan phase
|
|
// reported LESS than a partially-orphan one (which falls through to the
|
|
// general path further down and correctly surfaces its ghost rows).
|
|
if (items.length === 0 && ghostReqIds.length === 0) {
|
|
return {
|
|
enabled: true,
|
|
rows: [],
|
|
table: '## Post-Planning Gap Analysis\n\nNo requirements or decisions to check.\n',
|
|
summary: 'no requirements or decisions to check',
|
|
counts: { total: 0, covered: 0, uncovered: 0 },
|
|
phase_dir_read_error: phaseDirReadError,
|
|
};
|
|
}
|
|
|
|
const rows = sortRows([
|
|
...detectCoverage(items, planText),
|
|
...ghostReqIds.map(id => ({ source: 'REQUIREMENTS.md', item: id, status: 'Missing from REQUIREMENTS.md' })),
|
|
]);
|
|
const covered = rows.filter(r => r.status === 'Covered').length;
|
|
const uncovered = rows.length - covered;
|
|
|
|
const summary = uncovered === 0
|
|
? `✓ All ${rows.length} items covered by plans`
|
|
: `⚠ ${uncovered} of ${rows.length} items not covered by any plan`;
|
|
|
|
return {
|
|
enabled: true,
|
|
rows,
|
|
table: formatGapTable(rows) + '\n' + summary + '\n',
|
|
summary,
|
|
counts: { total: rows.length, covered, uncovered },
|
|
phase_dir_read_error: phaseDirReadError,
|
|
};
|
|
}
|
|
|
|
function cmdGapAnalysis(cwd: string, args: string[], raw: boolean): void {
|
|
const idx = args.indexOf('--phase-dir');
|
|
if (idx === -1 || !args[idx + 1]) {
|
|
error('Usage: gap-analysis --phase-dir <path-to-phase-directory>');
|
|
}
|
|
const phaseDir = args[idx + 1];
|
|
|
|
// Optional --phase-req-ids scopes the requirements comparison (#447).
|
|
// Absent → compare the whole REQUIREMENTS.md (back-compat).
|
|
const reqIdx = args.indexOf('--phase-req-ids');
|
|
const phaseReqIds = reqIdx === -1 ? undefined : (args[reqIdx + 1] ?? '');
|
|
|
|
const result = runGapAnalysis(cwd, phaseDir, { phaseReqIds });
|
|
output(result, raw, result.table || result.summary);
|
|
}
|
|
|
|
export = {
|
|
parseRequirements,
|
|
detectCoverage,
|
|
formatGapTable,
|
|
sortRows,
|
|
normalizePhaseReqIds,
|
|
runGapAnalysis,
|
|
cmdGapAnalysis,
|
|
};
|