Files
msd-core/src/uat-predicate.cts
Jeremy McSpadden 77c7b4fc9d fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition

* no-mistakes(review): Fix canonical verification closeout gates

* no-mistakes(review): Fix verify-work frontmatter promotion command

* no-mistakes(review): Fix stale verification gates

* no-mistakes(review): Fix canonical verification routing gates

* no-mistakes(review): Fix verification dependency and runtime routing gates

* no-mistakes(review): Block stale verification bypasses

* fix: handle large init manager outputs in verification workflows

* chore: update changeset pr number

* fix(verify-work): use fresh verification.status for stale gate

The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.

* fix(init): skip roadmap-checked phases when selecting next_phase

Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.

* fix: gaps_found not overridden by stale, transition uses canonical verification

- verification.cts: check gaps_found before stale so gap-closure routing
  is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
  already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
  to avoid false-positive blocks from body text matching

* ci: retrigger tests after rebase

* fix(transition): replace gsd_run advisory check with awk frontmatter extraction

The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.

Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.

The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.

Also update workflow-size-baseline.json for the updated transition.md size.

Fixes: runtime-launcher-parity test (B)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: re-check verification under planning lock in phase complete

Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.

* fix(transition): gate on canonical verification.status including stale

Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).

* Fix workflow verification gates for yolo transition and stale routing

Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.

* fix(transition): use verification.status query for stale-aware advisory check

The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.

Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.

Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.

Also update workflow-size-baseline.json for the updated transition.md size.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* ci: trigger test matrix for 525b946

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(transition): restore awk frontmatter extraction for pre-shim verification check

The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(#1522): clarify transition verification gate wording

* fix(#1522): update transition workflow size baseline

* fix(#1522): update workflow-size-baseline after rebase onto next

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)

Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 13:19:10 -04:00

349 lines
14 KiB
TypeScript

/**
* UAT Predicate — Pure-computation UAT pass/fail evaluation
*
* Evaluates all *-UAT.md and *-VERIFICATION.md files in a phase directory and
* returns a typed report. Used by `phase uat-passed` to harden against the
* naive whole-file regex in cmdPhaseComplete which false-matches `result:` lines
* inside frontmatter, fenced code blocks, blockquotes, and HTML comments.
*
* Issue #247 — phase uat-passed predicate
*
* ADR-457 build-at-publish: compiled by tsc to gsd-core/bin/lib/uat-predicate.cjs.
*/
import fs from 'node:fs';
import path from 'node:path';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import frontmatter = require('./frontmatter.cjs');
const { extractFrontmatter } = frontmatter;
// eslint-disable-next-line @typescript-eslint/no-require-imports
import markdownSectionizer = require('./markdown-sectionizer.cjs');
const { stripFencedCode } = markdownSectionizer;
// eslint-disable-next-line @typescript-eslint/no-require-imports
import verification = require('./verification.cjs');
const { readVerificationStatus } = verification;
// ─── Types ────────────────────────────────────────────────────────────────────
interface UatCheckItem {
file: string;
test: number;
name: string;
result: string;
passing: boolean;
}
interface UatPassedReport {
passed: boolean;
uat_files: string[];
verification_files: string[];
checks: UatCheckItem[];
blockers: string[];
no_uat_artifacts: boolean;
policy: {
require_verification: boolean;
};
}
// ─── Blocking state sets (documented for maintainability) ─────────────────────
// UAT file frontmatter `status` values that indicate the file is not fully done
const BLOCKING_UAT_FM_STATUSES = new Set([
'partial', 'diagnosed', 'pending', 'blocked', 'in_progress', 'failed',
]);
// UAT file frontmatter `result` values that indicate failure
const BLOCKING_UAT_FM_RESULTS = new Set(['pending', 'blocked', 'failed']);
// Canonical VERIFICATION frontmatter `status` value that indicates passing.
const PASSING_VERIFICATION_STATUSES = new Set(['passed']);
// VERIFICATION file frontmatter `status` values that explicitly block
const BLOCKING_VERIFICATION_FM_STATUSES = new Set([
'human_needed', 'gaps_found', 'pending', 'blocked', 'partial',
'failed', 'in_progress',
]);
// UAT test-item `result` values that count as passing
const PASSING_RESULTS = new Set(['passed', 'pass']);
// ─── stripFalsePositiveContexts ───────────────────────────────────────────────
/**
* Remove contexts that can contain `result: ...` lines that are NOT real test results:
* (a) leading frontmatter block at byte 0
* (b) HTML comments (unterminated comments swallow to EOF — fail-closed)
* (c) fenced code blocks (backtick and tilde, indented too) via CommonMark state machine
* (d) blockquote lines
*
* Each step is a composable function (Kernighan's Law — independently testable).
* Returns surviving lines joined by '\n'. Robust to CRLF input.
*/
function stripFalsePositiveContexts(content: string): string {
// Step (a): strip leading frontmatter block only at byte 0
let stripped = content.replace(/^---\r?\n[\s\S]*?\r?\n---[ \t]*(\r?\n|$)/, '');
// Step (b): remove HTML comments anywhere; unterminated comment swallows to EOF
stripped = stripped.replace(/<!--[\s\S]*?(?:-->|$)/g, '');
// Step (c): remove fenced code blocks via the canonical seam (ADR-1372 T5)
stripped = stripFencedCode(stripped).text;
// Step (d): remove blockquote lines
stripped = stripped
.split('\n')
.filter(line => !/^\s*>/.test(line))
.join('\n');
return stripped;
}
/**
* Analyse raw markdown for structural anomalies (unterminated fence / comment).
* Exported for unit-testability and used by evaluateUatPassed for per-file malformed detection.
*
* FIX C: properly balanced comments are stripped before checking for a dangling <!--,
* so an earlier closed comment does not mask a later unterminated one.
*/
function analyzeMarkdown(raw: string): { unterminatedFence: boolean; unterminatedComment: boolean } {
// Detect an unterminated HTML comment via a paired scan: every `<!--` must
// have a following `-->`. Using indexOf (not a regex .replace of the comment
// token) avoids the js/incomplete-multi-character-sanitization pattern — and
// is exact: a closed earlier comment never masks a later unterminated one.
let unterminatedComment = false;
for (let i = 0; ; ) {
const open = raw.indexOf('<!--', i);
if (open === -1) break;
const close = raw.indexOf('-->', open + 4);
if (close === -1) { unterminatedComment = true; break; }
i = close + 3;
}
// Fence state machine gives the accurate unterminated-fence signal (seam, ADR-1372 T5).
const { unterminatedFence } = stripFencedCode(raw);
return { unterminatedFence, unterminatedComment };
}
// ─── parseUatResultItems ──────────────────────────────────────────────────────
/**
* HEADING-BLOCK parser: scan the CLEANED body (after stripFalsePositiveContexts)
* for UAT test blocks.
*
* For each ### N. Name heading, the block spans until the next ### heading or EOF.
* Within each block, find a column-0 anchored result line (rejects indented YAML
* block-scalar bodies and inline/quoted fakes).
*
* - If a heading block has NO column-0 result line → emit result:'missing' (blocker).
* - Support bracketed [passed] and bare passed (#2273).
* - Returns ALL items (both passing and non-passing).
*/
function parseUatResultItems(cleanContent: string): Array<{ test: number; name: string; result: string }> {
const items: Array<{ test: number; name: string; result: string }> = [];
// Find all ### N. Name headings (line-anchored)
const headingPattern = /^###\s*(\d+)\.\s*(.+)$/gm;
const headings: Array<{ index: number; test: number; name: string }> = [];
let hMatch: RegExpExecArray | null;
while ((hMatch = headingPattern.exec(cleanContent)) !== null) {
headings.push({
index: hMatch.index + hMatch[0].length,
test: parseInt(hMatch[1], 10),
name: hMatch[2].trim(),
});
}
for (let i = 0; i < headings.length; i++) {
const h = headings[i];
const blockStart = h.index;
// More precise: find next heading's position in the original string
// We'll slice from current heading end to the position just before next heading's "###"
const nextHeadingMatch = i + 1 < headings.length
? cleanContent.lastIndexOf('\n###', headings[i + 1].index)
: -1;
const blockContent = nextHeadingMatch >= blockStart
? cleanContent.slice(blockStart, nextHeadingMatch)
: cleanContent.slice(blockStart);
// Column-0 anchored result line: /^result:[ \t]*\[?([\w-]+)\]?/mi
// Uses [ \t]* (not \s*) so the captured value must sit on the SAME line as result:.
// A result: key with the value on a subsequent line yields no match → 'missing' (blocker).
const resultMatch = /^result:[ \t]*\[?([\w-]+)\]?/mi.exec(blockContent);
if (resultMatch) {
items.push({
test: h.test,
name: h.name,
result: resultMatch[1].toLowerCase(),
});
} else {
// No column-0 result line → emit 'missing' (a non-passing state)
items.push({
test: h.test,
name: h.name,
result: 'missing',
});
}
}
return items;
}
// ─── evaluateUatPassed ────────────────────────────────────────────────────────
/**
* Evaluate all UAT/VERIFICATION files in a phase directory.
* Returns a UatPassedReport with the locked, stable shape defined by the interface.
*
* FAIL-CLOSED: any absence/ambiguity/malformed input → NOT passed.
* Pass requires at least one real passing check AND no blockers.
*/
function evaluateUatPassed(
phaseFullDir: string,
opts?: { policy?: { requireVerification?: boolean } },
): UatPassedReport {
const requireVerification = opts?.policy?.requireVerification === true;
const blockers: string[] = [];
const checks: UatCheckItem[] = [];
const uatFiles: string[] = [];
const verificationFiles: string[] = [];
// Read the directory — if unreadable, treat as no files (fail-closed: no artifacts → not passed)
let dirEntries: string[] = [];
try {
dirEntries = fs.readdirSync(phaseFullDir);
} catch {
// Unreadable dir — no_uat_artifacts:true, passed:false
const no_uat_artifacts = true;
if (requireVerification) {
blockers.push('policy: verification required but no passing *-VERIFICATION.md found');
}
return {
passed: false,
uat_files: [],
verification_files: [],
checks: [],
blockers,
no_uat_artifacts,
policy: { require_verification: requireVerification },
};
}
// Filter UAT and VERIFICATION files using the same filter as cmdPhaseComplete
const uatFileNames = dirEntries.filter(f => f.includes('-UAT') && f.endsWith('.md'));
const verFileNames = dirEntries.filter(f => f.includes('-VERIFICATION') && f.endsWith('.md'));
// ── Process UAT files ──────────────────────────────────────────────────────
for (const file of uatFileNames) {
uatFiles.push(file);
let raw = '';
try {
raw = fs.readFileSync(path.join(phaseFullDir, file), 'utf-8');
} catch {
blockers.push(`${file}: could not read file`);
continue;
}
// ── Per-file malformed markdown guard ──────────────────────────────────
// FIX D: use accurate signals from analyzeMarkdown instead of heuristics.
// unterminatedFence: CommonMark state machine detects a genuinely unclosed fence.
// unterminatedComment: strips balanced comments first, then checks for leftover <!--.
const { unterminatedFence, unterminatedComment } = analyzeMarkdown(raw);
if (unterminatedFence || unterminatedComment) {
blockers.push(`${file}: malformed markdown (unterminated fence or comment)`);
}
const fm = extractFrontmatter(raw) as Record<string, unknown>;
// File-level frontmatter status check
if (fm['status'] && BLOCKING_UAT_FM_STATUSES.has(fm['status'] as string)) {
blockers.push(`${file}: frontmatter status=${fm['status'] as string}`);
}
// File-level frontmatter result check
if (fm['result'] && BLOCKING_UAT_FM_RESULTS.has(fm['result'] as string)) {
blockers.push(`${file}: frontmatter result=${fm['result'] as string}`);
}
// Parse test items from the cleaned body (hardened against false positives)
const cleanContent = stripFalsePositiveContexts(raw);
const items = parseUatResultItems(cleanContent);
for (const item of items) {
const passing = PASSING_RESULTS.has(item.result);
checks.push({
file,
test: item.test,
name: item.name,
result: item.result,
passing,
});
if (!passing) {
blockers.push(`${file}: test ${item.test} (${item.result})`);
}
}
}
// ── Process VERIFICATION files ─────────────────────────────────────────────
let hasPassingVerification = false;
for (const file of verFileNames) {
verificationFiles.push(file);
let raw = '';
try {
raw = fs.readFileSync(path.join(phaseFullDir, file), 'utf-8');
} catch {
blockers.push(`${file}: could not read verification file`);
continue;
}
const vfm = extractFrontmatter(raw) as Record<string, unknown>;
const vStatus = vfm['status'] as string | undefined;
if (vStatus && BLOCKING_VERIFICATION_FM_STATUSES.has(vStatus)) {
blockers.push(`${file}: verification status=${vStatus}`);
} else if (vStatus && PASSING_VERIFICATION_STATUSES.has(vStatus)) {
// Allowlist: only explicitly-passing statuses count
hasPassingVerification = true;
}
// Missing or unknown status: does NOT count as passing, does NOT push a blocker
// (handled by the requireVerification policy check below if needed)
}
// ── Policy: requireVerification ───────────────────────────────────────────
if (requireVerification) {
const verificationStatus = readVerificationStatus(phaseFullDir).status;
if (verificationStatus === 'stale') {
blockers.push('policy: verification status=stale');
} else if (verificationStatus !== 'passed' || !hasPassingVerification) {
blockers.push('policy: verification required but no passing *-VERIFICATION.md found');
}
}
// ── Determine no_uat_artifacts and passed ─────────────────────────────────
// no_uat_artifacts: true when no real UAT test items were parsed from any file
const no_uat_artifacts = checks.length === 0;
// FIX 1: require positive passing evidence; no vacuous pass
// passed = no blockers AND at least one check AND all checks passing
const passed = blockers.length === 0 && checks.length > 0 && checks.every(c => c.passing);
return {
passed,
uat_files: uatFiles,
verification_files: verificationFiles,
checks,
blockers,
no_uat_artifacts,
policy: {
require_verification: requireVerification,
},
};
}
export = {
stripFalsePositiveContexts,
parseUatResultItems,
analyzeMarkdown,
evaluateUatPassed,
};