diff --git a/.changeset/happy-otters-leap.md b/.changeset/happy-otters-leap.md new file mode 100644 index 000000000..80f3dcbcb --- /dev/null +++ b/.changeset/happy-otters-leap.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 1063 +--- +**`phase uat-passed` predicate** — new runtime-neutral command evaluates HUMAN-UAT results with markdown-aware parsing (ignores frontmatter, fenced code, blockquotes, and HTML comments) and reports pass only when every required check passes. (#1063) diff --git a/.gitignore b/.gitignore index 82faf6c3c..08e407f21 100644 --- a/.gitignore +++ b/.gitignore @@ -173,6 +173,7 @@ build/ /gsd-core/bin/lib/verify.cjs /gsd-core/bin/lib/init.cjs /gsd-core/bin/lib/uat.cjs +/gsd-core/bin/lib/uat-predicate.cjs /gsd-core/bin/lib/workstream.cjs /gsd-core/bin/lib/roadmap.cjs /gsd-core/bin/lib/audit.cjs diff --git a/CONTEXT.md b/CONTEXT.md index 8cc1d93fb..4dad44981 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -198,6 +198,9 @@ The GSD-RESEARCH capability behind an L2-hybrid seam: code owns cache + provider - `GSD-RESEARCH.CONTEXT-DISCIPLINE=less-context levers: subagent isolation + compact provider output + fetches-to-disk + cache-returns-digest; API clear_tool_uses/memory tool are the conceptual model, not a Claude Code harness knob` - `DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT=provider waterfall duplicated across N researcher agent .md files drifts independently (META.RULE.brief-no-paraphrase); fix-forward=research-provider.cjs single source of truth + generated agents (#657)` +### UAT-Passed Predicate +Runtime-neutral predicate evaluating `*-UAT.md` / `*-VERIFICATION.md` result fields with markdown-aware parsing that ignores false-positive contexts (frontmatter body, fenced code, HTML comments, blockquotes). Returns `passed: true` only when all required checks pass; supports `--require-verification` to demand at least one VERIFICATION.md file alongside UAT results. Output envelope: `{ passed, uat_files[], verification_files[], checks[], blockers[], policy }`. Source: `gsd-core/bin/lib/uat-predicate.cjs` (generated from `src/uat-predicate.cts`). Wired via `phase uat-passed` alias → `phase-command-router` → `cmdPhaseUatPassed`. + ### MVP Mode Phase-level planning mode that frames work as a vertical slice (UI → API → DB) of one user-visible capability instead of horizontal layers. Resolved at workflow init via the precedence chain: `--mvp` CLI flag → ROADMAP.md `**Mode:** mvp` field → `workflow.mvp_mode` config → false. All-or-nothing per phase (PRD #2826 Q1). Surfaced as `MVP_MODE=true|false` to the planner, executor, verifier, and discovery surfaces (progress, stats, graphify). Canonical parser: `roadmap.cjs` `**Mode:**` field; canonical resolution chain documented in `workflows/plan-phase.md`. Concept index: `references/mvp-concepts.md`. diff --git a/docs/CLI-TOOLS.md b/docs/CLI-TOOLS.md index 14aaba631..84f900b1a 100644 --- a/docs/CLI-TOOLS.md +++ b/docs/CLI-TOOLS.md @@ -118,6 +118,10 @@ node gsd-tools.cjs phase remove [--force] # Mark phase complete, update state + roadmap node gsd-tools.cjs phase complete +# Evaluate HUMAN-UAT results for a phase (markdown-aware; ignores false-positive contexts) +# Returns JSON: { passed, uat_files[], verification_files[], checks[], blockers[], policy } +node gsd-tools.cjs phase uat-passed [--require-verification] + # Index plans with waves and status node gsd-tools.cjs phase-plan-index diff --git a/docs/COMMANDS.md b/docs/COMMANDS.md index ea89b087d..514949128 100644 --- a/docs/COMMANDS.md +++ b/docs/COMMANDS.md @@ -490,6 +490,37 @@ Retroactively audit and fill Nyquist validation gaps. --- +### `phase uat-passed [--require-verification]` + +Runtime-neutral predicate that evaluates HUMAN-UAT results for a phase and reports whether all required checks passed. Uses markdown-aware parsing that ignores false-positive contexts (YAML frontmatter, fenced code blocks, HTML comments, and blockquotes), so incomplete checkbox fragments in prose sections never trigger a false pass. + +| Argument | Required | Description | +|----------|----------|-------------| +| `N` | **Yes** | Phase number to evaluate | +| `--require-verification` | No | Require at least one `*-VERIFICATION.md` file alongside UAT results; fails if none are found | + +**Output fields (JSON):** + +| Field | Type | Description | +|-------|------|-------------| +| `passed` | `boolean` | `true` only when at least one check exists AND all checks pass AND no blockers — fail-closed (no vacuous pass) | +| `uat_files` | `string[]` | Filenames of `*-UAT.md` files evaluated | +| `verification_files` | `string[]` | Filenames of `*-VERIFICATION.md` files evaluated | +| `checks[]` | `{ file, test, name, result, passing }[]` | Per-item evaluation results parsed from heading blocks | +| `blockers[]` | `string[]` | Human-readable reasons for failure (frontmatter issues, failing/missing test items, policy violations, malformed markdown) — NOT a subset of `checks[]` | +| `no_uat_artifacts` | `boolean` | `true` when no real UAT test items were parsed (no `*-UAT.md` files, unreadable dir, or files with no test blocks); when `true`, `passed` is always `false` | +| `policy.require_verification` | `boolean` | Whether `--require-verification` was active | + +**Programmatic access:** `node gsd-tools.cjs phase uat-passed [--require-verification] [--raw]` — see [CLI Tools Reference](CLI-TOOLS.md) + +```bash +node gsd-tools.cjs phase uat-passed 3 # Evaluate UAT for phase 3 +node gsd-tools.cjs phase uat-passed 3 --require-verification # Also require VERIFICATION.md +node gsd-tools.cjs phase uat-passed 3 --raw # Machine-readable JSON output +``` + +--- + ## Navigation Commands ### `/gsd-progress` diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 952cc4503..f86733150 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -164,6 +164,7 @@ - [Statusline Context Position](#140-statusline-context-position) - [Milestone Tag Creation Toggle](#141-milestone-tag-creation-toggle) - [Structured JSON Error Mode](#142-structured-json-error-mode) + - [UAT-Passed Predicate](#143-uat-passed-predicate) --- @@ -3046,6 +3047,38 @@ explicit reviewer flags -> --all -> review.default_reviewers -> all detected rev --- +### 143. UAT-Passed Predicate + +**CLI:** `node gsd-tools.cjs phase uat-passed [--require-verification]` + +**Purpose:** Provide a runtime-neutral, automatable predicate that evaluates HUMAN-UAT results for a phase and returns a structured pass/fail verdict with full diagnostic detail. + +**Behavior:** Locates `*-UAT.md` and optionally `*-VERIFICATION.md` files for the given phase, parses UAT test blocks (heading-block parser, column-0 result lines) with a markdown-aware stripper that removes false-positive contexts (YAML frontmatter, fenced code blocks, HTML comments, and blockquotes). Returns `passed: true` only when at least one check exists AND all checks pass AND no blockers — fail-closed, no vacuous pass. The `--require-verification` flag requires at least one `*-VERIFICATION.md` with an allowlisted passing status; the command fails without one. + +**Output envelope:** `{ passed, uat_files[], verification_files[], checks[], blockers[], no_uat_artifacts, policy: { require_verification } }` + +| Field | Type | Description | +|-------|------|-------------| +| `passed` | `boolean` | `true` only when ≥1 check exists AND all passing AND no blockers | +| `uat_files` | `string[]` | Filenames of `*-UAT.md` files evaluated | +| `verification_files` | `string[]` | Filenames of `*-VERIFICATION.md` files evaluated | +| `checks[]` | `{ file, test, name, result, passing }[]` | Per-item results from heading blocks | +| `blockers[]` | `string[]` | Human-readable failure reasons (frontmatter, failing/missing items, policy, malformed markdown) | +| `no_uat_artifacts` | `boolean` | `true` when no test items were parsed; `passed` is always `false` when `true` | +| `policy.require_verification` | `boolean` | Whether `--require-verification` was active | + +**Requirements:** +- REQ-UAT-PRED-01: The predicate MUST ignore result lines inside YAML frontmatter, fenced code blocks, HTML comments, and blockquotes. +- REQ-UAT-PRED-02: `passed: true` MUST require at least one check AND all checks passing AND no blockers (fail-closed, no vacuous pass). +- REQ-UAT-PRED-03: `--require-verification` MUST cause the command to fail when no `*-VERIFICATION.md` file with an allowlisted passing status is found. +- REQ-UAT-PRED-04: `blockers[]` contains all human-readable failure reasons including frontmatter issues, policy violations, and malformed markdown — NOT limited to a subset of `checks[]`. +- REQ-UAT-PRED-05: The module MUST be runtime-neutral (no runtime-specific env checks or exit shortcuts). +- REQ-UAT-PRED-06: A heading block with no column-0 `result:` line emits `result:'missing'` (blocker); test items are never silently dropped. + +**Reference:** [Phase Management Commands](COMMANDS.md#phase-uat-passed-n---require-verification) + +--- + ## Related - [Commands](COMMANDS.md) diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index 9cf874497..9af98f83a 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -356,6 +356,7 @@ "surface.cjs", "task-command-router.cjs", "template.cjs", + "uat-predicate.cjs", "uat.cjs", "ui-safety-gate.cjs", "update-context.cjs", diff --git a/docs/INVENTORY.md b/docs/INVENTORY.md index 85ec0afb1..f4374d07d 100644 --- a/docs/INVENTORY.md +++ b/docs/INVENTORY.md @@ -371,7 +371,7 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t --- -## CLI Modules (105 shipped) +## CLI Modules (106 shipped) Full listing: `gsd-core/bin/lib/*.cjs`. @@ -468,6 +468,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`. | `task-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools task` | | `template.cjs` | Template selection and filling with variable substitution | | `uat.cjs` | UAT file parsing, verification debt tracking, audit-uat support | +| `uat-predicate.cjs` | UAT-passed predicate — markdown-aware evaluation of HUMAN-UAT results; returns pass only when all required checks pass; ignores false-positive contexts (frontmatter, fenced code, blockquotes, HTML comments) | | `ui-safety-gate.cjs` | Shell-free word-boundary UI token detector (#3706, #3718); reads phase-section text from stdin, exits 0 (UI found) or 1 (no UI); also deployed to `gsd-core/bin/lib/` so the GSD installer ships it to `$RUNTIME_DIR` (#448) | | `update-context.cjs` | Pure install-context resolver for `/gsd:update` — runtime/scope/config-dir/version detection (LOCAL/GLOBAL/UNKNOWN) ported from update.md bash; backs `gsd-tools update-context` (#498) | | `validate-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools validate` | diff --git a/eslint.config.mjs b/eslint.config.mjs index 6bd3e929e..2e09d49bf 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -137,6 +137,7 @@ export default tseslint.config( 'gsd-core/bin/lib/profile-pipeline.cjs', 'gsd-core/bin/lib/template.cjs', 'gsd-core/bin/lib/uat.cjs', + 'gsd-core/bin/lib/uat-predicate.cjs', 'gsd-core/bin/lib/workstream.cjs', 'gsd-core/bin/lib/roadmap.cjs', 'gsd-core/bin/lib/audit.cjs', diff --git a/src/phase-command-router.cts b/src/phase-command-router.cts index ab297e917..f9eb45eb0 100644 --- a/src/phase-command-router.cts +++ b/src/phase-command-router.cts @@ -32,6 +32,7 @@ interface PhaseHandlers { cmdPhaseInsert: (cwd: string, pos: string | undefined, desc: string, raw: boolean) => void; cmdPhaseRemove: (cwd: string, phaseNum: string, opts: { force: boolean }, raw: boolean) => void; cmdPhaseComplete: (cwd: string, phaseNum: string | undefined, raw: boolean) => void; + cmdPhaseUatPassed: (cwd: string, phaseNum: string | undefined, raw: boolean, opts?: { policy?: { requireVerification?: boolean } }) => void; } interface RoutePhaseCommandOptions { @@ -164,6 +165,23 @@ function routePhaseCommand({ phase, args, cwd, raw, error }: RoutePhaseCommandOp phase.cmdPhaseComplete(cwd, args[2], raw); return { ok: true as const, data: null }; }, + 'uat-passed': (_ctx: Record): { ok: true; data: null } => { + let requireVerification = false; + const positional: string[] = []; + for (const token of args.slice(2)) { + if (token === '--require-verification') { + requireVerification = true; + } else if (token === '--raw') { + // --raw is handled by the outer CLI layer; accepted here silently + } else if (token.startsWith('--')) { + return makeInvalidArgs(token, `phase uat-passed does not support ${token}`) as never; + } else { + positional.push(token); + } + } + phase.cmdPhaseUatPassed(cwd, positional[0], raw, { policy: { requireVerification } }); + return { ok: true as const, data: null }; + }, }, }; diff --git a/src/phase.cts b/src/phase.cts index 408f08ae4..d8ea084a2 100644 --- a/src/phase.cts +++ b/src/phase.cts @@ -29,6 +29,9 @@ import stateMod = require('./state.cjs'); import { platformWriteSync, platformReadSync, platformEnsureDir } from './shell-command-projection.cjs'; import { formatGsdSlash, resolveRuntime } from './runtime-slash.cjs'; import { deriveProgressFromRoadmap, clampPercent } from './phase-lifecycle.cjs'; +// eslint-disable-next-line @typescript-eslint/no-require-imports -- uat-predicate.cjs is an export= CommonJS module +import uatPredicate = require('./uat-predicate.cjs'); +const { evaluateUatPassed } = uatPredicate; const { escapeRegex, @@ -1712,6 +1715,28 @@ function cmdPhaseComplete(cwd: string, phaseNum: string, raw: boolean): void { output(result, raw); } +function cmdPhaseUatPassed( + cwd: string, + phaseNum: string | undefined, + raw: boolean, + opts: { policy?: { requireVerification?: boolean } } = {}, +): void { + if (!phaseNum) { + error('phase number required for phase uat-passed'); + } + + const phaseInfoRaw = findPhaseInternal(cwd, phaseNum!); + if (!phaseInfoRaw) { + error(`Phase ${phaseNum} not found`); + } + const phaseInfo = phaseInfoRaw as unknown as Record; + const phaseFullDir = path.join(cwd, phaseInfo['directory'] as string); + + const report = evaluateUatPassed(phaseFullDir, { policy: opts.policy }); + + output({ phase: phaseNum, ...report }, raw); +} + export = { cmdPhasesList, cmdPhaseNextDecimal, @@ -1723,5 +1748,6 @@ export = { cmdPhaseInsert, cmdPhaseRemove, cmdPhaseComplete, + cmdPhaseUatPassed, computeDependencyLevels, }; diff --git a/src/uat-predicate.cts b/src/uat-predicate.cts new file mode 100644 index 000000000..b0741e8d8 --- /dev/null +++ b/src/uat-predicate.cts @@ -0,0 +1,394 @@ +/** + * UAT Predicate — Pure-computation UAT pass/fail evaluation + * + * Evaluates all *-UAT.md and *-VERIFICATION.md files in a phase directory and + * returns a typed report. Used by `phase uat-passed` to harden against the + * naive whole-file regex in cmdPhaseComplete which false-matches `result:` lines + * inside frontmatter, fenced code blocks, blockquotes, and HTML comments. + * + * Issue #247 — phase uat-passed predicate + * + * ADR-457 build-at-publish: compiled by tsc to gsd-core/bin/lib/uat-predicate.cjs. + */ + +import fs from 'node:fs'; +import path from 'node:path'; +// eslint-disable-next-line @typescript-eslint/no-require-imports +import frontmatter = require('./frontmatter.cjs'); +const { extractFrontmatter } = frontmatter; + +// ─── Types ──────────────────────────────────────────────────────────────────── + +interface UatCheckItem { + file: string; + test: number; + name: string; + result: string; + passing: boolean; +} + +interface UatPassedReport { + passed: boolean; + uat_files: string[]; + verification_files: string[]; + checks: UatCheckItem[]; + blockers: string[]; + no_uat_artifacts: boolean; + policy: { + require_verification: boolean; + }; +} + +// ─── Blocking state sets (documented for maintainability) ───────────────────── + +// UAT file frontmatter `status` values that indicate the file is not fully done +const BLOCKING_UAT_FM_STATUSES = new Set([ + 'partial', 'diagnosed', 'pending', 'blocked', 'in_progress', 'failed', +]); + +// UAT file frontmatter `result` values that indicate failure +const BLOCKING_UAT_FM_RESULTS = new Set(['pending', 'blocked', 'failed']); + +// VERIFICATION file frontmatter `status` values that indicate passing +const PASSING_VERIFICATION_STATUSES = new Set([ + 'complete', 'verified', 'passed', 'human_passed', +]); + +// VERIFICATION file frontmatter `status` values that explicitly block +const BLOCKING_VERIFICATION_FM_STATUSES = new Set([ + 'human_needed', 'gaps_found', 'pending', 'blocked', 'partial', + 'failed', 'in_progress', +]); + +// UAT test-item `result` values that count as passing +const PASSING_RESULTS = new Set(['passed', 'pass']); + +// ─── stripFalsePositiveContexts ─────────────────────────────────────────────── + +/** + * Remove contexts that can contain `result: ...` lines that are NOT real test results: + * (a) leading frontmatter block at byte 0 + * (b) HTML comments (unterminated comments swallow to EOF — fail-closed) + * (c) fenced code blocks (backtick and tilde, indented too) via CommonMark state machine + * (d) blockquote lines + * + * Each step is a composable function (Kernighan's Law — independently testable). + * Returns surviving lines joined by '\n'. Robust to CRLF input. + */ +function stripFalsePositiveContexts(content: string): string { + // Step (a): strip leading frontmatter block only at byte 0 + let stripped = content.replace(/^---\r?\n[\s\S]*?\r?\n---[ \t]*(\r?\n|$)/, ''); + + // Step (b): remove HTML comments anywhere; unterminated comment swallows to EOF + stripped = stripped.replace(/|$)/g, ''); + + // Step (c): remove fenced code blocks via CommonMark-style state machine (handles CRLF + indented fences) + stripped = _stripFencedBlocks(stripped).text; + + // Step (d): remove blockquote lines + stripped = stripped + .split('\n') + .filter(line => !/^\s*>/.test(line)) + .join('\n'); + + return stripped; +} + +interface FenceState { + char: '`' | '~'; + len: number; +} + +interface StripFencedResult { + text: string; + unterminatedFence: boolean; +} + +/** + * CommonMark-style fenced-code-block stripper. + * Tracks the opening delimiter char and length so that a ~~~ line inside a + * ``` fence is correctly treated as fence content, not a closing delimiter. + * + * Opening rule: first delimiter line with char+len sets openFence. + * Closing rule: delimiter line with SAME char, run length >= openFence.len, + * and NO trailing non-whitespace text closes the fence. + * All delimiter and content lines are dropped; non-fence lines are kept. + * Returns the kept text plus unterminatedFence:true if EOF inside a fence. + */ +function _stripFencedBlocks(content: string): StripFencedResult { + const lines = content.split('\n'); + const kept: string[] = []; + let openFence: FenceState | null = null; + const delimRe = /^(\s*)(`{3,}|~{3,})(.*)$/; + + for (const rawLine of lines) { + // Tolerate CRLF: strip trailing \r for matching, but we work on split-by-\n lines + // (the outer caller joined by \n already; we just handle a stray \r in the last char) + const line = rawLine.replace(/\r$/, ''); + const m = delimRe.exec(line); + if (m) { + const char = m[2][0] as '`' | '~'; + const len = m[2].length; + const trailing = m[3]; + if (openFence === null) { + // Opening delimiter — drop this line and record the fence + openFence = { char, len }; + } else if (char === openFence.char && len >= openFence.len && /^\s*$/.test(trailing)) { + // Closing delimiter (same char, sufficient length, no trailing text) — drop and close + openFence = null; + } + // else: mismatched delimiter inside fence (e.g. ~~~ inside ```) — drop as content + continue; // delimiter lines are always dropped + } + if (openFence === null) { + kept.push(rawLine); + } + // Lines inside fence are dropped + } + + return { text: kept.join('\n'), unterminatedFence: openFence !== null }; +} + +/** + * Analyse raw markdown for structural anomalies (unterminated fence / comment). + * Exported for unit-testability and used by evaluateUatPassed for per-file malformed detection. + * + * FIX C: properly balanced comments are stripped before checking for a dangling `. Using indexOf (not a regex .replace of the comment + // token) avoids the js/incomplete-multi-character-sanitization pattern — and + // is exact: a closed earlier comment never masks a later unterminated one. + let unterminatedComment = false; + for (let i = 0; ; ) { + const open = raw.indexOf('', open + 4); + if (close === -1) { unterminatedComment = true; break; } + i = close + 3; + } + + // Fence state machine gives the accurate unterminated-fence signal. + const { unterminatedFence } = _stripFencedBlocks(raw); + + return { unterminatedFence, unterminatedComment }; +} + +// ─── parseUatResultItems ────────────────────────────────────────────────────── + +/** + * HEADING-BLOCK parser: scan the CLEANED body (after stripFalsePositiveContexts) + * for UAT test blocks. + * + * For each ### N. Name heading, the block spans until the next ### heading or EOF. + * Within each block, find a column-0 anchored result line (rejects indented YAML + * block-scalar bodies and inline/quoted fakes). + * + * - If a heading block has NO column-0 result line → emit result:'missing' (blocker). + * - Support bracketed [passed] and bare passed (#2273). + * - Returns ALL items (both passing and non-passing). + */ +function parseUatResultItems(cleanContent: string): Array<{ test: number; name: string; result: string }> { + const items: Array<{ test: number; name: string; result: string }> = []; + + // Find all ### N. Name headings (line-anchored) + const headingPattern = /^###\s*(\d+)\.\s*(.+)$/gm; + const headings: Array<{ index: number; test: number; name: string }> = []; + let hMatch: RegExpExecArray | null; + while ((hMatch = headingPattern.exec(cleanContent)) !== null) { + headings.push({ + index: hMatch.index + hMatch[0].length, + test: parseInt(hMatch[1], 10), + name: hMatch[2].trim(), + }); + } + + for (let i = 0; i < headings.length; i++) { + const h = headings[i]; + const blockStart = h.index; + // More precise: find next heading's position in the original string + // We'll slice from current heading end to the position just before next heading's "###" + const nextHeadingMatch = i + 1 < headings.length + ? cleanContent.lastIndexOf('\n###', headings[i + 1].index) + : -1; + const blockContent = nextHeadingMatch >= blockStart + ? cleanContent.slice(blockStart, nextHeadingMatch) + : cleanContent.slice(blockStart); + + // Column-0 anchored result line: /^result:[ \t]*\[?([\w-]+)\]?/mi + // Uses [ \t]* (not \s*) so the captured value must sit on the SAME line as result:. + // A result: key with the value on a subsequent line yields no match → 'missing' (blocker). + const resultMatch = /^result:[ \t]*\[?([\w-]+)\]?/mi.exec(blockContent); + if (resultMatch) { + items.push({ + test: h.test, + name: h.name, + result: resultMatch[1].toLowerCase(), + }); + } else { + // No column-0 result line → emit 'missing' (a non-passing state) + items.push({ + test: h.test, + name: h.name, + result: 'missing', + }); + } + } + + return items; +} + +// ─── evaluateUatPassed ──────────────────────────────────────────────────────── + +/** + * Evaluate all UAT/VERIFICATION files in a phase directory. + * Returns a UatPassedReport with the locked, stable shape defined by the interface. + * + * FAIL-CLOSED: any absence/ambiguity/malformed input → NOT passed. + * Pass requires at least one real passing check AND no blockers. + */ +function evaluateUatPassed( + phaseFullDir: string, + opts?: { policy?: { requireVerification?: boolean } }, +): UatPassedReport { + const requireVerification = opts?.policy?.requireVerification === true; + + const blockers: string[] = []; + const checks: UatCheckItem[] = []; + const uatFiles: string[] = []; + const verificationFiles: string[] = []; + + // Read the directory — if unreadable, treat as no files (fail-closed: no artifacts → not passed) + let dirEntries: string[] = []; + try { + dirEntries = fs.readdirSync(phaseFullDir); + } catch { + // Unreadable dir — no_uat_artifacts:true, passed:false + const no_uat_artifacts = true; + if (requireVerification) { + blockers.push('policy: verification required but no passing *-VERIFICATION.md found'); + } + return { + passed: false, + uat_files: [], + verification_files: [], + checks: [], + blockers, + no_uat_artifacts, + policy: { require_verification: requireVerification }, + }; + } + + // Filter UAT and VERIFICATION files using the same filter as cmdPhaseComplete + const uatFileNames = dirEntries.filter(f => f.includes('-UAT') && f.endsWith('.md')); + const verFileNames = dirEntries.filter(f => f.includes('-VERIFICATION') && f.endsWith('.md')); + + // ── Process UAT files ────────────────────────────────────────────────────── + for (const file of uatFileNames) { + uatFiles.push(file); + let raw = ''; + try { + raw = fs.readFileSync(path.join(phaseFullDir, file), 'utf-8'); + } catch { + blockers.push(`${file}: could not read file`); + continue; + } + + // ── Per-file malformed markdown guard ────────────────────────────────── + // FIX D: use accurate signals from analyzeMarkdown instead of heuristics. + // unterminatedFence: CommonMark state machine detects a genuinely unclosed fence. + // unterminatedComment: strips balanced comments first, then checks for leftover \nAfter'; + const out = stripFalsePositiveContexts(input); + assert.ok(!out.includes('result: pending'), 'HTML comment content must be stripped'); + assert.ok(out.includes('Before'), 'content before comment must survive'); + assert.ok(out.includes('After'), 'content after comment must survive'); + }); + + test('removes multi-line HTML comment', () => { + const input = 'A\n\nB'; + const out = stripFalsePositiveContexts(input); + assert.ok(!out.includes('result: pending'), 'multi-line HTML comment content must be stripped'); + assert.ok(out.includes('A'), 'content before comment must survive'); + assert.ok(out.includes('B'), 'content after comment must survive'); + }); + + test('unterminated HTML comment swallows to EOF (fail-closed)', () => { + const input = 'Before\n', + ].join('\n'); + const clean = stripFalsePositiveContexts(rawContent); + const items = parseUatResultItems(clean); + assert.strictEqual(items.length, 0, 'Fake result inside HTML comment must produce no items'); + + writeFile(tmpDir, 'phase-UAT.md', rawContent); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false); + assert.strictEqual(report.no_uat_artifacts, true); + }); + + test('#247: result:passed inside frontmatter: parseUatResultItems returns [] + evaluateUatPassed → passed:false + no_uat_artifacts:true', () => { + const rawContent = [ + '---', + 'example_result: passed', + '---', + '', + 'No real test blocks here.', + ].join('\n'); + const clean = stripFalsePositiveContexts(rawContent); + const items = parseUatResultItems(clean); + assert.strictEqual(items.length, 0, 'Fake result inside frontmatter must produce no items'); + + writeFile(tmpDir, 'phase-UAT.md', rawContent); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false); + assert.strictEqual(report.no_uat_artifacts, true); + }); + + test('#247: result:passed inside fenced block is NOT treated as passing test (with real failing test)', () => { + const content = [ + '---', + 'status: partial', + '---', + '', + '# Example', + '', + '```', + '### 1. Test', + 'expected: Example output', + 'result: passed', + '```', + '', + '### 1. Real Test', + 'expected: Something', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false, + 'result:passed inside a fenced block must not flip passed to true'); + assert.ok(report.checks.some(c => c.result === 'pending' && !c.passing), + 'Real pending test must be captured'); + }); + + test('#247: result:passed inside blockquote is NOT treated as passing test', () => { + const content = [ + '---', + 'status: partial', + '---', + '', + '> ### 1. Test', + '> expected: Example', + '> result: passed', + '', + '### 1. Real Test', + 'expected: Something', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false, + 'result:passed inside a blockquote must not flip passed to true'); + }); + + test('#247: result:passed inside HTML comment is NOT treated as passing test', () => { + const content = [ + '---', + 'status: partial', + '---', + '', + '', + '', + '### 1. Real Test', + 'expected: Something', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false, + 'result:passed inside an HTML comment must not flip passed to true'); + }); + + test('#247: result:passed inside frontmatter example is NOT treated as passing test', () => { + const content = [ + '---', + 'status: partial', + 'example_result: passed', + '---', + '', + '### 1. Real Test', + 'expected: Something', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false, + 'result:passed in frontmatter must not flip passed to true'); + }); + + test('#247: real passing test block outside all contexts → passed:true', () => { + const content = [ + '---', + 'status: passed', + '---', + '', + '# UAT Results', + '', + '### 1. Login works', + 'expected: User logs in', + 'result: passed', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, true, + 'A real passing test outside false-positive contexts must pass'); + }); + + test('block-scalar expected: followed by blank line + result: pending → parsed as blocker (not dropped)', () => { + const content = [ + '### 1. Test A', + 'expected: |', + ' multi', + ' line expected output', + '', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false, + 'Block-scalar expected: with result: pending must be captured as a blocker'); + assert.ok(report.checks.some(c => c.result === 'pending' && !c.passing), + `Expected pending check, got: ${JSON.stringify(report.checks)}`); + }); +}); + +// ─── evaluateUatPassed — output shape contract ─────────────────────────────── + +describe('evaluateUatPassed — output shape (Hyrum\'s Law contract)', () => { + let tmpDir; + + beforeEach(() => { + tmpDir = makeTmpDir(); + }); + + afterEach(() => { + rmDir(tmpDir); + }); + + test('returns all required fields in the locked shape including no_uat_artifacts', () => { + writeFile(tmpDir, 'phase-UAT.md', makePassingUat(1)); + const report = evaluateUatPassed(tmpDir); + + // Locked field names + assert.ok('passed' in report, 'report.passed must exist'); + assert.ok('uat_files' in report, 'report.uat_files must exist'); + assert.ok('verification_files' in report, 'report.verification_files must exist'); + assert.ok('checks' in report, 'report.checks must exist'); + assert.ok('blockers' in report, 'report.blockers must exist'); + assert.ok('no_uat_artifacts' in report, 'report.no_uat_artifacts must exist'); + assert.ok('policy' in report, 'report.policy must exist'); + assert.ok('require_verification' in report.policy, 'report.policy.require_verification must exist'); + + // checks item shape + if (report.checks.length > 0) { + const c = report.checks[0]; + assert.ok('file' in c, 'check.file must exist'); + assert.ok('test' in c, 'check.test must exist'); + assert.ok('name' in c, 'check.name must exist'); + assert.ok('result' in c, 'check.result must exist'); + assert.ok('passing' in c, 'check.passing must exist'); + } + }); + + test('no_uat_artifacts is false when real checks exist', () => { + writeFile(tmpDir, 'phase-UAT.md', makePassingUat(1)); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.no_uat_artifacts, false); + }); + + test('no_uat_artifacts is true when no checks exist', () => { + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.no_uat_artifacts, true); + }); + + test('uat_files contains the filename', () => { + writeFile(tmpDir, 'my-UAT.md', makePassingUat(1)); + const report = evaluateUatPassed(tmpDir); + assert.ok(report.uat_files.includes('my-UAT.md'), + `uat_files should include 'my-UAT.md', got: ${JSON.stringify(report.uat_files)}`); + }); + + test('verification_files contains the filename', () => { + writeFile(tmpDir, 'phase-UAT.md', makePassingUat(1)); + writeFile(tmpDir, 'phase-VERIFICATION.md', '---\nstatus: passed\n---\n'); + const report = evaluateUatPassed(tmpDir); + assert.ok(report.verification_files.includes('phase-VERIFICATION.md'), + `verification_files should include 'phase-VERIFICATION.md', got: ${JSON.stringify(report.verification_files)}`); + }); +}); + +// ─── FIX A regression: nested-fence (~~~ inside ```) ───────────────────────── + +describe('FIX A — nested fence: ~~~ inside ``` does not prematurely close outer fence', () => { + let tmpDir; + + beforeEach(() => { tmpDir = makeTmpDir(); }); + afterEach(() => { rmDir(tmpDir); }); + + test('parseUatResultItems sees [] for fake inside ``` that encloses ~~~', () => { + // A backtick fence that contains an inner ~~~ fence with a fake test block. + // The ~~~ must NOT close the ``` fence — the whole interior is content and is dropped. + const raw = [ + '```', + '~~~', + '### 1. Fake', + 'expected: X', + 'result: passed', + '~~~', + '```', + ].join('\n'); + const clean = stripFalsePositiveContexts(raw); + const items = parseUatResultItems(clean); + assert.strictEqual(items.length, 0, 'fake inside nested fence must not leak through'); + }); + + test('evaluateUatPassed → passed:false + no_uat_artifacts:true for nested-fence-only file', () => { + const raw = [ + '```', + '~~~', + '### 1. Fake', + 'expected: X', + 'result: passed', + '~~~', + '```', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', raw); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false, 'nested-fence fake must not flip passed'); + assert.strictEqual(report.no_uat_artifacts, true, 'no real items → no_uat_artifacts:true'); + assert.ok(!report.checks.some(c => c.name === 'Fake'), 'fake must not appear in checks'); + }); + + test('balanced nested fence (``` inside ~~~) is NOT flagged as malformed', () => { + // ~~~ outer, ``` inner — properly closed — should not trigger unterminatedFence + const raw = [ + '~~~', + '```', + 'code', + '```', + '~~~', + '', + '### 1. Real Test', + 'expected: Works', + 'result: passed', + ].join('\n'); + const { unterminatedFence } = analyzeMarkdown(raw); + assert.strictEqual(unterminatedFence, false, 'balanced nested fence must not be flagged'); + const clean = stripFalsePositiveContexts(raw); + const items = parseUatResultItems(clean); + assert.strictEqual(items.length, 1, 'real test outside fence must still be found'); + assert.strictEqual(items[0].result, 'passed'); + }); + + test('real result:passed outside ``` that encloses ~~~ → passed:true (no false-block)', () => { + const raw = [ + '---', + 'status: passed', + '---', + '', + '```', + '~~~', + '### 1. Fake', + 'result: passed', + '~~~', + '```', + '', + '### 1. Real Test', + 'expected: Works', + 'result: passed', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', raw); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, true, 'real result outside nested fence must still pass'); + assert.ok(!report.checks.some(c => c.name === 'Fake'), 'fake must not appear in checks'); + }); +}); + +// ─── FIX B regression: cross-line result value ──────────────────────────────── + +describe('FIX B — cross-line result: value must be on the same line', () => { + test('result: with value on next line → result:missing (not passed)', () => { + const content = [ + '### 1. Cross-line Test', + 'expected: Y', + 'result:', + '', + 'passed', + ].join('\n'); + const items = parseUatResultItems(content); + assert.strictEqual(items.length, 1); + assert.notStrictEqual(items[0].result, 'passed', + 'result value on a subsequent line must not be captured as passed'); + assert.strictEqual(items[0].result, 'missing', + 'cross-line result must yield missing (blocker)'); + }); + + test('evaluateUatPassed → passed:false for cross-line result:passed', () => { + const tmpDir = makeTmpDir(); + try { + const content = [ + '### 1. Cross-line Test', + 'expected: Y', + 'result:', + '', + 'passed', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false); + assert.ok(!report.checks.some(c => c.result === 'passed'), + 'cross-line result must not produce a passing check'); + } finally { + rmDir(tmpDir); + } + }); +}); + +// ─── FIX C regression: masked-comment (earlier closed comment + later unterminated) ── + +describe('FIX C — dangling comment survives earlier balanced comment', () => { + test('analyzeMarkdown detects unterminated comment after a properly closed one', () => { + const raw = [ + '', + 'Some text', + '', + '', + '', + 'Some text', + '', + '', + '### 1. Real Test', + 'result: passed', + ].join('\n'); + const { unterminatedComment } = analyzeMarkdown(raw); + assert.strictEqual(unterminatedComment, false, + 'only balanced comments must not be flagged as unterminated'); + }); +}); + +// ─── FIX D regression: balanced mixed fences are NOT flagged ───────────────── + +describe('FIX D — odd-fence-count heuristic replaced: balanced mixed fences not flagged', () => { + test('analyzeMarkdown: ``` followed by ~~~ (both balanced) → unterminatedFence:false', () => { + const raw = [ + '```', + 'code', + '```', + '', + '~~~', + 'more code', + '~~~', + ].join('\n'); + const { unterminatedFence } = analyzeMarkdown(raw); + assert.strictEqual(unterminatedFence, false, + 'two separate balanced fences must not trigger unterminatedFence'); + }); + + test('evaluateUatPassed: multiple balanced fences + real passing test → passed:true, no malformed blocker', () => { + const tmpDir = makeTmpDir(); + try { + const raw = [ + '---', + 'status: passed', + '---', + '', + '```', + 'result: fake', + '```', + '', + '~~~', + 'result: also fake', + '~~~', + '', + '### 1. Real Test', + 'expected: Works', + 'result: passed', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', raw); + const report = evaluateUatPassed(tmpDir); + assert.ok(!report.blockers.some(b => /malformed/i.test(b)), + `Balanced mixed fences must not produce malformed blocker, got: ${JSON.stringify(report.blockers)}`); + assert.strictEqual(report.passed, true, + 'real test outside balanced fences must still pass'); + } finally { + rmDir(tmpDir); + } + }); +}); + +// ─── FIX E extra: fake-not-in-checks assertions for existing mixed tests ────── + +describe('FIX E — fake items must NOT appear in checks (absence assertions)', () => { + let tmpDir; + + beforeEach(() => { tmpDir = makeTmpDir(); }); + afterEach(() => { rmDir(tmpDir); }); + + test('fake inside fence + real pending: fake NOT in checks', () => { + const content = [ + '---', 'status: partial', '---', '', + '```', + '### 10. Fake', + 'expected: Fake', + 'result: passed', + '```', + '', + '### 1. Real Test', + 'expected: Something', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false); + assert.ok(!report.checks.some(c => c.name === 'Fake'), + 'fake test inside fence must not appear in checks'); + }); + + test('fake inside blockquote + real pending: fake NOT in checks', () => { + const content = [ + '---', 'status: partial', '---', '', + '> ### 10. Fake', + '> expected: Fake', + '> result: passed', + '', + '### 1. Real Test', + 'expected: Something', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false); + assert.ok(!report.checks.some(c => c.name === 'Fake'), + 'fake test inside blockquote must not appear in checks'); + }); + + test('fake inside HTML comment + real pending: fake NOT in checks', () => { + const content = [ + '---', 'status: partial', '---', '', + '', + '', + '### 1. Real Test', + 'expected: Something', + 'result: pending', + '', + ].join('\n'); + writeFile(tmpDir, 'phase-UAT.md', content); + const report = evaluateUatPassed(tmpDir); + assert.strictEqual(report.passed, false); + assert.ok(!report.checks.some(c => c.name === 'Fake'), + 'fake test inside HTML comment must not appear in checks'); + }); +}); + +// ─── Property-based test (fast-check) ───────────────────────────────────────── + +describe('evaluateUatPassed — property: wrapping in false-positive context never flips to passed', () => { + let tmpDir; + + beforeEach(() => { + tmpDir = makeTmpDir(); + }); + + afterEach(() => { + rmDir(tmpDir); + }); + + test('fc: inserting result:passed inside wrapper context never flips a failing UAT to passed', () => { + // A baseline UAT file that has a pending item — it must always evaluate to passed:false + // regardless of how many "result: passed" lines we inject inside fenced/blockquote/comment wrappers. + const baseFailingBody = [ + '### 1. Real Test', + 'expected: It works', + 'result: pending', + '', + ].join('\n'); + + const wrappers = fc.constantFrom( + // backtick fence + (inner) => '```\n' + inner + '\n```', + // tilde fence + (inner) => '~~~\n' + inner + '\n~~~', + // HTML comment + (inner) => '', + // blockquote — prefix each line + (inner) => inner.split('\n').map(l => '> ' + l).join('\n'), + ); + + fc.assert( + fc.property(wrappers, fc.nat(3), (wrap, extraCount) => { + // Build a "fake passing block" that would fool a naive regex + const fakePassingLines = Array.from({ length: extraCount + 1 }, (_, i) => + `### ${i + 10}. Fake Test ${i + 10}\nexpected: Fake\nresult: passed` + ).join('\n'); + + const fullContent = [ + '---', + 'status: partial', + '---', + '', + wrap(fakePassingLines), + '', + baseFailingBody, + ].join('\n'); + + // Write to a unique tmp file to avoid cross-test state + const fcDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-fc-uat-')); + try { + fs.writeFileSync(path.join(fcDir, 'feature-UAT.md'), fullContent, 'utf-8'); + const report = evaluateUatPassed(fcDir); + // The pending item must always keep passed:false + // AND fake items injected via wrappers must never appear in checks + const hasFakeInChecks = report.checks.some(c => c.name.startsWith('Fake Test')); + return report.passed === false && !hasFakeInChecks; + } finally { + cleanup(fcDir); + } + }), + { numRuns: 50 } + ); + }); +});