* test(#3413): failing-first suite for the line-terminator seam Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Tests only — src/text-lines.cts does not exist yet, so tests/text-lines.test.cjs fails with MODULE_NOT_FOUND at its require line, which is the intended RED. The frontmatter.test.cjs additions drive #3360 (confirmed-bug) fail-first: parseMustHavesBlock currently returns [] for every must_haves block on a CRLF-authored plan file, because \r is its own LineTerminator in ECMAScript and two /m-anchored \s* patterns can absorb it, inflating a captured indent by one character and tripping the "not nested under must_haves" guard. Verified locally against the current (unfixed) compiled module: both the direct repro and the silent-exit "blank line before must_haves:" variant return [] today. A parity property test (crlf vs lf must deep-equal for every block name) matches a pattern this maintainer has required repeatedly for prior CRLF fixes in this codebase (Cortex-recorded, verify_intent=held). The no-crlf-fragile-split.rule.test.cjs additions lock the eslint rule's future fix-hint text (pointing at splitLines()) and its self-reference non-violation (the seam's own correct \r?\n split must never flag itself). Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md * chore(#3413): src/text-lines.cts owns line-terminator handling Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Adds splitLines/normalizeEol/ detectEol/joinLines and migrates frontmatter.cts onto it. parseMustHavesBlock (#3360, confirmed-bug) returned [] for every must_haves block on a CRLF plan file. Root cause: \r is its own LineTerminator in ECMAScript, so under /m two \s*-anchored indentation lookups could match at the position INSIDE a \r\n pair and absorb the terminator, inflating the captured indent by one character and tripping the "not nested under must_haves" guard. Two silent exits, one with a diagnostic and one without (a blank line before must_haves: hits the silent path). Fixed by converting both lookups from a whole-string /m match to split-then-scan — splitLines first, then a per-line, non-/m match — the same structural pattern parseYamlRegion (30 lines away in the same file) already used safely. Nothing downstream of the two lookups changed; blockLines is now sliced from the already-split array instead of re-splitting a substring, but its contents are unchanged for LF input, and the per-line dash/kv parsing loop is untouched. A parity property test (CRLF and LF plans parse to identical must_haves for every block name) matches a pattern this maintainer has required repeatedly for prior CRLF fixes in this file's neighborhood (Cortex: 7 recorded decisions, verify_intent -> held). frontmatter.cts's other .split(/\r?\n/) call sites (parseYamlRegion, isFrontmatterShaped, sliceTopLevelFrontmatterSegments, spliceFrontmatter) are rerouted onto splitLines — a literal 1:1 substitution, zero behavior change, since splitLines IS that same regex plus a type guard. The 4 scripts/normalizeLineEndings copies (gen-registry, gen-loop-host- contract, gen-capability-registry, gen-context-index) are deleted and rerouted onto normalizeEol, which strips a bare unpaired \r exactly like the deleted copies did (not just \r\n pairs) -- verified against each script's own --check mode against its real generated output. local/no-crlf-fragile-split widens from tests/ to src/**/*.cts, with its fix-hint message now naming splitLines() instead of the raw regex -- the prohibition finally has a primitive to point at. Detection logic unchanged in this phase (deliberate scope limit, see design doc Known limits: the rule doesn't yet recognize safeReadFile/platformReadSync as a content source, and has no detector for the \s-adjacent-to-anchor shape that is #3360's actual mechanism -- the CLASS is converged by the direct fix + regression test regardless). joinLines/detectEol are NOT wired into frontmatter.cts's own write path (cmdFrontmatterSet/Merge -> platformWriteSync) -- verified that platformWriteSync already, unconditionally converts CRLF->LF on every .md write today as a pre-existing policy owned by a different module, and ADR-3212's backward-compatibility clause rules out a file-format change in any phase. Stated explicitly in Known limits rather than left for a reader to discover. Six-gate ripple: .gitignore, eslint.config.mjs (src/**/*.cts block), docs/INVENTORY.md + INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary (Text Lines Module, mirroring Phase 1's Pattern Module entry). Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md * fix(#3413): fix 13 pre-existing CRLF-fragile splits the widened rule found Widening local/no-crlf-fragile-split from tests/ to src/**/*.cts (the previous commit) immediately surfaced 13 real, pre-existing violations across 10 files -- undetected until now because the rule never scanned src/. This is the exact defect class ADR-3212 exists to close, playing out again one phase after Phase 1 hit the same shape ("the new lint rule -- once live -- found 27 more"). Per CLAUDE.md's no-defer rule, fixed inline rather than deferred or suppressed; there is no established suppression convention for this rule in src/ and inventing one now would undermine the point of widening it. audit.cts, broken-windows.cts, core-utils.cts, init.cts, milestone.cts, phase.cts (x3), profile-output.cts, roadmap.cts (x2): bare-\n splits or regex character classes widened to \r?\n / [^\r\n], each following the same pattern already established migrating frontmatter.cts. phase-estimation.cts: `\r?(?:\n|$)` restructured to `(?:\r?\n|\r?$)` -- already semantically CRLF-safe, but the rule's lexical scanner doesn't recognize \r? guarding a group (only \r? immediately before a literal \n). Verified the two forms are equivalent across all four EOL/EOF cases before restructuring, not assumed. roadmap-upgrade.cts needed two coupled sites, not the one flagged line: computeMigrationPlan and applyMigration must agree on line representation for the lines[edit.lineIndex] === edit.from equality check to hold, and the write-back needed joinLines + detectEol -- a plain lines.join('\n') was silently flattening a CRLF ROADMAP.md to LF wholesale on every migration. This is the first real production consumer of joinLines/detectEol in this epic (frontmatter.cts's own write path doesn't use them -- see the previous commit's Known limits). Fixing the 13 flagged sites surfaced 4 more adjacent same-shape sites the rule doesn't track (.search() and new RegExp(dynamicString) aren't in its tracked call/construction set). Investigated each empirically -- hand-tracing this exact bug class already produced one wrong conclusion earlier in this phase (a detectEol design-doc arithmetic error), so these were verified with real CRLF fixtures rather than reasoned about on paper: - audit.cts (scanTodos): REAL bug, fixed. `bodyMatch.trim().split ('\n')[0]` leaked a trailing \r into a user-visible todo summary on CRLF input -- .trim() only strips the string's outer edges, not a \r sitting mid-string before the first bare \n. Now splitLines(...) [0]. - phase.cts (cmdPhaseInsert, bullet-style branch): REAL bug, fixed. [^\n]* in targetBulletPattern swallowed a line's trailing \r on CRLF input, shifting the computed insert position to land INSIDE the \r\n pair; combined with a hardcoded '\n' bullet separator, a CRLF ROADMAP.md ended up with a mixed CRLF/LF result after an insert. Fixed with two coupled changes (either alone still corrupts, verified both ways): [^\r\n]* in the pattern, and the new bullet's leading terminator now comes from detectEol(rawContent). - roadmap.cts (cmdRoadmapAnnotateDependencies phase-boundary scan): investigated, genuinely safe, left untouched. The .search(/\n#{2,4} .../) boundary-finder and the [^\n]*-based heading match were empirically verified on a 3-phase CRLF fixture -- the only stray \r ends up at the tail of an intermediate phaseSection string that is only ever used for .test()-based idempotency checks, never for an exact-match comparison or written back to disk. No corruption on round-trip. Every fix re-verified: npm run build:lib clean, npx eslint 'src/**/*.cts' --no-cache reports 0 problems (was 13), and each fixed function's existing LF-input tests were spot-checked unchanged. * fix(#3413): apply orthogonal review findings Two isolated review engines (correctness + security) ran against the full diff and found three majors, one real security issue, and several disclosure-worthy minors. All fixed or explicitly disclosed with evidence; nothing deferred. MAJOR — detectEol's tie-break contradicted its own documented contract. Code returned '\n' on a 1:1 crlf/bare-LF tie; every doc (design doc, CONTEXT.md, the function's own comment) says ties resolve to '\r\n'. The existing test masked this by reusing the same tie fixture the buggy code happened to satisfy, rather than a genuine LF-majority case. Root cause: an Edit attempted earlier in this phase to fix this exact arithmetic error was blocked by the tier guard, and a later dispatch was incorrectly told it had already landed. Fixed: condition is now crlfCount >= bareLfCount; the test fixture corrected to a genuine 2:1 majority, with a new explicit tie-case test. MAJOR — phase.cts's cmdPhaseInsert built an EOL-aware bulletEntry via detectEol(rawContent), justified by a comment claiming a hardcoded '\n' corrupts a CRLF ROADMAP.md. False: this write goes through platformWriteSync, whose normalizeContent/_normalizeMd unconditionally converts CRLF->LF for any .md target — the templating was inert dead code, erased before the file is ever written. Reverted to hardcoded '\n', comment corrected to state the true reasoning. The separate [^\n]* -> [^\r\n]* widening one function up (a real splice-position fix, independent of final EOL) was kept. MAJOR — roadmap-upgrade.cts's stated rationale for switching onto splitLines/joinLines was wrong (both functions always agreed on line representation, before and after — the claimed equality-check risk never existed), and the change it justified introduced a real regression: forcing every line onto one dominant terminator silently rewrites untouched lines' EOL on a mixed-CRLF/LF ROADMAP.md. This write path uses raw fs.writeFileSync, not platformWriteSync, so unlike the phase.cts case above the regression is genuinely live. Fixing this took two attempts. The first attempt (revert to split('\n')/join('\n') plus a suppression comment) was correctly blocked by an agent that discovered local/no-crlf-fragile-split is a PROTECTED_RULES entry in tests/portability-rule-disable-ban.test.cjs — a hard, out-of-band, ADR-1703-governed guardrail banning any eslint-disable of this rule anywhere in src/**/*.cts. That agent also detected and correctly disregarded an injected instruction that appeared in tool output during a git operation, per this session's untrusted-content policy. The actual fix: computeMigrationPlan reverted to roadmapContent.split('\n') (confirmed lint-clean — the rule's data-flow tracking only follows a variable's initializer, and this one is declared empty then reassigned in a try block). applyMigration's write-back now splices edits against the ORIGINAL content string via indexOf('\n', pos) boundary-walking instead of a full split/rejoin, so every untouched character — including every line's own terminator — is copied byte-for-byte. A capture-group split (/(\r\n|\n)/, preserving terminators inline) was tried first and empirically confirmed to still trip the rule before this approach was chosen instead. MINOR (security) — roadmap.cts's cmdRoadmapAnnotateDependencies used the STRING form of String#replace, so $&, $`, $', $1-$9 inside must_haves.truths content (author-controlled) were interpreted as replacement directives, splicing unrelated ROADMAP.md text into the result. Fixed with the function-replacement form, which is never pattern-interpreted. Verified before/after with the reviewer's exact repro. Also disclosed rather than silently left: test matrix row 31 (four planned CRLF-materialized regression tests) was never implemented as separate files — corrected to record the actual verification (a manual --check run plus incidental existing coverage via each script's normalizeLineEndings: normalizeEol alias). parseMustHavesBlock's LF behavior was claimed byte-for-byte unchanged but the old yaml.indexOf(blockMatch[0]) substring search could match an unrelated earlier occurrence of the header text (e.g. inside a quoted value) — the split-then-scan fix incidentally also closes this, a strict improvement now recorded in the design doc rather than left implicit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3413): checkpoint 2 red — missing eslint ignore entry, RuleTester config error Checkpoint 2 came back red with 5 failures on the reviewed sha, both gaps genuinely undetectable by any local gate. eslint.config.mjs was missing the 'gsd-core/bin/lib/text-lines.cjs' ignores-list entry (ADR-457: generated .cjs artifacts are excluded from direct type-aware linting). Phase 1's sibling entry (pattern.cjs) sits two lines above it and was the exact precedent read while researching the six-gate ripple for this module -- missed anyway. Caught by tests/repo-invariants.test.cjs's bin/lib coverage-tracking test, which only runs on the remote suite. tests/no-crlf-fragile-split.rule.test.cjs's row-32 case specified both `messageId` and `message` on the same RuleTester error assertion -- ESLint's RuleTester rejects that combination outright. This existed since the test was first authored and was never caught locally: `npx eslint` only lints the file's syntax, it does not execute RuleTester, and local `node --test` is hard-blocked in this repo -- the assertion had never actually RUN before this checkpoint. It was even present in checkpoint 1's failure list, listed there as one of the "expected RED" tests; I matched it against my expected-failures list by test NAME only and never inspected the actual failure detail closely enough to notice it was failing for the wrong reason (a RuleTester config error, not the intended message-text mismatch). Fixed by keeping `message` (the exact-text assertion the test exists to make) and dropping `messageId`. Verified the crlfFragileSplit message string in eslint-rules/no-crlf-fragile-split.cjs matches this assertion character-for-character, and swept every other invalid case in the file for the same double-specification bug (none found -- all pre-existing cases use messageId alone). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3413): add Fixed changeset for the #3360 CRLF parsing fix The sole user-visible effect of this phase. No breaking-change label or Changed fragment needed — ADR-3212's Backward Compatibility section names the Node floor (Phase 1, already shipped) as the epic's only breaking change; Phase 2 has none. * chore(#3413): backfill changeset pr number to 3420 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
434 lines
19 KiB
TypeScript
434 lines
19 KiB
TypeScript
/**
|
||
* Core Utilities — Shared low-level utility primitives
|
||
*
|
||
* ADR-857 rollout phase 2c: extracted from core.cts (issue #877).
|
||
* Owns POSIX path normalization, sub-repo/subdirectory scanning,
|
||
* phase file stats, slug/one-liner/plan-id helpers, and time-ago.
|
||
* Behaviour is preserved byte-for-behaviour from the prior location;
|
||
* only the module boundary moved. core.cjs re-exports every public symbol
|
||
* here under its own `export =` object so existing consumers are unaffected.
|
||
*
|
||
* New imports should pull core-utils helpers from core-utils.cjs directly.
|
||
*
|
||
* Dependencies (leaf modules only — no core.cjs, no loadConfig):
|
||
* - node:fs / node:path (stdlib)
|
||
* - ./phase-id.cjs (comparePhaseNum, used by readSubdirectories)
|
||
* - ./planning-workspace.cjs (findContextMdIn, used by getPhaseFileStats)
|
||
*/
|
||
|
||
import fs from 'node:fs';
|
||
import path from 'node:path';
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||
import phaseIdModule = require('./phase-id.cjs');
|
||
const { comparePhaseNum } = phaseIdModule;
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||
import planningWorkspace = require('./planning-workspace.cjs');
|
||
const { findContextMdIn } = planningWorkspace;
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||
import shellCommandProjection = require('./shell-command-projection.cjs');
|
||
|
||
// ─── Path helpers ────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Normalize a relative path to always use forward slashes (cross-platform).
|
||
* Delegates to the single separator seam in shell-command-projection so there is
|
||
* exactly one implementation of native→POSIX conversion across the codebase.
|
||
*/
|
||
function toPosixPath(p: string): string {
|
||
return shellCommandProjection.toPosixPath(p);
|
||
}
|
||
|
||
/**
|
||
* Scan immediate child directories for separate git repos.
|
||
* Returns a sorted array of directory names that have their own `.git`.
|
||
* Excludes hidden directories and node_modules.
|
||
*/
|
||
function detectSubRepos(cwd: string): string[] {
|
||
const results: string[] = [];
|
||
try {
|
||
const entries = fs.readdirSync(cwd, { withFileTypes: true });
|
||
for (const entry of entries) {
|
||
if (!entry.isDirectory()) continue;
|
||
if (entry.name.startsWith('.') || entry.name === 'node_modules') continue;
|
||
const gitPath = path.join(cwd, entry.name, '.git');
|
||
try {
|
||
if (fs.existsSync(gitPath)) {
|
||
results.push(entry.name);
|
||
}
|
||
} catch { /* ignore */ }
|
||
}
|
||
} catch { /* ignore */ }
|
||
return results.sort();
|
||
}
|
||
|
||
// ─── Summary body helpers ─────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Extract a one-liner from the summary body when it's not in frontmatter.
|
||
*/
|
||
function extractOneLinerFromBody(content: string | null | undefined): string | null {
|
||
if (!content) return null;
|
||
const normalized = content.replace(/\r\n/g, '\n').replace(/\r/g, '\n');
|
||
const body = normalized.replace(/^---\r?\n[\s\S]*?\r?\n---\r?\n*/, '');
|
||
// #3170: anchor to a summary-shaped heading (Summary / Overview /
|
||
// Accomplishments) so an incidental first heading (a rule list, task
|
||
// breakdown, deviation note) does not contribute its first bold run as the
|
||
// deliverable one-liner. Iterate headings in document order and extract from
|
||
// the first summary-shaped one that has a bold run; fall back to null (not the
|
||
// wrong text) when no such heading exists.
|
||
const headingRe = /^#+\s*([^\n]*)\n+\*\*([^*\n]+)\*\*([^\n]*)/gm;
|
||
let match: RegExpExecArray | null;
|
||
while ((match = headingRe.exec(body)) !== null) {
|
||
if (!/summary|overview|accomplish/i.test(match[1])) continue;
|
||
const boldInner = match[2].trim();
|
||
const afterBold = match[3];
|
||
if (/:\s*$/.test(boldInner)) {
|
||
const prose = afterBold.trim();
|
||
if (prose.length > 0) return prose;
|
||
} else if (boldInner.length > 0) {
|
||
return boldInner;
|
||
}
|
||
}
|
||
return null;
|
||
}
|
||
|
||
// ─── Misc utilities ───────────────────────────────────────────────────────────
|
||
|
||
function pathExistsInternal(cwd: string, targetPath: string): boolean {
|
||
const fullPath = path.isAbsolute(targetPath) ? targetPath : path.join(cwd, targetPath);
|
||
try {
|
||
fs.statSync(fullPath);
|
||
return true;
|
||
} catch {
|
||
return false;
|
||
}
|
||
}
|
||
|
||
function generateSlugInternal(text: string | null | undefined): string | null {
|
||
if (!text) return null;
|
||
// #2849: strip leading/trailing hyphens AFTER truncation, not only before.
|
||
// .substring(0, 60) can land on a separator, re-introducing a trailing hyphen
|
||
// the strip step exists to prevent. Truncation cannot add a leading hyphen, so
|
||
// running the full ^-+|-+$ pass last is equivalent for leading hyphens and
|
||
// fixes the trailing-hyphen-after-truncation case.
|
||
return transliterateForSlug(text).replace(/[^a-z0-9]+/g, '-').substring(0, 60).replace(/^-+|-+$/g, '');
|
||
}
|
||
|
||
// ─── Transliteration (#2848) ─────────────────────────────────────────────────
|
||
//
|
||
// Non-Latin titles used to reduce to an empty slug: the `[^a-z0-9]+` strip
|
||
// removed every character of an all-Cyrillic title and the hyphen cleanup left
|
||
// "". Callers then created unnamed phase directories (`01-`) and empty
|
||
// `milestone_slug` init JSON. The fix transliterates Cyrillic to ASCII BEFORE
|
||
// the existing ASCII filter, so a non-Latin title yields a usable ASCII slug
|
||
// while Latin-script text (which hits zero map entries) is byte-for-byte
|
||
// unchanged — the negative control is satisfied by construction.
|
||
//
|
||
// Multi-letter mappings (ж→zh, ч→ch, ш→sh, щ→sch, ю→yu, я→ya) are applied as a
|
||
// single pass; soft/hard signs (ъ, ь) drop to nothing rather than a hyphen.
|
||
// Scope is Cyrillic (Russian + the reported Ukrainian/Belarusian extras
|
||
// і ї є ґ ў) per the issue's confirmed-working patch. CJK and other
|
||
// non-transliterated scripts keep the existing strip-to-ASCII behavior.
|
||
const CYRILLIC_TRANSLITERATION: Readonly<Record<string, string>> = {
|
||
// multi-letter first (longest-match-safe within a single pass via ordered keys)
|
||
а: 'a', б: 'b', в: 'v', г: 'g', д: 'd', е: 'e', ё: 'e', ж: 'zh',
|
||
з: 'z', и: 'i', й: 'y', к: 'k', л: 'l', м: 'm', н: 'n', о: 'o',
|
||
п: 'p', р: 'r', с: 's', т: 't', у: 'u', ф: 'f', х: 'h', ц: 'ts',
|
||
ч: 'ch', ш: 'sh', щ: 'sch', ъ: '', ы: 'y', ь: '', э: 'e', ю: 'yu',
|
||
я: 'ya',
|
||
// Ukrainian / Belarusian extras reported in #2848
|
||
є: 'ye', і: 'i', ї: 'yi', ґ: 'g', ў: 'u',
|
||
};
|
||
|
||
const CYRILLIC_TRANSLITERATION_KEYS = Object.keys(CYRILLIC_TRANSLITERATION);
|
||
|
||
/**
|
||
* Lowercase + transliterate Cyrillic characters to ASCII. The output still
|
||
* contains non-ASCII for scripts outside the map (CJK, etc.) — the caller's
|
||
* existing `[^a-z0-9]+` filter handles those. Latin-script input is returned
|
||
* lowercased with no other change.
|
||
*
|
||
* Shared by `generateSlugInternal` (core-utils) and `slugify` (gsd2-import) so
|
||
* the transliteration step is not duplicated across the two slug helpers (#2848
|
||
* explicitly requires both be fixed).
|
||
*/
|
||
function transliterateForSlug(text: string): string {
|
||
const lowered = text.toLowerCase();
|
||
let out = '';
|
||
for (const ch of lowered) {
|
||
out += CYRILLIC_TRANSLITERATION_KEYS.includes(ch)
|
||
? CYRILLIC_TRANSLITERATION[ch]
|
||
: ch;
|
||
}
|
||
return out;
|
||
}
|
||
|
||
// ─── Phase file helpers ──────────────────────────────────────────────────────
|
||
|
||
interface PhaseFileStats {
|
||
plans: string[];
|
||
summaries: string[];
|
||
hasResearch: boolean;
|
||
hasContext: boolean;
|
||
hasVerification: boolean;
|
||
hasReviews: boolean;
|
||
scope: string;
|
||
}
|
||
|
||
// Minimal shape this module needs from plan-scan.cjs's scanPhasePlans result.
|
||
interface PlanScanResultShape {
|
||
planFiles: string[];
|
||
summaryFiles: string[];
|
||
scope: string;
|
||
}
|
||
|
||
/**
|
||
* Read a phase directory and return counts/flags for common file types.
|
||
*
|
||
* #3183 (ADR-3180 Decision 2): `plans`/`summaries` are derived from the
|
||
* canonical `scanPhasePlans` rather than a local re-derivation, so this
|
||
* primitive can no longer diverge from the single owner of live-plan
|
||
* counting. `scanPhasePlans`
|
||
* lives in plan-scan.cjs, which itself imports `countMatchedSummaries` from
|
||
* THIS module — a top-level import here would be circular, so the require
|
||
* is deferred (lazy, inside the function body) to break the cycle at load
|
||
* time. This mirrors the lazy-require seam already used elsewhere in this
|
||
* repo (see src/audit-command-router.cts) for the same "module A needs
|
||
* module B which needs module A" shape.
|
||
*
|
||
* `hasResearch`/`hasContext`/`hasVerification`/`hasReviews` stay on the raw
|
||
* `readdirSync` listing — they are not plan-scan concerns.
|
||
*
|
||
* Degrades on an unreadable directory instead of throwing: empty arrays,
|
||
* every flag false, scope UNREADABLE (mirroring scanPhasePlans's own
|
||
* degrade path).
|
||
*/
|
||
function getPhaseFileStats(phaseDir: string): PhaseFileStats {
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports, @typescript-eslint/no-unsafe-assignment
|
||
const scanPhasePlans: (dir: string) => PlanScanResultShape = require('./plan-scan.cjs');
|
||
const scan = scanPhasePlans(phaseDir);
|
||
|
||
let files: string[];
|
||
try {
|
||
files = fs.readdirSync(phaseDir);
|
||
} catch {
|
||
return {
|
||
plans: scan.planFiles,
|
||
summaries: scan.summaryFiles,
|
||
hasResearch: false,
|
||
hasContext: false,
|
||
hasVerification: false,
|
||
hasReviews: false,
|
||
scope: scan.scope,
|
||
};
|
||
}
|
||
|
||
return {
|
||
plans: scan.planFiles,
|
||
summaries: scan.summaryFiles,
|
||
hasResearch: files.some(f => f.endsWith('-RESEARCH.md') || f === 'RESEARCH.md'),
|
||
hasContext: findContextMdIn(files) !== null,
|
||
hasVerification: files.some(f => f.endsWith('-VERIFICATION.md') || f === 'VERIFICATION.md'),
|
||
hasReviews: files.some(f => f.endsWith('-REVIEWS.md') || f === 'REVIEWS.md'),
|
||
scope: scan.scope,
|
||
};
|
||
}
|
||
|
||
/**
|
||
* Read immediate child directories from a path.
|
||
* Returns [] if the path doesn't exist or can't be read.
|
||
* Pass sort=true to apply comparePhaseNum ordering.
|
||
*/
|
||
function readSubdirectories(dirPath: string, sort = false): string[] {
|
||
try {
|
||
const entries = fs.readdirSync(dirPath, { withFileTypes: true });
|
||
const dirs = entries.filter(e => e.isDirectory()).map(e => e.name);
|
||
return sort ? dirs.sort((a, b) => comparePhaseNum(a, b)) : dirs;
|
||
} catch {
|
||
return [];
|
||
}
|
||
}
|
||
|
||
/**
|
||
* Format a Date as a fuzzy relative time string (e.g. "5 minutes ago").
|
||
*/
|
||
function timeAgo(date: Date): string {
|
||
const seconds = Math.floor((Date.now() - date.getTime()) / 1000);
|
||
if (seconds < 5) return 'just now';
|
||
if (seconds < 60) return `${seconds} seconds ago`;
|
||
const minutes = Math.floor(seconds / 60);
|
||
if (minutes === 1) return '1 minute ago';
|
||
if (minutes < 60) return `${minutes} minutes ago`;
|
||
const hours = Math.floor(minutes / 60);
|
||
if (hours === 1) return '1 hour ago';
|
||
if (hours < 24) return `${hours} hours ago`;
|
||
const days = Math.floor(hours / 24);
|
||
if (days === 1) return '1 day ago';
|
||
if (days < 30) return `${days} days ago`;
|
||
const months = Math.floor(days / 30);
|
||
if (months === 1) return '1 month ago';
|
||
if (months < 12) return `${months} months ago`;
|
||
const years = Math.floor(days / 365);
|
||
if (years === 1) return '1 year ago';
|
||
return `${years} years ago`;
|
||
}
|
||
|
||
// ─── Plan ID helpers ─────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Extract the canonical plan ID from a filename.
|
||
* Private to the core cluster — exported so core.cjs:searchPhaseInDir can
|
||
* import it from this leaf without circular dependency, but NOT re-exported
|
||
* from core.cjs's public `export =` block.
|
||
*/
|
||
function extractCanonicalPlanId(filename: string): string {
|
||
const base = filename.replace(/-PLAN\.md$/i, '').replace(/-SUMMARY\.md$/i, '').replace(/\.md$/i, '');
|
||
const parts = base.split('-').filter(Boolean);
|
||
// #2043: a phase/plan token component is either a zero-padded number (≥2 digits)
|
||
// or a single-digit-plus-letter id ("3A"); a *bare* single digit is a slug word,
|
||
// so "46-6-rs-…" is not paired into a "46-6" id while "3A-01" stays intact.
|
||
const tokenRe = /^(?:\d{2,}[A-Z]?|\d[A-Z])(?:\.\d+)*$/i;
|
||
// #2232: the PAIRED plan component is a zero-padded continuation segment
|
||
// (exactly 2 digits), so a ≥3-digit slug word (a year) is not paired into a
|
||
// bogus "14-2026" id. The leading phase component keeps tokenRe's unbounded
|
||
// \d{2,} — phase numbers ≥100 are legitimate; only continuations are capped.
|
||
const planTokenRe = new RegExp(
|
||
`^(?:${phaseIdModule.PHASE_CONTINUATION_SEGMENT_SOURCE}[A-Z]?|\\d[A-Z])(?:\\.\\d+)*$`,
|
||
'i',
|
||
);
|
||
const phaseIdx = parts.findIndex(p => tokenRe.test(p));
|
||
if (phaseIdx >= 0 && phaseIdx + 1 < parts.length && planTokenRe.test(parts[phaseIdx + 1])) {
|
||
return `${parts[phaseIdx]}-${parts[phaseIdx + 1]}`;
|
||
}
|
||
return base;
|
||
}
|
||
|
||
/**
|
||
* Count summaries that correspond to a real plan (#1988).
|
||
*
|
||
* A summary counts toward phase completion iff it pairs with an existing plan
|
||
* file. This excludes stray non-plan summaries — e.g. `30-FIX-CR02-SUMMARY.md`,
|
||
* `30-GAPCLOSURE-SUMMARY.md` — that inflate the raw `*-SUMMARY.md` count and
|
||
* silently flip a phase to Complete when plans are actually missing summaries.
|
||
*
|
||
* Pairing is layout-agnostic. For each plan, up to three candidate summary
|
||
* filenames are generated and any match suffices:
|
||
* 1. marker swap `PLAN`→`SUMMARY` on the basename — root padded
|
||
* (`30-01-PLAN.md`↔`30-01-SUMMARY.md`), nested (`PLAN-01.md`↔
|
||
* `SUMMARY-01.md`, incl. a `plans/` prefix), and bare (`PLAN.md`↔
|
||
* `SUMMARY.md`);
|
||
* 2. `<stem>-SUMMARY.md` — bare (`PLAN.md`↔`PLAN-SUMMARY.md`) and legacy
|
||
* (`14-PLAN-01.md`↔`14-PLAN-01-SUMMARY.md`);
|
||
* 3. extended `<n>-PLAN-<m>…`→`<n>-<m>-SUMMARY.md`
|
||
* (`3-PLAN-01-setup.md`↔`3-01-SUMMARY.md`).
|
||
* The swap is applied to the basename only so a lowercase `plans/` dir prefix
|
||
* isn't corrupted to `SUMMARYs/…`.
|
||
*/
|
||
function countMatchedSummaries(planFiles: string[], summaryFiles: string[]): number {
|
||
const summarySet = new Set(summaryFiles);
|
||
let matched = 0;
|
||
for (const plan of planFiles) {
|
||
if (summaryCandidates(plan).some((c) => summarySet.has(c))) matched++;
|
||
}
|
||
return matched;
|
||
}
|
||
|
||
/**
|
||
* The candidate `*-SUMMARY.md` filenames a single plan's completion record
|
||
* could take, per the three naming conventions documented above
|
||
* `countMatchedSummaries`. Extracted so `findUnsummarizedPlans` can reuse the
|
||
* exact same matching rule without duplicating it (a divergence between the
|
||
* count and the list would let a plan be counted as matched while still
|
||
* appearing in the unsummarized set, or vice versa).
|
||
*/
|
||
function summaryCandidates(plan: string): string[] {
|
||
const slashIdx = plan.lastIndexOf('/');
|
||
const dir = slashIdx >= 0 ? plan.slice(0, slashIdx + 1) : '';
|
||
const base = (dir ? plan.slice(dir.length) : plan).replace(/\.md$/i, '');
|
||
const candidates: string[] = [
|
||
dir + base.replace(/PLAN/i, 'SUMMARY') + '.md',
|
||
dir + base + '-SUMMARY.md',
|
||
];
|
||
const extended = base.match(/^(\d+)-PLAN-(\d+)/i);
|
||
if (extended) candidates.push(dir + extended[1] + '-' + extended[2] + '-SUMMARY.md');
|
||
// #3183: canonical-id form. Restores the coverage of the pre-migration
|
||
// bespoke I001 rule (verify.cts, pre-#3183, via validate.cjs's now-unused
|
||
// `canonicalPlanStem` — behaviourally identical to `extractCanonicalPlanId`,
|
||
// confirmed empirically), which matched a plan carrying a descriptive slug
|
||
// after its <phase>-<plan> id — e.g. `68-01-scaffolding-PLAN.md` — against
|
||
// a summary named only by the bare id — `68-01-SUMMARY.md`. None of the
|
||
// three candidates above produce that filename.
|
||
//
|
||
// Narrowed to the case `extractCanonicalPlanId` actually extracted an
|
||
// <id>-<id> pair (its result differs from the plan's own PLAN-stripped
|
||
// base). When no pair is found it falls back to returning that same base
|
||
// unchanged, which would otherwise push a redundant candidate identical to
|
||
// the `<stem>-SUMMARY.md` form above (e.g. `setup-PLAN.md` -> canonical
|
||
// 'setup' -> 'setup-SUMMARY.md', already candidate #2) rather than the
|
||
// original rule's actual behavior of matching only real id pairs.
|
||
//
|
||
// Collision, matching the original rule byte-for-behaviour: two plans that
|
||
// share the same <phase>-<plan> id but differ only in their descriptive
|
||
// slug (`68-01-alpha-PLAN.md` + `68-01-beta-PLAN.md`) both generate the
|
||
// SAME candidate `68-01-SUMMARY.md` and therefore BOTH read as summarized
|
||
// off one shared summary file. This is not a new regression: the
|
||
// pre-migration bespoke rule collapsed the same way (it populated one
|
||
// `summaryBases` Set keyed by canonical stem, so any plan whose canonical
|
||
// stem hit the set counted as matched, with no cardinality check against
|
||
// how many plans shared that stem).
|
||
const planStem = base.replace(/-PLAN$/i, '');
|
||
const canonicalId = extractCanonicalPlanId(base + '.md');
|
||
if (canonicalId !== planStem) candidates.push(dir + canonicalId + '-SUMMARY.md');
|
||
return candidates;
|
||
}
|
||
|
||
/**
|
||
* #2648: the plan files in `planFiles` that have NO matching completion record
|
||
* in `summaryFiles`, using the identical matching rule as `countMatchedSummaries`
|
||
* (so the count and the named list can never disagree). Callers that must NAME
|
||
* the missing plans — e.g. phase.complete's fail-closed coverage gate, which
|
||
* refuses completion when any non-retired plan lacks a SUMMARY — need the list,
|
||
* not just the count. `planFiles` is expected to be already superseded-filtered
|
||
* (the caller passes `scanPhasePlans(...).planFiles`, which drops
|
||
* `status: superseded` plans), so a deliberately-retired plan never appears
|
||
* here and never blocks completion.
|
||
*/
|
||
function findUnsummarizedPlans(planFiles: string[], summaryFiles: string[]): string[] {
|
||
const summarySet = new Set(summaryFiles);
|
||
return planFiles.filter((plan) => !summaryCandidates(plan).some((c) => summarySet.has(c)));
|
||
}
|
||
|
||
/**
|
||
* #3183: the mirror image of `findUnsummarizedPlans` — the summary files in
|
||
* `summaryFiles` that do NOT pair with ANY plan in `planFiles`, using the
|
||
* identical `summaryCandidates` matching rule as `countMatchedSummaries` /
|
||
* `findUnsummarizedPlans`. Callers that must name orphaned summaries (a
|
||
* stray non-plan summary, or a summary whose plan was renamed/removed) need
|
||
* this instead of a bespoke exact-suffix Set-diff, which cannot recognize
|
||
* the nested or extended naming forms `summaryCandidates` already handles —
|
||
* a divergence that produced false "orphan summary" warnings.
|
||
*/
|
||
function findOrphanSummaries(planFiles: string[], summaryFiles: string[]): string[] {
|
||
const claimed = new Set<string>();
|
||
for (const plan of planFiles) {
|
||
for (const candidate of summaryCandidates(plan)) claimed.add(candidate);
|
||
}
|
||
return summaryFiles.filter((s) => !claimed.has(s));
|
||
}
|
||
|
||
export = {
|
||
toPosixPath,
|
||
detectSubRepos,
|
||
extractOneLinerFromBody,
|
||
pathExistsInternal,
|
||
generateSlugInternal,
|
||
transliterateForSlug,
|
||
getPhaseFileStats,
|
||
readSubdirectories,
|
||
timeAgo,
|
||
extractCanonicalPlanId,
|
||
countMatchedSummaries,
|
||
findUnsummarizedPlans,
|
||
findOrphanSummaries,
|
||
};
|