Files
msd-core/src/core-utils.cts
Tom Boucher 0624c5da6f chore(#3212): src/text-lines.cts is the sole owner of line-terminator handling — Phase 2 (#3420)
* test(#3413): failing-first suite for the line-terminator seam

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Tests only — src/text-lines.cts
does not exist yet, so tests/text-lines.test.cjs fails with MODULE_NOT_FOUND
at its require line, which is the intended RED.

The frontmatter.test.cjs additions drive #3360 (confirmed-bug) fail-first:
parseMustHavesBlock currently returns [] for every must_haves block on a
CRLF-authored plan file, because \r is its own LineTerminator in ECMAScript
and two /m-anchored \s* patterns can absorb it, inflating a captured indent
by one character and tripping the "not nested under must_haves" guard.
Verified locally against the current (unfixed) compiled module: both the
direct repro and the silent-exit "blank line before must_haves:" variant
return [] today. A parity property test (crlf vs lf must deep-equal for
every block name) matches a pattern this maintainer has required repeatedly
for prior CRLF fixes in this codebase (Cortex-recorded, verify_intent=held).

The no-crlf-fragile-split.rule.test.cjs additions lock the eslint rule's
future fix-hint text (pointing at splitLines()) and its self-reference
non-violation (the seam's own correct \r?\n split must never flag itself).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* chore(#3413): src/text-lines.cts owns line-terminator handling

Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Adds splitLines/normalizeEol/
detectEol/joinLines and migrates frontmatter.cts onto it.

parseMustHavesBlock (#3360, confirmed-bug) returned [] for every
must_haves block on a CRLF plan file. Root cause: \r is its own
LineTerminator in ECMAScript, so under /m two \s*-anchored indentation
lookups could match at the position INSIDE a \r\n pair and absorb the
terminator, inflating the captured indent by one character and tripping
the "not nested under must_haves" guard. Two silent exits, one with a
diagnostic and one without (a blank line before must_haves: hits the
silent path). Fixed by converting both lookups from a whole-string /m
match to split-then-scan — splitLines first, then a per-line, non-/m
match — the same structural pattern parseYamlRegion (30 lines away in
the same file) already used safely. Nothing downstream of the two
lookups changed; blockLines is now sliced from the already-split array
instead of re-splitting a substring, but its contents are unchanged for
LF input, and the per-line dash/kv parsing loop is untouched.

A parity property test (CRLF and LF plans parse to identical must_haves
for every block name) matches a pattern this maintainer has required
repeatedly for prior CRLF fixes in this file's neighborhood (Cortex:
7 recorded decisions, verify_intent -> held).

frontmatter.cts's other .split(/\r?\n/) call sites (parseYamlRegion,
isFrontmatterShaped, sliceTopLevelFrontmatterSegments, spliceFrontmatter)
are rerouted onto splitLines — a literal 1:1 substitution, zero behavior
change, since splitLines IS that same regex plus a type guard.

The 4 scripts/normalizeLineEndings copies (gen-registry, gen-loop-host-
contract, gen-capability-registry, gen-context-index) are deleted and
rerouted onto normalizeEol, which strips a bare unpaired \r exactly like
the deleted copies did (not just \r\n pairs) -- verified against each
script's own --check mode against its real generated output.

local/no-crlf-fragile-split widens from tests/ to src/**/*.cts, with its
fix-hint message now naming splitLines() instead of the raw regex --
the prohibition finally has a primitive to point at. Detection logic
unchanged in this phase (deliberate scope limit, see design doc Known
limits: the rule doesn't yet recognize safeReadFile/platformReadSync as
a content source, and has no detector for the \s-adjacent-to-anchor
shape that is #3360's actual mechanism -- the CLASS is converged by the
direct fix + regression test regardless).

joinLines/detectEol are NOT wired into frontmatter.cts's own write path
(cmdFrontmatterSet/Merge -> platformWriteSync) -- verified that
platformWriteSync already, unconditionally converts CRLF->LF on every
.md write today as a pre-existing policy owned by a different module,
and ADR-3212's backward-compatibility clause rules out a file-format
change in any phase. Stated explicitly in Known limits rather than left
for a reader to discover.

Six-gate ripple: .gitignore, eslint.config.mjs (src/**/*.cts block),
docs/INVENTORY.md + INVENTORY-MANIFEST.json (regenerated), CONTEXT.md
glossary (Text Lines Module, mirroring Phase 1's Pattern Module entry).

Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md
Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md

* fix(#3413): fix 13 pre-existing CRLF-fragile splits the widened rule found

Widening local/no-crlf-fragile-split from tests/ to src/**/*.cts (the
previous commit) immediately surfaced 13 real, pre-existing violations
across 10 files -- undetected until now because the rule never scanned
src/. This is the exact defect class ADR-3212 exists to close, playing
out again one phase after Phase 1 hit the same shape ("the new lint
rule -- once live -- found 27 more"). Per CLAUDE.md's no-defer rule,
fixed inline rather than deferred or suppressed; there is no
established suppression convention for this rule in src/ and inventing
one now would undermine the point of widening it.

audit.cts, broken-windows.cts, core-utils.cts, init.cts, milestone.cts,
phase.cts (x3), profile-output.cts, roadmap.cts (x2): bare-\n splits or
regex character classes widened to \r?\n / [^\r\n], each following the
same pattern already established migrating frontmatter.cts.

phase-estimation.cts: `\r?(?:\n|$)` restructured to `(?:\r?\n|\r?$)` --
already semantically CRLF-safe, but the rule's lexical scanner doesn't
recognize \r? guarding a group (only \r? immediately before a literal
\n). Verified the two forms are equivalent across all four EOL/EOF
cases before restructuring, not assumed.

roadmap-upgrade.cts needed two coupled sites, not the one flagged line:
computeMigrationPlan and applyMigration must agree on line
representation for the lines[edit.lineIndex] === edit.from equality
check to hold, and the write-back needed joinLines + detectEol -- a
plain lines.join('\n') was silently flattening a CRLF ROADMAP.md to LF
wholesale on every migration. This is the first real production
consumer of joinLines/detectEol in this epic (frontmatter.cts's own
write path doesn't use them -- see the previous commit's Known limits).

Fixing the 13 flagged sites surfaced 4 more adjacent same-shape sites
the rule doesn't track (.search() and new RegExp(dynamicString) aren't
in its tracked call/construction set). Investigated each empirically --
hand-tracing this exact bug class already produced one wrong conclusion
earlier in this phase (a detectEol design-doc arithmetic error), so
these were verified with real CRLF fixtures rather than reasoned about
on paper:

  - audit.cts (scanTodos): REAL bug, fixed. `bodyMatch.trim().split
    ('\n')[0]` leaked a trailing \r into a user-visible todo summary on
    CRLF input -- .trim() only strips the string's outer edges, not a
    \r sitting mid-string before the first bare \n. Now splitLines(...)
    [0].
  - phase.cts (cmdPhaseInsert, bullet-style branch): REAL bug, fixed.
    [^\n]* in targetBulletPattern swallowed a line's trailing \r on
    CRLF input, shifting the computed insert position to land INSIDE
    the \r\n pair; combined with a hardcoded '\n' bullet separator, a
    CRLF ROADMAP.md ended up with a mixed CRLF/LF result after an
    insert. Fixed with two coupled changes (either alone still
    corrupts, verified both ways): [^\r\n]* in the pattern, and the new
    bullet's leading terminator now comes from detectEol(rawContent).
  - roadmap.cts (cmdRoadmapAnnotateDependencies phase-boundary scan):
    investigated, genuinely safe, left untouched. The .search(/\n#{2,4}
    .../) boundary-finder and the [^\n]*-based heading match were
    empirically verified on a 3-phase CRLF fixture -- the only stray \r
    ends up at the tail of an intermediate phaseSection string that is
    only ever used for .test()-based idempotency checks, never for an
    exact-match comparison or written back to disk. No corruption on
    round-trip.

Every fix re-verified: npm run build:lib clean, npx eslint
'src/**/*.cts' --no-cache reports 0 problems (was 13), and each
fixed function's existing LF-input tests were spot-checked unchanged.

* fix(#3413): apply orthogonal review findings

Two isolated review engines (correctness + security) ran against the
full diff and found three majors, one real security issue, and several
disclosure-worthy minors. All fixed or explicitly disclosed with
evidence; nothing deferred.

MAJOR — detectEol's tie-break contradicted its own documented contract.
Code returned '\n' on a 1:1 crlf/bare-LF tie; every doc (design doc,
CONTEXT.md, the function's own comment) says ties resolve to '\r\n'.
The existing test masked this by reusing the same tie fixture the
buggy code happened to satisfy, rather than a genuine LF-majority
case. Root cause: an Edit attempted earlier in this phase to fix this
exact arithmetic error was blocked by the tier guard, and a later
dispatch was incorrectly told it had already landed. Fixed: condition
is now crlfCount >= bareLfCount; the test fixture corrected to a
genuine 2:1 majority, with a new explicit tie-case test.

MAJOR — phase.cts's cmdPhaseInsert built an EOL-aware bulletEntry via
detectEol(rawContent), justified by a comment claiming a hardcoded
'\n' corrupts a CRLF ROADMAP.md. False: this write goes through
platformWriteSync, whose normalizeContent/_normalizeMd unconditionally
converts CRLF->LF for any .md target — the templating was inert dead
code, erased before the file is ever written. Reverted to hardcoded
'\n', comment corrected to state the true reasoning. The separate
[^\n]* -> [^\r\n]* widening one function up (a real splice-position
fix, independent of final EOL) was kept.

MAJOR — roadmap-upgrade.cts's stated rationale for switching onto
splitLines/joinLines was wrong (both functions always agreed on line
representation, before and after — the claimed equality-check risk
never existed), and the change it justified introduced a real
regression: forcing every line onto one dominant terminator silently
rewrites untouched lines' EOL on a mixed-CRLF/LF ROADMAP.md. This
write path uses raw fs.writeFileSync, not platformWriteSync, so unlike
the phase.cts case above the regression is genuinely live.

Fixing this took two attempts. The first attempt (revert to
split('\n')/join('\n') plus a suppression comment) was correctly
blocked by an agent that discovered local/no-crlf-fragile-split is a
PROTECTED_RULES entry in tests/portability-rule-disable-ban.test.cjs —
a hard, out-of-band, ADR-1703-governed guardrail banning any
eslint-disable of this rule anywhere in src/**/*.cts. That agent also
detected and correctly disregarded an injected instruction that
appeared in tool output during a git operation, per this session's
untrusted-content policy. The actual fix: computeMigrationPlan
reverted to roadmapContent.split('\n') (confirmed lint-clean — the
rule's data-flow tracking only follows a variable's initializer, and
this one is declared empty then reassigned in a try block).
applyMigration's write-back now splices edits against the ORIGINAL
content string via indexOf('\n', pos) boundary-walking instead of a
full split/rejoin, so every untouched character — including every
line's own terminator — is copied byte-for-byte. A capture-group split
(/(\r\n|\n)/, preserving terminators inline) was tried first and
empirically confirmed to still trip the rule before this approach was
chosen instead.

MINOR (security) — roadmap.cts's cmdRoadmapAnnotateDependencies used
the STRING form of String#replace, so $&, $`, $', $1-$9 inside
must_haves.truths content (author-controlled) were interpreted as
replacement directives, splicing unrelated ROADMAP.md text into the
result. Fixed with the function-replacement form, which is never
pattern-interpreted. Verified before/after with the reviewer's exact
repro.

Also disclosed rather than silently left: test matrix row 31 (four
planned CRLF-materialized regression tests) was never implemented as
separate files — corrected to record the actual verification (a
manual --check run plus incidental existing coverage via each script's
normalizeLineEndings: normalizeEol alias). parseMustHavesBlock's LF
behavior was claimed byte-for-byte unchanged but the old
yaml.indexOf(blockMatch[0]) substring search could match an unrelated
earlier occurrence of the header text (e.g. inside a quoted value) —
the split-then-scan fix incidentally also closes this, a strict
improvement now recorded in the design doc rather than left implicit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3413): checkpoint 2 red — missing eslint ignore entry, RuleTester config error

Checkpoint 2 came back red with 5 failures on the reviewed sha, both
gaps genuinely undetectable by any local gate.

eslint.config.mjs was missing the 'gsd-core/bin/lib/text-lines.cjs'
ignores-list entry (ADR-457: generated .cjs artifacts are excluded from
direct type-aware linting). Phase 1's sibling entry (pattern.cjs) sits
two lines above it and was the exact precedent read while researching
the six-gate ripple for this module -- missed anyway. Caught by
tests/repo-invariants.test.cjs's bin/lib coverage-tracking test, which
only runs on the remote suite.

tests/no-crlf-fragile-split.rule.test.cjs's row-32 case specified both
`messageId` and `message` on the same RuleTester error assertion --
ESLint's RuleTester rejects that combination outright. This existed
since the test was first authored and was never caught locally: `npx
eslint` only lints the file's syntax, it does not execute RuleTester,
and local `node --test` is hard-blocked in this repo -- the assertion
had never actually RUN before this checkpoint. It was even present in
checkpoint 1's failure list, listed there as one of the "expected RED"
tests; I matched it against my expected-failures list by test NAME
only and never inspected the actual failure detail closely enough to
notice it was failing for the wrong reason (a RuleTester config error,
not the intended message-text mismatch). Fixed by keeping `message`
(the exact-text assertion the test exists to make) and dropping
`messageId`. Verified the crlfFragileSplit message string in
eslint-rules/no-crlf-fragile-split.cjs matches this assertion
character-for-character, and swept every other invalid case in the
file for the same double-specification bug (none found -- all
pre-existing cases use messageId alone).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3413): add Fixed changeset for the #3360 CRLF parsing fix

The sole user-visible effect of this phase. No breaking-change label
or Changed fragment needed — ADR-3212's Backward Compatibility section
names the Node floor (Phase 1, already shipped) as the epic's only
breaking change; Phase 2 has none.

* chore(#3413): backfill changeset pr number to 3420

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 20:27:48 -04:00

434 lines
19 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* Core Utilities — Shared low-level utility primitives
*
* ADR-857 rollout phase 2c: extracted from core.cts (issue #877).
* Owns POSIX path normalization, sub-repo/subdirectory scanning,
* phase file stats, slug/one-liner/plan-id helpers, and time-ago.
* Behaviour is preserved byte-for-behaviour from the prior location;
* only the module boundary moved. core.cjs re-exports every public symbol
* here under its own `export =` object so existing consumers are unaffected.
*
* New imports should pull core-utils helpers from core-utils.cjs directly.
*
* Dependencies (leaf modules only — no core.cjs, no loadConfig):
* - node:fs / node:path (stdlib)
* - ./phase-id.cjs (comparePhaseNum, used by readSubdirectories)
* - ./planning-workspace.cjs (findContextMdIn, used by getPhaseFileStats)
*/
import fs from 'node:fs';
import path from 'node:path';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import phaseIdModule = require('./phase-id.cjs');
const { comparePhaseNum } = phaseIdModule;
// eslint-disable-next-line @typescript-eslint/no-require-imports
import planningWorkspace = require('./planning-workspace.cjs');
const { findContextMdIn } = planningWorkspace;
// eslint-disable-next-line @typescript-eslint/no-require-imports
import shellCommandProjection = require('./shell-command-projection.cjs');
// ─── Path helpers ────────────────────────────────────────────────────────────
/**
* Normalize a relative path to always use forward slashes (cross-platform).
* Delegates to the single separator seam in shell-command-projection so there is
* exactly one implementation of native→POSIX conversion across the codebase.
*/
function toPosixPath(p: string): string {
return shellCommandProjection.toPosixPath(p);
}
/**
* Scan immediate child directories for separate git repos.
* Returns a sorted array of directory names that have their own `.git`.
* Excludes hidden directories and node_modules.
*/
function detectSubRepos(cwd: string): string[] {
const results: string[] = [];
try {
const entries = fs.readdirSync(cwd, { withFileTypes: true });
for (const entry of entries) {
if (!entry.isDirectory()) continue;
if (entry.name.startsWith('.') || entry.name === 'node_modules') continue;
const gitPath = path.join(cwd, entry.name, '.git');
try {
if (fs.existsSync(gitPath)) {
results.push(entry.name);
}
} catch { /* ignore */ }
}
} catch { /* ignore */ }
return results.sort();
}
// ─── Summary body helpers ─────────────────────────────────────────────────
/**
* Extract a one-liner from the summary body when it's not in frontmatter.
*/
function extractOneLinerFromBody(content: string | null | undefined): string | null {
if (!content) return null;
const normalized = content.replace(/\r\n/g, '\n').replace(/\r/g, '\n');
const body = normalized.replace(/^---\r?\n[\s\S]*?\r?\n---\r?\n*/, '');
// #3170: anchor to a summary-shaped heading (Summary / Overview /
// Accomplishments) so an incidental first heading (a rule list, task
// breakdown, deviation note) does not contribute its first bold run as the
// deliverable one-liner. Iterate headings in document order and extract from
// the first summary-shaped one that has a bold run; fall back to null (not the
// wrong text) when no such heading exists.
const headingRe = /^#+\s*([^\n]*)\n+\*\*([^*\n]+)\*\*([^\n]*)/gm;
let match: RegExpExecArray | null;
while ((match = headingRe.exec(body)) !== null) {
if (!/summary|overview|accomplish/i.test(match[1])) continue;
const boldInner = match[2].trim();
const afterBold = match[3];
if (/:\s*$/.test(boldInner)) {
const prose = afterBold.trim();
if (prose.length > 0) return prose;
} else if (boldInner.length > 0) {
return boldInner;
}
}
return null;
}
// ─── Misc utilities ───────────────────────────────────────────────────────────
function pathExistsInternal(cwd: string, targetPath: string): boolean {
const fullPath = path.isAbsolute(targetPath) ? targetPath : path.join(cwd, targetPath);
try {
fs.statSync(fullPath);
return true;
} catch {
return false;
}
}
function generateSlugInternal(text: string | null | undefined): string | null {
if (!text) return null;
// #2849: strip leading/trailing hyphens AFTER truncation, not only before.
// .substring(0, 60) can land on a separator, re-introducing a trailing hyphen
// the strip step exists to prevent. Truncation cannot add a leading hyphen, so
// running the full ^-+|-+$ pass last is equivalent for leading hyphens and
// fixes the trailing-hyphen-after-truncation case.
return transliterateForSlug(text).replace(/[^a-z0-9]+/g, '-').substring(0, 60).replace(/^-+|-+$/g, '');
}
// ─── Transliteration (#2848) ─────────────────────────────────────────────────
//
// Non-Latin titles used to reduce to an empty slug: the `[^a-z0-9]+` strip
// removed every character of an all-Cyrillic title and the hyphen cleanup left
// "". Callers then created unnamed phase directories (`01-`) and empty
// `milestone_slug` init JSON. The fix transliterates Cyrillic to ASCII BEFORE
// the existing ASCII filter, so a non-Latin title yields a usable ASCII slug
// while Latin-script text (which hits zero map entries) is byte-for-byte
// unchanged — the negative control is satisfied by construction.
//
// Multi-letter mappings (ж→zh, ч→ch, ш→sh, щ→sch, ю→yu, я→ya) are applied as a
// single pass; soft/hard signs (ъ, ь) drop to nothing rather than a hyphen.
// Scope is Cyrillic (Russian + the reported Ukrainian/Belarusian extras
// і ї є ґ ў) per the issue's confirmed-working patch. CJK and other
// non-transliterated scripts keep the existing strip-to-ASCII behavior.
const CYRILLIC_TRANSLITERATION: Readonly<Record<string, string>> = {
// multi-letter first (longest-match-safe within a single pass via ordered keys)
а: 'a', б: 'b', в: 'v', г: 'g', д: 'd', е: 'e', ё: 'e', ж: 'zh',
з: 'z', и: 'i', й: 'y', к: 'k', л: 'l', м: 'm', н: 'n', о: 'o',
п: 'p', р: 'r', с: 's', т: 't', у: 'u', ф: 'f', х: 'h', ц: 'ts',
ч: 'ch', ш: 'sh', щ: 'sch', ъ: '', ы: 'y', ь: '', э: 'e', ю: 'yu',
я: 'ya',
// Ukrainian / Belarusian extras reported in #2848
є: 'ye', і: 'i', ї: 'yi', ґ: 'g', ў: 'u',
};
const CYRILLIC_TRANSLITERATION_KEYS = Object.keys(CYRILLIC_TRANSLITERATION);
/**
* Lowercase + transliterate Cyrillic characters to ASCII. The output still
* contains non-ASCII for scripts outside the map (CJK, etc.) — the caller's
* existing `[^a-z0-9]+` filter handles those. Latin-script input is returned
* lowercased with no other change.
*
* Shared by `generateSlugInternal` (core-utils) and `slugify` (gsd2-import) so
* the transliteration step is not duplicated across the two slug helpers (#2848
* explicitly requires both be fixed).
*/
function transliterateForSlug(text: string): string {
const lowered = text.toLowerCase();
let out = '';
for (const ch of lowered) {
out += CYRILLIC_TRANSLITERATION_KEYS.includes(ch)
? CYRILLIC_TRANSLITERATION[ch]
: ch;
}
return out;
}
// ─── Phase file helpers ──────────────────────────────────────────────────────
interface PhaseFileStats {
plans: string[];
summaries: string[];
hasResearch: boolean;
hasContext: boolean;
hasVerification: boolean;
hasReviews: boolean;
scope: string;
}
// Minimal shape this module needs from plan-scan.cjs's scanPhasePlans result.
interface PlanScanResultShape {
planFiles: string[];
summaryFiles: string[];
scope: string;
}
/**
* Read a phase directory and return counts/flags for common file types.
*
* #3183 (ADR-3180 Decision 2): `plans`/`summaries` are derived from the
* canonical `scanPhasePlans` rather than a local re-derivation, so this
* primitive can no longer diverge from the single owner of live-plan
* counting. `scanPhasePlans`
* lives in plan-scan.cjs, which itself imports `countMatchedSummaries` from
* THIS module — a top-level import here would be circular, so the require
* is deferred (lazy, inside the function body) to break the cycle at load
* time. This mirrors the lazy-require seam already used elsewhere in this
* repo (see src/audit-command-router.cts) for the same "module A needs
* module B which needs module A" shape.
*
* `hasResearch`/`hasContext`/`hasVerification`/`hasReviews` stay on the raw
* `readdirSync` listing — they are not plan-scan concerns.
*
* Degrades on an unreadable directory instead of throwing: empty arrays,
* every flag false, scope UNREADABLE (mirroring scanPhasePlans's own
* degrade path).
*/
function getPhaseFileStats(phaseDir: string): PhaseFileStats {
// eslint-disable-next-line @typescript-eslint/no-require-imports, @typescript-eslint/no-unsafe-assignment
const scanPhasePlans: (dir: string) => PlanScanResultShape = require('./plan-scan.cjs');
const scan = scanPhasePlans(phaseDir);
let files: string[];
try {
files = fs.readdirSync(phaseDir);
} catch {
return {
plans: scan.planFiles,
summaries: scan.summaryFiles,
hasResearch: false,
hasContext: false,
hasVerification: false,
hasReviews: false,
scope: scan.scope,
};
}
return {
plans: scan.planFiles,
summaries: scan.summaryFiles,
hasResearch: files.some(f => f.endsWith('-RESEARCH.md') || f === 'RESEARCH.md'),
hasContext: findContextMdIn(files) !== null,
hasVerification: files.some(f => f.endsWith('-VERIFICATION.md') || f === 'VERIFICATION.md'),
hasReviews: files.some(f => f.endsWith('-REVIEWS.md') || f === 'REVIEWS.md'),
scope: scan.scope,
};
}
/**
* Read immediate child directories from a path.
* Returns [] if the path doesn't exist or can't be read.
* Pass sort=true to apply comparePhaseNum ordering.
*/
function readSubdirectories(dirPath: string, sort = false): string[] {
try {
const entries = fs.readdirSync(dirPath, { withFileTypes: true });
const dirs = entries.filter(e => e.isDirectory()).map(e => e.name);
return sort ? dirs.sort((a, b) => comparePhaseNum(a, b)) : dirs;
} catch {
return [];
}
}
/**
* Format a Date as a fuzzy relative time string (e.g. "5 minutes ago").
*/
function timeAgo(date: Date): string {
const seconds = Math.floor((Date.now() - date.getTime()) / 1000);
if (seconds < 5) return 'just now';
if (seconds < 60) return `${seconds} seconds ago`;
const minutes = Math.floor(seconds / 60);
if (minutes === 1) return '1 minute ago';
if (minutes < 60) return `${minutes} minutes ago`;
const hours = Math.floor(minutes / 60);
if (hours === 1) return '1 hour ago';
if (hours < 24) return `${hours} hours ago`;
const days = Math.floor(hours / 24);
if (days === 1) return '1 day ago';
if (days < 30) return `${days} days ago`;
const months = Math.floor(days / 30);
if (months === 1) return '1 month ago';
if (months < 12) return `${months} months ago`;
const years = Math.floor(days / 365);
if (years === 1) return '1 year ago';
return `${years} years ago`;
}
// ─── Plan ID helpers ─────────────────────────────────────────────────────────
/**
* Extract the canonical plan ID from a filename.
* Private to the core cluster — exported so core.cjs:searchPhaseInDir can
* import it from this leaf without circular dependency, but NOT re-exported
* from core.cjs's public `export =` block.
*/
function extractCanonicalPlanId(filename: string): string {
const base = filename.replace(/-PLAN\.md$/i, '').replace(/-SUMMARY\.md$/i, '').replace(/\.md$/i, '');
const parts = base.split('-').filter(Boolean);
// #2043: a phase/plan token component is either a zero-padded number (≥2 digits)
// or a single-digit-plus-letter id ("3A"); a *bare* single digit is a slug word,
// so "46-6-rs-…" is not paired into a "46-6" id while "3A-01" stays intact.
const tokenRe = /^(?:\d{2,}[A-Z]?|\d[A-Z])(?:\.\d+)*$/i;
// #2232: the PAIRED plan component is a zero-padded continuation segment
// (exactly 2 digits), so a ≥3-digit slug word (a year) is not paired into a
// bogus "14-2026" id. The leading phase component keeps tokenRe's unbounded
// \d{2,} — phase numbers ≥100 are legitimate; only continuations are capped.
const planTokenRe = new RegExp(
`^(?:${phaseIdModule.PHASE_CONTINUATION_SEGMENT_SOURCE}[A-Z]?|\\d[A-Z])(?:\\.\\d+)*$`,
'i',
);
const phaseIdx = parts.findIndex(p => tokenRe.test(p));
if (phaseIdx >= 0 && phaseIdx + 1 < parts.length && planTokenRe.test(parts[phaseIdx + 1])) {
return `${parts[phaseIdx]}-${parts[phaseIdx + 1]}`;
}
return base;
}
/**
* Count summaries that correspond to a real plan (#1988).
*
* A summary counts toward phase completion iff it pairs with an existing plan
* file. This excludes stray non-plan summaries — e.g. `30-FIX-CR02-SUMMARY.md`,
* `30-GAPCLOSURE-SUMMARY.md` — that inflate the raw `*-SUMMARY.md` count and
* silently flip a phase to Complete when plans are actually missing summaries.
*
* Pairing is layout-agnostic. For each plan, up to three candidate summary
* filenames are generated and any match suffices:
* 1. marker swap `PLAN`→`SUMMARY` on the basename — root padded
* (`30-01-PLAN.md`↔`30-01-SUMMARY.md`), nested (`PLAN-01.md`↔
* `SUMMARY-01.md`, incl. a `plans/` prefix), and bare (`PLAN.md`↔
* `SUMMARY.md`);
* 2. `<stem>-SUMMARY.md` — bare (`PLAN.md`↔`PLAN-SUMMARY.md`) and legacy
* (`14-PLAN-01.md`↔`14-PLAN-01-SUMMARY.md`);
* 3. extended `<n>-PLAN-<m>…`→`<n>-<m>-SUMMARY.md`
* (`3-PLAN-01-setup.md`↔`3-01-SUMMARY.md`).
* The swap is applied to the basename only so a lowercase `plans/` dir prefix
* isn't corrupted to `SUMMARYs/…`.
*/
function countMatchedSummaries(planFiles: string[], summaryFiles: string[]): number {
const summarySet = new Set(summaryFiles);
let matched = 0;
for (const plan of planFiles) {
if (summaryCandidates(plan).some((c) => summarySet.has(c))) matched++;
}
return matched;
}
/**
* The candidate `*-SUMMARY.md` filenames a single plan's completion record
* could take, per the three naming conventions documented above
* `countMatchedSummaries`. Extracted so `findUnsummarizedPlans` can reuse the
* exact same matching rule without duplicating it (a divergence between the
* count and the list would let a plan be counted as matched while still
* appearing in the unsummarized set, or vice versa).
*/
function summaryCandidates(plan: string): string[] {
const slashIdx = plan.lastIndexOf('/');
const dir = slashIdx >= 0 ? plan.slice(0, slashIdx + 1) : '';
const base = (dir ? plan.slice(dir.length) : plan).replace(/\.md$/i, '');
const candidates: string[] = [
dir + base.replace(/PLAN/i, 'SUMMARY') + '.md',
dir + base + '-SUMMARY.md',
];
const extended = base.match(/^(\d+)-PLAN-(\d+)/i);
if (extended) candidates.push(dir + extended[1] + '-' + extended[2] + '-SUMMARY.md');
// #3183: canonical-id form. Restores the coverage of the pre-migration
// bespoke I001 rule (verify.cts, pre-#3183, via validate.cjs's now-unused
// `canonicalPlanStem` — behaviourally identical to `extractCanonicalPlanId`,
// confirmed empirically), which matched a plan carrying a descriptive slug
// after its <phase>-<plan> id — e.g. `68-01-scaffolding-PLAN.md` — against
// a summary named only by the bare id — `68-01-SUMMARY.md`. None of the
// three candidates above produce that filename.
//
// Narrowed to the case `extractCanonicalPlanId` actually extracted an
// <id>-<id> pair (its result differs from the plan's own PLAN-stripped
// base). When no pair is found it falls back to returning that same base
// unchanged, which would otherwise push a redundant candidate identical to
// the `<stem>-SUMMARY.md` form above (e.g. `setup-PLAN.md` -> canonical
// 'setup' -> 'setup-SUMMARY.md', already candidate #2) rather than the
// original rule's actual behavior of matching only real id pairs.
//
// Collision, matching the original rule byte-for-behaviour: two plans that
// share the same <phase>-<plan> id but differ only in their descriptive
// slug (`68-01-alpha-PLAN.md` + `68-01-beta-PLAN.md`) both generate the
// SAME candidate `68-01-SUMMARY.md` and therefore BOTH read as summarized
// off one shared summary file. This is not a new regression: the
// pre-migration bespoke rule collapsed the same way (it populated one
// `summaryBases` Set keyed by canonical stem, so any plan whose canonical
// stem hit the set counted as matched, with no cardinality check against
// how many plans shared that stem).
const planStem = base.replace(/-PLAN$/i, '');
const canonicalId = extractCanonicalPlanId(base + '.md');
if (canonicalId !== planStem) candidates.push(dir + canonicalId + '-SUMMARY.md');
return candidates;
}
/**
* #2648: the plan files in `planFiles` that have NO matching completion record
* in `summaryFiles`, using the identical matching rule as `countMatchedSummaries`
* (so the count and the named list can never disagree). Callers that must NAME
* the missing plans — e.g. phase.complete's fail-closed coverage gate, which
* refuses completion when any non-retired plan lacks a SUMMARY — need the list,
* not just the count. `planFiles` is expected to be already superseded-filtered
* (the caller passes `scanPhasePlans(...).planFiles`, which drops
* `status: superseded` plans), so a deliberately-retired plan never appears
* here and never blocks completion.
*/
function findUnsummarizedPlans(planFiles: string[], summaryFiles: string[]): string[] {
const summarySet = new Set(summaryFiles);
return planFiles.filter((plan) => !summaryCandidates(plan).some((c) => summarySet.has(c)));
}
/**
* #3183: the mirror image of `findUnsummarizedPlans` — the summary files in
* `summaryFiles` that do NOT pair with ANY plan in `planFiles`, using the
* identical `summaryCandidates` matching rule as `countMatchedSummaries` /
* `findUnsummarizedPlans`. Callers that must name orphaned summaries (a
* stray non-plan summary, or a summary whose plan was renamed/removed) need
* this instead of a bespoke exact-suffix Set-diff, which cannot recognize
* the nested or extended naming forms `summaryCandidates` already handles —
* a divergence that produced false "orphan summary" warnings.
*/
function findOrphanSummaries(planFiles: string[], summaryFiles: string[]): string[] {
const claimed = new Set<string>();
for (const plan of planFiles) {
for (const candidate of summaryCandidates(plan)) claimed.add(candidate);
}
return summaryFiles.filter((s) => !claimed.has(s));
}
export = {
toPosixPath,
detectSubRepos,
extractOneLinerFromBody,
pathExistsInternal,
generateSlugInternal,
transliterateForSlug,
getPhaseFileStats,
readSubdirectories,
timeAgo,
extractCanonicalPlanId,
countMatchedSummaries,
findUnsummarizedPlans,
findOrphanSummaries,
};