* test(#3707): failing-first coverage for reverting the fence-shortfall fold shield Pins the post-revert contract: a phase whose only gap is a fence shortfall must degrade the fold and withhold the milestone percentages, like every other gap class. Five of the eight rows are CONTROLS that pass before the change, and they carry more weight than the failing row. The failure mode of this revert is degrading TOO MUCH: a revert that sets foldScope outside the headingsSeen > 0 branch would withhold every percentage in the project, and only the no-gap control catches that. Another control catches a revert that collapses the two scopes into one and loses the distinction between what a phase reports and what the fold folds -- uat.scope must stay TRUNCATED for every gap either way, which it already is. The row that pinned the shielded behavior is rewritten rather than deleted. Deleting a test because the behavior it asserts is being reversed leaves the reversal unguarded. * fix(#3707): degrade the fold for every UAT gap class, reverting the fence-shortfall shield Maintainer decision. The two orthogonal engines split on this during #3707 and neither filed it as blocking, so it shipped in the shape the engine that raised the objection endorsed after verifying seven fixtures. The call has now gone the other way, restoring the fail-safe direction chosen twice already on this issue. The shield exempted one gap class from the fold's teeth. It could not do that safely: shortfallBlocks is a single tally incremented at exactly one site and spans BOTH a harmless fenced documentation sample AND a genuinely fence-straddled result: blocked row. Exempting it therefore could not exempt only the harmless case -- it also published a milestone percentage over a real, unread outstanding row. SCOPE.TRUNCATED means the scan could not SEE part of the evidence, which is exactly that case. scope and foldScope now agree: every gap class degrades both. The accepted over-report documented in uat.cts is unchanged and still documented there; what changed is only that it no longer buys an exemption from the fold. The comment block above it argued FOR the shield and is rewritten, because a comment defending behavior the code no longer has is worse than no comment. shortfallBlocks leaves this function's destructure but is untouched upstream, where audit-uat still consumes it. * fix(#3707): correct the caller comment, add the changeset, and name what the order tests guard Review found a SECOND comment still documenting the removed shield -- the caller's, beside the worstScope fold, stating that foldScope differs from scope for exactly one case which must not raise phase_scope_degraded or withhold the milestone's percentages. That is now the opposite of what the code does. I rewrote the buildUatRows comment in the previous commit and asserted in its message that a comment defending behavior the code no longer has is worse than no comment, then left exactly that one standing a few hundred lines away. The change had no changeset. It is user-visible: a milestone's percentage goes from published to withheld whenever any phase has a fence-shortfall-only gap. PR gates hard-fail a user-facing code diff without one. The two scopes are now identical at every return site. They are NOT collapsed -- that would change the return shape and the caller on what is meant to be a one-condition revert, and the seam is worth keeping if the distinction is ever wanted again -- but the declaration now says plainly that they agree by decision rather than by accident, so a reader does not have to re-derive it. The two order-independence tests were renamed. foldScope is monotonic with no reset path, so file order is structurally irrelevant and those rows could never have failed for the ordering reason their names promised. They do guard something real -- a multi-file phase degrading when any one file has a shortfall-only gap -- so they now say that instead. * test(#3707): failing-first coverage for the lone-CR UAT false-clean The parser splits on newline only, and the heading tokenizer agrees with it, so a lone carriage return is not a line boundary anywhere in it. CommonMark treats a lone CR as a line ending, so such a row renders to a human reader while being invisible to BOTH sides of the parser's symmetry invariant: no item, no shortfall, no headingsSeen. A phase hiding a result: blocked row this way reports 100 percent with zero diagnostics. Found by the security review of the fold-shield revert. It is the one false-clean class that revert does not reach, and it is the same bug class this issue exists to fix -- an unreadable row reported as clean. Nine rows. The LF control is what proves this is a separator defect rather than a content defect: identical bodies, one separator apart, and only one of them hides the row. CRLF and CR-inside-a-fence controls guard the coming normalization against double-counting or tearing content that legitimately contains a carriage return. Two further manifestations turned up while writing them: a leading CR breaks column-0 anchoring of the first heading, and an all-CR document flags a shortfall it cannot attribute to any row. * fix(#3707): treat a lone carriage return as a line ending in the UAT parser A lone CR was not a line boundary anywhere in the parser -- it split on newline only, and the heading tokenizer agreed with it. CommonMark treats a lone CR as a line ending, so such a row rendered to a human reader while being invisible to BOTH sides of the parser's symmetry invariant: no item, no shortfall, no headingsSeen. A phase hiding a result: blocked row that way reported 100 percent with zero diagnostics. Line endings are now normalized once at document ingress -- CRLF and lone CR both to newline -- at the two independent entry points, rather than teaching each split site about CR. Every downstream scan, offset and span therefore reads one convention. That single-frame property is deliberate: this issue already cost a HIGH when two scans read the same document through different frames. MY OWN END-TO-END TEST WAS WRONG and is replaced rather than weakened. It asserted that a lone-CR document must withhold its percentage, which reasons from the pre-fix symptom: after the fix the row is not hidden, it is surfaced, and this module deliberately keeps visible outstanding UAT work separate from completion percentages -- only unreadable evidence degrades scope. The success of the fix is what made the assertion false. The implementing agent refused to satisfy both it and the architecture and asked instead of bending either; it was right. What replaces it is a stronger contract: a lone-CR document and its LF twin, built from one source, must produce identical audit output -- scope, percent, every unresolved row by identity, and the diagnostic set. That is what 'a line-ending convention must not change what the audit reports' actually means, and it carries a non-vacuity check so it cannot pass with both sides empty. shortfallBlocks keeps being returned, now documented as currently unconsumed. An earlier reviewer told me audit-uat still consumed it and I passed that on as an instruction; it was wrong, and it was caught by checking rather than by me. * fix(#3707): normalize at the document read boundary, not at two call sites The lone-CR fix was half-applied and both review engines caught it independently. cmdAuditUat has four document ingresses, not the two I normalized: VERIFICATION.md and deferred-items.md still handed raw text to newline-only splitters, and the frontmatter extract in the UAT loop read raw content while its parser read normalized -- one audit entry mixing the two frames the fix exists to unify. Measured: a phase written twice from one source gave total_files 2 / total_items 4 under LF and results [] / total_items 0 under lone CR, with zero diagnostics. Normalizing two call sites and declaring it done is exactly why two were missed, so this moves it to the read boundary: every document now enters through a helper that normalizes, in audit-uat, in planning-inspect's readDocument, and in the shared verification-status read. Future parsers downstream get normalized text by construction rather than because someone remembered. That last seam also fixes an under-reporting case of the same root: a lone-CR VERIFICATION.md saying status: passed was read as missing, telling the user a verify step that had completed never ran. The parity test's load-bearing assertion is now marked as such. Four of its five equality checks still pass with the bug present -- only the unresolved-row identity differs -- so trimming that one as redundant would make the row vacuous. Second changeset added: the CR fix is user-visible independently of the fold revert, and one fragment covering both would have described neither. * test(#3707): failing-first coverage for the U+2028 and duplicate-result false-cleans Two more of the same class, both found by the security review of this branch and both reproduced before writing a line. normalizeLineEndings folds only carriage returns, but a JS /m anchor also treats U+2028 and U+2029 as line terminators while split on newline does not. That is the identical asymmetry the carriage-return bug exploited, one separator over, and worse in one respect: these are not CommonMark line endings, so a reader still sees the column-0 result: blocked that the tool discards. Measured: a scalar-internal result: pass placed after U+2028 wins over the real blocked line and the row disappears with no gap raised. Separately, and independent of any separator, a block with two column-0 result: lines resolves to the first with no ambiguity signalled. Prepending result: pass to a block therefore deletes an outstanding row silently; reversing the order surfaces it. Order deciding meaning is the defect, so the pair of rows pins the contract as ambiguity-is-a-gap rather than last-one-wins, leaving the fix room to implement the gap sensibly. Four controls: an ordinary marker in the same position (proving separator not content), legitimate U+2028 inside prose that must not be torn, a single result line, and a result line inside a fence that must not count as a second occurrence. * fix(#3707): scan result lines by split, not by a multiline anchor Two more false-cleans from the security review, both closed by the same change. A JS /m anchor treats U+2028 and U+2029 as line terminators while split on newline does not. A scalar-internal result: pass placed after one of those separators therefore matched as a line start and beat the real column-0 result: blocked, and the row vanished at 100 percent with no gap. Worse than the carriage-return case in one respect: these are not CommonMark line endings, so a reader still saw the blocked row the tool discarded. Separately, the non-global match returned the leftmost hit, so a block with two column-0 result: lines silently resolved to the first. Prepending result: pass deleted an outstanding row; reversing the order surfaced it. Order deciding meaning was the defect. Both close by scanning lines produced by split rather than by anchoring a regex inside the whole document: each line is tested on its own, and a count other than exactly one is reported as a parse gap instead of resolved to either candidate. I asked for U+2028 to be folded in normalizeLineEndings and that was wrong. Folding is length-preserving, so it would have made the U+2028 fixture byte-identical to the genuine two-result-line fixture -- while one requires a confident item and the other requires an ambiguity gap. No implementation can satisfy both once the distinguishing character is erased. The agent proved that and deviated rather than forcing it, which is why normalizeLineEndings still folds only carriage returns, now with a comment saying why. * fix(#3707): bound the ambiguity scan at the next heading-shaped line The split-based result scan regressed four pre-existing #3078/#3707 guards, each off by exactly one gap. My diagnosis was wrong. I read the off-by-one as double counting -- zero-result blocks taking both the new path and the pre-existing one -- and said to change the ambiguity condition from not-equal-one to greater-than-one. The agent checked and refused: the zero path was never duplicated. The real cause is double ATTRIBUTION. A block is sliced to the next TOKENIZED heading, so when the next row is untokenized -- hidden by a straddling fence, or indented and already counted by the shortfall scan -- that row's own result: line is absorbed into the previous block. The scan then saw two result lines across what are really two rows and raised a second, redundant gap on top of the one already counted elsewhere. Had the greater-than-one change gone in, the counts would have matched while the double attribution stayed. That is the compensating-adjustment failure I had asked it to refuse, and it did. The scan is now bounded at the first following heading-shaped line, either indent class, so a genuine same-block ambiguity is untouched while spillover from a row counted elsewhere is excluded. * fix(#3707): keep the U+2028 immunity, revert the ambiguity detection The ambiguity half of this change regressed the suite twice and is coming out. Attempt one double-attributed: a block is sliced to the next TOKENIZED heading, so when the real next row is untokenized its result: line was absorbed into the previous block and raised a second gap on a row already counted elsewhere. Four guards broke. Attempt two bounded the scan at the next heading-shaped line and broke thirty. An indented ### N. inside a block scalar is legitimate scalar CONTENT, not a heading, and truncating there defeats every #3078 guard that exists to stop scalar bodies being read as rows. Telling a genuinely hidden indented row apart from indented scalar text is a classification countUnattributedIndentedRows already owns; a raw regex does not have that information. What survives is the half that is sound and was never implicated in either regression: the result scan tests each line produced by split rather than anchoring a regex with the multiline flag over the whole block. split never treats U+2028 or U+2029 as a delimiter, so those separators can no longer manufacture a line start and steal a row. Everything else returns to first-match-wins, byte-identical to origin/next. The two tests pinning ambiguity-as-a-gap are removed with it, since the contract is no longer implemented here. The defect they described is real, pre-existing and independent of any separator -- result: pass before result: blocked silently deletes an outstanding row -- and it needs its own change with a scalar-aware counter rather than being wedged into a branch already carrying three fixes. * fix(#3707): correct the shared-seam rationale and restore U+2028 trailing text The revert left a stale rationale in core-utils, justifying the decision not to fold U+2028 by claiming uat.cts must tell a fake line start apart from a real second column-0 result: declaration that gets flagged as ambiguous. Nothing flags ambiguity any more; that behavior was reverted and the same file says so a few lines away. The decision is still right, the stated reason was false. This is the third stale comment this branch has shipped and had to fix, and the worst placed of them: core-utils is a shared leaf that every future document consumer will read for guidance. Rewritten to the true reason -- the scan tests each split line individually rather than anchoring over the block, so an exotic separator cannot manufacture a line start and folding is unnecessary. Also a real behavior delta I had not noticed. Dropping the multiline flag left the pattern's trailing .*$ in place, and dot never matches U+2028, so a genuine column-0 result: blocked whose TRAILING text contained one stopped parsing entirely -- a visible parse gap rather than a false clean, so fail-safe, but a regression against origin/next that nothing pinned. The trailing portion now matches any character and a test pins it by identity against its plain-LF twin. Plus the JSDoc orphaned when normalizeLineEndings moved to core-utils, and the changeset, which described neither the separator fix nor planning-inspect surfacing lone-CR rows. * fix(#3707): harden the acceptance gate, which had both halves of the same bug uat-predicate is a SECOND, independent UAT parser, and it is the one that decides phase uat-passed. It read raw bytes and anchored a multiline regex over unsplit text -- exactly the two defects this branch closed one module away in uat.cts. The consequence is worse than the audit surface it mirrors. Measured on identical bytes: a U+2028 scalar injection made the gate return passed true while planning inspect reported the same row as blocked and outstanding. The hardened surface and the gate disagreed, and the gate was the permissive one -- so a phase could be accepted over a row the audit could see and the gate could not. Both raw reads now go through the shared normalize seam and both scans test lines produced by split rather than anchoring over the document. First-match-wins, matching uat.cts; no ambiguity counting is reintroduced. Tests assert the AGREEMENT between the two surfaces rather than each separately, because divergence is the defect. Also finishes the same root cause one module over: phase complete's advisory pre-scan read raw bytes, so a lone-CR VERIFICATION.md lost its human_needed or gaps_found warning -- the fix verification.cts already got on this branch. And narrows the core-utils rationale I reworded last commit, which claimed consumers already avoid multiline anchors. uat.cts still has five over unsplit text. That is the fourth comment on this branch to assert something the code does not do, so it now states only what is true of core-utils itself. * fix(#3707): give structure and attribution different line frames, normalize the close audit Two more from review, and the first was a regression I introduced one commit earlier. Converting the gate's heading scan to split-then-match removed a detection origin/next had: a ### N. heading delimited by U+2028 was found by the old multiline scan and was not found after. So hardening the result scan quietly weakened the heading scan, and the gate stopped blocking on rows origin/next blocked -- the permissive direction, on the surface that decides acceptance. The insight I had missed is that the two scans need DIFFERENT frames. Heading detection is structure: there is no distinction to preserve, so it splits on newline or either exotic separator and finds a heading however it is delimited. The result scan is attribution: the newline-only frame is exactly what stops a scalar-internal result: from being read as a column-0 line, so it stays. One frame applied uniformly was the error. Second, a THIRD unnormalized parser family: the milestone-close audit read every artifact raw. A lone-CR VERIFICATION.md degraded to status unknown and was skipped, and deferred entries vanished outright -- measured as three items requiring decisions under LF and one under CR, on identical bytes. All nine scanner reads now normalize; six of them had the identical defect beyond the three review named. The acknowledge path stays deliberately raw, since it splices by byte offset, and now says so. Also pins the cross-newline result: divergence, and replaces three raw U+2028 literals in test source with escapes. A raw separator in a fixture is one formatter away from becoming an ordinary-character control that still passes -- vacuous in the only test pinning the separator fix. * fix(#3707): share one frame between the acknowledge writer and the audit reader Normalizing the audit scanners left the writer and the reader on different frames. cmdAuditAcknowledge derives its stored snapshot values from raw content -- correct for the SPLICE, which rewrites by byte offset -- but scanUatGaps and scanContextQuestions now recompute those same values from normalized content. For a lone-CR artifact the two can never match, so an acknowledgement never suppresses its item and it resurfaces on every audit: acknowledge became a silent no-op. Fail-safe in direction, since the item stays visible rather than being wrongly suppressed, but it is the writer and reader disagreeing about what a line is -- the exact class this branch exists to eliminate, and the fourth instance of it here. The derive functions now read a normalized copy while the splice keeps raw bytes and raw offsets, so both sides share one frame and the byte-offset rewrite is untouched. Round trip pinned for lone-CR and LF, with an existing LF marker asserted still recognised so the change cannot silently invalidate acknowledgements already in users' files. Also tightens an assertion that pinned this branch's own heading fix with a proxy: notStrictEqual against 'passed' also passes on 'pass', which IS a passing token, so it could not have caught a regression attributing a passing result to the recovered heading. It now pins the exact token. * chore(#3707): backfill changeset pr numbers Both fragments still carried the pr: 0 placeholder, which failed changeset-lint and docs-lint on PR 3903. The review had flagged the backfill as pending and I opened the PR without doing it. --------- Co-authored-by: sim <sim@local>
526 lines
25 KiB
TypeScript
526 lines
25 KiB
TypeScript
/**
|
||
* Core Utilities — Shared low-level utility primitives
|
||
*
|
||
* ADR-857 rollout phase 2c: extracted from core.cts (issue #877).
|
||
* Owns POSIX path normalization, sub-repo/subdirectory scanning,
|
||
* phase file stats, slug/one-liner/plan-id helpers, and time-ago.
|
||
* Behaviour is preserved byte-for-behaviour from the prior location;
|
||
* only the module boundary moved. core.cjs re-exports every public symbol
|
||
* here under its own `export =` object so existing consumers are unaffected.
|
||
*
|
||
* New imports should pull core-utils helpers from core-utils.cjs directly.
|
||
*
|
||
* Dependencies (leaf modules only — no core.cjs, no loadConfig):
|
||
* - node:fs / node:path (stdlib)
|
||
* - ./phase-id.cjs (comparePhaseNum, used by readSubdirectories)
|
||
* - ./planning-workspace.cjs (findContextMdIn, used by getPhaseFileStats)
|
||
*
|
||
* #3883 (ADR-3473 §8.3): two of this module's cyclic partners require
|
||
* generateSlugInternal, the canonical slug formula:
|
||
* - phase-id.cjs requires this module directly.
|
||
* - planning-workspace.cjs is a cyclic partner via a longer path:
|
||
* core-utils.cjs -> planning-workspace.cjs -> active-workstream-store.cjs
|
||
* -> workstream-name-policy.cjs -> core-utils.cjs.
|
||
* Both are genuine circular requires. They are safe ONLY because every side
|
||
* accesses the other's exports lazily, through the live module-namespace
|
||
* object (`phaseIdModule.foo(...)` / `planningWorkspace.foo(...)`) inside a
|
||
* function body, never via a top-level destructure — a top-level
|
||
* `const { foo } = require(...)` copies the binding at import time and would
|
||
* silently capture `undefined` whichever module loses the load-order race.
|
||
* This is an absolute rule with no exception in this file: every cyclic
|
||
* partner's export is accessed through its module-namespace object, never
|
||
* destructured at the top level.
|
||
*/
|
||
|
||
import fs from 'node:fs';
|
||
import path from 'node:path';
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||
import phaseIdModule = require('./phase-id.cjs');
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||
import planningWorkspace = require('./planning-workspace.cjs');
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
||
import shellCommandProjection = require('./shell-command-projection.cjs');
|
||
|
||
// ─── Line-ending normalization ─────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Normalize every line ending in `content` to a bare `\n`, ONCE — the shared
|
||
* seam every document-READ boundary in this codebase should route through
|
||
* (#3707-CR follow-up MAJOR).
|
||
*
|
||
* CommonMark treats a lone CR (no paired LF) as a line ending — such a
|
||
* document RENDERS as separate lines to a human reader — but a parser that
|
||
* splits/tokenizes/scans on `\n` alone treats a lone-CR-separated document as
|
||
* ONE unbroken line, hiding every row boundary in it. `src/uat.cts` originally
|
||
* carried this exact fix as a PRIVATE, unexported function applied inside two
|
||
* of its own parse functions (`parseUatItemsWithStats`, `parseCurrentTest`) —
|
||
* which is why two OTHER read sites in the same module (`cmdAuditUat`'s
|
||
* VERIFICATION and deferred-items.md ingresses) were missed: normalizing
|
||
* per-parser means every new parser must remember to call it. Promoted here,
|
||
* to the shared leaf module every document consumer can reach without a new
|
||
* dependency edge, so normalization can be applied at the READ boundary
|
||
* instead — every current and future parser fed from a boundary that calls
|
||
* this gets normalized text by construction.
|
||
*
|
||
* `/\r\n?/g` is deliberately ONE alternation, not two separate replaces: a
|
||
* two-pass `replace(/\r\n/g,'\n').replace(/\r/g,'\n')` is equivalent here
|
||
* because the first pass already consumes every CRLF pair before the second
|
||
* pass ever runs, but a single regex avoids relying on pass ORDER and matches
|
||
* greedily left-to-right in one scan, so a CRLF is always consumed as ONE
|
||
* unit (never left as a stray trailing `\r` after the `\n` half is matched
|
||
* first) and a lone CR — including one immediately followed by nothing, i.e.
|
||
* at EOF, or by another lone CR — is still replaced.
|
||
*
|
||
* This is deliberately NOT a length-preserving transform (CRLF, two UTF-16
|
||
* units, becomes LF, one), so any offsets a caller computes must be compared
|
||
* only against THIS normalized string, never against the original raw text.
|
||
*
|
||
* U+2028 LINE SEPARATOR / U+2029 PARAGRAPH SEPARATOR (#3078-CR) are
|
||
* DELIBERATELY NOT folded here, unlike `\r`/`\r\n`. Folding is unnecessary:
|
||
* `String.prototype.split('\n')` never treats U+2028/U+2029 as a delimiter,
|
||
* so an exotic separator can never manufacture a fake line start for a
|
||
* consumer that scans lines produced by `split('\n')`, rather than anchoring
|
||
* a multiline (`/m`) regex directly over unsplit text. Only the latter
|
||
* pattern is vulnerable to the ECMA-262 LineTerminator set including these
|
||
* two code points. This module performs no line-anchored matching of its
|
||
* own; a caller that scans lines should split first and match per-line
|
||
* rather than anchor `/m` over unsplit text — this comment makes no claim
|
||
* about whether any particular caller currently does so.
|
||
*/
|
||
function normalizeLineEndings(content: string): string {
|
||
return content.replace(/\r\n?/g, '\n');
|
||
}
|
||
|
||
// ─── Path helpers ────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Normalize a relative path to always use forward slashes (cross-platform).
|
||
* Delegates to the single separator seam in shell-command-projection so there is
|
||
* exactly one implementation of native→POSIX conversion across the codebase.
|
||
*/
|
||
function toPosixPath(p: string): string {
|
||
return shellCommandProjection.toPosixPath(p);
|
||
}
|
||
|
||
/**
|
||
* Scan immediate child directories for separate git repos.
|
||
* Returns a sorted array of directory names that have their own `.git`.
|
||
* Excludes hidden directories and node_modules.
|
||
*/
|
||
function detectSubRepos(cwd: string): string[] {
|
||
const results: string[] = [];
|
||
try {
|
||
const entries = fs.readdirSync(cwd, { withFileTypes: true });
|
||
for (const entry of entries) {
|
||
if (!entry.isDirectory()) continue;
|
||
if (entry.name.startsWith('.') || entry.name === 'node_modules') continue;
|
||
const gitPath = path.join(cwd, entry.name, '.git');
|
||
try {
|
||
if (fs.existsSync(gitPath)) {
|
||
results.push(entry.name);
|
||
}
|
||
} catch { /* ignore */ }
|
||
}
|
||
} catch { /* ignore */ }
|
||
return results.sort();
|
||
}
|
||
|
||
// ─── Summary body helpers ─────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Extract a one-liner from the summary body when it's not in frontmatter.
|
||
*/
|
||
function extractOneLinerFromBody(content: string | null | undefined): string | null {
|
||
if (!content) return null;
|
||
const normalized = content.replace(/\r\n/g, '\n').replace(/\r/g, '\n');
|
||
const body = normalized.replace(/^---\r?\n[\s\S]*?\r?\n---\r?\n*/, '');
|
||
// #3170: anchor to a summary-shaped heading (Summary / Overview /
|
||
// Accomplishments) so an incidental first heading (a rule list, task
|
||
// breakdown, deviation note) does not contribute its first bold run as the
|
||
// deliverable one-liner. Iterate headings in document order and extract from
|
||
// the first summary-shaped one that has a bold run; fall back to null (not the
|
||
// wrong text) when no such heading exists.
|
||
const headingRe = /^#+\s*([^\n]*)\n+\*\*([^*\n]+)\*\*([^\n]*)/gm;
|
||
let match: RegExpExecArray | null;
|
||
while ((match = headingRe.exec(body)) !== null) {
|
||
if (!/summary|overview|accomplish/i.test(match[1])) continue;
|
||
const boldInner = match[2].trim();
|
||
const afterBold = match[3];
|
||
if (/:\s*$/.test(boldInner)) {
|
||
const prose = afterBold.trim();
|
||
if (prose.length > 0) return prose;
|
||
} else if (boldInner.length > 0) {
|
||
return boldInner;
|
||
}
|
||
}
|
||
return null;
|
||
}
|
||
|
||
// ─── Misc utilities ───────────────────────────────────────────────────────────
|
||
|
||
function pathExistsInternal(cwd: string, targetPath: string): boolean {
|
||
const fullPath = path.isAbsolute(targetPath) ? targetPath : path.join(cwd, targetPath);
|
||
try {
|
||
fs.statSync(fullPath);
|
||
return true;
|
||
} catch {
|
||
return false;
|
||
}
|
||
}
|
||
|
||
/**
|
||
* #3883 (ADR-3473 §8.3 remediation): `maxLen` lets a caller state its own
|
||
* truncation contract instead of being forced into this function's
|
||
* historical 60-char cap. Some call sites truncated at 60 before the #3883
|
||
* consolidation (commands.cts:cmdGenerateSlug) and some never truncated at
|
||
* all (phase-id.cts toDir/getPhaseDirFromPhaseId, the init.cts/phase-locator
|
||
* phase_slug sites, workstream-name-policy.cts toWorkstreamSlug) — collapsing
|
||
* every caller onto a single hard-coded 60 introduced two identity
|
||
* collisions (distinct >60-char names/phase-slugs truncating to the same
|
||
* value) that did not exist pre-migration. `maxLen: 60` remains the default
|
||
* so untouched callers keep prior behavior; pass `null` for no truncation.
|
||
*/
|
||
function generateSlugInternal(text: string | null | undefined, maxLen: number | null = 60): string | null {
|
||
if (!text) return null;
|
||
// #2849: strip leading/trailing hyphens AFTER truncation, not only before.
|
||
// .substring(0, 60) can land on a separator, re-introducing a trailing hyphen
|
||
// the strip step exists to prevent. Truncation cannot add a leading hyphen, so
|
||
// running the full ^-+|-+$ pass last is equivalent for leading hyphens and
|
||
// fixes the trailing-hyphen-after-truncation case.
|
||
const collapsed = transliterateForSlug(text).replace(/[^a-z0-9]+/g, '-');
|
||
const truncated = maxLen === null ? collapsed : collapsed.substring(0, maxLen);
|
||
return truncated.replace(/^-+|-+$/g, '');
|
||
}
|
||
|
||
// ─── Transliteration (#2848) ─────────────────────────────────────────────────
|
||
//
|
||
// Non-Latin titles used to reduce to an empty slug: the `[^a-z0-9]+` strip
|
||
// removed every character of an all-Cyrillic title and the hyphen cleanup left
|
||
// "". Callers then created unnamed phase directories (`01-`) and empty
|
||
// `milestone_slug` init JSON. The fix transliterates Cyrillic to ASCII BEFORE
|
||
// the existing ASCII filter, so a non-Latin title yields a usable ASCII slug
|
||
// while Latin-script text (which hits zero map entries) is byte-for-byte
|
||
// unchanged — the negative control is satisfied by construction.
|
||
//
|
||
// Multi-letter mappings (ж→zh, ч→ch, ш→sh, щ→sch, ю→yu, я→ya) are applied as a
|
||
// single pass; soft/hard signs (ъ, ь) drop to nothing rather than a hyphen.
|
||
// Scope is Cyrillic (Russian + the reported Ukrainian/Belarusian extras
|
||
// і ї є ґ ў) per the issue's confirmed-working patch. CJK and other
|
||
// non-transliterated scripts keep the existing strip-to-ASCII behavior.
|
||
const CYRILLIC_TRANSLITERATION: Readonly<Record<string, string>> = {
|
||
// multi-letter first (longest-match-safe within a single pass via ordered keys)
|
||
а: 'a', б: 'b', в: 'v', г: 'g', д: 'd', е: 'e', ё: 'e', ж: 'zh',
|
||
з: 'z', и: 'i', й: 'y', к: 'k', л: 'l', м: 'm', н: 'n', о: 'o',
|
||
п: 'p', р: 'r', с: 's', т: 't', у: 'u', ф: 'f', х: 'h', ц: 'ts',
|
||
ч: 'ch', ш: 'sh', щ: 'sch', ъ: '', ы: 'y', ь: '', э: 'e', ю: 'yu',
|
||
я: 'ya',
|
||
// Ukrainian / Belarusian extras reported in #2848
|
||
є: 'ye', і: 'i', ї: 'yi', ґ: 'g', ў: 'u',
|
||
};
|
||
|
||
const CYRILLIC_TRANSLITERATION_KEYS = Object.keys(CYRILLIC_TRANSLITERATION);
|
||
|
||
/**
|
||
* Lowercase + transliterate Cyrillic characters to ASCII. The output still
|
||
* contains non-ASCII for scripts outside the map (CJK, etc.) — the caller's
|
||
* existing `[^a-z0-9]+` filter handles those. Latin-script input is returned
|
||
* lowercased with no other change.
|
||
*
|
||
* Shared by `generateSlugInternal` (core-utils) and `slugify` (gsd2-import) so
|
||
* the transliteration step is not duplicated across the two slug helpers (#2848
|
||
* explicitly requires both be fixed).
|
||
*/
|
||
function transliterateForSlug(text: string): string {
|
||
const lowered = text.toLowerCase();
|
||
let out = '';
|
||
for (const ch of lowered) {
|
||
out += CYRILLIC_TRANSLITERATION_KEYS.includes(ch)
|
||
? CYRILLIC_TRANSLITERATION[ch]
|
||
: ch;
|
||
}
|
||
return out;
|
||
}
|
||
|
||
// ─── Phase file helpers ──────────────────────────────────────────────────────
|
||
|
||
interface PhaseFileStats {
|
||
plans: string[];
|
||
summaries: string[];
|
||
hasResearch: boolean;
|
||
hasContext: boolean;
|
||
hasVerification: boolean;
|
||
hasReviews: boolean;
|
||
scope: string;
|
||
}
|
||
|
||
// Minimal shape this module needs from plan-scan.cjs's scanPhasePlans result.
|
||
interface PlanScanResultShape {
|
||
planFiles: string[];
|
||
summaryFiles: string[];
|
||
scope: string;
|
||
}
|
||
|
||
/**
|
||
* Read a phase directory and return counts/flags for common file types.
|
||
*
|
||
* #3183 (ADR-3180 Decision 2): `plans`/`summaries` are derived from the
|
||
* canonical `scanPhasePlans` rather than a local re-derivation, so this
|
||
* primitive can no longer diverge from the single owner of live-plan
|
||
* counting. `scanPhasePlans`
|
||
* lives in plan-scan.cjs, which itself imports `countMatchedSummaries` from
|
||
* THIS module — a top-level import here would be circular, so the require
|
||
* is deferred (lazy, inside the function body) to break the cycle at load
|
||
* time. This mirrors the lazy-require seam already used elsewhere in this
|
||
* repo (see src/audit-command-router.cts) for the same "module A needs
|
||
* module B which needs module A" shape.
|
||
*
|
||
* `hasResearch`/`hasContext`/`hasVerification`/`hasReviews` stay on the raw
|
||
* `readdirSync` listing — they are not plan-scan concerns.
|
||
*
|
||
* #3511 BLOCKER-2: the raw listing is scoped through `scopeToPhase` (keyed on
|
||
* `path.basename(phaseDir)`) before any of the four artifact predicates run,
|
||
* so a stray cross-phase file (e.g. `04-VERIFICATION.md` sitting inside phase
|
||
* 03's directory) cannot flip `hasResearch`/`hasContext`/`hasVerification`/
|
||
* `hasReviews` true for a phase it does not belong to — the same membership
|
||
* rule every other aggregate phase-directory scan (`uat.cts`, `audit.cts`,
|
||
* `phase.cts`, `state.cts`) already routes through. `hasContext` is scoped by
|
||
* passing the already-scoped array into `findContextMdIn` at this call site
|
||
* only — `findContextMdIn` itself is unchanged and its other 4 call sites
|
||
* (roadmap.cts, gap-checker.cts, init.cts) are unaffected.
|
||
*
|
||
* Degrades on an unreadable directory instead of throwing: empty arrays,
|
||
* every flag false, scope UNREADABLE (mirroring scanPhasePlans's own
|
||
* degrade path).
|
||
*/
|
||
function getPhaseFileStats(phaseDir: string): PhaseFileStats {
|
||
// eslint-disable-next-line @typescript-eslint/no-require-imports, @typescript-eslint/no-unsafe-assignment
|
||
const scanPhasePlans: (dir: string) => PlanScanResultShape = require('./plan-scan.cjs');
|
||
const scan = scanPhasePlans(phaseDir);
|
||
|
||
let files: string[];
|
||
try {
|
||
files = fs.readdirSync(phaseDir);
|
||
} catch {
|
||
return {
|
||
plans: scan.planFiles,
|
||
summaries: scan.summaryFiles,
|
||
hasResearch: false,
|
||
hasContext: false,
|
||
hasVerification: false,
|
||
hasReviews: false,
|
||
scope: scan.scope,
|
||
};
|
||
}
|
||
|
||
const scopedFiles = phaseIdModule.scopeToPhase(files, path.basename(phaseDir));
|
||
|
||
return {
|
||
plans: scan.planFiles,
|
||
summaries: scan.summaryFiles,
|
||
hasResearch: scopedFiles.some(f => f.endsWith('-RESEARCH.md') || f === 'RESEARCH.md'),
|
||
hasContext: planningWorkspace.findContextMdIn(scopedFiles) !== null,
|
||
hasVerification: scopedFiles.some(f => f.endsWith('-VERIFICATION.md') || f === 'VERIFICATION.md'),
|
||
hasReviews: scopedFiles.some(f => f.endsWith('-REVIEWS.md') || f === 'REVIEWS.md'),
|
||
scope: scan.scope,
|
||
};
|
||
}
|
||
|
||
/**
|
||
* Read immediate child directories from a path.
|
||
* Returns [] if the path doesn't exist or can't be read.
|
||
* Pass sort=true to apply comparePhaseNum ordering.
|
||
*/
|
||
function readSubdirectories(dirPath: string, sort = false): string[] {
|
||
try {
|
||
const entries = fs.readdirSync(dirPath, { withFileTypes: true });
|
||
const dirs = entries.filter(e => e.isDirectory()).map(e => e.name);
|
||
return sort ? dirs.sort((a, b) => phaseIdModule.comparePhaseNum(a, b)) : dirs;
|
||
} catch {
|
||
return [];
|
||
}
|
||
}
|
||
|
||
/**
|
||
* Format a Date as a fuzzy relative time string (e.g. "5 minutes ago").
|
||
*/
|
||
function timeAgo(date: Date): string {
|
||
const seconds = Math.floor((Date.now() - date.getTime()) / 1000);
|
||
if (seconds < 5) return 'just now';
|
||
if (seconds < 60) return `${seconds} seconds ago`;
|
||
const minutes = Math.floor(seconds / 60);
|
||
if (minutes === 1) return '1 minute ago';
|
||
if (minutes < 60) return `${minutes} minutes ago`;
|
||
const hours = Math.floor(minutes / 60);
|
||
if (hours === 1) return '1 hour ago';
|
||
if (hours < 24) return `${hours} hours ago`;
|
||
const days = Math.floor(hours / 24);
|
||
if (days === 1) return '1 day ago';
|
||
if (days < 30) return `${days} days ago`;
|
||
const months = Math.floor(days / 30);
|
||
if (months === 1) return '1 month ago';
|
||
if (months < 12) return `${months} months ago`;
|
||
const years = Math.floor(days / 365);
|
||
if (years === 1) return '1 year ago';
|
||
return `${years} years ago`;
|
||
}
|
||
|
||
// ─── Plan ID helpers ─────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Extract the canonical plan ID from a filename.
|
||
* Private to the core cluster — exported so core.cjs:searchPhaseInDir can
|
||
* import it from this leaf without circular dependency, but NOT re-exported
|
||
* from core.cjs's public `export =` block.
|
||
*/
|
||
function extractCanonicalPlanId(filename: string): string {
|
||
const base = filename.replace(/-PLAN\.md$/i, '').replace(/-SUMMARY\.md$/i, '').replace(/\.md$/i, '');
|
||
const parts = base.split('-').filter(Boolean);
|
||
// #2043: a phase/plan token component is either a zero-padded number (≥2 digits)
|
||
// or a single-digit-plus-letter id ("3A"); a *bare* single digit is a slug word,
|
||
// so "46-6-rs-…" is not paired into a "46-6" id while "3A-01" stays intact.
|
||
const tokenRe = /^(?:\d{2,}[A-Z]?|\d[A-Z])(?:\.\d+)*$/i;
|
||
// #2232: the PAIRED plan component is a zero-padded continuation segment
|
||
// (exactly 2 digits), so a ≥3-digit slug word (a year) is not paired into a
|
||
// bogus "14-2026" id. The leading phase component keeps tokenRe's unbounded
|
||
// \d{2,} — phase numbers ≥100 are legitimate; only continuations are capped.
|
||
const planTokenRe = new RegExp(
|
||
`^(?:${phaseIdModule.PHASE_CONTINUATION_SEGMENT_SOURCE}[A-Z]?|\\d[A-Z])(?:\\.\\d+)*$`,
|
||
'i',
|
||
);
|
||
const phaseIdx = parts.findIndex(p => tokenRe.test(p));
|
||
if (phaseIdx >= 0 && phaseIdx + 1 < parts.length && planTokenRe.test(parts[phaseIdx + 1])) {
|
||
return `${parts[phaseIdx]}-${parts[phaseIdx + 1]}`;
|
||
}
|
||
return base;
|
||
}
|
||
|
||
/**
|
||
* Count summaries that correspond to a real plan (#1988).
|
||
*
|
||
* A summary counts toward phase completion iff it pairs with an existing plan
|
||
* file. This excludes stray non-plan summaries — e.g. `30-FIX-CR02-SUMMARY.md`,
|
||
* `30-GAPCLOSURE-SUMMARY.md` — that inflate the raw `*-SUMMARY.md` count and
|
||
* silently flip a phase to Complete when plans are actually missing summaries.
|
||
*
|
||
* Pairing is layout-agnostic. For each plan, up to three candidate summary
|
||
* filenames are generated and any match suffices:
|
||
* 1. marker swap `PLAN`→`SUMMARY` on the basename — root padded
|
||
* (`30-01-PLAN.md`↔`30-01-SUMMARY.md`), nested (`PLAN-01.md`↔
|
||
* `SUMMARY-01.md`, incl. a `plans/` prefix), and bare (`PLAN.md`↔
|
||
* `SUMMARY.md`);
|
||
* 2. `<stem>-SUMMARY.md` — bare (`PLAN.md`↔`PLAN-SUMMARY.md`) and legacy
|
||
* (`14-PLAN-01.md`↔`14-PLAN-01-SUMMARY.md`);
|
||
* 3. extended `<n>-PLAN-<m>…`→`<n>-<m>-SUMMARY.md`
|
||
* (`3-PLAN-01-setup.md`↔`3-01-SUMMARY.md`).
|
||
* The swap is applied to the basename only so a lowercase `plans/` dir prefix
|
||
* isn't corrupted to `SUMMARYs/…`.
|
||
*/
|
||
function countMatchedSummaries(planFiles: string[], summaryFiles: string[]): number {
|
||
const summarySet = new Set(summaryFiles);
|
||
let matched = 0;
|
||
for (const plan of planFiles) {
|
||
if (summaryCandidates(plan).some((c) => summarySet.has(c))) matched++;
|
||
}
|
||
return matched;
|
||
}
|
||
|
||
/**
|
||
* The candidate `*-SUMMARY.md` filenames a single plan's completion record
|
||
* could take, per the three naming conventions documented above
|
||
* `countMatchedSummaries`. Extracted so `findUnsummarizedPlans` can reuse the
|
||
* exact same matching rule without duplicating it (a divergence between the
|
||
* count and the list would let a plan be counted as matched while still
|
||
* appearing in the unsummarized set, or vice versa).
|
||
*/
|
||
function summaryCandidates(plan: string): string[] {
|
||
const slashIdx = plan.lastIndexOf('/');
|
||
const dir = slashIdx >= 0 ? plan.slice(0, slashIdx + 1) : '';
|
||
const base = (dir ? plan.slice(dir.length) : plan).replace(/\.md$/i, '');
|
||
const candidates: string[] = [
|
||
dir + base.replace(/PLAN/i, 'SUMMARY') + '.md',
|
||
dir + base + '-SUMMARY.md',
|
||
];
|
||
const extended = base.match(/^(\d+)-PLAN-(\d+)/i);
|
||
if (extended) candidates.push(dir + extended[1] + '-' + extended[2] + '-SUMMARY.md');
|
||
// #3183: canonical-id form. Restores the coverage of the pre-migration
|
||
// bespoke I001 rule (verify.cts, pre-#3183, via validate.cjs's now-unused
|
||
// `canonicalPlanStem` — behaviourally identical to `extractCanonicalPlanId`,
|
||
// confirmed empirically), which matched a plan carrying a descriptive slug
|
||
// after its <phase>-<plan> id — e.g. `68-01-scaffolding-PLAN.md` — against
|
||
// a summary named only by the bare id — `68-01-SUMMARY.md`. None of the
|
||
// three candidates above produce that filename.
|
||
//
|
||
// Narrowed to the case `extractCanonicalPlanId` actually extracted an
|
||
// <id>-<id> pair (its result differs from the plan's own PLAN-stripped
|
||
// base). When no pair is found it falls back to returning that same base
|
||
// unchanged, which would otherwise push a redundant candidate identical to
|
||
// the `<stem>-SUMMARY.md` form above (e.g. `setup-PLAN.md` -> canonical
|
||
// 'setup' -> 'setup-SUMMARY.md', already candidate #2) rather than the
|
||
// original rule's actual behavior of matching only real id pairs.
|
||
//
|
||
// Collision, matching the original rule byte-for-behaviour: two plans that
|
||
// share the same <phase>-<plan> id but differ only in their descriptive
|
||
// slug (`68-01-alpha-PLAN.md` + `68-01-beta-PLAN.md`) both generate the
|
||
// SAME candidate `68-01-SUMMARY.md` and therefore BOTH read as summarized
|
||
// off one shared summary file. This is not a new regression: the
|
||
// pre-migration bespoke rule collapsed the same way (it populated one
|
||
// `summaryBases` Set keyed by canonical stem, so any plan whose canonical
|
||
// stem hit the set counted as matched, with no cardinality check against
|
||
// how many plans shared that stem).
|
||
const planStem = base.replace(/-PLAN$/i, '');
|
||
const canonicalId = extractCanonicalPlanId(base + '.md');
|
||
if (canonicalId !== planStem) candidates.push(dir + canonicalId + '-SUMMARY.md');
|
||
return candidates;
|
||
}
|
||
|
||
/**
|
||
* #2648: the plan files in `planFiles` that have NO matching completion record
|
||
* in `summaryFiles`, using the identical matching rule as `countMatchedSummaries`
|
||
* (so the count and the named list can never disagree). Callers that must NAME
|
||
* the missing plans — e.g. phase.complete's fail-closed coverage gate, which
|
||
* refuses completion when any non-retired plan lacks a SUMMARY — need the list,
|
||
* not just the count. `planFiles` is expected to be already superseded-filtered
|
||
* (the caller passes `scanPhasePlans(...).planFiles`, which drops
|
||
* `status: superseded` plans), so a deliberately-retired plan never appears
|
||
* here and never blocks completion.
|
||
*/
|
||
function findUnsummarizedPlans(planFiles: string[], summaryFiles: string[]): string[] {
|
||
const summarySet = new Set(summaryFiles);
|
||
return planFiles.filter((plan) => !summaryCandidates(plan).some((c) => summarySet.has(c)));
|
||
}
|
||
|
||
/**
|
||
* #3183: the mirror image of `findUnsummarizedPlans` — the summary files in
|
||
* `summaryFiles` that do NOT pair with ANY plan in `planFiles`, using the
|
||
* identical `summaryCandidates` matching rule as `countMatchedSummaries` /
|
||
* `findUnsummarizedPlans`. Callers that must name orphaned summaries (a
|
||
* stray non-plan summary, or a summary whose plan was renamed/removed) need
|
||
* this instead of a bespoke exact-suffix Set-diff, which cannot recognize
|
||
* the nested or extended naming forms `summaryCandidates` already handles —
|
||
* a divergence that produced false "orphan summary" warnings.
|
||
*/
|
||
function findOrphanSummaries(planFiles: string[], summaryFiles: string[]): string[] {
|
||
const claimed = new Set<string>();
|
||
for (const plan of planFiles) {
|
||
for (const candidate of summaryCandidates(plan)) claimed.add(candidate);
|
||
}
|
||
return summaryFiles.filter((s) => !claimed.has(s));
|
||
}
|
||
|
||
export = {
|
||
toPosixPath,
|
||
normalizeLineEndings,
|
||
detectSubRepos,
|
||
extractOneLinerFromBody,
|
||
pathExistsInternal,
|
||
generateSlugInternal,
|
||
transliterateForSlug,
|
||
getPhaseFileStats,
|
||
readSubdirectories,
|
||
timeAgo,
|
||
extractCanonicalPlanId,
|
||
countMatchedSummaries,
|
||
findUnsummarizedPlans,
|
||
findOrphanSummaries,
|
||
};
|