* chore(#2143): prohibition-with-teeth + migrate remaining table sites — Phase 4 Phase 4 of epic #2143 (ADR-2143 §7). Completes the markdown table/mutation consolidation by (a) giving the ad-hoc-parsing prohibition teeth and (b) migrating the last ad-hoc table sites onto the shared seam. - src/markdown-table.cts: new formatting-preserving `updateTableCell` primitive (self-contained, ragged-row-tolerant header/delimiter/cell-range scan; splices only the target cell's raw span, preserving all other bytes incl. padding/CRLF; no-op-preserves-padding when a transformer returns the current value). Exports splitTableRow/isDelimiterRow/findTableStartOffset for tolerant reuse. - eslint-rules/no-adhoc-markdown-parsing.cjs: TABLE-REGEX detector extended to `new RegExp(<literal|static-template>)`; new `.replace()`-mutation detector for roadmap/state/content receivers with a table/section-shaped pattern. - scripts/lint-table-schema-drift.cjs (wired into lint:ci): fails if a TABLE_SCHEMA header drifts from its authored table; tests import its logic (single source). - Migrated onto the seam (behaviour-preserving vs pre-Phase-4 HEAD, verified byte-diff old-vs-new): roadmap.cts cmdRoadmapUpdatePlanProgress, phase.cts cmdPhaseComplete + traceability, milestone.cts cmdRequirementsMarkComplete, uat.cts read path, state.cts metrics/decisions/By-Phase. - Incidental correctness gains from the migration: a decoy table can no longer swallow a phase-progress update (## Progress scoping); a ragged neighbouring row no longer silently aborts an edit; completing integer phase N no longer touches a decimal sub-phase N.x row; record-metric no longer drops trailing section content or duplicates the ## Performance Metrics section. - Kept justified allow-adhoc-markdown markers only where genuinely not a table (security.cts <|role|> token) or a loose non-GFM section (uat human-verify). Two orthogonal isolated reviews (correctness/adversarial + security) passed; correctness found 4 behaviour regressions in the first migration pass, all fixed and re-verified byte-identical-or-better vs OLD. Surfaced for maintainer (pre-existing, ambiguous domain logic, NOT changed here): templates/state.md places a By-Phase table under ## Performance Metrics while cmdStateRecordMetric assumes a Plan|Duration|Tasks|Files table. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): match traceability row by first-cell value, not Requirement header Phase 4's migration matched the REQUIREMENTS.md traceability row by a column literally named `Requirement` (`row['Requirement']`), but real tables head that column `REQ-ID`. The by-name lookup found nothing, so `phase complete` and `requirements mark-complete` left the Status cell `Pending` (regressed #2769 / #2203, caught by gsd-test — 8 failures, both node 22/24). - src/phase.cts, src/milestone.cts: match the row by its FIRST cell's value (the requirement-ID column) regardless of that column's HEADER name, via `Object.values(row)[0]` (updateTableCell builds the record in header order). This mirrors OLD's first-cell `\|\s*<id>\s*\|` anchor, restoring header-name independence while keeping the seam. - src/milestone.cts hasTable: broadened from `Requirement`-only to also recognize `Requirement ID` / `REQ-ID` / `REQ ID` headers, kept in sync with the now-positional rowMatch/hasRow so a REQ-ID-headed table participates in the ADR-2143 §6 write-set and the #2140 table_unmatched drift check (it was silently omitted before — a checkbox-only partial reconcile against a REQ-ID table could report as fully reconciled). The `Requirement`-headed path is byte-identical to OLD. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2143): replace stale structural milestone guards with behavioural suite The `milestone.cjs regex global state fix` block was a source-structure guard (allow-test-rule: structural-regression-guard) — it readFileSync'd the compiled milestone.cjs and asserted removed regex idioms (`tablePattern.test`, `afterTable !== reqContent`, `doneTable = new RegExp(...)`). Phase 4's migration deleted those regexes (table update is now updateTableCell), making the assertions obsolete. Per the Test Cleanup rule, replace them in-PR with a behavioural suite driving the compiled CLI: - multi-ID mark-complete flips all IDs (guards the lastIndex/global-state class), - Pending->Complete flip under both `REQ-ID` and `Requirement` headers (#2769), - idempotent already_complete detection with no corruption, - REQ-ID-headed table participates in write_set (traceability entry, applied), - REQ-ID-headed table trips #2140 table_unmatched drift on a missing row. Pruned the now-nonexistent structural-regression-guard entry from the lint-allow-test-rule-refs allowlist (the source-text-is-the-product entry for the same file remains valid). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): backfill PR number 2253 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): record-metric targets its own metrics table, not By-Phase velocity `state record-metric` appended its per-plan row (`| Phase 1 P1 | 5min | 3 tasks | 4 files |`) into the FIRST table under `## Performance Metrics` — which on a real template-derived STATE.md is the By-Phase velocity table `| Phase | Plans | Total | Avg/Plan |`, polluting it on EVERY plan completion (execute-plan.md:414 is a per-plan call). The command's own metrics table is `| Plan | Duration | Tasks | Files |`, which the template does not ship, so the row never reached it; the scaffold branch also emitted a wrong `| Phase | Plan | Duration | Notes |` header matching neither the row nor the canonical table. Pre-existing (predates Phase 4); surfaced while migrating this site and fixed here per no-defer, on the user's explicit go-ahead. - src/state.cts cmdStateRecordMetric: locate the metrics table by its own header shape (`Plan|Duration|Tasks|Files`, via splitTableRow/isDelimiterRow) rather than "first table in the section". When the section exists but has no metrics table (only the By-Phase table), self-heal by appending a fresh **Per-Plan Metrics:** table to the END of the section body — By-Phase table, Recent Trend and footer preserved verbatim, no duplicate `## Performance Metrics` heading, created stays false. Absent-section scaffold header corrected to the canonical `| Plan | Duration | Tasks | Files |`. Ragged-tolerance + None-yet preserved. - Not touching templates/state.md (golden-install-parity hashed) — record-metric self-creates the table on first use instead. Failing-first regression test (tests/state.test.cjs) demonstrates the By-Phase pollution on the pre-fix build, then green after. Verified: no pollution, self- heal idempotency, both-tables isolation, content/heading preservation, flags, None-yet, corrected scaffold header (23-check adversarial harness + all existing record-metric scenarios). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#2143): deleteSection seam primitive (level-bounded whole-section removal) ADR-2143 §4 shipped withSection/collectSection (replace a section BODY) but no way to DELETE a section (heading + body). Phase 4 suppressed the phase-remove section delete instead of building it. deleteSection(content, predicate, opts) locates the section via the collectSection machinery and splices out from the heading's start offset to the next same-or-higher-level heading — so a level-3 `### Phase N` delete stops at a following level-2 `## Progress`, never past it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase remove no longer deletes ## Progress on last-phase removal updateRoadmapAfterPhaseRemoval deleted a `### Phase N` detail section with a greedy raw regex whose lazy scan, on the LAST phase, ran to EOF and destroyed the following `## Progress` heading and its entire tracking table — silent data loss, uncovered by tests (removal tests only exercised a middle phase). Migrated onto the new deleteSection seam (level-bounded, stops at `## Progress`); dropped the allow-adhoc-markdown SECTION-DELETION suppression. Failing-first regression (tests/phase.test.cjs) removes the LAST phase and asserts the ## Progress heading + table survive; middle-phase removal is byte-identical. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#2143): deleteTableRow seam primitive (row removal, ragged-tolerant) Sibling of updateTableCell: locates the first GFM table, matches a DATA row by predicate (ragged-tolerant record build, header order), and splices out that row's whole line preserving every other byte. Returns {ok:false,reason} on no table / no match. Enables migrating the phase-remove Progress-table row delete off its ad-hoc regex (ADR-2143 §7 — the "future row-delete seam" Phase 4 punted). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase remove deletes the Progress row via deleteTableRow The Progress-table row delete used a whole-document regex with two defects: (a) `\.?\s` required whitespace after the phase number, so a COMPACT row `|2|Beta|` was never deleted (stale row left behind); (b) unscoped — it could strike a row in a different table (e.g. an earlier `| Phase | Requirements |` table). Migrated onto deleteTableRow, scoped to the `## Progress` section (mirrors deriveProgressFromRoadmap), matching the row by first-cell phase number (integer zero-pad-insensitive; decimal exact; removing `2` never touches `2.5`). Both allow-adhoc-markdown suppressions removed. New behavioural tests: compact unpadded row deleted; padded byte-parity on the surviving rows (their ordinal correctly renumbers via the pre-existing renumber block). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): deleteTableRow leaves no dangling newline on last EOL-less row Deleting the final row of a table with no trailing EOL sliced from the row's start to end-of-string, stranding the newline that terminated the previous line. Back rowStart over the preceding \r?\n in that branch so the table ends cleanly. (Caught by the primitive's own unit test on gsd-test; local scenario checks missed the no-trailing-EOL edge.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): migrate read-only section-collects onto collectSection Six hand-rolled `## Section` read-extract regexes replaced by the collectSection seam (behaviour-preserving; extracted bodies feed the same downstream parsers): state.cts matchSessionSection (## Session / ## Session Continuity) + ## Blockers, smart-entry.cts ## Blockers, audit.cts ## Current Focus + ## Open Questions. Removes 6 allow-adhoc-markdown "pending #1372" suppressions. Incidental fix: the old Session regex `## Session[ \t]*\n` silently failed on a CRLF `## Session\r\n` heading (Windows STATE.md), nulling all session fields; collectSection is CRLF-safe, so session state now resolves on Windows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): fence-safe state-transition section writes + dedup stripFrontmatter - milestoneCompleteCore's `## Current Position` and `## Operator Next Steps` section resets used fence-blind raw regexes that a fenced `##` inside the body could truncate/mis-target (#2130/#2067/#2080 class). Migrated onto a fence-aware tokenizeHeadings-based helper (resetSectionVerbatim) that is byte-identical to the old output on the canonical path (9/9 fixtures) and correctly ignores a fenced fake heading (proven robustness gain). - mutateCurrentPositionFirstTime: hand-rolled locate+splice → collectSection + replaceSection (byte-parity). - stripFrontmatter was inlined byte-identically in state.cts AND state-transition.cts; hoisted the single canonical copy into frontmatter.cts (both call sites now import it) + unit tests — eliminates the divergence risk per CLAUDE.md "Generative Fix Divergence". Removes 3 allow-adhoc-markdown / #1372 markers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): name-address By-Phase sum + uat parse, eslint recall hole, catches - state.cts By-Phase "Total plans completed" sum: positional 2nd-cell regex → name-addressed splitTableRow read (correct on a reordered header, where the old code silently summed the wrong column). Marker removed. - uat.cts parseVerificationItems: loose pipe regex → splitTableRow within the existing table/numbered/bullet union scan (item list byte-identical; does NOT reintroduce the reverted strict-parseMarkdownTable item-drop). Marker removed. - eslint no-adhoc-markdown-parsing: close the `new RegExp(identifier)` recall hole — resolve a const-declared table-shaped regex identifier (mirrors the .replace() detector) + RuleTester cases; param/call args stay out (boundary). - commands.cts: delete a lying comment that claimed the scaffold date "stays on raw UTC / deferred" — #2136 already moved it to realClock.localToday(). - Empty catches (classified, not blind-swept): removed 4 dead try/catch; fixed 3 error-hiding (phase-insert decimal-dir I/O collision now fails loud; phase-remove rename partial-failure surfaced; milestone-archive true count via finally); left best-effort swallows with justification comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): extractFencedBlock seam + migrate api-coverage named fence parseCoverageMatrix extracted its ```coverage fenced block with an ad-hoc regex (the last real allow-adhoc-markdown suppression). Added extractFencedBlock to the markdown-sectionizer seam (reuses stripFencedCode's CommonMark fence engine — info-string match, ~~~/backtick, nesting, indent) and migrated onto it; byte- parity on the parsed CoverageMatrix across 8 fixtures. Only security.cts:367 (a genuine `<|role|>` protocol-token false-positive, not a GFM table) remains marked in src/ — the "prohibition with teeth" goal (nothing grandfathered but a true FP) is met. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): By-Phase row insert is name-addressed (insertTableRow seam) updatePerformanceMetricsSection's INSERT-new-row branch located the By-Phase table with a canonical-column-order-only regex + a hardcoded positional row literal, so on a reordered header it silently inserted nothing — inconsistent with the now name-addressed UPDATE and SUM halves of the same function. Added insertTableRow (markdown-table seam sibling of updateTableCell/deleteTableRow: name-addressed, header-order-agnostic, EOL-preserving) and migrated the branch onto it, mapping By-Phase values by column NAME. Canonical-order output is byte-identical; a reordered header now inserts a correctly-mapped row; a pre-existing CRLF mixed-EOL splice glitch is incidentally fixed. Retired the now-dead byPhaseTablePattern const. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase-list checkbox flip via updateBullet seam Added updateBullet (markdown-sectionizer): a fence-aware, offset-tracked single-bullet write primitive (GFM 1–4-space marker tolerance) — the write counterpart to read-only iterateBullets. Migrated mutateMilestonePhase's phase-list checkbox flip (`- [ ] Phase N …` → `- [x] … (completed <date>)`) off its whole-slice regex onto it, same milestone-slice scope + clock seam. Byte-identical across simple / idempotent / metachar-title / double-space / CRLF scenarios. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): scope the Progress-ordinal renumber to ## Progress via seam phase remove's integer-renumber decremented Progress-table phase ordinals with a whole-document `content.replace(/(\|\s*)(\d+)(\.\s)/g, …)` — unscoped, so it also rewrote any `| N. …` cell in an unrelated/decoy table (same class as the batch-2 row-delete scoping bug). Migrated onto updateTableCell, scoped to the ## Progress section, decrementing each affected row's leading phase ordinal by column name. Byte-identical on canonical Progress tables + multi-row + decimal-sibling cases; a decoy `| 3. … |` row before ## Progress is now correctly left untouched. The sibling heading / checkbox-bullet / PLAN.md-filename / Depends-on-prose renumbers are not GFM-table mutations (outside ADR-2143's table/section mandate) — left as-is. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): review fixes — scope traceability write, restore Current Position H3-stop Adversarial review of the remediation (BLOCK verdict) — all 9 findings fixed: - F1 (BLOCKER): requirements mark-complete / phase complete flipped the checkbox but NOT the traceability row on the shipped template, because updateTableCell bound to the FIRST table (## Out of Scope, no Status column) instead of the ## Traceability table — the #2140 silent-divergence class, re-introduced by the seam migration and missed by tests (fixtures had Traceability first). Scoped the write + hasRow probe to the ## Traceability section slice (updateTraceability Cell helper) in milestone.cts + phase.cts. Failing-first tests on the Out-of-Scope-before-Traceability layout; the #2769 first-cell match preserved. - F2 (MAJOR): mutateCurrentPositionFirstTime restored to locateCurrentPosition (STOP_H2_PLUS) — collectSection's default H2-stop swallowed a level-3 subsection and the field regexes clobbered it (#2130 class). - F3/F8: Progress-ordinal renumber re-escapes via escapeCell + keys padding recovery by row index (was de-escaping `\|` and losing padding on dup values). - F4: insertTableRow escapes cell values internally. - F5: updateBullet accepts a tab after the marker (`[ \t]{1,4}`). - F7: resetSectionVerbatim consumes CRLF blank lines (byte-parity on CRLF). - F6/F9: corrected two misleading comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): data-loss + CRLF-session user-facing fixes (#2253) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2143): de-flake the G10 windsurf ReDoS-guard wall-clock assertion The G10 test asserted `elapsedMs < 1000` for a 200k-char payload — a wall-clock assertion (CLAUDE.md: never assert on wall-clock time) that flaked on a loaded node24 bench at ~1.1s. It was redundant: runHook's spawnSync `timeout: 10000` already SIGKILLs a catastrophic-backtracking hook, so the exit-0 assertion is the real ReDoS guard. Removed the timing assertion; kept exit-0 + documented the subprocess-timeout mechanism. Surfaced (not caused) by this branch's gsd-test runs loading the bench; unrelated to the markdown-parsing changes but fixed in place per the no-flaky-tests rule. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
1016 lines
43 KiB
TypeScript
1016 lines
43 KiB
TypeScript
/**
|
||
* Markdown Sectionizer — canonical markdown-structure parsing seam
|
||
*
|
||
* Pure functions, Node built-ins only (no external deps). String-in → value-out, no I/O.
|
||
* Promoted from `uat-predicate.cts` `_stripFencedBlocks` (CommonMark-correct state machine)
|
||
* and extended with heading tokenisation, section collection, and bullet iteration.
|
||
*
|
||
* ADR-1372 — T0 foundational seam. Migration tiers T1–T7 progressively adopt this seam.
|
||
*
|
||
* ADR-457 build-at-publish: compiled by tsc to gsd-core/bin/lib/markdown-sectionizer.cjs.
|
||
*/
|
||
|
||
// ─── Types ────────────────────────────────────────────────────────────────────
|
||
|
||
/** Result of stripping fenced code blocks from markdown content. */
|
||
export interface StripFencedResult {
|
||
/** Content with all fenced code blocks removed (delimiters and body lines). */
|
||
text: string;
|
||
/**
|
||
* True when the input contained an unterminated fence (EOF inside a fence).
|
||
* Callers that wish to signal malformed input to the user should inspect this.
|
||
*/
|
||
unterminatedFence: boolean;
|
||
}
|
||
|
||
/** An ATX heading extracted by `tokenizeHeadings`. */
|
||
export interface HeadingToken {
|
||
/** Heading depth: 1 = `#`, 2 = `##`, 3 = `###`, etc. */
|
||
level: number;
|
||
/** Heading text with surrounding whitespace trimmed. */
|
||
text: string;
|
||
/** 1-based line number of the heading in the original content. */
|
||
line: number;
|
||
/** Character (string-index) offset of the `#` character in the original content string. */
|
||
offset: number;
|
||
}
|
||
|
||
/** A collected markdown section (heading + body). */
|
||
export interface Section {
|
||
/** The heading that opened this section. */
|
||
heading: HeadingToken;
|
||
/** All lines between this heading and the next stop, joined by `\n`. */
|
||
body: string;
|
||
/**
|
||
* Character (string-index) offset in the ORIGINAL content string where the
|
||
* section body begins (first character after the heading line's trailing newline).
|
||
* Populated by `collectSections` and `collectSection`.
|
||
* Used by `replaceSection` for a clean pure splice.
|
||
*
|
||
* INVARIANT: `content.slice(bodyStart, bodyEnd) === body` for every Section
|
||
* returned by `collectSection` and `collectSections`.
|
||
*/
|
||
bodyStart: number;
|
||
/**
|
||
* Character (string-index) offset in the ORIGINAL content string where the
|
||
* section body ends (exclusive). Because `body` is `trimEnd()`-ed, this equals
|
||
* `bodyStart + body.length` — NOT the start of the next heading line.
|
||
*
|
||
* INVARIANT: `content.slice(bodyStart, bodyEnd) === body`.
|
||
* This guarantees `replaceSection(content, section, section.body) === content`.
|
||
*/
|
||
bodyEnd: number;
|
||
}
|
||
|
||
/** Options shared by `collectSection` and `withSection` (see `collectSection`'s doc comment). */
|
||
export interface CollectSectionOptions {
|
||
levelBounded?: boolean;
|
||
stopAtLevel?: number;
|
||
stripFences?: boolean;
|
||
}
|
||
|
||
/** Recognised bullet markers. */
|
||
export type BulletMarker = 'dash' | 'checkbox-unchecked' | 'checkbox-checked' | 'numbered';
|
||
|
||
/** A single bullet item from `iterateBullets`. */
|
||
export interface BulletItem {
|
||
/** Which marker shape was recognised. */
|
||
marker: BulletMarker;
|
||
/** Full bullet text including all indented continuation lines, whitespace-trimmed. */
|
||
text: string;
|
||
/** Raw indentation prefix of the opening bullet line. */
|
||
indent: string;
|
||
/** Checkbox state — `true` for `[x]`, `false` for `[ ]`, `null` for non-checkbox. */
|
||
checked: boolean | null;
|
||
}
|
||
|
||
// ─── Internal types ───────────────────────────────────────────────────────────
|
||
|
||
interface FenceState {
|
||
char: '`' | '~';
|
||
len: number;
|
||
}
|
||
|
||
// ─── stripFencedCode ──────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* CommonMark-correct fenced-code-block stripper.
|
||
*
|
||
* Ported from `uat-predicate.cts` `_stripFencedBlocks` — the reference
|
||
* implementation for the repo. DO NOT modify `uat-predicate.cts` (its
|
||
* migration is T5); this is a tracked duplication until T5 lands.
|
||
*
|
||
* Rules:
|
||
* - Opening delimiter: a line whose non-indent portion begins with ≥3 backticks
|
||
* or tildes (≤3 leading spaces tolerated per CommonMark §4.5).
|
||
* - Closing delimiter: same character, run length ≥ opening, no trailing
|
||
* non-whitespace text.
|
||
* - A tilde fence inside a backtick fence (or vice versa) is fence *content*,
|
||
* not a closing delimiter — delimiter char must match.
|
||
* - Both delimiter lines and all content lines are dropped from the output.
|
||
* - CRLF-safe: trailing `\r` is stripped before delimiter matching; the kept
|
||
* non-fence lines are returned as-is (including any `\r`).
|
||
* - `unterminatedFence` signals EOF inside an open fence.
|
||
*/
|
||
export function stripFencedCode(content: string): StripFencedResult {
|
||
if (typeof content !== 'string') {
|
||
return { text: '', unterminatedFence: false };
|
||
}
|
||
const lines = content.split('\n');
|
||
const kept: string[] = [];
|
||
let openFence: FenceState | null = null;
|
||
|
||
// Matches: optional indent (≤3 spaces per CommonMark), fence run, optional info string
|
||
const delimRe = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||
|
||
for (const rawLine of lines) {
|
||
// Strip trailing \r for delimiter matching (CRLF safety)
|
||
const line = rawLine.replace(/\r$/, '');
|
||
const m = delimRe.exec(line);
|
||
if (m) {
|
||
const char = m[2][0] as '`' | '~';
|
||
const len = m[2].length;
|
||
const trailing = m[3];
|
||
if (openFence === null) {
|
||
// CommonMark §4.5: backtick fence info string must not contain a backtick.
|
||
// If it does, this line is NOT a valid fence opener (treat as ordinary content).
|
||
if (char === '`' && trailing.includes('`')) {
|
||
kept.push(rawLine);
|
||
continue;
|
||
}
|
||
// Opening delimiter — record fence state, drop this line
|
||
openFence = { char, len };
|
||
} else if (char === openFence.char && len >= openFence.len && /^\s*$/.test(trailing)) {
|
||
// Closing delimiter (same char, sufficient length, no trailing content) — close and drop
|
||
openFence = null;
|
||
}
|
||
// else: mismatched delimiter inside fence — treat as content, still drop (it's a fence line)
|
||
continue; // all delimiter lines are dropped
|
||
}
|
||
|
||
if (openFence === null) {
|
||
kept.push(rawLine); // non-fence content: keep as-is (preserve original \r if any)
|
||
}
|
||
// Lines inside a fence are silently dropped
|
||
}
|
||
|
||
return { text: kept.join('\n'), unterminatedFence: openFence !== null };
|
||
}
|
||
|
||
// ─── extractFencedBlock ───────────────────────────────────────────────────────
|
||
|
||
/** A fenced code block located by `scanFencedBlocks`: line-index span + info string. */
|
||
interface FencedBlockRecord {
|
||
/** Fence delimiter character (`` ` `` or `~`). */
|
||
char: '`' | '~';
|
||
/** Fence delimiter run length (≥3). */
|
||
len: number;
|
||
/** Opening line's trailing text (untrimmed) — the CommonMark "info string". */
|
||
infoString: string;
|
||
/** 0-based index (into the `lines` array) of the OPENING delimiter line. */
|
||
openLineIdx: number;
|
||
/**
|
||
* 0-based index of the CLOSING delimiter line, or `-1` when the fence is
|
||
* unterminated (EOF reached while still open — mirrors `stripFencedCode`'s
|
||
* `unterminatedFence` signal).
|
||
*/
|
||
closeLineIdx: number;
|
||
}
|
||
|
||
/**
|
||
* Shared low-level fence-scanning engine. Walks `lines` and returns every
|
||
* fenced block found, applying the EXACT SAME CommonMark delimiter rules as
|
||
* `stripFencedCode` (≥3 backticks/tildes, ≤3-space indent tolerance, a closer
|
||
* must be the same delimiter char with run length ≥ the opener and no
|
||
* trailing non-whitespace text; a mismatched delimiter char — or a same-char
|
||
* run that is too short or carries trailing text — encountered while a fence
|
||
* is already open is fence CONTENT, not a new open/close event). This is the
|
||
* "engine" `extractFencedBlock` reuses instead of an ad-hoc regex, so a
|
||
* different-info-string fence, a fence nested/indented inside another fence,
|
||
* and a `~~~` fence are all classified exactly as `stripFencedCode` would.
|
||
*
|
||
* Tracked duplication (same status as `tokenizeHeadings`'s copy, see its
|
||
* comment above): this is a second independent copy of the fence state
|
||
* machine, pending a T-tier consolidation.
|
||
*/
|
||
function scanFencedBlocks(lines: string[]): FencedBlockRecord[] {
|
||
const delimRe = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||
const blocks: FencedBlockRecord[] = [];
|
||
let open: { char: '`' | '~'; len: number; infoString: string; openLineIdx: number } | null = null;
|
||
|
||
for (let i = 0; i < lines.length; i++) {
|
||
const line = lines[i].replace(/\r$/, '');
|
||
const m = delimRe.exec(line);
|
||
if (!m) continue;
|
||
|
||
const char = m[2][0] as '`' | '~';
|
||
const len = m[2].length;
|
||
const trailing = m[3];
|
||
|
||
if (open === null) {
|
||
// CommonMark §4.5: backtick fence info string must not contain a backtick.
|
||
if (char === '`' && trailing.includes('`')) continue; // not a valid opener — ordinary content
|
||
open = { char, len, infoString: trailing.trim(), openLineIdx: i };
|
||
} else if (char === open.char && len >= open.len && /^\s*$/.test(trailing)) {
|
||
blocks.push({
|
||
char: open.char,
|
||
len: open.len,
|
||
infoString: open.infoString,
|
||
openLineIdx: open.openLineIdx,
|
||
closeLineIdx: i,
|
||
});
|
||
open = null;
|
||
}
|
||
// else: mismatched/insufficient delimiter while a fence is open — content, not a boundary.
|
||
}
|
||
|
||
if (open !== null) {
|
||
blocks.push({
|
||
char: open.char,
|
||
len: open.len,
|
||
infoString: open.infoString,
|
||
openLineIdx: open.openLineIdx,
|
||
closeLineIdx: -1,
|
||
});
|
||
}
|
||
|
||
return blocks;
|
||
}
|
||
|
||
/**
|
||
* Return the INNER text (the lines between the delimiters, joined by `\n`) of
|
||
* the FIRST fenced code block whose opening info string — trimmed,
|
||
* case-insensitive — equals `infoString`. Returns `null` when no such block
|
||
* exists, including when the only matching-name fence is left unterminated
|
||
* (EOF inside the fence — there is no well-defined inner span to return,
|
||
* matching a non-greedy `\n```-anchored` regex's behaviour of also failing to
|
||
* match an unclosed fence).
|
||
*
|
||
* Built on `scanFencedBlocks`, the same CommonMark fence-tracking engine
|
||
* `stripFencedCode` uses — so a fence of a DIFFERENT info string, a fence
|
||
* nested/indented inside another fence, and a `~~~` fence are all handled
|
||
* exactly as `stripFencedCode` would classify them; this is not a fresh
|
||
* ad-hoc regex.
|
||
*
|
||
* Migrated from `api-coverage.cts`'s bespoke
|
||
* `` /```coverage\s*\n([\s\S]*?)\n```/i `` (ADR-1372 tier migration, #2143 audit).
|
||
*/
|
||
export function extractFencedBlock(content: string, infoString: string): string | null {
|
||
if (typeof content !== 'string' || content.length === 0) return null;
|
||
if (typeof infoString !== 'string') return null;
|
||
|
||
const target = infoString.trim().toLowerCase();
|
||
const lines = content.split('\n');
|
||
const blocks = scanFencedBlocks(lines);
|
||
|
||
for (const block of blocks) {
|
||
if (block.closeLineIdx === -1) continue; // unterminated — no well-defined inner span
|
||
if (block.infoString.trim().toLowerCase() !== target) continue;
|
||
return lines.slice(block.openLineIdx + 1, block.closeLineIdx).join('\n');
|
||
}
|
||
|
||
return null;
|
||
}
|
||
|
||
// ─── tokenizeHeadings ─────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Extract all ATX headings from `content` in document order.
|
||
*
|
||
* Only headings OUTSIDE fenced code blocks are returned — `stripFencedCode` is
|
||
* applied first so that a `## heading` inside a ``` fence is not tokenised.
|
||
*
|
||
* Each token records `{ level, text, line, offset }` where `offset` is relative
|
||
* to the ORIGINAL `content` (before fence-stripping), enabling callers to use
|
||
* `collectSection` on the original string.
|
||
*/
|
||
export function tokenizeHeadings(content: string): HeadingToken[] {
|
||
if (typeof content !== 'string' || content.length === 0) return [];
|
||
|
||
// Strip fences first so headings inside code blocks are ignored.
|
||
// We need the original line positions, so we map stripped-text line numbers
|
||
// back to original by tracking which original lines survived stripping.
|
||
const originalLines = content.split('\n');
|
||
const tokens: HeadingToken[] = [];
|
||
|
||
// We re-run the fence state machine to know which lines are "kept", so we
|
||
// can map line index in original to whether it survived.
|
||
const delimRe = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||
let openFence: FenceState | null = null;
|
||
|
||
// Accumulate character offset as we iterate lines
|
||
let charOffset = 0;
|
||
|
||
for (let i = 0; i < originalLines.length; i++) {
|
||
const rawLine = originalLines[i];
|
||
const line = rawLine.replace(/\r$/, '');
|
||
|
||
const dm = delimRe.exec(line);
|
||
if (dm) {
|
||
const char = dm[2][0] as '`' | '~';
|
||
const len = dm[2].length;
|
||
const trailing = dm[3];
|
||
if (openFence === null) {
|
||
// CommonMark §4.5: backtick fence info string must not contain a backtick.
|
||
if (char === '`' && trailing.includes('`')) {
|
||
// Not a valid fence opener — check for heading on this line (will fall through)
|
||
} else {
|
||
openFence = { char, len };
|
||
charOffset += rawLine.length + 1;
|
||
continue;
|
||
}
|
||
} else if (char === openFence.char && len >= openFence.len && /^\s*$/.test(trailing)) {
|
||
openFence = null;
|
||
charOffset += rawLine.length + 1;
|
||
continue;
|
||
} else {
|
||
// Mismatched/invalid delimiter inside fence — treat as content (still inside fence), skip heading check
|
||
charOffset += rawLine.length + 1;
|
||
continue;
|
||
}
|
||
}
|
||
|
||
if (openFence === null) {
|
||
// This line is outside any fence — check for ATX heading.
|
||
// CommonMark: ≤3 leading spaces, then 1–6 `#`, then either EOF (empty heading)
|
||
// or at least one space/tab followed by optional text, with optional closing `#` sequence.
|
||
const headingMatch = /^( {0,3})(#{1,6})([ \t]+.*|[ \t]*)?$/.exec(line);
|
||
if (headingMatch) {
|
||
const hashes = headingMatch[2];
|
||
const rest = headingMatch[3] ?? '';
|
||
// Strip optional closing `#` sequence: trailing whitespace + one or more `#` + optional whitespace
|
||
const rawText = rest.replace(/^[ \t]+/, '').replace(/[ \t]+#+[ \t]*$/, '').replace(/^#+[ \t]*$/, '');
|
||
tokens.push({
|
||
level: hashes.length,
|
||
text: rawText.trim(),
|
||
line: i + 1, // 1-based
|
||
offset: charOffset,
|
||
});
|
||
}
|
||
}
|
||
|
||
charOffset += rawLine.length + 1;
|
||
}
|
||
|
||
return tokens;
|
||
}
|
||
|
||
// ─── collectSections ─────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Collect sections from `content`, calling `stopPredicate` on each heading to
|
||
* decide where sections end.
|
||
*
|
||
* Returns an array of `Section` objects, one per matched heading. The `body`
|
||
* of each section runs from the line after the heading up to (but not
|
||
* including) the next heading that satisfies `stopPredicate`, or EOF.
|
||
*
|
||
* Unlike a greedy-regex approach, this is a line-by-line walk — compatible
|
||
* with the repo's "line-by-line section collection" pattern.
|
||
*/
|
||
export function collectSections(
|
||
content: string,
|
||
stopPredicate: (heading: HeadingToken) => boolean,
|
||
): Section[] {
|
||
if (typeof content !== 'string' || content.length === 0) return [];
|
||
|
||
const headings = tokenizeHeadings(content);
|
||
if (headings.length === 0) return [];
|
||
|
||
const lines = content.split('\n');
|
||
const sections: Section[] = [];
|
||
|
||
// Build a set of line numbers (1-based) that are heading lines
|
||
const headingsByLine = new Map<number, HeadingToken>();
|
||
for (const h of headings) {
|
||
headingsByLine.set(h.line, h);
|
||
}
|
||
|
||
// Build a byte-offset table: lineOffsets[i] = byte offset of the start of line i+1 (1-based: i=0 → line 1)
|
||
// The body of a section starts at the byte after the heading line's trailing '\n'.
|
||
const lineOffsets: number[] = new Array<number>(lines.length);
|
||
let acc = 0;
|
||
for (let i = 0; i < lines.length; i++) {
|
||
lineOffsets[i] = acc;
|
||
acc += lines[i].length + 1; // +1 for the '\n' we split on
|
||
}
|
||
// lineOffsets[i] is the byte offset of line (i+1) (1-based). EOF sentinel:
|
||
const eofOffset = acc; // === content.length + (content.endsWith('\n') ? 0 : 0) ≈ content.length
|
||
|
||
let currentHeading: HeadingToken | null = null;
|
||
let currentBodyStart = 0;
|
||
let bodyLines: string[] = [];
|
||
|
||
const flush = (_bodyEndOffset: number): void => {
|
||
if (currentHeading !== null) {
|
||
const rawBody = bodyLines.join('\n');
|
||
const body = rawBody.trimEnd();
|
||
// INVARIANT: content.slice(bodyStart, bodyEnd) === body
|
||
// bodyEnd is derived from body.length, NOT from the raw separator offset,
|
||
// so round-trips via replaceSection(content, section, section.body) are exact.
|
||
sections.push({
|
||
heading: currentHeading,
|
||
body,
|
||
bodyStart: currentBodyStart,
|
||
bodyEnd: currentBodyStart + body.length,
|
||
});
|
||
currentHeading = null;
|
||
bodyLines = [];
|
||
}
|
||
};
|
||
|
||
for (let i = 0; i < lines.length; i++) {
|
||
const lineNo = i + 1; // 1-based
|
||
const h = headingsByLine.get(lineNo);
|
||
if (h !== undefined && stopPredicate(h)) {
|
||
// This heading is a stop boundary — flush current section, start new one.
|
||
// The body ends at the start of this heading line.
|
||
flush(lineOffsets[i]);
|
||
currentHeading = h;
|
||
// Body starts at the beginning of the line AFTER the heading line
|
||
const headingLineIdx = h.line - 1; // 0-based
|
||
currentBodyStart = lineOffsets[headingLineIdx] + lines[headingLineIdx].length + 1;
|
||
} else if (currentHeading !== null) {
|
||
bodyLines.push(lines[i]);
|
||
}
|
||
}
|
||
flush(eofOffset);
|
||
|
||
return sections;
|
||
}
|
||
|
||
// ─── collectSection ───────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Collect a single section whose heading satisfies `headingPredicate`.
|
||
*
|
||
* Options:
|
||
* - `levelBounded` (default: `true`): the section ends at the next heading of
|
||
* the same or higher level (lower level number = higher in the hierarchy).
|
||
* When `false`, the section body runs until any heading or EOF.
|
||
* Ignored when `stopAtLevel` is provided.
|
||
* - `stopAtLevel` (optional): when provided, the section ends at the next heading
|
||
* whose `level <= stopAtLevel`, regardless of the opener's level. This enables
|
||
* modeling sections like a `##`-opened section that also stops at `###`
|
||
* (pass `stopAtLevel: 3`). Takes precedence over `levelBounded` when set.
|
||
* - `stripFences` (default: `false`): apply `stripFencedCode` to the body
|
||
* before returning. The `heading` in the result always refers to the original
|
||
* heading (pre-strip).
|
||
*
|
||
* Returns `null` when no matching heading is found.
|
||
*/
|
||
export function collectSection(
|
||
content: string,
|
||
headingPredicate: (heading: HeadingToken) => boolean,
|
||
opts: CollectSectionOptions = {},
|
||
): Section | null {
|
||
if (typeof content !== 'string' || content.length === 0) return null;
|
||
|
||
const { levelBounded = true, stopAtLevel, stripFences = false } = opts;
|
||
|
||
const headings = tokenizeHeadings(content);
|
||
const targetIdx = headings.findIndex(headingPredicate);
|
||
if (targetIdx === -1) return null;
|
||
|
||
const target = headings[targetIdx];
|
||
const lines = content.split('\n');
|
||
|
||
// Determine which headings act as stops after the target
|
||
const bodyStartLine = target.line + 1; // 1-based, first line of body
|
||
let bodyEndLine = lines.length + 1; // 1-based, exclusive (default: EOF+1)
|
||
|
||
for (let j = targetIdx + 1; j < headings.length; j++) {
|
||
const next = headings[j];
|
||
let isStop: boolean;
|
||
if (stopAtLevel !== undefined) {
|
||
// stopAtLevel: stop at the next heading whose level <= stopAtLevel
|
||
isStop = next.level <= stopAtLevel;
|
||
} else {
|
||
isStop = levelBounded ? next.level <= target.level : true;
|
||
}
|
||
if (isStop) {
|
||
bodyEndLine = next.line; // stop before this line (1-based)
|
||
break;
|
||
}
|
||
}
|
||
|
||
// Compute character offsets for bodyStart.
|
||
// lineOffsets[i] = character offset of line (i+1) in content (1-based).
|
||
const lineOffsets: number[] = new Array<number>(lines.length);
|
||
let acc = 0;
|
||
for (let i = 0; i < lines.length; i++) {
|
||
lineOffsets[i] = acc;
|
||
acc += lines[i].length + 1; // +1 for the '\n' separator
|
||
}
|
||
const eofOffset = acc; // byte offset past the last line
|
||
|
||
// bodyStart: character offset of first line of body (bodyStartLine is 1-based)
|
||
const bodyStartOffset = bodyStartLine <= lines.length ? lineOffsets[bodyStartLine - 1] : eofOffset;
|
||
|
||
// Slice body lines (0-based array: bodyStartLine-1 to bodyEndLine-2 inclusive)
|
||
const bodyRaw = lines.slice(bodyStartLine - 1, bodyEndLine - 1).join('\n').trimEnd();
|
||
const body = stripFences ? stripFencedCode(bodyRaw).text : bodyRaw;
|
||
|
||
// INVARIANT: content.slice(bodyStart, bodyEnd) === body
|
||
// bodyEnd is derived from body.length so that replaceSection(content, section, section.body) === content.
|
||
return { heading: target, body, bodyStart: bodyStartOffset, bodyEnd: bodyStartOffset + body.length };
|
||
}
|
||
|
||
// ─── iterateBullets ───────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Extract bullet items from `sectionText`.
|
||
*
|
||
* Recognises three marker families:
|
||
* - **Checkbox**: `- [ ] text` (unchecked) and `- [x] text` / `- [X] text` (checked)
|
||
* - **Dash**: `- text`, `* text`, `+ text` (plain unordered list item)
|
||
* - **Numbered**: `1. text`, `42. text` (ordered list item)
|
||
*
|
||
* Indented continuation lines (lines that are not themselves bullet openers and
|
||
* have at least one leading space or tab) are accumulated into the current
|
||
* bullet's `text`.
|
||
*
|
||
* Blank lines terminate the current bullet (consistent with CommonMark block
|
||
* handling and the repo's existing bullet parsers).
|
||
*/
|
||
export function iterateBullets(sectionText: string): BulletItem[] {
|
||
if (typeof sectionText !== 'string' || sectionText.length === 0) return [];
|
||
|
||
const lines = sectionText.split('\n');
|
||
const items: BulletItem[] = [];
|
||
|
||
// Checkbox bullet: `<indent>- [ ] text` or `<indent>- [x] text`
|
||
const checkboxRe = /^(\s*)- \[([xX ])\] (.*)$/;
|
||
// Plain dash/asterisk/plus bullet: `<indent>- text`, `<indent>* text`, `<indent>+ text`
|
||
const dashRe = /^(\s*)[-*+] (.*)$/;
|
||
// Numbered bullet: `<indent>1. text`
|
||
const numberedRe = /^(\s*)\d+\. (.*)$/;
|
||
// Continuation: non-empty, indented, NOT a bullet opener
|
||
const continuationRe = /^[ \t]/;
|
||
|
||
let current: BulletItem | null = null;
|
||
|
||
const flush = (): void => {
|
||
if (current !== null) {
|
||
current.text = current.text.trim();
|
||
items.push(current);
|
||
current = null;
|
||
}
|
||
};
|
||
|
||
for (const rawLine of lines) {
|
||
// Strip trailing \r (CRLF safety)
|
||
const line = rawLine.replace(/\r$/, '');
|
||
const trimmed = line.trim();
|
||
|
||
// Blank line terminates current bullet
|
||
if (trimmed === '') {
|
||
flush();
|
||
continue;
|
||
}
|
||
|
||
// Checkbox bullet (checked or unchecked) — must test before dashRe
|
||
const cbm = checkboxRe.exec(line);
|
||
if (cbm) {
|
||
flush();
|
||
const stateChar = cbm[2];
|
||
const checked = stateChar === 'x' || stateChar === 'X';
|
||
current = {
|
||
marker: checked ? 'checkbox-checked' : 'checkbox-unchecked',
|
||
text: cbm[3],
|
||
indent: cbm[1],
|
||
checked,
|
||
};
|
||
continue;
|
||
}
|
||
|
||
// Numbered bullet
|
||
const nm = numberedRe.exec(line);
|
||
if (nm) {
|
||
flush();
|
||
current = {
|
||
marker: 'numbered',
|
||
text: nm[2],
|
||
indent: nm[1],
|
||
checked: null,
|
||
};
|
||
continue;
|
||
}
|
||
|
||
// Plain dash / asterisk / plus bullet
|
||
const dm = dashRe.exec(line);
|
||
if (dm) {
|
||
flush();
|
||
current = {
|
||
marker: 'dash',
|
||
text: dm[2],
|
||
indent: dm[1],
|
||
checked: null,
|
||
};
|
||
continue;
|
||
}
|
||
|
||
// Continuation line (indented, non-bullet) — append to current bullet
|
||
if (current !== null && continuationRe.test(line)) {
|
||
current.text += ' ' + trimmed;
|
||
continue;
|
||
}
|
||
|
||
// Non-bullet, non-continuation line (e.g. a paragraph, heading) — flush
|
||
flush();
|
||
}
|
||
flush();
|
||
|
||
return items;
|
||
}
|
||
|
||
// ─── updateBullet ─────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Locate the FIRST top-level bullet-opening line — checkbox (`- [ ]`/`- [x]`),
|
||
* dash/asterisk/plus (`- `/`* `/`+ `), or numbered (`1. `) — whose bullet text
|
||
* satisfies `match(bulletText, rawLine)`, replace that ONE physical line with
|
||
* `transform(rawLine)`, and return the resulting full content string. Every
|
||
* other byte in `content` — surrounding bullets, indentation, EOL style — is
|
||
* left untouched: this is a pure single-line splice, not a document-wide
|
||
* regex `.replace()`.
|
||
*
|
||
* Unlike `iterateBullets` (read-only, no offsets, and not itself fence-aware
|
||
* — callers pre-strip fences when that matters), `updateBullet` tracks
|
||
* character offsets itself so it can splice the transformed line back into
|
||
* the ORIGINAL `content`, and is fence-aware on its own: a bullet-shaped line
|
||
* inside a fenced code block (``` / ~~~, same CommonMark delimiter rules as
|
||
* `stripFencedCode`) is never offered to `match`/`transform`. (Tracked
|
||
* duplication of the fence state machine — same status as `tokenizeHeadings`'s
|
||
* copy, see its doc comment — pending a T-tier consolidation.)
|
||
*
|
||
* `rawLine` (second argument to both `match` and `transform`) is the
|
||
* UNMODIFIED physical line exactly as it appears between `\n` separators — so
|
||
* on a CRLF document its trailing `\r` is included, matching what a
|
||
* hand-rolled `^...[^\n]*`-shaped, `m`-flagged regex applied to the whole
|
||
* document would have seen. `bulletText` (first argument to `match`) is the
|
||
* bullet's own text with marker/checkbox stripped and any trailing `\r`
|
||
* removed — the same extraction `iterateBullets` uses for `BulletItem.text`.
|
||
*
|
||
* Only the OPENING line of a (possibly multi-line) bullet is ever matched or
|
||
* replaced — indented continuation lines are never presented to `match` or
|
||
* `transform`.
|
||
*
|
||
* The gap between the marker and its content tolerates 1 or more spaces — not
|
||
* only exactly one — mirroring CommonMark/GFM's 1–4-space allowance for
|
||
* list-marker spacing (`checkboxRe`/`numberedRe`/`dashRe`'s own dedicated
|
||
* quantifier caps at 4 per GFM; a wider run still recognises the line as a
|
||
* bullet opener via the uncapped `dashRe` fallback catching the excess as
|
||
* ordinary bullet text). So `- [ ] text` (two spaces), `1. text` (three
|
||
* spaces), and even a pathologically wide run are all recognised bullet
|
||
* openers, just as the canonical single-space `- [ ] text` / `1. text` are.
|
||
*
|
||
* Bounded no-op: if no bullet-opening line satisfies `match`, or `transform`
|
||
* returns a non-string, `content` is returned completely unchanged.
|
||
*/
|
||
export function updateBullet(
|
||
content: string,
|
||
match: (bulletText: string, rawLine: string) => boolean,
|
||
transform: (rawLine: string) => string,
|
||
): string {
|
||
if (typeof content !== 'string' || content.length === 0) return content;
|
||
|
||
const lines = content.split('\n');
|
||
|
||
// Marker-to-content gap: CommonMark/GFM tolerates 1–4 spaces between a list
|
||
// marker and its content (5+ pushes the content into indented-code-block
|
||
// territory) — so `- [ ] Phase 1: Foo` (two spaces) is still a valid
|
||
// bullet opener, not just the single-space `- [ ] …` shape. A hand-rolled
|
||
// single-space-only regex (e.g. the OLD `mutateMilestonePhase` checkbox
|
||
// regex before its `updateBullet` migration, which used `-\s*\[` — no cap,
|
||
// but at least 0+) would flip such a line; matching that requires this
|
||
// primitive's own bullet-opening recognition to tolerate the same gap,
|
||
// otherwise a wider-spaced bullet is silently never offered to `match`.
|
||
// F5 (#2245 review, nit): the gap also tolerates a literal TAB (`\t`), not
|
||
// only spaces — the OLD `-\s*\[` regex's `\s` class matched a tab too, so a
|
||
// `-\t[ ] text` bullet (tab-separated marker) must still be recognised here.
|
||
// Checkbox bullet: `<indent>- [ ] text` or `<indent>- [x] text`
|
||
const checkboxRe = /^(\s*)-[ \t]{1,4}\[([xX ])\] (.*)$/;
|
||
// Plain dash/asterisk/plus bullet: `<indent>- text`, `<indent>* text`, `<indent>+ text`
|
||
const dashRe = /^(\s*)[-*+][ \t]{1,4}(.*)$/;
|
||
// Numbered bullet: `<indent>1. text`
|
||
const numberedRe = /^(\s*)\d+\.[ \t]{1,4}(.*)$/;
|
||
|
||
// Fence tracking — same CommonMark delimiter rules as stripFencedCode
|
||
// (tracked duplication, see doc comment above).
|
||
const delimRe = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||
let openFence: FenceState | null = null;
|
||
|
||
let offset = 0;
|
||
for (let i = 0; i < lines.length; i++) {
|
||
const rawLine = lines[i];
|
||
const line = rawLine.replace(/\r$/, '');
|
||
|
||
const dm = delimRe.exec(line);
|
||
if (dm) {
|
||
const char = dm[2][0] as '`' | '~';
|
||
const len = dm[2].length;
|
||
const trailing = dm[3];
|
||
if (openFence === null) {
|
||
// CommonMark §4.5: backtick fence info string must not contain a backtick.
|
||
if (!(char === '`' && trailing.includes('`'))) {
|
||
// Valid opener — record fence state; this delimiter line is not a bullet.
|
||
openFence = { char, len };
|
||
offset += rawLine.length + 1;
|
||
continue;
|
||
}
|
||
// else: not a valid opener — falls through to the bullet check below.
|
||
} else if (char === openFence.char && len >= openFence.len && /^\s*$/.test(trailing)) {
|
||
// Closing delimiter — close the fence; this line is not a bullet.
|
||
openFence = null;
|
||
offset += rawLine.length + 1;
|
||
continue;
|
||
} else {
|
||
// Mismatched/insufficient delimiter while a fence is open — fence content.
|
||
offset += rawLine.length + 1;
|
||
continue;
|
||
}
|
||
}
|
||
|
||
if (openFence !== null) {
|
||
// Inside a fence — never a bullet candidate.
|
||
offset += rawLine.length + 1;
|
||
continue;
|
||
}
|
||
|
||
let bulletText: string | null = null;
|
||
const cbm = checkboxRe.exec(line);
|
||
if (cbm) {
|
||
bulletText = cbm[3];
|
||
} else {
|
||
const dm2 = dashRe.exec(line);
|
||
if (dm2) {
|
||
bulletText = dm2[2];
|
||
} else {
|
||
const nm = numberedRe.exec(line);
|
||
if (nm) bulletText = nm[2];
|
||
}
|
||
}
|
||
|
||
if (bulletText !== null && match(bulletText, rawLine)) {
|
||
const newLine = transform(rawLine);
|
||
if (typeof newLine !== 'string') return content;
|
||
return content.slice(0, offset) + newLine + content.slice(offset + rawLine.length);
|
||
}
|
||
|
||
offset += rawLine.length + 1;
|
||
}
|
||
|
||
return content;
|
||
}
|
||
|
||
// ─── extractTaggedBlocks ──────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Return the inner text of every `<tagName>…</tagName>` block in `content`,
|
||
* in document order.
|
||
*
|
||
* Designed for extracting structured XML-like annotation blocks that live in
|
||
* markdown prose (e.g. `<decisions>…</decisions>`, `<requirements>…</requirements>`).
|
||
* Returns `[]` when no matching blocks are found.
|
||
*
|
||
* The `tagName` argument is regex-escaped, so names that contain regex
|
||
* metacharacters (e.g. `foo.bar`, `my+tag`) are matched literally.
|
||
*
|
||
* **Input contract:** the caller decides whether to pass raw or fence-stripped
|
||
* content. `extractTaggedBlocks` is a pure block extractor — it does NOT strip
|
||
* fenced code blocks itself. If a `<tagName>` block appears inside a fenced code
|
||
* block and should be excluded, the caller should apply `stripFencedCode` first.
|
||
*
|
||
* **Nested tags are NOT supported.** The body scan terminates at the NEXT
|
||
* opening of the same tag (the ReDoS-safe boundary, #2128). Given
|
||
* `<x><x>inner</x></x>`, `extractTaggedBlocks(content, 'x')` returns `['inner']`
|
||
* — the well-formed inner block; the unterminated outer `<x>` is skipped.
|
||
* Callers that need true nesting must use a proper XML/HTML parser.
|
||
*
|
||
* `allowAttributes` (default `false`): when `true`, the opening tag may carry
|
||
* bounded attributes (`<tag foo="x">`) — needed for `<task type="…">` blocks.
|
||
* Leave `false` for tags that must match exactly (e.g. `<decisions>`), and never
|
||
* enable it for a tag where an attributed form is semantically distinct.
|
||
*
|
||
* Generalises `decisions.cts`'s bespoke `matchAll(/<decisions>([\s\S]*?)<\/decisions>/g)`
|
||
* so tier T1 can drop its own copy (tracked duplication until T1 lands).
|
||
*/
|
||
export function extractTaggedBlocks(content: string, tagName: string, allowAttributes = false): string[] {
|
||
if (typeof content !== 'string' || content.length === 0) return [];
|
||
if (typeof tagName !== 'string' || tagName.length === 0) return [];
|
||
|
||
const pattern = taggedBlockPattern(tagName, 'g', allowAttributes);
|
||
const results: string[] = [];
|
||
let match: RegExpExecArray | null;
|
||
while ((match = pattern.exec(content)) !== null) {
|
||
results.push(match[1]);
|
||
}
|
||
return results;
|
||
}
|
||
|
||
/**
|
||
* Build the single, ReDoS-safe `<tag>…</tag>` block regex shared by
|
||
* `extractTaggedBlocks` (extract bodies) and `stripTaggedBlocks` (remove blocks).
|
||
*
|
||
* Safety: the body terminates at the NEXT opening of this tag (stop-at-next-open)
|
||
* instead of lazily rescanning the whole remaining document for a `</tag>` that
|
||
* may never appear — so a document full of unclosed `<tag>` openings scans
|
||
* LINEARLY, not quadratically (#2128). Group 1 is the block body.
|
||
*
|
||
* `allowAttributes`: when `true`, the opener accepts bounded attributes
|
||
* (`<tag foo="x">`) and the body boundary is `<tag` followed by a space or `>`.
|
||
* When `false`, the opener is the EXACT `<tag>` and the boundary is exact `<tag>`,
|
||
* so an attributed `<tag foo>` is neither an opener nor a boundary — it is body
|
||
* content. That exact form is load-bearing for `<details>` stripping: `<details
|
||
* open>` marks the ACTIVE milestone and must be preserved, not stripped (#557).
|
||
*/
|
||
function taggedBlockPattern(tagName: string, flags: string, allowAttributes: boolean): RegExp {
|
||
const esc = tagName.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
||
const open = allowAttributes ? `<${esc}(?:\\s[^>]{0,1000})?>` : `<${esc}>`;
|
||
const boundary = allowAttributes ? `<${esc}[\\s>]` : `<${esc}>`;
|
||
return new RegExp(`${open}((?:(?!${boundary})[\\s\\S])*?)</${esc}>`, flags);
|
||
}
|
||
|
||
/**
|
||
* Remove every `<tagName>…</tagName>` block (opening tag, body, and closing tag)
|
||
* from `content`. The ReDoS-safe counterpart to `extractTaggedBlocks` — same
|
||
* hardened pattern, `.replace(…, '')` instead of body extraction. `allowAttributes`
|
||
* defaults to `false` so `<details open>` (active milestone) is preserved (#557);
|
||
* case-insensitive by default (matching the `<details>` strip call sites), pass
|
||
* `caseSensitive` to force exact-case matching.
|
||
*/
|
||
export function stripTaggedBlocks(content: string, tagName: string, allowAttributes = false, caseSensitive = false): string {
|
||
if (typeof content !== 'string' || content.length === 0) return '';
|
||
if (typeof tagName !== 'string' || tagName.length === 0) return content;
|
||
return content.replace(taggedBlockPattern(tagName, caseSensitive ? 'g' : 'gi', allowAttributes), '');
|
||
}
|
||
|
||
// ─── replaceSection ───────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Splice `newBody` in place of a section's body and return the resulting
|
||
* full content string.
|
||
*
|
||
* Uses the `bodyStart`/`bodyEnd` character offsets carried by the `Section`
|
||
* type to perform a pure string splice — no regex, no line-counting. The
|
||
* heading is preserved verbatim; only the bytes between `bodyStart` and
|
||
* `bodyEnd` are replaced.
|
||
*
|
||
* The `newBody` is inserted as-is between `content.slice(0, bodyStart)` and
|
||
* `content.slice(bodyEnd)`. If `newBody` should end with a trailing newline
|
||
* before the next section's heading, the caller is responsible for including
|
||
* it (consistent with how `trimEnd()` is applied to collected bodies — see
|
||
* `collectSections`/`collectSection`).
|
||
*
|
||
* Typical read-modify-write pattern (T6 state.cts use case):
|
||
* ```
|
||
* const section = collectSection(content, h => h.text === 'Name');
|
||
* if (section) {
|
||
* content = replaceSection(content, section, newBody);
|
||
* }
|
||
* ```
|
||
*
|
||
* CRLF-safe: the splice is purely character-offset-based, so CRLF sequences
|
||
* are preserved in the surrounding content unchanged.
|
||
*/
|
||
export function replaceSection(content: string, section: Section, newBody: string): string {
|
||
if (typeof content !== 'string') return content;
|
||
if (typeof newBody !== 'string') return content;
|
||
return content.slice(0, section.bodyStart) + newBody + content.slice(section.bodyEnd);
|
||
}
|
||
|
||
// ─── withSection ──────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Locate the section whose heading matches `target`, run `edit` against ONLY
|
||
* that section's body, and splice the result back into `content`.
|
||
*
|
||
* `target` is either an exact (trimmed) heading-text match or a predicate
|
||
* function over `HeadingToken`. `edit` receives ONLY the section body — so any
|
||
* regex it runs is physically confined to that section — an edit cannot cross
|
||
* a section boundary (ADR-2143 §4, structurally retires the #2130/#2067/#2080
|
||
* boundary-crossing class, where a hand-rolled regex escaped its intended
|
||
* section and mutated a sibling/shipped/backticked-literal occurrence instead).
|
||
*
|
||
* Bounded no-op behaviour (Phase 3 of ADR-2143 adds fail-loud diagnostics on
|
||
* top of this):
|
||
* - No heading matches `target` → `content` is returned unchanged.
|
||
* - `edit` returns a non-string, or returns the same string it was given →
|
||
* `content` is returned unchanged (no-op splice avoided).
|
||
*
|
||
* `opts` is forwarded verbatim to `collectSection` (see its doc comment for
|
||
* `levelBounded` / `stopAtLevel` / `stripFences` semantics) — it lets a caller
|
||
* whose heading levels are non-uniform (e.g. a mix of `###`/`####` phase
|
||
* headings) choose the correct section-end rule instead of relying on the
|
||
* `levelBounded: true` default.
|
||
*/
|
||
export function withSection(
|
||
content: string,
|
||
target: string | ((h: HeadingToken) => boolean),
|
||
edit: (body: string) => string,
|
||
opts: CollectSectionOptions = {},
|
||
): string {
|
||
if (typeof content !== 'string') return content;
|
||
const predicate = typeof target === 'function'
|
||
? target
|
||
: (h: HeadingToken) => h.text.trim() === target.trim();
|
||
const section = collectSection(content, predicate, opts);
|
||
if (!section) return content; // bounded no-op on miss (Phase 3 adds fail-loud)
|
||
const newBody = edit(section.body);
|
||
if (typeof newBody !== 'string' || newBody === section.body) return content;
|
||
return replaceSection(content, section, newBody);
|
||
}
|
||
|
||
// ─── deleteSection ────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Delete an entire section — the matching heading line ITSELF plus its body —
|
||
* and return the resulting full content string.
|
||
*
|
||
* Locates the target heading via the SAME machinery `collectSection` uses
|
||
* (`tokenizeHeadings` + `headingPredicate`), then determines the stop boundary
|
||
* with the SAME level-bounding rule (`levelBounded` / `stopAtLevel`, see
|
||
* `CollectSectionOptions`): the deleted range runs from the target heading's
|
||
* OWN start offset up to (but not including) the next heading whose level is
|
||
* the same-or-higher (lower level number) than the target's — so a level-3
|
||
* `### Phase N` section deletes through any nested `####` content but STOPS at
|
||
* the next `##`/`###` sibling, whatever that heading's text is (unlike a
|
||
* hand-rolled regex anchored to a specific heading TEXT pattern, which keeps
|
||
* scanning past an unrelated heading and can run away to EOF when no further
|
||
* heading of that specific text shape follows — the whole-section-deletion
|
||
* data-loss class this primitive retires).
|
||
*
|
||
* Unlike `collectSection`/`withSection` (which operate on a section's BODY
|
||
* only, leaving the heading line untouched), `deleteSection` removes the
|
||
* heading line too — the counterpart for "delete section" call sites that
|
||
* `withSection` structurally cannot serve.
|
||
*
|
||
* Collapses at most one resulting blank-line seam: if removing the section
|
||
* leaves 2+ blank lines immediately at the splice point (e.g. the original
|
||
* document already had a double-blank separator immediately before the
|
||
* deleted heading), the seam is normalized down to a single blank line so no
|
||
* double-blank gap accumulates where the section used to sit. Content
|
||
* elsewhere in the document is never touched.
|
||
*
|
||
* Returns `content` unchanged when no heading matches `headingPredicate`
|
||
* (bounded no-op, mirroring `withSection`'s miss behaviour).
|
||
*/
|
||
export function deleteSection(
|
||
content: string,
|
||
headingPredicate: (heading: HeadingToken) => boolean,
|
||
opts: CollectSectionOptions = {},
|
||
): string {
|
||
if (typeof content !== 'string') return content;
|
||
|
||
const { levelBounded = true, stopAtLevel } = opts;
|
||
|
||
const headings = tokenizeHeadings(content);
|
||
const targetIdx = headings.findIndex(headingPredicate);
|
||
if (targetIdx === -1) return content;
|
||
|
||
const target = headings[targetIdx];
|
||
const lines = content.split('\n');
|
||
|
||
// Determine the stop line using the SAME level-bounding rule collectSection uses.
|
||
let stopLine = lines.length + 1; // 1-based, exclusive (default: EOF+1)
|
||
for (let j = targetIdx + 1; j < headings.length; j++) {
|
||
const next = headings[j];
|
||
let isStop: boolean;
|
||
if (stopAtLevel !== undefined) {
|
||
isStop = next.level <= stopAtLevel;
|
||
} else {
|
||
isStop = levelBounded ? next.level <= target.level : true;
|
||
}
|
||
if (isStop) {
|
||
stopLine = next.line;
|
||
break;
|
||
}
|
||
}
|
||
|
||
// Character offsets — same line-offset table collectSection builds.
|
||
const lineOffsets: number[] = new Array<number>(lines.length);
|
||
let acc = 0;
|
||
for (let i = 0; i < lines.length; i++) {
|
||
lineOffsets[i] = acc;
|
||
acc += lines[i].length + 1; // +1 for the '\n' separator
|
||
}
|
||
const eofOffset = acc;
|
||
|
||
const sectionStart = lineOffsets[target.line - 1]; // start of the target heading LINE itself
|
||
const sectionEnd = stopLine <= lines.length ? lineOffsets[stopLine - 1] : eofOffset;
|
||
|
||
const before = content.slice(0, sectionStart);
|
||
const after = content.slice(sectionEnd);
|
||
|
||
// Collapse a resulting blank-line seam to at most one blank line (2 newlines).
|
||
// Only the tail of `before` (immediately at the splice point) is touched —
|
||
// this never reaches into unrelated content elsewhere in the document.
|
||
const collapsedBefore = before.replace(/(?:\r\n|\n){3,}$/, (m) => (m.includes('\r\n') ? '\r\n\r\n' : '\n\n'));
|
||
|
||
return collapsedBefore + after;
|
||
}
|
||
|
||
// Consumers: require('../gsd-core/bin/lib/markdown-sectionizer.cjs')
|
||
// Named CJS exports are the canonical surface (ADR-457 .cts → .cjs build-at-publish).
|