diff --git a/.changeset/eager-orcas-chatter.md b/.changeset/eager-orcas-chatter.md new file mode 100644 index 000000000..d519a1a4b --- /dev/null +++ b/.changeset/eager-orcas-chatter.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 3739 +--- +**`deferred-items.md` entries written with `*`, `+` or an ordered marker are no longer silently dropped** — the parser recognised only `- `, so a deferred list written with any other standard Markdown list marker contributed zero entries and reported as a clean zero to `audit-open`. An ordered list counts when it starts at `0.` or `1.` or continues a list already open at its level; `1)` is not a marker. Also fixed: a `status:` line inside a fenced code block indented four or more spaces — what a fence under a nested bullet looks like — was not treated as fenced, so a `status: resolved` written as documentation resolved the entry containing it; and an unclosed fence now ends with its own entry instead of hiding every entry after it. `audit-open acknowledge` reads and writes the same grammar — its line selection goes through the reader's own classifier on the headless and (since #3781) the heading-delimited shape alike — so an entry the audit surfaces under any of these markers can be acknowledged, and a heading-delimited file written with one of the newly recognised markers now surfaces its entries for `complete-milestone` to acknowledge in place, where it previously closed over them unseen. (#3702) diff --git a/gsd-core/references/executor-examples.md b/gsd-core/references/executor-examples.md index 4b3953956..0e75b1ff2 100644 --- a/gsd-core/references/executor-examples.md +++ b/gsd-core/references/executor-examples.md @@ -69,6 +69,48 @@ - MAYBE → Rule 4 (ask the user) - NO → Out of scope (log to deferred-items.md) +### Writing `deferred-items.md` + +The file has no template — write it by hand, as a Markdown list under a +`## Deferred Items` heading. What counts as one entry: + +- One entry per top-level list item. `-`, `*` and `+` all count, and so does a + dot-terminated ordered marker (`1.`) when the list starts at `0.` or `1.`, or + continues a list already open at that level — a sentence that merely opens + with a number (`2026. was a bad year`) is prose, not an item, and so is a + list numbered from `2.` upward until its first `0.`/`1.` line. `1)` is not a + marker here, and neither is an ordinal past nine digits (`999999999.` + counts, `1234567890.` does not). +- Continuation lines indent beneath their entry. Fields go on those lines: + `status: resolved`, or the bolded `**Status:** resolved` convention. **The + BARE key is lower-case only** — write `Status: resolved` without the bold and + the field is not read, so the entry stays open with no warning. Bold it or + lower-case it. The bolded form matches the key case-insensitively, and the + VALUE is case-insensitive in both forms. +- A `* * *` or `- - -` separator closes the list rather than opening an entry, + and nothing inside a fenced code block is an entry or a field — at any indent, + including one deeper than CommonMark's three-space cap, which is what a fence + written under a nested bullet looks like. A fence that is never closed runs to + the end of its own entry and no further, so an unclosed delimiter cannot hide + the entries after it — a closed pair of delimiters is a fence, whatever sits + between them. + +An entry is RESOLVED only if it carries an explicit `status: resolved`. Anything +else — including an entry with no `status:` at all — stays open and will surface +in `audit-open`, `audit-uat` and `complete-milestone`. That is deliberate: the +scanner never silently drops a possibly-open item. What it reads as something +other than an item is the short list above — a fenced line, a separator, and an +ordered list numbered from `2.` upward at a paragraph position — and nothing +else. + +```markdown +## Deferred Items + +- Retry budget is hardcoded at 3 + status: open + **What:** `fetchWithRetry` ignores the configured budget. +``` + ## Checkpoint Examples ### Good checkpoint placement diff --git a/gsd-core/workflows/progress/steps/forensic-audit.md b/gsd-core/workflows/progress/steps/forensic-audit.md index b07bb17a3..69bd64b5f 100644 --- a/gsd-core/workflows/progress/steps/forensic-audit.md +++ b/gsd-core/workflows/progress/steps/forensic-audit.md @@ -89,7 +89,7 @@ Glob every phase directory's SCOPE BOUNDARY log (executor writes out-of-scope di ls .planning/phases/*/deferred-items.md 2>/dev/null || true ``` -For each `deferred-items.md` found, read its entries (bullet list, one entry per top-level `- ` line, continuation lines indented beneath it). An entry is RESOLVED only if it carries an explicit `status: resolved` field (case-insensitive) on one of its lines; every other entry — including one with no `status:` field at all — is UNRESOLVED and must be surfaced (fail-safe: never silently drop a possibly-open item). +For each `deferred-items.md` found, read its entries (bullet list, one entry per top-level list item — `-`, `*`, `+` or an ordered marker of one to nine digits terminated by a DOT, where an ordered list counts when it starts at `0.` or `1.` or continues a list already open at its level; `1)` is not a marker and neither is a ten-digit ordinal (`999999999.` counts, `1234567890.` does not) — with continuation lines indented beneath it; fenced code blocks at ANY indent, including deeper than CommonMark's three-space cap, and `* * *`-style separators are not entries; an unclosed fence ends with its own entry and never hides the entries after it; a closed pair of delimiters is a fence whatever sits between them). An entry is RESOLVED only if it carries an explicit `status: resolved` field on one of its lines — the bare key lower-case only, or the bolded `**Status:**` form in any case; the VALUE is case-insensitive in both; every other entry — including one with no `status:` field at all — is UNRESOLVED and must be surfaced (fail-safe: never silently drop a possibly-open item). Emit: - ✓ `No unresolved deferred items` — if no `deferred-items.md` files exist, or every entry in every file is `status: resolved` diff --git a/src/uat.cts b/src/uat.cts index 4acba0713..4532367a3 100644 --- a/src/uat.cts +++ b/src/uat.cts @@ -1877,7 +1877,7 @@ function parseGapsTableItems(sectionBody: string): UatItem[] { * surfaced. * * #3457: when the section body contains headings, entries are delimited by - * LEAF headings (see `splitDeferredHeadingEntries`) rather than by bullets — + * LEAF headings (see `splitDeferredHeadingEntriesDetailed`) rather than by bullets — * the executor convention writes one deferred item as a heading followed by * sibling `- **Field:** …` bullets, which the bullet-only split mis-counted as * one item PER BULLET. A body with no headings keeps the original @@ -1909,19 +1909,28 @@ function parseDeferredItemsWithStatus(content: string): Array<{ name: string; st // before field extraction, not just line 0 (which `extractGapEntryFields` // does for the headless/Gaps shape, where a later `- ` line is a nested // sub-list, not a field). - const headingEntries = splitDeferredHeadingEntries(sectionBody); + const headingEntries = splitDeferredHeadingEntriesDetailed(sectionBody); + // The opener flags are HANDED DOWN rather than pre-applied (#3702 round 3, + // m7/m8). Marker-stripping the lines here and passing the result meant the + // reader's fence scan ran over text the splitter never saw, and the namer + // stripped a marker off the heading TEXT. Both consumers now take the raw + // lines plus the splitter's own per-line verdict — a rejected ordinal + // ("3. status: resolved" as prose) still keeps its `3. ` and yields no field, + // because that verdict is what carries the rejection. const entries = headingEntries !== null - ? headingEntries.map((entryLines) => ({ - lines: entryLines, - fields: extractGapEntryFields(entryLines.map(stripLeadingBulletMarker)), + ? headingEntries.map((entry) => ({ + lines: entry.lines, + opener: entry.opener, + fields: extractGapEntryFields(entry.lines, DEFERRED_BULLET_MARKERS, entry.opener), })) - : splitGapsEntries(sectionBody).map((entryLines) => ({ + : splitGapsEntries(sectionBody, DEFERRED_BULLET_MARKERS).map((entryLines) => ({ lines: entryLines, - fields: extractGapEntryFields(entryLines), + opener: undefined, + fields: extractGapEntryFields(entryLines, DEFERRED_BULLET_MARKERS), })); - for (const { lines: entryLines, fields } of entries) { - const text = rawGapEntryText(entryLines); + for (const { lines: entryLines, opener, fields } of entries) { + const text = rawGapEntryText(entryLines, DEFERRED_BULLET_MARKERS, opener); if (!text) continue; items.push({ name: text, status: fields.status || '' }); @@ -1953,6 +1962,48 @@ function parseDeferredItems(content: string): UatItem[] { })); } +/** + * The line ending for an entry that ends the FILE, where the entry is a single + * line and therefore carries no terminator of its own to copy. No entry-local + * evidence exists here — the separator before the entry terminates the + * PREVIOUS line, not this one — so this asks the weaker question that CAN be + * answered: does anything before the entry, within the scope the caller passes, + * contradict CRLF? Uniform CRLF across that scope is the one case where + * appending a `\r\n` cannot make the file more irregular. It fails CLOSED: + * any bare `\n` in scope, or no scope at all, yields LF. + * + * Adopted from #3773 (`crlfAtEof`), whose four counterexamples fixed the scope + * and are ported alongside it. Every simpler choice is refuted by a named test: + * the separator immediately PRECEDING the entry propagates an isolated CRLF + * into an LF-dominant list, because it terminates the previous line rather than + * this one — that is the algorithm this PR shipped through round 3 and it is + * withdrawn here. The whole DOCUMENT rejects CRLF over an unrelated bare `\n` + * elsewhere, inside a fenced block say. The deferred-items SECTION body is + * right when a heading delimits one, and becomes the whole document when it + * does not. + * + * Scope, therefore: the section body when `## Deferred Items` delimits one (its + * own preamble belongs to that section), else the entry-list region, where only + * the entries can be trusted. + * + * WITH ONE CORRECTION to #3773, which is its B4. The entry-list region goes + * EMPTY exactly when the list is undelimited AND holds a single entry, since + * the region runs from the first entry's start to the insertion point and those + * coincide. `crlfAtEof('')` is `false`, so a bare `\n` was inserted into a CRLF + * document — `'preamble\r\n\r\n- alpha'` gained one — which is the very defect + * the fallback exists to close, and it breaks the fix's own uniform-CRLF + * invariant. When the preferred region is empty the caller widens to everything + * preceding the insertion point rather than asserting LF from no evidence. That + * can only ever loosen a scope that was carrying zero information, and the + * predicate stays fail-closed over the wider one, so a contradicting bare `\n` + * still yields LF. An entry at offset 0 of an undelimited document has no + * evidence under either scope and stays LF, rather than inventing an ending + * from nothing. + */ +function crlfAtEof(before: string): boolean { + return before.length > 0 && !/(^|[^\r])\n/.test(before); +} + // ─── acknowledgeDeferredItem ─────────────────────────────────────────────────── /** Result of `acknowledgeDeferredItem`. */ @@ -1975,25 +2026,27 @@ interface AcknowledgeDeferredItemResult { * `status:` away from `acknowledged` (or delete the field) and it resurfaces * with no separate cleanup step, exactly like every other category's marker. * - * #3781: the heading-delimited (#3457) entry shape is SUPPORTED, via - * `splitDeferredHeadingEntriesWithSpans` — a span-carrying sibling of the - * reader's walk that records each entry's (start, end) character span in the - * SAME pass that groups its lines (the identical technique - * `splitGapsEntriesWithSpans` uses for the headless shape). Leaf entries keep - * their RAW heading line as `lines[0]` so `sectionBody.slice(start, end)` is - * byte-verbatim; pending (preamble / container-direct) regions are contiguous - * slices handed to `splitGapsEntriesWithSpans` with a baseOffset translation. - * Two write rules differ from the headless path on this shape: the status - * search runs over the READER-form lines (what the reader actually parses — - * including the leaf line-0 corner where the heading text itself parses as a - * status field, whose raw line is rewritten with its ATX prefix preserved), - * and the insert branch inserts after the entry's LAST NON-BLANK line — a - * heading entry's body is frequently a soft-wrapped sentence, and splicing - * after line 0 would split it in half (#3781's sentence-split trap). - * Entries whose span embeds a GFM table row are non-contiguous (table lines - * are excluded from entries) and still refuse (`unsupported_heading_shape`) - * rather than risk a wrong-entry write; the fully-headless shape below is - * byte-for-byte the pre-#3781 path. + * #3781: the heading-delimited (#3457) entry shape is SUPPORTED. The + * reader's own walk, `splitDeferredHeadingEntriesDetailed`, records each + * entry's (start, end) character span in the SAME pass that groups its + * lines — the technique `splitGapsEntriesWithSpans` already uses for the + * headless shape — so there is no second walk for the writer to drift from + * (#3702 round 5: upstream's fix shipped a hyphen-only sibling walk, and this + * PR's widened grammar would have left it reading a different set of + * entries than the reader; folding the spans into the one walk is what + * keeps the writer and the reader on one grammar). The heading half, + * `acknowledgeHeadingShapedEntry`, shares this function's guards and its + * rewrite/insert machinery through the same `entryFieldLines` seam, with two + * shape-specific rules: the status search runs over the READER-form lines + * (the heading TEXT on a leaf's line 0 — including the corner where that + * text itself parses as a status field, rewritten with its ATX prefix + * preserved), and the insert branch inserts after the entry's LAST NON-BLANK + * line, because a heading entry's body is frequently a soft-wrapped + * sentence and splicing after line 0 would split it (#3781's sentence-split + * trap). Entries whose span embeds a GFM table row are non-contiguous (table + * lines are excluded from entries) and still refuse + * (`unsupported_heading_shape`) rather than risk a wrong-entry write; the + * fully-headless shape below is byte-for-byte the pre-#3781 path. * * Also refuses `ambiguous` (2+ entries share the exact same text — status must * be unique to identify one) and `not_found`, and is a no-op @@ -2040,16 +2093,16 @@ function acknowledgeDeferredItem(content: string, targetText: string): Acknowled ); const sectionBody = deferredSection ? deferredSection.body : content; - // #3781: the heading-delimited shape carries its own span walk; the - // headless path below is unchanged. - const headingEntries = splitDeferredHeadingEntriesWithSpans(sectionBody); + // #3781: the heading-delimited shape carries its own spans, recorded by + // the reader's walk; the headless path below is unchanged. + const headingEntries = splitDeferredHeadingEntriesDetailed(sectionBody); if (headingEntries !== null) { return acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, headingEntries, targetText }); } - const entries = splitGapsEntriesWithSpans(sectionBody); + const entries = splitGapsEntriesWithSpans(sectionBody, DEFERRED_BULLET_MARKERS); const matches = entries - .map((entry) => ({ entry, text: rawGapEntryText(entry.lines) })) + .map((entry) => ({ entry, text: rawGapEntryText(entry.lines, DEFERRED_BULLET_MARKERS) })) .filter((e) => e.text === targetText); if (matches.length === 0) return { content, status: 'not_found' }; @@ -2057,7 +2110,7 @@ function acknowledgeDeferredItem(content: string, targetText: string): Acknowled const { entry } = matches[0]; const { lines: entryLines, start, end } = entry; - const fields = extractGapEntryFields(entryLines); + const fields = extractGapEntryFields(entryLines, DEFERRED_BULLET_MARKERS); if (fields.status && fields.status.toLowerCase() === 'resolved') { return { content, status: 'already_resolved' }; } @@ -2074,100 +2127,139 @@ function acknowledgeDeferredItem(content: string, targetText: string): Acknowled // comparison that selected this entry — this catches real drift between // the two rather than a regex trivially guaranteed to agree with itself. const strippedForVerify = matchedLines.map((l) => l.replace(/\r$/, '')); - if (rawGapEntryText(strippedForVerify) !== targetText) { + if (rawGapEntryText(strippedForVerify, DEFERRED_BULLET_MARKERS) !== targetText) { return { content, status: 'match_verification_failed' }; } const matchIndexInContent = sectionOffset + start; - // #3740: the search must mirror the reader exactly. extractGapEntryFields - // strips a bullet marker on line 0 ALONE — a later `- ` line is a nested - // sub-list, never a field line — so a marker-prefixed match on any - // continuation line would rewrite a line no reader reads and report `ok` - // while the entry stays outstanding. Line 0 KEEPS the marker-optional - // form: the reader de-bullets it, so `- status: open` as the entry line is - // a real field there (and first-wins means the insert branch could not - // outrank it). Everything else falls through to the insert branch below, - // which the marker-free and no-status controls already round-trip. + // Locate the status line with the READER'S OWN classifier, never a + // writer-side regex (#3702 round 3, B1/B3 — the shape is #3773's, + // parameterised here by the widened marker set per the round-3 review's + // prescribed end state). The only line worth rewriting in place is one the + // reader will read back as `fields.status`. // - // #3775: the CASE axis of the same rule. The reader lowercases BOLDED - // keys only; bare keys keep their literal case (#3457 design), so a bare - // `Status:`/`STATUS:` line is stored under key `Status` and never read as - // fields.status. The search therefore matches a bolded key in ANY case - // and a bare key in LOWERCASE only — never a bare Title-case/UPPER line, - // which must fall through to the insert branch whose lowercase output the - // reader consumes (leaving any human `Status: resolved` untouched). - const statusFieldBoldedRe = /^\s*\*+status:\*+/i; - const statusFieldBareRe = /^\s*status:/; - const statusFieldBoldedReLine0 = /^\s*(?:-\s+)?\*+status:\*+/i; - const statusFieldBareReLine0 = /^\s*(?:-\s+)?status:/; - const statusLineIdx = matchedLines.findIndex((rawLine, idx) => { - const line = rawLine.replace(/\r$/, ''); - return idx === 0 - ? (statusFieldBoldedReLine0.test(line) || statusFieldBareReLine0.test(line)) - : (statusFieldBoldedRe.test(line) || statusFieldBareRe.test(line)); - }); + // What this closes: round 2 widened the writer's finder to the deferred + // marker set while `extractGapEntryFields` still read a marker only on line + // 0. A nested ` * status: pending` was therefore SELECTED by the writer and + // invisible to the reader — acknowledge rewrote it, returned `ok`, and the + // item stayed outstanding forever. Measured against a `next` build: `*`, `+` + // and `1.` all resolved on base and stopped resolving here, so it was a + // regression, not a gap in new behaviour. The hyphen form of the same shape + // (` - status:`) was already broken on `next`; it is fixed here too, since + // one classifier cannot be right for three markers and wrong for the fourth. + // + // A line the reader skips falls through to the INSERT branch, which writes a + // line the reader does read — the fail-safe direction. That covers a bare + // capitalised `Status:` (the reader stores it under `Status`, not `status`) + // and a `status:` line inside a fenced block, both of which the writer must + // NOT rewrite in place. Selecting either one produced an entry that could not + // be acknowledged at all; that is why the selection goes through + // `entryFieldLines` rather than the classifier directly. + const statusLineIdx = entryFieldLines(matchedLines, DEFERRED_BULLET_MARKERS) + .findIndex((field) => field?.key === 'status'); - // No CRLF-preservation branch here (WARNING 1, #3458 follow-up review): - // every write goes through `platformWriteSync` → `normalizeContent`, which - // for a `.md` path unconditionally runs `_normalizeMd` — whole-file - // `\r\n` → `\n`, plus blank-line normalization around headings/lists — on - // EVERY write, not just this one. That is this codebase's single, - // deliberate OS-facing I/O seam (`shell-command-projection.cts`), applied - // uniformly to every `.md` writer; carving out one exception here would - // fight it rather than follow it, for a guarantee (byte-identical CRLF on - // disk) the seam already makes impossible. A marker write on a CRLF - // `deferred-items.md` normalizes the WHOLE file to LF, same as any other - // `.md` write in this codebase — expected, not a regression to guard - // against. Where a source line still carries a trailing `\r` (read from an - // on-disk CRLF document before normalization), `String.prototype.replace` - // consumes it as part of `.*$` and the replacement text does not - // reproduce it, so it is dropped here too — consistent with the eventual - // whole-file normalization rather than duplicating it. + // Per-line CRLF preservation is the honest in-memory contract. The lines + // here are RAW — `audit acknowledge` hands this function the `readFileSync` + // content of an on-disk `deferred-items.md`, and `_normalizeMd` runs only on + // WRITE — so on a CRLF document every line but the span's last still carries + // its `\r`. The write path's whole-file normalization still decides what + // reaches disk; this function does not duplicate that decision, and it no + // longer silently drops the `\r` either. (Round 2's comment here argued the + // opposite contract. It is withdrawn: #3773 documents per-line preservation + // in this same function, and two opposite contracts in one function was a + // round-3 blocker in its own right.) let newMatchedLines: string[]; if (statusLineIdx === -1) { - const bulletIndentMatch = matchedLines[0].match(/^(\s*)-\s+/); - const continuationIndent = ' '.repeat((bulletIndentMatch ? bulletIndentMatch[1].length : 0) + 2); + const bulletIndentMatch = matchedLines[0].replace(/\r$/, '').match(DEFERRED_BULLET_MARKERS.strip); + // The entry's own indent CHARACTERS, never a count of them — a tab counted + // as one column and re-emitted as one space puts a 3-space continuation + // under a tab-indented bullet. Identical output for all-space indents. + const continuationIndent = `${bulletIndentMatch ? bulletIndentMatch[1] : ''} `; + // The new line goes right after line 0, so line 0 stops being the span's + // last line. Under CRLF the span's last line is the one WITHOUT a `\r` + // (the file's own `\r\n` follows the span), so the ending is read from the + // entry's OWN boundary and never sniffed from the whole document — a + // mixed-ending file must keep its LF opener. At end-of-file there is no + // following separator, so the boundary immediately PRECEDING the entry is + // the remaining local evidence; an entry at offset 0 has neither and stays + // LF rather than inventing an ending from nothing. + const line0HadCr = matchedLines[0].endsWith('\r'); + // A single-line entry's line 0 IS the span's last line, so its own + // terminator sits OUTSIDE the span and the separator FOLLOWING the span is + // the evidence. At end-of-file there is no such separator, and reading its + // absence as "not CRLF" is what joined a CRLF entry to its inserted line + // with a bare `\n`. `crlfAtEof` answers the weaker question that remains, + // over the entry's own SECTION when one is delimited and the entry list + // alone when none is — widening to everything before the insertion point + // only where that region is empty, which is #3773's B4. See its doc comment. + const spanEnd = matchIndexInContent + (end - start); + const eofScope = sectionBody.slice(deferredSection ? 0 : entries[0].start, start) + || sectionBody.slice(0, start); + const crlf = matchedLines.length > 1 + ? line0HadCr + : content.startsWith('\r\n', spanEnd) + || (spanEnd >= content.length && crlfAtEof(eofScope)); newMatchedLines = [ - matchedLines[0], - `${continuationIndent}status: acknowledged`, + crlf ? `${matchedLines[0].replace(/\r$/, '')}\r` : matchedLines[0], + `${continuationIndent}status: acknowledged${line0HadCr ? '\r' : ''}`, ...matchedLines.slice(1), ]; } else { - const original = matchedLines[statusLineIdx]; - const replaced = original.replace( - /^(\s*(?:-\s+)?)(\*+status:\*+|status:)(\s*).*$/i, - (_m, indent: string, key: string, ws: string) => `${indent}${key}${ws}acknowledged`, - ); + // Rewrite at the offset the CLASSIFIER reported, rather than through a + // second regex of the writer's own. This is what makes the selection and + // the rewrite structurally incapable of disagreeing: a status line the + // classifier can select is one whose value offset it has already + // computed, so there is no shape it selects and then fails to rewrite. + // (A widened classifier over a hyphen-only rewrite regex is exactly that + // failure — it would select `* status: open` and hand back the line + // untouched.) The key's own spelling and any `**bold**` wrapper survive + // because only the value is replaced. + const raw = matchedLines[statusLineIdx]; + const cr = raw.endsWith('\r') ? '\r' : ''; + const line = raw.slice(0, raw.length - cr.length); + const field = parseGapEntryFieldLine(line, DEFERRED_BULLET_MARKERS, statusLineIdx === 0)!; + const prefix = line.slice(0, field.valueStart); + // `status:acknowledged` reads back fine, but a bare colon with no + // separator is not what this file's convention looks like; supply one only + // when the source had none. + const sep = /[ \t]$/.test(prefix) ? '' : ' '; newMatchedLines = matchedLines.slice(); - newMatchedLines[statusLineIdx] = replaced; + newMatchedLines[statusLineIdx] = `${prefix}${sep}acknowledged${cr}`; } + // NO post-write read-back guard here, deliberately (round 4, B3). Round 3 + // added one — `rewrite_not_readable` — after a fenced `status:` line proved + // the writer could select a line the reader would not read back. Round 3 + // then closed that divergence STRUCTURALLY, by routing the writer's line + // selection and the reader's field extraction through the one + // `entryFieldLines` seam above, and the guard became unreachable from the + // public API: 21 document shapes were driven against it (fence openers on + // the bullet line for every marker, duplicate and triplicate `status:` + // lines, bolded and nested variants, fences between duplicates) and none + // reached it. + // + // An unreachable branch is not free here. This repo's own + // `RULESET.TESTS.mutation-score` runs Stryker incrementally over changed + // files at an 80% threshold and says to "treat surviving mutant as a failing + // test specification"; an undriven `if` is exactly that. The only seam that + // would drive it is routing this call through the module's exports so a test + // could stub it — production surface reshaped for a test, which is a worse + // trade than the guard is worth now that construction, not assertion, + // enforces the invariant. + // + // What that gives up, stated plainly rather than hidden: if a future change + // re-splits the writer's selection from the reader's extraction, this + // function returns `ok` over an item that stays outstanding — the original + // #3702 defect class. `match_verification_failed` does NOT backfill it; that + // check runs BEFORE the write and compares the matched span to the target, + // so it cannot see a post-write read-back failure. The protection against + // re-splitting is the shared seam plus the round-3 tests that pin it, not a + // runtime assertion. + const newContent = content.slice(0, matchIndexInContent) + newMatchedLines.join('\n') + content.slice(matchIndexInContent + (end - start)); return { content: newContent, status: 'ok' }; } -/** - * #3781 — one entry of the heading-delimited deferred shape, carrying the - * exact character span it occupies within the `sectionBody` it was derived - * from. `lines[0]` of a LEAF entry is the RAW heading line (hashes intact) so - * `sectionBody.slice(start, end)` is byte-verbatim; `readerLines` is the - * reader-form of the same lines (ATX-stripped line 0, bullet-stripped body) - * the identity text and field extraction are computed over. `embeddedTable` - * marks entries whose span contains a GFM table line — table lines are - * excluded from entries, so such a span is non-contiguous and its entries - * refuse rather than risk a wrong write. - */ -interface DeferredHeadingEntrySpan { - kind: 'leaf' | 'pending'; - lines: string[]; - readerLines: string[]; - text: string; - fields: Record; - start: number; - end: number; - embeddedTable: boolean; -} - /** * #3781 — strip an ATX heading prefix, mirroring `tokenizeHeadings`' own ATX * regex (≤3 leading spaces, 1–6 `#`, space/tab separator, optional closing @@ -2183,133 +2275,6 @@ function stripAtxPrefix(line: string): string | null { : m[3].replace(/^[ \t]+/, '').replace(/[ \t]+#+[ \t]*$/, '').replace(/^#+[ \t]*$/, '').trim(); } -/** - * #3781 — span-carrying sibling of `splitDeferredHeadingEntries`: ONE walk, - * identical grouping rules (leaf = childless heading whose body carries a - * bullet; container = next heading deeper; preamble/container-direct lines → - * headless entries; table lines excluded), additionally recording each - * entry's (start, end) character span within `sectionBody`. Returns null when - * the body contains no heading at all — the caller then takes the unchanged - * fully-headless path. - */ -function splitDeferredHeadingEntriesWithSpans(sectionBody: string): DeferredHeadingEntrySpan[] | null { - const headings = tokenizeHeadings(sectionBody); - if (headings.length === 0) return null; - - const lines = sectionBody.split('\n'); - const lineStarts: number[] = []; - const lineEnds: number[] = []; - let cursor = 0; - for (const rawLine of lines) { - lineStarts.push(cursor); - cursor += rawLine.length; - lineEnds.push(cursor); - cursor += 1; - } - - const headingByLine = new Map(); - for (let i = 0; i < headings.length; i++) { - const isContainer = i + 1 < headings.length && headings[i + 1].level > headings[i].level; - headingByLine.set(headings[i].line, { text: headings[i].text, isContainer }); - } - - const isTableLine = (l: string): boolean => /^\s*\|/.test(l.replace(/\r$/, '')); - const isBulletLine = (l: string): boolean => /^\s*-\s/.test(l.replace(/\r$/, '')); - - const entries: DeferredHeadingEntrySpan[] = []; - let current: string[] | null = null; - let currentReaderLine0: string | null = null; - let currentStartLine = -1; - let currentEndLine = -1; - let currentHasBullet = false; - let currentTable = false; - let pendingStartLine = -1; - let pendingEndLine = -1; - - const flushCurrent = (): void => { - if (current !== null && currentHasBullet && currentReaderLine0 !== null) { - const bodyReader = current.slice(1).map(stripLeadingBulletMarker); - entries.push({ - kind: 'leaf', - lines: current, - readerLines: [currentReaderLine0, ...bodyReader], - text: rawGapEntryText([currentReaderLine0, ...current.slice(1)]), - fields: extractGapEntryFields([currentReaderLine0, ...bodyReader]), - start: lineStarts[currentStartLine], - end: lineEnds[currentEndLine], - embeddedTable: currentTable, - }); - } - current = null; - currentReaderLine0 = null; - currentStartLine = -1; - currentEndLine = -1; - currentHasBullet = false; - currentTable = false; - }; - const flushPending = (): void => { - if (pendingStartLine === -1) return; - // The pending region is contiguous (a heading flushes it), but table lines - // inside it were skipped by the walk: the reader's identity for this - // region is computed over the table-FILTERED join, which may merge - // entries across the gap, so spans cannot be translated faithfully — - // mark the region's entries as refusing instead. - let regionTable = false; - for (let i = pendingStartLine; i <= pendingEndLine; i++) { - if (isTableLine(lines[i])) regionTable = true; - } - const base = lineStarts[pendingStartLine]; - const regionText = sectionBody.slice(lineStarts[pendingStartLine], lineEnds[pendingEndLine]); - for (const e of splitGapsEntriesWithSpans(regionText)) { - entries.push({ - kind: 'pending', - lines: e.lines, - readerLines: e.lines, - text: rawGapEntryText(e.lines), - fields: extractGapEntryFields(e.lines), - start: base + e.start, - end: base + e.end, - embeddedTable: regionTable, - }); - } - pendingStartLine = -1; - pendingEndLine = -1; - }; - - for (let i = 0; i < lines.length; i++) { - const heading = headingByLine.get(i + 1); - if (heading !== undefined) { - flushCurrent(); - flushPending(); - if (!heading.isContainer) { - current = [lines[i]]; - currentReaderLine0 = heading.text; - currentStartLine = i; - currentEndLine = i; - currentHasBullet = false; - currentTable = false; - } - continue; - } - if (isTableLine(lines[i])) { - if (current !== null) currentTable = true; - continue; - } - if (current !== null) { - current.push(lines[i]); - currentEndLine = i; - if (isBulletLine(lines[i])) currentHasBullet = true; - } else { - if (pendingStartLine === -1) pendingStartLine = i; - pendingEndLine = i; - } - } - flushCurrent(); - flushPending(); - - return entries; -} - /** * #3781 — the heading-shaped half of `acknowledgeDeferredItem`, sharing the * headless path's guards (not_found / ambiguous / already_resolved / @@ -2317,58 +2282,60 @@ function splitDeferredHeadingEntriesWithSpans(sectionBody: string): DeferredHead * shape-specific rules documented on `acknowledgeDeferredItem` (reader-form * status search incl. the leaf line-0 ATX corner; insert after the entry's * last non-blank line). Extracted so the headless path stays byte-identical. + * + * Every question about an entry is asked of the READER'S OWN answer (#3702 + * round 3, B1/B3 — restated here rather than re-implemented): identity is + * `rawGapEntryText` over the walk's lines and opener flags, exactly as + * `parseDeferredItemsWithStatus` names the entry; the status line is whichever + * line `entryFieldLines` classifies as `status`, so a fenced `status:` or a + * rejected-ordinal prose line is never selected; and the rewrite lands at the + * offset the classifier reported. Upstream's #3781 carried its own + * hyphen-only walk and its own status regexes for this shape; under the + * widened marker grammar those would have read a different entry set than + * the reader and re-opened the writer/reader drift this PR closes. */ function acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, headingEntries, targetText }: { content: string; sectionBody: string; deferredSection: { body: string; bodyStart: number } | null; - headingEntries: DeferredHeadingEntrySpan[]; + headingEntries: DeferredHeadingEntry[]; targetText: string; }): AcknowledgeDeferredItemResult { - const matches = headingEntries.filter((e) => e.text === targetText); + const matches = headingEntries.filter( + (e) => rawGapEntryText(e.lines, DEFERRED_BULLET_MARKERS, e.opener) === targetText, + ); if (matches.length === 0) return { content, status: 'not_found' }; if (matches.length > 1) return { content, status: 'ambiguous' }; const entry = matches[0]; + // A table row inside the span: the walk skipped it, so `lines` is not 1:1 + // with the raw slice and no write can be anchored. Refuse, as before #3781. if (entry.embeddedTable) return { content, status: 'unsupported_heading_shape' }; - if (entry.fields.status && entry.fields.status.toLowerCase() === 'resolved') { + const fields = extractGapEntryFields(entry.lines, DEFERRED_BULLET_MARKERS, entry.opener); + if (fields.status && fields.status.toLowerCase() === 'resolved') { return { content, status: 'already_resolved' }; } const sectionOffset = deferredSection ? deferredSection.bodyStart : 0; - const rawSlice = sectionBody.slice(entry.start, entry.end); - const rawSliceLines = rawSlice.split('\n'); - - // Genuine invariant re-verification: re-derive the entry's identity from - // the span's own bytes and compare against the targetText that selected it - // — the span was recorded by an offset bookkeeping independent of the - // identity comparison above. - const verifyText = entry.kind === 'leaf' - ? (() => { - const stripped = stripAtxPrefix(rawSliceLines[0]); - return stripped === null ? null : rawGapEntryText([stripped, ...rawSliceLines.slice(1)]); - })() - : rawGapEntryText(rawSliceLines.map((l) => l.replace(/\r$/, ''))); - if (verifyText !== targetText) { + const rawLines = sectionBody.slice(entry.start, entry.end).split('\n'); + // The reader-form of the raw slice, index-aligned with it: a leaf's line 0 + // is the heading TEXT (re-derived from the span's own bytes, not copied + // from the walk, so the verification below is genuine), every other line + // CR-stripped. Markers stay on the lines — the classifier strips them per + // the opener flags, exactly as the reader does. + const readerLines = rawLines.map((raw, i) => { + const line = raw.replace(/\r$/, ''); + return i === 0 && entry.kind === 'leaf' ? (stripAtxPrefix(line) ?? line) : line; + }); + // Genuine invariant re-verification: the identity re-derived from the + // span's bytes must be the identity that selected the entry — the span was + // recorded by offset bookkeeping independent of that comparison. + if (readerLines.length !== entry.lines.length + || rawGapEntryText(readerLines, DEFERRED_BULLET_MARKERS, entry.opener) !== targetText) { return { content, status: 'match_verification_failed' }; } - // Status search over the READER-form lines — the exact set the reader - // parses (bolded any case + bare lowercase, per #3775; line-0 forms per - // #3740). Reader lines are index-aligned 1:1 with the raw lines. - const statusFieldBoldedRe = /^\s*\*+status:\*+/i; - const statusFieldBareRe = /^\s*status:/; - const statusFieldBoldedReLine0 = /^\s*(?:-\s+)?\*+status:\*+/i; - const statusFieldBareReLine0 = /^\s*(?:-\s+)?status:/; - const readerLines = entry.kind === 'leaf' - ? [ - stripAtxPrefix(rawSliceLines[0]) ?? rawSliceLines[0].replace(/\r$/, ''), - ...rawSliceLines.slice(1).map((l) => stripLeadingBulletMarker(l.replace(/\r$/, ''))), - ] - : rawSliceLines.map((l) => l.replace(/\r$/, '')); - const statusLineIdx = readerLines.findIndex((line, idx) => - idx === 0 - ? (statusFieldBoldedReLine0.test(line) || statusFieldBareReLine0.test(line)) - : (statusFieldBoldedRe.test(line) || statusFieldBareRe.test(line))); + const statusLineIdx = entryFieldLines(readerLines, DEFERRED_BULLET_MARKERS, entry.opener) + .findIndex((field) => field?.key === 'status'); let newRawLines: string[]; if (statusLineIdx === -1) { @@ -2376,43 +2343,53 @@ function acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, // entry's body is frequently a soft-wrapped sentence, and splicing after // line 0 would split it in half (#3781's sentence trap). The headless // (no-heading-anywhere) path keeps its own splice-after-line-0 shape. - let lastNonBlank = rawSliceLines.length - 1; - while (lastNonBlank > 0 && rawSliceLines[lastNonBlank].replace(/\r$/, '').trim() === '') { - lastNonBlank--; - } + // …and never INSIDE a fence (round 5, RV6.5 review): an entry whose body + // ends in a fenced block — closed or, worse, unclosed and so running to + // the entry's end — would otherwise receive its marker as fence content, + // a line the reader never reads: `ok` returned, item still outstanding. + // Walk back over blank and fenced lines alike, classified exactly as the + // reader classifies them, so the marker lands on a line the reader reads. + const fencedInEntry = entryFencedLines(readerLines, DEFERRED_BULLET_MARKERS, entry.opener); + let last = rawLines.length - 1; + while (last > 0 && (readerLines[last].trim() === '' || fencedInEntry.has(last))) last--; + // A pending entry's continuation sits two columns inside its own marker + // indent (the entry's indent CHARACTERS, as the headless path does); a + // leaf's body lines are sibling bullets, and an indented bare field line + // among them is what the reader reads on that shape. const indent = entry.kind === 'pending' - ? (() => { - const bulletIndentMatch = rawSliceLines[0].match(/^(\s*)-\s+/); - return ' '.repeat((bulletIndentMatch ? bulletIndentMatch[1].length : 0) + 2); - })() + ? `${rawLines[0].replace(/\r$/, '').match(DEFERRED_BULLET_MARKERS.strip)?.[1] ?? ''} ` : ' '; - newRawLines = [ - ...rawSliceLines.slice(0, lastNonBlank + 1), - `${indent}status: acknowledged`, - ...rawSliceLines.slice(lastNonBlank + 1), - ]; + // The inserted line copies the ending of the line it follows. When that + // line is the span's LAST, its terminator sits outside the span: the + // separator following the span decides, else (end of file) the section's + // own evidence — the same rule the headless path applies to line 0. + const followsLast = last === rawLines.length - 1; + const spanEnd = sectionOffset + entry.end; + const prevCr = followsLast + ? content.startsWith('\r\n', spanEnd) || (spanEnd >= content.length && crlfAtEof(sectionBody.slice(0, entry.start))) + : rawLines[last].endsWith('\r'); + newRawLines = rawLines.slice(); + if (followsLast && prevCr) newRawLines[last] = `${rawLines[last].replace(/\r$/, '')}\r`; + newRawLines.splice(last + 1, 0, `${indent}status: acknowledged${!followsLast && prevCr ? '\r' : ''}`); } else { - newRawLines = rawSliceLines.slice(); - if (entry.kind === 'leaf' && statusLineIdx === 0) { - // Leaf line 0 is the RAW heading line — rewrite the heading-text portion - // the reader treats as a field, with the ATX prefix preserved. - const replacedReader = readerLines[statusLineIdx].replace( - /^(\s*(?:-\s+)?)(\*+status:\*+|status:)(\s*).*$/i, - (_m, indent: string, key: string, ws: string) => `${indent}${key}${ws}acknowledged`, - ); - const atxMatch = /^(\s*#+[ \t]*)(.*)$/.exec(rawSliceLines[0].replace(/\r$/, '')); - newRawLines[0] = atxMatch ? atxMatch[1] + replacedReader : replacedReader; - } else { - // Every other status line is rewritten on its RAW line, so the bullet - // marker and indent survive the write (the reader-form line has the - // marker stripped — writing it back would mangle the markdown shape and - // change the entry's identity text). Same replacement regex as the - // headless path. - newRawLines[statusLineIdx] = rawSliceLines[statusLineIdx].replace( - /^(\s*(?:-\s+)?)(\*+status:\*+|status:)(\s*).*$/i, - (_m, indent: string, key: string, ws: string) => `${indent}${key}${ws}acknowledged`, - ); - } + // Rewrite at the offset the CLASSIFIER reported, on the RAW line — the + // marker, the indent, the key's spelling and any `**bold**` wrapper all + // survive because only the value is replaced. A leaf's line 0 is the + // heading line, so its ATX prefix is put back in front of the rewritten + // text (the reader reads the heading text itself as the field there). + const raw = rawLines[statusLineIdx]; + const cr = raw.endsWith('\r') ? '\r' : ''; + const line = raw.slice(0, raw.length - cr.length); + const reader = readerLines[statusLineIdx]; + const field = parseGapEntryFieldLine(reader, DEFERRED_BULLET_MARKERS, stripsMarkerAt(statusLineIdx, entry.opener))!; + const prefix = reader.slice(0, field.valueStart); + const sep = /[ \t]$/.test(prefix) ? '' : ' '; + const leafLine0 = statusLineIdx === 0 && entry.kind === 'leaf'; + const atx = leafLine0 ? (/^( {0,3}#{1,6}[ \t]+)/.exec(line)?.[1] ?? '') : ''; + // A closing `#` sequence is Markdown the reader ignores; keep it (RV6.5). + const closing = leafLine0 ? (/[ \t]+#+[ \t]*$/.exec(line)?.[0] ?? '') : ''; + newRawLines = rawLines.slice(); + newRawLines[statusLineIdx] = `${atx}${prefix}${sep}acknowledged${closing}${cr}`; } const matchIndexInContent = sectionOffset + entry.start; @@ -2421,14 +2398,334 @@ function acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, } /** - * Strip one leading `- ` bullet marker (#3457). Heading-delimited deferred - * entries carry their fields as sibling bullets; `extractGapEntryFields` only - * de-bullets line 0 (Gaps-protective — there, a later `- ` line is a nested - * sub-list), so the deferred heading path de-bullets every line itself before - * field extraction. Non-bullet lines pass through untouched. + * The two regex shapes a bullet-aware walk needs over one marker set: `open` + * detects an entry-opening marker, `strip` additionally captures the line's + * remaining content. Carrying both in ONE object is what keeps a detection + * site and its matching strip site from drifting to different marker sets — + * the asymmetry #3702 warns about (widen what OPENS an entry without widening + * what is STRIPPED before field extraction, and a status field written under + * the new marker never resolves its entry). */ -function stripLeadingBulletMarker(line: string): string { - return line.replace(/^(\s*)-\s+/, ''); +interface BulletMarkers { + /** `^(indent)()\s` — group 1 is the indent, group 2 the marker token. */ + open: RegExp; + /** `^(indent)\s+(rest)$` — group 1 indent, group 2 rest-of-line. */ + strip: RegExp; + /** + * Honour Markdown block structure — thematic breaks close the list, fenced + * code neither opens an entry nor carries a field (#3702 round 2, M1/M2), + * and indents are measured in CommonMark columns rather than raw characters. + * `false` for the Gaps grammar, which #3702's ruling keeps byte-for-byte on + * its `next` behaviour; `true` for the deferred-items grammar. Carried on the + * marker set rather than as a second parameter so a caller cannot widen one + * without the other. + */ + blockStructure: boolean; +} + +/** + * Hyphen-only markers — the `## Gaps` form, unchanged by #3702. Gaps entries + * come from a template that mandates the hyphen YAML-lite shape, so widening + * that section's grammar is not what the deferred-items ruling required; the + * shared splitting seam is parameterised rather than widened wholesale so the + * Gaps path stays byte-for-byte on its existing behaviour. + */ +const HYPHEN_BULLET_MARKERS: BulletMarkers = { + open: /^(\s*)(-)\s/, + strip: /^(\s*)-\s+(.*)$/, + blockStructure: false, +}; + +/** + * Deferred-items markers (#3702): the standard Markdown list markers, not the + * hyphen alone. `deferred-items.md` has NO template and no mandated shape — + * executors write it by hand (the same premise that justified the #2766 table + * union) — so an author reaching for `*`, `+` or `1.` wrote a list by every + * Markdown definition while this parser contributed ZERO entries for it. The + * hyphen restriction was a regex literal inherited from the Gaps seam, never a + * stated decision: measured in the wild, non-empty records parsed to a clean + * zero, and a MIXED file dropped its non-hyphen entries while keeping their + * hyphenated siblings — under-reporting without ever looking empty. + * + * Deliberately NOT widened to prose: "prose is not an item" is this parser's + * pre-existing, test-asserted contract (the `# Notes` case) and is untouched + * here. An asterisk bullet is not prose, and a `|` row is not a list marker — + * table lines are still skipped before the body-bullet flag can be set, so + * `parseDeferredTableItems` keeps sole ownership of table bodies and the + * #2766 anti-double-count property holds unchanged. + * + * The paren-terminated ordered form (`1)`) is out of scope for this fix: the + * #3702 ruling scopes the widening to `*`, `+` and the dot-terminated ordered + * marker. + * + * `DEFERRED_MARKER_ALT` is THE source every deferred-items marker regex is + * built from (#3702 round 2, M3). Since round 3 that is the splitter's + * `open`/`strip` pair here and nothing else: `acknowledgeDeferredItem`'s two + * status-line shapes used to be derived from it too, and are now deleted in + * favour of the reader's classifier. CommonMark + * §5.2: bullet markers `-`, `*`, `+`; an ordered marker is 1-9 digits and a + * `.` (round-1's `\d+` was uncapped). The marker is followed by a space or a + * tab — `[ \t]`, where round 1 wrote `\s`, which also accepted `\r`. + * `markdown-sectionizer`'s `iterateBullets` is the repo's other list-marker + * grammar; the `#3702 round 2: marker-grammar parity` test pins this one to + * it on the shared vocabulary and names the two points they deliberately + * differ (tab after the marker, the 9-digit cap). + */ +const DEFERRED_MARKER_ALT = '(?:[-*+]|\\d{1,9}\\.)'; +const DEFERRED_BULLET_MARKERS: BulletMarkers = { + open: new RegExp(`^(\\s*)(${DEFERRED_MARKER_ALT})[ \\t]`), + strip: new RegExp(`^(\\s*)${DEFERRED_MARKER_ALT}[ \\t]+(.*?)\\r?$`), + blockStructure: true, +}; + +// `acknowledgeDeferredItem` carries NO status-line regex of its own (#3702 +// round 3, B1/B3). It used to hold two — a finder and a rewrite — derived +// from `DEFERRED_MARKER_ALT` so the two WRITER shapes could not drift from +// each other. That kept the wrong pair in step: the finder's peer is the +// READER, and widening detection without widening the read is what made a +// nested ` * status:` line selectable by the writer and invisible to +// `extractGapEntryFields`. Both are gone; the writer now locates its line +// through `parseGapEntryFieldLine`, the reader's own classifier, and rewrites +// at the offset that classifier reports. See `parseGapEntryFieldLine`. + +/** + * CommonMark §4.1 thematic break: up to 3 spaces of indent, then three or + * more of the SAME `-`, `*` or `_`, optionally space/tab-separated, and + * nothing else. `- - -`, `* * *` and `+ + +` all also match a list opener — + * `- - -` was a phantom `"- -"` entry on base, and #3702's widening added the + * other two (#3702 round 2, M1). `+ + +` is not a CommonMark break, but it is + * the same authoring gesture and no less garbage as an entry name, so the + * class here is "three-or-more of one marker character, nothing else". The + * indent is unbounded, not CommonMark's `{0,3}`: this parser reads a list at + * any indent (see `#3702 round 2` m1), so a separator drawn at any indent is + * a separator too — otherwise ` * * *` is a phantom entry named `* *`. + */ +const THEMATIC_BREAK_RE = /^[ \t]*([-*+_])(?:[ \t]*\1){2,}[ \t]*$/; + +/** + * The deferred grammar's line view for fence classification: the SAME lines, + * with leading whitespace removed (#3702 round 4, M2). + * + * `scanFencedBlocks` is CommonMark, and CommonMark caps a fence delimiter's + * indent at three spaces — a fourth makes it an indented code block instead. + * The deferred grammar deliberately opted out of that cliff everywhere else: + * an entry opener is `[ \t]*`-indented and `THEMATIC_BREAK_RE` is + * `^[ \t]*`. Leaving the fence rule at CommonMark's cap while items and + * breaks are unbounded is not a conservative choice, it is an inconsistent + * one, and it is REACHED BY ORDINARY DOCUMENTS: a fenced block written under a + * nested bullet sits at four spaces, so its `status: resolved` line resolved + * the entry containing it. That is the #3702 silent-resolution defect class in + * a new place — driven, at indents 4, 5, 8 and a leading tab, before this fix. + * + * Still NO second fence dialect (the rule `blankIndentedFenceDelimiters` + * states): the classification is done by `scanFencedBlocks`, the one exported + * CommonMark state machine, over a de-indented view. Run lengths, backtick + * vs tilde, closer-must-match-and-not-trail and info-string rules are all + * still that engine's answers, not re-derived here; the unterminated case is + * its answer too, bounded by the walk (round 5, B1 — see `scanFencesFrom`). Indent is the only dimension this hides from it, and it is + * the exact dimension the deferred grammar has already declared it does not + * measure. Index alignment is 1:1 by construction — `map` preserves length — + * so every line index the engine returns still addresses the original line. + * + * Scope: the deferred grammar only. Both marker-parameterised call sites gate + * on `markers.blockStructure`, which the `## Gaps` set does not set, so Gaps + * reaches an empty set and is untouched by this — the same opt-out + * `indentWidth` documents for the indent half. + */ +function deindentedForFences(lines: string[]): string[] { + return lines.map((line) => line.replace(/^[ \t]+/, '')); +} + +/** + * Indices (into `lines`) of every line that sits inside a fenced code block, + * delimiters included — by the sectionizer's own fence state machine, so a + * `~~~` fence, an indented fence and an unterminated fence (runs to the end) + * are classified exactly as `stripFencedCode` would (#3702 round 2, M2), at + * ANY indent (round 4, M2 — see `deindentedForFences`). + * #3702's wild records carry reproduction blocks; `+`-prefixed diff lines and + * `1.`-numbered steps are their normal content, not entries. + * + * ENTRY-scoped: `lines` are ONE entry's lines (`entryFieldLines`), so an + * unterminated fence "running to the end" runs to the end of that entry — + * exactly the bound the section-level walks give it (round 5, B1; see + * `scanFencesFrom`). The two classifications agree by construction. + */ +function fencedLineSet(lines: string[]): Set { + const fenced = new Set(); + for (const block of scanFencedBlocks(deindentedForFences(lines))) { + const last = block.closeLineIdx === -1 ? lines.length - 1 : block.closeLineIdx; + for (let i = block.openLineIdx; i <= last; i++) fenced.add(i); + } + return fenced; +} + +/** + * Fence classification for a SECTION-level walk (round 5, B1). Terminated + * fences are classified wholesale, as `fencedLineSet` does. The UNTERMINATED + * fence is the difference: `scanFencedBlocks` runs it to the end of the + * document, and so did round 4 — at any indent, since round 4's M2 — so one + * stray delimiter swallowed every entry after it into the entry before it. + * `- a` / blank / four-space ``` / blank / `- b` yielded ONE entry where + * `next` yields two: a widening that made an already-counted item vanish, + * in the direction of silent data loss, on the exact mixed-file shape #3702 + * exists to close. + * + * The rule now: an unterminated fence runs to the end of its ENTRY, never + * past it. That is CommonMark's own answer for a fence inside a list item — + * the fence closes when its container does, and the container closes at the + * next item at its level — extended to a document-level stray delimiter, + * where CommonMark would swallow to end-of-file and this parser's fail-safe + * rule (surface a questionable entry rather than drop a real one) will not. + * The scan therefore reports the unterminated opener's index and leaves the + * bound to the walk, which knows what a top-level item is: on reaching one, + * the walk RESCANS from that line, so a delimiter after the bound is read on + * its own terms rather than as the content of a fence that has ended. Still + * no second fence dialect — every block boundary is `scanFencedBlocks`' + * answer; the walk contributes only the bound. + */ +interface FenceScan { + /** Lines inside a TERMINATED fence, delimiters included. */ + fenced: Set; + /** Every fence opener, terminated or not — an opener ends the list runs at its level, as a paragraph does. */ + openers: Set; + /** Opener index of the one unterminated fence in scope, else -1. Lines from it onward are fenced until the walk bounds it. */ + unterminatedFrom: number; +} + +function scanFencesFrom(lines: string[], from: number): FenceScan { + const scan: FenceScan = { fenced: new Set(), openers: new Set(), unterminatedFrom: -1 }; + for (const block of scanFencedBlocks(deindentedForFences(lines.slice(from)))) { + const open = block.openLineIdx + from; + scan.openers.add(open); + if (block.closeLineIdx === -1) { + scan.unterminatedFrom = open; // always the scan's last block + break; + } + for (let i = open; i <= block.closeLineIdx + from; i++) scan.fenced.add(i); + } + return scan; +} + +/** The scan a grammar without block structure (`## Gaps`) walks under: nothing is fenced. Never mutated. */ +const NO_FENCES: FenceScan = { fenced: new Set(), openers: new Set(), unterminatedFrom: -1 }; + +/** + * Does `line` LOOK like a top-level list item under `markers` — a marker at + * or above the base indent, start value ignored? The bound an unterminated + * fence runs to (round 5, B1; see `scanFencesFrom`). Shape rather than the + * ordered-start rule, because the list memory inside a fence is not evidence + * of anything, and closing a stray fence one line early errs in the + * surfacing direction. + */ +function topLevelItemShape(line: string, markers: BulletMarkers, baseIndent: number | null): boolean { + const m = line.match(markers.open); + return m !== null && (baseIndent === null || indentWidth(m[1], markers) <= baseIndent); +} + +/** + * Per-indent LIST memory (#3702 round 2, round review; widened round 5, M2): + * each list level remembers whether a list is OPEN there, so an ordered + * marker that does not start at `0.`/`1.` is an item when it continues or + * follows a list at its level, and prose otherwise. A new opener at indent + * `d` resets every deeper level (a new item starts new sub-lists); a + * paragraph after a blank at indent `d` ends the lists at `d` and deeper; a + * thematic break or a heading clears everything. + * + * Round 2 keyed this on whether the previous opener was ORDERED, so a bullet + * item closed the run and `1. a` / `- b` / `5. c` folded `5. c` into `b` — + * the mixed-file under-report #3702 names as the shape that bites. In + * CommonMark `5. c` there opens a fresh ordered list (`start=5`): a non-1 + * ordinal is refused only where it would INTERRUPT A PARAGRAPH (§5.3), and + * after a list item it interrupts nothing. Keying on "a list is open here" + * is that rule as far as this parser can state it without a paragraph model. + */ +class ListRuns { + private readonly byIndent = new Set(); + at(indent: number): boolean { return this.byIndent.has(indent); } + opened(indent: number): void { + for (const d of [...this.byIndent]) if (d > indent) this.byIndent.delete(d); + this.byIndent.add(indent); + } + endedAt(indent: number): void { + for (const d of [...this.byIndent]) if (d >= indent) this.byIndent.delete(d); + } + clear(): void { this.byIndent.clear(); } +} + +/** + * Leading-whitespace width of a line in CommonMark COLUMNS (§2.2: a tab + * advances to the next multiple of 4), so `\t` and ` ` are different levels + * and `\t` equals four spaces — character counting aliased them. + */ +/** + * Indent WIDTH under a grammar (#3702 round 2, review round 6). The deferred + * grammar measures CommonMark columns; the Gaps grammar keeps `next`'s raw + * character count, because its `blockStructure: false` opt-out promises + * byte-for-byte parity and a column measure silently breaks it — a + * tab-indented Gaps item followed by a two-space one split into two entries + * where `next` folded them into one, and the reverse pair folded where `next` + * split. The opt-out now covers indent semantics, not only fences and breaks. + */ +function indentWidth(indent: string, markers: BulletMarkers): number { + return markers.blockStructure ? indentOf(indent) : indent.length; +} + +function indentOf(line: string): number { + let col = 0; + for (const ch of line) { + if (ch === ' ') col += 1; + else if (ch === '\t') col += 4 - (col % 4); + else break; + } + return col; +} + +/** A line classified as a list-item opener by `matchListOpener`. */ +interface ListOpener { + /** Width of the indent before the marker. */ + indent: number; +} + +/** + * Classify `line` as a list-item opener under `markers`, applying the + * ORDERED-START rule (#3702 round 2, B2; round 5, M1/M2): a dot-terminated + * ordered marker opens an item when it starts at `0.` or `1.` (`01.` + * included), or when a list is already open at its level (`inList` — the + * caller's per-indent memory). + * + * Why: `\d{1,9}\.` alone reads ordinary prose as a list. "2026. was a bad + * year for this module" and, under a `### Notes` heading, "3. is the number + * of retries we settled on." are both sentences, and both opened an entry on + * round 1 — the second one straight through the "prose is not an item" + * contract that round claimed to preserve. CommonMark §5.3 faces the same + * ambiguity when an ordered list would interrupt a paragraph and resolves it + * the same way: the list must start with 1. This parser has no paragraph + * model, so it applies that rule wherever NO list is open at the line's + * level — the positions a sentence can occupy. Where a list IS open, + * CommonMark accepts any start (a list item interrupts no paragraph), and so + * does this. Numbers after the first are ignored, as CommonMark ignores + * them, so `1. / 3. / 7.` is a three-item run. + * + * `0.` is accepted as a start (round 5, M1): CommonMark §5.2 permits any + * 1-9-digit start number and a `0.`-numbered list is ordinary; refusing it + * dropped ONLY the first item, since the run then started at `1.` — the + * under-report that looks like a clean parse. A sentence opening with "0." + * is not a shape anyone writes. + * + * Stated cost, pinned by test: a list whose first ordinal is 2 or more, at a + * paragraph position, reads as prose UNTIL its first `0.`/`1.` line — the + * loss is that prefix, not the whole list. Every ordered record the #3702 + * scan found starts at 1, and the hyphen-style `- ` alternative loses + * nothing, so the trade buys the prose contract back at no measured cost. + * + * Bullet markers carry no rule — an asterisk bullet is not prose. + */ +function matchListOpener(line: string, markers: BulletMarkers, inList: boolean): ListOpener | null { + const m = line.match(markers.open); + if (!m) return null; + const token = m[2]; + if (/^\d/.test(token) && !inList && parseInt(token, 10) > 1) return null; + return { indent: indentWidth(m[1], markers) }; } /** @@ -2450,8 +2747,10 @@ function stripLeadingBulletMarker(line: string): string { * level; deepest level) each mis-count one of these shapes. * * A leaf entry is [heading text, ...body lines up to the next heading] and is - * kept only when its body (minus table lines) contains at least one `- ` - * bullet: + * kept only when its body (minus table lines) contains at least one list item + * — any of `-`, `*`, `+` or a dot-terminated ordered marker (#3702; the + * hyphen-only form silently dropped `*`/`+`/ordered bodies), where an + * ordered marker counts only under `matchListOpener`'s ordered-start rule: * - a prose-only or bare heading contributes nothing — "prose is not an item" * is this parser's pre-existing contract (see the `# Notes` case); * - a table-only body is left entirely to `parseDeferredTableItems`, which @@ -2463,11 +2762,55 @@ function stripLeadingBulletMarker(line: string): string { * `splitGapsEntries` — headless parity, so loose bullets before a later * heading group (the mixed shape) stay one item each. */ -function splitDeferredHeadingEntries(sectionBody: string): string[][] | null { +/** + * A heading-path entry: its lines, the splitter's per-line opener verdict, + * and — since #3781 — where it sits in the section body, so the writer can + * anchor a rewrite without a second walk. + */ +interface DeferredHeadingEntry { + lines: string[]; + /** `opener[i]` — line `i` opened a list item under `matchListOpener`'s rules. */ + opener: boolean[]; + /** `leaf` — a childless heading's entry, `lines[0]` its heading TEXT; `pending` — a headless bullet in the preamble or directly under a container heading. */ + kind: 'leaf' | 'pending'; + /** Character span within the section body — `sectionBody.slice(start, end)` is the entry's own raw text, CRLF intact, its lines 1:1 with `lines`. Both -1 when `embeddedTable`. */ + start: number; + end: number; + /** A GFM table row sat inside the span. Table lines belong to `parseDeferredTableItems` and are never entry lines, so the raw slice and `lines` disagree and no write can be anchored (#3781). */ + embeddedTable: boolean; +} + +/** Character offset of each line's start and end within the text they were split from (on `\n`). */ +function lineOffsets(lines: string[]): { lineStarts: number[]; lineEnds: number[] } { + const lineStarts: number[] = []; + const lineEnds: number[] = []; + let cursor = 0; + for (const line of lines) { + lineStarts.push(cursor); + cursor += line.length; + lineEnds.push(cursor); + cursor += 1; // the '\n' separator — absent after the final line, but nothing reads past it + } + return { lineStarts, lineEnds }; +} + +/** + * The heading-delimited split, carrying the per-line opener flags the deferred + * field-extraction path needs (#3702 round 2, round review): the heading path + * strips the marker off EVERY body line before field extraction (#3457), and a + * line whose ordinal `matchListOpener` REJECTED must not be stripped — or + * "3. status: resolved" as prose loses its `3. ` and reads as a resolved field. + * + * Since #3781 it also records each entry's character span (see + * `DeferredHeadingEntry`) in this same pass — the reader's walk IS the + * writer's walk, so there is no second copy of the grouping rules to drift. + */ +function splitDeferredHeadingEntriesDetailed(sectionBody: string): DeferredHeadingEntry[] | null { const headings = tokenizeHeadings(sectionBody); if (headings.length === 0) return null; const lines = sectionBody.split('\n'); + const { lineStarts, lineEnds } = lineOffsets(lines); const headingByLine = new Map(); for (let i = 0; i < headings.length; i++) { // Container iff the next heading is deeper (see doc comment). An empty @@ -2477,21 +2820,92 @@ function splitDeferredHeadingEntries(sectionBody: string): string[][] | null { headingByLine.set(headings[i].line, { text: headings[i].text, isContainer }); } - const entries: string[][] = []; - let current: string[] | null = null; // accumulating a leaf heading's entry - let pending: string[] = []; // preamble / container-heading body lines + const entries: DeferredHeadingEntry[] = []; + // The leaf entry being accumulated, with the raw line range it spans. + let current: { lines: string[]; opener: boolean[]; startLine: number; endLine: number } | null = null; let currentHasBullet = false; + // The headless-shaped region being accumulated (preamble / a container + // heading's direct lines): the reader's table-filtered, CR-stripped view, + // plus the raw line range it spans. + let pending: string[] = []; + let pendingStartLine = -1; + let pendingEndLine = -1; + // Table lines are never entry lines; where one sits INSIDE an entry's raw + // range, that entry's span is non-contiguous (#3781, `embeddedTable`). + const tableLines: number[] = []; + const tableWithin = (from: number, to: number): boolean => tableLines.some((t) => t >= from && t <= to); + // List memory for the leaf body being accumulated (#3702 round 2, B2) — + // reset at every heading, so `### Notes` + "3. is the number…" is prose + // while `### Steps` + "1. do / 2. then" is a list. A blank line then a + // non-indented non-list line is a PARAGRAPH, which ends the list + // (CommonMark §5.3); a non-indented line with no blank before it is lazy + // continuation and keeps it open. + const runs = new ListRuns(); + let blankSeen = false; + // Same level rule as the headless splitter: the first opener in a leaf body + // sets the base, and every indent at or shallower than it is one level. + let bodyBase: number | null = null; + const levelOf = (line: string): number => { + const ind = indentOf(line); + return bodyBase !== null && ind <= bodyBase ? bodyBase : ind; + }; + let scan = scanFencesFrom(lines, 0); const flushCurrent = (): void => { - // Keep the leaf entry only when its body carries a bullet; the heading + // Keep the leaf entry only when its body carries a list item; the heading // text line itself (element 0) never counts as one. - if (current !== null && currentHasBullet) entries.push(current); + if (current !== null && currentHasBullet) { + const table = tableWithin(current.startLine, current.endLine); + entries.push({ + lines: current.lines, + opener: current.opener, + kind: 'leaf', + start: table ? -1 : lineStarts[current.startLine], + end: table ? -1 : lineEnds[current.endLine], + embeddedTable: table, + }); + } current = null; currentHasBullet = false; }; const flushPending = (): void => { - entries.push(...splitGapsEntries(pending.join('\n'))); + if (pendingStartLine !== -1) { + // Headless-region entries carry the splitter's own opener flags — the + // same run state (ordered start, paragraph reset) that split them. The + // region is contiguous (a heading flushes it), so the core's + // region-relative spans translate by the region's own offset — unless a + // table row was skipped inside it, where the reader's view and the raw + // region disagree and no span is claimed. Both views split identically + // otherwise: the core CR-strips per line, and a table row is the only + // line the reader's view omits. + const table = tableWithin(pendingStartLine, pendingEndLine); + const region = table ? pending.join('\n') : lines.slice(pendingStartLine, pendingEndLine + 1).join('\n'); + const base = lineStarts[pendingStartLine]; + for (const { lines: entryLines, opener, start, end } of splitGapsEntriesCore(region, DEFERRED_BULLET_MARKERS)) { + entries.push({ + lines: entryLines, + opener, + kind: 'pending', + start: table ? -1 : base + start, + end: table ? -1 : base + end, + embeddedTable: table, + }); + } + } pending = []; + pendingStartLine = -1; + pendingEndLine = -1; + }; + const push = (line: string, i: number, opener: boolean): void => { + if (current !== null) { + current.lines.push(line); + current.opener.push(opener); + current.endLine = i; + } else { + pending.push(line); + if (pendingStartLine === -1) pendingStartLine = i; + pendingEndLine = i; + } }; for (let i = 0; i < lines.length; i++) { @@ -2503,20 +2917,69 @@ function splitDeferredHeadingEntries(sectionBody: string): string[][] | null { // ANY heading; flushing here keeps entries in document order even when // a container's direct bullets precede its first child entry. flushPending(); + runs.clear(); + blankSeen = false; + bodyBase = null; + // A heading ends the entry, and with it any unterminated fence (B1). + scan = scanFencesFrom(lines, i + 1); if (!heading.isContainer) { // Leaf heading: open an entry with the heading text as line 0. - current = [heading.text]; + current = { lines: [heading.text], opener: [false], startLine: i, endLine: i }; currentHasBullet = false; } continue; } + // CR-strip ONCE and carry the stripped line everywhere below — into the + // entry itself included (#3702 round 2, B1). `collectSection` slices raw + // `\n`-split lines, so on a CRLF file every body line but the last still + // carries its `\r`; the per-line marker strip feeding field extraction is + // `$`-anchored and fails on such a line, the marker survives into + // `extractGapEntryFields`, and the field is silently lost — a `**Status:**` + // that is not the file's final line then resurfaces its entry as open. + // The headless path (`splitGapsEntriesCore`) already stores stripped lines. + const line = lines[i].replace(/\r$/, ''); + // B1: an unterminated fence runs to the end of its entry. The next line + // shaped like a top-level item ends it — rescan from there, so a later + // delimiter is read on its own terms (see `scanFencesFrom`). + if (scan.unterminatedFrom !== -1 && i > scan.unterminatedFrom && topLevelItemShape(line, DEFERRED_BULLET_MARKERS, bodyBase)) { + scan = scanFencesFrom(lines, i); + } + if (scan.fenced.has(i) || (scan.unterminatedFrom !== -1 && i >= scan.unterminatedFrom)) { + // Fence content is body text, never list-item evidence (M2) — and + // never an opener, so it is never marker-stripped for fields either. + // Its opener ends the runs at its level and deeper, as a paragraph does. + if (scan.openers.has(i)) runs.endedAt(levelOf(line)); + push(line, i, false); + continue; + } + // A thematic break is a separator: not evidence, and it clears the list + // memory (M1). It stays a BODY line — the entry's span must stay + // contiguous for the writer, and the entry's name stays what `next` + // reported for a body containing one (round 5, m3). + if (THEMATIC_BREAK_RE.test(line)) { + runs.clear(); + blankSeen = false; + push(line, i, false); + continue; + } // Table lines belong to parseDeferredTableItems, never to a heading entry. - if (/^\s*\|/.test(lines[i].replace(/\r$/, ''))) continue; + if (/^\s*\|/.test(line)) { + tableLines.push(i); + continue; + } if (current !== null) { - current.push(lines[i]); - if (/^\s*-\s/.test(lines[i].replace(/\r$/, ''))) currentHasBullet = true; + const opener = matchListOpener(line, DEFERRED_BULLET_MARKERS, runs.at(levelOf(line))); + push(line, i, opener !== null); + if (opener !== null) { + currentHasBullet = true; + if (bodyBase === null) bodyBase = opener.indent; + runs.opened(levelOf(line)); + } else if (blankSeen && line.trim() !== '') { + runs.endedAt(levelOf(line)); // a paragraph after a blank line ends the lists at its level and deeper + } + blankSeen = line.trim() === ''; } else { - pending.push(lines[i]); + push(line, i, false); // the core derives its own opener verdicts for the region } } flushCurrent(); @@ -2575,6 +3038,8 @@ function parseDeferredTableItems(sectionBody: string): UatItem[] { */ interface GapsEntrySpan { lines: string[]; + /** `opener[i]` — line `i` was an ACCEPTED list opener (line 0 always; a nested one when accepted). */ + opener: boolean[]; start: number; end: number; } @@ -2588,75 +3053,145 @@ interface GapsEntrySpan { * boundary — a second, independently-written grouping pass is exactly how a * span-carrying sibling could disagree with the plain-lines version it is * supposed to be span-annotating. + * + * `markers` selects the marker set an entry may OPEN with (#3702). It defaults + * to the hyphen-only Gaps form, so every pre-existing caller is unaffected; + * the deferred-items callers pass `DEFERRED_BULLET_MARKERS`. Parameterising + * the shared seam — rather than widening it in place — is what keeps the + * template-mandated Gaps grammar out of the deferred-items ruling's blast + * radius while still leaving exactly ONE grouping pass in the module. */ -function splitGapsEntriesCore(sectionBody: string): GapsEntrySpan[] { +function splitGapsEntriesCore( + sectionBody: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, +): GapsEntrySpan[] { const rawLines = sectionBody.split('\n'); - const lineStarts: number[] = []; - const lineEnds: number[] = []; - let cursor = 0; - for (const rawLine of rawLines) { - lineStarts.push(cursor); - cursor += rawLine.length; - lineEnds.push(cursor); - cursor += 1; // the '\n' separator — absent after the final line, but nothing reads past it - } + const { lineStarts, lineEnds } = lineOffsets(rawLines); const entries: GapsEntrySpan[] = []; let current: string[] | null = null; let currentStartLine = -1; let currentEndLine = -1; let baseIndent: number | null = null; + // List memory per indent (#3702 round 2, B2 + round review; round 5, M2) — + // the top level decides entry boundaries; nested levels decide only which + // continuation lines count as accepted openers for field stripping. + const runs = new ListRuns(); + // Per-line opener flags for `current`, recorded HERE — the one place the + // run state is known — so the heading path's strip-only-openers rule reads + // the splitter's own verdict instead of re-deriving it (round review: a + // re-derivation without the paragraph reset re-accepted a rejected ordinal). + let currentOpeners: boolean[] = []; const flush = (): void => { if (current !== null) { - entries.push({ lines: current, start: lineStarts[currentStartLine], end: lineEnds[currentEndLine] }); + entries.push({ lines: current, opener: currentOpeners, start: lineStarts[currentStartLine], end: lineEnds[currentEndLine] }); } }; - // #3898: a spaced-hyphen thematic break (`- - -`, `- -`, `- - -`, …) is a - // SEPARATOR, not an entry. The opener regex below matches it (hyphen + - // whitespace), which fabricated a gap named `- -` with result 'unknown' — - // an item that cannot be cleared by editing any entry, because there is no - // entry, only the separator the author wrote deliberately. A line whose - // content after the opening marker consists solely of hyphens and spaces - // (with at least one further hyphen) is skipped entirely: it neither opens - // an entry nor is folded into the current one. This is deliberately NOT a - // full thematic-break concept (option 2 in the issue): a break does not - // close the Gaps list — entries after it keep parsing. + // #3898 (from `next`): a spaced-hyphen thematic break (`- - -`, `- -`, + // `- - -`, …) in `## Gaps` is a SEPARATOR, not an entry. The hyphen opener + // matches it (hyphen + whitespace), which fabricated a gap named `- -` with + // result 'unknown' — an item no edit can clear, because there is no entry, + // only the separator the author wrote deliberately. A line whose content + // after the opening marker is solely hyphens and spaces (with at least one + // further hyphen) is skipped: it neither opens an entry nor folds into the + // current one. Deliberately NOT a full thematic-break concept (option 2 in + // the issue): a break does not close the Gaps list — entries after it keep + // parsing. The deferred grammar (`blockStructure`) has its own, CommonMark + // reading of the same line through THEMATIC_BREAK_RE below, where a break + // CLOSES the list; this helper is consulted only for the Gaps set. const isSeparatorShaped = (line: string, bulletPrefixLen: number): boolean => { const remainder = line.slice(bulletPrefixLen); return /^[-\s]*$/.test(remainder) && remainder.includes('-'); }; + // Block structure (M1/M2 + column indents) is a property of the GRAMMAR, + // not of this seam: the Gaps set opts out and stays byte-for-byte on its + // `next` behaviour — see `indentWidth` for the indent half of that opt-out. + let scan = markers.blockStructure ? scanFencesFrom(rawLines, 0) : NO_FENCES; + let blankSeen = false; + // The run LEVEL of a line: every indent at or shallower than the list's + // base is the one top level (a dedenting list keeps its entry boundaries); + // deeper indents are their own nested levels. + const levelOf = (line: string): number => { + const ind = indentWidth(line.match(/^[ \t]*/)![0], markers); + return baseIndent !== null && ind <= baseIndent ? baseIndent : ind; + }; rawLines.forEach((rawLine, idx) => { const line = rawLine.replace(/\r$/, ''); - const bulletMatch = line.match(/^(\s*)-\s/); - // Narrowed skip (review disposition a): a separator-shaped line is skipped - // only when it sits BETWEEN entries (nothing open yet, or it would open a - // top-level entry — where the phantom came from). One landing strictly - // INSIDE a live entry (indent > baseIndent) folds back as a continuation - // line, so the entry's GapsEntrySpan stays byte-contiguous — the span - // invariant below and the ack writer's identity re-verification both hold. - if ( - bulletMatch && isSeparatorShaped(line, bulletMatch[0].length) && - (current === null || bulletMatch[1].length <= (baseIndent ?? 0)) - ) { - return; // separator line between entries — neither an opener nor a continuation + // B1: an unterminated fence runs to the end of its entry — the next line + // shaped like a top-level item ends it; rescan from there so a later + // delimiter is read on its own terms (see `scanFencesFrom`). + if (scan.unterminatedFrom !== -1 && idx > scan.unterminatedFrom && topLevelItemShape(line, markers, baseIndent)) { + scan = scanFencesFrom(rawLines, idx); } - if (bulletMatch) { - const indent = bulletMatch[1].length; + if (scan.fenced.has(idx) || (scan.unterminatedFrom !== -1 && idx >= scan.unterminatedFrom)) { + // Fence content never opens an entry (M2). Inside an open entry it is + // continuation — pushed, so the span invariant `acknowledgeDeferredItem` + // re-verifies still holds; before the first entry it is discarded. A + // fence is a non-list block: its opener ends the runs at its level and + // deeper, exactly as a paragraph does. + if (scan.openers.has(idx)) runs.endedAt(levelOf(line)); + if (current !== null) { + current.push(line); + currentOpeners.push(false); + currentEndLine = idx; + } + return; + } + if (markers.blockStructure && THEMATIC_BREAK_RE.test(line)) { + // A thematic break closes the list (M1): the open entry ends here, the + // break itself is neither an item nor a continuation, and nothing after + // it joins the closed entry — the next opener starts fresh. + flush(); + current = null; + runs.clear(); + blankSeen = false; + return; + } + // #3898 narrowed skip (review disposition a), Gaps set only: a + // separator-shaped line is skipped when it sits BETWEEN entries (nothing + // open yet, or it would open a top-level entry — where the phantom came + // from). One landing strictly INSIDE a live entry (indent > baseIndent) + // folds back as a continuation line, so the entry's GapsEntrySpan stays + // byte-contiguous — the span invariant below and the ack writer's identity + // re-verification both hold. The indent compare is the raw character + // count, which is what `indentWidth` measures for the Gaps set. + if (!markers.blockStructure) { + const bulletMatch = line.match(/^(\s*)-\s/); + if ( + bulletMatch && isSeparatorShaped(line, bulletMatch[0].length) && + (current === null || bulletMatch[1].length <= (baseIndent ?? 0)) + ) { + return; // separator line between entries — neither an opener nor a continuation + } + } + const opener = matchListOpener(line, markers, runs.at(levelOf(line))); + if (opener !== null) { + const { indent } = opener; if (baseIndent === null) baseIndent = indent; + runs.opened(levelOf(line)); if (indent <= baseIndent) { flush(); current = [line]; + currentOpeners = [true]; currentStartLine = idx; currentEndLine = idx; + blankSeen = false; // an opener is not blank — the memory must not survive it return; } } if (current !== null) { current.push(line); + currentOpeners.push(opener !== null); currentEndLine = idx; + // A blank line then a top-level non-list line is a PARAGRAPH: the list + // is over (CommonMark §5.3) and a later `5. x` is prose. Without the + // blank it is lazy continuation and the run stays open. + const blank = line.trim() === ''; + if (!blank && blankSeen && opener === null) runs.endedAt(levelOf(line)); + blankSeen = opener === null && blank; } // else: pre-first-bullet content (e.g. the template's HTML comment) — discarded. }); @@ -2667,7 +3202,8 @@ function splitGapsEntriesCore(sectionBody: string): GapsEntrySpan[] { /** * Split a `## Gaps` section body into per-entry line groups on TOP-LEVEL - * `- ` bullet openers. + * bullet openers — `- ` for Gaps, or whichever set `markers` names (#3702: + * the deferred-items callers pass the widened CommonMark set). * * The indentation of the FIRST bullet line encountered establishes the * "top-level" indent for the whole section; any subsequent `- `-opening line @@ -2680,17 +3216,21 @@ function splitGapsEntriesCore(sectionBody: string): GapsEntrySpan[] { * * Lines before the first bullet (e.g. the `` comment * the template emits) are discarded. An empty/whitespace-only section body - * (heading present, no bullets) returns `[]`. + * (heading present, no bullets) returns `[]`. Fenced code never opens an + * entry and a thematic break closes the open one (#3702 round 2, M1/M2). */ -function splitGapsEntries(sectionBody: string): string[][] { - return splitGapsEntriesCore(sectionBody).map((entry) => entry.lines); +function splitGapsEntries( + sectionBody: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, +): string[][] { + return splitGapsEntriesCore(sectionBody, markers).map((entry) => entry.lines); } /** * Sibling of `splitGapsEntries` (F1, #3458 follow-up review) that ADDITIVELY * carries each entry's character span — every existing `splitGapsEntries` * caller (`parseGapsItems`, `parseDeferredItemsWithStatus`, - * `splitDeferredHeadingEntries`'s `flushPending`) is unaffected and keeps + * `splitDeferredHeadingEntriesDetailed`'s `flushPending`) is unaffected and keeps * using the plain `lines`-only shape. `acknowledgeDeferredItem` is the one * caller that needs a span: it used to select an entry via `splitGapsEntries` * and then RE-FIND that entry's location with a fresh regex search over @@ -2702,8 +3242,11 @@ function splitGapsEntries(sectionBody: string): string[][] { * one. Carrying the span out of THIS same pass — the one that already knows * exactly where the entry lives — removes the re-derivation step entirely. */ -function splitGapsEntriesWithSpans(sectionBody: string): GapsEntrySpan[] { - return splitGapsEntriesCore(sectionBody); +function splitGapsEntriesWithSpans( + sectionBody: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, +): GapsEntrySpan[] { + return splitGapsEntriesCore(sectionBody, markers); } /** @@ -2732,38 +3275,175 @@ function splitGapsEntriesWithSpans(sectionBody: string): GapsEntrySpan[] { * keep their literal case, and mid-line emphasis is untouched, preserving the * start-anchored decoy invariant above. */ -function extractGapEntryFields(entryLines: string[]): Record { +function extractGapEntryFields( + entryLines: string[], + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + openerFlags?: boolean[], +): Record { const fields: Record = {}; - const fieldLineRe = /^([A-Za-z_][A-Za-z0-9_-]*):\s*(.*)$/; - const boldedKeyRe = /^\*+([A-Za-z_][A-Za-z0-9_-]*):\*+/; - - entryLines.forEach((rawLine, idx) => { - const line = rawLine.replace(/\r$/, ''); - // Strip ONLY the entry-opening bullet marker (idx 0); a bullet marker on - // a later line belongs to a nested sub-list and is handled by - // `splitGapsEntries` already folding it in — it is not itself a field - // line unless it independently matches `key: value` after stripping. - const bulletStripped = line.match(/^(\s*)-\s+(.*)$/); - const content = (idx === 0 && bulletStripped ? bulletStripped[2] : line.trim()) - .replace(boldedKeyRe, (_m, key: string) => `${key.toLowerCase()}:`); - - const m = fieldLineRe.exec(content); - if (!m) return; - const key = m[1]; - let value = m[2].trim(); - if (value.startsWith('"') && value.endsWith('"') && value.length >= 2) { - value = value.slice(1, -1); - } - if (!(key in fields)) fields[key] = value; + // A fenced line is content, not a field (#3702 round 2, round review): the + // splitters already keep fence lines from OPENING an entry, and a + // `status: resolved` quoted inside a code block must not resolve one either. + // An entry is a contiguous slice and a fence never spans two entries (the + // opener of the next entry would be fence content), so scanning the entry's + // own lines classifies exactly what the splitter classified. + // + // RAW lines, and that is the fix for #3702 round 3, m7. The heading path + // used to marker-strip its lines BEFORE calling this function, so the scan + // below ran over text the splitter never saw: `- ```sh` is an ordinary + // bullet to the splitter, but strips to ```` ```sh ````, which opens a + // fence here that exists in no other pass. A `**Status:** resolved` line + // after it was then suppressed as fence content and its resolved entry + // resurfaced as open. The stripping now happens INSIDE this function, after + // the fence scan, driven by the splitter's own per-line opener verdict. + entryFieldLines(entryLines, markers, openerFlags).forEach((field) => { + if (!field) return; + if (!(field.key in fields)) fields[field.key] = field.value; }); return fields; } -/** Fallback display text for a Gaps entry with no parseable `truth:` field. */ -function rawGapEntryText(entryLines: string[]): string { +/** + * Per line of an entry, the field it declares — or `null` where it declares + * none, INCLUDING because it is fenced. + * + * This is the seam, and it exists because `parseGapEntryFieldLine` alone was + * not it (#3702 round 3, pre-push review). The reader applied the fence gate + * before classifying and the acknowledge writer did not, so a `status:` line + * inside a fenced block was selected by the writer and skipped by the reader: + * the write produced a line nothing reads, the read-back guard refused it, and + * the entry became impossible to acknowledge at all — `audit acknowledge` + * surfaced an internal error and `complete-milestone` halted on it. That shape + * acknowledged cleanly on `next`, so it was a regression introduced by the fix + * for the nested-marker one, and the claim "the writer cannot select a line the + * reader will not read back" was false while the fence gate lived on one side. + * + * Both sides call this now, so the claim is structural rather than asserted. + */ +function entryFieldLines( + entryLines: string[], + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + openerFlags?: boolean[], +): ({ key: string; value: string; valueStart: number } | null)[] { + const fenced = entryFencedLines(entryLines, markers, openerFlags); + return entryLines.map((rawLine, idx) => ( + fenced.has(idx) ? null : parseGapEntryFieldLine(rawLine, markers, stripsMarkerAt(idx, openerFlags)) + )); +} + +/** + * The fenced lines of ONE entry, as the reader and the writer both see them. + * A leaf's line 0 is its heading TEXT, not a Markdown line: a heading that + * reads ``` or ~~~ is a heading, and must not open a fence over the body + * beneath it (round 5, RV6.5 — it fenced every field line, so the reader + * read nothing and the writer's marker landed on a line nothing reads). + * `openerFlags[0] === false` is the leaf tell: a pending or headless entry's + * line 0 is an accepted opener, and a marker line is never a delimiter. + */ +function entryFencedLines(entryLines: string[], markers: BulletMarkers, openerFlags?: boolean[]): Set { + if (!markers.blockStructure) return new Set(); + const leaf = openerFlags !== undefined && openerFlags[0] === false; + return fencedLineSet(leaf ? ['', ...entryLines.slice(1)] : entryLines); +} + +/** + * Which lines of an entry carry an entry-opening marker to be stripped before + * the line is read as a field. + * + * Without flags — the headless and `## Gaps` shapes — that is line 0 alone: a + * marker on a later line belongs to a nested sub-list (`splitGapsEntries` + * already folded it in) and is not a field line unless it independently + * matches `key: value` after a plain trim. + * + * With flags — the heading shape — it is whichever lines the SPLITTER accepted + * as list openers, because there every body line may be a sibling bullet + * carrying a field (#3457) while line 0 is the heading TEXT and carries no + * marker at all. Reading the splitter's verdict rather than re-deriving it is + * what keeps a rejected ordinal (`3. status: resolved` as prose) from being + * stripped into a field. + */ +function stripsMarkerAt(idx: number, openerFlags?: boolean[]): boolean { + return openerFlags ? openerFlags[idx] === true : idx === 0; +} + +/** + * The ONE place an entry line is classified as a `key: value` field line. + * `extractGapEntryFields` reads through it, and `acknowledgeDeferredItem` + * locates the line it will rewrite through it. + * + * Sharing the classifier is what makes the writer structurally unable to + * select a line the reader will not read back (#3702 round 3, B1; the shape + * is #3773's, parameterised here by `markers` per the round-3 review's + * prescribed end state). The writer used to carry its own marker-widened + * status regex, so a nested ` * status: pending` was selectable by the + * writer and invisible to this reader: acknowledge rewrote it in place, + * returned `ok`, and the item stayed outstanding forever. A single classifier + * has no second copy to drift from. + * + * `valueStart` is the offset, in the CR-stripped line, at which the VALUE + * begins — so a rewrite can replace the value without a second regex of its + * own. The bolded-key unwrap below is a PREFIX rewrite, so the tail of the + * rewritten content is byte-identical to the tail of the original and the + * offset maps back directly. + * + * Returns `null` for a non-field line. + */ +function parseGapEntryFieldLine( + rawLine: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + stripMarker = true, +): { key: string; value: string; valueStart: number } | null { + const fieldLineRe = /^([A-Za-z_][A-Za-z0-9_-]*):\s*(.*)$/; + const boldedKeyRe = /^\*+([A-Za-z_][A-Za-z0-9_-]*):\*+/; + + const line = rawLine.replace(/\r$/, ''); + const bulletStripped = stripMarker ? line.match(markers.strip) : null; + const bare = bulletStripped ? bulletStripped[2] : line.trim(); + // Where `bare` begins in `line`. The two branches differ: the marker strip's + // group 2 runs to end-of-line, so it is a plain suffix; `trim()` also cuts + // the tail, so its offset is the LEADING run alone. Computing one from the + // other's shape under-counts by the trailing whitespace. + const headLen = bulletStripped ? line.length - bare.length : line.length - line.trimStart().length; + const content = bare.replace(boldedKeyRe, (_m, key: string) => `${key.toLowerCase()}:`); + + const m = fieldLineRe.exec(content); + if (!m) return null; + let value = m[2].trim(); + if (value.startsWith('"') && value.endsWith('"') && value.length >= 2) { + value = value.slice(1, -1); + } + return { key: m[1], value, valueStart: headLen + (bare.length - m[2].length) }; +} + +/** + * Fallback display text for a Gaps entry with no parseable `truth:` field. + * + * `markers` selects which opening marker is stripped (#3702) — hyphen-only by + * default, the widened set for deferred-items callers, so a `*`-opened entry + * renders the same name its hyphen twin would. That name is the key + * `acknowledgeDeferredItem` matches on, so the two MUST use the same set: + * rendering `* alpha` where the parse surfaced `alpha` would make the entry + * un-acknowledgeable. + * + * `openerFlags` decides WHICH lines are stripped, and on the heading shape + * line 0 is not one of them (#3702 round 3, m8). There line 0 is the heading + * TEXT, so an unconditional strip renamed `### 1. Race in the writer` to + * `Race in the writer` and `### * starred title` to `starred title` — both + * silent renames of the very key acknowledge matches on, and both a change + * from this parser's behaviour on `next`. + */ +function rawGapEntryText( + entryLines: string[], + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + openerFlags?: boolean[], +): string { return entryLines - .map((l, i) => (i === 0 ? l.replace(/^(\s*)-\s+/, '') : l.trim())) + // Line 0 ONLY, and only if the splitter accepted it as an opener. The + // opener flags say which lines carry a marker; the entry's NAME is a + // different question, and stripping a body line's marker out of it changes + // the key `acknowledgeDeferredItem` matches on. + .map((l, i) => (i === 0 && stripsMarkerAt(0, openerFlags) ? l.replace(markers.strip, '$2') : l.trim())) .join(' ') .trim(); } @@ -2998,4 +3678,13 @@ export = { parseDeferredItems, parseDeferredItemsWithStatus, acknowledgeDeferredItem, + // #3702 round 2 (M3): exposed for the marker-grammar parity test only. + // Narrowed in round 3 (M6): the two status-line regexes are gone from the + // module, so nothing exports them, and the parity test they were widened for + // could not reach the defect it was meant to guard anyway — it asserted the + // four WRITER regexes shared a source string, which is true of a detect/read + // asymmetry too. The pair below is what the behavioural parity test against + // `iterateBullets` actually reads. + DEFERRED_MARKER_ALT, + DEFERRED_BULLET_MARKERS, }; diff --git a/tests/audit-command-cutover.test.cjs b/tests/audit-command-cutover.test.cjs index d4ee199e7..0d78fe7f5 100644 --- a/tests/audit-command-cutover.test.cjs +++ b/tests/audit-command-cutover.test.cjs @@ -2094,10 +2094,12 @@ describe('bug #950: quick-task SUMMARY must carry status: complete', () => { assert.match(fs.readFileSync(filePath, 'utf-8'), /status: acknowledged/); // A table row embedded in the entry's span is still refused — a - // non-contiguous span cannot anchor a safe write. - const tabled = ['## Deferred Items', '', '### Tabled finding', '', '- a bullet', '| x | y |', ''].join('\n'); + // non-contiguous span cannot anchor a safe write. The row sits BETWEEN + // two entry lines: a row after the entry's last line is outside its + // span, and that write is anchored and safe (#3702 round 5). + const tabled = ['## Deferred Items', '', '### Tabled finding', '', '- a bullet', '| x | y |', '- more evidence', ''].join('\n'); fs.writeFileSync(filePath, tabled); - const refuse = ack(tmpDir, ['--category', 'deferred_items', '--phase', '01', '--file', 'deferred-items.md', '--text', 'Tabled finding - a bullet', '--milestone', 'v1.0']); + const refuse = ack(tmpDir, ['--category', 'deferred_items', '--phase', '01', '--file', 'deferred-items.md', '--text', 'Tabled finding - a bullet - more evidence', '--milestone', 'v1.0']); assert.equal(refuse.success, false, 'table-embedded spans must still refuse'); assert.match(refuse.error, /GFM table row/i); assert.equal(fs.readFileSync(filePath, 'utf-8'), tabled, 'file must be byte-identical — nothing written on refusal'); @@ -2112,6 +2114,43 @@ describe('bug #950: quick-task SUMMARY must carry status: complete', () => { assert.equal(fs.readFileSync(filePath, 'utf-8'), proseOnly, 'file must be byte-identical'); }); + test('#3702 round 2 (m3, updated for #3781): a heading-delimited file written with `*`, `+` or `1.` SURFACES its entries, and the CLI writer acknowledges them', () => { + // Before #3702 such a file parsed to zero entries and `complete-milestone` + // closed over it silently. It now yields entries; under #3781 (merged in + // round 5) the heading shape is acknowledgeable, so the milestone loop + // suppresses them through the same two CLI calls it makes for a hyphen + // file: the listing that puts the entry in front of the writer, and the + // acknowledge that takes it. (Until #3781 the writer refused the shape + // and the loop halted — that contract is gone from `next`.) + const fixtures = [['*', '02-star'], ['+', '03-plus'], ['1.', '04-ordered']]; + const files = new Map(); + for (const [marker, slug] of fixtures) { + const phaseDir = planningPath('phases', slug); + fs.mkdirSync(phaseDir, { recursive: true }); + const filePath = path.join(phaseDir, 'deferred-items.md'); + const before = ['## Deferred Items', '', `### Out of scope under ${slug}`, '', `${marker} **What:** detail.`, ''].join('\n'); + fs.writeFileSync(filePath, before); + files.set(slug, { filePath, before }); + } + + const listed = audit(tmpDir).items.deferred_items.filter((i) => /^Out of scope under /.test(i.text)); + assert.equal(listed.length, fixtures.length, `every marker's heading entry must be listed: ${JSON.stringify(listed)}`); + + for (const item of listed) { + const slug = /under (\S+)/.exec(item.text)[1]; + const { filePath, before } = files.get(slug); + // `--text` is the audit's own reported text, exactly as the loop passes it. + const result = ack(tmpDir, ['--category', 'deferred_items', '--phase', slug.slice(0, 2), '--file', 'deferred-items.md', '--text', item.text, '--milestone', 'v1.0']); + assert.equal(result.success, true, `${slug}: a widened-marker heading entry must acknowledge; stderr: ${result.error}`); + const after = fs.readFileSync(filePath, 'utf-8'); + assert.notEqual(after, before, `${slug}: the file must carry the marker`); + assert.match(after, /status: acknowledged/, slug); + } + // And the loop's next listing no longer surfaces them. + const relisted = audit(tmpDir).items.deferred_items.filter((i) => /^Out of scope under /.test(i.text)); + assert.equal(relisted.length, 0, `acknowledged entries must leave the listing: ${JSON.stringify(relisted)}`); + }); + // ── F1 (#3458 follow-up review, HIGH): the writer must splice by the // SELECTED entry's own carried span, never re-find it by searching — // otherwise a byte-identical substring living inside an EARLIER entry diff --git a/tests/uat.test.cjs b/tests/uat.test.cjs index d50a95bb2..8859f2d8e 100644 --- a/tests/uat.test.cjs +++ b/tests/uat.test.cjs @@ -20,7 +20,10 @@ const { acknowledgeDeferredItem, parseUatItems, parseUatItemsWithStats, + DEFERRED_MARKER_ALT, + DEFERRED_BULLET_MARKERS, } = require('../gsd-core/bin/lib/uat.cjs'); +const { iterateBullets } = require('../gsd-core/bin/lib/markdown-sectionizer.cjs'); describe('audit-uat command', () => { let tmpDir; @@ -1801,9 +1804,9 @@ describe('#2287 progress.md forensic_audit: deferred-items.md contract', () => { }); }); -// ─── parseDeferredItems property test (#2287) ────────────────────────────── +// ─── parseDeferredItems property test (#2287, widened by #3702 round 2) ───── -describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () => { +describe('#2287 parseDeferredItems: property (status: resolved fail-safe) × marker × shape × line ending', () => { // Single-line entry text: no newlines (would break bullet-entry splitting), // non-empty after trim, and never itself SHAPED like a `status:` field line // (that would be indistinguishable from a real field regardless of intent). @@ -1818,43 +1821,88 @@ describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () = const decoyText = plainText.map((s) => `${s} status: resolved trailing note`); const textArb = fc.oneof(plainText, decoyText); - const entryArb = fc.record({ text: textArb, resolved: fc.boolean() }); + // `statusFirst` matters on the heading shape: round 1's only CRLF test put + // `**Status:**` LAST, the one line `collectSection`'s `.trimEnd()` had + // already de-CR'd, and the B1 regression hid behind it. + // `decoy` adds a prose line that BEGINS with a non-1 ordinal (`7. …`): under + // the start-at-1 rule (B2) it is never an item, never evidence and never a + // field — round 1 read it as an opener. Placed where it cannot end a run: + // before the first headless entry, and first in a heading body. + const entryArb = fc.record({ text: textArb, resolved: fc.boolean(), statusFirst: fc.boolean(), decoy: fc.boolean() }); + const decoyOrdinal = fc.integer({ min: 2, max: 999999999 }); + + // #3702 round 2 (B3): the marker set is an enumerated domain — exactly what a + // property is for. Ordered markers are numbered from 1 (the ordered-run + // rule, B2); the two shapes exercise both splitters; CRLF exercises the + // heading path's CR handling (B1). + const markerArb = fc.constantFrom('-', '*', '+', 'ordered'); + const shapeArb = fc.constantFrom('headless', 'heading'); + const eolArb = fc.constantFrom('\n', '\r\n'); + const mk = (marker, i) => (marker === 'ordered' ? `${i + 1}.` : marker); + + const render = (entries, marker, shape, eol, ordinal = 7) => { + const lines = ['## Deferred Items', '']; + if (shape === 'headless' && entries.some((e) => e.decoy)) { + // Pre-first-entry prose: a rejected ordinal is discarded, an accepted + // one (round 1) opens a phantom entry and breaks the count. + lines.push(`${ordinal}. ${entries.find((e) => e.decoy).text} status: resolved`, ''); + } + entries.forEach((e, i) => { + if (shape === 'headless') { + lines.push(`${mk(marker, i)} ${e.text}`); + if (e.resolved) lines.push(' status: resolved'); + } else { + lines.push(`### ${e.text}`, ''); + // A decoy ordinal line FIRST in the body: never stripped, so the + // `status: resolved` after it can never become a field. + if (e.decoy) lines.push(`${ordinal}. ${e.text} status: resolved`); + const what = (n) => `${mk(marker, n)} **What:** ${e.text}`; + const status = (n) => `${mk(marker, n)} **Status:** resolved`; + if (!e.resolved) lines.push(what(0)); + else if (e.statusFirst) lines.push(status(0), what(1)); + else lines.push(what(0), status(1)); + lines.push(''); + } + }); + return lines.join(eol); + }; + const idOf = (name) => { const m = /E(\d+)_/.exec(name); return m ? Number(m[1]) : -1; }; test('property: an entry is surfaced iff it is NOT marked status: resolved; surfaced count == non-resolved count', () => { fc.assert( fc.property( - fc.array(entryArb, { maxLength: 20 }), - (rawEntries) => { + fc.array(entryArb, { maxLength: 20 }), markerArb, shapeArb, eolArb, decoyOrdinal, + (rawEntries, marker, shape, eol, ordinal) => { // Index-prefix for uniqueness so surfaced items can be mapped back // to their source entry unambiguously even with colliding random text. - const entries = rawEntries.map((e, i) => ({ text: `E${i}_${e.text}`, resolved: e.resolved })); - - const lines = ['## Deferred Items', '']; - for (const e of entries) { - lines.push(`- ${e.text}`); - if (e.resolved) lines.push(' status: resolved'); - } - const content = lines.join('\n'); + const entries = rawEntries.map((e, i) => ({ ...e, text: `E${i}_${e.text}` })); + const content = render(entries, marker, shape, eol, ordinal); + const where = `${shape} ${JSON.stringify(marker)} ${JSON.stringify(eol)} decoy=${ordinal}`; const items = parseDeferredItems(content); - const surfacedNames = new Set(items.map((it) => it.name)); + const surfacedIds = new Set(items.map((it) => idOf(it.name))); const expectedUnresolved = entries.filter((e) => !e.resolved); - const expectedResolved = entries.filter((e) => e.resolved); // Total surfaced count equals the count of non-resolved entries. - assert.strictEqual(items.length, expectedUnresolved.length); + assert.strictEqual(items.length, expectedUnresolved.length, where); // Every non-resolved entry IS surfaced (including status:-shaped // decoy substrings embedded mid-line — those must not flip the // outcome). - for (const e of expectedUnresolved) { - assert.ok(surfacedNames.has(e.text), `expected unresolved entry to surface: ${e.text}`); + for (const [i, e] of entries.entries()) { + assert.strictEqual(surfacedIds.has(i), !e.resolved, `${where}: ${e.resolved ? 'resolved entry must never surface' : 'unresolved entry must surface'}: ${e.text}`); } - // No status:-resolved entry is EVER surfaced. - for (const e of expectedResolved) { - assert.ok(!surfacedNames.has(e.text), `status: resolved entry must never surface: ${e.text}`); + // Headless: the surfaced NAME is the entry text with the marker gone, + // whichever marker it was (the name is what acknowledge matches on). + if (shape === 'headless') { + for (const it of items) assert.strictEqual(it.name, entries[idOf(it.name)].text, where); + } + // A heading body carrying ONLY a decoy ordinal line is prose, not an entry. + if (shape === 'heading' && entries.length > 0) { + const decoyOnly = `## Deferred Items${eol}${eol}### only-decoy${eol}${eol}${ordinal}. ${entries[0].text} status: resolved${eol}`; + assert.deepStrictEqual(parseDeferredItems(decoyOnly), [], `${where}: decoy-only body`); } // Every returned item carries the fixed deferred category/result shape. @@ -1866,6 +1914,39 @@ describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () = ) ); }); + + test('property: acknowledge reaches and rewrites every unresolved headless entry, whichever marker or line ending', () => { + // The writer refuses the heading shape by design (`unsupported_heading_shape`), + // so this ranges over the headless shape only. It is the property that + // reaches M4 (CRLF rewrite reported ok and wrote nothing) and m2 (indent). + fc.assert( + fc.property( + fc.array(entryArb, { minLength: 1, maxLength: 12 }), markerArb, eolArb, + (rawEntries, marker, eol) => { + const entries = rawEntries.map((e, i) => ({ ...e, text: `E${i}_${e.text}` })); + let content = render(entries, marker, 'headless', eol); + const where = `${JSON.stringify(marker)} ${JSON.stringify(eol)}`; + + for (const e of entries.filter((x) => !x.resolved)) { + const got = acknowledgeDeferredItem(content, e.text); + assert.strictEqual(got.status, 'ok', `${where}: ${e.text}`); + assert.notStrictEqual(got.content, content, `${where}: an ok must have written: ${e.text}`); + content = got.content; + } + const after = parseDeferredItemsWithStatus(content); + assert.strictEqual(after.length, entries.length, where); + for (const it of after) { + const e = entries[idOf(it.name)]; + assert.strictEqual(it.status, e.resolved ? 'resolved' : 'acknowledged', `${where}: ${it.name}`); + } + // `acknowledged` is suppressed at the AUDIT layer, not the parser's: + // only `resolved` leaves the outstanding list here, so the count is + // unchanged by the writes above. + assert.strictEqual(parseDeferredItems(content).length, entries.filter((x) => !x.resolved).length, where); + } + ) + ); + }); }); // ─── #2766: archived phase dirs, and GFM-table-shaped deferred/gaps ──────── @@ -2313,6 +2394,1235 @@ describe('#3457 parseDeferredItems: heading-delimited entries', () => { }); }); +describe('#3702 parseDeferredItems: list-marker grammar', () => { + // `deferred-items.md` has no template and no mandated shape, but the parser + // recognised only the hyphen marker — so `*`, `+` and ordered lists (all + // lists in CommonMark and GFM) contributed ZERO entries on both the headless + // and the heading-delimited path. A mixed file dropped its non-hyphen entries + // while keeping their hyphenated siblings, under-reporting without ever + // looking empty. + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + const count = (md) => names(md).length; + + // The marker set the ruling widened to. `1)` is deliberately absent — the + // paren-terminated ordered form is out of scope for this fix, and the + // `1) still yields zero` case below pins that as intended, not as an oversight. + const MARKERS = ['-', '*', '+', '1.']; + + test('AC1 headless: every marker yields the same count as the hyphen form', () => { + const shape = (m) => `${m} alpha\n${m} beta\n`; + const hyphen = count(shape('-')); + + assert.strictEqual(hyphen, 2, 'baseline: the hyphen form must yield 2'); + for (const m of MARKERS) { + assert.strictEqual(count(shape(m)), hyphen, `marker ${JSON.stringify(m)}: ${JSON.stringify(names(shape(m)))}`); + } + }); + + test('AC1 headless: the entry NAME drops the marker, whichever marker it is', () => { + // rawGapEntryText renders the name acknowledgeDeferredItem later matches on, + // so a marker left in the rendered name would make the entry unreachable. + for (const m of MARKERS) { + assert.deepStrictEqual(names(`${m} alpha\n`), ['alpha'], `marker ${JSON.stringify(m)}`); + } + }); + + test('AC2 heading-delimited: a body carrying any marker is KEPT (was dropped)', () => { + const shape = (m) => `### Entry\n\n${m} **What:** x.\n`; + + for (const m of MARKERS) { + assert.strictEqual(count(shape(m)), 1, `marker ${JSON.stringify(m)}: ${JSON.stringify(names(shape(m)))}`); + } + }); + + test('AC2 heading-delimited: a mixed file no longer drops its non-hyphen entry', () => { + // The row that bites hardest in the wild: the file never looks empty, it + // just silently under-reports. + const got = names('### Hyphen entry\n\n- x.\n\n### Asterisk entry\n\n* y.\n'); + + assert.strictEqual(got.length, 2, JSON.stringify(got)); + assert.match(got[0], /^Hyphen entry/); + assert.match(got[1], /^Asterisk entry/); + }); + + test('AC3: a resolved-status field under any marker resolves its entry', () => { + // The lockstep property: widening what OPENS an entry without widening the + // marker STRIP feeding field extraction would surface the entry and then + // never resolve it — permanently unresolved, which is worse than dropped. + // + // Asserted through parseDeferredItemsWithStatus, NOT through an empty + // parseDeferredItems: "no outstanding item" is also what a DROPPED entry + // looks like, so the weaker form passes against the unfixed parser for + // precisely the reason under test. The entry must exist AND read resolved. + for (const m of MARKERS) { + for (const [shape, md] of [ + ['heading', `${SECTION}### Entry\n\n${m} **What:** x.\n${m} **Status:** resolved\n`], + ['headless', `${SECTION}${m} alpha\n status: resolved\n`], + ]) { + const where = `${shape} shape, marker ${JSON.stringify(m)}`; + const withStatus = parseDeferredItemsWithStatus(md); + + assert.strictEqual(withStatus.length, 1, `${where}: entry must be parsed at all — ${JSON.stringify(withStatus)}`); + assert.strictEqual(withStatus[0].status, 'resolved', `${where}: ${JSON.stringify(withStatus)}`); + assert.deepStrictEqual(parseDeferredItems(md), [], `${where}: resolved entries are not outstanding`); + } + } + }); + + test('AC3: the acknowledge writer reaches an entry written under any marker', () => { + for (const m of MARKERS) { + const content = `${SECTION}${m} alpha\n`; + const got = acknowledgeDeferredItem(content, 'alpha'); + + assert.strictEqual(got.status, 'ok', `marker ${JSON.stringify(m)}`); + assert.match(got.content, /status: acknowledged/, `marker ${JSON.stringify(m)}`); + assert.strictEqual( + parseDeferredItemsWithStatus(got.content)[0].status, + 'acknowledged', + `marker ${JSON.stringify(m)}: the written marker must parse back`, + ); + } + }); + + test('AC3: an already-acknowledged entry under any marker is not double-written', () => { + for (const m of MARKERS) { + const original = `${SECTION}${m} alpha\n`; + const once = acknowledgeDeferredItem(original, 'alpha').content; + const twice = acknowledgeDeferredItem(once, 'alpha').content; + + // Anti-vacuity: an unreachable entry is also idempotent, so pin that the + // first call actually wrote before pinning that the second did not. + assert.notStrictEqual(once, original, `marker ${JSON.stringify(m)}: first acknowledge must write`); + assert.strictEqual(twice, once, `marker ${JSON.stringify(m)}`); + } + }); + + test('AC4: prose-only and bare headings still contribute nothing', () => { + // The "prose is not an item" contract is untouched: an asterisk bullet is + // not prose, so widening the marker set cannot start counting prose. + assert.deepStrictEqual(names('### Musings\n\njust prose here.\n'), []); + assert.deepStrictEqual(names('### A bare heading with no body\n'), []); + assert.deepStrictEqual(names('### Notes\n\nwe considered * and + as options.\n'), []); + }); + + test('AC4: a bolded field key is not mistaken for an asterisk bullet', () => { + // `**Status:**` opens with `*` but supplies no whitespace after it, so the + // widened marker declines and the bolded-key path still owns the line. + const got = parseDeferredItemsWithStatus(`${SECTION}### Entry\n\n- **What:** x.\n**Status:** resolved\n`); + + assert.strictEqual(got.length, 1, JSON.stringify(got)); + assert.strictEqual(got[0].status, 'resolved', JSON.stringify(got)); + }); + + test('AC5: a table under a leaf heading still yields exactly its rows', () => { + // The anti-double-count property (#2766): table lines are skipped before + // the body-marker flag can be set, and a `|` row is not a list marker, so + // the heading still contributes no phantom entry. + const oneRow = names('### Discovered\n\n| Test | Seeds |\n|---|---|\n| test_a | 0, 1 |\n'); + assert.deepStrictEqual(oneRow, ['test_a — 0, 1'], JSON.stringify(oneRow)); + + const twoRows = names('### Discovered\n\n| Test | Seeds |\n|---|---|\n| test_a | 0 |\n| test_b | 1 |\n'); + assert.strictEqual(twoRows.length, 2, JSON.stringify(twoRows)); + }); + + test('the paren-terminated ordered marker `1)` remains out of scope', () => { + // Pinned so a later reader sees this as the ruling's scope, not a miss. + assert.deepStrictEqual(names('1) alpha\n2) beta\n'), []); + }); + + test('CRLF files: widened markers split and resolve identically', () => { + for (const m of MARKERS) { + const crlf = `## Deferred Items\r\n\r\n### Entry\r\n\r\n${m} **What:** x.\r\n${m} **Status:** resolved\r\n`; + const withStatus = parseDeferredItemsWithStatus(crlf); + + // Same anti-vacuity as AC3: an empty outstanding list would also be + // satisfied by the entry never being parsed. + assert.strictEqual(withStatus.length, 1, `marker ${JSON.stringify(m)}: ${JSON.stringify(withStatus)}`); + assert.strictEqual(withStatus[0].status, 'resolved', `marker ${JSON.stringify(m)}`); + assert.deepStrictEqual(parseDeferredItems(crlf), [], `marker ${JSON.stringify(m)}`); + } + }); + + test('nested sub-lists under any marker stay folded into their parent entry', () => { + // splitGapsEntries' indent rule (#2286) is marker-agnostic: only a marker at + // or shallower than the first one seen opens a new entry. + for (const m of MARKERS) { + const got = names(`${m} alpha\n ${m} nested one\n ${m} nested two\n${m} beta\n`); + assert.strictEqual(got.length, 2, `marker ${JSON.stringify(m)}: ${JSON.stringify(got)}`); + } + }); +}); + +describe('#3702 round 2: CRLF on the heading path and in the acknowledge writer', () => { + const MARKERS = ['-', '*', '+', '1.']; + const CRLF_SECTION = '## Deferred Items\r\n\r\n'; + + test('B1: a CRLF heading entry resolves when **Status:** is NOT the last line', () => { + // Round-1's CRLF test put `**Status:**` on the fixture's LAST line, where + // `collectSection`'s `.trimEnd()` had already removed the one `\r` that + // mattered — a false green. Every other line of a CRLF body still carries + // its `\r`, and a `$`-anchored marker strip fails on it, so the marker + // survived into field extraction and the field was silently lost. + for (const m of MARKERS) { + const statusFirst = `${CRLF_SECTION}### Entry\r\n\r\n${m} **Status:** resolved\r\n${m} **What:** x.\r\n`; + const withStatus = parseDeferredItemsWithStatus(statusFirst); + + assert.strictEqual(withStatus.length, 1, `marker ${JSON.stringify(m)}: ${JSON.stringify(withStatus)}`); + assert.strictEqual(withStatus[0].status, 'resolved', `marker ${JSON.stringify(m)}: status-first CRLF must resolve`); + assert.deepStrictEqual(parseDeferredItems(statusFirst), [], `marker ${JSON.stringify(m)}`); + + // And the mirror: a field ABOVE a trailing status line is not lost either. + const whatFirst = `${CRLF_SECTION}### Entry\r\n\r\n${m} **What:** x.\r\n${m} **Status:** resolved\r\n${m} **Why:** y.\r\n`; + assert.strictEqual(parseDeferredItemsWithStatus(whatFirst)[0].status, 'resolved', `marker ${JSON.stringify(m)}: mid-body status`); + } + }); + + test('B1: a CRLF heading entry parses byte-for-byte like its LF twin', () => { + for (const m of MARKERS) { + const body = `### Entry\n\n${m} **Status:** resolved\n${m} **What:** x.\n`; + const lf = parseDeferredItemsWithStatus(`## Deferred Items\n\n${body}`); + const crlf = parseDeferredItemsWithStatus(`${CRLF_SECTION}${body.replace(/\n/g, '\r\n')}`); + assert.deepStrictEqual(crlf, lf, `marker ${JSON.stringify(m)}`); + } + }); + + test('M4: acknowledge REWRITES an existing status line on a CRLF file (was: ok + no write)', () => { + // Pre-existing on `next`: the finder tested a CR-stripped copy, the rewrite + // ran on the raw `\r`-terminated line with a `$`-anchored regex, `replace` + // returned the input unchanged, and the writer reported `ok` over content + // that was byte-identical — the item then resurfaced on every audit. + for (const m of MARKERS) { + const content = `${CRLF_SECTION}${m} alpha\r\n status: pending\r\n${m} beta\r\n`; + const target = parseDeferredItemsWithStatus(content)[0].name; + const got = acknowledgeDeferredItem(content, target); + + assert.strictEqual(got.status, 'ok', `marker ${JSON.stringify(m)}`); + assert.notStrictEqual(got.content, content, `marker ${JSON.stringify(m)}: an ok must have written`); + assert.strictEqual(parseDeferredItemsWithStatus(got.content)[0].status, 'acknowledged', `marker ${JSON.stringify(m)}`); + } + }); + + test('m2: the inserted status line takes the entry indent on a CRLF file', () => { + for (const m of MARKERS) { + const content = `${CRLF_SECTION} ${m} alpha\r\n ${m} beta\r\n`; + const got = acknowledgeDeferredItem(content, 'alpha'); + + assert.strictEqual(got.status, 'ok', `marker ${JSON.stringify(m)}`); + assert.match(got.content, /\n {6}status: acknowledged/, `marker ${JSON.stringify(m)}: indent 4 + 2, not the indent-0 fallback`); + } + }); +}); + +describe('#3702 round 3: detect/strip symmetry on the acknowledge path (B1, B2)', () => { + // The round-3 blocker. Round 2 widened the WRITER's status-line finder to + // the deferred marker set while the reader still de-bulleted line 0 only, so + // a nested ` * status:` was selectable by the writer and invisible to the + // reader: acknowledge rewrote it, returned `ok`, and the item stayed + // outstanding on every later audit. Measured against a `next` build, `*`, + // `+` and `1.` all resolved on base and stopped resolving at round 2's head + // — a regression, not a gap in new behaviour. + for (const marker of ['-', '*', '+', '1.']) { + test(`a nested "${marker} status:" line acknowledges and READS BACK`, () => { + const doc = `## Deferred Items\n\n- alpha thing\n ${marker} status: pending\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok', `${marker}: acknowledge reported`); + assert.notStrictEqual(got.content, doc, `${marker}: content actually changed`); + const after = parseDeferredItemsWithStatus(got.content); + assert.strictEqual( + after[0].status, 'acknowledged', + `${marker}: the entry must read back as acknowledged — reporting ok over a line the reader skips is the defect`, + ); + }); + } + + // The hyphen row above is NOT a widened marker: it was already broken on + // `next`, for the same reason. One classifier cannot be right for three + // markers and wrong for the fourth, so it is fixed here rather than left as + // a pre-existing defect found while working. + test('a bare capitalised "Status:" resolves rather than reporting a write nothing reads', () => { + // The reader stores a bare key case-sensitively, so `Status:` is not + // `status:`. The writer must therefore NOT select it — it falls to the + // insert branch, which writes a line the reader does read. + const doc = '## Deferred Items\n\n- alpha\n Status: pending\n'; + const items = parseDeferredItemsWithStatus(doc); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(got.content)[0].status, 'acknowledged'); + }); + + test('a fenced "status:" line does not make the entry un-acknowledgeable', () => { + // Found by the pre-push review, and a regression THIS round introduced: + // the reader applied the fence gate before classifying and the writer did + // not, so the writer selected a fenced `status:` line the reader skips. + // The read-back guard then refused the write and the entry could not be + // acknowledged at all — `audit acknowledge` raised an internal error and + // `complete-milestone` halted. It acknowledged cleanly on `next`. + const F = '`'.repeat(3); + const doc = `## Deferred Items\n\n- alpha\n ${F}\n status: pending\n ${F}\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok', 'a fenced status line must not refuse the write'); + assert.strictEqual( + parseDeferredItemsWithStatus(got.content)[0].status, 'acknowledged', + 'the insert branch must write a line the reader reads, rather than rewriting one it skips', + ); + }); + + test('CONTROL: a bolded line-0 key keeps its ** wrapper and spelling through the rewrite', () => { + // Held before this round too — it is a control on the NEW offset-based + // rewrite, not a regression test for a reported defect. The rewrite + // replaces the VALUE at the offset the classifier reported, so the key is + // untouched by construction; a key-matching regex would have to reproduce + // the wrapper to preserve it, which is the thing that could regress. + const doc = '## Deferred Items\n\n- **Status:** pending alpha\n'; + const items = parseDeferredItemsWithStatus(doc); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok'); + assert.match(got.content, /- \*\*Status:\*\* acknowledged/); + }); + + test('acknowledge is idempotent across a re-read', () => { + // The failure this guards is the one the defect actually produced: the + // item resurfaces, gets acknowledged again, and never settles. + const doc = '## Deferred Items\n\n- alpha\n * status: pending\n'; + const first = acknowledgeDeferredItem(doc, parseDeferredItemsWithStatus(doc)[0].name); + const reread = parseDeferredItemsWithStatus(first.content); + assert.strictEqual(reread[0].status, 'acknowledged'); + const second = acknowledgeDeferredItem(first.content, reread[0].name); + assert.strictEqual(second.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(second.content)[0].status, 'acknowledged'); + }); +}); + +describe('#3702 round 4: ONE end-of-file CRLF algorithm, adopted from #3773 with its B4 closed (M1)', () => { + // Two open PRs shipped two different answers to "what line ending does an + // entry that ENDS THE FILE get?", and the maintainer's round-4 ruling was + // that the disagreement "needs one answer, not two". Neither shipped answer + // was that one. Measured, on builds of both heads, over the fixtures below: + // + // case #3739 r3 #3773 here + // undelimited single entry, CRLF preamble pass FAIL pass + // LF-dominant list, one stray CRLF at EOF FAIL pass pass + // (the other five) pass pass pass + // + // #3739's `content.endsWith('\r\n', matchIndexInContent)` reads the + // terminator of the PREVIOUS line, so it propagated an isolated CRLF into an + // LF-dominant list. #3773's `crlfAtEof` asks the right question but over a + // scope that goes EMPTY for an undelimited single-entry list, so it inserted + // a bare `\n` into a CRLF document — its own B4, and a violation of the + // uniform-CRLF invariant the fix exists to hold. What lands here is + // `crlfAtEof`'s semantics over a scope that widens instead of going empty. + // + // Every fixture below is a counterexample that killed a simpler algorithm, + // four of them ported from #3773 along with the function. They are not + // decoration: drop any one and a refuted algorithm passes again. + const endings = (text) => (text.match(/\r?\n/g) || []).map((b) => (b === '\r\n' ? 'CRLF' : 'LF')); + const assertUniformCrlf = (text) => { + assert.ok(!/[^\r]\n/.test(text) && !/^\n/.test(text), `mixed line endings: ${JSON.stringify(text)}`); + }; + + test('B4: an UNDELIMITED single-entry list at EOF inserts CRLF, not a bare LF', () => { + // #3773's B4, the defect this PR must not inherit while absorbing that PR. + // With no `## Deferred Items` heading and exactly ONE entry, the entry-list + // region runs from the first entry's start to the insertion point — and + // those are the same offset, so the region is empty and `crlfAtEof('')` is + // `false` by its own `before.length > 0` guard. The scope widens to + // everything before the insertion point rather than asserting LF from no + // evidence at all. + const content = 'preamble\r\n\r\n- alpha'; + const ack = acknowledgeDeferredItem(content, 'alpha'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, 'preamble\r\n\r\n- alpha\r\n status: acknowledged'); + assert.strictEqual(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + assertUniformCrlf(ack.content); + }); + + test('B4 control: the same shape under LF stays LF', () => { + const content = 'preamble\n\n- alpha'; + const ack = acknowledgeDeferredItem(content, 'alpha'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, 'preamble\n\n- alpha\n status: acknowledged'); + }); + + test('an LF-dominant document with ONE stray CRLF: the EOF entry does not inherit it', () => { + // Ported from #3773, and the fixture that refutes THIS PR's round-3 + // algorithm. `delta`'s CRLF terminates DELTA, not the unterminated `beta` + // after it, so copying the preceding separator propagates an isolated CRLF + // into an otherwise-LF file — strictly worse than the bare `\n` it + // replaced. Only a scope with no contradicting bare `\n` may assert CRLF. + const content = '## Deferred Items\n\n- alpha\n- gamma\n- delta\r\n- beta'; + const ack = acknowledgeDeferredItem(content, 'beta'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(ack.content)[3].status, 'acknowledged'); + assert.strictEqual(ack.content, '## Deferred Items\n\n- alpha\n- gamma\n- delta\r\n- beta\n status: acknowledged'); + assert.deepStrictEqual(endings(ack.content), ['LF', 'LF', 'LF', 'LF', 'CRLF', 'LF'], + 'the inserted break must not duplicate the unrelated CRLF above it'); + }); + + test('a DELIMITED single-entry list still has evidence: the section preamble is in scope', () => { + // Ported from #3773. The scope must not shrink to the entry list alone; a + // delimited section's own preamble belongs to that section and counts. + const content = '## Deferred Items\r\n\r\n- alpha thing'; + const ack = acknowledgeDeferredItem(content, 'alpha thing'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, '## Deferred Items\r\n\r\n- alpha thing\r\n status: acknowledged'); + assertUniformCrlf(ack.content); + const lf = '## Deferred Items\n\n- alpha thing'; + assert.strictEqual( + acknowledgeDeferredItem(lf, 'alpha thing').content, + '## Deferred Items\n\n- alpha thing\n status: acknowledged', + ); + }); + + test('a bare LF OUTSIDE the deferred section does not veto the CRLF insert', () => { + // Ported from #3773, and the fixture that rules out scanning the whole + // DOCUMENT: an unrelated bare LF inside a fenced block in another section + // would reject CRLF and drop an isolated LF into an otherwise-CRLF list. + const content = '# Notes\r\n\r\n```text\r\nfirst\nsecond\r\n```\r\n\r\n' + + '## Deferred Items\r\n\r\n- alpha\r\n- beta'; + const ack = acknowledgeDeferredItem(content, 'beta'); + assert.strictEqual(ack.status, 'ok'); + assert.match(ack.content, /- beta\r\n {2}status: acknowledged$/); + const section = ack.content.slice(ack.content.indexOf('## Deferred Items')); + assert.ok(!/(^|[^\r])\n/.test(section), `bare LF in the deferred section: ${JSON.stringify(section)}`); + assert.ok(ack.content.includes('first\nsecond'), 'the unrelated bare LF must not be rewritten'); + }); + + test('an UNDELIMITED list with entries reads only the entry list, not the preamble', () => { + // Ported from #3773, and the fixture that rules out the SECTION as the + // scope: with no heading the section body IS the whole document, so a + // section-scoped scan silently becomes the whole-document scan the case + // above already refuted. The preamble is consulted ONLY when the entry + // list is empty (the B4 case) — here it is not, so the fence's bare LF is + // out of scope and must not veto. + const content = '```text\r\nfirst\nsecond\r\n```\r\n\r\n- alpha\r\n- beta'; + const ack = acknowledgeDeferredItem(content, 'beta'); + assert.strictEqual(ack.status, 'ok'); + assert.match(ack.content, /- beta\r\n {2}status: acknowledged$/); + assert.ok(ack.content.includes('first\nsecond'), 'the unrelated bare LF must not be rewritten'); + }); + + test('no evidence under EITHER scope: an entry at offset 0 stays LF', () => { + // The widened scope terminates rather than recursing outward forever. A + // document that is exactly one unterminated entry has no line ending + // anywhere; LF is the floor, not an invented CRLF. + const ack = acknowledgeDeferredItem('- alpha', 'alpha'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, '- alpha\n status: acknowledged'); + }); + + test('mixed-ending document: the insert branch reads the ENTRY\'s ending, not the file\'s', () => { + // Ported from #3773. M1 asks for the mixed-ending fixture this PR lacked: + // every CRLF test it shipped used UNIFORM CRLF, so the motivating case + // could not fail. A CRLF heading over an LF entry — the opener's own `\n` + // must survive, and a document-wide sniff would rewrite it. + const content = '## Deferred Items\r\n\r\n- alpha\n reason: x\n'; + const ack = acknowledgeDeferredItem(content, parseDeferredItemsWithStatus(content)[0].name); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, '## Deferred Items\r\n\r\n- alpha\n status: acknowledged\n reason: x\n'); + // The single-line twin: a CRLF entry whose only evidence is the separator + // that FOLLOWS it, inside an otherwise-LF document. + const single = '## Deferred Items\n\n- alpha\r\n'; + assert.strictEqual( + acknowledgeDeferredItem(single, parseDeferredItemsWithStatus(single)[0].name).content, + '## Deferred Items\n\n- alpha\r\n status: acknowledged\r\n', + ); + }); + + test('the widened EOF grammar carries the CRLF rule for every marker', () => { + // The EOF fixtures above are all hyphen-shaped because they were inherited + // from a hyphen-only PR. This PR's whole subject is the widened marker set, + // so the EOF rule has to hold across it or the two changes are only + // accidentally compatible. + for (const marker of ['-', '*', '+', '1.']) { + const content = `preamble\r\n\r\n${marker} alpha`; + const ack = acknowledgeDeferredItem(content, 'alpha'); + assert.strictEqual(ack.status, 'ok', `marker ${JSON.stringify(marker)}`); + assert.strictEqual(ack.content, `preamble\r\n\r\n${marker} alpha\r\n status: acknowledged`, + `marker ${JSON.stringify(marker)}: EOF insert must be CRLF`); + assertUniformCrlf(ack.content); + } + }); + + test('acknowledging twice at EOF is idempotent', () => { + const first = acknowledgeDeferredItem('preamble\r\n\r\n- alpha', 'alpha'); + const again = acknowledgeDeferredItem(first.content, parseDeferredItemsWithStatus(first.content)[0].name); + assert.strictEqual(again.status, 'ok'); + assert.strictEqual(again.content, first.content, 'a second acknowledge must not append a second line ending'); + assert.deepStrictEqual(endings(again.content), ['CRLF', 'CRLF', 'CRLF']); + }); +}); + +describe('#3702 round 4: the fence gate is indent-unbounded, like the rest of the grammar (M2)', () => { + // `scanFencedBlocks` is CommonMark, which caps a fence delimiter's indent at + // three spaces. This grammar opted out of that cliff for entry openers + // (`[ \t]*`) and thematic breaks (`^[ \t]*`) but not for fences, so a fence + // at four spaces was not a fence to the gate and a `status: resolved` inside + // it RESOLVED the entry containing it. That is reached by ordinary + // documents, not exotic ones: a fenced block written under a nested bullet + // sits at four spaces. Driven before the fix at indents 4, 5, 8 and a + // leading tab — all four silently resolved. + // + // `gsd-core/references/executor-examples.md` states flatly that "nothing + // inside a fenced code block is an entry or a field". These pin that claim + // instead of quietly bounding it. + const H = '## Deferred Items\n\n'; + const fencedStatusDoc = (pad, delim = '```') => + `${H}- alpha\n${pad}${delim}text\n${pad}status: resolved\n${pad}${delim}\n`; + + for (const [label, pad] of [ + ['0 spaces', ''], ['1 space', ' '], ['2 spaces', ' '], ['3 spaces (CommonMark cap)', ' '], + ['4 spaces (past the cap)', ' '], ['5 spaces', ' '], ['8 spaces', ' '], + ['a leading tab', '\t'], + ]) { + test(`a fenced "status: resolved" at ${label} does not resolve the entry`, () => { + const items = parseDeferredItemsWithStatus(fencedStatusDoc(pad)); + assert.strictEqual(items.length, 1, `${label}: expected exactly one entry`); + assert.notStrictEqual(String(items[0].status || '').toLowerCase(), 'resolved', + `${label}: the fenced status line must not resolve the entry`); + }); + } + + test('the same rule holds for tilde fences past the cap', () => { + const items = parseDeferredItemsWithStatus(fencedStatusDoc(' ', '~~~')); + assert.strictEqual(items.length, 1); + assert.notStrictEqual(String(items[0].status || '').toLowerCase(), 'resolved'); + }); + + test('an entry-shaped line inside a deep fence is content, not an entry', () => { + // The other half of the doc's claim: not an entry EITHER. #3702's wild + // records carry reproduction blocks whose `- ` and `1. ` lines are prose. + const doc = `${H}- alpha\n \`\`\`diff\n - not an entry\n 1. also not an entry\n \`\`\`\n- beta\n`; + const names = parseDeferredItems(doc).map((i) => i.name); + assert.ok(names.some((n) => n.startsWith('alpha')), `alpha missing: ${JSON.stringify(names)}`); + assert.ok(names.includes('beta'), `beta missing: ${JSON.stringify(names)}`); + assert.ok(!names.includes('not an entry'), `fenced content became an entry: ${JSON.stringify(names)}`); + assert.ok(!names.includes('also not an entry'), `fenced content became an entry: ${JSON.stringify(names)}`); + }); + + test('an UNTERMINATED deep fence runs to the end of its ENTRY (round 5, B1) — the status inside it stays gated, the next entry does not', () => { + // Round 4 titled this "runs to end-of-file, exactly as CommonMark says"; + // CommonMark bounds a fence by its container, and the review's B1 showed + // the end-of-file reading swallowing the entries after a stray delimiter. + // The engine still classifies; the walk bounds it at the entry. + const doc = `${H}- alpha\n \`\`\`text\n status: resolved\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + assert.notStrictEqual(String(items[0].status || '').toLowerCase(), 'resolved'); + const two = parseDeferredItemsWithStatus(`${doc}- beta\n status: resolved\n`); + assert.deepStrictEqual(two.map((i) => [i.name, i.status]), [['alpha ```text status: resolved', ''], ['beta status: resolved', 'resolved']]); + }); + + test('a deep fence still CLOSES: fields after it are read again', () => { + // The gate must not swallow the rest of the entry. If the closer at the + // same deep indent were not recognised, everything after it would stay + // fenced to EOF and the real status line would be invisible. + const doc = `${H}- alpha\n \`\`\`text\n status: resolved\n \`\`\`\n status: acknowledged\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + assert.strictEqual(String(items[0].status).toLowerCase(), 'acknowledged', + 'the post-fence status line must still be read'); + }); + + test('acknowledge agrees with the reader about a deep fence', () => { + // Reader and writer share one classifier; a fence the reader honours must + // be one the writer refuses to rewrite into. Otherwise acknowledge writes + // inside the code block and reports ok. + const doc = `${H}- alpha\n \`\`\`text\n status: pending\n \`\`\`\n`; + const ack = acknowledgeDeferredItem(doc, parseDeferredItemsWithStatus(doc)[0].name); + assert.strictEqual(ack.status, 'ok'); + assert.ok(ack.content.includes(' status: pending'), + 'the fenced line must be left exactly as written'); + assert.strictEqual(String(parseDeferredItemsWithStatus(ack.content)[0].status).toLowerCase(), 'acknowledged', + 'the entry must read back acknowledged through an inserted line, not a rewritten fenced one'); + }); + + test('GUARD: `## Gaps` is untouched — it opts out of block structure entirely', () => { + // Both marker-parameterised call sites gate on `markers.blockStructure`, + // which the Gaps set does not set, so Gaps reaches an EMPTY fenced set and + // this change cannot reach it. Pinned rather than asserted: a future + // "consistency" edit that dropped that gate would silently move Gaps, and + // this PR's central claim is that Gaps keeps its `next` behaviour. + // + // The observable form of "no fence gate": Gaps folds the fenced lines into + // the entry's NAME rather than hiding them, and does so identically at an + // indent inside CommonMark's cap and one past it. If the deferred gate + // ever leaked into Gaps, the deep form would start differing from the + // shallow one. + const deep = '## Gaps\n\n- alpha\n ```text\n status: open\n ```\n'; + const shallow = '## Gaps\n\n- alpha\n ```text\n status: open\n ```\n'; + assert.deepStrictEqual(parseUatItems(deep), parseUatItems(shallow), + 'Gaps must read a fence at 4 spaces exactly as it reads one at 2'); + assert.deepStrictEqual(parseUatItems(deep), [ + { name: 'alpha ```text status: open ```', result: 'open', category: 'unknown' }, + ], 'Gaps folds fenced lines into the entry name — no fence gate at any indent'); + + // And the entry-shaped twin: a `- ` line inside a deep fence is still not + // a separate Gaps entry, because Gaps never split on it to begin with. + const entryShaped = '## Gaps\n\n- alpha\n ```diff\n - fenced line\n ```\n- beta\n'; + assert.deepStrictEqual(parseUatItems(entryShaped).map((i) => i.name), + ['alpha ```diff - fenced line ```', 'beta']); + }); +}); + +describe('#3702 round 4: the minors (m1, m2, m3, m5)', () => { + const H = '## Deferred Items\n\n'; + const names = (doc) => parseDeferredItems(doc).map((i) => i.name); + + // ── m1: the ordinal digit cap needs the limit-1 case ────────────────────── + // RULESET.TESTS.boundary-coverage asks for N ∈ {limit-1, limit, limit+1}. + // The limit (`999999999.`) and limit+1 (`1234567890.`) were already pinned; + // limit-1 was the missing third. It is not a formality: an off-by-one in the + // `\d{1,9}` bound shows up at eight digits, not at nine. + test('m1: an EIGHT-digit ordinal (limit-1) is a marker', () => { + assert.deepStrictEqual(names(`${H}1. a\n12345678. b\n`), ['a', 'b']); + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('12345678. x'), true); + }); + + test('m1: the three boundary points read consistently through the marker set', () => { + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('12345678. x'), true, 'limit-1'); + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('999999999. x'), true, 'limit'); + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('1234567890. x'), false, 'limit+1'); + }); + + // ── m2: the hand-rolled CommonMark copies, checked against CommonMark ───── + // `THEMATIC_BREAK_RE` and the tab-expanding indent counter are fresh + // implementations of rules CommonMark already specifies, and this repo has + // no sectionizer helper for either to be compared against. So the parity + // assertion is against the SPEC, with the two deliberate divergences named + // rather than left for a later reader to discover and "fix" back. + test('m2: THEMATIC_BREAK_RE agrees with CommonMark on `-`, `*` and `_`', () => { + const isBreak = (line) => { + const doc = `${H}- alpha\n${line}\n- beta\n`; + // A separator closes the list, so `beta` opens a NEW list rather than + // continuing alpha's. If the line is not a break, it is swallowed as + // alpha's continuation text. + return names(doc).length === 2 && names(doc)[0] === 'alpha'; + }; + // `-- -` belongs in the YES list: CommonMark asks for three or more + // matching characters "each followed optionally by any number of spaces or + // tabs", so the spacing between them is free. It was in the NO list on the + // first cut of this test and the parser was right, not the fixture. + for (const yes of ['---', '***', '___', '- - -', '* * *', '_ _ _', '-----', ' ---', '-- -']) { + assert.strictEqual(isBreak(yes), true, `CommonMark thematic break not recognised: ${JSON.stringify(yes)}`); + } + for (const no of ['--', '**', '__', '- -', 'a---']) { + assert.strictEqual(isBreak(no), false, `not a CommonMark thematic break, but treated as one: ${JSON.stringify(no)}`); + } + }); + + test('m2: the two DELIBERATE divergences from CommonMark are pinned, not accidental', () => { + const isBreak = (line) => { + const doc = `${H}- alpha\n${line}\n- beta\n`; + return names(doc).length === 2 && names(doc)[0] === 'alpha'; + }; + // (1) `+` is NOT a CommonMark thematic-break character. It is one here, + // because `+` IS a list marker in this grammar, so `+ + +` would otherwise + // open a phantom entry named `+ +` — the same reason `* * *` is a break + // rather than an entry named `* *`. + assert.strictEqual(isBreak('+ + +'), true, + '`+ + +` must be a separator here even though CommonMark says otherwise'); + // (2) Indent is UNBOUNDED. CommonMark stops recognising a thematic break at + // four spaces (it becomes indented code); this grammar reads entries and + // breaks at any indent, so the break must follow the entries. + assert.strictEqual(isBreak(' ---'), true, 'a 7-space break must still close the list'); + assert.strictEqual(isBreak('\t---'), true, 'a tab-indented break must still close the list'); + }); + + test('m2: the indent counter expands tabs to CommonMark 4-column stops', () => { + // Not directly exported, so it is measured through its observable effect: + // a continuation line must land at or past its opener's indent column. A + // tab counted as ONE column instead of expanding to the next stop of four + // puts a 3-space continuation "outside" a tab-indented bullet. + const tabOpener = `${H}\t- alpha\n\t status: acknowledged\n`; + assert.deepStrictEqual( + parseDeferredItemsWithStatus(tabOpener).map((i) => i.status), + ['acknowledged'], + 'a tab-indented entry must read its own tab-indented field', + ); + // A 4-space opener and a tab opener occupy the same column, so the same + // continuation depth works for both. + const spaceOpener = `${H} - alpha\n status: acknowledged\n`; + assert.deepStrictEqual( + parseDeferredItemsWithStatus(spaceOpener).map((i) => i.status), + ['acknowledged'], + ); + }); + + // ── m3: the result union is declared twice; check it BEHAVIOURALLY ──────── + // `src/audit.cts` carries a hand-written structural view of `uat.cjs`, + // including its own copy of the result union. That decoupling is deliberate + // (the CLI does not take a type dependency on the module it `require()`s), + // so the fix is not to delete one copy but to make drift observable: every + // status the writer can actually PRODUCE is driven here, so a status added + // to one union and not the other shows up as a fixture with no counterpart + // rather than as a silent fall-through to the write. + // + // Four of these had no assertion anywhere in the suite before this test. + test('m3: every reachable acknowledge status is driven from a fixture', () => { + const reached = new Map(); + const drive = (label, doc, target) => { + const r = acknowledgeDeferredItem(doc, target); + reached.set(r.status, label); + return r; + }; + + drive('ok', `${H}- alpha\n`, 'alpha'); + drive('not_found', `${H}- alpha\n`, 'no such entry'); + drive('ambiguous', `${H}- alpha\n- alpha\n`, 'alpha'); + const resolvedDoc = `${H}- alpha\n status: resolved\n`; + drive('already_resolved', resolvedDoc, parseDeferredItemsWithStatus(resolvedDoc)[0].name); + // #3781 (merged round 5): the heading shape acks; the refusal is reached + // only by a span with a GFM table row INSIDE it, which no write can anchor. + const headingDoc = `${H}### Entry one\n\n- **Status:** open\n| a | b |\n- more\n`; + drive('unsupported_heading_shape', headingDoc, parseDeferredItemsWithStatus(headingDoc)[0].name); + + assert.deepStrictEqual( + [...reached.keys()].sort(), + ['already_resolved', 'ambiguous', 'not_found', 'ok', 'unsupported_heading_shape'], + 'a reachable status stopped being reachable, or a new one appeared', + ); + + // `match_verification_failed` is the sixth member and is NOT driven here. + // It is a defensive re-verify of the matched span against the target text, + // computed by an independent code path, and no fixture reaches it — the + // same position `rewrite_not_readable` was in before round 4 removed it. + // Stated rather than quietly omitted: if this union is ever pruned to what + // tests reach, that member is the one to weigh, and the answer is the same + // one round 4 gave for the other — an assertion nothing can drive is a + // specification nothing holds. + }); + + test('m3: each reachable status produces a DISTINCT outcome, so a merge of two would show', () => { + const doc = `${H}- alpha\n`; + const ok = acknowledgeDeferredItem(doc, 'alpha'); + const notFound = acknowledgeDeferredItem(doc, 'nope'); + assert.notStrictEqual(ok.content, doc, 'ok must change the content'); + assert.strictEqual(notFound.content, doc, 'a refusal must return the content untouched'); + // Every non-ok status returns the ORIGINAL content: that is the property + // the CLI relies on when it refuses rather than writes. + for (const [label, d, t] of [ + ['ambiguous', `${H}- alpha\n- alpha\n`, 'alpha'], + ['not_found', doc, 'nope'], + ]) { + const r = acknowledgeDeferredItem(d, t); + assert.strictEqual(r.content, d, `${label}: refusal must not mutate content`); + } + }); + + // ── m5: DECLINED, with the measurement that refutes it ─────────────────── + // The review is right that `markers.open`'s `(\s*)` indent group and the + // `/^[ \t]*/` reader disagree about `\f`, `\v` and NBSP. Its prescribed fix + // — "narrow to `([ \t]*)`" — was implemented, measured, and REVERTED, + // because the disagreement is not currently doing any harm and the narrowing + // is. + // + // Driven, both directions: + // indent as shipped narrowed to [ \t]* + // \f entry, status open, ack ok, reads back NO ENTRY + // \v entry, status open, ack ok, reads back NO ENTRY + // NBSP entry, status open, ack ok, reads back NO ENTRY + // + // The whole round-trip already works for these: the entry surfaces, its + // field parses, `acknowledgeDeferredItem` returns ok and the result reads + // back acknowledged. Narrowing turns three working shapes into three + // SILENTLY DROPPED ones — which is the #3702 defect class itself, and the + // opposite of the fail-safe rule this file states ("never silently drop a + // possibly-open item"). A latent inconsistency in the safe direction is not + // worth trading for a live regression in the unsafe one. + // + // Pinned here so the prescription cannot be re-applied without failing a + // test that explains why. If the inconsistency is ever to be closed, the + // direction is to make the READERS agree with the opener — `indentOf` counts + // non-tab whitespace as a column instead of terminating on it — not to make + // the opener reject lines it currently accepts. + for (const [label, indent] of [['a form feed', '\f'], ['a vertical tab', '\v'], ['an NBSP', '\u00a0']]) { + test(`m5: an entry indented with ${label} still round-trips (narrowing would drop it)`, () => { + const doc = `${H}${indent}- alpha\n status: open\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1, `${label}: the entry must surface`); + assert.strictEqual(items[0].status, 'open', `${label}: its field must parse`); + const ack = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(ack.status, 'ok', `${label}: it must be acknowledgeable`); + assert.strictEqual( + String(parseDeferredItemsWithStatus(ack.content)[0].status).toLowerCase(), + 'acknowledged', + `${label}: and the acknowledgement must read back`, + ); + }); + } + + test('m5 GUARD: `## Gaps` and the deferred set read exotic indent IDENTICALLY today', () => { + // The two sets differing here is what a narrowing would introduce. Both + // use `\s*` now, so a form-feed-indented bullet is an item on both paths. + // If a later edit narrows only one, this fails. + const gaps = parseUatItems('## Gaps\n\n\f- alpha\n').map((i) => i.name); + const deferred = parseDeferredItems('## Deferred Items\n\n\f- alpha\n').map((i) => i.name); + assert.deepStrictEqual(gaps, ['alpha'], 'Gaps must read a form-feed-indented bullet'); + assert.deepStrictEqual(deferred, ['alpha'], 'the deferred set must read it too'); + }); +}); + +describe('#3702 round 3: heading-path reader/namer inputs (m7, m8)', () => { + test('a bullet whose CONTENT is a fence opener does not suppress the entry\'s fields', () => { + // m7: the fence re-scan used to run on already-marker-stripped lines, so + // `- ```sh` — an ordinary bullet to the splitter — stripped to a fence + // opener that existed in no other pass, and the `**Status:** resolved` + // after it was suppressed as fence content. A RESOLVED entry resurfaced. + const doc = '## Deferred Items\n\n### Entry\n\n- ```sh\n- **Status:** resolved\n- ```\n'; + // The status field must be READ (the round-2 fence suppression is for a + // fence the SPLITTER saw, which this is not) ... + assert.strictEqual(parseDeferredItemsWithStatus(doc)[0].status, 'resolved'); + // ... and a resolved entry must therefore not surface as outstanding. + assert.deepStrictEqual(parseDeferredItems(doc), []); + }); + + test('a heading that begins with a list marker keeps it in the entry name', () => { + // m8: line 0 of a heading entry is the heading TEXT, not a bullet. + // Stripping a marker off it silently renamed the entry — and the name is + // the key acknowledge matches on. + for (const [heading, expected] of [ + ['### 1. Race in the writer', '1. Race in the writer'], + ['### * starred title', '* starred title'], + ['### Race in the writer', 'Race in the writer'], + ]) { + const doc = `## Deferred Items\n\n${heading}\n\n- **What:** x\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1, heading); + assert.ok(items[0].name.startsWith(expected), `${heading} -> ${items[0].name}`); + } + }); + + test('CONTROL: a body bullet keeps its marker in the entry name', () => { + // Held before this round too — a control on the opener-flag threading, not + // a reported defect. The flags say which lines CARRY a marker, but only + // line 0's is part of the entry's identity; wiring the flags into the namer + // wholesale strips the body lines too, which is a rename of the key + // acknowledge matches on. This pins that it did not happen. + const doc = '## Deferred Items\n\n### Entry\n\n- **What:** x\n'; + assert.strictEqual(parseDeferredItemsWithStatus(doc)[0].name, 'Entry - **What:** x'); + }); +}); + +describe('#3702 round 2: marker-grammar parity (M3, N1, N2)', () => { + test('every deferred-items marker regex derives from the one alternation source', () => { + // Structural, not behavioural: both splitter regexes embed the SAME + // source string, so a marker added to one cannot be absent from the other. + // Round 2 ran this over four regexes; the two writer-side ones are gone. + for (const [name, re] of [ + ['open', DEFERRED_BULLET_MARKERS.open], + ['strip', DEFERRED_BULLET_MARKERS.strip], + ]) { + assert.ok(re.source.includes(DEFERRED_MARKER_ALT), `${name}: ${re.source}`); + } + }); + + // #3702 round 3 (M6): the structural assertion above is kept for the two + // SPLITTER regexes, which really are two copies of one alternation. It is no + // longer asked to stand in for the writer/reader agreement — it never could. + // Sharing a source string says nothing about whether the line the writer + // selects is a line the reader reads, which is precisely the asymmetry that + // shipped. The replacement is BEHAVIOURAL and drives the real seam. + test('every marker that OPENS an entry also resolves it through acknowledge', () => { + for (const m of ['-', '*', '+', '1.']) { + assert.ok(DEFERRED_BULLET_MARKERS.open.test(`${m} x`), `open: ${m}`); + // The marker goes on BOTH the opener and the nested status line. Round + // 2's structural test could not reach B1; a replacement that only marks + // the opener cannot either — it is green on the defective build. + const doc = `## Deferred Items\n\n${m} alpha\n ${m} status: pending\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1, `entry surfaced: ${m}`); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok', `ack ok: ${m}`); + const after = parseDeferredItemsWithStatus(got.content); + assert.strictEqual( + after[0].status, 'acknowledged', + `${m}: acknowledge must be READ BACK, not merely written — a status line the reader skips leaves the item outstanding forever`, + ); + } + }); + + test('parity with markdown-sectionizer iterateBullets on the shared vocabulary', () => { + // `iterateBullets` is the repo's other list-marker grammar. On everything + // both grammars are meant to agree on, they do — including the negatives. + const opens = (line) => DEFERRED_BULLET_MARKERS.open.test(line); + const sectionizerOpens = (line) => iterateBullets(line).length === 1; + const shared = [ + ['- x', true], ['* x', true], ['+ x', true], ['1. x', true], ['12. x', true], ['01. x', true], + [' - x', true], ['- [ ] x', true], ['- [x] x', true], + ['**Status:** x', false], ['1) x', false], ['-x', false], ['*x', false], ['1.x', false], + ['prose', false], ['| a | b |', false], ['2026 was a year', false], + ]; + for (const [line, expected] of shared) { + assert.strictEqual(opens(line), expected, `deferred: ${JSON.stringify(line)}`); + assert.strictEqual(sectionizerOpens(line), expected, `sectionizer: ${JSON.stringify(line)}`); + } + }); + + test('the two deliberate divergences from iterateBullets are exactly these', () => { + // N2 — a tab after the marker is CommonMark-legal; `iterateBullets` + // requires a literal space. Kept, and pinned so the difference is visible. + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('-\tx'), true); + assert.strictEqual(iterateBullets('-\tx').length, 0); + // N1 — CommonMark caps an ordered start at 9 digits; `iterateBullets` is + // `\d+`. Ten digits is not a list marker here. + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('1234567890. x'), false); + assert.strictEqual(iterateBullets('1234567890. x').length, 1); + // And `\r` is no longer whitespace after a marker (round 1's `\s` was) — + // nor are NBSP, form-feed or vertical-tab, which `\s` also accepted and + // CommonMark does not: only a space or a tab follows a marker. This is + // the assertion that fails on a `[ \t]` → `\s` revert on its own. + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('-\r'), false); + for (const ws of ['\u00a0', '\f', '\v']) { + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test(`-${ws}x`), false, JSON.stringify(ws)); + } + }); +}); + +describe('#3702 round 2: the ordered marker and the prose contract (B2, m1)', () => { + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + + test('B2: a sentence that happens to open with `.` is prose, not an item', () => { + // Both were items on round 1 — `\d+\.` accepted any digit run. CommonMark + // §5.3's own prose/list discriminator is that an ordered list interrupting + // a paragraph must START AT 1; this parser applies that rule everywhere + // an ordered marker is seen (see `matchListOpener`). + assert.deepStrictEqual(names('2026. was a bad year for this module\n'), []); + assert.deepStrictEqual(names('### Notes\n\n3. is the number of retries we settled on.\n'), []); + // And the mixed heading case: prose under one heading, a list under another. + const got = names('### Notes\n\n3. is the number of retries.\n\n### Steps\n\n1. do this\n2. then this\n'); + assert.strictEqual(got.length, 1, JSON.stringify(got)); + assert.match(got[0], /^Steps/); + }); + + test('B2: an ordered list that starts at 1. counts, at any later number in the run', () => { + assert.deepStrictEqual(names('1. alpha\n2. beta\n3. gamma\n'), ['alpha', 'beta', 'gamma']); + // CommonMark ignores the numbers after the first — so does the run. + assert.deepStrictEqual(names('1. alpha\n3. gamma\n7. delta\n'), ['alpha', 'gamma', 'delta']); + assert.deepStrictEqual(names('01. alpha\n02. beta\n'), ['alpha', 'beta']); + // Heading shape: the run is per entry body. + assert.strictEqual(names('### Steps\n\n1. do\n2. then\n').length, 1); + // Status fields under an ordered run still resolve their entry. + assert.deepStrictEqual(names('1. alpha\n status: resolved\n2. beta\n'), ['beta']); + }); + + test('B2: the rule\'s stated cost — a run that does not start at 1 is prose', () => { + // Pinned so the trade is visible: a hand-numbered list starting at 2 is + // read as prose, the same way CommonMark refuses it as a paragraph + // interruption. The wild records (#3702) all start at 1. + assert.deepStrictEqual(names('2. alpha\n3. beta\n'), []); + // A bullet does NOT end the list for this purpose (round 5, M2): a list is + // open at the level, so the following non-1 ordinal is an item — CommonMark + // reads `2. gamma` there as a fresh ordered list (start=2), not as text. + assert.deepStrictEqual(names('1. alpha\n- beta\n2. gamma\n'), ['alpha', 'beta', 'gamma']); + }); + + test('round 5 (M1): the ordered-start threshold — `0.` and `1.` open a list, `2.` does not', () => { + // CommonMark §5.2 permits any 1-9-digit start; a `0.`-numbered list is + // ordinary. Refusing `0.` dropped ONLY the first item, because the run + // then started at `1.` — the mixed under-report that looks like a clean + // parse. Boundary: limit-1 / limit / limit+1 of the threshold itself. + assert.deepStrictEqual(names('0. alpha\n1. beta\n2. gamma\n'), ['alpha', 'beta', 'gamma']); + assert.deepStrictEqual(names('1. alpha\n2. beta\n'), ['alpha', 'beta']); + assert.deepStrictEqual(names('2. alpha\n3. beta\n'), []); + assert.deepStrictEqual(names('0. only\n'), ['only']); + assert.deepStrictEqual(names('00. alpha\n01. beta\n'), ['alpha', 'beta']); + // The cost, stated accurately: a list starting at 2 or more reads as + // prose UNTIL its first `0.`/`1.` line — the loss is the prefix. + assert.deepStrictEqual(names('2. alpha\n3. beta\n1. gamma\n'), ['gamma']); + }); + + test('round 5 (M2): a non-1 ordinal is an item wherever a list is already open at its level', () => { + // `1. a` / `- b` / `5. c` folded `5. c` into `b` on round 4 — the + // ordered-run memory was cleared by the bullet. In CommonMark `5. c` there + // is a fresh ordered list (start=5): a non-1 start is refused only where + // it would interrupt a PARAGRAPH, and after a list item it interrupts none. + assert.deepStrictEqual(names('1. a\n- b\n5. c\n'), ['a', 'b', 'c']); + assert.deepStrictEqual(names('- a\n5. c\n'), ['a', 'c']); + // Across a blank line the list is still open (CommonMark: a loose list). + assert.deepStrictEqual(names('- a\n\n5. c\n'), ['a', 'c']); + // A paragraph after the blank ENDS the list; the ordinal after it is prose + // — the round-2 B2 contract, now placed where CommonMark places it. + assert.deepStrictEqual(names('- a\n\nprose here\n5. c\n'), ['a prose here 5. c']); + // Doc start and after a heading are paragraph positions: the B2 pins hold. + assert.deepStrictEqual(names('5. c\n'), []); + assert.deepStrictEqual(names('### Notes\n\n5. c\n'), []); + }); + + test('m1: the 9-digit boundary of an ordered start', () => { + // `999999999.` is a legal CommonMark ordered marker; ten digits is not. + assert.deepStrictEqual(names('1. a\n999999999. b\n'), ['a', 'b']); + // Ten digits: not a marker at all — a lazy continuation of the open item. + assert.deepStrictEqual(names('1. a\n1234567890. b\n'), ['a 1234567890. b']); + // And a ten-digit line cannot open a run on its own. + assert.deepStrictEqual(names('1234567890. b\n'), []); + }); + + test('m1: the indentation cliff is deliberately NOT applied — indent-lenient by design', () => { + // CommonMark reads a 4-space-indented line outside a list as indented + // code. This parser does not: `deferred-items.md` is hand-written with no + // mandated shape, and surfacing a questionable entry beats dropping a real + // one (the #2766 stance). Pinned as a decision, with the 3-space twin that + // both readings agree on. + assert.deepStrictEqual(names(' - x\n'), ['x']); + assert.deepStrictEqual(names(' - x\n'), ['x']); + // Nesting still folds by the indent rule at the 2-space depth executors + // actually write, not only at the 4-space depth round 1 tested. + assert.deepStrictEqual(names('- alpha\n - nested\n- beta\n'), ['alpha - nested', 'beta']); + }); +}); + +describe('#3702 round 2: thematic breaks and fenced code are not list items (M1, M2)', () => { + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + + test('M1: a thematic break opens no entry, whichever character it is drawn with', () => { + // `- - -` was a phantom `"- -"` entry on base already; round 1 added + // `* * *` and `+ + +` to the class — and `* * *` is the separator an + // author writing in the `*` style is most likely to use. + for (const hr of ['* * *', '+ + +', '- - -', '***', '---', '___', ' * * *', '* * * ', '- - - - -']) { + assert.deepStrictEqual(names(`${hr}\n`), [], JSON.stringify(hr)); + assert.deepStrictEqual(names(`### Entry\n\n${hr}\n`), [], `heading: ${JSON.stringify(hr)}`); + } + }); + + test('M1: a thematic break ENDS the open entry rather than joining it', () => { + // CommonMark: a thematic break closes the list. The separator is neither a + // phantom item nor a continuation line of the item above it. + assert.deepStrictEqual(names('- alpha\n\n* * *\n\n- beta\n'), ['alpha', 'beta']); + assert.deepStrictEqual(names('* alpha\n* * *\n* beta\n'), ['alpha', 'beta']); + // Under a heading the break stays a BODY line (round 5, m3): the entry's + // name is what `next` reports for that body, and its span stays contiguous + // for the writer. It is still not evidence — a break alone is no entry. + assert.deepStrictEqual(names('### Entry\n\n- **What:** x.\n\n* * *\n'), ['Entry - **What:** x. * * *']); + assert.deepStrictEqual(names('### Entry\n\n* * *\n'), []); + // `- - - x` is NOT a break (trailing text); it is a `- ` item whose text is `- - x`. + assert.deepStrictEqual(names('- - - x\n'), ['- - x']); + }); + + test('M2: lines inside a fenced code block never open an entry', () => { + // #3702's wild records carry reproduction blocks — `+`-prefixed diff lines + // and `1.`-numbered repro steps are the NORMAL content of such a file. + assert.deepStrictEqual(names('### Entry\n\n```sh\n1. run this\n2. then this\n```\n'), []); + assert.deepStrictEqual(names('```diff\n+ added\n- removed\n```\n'), []); + assert.deepStrictEqual(names('~~~\n* not an item\n~~~\n'), []); + // An unterminated fence runs to the end of its ENTRY, never past it + // (round 5, B1): a stray delimiter before the first item hides nothing. + assert.deepStrictEqual(names('```\n- still fenced\n'), ['still fenced']); + }); + + test('B1 (round 5): an unterminated fence runs to the end of its entry — a stray delimiter cannot swallow later entries', () => { + // The exact review reproduction, at indent 0 and at 4: `next` reports two + // entries and round 4 reported ONE, with `- b` swallowed into `a`'s name. + assert.deepStrictEqual(names('- a\n\n```\n\n- b\n'), ['a ```', 'b']); + assert.deepStrictEqual(names('- a\n\n ```\n\n- b\n'), ['a ```', 'b']); + // A TERMINATED deep fence still gates what it encloses (round 4, M2 holds). + assert.deepStrictEqual(names('- a\n ```\n - not b\n ```\n- b\n'), ['a ``` - not b ```', 'b']); + // Inside its own entry the stray fence still gates: a `status:` under it + // does not resolve the entry, and the NEXT entry is read on its own terms. + const statusesOf = (md) => parseDeferredItemsWithStatus(SECTION + md).map((i) => i.status); + assert.deepStrictEqual(statusesOf('- alpha\n ```\n status: resolved\n- beta\n status: resolved\n'), ['', 'resolved']); + // A delimiter after the bound is a fence in its own right: without the + // rescan the `~~~` pair below would be invisible and beta would resolve. + assert.deepStrictEqual(statusesOf('- a\n```\n- b\n ~~~\n status: resolved\n ~~~\n- c\n'), ['', '', '']); + assert.deepStrictEqual(names('- a\n```\n- b\n ~~~\n status: resolved\n ~~~\n- c\n'), ['a ```', 'b ~~~ status: resolved ~~~', 'c']); + // Heading shape: a heading ends the entry, and the fence with it. (At + // indent 0 the heading TOKENIZER applies CommonMark's own fence rule, so + // `### Next` after a stray delimiter is body text there, exactly as on + // `next`; at four spaces the tokenizer sees no fence and the heading holds.) + assert.deepStrictEqual(names('### Entry\n\n- **What:** x\n\n```\n\n### Next\n\n- y\n'), ['Entry - **What:** x ``` ### Next - y']); + assert.deepStrictEqual(names('### Entry\n\n- **What:** x\n\n ```\n\n### Next\n\n- y\n'), ['Entry - **What:** x ```', 'Next - y']); + // And in the headless region before the first heading (four spaces, for + // the tokenizer reason above — at indent 0 the section has no heading). + assert.deepStrictEqual(names(' ```\n- a\n\n### Entry\n\n- b\n'), ['a', 'Entry - b']); + assert.deepStrictEqual(names('```\n- a\n\n### Entry\n\n- b\n'), ['a ### Entry', 'b']); + }); + + test('M2: a fence inside an entry is continuation, and the entry still parses around it', () => { + const md = '- alpha\n ```sh\n 1. step\n + diff\n ```\n status: resolved\n- beta\n'; + const withStatus = parseDeferredItemsWithStatus(SECTION + md); + assert.strictEqual(withStatus.length, 2, JSON.stringify(withStatus)); + assert.strictEqual(withStatus[0].status, 'resolved'); + assert.deepStrictEqual(names(md), ['beta']); + // Heading shape: the fenced lines are body text, not evidence — the `-` + // line outside the fence is what keeps the entry. + assert.strictEqual(names('### Entry\n\n- **What:** x.\n\n```\n1. repro\n```\n').length, 1); + // And the acknowledge writer's span survives a fenced continuation. + const ack = acknowledgeDeferredItem(SECTION + '- alpha\n ```\n + diff\n ```\n- beta\n', 'alpha ``` + diff ```'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + }); +}); + +describe('#3702 round 2: indent measure is grammar-scoped (review round 6)', () => { + // The deferred grammar measures CommonMark columns (a tab is a jump to the + // next multiple of 4); the Gaps grammar keeps `next`'s raw character count. + // Sharing one measure silently changed Gaps entry boundaries in BOTH + // directions on tab-indented input, breaking the `blockStructure: false` + // opt-out's byte-for-byte promise. + const gapsNames = (body) => parseUatItems(['# UAT', '', '## Gaps', '', body, ''].join('\n')).map((i) => i.name); + const deferredNames = (body) => parseDeferredItems('## Deferred Items\n\n' + body + '\n').map((i) => i.name); + + test('Gaps: a tab-indented item followed by a two-space one stays ONE entry, as on `next`', () => { + assert.deepEqual(gapsNames('\t- first item\n - second item'), ['first item - second item']); + }); + + test('Gaps: a two-space item followed by a tab-indented one stays TWO entries, as on `next`', () => { + assert.deepEqual(gapsNames(' - first item\n\t- second item'), ['first item', 'second item']); + }); + + test('Gaps: four spaces then two spaces splits, and a tab pair splits — unchanged either way', () => { + assert.deepEqual(gapsNames(' - first item\n - second item'), ['first item', 'second item']); + assert.deepEqual(gapsNames('\t- first item\n\t- second item'), ['first item', 'second item']); + }); + + test('deferred: the SAME tab/space pairs measure in columns — the opposite verdict, by design', () => { + assert.deepEqual(deferredNames('\t- first\n - second'), ['first', 'second']); + assert.deepEqual(deferredNames(' - first\n\t- second'), ['first - second']); + }); +}); + +describe('#3702 round 2: round-review refinements (ordered run, rejected ordinals, breaks, fenced fields, Gaps scope)', () => { + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + const statuses = (md) => parseDeferredItemsWithStatus(SECTION + md).map((i) => i.status); + + test('an ordered run ENDS at a paragraph that follows a blank line (CommonMark §5.3), but survives lazy continuation', () => { + // A blank line then a non-indented, non-list line is a paragraph: the list + // is over, and `5. x` after it is prose folded into the open entry. + const got = names('1. a\n\nparagraph\n\n5. x\n'); + assert.strictEqual(got.length, 1, JSON.stringify(got)); + assert.match(got[0], /^a/); + // No blank line → lazy continuation → the list is still open and `2. b` is an item. + assert.deepStrictEqual(names('1. a\nlazy continuation\n2. b\n'), ['a lazy continuation', 'b']); + // Heading shape carries the same rule per body. + assert.strictEqual(names('### Steps\n\n1. do\n\nsome prose.\n\n4. not an item\n').length, 1); + }); + + test('an accepted opener clears the blank-line memory — lazy continuation right after it keeps the run', () => { + // Round-review continuation: `blankSeen` survived the opener branch, so + // `2. b` + a lazy line ended the run and `3. c` folded into `b`. + assert.deepStrictEqual(names('1. a\n\n2. b\nlazy continuation\n3. c\n'), ['a', 'b lazy continuation', 'c']); + }); + + test('a headless region of a heading-shaped file applies the SAME paragraph reset — opener flags come from the splitter, not a re-derivation', () => { + // Round-review continuation: a re-derived flag set re-accepted `3.` under a + // stale run after the paragraph had ended it, and stripped it into a field. + const md = '1. alpha\n\nparagraph\n\n3. status: resolved\n\n### Entry\n\n- **What:** x\n'; + assert.deepStrictEqual(statuses(md), ['', '']); + assert.strictEqual(names(md).length, 2); + }); + + test('an ordered run is per INDENT: a nested `1. / 2.` run resolves (round-1 parity), and a nested ordinal after a nested bullet continues that level\'s list', () => { + // Round-review continuation 2: nested openers read the top-level run and + // never wrote their own. + for (const eol of ['\n', '\r\n']) { + const nestedRun = '- alpha\n 1. what: detail\n 2. status: resolved\n\n### Entry\n\n- **What:** x\n'.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(nestedRun), ['resolved', ''], JSON.stringify(eol)); + // Round 5 (M2): a list is open at the nested level, so `3.` is an item + // there — a fresh ordered list in CommonMark — and a nested item that + // reads `status: resolved` is a field line, as `- status: resolved` is. + const nestedAfterBullet = '1. alpha\n - nested item\n 3. status: resolved\n\n### Entry\n\n- **What:** x\n'.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(nestedAfterBullet), ['resolved', ''], JSON.stringify(eol)); + // With NO list open at the nested level the ordinal is prose (round 2). + const leak = '1. alpha\n nested prose\n 3. status: resolved\n\n### Entry\n\n- **What:** x\n'.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(leak), ['', ''], JSON.stringify(eol)); + } + // A new top-level item resets the nested levels: `2.` under beta does not continue alpha's nested run. + assert.deepStrictEqual(statuses('- alpha\n 1. a\n- beta\n 2. status: resolved\n'), ['', '']); + // Under a heading the same per-indent rule applies. + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n 1. step\n 2. **Status:** resolved\n'), ['resolved']); + }); + + test('a DEDENTING top-level list keeps its entry boundaries — every indent at or above the base is one level', () => { + // Round-review continuation 3: the exact-indent run lookup rejected the + // shallower ordinals, collapsing three entries into one. + assert.deepStrictEqual(names(' 1. alpha\n 2. beta\n3. gamma\n'), ['alpha', 'beta', 'gamma']); + assert.deepStrictEqual(names(' - alpha\n- beta\n - gamma\n'), ['alpha', 'beta - gamma']); + }); + + test('indent is measured in CommonMark COLUMNS (a tab advances to the next multiple of 4), so a tab and a space are different levels', () => { + // Round-review continuation 3: character counting aliased `\t` and ` `. + // Heading shape, where every ACCEPTED nested opener is marker-stripped + // before field extraction (the headless path strips line 0 only — #3740). + for (const eol of ['\n', '\r\n']) { + const body = (nested) => `### Entry\n\n- **What:** x\n${nested}`.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(body('\t1. nested\n 2. **Status:** resolved\n')), [''], JSON.stringify(eol)); + assert.deepStrictEqual(statuses(body('\t1. nested\n\t2. **Status:** resolved\n')), ['resolved'], JSON.stringify(eol)); + assert.deepStrictEqual(statuses(body(' 1. nested\n\t2. **Status:** resolved\n')), ['resolved'], JSON.stringify(eol)); + } + }); + + test('a fenced block ends the runs at its indent and deeper, like a paragraph does', () => { + // Round-review continuation 3: a nested run stayed open across a fence, + // so a post-fence `2. status: resolved` resolved the entry. Heading shape, + // for the reason the columns test states. + const body = (nested) => `### Entry\n\n- **What:** x\n${nested}`; + assert.deepStrictEqual(statuses(body(' 1. a\n ```\n code\n ```\n 2. **Status:** resolved\n')), ['']); + // A deeper fence (3 spaces — the sectionizer's CommonMark `{0,3}` limit) leaves the shallower run alone. + assert.deepStrictEqual(statuses(body(' 1. a\n ```\n code\n ```\n 2. **Status:** resolved\n')), ['resolved']); + // Control: without the fence the run continues and resolves. + assert.deepStrictEqual(statuses(body(' 1. a\n 2. **Status:** resolved\n')), ['resolved']); + }); + + test('a REJECTED ordinal line under a heading is not marker-stripped, so it cannot manufacture a field', () => { + // `3. status: resolved` at a PARAGRAPH position is prose by the + // ordered-start rule; before this fix the heading path stripped its marker + // anyway and read a resolved field off it. (Round 5, M2: directly after a + // list item it is an item instead — a list is open there.) + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n\nSome prose.\n3. status: resolved\n'), ['']); + assert.strictEqual(names('### Entry\n\n- **What:** x\n\nSome prose.\n3. status: resolved\n').length, 1); + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n3. status: resolved\n'), ['resolved']); + // An ACCEPTED ordered status line still resolves, as `- status: resolved` does. + assert.deepStrictEqual(statuses('### Entry\n\n1. **What:** x\n2. **Status:** resolved\n'), ['resolved']); + // Same rule in a headless region of a heading-shaped file. + assert.deepStrictEqual(statuses('- alpha\n 3. status: resolved\n\n### Entry\n\n- **What:** x\n'), ['', '']); + }); + + test('a thematic break is recognised at any indent — the parser is indent-lenient for breaks as it is for items', () => { + assert.deepStrictEqual(names(' * * *\n'), []); + assert.deepStrictEqual(names('- alpha\n\n - - -\n\n- beta\n'), ['alpha', 'beta']); + }); + + test('a status line orphaned after a break leaves its entry OPEN — the fail-safe polarity, pinned', () => { + // `- alpha\n---` is a list then a thematic break in CommonMark; the indented + // line after it belongs to nothing. Surfacing alpha is the safe direction. + assert.deepStrictEqual(names('- alpha\n---\n status: resolved\n'), ['alpha']); + }); + + test('fenced lines carry no FIELDS either — a fenced `status: resolved` does not resolve the entry', () => { + assert.deepStrictEqual(statuses('- alpha\n```\nstatus: resolved\n```\n'), ['']); + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n```\n- **Status:** resolved\n```\n'), ['']); + assert.deepStrictEqual(statuses('- alpha\n ```yaml\n status: resolved\n ```\n status: acknowledged\n'), ['acknowledged']); + }); + + test('`## Gaps` keeps its round-1 grammar byte-for-byte: no fence or break awareness there', () => { + // Block structure (M1/M2) is scoped to the deferred grammar via + // `BulletMarkers.blockStructure`; the Gaps section is template-mandated and + // out of #3702's blast radius, so a fenced hyphen line still counts there, + // and a fenced field is still read — exactly as on `next`. + // + // THE SECOND ASSERTION tracks `next`'s #3898 fix, which landed in the + // base range this branch merged (b431ae9f0): a spaced hyphen thematic + // break in `## Gaps` is a SEPARATOR — skipped between entries — so + // `- - -` no longer surfaces a phantom open gap named `- -`. Until that + // fix this line pinned the phantom on purpose (round 4, m4), because it + // was the only assertion that would notice the Gaps path moving; it still + // is, and it now pins the fixed reading. The full separator table lives in + // the #3898 describe block above; this one keeps the Gaps opt-out honest. + const uat = ['---', 'status: partial', 'phase: 01-x', '---', '', '## Gaps', '', '```', '- truth: phantom', ' status: open', '```', ''].join('\n'); + const got = parseUatItems(uat); + assert.deepStrictEqual(got.map((i) => i.name), ['phantom'], JSON.stringify(got)); + const withBreak = ['---', 'status: partial', 'phase: 01-x', '---', '', '## Gaps', '', '- - -', '- truth: real', ' status: open', ''].join('\n'); + assert.deepStrictEqual(parseUatItems(withBreak).map((i) => i.name), ['real']); + }); +}); + // ─── Bug 3: table-shaped ## Gaps section ────────────────────────────────────── describe('#2766 parseGapsItems: GFM table shape', () => { @@ -6858,6 +8168,66 @@ describe('#3781: acknowledge supports the heading-delimited entry shape', () => assert.equal(ack.content, doc, 'file unchanged'); }); + test('a table row AFTER the entry\'s last line is outside its span — the write lands above it', () => { + // #3781 refused any leaf whose heading-to-next-heading range held a table + // row. The refusal exists because a row INSIDE the span makes it + // non-contiguous; a row after the last entry line is not inside anything, + // and refusing it halts `complete-milestone` over a write that is safe. + const doc = '## Deferred Items\n\n### Finding one\n- did a thing\n| x | y |\n'; + const before = parseDeferredItemsWithStatus(doc); + assert.equal(before.length, 2, 'fixture self-check: leaf entry + table row'); + const ack = acknowledgeDeferredItem(doc, before[0].name); + assert.equal(ack.status, 'ok'); + assert.equal(ack.content, '## Deferred Items\n\n### Finding one\n- did a thing\n status: acknowledged\n| x | y |\n'); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + }); + + test('an entry whose body ENDS in a fence acks readably — the marker lands before the fence, never inside it (round 5, RV6.5)', () => { + // Unclosed fence: it runs to the entry's end, so "after the last + // non-blank line" is fence content — the reader never sees the marker. + const unclosed = '## Deferred Items\n\n### E\n- **What:** x\n```\ncode\n'; + let ack = acknowledgeDeferredItem(unclosed, parseDeferredItemsWithStatus(unclosed)[0].name); + assert.equal(ack.status, 'ok'); + assert.equal(ack.content, '## Deferred Items\n\n### E\n- **What:** x\n status: acknowledged\n```\ncode\n'); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged', 'the marker must be read back'); + // Closed fence at the end of the body: same placement, for the same reason. + const closed = '## Deferred Items\n\n### E\n- **What:** x\n```\ncode\n```\n'; + ack = acknowledgeDeferredItem(closed, parseDeferredItemsWithStatus(closed)[0].name); + assert.equal(ack.status, 'ok'); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + assert.ok(ack.content.includes('- **What:** x\n status: acknowledged\n```'), ack.content); + // A pending (preamble) entry ending in an unclosed fence before a heading. + const pending = '## Deferred Items\n\n- a\n```\ncode\n\n### E\n- b\n'; + const items = parseDeferredItemsWithStatus(pending); + assert.equal(items.length, 2, JSON.stringify(items)); + ack = acknowledgeDeferredItem(pending, items[0].name); + assert.equal(ack.status, 'ok'); + assert.deepStrictEqual(parseDeferredItemsWithStatus(ack.content).map((i) => i.status), ['acknowledged', '']); + assert.ok(ack.content.startsWith('## Deferred Items\n\n- a\n status: acknowledged\n```\ncode\n'), ack.content); + }); + + test('a heading whose TEXT is a fence delimiter is a heading, not a fence (round 5, RV6.5)', () => { + // The entry-level fence scan saw line 0 (`\`\`\``, the heading text) as an + // opener: every body line was fenced, the reader read no field, and the + // writer's marker landed on a line nothing reads. + for (const delim of ['```', '~~~']) { + const doc = `## Deferred Items\n\n### ${delim}\n- x\n status: resolved\n`; + assert.deepStrictEqual(parseDeferredItemsWithStatus(doc).map((i) => i.status), ['resolved'], delim); + const open = `## Deferred Items\n\n### ${delim}\n- x\n`; + const ack = acknowledgeDeferredItem(open, parseDeferredItemsWithStatus(open)[0].name); + assert.equal(ack.status, 'ok', delim); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged', delim); + } + }); + + test('leaf line-0 rewrite keeps a closing `#` sequence (round 5, RV6.5)', () => { + const doc = '## Deferred Items\n\n### status: open ###\n- did a thing\n'; + const ack = acknowledgeDeferredItem(doc, parseDeferredItemsWithStatus(doc)[0].name); + assert.equal(ack.status, 'ok'); + assert.ok(ack.content.includes('### status: acknowledged ###'), ack.content); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + }); + test('leaf line-0 corner: a heading whose text parses as a status field', () => { const doc = '## Deferred Items\n\n### status: open\n- did a thing\n'; const before = parseDeferredItemsWithStatus(doc);