From ef9ce3e5982e50aca8233504abb0be1a76a1291a Mon Sep 17 00:00:00 2001 From: 0xdhx Date: Mon, 31 Aug 2026 12:44:47 -0500 Subject: [PATCH] fix(#3702): count asterisk, plus and ordered markers as deferred-items entries (#3739) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#3702): deferred-items counts `*`, `+` and ordered markers as list items `deferred-items.md` has no template and no mandated shape, but its parser recognised only the `- ` hyphen marker. Asterisk bullets, plus bullets and dot-terminated ordered lists — all lists in CommonMark and GFM — contributed ZERO entries on both the headless and the heading-delimited path, and a mixed file dropped its non-hyphen entries while keeping their hyphenated siblings, under-reporting without ever looking empty. The restriction was a regex literal inherited from the Gaps seam, where the template genuinely mandates the hyphen YAML-lite form; nothing in the module's stated rationale distinguishes `*` from `-`. Widened on the deferred path only: - `splitGapsEntriesCore`'s entry opener, `extractGapEntryFields`' line-0 strip and `rawGapEntryText`'s line-0 strip take a `BulletMarkers` parameter that DEFAULTS to the hyphen-only set, so `## Gaps` keeps its template-mandated grammar byte-for-byte and the module still has exactly one grouping pass. - `splitDeferredHeadingEntries`' body-bullet test, `stripLeadingBulletMarker` and `acknowledgeDeferredItem`'s status-field regexes move in lockstep — widening what OPENS an entry without widening what is STRIPPED before field extraction would surface an entry that can never resolve. Unchanged, and pinned by tests: prose-only and bare headings still contribute nothing ("prose is not an item"); a table under a leaf heading still yields exactly its rows, since table lines are skipped before the body-marker flag can be set and a `|` row is not a list marker; the paren-terminated ordered form `1)` is out of this fix's scope. * docs(#3702): changeset fragment (pr: 0 placeholder pre-create) * fix(#3702): widen the forensic-audit prose entry rule to match the parser Sibling site of the same defect class, found by a defect-class sweep of the deferred-items consumers. `/gsd-progress` check 7 does NOT go through `gsd-tools query` — it globs `deferred-items.md` and has the model read entries by a prose rule that mandated "one entry per top-level `- ` line". Left as-is, the marker widening would hold on the CLI path while the one consumer that bypasses the parser kept reporting "No unresolved deferred items" for a file written with `*`, `+` or an ordered marker: the same false negative, surviving in the only place the fix could not reach by code. Also pass DEFERRED_BULLET_MARKERS explicitly where the heading path extracts fields. It was already correct — stripLeadingBulletMarker pre-strips the widened set from every line, so the default hyphen strip is a no-op there — but relying on that leaves a detection site and a strip site nominally on different marker sets, which is exactly the asymmetry the BulletMarkers doc comment warns about. Explicit is local; inferred is a trap for whoever edits the strip next. Out of scope, noted rather than fixed: forensic-audit.md globs only `.planning/phases/*/` and so misses archived milestone phases that `scanDeferredItems` covers. Pre-existing, a different defect, and not this issue's ruling. * docs(#3702): note the milestone-close halt for heading-shape non-hyphen files in the changeset A heading-delimited deferred-items.md written with */+/ordered markers previously parsed to zero and closed silently; it now yields entries whose heading shape acknowledgeDeferredItem refuses, halting complete-milestone until hand-edited. User-visible, so the fragment states it. * chore(#3702): set changeset fragment pr to 3739 * fix(#3702): CR-normalise the heading path and the acknowledge writer (review B1, M4, m2) B1 — `splitDeferredHeadingEntries` stored RAW lines; on a CRLF file every body line but the last still carried its `\r`, the `$`-anchored marker strip failed on it, the marker survived into field extraction and the field was lost — a `**Status:** resolved` that was not the file's final line resurfaced its entry as open. The heading path now stores CR-stripped lines like the headless path already did, and the strip regex tolerates a trailing CR on its own. Round 1's CRLF test put `**Status:**` on the last line, the one position `collectSection`'s `.trimEnd()` had already de-CR'd; the new tests put it first and mid-body. M4 (pre-existing on `next`) — `acknowledgeDeferredItem` found the status line on a CR-stripped copy but rewrote the raw line with a `$`-anchored `.*`, which cannot consume `\r`; `replace` returned its input, and the writer reported `ok` over byte-identical content. The rewrite now runs on a CR-stripped line. The comment that claimed `.*$` consumed the `\r` is corrected — it was the bug, stated as the design. m2 — the indent probe for an inserted `status:` line ran on the raw line and fell back to indent 0 on CRLF; it is CR-stripped too. * fix(#3702): derive every deferred-items marker regex from one source (review M3, N1, N2) M3 — round 1 carried the marker alternation in FOUR places: the `BulletMarkers` pair and two inline literals inside `acknowledgeDeferredItem`, under a doc comment saying the interface existed so a detection site and its strip site could not drift. All four now derive from `DEFERRED_MARKER_ALT`; drift is impossible rather than discouraged. A parity test pins the vocabulary against `markdown-sectionizer`'s `iterateBullets` on everything the two grammars are meant to agree on, and names the two points they deliberately differ. N1 — the ordered marker is `\d{1,9}\.` (CommonMark §5.2), not `\d+\.`. N2 — the marker is followed by `[ \t]`, not `\s`, which also accepted `\r`; the tab remains accepted (CommonMark-legal) and the divergence from `iterateBullets`' literal space is pinned rather than papered over. The four regexes are exported for the parity test only. * fix(#3702): an ordered marker opens an entry only from `1.` or inside a run (review B2, m1) B2 — `\d+\.` alone read ordinary prose as a list: "2026. was a bad year for this module" and, under `### Notes`, "3. is the number of retries we settled on." both opened an entry on round 1, the second straight through the "prose is not an item" contract that round's AC4 claimed to preserve. CommonMark §5.3 faces the same ambiguity when an ordered list would interrupt a paragraph and resolves it by requiring the list to start with 1; `matchListOpener` applies that rule wherever an ordered marker is seen, with the run carried per list (headless) or per leaf-heading body. Numbers after the first are ignored, as CommonMark ignores them. Stated cost, pinned: a hand-numbered list starting at 2 reads as prose — every ordered record in the #3702 scan starts at 1. Both reviewer cases are pinned as prose; the ruling's `1. alpha / 2. beta` shape still counts. m1 — the 9-digit boundary is pinned at both sides (`999999999.` opens, ten digits is not a marker), and the 3-vs-4-space indentation cliff is pinned as deliberately NOT applied: the parser is indent-lenient because surfacing a questionable hand-written entry beats dropping a real one. * fix(#3702): thematic breaks close the list and fenced code never opens an entry (review M1, M2) M1 — `- - -` was a phantom `"- -"` entry on base; widening the marker set added `* * *` and `+ + +` to the class, and `* * *` is the separator an author writing in the `*` style is most likely to use. A CommonMark §4.1 thematic break (plus the `+ + +` gesture, which is the same garbage as an entry name) now closes the open entry on the headless path and is dropped from the body on the heading path — neither an item nor a continuation. M2 — neither splitter was fence-aware, so `+ `-prefixed diff lines and `1.`-numbered repro steps inside a code block counted as entries; #3702's wild records carry exactly those blocks. Both splitters now classify lines by the sectionizer's own `scanFencedBlocks` (so `~~~`, indented and unterminated fences behave as `stripFencedCode` would): fence content never opens an entry, is continuation inside an open one — keeping the span invariant `acknowledgeDeferredItem` re-verifies — and is discarded before the first. * test(#3702): range the #2287 deferred-items property over marker × shape × line ending (review B3) The `#2287` property hard-coded `- ` and filtered `\r\n` out of its arbitraries, so the widened marker set — an enumerated domain, exactly what a property is for — was never under it. It now ranges over `{-, *, +, ordered}` × `{headless, heading}` × `{LF, CRLF}`, with the heading shape placing `**Status:**` first or last: the review's prescription (markers × line endings) would not have reached B1, which lives on the heading path only, so the shape axis is the load-bearing addition. Ordered entries are numbered from 1, so the B2 run rule is under the property too. A second property drives `acknowledgeDeferredItem` over every unresolved headless entry across the same marker × line-ending grid — the one that reaches M4 (a CRLF rewrite that reported `ok` and wrote nothing) and m2. * test(#3702): pin the milestone-close halt on a heading-delimited `*`/`+`/`1.` file (review m3) A heading-delimited `deferred-items.md` written with a non-hyphen marker previously parsed to zero entries and let `complete-milestone` close silently; it now yields entries whose heading shape `acknowledgeDeferredItem` refuses, which the milestone loop turns into `record_ack_failure` → exit 1. The loop is prose in a workflow, so the test drives the two CLI calls it makes: `audit-open --json` must list the entry, and `audit-open acknowledge --text ` must refuse with the heading-delimited message and write nothing. * docs(#3702): changeset and forensic-audit prose carry the round-2 grammar The changeset names the CRLF fixes, the ordered start-at-1 rule, thematic breaks and fences. The `/gsd-progress` forensic-audit step is the one prose parser of this file and must state the same grammar the code has. * fix(#3702): round-review refinements — run ends at a paragraph, rejected ordinals unstripped, breaks at any indent, fenced fields, `## Gaps` scope Findings from the pre-push adversarial review of round 2, each pinned: - An ordered run ENDS at a paragraph that follows a blank line (CommonMark §5.3); a non-indented line with no blank before it is lazy continuation and keeps the run open. `1. a` / blank / `paragraph` / blank / `5. x` is one entry, not two. - The heading path strips the marker off every body line before field extraction (#3457); a line whose ordinal `matchListOpener` REJECTED must not be stripped, or "3. status: resolved" as prose loses its `3. ` and reads as a resolved field. `splitDeferredHeadingEntriesDetailed` now carries a per-line opener flag and only accepted openers are stripped — in headless regions of a heading-shaped file too. - A thematic break is recognised at any indent, matching the parser's indent-lenient reading of items; ` * * *` was a phantom `* *`. - Fenced lines carry no FIELDS either: a `status: resolved` quoted inside a code block no longer resolves its entry on either path. - Block structure (breaks, fences) is a property of the GRAMMAR, carried as `BulletMarkers.blockStructure`: the deferred set opts in, the Gaps set does not, so `## Gaps` is byte-for-byte on its `next` behaviour — the round-2 M1/M2 change had reached it through the shared splitter. * test(#3702): the property exercises the rejected-ordinal branch; the N2 control is independent Round review: the widened #2287 property numbered every ordered run from 1 and so never generated an ordinal the start-at-1 rule rejects — it could not tell round 1 from round 2 on B2. Each entry may now carry a decoy prose line beginning with a non-1 ordinal, placed where it cannot end a run (before the first headless entry; first in a heading body), followed by a `status: resolved` that must never become a field; and a decoy-only heading body must yield no entry. The N2 assertion accepted a tab, which round 1's `\s` accepted too, so a `[ \t]` → `\s` revert alone stayed green. NBSP, form-feed and vertical-tab are now asserted refused — the assertion that fails on that revert on its own, and the disclosure that `[ \t]` narrows what round 1 accepted. * fix(#3702): the splitter records its own opener flags; an opener clears the blank-line memory Round-review continuation, two state defects in the ordered-run logic: - `blankSeen` survived the headless splitter's opener branch, so an opener followed by a lazy continuation line read as "paragraph after a blank" and ended the run — `1. a` / blank / `2. b` / lazy / `3. c` folded `c` into `b`. The opener branch now clears it. - The heading path re-derived per-line opener flags for headless regions without the paragraph reset, re-accepting a rejected `3. status: resolved` under a stale run and stripping it into a field. `GapsEntrySpan` now carries the flags the splitter itself computed, and the heading path reads them; the re-derivation is deleted. * fix(#3702): ordered-run memory is per indent — nested runs resolve, nested ordinals never inherit the top-level run Round-review continuation 2: nested openers consulted the TOP-LEVEL run flag and never wrote their own, so a nested `1. / 2.` run under a hyphen entry rejected its `2. status: resolved` (round 1 resolved it), while a nested `3. status: resolved` under a nested `- ` bullet inherited an open top-level run and was stripped into a false field. `OrderedRuns` keys the memory by indent: a new opener at indent d resets every deeper level, a paragraph after a blank at indent d ends the runs at d and deeper, a thematic break or a heading clears all. Both splitters use it; the top level still decides entry boundaries, nested levels decide only which continuation lines are accepted openers for field stripping. Pinned for LF and CRLF. * fix(#3702): run levels — one top level at or above the base, CommonMark column indents, a fence ends its level's runs Round-review continuation 3: - A dedenting top-level list (` 1.` / ` 2.` / `3.`) lost its entry boundaries: the exact-indent run lookup rejected the shallower ordinals before the boundary check ran. Every indent at or shallower than the list's base is now ONE level, in both splitters. - `indentOf` counted characters, so a tab and a space aliased to one level and `\t1. nested` / ` 2. status: resolved` resolved falsely. Indent is now measured in CommonMark columns (§2.2: a tab advances to the next multiple of 4), for the run level and the entry-boundary check alike. - A nested run survived a fenced block. A fence is a non-list block: its opening delimiter ends the runs at its level and deeper, exactly as a paragraph after a blank does. * fix(#3702): the indent measure is grammar-scoped — Gaps keeps next's character count `blockStructure: false` promised the Gaps grammar byte-for-byte parity with `next`, but the CommonMark-column indent measure added for the deferred grammar was shared by the whole splitter core, so tab-indented Gaps input changed entry boundaries in BOTH directions: `\t- a` / ` - b` — next folded into one entry, HEAD split into two ` - a` / `\t- b` — next split into two, HEAD folded into one `indentWidth` now keys the measure on the grammar: columns for the deferred set, raw character count for Gaps. The opt-out covers indent semantics, not only fences and thematic breaks. Four cases pin both halves — the two flipped Gaps pairs, the two Gaps pairs that never moved, and the same tab/space pairs on the deferred path returning the opposite (column-measured) verdict by design. * fix(#3702): the acknowledge path reads and writes through one classifier Round 3, Blockers 1 and 3, and Minors 7 and 8 — one mechanism, so one commit. Every consumer of an entry's lines now reads the splitter's own per-line verdict instead of a re-derivation of it. B1. Round 2 widened the WRITER's status-line finder to the deferred marker set while `extractGapEntryFields` still de-bulleted line 0 only. A nested ` * status: pending` was therefore selectable by the writer and invisible to the reader: acknowledge rewrote it in place, returned `ok`, and the item stayed outstanding on every later audit. Measured against a `next` build, `*`, `+` and `1.` each resolved on base and stopped resolving at round 2's head — a regression, not a gap in new behaviour. The hyphen form of the same shape was already broken on `next` and is fixed here too: one classifier cannot be right for three markers and wrong for the fourth. `parseGapEntryFieldLine` is now the single place a line is classified as a field, and it reports the offset at which the VALUE begins. The rewrite happens at that offset rather than through a second regex, so a line the classifier can select is one whose rewrite it has already located — the selection and the rewrite cannot disagree. Both `DEFERRED_STATUS_FIELD_RE` and `DEFERRED_STATUS_REWRITE_RE` are deleted rather than widened. A read-back guard returns `rewrite_not_readable` rather than `ok`; it is unreachable by construction today and is the fail-loud floor under the next divergence. B3. This is the end state the round-3 review prescribed on both #3739 and #3773: #3773's shared classifier, parameterised by this PR's marker set, with this PR's two status regexes deleted. #3773 lands first. Its hyphen-only strip is consistent with `next`'s hyphen-only splitter today, so the writer/reader divergence is created by THIS merge, which is why widening every consumer belongs to the PR that widens the domain. m7. The heading path marker-stripped its lines before calling the reader, so the reader's fence scan ran over text the splitter never saw: `- ```sh` is an ordinary bullet to the splitter but strips to a fence opener, and a `**Status:** resolved` after it was suppressed as fence content — a resolved entry resurfaced as open. Stripping now happens inside the reader, after the fence scan. m8. `rawGapEntryText` stripped a marker off line 0 unconditionally, but on the heading shape line 0 is the heading TEXT: `### 1. Race in the writer` was silently renamed to `Race in the writer`, and the name is the key acknowledge matches on. Line 0 is stripped only when the splitter accepted it as an opener. Also removed: `splitDeferredHeadingEntries`, whose sole caller only null-checked it (round 3, M4 — the claim was zero callers, which was wrong; the wrapper's `.map` was waste at the one call site), and `stripLeadingBulletMarker`, which this change leaves with no callers at all. The export surface narrows to the two splitter regexes the behavioural parity test reads (M6). [PEER-ASK pr-order-12d5] q: Reviewer blocked both on merge order. I'm declaring #3773 lands first and building the end-state shape into #3739 now (both my status regexes deleted). Does that match your plan? reply: CONFIRMED - same order, derived independently. #3773 cannot carry the fold: `DEFERRED_BULLET_MARKERS`/`BulletMarkers` have zero occurrences at `next` (verified), so the prescribed end state is not executable inside #3773 without absorbing this PR's work. deadline: 03:55 UTC (answered before it) fallback: declare #3773 first, adopt end-state shape in #3739, push+comment decision: proceeded as stated; #3773 lands first, this PR carries the widening of every consumer. Refs #3740 * test(#3702): pin the detect/strip symmetry, and drop a white-box test that could not reach it Round 3, Blocker 2 and Minors 6 and 9. B2. The regression shipped green because no fixture put a marker on a nested status line. Four markers x {nested status line}, each asserting the entry READS BACK as acknowledged rather than that acknowledge merely reported `ok` — reporting `ok` over a line the reader skips is the whole defect. Plus the bare capitalised `Status:` case (the reader stores it case-sensitively, so the writer must not select it), and an idempotence test, which is the failure the defect actually produced: the item resurfaces, is acknowledged again, and never settles. Each of these was run against the pre-fix build first: all five fail there and pass here. Two further assertions in the block are labelled CONTROL because they held pre-fix — they guard the new offset-based rewrite and the opener-flag threading against regressing, and calling them regression tests for a reported defect would overclaim. M6. The round-2 parity test asserted that four writer-side regexes embedded the same source string. That is true of a detect/read asymmetry too, so it could not have caught B1 — and two of the four regexes were widened into `export =` purely to let it read them. Replaced with a behavioural test that drives the real seam: every marker that opens an entry must also resolve it through acknowledge. The structural assertion is kept for the two splitter regexes, which really are two copies of one alternation. m9. `expectedResolved` was computed and immediately voided; the loop beneath it already asserts both polarities. m7/m8 coverage lands here too: a bullet whose content is a fence opener must not suppress the entry's fields, and a heading beginning with a list marker must keep it in the entry name. * docs(#3702): document the deferred-items entry shape where the file is written Round 3, Major 5, and #3702's own item 2. The widened grammar was documented in the reader (`forensic-audit.md`) but not at the write site, where `executor-examples.md` still said only "log to deferred-items.md" — so the question the issue actually raised, which shapes count, remained unanswered anywhere a human writes the file. States what opens an entry (`-`, `*`, `+`, and `1.` when the list starts at `1.`), that `1)` is not a marker here, that a separator closes the list and fenced content is never an entry or a field, and that an entry without an explicit `status: resolved` stays open by design. * chore(#3702): regenerate the changeset through the generator Round 3, Minor 10. The fragment was hand-named against 64 generated names on `next`, and its body ran ~250 words against CONTRIBUTING's one-sentence form. Regenerated via `npm run changeset`, which is also what the random three-word name is for: concurrent PRs never collide. * fix(#3702): the fence gate lives on the seam both sides call, not just the reader Found by the pre-push adversarial review of this round, and it is a regression this round introduced rather than a pre-existing one. `extractGapEntryFields` applied `fencedLineSet` before classifying; the acknowledge writer's status-line search did not. So a `status:` line inside a fenced block was SELECTED by the writer and SKIPPED by the reader — the write produced a line nothing reads, the read-back guard refused it, and the entry became impossible to acknowledge at all: `audit acknowledge` raised an internal error and `complete-milestone` halted on it. Measured, `- alpha` / fence / ` status: pending` / fence: next ack=ok -> reads back "acknowledged" round-2 head ack=ok -> reads back "" (the B1 defect) before this ack=rewrite_not_readable -> refuses entirely (worse than next) `entryFieldLines` is now the seam — per line of an entry, the field it declares or `null`, fences included — and the reader and the writer both go through it. That makes "the writer cannot select a line the reader will not read back" structural rather than asserted, which is what the previous commit's message claimed while a second read-side filter still lived outside the classifier. Two comments corrected with it. The read-back guard is NOT "unreachable by construction": this round shipped a reachable path to it, which is precisely what an invariant asserted in a comment is worth. And the M6 replacement test put its marker only on the entry opener, so it passed against the defective build — the exact weakness it was introduced to fix in round 2's test. It now marks the nested status line too, and fails pre-fix like the rest. Round-3 tests against the pre-fix build: 10 of 12 fail there, and the 2 that hold are labelled CONTROL because they guard this round's new code rather than pin a reported defect. * fix(#3702): one end-of-file CRLF algorithm, adopting #3773's with its B4 closed Round-4 M1. Two open PRs shipped two different answers to "what line ending does an entry that ENDS THE FILE get?", and the review's ruling was that the disagreement needs one answer, not two. Neither shipped answer was that one. Measured on builds of both heads: case #3739 r3 #3773 here undelimited single entry, CRLF preamble pass FAIL pass LF-dominant list, one stray CRLF at EOF FAIL pass pass (the other five) pass pass pass This PR's content.endsWith('\r\n', matchIndexInContent) reads the terminator of the PREVIOUS line, so it propagated an isolated CRLF into an LF-dominant list -- refuted by #3773's own LF-dominant fixture, ported here. Withdrawn. #3773's crlfAtEof asks the right question -- does anything before the entry, within scope, contradict CRLF -- and fails closed. But its scope goes EMPTY for an undelimited single-entry list, because the entry-list region runs from the first entry's start to the insertion point and those coincide; crlfAtEof('') is false by its own before.length > 0 guard, so 'preamble\r\n\r\n- alpha' gained a bare \n in a CRLF document. That is #3773's B4, verified by driving its head. Adopted here with the scope widened to everything preceding the insertion point where the preferred region is empty, rather than asserting LF from no evidence. That only ever loosens a scope carrying zero information, and the predicate stays fail-closed over the wider one. An entry at offset 0 of an undelimited document has no evidence under either scope and stays LF. Tests: 10 added. Negative control, driven -- 1 of the 10 fails against this branch's own pre-fix head (the stray-CRLF fixture); B4 fails against #3773's head; the remaining 8 are the scope counterexamples ported with the function, which were regression pins in #3773 and are guards here. Each still kills a simpler algorithm: drop any one and a refuted scope passes again. Four deferred-items suites 450/450, 0 skipped. npm run lint:ci exit 0. * fix(#3702): drop the unreachable rewrite_not_readable guard (B3) Round-4 B3: the status had zero test coverage in either file. The review offered two branches -- drive it from a test, or delete it and stop carrying an untested terminal status. Taking the second, with the reason stated rather than assumed. Why it cannot be driven. Round 3 added the guard after a fenced `status:` line proved the writer could select a line the reader would not read back. Round 3 then closed that divergence STRUCTURALLY, by routing the writer's line selection and the reader's field extraction through one entryFieldLines seam. The guard now detects a state construction prevents: 21 document shapes were driven against it -- fence openers on the bullet line for every marker in the widened set, duplicate and triplicate status lines, bolded and nested variants, fences between duplicates -- and none reached it. The only seam that would is routing the internal call through the module's exports so a test could stub it, which reshapes production surface for a test. Why leaving it undriven is not free. RULESET.TESTS.mutation-score runs Stryker incrementally over changed files at an 80% threshold and says to treat a surviving mutant as a failing test specification. An undriven `if` on a changed file is exactly that, on both the condition and the .toLowerCase() comparison. What this gives up, stated rather than hidden: if a future change re-splits the writer's selection from the reader's extraction, acknowledgeDeferredItem returns ok over an item that stays outstanding -- the original #3702 defect class. One correction to the review's framing: match_verification_failed does NOT backfill it. That check runs BEFORE the write and compares the matched span to the target, so it cannot see a post-write read-back failure. The protection against re-splitting is the shared seam and the round-3 tests that pin it, not a runtime assertion. A comment at the removal site records all of this. Removing it also drops the union member from both files, which resolves the PR body's internal contradiction (it claimed no type-signature changes while adding one) and the duplicate-status surface #3773 collides on. No test changed behaviour: 450/450 across the four deferred-items suites, 149/149 across the audit suites, npm run lint:ci exit 0 -- the same figures as before the removal, which is itself the evidence that nothing exercised the branch. * fix(#3702): the deferred fence gate is indent-unbounded, like the rest of the grammar (M2) Round-4 M2. scanFencedBlocks is CommonMark, which caps a fence delimiter's indent at three spaces -- a fourth makes it an indented code block instead. This grammar had already opted out of that cliff for entry openers ([ \t]*) and for THEMATIC_BREAK_RE (^[ \t]*), but not for fences. So a fence at four spaces was not a fence to the gate, and a `status: resolved` line inside it RESOLVED the entry containing it. That is not an exotic shape. A fenced block written under a NESTED bullet sits at four spaces, so ordinary hand-written deferred-items.md files reach it. Driven before the fix at indents 4, 5, 8 and a leading tab: all four silently resolved. It is the #3702 silent-resolution defect class in a new place. gsd-core/references/executor-examples.md, added by this PR, states flatly that "nothing inside a fenced code block is an entry or a field". The review offered fixing the parser or bounding that claim in three places. Fixing it -- the claim is the one users will rely on, and the grammar had already chosen unbounded indent everywhere else. NO second fence dialect (the rule blankIndentedFenceDelimiters states). The classification is still done by scanFencedBlocks, the one exported CommonMark state machine, over a de-indented VIEW of the same lines. Run lengths, backtick vs tilde, closer-must-match-and-not-trail, info-string rules and the unterminated-at-EOF case remain that engine's answers. Indent is the only dimension hidden from it, and it is exactly the dimension this grammar has already declared it does not measure. Index alignment is 1:1 -- map preserves length -- so every returned line index still addresses the original line. Scope is the deferred grammar only. Both marker-parameterised call sites gate on markers.blockStructure, which the Gaps set does not set, so Gaps reaches an empty set. Verified, not asserted: the 47-fixture Gaps differential (marker x line-ending x separator x fence x break x key-shape x list-shape) is BYTE-IDENTICAL across this change, 8033 bytes both sides. Tests: 14 added, of which 8 fail against the pre-fix source and pass here; the other 6 are the deliberate controls -- indents 0 through 3, which must NOT move, and the Gaps opt-out guard. Four deferred-items suites green; the 58 suites touching uat/deferred/sectionizer run 6045 tests with an IDENTICAL failing set before and after this change (17 pre-existing environment failures -- installs and an unpinned GSD_EMITTED_BASE; emitted-attribution passes 259/259 in isolation with its base pinned). lint:ci exit 0. * fix(#3702): changeset, both prose parsers, and the minors (M3, M4, m1-m3, m5, n1-n2) M3 -- the changeset omitted a user-BREAKING change. Measured against next: a heading-delimited deferred-items.md written with `*`, `+` or `1.` went from "0 entries, so complete-milestone has nothing to acknowledge and closes" to "1 entry, the CLI writer refuses the heading shape, ACK_FAILURES accumulates, exit 1". The `-` form already halted and is unchanged. That is release-note material: a close that used to succeed now fails, and the correct response is to fix the file, not revert. Also names the fence-indent fix below, and adds #3740 so #3773's issue is attributed here as it is absorbed. M4 -- gsd-core/workflows/progress/steps/forensic-audit.md is a SECOND, model-executed parser of the same grammar, and prose cannot carry a parity test. Its widened text stated the start-at-1 rule, fences and separators but not the `1)` exclusion nor the nine-digit ordinal cap, both enforced in code with pinned tests. Both stated now, along with the round-4 fence-indent rule. (No ack fragment: the size ratchet's currentSizes does a NON-recursive readdirSync of gsd-core/workflows and agents, so a file under workflows/progress/steps/ is outside its scope -- verified by reading the helper, not by the green.) n1 -- executor-examples.md documented that the BOLDED status key is matched case-insensitively and left the bare key's rule to inference. Driven: bare `Status: resolved` is NOT read, so the entry stays open with no warning, while `**Status:**` is. Stated explicitly, with the digit cap and the any-indent fence rule (n2). m1 -- boundary coverage was 2/3. limit (999999999.) and limit+1 (1234567890.) were pinned; limit-1 (12345678.) added, per RULESET.TESTS.boundary-coverage. m2 -- THEMATIC_BREAK_RE and the tab-expanding indent counter are hand-rolled CommonMark rules with no in-repo peer to compare against, so the parity assertion is against the SPEC: eight positive and five negative fixtures, plus the two DELIBERATE divergences pinned as deliberate (`+` is a separator here but not in CommonMark, because `+` is a list marker in this grammar and `+ + +` would otherwise be a phantom entry; indent is unbounded). One fixture was initially wrong -- `-- -` IS a CommonMark break, since the spec allows free spacing between the three characters -- and the parser was right. m3 -- the result union is hand-duplicated in audit.cts as part of a deliberate structural view of uat.cjs, so the fix is not to delete a copy but to make drift observable. Every REACHABLE status is now driven from a fixture; four of the six (ambiguous, unsupported_heading_shape, already_resolved, match_verification_failed) had no assertion anywhere in the suite before this. match_verification_failed is still undriven and the test says so rather than omitting it. m5 -- DECLINED, with the measurement. The review is right that `(\s*)` in the opener and `/^[ \t]*/` in the reader disagree about \f, \v and NBSP, but its prescribed narrowing was implemented, driven and REVERTED: as shipped, an entry indented with any of those surfaces, parses its status field, acknowledges, and reads back acknowledged -- a complete round-trip. Narrowing turns all three into SILENTLY DROPPED entries, which is the #3702 defect class itself and the opposite of this file's stated fail-safe rule. A latent inconsistency in the safe direction is not worth a live regression in the unsafe one. Pinned by three round-trip tests so the prescription cannot be re-applied silently; if it is ever closed, the direction is to make the readers agree with the opener, not to make the opener reject lines it accepts today. Four deferred-items suites 475/475, 0 skipped. lint:ci and lint:changeset exit 0. The 47-fixture Gaps differential is byte-identical at 8033 bytes. * fix(#3702): the pinned `## Gaps` phantom now cites its issue (m4) Round-4 m4. The second assertion in the Gaps byte-for-byte test pins a real defect as expected output: a spaced hyphen thematic break in `## Gaps` is read as an ITEM, so `- - -` surfaces a phantom open gap named `- -`. Reproduced on pristine next at 389bc86e0 across nine separator shapes -- every spaced hyphen form is affected, `---`/`----`/`* * *`/`___` are not, and the dividing line is a space after the first hyphen (the Gaps opener is /^(\s*)(-)\s/ with no thematic-break concept at all). Filed as open-gsd/gsd-core#3898. The pin stays: scope-limiting Gaps is the point of the blockStructure opt-out, and this assertion is the only thing that would notice the Gaps path moving. What was missing was the tracking -- a pinned defect with no issue behind it reads as intended behaviour to the next reader. The comment now says which it is and what the expectation becomes when #3898 lands. * fix(#3702): an unterminated fence runs to the end of its entry, never past it (B1, B2) Round 4 de-indented every line before `scanFencedBlocks`, so a fence opened at any indent — and `scanFencedBlocks` runs an unterminated fence to end-of-document — so one stray delimiter swallowed every entry after it into the entry before it. `- a` / blank / four-space ``` / blank / `- b` yielded ONE entry where `next` yields two: a widening that made an already-counted item vanish, on the mixed-file shape #3702 exists to close. Reproduces at indent 0 as well. The bound is the entry. CommonMark closes a fence with its container and a container at the next item at its level; this parser extends that to a document-level stray delimiter, where CommonMark would swallow to EOF and the fail-safe rule (surface, don't drop) will not. `scanFencesFrom` reports the unterminated opener and the walk supplies the bound — the next line shaped like a top-level item — then RESCANS from it, so a later delimiter is read on its own terms. Still one fence dialect: every block boundary is `scanFencedBlocks`' answer. Entry-scoped `fencedLineSet` (the field reader) already ran an unterminated fence to the end of its lines, so reader and walk agree by construction. Tests: the M2 pin that asserted `[]` for a stray fence before an item flips (the item counts); the round-4 "runs to end-of-file, exactly as CommonMark says" test is retitled — its assertion stands because the bound is the entry — and extended with the next entry; a new block pins the review reproduction at both indents, a terminated deep fence still gating, the gated status inside the bounded fence, the rescan case, and the heading-tokenizer caveat (at indent 0 the tokenizer applies CommonMark's own fence rule, so a heading after a stray delimiter is body text there, exactly as on `next`). Reverted in isolation against the final tree: 3 named tests fail. * fix(#3702): `0.` starts an ordered list (M1) The start-at-1 rule applied unconditionally dropped ONLY the first item of a `0.`-numbered list — the run then started at `1.` — which is the mixed under-report that looks like a clean parse. CommonMark §5.2 permits any 1-9-digit start and a `0.` list is ordinary; a sentence opening with "0." is not a shape anyone writes. The threshold is now `> 1`. The cost is restated accurately in the doc comment and pinned: a list starting at 2 or more, at a paragraph position, reads as prose until its first `0.`/`1.` line — the prefix, not the whole list. Boundary tests at the threshold itself: `0.`, `1.`, `2.` starts, `00.`/ `01.`, and the prefix-loss case. Reverted in isolation: 1 named test fails. * fix(#3702): a non-1 ordinal is an item wherever a list is already open at its level (M2) The per-indent run memory recorded whether the previous opener was ORDERED, so a bullet item closed the run and `1. a` / `- b` / `5. c` folded `5. c` into `b` — another mixed-file under-report. In CommonMark `5. c` there opens a fresh ordered list (start=5): a non-1 start is refused only where it would interrupt a PARAGRAPH (§5.3), and after a list item it interrupts nothing. `ListRuns` now records "a list is open here"; the start rule applies where no list is open at the line's level — the positions a sentence can occupy — so the round-2 B2 pins (doc start, after a heading, after a paragraph) hold unchanged. Two round-2 pins move with it, both CommonMark-backed: `1. alpha` / `- beta` / `2. gamma` is three items, and a nested `3. status:` after a nested bullet is a nested item (a field line, as `- status:` would be); the "rejected ordinal is not stripped" pin is re-anchored at a paragraph position, where it still holds. Reverted in isolation: 4 named tests fail. * docs(#3702): the two runtime-loaded docs state the grammar the parser ships (B3, M4 parity) `executor-examples.md` (the write-site doc) and `forensic-audit.md` check 7 (the model-executed parser) both asserted "never silently drops a possibly-open item" over a grammar that dropped three measured shapes. Both now carry the round-5 grammar — `0.`/`1.` starts, a non-1 ordinal inside an open list, an unclosed fence ending with its own entry — and the fail-safe sentence is kept with what it does NOT cover named beside it: a fenced line, a separator, and an ordered list numbered from `2.` upward at a paragraph position, and nothing else. * docs(#3702): changeset reflects the merged contract The "Breaking, and deliberate: … HALTS complete-milestone" paragraph described a refusal that #3781 removed from `next`; a heading-shaped file written with a newly recognised marker now surfaces its entries and `complete-milestone` acknowledges them in place. The fragment cites #3702 alone — #3740 and #3775 closed on `next` through #3940 and #3989; this PR's shared reader/writer classifier subsumes both fixes rather than closing either issue. The round-5 grammar (ordered start, unclosed fence bound) is stated in the user-facing sentence. * fix(#3702): the heading-shape insert lands on a line the reader reads, and keeps a closing `#` sequence Found by the round's pre-push adversarial review. An entry whose body ends in a fenced block — closed, or unclosed and therefore running to the entry's end — received `status: acknowledged` AFTER its last non-blank line, i.e. as fence content: the writer returned `ok` and the reader never saw the marker, the item stayed outstanding. That is the #3702 class itself (a write nothing reads), on the shape #3781 just opened. The insert now walks back over blank AND fenced lines, classified by the reader's own `fencedLineSet`, so the marker lands on a line the reader reads; pinned as round-trips for an unclosed fence, a closed fence, and a pending entry ending in an unclosed fence before a heading. Reverted in isolation: the round-trip test fails. Separately, the leaf line-0 rewrite (`### status: open ###`) dropped the closing `#` sequence; it is kept now. Cosmetic, pinned. * docs(#3702): the prose parser states the bare-key case rule; both docs say what an unclosed fence does, no more `forensic-audit.md` check 7 called `status: resolved` case-insensitive where the code reads a bare key lower-case only (the bolded form in any case; the value case-insensitively) — `executor-examples.md` already said so, the model-executed parser did not. And both docs claimed "a stray delimiter cannot hide the entries after it", which overstates B1: an UNCLOSED fence ends with its entry; a closed pair of delimiters is a fence, whatever sits between them, as CommonMark reads it. Found by the round's pre-push review. * fix(#3702): a heading whose text is a fence delimiter is a heading, not a fence Second finding of the round's pre-push review, one door over from the first: for a leaf headed `### ```` (or `~~~`) the entry-level fence scan read line 0 — the heading TEXT, not a Markdown line — as a fence opener, so every body line was fenced: the reader read no field under it, and the writer's marker (placed by the same scan) landed on a line nothing reads — `ok`, item outstanding. `entryFencedLines` now owns the entry's fence view for reader and writer alike, and a leaf's line 0 never opens a fence (the leaf tell is `openerFlags[0] === false`; a pending or headless entry's line 0 is a marker line, never a delimiter). Pinned for both delimiters, read and write; reverted in isolation the pin fails. --------- Co-authored-by: CI Rebase Check Co-authored-by: Tom Boucher --- .changeset/eager-orcas-chatter.md | 5 + gsd-core/references/executor-examples.md | 42 + .../progress/steps/forensic-audit.md | 2 +- src/uat.cts | 1473 ++++++++++++----- tests/audit-command-cutover.test.cjs | 45 +- tests/uat.test.cjs | 1412 +++++++++++++++- 6 files changed, 2562 insertions(+), 417 deletions(-) create mode 100644 .changeset/eager-orcas-chatter.md diff --git a/.changeset/eager-orcas-chatter.md b/.changeset/eager-orcas-chatter.md new file mode 100644 index 000000000..d519a1a4b --- /dev/null +++ b/.changeset/eager-orcas-chatter.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 3739 +--- +**`deferred-items.md` entries written with `*`, `+` or an ordered marker are no longer silently dropped** — the parser recognised only `- `, so a deferred list written with any other standard Markdown list marker contributed zero entries and reported as a clean zero to `audit-open`. An ordered list counts when it starts at `0.` or `1.` or continues a list already open at its level; `1)` is not a marker. Also fixed: a `status:` line inside a fenced code block indented four or more spaces — what a fence under a nested bullet looks like — was not treated as fenced, so a `status: resolved` written as documentation resolved the entry containing it; and an unclosed fence now ends with its own entry instead of hiding every entry after it. `audit-open acknowledge` reads and writes the same grammar — its line selection goes through the reader's own classifier on the headless and (since #3781) the heading-delimited shape alike — so an entry the audit surfaces under any of these markers can be acknowledged, and a heading-delimited file written with one of the newly recognised markers now surfaces its entries for `complete-milestone` to acknowledge in place, where it previously closed over them unseen. (#3702) diff --git a/gsd-core/references/executor-examples.md b/gsd-core/references/executor-examples.md index 4b3953956..0e75b1ff2 100644 --- a/gsd-core/references/executor-examples.md +++ b/gsd-core/references/executor-examples.md @@ -69,6 +69,48 @@ - MAYBE → Rule 4 (ask the user) - NO → Out of scope (log to deferred-items.md) +### Writing `deferred-items.md` + +The file has no template — write it by hand, as a Markdown list under a +`## Deferred Items` heading. What counts as one entry: + +- One entry per top-level list item. `-`, `*` and `+` all count, and so does a + dot-terminated ordered marker (`1.`) when the list starts at `0.` or `1.`, or + continues a list already open at that level — a sentence that merely opens + with a number (`2026. was a bad year`) is prose, not an item, and so is a + list numbered from `2.` upward until its first `0.`/`1.` line. `1)` is not a + marker here, and neither is an ordinal past nine digits (`999999999.` + counts, `1234567890.` does not). +- Continuation lines indent beneath their entry. Fields go on those lines: + `status: resolved`, or the bolded `**Status:** resolved` convention. **The + BARE key is lower-case only** — write `Status: resolved` without the bold and + the field is not read, so the entry stays open with no warning. Bold it or + lower-case it. The bolded form matches the key case-insensitively, and the + VALUE is case-insensitive in both forms. +- A `* * *` or `- - -` separator closes the list rather than opening an entry, + and nothing inside a fenced code block is an entry or a field — at any indent, + including one deeper than CommonMark's three-space cap, which is what a fence + written under a nested bullet looks like. A fence that is never closed runs to + the end of its own entry and no further, so an unclosed delimiter cannot hide + the entries after it — a closed pair of delimiters is a fence, whatever sits + between them. + +An entry is RESOLVED only if it carries an explicit `status: resolved`. Anything +else — including an entry with no `status:` at all — stays open and will surface +in `audit-open`, `audit-uat` and `complete-milestone`. That is deliberate: the +scanner never silently drops a possibly-open item. What it reads as something +other than an item is the short list above — a fenced line, a separator, and an +ordered list numbered from `2.` upward at a paragraph position — and nothing +else. + +```markdown +## Deferred Items + +- Retry budget is hardcoded at 3 + status: open + **What:** `fetchWithRetry` ignores the configured budget. +``` + ## Checkpoint Examples ### Good checkpoint placement diff --git a/gsd-core/workflows/progress/steps/forensic-audit.md b/gsd-core/workflows/progress/steps/forensic-audit.md index b07bb17a3..69bd64b5f 100644 --- a/gsd-core/workflows/progress/steps/forensic-audit.md +++ b/gsd-core/workflows/progress/steps/forensic-audit.md @@ -89,7 +89,7 @@ Glob every phase directory's SCOPE BOUNDARY log (executor writes out-of-scope di ls .planning/phases/*/deferred-items.md 2>/dev/null || true ``` -For each `deferred-items.md` found, read its entries (bullet list, one entry per top-level `- ` line, continuation lines indented beneath it). An entry is RESOLVED only if it carries an explicit `status: resolved` field (case-insensitive) on one of its lines; every other entry — including one with no `status:` field at all — is UNRESOLVED and must be surfaced (fail-safe: never silently drop a possibly-open item). +For each `deferred-items.md` found, read its entries (bullet list, one entry per top-level list item — `-`, `*`, `+` or an ordered marker of one to nine digits terminated by a DOT, where an ordered list counts when it starts at `0.` or `1.` or continues a list already open at its level; `1)` is not a marker and neither is a ten-digit ordinal (`999999999.` counts, `1234567890.` does not) — with continuation lines indented beneath it; fenced code blocks at ANY indent, including deeper than CommonMark's three-space cap, and `* * *`-style separators are not entries; an unclosed fence ends with its own entry and never hides the entries after it; a closed pair of delimiters is a fence whatever sits between them). An entry is RESOLVED only if it carries an explicit `status: resolved` field on one of its lines — the bare key lower-case only, or the bolded `**Status:**` form in any case; the VALUE is case-insensitive in both; every other entry — including one with no `status:` field at all — is UNRESOLVED and must be surfaced (fail-safe: never silently drop a possibly-open item). Emit: - ✓ `No unresolved deferred items` — if no `deferred-items.md` files exist, or every entry in every file is `status: resolved` diff --git a/src/uat.cts b/src/uat.cts index 4acba0713..4532367a3 100644 --- a/src/uat.cts +++ b/src/uat.cts @@ -1877,7 +1877,7 @@ function parseGapsTableItems(sectionBody: string): UatItem[] { * surfaced. * * #3457: when the section body contains headings, entries are delimited by - * LEAF headings (see `splitDeferredHeadingEntries`) rather than by bullets — + * LEAF headings (see `splitDeferredHeadingEntriesDetailed`) rather than by bullets — * the executor convention writes one deferred item as a heading followed by * sibling `- **Field:** …` bullets, which the bullet-only split mis-counted as * one item PER BULLET. A body with no headings keeps the original @@ -1909,19 +1909,28 @@ function parseDeferredItemsWithStatus(content: string): Array<{ name: string; st // before field extraction, not just line 0 (which `extractGapEntryFields` // does for the headless/Gaps shape, where a later `- ` line is a nested // sub-list, not a field). - const headingEntries = splitDeferredHeadingEntries(sectionBody); + const headingEntries = splitDeferredHeadingEntriesDetailed(sectionBody); + // The opener flags are HANDED DOWN rather than pre-applied (#3702 round 3, + // m7/m8). Marker-stripping the lines here and passing the result meant the + // reader's fence scan ran over text the splitter never saw, and the namer + // stripped a marker off the heading TEXT. Both consumers now take the raw + // lines plus the splitter's own per-line verdict — a rejected ordinal + // ("3. status: resolved" as prose) still keeps its `3. ` and yields no field, + // because that verdict is what carries the rejection. const entries = headingEntries !== null - ? headingEntries.map((entryLines) => ({ - lines: entryLines, - fields: extractGapEntryFields(entryLines.map(stripLeadingBulletMarker)), + ? headingEntries.map((entry) => ({ + lines: entry.lines, + opener: entry.opener, + fields: extractGapEntryFields(entry.lines, DEFERRED_BULLET_MARKERS, entry.opener), })) - : splitGapsEntries(sectionBody).map((entryLines) => ({ + : splitGapsEntries(sectionBody, DEFERRED_BULLET_MARKERS).map((entryLines) => ({ lines: entryLines, - fields: extractGapEntryFields(entryLines), + opener: undefined, + fields: extractGapEntryFields(entryLines, DEFERRED_BULLET_MARKERS), })); - for (const { lines: entryLines, fields } of entries) { - const text = rawGapEntryText(entryLines); + for (const { lines: entryLines, opener, fields } of entries) { + const text = rawGapEntryText(entryLines, DEFERRED_BULLET_MARKERS, opener); if (!text) continue; items.push({ name: text, status: fields.status || '' }); @@ -1953,6 +1962,48 @@ function parseDeferredItems(content: string): UatItem[] { })); } +/** + * The line ending for an entry that ends the FILE, where the entry is a single + * line and therefore carries no terminator of its own to copy. No entry-local + * evidence exists here — the separator before the entry terminates the + * PREVIOUS line, not this one — so this asks the weaker question that CAN be + * answered: does anything before the entry, within the scope the caller passes, + * contradict CRLF? Uniform CRLF across that scope is the one case where + * appending a `\r\n` cannot make the file more irregular. It fails CLOSED: + * any bare `\n` in scope, or no scope at all, yields LF. + * + * Adopted from #3773 (`crlfAtEof`), whose four counterexamples fixed the scope + * and are ported alongside it. Every simpler choice is refuted by a named test: + * the separator immediately PRECEDING the entry propagates an isolated CRLF + * into an LF-dominant list, because it terminates the previous line rather than + * this one — that is the algorithm this PR shipped through round 3 and it is + * withdrawn here. The whole DOCUMENT rejects CRLF over an unrelated bare `\n` + * elsewhere, inside a fenced block say. The deferred-items SECTION body is + * right when a heading delimits one, and becomes the whole document when it + * does not. + * + * Scope, therefore: the section body when `## Deferred Items` delimits one (its + * own preamble belongs to that section), else the entry-list region, where only + * the entries can be trusted. + * + * WITH ONE CORRECTION to #3773, which is its B4. The entry-list region goes + * EMPTY exactly when the list is undelimited AND holds a single entry, since + * the region runs from the first entry's start to the insertion point and those + * coincide. `crlfAtEof('')` is `false`, so a bare `\n` was inserted into a CRLF + * document — `'preamble\r\n\r\n- alpha'` gained one — which is the very defect + * the fallback exists to close, and it breaks the fix's own uniform-CRLF + * invariant. When the preferred region is empty the caller widens to everything + * preceding the insertion point rather than asserting LF from no evidence. That + * can only ever loosen a scope that was carrying zero information, and the + * predicate stays fail-closed over the wider one, so a contradicting bare `\n` + * still yields LF. An entry at offset 0 of an undelimited document has no + * evidence under either scope and stays LF, rather than inventing an ending + * from nothing. + */ +function crlfAtEof(before: string): boolean { + return before.length > 0 && !/(^|[^\r])\n/.test(before); +} + // ─── acknowledgeDeferredItem ─────────────────────────────────────────────────── /** Result of `acknowledgeDeferredItem`. */ @@ -1975,25 +2026,27 @@ interface AcknowledgeDeferredItemResult { * `status:` away from `acknowledged` (or delete the field) and it resurfaces * with no separate cleanup step, exactly like every other category's marker. * - * #3781: the heading-delimited (#3457) entry shape is SUPPORTED, via - * `splitDeferredHeadingEntriesWithSpans` — a span-carrying sibling of the - * reader's walk that records each entry's (start, end) character span in the - * SAME pass that groups its lines (the identical technique - * `splitGapsEntriesWithSpans` uses for the headless shape). Leaf entries keep - * their RAW heading line as `lines[0]` so `sectionBody.slice(start, end)` is - * byte-verbatim; pending (preamble / container-direct) regions are contiguous - * slices handed to `splitGapsEntriesWithSpans` with a baseOffset translation. - * Two write rules differ from the headless path on this shape: the status - * search runs over the READER-form lines (what the reader actually parses — - * including the leaf line-0 corner where the heading text itself parses as a - * status field, whose raw line is rewritten with its ATX prefix preserved), - * and the insert branch inserts after the entry's LAST NON-BLANK line — a - * heading entry's body is frequently a soft-wrapped sentence, and splicing - * after line 0 would split it in half (#3781's sentence-split trap). - * Entries whose span embeds a GFM table row are non-contiguous (table lines - * are excluded from entries) and still refuse (`unsupported_heading_shape`) - * rather than risk a wrong-entry write; the fully-headless shape below is - * byte-for-byte the pre-#3781 path. + * #3781: the heading-delimited (#3457) entry shape is SUPPORTED. The + * reader's own walk, `splitDeferredHeadingEntriesDetailed`, records each + * entry's (start, end) character span in the SAME pass that groups its + * lines — the technique `splitGapsEntriesWithSpans` already uses for the + * headless shape — so there is no second walk for the writer to drift from + * (#3702 round 5: upstream's fix shipped a hyphen-only sibling walk, and this + * PR's widened grammar would have left it reading a different set of + * entries than the reader; folding the spans into the one walk is what + * keeps the writer and the reader on one grammar). The heading half, + * `acknowledgeHeadingShapedEntry`, shares this function's guards and its + * rewrite/insert machinery through the same `entryFieldLines` seam, with two + * shape-specific rules: the status search runs over the READER-form lines + * (the heading TEXT on a leaf's line 0 — including the corner where that + * text itself parses as a status field, rewritten with its ATX prefix + * preserved), and the insert branch inserts after the entry's LAST NON-BLANK + * line, because a heading entry's body is frequently a soft-wrapped + * sentence and splicing after line 0 would split it (#3781's sentence-split + * trap). Entries whose span embeds a GFM table row are non-contiguous (table + * lines are excluded from entries) and still refuse + * (`unsupported_heading_shape`) rather than risk a wrong-entry write; the + * fully-headless shape below is byte-for-byte the pre-#3781 path. * * Also refuses `ambiguous` (2+ entries share the exact same text — status must * be unique to identify one) and `not_found`, and is a no-op @@ -2040,16 +2093,16 @@ function acknowledgeDeferredItem(content: string, targetText: string): Acknowled ); const sectionBody = deferredSection ? deferredSection.body : content; - // #3781: the heading-delimited shape carries its own span walk; the - // headless path below is unchanged. - const headingEntries = splitDeferredHeadingEntriesWithSpans(sectionBody); + // #3781: the heading-delimited shape carries its own spans, recorded by + // the reader's walk; the headless path below is unchanged. + const headingEntries = splitDeferredHeadingEntriesDetailed(sectionBody); if (headingEntries !== null) { return acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, headingEntries, targetText }); } - const entries = splitGapsEntriesWithSpans(sectionBody); + const entries = splitGapsEntriesWithSpans(sectionBody, DEFERRED_BULLET_MARKERS); const matches = entries - .map((entry) => ({ entry, text: rawGapEntryText(entry.lines) })) + .map((entry) => ({ entry, text: rawGapEntryText(entry.lines, DEFERRED_BULLET_MARKERS) })) .filter((e) => e.text === targetText); if (matches.length === 0) return { content, status: 'not_found' }; @@ -2057,7 +2110,7 @@ function acknowledgeDeferredItem(content: string, targetText: string): Acknowled const { entry } = matches[0]; const { lines: entryLines, start, end } = entry; - const fields = extractGapEntryFields(entryLines); + const fields = extractGapEntryFields(entryLines, DEFERRED_BULLET_MARKERS); if (fields.status && fields.status.toLowerCase() === 'resolved') { return { content, status: 'already_resolved' }; } @@ -2074,100 +2127,139 @@ function acknowledgeDeferredItem(content: string, targetText: string): Acknowled // comparison that selected this entry — this catches real drift between // the two rather than a regex trivially guaranteed to agree with itself. const strippedForVerify = matchedLines.map((l) => l.replace(/\r$/, '')); - if (rawGapEntryText(strippedForVerify) !== targetText) { + if (rawGapEntryText(strippedForVerify, DEFERRED_BULLET_MARKERS) !== targetText) { return { content, status: 'match_verification_failed' }; } const matchIndexInContent = sectionOffset + start; - // #3740: the search must mirror the reader exactly. extractGapEntryFields - // strips a bullet marker on line 0 ALONE — a later `- ` line is a nested - // sub-list, never a field line — so a marker-prefixed match on any - // continuation line would rewrite a line no reader reads and report `ok` - // while the entry stays outstanding. Line 0 KEEPS the marker-optional - // form: the reader de-bullets it, so `- status: open` as the entry line is - // a real field there (and first-wins means the insert branch could not - // outrank it). Everything else falls through to the insert branch below, - // which the marker-free and no-status controls already round-trip. + // Locate the status line with the READER'S OWN classifier, never a + // writer-side regex (#3702 round 3, B1/B3 — the shape is #3773's, + // parameterised here by the widened marker set per the round-3 review's + // prescribed end state). The only line worth rewriting in place is one the + // reader will read back as `fields.status`. // - // #3775: the CASE axis of the same rule. The reader lowercases BOLDED - // keys only; bare keys keep their literal case (#3457 design), so a bare - // `Status:`/`STATUS:` line is stored under key `Status` and never read as - // fields.status. The search therefore matches a bolded key in ANY case - // and a bare key in LOWERCASE only — never a bare Title-case/UPPER line, - // which must fall through to the insert branch whose lowercase output the - // reader consumes (leaving any human `Status: resolved` untouched). - const statusFieldBoldedRe = /^\s*\*+status:\*+/i; - const statusFieldBareRe = /^\s*status:/; - const statusFieldBoldedReLine0 = /^\s*(?:-\s+)?\*+status:\*+/i; - const statusFieldBareReLine0 = /^\s*(?:-\s+)?status:/; - const statusLineIdx = matchedLines.findIndex((rawLine, idx) => { - const line = rawLine.replace(/\r$/, ''); - return idx === 0 - ? (statusFieldBoldedReLine0.test(line) || statusFieldBareReLine0.test(line)) - : (statusFieldBoldedRe.test(line) || statusFieldBareRe.test(line)); - }); + // What this closes: round 2 widened the writer's finder to the deferred + // marker set while `extractGapEntryFields` still read a marker only on line + // 0. A nested ` * status: pending` was therefore SELECTED by the writer and + // invisible to the reader — acknowledge rewrote it, returned `ok`, and the + // item stayed outstanding forever. Measured against a `next` build: `*`, `+` + // and `1.` all resolved on base and stopped resolving here, so it was a + // regression, not a gap in new behaviour. The hyphen form of the same shape + // (` - status:`) was already broken on `next`; it is fixed here too, since + // one classifier cannot be right for three markers and wrong for the fourth. + // + // A line the reader skips falls through to the INSERT branch, which writes a + // line the reader does read — the fail-safe direction. That covers a bare + // capitalised `Status:` (the reader stores it under `Status`, not `status`) + // and a `status:` line inside a fenced block, both of which the writer must + // NOT rewrite in place. Selecting either one produced an entry that could not + // be acknowledged at all; that is why the selection goes through + // `entryFieldLines` rather than the classifier directly. + const statusLineIdx = entryFieldLines(matchedLines, DEFERRED_BULLET_MARKERS) + .findIndex((field) => field?.key === 'status'); - // No CRLF-preservation branch here (WARNING 1, #3458 follow-up review): - // every write goes through `platformWriteSync` → `normalizeContent`, which - // for a `.md` path unconditionally runs `_normalizeMd` — whole-file - // `\r\n` → `\n`, plus blank-line normalization around headings/lists — on - // EVERY write, not just this one. That is this codebase's single, - // deliberate OS-facing I/O seam (`shell-command-projection.cts`), applied - // uniformly to every `.md` writer; carving out one exception here would - // fight it rather than follow it, for a guarantee (byte-identical CRLF on - // disk) the seam already makes impossible. A marker write on a CRLF - // `deferred-items.md` normalizes the WHOLE file to LF, same as any other - // `.md` write in this codebase — expected, not a regression to guard - // against. Where a source line still carries a trailing `\r` (read from an - // on-disk CRLF document before normalization), `String.prototype.replace` - // consumes it as part of `.*$` and the replacement text does not - // reproduce it, so it is dropped here too — consistent with the eventual - // whole-file normalization rather than duplicating it. + // Per-line CRLF preservation is the honest in-memory contract. The lines + // here are RAW — `audit acknowledge` hands this function the `readFileSync` + // content of an on-disk `deferred-items.md`, and `_normalizeMd` runs only on + // WRITE — so on a CRLF document every line but the span's last still carries + // its `\r`. The write path's whole-file normalization still decides what + // reaches disk; this function does not duplicate that decision, and it no + // longer silently drops the `\r` either. (Round 2's comment here argued the + // opposite contract. It is withdrawn: #3773 documents per-line preservation + // in this same function, and two opposite contracts in one function was a + // round-3 blocker in its own right.) let newMatchedLines: string[]; if (statusLineIdx === -1) { - const bulletIndentMatch = matchedLines[0].match(/^(\s*)-\s+/); - const continuationIndent = ' '.repeat((bulletIndentMatch ? bulletIndentMatch[1].length : 0) + 2); + const bulletIndentMatch = matchedLines[0].replace(/\r$/, '').match(DEFERRED_BULLET_MARKERS.strip); + // The entry's own indent CHARACTERS, never a count of them — a tab counted + // as one column and re-emitted as one space puts a 3-space continuation + // under a tab-indented bullet. Identical output for all-space indents. + const continuationIndent = `${bulletIndentMatch ? bulletIndentMatch[1] : ''} `; + // The new line goes right after line 0, so line 0 stops being the span's + // last line. Under CRLF the span's last line is the one WITHOUT a `\r` + // (the file's own `\r\n` follows the span), so the ending is read from the + // entry's OWN boundary and never sniffed from the whole document — a + // mixed-ending file must keep its LF opener. At end-of-file there is no + // following separator, so the boundary immediately PRECEDING the entry is + // the remaining local evidence; an entry at offset 0 has neither and stays + // LF rather than inventing an ending from nothing. + const line0HadCr = matchedLines[0].endsWith('\r'); + // A single-line entry's line 0 IS the span's last line, so its own + // terminator sits OUTSIDE the span and the separator FOLLOWING the span is + // the evidence. At end-of-file there is no such separator, and reading its + // absence as "not CRLF" is what joined a CRLF entry to its inserted line + // with a bare `\n`. `crlfAtEof` answers the weaker question that remains, + // over the entry's own SECTION when one is delimited and the entry list + // alone when none is — widening to everything before the insertion point + // only where that region is empty, which is #3773's B4. See its doc comment. + const spanEnd = matchIndexInContent + (end - start); + const eofScope = sectionBody.slice(deferredSection ? 0 : entries[0].start, start) + || sectionBody.slice(0, start); + const crlf = matchedLines.length > 1 + ? line0HadCr + : content.startsWith('\r\n', spanEnd) + || (spanEnd >= content.length && crlfAtEof(eofScope)); newMatchedLines = [ - matchedLines[0], - `${continuationIndent}status: acknowledged`, + crlf ? `${matchedLines[0].replace(/\r$/, '')}\r` : matchedLines[0], + `${continuationIndent}status: acknowledged${line0HadCr ? '\r' : ''}`, ...matchedLines.slice(1), ]; } else { - const original = matchedLines[statusLineIdx]; - const replaced = original.replace( - /^(\s*(?:-\s+)?)(\*+status:\*+|status:)(\s*).*$/i, - (_m, indent: string, key: string, ws: string) => `${indent}${key}${ws}acknowledged`, - ); + // Rewrite at the offset the CLASSIFIER reported, rather than through a + // second regex of the writer's own. This is what makes the selection and + // the rewrite structurally incapable of disagreeing: a status line the + // classifier can select is one whose value offset it has already + // computed, so there is no shape it selects and then fails to rewrite. + // (A widened classifier over a hyphen-only rewrite regex is exactly that + // failure — it would select `* status: open` and hand back the line + // untouched.) The key's own spelling and any `**bold**` wrapper survive + // because only the value is replaced. + const raw = matchedLines[statusLineIdx]; + const cr = raw.endsWith('\r') ? '\r' : ''; + const line = raw.slice(0, raw.length - cr.length); + const field = parseGapEntryFieldLine(line, DEFERRED_BULLET_MARKERS, statusLineIdx === 0)!; + const prefix = line.slice(0, field.valueStart); + // `status:acknowledged` reads back fine, but a bare colon with no + // separator is not what this file's convention looks like; supply one only + // when the source had none. + const sep = /[ \t]$/.test(prefix) ? '' : ' '; newMatchedLines = matchedLines.slice(); - newMatchedLines[statusLineIdx] = replaced; + newMatchedLines[statusLineIdx] = `${prefix}${sep}acknowledged${cr}`; } + // NO post-write read-back guard here, deliberately (round 4, B3). Round 3 + // added one — `rewrite_not_readable` — after a fenced `status:` line proved + // the writer could select a line the reader would not read back. Round 3 + // then closed that divergence STRUCTURALLY, by routing the writer's line + // selection and the reader's field extraction through the one + // `entryFieldLines` seam above, and the guard became unreachable from the + // public API: 21 document shapes were driven against it (fence openers on + // the bullet line for every marker, duplicate and triplicate `status:` + // lines, bolded and nested variants, fences between duplicates) and none + // reached it. + // + // An unreachable branch is not free here. This repo's own + // `RULESET.TESTS.mutation-score` runs Stryker incrementally over changed + // files at an 80% threshold and says to "treat surviving mutant as a failing + // test specification"; an undriven `if` is exactly that. The only seam that + // would drive it is routing this call through the module's exports so a test + // could stub it — production surface reshaped for a test, which is a worse + // trade than the guard is worth now that construction, not assertion, + // enforces the invariant. + // + // What that gives up, stated plainly rather than hidden: if a future change + // re-splits the writer's selection from the reader's extraction, this + // function returns `ok` over an item that stays outstanding — the original + // #3702 defect class. `match_verification_failed` does NOT backfill it; that + // check runs BEFORE the write and compares the matched span to the target, + // so it cannot see a post-write read-back failure. The protection against + // re-splitting is the shared seam plus the round-3 tests that pin it, not a + // runtime assertion. + const newContent = content.slice(0, matchIndexInContent) + newMatchedLines.join('\n') + content.slice(matchIndexInContent + (end - start)); return { content: newContent, status: 'ok' }; } -/** - * #3781 — one entry of the heading-delimited deferred shape, carrying the - * exact character span it occupies within the `sectionBody` it was derived - * from. `lines[0]` of a LEAF entry is the RAW heading line (hashes intact) so - * `sectionBody.slice(start, end)` is byte-verbatim; `readerLines` is the - * reader-form of the same lines (ATX-stripped line 0, bullet-stripped body) - * the identity text and field extraction are computed over. `embeddedTable` - * marks entries whose span contains a GFM table line — table lines are - * excluded from entries, so such a span is non-contiguous and its entries - * refuse rather than risk a wrong write. - */ -interface DeferredHeadingEntrySpan { - kind: 'leaf' | 'pending'; - lines: string[]; - readerLines: string[]; - text: string; - fields: Record; - start: number; - end: number; - embeddedTable: boolean; -} - /** * #3781 — strip an ATX heading prefix, mirroring `tokenizeHeadings`' own ATX * regex (≤3 leading spaces, 1–6 `#`, space/tab separator, optional closing @@ -2183,133 +2275,6 @@ function stripAtxPrefix(line: string): string | null { : m[3].replace(/^[ \t]+/, '').replace(/[ \t]+#+[ \t]*$/, '').replace(/^#+[ \t]*$/, '').trim(); } -/** - * #3781 — span-carrying sibling of `splitDeferredHeadingEntries`: ONE walk, - * identical grouping rules (leaf = childless heading whose body carries a - * bullet; container = next heading deeper; preamble/container-direct lines → - * headless entries; table lines excluded), additionally recording each - * entry's (start, end) character span within `sectionBody`. Returns null when - * the body contains no heading at all — the caller then takes the unchanged - * fully-headless path. - */ -function splitDeferredHeadingEntriesWithSpans(sectionBody: string): DeferredHeadingEntrySpan[] | null { - const headings = tokenizeHeadings(sectionBody); - if (headings.length === 0) return null; - - const lines = sectionBody.split('\n'); - const lineStarts: number[] = []; - const lineEnds: number[] = []; - let cursor = 0; - for (const rawLine of lines) { - lineStarts.push(cursor); - cursor += rawLine.length; - lineEnds.push(cursor); - cursor += 1; - } - - const headingByLine = new Map(); - for (let i = 0; i < headings.length; i++) { - const isContainer = i + 1 < headings.length && headings[i + 1].level > headings[i].level; - headingByLine.set(headings[i].line, { text: headings[i].text, isContainer }); - } - - const isTableLine = (l: string): boolean => /^\s*\|/.test(l.replace(/\r$/, '')); - const isBulletLine = (l: string): boolean => /^\s*-\s/.test(l.replace(/\r$/, '')); - - const entries: DeferredHeadingEntrySpan[] = []; - let current: string[] | null = null; - let currentReaderLine0: string | null = null; - let currentStartLine = -1; - let currentEndLine = -1; - let currentHasBullet = false; - let currentTable = false; - let pendingStartLine = -1; - let pendingEndLine = -1; - - const flushCurrent = (): void => { - if (current !== null && currentHasBullet && currentReaderLine0 !== null) { - const bodyReader = current.slice(1).map(stripLeadingBulletMarker); - entries.push({ - kind: 'leaf', - lines: current, - readerLines: [currentReaderLine0, ...bodyReader], - text: rawGapEntryText([currentReaderLine0, ...current.slice(1)]), - fields: extractGapEntryFields([currentReaderLine0, ...bodyReader]), - start: lineStarts[currentStartLine], - end: lineEnds[currentEndLine], - embeddedTable: currentTable, - }); - } - current = null; - currentReaderLine0 = null; - currentStartLine = -1; - currentEndLine = -1; - currentHasBullet = false; - currentTable = false; - }; - const flushPending = (): void => { - if (pendingStartLine === -1) return; - // The pending region is contiguous (a heading flushes it), but table lines - // inside it were skipped by the walk: the reader's identity for this - // region is computed over the table-FILTERED join, which may merge - // entries across the gap, so spans cannot be translated faithfully — - // mark the region's entries as refusing instead. - let regionTable = false; - for (let i = pendingStartLine; i <= pendingEndLine; i++) { - if (isTableLine(lines[i])) regionTable = true; - } - const base = lineStarts[pendingStartLine]; - const regionText = sectionBody.slice(lineStarts[pendingStartLine], lineEnds[pendingEndLine]); - for (const e of splitGapsEntriesWithSpans(regionText)) { - entries.push({ - kind: 'pending', - lines: e.lines, - readerLines: e.lines, - text: rawGapEntryText(e.lines), - fields: extractGapEntryFields(e.lines), - start: base + e.start, - end: base + e.end, - embeddedTable: regionTable, - }); - } - pendingStartLine = -1; - pendingEndLine = -1; - }; - - for (let i = 0; i < lines.length; i++) { - const heading = headingByLine.get(i + 1); - if (heading !== undefined) { - flushCurrent(); - flushPending(); - if (!heading.isContainer) { - current = [lines[i]]; - currentReaderLine0 = heading.text; - currentStartLine = i; - currentEndLine = i; - currentHasBullet = false; - currentTable = false; - } - continue; - } - if (isTableLine(lines[i])) { - if (current !== null) currentTable = true; - continue; - } - if (current !== null) { - current.push(lines[i]); - currentEndLine = i; - if (isBulletLine(lines[i])) currentHasBullet = true; - } else { - if (pendingStartLine === -1) pendingStartLine = i; - pendingEndLine = i; - } - } - flushCurrent(); - flushPending(); - - return entries; -} - /** * #3781 — the heading-shaped half of `acknowledgeDeferredItem`, sharing the * headless path's guards (not_found / ambiguous / already_resolved / @@ -2317,58 +2282,60 @@ function splitDeferredHeadingEntriesWithSpans(sectionBody: string): DeferredHead * shape-specific rules documented on `acknowledgeDeferredItem` (reader-form * status search incl. the leaf line-0 ATX corner; insert after the entry's * last non-blank line). Extracted so the headless path stays byte-identical. + * + * Every question about an entry is asked of the READER'S OWN answer (#3702 + * round 3, B1/B3 — restated here rather than re-implemented): identity is + * `rawGapEntryText` over the walk's lines and opener flags, exactly as + * `parseDeferredItemsWithStatus` names the entry; the status line is whichever + * line `entryFieldLines` classifies as `status`, so a fenced `status:` or a + * rejected-ordinal prose line is never selected; and the rewrite lands at the + * offset the classifier reported. Upstream's #3781 carried its own + * hyphen-only walk and its own status regexes for this shape; under the + * widened marker grammar those would have read a different entry set than + * the reader and re-opened the writer/reader drift this PR closes. */ function acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, headingEntries, targetText }: { content: string; sectionBody: string; deferredSection: { body: string; bodyStart: number } | null; - headingEntries: DeferredHeadingEntrySpan[]; + headingEntries: DeferredHeadingEntry[]; targetText: string; }): AcknowledgeDeferredItemResult { - const matches = headingEntries.filter((e) => e.text === targetText); + const matches = headingEntries.filter( + (e) => rawGapEntryText(e.lines, DEFERRED_BULLET_MARKERS, e.opener) === targetText, + ); if (matches.length === 0) return { content, status: 'not_found' }; if (matches.length > 1) return { content, status: 'ambiguous' }; const entry = matches[0]; + // A table row inside the span: the walk skipped it, so `lines` is not 1:1 + // with the raw slice and no write can be anchored. Refuse, as before #3781. if (entry.embeddedTable) return { content, status: 'unsupported_heading_shape' }; - if (entry.fields.status && entry.fields.status.toLowerCase() === 'resolved') { + const fields = extractGapEntryFields(entry.lines, DEFERRED_BULLET_MARKERS, entry.opener); + if (fields.status && fields.status.toLowerCase() === 'resolved') { return { content, status: 'already_resolved' }; } const sectionOffset = deferredSection ? deferredSection.bodyStart : 0; - const rawSlice = sectionBody.slice(entry.start, entry.end); - const rawSliceLines = rawSlice.split('\n'); - - // Genuine invariant re-verification: re-derive the entry's identity from - // the span's own bytes and compare against the targetText that selected it - // — the span was recorded by an offset bookkeeping independent of the - // identity comparison above. - const verifyText = entry.kind === 'leaf' - ? (() => { - const stripped = stripAtxPrefix(rawSliceLines[0]); - return stripped === null ? null : rawGapEntryText([stripped, ...rawSliceLines.slice(1)]); - })() - : rawGapEntryText(rawSliceLines.map((l) => l.replace(/\r$/, ''))); - if (verifyText !== targetText) { + const rawLines = sectionBody.slice(entry.start, entry.end).split('\n'); + // The reader-form of the raw slice, index-aligned with it: a leaf's line 0 + // is the heading TEXT (re-derived from the span's own bytes, not copied + // from the walk, so the verification below is genuine), every other line + // CR-stripped. Markers stay on the lines — the classifier strips them per + // the opener flags, exactly as the reader does. + const readerLines = rawLines.map((raw, i) => { + const line = raw.replace(/\r$/, ''); + return i === 0 && entry.kind === 'leaf' ? (stripAtxPrefix(line) ?? line) : line; + }); + // Genuine invariant re-verification: the identity re-derived from the + // span's bytes must be the identity that selected the entry — the span was + // recorded by offset bookkeeping independent of that comparison. + if (readerLines.length !== entry.lines.length + || rawGapEntryText(readerLines, DEFERRED_BULLET_MARKERS, entry.opener) !== targetText) { return { content, status: 'match_verification_failed' }; } - // Status search over the READER-form lines — the exact set the reader - // parses (bolded any case + bare lowercase, per #3775; line-0 forms per - // #3740). Reader lines are index-aligned 1:1 with the raw lines. - const statusFieldBoldedRe = /^\s*\*+status:\*+/i; - const statusFieldBareRe = /^\s*status:/; - const statusFieldBoldedReLine0 = /^\s*(?:-\s+)?\*+status:\*+/i; - const statusFieldBareReLine0 = /^\s*(?:-\s+)?status:/; - const readerLines = entry.kind === 'leaf' - ? [ - stripAtxPrefix(rawSliceLines[0]) ?? rawSliceLines[0].replace(/\r$/, ''), - ...rawSliceLines.slice(1).map((l) => stripLeadingBulletMarker(l.replace(/\r$/, ''))), - ] - : rawSliceLines.map((l) => l.replace(/\r$/, '')); - const statusLineIdx = readerLines.findIndex((line, idx) => - idx === 0 - ? (statusFieldBoldedReLine0.test(line) || statusFieldBareReLine0.test(line)) - : (statusFieldBoldedRe.test(line) || statusFieldBareRe.test(line))); + const statusLineIdx = entryFieldLines(readerLines, DEFERRED_BULLET_MARKERS, entry.opener) + .findIndex((field) => field?.key === 'status'); let newRawLines: string[]; if (statusLineIdx === -1) { @@ -2376,43 +2343,53 @@ function acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, // entry's body is frequently a soft-wrapped sentence, and splicing after // line 0 would split it in half (#3781's sentence trap). The headless // (no-heading-anywhere) path keeps its own splice-after-line-0 shape. - let lastNonBlank = rawSliceLines.length - 1; - while (lastNonBlank > 0 && rawSliceLines[lastNonBlank].replace(/\r$/, '').trim() === '') { - lastNonBlank--; - } + // …and never INSIDE a fence (round 5, RV6.5 review): an entry whose body + // ends in a fenced block — closed or, worse, unclosed and so running to + // the entry's end — would otherwise receive its marker as fence content, + // a line the reader never reads: `ok` returned, item still outstanding. + // Walk back over blank and fenced lines alike, classified exactly as the + // reader classifies them, so the marker lands on a line the reader reads. + const fencedInEntry = entryFencedLines(readerLines, DEFERRED_BULLET_MARKERS, entry.opener); + let last = rawLines.length - 1; + while (last > 0 && (readerLines[last].trim() === '' || fencedInEntry.has(last))) last--; + // A pending entry's continuation sits two columns inside its own marker + // indent (the entry's indent CHARACTERS, as the headless path does); a + // leaf's body lines are sibling bullets, and an indented bare field line + // among them is what the reader reads on that shape. const indent = entry.kind === 'pending' - ? (() => { - const bulletIndentMatch = rawSliceLines[0].match(/^(\s*)-\s+/); - return ' '.repeat((bulletIndentMatch ? bulletIndentMatch[1].length : 0) + 2); - })() + ? `${rawLines[0].replace(/\r$/, '').match(DEFERRED_BULLET_MARKERS.strip)?.[1] ?? ''} ` : ' '; - newRawLines = [ - ...rawSliceLines.slice(0, lastNonBlank + 1), - `${indent}status: acknowledged`, - ...rawSliceLines.slice(lastNonBlank + 1), - ]; + // The inserted line copies the ending of the line it follows. When that + // line is the span's LAST, its terminator sits outside the span: the + // separator following the span decides, else (end of file) the section's + // own evidence — the same rule the headless path applies to line 0. + const followsLast = last === rawLines.length - 1; + const spanEnd = sectionOffset + entry.end; + const prevCr = followsLast + ? content.startsWith('\r\n', spanEnd) || (spanEnd >= content.length && crlfAtEof(sectionBody.slice(0, entry.start))) + : rawLines[last].endsWith('\r'); + newRawLines = rawLines.slice(); + if (followsLast && prevCr) newRawLines[last] = `${rawLines[last].replace(/\r$/, '')}\r`; + newRawLines.splice(last + 1, 0, `${indent}status: acknowledged${!followsLast && prevCr ? '\r' : ''}`); } else { - newRawLines = rawSliceLines.slice(); - if (entry.kind === 'leaf' && statusLineIdx === 0) { - // Leaf line 0 is the RAW heading line — rewrite the heading-text portion - // the reader treats as a field, with the ATX prefix preserved. - const replacedReader = readerLines[statusLineIdx].replace( - /^(\s*(?:-\s+)?)(\*+status:\*+|status:)(\s*).*$/i, - (_m, indent: string, key: string, ws: string) => `${indent}${key}${ws}acknowledged`, - ); - const atxMatch = /^(\s*#+[ \t]*)(.*)$/.exec(rawSliceLines[0].replace(/\r$/, '')); - newRawLines[0] = atxMatch ? atxMatch[1] + replacedReader : replacedReader; - } else { - // Every other status line is rewritten on its RAW line, so the bullet - // marker and indent survive the write (the reader-form line has the - // marker stripped — writing it back would mangle the markdown shape and - // change the entry's identity text). Same replacement regex as the - // headless path. - newRawLines[statusLineIdx] = rawSliceLines[statusLineIdx].replace( - /^(\s*(?:-\s+)?)(\*+status:\*+|status:)(\s*).*$/i, - (_m, indent: string, key: string, ws: string) => `${indent}${key}${ws}acknowledged`, - ); - } + // Rewrite at the offset the CLASSIFIER reported, on the RAW line — the + // marker, the indent, the key's spelling and any `**bold**` wrapper all + // survive because only the value is replaced. A leaf's line 0 is the + // heading line, so its ATX prefix is put back in front of the rewritten + // text (the reader reads the heading text itself as the field there). + const raw = rawLines[statusLineIdx]; + const cr = raw.endsWith('\r') ? '\r' : ''; + const line = raw.slice(0, raw.length - cr.length); + const reader = readerLines[statusLineIdx]; + const field = parseGapEntryFieldLine(reader, DEFERRED_BULLET_MARKERS, stripsMarkerAt(statusLineIdx, entry.opener))!; + const prefix = reader.slice(0, field.valueStart); + const sep = /[ \t]$/.test(prefix) ? '' : ' '; + const leafLine0 = statusLineIdx === 0 && entry.kind === 'leaf'; + const atx = leafLine0 ? (/^( {0,3}#{1,6}[ \t]+)/.exec(line)?.[1] ?? '') : ''; + // A closing `#` sequence is Markdown the reader ignores; keep it (RV6.5). + const closing = leafLine0 ? (/[ \t]+#+[ \t]*$/.exec(line)?.[0] ?? '') : ''; + newRawLines = rawLines.slice(); + newRawLines[statusLineIdx] = `${atx}${prefix}${sep}acknowledged${closing}${cr}`; } const matchIndexInContent = sectionOffset + entry.start; @@ -2421,14 +2398,334 @@ function acknowledgeHeadingShapedEntry({ content, sectionBody, deferredSection, } /** - * Strip one leading `- ` bullet marker (#3457). Heading-delimited deferred - * entries carry their fields as sibling bullets; `extractGapEntryFields` only - * de-bullets line 0 (Gaps-protective — there, a later `- ` line is a nested - * sub-list), so the deferred heading path de-bullets every line itself before - * field extraction. Non-bullet lines pass through untouched. + * The two regex shapes a bullet-aware walk needs over one marker set: `open` + * detects an entry-opening marker, `strip` additionally captures the line's + * remaining content. Carrying both in ONE object is what keeps a detection + * site and its matching strip site from drifting to different marker sets — + * the asymmetry #3702 warns about (widen what OPENS an entry without widening + * what is STRIPPED before field extraction, and a status field written under + * the new marker never resolves its entry). */ -function stripLeadingBulletMarker(line: string): string { - return line.replace(/^(\s*)-\s+/, ''); +interface BulletMarkers { + /** `^(indent)()\s` — group 1 is the indent, group 2 the marker token. */ + open: RegExp; + /** `^(indent)\s+(rest)$` — group 1 indent, group 2 rest-of-line. */ + strip: RegExp; + /** + * Honour Markdown block structure — thematic breaks close the list, fenced + * code neither opens an entry nor carries a field (#3702 round 2, M1/M2), + * and indents are measured in CommonMark columns rather than raw characters. + * `false` for the Gaps grammar, which #3702's ruling keeps byte-for-byte on + * its `next` behaviour; `true` for the deferred-items grammar. Carried on the + * marker set rather than as a second parameter so a caller cannot widen one + * without the other. + */ + blockStructure: boolean; +} + +/** + * Hyphen-only markers — the `## Gaps` form, unchanged by #3702. Gaps entries + * come from a template that mandates the hyphen YAML-lite shape, so widening + * that section's grammar is not what the deferred-items ruling required; the + * shared splitting seam is parameterised rather than widened wholesale so the + * Gaps path stays byte-for-byte on its existing behaviour. + */ +const HYPHEN_BULLET_MARKERS: BulletMarkers = { + open: /^(\s*)(-)\s/, + strip: /^(\s*)-\s+(.*)$/, + blockStructure: false, +}; + +/** + * Deferred-items markers (#3702): the standard Markdown list markers, not the + * hyphen alone. `deferred-items.md` has NO template and no mandated shape — + * executors write it by hand (the same premise that justified the #2766 table + * union) — so an author reaching for `*`, `+` or `1.` wrote a list by every + * Markdown definition while this parser contributed ZERO entries for it. The + * hyphen restriction was a regex literal inherited from the Gaps seam, never a + * stated decision: measured in the wild, non-empty records parsed to a clean + * zero, and a MIXED file dropped its non-hyphen entries while keeping their + * hyphenated siblings — under-reporting without ever looking empty. + * + * Deliberately NOT widened to prose: "prose is not an item" is this parser's + * pre-existing, test-asserted contract (the `# Notes` case) and is untouched + * here. An asterisk bullet is not prose, and a `|` row is not a list marker — + * table lines are still skipped before the body-bullet flag can be set, so + * `parseDeferredTableItems` keeps sole ownership of table bodies and the + * #2766 anti-double-count property holds unchanged. + * + * The paren-terminated ordered form (`1)`) is out of scope for this fix: the + * #3702 ruling scopes the widening to `*`, `+` and the dot-terminated ordered + * marker. + * + * `DEFERRED_MARKER_ALT` is THE source every deferred-items marker regex is + * built from (#3702 round 2, M3). Since round 3 that is the splitter's + * `open`/`strip` pair here and nothing else: `acknowledgeDeferredItem`'s two + * status-line shapes used to be derived from it too, and are now deleted in + * favour of the reader's classifier. CommonMark + * §5.2: bullet markers `-`, `*`, `+`; an ordered marker is 1-9 digits and a + * `.` (round-1's `\d+` was uncapped). The marker is followed by a space or a + * tab — `[ \t]`, where round 1 wrote `\s`, which also accepted `\r`. + * `markdown-sectionizer`'s `iterateBullets` is the repo's other list-marker + * grammar; the `#3702 round 2: marker-grammar parity` test pins this one to + * it on the shared vocabulary and names the two points they deliberately + * differ (tab after the marker, the 9-digit cap). + */ +const DEFERRED_MARKER_ALT = '(?:[-*+]|\\d{1,9}\\.)'; +const DEFERRED_BULLET_MARKERS: BulletMarkers = { + open: new RegExp(`^(\\s*)(${DEFERRED_MARKER_ALT})[ \\t]`), + strip: new RegExp(`^(\\s*)${DEFERRED_MARKER_ALT}[ \\t]+(.*?)\\r?$`), + blockStructure: true, +}; + +// `acknowledgeDeferredItem` carries NO status-line regex of its own (#3702 +// round 3, B1/B3). It used to hold two — a finder and a rewrite — derived +// from `DEFERRED_MARKER_ALT` so the two WRITER shapes could not drift from +// each other. That kept the wrong pair in step: the finder's peer is the +// READER, and widening detection without widening the read is what made a +// nested ` * status:` line selectable by the writer and invisible to +// `extractGapEntryFields`. Both are gone; the writer now locates its line +// through `parseGapEntryFieldLine`, the reader's own classifier, and rewrites +// at the offset that classifier reports. See `parseGapEntryFieldLine`. + +/** + * CommonMark §4.1 thematic break: up to 3 spaces of indent, then three or + * more of the SAME `-`, `*` or `_`, optionally space/tab-separated, and + * nothing else. `- - -`, `* * *` and `+ + +` all also match a list opener — + * `- - -` was a phantom `"- -"` entry on base, and #3702's widening added the + * other two (#3702 round 2, M1). `+ + +` is not a CommonMark break, but it is + * the same authoring gesture and no less garbage as an entry name, so the + * class here is "three-or-more of one marker character, nothing else". The + * indent is unbounded, not CommonMark's `{0,3}`: this parser reads a list at + * any indent (see `#3702 round 2` m1), so a separator drawn at any indent is + * a separator too — otherwise ` * * *` is a phantom entry named `* *`. + */ +const THEMATIC_BREAK_RE = /^[ \t]*([-*+_])(?:[ \t]*\1){2,}[ \t]*$/; + +/** + * The deferred grammar's line view for fence classification: the SAME lines, + * with leading whitespace removed (#3702 round 4, M2). + * + * `scanFencedBlocks` is CommonMark, and CommonMark caps a fence delimiter's + * indent at three spaces — a fourth makes it an indented code block instead. + * The deferred grammar deliberately opted out of that cliff everywhere else: + * an entry opener is `[ \t]*`-indented and `THEMATIC_BREAK_RE` is + * `^[ \t]*`. Leaving the fence rule at CommonMark's cap while items and + * breaks are unbounded is not a conservative choice, it is an inconsistent + * one, and it is REACHED BY ORDINARY DOCUMENTS: a fenced block written under a + * nested bullet sits at four spaces, so its `status: resolved` line resolved + * the entry containing it. That is the #3702 silent-resolution defect class in + * a new place — driven, at indents 4, 5, 8 and a leading tab, before this fix. + * + * Still NO second fence dialect (the rule `blankIndentedFenceDelimiters` + * states): the classification is done by `scanFencedBlocks`, the one exported + * CommonMark state machine, over a de-indented view. Run lengths, backtick + * vs tilde, closer-must-match-and-not-trail and info-string rules are all + * still that engine's answers, not re-derived here; the unterminated case is + * its answer too, bounded by the walk (round 5, B1 — see `scanFencesFrom`). Indent is the only dimension this hides from it, and it is + * the exact dimension the deferred grammar has already declared it does not + * measure. Index alignment is 1:1 by construction — `map` preserves length — + * so every line index the engine returns still addresses the original line. + * + * Scope: the deferred grammar only. Both marker-parameterised call sites gate + * on `markers.blockStructure`, which the `## Gaps` set does not set, so Gaps + * reaches an empty set and is untouched by this — the same opt-out + * `indentWidth` documents for the indent half. + */ +function deindentedForFences(lines: string[]): string[] { + return lines.map((line) => line.replace(/^[ \t]+/, '')); +} + +/** + * Indices (into `lines`) of every line that sits inside a fenced code block, + * delimiters included — by the sectionizer's own fence state machine, so a + * `~~~` fence, an indented fence and an unterminated fence (runs to the end) + * are classified exactly as `stripFencedCode` would (#3702 round 2, M2), at + * ANY indent (round 4, M2 — see `deindentedForFences`). + * #3702's wild records carry reproduction blocks; `+`-prefixed diff lines and + * `1.`-numbered steps are their normal content, not entries. + * + * ENTRY-scoped: `lines` are ONE entry's lines (`entryFieldLines`), so an + * unterminated fence "running to the end" runs to the end of that entry — + * exactly the bound the section-level walks give it (round 5, B1; see + * `scanFencesFrom`). The two classifications agree by construction. + */ +function fencedLineSet(lines: string[]): Set { + const fenced = new Set(); + for (const block of scanFencedBlocks(deindentedForFences(lines))) { + const last = block.closeLineIdx === -1 ? lines.length - 1 : block.closeLineIdx; + for (let i = block.openLineIdx; i <= last; i++) fenced.add(i); + } + return fenced; +} + +/** + * Fence classification for a SECTION-level walk (round 5, B1). Terminated + * fences are classified wholesale, as `fencedLineSet` does. The UNTERMINATED + * fence is the difference: `scanFencedBlocks` runs it to the end of the + * document, and so did round 4 — at any indent, since round 4's M2 — so one + * stray delimiter swallowed every entry after it into the entry before it. + * `- a` / blank / four-space ``` / blank / `- b` yielded ONE entry where + * `next` yields two: a widening that made an already-counted item vanish, + * in the direction of silent data loss, on the exact mixed-file shape #3702 + * exists to close. + * + * The rule now: an unterminated fence runs to the end of its ENTRY, never + * past it. That is CommonMark's own answer for a fence inside a list item — + * the fence closes when its container does, and the container closes at the + * next item at its level — extended to a document-level stray delimiter, + * where CommonMark would swallow to end-of-file and this parser's fail-safe + * rule (surface a questionable entry rather than drop a real one) will not. + * The scan therefore reports the unterminated opener's index and leaves the + * bound to the walk, which knows what a top-level item is: on reaching one, + * the walk RESCANS from that line, so a delimiter after the bound is read on + * its own terms rather than as the content of a fence that has ended. Still + * no second fence dialect — every block boundary is `scanFencedBlocks`' + * answer; the walk contributes only the bound. + */ +interface FenceScan { + /** Lines inside a TERMINATED fence, delimiters included. */ + fenced: Set; + /** Every fence opener, terminated or not — an opener ends the list runs at its level, as a paragraph does. */ + openers: Set; + /** Opener index of the one unterminated fence in scope, else -1. Lines from it onward are fenced until the walk bounds it. */ + unterminatedFrom: number; +} + +function scanFencesFrom(lines: string[], from: number): FenceScan { + const scan: FenceScan = { fenced: new Set(), openers: new Set(), unterminatedFrom: -1 }; + for (const block of scanFencedBlocks(deindentedForFences(lines.slice(from)))) { + const open = block.openLineIdx + from; + scan.openers.add(open); + if (block.closeLineIdx === -1) { + scan.unterminatedFrom = open; // always the scan's last block + break; + } + for (let i = open; i <= block.closeLineIdx + from; i++) scan.fenced.add(i); + } + return scan; +} + +/** The scan a grammar without block structure (`## Gaps`) walks under: nothing is fenced. Never mutated. */ +const NO_FENCES: FenceScan = { fenced: new Set(), openers: new Set(), unterminatedFrom: -1 }; + +/** + * Does `line` LOOK like a top-level list item under `markers` — a marker at + * or above the base indent, start value ignored? The bound an unterminated + * fence runs to (round 5, B1; see `scanFencesFrom`). Shape rather than the + * ordered-start rule, because the list memory inside a fence is not evidence + * of anything, and closing a stray fence one line early errs in the + * surfacing direction. + */ +function topLevelItemShape(line: string, markers: BulletMarkers, baseIndent: number | null): boolean { + const m = line.match(markers.open); + return m !== null && (baseIndent === null || indentWidth(m[1], markers) <= baseIndent); +} + +/** + * Per-indent LIST memory (#3702 round 2, round review; widened round 5, M2): + * each list level remembers whether a list is OPEN there, so an ordered + * marker that does not start at `0.`/`1.` is an item when it continues or + * follows a list at its level, and prose otherwise. A new opener at indent + * `d` resets every deeper level (a new item starts new sub-lists); a + * paragraph after a blank at indent `d` ends the lists at `d` and deeper; a + * thematic break or a heading clears everything. + * + * Round 2 keyed this on whether the previous opener was ORDERED, so a bullet + * item closed the run and `1. a` / `- b` / `5. c` folded `5. c` into `b` — + * the mixed-file under-report #3702 names as the shape that bites. In + * CommonMark `5. c` there opens a fresh ordered list (`start=5`): a non-1 + * ordinal is refused only where it would INTERRUPT A PARAGRAPH (§5.3), and + * after a list item it interrupts nothing. Keying on "a list is open here" + * is that rule as far as this parser can state it without a paragraph model. + */ +class ListRuns { + private readonly byIndent = new Set(); + at(indent: number): boolean { return this.byIndent.has(indent); } + opened(indent: number): void { + for (const d of [...this.byIndent]) if (d > indent) this.byIndent.delete(d); + this.byIndent.add(indent); + } + endedAt(indent: number): void { + for (const d of [...this.byIndent]) if (d >= indent) this.byIndent.delete(d); + } + clear(): void { this.byIndent.clear(); } +} + +/** + * Leading-whitespace width of a line in CommonMark COLUMNS (§2.2: a tab + * advances to the next multiple of 4), so `\t` and ` ` are different levels + * and `\t` equals four spaces — character counting aliased them. + */ +/** + * Indent WIDTH under a grammar (#3702 round 2, review round 6). The deferred + * grammar measures CommonMark columns; the Gaps grammar keeps `next`'s raw + * character count, because its `blockStructure: false` opt-out promises + * byte-for-byte parity and a column measure silently breaks it — a + * tab-indented Gaps item followed by a two-space one split into two entries + * where `next` folded them into one, and the reverse pair folded where `next` + * split. The opt-out now covers indent semantics, not only fences and breaks. + */ +function indentWidth(indent: string, markers: BulletMarkers): number { + return markers.blockStructure ? indentOf(indent) : indent.length; +} + +function indentOf(line: string): number { + let col = 0; + for (const ch of line) { + if (ch === ' ') col += 1; + else if (ch === '\t') col += 4 - (col % 4); + else break; + } + return col; +} + +/** A line classified as a list-item opener by `matchListOpener`. */ +interface ListOpener { + /** Width of the indent before the marker. */ + indent: number; +} + +/** + * Classify `line` as a list-item opener under `markers`, applying the + * ORDERED-START rule (#3702 round 2, B2; round 5, M1/M2): a dot-terminated + * ordered marker opens an item when it starts at `0.` or `1.` (`01.` + * included), or when a list is already open at its level (`inList` — the + * caller's per-indent memory). + * + * Why: `\d{1,9}\.` alone reads ordinary prose as a list. "2026. was a bad + * year for this module" and, under a `### Notes` heading, "3. is the number + * of retries we settled on." are both sentences, and both opened an entry on + * round 1 — the second one straight through the "prose is not an item" + * contract that round claimed to preserve. CommonMark §5.3 faces the same + * ambiguity when an ordered list would interrupt a paragraph and resolves it + * the same way: the list must start with 1. This parser has no paragraph + * model, so it applies that rule wherever NO list is open at the line's + * level — the positions a sentence can occupy. Where a list IS open, + * CommonMark accepts any start (a list item interrupts no paragraph), and so + * does this. Numbers after the first are ignored, as CommonMark ignores + * them, so `1. / 3. / 7.` is a three-item run. + * + * `0.` is accepted as a start (round 5, M1): CommonMark §5.2 permits any + * 1-9-digit start number and a `0.`-numbered list is ordinary; refusing it + * dropped ONLY the first item, since the run then started at `1.` — the + * under-report that looks like a clean parse. A sentence opening with "0." + * is not a shape anyone writes. + * + * Stated cost, pinned by test: a list whose first ordinal is 2 or more, at a + * paragraph position, reads as prose UNTIL its first `0.`/`1.` line — the + * loss is that prefix, not the whole list. Every ordered record the #3702 + * scan found starts at 1, and the hyphen-style `- ` alternative loses + * nothing, so the trade buys the prose contract back at no measured cost. + * + * Bullet markers carry no rule — an asterisk bullet is not prose. + */ +function matchListOpener(line: string, markers: BulletMarkers, inList: boolean): ListOpener | null { + const m = line.match(markers.open); + if (!m) return null; + const token = m[2]; + if (/^\d/.test(token) && !inList && parseInt(token, 10) > 1) return null; + return { indent: indentWidth(m[1], markers) }; } /** @@ -2450,8 +2747,10 @@ function stripLeadingBulletMarker(line: string): string { * level; deepest level) each mis-count one of these shapes. * * A leaf entry is [heading text, ...body lines up to the next heading] and is - * kept only when its body (minus table lines) contains at least one `- ` - * bullet: + * kept only when its body (minus table lines) contains at least one list item + * — any of `-`, `*`, `+` or a dot-terminated ordered marker (#3702; the + * hyphen-only form silently dropped `*`/`+`/ordered bodies), where an + * ordered marker counts only under `matchListOpener`'s ordered-start rule: * - a prose-only or bare heading contributes nothing — "prose is not an item" * is this parser's pre-existing contract (see the `# Notes` case); * - a table-only body is left entirely to `parseDeferredTableItems`, which @@ -2463,11 +2762,55 @@ function stripLeadingBulletMarker(line: string): string { * `splitGapsEntries` — headless parity, so loose bullets before a later * heading group (the mixed shape) stay one item each. */ -function splitDeferredHeadingEntries(sectionBody: string): string[][] | null { +/** + * A heading-path entry: its lines, the splitter's per-line opener verdict, + * and — since #3781 — where it sits in the section body, so the writer can + * anchor a rewrite without a second walk. + */ +interface DeferredHeadingEntry { + lines: string[]; + /** `opener[i]` — line `i` opened a list item under `matchListOpener`'s rules. */ + opener: boolean[]; + /** `leaf` — a childless heading's entry, `lines[0]` its heading TEXT; `pending` — a headless bullet in the preamble or directly under a container heading. */ + kind: 'leaf' | 'pending'; + /** Character span within the section body — `sectionBody.slice(start, end)` is the entry's own raw text, CRLF intact, its lines 1:1 with `lines`. Both -1 when `embeddedTable`. */ + start: number; + end: number; + /** A GFM table row sat inside the span. Table lines belong to `parseDeferredTableItems` and are never entry lines, so the raw slice and `lines` disagree and no write can be anchored (#3781). */ + embeddedTable: boolean; +} + +/** Character offset of each line's start and end within the text they were split from (on `\n`). */ +function lineOffsets(lines: string[]): { lineStarts: number[]; lineEnds: number[] } { + const lineStarts: number[] = []; + const lineEnds: number[] = []; + let cursor = 0; + for (const line of lines) { + lineStarts.push(cursor); + cursor += line.length; + lineEnds.push(cursor); + cursor += 1; // the '\n' separator — absent after the final line, but nothing reads past it + } + return { lineStarts, lineEnds }; +} + +/** + * The heading-delimited split, carrying the per-line opener flags the deferred + * field-extraction path needs (#3702 round 2, round review): the heading path + * strips the marker off EVERY body line before field extraction (#3457), and a + * line whose ordinal `matchListOpener` REJECTED must not be stripped — or + * "3. status: resolved" as prose loses its `3. ` and reads as a resolved field. + * + * Since #3781 it also records each entry's character span (see + * `DeferredHeadingEntry`) in this same pass — the reader's walk IS the + * writer's walk, so there is no second copy of the grouping rules to drift. + */ +function splitDeferredHeadingEntriesDetailed(sectionBody: string): DeferredHeadingEntry[] | null { const headings = tokenizeHeadings(sectionBody); if (headings.length === 0) return null; const lines = sectionBody.split('\n'); + const { lineStarts, lineEnds } = lineOffsets(lines); const headingByLine = new Map(); for (let i = 0; i < headings.length; i++) { // Container iff the next heading is deeper (see doc comment). An empty @@ -2477,21 +2820,92 @@ function splitDeferredHeadingEntries(sectionBody: string): string[][] | null { headingByLine.set(headings[i].line, { text: headings[i].text, isContainer }); } - const entries: string[][] = []; - let current: string[] | null = null; // accumulating a leaf heading's entry - let pending: string[] = []; // preamble / container-heading body lines + const entries: DeferredHeadingEntry[] = []; + // The leaf entry being accumulated, with the raw line range it spans. + let current: { lines: string[]; opener: boolean[]; startLine: number; endLine: number } | null = null; let currentHasBullet = false; + // The headless-shaped region being accumulated (preamble / a container + // heading's direct lines): the reader's table-filtered, CR-stripped view, + // plus the raw line range it spans. + let pending: string[] = []; + let pendingStartLine = -1; + let pendingEndLine = -1; + // Table lines are never entry lines; where one sits INSIDE an entry's raw + // range, that entry's span is non-contiguous (#3781, `embeddedTable`). + const tableLines: number[] = []; + const tableWithin = (from: number, to: number): boolean => tableLines.some((t) => t >= from && t <= to); + // List memory for the leaf body being accumulated (#3702 round 2, B2) — + // reset at every heading, so `### Notes` + "3. is the number…" is prose + // while `### Steps` + "1. do / 2. then" is a list. A blank line then a + // non-indented non-list line is a PARAGRAPH, which ends the list + // (CommonMark §5.3); a non-indented line with no blank before it is lazy + // continuation and keeps it open. + const runs = new ListRuns(); + let blankSeen = false; + // Same level rule as the headless splitter: the first opener in a leaf body + // sets the base, and every indent at or shallower than it is one level. + let bodyBase: number | null = null; + const levelOf = (line: string): number => { + const ind = indentOf(line); + return bodyBase !== null && ind <= bodyBase ? bodyBase : ind; + }; + let scan = scanFencesFrom(lines, 0); const flushCurrent = (): void => { - // Keep the leaf entry only when its body carries a bullet; the heading + // Keep the leaf entry only when its body carries a list item; the heading // text line itself (element 0) never counts as one. - if (current !== null && currentHasBullet) entries.push(current); + if (current !== null && currentHasBullet) { + const table = tableWithin(current.startLine, current.endLine); + entries.push({ + lines: current.lines, + opener: current.opener, + kind: 'leaf', + start: table ? -1 : lineStarts[current.startLine], + end: table ? -1 : lineEnds[current.endLine], + embeddedTable: table, + }); + } current = null; currentHasBullet = false; }; const flushPending = (): void => { - entries.push(...splitGapsEntries(pending.join('\n'))); + if (pendingStartLine !== -1) { + // Headless-region entries carry the splitter's own opener flags — the + // same run state (ordered start, paragraph reset) that split them. The + // region is contiguous (a heading flushes it), so the core's + // region-relative spans translate by the region's own offset — unless a + // table row was skipped inside it, where the reader's view and the raw + // region disagree and no span is claimed. Both views split identically + // otherwise: the core CR-strips per line, and a table row is the only + // line the reader's view omits. + const table = tableWithin(pendingStartLine, pendingEndLine); + const region = table ? pending.join('\n') : lines.slice(pendingStartLine, pendingEndLine + 1).join('\n'); + const base = lineStarts[pendingStartLine]; + for (const { lines: entryLines, opener, start, end } of splitGapsEntriesCore(region, DEFERRED_BULLET_MARKERS)) { + entries.push({ + lines: entryLines, + opener, + kind: 'pending', + start: table ? -1 : base + start, + end: table ? -1 : base + end, + embeddedTable: table, + }); + } + } pending = []; + pendingStartLine = -1; + pendingEndLine = -1; + }; + const push = (line: string, i: number, opener: boolean): void => { + if (current !== null) { + current.lines.push(line); + current.opener.push(opener); + current.endLine = i; + } else { + pending.push(line); + if (pendingStartLine === -1) pendingStartLine = i; + pendingEndLine = i; + } }; for (let i = 0; i < lines.length; i++) { @@ -2503,20 +2917,69 @@ function splitDeferredHeadingEntries(sectionBody: string): string[][] | null { // ANY heading; flushing here keeps entries in document order even when // a container's direct bullets precede its first child entry. flushPending(); + runs.clear(); + blankSeen = false; + bodyBase = null; + // A heading ends the entry, and with it any unterminated fence (B1). + scan = scanFencesFrom(lines, i + 1); if (!heading.isContainer) { // Leaf heading: open an entry with the heading text as line 0. - current = [heading.text]; + current = { lines: [heading.text], opener: [false], startLine: i, endLine: i }; currentHasBullet = false; } continue; } + // CR-strip ONCE and carry the stripped line everywhere below — into the + // entry itself included (#3702 round 2, B1). `collectSection` slices raw + // `\n`-split lines, so on a CRLF file every body line but the last still + // carries its `\r`; the per-line marker strip feeding field extraction is + // `$`-anchored and fails on such a line, the marker survives into + // `extractGapEntryFields`, and the field is silently lost — a `**Status:**` + // that is not the file's final line then resurfaces its entry as open. + // The headless path (`splitGapsEntriesCore`) already stores stripped lines. + const line = lines[i].replace(/\r$/, ''); + // B1: an unterminated fence runs to the end of its entry. The next line + // shaped like a top-level item ends it — rescan from there, so a later + // delimiter is read on its own terms (see `scanFencesFrom`). + if (scan.unterminatedFrom !== -1 && i > scan.unterminatedFrom && topLevelItemShape(line, DEFERRED_BULLET_MARKERS, bodyBase)) { + scan = scanFencesFrom(lines, i); + } + if (scan.fenced.has(i) || (scan.unterminatedFrom !== -1 && i >= scan.unterminatedFrom)) { + // Fence content is body text, never list-item evidence (M2) — and + // never an opener, so it is never marker-stripped for fields either. + // Its opener ends the runs at its level and deeper, as a paragraph does. + if (scan.openers.has(i)) runs.endedAt(levelOf(line)); + push(line, i, false); + continue; + } + // A thematic break is a separator: not evidence, and it clears the list + // memory (M1). It stays a BODY line — the entry's span must stay + // contiguous for the writer, and the entry's name stays what `next` + // reported for a body containing one (round 5, m3). + if (THEMATIC_BREAK_RE.test(line)) { + runs.clear(); + blankSeen = false; + push(line, i, false); + continue; + } // Table lines belong to parseDeferredTableItems, never to a heading entry. - if (/^\s*\|/.test(lines[i].replace(/\r$/, ''))) continue; + if (/^\s*\|/.test(line)) { + tableLines.push(i); + continue; + } if (current !== null) { - current.push(lines[i]); - if (/^\s*-\s/.test(lines[i].replace(/\r$/, ''))) currentHasBullet = true; + const opener = matchListOpener(line, DEFERRED_BULLET_MARKERS, runs.at(levelOf(line))); + push(line, i, opener !== null); + if (opener !== null) { + currentHasBullet = true; + if (bodyBase === null) bodyBase = opener.indent; + runs.opened(levelOf(line)); + } else if (blankSeen && line.trim() !== '') { + runs.endedAt(levelOf(line)); // a paragraph after a blank line ends the lists at its level and deeper + } + blankSeen = line.trim() === ''; } else { - pending.push(lines[i]); + push(line, i, false); // the core derives its own opener verdicts for the region } } flushCurrent(); @@ -2575,6 +3038,8 @@ function parseDeferredTableItems(sectionBody: string): UatItem[] { */ interface GapsEntrySpan { lines: string[]; + /** `opener[i]` — line `i` was an ACCEPTED list opener (line 0 always; a nested one when accepted). */ + opener: boolean[]; start: number; end: number; } @@ -2588,75 +3053,145 @@ interface GapsEntrySpan { * boundary — a second, independently-written grouping pass is exactly how a * span-carrying sibling could disagree with the plain-lines version it is * supposed to be span-annotating. + * + * `markers` selects the marker set an entry may OPEN with (#3702). It defaults + * to the hyphen-only Gaps form, so every pre-existing caller is unaffected; + * the deferred-items callers pass `DEFERRED_BULLET_MARKERS`. Parameterising + * the shared seam — rather than widening it in place — is what keeps the + * template-mandated Gaps grammar out of the deferred-items ruling's blast + * radius while still leaving exactly ONE grouping pass in the module. */ -function splitGapsEntriesCore(sectionBody: string): GapsEntrySpan[] { +function splitGapsEntriesCore( + sectionBody: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, +): GapsEntrySpan[] { const rawLines = sectionBody.split('\n'); - const lineStarts: number[] = []; - const lineEnds: number[] = []; - let cursor = 0; - for (const rawLine of rawLines) { - lineStarts.push(cursor); - cursor += rawLine.length; - lineEnds.push(cursor); - cursor += 1; // the '\n' separator — absent after the final line, but nothing reads past it - } + const { lineStarts, lineEnds } = lineOffsets(rawLines); const entries: GapsEntrySpan[] = []; let current: string[] | null = null; let currentStartLine = -1; let currentEndLine = -1; let baseIndent: number | null = null; + // List memory per indent (#3702 round 2, B2 + round review; round 5, M2) — + // the top level decides entry boundaries; nested levels decide only which + // continuation lines count as accepted openers for field stripping. + const runs = new ListRuns(); + // Per-line opener flags for `current`, recorded HERE — the one place the + // run state is known — so the heading path's strip-only-openers rule reads + // the splitter's own verdict instead of re-deriving it (round review: a + // re-derivation without the paragraph reset re-accepted a rejected ordinal). + let currentOpeners: boolean[] = []; const flush = (): void => { if (current !== null) { - entries.push({ lines: current, start: lineStarts[currentStartLine], end: lineEnds[currentEndLine] }); + entries.push({ lines: current, opener: currentOpeners, start: lineStarts[currentStartLine], end: lineEnds[currentEndLine] }); } }; - // #3898: a spaced-hyphen thematic break (`- - -`, `- -`, `- - -`, …) is a - // SEPARATOR, not an entry. The opener regex below matches it (hyphen + - // whitespace), which fabricated a gap named `- -` with result 'unknown' — - // an item that cannot be cleared by editing any entry, because there is no - // entry, only the separator the author wrote deliberately. A line whose - // content after the opening marker consists solely of hyphens and spaces - // (with at least one further hyphen) is skipped entirely: it neither opens - // an entry nor is folded into the current one. This is deliberately NOT a - // full thematic-break concept (option 2 in the issue): a break does not - // close the Gaps list — entries after it keep parsing. + // #3898 (from `next`): a spaced-hyphen thematic break (`- - -`, `- -`, + // `- - -`, …) in `## Gaps` is a SEPARATOR, not an entry. The hyphen opener + // matches it (hyphen + whitespace), which fabricated a gap named `- -` with + // result 'unknown' — an item no edit can clear, because there is no entry, + // only the separator the author wrote deliberately. A line whose content + // after the opening marker is solely hyphens and spaces (with at least one + // further hyphen) is skipped: it neither opens an entry nor folds into the + // current one. Deliberately NOT a full thematic-break concept (option 2 in + // the issue): a break does not close the Gaps list — entries after it keep + // parsing. The deferred grammar (`blockStructure`) has its own, CommonMark + // reading of the same line through THEMATIC_BREAK_RE below, where a break + // CLOSES the list; this helper is consulted only for the Gaps set. const isSeparatorShaped = (line: string, bulletPrefixLen: number): boolean => { const remainder = line.slice(bulletPrefixLen); return /^[-\s]*$/.test(remainder) && remainder.includes('-'); }; + // Block structure (M1/M2 + column indents) is a property of the GRAMMAR, + // not of this seam: the Gaps set opts out and stays byte-for-byte on its + // `next` behaviour — see `indentWidth` for the indent half of that opt-out. + let scan = markers.blockStructure ? scanFencesFrom(rawLines, 0) : NO_FENCES; + let blankSeen = false; + // The run LEVEL of a line: every indent at or shallower than the list's + // base is the one top level (a dedenting list keeps its entry boundaries); + // deeper indents are their own nested levels. + const levelOf = (line: string): number => { + const ind = indentWidth(line.match(/^[ \t]*/)![0], markers); + return baseIndent !== null && ind <= baseIndent ? baseIndent : ind; + }; rawLines.forEach((rawLine, idx) => { const line = rawLine.replace(/\r$/, ''); - const bulletMatch = line.match(/^(\s*)-\s/); - // Narrowed skip (review disposition a): a separator-shaped line is skipped - // only when it sits BETWEEN entries (nothing open yet, or it would open a - // top-level entry — where the phantom came from). One landing strictly - // INSIDE a live entry (indent > baseIndent) folds back as a continuation - // line, so the entry's GapsEntrySpan stays byte-contiguous — the span - // invariant below and the ack writer's identity re-verification both hold. - if ( - bulletMatch && isSeparatorShaped(line, bulletMatch[0].length) && - (current === null || bulletMatch[1].length <= (baseIndent ?? 0)) - ) { - return; // separator line between entries — neither an opener nor a continuation + // B1: an unterminated fence runs to the end of its entry — the next line + // shaped like a top-level item ends it; rescan from there so a later + // delimiter is read on its own terms (see `scanFencesFrom`). + if (scan.unterminatedFrom !== -1 && idx > scan.unterminatedFrom && topLevelItemShape(line, markers, baseIndent)) { + scan = scanFencesFrom(rawLines, idx); } - if (bulletMatch) { - const indent = bulletMatch[1].length; + if (scan.fenced.has(idx) || (scan.unterminatedFrom !== -1 && idx >= scan.unterminatedFrom)) { + // Fence content never opens an entry (M2). Inside an open entry it is + // continuation — pushed, so the span invariant `acknowledgeDeferredItem` + // re-verifies still holds; before the first entry it is discarded. A + // fence is a non-list block: its opener ends the runs at its level and + // deeper, exactly as a paragraph does. + if (scan.openers.has(idx)) runs.endedAt(levelOf(line)); + if (current !== null) { + current.push(line); + currentOpeners.push(false); + currentEndLine = idx; + } + return; + } + if (markers.blockStructure && THEMATIC_BREAK_RE.test(line)) { + // A thematic break closes the list (M1): the open entry ends here, the + // break itself is neither an item nor a continuation, and nothing after + // it joins the closed entry — the next opener starts fresh. + flush(); + current = null; + runs.clear(); + blankSeen = false; + return; + } + // #3898 narrowed skip (review disposition a), Gaps set only: a + // separator-shaped line is skipped when it sits BETWEEN entries (nothing + // open yet, or it would open a top-level entry — where the phantom came + // from). One landing strictly INSIDE a live entry (indent > baseIndent) + // folds back as a continuation line, so the entry's GapsEntrySpan stays + // byte-contiguous — the span invariant below and the ack writer's identity + // re-verification both hold. The indent compare is the raw character + // count, which is what `indentWidth` measures for the Gaps set. + if (!markers.blockStructure) { + const bulletMatch = line.match(/^(\s*)-\s/); + if ( + bulletMatch && isSeparatorShaped(line, bulletMatch[0].length) && + (current === null || bulletMatch[1].length <= (baseIndent ?? 0)) + ) { + return; // separator line between entries — neither an opener nor a continuation + } + } + const opener = matchListOpener(line, markers, runs.at(levelOf(line))); + if (opener !== null) { + const { indent } = opener; if (baseIndent === null) baseIndent = indent; + runs.opened(levelOf(line)); if (indent <= baseIndent) { flush(); current = [line]; + currentOpeners = [true]; currentStartLine = idx; currentEndLine = idx; + blankSeen = false; // an opener is not blank — the memory must not survive it return; } } if (current !== null) { current.push(line); + currentOpeners.push(opener !== null); currentEndLine = idx; + // A blank line then a top-level non-list line is a PARAGRAPH: the list + // is over (CommonMark §5.3) and a later `5. x` is prose. Without the + // blank it is lazy continuation and the run stays open. + const blank = line.trim() === ''; + if (!blank && blankSeen && opener === null) runs.endedAt(levelOf(line)); + blankSeen = opener === null && blank; } // else: pre-first-bullet content (e.g. the template's HTML comment) — discarded. }); @@ -2667,7 +3202,8 @@ function splitGapsEntriesCore(sectionBody: string): GapsEntrySpan[] { /** * Split a `## Gaps` section body into per-entry line groups on TOP-LEVEL - * `- ` bullet openers. + * bullet openers — `- ` for Gaps, or whichever set `markers` names (#3702: + * the deferred-items callers pass the widened CommonMark set). * * The indentation of the FIRST bullet line encountered establishes the * "top-level" indent for the whole section; any subsequent `- `-opening line @@ -2680,17 +3216,21 @@ function splitGapsEntriesCore(sectionBody: string): GapsEntrySpan[] { * * Lines before the first bullet (e.g. the `` comment * the template emits) are discarded. An empty/whitespace-only section body - * (heading present, no bullets) returns `[]`. + * (heading present, no bullets) returns `[]`. Fenced code never opens an + * entry and a thematic break closes the open one (#3702 round 2, M1/M2). */ -function splitGapsEntries(sectionBody: string): string[][] { - return splitGapsEntriesCore(sectionBody).map((entry) => entry.lines); +function splitGapsEntries( + sectionBody: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, +): string[][] { + return splitGapsEntriesCore(sectionBody, markers).map((entry) => entry.lines); } /** * Sibling of `splitGapsEntries` (F1, #3458 follow-up review) that ADDITIVELY * carries each entry's character span — every existing `splitGapsEntries` * caller (`parseGapsItems`, `parseDeferredItemsWithStatus`, - * `splitDeferredHeadingEntries`'s `flushPending`) is unaffected and keeps + * `splitDeferredHeadingEntriesDetailed`'s `flushPending`) is unaffected and keeps * using the plain `lines`-only shape. `acknowledgeDeferredItem` is the one * caller that needs a span: it used to select an entry via `splitGapsEntries` * and then RE-FIND that entry's location with a fresh regex search over @@ -2702,8 +3242,11 @@ function splitGapsEntries(sectionBody: string): string[][] { * one. Carrying the span out of THIS same pass — the one that already knows * exactly where the entry lives — removes the re-derivation step entirely. */ -function splitGapsEntriesWithSpans(sectionBody: string): GapsEntrySpan[] { - return splitGapsEntriesCore(sectionBody); +function splitGapsEntriesWithSpans( + sectionBody: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, +): GapsEntrySpan[] { + return splitGapsEntriesCore(sectionBody, markers); } /** @@ -2732,38 +3275,175 @@ function splitGapsEntriesWithSpans(sectionBody: string): GapsEntrySpan[] { * keep their literal case, and mid-line emphasis is untouched, preserving the * start-anchored decoy invariant above. */ -function extractGapEntryFields(entryLines: string[]): Record { +function extractGapEntryFields( + entryLines: string[], + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + openerFlags?: boolean[], +): Record { const fields: Record = {}; - const fieldLineRe = /^([A-Za-z_][A-Za-z0-9_-]*):\s*(.*)$/; - const boldedKeyRe = /^\*+([A-Za-z_][A-Za-z0-9_-]*):\*+/; - - entryLines.forEach((rawLine, idx) => { - const line = rawLine.replace(/\r$/, ''); - // Strip ONLY the entry-opening bullet marker (idx 0); a bullet marker on - // a later line belongs to a nested sub-list and is handled by - // `splitGapsEntries` already folding it in — it is not itself a field - // line unless it independently matches `key: value` after stripping. - const bulletStripped = line.match(/^(\s*)-\s+(.*)$/); - const content = (idx === 0 && bulletStripped ? bulletStripped[2] : line.trim()) - .replace(boldedKeyRe, (_m, key: string) => `${key.toLowerCase()}:`); - - const m = fieldLineRe.exec(content); - if (!m) return; - const key = m[1]; - let value = m[2].trim(); - if (value.startsWith('"') && value.endsWith('"') && value.length >= 2) { - value = value.slice(1, -1); - } - if (!(key in fields)) fields[key] = value; + // A fenced line is content, not a field (#3702 round 2, round review): the + // splitters already keep fence lines from OPENING an entry, and a + // `status: resolved` quoted inside a code block must not resolve one either. + // An entry is a contiguous slice and a fence never spans two entries (the + // opener of the next entry would be fence content), so scanning the entry's + // own lines classifies exactly what the splitter classified. + // + // RAW lines, and that is the fix for #3702 round 3, m7. The heading path + // used to marker-strip its lines BEFORE calling this function, so the scan + // below ran over text the splitter never saw: `- ```sh` is an ordinary + // bullet to the splitter, but strips to ```` ```sh ````, which opens a + // fence here that exists in no other pass. A `**Status:** resolved` line + // after it was then suppressed as fence content and its resolved entry + // resurfaced as open. The stripping now happens INSIDE this function, after + // the fence scan, driven by the splitter's own per-line opener verdict. + entryFieldLines(entryLines, markers, openerFlags).forEach((field) => { + if (!field) return; + if (!(field.key in fields)) fields[field.key] = field.value; }); return fields; } -/** Fallback display text for a Gaps entry with no parseable `truth:` field. */ -function rawGapEntryText(entryLines: string[]): string { +/** + * Per line of an entry, the field it declares — or `null` where it declares + * none, INCLUDING because it is fenced. + * + * This is the seam, and it exists because `parseGapEntryFieldLine` alone was + * not it (#3702 round 3, pre-push review). The reader applied the fence gate + * before classifying and the acknowledge writer did not, so a `status:` line + * inside a fenced block was selected by the writer and skipped by the reader: + * the write produced a line nothing reads, the read-back guard refused it, and + * the entry became impossible to acknowledge at all — `audit acknowledge` + * surfaced an internal error and `complete-milestone` halted on it. That shape + * acknowledged cleanly on `next`, so it was a regression introduced by the fix + * for the nested-marker one, and the claim "the writer cannot select a line the + * reader will not read back" was false while the fence gate lived on one side. + * + * Both sides call this now, so the claim is structural rather than asserted. + */ +function entryFieldLines( + entryLines: string[], + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + openerFlags?: boolean[], +): ({ key: string; value: string; valueStart: number } | null)[] { + const fenced = entryFencedLines(entryLines, markers, openerFlags); + return entryLines.map((rawLine, idx) => ( + fenced.has(idx) ? null : parseGapEntryFieldLine(rawLine, markers, stripsMarkerAt(idx, openerFlags)) + )); +} + +/** + * The fenced lines of ONE entry, as the reader and the writer both see them. + * A leaf's line 0 is its heading TEXT, not a Markdown line: a heading that + * reads ``` or ~~~ is a heading, and must not open a fence over the body + * beneath it (round 5, RV6.5 — it fenced every field line, so the reader + * read nothing and the writer's marker landed on a line nothing reads). + * `openerFlags[0] === false` is the leaf tell: a pending or headless entry's + * line 0 is an accepted opener, and a marker line is never a delimiter. + */ +function entryFencedLines(entryLines: string[], markers: BulletMarkers, openerFlags?: boolean[]): Set { + if (!markers.blockStructure) return new Set(); + const leaf = openerFlags !== undefined && openerFlags[0] === false; + return fencedLineSet(leaf ? ['', ...entryLines.slice(1)] : entryLines); +} + +/** + * Which lines of an entry carry an entry-opening marker to be stripped before + * the line is read as a field. + * + * Without flags — the headless and `## Gaps` shapes — that is line 0 alone: a + * marker on a later line belongs to a nested sub-list (`splitGapsEntries` + * already folded it in) and is not a field line unless it independently + * matches `key: value` after a plain trim. + * + * With flags — the heading shape — it is whichever lines the SPLITTER accepted + * as list openers, because there every body line may be a sibling bullet + * carrying a field (#3457) while line 0 is the heading TEXT and carries no + * marker at all. Reading the splitter's verdict rather than re-deriving it is + * what keeps a rejected ordinal (`3. status: resolved` as prose) from being + * stripped into a field. + */ +function stripsMarkerAt(idx: number, openerFlags?: boolean[]): boolean { + return openerFlags ? openerFlags[idx] === true : idx === 0; +} + +/** + * The ONE place an entry line is classified as a `key: value` field line. + * `extractGapEntryFields` reads through it, and `acknowledgeDeferredItem` + * locates the line it will rewrite through it. + * + * Sharing the classifier is what makes the writer structurally unable to + * select a line the reader will not read back (#3702 round 3, B1; the shape + * is #3773's, parameterised here by `markers` per the round-3 review's + * prescribed end state). The writer used to carry its own marker-widened + * status regex, so a nested ` * status: pending` was selectable by the + * writer and invisible to this reader: acknowledge rewrote it in place, + * returned `ok`, and the item stayed outstanding forever. A single classifier + * has no second copy to drift from. + * + * `valueStart` is the offset, in the CR-stripped line, at which the VALUE + * begins — so a rewrite can replace the value without a second regex of its + * own. The bolded-key unwrap below is a PREFIX rewrite, so the tail of the + * rewritten content is byte-identical to the tail of the original and the + * offset maps back directly. + * + * Returns `null` for a non-field line. + */ +function parseGapEntryFieldLine( + rawLine: string, + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + stripMarker = true, +): { key: string; value: string; valueStart: number } | null { + const fieldLineRe = /^([A-Za-z_][A-Za-z0-9_-]*):\s*(.*)$/; + const boldedKeyRe = /^\*+([A-Za-z_][A-Za-z0-9_-]*):\*+/; + + const line = rawLine.replace(/\r$/, ''); + const bulletStripped = stripMarker ? line.match(markers.strip) : null; + const bare = bulletStripped ? bulletStripped[2] : line.trim(); + // Where `bare` begins in `line`. The two branches differ: the marker strip's + // group 2 runs to end-of-line, so it is a plain suffix; `trim()` also cuts + // the tail, so its offset is the LEADING run alone. Computing one from the + // other's shape under-counts by the trailing whitespace. + const headLen = bulletStripped ? line.length - bare.length : line.length - line.trimStart().length; + const content = bare.replace(boldedKeyRe, (_m, key: string) => `${key.toLowerCase()}:`); + + const m = fieldLineRe.exec(content); + if (!m) return null; + let value = m[2].trim(); + if (value.startsWith('"') && value.endsWith('"') && value.length >= 2) { + value = value.slice(1, -1); + } + return { key: m[1], value, valueStart: headLen + (bare.length - m[2].length) }; +} + +/** + * Fallback display text for a Gaps entry with no parseable `truth:` field. + * + * `markers` selects which opening marker is stripped (#3702) — hyphen-only by + * default, the widened set for deferred-items callers, so a `*`-opened entry + * renders the same name its hyphen twin would. That name is the key + * `acknowledgeDeferredItem` matches on, so the two MUST use the same set: + * rendering `* alpha` where the parse surfaced `alpha` would make the entry + * un-acknowledgeable. + * + * `openerFlags` decides WHICH lines are stripped, and on the heading shape + * line 0 is not one of them (#3702 round 3, m8). There line 0 is the heading + * TEXT, so an unconditional strip renamed `### 1. Race in the writer` to + * `Race in the writer` and `### * starred title` to `starred title` — both + * silent renames of the very key acknowledge matches on, and both a change + * from this parser's behaviour on `next`. + */ +function rawGapEntryText( + entryLines: string[], + markers: BulletMarkers = HYPHEN_BULLET_MARKERS, + openerFlags?: boolean[], +): string { return entryLines - .map((l, i) => (i === 0 ? l.replace(/^(\s*)-\s+/, '') : l.trim())) + // Line 0 ONLY, and only if the splitter accepted it as an opener. The + // opener flags say which lines carry a marker; the entry's NAME is a + // different question, and stripping a body line's marker out of it changes + // the key `acknowledgeDeferredItem` matches on. + .map((l, i) => (i === 0 && stripsMarkerAt(0, openerFlags) ? l.replace(markers.strip, '$2') : l.trim())) .join(' ') .trim(); } @@ -2998,4 +3678,13 @@ export = { parseDeferredItems, parseDeferredItemsWithStatus, acknowledgeDeferredItem, + // #3702 round 2 (M3): exposed for the marker-grammar parity test only. + // Narrowed in round 3 (M6): the two status-line regexes are gone from the + // module, so nothing exports them, and the parity test they were widened for + // could not reach the defect it was meant to guard anyway — it asserted the + // four WRITER regexes shared a source string, which is true of a detect/read + // asymmetry too. The pair below is what the behavioural parity test against + // `iterateBullets` actually reads. + DEFERRED_MARKER_ALT, + DEFERRED_BULLET_MARKERS, }; diff --git a/tests/audit-command-cutover.test.cjs b/tests/audit-command-cutover.test.cjs index d4ee199e7..0d78fe7f5 100644 --- a/tests/audit-command-cutover.test.cjs +++ b/tests/audit-command-cutover.test.cjs @@ -2094,10 +2094,12 @@ describe('bug #950: quick-task SUMMARY must carry status: complete', () => { assert.match(fs.readFileSync(filePath, 'utf-8'), /status: acknowledged/); // A table row embedded in the entry's span is still refused — a - // non-contiguous span cannot anchor a safe write. - const tabled = ['## Deferred Items', '', '### Tabled finding', '', '- a bullet', '| x | y |', ''].join('\n'); + // non-contiguous span cannot anchor a safe write. The row sits BETWEEN + // two entry lines: a row after the entry's last line is outside its + // span, and that write is anchored and safe (#3702 round 5). + const tabled = ['## Deferred Items', '', '### Tabled finding', '', '- a bullet', '| x | y |', '- more evidence', ''].join('\n'); fs.writeFileSync(filePath, tabled); - const refuse = ack(tmpDir, ['--category', 'deferred_items', '--phase', '01', '--file', 'deferred-items.md', '--text', 'Tabled finding - a bullet', '--milestone', 'v1.0']); + const refuse = ack(tmpDir, ['--category', 'deferred_items', '--phase', '01', '--file', 'deferred-items.md', '--text', 'Tabled finding - a bullet - more evidence', '--milestone', 'v1.0']); assert.equal(refuse.success, false, 'table-embedded spans must still refuse'); assert.match(refuse.error, /GFM table row/i); assert.equal(fs.readFileSync(filePath, 'utf-8'), tabled, 'file must be byte-identical — nothing written on refusal'); @@ -2112,6 +2114,43 @@ describe('bug #950: quick-task SUMMARY must carry status: complete', () => { assert.equal(fs.readFileSync(filePath, 'utf-8'), proseOnly, 'file must be byte-identical'); }); + test('#3702 round 2 (m3, updated for #3781): a heading-delimited file written with `*`, `+` or `1.` SURFACES its entries, and the CLI writer acknowledges them', () => { + // Before #3702 such a file parsed to zero entries and `complete-milestone` + // closed over it silently. It now yields entries; under #3781 (merged in + // round 5) the heading shape is acknowledgeable, so the milestone loop + // suppresses them through the same two CLI calls it makes for a hyphen + // file: the listing that puts the entry in front of the writer, and the + // acknowledge that takes it. (Until #3781 the writer refused the shape + // and the loop halted — that contract is gone from `next`.) + const fixtures = [['*', '02-star'], ['+', '03-plus'], ['1.', '04-ordered']]; + const files = new Map(); + for (const [marker, slug] of fixtures) { + const phaseDir = planningPath('phases', slug); + fs.mkdirSync(phaseDir, { recursive: true }); + const filePath = path.join(phaseDir, 'deferred-items.md'); + const before = ['## Deferred Items', '', `### Out of scope under ${slug}`, '', `${marker} **What:** detail.`, ''].join('\n'); + fs.writeFileSync(filePath, before); + files.set(slug, { filePath, before }); + } + + const listed = audit(tmpDir).items.deferred_items.filter((i) => /^Out of scope under /.test(i.text)); + assert.equal(listed.length, fixtures.length, `every marker's heading entry must be listed: ${JSON.stringify(listed)}`); + + for (const item of listed) { + const slug = /under (\S+)/.exec(item.text)[1]; + const { filePath, before } = files.get(slug); + // `--text` is the audit's own reported text, exactly as the loop passes it. + const result = ack(tmpDir, ['--category', 'deferred_items', '--phase', slug.slice(0, 2), '--file', 'deferred-items.md', '--text', item.text, '--milestone', 'v1.0']); + assert.equal(result.success, true, `${slug}: a widened-marker heading entry must acknowledge; stderr: ${result.error}`); + const after = fs.readFileSync(filePath, 'utf-8'); + assert.notEqual(after, before, `${slug}: the file must carry the marker`); + assert.match(after, /status: acknowledged/, slug); + } + // And the loop's next listing no longer surfaces them. + const relisted = audit(tmpDir).items.deferred_items.filter((i) => /^Out of scope under /.test(i.text)); + assert.equal(relisted.length, 0, `acknowledged entries must leave the listing: ${JSON.stringify(relisted)}`); + }); + // ── F1 (#3458 follow-up review, HIGH): the writer must splice by the // SELECTED entry's own carried span, never re-find it by searching — // otherwise a byte-identical substring living inside an EARLIER entry diff --git a/tests/uat.test.cjs b/tests/uat.test.cjs index d50a95bb2..8859f2d8e 100644 --- a/tests/uat.test.cjs +++ b/tests/uat.test.cjs @@ -20,7 +20,10 @@ const { acknowledgeDeferredItem, parseUatItems, parseUatItemsWithStats, + DEFERRED_MARKER_ALT, + DEFERRED_BULLET_MARKERS, } = require('../gsd-core/bin/lib/uat.cjs'); +const { iterateBullets } = require('../gsd-core/bin/lib/markdown-sectionizer.cjs'); describe('audit-uat command', () => { let tmpDir; @@ -1801,9 +1804,9 @@ describe('#2287 progress.md forensic_audit: deferred-items.md contract', () => { }); }); -// ─── parseDeferredItems property test (#2287) ────────────────────────────── +// ─── parseDeferredItems property test (#2287, widened by #3702 round 2) ───── -describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () => { +describe('#2287 parseDeferredItems: property (status: resolved fail-safe) × marker × shape × line ending', () => { // Single-line entry text: no newlines (would break bullet-entry splitting), // non-empty after trim, and never itself SHAPED like a `status:` field line // (that would be indistinguishable from a real field regardless of intent). @@ -1818,43 +1821,88 @@ describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () = const decoyText = plainText.map((s) => `${s} status: resolved trailing note`); const textArb = fc.oneof(plainText, decoyText); - const entryArb = fc.record({ text: textArb, resolved: fc.boolean() }); + // `statusFirst` matters on the heading shape: round 1's only CRLF test put + // `**Status:**` LAST, the one line `collectSection`'s `.trimEnd()` had + // already de-CR'd, and the B1 regression hid behind it. + // `decoy` adds a prose line that BEGINS with a non-1 ordinal (`7. …`): under + // the start-at-1 rule (B2) it is never an item, never evidence and never a + // field — round 1 read it as an opener. Placed where it cannot end a run: + // before the first headless entry, and first in a heading body. + const entryArb = fc.record({ text: textArb, resolved: fc.boolean(), statusFirst: fc.boolean(), decoy: fc.boolean() }); + const decoyOrdinal = fc.integer({ min: 2, max: 999999999 }); + + // #3702 round 2 (B3): the marker set is an enumerated domain — exactly what a + // property is for. Ordered markers are numbered from 1 (the ordered-run + // rule, B2); the two shapes exercise both splitters; CRLF exercises the + // heading path's CR handling (B1). + const markerArb = fc.constantFrom('-', '*', '+', 'ordered'); + const shapeArb = fc.constantFrom('headless', 'heading'); + const eolArb = fc.constantFrom('\n', '\r\n'); + const mk = (marker, i) => (marker === 'ordered' ? `${i + 1}.` : marker); + + const render = (entries, marker, shape, eol, ordinal = 7) => { + const lines = ['## Deferred Items', '']; + if (shape === 'headless' && entries.some((e) => e.decoy)) { + // Pre-first-entry prose: a rejected ordinal is discarded, an accepted + // one (round 1) opens a phantom entry and breaks the count. + lines.push(`${ordinal}. ${entries.find((e) => e.decoy).text} status: resolved`, ''); + } + entries.forEach((e, i) => { + if (shape === 'headless') { + lines.push(`${mk(marker, i)} ${e.text}`); + if (e.resolved) lines.push(' status: resolved'); + } else { + lines.push(`### ${e.text}`, ''); + // A decoy ordinal line FIRST in the body: never stripped, so the + // `status: resolved` after it can never become a field. + if (e.decoy) lines.push(`${ordinal}. ${e.text} status: resolved`); + const what = (n) => `${mk(marker, n)} **What:** ${e.text}`; + const status = (n) => `${mk(marker, n)} **Status:** resolved`; + if (!e.resolved) lines.push(what(0)); + else if (e.statusFirst) lines.push(status(0), what(1)); + else lines.push(what(0), status(1)); + lines.push(''); + } + }); + return lines.join(eol); + }; + const idOf = (name) => { const m = /E(\d+)_/.exec(name); return m ? Number(m[1]) : -1; }; test('property: an entry is surfaced iff it is NOT marked status: resolved; surfaced count == non-resolved count', () => { fc.assert( fc.property( - fc.array(entryArb, { maxLength: 20 }), - (rawEntries) => { + fc.array(entryArb, { maxLength: 20 }), markerArb, shapeArb, eolArb, decoyOrdinal, + (rawEntries, marker, shape, eol, ordinal) => { // Index-prefix for uniqueness so surfaced items can be mapped back // to their source entry unambiguously even with colliding random text. - const entries = rawEntries.map((e, i) => ({ text: `E${i}_${e.text}`, resolved: e.resolved })); - - const lines = ['## Deferred Items', '']; - for (const e of entries) { - lines.push(`- ${e.text}`); - if (e.resolved) lines.push(' status: resolved'); - } - const content = lines.join('\n'); + const entries = rawEntries.map((e, i) => ({ ...e, text: `E${i}_${e.text}` })); + const content = render(entries, marker, shape, eol, ordinal); + const where = `${shape} ${JSON.stringify(marker)} ${JSON.stringify(eol)} decoy=${ordinal}`; const items = parseDeferredItems(content); - const surfacedNames = new Set(items.map((it) => it.name)); + const surfacedIds = new Set(items.map((it) => idOf(it.name))); const expectedUnresolved = entries.filter((e) => !e.resolved); - const expectedResolved = entries.filter((e) => e.resolved); // Total surfaced count equals the count of non-resolved entries. - assert.strictEqual(items.length, expectedUnresolved.length); + assert.strictEqual(items.length, expectedUnresolved.length, where); // Every non-resolved entry IS surfaced (including status:-shaped // decoy substrings embedded mid-line — those must not flip the // outcome). - for (const e of expectedUnresolved) { - assert.ok(surfacedNames.has(e.text), `expected unresolved entry to surface: ${e.text}`); + for (const [i, e] of entries.entries()) { + assert.strictEqual(surfacedIds.has(i), !e.resolved, `${where}: ${e.resolved ? 'resolved entry must never surface' : 'unresolved entry must surface'}: ${e.text}`); } - // No status:-resolved entry is EVER surfaced. - for (const e of expectedResolved) { - assert.ok(!surfacedNames.has(e.text), `status: resolved entry must never surface: ${e.text}`); + // Headless: the surfaced NAME is the entry text with the marker gone, + // whichever marker it was (the name is what acknowledge matches on). + if (shape === 'headless') { + for (const it of items) assert.strictEqual(it.name, entries[idOf(it.name)].text, where); + } + // A heading body carrying ONLY a decoy ordinal line is prose, not an entry. + if (shape === 'heading' && entries.length > 0) { + const decoyOnly = `## Deferred Items${eol}${eol}### only-decoy${eol}${eol}${ordinal}. ${entries[0].text} status: resolved${eol}`; + assert.deepStrictEqual(parseDeferredItems(decoyOnly), [], `${where}: decoy-only body`); } // Every returned item carries the fixed deferred category/result shape. @@ -1866,6 +1914,39 @@ describe('#2287 parseDeferredItems: property (status: resolved fail-safe)', () = ) ); }); + + test('property: acknowledge reaches and rewrites every unresolved headless entry, whichever marker or line ending', () => { + // The writer refuses the heading shape by design (`unsupported_heading_shape`), + // so this ranges over the headless shape only. It is the property that + // reaches M4 (CRLF rewrite reported ok and wrote nothing) and m2 (indent). + fc.assert( + fc.property( + fc.array(entryArb, { minLength: 1, maxLength: 12 }), markerArb, eolArb, + (rawEntries, marker, eol) => { + const entries = rawEntries.map((e, i) => ({ ...e, text: `E${i}_${e.text}` })); + let content = render(entries, marker, 'headless', eol); + const where = `${JSON.stringify(marker)} ${JSON.stringify(eol)}`; + + for (const e of entries.filter((x) => !x.resolved)) { + const got = acknowledgeDeferredItem(content, e.text); + assert.strictEqual(got.status, 'ok', `${where}: ${e.text}`); + assert.notStrictEqual(got.content, content, `${where}: an ok must have written: ${e.text}`); + content = got.content; + } + const after = parseDeferredItemsWithStatus(content); + assert.strictEqual(after.length, entries.length, where); + for (const it of after) { + const e = entries[idOf(it.name)]; + assert.strictEqual(it.status, e.resolved ? 'resolved' : 'acknowledged', `${where}: ${it.name}`); + } + // `acknowledged` is suppressed at the AUDIT layer, not the parser's: + // only `resolved` leaves the outstanding list here, so the count is + // unchanged by the writes above. + assert.strictEqual(parseDeferredItems(content).length, entries.filter((x) => !x.resolved).length, where); + } + ) + ); + }); }); // ─── #2766: archived phase dirs, and GFM-table-shaped deferred/gaps ──────── @@ -2313,6 +2394,1235 @@ describe('#3457 parseDeferredItems: heading-delimited entries', () => { }); }); +describe('#3702 parseDeferredItems: list-marker grammar', () => { + // `deferred-items.md` has no template and no mandated shape, but the parser + // recognised only the hyphen marker — so `*`, `+` and ordered lists (all + // lists in CommonMark and GFM) contributed ZERO entries on both the headless + // and the heading-delimited path. A mixed file dropped its non-hyphen entries + // while keeping their hyphenated siblings, under-reporting without ever + // looking empty. + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + const count = (md) => names(md).length; + + // The marker set the ruling widened to. `1)` is deliberately absent — the + // paren-terminated ordered form is out of scope for this fix, and the + // `1) still yields zero` case below pins that as intended, not as an oversight. + const MARKERS = ['-', '*', '+', '1.']; + + test('AC1 headless: every marker yields the same count as the hyphen form', () => { + const shape = (m) => `${m} alpha\n${m} beta\n`; + const hyphen = count(shape('-')); + + assert.strictEqual(hyphen, 2, 'baseline: the hyphen form must yield 2'); + for (const m of MARKERS) { + assert.strictEqual(count(shape(m)), hyphen, `marker ${JSON.stringify(m)}: ${JSON.stringify(names(shape(m)))}`); + } + }); + + test('AC1 headless: the entry NAME drops the marker, whichever marker it is', () => { + // rawGapEntryText renders the name acknowledgeDeferredItem later matches on, + // so a marker left in the rendered name would make the entry unreachable. + for (const m of MARKERS) { + assert.deepStrictEqual(names(`${m} alpha\n`), ['alpha'], `marker ${JSON.stringify(m)}`); + } + }); + + test('AC2 heading-delimited: a body carrying any marker is KEPT (was dropped)', () => { + const shape = (m) => `### Entry\n\n${m} **What:** x.\n`; + + for (const m of MARKERS) { + assert.strictEqual(count(shape(m)), 1, `marker ${JSON.stringify(m)}: ${JSON.stringify(names(shape(m)))}`); + } + }); + + test('AC2 heading-delimited: a mixed file no longer drops its non-hyphen entry', () => { + // The row that bites hardest in the wild: the file never looks empty, it + // just silently under-reports. + const got = names('### Hyphen entry\n\n- x.\n\n### Asterisk entry\n\n* y.\n'); + + assert.strictEqual(got.length, 2, JSON.stringify(got)); + assert.match(got[0], /^Hyphen entry/); + assert.match(got[1], /^Asterisk entry/); + }); + + test('AC3: a resolved-status field under any marker resolves its entry', () => { + // The lockstep property: widening what OPENS an entry without widening the + // marker STRIP feeding field extraction would surface the entry and then + // never resolve it — permanently unresolved, which is worse than dropped. + // + // Asserted through parseDeferredItemsWithStatus, NOT through an empty + // parseDeferredItems: "no outstanding item" is also what a DROPPED entry + // looks like, so the weaker form passes against the unfixed parser for + // precisely the reason under test. The entry must exist AND read resolved. + for (const m of MARKERS) { + for (const [shape, md] of [ + ['heading', `${SECTION}### Entry\n\n${m} **What:** x.\n${m} **Status:** resolved\n`], + ['headless', `${SECTION}${m} alpha\n status: resolved\n`], + ]) { + const where = `${shape} shape, marker ${JSON.stringify(m)}`; + const withStatus = parseDeferredItemsWithStatus(md); + + assert.strictEqual(withStatus.length, 1, `${where}: entry must be parsed at all — ${JSON.stringify(withStatus)}`); + assert.strictEqual(withStatus[0].status, 'resolved', `${where}: ${JSON.stringify(withStatus)}`); + assert.deepStrictEqual(parseDeferredItems(md), [], `${where}: resolved entries are not outstanding`); + } + } + }); + + test('AC3: the acknowledge writer reaches an entry written under any marker', () => { + for (const m of MARKERS) { + const content = `${SECTION}${m} alpha\n`; + const got = acknowledgeDeferredItem(content, 'alpha'); + + assert.strictEqual(got.status, 'ok', `marker ${JSON.stringify(m)}`); + assert.match(got.content, /status: acknowledged/, `marker ${JSON.stringify(m)}`); + assert.strictEqual( + parseDeferredItemsWithStatus(got.content)[0].status, + 'acknowledged', + `marker ${JSON.stringify(m)}: the written marker must parse back`, + ); + } + }); + + test('AC3: an already-acknowledged entry under any marker is not double-written', () => { + for (const m of MARKERS) { + const original = `${SECTION}${m} alpha\n`; + const once = acknowledgeDeferredItem(original, 'alpha').content; + const twice = acknowledgeDeferredItem(once, 'alpha').content; + + // Anti-vacuity: an unreachable entry is also idempotent, so pin that the + // first call actually wrote before pinning that the second did not. + assert.notStrictEqual(once, original, `marker ${JSON.stringify(m)}: first acknowledge must write`); + assert.strictEqual(twice, once, `marker ${JSON.stringify(m)}`); + } + }); + + test('AC4: prose-only and bare headings still contribute nothing', () => { + // The "prose is not an item" contract is untouched: an asterisk bullet is + // not prose, so widening the marker set cannot start counting prose. + assert.deepStrictEqual(names('### Musings\n\njust prose here.\n'), []); + assert.deepStrictEqual(names('### A bare heading with no body\n'), []); + assert.deepStrictEqual(names('### Notes\n\nwe considered * and + as options.\n'), []); + }); + + test('AC4: a bolded field key is not mistaken for an asterisk bullet', () => { + // `**Status:**` opens with `*` but supplies no whitespace after it, so the + // widened marker declines and the bolded-key path still owns the line. + const got = parseDeferredItemsWithStatus(`${SECTION}### Entry\n\n- **What:** x.\n**Status:** resolved\n`); + + assert.strictEqual(got.length, 1, JSON.stringify(got)); + assert.strictEqual(got[0].status, 'resolved', JSON.stringify(got)); + }); + + test('AC5: a table under a leaf heading still yields exactly its rows', () => { + // The anti-double-count property (#2766): table lines are skipped before + // the body-marker flag can be set, and a `|` row is not a list marker, so + // the heading still contributes no phantom entry. + const oneRow = names('### Discovered\n\n| Test | Seeds |\n|---|---|\n| test_a | 0, 1 |\n'); + assert.deepStrictEqual(oneRow, ['test_a — 0, 1'], JSON.stringify(oneRow)); + + const twoRows = names('### Discovered\n\n| Test | Seeds |\n|---|---|\n| test_a | 0 |\n| test_b | 1 |\n'); + assert.strictEqual(twoRows.length, 2, JSON.stringify(twoRows)); + }); + + test('the paren-terminated ordered marker `1)` remains out of scope', () => { + // Pinned so a later reader sees this as the ruling's scope, not a miss. + assert.deepStrictEqual(names('1) alpha\n2) beta\n'), []); + }); + + test('CRLF files: widened markers split and resolve identically', () => { + for (const m of MARKERS) { + const crlf = `## Deferred Items\r\n\r\n### Entry\r\n\r\n${m} **What:** x.\r\n${m} **Status:** resolved\r\n`; + const withStatus = parseDeferredItemsWithStatus(crlf); + + // Same anti-vacuity as AC3: an empty outstanding list would also be + // satisfied by the entry never being parsed. + assert.strictEqual(withStatus.length, 1, `marker ${JSON.stringify(m)}: ${JSON.stringify(withStatus)}`); + assert.strictEqual(withStatus[0].status, 'resolved', `marker ${JSON.stringify(m)}`); + assert.deepStrictEqual(parseDeferredItems(crlf), [], `marker ${JSON.stringify(m)}`); + } + }); + + test('nested sub-lists under any marker stay folded into their parent entry', () => { + // splitGapsEntries' indent rule (#2286) is marker-agnostic: only a marker at + // or shallower than the first one seen opens a new entry. + for (const m of MARKERS) { + const got = names(`${m} alpha\n ${m} nested one\n ${m} nested two\n${m} beta\n`); + assert.strictEqual(got.length, 2, `marker ${JSON.stringify(m)}: ${JSON.stringify(got)}`); + } + }); +}); + +describe('#3702 round 2: CRLF on the heading path and in the acknowledge writer', () => { + const MARKERS = ['-', '*', '+', '1.']; + const CRLF_SECTION = '## Deferred Items\r\n\r\n'; + + test('B1: a CRLF heading entry resolves when **Status:** is NOT the last line', () => { + // Round-1's CRLF test put `**Status:**` on the fixture's LAST line, where + // `collectSection`'s `.trimEnd()` had already removed the one `\r` that + // mattered — a false green. Every other line of a CRLF body still carries + // its `\r`, and a `$`-anchored marker strip fails on it, so the marker + // survived into field extraction and the field was silently lost. + for (const m of MARKERS) { + const statusFirst = `${CRLF_SECTION}### Entry\r\n\r\n${m} **Status:** resolved\r\n${m} **What:** x.\r\n`; + const withStatus = parseDeferredItemsWithStatus(statusFirst); + + assert.strictEqual(withStatus.length, 1, `marker ${JSON.stringify(m)}: ${JSON.stringify(withStatus)}`); + assert.strictEqual(withStatus[0].status, 'resolved', `marker ${JSON.stringify(m)}: status-first CRLF must resolve`); + assert.deepStrictEqual(parseDeferredItems(statusFirst), [], `marker ${JSON.stringify(m)}`); + + // And the mirror: a field ABOVE a trailing status line is not lost either. + const whatFirst = `${CRLF_SECTION}### Entry\r\n\r\n${m} **What:** x.\r\n${m} **Status:** resolved\r\n${m} **Why:** y.\r\n`; + assert.strictEqual(parseDeferredItemsWithStatus(whatFirst)[0].status, 'resolved', `marker ${JSON.stringify(m)}: mid-body status`); + } + }); + + test('B1: a CRLF heading entry parses byte-for-byte like its LF twin', () => { + for (const m of MARKERS) { + const body = `### Entry\n\n${m} **Status:** resolved\n${m} **What:** x.\n`; + const lf = parseDeferredItemsWithStatus(`## Deferred Items\n\n${body}`); + const crlf = parseDeferredItemsWithStatus(`${CRLF_SECTION}${body.replace(/\n/g, '\r\n')}`); + assert.deepStrictEqual(crlf, lf, `marker ${JSON.stringify(m)}`); + } + }); + + test('M4: acknowledge REWRITES an existing status line on a CRLF file (was: ok + no write)', () => { + // Pre-existing on `next`: the finder tested a CR-stripped copy, the rewrite + // ran on the raw `\r`-terminated line with a `$`-anchored regex, `replace` + // returned the input unchanged, and the writer reported `ok` over content + // that was byte-identical — the item then resurfaced on every audit. + for (const m of MARKERS) { + const content = `${CRLF_SECTION}${m} alpha\r\n status: pending\r\n${m} beta\r\n`; + const target = parseDeferredItemsWithStatus(content)[0].name; + const got = acknowledgeDeferredItem(content, target); + + assert.strictEqual(got.status, 'ok', `marker ${JSON.stringify(m)}`); + assert.notStrictEqual(got.content, content, `marker ${JSON.stringify(m)}: an ok must have written`); + assert.strictEqual(parseDeferredItemsWithStatus(got.content)[0].status, 'acknowledged', `marker ${JSON.stringify(m)}`); + } + }); + + test('m2: the inserted status line takes the entry indent on a CRLF file', () => { + for (const m of MARKERS) { + const content = `${CRLF_SECTION} ${m} alpha\r\n ${m} beta\r\n`; + const got = acknowledgeDeferredItem(content, 'alpha'); + + assert.strictEqual(got.status, 'ok', `marker ${JSON.stringify(m)}`); + assert.match(got.content, /\n {6}status: acknowledged/, `marker ${JSON.stringify(m)}: indent 4 + 2, not the indent-0 fallback`); + } + }); +}); + +describe('#3702 round 3: detect/strip symmetry on the acknowledge path (B1, B2)', () => { + // The round-3 blocker. Round 2 widened the WRITER's status-line finder to + // the deferred marker set while the reader still de-bulleted line 0 only, so + // a nested ` * status:` was selectable by the writer and invisible to the + // reader: acknowledge rewrote it, returned `ok`, and the item stayed + // outstanding on every later audit. Measured against a `next` build, `*`, + // `+` and `1.` all resolved on base and stopped resolving at round 2's head + // — a regression, not a gap in new behaviour. + for (const marker of ['-', '*', '+', '1.']) { + test(`a nested "${marker} status:" line acknowledges and READS BACK`, () => { + const doc = `## Deferred Items\n\n- alpha thing\n ${marker} status: pending\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok', `${marker}: acknowledge reported`); + assert.notStrictEqual(got.content, doc, `${marker}: content actually changed`); + const after = parseDeferredItemsWithStatus(got.content); + assert.strictEqual( + after[0].status, 'acknowledged', + `${marker}: the entry must read back as acknowledged — reporting ok over a line the reader skips is the defect`, + ); + }); + } + + // The hyphen row above is NOT a widened marker: it was already broken on + // `next`, for the same reason. One classifier cannot be right for three + // markers and wrong for the fourth, so it is fixed here rather than left as + // a pre-existing defect found while working. + test('a bare capitalised "Status:" resolves rather than reporting a write nothing reads', () => { + // The reader stores a bare key case-sensitively, so `Status:` is not + // `status:`. The writer must therefore NOT select it — it falls to the + // insert branch, which writes a line the reader does read. + const doc = '## Deferred Items\n\n- alpha\n Status: pending\n'; + const items = parseDeferredItemsWithStatus(doc); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(got.content)[0].status, 'acknowledged'); + }); + + test('a fenced "status:" line does not make the entry un-acknowledgeable', () => { + // Found by the pre-push review, and a regression THIS round introduced: + // the reader applied the fence gate before classifying and the writer did + // not, so the writer selected a fenced `status:` line the reader skips. + // The read-back guard then refused the write and the entry could not be + // acknowledged at all — `audit acknowledge` raised an internal error and + // `complete-milestone` halted. It acknowledged cleanly on `next`. + const F = '`'.repeat(3); + const doc = `## Deferred Items\n\n- alpha\n ${F}\n status: pending\n ${F}\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok', 'a fenced status line must not refuse the write'); + assert.strictEqual( + parseDeferredItemsWithStatus(got.content)[0].status, 'acknowledged', + 'the insert branch must write a line the reader reads, rather than rewriting one it skips', + ); + }); + + test('CONTROL: a bolded line-0 key keeps its ** wrapper and spelling through the rewrite', () => { + // Held before this round too — it is a control on the NEW offset-based + // rewrite, not a regression test for a reported defect. The rewrite + // replaces the VALUE at the offset the classifier reported, so the key is + // untouched by construction; a key-matching regex would have to reproduce + // the wrapper to preserve it, which is the thing that could regress. + const doc = '## Deferred Items\n\n- **Status:** pending alpha\n'; + const items = parseDeferredItemsWithStatus(doc); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok'); + assert.match(got.content, /- \*\*Status:\*\* acknowledged/); + }); + + test('acknowledge is idempotent across a re-read', () => { + // The failure this guards is the one the defect actually produced: the + // item resurfaces, gets acknowledged again, and never settles. + const doc = '## Deferred Items\n\n- alpha\n * status: pending\n'; + const first = acknowledgeDeferredItem(doc, parseDeferredItemsWithStatus(doc)[0].name); + const reread = parseDeferredItemsWithStatus(first.content); + assert.strictEqual(reread[0].status, 'acknowledged'); + const second = acknowledgeDeferredItem(first.content, reread[0].name); + assert.strictEqual(second.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(second.content)[0].status, 'acknowledged'); + }); +}); + +describe('#3702 round 4: ONE end-of-file CRLF algorithm, adopted from #3773 with its B4 closed (M1)', () => { + // Two open PRs shipped two different answers to "what line ending does an + // entry that ENDS THE FILE get?", and the maintainer's round-4 ruling was + // that the disagreement "needs one answer, not two". Neither shipped answer + // was that one. Measured, on builds of both heads, over the fixtures below: + // + // case #3739 r3 #3773 here + // undelimited single entry, CRLF preamble pass FAIL pass + // LF-dominant list, one stray CRLF at EOF FAIL pass pass + // (the other five) pass pass pass + // + // #3739's `content.endsWith('\r\n', matchIndexInContent)` reads the + // terminator of the PREVIOUS line, so it propagated an isolated CRLF into an + // LF-dominant list. #3773's `crlfAtEof` asks the right question but over a + // scope that goes EMPTY for an undelimited single-entry list, so it inserted + // a bare `\n` into a CRLF document — its own B4, and a violation of the + // uniform-CRLF invariant the fix exists to hold. What lands here is + // `crlfAtEof`'s semantics over a scope that widens instead of going empty. + // + // Every fixture below is a counterexample that killed a simpler algorithm, + // four of them ported from #3773 along with the function. They are not + // decoration: drop any one and a refuted algorithm passes again. + const endings = (text) => (text.match(/\r?\n/g) || []).map((b) => (b === '\r\n' ? 'CRLF' : 'LF')); + const assertUniformCrlf = (text) => { + assert.ok(!/[^\r]\n/.test(text) && !/^\n/.test(text), `mixed line endings: ${JSON.stringify(text)}`); + }; + + test('B4: an UNDELIMITED single-entry list at EOF inserts CRLF, not a bare LF', () => { + // #3773's B4, the defect this PR must not inherit while absorbing that PR. + // With no `## Deferred Items` heading and exactly ONE entry, the entry-list + // region runs from the first entry's start to the insertion point — and + // those are the same offset, so the region is empty and `crlfAtEof('')` is + // `false` by its own `before.length > 0` guard. The scope widens to + // everything before the insertion point rather than asserting LF from no + // evidence at all. + const content = 'preamble\r\n\r\n- alpha'; + const ack = acknowledgeDeferredItem(content, 'alpha'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, 'preamble\r\n\r\n- alpha\r\n status: acknowledged'); + assert.strictEqual(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + assertUniformCrlf(ack.content); + }); + + test('B4 control: the same shape under LF stays LF', () => { + const content = 'preamble\n\n- alpha'; + const ack = acknowledgeDeferredItem(content, 'alpha'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, 'preamble\n\n- alpha\n status: acknowledged'); + }); + + test('an LF-dominant document with ONE stray CRLF: the EOF entry does not inherit it', () => { + // Ported from #3773, and the fixture that refutes THIS PR's round-3 + // algorithm. `delta`'s CRLF terminates DELTA, not the unterminated `beta` + // after it, so copying the preceding separator propagates an isolated CRLF + // into an otherwise-LF file — strictly worse than the bare `\n` it + // replaced. Only a scope with no contradicting bare `\n` may assert CRLF. + const content = '## Deferred Items\n\n- alpha\n- gamma\n- delta\r\n- beta'; + const ack = acknowledgeDeferredItem(content, 'beta'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(ack.content)[3].status, 'acknowledged'); + assert.strictEqual(ack.content, '## Deferred Items\n\n- alpha\n- gamma\n- delta\r\n- beta\n status: acknowledged'); + assert.deepStrictEqual(endings(ack.content), ['LF', 'LF', 'LF', 'LF', 'CRLF', 'LF'], + 'the inserted break must not duplicate the unrelated CRLF above it'); + }); + + test('a DELIMITED single-entry list still has evidence: the section preamble is in scope', () => { + // Ported from #3773. The scope must not shrink to the entry list alone; a + // delimited section's own preamble belongs to that section and counts. + const content = '## Deferred Items\r\n\r\n- alpha thing'; + const ack = acknowledgeDeferredItem(content, 'alpha thing'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, '## Deferred Items\r\n\r\n- alpha thing\r\n status: acknowledged'); + assertUniformCrlf(ack.content); + const lf = '## Deferred Items\n\n- alpha thing'; + assert.strictEqual( + acknowledgeDeferredItem(lf, 'alpha thing').content, + '## Deferred Items\n\n- alpha thing\n status: acknowledged', + ); + }); + + test('a bare LF OUTSIDE the deferred section does not veto the CRLF insert', () => { + // Ported from #3773, and the fixture that rules out scanning the whole + // DOCUMENT: an unrelated bare LF inside a fenced block in another section + // would reject CRLF and drop an isolated LF into an otherwise-CRLF list. + const content = '# Notes\r\n\r\n```text\r\nfirst\nsecond\r\n```\r\n\r\n' + + '## Deferred Items\r\n\r\n- alpha\r\n- beta'; + const ack = acknowledgeDeferredItem(content, 'beta'); + assert.strictEqual(ack.status, 'ok'); + assert.match(ack.content, /- beta\r\n {2}status: acknowledged$/); + const section = ack.content.slice(ack.content.indexOf('## Deferred Items')); + assert.ok(!/(^|[^\r])\n/.test(section), `bare LF in the deferred section: ${JSON.stringify(section)}`); + assert.ok(ack.content.includes('first\nsecond'), 'the unrelated bare LF must not be rewritten'); + }); + + test('an UNDELIMITED list with entries reads only the entry list, not the preamble', () => { + // Ported from #3773, and the fixture that rules out the SECTION as the + // scope: with no heading the section body IS the whole document, so a + // section-scoped scan silently becomes the whole-document scan the case + // above already refuted. The preamble is consulted ONLY when the entry + // list is empty (the B4 case) — here it is not, so the fence's bare LF is + // out of scope and must not veto. + const content = '```text\r\nfirst\nsecond\r\n```\r\n\r\n- alpha\r\n- beta'; + const ack = acknowledgeDeferredItem(content, 'beta'); + assert.strictEqual(ack.status, 'ok'); + assert.match(ack.content, /- beta\r\n {2}status: acknowledged$/); + assert.ok(ack.content.includes('first\nsecond'), 'the unrelated bare LF must not be rewritten'); + }); + + test('no evidence under EITHER scope: an entry at offset 0 stays LF', () => { + // The widened scope terminates rather than recursing outward forever. A + // document that is exactly one unterminated entry has no line ending + // anywhere; LF is the floor, not an invented CRLF. + const ack = acknowledgeDeferredItem('- alpha', 'alpha'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, '- alpha\n status: acknowledged'); + }); + + test('mixed-ending document: the insert branch reads the ENTRY\'s ending, not the file\'s', () => { + // Ported from #3773. M1 asks for the mixed-ending fixture this PR lacked: + // every CRLF test it shipped used UNIFORM CRLF, so the motivating case + // could not fail. A CRLF heading over an LF entry — the opener's own `\n` + // must survive, and a document-wide sniff would rewrite it. + const content = '## Deferred Items\r\n\r\n- alpha\n reason: x\n'; + const ack = acknowledgeDeferredItem(content, parseDeferredItemsWithStatus(content)[0].name); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(ack.content, '## Deferred Items\r\n\r\n- alpha\n status: acknowledged\n reason: x\n'); + // The single-line twin: a CRLF entry whose only evidence is the separator + // that FOLLOWS it, inside an otherwise-LF document. + const single = '## Deferred Items\n\n- alpha\r\n'; + assert.strictEqual( + acknowledgeDeferredItem(single, parseDeferredItemsWithStatus(single)[0].name).content, + '## Deferred Items\n\n- alpha\r\n status: acknowledged\r\n', + ); + }); + + test('the widened EOF grammar carries the CRLF rule for every marker', () => { + // The EOF fixtures above are all hyphen-shaped because they were inherited + // from a hyphen-only PR. This PR's whole subject is the widened marker set, + // so the EOF rule has to hold across it or the two changes are only + // accidentally compatible. + for (const marker of ['-', '*', '+', '1.']) { + const content = `preamble\r\n\r\n${marker} alpha`; + const ack = acknowledgeDeferredItem(content, 'alpha'); + assert.strictEqual(ack.status, 'ok', `marker ${JSON.stringify(marker)}`); + assert.strictEqual(ack.content, `preamble\r\n\r\n${marker} alpha\r\n status: acknowledged`, + `marker ${JSON.stringify(marker)}: EOF insert must be CRLF`); + assertUniformCrlf(ack.content); + } + }); + + test('acknowledging twice at EOF is idempotent', () => { + const first = acknowledgeDeferredItem('preamble\r\n\r\n- alpha', 'alpha'); + const again = acknowledgeDeferredItem(first.content, parseDeferredItemsWithStatus(first.content)[0].name); + assert.strictEqual(again.status, 'ok'); + assert.strictEqual(again.content, first.content, 'a second acknowledge must not append a second line ending'); + assert.deepStrictEqual(endings(again.content), ['CRLF', 'CRLF', 'CRLF']); + }); +}); + +describe('#3702 round 4: the fence gate is indent-unbounded, like the rest of the grammar (M2)', () => { + // `scanFencedBlocks` is CommonMark, which caps a fence delimiter's indent at + // three spaces. This grammar opted out of that cliff for entry openers + // (`[ \t]*`) and thematic breaks (`^[ \t]*`) but not for fences, so a fence + // at four spaces was not a fence to the gate and a `status: resolved` inside + // it RESOLVED the entry containing it. That is reached by ordinary + // documents, not exotic ones: a fenced block written under a nested bullet + // sits at four spaces. Driven before the fix at indents 4, 5, 8 and a + // leading tab — all four silently resolved. + // + // `gsd-core/references/executor-examples.md` states flatly that "nothing + // inside a fenced code block is an entry or a field". These pin that claim + // instead of quietly bounding it. + const H = '## Deferred Items\n\n'; + const fencedStatusDoc = (pad, delim = '```') => + `${H}- alpha\n${pad}${delim}text\n${pad}status: resolved\n${pad}${delim}\n`; + + for (const [label, pad] of [ + ['0 spaces', ''], ['1 space', ' '], ['2 spaces', ' '], ['3 spaces (CommonMark cap)', ' '], + ['4 spaces (past the cap)', ' '], ['5 spaces', ' '], ['8 spaces', ' '], + ['a leading tab', '\t'], + ]) { + test(`a fenced "status: resolved" at ${label} does not resolve the entry`, () => { + const items = parseDeferredItemsWithStatus(fencedStatusDoc(pad)); + assert.strictEqual(items.length, 1, `${label}: expected exactly one entry`); + assert.notStrictEqual(String(items[0].status || '').toLowerCase(), 'resolved', + `${label}: the fenced status line must not resolve the entry`); + }); + } + + test('the same rule holds for tilde fences past the cap', () => { + const items = parseDeferredItemsWithStatus(fencedStatusDoc(' ', '~~~')); + assert.strictEqual(items.length, 1); + assert.notStrictEqual(String(items[0].status || '').toLowerCase(), 'resolved'); + }); + + test('an entry-shaped line inside a deep fence is content, not an entry', () => { + // The other half of the doc's claim: not an entry EITHER. #3702's wild + // records carry reproduction blocks whose `- ` and `1. ` lines are prose. + const doc = `${H}- alpha\n \`\`\`diff\n - not an entry\n 1. also not an entry\n \`\`\`\n- beta\n`; + const names = parseDeferredItems(doc).map((i) => i.name); + assert.ok(names.some((n) => n.startsWith('alpha')), `alpha missing: ${JSON.stringify(names)}`); + assert.ok(names.includes('beta'), `beta missing: ${JSON.stringify(names)}`); + assert.ok(!names.includes('not an entry'), `fenced content became an entry: ${JSON.stringify(names)}`); + assert.ok(!names.includes('also not an entry'), `fenced content became an entry: ${JSON.stringify(names)}`); + }); + + test('an UNTERMINATED deep fence runs to the end of its ENTRY (round 5, B1) — the status inside it stays gated, the next entry does not', () => { + // Round 4 titled this "runs to end-of-file, exactly as CommonMark says"; + // CommonMark bounds a fence by its container, and the review's B1 showed + // the end-of-file reading swallowing the entries after a stray delimiter. + // The engine still classifies; the walk bounds it at the entry. + const doc = `${H}- alpha\n \`\`\`text\n status: resolved\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + assert.notStrictEqual(String(items[0].status || '').toLowerCase(), 'resolved'); + const two = parseDeferredItemsWithStatus(`${doc}- beta\n status: resolved\n`); + assert.deepStrictEqual(two.map((i) => [i.name, i.status]), [['alpha ```text status: resolved', ''], ['beta status: resolved', 'resolved']]); + }); + + test('a deep fence still CLOSES: fields after it are read again', () => { + // The gate must not swallow the rest of the entry. If the closer at the + // same deep indent were not recognised, everything after it would stay + // fenced to EOF and the real status line would be invisible. + const doc = `${H}- alpha\n \`\`\`text\n status: resolved\n \`\`\`\n status: acknowledged\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1); + assert.strictEqual(String(items[0].status).toLowerCase(), 'acknowledged', + 'the post-fence status line must still be read'); + }); + + test('acknowledge agrees with the reader about a deep fence', () => { + // Reader and writer share one classifier; a fence the reader honours must + // be one the writer refuses to rewrite into. Otherwise acknowledge writes + // inside the code block and reports ok. + const doc = `${H}- alpha\n \`\`\`text\n status: pending\n \`\`\`\n`; + const ack = acknowledgeDeferredItem(doc, parseDeferredItemsWithStatus(doc)[0].name); + assert.strictEqual(ack.status, 'ok'); + assert.ok(ack.content.includes(' status: pending'), + 'the fenced line must be left exactly as written'); + assert.strictEqual(String(parseDeferredItemsWithStatus(ack.content)[0].status).toLowerCase(), 'acknowledged', + 'the entry must read back acknowledged through an inserted line, not a rewritten fenced one'); + }); + + test('GUARD: `## Gaps` is untouched — it opts out of block structure entirely', () => { + // Both marker-parameterised call sites gate on `markers.blockStructure`, + // which the Gaps set does not set, so Gaps reaches an EMPTY fenced set and + // this change cannot reach it. Pinned rather than asserted: a future + // "consistency" edit that dropped that gate would silently move Gaps, and + // this PR's central claim is that Gaps keeps its `next` behaviour. + // + // The observable form of "no fence gate": Gaps folds the fenced lines into + // the entry's NAME rather than hiding them, and does so identically at an + // indent inside CommonMark's cap and one past it. If the deferred gate + // ever leaked into Gaps, the deep form would start differing from the + // shallow one. + const deep = '## Gaps\n\n- alpha\n ```text\n status: open\n ```\n'; + const shallow = '## Gaps\n\n- alpha\n ```text\n status: open\n ```\n'; + assert.deepStrictEqual(parseUatItems(deep), parseUatItems(shallow), + 'Gaps must read a fence at 4 spaces exactly as it reads one at 2'); + assert.deepStrictEqual(parseUatItems(deep), [ + { name: 'alpha ```text status: open ```', result: 'open', category: 'unknown' }, + ], 'Gaps folds fenced lines into the entry name — no fence gate at any indent'); + + // And the entry-shaped twin: a `- ` line inside a deep fence is still not + // a separate Gaps entry, because Gaps never split on it to begin with. + const entryShaped = '## Gaps\n\n- alpha\n ```diff\n - fenced line\n ```\n- beta\n'; + assert.deepStrictEqual(parseUatItems(entryShaped).map((i) => i.name), + ['alpha ```diff - fenced line ```', 'beta']); + }); +}); + +describe('#3702 round 4: the minors (m1, m2, m3, m5)', () => { + const H = '## Deferred Items\n\n'; + const names = (doc) => parseDeferredItems(doc).map((i) => i.name); + + // ── m1: the ordinal digit cap needs the limit-1 case ────────────────────── + // RULESET.TESTS.boundary-coverage asks for N ∈ {limit-1, limit, limit+1}. + // The limit (`999999999.`) and limit+1 (`1234567890.`) were already pinned; + // limit-1 was the missing third. It is not a formality: an off-by-one in the + // `\d{1,9}` bound shows up at eight digits, not at nine. + test('m1: an EIGHT-digit ordinal (limit-1) is a marker', () => { + assert.deepStrictEqual(names(`${H}1. a\n12345678. b\n`), ['a', 'b']); + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('12345678. x'), true); + }); + + test('m1: the three boundary points read consistently through the marker set', () => { + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('12345678. x'), true, 'limit-1'); + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('999999999. x'), true, 'limit'); + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('1234567890. x'), false, 'limit+1'); + }); + + // ── m2: the hand-rolled CommonMark copies, checked against CommonMark ───── + // `THEMATIC_BREAK_RE` and the tab-expanding indent counter are fresh + // implementations of rules CommonMark already specifies, and this repo has + // no sectionizer helper for either to be compared against. So the parity + // assertion is against the SPEC, with the two deliberate divergences named + // rather than left for a later reader to discover and "fix" back. + test('m2: THEMATIC_BREAK_RE agrees with CommonMark on `-`, `*` and `_`', () => { + const isBreak = (line) => { + const doc = `${H}- alpha\n${line}\n- beta\n`; + // A separator closes the list, so `beta` opens a NEW list rather than + // continuing alpha's. If the line is not a break, it is swallowed as + // alpha's continuation text. + return names(doc).length === 2 && names(doc)[0] === 'alpha'; + }; + // `-- -` belongs in the YES list: CommonMark asks for three or more + // matching characters "each followed optionally by any number of spaces or + // tabs", so the spacing between them is free. It was in the NO list on the + // first cut of this test and the parser was right, not the fixture. + for (const yes of ['---', '***', '___', '- - -', '* * *', '_ _ _', '-----', ' ---', '-- -']) { + assert.strictEqual(isBreak(yes), true, `CommonMark thematic break not recognised: ${JSON.stringify(yes)}`); + } + for (const no of ['--', '**', '__', '- -', 'a---']) { + assert.strictEqual(isBreak(no), false, `not a CommonMark thematic break, but treated as one: ${JSON.stringify(no)}`); + } + }); + + test('m2: the two DELIBERATE divergences from CommonMark are pinned, not accidental', () => { + const isBreak = (line) => { + const doc = `${H}- alpha\n${line}\n- beta\n`; + return names(doc).length === 2 && names(doc)[0] === 'alpha'; + }; + // (1) `+` is NOT a CommonMark thematic-break character. It is one here, + // because `+` IS a list marker in this grammar, so `+ + +` would otherwise + // open a phantom entry named `+ +` — the same reason `* * *` is a break + // rather than an entry named `* *`. + assert.strictEqual(isBreak('+ + +'), true, + '`+ + +` must be a separator here even though CommonMark says otherwise'); + // (2) Indent is UNBOUNDED. CommonMark stops recognising a thematic break at + // four spaces (it becomes indented code); this grammar reads entries and + // breaks at any indent, so the break must follow the entries. + assert.strictEqual(isBreak(' ---'), true, 'a 7-space break must still close the list'); + assert.strictEqual(isBreak('\t---'), true, 'a tab-indented break must still close the list'); + }); + + test('m2: the indent counter expands tabs to CommonMark 4-column stops', () => { + // Not directly exported, so it is measured through its observable effect: + // a continuation line must land at or past its opener's indent column. A + // tab counted as ONE column instead of expanding to the next stop of four + // puts a 3-space continuation "outside" a tab-indented bullet. + const tabOpener = `${H}\t- alpha\n\t status: acknowledged\n`; + assert.deepStrictEqual( + parseDeferredItemsWithStatus(tabOpener).map((i) => i.status), + ['acknowledged'], + 'a tab-indented entry must read its own tab-indented field', + ); + // A 4-space opener and a tab opener occupy the same column, so the same + // continuation depth works for both. + const spaceOpener = `${H} - alpha\n status: acknowledged\n`; + assert.deepStrictEqual( + parseDeferredItemsWithStatus(spaceOpener).map((i) => i.status), + ['acknowledged'], + ); + }); + + // ── m3: the result union is declared twice; check it BEHAVIOURALLY ──────── + // `src/audit.cts` carries a hand-written structural view of `uat.cjs`, + // including its own copy of the result union. That decoupling is deliberate + // (the CLI does not take a type dependency on the module it `require()`s), + // so the fix is not to delete one copy but to make drift observable: every + // status the writer can actually PRODUCE is driven here, so a status added + // to one union and not the other shows up as a fixture with no counterpart + // rather than as a silent fall-through to the write. + // + // Four of these had no assertion anywhere in the suite before this test. + test('m3: every reachable acknowledge status is driven from a fixture', () => { + const reached = new Map(); + const drive = (label, doc, target) => { + const r = acknowledgeDeferredItem(doc, target); + reached.set(r.status, label); + return r; + }; + + drive('ok', `${H}- alpha\n`, 'alpha'); + drive('not_found', `${H}- alpha\n`, 'no such entry'); + drive('ambiguous', `${H}- alpha\n- alpha\n`, 'alpha'); + const resolvedDoc = `${H}- alpha\n status: resolved\n`; + drive('already_resolved', resolvedDoc, parseDeferredItemsWithStatus(resolvedDoc)[0].name); + // #3781 (merged round 5): the heading shape acks; the refusal is reached + // only by a span with a GFM table row INSIDE it, which no write can anchor. + const headingDoc = `${H}### Entry one\n\n- **Status:** open\n| a | b |\n- more\n`; + drive('unsupported_heading_shape', headingDoc, parseDeferredItemsWithStatus(headingDoc)[0].name); + + assert.deepStrictEqual( + [...reached.keys()].sort(), + ['already_resolved', 'ambiguous', 'not_found', 'ok', 'unsupported_heading_shape'], + 'a reachable status stopped being reachable, or a new one appeared', + ); + + // `match_verification_failed` is the sixth member and is NOT driven here. + // It is a defensive re-verify of the matched span against the target text, + // computed by an independent code path, and no fixture reaches it — the + // same position `rewrite_not_readable` was in before round 4 removed it. + // Stated rather than quietly omitted: if this union is ever pruned to what + // tests reach, that member is the one to weigh, and the answer is the same + // one round 4 gave for the other — an assertion nothing can drive is a + // specification nothing holds. + }); + + test('m3: each reachable status produces a DISTINCT outcome, so a merge of two would show', () => { + const doc = `${H}- alpha\n`; + const ok = acknowledgeDeferredItem(doc, 'alpha'); + const notFound = acknowledgeDeferredItem(doc, 'nope'); + assert.notStrictEqual(ok.content, doc, 'ok must change the content'); + assert.strictEqual(notFound.content, doc, 'a refusal must return the content untouched'); + // Every non-ok status returns the ORIGINAL content: that is the property + // the CLI relies on when it refuses rather than writes. + for (const [label, d, t] of [ + ['ambiguous', `${H}- alpha\n- alpha\n`, 'alpha'], + ['not_found', doc, 'nope'], + ]) { + const r = acknowledgeDeferredItem(d, t); + assert.strictEqual(r.content, d, `${label}: refusal must not mutate content`); + } + }); + + // ── m5: DECLINED, with the measurement that refutes it ─────────────────── + // The review is right that `markers.open`'s `(\s*)` indent group and the + // `/^[ \t]*/` reader disagree about `\f`, `\v` and NBSP. Its prescribed fix + // — "narrow to `([ \t]*)`" — was implemented, measured, and REVERTED, + // because the disagreement is not currently doing any harm and the narrowing + // is. + // + // Driven, both directions: + // indent as shipped narrowed to [ \t]* + // \f entry, status open, ack ok, reads back NO ENTRY + // \v entry, status open, ack ok, reads back NO ENTRY + // NBSP entry, status open, ack ok, reads back NO ENTRY + // + // The whole round-trip already works for these: the entry surfaces, its + // field parses, `acknowledgeDeferredItem` returns ok and the result reads + // back acknowledged. Narrowing turns three working shapes into three + // SILENTLY DROPPED ones — which is the #3702 defect class itself, and the + // opposite of the fail-safe rule this file states ("never silently drop a + // possibly-open item"). A latent inconsistency in the safe direction is not + // worth trading for a live regression in the unsafe one. + // + // Pinned here so the prescription cannot be re-applied without failing a + // test that explains why. If the inconsistency is ever to be closed, the + // direction is to make the READERS agree with the opener — `indentOf` counts + // non-tab whitespace as a column instead of terminating on it — not to make + // the opener reject lines it currently accepts. + for (const [label, indent] of [['a form feed', '\f'], ['a vertical tab', '\v'], ['an NBSP', '\u00a0']]) { + test(`m5: an entry indented with ${label} still round-trips (narrowing would drop it)`, () => { + const doc = `${H}${indent}- alpha\n status: open\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1, `${label}: the entry must surface`); + assert.strictEqual(items[0].status, 'open', `${label}: its field must parse`); + const ack = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(ack.status, 'ok', `${label}: it must be acknowledgeable`); + assert.strictEqual( + String(parseDeferredItemsWithStatus(ack.content)[0].status).toLowerCase(), + 'acknowledged', + `${label}: and the acknowledgement must read back`, + ); + }); + } + + test('m5 GUARD: `## Gaps` and the deferred set read exotic indent IDENTICALLY today', () => { + // The two sets differing here is what a narrowing would introduce. Both + // use `\s*` now, so a form-feed-indented bullet is an item on both paths. + // If a later edit narrows only one, this fails. + const gaps = parseUatItems('## Gaps\n\n\f- alpha\n').map((i) => i.name); + const deferred = parseDeferredItems('## Deferred Items\n\n\f- alpha\n').map((i) => i.name); + assert.deepStrictEqual(gaps, ['alpha'], 'Gaps must read a form-feed-indented bullet'); + assert.deepStrictEqual(deferred, ['alpha'], 'the deferred set must read it too'); + }); +}); + +describe('#3702 round 3: heading-path reader/namer inputs (m7, m8)', () => { + test('a bullet whose CONTENT is a fence opener does not suppress the entry\'s fields', () => { + // m7: the fence re-scan used to run on already-marker-stripped lines, so + // `- ```sh` — an ordinary bullet to the splitter — stripped to a fence + // opener that existed in no other pass, and the `**Status:** resolved` + // after it was suppressed as fence content. A RESOLVED entry resurfaced. + const doc = '## Deferred Items\n\n### Entry\n\n- ```sh\n- **Status:** resolved\n- ```\n'; + // The status field must be READ (the round-2 fence suppression is for a + // fence the SPLITTER saw, which this is not) ... + assert.strictEqual(parseDeferredItemsWithStatus(doc)[0].status, 'resolved'); + // ... and a resolved entry must therefore not surface as outstanding. + assert.deepStrictEqual(parseDeferredItems(doc), []); + }); + + test('a heading that begins with a list marker keeps it in the entry name', () => { + // m8: line 0 of a heading entry is the heading TEXT, not a bullet. + // Stripping a marker off it silently renamed the entry — and the name is + // the key acknowledge matches on. + for (const [heading, expected] of [ + ['### 1. Race in the writer', '1. Race in the writer'], + ['### * starred title', '* starred title'], + ['### Race in the writer', 'Race in the writer'], + ]) { + const doc = `## Deferred Items\n\n${heading}\n\n- **What:** x\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1, heading); + assert.ok(items[0].name.startsWith(expected), `${heading} -> ${items[0].name}`); + } + }); + + test('CONTROL: a body bullet keeps its marker in the entry name', () => { + // Held before this round too — a control on the opener-flag threading, not + // a reported defect. The flags say which lines CARRY a marker, but only + // line 0's is part of the entry's identity; wiring the flags into the namer + // wholesale strips the body lines too, which is a rename of the key + // acknowledge matches on. This pins that it did not happen. + const doc = '## Deferred Items\n\n### Entry\n\n- **What:** x\n'; + assert.strictEqual(parseDeferredItemsWithStatus(doc)[0].name, 'Entry - **What:** x'); + }); +}); + +describe('#3702 round 2: marker-grammar parity (M3, N1, N2)', () => { + test('every deferred-items marker regex derives from the one alternation source', () => { + // Structural, not behavioural: both splitter regexes embed the SAME + // source string, so a marker added to one cannot be absent from the other. + // Round 2 ran this over four regexes; the two writer-side ones are gone. + for (const [name, re] of [ + ['open', DEFERRED_BULLET_MARKERS.open], + ['strip', DEFERRED_BULLET_MARKERS.strip], + ]) { + assert.ok(re.source.includes(DEFERRED_MARKER_ALT), `${name}: ${re.source}`); + } + }); + + // #3702 round 3 (M6): the structural assertion above is kept for the two + // SPLITTER regexes, which really are two copies of one alternation. It is no + // longer asked to stand in for the writer/reader agreement — it never could. + // Sharing a source string says nothing about whether the line the writer + // selects is a line the reader reads, which is precisely the asymmetry that + // shipped. The replacement is BEHAVIOURAL and drives the real seam. + test('every marker that OPENS an entry also resolves it through acknowledge', () => { + for (const m of ['-', '*', '+', '1.']) { + assert.ok(DEFERRED_BULLET_MARKERS.open.test(`${m} x`), `open: ${m}`); + // The marker goes on BOTH the opener and the nested status line. Round + // 2's structural test could not reach B1; a replacement that only marks + // the opener cannot either — it is green on the defective build. + const doc = `## Deferred Items\n\n${m} alpha\n ${m} status: pending\n`; + const items = parseDeferredItemsWithStatus(doc); + assert.strictEqual(items.length, 1, `entry surfaced: ${m}`); + const got = acknowledgeDeferredItem(doc, items[0].name); + assert.strictEqual(got.status, 'ok', `ack ok: ${m}`); + const after = parseDeferredItemsWithStatus(got.content); + assert.strictEqual( + after[0].status, 'acknowledged', + `${m}: acknowledge must be READ BACK, not merely written — a status line the reader skips leaves the item outstanding forever`, + ); + } + }); + + test('parity with markdown-sectionizer iterateBullets on the shared vocabulary', () => { + // `iterateBullets` is the repo's other list-marker grammar. On everything + // both grammars are meant to agree on, they do — including the negatives. + const opens = (line) => DEFERRED_BULLET_MARKERS.open.test(line); + const sectionizerOpens = (line) => iterateBullets(line).length === 1; + const shared = [ + ['- x', true], ['* x', true], ['+ x', true], ['1. x', true], ['12. x', true], ['01. x', true], + [' - x', true], ['- [ ] x', true], ['- [x] x', true], + ['**Status:** x', false], ['1) x', false], ['-x', false], ['*x', false], ['1.x', false], + ['prose', false], ['| a | b |', false], ['2026 was a year', false], + ]; + for (const [line, expected] of shared) { + assert.strictEqual(opens(line), expected, `deferred: ${JSON.stringify(line)}`); + assert.strictEqual(sectionizerOpens(line), expected, `sectionizer: ${JSON.stringify(line)}`); + } + }); + + test('the two deliberate divergences from iterateBullets are exactly these', () => { + // N2 — a tab after the marker is CommonMark-legal; `iterateBullets` + // requires a literal space. Kept, and pinned so the difference is visible. + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('-\tx'), true); + assert.strictEqual(iterateBullets('-\tx').length, 0); + // N1 — CommonMark caps an ordered start at 9 digits; `iterateBullets` is + // `\d+`. Ten digits is not a list marker here. + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('1234567890. x'), false); + assert.strictEqual(iterateBullets('1234567890. x').length, 1); + // And `\r` is no longer whitespace after a marker (round 1's `\s` was) — + // nor are NBSP, form-feed or vertical-tab, which `\s` also accepted and + // CommonMark does not: only a space or a tab follows a marker. This is + // the assertion that fails on a `[ \t]` → `\s` revert on its own. + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test('-\r'), false); + for (const ws of ['\u00a0', '\f', '\v']) { + assert.strictEqual(DEFERRED_BULLET_MARKERS.open.test(`-${ws}x`), false, JSON.stringify(ws)); + } + }); +}); + +describe('#3702 round 2: the ordered marker and the prose contract (B2, m1)', () => { + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + + test('B2: a sentence that happens to open with `.` is prose, not an item', () => { + // Both were items on round 1 — `\d+\.` accepted any digit run. CommonMark + // §5.3's own prose/list discriminator is that an ordered list interrupting + // a paragraph must START AT 1; this parser applies that rule everywhere + // an ordered marker is seen (see `matchListOpener`). + assert.deepStrictEqual(names('2026. was a bad year for this module\n'), []); + assert.deepStrictEqual(names('### Notes\n\n3. is the number of retries we settled on.\n'), []); + // And the mixed heading case: prose under one heading, a list under another. + const got = names('### Notes\n\n3. is the number of retries.\n\n### Steps\n\n1. do this\n2. then this\n'); + assert.strictEqual(got.length, 1, JSON.stringify(got)); + assert.match(got[0], /^Steps/); + }); + + test('B2: an ordered list that starts at 1. counts, at any later number in the run', () => { + assert.deepStrictEqual(names('1. alpha\n2. beta\n3. gamma\n'), ['alpha', 'beta', 'gamma']); + // CommonMark ignores the numbers after the first — so does the run. + assert.deepStrictEqual(names('1. alpha\n3. gamma\n7. delta\n'), ['alpha', 'gamma', 'delta']); + assert.deepStrictEqual(names('01. alpha\n02. beta\n'), ['alpha', 'beta']); + // Heading shape: the run is per entry body. + assert.strictEqual(names('### Steps\n\n1. do\n2. then\n').length, 1); + // Status fields under an ordered run still resolve their entry. + assert.deepStrictEqual(names('1. alpha\n status: resolved\n2. beta\n'), ['beta']); + }); + + test('B2: the rule\'s stated cost — a run that does not start at 1 is prose', () => { + // Pinned so the trade is visible: a hand-numbered list starting at 2 is + // read as prose, the same way CommonMark refuses it as a paragraph + // interruption. The wild records (#3702) all start at 1. + assert.deepStrictEqual(names('2. alpha\n3. beta\n'), []); + // A bullet does NOT end the list for this purpose (round 5, M2): a list is + // open at the level, so the following non-1 ordinal is an item — CommonMark + // reads `2. gamma` there as a fresh ordered list (start=2), not as text. + assert.deepStrictEqual(names('1. alpha\n- beta\n2. gamma\n'), ['alpha', 'beta', 'gamma']); + }); + + test('round 5 (M1): the ordered-start threshold — `0.` and `1.` open a list, `2.` does not', () => { + // CommonMark §5.2 permits any 1-9-digit start; a `0.`-numbered list is + // ordinary. Refusing `0.` dropped ONLY the first item, because the run + // then started at `1.` — the mixed under-report that looks like a clean + // parse. Boundary: limit-1 / limit / limit+1 of the threshold itself. + assert.deepStrictEqual(names('0. alpha\n1. beta\n2. gamma\n'), ['alpha', 'beta', 'gamma']); + assert.deepStrictEqual(names('1. alpha\n2. beta\n'), ['alpha', 'beta']); + assert.deepStrictEqual(names('2. alpha\n3. beta\n'), []); + assert.deepStrictEqual(names('0. only\n'), ['only']); + assert.deepStrictEqual(names('00. alpha\n01. beta\n'), ['alpha', 'beta']); + // The cost, stated accurately: a list starting at 2 or more reads as + // prose UNTIL its first `0.`/`1.` line — the loss is the prefix. + assert.deepStrictEqual(names('2. alpha\n3. beta\n1. gamma\n'), ['gamma']); + }); + + test('round 5 (M2): a non-1 ordinal is an item wherever a list is already open at its level', () => { + // `1. a` / `- b` / `5. c` folded `5. c` into `b` on round 4 — the + // ordered-run memory was cleared by the bullet. In CommonMark `5. c` there + // is a fresh ordered list (start=5): a non-1 start is refused only where + // it would interrupt a PARAGRAPH, and after a list item it interrupts none. + assert.deepStrictEqual(names('1. a\n- b\n5. c\n'), ['a', 'b', 'c']); + assert.deepStrictEqual(names('- a\n5. c\n'), ['a', 'c']); + // Across a blank line the list is still open (CommonMark: a loose list). + assert.deepStrictEqual(names('- a\n\n5. c\n'), ['a', 'c']); + // A paragraph after the blank ENDS the list; the ordinal after it is prose + // — the round-2 B2 contract, now placed where CommonMark places it. + assert.deepStrictEqual(names('- a\n\nprose here\n5. c\n'), ['a prose here 5. c']); + // Doc start and after a heading are paragraph positions: the B2 pins hold. + assert.deepStrictEqual(names('5. c\n'), []); + assert.deepStrictEqual(names('### Notes\n\n5. c\n'), []); + }); + + test('m1: the 9-digit boundary of an ordered start', () => { + // `999999999.` is a legal CommonMark ordered marker; ten digits is not. + assert.deepStrictEqual(names('1. a\n999999999. b\n'), ['a', 'b']); + // Ten digits: not a marker at all — a lazy continuation of the open item. + assert.deepStrictEqual(names('1. a\n1234567890. b\n'), ['a 1234567890. b']); + // And a ten-digit line cannot open a run on its own. + assert.deepStrictEqual(names('1234567890. b\n'), []); + }); + + test('m1: the indentation cliff is deliberately NOT applied — indent-lenient by design', () => { + // CommonMark reads a 4-space-indented line outside a list as indented + // code. This parser does not: `deferred-items.md` is hand-written with no + // mandated shape, and surfacing a questionable entry beats dropping a real + // one (the #2766 stance). Pinned as a decision, with the 3-space twin that + // both readings agree on. + assert.deepStrictEqual(names(' - x\n'), ['x']); + assert.deepStrictEqual(names(' - x\n'), ['x']); + // Nesting still folds by the indent rule at the 2-space depth executors + // actually write, not only at the 4-space depth round 1 tested. + assert.deepStrictEqual(names('- alpha\n - nested\n- beta\n'), ['alpha - nested', 'beta']); + }); +}); + +describe('#3702 round 2: thematic breaks and fenced code are not list items (M1, M2)', () => { + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + + test('M1: a thematic break opens no entry, whichever character it is drawn with', () => { + // `- - -` was a phantom `"- -"` entry on base already; round 1 added + // `* * *` and `+ + +` to the class — and `* * *` is the separator an + // author writing in the `*` style is most likely to use. + for (const hr of ['* * *', '+ + +', '- - -', '***', '---', '___', ' * * *', '* * * ', '- - - - -']) { + assert.deepStrictEqual(names(`${hr}\n`), [], JSON.stringify(hr)); + assert.deepStrictEqual(names(`### Entry\n\n${hr}\n`), [], `heading: ${JSON.stringify(hr)}`); + } + }); + + test('M1: a thematic break ENDS the open entry rather than joining it', () => { + // CommonMark: a thematic break closes the list. The separator is neither a + // phantom item nor a continuation line of the item above it. + assert.deepStrictEqual(names('- alpha\n\n* * *\n\n- beta\n'), ['alpha', 'beta']); + assert.deepStrictEqual(names('* alpha\n* * *\n* beta\n'), ['alpha', 'beta']); + // Under a heading the break stays a BODY line (round 5, m3): the entry's + // name is what `next` reports for that body, and its span stays contiguous + // for the writer. It is still not evidence — a break alone is no entry. + assert.deepStrictEqual(names('### Entry\n\n- **What:** x.\n\n* * *\n'), ['Entry - **What:** x. * * *']); + assert.deepStrictEqual(names('### Entry\n\n* * *\n'), []); + // `- - - x` is NOT a break (trailing text); it is a `- ` item whose text is `- - x`. + assert.deepStrictEqual(names('- - - x\n'), ['- - x']); + }); + + test('M2: lines inside a fenced code block never open an entry', () => { + // #3702's wild records carry reproduction blocks — `+`-prefixed diff lines + // and `1.`-numbered repro steps are the NORMAL content of such a file. + assert.deepStrictEqual(names('### Entry\n\n```sh\n1. run this\n2. then this\n```\n'), []); + assert.deepStrictEqual(names('```diff\n+ added\n- removed\n```\n'), []); + assert.deepStrictEqual(names('~~~\n* not an item\n~~~\n'), []); + // An unterminated fence runs to the end of its ENTRY, never past it + // (round 5, B1): a stray delimiter before the first item hides nothing. + assert.deepStrictEqual(names('```\n- still fenced\n'), ['still fenced']); + }); + + test('B1 (round 5): an unterminated fence runs to the end of its entry — a stray delimiter cannot swallow later entries', () => { + // The exact review reproduction, at indent 0 and at 4: `next` reports two + // entries and round 4 reported ONE, with `- b` swallowed into `a`'s name. + assert.deepStrictEqual(names('- a\n\n```\n\n- b\n'), ['a ```', 'b']); + assert.deepStrictEqual(names('- a\n\n ```\n\n- b\n'), ['a ```', 'b']); + // A TERMINATED deep fence still gates what it encloses (round 4, M2 holds). + assert.deepStrictEqual(names('- a\n ```\n - not b\n ```\n- b\n'), ['a ``` - not b ```', 'b']); + // Inside its own entry the stray fence still gates: a `status:` under it + // does not resolve the entry, and the NEXT entry is read on its own terms. + const statusesOf = (md) => parseDeferredItemsWithStatus(SECTION + md).map((i) => i.status); + assert.deepStrictEqual(statusesOf('- alpha\n ```\n status: resolved\n- beta\n status: resolved\n'), ['', 'resolved']); + // A delimiter after the bound is a fence in its own right: without the + // rescan the `~~~` pair below would be invisible and beta would resolve. + assert.deepStrictEqual(statusesOf('- a\n```\n- b\n ~~~\n status: resolved\n ~~~\n- c\n'), ['', '', '']); + assert.deepStrictEqual(names('- a\n```\n- b\n ~~~\n status: resolved\n ~~~\n- c\n'), ['a ```', 'b ~~~ status: resolved ~~~', 'c']); + // Heading shape: a heading ends the entry, and the fence with it. (At + // indent 0 the heading TOKENIZER applies CommonMark's own fence rule, so + // `### Next` after a stray delimiter is body text there, exactly as on + // `next`; at four spaces the tokenizer sees no fence and the heading holds.) + assert.deepStrictEqual(names('### Entry\n\n- **What:** x\n\n```\n\n### Next\n\n- y\n'), ['Entry - **What:** x ``` ### Next - y']); + assert.deepStrictEqual(names('### Entry\n\n- **What:** x\n\n ```\n\n### Next\n\n- y\n'), ['Entry - **What:** x ```', 'Next - y']); + // And in the headless region before the first heading (four spaces, for + // the tokenizer reason above — at indent 0 the section has no heading). + assert.deepStrictEqual(names(' ```\n- a\n\n### Entry\n\n- b\n'), ['a', 'Entry - b']); + assert.deepStrictEqual(names('```\n- a\n\n### Entry\n\n- b\n'), ['a ### Entry', 'b']); + }); + + test('M2: a fence inside an entry is continuation, and the entry still parses around it', () => { + const md = '- alpha\n ```sh\n 1. step\n + diff\n ```\n status: resolved\n- beta\n'; + const withStatus = parseDeferredItemsWithStatus(SECTION + md); + assert.strictEqual(withStatus.length, 2, JSON.stringify(withStatus)); + assert.strictEqual(withStatus[0].status, 'resolved'); + assert.deepStrictEqual(names(md), ['beta']); + // Heading shape: the fenced lines are body text, not evidence — the `-` + // line outside the fence is what keeps the entry. + assert.strictEqual(names('### Entry\n\n- **What:** x.\n\n```\n1. repro\n```\n').length, 1); + // And the acknowledge writer's span survives a fenced continuation. + const ack = acknowledgeDeferredItem(SECTION + '- alpha\n ```\n + diff\n ```\n- beta\n', 'alpha ``` + diff ```'); + assert.strictEqual(ack.status, 'ok'); + assert.strictEqual(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + }); +}); + +describe('#3702 round 2: indent measure is grammar-scoped (review round 6)', () => { + // The deferred grammar measures CommonMark columns (a tab is a jump to the + // next multiple of 4); the Gaps grammar keeps `next`'s raw character count. + // Sharing one measure silently changed Gaps entry boundaries in BOTH + // directions on tab-indented input, breaking the `blockStructure: false` + // opt-out's byte-for-byte promise. + const gapsNames = (body) => parseUatItems(['# UAT', '', '## Gaps', '', body, ''].join('\n')).map((i) => i.name); + const deferredNames = (body) => parseDeferredItems('## Deferred Items\n\n' + body + '\n').map((i) => i.name); + + test('Gaps: a tab-indented item followed by a two-space one stays ONE entry, as on `next`', () => { + assert.deepEqual(gapsNames('\t- first item\n - second item'), ['first item - second item']); + }); + + test('Gaps: a two-space item followed by a tab-indented one stays TWO entries, as on `next`', () => { + assert.deepEqual(gapsNames(' - first item\n\t- second item'), ['first item', 'second item']); + }); + + test('Gaps: four spaces then two spaces splits, and a tab pair splits — unchanged either way', () => { + assert.deepEqual(gapsNames(' - first item\n - second item'), ['first item', 'second item']); + assert.deepEqual(gapsNames('\t- first item\n\t- second item'), ['first item', 'second item']); + }); + + test('deferred: the SAME tab/space pairs measure in columns — the opposite verdict, by design', () => { + assert.deepEqual(deferredNames('\t- first\n - second'), ['first', 'second']); + assert.deepEqual(deferredNames(' - first\n\t- second'), ['first - second']); + }); +}); + +describe('#3702 round 2: round-review refinements (ordered run, rejected ordinals, breaks, fenced fields, Gaps scope)', () => { + const SECTION = '## Deferred Items\n\n'; + const names = (md) => parseDeferredItems(SECTION + md).map((i) => i.name); + const statuses = (md) => parseDeferredItemsWithStatus(SECTION + md).map((i) => i.status); + + test('an ordered run ENDS at a paragraph that follows a blank line (CommonMark §5.3), but survives lazy continuation', () => { + // A blank line then a non-indented, non-list line is a paragraph: the list + // is over, and `5. x` after it is prose folded into the open entry. + const got = names('1. a\n\nparagraph\n\n5. x\n'); + assert.strictEqual(got.length, 1, JSON.stringify(got)); + assert.match(got[0], /^a/); + // No blank line → lazy continuation → the list is still open and `2. b` is an item. + assert.deepStrictEqual(names('1. a\nlazy continuation\n2. b\n'), ['a lazy continuation', 'b']); + // Heading shape carries the same rule per body. + assert.strictEqual(names('### Steps\n\n1. do\n\nsome prose.\n\n4. not an item\n').length, 1); + }); + + test('an accepted opener clears the blank-line memory — lazy continuation right after it keeps the run', () => { + // Round-review continuation: `blankSeen` survived the opener branch, so + // `2. b` + a lazy line ended the run and `3. c` folded into `b`. + assert.deepStrictEqual(names('1. a\n\n2. b\nlazy continuation\n3. c\n'), ['a', 'b lazy continuation', 'c']); + }); + + test('a headless region of a heading-shaped file applies the SAME paragraph reset — opener flags come from the splitter, not a re-derivation', () => { + // Round-review continuation: a re-derived flag set re-accepted `3.` under a + // stale run after the paragraph had ended it, and stripped it into a field. + const md = '1. alpha\n\nparagraph\n\n3. status: resolved\n\n### Entry\n\n- **What:** x\n'; + assert.deepStrictEqual(statuses(md), ['', '']); + assert.strictEqual(names(md).length, 2); + }); + + test('an ordered run is per INDENT: a nested `1. / 2.` run resolves (round-1 parity), and a nested ordinal after a nested bullet continues that level\'s list', () => { + // Round-review continuation 2: nested openers read the top-level run and + // never wrote their own. + for (const eol of ['\n', '\r\n']) { + const nestedRun = '- alpha\n 1. what: detail\n 2. status: resolved\n\n### Entry\n\n- **What:** x\n'.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(nestedRun), ['resolved', ''], JSON.stringify(eol)); + // Round 5 (M2): a list is open at the nested level, so `3.` is an item + // there — a fresh ordered list in CommonMark — and a nested item that + // reads `status: resolved` is a field line, as `- status: resolved` is. + const nestedAfterBullet = '1. alpha\n - nested item\n 3. status: resolved\n\n### Entry\n\n- **What:** x\n'.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(nestedAfterBullet), ['resolved', ''], JSON.stringify(eol)); + // With NO list open at the nested level the ordinal is prose (round 2). + const leak = '1. alpha\n nested prose\n 3. status: resolved\n\n### Entry\n\n- **What:** x\n'.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(leak), ['', ''], JSON.stringify(eol)); + } + // A new top-level item resets the nested levels: `2.` under beta does not continue alpha's nested run. + assert.deepStrictEqual(statuses('- alpha\n 1. a\n- beta\n 2. status: resolved\n'), ['', '']); + // Under a heading the same per-indent rule applies. + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n 1. step\n 2. **Status:** resolved\n'), ['resolved']); + }); + + test('a DEDENTING top-level list keeps its entry boundaries — every indent at or above the base is one level', () => { + // Round-review continuation 3: the exact-indent run lookup rejected the + // shallower ordinals, collapsing three entries into one. + assert.deepStrictEqual(names(' 1. alpha\n 2. beta\n3. gamma\n'), ['alpha', 'beta', 'gamma']); + assert.deepStrictEqual(names(' - alpha\n- beta\n - gamma\n'), ['alpha', 'beta - gamma']); + }); + + test('indent is measured in CommonMark COLUMNS (a tab advances to the next multiple of 4), so a tab and a space are different levels', () => { + // Round-review continuation 3: character counting aliased `\t` and ` `. + // Heading shape, where every ACCEPTED nested opener is marker-stripped + // before field extraction (the headless path strips line 0 only — #3740). + for (const eol of ['\n', '\r\n']) { + const body = (nested) => `### Entry\n\n- **What:** x\n${nested}`.replace(/\n/g, eol); + assert.deepStrictEqual(statuses(body('\t1. nested\n 2. **Status:** resolved\n')), [''], JSON.stringify(eol)); + assert.deepStrictEqual(statuses(body('\t1. nested\n\t2. **Status:** resolved\n')), ['resolved'], JSON.stringify(eol)); + assert.deepStrictEqual(statuses(body(' 1. nested\n\t2. **Status:** resolved\n')), ['resolved'], JSON.stringify(eol)); + } + }); + + test('a fenced block ends the runs at its indent and deeper, like a paragraph does', () => { + // Round-review continuation 3: a nested run stayed open across a fence, + // so a post-fence `2. status: resolved` resolved the entry. Heading shape, + // for the reason the columns test states. + const body = (nested) => `### Entry\n\n- **What:** x\n${nested}`; + assert.deepStrictEqual(statuses(body(' 1. a\n ```\n code\n ```\n 2. **Status:** resolved\n')), ['']); + // A deeper fence (3 spaces — the sectionizer's CommonMark `{0,3}` limit) leaves the shallower run alone. + assert.deepStrictEqual(statuses(body(' 1. a\n ```\n code\n ```\n 2. **Status:** resolved\n')), ['resolved']); + // Control: without the fence the run continues and resolves. + assert.deepStrictEqual(statuses(body(' 1. a\n 2. **Status:** resolved\n')), ['resolved']); + }); + + test('a REJECTED ordinal line under a heading is not marker-stripped, so it cannot manufacture a field', () => { + // `3. status: resolved` at a PARAGRAPH position is prose by the + // ordered-start rule; before this fix the heading path stripped its marker + // anyway and read a resolved field off it. (Round 5, M2: directly after a + // list item it is an item instead — a list is open there.) + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n\nSome prose.\n3. status: resolved\n'), ['']); + assert.strictEqual(names('### Entry\n\n- **What:** x\n\nSome prose.\n3. status: resolved\n').length, 1); + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n3. status: resolved\n'), ['resolved']); + // An ACCEPTED ordered status line still resolves, as `- status: resolved` does. + assert.deepStrictEqual(statuses('### Entry\n\n1. **What:** x\n2. **Status:** resolved\n'), ['resolved']); + // Same rule in a headless region of a heading-shaped file. + assert.deepStrictEqual(statuses('- alpha\n 3. status: resolved\n\n### Entry\n\n- **What:** x\n'), ['', '']); + }); + + test('a thematic break is recognised at any indent — the parser is indent-lenient for breaks as it is for items', () => { + assert.deepStrictEqual(names(' * * *\n'), []); + assert.deepStrictEqual(names('- alpha\n\n - - -\n\n- beta\n'), ['alpha', 'beta']); + }); + + test('a status line orphaned after a break leaves its entry OPEN — the fail-safe polarity, pinned', () => { + // `- alpha\n---` is a list then a thematic break in CommonMark; the indented + // line after it belongs to nothing. Surfacing alpha is the safe direction. + assert.deepStrictEqual(names('- alpha\n---\n status: resolved\n'), ['alpha']); + }); + + test('fenced lines carry no FIELDS either — a fenced `status: resolved` does not resolve the entry', () => { + assert.deepStrictEqual(statuses('- alpha\n```\nstatus: resolved\n```\n'), ['']); + assert.deepStrictEqual(statuses('### Entry\n\n- **What:** x\n```\n- **Status:** resolved\n```\n'), ['']); + assert.deepStrictEqual(statuses('- alpha\n ```yaml\n status: resolved\n ```\n status: acknowledged\n'), ['acknowledged']); + }); + + test('`## Gaps` keeps its round-1 grammar byte-for-byte: no fence or break awareness there', () => { + // Block structure (M1/M2) is scoped to the deferred grammar via + // `BulletMarkers.blockStructure`; the Gaps section is template-mandated and + // out of #3702's blast radius, so a fenced hyphen line still counts there, + // and a fenced field is still read — exactly as on `next`. + // + // THE SECOND ASSERTION tracks `next`'s #3898 fix, which landed in the + // base range this branch merged (b431ae9f0): a spaced hyphen thematic + // break in `## Gaps` is a SEPARATOR — skipped between entries — so + // `- - -` no longer surfaces a phantom open gap named `- -`. Until that + // fix this line pinned the phantom on purpose (round 4, m4), because it + // was the only assertion that would notice the Gaps path moving; it still + // is, and it now pins the fixed reading. The full separator table lives in + // the #3898 describe block above; this one keeps the Gaps opt-out honest. + const uat = ['---', 'status: partial', 'phase: 01-x', '---', '', '## Gaps', '', '```', '- truth: phantom', ' status: open', '```', ''].join('\n'); + const got = parseUatItems(uat); + assert.deepStrictEqual(got.map((i) => i.name), ['phantom'], JSON.stringify(got)); + const withBreak = ['---', 'status: partial', 'phase: 01-x', '---', '', '## Gaps', '', '- - -', '- truth: real', ' status: open', ''].join('\n'); + assert.deepStrictEqual(parseUatItems(withBreak).map((i) => i.name), ['real']); + }); +}); + // ─── Bug 3: table-shaped ## Gaps section ────────────────────────────────────── describe('#2766 parseGapsItems: GFM table shape', () => { @@ -6858,6 +8168,66 @@ describe('#3781: acknowledge supports the heading-delimited entry shape', () => assert.equal(ack.content, doc, 'file unchanged'); }); + test('a table row AFTER the entry\'s last line is outside its span — the write lands above it', () => { + // #3781 refused any leaf whose heading-to-next-heading range held a table + // row. The refusal exists because a row INSIDE the span makes it + // non-contiguous; a row after the last entry line is not inside anything, + // and refusing it halts `complete-milestone` over a write that is safe. + const doc = '## Deferred Items\n\n### Finding one\n- did a thing\n| x | y |\n'; + const before = parseDeferredItemsWithStatus(doc); + assert.equal(before.length, 2, 'fixture self-check: leaf entry + table row'); + const ack = acknowledgeDeferredItem(doc, before[0].name); + assert.equal(ack.status, 'ok'); + assert.equal(ack.content, '## Deferred Items\n\n### Finding one\n- did a thing\n status: acknowledged\n| x | y |\n'); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + }); + + test('an entry whose body ENDS in a fence acks readably — the marker lands before the fence, never inside it (round 5, RV6.5)', () => { + // Unclosed fence: it runs to the entry's end, so "after the last + // non-blank line" is fence content — the reader never sees the marker. + const unclosed = '## Deferred Items\n\n### E\n- **What:** x\n```\ncode\n'; + let ack = acknowledgeDeferredItem(unclosed, parseDeferredItemsWithStatus(unclosed)[0].name); + assert.equal(ack.status, 'ok'); + assert.equal(ack.content, '## Deferred Items\n\n### E\n- **What:** x\n status: acknowledged\n```\ncode\n'); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged', 'the marker must be read back'); + // Closed fence at the end of the body: same placement, for the same reason. + const closed = '## Deferred Items\n\n### E\n- **What:** x\n```\ncode\n```\n'; + ack = acknowledgeDeferredItem(closed, parseDeferredItemsWithStatus(closed)[0].name); + assert.equal(ack.status, 'ok'); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + assert.ok(ack.content.includes('- **What:** x\n status: acknowledged\n```'), ack.content); + // A pending (preamble) entry ending in an unclosed fence before a heading. + const pending = '## Deferred Items\n\n- a\n```\ncode\n\n### E\n- b\n'; + const items = parseDeferredItemsWithStatus(pending); + assert.equal(items.length, 2, JSON.stringify(items)); + ack = acknowledgeDeferredItem(pending, items[0].name); + assert.equal(ack.status, 'ok'); + assert.deepStrictEqual(parseDeferredItemsWithStatus(ack.content).map((i) => i.status), ['acknowledged', '']); + assert.ok(ack.content.startsWith('## Deferred Items\n\n- a\n status: acknowledged\n```\ncode\n'), ack.content); + }); + + test('a heading whose TEXT is a fence delimiter is a heading, not a fence (round 5, RV6.5)', () => { + // The entry-level fence scan saw line 0 (`\`\`\``, the heading text) as an + // opener: every body line was fenced, the reader read no field, and the + // writer's marker landed on a line nothing reads. + for (const delim of ['```', '~~~']) { + const doc = `## Deferred Items\n\n### ${delim}\n- x\n status: resolved\n`; + assert.deepStrictEqual(parseDeferredItemsWithStatus(doc).map((i) => i.status), ['resolved'], delim); + const open = `## Deferred Items\n\n### ${delim}\n- x\n`; + const ack = acknowledgeDeferredItem(open, parseDeferredItemsWithStatus(open)[0].name); + assert.equal(ack.status, 'ok', delim); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged', delim); + } + }); + + test('leaf line-0 rewrite keeps a closing `#` sequence (round 5, RV6.5)', () => { + const doc = '## Deferred Items\n\n### status: open ###\n- did a thing\n'; + const ack = acknowledgeDeferredItem(doc, parseDeferredItemsWithStatus(doc)[0].name); + assert.equal(ack.status, 'ok'); + assert.ok(ack.content.includes('### status: acknowledged ###'), ack.content); + assert.equal(parseDeferredItemsWithStatus(ack.content)[0].status, 'acknowledged'); + }); + test('leaf line-0 corner: a heading whose text parses as a status field', () => { const doc = '## Deferred Items\n\n### status: open\n- did a thing\n'; const before = parseDeferredItemsWithStatus(doc);