* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0) Establishes the canonical markdown-structure parsing seam per ADR-1372. No existing parsers are modified; this is the foundational T0 tier only. - docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the seam interface, the tiered migration plan (T0-T7), and the prohibition enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7). - src/markdown-sectionizer.cts: Pure module, Node built-ins only. Exports: stripFencedCode (CommonMark-correct state machine ported from uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence signal), tokenizeHeadings (ATX headings outside fenced blocks), collectSections (line-by-line predicate-driven section collection), collectSection (single named section, levelBounded stop, optional stripFences), iterateBullets (dash/checkbox/numbered + continuation). - tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences, unterminated fences, nested levels, all bullet markers, continuation lines, empty/non-string input) plus 4 fast-check property tests (idempotence, output shape, never-throws, length monotonicity). - CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate). Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory - src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets; add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks, tagName regex-escaped, caller decides fence-stripping) and replaceSection(content, section, newBody) (pure character-offset splice for read-modify-write callers); update collectSections/collectSection to populate bodyStart/bodyEnd. - tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a shared 9-item corpus; documents the known 4-space-indent divergence. - docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and replaceSection in §"The seam". - CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports. - docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between loop-resolver and milestone). - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of the raw stop-line offset, eliminating the trailing-newline overcounting that caused replaceSection to drop separator newlines (## A\nbody## B gluing bug). FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty ATX headings (## / ## ), text=''. 4-space indent correctly excluded. FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections that also stop at ### without abusing levelBounded. FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark). Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences unaffected. FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section doc comments. FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks the non-greedy close-at-first-</tag> behavior. FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3). Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish) The seam's compiled artifact must be a build-at-publish output like every other src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
11 KiB
ADR-1372: Canonical markdown-structure parsing — the markdown-sectionizer seam
- Status: Accepted
- Date: 2026-06-17
- Issue: #1372 (epic)
- Resolves (via tier T1): #1364, #1365
- Relates: #1343, #1324, #447 — prior single-parser markdown bugs
- Pattern precedent: ADR-857 / epic #1267 (retire a duplicated spine via tiered children)
Context
GSD parses a lot of structured markdown — CONTEXT.md, ROADMAP.md, STATE.md, *-PLAN.md, UAT files, ADRs, frontmatter. There is no shared primitive for the three operations every one of these parsers needs (strip fenced code, tokenize headings into sections, iterate bullets), so each module hand-rolls them. A grounded map of src/*.cts found:
- 8+ independent markdown parsers:
decisions,gap-checker,roadmap-parser,state,uat,uat-predicate,adr-parser,check-command-router. - 3–4 independent fenced-code strippers of different fidelity:
decisions.cts(fragile regex, no unclosed-fence handling),roadmap-parser.ctsstripFencedLines(state machine, duplicated 3× in one file),uat-predicate.cts_stripFencedBlocks(CommonMark-correct, CRLF-safe, signals an unterminated fence),check-command-router.ctsstripCommentsAndFences(another regex copy). - ~20 hand-rolled section-collects, with
state.ctsalone re-implementing the same/(###?\s*<Name>\s*\n)([\s\S]*?)(?=\n###?|$)/ishape 13 times.
The consequence is a recurring maintenance game: every "the parser missed structure X" report (#1343 bullet-before-colon, #1364 markdown-header + em-dash, #1324 glued phase tokens, #447 gap scoping) is fixed locally with another regex, and the same class of bug re-opens in the next parser. The fixes do not compound — they accrete. Worse, in the decision-coverage case the failure mode is silent: a blocking gate that cannot parse its input reports passed:true, covered 0/0 and ships the phase with its decisions unchecked.
There are two root causes, and a durable fix must address both:
- No canonical structure primitive — so structural correctness (fences, CRLF, heading levels, Unicode, bullet shapes) is re-litigated per module and tested unevenly.
- Nothing prevents the next ad-hoc parser — a new PR can add a fourth fence stripper and no gate objects, so the divergence regrows even after a cleanup.
Decision
Establish a single canonical markdown-structure seam and make ad-hoc markdown scanning a lint-enforced prohibition. Migrate every existing parser onto the seam incrementally, tracked as tiered children of epic #1372.
1. The seam — src/markdown-sectionizer.cts (pure, Node built-ins only)
No external markdown library (the "no external dependencies in core" rule stands). Pure functions, string-in → value-out, no I/O:
stripFencedCode(content) → { text, unterminatedFence }— the CommonMark-correct state machine promoted fromuat-predicate.cts_stripFencedBlocks(CRLF-safe; ≤3-space indent tolerated; closes only on a same-or-longer fence run).unterminatedFenceis a reusable malformed-input diagnostic.tokenizeHeadings(content) → HeadingToken[]— ATX headings{ level, text, line, offset }in document order.collectSections(content, stopPredicate)andcollectSection(content, headingPredicate, { levelBounded, stripFences })— line-by-line (not greedy-regex) section collection;levelBoundedencodes the dominant "stop at same-or-higher-level heading" pattern. Both populatebodyStart/bodyEndcharacter offsets on the returnedSectionfor use byreplaceSection.iterateBullets(sectionText) → BulletItem[]— dash/asterisk/plus, checkbox (- [ ]/- [x]), and numbered markers, with indented continuation-line accumulation.extractTaggedBlocks(content, tagName) → string[]— returns the inner text of every<tagName>…</tagName>block in document order;tagNameis regex-escaped; the caller decides ordering (does not strip fences). Generalisesdecisions.cts's bespoke<decisions>extractor for T1 adoption.replaceSection(content, section, newBody) → string— pure character-offset splice usingsection.bodyStart/bodyEnd; replaces a section body in a read-modify-write workflow (e.g.state.cts's 7× inlinecontent.replace(/(##\s*Name\s*\n)([\s\S]*?)(?=\n##|$)/, ...)pattern). CRLF-safe.
The seam is fully tested against the parser QA matrix (CRLF, Unicode headings, headings-inside-fences, unterminated fences, nested levels, malformed bullets) once, so every adopter inherits that correctness instead of re-deriving it.
2. Prohibition + enforcement — local/no-adhoc-markdown-parsing
A new ESLint rule in eslint-rules/no-adhoc-markdown-parsing.cjs (wired in eslint.config.mjs, mirroring local/no-source-grep) flags new hand-rolled markdown-structure scanning outside the seam — fenced-code strip regexes, split(/\r?\n/) + heading-regex section walks, and D-/checkbox bullet regexes — in src/*.cts. Existing sites are grandfathered by an explicit allowlist that is burned down as each tier migrates (the same grandfathering pattern no-source-grep uses). New code must import the seam. This is the part that stops the game permanently: after this rule lands, a PR cannot introduce a fourth fence stripper without a reviewer-visible failure.
3. Decisions realization (tier T1) — typed result + fail-loud gate
The first behavioral adopter, which also resolves the two open bugs. decisions.cts is rewritten onto the seam, and a typed result distinguishes the states the blocking gate cares about:
type DecisionExtraction = {
decisions: Decision[];
outcome: 'parsed' | 'none-present' | 'could-not-parse';
};
parseDecisions(content): Decision[] is preserved as a thin delegate (consumers untouched); extractDecisions(content): DecisionExtraction is the typed entry point. cmdDecisionCoveragePlan (blocking) treats could-not-parse — content is decision-shaped (a <decisions> block, a /decisions?/i heading, \bD- tokens, or unterminatedFence) yet 0 decisions extracted — as a WARN/fail ("could not parse decisions — possible format mismatch") instead of a green pass (resolves #1365). Routing through the seam recognises the markdown-header + em-dash variants (resolves #1364). Recall-first by design: a false "could-not-parse" is a loud warning a human clears; a false "none-present" is the silent bypass we are deleting.
4. Migration tiers (epic #1372 children)
Each tier is its own issue + PR (issue-first; one concern per PR), behaviour-preserving except T1, each separately tested, each burning down the no-adhoc-markdown-parsing grandfather list for the files it touches.
| Tier | Scope | Risk | Notes |
|---|---|---|---|
| T0 | Seam foundation: markdown-sectionizer.cts + QA-matrix tests |
none | No migration, no behavior change. Foundational. |
| T1 | decisions.cts + coverage gate: adopt seam, typed result, fail-loud |
low–med | Resolves #1364, #1365. First behavioral adopter. |
| T2 | adr-parser.cts: parseSections/splitEntries → seam |
none | CLI-only, no in-process callers — the safe prototype; its parseSections is the API shape the seam generalizes. |
| T3 | check-command-router.cts + gap-checker.cts: dedupe stripCommentsAndFences, designated-section walk, requirements bullets |
low | Gate-adjacent; covered by existing gate tests. |
| T4 | roadmap-parser.cts: collapse the 3× inline fence loop + computeSectionEnd |
med | Heavily tested; watch milestone-section boundaries. |
| T5 | uat.cts + uat-predicate.cts: donate the canonical stripper, migrate heading/section scans |
med | _stripFencedBlocks becomes the seam's source in T0; T5 removes the local copy. |
| T6 | state.cts: 13 inline section-collects → collectSection |
high | Highest payoff, highest risk — load-bearing for STATE.md mutation. Surgical, full regression, last. |
| T7 | Enforcement: no-adhoc-markdown-parsing ESLint rule + grandfather burn-down |
low | Lands once enough tiers are migrated that the grandfather list is small; thereafter new ad-hoc parsing is blocked. |
frontmatter.cts stays as-is — YAML frontmatter is a different grammar with its own well-used shared parser (extractFrontmatter); it is out of scope.
Backward compatibility
No user-facing or authoring change. Behaviour-preserving migrations (T2–T6) keep each parser's outputs byte-identical (verified by each parser's existing tests + added characterization tests). T1 is the only behavior change: additive decision recall + the could-not-parse WARN; the <decisions> block stays canonical and parses identically (block presence still takes precedence). Internal API churn is contained per-tier; public CLI contracts are unchanged.
Consequences
Positive: structural correctness (fences/CRLF/levels/bullets) is solved and tested once; the silent fail-open class is eliminated for the blocking gate; the per-module regex pile stops growing and is prohibited from regrowing; future markdown parsers inherit correctness for free; the change models the repo's own typed-IR / no-source-grep philosophy. Retires 3–4 duplicate strippers and ~20 inline section-collects.
Negative / risks: a large surface migrated incrementally — mitigated by tiering (zero-risk T2 prototype first, high-risk state.cts last, behavior-preserving with characterization tests, the epic visible end-to-end). A new shared module is a dependency for adopters — mitigated by purity + exhaustive tests. The recall-first "could-not-parse" heuristic may occasionally warn on decision-shaped-but-empty content — acceptable and tunable; a loud false alarm beats the silent miss it replaces. The enforcement rule (T7) must grandfather precisely to avoid blocking unrelated PRs mid-migration.
Alternatives considered
- Point-fix each parser bug as it's reported (status quo). Rejected — this is the game we are ending; fixes accrete instead of compounding and the same class recurs in the next parser. The maintainer's explicit directive is a solution-wide structural fix, not another file edit.
- Consolidate the primitive but skip the enforcement rule. Rejected — without the lint guard the divergence regrows; the next PR adds a fifth stripper and no gate objects. The prohibition is what makes the consolidation durable.
- External markdown library (remark/markdown-it/unified). Rejected — "no external dependencies in core" is a hard rule.
- LLM / semantic extraction. Rejected —
gsd-toolsis a deterministic, no-LLM, zero-dependency CLI with regression-tested pureResultfunctions; an LLM breaks the determinism/testability a CI gate requires and contradicts the repo's no-LLM precedent. - One big-bang PR migrating every parser. Rejected —
state.ctsalone is load-bearing and high-risk; a single PR would be unreviewable and unmergeable. Gall's Law: the working complex system is grown from a working simple seam (T0) plus incremental, individually-verified migrations.