* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0) Establishes the canonical markdown-structure parsing seam per ADR-1372. No existing parsers are modified; this is the foundational T0 tier only. - docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the seam interface, the tiered migration plan (T0-T7), and the prohibition enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7). - src/markdown-sectionizer.cts: Pure module, Node built-ins only. Exports: stripFencedCode (CommonMark-correct state machine ported from uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence signal), tokenizeHeadings (ATX headings outside fenced blocks), collectSections (line-by-line predicate-driven section collection), collectSection (single named section, levelBounded stop, optional stripFences), iterateBullets (dash/checkbox/numbered + continuation). - tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences, unterminated fences, nested levels, all bullet markers, continuation lines, empty/non-string input) plus 4 fast-check property tests (idempotence, output shape, never-throws, length monotonicity). - CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate). Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory - src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets; add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks, tagName regex-escaped, caller decides fence-stripping) and replaceSection(content, section, newBody) (pure character-offset splice for read-modify-write callers); update collectSections/collectSection to populate bodyStart/bodyEnd. - tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a shared 9-item corpus; documents the known 4-space-indent divergence. - docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and replaceSection in §"The seam". - CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports. - docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between loop-resolver and milestone). - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of the raw stop-line offset, eliminating the trailing-newline overcounting that caused replaceSection to drop separator newlines (## A\nbody## B gluing bug). FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty ATX headings (## / ## ), text=''. 4-space indent correctly excluded. FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections that also stop at ### without abusing levelBounded. FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark). Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences unaffected. FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section doc comments. FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks the non-greedy close-at-first-</tag> behavior. FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3). Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish) The seam's compiled artifact must be a build-at-publish output like every other src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
1
.gitignore
vendored
1
.gitignore
vendored
@@ -67,6 +67,7 @@ build/
|
||||
# by `npm run build:lib`). Source of truth is src/; these are emitted, never edited.
|
||||
# Published via prepublishOnly; built before test via pretest. Grows as modules migrate.
|
||||
/tsconfig.build.tsbuildinfo
|
||||
/gsd-core/bin/lib/markdown-sectionizer.cjs
|
||||
/gsd-core/bin/lib/research-store.cjs
|
||||
/gsd-core/bin/lib/research-provider.cjs
|
||||
/gsd-core/bin/lib/package-legitimacy.cjs
|
||||
|
||||
@@ -118,6 +118,9 @@ Primary installer for all runtimes. Single production file: `bin/install.js` (ge
|
||||
### I/O Module
|
||||
Module owning the tool's CLI I/O primitives: `output()` result emission (with large-payload temp-file spillover via `GSD_TEMP_DIR`/`ensureGsdTempDir`/`reapStaleTempFiles`), `error()` stderr emission with exit-code mapping, and the JSON-error-mode toggle (`setJsonErrorMode`/`getJsonErrorMode`, `ERROR_REASON`). Extracted from the Core module per ADR-857 rollout phase 1 (#859) so feature modules (`graphify`, `intel`, `audit`, `profile-pipeline`) depend on a small I/O seam instead of the core god-module; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. Source of truth: `gsd-core/bin/lib/io.cjs` (generated from `src/io.cts`).
|
||||
|
||||
### Markdown Sectionizer
|
||||
Canonical markdown-structure parsing seam (`gsd-core/bin/lib/markdown-sectionizer.cjs`, generated from `src/markdown-sectionizer.cts`). Pure functions, Node built-ins only. Exports: `stripFencedCode(content) → { text, unterminatedFence }` (CommonMark-correct state machine, CRLF-safe, signals unterminated fences); `tokenizeHeadings(content) → HeadingToken[]` (ATX headings outside fenced blocks, `{ level, text, line, offset }`); `collectSections(content, stopPredicate) → Section[]` (line-by-line section collection driven by a heading predicate); `collectSection(content, headingPredicate, { levelBounded, stripFences }) → Section | null` (single named section with level-bounded stop); `iterateBullets(sectionText) → BulletItem[]` (dash/checkbox/numbered markers with indented continuation); `extractTaggedBlocks(content, tagName) → string[]` (inner text of every `<tagName>…</tagName>` block in document order, tagName regex-escaped, caller decides fence-stripping — generalises `decisions.cts`'s bespoke extractor for T1); `replaceSection(content, section, newBody) → string` (pure character-offset splice using `Section.bodyStart`/`bodyEnd` for read-modify-write callers — eliminates T6 `state.cts`'s 7× inline `content.replace` pattern). `Section` carries `bodyStart`/`bodyEnd` offsets for `replaceSection`. ADR-1372 (epic #1372) establishes this seam and a tiered migration plan (T0–T7) to retire the 8+ ad-hoc markdown parsers and ~20 inline section-collects across `src/*.cts`. New `src/*.cts` modules must import this seam instead of hand-rolling fence strippers or heading-regex section walks (enforced by the `no-adhoc-markdown-parsing` ESLint rule landing in tier T7).
|
||||
|
||||
### Roadmap Parser Module
|
||||
Module owning ROADMAP.md parsing: shipped-milestone slicing, current-milestone extraction, milestone/phase lookups, and milestone-phase filtering (`stripShippedMilestones`, `extractCurrentMilestone`, `replaceInCurrentMilestone`, `getRoadmapPhaseInternal`, `getMilestoneInfo`, `getMilestonePhaseFilter`). Depends only on leaf modules (`phase-id`, `planning-workspace`, `shell-command-projection`) — no `loadConfig`, no other core dependency. Extracted from the Core module per ADR-857 rollout phase 2b (#870), resolving the ROADMAP.md parse/write straddle so the Roadmap module (`roadmap.cjs`, which owns ROADMAP.md mutation) imports parsing directly instead of through Core; the `core.cjs` re-export spine was retired in epic #1267, so callers import this leaf directly. Source of truth: `gsd-core/bin/lib/roadmap-parser.cjs` (generated from `src/roadmap-parser.cts`).
|
||||
|
||||
|
||||
@@ -325,6 +325,7 @@
|
||||
"legacy-cleanup.cjs",
|
||||
"loop-host-contract.cjs",
|
||||
"loop-resolver.cjs",
|
||||
"markdown-sectionizer.cjs",
|
||||
"milestone.cjs",
|
||||
"model-catalog.cjs",
|
||||
"model-profiles.cjs",
|
||||
|
||||
@@ -437,6 +437,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
|
||||
| `legacy-cleanup.cjs` | Detect and remove leftover get-shit-done-cc artifacts; exports `planLegacyCleanup` (pure scan) and `applyLegacyCleanup` (thin IO applier) that root out stale files from the old package across every GSD-managed runtime config directory (#607) |
|
||||
| `loop-host-contract.cjs` | Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts for the five-step pipeline (discuss/plan/execute/verify/ship); emitted by `scripts/gen-loop-host-contract.cjs --write` (ADR-894 §3); consumed by `gen-capability-registry.cjs` |
|
||||
| `loop-resolver.cjs` | Loop Extension Point resolver — ADR-857 phase 3c/6 registry-consuming query; given a canonical loop point, filters `byLoopPoint` by resolved Capability State plus config activation (`when` key traversal with prototype-pollution guard), returns `{ point, activeHooks, rendered }` envelope; `resolveLoopHooks` and `renderLoopHooks` are pure (no I/O); command surface: `gsd-tools loop render-hooks <point> [--config-dir <path>]` |
|
||||
| `markdown-sectionizer.cjs` | Canonical markdown-structure parsing seam (ADR-1372, epic #1372) — pure, Node built-ins only; exports `stripFencedCode` (CommonMark-correct fence stripper, CRLF-safe), `tokenizeHeadings` (ATX headings outside fenced blocks), `collectSections`/`collectSection` (line-by-line section collection with `bodyStart`/`bodyEnd` offsets), `iterateBullets` (dash/checkbox/numbered markers), `extractTaggedBlocks` (inner text of `<tag>…</tag>` blocks, caller decides fence-stripping), and `replaceSection` (pure character-offset body splice for read-modify-write callers); foundation for T0–T7 migration tiers retiring 8+ ad-hoc parsers |
|
||||
| `milestone.cjs` | Milestone archival, requirements marking |
|
||||
| `model-catalog.cjs` | CJS adapter over the shared model catalog JSON; exports canonical runtime tier defaults, agent profile maps, alias maps, and routing metadata for all CLI consumers |
|
||||
| `model-profiles.cjs` | Backward-compatible profile helpers derived from `model-catalog.cjs`; no longer owns its own model table |
|
||||
|
||||
92
docs/adr/1372-markdown-sectionizer-seam.md
Normal file
92
docs/adr/1372-markdown-sectionizer-seam.md
Normal file
@@ -0,0 +1,92 @@
|
||||
# ADR-1372: Canonical markdown-structure parsing — the `markdown-sectionizer` seam
|
||||
|
||||
- **Status:** Accepted
|
||||
- **Date:** 2026-06-17
|
||||
- **Issue:** [#1372](https://github.com/open-gsd/gsd-core/issues/1372) (epic)
|
||||
- **Resolves (via tier T1):** [#1364](https://github.com/open-gsd/gsd-core/issues/1364), [#1365](https://github.com/open-gsd/gsd-core/issues/1365)
|
||||
- **Relates:** [#1343](https://github.com/open-gsd/gsd-core/issues/1343), [#1324](https://github.com/open-gsd/gsd-core/issues/1324), [#447](https://github.com/open-gsd/gsd-core/issues/447) — prior single-parser markdown bugs
|
||||
- **Pattern precedent:** [ADR-857](857-capability-system.md) / epic [#1267](https://github.com/open-gsd/gsd-core/issues/1267) (retire a duplicated spine via tiered children)
|
||||
|
||||
## Context
|
||||
|
||||
GSD parses a lot of structured markdown — `CONTEXT.md`, `ROADMAP.md`, `STATE.md`, `*-PLAN.md`, UAT files, ADRs, frontmatter. There is **no shared primitive** for the three operations every one of these parsers needs (strip fenced code, tokenize headings into sections, iterate bullets), so each module hand-rolls them. A grounded map of `src/*.cts` found:
|
||||
|
||||
- **8+ independent markdown parsers**: `decisions`, `gap-checker`, `roadmap-parser`, `state`, `uat`, `uat-predicate`, `adr-parser`, `check-command-router`.
|
||||
- **3–4 independent fenced-code strippers of different fidelity**: `decisions.cts` (fragile regex, no unclosed-fence handling), `roadmap-parser.cts` `stripFencedLines` (state machine, **duplicated 3× in one file**), `uat-predicate.cts` `_stripFencedBlocks` (CommonMark-correct, CRLF-safe, signals an unterminated fence), `check-command-router.cts` `stripCommentsAndFences` (another regex copy).
|
||||
- **~20 hand-rolled section-collects**, with `state.cts` alone re-implementing the same `/(###?\s*<Name>\s*\n)([\s\S]*?)(?=\n###?|$)/i` shape **13 times**.
|
||||
|
||||
The consequence is a recurring maintenance game: every "the parser missed structure X" report (#1343 bullet-before-colon, #1364 markdown-header + em-dash, #1324 glued phase tokens, #447 gap scoping) is fixed *locally* with another regex, and the same class of bug re-opens in the next parser. The fixes do not compound — they accrete. Worse, in the decision-coverage case the failure mode is **silent**: a blocking gate that cannot parse its input reports `passed:true, covered 0/0` and ships the phase with its decisions unchecked.
|
||||
|
||||
There are two root causes, and a durable fix must address both:
|
||||
|
||||
1. **No canonical structure primitive** — so structural correctness (fences, CRLF, heading levels, Unicode, bullet shapes) is re-litigated per module and tested unevenly.
|
||||
2. **Nothing prevents the next ad-hoc parser** — a new PR can add a fourth fence stripper and no gate objects, so the divergence regrows even after a cleanup.
|
||||
|
||||
## Decision
|
||||
|
||||
Establish a single canonical markdown-structure seam and make ad-hoc markdown scanning a lint-enforced prohibition. Migrate every existing parser onto the seam incrementally, tracked as tiered children of epic #1372.
|
||||
|
||||
### 1. The seam — `src/markdown-sectionizer.cts` (pure, Node built-ins only)
|
||||
|
||||
No external markdown library (the "no external dependencies in core" rule stands). Pure functions, string-in → value-out, no I/O:
|
||||
|
||||
- `stripFencedCode(content) → { text, unterminatedFence }` — the CommonMark-correct state machine promoted from `uat-predicate.cts` `_stripFencedBlocks` (CRLF-safe; ≤3-space indent tolerated; closes only on a same-or-longer fence run). `unterminatedFence` is a reusable malformed-input diagnostic.
|
||||
- `tokenizeHeadings(content) → HeadingToken[]` — ATX headings `{ level, text, line, offset }` in document order.
|
||||
- `collectSections(content, stopPredicate)` and `collectSection(content, headingPredicate, { levelBounded, stripFences })` — line-by-line (not greedy-regex) section collection; `levelBounded` encodes the dominant "stop at same-or-higher-level heading" pattern. Both populate `bodyStart`/`bodyEnd` character offsets on the returned `Section` for use by `replaceSection`.
|
||||
- `iterateBullets(sectionText) → BulletItem[]` — dash/asterisk/plus, checkbox (`- [ ]`/`- [x]`), and numbered markers, with indented continuation-line accumulation.
|
||||
- `extractTaggedBlocks(content, tagName) → string[]` — returns the inner text of every `<tagName>…</tagName>` block in document order; `tagName` is regex-escaped; the caller decides ordering (does not strip fences). Generalises `decisions.cts`'s bespoke `<decisions>` extractor for T1 adoption.
|
||||
- `replaceSection(content, section, newBody) → string` — pure character-offset splice using `section.bodyStart`/`bodyEnd`; replaces a section body in a read-modify-write workflow (e.g. `state.cts`'s 7× inline `content.replace(/(##\s*Name\s*\n)([\s\S]*?)(?=\n##|$)/, ...)` pattern). CRLF-safe.
|
||||
|
||||
The seam is fully tested against the parser QA matrix (CRLF, Unicode headings, headings-inside-fences, unterminated fences, nested levels, malformed bullets) **once**, so every adopter inherits that correctness instead of re-deriving it.
|
||||
|
||||
### 2. Prohibition + enforcement — `local/no-adhoc-markdown-parsing`
|
||||
|
||||
A new ESLint rule in `eslint-rules/no-adhoc-markdown-parsing.cjs` (wired in `eslint.config.mjs`, mirroring `local/no-source-grep`) flags new hand-rolled markdown-structure scanning outside the seam — fenced-code strip regexes, `split(/\r?\n/)` + heading-regex section walks, and `D-`/checkbox bullet regexes — in `src/*.cts`. Existing sites are **grandfathered** by an explicit allowlist that is burned down as each tier migrates (the same grandfathering pattern `no-source-grep` uses). New code must import the seam. This is the part that stops the game permanently: after this rule lands, a PR cannot introduce a fourth fence stripper without a reviewer-visible failure.
|
||||
|
||||
### 3. Decisions realization (tier T1) — typed result + fail-loud gate
|
||||
|
||||
The first behavioral adopter, which also resolves the two open bugs. `decisions.cts` is rewritten onto the seam, and a typed result distinguishes the states the blocking gate cares about:
|
||||
|
||||
```
|
||||
type DecisionExtraction = {
|
||||
decisions: Decision[];
|
||||
outcome: 'parsed' | 'none-present' | 'could-not-parse';
|
||||
};
|
||||
```
|
||||
|
||||
`parseDecisions(content): Decision[]` is preserved as a thin delegate (consumers untouched); `extractDecisions(content): DecisionExtraction` is the typed entry point. `cmdDecisionCoveragePlan` (blocking) treats `could-not-parse` — content is decision-shaped (a `<decisions>` block, a `/decisions?/i` heading, `\bD-` tokens, or `unterminatedFence`) yet 0 decisions extracted — as a **WARN/fail** ("could not parse decisions — possible format mismatch") instead of a green pass (**resolves #1365**). Routing through the seam recognises the markdown-header + em-dash variants (**resolves #1364**). Recall-first by design: a false "could-not-parse" is a loud warning a human clears; a false "none-present" is the silent bypass we are deleting.
|
||||
|
||||
### 4. Migration tiers (epic #1372 children)
|
||||
|
||||
Each tier is its own issue + PR (issue-first; one concern per PR), behaviour-preserving except T1, each separately tested, each burning down the `no-adhoc-markdown-parsing` grandfather list for the files it touches.
|
||||
|
||||
| Tier | Scope | Risk | Notes |
|
||||
|---|---|---|---|
|
||||
| **T0** | Seam foundation: `markdown-sectionizer.cts` + QA-matrix tests | none | No migration, no behavior change. Foundational. |
|
||||
| **T1** | `decisions.cts` + coverage gate: adopt seam, typed result, fail-loud | low–med | **Resolves #1364, #1365.** First behavioral adopter. |
|
||||
| **T2** | `adr-parser.cts`: `parseSections`/`splitEntries` → seam | none | CLI-only, no in-process callers — the safe prototype; its `parseSections` is the API shape the seam generalizes. |
|
||||
| **T3** | `check-command-router.cts` + `gap-checker.cts`: dedupe `stripCommentsAndFences`, designated-section walk, requirements bullets | low | Gate-adjacent; covered by existing gate tests. |
|
||||
| **T4** | `roadmap-parser.cts`: collapse the 3× inline fence loop + `computeSectionEnd` | med | Heavily tested; watch milestone-section boundaries. |
|
||||
| **T5** | `uat.cts` + `uat-predicate.cts`: donate the canonical stripper, migrate heading/section scans | med | `_stripFencedBlocks` becomes the seam's source in T0; T5 removes the local copy. |
|
||||
| **T6** | `state.cts`: 13 inline section-collects → `collectSection` | high | Highest payoff, highest risk — load-bearing for STATE.md mutation. Surgical, full regression, last. |
|
||||
| **T7** | Enforcement: `no-adhoc-markdown-parsing` ESLint rule + grandfather burn-down | low | Lands once enough tiers are migrated that the grandfather list is small; thereafter new ad-hoc parsing is blocked. |
|
||||
|
||||
`frontmatter.cts` stays as-is — YAML frontmatter is a different grammar with its own well-used shared parser (`extractFrontmatter`); it is out of scope.
|
||||
|
||||
## Backward compatibility
|
||||
|
||||
No user-facing or authoring change. Behaviour-preserving migrations (T2–T6) keep each parser's outputs byte-identical (verified by each parser's existing tests + added characterization tests). T1 is the only behavior change: additive decision recall + the could-not-parse WARN; the `<decisions>` block stays canonical and parses identically (block presence still takes precedence). Internal API churn is contained per-tier; public CLI contracts are unchanged.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive:** structural correctness (fences/CRLF/levels/bullets) is solved and tested once; the silent fail-open class is eliminated for the blocking gate; the per-module regex pile stops growing *and* is prohibited from regrowing; future markdown parsers inherit correctness for free; the change models the repo's own typed-IR / no-source-grep philosophy. Retires 3–4 duplicate strippers and ~20 inline section-collects.
|
||||
|
||||
**Negative / risks:** a large surface migrated incrementally — mitigated by tiering (zero-risk T2 prototype first, high-risk `state.cts` last, behavior-preserving with characterization tests, the epic visible end-to-end). A new shared module is a dependency for adopters — mitigated by purity + exhaustive tests. The recall-first "could-not-parse" heuristic may occasionally warn on decision-shaped-but-empty content — acceptable and tunable; a loud false alarm beats the silent miss it replaces. The enforcement rule (T7) must grandfather precisely to avoid blocking unrelated PRs mid-migration.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Point-fix each parser bug as it's reported (status quo).** Rejected — this is the game we are ending; fixes accrete instead of compounding and the same class recurs in the next parser. The maintainer's explicit directive is a solution-wide structural fix, not another file edit.
|
||||
- **Consolidate the primitive but skip the enforcement rule.** Rejected — without the lint guard the divergence regrows; the next PR adds a fifth stripper and no gate objects. The prohibition is what makes the consolidation durable.
|
||||
- **External markdown library (remark/markdown-it/unified).** Rejected — "no external dependencies in core" is a hard rule.
|
||||
- **LLM / semantic extraction.** Rejected — `gsd-tools` is a deterministic, no-LLM, zero-dependency CLI with regression-tested pure `Result` functions; an LLM breaks the determinism/testability a CI gate requires and contradicts the repo's no-LLM precedent.
|
||||
- **One big-bang PR migrating every parser.** Rejected — `state.cts` alone is load-bearing and high-risk; a single PR would be unreviewable and unmergeable. Gall's Law: the working complex system is grown from a working simple seam (T0) plus incremental, individually-verified migrations.
|
||||
@@ -160,6 +160,8 @@ export default tseslint.config(
|
||||
'gsd-core/bin/lib/capability-writer.cjs',
|
||||
// issue #1355: tsc-generated runtime artifact — lint the src/teams-status.cts source.
|
||||
'gsd-core/bin/lib/teams-status.cjs',
|
||||
// ADR-1372: tsc-generated runtime artifact — lint the src/markdown-sectionizer.cts source.
|
||||
'gsd-core/bin/lib/markdown-sectionizer.cjs',
|
||||
],
|
||||
},
|
||||
|
||||
|
||||
585
src/markdown-sectionizer.cts
Normal file
585
src/markdown-sectionizer.cts
Normal file
@@ -0,0 +1,585 @@
|
||||
/**
|
||||
* Markdown Sectionizer — canonical markdown-structure parsing seam
|
||||
*
|
||||
* Pure functions, Node built-ins only (no external deps). String-in → value-out, no I/O.
|
||||
* Promoted from `uat-predicate.cts` `_stripFencedBlocks` (CommonMark-correct state machine)
|
||||
* and extended with heading tokenisation, section collection, and bullet iteration.
|
||||
*
|
||||
* ADR-1372 — T0 foundational seam. Migration tiers T1–T7 progressively adopt this seam.
|
||||
*
|
||||
* ADR-457 build-at-publish: compiled by tsc to gsd-core/bin/lib/markdown-sectionizer.cjs.
|
||||
*/
|
||||
|
||||
// ─── Types ────────────────────────────────────────────────────────────────────
|
||||
|
||||
/** Result of stripping fenced code blocks from markdown content. */
|
||||
export interface StripFencedResult {
|
||||
/** Content with all fenced code blocks removed (delimiters and body lines). */
|
||||
text: string;
|
||||
/**
|
||||
* True when the input contained an unterminated fence (EOF inside a fence).
|
||||
* Callers that wish to signal malformed input to the user should inspect this.
|
||||
*/
|
||||
unterminatedFence: boolean;
|
||||
}
|
||||
|
||||
/** An ATX heading extracted by `tokenizeHeadings`. */
|
||||
export interface HeadingToken {
|
||||
/** Heading depth: 1 = `#`, 2 = `##`, 3 = `###`, etc. */
|
||||
level: number;
|
||||
/** Heading text with surrounding whitespace trimmed. */
|
||||
text: string;
|
||||
/** 1-based line number of the heading in the original content. */
|
||||
line: number;
|
||||
/** Character (string-index) offset of the `#` character in the original content string. */
|
||||
offset: number;
|
||||
}
|
||||
|
||||
/** A collected markdown section (heading + body). */
|
||||
export interface Section {
|
||||
/** The heading that opened this section. */
|
||||
heading: HeadingToken;
|
||||
/** All lines between this heading and the next stop, joined by `\n`. */
|
||||
body: string;
|
||||
/**
|
||||
* Character (string-index) offset in the ORIGINAL content string where the
|
||||
* section body begins (first character after the heading line's trailing newline).
|
||||
* Populated by `collectSections` and `collectSection`.
|
||||
* Used by `replaceSection` for a clean pure splice.
|
||||
*
|
||||
* INVARIANT: `content.slice(bodyStart, bodyEnd) === body` for every Section
|
||||
* returned by `collectSection` and `collectSections`.
|
||||
*/
|
||||
bodyStart: number;
|
||||
/**
|
||||
* Character (string-index) offset in the ORIGINAL content string where the
|
||||
* section body ends (exclusive). Because `body` is `trimEnd()`-ed, this equals
|
||||
* `bodyStart + body.length` — NOT the start of the next heading line.
|
||||
*
|
||||
* INVARIANT: `content.slice(bodyStart, bodyEnd) === body`.
|
||||
* This guarantees `replaceSection(content, section, section.body) === content`.
|
||||
*/
|
||||
bodyEnd: number;
|
||||
}
|
||||
|
||||
/** Recognised bullet markers. */
|
||||
export type BulletMarker = 'dash' | 'checkbox-unchecked' | 'checkbox-checked' | 'numbered';
|
||||
|
||||
/** A single bullet item from `iterateBullets`. */
|
||||
export interface BulletItem {
|
||||
/** Which marker shape was recognised. */
|
||||
marker: BulletMarker;
|
||||
/** Full bullet text including all indented continuation lines, whitespace-trimmed. */
|
||||
text: string;
|
||||
/** Raw indentation prefix of the opening bullet line. */
|
||||
indent: string;
|
||||
/** Checkbox state — `true` for `[x]`, `false` for `[ ]`, `null` for non-checkbox. */
|
||||
checked: boolean | null;
|
||||
}
|
||||
|
||||
// ─── Internal types ───────────────────────────────────────────────────────────
|
||||
|
||||
interface FenceState {
|
||||
char: '`' | '~';
|
||||
len: number;
|
||||
}
|
||||
|
||||
// ─── stripFencedCode ──────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* CommonMark-correct fenced-code-block stripper.
|
||||
*
|
||||
* Ported from `uat-predicate.cts` `_stripFencedBlocks` — the reference
|
||||
* implementation for the repo. DO NOT modify `uat-predicate.cts` (its
|
||||
* migration is T5); this is a tracked duplication until T5 lands.
|
||||
*
|
||||
* Rules:
|
||||
* - Opening delimiter: a line whose non-indent portion begins with ≥3 backticks
|
||||
* or tildes (≤3 leading spaces tolerated per CommonMark §4.5).
|
||||
* - Closing delimiter: same character, run length ≥ opening, no trailing
|
||||
* non-whitespace text.
|
||||
* - A tilde fence inside a backtick fence (or vice versa) is fence *content*,
|
||||
* not a closing delimiter — delimiter char must match.
|
||||
* - Both delimiter lines and all content lines are dropped from the output.
|
||||
* - CRLF-safe: trailing `\r` is stripped before delimiter matching; the kept
|
||||
* non-fence lines are returned as-is (including any `\r`).
|
||||
* - `unterminatedFence` signals EOF inside an open fence.
|
||||
*/
|
||||
export function stripFencedCode(content: string): StripFencedResult {
|
||||
if (typeof content !== 'string') {
|
||||
return { text: '', unterminatedFence: false };
|
||||
}
|
||||
const lines = content.split('\n');
|
||||
const kept: string[] = [];
|
||||
let openFence: FenceState | null = null;
|
||||
|
||||
// Matches: optional indent (≤3 spaces per CommonMark), fence run, optional info string
|
||||
const delimRe = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||||
|
||||
for (const rawLine of lines) {
|
||||
// Strip trailing \r for delimiter matching (CRLF safety)
|
||||
const line = rawLine.replace(/\r$/, '');
|
||||
const m = delimRe.exec(line);
|
||||
if (m) {
|
||||
const char = m[2][0] as '`' | '~';
|
||||
const len = m[2].length;
|
||||
const trailing = m[3];
|
||||
if (openFence === null) {
|
||||
// CommonMark §4.5: backtick fence info string must not contain a backtick.
|
||||
// If it does, this line is NOT a valid fence opener (treat as ordinary content).
|
||||
if (char === '`' && trailing.includes('`')) {
|
||||
kept.push(rawLine);
|
||||
continue;
|
||||
}
|
||||
// Opening delimiter — record fence state, drop this line
|
||||
openFence = { char, len };
|
||||
} else if (char === openFence.char && len >= openFence.len && /^\s*$/.test(trailing)) {
|
||||
// Closing delimiter (same char, sufficient length, no trailing content) — close and drop
|
||||
openFence = null;
|
||||
}
|
||||
// else: mismatched delimiter inside fence — treat as content, still drop (it's a fence line)
|
||||
continue; // all delimiter lines are dropped
|
||||
}
|
||||
|
||||
if (openFence === null) {
|
||||
kept.push(rawLine); // non-fence content: keep as-is (preserve original \r if any)
|
||||
}
|
||||
// Lines inside a fence are silently dropped
|
||||
}
|
||||
|
||||
return { text: kept.join('\n'), unterminatedFence: openFence !== null };
|
||||
}
|
||||
|
||||
// ─── tokenizeHeadings ─────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Extract all ATX headings from `content` in document order.
|
||||
*
|
||||
* Only headings OUTSIDE fenced code blocks are returned — `stripFencedCode` is
|
||||
* applied first so that a `## heading` inside a ``` fence is not tokenised.
|
||||
*
|
||||
* Each token records `{ level, text, line, offset }` where `offset` is relative
|
||||
* to the ORIGINAL `content` (before fence-stripping), enabling callers to use
|
||||
* `collectSection` on the original string.
|
||||
*/
|
||||
export function tokenizeHeadings(content: string): HeadingToken[] {
|
||||
if (typeof content !== 'string' || content.length === 0) return [];
|
||||
|
||||
// Strip fences first so headings inside code blocks are ignored.
|
||||
// We need the original line positions, so we map stripped-text line numbers
|
||||
// back to original by tracking which original lines survived stripping.
|
||||
const originalLines = content.split('\n');
|
||||
const tokens: HeadingToken[] = [];
|
||||
|
||||
// We re-run the fence state machine to know which lines are "kept", so we
|
||||
// can map line index in original to whether it survived.
|
||||
const delimRe = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||||
let openFence: FenceState | null = null;
|
||||
|
||||
// Accumulate character offset as we iterate lines
|
||||
let charOffset = 0;
|
||||
|
||||
for (let i = 0; i < originalLines.length; i++) {
|
||||
const rawLine = originalLines[i];
|
||||
const line = rawLine.replace(/\r$/, '');
|
||||
|
||||
const dm = delimRe.exec(line);
|
||||
if (dm) {
|
||||
const char = dm[2][0] as '`' | '~';
|
||||
const len = dm[2].length;
|
||||
const trailing = dm[3];
|
||||
if (openFence === null) {
|
||||
// CommonMark §4.5: backtick fence info string must not contain a backtick.
|
||||
if (char === '`' && trailing.includes('`')) {
|
||||
// Not a valid fence opener — check for heading on this line (will fall through)
|
||||
} else {
|
||||
openFence = { char, len };
|
||||
charOffset += rawLine.length + 1;
|
||||
continue;
|
||||
}
|
||||
} else if (char === openFence.char && len >= openFence.len && /^\s*$/.test(trailing)) {
|
||||
openFence = null;
|
||||
charOffset += rawLine.length + 1;
|
||||
continue;
|
||||
} else {
|
||||
// Mismatched/invalid delimiter inside fence — treat as content (still inside fence), skip heading check
|
||||
charOffset += rawLine.length + 1;
|
||||
continue;
|
||||
}
|
||||
}
|
||||
|
||||
if (openFence === null) {
|
||||
// This line is outside any fence — check for ATX heading.
|
||||
// CommonMark: ≤3 leading spaces, then 1–6 `#`, then either EOF (empty heading)
|
||||
// or at least one space/tab followed by optional text, with optional closing `#` sequence.
|
||||
const headingMatch = /^( {0,3})(#{1,6})([ \t]+.*|[ \t]*)?$/.exec(line);
|
||||
if (headingMatch) {
|
||||
const hashes = headingMatch[2];
|
||||
const rest = headingMatch[3] ?? '';
|
||||
// Strip optional closing `#` sequence: trailing whitespace + one or more `#` + optional whitespace
|
||||
const rawText = rest.replace(/^[ \t]+/, '').replace(/[ \t]+#+[ \t]*$/, '').replace(/^#+[ \t]*$/, '');
|
||||
tokens.push({
|
||||
level: hashes.length,
|
||||
text: rawText.trim(),
|
||||
line: i + 1, // 1-based
|
||||
offset: charOffset,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
charOffset += rawLine.length + 1;
|
||||
}
|
||||
|
||||
return tokens;
|
||||
}
|
||||
|
||||
// ─── collectSections ─────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Collect sections from `content`, calling `stopPredicate` on each heading to
|
||||
* decide where sections end.
|
||||
*
|
||||
* Returns an array of `Section` objects, one per matched heading. The `body`
|
||||
* of each section runs from the line after the heading up to (but not
|
||||
* including) the next heading that satisfies `stopPredicate`, or EOF.
|
||||
*
|
||||
* Unlike a greedy-regex approach, this is a line-by-line walk — compatible
|
||||
* with the repo's "line-by-line section collection" pattern.
|
||||
*/
|
||||
export function collectSections(
|
||||
content: string,
|
||||
stopPredicate: (heading: HeadingToken) => boolean,
|
||||
): Section[] {
|
||||
if (typeof content !== 'string' || content.length === 0) return [];
|
||||
|
||||
const headings = tokenizeHeadings(content);
|
||||
if (headings.length === 0) return [];
|
||||
|
||||
const lines = content.split('\n');
|
||||
const sections: Section[] = [];
|
||||
|
||||
// Build a set of line numbers (1-based) that are heading lines
|
||||
const headingsByLine = new Map<number, HeadingToken>();
|
||||
for (const h of headings) {
|
||||
headingsByLine.set(h.line, h);
|
||||
}
|
||||
|
||||
// Build a byte-offset table: lineOffsets[i] = byte offset of the start of line i+1 (1-based: i=0 → line 1)
|
||||
// The body of a section starts at the byte after the heading line's trailing '\n'.
|
||||
const lineOffsets: number[] = new Array<number>(lines.length);
|
||||
let acc = 0;
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
lineOffsets[i] = acc;
|
||||
acc += lines[i].length + 1; // +1 for the '\n' we split on
|
||||
}
|
||||
// lineOffsets[i] is the byte offset of line (i+1) (1-based). EOF sentinel:
|
||||
const eofOffset = acc; // === content.length + (content.endsWith('\n') ? 0 : 0) ≈ content.length
|
||||
|
||||
let currentHeading: HeadingToken | null = null;
|
||||
let currentBodyStart = 0;
|
||||
let bodyLines: string[] = [];
|
||||
|
||||
const flush = (_bodyEndOffset: number): void => {
|
||||
if (currentHeading !== null) {
|
||||
const rawBody = bodyLines.join('\n');
|
||||
const body = rawBody.trimEnd();
|
||||
// INVARIANT: content.slice(bodyStart, bodyEnd) === body
|
||||
// bodyEnd is derived from body.length, NOT from the raw separator offset,
|
||||
// so round-trips via replaceSection(content, section, section.body) are exact.
|
||||
sections.push({
|
||||
heading: currentHeading,
|
||||
body,
|
||||
bodyStart: currentBodyStart,
|
||||
bodyEnd: currentBodyStart + body.length,
|
||||
});
|
||||
currentHeading = null;
|
||||
bodyLines = [];
|
||||
}
|
||||
};
|
||||
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const lineNo = i + 1; // 1-based
|
||||
const h = headingsByLine.get(lineNo);
|
||||
if (h !== undefined && stopPredicate(h)) {
|
||||
// This heading is a stop boundary — flush current section, start new one.
|
||||
// The body ends at the start of this heading line.
|
||||
flush(lineOffsets[i]);
|
||||
currentHeading = h;
|
||||
// Body starts at the beginning of the line AFTER the heading line
|
||||
const headingLineIdx = h.line - 1; // 0-based
|
||||
currentBodyStart = lineOffsets[headingLineIdx] + lines[headingLineIdx].length + 1;
|
||||
} else if (currentHeading !== null) {
|
||||
bodyLines.push(lines[i]);
|
||||
}
|
||||
}
|
||||
flush(eofOffset);
|
||||
|
||||
return sections;
|
||||
}
|
||||
|
||||
// ─── collectSection ───────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Collect a single section whose heading satisfies `headingPredicate`.
|
||||
*
|
||||
* Options:
|
||||
* - `levelBounded` (default: `true`): the section ends at the next heading of
|
||||
* the same or higher level (lower level number = higher in the hierarchy).
|
||||
* When `false`, the section body runs until any heading or EOF.
|
||||
* Ignored when `stopAtLevel` is provided.
|
||||
* - `stopAtLevel` (optional): when provided, the section ends at the next heading
|
||||
* whose `level <= stopAtLevel`, regardless of the opener's level. This enables
|
||||
* modeling sections like a `##`-opened section that also stops at `###`
|
||||
* (pass `stopAtLevel: 3`). Takes precedence over `levelBounded` when set.
|
||||
* - `stripFences` (default: `false`): apply `stripFencedCode` to the body
|
||||
* before returning. The `heading` in the result always refers to the original
|
||||
* heading (pre-strip).
|
||||
*
|
||||
* Returns `null` when no matching heading is found.
|
||||
*/
|
||||
export function collectSection(
|
||||
content: string,
|
||||
headingPredicate: (heading: HeadingToken) => boolean,
|
||||
opts: { levelBounded?: boolean; stopAtLevel?: number; stripFences?: boolean } = {},
|
||||
): Section | null {
|
||||
if (typeof content !== 'string' || content.length === 0) return null;
|
||||
|
||||
const { levelBounded = true, stopAtLevel, stripFences = false } = opts;
|
||||
|
||||
const headings = tokenizeHeadings(content);
|
||||
const targetIdx = headings.findIndex(headingPredicate);
|
||||
if (targetIdx === -1) return null;
|
||||
|
||||
const target = headings[targetIdx];
|
||||
const lines = content.split('\n');
|
||||
|
||||
// Determine which headings act as stops after the target
|
||||
const bodyStartLine = target.line + 1; // 1-based, first line of body
|
||||
let bodyEndLine = lines.length + 1; // 1-based, exclusive (default: EOF+1)
|
||||
|
||||
for (let j = targetIdx + 1; j < headings.length; j++) {
|
||||
const next = headings[j];
|
||||
let isStop: boolean;
|
||||
if (stopAtLevel !== undefined) {
|
||||
// stopAtLevel: stop at the next heading whose level <= stopAtLevel
|
||||
isStop = next.level <= stopAtLevel;
|
||||
} else {
|
||||
isStop = levelBounded ? next.level <= target.level : true;
|
||||
}
|
||||
if (isStop) {
|
||||
bodyEndLine = next.line; // stop before this line (1-based)
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
// Compute character offsets for bodyStart.
|
||||
// lineOffsets[i] = character offset of line (i+1) in content (1-based).
|
||||
const lineOffsets: number[] = new Array<number>(lines.length);
|
||||
let acc = 0;
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
lineOffsets[i] = acc;
|
||||
acc += lines[i].length + 1; // +1 for the '\n' separator
|
||||
}
|
||||
const eofOffset = acc; // byte offset past the last line
|
||||
|
||||
// bodyStart: character offset of first line of body (bodyStartLine is 1-based)
|
||||
const bodyStartOffset = bodyStartLine <= lines.length ? lineOffsets[bodyStartLine - 1] : eofOffset;
|
||||
|
||||
// Slice body lines (0-based array: bodyStartLine-1 to bodyEndLine-2 inclusive)
|
||||
const bodyRaw = lines.slice(bodyStartLine - 1, bodyEndLine - 1).join('\n').trimEnd();
|
||||
const body = stripFences ? stripFencedCode(bodyRaw).text : bodyRaw;
|
||||
|
||||
// INVARIANT: content.slice(bodyStart, bodyEnd) === body
|
||||
// bodyEnd is derived from body.length so that replaceSection(content, section, section.body) === content.
|
||||
return { heading: target, body, bodyStart: bodyStartOffset, bodyEnd: bodyStartOffset + body.length };
|
||||
}
|
||||
|
||||
// ─── iterateBullets ───────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Extract bullet items from `sectionText`.
|
||||
*
|
||||
* Recognises three marker families:
|
||||
* - **Checkbox**: `- [ ] text` (unchecked) and `- [x] text` / `- [X] text` (checked)
|
||||
* - **Dash**: `- text`, `* text`, `+ text` (plain unordered list item)
|
||||
* - **Numbered**: `1. text`, `42. text` (ordered list item)
|
||||
*
|
||||
* Indented continuation lines (lines that are not themselves bullet openers and
|
||||
* have at least one leading space or tab) are accumulated into the current
|
||||
* bullet's `text`.
|
||||
*
|
||||
* Blank lines terminate the current bullet (consistent with CommonMark block
|
||||
* handling and the repo's existing bullet parsers).
|
||||
*/
|
||||
export function iterateBullets(sectionText: string): BulletItem[] {
|
||||
if (typeof sectionText !== 'string' || sectionText.length === 0) return [];
|
||||
|
||||
const lines = sectionText.split('\n');
|
||||
const items: BulletItem[] = [];
|
||||
|
||||
// Checkbox bullet: `<indent>- [ ] text` or `<indent>- [x] text`
|
||||
const checkboxRe = /^(\s*)- \[([xX ])\] (.*)$/;
|
||||
// Plain dash/asterisk/plus bullet: `<indent>- text`, `<indent>* text`, `<indent>+ text`
|
||||
const dashRe = /^(\s*)[-*+] (.*)$/;
|
||||
// Numbered bullet: `<indent>1. text`
|
||||
const numberedRe = /^(\s*)\d+\. (.*)$/;
|
||||
// Continuation: non-empty, indented, NOT a bullet opener
|
||||
const continuationRe = /^[ \t]/;
|
||||
|
||||
let current: BulletItem | null = null;
|
||||
|
||||
const flush = (): void => {
|
||||
if (current !== null) {
|
||||
current.text = current.text.trim();
|
||||
items.push(current);
|
||||
current = null;
|
||||
}
|
||||
};
|
||||
|
||||
for (const rawLine of lines) {
|
||||
// Strip trailing \r (CRLF safety)
|
||||
const line = rawLine.replace(/\r$/, '');
|
||||
const trimmed = line.trim();
|
||||
|
||||
// Blank line terminates current bullet
|
||||
if (trimmed === '') {
|
||||
flush();
|
||||
continue;
|
||||
}
|
||||
|
||||
// Checkbox bullet (checked or unchecked) — must test before dashRe
|
||||
const cbm = checkboxRe.exec(line);
|
||||
if (cbm) {
|
||||
flush();
|
||||
const stateChar = cbm[2];
|
||||
const checked = stateChar === 'x' || stateChar === 'X';
|
||||
current = {
|
||||
marker: checked ? 'checkbox-checked' : 'checkbox-unchecked',
|
||||
text: cbm[3],
|
||||
indent: cbm[1],
|
||||
checked,
|
||||
};
|
||||
continue;
|
||||
}
|
||||
|
||||
// Numbered bullet
|
||||
const nm = numberedRe.exec(line);
|
||||
if (nm) {
|
||||
flush();
|
||||
current = {
|
||||
marker: 'numbered',
|
||||
text: nm[2],
|
||||
indent: nm[1],
|
||||
checked: null,
|
||||
};
|
||||
continue;
|
||||
}
|
||||
|
||||
// Plain dash / asterisk / plus bullet
|
||||
const dm = dashRe.exec(line);
|
||||
if (dm) {
|
||||
flush();
|
||||
current = {
|
||||
marker: 'dash',
|
||||
text: dm[2],
|
||||
indent: dm[1],
|
||||
checked: null,
|
||||
};
|
||||
continue;
|
||||
}
|
||||
|
||||
// Continuation line (indented, non-bullet) — append to current bullet
|
||||
if (current !== null && continuationRe.test(line)) {
|
||||
current.text += ' ' + trimmed;
|
||||
continue;
|
||||
}
|
||||
|
||||
// Non-bullet, non-continuation line (e.g. a paragraph, heading) — flush
|
||||
flush();
|
||||
}
|
||||
flush();
|
||||
|
||||
return items;
|
||||
}
|
||||
|
||||
// ─── extractTaggedBlocks ──────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Return the inner text of every `<tagName>…</tagName>` block in `content`,
|
||||
* in document order.
|
||||
*
|
||||
* Designed for extracting structured XML-like annotation blocks that live in
|
||||
* markdown prose (e.g. `<decisions>…</decisions>`, `<requirements>…</requirements>`).
|
||||
* Returns `[]` when no matching blocks are found.
|
||||
*
|
||||
* The `tagName` argument is regex-escaped, so names that contain regex
|
||||
* metacharacters (e.g. `foo.bar`, `my+tag`) are matched literally.
|
||||
*
|
||||
* **Input contract:** the caller decides whether to pass raw or fence-stripped
|
||||
* content. `extractTaggedBlocks` is a pure block extractor — it does NOT strip
|
||||
* fenced code blocks itself. If a `<tagName>` block appears inside a fenced code
|
||||
* block and should be excluded, the caller should apply `stripFencedCode` first.
|
||||
*
|
||||
* **Nested tags are NOT supported.** The underlying regex uses a non-greedy
|
||||
* `[\s\S]*?` match, which means it closes at the FIRST `</tagName>` encountered.
|
||||
* Given `<x><x>inner</x></x>`, `extractTaggedBlocks(content, 'x')` returns
|
||||
* `['<x>inner']` — the inner `<x>` is captured as literal text, and the second
|
||||
* `</x>` is left unmatched (or matched as a second block with empty inner text
|
||||
* if another `<x>` follows). Callers that need to handle nested tags must
|
||||
* pre-process the input or use a proper XML/HTML parser.
|
||||
*
|
||||
* Generalises `decisions.cts`'s bespoke `matchAll(/<decisions>([\s\S]*?)<\/decisions>/g)`
|
||||
* so tier T1 can drop its own copy (tracked duplication until T1 lands).
|
||||
*/
|
||||
export function extractTaggedBlocks(content: string, tagName: string): string[] {
|
||||
if (typeof content !== 'string' || content.length === 0) return [];
|
||||
if (typeof tagName !== 'string' || tagName.length === 0) return [];
|
||||
|
||||
// Escape the tag name for safe interpolation into a RegExp.
|
||||
const escapedTag = tagName.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
|
||||
const pattern = new RegExp(`<${escapedTag}>([\\s\\S]*?)</${escapedTag}>`, 'g');
|
||||
|
||||
const results: string[] = [];
|
||||
let match: RegExpExecArray | null;
|
||||
while ((match = pattern.exec(content)) !== null) {
|
||||
results.push(match[1]);
|
||||
}
|
||||
return results;
|
||||
}
|
||||
|
||||
// ─── replaceSection ───────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Splice `newBody` in place of a section's body and return the resulting
|
||||
* full content string.
|
||||
*
|
||||
* Uses the `bodyStart`/`bodyEnd` character offsets carried by the `Section`
|
||||
* type to perform a pure string splice — no regex, no line-counting. The
|
||||
* heading is preserved verbatim; only the bytes between `bodyStart` and
|
||||
* `bodyEnd` are replaced.
|
||||
*
|
||||
* The `newBody` is inserted as-is between `content.slice(0, bodyStart)` and
|
||||
* `content.slice(bodyEnd)`. If `newBody` should end with a trailing newline
|
||||
* before the next section's heading, the caller is responsible for including
|
||||
* it (consistent with how `trimEnd()` is applied to collected bodies — see
|
||||
* `collectSections`/`collectSection`).
|
||||
*
|
||||
* Typical read-modify-write pattern (T6 state.cts use case):
|
||||
* ```
|
||||
* const section = collectSection(content, h => h.text === 'Name');
|
||||
* if (section) {
|
||||
* content = replaceSection(content, section, newBody);
|
||||
* }
|
||||
* ```
|
||||
*
|
||||
* CRLF-safe: the splice is purely character-offset-based, so CRLF sequences
|
||||
* are preserved in the surrounding content unchanged.
|
||||
*/
|
||||
export function replaceSection(content: string, section: Section, newBody: string): string {
|
||||
if (typeof content !== 'string') return content;
|
||||
if (typeof newBody !== 'string') return content;
|
||||
return content.slice(0, section.bodyStart) + newBody + content.slice(section.bodyEnd);
|
||||
}
|
||||
|
||||
// Consumers: require('../gsd-core/bin/lib/markdown-sectionizer.cjs')
|
||||
// Named CJS exports are the canonical surface (ADR-457 .cts → .cjs build-at-publish).
|
||||
1142
tests/markdown-sectionizer.test.cjs
Normal file
1142
tests/markdown-sectionizer.test.cjs
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user