diff --git a/.changeset/rapid-zebras-zip.md b/.changeset/rapid-zebras-zip.md new file mode 100644 index 000000000..910769132 --- /dev/null +++ b/.changeset/rapid-zebras-zip.md @@ -0,0 +1,5 @@ +--- +type: Added +pr: 4918 +--- +**Planning documents are now read and written through one parse → mutate → serialize seam** — a new internal `PlanningDoc` layer composes the existing markdown-sectionizer, markdown-table and frontmatter seams so a verb can no longer bring its own regex to a `.planning/` artifact. A field write reaches only its own value token, leaving hand-written prose on the same line intact by construction; serialization splices into the original buffer, so untouched regions stay byte-identical; and a write refuses outright on a document containing any region the parser could not read. No command changes behavior yet — this phase adds the seam and migrates no call site. (#4906) diff --git a/.gitignore b/.gitignore index 1503d3824..079e0a5d8 100644 --- a/.gitignore +++ b/.gitignore @@ -213,6 +213,7 @@ build/ /gsd-core/bin/lib/planning-workspace.cjs /gsd-core/bin/lib/planning-scope.cjs /gsd-core/bin/lib/planning-snapshot.cjs +/gsd-core/bin/lib/planning-document.cjs /gsd-core/bin/lib/planning-inspect.cjs /gsd-core/bin/lib/planning-command-router.cjs /gsd-core/bin/lib/plan-document.cjs diff --git a/CONTEXT.md b/CONTEXT.md index ae144786f..823d94de9 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -159,6 +159,9 @@ Leaf module owning the parse of a `*-PLAN.md` document BODY: `` extra ### Task Content Resolution Module Given a task's `tracker-id` attribute (Plan Document Module) and the set of installed capabilities' `taskContentResolver` declarations, resolves that task's content — `action`/`verify`/`acceptanceCriteria`/`readFirst`/`done` — from the matching external tracker via one bounded subprocess call (ADR-3646 Phase 1, #3970), or reports that no resolution applies. Pure/impure split: `splitTrackerId` (first-`:`-only split, colons after the first stay in the id verbatim), `findResolver` (matches a capability's `trackerPrefix`, returns `null`/one match/`'ambiguous'`, never silently picks one), and `buildInvocation` (expands `invoke.args`' `{{id}}` placeholder) are pure and total, never throwing even on hostile third-party-shaped capability input; `resolveTaskContent` is the one impure boundary, spawning through an injectable `execFn` (`spawnSync` with a bounded `timeout`) so tests never spawn a real process or wait a real timeout. **Four non-throwing outcomes** (`TASK_CONTENT_RESULT`: `not-applicable`, `no-resolver`, `resolved`, `empty`) and **HARD-HALT on four throwing outcomes** (`ResolverAmbiguousError`, `ResolverFailedError`, `ResolverTimeoutError`, `ResolverMalformedOutputError`) — ADR-3646 Decision 4: an ambiguous resolver match, a non-zero resolver exit, a timeout, or malformed resolver stdout must never degrade to a silently-empty or silently-picked result. Wired at `task resolve-content --plan --task-id --raw` (`task-command-router.cts`'s `routeResolveContent`), which turns each of the four throws into the CLI's own non-zero exit rather than swallowing them into a `{resolved: false}` JSON answer; capabilities are loaded through the merged first-party + validated-installed-overlay registry (ADR-1244 D2) so a third-party capability's resolver is honored. Source of truth: `gsd-core/bin/lib/task-content-resolution.cjs` (generated from `src/task-content-resolution.cts`). +### Planning Document Module +Canonical parse → mutate → serialize seam for a `.planning/` artifact as ONE object (`gsd-core/bin/lib/planning-document.cjs`; ADR-4910, epic #4906 Phase 1, #4917). Pure functions, Node built-ins only, string-in/value-out, no I/O. It is the **composition** layer the three existing seams never covered: the Markdown Sectionizer supplies structure (fences, headings, sections — never re-scanned here), the Markdown Table Model supplies tables, the Frontmatter Module supplies frontmatter, and the Write-Set Module supplies `Result`; this module owns a planning document as frontmatter + section tree + typed field nodes (`boldField`, `table`, `checklist`) and the write boundary over them. Exports: `parsePlanningDoc(source, artifact) → Result`; `findField(doc, label) → NodeId | null` (handle minting — a caller cannot address a node the parser did not find, ADR-4910 §2's replacement for path-string addressing, which would be a grammar needing a parser); `readNode(doc, id)` (node-scoped `could-not-parse` with the offending span, so an unreadable sibling never makes a readable node unreadable, §5); `setFieldValue(doc, id, value) → Result` (immutable — returns a NEW doc with the edit STAGED, writing only into that node's `valueSpan`; `trailingSpan` is readable and structurally unwritable, which is how the #2853/#3584/#4852 "the verb owns the count token ONLY" rule stops depending on an author reading a comment); `hasUnreadableNodes(doc)`; `serialize(doc) → SerializeOutcome` (splices staged spans into the ORIGINAL buffer, so an untouched region is the original bytes and a no-edit serialize is byte-identical — §3's Hyrum's-Law commitment, which closes #4499 without a targeted fix, and which also makes an already-escaped table cell impossible to double-escape on a round trip). Per ADR-4910's 2026-09-21 amendment the serializer **refuses** when any node carries a parse error, even with zero staged edits: a read answers a bounded question from a bounded region, but a write re-emits the whole file and therefore asserts "this is the document I read", which a writer cannot claim about a region it could not parse. That refusal is a `SerializeOutcome`, a different axis from §5's document-level `Result` parse failure (reserved for "not a planning artifact at all"), so a document can be valid to OPEN and still refuse to be WRITTEN. **Stated limit (amendment):** the rule makes a write *whole*, not *correct* — a value sourced from outside the document (e.g. `phase.complete`'s plan/summary counts, read from a filesystem scan at `src/phase.cts:3429-3434`, not from any node) is outside what this seam can check. The artifact registry derives from `artifacts.cts`'s `isCanonicalPlanningFile` rather than inventing a second notion of what a planning document is, with a parity test against divergence. Phase 1 parses every node kind but makes only `boldField` writable; table/checklist writers arrive with Phase 3's shared reader/writer escaping. Not to be confused with the **Plan Document Module** (`plan-document.cjs`), which parses one `*-PLAN.md` BODY and is absorbed beneath this seam as a typed field reader in Phase 6. + ### Planning Inspect Module Module owning the **schema-v1 canonical planning snapshot** emitted by the read-only `planning inspect` query (#2790), for downstream harness UIs that need truthful `.planning/` state without parsing ROADMAP/REQUIREMENTS/PLAN/SUMMARY Markdown a second time. `buildPlanningInspect(cwd) → payload`; `cmdPlanningInspect(cwd, raw)` emits it through `output()` (so the existing >50 KB `@file:` spill seam applies unchanged). `PLANNING_INSPECT_SCHEMA_VERSION = 1` is the wire contract — a consumer MUST reject any other value rather than best-effort-parse an unknown shape. Composes, never re-derives: milestone identity/windowing and phase enumeration via `buildPlanningSnapshot` (Planning Snapshot Module), completion via `isPhaseComplete` (§7.4, disk-strict), live-plan counting via `scanPhasePlans` (§7.5), percent via `clampPercent` (§7.6), STATE fields via `stateFieldValue`/`stateCurrentPositionSlice` (§7.7), plan bodies via `parsePlanDocument`, requirement IDs via `parseRequirements`, UAT items via `parseUatItems`/`selectPhaseUatFiles`. **It deliberately does NOT serialize `PlanningSnapshot`**: that shape is the §8.1 diagnostic-rule subject and is explicitly additive/growing (4 fields at Phase 10, 20+ by Phase 12), so handing it to external consumers would freeze an internal contract by accident (Hyrum's Law) — this module declares its own flat schema and maps into it, and a field added to `PlanningSnapshot` must never change schema-v1 output. Three frozen enums carry every non-answer — `INSPECT_DIAGNOSTIC`, `TASK_STATUS` (`done|pending|unknown`), `PROVENANCE` (`task_scoped|plan_scoped|absent`) and `AGREEMENT` (`agreed|conflicting|unknown`) — because unknown or conflicting evidence serializes as `unknown` plus a diagnostic and is **never inferred, reconciled, or defaulted**; keys are always present, `null` is the explicit non-answer. Roadmap acceptance, verification and UAT are reported **side by side and never folded into one verdict**, and a ROADMAP checkbox is emitted with `authoritative: false` per §7.4. **Not a diagnostic rule**, and deliberately NOT registered in `scripts/lint-planning-snapshot-bypass-drift.cjs`, which is `DIAGNOSTIC_RULE_FUNCTIONS`-scoped and must remain prunable to zero when #3309 lands. Dispatched by the Planning Command Router (`src/planning-command-router.cts`, family `planning`, subcommand `inspect`, no arguments in v1 — a stray positional or unknown flag is a fail-loud `ERROR_REASON.USAGE`). Source of truth: `gsd-core/bin/lib/planning-inspect.cjs` (generated from `src/planning-inspect.cts`). Design: `.gsd/phase/feat-2790-planning-inspect/40-design.md`. diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index fcb1ec693..0711da028 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -504,6 +504,7 @@ "plan-drift-guard.cjs", "plan-scan.cjs", "planning-command-router.cjs", + "planning-document.cjs", "planning-inspect.cjs", "planning-scope.cjs", "planning-snapshot.cjs", diff --git a/docs/INVENTORY.md b/docs/INVENTORY.md index fa5da992e..413cdeeba 100644 --- a/docs/INVENTORY.md +++ b/docs/INVENTORY.md @@ -641,6 +641,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`. | `plan-drift-guard.cjs` | Classifies symbol-verification severity against the ADR-22 authority ladder | | `plan-scan.cjs` | Canonical phase-plan scanner for detecting plan and summary files in flat and nested layouts (k014) | | `planning-command-router.cjs` | Thin CJS subcommand router for `gsd-tools planning` (compiled from `src/planning-command-router.cts`, gitignored; #2790) — one subcommand, `inspect`; v1 accepts no arguments, so a stray positional or unknown flag is a fail-loud `ERROR_REASON.USAGE` rather than a silently-ignored one | +| `planning-document.cjs` | Canonical parse → mutate → serialize seam for a `.planning/` artifact as ONE object (compiled from `src/planning-document.cts`, gitignored; ADR-4910, epic #4906 Phase 1, #4917) — the composition layer over the Markdown Sectionizer, Markdown Table Model, Frontmatter and Write-Set seams. `parsePlanningDoc` yields frontmatter + section tree + typed field nodes; `findField` mints an opaque handle so a caller cannot address a node the parser did not find; `setFieldValue` writes only into that node's `valueSpan`, leaving `trailingSpan` structurally unwritable (the #2853/#3584/#4852 rule, by construction rather than by comment); `serialize` splices staged spans into the ORIGINAL buffer, so a no-edit serialize is byte-identical (#4499). Per ADR-4910's 2026-09-21 amendment the serializer refuses whenever any node carries a parse error — reads are node-scoped, writes are document-scoped, because serialization re-emits the whole file | | `planning-inspect.cjs` | Read-only schema-v1 canonical planning snapshot (compiled from `src/planning-inspect.cts`, gitignored; #2790) — `buildPlanningInspect(cwd)` / `cmdPlanningInspect`; `PLANNING_INSPECT_SCHEMA_VERSION = 1` is the wire contract consumers must reject other values of. Composed from the ADR-3180 §7 owners plus `parsePlanDocument`/`parseRequirements`/`parseUatItems`, and read through the Markdown Sectionizer + Markdown Table Model seams; deliberately declares its own flat external schema rather than serializing the still-growing `PlanningSnapshot`. Frozen `INSPECT_DIAGNOSTIC`/`TASK_STATUS`/`PROVENANCE`/`AGREEMENT` enums carry every non-answer — unknown or conflicting evidence is reported with a coded diagnostic, never inferred | | `planning-scope.cjs` | Frozen `SCOPE` discriminator (`COMPLETE`/`TRUNCATED`/`UNSCOPED`/`UNREADABLE`) distinguishing a genuinely-empty derivation from one computed over a truncated or unscoped input, so callers can branch on the difference instead of reading a plausible zero (ADR-3180) | | `planning-snapshot.cjs` | Parsed projection of `.planning/` composed exclusively from the ADR-3180 §7 owners (milestone identity, phase enumeration, phase completion, plan/summary counting, STATE.md current-phase) — exposes only scope-carrying parsed values, never raw document text, so a diagnostic rule cannot re-derive a field's location (ADR-3180 §8.1) | diff --git a/eslint.config.mjs b/eslint.config.mjs index 83bf8a269..a56b0545c 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -126,6 +126,8 @@ export default tseslint.config( 'gsd-core/bin/lib/resolution.cjs', 'gsd-core/bin/lib/unusable-input.cjs', 'gsd-core/bin/lib/plan-drift-guard.cjs', + // #4917 (epic #4906 Phase 1, ADR-4910): lint src/planning-document.cts, not this. + 'gsd-core/bin/lib/planning-document.cjs', // #2401: tsc-generated runtime artifact — lint the src/verify-command-grounding.cts source. 'gsd-core/bin/lib/verify-command-grounding.cjs', 'gsd-core/bin/lib/cli-exit.cjs', diff --git a/src/frontmatter.cts b/src/frontmatter.cts index 44c097ffc..757e8418e 100644 --- a/src/frontmatter.cts +++ b/src/frontmatter.cts @@ -1597,4 +1597,9 @@ export = { cmdFrontmatterMerge, cmdFrontmatterValidate, propagateCommentChannel, + // #4917 / ADR-4910 Decision 1: additive-only export so `planning-document.cts` + // can COMPOSE this seam's fence-detection grammar (byte-0 rule, BOM strip, + // CR handling) instead of reimplementing it. No behavior change — same + // function `extractFrontmatter`/`frontmatterListEntries` already call. + frontmatterRegion, }; diff --git a/src/planning-document.cts b/src/planning-document.cts new file mode 100644 index 000000000..758aef507 --- /dev/null +++ b/src/planning-document.cts @@ -0,0 +1,577 @@ +/** + * Planning Document — the parse -> mutate -> serialize seam for a `.planning/` + * root artifact BODY (ADR-4910, epic #4906 Phase 1, #4917). + * + * Composes the existing structural seams — never reimplements them: + * - `markdown-sectionizer.cjs` (`tokenizeHeadings`, `collectSections`, + * `scanFencedBlocks`, `scanInlineCodeSpans`) for headings/sections and + * fence/inline-code awareness. + * - `markdown-table.cjs` (`splitTableRow`, `isDelimiterRow`, + * `parseMarkdownTable`) for GFM table detection and validation. + * - `artifacts.cjs` (`isCanonicalPlanningFile`) for the artifact-kind gate. + * + * This phase migrates NO call site — it is purely additive (ADR-4910 §7). + * Only `boldField` nodes are writable; `table`/`checklist` nodes parse and + * read only (their writers are Phase 3's escaping work). + * + * Hyrum's Law commitment (row 3 of the design's behaviour table): `serialize` + * with zero staged edits returns `doc.source` BYTE-IDENTICAL — never a + * re-render (#4499's root cause). Every byte outside an edited `valueSpan` is + * the ORIGINAL source, spliced, never regenerated. + * + * ADR-457 build-at-publish: source in src/planning-document.cts, compiled to + * gsd-core/bin/lib/planning-document.cjs (gitignored). + */ + +import { tokenizeHeadings, collectSections, scanFencedBlocks, iterateBullets } from './markdown-sectionizer.cjs'; +import { splitTableRow, isDelimiterRow, parseMarkdownTable } from './markdown-table.cjs'; +import { isCanonicalPlanningFile, CANONICAL_EXACT } from './artifacts.cjs'; +// `frontmatter.cts` uses `export =` (CJS-style single export object), so it +// is imported as a default import (esModuleInterop), not a named import. +import frontmatterModule from './frontmatter.cjs'; +import type { Result } from './write-set.cjs'; + +const { frontmatterRegion } = frontmatterModule; + +// ─── Types ──────────────────────────────────────────────────────────────────── + +/** Character offsets into the ORIGINAL `PlanningDoc.source` string. */ +export interface Span { + start: number; + end: number; +} + +/** Why a node failed to parse, and exactly where the bad span is. */ +export interface NodeError { + reason: string; + span: Span; +} + +export type NodeKind = 'frontmatter' | 'section' | 'boldField' | 'table' | 'checklist'; + +/** Opaque node handle, minted by `parsePlanningDoc`. A caller cannot name a + * node the parser did not find — `findField` is the only way to obtain one. */ +export type NodeId = string; + +/** Common shape every planning node carries. */ +interface BaseNode { + id: NodeId; + span: Span; + error: NodeError | null; +} + +/** The `---\n...\n---` YAML frontmatter block, span-only (this phase does not + * parse the YAML itself — that is `frontmatter.cts`'s job). */ +export interface FrontmatterNode extends BaseNode { + kind: 'frontmatter'; +} + +/** One heading + body region, per `collectSections`. */ +export interface SectionNode extends BaseNode { + kind: 'section'; + heading: string; + level: number; +} + +/** + * A `**Label:** value` line. `valueSpan` is the WRITE BOUNDARY — the only + * span any exported function accepts as a write target. `trailingSpan` is + * the rest of the line (a hand-written annotation, an em-dash note, etc.): + * readable via `readNode`'s reconstructed text is not exposed for it, but + * the field itself is public data on this node — there is deliberately no + * exported function that writes into it (ADR-4910 §1's structural rule). + */ +export interface BoldFieldNode extends BaseNode { + kind: 'boldField'; + label: string; + labelSpan: Span; + valueSpan: Span; + trailingSpan: Span; + value: string; +} + +/** A GFM pipe table found in the document body. */ +export interface TableNode extends BaseNode { + kind: 'table'; + columns: string[] | null; +} + +/** A contiguous run of checkbox-bullet lines (`- [ ] ...` / `- [x] ...`). */ +export interface ChecklistNode extends BaseNode { + kind: 'checklist'; + items: number; +} + +export type PlanningNode = FrontmatterNode | SectionNode | BoldFieldNode | TableNode | ChecklistNode; + +export interface PlanningDoc { + readonly source: string; + readonly artifact: string; + readonly nodes: readonly PlanningNode[]; + readonly staged: ReadonlyMap; +} + +export type NodeRead = { ok: true; value: string } | { ok: false; reason: string; span: Span }; + +export type SerializeOutcome = + | { ok: true; value: string } + | { + ok: false; + reason: 'unreadable-nodes'; + nodes: Array<{ id: NodeId; kind: NodeKind; span: Span; reason: string }>; + }; + +/** + * Canonical `.planning/` root artifact basenames this seam recognises, + * derived from the SAME registry `isCanonicalPlanningFile` consults + * (`artifacts.cts`'s `CANONICAL_EXACT`) — never a second, independently + * maintained list. + * + * Filtered to `.md` names only: `CANONICAL_EXACT` also carries non-markdown + * artifacts (`config.json`, `state.json`, `milestone.lock`, …) that this + * parser has no grammar for. Handing that JSON/lock content to the markdown + * parser below returns a successful EMPTY document (`nodes: []`), which reads + * as "this document records nothing" when the truth is "wrong kind entirely" + * — the empty-vs-error confusion #4917 / ADR-4910 §5 exists to eliminate. Do + * NOT remove this filter to "restore" the full registry. + */ +export const PLANNING_ARTIFACTS: readonly string[] = Object.freeze( + Array.from(CANONICAL_EXACT).filter((name) => name.endsWith('.md')), +); + +// ─── Internal helpers ─────────────────────────────────────────────────────── + +let nodeCounter = 0; +function mintId(kind: NodeKind): NodeId { + nodeCounter += 1; + return `${kind}-${nodeCounter}-${Math.random().toString(36).slice(2, 8)}`; +} + +interface LineInfo { + /** Line text WITHOUT a trailing `\r` (CRLF-safe). */ + text: string; + /** Absolute char offset of this line's first character in `source`. */ + start: number; + /** Absolute char offset one past this line's last content char, BEFORE + * any `\r`/`\n` — i.e. `source.slice(start, end) === text`. */ + end: number; +} + +function splitLinesInfo(source: string): LineInfo[] { + const out: LineInfo[] = []; + let offset = 0; + const rawLines = source.split('\n'); + for (let i = 0; i < rawLines.length; i++) { + const raw = rawLines[i]; + const hasCR = raw.endsWith('\r'); + const text = hasCR ? raw.slice(0, -1) : raw; + out.push({ text, start: offset, end: offset + text.length }); + offset += raw.length + 1; // +1 for the '\n' split on ('\r' already counted in raw.length) + } + return out; +} + +/** + * Locate the frontmatter block, if any, by COMPOSING `frontmatter.cts`'s + * `frontmatterRegion` — the fence-detection grammar (byte-0 rule, BOM strip, + * `\n---` search, CR handling) lives there, once, and this seam never + * re-derives it (ADR-4910 Decision 1). + * + * `frontmatterRegion` reports the YAML body's own bounds (`region`, + * `terminated`, and the possibly BOM-stripped `content`), not this seam's + * `Span` shape (an absolute byte range into the UNSTRIPPED `source`, + * inclusive of both fences). This adapter translates one into the other by + * reading ONLY the two boundary characters `frontmatterRegion` already + * anchored (whether the YAML end / closing fence sit on a CRLF line) — it + * does not re-scan for the fences themselves. + */ +function findFrontmatterSpan(source: string): { span: Span; terminated: boolean } | null { + const found = frontmatterRegion(source); + if (!found) return null; + + // `found.content` may be `source` with a single leading BOM stripped; + // every offset below is relative to `found.content`, so translate back to + // `source` coordinates by the same delta. + const bomDelta = source.length - found.content.length; + const content = found.content; + + if (!found.terminated) { + return { span: { start: bomDelta, end: bomDelta + content.length }, terminated: false }; + } + + // `frontmatterRegion` already did fence DETECTION — `found` being non-null + // and `terminated` IS that result. It reports only the YAML body's bounds + // (`region`), not an absolute span, so recover the closing fence's end + // from `region`'s length. The one thing still read directly here is the + // opening fence's fixed-width line ending (`\n` vs `\r\n`), needed to + // translate `region`'s length into a `content` offset — not a re-scan for + // the fence itself. + const headerEnd = content.startsWith('---\r\n') ? 5 : 4; + const yamlEnd = headerEnd + found.region.length; + const closingLineStart = content[yamlEnd] === '\r' ? yamlEnd + 1 : yamlEnd; + const fenceLineStart = closingLineStart + 1; + let fenceEnd = fenceLineStart + 3; + if (content[fenceEnd] === '\r') fenceEnd += 1; + return { span: { start: bomDelta, end: bomDelta + fenceEnd }, terminated: true }; +} + +/** Build the set of 0-based line indices that fall inside a fenced code + * block (opening/closing delimiter lines included), so `**Label:**`/table/ + * checklist scanning never treats fenced content as a node (rows 9/14). */ +function fencedLineIndices(lines: LineInfo[]): Set { + const raw = lines.map((l) => l.text); + const blocks = scanFencedBlocks(raw); + const set = new Set(); + for (const b of blocks) { + const end = b.closeLineIdx === -1 ? raw.length - 1 : b.closeLineIdx; + for (let i = b.openLineIdx; i <= end; i++) set.add(i); + } + return set; +} + +const BOLD_FIELD_RE = /^(\s*)(\*\*[^*\r\n]+:\*\*)([ \t]*)([^\r\n]*)$/; +/** Boundary marking a hand-written trailing annotation on a field line — + * the token owner must never destroy prose past this separator. */ +const TRAILING_SEPARATOR_RE = / — /; + +function parseBoldFieldLine(line: LineInfo): BoldFieldNode | null { + const m = BOLD_FIELD_RE.exec(line.text); + if (!m) return null; + const [, leading, token, spacing, rest] = m; + const labelStart = line.start + leading.length; + const labelSpan: Span = { start: labelStart, end: labelStart + token.length }; + const label = token.slice(2, -3); + const restStart = labelSpan.end + spacing.length; + + const sepMatch = TRAILING_SEPARATOR_RE.exec(rest); + const valueRaw = sepMatch ? rest.slice(0, sepMatch.index) : rest; + const trimmedValue = valueRaw.replace(/\s+$/, ''); + const valueSpan: Span = { start: restStart, end: restStart + trimmedValue.length }; + const trailingSpan: Span = { start: valueSpan.end, end: line.end }; + + return { + kind: 'boldField', + id: mintId('boldField'), + span: { start: labelSpan.start, end: line.end }, + error: null, + label, + labelSpan, + valueSpan, + trailingSpan, + value: trimmedValue, + }; +} + +/** A checklist line is one whose SOLE bullet, per `iterateBullets` (the same + * grammar the repo's other bullet consumers use), is a checkbox marker, OR + * whose bullet TEXT begins with a task-list marker. + * + * `iterateBullets` owns bullet *structure* — is this a bullet, where does its + * text start — and continues to own that here unchanged. It only classifies + * `-`-prefixed bullets as `checkbox-checked`/`checkbox-unchecked`; GFM also + * permits `*` and `+` as bullet markers, and `* [ ] x` / `+ [x] y` are valid + * GFM task-list items that `iterateBullets` reports as plain `dash`-family + * bullets with the `[ ]`/`[x]` left in the bullet's own text. Widening + * `iterateBullets` itself is forbidden by ADR-2143 §2's extend-never-mutate + * lock (inherited by this epic), so the task-list-marker interpretation is + * layered on here, over the bullet's already-extracted text — never by + * re-scanning the raw line with a new hand-rolled regex. + * + * Known limit inherited from `iterateBullets`, not introduced here: + * `-\t[ ] text` (a tab between the marker and the text) is not recognised as + * a bullet at all, so it can never become a checklist line. That is a + * pre-existing `markdown-sectionizer` boundary affecting every consumer of + * `iterateBullets`, and fixing it would mean altering the locked seam. */ +function isChecklistLine(text: string): boolean { + const items = iterateBullets(text); + if (items.length !== 1) return false; + const item = items[0]; + if (item.marker === 'checkbox-checked' || item.marker === 'checkbox-unchecked') return true; + return /^\[[ xX]\] /.test(item.text); +} + +/** + * Scan the document body (everything outside the frontmatter block and + * outside fenced code) for `boldField`, `table`, and `checklist` nodes, in + * document order. + */ +function scanBodyNodes(source: string, lines: LineInfo[], frontmatterEnd: number): PlanningNode[] { + const fenced = fencedLineIndices(lines); + const nodes: PlanningNode[] = []; + let i = 0; + while (i < lines.length) { + const line = lines[i]; + if (fenced.has(i) || line.start < frontmatterEnd) { + i += 1; + continue; + } + + const trimmed = line.text.trim(); + + // Table: a pipe-shaped header line followed by a valid delimiter row. + if (trimmed.startsWith('|') && trimmed.indexOf('|', 1) !== -1 && i + 1 < lines.length) { + const delimiterLine = lines[i + 1]; + const delimiterCells = splitTableRow(delimiterLine.text); + const headerCells = splitTableRow(line.text); + if ( + delimiterLine.text.trim().startsWith('|') + && isDelimiterRow(delimiterCells) + && delimiterCells.length === headerCells.length + && !fenced.has(i + 1) + ) { + let last = i + 1; + while (last + 1 < lines.length && lines[last + 1].text.trim().startsWith('|') && !fenced.has(last + 1)) { + last += 1; + } + const span: Span = { start: line.start, end: lines[last].end }; + const tableText = source.slice(span.start, span.end); + const parsed = parseMarkdownTable(tableText); + nodes.push( + parsed.ok + ? { + kind: 'table', + id: mintId('table'), + span, + error: null, + columns: parsed.value.columns, + } + : { + kind: 'table', + id: mintId('table'), + span, + error: { reason: parsed.reason, span }, + columns: null, + }, + ); + i = last + 1; + continue; + } + } + + // Checklist: a contiguous run of checkbox-bullet lines. + if (isChecklistLine(line.text)) { + let last = i; + let count = 0; + while (last < lines.length && !fenced.has(last) && isChecklistLine(lines[last].text)) { + count += 1; + last += 1; + } + last -= 1; + const span: Span = { start: line.start, end: lines[last].end }; + nodes.push({ kind: 'checklist', id: mintId('checklist'), span, error: null, items: count }); + i = last + 1; + continue; + } + + // Bold field. + const field = parseBoldFieldLine(line); + if (field) { + nodes.push(field); + i += 1; + continue; + } + + i += 1; + } + return nodes; +} + +// ─── Public API ───────────────────────────────────────────────────────────── + +/** + * Parse `source` (the raw text of a `.planning/` root artifact) into a + * `PlanningDoc`. Document-level `Result` failure is reserved for: `artifact` + * not a recognised planning artifact kind, `source` not a readable string, or + * an opened-but-never-closed frontmatter fence (ADR-4910 §5's reservation). + * A malformed SUB-structure (a ragged table, say) never fails the whole + * document — it is recorded as that one node's `error`, and every sibling + * node stays readable (row 7). `nodes: []` on a genuinely empty document is + * success, not an error (row 15). + */ +export function parsePlanningDoc(source: string, artifact: string): Result { + if (typeof source !== 'string') { + return { ok: false, reason: 'unreadable: source is not a string' }; + } + if ( + typeof artifact !== 'string' || + !isCanonicalPlanningFile(artifact) || + !PLANNING_ARTIFACTS.includes(artifact) + ) { + return { + ok: false, + reason: `not a markdown planning document (artifact: ${String(artifact)})`, + }; + } + + const nodes: PlanningNode[] = []; + let frontmatterEnd = 0; + + const fm = findFrontmatterSpan(source); + if (fm) { + if (!fm.terminated) { + return { ok: false, reason: 'no frontmatter terminator' }; + } + nodes.push({ kind: 'frontmatter', id: mintId('frontmatter'), span: fm.span, error: null }); + frontmatterEnd = fm.span.end; + } + + if (source.length === 0) { + return { ok: true, value: { source, artifact, nodes: [], staged: new Map() } }; + } + + const lines = splitLinesInfo(source); + + // Sections: one per heading, in document order — every heading is its own + // boundary (`collectSections(source, () => true)`), so a nested `####` + // still gets its own SectionNode rather than being folded into its parent. + const headings = tokenizeHeadings(source); + if (headings.length > 0) { + const sections = collectSections(source, () => true); + for (const s of sections) { + nodes.push({ + kind: 'section', + id: mintId('section'), + span: { start: s.heading.offset, end: s.bodyEnd }, + error: null, + heading: s.heading.text, + level: s.heading.level, + }); + } + } + + nodes.push(...scanBodyNodes(source, lines, frontmatterEnd)); + + nodes.sort((a, b) => a.span.start - b.span.start); + + return { ok: true, value: { source, artifact, nodes, staged: new Map() } }; +} + +/** Find the id of the (first, document-order) `boldField` node whose label + * exactly matches `label`, or `null` when none does. */ +export function findField(doc: PlanningDoc, label: string): NodeId | null { + for (const n of doc.nodes) { + if (n.kind === 'boldField' && n.label === label) return n.id; + } + return null; +} + +/** Read a node by id. Node-scoped failure only — an unknown id or a node + * that failed to parse never throws. */ +export function readNode(doc: PlanningDoc, id: NodeId): NodeRead { + const node = doc.nodes.find((n) => n.id === id); + if (!node) { + return { ok: false, reason: 'unknown node id', span: { start: 0, end: 0 } }; + } + if (node.error) { + return { ok: false, reason: node.error.reason, span: node.error.span }; + } + if (node.kind === 'boldField') { + return { ok: true, value: doc.staged.get(id) ?? node.value }; + } + return { ok: true, value: doc.source.slice(node.span.start, node.span.end) }; +} + +/** + * Stage a new value for a `boldField` node, returning a NEW `PlanningDoc` + * (immutable — `doc` itself is never mutated). Refuses an id this doc did + * not mint, and refuses any node kind other than `boldField` — only the + * `valueSpan` is ever writable this phase (ADR-4910 §1). + */ +export function setFieldValue(doc: PlanningDoc, id: NodeId, value: string): Result { + const node = doc.nodes.find((n) => n.id === id); + if (!node) { + return { ok: false, reason: 'unknown node id' }; + } + if (node.kind !== 'boldField') { + return { ok: false, reason: `node kind '${node.kind}' is not writable this phase` }; + } + // #4917 / ADR-4910 Decision 2 & 4: a boldField's token boundary is a LINE + // boundary, not just an offset range — a value containing \n or \r escapes + // the field's own span and reparses as sibling structure (a forged field) + // once spliced back into the source. Decision 4 licenses refusal for any + // value the grammar cannot represent; Phase 3 may widen this to escaping, + // but Phase 1 refuses outright. Do not remove this as an over-restriction. + if (/[\r\n]/.test(value)) { + return { ok: false, reason: 'field value must not contain a line break (\\r or \\n)' }; + } + // #4917 / ADR-4910 Decision 4: "a value that cannot be represented in the + // grammar is refused by the writer, with a report." This is a GENERAL + // round-trip representability check, not a blacklist of forbidden + // substrings — the `\r`/`\n` guard above is a narrower special case kept + // for its clearer message, but THIS check is the backstop. It rebuilds the + // line exactly as it would be written (existing leading/label/spacing + + // the new value + the existing trailing text) and re-parses that line + // through the SAME `parseBoldFieldLine` grammar the reader uses. If the + // value the grammar reads back is not byte-identical to what the caller + // staged, the grammar cannot represent this value (e.g. it contains the + // ` — ` trailing-separator token, which would silently reclassify the + // rest of the value as trailing prose) and the write is refused. Do NOT + // replace this with a list of forbidden characters/substrings — the next + // separator the grammar grows would silently slip past a blacklist. + const leadingText = doc.source.slice(node.span.start, node.labelSpan.start); + const tokenText = doc.source.slice(node.labelSpan.start, node.labelSpan.end); + const spacingText = doc.source.slice(node.labelSpan.end, node.valueSpan.start); + const trailingText = doc.source.slice(node.trailingSpan.start, node.trailingSpan.end); + const candidateLine = `${leadingText}${tokenText}${spacingText}${value}${trailingText}`; + const candidateInfo: LineInfo = { text: candidateLine, start: 0, end: candidateLine.length }; + const reparsed = parseBoldFieldLine(candidateInfo); + if (!reparsed || reparsed.value !== value) { + return { + ok: false, + reason: 'field value is not representable in the boldField grammar (would not round-trip)', + }; + } + const staged = new Map(doc.staged); + staged.set(id, value); + return { ok: true, value: { source: doc.source, artifact: doc.artifact, nodes: doc.nodes, staged } }; +} + +/** True when any node in `doc` failed to parse. */ +export function hasUnreadableNodes(doc: PlanningDoc): boolean { + return doc.nodes.some((n) => n.error !== null); +} + +/** + * Splice every staged edit into `doc.source` and return the resulting text. + * With zero staged edits, returns `doc.source` BYTE-IDENTICAL — never a + * re-render (row 3). Refuses outright — even with zero staged edits — when + * `hasUnreadableNodes(doc)` is true (the ADR-4910 amendment): `serialize` + * re-emits the WHOLE document, so the refusal is document-scoped, not + * mutation-scoped. + */ +export function serialize(doc: PlanningDoc): SerializeOutcome { + if (hasUnreadableNodes(doc)) { + return { + ok: false, + reason: 'unreadable-nodes', + nodes: doc.nodes + .filter((n): n is PlanningNode & { error: NodeError } => n.error !== null) + .map((n) => ({ id: n.id, kind: n.kind, span: n.error.span, reason: n.error.reason })), + }; + } + + if (doc.staged.size === 0) { + return { ok: true, value: doc.source }; + } + + const edits: Array<{ start: number; end: number; value: string }> = []; + for (const [id, value] of doc.staged) { + const node = doc.nodes.find((n) => n.id === id); + if (!node || node.kind !== 'boldField') continue; // unreachable: setFieldValue already gated this + edits.push({ start: node.valueSpan.start, end: node.valueSpan.end, value }); + } + edits.sort((a, b) => a.start - b.start); + + let out = ''; + let cursor = 0; + for (const e of edits) { + out += doc.source.slice(cursor, e.start) + e.value; + cursor = e.end; + } + out += doc.source.slice(cursor); + + return { ok: true, value: out }; +} + +// Consumers: require('../gsd-core/bin/lib/planning-document.cjs') +// Named CJS exports are the canonical surface (ADR-457 .cts → .cjs build-at-publish). diff --git a/tests/planning-document.test.cjs b/tests/planning-document.test.cjs new file mode 100644 index 000000000..5a51b39a3 --- /dev/null +++ b/tests/planning-document.test.cjs @@ -0,0 +1,781 @@ +/** + * PlanningDoc seam — parse -> mutate -> serialize (ADR-4910, epic #4906 Phase 1, + * #4917). Covers every row of `.gsd/phase/feat-4917-planning-document-seam/50-test-matrix.md`. + * + * All assertions are structural: either a typed Result/NodeRead/SerializeOutcome + * shape, or a byte-range comparison computed from the node's OWN `Span` offsets + * (never a hardcoded literal expectation of rendered prose) — per CONTRIBUTING.md + * "Prohibited: Raw Text Matching on Test Outputs". Full-string byte-identity + * comparisons (`assert.strictEqual(result, source)`) are used only where the + * module's contract IS byte equality: rows 3, 15, 21, 22, 23. + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fc = require('fast-check'); + +const MODULE_PATH = '../gsd-core/bin/lib/planning-document.cjs'; +const ARTIFACTS_PATH = '../gsd-core/bin/lib/artifacts.cjs'; + +let mod; +try { + mod = require(MODULE_PATH); +} catch (err) { + throw new Error(`Could not require ${MODULE_PATH}. Run "npm run build:lib" first. Underlying: ${err.message}`); +} +const { parsePlanningDoc, findField, readNode, setFieldValue, hasUnreadableNodes, serialize, PLANNING_ARTIFACTS } = mod; + +const artifactsMod = require(ARTIFACTS_PATH); +const { isCanonicalPlanningFile, CANONICAL_EXACT } = artifactsMod; + +const ARTIFACT = 'STATE.md'; +const EM_DASH = '—'; + +/** Parse and assert success, returning the PlanningDoc value. */ +function parseOk(source, artifact = ARTIFACT) { + const result = parsePlanningDoc(source, artifact); + assert.strictEqual(result.ok, true, `expected parse to succeed: ${JSON.stringify(result)}`); + return result.value; +} + +// ─── Row 1 / 2: happy path — byte-range preservation ─────────────────────────── + +describe('row 1-2: setFieldValue + serialize preserve every byte outside valueSpan', () => { + test('setFieldValue preserves every byte outside the value span', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + '**Plans:** initial value', + '**Status:** pending', + '', + ].join('\n'); + + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + assert.ok(id); + const node = doc.nodes.find((n) => n.id === id); + + const staged = setFieldValue(doc, id, 'updated value'); + assert.strictEqual(staged.ok, true); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + + const before = source.slice(0, node.valueSpan.start); + const after = source.slice(node.valueSpan.end); + assert.strictEqual(outcome.value.slice(0, node.valueSpan.start), before); + assert.strictEqual(outcome.value.slice(node.valueSpan.start + 'updated value'.length), after); + }); + + test('a field write does not touch trailing prose on the same line (#4852/#4862)', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + `**Summary:** short summary ${EM_DASH} hand-written annotation stays`, + '', + ].join('\n'); + + const doc = parseOk(source); + const id = findField(doc, 'Summary'); + const node = doc.nodes.find((n) => n.id === id); + assert.strictEqual(node.kind, 'boldField'); + + // The trailing annotation lives entirely past valueSpan.end. + const trailingBefore = source.slice(node.valueSpan.end, node.trailingSpan.end); + + const staged = setFieldValue(doc, id, 'new short summary'); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + + const trailingAfter = outcome.value.slice( + node.valueSpan.start + 'new short summary'.length, + node.valueSpan.start + 'new short summary'.length + trailingBefore.length, + ); + assert.strictEqual(trailingAfter, trailingBefore); + }); +}); + +// ─── Row 3/4/5: boundary — limit-1 (0 edits), limit (1 edit), limit+1 (2 edits) ─ + +describe('row 3-5: staged-edit count boundary', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + '**Alpha:** one', + '**Beta:** two', + '', + ].join('\n'); + + test('serialize with no edits returns the original bytes (limit-1)', () => { + const doc = parseOk(source); + const outcome = serialize(doc); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value, source); + }); + + test('a single staged edit splices exactly one span (limit)', () => { + const doc = parseOk(source); + const id = findField(doc, 'Alpha'); + const node = doc.nodes.find((n) => n.id === id); + const staged = setFieldValue(doc, id, 'ONE'); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + const expectedLength = source.length - (node.valueSpan.end - node.valueSpan.start) + 'ONE'.length; + assert.strictEqual(outcome.value.length, expectedLength); + assert.strictEqual(outcome.value.slice(0, node.valueSpan.start), source.slice(0, node.valueSpan.start)); + assert.strictEqual(outcome.value.slice(node.valueSpan.start + 'ONE'.length), source.slice(node.valueSpan.end)); + }); + + test('two edits on different nodes both apply without offset drift (limit+1)', () => { + const doc = parseOk(source); + const idA = findField(doc, 'Alpha'); + const idB = findField(doc, 'Beta'); + const nodeA = doc.nodes.find((n) => n.id === idA); + const nodeB = doc.nodes.find((n) => n.id === idB); + + const staged1 = setFieldValue(doc, idA, 'ALPHA-NEW'); + const staged2 = setFieldValue(staged1.value, idB, 'BETA-NEW'); + const outcome = serialize(staged2.value); + assert.strictEqual(outcome.ok, true); + + // Region strictly between the two fields is untouched. + const between = source.slice(nodeA.valueSpan.end, nodeB.valueSpan.start); + const resultBetween = outcome.value.slice( + nodeA.valueSpan.start + 'ALPHA-NEW'.length, + nodeA.valueSpan.start + 'ALPHA-NEW'.length + between.length, + ); + assert.strictEqual(resultBetween, between); + + // Tail after Beta's original span is untouched, offset by both edits' delta. + const tailBefore = source.slice(nodeB.valueSpan.end); + const deltaA = 'ALPHA-NEW'.length - (nodeA.valueSpan.end - nodeA.valueSpan.start); + const deltaB = 'BETA-NEW'.length - (nodeB.valueSpan.end - nodeB.valueSpan.start); + const tailStart = nodeB.valueSpan.end + deltaA + deltaB; + assert.strictEqual(outcome.value.slice(tailStart), tailBefore); + }); +}); + +// ─── Row 6: duplicate/conflicting — same node edited twice ───────────────────── + +describe('row 6: a second edit to one node replaces the first', () => { + test('last write wins; span applied once', () => { + const source = ['---', 'title: Fixture', '---', '', '**Alpha:** one', ''].join('\n'); + const doc = parseOk(source); + const id = findField(doc, 'Alpha'); + const node = doc.nodes.find((n) => n.id === id); + + const staged1 = setFieldValue(doc, id, 'first'); + const staged2 = setFieldValue(staged1.value, id, 'second'); + + // Staged map collapses to exactly one entry for this id. + assert.strictEqual(staged2.value.staged.size, 1); + assert.strictEqual(staged2.value.staged.get(id), 'second'); + + const outcome = serialize(staged2.value); + assert.strictEqual(outcome.ok, true); + const expectedLength = source.length - (node.valueSpan.end - node.valueSpan.start) + 'second'.length; + assert.strictEqual(outcome.value.length, expectedLength); + assert.strictEqual(readNode(staged2.value, id).value, 'second'); + }); +}); + +// ─── Row 7: negative — id not from this doc ──────────────────────────────────── + +describe('row 7: setFieldValue refuses a node id this document did not mint', () => { + test('an id minted by a different PlanningDoc is refused', () => { + const source = ['---', 'title: Fixture', '---', '', '**Alpha:** one', ''].join('\n'); + const docA = parseOk(source); + const docB = parseOk(source); + const idFromB = findField(docB, 'Alpha'); + + const result = setFieldValue(docA, idFromB, 'x'); + assert.strictEqual(result.ok, false); + }); + + test('a wholly unknown id is refused', () => { + const source = ['---', 'title: Fixture', '---', '', '**Alpha:** one', ''].join('\n'); + const doc = parseOk(source); + const result = setFieldValue(doc, 'not-a-real-id', 'x'); + assert.strictEqual(result.ok, false); + }); +}); + +// ─── Row 8/9/10: ragged table — unreadable node + sibling readability ────────── + +function raggedTableSource() { + return [ + '---', + 'title: Fixture', + '---', + '', + '**Before:** sibling one', + '', + '| A | B |', + '|---|---|', + '| onlyone |', + '', + '**After:** sibling two', + '', + ].join('\n'); +} + +describe('row 8: an unreadable node does not make its siblings unreadable', () => { + test('the ragged table node carries an error; both bordering fields stay readable', () => { + const doc = parseOk(raggedTableSource()); + const tableNode = doc.nodes.find((n) => n.kind === 'table'); + assert.ok(tableNode); + assert.ok(tableNode.error); + assert.strictEqual(hasUnreadableNodes(doc), true); + + const beforeId = findField(doc, 'Before'); + const afterId = findField(doc, 'After'); + assert.strictEqual(readNode(doc, beforeId).ok, true); + assert.strictEqual(readNode(doc, afterId).ok, true); + + const tableRead = readNode(doc, tableNode.id); + assert.strictEqual(tableRead.ok, false); + }); +}); + +describe('row 9-10: serialize refuses on any unreadable node, edit or not', () => { + test('serialize refuses when any node is unreadable, with a staged edit', () => { + const doc = parseOk(raggedTableSource()); + const beforeId = findField(doc, 'Before'); + const staged = setFieldValue(doc, beforeId, 'edited'); + assert.strictEqual(staged.ok, true); + + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, false); + assert.strictEqual(outcome.reason, 'unreadable-nodes'); + assert.strictEqual(outcome.nodes.length, 1); + assert.strictEqual(outcome.nodes[0].kind, 'table'); + }); + + test('the refusal is about the document, not the mutation (no staged edit)', () => { + const doc = parseOk(raggedTableSource()); + assert.strictEqual(doc.staged.size, 0); + + const outcome = serialize(doc); + assert.strictEqual(outcome.ok, false); + assert.strictEqual(outcome.reason, 'unreadable-nodes'); + assert.strictEqual(outcome.nodes.length, 1); + assert.strictEqual(outcome.nodes[0].kind, 'table'); + }); +}); + +// ─── Row 11/12: empty / whitespace-only input ────────────────────────────────── + +describe('row 11-12: empty and whitespace-only documents', () => { + test('an empty document parses to no nodes, not to could-not-parse', () => { + const doc = parseOk(''); + assert.strictEqual(doc.nodes.length, 0); + }); + + test('a whitespace-only document is empty, not unparseable', () => { + const doc = parseOk(' \n\t\n \n'); + assert.strictEqual(doc.nodes.length, 0); + }); +}); + +// ─── Row 13/14: document-level malformed input ───────────────────────────────── + +describe('row 13-14: document-level Result failure', () => { + test('a document with no frontmatter terminator fails at the document level', () => { + const source = ['---', 'title: Fixture', 'this fence is never closed', ''].join('\n'); + const result = parsePlanningDoc(source, ARTIFACT); + assert.strictEqual(result.ok, false); + }); + + test('parsing a non-planning document fails at the document level', () => { + const source = ['---', 'title: Fixture', '---', '', '**Alpha:** one', ''].join('\n'); + // The identical source parses OK under a canonical artifact name... + const good = parsePlanningDoc(source, ARTIFACT); + assert.strictEqual(good.ok, true); + // ...and fails purely because of the artifact kind under a bogus one. + const bad = parsePlanningDoc(source, 'NOT-A-PLANNING-FILE.md'); + assert.strictEqual(bad.ok, false); + }); + + test('a non-markdown canonical planning file is refused at the document level', () => { + // config.json/state.json/milestone.lock are canonical per artifacts.cjs but + // carry no markdown grammar — this must be a document-level refusal, never + // a successful empty document (row 28 / design.md row 16). + assert.ok(isCanonicalPlanningFile('config.json')); + const result = parsePlanningDoc('{}', 'config.json'); + assert.strictEqual(result.ok, false); + }); +}); + +// ─── Row 15: CRLF round-trip ──────────────────────────────────────────────────── + +describe('row 15: CRLF documents round-trip byte-identically', () => { + const source = ['---', 'title: Fixture', '---', '', '**Alpha:** one', '**Beta:** two', ''].join('\r\n'); + + test('no-op serialize on a CRLF document is byte-identical', () => { + const doc = parseOk(source); + const outcome = serialize(doc); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value, source); + }); + + test('a single edit on a CRLF document preserves every byte outside valueSpan', () => { + const doc = parseOk(source); + const id = findField(doc, 'Alpha'); + const node = doc.nodes.find((n) => n.id === id); + const staged = setFieldValue(doc, id, 'ALPHA'); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value.slice(0, node.valueSpan.start), source.slice(0, node.valueSpan.start)); + assert.strictEqual(outcome.value.slice(node.valueSpan.start + 'ALPHA'.length), source.slice(node.valueSpan.end)); + }); +}); + +// ─── Row 16: hostile field value — markdown metacharacters ───────────────────── + +describe('row 16: a field value containing markdown metacharacters round-trips', () => { + test('pipes, backticks, and bold markers read back verbatim', () => { + const rawValue = 'a | b `code` **bold** end'; + const source = ['---', 'title: Fixture', '---', '', `**Data:** ${rawValue}`, ''].join('\n'); + const doc = parseOk(source); + const id = findField(doc, 'Data'); + assert.strictEqual(readNode(doc, id).value, rawValue); + + const node = doc.nodes.find((n) => n.id === id); + const staged = setFieldValue(doc, id, rawValue); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value.slice(0, node.valueSpan.start), source.slice(0, node.valueSpan.start)); + assert.strictEqual(outcome.value.slice(node.valueSpan.start + rawValue.length), source.slice(node.valueSpan.end)); + }); +}); + +// ─── Row 17-20: negative space ────────────────────────────────────────────────── + +describe('row 17-20: negative space (looks like the target, is not)', () => { + test('bold emphasis in prose is not a field label', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + 'This is **emphasis** in prose, not a field.', + '**Bold statement** without a colon.', + '', + ].join('\n'); + const doc = parseOk(source); + assert.strictEqual(doc.nodes.some((n) => n.kind === 'boldField'), false); + }); + + test('a field-shaped line inside a fenced block is not a node', () => { + const source = ['---', 'title: Fixture', '---', '', '```', '**Label:** value', '```', ''].join('\n'); + const doc = parseOk(source); + assert.strictEqual(findField(doc, 'Label'), null); + assert.strictEqual(doc.nodes.some((n) => n.kind === 'boldField'), false); + }); + + test('a field-shaped line inside an inline code span is not a node', () => { + const source = ['---', 'title: Fixture', '---', '', '`**Label:** value`', ''].join('\n'); + const doc = parseOk(source); + assert.strictEqual(findField(doc, 'Label'), null); + }); + + test('a horizontal rule is not a frontmatter terminator', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + '## Section', + '', + 'Some text.', + '', + '---', + '', + 'More text after the rule.', + ].join('\n'); + const doc = parseOk(source); + const frontmatterNodes = doc.nodes.filter((n) => n.kind === 'frontmatter'); + assert.strictEqual(frontmatterNodes.length, 1); + // The frontmatter span ends at the FIRST closing fence, well before the + // horizontal rule further down the document. + const secondRuleOffset = source.lastIndexOf('\n---\n'); + assert.ok(frontmatterNodes[0].span.end < secondRuleOffset); + }); +}); + +// ─── Row 21: hostile round-trip — pre-escaped table cell ─────────────────────── + +describe('row 21: an untouched escaped cell is not re-escaped', () => { + test('a pre-escaped pipe in a table cell is byte-identical with zero edits', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + '| Col A | Col B |', + '|---|---|', + '| has \\| pipe | plain |', + '', + ].join('\n'); + const doc = parseOk(source); + assert.strictEqual(hasUnreadableNodes(doc), false); + const outcome = serialize(doc); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value, source); + }); +}); + +// ─── Row 22/23: document-shaped fast-check properties ────────────────────────── +// +// Per CONTRIBUTING.md fixture-provenance (#2371): the generator below builds +// documents directly from arbitrary frontmatter/heading/label/value TEXT +// pieces assembled with `.join('\n')` — it never calls `serialize` (or any +// other function from the module under test) to produce its fixtures. Seeding +// from the module's own writer would make the document shape a constant and +// the property could never explore a shape the writer wouldn't itself emit. + +const safeLabelArb = fc + .stringMatching(/^[A-Za-z][A-Za-z0-9 ]{0,12}$/) + .map((s) => s.trim()) + .filter((s) => s.length > 0); + +// Values are deliberately widened to include the grammar's own metacharacters +// (em-dash, hyphen, `*_|[]#:`, bare \n/\r, non-ASCII letters) — this is the +// exact input space where the setFieldValue representability defects lived +// (forged sibling via \n; silent truncation on the trailing " — " separator). +// Labels stay narrow: they are a different, narrower grammar. +// +// safeValueArb feeds an INITIAL field's own literal text on ONE physical +// line of the generated document (`**label:** value`) — it excludes bare +// \n/\r because embedding a line break there would corrupt the generator's +// own one-line-per-field assumption (the field would not parse as a +// boldField at all, which is a generator bug, not a module defect). newValue +// is only ever passed AS AN ARGUMENT to setFieldValue, never spliced +// directly into document text, so it carries the full alphabet including +// \n/\r — exactly what setFieldValue must correctly refuse. +const HOSTILE_LINE_SAFE_CHARS = 'A-Za-z0-9 .,!?—\\-*_`|\\[\\]#:\\u00C0-\\u024F\\u3040-\\u30FF\\u4E00-\\u9FFF'; +const HOSTILE_VALUE_CHARS = `${HOSTILE_LINE_SAFE_CHARS}\\n\\r`; +const hostileLineSafeArb = (max) => fc.stringMatching(new RegExp(`^[${HOSTILE_LINE_SAFE_CHARS}]{0,${max}}$`)); +const hostileValueArb = (max) => fc.stringMatching(new RegExp(`^[${HOSTILE_VALUE_CHARS}]{0,${max}}$`)); + +const safeValueArb = hostileLineSafeArb(20); + +const fieldArb = fc.record({ label: safeLabelArb, value: safeValueArb }); + +const documentPiecesArb = fc.record({ + heading: fc + .stringMatching(/^[A-Za-z][A-Za-z0-9 ]{0,15}$/) + .map((s) => s.trim()) + .filter((s) => s.length > 0), + fields: fc.uniqueArray(fieldArb, { minLength: 1, maxLength: 5, selector: (r) => r.label.toLowerCase() }), + mutateIndex: fc.nat(), + newValue: hostileValueArb(25), +}); + +/** Assemble a document TEXT from arbitrary document-shaped pieces (never via + * the module's own serializer — see the #2371 note above). */ +function buildDocumentText({ heading, fields }) { + return [ + '---', + 'title: generated fixture', + '---', + '', + `## ${heading}`, + '', + ...fields.map((f) => `**${f.label}:** ${f.value}`), + '', + 'Trailing prose line unrelated to any field.', + ].join('\n'); +} + +describe('row 22-23: document-shaped fast-check properties', () => { + test('property: a single node mutation leaves every other byte identical', () => { + fc.assert( + fc.property(documentPiecesArb, (pieces) => { + const source = buildDocumentText(pieces); + const parsed = parsePlanningDoc(source, ARTIFACT); + assert.strictEqual(parsed.ok, true); + const doc = parsed.value; + + const idx = pieces.mutateIndex % pieces.fields.length; + const label = pieces.fields[idx].label; + const id = findField(doc, label); + assert.ok(id, `expected to find field ${JSON.stringify(label)}`); + + const node = doc.nodes.find((n) => n.id === id); + const staged = setFieldValue(doc, id, pieces.newValue); + if (!staged.ok) return; // a representability refusal is a valid outcome, not a failure + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + + const before = source.slice(0, node.valueSpan.start); + const after = source.slice(node.valueSpan.end); + assert.strictEqual(outcome.value.slice(0, node.valueSpan.start), before); + assert.strictEqual(outcome.value.slice(node.valueSpan.start + pieces.newValue.length), after); + }), + { seed: 20260921, numRuns: 200 }, + ); + }); + + test('property: serialize with no edits is the identity function', () => { + fc.assert( + fc.property(documentPiecesArb, (pieces) => { + const source = buildDocumentText(pieces); + const parsed = parsePlanningDoc(source, ARTIFACT); + assert.strictEqual(parsed.ok, true); + const outcome = serialize(parsed.value); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value, source); + }), + { seed: 20260921, numRuns: 200 }, + ); + }); +}); + +// ─── Row 31: representability — every ACCEPTED value round-trips identically ─── + +describe('row 31: every value setFieldValue accepts round-trips identically', () => { + test('property: every value setFieldValue ACCEPTS round-trips identically', () => { + fc.assert( + fc.property(documentPiecesArb, (pieces) => { + const source = buildDocumentText(pieces); + const parsed = parsePlanningDoc(source, ARTIFACT); + assert.strictEqual(parsed.ok, true); + const doc = parsed.value; + + const idx = pieces.mutateIndex % pieces.fields.length; + const label = pieces.fields[idx].label; + const id = findField(doc, label); + assert.ok(id, `expected to find field ${JSON.stringify(label)}`); + + const staged = setFieldValue(doc, id, pieces.newValue); + if (!staged.ok) return; // refusing is a PASS — the whole point of the representability check + + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + + const reparsed = parsePlanningDoc(outcome.value, ARTIFACT); + assert.strictEqual(reparsed.ok, true); + const reId = findField(reparsed.value, label); + assert.ok(reId, `expected to re-find field ${JSON.stringify(label)} after round-trip`); + const read = readNode(reparsed.value, reId); + assert.strictEqual(read.ok, true); + assert.strictEqual(read.value, pieces.newValue); + }), + { seed: 20260921, numRuns: 300 }, + ); + }); +}); + +// ─── Row 24: registry parity — Generative-Fix-Divergence guard ───────────────── + +describe('row 24: the artifact registry does not diverge from isCanonicalPlanningFile', () => { + test('every PLANNING_ARTIFACTS entry is a real canonical planning file', () => { + assert.ok(Array.isArray(PLANNING_ARTIFACTS)); + assert.ok(PLANNING_ARTIFACTS.length > 0); + for (const name of PLANNING_ARTIFACTS) { + assert.strictEqual(isCanonicalPlanningFile(name), true, `${name} should be canonical`); + } + }); + + test('PLANNING_ARTIFACTS is exactly the markdown subset of CANONICAL_EXACT', () => { + const expected = Array.from(CANONICAL_EXACT).filter((name) => name.endsWith('.md')); + assert.deepStrictEqual(new Set(PLANNING_ARTIFACTS), new Set(expected)); + }); + + test('a non-markdown canonical entry is excluded from PLANNING_ARTIFACTS', () => { + assert.ok(CANONICAL_EXACT.has('config.json')); + assert.strictEqual(PLANNING_ARTIFACTS.includes('config.json'), false); + }); +}); + +// ─── Row 25: one positive control per declared grammar ───────────────────────── + +describe('row 25: every declared grammar has a positive control', () => { + test('frontmatter, section, boldField, table, and checklist each parse to a node of that kind', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + '## A Section', + '', + '**Field:** a value', + '', + '| Col A | Col B |', + '|---|---|', + '| x | y |', + '', + '- [ ] todo one', + '- [x] todo two', + '', + ].join('\n'); + const doc = parseOk(source); + const kinds = new Set(doc.nodes.map((n) => n.kind)); + for (const expectedKind of ['frontmatter', 'section', 'boldField', 'table', 'checklist']) { + assert.ok(kinds.has(expectedKind), `expected a ${expectedKind} node; got kinds: ${[...kinds].join(', ')}`); + } + }); +}); + +// ─── Row 26: unicode heading / label — byte-honest offsets ───────────────────── + +describe('row 26: unicode labels and headings keep byte-honest offsets', () => { + test('a unicode field label and value parse and round-trip correctly', () => { + const source = ['---', 'title: Fixture', '---', '', '## Resume 日本語', '', '**日本語:** 値', ''].join( + '\n', + ); + const doc = parseOk(source); + const id = findField(doc, '日本語'); + assert.ok(id); + assert.strictEqual(readNode(doc, id).value, '値'); + + const node = doc.nodes.find((n) => n.id === id); + const staged = setFieldValue(doc, id, '新しい値'); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value.slice(0, node.valueSpan.start), source.slice(0, node.valueSpan.start)); + assert.strictEqual( + outcome.value.slice(node.valueSpan.start + '新しい値'.length), + source.slice(node.valueSpan.end), + ); + }); +}); + +// ─── Row 27: bounded-size — large document splices without offset corruption ─── + +describe('row 27: a large document splices without offset corruption', () => { + test('a document with thousands of lines still splices exactly one span', () => { + const lineCount = 4000; + const lines = ['---', 'title: Large Fixture', '---', '']; + for (let i = 0; i < lineCount; i++) { + lines.push(`Prose line number ${i} filling out the document body.`); + } + lines.push('**Target:** the value to mutate'); + for (let i = 0; i < lineCount; i++) { + lines.push(`More prose line number ${i} after the target field.`); + } + const source = lines.join('\n'); + + const doc = parseOk(source); + const id = findField(doc, 'Target'); + assert.ok(id); + const node = doc.nodes.find((n) => n.id === id); + + const staged = setFieldValue(doc, id, 'MUTATED'); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value.slice(0, node.valueSpan.start), source.slice(0, node.valueSpan.start)); + assert.strictEqual(outcome.value.slice(node.valueSpan.start + 'MUTATED'.length), source.slice(node.valueSpan.end)); + assert.strictEqual(outcome.value.length, source.length - (node.valueSpan.end - node.valueSpan.start) + 'MUTATED'.length); + }); +}); + +// ─── Row 29: negative — a line-break value must not become a forged sibling ─── + +describe('row 29: setFieldValue refuses a value carrying a line break', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + '**Plans:** initial value', + '**Owner:** alice', + '', + ].join('\n'); + + test('a value containing \\n is refused', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + const result = setFieldValue(doc, id, '1/1\n**Owner:** mallory'); + assert.strictEqual(result.ok, false); + }); + + test('a value containing a bare \\r is refused', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + const result = setFieldValue(doc, id, '1/1\r**Owner:** mallory'); + assert.strictEqual(result.ok, false); + }); + + test('a value containing \\r\\n is refused', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + const result = setFieldValue(doc, id, '1/1\r\n**Owner:** mallory'); + assert.strictEqual(result.ok, false); + }); + + test('the negative proof: a refused write leaves the original document untouched', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + const result = setFieldValue(doc, id, '1/1\n**Owner:** mallory'); + assert.strictEqual(result.ok, false); + + const outcome = serialize(doc); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value, source); + assert.strictEqual(outcome.value.includes('mallory'), false); + assert.strictEqual(readNode(doc, findField(doc, 'Owner')).value, 'alice'); + }); + + test('legitimate values still stage successfully: plain, empty, **, and |', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + + for (const value of ['plain value', '', '**bold marker**', 'a | b']) { + const result = setFieldValue(doc, id, value); + assert.strictEqual(result.ok, true); + } + }); + + test('a full round trip with a legitimate value leaves the sibling field intact exactly once', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + const staged = setFieldValue(doc, id, '2/2'); + assert.strictEqual(staged.ok, true); + const outcome = serialize(staged.value); + assert.strictEqual(outcome.ok, true); + + const ownerMatches = outcome.value.match(/\*\*Owner:\*\*/g); + assert.strictEqual(ownerMatches.length, 1); + const ownerId = findField(staged.value, 'Owner'); + assert.strictEqual(readNode(staged.value, ownerId).value, 'alice'); + }); +}); + +// ─── Row 30: negative — a value carrying the trailing separator ─────────────── + +describe('row 30: setFieldValue refuses a value containing the trailing separator', () => { + const source = [ + '---', + 'title: Fixture', + '---', + '', + '**Plans:** initial value', + '**Owner:** alice', + '', + ].join('\n'); + + test('a value containing the trailing " — " separator is refused', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + const result = setFieldValue(doc, id, `sneaky ${EM_DASH} annotation`); + assert.strictEqual(result.ok, false); + }); + + test('after refusal, the document serializes byte-identically to the source', () => { + const doc = parseOk(source); + const id = findField(doc, 'Plans'); + const result = setFieldValue(doc, id, `sneaky ${EM_DASH} annotation`); + assert.strictEqual(result.ok, false); + + const outcome = serialize(doc); + assert.strictEqual(outcome.ok, true); + assert.strictEqual(outcome.value, source); + }); +});