fix(#2703): strip GSD-2 frontmatter with the canonical parser (#3027)

* test(#2703): failing-first coverage for CRLF frontmatter strip in SUMMARY.md

Drives the exported buildPlanningArtifacts seam. Rows for CRLF/LF parity,
stacked blocks and a leading BOM fail against the current hand-rolled
regex; the negative-space rows pin behavior that must not change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2703): strip GSD-2 frontmatter with the canonical parser

buildSummaryMd matched the closing delimiter with a hardcoded bare \n, so a
CRLF-authored task summary never matched and fell through to the raw-passthrough
branch. The function then prepended its own block, emitting a SUMMARY.md with two
stacked frontmatter blocks and no warning.

Delegates to stripFrontmatter from frontmatter.cts -- the canonical, line-ending
tolerant primitive this repo already deduplicated once (#2143) -- instead of
adding another hand-rolled variant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2703): strip only the first frontmatter block in gsd2 import

Adversarial review caught a regression in the first cut: stripFrontmatter
loops by design, so a summary body opening with a thematic-break-delimited
section (--- / heading / ---) had that section silently deleted. The old
pre-#2703 regex preserved it, so shipping the loop would have traded one
silent corruption for another.

Adds an explicit { once } option to the canonical primitive -- default
behavior and the two existing callers are unchanged -- and has buildSummaryMd
opt in. A GSD-2 summary is an arbitrary user document, not a GSD artifact with
a known doubling failure mode, so a second block there is body content.

This also makes the acceptance criterion exact: CRLF now produces the same
result LF already produced, rather than a new result for both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2703): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-08-03 15:21:30 -04:00
committed by GitHub
parent 8ad7845f16
commit fd64389616
5 changed files with 230 additions and 9 deletions

View File

@@ -0,0 +1,5 @@
---
type: Fixed
pr: 3027
---
**GSD-2 import no longer duplicates frontmatter in the generated SUMMARY.md** — importing a GSD-2 project whose task summaries were authored with CRLF line endings emitted the original GSD-2 frontmatter a second time, as body text, below the new one. Stripping now goes through the canonical line-ending-tolerant parser. (#2703)

View File

@@ -642,22 +642,29 @@ const FRONTMATTER_SCHEMAS: Record<string, { required: string[]; requiredValues?:
};
/**
* Strip ALL frontmatter blocks from the start of `content`.
* Strip frontmatter blocks from the start of `content`.
*
* Handles CRLF line endings and multiple stacked blocks (corruption
* recovery): greedily strips consecutive `---...---` blocks separated by
* optional whitespace, so a doubled/tripled frontmatter header (e.g. from a
* botched merge) is fully removed, not just the first block.
* Handles CRLF line endings and, by default, multiple stacked blocks
* (corruption recovery): greedily strips consecutive `---...---` blocks
* separated by optional whitespace, so a doubled/tripled frontmatter header
* (e.g. from a botched merge) is fully removed, not just the first block.
*
* Pass `{ once: true }` to stop after the first block. Callers whose input is
* an arbitrary user-authored document — rather than a GSD artefact with a
* known doubling failure mode — need this: a body that opens with a
* thematic-break-delimited section is lexically indistinguishable from a
* second frontmatter block, and the greedy loop deletes it silently (#2703).
*
* Canonical home for this primitive (#2143 audit dedup): previously
* duplicated byte-identically in both `state.cts` and `state-transition.cts`.
*/
function stripFrontmatter(content: string): string {
function stripFrontmatter(content: string, opts: { once?: boolean } = {}): string {
let result = content;
while (true) {
const stripped = result.replace(/^\s*---\r?\n[\s\S]*?\r?\n---\s*/, '');
if (stripped === result) break;
result = stripped;
if (opts.once) break;
}
return result;
}

View File

@@ -28,8 +28,11 @@ import { realClock } from './clock.cjs';
import coreUtilsMod = require('./core-utils.cjs');
// eslint-disable-next-line @typescript-eslint/no-require-imports
import ioMod = require('./io.cjs');
// eslint-disable-next-line @typescript-eslint/no-require-imports -- frontmatter.cjs is an export= CommonJS module
import frontmatterMod = require('./frontmatter.cjs');
const { output } = ioMod;
const { transliterateForSlug } = coreUtilsMod;
const { stripFrontmatter } = frontmatterMod;
// ─── Types ───────────────────────────────────────────────────────────────────
@@ -290,9 +293,22 @@ function buildPlanMd(task: TaskInfo, phasePrefix: string, planPrefix: string, ph
*/
function buildSummaryMd(task: TaskInfo, phasePrefix: string, planPrefix: string): string {
const raw = task.summary || '';
// Strip GSD-2 frontmatter block (--- ... ---) if present
const bodyMatch = raw.match(/^---[\s\S]*?---\n+([\s\S]*)$/);
const body = bodyMatch ? bodyMatch[1].trim() : raw.trim();
// Strip the GSD-2 frontmatter block via the canonical primitive (#2703). The
// previous local regex required a bare `\n` after the closing `---`, so a
// CRLF-authored summary never matched, fell through to the untouched-raw
// branch, and had its frontmatter emitted a second time inside the body of
// the document this function then wrapped in a fresh v1 block.
//
// `extractFrontmatter` — which the issue names — returns only the parsed
// object and never the body, so it cannot serve this call site;
// `stripFrontmatter` is the same module's canonical body primitive.
//
// `once` is load-bearing. A GSD-2 summary is an arbitrary user-authored
// document, not a GSD artefact with a known frontmatter-doubling failure
// mode, so a body opening with a thematic-break-delimited section
// (`---` / heading / `---`) is far likelier than a corrupt second header —
// and the default greedy loop would delete it without a trace.
const body = stripFrontmatter(raw, { once: true }).trim();
return [
'---',

View File

@@ -21,6 +21,7 @@ const {
extractFrontmatter,
reconstructFrontmatter,
spliceFrontmatter,
stripFrontmatter,
noOpObjectListSetError,
parseMustHavesBlock,
FRONTMATTER_SCHEMAS,
@@ -1357,3 +1358,54 @@ describe('noOpObjectListSetError (#1660)', () => {
assert.ok(msg.includes('Edit the file directly'), msg);
});
});
// ─── stripFrontmatter ─────────────────────────────────────────────────────────
describe('stripFrontmatter', () => {
const stacked = ['---', 'a: 1', '---', '---', 'b: 2', '---', '', 'Real body.'].join('\n');
test('strips a single block', () => {
assert.strictEqual(stripFrontmatter(['---', 'a: 1', '---', '', 'Body.'].join('\n')), 'Body.');
});
test('is CRLF-tolerant', () => {
const crlf = ['---', 'a: 1', '---', '', 'Body.'].join('\r\n');
assert.strictEqual(stripFrontmatter(crlf), 'Body.');
});
test('defaults to stripping every stacked block (corruption recovery)', () => {
assert.strictEqual(stripFrontmatter(stacked), 'Real body.');
});
test('an omitted options argument keeps the greedy default', () => {
// Back-compat: state.cts and state-transition.cts call this with one arg.
assert.strictEqual(stripFrontmatter(stacked, {}), 'Real body.');
});
test('once: true stops after the first block', () => {
assert.strictEqual(
stripFrontmatter(stacked, { once: true }),
['---', 'b: 2', '---', '', 'Real body.'].join('\n'),
);
});
test('once: false is the greedy default', () => {
assert.strictEqual(stripFrontmatter(stacked, { once: false }), 'Real body.');
});
test('returns content unchanged when there is no frontmatter', () => {
const plain = ['Just prose.', '', 'More prose.'].join('\n');
assert.strictEqual(stripFrontmatter(plain), plain);
assert.strictEqual(stripFrontmatter(plain, { once: true }), plain);
});
test('leaves an unterminated block alone under both modes', () => {
const unterminated = ['---', 'a: 1', 'b: 2'].join('\n');
assert.strictEqual(stripFrontmatter(unterminated), unterminated);
assert.strictEqual(stripFrontmatter(unterminated, { once: true }), unterminated);
});
test('empty string round-trips', () => {
assert.strictEqual(stripFrontmatter(''), '');
});
});

View File

@@ -9,6 +9,7 @@ const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { createTempDir, cleanup, runGsdTools } = require('./helpers.cjs');
const fc = require('./helpers/fast-check-setup.cjs');
const {
parseSlicesFromRoadmap,
@@ -578,3 +579,143 @@ describe('gsd-tools from-gsd2 CLI', () => {
assert.ok(!fs.existsSync(path.join(tmpDir, '.planning', 'phases', '02-auth-system', '02-01-SUMMARY.md')));
});
});
// ─── SUMMARY.md frontmatter stripping (#2703) ──────────────────────────────
/**
* The emitted artifact under test. `buildSummaryMd` is module-private;
* `buildPlanningArtifacts` is its only caller and is the shape `cmdFromGsd2`
* drives in production, so every row below asserts through that seam rather
* than against the private function.
*/
const SUMMARY_KEY = 'phases/01-setup/01-01-SUMMARY.md';
function emitSummary(summary) {
return buildPlanningArtifacts({
projectContent: '# P\n',
requirements: null,
milestones: [{
id: 'M001',
title: 'Foundation',
slices: [{
done: true,
id: 'S01',
title: 'Setup',
tasks: [{ done: true, id: 'T01', title: 'Init', description: '', mustHaves: [], summary }],
}],
}],
}).get(SUMMARY_KEY);
}
/** Re-encode an LF document with CRLF line endings. */
const crlf = (s) => s.replace(/\n/g, '\r\n');
/** The whole document `buildSummaryMd` is expected to emit for a given body. */
const expectedDoc = (body) => ['---', 'phase: "01"', 'plan: "01"', '---', '', body, ''].join('\n');
/**
* Assert that `summary` — and its CRLF re-encoding — both emit the document
* built from `body`. CRLF/LF parity is the invariant this whole block exists
* to protect, so every row asserts it the same way rather than restating the
* pair by hand.
*/
function assertBothEncodings(summary, body) {
assert.strictEqual(emitSummary(summary), expectedDoc(body));
assert.strictEqual(emitSummary(crlf(summary)), expectedDoc(crlf(body)));
}
describe('buildSummaryMd frontmatter stripping (#2703)', () => {
test('strips GSD-2 frontmatter identically under CRLF and LF (#2703)', () => {
const summary = ['---', 'task: T01', 'status: done', '---', '', 'The task body.'].join('\n');
// The bug: the CRLF emission carried a second, unstripped frontmatter block.
assert.strictEqual(emitSummary(crlf(summary)), emitSummary(summary));
assertBothEncodings(summary, 'The task body.');
});
test('strips a single frontmatter block', () => {
assertBothEncodings(['---', 'task: T01', '---', '', 'Body one.'].join('\n'), 'Body one.');
});
test('passes through a summary that has no frontmatter', () => {
// Body line endings are passed through untouched — stripping frontmatter
// must not silently re-encode the author's prose.
const summary = ['Just prose.', '', 'No frontmatter here.'].join('\n');
assertBothEncodings(summary, summary);
});
test('emits no SUMMARY.md at all for an empty summary', () => {
// buildPlanningArtifacts guards on `task.done && task.summary`, so an empty
// summary produces no artifact rather than a default-bodied one.
assert.strictEqual(emitSummary(''), undefined);
});
test('falls back to the migration default for a whitespace-only summary', () => {
assert.strictEqual(emitSummary(' \n \n'), expectedDoc('Task completed (migrated from GSD-2).'));
});
test('preserves a lone thematic break that opens the body', () => {
const summary = ['---', 'task: T01', '---', '', '---', '', 'Body after a rule.'].join('\n');
assertBothEncodings(summary, ['---', '', 'Body after a rule.'].join('\n'));
});
test('does not eat a thematic-break-delimited section that opens the body', () => {
// Regression guard. The canonical stripper's DEFAULT greedy loop deletes
// `Some Heading` outright, because `---` / text / `---` is lexically a
// second frontmatter block. buildSummaryMd therefore passes `once: true`.
// Caught by adversarial review of the first cut of this fix; the old
// pre-#2703 regex preserved this content, so eating it would have been a
// silent regression shipped alongside the CRLF fix.
const summary = ['---', 'task: T01', '---', '---', 'Some Heading', '---', '', 'Body content below.'].join('\n');
assertBothEncodings(summary, ['---', 'Some Heading', '---', '', 'Body content below.'].join('\n'));
});
test('strips only the first block when two frontmatter-shaped blocks lead', () => {
// Same `once` semantics stated for the YAML-shaped case: a GSD-2 summary is
// an arbitrary user document, so a second block is body content, not a
// corrupt duplicate header to be recovered from.
const summary = ['---', 'a: 1', '---', '---', 'b: 2', '---', '', 'Real body.'].join('\n');
assertBothEncodings(summary, ['---', 'b: 2', '---', '', 'Real body.'].join('\n'));
});
test('does not strip a --- that appears mid-body', () => {
const summary = ['Intro prose.', '', '---', '', 'More prose.'].join('\n');
assertBothEncodings(summary, summary);
});
test('leaves an unterminated frontmatter block as body text', () => {
const summary = ['---', 'task: T01', 'status: done'].join('\n');
assertBothEncodings(summary, summary);
});
test('strips frontmatter behind a leading BOM', () => {
const summary = '' + ['---', 'task: T01', '---', '', 'Body.'].join('\n');
assertBothEncodings(summary, 'Body.');
});
test('property: CRLF and LF summaries emit the same SUMMARY.md (#2703)', () => {
// Body lines are drawn from a charset with no `-`, so a generated body can
// never accidentally form a second frontmatter block.
const yamlKey = fc.stringMatching(/^[a-z][a-z0-9_]{0,10}$/);
const yamlValue = fc.stringMatching(/^[a-zA-Z0-9 ._]{1,20}$/);
const bodyLine = fc.stringMatching(/^[a-zA-Z0-9 ._]{1,30}$/);
fc.assert(
fc.property(
fc.array(fc.tuple(yamlKey, yamlValue), { minLength: 1, maxLength: 4 }),
fc.array(bodyLine, { minLength: 1, maxLength: 4 }),
(pairs, bodyLines) => {
const doc = [
'---',
...pairs.map(([k, v]) => `${k}: ${v}`),
'---',
'',
...bodyLines,
].join('\n');
const fromLf = emitSummary(doc);
const fromCrlf = emitSummary(crlf(doc));
assert.strictEqual(fromCrlf.replace(/\r\n/g, '\n'), fromLf);
},
),
);
});
});