* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local>
229 lines
10 KiB
JavaScript
229 lines
10 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* Property-based tests for src/context-predicates.cts (compiled to
|
|
* gsd-core/bin/lib/context-predicates.cjs).
|
|
*
|
|
* Document-shaped generators (CONTRIBUTING.md "Fixture provenance #2371"):
|
|
* these generators build arbitrary markdown documents out of prose lines,
|
|
* fences of varying tick-length, list items with varying markers, and
|
|
* predicate-shaped / predicate-lookalike lines — they are NOT seeded from
|
|
* this module's own writer/serializer. Seeding a property generator from the
|
|
* code under test's own render function make the document shape a constant
|
|
* and the property unable to fail; see 50-test-matrix.md Step 1.
|
|
*
|
|
* Deterministic per CONTRIBUTING.md: seed and numRuns are pinned by
|
|
* tests/helpers/fast-check-setup.cjs (seed 42, numRuns 200); failures print
|
|
* replay data via fast-check's own counterexample + seed reporting.
|
|
*/
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fc = require('./helpers/fast-check-setup.cjs');
|
|
|
|
const { parsePredicates, selectPredicates, buildIndex } = require('../gsd-core/bin/lib/context-predicates.cjs');
|
|
|
|
// ─── Document-shaped generators ────────────────────────────────────────────
|
|
|
|
// A predicate-shaped line: `CLASS.subkey=value` in one of the recognized
|
|
// declaration forms. `asDeclared` controls whether the caller wants this
|
|
// specific fixture to be a genuinely-recognized declaration (bare or `- `
|
|
// list item — the two forms even today's implementation accepts) so property
|
|
// H1/H4 can reason about "predicates the parser actually extracts" without
|
|
// depending on the A4-A7 defects under test elsewhere.
|
|
const idClassArb = fc.stringMatching(/^[A-Z][A-Z0-9_-]{0,8}$/);
|
|
const idSubkeyArb = fc.stringMatching(/^[A-Za-z0-9_-]{1,8}$/);
|
|
const valueArb = fc.stringMatching(/^[A-Za-z0-9 _.-]{1,20}$/);
|
|
|
|
const declaredPredicateLineArb = fc
|
|
.tuple(idClassArb, fc.option(idSubkeyArb, { nil: undefined }), valueArb, fc.boolean())
|
|
.map(([klass, subkey, value, asListItem]) => {
|
|
const id = subkey ? `${klass}.${subkey}` : klass;
|
|
const inner = `\`${id}=${value}\``;
|
|
return { text: asListItem ? `- ${inner}` : inner, id, klass, value };
|
|
});
|
|
|
|
// A prose line that is NOT a predicate declaration: plain text, a heading, a
|
|
// blockquote line, a table row, or an inline mid-prose backtick reference.
|
|
const proseLineArb = fc.oneof(
|
|
fc.stringMatching(/^[A-Za-z0-9 .,'"()-]{0,40}$/),
|
|
idClassArb.map((k) => `# ${k} heading`),
|
|
idClassArb.map((k) => `> quoting ${k}`),
|
|
idClassArb.map((k) => `| cell | \`${k}.x=y\` |`),
|
|
declaredPredicateLineArb.map((p) => `see ${p.text.replace(/^- /, '')} for details`),
|
|
);
|
|
|
|
const fenceTickArb = fc.constantFrom('```', '~~~', '````');
|
|
|
|
// A whole document assembled from a mix of prose lines and declared
|
|
// predicate lines, optionally wrapping a contiguous run in a fence.
|
|
function documentArb() {
|
|
return fc
|
|
.array(fc.oneof({ arbitrary: declaredPredicateLineArb, weight: 2 }, { arbitrary: proseLineArb, weight: 3 }), {
|
|
minLength: 0,
|
|
maxLength: 12,
|
|
})
|
|
.map((items) => items);
|
|
}
|
|
|
|
// ─── H1: index preserves every parsed predicate ────────────────────────────
|
|
|
|
describe('property: parse <-> index contract', () => {
|
|
test('indexPreservesEveryParsedPredicate', () => {
|
|
fc.assert(
|
|
fc.property(documentArb(), (items) => {
|
|
const md = items.map((it) => (typeof it === 'string' ? it : it.text)).join('\n');
|
|
const parsed = parsePredicates(md);
|
|
const index = buildIndex(parsed.predicates);
|
|
|
|
assert.equal(index.count, parsed.predicates.length);
|
|
|
|
const parsedIds = parsed.predicates.map((p) => p.id).sort();
|
|
const indexIds = index.predicates.map((p) => p.id).sort();
|
|
assert.deepEqual(indexIds, parsedIds, 'buildIndex must invent or lose no predicate id');
|
|
|
|
for (const p of parsed.predicates) {
|
|
const inIndex = index.predicates.find((ip) => ip.id === p.id && ip.value === p.value);
|
|
assert.ok(inIndex, `predicate ${p.id}=${p.value} from the parse must appear in the index`);
|
|
}
|
|
}),
|
|
);
|
|
});
|
|
|
|
// ─── H2: buildIndex is deterministic and order-independent ────────────────
|
|
|
|
test('indexSerializationIsOrderIndependentAndDeterministic', () => {
|
|
fc.assert(
|
|
fc.property(
|
|
fc.array(declaredPredicateLineArb, { minLength: 0, maxLength: 10 }),
|
|
(decls) => {
|
|
// Dedupe by id so this property is not entangled with duplicate
|
|
// semantics (covered separately by the E-row unit tests) — the
|
|
// property under test here is pure ordering independence.
|
|
const seen = new Set();
|
|
const uniqueDecls = decls.filter((d) => (seen.has(d.id) ? false : (seen.add(d.id), true)));
|
|
|
|
const forwardMd = uniqueDecls.map((d) => d.text).join('\n');
|
|
const reverseMd = uniqueDecls
|
|
.slice()
|
|
.reverse()
|
|
.map((d) => d.text)
|
|
.join('\n');
|
|
|
|
const forwardIndex = buildIndex(parsePredicates(forwardMd).predicates);
|
|
const reverseIndex = buildIndex(parsePredicates(reverseMd).predicates);
|
|
|
|
assert.deepEqual(forwardIndex, reverseIndex, 'index must not depend on source declaration order');
|
|
|
|
// Determinism: building twice from the same parsed predicates must
|
|
// produce byte-identical JSON serialization.
|
|
const parsed = parsePredicates(forwardMd);
|
|
const a = JSON.stringify(buildIndex(parsed.predicates));
|
|
const b = JSON.stringify(buildIndex(parsed.predicates));
|
|
assert.equal(a, b);
|
|
},
|
|
),
|
|
);
|
|
});
|
|
|
|
// ─── H3: fencing a region never increases the predicate count ─────────────
|
|
|
|
test('fencingRegionNeverIncreasesPredicateCount', () => {
|
|
fc.assert(
|
|
fc.property(
|
|
fc.array(declaredPredicateLineArb, { minLength: 1, maxLength: 6 }),
|
|
fc.nat({ max: 5 }),
|
|
fenceTickArb,
|
|
(decls, wrapAt, fence) => {
|
|
const lines = decls.map((d) => d.text);
|
|
const before = parsePredicates(lines.join('\n')).predicates.length;
|
|
|
|
const cut = Math.min(wrapAt, lines.length);
|
|
const fenced = [...lines.slice(0, cut), fence, ...lines.slice(cut), fence].join('\n');
|
|
const after = parsePredicates(fenced).predicates.length;
|
|
|
|
assert.ok(after <= before, `fencing must never increase the parsed count (before=${before}, after=${after})`);
|
|
},
|
|
),
|
|
);
|
|
});
|
|
});
|
|
|
|
// ─── H5: comment/fence mutual precedence (DEFECT.CONTEXT-PREDICATES-COMMENT-
|
|
// FENCE-BLIND, #2928 review) — a predicate genuinely OUTSIDE a wrapper
|
|
// (comment or fence) is always parsed live, and a predicate genuinely INSIDE
|
|
// it is never parsed live, even when the wrapper's own content contains
|
|
// tokens that LOOK LIKE the OTHER construct (fence delimiters inside a
|
|
// comment, or comment tokens inside a fence — the two directions the
|
|
// interleaved single-pass in `computeSkippedLineFlags` must both get right).
|
|
// Document-shaped generator: wraps a marked "inside" predicate between an
|
|
// open/close pair of ONE kind, salted with lookalike noise from the OTHER
|
|
// kind, with unrelated real predicates before/after the wrapper. ──────────
|
|
|
|
// Noise lines that look like the OTHER construct's tokens, keyed by which
|
|
// kind is being used as the OUTER wrapper for a given run.
|
|
const FENCE_NOISE_LINES = ['<!-- not a real comment, just fence content', 'text mentioning --> mid-line', '<!-- nested-looking'];
|
|
const COMMENT_NOISE_LINES = ['```', '~~~', '``` info string'];
|
|
|
|
const wrapperKindArb = fc.constantFrom('fence', 'comment');
|
|
|
|
describe('property: comment/fence mutual precedence', () => {
|
|
test('predicateOutsideWrapperIsAlwaysLiveAndInsideIsNeverLive', () => {
|
|
fc.assert(
|
|
fc.property(
|
|
wrapperKindArb,
|
|
fc.array(fc.constantFrom(0, 1, 2), { minLength: 0, maxLength: 3 }),
|
|
fc.array(fc.constantFrom(0, 1, 2), { minLength: 0, maxLength: 3 }),
|
|
(wrapperKind, beforeNoiseIdx, afterNoiseIdx) => {
|
|
const noisePool = wrapperKind === 'fence' ? FENCE_NOISE_LINES : COMMENT_NOISE_LINES;
|
|
const open = wrapperKind === 'fence' ? '```' : '<!-- wrapper open';
|
|
const close = wrapperKind === 'fence' ? '```' : '-->';
|
|
|
|
const lines = [
|
|
'`OUTSIDE_BEFORE=1`',
|
|
open,
|
|
...beforeNoiseIdx.map((i) => noisePool[i]),
|
|
'`INSIDE_MARKER=2`',
|
|
...afterNoiseIdx.map((i) => noisePool[i]),
|
|
close,
|
|
'`OUTSIDE_AFTER=3`',
|
|
];
|
|
|
|
const md = lines.join('\n');
|
|
const r = parsePredicates(md);
|
|
const ids = new Set(r.predicates.map((p) => p.id));
|
|
|
|
assert.ok(ids.has('OUTSIDE_BEFORE'), `OUTSIDE_BEFORE must always be live (wrapper=${wrapperKind})`);
|
|
assert.ok(ids.has('OUTSIDE_AFTER'), `OUTSIDE_AFTER must always be live (wrapper=${wrapperKind})`);
|
|
assert.ok(
|
|
!ids.has('INSIDE_MARKER'),
|
|
`INSIDE_MARKER must never be live inside a real ${wrapperKind}, even with ${
|
|
wrapperKind === 'fence' ? 'comment' : 'fence'
|
|
}-lookalike noise around it`,
|
|
);
|
|
},
|
|
),
|
|
);
|
|
});
|
|
});
|
|
|
|
describe('property: selectPredicates subset invariant', () => {
|
|
test('selectorReturnsOnlyMatchingSubset', () => {
|
|
fc.assert(
|
|
fc.property(fc.array(declaredPredicateLineArb, { minLength: 0, maxLength: 10 }), idClassArb, (decls, klass) => {
|
|
const md = decls.map((d) => d.text).join('\n');
|
|
const all = parsePredicates(md).predicates;
|
|
const selected = selectPredicates(all, { klass });
|
|
|
|
for (const p of selected) {
|
|
assert.ok(
|
|
all.includes(p),
|
|
'every selected predicate must be a reference from the original array (subset, not a copy with invented members)',
|
|
);
|
|
assert.equal(p.klass, klass, `selectPredicates({klass}) must only return predicates whose klass === ${klass}`);
|
|
}
|
|
}),
|
|
);
|
|
});
|
|
});
|