fix(#2944): remove the catastrophic-backtracking regex from the ADR-1671 example (#2950)

* fix(#2944): remove the catastrophic-backtracking regex from the example

The non-shipping Option-E reference example carried its own copy of the
predicate-id regex, which nested a dot-containing character class inside a
dot-prefixed repeat. A run of N consecutive dots therefore had exponentially
many partitions. Measured on next before this change: 30 dots 54ms, 35 66ms,
40 807ms — so roughly 55-60 dots hangs for hours.

Not exploitable where it sits: the example is outside tsconfig.build.json,
outside the npm package files list, outside the installer and outside tests,
so no build step or CI job parses anything with it. Fixed because the entire
point of a reference example is that people copy it forward, and ADR-1671
presents this one as the pattern for the platform.

Ports the linear per-segment validation that #2928 gave the production module,
so the two copies agree: both parse the real CONTEXT.md to 415 predicates
across 20 classes with 0 duplicates. Doubled-dot ids are now rejected here
too, matching production, and the grammar comment records it.

Also refreshes the example's committed index, which #2928 made stale when it
removed the duplicate predicate from CONTEXT.md.

Closes #2944

* test(#2944): guard predicate-index sync and example/production parity

Two regression tests for the two defects in this PR.

Index sync: asserts the committed docs/CONTEXT-INDEX.json equals a fresh parse
of CONTEXT.md, naming any diverging predicate ids. The merge race that reddened
next was invisible to both PRs involved and only surfaced on the next PR to run
lint:ci; this puts the same check inside the suite, which runs on every PR, and
a mutation test proves the assertion is not vacuous.

Example/production parity: asserts both copies of the parser report the same
count, classes and duplicates for the real CONTEXT.md, and agree verdict-for-
verdict over a table of id shapes. The divergence WAS the bug — production went
linear-time while the example kept the backtracking regex, with nothing
asserting they agreed. Also pins the example rejecting a 60-dot id, with the
clean rejection as the binding assertion and wall-clock only as a smoke check.

Notes a real tension rather than hiding it: ADR-1671 says the example sits
outside tests/, and this imports it. The ADR's intent is that the example is
not compiled, packaged or installed — not that it may silently rot. A parity
guard does not ship it. The file states this so a reviewer can object.

* fix(#2944): address both isolated review passes

Two independent reviewers (correctness and security axes, neither the author).
Security found nothing — it measured linearity to 100k chars across dots,
hyphens, underscores and mixed classes, and showed prototype pollution is
structurally unreachable because the first-segment pattern forbids
lowercase and underscore-leading ids. The correctness pass found three
blockers, all real.

Blocker: the parity test violated ADR-1671 verbatim. The ADR lists FOUR
exclusions for the reference example, the fourth being the CI test suite, and
the test imported it from tests/ while its own justification comment cited only
three -- constructing a rationale around the exclusion it broke. Moved to
scripts/lint-example-parser-parity.cjs wired into lint:ci; a lint script is not
the test suite, so the exclusion stands. The test file keeps only the
docs/CONTEXT-INDEX.json sync check.

Blocker: the mutation test leaked its temp dir. Its callback took no `t`, so a
failing assertion skipped the bare cleanup call. Now registered via t.after(),
matching the convention adr-index-gate.test.cjs documents.

Blocker: the example's own committed index carries the identical merge-race
staleness this PR fixes for the production one, and nothing guarded it.
Deliberately NOT fixed by wiring the example's --check into CI: that artifact
bakes line numbers, so it re-drifts on any unrelated CONTEXT.md line shift --
exactly ADR-1671 open question 4 -- and would make CI routinely red. The new
lint asserts the line-INDEPENDENT facts instead: count, class map, duplicate
set, and every (id, value) pair. Proven non-vacuous both ways: mutating a value
fails and names the id, mutating only a line number passes.

Major: a real divergence the parity claim would have missed. Production rejects
values containing an embedded CR, LF, U+2028 or U+2029; the example did not, so
a value with an embedded lone CR was rejected by one copy and accepted by the
other. Ported, and now covered by the parity table.

Also, found while verifying rather than reported: malformed diagnostics covered
only empty values. A doubled-dot id, a space in an id, and a lowercase-leading
id were all dropped silently. That contradicts the module's own intent -- a
typo should be diagnosable, and a space in an id is a likely one -- and
predicates are contractually cited, so a silently vanished predicate is the
failure mode that matters. Each rejection class now carries a named reason in
both copies, while ordinary inline code still yields none.

Trues up counts my own change staled: the example README and ADR-1671's
prototype figures said 416 and 393/18 against a real 415/20/0.

Closes #2944

* chore(#2944): backfill changeset PR number 2950

---------

Co-authored-by: sim <sim@local>
This commit is contained in:
Tom Boucher
2026-07-31 15:44:19 -04:00
committed by GitHub
parent 07603df8f2
commit 5a0a9f0972
10 changed files with 1174 additions and 407 deletions

View File

@@ -0,0 +1,5 @@
---
type: Fixed
pr: 2950
---
**Malformed predicate declarations are now reported instead of silently dropped.** A doubled-dot id, a space in an id, a lowercase-leading class, and a value with an embedded CR/LF are each surfaced as a distinct `malformed` diagnostic reason instead of vanishing with no trace; the example parser (examples/dynamic-context-management/) was also brought back into parity with production and its own index is now drift-guarded by a new lint script. (#2944)

View File

@@ -103,14 +103,14 @@ A working prototype proves the platform pattern end-to-end. It ships as a **refe
- `examples/dynamic-context-management/context-predicates.cjs` — pure parser/selector: `parsePredicates(markdown)` (handles bare and list-item backtick predicate forms, splits on first `=`, skips fenced code / blockquote prose, detects duplicate IDs), `selectPredicates(predicates, {klass, prefix, contains})` (the JIT "task → predicate set" selector), and `buildIndex(predicates)` (deterministic, sorted).
- `examples/dynamic-context-management/gen-context-index.cjs` — self-contained CLI with `--check`/`--write` drift-guard plus a `--select <query>` mode demonstrating JIT brief assembly.
- `examples/dynamic-context-management/CONTEXT-INDEX.json` — sample generated index: **416 predicates, 20 classes** (regenerated 2026-07-31). Originally committed as **393 predicates, 18 classes** (2026-06-24); `CONTEXT.md` has since gained the `PROBE` (11) and `PROHIB` (10) classes, with `DEFECT` 161→167 and `RULESET` 59→56. The committed artifact had gone stale (`--check` exited 1) and was regenerated with `--write`.
- `examples/dynamic-context-management/CONTEXT-INDEX.json` — sample generated index: **415 predicates, 20 classes** (verified 2026-07-31; down from 416 after #2928/PR #2938 reconciled the last duplicate predicate ID, `RULESET.WORKFLOW_MARKDOWN.FENCES`). Originally committed as **393 predicates, 18 classes** (2026-06-24); `CONTEXT.md` has since gained the `PROBE` (11) and `PROHIB` (10) classes, with `DEFECT` 161→167 and `RULESET` 59→56→55. The committed artifact had gone stale (`--check` exited 1) and was regenerated with `--write`.
- `examples/dynamic-context-management/demo.cjs` + `README.md` — runnable usage example and notes.
During research the slice was validated with 42 behavioral tests (predicate forms, fenced-code / prose skipping, duplicate-id detection, the selector, a deterministic index, and a fast-check property test); those return as CI tests under `tests/` with the production implementation.
During research the slice was validated with 42 behavioral tests (predicate forms, fenced-code / prose skipping, duplicate-id detection, the selector, a deterministic index, and a fast-check property test); those landed as CI tests under `tests/` with the production implementation (#2928/PR #2938).
The prototype immediately surfaced **3 latent duplicate predicate IDs** in `CONTEXT.md` (`RULESET.WORKFLOW_MARKDOWN.FENCES`, `RULESET.GEMINI.TOOLS.ask_user`, `RULESET.GEMINI.TEST_SENTINEL`) — integrity drift no existing tool catches.
**Re-checked 2026-07-31:** only **one** remains — `RULESET.WORKFLOW_MARKDOWN.FENCES`. The two `RULESET.GEMINI.*` duplicates were removed along with the Gemini runtime, not reconciled deliberately. Production `--check` can be made to fail on *new* duplicates once that single remaining ID is reconciled.
**Re-checked 2026-07-31:** the two `RULESET.GEMINI.*` duplicates were removed along with the Gemini runtime, not reconciled deliberately; the remaining `RULESET.WORKFLOW_MARKDOWN.FENCES` duplicate was reconciled deliberately in #2928/PR #2938, which also productionized `--check` into CI so it now fails closed on any *new* duplicate ID.
**Phase 0 acceptance status (2026-07-31).** The epic's Phase 0 criterion — "`gen-context-index --check` green in CI" — was **unmet**: `--check` exited 1 against `next`, and no CI job failed, because the example sits deliberately outside `tests/` — the red was invisible to the pipeline. The index has now been regenerated and `--check` exits 0.

File diff suppressed because one or more lines are too long

View File

@@ -21,7 +21,7 @@ hand-citing a 200 KB file.
- `context-predicates.cjs` — parser + selector + deterministic index builder (self-contained).
- `gen-context-index.cjs` — `--check` / `--write` drift-guarded generator + `--select`.
- `CONTEXT-INDEX.json` — sample generated output (393 predicates, 18 classes).
- `CONTEXT-INDEX.json` — sample generated output (415 predicates, 20 classes).
- `demo.cjs` — runnable usage example.
## Run (from the repo root)
@@ -36,9 +36,13 @@ node examples/dynamic-context-management/gen-context-index.cjs --check
During research this slice was validated with 42 behavioral tests — predicate
forms, fenced-code / prose skipping, duplicate-id detection, the selector, a
deterministic index, and a fast-check property test. Those return as CI tests
under `tests/` when the production implementation lands.
deterministic index, and a fast-check property test. The production
implementation has since landed, with `tests/context-predicates.test.cjs` and
`tests/context-index-sync.test.cjs` as its behavioral tests under `tests/`,
and `scripts/lint-example-parser-parity.cjs` (wired into `npm run lint:ci`)
asserting this example and production agree.
It also surfaced 3 latent duplicate predicate IDs in `CONTEXT.md`
Research also surfaced 3 latent duplicate predicate IDs in `CONTEXT.md`
(`RULESET.WORKFLOW_MARKDOWN.FENCES`, `RULESET.GEMINI.TOOLS.ask_user`,
`RULESET.GEMINI.TEST_SENTINEL`), recorded in the index `duplicates` field.
`RULESET.GEMINI.TEST_SENTINEL`) at the time; all three have since been
resolved and the current index carries 0 duplicate ids.

View File

@@ -17,6 +17,8 @@
*
* ID grammar: CLASS(.subkey)* where CLASS = first dot-separated segment.
* ID chars: [A-Za-z0-9._-] (CLASS always uppercase; subkeys may be mixed).
* A doubled dot (empty segment, e.g. `A..b`) is REJECTED — see the
* ID-validation comment below for why.
* Split on FIRST '=' only; everything before is the ID, everything after is
* the value (up to the closing backtick).
*
@@ -25,13 +27,105 @@
* - Prose lines (headings, blank lines, list items without a predicate)
* - The "PR fix discipline" section (pure prose, no predicates)
* - Session-log blockquote preamble
*
* `ParseResult.malformed` collects backtick lines that look like a predicate
* declaration attempt (contains a backtick-wrapped `id=value`-shaped inner
* with an `=` at index >= 1) but are rejected, with a distinct named `reason`
* per rejection class: `empty-value`, `empty-segment` (doubled dot),
* `invalid-id-chars` (disallowed characters, e.g. a space),
* `lowercase-leading-class` (id's first segment starts lowercase), and
* `value-contains-newline` (embedded CR/LF/U+2028/U+2029 in the value). A
* line with no `=` at all (ordinary inline code, e.g. `` `ID` ``) never
* produces a diagnostic. Mirrors production's src/context-predicates.cts.
*/
// Regex matching the predicate ID grammar: one or more dot-separated segments.
// First segment must start with an uppercase letter (CLASS).
// Subsequent segments may start with letter/digit and include hyphens/underscores.
// We intentionally allow lowercase-starting sub-segments (e.g. PRED.k320.rule).
const ID_RE = /^([A-Z][A-Z0-9_-]*(?:\.[A-Za-z0-9_.-]+)*)=(.+)$/;
// ID grammar, validated STRUCTURALLY rather than by a single regex
// (DEFECT.CONTEXT-PREDICATES-ID-REDOS, #2928 review). The formerly-used regex
// `^([A-Z][A-Z0-9_-]*(?:\.[A-Za-z0-9_.-]+)*)=(.+)$` is exponential: the group
// `(?:\.[A-Za-z0-9_.-]+)*` is ambiguous because its own character class
// contains `.`, so N consecutive dots have exponentially many
// backtick-partitionings for the regex engine to try on a failed match
// (measured: ~565ms for 40 consecutive dots, doubling roughly every 5).
//
// Fix: split the candidate id on '.' and validate each segment with a
// simple, non-backtracking, per-segment pattern — linear in id length, no
// ambiguous quantifier. First segment (CLASS) must start with an uppercase
// letter; subsequent segments may start with letter/digit and include
// hyphens/underscores. We intentionally allow lowercase-starting
// sub-segments (e.g. PRED.k320.rule).
//
// Behavior change vs. the old regex: an EMPTY segment (a doubled dot, e.g.
// `A..b`) now REJECTS — the old regex accepted it because `.` was inside the
// subsequent-segment character class, so `.` itself could satisfy
// `[A-Za-z0-9_.-]+` with a single character.
const ID_FIRST_SEGMENT_RE = /^[A-Z][A-Z0-9_-]*$/;
const ID_SUBSEQUENT_SEGMENT_RE = /^[A-Za-z0-9_-]+$/;
/**
* Structurally validate a candidate predicate id (linear time — no ambiguous
* backtracking quantifier; see the ID grammar comment above), returning WHY it
* is invalid so malformed diagnostics can name the exact rejection class.
*
* @param {string} id - candidate id (everything before the first '=')
* @returns {{ valid: boolean, reason?: string }}
*/
function validateIdDetailed(id) {
const segments = id.split('.');
// A doubled dot (or leading/trailing dot) produces an empty segment.
if (segments.some((seg) => seg === '')) return { valid: false, reason: 'empty-segment' };
const first = segments[0];
if (!ID_FIRST_SEGMENT_RE.test(first)) {
// Distinguish "starts lowercase" (a highly plausible typo, e.g.
// `foo.bar=1`) from any other first-segment character-set violation
// (e.g. a space, `FOO BAR=1`).
if (/^[a-z]/.test(first)) return { valid: false, reason: 'lowercase-leading-class' };
return { valid: false, reason: 'invalid-id-chars' };
}
for (let i = 1; i < segments.length; i++) {
if (!ID_SUBSEQUENT_SEGMENT_RE.test(segments[i])) return { valid: false, reason: 'invalid-id-chars' };
}
return { valid: true };
}
/**
* Structurally validate a candidate predicate id (linear time — no ambiguous
* backtracking quantifier; see the ID grammar comment above).
*
* @param {string} id - candidate id (everything before the first '=')
* @returns {boolean}
*/
function isValidId(id) {
return validateIdDetailed(id).valid;
}
/**
* Strip a source line down to its backtick-wrapped "inner" content, if any.
* Handles both line forms:
* 1. Bare backtick line: `ID=value` (starts with backtick at column 0)
* 2. List-item backtick: - `ID=value` (list-item with leading "- ")
* Also tolerates " - `ID=value`" (indented list item — observed in CONTEXT.md).
*
* @param {string} raw - the original source line (with newline stripped)
* @returns {string | null}
*/
function extractInner(raw) {
const line = raw.trimEnd();
if (line.startsWith('`') && line.endsWith('`') && line.length > 2) {
// bare backtick line
return line.slice(1, -1);
}
// strip optional leading whitespace + "- " then check for backtick wrapping
const stripped = line.replace(/^\s*-\s+/, '');
if (stripped.startsWith('`') && stripped.endsWith('`') && stripped.length > 2) {
return stripped.slice(1, -1);
}
return null;
}
/**
* Parse a single source line and return a raw {id, value} if it is a predicate,
@@ -41,24 +135,7 @@ const ID_RE = /^([A-Z][A-Z0-9_-]*(?:\.[A-Za-z0-9_.-]+)*)=(.+)$/;
* @returns {{ id: string, value: string } | null}
*/
function extractPredicate(raw) {
const line = raw.trimEnd();
// Form 1: `ID=value` (starts with backtick at column 0)
// Form 2: - `ID=value` (list-item with leading "- ")
// Also tolerate " - `ID=value`" (indented list item — observed in CONTEXT.md).
let inner = null;
if (line.startsWith('`') && line.endsWith('`') && line.length > 2) {
// bare backtick line
inner = line.slice(1, -1);
} else {
// strip optional leading whitespace + "- " then check for backtick wrapping
const stripped = line.replace(/^\s*-\s+/, '');
if (stripped.startsWith('`') && stripped.endsWith('`') && stripped.length > 2) {
inner = stripped.slice(1, -1);
}
}
const inner = extractInner(raw);
if (inner === null) return null;
// Now match the ID grammar. Split on FIRST '=' only.
@@ -68,12 +145,52 @@ function extractPredicate(raw) {
const id = inner.slice(0, eqIdx);
const value = inner.slice(eqIdx + 1);
// Validate ID — must match the grammar (no spaces, correct char set).
if (!ID_RE.test(inner)) return null;
// Value must be non-empty and must contain no embedded ECMAScript
// LineTerminator character (LF, CR, U+2028 LINE SEPARATOR, U+2029 PARAGRAPH
// SEPARATOR) — mirrors production's src/context-predicates.cts. And the id
// must match the structural grammar (no spaces, correct char set, no empty
// segment — see isValidId's doc comment).
if (value === '' || /[\n\r\u2028\u2029]/.test(value) || !isValidId(id)) return null;
return { id, value };
}
/**
* Detect the "looks like a predicate declaration attempt but is rejected"
* malformed case for a line that {@link extractPredicate} already rejected —
* naming WHY. Only fires when the line is backtick-wrapped AND contains an
* `=` at index >= 1 — a plain inline-code line with no `=` at all (e.g.
* `` `ID` ``) is not a declaration attempt and never produces a diagnostic.
* Does not change any accept/reject decision — diagnostic only. Mirrors
* production's src/context-predicates.cts detectMalformed.
*
* @param {string} raw - the original source line (with newline stripped)
* @returns {{ text: string, reason: string } | null}
*/
function detectMalformed(raw) {
const inner = extractInner(raw);
if (inner === null) return null;
const eqIdx = inner.indexOf('=');
if (eqIdx < 1) return null;
const id = inner.slice(0, eqIdx);
const value = inner.slice(eqIdx + 1);
const idCheck = validateIdDetailed(id);
if (!idCheck.valid) {
return { text: raw.trimEnd(), reason: idCheck.reason };
}
if (value === '') {
return { text: raw.trimEnd(), reason: 'empty-value' };
}
if (/[\n\r\u2028\u2029]/.test(value)) {
return { text: raw.trimEnd(), reason: 'value-contains-newline' };
}
return null;
}
/**
* Parse all predicates from a CONTEXT.md markdown string.
*
@@ -81,12 +198,14 @@ function extractPredicate(raw) {
* @returns {{
* predicates: Array<{ id: string, klass: string, value: string, line: number, section: string }>,
* duplicates: Array<{ id: string, lines: number[] }>,
* malformed: Array<{ line: number, text: string, reason: string }>,
* skippedSections: string[]
* }}
*/
function parsePredicates(markdown) {
const lines = markdown.split('\n');
const predicates = [];
const malformed = [];
// Track id -> list of line numbers for duplicate detection
const idLines = new Map(); // id -> number[]
@@ -131,7 +250,11 @@ function parsePredicates(markdown) {
// Attempt extraction.
const pred = extractPredicate(raw);
if (!pred) continue;
if (!pred) {
const bad = detectMalformed(raw);
if (bad) malformed.push({ line: lineNo, text: bad.text, reason: bad.reason });
continue;
}
const klass = pred.id.split('.')[0];
predicates.push({
@@ -164,7 +287,7 @@ function parsePredicates(markdown) {
const activeSections = new Set(predicates.map((p) => p.section));
const skippedSections = allSections.filter((s) => !activeSections.has(s));
return { predicates, duplicates, skippedSections };
return { predicates, duplicates, malformed, skippedSections };
}
/**

View File

@@ -105,7 +105,7 @@
"lint": "eslint . --cache --cache-location node_modules/.cache/eslint/",
"lint:fix": "eslint . --fix",
"lint:table-schema-drift": "node scripts/lint-table-schema-drift.cjs",
"lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-emitted-drift-ack.cjs && node scripts/lint-portable-timeout.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-test.cjs",
"lint:ci": "npm run lint && npm run lint:skill-deps && npm run lint:generated-sync && node scripts/lint-test-file-count.cjs && node scripts/lint-command-contract.cjs && node scripts/lint-pr-check-project-dir.cjs && npm run lint:legacy-name && node scripts/lint-regression-test-names.cjs && node scripts/lint-allow-test-rule-refs.cjs && node scripts/lint-resolution-provenance.cjs && node scripts/lint-emitted-drift-ack.cjs && node scripts/lint-portable-timeout.cjs && node scripts/validate-registry.cjs && node scripts/lint-table-schema-drift.cjs && node scripts/lint-fix-has-regression-test.cjs && node scripts/lint-example-parser-parity.cjs",
"lint:allow-test-rule-refs": "node scripts/lint-allow-test-rule-refs.cjs",
"lint:regression-names": "node scripts/lint-regression-test-names.cjs",
"lint:descriptions": "node scripts/lint-descriptions.cjs",

View File

@@ -0,0 +1,395 @@
#!/usr/bin/env node
'use strict';
/**
* lint-example-parser-parity.cjs — asserts examples/dynamic-context-management/
* context-predicates.cjs (reference prototype, ADR-1671) and production's
* src/context-predicates.cts (compiled to gsd-core/bin/lib/context-predicates.cjs)
* agree on parsing behavior, and that the example's own committed
* CONTEXT-INDEX.json has not silently drifted from a fresh parse.
*
* Lives OUTSIDE tests/ deliberately. ADR-1671
* (docs/adr/1671-dynamic-context-management-platform.md:102) places
* examples/dynamic-context-management/ outside FOUR surfaces: the build
* (src/ -> bin/lib/), the npm package files[], the installer, and the CI test
* suite (tests/). A lint script wired into `npm run lint:ci` is none of
* those four — it does not compile, package, install, or test-suite-import
* the example; it only READS both modules from a repo-root script and
* asserts they agree, the same shape as every other scripts/lint-*.cjs in
* this repo that reads source it does not own (see e.g.
* lint-compiled-artifact-sync.cjs).
*
* Two parity surfaces:
* 1. production vs. example, parsed against the real repo-root CONTEXT.md:
* predicate count, class map, duplicate-id set, the full set of
* (id, value) pairs, and a representative accept/reject id-shape table
* (including the embedded-CR-in-value case production already rejected
* and the example silently accepted before this fix).
* 2. the example's OWN committed CONTEXT-INDEX.json vs. a fresh parse by
* the example's own parser of the real CONTEXT.md — LINE-NUMBER-
* INDEPENDENT (see below). This guards the same #2944 merge-race
* staleness class this PR fixed for docs/CONTEXT-INDEX.json, but for the
* example's own copy, which nothing previously guarded at all.
*
* Why `line` is excluded from surface 2 (ADR-1671 open question 4): the
* example's CONTEXT-INDEX.json bakes source line numbers into every
* predicate entry. Any unrelated CONTEXT.md line-shift (e.g. inserting a
* sentence above the predicates) re-drifts that artifact even though every
* predicate id/value/class is byte-identical. Wiring the example's own
* `gen-context-index.cjs --check` into lint:ci — which DOES compare line
* numbers — would make CI routinely red for reasons unrelated to predicate
* integrity. This lint instead derives the LINE-INDEPENDENT facts from the
* committed index (count, class map, duplicate-id set, (id,value) pairs) and
* compares those to a fresh parse: genuine content drift (the #2944
* merge-race defect class) is still caught; a pure line-shift is not.
*
* `--context-path` / `--example-index-path` override the two hardcoded paths
* (mirrors scripts/gen-context-index.cjs) so this script's non-vacuousness
* can be proven against a throwaway temp fixture without ever touching the
* real committed artifacts.
*/
const fs = require('node:fs');
const path = require('node:path');
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
const ROOT = path.resolve(__dirname, '..');
const PROD_PREDICATES_PATH = path.join(ROOT, 'gsd-core', 'bin', 'lib', 'context-predicates.cjs');
const EXAMPLE_DIR = path.join(ROOT, 'examples', 'dynamic-context-management');
const EXAMPLE_PREDICATES_PATH = path.join(EXAMPLE_DIR, 'context-predicates.cjs');
const DEFAULT_EXAMPLE_INDEX_PATH = path.join(EXAMPLE_DIR, 'CONTEXT-INDEX.json');
const DEFAULT_CONTEXT_PATH = path.join(ROOT, 'CONTEXT.md');
// ─── Loaders ──────────────────────────────────────────────────────────────────
/**
* Load the compiled production predicates library. Throws a clean ExitError
* (never a bare MODULE_NOT_FOUND stack) naming the remedy when it is missing.
*/
function loadProdPredicates() {
try {
delete require.cache[require.resolve(PROD_PREDICATES_PATH)];
return require(PROD_PREDICATES_PATH);
} catch (err) {
throw new ExitError(
1,
`Cannot load ${path.relative(ROOT, PROD_PREDICATES_PATH)}: ${err && err.message}\n` +
'Run:\n npm run build:lib\n',
);
}
}
/** Load the example's self-contained predicates module (no build step). */
function loadExamplePredicates() {
delete require.cache[require.resolve(EXAMPLE_PREDICATES_PATH)];
return require(EXAMPLE_PREDICATES_PATH);
}
/**
* Read a CONTEXT.md-shaped markdown file. Throws a clean ExitError naming the
* path when it is missing or unreadable.
*
* @param {string} contextPath
* @returns {string}
*/
function readContextMarkdown(contextPath) {
try {
return fs.readFileSync(contextPath, 'utf8');
} catch (err) {
throw new ExitError(1, `Cannot read ${path.relative(ROOT, contextPath)}: ${err && err.message}`);
}
}
// ─── Diffing helpers ────────────────────────────────────────────────────────
/**
* Diff two predicate-shaped arrays by id, comparing ONLY `klass`/`value`
* (never `line`/`section` — see the module doc comment's "Why `line` is
* excluded"). Returns human-readable divergence strings naming the exact
* predicate id(s), so a failure reads as an actionable list, never "objects
* differ".
*
* @param {Array<{id:string,klass:string,value:string}>} leftPredicates
* @param {Array<{id:string,klass:string,value:string}>} rightPredicates
* @param {string} leftLabel
* @param {string} rightLabel
* @returns {string[]}
*/
function diffPredicatesById(leftPredicates, rightPredicates, leftLabel, rightLabel) {
const leftMap = new Map(leftPredicates.map((p) => [p.id, p]));
const rightMap = new Map(rightPredicates.map((p) => [p.id, p]));
const allIds = new Set([...leftMap.keys(), ...rightMap.keys()]);
const diffs = [];
for (const id of Array.from(allIds).sort()) {
const l = leftMap.get(id);
const r = rightMap.get(id);
if (l && !r) {
diffs.push(`${id}: present in ${leftLabel} but absent from ${rightLabel}`);
} else if (!l && r) {
diffs.push(`${id}: present in ${rightLabel} but absent from ${leftLabel}`);
} else if (l.value !== r.value) {
diffs.push(
`${id}: value diverged (${leftLabel}=${JSON.stringify(l.value)}, ${rightLabel}=${JSON.stringify(r.value)})`,
);
} else if (l.klass !== r.klass) {
diffs.push(
`${id}: klass diverged (${leftLabel}=${JSON.stringify(l.klass)}, ${rightLabel}=${JSON.stringify(r.klass)})`,
);
}
}
return diffs;
}
/** Sorted, deduplicated id list from a `duplicates` array (either `{id,count}` or `{id,lines}` shape). */
function duplicateIds(duplicates) {
return duplicates.map((d) => d.id).sort();
}
/** Build a single-line `` `id=value` `` markdown fixture wrapping one candidate id=value pair. */
function backtickLine(idEqualsValue) {
return '`' + idEqualsValue + '`';
}
// Representative accept/reject id-shape table both parsers must agree on —
// this is what PROVES the parity claim rather than merely asserting it.
// FINDING 4 (ADR-1671 PR review): the embedded-CR-in-value case below was the
// one real divergence the earlier parity claim missed — production rejected
// it, the example silently accepted it. Both now reject it (ported fix).
const ID_SHAPE_CASES = [
{ idEqualsValue: 'A=1', expectAccept: true },
{ idEqualsValue: 'FOO=x', expectAccept: true },
{ idEqualsValue: 'PRED.k320.rule=x', expectAccept: true },
{ idEqualsValue: 'RELEASE-NOTES.x=y', expectAccept: true },
{ idEqualsValue: 'A..b=1', expectAccept: false },
{ idEqualsValue: 'foo.bar=x', expectAccept: false },
{ idEqualsValue: 'FOO BAR=x', expectAccept: false },
{ idEqualsValue: '=v', expectAccept: false },
{ idEqualsValue: 'ID', expectAccept: false },
];
/**
* The embedded-CR case needs its own fixture, not the plain
* `` `ID=value` `` shape backtickLine() builds: a lone CR with no following
* LF never reaches a closing backtick at all (a separate, documented parser
* limit — see src/context-predicates.cts's module doc comment), so the CR
* must sit INSIDE an otherwise well-formed, LF-terminated backtick line.
*/
const CR_VALUE_MARKDOWN = '`ID=ab\rcd`\n';
// ─── Parity checks (surface 1: production vs. example) ────────────────────────
/**
* @param {object} prodPredicates
* @param {object} examplePredicates
* @param {string} markdown
* @returns {string[]} diagnostics; empty when the two parsers fully agree
*/
function checkProdExampleParity(prodPredicates, examplePredicates, markdown) {
const diagnostics = [];
const prodParsed = prodPredicates.parsePredicates(markdown);
const exampleParsed = examplePredicates.parsePredicates(markdown);
const prodIndex = prodPredicates.buildIndex(prodParsed.predicates);
const exampleIndex = examplePredicates.buildIndex(exampleParsed.predicates);
if (exampleIndex.count !== prodIndex.count) {
diagnostics.push(`predicate count diverged: example=${exampleIndex.count} production=${prodIndex.count}`);
}
const prodClassKeys = Object.keys(prodIndex.classes).sort();
const exampleClassKeys = Object.keys(exampleIndex.classes).sort();
if (JSON.stringify(exampleClassKeys) !== JSON.stringify(prodClassKeys)) {
diagnostics.push(
`class set diverged: example=${JSON.stringify(exampleClassKeys)} production=${JSON.stringify(prodClassKeys)}`,
);
} else if (JSON.stringify(exampleIndex.classes) !== JSON.stringify(prodIndex.classes)) {
diagnostics.push(
`per-class predicate counts diverged: example=${JSON.stringify(exampleIndex.classes)} ` +
`production=${JSON.stringify(prodIndex.classes)}`,
);
}
const prodDupIds = duplicateIds(prodIndex.duplicates);
const exampleDupIds = duplicateIds(exampleIndex.duplicates);
if (JSON.stringify(exampleDupIds) !== JSON.stringify(prodDupIds)) {
diagnostics.push(
`duplicate-id set diverged: example=${JSON.stringify(exampleDupIds)} production=${JSON.stringify(prodDupIds)}`,
);
}
diagnostics.push(...diffPredicatesById(exampleIndex.predicates, prodIndex.predicates, 'example', 'production'));
for (const { idEqualsValue, expectAccept } of ID_SHAPE_CASES) {
const line = backtickLine(idEqualsValue);
const exampleCount = examplePredicates.parsePredicates(line).predicates.length;
const prodCount = prodPredicates.parsePredicates(line).predicates.length;
if (exampleCount !== prodCount) {
diagnostics.push(
`id-shape verdict diverged for ${JSON.stringify(idEqualsValue)}: ` +
`example found ${exampleCount} predicate(s), production found ${prodCount}`,
);
continue;
}
const actualAccept = prodCount === 1;
if (actualAccept !== expectAccept) {
diagnostics.push(
`id-shape verdict wrong for ${JSON.stringify(idEqualsValue)}: expected ` +
`${expectAccept ? 'ACCEPT' : 'REJECT'}, both modules agreed on ${actualAccept ? 'ACCEPT' : 'REJECT'} instead`,
);
}
}
const exampleCrCount = examplePredicates.parsePredicates(CR_VALUE_MARKDOWN).predicates.length;
const prodCrCount = prodPredicates.parsePredicates(CR_VALUE_MARKDOWN).predicates.length;
if (exampleCrCount !== prodCrCount) {
diagnostics.push(
`id-shape verdict diverged for embedded-CR-in-value: example found ${exampleCrCount} predicate(s), ` +
`production found ${prodCrCount}`,
);
} else if (prodCrCount !== 0) {
diagnostics.push(
`embedded-CR-in-value must be REJECTED by both parsers; both instead ACCEPTED it (${prodCrCount} predicate(s))`,
);
}
return diagnostics;
}
// ─── Parity check (surface 2: example's committed index vs. fresh parse) ──────
/** Strip a predicate down to the line-number-independent fields used for comparison. */
function toLineIndependent(p) {
return { id: p.id, klass: p.klass, value: p.value };
}
/**
* @param {object} examplePredicates
* @param {string} markdown
* @param {string} exampleIndexPath
* @returns {string[]} diagnostics; empty when the committed index and a fresh
* parse agree on every LINE-NUMBER-INDEPENDENT fact
*/
function checkExampleIndexParity(examplePredicates, markdown, exampleIndexPath) {
const diagnostics = [];
let committed;
try {
committed = JSON.parse(fs.readFileSync(exampleIndexPath, 'utf8'));
} catch (err) {
diagnostics.push(`cannot read/parse ${exampleIndexPath}: ${err && err.message}`);
return diagnostics;
}
if (!committed || !Array.isArray(committed.predicates)) {
diagnostics.push(`${exampleIndexPath} does not have the expected ContextIndex shape`);
return diagnostics;
}
const { predicates } = examplePredicates.parsePredicates(markdown);
const fresh = examplePredicates.buildIndex(predicates);
if (committed.count !== fresh.count) {
diagnostics.push(`example index predicate count diverged: committed=${committed.count} fresh=${fresh.count}`);
}
if (JSON.stringify(committed.classes) !== JSON.stringify(fresh.classes)) {
diagnostics.push(
`example index class-count map diverged: committed=${JSON.stringify(committed.classes)} ` +
`fresh=${JSON.stringify(fresh.classes)}`,
);
}
const committedDupIds = duplicateIds(committed.duplicates || []);
const freshDupIds = duplicateIds(fresh.duplicates);
if (JSON.stringify(committedDupIds) !== JSON.stringify(freshDupIds)) {
diagnostics.push(
`example index duplicate-id set diverged: committed=${JSON.stringify(committedDupIds)} ` +
`fresh=${JSON.stringify(freshDupIds)}`,
);
}
const committedIndependent = committed.predicates.map(toLineIndependent);
const freshIndependent = fresh.predicates.map(toLineIndependent);
const diffs = diffPredicatesById(committedIndependent, freshIndependent, 'committed example index', 'fresh parse');
diagnostics.push(...diffs.map((d) => `example index: ${d}`));
return diagnostics;
}
// ─── Argument parsing ─────────────────────────────────────────────────────────
/**
* @param {string[]} argv - process.argv.slice(2)
* @returns {{ contextPath: string, exampleIndexPath: string }}
*/
function parseArgs(argv) {
const opts = { contextPath: DEFAULT_CONTEXT_PATH, exampleIndexPath: DEFAULT_EXAMPLE_INDEX_PATH };
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
if (arg === '--context-path') {
opts.contextPath = path.resolve(argv[++i] ?? '');
} else if (arg === '--example-index-path') {
opts.exampleIndexPath = path.resolve(argv[++i] ?? '');
} else {
throw new ExitError(
1,
`Unknown argument: ${arg}\n` +
'Usage: lint-example-parser-parity.cjs [--context-path <path>] [--example-index-path <path>]\n',
);
}
}
return opts;
}
// ─── Main ─────────────────────────────────────────────────────────────────────
function main() {
const opts = parseArgs(process.argv.slice(2));
const prodPredicates = loadProdPredicates();
const examplePredicates = loadExamplePredicates();
const markdown = readContextMarkdown(opts.contextPath);
const diagnostics = [
...checkProdExampleParity(prodPredicates, examplePredicates, markdown),
...checkExampleIndexParity(examplePredicates, markdown, opts.exampleIndexPath),
];
if (diagnostics.length > 0) {
process.stderr.write(
`\nERROR lint-example-parser-parity: ${diagnostics.length} divergence(s) found\n\n`,
);
for (const d of diagnostics) process.stderr.write(` - ${d}\n`);
process.stderr.write(
'\nexamples/dynamic-context-management/context-predicates.cjs and ' +
'src/context-predicates.cts (gsd-core/bin/lib/context-predicates.cjs) must agree; ' +
'the committed examples/dynamic-context-management/CONTEXT-INDEX.json must match a ' +
'fresh parse of CONTEXT.md on every line-number-independent fact (run:\n' +
' node examples/dynamic-context-management/gen-context-index.cjs --write\n' +
').\n',
);
throw new ExitError(1);
}
process.stdout.write(
'ok lint-example-parser-parity: production and example parsers agree on the real CONTEXT.md ' +
`(count/classes/duplicates/(id,value) pairs), on ${ID_SHAPE_CASES.length + 1} representative ` +
'accept/reject id-shape cases, and the committed example index matches a fresh parse ' +
'(line-number-independent)\n',
);
}
module.exports = {
loadProdPredicates,
loadExamplePredicates,
readContextMarkdown,
diffPredicatesById,
checkProdExampleParity,
checkExampleIndexParity,
parseArgs,
ID_SHAPE_CASES,
CR_VALUE_MARKDOWN,
};
if (require.main === module) {
runMain(main);
}

View File

@@ -18,10 +18,17 @@
* `line` and `section`.
*
* One addition beyond the prototype: `ParseResult.malformed` collects
* backtick lines that look like a predicate declaration but are rejected for
* having an empty value (e.g. `` `ID=` ``), so the empty-value case is
* surfaced as a diagnostic instead of being silently dropped. This does not
* change any accept/reject outcome — only adds a diagnostic.
* backtick lines that look like a predicate declaration attempt (contains a
* backtick-wrapped `id=value`-shaped inner with an `=` at index >= 1) but are
* rejected, with a distinct named `reason` per rejection class: `empty-value`
* (e.g. `` `ID=` ``), `empty-segment` (a doubled dot in the id, e.g.
* `` `A..b=1` ``), `invalid-id-chars` (disallowed characters in the id, e.g. a
* space: `` `FOO BAR=1` ``), `lowercase-leading-class` (id's first segment
* starts lowercase, e.g. `` `foo.bar=1` ``), and `value-contains-newline` (an
* embedded CR/LF/U+2028/U+2029 in the value). A line with no `=` at all (e.g.
* `` `ID` ``, ordinary inline code) is NOT a declaration attempt and never
* produces a diagnostic. This does not change any accept/reject outcome —
* only adds diagnostics.
*
* Grammar (from discovery facts):
* Two line forms, each on exactly one source line:
@@ -184,6 +191,39 @@ export interface ContextIndex {
const ID_FIRST_SEGMENT_RE = /^[A-Z][A-Z0-9_-]*$/;
const ID_SUBSEQUENT_SEGMENT_RE = /^[A-Za-z0-9_-]+$/;
/** Result of {@link validateIdDetailed}: valid, or invalid with a named reason. */
interface IdValidation {
valid: boolean;
reason?: 'empty-segment' | 'invalid-id-chars' | 'lowercase-leading-class';
}
/**
* Structurally validate a candidate predicate id (linear time — no ambiguous
* backtracking quantifier; see the ID grammar comment above), returning WHY it
* is invalid so malformed diagnostics can name the exact rejection class.
*
* @param id - candidate id (everything before the first '=')
*/
function validateIdDetailed(id: string): IdValidation {
const segments = id.split('.');
// A doubled dot (or leading/trailing dot) produces an empty segment.
if (segments.some((seg) => seg === '')) return { valid: false, reason: 'empty-segment' };
const first = segments[0];
if (!ID_FIRST_SEGMENT_RE.test(first)) {
// Distinguish "starts lowercase" (a highly plausible typo, e.g.
// `foo.bar=1`) from any other first-segment character-set violation
// (e.g. a space, `FOO BAR=1`).
if (/^[a-z]/.test(first)) return { valid: false, reason: 'lowercase-leading-class' };
return { valid: false, reason: 'invalid-id-chars' };
}
for (let i = 1; i < segments.length; i++) {
if (!ID_SUBSEQUENT_SEGMENT_RE.test(segments[i])) return { valid: false, reason: 'invalid-id-chars' };
}
return { valid: true };
}
/**
* Structurally validate a candidate predicate id (linear time — no ambiguous
* backtracking quantifier; see the ID grammar comment above).
@@ -191,12 +231,7 @@ const ID_SUBSEQUENT_SEGMENT_RE = /^[A-Za-z0-9_-]+$/;
* @param id - candidate id (everything before the first '=')
*/
function isValidId(id: string): boolean {
const segments = id.split('.');
if (!ID_FIRST_SEGMENT_RE.test(segments[0])) return false;
for (let i = 1; i < segments.length; i++) {
if (!ID_SUBSEQUENT_SEGMENT_RE.test(segments[i])) return false;
}
return true;
return validateIdDetailed(id).valid;
}
// List markers recognized ahead of a backtick-wrapped declaration:
@@ -273,11 +308,18 @@ function extractPredicate(raw: string): { id: string; value: string } | null {
}
/**
* Detect the "looks like a declaration but has an empty value" malformed
* case for a line that {@link extractPredicate} already rejected. Only
* fires when the ID portion is grammatically valid on its own and the value
* after the first '=' is empty (e.g. `` `ID=` ``). Does not change any
* accept/reject decision — diagnostic only.
* Detect the "looks like a predicate declaration attempt but is rejected"
* malformed case for a line that {@link extractPredicate} already rejected —
* naming WHY, so a maintainer's typo is diagnosable instead of silently
* vanishing. Only fires when the line is backtick-wrapped (in either
* recognized form) AND contains an `=` at index >= 1 — a plain inline-code
* line with no `=` at all (e.g. `` `ID` ``) is not a declaration attempt and
* never produces a diagnostic. Does not change any accept/reject decision —
* diagnostic only.
*
* Reason precedence when a line fails more than one check at once: id
* validity is checked first (an invalid id makes the value irrelevant), then
* empty-value, then embedded-newline.
*
* @param raw - the original source line (with newline stripped)
*/
@@ -291,9 +333,16 @@ function detectMalformed(raw: string): { text: string; reason: string } | null {
const id = inner.slice(0, eqIdx);
const value = inner.slice(eqIdx + 1);
if (value === '' && isValidId(id)) {
const idCheck = validateIdDetailed(id);
if (!idCheck.valid) {
return { text: raw.trimEnd(), reason: idCheck.reason as string };
}
if (value === '') {
return { text: raw.trimEnd(), reason: 'empty-value' };
}
if (/[\n\r\u2028\u2029]/.test(value)) {
return { text: raw.trimEnd(), reason: 'value-contains-newline' };
}
return null;
}

View File

@@ -0,0 +1,153 @@
'use strict';
/**
* Regression test for #2944 — 85140eac8 "refresh the predicate index staled
* by a merge race on next" (docs/CONTEXT-INDEX.json) shipped with zero
* behavioral test coverage.
*
* docs/CONTEXT-INDEX.json is a committed generated artifact guarded by
* `scripts/gen-context-index.cjs --check` in `lint:generated-sync`. A
* concurrent-merge race landed one PR's CONTEXT.md alongside another PR's
* index; each PR was green alone (lint only runs on the PR's own diff), the
* combination was red, and it only surfaced on the NEXT PR to run `lint:ci`.
* This file puts the same drift check inside the test suite (which
* `gsd-test` runs on every PR), so a future racing merge fails here directly
* instead of waiting for a downstream PR to trip over `lint:generated-sync`.
*
* The example/production parser parity guard (the OTHER defect #2944
* introduced regression coverage for) lives in
* scripts/lint-example-parser-parity.cjs, wired into `npm run lint:ci` — NOT
* here. ADR-1671 ("Dynamic context management platform",
* docs/adr/1671-dynamic-context-management-platform.md:102) places
* examples/dynamic-context-management/ deliberately outside four surfaces:
* the build (src/ -> bin/lib/), the npm package files[], the installer, and
* the CI test suite (tests/) — this file. A `require()` of the example from
* inside tests/ would violate that fourth exclusion directly; a lint script
* that only reads both modules from a repo-root script does not.
*/
process.env.GSD_TEST_MODE = '1';
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const os = require('node:os');
const path = require('node:path');
const helpers = require('./helpers.cjs');
const REPO_ROOT = path.resolve(__dirname, '..');
const LIB_DIR = path.join(REPO_ROOT, 'gsd-core', 'bin', 'lib');
const PROD_PREDICATES_PATH = path.join(LIB_DIR, 'context-predicates.cjs');
const CONTEXT_PATH = path.join(REPO_ROOT, 'CONTEXT.md');
const INDEX_PATH = path.join(REPO_ROOT, 'docs', 'CONTEXT-INDEX.json');
const prodPredicates = require(PROD_PREDICATES_PATH);
function readContextMarkdown() {
return fs.readFileSync(CONTEXT_PATH, 'utf8');
}
function readCommittedIndex() {
return JSON.parse(fs.readFileSync(INDEX_PATH, 'utf8'));
}
/**
* Diff two ContextIndex-shaped `{predicates:[{id,klass,value}]}` objects by
* id, returning human-readable divergence strings naming the exact
* predicate id(s) involved — so a failure reads as an actionable list, never
* "objects differ".
*/
function diffPredicatesById(leftPredicates, rightPredicates, leftLabel, rightLabel) {
const leftMap = new Map(leftPredicates.map((p) => [p.id, p]));
const rightMap = new Map(rightPredicates.map((p) => [p.id, p]));
const allIds = new Set([...leftMap.keys(), ...rightMap.keys()]);
const diffs = [];
for (const id of Array.from(allIds).sort()) {
const l = leftMap.get(id);
const r = rightMap.get(id);
if (l && !r) {
diffs.push(`${id}: present in ${leftLabel} but absent from ${rightLabel}`);
} else if (!l && r) {
diffs.push(`${id}: present in ${rightLabel} but absent from ${leftLabel}`);
} else if (l.value !== r.value) {
diffs.push(
`${id}: value diverged (${leftLabel}=${JSON.stringify(l.value)}, ${rightLabel}=${JSON.stringify(r.value)})`,
);
} else if (l.klass !== r.klass) {
diffs.push(
`${id}: klass diverged (${leftLabel}=${JSON.stringify(l.klass)}, ${rightLabel}=${JSON.stringify(r.klass)})`,
);
}
}
return diffs;
}
// ─── docs/CONTEXT-INDEX.json vs. a fresh CONTEXT.md parse ────────────────────
describe('docs/CONTEXT-INDEX.json sync with CONTEXT.md (#2944 concurrent-merge regression)', () => {
test('committed index matches a fresh parse of the real CONTEXT.md, predicate-by-predicate', () => {
const { parsePredicates, buildIndex } = prodPredicates;
const markdown = readContextMarkdown();
const { predicates } = parsePredicates(markdown);
const fresh = buildIndex(predicates);
const committed = readCommittedIndex();
const diffs = diffPredicatesById(
committed.predicates,
fresh.predicates,
'committed docs/CONTEXT-INDEX.json',
'fresh CONTEXT.md parse',
);
assert.deepEqual(
diffs,
[],
'docs/CONTEXT-INDEX.json has drifted from CONTEXT.md for predicate id(s) ' +
'(run `node scripts/gen-context-index.cjs --write` to refresh):\n' +
diffs.join('\n'),
);
assert.equal(
committed.count,
fresh.count,
`predicate count diverged: committed=${committed.count} fresh=${fresh.count}`,
);
assert.deepEqual(
committed.classes,
fresh.classes,
'class-count map diverged between committed index and fresh CONTEXT.md parse:\n' +
`committed=${JSON.stringify(committed.classes)}\nfresh=${JSON.stringify(fresh.classes)}`,
);
});
test('divergence detector is not vacuous: catches a mutated index in a throwaway temp copy (never the real artifact)', (t) => {
const { parsePredicates, buildIndex } = prodPredicates;
const markdown = readContextMarkdown();
const { predicates } = parsePredicates(markdown);
const fresh = buildIndex(predicates);
// Mutate a COPY of the fresh index in a fresh temp dir. docs/CONTEXT-
// INDEX.json itself is never opened for write anywhere in this file.
const mutated = JSON.parse(JSON.stringify(fresh));
const target =
mutated.predicates.find((p) => p.id === 'RULESET.WORKFLOW_SIZE_BUDGET') || mutated.predicates[0];
target.value = target.value + ' MUTATED-FOR-TEST';
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'context-index-sync-'));
t.after(() => helpers.cleanup(tmpDir));
const tmpIndexPath = path.join(tmpDir, 'CONTEXT-INDEX.json');
fs.writeFileSync(tmpIndexPath, JSON.stringify(mutated, null, 2) + '\n', 'utf8');
const mutatedCommitted = JSON.parse(fs.readFileSync(tmpIndexPath, 'utf8'));
const diffs = diffPredicatesById(
mutatedCommitted.predicates,
fresh.predicates,
'mutated temp-copy index',
'fresh CONTEXT.md parse',
);
assert.ok(diffs.length > 0, 'expected the mutated temp copy to diverge from the fresh parse');
assert.ok(
diffs.some((d) => d.startsWith(`${target.id}:`)),
`expected the diff to name ${target.id}; got:\n${diffs.join('\n')}`,
);
});
});

View File

@@ -320,6 +320,11 @@ describe('parsePredicates: ID/value grammar boundaries (C)', () => {
test('rejectsInlineCodeWithoutEqualsSign', () => {
const r = parsePredicates('`ID`');
assert.equal(r.predicates.length, 0);
assert.equal(
r.malformed.length,
0,
'ordinary inline code with no "=" at all is not a declaration attempt and must not be diagnosed',
);
});
test('reportsEmptyValueAsMalformedRatherThanDroppingSilently', () => {
@@ -523,6 +528,53 @@ describe('parsePredicates: duplicate detection + validation (E)', () => {
});
});
// ─── E2. Malformed diagnostics — one distinct reason per rejection class ───
describe('parsePredicates: malformed diagnostics name the exact rejection reason (E2)', () => {
test('reportsEmptySegmentForDoubledDot', () => {
const r = parsePredicates('`A..b=1`');
assert.equal(r.predicates.length, 0);
assert.equal(r.malformed.length, 1);
assert.equal(r.malformed[0].reason, 'empty-segment');
});
test('reportsInvalidIdCharsForSpaceInId', () => {
const r = parsePredicates('`FOO BAR=1`');
assert.equal(r.predicates.length, 0);
assert.equal(r.malformed.length, 1);
assert.equal(r.malformed[0].reason, 'invalid-id-chars');
});
test('reportsLowercaseLeadingClassForLowercaseFirstSegment', () => {
const r = parsePredicates('`foo.bar=1`');
assert.equal(r.predicates.length, 0);
assert.equal(r.malformed.length, 1);
assert.equal(r.malformed[0].reason, 'lowercase-leading-class');
});
test('reportsValueContainsNewlineForEmbeddedCr', () => {
// A lone embedded CR inside an otherwise well-formed, LF-terminated
// backtick line (distinct from the documented lone-CR-only-line limit
// pinned by yieldsNoPredicatesForLoneCrDocumentAsDocumentedLimit above,
// which never reaches a closing backtick at all).
const md = '`ID=ab\rcd`\n';
const r = parsePredicates(md);
assert.equal(r.predicates.length, 0, 'a value with an embedded CR must be rejected, not silently accepted');
assert.equal(r.malformed.length, 1);
assert.equal(r.malformed[0].reason, 'value-contains-newline');
});
test('idCheckTakesPrecedenceOverEmptyValueWhenBothFail', () => {
// `foo.bar=` fails BOTH the id (lowercase-leading) and the value (empty)
// checks — id validity is checked first per detectMalformed's documented
// precedence.
const r = parsePredicates('`foo.bar=`');
assert.equal(r.predicates.length, 0);
assert.equal(r.malformed.length, 1);
assert.equal(r.malformed[0].reason, 'lowercase-leading-class');
});
});
// ─── I. Independence + real-corpus regression ──────────────────────────────
describe('parsePredicates: independence + real-corpus regression (I)', () => {