* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local>
This commit is contained in:
5
.changeset/zesty-koalas-sing.md
Normal file
5
.changeset/zesty-koalas-sing.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Added
|
||||
pr: 2938
|
||||
---
|
||||
**`gsd_run query context-predicates` — targeted lookups against the `CONTEXT.md` fact-store** — search predicates live by class, id prefix, or substring instead of reading the whole file, with a CI-guarded `docs/CONTEXT-INDEX.json` index kept in sync automatically. (#2928)
|
||||
1
.gitignore
vendored
1
.gitignore
vendored
@@ -84,6 +84,7 @@ build/
|
||||
/gsd-core/bin/lib/mcp-server.cjs
|
||||
/gsd-core/bin/lib/external-descriptor-trust.cjs
|
||||
/gsd-core/bin/lib/cli-skew-check.cjs
|
||||
/gsd-core/bin/lib/context-predicates.cjs
|
||||
/gsd-core/bin/lib/capability-loader.cjs
|
||||
/gsd-core/bin/lib/capability-source.cjs
|
||||
/gsd-core/bin/lib/capability-ledger.cjs
|
||||
|
||||
@@ -524,7 +524,6 @@ The prompt-level data/instruction isolation seam for untrusted web/document ingr
|
||||
`RULESET.CODERABBIT.GUARD.RESOLVE=fix validated finding -> focused tests -> commit/push -> resolveReviewThread(threadId) -> wait CI/CodeRabbit -> final unresolved_count query`
|
||||
`RULESET.CODERABBIT.GUARD.SCOPE=if a new @me open PR appears during final list, include it in the same guard pass before declaring all-open-PRs complete`
|
||||
`RULESET.TESTS.CODERABBIT_FIX=prefer exported-function behavioral tests over source-grep; lint-no-source-grep rejects readFileSync source assertions without allow-test-rule`
|
||||
`RULESET.WORKFLOW_MARKDOWN.FENCES=when editing shell snippets inside workflow markdown, preserve the opening language fence; malformed fence can create fresh CodeRabbit threads`
|
||||
`CI.GATE.issue-link-required=hard-fail if PR body lacks closes/fixes/resolves #<issue>`
|
||||
`CI.GATE.changeset-lint=hard-fail for user-facing code diffs unless .changeset/* or PR has no-changelog label`
|
||||
`CI.GATE.repair-sequence(PR)=create issue -> apply approval label -> edit PR body w/ closing keyword -> apply no-changelog if appropriate -> re-run checks`
|
||||
|
||||
@@ -347,6 +347,14 @@ gsd-tools query research-plan ← Research Provider: check cache, build
|
||||
|
||||
Agents always return a `RESEARCH.md` path, never raw fetched content. Context discipline is enforced through subagent isolation, compact provider output, and fetch-to-disk. See [ADR-0656](adr/0656-research-module-seam.md).
|
||||
|
||||
### Context Predicate Fact-Store (`src/context-predicates.cts`, ADR-1671)
|
||||
|
||||
The `CONTEXT.md` predicate fact-store — every backtick-wrapped `CLASS.subkey=value` declaration in the repo-root `CONTEXT.md` — has a compiled parser/selector seam (generated to `gsd-core/bin/lib/context-predicates.cjs` per ADR-457) reachable live via `gsd-tools query context-predicates --class|--prefix|--contains`. Fence-aware line skipping mirrors `markdown-sectionizer.cts`'s exported `scanFencedBlocks` delimiter-matching rule exactly (proven by a fence-skip parity test suite), but is scanned by a LOCAL, interleaved single pass rather than a call into that seam directly: fences and HTML comments must mutually suppress each other's open/close detection while either is active (a fence delimiter inside a real comment, or a comment token inside a real fence, must not falsely toggle the other construct), and that precedence cannot be resolved by two independent passes over `scanFencedBlocks`'s comment-blind output — see `src/context-predicates.cts`'s module doc comment.
|
||||
|
||||
`scripts/gen-context-index.cjs --check` is the CI drift-guard for the committed `docs/CONTEXT-INDEX.json` artifact: it fails on staleness between a fresh parse of `CONTEXT.md` and the committed file, and on any duplicate predicate ID. It is wired into `lint:generated-sync` (so `lint:ci`, so CI). `docs/CONTEXT-INDEX.json` is **generated — never hand-edit it**; regenerate with `gen-context-index.cjs --write` (also wired into `build`, after `build:lib`, and into `regen:derived`). The generator `require()`s the compiled `context-predicates.cjs`, so it must run after `build:lib` in any pipeline; `.github/workflows/test.yml` does this.
|
||||
|
||||
The committed index intentionally carries **no `line` field** for any predicate (ADR-1671 open question 4, resolved by #2928) — committed-but-uncompared metadata goes silently stale, the same defect class the drift-guard exists to catch, with the alarm removed. The live `gsd-tools query context-predicates` parse still returns `line`/`section` for callers that want to cite a source location. See [ADR-1671](adr/1671-dynamic-context-management-platform.md) and [CLI Tools Reference](CLI-TOOLS.md#query-context-predicates).
|
||||
|
||||
### CLI Tools (`gsd-core/bin/`)
|
||||
|
||||
Node.js CLI utility (`gsd-tools.cjs`) with domain modules split across `gsd-core/bin/lib/` (see [`docs/INVENTORY.md`](INVENTORY.md#cli-modules) for the authoritative roster):
|
||||
@@ -382,6 +390,7 @@ Node.js CLI utility (`gsd-tools.cjs`) with domain modules split across `gsd-core
|
||||
| `schema-detect.cjs` | Schema-drift detection for ORM patterns (Prisma, Drizzle, etc.) |
|
||||
| `profile-pipeline.cjs` | User behavioral profiling data pipeline, session file scanning |
|
||||
| `profile-output.cjs` | Profile rendering, USER-PROFILE.md and dev-preferences.md generation |
|
||||
| `context-predicates.cjs` | `CONTEXT.md` predicate fact-store parser/selector (ADR-1671, #2928); backs `query context-predicates` and `scripts/gen-context-index.cjs`'s `docs/CONTEXT-INDEX.json` drift guard; compiled from `src/context-predicates.cts` |
|
||||
| `loop-host-contract.cjs` | Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts; emitted by `scripts/gen-loop-host-contract.cjs` from workflow markers (ADR-894 §3); consumed by `gen-capability-registry.cjs` |
|
||||
| `capability-loader.cjs` | Runtime registry overlay loader (ADR-1244 D2) — `loadRegistry({ includeInstalled })` composes the frozen first-party registry with a validated installed overlay of third-party capability manifests read from global `$GSD_HOME/.gsd/capabilities/` and project `<projectRoot>/.gsd/capabilities/`; first-party always wins; load-time `engines.gsd` re-gate skips incompatible overlays with a warning; gate-kind hooks on skipped capabilities fail OPEN — no gate is injected; a loud warning (stderr + envelope `warnings`) names the load failure and the `gsd capability remove <id>` remediation (#2009) |
|
||||
| `capability-registry.cjs` | Generated central Capability Registry — role-partitioned index of all co-located capability declarations; emitted by `scripts/gen-capability-registry.cjs` (ADR-894 §5) |
|
||||
|
||||
@@ -322,6 +322,47 @@ This command is strictly read-only — no config writes, no disk mutation.
|
||||
|
||||
---
|
||||
|
||||
### `query context-predicates`
|
||||
|
||||
```bash
|
||||
node gsd-tools.cjs query context-predicates --class <CLASS> | --prefix <dotted.prefix> | --contains <text>
|
||||
```
|
||||
|
||||
Selector surface for the `CONTEXT.md` predicate fact-store (ADR-1671, #2928). Parses the repo-root `CONTEXT.md` **live** on every call via the compiled `context-predicates.cjs` — it never reads the committed `docs/CONTEXT-INDEX.json` (that artifact is a CI drift-guard byproduct, not a query source, so it can never go stale relative to the live predicates it answers about).
|
||||
|
||||
**Selectors** (at least one required; when more than one is given they are ANDed together):
|
||||
|
||||
| Flag | Type | Description |
|
||||
|---|---|---|
|
||||
| `--class <CLASS>` | string | Exact match on the predicate's class (the segment before the first `.`) |
|
||||
| `--prefix <dotted.prefix>` | string | Match predicate ids starting with this dotted prefix |
|
||||
| `--contains <text>` | string | Case-insensitive substring match against `id + ' ' + value` |
|
||||
|
||||
Each flag also accepts the inline-assignment form (`--contains=<text>`), which is the escape
|
||||
hatch for a flag-shaped value the space-separated form cannot express — e.g.
|
||||
`--contains=--dry-run` to search for the literal substring `--dry-run`. The space-separated form
|
||||
(`--contains --dry-run`) always reads a following `--...` token as a missing value, by design.
|
||||
|
||||
**Output JSON:**
|
||||
|
||||
```json
|
||||
{
|
||||
"matched": 2,
|
||||
"predicates": [
|
||||
{ "id": "RULESET.EXAMPLE", "klass": "RULESET", "value": "…", "line": 42, "section": "Glossary" }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Description |
|
||||
|---|---|---|
|
||||
| `matched` | number | Count of predicates satisfying all given selectors |
|
||||
| `predicates` | array | Each entry is a live `Predicate` — `id`, `klass`, `value`, `line` (1-based source line), `section` (nearest enclosing heading) |
|
||||
|
||||
This command is strictly read-only — no config writes, no disk mutation. See [ADR-1671](adr/1671-dynamic-context-management-platform.md) and [Architecture — CLI Tools](ARCHITECTURE.md#cli-tools-gsd-corebin).
|
||||
|
||||
---
|
||||
|
||||
## Model Resolution
|
||||
|
||||
```bash
|
||||
@@ -739,6 +780,7 @@ User-facing entry point: `/gsd-graphify` (see [Command Reference](COMMANDS.md#gs
|
||||
| Audit | `lib/audit.cjs` | Phase/milestone audit queue handlers; `audit-open` helper |
|
||||
| GSD2 Import | `lib/gsd2-import.cjs` | Reverse-migration importer from GSD-2 projects (backs `/gsd-import --from-gsd2`) |
|
||||
| Intel | `lib/intel.cjs` | Queryable codebase intelligence index (backs `/gsd-map-codebase --query`) |
|
||||
| Context Predicates | `lib/context-predicates.cjs` | `CONTEXT.md` predicate fact-store parser/selector (ADR-1671, #2928) — backs `query context-predicates` and `scripts/gen-context-index.cjs`'s `docs/CONTEXT-INDEX.json` drift guard |
|
||||
| Capability State | `lib/capability-state.cjs` | Capability-state resolver — composes install profile, surface, and config into per-capability `enabled`/`active` view |
|
||||
| Capability Writer | `lib/capability-writer.cjs` | Capability-state writer (ADR-1213) — write-side inverse; projects `--on`/`--off`/`--gate` onto surface + config substrates then re-resolves |
|
||||
| Worktree Base Ref | `lib/worktree-base-ref.cjs` | Worktree fork-base detection and `worktree base-check` / `set-baseref` commands (#683) |
|
||||
|
||||
2104
docs/CONTEXT-INDEX.json
Normal file
2104
docs/CONTEXT-INDEX.json
Normal file
File diff suppressed because it is too large
Load Diff
@@ -346,6 +346,7 @@
|
||||
"config-types.cjs",
|
||||
"config.cjs",
|
||||
"configuration.cjs",
|
||||
"context-predicates.cjs",
|
||||
"context-utilization.cjs",
|
||||
"core-utils.cjs",
|
||||
"coverage.cjs",
|
||||
|
||||
@@ -446,6 +446,7 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
|
||||
| `config-types.cjs` | TypeScript type definitions for the `model_policy` config block — `ModelPolicyConfig`, `TierEntry`, `RuntimeTiers`; compiled from `src/config-types.cts` at publish time (ADR-457) |
|
||||
| `config.cjs` | `config.json` read/write, section initialization; imports validator from `config-schema.cjs` |
|
||||
| `configuration.cjs` | Configuration Module — legacy-key normalization, defaults merge, and explicit on-disk migration; pure normalization primitives consumed by `config-loader.cjs` and `config-schema.cjs` (loadConfig extracted to config-loader per ADR-857 #885) |
|
||||
| `context-predicates.cjs` | CONTEXT.md predicate fact-store parser (ADR-1671, #2928) — pure `parsePredicates` (extracts every backtick-wrapped `CLASS.subkey=value` declaration, fence/HTML-comment-aware), `selectPredicates` (class/prefix/contains selectors, ANDed), and `buildIndex` (deterministic, line-free artifact shape); backs both `gsd_run query context-predicates` and `scripts/gen-context-index.cjs`'s docs/CONTEXT-INDEX.json drift guard. Compiled from `src/context-predicates.cts` |
|
||||
| `context-utilization.cjs` | Pure classifier for `gsd-health --context` — turns (tokensUsed, contextWindow) into a `{ percent, state }` triage result against the 60%/70% fracture-point thresholds (#2792) |
|
||||
| `core-utils.cjs` | Shared low-level utilities — POSIX path normalization, sub-repo/subdirectory scanning, phase file stats, slug/one-liner/plan-id helpers, time-ago (extracted from `core.cjs`, ADR-857) |
|
||||
| `core.cjs` | Shared utilities and runtime fallbacks; compatibility re-exports for planning-workspace and I/O (`io.cjs`) helpers |
|
||||
|
||||
@@ -118,12 +118,15 @@ That is a point-in-time true-up, not a fix. Per Open question 4, the index is ke
|
||||
|
||||
Prototype scope notes: the parser is intentionally self-contained for the example; production should consume the compiled `markdown-sectionizer` seam, live under `src/` → `bin/lib/`, and be drift-guarded by a generator wired into the build **after** `build:lib`.
|
||||
|
||||
**Done (#2928).** Production landed under `src/context-predicates.cts` → `gsd-core/bin/lib/context-predicates.cjs` (ADR-457 build-at-publish). Fence-aware line skipping mirrors `markdown-sectionizer.cts`'s exported `scanFencedBlocks` delimiter-matching rule exactly (byte-for-behavior parity proven by a dedicated test suite) via a LOCAL, interleaved single pass, rather than a call into that seam directly: a two-pass design (mask comments, then call `scanFencedBlocks`, or the reverse) cannot correctly resolve mutual precedence between HTML comments and fences in both directions — a fence delimiter inside a real comment (with no later real closer) was found to falsely skip the rest of the file to EOF, and the converse ordering falsely let a comment token inside a real fence leak past the fence's own close — so the two constructs are scanned together, each suppressing the other's open/close detection while active (post-#2928-review fix; see `src/context-predicates.cts`'s module doc comment). `scripts/gen-context-index.cjs --check`/`--write` is wired into `lint:generated-sync` (so `lint:ci`, CI-gated) and into `build` (after `build:lib`) and `regen:derived`; the selector is exposed live via `gsd-tools query context-predicates --class|--prefix|--contains`.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Fragment unit: separate files vs in-file section markers?
|
||||
2. Build-time emission vs run-time assembly as the primary surface during migration (double-write vs per-workflow cutover)?
|
||||
3. Whether/when to invest in per-runtime native channels (skills, MCP) above the universal file floor.
|
||||
4. **Index keying: stable IDs vs baked `line` numbers.** `CONTEXT-INDEX.json` stores each predicate's `line`, so `--check` re-drifts on *any* `CONTEXT.md` line shift — a typo fix three sections up fails the gate. Phase 1 promotes `--check` to a CI gate, where that makes it routinely red for reasons unrelated to predicate integrity. Keying the comparison on stable IDs, with `line` retained as non-compared metadata, is the candidate fix. Raised by @davesienkowski (#1671, 2026-06-25) and confirmed on `next` 2026-07-31.
|
||||
|
||||
**Resolved by #2928 — index keying: stable IDs, with no `line` field at all.** Question 4 asked stable IDs vs baked `line` numbers: `CONTEXT-INDEX.json` stored each predicate's `line`, so `--check` re-drifted on *any* `CONTEXT.md` line shift — a typo fix three sections up failed the gate. Raised by @davesienkowski (#1671, 2026-06-25). The shipped resolution is **stronger than the option originally proposed** (keying the comparison on stable IDs with `line` retained as non-compared metadata): the committed `ContextIndex.predicates` entries carry **no `line` field at all**. Committed-but-uncompared metadata goes silently stale — the same defect class the drift-guard exists to catch, with the alarm removed — so it was dropped from the committed artifact rather than merely excluded from the comparison. `line` is still returned by the live `parsePredicates`/`gsd-tools query context-predicates` result for callers that want to cite a source location; only the committed `docs/CONTEXT-INDEX.json` shape omits it.
|
||||
|
||||
**Resolved by other work — not carried as open.** A fourth question was proposed in review (#1671, 2026-06-25): *what populates the eval-gate assertion set, and is it graded exogenously?* Since that review, the answer has landed as first-class predicate classes rather than remaining a design gap: `PROBE.principle` (`verifier-reach-equals-spec-reach`), `PROBE.family` (edge-probe + prohibition-probe + ui-consideration-probe), `PROBE.protocol` (recall → precision), and `PROHIB.judgment-tier` (exogenous grading) — see ADR-550 D4/D7 and ADR-1606. The `PROHIB.*` predicates live in the same `CONTEXT.md` store this ADR formalizes, which is the single-store property that review asked for.
|
||||
|
||||
|
||||
@@ -163,6 +163,18 @@ Before any AI agent writes a single line of code or docs, it must read:
|
||||
|
||||
If you are dispatching an AI agent, include these reads in the agent's prompt explicitly. An agent that invents synonyms for `CONTEXT.md` vocabulary or contradicts an accepted ADR without flagging it has failed the pre-work requirement.
|
||||
|
||||
**Citing a machine-oriented predicate in a brief.** `CONTEXT.md`'s `KEY.SUBKEY=value` predicates
|
||||
(see below) must be cited by ID verbatim, never paraphrased (`META.RULE.brief-must-cite-doc`,
|
||||
`META.RULE.brief-no-paraphrase`). Rather than grepping the file by eye for the predicate set a
|
||||
brief needs, pull it with the selector, which parses the live file on every call:
|
||||
|
||||
```bash
|
||||
node gsd-tools.cjs query context-predicates --class <CLASS> | --prefix <dotted.prefix> | --contains <text>
|
||||
```
|
||||
|
||||
See [`query context-predicates`](CLI-TOOLS.md#query-context-predicates) for the full flag and
|
||||
output reference.
|
||||
|
||||
**In the PR body**, state which ADR or standards section was followed. If using an AI assistant, this statement is your responsibility as the author — not the agent's.
|
||||
|
||||
### Worktree isolation
|
||||
|
||||
@@ -234,6 +234,8 @@ export default tseslint.config(
|
||||
'gsd-core/bin/lib/state-io.cjs',
|
||||
'gsd-core/bin/lib/external-descriptor-trust.cjs',
|
||||
'gsd-core/bin/lib/mcp-server.cjs',
|
||||
// ADR-1671 (#2928): tsc-generated runtime artifact — lint the src/context-predicates.cts source.
|
||||
'gsd-core/bin/lib/context-predicates.cjs',
|
||||
],
|
||||
},
|
||||
|
||||
|
||||
@@ -2605,6 +2605,133 @@ function dispatchOverlayCapabilityCommand({ command, args, cwd, raw, error, load
|
||||
}
|
||||
}
|
||||
|
||||
// `gsd_run query context-predicates` — selector surface for the CONTEXT.md
|
||||
// predicate fact-store (ADR-1671, #2928 Phase 1 row S9). Parses the
|
||||
// repo-root CONTEXT.md LIVE via the compiled context-predicates.cjs on
|
||||
// every call — it never reads the committed docs/CONTEXT-INDEX.json (that
|
||||
// artifact is a CI drift-guard byproduct, not a query source, so it can
|
||||
// never go stale relative to the live predicates it answers about).
|
||||
//
|
||||
// Selectors: --class <CLASS>, --prefix <dotted.prefix>, --contains <text>.
|
||||
// At least one is required. When more than one is given they are ANDed
|
||||
// together — the same documented precedence selectPredicates() itself
|
||||
// implements (see context-predicates.cjs doc comment: "Select predicates
|
||||
// by one or more optional criteria (ANDed together)"); no selector is
|
||||
// silently dropped or overridden by another.
|
||||
//
|
||||
// Flag parsing mirrors routePromptBudget's Map-based flagMap: the three
|
||||
// known flags are recognized in both the space-separated `--flag value`
|
||||
// form and the inline-assignment `--flag=value` form (the latter is the
|
||||
// escape hatch for a flag-shaped selector value, e.g. `--contains=--dry-run`
|
||||
// — #2928 review finding C; the space-separated form has no such escape by
|
||||
// design, since a following `--...` token always reads as a missing value).
|
||||
// `--class=` (empty value) and `--class==A` (double-equals typo shape)
|
||||
// are rejected the same way under either form. On a duplicate flag the
|
||||
// FIRST occurrence wins (`Map.set` only fires when the key is absent),
|
||||
// which is deterministic across repeated invocations with identical argv.
|
||||
//
|
||||
// Prototype-pollution safety: selector values are only ever compared via
|
||||
// `===`/`.startsWith()`/`.includes()` against ordinary string fields — this
|
||||
// route never uses a user-supplied string as an object property key
|
||||
// (`obj[userValue] = ...`), so `--class __proto__` / `constructor` /
|
||||
// `prototype` are just non-matching ordinary strings, not property-access
|
||||
// vectors. `flagMap` itself is a `Map`, immune to prototype pollution by
|
||||
// construction.
|
||||
function routeContextPredicates({ args, cwd, raw, error }) {
|
||||
const { parsePredicates, selectPredicates } = require('./lib/context-predicates.cjs');
|
||||
|
||||
const KNOWN_FLAGS = new Set(['--class', '--prefix', '--contains']);
|
||||
const flagMap = new Map();
|
||||
for (let i = 1; i < args.length; i++) {
|
||||
const current = args[i];
|
||||
if (typeof current !== 'string' || !current.startsWith('--')) continue;
|
||||
|
||||
// Inline-assignment escape hatch (`--flag=value`, mirrors the `--config-dir=`/
|
||||
// `--runtime=` convention in routeUpdateContext elsewhere in this file). This is
|
||||
// the ONLY way to pass a flag-shaped selector value (e.g. searching CONTEXT.md
|
||||
// for the literal substring "--dry-run"): the space-separated form below always
|
||||
// treats a following `--...` token as a missing value, by design, so it has no
|
||||
// escape hatch on its own (#2928 review finding C).
|
||||
const eqFlag = [...KNOWN_FLAGS].find((f) => current.startsWith(`${f}=`));
|
||||
if (eqFlag) {
|
||||
const value = current.slice(eqFlag.length + 1);
|
||||
// Reject an empty value (`--class=`) and the `--class==A` double-equals typo
|
||||
// shape (a value starting with `=`) the same way the pre-existing malformed-
|
||||
// assignment behavior did — never silently accept "=A" as a literal value.
|
||||
if (value === '' || value.startsWith('=')) {
|
||||
error(`context-predicates: ${eqFlag} requires a non-empty value`, ERROR_REASON.USAGE);
|
||||
return;
|
||||
}
|
||||
if (!flagMap.has(eqFlag)) flagMap.set(eqFlag, value);
|
||||
continue;
|
||||
}
|
||||
|
||||
if (!KNOWN_FLAGS.has(current)) {
|
||||
error(`Unknown flag for context-predicates: ${current}`, ERROR_REASON.USAGE);
|
||||
return;
|
||||
}
|
||||
const next = args[i + 1];
|
||||
if (next === undefined || next.startsWith('--')) {
|
||||
if (!flagMap.has(current)) flagMap.set(current, null);
|
||||
continue;
|
||||
}
|
||||
if (!flagMap.has(current)) flagMap.set(current, next);
|
||||
i++;
|
||||
}
|
||||
|
||||
const hasClass = flagMap.has('--class');
|
||||
const hasPrefix = flagMap.has('--prefix');
|
||||
const hasContains = flagMap.has('--contains');
|
||||
|
||||
if (!hasClass && !hasPrefix && !hasContains) {
|
||||
error(
|
||||
'Usage: gsd-tools query context-predicates --class <CLASS> | --prefix <dotted.prefix> | --contains <text> ' +
|
||||
'(selectors are ANDed when combined)',
|
||||
ERROR_REASON.USAGE,
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
const requireNonEmpty = (flagName, rawValue) => {
|
||||
if (rawValue === null || rawValue === undefined || rawValue.trim() === '') {
|
||||
error(`context-predicates: ${flagName} requires a non-empty value`, ERROR_REASON.USAGE);
|
||||
return null;
|
||||
}
|
||||
return rawValue;
|
||||
};
|
||||
|
||||
const opts = {};
|
||||
if (hasClass) {
|
||||
const v = requireNonEmpty('--class', flagMap.get('--class'));
|
||||
if (v === null) return;
|
||||
opts.klass = v;
|
||||
}
|
||||
if (hasPrefix) {
|
||||
const v = requireNonEmpty('--prefix', flagMap.get('--prefix'));
|
||||
if (v === null) return;
|
||||
opts.prefix = v;
|
||||
}
|
||||
if (hasContains) {
|
||||
const v = requireNonEmpty('--contains', flagMap.get('--contains'));
|
||||
if (v === null) return;
|
||||
opts.contains = v;
|
||||
}
|
||||
|
||||
const contextMdPath = path.join(__dirname, '..', '..', 'CONTEXT.md');
|
||||
let markdown;
|
||||
try {
|
||||
markdown = fs.readFileSync(contextMdPath, 'utf8');
|
||||
} catch (err) {
|
||||
error(`context-predicates: cannot read ${contextMdPath}: ${err && err.message}`, ERROR_REASON.USAGE);
|
||||
return;
|
||||
}
|
||||
|
||||
const { predicates } = parsePredicates(markdown);
|
||||
const matches = selectPredicates(predicates, opts);
|
||||
|
||||
output({ matched: matches.length, predicates: matches }, raw);
|
||||
}
|
||||
|
||||
function routeUpdateContext({ args, cwd, raw, error }) {
|
||||
// #498: resolve the installed GSD version, scope, runtime, and config dir
|
||||
// for /gsd:update. Replaces ~280 lines of inline bash in update.md with a
|
||||
@@ -2940,6 +3067,7 @@ const HOST_COMMAND_ROUTERS = {
|
||||
'restore-custom-files': routeRestoreCustomFiles,
|
||||
'from-gsd2': routeFromGsd2,
|
||||
'prompt-budget': routePromptBudget,
|
||||
'context-predicates': routeContextPredicates,
|
||||
'review-lane': routeReviewLane,
|
||||
'update-context': routeUpdateContext,
|
||||
'classify-confidence': routeClassifyConfidence,
|
||||
@@ -3143,6 +3271,90 @@ function runWithTimeout(argv) {
|
||||
|
||||
// ─── CLI Router ───────────────────────────────────────────────────────────────
|
||||
|
||||
// Top-level usage string — emitted by `gsd-tools` (no args) and by
|
||||
// `gsd-tools --help` / any `--help` request below.
|
||||
// CR feedback: the command list must enumerate every top-level command
|
||||
// supported by the dispatcher so `--help` is actually useful for
|
||||
// discovery; previously it was a partial subset that didn't include
|
||||
// phase / roadmap / milestone / progress / etc.
|
||||
//
|
||||
// Module-scoped (not function-local) so it can be exported and compared
|
||||
// against HOST_COMMAND_ROUTERS in a parity test (DEFECT.GENERATIVE-FIX) —
|
||||
// this string and HOST_COMMAND_ROUTERS/SKIP_ROOT_RESOLUTION are three
|
||||
// independently hand-maintained sites and nothing previously caught them
|
||||
// drifting apart when a query command was added to only one or two.
|
||||
const TOP_LEVEL_USAGE = 'Usage: gsd-tools <command> [args] [--raw] [--pick <field>] [--cwd <path>] [--ws <name>] [--json-errors]\n' +
|
||||
'Commands: agent, agent-skills, assumption-delta, audit-open, audit-uat, check, check-commit, commit, commit-to-subrepo, pr-subrepo, ' +
|
||||
'config-ensure-section, config-get, config-new-project, config-path, config-set, migrate-config, normalize-test-command, ' +
|
||||
'context-predicates, current-timestamp, detect-custom-files, docs-init, drift-guard, effort, extract-messages, find-phase, ' +
|
||||
'from-gsd2, frontmatter, gap-analysis, generate-claude-md, generate-claude-profile, ' +
|
||||
'generate-dev-preferences, generate-slug, graphify, history-digest, init, intel, ' +
|
||||
'capability, classify-confidence, git, learnings, list-seeds, list-todos, loop, milestone, package-legitimacy, phase, phase-plan-index, phases, profile-questionnaire, ' +
|
||||
'profile-sample, progress, project-instruction-file, prompt-budget, quick-tasks-append, requirements, research-plan, research-store, resolve-granularity, resolve-model, restore-custom-files, roadmap, scaffold, smart-entry, state, ' +
|
||||
'config-set-model-profile, dispatch-isolation, dispatch-should-flatten, estimate-calibrate, estimate-calibration, estimate-check, resolve-dispatch-type, ' +
|
||||
'resolve-execution, review-lane, skill-manifest, state-snapshot, stats, summary-extract, teams-status, todo, uat, update-context, verification, websearch, windows, ' +
|
||||
'task, template, user-story, validate, verify, verify-path-exists, verify-summary, eval, workstream, worktree\n\n' +
|
||||
'Global flags:\n' +
|
||||
' --raw Emit raw output without post-processing\n' +
|
||||
' --pick <field> Extract a single field from JSON output (dot/bracket notation)\n' +
|
||||
' --cwd <path> Override working directory for project-root resolution\n' +
|
||||
' --ws <name> Override active workstream (or set GSD_WORKSTREAM)\n' +
|
||||
' --json-errors Emit structured JSON error objects on stderr (or set GSD_JSON_ERRORS=1)\n\n' +
|
||||
'For command-specific argument requirements, invoke the command without args ' +
|
||||
'(e.g. `gsd-tools phase add`) — the resulting error lists what is required.';
|
||||
|
||||
// Multi-repo guard: resolve project root for commands that read/write .planning/.
|
||||
// Skip for pure-utility commands that don't touch .planning/ to avoid unnecessary
|
||||
// filesystem traversal on every invocation.
|
||||
// 'loop' and 'capability' are intentionally NOT in SKIP_ROOT_RESOLUTION.
|
||||
// Both are registry/config queries that resolve activation via
|
||||
// .planning/config.json; they need the project root (cwd) for correct
|
||||
// `when` key resolution. If one is ever moved to SKIP_ROOT_RESOLUTION,
|
||||
// move the other at the same time (keep them consistent).
|
||||
//
|
||||
// Module-scoped for the same reason as TOP_LEVEL_USAGE above — kept
|
||||
// module-private and exposed to the dispatch-table/help-string/skip-list
|
||||
// parity test only through the read-only skipsRootResolution() predicate
|
||||
// below (never as the live Set itself; see that function's doc comment).
|
||||
const SKIP_ROOT_RESOLUTION = new Set([
|
||||
'generate-slug', 'current-timestamp', 'verify-path-exists',
|
||||
// #2844: verify-summary was previously skipped, leaving relative file-claim
|
||||
// paths resolved against the raw process.cwd() — invoking from a subdirectory
|
||||
// manufactured "missing files" on an otherwise-correct SUMMARY. It now goes
|
||||
// through findProjectRoot so claims resolve against the project root.
|
||||
'template', 'frontmatter', 'detect-custom-files',
|
||||
// #1854: restore-custom-files operates on a runtime config dir passed
|
||||
// explicitly via --config-dir; it never reads .planning/.
|
||||
'restore-custom-files',
|
||||
'worktree', 'prompt-budget',
|
||||
// context-predicates is a pure repo-root CONTEXT.md read (like
|
||||
// prompt-budget); it never touches .planning/, so it needs no project
|
||||
// root resolution and must work from any cwd (including one with no
|
||||
// .planning/ directory at all).
|
||||
'context-predicates',
|
||||
'research-store', 'research-plan', 'package-legitimacy', 'classify-confidence',
|
||||
'user-story', // pure string validation — no .planning/ access needed
|
||||
// #1529: pure runtime→filename projection via getProjectInstructionFile; no
|
||||
// .planning/ access needed, and resolving project root would break workflow
|
||||
// invocations that run before .planning/ exists (new-project Step 1).
|
||||
'project-instruction-file',
|
||||
// #1579: eval.score is pure arithmetic (covered/total + infra weights); it
|
||||
// needs no .planning/ access, so skip the findProjectRoot traversal.
|
||||
'eval',
|
||||
]);
|
||||
|
||||
// Read-only accessor for SKIP_ROOT_RESOLUTION (DEFECT.MUTABLE-EXPORTED-SET,
|
||||
// #2928 review). The Set above stays module-private and mutable internally
|
||||
// (main() only ever calls .has() on it), but exporting the live Set directly
|
||||
// would let any importer call .add()/.delete() on it — Object.freeze() does
|
||||
// not lock Set.prototype.add/delete, so freezing the instance would not have
|
||||
// closed this — and silently change dispatch behavior for every caller in the
|
||||
// process. Export this predicate instead; it exposes membership without
|
||||
// exposing a mutation surface.
|
||||
function skipsRootResolution(command) {
|
||||
return SKIP_ROOT_RESOLUTION.has(command);
|
||||
}
|
||||
|
||||
async function main() {
|
||||
let args = process.argv.slice(2);
|
||||
|
||||
@@ -3276,30 +3488,6 @@ async function main() {
|
||||
}
|
||||
}
|
||||
|
||||
// Top-level usage string — emitted by `gsd-tools` (no args) and by
|
||||
// `gsd-tools --help` / any `--help` request below.
|
||||
// CR feedback: the command list must enumerate every top-level command
|
||||
// supported by the dispatcher so `--help` is actually useful for
|
||||
// discovery; previously it was a partial subset that didn't include
|
||||
// phase / roadmap / milestone / progress / etc.
|
||||
const TOP_LEVEL_USAGE = 'Usage: gsd-tools <command> [args] [--raw] [--pick <field>] [--cwd <path>] [--ws <name>] [--json-errors]\n' +
|
||||
'Commands: agent, agent-skills, assumption-delta, audit-open, audit-uat, check, check-commit, commit, commit-to-subrepo, pr-subrepo, ' +
|
||||
'config-ensure-section, config-get, config-new-project, config-path, config-set, migrate-config, normalize-test-command, ' +
|
||||
'current-timestamp, detect-custom-files, docs-init, drift-guard, effort, extract-messages, find-phase, ' +
|
||||
'from-gsd2, frontmatter, gap-analysis, generate-claude-md, generate-claude-profile, ' +
|
||||
'generate-dev-preferences, generate-slug, graphify, history-digest, init, intel, ' +
|
||||
'capability, classify-confidence, git, learnings, list-seeds, list-todos, loop, milestone, package-legitimacy, phase, phase-plan-index, phases, profile-questionnaire, ' +
|
||||
'profile-sample, progress, project-instruction-file, prompt-budget, quick-tasks-append, requirements, research-plan, research-store, resolve-granularity, resolve-model, restore-custom-files, roadmap, scaffold, smart-entry, state, ' +
|
||||
'task, template, user-story, validate, verify, verify-path-exists, verify-summary, eval, workstream, worktree\n\n' +
|
||||
'Global flags:\n' +
|
||||
' --raw Emit raw output without post-processing\n' +
|
||||
' --pick <field> Extract a single field from JSON output (dot/bracket notation)\n' +
|
||||
' --cwd <path> Override working directory for project-root resolution\n' +
|
||||
' --ws <name> Override active workstream (or set GSD_WORKSTREAM)\n' +
|
||||
' --json-errors Emit structured JSON error objects on stderr (or set GSD_JSON_ERRORS=1)\n\n' +
|
||||
'For command-specific argument requirements, invoke the command without args ' +
|
||||
'(e.g. `gsd-tools phase add`) — the resulting error lists what is required.';
|
||||
|
||||
if (!command) {
|
||||
error(TOP_LEVEL_USAGE);
|
||||
}
|
||||
@@ -3326,35 +3514,6 @@ async function main() {
|
||||
}
|
||||
}
|
||||
|
||||
// Multi-repo guard: resolve project root for commands that read/write .planning/.
|
||||
// Skip for pure-utility commands that don't touch .planning/ to avoid unnecessary
|
||||
// filesystem traversal on every invocation.
|
||||
// 'loop' and 'capability' are intentionally NOT in SKIP_ROOT_RESOLUTION.
|
||||
// Both are registry/config queries that resolve activation via
|
||||
// .planning/config.json; they need the project root (cwd) for correct
|
||||
// `when` key resolution. If one is ever moved to SKIP_ROOT_RESOLUTION,
|
||||
// move the other at the same time (keep them consistent).
|
||||
const SKIP_ROOT_RESOLUTION = new Set([
|
||||
'generate-slug', 'current-timestamp', 'verify-path-exists',
|
||||
// #2844: verify-summary was previously skipped, leaving relative file-claim
|
||||
// paths resolved against the raw process.cwd() — invoking from a subdirectory
|
||||
// manufactured "missing files" on an otherwise-correct SUMMARY. It now goes
|
||||
// through findProjectRoot so claims resolve against the project root.
|
||||
'template', 'frontmatter', 'detect-custom-files',
|
||||
// #1854: restore-custom-files operates on a runtime config dir passed
|
||||
// explicitly via --config-dir; it never reads .planning/.
|
||||
'restore-custom-files',
|
||||
'worktree', 'prompt-budget',
|
||||
'research-store', 'research-plan', 'package-legitimacy', 'classify-confidence',
|
||||
'user-story', // pure string validation — no .planning/ access needed
|
||||
// #1529: pure runtime→filename projection via getProjectInstructionFile; no
|
||||
// .planning/ access needed, and resolving project root would break workflow
|
||||
// invocations that run before .planning/ exists (new-project Step 1).
|
||||
'project-instruction-file',
|
||||
// #1579: eval.score is pure arithmetic (covered/total + infra weights); it
|
||||
// needs no .planning/ access, so skip the findProjectRoot traversal.
|
||||
'eval',
|
||||
]);
|
||||
if (!SKIP_ROOT_RESOLUTION.has(command)) {
|
||||
cwd = findProjectRoot(cwd);
|
||||
}
|
||||
@@ -3511,5 +3670,13 @@ if (require.main === module) {
|
||||
// synthetic registry + requireModule injections.
|
||||
// ADR-1244 Phase 5: export dispatchOverlayCapabilityCommand + defaultRequireFromInstallRoot for
|
||||
// the third-party overlay dispatch + install-root confinement tests.
|
||||
module.exports = { dispatchCapabilityCommand, dispatchOverlayCapabilityCommand, defaultRequireFromInstallRoot, dispatchHostCommand, HOST_COMMAND_ROUTERS };
|
||||
module.exports = {
|
||||
dispatchCapabilityCommand,
|
||||
dispatchOverlayCapabilityCommand,
|
||||
defaultRequireFromInstallRoot,
|
||||
dispatchHostCommand,
|
||||
HOST_COMMAND_ROUTERS,
|
||||
TOP_LEVEL_USAGE,
|
||||
skipsRootResolution,
|
||||
};
|
||||
|
||||
|
||||
@@ -84,16 +84,17 @@
|
||||
"check:identity-drift": "node scripts/lint-package-identity-drift.cjs",
|
||||
"check:phase-id-drift": "node scripts/lint-phase-id-drift.cjs",
|
||||
"check:integrity": "node scripts/check-npm-integrity.cjs",
|
||||
"build": "npm run generate:identity && npm run build:lib && npm run gen:plugin-skills && npm run gen:loop-host-contract && npm run gen:capability-registry && npm run build:hooks",
|
||||
"build": "npm run generate:identity && npm run build:lib && npm run gen:context-index && npm run gen:plugin-skills && npm run gen:loop-host-contract && npm run gen:capability-registry && npm run build:hooks",
|
||||
"build:hooks": "node scripts/build-hooks.js",
|
||||
"build:lib": "tsc -p tsconfig.build.json",
|
||||
"generate:identity": "node scripts/generate-package-identity.cjs",
|
||||
"gen:context-index": "node scripts/gen-context-index.cjs --write",
|
||||
"gen:loop-host-contract": "node scripts/gen-loop-host-contract.cjs --write",
|
||||
"gen:plugin-skills": "node scripts/gen-plugin-skills.cjs --write",
|
||||
"gen:capability-registry": "node scripts/gen-capability-registry.cjs --write",
|
||||
"gen:registry": "node scripts/gen-registry.cjs --write",
|
||||
"gen:install-tree": "node scripts/gen-install-tree-fixtures.cjs",
|
||||
"regen:derived": "npm run build && npm run gen:registry && node scripts/gen-adr-index.cjs --write && node scripts/gen-capability-matrix.cjs --write && node scripts/gen-inventory-manifest.cjs --write && node scripts/sync-manifest-versions.cjs && npm run gen:install-tree",
|
||||
"regen:derived": "npm run build && npm run gen:registry && node scripts/gen-adr-index.cjs --write && node scripts/gen-capability-matrix.cjs --write && node scripts/gen-inventory-manifest.cjs --write && node scripts/gen-context-index.cjs --write && node scripts/sync-manifest-versions.cjs && npm run gen:install-tree",
|
||||
"validate:registry": "node scripts/validate-registry.cjs",
|
||||
"prepack": "npm run build:lib",
|
||||
"prepare": "npm run build:lib",
|
||||
@@ -112,7 +113,7 @@
|
||||
"lint:test-file-count": "node scripts/lint-test-file-count.cjs",
|
||||
"lint:pr-checks": "node scripts/lint-pr-check-project-dir.cjs",
|
||||
"lint:changeset": "node scripts/changeset/lint.cjs",
|
||||
"lint:generated-sync": "node scripts/gen-capability-registry.cjs --check && node scripts/gen-loop-host-contract.cjs --check && node scripts/gen-capability-matrix.cjs --check && node scripts/sync-manifest-versions.cjs --check && node scripts/gen-inventory-manifest.cjs --check && node scripts/generate-package-identity.cjs --check && node scripts/gen-plugin-skills.cjs --check && node scripts/gen-registry.cjs --check && node scripts/gen-adr-index.cjs --check && node scripts/check-glossary-refs.cjs --check && node scripts/lint-compiled-artifact-sync.cjs --check",
|
||||
"lint:generated-sync": "node scripts/gen-capability-registry.cjs --check && node scripts/gen-loop-host-contract.cjs --check && node scripts/gen-capability-matrix.cjs --check && node scripts/sync-manifest-versions.cjs --check && node scripts/gen-inventory-manifest.cjs --check && node scripts/generate-package-identity.cjs --check && node scripts/gen-plugin-skills.cjs --check && node scripts/gen-registry.cjs --check && node scripts/gen-adr-index.cjs --check && node scripts/check-glossary-refs.cjs --check && node scripts/lint-compiled-artifact-sync.cjs --check && node scripts/gen-context-index.cjs --check",
|
||||
"lint:docs": "node scripts/lint-docs-required.cjs",
|
||||
"lint:legacy-name": "node scripts/lint-legacy-dir-name.cjs",
|
||||
"ci:test-scope": "node scripts/ci-test-scope.cjs",
|
||||
|
||||
448
scripts/gen-context-index.cjs
Normal file
448
scripts/gen-context-index.cjs
Normal file
@@ -0,0 +1,448 @@
|
||||
#!/usr/bin/env node
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* gen-context-index.cjs — generates docs/CONTEXT-INDEX.json from the
|
||||
* predicate declarations (`` `CLASS.subkey=value` `` lines) in the
|
||||
* repo-root CONTEXT.md.
|
||||
*
|
||||
* Usage:
|
||||
* node scripts/gen-context-index.cjs # print to stdout
|
||||
* node scripts/gen-context-index.cjs --write # write docs/CONTEXT-INDEX.json
|
||||
* node scripts/gen-context-index.cjs --check # exit 1 if committed file is stale
|
||||
* node scripts/gen-context-index.cjs --check --json # same, + typed report on stdout
|
||||
* node scripts/gen-context-index.cjs --write --context-path <p> --index-path <p>
|
||||
* # override the two hardcoded
|
||||
* # repo-root paths (tests use
|
||||
* # this to point the real CLI at
|
||||
* # a temp fixture tree with no fs
|
||||
* # monkeypatching required)
|
||||
*
|
||||
* ADR-1671 ("Dynamic context management platform", #2928) Phase 1 commits 1+3.
|
||||
* The generated artifact is plain JSON (docs/CONTEXT-INDEX.json), mirroring
|
||||
* docs/INVENTORY-MANIFEST.json's precedent: a committed, generated,
|
||||
* `--check`-guarded JSON manifest that is NOT runtime code. It previously
|
||||
* lived at gsd-core/bin/lib/context-index.cjs — a shipped runtime module is
|
||||
* the wrong place for ~120 KB of arbitrary CONTEXT.md prose: it tripped both
|
||||
* tests/cline-install.test.cjs's leaked-`.claude`-path guard and
|
||||
* tests/package-name-single-source.test.cjs's hardcoded-package-name guard,
|
||||
* both true positives against runtime-code content scanning. Moving the
|
||||
* artifact to docs/ (never scanned as runtime code) fixes both without
|
||||
* weakening either guard.
|
||||
*
|
||||
* Depends on the COMPILED gsd-core/bin/lib/context-predicates.cjs
|
||||
* (src/context-predicates.cts, built by `npm run build:lib`). This is safe
|
||||
* for CI: `.github/workflows/test.yml` runs `build:lib` before `lint:ci`.
|
||||
*
|
||||
* `--check --json` (CONTRIBUTING.md "Prohibited: Raw Text Matching on Test
|
||||
* Outputs"): emits `{ ok, reason, duplicates, count, classes }` — `reason` is
|
||||
* always one of the frozen `REASON` values (exported below) so tests assert
|
||||
* `report.reason === REASON.FAIL_X` instead of regex-matching the prose the
|
||||
* non-JSON mode still prints for human operators. The human-readable output
|
||||
* is unchanged.
|
||||
*/
|
||||
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const CONTEXT_PREDICATES_LIB_PATH = path.join(ROOT, 'gsd-core', 'bin', 'lib', 'context-predicates.cjs');
|
||||
const CONTEXT_PATH = path.join(ROOT, 'CONTEXT.md');
|
||||
const INDEX_PATH = path.join(ROOT, 'docs', 'CONTEXT-INDEX.json');
|
||||
|
||||
// ─── Typed reason enum (CONTRIBUTING.md "Prohibited: Raw Text Matching") ───────
|
||||
|
||||
/**
|
||||
* Stable reason codes for `checkReport`'s `reason` field. Tests assert via
|
||||
* `assert.equal(report.reason, REASON.X)` rather than regex-matching the
|
||||
* human-readable prose the non-JSON `--check` mode still writes to
|
||||
* stdout/stderr, so the diagnostic surface is a typed enum, not free text.
|
||||
*
|
||||
* Adding a new reason requires updating this map AND the tests' shape
|
||||
* assertion that locks the documented set of codes
|
||||
* (`Object.keys(REASON).sort()`).
|
||||
*/
|
||||
const REASON = Object.freeze({
|
||||
OK_UP_TO_DATE: 'ok_up_to_date',
|
||||
FAIL_STALE: 'fail_stale',
|
||||
FAIL_INDEX_MISSING: 'fail_index_missing',
|
||||
FAIL_INDEX_UNPARSEABLE: 'fail_index_unparseable',
|
||||
FAIL_DUPLICATE_IDS: 'fail_duplicate_ids',
|
||||
FAIL_CONTEXT_MISSING: 'fail_context_missing',
|
||||
FAIL_CONTEXT_UNREADABLE: 'fail_context_unreadable',
|
||||
FAIL_LIB_NOT_BUILT: 'fail_lib_not_built',
|
||||
});
|
||||
|
||||
// ─── Loaders ──────────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Load the compiled context-predicates library. The artifact is a gitignored
|
||||
* tsc build output of src/context-predicates.cts and only exists after
|
||||
* `npm run build:lib`. Throws a clean ExitError (never a bare
|
||||
* MODULE_NOT_FOUND stack) naming the remedy when it is missing.
|
||||
*
|
||||
* @returns {{ parsePredicates: Function, selectPredicates: Function, buildIndex: Function }}
|
||||
*/
|
||||
function loadContextPredicatesLib() {
|
||||
try {
|
||||
delete require.cache[require.resolve(CONTEXT_PREDICATES_LIB_PATH)];
|
||||
return require(CONTEXT_PREDICATES_LIB_PATH);
|
||||
} catch (err) {
|
||||
throw new ExitError(
|
||||
1,
|
||||
`Cannot load ${path.relative(ROOT, CONTEXT_PREDICATES_LIB_PATH)}: ${err && err.message}\n` +
|
||||
'Run:\n npm run build:lib\n',
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Read a CONTEXT.md-shaped markdown file. Throws a clean ExitError naming the
|
||||
* path (never a bare stack trace) when it is missing or unreadable.
|
||||
*
|
||||
* @param {string} [contextPath] - defaults to the real repo-root CONTEXT.md.
|
||||
* @returns {string}
|
||||
*/
|
||||
function readContextMarkdown(contextPath = CONTEXT_PATH) {
|
||||
try {
|
||||
return fs.readFileSync(contextPath, 'utf8');
|
||||
} catch (err) {
|
||||
throw new ExitError(1, `Cannot read ${path.relative(ROOT, contextPath)}: ${err && err.message}`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Build a fresh ContextIndex from the given CONTEXT.md-shaped content.
|
||||
*
|
||||
* @param {string} [contextPath] - defaults to the real repo-root CONTEXT.md.
|
||||
* @returns {object}
|
||||
*/
|
||||
function buildFreshIndex(contextPath = CONTEXT_PATH) {
|
||||
const { parsePredicates, buildIndex } = loadContextPredicatesLib();
|
||||
const markdown = readContextMarkdown(contextPath);
|
||||
const { predicates } = parsePredicates(markdown);
|
||||
return buildIndex(predicates);
|
||||
}
|
||||
|
||||
// ─── Serialization ────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Serialize a ContextIndex to the committed plain-JSON artifact text
|
||||
* (docs/CONTEXT-INDEX.json). Plain JSON — not a CommonJS module — because
|
||||
* this is a generated data manifest (mirroring docs/INVENTORY-MANIFEST.json),
|
||||
* not runtime code: it must never be `require()`-able from a shipped
|
||||
* gsd-core/bin/lib/*.cjs module, which is exactly the mistake that leaked
|
||||
* ~120 KB of CONTEXT.md prose (including `.claude/hooks/...` path literals
|
||||
* and hardcoded package-name strings) into runtime-code content scanning.
|
||||
*
|
||||
* @param {object} index
|
||||
* @returns {string}
|
||||
*/
|
||||
function serializeIndex(index) {
|
||||
return JSON.stringify(index, null, 2) + '\n';
|
||||
}
|
||||
|
||||
/**
|
||||
* Normalize line endings to LF for CRLF-agnostic full-content comparison.
|
||||
*
|
||||
* @param {string} content
|
||||
* @returns {string}
|
||||
*/
|
||||
function normalizeLineEndings(content) {
|
||||
return content.replace(/\r/g, '');
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract the sorted list of duplicate predicate ids from a built index.
|
||||
*
|
||||
* @param {{ duplicates: Array<{ id: string, count: number }> }} index
|
||||
* @returns {string[]}
|
||||
*/
|
||||
function duplicateIds(index) {
|
||||
return index.duplicates.map((d) => d.id);
|
||||
}
|
||||
|
||||
// ─── Typed check report ───────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Empty-report shape shared by every early-exit branch below, so callers
|
||||
* (JSON mode, tests) always see the same four data fields regardless of
|
||||
* which reason fired.
|
||||
*
|
||||
* @returns {{ duplicates: Array, count: number, classes: object }}
|
||||
*/
|
||||
function emptyReportFields() {
|
||||
return { duplicates: [], count: 0, classes: {} };
|
||||
}
|
||||
|
||||
/**
|
||||
* Compute the full `--check` result as a typed, non-throwing report — the
|
||||
* structured intermediate representation CONTRIBUTING.md's "Prohibited: Raw
|
||||
* Text Matching on Test Outputs" section requires alongside the human prose
|
||||
* `main()` still prints. Never throws; every failure mode is a `reason` code
|
||||
* from the frozen `REASON` enum.
|
||||
*
|
||||
* @param {string} [contextPath] - defaults to the real repo-root CONTEXT.md.
|
||||
* @param {string} [indexPath] - defaults to the real committed index.
|
||||
* @returns {{ ok: boolean, reason: string, duplicates: Array<{id:string,count:number}>, count: number, classes: object, message: string }}
|
||||
*/
|
||||
function checkReport(contextPath = CONTEXT_PATH, indexPath = INDEX_PATH) {
|
||||
let markdown;
|
||||
try {
|
||||
markdown = fs.readFileSync(contextPath, 'utf8');
|
||||
} catch (err) {
|
||||
const reason = err && err.code === 'ENOENT' ? REASON.FAIL_CONTEXT_MISSING : REASON.FAIL_CONTEXT_UNREADABLE;
|
||||
return {
|
||||
ok: false,
|
||||
reason,
|
||||
...emptyReportFields(),
|
||||
message: `Cannot read ${path.relative(ROOT, contextPath)}: ${err && err.message}`,
|
||||
};
|
||||
}
|
||||
|
||||
let parsePredicates;
|
||||
let buildIndex;
|
||||
try {
|
||||
delete require.cache[require.resolve(CONTEXT_PREDICATES_LIB_PATH)];
|
||||
({ parsePredicates, buildIndex } = require(CONTEXT_PREDICATES_LIB_PATH));
|
||||
} catch (err) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: REASON.FAIL_LIB_NOT_BUILT,
|
||||
...emptyReportFields(),
|
||||
message: `Cannot load ${path.relative(ROOT, CONTEXT_PREDICATES_LIB_PATH)}: ${err && err.message}\n` +
|
||||
'Run:\n npm run build:lib\n',
|
||||
};
|
||||
}
|
||||
|
||||
const { predicates } = parsePredicates(markdown);
|
||||
const live = buildIndex(predicates);
|
||||
const dups = duplicateIds(live);
|
||||
|
||||
if (dups.length > 0) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: REASON.FAIL_DUPLICATE_IDS,
|
||||
duplicates: live.duplicates,
|
||||
count: live.count,
|
||||
classes: live.classes,
|
||||
message: 'CONTEXT.md has duplicate predicate id(s): ' + dups.join(', ') + '\n' +
|
||||
'Each predicate id must be declared exactly once.\n',
|
||||
};
|
||||
}
|
||||
|
||||
if (!fs.existsSync(indexPath)) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: REASON.FAIL_INDEX_MISSING,
|
||||
duplicates: live.duplicates,
|
||||
count: live.count,
|
||||
classes: live.classes,
|
||||
message: `${path.relative(ROOT, indexPath)} does not exist. Run:\n node scripts/gen-context-index.cjs --write\n`,
|
||||
};
|
||||
}
|
||||
|
||||
let committedIndex;
|
||||
try {
|
||||
const committedText = fs.readFileSync(indexPath, 'utf8');
|
||||
committedIndex = JSON.parse(committedText);
|
||||
if (!committedIndex || typeof committedIndex !== 'object' || !Array.isArray(committedIndex.predicates)) {
|
||||
throw new Error('parsed JSON does not have the expected ContextIndex shape');
|
||||
}
|
||||
} catch (err) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: REASON.FAIL_INDEX_UNPARSEABLE,
|
||||
duplicates: live.duplicates,
|
||||
count: live.count,
|
||||
classes: live.classes,
|
||||
message: `${path.relative(ROOT, indexPath)} is unparseable: ${err && err.message}\n` +
|
||||
'Run:\n node scripts/gen-context-index.cjs --write\n',
|
||||
};
|
||||
}
|
||||
|
||||
// Comparing parsed JSON (rather than raw file text) is inherently
|
||||
// CRLF-agnostic: JSON.parse treats \r\n and \n as equivalent insignificant
|
||||
// whitespace between tokens, and predicate values never contain embedded
|
||||
// newlines (the parser only extracts single-physical-line declarations).
|
||||
if (JSON.stringify(committedIndex) !== JSON.stringify(live)) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: REASON.FAIL_STALE,
|
||||
duplicates: live.duplicates,
|
||||
count: live.count,
|
||||
classes: live.classes,
|
||||
message: `${path.relative(ROOT, indexPath)} is stale. Run:\n node scripts/gen-context-index.cjs --write\n`,
|
||||
};
|
||||
}
|
||||
|
||||
return {
|
||||
ok: true,
|
||||
reason: REASON.OK_UP_TO_DATE,
|
||||
duplicates: live.duplicates,
|
||||
count: live.count,
|
||||
classes: live.classes,
|
||||
message: `${path.relative(ROOT, indexPath)} is up to date.\n`,
|
||||
};
|
||||
}
|
||||
|
||||
// ─── Argument parsing ─────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* True when `value` cannot be accepted as a `--context-path`/`--index-path`
|
||||
* argument: absent (no more argv), empty string, or flag-shaped (starts with
|
||||
* `-`, so a dangling `--context-path` immediately followed by the NEXT flag
|
||||
* is rejected rather than silently swallowing that flag as a literal path).
|
||||
*
|
||||
* @param {string|undefined} value
|
||||
* @returns {boolean}
|
||||
*/
|
||||
function isMissingPathValue(value) {
|
||||
return value === undefined || value === '' || value.startsWith('-');
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse CLI arguments into a structured options object.
|
||||
*
|
||||
* `--context-path` / `--index-path` override the two hardcoded repo-root
|
||||
* paths — added so tests can point the real CLI at a temp fixture tree
|
||||
* directly, with no fs monkeypatching required.
|
||||
*
|
||||
* Two usage-error conditions (DEFECT.GEN-CONTEXT-INDEX-PARSEARGS-GATE-BYPASS,
|
||||
* MAJOR review finding) collapse into `mode: 'unknown'`, the same clean,
|
||||
* no-stack-trace usage-error path `main()` already uses for an unrecognized
|
||||
* flag:
|
||||
* (a) `--check` and `--write` given together — previously the LAST one
|
||||
* seen silently won, so `--check --write` exited 0 and REWROTE the
|
||||
* committed index instead of gating. Conflicting mode flags are now a
|
||||
* hard usage error regardless of order.
|
||||
* (b) a missing/empty/flag-shaped value for `--context-path` /
|
||||
* `--index-path` — previously `path.resolve(argv[++i] ?? '')`
|
||||
* resolved to the current working directory, and in `--write` mode
|
||||
* that later threw an uncaught, uncleaned `EISDIR` stack trace from
|
||||
* `fs.writeFileSync` (CONTRIBUTING.md: no stack trace in non-debug
|
||||
* failure output). Rejected up front instead, before any I/O.
|
||||
*
|
||||
* @param {string[]} argv - process.argv.slice(2)
|
||||
* @returns {{ mode: 'check'|'write'|'default'|'unknown', json: boolean, contextPath: string, indexPath: string, unknownArg?: string, usageMessage?: string }}
|
||||
*/
|
||||
function parseArgs(argv) {
|
||||
const opts = { mode: 'default', json: false, contextPath: CONTEXT_PATH, indexPath: INDEX_PATH };
|
||||
let sawCheck = false;
|
||||
let sawWrite = false;
|
||||
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const arg = argv[i];
|
||||
if (arg === '--check') {
|
||||
sawCheck = true;
|
||||
opts.mode = 'check';
|
||||
} else if (arg === '--write') {
|
||||
sawWrite = true;
|
||||
opts.mode = 'write';
|
||||
} else if (arg === '--json') {
|
||||
opts.json = true;
|
||||
} else if (arg === '--context-path' || arg === '--index-path') {
|
||||
const value = argv[i + 1];
|
||||
if (isMissingPathValue(value)) {
|
||||
return {
|
||||
...opts,
|
||||
mode: 'unknown',
|
||||
unknownArg: arg,
|
||||
usageMessage: `${arg} requires a non-empty path argument (got ${value === undefined ? 'nothing' : JSON.stringify(value)})`,
|
||||
};
|
||||
}
|
||||
i++;
|
||||
if (arg === '--context-path') opts.contextPath = path.resolve(value);
|
||||
else opts.indexPath = path.resolve(value);
|
||||
} else {
|
||||
opts.mode = 'unknown';
|
||||
opts.unknownArg = arg;
|
||||
}
|
||||
}
|
||||
|
||||
if (sawCheck && sawWrite) {
|
||||
return {
|
||||
...opts,
|
||||
mode: 'unknown',
|
||||
usageMessage: '--check and --write are mutually exclusive',
|
||||
};
|
||||
}
|
||||
|
||||
return opts;
|
||||
}
|
||||
|
||||
// ─── Main ─────────────────────────────────────────────────────────────────────
|
||||
|
||||
function main() {
|
||||
const opts = parseArgs(process.argv.slice(2));
|
||||
|
||||
if (opts.mode === 'unknown') {
|
||||
process.stderr.write('Usage: gen-context-index.cjs [--write|--check] [--json] [--context-path <path>] [--index-path <path>]\n');
|
||||
if (opts.usageMessage) process.stderr.write(`${opts.usageMessage}\n`);
|
||||
throw new ExitError(1);
|
||||
}
|
||||
|
||||
if (opts.mode === 'default') {
|
||||
process.stdout.write(serializeIndex(buildFreshIndex(opts.contextPath)) + '\n');
|
||||
return;
|
||||
}
|
||||
|
||||
if (opts.mode === 'check') {
|
||||
const report = checkReport(opts.contextPath, opts.indexPath);
|
||||
|
||||
if (opts.json) {
|
||||
process.stdout.write(JSON.stringify({
|
||||
ok: report.ok,
|
||||
reason: report.reason,
|
||||
duplicates: report.duplicates,
|
||||
count: report.count,
|
||||
classes: report.classes,
|
||||
}) + '\n');
|
||||
} else if (report.ok) {
|
||||
process.stdout.write(report.message);
|
||||
}
|
||||
|
||||
if (!report.ok) {
|
||||
// JSON mode already carries the structured report on stdout — do not
|
||||
// duplicate the prose onto stderr, but the exit code must still be 1.
|
||||
throw new ExitError(1, opts.json ? undefined : report.message);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
// opts.mode === 'write'
|
||||
const index = buildFreshIndex(opts.contextPath);
|
||||
fs.mkdirSync(path.dirname(opts.indexPath), { recursive: true });
|
||||
fs.writeFileSync(opts.indexPath, serializeIndex(index), 'utf8');
|
||||
const dupCount = index.duplicates.length;
|
||||
process.stdout.write(
|
||||
`Wrote ${path.relative(ROOT, opts.indexPath)}\n` +
|
||||
` ${index.count} predicates, ${Object.keys(index.classes).length} classes, ` +
|
||||
`${dupCount} duplicate id${dupCount !== 1 ? 's' : ''}\n`,
|
||||
);
|
||||
}
|
||||
|
||||
// ─── Exports (for tests) ──────────────────────────────────────────────────────
|
||||
|
||||
module.exports = {
|
||||
loadContextPredicatesLib,
|
||||
readContextMarkdown,
|
||||
buildFreshIndex,
|
||||
serializeIndex,
|
||||
normalizeLineEndings,
|
||||
duplicateIds,
|
||||
checkReport,
|
||||
parseArgs,
|
||||
REASON,
|
||||
CONTEXT_PREDICATES_LIB_PATH,
|
||||
CONTEXT_PATH,
|
||||
INDEX_PATH,
|
||||
};
|
||||
|
||||
// ─── CLI entry point ──────────────────────────────────────────────────────────
|
||||
|
||||
if (require.main === module) {
|
||||
runMain(main);
|
||||
}
|
||||
542
src/context-predicates.cts
Normal file
542
src/context-predicates.cts
Normal file
@@ -0,0 +1,542 @@
|
||||
/**
|
||||
* Context Predicates — CONTEXT.md predicate fact-store parser.
|
||||
*
|
||||
* Ported from the reference prototype
|
||||
* `examples/dynamic-context-management/context-predicates.cjs` (ADR-1671,
|
||||
* "Dynamic context management platform", Option-E predicate fact-store).
|
||||
* Behavior is preserved BYTE-FOR-BEHAVIOR from the prototype for the parts it
|
||||
* shares: `ID_RE`, the first-`=` split, the naive triple-backtick fence
|
||||
* toggle, and the `- ` list-item-only form. Known prototype defects are
|
||||
* DELIBERATELY carried forward here — a later commit fixes them behind a
|
||||
* failing-first test.
|
||||
*
|
||||
* Two intentional deviations from the prototype (ADR-1671 open question 4):
|
||||
* - `Duplicate` carries `count`, not `lines: number[]`.
|
||||
* - `ContextIndex.predicates` entries carry `{id, klass, value}` with NO
|
||||
* `line` field — a committed artifact with no line numbers cannot drift
|
||||
* on a line shift. The live parse result (`Predicate`) still carries
|
||||
* `line` and `section`.
|
||||
*
|
||||
* One addition beyond the prototype: `ParseResult.malformed` collects
|
||||
* backtick lines that look like a predicate declaration but are rejected for
|
||||
* having an empty value (e.g. `` `ID=` ``), so the empty-value case is
|
||||
* surfaced as a diagnostic instead of being silently dropped. This does not
|
||||
* change any accept/reject outcome — only adds a diagnostic.
|
||||
*
|
||||
* Grammar (from discovery facts):
|
||||
* Two line forms, each on exactly one source line:
|
||||
* 1. Bare backtick-wrapped, optionally indented: `ID=value`, ` `ID=value``
|
||||
* 2. List-item backtick: `-`/`*`/`+`/`N.` marker followed by `ID=value`
|
||||
*
|
||||
* ID grammar: CLASS(.subkey)* where CLASS = first dot-separated segment.
|
||||
* ID chars: [A-Za-z0-9._-] (CLASS always uppercase; subkeys may be mixed).
|
||||
* Split on FIRST '=' only; everything before is the ID, everything after is
|
||||
* the value (up to the closing backtick).
|
||||
*
|
||||
* Skip:
|
||||
* - Fenced code blocks: ``` or ~~~, fence-length- and fence-char-aware
|
||||
* (a longer fence containing a shorter same-char fence line stays a
|
||||
* single skipped region; mismatched-char lines are fence content, not
|
||||
* a toggle)
|
||||
* - HTML comments (`<!-- ... -->`), including multi-line
|
||||
* - Prose lines (headings, blank lines, list items without a predicate)
|
||||
* - Blockquote lines (session-log preamble, etc.)
|
||||
*
|
||||
* Fence-length-awareness note: `src/markdown-sectionizer.cts`'s
|
||||
* `stripFencedCode` is the repo's canonical CommonMark fence-stripper, but its
|
||||
* `StripFencedResult.text` DROPS fence delimiter and content lines from the
|
||||
* output — it does not preserve original line numbers. This parser reports
|
||||
* `Predicate.line`/`Malformed.line` as 1-based SOURCE line numbers, which
|
||||
* callers assert on — so line-accurate skip detection is required, and
|
||||
* `stripFencedCode` cannot serve it directly.
|
||||
*
|
||||
* Comment/fence precedence (DEFECT.CONTEXT-PREDICATES-COMMENT-FENCE-BLIND,
|
||||
* #2928 review): `computeSkippedLineFlags` previously ran the HTML-comment
|
||||
* scan and a delegated call to `markdown-sectionizer.cts`'s exported
|
||||
* `scanFencedBlocks` seam as two INDEPENDENT passes over the raw lines. That
|
||||
* is unsound: `scanFencedBlocks` is comment-blind, so a fence delimiter
|
||||
* appearing INSIDE an HTML comment (with no later matching close in the
|
||||
* file) was treated as a real *unterminated* fence — silently skipping every
|
||||
* remaining line to EOF and permanently dropping later predicates. The
|
||||
* converse is also unsound the other way: a bare two-pass ordering that
|
||||
* resolves comments first and only then masks-and-rescans for fences
|
||||
* mis-handles a `<!--`/`-->` token that appears *inside a genuine fenced
|
||||
* block* (proven while fixing this: guarding the comment pass by a
|
||||
* comment-blind fence pass, or vice versa, always breaks one of the two
|
||||
* directions — the two constructs must suppress each other's
|
||||
* open/close detection while active, which only a single interleaved
|
||||
* left-to-right pass can guarantee).
|
||||
*
|
||||
* Chosen precedence (documented per the review's requirement): the two
|
||||
* constructs are scanned in ONE forward pass with two mutually-exclusive
|
||||
* states, `fence: {char,len} | null` and `inHtmlComment: boolean`.
|
||||
* - While a fence is open, only a fence-close delimiter (same char, run
|
||||
* length >= the opener's) can close it; any `<!--`/`-->` token on a
|
||||
* fenced-content line is fence content, never a comment boundary.
|
||||
* - While NEITHER is open and an HTML comment opens, only a `-->` token
|
||||
* can close it; any fence delimiter seen while inside a comment is
|
||||
* comment content, never a fence boundary.
|
||||
* - When neither is open, a comment opener (`<!--`) is checked BEFORE a
|
||||
* fence opener on the same line — HTML comments are lexically outermost
|
||||
* in this document's grammar — so a fence-shaped info string that
|
||||
* happens to also look like `<!--...-->` never spuriously opens/closes
|
||||
* anything once the line has already been claimed as a fence opener (the
|
||||
* inverse: a genuine one-line comment `<!-- ``` -->` is masked whole and
|
||||
* never interpreted as a fence delimiter).
|
||||
* This single-pass design reuses `markdown-sectionizer.cts`'s exact
|
||||
* delimiter-matching rule (same regex, same `len >= open.len` /
|
||||
* mismatched-char-is-content / backtick-info-string-cannot-contain-backtick
|
||||
* semantics as `scanFencedBlocks`), so non-comment-interacting documents are
|
||||
* byte-for-byte identical to delegating to `scanFencedBlocks` (see
|
||||
* `tests/context-predicates.test.cjs`'s fence-skip parity suite) — the
|
||||
* interleaving is required ONLY to resolve the comment/fence interaction,
|
||||
* not to change fence semantics themselves. This is therefore a second,
|
||||
* necessarily local copy of the (tiny) delimiter-match condition — the
|
||||
* scanFencedBlocks seam cannot serve both scans at once, because the correct
|
||||
* boundary decision for either construct depends on the OTHER construct's
|
||||
* live state at that exact line, not just on a static, comment-blind
|
||||
* pre-scan of the raw lines.
|
||||
*
|
||||
* ADR-457 build-at-publish: compiled by tsc to
|
||||
* gsd-core/bin/lib/context-predicates.cjs (gitignored).
|
||||
*/
|
||||
|
||||
/** A single parsed predicate fact from CONTEXT.md. */
|
||||
export interface Predicate {
|
||||
id: string;
|
||||
klass: string;
|
||||
value: string;
|
||||
line: number;
|
||||
section: string;
|
||||
}
|
||||
|
||||
/** A predicate id that occurs more than once. */
|
||||
export interface Duplicate {
|
||||
id: string;
|
||||
count: number;
|
||||
}
|
||||
|
||||
/** A backtick-wrapped line that looked like a predicate declaration but was rejected. */
|
||||
export interface Malformed {
|
||||
line: number;
|
||||
text: string;
|
||||
reason: string;
|
||||
}
|
||||
|
||||
/** Result of parsing a CONTEXT.md markdown string for predicates. */
|
||||
export interface ParseResult {
|
||||
predicates: Predicate[];
|
||||
duplicates: Duplicate[];
|
||||
malformed: Malformed[];
|
||||
skippedSections: string[];
|
||||
}
|
||||
|
||||
/** Criteria for {@link selectPredicates}, ANDed together. */
|
||||
export interface SelectOptions {
|
||||
klass?: string;
|
||||
prefix?: string;
|
||||
contains?: string;
|
||||
}
|
||||
|
||||
/** A deterministic, committed index entry — no `line` (see module doc). */
|
||||
export interface ContextIndexPredicate {
|
||||
id: string;
|
||||
klass: string;
|
||||
value: string;
|
||||
}
|
||||
|
||||
/** Deterministic index built from a parsed predicates array. */
|
||||
export interface ContextIndex {
|
||||
schemaVersion: 1;
|
||||
count: number;
|
||||
classes: Record<string, number>;
|
||||
predicates: ContextIndexPredicate[];
|
||||
duplicates: Duplicate[];
|
||||
}
|
||||
|
||||
// ID grammar, validated STRUCTURALLY rather than by a single regex
|
||||
// (DEFECT.CONTEXT-PREDICATES-ID-REDOS, #2928 review). The formerly-used regex
|
||||
// `^([A-Z][A-Z0-9_-]*(?:\.[A-Za-z0-9_.-]+)*)=(.+)$` is exponential: the
|
||||
// group `(?:\.[A-Za-z0-9_.-]+)*` is ambiguous because its own character
|
||||
// class contains `.`, so N consecutive dots have exponentially many
|
||||
// backtick-partitionings for the regex engine to try on a failed match
|
||||
// (measured: ~565ms for 40 consecutive dots, doubling roughly every 5).
|
||||
// `.github/workflows/test.yml` runs `lint:ci` -> `lint:generated-sync` ->
|
||||
// `gen-context-index.cjs --check` on `pull_request`, which parses the PR's
|
||||
// own CONTEXT.md — so any external contributor could hang the shared CI
|
||||
// runner with one crafted line, no write access required.
|
||||
//
|
||||
// Fix: split the candidate id on '.' and validate each segment with a
|
||||
// simple, non-backtracking, per-segment pattern — linear in id length, no
|
||||
// ambiguous quantifier. First segment (CLASS) must start with an uppercase
|
||||
// letter; subsequent segments may start with letter/digit and include
|
||||
// hyphens/underscores. We intentionally allow lowercase-starting
|
||||
// sub-segments (e.g. PRED.k320.rule).
|
||||
//
|
||||
// Behavior change vs. the old regex: an EMPTY segment (a doubled dot, e.g.
|
||||
// `A..b`) now REJECTS — the old regex accepted it because `.` was inside the
|
||||
// subsequent-segment character class, so `.` itself could satisfy
|
||||
// `[A-Za-z0-9_.-]+` with a single character. The real repo CONTEXT.md was
|
||||
// checked (`grep -nE '\`[A-Z][A-Za-z0-9_.-]*\.\.[A-Za-z0-9_.-]*='
|
||||
// CONTEXT.md`) and contains ZERO ids with a doubled dot, so rejecting the
|
||||
// empty-segment case is the correct, stricter grammar with no behavior loss
|
||||
// against real data (pinned by a dedicated test below).
|
||||
const ID_FIRST_SEGMENT_RE = /^[A-Z][A-Z0-9_-]*$/;
|
||||
const ID_SUBSEQUENT_SEGMENT_RE = /^[A-Za-z0-9_-]+$/;
|
||||
|
||||
/**
|
||||
* Structurally validate a candidate predicate id (linear time — no ambiguous
|
||||
* backtracking quantifier; see the ID grammar comment above).
|
||||
*
|
||||
* @param id - candidate id (everything before the first '=')
|
||||
*/
|
||||
function isValidId(id: string): boolean {
|
||||
const segments = id.split('.');
|
||||
if (!ID_FIRST_SEGMENT_RE.test(segments[0])) return false;
|
||||
for (let i = 1; i < segments.length; i++) {
|
||||
if (!ID_SUBSEQUENT_SEGMENT_RE.test(segments[i])) return false;
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
// List markers recognized ahead of a backtick-wrapped declaration:
|
||||
// `-`, `*`, `+`, or a numbered marker (`1.`, `42.`), each followed by
|
||||
// whitespace. Mirrors the marker family `iterateBullets`/`updateBullet`
|
||||
// (markdown-sectionizer.cts) recognize, widened here beyond the
|
||||
// prototype-carried-forward dash-only form (ADR-1671 Phase 1 commit 3).
|
||||
const LIST_MARKER_RE = /^[ \t]*(?:[-*+]|\d+\.)[ \t]+/;
|
||||
|
||||
/**
|
||||
* Strip a source line down to its backtick-wrapped "inner" content, if any.
|
||||
* Handles both line forms:
|
||||
* 1. Bare backtick line, optionally indented: `ID=value`, ` `ID=value``
|
||||
* 2. List-item backtick, any of `-`/`*`/`+`/`N.`, optionally indented:
|
||||
* `- `ID=value``, `* `ID=value``, `+ `ID=value``, `1. `ID=value``
|
||||
*
|
||||
* @param raw - the original source line (with newline stripped)
|
||||
* @returns the inner content between the backticks, or null if the line is
|
||||
* not backtick-wrapped in either recognized form
|
||||
*/
|
||||
function extractInner(raw: string): string | null {
|
||||
const line = raw.trimEnd();
|
||||
|
||||
// Bare backtick-wrapped, tolerating leading indentation — shape decides
|
||||
// the bare form, not column 0 (Postel: CONTEXT.md authors indent freely).
|
||||
const bareTrimmed = line.replace(/^[ \t]+/, '');
|
||||
if (bareTrimmed.startsWith('`') && bareTrimmed.endsWith('`') && bareTrimmed.length > 2) {
|
||||
return bareTrimmed.slice(1, -1);
|
||||
}
|
||||
|
||||
// List-item form: strip optional leading whitespace + list marker, then
|
||||
// check for backtick wrapping. `stripped !== line` guards against a line
|
||||
// with no marker at all (LIST_MARKER_RE.replace would otherwise no-op and
|
||||
// re-check the same failed bare-form test).
|
||||
const stripped = line.replace(LIST_MARKER_RE, '');
|
||||
if (stripped !== line && stripped.startsWith('`') && stripped.endsWith('`') && stripped.length > 2) {
|
||||
return stripped.slice(1, -1);
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a single source line and return a raw {id, value} if it is a
|
||||
* predicate, or null otherwise.
|
||||
*
|
||||
* @param raw - the original source line (with newline stripped)
|
||||
*/
|
||||
function extractPredicate(raw: string): { id: string; value: string } | null {
|
||||
const inner = extractInner(raw);
|
||||
if (inner === null) return null;
|
||||
|
||||
// Split on FIRST '=' only.
|
||||
const eqIdx = inner.indexOf('=');
|
||||
if (eqIdx < 1) return null;
|
||||
|
||||
const id = inner.slice(0, eqIdx);
|
||||
const value = inner.slice(eqIdx + 1);
|
||||
|
||||
// Value must be non-empty (the old ID_RE's trailing `(.+)$` requirement)
|
||||
// and must contain no embedded ECMAScript LineTerminator character (LF,
|
||||
// CR, U+2028 LINE SEPARATOR, U+2029 PARAGRAPH SEPARATOR) -- the old
|
||||
// regex's `.` metachar excludes exactly those four characters and carried
|
||||
// no `s`/`m` flag, so a value spanning an embedded \r (possible only via
|
||||
// the documented lone-CR limit: a "line" with no real \n at all still
|
||||
// carries a mid-string \r joining what the author intended as two
|
||||
// separate lines) could never satisfy `(.+)$`. Preserved byte-for-behavior
|
||||
// here so `yieldsNoPredicatesForLoneCrDocumentAsDocumentedLimit` stays
|
||||
// pinned. And the id must match the structural grammar (no spaces,
|
||||
// correct char set, no empty segment -- see isValidId's doc comment).
|
||||
if (value === '' || /[\n\r\u2028\u2029]/.test(value) || !isValidId(id)) return null;
|
||||
|
||||
return { id, value };
|
||||
}
|
||||
|
||||
/**
|
||||
* Detect the "looks like a declaration but has an empty value" malformed
|
||||
* case for a line that {@link extractPredicate} already rejected. Only
|
||||
* fires when the ID portion is grammatically valid on its own and the value
|
||||
* after the first '=' is empty (e.g. `` `ID=` ``). Does not change any
|
||||
* accept/reject decision — diagnostic only.
|
||||
*
|
||||
* @param raw - the original source line (with newline stripped)
|
||||
*/
|
||||
function detectMalformed(raw: string): { text: string; reason: string } | null {
|
||||
const inner = extractInner(raw);
|
||||
if (inner === null) return null;
|
||||
|
||||
const eqIdx = inner.indexOf('=');
|
||||
if (eqIdx < 1) return null;
|
||||
|
||||
const id = inner.slice(0, eqIdx);
|
||||
const value = inner.slice(eqIdx + 1);
|
||||
|
||||
if (value === '' && isValidId(id)) {
|
||||
return { text: raw.trimEnd(), reason: 'empty-value' };
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
// Fence delimiter line matcher — mirrors `markdown-sectionizer.cts`'s
|
||||
// `scanFencedBlocks` regex exactly (≥3 backticks/tildes, ≤3-space indent
|
||||
// tolerance). Kept local so the single interleaved pass below can decide,
|
||||
// line by line, whether a delimiter is a REAL fence boundary given the
|
||||
// comment state AT THAT LINE — see the module doc comment's "Comment/fence
|
||||
// precedence" section for why this can't be a call-then-mask over
|
||||
// `scanFencedBlocks`'s output.
|
||||
const FENCE_DELIM_RE = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||||
|
||||
/**
|
||||
* Compute, per source line, whether that line falls inside a fenced code
|
||||
* block or an HTML comment (`<!-- ... -->`, single- or multi-line).
|
||||
* LINE-PRESERVING: returns one boolean per input line (no lines dropped or
|
||||
* collapsed) — see the module doc comment for why that distinction is
|
||||
* load-bearing here.
|
||||
*
|
||||
* Single interleaved forward pass over two mutually-exclusive states —
|
||||
* `fence` (open fence delimiter char + run length, or null) and
|
||||
* `inHtmlComment` — so each construct suppresses the OTHER's open/close
|
||||
* detection while it is active (module doc comment's "Comment/fence
|
||||
* precedence"). This is the fix for DEFECT.CONTEXT-PREDICATES-COMMENT-FENCE-
|
||||
* BLIND: a fence delimiter inside a real HTML comment is comment content
|
||||
* (never opens a fence), and a `<!--`/`-->` token inside a real fenced block
|
||||
* is fence content (never opens/closes a comment).
|
||||
*
|
||||
* @param lines - source lines (as produced by `markdown.split('\n')`)
|
||||
*/
|
||||
function computeSkippedLineFlags(lines: string[]): boolean[] {
|
||||
const skip = new Array<boolean>(lines.length).fill(false);
|
||||
|
||||
let fence: { char: '`' | '~'; len: number } | null = null;
|
||||
let inHtmlComment = false;
|
||||
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
// Strip trailing \r (CRLF safety), mirroring stripFencedCode's/
|
||||
// scanFencedBlocks's own `rawLine.replace(/\r$/, '')`.
|
||||
const line = lines[i].replace(/\r$/, '');
|
||||
|
||||
if (fence !== null) {
|
||||
// Inside a real fence: only a matching closer can end it. Any
|
||||
// `<!--`/`-->` on this line is fence content, not a comment boundary
|
||||
// (converse precedence).
|
||||
skip[i] = true;
|
||||
const m = FENCE_DELIM_RE.exec(line);
|
||||
if (m) {
|
||||
const char = m[2][0] as '`' | '~';
|
||||
const len = m[2].length;
|
||||
const trailing = m[3];
|
||||
if (char === fence.char && len >= fence.len && /^\s*$/.test(trailing)) {
|
||||
fence = null;
|
||||
}
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
if (inHtmlComment) {
|
||||
// Inside a real comment: only '-->' can end it. Any fence delimiter on
|
||||
// this line is comment content, not a fence boundary (primary
|
||||
// precedence — the DEFECT.CONTEXT-PREDICATES-COMMENT-FENCE-BLIND
|
||||
// repro: a fence delimiter with no later real closer must not skip to
|
||||
// EOF just because it happened to appear inside a comment).
|
||||
skip[i] = true;
|
||||
if (line.includes('-->')) inHtmlComment = false;
|
||||
continue;
|
||||
}
|
||||
|
||||
// Neither construct open: HTML comments are lexically outermost in this
|
||||
// document's grammar, so a comment opener is checked BEFORE a fence
|
||||
// opener on the same line.
|
||||
const trimmed = line.trim();
|
||||
if (trimmed.startsWith('<!--')) {
|
||||
skip[i] = true;
|
||||
if (!trimmed.includes('-->')) {
|
||||
inHtmlComment = true; // multi-line: stays open until a later '-->'
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
const m = FENCE_DELIM_RE.exec(line);
|
||||
if (m) {
|
||||
const char = m[2][0] as '`' | '~';
|
||||
const trailing = m[3];
|
||||
// CommonMark §4.5: a backtick fence opener's info string must not
|
||||
// itself contain a backtick — such a line is ordinary content, not a
|
||||
// valid opener (mirrors scanFencedBlocks).
|
||||
if (!(char === '`' && trailing.includes('`'))) {
|
||||
skip[i] = true;
|
||||
fence = { char, len: m[2].length };
|
||||
continue;
|
||||
}
|
||||
}
|
||||
|
||||
skip[i] = false;
|
||||
}
|
||||
|
||||
return skip;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse all predicates from a CONTEXT.md markdown string.
|
||||
*
|
||||
* @param markdown
|
||||
*/
|
||||
export function parsePredicates(markdown: string): ParseResult {
|
||||
const lines = markdown.split('\n');
|
||||
const predicates: Predicate[] = [];
|
||||
const malformed: Malformed[] = [];
|
||||
// Track id -> occurrence count for duplicate detection
|
||||
const idCounts = new Map<string, number>();
|
||||
|
||||
const skippedLines = computeSkippedLineFlags(lines);
|
||||
let currentSection = '';
|
||||
const allSections: string[] = [];
|
||||
const seenSections = new Set<string>();
|
||||
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const raw = lines[i];
|
||||
const lineNo = i + 1; // 1-based
|
||||
|
||||
// Fenced code blocks and HTML comments (line-preserving; see
|
||||
// computeSkippedLineFlags's doc comment).
|
||||
if (skippedLines[i]) continue;
|
||||
|
||||
// Track section headings for the section field.
|
||||
if (raw.startsWith('#')) {
|
||||
currentSection = raw.replace(/^#+\s*/, '').trim();
|
||||
if (currentSection && !seenSections.has(currentSection)) {
|
||||
seenSections.add(currentSection);
|
||||
allSections.push(currentSection);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
// Blockquote lines (start with ">") are prose — skip.
|
||||
if (raw.trimStart().startsWith('>')) continue;
|
||||
|
||||
// Attempt extraction.
|
||||
const pred = extractPredicate(raw);
|
||||
if (!pred) {
|
||||
const bad = detectMalformed(raw);
|
||||
if (bad) {
|
||||
malformed.push({ line: lineNo, text: bad.text, reason: bad.reason });
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
const klass = pred.id.split('.')[0];
|
||||
predicates.push({
|
||||
id: pred.id,
|
||||
klass,
|
||||
value: pred.value,
|
||||
line: lineNo,
|
||||
section: currentSection,
|
||||
});
|
||||
|
||||
idCounts.set(pred.id, (idCounts.get(pred.id) || 0) + 1);
|
||||
}
|
||||
|
||||
// Build duplicates list: ids with >1 occurrence.
|
||||
const duplicates: Duplicate[] = [];
|
||||
for (const [id, count] of idCounts) {
|
||||
if (count > 1) duplicates.push({ id, count });
|
||||
}
|
||||
// Sort duplicates by id for determinism.
|
||||
duplicates.sort((a, b) => (a.id < b.id ? -1 : a.id > b.id ? 1 : 0));
|
||||
|
||||
// Skipped sections: headings that yielded zero predicates (pure prose).
|
||||
const activeSections = new Set(predicates.map((p) => p.section));
|
||||
const skippedSections = allSections.filter((s) => !activeSections.has(s));
|
||||
|
||||
return { predicates, duplicates, malformed, skippedSections };
|
||||
}
|
||||
|
||||
/**
|
||||
* Select predicates by one or more optional criteria (ANDed together).
|
||||
*
|
||||
* @param predicates
|
||||
* @param opts
|
||||
*/
|
||||
export function selectPredicates(predicates: Predicate[], opts: SelectOptions = {}): Predicate[] {
|
||||
const { klass, prefix, contains } = opts;
|
||||
const containsLower = contains ? contains.toLowerCase() : null;
|
||||
|
||||
return predicates.filter((p) => {
|
||||
if (klass !== undefined && p.klass !== klass) return false;
|
||||
if (prefix !== undefined && !p.id.startsWith(prefix)) return false;
|
||||
if (containsLower !== null) {
|
||||
const haystack = (p.id + ' ' + p.value).toLowerCase();
|
||||
if (!haystack.includes(containsLower)) return false;
|
||||
}
|
||||
return true;
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Build a deterministic index object from a parsed predicates array.
|
||||
*
|
||||
* @param predicates
|
||||
*/
|
||||
export function buildIndex(predicates: Predicate[]): ContextIndex {
|
||||
// Count per class.
|
||||
const classCounts: Record<string, number> = {};
|
||||
for (const p of predicates) {
|
||||
classCounts[p.klass] = (classCounts[p.klass] || 0) + 1;
|
||||
}
|
||||
|
||||
// Sort classes object by key for determinism.
|
||||
const classes: Record<string, number> = {};
|
||||
for (const k of Object.keys(classCounts).sort()) {
|
||||
classes[k] = classCounts[k];
|
||||
}
|
||||
|
||||
// Sort predicates by id then by line number (line used for ordering only —
|
||||
// the committed index entry itself omits `line`; see module doc).
|
||||
const sortedPredicates = predicates
|
||||
.slice()
|
||||
.sort((a, b) => {
|
||||
if (a.id < b.id) return -1;
|
||||
if (a.id > b.id) return 1;
|
||||
return a.line - b.line;
|
||||
})
|
||||
.map(({ id, klass, value }) => ({ id, klass, value }));
|
||||
|
||||
// Rebuild duplicates from the (sorted-by-id) predicates for determinism.
|
||||
const idCounts = new Map<string, number>();
|
||||
for (const p of predicates) {
|
||||
idCounts.set(p.id, (idCounts.get(p.id) || 0) + 1);
|
||||
}
|
||||
const duplicates: Duplicate[] = [];
|
||||
for (const [id, count] of idCounts) {
|
||||
if (count > 1) duplicates.push({ id, count });
|
||||
}
|
||||
duplicates.sort((a, b) => (a.id < b.id ? -1 : a.id > b.id ? 1 : 0));
|
||||
|
||||
return {
|
||||
schemaVersion: 1,
|
||||
count: predicates.length,
|
||||
classes,
|
||||
predicates: sortedPredicates,
|
||||
duplicates,
|
||||
};
|
||||
}
|
||||
@@ -269,7 +269,7 @@ function stripInlineCodeLine(line: string): string {
|
||||
// ─── extractFencedBlock ───────────────────────────────────────────────────────
|
||||
|
||||
/** A fenced code block located by `scanFencedBlocks`: line-index span + info string. */
|
||||
interface FencedBlockRecord {
|
||||
export interface FencedBlockRecord {
|
||||
/** Fence delimiter character (`` ` `` or `~`). */
|
||||
char: '`' | '~';
|
||||
/** Fence delimiter run length (≥3). */
|
||||
@@ -301,8 +301,13 @@ interface FencedBlockRecord {
|
||||
* Tracked duplication (same status as `tokenizeHeadings`'s copy, see its
|
||||
* comment above): this is a second independent copy of the fence state
|
||||
* machine, pending a T-tier consolidation.
|
||||
*
|
||||
* Exported so `context-predicates.cts` can consume this seam directly for its
|
||||
* line-preserving fenced-line skip detection, instead of carrying a third
|
||||
* independent copy of the fence state machine (see that module's doc
|
||||
* comment).
|
||||
*/
|
||||
function scanFencedBlocks(lines: string[]): FencedBlockRecord[] {
|
||||
export function scanFencedBlocks(lines: string[]): FencedBlockRecord[] {
|
||||
const delimRe = /^( {0,3})(`{3,}|~{3,})(.*)$/;
|
||||
const blocks: FencedBlockRecord[] = [];
|
||||
let open: { char: '`' | '~'; len: number; infoString: string; openLineIdx: number } | null = null;
|
||||
|
||||
@@ -3763,3 +3763,78 @@ describe('#2279: map-codebase date stamp instructions overwrite existing dates',
|
||||
'workflow must instruct agents to overwrite existing dates, not just replace [YYYY-MM-DD] placeholders');
|
||||
});
|
||||
});
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// DEFECT.GENERATIVE-FIX parity guard: HOST_COMMAND_ROUTERS vs TOP_LEVEL_USAGE
|
||||
// vs SKIP_ROOT_RESOLUTION (#2928 S9)
|
||||
//
|
||||
// gsd-tools.cjs's query-command surface is declared across THREE
|
||||
// independently hand-maintained sites in the same file with no prior parity
|
||||
// gate between them: the dispatch table (HOST_COMMAND_ROUTERS), the
|
||||
// `--help` command list (TOP_LEVEL_USAGE), and the project-root-skip list
|
||||
// (SKIP_ROOT_RESOLUTION). Nothing previously caught a command being wired
|
||||
// into the dispatch table but omitted from the help string (or vice versa)
|
||||
// — exactly the generative-fix-divergence shape CLAUDE.md's
|
||||
// "Generative Fix Divergence" anti-pattern names ("add a parity assertion
|
||||
// test that fails if the shared constants/arrays/parsers diverge").
|
||||
//
|
||||
// This is a STRUCTURAL comparison against the exported constants, not a
|
||||
// source-text/string-match test, so it stays correct across reformatting
|
||||
// and is immune to the no-source-grep concern.
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
describe('gsd-tools.cjs dispatch/help/skip-list parity (DEFECT.GENERATIVE-FIX, #2928 S9)', () => {
|
||||
const { HOST_COMMAND_ROUTERS, TOP_LEVEL_USAGE, skipsRootResolution } = require('../gsd-core/bin/gsd-tools.cjs');
|
||||
|
||||
// Parse the "Commands: a, b, c\n\nGlobal flags:" line out of the usage
|
||||
// string rather than hardcoding a copy of it here — this test must fail
|
||||
// when the two sites diverge, not silently pass because it re-embeds its
|
||||
// own stale expectation.
|
||||
function parseHelpCommandNames(usage) {
|
||||
const match = usage.match(/Commands: ([\s\S]*?)\n\nGlobal flags:/);
|
||||
assert.ok(match, 'TOP_LEVEL_USAGE must contain a "Commands: ...\\n\\nGlobal flags:" block');
|
||||
return match[1]
|
||||
.split(',')
|
||||
.map((s) => s.trim())
|
||||
.filter(Boolean);
|
||||
}
|
||||
|
||||
test('every HOST_COMMAND_ROUTERS entry is listed in the --help command string', () => {
|
||||
const helpNames = new Set(parseHelpCommandNames(TOP_LEVEL_USAGE));
|
||||
const missing = Object.keys(HOST_COMMAND_ROUTERS).filter((name) => !helpNames.has(name));
|
||||
assert.deepEqual(
|
||||
missing,
|
||||
[],
|
||||
`command(s) registered in HOST_COMMAND_ROUTERS but missing from TOP_LEVEL_USAGE's ` +
|
||||
`"Commands:" list: ${missing.join(', ')}`,
|
||||
);
|
||||
});
|
||||
|
||||
test('context-predicates is registered in all three hand-maintained sites', () => {
|
||||
// Concrete regression pin for the command this parity test was added
|
||||
// alongside (#2928 S9) — a generic diff-based assertion alone would not
|
||||
// fail if ALL THREE sites were missing an entry simultaneously.
|
||||
assert.ok(
|
||||
Object.prototype.hasOwnProperty.call(HOST_COMMAND_ROUTERS, 'context-predicates'),
|
||||
'context-predicates must be registered in HOST_COMMAND_ROUTERS',
|
||||
);
|
||||
assert.ok(
|
||||
parseHelpCommandNames(TOP_LEVEL_USAGE).includes('context-predicates'),
|
||||
'context-predicates must be listed in TOP_LEVEL_USAGE',
|
||||
);
|
||||
assert.ok(
|
||||
skipsRootResolution('context-predicates'),
|
||||
'context-predicates must be in SKIP_ROOT_RESOLUTION (it is a pure repo-root CONTEXT.md ' +
|
||||
'read, like prompt-budget, and must work with no .planning/ directory present)',
|
||||
);
|
||||
});
|
||||
|
||||
test('SKIP_ROOT_RESOLUTION is not exported as a mutable live Set (DEFECT.MUTABLE-EXPORTED-SET, #2928)', () => {
|
||||
const gsdTools = require('../gsd-core/bin/gsd-tools.cjs');
|
||||
assert.equal(
|
||||
gsdTools.SKIP_ROOT_RESOLUTION,
|
||||
undefined,
|
||||
'the live Set must not be exported directly — only the read-only skipsRootResolution() predicate',
|
||||
);
|
||||
assert.equal(typeof gsdTools.skipsRootResolution, 'function');
|
||||
});
|
||||
});
|
||||
|
||||
278
tests/context-predicates-query.test.cjs
Normal file
278
tests/context-predicates-query.test.cjs
Normal file
@@ -0,0 +1,278 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* Integration tests for `gsd_run query context-predicates` — the selector
|
||||
* surface for the CONTEXT.md predicate fact-store (ADR-1671 S9, #2928 Phase
|
||||
* 1, rows G1-G17).
|
||||
*
|
||||
* NET-NEW COMMAND, EXPECTED RED: `context-predicates` is not yet registered
|
||||
* in gsd-core/bin/gsd-tools.cjs's HOST_COMMAND_ROUTERS / TOP_LEVEL_USAGE /
|
||||
* SKIP_ROOT_RESOLUTION (S9's three hand-maintained sites). Every test in this
|
||||
* file targets the REQUIRED behavior from 40-design.md rows 38-43 and
|
||||
* currently fails because the command does not exist — that is expected and
|
||||
* correct for this commit (a failing-first regression matrix), not a
|
||||
* defect in the test.
|
||||
*
|
||||
* Invocation shape: `gsd_run query <command> [args]` maps to
|
||||
* `node gsd-tools.cjs query <command> [args]` — `query` is a meta-prefix the
|
||||
* dispatcher strips (gsd-core/bin/gsd-tools.cjs, "Accept `query` as a
|
||||
* meta-prefix"). This is the same shape every existing query consumer in
|
||||
* gsd-core/workflows/*.md uses (e.g. `gsd_run query stats.json`,
|
||||
* `gsd_run query prompt-budget ...`) — mirrored here since no real caller of
|
||||
* `context-predicates` is wired yet (net-new capability).
|
||||
*
|
||||
* Structured assertions: success/failure is asserted via exit code, and via
|
||||
* `--json-errors` (a real, already-shipped global gsd-tools flag) parsed as
|
||||
* `{ok, reason, message}` — `reason` is compared against the frozen
|
||||
* ERROR_REASON enum (gsd-core/bin/lib/io.cjs), never a substring match on
|
||||
* `message` prose. On success, output is asserted via JSON.parse of stdout
|
||||
* (gsd-tools' shared `output()` helper always serializes JSON to stdout
|
||||
* unless --raw is passed) — never a stdout regex.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const { runGsdTools, createTempDir, cleanup } = require('./helpers.cjs');
|
||||
const { ERROR_REASON } = require('../gsd-core/bin/lib/io.cjs');
|
||||
const { parsePredicates, selectPredicates } = require('../gsd-core/bin/lib/context-predicates.cjs');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const STACK_FRAME_RE = /\n\s+at\s+\S+\s+\(.*:\d+:\d+\)/;
|
||||
|
||||
function queryContextPredicates(args, cwd = ROOT, env = {}) {
|
||||
return runGsdTools(['query', 'context-predicates', ...args], cwd, env);
|
||||
}
|
||||
|
||||
function queryContextPredicatesJsonErrors(args, cwd = ROOT) {
|
||||
const r = queryContextPredicates([...args, '--json-errors'], cwd);
|
||||
let parsedError = null;
|
||||
try {
|
||||
parsedError = JSON.parse(r.error);
|
||||
} catch {
|
||||
// r.error was not JSON (e.g. a resource-starvation message) — leave null.
|
||||
}
|
||||
return { ...r, parsedError };
|
||||
}
|
||||
|
||||
// Independently computed expectation set, from the real CONTEXT.md, via the
|
||||
// exported pure parser/selector — NOT by parsing the CLI's own rendered text.
|
||||
const REAL_PREDICATES = parsePredicates(fs.readFileSync(path.join(ROOT, 'CONTEXT.md'), 'utf8')).predicates;
|
||||
|
||||
describe('gsd_run query context-predicates (G)', () => {
|
||||
test('queryReturnsPredicatesForClass', () => {
|
||||
const expected = selectPredicates(REAL_PREDICATES, { klass: 'META' });
|
||||
assert.ok(expected.length > 0, 'fixture sanity: META class must exist in the real CONTEXT.md');
|
||||
|
||||
const r = queryContextPredicates(['--class', 'META']);
|
||||
assert.equal(r.success, true, 'query --class META must succeed once the command is wired');
|
||||
const parsed = JSON.parse(r.output);
|
||||
const ids = (Array.isArray(parsed) ? parsed : parsed.predicates).map((p) => p.id).sort();
|
||||
assert.deepEqual(ids, expected.map((p) => p.id).sort());
|
||||
});
|
||||
|
||||
test('queryReturnsPredicatesForDottedPrefix', () => {
|
||||
const expected = selectPredicates(REAL_PREDICATES, { prefix: 'RULESET.TESTS' });
|
||||
assert.ok(expected.length > 0, 'fixture sanity: RULESET.TESTS.* must exist in the real CONTEXT.md');
|
||||
|
||||
const r = queryContextPredicates(['--prefix', 'RULESET.TESTS']);
|
||||
assert.equal(r.success, true);
|
||||
const parsed = JSON.parse(r.output);
|
||||
const ids = (Array.isArray(parsed) ? parsed : parsed.predicates).map((p) => p.id).sort();
|
||||
assert.deepEqual(ids, expected.map((p) => p.id).sort());
|
||||
});
|
||||
|
||||
test('queryReturnsPredicatesForFreeTextContains', () => {
|
||||
const expected = selectPredicates(REAL_PREDICATES, { contains: 'changeset' });
|
||||
assert.ok(expected.length > 0, 'fixture sanity: "changeset" must appear in some real predicate id/value');
|
||||
|
||||
const r = queryContextPredicates(['--contains', 'changeset']);
|
||||
assert.equal(r.success, true);
|
||||
const parsed = JSON.parse(r.output);
|
||||
const ids = (Array.isArray(parsed) ? parsed : parsed.predicates).map((p) => p.id).sort();
|
||||
assert.deepEqual(ids, expected.map((p) => p.id).sort());
|
||||
});
|
||||
|
||||
test('queryReportsZeroMatchesStructurally', () => {
|
||||
const r = queryContextPredicates(['--class', 'ZZZ-NO-SUCH-CLASS-EXISTS']);
|
||||
assert.equal(r.success, true, 'a no-match query is not itself a failure');
|
||||
const parsed = JSON.parse(r.output);
|
||||
assert.equal(parsed.matched, 0, 'a no-match result must be structurally distinguishable ({matched: 0}), not silent success or failure');
|
||||
});
|
||||
|
||||
test('queryWithoutSelectorExitsWithUsage', () => {
|
||||
const r = queryContextPredicatesJsonErrors([]);
|
||||
assert.equal(r.success, false);
|
||||
assert.notEqual(r.exitCode, 0);
|
||||
assert.ok(r.parsedError, 'failure must be structured JSON under --json-errors');
|
||||
assert.equal(r.parsedError.reason, ERROR_REASON.USAGE, 'missing selector must be a usage error, not an internal/unknown-command error');
|
||||
});
|
||||
|
||||
test('queryWithEmptyClassExitsNonZero', () => {
|
||||
const r = queryContextPredicatesJsonErrors(['--class', '']);
|
||||
assert.equal(r.success, false);
|
||||
assert.equal(r.parsedError && r.parsedError.reason, ERROR_REASON.USAGE);
|
||||
});
|
||||
|
||||
test('queryWithWhitespaceOnlyClassExitsNonZero', () => {
|
||||
const r = queryContextPredicatesJsonErrors(['--class', ' ']);
|
||||
assert.equal(r.success, false);
|
||||
assert.equal(r.parsedError && r.parsedError.reason, ERROR_REASON.USAGE);
|
||||
});
|
||||
|
||||
test('queryWithDuplicateClassFlagsResolvesDeterministically', () => {
|
||||
const first = queryContextPredicates(['--class', 'META', '--class', 'RULESET']);
|
||||
const second = queryContextPredicates(['--class', 'META', '--class', 'RULESET']);
|
||||
assert.equal(first.exitCode, second.exitCode, 'duplicate flags must resolve the same way on every invocation');
|
||||
assert.equal(first.output, second.output, 'duplicate-flag resolution must be deterministic, not order-of-parse-dependent');
|
||||
});
|
||||
|
||||
test('queryWithConflictingSelectorsAppliesDocumentedPrecedence', () => {
|
||||
const r = queryContextPredicates(['--class', 'META', '--prefix', 'RULESET.TESTS']);
|
||||
assert.doesNotMatch(r.error || '', STACK_FRAME_RE, 'conflicting selectors must resolve via documented precedence, never crash');
|
||||
const repeat = queryContextPredicates(['--class', 'META', '--prefix', 'RULESET.TESTS']);
|
||||
assert.equal(r.exitCode, repeat.exitCode, 'precedence resolution must be deterministic');
|
||||
});
|
||||
|
||||
test('queryWithMalformedAssignmentExitsNonZero', () => {
|
||||
const eq = queryContextPredicatesJsonErrors(['--class=']);
|
||||
assert.equal(eq.success, false);
|
||||
assert.equal(eq.parsedError && eq.parsedError.reason, ERROR_REASON.USAGE);
|
||||
|
||||
const doubleEq = queryContextPredicatesJsonErrors(['--class==A']);
|
||||
assert.equal(doubleEq.success, false);
|
||||
assert.equal(doubleEq.parsedError && doubleEq.parsedError.reason, ERROR_REASON.USAGE);
|
||||
});
|
||||
|
||||
test('queryTreatsFlagLikeValueAsMissingValue', () => {
|
||||
const r = queryContextPredicatesJsonErrors(['--class', '--weird']);
|
||||
assert.equal(r.success, false, '--weird must not be silently consumed as the --class value');
|
||||
assert.equal(r.parsedError && r.parsedError.reason, ERROR_REASON.USAGE);
|
||||
});
|
||||
|
||||
// #2928 review finding C: a flag-shaped selector value (e.g. searching CONTEXT.md
|
||||
// for the literal text "--since") was previously unmatchable — the space-separated
|
||||
// form always reads a following `--...` token as a missing value, with no escape hatch.
|
||||
describe('inline-assignment escape hatch for flag-shaped values (#2928 finding C)', () => {
|
||||
test('contains=value form matches a flag-shaped selector value', () => {
|
||||
const expected = selectPredicates(REAL_PREDICATES, { contains: '--since' });
|
||||
assert.ok(expected.length > 0, 'fixture sanity: "--since" must appear in a real CONTEXT.md predicate value');
|
||||
|
||||
const r = queryContextPredicates(['--contains=--since']);
|
||||
assert.equal(r.success, true, '--contains=--since must be accepted, not read as an unknown flag');
|
||||
const parsed = JSON.parse(r.output);
|
||||
const ids = parsed.predicates.map((p) => p.id).sort();
|
||||
assert.deepEqual(ids, expected.map((p) => p.id).sort());
|
||||
});
|
||||
|
||||
test('space-separated form still treats a flag-shaped value as missing (no escape without =)', () => {
|
||||
const r = queryContextPredicatesJsonErrors(['--contains', '--since']);
|
||||
assert.equal(r.success, false, '--contains --since (space-separated) must remain a usage error');
|
||||
assert.equal(r.parsedError && r.parsedError.reason, ERROR_REASON.USAGE);
|
||||
});
|
||||
|
||||
test('contains= with an empty value is still rejected', () => {
|
||||
const r = queryContextPredicatesJsonErrors(['--contains=']);
|
||||
assert.equal(r.success, false);
|
||||
assert.equal(r.parsedError && r.parsedError.reason, ERROR_REASON.USAGE);
|
||||
});
|
||||
|
||||
test('contains== (double-equals) is still rejected, not accepted as literal "=x"', () => {
|
||||
const r = queryContextPredicatesJsonErrors(['--contains==x']);
|
||||
assert.equal(r.success, false);
|
||||
assert.equal(r.parsedError && r.parsedError.reason, ERROR_REASON.USAGE);
|
||||
});
|
||||
|
||||
test('class=/prefix= inline-assignment form also works (not contains-only)', () => {
|
||||
const expected = selectPredicates(REAL_PREDICATES, { klass: 'META' });
|
||||
const r = queryContextPredicates(['--class=META']);
|
||||
assert.equal(r.success, true);
|
||||
const parsed = JSON.parse(r.output);
|
||||
assert.deepEqual(
|
||||
parsed.predicates.map((p) => p.id).sort(),
|
||||
expected.map((p) => p.id).sort(),
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
test('queryRejectsPrototypePollutingSelectorKeys', () => {
|
||||
for (const hostile of ['__proto__', 'constructor', 'prototype']) {
|
||||
const r = queryContextPredicates(['--class', hostile]);
|
||||
assert.equal(r.success, true, `${hostile} must be treated as an ordinary (non-matching) class name, not crash`);
|
||||
const parsed = JSON.parse(r.output);
|
||||
assert.equal(parsed.matched, 0);
|
||||
assert.equal(Object.prototype.toString.call({}), '[object Object]', 'sanity: global Object.prototype must be untouched');
|
||||
}
|
||||
});
|
||||
|
||||
test('queryDoesNotInterpolateShellMetacharacters', () => {
|
||||
const markerPath = path.join(ROOT, '.gsd-test-shell-injection-marker-2928');
|
||||
assert.equal(fs.existsSync(markerPath), false, 'fixture sanity: marker must not pre-exist');
|
||||
const hostileValue = `; touch ${markerPath} #`;
|
||||
const r = queryContextPredicates(['--contains', hostileValue]);
|
||||
assert.equal(fs.existsSync(markerPath), false, 'a shell metacharacter payload must never be interpolated into a shell');
|
||||
assert.doesNotMatch(r.error || '', STACK_FRAME_RE);
|
||||
});
|
||||
|
||||
test('queryHandlesVeryLongSelectorValue', () => {
|
||||
const longValue = 'x'.repeat(32 * 1024);
|
||||
const r = queryContextPredicates(['--contains', longValue]);
|
||||
assert.equal(typeof r.exitCode, 'number', 'a 32K selector value must not hang or crash the process');
|
||||
assert.doesNotMatch(r.error || '', STACK_FRAME_RE);
|
||||
});
|
||||
|
||||
test('queryHandlesUnicodeSelectorValue', () => {
|
||||
const unicodeValue = String.fromCodePoint(0x9884, 0x6d4b, 0x5909, 0x6570, 0x1f525);
|
||||
const r = queryContextPredicates(['--contains', unicodeValue]);
|
||||
assert.equal(typeof r.exitCode, 'number', 'a Unicode selector value must not crash the process');
|
||||
assert.doesNotMatch(r.error || '', STACK_FRAME_RE);
|
||||
});
|
||||
|
||||
test('queryWorksFromSubdirectoryWithoutPlanningDir', (t) => {
|
||||
const tmp = createTempDir('gci-query-subdir-');
|
||||
t.after(() => cleanup(tmp));
|
||||
const r = queryContextPredicates(['--class', 'META'], tmp);
|
||||
assert.equal(r.success, true, 'the read-only selector must not require a .planning/ directory');
|
||||
});
|
||||
|
||||
test('queryFailureOutputContainsNoStackTrace', () => {
|
||||
const r = queryContextPredicates([]);
|
||||
assert.equal(r.success, false);
|
||||
assert.doesNotMatch(r.error || '', STACK_FRAME_RE, 'no bare stack trace in non-debug failure output');
|
||||
});
|
||||
});
|
||||
|
||||
// allow-test-rule: source-text-is-the-product #2928
|
||||
// Wiring check for #2928's acceptance criterion — "an agent/orchestrator can invoke the
|
||||
// selector through a wired `gsd_run query` surface to assemble a brief". Before this, nothing
|
||||
// in the repo called `context-predicates`; it was a stranded CLI. docs/contributor-standards.md
|
||||
// §"Pre-work requirements" is the real, tested brief-assembly site governing how an AI-agent
|
||||
// prompt must cite CONTEXT.md's META.RULE predicates — this asserts it routes through the
|
||||
// selector instead of a grep-by-eye read, and that the selector actually returns the predicate
|
||||
// set that site names.
|
||||
describe('context-predicates wired into a real brief-assembly site (#2928)', () => {
|
||||
const STANDARDS_DOC = path.join(ROOT, 'docs', 'contributor-standards.md');
|
||||
const standardsText = fs.readFileSync(STANDARDS_DOC, 'utf8');
|
||||
|
||||
test('contributorStandardsRoutesPredicateCitationThroughTheQuerySelector', () => {
|
||||
assert.ok(
|
||||
standardsText.includes('node gsd-tools.cjs query context-predicates --class'),
|
||||
'docs/contributor-standards.md must invoke the context-predicates selector rather than instructing a grep-by-eye read'
|
||||
);
|
||||
assert.ok(
|
||||
standardsText.includes('META.RULE.brief-must-cite-doc') && standardsText.includes('META.RULE.brief-no-paraphrase'),
|
||||
'the wiring site must name the predicates it expects the selector to surface'
|
||||
);
|
||||
});
|
||||
|
||||
test('theSelectorActuallyReturnsThePredicatesTheWiringSiteNames', () => {
|
||||
const r = queryContextPredicates(['--class', 'META']);
|
||||
assert.equal(r.success, true);
|
||||
const parsed = JSON.parse(r.output);
|
||||
const ids = parsed.predicates.map((p) => p.id);
|
||||
assert.ok(ids.includes('META.RULE.brief-must-cite-doc'));
|
||||
assert.ok(ids.includes('META.RULE.brief-no-paraphrase'));
|
||||
});
|
||||
});
|
||||
228
tests/context-predicates.property.test.cjs
Normal file
228
tests/context-predicates.property.test.cjs
Normal file
@@ -0,0 +1,228 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* Property-based tests for src/context-predicates.cts (compiled to
|
||||
* gsd-core/bin/lib/context-predicates.cjs).
|
||||
*
|
||||
* Document-shaped generators (CONTRIBUTING.md "Fixture provenance #2371"):
|
||||
* these generators build arbitrary markdown documents out of prose lines,
|
||||
* fences of varying tick-length, list items with varying markers, and
|
||||
* predicate-shaped / predicate-lookalike lines — they are NOT seeded from
|
||||
* this module's own writer/serializer. Seeding a property generator from the
|
||||
* code under test's own render function make the document shape a constant
|
||||
* and the property unable to fail; see 50-test-matrix.md Step 1.
|
||||
*
|
||||
* Deterministic per CONTRIBUTING.md: seed and numRuns are pinned by
|
||||
* tests/helpers/fast-check-setup.cjs (seed 42, numRuns 200); failures print
|
||||
* replay data via fast-check's own counterexample + seed reporting.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fc = require('./helpers/fast-check-setup.cjs');
|
||||
|
||||
const { parsePredicates, selectPredicates, buildIndex } = require('../gsd-core/bin/lib/context-predicates.cjs');
|
||||
|
||||
// ─── Document-shaped generators ────────────────────────────────────────────
|
||||
|
||||
// A predicate-shaped line: `CLASS.subkey=value` in one of the recognized
|
||||
// declaration forms. `asDeclared` controls whether the caller wants this
|
||||
// specific fixture to be a genuinely-recognized declaration (bare or `- `
|
||||
// list item — the two forms even today's implementation accepts) so property
|
||||
// H1/H4 can reason about "predicates the parser actually extracts" without
|
||||
// depending on the A4-A7 defects under test elsewhere.
|
||||
const idClassArb = fc.stringMatching(/^[A-Z][A-Z0-9_-]{0,8}$/);
|
||||
const idSubkeyArb = fc.stringMatching(/^[A-Za-z0-9_-]{1,8}$/);
|
||||
const valueArb = fc.stringMatching(/^[A-Za-z0-9 _.-]{1,20}$/);
|
||||
|
||||
const declaredPredicateLineArb = fc
|
||||
.tuple(idClassArb, fc.option(idSubkeyArb, { nil: undefined }), valueArb, fc.boolean())
|
||||
.map(([klass, subkey, value, asListItem]) => {
|
||||
const id = subkey ? `${klass}.${subkey}` : klass;
|
||||
const inner = `\`${id}=${value}\``;
|
||||
return { text: asListItem ? `- ${inner}` : inner, id, klass, value };
|
||||
});
|
||||
|
||||
// A prose line that is NOT a predicate declaration: plain text, a heading, a
|
||||
// blockquote line, a table row, or an inline mid-prose backtick reference.
|
||||
const proseLineArb = fc.oneof(
|
||||
fc.stringMatching(/^[A-Za-z0-9 .,'"()-]{0,40}$/),
|
||||
idClassArb.map((k) => `# ${k} heading`),
|
||||
idClassArb.map((k) => `> quoting ${k}`),
|
||||
idClassArb.map((k) => `| cell | \`${k}.x=y\` |`),
|
||||
declaredPredicateLineArb.map((p) => `see ${p.text.replace(/^- /, '')} for details`),
|
||||
);
|
||||
|
||||
const fenceTickArb = fc.constantFrom('```', '~~~', '````');
|
||||
|
||||
// A whole document assembled from a mix of prose lines and declared
|
||||
// predicate lines, optionally wrapping a contiguous run in a fence.
|
||||
function documentArb() {
|
||||
return fc
|
||||
.array(fc.oneof({ arbitrary: declaredPredicateLineArb, weight: 2 }, { arbitrary: proseLineArb, weight: 3 }), {
|
||||
minLength: 0,
|
||||
maxLength: 12,
|
||||
})
|
||||
.map((items) => items);
|
||||
}
|
||||
|
||||
// ─── H1: index preserves every parsed predicate ────────────────────────────
|
||||
|
||||
describe('property: parse <-> index contract', () => {
|
||||
test('indexPreservesEveryParsedPredicate', () => {
|
||||
fc.assert(
|
||||
fc.property(documentArb(), (items) => {
|
||||
const md = items.map((it) => (typeof it === 'string' ? it : it.text)).join('\n');
|
||||
const parsed = parsePredicates(md);
|
||||
const index = buildIndex(parsed.predicates);
|
||||
|
||||
assert.equal(index.count, parsed.predicates.length);
|
||||
|
||||
const parsedIds = parsed.predicates.map((p) => p.id).sort();
|
||||
const indexIds = index.predicates.map((p) => p.id).sort();
|
||||
assert.deepEqual(indexIds, parsedIds, 'buildIndex must invent or lose no predicate id');
|
||||
|
||||
for (const p of parsed.predicates) {
|
||||
const inIndex = index.predicates.find((ip) => ip.id === p.id && ip.value === p.value);
|
||||
assert.ok(inIndex, `predicate ${p.id}=${p.value} from the parse must appear in the index`);
|
||||
}
|
||||
}),
|
||||
);
|
||||
});
|
||||
|
||||
// ─── H2: buildIndex is deterministic and order-independent ────────────────
|
||||
|
||||
test('indexSerializationIsOrderIndependentAndDeterministic', () => {
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.array(declaredPredicateLineArb, { minLength: 0, maxLength: 10 }),
|
||||
(decls) => {
|
||||
// Dedupe by id so this property is not entangled with duplicate
|
||||
// semantics (covered separately by the E-row unit tests) — the
|
||||
// property under test here is pure ordering independence.
|
||||
const seen = new Set();
|
||||
const uniqueDecls = decls.filter((d) => (seen.has(d.id) ? false : (seen.add(d.id), true)));
|
||||
|
||||
const forwardMd = uniqueDecls.map((d) => d.text).join('\n');
|
||||
const reverseMd = uniqueDecls
|
||||
.slice()
|
||||
.reverse()
|
||||
.map((d) => d.text)
|
||||
.join('\n');
|
||||
|
||||
const forwardIndex = buildIndex(parsePredicates(forwardMd).predicates);
|
||||
const reverseIndex = buildIndex(parsePredicates(reverseMd).predicates);
|
||||
|
||||
assert.deepEqual(forwardIndex, reverseIndex, 'index must not depend on source declaration order');
|
||||
|
||||
// Determinism: building twice from the same parsed predicates must
|
||||
// produce byte-identical JSON serialization.
|
||||
const parsed = parsePredicates(forwardMd);
|
||||
const a = JSON.stringify(buildIndex(parsed.predicates));
|
||||
const b = JSON.stringify(buildIndex(parsed.predicates));
|
||||
assert.equal(a, b);
|
||||
},
|
||||
),
|
||||
);
|
||||
});
|
||||
|
||||
// ─── H3: fencing a region never increases the predicate count ─────────────
|
||||
|
||||
test('fencingRegionNeverIncreasesPredicateCount', () => {
|
||||
fc.assert(
|
||||
fc.property(
|
||||
fc.array(declaredPredicateLineArb, { minLength: 1, maxLength: 6 }),
|
||||
fc.nat({ max: 5 }),
|
||||
fenceTickArb,
|
||||
(decls, wrapAt, fence) => {
|
||||
const lines = decls.map((d) => d.text);
|
||||
const before = parsePredicates(lines.join('\n')).predicates.length;
|
||||
|
||||
const cut = Math.min(wrapAt, lines.length);
|
||||
const fenced = [...lines.slice(0, cut), fence, ...lines.slice(cut), fence].join('\n');
|
||||
const after = parsePredicates(fenced).predicates.length;
|
||||
|
||||
assert.ok(after <= before, `fencing must never increase the parsed count (before=${before}, after=${after})`);
|
||||
},
|
||||
),
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── H5: comment/fence mutual precedence (DEFECT.CONTEXT-PREDICATES-COMMENT-
|
||||
// FENCE-BLIND, #2928 review) — a predicate genuinely OUTSIDE a wrapper
|
||||
// (comment or fence) is always parsed live, and a predicate genuinely INSIDE
|
||||
// it is never parsed live, even when the wrapper's own content contains
|
||||
// tokens that LOOK LIKE the OTHER construct (fence delimiters inside a
|
||||
// comment, or comment tokens inside a fence — the two directions the
|
||||
// interleaved single-pass in `computeSkippedLineFlags` must both get right).
|
||||
// Document-shaped generator: wraps a marked "inside" predicate between an
|
||||
// open/close pair of ONE kind, salted with lookalike noise from the OTHER
|
||||
// kind, with unrelated real predicates before/after the wrapper. ──────────
|
||||
|
||||
// Noise lines that look like the OTHER construct's tokens, keyed by which
|
||||
// kind is being used as the OUTER wrapper for a given run.
|
||||
const FENCE_NOISE_LINES = ['<!-- not a real comment, just fence content', 'text mentioning --> mid-line', '<!-- nested-looking'];
|
||||
const COMMENT_NOISE_LINES = ['```', '~~~', '``` info string'];
|
||||
|
||||
const wrapperKindArb = fc.constantFrom('fence', 'comment');
|
||||
|
||||
describe('property: comment/fence mutual precedence', () => {
|
||||
test('predicateOutsideWrapperIsAlwaysLiveAndInsideIsNeverLive', () => {
|
||||
fc.assert(
|
||||
fc.property(
|
||||
wrapperKindArb,
|
||||
fc.array(fc.constantFrom(0, 1, 2), { minLength: 0, maxLength: 3 }),
|
||||
fc.array(fc.constantFrom(0, 1, 2), { minLength: 0, maxLength: 3 }),
|
||||
(wrapperKind, beforeNoiseIdx, afterNoiseIdx) => {
|
||||
const noisePool = wrapperKind === 'fence' ? FENCE_NOISE_LINES : COMMENT_NOISE_LINES;
|
||||
const open = wrapperKind === 'fence' ? '```' : '<!-- wrapper open';
|
||||
const close = wrapperKind === 'fence' ? '```' : '-->';
|
||||
|
||||
const lines = [
|
||||
'`OUTSIDE_BEFORE=1`',
|
||||
open,
|
||||
...beforeNoiseIdx.map((i) => noisePool[i]),
|
||||
'`INSIDE_MARKER=2`',
|
||||
...afterNoiseIdx.map((i) => noisePool[i]),
|
||||
close,
|
||||
'`OUTSIDE_AFTER=3`',
|
||||
];
|
||||
|
||||
const md = lines.join('\n');
|
||||
const r = parsePredicates(md);
|
||||
const ids = new Set(r.predicates.map((p) => p.id));
|
||||
|
||||
assert.ok(ids.has('OUTSIDE_BEFORE'), `OUTSIDE_BEFORE must always be live (wrapper=${wrapperKind})`);
|
||||
assert.ok(ids.has('OUTSIDE_AFTER'), `OUTSIDE_AFTER must always be live (wrapper=${wrapperKind})`);
|
||||
assert.ok(
|
||||
!ids.has('INSIDE_MARKER'),
|
||||
`INSIDE_MARKER must never be live inside a real ${wrapperKind}, even with ${
|
||||
wrapperKind === 'fence' ? 'comment' : 'fence'
|
||||
}-lookalike noise around it`,
|
||||
);
|
||||
},
|
||||
),
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
describe('property: selectPredicates subset invariant', () => {
|
||||
test('selectorReturnsOnlyMatchingSubset', () => {
|
||||
fc.assert(
|
||||
fc.property(fc.array(declaredPredicateLineArb, { minLength: 0, maxLength: 10 }), idClassArb, (decls, klass) => {
|
||||
const md = decls.map((d) => d.text).join('\n');
|
||||
const all = parsePredicates(md).predicates;
|
||||
const selected = selectPredicates(all, { klass });
|
||||
|
||||
for (const p of selected) {
|
||||
assert.ok(
|
||||
all.includes(p),
|
||||
'every selected predicate must be a reference from the original array (subset, not a copy with invented members)',
|
||||
);
|
||||
assert.equal(p.klass, klass, `selectPredicates({klass}) must only return predicates whose klass === ${klass}`);
|
||||
}
|
||||
}),
|
||||
);
|
||||
});
|
||||
});
|
||||
684
tests/context-predicates.test.cjs
Normal file
684
tests/context-predicates.test.cjs
Normal file
@@ -0,0 +1,684 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* Unit tests for src/context-predicates.cts (compiled to
|
||||
* gsd-core/bin/lib/context-predicates.cjs) — the CONTEXT.md predicate
|
||||
* fact-store parser/validator/selector (ADR-1671, #2928 Phase 1).
|
||||
*
|
||||
* FAILING-FIRST: this file targets the REQUIRED production behavior from
|
||||
* .gsd/phase/chore-2928-context-predicate-store/40-design.md's behavior
|
||||
* table, not today's prototype-carried-forward behavior. Eight row-groups
|
||||
* are measured, provable defects in the current implementation and MUST be
|
||||
* RED until a later commit fixes them: A4, A5, A6, A7 (indented-bare / `*` /
|
||||
* `+` / numbered-list declaration forms are dropped), B2 (tilde fence not
|
||||
* skipped), B3 (4-backtick fence containing a 3-backtick line mis-toggles),
|
||||
* B5 (multi-line HTML comment not skipped).
|
||||
*
|
||||
* Fixture provenance (CONTRIBUTING.md #2371): the must-NOT-parse corpus
|
||||
* (rows A8-A11, negative space, I4) is drawn from real repo documents that
|
||||
* predate/ignore this grammar — the real CONTEXT.md and the real
|
||||
* CONTRIBUTING.md — never hand-authored from the grammar spec itself.
|
||||
*/
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const {
|
||||
parsePredicates,
|
||||
selectPredicates,
|
||||
buildIndex,
|
||||
} = require('../gsd-core/bin/lib/context-predicates.cjs');
|
||||
|
||||
const { scanFencedBlocks } = require('../gsd-core/bin/lib/markdown-sectionizer.cjs');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const CRLF = '\r\n';
|
||||
|
||||
// ─── A. Declaration recognition ───────────────────────────────────────────
|
||||
|
||||
describe('parsePredicates: declaration recognition (A)', () => {
|
||||
test('parsesBareBacktickDeclaration', () => {
|
||||
const r = parsePredicates('`FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.deepEqual(
|
||||
{ id: r.predicates[0].id, klass: r.predicates[0].klass, value: r.predicates[0].value },
|
||||
{ id: 'FOO', klass: 'FOO', value: 'bar' },
|
||||
);
|
||||
});
|
||||
|
||||
test('parsesDashListItemDeclaration', () => {
|
||||
const r = parsePredicates('- `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
|
||||
test('parsesIndentedDashListItemDeclaration', () => {
|
||||
const r = parsePredicates(' - `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
|
||||
test('parsesIndentedBareDeclaration', () => {
|
||||
// RED (measured defect): the prototype's bare-form check requires column
|
||||
// 0 (`line.startsWith('`')` on a line whose only trim is trimEnd()), so
|
||||
// leading whitespace with no list marker is silently dropped today.
|
||||
const r = parsePredicates(' `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1, 'indented bare declaration must be tolerated (Postel honored on shape)');
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
|
||||
test('parsesStarListItemDeclaration', () => {
|
||||
// RED (measured defect): today's list-item stripper only recognizes `-`.
|
||||
const r = parsePredicates('* `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
|
||||
test('parsesPlusListItemDeclaration', () => {
|
||||
// RED (measured defect): same as `*`, `+` is not recognized today.
|
||||
const r = parsePredicates('+ `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
|
||||
test('parsesNumberedListItemDeclaration', () => {
|
||||
// RED (measured defect): numbered list markers are not recognized today.
|
||||
const r = parsePredicates('1. `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
|
||||
test('ignoresInlineReferenceInsideProse', () => {
|
||||
const r = parsePredicates('see `FOO=bar` for details');
|
||||
assert.equal(r.predicates.length, 0, 'an inline mid-prose mention is a reference, not a declaration');
|
||||
});
|
||||
|
||||
test('ignoresPredicateShapeInTableCell', () => {
|
||||
const r = parsePredicates('| x | `FOO=bar` |');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('ignoresPredicateShapeInHeading', () => {
|
||||
const md = ['# `FOO=bar`', '`REAL.one=value`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 1, 'the heading itself must never yield a predicate');
|
||||
assert.equal(r.predicates[0].id, 'REAL.one');
|
||||
assert.equal(r.predicates[0].section, '`FOO=bar`', 'the heading text (predicate-shaped or not) becomes the tracked section');
|
||||
});
|
||||
|
||||
test('ignoresPredicateShapeInBlockquote', () => {
|
||||
const r = parsePredicates('> `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('tracksNearestPrecedingSectionHeading', () => {
|
||||
const md = ['# My Section', '`FOO=bar`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].section, 'My Section');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── B. Fenced / commented regions ────────────────────────────────────────
|
||||
|
||||
describe('parsePredicates: fenced/commented regions (B)', () => {
|
||||
test('ignoresDeclarationInsideBacktickFence', () => {
|
||||
const md = ['```', '`FOO=bar`', '```'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('ignoresDeclarationInsideTildeFence', () => {
|
||||
// RED (measured defect): the naive toggle only matches triple-backtick lines.
|
||||
const md = ['~~~', '`FOO=bar`', '~~~'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0, 'a tilde fence must skip its contents exactly like a backtick fence');
|
||||
});
|
||||
|
||||
test('ignoresDeclarationInsideLongerFenceContainingShorterFence', () => {
|
||||
// RED (measured defect): a naive backtick-count-agnostic toggle mis-flips
|
||||
// on the inner 3-backtick line and un-skips the remainder.
|
||||
const md = ['````', '```', '`FOO=bar`', '```', '````'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0, 'fence-length awareness must prevent the inner fence from un-skipping the outer one');
|
||||
});
|
||||
|
||||
test('ignoresDeclarationInsideLanguageTaggedFence', () => {
|
||||
const md = ['```bash', '`FOO=bar`', '```'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('ignoresDeclarationInsideMultiLineHtmlComment', () => {
|
||||
// RED (measured defect): the prototype has no HTML-comment awareness at
|
||||
// all — a predicate-shaped line between `<!--` and `-->` on its own line
|
||||
// parses as live today.
|
||||
const md = ['<!--', '`FOO=bar`', '-->'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0, 'a commented-out predicate must not be read as live');
|
||||
});
|
||||
|
||||
test('ignoresDeclarationInsideSingleLineHtmlComment', () => {
|
||||
const r = parsePredicates('<!-- `FOO=bar` -->');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('treatsUnclosedFenceAsSkippedToEndOfFile', () => {
|
||||
const md = ['```', '`FOO=bar`', '`BAZ=qux`'].join('\n');
|
||||
assert.doesNotThrow(() => parsePredicates(md));
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0, 'everything after an unclosed fence must be treated as skipped, not crash or leak');
|
||||
});
|
||||
|
||||
test('parsesDeclarationInFourSpaceIndentedBlockAsDocumentedLimit', () => {
|
||||
// Pins the documented limit (design known-limit 1): 4-space indentation
|
||||
// is NOT treated as a code block, because CONTEXT.md authors real
|
||||
// predicates as indented list items at that depth.
|
||||
const r = parsePredicates(' - `FOO=bar`');
|
||||
assert.equal(r.predicates.length, 1, 'documented limit: 4-space indent is not code, so this must still parse');
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── B2. Comment/fence mutual precedence
|
||||
// (DEFECT.CONTEXT-PREDICATES-COMMENT-FENCE-BLIND — BLOCKER review finding) ──
|
||||
//
|
||||
// The HTML-comment scan and the fence scan previously ran as two
|
||||
// INDEPENDENT passes: `scanFencedBlocks` is comment-blind, so a fence
|
||||
// delimiter appearing INSIDE an HTML comment (with no later matching close
|
||||
// in the file) was treated as a real *unterminated* fence — silently
|
||||
// dropping every remaining predicate to EOF. This is invisible to `--check`
|
||||
// because `--check` diffs against a baseline generated by the same
|
||||
// corrupted parse. This suite locks the chosen precedence (module doc
|
||||
// comment, `computeSkippedLineFlags`): the two constructs are scanned in one
|
||||
// interleaved pass and mutually suppress each other's open/close detection
|
||||
// while active.
|
||||
|
||||
describe('parsePredicates: comment/fence mutual precedence (B2)', () => {
|
||||
test('fenceDelimiterInsideCommentWithNoLaterFenceDoesNotSwallowRemainingFile', () => {
|
||||
// The exact BLOCKER repro: a fence delimiter inside an HTML comment, with
|
||||
// no later real fence anywhere in the document, must not be treated as
|
||||
// an unterminated fence — the later real predicate must still parse.
|
||||
const md = ['<!-- example:', '```', '-->', '`RULESET.REAL.PREDICATE=x`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 1, 'the fence delimiter inside the comment must not swallow the rest of the file');
|
||||
assert.equal(r.predicates[0].id, 'RULESET.REAL.PREDICATE');
|
||||
});
|
||||
|
||||
test('fenceDelimiterInsideCommentWithLaterRealFenceStillBehavesCorrectly', () => {
|
||||
const md = [
|
||||
'<!-- example:',
|
||||
'```',
|
||||
'-->',
|
||||
'`BEFORE=1`',
|
||||
'```',
|
||||
'`INSIDE=2`',
|
||||
'```',
|
||||
'`AFTER=3`',
|
||||
].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
const ids = r.predicates.map((p) => p.id);
|
||||
assert.deepEqual(ids.sort(), ['AFTER', 'BEFORE'], 'the real fence after the comment must still skip its own content');
|
||||
});
|
||||
|
||||
test('arrowInsideRealFencedBlockDoesNotTerminateAnythingAndFenceStillSkipsItsContent', () => {
|
||||
const md = ['```', '`INSIDE=1`', 'noise line with --> token', '```', '`AFTER=2`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
const ids = r.predicates.map((p) => p.id);
|
||||
assert.deepEqual(ids, ['AFTER'], '--> inside real fence content must not open/close a comment; fence must still skip INSIDE');
|
||||
});
|
||||
|
||||
test('commentOpenerInsideRealFencedBlockDoesNotOpenAComment', () => {
|
||||
const md = ['```', '<!-- looks like a comment opener, but this is fence content', '`INSIDE=1`', '```', '`AFTER=2`'].join(
|
||||
'\n',
|
||||
);
|
||||
const r = parsePredicates(md);
|
||||
const ids = r.predicates.map((p) => p.id);
|
||||
assert.deepEqual(ids, ['AFTER'], '<!-- inside real fence content must not open a comment that leaks past the fence close');
|
||||
});
|
||||
|
||||
test('unterminatedCommentAtEofSkipsRemainingLinesWithoutCrashing', () => {
|
||||
const md = ['`BEFORE=1`', '<!-- unterminated', '`INSIDE=2`', '`ALSOINSIDE=3`'].join('\n');
|
||||
assert.doesNotThrow(() => parsePredicates(md));
|
||||
const r = parsePredicates(md);
|
||||
const ids = r.predicates.map((p) => p.id);
|
||||
assert.deepEqual(ids, ['BEFORE'], 'an unterminated comment must skip to EOF, not crash and not leak later predicates');
|
||||
});
|
||||
|
||||
test('commentAndFenceTokensOnTheSameLineResolveWithoutInterference', () => {
|
||||
// A fence-opener line whose CommonMark info string happens to contain a
|
||||
// full inline HTML comment: the line still starts with backticks, not
|
||||
// `<!--`, so it opens a real fence (comment text is just info-string
|
||||
// trailing content) — the fence still closes and skips its own content.
|
||||
const infoStringMd = ['`BEFORE=1`', '``` <!-- note -->', '`INSIDE=2`', '```', '`AFTER=3`'].join('\n');
|
||||
const infoStringResult = parsePredicates(infoStringMd);
|
||||
assert.deepEqual(
|
||||
infoStringResult.predicates.map((p) => p.id).sort(),
|
||||
['AFTER', 'BEFORE'],
|
||||
'a fence opener whose info string contains comment-like text must still open/close as a real fence',
|
||||
);
|
||||
|
||||
// A self-closing single-line HTML comment whose content happens to
|
||||
// contain a fence-delimiter-shaped token: the line still starts with
|
||||
// `<!--`, so the whole line is comment content and the embedded
|
||||
// backticks never open a fence.
|
||||
const commentMd = ['`BEFORE=1`', '<!-- see ``` for an example -->', '`AFTER=2`'].join('\n');
|
||||
const commentResult = parsePredicates(commentMd);
|
||||
assert.deepEqual(
|
||||
commentResult.predicates.map((p) => p.id).sort(),
|
||||
['AFTER', 'BEFORE'],
|
||||
'a single-line comment whose content contains fence-shaped text must not open a fence',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── C. ID / value grammar boundaries ──────────────────────────────────────
|
||||
|
||||
describe('parsePredicates: ID/value grammar boundaries (C)', () => {
|
||||
test('rejectsBacktickContentAtLengthTwo', () => {
|
||||
// `A=` — backtick content length 2 (limit-1 of the `inner.length > 2` guard).
|
||||
const r = parsePredicates('`A=`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('acceptsMinimalBacktickContentAtLengthThree', () => {
|
||||
// `A=1` — length 3 (limit).
|
||||
const r = parsePredicates('`A=1`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.deepEqual({ id: r.predicates[0].id, value: r.predicates[0].value }, { id: 'A', value: '1' });
|
||||
});
|
||||
|
||||
test('acceptsBacktickContentAboveMinimumLength', () => {
|
||||
// `A=12` — length 4 (limit+1).
|
||||
const r = parsePredicates('`A=12`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].value, '12');
|
||||
});
|
||||
|
||||
test('rejectsEqualsSignAtIndexZero', () => {
|
||||
// `=value` — eqIdx 0 (limit-1 of the `eqIdx < 1` guard).
|
||||
const r = parsePredicates('`=value`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('acceptsEqualsSignAtIndexOne', () => {
|
||||
// `A=value` — eqIdx 1 (limit).
|
||||
const r = parsePredicates('`A=value`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'A');
|
||||
});
|
||||
|
||||
test('acceptsEqualsSignAboveIndexOne', () => {
|
||||
// `AB=value` — eqIdx 2 (limit+1).
|
||||
const r = parsePredicates('`AB=value`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'AB');
|
||||
});
|
||||
|
||||
test('rejectsInlineCodeWithoutEqualsSign', () => {
|
||||
const r = parsePredicates('`ID`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('reportsEmptyValueAsMalformedRatherThanDroppingSilently', () => {
|
||||
const r = parsePredicates('`ID=`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
assert.equal(r.malformed.length, 1, 'an empty value must surface as a diagnostic, not vanish silently');
|
||||
assert.equal(r.malformed[0].reason, 'empty-value');
|
||||
});
|
||||
|
||||
test('splitsOnFirstEqualsSignOnly', () => {
|
||||
const r = parsePredicates('`ID=a=b=c`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.deepEqual({ id: r.predicates[0].id, value: r.predicates[0].value }, { id: 'ID', value: 'a=b=c' });
|
||||
});
|
||||
|
||||
test('rejectsLowercaseLeadingIdentifier', () => {
|
||||
const r = parsePredicates('`foo.bar=x`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('acceptsLowercaseSubSegments', () => {
|
||||
const r = parsePredicates('`PRED.k320.rule=x`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].klass, 'PRED');
|
||||
});
|
||||
|
||||
test('acceptsHyphenatedClassSegment', () => {
|
||||
const r = parsePredicates('`RELEASE-NOTES.x=y`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].klass, 'RELEASE-NOTES');
|
||||
});
|
||||
|
||||
test('rejectsWhitespaceInIdentifier', () => {
|
||||
const r = parsePredicates('`FOO BAR=x`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('acceptsSingleSegmentIdentifier', () => {
|
||||
const r = parsePredicates('`FOO=x`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].klass, r.predicates[0].id);
|
||||
});
|
||||
|
||||
test('preservesTrailingWhitespaceInValueVerbatim', () => {
|
||||
const r = parsePredicates('`ID=1 `');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].value, '1 ', 'trailing whitespace in the value must not be trimmed');
|
||||
});
|
||||
|
||||
test('acceptsBacktickWithinValue', () => {
|
||||
const r = parsePredicates('`ID=a`b`');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].value, 'a`b');
|
||||
});
|
||||
|
||||
test('rejectsPathAndQueryCharactersInIdentifier', () => {
|
||||
const r = parsePredicates('`A/B?c=d`');
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('rejectsDoubledDotEmptySegment', () => {
|
||||
// DEFECT.CONTEXT-PREDICATES-ID-REDOS (MAJOR review finding): the
|
||||
// structural id validator (isValidId) splits on '.' and rejects any
|
||||
// empty segment. This is a deliberate behavior CHANGE vs. the old
|
||||
// `ID_RE` regex, which accepted `A..b` because `.` was inside the
|
||||
// subsequent-segment character class. The real repo CONTEXT.md was
|
||||
// checked (`grep -nE '`[A-Z][A-Za-z0-9_.-]*\\.\\.[A-Za-z0-9_.-]*='
|
||||
// CONTEXT.md`) and contains ZERO ids with a doubled dot, so this is a
|
||||
// pure grammar-tightening with no behavior loss against real data —
|
||||
// pinned here so a future change cannot silently re-loosen it.
|
||||
const r = parsePredicates('`A..b=value`');
|
||||
assert.equal(r.predicates.length, 0, 'a doubled dot (empty segment) must be rejected, not silently accepted');
|
||||
});
|
||||
|
||||
test('idGrammarValidationIsLinearTimeAgainstManyConsecutiveDots', () => {
|
||||
// DEFECT.CONTEXT-PREDICATES-ID-REDOS (MAJOR review finding): the old
|
||||
// `ID_RE`'s `(?:\.[A-Za-z0-9_.-]+)*` group was exponential in the number
|
||||
// of consecutive dots (measured: ~565ms for 40 dots). The structural
|
||||
// per-segment validator is linear. A generous wall-clock bound is used
|
||||
// only as a smoke check; the load-bearing assertion is that the result
|
||||
// is a clean rejection (an id-shaped line with 60 consecutive dots has
|
||||
// an empty segment at every step and must not parse).
|
||||
const dots = '.'.repeat(60);
|
||||
const md = `\`A${dots}x=value\``;
|
||||
const start = Date.now();
|
||||
const r = parsePredicates(md);
|
||||
const elapsedMs = Date.now() - start;
|
||||
assert.equal(r.predicates.length, 0, 'an id with 60 consecutive dots has empty segments and must cleanly reject');
|
||||
assert.ok(elapsedMs < 1000, `expected well under 1s (linear time), got ${elapsedMs}ms — possible ReDoS regression`);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── D. CRLF / newline fidelity ────────────────────────────────────────────
|
||||
|
||||
describe('parsePredicates: CRLF/newline fidelity (D)', () => {
|
||||
test('parsesBareDeclarationUnderCrlf', () => {
|
||||
const r = parsePredicates('`FOO=bar`' + CRLF);
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].id, 'FOO');
|
||||
});
|
||||
|
||||
test('parsesListItemDeclarationUnderCrlf', () => {
|
||||
const r = parsePredicates('- `FOO=bar`' + CRLF);
|
||||
assert.equal(r.predicates.length, 1);
|
||||
});
|
||||
|
||||
test('parsesIndentedListDeclarationUnderCrlf', () => {
|
||||
const r = parsePredicates(' - `FOO=bar`' + CRLF);
|
||||
assert.equal(r.predicates.length, 1);
|
||||
});
|
||||
|
||||
test('skipsFencedDeclarationUnderCrlf', () => {
|
||||
const md = ['```', '`FOO=bar`', '```', ''].join(CRLF);
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('skipsBlockquoteDeclarationUnderCrlf', () => {
|
||||
const r = parsePredicates('> `FOO=bar`' + CRLF);
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
|
||||
test('tracksSectionAndLineNumbersUnderCrlf', () => {
|
||||
const md = ['# Section Name', '`FOO=bar`', ''].join(CRLF);
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 1);
|
||||
assert.equal(r.predicates[0].section, 'Section Name');
|
||||
assert.equal(r.predicates[0].line, 2);
|
||||
});
|
||||
|
||||
test('parsesMixedLfAndCrlfDocument', () => {
|
||||
const md = '`FOO=bar`' + CRLF + '- `BAZ=qux`' + '\n';
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 2);
|
||||
assert.deepEqual(r.predicates.map((p) => p.id).sort(), ['BAZ', 'FOO']);
|
||||
});
|
||||
|
||||
test('yieldsNoPredicatesForLoneCrDocumentAsDocumentedLimit', () => {
|
||||
// Pins the documented limit (design known-limit 2): lone-CR-only line
|
||||
// endings are unsupported; no `\r`-only file exists in this repo.
|
||||
const md = '`FOO=bar`\r`BAZ=qux`\r';
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── E. Duplicate detection + validation ───────────────────────────────────
|
||||
|
||||
describe('parsePredicates: duplicate detection + validation (E)', () => {
|
||||
test('reportsDuplicateIdentifierWithDifferentValues', () => {
|
||||
const md = ['`FOO=a`', '`FOO=b`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 1);
|
||||
assert.deepEqual(r.duplicates[0], { id: 'FOO', count: 2 });
|
||||
});
|
||||
|
||||
test('reportsDuplicateIdentifierEvenWhenValuesAreIdentical', () => {
|
||||
const md = ['`FOO=a`', '`FOO=a`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 1, 'identical-value duplicates must not be silently deduped');
|
||||
assert.deepEqual(r.duplicates[0], { id: 'FOO', count: 2 });
|
||||
});
|
||||
|
||||
test('reportsDuplicateCountForThreeOccurrences', () => {
|
||||
const md = ['`FOO=a`', '`FOO=b`', '`FOO=c`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 1);
|
||||
assert.equal(r.duplicates[0].count, 3);
|
||||
});
|
||||
|
||||
test('reportsNoDuplicateForSingleOccurrence', () => {
|
||||
const r = parsePredicates('`FOO=a`');
|
||||
assert.equal(r.duplicates.length, 0);
|
||||
});
|
||||
|
||||
test('doesNotCountFenceSkippedOccurrenceAsDuplicate', () => {
|
||||
const md = ['`FOO=a`', '```', '`FOO=b`', '```'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 0, 'a fence-skipped occurrence is not a declaration');
|
||||
assert.equal(r.predicates.length, 1);
|
||||
});
|
||||
|
||||
test('doesNotCountCommentedOccurrenceAsDuplicate', () => {
|
||||
const md = ['`FOO=a`', '<!-- `FOO=b` -->'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 0);
|
||||
assert.equal(r.predicates.length, 1);
|
||||
});
|
||||
|
||||
test('reportsMalformedAndDuplicateDiagnosticsTogether', () => {
|
||||
const md = ['`FOO=a`', '`FOO=b`', '`BAR=`'].join('\n');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 1, 'the duplicate path must not be suppressed by the malformed path');
|
||||
assert.equal(r.malformed.length, 1, 'the malformed path must not be suppressed by the duplicate path');
|
||||
});
|
||||
|
||||
test('realContextMdHasNoDuplicateIdentifiers', () => {
|
||||
const md = fs.readFileSync(path.join(ROOT, 'CONTEXT.md'), 'utf8');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 0, 'the real CONTEXT.md must carry no duplicate predicate ids');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── I. Independence + real-corpus regression ──────────────────────────────
|
||||
|
||||
describe('parsePredicates: independence + real-corpus regression (I)', () => {
|
||||
test('parserHasNoCrossTestSharedState', () => {
|
||||
// Run in an order that would surface a leaking module-level cache: parse
|
||||
// a document with a duplicate, then a clean document, then re-parse the
|
||||
// first — results must be identical each time, regardless of call order.
|
||||
const dupMd = ['`FOO=a`', '`FOO=b`'].join('\n');
|
||||
const cleanMd = '`BAR=x`';
|
||||
|
||||
const firstPass = parsePredicates(dupMd);
|
||||
const cleanPass = parsePredicates(cleanMd);
|
||||
const secondPass = parsePredicates(dupMd);
|
||||
|
||||
assert.equal(cleanPass.duplicates.length, 0, 'a clean parse must never see the previous call\'s duplicate');
|
||||
assert.deepEqual(secondPass.duplicates, firstPass.duplicates, 'repeating the same input must repeat the same result');
|
||||
assert.equal(secondPass.predicates.length, firstPass.predicates.length);
|
||||
});
|
||||
|
||||
test('realContributingMdYieldsNoPredicates', () => {
|
||||
// Fixture provenance (#2371): CONTRIBUTING.md contains
|
||||
// `GITHUB_BASE_REF=next …` (inside a blockquote) and
|
||||
// `export GSD_BLOCKED_AUTHOR_REGEX='@example-corp\.com$'` (inside a
|
||||
// fenced bash block) — uppercase, `=`-bearing, backtick-adjacent text
|
||||
// authored by someone who never heard of this grammar.
|
||||
const md = fs.readFileSync(path.join(ROOT, 'CONTRIBUTING.md'), 'utf8');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.predicates.length, 0, 'a document that never heard of the grammar must yield zero predicates');
|
||||
});
|
||||
|
||||
test('realContextMdParsesFullPredicateSet', () => {
|
||||
const md = fs.readFileSync(path.join(ROOT, 'CONTEXT.md'), 'utf8');
|
||||
const r = parsePredicates(md);
|
||||
assert.equal(r.duplicates.length, 0);
|
||||
assert.ok(r.predicates.length > 0, 'the real CONTEXT.md must yield a non-empty predicate set');
|
||||
const classes = new Set(r.predicates.map((p) => p.klass));
|
||||
assert.ok(classes.size >= 20, `expected >= 20 classes in the live predicate set, got ${classes.size}`);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── selectPredicates / buildIndex smoke coverage (not in the row matrix,
|
||||
// but exercised here since they are pure, no-I/O functions covered by unit
|
||||
// tests per the risk table — the CLI/query surfaces are covered separately). ───
|
||||
|
||||
describe('selectPredicates + buildIndex: pure-function smoke coverage', () => {
|
||||
test('selectPredicates filters by klass/prefix/contains independently', () => {
|
||||
const r = parsePredicates(['`FOO.a=hello world`', '`FOO.b=other`', '`BAR.a=hello`'].join('\n'));
|
||||
const byKlass = selectPredicates(r.predicates, { klass: 'FOO' });
|
||||
assert.equal(byKlass.length, 2);
|
||||
const byPrefix = selectPredicates(r.predicates, { prefix: 'FOO.a' });
|
||||
assert.equal(byPrefix.length, 1);
|
||||
const byContains = selectPredicates(r.predicates, { contains: 'hello' });
|
||||
assert.equal(byContains.length, 2);
|
||||
});
|
||||
|
||||
test('buildIndex omits the line field from every entry (S5)', () => {
|
||||
const r = parsePredicates('`FOO=bar`');
|
||||
const index = buildIndex(r.predicates);
|
||||
assert.equal(index.predicates.length, 1);
|
||||
assert.ok(!Object.prototype.hasOwnProperty.call(index.predicates[0], 'line'));
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Parity: fenced-line determination vs markdown-sectionizer.scanFencedBlocks
|
||||
// (DEFECT.GENERATIVE-FIX) ───────────────────────────────────────────────────
|
||||
//
|
||||
// context-predicates.cts derives its fenced-line skip flags from
|
||||
// markdown-sectionizer.cts's exported `scanFencedBlocks` seam rather than
|
||||
// carrying its own copy of the fence state machine. This suite asserts that
|
||||
// parsePredicates' observable skip/keep decision for every ID-marker line
|
||||
// agrees with what `scanFencedBlocks` independently reports for that same
|
||||
// `lines` array, across the fence-shape table below — so a future change to
|
||||
// the shared scanner cannot silently diverge from predicate parsing.
|
||||
|
||||
describe('parsePredicates: fence-skip parity with markdown-sectionizer.scanFencedBlocks', () => {
|
||||
const BARE_ID_RE = /^`([A-Za-z][A-Za-z0-9._-]*)=/;
|
||||
|
||||
function assertFenceParity(name, lines) {
|
||||
test(name, () => {
|
||||
const md = lines.join('\n');
|
||||
const blocks = scanFencedBlocks(lines);
|
||||
const expectedSkip = new Array(lines.length).fill(false);
|
||||
for (const block of blocks) {
|
||||
const end = block.closeLineIdx === -1 ? lines.length - 1 : block.closeLineIdx;
|
||||
for (let i = block.openLineIdx; i <= end; i++) expectedSkip[i] = true;
|
||||
}
|
||||
|
||||
const r = parsePredicates(md);
|
||||
const parsedIds = new Set(r.predicates.map((p) => p.id));
|
||||
|
||||
let sawMarker = false;
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const m = BARE_ID_RE.exec(lines[i].trim());
|
||||
if (!m) continue;
|
||||
sawMarker = true;
|
||||
const id = m[1];
|
||||
if (expectedSkip[i]) {
|
||||
assert.ok(
|
||||
!parsedIds.has(id),
|
||||
`${name}: line ${i} (${id}) is inside a scanFencedBlocks fence and must not be parsed as live`,
|
||||
);
|
||||
} else {
|
||||
assert.ok(
|
||||
parsedIds.has(id),
|
||||
`${name}: line ${i} (${id}) is outside any scanFencedBlocks fence and must be parsed as live`,
|
||||
);
|
||||
}
|
||||
}
|
||||
assert.ok(sawMarker, `${name}: fixture must contain at least one ID marker line`);
|
||||
});
|
||||
}
|
||||
|
||||
assertFenceParity('3-backtick fence', ['`BEFORE=1`', '```', '`INSIDE=2`', '```', '`AFTER=3`']);
|
||||
|
||||
assertFenceParity('3-tilde fence', ['`BEFORE=1`', '~~~', '`INSIDE=2`', '~~~', '`AFTER=3`']);
|
||||
|
||||
assertFenceParity('4-backtick fence containing a nested 3-backtick fence', [
|
||||
'`BEFORE=1`',
|
||||
'````',
|
||||
'```',
|
||||
'`INSIDE=2`',
|
||||
'```',
|
||||
'````',
|
||||
'`AFTER=3`',
|
||||
]);
|
||||
|
||||
assertFenceParity('language-tagged fence', [
|
||||
'`BEFORE=1`',
|
||||
'```bash',
|
||||
'`INSIDE=2`',
|
||||
'```',
|
||||
'`AFTER=3`',
|
||||
]);
|
||||
|
||||
assertFenceParity('info string containing a backtick is not a valid opener', [
|
||||
'`BEFORE=1`',
|
||||
'``` `evil` ',
|
||||
'`STILL=2`',
|
||||
'```',
|
||||
'`INSIDE=3`',
|
||||
'```',
|
||||
'`AFTER=4`',
|
||||
]);
|
||||
|
||||
assertFenceParity('indented (<=3 space) fence', [
|
||||
'`BEFORE=1`',
|
||||
' ```',
|
||||
'`INSIDE=2`',
|
||||
' ```',
|
||||
'`AFTER=3`',
|
||||
]);
|
||||
|
||||
assertFenceParity('unterminated fence skips to end of file', [
|
||||
'`BEFORE=1`',
|
||||
'```',
|
||||
'`INSIDE=2`',
|
||||
'`ALSOINSIDE=3`',
|
||||
]);
|
||||
});
|
||||
445
tests/gen-context-index.test.cjs
Normal file
445
tests/gen-context-index.test.cjs
Normal file
@@ -0,0 +1,445 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* Integration tests for scripts/gen-context-index.cjs — the CI gate that
|
||||
* keeps docs/CONTEXT-INDEX.json in sync with the predicates declared in the
|
||||
* repo-root CONTEXT.md (ADR-1671, #2928 Phase 1, rows F1-F17).
|
||||
*
|
||||
* The committed artifact is plain JSON (not a `.cjs` CommonJS module): a
|
||||
* shipped runtime module is the wrong place for ~120 KB of arbitrary
|
||||
* CONTEXT.md prose, and embedding it there tripped both
|
||||
* tests/cline-install.test.cjs (leaked `.claude/hooks/...` path literals) and
|
||||
* tests/package-name-single-source.test.cjs (hardcoded package-name
|
||||
* literals) — both true positives against runtime-code content scanning.
|
||||
* docs/CONTEXT-INDEX.json mirrors docs/INVENTORY-MANIFEST.json's precedent:
|
||||
* a committed, generated, `--check`-guarded JSON manifest that is not
|
||||
* runtime code.
|
||||
*
|
||||
* Fixture isolation (ADR-1671 Phase 1 commit 3): gen-context-index.cjs now
|
||||
* accepts `--context-path <p>` / `--index-path <p>` CLI overrides (and the
|
||||
* same-named parameters on the exported `checkReport`/`buildFreshIndex`
|
||||
* pure functions), so every test here spawns the real CLI (spawnSync, not an
|
||||
* engine-direct call — an engine-direct call is false-green for CLI behavior
|
||||
* per the design's own risk analysis) pointed directly at temp fixture
|
||||
* files, with NO fs monkeypatching. The prior `--require` preload
|
||||
* (tests/helpers/gen-context-index-fs-fixture.cjs) redirected two hardcoded
|
||||
* absolute paths by patching fs.readFileSync/existsSync/writeFileSync — that
|
||||
* indirection is no longer needed now that the paths are directly
|
||||
* injectable, and the preload has been deleted.
|
||||
*
|
||||
* Prohibited: Raw Text Matching on Test Outputs (CONTRIBUTING.md). This
|
||||
* generator's `--check --json` mode emits a typed `{ ok, reason, duplicates,
|
||||
* count, classes }` report — `reason` is always one of the frozen `REASON`
|
||||
* enum values. Rows F7-F10 assert on `report.reason === REASON.FAIL_X`
|
||||
* (and, for F7, that `report.duplicates` names the duplicate id) instead of
|
||||
* exit-code-only / stderr-substring assertions.
|
||||
*/
|
||||
|
||||
const { describe, test, beforeEach, afterEach } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { execFileSync } = require('node:child_process');
|
||||
|
||||
const { createTempDir, cleanup } = require('./helpers.cjs');
|
||||
const { serializeIndex, buildFreshIndex, checkReport, REASON } = require('../scripts/gen-context-index.cjs');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const SCRIPT = path.join(ROOT, 'scripts', 'gen-context-index.cjs');
|
||||
const REAL_CONTEXT_PATH = path.join(ROOT, 'CONTEXT.md');
|
||||
|
||||
const STACK_FRAME_RE = /\n\s+at\s+\S+\s+\(.*:\d+:\d+\)/;
|
||||
|
||||
/**
|
||||
* Spawn the real gen-context-index.cjs CLI with explicit `--context-path` /
|
||||
* `--index-path` overrides — no fs monkeypatching, no `--require` preload.
|
||||
*
|
||||
* @param {string[]} args - CLI args (e.g. ['--check', '--json']).
|
||||
* @param {{contextPath?: string, indexPath?: string}} [paths] - absolute
|
||||
* fixture paths to pass via `--context-path`/`--index-path`. Omit a key to
|
||||
* leave that seam at its real-repo default (read-only, untouched).
|
||||
* @returns {{code: number, stdout: string, stderr: string}}
|
||||
*/
|
||||
function runGenContextIndex(args, paths = {}) {
|
||||
const fullArgs = [...args];
|
||||
if (paths.contextPath !== undefined) fullArgs.push('--context-path', paths.contextPath);
|
||||
if (paths.indexPath !== undefined) fullArgs.push('--index-path', paths.indexPath);
|
||||
|
||||
try {
|
||||
const stdout = execFileSync(process.execPath, [SCRIPT, ...fullArgs], {
|
||||
cwd: ROOT,
|
||||
encoding: 'utf8',
|
||||
stdio: ['pipe', 'pipe', 'pipe'],
|
||||
timeout: 30000,
|
||||
});
|
||||
return { code: 0, stdout, stderr: '' };
|
||||
} catch (err) {
|
||||
return {
|
||||
code: err.status ?? 1,
|
||||
stdout: err.stdout ? err.stdout.toString() : '',
|
||||
stderr: err.stderr ? err.stderr.toString() : '',
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse the single JSON line `--check --json` writes to stdout.
|
||||
*
|
||||
* @param {string} stdout
|
||||
* @returns {object}
|
||||
*/
|
||||
function parseJsonReport(stdout) {
|
||||
return JSON.parse(stdout.trim());
|
||||
}
|
||||
|
||||
describe('gen-context-index.cjs REASON enum (three-coordinated-changes lock)', () => {
|
||||
test('REASON key set is exactly the documented set', () => {
|
||||
// Locks the documented enum shape (CONTRIBUTING.md three-coordinated-
|
||||
// changes pattern): adding a reason requires updating this assertion
|
||||
// too, so the typed surface cannot silently drift from what tests expect.
|
||||
assert.deepEqual(Object.keys(REASON).sort(), [
|
||||
'FAIL_CONTEXT_MISSING',
|
||||
'FAIL_CONTEXT_UNREADABLE',
|
||||
'FAIL_DUPLICATE_IDS',
|
||||
'FAIL_INDEX_MISSING',
|
||||
'FAIL_INDEX_UNPARSEABLE',
|
||||
'FAIL_LIB_NOT_BUILT',
|
||||
'FAIL_STALE',
|
||||
'OK_UP_TO_DATE',
|
||||
]);
|
||||
});
|
||||
|
||||
test('REASON is frozen', () => {
|
||||
assert.ok(Object.isFrozen(REASON));
|
||||
});
|
||||
});
|
||||
|
||||
describe('gen-context-index.cjs --check (F)', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempDir('gen-context-index-');
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
test('checkExitsZeroWhenIndexIsFresh', () => {
|
||||
// Read-only against the real, already-fresh repo state — no override
|
||||
// needed, and nothing is mutated.
|
||||
const r = runGenContextIndex(['--check']);
|
||||
assert.equal(r.code, 0);
|
||||
});
|
||||
|
||||
test('checkExitsZeroAfterPureLineShift', () => {
|
||||
// S5: the committed artifact carries no `line` field, so a pure line
|
||||
// shift in CONTEXT.md must not perturb the byte-identical serialization.
|
||||
const real = fs.readFileSync(REAL_CONTEXT_PATH, 'utf8');
|
||||
const lines = real.split(/\r?\n/);
|
||||
const shifted = [lines[0], '', '', ...lines.slice(1)].join('\n');
|
||||
const shiftedPath = path.join(tmpDir, 'CONTEXT-shifted.md');
|
||||
fs.writeFileSync(shiftedPath, shifted, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check'], { contextPath: shiftedPath });
|
||||
assert.equal(r.code, 0, 'a pure line shift must not fail the gate (Q4 resolution, S5)');
|
||||
});
|
||||
|
||||
test('checkExitsOneWhenPredicateValueChanged', () => {
|
||||
const real = fs.readFileSync(REAL_CONTEXT_PATH, 'utf8');
|
||||
const modified = real.replace(
|
||||
/`RULESET\.PR-SCOPE\.one-concern-per-pr=[^`]*`/,
|
||||
'`RULESET.PR-SCOPE.one-concern-per-pr=CHANGED VALUE FOR TEST`',
|
||||
);
|
||||
assert.notEqual(modified, real, 'fixture setup sanity: the substitution must actually apply');
|
||||
const modifiedPath = path.join(tmpDir, 'CONTEXT-value-changed.md');
|
||||
fs.writeFileSync(modifiedPath, modified, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check'], { contextPath: modifiedPath });
|
||||
assert.equal(r.code, 1);
|
||||
});
|
||||
|
||||
test('checkExitsOneWhenPredicateAdded', () => {
|
||||
const real = fs.readFileSync(REAL_CONTEXT_PATH, 'utf8');
|
||||
const added = real + '\n`ZZZTEST.added-by-test=value`\n';
|
||||
const addedPath = path.join(tmpDir, 'CONTEXT-added.md');
|
||||
fs.writeFileSync(addedPath, added, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check'], { contextPath: addedPath });
|
||||
assert.equal(r.code, 1);
|
||||
});
|
||||
|
||||
test('checkExitsOneWhenPredicateRemoved', () => {
|
||||
const real = fs.readFileSync(REAL_CONTEXT_PATH, 'utf8');
|
||||
const removed = real.replace(/`RULESET\.PR-SCOPE\.one-concern-per-pr=[^`]*`\r?\n/, '');
|
||||
assert.notEqual(removed, real, 'fixture setup sanity: the removal must actually apply');
|
||||
const removedPath = path.join(tmpDir, 'CONTEXT-removed.md');
|
||||
fs.writeFileSync(removedPath, removed, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check'], { contextPath: removedPath });
|
||||
assert.equal(r.code, 1);
|
||||
});
|
||||
|
||||
test('checkExitsOneWhenClassSetChanged', () => {
|
||||
const real = fs.readFileSync(REAL_CONTEXT_PATH, 'utf8');
|
||||
const classGained = real + '\n`BRANDNEWCLASSFORTEST.x=y`\n';
|
||||
const classGainedPath = path.join(tmpDir, 'CONTEXT-class-gained.md');
|
||||
fs.writeFileSync(classGainedPath, classGained, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check'], { contextPath: classGainedPath });
|
||||
assert.equal(r.code, 1);
|
||||
});
|
||||
|
||||
test('checkExitsOneAndNamesDuplicateIdentifier (F7)', () => {
|
||||
const real = fs.readFileSync(REAL_CONTEXT_PATH, 'utf8');
|
||||
const dupPath = path.join(tmpDir, 'CONTEXT-dup.md');
|
||||
const dupIntroduced = real + '\n`RULESET.PR-SCOPE.one-concern-per-pr=duplicate copy for test`\n';
|
||||
fs.writeFileSync(dupPath, dupIntroduced, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check', '--json'], { contextPath: dupPath });
|
||||
assert.equal(r.code, 1);
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE, 'no bare stack trace in non-debug failure output');
|
||||
|
||||
const report = parseJsonReport(r.stdout);
|
||||
assert.equal(report.ok, false);
|
||||
assert.equal(report.reason, REASON.FAIL_DUPLICATE_IDS);
|
||||
assert.ok(
|
||||
report.duplicates.some((d) => d.id === 'RULESET.PR-SCOPE.one-concern-per-pr'),
|
||||
'report.duplicates must name the duplicate id',
|
||||
);
|
||||
});
|
||||
|
||||
test('checkExitsOneWithRemedyWhenIndexMissing (F8)', () => {
|
||||
const missingIndexPath = path.join(tmpDir, 'does-not-exist.json');
|
||||
const r = runGenContextIndex(['--check', '--json'], { indexPath: missingIndexPath });
|
||||
assert.equal(r.code, 1);
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE);
|
||||
|
||||
const report = parseJsonReport(r.stdout);
|
||||
assert.equal(report.ok, false);
|
||||
assert.equal(report.reason, REASON.FAIL_INDEX_MISSING);
|
||||
});
|
||||
|
||||
test('checkExitsOneWithNamedReasonWhenIndexCorrupt (F9)', () => {
|
||||
const corruptIndexPath = path.join(tmpDir, 'corrupt-index.json');
|
||||
fs.writeFileSync(corruptIndexPath, 'this is not { valid javascript', 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check', '--json'], { indexPath: corruptIndexPath });
|
||||
assert.equal(r.code, 1);
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE, 'no bare stack trace for a corrupt committed index');
|
||||
|
||||
const report = parseJsonReport(r.stdout);
|
||||
assert.equal(report.ok, false);
|
||||
assert.equal(report.reason, REASON.FAIL_INDEX_UNPARSEABLE);
|
||||
});
|
||||
|
||||
test('checkExitsOneWhenContextMdMissing (F10)', () => {
|
||||
const missingContextPath = path.join(tmpDir, 'does-not-exist.md');
|
||||
const r = runGenContextIndex(['--check', '--json'], { contextPath: missingContextPath });
|
||||
assert.equal(r.code, 1);
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE, 'no bare stack trace when CONTEXT.md is missing');
|
||||
|
||||
const report = parseJsonReport(r.stdout);
|
||||
assert.equal(report.ok, false);
|
||||
assert.equal(report.reason, REASON.FAIL_CONTEXT_MISSING);
|
||||
});
|
||||
|
||||
test('checkExitsOneWhenContextMdUnreadable', () => {
|
||||
// Fault injection via the mandated technique (CONTRIBUTING.md /
|
||||
// CLAUDE.md cross-platform IO-failure rule): monkeypatch fs.readFileSync
|
||||
// to throw an injected EACCES for one specific fixture path, restore in
|
||||
// `finally`. Never chmod 0o000 (root bypasses mode bits). This is an
|
||||
// in-process call to the exported `checkReport` pure function rather
|
||||
// than a subprocess spawn — a subprocess's fs cannot be monkeypatched
|
||||
// from the parent test process without a `--require` preload, and
|
||||
// `checkReport` IS the typed surface under test here, so calling it
|
||||
// directly is not an engine-direct false-green for CLI *argv* behavior
|
||||
// (that risk is covered by the spawned-CLI tests above); it is the
|
||||
// correct level to exercise a fault the CLI itself cannot inject.
|
||||
const fixtureContextPath = path.join(tmpDir, 'unreadable-context.md');
|
||||
fs.writeFileSync(fixtureContextPath, '`FOO=bar`\n', 'utf8');
|
||||
|
||||
const origReadFileSync = fs.readFileSync;
|
||||
fs.readFileSync = function patchedReadFileSync(p, ...rest) {
|
||||
if (p === fixtureContextPath) {
|
||||
const err = new Error(`EACCES: permission denied, open '${p}' (injected by test, never a real fs fault)`);
|
||||
err.code = 'EACCES';
|
||||
throw err;
|
||||
}
|
||||
return origReadFileSync.call(fs, p, ...rest);
|
||||
};
|
||||
try {
|
||||
const report = checkReport(fixtureContextPath, path.join(tmpDir, 'unused-index.json'));
|
||||
assert.equal(report.ok, false);
|
||||
assert.equal(report.reason, REASON.FAIL_CONTEXT_UNREADABLE);
|
||||
} finally {
|
||||
fs.readFileSync = origReadFileSync;
|
||||
}
|
||||
});
|
||||
|
||||
test('checkIsCrlfAgnostic (F17)', () => {
|
||||
// F17: a CRLF-committed index compared against the (LF) fresh real
|
||||
// CONTEXT.md must still exit 0 — comparison is CRLF-normalized.
|
||||
const freshSerialized = serializeIndex(buildFreshIndex());
|
||||
const crlfSerialized = freshSerialized.replace(/\n/g, '\r\n');
|
||||
const crlfIndexPath = path.join(tmpDir, 'context-index-crlf.json');
|
||||
fs.writeFileSync(crlfIndexPath, crlfSerialized, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check'], { indexPath: crlfIndexPath });
|
||||
assert.equal(r.code, 0, 'CRLF-vs-LF committed/fresh comparison must be normalized, not a false failure');
|
||||
});
|
||||
});
|
||||
|
||||
describe('gen-context-index.cjs --write / default / usage (F)', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempDir('gen-context-index-');
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
test('writeThenCheckIsClean', () => {
|
||||
const writeTarget = path.join(tmpDir, 'write-target.json');
|
||||
const w = runGenContextIndex(['--write'], { indexPath: writeTarget });
|
||||
assert.equal(w.code, 0);
|
||||
assert.ok(fs.existsSync(writeTarget), '--write must create the fixture-redirected index file');
|
||||
|
||||
const c = runGenContextIndex(['--check'], { indexPath: writeTarget });
|
||||
assert.equal(c.code, 0, '--check must be clean immediately after --write');
|
||||
});
|
||||
|
||||
test('writeIsByteIdenticalAcrossRuns', () => {
|
||||
const target1 = path.join(tmpDir, 'w1.json');
|
||||
const target2 = path.join(tmpDir, 'w2.json');
|
||||
assert.equal(runGenContextIndex(['--write'], { indexPath: target1 }).code, 0);
|
||||
assert.equal(runGenContextIndex(['--write'], { indexPath: target2 }).code, 0);
|
||||
|
||||
const content1 = fs.readFileSync(target1, 'utf8');
|
||||
const content2 = fs.readFileSync(target2, 'utf8');
|
||||
assert.equal(content1, content2, '--write must be deterministic across independent runs');
|
||||
});
|
||||
|
||||
test('writtenIndexContainsNoLineField', () => {
|
||||
const target = path.join(tmpDir, 'no-line-field.json');
|
||||
assert.equal(runGenContextIndex(['--write'], { indexPath: target }).code, 0);
|
||||
const content = fs.readFileSync(target, 'utf8');
|
||||
assert.equal(content.includes('"line"'), false, 'the committed artifact must carry no `line` field anywhere (S5)');
|
||||
});
|
||||
|
||||
test('defaultInvocationPrintsIndexToStdout', () => {
|
||||
// Fully safe against the real repo: default mode only reads CONTEXT.md
|
||||
// and the compiled predicates lib (read-only) and writes nothing.
|
||||
const r = runGenContextIndex([]);
|
||||
assert.equal(r.code, 0);
|
||||
assert.ok(r.stdout.length > 0);
|
||||
// Compare against the exact expected serialization (computed the same
|
||||
// way the CLI does, via the exported pure functions) rather than
|
||||
// hand-parsing the rendered text — avoids brittle delimiter-scanning
|
||||
// over a JSON payload that legitimately contains ';' inside string values.
|
||||
const expected = serializeIndex(buildFreshIndex()) + '\n';
|
||||
assert.equal(r.stdout, expected);
|
||||
});
|
||||
|
||||
test('unknownFlagExitsWithUsage', () => {
|
||||
// Safe against the real repo: the unknown-flag branch never reads
|
||||
// CONTEXT.md or the committed index at all.
|
||||
const r = runGenContextIndex(['--totally-bogus-flag']);
|
||||
assert.notEqual(r.code, 0);
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE, 'usage output must never be a bare stack trace');
|
||||
});
|
||||
|
||||
// ─── DEFECT.GEN-CONTEXT-INDEX-PARSEARGS-GATE-BYPASS (MAJOR review finding):
|
||||
// conflicting `--check --write` must be a hard usage error, not a silent
|
||||
// `--write` win, and a missing/flag-shaped path value must never resolve
|
||||
// to the cwd (which previously leaked a raw EISDIR stack trace). ────────
|
||||
|
||||
test('checkAndWriteTogetherIsUsageErrorNotASilentWrite (a)', () => {
|
||||
const target = path.join(tmpDir, 'should-not-be-written.json');
|
||||
const r = runGenContextIndex(['--check', '--write'], { indexPath: target });
|
||||
assert.notEqual(r.code, 0, '--check --write together must not silently exit 0 as a write');
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE, 'usage output must never be a bare stack trace');
|
||||
assert.equal(fs.existsSync(target), false, '--write must never win over --check and rewrite the index');
|
||||
});
|
||||
|
||||
test('writeAndCheckReversedOrderIsAlsoAUsageError (a)', () => {
|
||||
const target = path.join(tmpDir, 'should-also-not-be-written.json');
|
||||
const r = runGenContextIndex(['--write', '--check'], { indexPath: target });
|
||||
assert.notEqual(r.code, 0, 'conflicting mode flags must be a usage error regardless of order');
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE);
|
||||
assert.equal(fs.existsSync(target), false);
|
||||
});
|
||||
|
||||
test('missingTrailingValueForContextPathIsUsageErrorNotEisdirStackTrace (b)', () => {
|
||||
const target = path.join(tmpDir, 'should-not-be-written-2.json');
|
||||
// `--context-path` is the LAST arg: argv[i+1] is undefined, which used
|
||||
// to resolve to the cwd via `path.resolve(undefined ?? '')`.
|
||||
const r = runGenContextIndex(['--write', '--index-path', target, '--context-path']);
|
||||
assert.notEqual(r.code, 0, 'a missing --context-path value must be a usage error');
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE, 'must never leak a raw EISDIR (or any) stack trace');
|
||||
assert.equal(fs.existsSync(target), false, 'no write must happen when the path argument is rejected');
|
||||
});
|
||||
|
||||
test('flagShapedValueForIndexPathIsUsageErrorNotSwallowedAsALiteralPath (b)', () => {
|
||||
// `--index-path` is immediately followed by another flag rather than a
|
||||
// path — must be rejected, not silently swallowed as the literal path
|
||||
// "--json".
|
||||
const r = runGenContextIndex(['--write', '--index-path', '--json']);
|
||||
assert.notEqual(r.code, 0, 'a flag-shaped --index-path value must be a usage error');
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── DEFECT.GEN-CONTEXT-INDEX-DUPLICATE-GATE-UNPROVEN (MAJOR review finding):
|
||||
// `FAIL_DUPLICATE_IDS` was only ever proven against synthetic fixtures — this
|
||||
// branch hand-deleted the ONE live duplicate
|
||||
// (`RULESET.WORKFLOW_MARKDOWN.FENCES`) from the real CONTEXT.md, so the
|
||||
// committed docs/CONTEXT-INDEX.json ships `duplicates: []` and the gate has
|
||||
// never been shown to catch a REAL duplicate in the real document. This
|
||||
// suite re-inserts the exact deleted line (recovered from
|
||||
// `git show origin/next:CONTEXT.md`) into a copy of the REAL CONTEXT.md and
|
||||
// runs the real generator CLI against it. ───────────────────────────────────
|
||||
|
||||
describe('gen-context-index.cjs --check against a real-CONTEXT.md duplicate (real-data proof)', () => {
|
||||
let tmpDir;
|
||||
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempDir('gen-context-index-real-dup-');
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
cleanup(tmpDir);
|
||||
});
|
||||
|
||||
test('reinsertingTheDeletedRulesetWorkflowMarkdownFencesLineFailsWithNamedDuplicate', () => {
|
||||
// The exact line this branch deleted from the real CONTEXT.md (does NOT
|
||||
// mention MD040 — the live line that replaced it does).
|
||||
const deletedLine =
|
||||
'`RULESET.WORKFLOW_MARKDOWN.FENCES=when editing shell snippets inside workflow markdown, preserve the opening language fence; malformed fence can create fresh CodeRabbit threads`';
|
||||
|
||||
const real = fs.readFileSync(REAL_CONTEXT_PATH, 'utf8');
|
||||
assert.ok(
|
||||
real.includes('RULESET.WORKFLOW_MARKDOWN.FENCES'),
|
||||
'sanity: the real CONTEXT.md must still carry the live (MD040) FENCES line',
|
||||
);
|
||||
assert.ok(!real.includes(deletedLine), 'sanity: the deleted line must not already be present verbatim');
|
||||
|
||||
const reinserted = real + '\n' + deletedLine + '\n';
|
||||
const fixturePath = path.join(tmpDir, 'CONTEXT-real-with-reinserted-duplicate.md');
|
||||
fs.writeFileSync(fixturePath, reinserted, 'utf8');
|
||||
|
||||
const r = runGenContextIndex(['--check', '--json'], { contextPath: fixturePath });
|
||||
assert.equal(r.code, 1, 'a real duplicate reintroduced into the real CONTEXT.md must fail the gate');
|
||||
assert.doesNotMatch(r.stderr, STACK_FRAME_RE);
|
||||
|
||||
const report = parseJsonReport(r.stdout);
|
||||
assert.equal(report.ok, false);
|
||||
assert.equal(report.reason, REASON.FAIL_DUPLICATE_IDS);
|
||||
assert.ok(
|
||||
report.duplicates.some((d) => d.id === 'RULESET.WORKFLOW_MARKDOWN.FENCES'),
|
||||
'report.duplicates must name RULESET.WORKFLOW_MARKDOWN.FENCES as the real duplicate',
|
||||
);
|
||||
});
|
||||
});
|
||||
@@ -33,6 +33,7 @@ const fc = require('./helpers/fast-check-setup.cjs');
|
||||
const {
|
||||
stripFencedCode,
|
||||
extractFencedBlock,
|
||||
scanFencedBlocks,
|
||||
tokenizeHeadings,
|
||||
collectSections,
|
||||
collectSection,
|
||||
@@ -350,6 +351,160 @@ describe('extractFencedBlock', () => {
|
||||
});
|
||||
});
|
||||
|
||||
// ─── scanFencedBlocks (public export) ─────────────────────────────────────────
|
||||
// #2928: `export` was newly added to this function and its `FencedBlockRecord`
|
||||
// interface, making it public API for the first time (context-predicates.cts
|
||||
// consumes it directly). Hyrum's Law: a newly-public contract needs its own
|
||||
// lock-in test, independent of the consumer that motivated exporting it.
|
||||
|
||||
describe('scanFencedBlocks (public export)', () => {
|
||||
test('is actually exported from the compiled module as a function', () => {
|
||||
assert.equal(typeof scanFencedBlocks, 'function');
|
||||
});
|
||||
|
||||
test('returns {char, len, infoString, openLineIdx, closeLineIdx} with 0-based line indices', () => {
|
||||
const lines = [
|
||||
'before',
|
||||
'```js',
|
||||
'const x = 1;',
|
||||
'```',
|
||||
'after',
|
||||
];
|
||||
const blocks = scanFencedBlocks(lines);
|
||||
assert.equal(blocks.length, 1);
|
||||
assert.deepEqual(blocks[0], {
|
||||
char: '`',
|
||||
len: 3,
|
||||
infoString: 'js',
|
||||
openLineIdx: 1,
|
||||
closeLineIdx: 3,
|
||||
});
|
||||
});
|
||||
|
||||
test('closeLineIdx is -1 for an unterminated fence (EOF while open)', () => {
|
||||
const lines = [
|
||||
'before',
|
||||
'```',
|
||||
'body, never closed',
|
||||
];
|
||||
const blocks = scanFencedBlocks(lines);
|
||||
assert.equal(blocks.length, 1);
|
||||
assert.equal(blocks[0].openLineIdx, 1);
|
||||
assert.equal(blocks[0].closeLineIdx, -1);
|
||||
});
|
||||
|
||||
test('requires >=3 backticks or >=3 tildes to open a fence', () => {
|
||||
assert.deepEqual(scanFencedBlocks(['``', 'not a fence']), []);
|
||||
assert.deepEqual(scanFencedBlocks(['~~', 'not a fence']), []);
|
||||
const backtickBlocks = scanFencedBlocks(['```', 'x', '```']);
|
||||
assert.equal(backtickBlocks.length, 1);
|
||||
assert.equal(backtickBlocks[0].char, '`');
|
||||
const tildeBlocks = scanFencedBlocks(['~~~', 'x', '~~~']);
|
||||
assert.equal(tildeBlocks.length, 1);
|
||||
assert.equal(tildeBlocks[0].char, '~');
|
||||
});
|
||||
|
||||
test('tolerates up to 3 spaces of indent on the opening delimiter', () => {
|
||||
const blocks = scanFencedBlocks([' ```', 'body', '```']);
|
||||
assert.equal(blocks.length, 1);
|
||||
assert.equal(blocks[0].openLineIdx, 0);
|
||||
assert.equal(blocks[0].closeLineIdx, 2);
|
||||
});
|
||||
|
||||
test('4+ spaces of indent is not recognised as a fence delimiter', () => {
|
||||
const blocks = scanFencedBlocks([' ```', 'still not a fence']);
|
||||
assert.deepEqual(blocks, []);
|
||||
});
|
||||
|
||||
test('a closer must be the same delimiter char with run length >= the opener, and no trailing text', () => {
|
||||
// Same char, longer run: valid closer.
|
||||
const longerCloser = scanFencedBlocks(['```', 'body', '`````']);
|
||||
assert.equal(longerCloser.length, 1);
|
||||
assert.equal(longerCloser[0].closeLineIdx, 2);
|
||||
|
||||
// Same char, shorter run: not a valid closer -> content, fence stays open (EOF -> -1).
|
||||
const shorterCloser = scanFencedBlocks(['````', 'body', '```']);
|
||||
assert.equal(shorterCloser.length, 1);
|
||||
assert.equal(shorterCloser[0].closeLineIdx, -1);
|
||||
});
|
||||
|
||||
test('mismatched delimiter char while a fence is open is CONTENT, not a boundary — a 3-backtick line inside a 4-backtick fence does not close it', () => {
|
||||
const lines = [
|
||||
'````outer',
|
||||
'```coverage',
|
||||
'nested body',
|
||||
'```',
|
||||
'````',
|
||||
];
|
||||
const blocks = scanFencedBlocks(lines);
|
||||
assert.equal(blocks.length, 1, 'only the outer 4-backtick fence is a real block');
|
||||
assert.equal(blocks[0].char, '`');
|
||||
assert.equal(blocks[0].len, 4);
|
||||
assert.equal(blocks[0].openLineIdx, 0);
|
||||
assert.equal(blocks[0].closeLineIdx, 4);
|
||||
});
|
||||
|
||||
test('a same-char run that is too short, encountered while open, is content not a closer', () => {
|
||||
const lines = [
|
||||
'````',
|
||||
'```',
|
||||
'still inside',
|
||||
'````',
|
||||
];
|
||||
const blocks = scanFencedBlocks(lines);
|
||||
assert.equal(blocks.length, 1);
|
||||
assert.equal(blocks[0].openLineIdx, 0);
|
||||
assert.equal(blocks[0].closeLineIdx, 3);
|
||||
});
|
||||
|
||||
test('a same-char, sufficient-length run carrying trailing non-whitespace text, encountered while open, is content not a closer', () => {
|
||||
const lines = [
|
||||
'```',
|
||||
'``` still inside (has trailing text)',
|
||||
'```',
|
||||
'after',
|
||||
];
|
||||
const blocks = scanFencedBlocks(lines);
|
||||
assert.equal(blocks.length, 1);
|
||||
assert.equal(blocks[0].openLineIdx, 0);
|
||||
assert.equal(blocks[0].closeLineIdx, 2);
|
||||
});
|
||||
|
||||
test('CommonMark §4.5: a backtick fence info string must not contain a backtick — such a line is not a valid opener', () => {
|
||||
const blocks = scanFencedBlocks(['``` has ` a backtick', 'more text']);
|
||||
assert.deepEqual(blocks, [], 'a backtick in the info string means the line is not a valid opener at all');
|
||||
});
|
||||
|
||||
test('infoString is the trimmed trailing text of the opener', () => {
|
||||
const blocks = scanFencedBlocks(['``` js and stuff ', 'body', '```']);
|
||||
assert.equal(blocks.length, 1);
|
||||
assert.equal(blocks[0].infoString, 'js and stuff');
|
||||
});
|
||||
|
||||
test('multiple sequential blocks in one document are all returned, in order', () => {
|
||||
const lines = [
|
||||
'a',
|
||||
'```',
|
||||
'code1',
|
||||
'```',
|
||||
'b',
|
||||
'~~~py',
|
||||
'code2',
|
||||
'~~~',
|
||||
'c',
|
||||
];
|
||||
const blocks = scanFencedBlocks(lines);
|
||||
assert.equal(blocks.length, 2);
|
||||
assert.equal(blocks[0].char, '`');
|
||||
assert.equal(blocks[0].openLineIdx, 1);
|
||||
assert.equal(blocks[0].closeLineIdx, 3);
|
||||
assert.equal(blocks[1].char, '~');
|
||||
assert.equal(blocks[1].infoString, 'py');
|
||||
assert.equal(blocks[1].openLineIdx, 5);
|
||||
assert.equal(blocks[1].closeLineIdx, 7);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── tokenizeHeadings ─────────────────────────────────────────────────────────
|
||||
|
||||
describe('tokenizeHeadings', () => {
|
||||
|
||||
Reference in New Issue
Block a user