Phase 0 of the Dynamic Context Management epic (#1671): the ADR (documentation) plus a non-shipping reference example under examples/dynamic-context-management/ — the Option-E predicate fact-store parser/selector, a self-contained --check/--write/--select index generator, a sample index, and a runnable demo. The example is intentionally excluded from the build (src/->bin/lib/), the npm package files[], the installer, and the CI test suite (tests/) — nothing here is compiled into or installed with GSD. Production implementation lands in a later phase. Resolves #1672 Refs #1671 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
125
docs/adr/1671-dynamic-context-management-platform.md
Normal file
125
docs/adr/1671-dynamic-context-management-platform.md
Normal file
@@ -0,0 +1,125 @@
|
|||||||
|
# ADR-1671: Dynamic context management platform
|
||||||
|
|
||||||
|
- **Status:** Proposed
|
||||||
|
- **Date:** 2026-06-24
|
||||||
|
- **Extends:** ADR-0002 (Command Contract Validation Module), ADR-457 (build-at-publish generation model for `bin/lib/*.cjs`)
|
||||||
|
- **Relates:** ADR-857 §7 (Connected-Capability / MCP contract — kept deferred by this ADR)
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
GSD ships command and workflow content as large, hand-edited Markdown files. Two structural problems compound:
|
||||||
|
|
||||||
|
1. **Authoring is monolithic.** A single workflow body carries every branch inline. `gsd-core/workflows/plan-phase.md` is 93,973 bytes / 1,770 lines; `execute-phase.md` is 93,426 bytes. Mutually-exclusive paths (`--prd`, `--ingest`, `--mvp`, `--reviews`) all live in the same file, so a runtime loads guidance for branches a given invocation will never take.
|
||||||
|
|
||||||
|
2. **One payload ships to every runtime.** Install copies the whole `gsd-core/` tree (3.4 MB, 89 workflows, 1.7 MB) **byte-identical to all 15 runtimes** via `copyWithPathReplacement` (`bin/install.js`). The only per-runtime work is string rewrites and description truncation. There is **no per-runtime trimming or splitting**.
|
||||||
|
|
||||||
|
The result is constant pressure against size caps, enforced today only against *source* files (not emitted output) by a two-part guard (issue #1074): a per-file baseline ratchet plus per-tier hard caps (workflows XL 96 KiB / LARGE 60 KiB / DEFAULT 40 KiB; agents XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB). Several files have almost no headroom — `agents/gsd-verifier.md` has **293 bytes**. The one true emission-time cap, Windsurf's 12,000-byte limit (`src/runtime-artifact-conversion.cts`), is a hard `throw` with no graceful fallback. Adding one rule to a tight file forces an extract-to-`references/` refactor (`DEFECT.AGENT-FILE-SIZE-CAP-BREACH`), turning a one-line edit into a multi-file change that ripples across stub frontmatter, the workflow body, reference fragments, and `docs/` — each guarded by a different lint.
|
||||||
|
|
||||||
|
A separate but related pain: the repo-root `CONTEXT.md` predicate fact-store (~935 lines, ~200 KB of `CLASS.subkey=value` predicates that agent briefs are required to "cite verbatim") has **no programmatic reader, validator, or selector**. Briefs are hand-assembled, and `META.RULE.brief-must-cite-doc` is enforced only socially — paraphrasing from memory has caused real violations (5/8 agents in one documented batch).
|
||||||
|
|
||||||
|
### The machinery already exists, in silos
|
||||||
|
|
||||||
|
Research into the codebase found that most JIT primitives are already present and proven; they are just single-purpose and not composed:
|
||||||
|
|
||||||
|
- **Lazy reference loading** — the init bundle. `gsd_run query init.<cmd>` (`src/init.cts`) returns JSON of *paths + flags, not contents*; the model reads only the files it needs ("paths only to minimize orchestrator context", `plan-phase.md:66`). This is Anthropic's recommended "lightweight identifiers over payloads" pattern, in production.
|
||||||
|
- **Progressive disclosure** — `gsd-core/workflows/help.md` reads only the one mode file matching the argument (`brief` 0.9 KB / `default` 1.9 KB / `full` 34 KB).
|
||||||
|
- **Token-budgeted assembly** — `src/prompt-budget.cts` `applyBudget()` already does priority-ordered, budget-trimmed composition with an omission note — but it is walled into the cross-AI review pipeline only.
|
||||||
|
- **Pointer-passing channel** — `src/io.cts` spills any payload > 50 KB to a tmpfile and returns `@file:<path>`.
|
||||||
|
- **A codegen factory + drift-guard harness** — 13 generators share one `--check`/`--write` idiom (derive fresh, diff committed, exit 1 on drift). `scripts/gen-plugin-skills.cjs` already generates 69 shipped `SKILL.md` files from `commands/gsd/*.md`.
|
||||||
|
- **A reusable structured-markdown parser** — `src/markdown-sectionizer.cts`, already powering the per-phase `<decisions>` fact-store reader (`src/decisions.cts`).
|
||||||
|
|
||||||
|
### External practice
|
||||||
|
|
||||||
|
The closest external analogs are Anthropic Agent Skills' three-tier progressive disclosure (metadata → `SKILL.md` → bundled references), MCP resources/prompts/deferred-tools (list-then-fetch JIT), and priority/token-budget prompt renderers (Priompt, VS Code `@vscode/prompt-tsx`) that include the highest-priority fragments that fit a budget via a binary-search cutoff, with `flexReserve` floors for load-bearing content and `<isolate>` for a stable cacheable prefix. The portability catch is real and load-bearing: only the Skills *format* (directory + `SKILL.md` + frontmatter) is an open standard; native lazy loading is Claude-specific, and GSD's 15 runtimes do not all support skills or MCP (cf. surface-mismatch bugs #1614 antigravity, #1615 windsurf).
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
Adopt a **dynamic context management platform** built on a hybrid of build-time and run-time assembly, reusing the existing seams rather than inventing new infrastructure:
|
||||||
|
|
||||||
|
1. **Fragment store (authoring model).** Author workflow content as composable, priority-tagged fragments (workflow sections + shared `references/` + predicate-derived blocks), each carrying an applicability condition (which flags / capabilities / runtimes require it). This is the net-new authoring discipline.
|
||||||
|
|
||||||
|
2. **Build-time composer + per-runtime budget emission (the universal floor).** Generalize `prompt-budget.cts` out of the review silo into a shared `context-composer` seam (`src/*.cts` → `build:lib` → `bin/lib/*.cjs`). At build/install time, for each command × runtime, the composer selects the needed fragments and trims by priority to fit that runtime's measured cap (`scripts/workflow-size.cjs` `lfByteCount`), emitting a right-sized artifact through the existing converter. Caps move from *source* to *emitted output*; the Windsurf 12 KB `throw` becomes a graceful auto-trim/auto-extract. This is what makes caps stop biting on non-lazy runtimes, and it requires no runtime feature — so it is the universal floor.
|
||||||
|
|
||||||
|
3. **Progressive disclosure where the host supports it.** On lazy-loading hosts (Claude Code and the Agent SDK), keep the stub + `@-ref` model and let the init bundle name exactly which files to read; the body and references load on demand.
|
||||||
|
|
||||||
|
4. **Run-time selection via the init seam (per-request precision).** Extend the init bundle / `command-routing-hub` dispatch (`src/command-routing-hub.cts`) to emit a typed manifest of which sections / references / predicates a *specific* invocation needs (given parsed args, flags, phase state, active capabilities), reusing the `@file:` spill channel for assembled fragments. This is layered on top of the fragment store.
|
||||||
|
|
||||||
|
5. **Formalize the `CONTEXT.md` predicate fact-store → JIT selector.** Give the predicate grammar a parser (on `markdown-sectionizer`), an ID-uniqueness validator, a `--check`/`--write` drift-guard, and a `task → relevant predicate set` selector. This converts hand-assembled briefs into JIT-generated context and attacks the maintainer-side "edit a 200 KB file by hand" pain directly. **This is sequenced first** (see Prototype) because it is the smallest, lowest-risk piece that proves the whole pattern.
|
||||||
|
|
||||||
|
6. **Defer MCP (Connected-Capability).** Per ADR-857 §7 / #956, a served MCP catalog (resources/prompts/deferred-tools) remains an additive future enhancement for MCP-capable runtimes — never a replacement for the file-copy floor. Not in scope here.
|
||||||
|
|
||||||
|
### Options considered
|
||||||
|
|
||||||
|
| Option | Summary | Fixes caps? | Runtime compat | Decision |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| A. Progressive-disclosure authoring | Metadata-first files + one-level references; lean on host lazy-load | Partial; needs host lazy-load | Authoring universal; native JIT Claude-first | Adopt as a layer |
|
||||||
|
| B. Build-time composer + per-runtime budget emission | Composer trims fragments to each runtime cap, emits right-sized files | Yes — measured before write | Universal floor | **Adopt as core** |
|
||||||
|
| C. Run-time selection via init seam | Init bundle names which slices this invocation needs | Reduces per-invocation context | Broad (the `gsd_run` shim is universal) | Adopt after B |
|
||||||
|
| D. MCP served catalog | Serve content as resources/prompts/deferred-tools | For MCP hosts only | Partial; needs 2nd channel | Defer (ADR-857 §7) |
|
||||||
|
| E. Predicate fact-store → JIT selector | Parse/validate/select `CONTEXT.md` predicates | Maintainer-side big-file pain | N/A (build + orchestrator) | **Adopt first** |
|
||||||
|
|
||||||
|
Pure Agent Skills (A alone) and pure MCP (D alone) were rejected as the foundation because both are runtime-partial; only build-time emission (B) relieves caps on every runtime.
|
||||||
|
|
||||||
|
## Architecture and contracts
|
||||||
|
|
||||||
|
- **Fragment unit (open question, see below):** either separate files (clean lazy-load + INVENTORY rows) or in-file section markers (`<!-- gsd:section ... -->`, mirroring the existing `<!-- gsd:loop-host -->` markers consumed by `scripts/gen-loop-host-contract.cjs`).
|
||||||
|
- **Composer contract:** priority + binary-search cutoff to a per-runtime budget; `flexReserve`-style floors for load-bearing fragments (`META.RULE` citation rules, contribution gates, closing-keyword rules); a byte-stable canonical prefix (`<isolate>`) kept identical across runtimes to preserve KV-cache warmth and keep launcher-parity tests green.
|
||||||
|
- **Budget unit:** bytes for emission caps (matches `lfByteCount`, deterministic, offline-safe); a token estimate for run-time selection.
|
||||||
|
- **Determinism + drift-guard:** every generated artifact follows the universal `--check`/`--write` idiom and is committed; any constant shared between two surfaces gets a `DEFECT.GENERATIVE-FIX` parity assertion. Caps are asserted on **emitted per-runtime bytes** via real spawn-install tests (engine-direct tests are false-green for install behavior).
|
||||||
|
- **Boundary coverage:** the composer's budget logic is tested at `cap-1 / cap / cap+1` per `RULESET.TESTS.boundary-coverage`.
|
||||||
|
|
||||||
|
## Migration path
|
||||||
|
|
||||||
|
Sequenced to de-risk — prove the pattern on the smallest surface first, scale last:
|
||||||
|
|
||||||
|
1. **This ADR** establishes the platform, the fragment/composer contract, emission-time caps, and the drift-guard requirement.
|
||||||
|
2. **Prototype the predicate fact-store (Option E)** — *landed with this ADR as a non-shipping reference example* under `examples/dynamic-context-management/` (see Prototype below).
|
||||||
|
3. **Lift `prompt-budget.cts`** out of the review silo into a shared `context-composer` seam with fast-check property tests + boundary coverage.
|
||||||
|
4. **Pilot fragmentization on one XL workflow** (`plan-phase.md` or `execute-phase.md`): split into priority-tagged sections + applicability; composer emits per-runtime; prove byte-identical-or-smaller output and green `gsd-test` docker.
|
||||||
|
5. **Move caps from source to emitted output**; turn the Windsurf `throw` into graceful auto-trim; auto-regenerate size baselines on intentional edits.
|
||||||
|
6. **Wire the init bundle (C)** to emit a per-invocation sections manifest; workflows consume it.
|
||||||
|
7. **Roll out across LARGE/XL tiers**; update INVENTORY families + parity tests.
|
||||||
|
8. **(Deferred)** MCP served catalog (ADR-857 §7 / #956).
|
||||||
|
|
||||||
|
**Ordering landmine:** any generator consuming compiled output must run *after* `build:lib` (tsc), like `gen-plugin-skills` / `gen-capability-registry`; regenerating before `build:lib` silently drops unbuilt modules (`gsd-inventory-manifest-regen-needs-build`).
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
**Positive**
|
||||||
|
- Caps stop biting: each runtime's emitted artifact is measured and trimmed before write.
|
||||||
|
- A discovered fact lands in one fragment / predicate, not 4 hand-edited surfaces.
|
||||||
|
- Reuses the existing converter, drift-guard, boundary-test, and `markdown-sectionizer` infrastructure — the net-new pieces are only the fragment model and the composer.
|
||||||
|
- Opens a path to collapse the 10+ hand-written per-runtime body converters toward a data-driven spec.
|
||||||
|
|
||||||
|
**Negative / risks**
|
||||||
|
- Trimming a load-bearing fragment is a correctness hazard (history: paraphrased `META.RULE` → agent violations). Mitigate with `flexReserve` floors, a Promptfoo-style eval gate, and boundary tests.
|
||||||
|
- Per-runtime emission multiplies artifacts across the 15 × N matrix (inventory/parity surface).
|
||||||
|
- Build-order fragility (must run after `build:lib`).
|
||||||
|
- Dual-surface drift if any future MCP channel is added — requires parity assertions.
|
||||||
|
|
||||||
|
## Prototype (step 2, Option E) — non-shipping reference example
|
||||||
|
|
||||||
|
A working prototype proves the platform pattern end-to-end. It ships as a **reference example only**, under `examples/dynamic-context-management/` — deliberately outside the build (`src/` → `bin/lib/`), the npm package `files[]`, the installer, and the CI test suite (`tests/`). Nothing in it is compiled into or installed with GSD; the production implementation lands in a later phase.
|
||||||
|
|
||||||
|
- `examples/dynamic-context-management/context-predicates.cjs` — pure parser/selector: `parsePredicates(markdown)` (handles bare and list-item backtick predicate forms, splits on first `=`, skips fenced code / blockquote prose, detects duplicate IDs), `selectPredicates(predicates, {klass, prefix, contains})` (the JIT "task → predicate set" selector), and `buildIndex(predicates)` (deterministic, sorted).
|
||||||
|
- `examples/dynamic-context-management/gen-context-index.cjs` — self-contained CLI with `--check`/`--write` drift-guard plus a `--select <query>` mode demonstrating JIT brief assembly.
|
||||||
|
- `examples/dynamic-context-management/CONTEXT-INDEX.json` — sample generated index: **393 predicates, 18 classes**.
|
||||||
|
- `examples/dynamic-context-management/demo.cjs` + `README.md` — runnable usage example and notes.
|
||||||
|
|
||||||
|
During research the slice was validated with 42 behavioral tests (predicate forms, fenced-code / prose skipping, duplicate-id detection, the selector, a deterministic index, and a fast-check property test); those return as CI tests under `tests/` with the production implementation.
|
||||||
|
|
||||||
|
The prototype immediately surfaced **3 latent duplicate predicate IDs** in `CONTEXT.md` (`RULESET.WORKFLOW_MARKDOWN.FENCES`, `RULESET.GEMINI.TOOLS.ask_user`, `RULESET.GEMINI.TEST_SENTINEL`) — integrity drift no existing tool catches. Production `--check` can be made to fail on *new* duplicates once the existing three are reconciled.
|
||||||
|
|
||||||
|
Prototype scope notes: the parser is intentionally self-contained for the example; production should consume the compiled `markdown-sectionizer` seam, live under `src/` → `bin/lib/`, and be drift-guarded by a generator wired into the build **after** `build:lib`.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
1. Fragment unit: separate files vs in-file section markers?
|
||||||
|
2. Build-time emission vs run-time assembly as the primary surface during migration (double-write vs per-workflow cutover)?
|
||||||
|
3. Whether/when to invest in per-runtime native channels (skills, MCP) above the universal file floor.
|
||||||
|
|
||||||
|
## Related
|
||||||
|
|
||||||
|
- ADR-0002 — Command Contract Validation Module (the stub `<execution_context>` @-ref contract this platform's emission must keep satisfying).
|
||||||
|
- ADR-457 — build-at-publish generation model (the codegen + drift-guard precedent the composer extends).
|
||||||
|
- ADR-857 §7 — Connected-Capability / MCP contract (the deferred served-catalog channel).
|
||||||
2407
examples/dynamic-context-management/CONTEXT-INDEX.json
Normal file
2407
examples/dynamic-context-management/CONTEXT-INDEX.json
Normal file
File diff suppressed because it is too large
Load Diff
44
examples/dynamic-context-management/README.md
Normal file
44
examples/dynamic-context-management/README.md
Normal file
@@ -0,0 +1,44 @@
|
|||||||
|
# Dynamic context management — Option-E reference example
|
||||||
|
|
||||||
|
Reference example for [ADR-1671](../../docs/adr/1671-dynamic-context-management-platform.md),
|
||||||
|
"Dynamic context management platform."
|
||||||
|
|
||||||
|
> **This is a non-shipping reference example.** It lives outside the build
|
||||||
|
> (`src/` → `bin/lib/`), the npm package `files[]`, the installer, and the CI
|
||||||
|
> test suite (`tests/`). Nothing here is compiled into or installed with GSD.
|
||||||
|
> The production implementation lands in a later phase of the
|
||||||
|
> [Dynamic Context Management epic (#1671)](https://github.com/open-gsd/gsd-core/issues/1671).
|
||||||
|
|
||||||
|
## What it demonstrates
|
||||||
|
|
||||||
|
The **predicate fact-store → JIT selector** slice of the platform: parse the
|
||||||
|
repo-root `CONTEXT.md` `CLASS.subkey=value` predicates into structured records,
|
||||||
|
drift-guard a generated index, and select the relevant predicate subset for a
|
||||||
|
task — the building block for just-in-time agent-brief assembly instead of
|
||||||
|
hand-citing a 200 KB file.
|
||||||
|
|
||||||
|
## Files
|
||||||
|
|
||||||
|
- `context-predicates.cjs` — parser + selector + deterministic index builder (self-contained).
|
||||||
|
- `gen-context-index.cjs` — `--check` / `--write` drift-guarded generator + `--select`.
|
||||||
|
- `CONTEXT-INDEX.json` — sample generated output (393 predicates, 18 classes).
|
||||||
|
- `demo.cjs` — runnable usage example.
|
||||||
|
|
||||||
|
## Run (from the repo root)
|
||||||
|
|
||||||
|
```sh
|
||||||
|
node examples/dynamic-context-management/demo.cjs
|
||||||
|
node examples/dynamic-context-management/gen-context-index.cjs --select PRED.k320
|
||||||
|
node examples/dynamic-context-management/gen-context-index.cjs --check
|
||||||
|
```
|
||||||
|
|
||||||
|
## Validation
|
||||||
|
|
||||||
|
During research this slice was validated with 42 behavioral tests — predicate
|
||||||
|
forms, fenced-code / prose skipping, duplicate-id detection, the selector, a
|
||||||
|
deterministic index, and a fast-check property test. Those return as CI tests
|
||||||
|
under `tests/` when the production implementation lands.
|
||||||
|
|
||||||
|
It also surfaced 3 latent duplicate predicate IDs in `CONTEXT.md`
|
||||||
|
(`RULESET.WORKFLOW_MARKDOWN.FENCES`, `RULESET.GEMINI.TOOLS.ask_user`,
|
||||||
|
`RULESET.GEMINI.TEST_SENTINEL`), recorded in the index `duplicates` field.
|
||||||
248
examples/dynamic-context-management/context-predicates.cjs
Normal file
248
examples/dynamic-context-management/context-predicates.cjs
Normal file
@@ -0,0 +1,248 @@
|
|||||||
|
'use strict';
|
||||||
|
|
||||||
|
/**
|
||||||
|
* context-predicates.cjs — CONTEXT.md predicate fact-store parser.
|
||||||
|
*
|
||||||
|
* Self-contained CommonJS module (no dependency on build:lib output).
|
||||||
|
*
|
||||||
|
* Exports:
|
||||||
|
* parsePredicates(markdown) -> { predicates, duplicates, skippedSections }
|
||||||
|
* selectPredicates(predicates, { klass, prefix, contains }) -> filtered array
|
||||||
|
* buildIndex(predicates) -> deterministic plain object
|
||||||
|
*
|
||||||
|
* Grammar (from discovery facts):
|
||||||
|
* Two line forms, each on exactly one source line:
|
||||||
|
* 1. Bare backtick-wrapped: `ID=value`
|
||||||
|
* 2. List-item backtick: - `ID=value`
|
||||||
|
*
|
||||||
|
* ID grammar: CLASS(.subkey)* where CLASS = first dot-separated segment.
|
||||||
|
* ID chars: [A-Za-z0-9._-] (CLASS always uppercase; subkeys may be mixed).
|
||||||
|
* Split on FIRST '=' only; everything before is the ID, everything after is
|
||||||
|
* the value (up to the closing backtick).
|
||||||
|
*
|
||||||
|
* Skip:
|
||||||
|
* - Fenced code blocks (toggle on triple-backtick lines)
|
||||||
|
* - Prose lines (headings, blank lines, list items without a predicate)
|
||||||
|
* - The "PR fix discipline" section (pure prose, no predicates)
|
||||||
|
* - Session-log blockquote preamble
|
||||||
|
*/
|
||||||
|
|
||||||
|
// Regex matching the predicate ID grammar: one or more dot-separated segments.
|
||||||
|
// First segment must start with an uppercase letter (CLASS).
|
||||||
|
// Subsequent segments may start with letter/digit and include hyphens/underscores.
|
||||||
|
// We intentionally allow lowercase-starting sub-segments (e.g. PRED.k320.rule).
|
||||||
|
const ID_RE = /^([A-Z][A-Z0-9_-]*(?:\.[A-Za-z0-9_.-]+)*)=(.+)$/;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parse a single source line and return a raw {id, value} if it is a predicate,
|
||||||
|
* or null otherwise. Handles both line forms after stripping list markers.
|
||||||
|
*
|
||||||
|
* @param {string} raw - the original source line (with newline stripped)
|
||||||
|
* @returns {{ id: string, value: string } | null}
|
||||||
|
*/
|
||||||
|
function extractPredicate(raw) {
|
||||||
|
const line = raw.trimEnd();
|
||||||
|
|
||||||
|
// Form 1: `ID=value` (starts with backtick at column 0)
|
||||||
|
// Form 2: - `ID=value` (list-item with leading "- ")
|
||||||
|
// Also tolerate " - `ID=value`" (indented list item — observed in CONTEXT.md).
|
||||||
|
let inner = null;
|
||||||
|
|
||||||
|
if (line.startsWith('`') && line.endsWith('`') && line.length > 2) {
|
||||||
|
// bare backtick line
|
||||||
|
inner = line.slice(1, -1);
|
||||||
|
} else {
|
||||||
|
// strip optional leading whitespace + "- " then check for backtick wrapping
|
||||||
|
const stripped = line.replace(/^\s*-\s+/, '');
|
||||||
|
if (stripped.startsWith('`') && stripped.endsWith('`') && stripped.length > 2) {
|
||||||
|
inner = stripped.slice(1, -1);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (inner === null) return null;
|
||||||
|
|
||||||
|
// Now match the ID grammar. Split on FIRST '=' only.
|
||||||
|
const eqIdx = inner.indexOf('=');
|
||||||
|
if (eqIdx < 1) return null;
|
||||||
|
|
||||||
|
const id = inner.slice(0, eqIdx);
|
||||||
|
const value = inner.slice(eqIdx + 1);
|
||||||
|
|
||||||
|
// Validate ID — must match the grammar (no spaces, correct char set).
|
||||||
|
if (!ID_RE.test(inner)) return null;
|
||||||
|
|
||||||
|
return { id, value };
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Parse all predicates from a CONTEXT.md markdown string.
|
||||||
|
*
|
||||||
|
* @param {string} markdown
|
||||||
|
* @returns {{
|
||||||
|
* predicates: Array<{ id: string, klass: string, value: string, line: number, section: string }>,
|
||||||
|
* duplicates: Array<{ id: string, lines: number[] }>,
|
||||||
|
* skippedSections: string[]
|
||||||
|
* }}
|
||||||
|
*/
|
||||||
|
function parsePredicates(markdown) {
|
||||||
|
const lines = markdown.split('\n');
|
||||||
|
const predicates = [];
|
||||||
|
// Track id -> list of line numbers for duplicate detection
|
||||||
|
const idLines = new Map(); // id -> number[]
|
||||||
|
|
||||||
|
let inFencedCode = false;
|
||||||
|
let currentSection = '';
|
||||||
|
const allSections = [];
|
||||||
|
const seenSections = new Set();
|
||||||
|
|
||||||
|
// Section names that are known pure-prose (0 predicates) — we still scan them
|
||||||
|
// but track them as skipped if nothing is found. The parser is tolerant; it
|
||||||
|
// simply won't find predicates in prose sections.
|
||||||
|
// We do NOT hard-skip any section except fenced code — the grammar says "scan
|
||||||
|
// for backtick predicates everywhere but skip fenced code".
|
||||||
|
|
||||||
|
for (let i = 0; i < lines.length; i++) {
|
||||||
|
const raw = lines[i];
|
||||||
|
const lineNo = i + 1; // 1-based
|
||||||
|
|
||||||
|
// Track fenced code blocks (triple-backtick toggle).
|
||||||
|
// A fenced-code fence starts with ``` possibly followed by a language token.
|
||||||
|
// We use a simple heuristic: a line trimmed to /^```/ triggers the toggle.
|
||||||
|
const trimmed = raw.trimStart();
|
||||||
|
if (trimmed.startsWith('```')) {
|
||||||
|
inFencedCode = !inFencedCode;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (inFencedCode) continue;
|
||||||
|
|
||||||
|
// Track section headings for the section field.
|
||||||
|
if (raw.startsWith('#')) {
|
||||||
|
currentSection = raw.replace(/^#+\s*/, '').trim();
|
||||||
|
if (currentSection && !seenSections.has(currentSection)) {
|
||||||
|
seenSections.add(currentSection);
|
||||||
|
allSections.push(currentSection);
|
||||||
|
}
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Blockquote lines (start with ">") are prose — skip.
|
||||||
|
if (trimmed.startsWith('>')) continue;
|
||||||
|
|
||||||
|
// Attempt extraction.
|
||||||
|
const pred = extractPredicate(raw);
|
||||||
|
if (!pred) continue;
|
||||||
|
|
||||||
|
const klass = pred.id.split('.')[0];
|
||||||
|
predicates.push({
|
||||||
|
id: pred.id,
|
||||||
|
klass,
|
||||||
|
value: pred.value,
|
||||||
|
line: lineNo,
|
||||||
|
section: currentSection,
|
||||||
|
});
|
||||||
|
|
||||||
|
const existing = idLines.get(pred.id);
|
||||||
|
if (existing) {
|
||||||
|
existing.push(lineNo);
|
||||||
|
} else {
|
||||||
|
idLines.set(pred.id, [lineNo]);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Build duplicates list: ids with >1 occurrence.
|
||||||
|
const duplicates = [];
|
||||||
|
for (const [id, lns] of idLines) {
|
||||||
|
if (lns.length > 1) {
|
||||||
|
duplicates.push({ id, lines: lns });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
// Sort duplicates by id for determinism.
|
||||||
|
duplicates.sort((a, b) => a.id < b.id ? -1 : a.id > b.id ? 1 : 0);
|
||||||
|
|
||||||
|
// Skipped sections: headings that yielded zero predicates (pure prose).
|
||||||
|
const activeSections = new Set(predicates.map((p) => p.section));
|
||||||
|
const skippedSections = allSections.filter((s) => !activeSections.has(s));
|
||||||
|
|
||||||
|
return { predicates, duplicates, skippedSections };
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Select predicates by one or more optional criteria (ANDed together).
|
||||||
|
*
|
||||||
|
* @param {Array<{ id: string, klass: string, value: string, line: number, section: string }>} predicates
|
||||||
|
* @param {{ klass?: string, prefix?: string, contains?: string }} opts
|
||||||
|
* @returns {Array<{ id: string, klass: string, value: string, line: number, section: string }>}
|
||||||
|
*/
|
||||||
|
function selectPredicates(predicates, opts = {}) {
|
||||||
|
const { klass, prefix, contains } = opts;
|
||||||
|
const containsLower = contains ? contains.toLowerCase() : null;
|
||||||
|
|
||||||
|
return predicates.filter((p) => {
|
||||||
|
if (klass !== undefined && p.klass !== klass) return false;
|
||||||
|
if (prefix !== undefined && !p.id.startsWith(prefix)) return false;
|
||||||
|
if (containsLower !== null) {
|
||||||
|
const haystack = (p.id + ' ' + p.value).toLowerCase();
|
||||||
|
if (!haystack.includes(containsLower)) return false;
|
||||||
|
}
|
||||||
|
return true;
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Build a deterministic index object from a parsed predicates array.
|
||||||
|
*
|
||||||
|
* @param {Array<{ id: string, klass: string, value: string, line: number }>} predicates
|
||||||
|
* @returns {{
|
||||||
|
* schemaVersion: 1,
|
||||||
|
* count: number,
|
||||||
|
* classes: Record<string, number>,
|
||||||
|
* predicates: Array<{ id: string, klass: string, value: string, line: number }>,
|
||||||
|
* duplicates: Array<{ id: string, lines: number[] }>
|
||||||
|
* }}
|
||||||
|
*/
|
||||||
|
function buildIndex(predicates) {
|
||||||
|
// Count per class.
|
||||||
|
const classCounts = {};
|
||||||
|
for (const p of predicates) {
|
||||||
|
classCounts[p.klass] = (classCounts[p.klass] || 0) + 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Sort classes object by key for determinism.
|
||||||
|
const classes = {};
|
||||||
|
for (const k of Object.keys(classCounts).sort()) {
|
||||||
|
classes[k] = classCounts[k];
|
||||||
|
}
|
||||||
|
|
||||||
|
// Sort predicates by id then by line number.
|
||||||
|
const sortedPredicates = predicates
|
||||||
|
.map(({ id, klass, value, line }) => ({ id, klass, value, line }))
|
||||||
|
.sort((a, b) => {
|
||||||
|
if (a.id < b.id) return -1;
|
||||||
|
if (a.id > b.id) return 1;
|
||||||
|
return a.line - b.line;
|
||||||
|
});
|
||||||
|
|
||||||
|
// Rebuild duplicates from sorted predicates for determinism.
|
||||||
|
const idToLines = new Map();
|
||||||
|
for (const p of sortedPredicates) {
|
||||||
|
const arr = idToLines.get(p.id);
|
||||||
|
if (arr) arr.push(p.line);
|
||||||
|
else idToLines.set(p.id, [p.line]);
|
||||||
|
}
|
||||||
|
const duplicates = [];
|
||||||
|
for (const [id, lines] of idToLines) {
|
||||||
|
if (lines.length > 1) duplicates.push({ id, lines });
|
||||||
|
}
|
||||||
|
duplicates.sort((a, b) => a.id < b.id ? -1 : a.id > b.id ? 1 : 0);
|
||||||
|
|
||||||
|
return {
|
||||||
|
schemaVersion: 1,
|
||||||
|
count: predicates.length,
|
||||||
|
classes,
|
||||||
|
predicates: sortedPredicates,
|
||||||
|
duplicates,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
module.exports = { parsePredicates, selectPredicates, buildIndex };
|
||||||
30
examples/dynamic-context-management/demo.cjs
Normal file
30
examples/dynamic-context-management/demo.cjs
Normal file
@@ -0,0 +1,30 @@
|
|||||||
|
#!/usr/bin/env node
|
||||||
|
'use strict';
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Runnable usage example for the Option-E predicate fact-store (ADR-1671).
|
||||||
|
* Reference example only — not shipped, not part of the CI suite.
|
||||||
|
*
|
||||||
|
* node examples/dynamic-context-management/demo.cjs
|
||||||
|
*/
|
||||||
|
|
||||||
|
const fs = require('node:fs');
|
||||||
|
const path = require('node:path');
|
||||||
|
|
||||||
|
const { parsePredicates, selectPredicates } = require('./context-predicates.cjs');
|
||||||
|
|
||||||
|
const md = fs.readFileSync(path.resolve(__dirname, '..', '..', 'CONTEXT.md'), 'utf8');
|
||||||
|
const { predicates, duplicates } = parsePredicates(md);
|
||||||
|
const classes = new Set(predicates.map((p) => p.klass));
|
||||||
|
|
||||||
|
process.stdout.write(
|
||||||
|
`Parsed ${predicates.length} predicates across ${classes.size} classes; ` +
|
||||||
|
`${duplicates.length} duplicate id(s).\n`,
|
||||||
|
);
|
||||||
|
|
||||||
|
const slice = selectPredicates(predicates, { prefix: 'PRED.k320' });
|
||||||
|
process.stdout.write(
|
||||||
|
`\nselectPredicates({ prefix: 'PRED.k320' }) -> ${slice.length} matches ` +
|
||||||
|
`(a JIT brief slice):\n`,
|
||||||
|
);
|
||||||
|
for (const p of slice) process.stdout.write(` ${p.id} = ${p.value}\n`);
|
||||||
90
examples/dynamic-context-management/gen-context-index.cjs
Normal file
90
examples/dynamic-context-management/gen-context-index.cjs
Normal file
@@ -0,0 +1,90 @@
|
|||||||
|
#!/usr/bin/env node
|
||||||
|
'use strict';
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Reference example (NOT shipped, NOT compiled, NOT installed) for ADR-1671,
|
||||||
|
* "Dynamic context management platform" — the Option-E predicate fact-store.
|
||||||
|
*
|
||||||
|
* Builds a deterministic, drift-guarded index of every predicate fact in the
|
||||||
|
* repo-root CONTEXT.md, and demonstrates a JIT "task -> relevant predicates"
|
||||||
|
* selector. Self-contained: depends only on the sibling context-predicates.cjs.
|
||||||
|
*
|
||||||
|
* Usage (run from the repo root):
|
||||||
|
* node examples/dynamic-context-management/gen-context-index.cjs # print index to stdout
|
||||||
|
* node examples/dynamic-context-management/gen-context-index.cjs --write # write CONTEXT-INDEX.json (next to this file)
|
||||||
|
* node examples/dynamic-context-management/gen-context-index.cjs --check # exit 1 if the committed sample is stale
|
||||||
|
* node examples/dynamic-context-management/gen-context-index.cjs --select <query>
|
||||||
|
*
|
||||||
|
* --select <query> tries, in order: exact class ("PRED"), dotted prefix
|
||||||
|
* ("PRED.k320"), then free-text contains — the first non-empty match wins.
|
||||||
|
*/
|
||||||
|
|
||||||
|
const fs = require('node:fs');
|
||||||
|
const path = require('node:path');
|
||||||
|
|
||||||
|
const { parsePredicates, selectPredicates, buildIndex } = require('./context-predicates.cjs');
|
||||||
|
|
||||||
|
const REPO_ROOT = path.resolve(__dirname, '..', '..');
|
||||||
|
const CONTEXT_PATH = path.join(REPO_ROOT, 'CONTEXT.md');
|
||||||
|
const INDEX_PATH = path.join(__dirname, 'CONTEXT-INDEX.json');
|
||||||
|
|
||||||
|
function buildFreshIndex() {
|
||||||
|
const markdown = fs.readFileSync(CONTEXT_PATH, 'utf8');
|
||||||
|
const { predicates } = parsePredicates(markdown);
|
||||||
|
return buildIndex(predicates);
|
||||||
|
}
|
||||||
|
|
||||||
|
function main(args) {
|
||||||
|
const flag = args[0];
|
||||||
|
|
||||||
|
if (flag === '--check') {
|
||||||
|
const committed = JSON.parse(fs.readFileSync(INDEX_PATH, 'utf8'));
|
||||||
|
const live = buildFreshIndex();
|
||||||
|
if (JSON.stringify(committed, null, 2) !== JSON.stringify(live, null, 2)) {
|
||||||
|
process.stderr.write(
|
||||||
|
'CONTEXT-INDEX.json is stale. Run:\n' +
|
||||||
|
' node examples/dynamic-context-management/gen-context-index.cjs --write\n',
|
||||||
|
);
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
process.stdout.write('CONTEXT-INDEX.json is up to date.\n');
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (flag === '--write') {
|
||||||
|
const index = buildFreshIndex();
|
||||||
|
fs.writeFileSync(INDEX_PATH, JSON.stringify(index, null, 2) + '\n');
|
||||||
|
const dupNote = index.duplicates.length > 0
|
||||||
|
? ` (${index.duplicates.length} duplicate id${index.duplicates.length !== 1 ? 's' : ''})`
|
||||||
|
: '';
|
||||||
|
process.stdout.write(
|
||||||
|
`Wrote ${path.relative(REPO_ROOT, INDEX_PATH)}\n` +
|
||||||
|
` ${index.count} predicates, ${Object.keys(index.classes).length} classes${dupNote}\n`,
|
||||||
|
);
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (flag === '--select') {
|
||||||
|
const query = args[1];
|
||||||
|
if (!query) {
|
||||||
|
process.stderr.write('Usage: gen-context-index.cjs --select <query>\n');
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
const { predicates } = parsePredicates(fs.readFileSync(CONTEXT_PATH, 'utf8'));
|
||||||
|
let results = selectPredicates(predicates, { klass: query });
|
||||||
|
if (results.length === 0) results = selectPredicates(predicates, { prefix: query });
|
||||||
|
if (results.length === 0) results = selectPredicates(predicates, { contains: query });
|
||||||
|
if (results.length === 0) {
|
||||||
|
process.stdout.write(`No predicates matched: ${query}\n`);
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
for (const p of results) process.stdout.write(`${p.id} = ${p.value}\n`);
|
||||||
|
process.stdout.write(`\n(${results.length} predicate${results.length !== 1 ? 's' : ''} matched)\n`);
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
process.stdout.write(JSON.stringify(buildFreshIndex(), null, 2) + '\n');
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
process.exitCode = main(process.argv.slice(2));
|
||||||
Reference in New Issue
Block a user