* test(#2929): capture prompt-budget parity corpus pre-refactor Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared context-composer seam. Its success condition is that review-prompt output does not change, and the only authority on "did not change" is the behavior that shipped before the refactor. Capture that behavior now, while it is still the live implementation. 47 characterization cases, every `expected` value computed by executing the current implementation rather than hand-authored — the independence CONTRIBUTING.md "Fixture provenance (#2371)" asks for. A corpus is only worth what it can detect, so this one was validated by mutation rather than assumed. Five deliberate defects were injected and each must be caught by at least one case: - the note reserve deducted unconditionally instead of only under pressure - the pressure test relaxed from `>` to `>=` - a no-op head-shrink still setting the shrunk flag - the per-plan floor dropped from the proportional share - drop order reversed Two of those exposed real holes in the first cut of this corpus, and the cases that close them exist because of it: - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a floored plan group, and the 1024-char floor absorbed the entire trim, so the mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap, which makes the strict inequality observable as context kept vs omitted. - No case reached proportional-truncate at all — B6 and B7 both hard-failed the min-set pre-check first, leaving planTruncationPct at 0 across every case and the floor semantics entirely unexercised. Rebudgeted to 700 and 1100 so the min-set fits and the truncate step is actually reached; they now record 40.20% and 48.80%. The A4/A10 families sweep the pressure boundary from both sides, which is where this function has regressed before: CONTEXT.md's LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band, because the suite paired a trivially-fitting budget with a trivially-overflowing one and never sampled between them. A4 pins that nothing is trimmed from the cap down to 81 tokens under it; A10 pins that pressure fires at +1. Together with A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures. Two facts the corpus establishes that the design notes had wrong: - "" and null sections are NOT distinguished. applyBudget uses truthy checks throughout, so an empty-string section is treated as absent: not rendered, not dropped, never recorded in `omitted`. B13b pins this while the ladder is actively trimming, where only the non-empty `research` is dropped. - Sizing matters. B12/B13 were first written at a budget where both hard-failed the min-set check and returned "", so comparing them compared two empty strings and proved nothing. Committed as its own commit, ahead of the refactor, and regenerated against the pre-refactor implementation, so the oracle is demonstrably independent of the change it will adjudicate. Refs #2929 * refactor(#2929): extract the context-composer seam from prompt-budget Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer — per-runtime artifact emission — but it is walled inside the cross-AI review pipeline. Lift it into a shared seam so later phases can call it, without changing what the review pipeline emits. ADR-1671 specifies the composer as "priority + binary-search cutoff to a per-runtime budget". Read against the code it generalizes, that contract cannot express the thing being generalized. applyBudget is not a cutoff: it is a fixed five-step ladder in which each section carries its own shrink strategy, and only three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor; instructions and roadmap are never touched at all. A cutoff composer sorts by priority and discards the tail — it has no way to say "shrink this one", "truncate that one but never below 1 KB each", or "these three are the only droppables, in this order". Building to the literal contract and routing prompt-budget through it would have silently changed review-prompt output, which is the one outcome this phase forbids. So shrink strategies are the core abstraction here, and cutoff becomes one strategy among them — the right one for per-runtime emission in Phases 3-4, not for this ladder. That is an elaboration of the ADR's intent, not a departure from it, and ADR-1671 is updated to say so. Three decisions worth stating: - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan of surviving fragments and never a string. assemblePrompt's rendering is prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and owning it in the composer would force emission to adopt prompt-shaped rendering. The split is what lets one seam serve both consumers. - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires for emission caps. The existing code converts a token budget to a character budget with a hardcoded `* 4`; that assumption is now an explicit `charsPerUnit` inverse, which is precisely what a byte unit needs in order to reuse this. - The entry point is `composeWithinBudget`, not `applyBudget`. That name already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter being an unrelated graph-edge budget. A third would make every symbol search in this repo ambiguous, and it already misresolves: preflight and impact queries for "applyBudget" return graphify's. Behavior is unchanged and proven so: all 47 characterization cases reproduce byte-identically, and the corpus is mutation-validated rather than merely green (see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and from eighteen mutable accumulators to two, both inside a helper copied verbatim. estimateTokens deliberately stays in prompt-budget and keeps its exact math: src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins plan estimates and recorded actuals to that same scale, so moving or changing it would silently break the calibration loop. Refs #2929 * docs(#2929): document the context-composer seam and amend ADR-1671 Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new domain modules), and a mutation-matrix entry for the new module. The ADR amendment is the substantive part. ADR-1671 specified the composer as "priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2 established that a cutoff alone cannot express the function the platform generalizes, so the ADR now records shrink strategies as the core abstraction with cutoff as one strategy among them, reserved for per-runtime emission in Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned against that contract and would otherwise be planned against a mechanism that does not work. The mutation-matrix entry is not bookkeeping. Stryker scores per module against a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the extracted code unmeasured while prompt-budget's own score floated free of the logic it used to cover. context-composer gets its own entry at the same floor. Refs #2929 * test(#2929): pin the effectiveBudget rounding mode in the parity corpus An isolated correctness review found a real blind spot: mutating `Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator happened to produce a whole number, so floor, round and ceil all agreed and the rounding mode was entirely unpinned by a corpus whose whole job is to pin observable behavior. Three cases fix that by straddling the .5 boundary: A11 95 * 0.90 = 85.5 floor 85, round 86 -> the two disagree A12 97 * 0.90 = 87.3 floor and round agree; ceil (88) does not A13 93 * 0.85 = 79.05 same guard at a non-multiple-of-10 margin, so the margin arithmetic is exercised and not just the budget A11 alone catches the round mutation; all three catch ceil. Regenerated against the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so the expanded corpus keeps the independence property the original capture had. The corpus is now mutation-validated against seven injected defects, every one caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink setting its flag, the truncate floor ignored, drop order reversed, and both rounding-mode changes. Refs #2929 * feat(#2929): flexReserve floors and the byte-stable isolate prefix Two of issue #2929's "Done when" items were unimplemented rather than deferred, and an isolated review flagged them alongside my own audit. Both are part of ADR-1671's composer contract, so shipping the seam without them would have left Phases 3-4 building against a contract that does not exist yet. flexReserve is a per-fragment floor in measure units that every strategy must respect, which is what makes it different from the pre-existing floorChars: that one is a chars-denominated detail of proportional-truncate alone and is retained unchanged. A floored fragment is never dropped, is never head-shrunk below its floor, and raises its own proportional cap. A fragment already smaller than its floor is untouchable outright. Metadata gains `floored`, listing the ids whose floor actually prevented a trim — a guarantee no caller can observe is a guarantee no test can hold you to. isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed, never dropped, but still counted, because a prefix excluded from accounting would silently under-count real context. Metadata gains `isolatePrefix` so a caller can hash or assert on the exact bytes. Declaring an isolate fragment after a non-isolate one throws: a prefix that is not at the front is not a prefix, and accepting it would make the cross-runtime stability claim meaningless. Adds tests/context-composer.test.cjs for the exact new semantics and tests/context-composer.property.test.cjs for the five invariants, including the budget-monotonicity property the issue names explicitly. Both are registered in the mutation matrix, since coverage does not migrate with relocated code. prompt-budget uses neither feature, and its output is unchanged: all 50 corpus cases still reproduce byte-identically. Refs #2929 * chore(#2929): allowlist the prompt-budget parity suite The parity corpus needs its own test file and that makes prompt-budget a three-file module against a limit of two. The lint offers consolidation or an allowlist entry with justification; the entry is the right call here. Consolidation would mean folding the characterization suite into prompt-budget.test.cjs, which is the one thing that should not happen to it. The parity suite is a distinct concern with a distinct lifecycle: it is generated rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own scoring target, and its failure means something categorically different from a unit-test failure — not "this behavior is wrong" but "observable output moved". Burying it inside a general unit file would obscure exactly that signal. The allowlist is an identity ratchet, so this entry pins today's three exact filenames: adding a fourth still fails, and dropping back to two requires removing the entry. Refs #2929 * fix(#2929): register the new module with two gates it was missing The remote matrix caught three defects that no local check could, because the local runner is blocked in this repo and these suites had therefore never executed. Eight failures, identical on node22 and node24, so nothing environment-shaped. Two are the new-module ripple. A net-new src/*.cts lands in six places and this change had reached four of them — .gitignore, INVENTORY, the manifest, and the CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet baseline (a deliberate review-visible mirror of the matrix floors, which every COVERED module must carry). Both are now registered, the ratchet at the same floor of 66 the matrix declares. The third was a test asserting an outcome it had made impossible. It set budget:1 alongside a 400-char required fragment, so the group budget came out at -99 and the proportional-truncate step was skipped entirely — the deliberate "non-positive group budget is skipped, never clamped" rule inherited from the original ladder. Nothing was trimmed, and the test then asserted a truncation. Rebudgeted so the step actually runs, with the arithmetic written out in a comment so the next reader does not have to re-derive why 120 rather than 80. Fixing that surfaced a genuine bug in the composer. `floored` is documented as recording fragments whose flexReserve prevented a trim that would otherwise have happened, but the push sat in the else-branch of "content did not change", so it only fired when nothing was trimmed at all. A fragment truncated to a reserve-raised cap has also had a trim prevented — 40 characters' worth in the test above — and was silently absent from the field that exists to make the guarantee observable. The condition was already right; it was in the wrong branch. Now recorded on both paths: a drop prevented outright, and a truncation capped higher than the share alone would have allowed. Parity is unaffected — prompt-budget never sets flexReserve, so the branch is unreachable from every corpus path, and all 50 cases still match. Refs #2929 * chore(#2929): backfill changeset PR number (#2958) * chore(#2929): correct the corpus case count in the changeset fragment --------- Co-authored-by: sim <sim@local>
20 KiB
ADR-1671: Dynamic context management platform
- Status: Proposed
- Date: 2026-06-24
- Extends: ADR-0002 (Command Contract Validation Module), ADR-457 (build-at-publish generation model for
bin/lib/*.cjs) - Relates: ADR-857 §7 (Connected-Capability / MCP contract — kept deferred by this ADR)
Context
GSD ships command and workflow content as large, hand-edited Markdown files. Two structural problems compound:
-
Authoring is monolithic. A single workflow body carries every branch inline.
gsd-core/workflows/plan-phase.mdis 93,973 bytes / 1,770 lines;execute-phase.mdis 93,426 bytes. Mutually-exclusive paths (--prd,--ingest,--mvp,--reviews) all live in the same file, so a runtime loads guidance for branches a given invocation will never take. -
One payload ships to every runtime. Install copies the whole
gsd-core/tree (3.4 MB, 89 workflows, 1.7 MB) byte-identical to all 15 runtimes viacopyWithPathReplacement(bin/install.js). The only per-runtime work is string rewrites and description truncation. There is no per-runtime trimming or splitting.
The result is constant pressure against size caps, enforced today only against source files (not emitted output) by a two-part guard (issue #1074): a per-file baseline ratchet plus per-tier hard caps (workflows XL 96 KiB / LARGE 60 KiB / DEFAULT 40 KiB; agents XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB). Several files have almost no headroom — agents/gsd-verifier.md has 293 bytes. The one true emission-time cap, Windsurf's 12,000-byte limit (src/runtime-artifact-conversion.cts), is a hard throw with no graceful fallback. Adding one rule to a tight file forces an extract-to-references/ refactor (DEFECT.AGENT-FILE-SIZE-CAP-BREACH), turning a one-line edit into a multi-file change that ripples across stub frontmatter, the workflow body, reference fragments, and docs/ — each guarded by a different lint.
A separate but related pain: the repo-root CONTEXT.md predicate fact-store (~935 lines, ~200 KB of CLASS.subkey=value predicates that agent briefs are required to "cite verbatim") has no programmatic reader, validator, or selector. Briefs are hand-assembled, and META.RULE.brief-must-cite-doc is enforced only socially — paraphrasing from memory has caused real violations (5/8 agents in one documented batch).
The machinery already exists, in silos
Research into the codebase found that most JIT primitives are already present and proven; they are just single-purpose and not composed:
- Lazy reference loading — the init bundle.
gsd_run query init.<cmd>(src/init.cts) returns JSON of paths + flags, not contents; the model reads only the files it needs ("paths only to minimize orchestrator context",plan-phase.md:66). This is Anthropic's recommended "lightweight identifiers over payloads" pattern, in production. - Progressive disclosure —
gsd-core/workflows/help.mdreads only the one mode file matching the argument (brief0.9 KB /default1.9 KB /full34 KB). - Token-budgeted assembly —
src/prompt-budget.ctsapplyBudget()already does priority-ordered, budget-trimmed composition with an omission note — but it is walled into the cross-AI review pipeline only. - Pointer-passing channel —
src/io.ctsspills any payload > 50 KB to a tmpfile and returns@file:<path>. - A codegen factory + drift-guard harness — 13 generators share one
--check/--writeidiom (derive fresh, diff committed, exit 1 on drift).scripts/gen-plugin-skills.cjsalready generates 69 shippedSKILL.mdfiles fromcommands/gsd/*.md. - A reusable structured-markdown parser —
src/markdown-sectionizer.cts, already powering the per-phase<decisions>fact-store reader (src/decisions.cts).
External practice
The closest external analogs are Anthropic Agent Skills' three-tier progressive disclosure (metadata → SKILL.md → bundled references), MCP resources/prompts/deferred-tools (list-then-fetch JIT), and priority/token-budget prompt renderers (Priompt, VS Code @vscode/prompt-tsx) that include the highest-priority fragments that fit a budget via a binary-search cutoff, with flexReserve floors for load-bearing content and <isolate> for a stable cacheable prefix. The portability catch is real and load-bearing: only the Skills format (directory + SKILL.md + frontmatter) is an open standard; native lazy loading is Claude-specific, and GSD's 15 runtimes do not all support skills or MCP (cf. surface-mismatch bugs #1614 antigravity, #1615 windsurf).
Decision
Adopt a dynamic context management platform built on a hybrid of build-time and run-time assembly, reusing the existing seams rather than inventing new infrastructure:
-
Fragment store (authoring model). Author workflow content as composable, priority-tagged fragments (workflow sections + shared
references/+ predicate-derived blocks), each carrying an applicability condition (which flags / capabilities / runtimes require it). This is the net-new authoring discipline. -
Build-time composer + per-runtime budget emission (the universal floor). Generalize
prompt-budget.ctsout of the review silo into a sharedcontext-composerseam (src/*.cts→build:lib→bin/lib/*.cjs). At build/install time, for each command × runtime, the composer selects the needed fragments and trims by priority to fit that runtime's measured cap (scripts/workflow-size.cjslfByteCount), emitting a right-sized artifact through the existing converter. Caps move from source to emitted output; the Windsurf 12 KBthrowbecomes a graceful auto-trim/auto-extract — noting that thisthrowis currently duplicated byte-identically in two surfaces,bin/install.js:2796-2797andsrc/runtime-artifact-conversion.cts:1116-1117, so the change must land in both or they drift. This is what makes caps stop biting on non-lazy runtimes, and it requires no runtime feature — so it is the universal floor. -
Progressive disclosure where the host supports it. On lazy-loading hosts (Claude Code and the Agent SDK), keep the stub +
@-refmodel and let the init bundle name exactly which files to read; the body and references load on demand. -
Run-time selection via the init seam (per-request precision). Extend the init bundle /
command-routing-hubdispatch (src/command-routing-hub.cts) to emit a typed manifest of which sections / references / predicates a specific invocation needs (given parsed args, flags, phase state, active capabilities), reusing the@file:spill channel for assembled fragments. This is layered on top of the fragment store. -
Formalize the
CONTEXT.mdpredicate fact-store → JIT selector. Give the predicate grammar a parser (onmarkdown-sectionizer), an ID-uniqueness validator, a--check/--writedrift-guard, and atask → relevant predicate setselector. This converts hand-assembled briefs into JIT-generated context and attacks the maintainer-side "edit a 200 KB file by hand" pain directly. This is sequenced first (see Prototype) because it is the smallest, lowest-risk piece that proves the whole pattern. -
Defer MCP (Connected-Capability). Per ADR-857 §7 / #956, a served MCP catalog (resources/prompts/deferred-tools) remains an additive future enhancement for MCP-capable runtimes — never a replacement for the file-copy floor. Not in scope here.
Options considered
| Option | Summary | Fixes caps? | Runtime compat | Decision |
|---|---|---|---|---|
| A. Progressive-disclosure authoring | Metadata-first files + one-level references; lean on host lazy-load | Partial; needs host lazy-load | Authoring universal; native JIT Claude-first | Adopt as a layer |
| B. Build-time composer + per-runtime budget emission | Composer trims fragments to each runtime cap, emits right-sized files | Yes — measured before write | Universal floor | Adopt as core |
| C. Run-time selection via init seam | Init bundle names which slices this invocation needs | Reduces per-invocation context | Broad (the gsd_run shim is universal) |
Adopt after B |
| D. MCP served catalog | Serve content as resources/prompts/deferred-tools | For MCP hosts only | Partial; needs 2nd channel | Defer (ADR-857 §7) |
| E. Predicate fact-store → JIT selector | Parse/validate/select CONTEXT.md predicates |
Maintainer-side big-file pain | N/A (build + orchestrator) | Adopt first |
Pure Agent Skills (A alone) and pure MCP (D alone) were rejected as the foundation because both are runtime-partial; only build-time emission (B) relieves caps on every runtime.
Architecture and contracts
-
Fragment unit (open question, see below): either separate files (clean lazy-load + INVENTORY rows) or in-file section markers (
<!-- gsd:section ... -->, mirroring the existing<!-- gsd:loop-host -->markers consumed byscripts/gen-loop-host-contract.cjs). -
Composer contract: an ordered list of fragments, each carrying a shrink strategy; the closed set is
verbatim,head-shrink,proportional-truncate(with a per-fragment floor), anddrop.flexReserve-style floors for load-bearing fragments (META.RULEcitation rules, contribution gates, closing-keyword rules) generalize the existing per-plan 1024-byte floor. A byte-stable canonical prefix (<isolate>) is kept identical across runtimes to preserve KV-cache warmth and keep launcher-parity tests green.Amended by #2929 (Phase 2). This ADR originally specified the contract as "priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2 established that a cutoff alone cannot express the function this platform generalizes:
prompt-budget.applyBudgetis not a cutoff but a fixed five-step ladder in which each section carries its own shrink strategy, and only three of its eight sections are ever droppable —PROJECT.mdis head-shrunk to N lines and plans are proportionally tail-truncated with a per-plan floor, while instructions and roadmap are never trimmed at all. A cutoff composer sorts by priority and discards the tail; it has no way to say "shrink this one", "truncate that one but never below its floor", or "these three are the only droppables, in this order". Building to the literal wording and routingprompt-budgetthrough it would have silently changed review-prompt output. Shrink strategies are therefore the core abstraction, and binary-search cutoff becomes one strategy among them — the right one for per-runtime emission in Phases 3-4, not for this ladder. Ordering is declaration order rather than a numeric priority field. This is an elaboration of the decision's intent, not a reversal of it. -
Budget unit: bytes for emission caps (matches
lfByteCount, deterministic, offline-safe); a token estimate for run-time selection. -
Determinism + drift-guard: every generated artifact follows the universal
--check/--writeidiom and is committed; any constant shared between two surfaces gets aDEFECT.GENERATIVE-FIXparity assertion. Caps are asserted on emitted per-runtime bytes via real spawn-install tests (engine-direct tests are false-green for install behavior). -
Boundary coverage: the composer's budget logic is tested at
cap-1 / cap / cap+1perRULESET.TESTS.boundary-coverage.
Migration path
Sequenced to de-risk — prove the pattern on the smallest surface first, scale last:
- This ADR establishes the platform, the fragment/composer contract, emission-time caps, and the drift-guard requirement.
- Prototype the predicate fact-store (Option E) — landed with this ADR as a non-shipping reference example under
examples/dynamic-context-management/(see Prototype below). - Lift
prompt-budget.ctsout of the review silo into a sharedcontext-composerseam with fast-check property tests + boundary coverage. - Pilot fragmentization on one XL workflow (
plan-phase.mdorexecute-phase.md): split into priority-tagged sections + applicability; composer emits per-runtime; prove byte-identical-or-smaller output and greengsd-testdocker. - Move caps from source to emitted output; turn the Windsurf
throwinto graceful auto-trim; auto-regenerate size baselines on intentional edits. - Wire the init bundle (C) to emit a per-invocation sections manifest; workflows consume it.
- Roll out across LARGE/XL tiers; update INVENTORY families + parity tests.
- (Deferred) MCP served catalog (ADR-857 §7 / #956).
Ordering landmine: any generator consuming compiled output must run after build:lib (tsc), like gen-plugin-skills / gen-capability-registry; regenerating before build:lib silently drops unbuilt modules (gsd-inventory-manifest-regen-needs-build).
Consequences
Positive
- Caps stop biting: each runtime's emitted artifact is measured and trimmed before write.
- A discovered fact lands in one fragment / predicate, not 4 hand-edited surfaces.
- Reuses the existing converter, drift-guard, boundary-test, and
markdown-sectionizerinfrastructure — the net-new pieces are only the fragment model and the composer. - Opens a path to collapse the 10+ hand-written per-runtime body converters toward a data-driven spec.
Negative / risks
- Trimming a load-bearing fragment is a correctness hazard (history: paraphrased
META.RULE→ agent violations). Mitigate withflexReservefloors, a Promptfoo-style eval gate, and boundary tests. - Per-runtime emission multiplies artifacts across the 15 × N matrix (inventory/parity surface).
- Build-order fragility (must run after
build:lib). - Dual-surface drift if any future MCP channel is added — requires parity assertions.
Prototype (step 2, Option E) — non-shipping reference example
A working prototype proves the platform pattern end-to-end. It ships as a reference example only, under examples/dynamic-context-management/ — deliberately outside the build (src/ → bin/lib/), the npm package files[], the installer, and the CI test suite (tests/). Nothing in it is compiled into or installed with GSD; the production implementation lands in a later phase.
examples/dynamic-context-management/context-predicates.cjs— pure parser/selector:parsePredicates(markdown)(handles bare and list-item backtick predicate forms, splits on first=, skips fenced code / blockquote prose, detects duplicate IDs),selectPredicates(predicates, {klass, prefix, contains})(the JIT "task → predicate set" selector), andbuildIndex(predicates)(deterministic, sorted).examples/dynamic-context-management/gen-context-index.cjs— self-contained CLI with--check/--writedrift-guard plus a--select <query>mode demonstrating JIT brief assembly.examples/dynamic-context-management/CONTEXT-INDEX.json— sample generated index: 415 predicates, 20 classes (verified 2026-07-31; down from 416 after #2928/PR #2938 reconciled the last duplicate predicate ID,RULESET.WORKFLOW_MARKDOWN.FENCES). Originally committed as 393 predicates, 18 classes (2026-06-24);CONTEXT.mdhas since gained thePROBE(11) andPROHIB(10) classes, withDEFECT161→167 andRULESET59→56→55. The committed artifact had gone stale (--checkexited 1) and was regenerated with--write.examples/dynamic-context-management/demo.cjs+README.md— runnable usage example and notes.
During research the slice was validated with 42 behavioral tests (predicate forms, fenced-code / prose skipping, duplicate-id detection, the selector, a deterministic index, and a fast-check property test); those landed as CI tests under tests/ with the production implementation (#2928/PR #2938).
The prototype immediately surfaced 3 latent duplicate predicate IDs in CONTEXT.md (RULESET.WORKFLOW_MARKDOWN.FENCES, RULESET.GEMINI.TOOLS.ask_user, RULESET.GEMINI.TEST_SENTINEL) — integrity drift no existing tool catches.
Re-checked 2026-07-31: the two RULESET.GEMINI.* duplicates were removed along with the Gemini runtime, not reconciled deliberately; the remaining RULESET.WORKFLOW_MARKDOWN.FENCES duplicate was reconciled deliberately in #2928/PR #2938, which also productionized --check into CI so it now fails closed on any new duplicate ID.
Phase 0 acceptance status (2026-07-31). The epic's Phase 0 criterion — "gen-context-index --check green in CI" — was unmet: --check exited 1 against next, and no CI job failed, because the example sits deliberately outside tests/ — the red was invisible to the pipeline. The index has now been regenerated and --check exits 0.
That is a point-in-time true-up, not a fix. Per Open question 4, the index is keyed on baked line numbers, so it will re-drift on the next CONTEXT.md line shift. The criterion stays fragile until the keying changes — and it stays silently fragile for as long as the drift-guard remains outside CI.
Prototype scope notes: the parser is intentionally self-contained for the example; production should consume the compiled markdown-sectionizer seam, live under src/ → bin/lib/, and be drift-guarded by a generator wired into the build after build:lib.
Done (#2928). Production landed under src/context-predicates.cts → gsd-core/bin/lib/context-predicates.cjs (ADR-457 build-at-publish). Fence-aware line skipping mirrors markdown-sectionizer.cts's exported scanFencedBlocks delimiter-matching rule exactly (byte-for-behavior parity proven by a dedicated test suite) via a LOCAL, interleaved single pass, rather than a call into that seam directly: a two-pass design (mask comments, then call scanFencedBlocks, or the reverse) cannot correctly resolve mutual precedence between HTML comments and fences in both directions — a fence delimiter inside a real comment (with no later real closer) was found to falsely skip the rest of the file to EOF, and the converse ordering falsely let a comment token inside a real fence leak past the fence's own close — so the two constructs are scanned together, each suppressing the other's open/close detection while active (post-#2928-review fix; see src/context-predicates.cts's module doc comment). scripts/gen-context-index.cjs --check/--write is wired into lint:generated-sync (so lint:ci, CI-gated) and into build (after build:lib) and regen:derived; the selector is exposed live via gsd-tools query context-predicates --class|--prefix|--contains.
Open questions
- Fragment unit: separate files vs in-file section markers?
- Build-time emission vs run-time assembly as the primary surface during migration (double-write vs per-workflow cutover)?
- Whether/when to invest in per-runtime native channels (skills, MCP) above the universal file floor.
Resolved by #2928 — index keying: stable IDs, with no line field at all. Question 4 asked stable IDs vs baked line numbers: CONTEXT-INDEX.json stored each predicate's line, so --check re-drifted on any CONTEXT.md line shift — a typo fix three sections up failed the gate. Raised by @davesienkowski (#1671, 2026-06-25). The shipped resolution is stronger than the option originally proposed (keying the comparison on stable IDs with line retained as non-compared metadata): the committed ContextIndex.predicates entries carry no line field at all. Committed-but-uncompared metadata goes silently stale — the same defect class the drift-guard exists to catch, with the alarm removed — so it was dropped from the committed artifact rather than merely excluded from the comparison. line is still returned by the live parsePredicates/gsd-tools query context-predicates result for callers that want to cite a source location; only the committed docs/CONTEXT-INDEX.json shape omits it.
Resolved by other work — not carried as open. A fourth question was proposed in review (#1671, 2026-06-25): what populates the eval-gate assertion set, and is it graded exogenously? Since that review, the answer has landed as first-class predicate classes rather than remaining a design gap: PROBE.principle (verifier-reach-equals-spec-reach), PROBE.family (edge-probe + prohibition-probe + ui-consideration-probe), PROBE.protocol (recall → precision), and PROHIB.judgment-tier (exogenous grading) — see ADR-550 D4/D7 and ADR-1606. The PROHIB.* predicates live in the same CONTEXT.md store this ADR formalizes, which is the single-store property that review asked for.
Related
- ADR-0002 — Command Contract Validation Module (the stub
<execution_context>@-ref contract this platform's emission must keep satisfying). - ADR-457 — build-at-publish generation model (the codegen + drift-guard precedent the composer extends).
- ADR-857 §7 — Connected-Capability / MCP contract (the deferred served-catalog channel).