chore(#3065): build the deterministic load-bearing-fragment contract gate (#3068)

* test(#3065): build the load-bearing contract gate ADR-1671 promised

Epic #1671 Phase 7. A post-merge audit of every promise in ADR-1671 against the
merged tree found one mitigation asserted-but-absent and two stale records.

ADR-1671 names exactly one correctness risk — trimming a load-bearing fragment,
with the recorded history of a paraphrased META.RULE causing agent violations —
and #2931 amended its mitigation to a deterministic contract gate that proves no
load-bearing fragment was omitted or shrunk, treats a floored fragment as a
success, and asserts the isolate prefix survives byte-identical, with an explicit
anti-vacuity rule.

That gate did not exist. What existed was tests/context-composer.test.cjs:
synthetic unit tests of the composeWithinBudget primitive over invented
fragments, asserting nothing about real declared strategies. The ADR asserted a
mitigation that was never built, which is the promised-but-not-built shape the
epic's own coverage discipline exists to catch.

The gate derives its load-bearing set from declared verbatim strategies rather
than a hand-maintained list, so it cannot go stale as upstream changes. It sweeps
budgets from 4x total down to a quarter of total and asserts at every step that
no load-bearing id appears in omitted or shrunk, that isolatePrefix is
byte-identical, and that hardFailed is surfaced rather than silently passed.

Both anti-vacuity guards are EXECUTABLE, not comments. One proves the empty
load-bearing set guard actually throws. The other proves a sweep that never
applies pressure is rejected — because a gate that only ever runs unpressured is
exactly how the original mitigation went missing without anyone noticing.
Measured: underPressure true at 6 of 7 budgets, false only at 4x total.

Three ADR records corrected in the same change, all doc-vs-reality drift:

  - Decision item 2 describes a composer that trims by priority to fit a measured
    per-runtime cap. composeWorkflow in fact passes MAX_SAFE_INTEGER with every
    fragment verbatim (both verified in source), so no trimming happens there;
    the emitted-byte cap is a separate measure-and-fail gate and Windsurf's limit
    a bespoke truncation. The wording described an option as shipped behavior.
  - flag:--converge never reached a terminal state. #2992 withheld six atoms;
    five were resolved explicitly. This one was resolved in code by reusing
    state:plan-strategy-converge but recorded nowhere — the same gap #2995 closed
    for flag:--verify-only, and I closed five of six.
  - The open-questions list enumerated three questions while two Resolved-by
    blocks resolved an unlisted Question 4. It is now listed.

Refs #3065

* fix(#3065): make the gate assert over production, not a copy of it

The isolated review found a blocker, and it was fatal to the gate's purpose: it
hand-copied applyBudget's fragment array into the test, so flipping a strategy in
src/prompt-budget.cts — say roadmap from verbatim to drop — would leave the gate
computing from its own untouched copy and still passing. A guard built as an
instance of the very divergence class it exists to prevent
(DEFECT.GENERATIVE-FIX) is worse than no guard, because it reports green.

Fixed by eliminating the duplicate rather than adding a parity assertion, the
same resolution used for the FAMILIES table in #2996. applyBudget's inline
construction is extracted to an exported buildBudgetFragments(), which both
applyBudget and the gate now call; the 1024 plan floor is exported as
PLAN_FLOOR_CHARS instead of being re-declared in the test. The extraction is pure
— verified behavior-preserving at budget=2000: hardFailed false, omitted
['context'], projectMd shrunk, plan truncation ~27.8%, all headers present. There
is no longer a second copy to diverge from.

Also fixed a vacuous assertion the same review caught: isolatePrefix was pinned
across the sweep, but no production fragment sets isolate:true, so the value is
always '' and the check could never fail. The pinning assertion stays, with an
honest comment that nothing in production sets it today, and a second test now
constructs an isolate:true fragment set and proves the prefix is non-empty and
byte-identical across a roomy and a severely tight budget — which is what makes
the first assertion capable of detecting a real change.

Refs #3065

* chore(#3065): backfill changeset pr number to 3068

---------

Co-authored-by: sim <sim@local>
This commit is contained in:
Tom Boucher
2026-08-04 22:41:01 -04:00
committed by GitHub
parent 83a26ed1dc
commit c899f5ada3
4 changed files with 355 additions and 12 deletions

View File

@@ -186,6 +186,37 @@ Pure Agent Skills (A alone) and pure MCP (D alone) were rejected as the foundati
`DEFECT.AGENT-FILE-SIZE-CAP-BREACH` fix-forward), not by `when=` markers. Extending
gating to agents would require a per-agent manifest family and a dispatch-time seam to
read it; that is a separate decision, not an organic edit, and is not taken here.
**Amended by #3065 (Phase 7) — the promised contract gate is built, and three records are
corrected.** A post-merge audit of every promise in this ADR against the merged tree found one
mitigation asserted-but-absent and two stale records.
*The load-bearing contract gate now exists.* The Consequences section below claims, as amended by
#2931, that a deterministic gate proves no load-bearing fragment was omitted or shrunk. Until this
phase only synthetic unit tests of the `composeWithinBudget` primitive existed, over invented
fragments, asserting nothing about real content. `tests/load-bearing-contract-gate.test.cjs` now
derives the load-bearing set from declared `verbatim` strategies rather than a hand-maintained
list, sweeps a descending budget range, and carries both anti-vacuity guards as executable
assertions: an empty load-bearing set fails, and a sweep that never applies pressure fails.
*Decision item 2 overstated what shipped.* It describes a composer that "selects the needed
fragments and trims by priority to fit that runtime's measured cap". `composeWorkflow` in fact
calls `composeWithinBudget` with `budget: Number.MAX_SAFE_INTEGER` and every fragment
`{kind:'verbatim'}` — non-lossiness is a structural guarantee of the strategy set, not a
large-budget trick, and no per-runtime trimming happens there. The emitted-byte cap is enforced by
a separate measure-and-fail gate, and Windsurf's limit by a bespoke description truncation
(#2931), not by this composer. Per-runtime trimming remains available in the strategy set and
unused; the wording above describes an option, not shipped behavior.
*`flag:--converge` reaches a terminal state.* #2992 withheld six atoms and deferred them to the
rollout phase. Five were resolved explicitly. `flag:--converge` was resolved in code by reusing
`state:plan-strategy-converge` for `autonomous.md`'s converge sections, but that disposition was
recorded nowhere — the same undocumented-disposition gap #2995 closed for `flag:--verify-only`.
It is recorded here: **not admitted as its own atom; superseded by `state:plan-strategy-converge`.**
*Open-questions numbering is corrected.* The list enumerates three questions, while two "Resolved
by" blocks below resolve a "Question 4" that was never added to it. Question 4 — index keying,
stable ids vs baked line numbers — is now listed explicitly.
- **Budget unit:** bytes for emission caps (matches `lfByteCount`, deterministic, offline-safe); a token estimate for run-time selection.
**Corrected by #2931 (Phase 4) — the Windsurf cap was never load-bearing.** The Context
@@ -305,6 +336,7 @@ Prototype scope notes: the parser is intentionally self-contained for the exampl
1. Fragment unit: separate files vs in-file section markers?
2. Build-time emission vs run-time assembly as the primary surface during migration (double-write vs per-workflow cutover)?
3. Whether/when to invest in per-runtime native channels (skills, MCP) above the universal file floor.
4. Index keying: stable IDs vs baked `line` numbers? *(Resolved by #2928 — see below.)*
**Resolved by #2928 — index keying: stable IDs, with no `line` field at all.** Question 4 asked stable IDs vs baked `line` numbers: `CONTEXT-INDEX.json` stored each predicate's `line`, so `--check` re-drifted on *any* `CONTEXT.md` line shift — a typo fix three sections up failed the gate. Raised by @davesienkowski (#1671, 2026-06-25). The shipped resolution is **stronger than the option originally proposed** (keying the comparison on stable IDs with `line` retained as non-compared metadata): the committed `ContextIndex.predicates` entries carry **no `line` field at all**. Committed-but-uncompared metadata goes silently stale — the same defect class the drift-guard exists to catch, with the alarm removed — so it was dropped from the committed artifact rather than merely excluded from the comparison. `line` is still returned by the live `parsePredicates`/`gsd-tools query context-predicates` result for callers that want to cite a source location; only the committed `docs/CONTEXT-INDEX.json` shape omits it.