* test(#3065): build the load-bearing contract gate ADR-1671 promised Epic #1671 Phase 7. A post-merge audit of every promise in ADR-1671 against the merged tree found one mitigation asserted-but-absent and two stale records. ADR-1671 names exactly one correctness risk — trimming a load-bearing fragment, with the recorded history of a paraphrased META.RULE causing agent violations — and #2931 amended its mitigation to a deterministic contract gate that proves no load-bearing fragment was omitted or shrunk, treats a floored fragment as a success, and asserts the isolate prefix survives byte-identical, with an explicit anti-vacuity rule. That gate did not exist. What existed was tests/context-composer.test.cjs: synthetic unit tests of the composeWithinBudget primitive over invented fragments, asserting nothing about real declared strategies. The ADR asserted a mitigation that was never built, which is the promised-but-not-built shape the epic's own coverage discipline exists to catch. The gate derives its load-bearing set from declared verbatim strategies rather than a hand-maintained list, so it cannot go stale as upstream changes. It sweeps budgets from 4x total down to a quarter of total and asserts at every step that no load-bearing id appears in omitted or shrunk, that isolatePrefix is byte-identical, and that hardFailed is surfaced rather than silently passed. Both anti-vacuity guards are EXECUTABLE, not comments. One proves the empty load-bearing set guard actually throws. The other proves a sweep that never applies pressure is rejected — because a gate that only ever runs unpressured is exactly how the original mitigation went missing without anyone noticing. Measured: underPressure true at 6 of 7 budgets, false only at 4x total. Three ADR records corrected in the same change, all doc-vs-reality drift: - Decision item 2 describes a composer that trims by priority to fit a measured per-runtime cap. composeWorkflow in fact passes MAX_SAFE_INTEGER with every fragment verbatim (both verified in source), so no trimming happens there; the emitted-byte cap is a separate measure-and-fail gate and Windsurf's limit a bespoke truncation. The wording described an option as shipped behavior. - flag:--converge never reached a terminal state. #2992 withheld six atoms; five were resolved explicitly. This one was resolved in code by reusing state:plan-strategy-converge but recorded nowhere — the same gap #2995 closed for flag:--verify-only, and I closed five of six. - The open-questions list enumerated three questions while two Resolved-by blocks resolved an unlisted Question 4. It is now listed. Refs #3065 * fix(#3065): make the gate assert over production, not a copy of it The isolated review found a blocker, and it was fatal to the gate's purpose: it hand-copied applyBudget's fragment array into the test, so flipping a strategy in src/prompt-budget.cts — say roadmap from verbatim to drop — would leave the gate computing from its own untouched copy and still passing. A guard built as an instance of the very divergence class it exists to prevent (DEFECT.GENERATIVE-FIX) is worse than no guard, because it reports green. Fixed by eliminating the duplicate rather than adding a parity assertion, the same resolution used for the FAMILIES table in #2996. applyBudget's inline construction is extracted to an exported buildBudgetFragments(), which both applyBudget and the gate now call; the 1024 plan floor is exported as PLAN_FLOOR_CHARS instead of being re-declared in the test. The extraction is pure — verified behavior-preserving at budget=2000: hardFailed false, omitted ['context'], projectMd shrunk, plan truncation ~27.8%, all headers present. There is no longer a second copy to diverge from. Also fixed a vacuous assertion the same review caught: isolatePrefix was pinned across the sweep, but no production fragment sets isolate:true, so the value is always '' and the check could never fail. The pinning assertion stays, with an honest comment that nothing in production sets it today, and a second test now constructs an isolate:true fragment set and proves the prefix is non-empty and byte-identical across a roomy and a severely tight budget — which is what makes the first assertion capable of detecting a real change. Refs #3065 * chore(#3065): backfill changeset pr number to 3068 --------- Co-authored-by: sim <sim@local>
This commit is contained in:
@@ -186,6 +186,37 @@ Pure Agent Skills (A alone) and pure MCP (D alone) were rejected as the foundati
|
||||
`DEFECT.AGENT-FILE-SIZE-CAP-BREACH` fix-forward), not by `when=` markers. Extending
|
||||
gating to agents would require a per-agent manifest family and a dispatch-time seam to
|
||||
read it; that is a separate decision, not an organic edit, and is not taken here.
|
||||
|
||||
**Amended by #3065 (Phase 7) — the promised contract gate is built, and three records are
|
||||
corrected.** A post-merge audit of every promise in this ADR against the merged tree found one
|
||||
mitigation asserted-but-absent and two stale records.
|
||||
|
||||
*The load-bearing contract gate now exists.* The Consequences section below claims, as amended by
|
||||
#2931, that a deterministic gate proves no load-bearing fragment was omitted or shrunk. Until this
|
||||
phase only synthetic unit tests of the `composeWithinBudget` primitive existed, over invented
|
||||
fragments, asserting nothing about real content. `tests/load-bearing-contract-gate.test.cjs` now
|
||||
derives the load-bearing set from declared `verbatim` strategies rather than a hand-maintained
|
||||
list, sweeps a descending budget range, and carries both anti-vacuity guards as executable
|
||||
assertions: an empty load-bearing set fails, and a sweep that never applies pressure fails.
|
||||
|
||||
*Decision item 2 overstated what shipped.* It describes a composer that "selects the needed
|
||||
fragments and trims by priority to fit that runtime's measured cap". `composeWorkflow` in fact
|
||||
calls `composeWithinBudget` with `budget: Number.MAX_SAFE_INTEGER` and every fragment
|
||||
`{kind:'verbatim'}` — non-lossiness is a structural guarantee of the strategy set, not a
|
||||
large-budget trick, and no per-runtime trimming happens there. The emitted-byte cap is enforced by
|
||||
a separate measure-and-fail gate, and Windsurf's limit by a bespoke description truncation
|
||||
(#2931), not by this composer. Per-runtime trimming remains available in the strategy set and
|
||||
unused; the wording above describes an option, not shipped behavior.
|
||||
|
||||
*`flag:--converge` reaches a terminal state.* #2992 withheld six atoms and deferred them to the
|
||||
rollout phase. Five were resolved explicitly. `flag:--converge` was resolved in code by reusing
|
||||
`state:plan-strategy-converge` for `autonomous.md`'s converge sections, but that disposition was
|
||||
recorded nowhere — the same undocumented-disposition gap #2995 closed for `flag:--verify-only`.
|
||||
It is recorded here: **not admitted as its own atom; superseded by `state:plan-strategy-converge`.**
|
||||
|
||||
*Open-questions numbering is corrected.* The list enumerates three questions, while two "Resolved
|
||||
by" blocks below resolve a "Question 4" that was never added to it. Question 4 — index keying,
|
||||
stable ids vs baked line numbers — is now listed explicitly.
|
||||
- **Budget unit:** bytes for emission caps (matches `lfByteCount`, deterministic, offline-safe); a token estimate for run-time selection.
|
||||
|
||||
**Corrected by #2931 (Phase 4) — the Windsurf cap was never load-bearing.** The Context
|
||||
@@ -305,6 +336,7 @@ Prototype scope notes: the parser is intentionally self-contained for the exampl
|
||||
1. Fragment unit: separate files vs in-file section markers?
|
||||
2. Build-time emission vs run-time assembly as the primary surface during migration (double-write vs per-workflow cutover)?
|
||||
3. Whether/when to invest in per-runtime native channels (skills, MCP) above the universal file floor.
|
||||
4. Index keying: stable IDs vs baked `line` numbers? *(Resolved by #2928 — see below.)*
|
||||
|
||||
**Resolved by #2928 — index keying: stable IDs, with no `line` field at all.** Question 4 asked stable IDs vs baked `line` numbers: `CONTEXT-INDEX.json` stored each predicate's `line`, so `--check` re-drifted on *any* `CONTEXT.md` line shift — a typo fix three sections up failed the gate. Raised by @davesienkowski (#1671, 2026-06-25). The shipped resolution is **stronger than the option originally proposed** (keying the comparison on stable IDs with `line` retained as non-compared metadata): the committed `ContextIndex.predicates` entries carry **no `line` field at all**. Committed-but-uncompared metadata goes silently stale — the same defect class the drift-guard exists to catch, with the alarm removed — so it was dropped from the committed artifact rather than merely excluded from the comparison. `line` is still returned by the live `parsePredicates`/`gsd-tools query context-predicates` result for callers that want to cite a source location; only the committed `docs/CONTEXT-INDEX.json` shape omits it.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user