9e4f0e99ad05d3449be98d15fdd4b003d768febd
1 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9ac0dfad58 |
chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared context-composer seam. Its success condition is that review-prompt output does not change, and the only authority on "did not change" is the behavior that shipped before the refactor. Capture that behavior now, while it is still the live implementation. 47 characterization cases, every `expected` value computed by executing the current implementation rather than hand-authored — the independence CONTRIBUTING.md "Fixture provenance (#2371)" asks for. A corpus is only worth what it can detect, so this one was validated by mutation rather than assumed. Five deliberate defects were injected and each must be caught by at least one case: - the note reserve deducted unconditionally instead of only under pressure - the pressure test relaxed from `>` to `>=` - a no-op head-shrink still setting the shrunk flag - the per-plan floor dropped from the proportional share - drop order reversed Two of those exposed real holes in the first cut of this corpus, and the cases that close them exist because of it: - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a floored plan group, and the 1024-char floor absorbed the entire trim, so the mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap, which makes the strict inequality observable as context kept vs omitted. - No case reached proportional-truncate at all — B6 and B7 both hard-failed the min-set pre-check first, leaving planTruncationPct at 0 across every case and the floor semantics entirely unexercised. Rebudgeted to 700 and 1100 so the min-set fits and the truncate step is actually reached; they now record 40.20% and 48.80%. The A4/A10 families sweep the pressure boundary from both sides, which is where this function has regressed before: CONTEXT.md's LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band, because the suite paired a trivially-fitting budget with a trivially-overflowing one and never sampled between them. A4 pins that nothing is trimmed from the cap down to 81 tokens under it; A10 pins that pressure fires at +1. Together with A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures. Two facts the corpus establishes that the design notes had wrong: - "" and null sections are NOT distinguished. applyBudget uses truthy checks throughout, so an empty-string section is treated as absent: not rendered, not dropped, never recorded in `omitted`. B13b pins this while the ladder is actively trimming, where only the non-empty `research` is dropped. - Sizing matters. B12/B13 were first written at a budget where both hard-failed the min-set check and returned "", so comparing them compared two empty strings and proved nothing. Committed as its own commit, ahead of the refactor, and regenerated against the pre-refactor implementation, so the oracle is demonstrably independent of the change it will adjudicate. Refs #2929 * refactor(#2929): extract the context-composer seam from prompt-budget Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer — per-runtime artifact emission — but it is walled inside the cross-AI review pipeline. Lift it into a shared seam so later phases can call it, without changing what the review pipeline emits. ADR-1671 specifies the composer as "priority + binary-search cutoff to a per-runtime budget". Read against the code it generalizes, that contract cannot express the thing being generalized. applyBudget is not a cutoff: it is a fixed five-step ladder in which each section carries its own shrink strategy, and only three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor; instructions and roadmap are never touched at all. A cutoff composer sorts by priority and discards the tail — it has no way to say "shrink this one", "truncate that one but never below 1 KB each", or "these three are the only droppables, in this order". Building to the literal contract and routing prompt-budget through it would have silently changed review-prompt output, which is the one outcome this phase forbids. So shrink strategies are the core abstraction here, and cutoff becomes one strategy among them — the right one for per-runtime emission in Phases 3-4, not for this ladder. That is an elaboration of the ADR's intent, not a departure from it, and ADR-1671 is updated to say so. Three decisions worth stating: - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan of surviving fragments and never a string. assemblePrompt's rendering is prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and owning it in the composer would force emission to adopt prompt-shaped rendering. The split is what lets one seam serve both consumers. - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires for emission caps. The existing code converts a token budget to a character budget with a hardcoded `* 4`; that assumption is now an explicit `charsPerUnit` inverse, which is precisely what a byte unit needs in order to reuse this. - The entry point is `composeWithinBudget`, not `applyBudget`. That name already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter being an unrelated graph-edge budget. A third would make every symbol search in this repo ambiguous, and it already misresolves: preflight and impact queries for "applyBudget" return graphify's. Behavior is unchanged and proven so: all 47 characterization cases reproduce byte-identically, and the corpus is mutation-validated rather than merely green (see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and from eighteen mutable accumulators to two, both inside a helper copied verbatim. estimateTokens deliberately stays in prompt-budget and keeps its exact math: src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins plan estimates and recorded actuals to that same scale, so moving or changing it would silently break the calibration loop. Refs #2929 * docs(#2929): document the context-composer seam and amend ADR-1671 Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new domain modules), and a mutation-matrix entry for the new module. The ADR amendment is the substantive part. ADR-1671 specified the composer as "priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2 established that a cutoff alone cannot express the function the platform generalizes, so the ADR now records shrink strategies as the core abstraction with cutoff as one strategy among them, reserved for per-runtime emission in Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned against that contract and would otherwise be planned against a mechanism that does not work. The mutation-matrix entry is not bookkeeping. Stryker scores per module against a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the extracted code unmeasured while prompt-budget's own score floated free of the logic it used to cover. context-composer gets its own entry at the same floor. Refs #2929 * test(#2929): pin the effectiveBudget rounding mode in the parity corpus An isolated correctness review found a real blind spot: mutating `Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator happened to produce a whole number, so floor, round and ceil all agreed and the rounding mode was entirely unpinned by a corpus whose whole job is to pin observable behavior. Three cases fix that by straddling the .5 boundary: A11 95 * 0.90 = 85.5 floor 85, round 86 -> the two disagree A12 97 * 0.90 = 87.3 floor and round agree; ceil (88) does not A13 93 * 0.85 = 79.05 same guard at a non-multiple-of-10 margin, so the margin arithmetic is exercised and not just the budget A11 alone catches the round mutation; all three catch ceil. Regenerated against the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so the expanded corpus keeps the independence property the original capture had. The corpus is now mutation-validated against seven injected defects, every one caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink setting its flag, the truncate floor ignored, drop order reversed, and both rounding-mode changes. Refs #2929 * feat(#2929): flexReserve floors and the byte-stable isolate prefix Two of issue #2929's "Done when" items were unimplemented rather than deferred, and an isolated review flagged them alongside my own audit. Both are part of ADR-1671's composer contract, so shipping the seam without them would have left Phases 3-4 building against a contract that does not exist yet. flexReserve is a per-fragment floor in measure units that every strategy must respect, which is what makes it different from the pre-existing floorChars: that one is a chars-denominated detail of proportional-truncate alone and is retained unchanged. A floored fragment is never dropped, is never head-shrunk below its floor, and raises its own proportional cap. A fragment already smaller than its floor is untouchable outright. Metadata gains `floored`, listing the ids whose floor actually prevented a trim — a guarantee no caller can observe is a guarantee no test can hold you to. isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed, never dropped, but still counted, because a prefix excluded from accounting would silently under-count real context. Metadata gains `isolatePrefix` so a caller can hash or assert on the exact bytes. Declaring an isolate fragment after a non-isolate one throws: a prefix that is not at the front is not a prefix, and accepting it would make the cross-runtime stability claim meaningless. Adds tests/context-composer.test.cjs for the exact new semantics and tests/context-composer.property.test.cjs for the five invariants, including the budget-monotonicity property the issue names explicitly. Both are registered in the mutation matrix, since coverage does not migrate with relocated code. prompt-budget uses neither feature, and its output is unchanged: all 50 corpus cases still reproduce byte-identically. Refs #2929 * chore(#2929): allowlist the prompt-budget parity suite The parity corpus needs its own test file and that makes prompt-budget a three-file module against a limit of two. The lint offers consolidation or an allowlist entry with justification; the entry is the right call here. Consolidation would mean folding the characterization suite into prompt-budget.test.cjs, which is the one thing that should not happen to it. The parity suite is a distinct concern with a distinct lifecycle: it is generated rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own scoring target, and its failure means something categorically different from a unit-test failure — not "this behavior is wrong" but "observable output moved". Burying it inside a general unit file would obscure exactly that signal. The allowlist is an identity ratchet, so this entry pins today's three exact filenames: adding a fourth still fails, and dropping back to two requires removing the entry. Refs #2929 * fix(#2929): register the new module with two gates it was missing The remote matrix caught three defects that no local check could, because the local runner is blocked in this repo and these suites had therefore never executed. Eight failures, identical on node22 and node24, so nothing environment-shaped. Two are the new-module ripple. A net-new src/*.cts lands in six places and this change had reached four of them — .gitignore, INVENTORY, the manifest, and the CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet baseline (a deliberate review-visible mirror of the matrix floors, which every COVERED module must carry). Both are now registered, the ratchet at the same floor of 66 the matrix declares. The third was a test asserting an outcome it had made impossible. It set budget:1 alongside a 400-char required fragment, so the group budget came out at -99 and the proportional-truncate step was skipped entirely — the deliberate "non-positive group budget is skipped, never clamped" rule inherited from the original ladder. Nothing was trimmed, and the test then asserted a truncation. Rebudgeted so the step actually runs, with the arithmetic written out in a comment so the next reader does not have to re-derive why 120 rather than 80. Fixing that surfaced a genuine bug in the composer. `floored` is documented as recording fragments whose flexReserve prevented a trim that would otherwise have happened, but the push sat in the else-branch of "content did not change", so it only fired when nothing was trimmed at all. A fragment truncated to a reserve-raised cap has also had a trim prevented — 40 characters' worth in the test above — and was silently absent from the field that exists to make the guarantee observable. The condition was already right; it was in the wrong branch. Now recorded on both paths: a drop prevented outright, and a truncation capped higher than the share alone would have allowed. Parity is unaffected — prompt-budget never sets flexReserve, so the branch is unreachable from every corpus path, and all 50 cases still match. Refs #2929 * chore(#2929): backfill changeset PR number (#2958) * chore(#2929): correct the corpus case count in the changeset fragment --------- Co-authored-by: sim <sim@local> |