chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)

* fix(#2931): preserve protected regions and cap emitted per-runtime bytes

Route every runtime brand swap through applyClaudeCodeBrandSwap so
"Claude Code" survives verbatim inside <runtime_compatibility> regions
(#2284b). The fix existed only in bin/install.js's local copies; the
src/*.cts exports still used a naive replace, so binding install.js to
the single source -- as this phase does for the Windsurf family --
would have silently regressed those runtimes. A table-driven parity
guard now covers all nine brand-swapping converters.

De-duplicate the Windsurf converter family: delete the six local copies
in bin/install.js and bind the four exported ones by reference, guarded
by reference-identity assertions (the ADR-1508/#1675 pattern). The two
unexported helpers and an unused tool table go with them.

Replace the Windsurf 12,000-byte throw with description truncation,
matching the bound its sibling skill converter already applied. The
throw could only fire on an ~11.7 KB frontmatter description: the
largest emitted workflow is 311 bytes. Truncation makes the cap
unreachable by construction and leaves 12,000 in exactly one place,
eliminating the dual-surface duplication rather than testing for it.

Add the emitted-byte cap gate: buildEmittedSizes captures LF- and
<HOME>-normalized bytes from the walk buildParityManifest already
performs, and evaluateEmittedCaps asserts them against a per-runtime
cap table with dead-rule detection. buildParityManifest's return shape
is deliberately unchanged -- diffEmitted compares its values with
===, so making them objects would report all 8,529 emitted paths as
moved. A regression test pins the values as strings.

Add a deterministic trim-safety gate over composeWithinBudget's
omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity
rule, replacing the model-graded eval gate the issue described.

* docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract

* fix(#2931): bound the windsurf command name and single-source the brand swap

Review findings from the orthogonal passes, all fixed inline.

The claim that removing the 12,000-byte throw left total emission
"bounded by construction" was false. The #1615 regex constrains the
character class but not the length, and commandName is interpolated
three times into the emitted workflow: a 20,000-character name emitted
60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate,
clearly-labelled size control that THROWS -- commandName is the @-ref
path target, so truncating it would point the workflow at a file that
does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The
#1615 security regex is untouched and still runs first. 128 is generous:
the longest shipped name is gsd-plan-review-convergence at 27.

Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe
truncation helper. It still used a UTF-16 slice(0,177) -- the exact
surrogate-splitting bug the helper was written to avoid, in the very
sibling the helper's comment cites as its model. Bounds are unchanged,
so output is byte-identical for every shipped command (descriptions max
out at 99 chars).

Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting
the local copy. Adding it to the .cts left two unlinked implementations
of identical logic -- the drift class this change exists to remove.
Verified byte-identical across eight fixtures and five sequential calls
before merging, and guarded by a reference-identity assertion.

Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344),
add fast-check property coverage for the trim-safety contract, and use
fc.pre instead of a bare return in a property callback.

* test(#2931): fix three test-authoring bugs the remote matrix caught

The remote runner returned 8 unique failures on 6f15cdeb8. All three
causes were in the test files, not the modules under test -- local
harnesses exercise the modules directly, so nothing executed the test
bodies until the matrix did.

`{ __proto__: [...] }` in an object literal sets the prototype instead
of an own key, so the JSON round-trip erased it and the cap table never
saw a reserved runtime key. The production rejection was already
correct; the test could not reach it. Use a computed key.

Two cap fixtures tripped orthogonal error paths rather than the paths
they name: one declared windsurf in the cap table but omitted it from
sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern
matching nothing (a genuine dead rule). Both now include a compliant
artifact so the intended branch is what is asserted. The dead-rule and
unknown-runtime contracts are deliberate and unchanged.

`const { root } = makeSyntheticConfig({ ... `${root}` })` referenced
`root` from inside its own initializer -- a temporal dead zone error.
makeSyntheticConfig now optionally takes a (root) => files factory.

Also raise the npm pack --dry-run bound 60s -> 120s in the shipped-
scripts packaging test. That failure is NOT from this branch: the file
is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s
here, and the run recorded 60,637ms against a 60,000ms bound -- a
timeout under 28,948-test parallel contention, not a slowdown. Fixed
rather than deferred because a bound that tight is fragile regardless
of which branch trips it.

* chore(#2931): backfill changeset pr number to 2984

---------

Co-authored-by: sim <sim@local>
This commit is contained in:
Tom Boucher
2026-08-01 16:00:14 -04:00
committed by GitHub
parent 4df6d884b3
commit 628648d63a
15 changed files with 2284 additions and 221 deletions

View File

@@ -0,0 +1,120 @@
'use strict';
/**
* trim-safety.cjs — the trim-safety contract gate over `ComposeMetadata`
* (issue #2931, epic #1671, Phase 4).
*
* `composeWithinBudget` (src/context-composer.cts) is a pressure-aware
* budgeter: under load it may shrink, floor, or drop fragments entirely. This
* module is the gate a caller runs AFTER composing to prove that pressure
* never silently touched a fragment the caller has declared load-bearing —
* an id whose full, unshrunk content the caller is relying on to be present.
*
* Pure: no fs, no git, no clock. It only ever reads the `ComposeMetadata`
* shape (`omitted`, `shrunk`, `floored`, `isolatePrefix`, `hardFailed`,
* `hardFailReason`) and a caller-declared `loadBearingIds` list.
*
* ── The anti-vacuity rule is the most important rule in this module ────────
* A trim-safety gate that runs with an EMPTY `loadBearingIds` set would pass
* on every input, forever, having asserted nothing at all — the exact "looks
* like coverage and is not" failure this repo's test-matrix discipline
* exists to catch. `loadBearingIds` empty or absent is therefore a hard
* error (REASON.NO_LOAD_BEARING_DECLARED), not a vacuous pass.
*/
const REASON = Object.freeze({
LOAD_BEARING_OMITTED: 'load_bearing_omitted',
LOAD_BEARING_SHRUNK: 'load_bearing_shrunk',
ISOLATE_PREFIX_DRIFT: 'isolate_prefix_drift',
MINIMUM_SET_HARD_FAIL: 'minimum_set_hard_fail',
NO_LOAD_BEARING_DECLARED: 'no_load_bearing_declared',
});
/** Stable comparator over findings of possibly-different shapes: sort by
* `reason`, then by `id` (absent for ISOLATE_PREFIX_DRIFT, which sorts
* first within its reason bucket via the empty-string fallback). */
function byReasonThenId(a, b) {
if (a.reason !== b.reason) return a.reason < b.reason ? -1 : 1;
const aId = a.id || '';
const bId = b.id || '';
if (aId === bId) return 0;
return aId < bId ? -1 : 1;
}
/**
* Evaluate a compose result against a caller-declared load-bearing set.
*
* @param {object} opts
* @param {object} opts.metadata a `ComposeMetadata` from
* `composeWithinBudget` (src/context-composer.cts): `omitted`, `shrunk`,
* `floored` (string[] fragment ids), `isolatePrefix` (string),
* `hardFailed` (boolean), `hardFailReason` ('minimum-set' | null).
* @param {string[]} opts.loadBearingIds fragment ids the caller declares
* must survive intact. MUST be non-empty — see the anti-vacuity rule above.
* @param {string} [opts.expectedIsolatePrefix] when provided, the exact
* byte-for-byte prefix `metadata.isolatePrefix` must equal.
* @returns {{ findings: Array, errors: Array, ok: boolean }}
*/
function evaluateTrimSafety({ metadata, loadBearingIds, expectedIsolatePrefix } = {}) {
if (!Array.isArray(loadBearingIds) || loadBearingIds.length === 0) {
// Anti-vacuity: return immediately. Every other check below would be
// evaluated against a caller who declared nothing worth protecting, so
// computing them would dress up "proves nothing" as a real result.
return {
findings: [],
errors: [{
reason: REASON.NO_LOAD_BEARING_DECLARED,
message:
'loadBearingIds must be a non-empty array — a trim-safety gate with an empty '
+ 'assertion set proves nothing',
}],
ok: false,
};
}
const errors = [];
const findings = [];
const meta = metadata && typeof metadata === 'object' ? metadata : {};
const omitted = Array.isArray(meta.omitted) ? meta.omitted : [];
const shrunk = Array.isArray(meta.shrunk) ? meta.shrunk : [];
if (meta.hardFailed === true) {
errors.push({
reason: REASON.MINIMUM_SET_HARD_FAIL,
hardFailReason: meta.hardFailReason ?? null,
message: 'compose hard-failed (minimum required set could not fit the budget) — never a silent empty compose',
});
}
for (const id of loadBearingIds) {
if (omitted.includes(id)) {
findings.push({ reason: REASON.LOAD_BEARING_OMITTED, id });
}
if (shrunk.includes(id)) {
findings.push({ reason: REASON.LOAD_BEARING_SHRUNK, id });
}
// `floored` is deliberately never a finding: it means the floor did its
// job and the fragment's declared minimum survived.
}
if (expectedIsolatePrefix !== undefined) {
const actual = typeof meta.isolatePrefix === 'string' ? meta.isolatePrefix : '';
if (actual !== expectedIsolatePrefix) {
// Byte-identical means byte-identical: no trimming, no normalizing —
// a trailing-whitespace-only difference IS drift.
findings.push({ reason: REASON.ISOLATE_PREFIX_DRIFT, expected: expectedIsolatePrefix, actual });
}
}
findings.sort(byReasonThenId);
const ok = errors.length === 0 && findings.length === 0;
return { findings, errors, ok };
}
module.exports = {
REASON,
evaluateTrimSafety,
};