Files
msd-core/tests/runtime-brand-swap-parity.test.cjs
Tom Boucher 628648d63a chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)
* fix(#2931): preserve protected regions and cap emitted per-runtime bytes

Route every runtime brand swap through applyClaudeCodeBrandSwap so
"Claude Code" survives verbatim inside <runtime_compatibility> regions
(#2284b). The fix existed only in bin/install.js's local copies; the
src/*.cts exports still used a naive replace, so binding install.js to
the single source -- as this phase does for the Windsurf family --
would have silently regressed those runtimes. A table-driven parity
guard now covers all nine brand-swapping converters.

De-duplicate the Windsurf converter family: delete the six local copies
in bin/install.js and bind the four exported ones by reference, guarded
by reference-identity assertions (the ADR-1508/#1675 pattern). The two
unexported helpers and an unused tool table go with them.

Replace the Windsurf 12,000-byte throw with description truncation,
matching the bound its sibling skill converter already applied. The
throw could only fire on an ~11.7 KB frontmatter description: the
largest emitted workflow is 311 bytes. Truncation makes the cap
unreachable by construction and leaves 12,000 in exactly one place,
eliminating the dual-surface duplication rather than testing for it.

Add the emitted-byte cap gate: buildEmittedSizes captures LF- and
<HOME>-normalized bytes from the walk buildParityManifest already
performs, and evaluateEmittedCaps asserts them against a per-runtime
cap table with dead-rule detection. buildParityManifest's return shape
is deliberately unchanged -- diffEmitted compares its values with
===, so making them objects would report all 8,529 emitted paths as
moved. A regression test pins the values as strings.

Add a deterministic trim-safety gate over composeWithinBudget's
omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity
rule, replacing the model-graded eval gate the issue described.

* docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract

* fix(#2931): bound the windsurf command name and single-source the brand swap

Review findings from the orthogonal passes, all fixed inline.

The claim that removing the 12,000-byte throw left total emission
"bounded by construction" was false. The #1615 regex constrains the
character class but not the length, and commandName is interpolated
three times into the emitted workflow: a 20,000-character name emitted
60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate,
clearly-labelled size control that THROWS -- commandName is the @-ref
path target, so truncating it would point the workflow at a file that
does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The
#1615 security regex is untouched and still runs first. 128 is generous:
the longest shipped name is gsd-plan-review-convergence at 27.

Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe
truncation helper. It still used a UTF-16 slice(0,177) -- the exact
surrogate-splitting bug the helper was written to avoid, in the very
sibling the helper's comment cites as its model. Bounds are unchanged,
so output is byte-identical for every shipped command (descriptions max
out at 99 chars).

Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting
the local copy. Adding it to the .cts left two unlinked implementations
of identical logic -- the drift class this change exists to remove.
Verified byte-identical across eight fixtures and five sequential calls
before merging, and guarded by a reference-identity assertion.

Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344),
add fast-check property coverage for the trim-safety contract, and use
fc.pre instead of a bare return in a property callback.

* test(#2931): fix three test-authoring bugs the remote matrix caught

The remote runner returned 8 unique failures on 6f15cdeb8. All three
causes were in the test files, not the modules under test -- local
harnesses exercise the modules directly, so nothing executed the test
bodies until the matrix did.

`{ __proto__: [...] }` in an object literal sets the prototype instead
of an own key, so the JSON round-trip erased it and the cap table never
saw a reserved runtime key. The production rejection was already
correct; the test could not reach it. Use a computed key.

Two cap fixtures tripped orthogonal error paths rather than the paths
they name: one declared windsurf in the cap table but omitted it from
sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern
matching nothing (a genuine dead rule). Both now include a compliant
artifact so the intended branch is what is asserted. The dead-rule and
unknown-runtime contracts are deliberate and unchanged.

`const { root } = makeSyntheticConfig({ ... `${root}` })` referenced
`root` from inside its own initializer -- a temporal dead zone error.
makeSyntheticConfig now optionally takes a (root) => files factory.

Also raise the npm pack --dry-run bound 60s -> 120s in the shipped-
scripts packaging test. That failure is NOT from this branch: the file
is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s
here, and the run recorded 60,637ms against a 60,000ms bound -- a
timeout under 28,948-test parallel contention, not a slowdown. Fixed
rather than deferred because a bound that tight is fragile regardless
of which branch trips it.

* chore(#2931): backfill changeset pr number to 2984

---------

Co-authored-by: sim <sim@local>
2026-08-01 16:00:14 -04:00

133 lines
6.5 KiB
JavaScript

'use strict';
/**
* runtime-brand-swap-parity.test.cjs — DEFECT.GENERATIVE-FIX family parity
* guard for the #2284(b) protected-region fix.
*
* `applyClaudeCodeBrandSwap` (src/runtime-artifact-conversion.cts) rewrites
* bare "Claude Code" self-references to a runtime's brand name EXCEPT inside
* `<runtime_compatibility>...</runtime_compatibility>` blocks, which must
* survive byte-for-byte verbatim (a runtime-comparison table that says
* "Claude Code" is describing Claude Code's own behavior, not this
* runtime's — brand-swapping it mislabels the comparison). The Windsurf
* converter got this fix; every sibling markdown/agent converter that also
* performs a "Claude Code" -> brand swap needed the SAME fix (#2284b
* follow-up).
*
* A per-converter regression test would only catch a REintroduction in the
* converter it targets — the exact shape of divergence CONTEXT.md's
* DEFECT.GENERATIVE-FIX names. This file instead drives every brand-swapping
* converter through ONE assertion body from a declarative table, so a NEW
* runtime converter that skips `applyClaudeCodeBrandSwap` and reintroduces a
* naive `.replace(/\bClaude Code\b/g, brand)` fails the moment it is added
* to the table below — and the floor test catches an entry silently dropped
* from the table itself.
*/
process.env.GSD_TEST_MODE = '1';
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const conv = require('../gsd-core/bin/lib/runtime-artifact-conversion.cjs');
const capabilityRegistry = require('../gsd-core/bin/lib/capability-registry.cjs');
// Qwen/Hermes brand-swap the "Claude Code" literal to a descriptor-driven
// value (runtime.hostBehaviors.brandingRewrites), not a hardcoded string —
// read the SAME source the converters themselves read, so this test can
// never drift from the shipped descriptor.
function brandingRewrite(runtime, key) {
return capabilityRegistry.runtimes[runtime].runtime.hostBehaviors.brandingRewrites[key];
}
const QWEN_BRAND = brandingRewrite('qwen', 'Claude Code');
const HERMES_BRAND = brandingRewrite('hermes', 'Claude Code');
/**
* Every runtime converter in src/runtime-artifact-conversion.cts that
* performs a "Claude Code" -> brand swap, normalized to a uniform
* `(content) => string` shape regardless of the underlying function's real
* signature (fixed-brand markdown converters vs descriptor-driven
* agent/rewrite functions).
*/
const BRAND_SWAP_CONVERTERS = [
{ name: 'cursor', brand: 'Cursor', convert: (content) => conv.convertClaudeToCursorMarkdown(content) },
{ name: 'windsurf', brand: 'Windsurf', convert: (content) => conv.convertClaudeToWindsurfMarkdown(content) },
{ name: 'augment', brand: 'Augment', convert: (content) => conv.convertClaudeToAugmentMarkdown(content) },
{ name: 'trae', brand: 'Trae', convert: (content) => conv.convertClaudeToTraeMarkdown(content) },
{ name: 'codebuddy', brand: 'CodeBuddy', convert: (content) => conv.convertClaudeToCodebuddyMarkdown(content) },
{ name: 'cline', brand: 'Cline', convert: (content) => conv.convertClaudeToCliineMarkdown(content) },
// Dynamic (descriptor-driven) brand converters — same protected-region
// guard, brand value sourced from capability.json instead of a literal.
{ name: 'qwen-agent', brand: QWEN_BRAND, convert: (content) => conv.convertClaudeAgentToQwenAgent(content) },
{
name: 'qwen-runtime-rewrites',
brand: QWEN_BRAND,
convert: (content) => conv._applyRuntimeRewrites(content, 'qwen', '~/.qwen/', false, undefined),
},
{
name: 'hermes-runtime-rewrites',
brand: HERMES_BRAND,
convert: (content) => conv._applyRuntimeRewrites(content, 'hermes', '~/.hermes/', false, undefined),
},
];
// Floor, not an exact count (mirrors MINIMUM_MANIFEST_FAMILIES in
// tests/helpers/install-shared.cjs): the point of this table is that a NEW
// brand-swapping converter must be added here explicitly and reviewably.
// Lowering it is a deliberate act; this only guards it from silently
// shrinking underneath a refactor.
const MINIMUM_BRAND_SWAP_CONVERTER_COUNT = 9;
const PROTECTED_BLOCK =
'<runtime_compatibility>\n| Runtime | Claude Code | Other |\n|---|---|---|\n| x | Claude Code native | y |\n</runtime_compatibility>';
function buildFixture() {
return `Before: mentions Claude Code here.\n\n${PROTECTED_BLOCK}\n\nAfter: also mentions Claude Code here.\n`;
}
describe('everyRuntimeMarkdownConverterPreservesRuntimeCompatibilityBlocks', () => {
test('the brand-swap converter table has not silently shrunk below its floor', () => {
assert.ok(
BRAND_SWAP_CONVERTERS.length >= MINIMUM_BRAND_SWAP_CONVERTER_COUNT,
`expected at least ${MINIMUM_BRAND_SWAP_CONVERTER_COUNT} brand-swapping converters in the table, `
+ `got ${BRAND_SWAP_CONVERTERS.length}`,
);
});
for (const { name, brand, convert } of BRAND_SWAP_CONVERTERS) {
test(`${name}: <runtime_compatibility> block is preserved verbatim while outside text is brand-swapped`, () => {
// Guard the table entry itself before trusting the assertions below —
// an undefined brand (a descriptor lookup that silently returned
// nothing) would make every `includes()` check below vacuously
// meaningless.
assert.strictEqual(typeof brand, 'string', `${name}: table entry must declare a string brand`);
assert.ok(brand.length > 0, `${name}: table entry brand must be non-empty`);
const result = convert(buildFixture());
assert.ok(
result.includes(PROTECTED_BLOCK),
`${name}: <runtime_compatibility> block must survive byte-for-byte verbatim`,
);
assert.ok(
result.includes(`Before: mentions ${brand} here.`),
`${name}: text BEFORE the protected block must be brand-swapped to "${brand}"`,
);
assert.ok(
result.includes(`After: also mentions ${brand} here.`),
`${name}: text AFTER the protected block must be brand-swapped to "${brand}"`,
);
// The protected block's OWN "Claude Code" occurrences must not have
// been swapped anywhere in the output — a stronger form of the
// verbatim-block assertion above, independent of exact block framing.
const claudeCodeOccurrencesInResult = (result.match(/Claude Code/g) || []).length;
assert.strictEqual(
claudeCodeOccurrencesInResult, 2,
`${name}: expected exactly the 2 "Claude Code" occurrences inside the protected block to survive `
+ `(got ${claudeCodeOccurrencesInResult} — outside occurrences must be brand-swapped away)`,
);
});
}
});