f7df920681f233ae0fe064ee659550bdf41ff708
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3fac6e629f |
test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73. Timeouts are sized from evidence already in the tree rather than a house default, because this wave spawns installers rather than git plumbing and an undersized bound does not catch a hang -- it manufactures CI flake, which is worse, since a flake gets re-run instead of investigated. install.test.cjs records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while another lane passed the same commit in 12.7s, so full installs are bound at 120000ms against that recorded incident. Also adds an auditable escape to the guard's timeout ceiling. The 600000ms cap was set in #3143 from partial evidence, but fragment-single-edit- propagation carries a documented, load-tested 900000ms bound on a run that chains a full build plus eight generators -- the guard would have rejected a correct timeout the moment that file left the allowlist. A value above the ceiling is now permitted only with an inline allow-spawn-timeout-ceiling marker carrying a non-empty reason. It raises the ceiling; it never waives the requirement for a bound, which is asserted directly. install-shared.cjs keeps its hand-rolled assert rather than routing through throwIfFailed: its message embeds both streams, and throwIfFailed carries only a trimmed stderr. The message now also names the outcome, so a bounded timeout reads as such across its 38 importers instead of as expected null to equal 0. * test(#3145): extract class-norm timeouts and correct the build-hooks sizing A pre-PR review found 52 copies of four class-norm timeout constants across this wave. These are not per-suite fixture bindings -- they are shared facts about how long a class of subprocess takes, derived from a recorded bench incident. That norm already moved once (60000 to 120000 after a real ETIMEDOUT), and 52 copies would have drifted the next time it moved. Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and converts the copies. A site that genuinely differs -- a real tsc compile, or regen:derived -- keeps its own local constant with its own justification. Also corrects a misclassification: scripts/build-hooks.js was sized as a build at 120000 in twelve places and 60000 in another, but it compiles and bundles nothing. Its own header says no bundling needed; it copies pre-built files and syntax-checks them with vm. Three different values bounded one script; now there is one. * test(#3145): fix red CI — lint self-match and a Windows chunk overrun Two failures on PR 3176. lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The fixture exists to prove an unrelated marker does NOT suppress the rule, so it carries that marker's literal text as test data. Split via concatenation, the same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own self-match problem. The explanatory comment needed the same treatment. The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped seven minutes before the kill, so this was an overrun rather than a slow chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs regen:derived bounded at 900000ms, which is larger than the whole chunk budget, so the chunk killer always fires first and it can never complete there. Both the test and that bound predate this change; modifying the file pulled it into the Windows targeted set and exposed it. Skipped on Windows with the reason recorded; the Linux lanes cover it. The 900000 bound and its ceiling marker are unchanged -- they are correct. * test(#3145): refresh the stale test-timings cost table The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs packs chunks by measured duration from tests/test-timings.json, and an unknown file falls back to the table's median weight -- advisory by design, but it silently underweights exactly the files that matter. Four of the failing chunk's 22 files were absent from the table, including the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s (it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s. Both were weighted as average, so the chunk's total weight read 53.68 against a budget of 60 and the packer produced a single chunk. Regenerated from a passing full-suite run, per the remedy the script itself documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since gen-test-timings.cjs replaces the table wholesale rather than merging. Proven against the real packer: the same 22 files now weigh 103.91 and split into two chunks. No logic, budget, or timeout was changed; raising a budget to make a red gate pass is not a fix. --------- Co-authored-by: sim <sim@local> |
||
|
|
628648d63a |
chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)
* fix(#2931): preserve protected regions and cap emitted per-runtime bytes Route every runtime brand swap through applyClaudeCodeBrandSwap so "Claude Code" survives verbatim inside <runtime_compatibility> regions (#2284b). The fix existed only in bin/install.js's local copies; the src/*.cts exports still used a naive replace, so binding install.js to the single source -- as this phase does for the Windsurf family -- would have silently regressed those runtimes. A table-driven parity guard now covers all nine brand-swapping converters. De-duplicate the Windsurf converter family: delete the six local copies in bin/install.js and bind the four exported ones by reference, guarded by reference-identity assertions (the ADR-1508/#1675 pattern). The two unexported helpers and an unused tool table go with them. Replace the Windsurf 12,000-byte throw with description truncation, matching the bound its sibling skill converter already applied. The throw could only fire on an ~11.7 KB frontmatter description: the largest emitted workflow is 311 bytes. Truncation makes the cap unreachable by construction and leaves 12,000 in exactly one place, eliminating the dual-surface duplication rather than testing for it. Add the emitted-byte cap gate: buildEmittedSizes captures LF- and <HOME>-normalized bytes from the walk buildParityManifest already performs, and evaluateEmittedCaps asserts them against a per-runtime cap table with dead-rule detection. buildParityManifest's return shape is deliberately unchanged -- diffEmitted compares its values with ===, so making them objects would report all 8,529 emitted paths as moved. A regression test pins the values as strings. Add a deterministic trim-safety gate over composeWithinBudget's omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity rule, replacing the model-graded eval gate the issue described. * docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract * fix(#2931): bound the windsurf command name and single-source the brand swap Review findings from the orthogonal passes, all fixed inline. The claim that removing the 12,000-byte throw left total emission "bounded by construction" was false. The #1615 regex constrains the character class but not the length, and commandName is interpolated three times into the emitted workflow: a 20,000-character name emitted 60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate, clearly-labelled size control that THROWS -- commandName is the @-ref path target, so truncating it would point the workflow at a file that does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The #1615 security regex is untouched and still runs first. 128 is generous: the longest shipped name is gsd-plan-review-convergence at 27. Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe truncation helper. It still used a UTF-16 slice(0,177) -- the exact surrogate-splitting bug the helper was written to avoid, in the very sibling the helper's comment cites as its model. Bounds are unchanged, so output is byte-identical for every shipped command (descriptions max out at 99 chars). Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting the local copy. Adding it to the .cts left two unlinked implementations of identical logic -- the drift class this change exists to remove. Verified byte-identical across eight fixtures and five sequential calls before merging, and guarded by a reference-identity assertion. Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344), add fast-check property coverage for the trim-safety contract, and use fc.pre instead of a bare return in a property callback. * test(#2931): fix three test-authoring bugs the remote matrix caught The remote runner returned 8 unique failures on 6f15cdeb8. All three causes were in the test files, not the modules under test -- local harnesses exercise the modules directly, so nothing executed the test bodies until the matrix did. `{ __proto__: [...] }` in an object literal sets the prototype instead of an own key, so the JSON round-trip erased it and the cap table never saw a reserved runtime key. The production rejection was already correct; the test could not reach it. Use a computed key. Two cap fixtures tripped orthogonal error paths rather than the paths they name: one declared windsurf in the cap table but omitted it from sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern matching nothing (a genuine dead rule). Both now include a compliant artifact so the intended branch is what is asserted. The dead-rule and unknown-runtime contracts are deliberate and unchanged. `const { root } = makeSyntheticConfig({ ... `${root}` })` referenced `root` from inside its own initializer -- a temporal dead zone error. makeSyntheticConfig now optionally takes a (root) => files factory. Also raise the npm pack --dry-run bound 60s -> 120s in the shipped- scripts packaging test. That failure is NOT from this branch: the file is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s here, and the run recorded 60,637ms against a 60,000ms bound -- a timeout under 28,948-test parallel contention, not a slowdown. Fixed rather than deferred because a bound that tight is fragile regardless of which branch trips it. * chore(#2931): backfill changeset pr number to 2984 --------- Co-authored-by: sim <sim@local> |