682eaae3f047280d69ce3da75fc35afafeca02ae
63 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7c649a9970 |
fix(#3585): close raw-git bypasses of the commit_docs gate (#3590)
* test(#3585): repo-wide guard for unguarded .planning/ git add Replaces the two-file #1783 scan, which required .planning/ on the git add line and so was structurally blind to fast.md's `git add -A` and to new-milestone.md (never scanned). Extracts the shell tokenizer, comment-position rule and gsd-scan-ignore marker from the #2269 guard into tests/helpers/shipped-command-scan.cjs so both guards consume one implementation. Commit-specific logic stays in commit-files-pathspec.test.cjs; every pre-existing test there passes unedited. Fails RED on five sites: fast.md:58, new-milestone.md:262, spec-phase.md:480, eval-review.md:148, ai-integration-phase.md:263. The last three carry a markdown prose conditional outside the bash block it claims to guard. * fix(#3585): close raw-git bypasses of the commit_docs gate Five shipped workflow steps staged .planning/ with raw git. Two had no check at all; three had a markdown prose conditional sitting outside the bash block it claimed to guard, so the block ran unconditionally. spec-phase, eval-review and ai-integration-phase now route through the gsd_run query commit seam, which performs the commit_docs and gitignore checks internally and returns a skipped envelope -- this deletes the raw git pair rather than wrapping it. new-milestone stages directories for a later commit and cannot use the seam, so it takes the executable guard form, fail-open on a tooling error. fast writes no planning artifacts and has no gsd_run in scope at that point, so it excludes .planning via pathspec instead of reading config. Guard now reports 0 offenders. * test(#3585): pin skipped_gitignored to production behavior COMMIT_REASON was a test-local frozen enum joined to production only by a hand-maintained keep-in-sync comment -- the Generative Fix Divergence class, whose required remedy is a parity assertion. B1-B3 already pinned SKIPPED_COMMIT_DOCS_FALSE. SKIPPED_GITIGNORED was pinned by nothing: production could rename it and every test still passed. G1-G3 drive the gitignore auto-detect path and assert the canonical reason. The fixture must OMIT .planning/config.json entirely -- with config.json present the loader resolves commit_docs to false first and cmdCommit returns skipped_commit_docs_false, never reaching its own isGitIgnored branch. * docs(#3585): document the planning commit gate and its guard CONTEXT.md had zero commit_docs entries. Adds a Planning Commit Gate glossary entry covering the resolution chain, the typed skip envelope, the measured ordering of the two reason codes, and why the gate is enforceable only as a text guard. CONTRIBUTING.md gains the contributor rule for the new guard, with the prose-is-not-a-guard example that caused three of the five defects. * fix(#3585): address review findings in the planning-add guard Spec review (blocker): fast.md excluded .planning unconditionally, changing behavior for commit_docs=true users and violating epic AC4. Now gated -- the launcher preamble was MOVED from log_to_state into the commit block rather than copied, so gsd_run is in scope for +4 lines instead of +4KB, and the else branch is byte-identical to the previous git add -A. Security review (major): git -C <dir> add was a false negative because the flag-skip loop never modelled flags that consume a separate value. Fixed for -C/-c/--git-dir/--work-tree/--namespace. The fail-closed rule now also covers $(...) substitution args and --pathspec-from-file, which were opaque in the same way $VAR is. git commit -a/-am is now classified as reaching, since it stages every tracked modification. Self-review: isSkippable treated any NAME= token as a skippable prefix, so V=$(git add -A) escaped -- the exact divergence the shared-helper extraction existed to prevent. Adopted the sibling predicate verbatim. eval, xargs, one-line function bodies and line-continuation remain blind and are now enumerated as declared limits in the guard docblock and CONTRIBUTING. The ifDepth clamp is defensive only: a 200k-case differential fuzz found no reproducing input, so its test is labeled a pin, not a failing-first test. * test(#3585): acknowledge emitted growth in three workflow files emitted-attribution has two arms: hash attribution AND per-file growth. The growth arm needs an acknowledgment even when every moved byte is attributable to the diff, which is why the first remote run went red on it. fast.md +417: the launcher preamble moved into the commit block so gsd_run is in scope for the commit_docs guard, plus the guard itself. new-milestone.md +281: the executable guard plus one line recording that the unstaged archive move is deliberate. spec-phase.md +21: reworded prose describing the skipped envelope. eval-review.md and ai-integration-phase.md shrank; no entry needed. * test(#3585): drop duplicate spec-phase ack, shrink its prose instead The base already acknowledges spec-phase.md (from #2733), and two ack sources may never name the same path. But a base-side ack is SPENT -- it cannot clear new growth -- so the two gates were in direct conflict: attribution wanted an ack, the ack lint forbade one. Resolved by removing the growth rather than the conflict. spec-phase.md's +21 was purely a prose reword; rewritten shorter, the file now shrinks 36 bytes against base and needs no acknowledgment at all. fast.md and new-milestone.md have no base ack and keep theirs. * chore(#3585): backfill changeset pr number to 3590 --------- Co-authored-by: sim <sim@local> |
||
|
|
9448736872 |
fix(#3547): exercise the real global config-home shape in the install harness (#3567)
* test(#3547): failing-first regression for collapsed global install shape * fix(#3547): exercise the real global config-home shape in the install harness * test(#3547): align ripple suites with the real global install shape * test(#3547): update stale collapsed-shape pins in provenance and migration suites * fix(#3547): bump emitted-baseline schema version for the real install shape --------- Co-authored-by: sim <sim@local> |
||
|
|
147856040b |
fix(#2873): close review findings across fences, sanitizer and docs
Isolated security review found resolveSpecRootReference's fence tracker toggled on any delimiter, so a backtick fence could be closed by a tilde one and an include in the gap was rewritten inside a code block. Fixed by reusing scanFencedBlocks - the canonical engine already behind stripFencedCode and extractFencedBlock - rather than carrying a fourth copy of fence detection, which also closes the duplication the standards review flagged. sanitizeForRender now strips combining marks and zero-width characters alongside the ANSI, control and bidi classes it already handled. Adds the C, E and F matrix rows the spec review found missing, including installer-level coverage that spawns the real install rather than calling the report builder. Ships the how-to, the reference and command docs in five locales, the changeset, the inventory and glossary entries, and regenerates health.md for the new W028 rule. Refs #2873 |
||
|
|
dc3c81e93d |
chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both suites fail with MODULE_NOT_FOUND, which is the intended RED. Locks the measured behavior rather than the assumed behavior: RegExp.escape hex-escapes the leading character of nearly every string ("abc" -> "\x61bc"), so the suite asserts match-equivalence against an inlined historical oracle (the implementation being deleted) rather than byte-equivalence of pattern text — 200 seeded fast-check runs plus a fixed corpus, 0 mismatches. Also locks the latent character-class range bug this phase fixes as a side effect: a hyphen-bearing value interpolated into [...] currently forms a real range and matches an unintended character; post-migration it must not. * chore(#3412): src/pattern.cts owns runtime-value regex construction Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam delegating to the built-in RegExp.escape, deletes every hand-rolled copy, and raises the Node floor to the Active LTS line. The census was low, three times over. ADR-3212 counted 10 copies; a graph query found 12; the new lint rule — once live — found 27 more. The difference is that the census counted named helper FUNCTIONS while the rule counts the escape SHAPE, so inline .replace(<class>, '\$&') copies were never in scope. ADR §1's actual requirement is that no module outside the seam escapes a value for regex use, so all of them are, and CLAUDE.md's no-defer rule makes them this change's work. Fourth consecutive epic here whose copy count was low — the argument for ADR-3180 Amendment 3's "state N found by the guard" rule. Also corrected mid-implementation: the survey reported phase-id.cts's escapeRegex had 0 external importers. It had 8 production importers, making its removal a public-surface change to an ADR-2121-owned module and requiring an update to that ADR's locked-surface test. Blast radius revised Medium-High -> High. RegExp.escape is match-equivalent but NOT text-equivalent: it hex-escapes the leading char of nearly every string ("abc" -> "\x61bc"). Equivalence is proven by a seeded fast-check property test against the deleted implementation as oracle. It also fixes a latent bug: a hyphen-bearing value interpolated into a character class previously formed a real range and matched an unintended character. Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines, .nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate `required-tests` context is unchanged and no job was added or removed, so branch protection cannot be orphaned by the dropped lanes. Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with structural provenance for reviewed pattern-fragment constants rather than a name heuristic) plus a whole-tree companion guard covering the directories ESLint's globs miss. * fix(#3412): close the _SOURCE guard evasion, correct two false claims Three findings from the orthogonal review pass, all fixed. 1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier- name matching with no binding check, so `new RegExp(userInput_SOURCE)` — a function parameter — sailed past the guard. That is the same rename-evasion class issue #3410 documents, reopened by the very fallback meant to complement the structural check. Now bound to the identifier's actual binding kind: import, require-derived const, or module-scope const; parameters, `let`/`var`, and unresolvable bindings fail closed. Four RuleTester cases cover the evasion and prove the legitimate cross-module case still passes. 2. src/pattern.cts's own header carried the stale pre-correction counts (12 copies / 17 call sites) while CONTEXT.md and the design doc carried the corrected ones (~39 / ~44) — a self-contradiction inside the PR whose entire purpose is deleting divergent copies. Rewritten, preserving the durable lesson: a named-function census cannot see inline copies; only a shape-matching guard can. 3. The claim that all deleted copies threw TypeError on non-string was false. phase-id.cts's copy — the one with 8 external importers — did String(value).replace(...) and never threw. The seam's locked signature does not coerce, so this is a real, now-disclosed behavior change rather than the pure preservation the tests asserted. Audited all 32 invocations across the 8 importers and 6 in-file callers: every one is safe by construction (upstream truthy guard or a string-producing derivation), verified by runtime probe against the compiled modules rather than by TS compilation, which cannot see a runtime undefined. Corrected the false claim in both the test comment and the design doc, and added it to Known limits. * docs(#3412): add Changed changeset for the Node 24 floor The only user-visible break in this phase. The escape-behavior change is internal and match-equivalent, so it carries no user-facing note. * fix(#3412): resolve the seam's require graph in script fixtures and packaging Checkpoint 2 came back red with 90 failures on the node24 lane. Three distinct defects, all introduced by routing scripts/ through the new pattern seam, none reproducible by any local gate: 1. ~82 failures — tests/adr-index-gate.test.cjs and tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an mkdtemp fixture and spawn it there (necessary: those scripts resolve their scan root from __dirname/.., so running the real script would scan the real repo). Each harness hand-listed the dependencies to copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to gen-adr-index.cjs made both lists silently incomplete -> MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON' failures from the same crash. Fixed as a class, not an instance: new tests/helpers/copy-script- fixture.cjs walks a script's transitive static relative-require graph and copies it, so dependencies are derived and never re-declared. It throws (naming the unbuilt artifact) instead of letting the child die with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host- contract, sync-runtime-launcher. 2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so the new scripts/lint-no-adhoc-regex-escape.cjs would be MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from the tarball, matching the existing precedent for gen-emitted- baseline.cjs, which is excluded for the identical reason, and locked with a test modeled on that one. Confirmed against a real npm pack: 890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs present (so the other four scripts' requires are legitimate). 3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to the retired hand-rolled escaper but NOT text-equivalent: it hex- escapes the leading character and all hyphens ('0*\x329', '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match decisions across all three real interpolation prefixes, zero divergence. Those tests now compile each source into the same heading regex src/roadmap.cts's searchPhaseInContent builds and assert what matches and what does not, including the 'i'-flag canonicalization the hex escape has to preserve. Re-pinning the new literals would have rebuilt the same brittleness one layer down. Adds a test for the property the escape exists for: a dot in '1.2' must not act as a wildcard. Also shares one definition of 'a require' between the packaging guard and the fixture copier, so the two cannot disagree about what they scan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): refuse to copy a fixture dependency outside the fixture root copyScriptWithDeps resolved each relative require and joined the repo-relative result onto fixtureRoot. A require resolving OUTSIDE the repo yields a '../'-prefixed relative path, so path.join climbed out of the fixture and wrote into the surrounding temp dir (verified: repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd). No script in the tree does this today, so this closes an available escape rather than an active one. Refuses via the existing unresolved- require path so the failure names the offending specifier. Covered by a negative proof that the guard fires and that nothing lands outside the fixture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract Applies all findings from the second orthogonal review round, re-run because real code changed after round 1. HIGH (security) — extractRequires stripped BLOCK comments before LINE comments, so a '//' comment containing '/*' opened a phantom block comment, and a '//' inside a string literal truncated the line. Both hid real requires: 'const u="http://x"; require("./real.cjs")' returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were invisible. Replaced with a real AST parse via espree. This is ADR-3212's own Decision 4 — tokenizer-first for stateful grammars — applied to the case it describes; comment/string/regex nesting is exactly such a grammar, which is why the regex version was wrong. The function was moved byte-identical out of the #2858 packaging guard, so the bug PRE-DATES this branch and has been a live blind spot there: a shipped script could have required an unshipped path undetected. Fixing it makes that guard strictly stronger than on next. espree is promoted from a transitive eslint dependency to an explicit devDependency rather than relying on hoisting. The script parse attempt sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a function, making a top-level return legal — scripts/check-coverage-gate .cjs relies on it, and without the flag the guard throws on a file it is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a real npm pack, so the exact extractor does not newly fail the guard. MEDIUM (security) — the repo-containment check guarded dependencies but not the entry path. One escapesContainment predicate now guards both. LOW (security) — containment was lexical while fs follows symlinks, and a directory symlink could mint a fresh dedupe key per level. realpath now resolves both repoRoot and each dependency before the decision, and the realpath-derived path is the dedupe key. Destination layout still uses the original repo-relative path, so copied trees are unchanged. MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests lost the foreign-prefix contract: every assertion was satisfied by an impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599 bug class the exact-source prevents. The literal assertions it replaced were catching this. Now asserts the compiled regex REJECTS a different prefix with the same number. MAJOR (standards) — the test hand-duplicated production's heading regex with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed the parallel surface instead of policing it: src/roadmap.cts exports buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports it. Byte-identical .source and .flags verified for both escaped forms. MINOR — '..foo' no longer false-flagged as an escape; the inverted spurious-vs-missing doc claim corrected; the dead allow-test-rule header removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3412): backfill changeset pr number to 3416 * fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision Two CI failures on PR #3416, both in code this branch added. CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a bracket run be consumed EITHER by the character-class branch OR one character at a time by the trailing catch-all, so a failing match explored both parses of every pair. Measured on the real regex: n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script scans repo source, so a file with a long bracket run after '.replace(/' would hang CI outright — a guard against undisciplined pattern construction was itself the worst pattern in the diff. Fixed the way ADR-3212 already prescribes: the catch-all branch now excludes '[' and ']' so a bracket can only be consumed by the class branch (this is what makes it linear), and every quantifier is bounded (the locked bounded-quantifiers decision) as a second line of defense. Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the constant: a regex literal with a BARE unescaped ']' outside a class is no longer matched by this backstop. No census shape has that form, and the AST rule remains the primary detector. Verified the guard did not go blind doing it: a real census-shape violation is still reported, and an allow-adhoc-regex-escape suppression comment is still honored. Regression test drives the exported findViolations on a 2000-repetition adversarial input and asserts the RESULT. It makes no wall-clock assertion — elapsed-time tests are forbidden — so a regression surfaces as a harness timeout, which is the correct signal. Prompt injection scan — 'must not act as a regex wildcard' in a test comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if| my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a whole test file over one phrase would blunt the scanner permanently, and the comment has nothing to do with injection. Neither failure was reachable from the remote runner — CodeQL and the injection scan are not in that matrix, so the sha it passed was green and still wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7a7bf19fc1 |
enhance(#2872): record scope and runtime in the install manifest (#3323)
* enhance(#2872): record scope and runtime in the install manifest gsd-file-manifest.json gains manifestVersion, runtime and scope, and a new read-only Installed Surface Resolver Module reads both install scopes for a runtime in one call -- the first code path in the repo that does. Phase 3 of epic #2866 (ADR-2866). Blocks Phase 4 (#2873), which resolves #2218: the resolver's shadowedBy field is that defect expressed as a value for the first time. It ships computed-and-unread here. Installed-ness is decided by manifest PRESENCE, never by the new fields, so a manifest written by an older GSD stays fully functional and no user needs to reinstall. Recorded runtime/scope are corroboration: a disagreement with the probed config dir is reported as declaredScopeMatchesProbe: false, never silently corrected. readInstallManifest is widened additively -- version/timestamp/mode/files keep their exact names, types and meanings for all four existing callers. manifestVersion is a new field rather than a reinterpretation of version, which holds the package version and is read by the golden-parity fixtures. Stems are derived from the installed manifest's own file keys, the inverse of Phase 2's filename composition, guarded by a fast-check round-trip property plus a kebab-case charset check so a crafted manifest key cannot put a traversal segment, control character or ANSI escape into a trigger that Phase 4 renders back to the user. Also fixes two defects found while working: - bin/install.js hardcoded manifestVersion: 2 while the reader owned MANIFEST_SCHEMA_VERSION = 2. Now single-sourced, with a parity test. - docs/installer-migrations.md documented an install-state schema of five snake_case fields that have never been written; InstallState has only ever been { schemaVersion, appliedMigrations }. Corrected with a dated note. Verification runs on the remote runner. * fix(#2872): fold review findings from three independent engines Standards axis: - convert the manifest-schema suite from a hybrid setup(t) closure to beforeEach/afterEach (CONTRIBUTING.md:319-354 Pattern 1). The hybrid was neither approved pattern and a new test forgetting the call got no warning. - SCOPE_ORDER was declared twice with no parity test -- this repo's recorded generative-fix-divergence class. Give the ordering one owner: install-scope exports it frozen, the layout module and the resolver both import it, and a test locks it against scopeRank so the constant and the ranks cannot drift. - drop the defaultReadManifest passthrough (Middle Man). Spec axis: - add the VOLATILE_FILES exclusion test and source comment the acceptance table promised and did not deliver. gsd-file-manifest.json stays excluded: the new fields are deterministic, but timestamp -- the original reason -- is unchanged. Security axis: - bound the reported manifest runtime at 64 chars, matching the truncatePostureValue convention already used in this subsystem. It reached declaredRuntime unbounded while the adjacent stems were gated by SAFE_STEM; an inconsistent posture on the same attacker-influenceable document. The charset stays ungated on purpose -- declaredRuntimeMatchesProbe needs to see the real value -- so Phase 4 must sanitize before rendering, recorded in the design's Known limits. Both new parity tests were verified to FAIL when the two sides are made to disagree, then pass again on revert. Verification runs on the remote runner. * chore(#2872): backfill changeset pr number to 3323 * fix(#2872): give git fixture construction its own timeout class PR #3323's full test (windows-latest, 22, shard 2/3) failed with gitOrThrow: 'git init' failed -- outcome=timed_out exitCode=null gitOrThrow: 'git commit --allow-empty' failed -- outcome=timed_out from drift-detection.test.cjs's beforeEach, a file this branch never touched. Every other lane passed the same commit, including windows-latest node 24 on all three shards, and next is green. Root cause is a bound sized for the wrong class. DEFAULT_GIT_TIMEOUT_MS is 15000 and its own comment scopes it to plumbing READS -- rev-parse, branch, log -- against an existing repo. createFixture uses it for six sequential repo-CONSTRUCTION spawns: init, three config writes, add -A, commit. init and commit each write dozens of files, and on Windows every spawn is Defender-scanned. Sibling tests in the failing block took 15.6-22.0s against a 15000ms bound. This repo already diagnosed this exact shape once: timeouts.cjs's HOOK_FANOUT_TIMEOUT_MS records PR #3285 failing in the SAME job with the SAME outcome=timed_out exitCode=null signature at the SAME bound while every other lane passed, and concludes 'a bound sized for the wrong class, not a slow machine'. It was fixed by splitting out a heavier class-norm at 60000. Same remedy here: GIT_FIXTURE_TIMEOUT_MS = 60000, 4x the bound that failed and half INSTALL_TIMEOUT_MS. DEFAULT_GIT_TIMEOUT_MS deliberately stays at 15000 -- a blanket raise would stop a genuinely hung plumbing read from surfacing fast. Verified the value reaches the spawn rather than being an ignored option: spawnSync was monkeypatched before requiring the fixture module, and all six git construction calls were captured carrying timeout: 60000. This branch's two new test files shift shard composition, which is how a pre-existing fragility landed in the heaviest shard on the slowest lane. Fixed here rather than deferred, per the no-defer rule. Verification runs on the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
c28134ab39 |
fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host Adds local/no-duplicate-fold-marker, an AST rule that reports the second and every subsequent `folded:<name>` marker in a host file, plus RuleTester cases and a tree-wide regression assertion. Failing-first on purpose: the rule is registered at error and the 25 duplicated regions are still present, so eslint and the new tree-wide test are RED. The deletions land in the next commit. The marker key is the whitespace-delimited token after `folded:` — not the issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and feat-443-effort-fast-mode are two distinct folded suites. Refs #3271 * fix(#3271): delete 25 duplicated folded suites from three install hosts Three consolidated install suites each carried a verbatim second copy of a contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and green — each duplicated block registered and ran twice on every lane. tests/install.test.cjs 5981-9937 (3957 lines, 18 blocks) tests/install-minimal-hooks.test.cjs 2734-4015 (1282 lines, 5 blocks) tests/install-write-confinement.test.cjs 1754-2321 ( 568 lines, 2 blocks) Introduced by |
||
|
|
2e2b8ba4a7 |
enhance(#2704): resolve documentation links and compare H1 status brackets in the ADR gate (#3266)
* test(#2704): failing-first coverage for ADR link resolution and H1 status brackets Binds the gate to two assertions it does not yet make: every relative markdown link under docs/adr/ must resolve, and an H1 trailing status bracket must agree with the Status: field instead of being silently stripped. Covers all 51 rows of the phase test matrix across two altitudes - the pure extractLinks/maskCode IR for fence and inline-code-span boundaries, hostile input and the fast-check totality properties, and the real CLI verdict for the end-to-end classes. Includes the DEFECT.GENERATIVE-FIX parity test that iterates the exported STATUSES array so a sixth status is covered the day it is added. * feat(#2704): resolve ADR documentation links and compare H1 status brackets The ADR gate validated naming, relation symmetry and index freshness but never resolved a link target, and it stripped an ADR's trailing H1 status bracket for display rather than comparing it against that ADR's own Status: field. Both classes were structurally invisible: #2691 found five dangling references by manual audit roughly a year after they were introduced, one of which reached the published npm payload, while CI reported green throughout. Both are now assertions on the same --check path, using only node:fs and node:path - no dependency and no subprocess. Fenced blocks and inline code spans are masked before scanning, because markdown does not render a link inside code. That is not a policy choice: the corpus contains exactly two such sequences today and both are ordinary JavaScript. Masking preserves length and column positions so findings still name a real line. Resolution is case-exact on every platform - a link that resolves only through macOS or Windows case-folding still 404s on github.com and still fails the Linux lane - and a destination resolving outside the repository is reported before any filesystem call is made. Also single-sources two duplicated surfaces this change would otherwise have extended: the H1 bracket vocabulary (a second hand-written copy of STATUSES with nothing asserting agreement, a DEFECT.GENERATIVE-FIX instance) and the docs/adr directory traversal. Two tests added by #2691 that reimplemented link resolution and bracket comparison inside the test file are removed for the same reason; the corpus assertion is now made by running the real gate against the real corpus. * fix(#2704): reject symlinks that leave the repository and linearize code masking Four defects from the isolated adversarial security review, plus one it noted. BLOCKER - a symlink defeated path containment. path.relative(ROOT, abs) is purely lexical, but the case-exact walk then calls readdirSync, which follows symlinks at the OS level: a contributor-committed docs/adr/x -> /etc together with a link through it passed containment and listed the real external directory, and a wrong-case probe echoed a real external filename through the "Did you mean" hint into publicly-readable fork-PR logs. Every segment is now lstat'd before descent; a symlink is realpathed and re-checked against realpath(ROOT) - realpath on both sides, so a root under /var does not produce false escapes - and an escape emits no hint and reads nothing further. The same rule now governs which FILES are read: an ADR entry that is a symlink out of the repository is excluded and reported rather than parsed, closing the vector this change had widened by newly reading README.md, naming-violation files, and full bodies rather than only header fields. MAJOR - inline-span masking rescanned the line remainder per backtick run, roughly O(n^1.6) on adversarial input: 1.76s for an 800KB line. Rewritten as a single linear pass pairing runs through forward-only per-length cursors. Same input now takes 3.31ms, with behavior unchanged. MINOR - an unreadable or broken entry threw, and the generic handler wrote a raw stack trace carrying absolute CI paths to stderr. The scan is now fault-tolerant and reports excluded entries as ordinary violations. The status vocabulary is escaped before being interpolated into a dynamic RegExp - defence-in-depth, not a live bug. The containment predicate had reached three hand-written copies while fixing this; it is now the single escapesRoot() helper used by all four call sites. * feat(#2704): add a --json report so the gate's tests assert on typed values Maintainer-directed addition. CONTRIBUTING.md's "Prohibited: Raw Text Matching on Test Outputs" requires that a system under test producing text also expose a structured intermediate representation, and that tests assert on that IR rather than on rendered prose. This gate had no such surface, so its verdict tests matched on stderr. --json runs exactly the same validation as --check and writes a report to stdout with the same exit code, following the frozen-REASON-enum pattern already used by verify-reapply-patches.cjs. Every violation carries a stable reason code plus the fields a consumer needs, so nothing has to pattern-match an error message. Adding a reason stays three coordinated changes - the enum, the emitting site, and the test locking Object.keys(REASON).sort(). The human output is unchanged, deliberately: a large pre-existing suite asserts on it and migrating that is not this PR's concern. Verified by running the pre-change and post-change scripts against an identical violating corpus and diffing their stderr - character-for-character identical. This PR's own verdict tests now assert on parsed --json. Absence checks improve the most: "no bracket violation" is now a reason-code predicate rather than a negative regex over prose, which could pass for the wrong reason. The security assertions were strengthened rather than translated - no leaked filename may appear in ANY field of the serialized report. Unknown flags are now rejected instead of silently falling through to printing the index. * test(#2704): fix the status-parity fixture and guard hooks/dist before overlay builds Two failures from the matrix run of 79b29909. The status-parity fixture was mine. It built, per status token, an ADR whose H1 bracket and Status field both carried that token - but Superseded carries an obligation beyond the bracket: it must name its successor as a file link and be symmetric with it. The fixture declared a bare Superseded, tripped that unrelated invariant, and the test reported a bracket-parity failure for a reason that had nothing to do with bracket parity. The fixture now satisfies each token's own obligations in both the agreeing and contradicting corpora, derived from the status actually declared rather than special-cased on one name, so a future token carrying obligations is handled rather than silently skipped. The second failure was not mine but is fixed here rather than deferred. mcp-catalog-parity.install.test.cjs hardlinks hooks/dist/* while building its overlay, but hooks/dist is a gitignored build artifact produced only by build:hooks. The suite had no guard, so it passed only when some other suite happened to build it first - an execution-order dependency, which is why it failed on node22 and passed on node24 for identical code. install.test.cjs already documents this exact hazard and guards it. Six behaviorally identical copies of that guard existed across three files. Rather than add a seventh, they are now one canonical tests/helpers/hooks-dist.cjs - idempotent and bounded by the shared BUILD_TIMEOUT_MS class norm - which is the same single-sourcing this PR applies to the ADR gate itself. * docs(#2704): add a how-to for contributors the ADR gate rejects The reference and explanation quadrants were covered by Lifecycle rules 5 and 6, but the task-oriented one was thin: a contributor meets this gate because it failed on their PR, under pressure, and the rules told them what is checked without telling them what to do about it. Adds the command to reproduce the CI failure locally and a message-to-remedy table covering every reason code that can be hit - unresolved target, wrong case with the did-you-mean hint, repository escape, symlinked ADR file, bracket contradiction - plus the backtick escape hatch for illustrative links and the caveat that indented code blocks are not skipped. The table is itself written in backticked inline code, so the gate skips it: the escape hatch demonstrated on the page that documents it. * chore(#2704): backfill changeset PR number pr:0 placeholder replaced with the real PR number now that #3266 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
35df0891af |
fix(#3156): sandbox HOME on raw installer spawns — the one leak no scrub reaches
The strict live-config guard this PR ships went red on CI: the suite creates $HOME/.gsd/defaults.json. Diagnosed rather than suppressed, because the guard is right — this is #2665's class arriving through the one door the scrub set is structurally unable to close. bin/install.js writeNonClaudeDefaults() (#2834) writes path.join(os.homedir(), '.gsd', 'defaults.json') for every non-Claude runtime. os.homedir() consults NO GSD variable, so: - no entry in CONFIG_LOCATION_ENV_KEYS can reach it, however the set is derived; and - blanking GSD_HOME does not reach it either -- a blank GSD_HOME falls back to exactly that homedir(). Only a sandboxed HOME contains it, and HOME is deliberately excluded from TEST_ENV_BASE because blanking it would break far more than it fixed. So the containment belongs per-spawn, which is the discipline the suite already applies by hand -- install-minimal-hooks.test.cjs:576 carries a comment naming this exact hazard for the --codex spawn, while the parameterized --${runtime} spawn 130 lines below it does not. Instance fixed, class open. Rather than add a fifth hand-synced env shape to a PR whose subject is that hand-synced copies drift, this adds ONE export -- installSpawnEnv() in tests/helpers.cjs -- and routes every raw installer spawn through it, including the shared tests/helpers/install-shared.cjs installerEnv(), which every install suite already consumes. Callers passing an explicit { HOME, USERPROFILE } are unaffected: overrides spread last. Census (measured, not reasoned): of the 119 test files that spawn bin/install.js, exactly four wrote into a sandboxed $HOME before this commit -- install, copilot-install, install-minimal-hooks, opencode-plugin-adapter -- and zero do after. Only install.test.cjs sits in CI's targeted lane, which is why ubuntu went red on one file while the macOS full lanes went red on four. Attribution: the leak reproduces unchanged at upstream/next itself, so the defect is base-owned and pre-existing; only the detector is new. The guard found a real leak on next within one run. No new failures: the surviving names under a sandboxed HOME (folded:enh-2380-sync-skills, getGlobalConfigDir (Copilot)) fail at base too, and base additionally fails folded:bug-3288-model-catalog-install-path, which this tree does not. |
||
|
|
9faacc0c15 |
test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local> |
||
|
|
cbd180c5cd |
test(#3147): bound the lint/changeset/docs cluster onto the process seam (#3181)
* test(#3147): bound the lint/changeset/docs cluster onto the process seam Migrates 69 unbounded sync spawn sites across 24 files. Allowlist 73 to 49. Two shared helpers move: tests/helpers/graphify.cjs (6 importing suites) and tests/fixtures/index.cjs, whose three quoted-argument shell strings became single argv elements rather than whitespace splits. changeset-lint's throw-native git() helper routes to gitOrThrow; migrating it to bare runGit would have silently swallowed a failure that is loud today. ingest-docs goes the other way -- its catch never rethrew, it degraded failure into data every call site asserts on, so throwIfFailed would have thrown where the original returned. The design doc said otherwise and was corrected. tsconfig-noemit runs a real tsc --noEmit and takes a bespoke 180000ms per the ensure-runtime-build precedent, not the 30000ms build-hooks norm -- that norm is for a file copy, and sizing against a label rather than the work is the same error in the opposite direction. * test(#3147): add toLegacyResult and settle review findings The seam exposed a throwing adapter (throwIfFailed) but no non-throwing one, so eight files independently re-derived the same unwrap back to the legacy {status, stdout, stderr} shape. That is the third time this epic produced N copies of one mechanism -- seven throw wrappers in Wave 1, fifty-two timeout constants in Wave 2, eight result adapters here. The pattern is that whenever the seam does not expose a mechanism, every suite re-derives it. toLegacyResult now sits beside throwIfFailed, with its own tests. Two sites are deliberately NOT converted: changeset-cli's runRender and runRenderIn return {status, report, stderr} from parsed JSON and never a raw stdout, so they are a different shape family. lint-legacy-dir-name keeps its local GUARD_TIMEOUT_MS: 30000 matches the build norm numerically but bounds a lint probe, not hooks bundling, and importing it would encode a coincidence as a relationship. --------- Co-authored-by: sim <sim@local> |
||
|
|
3fac6e629f |
test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73. Timeouts are sized from evidence already in the tree rather than a house default, because this wave spawns installers rather than git plumbing and an undersized bound does not catch a hang -- it manufactures CI flake, which is worse, since a flake gets re-run instead of investigated. install.test.cjs records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while another lane passed the same commit in 12.7s, so full installs are bound at 120000ms against that recorded incident. Also adds an auditable escape to the guard's timeout ceiling. The 600000ms cap was set in #3143 from partial evidence, but fragment-single-edit- propagation carries a documented, load-tested 900000ms bound on a run that chains a full build plus eight generators -- the guard would have rejected a correct timeout the moment that file left the allowlist. A value above the ceiling is now permitted only with an inline allow-spawn-timeout-ceiling marker carrying a non-empty reason. It raises the ceiling; it never waives the requirement for a bound, which is asserted directly. install-shared.cjs keeps its hand-rolled assert rather than routing through throwIfFailed: its message embeds both streams, and throwIfFailed carries only a trimmed stderr. The message now also names the outcome, so a bounded timeout reads as such across its 38 importers instead of as expected null to equal 0. * test(#3145): extract class-norm timeouts and correct the build-hooks sizing A pre-PR review found 52 copies of four class-norm timeout constants across this wave. These are not per-suite fixture bindings -- they are shared facts about how long a class of subprocess takes, derived from a recorded bench incident. That norm already moved once (60000 to 120000 after a real ETIMEDOUT), and 52 copies would have drifted the next time it moved. Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and converts the copies. A site that genuinely differs -- a real tsc compile, or regen:derived -- keeps its own local constant with its own justification. Also corrects a misclassification: scripts/build-hooks.js was sized as a build at 120000 in twelve places and 60000 in another, but it compiles and bundles nothing. Its own header says no bundling needed; it copies pre-built files and syntax-checks them with vm. Three different values bounded one script; now there is one. * test(#3145): fix red CI — lint self-match and a Windows chunk overrun Two failures on PR 3176. lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The fixture exists to prove an unrelated marker does NOT suppress the rule, so it carries that marker's literal text as test data. Split via concatenation, the same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own self-match problem. The explanatory comment needed the same treatment. The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped seven minutes before the kill, so this was an overrun rather than a slow chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs regen:derived bounded at 900000ms, which is larger than the whole chunk budget, so the chunk killer always fires first and it can never complete there. Both the test and that bound predate this change; modifying the file pulled it into the Windows targeted set and exposed it. Skipped on Windows with the reason recorded; the Linux lanes cover it. The 900000 bound and its ceiling marker are unchanged -- they are correct. * test(#3145): refresh the stale test-timings cost table The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs packs chunks by measured duration from tests/test-timings.json, and an unknown file falls back to the table's median weight -- advisory by design, but it silently underweights exactly the files that matter. Four of the failing chunk's 22 files were absent from the table, including the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s (it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s. Both were weighted as average, so the chunk's total weight read 53.68 against a budget of 60 and the packer produced a single chunk. Regenerated from a passing full-suite run, per the remedy the script itself documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since gen-test-timings.cjs replaces the table wholesale rather than merging. Proven against the real packer: the same 22 files now weigh 103.91 and split into two chunks. No logic, budget, or timeout was changed; raising a budget to make a red gate pass is not a fix. --------- Co-authored-by: sim <sim@local> |
||
|
|
27aa40f65e |
fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir pi reserves <configDir>/hooks as its deprecated extension location and warns on every startup when it exists. Assert a pi install stages the shared hook bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/. Also adds pi to the local-scope dir table in install-shared.cjs: pi was in RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved path.join(root, undefined) and no local pi install could be exercised. Fails before the fix. Verified via the remote runner. * fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir pi reserves <configDir>/hooks as its now-deprecated extension location and warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs() guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged its shared hook bundle exactly there, and pi's advised remediation (move it to extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's extension auto-discovery. The bundle directory name is now runtime-descriptor-driven: hostBehaviors .sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path segment — separators, dot-only segments, trailing dots, absolute paths, NUL, and Windows reserved device names all fall back to the default, because the value is joined onto a user's config root and written to. Renamed in place rather than relocated: hook scripts resolve siblings via __dirname/.., so a depth change would silently break them. - install / uninstall / manifest sites all read the resolved name - pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded trees still resolve; the never-throws contract is preserved - new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new non-recursive remove-empty-dir engine primitive (rmdirSync only, symlink-refusing, containment-guarded); ADR-0008 amended accordingly - fixes two latent name-dependencies the rename exposed: the stale-hook scan and the injection scanner's self-exclusion both hardcoded 'hooks' Verified on the remote runner. Closes #3023 * fix(#3023): close review findings and align emitted provenance with the rename Adversarial review found two defects, and the remote runner found four failure clusters. All fixed here. Review BLOCKER — detect-custom-files was blind to the renamed bundle. GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the whole gsd-hooks/ tree was invisible to the custom-file scan and user-added files there were never backed up before the next update's clean-install wipe. The dir set now resolves via the .gsd-runtime marker plus the shipped capability registry (never bin/install.js, which is not shipped into installed trees), and falls back to scanning every known candidate when the runtime cannot be determined — over-scanning is safe, under-scanning is the data loss. Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir accepted any directory, so an interrupted install left gsd-hooks/ winning over a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now qualifies only if it is non-empty. Remote-runner clusters: - emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped rules pointing at the same sources the existing hooks/ rules use. The table is total, so an unattributed family is a hard failure by design. - pi tests in install-minimal-hooks and the install integration suite asserted the old layout; updated to derive the dir name from the descriptor rather than hardcoding either name. - 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after() runs in registration order, cleanup was registered before mock.restoreAll(), and node22's JS rimraf calls the public fs.rmdirSync while node24's native path does not — so the EACCES stub leaked process-wide on one lane. Restore now runs first. Verified on the remote runner. * fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent (packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty configHome.env, so a user with that variable set had GSD installed where pi never looks. Added the env name; the dot-home-nested resolver already handled the override, so no resolver logic changed. Also fixes expandTilde in the shared runtime-homes resolver, found while adding that: it hardcoded os.homedir() and ignored the opts.home every caller threads, so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi) silently resolved against the real home. That is a correctness bug and a test-escape hazard — a sandboxed test asserting on a tilde override reached the developer's actual home directory. Now threaded through every branch; behavior with no injected home is unchanged. Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location moved with the rename. The provenance rules satisfy the totality gate; the differential gate needs the ack because the hook sources are byte-unchanged — only the installer's target directory moved. The two hook files this branch genuinely edits stay attributed and are not double-acked. Note on piConfig.configDir: it is read from pi's OWN installed package.json (getPackageDir walks up from pi's __dirname), alongside piConfig.name — a white-label setting for a redistributed pi fork, not a per-project user setting. Documented accordingly rather than treated as an unsupported override. Verified on the remote runner. * fix(#3023): reject blank env overrides, pin adapter/descriptor parity Three review findings, all fixed. A whitespace-only config-dir override was accepted verbatim: the guard was `if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR=' ' resolved to a literal three-space directory name instead of falling back to the descriptor default. Fixed across every env-consuming branch — dot-home, dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's. Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working. pi/gsd.cjs's probe list and the descriptor were two independent sources of truth for the bundle directory name; a future rename would have desynced them silently and left every pi hook quiet with no error. The probe list stays deliberate — it must resolve in a dev checkout and a half-upgraded tree, where the registry's answer would be wrong — so this adds the parity assertion the repo's generative-fix-divergence rule calls for: the descriptor value must be the FIRST candidate, and the default must remain present. Changeset body rewritten to cover the two later user-facing fixes it had not caught up with. Verified on the remote runner. * chore(#3023): backfill changeset PR number * fix(#3023): anchor injection-scan patterns and fix a macOS detection hole CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the same fact as a genuinely empty or absent one'. The match was the 'act as a' INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending in act tripped it (fact, impact, contract, artifact, interact, redact, abstract). My four-line CONTEXT.md edit dragged the latent false positive into this PR because the scan is diff-scoped by file but reads whole files. Anchored with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would have left the class alive for the next PR touching any file saying 'fact as a'. Auditing the rest of the list for the same class surfaced a real detection hole: the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex escape. BSD/macOS grep reads it as four literal characters, so single-quoted eval('...')/exec('...') payloads were NEVER detected there while passing on GNU-grep CI. Replaced with a literal apostrophe class. Boundaries were added only where a real word-suffix collision exists; exec, jailbreak, developer mode and the role-manipulation family were audited and deliberately left unanchored. 22 new cases cover both directions — the false positives now scan clean, and every real payload still fires, including the quote/punctuation/start-of-line boundary forms. Also builds this branch's injection test fixture at runtime instead of carrying the literal phrase, so the payload keeps its teeth without tripping the scan. Verified on the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
1d208e5af6 |
test(#3144): bound the git/worktree cluster onto the process seam (#3152)
* test(#3144): bound the git/worktree cluster onto the process seam Migrates 180 unbounded sync spawn sites across 19 files. Every previously unbounded call now carries an explicit timeout with a comment giving the number and why. The migration is not a callee swap. execSync and execFileSync throw on a non-zero exit and the seam never does, so each site was classified first: sites that rely on the throw route to gitOrThrow, and sites that already read .status to detect an EXPECTED non-zero -- an intended cherry-pick conflict, a rev-parse outside a repo driving a skip -- route to the never-throwing runGit instead, which would otherwise throw on exactly the exit being probed for. Two same-named git() helpers in worktree-cleanup.test.cjs have different return contracts, one trimmed and one raw; both are preserved rather than unified. Collapses five hand-rolled throw wrappers onto one throwIfFailed in git-fixture.cjs, which gitOrThrow now also uses so the shape cannot drift. Allowlist drops 139 to 120; BASELINE lowered to match. * test(#3144): fix pre-PR review findings Documents throwIfFailed in the CONTEXT.md glossary and CONTRIBUTING.md -- it became the shared throw mechanism without either doc naming it. Routes the sixth and seventh hand-rolled copies of the throw shape through throwIfFailed (worktree-baseref-install, worktree-safety-reap); the first consolidation missed both. Converts ci-rebase-check's 8 fixture-setup calls from unchecked runGit to gitOrThrow so a failed setup step aborts where it fails rather than surfacing later as a confusing failure against the wrong subject. Adds 12 direct unit tests for throwIfFailed, which until now was only exercised transitively. Splits verify.test.cjs's non-git grep/sed bound off GIT_TIMEOUT_MS. --------- Co-authored-by: sim <sim@local> |
||
|
|
2afe17bbdb |
test(#3143): add the no-unbounded-spawn guard and throw-preserving git fixture (#3150)
* test(#3143): add no-unbounded-spawn guard and throw-preserving git fixture Adds the ESLint rule local/no-unbounded-spawn, wired into the tests/**/*.cjs block, plus an allowlist that only ratchets down: a listed file with zero violations reports its own entry as stale. The rule resolves renamed destructures and chained requires rather than matching literal callee names -- both forms exist in the suite today and a name-only matcher leaves them permanently invisible. It resolves an options object held in a single-write const, which is what keeps process-seam.cjs, the bounded reference implementation, from flagging itself. timeout: 0 and anything above the 600000ms ceiling are rejected as only nominally bounded. Adds tests/helpers/git-fixture.cjs so a migrated execSync call site keeps its throw-on-non-zero contract; process-seam.cjs is unchanged. * test(#3143): prove the allowlist guards can actually fail Extracts the D4/D6/D7/D8 checks into pure helpers and drives each against a synthetic fixture carrying an injected violation. Without this the suite only proved that today's clean data passes, which a deleted check would also satisfy. * fix(#3143): close two ceiling and alias escapes found in review Nested arithmetic bypassed the ceiling entirely: the numeric evaluator only resolved a flat literal, so `timeout: 60 * 60 * 1000` (3600000ms, six times the ceiling) fell through to trusted and reported nothing. The evaluator now recurses through arithmetic and unary signs with a depth cap. Alias resolution was traversal-order dependent, not scope dependent: a call textually above its own require destructure saw an empty alias map and reported clean. The map is now built in a Program pre-pass. Also: an explicit timeoutMs:undefined no longer overwrites the git fixture default via spread, adds the missing seam-routed rule test, and de-duplicates the repeated try/catch in the fixture tests. --------- Co-authored-by: sim <sim@local> |
||
|
|
7e1004d89e |
test(#3056): add the in-process fault-injection adapter, and normalize execGit's shape (#3077)
* refactor(#3071): normalize execGit's call and result shape, unify ExecGitFn ExecGitFn was declared four times. Three were hand-copies of one signature and two of those were wrong: they typed exitCode as number|null when _spawnResult returns `result.status ?? 1` and can never yield null, weakened signal from NodeJS.Signals to string, and widened error from Error to unknown. Only verification.cts got it right, via `typeof execGit`. The root cause was a missing export: SpawnResultOutput was declared without `export`, so no other module could name the return type of execGit. Three authors independently hand-copied it instead. Exported now. Normalizing the type alone would have left the pressure that caused the divergence, so the function is normalized on both sides. It now ACCEPTS every call its consumers make — worktree-safety's declaration could not express an env-carrying call at all — and RETURNS every result code they need: timedOut moves into _spawnResult, so execGit, execNpm and execTool all carry it and the one extension that justified a separate type disappears. All four sites are now `typeof execGit` with nothing left to restate. timedOut reuses the existing isSpawnTimeout predicate introduced by #3050 rather than re-deriving it. That predicate checks error.code === 'ETIMEDOUT' only; the signal === 'SIGTERM' conjunct was deliberately dropped there because Windows does not reliably report SIGTERM and requiring it risks a false negative. There is no false-positive risk, and a test proves it: an externally-delivered SIGTERM leaves error null, so it is still not reported as a timeout. No dead null-checks surfaced. Every exitCode comparison in the two affected modules is === 0, !== 0 or === 128 — never a null guard — so the nullable declaration had never been written against. Closes #3071 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3056): add the in-process fault-injection adapter Adds tests/helpers/faulty-deps.cjs — makeFaultyGit() and withFaultyFs() — so a module's error branch can be driven deterministically and its degraded verdict asserted, instead of a counter-test that only proves the call did not throw. makeFaultyGit returns a value structurally assignable to `typeof execGit`, so one stub satisfies all four seams that the #3071 normalization collapsed into that single shape. A parity test drives the same stub through a real injectable entry point in each of worktree-safety, git-base-branch, worktree-base-ref and verification; it fails the moment any of them re-grows its own shape. Faults are scoped rather than global — by argv predicate and by call ordinal — because a fault adapter that faults everything looks like it works and proves nothing, and because verification.cts's two-call error handling needs to fault the second call only. Invocations are recorded so a test can assert an exact call count. The timeout fault carries error.code === 'ETIMEDOUT', and a test asserts the real isSpawnTimeout predicate matches it, so the fixture cannot drift from the production definition of a timeout. A companion test asserts an externally delivered SIGTERM with a null error is still NOT reported as a timeout — the false-positive guard for #3050's dropped conjunct. withFaultyFs restores in a finally so a throwing body still restores, patches only the named methods, and nests without clobbering an outer saved original. It never uses chmod: that no-ops under root, so the test would pass with zero coverage in root Docker/CI. The adapter is in-process via deps only. The Phase 1 process seam is documented as deliberately not a fault-injection surface — it cannot distinguish an injected timeout from a genuine bench OOM and would retry it. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3071): add changeset fragment for the execGit normalization Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3071): route the last two timeout checks through the shared predicate An isolated review found this branch had normalised the timeout verdict but left two callers still hand-rolling the fragile version of it. check-command-router's runBoundedShell computed `timedOut: r.signal === 'SIGTERM'` while the correctly-derived `r.timedOut` sat on the same result object. graphify's execGraphify branched on `result.signal === 'SIGTERM'`, with a comment asserting the very premise isSpawnTimeout exists to reject. Both fail in both directions. On Windows a genuine timeout is not reliably reported as SIGTERM, so the guard silently fails to fire — the false negative #3050 was raised for. And an externally-delivered SIGTERM is not a timeout at all, so the check also fires when it should not; isSpawnTimeout avoids that because `error` is null in that case and it keys on error.code. Both now read the derived `timedOut`, and graphify's comment states the actual rule instead of the fragile assumption. Also replaces a vacuous test: "execGitDefault now accepts env" never called execGitDefault (it is unexported), called execGit — whose signature already accepted env before this branch — and asserted only that exitCode was a number, which would pass whether or not the change under test existed. It now proves env reaches the child by asserting `git var GIT_EDITOR` returns the injected sentinel. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3071): make the graphify timeout fixture faithful to a real timeout The remote matrix failed on both Linux lanes: graphify's "returns exitCode 124 on timeout" got 1 instead of 124. The fixture was wrong, not the production change. It stubbed spawnSync as { status: null, signal: 'SIGTERM', error: undefined }. That is not a timeout. A real spawnSync timeout also sets error.code 'ETIMEDOUT'; a SIGTERM with no error is an externally delivered signal — a kill. The old `result.signal === 'SIGTERM'` check accepted it as a timeout, which is the false positive the shared predicate exists to reject, so this test was locking that bug in rather than guarding against it. The fixture now carries a real ETIMEDOUT error and all three original assertions pass unchanged. A counter-test is added alongside it: an externally delivered SIGTERM with no error must NOT be reported as a timeout. That is the assertion whose absence let the false positive live. Swept every SIGTERM/SIGKILL stub under tests/ for the same unfaithful shape. No other instance: the worktree-safety, worktree-base-ref and commit-staging fixtures already set ETIMEDOUT, and the remaining hits are either deliberate external-kill tests or feed code that never consults timedOut. Two sites keep their own signal check deliberately and are NOT changed: capability-source.cts:1301,1386 fail closed on ANY abnormal termination, which is correct — reading timedOut there would stop it failing closed on a kill and let it parse stdout from a killed process. Their reason strings, and check-latest-version.cjs:115, label any signal as "timed out", which is imprecise wording over a correct verdict, not a silent failure. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3071): backfill changeset pr number to 3077 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5fd5c81042 |
test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync, each returning a typed discriminated union { outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }. Every call is timeout-bounded; there is no unbounded path. runGsdTools becomes an adapter over the seam. Its legacy { success, output, error, exitCode } shape and retry-once-on-kill behaviour are preserved byte-identically, so none of its 136 caller files change. Outcome discrimination was corrected against probed runtime behaviour rather than assumption: a timeout and a maxBuffer overflow are identical on both status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs ENOBUFS). Overflow is therefore classified before timeout. This fixes a live defect — the previous isKilled() treated an overflow as a kill, retried it for a second full 60s run, and then reported "host OOM or scheduler contention" for a child that had merely printed too much. Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs, which brought 31 previously unlinted shared helpers under the same rules their sibling test files already obey, and fixes the 5 violations that surfaced — including a bare npm invocation without shell:true in tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now routed through the existing portable runNpm helper. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): migrate every local spawn wrapper onto the process seam Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name, parameter list, return shape and post-processing (JSON parse, ANSI strip, env sanitising, field extraction) — only the spawn mechanism changes, so no test assertion moves. The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather than a fourth primitive; it is explicit rather than inferred from the file extension, because guessing an interpreter from a path fails silently when a script's name does not match its shebang. Seven wrappers were previously unbounded and now carry an explicit timeout sized to what each actually runs, not the seam default. Two of those seven (gsd-write-guard, lint-docs-command-form) were absent from the issue's inventory entirely and were found by scanning after the migration. Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md reference section covering the three primitives, the discriminated union, and the two rules the seam enforces. Scope disclosure recorded in the phase design notes: the issue scoped three identifier names. A scan for local helpers that spawn AND return the spawn result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded direct git call sites. This change bounds 25 of those. The remaining surface is the same defect class and is NOT closed by this PR. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify an externally-killed child as KILLED, not EXITED Blocker found in this branch's own diff, independently confirmed by an isolated reviewer. A child killed by an external signal — a genuine bench OOM kill — makes spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field. The seam's "no error implies EXITED" rule therefore classified it as a clean exit, and runGsdTools returned { success: false, exitCode: 1 } without retrying. That silently defeated the #969 kill-discrimination for precisely the case it was built for: the old isKilled() fired on `signal != null`, retried once, then threw a labelled resource-starvation error. A real OOM would have been reported as an ordinary assertion failure. Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes the adapter retry on TIMED_OUT or KILLED — reproducing the old `killed || signal != null || code === 'ETIMEDOUT'` condition exactly. SPAWN_FAILED still does not retry (matching the old behaviour, where signal was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate divergence: the old code retried it because signal was SIGTERM, burning a second 60s run on a child that had merely printed too much. All five outcomes verified against the live runtime rather than assumed: SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT), >1MB stdout -> buffer_overflow (ENOBUFS). Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): address standards-review findings on this branch Three findings from the standards axis of the review, all in this branch's own diff. The CONTEXT.md glossary entry this branch introduced was already stale on the branch's own last commit: it enumerated a 4-member OUTCOME while the code had 5, because the KILLED fix did not update it. That is precisely the drift the "module changes update Domain-terms" gate exists to catch, so the entry now lists all five and explains KILLED. api-coverage-gate-e2e compared an outcome against the raw string 'exited' rather than OUTCOME.EXITED, the only such outlier; the enum is now imported and used. A sweep for the other four outcome literals found no further comparison sites. Three call sites hand the literal bash flag '-c' to the seam's first parameter, which the JSDoc described as an absolute script path. Rather than add a fourth primitive, the contract is corrected to match reality: the parameter is renamed `target` and documented as the first argv element handed to the interpreter — normally a script path, but for an interpreter invoked with an inline program it may be that interpreter's own flag. No behaviour change. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): assert the cross-platform timeout contract, not the macOS one The remote runner failed on both Linux lanes (node 22 and node 24, identical) while the same tests passed locally on macOS. Two assertions encoded a platform-specific behaviour as a cross-platform guarantee. When spawnSync times out, macOS preserves the child's partial stdout/stderr; Linux discards it and returns empty strings. Verified on node v26.5.1 both ways. The seam passes through whatever spawnSync hands it and cannot manufacture output that was discarded, so the production code was correct — the tests were wrong. Both tests now assert the guarantee the seam actually makes on every platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being strings rather than undefined or a Buffer. The partial-content assertions are retained behind an explicit process.platform === 'darwin' guard so the macOS coverage is not lost, and the first test is renamed to say what it now guarantees. This is the failure mode the remote matrix exists to catch: local macOS verification would have shipped it. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout Windows CI caught two defects the Linux matrix could not. tests/context-predicates-query.test.cjs passes a 32K-char argv value. On Windows that exceeds the argv limit and spawnSync fails with code ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise, status === null implies TIMED_OUT" — swallowed it, so the adapter retried a spawn that can never succeed and then threw the resource-starvation error. The old isKilled() returned false for that shape and returned an ordinary failure result. TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set. Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG, E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform whose timeout errno differs classified correctly, so the greedy catch-all is no longer needed. The second defect is a contract regression I introduced and had claimed otherwise. That same test asserts `typeof r.exitCode === 'number'`, and toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the assertion failed on type. The old code returned `err.status ?? 1` on every non-retried failure path. The adapter now returns 1 again for both, and the comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is retracted: the seam keeps the richer truth (exitCode null plus a distinct outcome), the legacy adapter keeps the old numeric contract its callers actually depend on. Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT -> SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL -> KILLED; clean exit -> EXITED. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8f75e27554 |
fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b780cd2dc6 |
test(#2933): prove one fragment edit reaches every emitted artifact (#3046)
Epic #1671 Phase 6 "Done when" required a maintainer-reachable proof that a single-fragment edit propagates to every emitted per-runtime artifact with no second source surface needing an edit. No test referenced that surface at all. Adds tests/fragment-single-edit-propagation.install.test.cjs (20 rows): a hard-linked overlay repo overrides exactly ONE steps/ fragment, real installers are spawned per runtime, and the emitted artifacts are asserted directly. Expected runtime sets derive from RUNTIME_META at run time, never a hardcoded count, so a new runtime cannot be silently under-covered. Six negative controls keep it from being pass-always theater. An identity-stubbed composer must make marker bytes LEAK, proving the marker-absence assertion can fail. Each derived generator whose --check is used as evidence has its own red-path control driven by an override-only edit, each asserting the generator's own typed reason enum rather than matching prose. Coverage is disclosed, not implied. REGEN_STEPS_WITHOUT_CHECK_MODE names the regen:derived steps with no read-only mode; CONTENT_EDIT_INSENSITIVE_CHECKS names gen-inventory-manifest, whose --check derives from directory listings and is structurally blind to content edits. Both constants are pinned by a test so the disclosure cannot silently rot. Assertions check sentinel PRESENCE, not whole-file byte identity: partial-wave.md embeds the runtime-launcher snippet, so emitted fragments are legitimately rewritten per runtime (windsurf -> .windsurf, qwen -> .qwen, claude -> its absolute config dir). A dedicated row now locks that behavior in. The overlay tree-diff is labelled a harness self-check, not the no-cascade proof it cannot be: the overlay is built from the checkout with the override map applied, so that diff can only ever restate the test's own fixture. Extracts buildOverlayRepo into tests/helpers/overlay-repo.cjs so both install suites share one implementation instead of diverging copies, converts that sibling's six try/finally test bodies to t.after() per CONTRIBUTING.md, and frees each per-runtime temp install eagerly so peak disk stays bounded. Refs #2933 Co-authored-by: sim <sim@local> |
||
|
|
c6ce4d1d9a |
fix(#2755): resolve the kimi hooks-TOML root per runtime (#3032)
* test(#2755): failing-first coverage for per-runtime kimi hooks root Install/uninstall filesystem-shape rows over a sandbox HOME (no permission tricks) plus resolver unit rows. Covers both uninstall directions, which is where a fix applied only to the install call site would drift. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2755): resolve the kimi hooks-TOML root per runtime resolveKimiHooksTomlDir took no runtime argument and hardcoded ~/.kimi, but both kimi and kimi-code route through the single hooksSurface=kimi-hooks-toml branch. A --kimi-code install therefore wrote its [[hooks]] block, hook bundle and CommonJS marker into Kimi CLI's config file, and a --kimi-code uninstall stripped Kimi CLI's block. Adds a runtime selector to the resolver -- kimi keeps ~/.kimi + KIMI_SHARE_DIR, kimi-code gets ~/.kimi-code + KIMI_CODE_HOME, per Kimi Code's own upstream data-locations and hooks docs -- and passes the runtime at both the install and uninstall call sites. An omitted or unrecognized runtime still resolves ~/.kimi, so the exported no-arg contract is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2755): use centralized helpers and add a divergence guard Review findings, all fixed in-PR: - The new test block reimplemented runMinimalInstall, createTempDir and toPosixPath. Extends runMinimalInstall with optional root/extraEnv instead (back-compat: every existing caller passes neither) and uses the centralized helpers, per CONTRIBUTING's Use Centralized Test Helpers rule. - Adds a parity assertion between the capability registry and the resolver: a third runtime declaring hooksSurface kimi-hooks-toml would silently inherit ~/.kimi, re-creating this very defect. The guard fires the moment those two surfaces drift. - Adds an installer-level test proving KIMI_SHARE_DIR and KIMI_CODE_HOME do not interfere when both are set, which only the resolver unit covered before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2755): track the kimi-code hooks root in the emitted-artifact gates The remote runner caught a real ripple: moving kimi-code hooks to ~/.kimi-code made 31 emitted paths unattributable and 58 emitted hashes unexplained, because three parallel surfaces keyed on the literal .kimi path. - HOOK_CONFIG_RELATIVE_PATHS excluded only .kimi/config.toml, so kimi-code's config.toml became manifest-visible; it embeds a platform-varying node-runner command and must stay out for both products. - HOOKS_ROOTS, the package.json-marker branch and the synthesized-install-metadata pattern each named .kimi only. - tests/fixtures/install-tree/kimi-code.json still recorded the old paths; regenerated via gen:install-tree. Adds the per-PR drift acknowledgment for the 58 paths whose bytes are unchanged but whose destination moved - a ripple no source diff can show, since no hook script was edited. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2755): clear production-tree security advisories The remote runner's npm-integrity gate reported 2 high advisories in the production dependency tree. My diff touches neither package.json nor package-lock.json, so these come from the base -- but a red gate is not something to wave off as pre-existing, so it is fixed here rather than deferred. Lockfile-only, semver-in-range, via npm audit fix: fast-uri 3.1.4 -> 3.1.5 (host confusion via backslash authority introducer) ip-address 10.2.0 -> 10.4.0 (three SSRF / trust-boundary bypasses) hono 4.12.31 -> 4.13.0 (moderate; reverting it traded a high for a moderate, so the full remedy is taken) npm audit now reports 0 vulnerabilities at every severity, npm ci installs clean from the updated lockfile, and the build and the kimi behavior both re-verified afterwards. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2755): backfill changeset pr numbers Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cc3ee301a7 |
fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root installSharedHooksBundle wrote `{"type":"commonjs"}` over <configRoot>/package.json unconditionally — no existence check, no merge, no backup — on every install and every /gsd-update re-install. On the 11 affected runtimes that file is often user-owned; on OpenCode and Kilo it is the documented place to declare local-plugin npm dependencies, so a user's name/type/dependencies/scripts were destroyed on each run. The uninstall path already read the file and unlinked it only on an exact content match. That asymmetry was the defect: the discipline existed in the codebase, it just was not applied on the write side. Move the marker into the directories GSD creates and fills with its own .js files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop writing the config root entirely. New src/commonjs-marker.cts owns the marker string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker, so install and uninstall cannot drift apart again. Nothing else depended on the config-root marker: package identity is baked at build time (#378/#498) and version resolution prefers gsd-core/VERSION and already tolerates a missing root package.json (#1383) — Codex has installed without one all along. A package.json in plugins/ or extensions/ is inert to plugin discovery, which globs *.{ts,js} only (see installer-migration 006). Uninstall retires the pre-fix config-root marker, so upgrading users are cleaned up on removal, and still never touches a file it did not write. * fix(#2544): point the changeset fragment at the filed PR The fragment's `pr:` field is only knowable after `gh pr create` returns. * fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the linted source), so it belongs in the ADR-457 ignore list like its siblings. Clears the lint-tests no-var failure and the repo-invariants "linted xor ignored" migration-state test. * fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root The UPGRADE 1 test still asserted the pre-#2544 marker location (~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the directory GSD itself creates — matching the updated golden-install-parity and install-tree fixtures. Also asserts the root marker is NOT written. * fix(#2544): make the CommonJS marker write path non-fatal Review round 2, Major 3 + Minor 1 + the stagedHooks nit. ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole install with a raw stack trace. Every other marker interaction in the module is best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker swallows read failures — and this was the write path, i.e. the one most likely to fail on a locked-down config dir. It now returns a new 'failed' outcome and both call sites warn and continue. Sibling found while sweeping for the same defect class: fs.mkdirSync sat OUTSIDE the try block, so an unwritable parent threw past the guard entirely. Creating the directory is the same environmental hazard as writing into it, so it moved inside. Also in this file: - The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks alone. stagedHooks is computed from the SOURCE listing before the copy loop, so it stays true when the copies land but verifyInstalled() then fails — marking a hooks/ GSD did not successfully populate claims an ownership the install did not earn. - The uninstall rmdir of the native plugin dir is gated on GSD having actually removed something from it. Hoisting it out of the adapter-exists guard (so the marker-only case could prune) had silently widened it into deleting a user-created but empty plugins/ or extensions/ dir — the same "don't touch territory GSD didn't fill" principle this issue is about, inverted. - Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the same call site that writes its replacement. That path is outside kimi's configDir, so installer-migration 007 structurally cannot reach it. * fix(#2544): retire the stale config-root marker via installer-migration 007 Review round 2, Major 1 — the PR's headline claim was false for existing installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale {"type":"commonjs"} at the config root, so their config root stayed pinned to CommonJS and their dependency manifest stayed gone until they uninstalled. The migration is unusual in one way, and it is the part worth reviewing: the config-root marker was never recorded in gsd-file-manifest.json (writeManifest records hooks/, agents/, commands/, scripts/ and the native plugin, never a root package.json), so classifyArtifact answers 'unknown' for it and the planner's own guard downgrades a remove-managed on an 'unknown' classification to preserve-user. 007 therefore supplies the "purpose-built detector for an old GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions — exact content match, the same predicate removeCommonJsMarker has always used — and declares the resulting classification on the action. A package.json with any other content is left untouched, and there is deliberately no backup-and-remove branch: a non-matching file here is not a patched GSD artifact, it is somebody else's file. Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`: validateStringArray requires the field to be non-empty WHEN PRESENT, while the runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan time and the migration never runs. The metadata test pins this. Kimi is a deliberate carve-out, named in the migration's own header: its marker lived at ~/.kimi, outside kimi's configDir, and migration relPaths are structurally confined to configDir. It is retired by the installer instead. Registration: shipped-migrations table, .gitignore for the emitted .cjs, the EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not copied from migration 006 by rote — 006 needs no entry because it imports nothing, while 007 imports node builtins, so tsc emits its __importDefault helper and the `var` in it trips no-var. This is the same lint gate that made round 1 red. * test(#2544): fault-injection and multi-runtime marker coverage Review round 2, Major 2 + Minors 4 and 5. Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and the suite had no fs monkeypatching at all. Every branch now covered is one whose doc comment claims it as the module's safety posture: - classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule), with an ENOENT control alongside it so the test discriminates rather than just asserting one side - classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never downgrades to the permissive answer) — the fixture's bytes are exactly GSD's marker, so the test fails if the code ever answers on content it could not read - a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already covered with a real symlink, the directory case needs no injection at all) - the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx' - the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the mkdirSync that used to sit outside the guard - removeCommonJsMarker unlink throw -> false These save and restore fs methods in `finally` rather than using chmod 0o000, which does not fault under root and would pass vacuously in root Docker and CI. Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both kimi locations now have behavioral coverage, install and uninstall, each paired with a user-authored-file case proving GSD leaves it alone. Minor 5 — the stagedHooks gate had no assertion behind its stated reason. A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime that declares skipSharedHooksInstall and asserted to stay marker-free, with its user content intact. Also regression-tests the uninstall rmdir gate from the previous commit: an empty plugin dir GSD removed nothing from must survive. * docs(#2544): correct stale marker prose, register the module, document the trade-off Review round 2, Minors 2, 3 and 6. Minor 2 — six files asserted the installed ROOT ships the synthetic marker. None was load-bearing (all three walk-up consumers are VERSION-first with try/catch and the marker never carried a `version`), but ADR-457:52 is the rationale for keeping a generated module, so a future reader would mis-derive the constraint from it. Each site is corrected to what is now true: the installed tree carries no package.json with a .name at all, because the only ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own directories. Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js and the platform-gate test both described `require('../package.json').name` resolving to undefined; post-#2544 that require does not resolve at all, so the history is kept accurate and the present-tense claim corrected rather than just moved. And src/runtime-artifact-conversion.cts described the no-root-package.json case as Codex-only — it is now every runtime, which strengthens that comment's own argument for lazy resolution. The generated .cjs sibling needs no edit: it is gitignored build output, not a tracked file. Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added, including the fail-closed posture and the never-throws contract. Minor 6 — the plugins//extensions/ marker shadows the config root for all .js siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is exactly what #2544's Fix section prescribed and it is disclosed in the PR body, but the PR body is not documentation. It now lives in the OpenCode section of docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than a pure improvement, with the .ts mitigation and a fallback for ESM plugins. * test(#2544): attribute the CommonJS marker in the emitted-provenance rules The differential emitted-attribution gate (#2723, landed on `next` after this branch was cut) went red on the macOS shards once this PR rebased onto it. Two distinct causes, both real gaps rather than noise: 1. `plugins/package.json` and `extensions/package.json` matched NO rule — the `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an unattributed emitted family. 2. `hooks/package.json` fell through to `hooks-built`, which attributes an emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no `hooks/package.json` in the repo, so it resolved to a nonexistent path. Cause 2 is exactly the failure already documented three lines above it for Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix follows that precedent rather than inventing one: `package.json` is excluded from `hooks-built` the same way, and a dedicated `commonjs-marker` rule attributes the family across all four roots it can appear in (both hooks roots plus `plugins`/`extensions`) to the sources that actually emit it. Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for a one-off ripple and goes stale by design — the gate fails a stale ack precisely so it cannot pre-clear the next change on that path. These markers are a permanent part of the emitted tree from #2544 onward, so they need standing attribution. Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3 provenance errors + 12 unattributed paths before, 35/35 green after. * fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker #2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's module on the two properties that matter: - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a DANGLING one, so a dangling package.json symlink classified as absent and the write went straight through it. Demonstrated: against the pre-fix copy, ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory. - create: a plain writeFileSync leaves the classify->write window open, where commonjs-marker creates with flag:'wx' (O_EXCL). Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's own docstring already claimed was the single place these rules are enforced. Exported signatures are unchanged (still boolean), so bin/install.js and the #2717 tests are unaffected. The new subtest is the only coverage that fails if the duplicate is ever reintroduced — the two implementations agree on every non-adversarial input, so the existing suites pass against both. * test(#2544): pin the stagedHooks gate on zcode, not windsurf The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the marker was correctly absent. #2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via dedicated paths and get the marker beside those scripts. Measured on this tree, windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning behaviour that is now wrong, not the gate it was written for. ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin surface to spawn hooks, so GSD stages no .js there by either route (measured: 0 staged, no marker). The property under test is unchanged — a user-created hooks/ directory GSD never fills stays marker-free. * test(#2544): use the shared cleanup helper in the migration test Addresses the review's Major 1. The suppression's stated reason — "no helpers import available" — was not correct: tests/helpers.cjs exports cleanup, and the other test file added in this same PR imports it (tests/commonjs-marker.test.cjs). The local reimplementation dropped two protections that are live on this repo's windows-latest lane: the CWD guard (Windows cannot remove a directory that is the current working directory) and the 20 x 250ms retry budget that absorbs the deferred-scan handle Windows Defender holds on newly-written files. Local function and suppression both removed; local/no-raw-rmsync-in-tests now passes without one. * test(#2544): expect hooks/package.json for the #2717 runtimes The fresh-install contract table predates #2717, which stages cursor/windsurf/ codex .js hooks via dedicated paths and writes the CommonJS marker beside them. All three therefore now receive hooks/package.json legitimately. Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with the marker; cline/copilot/trae/zcode stage none and get none, so their contracts are unchanged. * fix(#2544): gate the #2717 marker writes on having staged something The three dedicated marker writers #2717 added ran unconditionally. Each one mkdirs hooks/ up front and stages its scripts conditionally on the source existing, so with an absent or empty hook source they created a directory, filled it with nothing, and marked it as GSD's anyway. That is the same write-into-someone-else's-territory this issue is about, and installSharedHooksBundle already guards the identical case with `stagedHooks`. The dedicated paths now carry the matching gate: - cursor / windsurf: `installedScripts.size > 0` - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY entry landed. Covered for cursor and windsurf by driving each writer against a src tree whose hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its trigger state needs a package tree where hooks/dist exists but holds none of the allowlist, which is not constructible from a real checkout. * test(#2544): scope the commonjs-marker sources per root The rule declared one flat source list for every marker root, so `extensions/package.json` was attributed to runtime-hooks-surface.cts (which never writes there) and `.kimi/hooks/package.json` to install-engine.cts. That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source, so a flat list containing bin/install.js let any change anywhere in that 13k-line file authorise marker drift for every root — the blanket escape hatch this file's own agents-verbatim comment refuses for exactly the same reason. Sources are now derived per root from ctx.rel. Note the rule ctx is `{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent every path down one branch silently. * test(#2544): state precisely what the zcode assertion pins The comment claimed the test pinned installSharedHooksBundle's `stagedHooks` gate. It does not, and neither did the windsurf version it replaced: zcode declares skipSharedHooksInstall, so the outer guard skips that helper entirely and the gate is never evaluated. The test passes on the runtime exclusion. What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays marker-free — is still worth having, and is what the review asked for. The two `staging zero hook scripts` tests are the ones that pin a real staged-nothing gate. Comment corrected rather than left implying coverage that is not there. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
628648d63a |
chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)
* fix(#2931): preserve protected regions and cap emitted per-runtime bytes Route every runtime brand swap through applyClaudeCodeBrandSwap so "Claude Code" survives verbatim inside <runtime_compatibility> regions (#2284b). The fix existed only in bin/install.js's local copies; the src/*.cts exports still used a naive replace, so binding install.js to the single source -- as this phase does for the Windsurf family -- would have silently regressed those runtimes. A table-driven parity guard now covers all nine brand-swapping converters. De-duplicate the Windsurf converter family: delete the six local copies in bin/install.js and bind the four exported ones by reference, guarded by reference-identity assertions (the ADR-1508/#1675 pattern). The two unexported helpers and an unused tool table go with them. Replace the Windsurf 12,000-byte throw with description truncation, matching the bound its sibling skill converter already applied. The throw could only fire on an ~11.7 KB frontmatter description: the largest emitted workflow is 311 bytes. Truncation makes the cap unreachable by construction and leaves 12,000 in exactly one place, eliminating the dual-surface duplication rather than testing for it. Add the emitted-byte cap gate: buildEmittedSizes captures LF- and <HOME>-normalized bytes from the walk buildParityManifest already performs, and evaluateEmittedCaps asserts them against a per-runtime cap table with dead-rule detection. buildParityManifest's return shape is deliberately unchanged -- diffEmitted compares its values with ===, so making them objects would report all 8,529 emitted paths as moved. A regression test pins the values as strings. Add a deterministic trim-safety gate over composeWithinBudget's omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity rule, replacing the model-graded eval gate the issue described. * docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract * fix(#2931): bound the windsurf command name and single-source the brand swap Review findings from the orthogonal passes, all fixed inline. The claim that removing the 12,000-byte throw left total emission "bounded by construction" was false. The #1615 regex constrains the character class but not the length, and commandName is interpolated three times into the emitted workflow: a 20,000-character name emitted 60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate, clearly-labelled size control that THROWS -- commandName is the @-ref path target, so truncating it would point the workflow at a file that does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The #1615 security regex is untouched and still runs first. 128 is generous: the longest shipped name is gsd-plan-review-convergence at 27. Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe truncation helper. It still used a UTF-16 slice(0,177) -- the exact surrogate-splitting bug the helper was written to avoid, in the very sibling the helper's comment cites as its model. Bounds are unchanged, so output is byte-identical for every shipped command (descriptions max out at 99 chars). Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting the local copy. Adding it to the .cts left two unlinked implementations of identical logic -- the drift class this change exists to remove. Verified byte-identical across eight fixtures and five sequential calls before merging, and guarded by a reference-identity assertion. Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344), add fast-check property coverage for the trim-safety contract, and use fc.pre instead of a bare return in a property callback. * test(#2931): fix three test-authoring bugs the remote matrix caught The remote runner returned 8 unique failures on 6f15cdeb8. All three causes were in the test files, not the modules under test -- local harnesses exercise the modules directly, so nothing executed the test bodies until the matrix did. `{ __proto__: [...] }` in an object literal sets the prototype instead of an own key, so the JSON round-trip erased it and the cap table never saw a reserved runtime key. The production rejection was already correct; the test could not reach it. Use a computed key. Two cap fixtures tripped orthogonal error paths rather than the paths they name: one declared windsurf in the cap table but omitted it from sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern matching nothing (a genuine dead rule). Both now include a compliant artifact so the intended branch is what is asserted. The dead-rule and unknown-runtime contracts are deliberate and unchanged. `const { root } = makeSyntheticConfig({ ... `${root}` })` referenced `root` from inside its own initializer -- a temporal dead zone error. makeSyntheticConfig now optionally takes a (root) => files factory. Also raise the npm pack --dry-run bound 60s -> 120s in the shipped- scripts packaging test. That failure is NOT from this branch: the file is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s here, and the run recorded 60,637ms against a 60,000ms bound -- a timeout under 28,948-test parallel contention, not a slowdown. Fixed rather than deferred because a bound that tight is fragile regardless of which branch trips it. * chore(#2931): backfill changeset pr number to 2984 --------- Co-authored-by: sim <sim@local> |
||
|
|
c043f2946c |
fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900. Every entry is scoped to the diff that introduced it (#2789), so once merged to next it is at the base by definition -- spent and inert. Its presence is still load-bearing though: each PR rewrites the paths map wholesale, making a persistent base copy a shared cell. Five of six conflicting PRs in the open queue collided on this file and nothing else. Deletes the stale document and adds a push-to-next guard asserting it stays absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check against the base is the #2768 shape #2789 exists to end. Closes #2914 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2914): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2914): per-PR ack fragments instead of one shared mutable file The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json whose paths map every PR rewrote wholesale. That is a shared mutable cell: any two PRs needing an ack edit the same lines and conflict. Five of six conflicting PRs in the open queue collided on this file and nothing else. Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same shape .changeset/ already uses to solve this exact problem. Two PRs pick different filenames, so they cannot collide, and fragments lingering on next are harmless rather than toxic. The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An earlier delete-only attempt failed verification twice: the ratchet lost the spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no ack. Relocating preserves every acknowledgment. The legacy single file is still READ (unioned with the fragments) because five open PRs carry it; dropping support would break all of them. A duplicate path key across sources is a hard error, never last-wins. The push-to-next guard is retargeted accordingly: it now asserts only that the legacy SHARED file never reappears on next. Fragments may persist harmlessly. Closes #2914 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f093738412 |
fix(#2891): normalize emitted version against the measured tree, not the measuring repo (#2894)
* fix(#2891): normalize emitted version against the measured tree, not the measuring repo buildParityManifest normalized the install-time {{GSD_VERSION}} stamp using PKG_VERSION, bound at module load from the MEASURING repo's package.json. Since #2767, currentManifests({repoRoot}) measures a DIFFERENT checkout, so during a release cut the baseline worktree (origin/next, 1.8.0) was normalized with the current tree's version (1.9.0) and its literal 1.8.0 stamp survived into the hash. All 364 emitted hook paths diverged and the differential attribution gate hard-failed every finalize/rc run. Normalize against the version of the tree that PRODUCED the emitted output: buildParityManifest takes an explicit pkgVersion, and currentManifests resolves it from the measured tree via a new fail-closed measuredPackageVersion(). * chore(#2891): backfill changeset pr number (#2894) --------- Co-authored-by: Test <test@example.com> |
||
|
|
aa19e3478c |
fix(#2854): pin the emitted gate to the base the tree was merged with (#2859)
* test(#2854): failing-first coverage for CI baseline export provenance Extracts the export decision out of main() behind injected IO so it is unit-testable, preserving today's export-whenever-present semantics, and adds the matrix that proves those semantics are wrong. The PR lane restores the emitted baseline keyed on the PR's recorded base sha while the gate resolves the base ref live, so the two drift whenever next advances mid-flight. The restore was published straight to GSD_EMITTED_BASELINE, where a mismatch is fatal, turning a recoverable cache into a hard failure on diffs that touched nothing related. Also renames the stale-env fixture from 'from-cache-restore.json' to an operator-pin name: that fixture asserted the exact conflation this bug is, documenting the defect as intended behavior. Refs #2854 * fix(#2854): validate a restored baseline before publishing it as an operator pin GSD_EMITTED_BASELINE is an operator pin: resolveBaseline() treats a mismatch there as a hard stop, because the operator said "use this one". CI published its cache restore to that same variable whenever the file merely existed, so a restore keyed on the PR's recorded base sha - while the gate resolves the base ref live - turned a recoverable cache into a fatal error whenever next advanced mid-flight. Required tests went red on diffs that touched nothing related, and named a test file the contributor never opened. The export step is the boundary, so it is the boundary that validates. It now publishes only a baseline already valid for the sha under test, judged by validateBaseline so the staleness rule keeps one definition. Anything refused is left to be found via the cache path, where a mismatch degrades to the in-job build exactly as ADR-2719 SS5 specifies. The operator hard stop is untouched, and the fast path still hits on a current cache. Also reports the sources actually reached rather than asserting all three ran: the failure message claimed an in-job build it had returned before calling, sending contributors after a rebuild that never happened. Fixes #2854 * fix(#2854): pin the emitted gate to the base the tree was actually merged with The differential compared a tree built on one commit against a baseline at a different one. "Rebase check" merges pull_request.base.sha, pinned by #2472 so all 12 matrix jobs agree on one tree, but resolveBase() fell through to origin/next, which fetch-depth 0 leaves at the live tip. Nothing set GSD_EMITTED_BASE, so whenever next advanced mid-flight the two disagreed. The cached baseline, keyed on base.sha, was correct for that tree and was rejected as STALE by a target that was not. Required tests went red on diffs that touched nothing related, naming a test file the contributor never opened. The near miss is the worse half: had resolution gotten past the baseline step, a baseline at the live tip would have attributed commits merged to next in between to the PR under test. The hard stop was shielding us from a wrong answer, so making it fall through would have made this worse. Pins GSD_EMITTED_BASE to the same expression as CI_REBASE_BASE_SHA in every rebase-merged job, with a parity test asserting the two cannot diverge. That parity check immediately caught a third lane, test-inert, that merges a pinned base and had been missed. Fixes #2854 * fix(#2854): keep the export step self-contained across the package boundary scripts/ ships in the npm tarball and tests/ does not, so requiring the validator across that boundary is MODULE_NOT_FOUND in a published install. The export step now reads the same GSD_EMITTED_BASE pin the gate resolves through, and applies a cheap self-contained precondition; validateBaseline remains the sole authority and still runs downstream on whatever is published, so there is no second opinion to drift. Reading the pin rather than re-deriving a base is the point: a second, divergent base lookup is exactly what caused this bug. The same hazard pre-exists in scripts/gen-emitted-baseline.cjs, which ships and requires three tests/ modules. Filed as #2858 rather than folded in: fixing it means relocating the shared helpers out of tests/ and updating every consumer, which would bury this change. Refs #2854 * fix(#2854): stop announcing the restored cache through the operator-pin door Two independent reviewers found the same blocker in the previous approach. Validating before publishing to GSD_EMITTED_BASELINE only narrowed the hole: the precondition gated on sha equality alone, so a document with a MATCHING sha but a wrong schema version or malformed manifests was still announced as an operator pin and still hard-stopped downstream. That reproduces this bug's own class, triggered by malformation instead of staleness. The step was never load-bearing. The cache restores to .gsd-cache/emitted-baseline.json, which is resolveBaseline's DEFAULT_CACHE_PATH and is read whether or not anything announces it. Publishing the same file to the pin door could only ever convert recoverable into fatal, so the step and its script are deleted rather than made cleverer. Every failure mode now degrades to the in-job build by construction, and validateBaseline is once again the only thing that judges a baseline. Coverage moves to where the behaviour lives: stale sha, wrong schema version, manifests array/absent, non-object documents, unreadable file, and the 39/40/41 hex boundary all assert degradation via the cache path. Adds the empty-pin case a reviewer flagged as untested - the pin is job-level env, so on push events it renders as an empty string, and only baseRefCandidates' truthy check keeps it out of the candidate list. Fixes #2854 --------- Co-authored-by: Test <test@example.com> |
||
|
|
4f6935e29b |
fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks (#2846)
* test(#2717): CommonJS marker for cursor/windsurf/codex staged .js hooks Cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) stage .js hook scripts via dedicated paths that bypass installSharedHooksBundle — the only writer of the {"type":"commonjs"} marker. Under a config root declaring {"type":"module"}, Node loaded those scripts as ESM and every require() failed with 'require is not defined', silently disabling the runtime's hooks. Adds regression tests (RED first, fix lands next commit): - parametrized cursor/windsurf/codex install asserts hooks/package.json exists with exactly GSD's marker content; - end-to-end: a cursor require()-using hook loads under a planted ESM-typed config root without the require-is-not-defined error; - the ensureCommonJsMarker / removeCommonJsMarkerIfGsdOwned contract: GSD markers are removed on uninstall, user-authored package.json is never touched. * fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks The {"type":"commonjs"} marker lived only inside installSharedHooksBundle, which cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) never reach. Their .js hooks are staged by dedicated paths, so under a config root declaring {"type":"module"} Node loaded them as ESM and every require() failed with 'require is not defined', silently disabling those runtimes' hooks. Decouple the marker write into a shared helper so any code path that stages .js hooks can ensure it lands in the SAME directory as the scripts: - src/runtime-hooks-surface.cts: add ensureCommonJsMarker(dir) + removeCommonJsMarkerIfGsdOwned(dir) (byte-identical content to installSharedHooksBundle's marker; preserves a user-authored package.json on both write and uninstall). Call ensureCommonJsMarker(hooksDir) from writeCursorHooksJson + writeWindsurfHooksJson; call removeCommonJsMarkerIfGsdOwned on their matching remove paths. Export both. - bin/install.js: call hooksSurface.ensureCommonJsMarker after the codex hook copy; call hooksSurface.removeCommonJsMarkerIfGsdOwned in the generic hooks-removal loop (safe no-op where no marker exists). No change to which runtimes receive the shared bundle, the !isCodex gate, skipSharedHooksInstall, or kimi/kimi-code/cline/copilot/trae/zcode (all unchanged — audit in the diagnosis). RED @ dbb7d2bb (6 failures: 3 missing markers + the ESM require error + missing helpers); GREEN pending. * docs(#2717): changeset fragment (pr:0, backfilled post-PR) * chore(#2717): regen codex/cursor/windsurf install-tree fixtures + attribution ack The fix adds hooks/package.json to those three runtimes' install trees (the new CommonJS marker), so the golden install-tree fixtures gain one path each (regenerated via npm run gen:install-tree). emitted-attribution (ADR-2719) flags the 3 emitted hooks/package.json paths under the hooks-built rule; acknowledge them. Also drops 5 spent ack entries left by now-merged PRs (#2694 code-review.md/code-review-fix.md, #2695 worker/registry, #2794 review.md) — they are stale on this branch (base already carries them). * fix(#2717): codex ESM-root behavioral test + hooks-built provenance for package.json Two review-driven follow-ups on the #2717 fix: - Adversarial review noted the ESM-root behavioral test covered only cursor; refactor it into a helper and add a codex case (the !isCodex-gated path most likely to regress, whose marker write lives in bin/install.js). gsd-check-update.js require()s at module load, so it surfaces the ESM failure immediately. - emitted-provenance flagged hooks/package.json as 'attributed source does not exist' — the marker is code-derived (a fixed literal emitted by ensureCommonJsMarker at install time), not built from a tracked source. Route the hooks-built rule's sources/transforms for package.json to the surface source file, mirroring the existing .cmd-shim sub-family. * chore(#2717): drop now-redundant hooks/package.json attribution ack The hooks-built provenance routing (prior commit) now self-attributes the emitted hooks/package.json to src/runtime-hooks-surface.cts, which IS in this diff — so the attribution is self-explaining and the emitted-drift-ack entry became stale. Delete the (now-empty) ack file per ADR-2719's empty-file rule. * docs(changeset): backfill #2717 PR number to 2846 |
||
|
|
1e3c995e6f |
fix(#2789): scope the emitted-drift ack to the diff that introduced it (#2803)
* fix(#2789): scope the emitted-drift ack to the diff that introduced it Every input to `diffEmitted` is base-relative -- `baseline` vs `current`, `changedPaths` from `git diff base...HEAD` -- except the ack set, which was read absolutely, from the working tree only. A differential machine consulting a non-differential input. So `staleAcks` asks exactly one question, "did a delta consume you?", and that cannot distinguish an ack that never explained anything (an authoring mistake) from one whose ripple is now absorbed into the base (the ack's SUCCESS condition). After merge an ack is in the second state but reports as the first. The trigger is ordinary. Actions sets GITHUB_BASE_REF on pull_request events only, so a push to `next` falls through to origin/next -- the very commit under test. Both sides build identical content, no deltas remain, and every live ack is reported stale. PR #2768 acked a deliberate 40866 -> 42020 byte growth, was green on its own lane, and reddened `next` the moment it merged. It also reds every PR branching off the poisoned base, and since publish-emitted-baseline is gated on the test job, it blocked baseline publication too. Give the ack the base side it was missing. `diffEmitted` now takes `baseAck` -- the same document at the base ref, via `readAckFileAtRef`. An entry already present there is SPENT: it may no longer consume a delta and is never reported stale, only surfaced as `spentAcks` for tidying. An entry new or reworded in this diff stays live, and if nothing consumes it that genuinely fails, with blame on the author who just wrote it. This closes a hazard the IMPLEMENTATION named but could not prevent -- a leftover ack silently pre-clearing the next ripple on its path. (ADR-2719 §3 asserted only that TOUCHING the file is the alarm; its residual-risk list never covered pre-clearing, and §3 now carries an amendment.) Verified against the two-PR laundering sequence -- land an innocuous ack, then change the artifact -- which passed silently before and now fails on both the hash pass and the size ratchet. Three things the design has to get right, each of which was wrong first: - A read failure on the base document THROWS; only absence-at-the-ref returns null. Returning null on error LOOKS armed (every entry stays live) but a live entry's defining power is that it CONSUMES a delta, so null is armed on the staleness axis and DISARMED on consumption -- silently the whole pre-#2789 gate. `git show` cannot tell absence from fault, so absence is established with `ls-tree`. - Re-arming a spent ack costs actual PROSE. Internal whitespace and the zero-width family collapse, and `runtime` is not compared: a doubled space, an invisible character, or a decorative field would otherwise re-arm an ack whose justification still describes the previous ripple, showing a reviewer nothing. - `baseAck` is REQUIRED once an ack declares entries -- omission is an error, not a silent "inherit nothing" -- so a dropped argument fails loudly instead of quietly restoring this bug with the suite green. Because a corrupt document ON THE BASE is expensive (the loud base-side failure reds every ack-carrying PR), scripts/lint-emitted-drift-ack.cjs blocks one from landing. It is standalone rather than importing parseAck -- scripts/ ships in the npm package and tests/ does not -- so a parity test runs both surfaces over one corpus and fails on divergence; it caught one immediately, a `null` document, now classed as policy rather than schema. Deadlock is separately foreclosed: a tree carrying no ack never reads the base, so the PR that DELETES a corrupt file still lands. `readAckFileAtRef` takes an injected git runner so all four branches are tested deterministically; it never executes in the remote runner, where the real-tree test skips for want of a base ref. It also refuses an option-shaped ref, since execFileSync's array form stops shell metacharacters but not git's own option parsing. Rejected: skipping the differential when base == HEAD. It treats the symptom, costs real coverage on the push-to-next lane, and does nothing about the downstream PRs the same flaw was reddening. Deletes the now-spent tests/emitted-drift-ack.json, and updates the CONTEXT.md canon and ADR-2719 §3: presence is no longer the alarm -- a LIVE entry is, and a spent one is inert. Closes #2789 * chore(#2789): backfill changeset PR number |
||
|
|
e276cc7f00 |
enhance(#2778): make the size-ratchet failure name its own remedy (#2780)
* fix(#2778): exempt intentionally-absent paths from the glossary gate check-glossary-refs asserts that every backticked tests/ token in CONTEXT.md resolves on disk. tests/emitted-drift-ack.json (ADR-2719 section 3) is absent on a healthy next BY DESIGN — it appears only inside a PR that needs it, which is what makes touching it the alarm. It passed before only by accident of backtick pairing: CONTEXT.md's RULESET entries are themselves backtick-wrapped and contain backticks, so the token happened to fall outside a code span. Any edit that shifted the parity exposed it. A gate that passes by luck is not passing. The exemption is exact, not a prefix hole: a sibling missing tests/ path still fails, and a test locks that. * feat(#2778): make the size-ratchet failure name its own remedy The growth branch stated a requirement and withheld the means of satisfying it: no ack file named, no schema, no key format, and no do-not-regenerate line — so the likeliest guess was to hunt for a baseline that #2724 deleted. Observed live on #2543. All remediation now comes from one frozen REMEDIATION export whose example document is rendered from ACK_VERSION, so the taught schema cannot drift from the schema parseAck accepts. A round-trip test feeds the printed document back through parseAck. The report is now built as a typed IR (buildReport) that formatReport renders, so tests assert on structure rather than prose, per CONTRIBUTING.md's raw-text-matching rule. Two defects found and fixed inline while building: - diffEmitted's validation early-return omitted newFileCapExceeded while formatReport reads its length, so the branch that reports a failed git diff threw a TypeError instead of naming the problem. - Printing one complete ack document per failing branch made each read as the whole file, so pasting the second over the first silently lost an acknowledgment. One document now covers the whole report. Closes #2778 * chore(#2778): backfill changeset pr number to 2780 |
||
|
|
1c1af70a4b |
refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator Removes the 19 committed path->hash manifests, the two per-file size baselines, tests/golden-install-parity.test.cjs, and scripts/gen-golden-install-parity-zcode.cjs. These were pure functions of the source tree (ADR-2719); the differential attribution check (tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs) is now the sole gate for emitted-artifact propagation. tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs are unchanged (ADR-2719 section 7 exception). Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's existence guard, .gitattributes, package.json scripts, the emitted-provenance totality guard's IO, the differential check's baseline acquisition, CI wiring to publish/restore the baseline artifact, and docs. * refactor(#2724): make the differential attribution check self-sufficient Three fixes required to delete the golden fixtures without breaking CI: - scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs from the three rules that named it. #2759's missingRuleTestFiles guard hard-throws at module load if a rule names a test file absent from disk, which would break the changes job on every PR the moment the fixture-deletion commit landed. - tests/helpers/emitted-provenance.cjs: loadManifests() read the committed golden fixture directory. With that directory deleted at every future ref, this would throw at module load forever, taking the Phase 2 totality guard down with it. Rebuilt from real installer spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest), the same shape emitted-runtime.cjs's currentManifests() already uses. - tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs: the real-tree test's baseline acquisition swaps from baselineManifestsAtRef(base) (git show at a ref that no longer carries fixtures) to resolveBaseline()'s documented precedence: env, then the on-disk cache, then an in-job build. The build fallback (buildBaselineAtRef, new) checks out base into a throwaway git worktree and runs the new scripts/gen-emitted-baseline.cjs there -- no npm ci needed, since bin/install.js and the test helper shells are Node-builtins-only. That script also publishes the baseline artifact from CI's push-to-next job (wired in a follow-up commit). * refactor(#2724): retire the merge-driver bridge and per-file size baselines The Phase 1 bridge (#2721) is retired now that the artifacts it guarded are deleted: scripts/git-merge-regen-driver.cjs, its test, and the 'setup:merge-driver' npm script are removed, and the .gitattributes merge=gsd-regen/linguist-generated block for the three deleted-path globs is dropped. tests/fixtures/install-tree/*.json keeps its normal merge behavior, unchanged (ADR-2719 section 7). scripts/update-size-baseline.cjs and its test are removed: their sole purpose was regenerating tests/workflow-size-baseline.json and tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm script and its step in 'regen:derived' go with it. The per-file baseline describe blocks in tests/workflow-size-budget.test.cjs and tests/agent-size-budget.test.cjs are removed for the same reason; the independent loose-tier hard caps are untouched. The differential attribution check's size ratchet (tests/emitted-diff.cjs, already shipped in #2723) is the replacement anti-creep mechanism. 'npm run gen:golden' is replaced by 'npm run gen:install-tree', which keeps regenerating tests/fixtures/install-tree/*.json (the one artifact family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's error messages point at the new command name. tests/golden-parity-single-source.test.cjs's anti-divergence guard (#2266) is retargeted from the two deleted golden-parity consumers to their two replacements (tests/helpers/emitted-runtime.cjs and tests/helpers/emitted-provenance.cjs), which import buildParityManifest the same way — the divergence risk the guard exists for is unchanged. Also wires CI: a new publish-emitted-baseline job runs scripts/gen-emitted-baseline.cjs after a push to next and caches the result keyed on the sha; the test and test-full jobs restore that cache on pull_request events, keyed on the PR's base sha, and export GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree test to pick up. * docs(#2724): flip ADR-2719 to Accepted and update contributor docs Status: Proposed -> Accepted. Regenerated docs/adr/README.md index. CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET. EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET. AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry) no longer point at the deleted golden-install-parity fixtures, size baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver / git-merge-regen-driver.cjs bridge. Editing shipped content now requires zero manual fixture regeneration, documented against the differential attribution check instead of the deleted commands. * docs(#2724): add changeset for removed golden-parity commands * fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs token in CONTEXT.md resolves to a real file. The RULESET. EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks, which the checker reads as a live reference, not historical prose. * test(#2724): retarget ci-test-scope tests off the deleted golden test tests/ci-test-scope.test.cjs asserted specific RULES entries select tests/golden-install-parity.test.cjs, and that every rule selecting it also selects both emitted gates. Both premises broke when the golden test was deleted (#2724): the deleted filename never re-appears in targeted_tests, and there was no longer a third file for the gates to travel alongside. Retargeted the two selection describe blocks to assert tests/emitted-provenance.test.cjs directly (the drift guard the golden gate's rules were retargeted to), and simplified the third block to assert the two emitted gates always travel together, without reference to the golden filename. * docs(#2724): repoint two contributor how-to guides at the differential check Both guides told contributors to regenerate a baseline against tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed at the differential attribution check (tests/emitted-attribution.test.cjs, ADR-2719), which needs no manual regeneration step. * fix(#2724): repair phase6-capstone-conformance's deleted-baseline read An independent orthogonal review caught a real regression this branch introduced into a test file the branch's diff never touched: tests/phase6-capstone-conformance.test.cjs read tests/workflow-size-baseline.json (deleted earlier in this branch) with no fallback, so the whole suite would throw ENOENT the moment this branch landed. The test's actual intent — prove the host-loop workflow files are real, tracked, non-empty docs — is preserved by asserting the live byte count via the same shared counter (scripts/workflow-size.cjs) the size guards already use, instead of a committed snapshot. Also, from the same review: a stale doc comment in scripts/workflow-size.cjs still named the deleted scripts/update-size-baseline.cjs as a consumer, and buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left two fs.rmSync calls unguarded against masking the primary result/error, inconsistent with the try/catch already wrapping the git cleanup beside them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef explaining why it (and its siblings) are kept despite having no production caller post-cutover — they still answer real questions about refs that predate the cutover. * fix(#2724): repair three real regressions found by remote verification 1. tests/emitted-provenance.test.cjs's two hostile-input tests (non-object manifest, unreadable fixture) drove loadManifests(tmp) and monkeypatched fs.readFileSync, both premised on the deleted fixture-directory read this branch already replaced with real installer spawns -- the negative assertions silently stopped firing. loadManifests() now accepts injected {families, install, build, clean} (defaulting to production values), giving the tests a real seam to drive a bad build result and a build failure through the ACTUAL loader instead of a reimplementation, and added coverage that clean() still runs on both paths. 2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE' steps hardcoded shell: bash, which is wrong on windows-latest (native pwsh) and on test-full's macos-latest legs (native zsh per that job's own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning .test.cjs) caught it. Replaced the inline bash script with scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a bare 'node <path>' command line has no shell-specific syntax, so it runs correctly under bash, zsh, and pwsh without a shell override. tests/phase6-capstone-conformance.test.cjs's deleted-baseline read (caught by the same remote run, at a commit prior to this one) was already fixed in d0c3b1242 and is not touched here; verified still passing after these changes. * fix(#2724): revive ADR-1610's new-file size cap inside the differential An isolated review caught a real regression: deleting tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP (ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor) with no successor. tests/helpers/emitted-diff.cjs's size ratchet already 'continue's past any file absent from sizeBaseline -- exactly the files this cap exists to bound -- so a brand-new workflow file sized 32,769-40,960 bytes passed CI clean and shipped, then risked silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted and never referenced anywhere in this branch. Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline, name) signal the growth check already computes -- 'new' is exactly 'present in sizeCurrent, absent from sizeBaseline'. Not ack-able, matching the tier hard caps it sits beside: the fix is extraction, not an acknowledgment entry. Documented, disclosed narrowing: the pure differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering (tests/workflow-size-budget.test.cjs's classification), so a legitimately large new file must extract rather than tier in, one release earlier than an existing file would need to. ADR-1610 itself is left unamended -- this restores its decision rather than re-litigating it. Also fixes a stale comment plus a redundant real 19-installer-spawn assertion left over from the pre-injection-seam version of tests/emitted-provenance.test.cjs's build-failure test, and annotates 3 of 4 stale golden-fixture citations in docs/reference/host-integration-capability-matrix.md as superseded (the 4th is an accurate historical PR narrative, left alone). * fix(#2724): repair three red CI defects on the golden-fixture cutover Windows-only provenance false attribution (defect A): the `hooks-built` provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are Windows-only installer output (ensureCodexHooksJsonSessionStart / ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so the self-attribution resolved to a path that exists on no platform. Only windows-latest ever emits the key, so this only failed there. Fixed by special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated rule) — a dedicated rule would match zero paths, and therefore report as a dead rule, on every non-Windows lane of the same totality guard. `sources` already supported per-match functions; `transforms` is extended to support the same shape so the attribution can vary by match within one rule. Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef` ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but that script is new in this PR and therefore absent at any base ref that predates it — every call failed closed with "Cannot find module". Fixed by running the PR checkout's own generator against the worktree via a new `--dir` parameter, decoupling "which copy of the script runs" from "which tree it measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override, threaded down to `runMinimalInstall`'s new `installScript` override). This is not just a bootstrap fix: a differential needs ONE measurement schema applied to both sides, or the two stop being comparable the moment that schema evolves — running each side's own copy would silently reintroduce that risk. Verified locally end-to-end against real origin/next: resolves a valid {version, sha, manifests, sizes} artifact with the correct sha and no leaked worktree. Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let docs-lint evaluate the fragment for the first time; it already passes (docs/TESTING-SUITES.md and friends already document the removed scripts). Also fixed while in this file: an eslint no-unused-vars warning surfaced by the changed lint run (unused `cleanup` import in tests/emitted-provenance.test.cjs). Added regression coverage for both A and B: a cross-platform spot-check that drives the real hooks-built rule against `.cmd` keys directly (not through a real Windows install), and a real-tree test that drives buildBaselineAtRef against a base ref verified (via git cat-file) to lack the generator, both skipping honestly rather than false-passing when their precondition does not hold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test Two isolated-review findings on PR #2767: - `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the wrapped `hooks/<name>.js` script, asserting a byte-provenance link that does not exist — traced against buildCodexHookWindowsShimIR (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that same file) flows into the .cmd bytes, never its content. Point `sources` at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention used elsewhere in the table. Since `sources` is checked before `transforms` in the differential, the wrong mapping silently excused any .cmd byte movement caused by editing the wrapped .js file. - The `buildBaselineAtRef` regression test skipped unless a resolvable base ref still lacked scripts/gen-emitted-baseline.cjs — true only until this PR merges, after which every base ref carries the file and the test skips forever with zero ongoing coverage. Rebuilt hermetically: synthesize the missing-generator condition in-place via git plumbing (a throwaway commit, child of HEAD, with just that one file removed from a scratch index), never touching the real working tree, HEAD, or index, and never depending on ambient history or remotes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path The runner container mounts the repo at a path owned by a different uid than the process running the suite, so git's dubious-ownership protection refuses every git operation there. GitHub Actions never hits this because actions/checkout registers the workspace as safe automatically; this runner's container does not. buildBaselineAtRef is the production build-fallback the sole remaining emitted gate depends on (resolveBaseline's in-job-build leg), not just a test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs) that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's worktree add/remove/prune, and the hermetic regression test added in the prior commit — funnels through, plus gen-emitted-baseline.cjs's own rev-parse (now reusing that same wrapper instead of a second execFileSync, so the fix has one source of truth). Each call declares -c safe.directory=<the exact directory it already operates on>, never the * wildcard. Audited every other helper on this surface (emitted-diff.cjs, emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them shell out to git at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
44707c2c5e |
fix(#2757): attribute transform-driven emitted changes and reclassify agents-verbatim (#2760)
* fix(#2757): let derived provenance rules attribute a transform change The Phase 2 provenance table (#2722) could only explain a moved emitted path via its `sources`, so a `kind: 'derived'` artifact whose TRANSFORM code changed (not its source) was unattributable by construction: 16 emitted agents/*.toml moved by PR #2566's runtime-artifact-conversion.cts change, with zero agents/*.md in the diff. Adds an optional per-rule `transforms: string[]`. A moved emitted path now attributes via its source OR its transform, reusing sourceSatisfiedBy so exact/prefix matching stays identical for both. Declared narrowly on agents-toml-derived and agents-verbatim (the two families verified to route through the same conversion pipeline) as src/runtime-artifact-conversion.cts, src/install-effort-resolver.cts, and src/model-catalog.cts — never bin/install.js, which would be a blanket escape hatch spanning every installer concern. Also corrects agents-verbatim from kind:'identity' to kind:'derived': measured against origin/next's own fixtures, the same agents/<name>.md hashes differently per runtime (e.g. codex vs claude), which a true verbatim copy cannot do. Verified empirically by installing every runtime and diffing against the raw repo source — none reproduce it byte-for-byte; every one rewrites frontmatter and the hardcoded .claude/ self-reference, and claude additionally gets an effort: line injected. A new assertNoIdentityTransforms invariant rejects an identity rule that declares a non-empty transforms list. Closes #2757 * docs(#2757): flag PR #2566's transform-role files for later verification src/agent-tools-contract.cts (added by #2566, absent on next) and src/agent-install-check.cts (modified +112/-1 by #2566, read-only today) were reviewed for prospective inclusion in AGENT_TRANSFORM_SRCS. Excluded for now: the nonexistent file would fail this fix's own "every declared transform path exists" hygiene test, and the modified file's future role cannot be verified without the unmerged PR's diff. Documents the reasoning, the interim ack-file safety net, and the verification method to apply once these files stabilize, so the next PR touching them has a pre-scoped one-line fix rather than a silent gap. Related to #2757 |
||
|
|
0f60266042 |
fix(#2723): reconcile emitted manifest families as a set, not a count (#2750)
* fix(#2723): reconcile emitted manifest families as a set, not a shared count EXPECTED_MANIFEST_COUNT was a single literal 19 asserted against both the baseline (built at the base ref) and the current tree (built at PR HEAD). Those sides legitimately differ by one family whenever a PR adds or removes a runtime, so no value satisfied both: 19 rejected the current side, 20 rejected the baseline side. Every runtime-adding PR was hard-blocked. Replace the shared literal with three independent signals - the derived family set, the recorded fixture set, and the families present at the base ref - reconciled as sets in both directions. A family may appear or vanish only when the diff plausibly touches the runtime registry, and the failure names the family rather than a count. An absolute floor catches the uniformly shrunken universe a same-count self-check passes vacuously. Found by tracing #2005 (Qoder runtime) through the gate during the ADR-2719 dual-run window. * fix(#2723): read the baseline family set from the ref, not HEAD's registry Review found three defects in the first cut. Blocker: baselineManifestsAtRef enumerated MANIFEST_FAMILIES, which is imported at module load and therefore describes PR HEAD. A runtime REMOVED by the PR is already absent from that list, so the base ref was never asked for it, the baseline silently omitted a family that genuinely existed, and the dropped-family check could never fire in production - while its unit tests passed, because they inject the baseline directly. Enumerate from the ref with git ls-tree instead. Also: narrow the registry-signal set to the two surfaces that actually define the family set, since every extra path widens what excuses an unattributed delta; drop the ack bypass, which was a one-sided escape hatch making removals easier to wave through than additions; and gate the derived/fixtures inputs so malformed values return a verdict rather than an unhandled TypeError. * fix(#2723): filter prototype-shaped family names read from git output baselineFamilyNamesAtRef derives object keys from git ls-tree output rather than a trusted constant, so a fixture committed as __proto__.json would turn the manifests[name] assignment into a prototype write. Compared inline rather than through a Set, which is the form the prototype-pollution analysis recognizes. * fix(#2723): require an exact capability path depth and make the ref test hermetic The remote runner went red on both linux lanes with two real defects. The capability signal matched by prefix+suffix, so 'capabilities/capability.json' (no runtime segment) and 'capabilities/a/b/capability.json' (wrong depth) both attributed a family change and would have excused an unattributed delta. Anchored to an exact single-segment pattern. The ref-derivation test reached for this repo's root commit, which is not stable: the remote runner shallow-clones, so rev-list --max-parents=0 returns the grafted boundary carrying every fixture, and this repo has two root commits locally anyway. It now builds its own git repo containing a family absent from the current registry - the real discriminator, and one the root-commit version could never assert. Lint then caught a third: the test called t.after() without declaring t. * fix(#2723): stop asserting ref enumeration against the ambient checkout The remote runner returned [] for the repo's own HEAD while every hermetic temp-repo assertion in the same test passed. That is this function's documented behavior when git cannot read the ref - the runner works from a shallow clone under a bind-mounted workdir - so the assertion was testing the checkout rather than the code. Dropped it. The temp repo already proves the property that matters, and proves it more strongly: it contains a family absent from the current registry, which a registry-derived implementation could never report. The ambient path stays covered by the real-tree test, which skips explicitly when no base ref is resolvable. A git failure is not silently permissive downstream: baselineManifestsAtRef returns null on an empty family set and the real-tree test asserts the baseline is non-empty. |
||
|
|
9138271b5f |
test(#2723): differential emitted-attribution check, dual-run beside the golden (#2737)
* test(#2723): differential emitted-attribution check, dual-run beside the golden Phase 3 of #2719. The conservation law itself, running BESIDE golden-install-parity.test.cjs -- both green, fixtures untouched. Every emitted path whose hash moved between next HEAD and PR HEAD must be attributable, through the Phase 2 table, to a path the PR actually changed. Unattributable deltas fail with the paths NAMED. The only way through is a committed acknowledgment, never a flag -- a contributor facing a red gate sets a flag, which is what UPDATE_GOLDEN=1 is today. The central decision is that the law is a PURE function (no fs, git, installer, or clock), with I/O confined to a separate resolver. The naive one-big-integration-test shape would need ~38 installer spawns per assertion, so #2723's four failing-first criteria would not in practice have been written -- which is exactly how a phase ships promised-but-not-built. Pure, they are millisecond table tests, and the Stryker gate can actually bite. Buckets are conserved: every moved path lands in exactly one of attributed | unattributable | acked, property-tested at 400 runs. A path the provenance table cannot resolve surfaces as an error, never a silent skip. Asymmetries that are deliberate, each with a test: - an ADDED emitted key is a ripple too, not just a modified one - synthesized paths are exempt; code-derived ones are NOT (Phase 2 refused to mark them exempt precisely because exempt means permanently blind) - shrinkage needs no ack; growth does. Gating shrinkage would punish exactly what the size ratchet wants - a STALE ack is a hard failure -- an ack outliving its ripple pre-clears the next one on that path - a failed `git diff` is an explicit error, never an empty changedPaths set; reading it as "nothing changed" would make everything unattributable and produce a failure storm that reads like a real finding - prefix sources are SEGMENT-aware, so `agents/` does not attribute `agentsfoo/x.md` Baseline is cached, not committed, keyed on the next sha. A stale key is refused rather than used: absence fails loudly and gets fixed, whereas staleness produces a confident wrong answer. An explicitly pointed-at GSD_EMITTED_BASELINE that is stale is a hard stop; a stale cache falls through to the in-job build. No baseline-unavailable path returns -- ADR-2719 section 6 names that trap, since in node:test a bare return is a PASS. Both the conservation property and the staleness gate were mutation-verified (injecting a swallowed key fails 9 tests; disabling the staleness comparison fails 5). Fixtures, generators, the merge driver and the ADR status are untouched -- those are Phase 4 (#2724). Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2723): compute stale acks once, after the size pass Self-review defect found while the reviewers were running. `staleAcks` was computed twice: once between the hash pass and the size pass, then again after. Only the second value was returned, so the first was dead code -- and the dead one was placed where it would have been WRONG. An acknowledgment can be consumed by either a hash move or a size growth. Computing staleness before the size pass reports a legitimate growth ack as stale, which is a false failure that pushes a contributor to delete the very ack that is doing its job. Now computed once, after both passes, with a regression test. Verified by mutation: restoring the early computation fails 2 tests. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2723): wire the attribution check to the real tree, not just synthetic input An isolated reviewer caught that the first cut was INTERFACE-ONLY: nothing read the ack file from disk, nothing shelled git, nothing built real manifests. Every test was true of hand-built inputs and none of the repo, so the acceptance criterion "both this check and golden-install-parity green on the same tree" was trivially true rather than meaningfully true. That is the promised-but-not-built failure this epic keeps finding in its predecessors, recurring one phase later for the wiring itself. Taken, not argued. Adds tests/helpers/emitted-runtime.cjs -- the only module that touches git, disk, or the installer -- and an integration test that runs the same pure law against reality: - CURRENT side: 19 real installer spawns via runMinimalInstall + buildParityManifest, the same machinery the golden harness uses. - BASELINE side: `git show origin/next:<fixture>`. That is next's RECORDED emitted state and it costs nothing. Deliberately NOT the working-tree fixtures, which are whatever this PR's author regenerated -- comparing against those would be vacuous. Phase 4 deletes the fixtures and swaps in resolveBaseline's cache path, already implemented and tested. - changed paths from real `git diff --name-only origin/next...HEAD`, with the git subprocess bounded at 30s per CLAUDE.md's unbounded-subprocess rule. - the real tests/emitted-drift-ack.json (absent is legal; present-but-empty or unparseable throws rather than being read as absent). Verified it can actually fail: an uncommitted edit to a shipped workflow moves emitted output but never appears in the committed diff, and the check names all 18 affected emitted paths with the message format ADR-2719 §1 specifies. Restores clean. Also from review: - readAckFile now has a real test exercising the SUT across absent / valid / empty / unparseable / unreadable. The previous test asserted fs behaviour rather than SUT behaviour, because no SUT ack-reading path existed yet. - formatReport's sampleLimit gains true limit-1/limit/limit+1 coverage at 19/20/21. A test was previously NAMED "(limit+1)" while testing no numeric limit at all, which is worse than no coverage because it reads as covered. Windows uses an explicit t.skip (install output is platform-specific there, mirroring the golden harness) -- never a bare return, which node:test scores as a PASS. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2723): cover the claude-local manifest family in the real-tree check Isolated adversarial review, MAJOR. The real-tree wiring enumerated Object.keys(RUNTIME_META) -- 18 entries -- while the emitted manifest set has 19 families. The 19th is claude-local: claude is the reference host and the only runtime with a distinct LOCAL "legacy flat-commands" layout (commands/gsd-*.md + agents/gsd-*.md at project scope), which golden-install-parity.test.cjs guards with a hand-coded test outside its RUNTIME_META loop (#2086). The family was dropped from BOTH sides, so the test's own self-check (current.length === baseline.length) passed vacuously at 18 === 18. A PR changing Claude's local-scope output would have failed the golden while this check reported ok -- and that disagreement is precisely what the dual-run window is designed to surface as a provenance-table hole. A wiring omission masquerading as one is the worst available failure here, because it would have been read as evidence about Phase 2 rather than a bug in Phase 3. Fixed by deriving MANIFEST_FAMILIES explicitly (18 global + claude-local at local scope) instead of inferring the set from RUNTIME_META. The self-check is also repaired: it now asserts both sides against the INDEPENDENT EXPECTED_MANIFEST_COUNT from the Phase 2 table, and asserts claude-local specifically. Comparing the two sides to each other can never catch a family missing from both -- the assertion has to come from outside. Verified by mutation: removing claude-local again fails the test. Also from the same review: - sourceSatisfiedBy returns the matched source string, so an empty-string source would return '' and the caller's `if (hit)` would silently discard a real match. Unreachable today (every rule source is a non-empty template) but a footgun for the next rule author; now `!== null`. - the purity fixture used a single-element changedPaths array, so an in-place sort would have been invisible. Now three elements in unsorted order. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * fix(#2723): resolve the base ref tolerantly instead of hard-requiring origin/next The first matrix run failed on both linux lanes: differential attribution over the real tree cannot resolve origin/next (Command failed: git rev-parse origin/next) Not a flake, and not an environment excuse -- a real defect in this diff. The gsd-test runner shallow-clones and merges base+head, so no origin/* remote- tracking refs exist in the container. My own fail-loud path fired correctly; what was wrong was hard-depending on that ref existing. GitHub Actions has the same shape by default, which is exactly why changeset-required.yml carries an explicit `git fetch origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"`. Now resolved through an ordered candidate list -- GSD_EMITTED_BASE (explicit lane override), then origin/$GITHUB_BASE_REF and $GITHUB_BASE_REF, then origin/next and next -- de-duplicated, each verified with `rev-parse --verify <ref>^{commit}`. When NO candidate resolves the test takes an explicit t.skip() naming every ref it tried and stating that the gate did not run here. That is the ADR-2719 section 6 distinction: t.skip is REPORTED as skipped, whereas a bare return is scored as a PASS. Hard-failing was the other option and is wrong -- it would make the suite permanently red wherever a base ref cannot exist by construction, which is a statement about the checkout, not a propagation finding. The candidate ordering is pinned by a unit test rather than left implicit, since the ordering IS the fix. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1f6822ccba |
test(#2722): emitted-artifact provenance table with a totality guard (#2735)
* test(#2722): emitted-artifact provenance table with a totality guard Adds the declarative emitted-path -> source-path table that ADR-2719 §2 specifies, plus the totality guard that keeps it honest. Phase 2 of #2719. Every emitted path across all 19 committed golden-parity manifests (8,524 paths) must match exactly one rule. Zero matches, two matches, and a rule matching nothing are all hard failures, so a new emitted family fails the build loudly instead of passing through unattributed. The measured surface is larger than #2722 estimated from claude.json alone (26 top-level families across 19 runtimes, not 13), which is itself what the totality guard exists to surface. It resolves to 19 rules. Building the table caught three false attributions that were total but resolved to repo files that do not exist -- Copilot's `<name>.agent.md` rename, Kimi's code-literal `agents/gsd.{yaml,md}` root agent, and Copilot's `hooks/gsd-session.json` registration. The "every attributed source exists" test is therefore a first-class gate, not a nicety. Notable correctness decisions: - Emitted shapes are hard-coded; deriving them from the installer would make the guard tautological (it would follow any installer change silently). Only source paths read a first-party descriptor, and only where the descriptor is the sole declaration (hostBehaviors.nativePlugin.source). - Emitted skills attribute to commands/gsd/*.md, NOT the repo skills/ dir -- that directory is generated from commands/gsd by gen-plugin-skills.cjs, so attributing to it would be false attribution that still passes totality. - Attribution is keyed on (rel, runtime): plugins/gsd-core.js has different sources for opencode and kilo. - Rule order carries no semantics (property-tested), since exactly-one matching is enforced rather than first-match-wins. Nothing here reads a git diff, builds a live manifest, or touches a fixture; the differential check, drift-ack file and size ratchet are #2723. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * docs(#2722): record the delivered provenance table in the CONTEXT.md glossary The `### Emitted Artifact Provenance` entry landed in #2721 describing the table as future work. Phase 2 delivers it, so the glossary now records what actually exists and the invariants #2723 must preserve: - where the table lives, its rule count, and that it is total over all 8,524 emitted paths across the 19 manifests - dead-rule detection, so table rot is loud in both directions - the corrected surface measurement (26 families, not the 13 estimated from claude.json alone) - the two invariants #2723 inherits: shapes hard-coded (deriving them would make the guard tautological), and attribution keyed on (rel, runtime) - the skills/ false-attribution trap, and that totality does NOT catch a wrong-source rule — the source-existence assertion is what does Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 * test(#2722): close five review findings on the provenance table Two orthogonal review passes plus an isolated adversarial reviewer returned findings at minor..major. All fixed; no blockers were raised. Standards axis (CONTEXT.md:456, RULESET.TESTS.guard-toplevel-readFileSync): - module-level loadManifests() threw at require time before any test() registered, turning a missing fixture dir into an opaque crash instead of one named failure. Now a memoized lazy accessor. - matchRules and assertTotality each carried their own copy of the matching loop; assertTotality now calls matchRules. That is the #2266 divergence class, and two copies could let the guard and the attributor disagree. - named the corpus stride constant; dropped an inline require. Spec axis: - the CONTEXT.md glossary carried a "26 families" figure that is not reproducible from the code and that no test pinned -- a hand-maintained number in permanent canon, i.e. exactly the silent drift this epic exists to end. All volatile counts are now removed from the glossary, with the reason stated inline: the guard recomputes them every run, so they belong in a failure message, not in prose. No test was added to pin the count, because that would rebuild the brittle committed number we are deleting. Isolated adversarial review: - `.+` tail captures let a `..` segment reach a constructed source path that resolves outside the repo. Not live-exploitable (fixtures are committed and the only consumer is an existsSync probe) but Phase 3 feeds these strings into a diff-consuming check, so assertSafeRelPath now fails closed once, in matchRules, rather than per-rule. - attributeEmittedPath's ambiguous-match branch was never exercised; only assertTotality's parallel path was. Now tested directly. - sampleLimit's truncation branch had no limit-1/limit/limit+1 coverage. - the fast-check property could not fail for the reason it was named for. That last one took two attempts and is the one worth reading. The property hand-rolled its shuffled side from the per-rule matchOne primitive, which is order-independent by construction, so it held for reasons unrelated to the shipped matchRules. Routing it through the real matchRules was still not enough: on an unambiguous table, first-match-wins and collect-all return identical results for every path (measured: 0 of 190 corpus paths differ). Order can only matter where more than one rule matches, so the property now also asserts that an intentionally ambiguous table reports BOTH hits as a set under every permutation. Verified by mutation -- injecting a `break` into matchRules makes it fail, and restoring makes it pass. Enabling all of the above: matchRules and attributeEmittedPath now take an injectable rules table, so tests can drive the real code path instead of re-implementing it by hand. Refs #2719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bf8f320083 |
feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI) PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting GSD's kimi support into two distinct products per the user's directive: - kimi (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python) - kimi-code (new): Moonshot's Node Kimi Code CLI (~/.kimi-code, runtime: node, KIMI_CODE_HOME env) Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json descriptors, not hardcoded branches in install.js. The new descriptor uses the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs). Critical Kimi Code constraint reflected in the descriptor: hostIntegration.dispatch.namedDispatch: false hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan'] hostBehaviors.namedSubagentsSupported: false Kimi Code's official docs confirm only 3 built-in subagents with NO custom- subagent registration (the [subagent] table only has timeout_ms). The kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in kimi-code's artifactLayout. Schema adjustments: - subagentToolkit set to 'undocumented' (the existing escape hatch); the schema enum (full/read-only) lacks a 'limited'/'built-in-only' value. A follow-up PR can extend the schema enum to add 'built-in-only' as a first-class axis value reflecting Kimi Code's documented model. Registration: - capabilities/kimi-code/capability.json (new descriptor, modeled on codex) - bin/install.js: allRuntimes array + --all list + --kimi-code flag - gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases (kimi-code, kimicode, kimi_code) - src/runtime-name-policy.cts: FALLBACK_ALIASES map - gsd-core/bin/lib/capability-registry.cjs: regenerated via scripts/gen-capability-registry.cjs --write Tests: - tests/multi-runtime-select.test.cjs updated for the new runtime count (18) + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19. Out of scope for PR 1 (follow-up PRs in the sequence): - Install-time decision logic (kimi vs kimi-code detection / prompt) - agent-install-check semantics for kimi-code (verify Agent Skills presence) - cmdAgentSkills fallback returning subagent prompt content - Workflow template mapping (named agents → built-in coder/explore/plan) - Migration guidance for users currently on 'kimi' who are actually on Kimi Code - Schema enum extension for subagentToolkit: 'built-in-only' Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS). * fix(#2454): complete drift-guard registrations for kimi-code runtime The drift guards caught every surface that pins runtime enumeration. Each update is mechanical, driven by the guard's named failure mode: - src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code - src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the isKimiCode predicate generator - bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19 - gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry (null/null/null — same as kimi, no model tier defaults until configured) - docs/reference/capability-matrix.md: regenerated via scripts/gen-capability-matrix.cjs --write (kimi-code row added) - tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP: kimi-code → '.kimi-code' - tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated) The capability-registry is already regenerated from the prior commit. * test(#2454): update drift-guard tests for kimi-code runtime registration Multiple drift guards pin runtime enumeration counts and option numbering. Each update is mechanical, driven by the guard's named failure mode: - tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17); 'all 16 flags' → 'all 17 flags' in test names + messages. - tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice test for kimi-code (option 11). Prompt test updated for new numbering. - tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs); EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per docs, same as Python kimi/opencode). - tests/global-config-home-fragment.test.cjs: table-count test renamed 13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier). * fix(#2454): empty artifactLayout for kimi-code (PR 1 scope) The skills kind requires a converter (existing converters are per-runtime like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only registers the descriptor; the actual Agent Skills converter (and a new 'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside the install-time decision logic. Empty artifactLayout.global is valid and means 'nothing to install yet via the layout seam'. Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs (localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as option 11 in install.js's buildRuntimePromptText (renumbered downstream options 11..17 → 12..18, All 18 → 19). * fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode) The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved) for the new kimi-code runtime id. Property names with hyphens are awkward for consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing the conventional PascalCase flag name. The 16 prior single-word runtime ids are unaffected (the regex finds no hyphens). * fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures - tests/runtime-flags.test.cjs drift guard: use proper kebab-case conversion (isKimiCode → kimi-code, not 'kimicode') so the registry comparison doesn't false-positive on hyphenated runtime ids. - tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae single-choice tests for the renumbered options (kilo 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16). - tests/install.test.cjs: Kilo integration option 11→12, prompt test regex updated. - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1. The kimi-code install produces the standard GSD install layout (skills, contexts, references, etc.) — 436 paths, same shape as other runtimes that have no custom converter yet. * fix(#2454): add kimi-code install contract + global config home fragment - src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path instead of falling through to the default '.claude'. - tests/installer-migration-install.integration.test.cjs RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for PR 1; PR 2 will specialize once the Agent Skills converter lands). - tests/multi-runtime-select.test.cjs: fix space-separated-choices test for the renumbered kilo option (11 → 12). - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: regenerated after rebasing onto current next (new planner-reversibility.md from #2471 etc. now included). * test(#2454): skip kimi-code install contract until PR 2 ships install layout The end-to-end install test (tests/installer-migration-install.integration .test.cjs) asserts every allRuntimes entry installs a runtime-specific artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the capability descriptor + flags + labels, but the install LAYOUT (Agent Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and self-removing — PR 2 removes the entry alongside adding the install surface, restoring the contract loop to full coverage. * fix(#2454): restore compact model-catalog.json format (M1 review) Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations' commit used python json.dump(indent=2) which inflated the file from 165→607 lines (every nested entry got expanded) and lost the trailing newline. The semantic change was just a 3-line kimi-code entry. Restored the original hybrid format (top-level indent=2 + inner entries' one-line style) and added kimi-code in matching form. Regenerated golden install parity + install tree fixtures since the model-catalog.json hash changed. * fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code) CI lint-tests job failed on the glossary drift guard (scripts/check-glossary-refs.cjs --check): ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but bin/install.js's allRuntimes array has 18. ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js (missing from CONTEXT.md's list: kimi-code). Missed in the prior commits because gsd-test does not run the glossary check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims to 18 values + kimi-code in the member list. * chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511) * docs(changeset): Phase 1 kimi-code runtime Added (#2511) * test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands * docs(changeset): backfill PR #2519 for Phase 1 (#2511) |
||
|
|
89b1bef881 |
refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267. |
||
|
|
6a474db3aa |
refactor(#2266): single-source golden-parity manifest builder + fixture correction (#2273)
Phase 1 of golden-install-parity redesign (epic #2264). Consolidates buildParityManifest + exclusion constants into tests/helpers/install-shared.cjs (fixes realRoot divergence), adds anti-divergence guard, corrects 12 stale golden fixtures to portable values. Closes #2266. |
||
|
|
51b6e3c35c |
test(#2136): add localToday regression + mirror fake/inline clock stubs
Adds a dedicated regression (tests/fix-2136-clock-local-today.test.cjs): - realClock.localToday() returns the LOCAL calendar day under a pinned instant + TZ (America/Chicago → 2020-06-14, the issue's repro instant). - realClock.today() is unchanged (UTC) — internal/cosmetic stamps stay UTC. - localToday === today on a UTC host (no spurious divergence). - makeFakeClock mirrors localToday (drop-in Clock substitute). - field-level: a 'sync' transition with a split clock (today≠localToday) writes the LOCAL day into Last Activity, proving the wiring uses localToday. Mirrors localToday() in the fake clock helper and the inline fixedClock stubs in state-transition.test.cjs / state-rebuild.test.cjs (the Clock interface now requires it). |
||
|
|
79d7657eff |
feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.
Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
slash-programmatic / active-model / native-extension / bun) + hostBehaviors
{nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
(global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
markdown); the 16 other fixtures + claude-local change only by the shared
model-catalog hash line.
Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
(every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
consumes params; getArgumentCompletions; before_provider_request active-model
steering (fail-open on null resolution); functional session_start /
before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
vocabulary.
Docs (host-integration matrix + how-to) + changeset (Added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
6e773d97df |
feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration Interface (declarative embedding adapter → engine surface dispatch) and fold every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven `runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated (tests/fixtures/golden-install-parity/codex.json); no other runtime changes. Three Context7-verified upgrades, each with a test on the user-reachable surface: - Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override, with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest, and the skill-manifest inventory to honor the override so --skills-root / sync-skills / the manifest report the real location. - Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact, PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall; extendedHookEvents reconciled [] -> the schema-valid wired subset. - Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own AgentsToml scalars (max_threads etc.) instead of dropping them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
69b309e4e0 |
feat(#1925): add ZCode (Z.ai) as a pluggable runtime descriptor
Add ZCode as a first-party runtime via a declarative capability descriptor (capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode' branches — exercising the de-hardcoded, data-driven runtime path that 1.7.0 (ADR-1016 / ADR-1239) enables. Descriptor (all axes sourced verbatim from zcode.z.ai docs): - configHome ~/.zcode; nested skills + flat commands/agents; profile-marker install - Claude-shaped skill format reuses convertClaudeCommandToClaudeSkill (no new converter) - hostIntegration: declarative / slash-file / mcp / electron; dispatch background=false (foreground-only per docs); nested+maxDepth undocumented; passive model mode Installer registration (data, not branches): --zcode flag, allRuntimes, runtimeMap, interactive menu, --all list. getGlobalConfigDir + resolveRuntimeArtifactLayout + ALLOWED_CONFIG_RUNTIMES + resolveInstallPlan all derive from the descriptor. Revamped the brittle per-runtime golden-master tests to be count-agnostic, descriptor-derived property tests (1.7.0 makes runtimes pluggable data, so pinning frozen '15 runtime' snapshots is the wrong invariant): getdirname, label-policy, config-adapter-registry (intent + install-plan golden master), capability-registry, host-integration-descriptors (counts derive from curated maps). Adding a runtime descriptor now extends coverage with zero edits to those suites. Welcome banner, --zcode help, supported-runtimes how-to, and the host-integration capability matrix (every axis cited) updated. |
||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fc5ca178a2 |
test(#1178): consolidate duplicated agent-roster helper into tests/helpers (#1420)
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites through it. Semantically-different sites (installed-dest dirs, absolute-path returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform multi-family inventory table) are left intact, each with a one-line comment. Test-only; no production code touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1178): note AGENTS_DIR export is for future call sites Review nit: clarify that the currently-unused AGENTS_DIR export is intentional — available for future tests needing the canonical source agents path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
fc2a7c0555 | fix(#1615): install Windsurf slash workflows | ||
|
|
33ccf5f89d |
fix(#1367): project-local install uses flat gsd-<cmd>.md layout (fixes /gsd: colon namespace) (#1489)
* fix(#1367): project-local install uses flat gsd-<cmd>.md layout Claude Code project-local installs now write command files as flat gsd-<cmd>.md at .claude/commands/ level instead of commands/gsd/<cmd>.md (subdirectory), so Claude Code registers /gsd-<cmd> (hyphen form) matching hooks, statusline, and all cross-command references. - capabilities/claude/capability.json: local destSubpath commands/gsd → commands - bin/install.js else branch: flat gsd-<stem>.md loop with runtime rewrites - bin/install.js uninstall (1c): remove flat files + legacy subdir cleanup - bin/install.js writeManifest: record flat commands/gsd-<cmd>.md keys - legacy migration: preserves dev-preferences.md across reinstall and uninstall - gsd-core/bin/lib/capability-registry.cjs: regenerated - 6 new regression tests (L0–L5) in bug-1367-*.test.cjs - Updated E suite in bug-3683 + bug-1736, layout + surface + descriptor tests - scripts/lint-regression-test-names.allowlist.json: grandfathered bug-1367 test Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1367): add issue reference to allow-test-rule comment Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
9e5d4b266b |
fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs (#1207)
* fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs via SessionStart hook Claude Code marketplace plugin installs unpack the package into the version-pinned plugin cache and never run bin/install.js, so ~/.claude/gsd-core/ is never created. Agents, commands, and templates markdown-@-include the canonical ~/.claude/gsd-core/... path (which expands ~ but NOT ${CLAUDE_PLUGIN_ROOT}), so every include resolved to nothing and agents (e.g. the executor) failed. Add a SessionStart hook (hooks/gsd-ensure-canonical-path.js) that, on a plugin install, symlinks the canonical path's immutable subdirs (bin, contexts, references, templates, workflows) to the plugin's bundled gsd-core/ tree. It changes zero @-references, is a no-op in classic installs, preserves user-generated files (USER-PROFILE.md, STATE.md), prunes stale links so it self-heals after `claude plugin update`, uses Windows junctions, and rejects bundled/canonical paths that escape the resolved plugin root (no traversal, no clobber). Registered in HOOKS_TO_COPY (build-hooks), MANAGED_HOOKS, hooks.json SessionStart (runs first, timeout 5), and BUNDLED_GSD_HOOK_FILES. Behavioral regression tests folded into issue-766-plugin-manifest.test.cjs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#997): backfill changeset PR number to #1207 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b4a7eabaae |
feat(#1085): migrate windsurf workspace skills to .devin/ + fix global content refs (#1093)
Fresh windsurf/devin-desktop workspace installs write skills under .devin/ (legacy .windsurf/ recognized); global ~/.codeium/windsurf/ unchanged. Also threads real isGlobal through _applyRuntimeRewrites so global skill content references the codeium path. Closes #1085. |
||
|
|
77a671ec53 |
feat(#791): migrate antigravity workspace base dir .agent → .agents (#1090)
Fresh antigravity workspace installs write under the canonical .agents/ (plural) base; legacy .agent/ stays recognized (dual-read). Global ~/.gemini/antigravity/ path unchanged. Closes #791. |
||
|
|
f2dd125593 | Merge branch 'next' into kimi-runtime-support | ||
|
|
b6199460ed |
refactor(#887): consolidate duplicated CHILD_ROUTER test maps into shared helper (#889)
Three test files copy-pasted the concrete-skill->namespace-router CHILD_ROUTER map verbatim, and two more duplicated an identical local parseRouterRequires regex. Introduce tests/helpers/nested-layout.cjs that derives the child->router map from the authoritative commands/gsd/ns-*.md requires: lists once, reusing the production parseRequires (now exported from install-profiles) plus a nestedSkillPath(skillsRoot, prefix, stem) helper. Refactor all five test files to import from it. Test-only + a single internal export addition; no production behavior change. Closes #887 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
54191f6c6c | Merge branch 'next' into kimi-runtime-support | ||
|
|
1b6bd66f2c |
feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) (#821)
* feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) Wire three new context-tracking events (SubagentStop, Stop, PreCompact) to gsd-context-monitor so context-headroom warnings surface at model-stop and subagent-finalisation moments — not just on PostToolUse. Add a new FileChanged hook (gsd-config-reload.js) that hot-reloads .planning/config.json context mid-session when the user edits it, injecting a config summary as hookSpecificOutput.additionalContext. Updates plugin manifest hooks.json, managed-hooks-registry, installer-migration-report allowlist, and shell-command-projection cleanup tables. Tests: 21 new assertions in enh-770-claude-hook-events.test.cjs; enh-788 and issue-766 test suites updated. Closes #770 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#770): document newly-registered Claude Code lifecycle hooks Add a Hook coverage table to the Claude Code npm installer section of docs/how-to/install-on-your-runtime.md describing SubagentStop, Stop, PreCompact, and the new FileChanged (gsd-config-reload.js) hook that hot-reloads .planning/config.json mid-session. Also fixes the changeset frontmatter (adds type: Added + pr: 821) so docs-lint can consume the fragment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): add gsd-config-reload.js to INVENTORY.md and regenerate manifest The feat commit added hooks/gsd-config-reload.js but did not bump the Hooks count in docs/INVENTORY.md (14→15) or add the new row, and did not regenerate docs/INVENTORY-MANIFEST.json. Both inventory-counts and inventory-manifest-sync tests failed across the full CI matrix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): make lifecycle-hook tests deterministic on scoped runner Replace the shared hooks/dist/ ensemble setup (ensureHooksDist / teardownHooksDist) in the Claude hook tests with per-test isolation: pre-populate each test's own tmpDir/.claude/hooks/ with stub files and pass installerMigrations:[] to install() so the first-time-baseline migration does not remove the stubs before the copy step can run. Root cause: hooks/dist/ is gitignored and absent on a fresh npm ci. ensureHooksDist() created it and teardownHooksDist() deleted it, but with --test-concurrency=4 both test files ran concurrently as separate Node.js worker processes sharing the same filesystem. One file's afterEach teardown deleted hooks/dist/ while the other file's install() was copying from it, producing an ENOENT (reproduced 2/10 runs locally). The additional issue: even with pre-placed stubs surviving the copy race, the 000-first-time-baseline migration classified hooks/gsd-*.js as bundled-gsd-hook artifacts, auto-removed them, and the copy step never re-ran (hooks/dist/ absent) — leaving contextMonitorFile missing and all hook registrations silently skipped (the 'got: []' symptom). Fix: pre-populate targetDir/hooks/ per-test (isolated temp dir) AND pass installerMigrations:[] so the baseline scan is skipped. The Qwen suites already used this pattern correctly; the Claude suites are aligned to it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): ship gsd-config-reload.js by adding it to build-hooks HOOKS_TO_COPY The #770 feature added hooks/gsd-config-reload.js and registered it in MANAGED_HOOKS, the installer, INVENTORY, and the test EXPECTED_ALL_HOOKS list — but never added it to scripts/build-hooks.js HOOKS_TO_COPY. As a result the hook was never copied into hooks/dist/ during the build, so: - the hook would never ship to users (real production bug — the FileChanged config-reload feature was dead-on-arrival), and - install-minimal-hooks.test.cjs #1755 ("all expected hooks are copied from hooks/dist/ to target", ".js hooks are executable after copy", "manifest contains .js hook entries") failed on any environment with a clean checkout (no pre-existing hooks/dist/): coverage, full test macos-22/macos-24, test ubuntu-24. The failures were masked locally only by a stale hooks/dist/ left from a prior build (build-hooks copies into dist without clearing it). On CI's fresh `npm ci` there is no dist, so the omission surfaced. Fix: add 'gsd-config-reload.js' to HOOKS_TO_COPY so build-hooks stages it into hooks/dist/ alongside the other JS hooks. Verified by removing hooks/dist/ and rerunning the full suite green (0 fail). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): make config prototype-pollution beforeEach deterministic on scoped runner Root cause: the #663 and alert-#26 prototype-pollution describe blocks seeded .planning/config.json in beforeEach via a bare runGsdTools('config-ensure-section') whose result was discarded. That command runs in a spawned gsd-tools child; on the scoped CI lane (--test-concurrency=4, config.test.cjs scheduled alongside the heavy install/tarball suites that #770 pulled into the targeted set) the child can be transiently killed under resource pressure (non-zero exit, empty stderr — an OS-level kill, not an app error). The swallowed failure left config.json absent, so the first subtest's readConfig() threw ENOENT opening <tmp>/.planning/config.json. Only 1 of 4 subtests failed, confirming a per-invocation transient, not a deterministic miss; the full suite schedules files differently so config.test.cjs did not collide with those heavy neighbors → passed there. Fix: add ensureConfigReady(tmpDir) which retries config-ensure-section on ANY failure or missing file and throws a clear diagnostic if it still cannot create config.json, then use it in both prototype-pollution beforeEach blocks. Setup is now deterministic under load; the #663/alert-#26 security assertions are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |