Commit Graph

63 Commits

Author SHA1 Message Date
Tom Boucher
7c649a9970 fix(#3585): close raw-git bypasses of the commit_docs gate (#3590)
* test(#3585): repo-wide guard for unguarded .planning/ git add

Replaces the two-file #1783 scan, which required .planning/ on the git add
line and so was structurally blind to fast.md's `git add -A` and to
new-milestone.md (never scanned).

Extracts the shell tokenizer, comment-position rule and gsd-scan-ignore
marker from the #2269 guard into tests/helpers/shipped-command-scan.cjs so
both guards consume one implementation. Commit-specific logic stays in
commit-files-pathspec.test.cjs; every pre-existing test there passes
unedited.

Fails RED on five sites: fast.md:58, new-milestone.md:262, spec-phase.md:480,
eval-review.md:148, ai-integration-phase.md:263. The last three carry a
markdown prose conditional outside the bash block it claims to guard.

* fix(#3585): close raw-git bypasses of the commit_docs gate

Five shipped workflow steps staged .planning/ with raw git. Two had no
check at all; three had a markdown prose conditional sitting outside the
bash block it claimed to guard, so the block ran unconditionally.

spec-phase, eval-review and ai-integration-phase now route through the
gsd_run query commit seam, which performs the commit_docs and gitignore
checks internally and returns a skipped envelope -- this deletes the raw
git pair rather than wrapping it.

new-milestone stages directories for a later commit and cannot use the
seam, so it takes the executable guard form, fail-open on a tooling error.

fast writes no planning artifacts and has no gsd_run in scope at that
point, so it excludes .planning via pathspec instead of reading config.

Guard now reports 0 offenders.

* test(#3585): pin skipped_gitignored to production behavior

COMMIT_REASON was a test-local frozen enum joined to production only by a
hand-maintained keep-in-sync comment -- the Generative Fix Divergence class,
whose required remedy is a parity assertion.

B1-B3 already pinned SKIPPED_COMMIT_DOCS_FALSE. SKIPPED_GITIGNORED was
pinned by nothing: production could rename it and every test still passed.

G1-G3 drive the gitignore auto-detect path and assert the canonical reason.
The fixture must OMIT .planning/config.json entirely -- with config.json
present the loader resolves commit_docs to false first and cmdCommit returns
skipped_commit_docs_false, never reaching its own isGitIgnored branch.

* docs(#3585): document the planning commit gate and its guard

CONTEXT.md had zero commit_docs entries. Adds a Planning Commit Gate
glossary entry covering the resolution chain, the typed skip envelope, the
measured ordering of the two reason codes, and why the gate is enforceable
only as a text guard.

CONTRIBUTING.md gains the contributor rule for the new guard, with the
prose-is-not-a-guard example that caused three of the five defects.

* fix(#3585): address review findings in the planning-add guard

Spec review (blocker): fast.md excluded .planning unconditionally, changing
behavior for commit_docs=true users and violating epic AC4. Now gated -- the
launcher preamble was MOVED from log_to_state into the commit block rather
than copied, so gsd_run is in scope for +4 lines instead of +4KB, and the
else branch is byte-identical to the previous git add -A.

Security review (major): git -C <dir> add was a false negative because the
flag-skip loop never modelled flags that consume a separate value. Fixed for
-C/-c/--git-dir/--work-tree/--namespace. The fail-closed rule now also covers
$(...) substitution args and --pathspec-from-file, which were opaque in the
same way $VAR is. git commit -a/-am is now classified as reaching, since it
stages every tracked modification.

Self-review: isSkippable treated any NAME= token as a skippable prefix, so
V=$(git add -A) escaped -- the exact divergence the shared-helper extraction
existed to prevent. Adopted the sibling predicate verbatim.

eval, xargs, one-line function bodies and line-continuation remain blind and
are now enumerated as declared limits in the guard docblock and CONTRIBUTING.
The ifDepth clamp is defensive only: a 200k-case differential fuzz found no
reproducing input, so its test is labeled a pin, not a failing-first test.

* test(#3585): acknowledge emitted growth in three workflow files

emitted-attribution has two arms: hash attribution AND per-file growth. The
growth arm needs an acknowledgment even when every moved byte is attributable
to the diff, which is why the first remote run went red on it.

fast.md +417: the launcher preamble moved into the commit block so gsd_run is
in scope for the commit_docs guard, plus the guard itself.
new-milestone.md +281: the executable guard plus one line recording that the
unstaged archive move is deliberate.
spec-phase.md +21: reworded prose describing the skipped envelope.

eval-review.md and ai-integration-phase.md shrank; no entry needed.

* test(#3585): drop duplicate spec-phase ack, shrink its prose instead

The base already acknowledges spec-phase.md (from #2733), and two ack sources
may never name the same path. But a base-side ack is SPENT -- it cannot clear
new growth -- so the two gates were in direct conflict: attribution wanted an
ack, the ack lint forbade one.

Resolved by removing the growth rather than the conflict. spec-phase.md's +21
was purely a prose reword; rewritten shorter, the file now shrinks 36 bytes
against base and needs no acknowledgment at all.

fast.md and new-milestone.md have no base ack and keep theirs.

* chore(#3585): backfill changeset pr number to 3590

---------

Co-authored-by: sim <sim@local>
2026-08-17 13:34:55 -04:00
Tom Boucher
9448736872 fix(#3547): exercise the real global config-home shape in the install harness (#3567)
* test(#3547): failing-first regression for collapsed global install shape

* fix(#3547): exercise the real global config-home shape in the install harness

* test(#3547): align ripple suites with the real global install shape

* test(#3547): update stale collapsed-shape pins in provenance and migration suites

* fix(#3547): bump emitted-baseline schema version for the real install shape

---------

Co-authored-by: sim <sim@local>
2026-08-16 01:14:26 -04:00
sim
147856040b fix(#2873): close review findings across fences, sanitizer and docs
Isolated security review found resolveSpecRootReference's fence tracker
toggled on any delimiter, so a backtick fence could be closed by a tilde
one and an include in the gap was rewritten inside a code block. Fixed by
reusing scanFencedBlocks - the canonical engine already behind
stripFencedCode and extractFencedBlock - rather than carrying a fourth
copy of fence detection, which also closes the duplication the standards
review flagged.

sanitizeForRender now strips combining marks and zero-width characters
alongside the ANSI, control and bidi classes it already handled.

Adds the C, E and F matrix rows the spec review found missing, including
installer-level coverage that spawns the real install rather than calling
the report builder. Ships the how-to, the reference and command docs in
five locales, the changeset, the inventory and glossary entries, and
regenerates health.md for the new W028 rule.

Refs #2873
2026-08-14 23:48:39 -04:00
Tom Boucher
dc3c81e93d chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam

Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts
and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both
suites fail with MODULE_NOT_FOUND, which is the intended RED.

Locks the measured behavior rather than the assumed behavior:
RegExp.escape hex-escapes the leading character of nearly every string
("abc" -> "\x61bc"), so the suite asserts match-equivalence against an
inlined historical oracle (the implementation being deleted) rather
than byte-equivalence of pattern text — 200 seeded fast-check runs plus
a fixed corpus, 0 mismatches. Also locks the latent character-class
range bug this phase fixes as a side effect: a hyphen-bearing value
interpolated into [...] currently forms a real range and matches an
unintended character; post-migration it must not.

* chore(#3412): src/pattern.cts owns runtime-value regex construction

Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam
delegating to the built-in RegExp.escape, deletes every hand-rolled
copy, and raises the Node floor to the Active LTS line.

The census was low, three times over. ADR-3212 counted 10 copies; a
graph query found 12; the new lint rule — once live — found 27 more.
The difference is that the census counted named helper FUNCTIONS while
the rule counts the escape SHAPE, so inline .replace(<class>, '\$&')
copies were never in scope. ADR §1's actual requirement is that no
module outside the seam escapes a value for regex use, so all of them
are, and CLAUDE.md's no-defer rule makes them this change's work.
Fourth consecutive epic here whose copy count was low — the argument
for ADR-3180 Amendment 3's "state N found by the guard" rule.

Also corrected mid-implementation: the survey reported phase-id.cts's
escapeRegex had 0 external importers. It had 8 production importers,
making its removal a public-surface change to an ADR-2121-owned module
and requiring an update to that ADR's locked-surface test. Blast
radius revised Medium-High -> High.

RegExp.escape is match-equivalent but NOT text-equivalent: it
hex-escapes the leading char of nearly every string ("abc" ->
"\x61bc"). Equivalence is proven by a seeded fast-check property test
against the deleted implementation as oracle. It also fixes a latent
bug: a hyphen-bearing value interpolated into a character class
previously formed a real range and matched an unintended character.

Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines,
.nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate
`required-tests` context is unchanged and no job was added or removed,
so branch protection cannot be orphaned by the dropped lanes.

Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with
structural provenance for reviewed pattern-fragment constants rather
than a name heuristic) plus a whole-tree companion guard covering the
directories ESLint's globs miss.

* fix(#3412): close the _SOURCE guard evasion, correct two false claims

Three findings from the orthogonal review pass, all fixed.

1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier-
   name matching with no binding check, so `new RegExp(userInput_SOURCE)`
   — a function parameter — sailed past the guard. That is the same
   rename-evasion class issue #3410 documents, reopened by the very
   fallback meant to complement the structural check. Now bound to the
   identifier's actual binding kind: import, require-derived const, or
   module-scope const; parameters, `let`/`var`, and unresolvable
   bindings fail closed. Four RuleTester cases cover the evasion and
   prove the legitimate cross-module case still passes.

2. src/pattern.cts's own header carried the stale pre-correction counts
   (12 copies / 17 call sites) while CONTEXT.md and the design doc
   carried the corrected ones (~39 / ~44) — a self-contradiction inside
   the PR whose entire purpose is deleting divergent copies. Rewritten,
   preserving the durable lesson: a named-function census cannot see
   inline copies; only a shape-matching guard can.

3. The claim that all deleted copies threw TypeError on non-string was
   false. phase-id.cts's copy — the one with 8 external importers — did
   String(value).replace(...) and never threw. The seam's locked
   signature does not coerce, so this is a real, now-disclosed behavior
   change rather than the pure preservation the tests asserted. Audited
   all 32 invocations across the 8 importers and 6 in-file callers:
   every one is safe by construction (upstream truthy guard or a
   string-producing derivation), verified by runtime probe against the
   compiled modules rather than by TS compilation, which cannot see a
   runtime undefined. Corrected the false claim in both the test comment
   and the design doc, and added it to Known limits.

* docs(#3412): add Changed changeset for the Node 24 floor

The only user-visible break in this phase. The escape-behavior change
is internal and match-equivalent, so it carries no user-facing note.

* fix(#3412): resolve the seam's require graph in script fixtures and packaging

Checkpoint 2 came back red with 90 failures on the node24 lane. Three
distinct defects, all introduced by routing scripts/ through the new
pattern seam, none reproducible by any local gate:

1. ~82 failures — tests/adr-index-gate.test.cjs and
   tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an
   mkdtemp fixture and spawn it there (necessary: those scripts resolve
   their scan root from __dirname/.., so running the real script would
   scan the real repo). Each harness hand-listed the dependencies to
   copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to
   gen-adr-index.cjs made both lists silently incomplete ->
   MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON'
   failures from the same crash.

   Fixed as a class, not an instance: new tests/helpers/copy-script-
   fixture.cjs walks a script's transitive static relative-require graph
   and copies it, so dependencies are derived and never re-declared. It
   throws (naming the unbuilt artifact) instead of letting the child die
   with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming
   scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host-
   contract, sync-runtime-launcher.

2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so
   the new scripts/lint-no-adhoc-regex-escape.cjs would be
   MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from
   the tarball, matching the existing precedent for gen-emitted-
   baseline.cjs, which is excluded for the identical reason, and locked
   with a test modeled on that one. Confirmed against a real npm pack:
   890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs
   present (so the other four scripts' requires are legitimate).

3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped
   source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to
   the retired hand-rolled escaper but NOT text-equivalent: it hex-
   escapes the leading character and all hyphens ('0*\x329',
   '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match
   decisions across all three real interpolation prefixes, zero
   divergence. Those tests now compile each source into the same heading
   regex src/roadmap.cts's searchPhaseInContent builds and assert what
   matches and what does not, including the 'i'-flag canonicalization
   the hex escape has to preserve. Re-pinning the new literals would
   have rebuilt the same brittleness one layer down. Adds a test for the
   property the escape exists for: a dot in '1.2' must not act as a
   wildcard.

Also shares one definition of 'a require' between the packaging guard
and the fixture copier, so the two cannot disagree about what they scan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3412): refuse to copy a fixture dependency outside the fixture root

copyScriptWithDeps resolved each relative require and joined the
repo-relative result onto fixtureRoot. A require resolving OUTSIDE the
repo yields a '../'-prefixed relative path, so path.join climbed out of
the fixture and wrote into the surrounding temp dir (verified:
repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd).

No script in the tree does this today, so this closes an available
escape rather than an active one. Refuses via the existing unresolved-
require path so the failure names the offending specifier. Covered by a
negative proof that the guard fires and that nothing lands outside the
fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract

Applies all findings from the second orthogonal review round, re-run
because real code changed after round 1.

HIGH (security) — extractRequires stripped BLOCK comments before LINE
comments, so a '//' comment containing '/*' opened a phantom block
comment, and a '//' inside a string literal truncated the line. Both
hid real requires: 'const u="http://x"; require("./real.cjs")'
returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were
invisible. Replaced with a real AST parse via espree.

This is ADR-3212's own Decision 4 — tokenizer-first for stateful
grammars — applied to the case it describes; comment/string/regex
nesting is exactly such a grammar, which is why the regex version was
wrong. The function was moved byte-identical out of the #2858 packaging
guard, so the bug PRE-DATES this branch and has been a live blind spot
there: a shipped script could have required an unshipped path
undetected. Fixing it makes that guard strictly stronger than on next.

espree is promoted from a transitive eslint dependency to an explicit
devDependency rather than relying on hoisting. The script parse attempt
sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a
function, making a top-level return legal — scripts/check-coverage-gate
.cjs relies on it, and without the flag the guard throws on a file it
is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js
under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a
real npm pack, so the exact extractor does not newly fail the guard.

MEDIUM (security) — the repo-containment check guarded dependencies but
not the entry path. One escapesContainment predicate now guards both.

LOW (security) — containment was lexical while fs follows symlinks, and
a directory symlink could mint a fresh dedupe key per level. realpath
now resolves both repoRoot and each dependency before the decision, and
the realpath-derived path is the dedupe key. Destination layout still
uses the original repo-relative path, so copied trees are unchanged.

MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests
lost the foreign-prefix contract: every assertion was satisfied by an
impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599
bug class the exact-source prevents. The literal assertions it replaced
were catching this. Now asserts the compiled regex REJECTS a different
prefix with the same number.

MAJOR (standards) — the test hand-duplicated production's heading regex
with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed
the parallel surface instead of policing it: src/roadmap.cts exports
buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports
it. Byte-identical .source and .flags verified for both escaped forms.

MINOR — '..foo' no longer false-flagged as an escape; the inverted
spurious-vs-missing doc claim corrected; the dead allow-test-rule
header removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3412): backfill changeset pr number to 3416

* fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision

Two CI failures on PR #3416, both in code this branch added.

CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a
bracket run be consumed EITHER by the character-class branch OR one
character at a time by the trailing catch-all, so a failing match
explored both parses of every pair. Measured on the real regex:
n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script
scans repo source, so a file with a long bracket run after '.replace(/'
would hang CI outright — a guard against undisciplined pattern
construction was itself the worst pattern in the diff.

Fixed the way ADR-3212 already prescribes: the catch-all branch now
excludes '[' and ']' so a bracket can only be consumed by the class
branch (this is what makes it linear), and every quantifier is bounded
(the locked bounded-quantifiers decision) as a second line of defense.
Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the
constant: a regex literal with a BARE unescaped ']' outside a class is
no longer matched by this backstop. No census shape has that form, and
the AST rule remains the primary detector.

Verified the guard did not go blind doing it: a real census-shape
violation is still reported, and an allow-adhoc-regex-escape
suppression comment is still honored.

Regression test drives the exported findViolations on a
2000-repetition adversarial input and asserts the RESULT. It makes no
wall-clock assertion — elapsed-time tests are forbidden — so a
regression surfaces as a harness timeout, which is the correct signal.

Prompt injection scan — 'must not act as a regex wildcard' in a test
comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if|
my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a
whole test file over one phrase would blunt the scanner permanently,
and the comment has nothing to do with injection.

Neither failure was reachable from the remote runner — CodeQL and the
injection scan are not in that matrix, so the sha it passed was green
and still wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 16:19:57 -04:00
Tom Boucher
7a7bf19fc1 enhance(#2872): record scope and runtime in the install manifest (#3323)
* enhance(#2872): record scope and runtime in the install manifest

gsd-file-manifest.json gains manifestVersion, runtime and scope, and a new
read-only Installed Surface Resolver Module reads both install scopes for a
runtime in one call -- the first code path in the repo that does.

Phase 3 of epic #2866 (ADR-2866). Blocks Phase 4 (#2873), which resolves
#2218: the resolver's shadowedBy field is that defect expressed as a value
for the first time. It ships computed-and-unread here.

Installed-ness is decided by manifest PRESENCE, never by the new fields, so
a manifest written by an older GSD stays fully functional and no user needs
to reinstall. Recorded runtime/scope are corroboration: a disagreement with
the probed config dir is reported as declaredScopeMatchesProbe: false, never
silently corrected.

readInstallManifest is widened additively -- version/timestamp/mode/files
keep their exact names, types and meanings for all four existing callers.
manifestVersion is a new field rather than a reinterpretation of version,
which holds the package version and is read by the golden-parity fixtures.

Stems are derived from the installed manifest's own file keys, the inverse
of Phase 2's filename composition, guarded by a fast-check round-trip
property plus a kebab-case charset check so a crafted manifest key cannot
put a traversal segment, control character or ANSI escape into a trigger
that Phase 4 renders back to the user.

Also fixes two defects found while working:
- bin/install.js hardcoded manifestVersion: 2 while the reader owned
  MANIFEST_SCHEMA_VERSION = 2. Now single-sourced, with a parity test.
- docs/installer-migrations.md documented an install-state schema of five
  snake_case fields that have never been written; InstallState has only ever
  been { schemaVersion, appliedMigrations }. Corrected with a dated note.

Verification runs on the remote runner.

* fix(#2872): fold review findings from three independent engines

Standards axis:
- convert the manifest-schema suite from a hybrid setup(t) closure to
  beforeEach/afterEach (CONTRIBUTING.md:319-354 Pattern 1). The hybrid was
  neither approved pattern and a new test forgetting the call got no warning.
- SCOPE_ORDER was declared twice with no parity test -- this repo's recorded
  generative-fix-divergence class. Give the ordering one owner: install-scope
  exports it frozen, the layout module and the resolver both import it, and a
  test locks it against scopeRank so the constant and the ranks cannot drift.
- drop the defaultReadManifest passthrough (Middle Man).

Spec axis:
- add the VOLATILE_FILES exclusion test and source comment the acceptance
  table promised and did not deliver. gsd-file-manifest.json stays excluded:
  the new fields are deterministic, but timestamp -- the original reason --
  is unchanged.

Security axis:
- bound the reported manifest runtime at 64 chars, matching the
  truncatePostureValue convention already used in this subsystem. It reached
  declaredRuntime unbounded while the adjacent stems were gated by SAFE_STEM;
  an inconsistent posture on the same attacker-influenceable document. The
  charset stays ungated on purpose -- declaredRuntimeMatchesProbe needs to see
  the real value -- so Phase 4 must sanitize before rendering, recorded in the
  design's Known limits.

Both new parity tests were verified to FAIL when the two sides are made to
disagree, then pass again on revert. Verification runs on the remote runner.

* chore(#2872): backfill changeset pr number to 3323

* fix(#2872): give git fixture construction its own timeout class

PR #3323's full test (windows-latest, 22, shard 2/3) failed with

  gitOrThrow: 'git init' failed -- outcome=timed_out exitCode=null
  gitOrThrow: 'git commit --allow-empty' failed -- outcome=timed_out

from drift-detection.test.cjs's beforeEach, a file this branch never touched.
Every other lane passed the same commit, including windows-latest node 24 on
all three shards, and next is green.

Root cause is a bound sized for the wrong class. DEFAULT_GIT_TIMEOUT_MS is
15000 and its own comment scopes it to plumbing READS -- rev-parse, branch,
log -- against an existing repo. createFixture uses it for six sequential
repo-CONSTRUCTION spawns: init, three config writes, add -A, commit. init and
commit each write dozens of files, and on Windows every spawn is
Defender-scanned. Sibling tests in the failing block took 15.6-22.0s against
a 15000ms bound.

This repo already diagnosed this exact shape once: timeouts.cjs's
HOOK_FANOUT_TIMEOUT_MS records PR #3285 failing in the SAME job with the SAME
outcome=timed_out exitCode=null signature at the SAME bound while every other
lane passed, and concludes 'a bound sized for the wrong class, not a slow
machine'. It was fixed by splitting out a heavier class-norm at 60000. Same
remedy here: GIT_FIXTURE_TIMEOUT_MS = 60000, 4x the bound that failed and half
INSTALL_TIMEOUT_MS.

DEFAULT_GIT_TIMEOUT_MS deliberately stays at 15000 -- a blanket raise would
stop a genuinely hung plumbing read from surfacing fast.

Verified the value reaches the spawn rather than being an ignored option:
spawnSync was monkeypatched before requiring the fixture module, and all six
git construction calls were captured carrying timeout: 60000.

This branch's two new test files shift shard composition, which is how a
pre-existing fragility landed in the heaviest shard on the slowest lane.
Fixed here rather than deferred, per the no-defer rule.

Verification runs on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-10 15:50:55 -04:00
Tom Boucher
c28134ab39 fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host

Adds local/no-duplicate-fold-marker, an AST rule that reports the second and
every subsequent `folded:<name>` marker in a host file, plus RuleTester cases
and a tree-wide regression assertion.

Failing-first on purpose: the rule is registered at error and the 25 duplicated
regions are still present, so eslint and the new tree-wide test are RED. The
deletions land in the next commit.

The marker key is the whitespace-delimited token after `folded:` — not the
issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on
tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and
feat-443-effort-fast-mode are two distinct folded suites.

Refs #3271

* fix(#3271): delete 25 duplicated folded suites from three install hosts

Three consolidated install suites each carried a verbatim second copy of a
contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and
green — each duplicated block registered and ran twice on every lane.

  tests/install.test.cjs                   5981-9937  (3957 lines, 18 blocks)
  tests/install-minimal-hooks.test.cjs     2734-4015  (1282 lines,  5 blocks)
  tests/install-write-confinement.test.cjs 1754-2321  ( 568 lines,  2 blocks)

Introduced by 6d072435d (#1975 re-applying #1970's hunks on a tree that already
had them, 2026-07-03) — one stale-base re-application, three files, one commit.
Verified by marker-count bisect: 1 at 4f779eda4 and 0cc7a1a42, 2 from 6d072435d
onward.

The later copy is deleted in each case, so every file returns to what its
authoring batch produced and blame on the surviving lines stays accurate.
local/no-duplicate-fold-marker, red on the previous commit, is now green.

tests/model-resolver.test.cjs is untouched: the issue lists it, but its two
blocks are folded from two different files and are not identical. It is a false
positive of the issue's own grep, whose `[a-z0-9-]*` key truncates at `.`.

Fixes #3271

* test(#3271): property-test marker identity and pin the alias non-goal

Three review findings, all fixed inline:

1. foldMarkerOf is a parser and carried no fast-check property test. Raised
   independently by the /code-review standards axis and the isolated adversarial
   pass; the file already establishes the fc.property-driving-ruleTester idiom
   for a sibling rule. Added, two arms over markers generated from [a-z0-9-._]:
   the same marker twice always reports exactly once against firstLine 1, and
   two distinct markers never collide. The alphabet includes `.` on purpose —
   an implementation keyed on the issue's [a-z0-9-]* slice passes arm 1 and
   fails arm 2, which is exactly the model-resolver false positive.

2. meta.docs.category was the novel value 'Test hygiene'; all 16 sibling local
   rules use 'Best Practices', 'Portability' or 'Reliability'. Now
   'Best Practices'.

3. A call through a further alias (const d = __foldDescribe) was unreported and
   undocumented — accidental rather than deliberate. It is now the fourth entry
   in the rule's documented non-goals, with the reason, and pinned by a valid
   RuleTester case so it cannot drift silently.

Refs #3271

* test(#3271): name the step and elapsed time when a baseline build fails

buildBaselineAtRef runs four bounded steps and, when one exceeded its bound,
threw a bare "spawnSync ETIMEDOUT" naming neither the step nor how long
anything took. Diagnosing one real failure took four separate experiments to
recover information the throw already had.

Each step is now timed, and any throw carries the breakdown: which step failed,
its elapsed time, the timings of every step that completed before it, all three
bounds, and the tail of the child's captured stdout/stderr.

The failure message is deliberately the carrier. On the remote runner the
captured output field comes back empty in failures.json while error and stack
survive verbatim, so the message is the only channel that reaches a reader of a
remote verdict.

Refs #3271

* fix(#3271): size the baseline generator bound for the machine it runs on

Instrumentation from a real remote-runner failure gave the breakdown:

  git-worktree-add=15.1s  npm-run-build-lib=19.8s  gen-emitted-baseline=FAILED@300.1s

Steps 1 and 2 are comfortable. Only the generator exceeds its bound, and it is
not hung — it needs more than 300s there.

Measured ladder for that step: ~22s idle in a container, ~39s end-to-end in a
clean container, ~142s with 8 CPU burners on 8 cores, and >300s under the real
suite. Its cost is 19 sequential installer spawns, and spawn latency is exactly
where a container degrades worst (3.9x slower than host, against 1.1x for file
IO) — which is why a CPU-only load test did not reproduce it and why four
earlier hypotheses (container slowness, network, shallow clone, CPU contention)
all measured clean.

The 300s bound was sized on an idle machine for a step that never runs on one.
Under the remote runner the on-disk baseline cache is structurally absent — CI
restores it via actions/cache keyed on github.event.pull_request.base.sha, a key
that exists only inside GitHub Actions — so this slow path runs on every remote
verification. The result: this gate has passed 0 times in 754 runs, failing 80
times and never once executing successfully.

Raised to the 600000ms ceiling that local/no-unbounded-spawn treats as the
largest meaningful bound; the other two bounds are untouched. This makes the
gate RUN, which is the point: the alternative considered and rejected was
degrading the timeout to a skip, and that was measured to turn the suite green
with the gate silently not running at all.

The real remedy is making the cache reachable from the remote runner so the
in-job build returns to being the rare fallback ADR-2719 §5 describes. That is a
gsd-test-runner change, not one this repo can make.

Refs #3271

* fix(#3271): tolerate an overlay source that vanishes mid-walk

Observed on the remote runner, three runs across three different branches:

  ENOENT: no such file or directory, link '/work/hooks/dist/gsd-config-reload.js'
    -> '/tmp/gsd-2930-overlay-6nOZay/hooks/dist/gsd-config-reload.js'

buildOverlayRepo enumerates names with readdirSync and then acts on each one, so
statSync, copyFileSync and linkSync all sit in a TOCTOU window. hooks/dist is
regenerated by an ATOMIC REPLACE (scripts/build-hooks.js unlinks and renames), so
any concurrently running test that rebuilds hooks retires a just-listed name
mid-walk and the overlay dies on it. linkOrCopyFile already tolerated EXDEV and
EPERM; ENOENT went straight through.

On ENOENT the source is now re-examined ONCE rather than slept on. An atomic
rename is a single syscall, so by the time the failure surfaces the successor is
either already in place (the retry succeeds) or the path has genuinely left the
tree, in which case there is nothing to mirror and the leaf is skipped. No sleep
and no spin: a timing-based wait here would be the very flake being fixed. Every
other errno still propagates untouched, so a real permission or IO fault stays a
hard failure.

Five tests hold the boundary: gone-for-good skips without retrying, mid-replace
retries exactly once and places the file, EACCES still throws, a real linkSync
ENOENT is injected by monkeypatching fs and restoring it in a finally (never a
mode-bit trick, which root bypasses), and isMissingPath accepts only ENOENT.

Refs #3271

* fix(#3271): order the timeout ladder inward-out and lock it

Two review blockers, both real.

The generator bound had been raised to 600000ms — exactly the whole-chunk timeout
in scripts/run-tests.cjs:973. A step bound equal to the chunk ceiling loses the
race: the chunk is killed first and the failure arrives as an opaque "no failed
step" kill, so the per-step diagnostic added a commit earlier was built and then
made unreachable in the same change.

Separately the #2767 test declared a per-test timeout of 300000ms, BELOW the
inner bound it was meant to permit, so it could still die at the exact 300s
ceiling this was supposed to lift — via node:test's timeout rather than
spawnSync's. Its sibling declared 900000ms, above the chunk ceiling, which is the
same opaque-kill hazard from the other direction.

The three bounds only produce a useful failure if they fire inward-out, so they
now do: step 360s, per-test 480s, chunk 600s. 360s is ~3x the passing observation
(91.6s / 115.8s) and 20% above the censored 300.1s timeout, while leaving 240s of
chunk headroom for every other file sharing it. Four tests lock the ordering,
including a drift guard on the exported values — without it, editing a call
site's literal timeout would leave the ordering assertions passing while the real
ladder inverted.

Also from review:

- err.gsdBaselineStep and err.gsdBaselineTimings were written and never read
  anywhere in the tree; only the rewritten message is consumed. Removed rather
  than kept as speculative surface.
- buildOverlayRepo discarded placeVanishableLeaf's boolean at both call sites, so
  a vanished leaf left the overlay with no accounting at all. It now collects the
  skipped paths and warns once. Not thrown: a source that left the tree really is
  not part of the snapshot, and throwing would reintroduce the crash the
  tolerance removes — but silence would let a dropped leaf resurface later as an
  unrelated missing-file assertion.
- The instrumentation commit shipped no test. One now drives a real failure and
  asserts the message names the step, its elapsed time, and the bounds.

Refs #3271

* chore(#3271): backfill the changeset PR number

* fix(#3271): bound a hook fan-out as its own class, not as a bare probe

CI failure on PR #3285, job full test (windows-latest, 22, shard 2/3) — every
other lane green, including windows-latest node 24 across all three shards:

  not ok 1 - blocks push when any to-be-pushed commit matches local blocked regex
    error: bash .githooks\pre-push failed — outcome=timed_out exitCode=null stderr=
    duration_ms: 15040.2168

A bound, not a hang: the test supplies stdin via input:, so the hook is not
blocked reading its ref list, and the duration lands exactly on the 15000ms
bound.

The site used PROBE_TIMEOUT_MS, which tests/helpers/timeouts.cjs documents as "a
single short CLI query or node -e probe against a temp fixture". This is not
that. It spawns bash running .githooks/pre-push, and the hook then invokes a MOCK
git that is itself a bash script, so one runHook is roughly four Git Bash spawns.
On Windows each is Defender-scanned and the first hook test in a file pays cold
start on top. That module's own docstring warns against precisely this: a call
site that differs from its class must not be forced onto a shared value that does
not describe it.

HOOK_FANOUT_TIMEOUT_MS is that missing class — 60000ms, 4x the bound that failed
and half INSTALL_TIMEOUT_MS, which is the right order: a hook fan-out is much
lighter than a full installer run and far heavier than reading back a version
string. Two tests lock the ordering against both neighbours, including one
asserting real margin over the censored 15040ms observation, since a bound that
merely matched what was measured would be the same defect again.

Scoped deliberately: the other ~360 runHook sites keep their current bounds. This
adds the norm and applies it where a real failure demonstrated the need, rather
than sweeping a value across sites with no evidence for any of them.

Refs #3271

---------

Co-authored-by: sim <sim@local>
2026-08-09 23:41:44 -04:00
Tom Boucher
2e2b8ba4a7 enhance(#2704): resolve documentation links and compare H1 status brackets in the ADR gate (#3266)
* test(#2704): failing-first coverage for ADR link resolution and H1 status brackets

Binds the gate to two assertions it does not yet make: every relative markdown
link under docs/adr/ must resolve, and an H1 trailing status bracket must agree
with the Status: field instead of being silently stripped.

Covers all 51 rows of the phase test matrix across two altitudes - the pure
extractLinks/maskCode IR for fence and inline-code-span boundaries, hostile
input and the fast-check totality properties, and the real CLI verdict for the
end-to-end classes. Includes the DEFECT.GENERATIVE-FIX parity test that iterates
the exported STATUSES array so a sixth status is covered the day it is added.

* feat(#2704): resolve ADR documentation links and compare H1 status brackets

The ADR gate validated naming, relation symmetry and index freshness but never
resolved a link target, and it stripped an ADR's trailing H1 status bracket for
display rather than comparing it against that ADR's own Status: field. Both
classes were structurally invisible: #2691 found five dangling references by
manual audit roughly a year after they were introduced, one of which reached the
published npm payload, while CI reported green throughout.

Both are now assertions on the same --check path, using only node:fs and
node:path - no dependency and no subprocess.

Fenced blocks and inline code spans are masked before scanning, because markdown
does not render a link inside code. That is not a policy choice: the corpus
contains exactly two such sequences today and both are ordinary JavaScript.
Masking preserves length and column positions so findings still name a real line.

Resolution is case-exact on every platform - a link that resolves only through
macOS or Windows case-folding still 404s on github.com and still fails the Linux
lane - and a destination resolving outside the repository is reported before any
filesystem call is made.

Also single-sources two duplicated surfaces this change would otherwise have
extended: the H1 bracket vocabulary (a second hand-written copy of STATUSES with
nothing asserting agreement, a DEFECT.GENERATIVE-FIX instance) and the docs/adr
directory traversal. Two tests added by #2691 that reimplemented link resolution
and bracket comparison inside the test file are removed for the same reason; the
corpus assertion is now made by running the real gate against the real corpus.

* fix(#2704): reject symlinks that leave the repository and linearize code masking

Four defects from the isolated adversarial security review, plus one it noted.

BLOCKER - a symlink defeated path containment. path.relative(ROOT, abs) is
purely lexical, but the case-exact walk then calls readdirSync, which follows
symlinks at the OS level: a contributor-committed docs/adr/x -> /etc together
with a link through it passed containment and listed the real external
directory, and a wrong-case probe echoed a real external filename through the
"Did you mean" hint into publicly-readable fork-PR logs. Every segment is now
lstat'd before descent; a symlink is realpathed and re-checked against
realpath(ROOT) - realpath on both sides, so a root under /var does not produce
false escapes - and an escape emits no hint and reads nothing further.

The same rule now governs which FILES are read: an ADR entry that is a symlink
out of the repository is excluded and reported rather than parsed, closing the
vector this change had widened by newly reading README.md, naming-violation
files, and full bodies rather than only header fields.

MAJOR - inline-span masking rescanned the line remainder per backtick run,
roughly O(n^1.6) on adversarial input: 1.76s for an 800KB line. Rewritten as a
single linear pass pairing runs through forward-only per-length cursors. Same
input now takes 3.31ms, with behavior unchanged.

MINOR - an unreadable or broken entry threw, and the generic handler wrote a
raw stack trace carrying absolute CI paths to stderr. The scan is now
fault-tolerant and reports excluded entries as ordinary violations. The status
vocabulary is escaped before being interpolated into a dynamic RegExp -
defence-in-depth, not a live bug.

The containment predicate had reached three hand-written copies while fixing
this; it is now the single escapesRoot() helper used by all four call sites.

* feat(#2704): add a --json report so the gate's tests assert on typed values

Maintainer-directed addition. CONTRIBUTING.md's "Prohibited: Raw Text Matching
on Test Outputs" requires that a system under test producing text also expose a
structured intermediate representation, and that tests assert on that IR rather
than on rendered prose. This gate had no such surface, so its verdict tests
matched on stderr.

--json runs exactly the same validation as --check and writes a report to stdout
with the same exit code, following the frozen-REASON-enum pattern already used
by verify-reapply-patches.cjs. Every violation carries a stable reason code plus
the fields a consumer needs, so nothing has to pattern-match an error message.
Adding a reason stays three coordinated changes - the enum, the emitting site,
and the test locking Object.keys(REASON).sort().

The human output is unchanged, deliberately: a large pre-existing suite asserts
on it and migrating that is not this PR's concern. Verified by running the
pre-change and post-change scripts against an identical violating corpus and
diffing their stderr - character-for-character identical.

This PR's own verdict tests now assert on parsed --json. Absence checks improve
the most: "no bracket violation" is now a reason-code predicate rather than a
negative regex over prose, which could pass for the wrong reason. The security
assertions were strengthened rather than translated - no leaked filename may
appear in ANY field of the serialized report.

Unknown flags are now rejected instead of silently falling through to printing
the index.

* test(#2704): fix the status-parity fixture and guard hooks/dist before overlay builds

Two failures from the matrix run of 79b29909.

The status-parity fixture was mine. It built, per status token, an ADR whose H1
bracket and Status field both carried that token - but Superseded carries an
obligation beyond the bracket: it must name its successor as a file link and be
symmetric with it. The fixture declared a bare Superseded, tripped that
unrelated invariant, and the test reported a bracket-parity failure for a reason
that had nothing to do with bracket parity. The fixture now satisfies each
token's own obligations in both the agreeing and contradicting corpora, derived
from the status actually declared rather than special-cased on one name, so a
future token carrying obligations is handled rather than silently skipped.

The second failure was not mine but is fixed here rather than deferred.
mcp-catalog-parity.install.test.cjs hardlinks hooks/dist/* while building its
overlay, but hooks/dist is a gitignored build artifact produced only by
build:hooks. The suite had no guard, so it passed only when some other suite
happened to build it first - an execution-order dependency, which is why it
failed on node22 and passed on node24 for identical code. install.test.cjs
already documents this exact hazard and guards it.

Six behaviorally identical copies of that guard existed across three files.
Rather than add a seventh, they are now one canonical
tests/helpers/hooks-dist.cjs - idempotent and bounded by the shared
BUILD_TIMEOUT_MS class norm - which is the same single-sourcing this PR applies
to the ADR gate itself.

* docs(#2704): add a how-to for contributors the ADR gate rejects

The reference and explanation quadrants were covered by Lifecycle rules 5 and 6,
but the task-oriented one was thin: a contributor meets this gate because it
failed on their PR, under pressure, and the rules told them what is checked
without telling them what to do about it.

Adds the command to reproduce the CI failure locally and a message-to-remedy
table covering every reason code that can be hit - unresolved target, wrong case
with the did-you-mean hint, repository escape, symlinked ADR file, bracket
contradiction - plus the backtick escape hatch for illustrative links and the
caveat that indented code blocks are not skipped.

The table is itself written in backticked inline code, so the gate skips it: the
escape hatch demonstrated on the page that documents it.

* chore(#2704): backfill changeset PR number

pr:0 placeholder replaced with the real PR number now that #3266 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-09 17:08:46 -04:00
0xdhx
35df0891af fix(#3156): sandbox HOME on raw installer spawns — the one leak no scrub reaches
The strict live-config guard this PR ships went red on CI: the suite creates
$HOME/.gsd/defaults.json. Diagnosed rather than suppressed, because the guard
is right — this is #2665's class arriving through the one door the scrub set is
structurally unable to close.

bin/install.js writeNonClaudeDefaults() (#2834) writes
path.join(os.homedir(), '.gsd', 'defaults.json') for every non-Claude runtime.
os.homedir() consults NO GSD variable, so:

  - no entry in CONFIG_LOCATION_ENV_KEYS can reach it, however the set is
    derived; and
  - blanking GSD_HOME does not reach it either -- a blank GSD_HOME falls back
    to exactly that homedir().

Only a sandboxed HOME contains it, and HOME is deliberately excluded from
TEST_ENV_BASE because blanking it would break far more than it fixed. So the
containment belongs per-spawn, which is the discipline the suite already
applies by hand -- install-minimal-hooks.test.cjs:576 carries a comment
naming this exact hazard for the --codex spawn, while the parameterized
--${runtime} spawn 130 lines below it does not. Instance fixed, class open.

Rather than add a fifth hand-synced env shape to a PR whose subject is that
hand-synced copies drift, this adds ONE export -- installSpawnEnv() in
tests/helpers.cjs -- and routes every raw installer spawn through it, including
the shared tests/helpers/install-shared.cjs installerEnv(), which every install
suite already consumes. Callers passing an explicit { HOME, USERPROFILE } are
unaffected: overrides spread last.

Census (measured, not reasoned): of the 119 test files that spawn bin/install.js,
exactly four wrote into a sandboxed $HOME before this commit -- install,
copilot-install, install-minimal-hooks, opencode-plugin-adapter -- and zero do
after. Only install.test.cjs sits in CI's targeted lane, which is why ubuntu
went red on one file while the macOS full lanes went red on four.

Attribution: the leak reproduces unchanged at upstream/next itself, so the
defect is base-owned and pre-existing; only the detector is new. The guard
found a real leak on next within one run.

No new failures: the surviving names under a sandboxed HOME
(folded:enh-2380-sync-skills, getGlobalConfigDir (Copilot)) fail at base too,
and base additionally fails folded:bug-3288-model-catalog-install-path, which
this tree does not.
2026-08-08 05:56:22 -05:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
cbd180c5cd test(#3147): bound the lint/changeset/docs cluster onto the process seam (#3181)
* test(#3147): bound the lint/changeset/docs cluster onto the process seam

Migrates 69 unbounded sync spawn sites across 24 files. Allowlist 73 to 49.

Two shared helpers move: tests/helpers/graphify.cjs (6 importing suites) and
tests/fixtures/index.cjs, whose three quoted-argument shell strings became
single argv elements rather than whitespace splits.

changeset-lint's throw-native git() helper routes to gitOrThrow; migrating it
to bare runGit would have silently swallowed a failure that is loud today.
ingest-docs goes the other way -- its catch never rethrew, it degraded failure
into data every call site asserts on, so throwIfFailed would have thrown where
the original returned. The design doc said otherwise and was corrected.

tsconfig-noemit runs a real tsc --noEmit and takes a bespoke 180000ms per the
ensure-runtime-build precedent, not the 30000ms build-hooks norm -- that norm
is for a file copy, and sizing against a label rather than the work is the
same error in the opposite direction.

* test(#3147): add toLegacyResult and settle review findings

The seam exposed a throwing adapter (throwIfFailed) but no non-throwing one,
so eight files independently re-derived the same unwrap back to the legacy
{status, stdout, stderr} shape. That is the third time this epic produced N
copies of one mechanism -- seven throw wrappers in Wave 1, fifty-two timeout
constants in Wave 2, eight result adapters here. The pattern is that whenever
the seam does not expose a mechanism, every suite re-derives it.

toLegacyResult now sits beside throwIfFailed, with its own tests.

Two sites are deliberately NOT converted: changeset-cli's runRender and
runRenderIn return {status, report, stderr} from parsed JSON and never a raw
stdout, so they are a different shape family. lint-legacy-dir-name keeps its
local GUARD_TIMEOUT_MS: 30000 matches the build norm numerically but bounds a
lint probe, not hooks bundling, and importing it would encode a coincidence
as a relationship.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:45:51 -04:00
Tom Boucher
3fac6e629f test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam

Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73.

Timeouts are sized from evidence already in the tree rather than a house
default, because this wave spawns installers rather than git plumbing and an
undersized bound does not catch a hang -- it manufactures CI flake, which is
worse, since a flake gets re-run instead of investigated. install.test.cjs
records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while
another lane passed the same commit in 12.7s, so full installs are bound at
120000ms against that recorded incident.

Also adds an auditable escape to the guard's timeout ceiling. The 600000ms
cap was set in #3143 from partial evidence, but fragment-single-edit-
propagation carries a documented, load-tested 900000ms bound on a run that
chains a full build plus eight generators -- the guard would have rejected a
correct timeout the moment that file left the allowlist. A value above the
ceiling is now permitted only with an inline allow-spawn-timeout-ceiling
marker carrying a non-empty reason. It raises the ceiling; it never waives
the requirement for a bound, which is asserted directly.

install-shared.cjs keeps its hand-rolled assert rather than routing through
throwIfFailed: its message embeds both streams, and throwIfFailed carries
only a trimmed stderr. The message now also names the outcome, so a bounded
timeout reads as such across its 38 importers instead of as
expected null to equal 0.

* test(#3145): extract class-norm timeouts and correct the build-hooks sizing

A pre-PR review found 52 copies of four class-norm timeout constants across
this wave. These are not per-suite fixture bindings -- they are shared facts
about how long a class of subprocess takes, derived from a recorded bench
incident. That norm already moved once (60000 to 120000 after a real
ETIMEDOUT), and 52 copies would have drifted the next time it moved.

Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and
converts the copies. A site that genuinely differs -- a real tsc compile, or
regen:derived -- keeps its own local constant with its own justification.

Also corrects a misclassification: scripts/build-hooks.js was sized as a
build at 120000 in twelve places and 60000 in another, but it compiles and
bundles nothing. Its own header says no bundling needed; it copies pre-built
files and syntax-checks them with vm. Three different values bounded one
script; now there is one.

* test(#3145): fix red CI — lint self-match and a Windows chunk overrun

Two failures on PR 3176.

lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The
fixture exists to prove an unrelated marker does NOT suppress the rule, so it
carries that marker's literal text as test data. Split via concatenation, the
same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own
self-match problem. The explanatory comment needed the same treatment.

The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped
seven minutes before the kill, so this was an overrun rather than a slow
chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs
regen:derived bounded at 900000ms, which is larger than the whole chunk
budget, so the chunk killer always fires first and it can never complete
there. Both the test and that bound predate this change; modifying the file
pulled it into the Windows targeted set and exposed it. Skipped on Windows
with the reason recorded; the Linux lanes cover it. The 900000 bound and its
ceiling marker are unchanged -- they are correct.

* test(#3145): refresh the stale test-timings cost table

The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs
packs chunks by measured duration from tests/test-timings.json, and an
unknown file falls back to the table's median weight -- advisory by design,
but it silently underweights exactly the files that matter.

Four of the failing chunk's 22 files were absent from the table, including
the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s
(it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s.
Both were weighted as average, so the chunk's total weight read 53.68 against
a budget of 60 and the packer produced a single chunk.

Regenerated from a passing full-suite run, per the remedy the script itself
documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since
gen-test-timings.cjs replaces the table wholesale rather than merging.

Proven against the real packer: the same 22 files now weigh 103.91 and split
into two chunks. No logic, budget, or timeout was changed; raising a budget
to make a red gate pass is not a fix.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:18:18 -04:00
Tom Boucher
27aa40f65e fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir

pi reserves <configDir>/hooks as its deprecated extension location and warns
on every startup when it exists. Assert a pi install stages the shared hook
bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/.

Also adds pi to the local-scope dir table in install-shared.cjs: pi was in
RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved
path.join(root, undefined) and no local pi install could be exercised.

Fails before the fix. Verified via the remote runner.

* fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir

pi reserves <configDir>/hooks as its now-deprecated extension location and
warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs()
guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged
its shared hook bundle exactly there, and pi's advised remediation (move it to
extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's
extension auto-discovery.

The bundle directory name is now runtime-descriptor-driven: hostBehaviors
.sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are
byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path
segment — separators, dot-only segments, trailing dots, absolute paths, NUL,
and Windows reserved device names all fall back to the default, because the
value is joined onto a user's config root and written to.

Renamed in place rather than relocated: hook scripts resolve siblings via
__dirname/.., so a depth change would silently break them.

- install / uninstall / manifest sites all read the resolved name
- pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded
  trees still resolve; the never-throws contract is preserved
- new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new
  non-recursive remove-empty-dir engine primitive (rmdirSync only,
  symlink-refusing, containment-guarded); ADR-0008 amended accordingly
- fixes two latent name-dependencies the rename exposed: the stale-hook scan
  and the injection scanner's self-exclusion both hardcoded 'hooks'

Verified on the remote runner.

Closes #3023

* fix(#3023): close review findings and align emitted provenance with the rename

Adversarial review found two defects, and the remote runner found four
failure clusters. All fixed here.

Review BLOCKER — detect-custom-files was blind to the renamed bundle.
GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the
whole gsd-hooks/ tree was invisible to the custom-file scan and user-added
files there were never backed up before the next update's clean-install wipe.
The dir set now resolves via the .gsd-runtime marker plus the shipped
capability registry (never bin/install.js, which is not shipped into installed
trees), and falls back to scanning every known candidate when the runtime
cannot be determined — over-scanning is safe, under-scanning is the data loss.

Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir
accepted any directory, so an interrupted install left gsd-hooks/ winning over
a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now
qualifies only if it is non-empty.

Remote-runner clusters:
- emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped
  rules pointing at the same sources the existing hooks/ rules use. The table is
  total, so an unattributed family is a hard failure by design.
- pi tests in install-minimal-hooks and the install integration suite asserted
  the old layout; updated to derive the dir name from the descriptor rather than
  hardcoding either name.
- 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after()
  runs in registration order, cleanup was registered before mock.restoreAll(),
  and node22's JS rimraf calls the public fs.rmdirSync while node24's native
  path does not — so the EACCES stub leaked process-wide on one lane. Restore
  now runs first.

Verified on the remote runner.

* fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde

pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent
(packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty
configHome.env, so a user with that variable set had GSD installed where pi
never looks. Added the env name; the dot-home-nested resolver already handled
the override, so no resolver logic changed.

Also fixes expandTilde in the shared runtime-homes resolver, found while adding
that: it hardcoded os.homedir() and ignored the opts.home every caller threads,
so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi)
silently resolved against the real home. That is a correctness bug and a
test-escape hazard — a sandboxed test asserting on a tilde override reached the
developer's actual home directory. Now threaded through every branch; behavior
with no injected home is unchanged.

Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location
moved with the rename. The provenance rules satisfy the totality gate; the
differential gate needs the ack because the hook sources are byte-unchanged —
only the installer's target directory moved. The two hook files this branch
genuinely edits stay attributed and are not double-acked.

Note on piConfig.configDir: it is read from pi's OWN installed package.json
(getPackageDir walks up from pi's __dirname), alongside piConfig.name — a
white-label setting for a redistributed pi fork, not a per-project user setting.
Documented accordingly rather than treated as an unsupported override.

Verified on the remote runner.

* fix(#3023): reject blank env overrides, pin adapter/descriptor parity

Three review findings, all fixed.

A whitespace-only config-dir override was accepted verbatim: the guard was
`if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR='   '
resolved to a literal three-space directory name instead of falling back to the
descriptor default. Fixed across every env-consuming branch — dot-home,
dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's.
Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working.

pi/gsd.cjs's probe list and the descriptor were two independent sources of truth
for the bundle directory name; a future rename would have desynced them silently
and left every pi hook quiet with no error. The probe list stays deliberate — it
must resolve in a dev checkout and a half-upgraded tree, where the registry's
answer would be wrong — so this adds the parity assertion the repo's
generative-fix-divergence rule calls for: the descriptor value must be the FIRST
candidate, and the default must remain present.

Changeset body rewritten to cover the two later user-facing fixes it had not
caught up with.

Verified on the remote runner.

* chore(#3023): backfill changeset PR number

* fix(#3023): anchor injection-scan patterns and fix a macOS detection hole

CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the
same fact as a genuinely empty or absent one'. The match was the 'act as a'
INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending
in act tripped it (fact, impact, contract, artifact, interact, redact,
abstract). My four-line CONTEXT.md edit dragged the latent false positive into
this PR because the scan is diff-scoped by file but reads whole files. Anchored
with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would
have left the class alive for the next PR touching any file saying 'fact as a'.

Auditing the rest of the list for the same class surfaced a real detection hole:
the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex
escape. BSD/macOS grep reads it as four literal characters, so single-quoted
eval('...')/exec('...') payloads were NEVER detected there while passing on
GNU-grep CI. Replaced with a literal apostrophe class.

Boundaries were added only where a real word-suffix collision exists; exec,
jailbreak, developer mode and the role-manipulation family were audited and
deliberately left unanchored. 22 new cases cover both directions — the false
positives now scan clean, and every real payload still fires, including the
quote/punctuation/start-of-line boundary forms.

Also builds this branch's injection test fixture at runtime instead of carrying
the literal phrase, so the payload keeps its teeth without tripping the scan.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-07 13:41:21 -04:00
Tom Boucher
1d208e5af6 test(#3144): bound the git/worktree cluster onto the process seam (#3152)
* test(#3144): bound the git/worktree cluster onto the process seam

Migrates 180 unbounded sync spawn sites across 19 files. Every previously
unbounded call now carries an explicit timeout with a comment giving the
number and why.

The migration is not a callee swap. execSync and execFileSync throw on a
non-zero exit and the seam never does, so each site was classified first:
sites that rely on the throw route to gitOrThrow, and sites that already read
.status to detect an EXPECTED non-zero -- an intended cherry-pick conflict, a
rev-parse outside a repo driving a skip -- route to the never-throwing runGit
instead, which would otherwise throw on exactly the exit being probed for.

Two same-named git() helpers in worktree-cleanup.test.cjs have different
return contracts, one trimmed and one raw; both are preserved rather than
unified.

Collapses five hand-rolled throw wrappers onto one throwIfFailed in
git-fixture.cjs, which gitOrThrow now also uses so the shape cannot drift.

Allowlist drops 139 to 120; BASELINE lowered to match.

* test(#3144): fix pre-PR review findings

Documents throwIfFailed in the CONTEXT.md glossary and CONTRIBUTING.md --
it became the shared throw mechanism without either doc naming it.

Routes the sixth and seventh hand-rolled copies of the throw shape through
throwIfFailed (worktree-baseref-install, worktree-safety-reap); the first
consolidation missed both.

Converts ci-rebase-check's 8 fixture-setup calls from unchecked runGit to
gitOrThrow so a failed setup step aborts where it fails rather than
surfacing later as a confusing failure against the wrong subject.

Adds 12 direct unit tests for throwIfFailed, which until now was only
exercised transitively.

Splits verify.test.cjs's non-git grep/sed bound off GIT_TIMEOUT_MS.

---------

Co-authored-by: sim <sim@local>
2026-08-07 10:58:34 -04:00
Tom Boucher
2afe17bbdb test(#3143): add the no-unbounded-spawn guard and throw-preserving git fixture (#3150)
* test(#3143): add no-unbounded-spawn guard and throw-preserving git fixture

Adds the ESLint rule local/no-unbounded-spawn, wired into the tests/**/*.cjs
block, plus an allowlist that only ratchets down: a listed file with zero
violations reports its own entry as stale.

The rule resolves renamed destructures and chained requires rather than
matching literal callee names -- both forms exist in the suite today and a
name-only matcher leaves them permanently invisible. It resolves an options
object held in a single-write const, which is what keeps process-seam.cjs,
the bounded reference implementation, from flagging itself.

timeout: 0 and anything above the 600000ms ceiling are rejected as only
nominally bounded.

Adds tests/helpers/git-fixture.cjs so a migrated execSync call site keeps
its throw-on-non-zero contract; process-seam.cjs is unchanged.

* test(#3143): prove the allowlist guards can actually fail

Extracts the D4/D6/D7/D8 checks into pure helpers and drives each against a
synthetic fixture carrying an injected violation. Without this the suite only
proved that today's clean data passes, which a deleted check would also
satisfy.

* fix(#3143): close two ceiling and alias escapes found in review

Nested arithmetic bypassed the ceiling entirely: the numeric evaluator only
resolved a flat literal, so `timeout: 60 * 60 * 1000` (3600000ms, six times
the ceiling) fell through to trusted and reported nothing. The evaluator now
recurses through arithmetic and unary signs with a depth cap.

Alias resolution was traversal-order dependent, not scope dependent: a call
textually above its own require destructure saw an empty alias map and
reported clean. The map is now built in a Program pre-pass.

Also: an explicit timeoutMs:undefined no longer overwrites the git fixture
default via spread, adds the missing seam-routed rule test, and de-duplicates
the repeated try/catch in the fixture tests.

---------

Co-authored-by: sim <sim@local>
2026-08-07 09:43:36 -04:00
Tom Boucher
7e1004d89e test(#3056): add the in-process fault-injection adapter, and normalize execGit's shape (#3077)
* refactor(#3071): normalize execGit's call and result shape, unify ExecGitFn

ExecGitFn was declared four times. Three were hand-copies of one signature and
two of those were wrong: they typed exitCode as number|null when _spawnResult
returns `result.status ?? 1` and can never yield null, weakened signal from
NodeJS.Signals to string, and widened error from Error to unknown. Only
verification.cts got it right, via `typeof execGit`.

The root cause was a missing export: SpawnResultOutput was declared without
`export`, so no other module could name the return type of execGit. Three
authors independently hand-copied it instead. Exported now.

Normalizing the type alone would have left the pressure that caused the
divergence, so the function is normalized on both sides. It now ACCEPTS every
call its consumers make — worktree-safety's declaration could not express an
env-carrying call at all — and RETURNS every result code they need: timedOut
moves into _spawnResult, so execGit, execNpm and execTool all carry it and the
one extension that justified a separate type disappears. All four sites are
now `typeof execGit` with nothing left to restate.

timedOut reuses the existing isSpawnTimeout predicate introduced by #3050
rather than re-deriving it. That predicate checks error.code === 'ETIMEDOUT'
only; the signal === 'SIGTERM' conjunct was deliberately dropped there because
Windows does not reliably report SIGTERM and requiring it risks a false
negative. There is no false-positive risk, and a test proves it: an
externally-delivered SIGTERM leaves error null, so it is still not reported as
a timeout.

No dead null-checks surfaced. Every exitCode comparison in the two affected
modules is === 0, !== 0 or === 128 — never a null guard — so the nullable
declaration had never been written against.

Closes #3071

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3056): add the in-process fault-injection adapter

Adds tests/helpers/faulty-deps.cjs — makeFaultyGit() and withFaultyFs() — so a
module's error branch can be driven deterministically and its degraded verdict
asserted, instead of a counter-test that only proves the call did not throw.

makeFaultyGit returns a value structurally assignable to `typeof execGit`, so
one stub satisfies all four seams that the #3071 normalization collapsed into
that single shape. A parity test drives the same stub through a real injectable
entry point in each of worktree-safety, git-base-branch, worktree-base-ref and
verification; it fails the moment any of them re-grows its own shape.

Faults are scoped rather than global — by argv predicate and by call ordinal —
because a fault adapter that faults everything looks like it works and proves
nothing, and because verification.cts's two-call error handling needs to fault
the second call only. Invocations are recorded so a test can assert an exact
call count.

The timeout fault carries error.code === 'ETIMEDOUT', and a test asserts the
real isSpawnTimeout predicate matches it, so the fixture cannot drift from the
production definition of a timeout. A companion test asserts an externally
delivered SIGTERM with a null error is still NOT reported as a timeout — the
false-positive guard for #3050's dropped conjunct.

withFaultyFs restores in a finally so a throwing body still restores, patches
only the named methods, and nests without clobbering an outer saved original.
It never uses chmod: that no-ops under root, so the test would pass with zero
coverage in root Docker/CI.

The adapter is in-process via deps only. The Phase 1 process seam is documented
as deliberately not a fault-injection surface — it cannot distinguish an
injected timeout from a genuine bench OOM and would retry it.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3071): add changeset fragment for the execGit normalization

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3071): route the last two timeout checks through the shared predicate

An isolated review found this branch had normalised the timeout verdict but
left two callers still hand-rolling the fragile version of it.

check-command-router's runBoundedShell computed `timedOut: r.signal ===
'SIGTERM'` while the correctly-derived `r.timedOut` sat on the same result
object. graphify's execGraphify branched on `result.signal === 'SIGTERM'`, with
a comment asserting the very premise isSpawnTimeout exists to reject.

Both fail in both directions. On Windows a genuine timeout is not reliably
reported as SIGTERM, so the guard silently fails to fire — the false negative
#3050 was raised for. And an externally-delivered SIGTERM is not a timeout at
all, so the check also fires when it should not; isSpawnTimeout avoids that
because `error` is null in that case and it keys on error.code.

Both now read the derived `timedOut`, and graphify's comment states the actual
rule instead of the fragile assumption.

Also replaces a vacuous test: "execGitDefault now accepts env" never called
execGitDefault (it is unexported), called execGit — whose signature already
accepted env before this branch — and asserted only that exitCode was a
number, which would pass whether or not the change under test existed. It now
proves env reaches the child by asserting `git var GIT_EDITOR` returns the
injected sentinel.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3071): make the graphify timeout fixture faithful to a real timeout

The remote matrix failed on both Linux lanes: graphify's "returns exitCode 124
on timeout" got 1 instead of 124. The fixture was wrong, not the production
change.

It stubbed spawnSync as { status: null, signal: 'SIGTERM', error: undefined }.
That is not a timeout. A real spawnSync timeout also sets error.code
'ETIMEDOUT'; a SIGTERM with no error is an externally delivered signal — a
kill. The old `result.signal === 'SIGTERM'` check accepted it as a timeout,
which is the false positive the shared predicate exists to reject, so this test
was locking that bug in rather than guarding against it.

The fixture now carries a real ETIMEDOUT error and all three original
assertions pass unchanged. A counter-test is added alongside it: an externally
delivered SIGTERM with no error must NOT be reported as a timeout. That is the
assertion whose absence let the false positive live.

Swept every SIGTERM/SIGKILL stub under tests/ for the same unfaithful shape.
No other instance: the worktree-safety, worktree-base-ref and commit-staging
fixtures already set ETIMEDOUT, and the remaining hits are either deliberate
external-kill tests or feed code that never consults timedOut.

Two sites keep their own signal check deliberately and are NOT changed:
capability-source.cts:1301,1386 fail closed on ANY abnormal termination, which
is correct — reading timedOut there would stop it failing closed on a kill and
let it parse stdout from a killed process. Their reason strings, and
check-latest-version.cjs:115, label any signal as "timed out", which is
imprecise wording over a correct verdict, not a silent failure.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3071): backfill changeset pr number to 3077

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 02:14:41 -04:00
Tom Boucher
5fd5c81042 test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it

Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync,
each returning a typed discriminated union
{ outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }.
Every call is timeout-bounded; there is no unbounded path.

runGsdTools becomes an adapter over the seam. Its legacy
{ success, output, error, exitCode } shape and retry-once-on-kill behaviour
are preserved byte-identically, so none of its 136 caller files change.

Outcome discrimination was corrected against probed runtime behaviour rather
than assumption: a timeout and a maxBuffer overflow are identical on both
status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs
ENOBUFS). Overflow is therefore classified before timeout. This fixes a live
defect — the previous isKilled() treated an overflow as a kill, retried it for
a second full 60s run, and then reported "host OOM or scheduler contention"
for a child that had merely printed too much.

Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs,
which brought 31 previously unlinted shared helpers under the same rules their
sibling test files already obey, and fixes the 5 violations that surfaced —
including a bare npm invocation without shell:true in
tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now
routed through the existing portable runNpm helper.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): migrate every local spawn wrapper onto the process seam

Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions
with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name,
parameter list, return shape and post-processing (JSON parse, ANSI strip, env
sanitising, field extraction) — only the spawn mechanism changes, so no test
assertion moves.

The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather
than a fourth primitive; it is explicit rather than inferred from the file
extension, because guessing an interpreter from a path fails silently when a
script's name does not match its shebang.

Seven wrappers were previously unbounded and now carry an explicit timeout
sized to what each actually runs, not the seam default. Two of those seven
(gsd-write-guard, lint-docs-command-form) were absent from the issue's
inventory entirely and were found by scanning after the migration.

Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md
reference section covering the three primitives, the discriminated union, and
the two rules the seam enforces.

Scope disclosure recorded in the phase design notes: the issue scoped three
identifier names. A scan for local helpers that spawn AND return the spawn
result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded
direct git call sites. This change bounds 25 of those. The remaining surface
is the same defect class and is NOT closed by this PR.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify an externally-killed child as KILLED, not EXITED

Blocker found in this branch's own diff, independently confirmed by an
isolated reviewer.

A child killed by an external signal — a genuine bench OOM kill — makes
spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field.
The seam's "no error implies EXITED" rule therefore classified it as a clean
exit, and runGsdTools returned { success: false, exitCode: 1 } without
retrying. That silently defeated the #969 kill-discrimination for precisely
the case it was built for: the old isKilled() fired on `signal != null`,
retried once, then threw a labelled resource-starvation error. A real OOM
would have been reported as an ordinary assertion failure.

Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes
the adapter retry on TIMED_OUT or KILLED — reproducing the old
`killed || signal != null || code === 'ETIMEDOUT'` condition exactly.
SPAWN_FAILED still does not retry (matching the old behaviour, where signal
was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate
divergence: the old code retried it because signal was SIGTERM, burning a
second 60s run on a child that had merely printed too much.

All five outcomes verified against the live runtime rather than assumed:
SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT),
>1MB stdout -> buffer_overflow (ENOBUFS).

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): address standards-review findings on this branch

Three findings from the standards axis of the review, all in this branch's
own diff.

The CONTEXT.md glossary entry this branch introduced was already stale on the
branch's own last commit: it enumerated a 4-member OUTCOME while the code had
5, because the KILLED fix did not update it. That is precisely the drift the
"module changes update Domain-terms" gate exists to catch, so the entry now
lists all five and explains KILLED.

api-coverage-gate-e2e compared an outcome against the raw string 'exited'
rather than OUTCOME.EXITED, the only such outlier; the enum is now imported
and used. A sweep for the other four outcome literals found no further
comparison sites.

Three call sites hand the literal bash flag '-c' to the seam's first
parameter, which the JSDoc described as an absolute script path. Rather than
add a fourth primitive, the contract is corrected to match reality: the
parameter is renamed `target` and documented as the first argv element handed
to the interpreter — normally a script path, but for an interpreter invoked
with an inline program it may be that interpreter's own flag. No behaviour
change.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): assert the cross-platform timeout contract, not the macOS one

The remote runner failed on both Linux lanes (node 22 and node 24, identical)
while the same tests passed locally on macOS. Two assertions encoded a
platform-specific behaviour as a cross-platform guarantee.

When spawnSync times out, macOS preserves the child's partial stdout/stderr;
Linux discards it and returns empty strings. Verified on node v26.5.1 both
ways. The seam passes through whatever spawnSync hands it and cannot
manufacture output that was discarded, so the production code was correct —
the tests were wrong.

Both tests now assert the guarantee the seam actually makes on every
platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being
strings rather than undefined or a Buffer. The partial-content assertions are
retained behind an explicit process.platform === 'darwin' guard so the macOS
coverage is not lost, and the first test is renamed to say what it now
guarantees.

This is the failure mode the remote matrix exists to catch: local macOS
verification would have shipped it.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout

Windows CI caught two defects the Linux matrix could not.

tests/context-predicates-query.test.cjs passes a 32K-char argv value. On
Windows that exceeds the argv limit and spawnSync fails with code
ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise,
status === null implies TIMED_OUT" — swallowed it, so the adapter retried a
spawn that can never succeed and then threw the resource-starvation error. The
old isKilled() returned false for that shape and returned an ordinary failure
result.

TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set.
Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG,
E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform
whose timeout errno differs classified correctly, so the greedy catch-all is no
longer needed.

The second defect is a contract regression I introduced and had claimed
otherwise. That same test asserts `typeof r.exitCode === 'number'`, and
toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the
assertion failed on type. The old code returned `err.status ?? 1` on every
non-retried failure path. The adapter now returns 1 again for both, and the
comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is
retracted: the seam keeps the richer truth (exitCode null plus a distinct
outcome), the legacy adapter keeps the old numeric contract its callers
actually depend on.

Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT ->
SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL ->
KILLED; clean exit -> EXITED.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 01:06:39 -04:00
Tom Boucher
8f75e27554 fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag

Every isolation gate already resolved correctly. The resolved value then reached
the executor through a prose instruction telling the model to substitute it into
a call the model composes itself, and nothing verified the substitution. When it
was dropped, the executor edited and committed in the user's primary checkout
with no consent and no warning.

A prose backstop would be the same class of artifact as the defect, so this is a
shipped PreToolUse hook on the Agent tool. It fires at the instant of the call
rather than being read once at the top of a workflow, which is the only placement
the model cannot skip.

The guard is inert unless it can positively establish that this is a GSD project,
that the project resolves to harness isolation, and that the dispatch targets an
executor. A non-GSD repo has no invariant to enforce. Where it cannot read the
configuration at all, it denies rather than assuming, with its own reason -- a
guard that cannot verify must not answer safe. A malformed payload allows rather
than throwing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3045): extend the isolation guard to Cursor

Cursor is the second of only two runtimes that resolve harness isolation, so
shipping the guard for Claude alone left half the exposed surface unguarded while
the changeset implied it was covered.

The two runtimes fail differently. On Claude the harness flag is a per-dispatch
kwarg the model must copy into a call it composes, and the defect is that it can
be dropped. On Cursor the flag is --worktree, which applies to the whole session,
and the subagent-start payload carries no isolation field at all. There is no
flag to check, so the guard verifies the effective state instead: whether the
workspace is genuinely running outside the user's primary checkout. That is a
stronger check than the Claude one because it tests reality rather than intent,
and it is commented so nobody later rewrites it into a flag check.

Isolation is established two ways, either sufficient: the workspace resolves to a
linked git worktree, or it sits under the worktree root Cursor manages. The
second matters because a directory Cursor placed there is a legitimate isolated
session even before it becomes a distinct git worktree, where linkage alone would
report no repository.

Detecting linkage required a new primitive rather than the existing context
resolver. That resolver short-circuits on finding a local .planning directory
before it ever compares the git directory to the common one -- and an isolation
worktree normally has its own checked-out .planning. Reusing it would have read a
correctly isolated session as unisolated and denied it, which is the failure
direction that gets a guard switched off. The comparison is now its own
shortcut-free function that the resolver delegates to after its own shortcut, so
existing behavior is unchanged, and the case that would have broken is pinned.

The subagent type is checked before any configuration is read, so an unreadable
config cannot deny a dispatch this guard would never have enforced against.

The input-schema comment on the Cursor hook documented only the fields common to
every event and omitted the ones specific to this one. That omission cost a
halt during this work; it now documents both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): enforce the resolved dispatch decision, not the host capability

The guard keyed on the registry's dispatch.isolation, which says only that a
runtime is CAPABLE of harness worktrees. The decision that actually governs a
dispatch is the one the workflow resolves after gating, and that legitimately
comes out as sequential in three documented cases: a project setting
use_worktrees false, a per-plan submodule intersection, and the base-check
auto-degrade. The workflow tells the model to omit the flag in exactly those
cases, and the guard was denying every one of them.

The third case matters most. The preceding fix made the base-check degrade on
git timeouts and a missing git binary, where it had previously answered "safe".
That correction is right, and it means a transient hang now degrades to
sequential far more often than before -- so the two changes composed into a trap
where the workflow behaved exactly as designed and the guard blocked it.

The workflow already resolves isolation in shell, deterministically, which is
what makes it a trustworthy source in a way the model-authored call is not. It
now records that resolved value through a dedicated verb, and both guards read
it first. A fresh record is authoritative, so sequential dispatches pass
untouched. Absent or stale, the guards fall back to the capability check
combined with the project's use_worktrees setting, which still covers the case
that never reaches the workflow.

Also widened the matcher to accept Task alongside Agent, since a host that names
the tool Task would otherwise leave the guard silently inert while implying
coverage; stopped assuming Claude when no runtime is declared, which is the
shipped default and would have demanded a Claude-only argument elsewhere; and
made a non-git project inert rather than denied, since advising a worktree
session is not actionable without a repository.

The original diagnosis never modeled sequential mode as legitimate. That
omission is what let this through, and it is now recorded there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): record at resolution and bind the record to its dispatch

Two independent reviews converged on the same failure: the guard was fail-open in
a default install, so it did not catch the defect it exists to catch. A shipped
project carries no runtime key, which made "runtime not confidently known" the
common case rather than a corner one. A record asserting that isolation was
required but carrying no flag then fell through to a capability lookup that
answered "none", and the dispatch was allowed. The flag itself only arrived from
a second shell block -- the same block a model dropping the argument would also
skip. A test had pinned that behavior as intended.

The record is now written by the resolver, as an unavoidable consequence of
asking for the value, rather than by a step the model is told in prose to go and
run. A guard against a prose-carried value cannot itself depend on prose. Mode,
flag and identifiers are written together and atomically, so the flagless window
is gone, and a record asserting isolation with no resolvable flag now denies
instead of degrading. Runtime is also resolved from the installer's own recorded
default, which makes confident resolution the normal case.

The per-plan submodule gate degrades after the phase-level decision and never
re-recorded, so a plan that legitimately ran sequentially was denied against a
still-fresh phase record. It now records its own, scoped to the plan.

A record also authorized any dispatch for four hours. One phase degrading to
sequential could silently license an unisolated dispatch in the next. Records
now carry phase and plan, the guards require them to match, and the window is
minutes rather than hours -- the resolver rewrites it before every dispatch, so
a long window bought nothing and only widened the hole.

The flag validator rejected any value beginning with two dashes, which is exactly
the form Cursor and Windsurf declare, so their real value could never have been
stored. Writer and reader also derived the record path differently and diverged
inside a linked worktree without local planning state.

The predictable path remains a way to silence the control without leaving a trace
in the diff. It grants no access an agent with shell does not already have, so it
is documented as accepted rather than redesigned around.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): correct the staleness boundary and unmask a vacuous parity test

The remote runner returned twenty failures. One was a real production defect the
boundary case existed to catch: a record whose age exactly equalled the staleness
window was treated as fresh, so it stayed authoritative for one tick past its own
expiry. Freshness is now strictly inside the window.

The parity test meant to stop the two guards' executor lists from drifting could
never have failed. Its project fixture was a bare directory rather than a
repository, so the non-git inert branch answered before the executor list was
ever consulted. It asserted agreement it never actually measured. The fixture is
now a real repository, like every sibling in the file.

A test also asserted that Windsurf declares the worktree flag. It does not --
Windsurf resolves to no isolation by design, having no named concurrent dispatch
to isolate. The test claimed a registry fact that was never true, and a comment
in the resolver repeated it. Both corrected, and the test now proves what it
should have all along: that the parser accepts any bare flag value, rather than
one runtime's supposed value.

The new guard was missing from the bundled-hook whitelist, which is the surface
that decides what actually ships, and the per-plan gate had gained calls to the
launcher without the preamble those calls require. The changeset carried
parenthetical product descriptions the purity rule forbids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3045): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3045): make the guard tests hold on Windows

Two tests redirect HOME to control where the installer-persisted runtime default
is read from. Node resolves the home directory from USERPROFILE on Windows and
never consults HOME, so both silently read the real runner profile, found no
recorded runtime, and asserted against a project the hook had not recognised. The
production code was already correct in asking the platform rather than the
variable; only the tests were wrong to assume one variable answers everywhere.
The helpers now mirror the override onto both.

The symlink spoofing test also created a directory symlink unconditionally, which
needs elevated privileges on Windows. It survived on this runner, but it would
fail on any host without them, so the creation is now attempted and the test
skips explicitly when it cannot be done -- a bare return would have counted as a
pass and hidden the gap.

Skipping alone would have left the platform uncovered, so the behaviour it proves
is now also driven in-process through an injected realpath, following the seam
already used for the clock. That case no longer depends on privileges at all, and
the end-to-end test keeps its original assertions wherever symlinks work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 23:42:16 -04:00
Tom Boucher
b780cd2dc6 test(#2933): prove one fragment edit reaches every emitted artifact (#3046)
Epic #1671 Phase 6 "Done when" required a maintainer-reachable proof that a
single-fragment edit propagates to every emitted per-runtime artifact with no
second source surface needing an edit. No test referenced that surface at all.

Adds tests/fragment-single-edit-propagation.install.test.cjs (20 rows): a
hard-linked overlay repo overrides exactly ONE steps/ fragment, real installers
are spawned per runtime, and the emitted artifacts are asserted directly.
Expected runtime sets derive from RUNTIME_META at run time, never a hardcoded
count, so a new runtime cannot be silently under-covered.

Six negative controls keep it from being pass-always theater. An
identity-stubbed composer must make marker bytes LEAK, proving the
marker-absence assertion can fail. Each derived generator whose --check is used
as evidence has its own red-path control driven by an override-only edit, each
asserting the generator's own typed reason enum rather than matching prose.

Coverage is disclosed, not implied. REGEN_STEPS_WITHOUT_CHECK_MODE names the
regen:derived steps with no read-only mode; CONTENT_EDIT_INSENSITIVE_CHECKS
names gen-inventory-manifest, whose --check derives from directory listings and
is structurally blind to content edits. Both constants are pinned by a test so
the disclosure cannot silently rot.

Assertions check sentinel PRESENCE, not whole-file byte identity: partial-wave.md
embeds the runtime-launcher snippet, so emitted fragments are legitimately
rewritten per runtime (windsurf -> .windsurf, qwen -> .qwen, claude -> its
absolute config dir). A dedicated row now locks that behavior in.

The overlay tree-diff is labelled a harness self-check, not the no-cascade
proof it cannot be: the overlay is built from the checkout with the override
map applied, so that diff can only ever restate the test's own fixture.

Extracts buildOverlayRepo into tests/helpers/overlay-repo.cjs so both install
suites share one implementation instead of diverging copies, converts that
sibling's six try/finally test bodies to t.after() per CONTRIBUTING.md, and
frees each per-runtime temp install eagerly so peak disk stays bounded.

Refs #2933

Co-authored-by: sim <sim@local>
2026-08-04 13:37:03 -04:00
Tom Boucher
c6ce4d1d9a fix(#2755): resolve the kimi hooks-TOML root per runtime (#3032)
* test(#2755): failing-first coverage for per-runtime kimi hooks root

Install/uninstall filesystem-shape rows over a sandbox HOME (no permission
tricks) plus resolver unit rows. Covers both uninstall directions, which is
where a fix applied only to the install call site would drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2755): resolve the kimi hooks-TOML root per runtime

resolveKimiHooksTomlDir took no runtime argument and hardcoded ~/.kimi, but
both kimi and kimi-code route through the single hooksSurface=kimi-hooks-toml
branch. A --kimi-code install therefore wrote its [[hooks]] block, hook bundle
and CommonJS marker into Kimi CLI's config file, and a --kimi-code uninstall
stripped Kimi CLI's block.

Adds a runtime selector to the resolver -- kimi keeps ~/.kimi + KIMI_SHARE_DIR,
kimi-code gets ~/.kimi-code + KIMI_CODE_HOME, per Kimi Code's own upstream
data-locations and hooks docs -- and passes the runtime at both the install and
uninstall call sites. An omitted or unrecognized runtime still resolves ~/.kimi,
so the exported no-arg contract is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2755): use centralized helpers and add a divergence guard

Review findings, all fixed in-PR:

- The new test block reimplemented runMinimalInstall, createTempDir and
  toPosixPath. Extends runMinimalInstall with optional root/extraEnv instead
  (back-compat: every existing caller passes neither) and uses the centralized
  helpers, per CONTRIBUTING's Use Centralized Test Helpers rule.

- Adds a parity assertion between the capability registry and the resolver: a
  third runtime declaring hooksSurface kimi-hooks-toml would silently inherit
  ~/.kimi, re-creating this very defect. The guard fires the moment those two
  surfaces drift.

- Adds an installer-level test proving KIMI_SHARE_DIR and KIMI_CODE_HOME do not
  interfere when both are set, which only the resolver unit covered before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2755): track the kimi-code hooks root in the emitted-artifact gates

The remote runner caught a real ripple: moving kimi-code hooks to ~/.kimi-code
made 31 emitted paths unattributable and 58 emitted hashes unexplained, because
three parallel surfaces keyed on the literal .kimi path.

- HOOK_CONFIG_RELATIVE_PATHS excluded only .kimi/config.toml, so kimi-code's
  config.toml became manifest-visible; it embeds a platform-varying node-runner
  command and must stay out for both products.
- HOOKS_ROOTS, the package.json-marker branch and the synthesized-install-metadata
  pattern each named .kimi only.
- tests/fixtures/install-tree/kimi-code.json still recorded the old paths;
  regenerated via gen:install-tree.

Adds the per-PR drift acknowledgment for the 58 paths whose bytes are unchanged
but whose destination moved - a ripple no source diff can show, since no hook
script was edited.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2755): clear production-tree security advisories

The remote runner's npm-integrity gate reported 2 high advisories in the
production dependency tree. My diff touches neither package.json nor
package-lock.json, so these come from the base -- but a red gate is not
something to wave off as pre-existing, so it is fixed here rather than deferred.

Lockfile-only, semver-in-range, via npm audit fix:
  fast-uri   3.1.4  -> 3.1.5   (host confusion via backslash authority introducer)
  ip-address 10.2.0 -> 10.4.0  (three SSRF / trust-boundary bypasses)
  hono       4.12.31 -> 4.13.0 (moderate; reverting it traded a high for a
                                moderate, so the full remedy is taken)

npm audit now reports 0 vulnerabilities at every severity, npm ci installs
clean from the updated lockfile, and the build and the kimi behavior both
re-verified afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2755): backfill changeset pr numbers

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 18:56:42 -04:00
0xdhx
cc3ee301a7 fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root

installSharedHooksBundle wrote `{"type":"commonjs"}` over
<configRoot>/package.json unconditionally — no existence check, no merge,
no backup — on every install and every /gsd-update re-install. On the 11
affected runtimes that file is often user-owned; on OpenCode and Kilo it is
the documented place to declare local-plugin npm dependencies, so a user's
name/type/dependencies/scripts were destroyed on each run.

The uninstall path already read the file and unlinked it only on an exact
content match. That asymmetry was the defect: the discipline existed in the
codebase, it just was not applied on the write side.

Move the marker into the directories GSD creates and fills with its own .js
files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the
nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop
writing the config root entirely. New src/commonjs-marker.cts owns the marker
string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed
on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker,
so install and uninstall cannot drift apart again.

Nothing else depended on the config-root marker: package identity is baked at
build time (#378/#498) and version resolution prefers gsd-core/VERSION and
already tolerates a missing root package.json (#1383) — Codex has installed
without one all along. A package.json in plugins/ or extensions/ is inert to
plugin discovery, which globs *.{ts,js} only (see installer-migration 006).

Uninstall retires the pre-fix config-root marker, so upgrading users are
cleaned up on removal, and still never touches a file it did not write.

* fix(#2544): point the changeset fragment at the filed PR

The fragment's `pr:` field is only knowable after `gh pr create` returns.

* fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set

bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the
linted source), so it belongs in the ADR-457 ignore list like its siblings.
Clears the lint-tests no-var failure and the repo-invariants
"linted xor ignored" migration-state test.

* fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root

The UPGRADE 1 test still asserted the pre-#2544 marker location
(~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the
directory GSD itself creates — matching the updated golden-install-parity
and install-tree fixtures. Also asserts the root marker is NOT written.

* fix(#2544): make the CommonJS marker write path non-fatal

Review round 2, Major 3 + Minor 1 + the stagedHooks nit.

ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site
caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole
install with a raw stack trace. Every other marker interaction in the module is
best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker
swallows read failures — and this was the write path, i.e. the one most likely
to fail on a locked-down config dir. It now returns a new 'failed' outcome and
both call sites warn and continue.

Sibling found while sweeping for the same defect class: fs.mkdirSync sat
OUTSIDE the try block, so an unwritable parent threw past the guard entirely.
Creating the directory is the same environmental hazard as writing into it, so
it moved inside.

Also in this file:

- The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks
  alone. stagedHooks is computed from the SOURCE listing before the copy loop,
  so it stays true when the copies land but verifyInstalled() then fails —
  marking a hooks/ GSD did not successfully populate claims an ownership the
  install did not earn.
- The uninstall rmdir of the native plugin dir is gated on GSD having actually
  removed something from it. Hoisting it out of the adapter-exists guard (so
  the marker-only case could prune) had silently widened it into deleting a
  user-created but empty plugins/ or extensions/ dir — the same "don't touch
  territory GSD didn't fill" principle this issue is about, inverted.
- Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the
  same call site that writes its replacement. That path is outside kimi's
  configDir, so installer-migration 007 structurally cannot reach it.

* fix(#2544): retire the stale config-root marker via installer-migration 007

Review round 2, Major 1 — the PR's headline claim was false for existing
installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale
{"type":"commonjs"} at the config root, so their config root stayed pinned to
CommonJS and their dependency manifest stayed gone until they uninstalled.

The migration is unusual in one way, and it is the part worth reviewing: the
config-root marker was never recorded in gsd-file-manifest.json (writeManifest
records hooks/, agents/, commands/, scripts/ and the native plugin, never a root
package.json), so classifyArtifact answers 'unknown' for it and the planner's
own guard downgrades a remove-managed on an 'unknown' classification to
preserve-user. 007 therefore supplies the "purpose-built detector for an old
GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions —
exact content match, the same predicate removeCommonJsMarker has always used —
and declares the resulting classification on the action. A package.json with any
other content is left untouched, and there is deliberately no backup-and-remove
branch: a non-matching file here is not a patched GSD artifact, it is somebody
else's file.

Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`:
validateStringArray requires the field to be non-empty WHEN PRESENT, while the
runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan
time and the migration never runs. The metadata test pins this.

Kimi is a deliberate carve-out, named in the migration's own header: its marker
lived at ~/.kimi, outside kimi's configDir, and migration relPaths are
structurally confined to configDir. It is retired by the installer instead.

Registration: shipped-migrations table, .gitignore for the emitted .cjs, the
EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not
copied from migration 006 by rote — 006 needs no entry because it imports
nothing, while 007 imports node builtins, so tsc emits its __importDefault
helper and the `var` in it trips no-var. This is the same lint gate that made
round 1 red.

* test(#2544): fault-injection and multi-runtime marker coverage

Review round 2, Major 2 + Minors 4 and 5.

Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and
the suite had no fs monkeypatching at all. Every branch now covered is one whose
doc comment claims it as the module's safety posture:

- classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule),
  with an ENOENT control alongside it so the test discriminates rather than
  just asserting one side
- classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never
  downgrades to the permissive answer) — the fixture's bytes are exactly GSD's
  marker, so the test fails if the code ever answers on content it could not read
- a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already
  covered with a real symlink, the directory case needs no injection at all)
- the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx'
- the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the
  mkdirSync that used to sit outside the guard
- removeCommonJsMarker unlink throw -> false

These save and restore fs methods in `finally` rather than using chmod 0o000,
which does not fault under root and would pass vacuously in root Docker and CI.

Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both
kimi locations now have behavioral coverage, install and uninstall, each paired
with a user-authored-file case proving GSD leaves it alone.

Minor 5 — the stagedHooks gate had no assertion behind its stated reason.
A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime
that declares skipSharedHooksInstall and asserted to stay marker-free, with its
user content intact.

Also regression-tests the uninstall rmdir gate from the previous commit: an
empty plugin dir GSD removed nothing from must survive.

* docs(#2544): correct stale marker prose, register the module, document the trade-off

Review round 2, Minors 2, 3 and 6.

Minor 2 — six files asserted the installed ROOT ships the synthetic marker.
None was load-bearing (all three walk-up consumers are VERSION-first with
try/catch and the marker never carried a `version`), but ADR-457:52 is the
rationale for keeping a generated module, so a future reader would mis-derive
the constraint from it. Each site is corrected to what is now true: the
installed tree carries no package.json with a .name at all, because the only
ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own
directories.

Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js
and the platform-gate test both described `require('../package.json').name`
resolving to undefined; post-#2544 that require does not resolve at all, so the
history is kept accurate and the present-tense claim corrected rather than just
moved. And src/runtime-artifact-conversion.cts described the no-root-package.json
case as Codex-only — it is now every runtime, which strengthens that comment's
own argument for lazy resolution. The generated .cjs sibling needs no edit: it
is gitignored build output, not a tracked file.

Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer
module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added,
including the fail-closed posture and the never-throws contract.

Minor 6 — the plugins//extensions/ marker shadows the config root for all .js
siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is
exactly what #2544's Fix section prescribed and it is disclosed in the PR body,
but the PR body is not documentation. It now lives in the OpenCode section of
docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than
a pure improvement, with the .ts mitigation and a fallback for ESM plugins.

* test(#2544): attribute the CommonJS marker in the emitted-provenance rules

The differential emitted-attribution gate (#2723, landed on `next` after this
branch was cut) went red on the macOS shards once this PR rebased onto it. Two
distinct causes, both real gaps rather than noise:

1. `plugins/package.json` and `extensions/package.json` matched NO rule — the
   `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an
   unattributed emitted family.
2. `hooks/package.json` fell through to `hooks-built`, which attributes an
   emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no
   `hooks/package.json` in the repo, so it resolved to a nonexistent path.

Cause 2 is exactly the failure already documented three lines above it for
Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix
follows that precedent rather than inventing one: `package.json` is excluded
from `hooks-built` the same way, and a dedicated `commonjs-marker` rule
attributes the family across all four roots it can appear in (both hooks roots
plus `plugins`/`extensions`) to the sources that actually emit it.

Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for
a one-off ripple and goes stale by design — the gate fails a stale ack precisely
so it cannot pre-clear the next change on that path. These markers are a
permanent part of the emitted tree from #2544 onward, so they need standing
attribution.

Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3
provenance errors + 12 unattributed paths before, 35/35 green after.

* fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker

#2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned
in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via
dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's
module on the two properties that matter:

  - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a
    DANGLING one, so a dangling package.json symlink classified as absent and
    the write went straight through it. Demonstrated: against the pre-fix copy,
    ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json
    symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory.
  - create: a plain writeFileSync leaves the classify->write window open, where
    commonjs-marker creates with flag:'wx' (O_EXCL).

Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's
own docstring already claimed was the single place these rules are enforced.
Exported signatures are unchanged (still boolean), so bin/install.js and the
#2717 tests are unaffected.

The new subtest is the only coverage that fails if the duplicate is ever
reintroduced — the two implementations agree on every non-adversarial input, so
the existing suites pass against both.

* test(#2544): pin the stagedHooks gate on zcode, not windsurf

The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall
kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the
marker was correctly absent.

#2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via
dedicated paths and get the marker beside those scripts. Measured on this tree,
windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning
behaviour that is now wrong, not the gate it was written for.

ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin
surface to spawn hooks, so GSD stages no .js there by either route (measured: 0
staged, no marker). The property under test is unchanged — a user-created hooks/
directory GSD never fills stays marker-free.

* test(#2544): use the shared cleanup helper in the migration test

Addresses the review's Major 1. The suppression's stated reason — "no helpers
import available" — was not correct: tests/helpers.cjs exports cleanup, and the
other test file added in this same PR imports it (tests/commonjs-marker.test.cjs).

The local reimplementation dropped two protections that are live on this repo's
windows-latest lane: the CWD guard (Windows cannot remove a directory that is the
current working directory) and the 20 x 250ms retry budget that absorbs the
deferred-scan handle Windows Defender holds on newly-written files.

Local function and suppression both removed; local/no-raw-rmsync-in-tests now
passes without one.

* test(#2544): expect hooks/package.json for the #2717 runtimes

The fresh-install contract table predates #2717, which stages cursor/windsurf/
codex .js hooks via dedicated paths and writes the CommonJS marker beside them.
All three therefore now receive hooks/package.json legitimately.

Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with
the marker; cline/copilot/trae/zcode stage none and get none, so their contracts
are unchanged.

* fix(#2544): gate the #2717 marker writes on having staged something

The three dedicated marker writers #2717 added ran unconditionally. Each one
mkdirs hooks/ up front and stages its scripts conditionally on the source
existing, so with an absent or empty hook source they created a directory,
filled it with nothing, and marked it as GSD's anyway.

That is the same write-into-someone-else's-territory this issue is about, and
installSharedHooksBundle already guards the identical case with `stagedHooks`.
The dedicated paths now carry the matching gate:

  - cursor / windsurf: `installedScripts.size > 0`
  - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that
    hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY
    entry landed.

Covered for cursor and windsurf by driving each writer against a src tree whose
hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its
trigger state needs a package tree where hooks/dist exists but holds none of the
allowlist, which is not constructible from a real checkout.

* test(#2544): scope the commonjs-marker sources per root

The rule declared one flat source list for every marker root, so
`extensions/package.json` was attributed to runtime-hooks-surface.cts (which
never writes there) and `.kimi/hooks/package.json` to install-engine.cts.

That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source,
so a flat list containing bin/install.js let any change anywhere in that
13k-line file authorise marker drift for every root — the blanket escape hatch
this file's own agents-verbatim comment refuses for exactly the same reason.

Sources are now derived per root from ctx.rel. Note the rule ctx is
`{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent
every path down one branch silently.

* test(#2544): state precisely what the zcode assertion pins

The comment claimed the test pinned installSharedHooksBundle's `stagedHooks`
gate. It does not, and neither did the windsurf version it replaced: zcode
declares skipSharedHooksInstall, so the outer guard skips that helper entirely
and the gate is never evaluated. The test passes on the runtime exclusion.

What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays
marker-free — is still worth having, and is what the review asked for. The two
`staging zero hook scripts` tests are the ones that pin a real staged-nothing
gate. Comment corrected rather than left implying coverage that is not there.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:00:23 -04:00
Tom Boucher
628648d63a chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)
* fix(#2931): preserve protected regions and cap emitted per-runtime bytes

Route every runtime brand swap through applyClaudeCodeBrandSwap so
"Claude Code" survives verbatim inside <runtime_compatibility> regions
(#2284b). The fix existed only in bin/install.js's local copies; the
src/*.cts exports still used a naive replace, so binding install.js to
the single source -- as this phase does for the Windsurf family --
would have silently regressed those runtimes. A table-driven parity
guard now covers all nine brand-swapping converters.

De-duplicate the Windsurf converter family: delete the six local copies
in bin/install.js and bind the four exported ones by reference, guarded
by reference-identity assertions (the ADR-1508/#1675 pattern). The two
unexported helpers and an unused tool table go with them.

Replace the Windsurf 12,000-byte throw with description truncation,
matching the bound its sibling skill converter already applied. The
throw could only fire on an ~11.7 KB frontmatter description: the
largest emitted workflow is 311 bytes. Truncation makes the cap
unreachable by construction and leaves 12,000 in exactly one place,
eliminating the dual-surface duplication rather than testing for it.

Add the emitted-byte cap gate: buildEmittedSizes captures LF- and
<HOME>-normalized bytes from the walk buildParityManifest already
performs, and evaluateEmittedCaps asserts them against a per-runtime
cap table with dead-rule detection. buildParityManifest's return shape
is deliberately unchanged -- diffEmitted compares its values with
===, so making them objects would report all 8,529 emitted paths as
moved. A regression test pins the values as strings.

Add a deterministic trim-safety gate over composeWithinBudget's
omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity
rule, replacing the model-graded eval gate the issue described.

* docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract

* fix(#2931): bound the windsurf command name and single-source the brand swap

Review findings from the orthogonal passes, all fixed inline.

The claim that removing the 12,000-byte throw left total emission
"bounded by construction" was false. The #1615 regex constrains the
character class but not the length, and commandName is interpolated
three times into the emitted workflow: a 20,000-character name emitted
60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate,
clearly-labelled size control that THROWS -- commandName is the @-ref
path target, so truncating it would point the workflow at a file that
does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The
#1615 security regex is untouched and still runs first. 128 is generous:
the longest shipped name is gsd-plan-review-convergence at 27.

Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe
truncation helper. It still used a UTF-16 slice(0,177) -- the exact
surrogate-splitting bug the helper was written to avoid, in the very
sibling the helper's comment cites as its model. Bounds are unchanged,
so output is byte-identical for every shipped command (descriptions max
out at 99 chars).

Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting
the local copy. Adding it to the .cts left two unlinked implementations
of identical logic -- the drift class this change exists to remove.
Verified byte-identical across eight fixtures and five sequential calls
before merging, and guarded by a reference-identity assertion.

Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344),
add fast-check property coverage for the trim-safety contract, and use
fc.pre instead of a bare return in a property callback.

* test(#2931): fix three test-authoring bugs the remote matrix caught

The remote runner returned 8 unique failures on 6f15cdeb8. All three
causes were in the test files, not the modules under test -- local
harnesses exercise the modules directly, so nothing executed the test
bodies until the matrix did.

`{ __proto__: [...] }` in an object literal sets the prototype instead
of an own key, so the JSON round-trip erased it and the cap table never
saw a reserved runtime key. The production rejection was already
correct; the test could not reach it. Use a computed key.

Two cap fixtures tripped orthogonal error paths rather than the paths
they name: one declared windsurf in the cap table but omitted it from
sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern
matching nothing (a genuine dead rule). Both now include a compliant
artifact so the intended branch is what is asserted. The dead-rule and
unknown-runtime contracts are deliberate and unchanged.

`const { root } = makeSyntheticConfig({ ... `${root}` })` referenced
`root` from inside its own initializer -- a temporal dead zone error.
makeSyntheticConfig now optionally takes a (root) => files factory.

Also raise the npm pack --dry-run bound 60s -> 120s in the shipped-
scripts packaging test. That failure is NOT from this branch: the file
is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s
here, and the run recorded 60,637ms against a 60,000ms bound -- a
timeout under 28,948-test parallel contention, not a slowdown. Fixed
rather than deferred because a bound that tight is fragile regardless
of which branch trips it.

* chore(#2931): backfill changeset pr number to 2984

---------

Co-authored-by: sim <sim@local>
2026-08-01 16:00:14 -04:00
Tom Boucher
c043f2946c fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next

tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900.
Every entry is scoped to the diff that introduced it (#2789), so once merged
to next it is at the base by definition -- spent and inert. Its presence is
still load-bearing though: each PR rewrites the paths map wholesale, making a
persistent base copy a shared cell. Five of six conflicting PRs in the open
queue collided on this file and nothing else.

Deletes the stale document and adds a push-to-next guard asserting it stays
absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check
against the base is the #2768 shape #2789 exists to end.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2914): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2914): per-PR ack fragments instead of one shared mutable file

The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json
whose paths map every PR rewrote wholesale. That is a shared mutable cell: any
two PRs needing an ack edit the same lines and conflict. Five of six conflicting
PRs in the open queue collided on this file and nothing else.

Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same
shape .changeset/ already uses to solve this exact problem. Two PRs pick
different filenames, so they cannot collide, and fragments lingering on next
are harmless rather than toxic.

The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An
earlier delete-only attempt failed verification twice: the ratchet lost the
spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no
ack. Relocating preserves every acknowledgment.

The legacy single file is still READ (unioned with the fragments) because five
open PRs carry it; dropping support would break all of them. A duplicate path
key across sources is a hard error, never last-wins.

The push-to-next guard is retargeted accordingly: it now asserts only that the
legacy SHARED file never reappears on next. Fragments may persist harmlessly.

Closes #2914

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:15:29 -04:00
Tom Boucher
f093738412 fix(#2891): normalize emitted version against the measured tree, not the measuring repo (#2894)
* fix(#2891): normalize emitted version against the measured tree, not the measuring repo

buildParityManifest normalized the install-time {{GSD_VERSION}} stamp using
PKG_VERSION, bound at module load from the MEASURING repo's package.json. Since
#2767, currentManifests({repoRoot}) measures a DIFFERENT checkout, so during a
release cut the baseline worktree (origin/next, 1.8.0) was normalized with the
current tree's version (1.9.0) and its literal 1.8.0 stamp survived into the
hash. All 364 emitted hook paths diverged and the differential attribution gate
hard-failed every finalize/rc run.

Normalize against the version of the tree that PRODUCED the emitted output:
buildParityManifest takes an explicit pkgVersion, and currentManifests resolves
it from the measured tree via a new fail-closed measuredPackageVersion().

* chore(#2891): backfill changeset pr number (#2894)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 22:14:47 -04:00
Tom Boucher
aa19e3478c fix(#2854): pin the emitted gate to the base the tree was merged with (#2859)
* test(#2854): failing-first coverage for CI baseline export provenance

Extracts the export decision out of main() behind injected IO so it is
unit-testable, preserving today's export-whenever-present semantics, and
adds the matrix that proves those semantics are wrong.

The PR lane restores the emitted baseline keyed on the PR's recorded base
sha while the gate resolves the base ref live, so the two drift whenever
next advances mid-flight. The restore was published straight to
GSD_EMITTED_BASELINE, where a mismatch is fatal, turning a recoverable
cache into a hard failure on diffs that touched nothing related.

Also renames the stale-env fixture from 'from-cache-restore.json' to an
operator-pin name: that fixture asserted the exact conflation this bug
is, documenting the defect as intended behavior.

Refs #2854

* fix(#2854): validate a restored baseline before publishing it as an operator pin

GSD_EMITTED_BASELINE is an operator pin: resolveBaseline() treats a mismatch
there as a hard stop, because the operator said "use this one". CI published
its cache restore to that same variable whenever the file merely existed, so
a restore keyed on the PR's recorded base sha - while the gate resolves the
base ref live - turned a recoverable cache into a fatal error whenever next
advanced mid-flight. Required tests went red on diffs that touched nothing
related, and named a test file the contributor never opened.

The export step is the boundary, so it is the boundary that validates. It now
publishes only a baseline already valid for the sha under test, judged by
validateBaseline so the staleness rule keeps one definition. Anything refused
is left to be found via the cache path, where a mismatch degrades to the
in-job build exactly as ADR-2719 SS5 specifies. The operator hard stop is
untouched, and the fast path still hits on a current cache.

Also reports the sources actually reached rather than asserting all three
ran: the failure message claimed an in-job build it had returned before
calling, sending contributors after a rebuild that never happened.

Fixes #2854

* fix(#2854): pin the emitted gate to the base the tree was actually merged with

The differential compared a tree built on one commit against a baseline at a
different one. "Rebase check" merges pull_request.base.sha, pinned by #2472 so
all 12 matrix jobs agree on one tree, but resolveBase() fell through to
origin/next, which fetch-depth 0 leaves at the live tip. Nothing set
GSD_EMITTED_BASE, so whenever next advanced mid-flight the two disagreed.

The cached baseline, keyed on base.sha, was correct for that tree and was
rejected as STALE by a target that was not. Required tests went red on diffs
that touched nothing related, naming a test file the contributor never opened.

The near miss is the worse half: had resolution gotten past the baseline step,
a baseline at the live tip would have attributed commits merged to next in
between to the PR under test. The hard stop was shielding us from a wrong
answer, so making it fall through would have made this worse.

Pins GSD_EMITTED_BASE to the same expression as CI_REBASE_BASE_SHA in every
rebase-merged job, with a parity test asserting the two cannot diverge. That
parity check immediately caught a third lane, test-inert, that merges a pinned
base and had been missed.

Fixes #2854

* fix(#2854): keep the export step self-contained across the package boundary

scripts/ ships in the npm tarball and tests/ does not, so requiring the
validator across that boundary is MODULE_NOT_FOUND in a published install.
The export step now reads the same GSD_EMITTED_BASE pin the gate resolves
through, and applies a cheap self-contained precondition; validateBaseline
remains the sole authority and still runs downstream on whatever is
published, so there is no second opinion to drift.

Reading the pin rather than re-deriving a base is the point: a second,
divergent base lookup is exactly what caused this bug.

The same hazard pre-exists in scripts/gen-emitted-baseline.cjs, which ships
and requires three tests/ modules. Filed as #2858 rather than folded in:
fixing it means relocating the shared helpers out of tests/ and updating
every consumer, which would bury this change.

Refs #2854

* fix(#2854): stop announcing the restored cache through the operator-pin door

Two independent reviewers found the same blocker in the previous approach.
Validating before publishing to GSD_EMITTED_BASELINE only narrowed the hole:
the precondition gated on sha equality alone, so a document with a MATCHING
sha but a wrong schema version or malformed manifests was still announced as
an operator pin and still hard-stopped downstream. That reproduces this bug's
own class, triggered by malformation instead of staleness.

The step was never load-bearing. The cache restores to
.gsd-cache/emitted-baseline.json, which is resolveBaseline's DEFAULT_CACHE_PATH
and is read whether or not anything announces it. Publishing the same file to
the pin door could only ever convert recoverable into fatal, so the step and
its script are deleted rather than made cleverer. Every failure mode now
degrades to the in-job build by construction, and validateBaseline is once
again the only thing that judges a baseline.

Coverage moves to where the behaviour lives: stale sha, wrong schema version,
manifests array/absent, non-object documents, unreadable file, and the 39/40/41
hex boundary all assert degradation via the cache path. Adds the empty-pin case
a reviewer flagged as untested - the pin is job-level env, so on push events it
renders as an empty string, and only baseRefCandidates' truthy check keeps it
out of the candidate list.

Fixes #2854

---------

Co-authored-by: Test <test@example.com>
2026-07-30 11:19:45 -04:00
Tom Boucher
4f6935e29b fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks (#2846)
* test(#2717): CommonJS marker for cursor/windsurf/codex staged .js hooks

Cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) stage .js
hook scripts via dedicated paths that bypass installSharedHooksBundle — the
only writer of the {"type":"commonjs"} marker. Under a config root declaring
{"type":"module"}, Node loaded those scripts as ESM and every require()
failed with 'require is not defined', silently disabling the runtime's hooks.

Adds regression tests (RED first, fix lands next commit):
- parametrized cursor/windsurf/codex install asserts hooks/package.json exists
  with exactly GSD's marker content;
- end-to-end: a cursor require()-using hook loads under a planted ESM-typed
  config root without the require-is-not-defined error;
- the ensureCommonJsMarker / removeCommonJsMarkerIfGsdOwned contract: GSD
  markers are removed on uninstall, user-authored package.json is never touched.

* fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks

The {"type":"commonjs"} marker lived only inside installSharedHooksBundle,
which cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) never
reach. Their .js hooks are staged by dedicated paths, so under a config root
declaring {"type":"module"} Node loaded them as ESM and every require()
failed with 'require is not defined', silently disabling those runtimes' hooks.

Decouple the marker write into a shared helper so any code path that stages
.js hooks can ensure it lands in the SAME directory as the scripts:

- src/runtime-hooks-surface.cts: add ensureCommonJsMarker(dir) +
  removeCommonJsMarkerIfGsdOwned(dir) (byte-identical content to
  installSharedHooksBundle's marker; preserves a user-authored package.json on
  both write and uninstall). Call ensureCommonJsMarker(hooksDir) from
  writeCursorHooksJson + writeWindsurfHooksJson; call
  removeCommonJsMarkerIfGsdOwned on their matching remove paths. Export both.
- bin/install.js: call hooksSurface.ensureCommonJsMarker after the codex hook
  copy; call hooksSurface.removeCommonJsMarkerIfGsdOwned in the generic
  hooks-removal loop (safe no-op where no marker exists).

No change to which runtimes receive the shared bundle, the !isCodex gate,
skipSharedHooksInstall, or kimi/kimi-code/cline/copilot/trae/zcode (all
unchanged — audit in the diagnosis). RED @ dbb7d2bb (6 failures: 3 missing
markers + the ESM require error + missing helpers); GREEN pending.

* docs(#2717): changeset fragment (pr:0, backfilled post-PR)

* chore(#2717): regen codex/cursor/windsurf install-tree fixtures + attribution ack

The fix adds hooks/package.json to those three runtimes' install trees (the
new CommonJS marker), so the golden install-tree fixtures gain one path each
(regenerated via npm run gen:install-tree). emitted-attribution (ADR-2719)
flags the 3 emitted hooks/package.json paths under the hooks-built rule;
acknowledge them. Also drops 5 spent ack entries left by now-merged PRs
(#2694 code-review.md/code-review-fix.md, #2695 worker/registry, #2794
review.md) — they are stale on this branch (base already carries them).

* fix(#2717): codex ESM-root behavioral test + hooks-built provenance for package.json

Two review-driven follow-ups on the #2717 fix:
- Adversarial review noted the ESM-root behavioral test covered only cursor;
  refactor it into a helper and add a codex case (the !isCodex-gated path most
  likely to regress, whose marker write lives in bin/install.js). gsd-check-update.js
  require()s at module load, so it surfaces the ESM failure immediately.
- emitted-provenance flagged hooks/package.json as 'attributed source does not
  exist' — the marker is code-derived (a fixed literal emitted by
  ensureCommonJsMarker at install time), not built from a tracked source. Route
  the hooks-built rule's sources/transforms for package.json to the surface
  source file, mirroring the existing .cmd-shim sub-family.

* chore(#2717): drop now-redundant hooks/package.json attribution ack

The hooks-built provenance routing (prior commit) now self-attributes the
emitted hooks/package.json to src/runtime-hooks-surface.cts, which IS in this
diff — so the attribution is self-explaining and the emitted-drift-ack entry
became stale. Delete the (now-empty) ack file per ADR-2719's empty-file rule.

* docs(changeset): backfill #2717 PR number to 2846
2026-07-29 22:18:48 -04:00
Tom Boucher
1e3c995e6f fix(#2789): scope the emitted-drift ack to the diff that introduced it (#2803)
* fix(#2789): scope the emitted-drift ack to the diff that introduced it

Every input to `diffEmitted` is base-relative -- `baseline` vs `current`,
`changedPaths` from `git diff base...HEAD` -- except the ack set, which
was read absolutely, from the working tree only. A differential machine
consulting a non-differential input.

So `staleAcks` asks exactly one question, "did a delta consume you?", and
that cannot distinguish an ack that never explained anything (an
authoring mistake) from one whose ripple is now absorbed into the base
(the ack's SUCCESS condition). After merge an ack is in the second state
but reports as the first.

The trigger is ordinary. Actions sets GITHUB_BASE_REF on pull_request
events only, so a push to `next` falls through to origin/next -- the very
commit under test. Both sides build identical content, no deltas remain,
and every live ack is reported stale. PR #2768 acked a deliberate 40866
-> 42020 byte growth, was green on its own lane, and reddened `next` the
moment it merged. It also reds every PR branching off the poisoned base,
and since publish-emitted-baseline is gated on the test job, it blocked
baseline publication too.

Give the ack the base side it was missing. `diffEmitted` now takes
`baseAck` -- the same document at the base ref, via `readAckFileAtRef`.
An entry already present there is SPENT: it may no longer consume a delta
and is never reported stale, only surfaced as `spentAcks` for tidying. An
entry new or reworded in this diff stays live, and if nothing consumes it
that genuinely fails, with blame on the author who just wrote it.

This closes a hazard the IMPLEMENTATION named but could not prevent -- a
leftover ack silently pre-clearing the next ripple on its path. (ADR-2719
§3 asserted only that TOUCHING the file is the alarm; its residual-risk
list never covered pre-clearing, and §3 now carries an amendment.)
Verified against the two-PR laundering sequence -- land an innocuous ack,
then change the artifact -- which passed silently before and now fails on
both the hash pass and the size ratchet.

Three things the design has to get right, each of which was wrong first:

  - A read failure on the base document THROWS; only absence-at-the-ref
    returns null. Returning null on error LOOKS armed (every entry stays
    live) but a live entry's defining power is that it CONSUMES a delta,
    so null is armed on the staleness axis and DISARMED on consumption --
    silently the whole pre-#2789 gate. `git show` cannot tell absence
    from fault, so absence is established with `ls-tree`.
  - Re-arming a spent ack costs actual PROSE. Internal whitespace and the
    zero-width family collapse, and `runtime` is not compared: a doubled
    space, an invisible character, or a decorative field would otherwise
    re-arm an ack whose justification still describes the previous
    ripple, showing a reviewer nothing.
  - `baseAck` is REQUIRED once an ack declares entries -- omission is an
    error, not a silent "inherit nothing" -- so a dropped argument fails
    loudly instead of quietly restoring this bug with the suite green.

Because a corrupt document ON THE BASE is expensive (the loud base-side
failure reds every ack-carrying PR), scripts/lint-emitted-drift-ack.cjs
blocks one from landing. It is standalone rather than importing parseAck
-- scripts/ ships in the npm package and tests/ does not -- so a parity
test runs both surfaces over one corpus and fails on divergence; it
caught one immediately, a `null` document, now classed as policy rather
than schema. Deadlock is separately foreclosed: a tree carrying no ack
never reads the base, so the PR that DELETES a corrupt file still lands.

`readAckFileAtRef` takes an injected git runner so all four branches are
tested deterministically; it never executes in the remote runner, where
the real-tree test skips for want of a base ref. It also refuses an
option-shaped ref, since execFileSync's array form stops shell
metacharacters but not git's own option parsing.

Rejected: skipping the differential when base == HEAD. It treats the
symptom, costs real coverage on the push-to-next lane, and does nothing
about the downstream PRs the same flaw was reddening.

Deletes the now-spent tests/emitted-drift-ack.json, and updates the
CONTEXT.md canon and ADR-2719 §3: presence is no longer the alarm -- a
LIVE entry is, and a spent one is inert.

Closes #2789

* chore(#2789): backfill changeset PR number
2026-07-28 21:28:10 -04:00
Tom Boucher
e276cc7f00 enhance(#2778): make the size-ratchet failure name its own remedy (#2780)
* fix(#2778): exempt intentionally-absent paths from the glossary gate

check-glossary-refs asserts that every backticked tests/ token in
CONTEXT.md resolves on disk. tests/emitted-drift-ack.json (ADR-2719
section 3) is absent on a healthy next BY DESIGN — it appears only
inside a PR that needs it, which is what makes touching it the alarm.

It passed before only by accident of backtick pairing: CONTEXT.md's
RULESET entries are themselves backtick-wrapped and contain backticks,
so the token happened to fall outside a code span. Any edit that
shifted the parity exposed it. A gate that passes by luck is not
passing.

The exemption is exact, not a prefix hole: a sibling missing tests/
path still fails, and a test locks that.

* feat(#2778): make the size-ratchet failure name its own remedy

The growth branch stated a requirement and withheld the means of
satisfying it: no ack file named, no schema, no key format, and no
do-not-regenerate line — so the likeliest guess was to hunt for a
baseline that #2724 deleted. Observed live on #2543.

All remediation now comes from one frozen REMEDIATION export whose
example document is rendered from ACK_VERSION, so the taught schema
cannot drift from the schema parseAck accepts. A round-trip test feeds
the printed document back through parseAck.

The report is now built as a typed IR (buildReport) that formatReport
renders, so tests assert on structure rather than prose, per
CONTRIBUTING.md's raw-text-matching rule.

Two defects found and fixed inline while building:
- diffEmitted's validation early-return omitted newFileCapExceeded
  while formatReport reads its length, so the branch that reports a
  failed git diff threw a TypeError instead of naming the problem.
- Printing one complete ack document per failing branch made each read
  as the whole file, so pasting the second over the first silently lost
  an acknowledgment. One document now covers the whole report.

Closes #2778

* chore(#2778): backfill changeset pr number to 2780
2026-07-28 18:26:55 -04:00
Tom Boucher
1c1af70a4b refactor(#2724): delete the committed golden fixtures and size baselines (#2767)
* test(#2724): delete golden-install-parity fixtures, test, and generator

Removes the 19 committed path->hash manifests, the two per-file size
baselines, tests/golden-install-parity.test.cjs, and
scripts/gen-golden-install-parity-zcode.cjs. These were pure functions
of the source tree (ADR-2719); the differential attribution check
(tests/emitted-attribution.test.cjs + tests/emitted-provenance.test.cjs)
is now the sole gate for emitted-artifact propagation.

tests/fixtures/install-tree/*.json and tests/golden-install-tree.test.cjs
are unchanged (ADR-2719 section 7 exception).

Follow-up commits fix the resulting bookkeeping: scripts/ci-test-scope.cjs's
existence guard, .gitattributes, package.json scripts, the emitted-provenance
totality guard's IO, the differential check's baseline acquisition, CI
wiring to publish/restore the baseline artifact, and docs.

* refactor(#2724): make the differential attribution check self-sufficient

Three fixes required to delete the golden fixtures without breaking CI:

- scripts/ci-test-scope.cjs: remove tests/golden-install-parity.test.cjs
  from the three rules that named it. #2759's missingRuleTestFiles guard
  hard-throws at module load if a rule names a test file absent from
  disk, which would break the changes job on every PR the moment the
  fixture-deletion commit landed.

- tests/helpers/emitted-provenance.cjs: loadManifests() read the
  committed golden fixture directory. With that directory deleted at
  every future ref, this would throw at module load forever, taking
  the Phase 2 totality guard down with it. Rebuilt from real installer
  spawns (MANIFEST_FAMILIES + runMinimalInstall + buildParityManifest),
  the same shape emitted-runtime.cjs's currentManifests() already uses.

- tests/emitted-attribution.test.cjs / tests/helpers/emitted-runtime.cjs:
  the real-tree test's baseline acquisition swaps from
  baselineManifestsAtRef(base) (git show at a ref that no longer carries
  fixtures) to resolveBaseline()'s documented precedence: env, then the
  on-disk cache, then an in-job build. The build fallback
  (buildBaselineAtRef, new) checks out base into a throwaway git
  worktree and runs the new scripts/gen-emitted-baseline.cjs there --
  no npm ci needed, since bin/install.js and the test helper shells are
  Node-builtins-only. That script also publishes the baseline artifact
  from CI's push-to-next job (wired in a follow-up commit).

* refactor(#2724): retire the merge-driver bridge and per-file size baselines

The Phase 1 bridge (#2721) is retired now that the artifacts it guarded
are deleted: scripts/git-merge-regen-driver.cjs, its test, and the
'setup:merge-driver' npm script are removed, and the .gitattributes
merge=gsd-regen/linguist-generated block for the three deleted-path
globs is dropped. tests/fixtures/install-tree/*.json keeps its normal
merge behavior, unchanged (ADR-2719 section 7).

scripts/update-size-baseline.cjs and its test are removed: their sole
purpose was regenerating tests/workflow-size-baseline.json and
tests/agent-size-baseline.json, both deleted. The 'size:baseline' npm
script and its step in 'regen:derived' go with it. The per-file
baseline describe blocks in tests/workflow-size-budget.test.cjs and
tests/agent-size-budget.test.cjs are removed for the same reason; the
independent loose-tier hard caps are untouched. The differential
attribution check's size ratchet (tests/emitted-diff.cjs, already
shipped in #2723) is the replacement anti-creep mechanism.

'npm run gen:golden' is replaced by 'npm run gen:install-tree', which
keeps regenerating tests/fixtures/install-tree/*.json (the one artifact
family ADR-2719 section 7 keeps committed); tests/golden-install-tree.test.cjs's
error messages point at the new command name.

tests/golden-parity-single-source.test.cjs's anti-divergence guard
(#2266) is retargeted from the two deleted golden-parity consumers to
their two replacements (tests/helpers/emitted-runtime.cjs and
tests/helpers/emitted-provenance.cjs), which import buildParityManifest
the same way — the divergence risk the guard exists for is unchanged.

Also wires CI: a new publish-emitted-baseline job runs
scripts/gen-emitted-baseline.cjs after a push to next and caches the
result keyed on the sha; the test and test-full jobs restore that cache
on pull_request events, keyed on the PR's base sha, and export
GSD_EMITTED_BASELINE for tests/emitted-attribution.test.cjs's real-tree
test to pick up.

* docs(#2724): flip ADR-2719 to Accepted and update contributor docs

Status: Proposed -> Accepted. Regenerated docs/adr/README.md index.

CONTRIBUTING.md, docs/TESTING-SUITES.md, and CONTEXT.md (RULESET.
EMITTED_ATTRIBUTION, RULESET.WORKFLOW_SIZE_BUDGET, RULESET.
AGENT_SIZE_BUDGET, and the Emitted Artifact Provenance glossary entry)
no longer point at the deleted golden-install-parity fixtures, size
baselines, gen:golden, UPDATE_GOLDEN, or the setup:merge-driver /
git-merge-regen-driver.cjs bridge. Editing shipped content now
requires zero manual fixture regeneration, documented against the
differential attribution check instead of the deleted commands.

* docs(#2724): add changeset for removed golden-parity commands

* fix(#2724): drop stale scripts/update-size-baseline.cjs glossary ref

check-glossary-refs.cjs verifies every backtick-wrapped scripts/*.cjs
token in CONTEXT.md resolves to a real file. The RULESET.
EMITTED_ATTRIBUTION rewrite named the deleted script inside backticks,
which the checker reads as a live reference, not historical prose.

* test(#2724): retarget ci-test-scope tests off the deleted golden test

tests/ci-test-scope.test.cjs asserted specific RULES entries select
tests/golden-install-parity.test.cjs, and that every rule selecting it
also selects both emitted gates. Both premises broke when the golden
test was deleted (#2724): the deleted filename never re-appears in
targeted_tests, and there was no longer a third file for the gates to
travel alongside. Retargeted the two selection describe blocks to
assert tests/emitted-provenance.test.cjs directly (the drift guard the
golden gate's rules were retargeted to), and simplified the third block
to assert the two emitted gates always travel together, without
reference to the golden filename.

* docs(#2724): repoint two contributor how-to guides at the differential check

Both guides told contributors to regenerate a baseline against
tests/golden-install-parity.test.cjs, which #2724 deletes. Repointed
at the differential attribution check (tests/emitted-attribution.test.cjs,
ADR-2719), which needs no manual regeneration step.

* fix(#2724): repair phase6-capstone-conformance's deleted-baseline read

An independent orthogonal review caught a real regression this branch
introduced into a test file the branch's diff never touched:
tests/phase6-capstone-conformance.test.cjs read
tests/workflow-size-baseline.json (deleted earlier in this branch) with
no fallback, so the whole suite would throw ENOENT the moment this
branch landed. The test's actual intent — prove the host-loop workflow
files are real, tracked, non-empty docs — is preserved by asserting the
live byte count via the same shared counter (scripts/workflow-size.cjs)
the size guards already use, instead of a committed snapshot.

Also, from the same review: a stale doc comment in
scripts/workflow-size.cjs still named the deleted
scripts/update-size-baseline.cjs as a consumer, and
buildBaselineAtRef's cleanup in tests/helpers/emitted-runtime.cjs left
two fs.rmSync calls unguarded against masking the primary result/error,
inconsistent with the try/catch already wrapping the git cleanup beside
them. Both fixed. A doc comment was added to baselineFamilyNamesAtRef
explaining why it (and its siblings) are kept despite having no
production caller post-cutover — they still answer real questions
about refs that predate the cutover.

* fix(#2724): repair three real regressions found by remote verification

1. tests/emitted-provenance.test.cjs's two hostile-input tests
   (non-object manifest, unreadable fixture) drove loadManifests(tmp)
   and monkeypatched fs.readFileSync, both premised on the deleted
   fixture-directory read this branch already replaced with real
   installer spawns -- the negative assertions silently stopped firing.
   loadManifests() now accepts injected {families, install, build,
   clean} (defaulting to production values), giving the tests a real
   seam to drive a bad build result and a build failure through the
   ACTUAL loader instead of a reimplementation, and added coverage that
   clean() still runs on both paths.

2. .github/workflows/test.yml's two 'Export GSD_EMITTED_BASELINE'
   steps hardcoded shell: bash, which is wrong on windows-latest (native
   pwsh) and on test-full's macos-latest legs (native zsh per that job's
   own matrix) -- the repo's H1 shell policy (tests/policy-shell-pinning
   .test.cjs) caught it. Replaced the inline bash script with
   scripts/ci-export-emitted-baseline-env.cjs, a plain Node script: a
   bare 'node <path>' command line has no shell-specific syntax, so it
   runs correctly under bash, zsh, and pwsh without a shell override.

tests/phase6-capstone-conformance.test.cjs's deleted-baseline read
(caught by the same remote run, at a commit prior to this one) was
already fixed in d0c3b1242 and is not touched here; verified still
passing after these changes.

* fix(#2724): revive ADR-1610's new-file size cap inside the differential

An isolated review caught a real regression: deleting
tests/workflow-size-baseline.json silently dropped NEW_FILE_CAP
(ADR-1610 Decision point 3, the Codex project_doc_max_bytes anchor)
with no successor. tests/helpers/emitted-diff.cjs's size ratchet
already 'continue's past any file absent from sizeBaseline -- exactly
the files this cap exists to bound -- so a brand-new workflow file
sized 32,769-40,960 bytes passed CI clean and shipped, then risked
silent truncation at the Codex anchor at runtime. ADR-1610 is Accepted
and never referenced anywhere in this branch.

Fix: NEW_FILE_CAP=32768 revived inside emitted-diff.cjs's own
size-ratchet loop, keyed off the SAME hasOwnProperty(sizeBaseline,
name) signal the growth check already computes -- 'new' is exactly
'present in sizeCurrent, absent from sizeBaseline'. Not ack-able,
matching the tier hard caps it sits beside: the fix is extraction, not
an acknowledgment entry. Documented, disclosed narrowing: the pure
differential module cannot see XL_WORKFLOWS/LARGE_WORKFLOWS tiering
(tests/workflow-size-budget.test.cjs's classification), so a
legitimately large new file must extract rather than tier in, one
release earlier than an existing file would need to. ADR-1610 itself is
left unamended -- this restores its decision rather than re-litigating
it.

Also fixes a stale comment plus a redundant real 19-installer-spawn
assertion left over from the pre-injection-seam version of
tests/emitted-provenance.test.cjs's build-failure test, and annotates
3 of 4 stale golden-fixture citations in
docs/reference/host-integration-capability-matrix.md as superseded
(the 4th is an accurate historical PR narrative, left alone).

* fix(#2724): repair three red CI defects on the golden-fixture cutover

Windows-only provenance false attribution (defect A): the `hooks-built`
provenance rule attributed `hooks/<name>.cmd` to itself. Those shims are
Windows-only installer output (ensureCodexHooksJsonSessionStart /
ensureCodexHooksJsonEvent, both in src/runtime-hooks-surface.cts) wrapping
the same-named `.js` hook — no `.cmd` file is ever tracked in the repo, so
the self-attribution resolved to a path that exists on no platform. Only
windows-latest ever emits the key, so this only failed there. Fixed by
special-casing `.cmd` inside the SAME `hooks-built` rule (not a dedicated
rule) — a dedicated rule would match zero paths, and therefore report as a
dead rule, on every non-Windows lane of the same totality guard. `sources`
already supported per-match functions; `transforms` is extended to support
the same shape so the attribution can vary by match within one rule.

Baseline bootstrap was structurally impossible (defect B): `buildBaselineAtRef`
ran `scripts/gen-emitted-baseline.cjs` from INSIDE the base-ref worktree, but
that script is new in this PR and therefore absent at any base ref that
predates it — every call failed closed with "Cannot find module". Fixed by
running the PR checkout's own generator against the worktree via a new `--dir`
parameter, decoupling "which copy of the script runs" from "which tree it
measures" (`currentManifests`/`currentSizes` gained a `repoRoot` override,
threaded down to `runMinimalInstall`'s new `installScript` override). This is
not just a bootstrap fix: a differential needs ONE measurement schema applied
to both sides, or the two stop being comparable the moment that schema
evolves — running each side's own copy would silently reintroduce that risk.
Verified locally end-to-end against real origin/next: resolves a valid
{version, sha, manifests, sizes} artifact with the correct sha and no leaked
worktree.

Changeset placeholder (defect C): `pr: 0` -> `pr: 2767`, which is what let
docs-lint evaluate the fragment for the first time; it already passes
(docs/TESTING-SUITES.md and friends already document the removed scripts).

Also fixed while in this file: an eslint no-unused-vars warning surfaced by
the changed lint run (unused `cleanup` import in
tests/emitted-provenance.test.cjs).

Added regression coverage for both A and B: a cross-platform spot-check that
drives the real hooks-built rule against `.cmd` keys directly (not through a
real Windows install), and a real-tree test that drives buildBaselineAtRef
against a base ref verified (via git cat-file) to lack the generator, both
skipping honestly rather than false-passing when their precondition does not
hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): repair false .cmd byte-provenance and a permanently-skipping regression test

Two isolated-review findings on PR #2767:

- `hooks-built`'s `.cmd` branch attributed the Windows shim's bytes to the
  wrapped `hooks/<name>.js` script, asserting a byte-provenance link that
  does not exist — traced against buildCodexHookWindowsShimIR
  (src/runtime-hooks-surface.cts), only the script's NAME (a literal in that
  same file) flows into the .cmd bytes, never its content. Point `sources`
  at HOOKS_WINDOWS_SHIM_SRC instead, matching the code-derived convention
  used elsewhere in the table. Since `sources` is checked before
  `transforms` in the differential, the wrong mapping silently excused any
  .cmd byte movement caused by editing the wrapped .js file.

- The `buildBaselineAtRef` regression test skipped unless a resolvable base
  ref still lacked scripts/gen-emitted-baseline.cjs — true only until this
  PR merges, after which every base ref carries the file and the test skips
  forever with zero ongoing coverage. Rebuilt hermetically: synthesize the
  missing-generator condition in-place via git plumbing (a throwaway commit,
  child of HEAD, with just that one file removed from a scratch index),
  never touching the real working tree, HEAD, or index, and never depending
  on ambient history or remotes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2724): tolerate the remote runner's dubious-ownership git mount in the emitted baseline path

The runner container mounts the repo at a path owned by a different uid than
the process running the suite, so git's dubious-ownership protection refuses
every git operation there. GitHub Actions never hits this because
actions/checkout registers the workspace as safe automatically; this
runner's container does not.

buildBaselineAtRef is the production build-fallback the sole remaining
emitted gate depends on (resolveBaseline's in-job-build leg), not just a
test helper, so the fix is in the shared git() wrapper (emitted-runtime.cjs)
that every caller — resolveChangedPaths, resolveBase, buildBaselineAtRef's
worktree add/remove/prune, and the hermetic regression test added in the
prior commit — funnels through, plus gen-emitted-baseline.cjs's own
rev-parse (now reusing that same wrapper instead of a second execFileSync,
so the fix has one source of truth). Each call declares -c
safe.directory=<the exact directory it already operates on>, never the *
wildcard.

Audited every other helper on this surface (emitted-diff.cjs,
emitted-baseline.cjs, install-shared.cjs) for the same gap: none of them
shell out to git at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:41:43 -04:00
Tom Boucher
44707c2c5e fix(#2757): attribute transform-driven emitted changes and reclassify agents-verbatim (#2760)
* fix(#2757): let derived provenance rules attribute a transform change

The Phase 2 provenance table (#2722) could only explain a moved emitted
path via its `sources`, so a `kind: 'derived'` artifact whose TRANSFORM
code changed (not its source) was unattributable by construction: 16
emitted agents/*.toml moved by PR #2566's runtime-artifact-conversion.cts
change, with zero agents/*.md in the diff.

Adds an optional per-rule `transforms: string[]`. A moved emitted path
now attributes via its source OR its transform, reusing sourceSatisfiedBy
so exact/prefix matching stays identical for both. Declared narrowly on
agents-toml-derived and agents-verbatim (the two families verified to
route through the same conversion pipeline) as
src/runtime-artifact-conversion.cts, src/install-effort-resolver.cts, and
src/model-catalog.cts — never bin/install.js, which would be a blanket
escape hatch spanning every installer concern.

Also corrects agents-verbatim from kind:'identity' to kind:'derived':
measured against origin/next's own fixtures, the same agents/<name>.md
hashes differently per runtime (e.g. codex vs claude), which a true
verbatim copy cannot do. Verified empirically by installing every
runtime and diffing against the raw repo source — none reproduce it
byte-for-byte; every one rewrites frontmatter and the hardcoded
.claude/ self-reference, and claude additionally gets an effort: line
injected. A new assertNoIdentityTransforms invariant rejects an
identity rule that declares a non-empty transforms list.

Closes #2757

* docs(#2757): flag PR #2566's transform-role files for later verification

src/agent-tools-contract.cts (added by #2566, absent on next) and
src/agent-install-check.cts (modified +112/-1 by #2566, read-only today)
were reviewed for prospective inclusion in AGENT_TRANSFORM_SRCS.

Excluded for now: the nonexistent file would fail this fix's own
"every declared transform path exists" hygiene test, and the modified
file's future role cannot be verified without the unmerged PR's diff.
Documents the reasoning, the interim ack-file safety net, and the
verification method to apply once these files stabilize, so the next
PR touching them has a pre-scoped one-line fix rather than a silent gap.

Related to #2757
2026-07-28 10:47:43 -04:00
Tom Boucher
0f60266042 fix(#2723): reconcile emitted manifest families as a set, not a count (#2750)
* fix(#2723): reconcile emitted manifest families as a set, not a shared count

EXPECTED_MANIFEST_COUNT was a single literal 19 asserted against both the
baseline (built at the base ref) and the current tree (built at PR HEAD).
Those sides legitimately differ by one family whenever a PR adds or removes
a runtime, so no value satisfied both: 19 rejected the current side, 20
rejected the baseline side. Every runtime-adding PR was hard-blocked.

Replace the shared literal with three independent signals - the derived
family set, the recorded fixture set, and the families present at the base
ref - reconciled as sets in both directions. A family may appear or vanish
only when the diff plausibly touches the runtime registry, and the failure
names the family rather than a count. An absolute floor catches the
uniformly shrunken universe a same-count self-check passes vacuously.

Found by tracing #2005 (Qoder runtime) through the gate during the ADR-2719
dual-run window.

* fix(#2723): read the baseline family set from the ref, not HEAD's registry

Review found three defects in the first cut.

Blocker: baselineManifestsAtRef enumerated MANIFEST_FAMILIES, which is imported
at module load and therefore describes PR HEAD. A runtime REMOVED by the PR is
already absent from that list, so the base ref was never asked for it, the
baseline silently omitted a family that genuinely existed, and the dropped-family
check could never fire in production - while its unit tests passed, because they
inject the baseline directly. Enumerate from the ref with git ls-tree instead.

Also: narrow the registry-signal set to the two surfaces that actually define the
family set, since every extra path widens what excuses an unattributed delta; drop
the ack bypass, which was a one-sided escape hatch making removals easier to wave
through than additions; and gate the derived/fixtures inputs so malformed values
return a verdict rather than an unhandled TypeError.

* fix(#2723): filter prototype-shaped family names read from git output

baselineFamilyNamesAtRef derives object keys from git ls-tree output rather
than a trusted constant, so a fixture committed as __proto__.json would turn
the manifests[name] assignment into a prototype write. Compared inline rather
than through a Set, which is the form the prototype-pollution analysis
recognizes.

* fix(#2723): require an exact capability path depth and make the ref test hermetic

The remote runner went red on both linux lanes with two real defects.

The capability signal matched by prefix+suffix, so 'capabilities/capability.json'
(no runtime segment) and 'capabilities/a/b/capability.json' (wrong depth) both
attributed a family change and would have excused an unattributed delta. Anchored
to an exact single-segment pattern.

The ref-derivation test reached for this repo's root commit, which is not stable:
the remote runner shallow-clones, so rev-list --max-parents=0 returns the grafted
boundary carrying every fixture, and this repo has two root commits locally anyway.
It now builds its own git repo containing a family absent from the current registry
- the real discriminator, and one the root-commit version could never assert.

Lint then caught a third: the test called t.after() without declaring t.

* fix(#2723): stop asserting ref enumeration against the ambient checkout

The remote runner returned [] for the repo's own HEAD while every hermetic
temp-repo assertion in the same test passed. That is this function's documented
behavior when git cannot read the ref - the runner works from a shallow clone
under a bind-mounted workdir - so the assertion was testing the checkout rather
than the code.

Dropped it. The temp repo already proves the property that matters, and proves it
more strongly: it contains a family absent from the current registry, which a
registry-derived implementation could never report. The ambient path stays covered
by the real-tree test, which skips explicitly when no base ref is resolvable.

A git failure is not silently permissive downstream: baselineManifestsAtRef returns
null on an empty family set and the real-tree test asserts the baseline is non-empty.
2026-07-28 07:41:13 -04:00
Tom Boucher
9138271b5f test(#2723): differential emitted-attribution check, dual-run beside the golden (#2737)
* test(#2723): differential emitted-attribution check, dual-run beside the golden

Phase 3 of #2719. The conservation law itself, running BESIDE
golden-install-parity.test.cjs -- both green, fixtures untouched.

Every emitted path whose hash moved between next HEAD and PR HEAD must be
attributable, through the Phase 2 table, to a path the PR actually changed.
Unattributable deltas fail with the paths NAMED. The only way through is a
committed acknowledgment, never a flag -- a contributor facing a red gate
sets a flag, which is what UPDATE_GOLDEN=1 is today.

The central decision is that the law is a PURE function (no fs, git,
installer, or clock), with I/O confined to a separate resolver. The naive
one-big-integration-test shape would need ~38 installer spawns per assertion,
so #2723's four failing-first criteria would not in practice have been
written -- which is exactly how a phase ships promised-but-not-built. Pure,
they are millisecond table tests, and the Stryker gate can actually bite.

Buckets are conserved: every moved path lands in exactly one of
attributed | unattributable | acked, property-tested at 400 runs. A path the
provenance table cannot resolve surfaces as an error, never a silent skip.

Asymmetries that are deliberate, each with a test:
- an ADDED emitted key is a ripple too, not just a modified one
- synthesized paths are exempt; code-derived ones are NOT (Phase 2 refused to
  mark them exempt precisely because exempt means permanently blind)
- shrinkage needs no ack; growth does. Gating shrinkage would punish exactly
  what the size ratchet wants
- a STALE ack is a hard failure -- an ack outliving its ripple pre-clears the
  next one on that path
- a failed `git diff` is an explicit error, never an empty changedPaths set;
  reading it as "nothing changed" would make everything unattributable and
  produce a failure storm that reads like a real finding
- prefix sources are SEGMENT-aware, so `agents/` does not attribute
  `agentsfoo/x.md`

Baseline is cached, not committed, keyed on the next sha. A stale key is
refused rather than used: absence fails loudly and gets fixed, whereas
staleness produces a confident wrong answer. An explicitly pointed-at
GSD_EMITTED_BASELINE that is stale is a hard stop; a stale cache falls
through to the in-job build. No baseline-unavailable path returns -- ADR-2719
section 6 names that trap, since in node:test a bare return is a PASS.

Both the conservation property and the staleness gate were mutation-verified
(injecting a swallowed key fails 9 tests; disabling the staleness comparison
fails 5).

Fixtures, generators, the merge driver and the ADR status are untouched --
those are Phase 4 (#2724).

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): compute stale acks once, after the size pass

Self-review defect found while the reviewers were running. `staleAcks` was
computed twice: once between the hash pass and the size pass, then again
after. Only the second value was returned, so the first was dead code -- and
the dead one was placed where it would have been WRONG.

An acknowledgment can be consumed by either a hash move or a size growth.
Computing staleness before the size pass reports a legitimate growth ack as
stale, which is a false failure that pushes a contributor to delete the very
ack that is doing its job.

Now computed once, after both passes, with a regression test. Verified by
mutation: restoring the early computation fails 2 tests.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): wire the attribution check to the real tree, not just synthetic input

An isolated reviewer caught that the first cut was INTERFACE-ONLY: nothing
read the ack file from disk, nothing shelled git, nothing built real
manifests. Every test was true of hand-built inputs and none of the repo, so
the acceptance criterion "both this check and golden-install-parity green on
the same tree" was trivially true rather than meaningfully true. That is the
promised-but-not-built failure this epic keeps finding in its predecessors,
recurring one phase later for the wiring itself. Taken, not argued.

Adds tests/helpers/emitted-runtime.cjs -- the only module that touches git,
disk, or the installer -- and an integration test that runs the same pure law
against reality:

- CURRENT side: 19 real installer spawns via runMinimalInstall +
  buildParityManifest, the same machinery the golden harness uses.
- BASELINE side: `git show origin/next:<fixture>`. That is next's RECORDED
  emitted state and it costs nothing. Deliberately NOT the working-tree
  fixtures, which are whatever this PR's author regenerated -- comparing
  against those would be vacuous. Phase 4 deletes the fixtures and swaps in
  resolveBaseline's cache path, already implemented and tested.
- changed paths from real `git diff --name-only origin/next...HEAD`, with the
  git subprocess bounded at 30s per CLAUDE.md's unbounded-subprocess rule.
- the real tests/emitted-drift-ack.json (absent is legal; present-but-empty
  or unparseable throws rather than being read as absent).

Verified it can actually fail: an uncommitted edit to a shipped workflow
moves emitted output but never appears in the committed diff, and the check
names all 18 affected emitted paths with the message format ADR-2719 §1
specifies. Restores clean.

Also from review:
- readAckFile now has a real test exercising the SUT across absent / valid /
  empty / unparseable / unreadable. The previous test asserted fs behaviour
  rather than SUT behaviour, because no SUT ack-reading path existed yet.
- formatReport's sampleLimit gains true limit-1/limit/limit+1 coverage at
  19/20/21. A test was previously NAMED "(limit+1)" while testing no numeric
  limit at all, which is worse than no coverage because it reads as covered.

Windows uses an explicit t.skip (install output is platform-specific there,
mirroring the golden harness) -- never a bare return, which node:test scores
as a PASS.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2723): cover the claude-local manifest family in the real-tree check

Isolated adversarial review, MAJOR. The real-tree wiring enumerated
Object.keys(RUNTIME_META) -- 18 entries -- while the emitted manifest set has
19 families. The 19th is claude-local: claude is the reference host and the
only runtime with a distinct LOCAL "legacy flat-commands" layout
(commands/gsd-*.md + agents/gsd-*.md at project scope), which
golden-install-parity.test.cjs guards with a hand-coded test outside its
RUNTIME_META loop (#2086).

The family was dropped from BOTH sides, so the test's own self-check
(current.length === baseline.length) passed vacuously at 18 === 18. A PR
changing Claude's local-scope output would have failed the golden while this
check reported ok -- and that disagreement is precisely what the dual-run
window is designed to surface as a provenance-table hole. A wiring omission
masquerading as one is the worst available failure here, because it would
have been read as evidence about Phase 2 rather than a bug in Phase 3.

Fixed by deriving MANIFEST_FAMILIES explicitly (18 global + claude-local at
local scope) instead of inferring the set from RUNTIME_META.

The self-check is also repaired: it now asserts both sides against the
INDEPENDENT EXPECTED_MANIFEST_COUNT from the Phase 2 table, and asserts
claude-local specifically. Comparing the two sides to each other can never
catch a family missing from both -- the assertion has to come from outside.
Verified by mutation: removing claude-local again fails the test.

Also from the same review:
- sourceSatisfiedBy returns the matched source string, so an empty-string
  source would return '' and the caller's `if (hit)` would silently discard a
  real match. Unreachable today (every rule source is a non-empty template)
  but a footgun for the next rule author; now `!== null`.
- the purity fixture used a single-element changedPaths array, so an in-place
  sort would have been invisible. Now three elements in unsorted order.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* fix(#2723): resolve the base ref tolerantly instead of hard-requiring origin/next

The first matrix run failed on both linux lanes:

  differential attribution over the real tree
  cannot resolve origin/next (Command failed: git rev-parse origin/next)

Not a flake, and not an environment excuse -- a real defect in this diff. The
gsd-test runner shallow-clones and merges base+head, so no origin/* remote-
tracking refs exist in the container. My own fail-loud path fired correctly;
what was wrong was hard-depending on that ref existing. GitHub Actions has the
same shape by default, which is exactly why changeset-required.yml carries an
explicit `git fetch origin "${BASE_REF}:refs/remotes/origin/${BASE_REF}"`.

Now resolved through an ordered candidate list -- GSD_EMITTED_BASE (explicit
lane override), then origin/$GITHUB_BASE_REF and $GITHUB_BASE_REF, then
origin/next and next -- de-duplicated, each verified with
`rev-parse --verify <ref>^{commit}`.

When NO candidate resolves the test takes an explicit t.skip() naming every
ref it tried and stating that the gate did not run here. That is the
ADR-2719 section 6 distinction: t.skip is REPORTED as skipped, whereas a bare
return is scored as a PASS. Hard-failing was the other option and is wrong --
it would make the suite permanently red wherever a base ref cannot exist by
construction, which is a statement about the checkout, not a propagation
finding.

The candidate ordering is pinned by a unit test rather than left implicit,
since the ordering IS the fix.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 23:22:37 -04:00
Tom Boucher
1f6822ccba test(#2722): emitted-artifact provenance table with a totality guard (#2735)
* test(#2722): emitted-artifact provenance table with a totality guard

Adds the declarative emitted-path -> source-path table that ADR-2719 §2
specifies, plus the totality guard that keeps it honest. Phase 2 of #2719.

Every emitted path across all 19 committed golden-parity manifests (8,524
paths) must match exactly one rule. Zero matches, two matches, and a rule
matching nothing are all hard failures, so a new emitted family fails the
build loudly instead of passing through unattributed.

The measured surface is larger than #2722 estimated from claude.json alone
(26 top-level families across 19 runtimes, not 13), which is itself what the
totality guard exists to surface. It resolves to 19 rules.

Building the table caught three false attributions that were total but
resolved to repo files that do not exist -- Copilot's `<name>.agent.md`
rename, Kimi's code-literal `agents/gsd.{yaml,md}` root agent, and Copilot's
`hooks/gsd-session.json` registration. The "every attributed source exists"
test is therefore a first-class gate, not a nicety.

Notable correctness decisions:
- Emitted shapes are hard-coded; deriving them from the installer would make
  the guard tautological (it would follow any installer change silently).
  Only source paths read a first-party descriptor, and only where the
  descriptor is the sole declaration (hostBehaviors.nativePlugin.source).
- Emitted skills attribute to commands/gsd/*.md, NOT the repo skills/ dir --
  that directory is generated from commands/gsd by gen-plugin-skills.cjs, so
  attributing to it would be false attribution that still passes totality.
- Attribution is keyed on (rel, runtime): plugins/gsd-core.js has different
  sources for opencode and kilo.
- Rule order carries no semantics (property-tested), since exactly-one
  matching is enforced rather than first-match-wins.

Nothing here reads a git diff, builds a live manifest, or touches a fixture;
the differential check, drift-ack file and size ratchet are #2723.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* docs(#2722): record the delivered provenance table in the CONTEXT.md glossary

The `### Emitted Artifact Provenance` entry landed in #2721 describing the
table as future work. Phase 2 delivers it, so the glossary now records what
actually exists and the invariants #2723 must preserve:

- where the table lives, its rule count, and that it is total over all 8,524
  emitted paths across the 19 manifests
- dead-rule detection, so table rot is loud in both directions
- the corrected surface measurement (26 families, not the 13 estimated from
  claude.json alone)
- the two invariants #2723 inherits: shapes hard-coded (deriving them would
  make the guard tautological), and attribution keyed on (rel, runtime)
- the skills/ false-attribution trap, and that totality does NOT catch a
  wrong-source rule — the source-existence assertion is what does

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

* test(#2722): close five review findings on the provenance table

Two orthogonal review passes plus an isolated adversarial reviewer returned
findings at minor..major. All fixed; no blockers were raised.

Standards axis (CONTEXT.md:456, RULESET.TESTS.guard-toplevel-readFileSync):
- module-level loadManifests() threw at require time before any test()
  registered, turning a missing fixture dir into an opaque crash instead of
  one named failure. Now a memoized lazy accessor.
- matchRules and assertTotality each carried their own copy of the matching
  loop; assertTotality now calls matchRules. That is the #2266 divergence
  class, and two copies could let the guard and the attributor disagree.
- named the corpus stride constant; dropped an inline require.

Spec axis:
- the CONTEXT.md glossary carried a "26 families" figure that is not
  reproducible from the code and that no test pinned -- a hand-maintained
  number in permanent canon, i.e. exactly the silent drift this epic exists
  to end. All volatile counts are now removed from the glossary, with the
  reason stated inline: the guard recomputes them every run, so they belong
  in a failure message, not in prose. No test was added to pin the count,
  because that would rebuild the brittle committed number we are deleting.

Isolated adversarial review:
- `.+` tail captures let a `..` segment reach a constructed source path that
  resolves outside the repo. Not live-exploitable (fixtures are committed and
  the only consumer is an existsSync probe) but Phase 3 feeds these strings
  into a diff-consuming check, so assertSafeRelPath now fails closed once, in
  matchRules, rather than per-rule.
- attributeEmittedPath's ambiguous-match branch was never exercised; only
  assertTotality's parallel path was. Now tested directly.
- sampleLimit's truncation branch had no limit-1/limit/limit+1 coverage.
- the fast-check property could not fail for the reason it was named for.

That last one took two attempts and is the one worth reading. The property
hand-rolled its shuffled side from the per-rule matchOne primitive, which is
order-independent by construction, so it held for reasons unrelated to the
shipped matchRules. Routing it through the real matchRules was still not
enough: on an unambiguous table, first-match-wins and collect-all return
identical results for every path (measured: 0 of 190 corpus paths differ).
Order can only matter where more than one rule matches, so the property now
also asserts that an intentionally ambiguous table reports BOTH hits as a set
under every permutation. Verified by mutation -- injecting a `break` into
matchRules makes it fail, and restoring makes it pass.

Enabling all of the above: matchRules and attributeEmittedPath now take an
injectable rules table, so tests can drive the real code path instead of
re-implementing it by hand.

Refs #2719

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W5kQs6ZufZDySC6zDJfYP6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:51:51 -04:00
Tom Boucher
bf8f320083 feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI)

PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting
GSD's kimi support into two distinct products per the user's directive:
- kimi       (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python)
- kimi-code  (new):      Moonshot's Node Kimi Code CLI (~/.kimi-code,
                         runtime: node, KIMI_CODE_HOME env)

Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json
descriptors, not hardcoded branches in install.js. The new descriptor uses
the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml
hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs).

Critical Kimi Code constraint reflected in the descriptor:
  hostIntegration.dispatch.namedDispatch: false
  hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan']
  hostBehaviors.namedSubagentsSupported: false
Kimi Code's official docs confirm only 3 built-in subagents with NO custom-
subagent registration (the [subagent] table only has timeout_ms). The
kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in
kimi-code's artifactLayout.

Schema adjustments:
- subagentToolkit set to 'undocumented' (the existing escape hatch); the
  schema enum (full/read-only) lacks a 'limited'/'built-in-only' value.
  A follow-up PR can extend the schema enum to add 'built-in-only' as a
  first-class axis value reflecting Kimi Code's documented model.

Registration:
- capabilities/kimi-code/capability.json (new descriptor, modeled on codex)
- bin/install.js: allRuntimes array + --all list + --kimi-code flag
- gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases
  (kimi-code, kimicode, kimi_code)
- src/runtime-name-policy.cts: FALLBACK_ALIASES map
- gsd-core/bin/lib/capability-registry.cjs: regenerated via
  scripts/gen-capability-registry.cjs --write

Tests:
- tests/multi-runtime-select.test.cjs updated for the new runtime count (18)
  + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19.

Out of scope for PR 1 (follow-up PRs in the sequence):
- Install-time decision logic (kimi vs kimi-code detection / prompt)
- agent-install-check semantics for kimi-code (verify Agent Skills presence)
- cmdAgentSkills fallback returning subagent prompt content
- Workflow template mapping (named agents → built-in coder/explore/plan)
- Migration guidance for users currently on 'kimi' who are actually on Kimi Code
- Schema enum extension for subagentToolkit: 'built-in-only'

Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS).

* fix(#2454): complete drift-guard registrations for kimi-code runtime

The drift guards caught every surface that pins runtime enumeration. Each
update is mechanical, driven by the guard's named failure mode:

- src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code
- src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the
  isKimiCode predicate generator
- bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream
  entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19
- gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry
  (null/null/null — same as kimi, no model tier defaults until configured)
- docs/reference/capability-matrix.md: regenerated via
  scripts/gen-capability-matrix.cjs --write (kimi-code row added)
- tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP:
  kimi-code → '.kimi-code'
- tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden
  (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated)

The capability-registry is already regenerated from the prior commit.

* test(#2454): update drift-guard tests for kimi-code runtime registration

Multiple drift guards pin runtime enumeration counts and option numbering.
Each update is mechanical, driven by the guard's named failure mode:

- tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17);
  'all 16 flags' → 'all 17 flags' in test names + messages.
- tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering
  cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15,
  trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice
  test for kimi-code (option 11). Prompt test updated for new numbering.
- tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains
  kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs);
  EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per
  docs, same as Python kimi/opencode).
- tests/global-config-home-fragment.test.cjs: table-count test renamed
  13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier).

* fix(#2454): empty artifactLayout for kimi-code (PR 1 scope)

The skills kind requires a converter (existing converters are per-runtime
like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only
registers the descriptor; the actual Agent Skills converter (and a new
'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside
the install-time decision logic. Empty artifactLayout.global is valid and
means 'nothing to install yet via the layout seam'.

Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs
(localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as
option 11 in install.js's buildRuntimePromptText (renumbered downstream
options 11..17 → 12..18, All 18 → 19).

* fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode)

The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved)
for the new kimi-code runtime id. Property names with hyphens are awkward for
consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new
runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing
the conventional PascalCase flag name. The 16 prior single-word runtime ids
are unaffected (the regex finds no hyphens).

* fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures

- tests/runtime-flags.test.cjs drift guard: use proper kebab-case
  conversion (isKimiCode → kimi-code, not 'kimicode') so the registry
  comparison doesn't false-positive on hyphenated runtime ids.
- tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae
  single-choice tests for the renumbered options (kilo 11→12, opencode
  12→13, pi 13→14, qwen 14→15, trae 15→16).
- tests/install.test.cjs: Kilo integration option 11→12, prompt test
  regex updated.
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1.
  The kimi-code install produces the standard GSD install layout (skills,
  contexts, references, etc.) — 436 paths, same shape as other runtimes
  that have no custom converter yet.

* fix(#2454): add kimi-code install contract + global config home fragment

- src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code
  → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path
  instead of falling through to the default '.claude'.
- tests/installer-migration-install.integration.test.cjs
  RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for
  PR 1; PR 2 will specialize once the Agent Skills converter lands).
- tests/multi-runtime-select.test.cjs: fix space-separated-choices test
  for the renumbered kilo option (11 → 12).
- tests/fixtures/golden-install-parity/kimi-code.json + install-tree/
  kimi-code.json: regenerated after rebasing onto current next (new
  planner-reversibility.md from #2471 etc. now included).

* test(#2454): skip kimi-code install contract until PR 2 ships install layout

The end-to-end install test (tests/installer-migration-install.integration
.test.cjs) asserts every allRuntimes entry installs a runtime-specific
artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the
capability descriptor + flags + labels, but the install LAYOUT (Agent
Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands
in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and
self-removing — PR 2 removes the entry alongside adding the install
surface, restoring the contract loop to full coverage.

* fix(#2454): restore compact model-catalog.json format (M1 review)

Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations'
commit used python json.dump(indent=2) which inflated the file from 165→607
lines (every nested entry got expanded) and lost the trailing newline. The
semantic change was just a 3-line kimi-code entry. Restored the original
hybrid format (top-level indent=2 + inner entries' one-line style) and
added kimi-code in matching form.

Regenerated golden install parity + install tree fixtures since the
model-catalog.json hash changed.

* fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code)

CI lint-tests job failed on the glossary drift guard
(scripts/check-glossary-refs.cjs --check):
  ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but
    bin/install.js's allRuntimes array has 18.
  ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js
    (missing from CONTEXT.md's list: kimi-code).

Missed in the prior commits because gsd-test does not run the glossary
check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims
to 18 values + kimi-code in the member list.

* chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511)

* docs(changeset): Phase 1 kimi-code runtime Added (#2511)

* test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands

* docs(changeset): backfill PR #2519 for Phase 1 (#2511)
2026-07-22 00:27:35 -04:00
Tom Boucher
89b1bef881 refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267.
2026-07-14 19:22:13 -04:00
Tom Boucher
6a474db3aa refactor(#2266): single-source golden-parity manifest builder + fixture correction (#2273)
Phase 1 of golden-install-parity redesign (epic #2264). Consolidates buildParityManifest + exclusion constants into tests/helpers/install-shared.cjs (fixes realRoot divergence), adds anti-divergence guard, corrects 12 stale golden fixtures to portable values. Closes #2266.
2026-07-14 17:10:04 -04:00
Tom Boucher
51b6e3c35c test(#2136): add localToday regression + mirror fake/inline clock stubs
Adds a dedicated regression (tests/fix-2136-clock-local-today.test.cjs):
- realClock.localToday() returns the LOCAL calendar day under a pinned instant
  + TZ (America/Chicago → 2020-06-14, the issue's repro instant).
- realClock.today() is unchanged (UTC) — internal/cosmetic stamps stay UTC.
- localToday === today on a UTC host (no spurious divergence).
- makeFakeClock mirrors localToday (drop-in Clock substitute).
- field-level: a 'sync' transition with a split clock (today≠localToday)
  writes the LOCAL day into Last Activity, proving the wiring uses localToday.

Mirrors localToday() in the fake clock helper and the inline fixedClock stubs
in state-transition.test.cjs / state-rebuild.test.cjs (the Clock interface now
requires it).
2026-07-12 15:34:59 -04:00
Tom Boucher
79d7657eff feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.

Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
  slash-programmatic / active-model / native-extension / bun) + hostBehaviors
  {nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
  RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
  payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
  by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
  EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
  surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
  from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
  (global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
  markdown); the 16 other fixtures + claude-local change only by the shared
  model-catalog hash line.

Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
  subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
  in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
  (every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
  gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
  a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
  shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
  query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
  consumes params; getArgumentCompletions; before_provider_request active-model
  steering (fail-open on null resolution); functional session_start /
  before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
  hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
  vocabulary.

Docs (host-integration matrix + how-to) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 21:07:38 -04:00
Tom Boucher
6e773d97df feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration
Interface (declarative embedding adapter → engine surface dispatch) and fold
every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven
`runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated
(tests/fixtures/golden-install-parity/codex.json); no other runtime changes.

Three Context7-verified upgrades, each with a test on the user-reachable surface:
- Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override,
  with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install
  and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest,
  and the skill-manifest inventory to honor the override so --skills-root /
  sync-skills / the manifest report the real location.
- Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact,
  PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall;
  extendedHookEvents reconciled [] -> the schema-valid wired subset.
- Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the
  negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a
  known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and
  unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own
  AgentsToml scalars (max_threads etc.) instead of dropping them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 21:42:11 -04:00
Tom Boucher
69b309e4e0 feat(#1925): add ZCode (Z.ai) as a pluggable runtime descriptor
Add ZCode as a first-party runtime via a declarative capability descriptor
(capabilities/zcode/capability.json) with zero hardcoded runtime === 'zcode'
branches — exercising the de-hardcoded, data-driven runtime path that 1.7.0
(ADR-1016 / ADR-1239) enables.

Descriptor (all axes sourced verbatim from zcode.z.ai docs):
- configHome ~/.zcode; nested skills + flat commands/agents; profile-marker install
- Claude-shaped skill format reuses convertClaudeCommandToClaudeSkill (no new converter)
- hostIntegration: declarative / slash-file / mcp / electron; dispatch background=false
  (foreground-only per docs); nested+maxDepth undocumented; passive model mode

Installer registration (data, not branches): --zcode flag, allRuntimes, runtimeMap,
interactive menu, --all list. getGlobalConfigDir + resolveRuntimeArtifactLayout +
ALLOWED_CONFIG_RUNTIMES + resolveInstallPlan all derive from the descriptor.

Revamped the brittle per-runtime golden-master tests to be count-agnostic,
descriptor-derived property tests (1.7.0 makes runtimes pluggable data, so pinning
frozen '15 runtime' snapshots is the wrong invariant): getdirname, label-policy,
config-adapter-registry (intent + install-plan golden master), capability-registry,
host-integration-descriptors (counts derive from curated maps). Adding a runtime
descriptor now extends coverage with zero edits to those suites.

Welcome banner, --zcode help, supported-runtimes how-to, and the host-integration
capability matrix (every axis cited) updated.
2026-07-06 08:31:51 -04:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Behruz Nassre Esfahani
fc5ca178a2 test(#1178): consolidate duplicated agent-roster helper into tests/helpers (#1420)
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers

The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated
across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and
inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting
listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites
through it. Semantically-different sites (installed-dest dirs, absolute-path
returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform
multi-family inventory table) are left intact, each with a one-line comment.

Test-only; no production code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1178): note AGENTS_DIR export is for future call sites

Review nit: clarify that the currently-unused AGENTS_DIR export is intentional
— available for future tests needing the canonical source agents path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-23 19:34:08 -04:00
Tom Boucher
fc2a7c0555 fix(#1615): install Windsurf slash workflows 2026-06-23 12:10:21 -04:00
Tom Boucher
33ccf5f89d fix(#1367): project-local install uses flat gsd-<cmd>.md layout (fixes /gsd: colon namespace) (#1489)
* fix(#1367): project-local install uses flat gsd-<cmd>.md layout

Claude Code project-local installs now write command files as flat
gsd-<cmd>.md at .claude/commands/ level instead of commands/gsd/<cmd>.md
(subdirectory), so Claude Code registers /gsd-<cmd> (hyphen form)
matching hooks, statusline, and all cross-command references.

- capabilities/claude/capability.json: local destSubpath commands/gsd → commands
- bin/install.js else branch: flat gsd-<stem>.md loop with runtime rewrites
- bin/install.js uninstall (1c): remove flat files + legacy subdir cleanup
- bin/install.js writeManifest: record flat commands/gsd-<cmd>.md keys
- legacy migration: preserves dev-preferences.md across reinstall and uninstall
- gsd-core/bin/lib/capability-registry.cjs: regenerated
- 6 new regression tests (L0–L5) in bug-1367-*.test.cjs
- Updated E suite in bug-3683 + bug-1736, layout + surface + descriptor tests
- scripts/lint-regression-test-names.allowlist.json: grandfathered bug-1367 test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1367): add issue reference to allow-test-rule comment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 13:37:27 -04:00
Tom Boucher
9e5d4b266b fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs (#1207)
* fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs via SessionStart hook

Claude Code marketplace plugin installs unpack the package into the
version-pinned plugin cache and never run bin/install.js, so
~/.claude/gsd-core/ is never created. Agents, commands, and templates
markdown-@-include the canonical ~/.claude/gsd-core/... path (which
expands ~ but NOT ${CLAUDE_PLUGIN_ROOT}), so every include resolved to
nothing and agents (e.g. the executor) failed.

Add a SessionStart hook (hooks/gsd-ensure-canonical-path.js) that, on a
plugin install, symlinks the canonical path's immutable subdirs (bin,
contexts, references, templates, workflows) to the plugin's bundled
gsd-core/ tree. It changes zero @-references, is a no-op in classic
installs, preserves user-generated files (USER-PROFILE.md, STATE.md),
prunes stale links so it self-heals after `claude plugin update`, uses
Windows junctions, and rejects bundled/canonical paths that escape the
resolved plugin root (no traversal, no clobber).

Registered in HOOKS_TO_COPY (build-hooks), MANAGED_HOOKS, hooks.json
SessionStart (runs first, timeout 5), and BUNDLED_GSD_HOOK_FILES.
Behavioral regression tests folded into issue-766-plugin-manifest.test.cjs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#997): backfill changeset PR number to #1207

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 09:37:41 -04:00
Tom Boucher
b4a7eabaae feat(#1085): migrate windsurf workspace skills to .devin/ + fix global content refs (#1093)
Fresh windsurf/devin-desktop workspace installs write skills under .devin/ (legacy .windsurf/ recognized); global ~/.codeium/windsurf/ unchanged. Also threads real isGlobal through _applyRuntimeRewrites so global skill content references the codeium path. Closes #1085.
2026-06-11 23:17:01 -04:00
Tom Boucher
77a671ec53 feat(#791): migrate antigravity workspace base dir .agent → .agents (#1090)
Fresh antigravity workspace installs write under the canonical .agents/ (plural) base; legacy .agent/ stays recognized (dual-read). Global ~/.gemini/antigravity/ path unchanged. Closes #791.
2026-06-11 22:37:55 -04:00
Viktorplus
f2dd125593 Merge branch 'next' into kimi-runtime-support 2026-06-08 22:11:07 +02:00
Tom Boucher
b6199460ed refactor(#887): consolidate duplicated CHILD_ROUTER test maps into shared helper (#889)
Three test files copy-pasted the concrete-skill->namespace-router CHILD_ROUTER
map verbatim, and two more duplicated an identical local parseRouterRequires
regex. Introduce tests/helpers/nested-layout.cjs that derives the child->router
map from the authoritative commands/gsd/ns-*.md requires: lists once, reusing
the production parseRequires (now exported from install-profiles) plus a
nestedSkillPath(skillsRoot, prefix, stem) helper. Refactor all five test files
to import from it.

Test-only + a single internal export addition; no production behavior change.

Closes #887

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 16:05:59 -04:00
Viktorplus
54191f6c6c Merge branch 'next' into kimi-runtime-support 2026-06-08 03:57:58 +02:00
Tom Boucher
1b6bd66f2c feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) (#821)
* feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged)

Wire three new context-tracking events (SubagentStop, Stop, PreCompact) to
gsd-context-monitor so context-headroom warnings surface at model-stop and
subagent-finalisation moments — not just on PostToolUse.  Add a new
FileChanged hook (gsd-config-reload.js) that hot-reloads .planning/config.json
context mid-session when the user edits it, injecting a config summary as
hookSpecificOutput.additionalContext.  Updates plugin manifest hooks.json,
managed-hooks-registry, installer-migration-report allowlist, and
shell-command-projection cleanup tables.  Tests: 21 new assertions in
enh-770-claude-hook-events.test.cjs; enh-788 and issue-766 test suites updated.

Closes #770

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#770): document newly-registered Claude Code lifecycle hooks

Add a Hook coverage table to the Claude Code npm installer section of
docs/how-to/install-on-your-runtime.md describing SubagentStop, Stop,
PreCompact, and the new FileChanged (gsd-config-reload.js) hook that
hot-reloads .planning/config.json mid-session. Also fixes the changeset
frontmatter (adds type: Added + pr: 821) so docs-lint can consume the
fragment.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): add gsd-config-reload.js to INVENTORY.md and regenerate manifest

The feat commit added hooks/gsd-config-reload.js but did not bump the
Hooks count in docs/INVENTORY.md (14→15) or add the new row, and did not
regenerate docs/INVENTORY-MANIFEST.json. Both inventory-counts and
inventory-manifest-sync tests failed across the full CI matrix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): make lifecycle-hook tests deterministic on scoped runner

Replace the shared hooks/dist/ ensemble setup (ensureHooksDist /
teardownHooksDist) in the Claude hook tests with per-test isolation:
pre-populate each test's own tmpDir/.claude/hooks/ with stub files and
pass installerMigrations:[] to install() so the first-time-baseline
migration does not remove the stubs before the copy step can run.

Root cause: hooks/dist/ is gitignored and absent on a fresh npm ci.
ensureHooksDist() created it and teardownHooksDist() deleted it, but
with --test-concurrency=4 both test files ran concurrently as separate
Node.js worker processes sharing the same filesystem.  One file's
afterEach teardown deleted hooks/dist/ while the other file's install()
was copying from it, producing an ENOENT (reproduced 2/10 runs locally).

The additional issue: even with pre-placed stubs surviving the copy race,
the 000-first-time-baseline migration classified hooks/gsd-*.js as
bundled-gsd-hook artifacts, auto-removed them, and the copy step never
re-ran (hooks/dist/ absent) — leaving contextMonitorFile missing and all
hook registrations silently skipped (the 'got: []' symptom).

Fix: pre-populate targetDir/hooks/ per-test (isolated temp dir) AND pass
installerMigrations:[] so the baseline scan is skipped.  The Qwen suites
already used this pattern correctly; the Claude suites are aligned to it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): ship gsd-config-reload.js by adding it to build-hooks HOOKS_TO_COPY

The #770 feature added hooks/gsd-config-reload.js and registered it in
MANAGED_HOOKS, the installer, INVENTORY, and the test EXPECTED_ALL_HOOKS
list — but never added it to scripts/build-hooks.js HOOKS_TO_COPY. As a
result the hook was never copied into hooks/dist/ during the build, so:

  - the hook would never ship to users (real production bug — the
    FileChanged config-reload feature was dead-on-arrival), and
  - install-minimal-hooks.test.cjs #1755 ("all expected hooks are copied
    from hooks/dist/ to target", ".js hooks are executable after copy",
    "manifest contains .js hook entries") failed on any environment with
    a clean checkout (no pre-existing hooks/dist/): coverage, full test
    macos-22/macos-24, test ubuntu-24.

The failures were masked locally only by a stale hooks/dist/ left from a
prior build (build-hooks copies into dist without clearing it). On CI's
fresh `npm ci` there is no dist, so the omission surfaced.

Fix: add 'gsd-config-reload.js' to HOOKS_TO_COPY so build-hooks stages it
into hooks/dist/ alongside the other JS hooks. Verified by removing
hooks/dist/ and rerunning the full suite green (0 fail).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#770): make config prototype-pollution beforeEach deterministic on scoped runner

Root cause: the #663 and alert-#26 prototype-pollution describe blocks
seeded .planning/config.json in beforeEach via a bare
runGsdTools('config-ensure-section') whose result was discarded. That
command runs in a spawned gsd-tools child; on the scoped CI lane
(--test-concurrency=4, config.test.cjs scheduled alongside the heavy
install/tarball suites that #770 pulled into the targeted set) the child
can be transiently killed under resource pressure (non-zero exit, empty
stderr — an OS-level kill, not an app error). The swallowed failure left
config.json absent, so the first subtest's readConfig() threw ENOENT
opening <tmp>/.planning/config.json. Only 1 of 4 subtests failed,
confirming a per-invocation transient, not a deterministic miss; the full
suite schedules files differently so config.test.cjs did not collide with
those heavy neighbors → passed there.

Fix: add ensureConfigReady(tmpDir) which retries config-ensure-section on
ANY failure or missing file and throws a clear diagnostic if it still
cannot create config.json, then use it in both prototype-pollution
beforeEach blocks. Setup is now deterministic under load; the #663/alert-#26
security assertions are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 20:36:11 -04:00