Commit Graph

5071 Commits

Author SHA1 Message Date
0xdhx
104fc76f70 fix(#2665): watch the hook bundle and the install markers the census found
Self-found by re-deriving the guard-shape census against bin/install.js's own
write sites, not by a review finding. Three artifacts a global install writes
into a live config ROOT were watched by nothing:

  hooks/gsd-check-update.js, hooks/gsd-context-monitor.js,
  hooks/gsd-update-banner.js   -- `hooks` was absent from GSD_PREFIXED_PARENTS
  .gsd-source, .gsd-profile    -- absent from GSD_OWNED_ENTRIES, and an
                                  exact-name list does not match a dot-prefixed
                                  name via the `gsd-` prefix rule

This is the SAME shape as the leak that motivated the prefixed-parent scan in
round 1 -- a gsd-prefixed child under a parent nobody had listed -- one parent
over. That it recurred is the argument for re-deriving this list from the
installer each round instead of trusting it: the enumeration is the weak point
of an enumerate-and-block mechanism, and it does not announce when it falls
behind.

Ownership is unchanged, only coverage: `hooks/` is shared with the host agent,
so only `gsd-`-prefixed children are watched. A test asserts a host-owned
hook is still ignored, because widening the parent list must not widen
ownership -- a guard that flags the host's own files gets switched off, and
then catches nothing at all.

Reverting the widening fails the new test.
2026-08-08 05:50:49 -05:00
0xdhx
e31f706ceb docs(#2665): document the two live-config-guard env vars
The changeset for this PR is typed `Added`, and CONTRIBUTING requires a docs/
change for that type. The only docs/ file in the diff was CONTEXT-INDEX.json --
a GENERATED index -- so the Docs Required gate passed while no human-readable
documentation existed for either new variable. A gate satisfied by a generated
artifact is satisfied vacuously.

docs/TESTING-SUITES.md now carries a section on the guard: what it watches and
why it is ownership-scoped rather than whole-root, the two env vars in a table,
why the default is report-only and what the promotion condition is, and what
each violation label means (including that UNVERIFIED is not clean).

GSD_SKIP_LIVE_CONFIG_GUARD is named explicitly because it is a bypass on a
safety check. An undocumented bypass is one people eventually set without
knowing what they turned off.

A test asserts both variables appear in that doc -- checked as permitted by
local/no-source-grep before writing it, rather than assumed forbidden. It fails
when the section is removed, so the doc cannot rot back to the state the review
found.

Addresses review finding: Major 4.
2026-08-08 05:50:49 -05:00
0xdhx
f0ef9063d5 test(#2665): restore the three agent-skills tests this PR deleted
Commit 2bed9fd8 ("replace the hand-synced TEST_ENV_BASE copies with the
canonical import") also removed markLocalGsdInstall and three behavioural tests
from tests/agent-skills.test.cjs -- 79 lines, no replacement, and no mention in
the commit message or the PR body:

  - unconfigured Codex reads its local companion agent from a descendant cwd
  - workstream runtime selects the local Codex companion when root config differs
  - unconfigured Claude remains empty when a local Codex companion exists

RULESET.TESTS.delete-bad-tests permits deleting a bad test only when it is
replaced with compliant tests in the same PR. Nothing was replaced, and these
were not bad tests -- they were collateral in a mechanical edit. A silent net
loss of behavioural coverage inside a PR whose subject is test hygiene is the
one thing that should not pass here, and the reviewer was right to block on it.

Restored verbatim. They need no adaptation to the canonical TEST_ENV_BASE
import: they pass their env explicitly, and an explicit env still spreads last
over the base. The file goes 80 -> 83 tests, all green.

Non-vacuity checked rather than assumed: dropping the local-install marker the
first two depend on fails both. The third is a negative assertion and correctly
stays green, which is why it is named here rather than counted as covered.

Addresses review finding: Blocker 1.
2026-08-08 05:50:49 -05:00
0xdhx
e4f79c32b0 fix(#2665): wire the fourth suite lane, and derive the lane list instead of naming it
qa-loop-walk runs `npm run test:qa`, which is `run-tests.cjs --suite qa` -- so it
runs the live-config guard like every other suite lane, and it set no
GSD_STRICT_LIVE_CONFIG_GUARD. A leak of exactly the class this PR closes would
have printed a warning there and left the lane green.

The test that is supposed to prove the guard is wired everywhere could not
detect that, because its job list was three literals (`test`, `test-full`,
`test-inert`). A hand-list certifying its own completeness is the defect this
whole PR is about, reproduced inside the test guarding the fix -- so the list is
now DERIVED from the workflow: every job with a step reaching run-tests.cjs,
directly or through an npm script resolved transitively through package.json.
The indirection is the load-bearing half; a grep for the filename alone is what
made qa-loop-walk invisible.

The derivation asserts a floor (>= 4 jobs) before ruling on any of them, so a
selector that silently matched nothing fails loudly instead of passing
vacuously. Windows lanes keep their carve-out, keyed on whether the job's matrix
mentions windows rather than on the job's name.

Negative-controlled: un-wiring qa-loop-walk fails the new test. The literal
version passed with that lane unwired, which is how it shipped.

Addresses review finding: Major 5.
2026-08-08 05:50:49 -05:00
0xdhx
4bc6b0a2a3 fix(#2665): defer the built-lib require so an unbuilt tree fails one test, not all
tests/helpers.cjs required gsd-core/bin/lib at module scope to derive the
config-location scrub set. That lib is BUILT, so on an unbuilt tree the require
threw inside `require('./helpers.cjs')` -- before a single test() had registered
-- turning one missing `npm run build:lib` into a whole-suite crash with no
message naming the remedy. This is the file ~370 test files import, so the blast
radius is the suite. `npm test` builds via its pretest hook; the shape that
reaches this is a direct `node --test` invocation, which is exactly what a
contributor reaches for when running one file.

The require is now memoized behind builtLib(), and the two derived exports
(TEST_ENV_BASE, CONFIG_LOCATION_ENV_KEYS) are enumerable lazy getters, so
destructuring and Object.keys() behave as before. Reading either is what forces
the build; a test file that needs neither now imports cleanly. When the build IS
missing, the error names `npm run build:lib` instead of surfacing a bare
MODULE_NOT_FOUND.

Verified by a cold-child probe rather than by inspection -- this process has
already loaded everything, so an in-process assertion would pass vacuously. The
probe checks require.cache before and after touching TEST_ENV_BASE, and fails
when the require is moved back to module scope.

Addresses review finding: Major 7.
2026-08-08 05:50:49 -05:00
0xdhx
4eb29b9741 test(#2665): pin the MAX_DEPTH boundary on both sides, not just above it
RULESET.TESTS.boundary-coverage asks for {limit-1, limit, limit+1}. The depth
bound was exercised only at limit+2, which pins neither side of the edge: an
off-by-one that truncated a tree sitting exactly AT MAX_DEPTH would have passed,
and a truncation is not a cosmetic miss here -- it reports `unverified`, which
in strict mode fails the run.

Negative-controlled by weakening the guard to `depth >= MAX_DEPTH`: the new
limit case fails, where the previous single limit+2 assertion did not.

Addresses review finding: Minor 8.
2026-08-08 05:50:49 -05:00
0xdhx
f054c85fb0 fix(#2665): budget the scan per target, so order stops deciding the verdict
MAX_ENTRIES was a single running budget threaded across every watch target. One
large early target exhausted it, and every target scanned afterwards reported
truncated -> `unverified` -- which under GSD_STRICT_LIVE_CONFIG_GUARD=1 is a
failed run. The guard's verdict therefore depended on directory iteration order
and on unrelated local state, neither of which says anything about whether the
suite leaked.

Each target now draws a fresh allotment, so a pathological tree truncates itself
and nothing else. MAX_TOTAL_ENTRIES keeps the aggregate bounded -- which is what
the single budget was actually for -- and when that ceiling engages, the targets
it curtails are still reported `unverified` rather than attested clean.

The limits are injectable so the boundary is testable without materialising
20000 entries, matching the `deps` seam the resolvers already use.

Two of the three new tests fail when the shared budget is restored; the third
asserts the retained global ceiling, which is deliberately unchanged behaviour.

Addresses review finding: Major 6.
2026-08-08 05:50:49 -05:00
0xdhx
e82a15a852 fix(#2665): watch the fallback root a scrubbing child actually resolves to
resolveLiveConfigRoots resolves what THIS process sees, and getGlobalConfigDir
is env-first -- so with an ambient CLAUDE_CONFIG_DIR the guard watched that
path. A spawned child does not see it: TEST_ENV_BASE blanks the config-location
vars precisely so the child cannot follow them, and a blanked var is falsy, so
the child resolves its HOME-derived root instead.

A child that blanks the var and does NOT also sandbox HOME therefore writes into
the developer's real ~/.claude, which the guard was not watching. That is this
PR's own escape route, taken one process deeper -- and the guard is the artifact
that is supposed to make it loud.

Both resolutions are now unioned: the ambient one, and the fallback one obtained
by handing the REAL descriptor resolver an EMPTY env. Deriving it that way is
deliberate -- a hand-listed copy of the scrub set inside the guard is a second
list to drift, which is the defect this PR spent three rounds closing one layer
up. grok resolves through a hardcoded branch rather than a descriptor, so its
fallback is stated explicitly for the same reason it is named in the ambient
loop.

Addresses review finding: Blocker 3.
2026-08-08 05:50:49 -05:00
0xdhx
6a1fbf96fd fix(#2665): let the guard see deletions, in both shapes it can take
diffLiveConfig walked `after` alone, so it had no branch for a path that
existed before the run and does not after. A test run that DELETES a file from
the developer's real config dir passed the guard silently -- the least
recoverable case in the threat model this guard exists to cover.

The review named the missing `pre.exists && !post.exists` branch. That branch is
necessary and not sufficient: deletion arrives in two shapes and it reaches only
one of them.

  - A FIXED owned entry (GSD_OWNED_ENTRIES x roots, plus every extra target) is
    recorded at both ends whether it exists or not, so a deletion reads
    {exists:true} -> {exists:false}. This is the shape the named branch fixes.
  - A gsd-prefixed child is DISCOVERED by readdirSync, so a deleted one is
    absent from `after` entirely and never enters an after-keyed loop at all.
    The named branch is unreachable for it.

So the walk is now over the UNION of both key sets, with the explicit branch for
the first shape and an `!post` branch for the second. Both are covered by a
test, and reverting the fix fails both -- the prefixed-child test is the one
that would still fail with only the prescribed branch in place.

Addresses review finding: Blocker 2.
2026-08-08 05:50:49 -05:00
0xdhx
ecea537194 docs(#2665): the guard watches config.toml but GSD also writes <root>/hooks/ there
Found pre-push by this round's third adversarial review pass. Not a rebase
regression — round 3 shipped it and #2755 doubled it.

resolveExtraWatchTargets watches one config.toml per non-registry descriptor,
and its comment asserted "GSD writes ONE named file into these third-party
roots". That is false: bin/install.js also calls installSharedHooksBundle on the
same root, populating <root>/hooks/ with GSD's hook scripts and a CommonJS
marker. So a suite-produced leak of a hook bundle into a developer's real
~/.kimi or ~/.kimi-code passes this guard silently — #2665's own hazard, in
#2665's own safety net.

Behaviour is deliberately unchanged and the gap is disclosed instead. Closing it
is a layout decision rather than one more path, for the same reason
getGlobalSkillsBase is already a deliberate non-target: the snapshot applies the
config-root layout beneath every root it is given, and these roots are not ours.
Happy to fix it here or take it as a separate issue — the maintainer's call.

The enumeration-relative test could not have caught this: it asserts one target
PER DESCRIPTOR and nothing about whether one per descriptor is enough, because
its expectation is derived from the same array it checks. That is exactly the
scope boundary round-2 Nit 7 asked to be marked, biting one layer up from where
it was marked; the test now says so.

479979c4's message says "there are three" — that is three WATCHED targets, not a
count of write surfaces. The hooks bundle is a fourth, and unwatched.

lint:ci rc=0; tests/live-config-guard.test.cjs 24/24. Comments and catalog only.
2026-08-08 05:50:49 -05:00
0xdhx
12cfd27f53 docs(#2665): the guard's own comments still described one kimi home, not two
Same drift as the CONTEXT.md seams, one layer over: #2755 took
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS from one entry to two, and five comments
across three files were left describing the one-entry world — "two live write
surfaces", "today's only entry", "today's single entry", and a <kimi>/config.toml
bullet naming only Kimi CLI's KIMI_SHARE_DIR.

The sharpest one was a wrong pointer rather than a stale count: run-tests.cjs
cited "scripts/lib/live-config-guard.cjs" for why the scope is narrow. That path
does not exist, and it names the one directory this module is deliberately NOT
in — the installer copies scripts/lib/ to users wholesale while uninstall removes
only an allowlist, which is the whole reason the guard lives one level up. A
reader following that pointer would have concluded the opposite of the decision.

Comments only; no behaviour change. lint:ci rc=0, tests/live-config-guard.test.cjs
24/24, tests/run-tests-harness.test.cjs 138/138.
2026-08-08 05:50:49 -05:00
0xdhx
664bab3b48 docs(#2665): the scrub-set and guard-target seams understated their own mechanism
Both predicates were written in round 3 and not updated when round 4 widened
what they describe, so the catalog that exists to stop a future omission had
two of its own.

CONFIG.LOCATION.SEAM.scrub-set said "four sources" and listed four. There are
five rungs: the two descriptor rungs each additionally walk skillsHome.env (the
maintainer's round-4 Minor 3), and the fifth is WRITE_ESCAPE_PERMISSION_ENV_KEYS,
which is a permission rather than a location and so is reachable by no other rung.

LIVE-CONFIG.GUARD.SEAM.non-root-targets said "the two live write surfaces" and
named a single config.toml via resolveKimiHooksTomlDir. Since #2755 landed
KIMI_CODE_HOOKS_TOML_DESCRIPTOR there are three, and the guard derives them by
iterating NON_REGISTRY_CONFIG_HOME_DESCRIPTORS rather than calling a named
resolver. Both bounds are stated rather than left open: skills bases are a
deliberate non-target (the config-root layout misfires beneath them), and a
further descriptor is only free if it owns the same NON_REGISTRY_OWNED_FILE —
the residual the guard already names at its own definition.

LIVE-CONFIG.GUARD.SEAM.scope's "gsd--prefixed" read as a two-hyphen prefix; the
selector is startsWith(GSD_ARTIFACT_PREFIX) where that constant is 'gsd-'.

CONFIG.LOCATION.SEAM.kimi-two-homes was checked and is NOT stale: #2755 made
Kimi Code a separate runtime, so Kimi CLI still declares exactly two.

Both CONTEXT-INDEX mirrors regenerated; predicate ids unchanged (425), only
their descriptions move.
2026-08-08 05:50:49 -05:00
0xdhx
d43f781530 chore(#2665): regenerate the example CONTEXT-INDEX mirror after the rebase
The rebase onto next conflicted in this generated file; per the
generated-file discipline the conflict was resolved arbitrarily and the
artifact regenerated from its producer rather than hand-merged.

Regen delta against the base's committed copy is exactly this PR's own
10 predicates (CONFIG.LOCATION.SEAM.* x4, LIVE-CONFIG.GUARD.SEAM.* x6),
415 -> 425 keys, zero removed and zero foreign — so the merge introduced
no drift the PR does not own.
2026-08-08 05:50:49 -05:00
0xdhx
c95b817cb6 fix(#2665): scrub GSD_ALLOW_SYMLINKED_DEST — a write-escape permission
Found by the pre-publication claim audit of this round's response comment,
which refuted the sentence "none of the unscrubbed env reads names a write
destination" on the grounds that naming a path is not the same property as
influencing where writes land.

GSD_ALLOW_SYMLINKED_DEST is boolean and names no path, so every rung of the
derivation is structurally incapable of reaching it: not a registry configHome,
not descriptor-shaped, not one of GSD's own location vars. It is still a #2665
leak vector. install-engine.cts reads it env-first (:214) and threads it as
allowOptInFollow into the symlink-escape guard at four call sites, each gating
a write (:361/:367, :416/:424, :785/:790, :927/:932). That guard is what stops
a write leaving the install root, so an ambient =1 disarms it for the whole
suite — the #2665 hazard arriving through a permission rather than a path.

Added as its own named family (WRITE_ESCAPE_PERMISSION_ENV_KEYS) rather than
folded into a location rung, for the same reason GSD_HOME got its own family in
round 3: the list should not misdescribe what its members are.

Blanking is fail-safe in the only direction that matters — '' is neither '1'
nor 'true', so a blanked value makes the guard stricter, never looser. That
asymmetry is what licenses scrubbing it wholesale rather than reasoning about
each call site.

The guard test names the variable literally rather than iterating the family
constant: a test that asserts over the constant shrinks its own expectation
when the family is emptied, which is the enumeration-relative failure that let
the kimi-code descriptor go unwatched earlier in this same round.

#2393's opt-in suite is unaffected (99/100, 0 fail): it sets process.env
directly in-process and never routes through scrubConfigLocationEnv, and an
explicit env argument still wins over TEST_ENV_BASE in childEnv.
2026-08-08 05:50:49 -05:00
0xdhx
2a98b6b0b1 fix(#2665): keep the dot-home type guarantee #2755 chose
The rebase resolution annotated both hoisted descriptors and the local in
resolveKimiHooksTomlDir as ConfigHomeDescriptor. #2755 used DotHomeDescriptor
there deliberately: the union permits xdg / dot-home-nested / generic-agents-root
shapes, so the broader annotation drops the compile-time guarantee that this
resolver selects a dot-home descriptor and nothing else.

Runtime behaviour was already correct — all eleven paths resolve identically —
so this restores a type-level property, not a behavioural one. Both exported
constants are now pinned to DotHomeDescriptor as well, which is stricter than
the ConfigHomeDescriptor they carried since round 3; the interface stays
unexported and NON_REGISTRY_CONFIG_HOME_DESCRIPTORS keeps its
ConfigHomeDescriptor[] type, which the narrower constants satisfy as subtypes.

Found by the pre-push adversarial review of this round, which is the one claim
of six it refuted.
2026-08-08 05:50:49 -05:00
0xdhx
d1c8b32689 fix(#2665): watch kimi-code's config.toml, and pin it by name
The rebase onto next brought in #2755, which added a SECOND Kimi config
home — kimi-code's `~/.kimi-code`, overridden by KIMI_CODE_HOME — declared
as an inline object literal inside resolveKimiHooksTomlDir's body. That is
the resolvable-but-not-enumerable shape round 3 hoisted KIMI_SHARE_DIR out
of, so the hoist is extended to cover both descriptors rather than reverting
#2755's parameterization.

The scrub set was already complete: KIMI_CODE_HOME is declared in
capabilities/kimi-code/capability.json, so the registry rung covered it and
CONFIG_LOCATION_ENV_KEYS is 28 keys both before and after the rebase. What
was NOT covered is the guard — resolveExtraWatchTargets iterates
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, so with only one entry it watched Kimi
CLI's config.toml and never Kimi Code's. Targets go 2 -> 3.

The existing 'extra targets are DERIVED from the descriptor array' test
cannot catch this: it builds its expectation FROM the array, so removing an
entry shrinks the expectation with it. Verified — with the kimi-code
descriptor removed that test still passes while the new named test fails.
This is the enumeration-relative scope boundary the suite already documents
one layer down, biting one layer up.

Also rewrites NON_REGISTRY_OWNED_FILE's docblock, which asserted "today's
only such descriptor is kimi's ~/.kimi". There are now two, and its named
residual is load-bearing rather than vacuous.
2026-08-08 05:50:49 -05:00
0xdhx
111733a057 chore(#2665): regenerate the examples CONTEXT-INDEX.json mirror too
lint-example-parser-parity holds examples/dynamic-context-management/
CONTEXT-INDEX.json to a fresh parse of CONTEXT.md; the round-4 seam-entry
rewrites needed this second derived artifact regenerated alongside
docs/CONTEXT-INDEX.json. Derived regen only.
2026-08-08 05:50:49 -05:00
0xdhx
123ba26fc5 chore(#2665): regenerate docs/CONTEXT-INDEX.json for the round-4 CONTEXT.md edits
The round-4 seam-entry rewrites (packaging fact on SEAM.module, strict-mode
state on SEAM.severity/ci-blind) left the derived index stale, which failed
lint:generated-sync and the #2944 sync test across ten CI contexts. Derived
artifact regen only; no content change beyond what CONTEXT.md already says.
2026-08-08 05:50:49 -05:00
0xdhx
652b99b71b test(#2665): reversion guards for all three round-4 fixes
None of the round-4 fixes had a test that fails on reversion: re-shipping
the test-instrumentation chain passes #2858 (everything ships, so every
require resolves), dropping the strict env from test.yml demotes the guard
to report-only with nothing red, and the skillsHome derivation tests are
enumeration-relative over declarations that are all empty today. One
guard each, every one negative-controlled against its reverted fix (fails
pre-fix, passes post-fix):

- packaging-shipped-scripts-require-only-shipped.test.cjs asserts the four
  chain files are absent from the npm pack file list (reuses the tarball
  set the #2858 gate already resolves — no second npm pack).
- live-config-guard.test.cjs asserts all three test jobs wire
  GSD_STRICT_LIVE_CONFIG_GUARD, matching the WHOLE expression anchored —
  a prefix match accepted both a Windows-silently-strict tail and a
  malformed one.
- helpers-process-isolation.test.cjs cold-requires helpers.cjs in a child
  with sentinel skillsHome env vars injected into both enumerations, so
  the walk itself is under test rather than today's empty declarations.
2026-08-08 05:50:49 -05:00
0xdhx
7b8c36f904 fix(#2665): walk skillsHome.env on both descriptor rungs of the scrub derivation
Review round 4, Minor 3. A configHome descriptor can nest a second,
independently-resolved descriptor (skillsHome -> resolveSkillsBaseFromDescriptor)
carrying its own env array, and the derivation walked configHome.env alone —
the identical walk-one-field gap-shape rounds 2-3 closed for the registry
and the non-registry set. Inert today (only kilo declares skillsHome, with
env: []), closed before it is live rather than after.

The guard's root enumeration deliberately does NOT gain the skills base:
getGlobalSkillsBase returns a skills directory (codex: ~/.agents/skills),
not a config root, and the snapshot applies the config-root layout beneath
every root — adding it false-positives on <skillsBase>/gsd-core while
missing a real <skillsBase>/gsd-help write (found by this round's pre-push
adversarial review). Watching skills bases needs its own layout, like
resolveExtraWatchTargets; a comment in resolveLiveConfigRoots records the
non-action.

New derivation test asserts both skillsHome rungs land in TEST_ENV_BASE,
with an anti-vacuity check that at least one runtime actually declares the
field.
2026-08-08 05:50:49 -05:00
0xdhx
253250580e fix(#2665): wire the live-config guard to strict mode on Linux/macOS CI lanes
Review round 4, Major 2, answering the explicit report-only-vs-strict
question: strict now. A future regression of the class this PR closes
should fail CI, not print a warning nobody reads — that is what the PR
title promises.

Scoped deliberately: GSD_STRICT_LIVE_CONFIG_GUARD=1 on the Linux/macOS
lanes of all three test jobs; Windows lanes stay report-only because the
guard's first run found pre-existing USERPROFILE leaks there (~190 test
sites sandbox HOME alone) — flipping them strict today reddens next on a
defect class this PR does not carry. Promote once that sweep lands (the
SEVERITY note in live-config-guard.cjs and the CONTEXT.md seam both now
record that state).
2026-08-08 05:50:49 -05:00
0xdhx
209f2fe983 fix(#2665): exclude the test-instrumentation chain from the npm tarball
Review round 4, Major 1: scripts/live-config-guard.cjs is pure test
instrumentation and was shipping to every npm install. The repo already
carries the exclusion convention (gen-emitted-baseline, qa-smell-ratchet)
in the same files[] array.

Excluding the guard alone would trip the #2858 shipped-requires-only-shipped
gate: run-tests.cjs (shipped) requires it at load time, and
affected-tests-lib.cjs / run-affected-tests.cjs sit on the same chain. The
four files are one closed require chain of test instrumentation, so the
exclusion covers the chain, not one link. The guard's LOCATION header cited
affected-tests-lib.cjs as "the precedent for a non-shipped helper", which
npm pack disproves — rewritten to the tarball-exclusion fact.
2026-08-08 05:50:49 -05:00
0xdhx
706bd2ab4e refactor(#2665): derive the guard's non-root targets from the descriptor array too
Follow-up to 38c9395d, found while fact-checking the round-3 response rather
than by a test.

That commit made TEST_ENV_BASE derive its keys from
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, but had the guard call
resolveKimiHooksTomlDir directly. Both halves covered kimi, so nothing was
broken — but only one of them would pick up a SECOND descriptor. That is the
same partial-enumeration defect that put KIMI_SHARE_DIR outside the scrub set,
reintroduced one layer over, in the very commit that closed it.

resolveExtraWatchTargets now iterates the array and resolves each descriptor
through resolveConfigHomeFromDescriptor, so the scrub set and the guard derive
from one source and cannot drift apart.

Verified: a synthetic second descriptor is picked up automatically (it was not
before); kimi's target is unchanged on both the default (~/.kimi/config.toml)
and KIMI_SHARE_DIR override paths.

The new test asserts one target per descriptor plus the store root. The COUNT
is the load-bearing half — every per-descriptor assertion passes vacuously
today with a single entry, so only the count fails when the array grows and the
guard does not follow.

NAMED RESIDUAL, documented at NON_REGISTRY_OWNED_FILE: this assumes every
non-registry descriptor is written the same way (config.toml). A descriptor
whose owned file differs needs a per-descriptor mapping. It fails toward
under-watching rather than false positives, so it is called out rather than
left to be discovered.
2026-08-08 05:50:49 -05:00
0xdhx
fb02fa5722 chore(#2665): add the changeset fragment this PR now needs
Caught by CI on the round-3 push, not by review: `changeset-lint` went from
OK_NO_USER_FACING_CHANGES to FAIL_MISSING_FRAGMENT.

Both prior rounds passed that gate because the diff was tests/ + scripts/ only,
and neither is a USER_FACING_PREFIX. Round 3 is the first commit in this PR to
touch src/ — the descriptor hoist in src/runtime-homes.cts — and src/ is on the
list. So the fragment requirement is a direct consequence of the remedy shape,
not something the earlier rounds missed.

src/runtime-homes.cts is the ONLY user-facing file in the whole PR diff
(23 changed files); everything else is tests/, scripts/, CONTEXT.md, the two
generated CONTEXT-INDEX.json artifacts, and examples/.

Typed `Added` rather than `Fixed` because that is what actually reaches a user:
new exported constants on a shipped module. The fix itself (#2665) is test
hermeticity, which changes no shipped behaviour — resolveKimiHooksTomlDir()
resolves identically before and after.

Worth noting the local lint disagreed with CI and was wrong: it diffs
`origin/main...HEAD`, which sweeps in the base's own committed fragments and
returns a false ok_fragment_present. CI diffs `origin/$GITHUB_BASE_REF...HEAD`
(next) and sees only this PR's files, which is the correct question.
2026-08-08 05:50:49 -05:00
0xdhx
434d71b03b docs(#2665): catalog the config-location and live-config-guard seams in CONTEXT.md
Round 2, Minor. "Workspace seams" carried WORKTREE.SEAM.* and
CONFIG.SEAM.loadConfig-context but nothing for this PR's mechanism or its env
vars, so the one machine-readable place a future author would look said nothing
about the class that has now recurred three times.

Ten predicates across two groups:

  CONFIG.LOCATION.SEAM.*  — the scrub set's four derivation sources and the rule
                            that a new var is made ENUMERABLE rather than
                            appended; the two-families distinction (runtime
                            configHomes vs GSD's own GSD_HOME/GSD_AGENTS_DIR)
                            that round 2 turned on; kimi's two config-location
                            vars; and the in-process scrub requirement, since
                            HOME sandboxing alone is the trap that produced
                            Blocker 1 twice.
  LIVE-CONFIG.GUARD.SEAM.* — module + exports + why it is scripts/ and not
                            scripts/lib/; ownership-based scope; the two
                            non-root targets and their asymmetric treatment;
                            the truncation contract; the report-not-fatal
                            severity ratchet; and that CI is structurally blind
                            here, so green CI is not evidence.

Both generated indexes regenerated. docs/CONTEXT-INDEX.json is checked by
lint:generated-sync (`gen-context-index.cjs --check`) LINE-NUMBER-SENSITIVELY,
and the example's own committed index is separately checked by
lint-example-parser-parity.cjs, which the first regen did not satisfy — editing
CONTEXT.md requires both, and only one of them says so in its error text.

Verified: parity lint rc=0, gen-context-index --check rc=0, full
lint:generated-sync rc=0. Regen diff audited — 10 predicates added, 0 removed,
0 values changed; the example index's remaining churn is line-number re-baking,
which is exactly why the parity lint excludes line numbers.
2026-08-08 05:50:49 -05:00
0xdhx
fec42e9a7d test(#2665): cover the widened derivation, and mark what these tests cannot prove
Two tests for round 3's change (round 2 Blockers 1 and 2): KIMI_SHARE_DIR must
arrive via NON_REGISTRY_CONFIG_HOME_DESCRIPTORS rather than a literal, and every
GSD_LOCATION_ENV_KEYS entry must be blanked. Both assert a floor on their source
first, so a renamed export fails loudly instead of passing vacuously.
Negative-controlled against the registry-only derivation: both fail there, and
only those two.

And the Nit, which is the more useful half. This block asserts that TEST_ENV_BASE
is not narrower than the enumerations it derives from. It cannot prove those
enumerations are complete — a var no enumeration carries is invisible to every
test here, and they stay green.

That is exactly how round 2 found GSD_HOME and KIMI_SHARE_DIR while this block
was fully green: one belonged to no enumeration at all, the other sat inside a
function body where nothing could enumerate it. So the scope boundary is now
written down at the top of the block, naming where the completeness question is
actually answered — a source census re-derived each round, and
live-config-guard.cjs observing real writes at runtime — so that a green run
here is not misread as "the set is exhaustive."

Deliberately NOT added: an assertion per reviewer-named variable. That is the
hand-maintained list wearing a test's clothes, and it fails the same way.
2026-08-08 05:50:36 -05:00
0xdhx
1f6d827e48 fix(#2665): scrub config-location env in the #2624 in-process install block
Self-found during the rebase onto next, not from the review.

The base range added `describe('#2624 .gsd-source marker is rewritten before
staging reads it')`, which calls the real `install(true, 'claude')` IN-PROCESS
and sandboxes HOME/USERPROFILE/GSD_EXPLICIT_CONFIG_DIR — but not
CLAUDE_CONFIG_DIR. That is exactly the Blocker-1 shape this PR exists to close,
reintroduced in new code written after the round-1 review.

Measured on the rebased tree, same file, same commit:

  CLAUDE_CONFIG_DIR unset -> 50/50 pass
  CLAUDE_CONFIG_DIR set   -> 47/50, and a COMPLETE global install lands in it
                             (gsd-core/, agents/, skills/, hooks/, scripts/,
                             gsd-file-manifest.json, gsd-install-state.json,
                             .gsd-source, .gsd-profile)

The three failures are the honest symptom rather than the problem: the install
goes to the ambient config dir, so the assertions look for a marker under
tmpRoot that was never written there.

With scrubConfigLocationEnv() wired into the block's beforeEach/afterEach —
the same pattern the sibling block at :753 already uses — the file is 50/50
under BOTH conditions and leaks zero entries.

Worth stating plainly: CI cannot catch this class, since CI never has these
vars set. It surfaced here only because the rebase brought the base's new tests
under an ambient CLAUDE_CONFIG_DIR, which is the condition #2665's own
acceptance criterion runs under.
2026-08-08 05:50:36 -05:00
0xdhx
a294ec2a2b test(#2665): widen the hermeticity guard to its two blind surfaces, and cover its budget
Round 2, both Majors. They are one defect seen twice: the recurrence guard did
not cover the surface it exists to guard.

Blind surfaces. resolveLiveConfigRoots enumerates getGlobalConfigDir per registry
runtime plus a hardcoded grok branch, so it can only ever see runtime config
ROOTS. Two live write surfaces are not roots and passed through silently:

  $GSD_HOME/.gsd     — GSD's user-owned store. Watched WHOLESALE: unlike ~/.claude
                       this root is exclusively ours, so the shared-root
                       false-positive trap the module documents does not apply.
  <kimi>/config.toml — the file GSD writes its native [[hooks]] block into. The
                       INVERSE case: ~/.kimi belongs to Kimi CLI, so only the one
                       file GSD writes is watched, never the root.

That asymmetry is why this is not a two-line "add two roots" patch — one target
needs the whole tree, the other needs exactly one file, and collapsing them
either under-watches the store or trips the guard's own documented
false-positive trap on a third party's directory.

Extras are passed to snapshotLiveConfig explicitly rather than resolved inside
it, so a caller snapshotting a fixture root cannot silently pull the developer's
real ~/.gsd into its own assertions. run-tests.cjs now snapshots when EITHER the
roots or the extras are non-empty — previously an unbuilt tree yielding zero
roots disabled the entire guard without saying so.

Budget coverage. The MAX_ENTRIES/MAX_DEPTH bound and the truncated -> 'unverified'
branch had zero tests, despite this module's own docstring naming "a truncated
scan reading as clean" as the safety-critical case. Added per
RULESET.TESTS.boundary-coverage (N in {limit-1, limit, limit+1}, exercised
through newestMtime's injected budget so the boundary is real without
materialising 20000 files) and RULESET.TESTS.property-based-testing (fast-check:
truncation is monotone in the budget; reported newest never exceeds the true
maximum). A regression flipping `truncated` to false on an exhausted budget now
breaks the property for every budget below the tree size.

Negative-controlled: neutering the extras wiring fails exactly the two
new-surface tests and nothing else. 21/21 green with it restored.
2026-08-08 05:50:36 -05:00
0xdhx
3c580b77dc fix(#2665): derive the second config-location family instead of hand-adding it
Review round 2 named GSD_HOME and KIMI_SHARE_DIR as missing from the derived
scrub set. Both premises confirmed; the prescribed remedy is not adopted
verbatim, because adding two more literals to a four-item hand list is the
pattern that reopened this bug three times. The census the derivation
generalizes over was partial, so the census is what widens.

Two structural gaps, both closed at the source:

1. KIMI_SHARE_DIR lived inside resolveKimiHooksTomlDir's body as an inline
   descriptor, resolvable but not ENUMERABLE. Hoisted to an exported
   KIMI_HOOKS_TOML_DESCRIPTOR and collected in
   NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, which TEST_ENV_BASE now derives from.
   kimi is the sharp case: it owns TWO config homes (KIMI_CONFIG_DIR, already
   registry-visible, and this one), so a registry-only derivation looks
   complete and is not.

2. GSD_HOME is a different FAMILY, not a missing registry entry. The registry
   describes where third-party runtimes keep config; GSD_HOME decides where GSD
   keeps its own user-owned state ($GSD_HOME/.gsd/ — consent.json,
   defaults.json, capability overlays), read env-first ahead of os.homedir() by
   capability-loader, capability-consent, capability-state, capability-writer,
   config-loader, install-profiles and bin/install.js. Named as
   GSD_LOCATION_ENV_KEYS rather than folded into the descriptor array, since it
   does not resolve through resolveConfigHomeFromDescriptor.

GSD_AGENTS_DIR joins the same family (round 2, Minor): env-first and
unconditional in getAgentsDir, misdirecting a read rather than a write.

A census of every env-first first-party location var — the guard-shape question
this PR owes each round — now yields exactly one remaining unguarded name,
GSD_MODEL_CATALOG, and it is dead by precedence: the co-located candidate is
index 0 and the loop breaks on first success, so the env var can only win on a
tree that is already broken, and it redirects a read even then.

Derived set: 39 -> 42 keys. resolveKimiHooksTomlDir behaviour unchanged on both
the default and the KIMI_SHARE_DIR override path.
2026-08-08 05:50:36 -05:00
0xdhx
e2eed1c58a test(#2665): ship the hermeticity guard at report level, not fatal
Its first CI run found PRE-EXISTING leaks on the Windows lane —
C:\Users\runneradmin\.claude\gsd-core and skills\gsd-dev-preferences — with all
1196 Windows tests otherwise passing. os.homedir() reads USERPROFILE on Windows,
and ~190 test sites across 31 files sandbox HOME alone, so the suite has been
installing GSD into the runner's real home directory invisibly. That is exactly
the class the guard exists to surface, and exactly the class this PR's review
said CI could never catch.

It is also a different defect from the one #2665 closes, and too large to fold in
here. A brand-new gate that immediately reds an unrelated lane gets bypassed or
reverted rather than obeyed, so the guard reports by default and fails only under
GSD_STRICT_LIVE_CONFIG_GUARD=1.

This is the repo's own established ratchet, not a hedge: the local/no-source-grep
ESLint rule shipped at `warn` and was promoted to `error` after its cleanup sweep
(ADR 452). Promote this the same way once the USERPROFILE sweep lands.
2026-08-08 05:50:17 -05:00
0xdhx
5863351281 test(#2665): separator-safe containment in the regression assertion
startsWith(ambientConfigDir) also matches a sibling like
<tmp>/ambient-live-config-2, so it can report a leak that did not happen. Use
path.relative and check for '..' or an absolute result, the repo's usual shape.

The readdirSync assertion already carried the test, so this is cosmetic.

Addresses review finding: Nit 9.
2026-08-08 05:50:17 -05:00
0xdhx
771980b661 test(#2665): restore fallback-branch coverage in the #2003 regression
This PR fixed the test's real defect -- it compared the child's answer against
the PARENT process's getGlobalConfigDir(), two different environments, agreeing
only because the child inherited the developer's ambient CLAUDE_CONFIG_DIR -- but
fixed it by INJECTING CLAUDE_CONFIG_DIR, which moved the test onto the env-first
branch and silently dropped the only coverage #2003 had of the home-derived
fallback.

Sandbox HOME/USERPROFILE and leave the config vars blank instead: the expectation
stays test-controlled AND the branch under test is unchanged.

Drops the notStrictEqual against the codex dir. It could not fail whenever the
strictEqual on the line above passed.

Addresses review finding: Minor 7.
2026-08-08 05:50:17 -05:00
0xdhx
a4efa4deda test(#2665): cover the derivation and the in-process scrub
The single regression test exercised CLAUDE_CONFIG_DIR only, so deleting
GSD_RUNTIME or CODEX_HOME from any literal broke nothing -- a mutation of 24 of
the 27 added key-value pairs survived.

Four tests here: every configHome env var the registry declares is scrubbed;
every scrubbed key is blanked rather than merely present; the four non-registry
vars are named explicitly so deleting one is a failure rather than a silent
narrowing; and scrubConfigLocationEnv round-trips both a set and an unset var
(restoring an originally-unset var as '' would itself be a leak).

Both derivation tests assert a floor on the registry first, so a renamed registry
shape fails loudly instead of making the assertions vacuously true.

Negative-controlled against the hand-written 3-key list this PR shipped: the
parity test fails there and names all 19 missing vars.

Addresses review findings: Major 6, Minor 8.
2026-08-08 05:50:17 -05:00
0xdhx
a02462e050 test(#2665): fail the suite when it writes into a live config dir
The recurrence guard, and #2665's own "Optional hardening". This class is silent
by construction: TEST_ENV_BASE cannot see an in-process caller, and CI cannot see
the class at all because CI never has these env vars set. It damages the
developer's machine and reports nothing -- which is how two prior authors each
diagnosed it and fixed only the instance in front of them.

run-tests.cjs snapshots GSD's install footprint in every live runtime config dir
before the suite and re-checks it after, failing the run on a create or a modify.
Roots come from the product's own getGlobalConfigDir, so the guard watches
wherever the product actually points, including through an ambient var.

Scope is ownership-based, not whole-root: the top-level install footprint plus
gsd-prefixed children of dirs GSD shares with the host agent. A config root like
~/.claude is shared, and watching it wholesale would false-positive on the host's
own history.jsonl or settings.json -- a guard that cries wolf gets disabled, and
then catches nothing. The prefix test is load-bearing: the first version watched
only the three top-level entries and MISSED a real leak into skills/gsd-*.

It earned its place immediately -- it is what found the fifth in-process leak in
runtime-artifact-layout.test.cjs, which no amount of reading the review would have
surfaced. Known gap documented in the module: a write to a file GSD does not own
is out of scope by construction.

Lives in scripts/, deliberately NOT scripts/lib/ -- the installer copies that dir
into every user's config dir wholesale while uninstall removes only an allowlist,
so a test-only module there would ship to users and survive uninstall.

Addresses review finding: Minor 8.
2026-08-08 05:50:17 -05:00
0xdhx
36f4f9479a fix(#2665): scrub config-location env on raw spawns that sandbox only HOME
Same class as the in-process leak, on the child-spawn side. These call spawnSync
directly rather than through runGsdTools, so TEST_ENV_BASE never applies to them
and an ambient config-location var survives into the child.

  - install-runtime-artifacts: the dev-preferences writer resolved env-first and
    wrote SKILL.md into the live config dir instead of its tmp HOME. This also
    fixes a test that FAILS today on any machine with CLAUDE_CONFIG_DIR set.
  - issue-766: the `claude` CLI is third-party and bootstraps its own config into
    whatever CLAUDE_CONFIG_DIR names, so even a bare --version probe wrote there.
    It gets an explicit throwaway dir rather than a blank value -- GSD's resolvers
    treat '' as falsy and fall back to the home dir, but a third-party binary
    offers no such guarantee (blanking it produced a stray backups/ in the repo
    root).

These two were the last writers standing between the suite and #2665's stated
acceptance criterion.
2026-08-08 05:50:17 -05:00
0xdhx
466db2f1a3 fix(#2665): scrub config-location env on the parent for in-process install()
The blocker the review said decides this PR. These tests call the real installer
IN-PROCESS with only HOME/USERPROFILE sandboxed. getGlobalConfigDir is env-first,
so an ambient CLAUDE_CONFIG_DIR beats the sandbox and a complete global install --
agents/, commands/, skills/, gsd-core/, manifest, settings -- lands in the
developer's live config dir. No child-env scrub can reach it; only clearing the
parent's env can.

The review named two describe blocks in install.test.cjs. A sweep for the shape
found four there (bug #3571 and bug #3288 each appear twice in the file), and the
post-suite guard added later in this series found a fifth in
runtime-artifact-layout.test.cjs, which the review did not name. All five now
save/clear/restore via scrubConfigLocationEnv().

Verified: codebuddy-install, cline-install and codex-config already guard their
own config-location var around in-process install(), so the class is closed.

Addresses review finding: Blocker 1.
2026-08-08 05:50:17 -05:00
0xdhx
bc5c362610 fix(#2665): replace the hand-synced TEST_ENV_BASE copies with the canonical import
Eight declarations were kept in sync by hand with no parity assertion. The drift
was already in the tree: api-coverage-gate-e2e and representative-corpus declared
TERM_SESSION, but the real variable is TERM_SESSION_ID, so both scrubbed nothing
for that slot and leaked TERM_SESSION_ID into every child. Deleting the copies
removes the dead key with them -- there is no longer a second place to get wrong.

run-tests-harness keeps its local session-identity literal: that helper mirrors
the production runGsdTools to prove its CONTRACT, so it must not re-import what
it is testing. The config-LOCATION keys are a safety scrub rather than part of
that contract, so it spreads the canonical derived set and keeps the rest local.

Addresses review findings: Blocker 2, Blocker 3.
2026-08-08 05:50:17 -05:00
0xdhx
0f97a26f76 fix(#2665): derive the config-location scrub set from the capability registry
TEST_ENV_BASE listed three config-location vars by hand. The resolver reaches
25: every runtime descriptor's configHome.env, plus GROK_AGENTS_HOME (a hardcoded
branch of getGlobalConfigDir), GSD_RUNTIME, and GSD_PROJECT/GSD_WORKSTREAM
(planningDir). A hand-written list can only ever be as complete as the author's
recall, and every one of those resolvers is env-FIRST, so a missing key is a live
escape hatch rather than a cosmetic gap -- which is why this bug has now been
diagnosed three times.

Derive the set from the same registry the resolver reads. The scrub list becomes
structurally incapable of being narrower than the surface it guards: adding a
capability that declares a new configHome env var extends it in the same commit.

Also export TEST_ENV_BASE (a parity test could not previously import the
canonical copy) and add scrubConfigLocationEnv(), the in-process counterpart --
TEST_ENV_BASE only ever reaches child processes.

Addresses review findings: Blocker 2, Major 4 (GSD_WORKSTREAM/GSD_PROJECT),
Major 5 (XDG_CONFIG_HOME).
2026-08-08 05:49:33 -05:00
0xdhx
08021b02c0 fix(#2665): scrub config-location env vars in every TEST_ENV_BASE declaration
TEST_ENV_BASE blanks session-identity variables but none of the three that
decide WHERE a child process writes: CLAUDE_CONFIG_DIR, GSD_RUNTIME and
CODEX_HOME. The config-home resolver is env-first (runtime-homes.cts, the
dot-home case consults the env var before the home-derived fallback), so an
ambient CLAUDE_CONFIG_DIR in the developer's shell beats a call site that
sandboxes only HOME. The suite then writes into the developer's real config
directory -- including a registered skill under <configDir>/skills/ whose
body carries behavioural directives that load into later sessions.

Blank all three alongside the session-identity vars. `...env` still spreads
last, so the five call sites that already constrain these locally keep
winning with their explicit values.

TEST_ENV_BASE is re-declared in nine files, so the three lines are added
nine times rather than once. Consolidating the nine into a single exported
constant -- and fixing the TERM_SESSION / TERM_SESSION_ID drift between the
copies -- is deliberately left out of this change; see the PR body.

One call site needed adjusting. capability-state.test.cjs's
`capability state --runtime claude` CLI test passed no env at all and
compared the CHILD's resolved config dir against the PARENT process's
getGlobalConfigDir('claude'). That agreed only because the child inherited
the developer's ambient CLAUDE_CONFIG_DIR -- i.e. it passed *because of*
the leak. It now redirects both runtime homes into the sandbox and asserts
against values the test controls, so it is hermetic with the variable set
or unset.

Regression case folded into the owning module's test file rather than a new
bug-NNNN file, per scripts/lint-regression-test-names.cjs. It sets the
variable on the PARENT process, which is the actual vector; setting it in
the per-call env argument would exercise a path that was never broken.
2026-08-08 05:49:33 -05:00
Tom Boucher
343835facc refactor(#3183): route live-plan counting through scanPhasePlans (#3199)
* refactor(#3183): route live-plan counting through scanPhasePlans

scanPhasePlans becomes the sole owner of the live-plan derivation. Twenty-one
independent re-derivations across seven modules now route through it, and
scripts/lint-plan-count-drift.cjs reports zero, scanning the whole repo rather
than an allowlist (ADR-3180 Decision 4a).

The epic scoped this at three copies. A whole-repo guard found twenty-six sites
across nine files, so Phase 1 absorbs every live-plan re-derivation and Phase 3
narrows to window plus sentinel enumeration.

Two sites are exempt with a documented reason rather than a bare allowlist:
audit.cts scans one quick task's own directory for a single completion record,
and gsd2-import.cts reads a foreign GSD-2 tasks/ layout during a one-time
import. Neither is a phase directory.

scanPhasePlans gains allPlanFiles (pre-supersession) alongside planFiles so one
owner answers both questions: verify.cts's numbering-gap check wants every plan
on disk, its pairing check wants the live set. Both fields are additive.

Highest-severity fix: cmdPhasePlanIndex, which feeds execute-phase wave
scheduling, was scheduling status:superseded plans into waves and reporting zero
plans for the post-#3139 nested layout.

filterPlanFiles and filterSummaryFiles are deleted; getPhaseFileStats orphaned
them and only their own tests still called them.

New leaf module src/planning-scope.cts carries the frozen SCOPE discriminator,
with its six-gate ripple closed: gitignore, inventory manifest, INVENTORY.md and
the CONTEXT.md glossary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* docs(#3183): amend ADR-3180 for the Phase 1/3 boundary re-slice

The contract held; the phase boundary did not. The whole-repo drift guard found
26 re-derivations across 9 files against the epic's estimate of 3, and
cmdProgressRender re-derives both enumeration and plan counting on adjacent
lines, so DW4 was unsatisfiable within Phase 1's original file scope.

Records the amended scope, scanPhasePlans's new allPlanFiles field,
findOrphanSummaries, the two documented exemptions, the re-derived Tier-2
table, and the describeNonCanonicalPlans trap for later phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): complete the canonical pairing rule and gate the naming diagnostic

The remote runner went red with 13 deterministic failures on both lanes,
and they were right: replacing verify.cts's canonicalPlanStem pairing with
summaryCandidates dropped a case the bespoke rule covered. A plan carrying a
descriptive slug after its id (68-01-scaffolding-PLAN.md) pairs with its
canonical-stem summary (68-01-SUMMARY.md), and summaryCandidates generated no
such candidate, so the plan read unsummarized.

The fix is to complete the one rule rather than restore a second:
summaryCandidates gains a canonical-id candidate, narrowed to fire only when an
id pair was actually extracted. countMatchedSummaries, findUnsummarizedPlans
and findOrphanSummaries all inherit it. The two-plans-one-summary collision
behaviour of the original rule is preserved deliberately and documented in
place.

Second defect, independently root-caused while verifying: routing the #2893
naming diagnostic through scanPhasePlans exposed it to the loose /PLAN/i
fallback, which is correct for counting and wrong for a naming check — a
non-canonically-named file was accepted as a valid plan and the diagnostic
went silent. cmdPhasesList, cmdFindPhase and cmdPhasePlanIndex now intersect
with a strict isCanonicalPlanFile predicate before reporting names.

Same class as the describeNonCanonicalPlans trap already recorded in ADR-3180:
a question about file naming wants the physical, strictly-matched set; only a
question about outstanding work wants the live set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): register planning-scope.cjs in the eslint migration list

tests/repo-invariants.test.cjs asserts every bin/lib/*.cjs is linted xor
ignored per its ADR-457 migration state. The new planning-scope module closed
five of the six .cts ripple gates - gitignore, inventory manifest, INVENTORY.md
and the CONTEXT.md glossary - but not eslint, because that one is enforced by a
test rather than by lint:ci, so the local pipeline stayed green while it was
missing.

Generated from src/planning-scope.cts, so the .cjs is ignored and the .cts is
linted, matching every other migrated module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): replace the plan-count drift detector with a literal tokenizer

CodeQL reported 4 high-severity js/redos alerts on REGEX_LITERAL_MD_RE, the
backtracking regex that finds "a regex literal mentioning PLAN/SUMMARY and an
escaped \.md". Five review rounds found it had two defects, not one:

  - EXPONENTIAL, then CUBIC. Its "any char" atom `(?:\\.|[^/\r\n])` let a `\.`
    pair be consumed either as one escape or as two class characters, which is
    exponential backtracking: 27,464ms on `"/\.mdplan" + "\.".repeat(28) + "X"`.
    Excluding `\` from the class killed that but left a cubic path — 23ms at
    N=200, 172ms at N=400, 1362ms at N=800 on `"/" + "PLAN\.md".repeat(N)` with
    no closing `/`. This guard is the last stage of `npm run lint:ci`, which CI
    runs on fork pull requests, so a crafted src/*.cts could stall the job.
  - A DETECTION HOLE. A character class holding a bare, unescaped `/` — e.g.
    `/SUMMARY[^/]*\.md$/`, an ordinary path-excluding filter — terminated the
    literal at that `/`, so the scan never reached `\.md` and the guard missed
    it entirely. (Classes holding an ESCAPED `\/` were already matched; the
    tests cover those separately as parity, not as regressions.)

Both defects have one root cause: regex-literal grammar — `\x` escapes, and
`/` inside `[...]` not terminating — is not expressible in a backtracking
regex. So the detector is now a tokenizer, not a regex.

readRegexLiteralAt reads the literal at a given `/` in a single left-to-right
pass with no backtracking, treating escapes as two-character units and
suppressing the `/` terminator inside a character class. findRegexLiteralMdMatch
restarts it at every `/` on the line, preserving the old "find anywhere"
behaviour; MAX_REGEX_LITERAL_LEN (400) bounds each read — including the
trailing-flag scan — which keeps the whole-line cost linear.

Results: cubic shape flat at 0.06-0.39ms out to N=3200 (25KB), exponential
shape 0.01ms at 28 reps and 0.00ms at 64, and the bare-`/` class shapes are now
caught. Differential against the old regex over 28,474 lines (those matching
FILENAME_TEST_RE but not PLAN_SUMMARY_LITERAL_RE, across src/tests/scripts/
gsd-core/bin/eslint-rules, excluding 265 lines with >6 backslashes on which the
old regex hangs): 6 differences, all the tokenizer returning the fuller or
newly-correct literal, 0 old-only misses. The `\.md` token stays
case-insensitive, matching the `/i` the old regex carried.

Also closes three holes in the same new file:

  - walk() tested entry.isFile(), false for a symlink, so a symlinked
    src/*.cts was silently unscanned — an evasion of a guard whose stated
    principle (ADR-3180 Decision 4a) is whole-repo discovery with no allowlist.
    It now resolves symlinks, but confined: file links must resolve inside the
    repo root, directory links inside the scanned dir itself. Every sibling
    drift guard in scripts/ uses the Dirent classification and never follows
    links, so following them unconfined would have made this the only linter
    able to read outside the tree — on fork PRs an arbitrary out-of-repo read
    whose matched fragments reach a public CI log. The narrower directory rule
    additionally stops `src/up -> ..` from sweeping the whole repo, and the
    skip list is now checked against resolved paths so `src/g -> ../.git`
    cannot reach .git/** or node_modules/**. Real paths are de-duplicated and
    files reported canonically, so a symlink alias cannot shift which
    FUNCTION_SCOPED_EXEMPTIONS key applies.
  - Both the reported fragment and the reported FILE PATH are attacker-
    controlled source text written straight to a CI log, and git permits
    control bytes in a filename. Both are now escaped — C0/C1/DEL plus the
    bidi and zero-width controls — so a crafted literal or filename cannot
    recolour the log, overwrite a line with CR, or fabricate a line that looks
    like this guard's own success output.

Regression coverage in tests/plan-count-single-owner.test.cjs: a child-process
probe over both pathological shapes (catastrophic backtracking is synchronous
and would freeze the suite rather than fail one test), the bare-`/` class
shapes verified to fail against the parent-commit blob, root-confinement tests
covering the outside-file, outside-directory, cycle, broken-link and duplicate
cases, direct isInsideRoot coverage including the sibling-prefix case that a
bare startsWith would let through, sanitizeForReport coverage, and
limit-1/limit/limit+1 coverage of MAX_REGEX_LITERAL_LEN derived from the
exported constant. The earlier structural assertion was dropped — it checked
for the substring `[^/`, which respelling the class as `[^\r\n/]` defeats
while staying exponential.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): backfill changeset PR number

Restores b77931869, which a force-push during the ReDoS remediation dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 01:20:55 -04:00
Tom Boucher
409da3223c Merge pull request #3202 from open-gsd/chore/backmerge-main-to-next-62f44f8d
chore: back-merge main → next (62f44f8d)
2026-08-08 01:08:13 -04:00
github-actions[bot]
9f22fc63e1 chore: back-merge main into next (62f44f8d) 2026-08-08 05:07:50 +00:00
Tom Boucher
97327b994e Merge pull request #3201 from open-gsd/chore/sync-next-version-1.10.0
chore: sync next package version to 1.10.0
2026-08-08 01:07:40 -04:00
github-actions[bot]
4f2a4932ce chore: sync next package version to 1.10.0 2026-08-08 05:07:32 +00:00
Tom Boucher
62f44f8d7c Merge pull request #3200 from open-gsd/release/1.10.0
chore: merge release v1.10.0 to main
2026-08-08 01:07:30 -04:00
github-actions[bot]
68a04ccf8e chore: promote CHANGELOG for v1.10.0 2026-08-08 05:06:54 +00:00
github-actions[bot]
9335dc00d2 chore: bump version to 1.10.0 for release 2026-08-08 04:30:53 +00:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
664d49e513 docs(#3182): ADR-3180 — planning semantic model single owner (#3196)
* docs(#3182): ADR-3180 — planning semantic model single owner

Phase 0 design lock for epic #3180. Names one canonical owner per
semantic derivation, specifies the frozen-enum scope contract that
distinguishes a genuinely-empty computation from a truncated or
unscoped one, and locks the drift-guard contract.

Ships no production code. Phases 1-5 execute against this ADR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* test(#3182): prune real issue 3182 from phantom-ref guard, fix empty-list regex

The guard's own header documents that its list rots: entries are phantom
only until the repo's shared issue/PR counter reaches them, and once the
counter passes an entry it must be deleted. Creating the Phase-0 sub-issue
advanced the counter past 3182, so the guard began rejecting a legitimate
citation of a real issue - the failure its header already records happening
twice, with PRs 2551 and 2361.

3182 was the last entry, and removing it exposed a latent bug: the regex
builder interpolated the list unconditionally, so an empty list yields
(?:#(?:)\b)|(?:issues/(?:)\b), whose empty alternation matches every issue
reference in the repo. Following the file's own maintenance instruction
would have turned a green guard into one failing on nearly every file.

buildRefRe() now returns null for an empty list and is exported, with
boundary coverage at 0/1/2 entries plus word-boundary and bare-digit
negative cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 20:25:12 -04:00
Tom Boucher
3f349e551d fix(#3024): route sync-skills through the shipped gsd-tools instead of an unshipped install.js (#3195)
* fix(#3024): sync-skills workflow uses gsd-tools query skills-root instead of unshipped install.js

The sync-skills workflow Step 2 shelled out to gsd-core/bin/install.js --skills-root,
but install.js is not shipped in installed trees (only in the npm tarball root bin/).
Every /gsd-update --sync invocation failed with MODULE_NOT_FOUND.

Fix: added 'gsd-tools query skills-root <runtime>' subcommand (gsd-tools IS shipped)
that calls the same getGlobalSkillsBase function install.js used. Updated the
workflow to call gsd_run query skills-root instead of the dead install.js path.

Also documented the #3025 verbatim-cp limitation in Step 5 with a workaround.

* test(#3024): failing-first guards for the three defects in the adopted fix

The cherry-picked commit came from an aborted run that never executed its own
tests. Its raw-path assertion fails as written, which is the clearest evidence
the work never reached verification.

Covers:
- --raw must emit a bare path, not JSON (output() takes a third rawValue arg
  that routeSkillsRoot omits, so the raw branch never fires)
- an unknown, empty, whitespace, traversing, or metacharacter-bearing runtime
  must be rejected, not silently resolved to claude's skills root
- sync-skills.md must contain zero references to the unshipped install.js,
  including the guard's remediation text — the issue's second reported defect
- parity across every runtime in the registry, not three hardcoded ones, so the
  two entry points cannot drift

Also converts the adopted tests off a hand-rolled spawnSync onto the bounded
process seam, per CONTRIBUTING.

Fails before the fix. Verified via the remote runner.

* fix(#3024): make the skills-root query actually work and reach non-Claude runtimes

The cherry-picked commit never ran its own tests. Six defects, all fixed here.

--raw was ignored: output() is output(result, raw, rawValue) and the third
argument was omitted, so the raw branch never fired and the workflow captured a
JSON blob as SRC_SKILLS_ROOT. Every downstream cp -r then resolved against a
nonexistent path — the command would have shipped still broken.

An unknown runtime silently resolved to claude's skills root, because
getGlobalSkillsBase falls back rather than returning null, leaving the existing
=== null guard dead. The runtime id is now validated at the CLI boundary against
the shipped registry, so a typo'd --from/--to fails instead of reading from or
writing into the wrong runtime's tree.

getGlobalSkillsBase('vscode') threw a raw TypeError. vscode is non-installable
by descriptor, so it has no skills root — null is the answer, not a crash. The
resolver now short-circuits configHome.kind 'none', which also fixes the same
latent crash in install.js --skills-root vscode. Every caller already gates on
=== null.

sync-skills.md used gsd_run WITHOUT the canonical launcher preamble, so gsd_run
was undefined on non-Claude runtimes — the fix would have been dead in exactly
the place the original bug bit. Preamble propagated via sync-runtime-launcher.

Also registers skills-root in TOP_LEVEL_USAGE (the help/dispatch parity guard
caught it), removes the last two install.js references including the guard's
remediation text (the issue's second reported defect), and updates the stale
assertion that still described the removed contract.

Verified on the remote runner.

* fix(#3024): align the documented runtime list with the registry and gate both entry points

Isolated review returned BLOCK on two findings.

The workflow's Supported-runtimes list and its --to all expansion named grok and
gemini, neither of which is a registered runtime. Once this branch added
validation, --to all — a documented first-class feature — aborted. The list was
hand-copied prose shadowing the registry, so correcting it alone would drift
again; a parity assertion now fails in BOTH directions if the doc and the
registry disagree. vscode is excluded by name: it is installSurface 'none', so
syncing skills to it is meaningless and would abort.

bin/install.js --skills-root reached getGlobalSkillsBase with no own-property
gate, so --skills-root __proto__ silently resolved to claude's skills root. This
branch had just hardened the OTHER entry point to the same function; leaving one
of two parallel surfaces open is the same divergence class as the first finding.
Both now call one shared isRegisteredRuntimeId() rather than a copied check, and
the parity test covers the hostile ids so the two can never disagree again.

Also guards the workflow's root resolution: neither command substitution checked
its exit status and only the source had an existence guard, so a failed
destination resolution left DEST_ROOT empty and turned rm -rf "$DEST_ROOT/$SKILL"
into an absolute path at filesystem root. Both resolutions are now checked, and
Step 5 requires both roots to be non-empty and absolute before any destructive
command.

Verified on the remote runner.

* test(#3024): anchor the runtime-list parity extractor to the list span

The extractor captured (.+) to end of line, so it swallowed the em-dash prose
that explains the vscode exclusion — and that sentence contains backticked
`runtimes` and `null`, which is where the three phantom ids came from. The
documented list was correct; the test was reading its own explanation back as
data. Anchored to the id-list span.

Both directions still fail as intended: proven by injecting a bogus id and by
removing a registered one.

* test(#3024): anchor the --to all extractor and fail loudly on empty captures

The workflow has three TO_RUNTIMES= assignments and the regex matched the first
one — an empty array initializer at line 28 — so the extractor captured nothing
and the assertion diffed [] against 18 ids as if that were data.

That is the same failure twice, so the fix is the general one: every extractor
in this test now asserts it captured a plausible list before comparing, naming
which extractor found nothing and what it was looking for. An extractor that
silently yields [] is a confident wrong answer, and a parity guard that reports
it as a data mismatch teaches the reader to loosen the assertion.

Verified against the real workflow and against doctored copies with each target
construct removed, plus both teeth directions.

* fix(#3024): merge duplicate process-seam import after rebase

The rebase applied cleanly but left runNode declared twice: next had gained its
own import of the seam while this branch added one carrying OUTCOME. A clean
rebase is not a correct one — the file no longer parsed. Merged into a single
import providing both.

* fix(#3024): bind DEST_ROOT per destination instead of a dangling map

Step 2 stored each destination's root into DEST_SKILLS_ROOTS, which nothing ever
read, while Steps 3 and 5 used a scalar DEST_ROOT that nothing ever assigned. The
array was also never declare -A'd, so on bash 3.2 — macOS system bash, which this
repo supports — every destination collapsed onto index 0.

The absolute-path guard added earlier was the only thing standing between that and
rm -rf "/$SKILL"; it turned a silent disaster into a hard stop, but the feature
still could not complete. Each destination now binds its own DEST_ROOT where it is
used, and the unread map is gone rather than replaced.

Step 2 keeps eager validation, so a bad runtime id in a multi-destination --to
aborts before any destination is written rather than after some already have been.

Verified on bash 3.2 with a two-destination run binding distinct roots, and with a
bad id aborting before any destructive call.

* fix(#3024): restore grok support broken by the registry gate

The registry gate added earlier rejected grok, and that was my error. I confirmed
grok was absent from the capability registry and concluded the hardcoded branch
was dead — without checking what it resolved to. It resolves to ~/.agents/skills,
a real grok-specific path, exactly as the pre-fix workflow documented ('grok uses
the ~/.agents layout'), and there is a support discussion doc for it. So a
working, documented runtime silently lost --skills-root and sync-skills support
as a side effect of prototype-pollution hardening — and the parity test I added
locked that in as correct.

gemini is the one that really was dead: it fell through to CLAUDE's skills root,
so rejecting it is right and it stays rejected, as do bogus ids, __proto__,
empty, whitespace and traversal.

The validator's real question is 'does this id have a genuine runtime-specific
resolution', not 'is it in the registry map'. Registry membership was a proxy
that happened to miss grok. Legacy non-registry runtimes with dedicated
resolution branches are now a named, documented set; enumerating every hardcoded
branch in getGlobalConfigDir against the registry confirms grok is the only one.

The new tests assert grok resolves UNDER .agents and specifically not to claude's
root. Allow-listing an id proves nothing about whether it resolves correctly —
that assertion is what would have caught my mistake.

Also uses the shared PROBE_TIMEOUT_MS instead of a duplicate literal, and guards
Step 3's DEST_ROOT re-resolution, which contradicted the file's own stated
guarantee.

Verified on the remote runner.

* test(#3024): guard against LEGACY_NON_REGISTRY_RUNTIME_IDS drifting

The named legacy set is a second hand-maintained proxy for the same predicate
the registry check got wrong — 'does this id resolve runtime-specifically'.
Nothing stopped a third hardcoded branch being added to getGlobalConfigDir
without updating the Set, reproducing the exact class of bug that broke grok.

Production stays explicit and greppable; the test derives the truth instead. It
resolves a sentinel id to learn the generic fallback, classifies every candidate
against it, and fails in both directions — an id resolving runtime-specifically
that is in neither the registry nor the Set, or a Set entry that no longer earns
its exemption. The failure message names the remedy.

Confirms grok resolves runtime-specifically and gemini does not, which is the
distinction the original registry check could not see.

Also reverts the shared-timeout swap: SKILLS_ROOT_PROBE_TIMEOUT_MS is
pre-existing on next and arrived by rebase, so changing it here was scope creep
into another issue's territory.

Verified on the remote runner.

* chore(#3024): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-07 18:53:25 -04:00