Self-found by re-deriving the guard-shape census against bin/install.js's own
write sites, not by a review finding. Three artifacts a global install writes
into a live config ROOT were watched by nothing:
hooks/gsd-check-update.js, hooks/gsd-context-monitor.js,
hooks/gsd-update-banner.js -- `hooks` was absent from GSD_PREFIXED_PARENTS
.gsd-source, .gsd-profile -- absent from GSD_OWNED_ENTRIES, and an
exact-name list does not match a dot-prefixed
name via the `gsd-` prefix rule
This is the SAME shape as the leak that motivated the prefixed-parent scan in
round 1 -- a gsd-prefixed child under a parent nobody had listed -- one parent
over. That it recurred is the argument for re-deriving this list from the
installer each round instead of trusting it: the enumeration is the weak point
of an enumerate-and-block mechanism, and it does not announce when it falls
behind.
Ownership is unchanged, only coverage: `hooks/` is shared with the host agent,
so only `gsd-`-prefixed children are watched. A test asserts a host-owned
hook is still ignored, because widening the parent list must not widen
ownership -- a guard that flags the host's own files gets switched off, and
then catches nothing at all.
Reverting the widening fails the new test.
The changeset for this PR is typed `Added`, and CONTRIBUTING requires a docs/
change for that type. The only docs/ file in the diff was CONTEXT-INDEX.json --
a GENERATED index -- so the Docs Required gate passed while no human-readable
documentation existed for either new variable. A gate satisfied by a generated
artifact is satisfied vacuously.
docs/TESTING-SUITES.md now carries a section on the guard: what it watches and
why it is ownership-scoped rather than whole-root, the two env vars in a table,
why the default is report-only and what the promotion condition is, and what
each violation label means (including that UNVERIFIED is not clean).
GSD_SKIP_LIVE_CONFIG_GUARD is named explicitly because it is a bypass on a
safety check. An undocumented bypass is one people eventually set without
knowing what they turned off.
A test asserts both variables appear in that doc -- checked as permitted by
local/no-source-grep before writing it, rather than assumed forbidden. It fails
when the section is removed, so the doc cannot rot back to the state the review
found.
Addresses review finding: Major 4.
qa-loop-walk runs `npm run test:qa`, which is `run-tests.cjs --suite qa` -- so it
runs the live-config guard like every other suite lane, and it set no
GSD_STRICT_LIVE_CONFIG_GUARD. A leak of exactly the class this PR closes would
have printed a warning there and left the lane green.
The test that is supposed to prove the guard is wired everywhere could not
detect that, because its job list was three literals (`test`, `test-full`,
`test-inert`). A hand-list certifying its own completeness is the defect this
whole PR is about, reproduced inside the test guarding the fix -- so the list is
now DERIVED from the workflow: every job with a step reaching run-tests.cjs,
directly or through an npm script resolved transitively through package.json.
The indirection is the load-bearing half; a grep for the filename alone is what
made qa-loop-walk invisible.
The derivation asserts a floor (>= 4 jobs) before ruling on any of them, so a
selector that silently matched nothing fails loudly instead of passing
vacuously. Windows lanes keep their carve-out, keyed on whether the job's matrix
mentions windows rather than on the job's name.
Negative-controlled: un-wiring qa-loop-walk fails the new test. The literal
version passed with that lane unwired, which is how it shipped.
Addresses review finding: Major 5.
RULESET.TESTS.boundary-coverage asks for {limit-1, limit, limit+1}. The depth
bound was exercised only at limit+2, which pins neither side of the edge: an
off-by-one that truncated a tree sitting exactly AT MAX_DEPTH would have passed,
and a truncation is not a cosmetic miss here -- it reports `unverified`, which
in strict mode fails the run.
Negative-controlled by weakening the guard to `depth >= MAX_DEPTH`: the new
limit case fails, where the previous single limit+2 assertion did not.
Addresses review finding: Minor 8.
MAX_ENTRIES was a single running budget threaded across every watch target. One
large early target exhausted it, and every target scanned afterwards reported
truncated -> `unverified` -- which under GSD_STRICT_LIVE_CONFIG_GUARD=1 is a
failed run. The guard's verdict therefore depended on directory iteration order
and on unrelated local state, neither of which says anything about whether the
suite leaked.
Each target now draws a fresh allotment, so a pathological tree truncates itself
and nothing else. MAX_TOTAL_ENTRIES keeps the aggregate bounded -- which is what
the single budget was actually for -- and when that ceiling engages, the targets
it curtails are still reported `unverified` rather than attested clean.
The limits are injectable so the boundary is testable without materialising
20000 entries, matching the `deps` seam the resolvers already use.
Two of the three new tests fail when the shared budget is restored; the third
asserts the retained global ceiling, which is deliberately unchanged behaviour.
Addresses review finding: Major 6.
resolveLiveConfigRoots resolves what THIS process sees, and getGlobalConfigDir
is env-first -- so with an ambient CLAUDE_CONFIG_DIR the guard watched that
path. A spawned child does not see it: TEST_ENV_BASE blanks the config-location
vars precisely so the child cannot follow them, and a blanked var is falsy, so
the child resolves its HOME-derived root instead.
A child that blanks the var and does NOT also sandbox HOME therefore writes into
the developer's real ~/.claude, which the guard was not watching. That is this
PR's own escape route, taken one process deeper -- and the guard is the artifact
that is supposed to make it loud.
Both resolutions are now unioned: the ambient one, and the fallback one obtained
by handing the REAL descriptor resolver an EMPTY env. Deriving it that way is
deliberate -- a hand-listed copy of the scrub set inside the guard is a second
list to drift, which is the defect this PR spent three rounds closing one layer
up. grok resolves through a hardcoded branch rather than a descriptor, so its
fallback is stated explicitly for the same reason it is named in the ambient
loop.
Addresses review finding: Blocker 3.
diffLiveConfig walked `after` alone, so it had no branch for a path that
existed before the run and does not after. A test run that DELETES a file from
the developer's real config dir passed the guard silently -- the least
recoverable case in the threat model this guard exists to cover.
The review named the missing `pre.exists && !post.exists` branch. That branch is
necessary and not sufficient: deletion arrives in two shapes and it reaches only
one of them.
- A FIXED owned entry (GSD_OWNED_ENTRIES x roots, plus every extra target) is
recorded at both ends whether it exists or not, so a deletion reads
{exists:true} -> {exists:false}. This is the shape the named branch fixes.
- A gsd-prefixed child is DISCOVERED by readdirSync, so a deleted one is
absent from `after` entirely and never enters an after-keyed loop at all.
The named branch is unreachable for it.
So the walk is now over the UNION of both key sets, with the explicit branch for
the first shape and an `!post` branch for the second. Both are covered by a
test, and reverting the fix fails both -- the prefixed-child test is the one
that would still fail with only the prescribed branch in place.
Addresses review finding: Blocker 2.
Found pre-push by this round's third adversarial review pass. Not a rebase
regression — round 3 shipped it and #2755 doubled it.
resolveExtraWatchTargets watches one config.toml per non-registry descriptor,
and its comment asserted "GSD writes ONE named file into these third-party
roots". That is false: bin/install.js also calls installSharedHooksBundle on the
same root, populating <root>/hooks/ with GSD's hook scripts and a CommonJS
marker. So a suite-produced leak of a hook bundle into a developer's real
~/.kimi or ~/.kimi-code passes this guard silently — #2665's own hazard, in
#2665's own safety net.
Behaviour is deliberately unchanged and the gap is disclosed instead. Closing it
is a layout decision rather than one more path, for the same reason
getGlobalSkillsBase is already a deliberate non-target: the snapshot applies the
config-root layout beneath every root it is given, and these roots are not ours.
Happy to fix it here or take it as a separate issue — the maintainer's call.
The enumeration-relative test could not have caught this: it asserts one target
PER DESCRIPTOR and nothing about whether one per descriptor is enough, because
its expectation is derived from the same array it checks. That is exactly the
scope boundary round-2 Nit 7 asked to be marked, biting one layer up from where
it was marked; the test now says so.
479979c4's message says "there are three" — that is three WATCHED targets, not a
count of write surfaces. The hooks bundle is a fourth, and unwatched.
lint:ci rc=0; tests/live-config-guard.test.cjs 24/24. Comments and catalog only.
Same drift as the CONTEXT.md seams, one layer over: #2755 took
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS from one entry to two, and five comments
across three files were left describing the one-entry world — "two live write
surfaces", "today's only entry", "today's single entry", and a <kimi>/config.toml
bullet naming only Kimi CLI's KIMI_SHARE_DIR.
The sharpest one was a wrong pointer rather than a stale count: run-tests.cjs
cited "scripts/lib/live-config-guard.cjs" for why the scope is narrow. That path
does not exist, and it names the one directory this module is deliberately NOT
in — the installer copies scripts/lib/ to users wholesale while uninstall removes
only an allowlist, which is the whole reason the guard lives one level up. A
reader following that pointer would have concluded the opposite of the decision.
Comments only; no behaviour change. lint:ci rc=0, tests/live-config-guard.test.cjs
24/24, tests/run-tests-harness.test.cjs 138/138.
The rebase onto next brought in #2755, which added a SECOND Kimi config
home — kimi-code's `~/.kimi-code`, overridden by KIMI_CODE_HOME — declared
as an inline object literal inside resolveKimiHooksTomlDir's body. That is
the resolvable-but-not-enumerable shape round 3 hoisted KIMI_SHARE_DIR out
of, so the hoist is extended to cover both descriptors rather than reverting
#2755's parameterization.
The scrub set was already complete: KIMI_CODE_HOME is declared in
capabilities/kimi-code/capability.json, so the registry rung covered it and
CONFIG_LOCATION_ENV_KEYS is 28 keys both before and after the rebase. What
was NOT covered is the guard — resolveExtraWatchTargets iterates
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, so with only one entry it watched Kimi
CLI's config.toml and never Kimi Code's. Targets go 2 -> 3.
The existing 'extra targets are DERIVED from the descriptor array' test
cannot catch this: it builds its expectation FROM the array, so removing an
entry shrinks the expectation with it. Verified — with the kimi-code
descriptor removed that test still passes while the new named test fails.
This is the enumeration-relative scope boundary the suite already documents
one layer down, biting one layer up.
Also rewrites NON_REGISTRY_OWNED_FILE's docblock, which asserted "today's
only such descriptor is kimi's ~/.kimi". There are now two, and its named
residual is load-bearing rather than vacuous.
None of the round-4 fixes had a test that fails on reversion: re-shipping
the test-instrumentation chain passes #2858 (everything ships, so every
require resolves), dropping the strict env from test.yml demotes the guard
to report-only with nothing red, and the skillsHome derivation tests are
enumeration-relative over declarations that are all empty today. One
guard each, every one negative-controlled against its reverted fix (fails
pre-fix, passes post-fix):
- packaging-shipped-scripts-require-only-shipped.test.cjs asserts the four
chain files are absent from the npm pack file list (reuses the tarball
set the #2858 gate already resolves — no second npm pack).
- live-config-guard.test.cjs asserts all three test jobs wire
GSD_STRICT_LIVE_CONFIG_GUARD, matching the WHOLE expression anchored —
a prefix match accepted both a Windows-silently-strict tail and a
malformed one.
- helpers-process-isolation.test.cjs cold-requires helpers.cjs in a child
with sentinel skillsHome env vars injected into both enumerations, so
the walk itself is under test rather than today's empty declarations.
Follow-up to 38c9395d, found while fact-checking the round-3 response rather
than by a test.
That commit made TEST_ENV_BASE derive its keys from
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, but had the guard call
resolveKimiHooksTomlDir directly. Both halves covered kimi, so nothing was
broken — but only one of them would pick up a SECOND descriptor. That is the
same partial-enumeration defect that put KIMI_SHARE_DIR outside the scrub set,
reintroduced one layer over, in the very commit that closed it.
resolveExtraWatchTargets now iterates the array and resolves each descriptor
through resolveConfigHomeFromDescriptor, so the scrub set and the guard derive
from one source and cannot drift apart.
Verified: a synthetic second descriptor is picked up automatically (it was not
before); kimi's target is unchanged on both the default (~/.kimi/config.toml)
and KIMI_SHARE_DIR override paths.
The new test asserts one target per descriptor plus the store root. The COUNT
is the load-bearing half — every per-descriptor assertion passes vacuously
today with a single entry, so only the count fails when the array grows and the
guard does not follow.
NAMED RESIDUAL, documented at NON_REGISTRY_OWNED_FILE: this assumes every
non-registry descriptor is written the same way (config.toml). A descriptor
whose owned file differs needs a per-descriptor mapping. It fails toward
under-watching rather than false positives, so it is called out rather than
left to be discovered.
Round 2, both Majors. They are one defect seen twice: the recurrence guard did
not cover the surface it exists to guard.
Blind surfaces. resolveLiveConfigRoots enumerates getGlobalConfigDir per registry
runtime plus a hardcoded grok branch, so it can only ever see runtime config
ROOTS. Two live write surfaces are not roots and passed through silently:
$GSD_HOME/.gsd — GSD's user-owned store. Watched WHOLESALE: unlike ~/.claude
this root is exclusively ours, so the shared-root
false-positive trap the module documents does not apply.
<kimi>/config.toml — the file GSD writes its native [[hooks]] block into. The
INVERSE case: ~/.kimi belongs to Kimi CLI, so only the one
file GSD writes is watched, never the root.
That asymmetry is why this is not a two-line "add two roots" patch — one target
needs the whole tree, the other needs exactly one file, and collapsing them
either under-watches the store or trips the guard's own documented
false-positive trap on a third party's directory.
Extras are passed to snapshotLiveConfig explicitly rather than resolved inside
it, so a caller snapshotting a fixture root cannot silently pull the developer's
real ~/.gsd into its own assertions. run-tests.cjs now snapshots when EITHER the
roots or the extras are non-empty — previously an unbuilt tree yielding zero
roots disabled the entire guard without saying so.
Budget coverage. The MAX_ENTRIES/MAX_DEPTH bound and the truncated -> 'unverified'
branch had zero tests, despite this module's own docstring naming "a truncated
scan reading as clean" as the safety-critical case. Added per
RULESET.TESTS.boundary-coverage (N in {limit-1, limit, limit+1}, exercised
through newestMtime's injected budget so the boundary is real without
materialising 20000 files) and RULESET.TESTS.property-based-testing (fast-check:
truncation is monotone in the budget; reported newest never exceeds the true
maximum). A regression flipping `truncated` to false on an exhausted budget now
breaks the property for every budget below the tree size.
Negative-controlled: neutering the extras wiring fails exactly the two
new-surface tests and nothing else. 21/21 green with it restored.
Its first CI run found PRE-EXISTING leaks on the Windows lane —
C:\Users\runneradmin\.claude\gsd-core and skills\gsd-dev-preferences — with all
1196 Windows tests otherwise passing. os.homedir() reads USERPROFILE on Windows,
and ~190 test sites across 31 files sandbox HOME alone, so the suite has been
installing GSD into the runner's real home directory invisibly. That is exactly
the class the guard exists to surface, and exactly the class this PR's review
said CI could never catch.
It is also a different defect from the one #2665 closes, and too large to fold in
here. A brand-new gate that immediately reds an unrelated lane gets bypassed or
reverted rather than obeyed, so the guard reports by default and fails only under
GSD_STRICT_LIVE_CONFIG_GUARD=1.
This is the repo's own established ratchet, not a hedge: the local/no-source-grep
ESLint rule shipped at `warn` and was promoted to `error` after its cleanup sweep
(ADR 452). Promote this the same way once the USERPROFILE sweep lands.
The recurrence guard, and #2665's own "Optional hardening". This class is silent
by construction: TEST_ENV_BASE cannot see an in-process caller, and CI cannot see
the class at all because CI never has these env vars set. It damages the
developer's machine and reports nothing -- which is how two prior authors each
diagnosed it and fixed only the instance in front of them.
run-tests.cjs snapshots GSD's install footprint in every live runtime config dir
before the suite and re-checks it after, failing the run on a create or a modify.
Roots come from the product's own getGlobalConfigDir, so the guard watches
wherever the product actually points, including through an ambient var.
Scope is ownership-based, not whole-root: the top-level install footprint plus
gsd-prefixed children of dirs GSD shares with the host agent. A config root like
~/.claude is shared, and watching it wholesale would false-positive on the host's
own history.jsonl or settings.json -- a guard that cries wolf gets disabled, and
then catches nothing. The prefix test is load-bearing: the first version watched
only the three top-level entries and MISSED a real leak into skills/gsd-*.
It earned its place immediately -- it is what found the fifth in-process leak in
runtime-artifact-layout.test.cjs, which no amount of reading the review would have
surfaced. Known gap documented in the module: a write to a file GSD does not own
is out of scope by construction.
Lives in scripts/, deliberately NOT scripts/lib/ -- the installer copies that dir
into every user's config dir wholesale while uninstall removes only an allowlist,
so a test-only module there would ship to users and survive uninstall.
Addresses review finding: Minor 8.