Commit Graph

5090 Commits

Author SHA1 Message Date
Tom Boucher
b9f51836e6 refactor(#3180): ADR-3180 behavior contract + cross-surface drift guardrails (#3223)
* refactor(#3180): one owner for completion ratio, a prompt-layer drift guard, and a written behavior contract

The 2026-08-08 coverage audit on #3180 found the epic's copy counts were a
lower bound for the third consecutive time, and that two derivation families
had never been named at all.

ADR-3180 gains Decision 7 — a normative behavior contract that says what the
right answer IS for each derivation, not merely who owns it. A reviewer with
no written rule can only ask "does this look like the others", which is how a
fifth copy passes review. Decision 4 gains (d) scan surface is every authored
surface and an owner FILE is never exempt, only its named functions; and (e)
a surface that cannot be consolidated today ships ratcheted, never unguarded.

Completion ratio: `clampPercent` sat exported and unused beside six hand-inlined
copies of its own body across five modules. All six now route through it;
`clampPercentFromFraction` is added for the one caller that already held a
fraction. Every migration is behaviour-identical — clampPercent's first line IS
the `total > 0 ? … : 0` ternary each copy carried. Guarded by
lint-completion-ratio-drift.cjs, which reports zero re-derivations with no
file-level exemption.

Prompt layer: workflow markdown re-derives live-plan counting in raw shell
(#1762), invisible to every `src/`-scoped guard. lint-planning-prompt-drift.cjs
scans it with a shrink-only baseline of the 7 sites that exist today — new
sites fail, and a baseline entry that stops firing fails too, so an
acknowledgment can never outlive the thing it describes.

lint-milestone-window-drift.cjs stops exempting its owner file wholesale; only
the four named canonical functions are exempt now. The blanket exemption was
pointed at the one file most likely to grow the next copy, and it had.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3180): link Phases 6-8 sub-issues (#3216, #3217, #3218) from ADR-3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): address orthogonal review — consumer-output identity tests, count-keyed ratchet, property coverage

Five findings from the two orthogonal review passes, all fixed.

Decision 4(c) breach: the completion-ratio identity test asserted at the
OWNER, which is exactly the bypass that decision exists to close — a consumer
can call clampPercent and then post-process locally, leaving both the lint and
an owner-level test green. It now drives `roadmap analyze`, `query progress`
and `stats` and asserts on their own output, over a fixture containing a
`status: superseded` plan so a consumer that re-counted raw files would report
60 where the owner reports 75.

Decision 4(e) breach: ratchet entries named the epic (#3180) rather than the
issue that removes them. They name Phase 8 (#3218) now.

The ratchet keyed on (file, text) alone, so plan-phase.md's two byte-identical
sites were one indistinguishable key and migrating either would have left the
guard green with the other alive. Entries carry an occurrence count; fewer than
acknowledged fails as a partial migration, more fails as a new copy.

Adds the missing MAX_REGEX_LITERAL_LEN boundary coverage the sibling guard's
test already had, and the fast-check property tests CONTRIBUTING requires for
clamp/budget-limit functions.

Refs #3180

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop wrapping a nested double-spawn in a 15s wall-clock budget (bug #641 probes)

`tests/ci-test-scope.test.cjs`'s `bug #641` block spawned `run-tests.cjs`
under PROBE_TIMEOUT_MS=15000; that child then spawned a nested `node --test`.
A fixed wall-clock budget around a double spawn, running inside a container
that is concurrently executing the full ~31k-test suite, fails by construction
under load.

Confirmed against three full matrix runs. Every failure was shaped
`null !== 0` — the child was KILLED, never an assertion about the thing under
test. One captured probe had already printed the correct resolution
(`suite="all" files=2: a.test.cjs b.test.cjs`) and was killed anyway. It
reproduces on `next` alone: 5 failures on linux-node22, 0 on linux-node24. The
victim subset varies by run and by lane.

What these tests are actually about is suite-token RESOLUTION — `unit` as a
bare token in --files/--files-from. Executing the seeded trivial files is
incidental and is the entire timeout surface, so the assertions move
in-process against the same functions `main()` calls, in the same order.
`parseArgs`, `selectExplicitFiles`, `selectFiles` and `walkTestFiles` are
exported for that; no behavior, signature or logic changed.

No coverage lost: `tests/run-tests-harness.test.cjs` already spawns the
harness for real and asserts exit codes end to end, on a 120s budget.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: delete the three elapsed-time assertions

CLAUDE.md forbids asserting on wall-clock time. Three assertions did, and all
three are load-sensitive: on a saturated bench each can fail while the code
under test is correct. In every case the load-bearing assertion sits on the
line above and the timing line adds no discrimination.

run-with-timeout: the stated worry — "was this 124 the cap firing or the 30s
harness backstop?" — is already answered by the assertion above it. A backstop
kills by signal, which surfaces as status null, never 124. Observed directly
this session: three matrix runs produced exactly that null shape from killed
children.

normalize-test-command and context-predicates: both bounded a ReDoS check.
A threshold only ever separates "fast" from "slightly slow", which is bench
load, not correctness — catastrophic backtracking on 800 KB of input does not
take 251ms, it does not finish at all. A real regression therefore shows up as
the suite being killed on that test, which is louder and more reliable than a
number. The structural assertions (returned unchanged; cleanly rejected) are
what actually carry those tests, and they stay.

The sweep now reports zero elapsed-time assertions in tests/. The remaining
Date.now() uses are unique-path suffixes, barrier deadlines, fixture
timestamps and fake mtimes — none of them assertions.

Pre-existing on `next`, fixed here rather than deferred.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3180): backfill changeset PR number (#3223)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3180): key the prompt-drift ratchet on POSIX paths so it works on Windows

The baseline keys on (file, trimmed text). `file` came from scanTree's
`path.relative()`, which uses NATIVE separators, while the committed baseline
stores POSIX. On Windows every violation was therefore unmatched — reported as
FRESH — and every baseline entry matched nothing — reported as STALE. The guard
failed 100% of the time there, on both CI shards:

  ✖ scanRepo(repoRoot) matches the baseline exactly: zero fresh AND zero stale
    + { file: 'gsd-core\\workflows\\execute-plan.md', ... }

The remote runner this repo gates on is Linux-only and cannot see this class at
all; the GitHub Actions Windows lane is what caught it.

Normalization is unconditional — never gated on process.platform. A
platform-conditional normalizer makes the POSIX path the special case and
leaves the Windows branch unexercised on every other OS, which is the same
blind spot in a different place. It is applied at one seam inside
findPromptDrift, which builds `file` on every returned violation, so the
baseline key, the --update writer, the stderr report and the tests all consume
one normalized value.

The regression tests drive a Windows-shaped relPath directly and run on every
OS rather than skipping off-Windows — a test that only runs on the platform
where the bug lives is why this escaped. They include a sanity check that
un-normalized input does NOT match, so the assertion cannot pass vacuously.

Audited the three sibling guards: none keys against a committed cross-platform
baseline, and their exemption keys are path.join-built, so producer and
consumer share the native convention. Left correct code alone rather than
making them look alike. scripts/lib/drift-scan.cjs is untouched — normalizing
there would break those three on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 16:05:17 -04:00
Tom Boucher
636ec92107 refactor(#3185): phase enumeration has one owner and a decidable scope (#3222)
* test(#3185): failing-first phase-enumeration single-owner suite

Covers the enumeration rows with direct code evidence: 999.* backlog dirs
listed by progress/stats, the phase-0 sentinel divergence, the #1324
letter-prefixed-decimal negative space, and the destructive-path find —
cmdPhasesClear carries a fifth sentinel copy (/^999(?:\.|$)/) that excludes
999 but not 0, so a 0-* directory roadmap.analyze preserves is deleted there.

Also covers the pass-all degrade, which is where the defect actually lives:
when the milestone window declares no phases the filter becomes a literal
() => true and its heading-side sentinel exclusion is unreachable. A fixture
carrying phase headings keeps the filter active and never reaches that path.

Named for the derivation, not a module: the suite drives commands, phase,
milestone, workstream-inventory and state, and both the phase and
phase-locator buckets are already at the per-module test-file cap.

Committed alone so the remote runner records the failure before the fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): phase enumeration has one owner and a decidable scope

Adds phase-locator.cts::listMilestonePhaseDirs as the single canonical owner
of "which phase directories belong to the current milestone". It applies the
milestone window AND the sentinel filter and returns a ScopedResult, so a
caller can tell a genuinely-empty milestone from an enumeration that could
not be scoped.

The sentinel test now runs against DIRECTORY NAMES and is unconditional.
getMilestonePhaseFilter excludes sentinels from its ROADMAP heading set, but
degrades to a literal () => true pass-all predicate when that set is empty --
at which point the heading set is never consulted and its sentinel exclusion
is unreachable exactly when it is needed. That degrade is the #3167 path, and
it is why stats already used the filter and still listed backlog directories.
The narrowing is sentinel-only: pass-all stays over-inclusive otherwise.

Sentinel copies deleted, canonical isSentinelPhaseId adopted:
  - cmdRoadmapAnalyze's local closure (parseInt === 0 || === 999), 2 call sites
  - cmdPhasesClear's /^999(?:\.|$)/ -- the DESTRUCTIVE path, which excluded
    999 but not 0, so a 0-* directory roadmap.analyze preserves was deleted

cmdStats also seeded rows from ROADMAP headings with no sentinel filter, so a
999 heading produced a row with no directory; that seed is filtered now.

cmdPhasesList routes only its ENUMERATION. --phase lookup searches the
physical set (scoping it would report an out-of-window phase as not found) and
--include-archived still merges archived dirs (they are by definition from
other milestones). Both exempt by documented reason, never a file allowlist.

Fixed inline, found while building: isDirInMilestone could not match a #1324
letter-prefixed-decimal directory (P0.0-foundation) to its own Phase P0.0
heading, so stats reported the phase with plans: 0 while its directory held
plan files. Defers to phase-id's extractPhaseToken rather than widening a
fourth bespoke regex; additive, so it can only admit directories.

Refs #3180. Closes #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): route the last two enumeration re-derivations

workstream-inventory countRoadmapPhases counted every `Phase` heading across
the whole ROADMAP -- no window, no sentinel filter -- so it counted 999.*
backlog and Phase 0 and spanned every milestone the document ever had. Its own
caller already resolved a currentVersion and passed it to getMilestonePhaseFilter
elsewhere in the same file; this was the sibling copy that never got the fix.

state.cts phaseInventoryProvider enumerated phase dirs with its own
/^(\d+)-(.+)$/ convention regex and neither filter, so a rebuilt STATE.md
inventory carried backlog and sentinel directories as current-milestone phases.
A non-COMPLETE enumeration scope now throws to the outer catch as a real scan
failure rather than reporting a confident undercount, mirroring the per-phase
scanPhasePlans contract beside it.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): consolidate 23 sentinel re-derivations onto one predicate

The whole-repo drift guard (ADR-3180 Decision 4a, no file allowlist) found the
sentinel rule re-implemented 23 times across 8 modules, in three regex variants
plus four integer-comparison forms. Most tested 999 only, so Phase 0 slipped
through them while roadmap.analyze and the engine-wide convention (#1580) both
treat 0 and 999 alike. That disagreement is the defect class this epic removes.

All 23 now call phase-id's isSentinelPhaseId (SENTINEL_RANGES [0,999]). Sites:
init recommended-actions and backlog counts, milestone phase scan, the
phase-lifecycle progress table, phase.cts used-number collection and the four
renumber-on-remove guards, roadmap-parser's heading and bullet milestone
counts, roadmap get-phase fallbacks, and state's heading denominator.

Excluding Phase 0 at these sites is a deliberate behavior change and the point
of the consolidation — several carried comments already saying 0 should be
excluded while the literal beside them caught only 999.

Adds scripts/lint-phase-enumeration-drift.cjs, wired into lint:ci. It scans the
whole src/ tree with no file allowlist and reports both shapes: an independent
phases-dir enumeration, and an independent sentinel literal. Exemptions are
function-scoped with a written reason. The guard is comment-aware — its first
pass flagged JSDoc and a comment documenting that the code below uses the
canonical owner, which would have trained readers to exempt prose.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): resolve every phases-dir enumeration; drift guard reports zero

Per-site triage of the 31 remaining whole-repo guard hits, applying the rule
generalized from #3183's Amendment 1: a LOOKUP, DIAGNOSTIC, ARCHIVAL or
MUTATION pass wants the physical set; only "which phases belong to this
milestone" wants the scoped set.

Routed (10): init new-milestone phase_dir_count, init milestone-op fallback
count, init manager, init progress, milestone complete stats/dry-run/archive
move, phase complete's next-phase scan, state update-progress, state
frontmatter stats, and uat audit's active set.

Exempt with a written function-scoped reason (never a file allowlist): the
audit/UAT/verification sweeps that deliberately scan every directory to report
gaps, phase create/insert/rename/renumber mutations, single-phase lookups,
roadmap-upgrade's cross-milestone migration, cmdPhasesClear's whole-tree
destructive pass, and the reads that list a phase dir's FILES rather than
enumerating the phases dir at all.

Latent defects fixed by the routing: sentinel directories leaked into
cmdInitNewMilestone's phase_dir_count, cmdMilestoneComplete's stats, dry-run
AND ARCHIVE MOVE, cmdStateUpdateProgress, buildStateFrontmatter and
cmdAuditUat's active set — every one of those hand-rolled an isDirInMilestone
filter with no sentinel exclusion, so `milestone complete` was archiving
backlog directories.

scripts/lint-phase-enumeration-drift.cjs now reports 0 re-derivations and
npm run lint:ci is green.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* docs(#3185): document milestone-scoped enumeration and record ADR Amendment 3

Changeset fragment (Changed), CLI-TOOLS/COMMANDS/USER-GUIDE updates for the
scoped output of progress, stats, phases list, phases clear and milestone
complete, the CONTEXT.md Phase Locator glossary entry naming
listMilestonePhaseDirs, and ADR-3180 Amendment 3.

Amendment 3 records: the SCOPE contract held unchanged; the declared deviation
from Decision 1's provisional signature (the window needs cwd/ws, which the
locked roadmapContent parameter cannot supply); the copy count being a lower
bound for the third consecutive phase (4 scoped vs 54 found); the load-bearing
finding that the sentinel exclusion sat on the heading set and was unreachable
under the pass-all degrade; the two destructive-path defects; and the
generalized exemption rule.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): wire scope to consumers; revert two wrong routings the suite caught

Review + remote runner findings, all fixed:

The three consumers computed the enumeration scope and threw it away, so
TRUNCATED/UNSCOPED/UNREADABLE collapsed into the same output as COMPLETE --
reproducing this epic's own output-identical-failure defect one layer up.
progress, stats and phases list now emit phase_scope (null on the phases list
--phase lookup path, which performs no enumeration).

Two routings were wrong and the suite proved it:

roadmap-parser's two milestone phase-count scans are reverted to the 999-only
literal. isSentinelPhaseId is BROADER than what it replaced: its legacy branch
runs /^0*(\d+)/ over "00.1", which backtracks to capture 0, so it read #2554's
decimal phase ids as sentinel milestone 0 and stopped counting them.

state.cts phaseInventoryProvider is reverted to the physical disk scan.
`state rebuild` is a RECONCILIATION pass -- scoping it made it throw on healthy
trees whose fixture resolves no window, swallowed the raw readdirSync fault
message #3057 B1 requires verbatim, and stopped it dropping orphan STATE.md
rows, which is the job.

Both are now function-scoped guard exemptions with written reasons, not
silent reverts. This is the consolidation trap named in the epic: a canonical
rule can cover MORE than the copy it replaces, and only real inputs show it.

Adds phases list coverage, a scope-branch test, and a drift-guard unit suite;
backports comment-awareness to the milestone-window and plan-count guards so
all three siblings share one false-positive profile; names #3161 alongside
#3167 in Amendment 3's subsumption record.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): correct isSentinelPhaseId's decimal-zero misclassification

An isolated security review caught this branch committing the epic's own sin:
the over-broad predicate was worked around at ONE call site and left live at
the destructive ones.

isSentinelPhaseId's legacy branch ran /^0*(\d+)/, which backtracks so any id
whose leading digit run is all zeros before a non-digit captures 0 -- "0.1",
"00.1" and "0.2554" all read as sentinel milestone 0. Two pinned contracts
disagree with that: #2554 requires "00.1" to be counted as a real phase, and
the 999 icebox is a whole reserved milestone so "999.1" must stay sentinel.

The rule is asymmetric and now says so explicitly: 999 is sentinel with or
without a decimal part; 0 is sentinel only when bare. A decimal phase under
either is a real phase for 0 and reserved for 999, because 999 reserves a
MILESTONE while 0 reserves a PHASE.

Fixing the owner lets the earlier workaround go: getMilestonePhaseFilter's two
scans route through isSentinelPhaseId again and the guard exemption that
existed only to accommodate the defect is deleted. The state.cts cmdStateRebuild
exemption stays -- that one is a genuine reconciliation-wants-the-physical-set
case.

Also corrects tests/adr-612-bracket-grammar.test.cjs, which asserted
isSentinelPhaseId('0.1') === true and so had encoded the defect as expected
behavior.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): keep isSentinelPhaseId's semantics — 0.x is layered, not wrong

Reverts the previous commit. The remote suite failed six tests proving it
wrong, and the reason is the sharpest finding of this phase.

An isolated security review observed that isSentinelPhaseId reads 0.1 and 00.1
as sentinel milestone 0 and judged that a defect against #2554. Correcting the
canonical predicate broke #2949. Both contracts are pinned and both are right,
because they ask different questions:

  #2554  is this dir part of the current milestone's phase SET?  -> count 00.1
  #2949  must this phase COMPLETE before the milestone closes?   -> 0.x sentinel

No single global predicate answers both. isSentinelPhaseId keeps its semantics
(0.x IS a sentinel, #2949), and the milestone-window layer keeps a narrower
999-only rule (#2554) as a function-scoped guard exemption with a written
reason — not a second silent copy.

That corrects how Decision 1 reads: "one owner per derivation" governs who
computes an answer, not how many questions share it. An over-broad canonical
rule is as much a defect as a divergent copy and fails worse, because it looks
like consolidation. Recorded in Amendment 3 as the lesson for Phases 4 and 5.

Where a review's inference about intent conflicts with a pinned contract, the
pinned contract wins; the finding is adjudicated, not fixed.

The boundary tables in the enumeration suite are corrected to assert 0.x IS a
sentinel, with the layering explained.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* chore(#3185): set changeset fragment pr to 3222

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 14:22:10 -04:00
Tom Boucher
bc3a3f170f docs(#3215): add ADR-3212 lexical seam consolidation (Phase 0) (#3221)
Design-lock ADR for epic #3212 — the lexical layer beneath the #1372
markdown-sectionizer and #2143 table/mutation seams.

Locks seven decisions Phases 1-4 execute against:
- src/pattern.cts as sole owner of dynamic regex construction,
  delegating to the built-in RegExp.escape; the ten private
  escapeRegex/escapeRegExp/escapeRe copies are deleted, not merged
- engines.node floor raised to the Active LTS line (>=24), which is
  what makes RegExp.escape reachable; delegating-shim alternative
  recorded and rejected
- src/text-lines.cts as sole owner of line-terminator handling —
  the primitive the existing no-crlf-fragile-split prohibition lacks;
  brings frontmatter.cts in from the #1372 exclusion
- tokenizer-first for stateful grammars, with a decidable five-condition
  test, generalizing the proven hooks/lib/git-cmd.js token-walk (#3129)
- bounded quantifiers over caller-supplied content
- extend-never-mutate (inherited from ADR-2143 §2)
- prohibition with teeth: no-adhoc-regex-escape, no-unbounded-quantifier,
  no-crlf-fragile-split widened to src/, plus a parity assertion

Explicit non-goal: no wholesale regex-to-parser rewrite. A census found
2,113 regex literals across 317 files; most are correct and stay.

Docs-only. No production code.

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 13:29:45 -04:00
Tom Boucher
63a677aed4 docs(#3198): record project-gitignore-for-third-party-agent-runtime as out-of-scope (#3219)
Records the 2026-08-08 triage denial of #3198 in the .out-of-scope/ knowledge
base so the prior-denial check on a future triage sweep can find it, rather
than the reasoning living only in a closed-issue comment.

The request asked to add .omc/ and .omo/ to "the init-template .gitignore".
gsd-core ships no such template: the __pycache__/*.pyc pair quoted in the
report is gsd-core's own repo .gitignore, and at init the only project
.gitignore write is a single .planning/ line under commit_docs = No. The
.omc and .omo directories have zero occurrences anywhere in the repo.

The entry's "What this does NOT cover" section keeps adjacent asks live --
notably any report that gsd-core itself writes churning runtime state outside
.planning/, which would be a real defect and is not denied here.

Closes #3198

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 13:08:08 -04:00
Tom Boucher
e705652ba1 Merge pull request #2677 from 0xdhx/fix/2665-test-env-base-config-location-vars
fix(#3156): derive the config-location scrub set and close the leaks it cannot reach
2026-08-08 11:25:44 -04:00
Tom Boucher
8f437cfb49 Merge pull request #3205 from 0xdhx/fix/3174-quick-verification-status-query
fix(#3174): read quick's verification status via the verification.status query
2026-08-08 11:25:41 -04:00
Tom Boucher
66a4940d6f Merge branch 'next' into fix/2665-test-env-base-config-location-vars 2026-08-08 10:47:09 -04:00
Tom Boucher
b421e95434 Merge branch 'next' into fix/3174-quick-verification-status-query 2026-08-08 10:47:06 -04:00
Tom Boucher
342590c70e refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite

Covers the 50 input classes in the phase test matrix: scope classification
(genuinely-empty vs truncated vs unscoped vs unreadable), the section-end
owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c),
the milestone.complete refusal with negative proof that no directory moved,
the version-token boundary defect, drift-guard behavior, and three fast-check
properties over document-shaped generators.

Committed alone so the remote runner records the failure before the fix lands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* refactor(#3184): milestone windowing routes through one owner

Three copies of the milestone section-end walk lived in roadmap-parser.cts —
two distinct computeSectionEnd function nodes plus an inline third in
getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is
now the sole owner and the other two are deleted, not kept in sync by comment.

The whole-repo drift guard found what the epic did not: state.cts held three
more re-derivations of the same vocabulary — two byte-identical milestone
bounding checks carrying a defect neither reported copy has (no boundary after
the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning
predicate. All three route through the owner now.

A composition-level duplicate appeared inside this change's own first pass:
getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out
of the owner's primitives, and had already diverged on whether to skip a closed
milestone heading. sliceMilestoneWindow is the one composition.

Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is
distinguishable from a genuinely empty milestone — those were output-identical,
which is the whole failure class. roadmap analyze emits it (#3165), and
milestone complete refuses to archive on anything but COMPLETE rather than
pass-all moving every phase directory on disk (#3166). The pass-all degrade is
preserved where its premise holds: making the filter deny-all would trade a
silent over-inclusive answer for a silent under-inclusive one on the read paths
that count with it.

extractCurrentMilestone keeps its signature — 200+ affected symbols across 41
files and 25 process flows — and is a one-line wrapper over the scoped owner.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): fence-aware phase detection and one heading-selection owner

Review fixes from the two orthogonal passes.

The blocker: hasPhaseEntries matched ATX phase headings fence-aware via
tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown,
so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely
empty milestone then classified TRUNCATED and milestone complete refused a
legitimate archive — a false positive in the destructive direction, worse than
the defect this phase set out to fix. Both that path and getMilestonePhaseFilter
own pre-existing bullet scan now run on stripFencedCode, since leaving one meant
the owner file gave two different answers to the same question.

The selection rule — locate, prefer the non-closed heading, else the first — had
been written three more times inside the file whose thesis is single ownership.
selectMilestoneHeading owns it; all three sites route through it. The copies were
behaviorally identical, so this is de-duplication with no observable change,
verified by probing that all three paths select the same heading.

roadmap analyze emitting a scope no consumer read left #3165's actual symptom
alive, so Route 0 in next.md now treats a non-complete scope as scan-failed
rather than as a clean empty scan, and the ADR amendment no longer overstates
what shipped.

Also: the scope refusal moved above the archive-directory create, so a refusal
leaves nothing on disk; the versionOverride comment names all four consumers;
COMMANDS.md documents the new guard beside its sibling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#2658): exclude the changelog from the malformed-path scan

The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts
none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships
into that tree, and its #2658 entry quotes both malformed paths while describing
the fix that removed them — so the release note documenting the fix trips the
fix's own regression test. Red on next before this branch.

The installer is correct: a probe over a real --trae --local install found 621
emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope
was the defect, not the product.

Excluded by exact relative path rather than by loosening the patterns or skipping
all markdown — the emitted agent and command markdown is precisely what #2658 was
about, so the gate stays strong everywhere it matters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#3184): regenerate install-tree fixtures for the shared drift scanner

scripts/lib/ ships in the npm package and installer, so extracting the shared
tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's
install tree. Regenerated via npm run gen:install-tree; the delta is exactly
that one path per fixture.

The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded),
so only the extracted library moves. This matches the existing
scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only
helper carried in the shipped tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal

The remote runner caught two regressions this branch introduced. Both were mine,
and neither review pass found them — only running the existing suite did.

The version-token boundary. I replaced locateMilestoneHeadings' \b with
(?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562
fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone
state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a
closed v8.0-A sibling (#730), and \b is what allows it while the stricter
boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to
\b; the state.cts consolidation is now a straight merge with no behavior change,
and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and
the design doc no longer claim otherwise.

The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is
about the TRUNCATED window specifically — the heading is found and the section
closes before the phase region, so pass-all archives everything. UNREADABLE and
UNSCOPED are pre-existing, legitimately handled states, and refusing on them
broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed
to TRUNCATED; docs corrected to match.

One of the new tests was also wrong: its fixture gave the shipped and current
milestones' phases the same numeric id, and the filter matches on that id, so it
could not have distinguished the two windows. Fixture corrected to exercise what
it claims to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* fix(#3184): enumerate drift-scan.cjs for uninstall

The installer copies scripts/lib/ wholesale, but uninstall removes an explicit
set — deliberately, so a user's own helpers in that directory survive. The
extracted drift-scan.cjs was copied in and never enumerated, so it outlived
uninstall, left the directory non-empty, and the rmdir that follows failed.

Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is
likewise a lint-only helper that ships there and is enumerated. Verified with a
real install-then-uninstall into a temp target: scripts/lib/ held exactly the
three GSD files and was gone afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset

Found while shipping this phase, and fixed here rather than noted.

install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE —
the comment at the copy site literally says "and any future lib helpers".
uninstall() removes them by hardcoded enumeration, deliberately, so a user's own
helpers in those directories survive. A wholesale writer paired with an
enumerated remover cannot stay in sync by construction: any file added to either
directory ships to every user and is then orphaned in their repo forever, since
it survives uninstall, leaves the directory non-empty, and the rmdir that follows
fails. Nothing reported this. 31,225 tests were green over it.

That is the same divergence class this epic exists to delete, sitting in the
installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity
assertion that fails the moment the two surfaces disagree. The test compares each
directory's real contents against its enumeration and names the offending file
plus the constant to add it to.

Both enumerations are hoisted to module scope and exported, so the test asserts
on the actual arrays rather than pattern-matching the installer's source — no
allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on
the current tree, correct report when an unenumerated file is injected.

scripts/changeset/ turned out to carry the identical defect and is covered too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

* chore(#3184): backfill changeset PR number

Also narrows the wording to match the shipped behavior: the refusal fires on a
truncated window specifically, not on any non-complete scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 09:50:35 -04:00
0xdhx
f3ce2dbab5 docs(#3156): name the isolation this helper does NOT provide
Found by the pre-push adversarial review of this round, and worth recording in
the code rather than only in the PR thread.

installSpawnHome() creates one sandbox home per test-FILE process, not one per
spawn, so two installer spawns in the same file share .gsd state. The
containment claim is unaffected -- nothing reaches the developer's real home --
and it is strictly better than the status quo it replaces, which shared the real
home and every byte of its state. But "contained" and "isolated from each other"
are different properties, and only the first is claimed.
2026-08-08 06:18:09 -05:00
0xdhx
ee266c318f chore(#3174): set changeset fragment pr to 3205 2026-08-08 06:06:44 -05:00
0xdhx
78330e505c fix(#3174): read quick's verification status via the verification.status query
`gsd-core/workflows/quick/steps/quick-verification.md` read the verifier's
result with `grep "^status:" F | cut -d: -f2 | tr -d ' '` and routed it through
a table whose only arms were passed / human_needed / gaps_found. That read
fails two ways. Driven against the old pipeline:

  never written (verifier died)      -> empty        -> no arm
  off-schema value                   -> weird_value  -> no arm
  `status:` in frontmatter AND prose -> two lines    -> no arm
  valid `passed` on a CRLF checkout  -> passed\r     -> no arm
  stale report still reading passed  -> passed       -> SUCCESS
  `status: passed` in the prose only -> passed       -> SUCCESS
  off-schema `passed:bogus`          -> passed       -> SUCCESS

The first four leave the orchestrating agent improvising at the moment the
pipeline failed. The last three are silent false passes: staleness was never
evaluated, the match was not anchored to frontmatter, and `cut -d: -f2` splits
an off-schema value at its own colon. The CRLF row is a pre-existing Windows
bug this change closes as a side effect.

The unanchored match is DEFECT.FRONTMATTER-SCALAR-BROAD-GREP, which the code
side already fixed by name — `readVerificationStatus` parses frontmatter only,
anchored at byte 0, and is total over its input space, returning `missing`,
`unknown` and `stale` sentinels. execute-phase.md, verify-work.md and
progress.md all read this same artifact through that query already; quick was
the remaining second mechanism.

Route quick's read through it and add an explicit terminal arm.

Three details a naive swap misses:

- The step file carries the runtime shim bootstrap itself. Step files are read
  and executed as their own units, so quick.md's bootstrap does not reach here.
  Copied byte-identically from gsd-core/workflows/_runtime-launcher.snippet.sh,
  the source sync-runtime-launcher.cjs generates every workflow's copy from.
  Without it the call resolves to nothing, 2>/dev/null swallows the error, and
  the fix degrades to a permanently-taken recovery arm.

- No jq. `--pick status` returns the bare value. Per #2589 a `| jq -r` pipe
  yields an EMPTY variable with no diagnostic wherever jq is absent — the
  Windows/Git-Bash default — which here would route a passing verification into
  the recovery arm, strictly worse than the grep being replaced.

- $VERIFICATION_STATUS is a DISPLAY string ("Verified" / "Needs Review" /
  "Gaps") consumed at quick.md:619 and quick.md:684, not the raw status. The
  raw value lands in $STATUS and the new arm sets both, so the failure path
  does not emit an empty index-table cell.

next_action / next_command are deliberately not surfaced. readVerificationStatus
discovers and parses shape-agnostically, which is what makes the status half
correct for ${QUICK_DIR}; but it also reads the directory basename as a phase
token to build those commands, and a quick dir is `${quick_id}-${slug}` with a
date-derived quick_id — so the projection carries the date as a phase argument.
Quick supplies its own recovery actions instead.

Adds tests/fix-3174-quick-verification-status-read.test.cjs under
`allow-test-rule: source-text-is-the-product` (CONTRIBUTING.md's exception
matrix; the pattern tests/verify-work-auto-transition.test.cjs already uses for
verify-work's status-query ordering). It pins five properties: the query
replaces the grep, the bootstrap precedes the call, the bootstrap matches the
canonical launcher snippet, the status-read fence is jq-free, and the terminal
arm names all three sentinels and sets the display string. Verified as a
negative control against pre-fix next: 0/5 pass there, 5/5 here.
2026-08-08 06:01:46 -05:00
0xdhx
6ac4d2e6ab fix(#3156): bound the cold-require probe the rebase turned into a violation
Post-rebase validity finding, not a new defect. The base range added
local/no-unbounded-spawn (#3143, 2afe17bb) and then DELETED its allowlist
(#3148, 9faacc0c), so the cold-require probe this PR added in 4bc6b0a2 --
legal when written -- is now an error. Bounded at 30s: a cold require is
sub-second, so that is ~30x headroom and still fails loudly rather than
hanging a lane.

Also routes this round's own teardown through cleanup() instead of raw
fs.rmSync, per local/no-raw-rmsync-in-tests, which carries the Windows-EBUSY
retry budget.

eslint clean on the file.
2026-08-08 05:58:35 -05:00
0xdhx
dbbbd7e89d chore(#3156): re-point the changeset at the successor issue
#2665 was closed by #3134 (the narrow three-variable scrub) while this PR was
open, so the changeset's issue refs pointed at resolved work. Re-pointed to
#3156, which is the class this PR actually closes. Fragment content is
otherwise unchanged; changeset-lint returns ok_fragment_present.
2026-08-08 05:57:42 -05:00
0xdhx
af39f13be2 test(#3156): pin the ambient-HOME leak the scrub set cannot reach
Two halves, and the second is the one that matters.

The contract half asserts installSpawnEnv() and installerEnv() both replace the
ambient HOME, keep USERPROFILE tracking it (os.homedir() reads that one on
Windows), and still let an explicit override win.

The behavioural half drives the REAL installer against a REAL ambient HOME: it
points process.env.HOME at a canary, runs `install.js --cursor --local`, and
asserts the canary gains no .gsd. No assertion about the scrub set can stand in
for this, because writeNonClaudeDefaults() resolves through os.homedir(), which
reads no GSD variable -- a set-membership test would pass with the bug fully
live.

Negative-controlled against the pre-fix tree rather than assumed: reverting
installerEnv() to `{ ...process.env, ...overrides }` fails BOTH halves, the
behavioural one reporting the actual artifact
("the installer wrote GSD's user store into the ambient HOME: defaults.json").
2026-08-08 05:57:28 -05:00
0xdhx
35df0891af fix(#3156): sandbox HOME on raw installer spawns — the one leak no scrub reaches
The strict live-config guard this PR ships went red on CI: the suite creates
$HOME/.gsd/defaults.json. Diagnosed rather than suppressed, because the guard
is right — this is #2665's class arriving through the one door the scrub set is
structurally unable to close.

bin/install.js writeNonClaudeDefaults() (#2834) writes
path.join(os.homedir(), '.gsd', 'defaults.json') for every non-Claude runtime.
os.homedir() consults NO GSD variable, so:

  - no entry in CONFIG_LOCATION_ENV_KEYS can reach it, however the set is
    derived; and
  - blanking GSD_HOME does not reach it either -- a blank GSD_HOME falls back
    to exactly that homedir().

Only a sandboxed HOME contains it, and HOME is deliberately excluded from
TEST_ENV_BASE because blanking it would break far more than it fixed. So the
containment belongs per-spawn, which is the discipline the suite already
applies by hand -- install-minimal-hooks.test.cjs:576 carries a comment
naming this exact hazard for the --codex spawn, while the parameterized
--${runtime} spawn 130 lines below it does not. Instance fixed, class open.

Rather than add a fifth hand-synced env shape to a PR whose subject is that
hand-synced copies drift, this adds ONE export -- installSpawnEnv() in
tests/helpers.cjs -- and routes every raw installer spawn through it, including
the shared tests/helpers/install-shared.cjs installerEnv(), which every install
suite already consumes. Callers passing an explicit { HOME, USERPROFILE } are
unaffected: overrides spread last.

Census (measured, not reasoned): of the 119 test files that spawn bin/install.js,
exactly four wrote into a sandboxed $HOME before this commit -- install,
copilot-install, install-minimal-hooks, opencode-plugin-adapter -- and zero do
after. Only install.test.cjs sits in CI's targeted lane, which is why ubuntu
went red on one file while the macOS full lanes went red on four.

Attribution: the leak reproduces unchanged at upstream/next itself, so the
defect is base-owned and pre-existing; only the detector is new. The guard
found a real leak on next within one run.

No new failures: the surviving names under a sandboxed HOME
(folded:enh-2380-sync-skills, getGlobalConfigDir (Copilot)) fail at base too,
and base additionally fails folded:bug-3288-model-catalog-install-path, which
this tree does not.
2026-08-08 05:56:22 -05:00
0xdhx
87f6af282a chore(#3156): rebake both CONTEXT-INDEX mirrors after the rebase
The rebase conflicted on both generated indexes. Resolved arbitrarily during
the replay and regenerated from their producers rather than hand-merged --
gen-context-index.cjs for docs/, and the example's own generator for the
mirror, which a separate lint checks.

426 predicates, 21 classes, 0 duplicate ids. lint:generated-sync and
lint-example-parser-parity both pass.
2026-08-08 05:51:57 -05:00
0xdhx
fae0c6ae1a fix(#2665): stop watching shared ground, and derive the artifact prefix too
The previous commit widened the guard's watch set and claimed the enumeration
was complete. Re-running the pre-push adversarial gate on that commit -- which I
should have done before pushing it, and did not -- refuted the claim on four
counts. All four were real.

1. FALSE POSITIVES, which is the worse polarity. `hooks/lib`,
   `hooks/package.json`, `scripts/lib` and `scripts/changeset` were watched
   WHOLESALE. The installer preserves foreign files in every one of them -- it
   removes the CommonJS marker only on an exact content match, because "a
   user-authored package.json is never deleted" -- so a user editing their own
   helper mid-suite tripped the guard. A driven probe produced four violations
   from touching only user-owned files. Watching shared ground is exactly what
   the module's SCOPE note refuses: a guard that cries wolf gets switched off,
   and then catches nothing at all. Now only exact GSD filenames inside those
   dirs are watched, and a test asserts foreign edits stay silent.

2. THE PREFIX WAS HARDCODED, which is this PR's own defect one level down. Each
   artifactLayout declares its OWN prefix, and kimi's `kimi-agents` layout
   declares `gsd` with no hyphen, writing `agents/gsd.yaml` and `agents/gsd.md`.
   A fixed `gsd-` scan is structurally blind to both, as it is to pi's
   `extensions/gsd.js`. The prefix is now derived per parent, as a SET -- the
   same destSubpath carries different prefixes across runtimes (`agents` appears
   with both `gsd` and `gsd-`). `extensions` joins the non-registry parents; pi
   declares no artifactLayout at all, so no registry walk could find it.

3. THE ENTRY BOUND FAILED OPEN on a non-finite limit: `Math.max(0, NaN)` is NaN,
   and every budget comparison against NaN is false, so the walk was unbounded --
   the single thing the constant exists to prevent. Clamped with Number.isFinite.
   The walk also kept invoking itself for every remaining sibling after the
   budget was gone; it now returns.

4. THE RESIDUAL LIST WAS WRONG AGAIN. `agents/subagents/**` (kimi stages under an
   unprefixed intermediate dir), the loose capability generators, and the
   `extensions`/`plugins` CommonJS markers are all unwatched and were unnamed.
   They are named now, and the four shared dirs are recorded as DELIBERATELY not
   watched -- a different thing from missed.

Each fix is negative-controlled and each control fires. The NaN control did not
fire on its first form: the test asserted `truncated: false`, which the broken
code also produces on a small tree, so it discriminated nothing. Repaired with a
NaN perTarget against a small finite ceiling, where the two behaviours differ.
2026-08-08 05:50:49 -05:00
0xdhx
766480967e fix(#2665): derive the guard's artifact targets, and close the fallback hole in the extras
A pre-push adversarial review refuted this round's own completeness claim, and it
was right on all three counts. Fixes, in the order they matter:

1. The watch enumeration was still a hand-list, and it was measurably incomplete.
   It missed kilo's SINGULAR `command/`, hermes' `skills/gsd` (a whole directory
   whose name carries no `gsd-` prefix, so no prefix rule could ever reach it),
   `plugins/gsd-core.js`, and the unprefixed subtrees the installer fills --
   `hooks/lib`, `hooks/package.json`, `scripts/lib`, `scripts/changeset`.
   The parents are now DERIVED from the capability registry's own
   artifactLayout.global destSubpath values, exactly as TEST_ENV_BASE derives its
   keys, plus a named list for the non-registry paths the installer writes
   directly. A capability declaring a new destination now extends the watch set
   in the commit that declares it. Scope note: only `global` is walked --
   `workflows` is declared LOCAL-only (windsurf) and is not a config-root parent.

2. resolveExtraWatchTargets carried the identical ambient-only defect that
   Blocker 3 closed one function over: it resolved $GSD_HOME/.gsd and each kimi
   descriptor from the ambient env alone, so a child that BLANKED those vars
   wrote to the HOME-derived fallback while the guard watched the override. Both
   legs are now unioned, matching resolveLiveConfigRoots.

3. The order-independence claim for the scan budget was too strong. It holds
   BELOW the global ceiling; once MAX_TOTAL_ENTRIES is exhausted, which targets
   get curtailed still depends on iteration order -- inherent to any shared
   aggregate bound. The residual is now named in the docblock and the test title
   says which regime it pins, instead of asserting the general claim. Negative
   limits are clamped at 0 so an injected value cannot masquerade as a scan bound.

The module's KNOWN GAP now names its remaining residuals (the loose generator
scripts, the kimi native-root hook bundle) rather than implying completeness --
an unqualified claim here just invites the same refutation next round. Both
under-watch, which fails quiet.

Reverting the derivation fails two tests; reverting the fallback leg fails a
third.
2026-08-08 05:50:49 -05:00
0xdhx
104fc76f70 fix(#2665): watch the hook bundle and the install markers the census found
Self-found by re-deriving the guard-shape census against bin/install.js's own
write sites, not by a review finding. Three artifacts a global install writes
into a live config ROOT were watched by nothing:

  hooks/gsd-check-update.js, hooks/gsd-context-monitor.js,
  hooks/gsd-update-banner.js   -- `hooks` was absent from GSD_PREFIXED_PARENTS
  .gsd-source, .gsd-profile    -- absent from GSD_OWNED_ENTRIES, and an
                                  exact-name list does not match a dot-prefixed
                                  name via the `gsd-` prefix rule

This is the SAME shape as the leak that motivated the prefixed-parent scan in
round 1 -- a gsd-prefixed child under a parent nobody had listed -- one parent
over. That it recurred is the argument for re-deriving this list from the
installer each round instead of trusting it: the enumeration is the weak point
of an enumerate-and-block mechanism, and it does not announce when it falls
behind.

Ownership is unchanged, only coverage: `hooks/` is shared with the host agent,
so only `gsd-`-prefixed children are watched. A test asserts a host-owned
hook is still ignored, because widening the parent list must not widen
ownership -- a guard that flags the host's own files gets switched off, and
then catches nothing at all.

Reverting the widening fails the new test.
2026-08-08 05:50:49 -05:00
0xdhx
e31f706ceb docs(#2665): document the two live-config-guard env vars
The changeset for this PR is typed `Added`, and CONTRIBUTING requires a docs/
change for that type. The only docs/ file in the diff was CONTEXT-INDEX.json --
a GENERATED index -- so the Docs Required gate passed while no human-readable
documentation existed for either new variable. A gate satisfied by a generated
artifact is satisfied vacuously.

docs/TESTING-SUITES.md now carries a section on the guard: what it watches and
why it is ownership-scoped rather than whole-root, the two env vars in a table,
why the default is report-only and what the promotion condition is, and what
each violation label means (including that UNVERIFIED is not clean).

GSD_SKIP_LIVE_CONFIG_GUARD is named explicitly because it is a bypass on a
safety check. An undocumented bypass is one people eventually set without
knowing what they turned off.

A test asserts both variables appear in that doc -- checked as permitted by
local/no-source-grep before writing it, rather than assumed forbidden. It fails
when the section is removed, so the doc cannot rot back to the state the review
found.

Addresses review finding: Major 4.
2026-08-08 05:50:49 -05:00
0xdhx
f0ef9063d5 test(#2665): restore the three agent-skills tests this PR deleted
Commit 2bed9fd8 ("replace the hand-synced TEST_ENV_BASE copies with the
canonical import") also removed markLocalGsdInstall and three behavioural tests
from tests/agent-skills.test.cjs -- 79 lines, no replacement, and no mention in
the commit message or the PR body:

  - unconfigured Codex reads its local companion agent from a descendant cwd
  - workstream runtime selects the local Codex companion when root config differs
  - unconfigured Claude remains empty when a local Codex companion exists

RULESET.TESTS.delete-bad-tests permits deleting a bad test only when it is
replaced with compliant tests in the same PR. Nothing was replaced, and these
were not bad tests -- they were collateral in a mechanical edit. A silent net
loss of behavioural coverage inside a PR whose subject is test hygiene is the
one thing that should not pass here, and the reviewer was right to block on it.

Restored verbatim. They need no adaptation to the canonical TEST_ENV_BASE
import: they pass their env explicitly, and an explicit env still spreads last
over the base. The file goes 80 -> 83 tests, all green.

Non-vacuity checked rather than assumed: dropping the local-install marker the
first two depend on fails both. The third is a negative assertion and correctly
stays green, which is why it is named here rather than counted as covered.

Addresses review finding: Blocker 1.
2026-08-08 05:50:49 -05:00
0xdhx
e4f79c32b0 fix(#2665): wire the fourth suite lane, and derive the lane list instead of naming it
qa-loop-walk runs `npm run test:qa`, which is `run-tests.cjs --suite qa` -- so it
runs the live-config guard like every other suite lane, and it set no
GSD_STRICT_LIVE_CONFIG_GUARD. A leak of exactly the class this PR closes would
have printed a warning there and left the lane green.

The test that is supposed to prove the guard is wired everywhere could not
detect that, because its job list was three literals (`test`, `test-full`,
`test-inert`). A hand-list certifying its own completeness is the defect this
whole PR is about, reproduced inside the test guarding the fix -- so the list is
now DERIVED from the workflow: every job with a step reaching run-tests.cjs,
directly or through an npm script resolved transitively through package.json.
The indirection is the load-bearing half; a grep for the filename alone is what
made qa-loop-walk invisible.

The derivation asserts a floor (>= 4 jobs) before ruling on any of them, so a
selector that silently matched nothing fails loudly instead of passing
vacuously. Windows lanes keep their carve-out, keyed on whether the job's matrix
mentions windows rather than on the job's name.

Negative-controlled: un-wiring qa-loop-walk fails the new test. The literal
version passed with that lane unwired, which is how it shipped.

Addresses review finding: Major 5.
2026-08-08 05:50:49 -05:00
0xdhx
4bc6b0a2a3 fix(#2665): defer the built-lib require so an unbuilt tree fails one test, not all
tests/helpers.cjs required gsd-core/bin/lib at module scope to derive the
config-location scrub set. That lib is BUILT, so on an unbuilt tree the require
threw inside `require('./helpers.cjs')` -- before a single test() had registered
-- turning one missing `npm run build:lib` into a whole-suite crash with no
message naming the remedy. This is the file ~370 test files import, so the blast
radius is the suite. `npm test` builds via its pretest hook; the shape that
reaches this is a direct `node --test` invocation, which is exactly what a
contributor reaches for when running one file.

The require is now memoized behind builtLib(), and the two derived exports
(TEST_ENV_BASE, CONFIG_LOCATION_ENV_KEYS) are enumerable lazy getters, so
destructuring and Object.keys() behave as before. Reading either is what forces
the build; a test file that needs neither now imports cleanly. When the build IS
missing, the error names `npm run build:lib` instead of surfacing a bare
MODULE_NOT_FOUND.

Verified by a cold-child probe rather than by inspection -- this process has
already loaded everything, so an in-process assertion would pass vacuously. The
probe checks require.cache before and after touching TEST_ENV_BASE, and fails
when the require is moved back to module scope.

Addresses review finding: Major 7.
2026-08-08 05:50:49 -05:00
0xdhx
4eb29b9741 test(#2665): pin the MAX_DEPTH boundary on both sides, not just above it
RULESET.TESTS.boundary-coverage asks for {limit-1, limit, limit+1}. The depth
bound was exercised only at limit+2, which pins neither side of the edge: an
off-by-one that truncated a tree sitting exactly AT MAX_DEPTH would have passed,
and a truncation is not a cosmetic miss here -- it reports `unverified`, which
in strict mode fails the run.

Negative-controlled by weakening the guard to `depth >= MAX_DEPTH`: the new
limit case fails, where the previous single limit+2 assertion did not.

Addresses review finding: Minor 8.
2026-08-08 05:50:49 -05:00
0xdhx
f054c85fb0 fix(#2665): budget the scan per target, so order stops deciding the verdict
MAX_ENTRIES was a single running budget threaded across every watch target. One
large early target exhausted it, and every target scanned afterwards reported
truncated -> `unverified` -- which under GSD_STRICT_LIVE_CONFIG_GUARD=1 is a
failed run. The guard's verdict therefore depended on directory iteration order
and on unrelated local state, neither of which says anything about whether the
suite leaked.

Each target now draws a fresh allotment, so a pathological tree truncates itself
and nothing else. MAX_TOTAL_ENTRIES keeps the aggregate bounded -- which is what
the single budget was actually for -- and when that ceiling engages, the targets
it curtails are still reported `unverified` rather than attested clean.

The limits are injectable so the boundary is testable without materialising
20000 entries, matching the `deps` seam the resolvers already use.

Two of the three new tests fail when the shared budget is restored; the third
asserts the retained global ceiling, which is deliberately unchanged behaviour.

Addresses review finding: Major 6.
2026-08-08 05:50:49 -05:00
0xdhx
e82a15a852 fix(#2665): watch the fallback root a scrubbing child actually resolves to
resolveLiveConfigRoots resolves what THIS process sees, and getGlobalConfigDir
is env-first -- so with an ambient CLAUDE_CONFIG_DIR the guard watched that
path. A spawned child does not see it: TEST_ENV_BASE blanks the config-location
vars precisely so the child cannot follow them, and a blanked var is falsy, so
the child resolves its HOME-derived root instead.

A child that blanks the var and does NOT also sandbox HOME therefore writes into
the developer's real ~/.claude, which the guard was not watching. That is this
PR's own escape route, taken one process deeper -- and the guard is the artifact
that is supposed to make it loud.

Both resolutions are now unioned: the ambient one, and the fallback one obtained
by handing the REAL descriptor resolver an EMPTY env. Deriving it that way is
deliberate -- a hand-listed copy of the scrub set inside the guard is a second
list to drift, which is the defect this PR spent three rounds closing one layer
up. grok resolves through a hardcoded branch rather than a descriptor, so its
fallback is stated explicitly for the same reason it is named in the ambient
loop.

Addresses review finding: Blocker 3.
2026-08-08 05:50:49 -05:00
0xdhx
6a1fbf96fd fix(#2665): let the guard see deletions, in both shapes it can take
diffLiveConfig walked `after` alone, so it had no branch for a path that
existed before the run and does not after. A test run that DELETES a file from
the developer's real config dir passed the guard silently -- the least
recoverable case in the threat model this guard exists to cover.

The review named the missing `pre.exists && !post.exists` branch. That branch is
necessary and not sufficient: deletion arrives in two shapes and it reaches only
one of them.

  - A FIXED owned entry (GSD_OWNED_ENTRIES x roots, plus every extra target) is
    recorded at both ends whether it exists or not, so a deletion reads
    {exists:true} -> {exists:false}. This is the shape the named branch fixes.
  - A gsd-prefixed child is DISCOVERED by readdirSync, so a deleted one is
    absent from `after` entirely and never enters an after-keyed loop at all.
    The named branch is unreachable for it.

So the walk is now over the UNION of both key sets, with the explicit branch for
the first shape and an `!post` branch for the second. Both are covered by a
test, and reverting the fix fails both -- the prefixed-child test is the one
that would still fail with only the prescribed branch in place.

Addresses review finding: Blocker 2.
2026-08-08 05:50:49 -05:00
0xdhx
ecea537194 docs(#2665): the guard watches config.toml but GSD also writes <root>/hooks/ there
Found pre-push by this round's third adversarial review pass. Not a rebase
regression — round 3 shipped it and #2755 doubled it.

resolveExtraWatchTargets watches one config.toml per non-registry descriptor,
and its comment asserted "GSD writes ONE named file into these third-party
roots". That is false: bin/install.js also calls installSharedHooksBundle on the
same root, populating <root>/hooks/ with GSD's hook scripts and a CommonJS
marker. So a suite-produced leak of a hook bundle into a developer's real
~/.kimi or ~/.kimi-code passes this guard silently — #2665's own hazard, in
#2665's own safety net.

Behaviour is deliberately unchanged and the gap is disclosed instead. Closing it
is a layout decision rather than one more path, for the same reason
getGlobalSkillsBase is already a deliberate non-target: the snapshot applies the
config-root layout beneath every root it is given, and these roots are not ours.
Happy to fix it here or take it as a separate issue — the maintainer's call.

The enumeration-relative test could not have caught this: it asserts one target
PER DESCRIPTOR and nothing about whether one per descriptor is enough, because
its expectation is derived from the same array it checks. That is exactly the
scope boundary round-2 Nit 7 asked to be marked, biting one layer up from where
it was marked; the test now says so.

479979c4's message says "there are three" — that is three WATCHED targets, not a
count of write surfaces. The hooks bundle is a fourth, and unwatched.

lint:ci rc=0; tests/live-config-guard.test.cjs 24/24. Comments and catalog only.
2026-08-08 05:50:49 -05:00
0xdhx
12cfd27f53 docs(#2665): the guard's own comments still described one kimi home, not two
Same drift as the CONTEXT.md seams, one layer over: #2755 took
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS from one entry to two, and five comments
across three files were left describing the one-entry world — "two live write
surfaces", "today's only entry", "today's single entry", and a <kimi>/config.toml
bullet naming only Kimi CLI's KIMI_SHARE_DIR.

The sharpest one was a wrong pointer rather than a stale count: run-tests.cjs
cited "scripts/lib/live-config-guard.cjs" for why the scope is narrow. That path
does not exist, and it names the one directory this module is deliberately NOT
in — the installer copies scripts/lib/ to users wholesale while uninstall removes
only an allowlist, which is the whole reason the guard lives one level up. A
reader following that pointer would have concluded the opposite of the decision.

Comments only; no behaviour change. lint:ci rc=0, tests/live-config-guard.test.cjs
24/24, tests/run-tests-harness.test.cjs 138/138.
2026-08-08 05:50:49 -05:00
0xdhx
664bab3b48 docs(#2665): the scrub-set and guard-target seams understated their own mechanism
Both predicates were written in round 3 and not updated when round 4 widened
what they describe, so the catalog that exists to stop a future omission had
two of its own.

CONFIG.LOCATION.SEAM.scrub-set said "four sources" and listed four. There are
five rungs: the two descriptor rungs each additionally walk skillsHome.env (the
maintainer's round-4 Minor 3), and the fifth is WRITE_ESCAPE_PERMISSION_ENV_KEYS,
which is a permission rather than a location and so is reachable by no other rung.

LIVE-CONFIG.GUARD.SEAM.non-root-targets said "the two live write surfaces" and
named a single config.toml via resolveKimiHooksTomlDir. Since #2755 landed
KIMI_CODE_HOOKS_TOML_DESCRIPTOR there are three, and the guard derives them by
iterating NON_REGISTRY_CONFIG_HOME_DESCRIPTORS rather than calling a named
resolver. Both bounds are stated rather than left open: skills bases are a
deliberate non-target (the config-root layout misfires beneath them), and a
further descriptor is only free if it owns the same NON_REGISTRY_OWNED_FILE —
the residual the guard already names at its own definition.

LIVE-CONFIG.GUARD.SEAM.scope's "gsd--prefixed" read as a two-hyphen prefix; the
selector is startsWith(GSD_ARTIFACT_PREFIX) where that constant is 'gsd-'.

CONFIG.LOCATION.SEAM.kimi-two-homes was checked and is NOT stale: #2755 made
Kimi Code a separate runtime, so Kimi CLI still declares exactly two.

Both CONTEXT-INDEX mirrors regenerated; predicate ids unchanged (425), only
their descriptions move.
2026-08-08 05:50:49 -05:00
0xdhx
d43f781530 chore(#2665): regenerate the example CONTEXT-INDEX mirror after the rebase
The rebase onto next conflicted in this generated file; per the
generated-file discipline the conflict was resolved arbitrarily and the
artifact regenerated from its producer rather than hand-merged.

Regen delta against the base's committed copy is exactly this PR's own
10 predicates (CONFIG.LOCATION.SEAM.* x4, LIVE-CONFIG.GUARD.SEAM.* x6),
415 -> 425 keys, zero removed and zero foreign — so the merge introduced
no drift the PR does not own.
2026-08-08 05:50:49 -05:00
0xdhx
c95b817cb6 fix(#2665): scrub GSD_ALLOW_SYMLINKED_DEST — a write-escape permission
Found by the pre-publication claim audit of this round's response comment,
which refuted the sentence "none of the unscrubbed env reads names a write
destination" on the grounds that naming a path is not the same property as
influencing where writes land.

GSD_ALLOW_SYMLINKED_DEST is boolean and names no path, so every rung of the
derivation is structurally incapable of reaching it: not a registry configHome,
not descriptor-shaped, not one of GSD's own location vars. It is still a #2665
leak vector. install-engine.cts reads it env-first (:214) and threads it as
allowOptInFollow into the symlink-escape guard at four call sites, each gating
a write (:361/:367, :416/:424, :785/:790, :927/:932). That guard is what stops
a write leaving the install root, so an ambient =1 disarms it for the whole
suite — the #2665 hazard arriving through a permission rather than a path.

Added as its own named family (WRITE_ESCAPE_PERMISSION_ENV_KEYS) rather than
folded into a location rung, for the same reason GSD_HOME got its own family in
round 3: the list should not misdescribe what its members are.

Blanking is fail-safe in the only direction that matters — '' is neither '1'
nor 'true', so a blanked value makes the guard stricter, never looser. That
asymmetry is what licenses scrubbing it wholesale rather than reasoning about
each call site.

The guard test names the variable literally rather than iterating the family
constant: a test that asserts over the constant shrinks its own expectation
when the family is emptied, which is the enumeration-relative failure that let
the kimi-code descriptor go unwatched earlier in this same round.

#2393's opt-in suite is unaffected (99/100, 0 fail): it sets process.env
directly in-process and never routes through scrubConfigLocationEnv, and an
explicit env argument still wins over TEST_ENV_BASE in childEnv.
2026-08-08 05:50:49 -05:00
0xdhx
2a98b6b0b1 fix(#2665): keep the dot-home type guarantee #2755 chose
The rebase resolution annotated both hoisted descriptors and the local in
resolveKimiHooksTomlDir as ConfigHomeDescriptor. #2755 used DotHomeDescriptor
there deliberately: the union permits xdg / dot-home-nested / generic-agents-root
shapes, so the broader annotation drops the compile-time guarantee that this
resolver selects a dot-home descriptor and nothing else.

Runtime behaviour was already correct — all eleven paths resolve identically —
so this restores a type-level property, not a behavioural one. Both exported
constants are now pinned to DotHomeDescriptor as well, which is stricter than
the ConfigHomeDescriptor they carried since round 3; the interface stays
unexported and NON_REGISTRY_CONFIG_HOME_DESCRIPTORS keeps its
ConfigHomeDescriptor[] type, which the narrower constants satisfy as subtypes.

Found by the pre-push adversarial review of this round, which is the one claim
of six it refuted.
2026-08-08 05:50:49 -05:00
0xdhx
d1c8b32689 fix(#2665): watch kimi-code's config.toml, and pin it by name
The rebase onto next brought in #2755, which added a SECOND Kimi config
home — kimi-code's `~/.kimi-code`, overridden by KIMI_CODE_HOME — declared
as an inline object literal inside resolveKimiHooksTomlDir's body. That is
the resolvable-but-not-enumerable shape round 3 hoisted KIMI_SHARE_DIR out
of, so the hoist is extended to cover both descriptors rather than reverting
#2755's parameterization.

The scrub set was already complete: KIMI_CODE_HOME is declared in
capabilities/kimi-code/capability.json, so the registry rung covered it and
CONFIG_LOCATION_ENV_KEYS is 28 keys both before and after the rebase. What
was NOT covered is the guard — resolveExtraWatchTargets iterates
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, so with only one entry it watched Kimi
CLI's config.toml and never Kimi Code's. Targets go 2 -> 3.

The existing 'extra targets are DERIVED from the descriptor array' test
cannot catch this: it builds its expectation FROM the array, so removing an
entry shrinks the expectation with it. Verified — with the kimi-code
descriptor removed that test still passes while the new named test fails.
This is the enumeration-relative scope boundary the suite already documents
one layer down, biting one layer up.

Also rewrites NON_REGISTRY_OWNED_FILE's docblock, which asserted "today's
only such descriptor is kimi's ~/.kimi". There are now two, and its named
residual is load-bearing rather than vacuous.
2026-08-08 05:50:49 -05:00
0xdhx
111733a057 chore(#2665): regenerate the examples CONTEXT-INDEX.json mirror too
lint-example-parser-parity holds examples/dynamic-context-management/
CONTEXT-INDEX.json to a fresh parse of CONTEXT.md; the round-4 seam-entry
rewrites needed this second derived artifact regenerated alongside
docs/CONTEXT-INDEX.json. Derived regen only.
2026-08-08 05:50:49 -05:00
0xdhx
123ba26fc5 chore(#2665): regenerate docs/CONTEXT-INDEX.json for the round-4 CONTEXT.md edits
The round-4 seam-entry rewrites (packaging fact on SEAM.module, strict-mode
state on SEAM.severity/ci-blind) left the derived index stale, which failed
lint:generated-sync and the #2944 sync test across ten CI contexts. Derived
artifact regen only; no content change beyond what CONTEXT.md already says.
2026-08-08 05:50:49 -05:00
0xdhx
652b99b71b test(#2665): reversion guards for all three round-4 fixes
None of the round-4 fixes had a test that fails on reversion: re-shipping
the test-instrumentation chain passes #2858 (everything ships, so every
require resolves), dropping the strict env from test.yml demotes the guard
to report-only with nothing red, and the skillsHome derivation tests are
enumeration-relative over declarations that are all empty today. One
guard each, every one negative-controlled against its reverted fix (fails
pre-fix, passes post-fix):

- packaging-shipped-scripts-require-only-shipped.test.cjs asserts the four
  chain files are absent from the npm pack file list (reuses the tarball
  set the #2858 gate already resolves — no second npm pack).
- live-config-guard.test.cjs asserts all three test jobs wire
  GSD_STRICT_LIVE_CONFIG_GUARD, matching the WHOLE expression anchored —
  a prefix match accepted both a Windows-silently-strict tail and a
  malformed one.
- helpers-process-isolation.test.cjs cold-requires helpers.cjs in a child
  with sentinel skillsHome env vars injected into both enumerations, so
  the walk itself is under test rather than today's empty declarations.
2026-08-08 05:50:49 -05:00
0xdhx
7b8c36f904 fix(#2665): walk skillsHome.env on both descriptor rungs of the scrub derivation
Review round 4, Minor 3. A configHome descriptor can nest a second,
independently-resolved descriptor (skillsHome -> resolveSkillsBaseFromDescriptor)
carrying its own env array, and the derivation walked configHome.env alone —
the identical walk-one-field gap-shape rounds 2-3 closed for the registry
and the non-registry set. Inert today (only kilo declares skillsHome, with
env: []), closed before it is live rather than after.

The guard's root enumeration deliberately does NOT gain the skills base:
getGlobalSkillsBase returns a skills directory (codex: ~/.agents/skills),
not a config root, and the snapshot applies the config-root layout beneath
every root — adding it false-positives on <skillsBase>/gsd-core while
missing a real <skillsBase>/gsd-help write (found by this round's pre-push
adversarial review). Watching skills bases needs its own layout, like
resolveExtraWatchTargets; a comment in resolveLiveConfigRoots records the
non-action.

New derivation test asserts both skillsHome rungs land in TEST_ENV_BASE,
with an anti-vacuity check that at least one runtime actually declares the
field.
2026-08-08 05:50:49 -05:00
0xdhx
253250580e fix(#2665): wire the live-config guard to strict mode on Linux/macOS CI lanes
Review round 4, Major 2, answering the explicit report-only-vs-strict
question: strict now. A future regression of the class this PR closes
should fail CI, not print a warning nobody reads — that is what the PR
title promises.

Scoped deliberately: GSD_STRICT_LIVE_CONFIG_GUARD=1 on the Linux/macOS
lanes of all three test jobs; Windows lanes stay report-only because the
guard's first run found pre-existing USERPROFILE leaks there (~190 test
sites sandbox HOME alone) — flipping them strict today reddens next on a
defect class this PR does not carry. Promote once that sweep lands (the
SEVERITY note in live-config-guard.cjs and the CONTEXT.md seam both now
record that state).
2026-08-08 05:50:49 -05:00
0xdhx
209f2fe983 fix(#2665): exclude the test-instrumentation chain from the npm tarball
Review round 4, Major 1: scripts/live-config-guard.cjs is pure test
instrumentation and was shipping to every npm install. The repo already
carries the exclusion convention (gen-emitted-baseline, qa-smell-ratchet)
in the same files[] array.

Excluding the guard alone would trip the #2858 shipped-requires-only-shipped
gate: run-tests.cjs (shipped) requires it at load time, and
affected-tests-lib.cjs / run-affected-tests.cjs sit on the same chain. The
four files are one closed require chain of test instrumentation, so the
exclusion covers the chain, not one link. The guard's LOCATION header cited
affected-tests-lib.cjs as "the precedent for a non-shipped helper", which
npm pack disproves — rewritten to the tarball-exclusion fact.
2026-08-08 05:50:49 -05:00
0xdhx
706bd2ab4e refactor(#2665): derive the guard's non-root targets from the descriptor array too
Follow-up to 38c9395d, found while fact-checking the round-3 response rather
than by a test.

That commit made TEST_ENV_BASE derive its keys from
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, but had the guard call
resolveKimiHooksTomlDir directly. Both halves covered kimi, so nothing was
broken — but only one of them would pick up a SECOND descriptor. That is the
same partial-enumeration defect that put KIMI_SHARE_DIR outside the scrub set,
reintroduced one layer over, in the very commit that closed it.

resolveExtraWatchTargets now iterates the array and resolves each descriptor
through resolveConfigHomeFromDescriptor, so the scrub set and the guard derive
from one source and cannot drift apart.

Verified: a synthetic second descriptor is picked up automatically (it was not
before); kimi's target is unchanged on both the default (~/.kimi/config.toml)
and KIMI_SHARE_DIR override paths.

The new test asserts one target per descriptor plus the store root. The COUNT
is the load-bearing half — every per-descriptor assertion passes vacuously
today with a single entry, so only the count fails when the array grows and the
guard does not follow.

NAMED RESIDUAL, documented at NON_REGISTRY_OWNED_FILE: this assumes every
non-registry descriptor is written the same way (config.toml). A descriptor
whose owned file differs needs a per-descriptor mapping. It fails toward
under-watching rather than false positives, so it is called out rather than
left to be discovered.
2026-08-08 05:50:49 -05:00
0xdhx
fb02fa5722 chore(#2665): add the changeset fragment this PR now needs
Caught by CI on the round-3 push, not by review: `changeset-lint` went from
OK_NO_USER_FACING_CHANGES to FAIL_MISSING_FRAGMENT.

Both prior rounds passed that gate because the diff was tests/ + scripts/ only,
and neither is a USER_FACING_PREFIX. Round 3 is the first commit in this PR to
touch src/ — the descriptor hoist in src/runtime-homes.cts — and src/ is on the
list. So the fragment requirement is a direct consequence of the remedy shape,
not something the earlier rounds missed.

src/runtime-homes.cts is the ONLY user-facing file in the whole PR diff
(23 changed files); everything else is tests/, scripts/, CONTEXT.md, the two
generated CONTEXT-INDEX.json artifacts, and examples/.

Typed `Added` rather than `Fixed` because that is what actually reaches a user:
new exported constants on a shipped module. The fix itself (#2665) is test
hermeticity, which changes no shipped behaviour — resolveKimiHooksTomlDir()
resolves identically before and after.

Worth noting the local lint disagreed with CI and was wrong: it diffs
`origin/main...HEAD`, which sweeps in the base's own committed fragments and
returns a false ok_fragment_present. CI diffs `origin/$GITHUB_BASE_REF...HEAD`
(next) and sees only this PR's files, which is the correct question.
2026-08-08 05:50:49 -05:00
0xdhx
434d71b03b docs(#2665): catalog the config-location and live-config-guard seams in CONTEXT.md
Round 2, Minor. "Workspace seams" carried WORKTREE.SEAM.* and
CONFIG.SEAM.loadConfig-context but nothing for this PR's mechanism or its env
vars, so the one machine-readable place a future author would look said nothing
about the class that has now recurred three times.

Ten predicates across two groups:

  CONFIG.LOCATION.SEAM.*  — the scrub set's four derivation sources and the rule
                            that a new var is made ENUMERABLE rather than
                            appended; the two-families distinction (runtime
                            configHomes vs GSD's own GSD_HOME/GSD_AGENTS_DIR)
                            that round 2 turned on; kimi's two config-location
                            vars; and the in-process scrub requirement, since
                            HOME sandboxing alone is the trap that produced
                            Blocker 1 twice.
  LIVE-CONFIG.GUARD.SEAM.* — module + exports + why it is scripts/ and not
                            scripts/lib/; ownership-based scope; the two
                            non-root targets and their asymmetric treatment;
                            the truncation contract; the report-not-fatal
                            severity ratchet; and that CI is structurally blind
                            here, so green CI is not evidence.

Both generated indexes regenerated. docs/CONTEXT-INDEX.json is checked by
lint:generated-sync (`gen-context-index.cjs --check`) LINE-NUMBER-SENSITIVELY,
and the example's own committed index is separately checked by
lint-example-parser-parity.cjs, which the first regen did not satisfy — editing
CONTEXT.md requires both, and only one of them says so in its error text.

Verified: parity lint rc=0, gen-context-index --check rc=0, full
lint:generated-sync rc=0. Regen diff audited — 10 predicates added, 0 removed,
0 values changed; the example index's remaining churn is line-number re-baking,
which is exactly why the parity lint excludes line numbers.
2026-08-08 05:50:49 -05:00
0xdhx
fec42e9a7d test(#2665): cover the widened derivation, and mark what these tests cannot prove
Two tests for round 3's change (round 2 Blockers 1 and 2): KIMI_SHARE_DIR must
arrive via NON_REGISTRY_CONFIG_HOME_DESCRIPTORS rather than a literal, and every
GSD_LOCATION_ENV_KEYS entry must be blanked. Both assert a floor on their source
first, so a renamed export fails loudly instead of passing vacuously.
Negative-controlled against the registry-only derivation: both fail there, and
only those two.

And the Nit, which is the more useful half. This block asserts that TEST_ENV_BASE
is not narrower than the enumerations it derives from. It cannot prove those
enumerations are complete — a var no enumeration carries is invisible to every
test here, and they stay green.

That is exactly how round 2 found GSD_HOME and KIMI_SHARE_DIR while this block
was fully green: one belonged to no enumeration at all, the other sat inside a
function body where nothing could enumerate it. So the scope boundary is now
written down at the top of the block, naming where the completeness question is
actually answered — a source census re-derived each round, and
live-config-guard.cjs observing real writes at runtime — so that a green run
here is not misread as "the set is exhaustive."

Deliberately NOT added: an assertion per reviewer-named variable. That is the
hand-maintained list wearing a test's clothes, and it fails the same way.
2026-08-08 05:50:36 -05:00
0xdhx
1f6d827e48 fix(#2665): scrub config-location env in the #2624 in-process install block
Self-found during the rebase onto next, not from the review.

The base range added `describe('#2624 .gsd-source marker is rewritten before
staging reads it')`, which calls the real `install(true, 'claude')` IN-PROCESS
and sandboxes HOME/USERPROFILE/GSD_EXPLICIT_CONFIG_DIR — but not
CLAUDE_CONFIG_DIR. That is exactly the Blocker-1 shape this PR exists to close,
reintroduced in new code written after the round-1 review.

Measured on the rebased tree, same file, same commit:

  CLAUDE_CONFIG_DIR unset -> 50/50 pass
  CLAUDE_CONFIG_DIR set   -> 47/50, and a COMPLETE global install lands in it
                             (gsd-core/, agents/, skills/, hooks/, scripts/,
                             gsd-file-manifest.json, gsd-install-state.json,
                             .gsd-source, .gsd-profile)

The three failures are the honest symptom rather than the problem: the install
goes to the ambient config dir, so the assertions look for a marker under
tmpRoot that was never written there.

With scrubConfigLocationEnv() wired into the block's beforeEach/afterEach —
the same pattern the sibling block at :753 already uses — the file is 50/50
under BOTH conditions and leaks zero entries.

Worth stating plainly: CI cannot catch this class, since CI never has these
vars set. It surfaced here only because the rebase brought the base's new tests
under an ambient CLAUDE_CONFIG_DIR, which is the condition #2665's own
acceptance criterion runs under.
2026-08-08 05:50:36 -05:00
0xdhx
a294ec2a2b test(#2665): widen the hermeticity guard to its two blind surfaces, and cover its budget
Round 2, both Majors. They are one defect seen twice: the recurrence guard did
not cover the surface it exists to guard.

Blind surfaces. resolveLiveConfigRoots enumerates getGlobalConfigDir per registry
runtime plus a hardcoded grok branch, so it can only ever see runtime config
ROOTS. Two live write surfaces are not roots and passed through silently:

  $GSD_HOME/.gsd     — GSD's user-owned store. Watched WHOLESALE: unlike ~/.claude
                       this root is exclusively ours, so the shared-root
                       false-positive trap the module documents does not apply.
  <kimi>/config.toml — the file GSD writes its native [[hooks]] block into. The
                       INVERSE case: ~/.kimi belongs to Kimi CLI, so only the one
                       file GSD writes is watched, never the root.

That asymmetry is why this is not a two-line "add two roots" patch — one target
needs the whole tree, the other needs exactly one file, and collapsing them
either under-watches the store or trips the guard's own documented
false-positive trap on a third party's directory.

Extras are passed to snapshotLiveConfig explicitly rather than resolved inside
it, so a caller snapshotting a fixture root cannot silently pull the developer's
real ~/.gsd into its own assertions. run-tests.cjs now snapshots when EITHER the
roots or the extras are non-empty — previously an unbuilt tree yielding zero
roots disabled the entire guard without saying so.

Budget coverage. The MAX_ENTRIES/MAX_DEPTH bound and the truncated -> 'unverified'
branch had zero tests, despite this module's own docstring naming "a truncated
scan reading as clean" as the safety-critical case. Added per
RULESET.TESTS.boundary-coverage (N in {limit-1, limit, limit+1}, exercised
through newestMtime's injected budget so the boundary is real without
materialising 20000 files) and RULESET.TESTS.property-based-testing (fast-check:
truncation is monotone in the budget; reported newest never exceeds the true
maximum). A regression flipping `truncated` to false on an exhausted budget now
breaks the property for every budget below the tree size.

Negative-controlled: neutering the extras wiring fails exactly the two
new-surface tests and nothing else. 21/21 green with it restored.
2026-08-08 05:50:36 -05:00
0xdhx
3c580b77dc fix(#2665): derive the second config-location family instead of hand-adding it
Review round 2 named GSD_HOME and KIMI_SHARE_DIR as missing from the derived
scrub set. Both premises confirmed; the prescribed remedy is not adopted
verbatim, because adding two more literals to a four-item hand list is the
pattern that reopened this bug three times. The census the derivation
generalizes over was partial, so the census is what widens.

Two structural gaps, both closed at the source:

1. KIMI_SHARE_DIR lived inside resolveKimiHooksTomlDir's body as an inline
   descriptor, resolvable but not ENUMERABLE. Hoisted to an exported
   KIMI_HOOKS_TOML_DESCRIPTOR and collected in
   NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, which TEST_ENV_BASE now derives from.
   kimi is the sharp case: it owns TWO config homes (KIMI_CONFIG_DIR, already
   registry-visible, and this one), so a registry-only derivation looks
   complete and is not.

2. GSD_HOME is a different FAMILY, not a missing registry entry. The registry
   describes where third-party runtimes keep config; GSD_HOME decides where GSD
   keeps its own user-owned state ($GSD_HOME/.gsd/ — consent.json,
   defaults.json, capability overlays), read env-first ahead of os.homedir() by
   capability-loader, capability-consent, capability-state, capability-writer,
   config-loader, install-profiles and bin/install.js. Named as
   GSD_LOCATION_ENV_KEYS rather than folded into the descriptor array, since it
   does not resolve through resolveConfigHomeFromDescriptor.

GSD_AGENTS_DIR joins the same family (round 2, Minor): env-first and
unconditional in getAgentsDir, misdirecting a read rather than a write.

A census of every env-first first-party location var — the guard-shape question
this PR owes each round — now yields exactly one remaining unguarded name,
GSD_MODEL_CATALOG, and it is dead by precedence: the co-located candidate is
index 0 and the loop breaks on first success, so the env var can only win on a
tree that is already broken, and it redirects a read even then.

Derived set: 39 -> 42 keys. resolveKimiHooksTomlDir behaviour unchanged on both
the default and the KIMI_SHARE_DIR override path.
2026-08-08 05:50:36 -05:00
0xdhx
e2eed1c58a test(#2665): ship the hermeticity guard at report level, not fatal
Its first CI run found PRE-EXISTING leaks on the Windows lane —
C:\Users\runneradmin\.claude\gsd-core and skills\gsd-dev-preferences — with all
1196 Windows tests otherwise passing. os.homedir() reads USERPROFILE on Windows,
and ~190 test sites across 31 files sandbox HOME alone, so the suite has been
installing GSD into the runner's real home directory invisibly. That is exactly
the class the guard exists to surface, and exactly the class this PR's review
said CI could never catch.

It is also a different defect from the one #2665 closes, and too large to fold in
here. A brand-new gate that immediately reds an unrelated lane gets bypassed or
reverted rather than obeyed, so the guard reports by default and fails only under
GSD_STRICT_LIVE_CONFIG_GUARD=1.

This is the repo's own established ratchet, not a hedge: the local/no-source-grep
ESLint rule shipped at `warn` and was promoted to `error` after its cleanup sweep
(ADR 452). Promote this the same way once the USERPROFILE sweep lands.
2026-08-08 05:50:17 -05:00
0xdhx
5863351281 test(#2665): separator-safe containment in the regression assertion
startsWith(ambientConfigDir) also matches a sibling like
<tmp>/ambient-live-config-2, so it can report a leak that did not happen. Use
path.relative and check for '..' or an absolute result, the repo's usual shape.

The readdirSync assertion already carried the test, so this is cosmetic.

Addresses review finding: Nit 9.
2026-08-08 05:50:17 -05:00