Commit Graph

5053 Commits

Author SHA1 Message Date
0xdhx
652b99b71b test(#2665): reversion guards for all three round-4 fixes
None of the round-4 fixes had a test that fails on reversion: re-shipping
the test-instrumentation chain passes #2858 (everything ships, so every
require resolves), dropping the strict env from test.yml demotes the guard
to report-only with nothing red, and the skillsHome derivation tests are
enumeration-relative over declarations that are all empty today. One
guard each, every one negative-controlled against its reverted fix (fails
pre-fix, passes post-fix):

- packaging-shipped-scripts-require-only-shipped.test.cjs asserts the four
  chain files are absent from the npm pack file list (reuses the tarball
  set the #2858 gate already resolves — no second npm pack).
- live-config-guard.test.cjs asserts all three test jobs wire
  GSD_STRICT_LIVE_CONFIG_GUARD, matching the WHOLE expression anchored —
  a prefix match accepted both a Windows-silently-strict tail and a
  malformed one.
- helpers-process-isolation.test.cjs cold-requires helpers.cjs in a child
  with sentinel skillsHome env vars injected into both enumerations, so
  the walk itself is under test rather than today's empty declarations.
2026-08-08 05:50:49 -05:00
0xdhx
7b8c36f904 fix(#2665): walk skillsHome.env on both descriptor rungs of the scrub derivation
Review round 4, Minor 3. A configHome descriptor can nest a second,
independently-resolved descriptor (skillsHome -> resolveSkillsBaseFromDescriptor)
carrying its own env array, and the derivation walked configHome.env alone —
the identical walk-one-field gap-shape rounds 2-3 closed for the registry
and the non-registry set. Inert today (only kilo declares skillsHome, with
env: []), closed before it is live rather than after.

The guard's root enumeration deliberately does NOT gain the skills base:
getGlobalSkillsBase returns a skills directory (codex: ~/.agents/skills),
not a config root, and the snapshot applies the config-root layout beneath
every root — adding it false-positives on <skillsBase>/gsd-core while
missing a real <skillsBase>/gsd-help write (found by this round's pre-push
adversarial review). Watching skills bases needs its own layout, like
resolveExtraWatchTargets; a comment in resolveLiveConfigRoots records the
non-action.

New derivation test asserts both skillsHome rungs land in TEST_ENV_BASE,
with an anti-vacuity check that at least one runtime actually declares the
field.
2026-08-08 05:50:49 -05:00
0xdhx
253250580e fix(#2665): wire the live-config guard to strict mode on Linux/macOS CI lanes
Review round 4, Major 2, answering the explicit report-only-vs-strict
question: strict now. A future regression of the class this PR closes
should fail CI, not print a warning nobody reads — that is what the PR
title promises.

Scoped deliberately: GSD_STRICT_LIVE_CONFIG_GUARD=1 on the Linux/macOS
lanes of all three test jobs; Windows lanes stay report-only because the
guard's first run found pre-existing USERPROFILE leaks there (~190 test
sites sandbox HOME alone) — flipping them strict today reddens next on a
defect class this PR does not carry. Promote once that sweep lands (the
SEVERITY note in live-config-guard.cjs and the CONTEXT.md seam both now
record that state).
2026-08-08 05:50:49 -05:00
0xdhx
209f2fe983 fix(#2665): exclude the test-instrumentation chain from the npm tarball
Review round 4, Major 1: scripts/live-config-guard.cjs is pure test
instrumentation and was shipping to every npm install. The repo already
carries the exclusion convention (gen-emitted-baseline, qa-smell-ratchet)
in the same files[] array.

Excluding the guard alone would trip the #2858 shipped-requires-only-shipped
gate: run-tests.cjs (shipped) requires it at load time, and
affected-tests-lib.cjs / run-affected-tests.cjs sit on the same chain. The
four files are one closed require chain of test instrumentation, so the
exclusion covers the chain, not one link. The guard's LOCATION header cited
affected-tests-lib.cjs as "the precedent for a non-shipped helper", which
npm pack disproves — rewritten to the tarball-exclusion fact.
2026-08-08 05:50:49 -05:00
0xdhx
706bd2ab4e refactor(#2665): derive the guard's non-root targets from the descriptor array too
Follow-up to 38c9395d, found while fact-checking the round-3 response rather
than by a test.

That commit made TEST_ENV_BASE derive its keys from
NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, but had the guard call
resolveKimiHooksTomlDir directly. Both halves covered kimi, so nothing was
broken — but only one of them would pick up a SECOND descriptor. That is the
same partial-enumeration defect that put KIMI_SHARE_DIR outside the scrub set,
reintroduced one layer over, in the very commit that closed it.

resolveExtraWatchTargets now iterates the array and resolves each descriptor
through resolveConfigHomeFromDescriptor, so the scrub set and the guard derive
from one source and cannot drift apart.

Verified: a synthetic second descriptor is picked up automatically (it was not
before); kimi's target is unchanged on both the default (~/.kimi/config.toml)
and KIMI_SHARE_DIR override paths.

The new test asserts one target per descriptor plus the store root. The COUNT
is the load-bearing half — every per-descriptor assertion passes vacuously
today with a single entry, so only the count fails when the array grows and the
guard does not follow.

NAMED RESIDUAL, documented at NON_REGISTRY_OWNED_FILE: this assumes every
non-registry descriptor is written the same way (config.toml). A descriptor
whose owned file differs needs a per-descriptor mapping. It fails toward
under-watching rather than false positives, so it is called out rather than
left to be discovered.
2026-08-08 05:50:49 -05:00
0xdhx
fb02fa5722 chore(#2665): add the changeset fragment this PR now needs
Caught by CI on the round-3 push, not by review: `changeset-lint` went from
OK_NO_USER_FACING_CHANGES to FAIL_MISSING_FRAGMENT.

Both prior rounds passed that gate because the diff was tests/ + scripts/ only,
and neither is a USER_FACING_PREFIX. Round 3 is the first commit in this PR to
touch src/ — the descriptor hoist in src/runtime-homes.cts — and src/ is on the
list. So the fragment requirement is a direct consequence of the remedy shape,
not something the earlier rounds missed.

src/runtime-homes.cts is the ONLY user-facing file in the whole PR diff
(23 changed files); everything else is tests/, scripts/, CONTEXT.md, the two
generated CONTEXT-INDEX.json artifacts, and examples/.

Typed `Added` rather than `Fixed` because that is what actually reaches a user:
new exported constants on a shipped module. The fix itself (#2665) is test
hermeticity, which changes no shipped behaviour — resolveKimiHooksTomlDir()
resolves identically before and after.

Worth noting the local lint disagreed with CI and was wrong: it diffs
`origin/main...HEAD`, which sweeps in the base's own committed fragments and
returns a false ok_fragment_present. CI diffs `origin/$GITHUB_BASE_REF...HEAD`
(next) and sees only this PR's files, which is the correct question.
2026-08-08 05:50:49 -05:00
0xdhx
434d71b03b docs(#2665): catalog the config-location and live-config-guard seams in CONTEXT.md
Round 2, Minor. "Workspace seams" carried WORKTREE.SEAM.* and
CONFIG.SEAM.loadConfig-context but nothing for this PR's mechanism or its env
vars, so the one machine-readable place a future author would look said nothing
about the class that has now recurred three times.

Ten predicates across two groups:

  CONFIG.LOCATION.SEAM.*  — the scrub set's four derivation sources and the rule
                            that a new var is made ENUMERABLE rather than
                            appended; the two-families distinction (runtime
                            configHomes vs GSD's own GSD_HOME/GSD_AGENTS_DIR)
                            that round 2 turned on; kimi's two config-location
                            vars; and the in-process scrub requirement, since
                            HOME sandboxing alone is the trap that produced
                            Blocker 1 twice.
  LIVE-CONFIG.GUARD.SEAM.* — module + exports + why it is scripts/ and not
                            scripts/lib/; ownership-based scope; the two
                            non-root targets and their asymmetric treatment;
                            the truncation contract; the report-not-fatal
                            severity ratchet; and that CI is structurally blind
                            here, so green CI is not evidence.

Both generated indexes regenerated. docs/CONTEXT-INDEX.json is checked by
lint:generated-sync (`gen-context-index.cjs --check`) LINE-NUMBER-SENSITIVELY,
and the example's own committed index is separately checked by
lint-example-parser-parity.cjs, which the first regen did not satisfy — editing
CONTEXT.md requires both, and only one of them says so in its error text.

Verified: parity lint rc=0, gen-context-index --check rc=0, full
lint:generated-sync rc=0. Regen diff audited — 10 predicates added, 0 removed,
0 values changed; the example index's remaining churn is line-number re-baking,
which is exactly why the parity lint excludes line numbers.
2026-08-08 05:50:49 -05:00
0xdhx
fec42e9a7d test(#2665): cover the widened derivation, and mark what these tests cannot prove
Two tests for round 3's change (round 2 Blockers 1 and 2): KIMI_SHARE_DIR must
arrive via NON_REGISTRY_CONFIG_HOME_DESCRIPTORS rather than a literal, and every
GSD_LOCATION_ENV_KEYS entry must be blanked. Both assert a floor on their source
first, so a renamed export fails loudly instead of passing vacuously.
Negative-controlled against the registry-only derivation: both fail there, and
only those two.

And the Nit, which is the more useful half. This block asserts that TEST_ENV_BASE
is not narrower than the enumerations it derives from. It cannot prove those
enumerations are complete — a var no enumeration carries is invisible to every
test here, and they stay green.

That is exactly how round 2 found GSD_HOME and KIMI_SHARE_DIR while this block
was fully green: one belonged to no enumeration at all, the other sat inside a
function body where nothing could enumerate it. So the scope boundary is now
written down at the top of the block, naming where the completeness question is
actually answered — a source census re-derived each round, and
live-config-guard.cjs observing real writes at runtime — so that a green run
here is not misread as "the set is exhaustive."

Deliberately NOT added: an assertion per reviewer-named variable. That is the
hand-maintained list wearing a test's clothes, and it fails the same way.
2026-08-08 05:50:36 -05:00
0xdhx
1f6d827e48 fix(#2665): scrub config-location env in the #2624 in-process install block
Self-found during the rebase onto next, not from the review.

The base range added `describe('#2624 .gsd-source marker is rewritten before
staging reads it')`, which calls the real `install(true, 'claude')` IN-PROCESS
and sandboxes HOME/USERPROFILE/GSD_EXPLICIT_CONFIG_DIR — but not
CLAUDE_CONFIG_DIR. That is exactly the Blocker-1 shape this PR exists to close,
reintroduced in new code written after the round-1 review.

Measured on the rebased tree, same file, same commit:

  CLAUDE_CONFIG_DIR unset -> 50/50 pass
  CLAUDE_CONFIG_DIR set   -> 47/50, and a COMPLETE global install lands in it
                             (gsd-core/, agents/, skills/, hooks/, scripts/,
                             gsd-file-manifest.json, gsd-install-state.json,
                             .gsd-source, .gsd-profile)

The three failures are the honest symptom rather than the problem: the install
goes to the ambient config dir, so the assertions look for a marker under
tmpRoot that was never written there.

With scrubConfigLocationEnv() wired into the block's beforeEach/afterEach —
the same pattern the sibling block at :753 already uses — the file is 50/50
under BOTH conditions and leaks zero entries.

Worth stating plainly: CI cannot catch this class, since CI never has these
vars set. It surfaced here only because the rebase brought the base's new tests
under an ambient CLAUDE_CONFIG_DIR, which is the condition #2665's own
acceptance criterion runs under.
2026-08-08 05:50:36 -05:00
0xdhx
a294ec2a2b test(#2665): widen the hermeticity guard to its two blind surfaces, and cover its budget
Round 2, both Majors. They are one defect seen twice: the recurrence guard did
not cover the surface it exists to guard.

Blind surfaces. resolveLiveConfigRoots enumerates getGlobalConfigDir per registry
runtime plus a hardcoded grok branch, so it can only ever see runtime config
ROOTS. Two live write surfaces are not roots and passed through silently:

  $GSD_HOME/.gsd     — GSD's user-owned store. Watched WHOLESALE: unlike ~/.claude
                       this root is exclusively ours, so the shared-root
                       false-positive trap the module documents does not apply.
  <kimi>/config.toml — the file GSD writes its native [[hooks]] block into. The
                       INVERSE case: ~/.kimi belongs to Kimi CLI, so only the one
                       file GSD writes is watched, never the root.

That asymmetry is why this is not a two-line "add two roots" patch — one target
needs the whole tree, the other needs exactly one file, and collapsing them
either under-watches the store or trips the guard's own documented
false-positive trap on a third party's directory.

Extras are passed to snapshotLiveConfig explicitly rather than resolved inside
it, so a caller snapshotting a fixture root cannot silently pull the developer's
real ~/.gsd into its own assertions. run-tests.cjs now snapshots when EITHER the
roots or the extras are non-empty — previously an unbuilt tree yielding zero
roots disabled the entire guard without saying so.

Budget coverage. The MAX_ENTRIES/MAX_DEPTH bound and the truncated -> 'unverified'
branch had zero tests, despite this module's own docstring naming "a truncated
scan reading as clean" as the safety-critical case. Added per
RULESET.TESTS.boundary-coverage (N in {limit-1, limit, limit+1}, exercised
through newestMtime's injected budget so the boundary is real without
materialising 20000 files) and RULESET.TESTS.property-based-testing (fast-check:
truncation is monotone in the budget; reported newest never exceeds the true
maximum). A regression flipping `truncated` to false on an exhausted budget now
breaks the property for every budget below the tree size.

Negative-controlled: neutering the extras wiring fails exactly the two
new-surface tests and nothing else. 21/21 green with it restored.
2026-08-08 05:50:36 -05:00
0xdhx
3c580b77dc fix(#2665): derive the second config-location family instead of hand-adding it
Review round 2 named GSD_HOME and KIMI_SHARE_DIR as missing from the derived
scrub set. Both premises confirmed; the prescribed remedy is not adopted
verbatim, because adding two more literals to a four-item hand list is the
pattern that reopened this bug three times. The census the derivation
generalizes over was partial, so the census is what widens.

Two structural gaps, both closed at the source:

1. KIMI_SHARE_DIR lived inside resolveKimiHooksTomlDir's body as an inline
   descriptor, resolvable but not ENUMERABLE. Hoisted to an exported
   KIMI_HOOKS_TOML_DESCRIPTOR and collected in
   NON_REGISTRY_CONFIG_HOME_DESCRIPTORS, which TEST_ENV_BASE now derives from.
   kimi is the sharp case: it owns TWO config homes (KIMI_CONFIG_DIR, already
   registry-visible, and this one), so a registry-only derivation looks
   complete and is not.

2. GSD_HOME is a different FAMILY, not a missing registry entry. The registry
   describes where third-party runtimes keep config; GSD_HOME decides where GSD
   keeps its own user-owned state ($GSD_HOME/.gsd/ — consent.json,
   defaults.json, capability overlays), read env-first ahead of os.homedir() by
   capability-loader, capability-consent, capability-state, capability-writer,
   config-loader, install-profiles and bin/install.js. Named as
   GSD_LOCATION_ENV_KEYS rather than folded into the descriptor array, since it
   does not resolve through resolveConfigHomeFromDescriptor.

GSD_AGENTS_DIR joins the same family (round 2, Minor): env-first and
unconditional in getAgentsDir, misdirecting a read rather than a write.

A census of every env-first first-party location var — the guard-shape question
this PR owes each round — now yields exactly one remaining unguarded name,
GSD_MODEL_CATALOG, and it is dead by precedence: the co-located candidate is
index 0 and the loop breaks on first success, so the env var can only win on a
tree that is already broken, and it redirects a read even then.

Derived set: 39 -> 42 keys. resolveKimiHooksTomlDir behaviour unchanged on both
the default and the KIMI_SHARE_DIR override path.
2026-08-08 05:50:36 -05:00
0xdhx
e2eed1c58a test(#2665): ship the hermeticity guard at report level, not fatal
Its first CI run found PRE-EXISTING leaks on the Windows lane —
C:\Users\runneradmin\.claude\gsd-core and skills\gsd-dev-preferences — with all
1196 Windows tests otherwise passing. os.homedir() reads USERPROFILE on Windows,
and ~190 test sites across 31 files sandbox HOME alone, so the suite has been
installing GSD into the runner's real home directory invisibly. That is exactly
the class the guard exists to surface, and exactly the class this PR's review
said CI could never catch.

It is also a different defect from the one #2665 closes, and too large to fold in
here. A brand-new gate that immediately reds an unrelated lane gets bypassed or
reverted rather than obeyed, so the guard reports by default and fails only under
GSD_STRICT_LIVE_CONFIG_GUARD=1.

This is the repo's own established ratchet, not a hedge: the local/no-source-grep
ESLint rule shipped at `warn` and was promoted to `error` after its cleanup sweep
(ADR 452). Promote this the same way once the USERPROFILE sweep lands.
2026-08-08 05:50:17 -05:00
0xdhx
5863351281 test(#2665): separator-safe containment in the regression assertion
startsWith(ambientConfigDir) also matches a sibling like
<tmp>/ambient-live-config-2, so it can report a leak that did not happen. Use
path.relative and check for '..' or an absolute result, the repo's usual shape.

The readdirSync assertion already carried the test, so this is cosmetic.

Addresses review finding: Nit 9.
2026-08-08 05:50:17 -05:00
0xdhx
771980b661 test(#2665): restore fallback-branch coverage in the #2003 regression
This PR fixed the test's real defect -- it compared the child's answer against
the PARENT process's getGlobalConfigDir(), two different environments, agreeing
only because the child inherited the developer's ambient CLAUDE_CONFIG_DIR -- but
fixed it by INJECTING CLAUDE_CONFIG_DIR, which moved the test onto the env-first
branch and silently dropped the only coverage #2003 had of the home-derived
fallback.

Sandbox HOME/USERPROFILE and leave the config vars blank instead: the expectation
stays test-controlled AND the branch under test is unchanged.

Drops the notStrictEqual against the codex dir. It could not fail whenever the
strictEqual on the line above passed.

Addresses review finding: Minor 7.
2026-08-08 05:50:17 -05:00
0xdhx
a4efa4deda test(#2665): cover the derivation and the in-process scrub
The single regression test exercised CLAUDE_CONFIG_DIR only, so deleting
GSD_RUNTIME or CODEX_HOME from any literal broke nothing -- a mutation of 24 of
the 27 added key-value pairs survived.

Four tests here: every configHome env var the registry declares is scrubbed;
every scrubbed key is blanked rather than merely present; the four non-registry
vars are named explicitly so deleting one is a failure rather than a silent
narrowing; and scrubConfigLocationEnv round-trips both a set and an unset var
(restoring an originally-unset var as '' would itself be a leak).

Both derivation tests assert a floor on the registry first, so a renamed registry
shape fails loudly instead of making the assertions vacuously true.

Negative-controlled against the hand-written 3-key list this PR shipped: the
parity test fails there and names all 19 missing vars.

Addresses review findings: Major 6, Minor 8.
2026-08-08 05:50:17 -05:00
0xdhx
a02462e050 test(#2665): fail the suite when it writes into a live config dir
The recurrence guard, and #2665's own "Optional hardening". This class is silent
by construction: TEST_ENV_BASE cannot see an in-process caller, and CI cannot see
the class at all because CI never has these env vars set. It damages the
developer's machine and reports nothing -- which is how two prior authors each
diagnosed it and fixed only the instance in front of them.

run-tests.cjs snapshots GSD's install footprint in every live runtime config dir
before the suite and re-checks it after, failing the run on a create or a modify.
Roots come from the product's own getGlobalConfigDir, so the guard watches
wherever the product actually points, including through an ambient var.

Scope is ownership-based, not whole-root: the top-level install footprint plus
gsd-prefixed children of dirs GSD shares with the host agent. A config root like
~/.claude is shared, and watching it wholesale would false-positive on the host's
own history.jsonl or settings.json -- a guard that cries wolf gets disabled, and
then catches nothing. The prefix test is load-bearing: the first version watched
only the three top-level entries and MISSED a real leak into skills/gsd-*.

It earned its place immediately -- it is what found the fifth in-process leak in
runtime-artifact-layout.test.cjs, which no amount of reading the review would have
surfaced. Known gap documented in the module: a write to a file GSD does not own
is out of scope by construction.

Lives in scripts/, deliberately NOT scripts/lib/ -- the installer copies that dir
into every user's config dir wholesale while uninstall removes only an allowlist,
so a test-only module there would ship to users and survive uninstall.

Addresses review finding: Minor 8.
2026-08-08 05:50:17 -05:00
0xdhx
36f4f9479a fix(#2665): scrub config-location env on raw spawns that sandbox only HOME
Same class as the in-process leak, on the child-spawn side. These call spawnSync
directly rather than through runGsdTools, so TEST_ENV_BASE never applies to them
and an ambient config-location var survives into the child.

  - install-runtime-artifacts: the dev-preferences writer resolved env-first and
    wrote SKILL.md into the live config dir instead of its tmp HOME. This also
    fixes a test that FAILS today on any machine with CLAUDE_CONFIG_DIR set.
  - issue-766: the `claude` CLI is third-party and bootstraps its own config into
    whatever CLAUDE_CONFIG_DIR names, so even a bare --version probe wrote there.
    It gets an explicit throwaway dir rather than a blank value -- GSD's resolvers
    treat '' as falsy and fall back to the home dir, but a third-party binary
    offers no such guarantee (blanking it produced a stray backups/ in the repo
    root).

These two were the last writers standing between the suite and #2665's stated
acceptance criterion.
2026-08-08 05:50:17 -05:00
0xdhx
466db2f1a3 fix(#2665): scrub config-location env on the parent for in-process install()
The blocker the review said decides this PR. These tests call the real installer
IN-PROCESS with only HOME/USERPROFILE sandboxed. getGlobalConfigDir is env-first,
so an ambient CLAUDE_CONFIG_DIR beats the sandbox and a complete global install --
agents/, commands/, skills/, gsd-core/, manifest, settings -- lands in the
developer's live config dir. No child-env scrub can reach it; only clearing the
parent's env can.

The review named two describe blocks in install.test.cjs. A sweep for the shape
found four there (bug #3571 and bug #3288 each appear twice in the file), and the
post-suite guard added later in this series found a fifth in
runtime-artifact-layout.test.cjs, which the review did not name. All five now
save/clear/restore via scrubConfigLocationEnv().

Verified: codebuddy-install, cline-install and codex-config already guard their
own config-location var around in-process install(), so the class is closed.

Addresses review finding: Blocker 1.
2026-08-08 05:50:17 -05:00
0xdhx
bc5c362610 fix(#2665): replace the hand-synced TEST_ENV_BASE copies with the canonical import
Eight declarations were kept in sync by hand with no parity assertion. The drift
was already in the tree: api-coverage-gate-e2e and representative-corpus declared
TERM_SESSION, but the real variable is TERM_SESSION_ID, so both scrubbed nothing
for that slot and leaked TERM_SESSION_ID into every child. Deleting the copies
removes the dead key with them -- there is no longer a second place to get wrong.

run-tests-harness keeps its local session-identity literal: that helper mirrors
the production runGsdTools to prove its CONTRACT, so it must not re-import what
it is testing. The config-LOCATION keys are a safety scrub rather than part of
that contract, so it spreads the canonical derived set and keeps the rest local.

Addresses review findings: Blocker 2, Blocker 3.
2026-08-08 05:50:17 -05:00
0xdhx
0f97a26f76 fix(#2665): derive the config-location scrub set from the capability registry
TEST_ENV_BASE listed three config-location vars by hand. The resolver reaches
25: every runtime descriptor's configHome.env, plus GROK_AGENTS_HOME (a hardcoded
branch of getGlobalConfigDir), GSD_RUNTIME, and GSD_PROJECT/GSD_WORKSTREAM
(planningDir). A hand-written list can only ever be as complete as the author's
recall, and every one of those resolvers is env-FIRST, so a missing key is a live
escape hatch rather than a cosmetic gap -- which is why this bug has now been
diagnosed three times.

Derive the set from the same registry the resolver reads. The scrub list becomes
structurally incapable of being narrower than the surface it guards: adding a
capability that declares a new configHome env var extends it in the same commit.

Also export TEST_ENV_BASE (a parity test could not previously import the
canonical copy) and add scrubConfigLocationEnv(), the in-process counterpart --
TEST_ENV_BASE only ever reaches child processes.

Addresses review findings: Blocker 2, Major 4 (GSD_WORKSTREAM/GSD_PROJECT),
Major 5 (XDG_CONFIG_HOME).
2026-08-08 05:49:33 -05:00
0xdhx
08021b02c0 fix(#2665): scrub config-location env vars in every TEST_ENV_BASE declaration
TEST_ENV_BASE blanks session-identity variables but none of the three that
decide WHERE a child process writes: CLAUDE_CONFIG_DIR, GSD_RUNTIME and
CODEX_HOME. The config-home resolver is env-first (runtime-homes.cts, the
dot-home case consults the env var before the home-derived fallback), so an
ambient CLAUDE_CONFIG_DIR in the developer's shell beats a call site that
sandboxes only HOME. The suite then writes into the developer's real config
directory -- including a registered skill under <configDir>/skills/ whose
body carries behavioural directives that load into later sessions.

Blank all three alongside the session-identity vars. `...env` still spreads
last, so the five call sites that already constrain these locally keep
winning with their explicit values.

TEST_ENV_BASE is re-declared in nine files, so the three lines are added
nine times rather than once. Consolidating the nine into a single exported
constant -- and fixing the TERM_SESSION / TERM_SESSION_ID drift between the
copies -- is deliberately left out of this change; see the PR body.

One call site needed adjusting. capability-state.test.cjs's
`capability state --runtime claude` CLI test passed no env at all and
compared the CHILD's resolved config dir against the PARENT process's
getGlobalConfigDir('claude'). That agreed only because the child inherited
the developer's ambient CLAUDE_CONFIG_DIR -- i.e. it passed *because of*
the leak. It now redirects both runtime homes into the sandbox and asserts
against values the test controls, so it is hermetic with the variable set
or unset.

Regression case folded into the owning module's test file rather than a new
bug-NNNN file, per scripts/lint-regression-test-names.cjs. It sets the
variable on the PARENT process, which is the actual vector; setting it in
the per-call env argument would exercise a path that was never broken.
2026-08-08 05:49:33 -05:00
Tom Boucher
343835facc refactor(#3183): route live-plan counting through scanPhasePlans (#3199)
* refactor(#3183): route live-plan counting through scanPhasePlans

scanPhasePlans becomes the sole owner of the live-plan derivation. Twenty-one
independent re-derivations across seven modules now route through it, and
scripts/lint-plan-count-drift.cjs reports zero, scanning the whole repo rather
than an allowlist (ADR-3180 Decision 4a).

The epic scoped this at three copies. A whole-repo guard found twenty-six sites
across nine files, so Phase 1 absorbs every live-plan re-derivation and Phase 3
narrows to window plus sentinel enumeration.

Two sites are exempt with a documented reason rather than a bare allowlist:
audit.cts scans one quick task's own directory for a single completion record,
and gsd2-import.cts reads a foreign GSD-2 tasks/ layout during a one-time
import. Neither is a phase directory.

scanPhasePlans gains allPlanFiles (pre-supersession) alongside planFiles so one
owner answers both questions: verify.cts's numbering-gap check wants every plan
on disk, its pairing check wants the live set. Both fields are additive.

Highest-severity fix: cmdPhasePlanIndex, which feeds execute-phase wave
scheduling, was scheduling status:superseded plans into waves and reporting zero
plans for the post-#3139 nested layout.

filterPlanFiles and filterSummaryFiles are deleted; getPhaseFileStats orphaned
them and only their own tests still called them.

New leaf module src/planning-scope.cts carries the frozen SCOPE discriminator,
with its six-gate ripple closed: gitignore, inventory manifest, INVENTORY.md and
the CONTEXT.md glossary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* docs(#3183): amend ADR-3180 for the Phase 1/3 boundary re-slice

The contract held; the phase boundary did not. The whole-repo drift guard found
26 re-derivations across 9 files against the epic's estimate of 3, and
cmdProgressRender re-derives both enumeration and plan counting on adjacent
lines, so DW4 was unsatisfiable within Phase 1's original file scope.

Records the amended scope, scanPhasePlans's new allPlanFiles field,
findOrphanSummaries, the two documented exemptions, the re-derived Tier-2
table, and the describeNonCanonicalPlans trap for later phases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): complete the canonical pairing rule and gate the naming diagnostic

The remote runner went red with 13 deterministic failures on both lanes,
and they were right: replacing verify.cts's canonicalPlanStem pairing with
summaryCandidates dropped a case the bespoke rule covered. A plan carrying a
descriptive slug after its id (68-01-scaffolding-PLAN.md) pairs with its
canonical-stem summary (68-01-SUMMARY.md), and summaryCandidates generated no
such candidate, so the plan read unsummarized.

The fix is to complete the one rule rather than restore a second:
summaryCandidates gains a canonical-id candidate, narrowed to fire only when an
id pair was actually extracted. countMatchedSummaries, findUnsummarizedPlans
and findOrphanSummaries all inherit it. The two-plans-one-summary collision
behaviour of the original rule is preserved deliberately and documented in
place.

Second defect, independently root-caused while verifying: routing the #2893
naming diagnostic through scanPhasePlans exposed it to the loose /PLAN/i
fallback, which is correct for counting and wrong for a naming check — a
non-canonically-named file was accepted as a valid plan and the diagnostic
went silent. cmdPhasesList, cmdFindPhase and cmdPhasePlanIndex now intersect
with a strict isCanonicalPlanFile predicate before reporting names.

Same class as the describeNonCanonicalPlans trap already recorded in ADR-3180:
a question about file naming wants the physical, strictly-matched set; only a
question about outstanding work wants the live set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): register planning-scope.cjs in the eslint migration list

tests/repo-invariants.test.cjs asserts every bin/lib/*.cjs is linted xor
ignored per its ADR-457 migration state. The new planning-scope module closed
five of the six .cts ripple gates - gitignore, inventory manifest, INVENTORY.md
and the CONTEXT.md glossary - but not eslint, because that one is enforced by a
test rather than by lint:ci, so the local pipeline stayed green while it was
missing.

Generated from src/planning-scope.cts, so the .cjs is ignored and the .cts is
linted, matching every other migrated module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* fix(#3183): replace the plan-count drift detector with a literal tokenizer

CodeQL reported 4 high-severity js/redos alerts on REGEX_LITERAL_MD_RE, the
backtracking regex that finds "a regex literal mentioning PLAN/SUMMARY and an
escaped \.md". Five review rounds found it had two defects, not one:

  - EXPONENTIAL, then CUBIC. Its "any char" atom `(?:\\.|[^/\r\n])` let a `\.`
    pair be consumed either as one escape or as two class characters, which is
    exponential backtracking: 27,464ms on `"/\.mdplan" + "\.".repeat(28) + "X"`.
    Excluding `\` from the class killed that but left a cubic path — 23ms at
    N=200, 172ms at N=400, 1362ms at N=800 on `"/" + "PLAN\.md".repeat(N)` with
    no closing `/`. This guard is the last stage of `npm run lint:ci`, which CI
    runs on fork pull requests, so a crafted src/*.cts could stall the job.
  - A DETECTION HOLE. A character class holding a bare, unescaped `/` — e.g.
    `/SUMMARY[^/]*\.md$/`, an ordinary path-excluding filter — terminated the
    literal at that `/`, so the scan never reached `\.md` and the guard missed
    it entirely. (Classes holding an ESCAPED `\/` were already matched; the
    tests cover those separately as parity, not as regressions.)

Both defects have one root cause: regex-literal grammar — `\x` escapes, and
`/` inside `[...]` not terminating — is not expressible in a backtracking
regex. So the detector is now a tokenizer, not a regex.

readRegexLiteralAt reads the literal at a given `/` in a single left-to-right
pass with no backtracking, treating escapes as two-character units and
suppressing the `/` terminator inside a character class. findRegexLiteralMdMatch
restarts it at every `/` on the line, preserving the old "find anywhere"
behaviour; MAX_REGEX_LITERAL_LEN (400) bounds each read — including the
trailing-flag scan — which keeps the whole-line cost linear.

Results: cubic shape flat at 0.06-0.39ms out to N=3200 (25KB), exponential
shape 0.01ms at 28 reps and 0.00ms at 64, and the bare-`/` class shapes are now
caught. Differential against the old regex over 28,474 lines (those matching
FILENAME_TEST_RE but not PLAN_SUMMARY_LITERAL_RE, across src/tests/scripts/
gsd-core/bin/eslint-rules, excluding 265 lines with >6 backslashes on which the
old regex hangs): 6 differences, all the tokenizer returning the fuller or
newly-correct literal, 0 old-only misses. The `\.md` token stays
case-insensitive, matching the `/i` the old regex carried.

Also closes three holes in the same new file:

  - walk() tested entry.isFile(), false for a symlink, so a symlinked
    src/*.cts was silently unscanned — an evasion of a guard whose stated
    principle (ADR-3180 Decision 4a) is whole-repo discovery with no allowlist.
    It now resolves symlinks, but confined: file links must resolve inside the
    repo root, directory links inside the scanned dir itself. Every sibling
    drift guard in scripts/ uses the Dirent classification and never follows
    links, so following them unconfined would have made this the only linter
    able to read outside the tree — on fork PRs an arbitrary out-of-repo read
    whose matched fragments reach a public CI log. The narrower directory rule
    additionally stops `src/up -> ..` from sweeping the whole repo, and the
    skip list is now checked against resolved paths so `src/g -> ../.git`
    cannot reach .git/** or node_modules/**. Real paths are de-duplicated and
    files reported canonically, so a symlink alias cannot shift which
    FUNCTION_SCOPED_EXEMPTIONS key applies.
  - Both the reported fragment and the reported FILE PATH are attacker-
    controlled source text written straight to a CI log, and git permits
    control bytes in a filename. Both are now escaped — C0/C1/DEL plus the
    bidi and zero-width controls — so a crafted literal or filename cannot
    recolour the log, overwrite a line with CR, or fabricate a line that looks
    like this guard's own success output.

Regression coverage in tests/plan-count-single-owner.test.cjs: a child-process
probe over both pathological shapes (catastrophic backtracking is synchronous
and would freeze the suite rather than fail one test), the bare-`/` class
shapes verified to fail against the parent-commit blob, root-confinement tests
covering the outside-file, outside-directory, cycle, broken-link and duplicate
cases, direct isInsideRoot coverage including the sibling-prefix case that a
bare startsWith would let through, sanitizeForReport coverage, and
limit-1/limit/limit+1 coverage of MAX_REGEX_LITERAL_LEN derived from the
exported constant. The earlier structural assertion was dropped — it checked
for the substring `[^/`, which respelling the class as `[^\r\n/]` defeats
while staying exponential.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* chore(#3183): backfill changeset PR number

Restores b77931869, which a force-push during the ReDoS remediation dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 01:20:55 -04:00
Tom Boucher
409da3223c Merge pull request #3202 from open-gsd/chore/backmerge-main-to-next-62f44f8d
chore: back-merge main → next (62f44f8d)
2026-08-08 01:08:13 -04:00
github-actions[bot]
9f22fc63e1 chore: back-merge main into next (62f44f8d) 2026-08-08 05:07:50 +00:00
Tom Boucher
97327b994e Merge pull request #3201 from open-gsd/chore/sync-next-version-1.10.0
chore: sync next package version to 1.10.0
2026-08-08 01:07:40 -04:00
github-actions[bot]
4f2a4932ce chore: sync next package version to 1.10.0 2026-08-08 05:07:32 +00:00
Tom Boucher
62f44f8d7c Merge pull request #3200 from open-gsd/release/1.10.0
chore: merge release v1.10.0 to main
2026-08-08 01:07:30 -04:00
github-actions[bot]
68a04ccf8e chore: promote CHANGELOG for v1.10.0 2026-08-08 05:06:54 +00:00
github-actions[bot]
9335dc00d2 chore: bump version to 1.10.0 for release 2026-08-08 04:30:53 +00:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
664d49e513 docs(#3182): ADR-3180 — planning semantic model single owner (#3196)
* docs(#3182): ADR-3180 — planning semantic model single owner

Phase 0 design lock for epic #3180. Names one canonical owner per
semantic derivation, specifies the frozen-enum scope contract that
distinguishes a genuinely-empty computation from a truncated or
unscoped one, and locks the drift-guard contract.

Ships no production code. Phases 1-5 execute against this ADR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

* test(#3182): prune real issue 3182 from phantom-ref guard, fix empty-list regex

The guard's own header documents that its list rots: entries are phantom
only until the repo's shared issue/PR counter reaches them, and once the
counter passes an entry it must be deleted. Creating the Phase-0 sub-issue
advanced the counter past 3182, so the guard began rejecting a legitimate
citation of a real issue - the failure its header already records happening
twice, with PRs 2551 and 2361.

3182 was the last entry, and removing it exposed a latent bug: the regex
builder interpolated the list unconditionally, so an empty list yields
(?:#(?:)\b)|(?:issues/(?:)\b), whose empty alternation matches every issue
reference in the repo. Following the file's own maintenance instruction
would have turned a green guard into one failing on nearly every file.

buildRefRe() now returns null for an empty list and is exported, with
boundary coverage at 0/1/2 entries plus word-boundary and bare-digit
negative cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 20:25:12 -04:00
Tom Boucher
3f349e551d fix(#3024): route sync-skills through the shipped gsd-tools instead of an unshipped install.js (#3195)
* fix(#3024): sync-skills workflow uses gsd-tools query skills-root instead of unshipped install.js

The sync-skills workflow Step 2 shelled out to gsd-core/bin/install.js --skills-root,
but install.js is not shipped in installed trees (only in the npm tarball root bin/).
Every /gsd-update --sync invocation failed with MODULE_NOT_FOUND.

Fix: added 'gsd-tools query skills-root <runtime>' subcommand (gsd-tools IS shipped)
that calls the same getGlobalSkillsBase function install.js used. Updated the
workflow to call gsd_run query skills-root instead of the dead install.js path.

Also documented the #3025 verbatim-cp limitation in Step 5 with a workaround.

* test(#3024): failing-first guards for the three defects in the adopted fix

The cherry-picked commit came from an aborted run that never executed its own
tests. Its raw-path assertion fails as written, which is the clearest evidence
the work never reached verification.

Covers:
- --raw must emit a bare path, not JSON (output() takes a third rawValue arg
  that routeSkillsRoot omits, so the raw branch never fires)
- an unknown, empty, whitespace, traversing, or metacharacter-bearing runtime
  must be rejected, not silently resolved to claude's skills root
- sync-skills.md must contain zero references to the unshipped install.js,
  including the guard's remediation text — the issue's second reported defect
- parity across every runtime in the registry, not three hardcoded ones, so the
  two entry points cannot drift

Also converts the adopted tests off a hand-rolled spawnSync onto the bounded
process seam, per CONTRIBUTING.

Fails before the fix. Verified via the remote runner.

* fix(#3024): make the skills-root query actually work and reach non-Claude runtimes

The cherry-picked commit never ran its own tests. Six defects, all fixed here.

--raw was ignored: output() is output(result, raw, rawValue) and the third
argument was omitted, so the raw branch never fired and the workflow captured a
JSON blob as SRC_SKILLS_ROOT. Every downstream cp -r then resolved against a
nonexistent path — the command would have shipped still broken.

An unknown runtime silently resolved to claude's skills root, because
getGlobalSkillsBase falls back rather than returning null, leaving the existing
=== null guard dead. The runtime id is now validated at the CLI boundary against
the shipped registry, so a typo'd --from/--to fails instead of reading from or
writing into the wrong runtime's tree.

getGlobalSkillsBase('vscode') threw a raw TypeError. vscode is non-installable
by descriptor, so it has no skills root — null is the answer, not a crash. The
resolver now short-circuits configHome.kind 'none', which also fixes the same
latent crash in install.js --skills-root vscode. Every caller already gates on
=== null.

sync-skills.md used gsd_run WITHOUT the canonical launcher preamble, so gsd_run
was undefined on non-Claude runtimes — the fix would have been dead in exactly
the place the original bug bit. Preamble propagated via sync-runtime-launcher.

Also registers skills-root in TOP_LEVEL_USAGE (the help/dispatch parity guard
caught it), removes the last two install.js references including the guard's
remediation text (the issue's second reported defect), and updates the stale
assertion that still described the removed contract.

Verified on the remote runner.

* fix(#3024): align the documented runtime list with the registry and gate both entry points

Isolated review returned BLOCK on two findings.

The workflow's Supported-runtimes list and its --to all expansion named grok and
gemini, neither of which is a registered runtime. Once this branch added
validation, --to all — a documented first-class feature — aborted. The list was
hand-copied prose shadowing the registry, so correcting it alone would drift
again; a parity assertion now fails in BOTH directions if the doc and the
registry disagree. vscode is excluded by name: it is installSurface 'none', so
syncing skills to it is meaningless and would abort.

bin/install.js --skills-root reached getGlobalSkillsBase with no own-property
gate, so --skills-root __proto__ silently resolved to claude's skills root. This
branch had just hardened the OTHER entry point to the same function; leaving one
of two parallel surfaces open is the same divergence class as the first finding.
Both now call one shared isRegisteredRuntimeId() rather than a copied check, and
the parity test covers the hostile ids so the two can never disagree again.

Also guards the workflow's root resolution: neither command substitution checked
its exit status and only the source had an existence guard, so a failed
destination resolution left DEST_ROOT empty and turned rm -rf "$DEST_ROOT/$SKILL"
into an absolute path at filesystem root. Both resolutions are now checked, and
Step 5 requires both roots to be non-empty and absolute before any destructive
command.

Verified on the remote runner.

* test(#3024): anchor the runtime-list parity extractor to the list span

The extractor captured (.+) to end of line, so it swallowed the em-dash prose
that explains the vscode exclusion — and that sentence contains backticked
`runtimes` and `null`, which is where the three phantom ids came from. The
documented list was correct; the test was reading its own explanation back as
data. Anchored to the id-list span.

Both directions still fail as intended: proven by injecting a bogus id and by
removing a registered one.

* test(#3024): anchor the --to all extractor and fail loudly on empty captures

The workflow has three TO_RUNTIMES= assignments and the regex matched the first
one — an empty array initializer at line 28 — so the extractor captured nothing
and the assertion diffed [] against 18 ids as if that were data.

That is the same failure twice, so the fix is the general one: every extractor
in this test now asserts it captured a plausible list before comparing, naming
which extractor found nothing and what it was looking for. An extractor that
silently yields [] is a confident wrong answer, and a parity guard that reports
it as a data mismatch teaches the reader to loosen the assertion.

Verified against the real workflow and against doctored copies with each target
construct removed, plus both teeth directions.

* fix(#3024): merge duplicate process-seam import after rebase

The rebase applied cleanly but left runNode declared twice: next had gained its
own import of the seam while this branch added one carrying OUTCOME. A clean
rebase is not a correct one — the file no longer parsed. Merged into a single
import providing both.

* fix(#3024): bind DEST_ROOT per destination instead of a dangling map

Step 2 stored each destination's root into DEST_SKILLS_ROOTS, which nothing ever
read, while Steps 3 and 5 used a scalar DEST_ROOT that nothing ever assigned. The
array was also never declare -A'd, so on bash 3.2 — macOS system bash, which this
repo supports — every destination collapsed onto index 0.

The absolute-path guard added earlier was the only thing standing between that and
rm -rf "/$SKILL"; it turned a silent disaster into a hard stop, but the feature
still could not complete. Each destination now binds its own DEST_ROOT where it is
used, and the unread map is gone rather than replaced.

Step 2 keeps eager validation, so a bad runtime id in a multi-destination --to
aborts before any destination is written rather than after some already have been.

Verified on bash 3.2 with a two-destination run binding distinct roots, and with a
bad id aborting before any destructive call.

* fix(#3024): restore grok support broken by the registry gate

The registry gate added earlier rejected grok, and that was my error. I confirmed
grok was absent from the capability registry and concluded the hardcoded branch
was dead — without checking what it resolved to. It resolves to ~/.agents/skills,
a real grok-specific path, exactly as the pre-fix workflow documented ('grok uses
the ~/.agents layout'), and there is a support discussion doc for it. So a
working, documented runtime silently lost --skills-root and sync-skills support
as a side effect of prototype-pollution hardening — and the parity test I added
locked that in as correct.

gemini is the one that really was dead: it fell through to CLAUDE's skills root,
so rejecting it is right and it stays rejected, as do bogus ids, __proto__,
empty, whitespace and traversal.

The validator's real question is 'does this id have a genuine runtime-specific
resolution', not 'is it in the registry map'. Registry membership was a proxy
that happened to miss grok. Legacy non-registry runtimes with dedicated
resolution branches are now a named, documented set; enumerating every hardcoded
branch in getGlobalConfigDir against the registry confirms grok is the only one.

The new tests assert grok resolves UNDER .agents and specifically not to claude's
root. Allow-listing an id proves nothing about whether it resolves correctly —
that assertion is what would have caught my mistake.

Also uses the shared PROBE_TIMEOUT_MS instead of a duplicate literal, and guards
Step 3's DEST_ROOT re-resolution, which contradicted the file's own stated
guarantee.

Verified on the remote runner.

* test(#3024): guard against LEGACY_NON_REGISTRY_RUNTIME_IDS drifting

The named legacy set is a second hand-maintained proxy for the same predicate
the registry check got wrong — 'does this id resolve runtime-specifically'.
Nothing stopped a third hardcoded branch being added to getGlobalConfigDir
without updating the Set, reproducing the exact class of bug that broke grok.

Production stays explicit and greppable; the test derives the truth instead. It
resolves a sentinel id to learn the generic fallback, classifies every candidate
against it, and fails in both directions — an id resolving runtime-specifically
that is in neither the registry nor the Set, or a Set entry that no longer earns
its exemption. The failure message names the remedy.

Confirms grok resolves runtime-specifically and gemini does not, which is the
distinction the original registry check could not see.

Also reverts the shared-timeout swap: SKILLS_ROOT_PROBE_TIMEOUT_MS is
pre-existing on next and arrived by rebase, so changing it here was scope creep
into another issue's territory.

Verified on the remote runner.

* chore(#3024): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-07 18:53:25 -04:00
Tom Boucher
cbd180c5cd test(#3147): bound the lint/changeset/docs cluster onto the process seam (#3181)
* test(#3147): bound the lint/changeset/docs cluster onto the process seam

Migrates 69 unbounded sync spawn sites across 24 files. Allowlist 73 to 49.

Two shared helpers move: tests/helpers/graphify.cjs (6 importing suites) and
tests/fixtures/index.cjs, whose three quoted-argument shell strings became
single argv elements rather than whitespace splits.

changeset-lint's throw-native git() helper routes to gitOrThrow; migrating it
to bare runGit would have silently swallowed a failure that is loud today.
ingest-docs goes the other way -- its catch never rethrew, it degraded failure
into data every call site asserts on, so throwIfFailed would have thrown where
the original returned. The design doc said otherwise and was corrected.

tsconfig-noemit runs a real tsc --noEmit and takes a bespoke 180000ms per the
ensure-runtime-build precedent, not the 30000ms build-hooks norm -- that norm
is for a file copy, and sizing against a label rather than the work is the
same error in the opposite direction.

* test(#3147): add toLegacyResult and settle review findings

The seam exposed a throwing adapter (throwIfFailed) but no non-throwing one,
so eight files independently re-derived the same unwrap back to the legacy
{status, stdout, stderr} shape. That is the third time this epic produced N
copies of one mechanism -- seven throw wrappers in Wave 1, fifty-two timeout
constants in Wave 2, eight result adapters here. The pattern is that whenever
the seam does not expose a mechanism, every suite re-derives it.

toLegacyResult now sits beside throwIfFailed, with its own tests.

Two sites are deliberately NOT converted: changeset-cli's runRender and
runRenderIn return {status, report, stderr} from parsed JSON and never a raw
stdout, so they are a different shape family. lint-legacy-dir-name keeps its
local GUARD_TIMEOUT_MS: 30000 matches the build norm numerically but bounds a
lint probe, not hooks bundling, and importing it would encode a coincidence
as a relationship.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:45:51 -04:00
Tom Boucher
3fac6e629f test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam

Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73.

Timeouts are sized from evidence already in the tree rather than a house
default, because this wave spawns installers rather than git plumbing and an
undersized bound does not catch a hang -- it manufactures CI flake, which is
worse, since a flake gets re-run instead of investigated. install.test.cjs
records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while
another lane passed the same commit in 12.7s, so full installs are bound at
120000ms against that recorded incident.

Also adds an auditable escape to the guard's timeout ceiling. The 600000ms
cap was set in #3143 from partial evidence, but fragment-single-edit-
propagation carries a documented, load-tested 900000ms bound on a run that
chains a full build plus eight generators -- the guard would have rejected a
correct timeout the moment that file left the allowlist. A value above the
ceiling is now permitted only with an inline allow-spawn-timeout-ceiling
marker carrying a non-empty reason. It raises the ceiling; it never waives
the requirement for a bound, which is asserted directly.

install-shared.cjs keeps its hand-rolled assert rather than routing through
throwIfFailed: its message embeds both streams, and throwIfFailed carries
only a trimmed stderr. The message now also names the outcome, so a bounded
timeout reads as such across its 38 importers instead of as
expected null to equal 0.

* test(#3145): extract class-norm timeouts and correct the build-hooks sizing

A pre-PR review found 52 copies of four class-norm timeout constants across
this wave. These are not per-suite fixture bindings -- they are shared facts
about how long a class of subprocess takes, derived from a recorded bench
incident. That norm already moved once (60000 to 120000 after a real
ETIMEDOUT), and 52 copies would have drifted the next time it moved.

Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and
converts the copies. A site that genuinely differs -- a real tsc compile, or
regen:derived -- keeps its own local constant with its own justification.

Also corrects a misclassification: scripts/build-hooks.js was sized as a
build at 120000 in twelve places and 60000 in another, but it compiles and
bundles nothing. Its own header says no bundling needed; it copies pre-built
files and syntax-checks them with vm. Three different values bounded one
script; now there is one.

* test(#3145): fix red CI — lint self-match and a Windows chunk overrun

Two failures on PR 3176.

lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The
fixture exists to prove an unrelated marker does NOT suppress the rule, so it
carries that marker's literal text as test data. Split via concatenation, the
same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own
self-match problem. The explanatory comment needed the same treatment.

The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped
seven minutes before the kill, so this was an overrun rather than a slow
chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs
regen:derived bounded at 900000ms, which is larger than the whole chunk
budget, so the chunk killer always fires first and it can never complete
there. Both the test and that bound predate this change; modifying the file
pulled it into the Windows targeted set and exposed it. Skipped on Windows
with the reason recorded; the Linux lanes cover it. The 900000 bound and its
ceiling marker are unchanged -- they are correct.

* test(#3145): refresh the stale test-timings cost table

The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs
packs chunks by measured duration from tests/test-timings.json, and an
unknown file falls back to the table's median weight -- advisory by design,
but it silently underweights exactly the files that matter.

Four of the failing chunk's 22 files were absent from the table, including
the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s
(it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s.
Both were weighted as average, so the chunk's total weight read 53.68 against
a budget of 60 and the packer produced a single chunk.

Regenerated from a passing full-suite run, per the remedy the script itself
documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since
gen-test-timings.cjs replaces the table wholesale rather than merging.

Proven against the real packer: the same 22 files now weigh 103.91 and split
into two chunks. No logic, budget, or timeout was changed; raising a budget
to make a red gate pass is not a fix.

---------

Co-authored-by: sim <sim@local>
2026-08-07 15:18:18 -04:00
Tom Boucher
27aa40f65e fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir

pi reserves <configDir>/hooks as its deprecated extension location and warns
on every startup when it exists. Assert a pi install stages the shared hook
bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/.

Also adds pi to the local-scope dir table in install-shared.cjs: pi was in
RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved
path.join(root, undefined) and no local pi install could be exercised.

Fails before the fix. Verified via the remote runner.

* fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir

pi reserves <configDir>/hooks as its now-deprecated extension location and
warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs()
guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged
its shared hook bundle exactly there, and pi's advised remediation (move it to
extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's
extension auto-discovery.

The bundle directory name is now runtime-descriptor-driven: hostBehaviors
.sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are
byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path
segment — separators, dot-only segments, trailing dots, absolute paths, NUL,
and Windows reserved device names all fall back to the default, because the
value is joined onto a user's config root and written to.

Renamed in place rather than relocated: hook scripts resolve siblings via
__dirname/.., so a depth change would silently break them.

- install / uninstall / manifest sites all read the resolved name
- pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded
  trees still resolve; the never-throws contract is preserved
- new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new
  non-recursive remove-empty-dir engine primitive (rmdirSync only,
  symlink-refusing, containment-guarded); ADR-0008 amended accordingly
- fixes two latent name-dependencies the rename exposed: the stale-hook scan
  and the injection scanner's self-exclusion both hardcoded 'hooks'

Verified on the remote runner.

Closes #3023

* fix(#3023): close review findings and align emitted provenance with the rename

Adversarial review found two defects, and the remote runner found four
failure clusters. All fixed here.

Review BLOCKER — detect-custom-files was blind to the renamed bundle.
GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the
whole gsd-hooks/ tree was invisible to the custom-file scan and user-added
files there were never backed up before the next update's clean-install wipe.
The dir set now resolves via the .gsd-runtime marker plus the shipped
capability registry (never bin/install.js, which is not shipped into installed
trees), and falls back to scanning every known candidate when the runtime
cannot be determined — over-scanning is safe, under-scanning is the data loss.

Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir
accepted any directory, so an interrupted install left gsd-hooks/ winning over
a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now
qualifies only if it is non-empty.

Remote-runner clusters:
- emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped
  rules pointing at the same sources the existing hooks/ rules use. The table is
  total, so an unattributed family is a hard failure by design.
- pi tests in install-minimal-hooks and the install integration suite asserted
  the old layout; updated to derive the dir name from the descriptor rather than
  hardcoding either name.
- 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after()
  runs in registration order, cleanup was registered before mock.restoreAll(),
  and node22's JS rimraf calls the public fs.rmdirSync while node24's native
  path does not — so the EACCES stub leaked process-wide on one lane. Restore
  now runs first.

Verified on the remote runner.

* fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde

pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent
(packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty
configHome.env, so a user with that variable set had GSD installed where pi
never looks. Added the env name; the dot-home-nested resolver already handled
the override, so no resolver logic changed.

Also fixes expandTilde in the shared runtime-homes resolver, found while adding
that: it hardcoded os.homedir() and ignored the opts.home every caller threads,
so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi)
silently resolved against the real home. That is a correctness bug and a
test-escape hazard — a sandboxed test asserting on a tilde override reached the
developer's actual home directory. Now threaded through every branch; behavior
with no injected home is unchanged.

Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location
moved with the rename. The provenance rules satisfy the totality gate; the
differential gate needs the ack because the hook sources are byte-unchanged —
only the installer's target directory moved. The two hook files this branch
genuinely edits stay attributed and are not double-acked.

Note on piConfig.configDir: it is read from pi's OWN installed package.json
(getPackageDir walks up from pi's __dirname), alongside piConfig.name — a
white-label setting for a redistributed pi fork, not a per-project user setting.
Documented accordingly rather than treated as an unsupported override.

Verified on the remote runner.

* fix(#3023): reject blank env overrides, pin adapter/descriptor parity

Three review findings, all fixed.

A whitespace-only config-dir override was accepted verbatim: the guard was
`if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR='   '
resolved to a literal three-space directory name instead of falling back to the
descriptor default. Fixed across every env-consuming branch — dot-home,
dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's.
Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working.

pi/gsd.cjs's probe list and the descriptor were two independent sources of truth
for the bundle directory name; a future rename would have desynced them silently
and left every pi hook quiet with no error. The probe list stays deliberate — it
must resolve in a dev checkout and a half-upgraded tree, where the registry's
answer would be wrong — so this adds the parity assertion the repo's
generative-fix-divergence rule calls for: the descriptor value must be the FIRST
candidate, and the default must remain present.

Changeset body rewritten to cover the two later user-facing fixes it had not
caught up with.

Verified on the remote runner.

* chore(#3023): backfill changeset PR number

* fix(#3023): anchor injection-scan patterns and fix a macOS detection hole

CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the
same fact as a genuinely empty or absent one'. The match was the 'act as a'
INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending
in act tripped it (fact, impact, contract, artifact, interact, redact,
abstract). My four-line CONTEXT.md edit dragged the latent false positive into
this PR because the scan is diff-scoped by file but reads whole files. Anchored
with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would
have left the class alive for the next PR touching any file saying 'fact as a'.

Auditing the rest of the list for the same class surfaced a real detection hole:
the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex
escape. BSD/macOS grep reads it as four literal characters, so single-quoted
eval('...')/exec('...') payloads were NEVER detected there while passing on
GNU-grep CI. Replaced with a literal apostrophe class.

Boundaries were added only where a real word-suffix collision exists; exec,
jailbreak, developer mode and the role-manipulation family were audited and
deliberately left unanchored. 22 new cases cover both directions — the false
positives now scan clean, and every real payload still fires, including the
quote/punctuation/start-of-line boundary forms.

Also builds this branch's injection test fixture at runtime instead of carrying
the literal phrase, so the payload keeps its teeth without tripping the scan.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-07 13:41:21 -04:00
Tom Boucher
86f72a57c3 Merge pull request #3173 from open-gsd/docs/3155-adr-default-off
docs(#3155): ADR-3128 — shipped default is off, not adaptive
2026-08-07 11:27:15 -04:00
sim
c643320cef docs(#3155): ADR-3128 — shipped default is off, not adaptive
The maintainer decided the default after the ADR first merged, by
consistency with the shipped workflow config: verification gates default
on (research, plan_check, verifier, nyquist_validation,
security_enforcement), agent autonomy defaults off (auto_advance,
research_before_questions, plan_bounce, cross_ai_execution). Installing
probes into tracked source without a second confirmation is autonomy,
not a gate.

Adds Decision 8, flips the legacy/absent-section default and the
precedence tail to off, and marks Open question 2 resolved.

Decision 1's justification is corrected rather than deleted. It rested
on 'adaptive carries no flag', which the new default makes false. The
conclusion is unchanged and the real reason is stronger: precedence
includes the saved session policy, so a resumed session that persisted
adaptive passes no flag either, and a flag-keyed atom would exclude the
section from exactly the sessions already running the protocol.

Closes #3155

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 11:24:10 -04:00
Tom Boucher
b7fd6e98c9 Merge pull request #3157 from open-gsd/docs/3155-adr-adaptive-runtime-evidence
docs(#3155): ADR-3128 adaptive runtime evidence — Phase 0 design lock
2026-08-07 11:13:30 -04:00
sim
55136a19e9 docs(#3155): ADR-3128 adaptive runtime evidence — Phase 0 design lock
Records the design decisions #3128's maintainer approval made a
condition: schema v1, the probe/artifact ownership model, and the
cleanup state machine that gates terminal transitions.

Also amends ADR-1671 with a RESERVED atom rather than a widening. The
vocabulary stays at 29 until #3128's implementation lands; the
reservation exists so the widening is a coordinated decision rather than
an organic edit found in review.

The load-bearing decision is the atom's shape. #3128's probe policy is
tri-state (adaptive|force|off), so gating on flag:--runtime-probes would
exclude the protocol section from every default invocation -- adaptive
carries no flag -- and the feature's primary mode could never activate.
That is admission gate (2)'s silent-exclusion failure arriving through a
different door: not a fact nobody computes, but a fact computed for only
one of three policies. The atom is therefore a resolved boolean folded
in cmdInitDebug.

Closes #3155

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 11:09:53 -04:00
Tom Boucher
1d208e5af6 test(#3144): bound the git/worktree cluster onto the process seam (#3152)
* test(#3144): bound the git/worktree cluster onto the process seam

Migrates 180 unbounded sync spawn sites across 19 files. Every previously
unbounded call now carries an explicit timeout with a comment giving the
number and why.

The migration is not a callee swap. execSync and execFileSync throw on a
non-zero exit and the seam never does, so each site was classified first:
sites that rely on the throw route to gitOrThrow, and sites that already read
.status to detect an EXPECTED non-zero -- an intended cherry-pick conflict, a
rev-parse outside a repo driving a skip -- route to the never-throwing runGit
instead, which would otherwise throw on exactly the exit being probed for.

Two same-named git() helpers in worktree-cleanup.test.cjs have different
return contracts, one trimmed and one raw; both are preserved rather than
unified.

Collapses five hand-rolled throw wrappers onto one throwIfFailed in
git-fixture.cjs, which gitOrThrow now also uses so the shape cannot drift.

Allowlist drops 139 to 120; BASELINE lowered to match.

* test(#3144): fix pre-PR review findings

Documents throwIfFailed in the CONTEXT.md glossary and CONTRIBUTING.md --
it became the shared throw mechanism without either doc naming it.

Routes the sixth and seventh hand-rolled copies of the throw shape through
throwIfFailed (worktree-baseref-install, worktree-safety-reap); the first
consolidation missed both.

Converts ci-rebase-check's 8 fixture-setup calls from unchecked runGit to
gitOrThrow so a failed setup step aborts where it fails rather than
surfacing later as a confusing failure against the wrong subject.

Adds 12 direct unit tests for throwIfFailed, which until now was only
exercised transitively.

Splits verify.test.cjs's non-git grep/sed bound off GIT_TIMEOUT_MS.

---------

Co-authored-by: sim <sim@local>
2026-08-07 10:58:34 -04:00
Tom Boucher
cfb37ccdfc Merge pull request #3154 from open-gsd/feat/3149-cmdinitdebug
enhance(#3149): dedicated init.debug entry point for /gsd:debug
2026-08-07 10:52:14 -04:00
Tom Boucher
28d6d20562 chore(#3149): backfill changeset PR number 2026-08-07 10:42:36 -04:00
sim
796bd2cb24 docs(#3149): use the hyphen slash form in reference docs
docs/ is never passed through the install-time slash-form converters, so
the colon form names a command no runtime registers. Caught by
lint-docs-command-form.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:10:23 -04:00
sim
8a0c1bce2e test(#3149): correct stale tdd_mode assertion and drop a marker-token collision
Two failures from the remote runner on 654b2cc10, both introduced here.

1. tests/mcp-catalog-parity.install.test.cjs greps emitted workflow files
   for the bare substring 'gsd:section' and treats its presence in a
   composed file as an un-stripped marker. debug.md's new Step 0 prose
   documented the field by writing that token literally, so the emitted
   file tripped the gate even though the parser correctly ignored it as
   prose. Reworded to 'applicability-section markers'. Same class as
   DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

2. tests/debug-session-management.test.cjs asserted debug.md contains the
   literal 'config-get workflow.tdd_mode'. That call is gone. The
   invariant it protected -- tdd_mode comes from the workflow.tdd_mode
   key, never a bare top-level one -- is unchanged, and now has a
   stronger behavioral home: init-debug.test.cjs row A9 asserts a bare
   key is ignored and the canonical key is honored, whatever the read
   mechanism.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:09:18 -04:00
Tom Boucher
2afe17bbdb test(#3143): add the no-unbounded-spawn guard and throw-preserving git fixture (#3150)
* test(#3143): add no-unbounded-spawn guard and throw-preserving git fixture

Adds the ESLint rule local/no-unbounded-spawn, wired into the tests/**/*.cjs
block, plus an allowlist that only ratchets down: a listed file with zero
violations reports its own entry as stale.

The rule resolves renamed destructures and chained requires rather than
matching literal callee names -- both forms exist in the suite today and a
name-only matcher leaves them permanently invisible. It resolves an options
object held in a single-write const, which is what keeps process-seam.cjs,
the bounded reference implementation, from flagging itself.

timeout: 0 and anything above the 600000ms ceiling are rejected as only
nominally bounded.

Adds tests/helpers/git-fixture.cjs so a migrated execSync call site keeps
its throw-on-non-zero contract; process-seam.cjs is unchanged.

* test(#3143): prove the allowlist guards can actually fail

Extracts the D4/D6/D7/D8 checks into pure helpers and drives each against a
synthetic fixture carrying an injected violation. Without this the suite only
proved that today's clean data passes, which a deleted check would also
satisfy.

* fix(#3143): close two ceiling and alias escapes found in review

Nested arithmetic bypassed the ceiling entirely: the numeric evaluator only
resolved a flat literal, so `timeout: 60 * 60 * 1000` (3600000ms, six times
the ceiling) fell through to trusted and reported nothing. The evaluator now
recurses through arithmetic and unary signs with a depth cap.

Alias resolution was traversal-order dependent, not scope dependent: a call
textually above its own require destructure saw an empty alias map and
reported clean. The map is now built in a Program pre-pass.

Also: an explicit timeoutMs:undefined no longer overwrites the git fixture
default via spread, adds the missing seam-routed rule test, and de-duplicates
the repeated try/catch in the fixture tests.

---------

Co-authored-by: sim <sim@local>
2026-08-07 09:43:36 -04:00
sim
654b2cc10c refactor(#3149): bind planningPaths once in cmdStateLoad
The debug_dir change introduced a second planningPaths(cwd) call in the
same function. Bind the struct once and read both .planning and .debug
from it; emitted output is byte-identical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:33:47 -04:00
sim
2bead6ca1d feat(#3149): add dedicated init.debug entry point for /gsd:debug
/gsd:debug was one of the last workflows with no cmdInit* of its own: its
Step 0 made three separate round-trips (state.load, resolve-model
gsd-debugger, config-get workflow.tdd_mode) to assemble one context. Because
no debug-scoped fact was computed at any entry point, ADR-1671 admission gate
(2) could never be satisfied for debug — an applicability atom naming such a
fact would evaluate FALSE forever and silently exclude its section.

Adds cmdInitDebug (init.debug), registers it in the init router and the
command-alias table, and collapses debug.md Step 0 to one call. Every field
resolves through the same primitive the call it replaces used: loadConfig for
commit_docs, withProjectRoot for response_language (#2402), planningPaths for
debug_dir, resolveModelInternal for debugger_model, and the existing
Boolean(workflow.tdd_mode) idiom for tdd_mode.

PlanningPaths gains a debug field so state.load and init.debug share ONE
debug-directory expression rather than two kept in sync by hand. state.load
keeps emitting debug_dir: it is a shipped query surface with its own test
anchor, so narrowing it would break unseen consumers for no gain.

No WHEN_VOCABULARY atom and no gsd:section marker: gate (1), a consuming
section of at least 400 bytes, belongs to the change that adds the section.

Closes #3149

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:30:27 -04:00
sim
0d93a5cfe5 test(#3149): failing-first coverage for the init.debug entry point
Adds the behavioral matrix for a dedicated init.debug handler before the
handler exists: equivalence cross-checks against the three calls debug.md
makes today (state.load, resolve-model, config-get workflow.tdd_mode),
bundle shape, the CLI negative/hostile argv matrix, section_manifest
null-vs-[] degradation, planningPaths.debug, and a guard that
WHEN_VOCABULARY stays closed at 29 entries.

Currently RED: the init router reports "Unknown init workflow: debug".

Matrix: .gsd/phase/feat-3149-cmdinitdebug/50-test-matrix.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:18:00 -04:00
Tom Boucher
4b66bf4560 fix(#3086): apply #2667 .cmd-shim gate to deps.spawn + surface errorCode in review lanes (#3142)
* fix(#3086): apply #2667 .cmd-shim gate to deps.spawn + surface errorCode in review lanes

deps.spawn used shell:false with a bare binary name — on Windows, npm-installed
CLIs (gemini, codex, etc.) are .cmd shims that CreateProcess cannot start,
producing ENOENT + empty stderr. The review path then wrote an empty err file
and emitted a generic 'failed or returned empty output' stub.

Two fixes:
1. deps.spawn: detect .cmd/.bat on win32 and mediate through cmd.exe /d /s /c
   (same gate as runWithTimeout #2667, same explicit argv array).
2. runSpawnLane: surface errorCode (ENOENT, ETIMEDOUT) in the err file so the
   stub explains WHY the lane produced nothing.

* chore(#3086): backfill changeset PR number 3142

---------

Co-authored-by: sim <sim@local>
2026-08-07 07:43:14 -04:00
Tom Boucher
7ab4556395 fix(#3079): query commit no longer resurrects deleted phase branches via silent switch (#3141)
* fix(#3079: query commit no longer resurrects deleted phase branches via silent switch

git checkout -b both created AND switched HEAD, so when branching_strategy:
'phase' deletes its branch on merge, a later query commit silently
recreated it and moved HEAD there. The commit landed on the wrong branch.

Fix: replace checkout -b with git rev-parse --verify + git branch (create-
only, no switch). The commit always lands on the current branch. Callers
that want to be on the phase branch use execute-phase's handle_branching.

Updated 3 existing tests that asserted the old switch behavior.

* chore(#3079): backfill changeset PR number 3141

---------

Co-authored-by: sim <sim@local>
2026-08-07 07:02:43 -04:00