Commit Graph

268 Commits

Author SHA1 Message Date
Tom Boucher
1e67ec9737 enhance(#3908): the scanners distinguish an empty diff from one they could not compute (#3937)
* feat(#3908): the scanners distinguish an empty diff from one they could not compute

collect_files ended 2>/dev/null || true, which destroyed the evidence three ways: the redirect discarded git's diagnostic, the pipe replaced git's status with grep's, and || true forced success regardless. Four distinct conditions - an established-empty diff, a bad ref, no repository, and a repository with no commits - all reported clean, and a secret scanner reporting clean because git failed is indistinguishable from an all-clear to any gate consuming it.

git now runs separately from the filter so its status and diagnostic both survive. An established-empty diff exits NO_INPUT; a scope that could not be established exits UNAVAILABLE; the usage sites move off 2 to USAGE. || true is retained on the filter alone, where it is correct: a diff of only images is empty, not failed.

Codes are sourced from a generated shell fragment rather than written into three scripts, so a re-allocation cannot desync them, and a missing fragment fails loudly instead of falling back to literals. The security workflow is updated in the same change: without it, a docs-only PR would newly fail the job.

* fix(#3908): keep scanner stderr out of the file list, and drop try/finally from test bodies

Capturing git and find output with 2>&1 was right for the failure path but wrong for the success path: a warning emitted alongside a successful diff flowed into the file list and was treated as a filename. stderr is now captured separately, forwarded as a warning on success and as the diagnostic on failure, and never folded into the list.

Also converts the control tests' try/finally blocks to t.after(), which CONTRIBUTING bans inside a test body because it masks failures.

* chore(#3908): backfill changeset pr number

* docs(#3908): record the scanners' four-outcome exit contract

SECURITY.md is root-level, so the docs gate correctly held: a Changed fragment owes a file under docs/. The contract also belongs where the feature is described, as REQ-SCAN-INJ-05.

docs/FEATURES.md is GENERATED from per-feature fragments (#3840) - the first edit went into the generated file and gen-features --check caught it, which is the same edit-the-output drift this epic exists to close. The fragment is the source; FEATURES.md is regenerated.

---------

Co-authored-by: sim <sim@local>
2026-08-27 13:11:13 -04:00
Tom Boucher
c5f2b94b27 enhance(#3907): gates report no-input instead of a verdict they never reached (#3932)
* feat(#3907): gates report no-input instead of asserting a verdict they never reached

The three stdin-reading gates bound 2 to a stdin read error only, with no arm for stdin closed at zero bytes - so empty input flowed into the detector, found nothing, and exited 1, which each module's own comment defines as a negative verdict. An unset PHASE_SECTION made the UI gate assert the phase has no UI. Empty and whitespace-only input now exit NO_INPUT, and a read error exits UNAVAILABLE rather than a locally-invented 2, both resolved through the registry and delivered by terminateNow.

The exit code was only half of it: under --json the same input emitted {detected:false}, byte-identical to the fabricated payload #3909 exists to fix, and the blocking coverage gate reads that payload. Empty input now emits the in-tree {skipped:true,reason} form with no detected key at all.

teams-status is excluded: it never reads stdin and has no invented 2, so the four-module framing in the issue and ADR is wrong. The dead root bin/lib/ui-safety-gate.cjs is deleted - no installer reference, no workflow invocation, and the live fallback chains are for other modules. Its removal restores the unit tests to the module that actually ships; they had been asserting the stale copy's two-field shape, which is why it drifted unnoticed.

* fix(#3907): drive gate tests through the process seam, and make removed-but-needed basename-precise

CONTRIBUTING requires every subprocess go through tests/helpers/process-seam.cjs; two of the three gate suites hand-rolled spawnSync while the third, added in the same change, used runNode correctly for the identical injection case. Converted the blocks this change added, leaving pre-existing ones alone.

Deleting one of two files sharing a basename made lint-removed-but-needed report 14 references that were all to the surviving canonical module - the false-positive class its own docstring names. It now matches on the deleted file's full path when a surviving file shares its basename, which is more precise rather than weaker: a genuine full-path reference still fails, and behaviour is unchanged when no basename collides. It immediately caught a docstring on this branch that spelled the deleted path.

* test(#3907): update the one existing assertion that pinned the old empty-stdin verdict

A pre-existing test asserted exit 1 on empty stdin - the defect this phase removes - and was missed because the change added new blocks without auditing existing ones pinning the old contract. Audited the rest: the other three status-1 assertions in that file all feed real input and are the genuine-negative controls that must keep returning 1, so exactly one was stale. The retired 2 is gone from the describe's contract comment too.

* chore(#3907): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-27 11:37:31 -04:00
Tom Boucher
941b62249e enhance(#3906): two terminators over one registry, with a versioned exit projection (#3924)
* feat(#3906): two terminators over one registry, with a versioned projection

Adds terminateNow (write-then-terminate, for callers that cannot wait for the event loop) beside runMain (drain-then-exit), both projecting through one shared function so they cannot disagree - the parity the ADR makes mandatory. A failed write does not change the exit code: letting it propagate would fail a hook open, which is what the fail-closed branches exist to prevent.

The projection is versioned. v1 reproduces today's integers, including keeping a payload-carried degraded result at exit 0 - ADR-2980 ratified that across 60 sites and declined normalizing it on measured blast radius. v2 applies the registry. --exit-contract=v2 or GSD_EXIT_CONTRACT=v2 selects it; an unrecognized version throws rather than silently defaulting.

The registry is now emitted beside both copies of the exit module, so it resolves as a sibling in the built tree and in the committed scripts/ copy that must load on an unbuilt clone.

* fix(#3906): actually restrict code 2 to terminateNow, and generate the registry's type

The claim that terminateNow is the only place 2 can be produced was false: runMain's outcome arm applied no guard, so runMain(()=>'HOOK_DENY') set exitCode 2 through the drain path - and the parity matrix demonstrated it while calling it parity. runMain now refuses any outcome projecting to the hook-protocol code, gated on the code rather than the name so an alias cannot slip past, and the matrix asserts the restriction instead of contradicting it.

The ambient type for the generated registry was hand-written with no gate against the generator's actual output - the declared-surface-diverges-from-runtime defect class this epic exists to close, reintroduced inside it. It is now a third generated artifact covered by the same --check. Also converts every test-body try/finally to t.after().

* test(#3906): derive the glossary fixture's dependencies instead of hand-listing them

Adding a require to scripts/lib/cli-exit.cjs broke 31 tests in one suite that built its fixture from a hand-written dependency list, so the new sibling was absent and the copied script could not load. copyScriptWithDeps walks the require graph and exists for exactly this class - #3412 paid the same bill when one new require broke 82 tests across two suites. Migrating rather than adding another copyFileSync line keeps the class closed. The other nine suites referencing that path were triaged; none copies-and-spawns, so none needed migrating.

* fix(#3906): enumerate the new shipped file, drop a vendor name from shipped data, and fix three test defects

install: scripts/lib/exit-code-registry.cjs was missing from GSD_SCRIPTS_LIB_FILES, so it shipped to every install and orphaned on uninstall.

The registry gave HOOK_DENY a meaning naming one harness, and that string ships into every runtime's tree - a guard correctly caught it leaking into the hermes and qwen installs. The registry is runtime-neutral infrastructure; the vendor name belongs in the ADR, not in shipped data.

Two more fixture harnesses built their trees from hand-listed dependencies and broke on the new require; both migrated to the derived helper, and all 23 copy-and-spawn candidates were enumerated so the class is closed rather than patched. One generator test used a fixture code that collided with a real allocation, so the generator correctly reported a duplicate where the test expected drift. The large-payload test embedded a 256KB literal in the child's argv, exceeding Linux's 128KiB MAX_ARG_STRLEN so the child never started - it now builds the payload inside the child.

* chore(#3906): backfill changeset pr number

* docs(#3906): document the exit-code contract selector

P2 is the first phase of this epic with a user-invocable surface, so the flag and env var owe a reference entry. Records what actually differs between v1 and v2 today (one outcome), that an unrecognized value is rejected rather than silently defaulted, and the fail-safe property that makes switching safe.

---------

Co-authored-by: sim <sim@local>
2026-08-27 03:31:02 -04:00
Tom Boucher
39673ae9ff fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config

Regression tests (RED first): --skills-root and gsd-tools query surfaces,
install-plan dest dirs, and converter skills-path rewrite.

* fix(#3738): antigravity global skills/agents install to ~/.gemini/config

Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents};
the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the
ADR-1239 skills/agents 'home' override on the antigravity global layout — the
same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references
in converted global content to ~/.gemini/config/skills/. configHome, settings,
probe/migration semantics, and the local .agents layout are unchanged.

* fix(#3738): retire deprecated configHome artifacts via installer migration 010

Next install converges an existing antigravity install: manifest-managed
skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does
not scan) are removed — modified files backed up first, unmanifested and
non-gsd entries preserved — and now-empty containers retired. Global scope
only; the local .agents surface is live. Docs + inventory updated.

* fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline

- bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/
  rewrite as src (ADR-1508 dual copy must stay in sync).
- Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so
  antigravity's emitted skills/agents stay differential-visible at their new
  install root; install-tree fixture regen confirms an unchanged key set.
- skills-from-commands rule declares the antigravity converter as a
  runtime-scoped transform; one ack fragment covers the identity-classed
  workflow whose antigravity copy embeds the old skills path.
- Migration 010 checksum baseline + home-override set doc updated; existing
  tests updated to the #3738 contract (global dest, golden parity via layout
  dest, integration expectations).

* fix(#3738): tolerate an absent extra emit root on baseline-side measurement

The base tree's installer predates the home override, so <HOME>/.gemini/config
does not exist there; walk() threw ENOENT and the in-job baseline build failed.
An absent extra root is the legitimate pre-override shape — skip it.

* fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment

- writeManifest resolves the agents-kind home override (_kindDestDirSafe), so
  the manifest records agents at their actual install root and drift detection
  keeps working (isolated review finding 1, major).
- Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing
  slash) divert to ~/.gemini/config/skills instead of falling through to the
  retired configHome path (finding 2).
- real-home-guard comment updated: antigravity's global agents kind is the
  first agents-kind home override (finding 3, doc-only).
- Regression tests for both behavioral findings.

* chore(#3738): changeset fragment (pr number backfilled after PR creation)

* chore(#3738): backfill changeset PR number (3921)

* fix(#3738): sandbox HOME in tests that install antigravity global artifacts

antigravity is the first home-override runtime in the golden-parity and
skills-wrapper suites (codex is not in their runtime lists), so those tests
never needed HOME sandboxing — the real-home guard now (correctly) refuses
their un-sandboxed global installs on CI, where HOME is the passwd home.

* fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans

K3's two back-to-back sandboxHome calls leave HOME pointing at the first
sandbox once the after-hooks restore (each call saves the env as it found
it, so the second saves the first's sandbox as 'original'). On the windows
matrix that leaked gsd-k3-qwen-* home into the L2 property, whose
antigravity/global run then (correctly) refused via the #3712 real-home
guard — antigravity is the runtime that made L2's plan escape into
os.homedir(). K3 now manages the env with a single restore; L2 sandboxes
HOME per run, mirroring L1.

* fix(#3738): L2 property's HOME sandbox must exist on disk

The #3712 guard's sandbox exemption fails closed when identify(effectiveHome)
is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir
under the real home) the antigravity/global run refused even with HOME
sandboxed. Create the per-run sandbox dir and clean it up.

---------

Co-authored-by: sim <sim@local>
2026-08-27 02:24:03 -04:00
Tom Boucher
e20744eacb enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick

ADR-3473 §8.4 says failure is a value. Three families currently encode failure as
success, and this commit pins each one RED before the fix lands.

Measured on this tree, 2026-08-26:

  gsd-tools generate-slug "test" --pick nonexistent
    -> empty stdout, exit 0                                     (#3365)

  gsd-tools audit-open --pick nonexistent_field
    -> dumps the entire human-readable audit report, exit 0

  gsd-tools generate-slug "Hello World" --raw --pick bogus
    -> prints "hello-world", another field's value, exit 0

  gsd-tools query state.planned-phase 3        (positional, no --phase)
    -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to
       "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted
       current_phase_name                                        (#3358)

tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract
("returns empty string for missing field", success === true). That assertion is
replaced by the required behavior rather than deleted.

The new parseNamedArgs block calls the spec-object signature that does not exist
yet, so it fails today by construction. The 11 existing behavior-lock tests are
left untouched here; they are corrected in the implementation commit.

C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180
Decision 4(b). A unit assertion on the parser would have passed throughout this
defect's life.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3884): failure is a value — strict argv, and --pick that signals absence

Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable
ways to say "I could not answer".

parseNamedArgs (src/command-arg-projection.cts)
  Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the
  hub's Result shape instead of a bare Record. Declaring the positional arity is what
  makes #3358's call site unrepresentable rather than merely detectable: an unrecognized
  flag or a token past the declared boundary is now InvalidArgs, naming the offending
  token and listing the accepted flags. The legacy positional-array call shape throws
  a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale
  hand-written .cjs call site fails loudly instead of destructuring undefined off a
  Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a
  projection over the one parser, not a second parser.

  Measured before, against a STATE.md with a populated phase-2 block:
    query state.planned-phase 3        (positional, no --phase)
    -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to
       "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted
       current_phase_name
  After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical.
  The flag form is unchanged and still updates STATE.md.

--pick <field> (gsd-core/bin/gsd-tools.cjs)
  extractField returns {found,value}, and the pick block no longer shares one catch
  between "output was not JSON" and "field was absent". An absent field exits 1 with
  pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1
  with pick_output_not_json instead of dumping the command's entire output. A field that
  is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer,
  not a failure, and it is what keeps `--pick count` printing 0 on a fresh project.

  Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable
  audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" —
  a different field's value, confidently, at exit 0.

  ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The
  sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the
  ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0
  would demote "could not answer" to "the answer is zero" — the hazard
  docs/how-to/resolve-unreachable-guard-findings.md already warns against.

Guard ledger (ADR-3473 Decision 6)
  scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a
  `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the
  correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls,
  a nullglob mechanism this change does not touch) is retained in full, as are the shared
  scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file
  is not deleted.

Call-site audit
  45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an
  if test, && chain, or a pipeline whose status is consumed, and no shell block in
  workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the
  prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind
  a prior found/existence check. No ADR-3409-class "field the command never produces"
  remains.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows

Two review findings, both fixed here rather than recorded as limits.

1. A newline in an untrusted token forged a second stderr line.

   Before, plain-text mode:
     $ gsd-tools query state.planned-phase $'foo\nError: forged second line'
     Error: unexpected positional argument "foo
     Error: forged second line"

   After:
     Error: unexpected positional argument "foo\nError: forged second line"

   --json-errors mode was never affected — io.error runs that payload through
   JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the
   three new InvalidArgs reasons plus the two new --pick diagnostics all
   interpolate a token that comes straight from argv.

   Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at
   every interpolation site — not a copy per site. It is deliberately NOT
   applied inside error() itself: several callers in this tree emit intentional
   multi-line diagnostics, and escaping newlines there would mangle them.

   The available-top-level-keys list needed the same treatment for a reason the
   review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user
   document and echoes that document's own keys into the diagnostic. Verified
   reachable — a frontmatter key containing a newline reaches the key list — so
   formatKeyForDiagnosticList is guarding a live path, not a hypothetical one.
   Ordinary keys still render plain and unquoted; a fix that merely dropped the
   key would also have passed a "one line" assertion, so the test pins the
   escaped key's presence too.

2. Five behavior-table rows were implemented but nothing pinned them:
   B7  a dotted path that dies partway
   B9  bracket syntax on a non-array
   B10 a negative array index, in and out of range
   B14 a JSON root that is not an object
   B17 an @file: payload over 50KB

   B17 is the load-bearing one. output() writes @file:<path> instead of inline
   JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no
   test, a future reordering of those two steps turns every large result into a
   false pick_output_not_json. The fixture seeds 1200 phase directories and
   measures the payload at 62474 characters, asserting the spill actually
   happened rather than assuming it.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): correct the strict-argv surface against a full verification run

The first full run came back with 90 failures across 12 files, none in the new
tests. They were the argv surface telling me what it actually is. Ten root
causes; each classified before anything was changed.

I over-implemented, and that is reverted.

  ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional
  tokens". It says nothing about a value flag whose value is missing. Making
  that an error was my design decision, not the rule, and it broke a
  deliberately recorded contract: `--prd` with no value resolving to null
  (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5;
  tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The
  "requires a value" branch is deleted outright rather than kept behind an
  option — an unused strictness mode is speculative generality. Unknown-flag
  and unexpected-positional rejection, which is what §8.4 actually mandates,
  is unchanged.

--wave needed a third flag kind the original design did not anticipate.

  `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the
  shipped workflow reconstructs and passes it (execute-phase.md:84), while
  #2932 records token-PRESENCE semantics: the CLI cares only that the flag
  appeared, and the value belongs to the workflow layer. That is neither a
  boolean flag nor a value flag, so `optionalValueFlags` now exists —
  presence-only in `data`, and the validation cursor consumes a following
  non-flag token so it is not reported as a stray positional. Every other
  declared boolean flag was checked against every argument-hint and prose
  usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one
  of this shape.

Five tests were pinning forms that never worked.

  tests/adr857-core-without-capabilities.test.cjs passed
  `init plan-phase --phase 01-stub`, but the documented form is positional
  (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form
  is the literal string "--phase". Measured on the pre-fix build against a
  real .planning/phases/01-stub/ directory:

    init plan-phase 01-stub          -> phase_found=true
    init plan-phase --phase 01-stub  -> phase_found=false

  The test asserted only exit 0 and key presence, so it had been green while
  proving nothing about phase resolution. Corrected to the documented form and
  strengthened to assert phase_found === true. Same class in state.test.cjs
  (`--plan-count`, a flag that does not exist; the real one is `--plans`),
  milestone-archive.test.cjs (`init new-milestone --json`, silently ignored),
  and concurrency-safety.test.cjs (a bare positional field name whose
  OR-assertion passed because a whole-document dump happens to contain the
  substring it looked for).

Six handlers had no argv validation at all — the same #3358 shape this phase
exists to close, found while fixing the rest: init verify-work / phase-op /
review / todos / remove-workspace read args[2] with nothing checking the rest,
and validate health read --repair/--backfill through a bare args.includes()
scan that bypassed the parser entirely. All now go through the seam, so the
flag has one owner.

tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must
NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's
local expectation does not override §8, so they are inverted and renamed —
a test still called "ignores an unrecognized flag" while asserting rejection
would be its own defect. Row C6's point is its PWNED canary; that assertion is
kept verbatim and only its exit-status expectation changed, because the
hostile token is now rejected rather than absorbed.

The blast-radius estimate in 40-design.md is corrected rather than quietly
left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was
accurate for what the graph can see — parseNamedArgs's callers. It cannot see
that those callers' handlers accept argv shapes wider than the code reading
args[2] suggests, which is where the real surface was.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert

Second full run: 46 failures, down from 90. Four causes, two of them mine.

Reverted `validate health` entirely — it was scope creep, and it broke a real flag.

  ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`.
  The previous commit routed `validate health` through the parser on the
  reasoning that a flag should have one owner. That was wrong twice over:
  §8.4 names parseNamedArgs and count queries, and `validate health` was never
  a parseNamedArgs call site — it read its flags, just not through the parser,
  so it had no silent-drop defect to fix. Tightening it omitted `--json`, which
  the health-diagnostic suites use heavily. The handler is now byte-for-behaviour
  back to its pre-branch form. `validate context` stays converted: it genuinely
  was a call site, and its `--json` is now declared rather than read by a second
  `args.includes` scan.

  The five handlers that had NO validation at all — init verify-work / phase-op /
  review / todos / remove-workspace — stay fixed. Those read args[2] with nothing
  checking the rest, which is the #3358 shape this phase owns.

Finished the A2/A3 revert. Three tests still encoded the deleted
"a value flag with a missing value is an error" rule, including one added by the
previous commit for that rule. All three now assert the reverted null contract,
and the ones whose titles said "rejected" are renamed — a test named for a
contract it no longer asserts is its own defect.

`--wave=` and `--wave --weird` are correctly rejected. Neither is documented in
commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and
neither is emitted by the shipped prompt layer, so both are unrecognized tokens
that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the
property it exists for — asserted directly now, at the parser, that `--wave` does
not swallow a following flag as its value — and only its exit-status expectation
changed.

A contradiction inside this branch, surfaced by the audit and resolved the safe way.

  Two pre-existing #3573 tests call `state begin-phase '2'` and
  `state planned-phase '2'` with a bare positional, relying on the old permissive
  parser to ignore it. This branch's own #3358 regression test requires that exact
  argv to be REJECTED. The two are mutually exclusive.

  Widening the router to accept a bare positional — mirroring complete-phase —
  would have silently re-opened #3358, and was verified to do exactly that: with
  the widened router, `query state.planned-phase 3` returned exit 0 and wrote
  current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and
  docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the
  two #3573 tests move to it. Their assertions were never about the call shape —
  only that total_phases survives the resync — and both still pass.

  complete-phase is untouched: its bare positional IS documented, and it keeps the
  dynamic boundary and the negative-space note that record why.

The audit that produced this is in the PR body: for every handler whose declaration
changed, the flags it reads anywhere in its body, the flags the shipped surface
documents, and the shapes the suite passes, compared. The `--json` miss was a
pattern, not an accident — declaring a handler's flags from its parseNamedArgs call
alone misses whatever it reads elsewhere.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3884): backfill the changeset PR number

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 00:12:13 -04:00
Tom Boucher
878f25025c enhance(#3905): the exit-code registry — one number, one meaning, enforced at build (#3920)
* feat(#3905): the exit-code registry — one number, one meaning, enforced at build

A generated registry replaces locally-invented exit codes. Every entry records code, name, meaning, owning module and the decision that authorized it. The generator refuses to build a table where two entries claim one code, two claim one name, a code falls in a range Node or the shell reserves, 2 is claimed by anything but the hook adapter, or an allocation carries no justification. exitCodeFor is pure and total: it throws rather than returning undefined, including for prototype-chain names.

Inert by design — nothing emits a registered code until #3906. Every registered code is non-zero, asserted over the whole table, so a caller testing for failure behaves identically for pass and trips for everything else.

* feat(#3905): make the registry generator's failures machine-readable

Adds a --json mode carrying {ok, reason, context, detail}, where context is a typed payload naming the specifics the prose embedded - which code collided and under which names, which band rejected a code, which field was missing. The tests now assert on that structure instead of regex-matching the generator's stderr, which CONTRIBUTING prohibits, and the CONTEXT.md glossary gains the entry the issue's scope requires.

* test(#3905): refresh the install-tree fixtures for the new declaration

The registry declaration ships in the install tree, so all 19 golden fixtures needed regenerating. Caught by the remote matrix, not by lint:ci - the install-tree goldens are verified by a test rather than a lint, so a newly shipped file clears every local gate and fails only under the suite.

* chore(#3905): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-27 00:10:11 -04:00
Tom Boucher
6b7df61938 enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises

ADR-3473 §8.1 carries a blocking open question with a forcing function: it must
be answered before any implementation PR for the rule opens. Answered here as (a),
a string-coercing adapter, with the measurement that settles it.

The sequencing note bet that §8.8's schema would make (b) tractable. Measured
against merged reality it does not: only 33 of extractFrontmatter's 78 non-test
call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists
survive real types, leaving ~31 lines across 3 call sites as the actual prize.

Also corrects three claims verified false while answering it. §8.1's justifying
sentence names #3349 and #3360 as defects a real parser would fix; both are
already fixed on next, confirmed by executing the compiled parser rather than
reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an
expected casualty of this rule, but it guards shell grep idioms in workflow bash
fences and never touches our parser. The same roster calls lint-vendored-deps.cjs
reusable as-is; it is hardcoded to re2js throughout.

The last two were caught by applying the rule this amendment records -- a factual
claim in this ADR is a hypothesis until the implementing phase executes it -- on
its first use.

Refs #3881

* docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable

An adversarial pass on the Phase 4 design established by execution that
extractFrontmatter is not a YAML parser but a line-oriented scanner whose output
is a function of raw source text. Four spellings of the same value collapse to
one js-yaml tree but produce four distinct legacy strings, one of them mangled.
No adapter over a tree can choose among outputs the tree does not distinguish,
so fork (a) -- keep a string-coercing adapter so the existing contract holds --
cannot be built. For any document with a non-scalar value, (a) collapses into
(b); about 26 percent of frontmatter-carrying documents have one.

Also records three design defects and one new attack surface, all confirmed by
execution: catching a parse failure and returning {} would delete the frontmatter
block on the next write at eight call sites that conflate empty with unparseable;
an empty value yields null where legacy yields {}, and reconstructFrontmatter
omits null-valued keys, so the shipped state template's empty progress key would
vanish; the #1882 truncation probe is parseYamlRegion itself rather than a
pre-parse heuristic, so it cannot both stay unchanged and survive that deletion;
and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB.

The rule is not deferred. The measurement is the deliverable and the re-scoping
is recorded as an open question with a forcing function, per section 8's own rule.

Refs #3881

* test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix

Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed.

Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states.

B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today.

B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today.

B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today.

Refs #3881

* chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest

Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496).

gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own.

src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare.

scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored.

docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds.

Refs #3881

* feat(#3881): parse .planning frontmatter with the vendored js-yaml

ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled
line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and
parseQuotedScalar are deleted (not patched); parsing now goes through the
vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA,
json: true }. Everything js-yaml does not do is layered on top, in one
place, carrying the seven design-doc consequences:

1. Empty value: a null js-yaml value is coerced to {} (matching legacy's
   own empty-value contract) so reconstructFrontmatter — which omits
   null-valued keys — still round-trips a bare `key:` line instead of
   deleting it. Verified live: progress: with no value survives
   parse -> reconstruct -> re-parse.

2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE
   Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS
   channel, is carried on the {} returned for malformed/refused YAML. Invisible
   to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that
   never inspect it are unaffected; wiring the 8 hasFrontmatter sites to
   consult it is a separate change, not done here.

3. Non-scalar object-list items (the four spellings of `- test: a b` that
   js-yaml collapses into one tree shape) are rendered as a canonical
   `key: value[, key2: value2]` string per item, keeping the existing
   array-of-strings value SHAPE. A full corpus differential over all 1702
   tracked markdown files found 11 residual divergences from the legacy
   parser (enumerated in the PR/report), most of them the parser now being
   MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key
   defect, a dropped Unicode key).

4. The #1882 truncation probe still runs the one real parser, but derives
   its key count from js-yaml's own thrown error and mark.line when the
   whole region doesn't parse cleanly (the dominant real truncation shape:
   fence opened, well-formed keys, no closing fence). Verified against both
   the clean-parse and the exception-fallback path.

5. The #3257 comment channel now attributes each pending column-0 comment
   against js-yaml's own parsed top-level key list (matched by literal key
   text, in document order) instead of the legacy ASCII-only key regex, so
   a comment above a Unicode key attaches correctly.

6. Anchors, aliases and merge keys are refused outright (a raw-text
   pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences
   today: zero. A 7-line billion-laughs fixture is verified refused rather
   than expanded.

7. A literal U+0000 is swapped for a private-use sentinel before the parse
   and restored in every resulting string afterward, since js-yaml rejects
   NUL unconditionally under every schema.

escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump()
(forced double-quoted style), with control-char hex escapes lowercased to
keep serialized output byte-stable (#1779 emitted lowercase); it keeps its
exported name and signature for its two other call sites (commands.cts,
runtime-artifact-conversion.cts), which need no change.

frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments,
regenerateFrontmatterKey's guard, noOpObjectListSetError and
parseMustHavesBlock are all unchanged — retiring them is fork (b) and is
not this phase.

Refs #3881

* fix(#3881): quote template placeholders and preserve unparseable frontmatter

SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as
bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax
under the vendored js-yaml parser, not the literal placeholder text
intended. Quote them so they parse as strings.

Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the
8 call sites in state.cts/state-transition.cts that compute
hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and
reassemble the document without a frontmatter block when false. That
check conflated 'no frontmatter' with 'unparseable frontmatter' (both
parse to {}), so a document with a merge-conflict marker or refused
alias in its frontmatter had that block silently dropped on write.
Each site now preserves the exact raw bytes stripFrontmatter removed
when the marker is set, leaving the genuinely-empty case unchanged.

Refs #3881

* test(#3881): consequence and boundary coverage for the js-yaml migration

Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix.

Refs #3881

* docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to

Refs #3881

* docs(#3881): correct the frontmatter glossary entry

Two errors in the entry as first written: it named parseYamlRegion as part of
the read path when that function is deleted, and it recorded the eight
hasFrontmatter call sites as unwired follow-on work when they were wired in
e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the
frontmatter block independently, so the marker binds at the transform layer.

Refs #3881

* docs(#3881): record the semantic-migration decision and the counted guard ledger

The maintainer chose the full semantic migration over splitting the rule into
its own epic or patching the scanner, so section 8.1 is answered as "the fork
was ill-posed and the migration is semantic" rather than as (a) or (b).

Also replaces the pre-implementation guess that this phase would shrink the
guard surface with the counted result: excluding vendored third-party lines the
hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines
despite four functions being deleted, because the compatibility layer over
js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is
therefore not delivered as written; what improved is the kind of code
maintained, not the amount. Decision 6 requires recording that rather than
netting it away.

Refs #3881

* chore(#3881): changeset for the vendored YAML parser migration

Refs #3881

* test(#3881): golden parity, round-trip property and packaging coverage

Refs #3881

* fix(#3881): refuse anchors structurally and fold in review findings

ADR-3473 §8.1 review findings, addressed inline:

Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched
only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping
({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor
mechanics while never matching that line shape, so the exact expansion the
guard exists to stop went straight through unrefused (a 303-byte quoted-key
bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener`
callback, which reports `state.anchor` for every event belonging to an
anchored node in every spelling, and throws from inside the callback to abort
before any expansion (~1-2ms vs full expand-then-discard). A merge key with
an alias is still refused (merge always requires a previously anchored node,
so the alias itself trips the listener); a bare merge key with NO alias is no
longer separately refused, documented as intentional: FAILSAFE_SCHEMA never
resolves `!!merge`, so it carries no expansion risk. Table-driven tests added
for all four bypass spellings + merge key, plus a quoted-key-spelled
billion-laughs fixture registered in the adversarial matrix and README.

Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/
aliases were "simply UNREACHABLE from typed code" through the twin. Corrected
to state the truth: anchor/alias resolution is document-level `load`
mechanics reachable through exactly the declared surface, and refusal is
enforced at RUNTIME (Finding 1's listener), not by the type surface.

Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was
non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree
back to NUL, including one the document author legitimately wrote, silently
corrupting it. Now refuses outright whenever the raw region already contains
U+E000 (consistent with the existing anchor/merge-key refusal path), making
the substitution provably injective. Tests added for a real NUL alone
(preserved), a pre-existing U+E000 alone (refused, not corrupted), and both
together (refused, not merged into one byte).

Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead
for a hand-authored row (only read inside the upstream-verbatim branch) —
exactly how Finding 2's stale docblock drifted unnoticed. Added
checkHandAuthoredTwin: every value-level export the twin DECLARES must be an
actual own property of the vendored runtime module at require-time. Tests
added, including a sensor that a declared-but-nonexistent export IS caught.

Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/
unparseableFm/reassemble preamble, copy-pasted at 7 sites in
state-transition.cts plus a sixth hand-inlined copy in state.cts's
cmdStateCompletePhase, is now one exported helper
(beginFrontmatterReassembly) every site routes through, including the
hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore)
keep a literal `body = stripFrontmatter(content)` assignment alongside the
helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward
scan (which does not chase aliases) still sees the strip; stripFrontmatter is
pure/idempotent so the extra call changes nothing observable.

Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a
separate change" claim (the 8 call sites are wired on this branch) and the
changeset's backlink from (#3473) to (#3881).

Finding 7: fixed the lint:ci failures blocking the gate — an
@typescript-eslint/only-throw-error violation from throwing a bare Symbol as
the anchor-detected signal (now a real Error subclass), unused-var warnings
left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by
two migration-specific test files (allowlisted with justification), and the
lint-state-write-path-drift false positive from Finding 5's helper (fixed
above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already
carried an explicit timeout; no change was needed there.

Golden fixture: added a golden entry for the new
anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line
scanner would also produce, since it independently dropped every quoted
top-level key). No other corpus document diverges: real .planning/ documents
carry zero anchors/aliases/merge keys/U+E000 today.

Refs #3881

* fix(#3881): fold in second-round review findings

Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration
ASCII-only key regex for the Unicode fixture; updated to require the 相
key's value now that js-yaml has no such restriction. Audited the rest of
the file for other pre-migration pins (block scalars, quoted keys,
flattened values, empty values, duplicate keys, unclosed blocks, null
bytes) by execution against real fixtures; found none regressed.

Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to
parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts
so no function still answers to the deleted hand-rolled scanner's name
(ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three
external call sites (src/commands.cts, src/runtime-artifact-conversion.cts)
updated in the same change — a mechanical rename, not an ADR-amendment
matter.

Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found
by execution. A top-level key named constructor/__proto__/toString/
valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read
resolving an inherited Object.prototype member); a key literally named
__proto__ was silently DROPPED entirely (bracket assignment on an
ordinary {} invoked the inherited __proto__ setter instead of creating a
data property). Fixed by building every parsed Frontmatter object with
Object.create(null), and replacing an `in` check with hasOwnProperty.call
in propagateCommentChannel. Added round-trip tests for all five hostile
keys, each with its own leading comment.

Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed
full byte-stability across the migration. Verified by execution: BEL/NUL/
NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw-
literal forms. Proved round-trip equivalence (each escape re-parses to the
exact source codepoint) and corrected the docstring. Found and fixed a
related real defect while verifying: a lone UTF-16 surrogate was emitted
BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely
unparseable YAML that silently collapsed to {} on re-read — extended
scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped
path.

Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real
truncation shapes (unquoted colon, open flow collection, mis-indented
sibling key, refused anchor). Root cause: the mark-based prefix recovery
excluded the very line whose key needed counting, and a mark-less refusal
never entered the recovery branch at all. Fixed by taking the max of two
lower bounds: the longest parser-verified line-prefix, and a raw-text
count of key-shaped lines (reusing the same key-shape pattern this file
already uses for isFrontmatterShaped). Extended test-matrix row A5
table-driven over all 4 regressed shapes.

Finding 6: the design doc's claim that no test owned the #3594 adversarial
fixture corpus was false — consolidation epic #1969 had already folded it
into tests/frontmatter.test.cjs. An earlier commit on this branch
re-created a standalone duplicate under that false premise; folded its
genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures,
block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the
duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the
ADR's §8.1 note, including the roadmap-sibling claim (no such file exists).

Finding 7: the golden serializer sorted object keys, making it structurally
blind to the key-order-parity invariant ADR-3473 §8.1 actually claims.
Made it order-preserving and regenerated the golden fixture from a
standalone compile of the legacy (pre-#3881) parser at ddde001af; the
current parser matches it with zero undocumented divergences, confirming
key-order parity genuinely holds. Extended row A2 table-driven across 6 of
the remaining 7 transitionCore kinds (all pass) plus documented, by
execution, a newly-discovered 8th-site regression: state.cts's
cmdStateCompletePhase calls the same preservation helper but its result is
clobbered by a later unconditional resync — filed as a distinct finding
rather than fixed here (touches syncAndPreserveStateMd, outside this
change's verified scope).

Refs #3881

* fix(#3881): preserve unparseable frontmatter through the CLI write path

Characterization (executed, before/after shown): case (b), not (a). The
frontmatter FENCE survives — `state complete-phase` on a conflict-marked
STATE.md returns success and a well-formed, freshly-derived frontmatter
block, not a document with no frontmatter at all. But the block's actual
content (the merge-conflict markers, and with them any signal to a human
that the document was in conflict) is silently discarded and replaced.

Root cause was two clobber sites, not one:

1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved
   `transformedContent` from readModifyWriteStateMd, finds {} + the
   FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh
   frontmatter block from the body anyway.
2. Even after (1) is fixed, applyPostSyncPreservation's own
   postFm/applyStatePreservation/authoritativeFm-reassertion machinery
   re-extracts frontmatter from syncedContent, restores curated fields
   from the pre-write snapshot, and reconstructs a NEW block again —
   confirmed live via `state begin-phase`, which still lost the markers
   after fixing (1) alone.

Both are now guarded by the same predicate (isUnparseableFrontmatter,
checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not
parse and the caller is not on ADR-3408 §8.3's closed "body wins" list,
both functions return their input content unchanged rather than
re-deriving over it. The closed list (cmdStateSync #905,
/gsd-health --repair's REGENERATE_STATE, both routed only through
writeStateMd, which never reaches applyPostSyncPreservation and passes
sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is
untouched — neither widened nor narrowed; verified by execution that
`state sync` still overwrites the conflict-marked block exactly as before.

Other verbs sharing the same readModifyWriteStateMd path were checked and
were equally affected before this fix: state update, query state.patch,
and state begin-phase all lost the conflict markers (RED, shown by
execution), and all three now preserve them (GREEN). Covered table-driven
in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe
block, which drives the real CLI verbs via runGsdTools — not just the pure
transitionCore layer the earlier A2 rows exercised — plus a control
asserting state sync's body-wins contract is unchanged.

Refs #3881

* fix(#3881): restore the parse surface's prototype and fix remote-runner failures

Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched.

Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write.

Refs #3881

* chore(#3881): backfill changeset PR number

Refs #3881

* test(#3881): make golden parity resistant to unrelated tree churn

A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus.

Refs #3881

* test(#3881): make the parser golden hermetic instead of tree-keyed

This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the
exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md,
docs/*.md) the prior golden pinned by tracked path. Any PR editing one of
those files' frontmatter for reasons unrelated to the parser (an
argument-hint addition, an allowed-tools tweak) turned the suite red, and the
reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to
catch a real regression. Excluding .changeset/** was not enough; the design
itself was wrong: a regression fixture must not be keyed to mutable repo
paths, and a single 376-entry JSON every such PR touches is also a
guaranteed merge-conflict surface.

Rebuilt the fixture to carry its own documents: each of 51 entries stores a
stable id, literal documentText (shrunk from a real ddde001af-era corpus
document), and an expectedParse captured independently from the
pre-migration legacy parser (git show ddde001af:src/frontmatter.cts,
compiled standalone against its byte-identical sibling modules). The test
reads no tracked path, shells out to no git command, and enumerates no tree
-- a PR editing commands/gsd/help.md cannot affect it. Every entry's
reconstruction was verified at capture time to reproduce both the current
and legacy parser's output on the original document; 0 of 51 candidates
were dropped by that check (1, the deliberately-unterminated
unclosed-block.md adversarial fixture, has no closing fence to truncate at
and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now
diverges:true entries) and the D2 order-preserving structural serializer
that keeps the comparison from passing vacuously; dropped the
tree-enumeration helpers, the coverage floor, the post-capture-skip logic,
and the vanished-file check -- all artifacts of the path-keyed design.

Refs #3881

* fix(#3881): resolve vendored-deps paths independently of cwd shape

Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on
windows-latest CI: the test passed absolute scratch-file paths into
compareFiles()/checkRow(), whose helpers joined every input onto ROOT
via path.join(ROOT, rel), producing garbage when the input was already
absolute. It surfaced on windows-latest specifically because GitHub's
Windows runners checkout the repo on a different drive than TEMP, so
path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged
(no relative traversal is representable across drives) rather than the
relative form the test assumed. The remote gsd-test runner this repo
gates pushes on is Linux-only and could never have caught this;
GitHub CI's windows-latest job is the only signal that does, and it did.

Fixed the helper itself (scripts/lint-vendored-deps.cjs's new
resolvePath()) to treat an already-absolute input as absolute-in,
absolute-out instead of silently mis-joining it, and updated the test
to pass the scratch file's absolute path directly rather than relying
on a relative conversion that is not always representable. Kept every
mutation-sensor assertion intact and added coverage proving
resolvePath is a no-op for relative inputs and correctly passes
absolute ones through unchanged.

Refs #3881

* fix(#3881): warn when state sync regenerates over unparseable frontmatter

state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly
overwrites an unparseable frontmatter block per its 'body wins'
contract — that overwrite behavior is unchanged here. The defect was
the silence: synced:true/exit 0 gave no signal that the existing
block (including git merge-conflict markers) could not be parsed and
was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be
reported as authoritative when the derivation dropped input it could
not resolve') and §8.4 ('failure is a value').

Adds a gsd: warning — ... (#3881) line on stderr, matching the
existing #3573 precedent, and surfaces the same disclosure in the
JSON result's existing changes[] array so a machine consumer sees it
too. Exit code and synced:true are left unchanged — sync did what its
contract says.

REGENERATE_STATE (/gsd-health --repair's sibling on the same
sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally
refused by applyRepairs's dispatcher before runRepairAction ever runs
(src/health-diagnostic.cts), so it is not a live path today and is not
in scope for this fix.

Refs #3881

* fix(#3881): exit non-zero when a state command returns an error

Refs #3881

* chore(#3881): changeset for the state exit-code fix

Refs #3881

* fix(#3881): honor the documented --project-dir flag

Refs #3881

* revert(#3881): restore exit-0 result envelopes for state errors

Reverts 9638f2936 and its changeset. The change was wrong and the revert is
the correction.

This repo distinguishes two error mechanisms deliberately. error() in
src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure
path. output({error: ...}) writes a JSON result envelope to stdout and returns
normally with exit 0. The reverted commit converted 23 result-envelope sites
into hard failures, which is a different contract, not a bug fix.

tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope
contract directly -- a failing command exits 0 with a JSON error envelope and
must not publish state.json -- and the remote matrix run caught it along with
four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that
the original commit rewrote were encoding that real contract, not the bug it
claimed; they are restored.

Whether an error envelope on stdout with exit 0 is the right CLI design is a
genuine question, and it is section 8.4's rule ('failure is a value') with its
own phase. It is not something to flip inside this PR.

Refs #3881

* chore(#3881): backfill changeset PR number for the project-dir fix

Refs #3881

* test(#3881): keep the frontmatter mutation shard inside its time budget

The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard
cap. Root cause is NOT row-level spawn overhead (contrast the #2790/
core-utils precedent): the three shard test files' own logic runs in
~413ms total (356+30+27ms) with all 392 assertions passing. Instead,
src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to
the vendored YAML parser, proportionally growing the mutant count Stryker
generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner
bills the full 'node --test <3 files>' invocation once per mutant, and
node:test's default per-file process isolation forks a child process for
each of the three files on every one of those invocations — pure fork
overhead multiplied by a much larger mutant population.

Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares
isolation: 'none', and .github/workflows/mutation.yml passes
--test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e.
unchanged behavior — for the other 8 shards, which were not individually
audited for cross-file state leakage under shared-process execution).
Measured locally via node:test's run() API on the exact 3-file set:
isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same
392 passing assertions. The true CI-shard number can only be confirmed
on the GitHub Actions run (Stryker cannot run locally, and 'node --test'
is hard-blocked in this environment).

Refs #3881

* test(#3881): register the vendored-parser tests in the frontmatter mutation shard

stryker.config.mjs's own rule ("Keep this list in sync with the tests
arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew
src/frontmatter.cts from ~825 to 1496 lines but its new tests
(tests/feat-3881-yaml-parser-consequences.test.cjs,
tests/frontmatter-golden-parity.test.cjs,
tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in
tests/frontmatter.test.cjs) were never added to the frontmatter shard's
tests array, so Stryker's mutants in the new vendored-js-yaml adapter had
nothing constraining them. PR #3888 measured 55.8% against the 65 floor
(748 killed / 593 survived / 17 timeout) and the shard was separately
cancelled at 15m04s against the 15-minute per-shard cap.

Registers all four files (each earns its slot on evidence of a unique
constraining assertion, documented inline), gives the shard a
measured/projected 180-minute budget via a new per-module
timeoutMinutes field threaded through mutation.yml's job-level
timeout-minutes the same way isolation is threaded, and removes the
prior isolation:'none' override (re-measured at this file-set size, its
savings are within run-to-run noise, not worth the unaudited
cross-file-state-leakage risk).

Refs #3881

* feat(#3881): derive the mutation test list and ratchet the score floor

Refs #3881

* test(#3881): ratchet five stale mutation floors and close the frontmatter gap

Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1):
config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78,
context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both
scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs
RATCHET_BASELINE in the same diff per the ratchet's own contract.

Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in
tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented
near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order
semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's
leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter),
repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via
extractFrontmatter), and the null-byte sentinel round-trip surviving at region
offset 1. Did not lower minScore.

Refs #3881

* test(#3881): decouple the ratchet test from real module floors

The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour.

Refs #3881

* refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency

Refs #3881

* fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs

A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives.

Closes #2571
Refs #3881

---------

Co-authored-by: sim <sim@local>
2026-08-26 19:29:32 -04:00
Tom Boucher
a3d5841117 fix(#3714): deliver an explicitly pinned model to the Codex worktree executor, and drop an unusable one (#3891)
* test(#3714): failing-first coverage for the dropped Codex worktree model override

Pins the argv contract for the orchestrator-worktree process dispatch: an explicit
model_overrides pin must reach the child as --model, while an unpinned, empty,
inherit, or profile-only configuration must emit no flag at all.

Five of the eight matrix rows are CONTROLS that pass before the fix. They carry the
weight here because this is an over-emission bug waiting to happen: resolve-model
returns 'sonnet' for the unpinned, empty and profile-only cases, so a fix that
threads its return value into argv would satisfy the positive row and emit
--model sonnet to Codex on every unpinned install -- the documented 400 that
ADR-2313 exists to prevent. The controls are what separate the correct fix from
the obvious one.

* fix(#3714): deliver an explicitly pinned model to the Codex worktree executor

resolveOrchestratorExec had no model input at all -- the descriptor carried only
command/args/cwdFlag/promptFlag -- so a resolved override had no way to reach the
spawned process even in principle. The baked gsd-executor.toml could not compensate
because this path spawns a process rather than dispatching a named agent.

Three parts, and the third is the load-bearing one.

The descriptor gains modelFlag (codex: --model), keeping the per-host knowledge as
descriptor data exactly as cwdFlag and promptFlag already are, so the scheduler
grows no per-host branch. No other runtime declares it.

The seam appends [modelFlag, model] and stays MECHANICAL: it does not know the
inherit sentinel, does not know which models Codex rejects, and reads no config.
Its sibling codex-agent-toml states that rule outright -- callers decide what to
strip. Argv order is baseArgs, model, cwd, prompt so the prompt remains the final
positional token. Omitting the model is byte-identical to before. A model starting
with '-' now fails closed as unsafe_leading_dash_model, the same hazard the prompt
and cwd guards already reject and which was silently accepted before.

The policy lives at the caller and passes ONLY an explicit, non-sentinel per-agent
pin. Passing null as the runtime resolver is what keeps profile and tier derived
models out of argv, which is what Codex's session-only model posture requires: the
model resolver returns 'sonnet' for the unpinned, empty and profile-only cases, and
emitting that revives the documented 400 from #2310/#2311 that ADR-2313 removed. So
the gate is the presence of an explicit pin, never that a value came back.

* fix(#3714): enforce the real-Codex value policy the sibling surface already applies

Review found one root cause behind a BLOCKER, two MAJORs, two MINORs and an
argv-injection finding: the dispatch path gated on the PRESENCE of an explicit pin
but never applied the VALUE policy that generateCodexAgentToml already applies to
the same config key. Textbook generative divergence -- and the parity row I wrote
tested resolver parity, not this policy, so it could never have caught it.

BLOCKER: a global ~/.gsd/defaults.json model_overrides.gsd-executor of 'sonnet',
'opus' or 'claude-sonnet-4-5' reached codex exec --model verbatim. That is the
documented 400 from #2310/#2311, arriving through the explicit-pin door rather
than the tier door. The issue asks for an explicit REAL-CODEX pin; real-Codex was
unenforced. Now dropped with a warning via the isAnthropicFlavoredModel predicate
#3241 single-sourced for exactly this reason.

Also: values are trimmed, so a whitespace-only pin is blank rather than
--model "   "; 'inherit' is matched case- and whitespace-insensitively, so
'Inherit' and ' inherit ' no longer reach the wire; and a value outside a model-id
charset is dropped with a warning. That last one closes the injection surface --
.planning/config.json travels with a clone, and values like
'gpt-5 -c approval_policy=never' or a command-substitution value previously reached argv verbatim,
where the spawner is an agent writing bash.

Every rejection DROPS AND WARNS rather than failing closed. An unusable exec is not
degraded, it is fatal: the dispatch step halts the wave after the worktree already
exists, so a config typo would have aborted execute-phase. The stale comment
claiming it degrades to sequential is corrected.

Separately, the seam's own empty-model handling contradicted the committed contract
and failed four tests on the remote runner. An absent, null or empty model is not an
error -- it means use the host default, the same degradation cwdFlag:null already
expresses. Unlike a prompt, where empty is a hang rather than a degraded run, so
that one stays fail-closed. Non-string values still fail closed.

Tests: the CLI rows were not HOME-hermetic and read the developer's real
~/.gsd/defaults.json, which is precisely the file the BLOCKER is about; HOME and
USERPROFILE are now sandboxed per call. Adds the global-pin regression, the
injection shapes, the case-variant inherit rows, and a real cross-surface
divergence guard.

* fix(#3714): close the flag-shaped pin, the case-variant alias, and the warning sink

Round two of review found three more defects, two of which both engines reached
independently, and all three were mine.

The charset allowlist put the dash INSIDE the character class, so a value made only
of allowed characters passed the pin policy silently and then tripped the seam's
leading-dash guard, producing exec:null. That is the wave-fatal path the whole
drop-and-warn design exists to avoid, reachable from a committed config file: -c,
--config, -p and --dangerously-skip-permissions all reproduced it. It also regressed
hosts with no model flag at all, where a dash pin turned a previously working
kimi-code dispatch into exec:null. The first character is now anchored, so a
flag-shaped value is dropped and warned like every other rejection, and the comment
that claimed this path was unreachable is corrected.

isAnthropicFlavoredModel folded case on its substring arm but not on its alias-set
arm, so SONNET, Sonnet, OPUS and HAIKU all reached Codex argv while lowercase sonnet
was correctly dropped -- the same 400 the drop exists to prevent. The predicate is
the one #3241 single-sourced so these surfaces cannot diverge, so folding case there
fixes the install-side .toml surface too.

The warning wrote the rejected value RAW to stderr. Every value that fails the
charset test contains by definition the characters the charset excludes, so it was a
guaranteed-reachable raw-to-terminal sink: an OSC sequence in a committed config
reached the operator's terminal byte for byte, and truncation could sever an escape
before its reset. The value is now sanitized before truncation.

Also from review: the allowlist rejected Vertex version pins like text-bison@002, a
false positive on a real model id; the policy ran host-neutrally so hosts with no
model flag printed a misleading drop warning on every dispatch; the invalid_model
branch had no test at all; and the changeset disclosed only that a pin is delivered,
not that an unusable one is now dropped with a warning.

* fix(#3714): single-source the model-id charset, bound the pin, keep the flag diagnosis

Round three found no blocking findings on either engine. These are the three
correctness items left in code I added.

The charset existed TWICE -- once to accept a pin, once to render a rejected one in
the warning -- and the two copies had already drifted inside a single commit: '@'
was added to the accept class and not the render class, so a Vertex-shaped value
rejected for some other reason rendered as text-bison?002. Both are now derived
from one definition, with a parity test asserting every character the matcher
accepts survives the sanitizer unchanged, so they cannot drift again.

A pin reached argv unbounded. CLAUDE.md documents the hazard: execFileSync aborts
on Windows above 32,767 characters of argv. A model id has no reason to be long, so
a pin over 200 characters is dropped and warned rather than truncated -- a truncated
model id is a different model id. Boundary rows at 199, 200 and 201.

The first character is now required to be alphanumeric, so '@evil' and '/c' no
longer reach argv. Security rates both inert on codex today, so this is hardening
rather than a live defect; it is here because it is one character of regex and
resolveOrchestratorExec documents itself as a general descriptor-to-argv seam that
other hosts may adopt.

Tightening the anchor made the leading-dash branch unreachable and, with it,
regressed the diagnosis: '-c' began reporting 'unsafe characters' instead of
'looks like a flag/option'. The dash check now runs before the charset test, which
both restores the actionable message for the most likely user typo and keeps the
branch live. A test pins the distinction between the three rejection messages so
the branch cannot silently die again.

* test(#3714): make the charset parity guard actually guard, and remove a false-green trap

Review proved by mutation that my parity test could not do what its own comment
claimed. It bound the expected character set to a local that was assigned and
discarded, and the shared definition was not exported, so widening that definition
left the test green. A comment overstating what a test guards is worse than no
comment, because the next reader trusts it. The shared body is now exported and
asserted equal, and I watched the assertion fail against a widened definition
before keeping it.

The sanitizer regex was module-scope, carried the g flag and was exported.
Production only uses it with replace, which resets lastIndex, so there was no live
bug -- but test() on a g-flagged regex alternates between calls, so any future test
reaching for it would false-green. The g-flagged copy is now internal to the one
call site that needs it and the exported companion carries no flags, which removes
the footgun rather than documenting it.

Also notes at the shared definition that it is interpolated into both a positive and
a negated character class, so only plain characters and ranges are safe to add.

No runtime behavior changes: the accept set, the anchor, the length cap, the
sanitizer output, the rejection ordering and all three message texts are
byte-identical, re-verified through the real CLI.

* chore(#3714): backfill changeset pr number

* chore(#3714): re-trigger CI after the GitHub Actions outage

The workflow runs for this branch were created during the Actions major outage on
2026-08-26 and never got scheduled. They are wedged: GitHub reports them queued,
refuses to cancel them, and refuses to rerun them because it believes the workflow
is already running. The Tests run is also pinned to a superseded sha, so no test run
exists for the current head at all.

Actions is operational again and the repo-wide queue has drained, so a fresh push is
what creates schedulable runs. This commit is empty on purpose: nothing about the
change is being altered, and the verified content is byte-identical to 7860c6ccc.

---------

Co-authored-by: sim <sim@local>
2026-08-26 17:01:09 -04:00
Tom Boucher
63abcface9 feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools

The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git.

The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap.

unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell.

Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true.

An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes.

Closes #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3146): stop sync:launcher relocating a deliberate preamble placement

Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins.

Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture.

Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3146): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3146): document the FEATURES.md section-numbering practice

The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases.

Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set.

Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914).

Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:57:16 -04:00
Tom Boucher
aaf47c5fc2 fix(#3691): let every reviewer lane take a prompt cap, and make the documented global resolve (#3832)
* test(#3691): failing-first coverage for the reviewer prompt budget

No prompt cap can reach any CLI reviewer lane, by any configuration. Two
independent defects compound: all nine `transport: spawn` lanes declare
`promptBudgetKey: null`, so `budgetFor` returns on its first line; and the
documented global `review.max_prompt_tokens` is advertised in the schema
manifest but declared nowhere, so the resolver never materializes it and
`budgetFor`'s fallback is dead code.

Adds to tests/reviewer-config-federation.test.cjs, which already owns the
per-reviewer budget config-set/config-get idiom:

- a CLI lane inherits the global cap (RED: reports null)
- an http lane with the -1 sentinel inherits the global cap (RED: reports null)
- the resolved review surface carries max_prompt_tokens at all (RED: absent)
- per-lane overrides the global on a CLI lane
- the sentinel boundary: -1 inherits, 0 means do-not-trim and must NOT read as
  unset, 1 is the smallest real budget — the regression budgetFor's own comment
  warns about
- anti-tightening pins that must stay green: an empty config leaves every lane
  null, the three existing budgeted lanes are unchanged, and config-set still
  rejects a per-reviewer key naming something that is not a declared lane
- a fast-check property over the resolution contract itself, with -1, 0 and
  non-finite inputs generated explicitly rather than left to chance

Every row was reproduced by hand against the real CLI before being written, so
the RED/GREEN split is observed rather than predicted.

Refs #3691

* fix(#3691): let every reviewer lane take a prompt cap, and make the global resolve

No prompt cap could reach any CLI reviewer lane, by any configuration. Two
independent defects compounded.

The nine spawn-transport lanes — claude, coderabbit, antigravity, cursor,
gemini, codex, kimi-code, opencode, qwen — declared `promptBudgetKey: null`, so
`budgetFor` returned on its first line and `review-lane plan` reported
`promptBudget: null` no matter what was configured. Each now declares
`review.max_prompt_tokens_per_reviewer.<slug>` with the same `-1`-is-unset
sentinel the three local-server lanes already use.

Separately, the central `review.max_prompt_tokens` was listed in the schema
manifest's validKeys and documented as a supported setting, but declared
nowhere — the resolved surface is built from capability declarations plus the
defaults manifest, and neither carried it. `configGet` returned undefined and
`budgetFor`'s documented fallback was dead code. It is now declared with a
`null` default, exactly as docs/CONFIGURATION.md already specified, so the
default behavior is unchanged: nothing configured means nothing trims.

Two things the diagnosis had not predicted, found and fixed while implementing:

- `REVIEWER_LANES` in src/review-lane-descriptor.cts is a second, hardcoded
  registration site that `mergeReviewerLanes` prefers over the capability
  registry on a slug collision. Editing only the capability files left every
  CLI lane still null. Both sites now agree.
- The generated `gsd-core/bin/lib/capability-registry.cjs` was stale and masked
  the capability edits; regenerated with `npm run gen:capability-registry`
  rather than hand-edited.

docs/CONFIGURATION.md said "Only lanes that declare a budget key accept one —
today ollama, lm_studio and llama_cpp". That is false as of this change and is
corrected rather than left to rot.

The trim-versus-refuse question the issue raises is deliberately not taken up
here: the refusal path already exists for the case that matters — a reviewer
whose minimum set exceeds its budget is skipped rather than sent a misleading
prompt — and trimming above that floor is the documented, shipped design of the
feature. Changing it would alter behavior for the three lanes that already
work, which is not what the issue asks for.

Fixes #3691

* fix(#3691): document the new global and narrow an invariant this change obsoleted

The full suite surfaced two consequences of giving every CLI lane a budget key.

`review.max_prompt_tokens` entered CONFIG_DEFAULTS without a matching entry in
the planning-config reference, which config-field-docs guards. Documented,
including the sentinel semantics a reader needs: a per-lane value overrides the
global, `-1` means unset and inherits it, and `0` means "do not trim that lane"
and is not unset.

The #2797 federation guard asserted that "a lane with no model flag and no host
owns no config keys". That held only because budget keys existed solely on the
three local-server lanes, all of which have hosts. A lane can now legitimately
own a config key for a third reason, so qwen tripped it.

The assertion is narrowed rather than weakened: such a lane must still own no
model key and no host key, and may own at most its own
`review.max_prompt_tokens_per_reviewer.<slug>` — never another lane's. That is
strictly more specific in the dimensions that still matter. Proven to still
bite: hypothetically giving qwen a `review.models.qwen` key fails it with
`model/host: review.models.qwen`. The name and comment cite #3691 for why the
premise changed, so a reader sees a deliberate narrowing, not erosion.

Checked the sibling assertions in that describe block; the other three do not
rest on the obsolete premise and are untouched.

Refs #3691

* fix(#3685): port the write-flag content-change contract to its three sibling sites

#3685 fixed `phase complete`'s `roadmap_updated` / `state_updated`, which
reported `fs.existsSync(path)` rather than whether the transaction wrote
anything. Three sibling sites carried the identical defect and are ported here.

- `cmdPhaseRemove` reported `roadmap_updated: true`, hardcoded.
  `updateRoadmapAfterPhaseRemoval` now returns whether the content changed and
  the flag reports it. #2640/#2974 already fixed `state_updated` at this same
  call site and left this one behind, so the correct shape was adjacent.
- `cmdMilestoneComplete` reported `state_updated: fs.existsSync(statePath)` —
  byte-identical to #3685's bug in a different command.
- `cmdMilestoneComplete` reported `milestones_updated: true`, hardcoded, never
  consulting the MILESTONES.md write.

`gsd-core/workflows/remove-phase.md:100` extracts `roadmap_updated` for display
and never branches on it, so the flip from always-true to content-based changes
no workflow behavior. Verified by reading the step, not assumed.

One trap found while implementing: the obvious in-memory
`finalContent !== originalStateContent` comparison — copying `cmdPhaseComplete`'s
shipped shape verbatim — gives a FALSE POSITIVE for milestone completion.
`platformWriteSync` normalizes Markdown at write time, and the milestone-closure
transform regenerates `## Current Position` fresh on every call, so its
pre-normalize output always differs from the already-normalized file on disk
even when the persisted bytes are identical. The comparison is therefore made
against the post-write on-disk content. `cmdPhaseComplete`'s own comparisons are
left untouched — their repeat-no-op tests pass, so they are not exposed to this
artifact.

`milestones_updated` has no reachable no-op: the MILESTONES.md write
unconditionally appends an entry every call. Only the true direction is pinned,
documented inline rather than faked with a passing test.

Refs #3685

* fix(#3685): compare write-flag content through the writer's own normalizer

An independent reviewer disproved a claim made while porting #3685's contract
to its sibling sites: that `cmdPhaseComplete`'s comparisons were not exposed to
the Markdown-normalization artifact already diagnosed in `cmdMilestoneComplete`.

`platformWriteSync` normalizes on write — CRLF stripped, blank-line runs
collapsed, a blank line inserted after a heading, a single trailing newline
enforced. Every flag that compares the PRE-normalization in-memory string
against the on-disk pre-image can therefore report a change when the persisted
bytes are identical. `cmdMilestoneComplete` had been worked around by re-reading
the file after the write; the other sites compared raw strings.

All of them now go through one exported seam,
`contentChangedAfterNormalize(filePath, before, after)`, which normalizes both
sides exactly as the writer does. That removes the extra disk read the milestone
workaround needed, and makes the sites agree by construction rather than by
four independent implementations of one rule — the divergence the repo names as
an anti-pattern.

Reachability, stated precisely rather than uniformly: the seam is load-bearing
at `cmdPhaseComplete`'s `roadmapUpdated`, `requirementsUpdated` and
`stateUpdated`, where section-rewrite logic genuinely regenerates content into a
different-but-normalization-equivalent shape. At
`updateRoadmapAfterPhaseRemoval` it is defense-in-depth: the no-match branch
never reassigns `content`, so the raw comparison was already correct there. The
first analysis claimed the reverse; this is the corrected finding.

Also fixes an unsound test premise the remote suite caught. The byte-identity
precondition in `roadmap_updated is false when ROADMAP.md comes out
byte-identical` asserted against a hand-authored, un-normalized fixture — so the
very first write reformatted it and the file could not come back identical. The
fixture is now written already-normalized, so the assertion compares a
normalized pre-image against a normalized post-image and still fails if the flag
regresses to a hardcoded `true`. Not platform-specific; it reproduces on macOS
too, and the earlier local check simply never exercised it.

The sibling true-direction and milestone tests were checked for the same premise
and do not share it — they assert `notEqual`, or compare two post-write states
produced through the same normalizing seam.

Refs #3685

* chore(changeset): backfill PR number for #3691 fragment

---------

Co-authored-by: sim <sim@local>
2026-08-24 19:39:51 -04:00
Tom Boucher
a2387a0545 feat(#3034): add opt-in parallel reviewer lanes (#3822)
* test(#3034): failing-first coverage for opt-in parallel reviewer lanes

Executes the real invoke_reviewers dispatch block from review.md against a
stubbed gsd_run seam rather than pattern-matching the workflow text, so the
two properties that actually carry risk are observable: that every lane is
joined before aggregation, and that concurrent lanes cannot tear a line in
gsd-review-lane-results.jsonl.

Concurrency is proven by a barrier fixture, not by elapsed time -- each stub
lane blocks until all lanes have checked in, which can only complete if they
overlap.

Red against the current sequential dispatch, by design.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3034): add opt-in parallel reviewer lanes

Reviewer lanes within one review pass inspect the same immutable plan
snapshot and have no dependency on one another, but were dispatched strictly
one at a time, so a multi-reviewer pass cost roughly the sum of its lanes.
The serialization is a deliberate protection against provider rate limits,
so it stays the default; review.parallel_lanes opts a project out of it.

The loop body is hoisted into run_review_lane so the sequential and
concurrent paths share one body -- two hand-synced dispatch bodies is the
divergence class ADR-2782 spent a phase deleting. Each lane writes a
slug-scoped result file, concatenated in selection order after the join:
concurrent O_APPEND is atomic only below PIPE_BUF, and write_reviews parses
that JSONL to render the models:/model_sources: frontmatter, so a torn line
is a broken REVIEWS.md rather than a cosmetic log defect. Aggregating in
selection order also keeps the artifact byte-identical between the two paths.

The guard is strict equality on "true" and falls back to sequential when
config-get fails -- the opposite polarity from the commit_docs guard,
because failing open here fires the very requests the default prevents.

Also corrects docs/COMMANDS.md and its four locale mirrors, which described
--all as running every configured reviewer in parallel when dispatch was in
fact sequential.

Closes #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3034): de-duplicate dispatch slugs and scope lane locals

Review finding (Standards axis): a slug repeated in SELECTED_REVIEWERS would
put two concurrent background jobs on the same > -truncated per-lane result
file. The shared-append form this replaced could not corrupt itself that way,
so de-duplicating is what keeps the concurrent path no worse than the
sequential one.

Selection de-dupes today -- the roster is a Set and review.default_reviewers
normalizes lowercase-unique -- but reachability analysis is not a contract,
which is the same reason the roster derivation itself is guarded.

Splitting once into DISPATCH_SLUGS also removes the duplicated tr-split the
same review flagged: the dispatch and aggregation loops now share one list,
which is what guarantees they walk the same slugs in the same order. A plain
string accumulator rather than an array, because zsh and bash disagree on
array indexing and this block runs under both.

Also scopes run_review_lane's locals. Not a live fix -- each dispatched call
already forks its own subshell -- but it makes the isolation a property of the
function rather than of the dispatch mechanism happening to fork.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3034): acknowledge review.md growth, drop spent 2295 ack

The differential attribution gate reported review.md growing 4173 bytes
(30712 -> 34885) with no live acknowledgment. Adds the per-PR fragment it
asks for, naming only the one path it reported.

Deleting tests/emitted-drift-acks/2295-resolved-model.json is required, not
opportunistic. That fragment declared review.md and nothing else, and its
ripple is already absorbed into the base, so it is spent -- it can no longer
clear anything, which is why the gate still reported review.md as
unacknowledged. It could not simply be left alone either: two ack sources may
never name the same path, so it blocked this PR's fragment outright.
CONTRIBUTING is explicit that a fragment whose last entry is removed gets
deleted with it, because an empty fragment signals nothing while its presence
reads as a live alarm.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3034): backfill changeset PR number

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 15:25:22 -04:00
Tom Boucher
004e9dd741 fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability

RED by construction. Binds to behavior renderEffortForRuntime does not yet
have: an optional third `model` argument, a per-model advertised-level table,
`max` passing through instead of clamping to `xhigh`, `minimal` clamping to
`low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/
`reason`) so a downgrade is legible from resolver output rather than silent.

Two of these pin defects that exist on next today:

- `max` is discarded. Both Codex models whose catalog entries are retrievable
  (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing.
- `minimal` is emitted to a model that refuses it. providerPresets.openai.
  haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's
  advertised floor is `low`. GSD is sending a value into a document Codex
  itself validates. The parity test is what pins that fixed, and it names the
  offending path/model/effort when it trips.

Also corrects tests/model-resolver.test.cjs:351, which asserted
renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as
though it were a contract. ADR-443 recorded "Codex has no max" as fact and it
was true when written; Codex has since added both `max` and `ultra`. That is a
stale premise, so the assertion is corrected here rather than worked around.

The property test asserts the invariant the whole change exists for: a rendered
effort is always a level the target model actually advertises, or an explicit
rejection. There is no third outcome.

* fix(#3007): resolve Codex effort per model, and make every clamp visible

Codex declares supported_reasoning_levels per MODEL and validates against it,
so a single per-runtime capability set cannot be right for all of them. GSD's
was wrong in both directions at once.

`max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped
max -> xhigh on that basis; it was accurate when written, and Codex has since
added both `max` and `ultra`. Every Codex model whose catalog entry is
retrievable advertises `max`, so the clamp was discarding a level the provider
supports, silently, on the most-used path.

`minimal` stops reaching Codex. No Codex model advertises it -- both retrievable
entries floor at `low` -- yet providerPresets.openai.haiku.low paired
gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the
receiver validates and refuses into a file the receiver reads. Being
unconservative in what you send is the half of Postel's rule with no defensible
reading, so that preset is corrected and a parity test pins it.

`ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum
reasoning with automatic task delegation": at ultra, effective_multi_agent_mode
returns Proactive and Codex spawns sub-agents on its own initiative, underneath
GSD's orchestration rather than inside it (#2167). It is a mode switch, not a
reasoning depth, so it is not added to the universal ladder -- which stays
provider-agnostic by ADR-443's design -- and it is rejected even for
gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered
and rejected: that silently discards what the user actually asked for.

Clamping is now visible. RenderedEffort carries requested/clamped/reason and
resolve-execution surfaces them. The previous table clamped correctly but
invisibly, so a user asking for `max` on Codex had no way to find out they were
getting `xhigh` -- exactly the failure mode the robustness principle's modern
critique warns about, and why "be liberal" has to mean "liberal and loud".

Also closes a latent trap found while reviewing the implementation: the clamp-up
loop walks the ladder upward, and for a future model advertising `ultra` but not
`max` it would have selected `ultra` as the clamp target -- re-entering by the
back door the mode the rejection above exists to keep out. A clamp may never
produce a value that a direct request for that value would refuse. Unreachable
with today's catalog, which is why no test caught it; a test now asserts the
invariant directly.

Signature stability is preserved: the third `model` argument is optional and the
two-argument form still resolves, against the family baseline. That form's
BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old
answer would have fixed the defect only where a model happened to be threaded
through and left it live everywhere else.

tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract
and is corrected here rather than worked around.

* fix(#3007): close every review finding on the Codex effort alignment

Two isolated reviewers, correctness and security. Both found the same two
blockers, and the per-model work was inert on every surface that matters until
this commit.

BLOCKER — resolve-execution never passed the model and discarded the clamp.
cmdResolveExecution called the two-argument form and emitted only
effort_rendered/effort_param/effort_propagation, so the per-model table was
unreachable from production code (tests were its only caller) and requested/
clamped/reason were computed and thrown away. Requested outcome 3 names "the
effective rendered effort in resolver output" specifically, so the feature was
unmet on the exact surface the issue asks for. Now passes the resolved model and
emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the
existing key convention rather than introducing a nested object.

BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed
a nested {"effort": ...} sample; the real result is flat and those keys were
absent entirely. A reference doc asserting a JSON path a reader can copy is worse
than no doc. Corrected against the actual emitted key set.

MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex
kept minimal in its supported set and still clamped max down to xhigh, so the
invocation-time and install-time channels disagreed about the same runtime's
capability: --host codex with max emitted xhigh while the generated TOML said
max. This is the repo's documented generative-fix-divergence class, so both
tables now cross-reference each other and a parity test fails if they ever
diverge again.

MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null
_baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback
never fired and every effort rendered as null. And a non-array value made the Set
constructor throw at module load — model-catalog.cjs is required across the whole
CLI, so one bad JSON value killed every command, not just codex effort. Guarded
on size and filtered to array values; both degrade to the hardcoded baseline.

MAJOR — value widened to a nullable string with two consumers left behind.
runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a
null effort key in generated frontmatter); install-effort-resolver still declared
a non-nullable return, a structural lie that silently defeated null checking.
Both corrected, both omitting the key on null — the same posture as 'inherit',
where omission means "follow the host default".

MAJOR — the per-model table is inert today, and the docs now say so. All three
shipped models advertise the same usable range and ultra (sol's only
differentiator) is rejected for every model, so no observable output differs by
model. The table stays because Codex declares capability per model and the sets
are free to diverge — a single per-runtime assumption is precisely what went
stale and produced this issue — but overselling it as a visible per-model feature
would have been the same class of error as the doc blocker above.

Tests: three passed under a full revert and are strengthened rather than deleted,
since each guards a real contract (#3533's inherit rule, the undeclared-host
rule, off-ladder handling) — they now also assert the clamp-visibility fields,
which only exist after this change. The fast-check property is kept for its
shrinking, and a deterministic nested loop over the full cross-product now sits
beside it so coverage is exhaustive rather than sampled.

Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg
form and would have written a literal null reasoning effort on the ultra path;
CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT.
The installer defect was found by the co-change gate, not by a reviewer —
install.js is a historical co-change partner of model-catalog.cts that this diff
had not touched.

* test(#3007): correct assertions that pinned Codex's stale effort premise

Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the
shipped commit. Every one is a stale pin, not a defect: each was probed against
the built module before its expectation was changed, and none failed for a
reason other than this premise correction.

Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale
by a production change must not ride inside another commit, because the
release-sdk hotfix cherry-pick filter routes by subject prefix and a correction
buried under the wrong prefix ships a half-state (v1.42.3, #3621).

The most valuable one was tests/model-resolver.test.cjs's cross-provider
validity invariant, which hardcoded the Codex enum as
`minimal|low|medium|high|xhigh` and failed with "real API would 400". That
message is now false in both directions: Codex accepts `max`, and rejects
`minimal`, which no model advertises. The enum is corrected to
`low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the
"would the real API refuse this" check worth having, and it was right to fail
here. It simply carried the stale fact in its own fixture.

Test NAMES were corrected alongside their assertions wherever the name asserted
the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal
passthrough". A renamed test that still claims the old thing is worse than a
failing one, and a green test whose name states a falsehood is how the next
reader inherits the wrong premise.

Both channels are covered: install-time (renderEffortForRuntime, and the
generated .toml in install-runtime-artifacts) and invocation-time argv
(effort-surface-axis). They were deliberately brought into agreement in this
change, so their assertions had to move together.

Each site carries a #3007 comment recording that Codex gained max/ultra and that
capability is declared per model, so a future reader can tell this was a
deliberate premise correction rather than a test bent to fit an implementation.

* test(#3007): separate the effort-precedence case from the clamp case

The previous stale-assertion pass over-corrected one test. It saw
`effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`,
assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to
"max passes through" and changed the expectation to `max`. The remote runner
disagreed.

Reproduced against the real CLI: with that config and `gsd-planner`, the
resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`,
`effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE —
gsd-planner is heavy/opus tier and its routing-tier default outranks
`effort.default` — so `max` never reaches the renderer at all. The test says
nothing about clamping and never did; it only looked like a clamp pin because
both mechanisms happened to yield the same string.

Restored to `xhigh` and renamed to say what it actually tests. It now also
asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is
what makes it impossible to mistake for a clamp pin again: those two fields prove
the value is what the resolver produced rather than something the renderer
downgraded. Before #3007 there was no way to tell the two apart from the output —
which is precisely why the previous pass could not tell them apart either.

Added the test that was actually missing: `effort.agent_overrides`, which
outranks the tier default, so the requested level genuinely reaches the renderer
and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified
by probe before asserting.

One test now pins the precedence rule and the other pins the #3007 behavior, and
neither can be read as the other. That the clamp-visibility fields are what
resolved this is a small argument for having added them.

* chore(#3007): backfill changeset pr number to 3765

* test(#3007): put model-catalog under the mutation gate

The Stryker shard showed as `skipping` on this PR despite the diff rewriting
model-catalog's effort logic. That was legitimate, not a detection bug:
`model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the
whole module — including everything #3007 touches — sat entirely outside
mutation scoring with has_work "false".

Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs
is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no
temp dirs. That shape is not stylistic — it is the #2790 precedent this file
already documents. Stryker's command runner treats a whole `node --test <file>`
invocation as ONE test costing whatever its slowest case costs, and re-runs it
per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses
runGsdTools throughout) would reproduce exactly the 15-minute shard-cap
cancellation #2790 hit. The integration file is unaffected and keeps running in
full in the normal test job.

Coverage spans the module rather than only the diff, because the score is
measured over the whole file: effort rendering across every model and ladder
level in both channels, the prototype-chain host guard, the exported enums and
maps, isAnthropicFlavoredModel's provider namespacings, the profile projections,
nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are
worth naming — every uncovered exported function is score given away, and
mergeEffortTierDefaults turned out to have a genuinely interesting contract
(#3531: a partial override merges over the built-ins rather than replacing them,
and isValid gates the VALUE, not the tier name, so an unknown tier key is still
merged in). Every expectation was probed against the built module before being
asserted.

minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment.
Floors in this repo are measured, not chosen — the existing entries sit at 94, 75
and 56 — and they can only be measured in CI, because mutation shards run
`node --test`, which is hard-blocked locally. The first CI run on this branch
reports the real number and the floor gets ratcheted to it before merge. A
placeholder of 1 reaching `next` would make the gate decorative: it would pass
whether or not a single mutant is ever killed.

Note the target is "never regress from measured", not a fixed 80 — planning-inspect
sits at 56 and is documented as an accepted ratchet candidate.

* test(#3007): bootstrap model-catalog's mutation floor legally

The placeholder floor was structurally illegal and the remote run said so.
tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module
must carry a matching RATCHET_BASELINE entry in the same diff, minScore must
EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three.
That is the ratchet working exactly as intended — a floor nobody can satisfy
accidentally is the point of it.

Bootstrapped at 50 in both places. Fifty is not a measured score and the comment
says so plainly: it is the minimum the guard permits, and it coincides with
Stryker's own configured `break` threshold, so it is the lowest legal starting
point for a module that has never been measured. It still must be ratcheted to
floor(measured) - 1 before this PR merges.

Also corrected a real defect in the file's own instructions. "HOW TO UPDATE"
step 1 read "Run the per-module Stryker shard locally" — which cannot be done
here, and which the same file contradicts eighty lines further down, where the
#2790 scores are recorded as "not a local run; mutation shards run `node --test`,
hard-blocked in this repo's local environment". stryker.config.mjs confirms the
command runner invokes `node --test` once per mutant, and
.claude/hooks/block-local-node-test.sh denies exactly that. So the documented
first step sends the next contributor at a wall. Rewritten to describe the path
that works — push, read the measured score off the CI shard, then set the floor
and its baseline together in one diff — and to say why local measurement is not
available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched.

The two-step is inherent to the environment rather than a shortcut: a floor
cannot be measured before the first CI run exists, and the guard rightly refuses
to accept an unmeasured one below its minimum.

* test(#3007): ratchet model-catalog's mutation floor to its measured score

The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no
timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this
file's own rule, minScore = floor(measured) - 1, which is the same arithmetic
every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94.

Both halves moved together, because the ratchet guard asserts minScore equals its
RATCHET_BASELINE entry and would reject them drifting apart.

The spawn-free unit surface is vindicated by the clock: 57 seconds, against a
15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the
whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing
the shard at tests/model-resolver.test.cjs — #2790 recorded shards being
CANCELLED at that cap when they targeted a runGsdTools-heavy integration file.

The registry comment is rewritten rather than deleted. It previously warned that
the floor was provisional and must not ship that way; leaving that text next to a
measured floor would make the file lie in the other direction. It now records the
measurement the way the sibling entries do, including that 59.62 sits below
TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 —
comfortably clear of its own floor with real room to grow. Raise it as the tests
improve; never lower it.

Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is
an honest floor for a module that had NO mutation coverage at all an hour ago,
and it is now pinned so it cannot silently regress.

---------

Co-authored-by: sim <sim@local>
2026-08-22 20:51:55 -04:00
Tom Boucher
2f86278b5e fix(#3003): opt-in mechanism for intentional deletions in worktree.cleanup-wave (#3757)
* test(#3003): failing-first suite for declared deletions in cleanup-wave

Binds the guard's opt-in before it exists, so the suite is RED against next.

The rows that carry the weight are the over-authorization set: a directory
declaration must not authorize its children, a glob declaration must authorize
nothing, and a declaration must not act as a string prefix of another path.
Each of those BLOCKS, and each would PASS under a prefix, glob, or startsWith
matcher — which is how a path list quietly degrades into the boolean opt-in
#3003 explicitly rejected. The glob row matters most: declaredScopePrefix
already returns null ("matches everything") for a glob-leading pattern, correct
for the advisory it serves and catastrophic for a gate.

Also pinned: a failed deletion check blocks on its own reason rather than being
filtered into a pass; the block detail names only the undeclared residue so the
operator is not misdirected by paths that were fine; an entry with no
declaration blocks exactly as before; junk and non-array declarations do not
authorize; and a blocked entry still isolates rather than aborting the wave
(#2852, which must stay fixed).

Two advisory rows cover an interaction found while designing: git diff
--name-only includes deleted paths, so without unioning the declaration into
the #2596 scope check, authorizing a deletion would raise
SCOPE_OUT_OF_DECLARED against the very path just authorized.

A seeded property states the whole invariant the three over-authorization rows
sample: a deletion merges iff its normalized path is in the declared set.

* feat(#3003): declared deletions opt-in for the cleanup-wave guard

The deletions guard blocked the merge-back of any executor branch whose diff
removed a file, with no way to say a removal was intended. A plan that folded
one test file into a sibling could not be merged by the tool meant to merge it,
forcing a manual --no-ff outside the tool -- strictly less safe than what the
guard protects against.

A plan now declares removals in its own frontmatter (files_deleted), and that
list rides the same path files_modified already travels: plan-document parse ->
phase plan JSON -> the per-plan worktree gate -> record-agent/create
--deletions -> declared_deletions on the manifest entry -> the guard. The guard
blocks only the deletions NOT in that list.

A path list rather than a boolean, per the pinned decision: a boolean disarms
the guard for the whole entry, so an unexpected deletion riding along with a
declared one would pass unnoticed. Matching is exact after the module's shared
normalizer -- never a prefix, never a glob. Both would let one declaration
authorize a whole set, which is the mass-deletion accident the guard exists to
catch. That also means declaredScopePrefix is deliberately NOT reused here: it
returns null ("matches everything") for a glob-leading pattern, which is right
for the advisory it serves and would silently disarm a gate.

The block detail now carries only the undeclared residue, so an operator is not
sent looking at paths that were fine. A failed deletion check still blocks on
its own reason and is never filtered into a pass. A blocked entry still
isolates rather than aborting the wave (#2852).

The #2596 scope advisory unions the declaration into its declared set --
git diff --name-only includes deleted paths, so without that, authorizing a
deletion would immediately warn that the same path was out of declared scope.

Optional and additive throughout: files_deleted is absent from
PLAN_REQUIRED_FIELDS, a manifest entry without declared_deletions keeps the
original unconditional block, and omitting --deletions leaves the on-disk entry
shape untouched.

Supersedes the spent #2856 emitted-drift ack entry for execute-phase.md, the
same supersede that entry performed on #3370 and #3370 on #3324.

* fix(#3003): wire --deletions on every dispatch surface, not just one

Review found the feature inert on two of three dispatch paths. execute-phase.md
(harness inline) passed --deletions, but the orchestrator-worktree path
(executor-isolation-dispatch.md, worktree.create) and the Fleet-parallel batch
path (capabilities/claude-orchestration/fragments/execute-wave-pre.md,
worktree.record-agent) still passed only --files. A plan declaring
files_deleted would have merged on one path and been blocked on the other two
-- the exact bug #3003 exists to fix, left unfixed where most of the isolation
actually runs.

Worse, per-plan-worktree-gate.md already claimed --deletions was passed 'on the
same worktree.record-agent / worktree.create calls', which was false for both
untouched sites. A doc asserting coverage that does not exist is how a gap
survives review.

All four surfaces now pass the flag, verified by sweeping every .md under
gsd-core/, capabilities/, commands/, skills/ and agents/ that invokes
worktree.record-agent or worktree.create: each one that passes --files now also
passes --deletions. The isolation-dispatch note explains why this flag, unlike
--files, is not advisory -- omitting it does not skip a check, it blocks a
merge the plan declared.

Regenerates capability-registry.cjs, which the fragment edit made stale.

Neither newly-grown file needs an emitted-drift ack: executor-isolation-dispatch.md
sits under workflows/execute-phase/steps/ and execute-wave-pre.md under
capabilities/, both outside currentSizes()'s non-recursive scan of
gsd-core/workflows/ and agents/.

* docs(#3003): document files_deleted where a plan author will actually find it

The feature's entire user surface is one plan-frontmatter field, and the
canonical reference for that frontmatter -- docs/reference/plan-md.md, the table
that documents every other key -- never mentioned it. A field nobody can
discover ships as a field nobody uses. Adds the files_deleted row and an example
entry in all five locales (en, ja-JP, zh-CN, ko-KR, pt-BR), stating the property
that makes the opt-in safe: matching is exact per path after separator
normalization, with no globs and no directory prefixes, so a declaration can
never authorize more than it literally lists, and omitting the field keeps the
guard's original unconditional block.

Also corrects two claims in the scope-conformance how-to that this change made
false. Its opening paragraph described the recorded declared scope as
files_modified alone; declared_deletions is now unioned into that comparison.
Its "Renames are not detected specially" bullet asserted the deletions guard
blocks any entry whose diff contains a deletion, full stop -- which was the
whole point of #3003 and is no longer true. Reworked to say what now decides a
rename's fate: declare the old path in files_deleted and both halves become
ordinary paths for the advisory check, which is also why the old path needs no
separate files_modified entry.

Documentation that describes the pre-change behavior of the thing being changed
is worse than no documentation, because a reader trusts it.

* fix(#3003): close every review finding on the declared-deletions opt-in

Two independent isolated reviewers, correctness and security. Neither found a
blocker; both found real defects, and the directive treats a finding at any
severity as blocking. All of them are fixed here.

MAJOR -- the submodule worktree gate could not see a deletion-only plan.
per-plan-worktree-gate.md intersected $SUBMODULE_PATHS against $PLAN_FILES
alone, while $PLAN_DELETIONS was extracted and then never used. Before
files_deleted existed, a path had to appear in files_modified to be planned at
all, so the gate saw it; the new field plus the new docs telling authors a
deleted path needs no files_modified entry opened a hole where a plan whose only
submodule touch is a removal kept worktree isolation on -- the exact case #2772
disabled it for. Both channels now feed the intersection. Note the posture is
deliberately the OPPOSITE of the cleanup-wave guard: there the channels stay
apart because a deletion AUTHORIZATION must never be inferred; here they merge
because a safety fallback must never MISS a touch.

MAJOR -- same-wave conflict detection could not see a deletion. The planner's
implicit-dependency rule compared files_modified only, so plan A editing
src/x.ts and plan B declaring files_deleted: [src/x.ts] scored as conflict-free
and ran in parallel: one branch removing what the other is writing, which is the
sharpest conflict there is. Overlap is now computed across both channels.

MINOR (both reviewers, one root cause) -- the advisory union gave one field two
matching rules. declared_deletions was unioned into the scope list handed to
planWaveScopeConformance, which reads it with prefix-and-glob semantics. So a
field that is exact-match-only at the gate silently became wider at the
advisory: ["*.md"], inert at the gate, yielded a null prefix meaning "matches
everything" and muted the advisory completely, and ["src"] muted all of src/.
The union also activated the advisory on plans that declared no modification
scope at all, warning on every modified path. Replaced with subtraction from the
findings, gated on files_modified alone. One field, one rule, everywhere.

MINOR -- core.quotepath made the feature silently inert for non-ASCII paths.
git emits "tests/\303\251.ts" C-escaped and quoted, which never equals the
declared plain path, so a correctly declared deletion of tests/é.ts would block
forever with nothing pointing at the encoding. Both diffs now pass
-c core.quotepath=false.

NIT -- flag() consumed a following flag as a value, so --deletions --files x
swallowed --files and dropped both. Now treated as a missing declaration, which
fails closed. Fixed at both call sites; the helper is duplicated verbatim in
cmdWorktreeRecordAgent and cmdWorktreeCreate and leaving one would reintroduce it.

TEST -- one test passed for the wrong reason. "a declared deletion is in scope
for the advisory" asserted only that warnings omit the deleted path; under a
full revert the entry blocks first, warnings come back empty, and the negative
assertion passes anyway. It now asserts the entry actually merged, which is the
load-bearing half. Four regressions added, one per fix above.

Docs corrected rather than extended. The rename bullet in the scope-conformance
how-to claimed a rename whose delete side is undeclared never reaches the
advisory. Verified false: git's rename detection is on by default, so a pure
rename is a single R entry that appears in no --diff-filter=D output and was
never gated, before or after #3003. Only a rename that edits enough to fall
below the similarity threshold decomposes into add+delete. The pre-existing
sentence made the same wrong claim; this restates it correctly instead of
sharpening the error. The localized plan-md.md reference edits are reverted:
the PR template requires docs content added here to be English, and the
translations already lag by three fields, so English-only is the repo's
standing posture, not an oversight.

Agent-file size caps respected: gsd-planner.md is XL-tier by bytes but carries a
separate 49152-LF-CHAR cap asserted by four suites, so its edit is deliberately
terse and lands at 49141 with 11 chars of headroom, with the rationale moved to
docs/reference/plan-md.md, which has no cap. gsd-plan-checker.md lands at 49107
bytes, 45 under the LARGE cap. Both acks merged into the existing fragments that
already name those paths, since two ack sources may never name the same path.

* fix(#3003): decode git's path quoting instead of changing the git argv

The previous commit's non-ASCII fix turned the remote suite red: 44 failures,
42 of them "unexpected git call: -c core.quotepath=false diff --diff-filter=D
--name-only ...". The suite's git mocks match on exact argv, so adding two
flags to the deletions diff and the advisory diff invalidated every existing
fixture in tests/worktree-safety.test.cjs. Rewriting dozens of fixtures to
accommodate one flag would be paying a large Hyrum's-law bill to fix a small
defect.

Both execGit calls are reverted to their original argv. The C-quoting is now
decoded in normalizeScopePath instead, via a new decodeGitQuotedPath helper.
That is the better fix on its own merits, not merely the cheaper one: the git
argv is untouched so no fixture moves, the decode lands on the ONE normalizer
already applied to both sides of the comparison so the declared and reported
paths cannot disagree, and it holds regardless of the user's own core.quotepath
setting rather than only when we remember to override it.

A value not wrapped in a leading AND trailing quote is returned completely
untouched, so the plain-ASCII path -- the overwhelmingly common case -- is
byte-identical to before. Escapes decode to BYTES collected into a Buffer and
UTF-8 decoded only at the end, because \303\251 is two bytes forming one
character and decoding them separately yields mojibake. Malformed input never
throws: a trailing lone backslash or a short octal escape degrades to the
literal character, since one bad path must not take down a cleanup wave.

Caught while reviewing the helper: the non-escape branch pushed a UTF-16 code
unit rather than UTF-8 bytes. Git always escapes non-ASCII so its own output was
fine, but this normalizer runs on the DECLARED side too, and an author may write
a quoted path holding a literal é -- pushing 0xE9 alone is invalid UTF-8, so the
declaration would decode to a replacement character and silently stop matching.
That is precisely the failure this change removes, reintroduced on the other
side of the comparison. Now converts whole code points, surrogate pairs intact.

The other 2 failures: tests/parallel-dependent-plans.test.cjs pins the exact
unbackticked substring "files_modified overlap" in gsd-planner.md, and rewording
that comment to "declared-scope overlap" deleted it. The comment is restored
verbatim and the files_deleted change rides in the pseudocode and the Rule
sentence instead. Recorded in the ack fragment so the next contributor does not
rediscover it the same way.

Four regression tests cover the decode through the public cleanup-wave seam
(the helper is module-private): a declared non-ASCII deletion merges against a
C-quoted git report, the symmetric case where the DECLARATION is the quoted
form, an undeclared non-ASCII deletion still blocks with the residue naming the
decoded path an operator can act on, and a path merely containing a quote is
left alone. Plain ASCII was already covered and is not duplicated.

* fix(#3003): revert the leading-dash flag guard, the review nit was wrong

The remote suite came back with 2 failures, down from 44, and both point at the
same thing: tests/worktree-safety.test.cjs:7045 already pins the opposite
contract, deliberately.

  test('a flag-shaped --files value is not re-parsed as a flag', ...)
    recordAgent(['--files', '--branch'])
    -> files_modified === ['--branch']
    -> branch === 'worktree-agent-a1'  ("the real --branch value must be untouched")

So consuming the next argv element positionally, whatever its shape, is the
tested intent of this parser, not an oversight. The security reviewer's nit
claimed --deletions --files x would "swallow --files and drop both". It does
not: each flag runs its own indexOf, so --deletions records the literal
'--files' while --files independently still resolves to x. And that literal is
a path git never reports as deleted, so it authorizes nothing -- already
fail-closed with no guard at all. The guard bought no safety and silently
changed --files behavior along the way, outside this issue's scope.

Reverted at both call sites, which are byte-identical again, along with the test
asserting the reverted behavior and the docs sentence describing it. The nit is
recorded as REJECTED in the review artifact with the reasoning above, rather
than as fixed -- a finding that turns out to be wrong should leave a trace of
why, or the next reviewer files it again.

docs/CLI-TOOLS.md now states the positional-read behavior plainly instead, so
the next person meets it as documented intent rather than rediscovering it
through a red suite.

* chore(#3003): backfill changeset pr number to 3757

* test(#3003): cover parsePlanDocument's filesDeleted branch to clear the mutation gate

CI's Stryker shard for plan-document failed at 73.28 against a break threshold
of 75: 170 killed, 62 survived, 232 total. Eight of those survivors are the
filesDeleted block this issue added to parsePlanDocument, which shipped with no
direct coverage at all -- the field was exercised end to end through the
cleanup-wave tests, but the parser itself was never called with a plan that
declares it, so every mutant in the block lived.

Four tests, each pinned to specific mutants rather than written for coverage
percentage:

- absent key yields exactly [] -- kills the array-literal seed
  (["Stryker was here"]) and the `fmDeleted = true` conditional, which would
  otherwise produce ["true"]
- a scalar underscore `files_deleted:` wraps into a one-element array -- kills
  `fmDeleted = false`, the `&&` logical-operator swap, the `fm[""]` string
  mutation on the first operand, the emptied if-block, and the ternary's
  non-array branch
- an array-valued hyphenated `files-deleted:` maps element-wise -- kills the
  `fm[""]` mutation on the SECOND operand (only reachable when the legacy
  hyphen alias is the one carrying the value) and the ternary's array branch
- an empty list yields [] -- boundary case, and a genuinely distinct one from
  the absent key: [] is truthy in JS so it ENTERS the if, and only
  Array.isArray's true branch mapping over nothing produces the same []

Threshold arithmetic: 174 of 232 are needed for 75%, and these take it to about
178, so the shard clears with margin rather than landing on the line.

Every expected value was confirmed by executing the built parser before being
asserted, not inferred from reading the source.

---------

Co-authored-by: sim <sim@local>
2026-08-22 13:17:51 -04:00
Tom Boucher
9a69a86f42 enhance(#2971): strict planning filter mode for /gsd-pr-branch (#3720)
* test(#2971): failing-first suite for the pr-branch planning-path filter

Binds the not-yet-built planning.pr_strict mode and the corrected filter recipe
for /gsd-pr-branch across six layers: pure classification and forbidden-path
predicates, real-git fixtures that run the cherry-pick filter loop end to end,
config-key registration through the real CLI and both manifests, the executed
worktree-materialization claim the issue's triage asked to establish, fast-check
properties over arbitrary path sets, and a drift guard over the shipped workflow.

Two live defects in today's shipped recipe are pinned as regressions, both
reproduced empirically first: `git rm -r --cached` stages a deletion of any
.planning/ path the target branch already tracks, so the generated PR removes the
base branch's planning files; and the same command leaves the cherry-picked file
untracked on disk, so a second commit touching that path aborts the pick with
"untracked working tree files would be overwritten" and every remaining commit is
silently dropped.

The test helper parses the canonical path lists out of gsd-core/workflows/pr-branch.md
rather than restating them, so the workflow stays the single source of truth and the
suite cannot drift from what ships.

Refs #2971

* feat(#2971): strict planning filter mode for /gsd-pr-branch

Adds planning.pr_strict — a boolean, default false, that selects what
/gsd-pr-branch means by "filtered". Default mode is unchanged: structural
planning state survives into the PR branch and the nine transient
subdirectories do not. Strict mode drops every .planning/ path, structural
files included, and carries a commit over only when it touches at least one
file outside .planning/.

Strict mode is what makes planning.commit_docs: true safe for a project that
versions its planning tree locally but publishes none of it. The alternative
posture, commit_docs: false, silently costs parallel executor isolation — a
worktree is checked out from a commit, so an untracked or ignored .planning/
is simply absent inside it and the executor has no PLAN.md to read. That claim
is now established by an executed fixture rather than inherited.

The two path lists are declared once and both projections derived from them,
so create_pr_branch and verify can no longer disagree about what the filter
promised. verify previously counted every .planning/ path against a documented
success criterion of zero while create_pr_branch was specified to preserve five
structural files, so a correct run reported itself as failed on every phase
that touched STATE.md — which is every phase. It now asserts against the active
mode, and names the .planning/ paths default mode deliberately keeps rather
than trading a wrong signal for silence.

Two verified defects in the same recipe are fixed alongside, because strict
mode would have amplified both. `git rm -r --cached` staged a deletion for any
.planning/ path the target branch already tracked, so the generated PR removed
the base branch's planning files — under strict mode that would have been the
entire tree. The same command left the picked file untracked on disk, so a
second commit touching that path aborted the cherry-pick with "untracked
working tree files would be overwritten" and every remaining commit was
silently dropped. Both were reproduced against real git before being fixed.
The filter now forces excluded paths back to what the PR branch's HEAD carries,
in the index and the working tree; a conflict outside the filter halts instead
of being improvised past; a commit left empty by filtering is skipped rather
than failing. A clean-working-tree precondition makes the worktree half safe.

Closes #2971

* fix(#2971): unwind the checkout on a conflict halt, and test the real recipe

Two review findings, both fixed in place.

The isolated adversarial pass found that the conflict-outside-the-filter branch
exited while leaving the user checked out on the half-built PR branch with
cherry-pick state still live — this loop runs in the user's own working
directory, so stranding them there is a real cost even though it is not a
vulnerability. The branch now aborts the pick, returns to the original branch,
removes the partial PR branch, and says so before exiting.

The standards pass found the L2 fixtures executed a hand-written mirror of the
cherry-pick filter recipe rather than the recipe itself, so a reordering in the
workflow would not have been caught — and the order is load-bearing, since
restoring a path from HEAD before removing it inverts the filter. The helper now
extracts the canonical loop from the shipped workflow and the fixtures execute
that verbatim, which also gives the conflict-halt unwind above real coverage.
The drift guard additionally pins the two commands' relative order and asserts
the workflow carries exactly one canonical loop.

Also records the publication gate in the CONTEXT.md glossary next to the commit
gate it is distinct from.

Refs #2971

* fix(#2971): make the conflict-halt unwind actually unwind, and use the colon slash form

The remote matrix caught two defects in the previous commit.

The halt path claimed to restore the original branch but did not. `git
cherry-pick --abort` does not apply to a single `--no-commit` pick with no
sequencer file, and the fallback left the unmerged index in place, which makes
`git checkout` refuse — a failure the `2>/dev/null || true` then swallowed, so
the user was told they had been restored while still sitting on the half-built
PR branch. The unwind now drops sequencer state, hard-resets the disposable PR
branch to clear the unmerged index, and only claims a restore when the checkout
actually succeeded; when it does not, it says where the user is and gives them
the two commands to finish it by hand. Verified against real git: exit 1, the
conflict named, HEAD back on the original branch, the partial branch gone, a
clean tree and no CHERRY_PICK_HEAD.

Two runtime-loaded source artifacts used the retired `/gsd-<cmd>` hyphen form,
which names a command no runtime registers. The canonical authoring token for
workflows and references is `/gsd:<cmd>`; docs keep the hyphen form, so the
documentation added in this branch is unaffected. The comment in src/config.cts
moves to the colon form too, since it propagates into the generated lib.

Refs #2971

* docs(#2971): backfill PR number into the changeset fragments (#3720)

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:40 -04:00
Tom Boucher
14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00
Tom Boucher
8da2dd3ad2 feat(#2790): add read-only planning.inspect schema-v1 snapshot query (#3708)
* feat(#2790): add read-only planning.inspect schema-v1 snapshot query

Adds a read-only query emitting a schema-versioned JSON projection of .planning/
so downstream harness UIs can consume planning state without parsing GSD's
Markdown a second time.

Composed strictly from the ADR-3180 section 7 owners plus parsePlanDocument,
parseRequirements and parseUatItems; markdown structure is read through the
Markdown Sectionizer and Markdown Table Model seams. It declares its own flat
external schema rather than serializing PlanningSnapshot, which is the
diagnostic-rule subject and still growing.

Extracts plan-document parsing out of cmdPhasePlanIndex into a shared leaf
module so phase.plan-index and planning.inspect cannot drift, including the
plan-id derivation both surfaces report.

Also fixes parseRequirements dropping the separator delimiter used by the
shipped requirements template, surfaced while wiring the requirement rows.

* fix(#2790): close spec gaps and a raw-text test assertion found in review

Review findings from the standards, spec and security passes:

- phases[] rows carry goal and dependencies, the two per-phase elements the
  issue Summary names that had no corresponding field. Goal is bounded to the
  section's leading prose so the Depends-on line, the Plans checklist and the
  wave annotations are not duplicated into it.
- requirement rows carry their own diagnostic codes, so a consumer no longer
  has to string-parse the global diagnostics subject to correlate.
- roadmap_acceptance.checkbox is looked up through the phase-id key owners.
  It was compared raw against the on-disk directory name, so it read null for
  every real-world slugged phase directory and the evidence channel was inert.
- the hostile-input test asserts the structured payload instead of matching the
  raw stdout string. The absence proof over raw stdout is kept deliberately.

* fix(#2790): register planning in the runtime usage list and repair fixtures

Remote runner reported 9 failures on 9b3f9aa. Two root causes, both fixed:

- gsd-tools.cjs registered the planning family in HOST_COMMAND_ROUTERS but
  never added it to TOP_LEVEL_USAGE's Commands list. Those are two surfaces a
  parity test guards, and the top-of-file block comment is not the runtime
  help string. A real wiring gap that every local gate and three review passes
  missed.

- the new suite's fixtures could not produce a resolvable phase set. STATE.md
  frontmatter omitted the milestone field, which ADR-3180 7.2 rule 1 makes the
  primary milestone selector, so the phase set scoped unscoped and every
  percentage was correctly withheld. Separately declarePhase returned a path
  without creating the directory, so a phase declared but never written to left
  phases empty. Both reproduced against the built module before fixing.

No assertion was weakened. The withholding path is still exercised and still
returns null when the roadmap is absent.

* chore(#2790): backfill changeset pr number

* test(#2790): cover every enumerated matrix row and contain a symlink escape

Reverses a silent deferral. An earlier revision left 23 of the 78 enumerated
matrix rows unimplemented and 7 more as one-off manual checks, with a paragraph
in the artifact and the PR body describing the gap. CLAUDE.md is explicit that
such a note is not a fix and is not surfacing. The rows are implemented instead
and the manual-evidence bucket is gone: 49 test cases become 88, covering all 78.

Writing the symlink row proved a real leak: a *-PLAN.md symlinked outside
.planning/ had its content emitted into the payload, confirmed via a direct call
and the spawned CLI. readDocument now resolves target and planning root with
realpathSync and rejects an escape, returning the ordinary unreadable-document
shape. Tested both ways, because a containment check that over-rejects is its own
defect: an escaping symlink leaks nothing and degrades that plan alone, while a
legitimately relocated .planning/ symlink stays fully readable.

The three new modules are registered in the mutation COVERED registry, which had
been reporting has_work false and skipping the Stryker gate entirely. Provisional
non-binding floors so the shards run and report; raised to the measured value
before merge, since the registry forbids calibrating from a local run.

* fix(#2790): satisfy the mutation ratchet contract and scope the 1MB test

Remote runner reported 16 failures on 8c451ed. Two causes.

The COVERED registry has a paired contract the earlier commit violated: every
module needs a matching RATCHET_BASELINE entry, and minScore must be between 50
and 100 with minScore === baseline. The provisional floor of 1 was illegal on
both counts. All three modules now sit at 50 — the registry's own enforced
minimum — with matching baselines. The score cannot be measured locally: the
shard runs node --test, which this repo hard-blocks, so CI is the only source.
Floors are raised to the measured value once this PR's shards report; a shard
below 50 means the tests need strengthening, since the floor cannot go lower.

The 1MB test was measuring the test harness rather than the product. The command
handles the oversized payload correctly by spilling to a tmpfile and resolving it
back, but the resolved stdout then exceeds runGsdTools' maxBuffer and the helper
reports ENOBUFS. It now uses --pick so stdout stays one byte while the full 1MB
document is still read and parsed end to end.

* fix(#2790): wire containment across every document read this command drives

An isolated security review of the containment control found the boundary logic
sound but not comprehensively wired: two content reads reached the filesystem
without it.

An escaped phase DIRECTORY could enumerate external filenames into the file
fields and diagnostic subjects. Both enumeration sites now containment-check the
directory before reading. Worth recording that the leak was already prevented one
layer earlier than the review claimed: Dirent#isDirectory() reports false for a
directory symlink, so such a directory never becomes a phase row at all. The
guard is defense-in-depth for a direct caller and for platforms where a reparse
point reports as a directory.

A *-VERIFICATION.md symlinked outside the root leaked one frontmatter value
verbatim, because readVerificationStatus does its own read and copies an
unrecognized status into the payload's next_action. Closed from the consumer
side through that function's existing fs injection seam, so src/verification.cts
keeps its signature and its other callers are untouched.

The reviewer additionally rated a forged status: passed as an integrity bypass.
It is not: anyone able to plant the symlink can plant a real VERIFICATION.md
saying the same thing. The incremental risk is confidentiality, which is what
these fixes close.

src/plan-scan.cts is deliberately unchanged: isPlanSuperseded reads
symlink-followed content but yields only a derived boolean, no document text.

* test(#2790): give the mutation shards an in-process surface

Two Stryker shards were CANCELLED at the 15-minute cap, not failed on score.
CI log: 640 mutants instrumented, and the dry run reported 'Ran 1 tests in 20
seconds' because the shards pointed at the integration suite, where nearly every
case spawns a gsd-tools subprocess and Stryker's command runner treats the whole
test-runner invocation as a single test. 640 x 20s cannot finish in 15 minutes;
at the kill it was 27/640 with an ETA over an hour.

Every other COVERED module points at a property or unit file, and the workflow's
own paths filter lists exactly those two patterns. In-process is the intended
mutation surface; the shards were pointed at the wrong shape of test.

Adds tests/planning-inspect.unit.test.cjs — 39 cases in 10 describes that spawn
nothing and call the built modules directly. plan-document and the router need no
filesystem at all, one being a pure content-to-object parser and the other taking
an injected mock. The three shards now point here. The 91-case integration suite
is untouched and still runs in the normal test job.

* chore(#2790): ratchet mutation floors to the measured CI scores

CI run 32392791843 measured all three shards, which is the only source the
registry accepts — local runs count timeouts as kills and inflate badly.

  planning-command-router  95.65 -> floor 94
  plan-document            76.58 -> floor 75
  planning-inspect         57.03 -> floor 56

Applied the registry's own rule, floor(score) - 1, and updated RATCHET_BASELINE
to match, since the ratchet test enforces equality.

planning-inspect sits well below the file's target of 80 and is the obvious
ratchet candidate as its tests improve. planning-command-router already exceeds
the target. The placeholder comment about floors pending measurement is removed
rather than left standing as a false statement.

---------

Co-authored-by: sim <sim@local>
2026-08-20 13:42:43 -04:00
Tom Boucher
adb46cdd85 feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker

Binds the contract before any hook change exists: a `state ~N commits back`
segment gated on the state_head stamp landed by #2622, firing at the same
advisory threshold /gsd-health's W024 uses rather than at > 0.

Covers all five acceptance criteria — threshold parity (19/20/21 boundaries),
both renderers including formatGsdStateCompact, an exact spawn-count assertion,
repo-pinning and sub_repos degradation, and behavioral parity against
readStateHeadFreshness rather than a source-grep of the two fence copies.

52 example-based tests plus 5 seeded fast-check properties. Red now by design.

* feat(#2734): surface STATE.md commit-age on the statusline

Adds an opt-in `state ~N commits back` marker to the GSD-state segment,
consuming the `state_head` stamp and freshness contract landed by #2622.
A solo developer returning to a project reads "Phase 4, executing" in
STATE.md and acts on it, without noticing the codebase moved 40 commits
since that line was written. /gsd-health reports it as W024, but only if
you think to run it; the statusline is the surface you see without asking.

Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses,
not at > 0: with commit_docs:true the commit carrying a STATE.md sync
advances HEAD by one, so > 0 would alarm permanently on a fresh project.

Costs exactly one bounded git subprocess per render and none when
disabled. `rev-list --left-right --count` answers ancestry and distance
together, and repo pinning is a filesystem check mirroring
projectOwnsItsRepo rather than a --show-toplevel compare, which is
unreliable on macOS /private/var and Windows 8.3 paths.

Every unresolvable input degrades to the tri-state unknown -- the marker
is absent, never a "fresh" claim the project cannot substantiate: a
malformed stamp, a root that does not own its .git, a sub_repos
workspace, history rewound past the stamp, or git being unavailable.

Also collapses statusline config resolution onto one resolveStatuslineOptions()
seam. runStatusline() and renderStatusline() duplicated it byte-for-byte;
one copy is what keeps a newly-added key from reaching only one of them.

* test(#2734): route the e2e spawn through the process seam and fix fixture leaks

Review findings from the two orthogonal passes:

- `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched
  its stdout to test a pure function. It now calls resolveStatuslineOptions()
  directly — no subprocess, no text matching.
- `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate
  lives in runStatusline, which reads stdin), so it now spawns through
  tests/helpers/process-seam.cjs and proves the negative with a filesystem fact:
  the git shim appends to a marker file on every invocation, and the assertion is
  that the marker never appears. Stronger than asserting text is missing, and it
  drops the last stdout substring match in the block.
- Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead
  of a trailing cleanup(dir), which leaked the temp repo on assertion failure.
  derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it
  binds each directory at scheduling time rather than cleaning only the last.

Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong
expectation rather than finding a code defect: `percent` drives the progress bar
too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends
after it, which is what the test exists to prove.

CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well
as the new key; both are now enumerated.

* docs(#2734): backfill changeset PR number (#3700)

---------

Co-authored-by: sim <sim@local>
2026-08-20 00:35:01 -04:00
Tom Boucher
2fca0e17e4 enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides

Binds the not-yet-built code-review-depth module: segment-aware path-prefix
matching of a changed-file set against ordered {paths,depth} rules, resolution
order flag > strongest matching rule > global > standard, typed validation
errors, and the large-scope downgrade boundary. Also proves behaviorally that
workflow.code_review_depth_overrides is not yet a registered config key.

Refs #2554

* feat(#2554): resolve code review depth from path-scoped override rules

Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth}
rules matched against a review's changed-file set by segment-aware path-prefix
comparison. Resolution order is --depth= flag, then the strongest matching rule,
then workflow.code_review_depth, then standard; a matching rule replaces the
global rather than being max'd with it, so quick and standard rules stay
meaningful. Glob metacharacters are a hard configuration error rather than sugar
for a prefix, and malformed rules halt the review instead of degrading to
standard. The resolver is pure and reports its own provenance, so the workflow
can print the resolved depth and the rule that matched. The pre-existing
>50-file deep-to-standard downgrade moves into the module and now names the rule
it overrode.

The key is registered centrally rather than as a capability config slice: the
federated slice channel admits only boolean/string/number/enum, so an array
slice would be dropped as malformed.

Closes #2554

* test(#2554): correct depth-provenance assertions and pin out-of-repo paths

Two corrections to the failing-first suite. The source assertion for a
non-matching rule with no global configured expected 'config'; with no global
set the depth comes from the default, and a companion assertion tolerated
either value, so both passed against an implementation that derived provenance
from whether any rules existed rather than from where the depth came from.

The out-of-repo absolute-path case used a home-directory path that matched
neither implementation, so it never exercised the defect it named. It now pins
the discriminating cases: an absolute path outside the repo root must not match
a repo-relative rule, and one under the root must.

* docs(#2554): document path-scoped code review depth overrides

Reference rows for workflow.code_review_depth_overrides in the configuration,
features and commands references plus the locale copies that carry those tables,
and in the planning-config reference. Explanation of why escalation is
whole-review rather than per-file and why v1 is prefix-only. New how-to for
scoping review depth by path, carrying the configuration-error reason table and
the distinction between nothing to report and could not look. CONTEXT.md
glossary entry and the INVENTORY row for the new CLI module.

ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and
pt-BR FEATURES.md carry no code-review config table, so those files are
deliberately untouched.

* fix(#2554): make the depth-misconfiguration halt executable and reject control chars

Three review findings, all in this change.

The misconfiguration halt was prose rather than shell: the error-printing fence
was followed by an unconditional extraction fence, so an ok:false result threw
and left the depth empty instead of stopping the review. Prose is not a guard —
the two fences are now one block with a real conditional, and anything that is
not the literal string true fails closed.

An interior control character in a rule path survived validation and reached the
provenance string and the summary box; rule paths now reject control characters
via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is
unchanged. That in turn makes the field record safe to delimit, so the seven
node invocations that each re-parsed the same result to read one field collapse
to one.

Also corrects the glossary entry's illustrative paths, which the glossary-ref
check read as real repository references.

* fix(#2554): use the fast-check v4 string API and acknowledge workflow growth

Two failures from the remote matrix on d3111f45, both this branch's.

The property block built its segment arbitrary with fc.stringOf, removed in
fast-check v4. Because the arbitrary is constructed in the describe body, the
throw took out all four property tests rather than one — they had never
executed. Rewritten to fc.string({unit, ...}), the form this repo already uses
in emitted-attribution.test.cjs. Every other fast-check helper in the file was
audited against the installed module.

The emitted-attribution growth arm needed an acknowledgment for code-review.md,
which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is
spent — its ripple was absorbed when #3503 merged, and the base file is exactly
the 34435-byte baseline this growth is measured against — so it cannot clear
anything, while the ack lint hard-fails on a duplicate key across two sources.
Removed it in favor of the new fragment, which is exactly how #3503 itself
replaced the spent 3191 fragment.

* docs(#2554): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-19 22:45:35 -04:00
Tom Boucher
ea594300d9 fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard

* fix(#3606): validate hook-kind coverage at call sites and dispatch generically

* fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction

* fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex

* chore(#3606): regenerate install-tree fixtures for new wave-post fragment

* chore(#3606): sync canonical launcher preamble into new fragment

* fix(#3606): keep fragment preamble ahead of first gsd_run mention

* fix(#3606): revert sync script's preamble move in explore.md

* chore(#3606): regenerate derived manifests post-rebase

* chore(#3606): allowlist peer test files - base was red on the count lane

* chore(#3606): regenerate inventory for peer's verify-command-grounding doc

* chore(#3606): grounding test maps to its own module by longest prefix

* chore(#3606): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 16:41:27 -04:00
github-actions[bot]
97ca69a53e chore: sync next package version to 1.11.0 2026-08-19 13:51:49 +00:00
Tom Boucher
1bf73d957b enhance(#2295): record the resolved model per reviewer in REVIEWS.md frontmatter (#3649)
* test(#2295): failing-first coverage for per-lane resolved-model recording

* feat(#2295): record the resolved model per reviewer lane

* docs(#2295): document the recorded reviewer model and its provenance

* fix(#2295): refuse control characters in a recorded model value

* test(#2295): correct watermark assertions for the widened mark shape

* fix(#2295): anchor the role-manipulation injection pattern at a word boundary

* feat(#2295): record the applied reasoning effort in the model value

* chore(#2295): backfill changeset pr number

* chore(#2295): restore em-dash in changeset body

---------

Co-authored-by: sim <sim@local>
2026-08-19 08:59:37 -04:00
Tom Boucher
9e4f0e99ad fix(#3631): exclude only __pycache__-resident bytecode from the consent digest (#3650)
* test(3631): failing-first coverage for bytecode-cache in the consent hash

bundleContentHash digests a walk with no exclusion, so a routine 'python3 -m unittest'
inside a Python-backed capability bundle writes __pycache__ under the bundle, the
recomputed hash stops matching the consent record, and the capability silently goes
inactive — no error, no warning, and loop render-hooks then omits its step and gate.

Two distinct triggers, and the second is the sharper one: collectBundleEntries pushes a
{kind:'dir'} entry for EVERY directory and the digest emits a TAG_DIR marker for it, so an
EMPTY __pycache__/ flips the hash before a single .pyc is written. A fix filtering only
*.pyc would leave that live. Verified by execution against the built lib: 5 of 7 probe
rows diverge from intent today, including the empty-directory row.

The anti-regression rows are the point of the shape: editing a real scripts/m.py and
adding node_modules/pkg/index.js must BOTH still change the hash. node_modules is
deliberately not excludable — its contents are required at runtime, so dropping it from
the digest would stop consent binding executable content. The symlink row pins ordering:
exclusion must apply after the lstat fail-closed rejection, never before.

Refs #3631

* fix(3631): exclude derived bytecode caches from the consent digest

RED proven at e5ba8f1fe on the remote runner: 8 failures, exactly the rows predicted to
fail, with the four anti-regression rows already green.

collectBundleEntries now skips a hardcoded, gitignore-independent set from the DIGEST:
basenames __pycache__, .pytest_cache, .DS_Store, and any .pyc/.pyo file. Matching is
byte-exact on the raw Buffer name (the walk never utf8-decodes) and case-sensitive, so the
digest does not vary with how a name happens to be spelled on a case-insensitive volume.

Three properties were preserved deliberately, each pinned by a test:

  - The filter runs AFTER the lstat symlink/non-regular fail-closed rejection. Filtering
    first would have turned the exclusion into a way to smuggle a symlink past the check;
    a symlink named x.pyc still throws.
  - Excluded entries still count toward BUNDLE_MAX_FILES and BUNDLE_MAX_TOTAL_BYTES. The
    caps guard the WALK; the digest answers a different question, and exclusion must not
    become an unbounded-bytes hole.
  - An excluded DIRECTORY is neither emitted as a TAG_DIR marker nor recursed into. The
    directory marker was the sharper half of this bug: an empty __pycache__ flipped the
    hash before any .pyc existed, so a *.pyc-only filter would have left it live.

The issue proposed either a gitignore-aware walk or a list including node_modules. Both
are rejected. A consent binding must not delegate its scope to a .gitignore the bundle
author does not control — one line there would drop arbitrary executable content out of
the hash. And node_modules holds code that is required at runtime; excluding it would stop
consent binding executable content, turning a usability bug into a supply-chain hole. What
makes __pycache__ different is that CPython validates each .pyc against its sibling
source, which remains hashed, so a real code change still invalidates consent.

Docs: CONTEXT.md's 'EVERY regular file AND directory' claim is corrected in place.
ADR-2363's residual-gap section said the walk had 'no exclusions' — per
docs/adr/README.md ('ADRs are append-only') that is corrected by a dated amendment rather
than an in-place edit. Its D4 argument is unaffected: skill bodies are .md and stay bound.

Fixes #3631

* fix(3631): narrow the digest exclusion after two isolated security reviews

The first cut of this fix passed the full suite and was still wrong. Both orthogonal
reviews rejected it, and the second one found a hole that has nothing to do with Python.

HIGH — an excluded DIRECTORY was 'continue'd before recursion, so its whole subtree was
permanently outside the digest. Declared hook script paths allow '_', '.' and '/' with no
directory or extension rule, so hooks:[{script:'__pycache__/run.js'}] installed, executed
via node, and its bytes could be rewritten forever without moving the hash. Ship benign
v1, collect consent, then own the machine. No Python involved.

FALSE RATIONALE — the justification I wrote into the code, CONTEXT.md, the ADR amendment
and the changeset claimed CPython validates a cached .pyc against its sibling source, so
the source staying hashed kept consent honest. That is not true, and I proved it by
execution rather than argument: default timestamp invalidation compares only the source's
mtime and size, both settable by anyone who can write the bundle. A forged pyc ran while
the .py was byte-identical.

Also wrong: '*.pyc' matched anywhere, but a legacy sourceless scripts/x.pyc IS importable,
so excluding it was a live vector.

Narrowed to what is actually defensible:
  - a DIRECTORY named __pycache__/.pytest_cache has only its TAG_DIR marker suppressed;
    the walk still recurses and hashes every non-excluded child.
  - .pyc/.pyo are excluded ONLY when the parent basename is exactly __pycache__.
  - a regular FILE named __pycache__, and a DIRECTORY named x.pyc, stay bound.
  - declared hook paths containing a __pycache__/.pytest_cache segment or a .pyc/.pyo
    basename are now rejected in both validator copies — a file named .pyc can contain
    perfectly valid JavaScript, so the exclusion must not be reachable from a declared
    surface.

Accepted residual risk, stated plainly in ADR-2363 and CONTEXT.md instead of explained
away: a forged __pycache__/mod.pyc matching an unmodified, still-hashed mod.py executes
without moving the digest. Before this change that write was detected. It is accepted to
stop routine bytecode caching from silently deactivating capabilities, and it is bounded —
the attacker needs post-consent write access, everything outside __pycache__/*.pyc stays
hashed, and no declared surface can point into the excluded space.

Known limitation, not papered over: .pytest_cache CONTENTS still move the digest. Only the
directory marker is suppressed. Excluding that subtree would reopen the HIGH finding.

Refs #3631

* fix(3631): drop the .DS_Store exclusion and pin what the caps actually bind

Second round of isolated review findings. The hardening closed the two original holes —
both re-reviews confirmed that by execution — but it introduced a new one of the same
shape, and left three claims unbacked.

HIGH, self-inflicted: .DS_Store was excluded from the digest at any depth, but the hook
path validator was hardened only for __pycache__/.pytest_cache/.pyc/.pyo. So
script:'hooks/.DS_Store' was ACCEPTED, runnableHookCommand emits the bare quoted path for
a non-.js name (the branch .sh hooks already use), and capability-source copies it with
its mode bit intact. Ship it +x with a benign shebang, take consent, then rewrite it
forever — the digest never moves. Fixed by DELETING the .DS_Store exclusion rather than
teaching the validator about it: .DS_Store has nothing to do with this issue's Python
bytecode symptom, and an excluded filename is a permanently unhashed name. The narrower
the exclusion, the smaller the hole.

The residual-risk bound in ADR-2363 and CONTEXT.md claimed declared surfaces cannot reach
excluded space. That is false and is now stated correctly: node resolves an unregistered
extension through the default .js handler, so a hashed, consent-covered hooks/run.js that
requires '../__pycache__/mod.pyc' reaches it in one hop. The validator guard raises the
bar for DECLARED surfaces; it does not contain the risk. The two bounds that are real —
post-consent write access required, everything outside __pycache__/*.pyc still hashed —
are kept.

The BUNDLE_MAX_FILES boundary test had gone vacuous: it padded with root-level *.pyc,
which the hardening made non-excluded, so it no longer proved anything about excluded
entries while the ADR claimed the caps were test-pinned. It now pads __pycache__/f{i}.pyc,
with the arithmetic re-derived by execution (capability.json + the still-counted
__pycache__ dir + N). BUNDLE_MAX_TOTAL_BYTES had zero coverage at all and is now pinned by
a sparse 32 MiB __pycache__/big.pyc that must still trip the size cap — the test that
proves exclusion did not become an unbounded-bytes hole.

Added the parity assertion CLAUDE.md's Generative Fix Divergence rule requires for the two
isSafeHookScriptPath copies, and proved it can fail: mutating one BUILT copy to drop .pyo
made the parity check report the divergence. Also pinned semantics that were correct but
untested and would have survived mutation — __pycache__/sub/x.pyc stays hashed (the parent
resets to sub, which is the recursion threading itself), .pytest_cache/y.pyc stays hashed,
and .pyo in both directions, which was a free surviving mutant.

Changeset rewritten: it still described the rejected wholesale-exclusion semantics.

Refs #3631

* chore(3631): backfill changeset PR number (#3650)

---------

Co-authored-by: sim <sim@local>
2026-08-18 22:35:54 -04:00
Tom Boucher
bf87dd4156 enhance(#3617): one canonical Windows binary resolver in the platform seam (epic #3411 Phase 1) (#3621)
* feat(#3411): one canonical Windows binary resolver in the platform seam

CONTEXT.md declares src/shell-command-projection.cts the single OS-facing seam,
but Windows binary resolution had grown four divergent implementations outside
it. #3445 folded two of them together — inside gsd-core/bin/gsd-tools.cjs, not
the seam — so the declaration stayed untrue and execTool still had no handling
at all.

Lift the resolver into the seam as resolveExecutableBinary, and export the half
that actually executes as projectSpawnInvocation: CreateProcess cannot run a
.cmd/.bat, so the cmd.exe mediation is inseparable from the lookup and splitting
them is how the copies accumulated. cmd.exe is invoked with an explicit argv
array, never shell:true — CVE-2024-27980's vector and Node 26's DEP0190.

execTool now resolves on win32. POSIX is a strict no-op by construction, which
matters: execTool rates CRITICAL blast radius (167 symbols, 53 files).
gsd-tools.cjs deletes its private scan and its private mediation and delegates.

Two semantics grown beyond #3445's resolver, both additive: a name already
carrying a PATHEXT-listed extension is tried as-is before the append loop, and a
suffix outside PATHEXT is not treated as an extension.

Refs #3411

* fix(#3411): keep mediating a declared .cmd that PATH resolution misses

Standards review caught a narrowing against the code this replaces. gsd-tools.cjs
computed `target = resolveSpawnBinary(binary) || binary` and keyed the shim test
on `target`, so a declared .cmd mediated whether or not PATH resolution found it.
That is load-bearing: resolveExecutableBinary scans PATH only, while `cmd.exe /c`
also finds a batch file in the current directory.

Mediation now keys on the target — resolved path, else declared name. The ENOENT
contract still holds for BARE unresolved names, which is the case it was written
for. P9/P10 pin both halves.

Spec review found E1/E2/E3/E5 promised by 50-test-matrix.md but never written;
added. E3 is the integration proof that the CVE-relevant mediation fires through
execTool, not only through projectSpawnInvocation in isolation.

Also adds the CONTEXT.md glossary entry for the seam's new resolution ownership
(a PR gate) and the changeset fragment.

Refs #3411

* fix(#3617): pass mediated cmd.exe arguments verbatim so metacharacters cannot inject

The isolated security pass found the mediation shape carried an argument-injection
surface. libuv's quote_cmd_arg force-quotes an argv element only when it contains
a space, tab, or quote — never for a cmd metacharacter — and cmd.exe re-parses
everything after /c. So an arg of a&calc arrived unquoted and cmd ran calc.
Node's own CVE-2024-27980 escaping cannot help: it fires only when the spawned
FILE is the .bat/.cmd, and here the file is cmd.exe.

Caret-escaping is not a fix. It is correct only when libuv does not quote, and
libuv quotes whenever the arg also contains a space — no per-arg transform is
right in both cases. So build the command line and pass it through verbatim, the
shape Rust's std adopted for the sibling CVE-2024-24576: one outer quote pair
that cmd /c strips, every token inside force-quoted, embedded quotes doubled.

An argument containing CR or LF is refused rather than mediated — a newline
cannot be represented in a Windows command line, so mediating would silently
truncate. Failing visibly is correct.

Known limit, documented at the seam: %VAR% still expands inside a /c string and
has no escape outside a batch file. That is information disclosure, not arbitrary
execution, and is the same limit Rust's std documents.

This was byte-for-byte the shape #3445 shipped, so the fix closes it for the
reviewer-lane spawn path too, not only for execTool's newly reachable route.

Refs #3411

* docs(#3617): document the subprocess-execution security posture

Adds Layer 4 to the security model: why GSD never uses shell:true for binary
invocation (CVE-2024-27980, Node 26 DEP0190), why resolution is explicit and
never tries the bare name on Windows (the npm extensionless-shim trap behind
#3275), and why .cmd/.bat mediation builds a verbatim force-quoted command line
rather than relying on default escaping — Node's own CVE protection cannot fire
once the started program is cmd.exe.

The residual %VAR% expansion limit is stated plainly under Trade-offs rather
than left implicit: it is information disclosure, not arbitrary execution, and
callers passing untrusted text to a Windows .cmd should not assume the value
arrives byte-identical.

Docs-only; no code change.

Refs #3411

* chore(#3617): backfill changeset pr number 3621

* fix(#3617): read PATH, PATHEXT and ComSpec case-insensitively

The Windows CI lane on #3621 failed E5, and the root cause was a defect in the
implementation, not the assertion.

Windows names the variable Path, not PATH. process.env is a case-insensitive
proxy, so process.env.PATH works — but execTool builds
{ ...process.env, ...opts.env } whenever a caller supplies opts.env, and
spreading discards the proxy while keeping the OS's actual casing. The exact-case
env['PATH'] lookup then returned undefined, the PATH scan saw zero segments,
resolution returned null, and the change degraded to precisely the spawn ENOENT
it exists to fix. ComSpec and PATHEXT had the same exposure.

#3445's tests never caught it because they pass uppercase keys explicitly, and
neither did the Linux remote runner — this is a defect only the Windows lane
could see.

_envGet resolves a variable by exact match first (so a canonical caller pays no
scan) and falls back to a case-insensitive sweep.

R23 and P16 pin it and were proven RED by execution: with the fix stashed and
build:lib re-run, R23 returned null and P16 returned the cmd.exe default.

R24 was rewritten because the first version was vacuous — it staged foo.CMD, so
the default PATHEXT already contained .CMD and it passed against the broken code
for the wrong reason. It now stages foo.XYZ, an extension absent from the
default, and carries a negative control asserting that dropping the Pathext key
yields null. Re-proven RED the same way.

E5's assertion was corrected alongside the fix: 'PATH' in options.env expressed
the wrong contract. It now checks case-insensitively for the key.

Refs #3411

* fix(#3617): execTool spawns the declared name unless mediation is required

The Windows full-test lane on #3621 failed tests/graphify.test.cjs — the python3
identity check asserted 'python3' and got the absolute resolved path
C:\hostedtoolcache\windows\Python\3.12.10\x64\python3.EXE instead.

Those tests are correct and the change was wrong. They pin a long-standing
contract — execTool spawns the program name it was given — by spying on
spawnSync's first argument, and routing every win32 call through the projected
invocation broke it.

Resolving a .exe buys nothing. libuv's CreateProcess path already performs
PATH + PATHEXT search, which is why spawning a bare 'node' has always worked on
Windows. The only case the OS genuinely cannot spawn is a .cmd/.bat. So execTool
now adopts the projection only when mediation actually happened —
windowsVerbatimArguments is exactly that flag — and otherwise passes the declared
program and args through untouched.

40-design.md already rejected gratuitous change for this reason: symmetry is not
worth a behavior change to 53 files that fixes nothing. That reasoning was
applied to POSIX and missed the win32 non-batch case. Rows 5 and 20 now record
it, and the CONTEXT.md glossary states the caller-choice rule.

deps.spawn deliberately still adopts the resolved path: its hasBinary probe
answers from the same resolver, so probe and spawn must agree on the exact file
(#3445). The asymmetry is now documented at both call sites rather than latent.

E7 pins the restored contract and was verified by executing execTool against a
monkeypatched spawnSync: python3 in, python3 spawned.

Refs #3411

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:12:44 -04:00
Tom Boucher
9de4d67118 fix(#3579): a pointer-less session inherits the repo active-workstream marker (#3616)
* test(3579): failing-first coverage for repo-marker inheritance

A session that carries an identity but has never run 'workstream use' reads an
absent session pointer, resolves null, and composes the flat .planning tree even
when .planning/active-workstream names a live workstream. These tests fail on that
and pin the invariants the fix must not break: a session with its own pointer is
never repointed, and a session that merely lacked a pointer must never clear the
shared marker on another session's behalf.

* fix(3579): a pointer-less session inherits the repo active-workstream marker

RED proven at 157cae26: the three inheritance tests failed while every isolation and
negative control passed on base — the gap, and nothing else.

pickActiveWorkstreamAdapter returned exactly ONE adapter: the session-scoped one
whenever a session key existed, so the shared .planning/active-workstream marker was
never consulted. getWorkstreamSessionKey resolves a key from ~13 env vars or the
controlling TTY, so on any normal interactive terminal a key almost always exists —
which is why a session that had never run 'workstream use' read an absent pointer,
resolved null, and composed the FLAT planning tree even though the repo marker named a
live workstream. Reads misreported; writes corrupted the superseded flat STATE. Silent,
because the stale tree is well-formed.

This was a genuine design fork, not an oversight: references/workstream-flag.md
documented step 4 as a fallback 'when no session key exists', and the session isolation
that buys is deliberate (#2850). The issue's Agent Brief left the choice open and said
the reference doc should match whatever semantics ship. The maintainer ruled in chat for
inheritance.

Resolution now walks an ORDERED chain — session adapter first, shared second — and only
a null from the session adapter falls through to the marker. Strictly additive: it can
only turn a null into a name, never change a name that already resolves.

The dangerous part is clear() ownership. resolveFromChain treats chain[0] as owned: only
it is ever cleared, and only under selfHeal (getActiveWorkstream, never peek). An
INHERITED marker is read-only — a stale value there resolves null and the file is left
alone. Without that, one pointer-less session's read would delete the repo marker for
every other session, which is a worse bug than the one being fixed. Covered by a test
that asserts the marker still exists on disk after such a read.

peekActiveWorkstream inherits but still mutates nothing (#2850 — the statusline draws on
every render).

references/workstream-flag.md's Resolution Priority is rewritten to match, keeping the
session-isolation rationale and noting that inheritance does not weaken it: a session
that owns a pointer is never repointed.

Fixes #3579

* fix(3579): correct the guard diagnostics and lock the clear-semantics

Three review passes; every finding fixed inline.

MISSING ACCEPTANCE CRITERION (spec pass). The brief requires refusal diagnostics that
distinguish 'marker present but the session lookup missed it' from 'no workstream set at
all', and the two workstream-mode fail-safe guards were byte-for-byte untouched — still
emitting a generic 'no active workstream is set' even when a marker exists and merely
names a missing directory. Both guards (cmdPhaseComplete, cmdInitProgress) now branch on
a new read-only diagnoseUnresolvedActiveWorkstream, which reuses the SAME
resolvesToExistingWorkstream predicate resolveFromChain uses, so the diagnosis and the
resolution cannot disagree. Two typed reasons added to ERROR_REASON; both arms still
refuse — the fail-closed behavior is unchanged, only the message is now true.

REAL TEST FAILURE, not a flake. The remote run failed 'clearing one session does not
clear another session pointer'. That describe uses before() rather than beforeEach, so
one tmpDir is shared and an earlier test writes active-workstream=beta into it; under
inheritance the just-cleared session picks that marker up and resolves beta instead of
null. The failure is a CORRECT consequence of Option A surfaced through an
order-dependent fixture. The test now establishes its own marker state explicitly — its
real intent (clearing A must not disturb B's pointer) is preserved and not weakened — and
a new test pins the semantic deliberately: clearing a session pointer returns that
session to INHERITING the marker, it does not force flat mode. Documented in
references/workstream-flag.md, including how to actually get flat behavior.

Also from review: partial activeWorkstreamAdapters injection no longer silently
synthesizes a REAL filesystem adapter for the missing half (a latent test-isolation
trap); the duplicated validate-then-existsSync logic is factored into one predicate; and
the two try/finally test bodies are converted to t.after per CONTRIBUTING.

New coverage: whitespace/empty shared marker; a session whose OWN pointer is stale while
the marker names a different valid workstream (must self-heal to null, never inherit —
the isolation guarantee at its sharpest); and both new diagnostic arms asserted on
structured --json-errors output rather than prose.

* fix(3579): read resolvability with the non-mutating peek, not the self-healing resolver

Three of our own new tests failed on 7f5e706a. All three had ONE root cause, and none
was fixed by relaxing an assertion.

gsd-tools.cjs's bootstrap called the MUTATING getActiveWorkstream unconditionally on
every invocation, purely to populate routing env. On an unresolvable pointer that
self-healed — cleared it — BEFORE the dispatched command ran its own resolution. A second
read in the same process then observed already-cleared state:

- Isolation violation: a session whose own pointer was stale had it cleared by the
  bootstrap, so cmdWorkstreamGet's own resolution found a pointer-LESS session and
  inherited the shared marker ('beta' instead of null). Exactly the guarantee #2850 exists
  to protect, defeated across two calls rather than within one.
- Guard diagnostics: the guards' own truthiness check also used the mutating resolver, so
  it cleared the invalid marker and the immediately-following read-only diagnosis found
  nothing and reported none_active instead of marker_unresolved.

So a single invocation's answer depended on how many times it resolved. The bootstrap
self-heal is PRE-EXISTING and was harmless while pointer-less meant flat — inheritance is
what made it answer-changing, so this fix belongs here.

Every call site that only CHECKS resolvability — the bootstrap, both fail-safe guards'
truthiness check, and two informational init report fields — now uses the non-mutating
peekActiveWorkstream. Self-heal is unchanged in active-workstream-store and still fires
exactly once, at whichever site actually consumes the workstream.

Verified by driving the real CLI against temp fixtures, since the suite cannot run
locally: stale-own-pointer resolves null with the marker intact; both guard arms report
marker_unresolved with missing_workstream_dir / invalid_name and the marker survives;
no-marker still reports none_active; identity-less self-heal still deletes an invalid
marker byte-identically to pre-#3579; and a session with a valid own pointer still wins.

* chore(3579): backfill changeset PR number (#3616)

* test(3579): kill the surviving mutants in the new resolution code

CI's Stryker gate failed: active-workstream-store scored 79.45% against a break
threshold of 80 — 259 killed, 67 survived, at 'Ran 1.00 tests per mutant on average'.
The survivors cluster in the code this PR added (pickActiveWorkstreamAdapterChain,
resolvesToExistingWorkstream, resolveFromChain, diagnoseUnresolvedActiveWorkstream):
the CLI-level tests exercise those paths but do not DISCRIMINATE their branches, which
is precisely what a surviving mutant means.

Raised by strengthening assertions, never by touching the threshold. 21 unit tests added
to the existing unit suite, each written to fail under a specific named mutant, using the
module's injected adapter seams and createMemoryPointerAdapter so they stay hermetic
under Stryker's per-mutant reruns:

- chain shape with and without a session key, asserting length AND element identity
  (kills the if(false), the ': []' array mutant, and the block removal)
- partial adapter injection, asserting the missing half is an inert memory adapter that
  never touches the filesystem (kills the three '??' -> '&&' mutants)
- both arms of '!name || !validateWorkstreamName(name)' as SEPARATE tests — an absent
  name and a non-empty invalid one — which is what kills the '||' -> '&&' mutant
- self-heal discrimination: getActiveWorkstream must clear an unresolvable owned pointer
  and peekActiveWorkstream must not, asserted on adapter state after each
  (kills if(selfHeal) -> if(true))
- fallback arm both ways: a fallback that resolves and one that does not
- diagnoseUnresolvedActiveWorkstream asserted as a full object per case, with the reason
  strings compared exactly (kills present:true -> false and both StringLiteral mutants)

One mutant is deliberately left: 'if (chain.length === 0)' -> 'if (false)'. The branch is
structurally unreachable — the only chain source always returns a 1- or 2-element array
literal — and resolveFromChain is not exported. Killing it would mean exporting an
internal or deleting a defensive guard; neither is worth doing for a mutant, and the
score clears 80 without it. Recorded here rather than left unexplained.

Every new assertion was evaluated against the built module with real fixtures before
committing, since the suite cannot run locally.

---------

Co-authored-by: sim <sim@local>
2026-08-18 11:37:44 -04:00
Tom Boucher
fe64704ace enhance(#3588): add an opt-in commit_docs pre-commit hook (#3609)
* feat(#3588): add an opt-in commit_docs pre-commit hook

Final phase of epic #2292, scope narrowed to opt-in by maintainer decision:
default-on installation and the bin/install.js wiring it would have required
are explicitly out of scope.

Enabling is an explicit verb call. The hook is written to the repo's real hooks
dir resolved via git rev-parse --git-path hooks, so a linked worktree or
submodule whose .git is a FILE works rather than getting a literal .git/hooks
path. It refuses rather than overwrite a foreign pre-commit, refuses to delete
one it did not write, and refuses outright when core.hooksPath is already set --
a written-but-ignored hook is worse than a refusal. Ownership is detected by
marker presence, not byte-equality, so a user who appends a line does not make
it unrecognizable.

Deliberately NOT included: teaching cmdCheckCommit the per-phase commit_docs
tier. #3587 was still unmerged when this landed, and implementing precedence
against helpers that did not yet exist would have meant a second copy of the
resolution chain -- the divergence class this epic has spent three phases
fighting. That follows as its own change now that #3587 is on next.

The ordering constraint is recorded in the design doc: this must not merge
before #3587, or the hook would block a commit cmdCommit itself allows.

* fix(#3588): teach the commit_docs guard the per-phase tier and -z paths

Part 1, deferred until #3587 merged.

cmdCheckCommit read only project-level commit_docs, so once #3587 landed, a
phase with phase_commit_docs true under project false was ALLOWED by
query commit and BLOCKED by this guard -- and the hook shipped in this same
branch shells out to it. It now derives the staged phase via the single-owner
detectPhaseNumberFromFiles and resolves through #3587's own
resolveCommitDocsPolicy rather than a second precedence copy.

Also fixes a proven false negative in the harm direction. git diff --cached
--name-only C-style-quotes any path with non-ASCII or special characters, so a
staged .planning/cafe.md was emitted as a quoted string, failed
startsWith('.planning/'), and slipped past the guard entirely under
commit_docs:false. Reading with -z and splitting on NUL removes the quoting at
the source. The f.startsWith('.planning\\') branch was dead code under that
read -- git emits /-separated paths on every platform -- and is removed rather
than left implying coverage it never provided.

The earlier C7 test pinned the buggy behavior as intended; it now asserts the
file is detected and the commit refused.

Self-caught: the commit-docs-guard verb was wired into the routers by this
branch's earlier pass but missing from the top-level help listing.

* test(#3588): replace try/finally with t.after, add negative-routing cases

Standards review findings.

CONTRIBUTING bans try/finally inside a test body outright -- it masks failures
-- and B8 used one for worktree cleanup. Now t.after(), assertions unchanged.

The new commit-docs-guard command family had zero negative-routing coverage,
which CONTRIBUTING requires for any change to command dispatch. B11-B15 cover
no subcommand, unknown, empty string, whitespace-only and a flag-shaped value,
each asserting non-zero exit, a structured error, no stack trace, and -- the
one that matters for a command that writes into a user's repo -- that NO hook
is written in any of them.

Those tests were verified to fail when routeCommitDocsGuard's else-branch is
neutered, so they exercise the routing guard rather than any convenient error
path.

Also made two error() calls' control flow explicit with a return; they were
safe only because error() is typed never two files away.

* chore(#3588): backfill changeset pr number to 3609

* test(#3588): skip Windows-unrepresentable fixtures on win32

CI's Windows shards caught two of my own tests: fixtures whose filenames
contain a quote and a backslash. Both are illegal on Windows -- backslash is
the path separator, quote is invalid on NTFS -- so fixture creation failed
before any assertion ran.

Test-portability defect, not a production one. Those inputs cannot exist on
that platform, so the guard has nothing to detect there.

Both now check process.platform FIRST, before any fs or git call, and use
t.skip() rather than a bare return -- a bare return registers as a PASS and
would hide the gap it is meant to record. Each carries a comment saying the
input is unrepresentable rather than unverified, so nobody later re-enables it.

No padding added: the cafe.md case already exercises git's C-quoting path on
every platform, since non-ASCII names are legal on NTFS.

This is exactly the coverage the Linux-only remote matrix cannot provide, which
the PR body already stated -- CI's Windows shards are what caught it.

---------

Co-authored-by: sim <sim@local>
2026-08-18 00:25:32 -04:00
Tom Boucher
debeabd524 enhance(#3587): add a per-phase commit_docs override (#3601)
* feat(#3587): add a per-phase commit_docs override

Delivers epic #2292's second user story: commit an architecture phase's
artifacts while execution phases stay local. commit_docs was project-wide and
binary, so the only choices were all phases or none.

Shape is a config dynamic key phase_commit_docs.<phase-id>, following the 14
existing dynamicKeyPatterns precedents rather than inventing a PLAN.md
frontmatter spec -- which #2292 itself flags as becoming its own maintenance
surface.

Tier 1 resolves in cmdCommit, NOT in loadConfig: loadConfig has no phase
context and is called by nearly every command, so threading one through it to
serve a single caller would be a far larger blast radius for no gain. The phase
comes from detectPhaseNumberFromFiles, which cmdCommit already computes for
branch naming and which is already hardened against the #2539 project-code bug.

Suppression by the per-phase tier returns its own reason rather than reusing
skipped_commit_docs_false -- telling a user their project setting is false when
it is true would be actively misleading. Additive; the two existing reason
strings that agents/gsd-executor.md matches on are unchanged.

The manifest's phase-id pattern is a hand-copy of PHASE_NUMBER_TOKEN_SOURCE
because the manifest is hand-maintained JSON, so a behavioral parity test
asserts both surfaces accept and reject the same token shapes.

* fix(#3587): fold tests, close review findings, update reference docs

Fold: the new tests were added as their own file, which required loosening a
grandfathered lint-test-file-count bucket 5-to-6. A ratchet exists to go down
only. commit-docs-bypass.test.cjs is the established commit_docs test home and
already hosts two folded suites, so the tests fold there as a third block and
the allowlist is reverted untouched.

Standards review: CONTEXT.md and the test header both cited a
phase-commit-docs-manifest-parity.test.cjs that never existed; a repo-wide
sweep found a fourth stale cite in the schema manifest description. All four
now name the real location.

Spec review: the issue's Scope of changes named planning-config.md and
git-planning-commit.md and neither was touched. Both now document the four-tier
precedence and the new skip reason.

Security review, minor and unproven: detectPhaseNumberFromFiles returns the
FIRST matching path's phase, so a --files list spanning two phases resolves the
override against whichever comes first. That helper is hardened and widely used,
so it is not changed; the behavior is pinned by a named test and disclosed in
the design and user docs. A pinned behavior is not a bug; an unpinned surprise
is.

* chore(#3587): backfill changeset pr number to 3601

---------

Co-authored-by: sim <sim@local>
2026-08-17 21:59:40 -04:00
Tom Boucher
3ab0007164 enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19)

preserveUserArtifacts held user files only in an in-memory Map across the
wipe, so any process death between preserve and restore lost them outright.

Seven call sites, not the four the issue records. Three of them never called
the helper at all - they open-coded the same read/wipe/write - so searching
for callers under-counted by construction; the extra sites were found by
sweeping for the pattern instead.

The worst is the mainline install path, where the crash window spans the
entire gsd-core tree copy rather than a single rmSync.

Adds src/user-artifact-staging.cts: durable on-disk staging with a record
written after the copies land as the commit point, plus recovery of orphaned
batches on the next run - without recovery the staged bytes survive but the
user's file is still gone, which would pass its own test while delivering
nothing.

Routes copyPreservingSymlink through installFs() so staging cannot bypass the
install fs seam, and reunites its symlink-safety docblock with the function it
documents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-3574 with four claims disproved by implementation

Implementing Phase 6 disproved four statements the ADR rests on. The central
decision - no single materializer - is unaffected and stands.

Corrected: decision 3 was already satisfied, so nothing was extracted; the
agents-bypass runtime set omitted claude, kilo and opencode, and closing it
needed three new pieces of descriptor contract rather than proceeding on its
own terms; three of the four blockers the layout comment names were already
stale; and F19 is seven call sites, not four.

Records the generalizable lesson: the defect is the pattern of holding user
data in memory across a wipe, not the helper, so searching for callers of the
helper under-counts by construction.

Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and
notes that copyPreservingSymlink needed routing through the install fs seam
before it could be reused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close dangling-symlink blind spot and harden staging recovery

An adversarial review found the F19 staging work shipped red and unsafe.

Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween
missed dangling symlinks in both its root check and its per-segment walk,
because it probed with existsSync, which is false for a link whose target does
not exist. Fixing only the new module would have reused a guard that was
itself blind. This guard protects the whole install tree.

Recovery no longer throws: it degrades per entry and per file, so one bad
batch cannot block the others. Previously an unrecoverable entry propagated
out of the first statement of install and uninstall, before the cleanup that
would have removed it - wedging the installer permanently.

Partial fs adapters now throw on any omitted method instead of silently
reaching the real filesystem, closing the trap that let a test poison list
pass while real IO happened.

Staged names must be flat, recovery refuses a dangling destination symlink,
and a batch whose recovery genuinely failed is no longer swept - it was
discarding the only durable copy of the file it had just failed to restore.

Replaces three tests that could not fail, including the one labelled negative
proof.

Known limitation, documented not closed: concurrent installs sharing a staging
key can still lose a batch. A real fix needs a cross-process lock.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enh(#2875): make the descriptor authoritative for the agents kind

Deletes the inline agent-staging loop in bin/install.js and the
_DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from
its capability descriptor instead of an inline hostBehaviors dispatch.

Closing it needed three pieces of contract the descriptor pipeline never had,
all reducible to one missing input - per-agent resolution context: a
frontmatter-extensions step for claude's effort and disallowedTools, per-agent
model-override resolution for kilo and opencode, and a named branding
converter for hermes, whose rewrite data was already declared.

Seven runtimes were on the loop, not the six the design recorded - kimi-code
was found by a golden fixture, not by analysis. claude-local and kimi-code
both silently lost their agents mid-change; the fixtures caught both and the
cause was fixed rather than the fixtures regenerated.

A parity harness gates the migration: both pipelines over identical inputs,
byte-identical output including filenames, per runtime. It is demonstrated
red before being trusted. Surface and install paths converge for all seven,
which also fixes surface previously writing no agents for these runtimes.

Codex's config.toml strip stays put - it mutates host config, which no
descriptor kind models.

Also routes install-model-override-resolver and install-effort-resolver
through the install fs seam. Both leaked real filesystem IO from the install
call tree; the stricter adapter is what exposed them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): record the agents-descriptor migration and correct the ADR count

The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host
integration guide told readers to join a set that is gone. Replaces that with
what is now true - declare an agents entry and it installs, on the surface
path as well as install - and points anyone needing a per-agent transform at
the three extension points rather than at a new inline branch.

Corrects the ADR amendment: seven runtimes were on the inline loop, not six.
kimi-code was found by a golden fixture going red, not by reading. That is the
third short count this phase, all from enumerating by symbol or set membership
when the thing that matters is a behavior.

Adds the Changed changeset for the surface-path convergence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): amend ADR-2866 - claude global always wrote agents on disk

The claude row's global=[skills] described what capability.json declared, not
what the installer wrote. bin/install.js's inline agent-staging loop was never
scope-gated and never consulted the descriptor, so a claude --global install
has always written agents/gsd-*.md.

Phase 6 closes the gap by deleting that loop and declaring agents on claude's
descriptor at global scope. On-disk bytes are unchanged - the golden fixtures
did not move, which is the evidence that the descriptor, not the installer,
was incomplete.

#2218 is unaffected: agents are not trigger-bearing, so the wider row does not
introduce a new shadowing case.

Records the warning that an incomplete descriptor is invisible while a second
code path silently does its work, and only surfaces when the two are forced
into agreement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close review findings across staging, agents and the parity harness

Two independent reviews of this branch found defects the local gates missed.

Security: a dangling symlink at a migration destination allowed writing
outside configDir - the same class this change claimed to close, missed at the
terminal write of the flow being added. The staging-root resolver threw as the
first statement of install and uninstall, so a hostile symlink bricked both,
and symlinked-configDir users lost uninstall as well as install; it now
degrades instead of aborting. Recovery gained a source-side symlink check and
now refuses a relative destDir, which resolved against cwd. Converter dispatch
gained a runtime allowlist - lint-time validation stopped mattering once this
branch promoted that dispatch from the surface path to real installs.

Correctness: claude --local --minimal exited 1 because the minimal profile
legitimately yields zero agents and the new path treated that as a failure.
cline --local silently lost its agents - its descriptor declared none while
the deleted loop wrote them unconditionally. The agents prune was widened to
any gsd-* entry and destroyed user files it never owned.

The parity harness, on which the migration's safety argument rested, drove a
synthetic registry and never byte-compared the shipped descriptors; two of its
trap rows could not fail. It now drives the real registry across 13
runtime-scope rows including kimi-code and cline-local, and its red-proof is
demonstrated by corrupting a live capability.json. Three goldens that had
encoded the cline regression as expected behavior were corrected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): close findings from both mandated review engines

/security-review found the staging source-side walk honouring
GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write
destination. A symlinked files/ component dereferenced because
copyPreservingSymlink lstats the leaf only, so an intermediate link is
followed. The source walk no longer honours the opt-in; the destination check
still does.

/code-review spec axis found this branch had reintroduced its own bug:
migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after
the legacy dir was wiped and before the staged batch was restored, so a
planted symlink bricked uninstall permanently and orphaned the batch. Refusal
kept, abort removed.

kimi-code local silently lost its agents, the same class as the cline bug, and
the parity harness recorded that exclusion as intentional - the third test in
this branch to pin a regression as correct.

--minimal now creates an empty agents/ dir that never existed. Behaviour
restored rather than softening the changeset, so its byte-identical claim
stays true.

Standards axis: try/finally removed from twelve test bodies, fast-check
properties added for parseOwnerPid, boundary coverage at the grace window and
the ancestor-probe depth, a parity assertion for the staging-root helper
duplicated across two files, and the 8-deep config walk deduplicated.

Records 60-review.json with every finding and disposition from five passes,
including the smells left unfixed and why.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2875): prune stale agents unconditionally in minimal mode

The previous round stopped an empty agents/ directory being created when the
resolved profile yields no agents. That was implemented by skipping the agents
kind entirely, which also skipped its stale-agent prune - so a full to minimal
downgrade left stale gsd-* agents behind.

The deleted inline loop pruned unconditionally and only skipped writing. Those
are three separate conditions, not one: prune always, write only when there is
something to write, create the directory only when writing.

Both call sites now run _removeGsdEntries before the empty-staged early exit.
The symlink-escape guard moved with it, since the prune also touches dest.
Codex .toml agents and the config.toml stanzas are cleaned again, and
user-owned agents are still preserved.

The agents/ directory is left in place after a prune empties it, matching
every sibling kind - none of them remove the destination directory itself.

Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh
install, so fixture generation is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2875): document interrupted-install recovery for user-owned files

The durable-staging fix is invisible to the user it protects. Someone whose
install died mid-flight has no way to know USER-PROFILE.md was staged before
the delete, that the next run restores it, or that recovery happens at the
start of that run rather than in the background.

Written as the task the user has - finish the interrupted command - rather
than as a description of the mechanism, and states what it will not do:
overwrite a file already present, or touch staging belonging to another
install still running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2875): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2875): assert the J8 model override without building a regex

CodeQL flagged incomplete string escaping: the assertion interpolated the
override value into a RegExp while escaping only forward slashes, which is
meaningless in a constructor, leaving real metacharacters unescaped.

The failure direction was the dangerous one - a metacharacter would have made
the match more permissive, so the row would pass when it should fail. That
matters here because J8 exists precisely because an earlier revision was a
tautology; the rewrite reintroduced a different way for the same assertion to
stop discriminating.

Replaced with a line-wise exact match, so no regex is constructed at all.
Swept the other test files this branch adds; no sibling instances.

lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full
metachar-escape copy, so a single slash replace slipped under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 17:25:53 -04:00
Tom Boucher
98ecb2ba8c enhance(#2142): archive quick tasks at milestone close-out (#3592)
* test(#2142): failing-first coverage for quick-task archival at milestone close-out

* enhance(#2142): archive quick tasks at milestone close-out

* fix(#2142): resolve review findings — readme injection, move/reset ordering, owned state write

* fix(#2142): fold archival under milestone namespace, expose index IR, dedupe reset decision

* test(#2142): assert archive-dir-relative summary path in index IR

* docs(#2142): backfill changeset pr number to 3592

* test(#2142): skip newline-fixture injection test on windows (control chars illegal in path names)

---------

Co-authored-by: sim <sim@local>
2026-08-17 14:51:00 -04:00
Tom Boucher
71180983a0 fix(#3423): standardize on <required_reading>, retire the files_to_read emit tag (#3432)
* fix(#3423): standardize on required_reading, retire files_to_read emit tag

* test(#3423): flip tag assertions, extend consistency guard to spawner surfaces

* fix(#3423): sweep capabilities fragments, regen registry+skills, anchor executor test

* chore(#3423): acknowledge tag-rename emitted ripples and workflow growth

* chore(#3423): broaden emitted-ripple acknowledgment to all embedders

* chore(#3423): settle emitted-drift acks post-rebase (merge 3004/1689-owned keys)

* chore(#3423): drop stale ripple acks, ack execute-phase growth

* chore(#3423): restore pristine 3004 fragment, keep only consumed appends

* chore(#3423): backfill changeset pr number

* chore(#3423): settle emitted-drift acks post-merge (move code-review-fix ripple into 3190, tag-rename ripples into 3191/3297)

* chore(#3423): re-arm 3324 ack for execute-phase.md tag-rename ripple

* fix(#3423): trim 8 bytes from execute-phase model note to hold ADR-857 margin, re-arm 3370 ack for net +4 growth

---------

Co-authored-by: sim <sim@local>
2026-08-14 16:03:48 -04:00
Tom Boucher
895d9df96d fix(#3477): run untrusted key_links patterns on a linear-time engine (#3496)
`cmdVerifyKeyLinks` compiled `must_haves.key_links[].pattern` from plan frontmatter with `new RegExp()` and tested it against whole file contents, so a nested-quantifier pattern such as `(a+)+$` hung `verify-phase` indefinitely (CWE-1333). JavaScript has no regex-execution timeout.

Untrusted patterns now run on RE2 (re2js), whose match time is linear in input length — the class is closed by the engine, not by a heuristic screen. The screen lost in the ADR-0174 consolidation was deliberately NOT restored: it never worked, since `(a|a)*$`, `((a+))+$`, `(a+){2,}$` and `(a{1,3})+$` all evade it. A refused pattern's matcher returns false for every input, so it cannot report a match no matter what the caller does.

The engine is vendored at gsd-core/bin/lib/vendor/re2js.cjs because gsd-core/bin/** is copied into installed trees with no node_modules; runtime dependencies are unchanged. New ESLint rule local/no-external-require-in-bin enforces that invariant, which had been documented in a comment since the #3024/#2071 bug class and enforced nowhere.

Backreferences and look-around are unsupported by RE2 by construction — disclosed in a Changed changeset.

Closes #3477

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 14:34:36 -04:00
Tom Boucher
967bddba37 fix(#3384): strip mcp__* tool grants from zcode-installed subagents (#3483)
* fix(#3384): strip mcp__* tool grants from zcode-installed subagents

ZCode's dispatcher treats every mcp__<server>__* entry in an agent's
tools: frontmatter as a required MCP server and hard-fails the subagent
spawn (CONFIGURATION_ERROR) when it is not connected, whereas Claude Code
treats the same grants as an optional allowlist. ZCode shared Claude's
verbatim agents copy (converter: null), so all 8 MCP-granted agents
failed to spawn out of the box with zero MCP servers configured.

Add convertClaudeAgentToZcodeAgent — a line-surgical converter that
filters mcp__* entries out of the frontmatter tools: grant list (both
inline comma and YAML block-list shapes) and preserves every other byte.
Declare it on both of zcode's capability.json agents entries and cut
zcode over to the descriptor-driven agents path
(_DESCRIPTOR_AGENTS_RUNTIMES) so the legacy inline loop stops
deleting+re-copying the converted agents raw. Claude Code, Kimi, and
Gemini install behavior is unchanged.

* chore(#3384): link changeset fragment to pr 3483

---------

Co-authored-by: sim <sim@local>
2026-08-14 11:49:01 -04:00
Tom Boucher
9dd240f634 fix(#3302): return per-agent worktree metadata from workflow scripts (#3450)
* fix(#3302): return per-agent worktree metadata from workflow scripts

* fix(#3302): link changeset fragment to pr 3450

---------

Co-authored-by: sim <sim@local>
2026-08-14 09:28:22 -04:00
Tom Boucher
f78a4b3b80 fix(#3275): resolve reviewer lane binaries via shared pathext scan (#3445)
* fix(#3275): resolve reviewer lane binaries via shared pathext scan

* fix(#3275): pin changeset fragment to pr 3445

* fix(#3275): stage p4 resolver fixture in pathext casing

---------

Co-authored-by: sim <sim@local>
2026-08-14 02:34:29 -04:00
Tom Boucher
7976b1ca0d feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing

Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing.

- src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint)
- agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names
- gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json)
- execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling)
- execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree)
- config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator)
- docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests)

* chore(#1689): backfill changeset PR number (#3417)

* chore(#1689): regenerate install-tree fixtures for new workflow fragment

* chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing)

* test(#1689): SPAWN contract allows parameterized subagent_type placeholder

agent-frontmatter's spawn-type checks scanned subagent_type="..." as a
concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}"
(a runtime placeholder resolved via resolve-agent, default gsd-executor).
Skip {TOKEN} placeholders in both the known-type and <available_agent_types>
checks; execute-phase still lists the built-in roster incl. gsd-executor.

* fix(#1689): CI conformance for the routing fragment

- per-plan-executor-routing.md: add the canonical runtime-launcher preamble to
  its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments.
- agent-install-check.cts: drop a literal ~/.claude/agents path from the
  resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs
  (cline install leak guard).

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:10:52 -04:00
Behruz Nassre Esfahani
f0abdb1b89 fix(#2486): do not recommend or persist Claude-only worktree isolation on non-Claude runtimes (#2531)
* fix(#2486): runtime-branch the settings worktrees question + W020 health diagnostic

On non-Claude runtimes /gsd:settings offered "Yes (Recommended)" for
worktree isolation and persisted workflow.use_worktrees: true — the exact
value the execution workflows fail closed on (#1521 guards). Branch the
question on the same stamped config-get runtime read the guards use:
Claude keeps the unchanged question; non-Claude offers only
"No (Recommended)" / "Leave unchanged", never persists true, and warns
when the config carries an inherited explicit true. /gsd:health gains
W020, surfacing such a config with the guards' own predicate before
execution-time failure. Docs state the runtime-conditional default.

Fixes #2486

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2486): add changeset for PR #2531

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): reassign the health worktrees check W020 -> W024 (verify.cts namespace collision)

The workflow-level check collided with the live W020 (git-worktree-list
health) emitted by cmdValidateHealth in src/verify.cts — invisible from
health.md's error_codes table, which stops at W019 and under-represents
the real namespace (W010-W017, W020-W023 all live). W024 verified free.
Adds a regression test pinning the chosen code against src/verify.cts so
a future assignment cannot silently collide, a table note naming the
namespace owner, and the changeset body reworded to house style.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): pre-select the recommended repair in the broken-inheritance case

Review round 2: at settings.md:142 the pre-selection rule left "Leave
unchanged" as the default when the config carried an explicit
non-false use_worktrees — the exact broken state the adjacent notice
warns about, so accepting the default kept a config that fails closed
at execution time. "Leave unchanged" is now the default only when the
key is absent (nothing to repair); explicit false AND explicit
non-false both pre-select "No (Recommended)", aligning the default,
the label, and the notice.

Pinned by two source-contract assertions in the #2486 regression
block. Goldens (settings.md hash x19) + size baseline regenerated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): gate the worktrees question on dispatch.isolation, not the runtime name

Review round 2: #2584 Phase 3 replaced the runtime-name test with a
declared `dispatch.isolation` capability, invalidating this PR's
premise. cursor declares harness-worktree and codex/opencode/kimi/
kimi-code declare orchestrator-worktree, so a `RUNTIME != claude` gate
blocked a supported configuration on five runtimes and false-warned in
health.

- settings.md + health.md read `query dispatch-isolation` and branch on
  `ISOLATION = none`; the runtime-name read is gone from both, and the
  capability read needs no per-runtime stamping (it fail-closes unknown/
  undocumented internally)
- all "Claude Code-only primitive" prose rewritten, including the two
  gates the shell-syntax check missed (config-key list, JSON schema
  comment)
- W024 reconciled across health.md + CONFIGURATION.md + planning-config.md
  (docs still said W020, which collides with a verify.cts code)
- health.md error-codes table fixed: the namespace note no longer sits
  between rows orphaning I001
- the asymmetry note for the two workflows #2584 has not migrated
  (quick.md, diagnose-issues.md) is enforced by a set-equality test with
  a self-check table, so it cannot go stale in either direction

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(#2486): restore the Executor isolation section clobbered by #2661

`46ba02ac` (feat(#2630), the current next tip) reverted docs/CONFIGURATION.md
to a pre-#2584 state: it restored the old "Non-Claude note" wording on the
workflow.use_worktrees row and deleted the whole "Executor isolation per
runtime" section. The change is unrelated to that PR's phase-estimation
feature and looks like a stale-copy edit.

This PR's use_worktrees row links to #executor-isolation-per-runtime, so the
deletion leaves a dangling anchor. Restored byte-for-byte from a40ee8a5 (the
text #2584 Phase 3 originally shipped). No link/anchor checker exists in
scripts/ or tests/, so CI would not have caught the dead link.

The reverted row wording is outside this PR's scope and is still wrong on
next; reported upstream as #2668.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2486): acknowledge the emitted size growth for settings.md and health.md

The #2723 differential attribution check (epic #2719 Phase 3) landed on next
after this branch opened and gates emitted-file growth behind a committed
acknowledgment. Both grown files are attributable to source this PR changes;
the ack names them and says why, per ADR-2719 §3.

Verified load-bearing: removing tests/emitted-drift-ack.json reproduces the
same failure; restoring it passes 59/59.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2486): reword the W024 remediation so it names no bare gsd-tools call

The W024 warning told the user to run `gsd-tools query config-set`, a
command-position bare invocation that fails with "command not found" on a
shim-only install (#2751) — the diagnostic sent the user into a second,
more confusing error than the one it reported. The remediation now names
only `/gsd:settings` and the config key itself, both of which work on
every install layout, and the #2751 command-position gate goes green.

Fixes #2486

* fix(#2486): drop the Known-asymmetry note and its guard test — #2728 migrates both workflows

Review Major 2: once #2728 lands, the note describes a gap that no longer
exists and instructs maintainers not to do the thing that was just done —
with the set-equality guard test pinning the stale prose green. This PR now
depends on #2728 (declared in the PR body), so the note and its guard go.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(#2486): make the #2728 merge order enforceable instead of advisory

settings.md recommends and persists `workflow.use_worktrees: true` for every
runtime whose declared dispatch.isolation is not `none` — cursor
(harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree).
quick.md and diagnose-issues.md consume that value at dispatch time and, while
they still gate on the runtime NAME, FATAL for all five. Merged first, this PR
reintroduces #2486's own shape for the exact runtimes it exists to help — on
the path it now labels "Recommended".

The dependency was stated only as prose in the PR body. A merge-order note is
not a gate: an automated batch merge never reads it. This adds the check that
makes the ordering structural — red while either sibling is still name-gated,
green the moment #2728 lands.

The predicate matches a RUNTIME-vs-"claude" comparison, not the legitimate
runtime-identity read, and accepts either the inline canonical
dispatch-isolation read or a reference to dispatch-isolation-gate.md, which is
the shape #2728 gives quick.md. Verified both directions against real sources:
red against the current workflows, green against #2728's.

This also closes W024's coverage window. W024 fires only when ISOLATION is
`none`, so it is structurally blind to these five runtimes — /gsd:health would
report healthy right up until quick.md FATALs. W024 cannot see the hazard, so
the hazard is prevented by making the unsafe ordering unmergeable rather than
by warning after the fact.

Separately, the W-code namespace-collision test now declares its source read
instead of passing lint silently: verify.cts emits its codes as inline string
literals across ~25 addIssue() calls and exports no enumerable registry, so
there is nothing to require() and assert against. The gap, and what it costs,
are stated in the test.

* docs(#2486): use the hyphen command form for /gsd-health in CONFIGURATION.md

docs/ is never passed through the install-time slash-form converters, so the
colon form names a command no runtime registers. Clears the sole
lint-docs-command-form violation attributable to this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2486): resolve isolation via new side-effect-free inspect-dispatch-isolation query; scope the settings change to isolation-none runtimes

Review round 4:

- B1: /gsd:health and /gsd:settings no longer call the recording
  dispatch-isolation query — on current next it persists the resolved
  decision to the executor-isolation sentinel as an unconditional #3045
  side effect, letting a read-only diagnostic hard-block executor
  dispatch for the sentinel's lifetime across sessions. Both surfaces now
  use inspect-dispatch-isolation, a new read-only verb sharing the exact
  resolution implementation (extracted resolveDispatchIsolationDecision)
  with zero writes. Behavioral tests pin: no sentinel write, per-runtime
  parity with the recording verb, recording knobs ignored, --json shape.

- B2: the #2728 merge-order interlock test is deleted — a repo test
  cannot sequence merges; it only made this PR unmergeable on its own
  schedule. The settings behavior change is scoped entirely to the
  ISOLATION=none branch, which needs nothing from #2728; the != none
  path is base behavior unchanged.

- M1: the W024-vs-verify.cts namespace test (an admitted source-grep) is
  deleted per RULESET.TESTS.delete-bad-tests, without a standing
  exemption; the namespace claim lives as guidance in health.md.

- M2: remaining allow-test-rule exemptions re-derived into documented
  categories (source-text-is-the-product, integration-test-input) with
  issue refs per ADR-456.

- Minor: settings.md's current-value read drops the stampable
  --default/fallback so key-absence stays distinguishable from an
  explicit false on non-Claude emits (the pre-selection rule depends on
  the tri-state); docs now present inspect-dispatch-isolation as the
  inspection command and name dispatch-isolation as the recording
  resolver.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(#2486): put the install-marker rung in the canonical runtime resolver

Round-7 review (independent cross-AI pass over the whole PR).

BLOCKER — the gate did not fire in the default case. `resolveRuntime`
stopped at GSD_RUNTIME > config.runtime > 'claude', and
`config-new-project` writes NO `runtime` key — so on a real non-Claude
install every consumer believed it was on Claude: isolation reported
`harness-worktree`, /gsd:settings still offered "Yes (Recommended)" and
W024 stayed silent. That is #2486's own symptom, surviving the fix meant
to remove it. Verified end-to-end on a real `--qwen` install with a
runtime-neutral config and GSD_RUNTIME unset: pre-fix `harness-worktree`,
fixed `none`.

The rung lives in `resolveRuntime` (src/runtime-slash.cts), the ONE
canonical resolver, not in the isolation call site. A per-consumer fix
forks precedence: `inspect-dispatch-isolation` would answer `cursor` while
`dispatch-should-flatten` and `resolve-dispatch-type` still answered
`claude` for the same install. All three now agree. `readInstallRuntimeMarker`,
its cache and its test seams MOVED from model-resolver.cts to
runtime-slash.cts, with model-resolver re-exporting the seams — one marker
read and one cache, not two that drift.

Deliberately NOT `resolveActiveRuntime`/`loadConfig`: an intermediate
revision of this fix routed through `loadConfig`, which normalizes and
rewrites legacy keys back to disk. That gave `inspect-dispatch-isolation`
— the verb whose entire purpose is being side-effect-free — a write side
effect, which is the defect the verb exists to avoid. `resolveRuntime`
reads .planning/config.json directly. Re-verified: the inspect query
against a real install creates no `.gsd/`.

Tests, both of which were too weak in the first attempt and are now
fail-first proven:
- The W024 behavioral stub returned success-with-empty-output for an absent
  key. Real `config-get` EXITS NON-ZERO, and the `|| echo "true"` fallback
  only triggers on failure — so reintroducing the fallback would have passed.
  The stub now returns 1 for the absent case.
- The marker regression asserted on the exported helper, so reverting
  gsd-tools.cjs to a marker-blind resolver still passed. It now drives a real
  `--qwen` install through the shipped `inspect-dispatch-isolation` query.
  (GSD_TEST_MODE must be cleared for that child, or install.js no-ops while
  still exiting 0 — a green test over an install that wrote nothing.)
  Spawned through `installSpawnEnv()` so ambient GSD_HOME cannot leak in.

Also: the three PR-added `try/finally` test bodies converted to `t.after()`
per CONTRIBUTING; CONTEXT.md's `worktree create` UNCONSUMED claim corrected
(Phase 3 calls it in executor-isolation-dispatch.md); docs/CONFIGURATION.md
now states the current non-Claude `use_worktrees` default rather than
describing capability-scoped stamping that lands with #2652; resolver
precedence comments updated to name the marker rung.

Refs #2486
Refs #2668

* fix(#2486): split the marker rung out; make the use_worktrees doc row order-independent

Round-8 review. trek-e's Blocker was procedural — commit 10fba7b8 moved the
per-install .gsd-runtime marker rung into the canonical resolver, a large
blast-radius change that arrived undisclosed and unreviewed. They offered
two remedies; taking the second: SPLIT IT OUT.

src/runtime-slash.cts and src/model-resolver.cts are reverted to their next
state and the marker regression test is removed. What remains is what this
PR was filed for: the settings.md / health.md / gsd-tools.cjs isolation
query, W024, and the #2668 docs restoration.

COST, stated plainly rather than buried: without that rung this PR's gate
resolves 'claude' on a non-Claude install whose project config carries no
`runtime` key — which is every config config-new-project writes. On those
installs W024 stays quiet and /gsd-settings still offers Worktrees. The gate
is correct whenever the runtime IS resolvable (GSD_RUNTIME set, or an
explicit config.runtime). That is why `Fixes #2486` is already downgraded to
`Refs` — #2486 must not close until the residual lands. A KNOWN LIMITATION
comment at the resolver call site names #2395 so this does not read as an
oversight.

The rung itself belongs to #2395, which reports this exact defect and was
closed by #2446 — a PR that touched only bin/install.js and fixtures and
never runtime-slash.cts, so it persisted the identity into
~/.gsd/defaults.json, a tier resolveRuntime does not read. Evidence and a
reopen request are posted there. That same tier is #2566's B1.

DOCS — the reviewer flagged that this row and #2728's are order-dependent:
whichever merges second falsifies the other. Removed the dependency instead
of picking an order. The "Current default … until #2652" paragraph is now a
plain troubleshooting note, true before and after #2728 lands.

CONFLICT — one hunk in CONTEXT.md: next added the #2596 scope-conformance
interface to the same Worktree-Safety paragraph where this PR corrected the
`worktree create` UNCONSUMED claim. Resolved keeping both.

Verified: lint:ci green. Full suite clean apart from the pre-existing
#1160 installed-runtime capability surface. (emitted-attribution also failed
until the fork's stale next was fast-forwarded — it defaults to origin/next,
which was 131 commits behind; 175/175 against the current base.)

Refs #2486
Refs #2668

* fix(#2486): address round-9 review — distinguish "cannot resolve" from "declares none", reject recording-only args on the read verb

Codex review of the whole PR, five findings, all verified against source first.

Major 3 (the one real defect). Both surfaces read isolation as
`ISOLATION=$(… || echo "none")`, so a resolver failure became indistinguishable
from a genuine `dispatch.isolation: none` declaration — and W024's text then
asserts the latter, telling a user their runtime declares no executor-isolation
primitive when GSD simply could not find out. health.md and settings.md now
capture the raw value and track ISOLATION_RESOLVED, the same shape
references/dispatch-isolation-gate.md already uses (#2652 review). W024 gained a
second message for the unresolved case; settings still fails closed there — it
must never persist a `true` it cannot justify — but reports what happened rather
than a verdict it never reached.

Major 4. `dispatch-isolation` applies --force-isolation AFTER the shared
resolver returns; inspection accepted the flag and silently ignored it, so the
same argv yielded 'none' from one verb and the declared capability from the
other. inspect-dispatch-isolation now rejects --force-isolation/--phase/--plan
as usage errors. Verified by mutation: disabling the guard reds exactly the
three new rejection tests.

Major 1. settings.md claimed this flow and the execution guards "always reach
the same verdict". False while quick.md still gates on the runtime name: an
orchestrator-worktree host can be offered a `true` that /gsd:quick rejects as
fatal, W024 silent because isolation is not none. Scoped the claim to the
capability gate, named #2728 as the conversion, and added the same caveat to the
CONFIGURATION.md use_worktrees row. Not a behavior change — the `!= none` branch
is untouched base behavior.

Major 2. "side-effect-free" overstated it: every gsd-tools invocation runs the
shared bootstrap, and getActiveWorkstream unlinks a stale workstream pointer.
That is pre-existing, verb-independent and cannot block a dispatch; writing the
sentinel can. Renamed the claim to "sentinel-free" everywhere and documented
precisely what is and is not asserted.

Minor 5. The JSON parity test asserted against a handwritten key list and never
invoked the recording verb, so the two could diverge and stay green. It now runs
`dispatch-isolation --json` in a separate project dir and deep-compares, with a
control asserting that verb DID write a sentinel. Added the orchestrator exec
branch (--cwd-target) the registry parity test never covered.

runtime-converters.test.cjs pinned the old `|| echo "none"` line as canonical —
that literal was the Major 3 defect. Repinned as two invariants (the raw read is
present; the collapsing fallback is absent) plus an ISOLATION_RESOLVED
requirement, so reformatting does not fail the test but a semantic regression
does.

settings.md sits at 40777 bytes against the 40960 DEFAULT hard cap. The
round-9 prose was compressed to fit rather than raising the cap.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d (#1160 _resolveManifest, and
the #3053 quick_id tests, which compare a local-time expectation against a
TZ=UTC child and so only pass on a UTC host).

* fix(#2486): second review pass — drop a dangling reference, correct the whole unresolved branch, pin the branch behaviorally

Codex re-review of the full PR after the first round-9 pass. It confirmed Majors
1, 2 and 4 and Minor 5 fixed, and found seven more. All verified against source.

Minor 5 was mine and the worst of them: settings.md and health.md pointed at
`gsd-core/references/dispatch-isolation-gate.md`, which exists only on the #2728
branch — not in this PR and not on the merge-base. Merging this alone would have
shipped three dangling canonical-source references. Removed; the surrounding
text now stands on its own.

Major 2. The unresolved branch corrected only the two option descriptions. The
pre-selection rationale still claimed an absent key already resolves to false on
this runtime, and the explicit-true notice still stated the runtime declares no
primitive — both capability verdicts that were never reached. The substitution
rule now covers every place in that branch that asserts one.

Major 3. The repinned invariants could not catch the mutation they existed to
catch: flipping the shipped block's ISOLATION_RESOLVED=true to false left all
three green, since they only assert the raw read is present, the token appears,
and the collapsing fallback is gone. Added a behavioral test that drives the
shipped W024 block twice — resolver answering vs resolver exiting non-zero — and
asserts the two emit different text, that the unresolved branch says "could not
resolve", and that it does NOT claim the runtime has no primitive. Verified: the
true->false mutation now reds it.

Major 1. The verb fail-closes an unknown or undeclared runtime to `none` and
exits 0, so ISOLATION_RESOLVED=true means "the query answered", not "the runtime
published a declaration" — the shell cannot see an internal fallback. Exposing
provenance is an API change and out of scope here, so W024's resolved-case text
is instead written to be true of every path that reaches it ("no usable
executor-isolation primitive — declared or fail-closed from an unknown value"),
and both workflows state the limit of the signal explicitly.

Minor 4. The rejection tests regex-matched human prose, so swapping
ERROR_REASON.USAGE for UNKNOWN would have kept them green while breaking machine
consumers. They now pass --json-errors and assert reason === 'usage' on the
parsed envelope.

Minor 6. CONTEXT.md still described the verb as side-effect-free. Now
sentinel-free, consistent with the router comment and the other two surfaces.

Minor 7. 183 bytes of headroom under the 40960 hard cap was called
unacceptable, and it was — a routine 184-byte edit would have failed CI.
Compressed this PR's own settings.md prose (the #2486 block was carrying ~7.1KB
of rationale) to 40499 bytes, 458 free. Two phrases other tests pin verbatim
were restored after the first compression pass reworded them.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): third review pass — placeholder table replaces the fragile substitution rule

Codex round 3 blocked the push on two Majors. Both verified and real.

Major 2 was the substantive one and my error. The round-2 fix told the model to
find and replace the string "this runtime declares no executor-isolation
primitive (dispatch.isolation: none)" — text that does not appear anywhere in
the branch. The first option description has no parenthetical, the second
asserts "absent, it already resolves to false on this runtime" (not covered at
all), and the pre-selection rationale and notice carried three more unsupported
assertions. A model following the instruction literally would have found no
match and changed nothing.

Replaced with three named placeholders — {FINDING}, {CONSEQUENCE}, {ABSENCE} —
and a two-column table giving each one's resolved and unresolved wording. Every
assertion in the branch now flows from the table, so none of them can outrun
what was actually established, and there is no string-matching to get wrong. It
is also shorter than the prose it replaced.

Major 1: the resolver catches thrown errors and returns none successfully, a
path the round-2 wording ("declared, or unknown/undeclared") did not cover.
Both surfaces now say "declared as none, or fail-closed because the capability
could not be determined", which is true of the thrown path too, and both state
that an internal resolution error is among the things the verb fail-closes.
Still not provenance — the verb cannot distinguish these for the caller — but no
longer a claim the code contradicts.

Minor 3: the behavioral test asserted only that the two branches differ and that
the resolved one omits "could not resolve", so arbitrary resolved text stayed
green. Now pins what it must positively say: the capability finding, the
fail-closed consequence, the offending key, and (both branches) the repair
command.

Minor 4: two stale artifacts the round-2 sweep missed — a test comment still
citing the #2728-only references/dispatch-isolation-gate.md, and the emitted
drift ack still describing inspection as side-effect-free. Also aligned the two
remaining router comments to sentinel-free.

settings.md 40420 bytes, 540 free under the cap.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): fourth review pass — stop recommending a repair whose effect depends on the emit

Codex round 4, one Major and three Minors. The Major is a real hole and worth
the round.

W024 advised "remove the key from .planning/config.json so the runtime default
(false) applies". That default is not false everywhere: execute-phase.md:102,
quick.md:155 and diagnose-issues.md:62 all read the key as
`|| echo "true"`, and only `_stampNonClaudeRuntimeDefaults` rewrites them to
`--default false` on a non-Claude emit. So on an un-stamped emit the advised
repair leaves use_worktrees resolving to TRUE against isolation none — exactly
the state W024 exists to flag. Reachable in the round-3 gap: a caught internal
resolver error returns none with rc=0, so ISOLATION_RESOLVED is true and the
confident branch fires on a host whose emit was never stamped.

Fixed by removing the dependence rather than the symptom: both surfaces now say
to set workflow.use_worktrees false explicitly, and say why deleting the key is
not equivalent. The {ABSENCE} placeholder no longer claims absence resolves to
false unconditionally — it is scoped to an emit that stamped that default.

Minor: the notice and W024 hardcoded .planning/config.json, which is the wrong
file in an active workstream (settings itself resolves $GSD_CONFIG_PATH
correctly). Both now say "the project config", naming the workstream case.

Minor: the new repair pin required the literal /gsd:settings, a form documented
as no longer routable. Relaxed to accept the canonical slash and $-prefixed
forms so a correct rewording cannot fail the test.

Nit: two test descriptions still said side-effect-free.

Not fixed, deliberately: the verb still cannot tell a caught resolver error from
a declared none, so ISOLATION_RESOLVED remains a signal about the CALL, not the
declaration. Both workflows now state that limit outright. Exposing provenance
is a change to the query contract that belongs with the runtime-resolution work
in #2395, not here.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d. The only change after that
suite run was removing a redundant regex escape flagged by eslint; that test was
re-run individually and eslint is clean.

* fix(#2486): fifth review pass — stop recommending "leave it absent" as the safe default

Codex round 5, one Major: the round-4 fix corrected the {ABSENCE} wording but
left the pre-selection rule built on the assumption it had just falsified. An
absent key still pre-selected "Leave unchanged" on the rationale that there was
nothing to repair. Absence resolves to false only where the emit stamped that
default; on an un-stamped emit it reads as true, so accepting the recommended
default could preserve the exact isolation-none-plus-true state this branch
exists to prevent.

The isolation-none branch now pre-selects "No (Recommended)" in every case,
including an absent key. Writing an explicit false is correct under either emit;
"Leave unchanged" stays available for the deliberate shared-config case but is
never the default. The test that pinned the old rule pinned a falsified
assumption, so it now pins the new one and asserts the old sentence is gone.

Also tightened the round-4 repair regex, which had been relaxed far enough to
accept `$gsd:settings` and `/gsd:settings-bogus`. It now matches only the
canonical `/gsd-settings`, `/gsd:settings` and `$gsd-settings` forms.

settings.md 40685 bytes, 275 free.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next @ 33fca50d.

* fix(#2486): make the inspection exemption block-scoped, and drop the last pre-#2728 claim

Codex review of the conflict resolution: one Major, two Minors, all real.

Major. The inspection-surface exemption I added to #2728's "every dispatch-site
degrade block re-records" guard was FILE-wide. Codex probed it by adding an
unrecorded executor block to health.md: the exemption assertions still passed,
the whole file was skipped, and the offender went unreported. That is the same
hole the hand-listed revision of this guard had — a promise of "every dispatch
site" with a silent carve-out.

Now scoped per block. A block earns the exemption only by resolving through
`inspect-dispatch-isolation` AND carrying no dispatch primitive (`Agent(`,
`harnessFlag`/`HARNESS_FLAG`, `isolation="worktree"`, or the recording verb).
The file-level assertions remain as a precondition on top. Verified with Codex's
own probe: the appended dispatch block is now reported at health.md:285, while
the legitimate inspection block still passes.

Minor. settings.md still carried the caveat saying `quick.md` gates on the
runtime name and "#2728 converts that surface". #2728 has landed. Same obsolete
premise already removed from docs/CONFIGURATION.md in the merge commit; this was
the copy I missed. Removing it also returns 413 bytes, so settings.md now has
688 free under the 40960 cap rather than 275.

Minor. Four scratch-directory prefixes still read `w024` after the renumber.
Non-functional, but the whole point of the rename was that W024 now means
someone else's warning.

Validated: lint:ci clean; full suite green except the two failures that
reproduce identically on pristine next.

* fix(#2486): name the diagnostics' state INSPECTED_ISOLATION and delete the exemption

Codex probed the per-block exemption and it was still escapable: DISPATCH_PRIMITIVE
matched only `Agent(`, two harness-flag spellings, `isolation="worktree"` and the
recording verb, so `Task(`, `spawn_agent(`, a `codex exec` process dispatch, and —
unfixably by any regex — the repo's normal split shape (isolation resolved in one
bash fence, dispatch in a later one) all still qualified as exempt.

Took Codex's suggested fix, which is better than the one it replaces. The
diagnostics now name their state `INSPECTED_ISOLATION` / `INSPECTED_RESOLVED` /
`_INSPECTED_RAW` instead of reusing the dispatch-site names. #2728's guard scans
for a literal `ISOLATION=none`, so health.md and settings.md fall outside it BY
CONSTRUCTION and the exemption is deleted outright — no carve-out to escape, and
a diagnostic that ever writes a real `ISOLATION=none` is caught like any other
site. The name is also just more accurate: an inspection result is not a dispatch
decision.

Pinned so the reasoning cannot be lost: the #2486 suite now asserts neither
diagnostic assigns the bare `ISOLATION` name, with the rationale in the comment.
Verified by mutation — renaming back makes #2728's guard flag health.md:107 AND
fails the new naming pin.

Net effect on the guard's coverage is positive: before this PR it scanned two
fewer files by exemption; now it scans everything.

settings.md 40352 bytes, 608 free.

Validated: lint:ci clean; full suite green except the two failures that reproduce
identically on pristine next.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 20:45:08 -04:00
0xdhx
0396d9cab1 enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection

The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.

That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.

Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.

review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.

* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines

Two CI failures from the first push, both mine:

1. lint-tests: the new regression test split readFileSync content on a
   literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
   "\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).

2. golden-install-parity / workflow-size-budget / workflow-compat: editing
   gsd-core/workflows/review.md changes its content hash and byte size, and
   both are pinned in committed baselines. Regenerated via the repo's own
   generators (npm run size:baseline, npm run gen:golden).

The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.

Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).

* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape

The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.

* enhance(#2483): also guard the claude leg against auto-memory injection

CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).

* enhance(#2483): match the claude binary in command position, not argument position

The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as 920a5f3f)
reshaped the effort-args lookup from

  --host claude 2>/dev/null | jq -r '.effort_argv_string // ""'

to

  --host claude --pick effort_argv_string

which put a flag immediately after `claude` and made the config query
read as a third claude dispatch, failing the count assertion.

The defect class is a binary name in *argument* position being read as a
command. Fixed at the class rather than the instance: tokenise the line
and skip any `claude` whose preceding token is a flag. That also covers
the latent sibling one line away in review.md (`command -v claude`),
which escaped today only because its next token is a redirect.

Negative-controlled four ways: stripping CLAUDE_CODE_DISABLE_AUTO_MEMORY=1
fails, stripping the whole env guard fails, adding a genuine third
unguarded dispatch (`timeout 900 claude --output-format text -p -`) still
fails — so the narrowing did not blind the matcher to reshapes, which is
the property the count assertion exists for — and the pre-#2589 jq form of
the lookup still passes, so the matcher is not pinned to today's base.

* enhance(#2483): carry the claude reviewer's memory guard as declared lane data

ADR-2782 Phase 5b replaced the hand-authored per-CLI dispatch legs in
review.md with the declared lane table, so the two `env`-prefixed shell
lines this PR previously added no longer have a surface to live on. The
guard is reimplemented where the lane contract now lives.

`SpawnInvoke` gains an optional `env`, the claude lane declares the pair,
the resolver folds own string-valued entries into `SpawnPlan.env` (absent
or empty resolves to null, so the runner has one shape to test), and the
runner passes it to spawn. Production merges it OVER `process.env` into a
fresh object for that one child, so nothing reaches the orchestrating
session or any other lane in the run.

Declared data rather than a handler (D6): the pairs are static per lane,
which is precisely what the manifest vocabulary is for. The capability
manifest carries the same field, because the lane-fidelity test compares
manifest and descriptor over the union of `invoke`'s keys.

The regression test is rewritten against the resolver and runner rather
than review.md's text. It gains the property the source-text assertions
could only approximate: that `process.env` is never mutated.

Scope boundary, asserted rather than left in prose: `env` is not part of
the trust-disclosure surface, which is safe only while no manifest body
reaches the resolver — the registry's reviewer bodies contribute slugs to
the parity check and execution resolves from `REVIEWER_LANES`. The new
test fails first if that ever changes.

* enhance(#2483): restate the guard's mechanism in the docs and changeset

Both described the fix as two `env`-prefixed dispatch lines, which is the
surface ADR-2782 Phase 5b removed. The user-visible behaviour is
unchanged; the carrier is not, and a changeset that ships a description
of a mechanism the tree does not have is a CHANGELOG entry nobody can
verify against the code.

* enhance(#2483): cover the production spawn wiring end to end

The unit tests stop at the runner's `deps.spawn` seam — every one injects
a spy. Production supplies that seam in `gsd-core/bin/gsd-tools.cjs` as a
hand-written object no test constructs, so the chain could be correct all
the way to `SpawnPlan.env` and the merge could still be wrong or absent
with the suite green. Deleting those four lines was the one mutation that
left every other control silent.

This runs the real `spawnSync` through `gsd-tools review-lane invoke`,
with a `claude` shim on PATH that records the environment it was handed.
It asserts both halves in one test: the pair arrives, and an unrelated
inherited variable survives — a wiring that REPLACED the environment
rather than merging over it would satisfy the first and break every
lane's PATH and HOME.

POSIX-only; mediating a Windows `.cmd` shim is a separate concern the
repo already tests on its own.

Noted rather than fixed: `timeout`, `killSignal`, `maxBuffer` and
`shell: false` on that same object are equally uncovered. That is the
epic's gap, not this change's, and closing it is not in scope here.

* enhance(#2483): validate the invoke.env shape and register it as spawn-only

`env` was the one spawn-invoke field with no shape enforcement: every sibling in
`validateSpawnInvoke` is checked, and a manifest declaring `env` as an array, a
string, a number, or an object with non-string values passed validation in
silence. That matters more than an ordinary schema gap here, because
`resolveLanePlan` DROPS a non-string value rather than coercing it — so an
unvalidated manifest declares a pair that never reaches the spawn, which is the
failure a memory guard can least afford.

Two registrations, not one. `env` was also absent from
`SPAWN_ONLY_INVOKE_FIELDS`, which is the list the openai-http arm rejects
against — so `invoke.env` was accepted on a transport that issues an HTTP POST
and has no child environment at all. It was the only spawn-shaped field accepted
there; the other six each produce two errors. Self-found while sweeping the
class, not raised in review.

Keys are held to the portable POSIX environment-name grammar. That is a policy,
not a claim about what an environment can hold: measured, only NUL is actually
rejected by `spawnSync`, while `=`, a leading digit, a dash and a space are all
carried through to the child (an `A=B` key arrives as the raw entry `A=B=value`).
They are refused because a name outside the grammar is not portably addressable
by the program meant to read it.

`__proto__` is refused for a different and concrete reason. It passes that
grammar and is a real own key once a manifest is JSON-parsed, but assigning it
onto a plain accumulator goes through the inherited `__proto__` setter rather
than creating an own property — and for the string values this field permits the
setter is a no-op that does not even change the prototype. The pair would
validate and then simply vanish before the spawn. (An environment CAN carry a
literal `__proto__` entry; this is about the resolver's accumulator, and the
error message says so.)

Deliberately narrower than the sibling reserved-name guards in this file, which
also reject `constructor`/`prototype`: those guard bracket lookups that resolve
prototype members, whereas this reads via `Object.keys` plus an own-value read,
where `constructor` assigns as an ordinary key the spawn could carry.

`effortChannel` is deliberately left in neither field list: ADR-2782 D2 defines
it for both transports, so it is shared rather than spawn-only.

Reversion-controlled, three mutations, all three fire a named test: dropping
`env` from the discriminator fails `httpTransportRejectsEnv`; removing the
`__proto__` arm fails `envRejectsProtoKeyThatWouldSilentlyVanish`; disabling
the block fails four.

(#2483)

* enhance(#2483): amend ADR-2782 D2 for the invoke.env vocabulary widening

D2 records the spawn `invoke` shape as a closed vocabulary, and its Amendments
section carries a dated entry for every prior widening (Phase 1 #2794, Phase 2
corrections #2795, Phase 5b #2799). This change extended that vocabulary in code
without touching the ADR governing it, so the ADR contradicted the
implementation — and the repo's own convention, recorded in CONTEXT.md, is that
the ADR is amended in the same PR precisely because the prior widenings did it
correctly.

Adds the `invoke.env` row to the D2 table and a dated Amendments entry.

The entry also corrects the authority this change cited. The source comment
pointed at D6, which governs the closed `handler` enum — imperative behavior
admitted first-party — and says nothing about the `invoke` field vocabulary.
That is D2's territory, so the citation never covered the gap.

Two claims are corrected rather than restated, both about the trust boundary
that justifies leaving `env` out of the D5 disclosure signature:

- The regression test does not enforce that boundary. On one forged lane it
  shows the resolver folds whatever it is handed, so a future path feeding it
  manifest lanes would not make any assertion in that test fail. Its comment
  claimed it "will fail first"; that was wrong, and both the comment and the
  ADR now say the boundary is a property of the production call chain instead.
- The ADR is internally inconsistent on whether third-party manifest lanes
  execute at all: Consequences says adding a reviewer needs "no core patch",
  while `gsd-tools.cjs` rejects every slug absent from the first-party
  REVIEWER_LANES map. CONTEXT.md, `workflows/review.md` and the resolver's own
  header take the first view. #2483 did not create that inconsistency and does
  not resolve it; the entry records it rather than settling it in its own favour.

(#2483)

* enhance(#2483): document invoke.env in the capability-manifest reference

ADR-2782 points capability and plugin authors at
`docs/reference/capability-manifest.md` as where the lane vocabulary must be
visible, and its `invoke` row enumerates the spawn sub-shape field by field.
`env` was absent from that table while being part of the real shape, so the one
document a third-party capability author would actually consult to learn the
field exists did not mention it.

Squarely Diataxis reference material — a field-by-field schema description — so
it goes here rather than in the user-facing prose, which was already updated.
States the constraints a manifest author can actually trip, and is explicit that
the name grammar is a portability policy rather than an OS limit, so a reader
does not take it for a claim about what an environment can hold.

(#2483)

* enhance(#2483): disclose and sign the reviewer lane's env and residual invoke fields

`invoke.env` was undisclosed at install time. That was defensible while manifest
lanes could not execute — the premise this PR's own ADR amendment recorded — and
#2927/#3062 retired it: `routeReviewLane` now merges installed overlay `reviewer`
bodies into its lane map via `mergeReviewerLanes`, which is a field-identical merge
by ADR-2782 D1 and deliberately does not deep-validate. An overlay's whole `invoke`
therefore reaches `resolveLanePlan`, and `env` reaches the spawned child. A consented
third-party capability could set `NODE_OPTIONS=--require ./evil.js` on a reviewer lane
with no install-time disclosure and no re-consent.

The same file already decided what `env` means in a manifest: MCP servers fold it into
the disclosure signature and render each key and value in the consent prompt, with an
inline rationale naming this exact shape. Reviewer lanes get the identical treatment.

`env` was the ninth unsigned invoke field, not the first. `defaultHost` (the manifest's
OWN fallback egress host, used whenever the config key resolves to nothing),
`path`, `outputChannel`/`outputArg`, `modelArg`, `effortChannel` and `modelDiscovery`
all reach `resolveLanePlan` and none was bound. Enumerating a ninth name leaves the
tenth open, so the lane signature carries a RESIDUAL of every other declared `invoke`
key — the completeness backstop `rawConfig` already gives the MCP line (#1459 finding 5),
and the "sign the whole object" remedy the recorded decision on this class prefers.

`defaultHost` is also rendered: `resolvedHost` comes from user config, so a lane whose
key is unset displayed "(unresolved …)" — which reads as "no destination" — while the
runtime egresses the plan and review text to the address the manifest picked.

D4.5 is preserved one level down: the extra element is appended ONLY when the lane
declares something beyond the eight already-bound fields, so an env-free lane's
signature stays byte-identical and no already-consented capability is re-prompted for
a field it does not use. A lane that does declare one re-consents, which is the point.

Execution-primitive env names are FLAGGED in the prompt, not refused in the validator.
A denylist cannot be the boundary here: `PATH` alone is a complete execution primitive
for a spawn lane and can never be refused, the child is an arbitrary third-party binary
so the true set spans every interpreter's injection vars, and the MCP `env` this mirrors
refuses nothing and discloses everything. Missing a name costs a quieter line, never a
boundary.

Refs #2483.

* enhance(#2483): exercise the real overlay merge path in the guard test

The test named for the manifest/first-party boundary did not test it. It built a
forged lane locally, handed it straight to `resolveLanePlan`, and asserted that
`REVIEWER_LANES` did not contain it — so no assertion in it depended on the claim its
name made, and a code path that fed manifest lanes to the resolver would not have made
it fail. Its own comment said as much, and named the production chain as the real
carrier of the guarantee: "gsd-tools.cjs builds its lane map solely from REVIEWER_LANES".

That sentence is now false. #3062 merged overlay reviewer bodies into that map, so the
test's premise and its subject both moved.

The replacement routes through `mergeReviewerLanes` — the real helper the production
path calls — and asserts the overlay lane is admitted, resolves, and carries its `env`
into `SpawnPlan.env`. That makes the security property falsifiable instead of narrated.
It then asserts what now backs it: the env is disclosed on the surface, rendered key
and value in the consent prompt, flagged when the name is an execution primitive, and
bound to the signature so a value change, an addition, or a removal each force
re-consent.

Three further cases, because the finding's generative half is what stops it recurring:
the residual backstop is asserted against five fields including one that does not exist
(`aFieldThatDoesNotExistYet`), so a future vocabulary widening cannot silently re-open
this; a fully-enumerated lane is pinned to its original 8-tuple, which is what keeps the
fix from re-prompting every consented capability; and an http lane's manifest-declared
`defaultHost` is asserted to reach both the prompt and the signature.

Reversion-controlled, seven mutations, all seven fail a named test: env dropped from the
surface, the prompt's env line removed, the execution-primitive warning removed, the
signature's extra element never appended, the residual emptied, the defaultHost line
removed, and the declares-something test un-widened. The last of those was SILENT on its
first run and its test was written in response, then the control re-run.

Refs #2483.

* enhance(#2483): correct the ADR amendment's manifest-lane premise

The amendment argued `env` needed no D5 disclosure because a manifest's `invoke`
fields never reach `resolveLanePlan`. That was true when written and #3062 retired it
22 hours after this branch's last commit: `routeReviewLane` now builds its lane map
from `mergeReviewerLanes(REVIEWER_LANES, loadRegistry({includeInstalled: true}))`, and
D1's no-translation-layer rule makes that a field-identical merge, so an overlay's
whole `invoke` reaches the resolver and executes.

The entry had named this exact trigger — "were manifest lanes ever made executable,
`env` must join the disclosed surface in that change, and nothing here will trip if it
does not." Nothing tripped. The premise is rewritten to current truth rather than
annotated, because an ADR is read in fragments and a superseded paragraph left standing
reads as live reasoning to the next author; a one-line dated tombstone points at git for
the withdrawn text.

The rewritten entry records four things the first draft could not: that the enumeration
itself was the defect (`env` was the ninth unbound `invoke` field, and `defaultHost` and
`path` are egress-relevant on their own), that the residual is what closes the class,
that D4.5's byte-identical-signature property is preserved by appending the residual only
when a lane declares something beyond the eight bound fields, and that consent — not
shape validation — is the boundary, since no honest env denylist can exclude `PATH`.

It also closes the internal inconsistency the previous entry could only record. This ADR,
`CONTEXT.md`, `gsd-core/workflows/review.md` and `resolveLanePlan`'s own header all said
overlay lanes reach the resolver while the runtime said otherwise; #3062 resolved that in
the documents' favour, which is what makes the disclosure mandatory rather than defensive.

Refs #2483.

* enhance(#2483): record in the manifest reference that invoke fields are consent-bound

`docs/reference/capability-manifest.md` is the field table ADR-2782 points capability
authors at, and it described `invoke` purely as a schema. A third-party author reading it
could not learn that everything they declare there is shown to the user at install and
bound to the consent signature — which is exactly what they need to know now that an
overlay reviewer lane executes (#2927/#3062).

States the two things the schema alone cannot: that `env` and `defaultHost` are named in
the consent prompt and the rest is covered by a residual, so any change to a declared
`invoke` field forces re-consent; and that `env`'s validation is a portability policy
rather than a safety boundary, since `PATH` is a complete execution primitive and cannot
be refused. Names that are execution primitives are highlighted in the prompt instead.

Refs #2483.

* enhance(#2483): add a Security changeset for the reviewer-lane disclosure

The existing fragment describes the enhancement this PR was opened for and stays as it
is. The disclosure fix is a separate user-visible change of a different type: a
capability declaring `invoke.env` or `invoke.defaultHost` will ask for consent once
more, and users are entitled to read why in the changelog rather than discover it as an
unexplained prompt.

Type is `Security` rather than `Changed` because the entry describes a closed
code-execution disclosure gap, not a behaviour adjustment.

Refs #2483.

* enhance(#2483): correct this round's own claim about who gets re-prompted

Self-found while auditing the round's claims before publishing them. The changeset and
the ADR entry both stated that a capability declaring `invoke.env` or `defaultHost`
"will ask for consent once more". That is wrong, and it overstated the cost of the fix
in the one direction a maintainer would have had to take on trust.

A code change to `disclosureSignature` re-prompts nobody. `hasProjectConsent` matches on
the recomputed bundle `contentHash` — the signature has not been the security binding
since #1459 CB-1/CB-2 — and the upgrade path's `executableSetChanged(old, new)` compares
two disclosures both computed by the CURRENT code, so widening the signature moves both
sides of that comparison equally. First-party capabilities never reach the path at all:
the install flow blocks a first-party id before trust evaluation.

What the widening actually buys is forward-looking, and is the real argument for it: an
upgrade whose manifest edits a declared `invoke` field now registers as an
executable-surface change and re-consents, where before it could change what the lane
runs in silence.

Also measured and recorded, because the D4.5 property was stated more strongly than it
deserved: of the twelve first-party reviewer capabilities, ZERO are in the
byte-identical-signature class — every real lane declares at least `effortChannel`. The
property is a guarantee about minimal lanes, not a description of the fleet, and the ADR
now says so.

Refs #2483.

* enhance(#2483): sign and disclose the probe binary and the lane's outer fields

Found by this round's own adversarial review, and it is the same defect one level out:
the `invoke` residual cannot reach the lane body's OUTER fields, and `probeLane` SPAWNS
`probe.binary` with `--help` before dispatch (`review-lane-runner.cts`, the
`command-exists`/`command-capability` arms). An overlay naming an arbitrary probe binary
therefore executes it — unsigned and undisclosed, exactly as `invoke.env` was, and
reachable on the same #3062 path.

The lane element now carries a second residual over the outer fields, and the probe
binary is shown in the consent prompt when it differs from the dispatch binary — it is a
program that runs, and the user is entitled to see it.

TWO fields stay excluded, and that is a decision rather than an omission:
`reviewsSection` and `timeoutFloorMs` are ADR-2782's cosmetic carve-outs (matrix
A10/A13), where re-consenting would present a prompt carrying no security information.
A test pins that they remain excluded, so a later widening cannot quietly reverse D4.5
while claiming to complete this fix.

Also corrects a miscount introduced by the previous commit: the source comment said the
enumeration had fallen behind by "seven fields" and omitted `fallbackModel`, while
asserting `env` was the ninth. `resolveLanePlan` reads twelve `inv.*` fields and four
were bound, so the number is eight. The comment now states the derivation rather than
just the total.

Reversion-controlled: emptying the outer residual fails "repointing the probe binary must
force re-consent"; removing the render line fails its own named assertion.

Refs #2483.

* enhance(#2483): refuse execution-primitive env names as defence in depth

Adopts the review's B5 after this round's own adversarial pass refuted my reason for
declining it. I had argued a denylist was worthless because `PATH` can never be refused.
That was wrong on the facts: no shipped reviewer manifest declares `PATH`, so it can be
refused, and it is the most complete primitive in the set — repoint it at a directory
holding a fake binary and the declared `invoke.binary` is irrelevant. A list that cannot
be exhaustive can still close the highest-confidence, lowest-legitimacy routes.

So the validator now rejects `PATH`, `NODE_OPTIONS`, `LD_PRELOAD`, `DYLD_INSERT_LIBRARIES`,
`BASH_ENV`, `PYTHONPATH`, `PERL5OPT`, `RUBYOPT`, `GIT_SSH_COMMAND`, `JAVA_TOOL_OPTIONS`
and their siblings on a reviewer lane. A lane needing a specific executable declares an
absolute `invoke.binary` instead of reshaping the child's environment.

The comment states plainly that this is defence in depth and NOT the boundary — the
boundary is install-time consent, which discloses every declared pair and binds it to the
signature, so an unlisted name is still SEEN before it runs. That framing is load-bearing:
a future reader who mistakes the denylist for the control will under-invest in the one
that is, which is the failure mode I was trying to avoid by declining it outright.

Two tests: the rejection itself across ten names, and a guard asserting no shipped
reviewer capability declares a denied key — so if the list ever outgrows its evidence,
that surfaces as a decision rather than a silent removal.

Refs #2483.

* enhance(#2483): fix two stale D5 enumerations elsewhere in the ADR

The previous commit rewrote the amendment's premise but swept only the amendment. Two
normative passages earlier in the same ADR still enumerated the old closed field list and
now contradicted it: the `executableSetChanged` trigger list, and the split-binding note
asserting the seven manifest-derived fields were "everything that is SHA-pinned".

That is the failure the rewrite-don't-annotate rule exists to prevent, one section over —
an ADR is read in fragments, and a fragment carries no supersession marker, so a reader
landing on either passage would have taken the superseded enumeration as current.

Both now name the residual as the mechanism rather than restating a list, which is also
what stops them going stale the next time the vocabulary widens.

Found by this round's adversarial review, which grepped the whole document rather than
the section under edit.

Refs #2483.

* enhance(#2483): stop the probe disclosure claiming a spawn that does not happen

The probe line added one commit ago rendered "probes by running: <binary> --help" for
every lane. That is false for `kind: "command-exists"`, which only calls `hasBinary` — a
PATH/filesystem scan that starts no process. Only `command-capability` spawns.

A false statement in a consent prompt is worse than a missing one: the prompt is the
surface a user is asked to trust, and this one overstated what a lane does. Worse, the
test I wrote to prove the fix used `command-exists` — the kind that does NOT spawn — so
it pinned the wrong claim and would have kept the error green forever.

The surface now carries `probeKind` and the two kinds render differently: a spawn is
described as a spawn, a presence check as a presence check. The test exercises both, and
asserts the `command-exists` path never emits the spawn wording.

Also corrects the field-count parenthetical to state its derivation unambiguously —
`resolveLanePlan` reads thirteen `inv.*` fields including `env` (twelve before this PR),
four were bound, so eight were unbound before `env` and nine including it. The bare
"twelve" was true only of the pre-PR tree and read as a claim about the current one.

And retires two comments that argued AGAINST the validator denylist this round then
shipped. Leaving them would have handed the next reader the reasoning for removing it.

Reversion-controlled: conflating the two probe kinds fails a named test.

Refs #2483.

* enhance(#2483): match the reviewer-lane env denylist case-insensitively

The denylist added one commit ago compared exact case, so `Path`, `path`, `node_options`
and `Node_Options` all passed it. Windows environment lookup is case-insensitive, so
those reach the child as `PATH` and `NODE_OPTIONS` — the exact inputs the list names.

An exactly-cased denylist is worse than none: it reads as a control while admitting the
input it was written to refuse, and the next reader has no reason to doubt it. Members
are stored uppercase and the key is folded before lookup; the name grammar already
constrains keys to ASCII, so a plain fold is sufficient.

Reversion-controlled: restoring the exact-case compare fails `envDenylistIsCaseInsensitive`
on `Path`.

Refs #2483.

* enhance(#2483): correct the docs that still described the denylist as absent

Both the ADR and the manifest reference still said `env` carries no denylist and that
`PATH` "can never be refused" — written when that was this round's position, and left
standing after the round reversed it. A reader landing on either passage would have taken
the superseded argument as current, which is precisely the failure the rewrite-don't-
annotate rule exists to prevent.

Both now describe the denylist, name `PATH`'s inclusion and the case-insensitive match,
and keep the limit explicit: the list cannot be complete against an arbitrary child and
disclosure runs before validation, so consent remains the boundary.

The ADR's byte-identical-signature claim is also corrected rather than softened. With the
outer residual in place, a lane producing no residual is one the validator rejects — it
declares no `flags`, `probe`, `emptyOutput`, `evidenceClass`, `requiresBinaries` or
`promptBudgetKey`. So the property is about the ENCODING, not a claim that any real
signature is unchanged, and it is not the argument for the change being safe. That
argument is that consent binds to the bundle contentHash and no existing consent is
invalidated at all.

Refs #2483.

* test(#2483): cover the three new lane disclosure fields in the injection-safety parity guard

The PARITY test in section N exists to catch a renderer field that skips
`renderValueForPrompt` (#3248). Its payload manifest is hand-maintained, so it
covers the fields that existed when it was written — slug, binary, args,
hostConfigKey, handler — and none of the fields this PR adds.

This PR renders three further manifest-supplied values into consent-prompt
lines: `invoke.env` (keys and values), `invoke.defaultHost` and `probe.binary`.
The gap was silent rather than theoretical: with the lane env line reverted to
the pre-#3248 raw form, the whole 948-test lane/capability/trust-disclosure
suite stayed green.

Two manifests, because the shapes render disjoint lines — `defaultHost` only on
the openai-http branch, `env`/`probe` only where declared, and the probe line
only when the probe binary differs from the dispatch binary.

Non-vacuity is asserted on the typed disclosure object and on structural line
counts, not by substring-matching rendered prose: CONTRIBUTING.md forbids raw
text matching on test output, and this section's own header promises structural
assertions only, so a prose match here would have made that promise false.

Negative-controlled three ways against the merged tree, each producing exactly
one named failure: env rendered raw, defaultHost rendered raw, probe binary
rendered raw.

* docs(#2483): extend the #3248 render-site comment to the fields this PR adds

The comment enumerates every manifest-supplied value that must pass through
`renderValueForPrompt`, and it stopped at `handler` — the reviewer-lane fields
that existed when #3248 landed. This PR renders three more (`defaultHost`, the
probe binary, and the env keys and values), so the list understated its own
contract in the one place a future author would check before adding a fourth.

A comment enumerating a closed set is a set that can silently fall behind the
code it describes; the parity test added alongside is what makes the omission
fail loudly rather than read as deliberate.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-11 17:42:41 -04:00
Tom Boucher
cf6de5e1c0 feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution

23 tests over the 50-test-matrix rows. RED by construction:
resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet,
and the validator silently ignores triggerPrecedence today.

Written in the per-runtime describe idiom the other four
runtime-artifact-layout suites use, not a table.

The rows that carry the weight: windsurf must NOT report a shadow it
does not have, since its global scope emits only agents and agents are
not trigger-bearing; agents and kimi-agents must be absent from the
output for every runtime; and reordering a runtime's triggerPrecedence
must flip the winner, which is the only assertion that proves the axis
is read rather than decorative.

Stems are injected, never scanned, so the surface is assertable with no
filesystem.

* feat(#2871): resolve triggers and host precedence, not just placement

resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger
a runtime emits, with the scope and kind that produced it, whether the
host registers it directly or only through a router, and which artifact
shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers
need placement only and the issue requires them unchanged.

AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The
host-integration matrix models command and dispatch as separate interface
points: an agent is invoked through the Agent tool's subagent_type, not
by typing a slash trigger, and _copyStaged never applies the kind prefix
to an agents entry. So agents and kimi-agents are excluded from the
surface entirely, and this commit amends ADR-2866 with a dated
correction. #2218's conclusion is unchanged -- the collision is strictly
commands-vs-skills, and claude's local /gsd-* trigger surface is still
fully shadowed -- but the ADR implied the local agents surface was lost
too, and it is not.

That correction is what makes windsurf come out right. Its global scope
emits only agents, so it has no global trigger and its local commands
are unshadowed. Model agents as trigger-bearing and windsurf falsely
reports a full shadow.

The triggerPrecedence axis lands on all 19 descriptors as an ordered
kind list, one value with one owner, rather than a numeric rank spread
across N kind entries with nothing keeping them consistent. Validation
uses a required-with-default shape that has no precedent in this
validator -- every existing axis is hard-required -- so a third-party
capability.json omitting the field still validates, which is what
ADR-894's additive-only contract promises.

Winner resolution reads Phase 1's scope rank first, then the kind
ordering. A test reorders the axis and asserts the winner flips, since
an axis that is added, validated and never consulted would pass every
other assertion.

shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this
issue's out-of-scope note.

Verified via the remote runner.

* fix(#2871): single-source namespacedByDir and close two test gaps

Four findings from the isolated adversarial review.

The namespacedByDir rule had reached three copies -- install-engine,
surface, and the new trigger resolver -- one of which carried a
hand-written keep-in-sync comment and no assertion. That is this repo's
generative-fix-divergence class. Extracted to one exported predicate all
three now call. Verified by diverging one copy deliberately: the existing
#816 parity test failed, and passes again on revert.

The omission test was vacuous. Row 16 asserted that a descriptor without
triggerPrecedence still validates, but built its fixture from claude's
shipped descriptor -- which this PR had just added the axis to. It now
clones and deletes the key, following the shippedDescriptorWithout
pattern, and asserts both that validation passes and that the resolver
still picks the right winner from the default. The second half is what
makes it prove anything.

resolveTriggerSurface silently dropped an unrecognized scope while every
sibling in this epic throws. Two phases of one epic should not disagree
about whether an invalid scope is an error, so it now rejects through the
same shared validator; an empty scope list still returns empty rather
than throwing.

The ADR amendment had been spliced into the middle of the References
list, orphaning its last bullet. Moved to the top, after the header
block, which is where ADR-3660 and ADR-1016 both put dated amendments.
No lint checks markdown structure, so this was green while malformed.

* fix(#2871): single-source the command filename composition too

The earlier fix shared the namespacedByDir boolean but left the
filename composition around it written twice -- once in _copyStaged as
what actually gets written, once in resolveTriggerSurface as what gets
predicted. The predictor could go stale silently.

One exported helper now composes it for both. The entry.name asymmetry
that looked like it would block extraction does not: entry.name is
filtered to end in .md and stem is entry.name minus those three
characters, so the two branches are the same string by construction.

Divergence proven to fail: injecting a marker into the helper broke the
trigger-surface suite; reverting restored 25/25. The four sibling layout
suites hold at 227 unchanged.

* docs(#2871): correct the ADR timing notes that this phase makes stale

The Amended by back-links on ADR-3660 and ADR-1016 were written in
Phase 0, when the widenings they describe had not shipped. Each carried
a forward-looking clause -- "the module changes at Phase 2, not before,
until then this module resolves placement only" -- which becomes false
the moment this PR merges. ADR-2866's own Amends header and its
reciprocal-notes section carried the same tense.

All four now describe what shipped. This is a tense and status
correction on Accepted ADRs, not a change to any decision.

Worth stating because it is the failure mode this epic keeps meeting:
gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI
would have caught either the missing back-link in Phase 0 or these stale
clauses now. They stay correct only because someone checks.

* chore(#2871): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:25:42 -04:00
Tom Boucher
b901d1e06f feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook

60 behavioral cases against src/complexity-trigger.cts, which does not exist yet:
decision-point counting, the comment/literal stripping leak surface, threshold and
jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics,
and fs fault injection via mock.method. Two fast-check properties assert that
stripping never manufactures a decision point and that comments and string
literals are score-neutral.

Also registers the refactor-trigger capability manifest (inert until
refactor.trigger_enabled) and regenerates the capability registry and matrix.

Verified RED on the remote runner before any implementation exists.

* feat(#1953): complexity-triggered refactor extension point

Adds the opt-in refactor-trigger capability. After a phase executes, an
execute:post step measures per-function complexity for the files the phase
touched and writes a scoped refactor proposal when a function crosses the
configured threshold or drifts past its recorded anchor.

Design notes worth carrying:

- The signal is computed in-core (decision-point counting over comment- and
  literal-stripped source, Node builtins only) rather than via Memtrace or a
  shelled-out analyzer. The hook fires as a deterministic CLI, not an agent
  with MCP tools, and core takes no external dependencies — this is the only
  option a behavioral test can bind to. The metric sits behind a seam.
- The baseline is a stable anchor, not a rolling value: set on first
  observation, moved only on disposition. A rolling baseline makes the delta
  the single-phase change, so a function creeping +2 per phase never trips a
  delta of 5 and the jump check adds nothing over the absolute threshold.
- Strict mode records an open deviation window in the broken-windows ledger
  rather than declaring its own ship:pre gate. ship.md has no generic ship:pre
  gate dispatch — only two hardcoded branches — so a third gate of any kind
  would be declared and never evaluated.
- The gate clears on the proposal being dispositioned, never on the score
  improving. A blocking complexity number is one an executor can satisfy by
  splitting a coherent function in two.

execute-phase.md gains a generic execute:post step-dispatch contract; it
previously matched only ref.skill == "code-review", so any other step
registered there was declared and never run. The code-review branch is
unchanged.

Full rationale in ADR-1953.

Closes #1953

* fix(#1953): close git option injection and symlink escape in the refactor hook

Three findings from the isolated security review, all fixed inline.

HIGH — changedFilesSince interpolated the --since value into a revision
token placed before the -- separator. A -- only stops PATHSPEC parsing of
arguments after it; git still option-parses what comes before. So
--since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts
as --output=<file> and uses to redirect diff output — an arbitrary write.
Fixed with --end-of-options before the revision range plus a conservative
ref validator. The validator deliberately permits ~ ^ @ { } because those
are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct
from ref-NAME syntax; --end-of-options is the actual barrier. The doc
comment asserting the trailing -- was sufficient was wrong and is corrected.

MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink
committed inside the repo passed the check (its own path is under cwd) and
readFileSync then followed it outside the root. Now lstat-checks for a
regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one
bad path skips one file and the run continues.

LOW — the new execute:post dispatch contract showed the gsd_run example
before the rule requiring ref.command be validated first. That prose is
executed by an agent, so textual order is execution order. Reordered.

Refs #1953

* fix(#1953): make the analyzer able to see TypeScript at all

Found by running the shipped analyzer over its own source: it reported
functions=1 for a 940-line module with 24 function forms. A return-type
annotation or a generic parameter list made a function invisible —
`function f(a): number {}` and `function f<T>(a: T): T {}` both detected as
zero. Since gsd-core is written in .cts and the capability declares
.ts/.cts/.mts analyzable, the feature silently found nothing in this repo's
own primary language while reporting success. A safety net that reports
"all clear" because it cannot see is worse than no safety net.

All 98 tests passed over this, because every fixture was plain JS — the
exact failure the test matrix's own "assert against the shape production
uses" warning describes. Adds a TypeScript-shapes suite covering return
types (including unions, generics, object literals and type predicates),
generic parameter lists (constrained and defaulted), export/async/generator
combinations, annotated arrows, class-method modifiers, and optional/
default/rest params — plus the two traps: an overload signature has no body
and must not count, and `a < b && c > d` is a comparison, not a generic.
Detection now reports 24/37/21 functions for the three source files, which
matches a hand count exactly.

Also from review:

- The strict-mode ledger dedup identified entries by parsing a prose
  description string. That is banned by CONTRIBUTING's raw-text-matching
  rule and was a real bug: the "exactly one window per untriaged proposal"
  guarantee rested on prose matching, so rewording a description or editing
  WINDOWS.md by hand silently produced duplicates. Now matches structurally
  on kind + phase + file + line.
- A property test asserted on the stripper's output text. Reframed to
  assert the same invariant through analyzeSource's score.
- nextBaseline's `candidates` parameter has been dead since the anchor
  change; removed from the signature and all call sites.
- Extracted the duplicated require-or-degrade and capability-check
  boilerplate.
- ADR-1953's Implementation bullet still named a `refactor.ship-gate` in
  check-command-router.cts — a leftover from the design cut D6 rejects.
  That file is untouched and no such gate exists. Removed.

Refs #1953

* fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test

Five of the seven remote-runner failures were one cause: the execute:post
dispatch contract, written out inline, grew execute-phase.md 1876 bytes
(93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack
does not clear that — tests/phase6-capstone-conformance.test.cjs and
tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is
literally under the cap.

The contract now lives in gsd-core/references/loop-hook-dispatch.md, which
already claimed to be the point-agnostic dispatch reference and already
documented ref.skill and ref.agent. It gains the ref.command shape, its
in-context validation rule, the advisory-by-construction statement, and a
note that a point whose workflow hand-rolls one kind is not implementing
this contract. execute-phase.md now defers to it in one line: 145 bytes of
growth, 55 B of headroom under the cap. Better placement than the first cut
— the reference was overstating its coverage, and this makes the claim true
rather than duplicating prose next to it.

Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather
than a new 1953-*.json: two ack sources may never name the same path, and
that fragment is already the accumulating ack for this file.

Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own
vacuity guard — the fixture generated ~480 KB against a `> 500000` assert,
so the guard fired and the three assertions after it never ran. The test
has been vacuous since it was written. The matrix row specifies ~1 MB, so
N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000.
Verified by reproducing the exact body against the compiled module: 1168888
bytes, 118 ms, all four assertions hold.

Refs #1953

* fix(#1953): fold the execute:post step deferral into the existing resolve line

The remaining two failures were one test: execute-phase.md carries a SECOND,
tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its
current size. The file cannot grow by a single byte. My previous fix got it
under 93,600 but not under 93,400, so it still failed. ("H." in the report is
just the parent describe of that same test, not a separate defect.)

Rather than add a paragraph, the deferral now REPLACES the existing hook
resolution line. It read:

  Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where
  `kind == "step"` and `ref.skill == "code-review"`.

which is the bug itself written down — only code-review was ever dispatched.
It now reads:

  Dispatch each `kind == "step"` hook per
  @gsd-core/references/loop-hook-dispatch.md. For `code-review`:

The following prose already begins "If no active code-review step hook
exists", so it reads correctly and the code-review handling is untouched.
Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin
assertion rather than merely under the ceiling.

That also removes the need for a drift-ack: the file shrank, so there is no
growth to acknowledge, and the append to the shared 2930-*.json fragment is
reverted. Leaving it would have shipped a claim of "145 bytes of growth"
that is no longer true, on a file six other issues share.

The test's own comment states the principle this ended up honoring: "the host
loop must stay small — optional-feature detail belongs in the capability
fragment, not the host workflow." Putting the dispatch contract in the
reference rather than inline is that rule, applied.

Refs #1953

* fix(#1953): keep the code-review hook literal the workflow test requires

tests/code-review.test.cjs extracts the <step name="code_review_gate"> block
and asserts it contains `ref.skill == "code-review"` verbatim. The previous
commit replaced the line carrying that literal, so the token vanished and the
test went red — a fair assertion: code-review IS the bespoke branch there and
the workflow should still name it.

Restored inside the same one-line deferral, which now reads:

  Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md.
  `ref.skill == "code-review"`:

93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below
the base, so the file continues to shrink rather than grow.

Because three consecutive runs were each reddened by a different assertion on
this one file, this change was verified by sweeping ALL of them at once rather
than one run at a time: every test under tests/ that reads execute-phase.md or
references/loop-hook-dispatch.md was located by resolving its path constants,
and each content/size assertion was evaluated directly against the working
tree — 22 assertions, plus two real executions (gen-section-manifest --check,
and emitted-attribution's full real-tree differential). All pass.

That sweep also confirms the earlier judgement call: the net change to
execute-phase.md is a SHRINK, and the size ratchet only gates growth, so
reverting the append to the shared 2930-*.json ack fragment was correct — an
ack would have been both unnecessary and factually wrong.

Refs #1953

* chore(#1953): backfill changeset pr number to 3261

* docs(#1953): add the missing how-to for acting on a refactor proposal

Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md,
FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and
that is the one a user reaches for. CONTRIBUTING's required-docs table is
'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap.

Enabling this feature is genuinely multi-step and no single page walked it:
turn it on, tune the threshold, understand advisory vs strict, discover
that strict needs a SECOND toggle on a DIFFERENT capability, and know what
to do when a proposal appears. The two-toggle subtlety in particular was a
footnote in a config table; here it is a section with both commands.

Follows the shape of its closest siblings, resolve-edge-coverage-findings
and resolve-prohibition-findings — both 'the loop surfaced a finding, here
is what to do with it'. Includes a reason-code table for the silent cases,
since the analyzer is deliberately quiet in six situations and a user who
expected a proposal needs to tell 'nothing to report' from 'could not look'.

Indexed from docs/README.md beside the other loop how-tos.

Docs-only: exempt from the push gate, no re-verification, pass marker on
2af188b4 untouched.

Refs #1953

* feat(#1953): warn when strict mode is on but nothing will actually block

Closes acceptance criterion 5, which I had wrongly marked satisfied.

refactor.trigger_strict records an untriaged proposal as an open deviation
window, but a ship only STOPS if workflow.windows_enforce is also on — a
toggle owned by the broken-windows capability that this feature neither sets
nor requires. So a user could enable strict, believe ship was gated, and find
out otherwise at ship time.

The split itself stays: requires:["broken-windows"] would force-install the
ledger on advisory users who never enable strict, and a ship:pre gate of our
own would never fire because ship.md has no generic ship:pre gate dispatch.
What was missing was discoverability, so that is what this fixes.

`refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning,
naming the exact remediation command, whenever strict is on and either
workflow.windows_enforce is off or broken-windows is unavailable. It fires
only on a run that produced a candidate — with nothing to block on there is
nothing to warn about, and warning every run would be noise.

Reads workflow.windows_enforce through the same resolveConfigKey walk the
router already uses for its own keys rather than a second config reader.
Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does
not, strict+ledger-absent warns, strict-off never warns.

Also corrects a user-facing message in this same file that told the user to
run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the
broken-windows capability both use `gsd config-set`, and gsd-tools is invoked
as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent
messages in this file now agree.

Refs #1953

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:52:47 -04:00
Tom Boucher
653f95e39f chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias (#3272)
* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal

Inverts the Phase 5a rows that assert the derived legacy alias still
contributes a reviewer slug, and adds the removal-warning coverage the
alias's exit needs (ADR-2782 D9).

RED against unmodified production code, by design: the six shipped
manifests still declare the key and collectReviewerWarnings emits nothing
for hostBehaviors.

Refs #2801

* chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias

ADR-2782 D9, Phase 7 — the final phase of epic #2782.

The derived legacy alias survived one release (Phase 5a shipped in 1.9.0;
1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer
body is now the only route onto the reviewer roster.

- deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli
- the key is stripped from the six manifests that carried it; each already
  declares a reviewer body whose slug equals its capability id, so the
  derived roster is unchanged at the same twelve slugs
- collectReviewerWarnings emits a presence-based, non-fatal removal notice
  for any manifest still declaring the key, reaching both the build-time
  registry generation and the third-party overlay load path. The check runs
  before the reviewer-body early-return, because the manifest it exists for
  is the alias-only one that has no body.
- hostBehaviors stays an open, unvalidated bag for its other 59 keys; this
  adds one keyed removal notice, not general validation

Refs #2801

* refactor(#2801): give the reviewer-warning channel a typed IR

Review finding: the new tests asserted with String#includes() on the
warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on
Test Outputs' bans in favor of a typed intermediate representation.

Adds the IR beside the renderer rather than replacing it, which is the
shape that section prescribes and bin/verify-reapply-patches.cjs already
models:

- REVIEWER_WARNING, a frozen code enum
- REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one
  symbol instead of duplicating a literal
- collectReviewerWarningRecords(cap), returning typed records

collectReviewerWarnings(cap) keeps its exact string[] contract as a thin
map over the records, so both production consumers are untouched. Every
section-K row now asserts on record.code/field/capId and none on the
rendered message. Locks the code surface, asserts the renderer stays
one-to-one with the records, and migrates the pre-existing Phase 2 test
on the same channel off prose matching.

Refs #2801

* test(#2801): invert the section F alias fall-through regression row

Caught by the remote runner: 2 unique failures on both Node lanes out of
31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's
isolated-security-review regressions — asserted that a blank reviewer.slug
falls through to the hostBehaviors.reviewerCli alias rather than dropping
the lane. That is the direct inverse of this phase's contract.

The original rationale held only while the alias existed. With it gone
there is nothing to fall through to: a blank body is not a declaration,
and a declaration is the only route onto the roster.

Inverted rather than deleted — the row carries the adversarial-review
provenance for the slug trim, and removing a security regression guard to
make a change pass is backwards. The duplicate row added earlier in
section C is dropped instead; section F is its canonical home.

Also corrects two count strings Phase 5b left at eleven while asserting
twelve, which would misreport on failure.

Refs #2801

* docs(#2801): give the removed reviewerCli flag a migration path

The Reference edit alone satisfied CI — a file under docs/ moved, so
lint-docs-required.cjs was green — while the task-oriented quadrant said
nothing about the removal. A maintainer whose lane had just gone silent
would have found the field documented as removed and no page telling them
what to do about it.

Adds a migration section to the how-to: the symptom, the verbatim warning
they will see, the before/after manifest, and the note to keep the
reviewer slug equal to the capability id so existing
review.default_reviewers entries and --<slug> flags survive.

Refs #2801

* chore(#2801): backfill changeset pr number to 3272

* feat(#2801): close the runtime.hostBehaviors vocabulary

ADR-1016 closes twelve descriptor axes and rejects an open escape hatch
in the descriptor. It never mentioned runtime.hostBehaviors, and that
silence was read as permission: 59 keys across 18 manifests, 39 of them
set by a single capability, validated by nothing. The reference docs went
further and attributed the open seam to ADR-1016, which does not mention
the field at all.

KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a
non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the
alias removal notice, reaching both build-time generation and overlay
install.

Warning, never error, for the reason this phase exists: an error would
hard-break an out-of-tree descriptor carrying a bespoke key with no
deprecation window, which is what reviewerCli was given a release to
avoid. Escalation is a separate decision.

reviewerCli is excluded from the unknown-key sweep so it keeps its own
notice with the migration pointer rather than drawing two records.

A parity test binds the vocabulary to the shipped manifests in both
directions, and a second asserts no shipped capability draws a notice, so
the closure is provably inert in-tree.

Records the decision and the miscitation as an ADR-1016 amendment.

Refs #2801

* fix(#2801): bound and sanitize the unknown-key diagnostics

Two findings from an isolated adversarial review of the closure commit,
both proven by execution rather than asserted.

MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep
had no ceiling. An installed third-party manifest is bounded only by
MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced
800,000 records and ~139MB of message text, retained for the registry's
lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a
summary carrying omittedCount, mirroring capability-loader's existing
slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars.

MINOR, newly reachable: manifest-supplied key names were interpolated raw.
Unlike cap.id, which validateCapability gates on KEBAB_RE before these
diagnostics run, hostBehaviors keys have no grammar check anywhere, so
ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New
describeKey replaces C0/C1 controls and clips at 80 chars. The file
already had describeValue for this and applied it only to values.

Both fixes land on the pre-existing reviewer.* sweep too — it carried the
identical pair, and fixing only the new copy would leave the same defect
one screen from its own fix.

Refs #2801

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:18:56 -04:00
Tom Boucher
b7431a9259 feat(#1956): flag cross-artifact fact drift in the plan drift guard (#3259)
* test(#1956): failing-first contract for cross-artifact fact-drift pass

* feat(#1956): flag cross-artifact fact drift in the plan drift guard

* fix(#1956): correct config-key assertion and bidirectional lifecycle-lag exemption

* docs(#1956): document the cross-artifact axis in the architecture reference

* feat(#1956): decide the phase-status drift axis deterministically

* fix(#1956): scope the progress-table lookup, abstain without a position section, rank deferred

* docs(#1956): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-09 14:23:35 -04:00
github-actions[bot]
4f2a4932ce chore: sync next package version to 1.10.0 2026-08-08 05:07:32 +00:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
3f349e551d fix(#3024): route sync-skills through the shipped gsd-tools instead of an unshipped install.js (#3195)
* fix(#3024): sync-skills workflow uses gsd-tools query skills-root instead of unshipped install.js

The sync-skills workflow Step 2 shelled out to gsd-core/bin/install.js --skills-root,
but install.js is not shipped in installed trees (only in the npm tarball root bin/).
Every /gsd-update --sync invocation failed with MODULE_NOT_FOUND.

Fix: added 'gsd-tools query skills-root <runtime>' subcommand (gsd-tools IS shipped)
that calls the same getGlobalSkillsBase function install.js used. Updated the
workflow to call gsd_run query skills-root instead of the dead install.js path.

Also documented the #3025 verbatim-cp limitation in Step 5 with a workaround.

* test(#3024): failing-first guards for the three defects in the adopted fix

The cherry-picked commit came from an aborted run that never executed its own
tests. Its raw-path assertion fails as written, which is the clearest evidence
the work never reached verification.

Covers:
- --raw must emit a bare path, not JSON (output() takes a third rawValue arg
  that routeSkillsRoot omits, so the raw branch never fires)
- an unknown, empty, whitespace, traversing, or metacharacter-bearing runtime
  must be rejected, not silently resolved to claude's skills root
- sync-skills.md must contain zero references to the unshipped install.js,
  including the guard's remediation text — the issue's second reported defect
- parity across every runtime in the registry, not three hardcoded ones, so the
  two entry points cannot drift

Also converts the adopted tests off a hand-rolled spawnSync onto the bounded
process seam, per CONTRIBUTING.

Fails before the fix. Verified via the remote runner.

* fix(#3024): make the skills-root query actually work and reach non-Claude runtimes

The cherry-picked commit never ran its own tests. Six defects, all fixed here.

--raw was ignored: output() is output(result, raw, rawValue) and the third
argument was omitted, so the raw branch never fired and the workflow captured a
JSON blob as SRC_SKILLS_ROOT. Every downstream cp -r then resolved against a
nonexistent path — the command would have shipped still broken.

An unknown runtime silently resolved to claude's skills root, because
getGlobalSkillsBase falls back rather than returning null, leaving the existing
=== null guard dead. The runtime id is now validated at the CLI boundary against
the shipped registry, so a typo'd --from/--to fails instead of reading from or
writing into the wrong runtime's tree.

getGlobalSkillsBase('vscode') threw a raw TypeError. vscode is non-installable
by descriptor, so it has no skills root — null is the answer, not a crash. The
resolver now short-circuits configHome.kind 'none', which also fixes the same
latent crash in install.js --skills-root vscode. Every caller already gates on
=== null.

sync-skills.md used gsd_run WITHOUT the canonical launcher preamble, so gsd_run
was undefined on non-Claude runtimes — the fix would have been dead in exactly
the place the original bug bit. Preamble propagated via sync-runtime-launcher.

Also registers skills-root in TOP_LEVEL_USAGE (the help/dispatch parity guard
caught it), removes the last two install.js references including the guard's
remediation text (the issue's second reported defect), and updates the stale
assertion that still described the removed contract.

Verified on the remote runner.

* fix(#3024): align the documented runtime list with the registry and gate both entry points

Isolated review returned BLOCK on two findings.

The workflow's Supported-runtimes list and its --to all expansion named grok and
gemini, neither of which is a registered runtime. Once this branch added
validation, --to all — a documented first-class feature — aborted. The list was
hand-copied prose shadowing the registry, so correcting it alone would drift
again; a parity assertion now fails in BOTH directions if the doc and the
registry disagree. vscode is excluded by name: it is installSurface 'none', so
syncing skills to it is meaningless and would abort.

bin/install.js --skills-root reached getGlobalSkillsBase with no own-property
gate, so --skills-root __proto__ silently resolved to claude's skills root. This
branch had just hardened the OTHER entry point to the same function; leaving one
of two parallel surfaces open is the same divergence class as the first finding.
Both now call one shared isRegisteredRuntimeId() rather than a copied check, and
the parity test covers the hostile ids so the two can never disagree again.

Also guards the workflow's root resolution: neither command substitution checked
its exit status and only the source had an existence guard, so a failed
destination resolution left DEST_ROOT empty and turned rm -rf "$DEST_ROOT/$SKILL"
into an absolute path at filesystem root. Both resolutions are now checked, and
Step 5 requires both roots to be non-empty and absolute before any destructive
command.

Verified on the remote runner.

* test(#3024): anchor the runtime-list parity extractor to the list span

The extractor captured (.+) to end of line, so it swallowed the em-dash prose
that explains the vscode exclusion — and that sentence contains backticked
`runtimes` and `null`, which is where the three phantom ids came from. The
documented list was correct; the test was reading its own explanation back as
data. Anchored to the id-list span.

Both directions still fail as intended: proven by injecting a bogus id and by
removing a registered one.

* test(#3024): anchor the --to all extractor and fail loudly on empty captures

The workflow has three TO_RUNTIMES= assignments and the regex matched the first
one — an empty array initializer at line 28 — so the extractor captured nothing
and the assertion diffed [] against 18 ids as if that were data.

That is the same failure twice, so the fix is the general one: every extractor
in this test now asserts it captured a plausible list before comparing, naming
which extractor found nothing and what it was looking for. An extractor that
silently yields [] is a confident wrong answer, and a parity guard that reports
it as a data mismatch teaches the reader to loosen the assertion.

Verified against the real workflow and against doctored copies with each target
construct removed, plus both teeth directions.

* fix(#3024): merge duplicate process-seam import after rebase

The rebase applied cleanly but left runNode declared twice: next had gained its
own import of the seam while this branch added one carrying OUTCOME. A clean
rebase is not a correct one — the file no longer parsed. Merged into a single
import providing both.

* fix(#3024): bind DEST_ROOT per destination instead of a dangling map

Step 2 stored each destination's root into DEST_SKILLS_ROOTS, which nothing ever
read, while Steps 3 and 5 used a scalar DEST_ROOT that nothing ever assigned. The
array was also never declare -A'd, so on bash 3.2 — macOS system bash, which this
repo supports — every destination collapsed onto index 0.

The absolute-path guard added earlier was the only thing standing between that and
rm -rf "/$SKILL"; it turned a silent disaster into a hard stop, but the feature
still could not complete. Each destination now binds its own DEST_ROOT where it is
used, and the unread map is gone rather than replaced.

Step 2 keeps eager validation, so a bad runtime id in a multi-destination --to
aborts before any destination is written rather than after some already have been.

Verified on bash 3.2 with a two-destination run binding distinct roots, and with a
bad id aborting before any destructive call.

* fix(#3024): restore grok support broken by the registry gate

The registry gate added earlier rejected grok, and that was my error. I confirmed
grok was absent from the capability registry and concluded the hardcoded branch
was dead — without checking what it resolved to. It resolves to ~/.agents/skills,
a real grok-specific path, exactly as the pre-fix workflow documented ('grok uses
the ~/.agents layout'), and there is a support discussion doc for it. So a
working, documented runtime silently lost --skills-root and sync-skills support
as a side effect of prototype-pollution hardening — and the parity test I added
locked that in as correct.

gemini is the one that really was dead: it fell through to CLAUDE's skills root,
so rejecting it is right and it stays rejected, as do bogus ids, __proto__,
empty, whitespace and traversal.

The validator's real question is 'does this id have a genuine runtime-specific
resolution', not 'is it in the registry map'. Registry membership was a proxy
that happened to miss grok. Legacy non-registry runtimes with dedicated
resolution branches are now a named, documented set; enumerating every hardcoded
branch in getGlobalConfigDir against the registry confirms grok is the only one.

The new tests assert grok resolves UNDER .agents and specifically not to claude's
root. Allow-listing an id proves nothing about whether it resolves correctly —
that assertion is what would have caught my mistake.

Also uses the shared PROBE_TIMEOUT_MS instead of a duplicate literal, and guards
Step 3's DEST_ROOT re-resolution, which contradicted the file's own stated
guarantee.

Verified on the remote runner.

* test(#3024): guard against LEGACY_NON_REGISTRY_RUNTIME_IDS drifting

The named legacy set is a second hand-maintained proxy for the same predicate
the registry check got wrong — 'does this id resolve runtime-specifically'.
Nothing stopped a third hardcoded branch being added to getGlobalConfigDir
without updating the Set, reproducing the exact class of bug that broke grok.

Production stays explicit and greppable; the test derives the truth instead. It
resolves a sentinel id to learn the generic fallback, classifies every candidate
against it, and fails in both directions — an id resolving runtime-specifically
that is in neither the registry nor the Set, or a Set entry that no longer earns
its exemption. The failure message names the remedy.

Confirms grok resolves runtime-specifically and gemini does not, which is the
distinction the original registry check could not see.

Also reverts the shared-timeout swap: SKILLS_ROOT_PROBE_TIMEOUT_MS is
pre-existing on next and arrived by rebase, so changing it here was scope creep
into another issue's territory.

Verified on the remote runner.

* chore(#3024): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-07 18:53:25 -04:00
Tom Boucher
27aa40f65e fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir

pi reserves <configDir>/hooks as its deprecated extension location and warns
on every startup when it exists. Assert a pi install stages the shared hook
bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/.

Also adds pi to the local-scope dir table in install-shared.cjs: pi was in
RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved
path.join(root, undefined) and no local pi install could be exercised.

Fails before the fix. Verified via the remote runner.

* fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir

pi reserves <configDir>/hooks as its now-deprecated extension location and
warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs()
guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged
its shared hook bundle exactly there, and pi's advised remediation (move it to
extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's
extension auto-discovery.

The bundle directory name is now runtime-descriptor-driven: hostBehaviors
.sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are
byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path
segment — separators, dot-only segments, trailing dots, absolute paths, NUL,
and Windows reserved device names all fall back to the default, because the
value is joined onto a user's config root and written to.

Renamed in place rather than relocated: hook scripts resolve siblings via
__dirname/.., so a depth change would silently break them.

- install / uninstall / manifest sites all read the resolved name
- pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded
  trees still resolve; the never-throws contract is preserved
- new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new
  non-recursive remove-empty-dir engine primitive (rmdirSync only,
  symlink-refusing, containment-guarded); ADR-0008 amended accordingly
- fixes two latent name-dependencies the rename exposed: the stale-hook scan
  and the injection scanner's self-exclusion both hardcoded 'hooks'

Verified on the remote runner.

Closes #3023

* fix(#3023): close review findings and align emitted provenance with the rename

Adversarial review found two defects, and the remote runner found four
failure clusters. All fixed here.

Review BLOCKER — detect-custom-files was blind to the renamed bundle.
GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the
whole gsd-hooks/ tree was invisible to the custom-file scan and user-added
files there were never backed up before the next update's clean-install wipe.
The dir set now resolves via the .gsd-runtime marker plus the shipped
capability registry (never bin/install.js, which is not shipped into installed
trees), and falls back to scanning every known candidate when the runtime
cannot be determined — over-scanning is safe, under-scanning is the data loss.

Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir
accepted any directory, so an interrupted install left gsd-hooks/ winning over
a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now
qualifies only if it is non-empty.

Remote-runner clusters:
- emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped
  rules pointing at the same sources the existing hooks/ rules use. The table is
  total, so an unattributed family is a hard failure by design.
- pi tests in install-minimal-hooks and the install integration suite asserted
  the old layout; updated to derive the dir name from the descriptor rather than
  hardcoding either name.
- 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after()
  runs in registration order, cleanup was registered before mock.restoreAll(),
  and node22's JS rimraf calls the public fs.rmdirSync while node24's native
  path does not — so the EACCES stub leaked process-wide on one lane. Restore
  now runs first.

Verified on the remote runner.

* fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde

pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent
(packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty
configHome.env, so a user with that variable set had GSD installed where pi
never looks. Added the env name; the dot-home-nested resolver already handled
the override, so no resolver logic changed.

Also fixes expandTilde in the shared runtime-homes resolver, found while adding
that: it hardcoded os.homedir() and ignored the opts.home every caller threads,
so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi)
silently resolved against the real home. That is a correctness bug and a
test-escape hazard — a sandboxed test asserting on a tilde override reached the
developer's actual home directory. Now threaded through every branch; behavior
with no injected home is unchanged.

Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location
moved with the rename. The provenance rules satisfy the totality gate; the
differential gate needs the ack because the hook sources are byte-unchanged —
only the installer's target directory moved. The two hook files this branch
genuinely edits stay attributed and are not double-acked.

Note on piConfig.configDir: it is read from pi's OWN installed package.json
(getPackageDir walks up from pi's __dirname), alongside piConfig.name — a
white-label setting for a redistributed pi fork, not a per-project user setting.
Documented accordingly rather than treated as an unsupported override.

Verified on the remote runner.

* fix(#3023): reject blank env overrides, pin adapter/descriptor parity

Three review findings, all fixed.

A whitespace-only config-dir override was accepted verbatim: the guard was
`if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR='   '
resolved to a literal three-space directory name instead of falling back to the
descriptor default. Fixed across every env-consuming branch — dot-home,
dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's.
Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working.

pi/gsd.cjs's probe list and the descriptor were two independent sources of truth
for the bundle directory name; a future rename would have desynced them silently
and left every pi hook quiet with no error. The probe list stays deliberate — it
must resolve in a dev checkout and a half-upgraded tree, where the registry's
answer would be wrong — so this adds the parity assertion the repo's
generative-fix-divergence rule calls for: the descriptor value must be the FIRST
candidate, and the default must remain present.

Changeset body rewritten to cover the two later user-facing fixes it had not
caught up with.

Verified on the remote runner.

* chore(#3023): backfill changeset PR number

* fix(#3023): anchor injection-scan patterns and fix a macOS detection hole

CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the
same fact as a genuinely empty or absent one'. The match was the 'act as a'
INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending
in act tripped it (fact, impact, contract, artifact, interact, redact,
abstract). My four-line CONTEXT.md edit dragged the latent false positive into
this PR because the scan is diff-scoped by file but reads whole files. Anchored
with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would
have left the class alive for the next PR touching any file saying 'fact as a'.

Auditing the rest of the list for the same class surfaced a real detection hole:
the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex
escape. BSD/macOS grep reads it as four literal characters, so single-quoted
eval('...')/exec('...') payloads were NEVER detected there while passing on
GNU-grep CI. Replaced with a literal apostrophe class.

Boundaries were added only where a real word-suffix collision exists; exec,
jailbreak, developer mode and the role-manipulation family were audited and
deliberately left unanchored. 22 new cases cover both directions — the false
positives now scan clean, and every real payload still fires, including the
quote/punctuation/start-of-line boundary forms.

Also builds this branch's injection test fixture at runtime instead of carrying
the literal phrase, so the payload keeps its teeth without tripping the scan.

Verified on the remote runner.

---------

Co-authored-by: sim <sim@local>
2026-08-07 13:41:21 -04:00
Tom Boucher
4b66bf4560 fix(#3086): apply #2667 .cmd-shim gate to deps.spawn + surface errorCode in review lanes (#3142)
* fix(#3086): apply #2667 .cmd-shim gate to deps.spawn + surface errorCode in review lanes

deps.spawn used shell:false with a bare binary name — on Windows, npm-installed
CLIs (gemini, codex, etc.) are .cmd shims that CreateProcess cannot start,
producing ENOENT + empty stderr. The review path then wrote an empty err file
and emitted a generic 'failed or returned empty output' stub.

Two fixes:
1. deps.spawn: detect .cmd/.bat on win32 and mediate through cmd.exe /d /s /c
   (same gate as runWithTimeout #2667, same explicit argv array).
2. runSpawnLane: surface errorCode (ENOENT, ETIMEDOUT) in the err file so the
   stub explains WHY the lane produced nothing.

* chore(#3086): backfill changeset PR number 3142

---------

Co-authored-by: sim <sim@local>
2026-08-07 07:43:14 -04:00
Tom Boucher
8f75e27554 fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag

Every isolation gate already resolved correctly. The resolved value then reached
the executor through a prose instruction telling the model to substitute it into
a call the model composes itself, and nothing verified the substitution. When it
was dropped, the executor edited and committed in the user's primary checkout
with no consent and no warning.

A prose backstop would be the same class of artifact as the defect, so this is a
shipped PreToolUse hook on the Agent tool. It fires at the instant of the call
rather than being read once at the top of a workflow, which is the only placement
the model cannot skip.

The guard is inert unless it can positively establish that this is a GSD project,
that the project resolves to harness isolation, and that the dispatch targets an
executor. A non-GSD repo has no invariant to enforce. Where it cannot read the
configuration at all, it denies rather than assuming, with its own reason -- a
guard that cannot verify must not answer safe. A malformed payload allows rather
than throwing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3045): extend the isolation guard to Cursor

Cursor is the second of only two runtimes that resolve harness isolation, so
shipping the guard for Claude alone left half the exposed surface unguarded while
the changeset implied it was covered.

The two runtimes fail differently. On Claude the harness flag is a per-dispatch
kwarg the model must copy into a call it composes, and the defect is that it can
be dropped. On Cursor the flag is --worktree, which applies to the whole session,
and the subagent-start payload carries no isolation field at all. There is no
flag to check, so the guard verifies the effective state instead: whether the
workspace is genuinely running outside the user's primary checkout. That is a
stronger check than the Claude one because it tests reality rather than intent,
and it is commented so nobody later rewrites it into a flag check.

Isolation is established two ways, either sufficient: the workspace resolves to a
linked git worktree, or it sits under the worktree root Cursor manages. The
second matters because a directory Cursor placed there is a legitimate isolated
session even before it becomes a distinct git worktree, where linkage alone would
report no repository.

Detecting linkage required a new primitive rather than the existing context
resolver. That resolver short-circuits on finding a local .planning directory
before it ever compares the git directory to the common one -- and an isolation
worktree normally has its own checked-out .planning. Reusing it would have read a
correctly isolated session as unisolated and denied it, which is the failure
direction that gets a guard switched off. The comparison is now its own
shortcut-free function that the resolver delegates to after its own shortcut, so
existing behavior is unchanged, and the case that would have broken is pinned.

The subagent type is checked before any configuration is read, so an unreadable
config cannot deny a dispatch this guard would never have enforced against.

The input-schema comment on the Cursor hook documented only the fields common to
every event and omitted the ones specific to this one. That omission cost a
halt during this work; it now documents both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): enforce the resolved dispatch decision, not the host capability

The guard keyed on the registry's dispatch.isolation, which says only that a
runtime is CAPABLE of harness worktrees. The decision that actually governs a
dispatch is the one the workflow resolves after gating, and that legitimately
comes out as sequential in three documented cases: a project setting
use_worktrees false, a per-plan submodule intersection, and the base-check
auto-degrade. The workflow tells the model to omit the flag in exactly those
cases, and the guard was denying every one of them.

The third case matters most. The preceding fix made the base-check degrade on
git timeouts and a missing git binary, where it had previously answered "safe".
That correction is right, and it means a transient hang now degrades to
sequential far more often than before -- so the two changes composed into a trap
where the workflow behaved exactly as designed and the guard blocked it.

The workflow already resolves isolation in shell, deterministically, which is
what makes it a trustworthy source in a way the model-authored call is not. It
now records that resolved value through a dedicated verb, and both guards read
it first. A fresh record is authoritative, so sequential dispatches pass
untouched. Absent or stale, the guards fall back to the capability check
combined with the project's use_worktrees setting, which still covers the case
that never reaches the workflow.

Also widened the matcher to accept Task alongside Agent, since a host that names
the tool Task would otherwise leave the guard silently inert while implying
coverage; stopped assuming Claude when no runtime is declared, which is the
shipped default and would have demanded a Claude-only argument elsewhere; and
made a non-git project inert rather than denied, since advising a worktree
session is not actionable without a repository.

The original diagnosis never modeled sequential mode as legitimate. That
omission is what let this through, and it is now recorded there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): record at resolution and bind the record to its dispatch

Two independent reviews converged on the same failure: the guard was fail-open in
a default install, so it did not catch the defect it exists to catch. A shipped
project carries no runtime key, which made "runtime not confidently known" the
common case rather than a corner one. A record asserting that isolation was
required but carrying no flag then fell through to a capability lookup that
answered "none", and the dispatch was allowed. The flag itself only arrived from
a second shell block -- the same block a model dropping the argument would also
skip. A test had pinned that behavior as intended.

The record is now written by the resolver, as an unavoidable consequence of
asking for the value, rather than by a step the model is told in prose to go and
run. A guard against a prose-carried value cannot itself depend on prose. Mode,
flag and identifiers are written together and atomically, so the flagless window
is gone, and a record asserting isolation with no resolvable flag now denies
instead of degrading. Runtime is also resolved from the installer's own recorded
default, which makes confident resolution the normal case.

The per-plan submodule gate degrades after the phase-level decision and never
re-recorded, so a plan that legitimately ran sequentially was denied against a
still-fresh phase record. It now records its own, scoped to the plan.

A record also authorized any dispatch for four hours. One phase degrading to
sequential could silently license an unisolated dispatch in the next. Records
now carry phase and plan, the guards require them to match, and the window is
minutes rather than hours -- the resolver rewrites it before every dispatch, so
a long window bought nothing and only widened the hole.

The flag validator rejected any value beginning with two dashes, which is exactly
the form Cursor and Windsurf declare, so their real value could never have been
stored. Writer and reader also derived the record path differently and diverged
inside a linked worktree without local planning state.

The predictable path remains a way to silence the control without leaving a trace
in the diff. It grants no access an agent with shell does not already have, so it
is documented as accepted rather than redesigned around.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3045): correct the staleness boundary and unmask a vacuous parity test

The remote runner returned twenty failures. One was a real production defect the
boundary case existed to catch: a record whose age exactly equalled the staleness
window was treated as fresh, so it stayed authoritative for one tick past its own
expiry. Freshness is now strictly inside the window.

The parity test meant to stop the two guards' executor lists from drifting could
never have failed. Its project fixture was a bare directory rather than a
repository, so the non-git inert branch answered before the executor list was
ever consulted. It asserted agreement it never actually measured. The fixture is
now a real repository, like every sibling in the file.

A test also asserted that Windsurf declares the worktree flag. It does not --
Windsurf resolves to no isolation by design, having no named concurrent dispatch
to isolate. The test claimed a registry fact that was never true, and a comment
in the resolver repeated it. Both corrected, and the test now proves what it
should have all along: that the parser accepts any bare flag value, rather than
one runtime's supposed value.

The new guard was missing from the bundled-hook whitelist, which is the surface
that decides what actually ships, and the per-plan gate had gained calls to the
launcher without the preamble those calls require. The changeset carried
parenthetical product descriptions the purity rule forbids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3045): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3045): make the guard tests hold on Windows

Two tests redirect HOME to control where the installer-persisted runtime default
is read from. Node resolves the home directory from USERPROFILE on Windows and
never consults HOME, so both silently read the real runner profile, found no
recorded runtime, and asserted against a project the hook had not recognised. The
production code was already correct in asking the platform rather than the
variable; only the tests were wrong to assume one variable answers everywhere.
The helpers now mirror the override onto both.

The symlink spoofing test also created a directory symlink unconditionally, which
needs elevated privileges on Windows. It survived on this runner, but it would
fail on any host without them, so the creation is now attempted and the test
skips explicitly when it cannot be done -- a bare return would have counted as a
pass and hidden the gap.

Skipping alone would have left the platform uncovered, so the behaviour it proves
is now also driven in-process through an injected realpath, following the seam
already used for the clock. That case no longer depends on privileges at all, and
the end-to-end test keeps its original assertions wherever symlinks work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 23:42:16 -04:00
Tom Boucher
c547e73a71 fix(#2927): merge installed overlay reviewer lanes into review-lane invocation (#3062)
* test(#2927): prove overlay reviewer lanes are invisible to review-lane

Failing-first regression for #2927. routeReviewLane builds its lane map from
the static REVIEWER_LANES array only, so an installed overlay reviewer lane
(role:"reviewer" capability) is roster-visible and disclosed at install but
never selectable, plannable, or invocable. The test exercises a pure
mergeReviewerLanes(firstParty, registry) helper that does not exist yet, so
every row fails at the require().

* fix(#2927): merge installed overlay reviewer lanes into review-lane invocation

routeReviewLane built its lane map exclusively from the frozen first-party
REVIEWER_LANES array, so an installed, consented third-party reviewer lane
(role:"reviewer" capability) was roster-visible and disclosed at install but
never selectable, plannable, or invocable — sections/flags/plan/invoke all
shared the one static map.

Add a pure, total mergeReviewerLanes(firstParty, registry) helper
(src/review-lane-descriptor.cts) implementing ADR-2782 D8: first-party ∪
installed overlay reviewer bodies, first-party winning on slug collision. The
overlay body is field-identical to ReviewerLane per ADR-2782 D1 ("no
translation layer"), so the helper MERGES rather than PROJECTS. Malformed
overlays (missing/non-object body, empty or grammar-invalid slug) are skipped,
never thrown — one bad third-party manifest cannot take the first-party lanes
down. routeReviewLane consults loadRegistry({includeInstalled:true}) and
degrades to the static set on any load failure.

* test(#2927): add CLI-seam coverage for the wiring defect + normalize slug

Two findings from the isolated adversarial review:

1. The test matrix's rows 9-10 (acceptance criteria #1-#3: overlay appears
   in sections/flags and plan resolves ok) were documented as covered but
   had no backing tests. The eight pure-helper tests would stay green if the
   one-line routeReviewLane wiring were reverted — the actual defect this PR
   closes had no regression guard. Add real end-to-end CLI tests that install
   a global-scope role:"reviewer" overlay and assert review-lane
   sections/flags/plan see it through loadRegistry -> mergeReviewerLanes.

2. mergeReviewerLanes trimmed the slug for the map key but stored the body
   with its untrimmed slug, diverging from deriveReviewerSlugs (which trims
   before adding to the roster). Normalize the stored lane's slug to the
   trimmed value so the two surfaces agree on the canonical key.

* test(#2927): correct CLI-seam fixtures for reviewer manifest shape

Two corrections from local CLI smoke-testing before the verification run:

1. role:"reviewer" manifests must omit feature-only fields (skills/agents/
   steps/contributions/gates/hooks/runtimeCompat) — the validator rejects them.
   Match the shipped capabilities/lm-studio shape.

2. The plan subcommand renders an ARRAY of {slug,ok,section,transport,...}
   (it strips the nested invocation plan object), so assert on the array
   element, not a top-level object. Also drop the malformed-flag-filter
   assertion: the capability validator enforces flag grammar at install time,
   so a lane with a malformed flag cannot be installed and never reaches the
   flags shape filter (which is defense-in-depth, not independently reachable).

* fix(#2927): drop unnecessary type assertion flagged by lint:ci

The `body as object` cast inside the spread is redundant — body is already
narrowed to object by the preceding typeof check. eslint no-unnecessary-type-
assertion flagged it; lint:ci is a merge gate.

* chore(#2927): add changeset fragment

pr:0 placeholder backfilled with the real PR number once the PR exists.

* fix(#2927): access runGsdTools result via .output in CLI-seam tests

runGsdTools returns {success, output, exitCode, error}, not a string. The CLI
tests (rows 9-10) passed the result object directly to JSON.parse/.split, which
string-coerced to "[object Object]" and threw under gsd-test (3 failures). My
local smoke test ran the CLI directly (string stdout), so it missed this — the
helper wraps execFileSync and returns a result object. Access .output and assert
.success explicitly, matching the established capability-cli.test.cjs convention.

* chore(#2927): backfill changeset PR number 3062

---------

Co-authored-by: sim <sim@local>
2026-08-04 18:59:19 -04:00
Tom Boucher
4eb8e3648c fix(#3050): consolidate the spawn-timeout predicate and propagate the unresolved-root reason (#3060)
* chore(#3050): changeset and review artifacts for the follow-up

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3050): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 17:21:24 -04:00
Dennis Kim
178ec00040 fix(#2787): clarify broken-windows ship blocking enforcement (#2814)
* fix(#2787): clarify broken-windows ship blocking enforcement

* fix(#2787): update renderLedger header to clarify opt-in enforcement

* fix(#2787): address maintainer scope and wording review

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:39:16 -04:00
𝚌𝚕𝚎𝚣𝚌𝚘𝚍𝚒𝚗𝚐
88f6d9bd1b fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu

* fix: preserve installer executable mode

* chore: add changeset for PR #2812

* test(#2644): acknowledge Cursor emission changes

* test(#2644): drop spent emitted drift acknowledgments

* fix(#2644): remove retired Cursor command converter

---------

Co-authored-by: clezcoding <clezcoding@users.noreply.github.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-03 12:05:45 -04:00