Commit Graph

202 Commits

Author SHA1 Message Date
Tom Boucher
b956bb7c67 fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs (#4902)
* test(#4834): failing-first launcher hijack regressions

* fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs

A gsd_run on PATH that cannot prove it is @opengsd/gsd-core (a foreign package, or a
release older than the runtime-identity verb) is no longer accepted by the launcher
snippet's PATH arm; resolution falls through to the hard error when no path-based
candidate matches. The runtime-config-home arm now precedes the PATH arm, restoring
the documented prefer-local-over-PATH order, so an installer-managed install wins
even against a genuine global. The 16-home probe list is factored into _gsd_homes()
and the identity gate into _gsd_id_ok(), keeping the per-copy delta at +141 bytes.

The files whose frozen ceilings had no headroom (gsd-executor, gsd-plan-checker,
gsd-verifier, gsd-planner, execute-phase, execute-plan) now load the resolver by
@-include from gsd-core/references/gsd-run-resolver.md (the onboard.md pattern)
instead of carrying an inline copy. Propagated to all other inlined workflow/agent
copies via scripts/sync-runtime-launcher.cjs; the resolver reference re-copied
byte-equal (parity B2); the hard-error text, docs/how-to/diagnose-a-foreign-gsd-tools.md,
and the CONTEXT.md launcher predicate updated to match (#4834); the quick-batch row-48
guard gains the canonical-preamble sweep carve-out (#4834, per its own #3730/#2529
precedents); the compact-content benchmark baseline regenerated.

Emitted-Drift-Ack-Growth: add-backlog.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-tests.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-todo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ai-integration-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-uat.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: autonomous.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: check-todos.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: cleanup.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: code-review-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: code-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: complete-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: debug.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: diagnose-issues.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: discuss-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: do.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: docs-update.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: edit-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: eval-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: explore.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: extract-learnings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: fast.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: forensics.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: graduation.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debugger.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-eval-auditor.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-eval-auditor.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-intel-updater.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-intel-updater.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-project-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-project-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-research-synthesizer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-research-synthesizer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-ui-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: health.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: import.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: inbox.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ingest-docs.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: insert-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: list-seeds.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: list-workspaces.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: map-codebase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: milestone-summary.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: mvp-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: next.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: pause-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plant-seed.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: pr-branch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: profile-user.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: progress.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: quick-batch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: quick.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: remove-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: remove-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: resume-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: scan.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: secure-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings-advanced.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings-integrations.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ship.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sketch-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sketch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: smart-entry.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spec-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spike-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spike.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: stats.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sync-skills.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: thread.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: transition.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ui-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ui-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ultraplan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: undo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: validate-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: verify-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)

* docs(#4834): backfill the changeset PR number

* test(#4834): regenerate the compact-content benchmark baseline after the rebase

---------

Co-authored-by: sim <sim@local>
2026-09-21 02:27:36 -04:00
0xdhx
88b5775dc8 enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI (#4477)
* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI

gsd-ui-auditor is chartered to audit interaction and handed a capture
driver with no interaction verb: `npx playwright screenshot` cannot
click, fill, hover, press or snapshot, so a hover state, an open menu,
a focus ring or a form's validation state never appears in its
evidence and every Experience Design finding degrades to code reading.

Implements the shape approved at triage, not a new capability:

- capabilities/ui/capability.json declares `workflow.ui_interaction_capture`
  (boolean, default false) on the capability that already owns the
  auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated.
- gsd-core/workflows/ui-review.md reads the key through gsd_run and hands
  it to the auditor as `interaction_capture:` in the spawn <config> block —
  the auditor carries no gsd_run resolver, so the key travels by value.
- agents/gsd-ui-auditor.md gains an anchored interaction-capture section
  AFTER the static block. With the key on and a Chrome binary resolved it
  starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on
  an --isolated profile, opens the dev URL the static block reached,
  takes the a11y snapshot for element uids, captures the baseline and a
  Tab focus-ring state, drives the UI-SPEC's interactive components, saves
  console output, and stops the daemon unconditionally. Key off, no dev
  server, or no Chrome: one status line, and the Playwright-only static
  path runs exactly as before — the static fence is untouched.

Needs only Bash: no MCP server, no tools: change. Chromium-only by
nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so
readiness is polled through evaluate_script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* test(#4223): bind the interaction-capture shape and containment

- manifest, generated registry, config schema and config-set/loadConfig
  all know workflow.ui_interaction_capture as a default-off boolean, and
  hand-written non-booleans fall to the slice default
- the orchestrator reads the key and hands it down; the auditor never
  grows a gsd_run dependency
- the static fence stays Playwright-only and the interaction fence
  chrome-devtools-only, so key-off is today's path
- the interaction fence runs under bash with a stub driver on PATH: key
  off / absent / no dev server / no Chrome invoke nothing; the happy path
  starts first and stops last on the [selected] pageId with the documented
  flags; a failed capture is removed and not counted; new_page and start
  failures still honour the stop-only-if-started rule; CHROME_BIN and
  CHROME_DEVTOOLS_MCP_VERSION overrides flow through
- docs/CONFIGURATION.md row shape; registered in the docs-guard lane

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* docs(#4223): document workflow.ui_interaction_capture and its how-to

- docs/CONFIGURATION.md: one row in the workflow.* table, default-off
- docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the
  interaction-capture section adds, skips and never claims
- docs/how-to/enable-ui-interaction-capture.md: turn it on, read the
  `**Interaction captures:**` outcomes, what it does not do, turn it off
- docs/README.md: index the how-to beside live-DOM verification

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* chore(#4223): add changeset

Added-type fragment; pr: carries the issue number until the PR exists.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose

Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace;
the hyphen form is retired there and the slash-command-namespace guard
rejects it. docs/ keep the hyphen form by convention.

Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered.
Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): per-run daemon session, bounded navigation, and step failures that count

Three findings from the pre-file adversarial review of the interaction
fence, folded in:

- `--sessionId <epoch>-<pid>` on every driver call. `start` restarts
  whatever daemon shares its session and --isolated isolates only the
  browser profile, so two concurrent audits — or an audit beside the
  operator's own CLI daemon — would otherwise stop each other. The CLI
  accepts hex and dashes only; the id is validated by the test stub.
- `new_page --timeout 30000`: the one verb that takes a bound, placed
  before every verb that does not, so a hung page is caught first.
- a failed take_snapshot or press_key now increments the failure count
  and is named on stdout; two clean screenshots can no longer read as
  `0 failed` after the step that gives the interactions their uids failed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal

Second review round, both reviewers:

- session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a
  subshell, so two audits forked from one parent in the same second
  shared an id and could stop each other's daemon (driven by the reviewer)
- `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver
  under Git Bash still matches the `$` anchor, and `|| true` on the
  assignment so a failed new_page cannot abort the block under
  `set -e -o pipefail` before the unconditional stop
- a failed take_snapshot removes any snapshot.txt it left or inherited
  from a reused directory, so stale uids never drive the interactions
- `<config>` placeholder is `{interaction_capture}`, lowercase like its
  `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt
  template, not a bash heredoc
- how-to: the `not captured` row no longer claims the daemon started

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges

Third review round:

- a new_page that prints a page line and then exits non-zero is a failed
  navigation, not a page id: the exit status is checked in an `if` before
  the output is parsed (driven by the reviewer against the previous
  `|| true`, which masked exactly that)
- regression cases for what the last two rounds added: CRLF driver
  output, a stale snapshot removed on failure, partial-output new_page
  failure, and the whole fence under `set -e -o pipefail` (both the
  failed-navigation path and the happy path)
- the harness whitelist gains `date`; the session-id assertion now
  requires all three parts, so a silently empty epoch cannot hide again

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder

- the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes,
  over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces;
  the interaction section's comments are tightened to the same content
  in fewer bytes (23559 now). No bash changed — the fence's own tests and
  the real-browser run are unchanged.
- .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which
  the post-create backfill rewrites to the PR's own number.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* chore(#4223): set changeset fragment pr to 4477

* test(#4223): compare the fence's status path with the separator the fence uses

On the windows-latest lane the happy-path case failed on `\interaction` vs
`/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a
literal slash, and the assertion built its expectation with path.join. Every
other case in the file passed on that lane, including the CRLF and
errexit/pipefail ones.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* test(#4223): drop the inert file-header allow-test-rule marker

Review round 1 on #4477: the `source-text-is-the-product` marker sat at
line 2, outside no-source-grep's 8-line lookahead of every readFileSync
site (the first is ~60 lines down), so it suppressed nothing. It was also
unnecessary: every read in this file is a .md/.json path, which the rule
does not trigger on. Deleted rather than relocated — there is no site to
relocate it to. Negative control: `eslint` on the file is clean without it.

* chore(#4223): regenerate the platform-conformance-tier lists for the new test

Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier
gate after this branch opened; its two committed lists must name every file
under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never
been in them. Once the branch was updated against next the lists were stale
and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier
--check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's
fragment-single-edit-propagation, which sees the same staleness as regen:derived
touching files beyond the fragment edit under test.

Regenerated with the repo's own generators, no hand-editing. The general tier
goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry;
both --check arms are clean. Verified the red is this PR's own file and not
base drift: at upstream/next both generators report "list matches" (546 / 196),
and our committed copies were byte-identical to next's before this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ

* fix(#4223): bound, confine and trap the chrome-devtools driver fence

Round 4 — three findings in one fence, interleaved on the same lines, so one
commit:

- Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a
  background job in its own process group (`set -m`) under a watchdog that kills
  the whole group at the ceiling — TERM, then KILL two seconds later. One pid is
  not enough: npm forwards SIGTERM only to its direct child, so killing `npx`
  alone leaves the client holding the fence's stdout and a `$(cdt … new_page)`
  capture blocked past the ceiling (driven against a real npx tree by the
  round's adversarial review; the pid-only first cut of this commit had exactly
  that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )`
  subshell: a subshell inherits bash's saved copies of the caller's stdio (the
  fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them
  open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a
  watchdog outlived its kill — measured as the intermittent 30 s test run the
  review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 --
  -pgid`, every 0.1 s) and stands down by itself once the group is empty;
  nothing ever signals it. The group, not the leader pid: a child can outlive
  the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll
  stood down at once and left the substitution open-ended (driven by the
  round's adversarial review at 6× the ceiling; a pgid cannot be reused while
  any member lives, which a bare pid can). The daemon `start` launches is
  spawned detached (its own session) and never in that group. Two platforms
  forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM;
  sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out
  the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git
  Bash a signal to a watchdog still starting up hung the fence's `wait` for it:
  18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs,
  while a fence slowed by xtrace, or three tests run alone, never hit it (a
  startup race; the mechanism is not pinned further). Polling: 0/300 slow calls
  and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3
  full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU,
  msys and busybox sleep; BSD sleep documents it). A clock that cannot launch
  (`sleep … || exit 0`) stands the watchdog down rather than firing at once
  and killing a healthy call — by design that leaves a hung call unbounded,
  the pre-round-4 behaviour, instead of failing a healthy one. A hung call
  returns once its group is gone: at the ceiling, plus up to the 2 s
  TERM-to-KILL grace. The KILL after the grace is sent only to a group that
  is still alive: a pgid freed during the grace can be reused, and an
  unconditional KILL could hit an unrelated group (the round's review).
  `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT
  (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent
  on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog
  rather than either.
- --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may
  write under the run's interaction/ directory and nowhere else. Relative, like
  every --filePath (unchanged from rounds 1-3): the daemon resolves both against
  one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and
  path.resolve()s both), and a relative path needs no dialect translation — an
  absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native
  daemon cannot resolve (CI's windows conformance shard caught the first cut).
  --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified),
  so the documented floor moves from ^1.8.0 to ^1.9.0, where
  --allowUnrestrictedPaths is deprecated.
- `stop` is owed by an EXIT trap after a successful `start`, not by position
  (it replaces any earlier EXIT trap — none exists in this file); the explicit
  call keeps it in order, a flag makes the trap a no-op afterwards, and only the
  shell that installed the trap may act: a subshell copy of the fence state
  carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced
  locally at 3/40 under load: the second `stop` came from a subshell pid, never
  main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo
  "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist
  and CI's macos conformance job showed the guard comparing empty to empty. The
  fence was driven under bash 3.2.57 for the injected-subshell, errexit
  failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy
  paths. A failed resize_page is a counted failed step now, not the one bare
  command an errexit runner could abort on.

Prose in the section is tightened to pay for the mechanism: 23559 -> 24517
bytes against the 24576 DEFAULT-tier cap.

Tests: the stub driver hangs as a real child tree (sh waiting on a child that
holds stdout — never an exec), so a pid-only kill fails the new
aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative-
controlled: it blocks for the harness's whole cap on the old wrapper). A hung
start and a hung capture are cut off within ceiling + grace + slack and still
reach stop; an injected bare failure under errexit reaches stop through the
trap, exactly once; an injected subshell call of cdt_stop issues nothing; the
happy path issues exactly one stop; every driver call site names a ceiling and
the only bare $CDT is the wrapper's own spawn; the start line carries
--workspace with the capture directory, every --filePath lies under it, and no
code line carries --allowUnrestrictedPaths. A driver whose leader exits at
once while a child keeps holding the capture pipe is still cut off at the
ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole
cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone
(negative-controlled: the trap form kills `start` in under 20 ms). The harness
EXPORTS its stub-only PATH — unexported, the exec'd
watchdog fell through to bash's compiled-in default PATH and never saw the stub
dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash,
pid-preserving).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

* fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file

Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the
accessibility tree, with entered form values) and console.txt (which can carry
tokens) were committable by `git add .`. The gate now ignores `interaction/` as
a directory — the next artifact type is covered by construction — and it appends
whatever an existing .gitignore lacks instead of writing once. The write-once
form was the same defect one step later: every project that had already run an
audit would never have received the new pattern at all.

Tests run the gate fence under bash: a fresh file carries every pattern; an
image-only file from an earlier audit gains interaction/ and keeps its own
header without duplicating present lines; a second run appends nothing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

* test(#4223): declare the interaction-capture anchor as a comment marker

The #4324 colon-token gate (slash-command-namespace) landed on next after this
branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an
unconvertible /gsd: command token. It is a section anchor of the same family as
gsd:live-dom-families and gsd:write-continue, so it is declared in
COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added
gates against the merged tree; CI at ca8d2508 predates the gate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-09-20 05:45:15 -04:00
Tom Boucher
9a41a95212 fix(#4717): consult the per-install runtime marker at both identity seams (#4861)
* test(#4717): add failing-first coverage for the two runtime-identity marker seams

* fix(#4717): consult the per-install runtime marker at both identity seams

resolveReportedRuntime (agent_runtime) and loadConfigResolved
(config.runtime) both ignored the per-install .gsd-runtime marker that
resolveRuntime and the model-resolver gate already read. On a
multi-runtime machine (e.g. a globally exported CODEX_HOME), host sniffing
misreported every Claude Code session as codex, and a shared
defaults.json stamped by the first non-Claude install leaked its runtime
to every other one.

Seam 1: the reported-runtime ladder becomes explicit > install marker >
host detection > claude. Seam 2: loadConfigResolved fills an empty
config.runtime from GSD_RUNTIME then the marker, copy-on-write (the
builtin-defaults branch returns a shared object). Explicit runtimes and
marker-less trees are unchanged.

* fix(#4717): a marker-detected runtime opts into its tier map (decision a)

* fix(#4717): stamped-defaults leg, marker fail-safe, docs, review fold-ins

* chore(#4717): backfill changeset PR number (4861)

---------

Co-authored-by: sim <sim@local>
2026-09-18 13:50:01 -04:00
Dennis Alexis Valin Dittrich
ad1477d659 enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap

* enhance(260903-m7p): validate configured entrypoints before success

* test(260903-m7p): require pre-success entrypoint validation

* enhance(260903-m7p): gate install success on entrypoints

* test(260903-m7p): cover configured entrypoints across runtimes

* enhance(260903-m7p): cover emitted runtime entrypoints

* fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number

- finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process
  before the new configured-entrypoint assertion throws. Without a HOME +
  config-location-env sandbox that write resolved through the ambient
  environment and landed in the developer's live ~/.gsd (confirmed absent on
  origin/next baseline, present only on this branch — full-suite HERMETICITY
  WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the
  duration of the test, matching the existing in-process finishInstall/
  install() pattern in tests/install.test.cjs (#2665).
- .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder
  (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the
  fork PR number until the upstream PR number is known.

* fix(260903-m7p): repair cross-platform and pre-existing shape fallout

- tests/configured-entrypoint-validation.test.cjs: the win32 branch of
  ensureCodexHooksJsonSessionStart writes a .cmd shim under
  <codexRoot>/hooks/; create that dir in the test (the real installer only
  calls this once hooks/gsd-check-update.js already exists) and assert the
  platform-common entrypoint shape instead of a fixed non-Windows array,
  since win32 legitimately emits two entries (cmd shim + script).
- tests/install.test.cjs: finishInstall's shared settings-json return now
  carries configuredEntrypoints/rollbackInstallerMigrations for every
  runtime on that path (trae included, not just Claude/Cursor/Windsurf);
  update the trae install() exact-shape assertion to match.

* fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash

configuredEntrypointsForHook's shell branch dropped interpreterCandidates
entirely when resolveBashExecutable returned null, unlike the sibling
portableHooks runner entry a few lines below (which correctly falls back to
the literal 'bash' token). Found via agy adversarial review; verified
unreachable through the current call graph (buildHookCommand's own
resolveBashRunner==null gate already short-circuits before
recordConfiguredHookCommand runs), so this is a defensive consistency fix,
not a live-bug patch — kept for the next caller that does not share that
gate.

* chore(260903-m7p): backfill changeset pr number to the opened upstream PR

.changeset/quick-wasps-sing.md carried the fork PR number (16) as a
placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now
open, so record its real number per CONTRIBUTING.md's changeset pr-field
convention.

* fix(#4154): track already-registered hooks for entrypoint validation on update

applySettingsJsonHooks registers each guard hook only if absent, so a hook
already present from a prior install keeps its stale on-disk command. The
new entrypoint tracker always records the freshly-computed command for it,
which never matches what is actually persisted, so the exact-string filter
in finishInstall silently dropped it from validation — the Blocker case
this feature exists to catch (an already-installed entrypoint going stale
between installs) was exactly the case it never validated.

Match on the managed script's basename instead, which the persisted
command carries either way, so an already-registered hook stays in the
validated set. Regression test forces this path by mutating a
freshly-installed hook's persisted command before a second install.

* fix(#4154): distinguish an unreadable script from a missing one

validateConfiguredEntrypoints folded an EACCES statSync failure into the
same 'missing' reason as ENOENT, misreporting a real permission problem as
an absent file. Check the error code and report 'unreadable' instead.

* docs(#4154): document entrypoint validation's rollback and PATH scope

CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had
no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite
bin/install.js x CONTEXT.md being this repo's strongest co-change pairing.

The update-gsd.md how-to overstated what a validation failure undoes: for
Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/
config.toml inside install() before the aggregate validation call runs, so
there is no rollback path for that write regardless of "where available"
phrasing. Also note that interpreter resolution checks the installer's own
PATH, not necessarily the PATH a hook fires under later (#2979 launchers).

* chore(#4154): point changeset pr field at the fork PR while CI runs there

Mirrors the branch's own prior backfill commit: pr: matches whichever PR
number changeset-lint is currently validating against (fork PR #16 during
the fork-first CI/review loop), flipped back to the upstream PR number
right before the final push to open-gsd/gsd-core.

* fix(#4249): address adversarial-review findings in entrypoint validation

An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR
found several real gaps beyond the human reviewer's Blocker, verified
against source before fixing:

- Codex's install() result bound rollbackInstallerMigrations to the narrow
  installer-migrations-only rollback instead of restoreCodexSnapshot (#3245),
  the full pre-install snapshot/restore Codex already owns for exactly this
  case — a validation failure discovered outside install() reverted nothing
  of the config.toml/hooks.json that call had already written.
- The register-only-if-absent basename match from the prior fix used a bare
  substring, which an unrelated user command mentioning the same filename
  could false-positive into GSD's validated set — anchored on the
  `/hooks/<basename>` path segment instead.
- nodeCandidates checked raw process.execPath (always true — we're running
  in that process) instead of normalizeNodePath's stable version-manager
  alias, the same one buildNodeRunnerChainToken bakes as its first choice —
  a false green regardless of whether that alias itself still resolves.
- An entry with no interpreterCandidates (Cline's PreToolUse hook, or a
  Windows-Claude .sh hook invoked without a bash runner) runs via its own
  shebang; validateConfiguredEntrypoints checked only file-type, never the
  execute bit. Cline's writer also never reported an entrypoint at all.
- Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor
  hook registered across several events) were validated once per duplicate.

Each fix is covered by a new or extended test; the Codex one required
inlining runCodexInstall's env sandboxing so the rollback closure — which
re-resolves the $HOME-relative skills root live — runs before the sandbox
is torn down, matching how installAllRuntimes' real aggregate gate calls it.

* docs(#4249): document the round-2 entrypoint-validation fixes

Runtime Hooks Surface Module and Installer Module entries now name
ConfiguredEntrypoint's not-executable reason, the normalizeNodePath
alignment, Cline's tracked hook, and which install() result the
finishInstall/installAllRuntimes rollback path actually reverts per
runtime (Codex's full snapshot vs. the others' narrow migrations-only
rollback).

* chore(#4249): point changeset pr field at the upstream PR now that fork CI is green

* fix(#4249): address agy adversarial-review findings

- validateConfiguredEntrypoints: statSync alone never detects a
  chmod-000 script (it only needs parent-dir search permission), so an
  interpreter-invoked entry with an unreadable script passed validation.
  Add an explicit R_OK check for the interpreterCandidates branch only —
  the candidate-less/shebang branch already has its own X_OK gate.
- docs/how-to/update-gsd.md: the blanket "does not revert" claim was
  false for Codex, which reverts config.toml/hooks.json via its full
  pre-install snapshot; qualify it per runtime.
- tests/codex-config.test.cjs: the #4249 rollback regression test
  asserted skills/ and VERSION were reverted but never asserted
  config.toml/hooks.json were too, despite the test's own stated intent.
- CONTEXT.md: qualify which interpreterCandidates entries get
  normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's
  Windows .cmd shim as a third candidate-less case that relies on
  extension dispatch, not a shebang.

* fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit

Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node',
so it needs the execute bit (like any shebang-invoked entry), but its
interpreter is looked up on PATH by 'env' at hook-fire time (unlike
every other GSD JS hook, which bakes an absolute node path specifically
to avoid that dependency). The candidate-less/interpreterCandidates
fork treated these as mutually exclusive, so Cline's entry silently
skipped interpreter resolution entirely — a completely missing 'node'
on PATH would still validate successfully.

Add an orthogonal selfExecutable flag so both checks run for entries
that need them. (CodeRabbit finding on the fork rehearsal PR.)

* fix(#4249): address second-round adversarial review findings (opus + agy)

- validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry,
  not just interpreterCandidates ones — a self-executable shebang script
  is still opened and read by its kernel-invoked interpreter, so X_OK
  alone never proved it was readable.
- selfExecutable is now the sole, explicit source of truth for the
  execute-bit check (every producer that needs it sets the flag) instead
  of being partly inferred from an absent interpreterCandidates, which
  Cline's hybrid entry also carries.
- The execute-bit check now skips explicitly on win32 (matching
  resolveExecutableBinary's own carve-out) instead of relying on Node's
  accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows
  machine and not a test that simulates win32 on a POSIX runner.
- bin/install.js: fixed a stale comment claiming no runtime's
  install()-time writes have a rollback path — Codex's does
  (restoreCodexSnapshot) — and added the omitted Cline to both that
  comment and CONTEXT.md's equivalent lists.
- CONTEXT.md: fixed the Cline description left stale by the previous
  commit's selfExecutable addition, and rewrote the validation-mechanism
  paragraph for clarity (writing-for-agents pass).
- docs/how-to/update-gsd.md: split an overloaded 4-clause sentence.
- Removed a fault-injection integration test that could not reliably
  exercise the real installAllRuntimes -> finalize -> rollback wiring
  without fighting the installer's own pre-registration existence
  guards; the constituent pieces remain covered individually.

* fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner

X_OK is a POSIX-only concept, skipped entirely when an entry's platform
is win32 (matching production). Two test entries omitted platform,
defaulting to process.platform — on an actual windows-latest CI runner
that silently skipped the very check they were meant to exercise,
turning 'not-executable' into a false pass. Pin platform: 'linux' so
these are deterministic regardless of which OS runs the suite.

* fix(#4249): classify EPERM the same as EACCES in statSync error handling

Windows raises EPERM (not EACCES) for a parent directory that couldn't
be traversed into — was falling through to 'missing', misreporting a
genuine permission problem as a nonexistent path.

* docs(#4249): address final CodeRabbit doc-completeness findings

- CONTEXT.md: install()'s documented result shape omitted
  configuredEntrypoints; the ConfiguredEntrypoint shape omitted
  selfExecutable.
- docs/how-to/update-gsd.md: the failure-mode sentence omitted
  unreadable and lacks-execute-permission, which the installer also
  rejects.

* fix(#4249): stop double-validating every configured entrypoint on install/update

installAllRuntimes' finalize() already runs assertConfiguredEntrypoints
once over the aggregate set; finishInstall then re-ran the identical
check per runtime in the printSummaries loop right after, so every
entrypoint paid its statSync/accessSync/interpreter-resolution cost
twice on every install and update. Add entrypointsAlreadyValidated to
skip the redundant pass specifically on that path, while leaving the
check intact for any caller that invokes finishInstall directly.

* chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there

* perf(#4249): memoize interpreter candidate resolution across entrypoints

resolveExecutableBinary walked PATH once per (entry, candidate) pair; a
typical install has a dozen-plus entries sharing the same few candidate
lists (process.execPath for JS hooks, bash for shell hooks). Cache by
(platform, candidate) so each distinct pair resolves once per validation
call instead of once per entry.

* chore(#4249): point changeset pr field at the rebased rehearsal fork PR

* fix(#4249): drop entrypoint tracking from the now-dead Codex event writer

#2586 (landed on next after this branch forked) removed install.js's
CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent
no longer runs during install or update. The ConfiguredEntrypoint records
this branch added inside it were therefore unreachable and untested. Restore
the function to its upstream shape; the entrypoints it used to report were
never collected by any caller.

* refactor(#4249): drop the revalidation bypass flag and the candidate cache

Both were this PR's own micro-optimisations over a set of roughly a dozen
entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall
gate off to save one statSync/accessSync pass; `resolvedCandidateCache`
memoised resolveExecutableBinary across entries that are already deduped by
(configPath, scriptPath). Neither is measurable, and the flag was the only
way to reach finishInstall with validation disabled. finishInstall now
always validates what it is given.

* chore(#4249): point the changeset pr field back at the upstream PR

* refactor(#4249): track settings.json entrypoints without the hooksSurface gate

The install-surface writer only tracked configured entrypoints when the
runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing
asserts that axis agrees with `installSurface`, so a descriptor that broke
the coupling would silently pass `configuredEntrypoints: undefined` and drop
that runtime out of the validation this PR adds — reintroducing the exact
'reports Done! over a broken entrypoint' failure #4154 exists to close.

Remove the dependence rather than test it: everything recorded on this path
lands in settings.json by construction, and the registered-command filter
already discards entries no persisted hook references.

* chore(#4249): put the changeset body in the documented two-part format

CONTRIBUTING.md and .changeset/README.md both show
`**<bold change>** — <symptom-led explanation>.`; the fragment was a single
unbolded sentence.

* chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there

* fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback

#3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-*
and gsd-core/VERSION. The install overwrites every other GSD-owned file too —
hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest
itself — before the entrypoint-validation gate runs, so a validation failure
left the new payload sitting on top of the restored old config.

Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims,
before runInstallerMigrations so the bytes are the true pre-install state, and
restore it from both Codex rollback closures ahead of the per-surface restores.
Files only the failed install introduced are removed, read from the manifest
now on disk. The manifest is already the authoritative record of what GSD owns,
so no second hand-written list can drift out of sync, and user-owned files are
never snapshotted or removed. Every path is confined through
resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into
an arbitrary-path write.

Non-Codex runtimes are unaffected: the snapshot is gated on the same
tomlConfigInstall + non-minimal condition as #3245's.

* fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest

Two follow-on defects in the previous commit's snapshot:

- The capture was gated on `!isMinimalMode`, copied from #3245. A core/
  --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the
  manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the
  snapshot came back empty while the rollback still ran — and its removal pass
  would have deleted every file the new manifest lists. Gate on
  tomlConfigInstall alone, matching where the rollback actually reaches.

- An unreadable or unparseable prior manifest was caught alongside ENOENT and
  treated as a fresh install. That is the same empty-snapshot state, so a failed
  update over a real install with a corrupt manifest could delete its prior
  payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means
  known-empty; any other read error or a parse failure means unknown, and the
  restore closure returns without touching anything, degrading to #3245's
  narrower rollback. Deliberately not fatal — a corrupt manifest has to stay
  repairable by reinstalling over it.

Both paths are covered by red-checked regression tests.

* fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too

commit removed from the manifest snapshot. restoreCodexSnapshot is reachable
for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-*
skill dir and gsd-* agent file the snapshot does not claim — so with an empty
minimal-mode snapshot a rollback deleted the whole skills/agents surface with
nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via
the ADR-1239 skills-kind home override, so this is also the reason manifest
`skills/` keys do not resolve under configDir: that surface belongs to this
snapshot, not to the manifest-driven one.

Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal
mode — doing nothing on an early failure is the non-destructive side.

Covered by a red-checked regression test that plants bytes in an alternate-home
skill file, reinstalls under the core profile marker, and asserts the rollback
restores it.

* fix(#4249): never remove on rollback unless a prior manifest proves what predates the install

Three defects in the manifest-driven Codex rollback, all in its removal half:

- ENOENT marked the snapshot usable, arming the removal pass on a FIRST
  install. GSD may have overwritten a user's file at a manifest-tracked path
  there, and no prior manifest records the difference — so rollback deleted it
  where before it merely left it overwritten. Absent, unreadable and malformed
  manifests now all leave the prior set UNKNOWN and skip removal entirely.

- Membership was tested against the map of files whose pre-install read
  SUCCEEDED, so a tracked file that existed but was unreadable read as
  introduced-by-this-install and was removed. Track the prior manifest's paths
  in their own Set and test against that.

- The unreachable "delete the manifest when there was no prior one" branch is
  gone: usable now implies a parsed prior manifest.

Also adds the end-to-end test the aggregate gate was missing — the four Codex
rollback tests drove the closure directly, proving the restore but not the
wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes
Cline's `env node` entry fail validation for real, and asserts Codex's payload
comes back.

Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper
instead of six copies. Both new tests are red-checked.

* test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test

lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its
Windows-EBUSY retry budget. That budget is for directory trees; this removes a
single file, which unlinkSync says more precisely and the rule does not flag.

* chore(#4249): point the changeset pr field back at the upstream PR

* fix(#4249): use an unambiguous dedup key and surface partial-restore failures

trek-e's 2026-09-08 adversarial pass flagged two findings in the new
entrypoint-validation/rollback code:
- assertConfiguredEntrypoints' dedup key already used a raw NUL
  separator (introduced in ceebb65f2d), but git/Read render NUL as a
  space, so the key looked like a plain-space join to every reviewer
  that read the diff. Replace it with JSON.stringify([configPath,
  scriptPath]) so the separator is visible and unambiguous.
- restoreManagedFileSnapshot's per-file restore catch block claimed to
  'surface the original error' but only swallowed it, matching (and
  widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add
  an actual console.warn using the existing best-effort-warning
  convention, scoped to just this PR's new function.

* fix(#4249): treat a files-less prior manifest as unknown, not known-empty

agy's gemini-3.8-flash-high adversarial pass (round 5) found and I
reproduced empirically: a structurally-valid manifest missing the
files key (e.g. {"version":1}) parses without throwing, so
Object.keys(undefined || {}) silently read as 'zero files predate
this install' instead of the UNKNOWN state the malformed-manifest
guard exists to produce. Rollback's removal pass then deleted every
GSD-owned file the failed install's own manifest listed, including
ones that predated it — the exact data loss the #4249 CodeRabbit
malformed-manifest fix was supposed to prevent, reachable through a
JSON.parse success instead of a failure. Route the shapeless case
into the same catch-all UNKNOWN path via an explicit shape check.
Regression test reproduces the deletion before the fix and confirms
the file survives after it.

Also extend restoreManagedFileSnapshot's removal-pass rmSync and
final manifest-rewrite catches with the same real console.warn
trek-e's round-4 review asked for on the per-file restore catch —
same rollback function, same operator-facing-signal gap.

* docs(#4249): correct which runtimes actually leave a written config on rollback

agy's completeness audit (round 5, holistic pass) caught this new
paragraph claiming 'for every other runtime, the configuration file(s)
already written during that update are left in place' — false for
Claude Code and other settings.json-based runtimes, whose write never
happens on failure (assertConfiguredEntrypoints runs before
finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline
actually match that description, since they persist their config file
inside install() ahead of the gate. Split the one sentence into the
three actual outcomes; matches the PR body's own accurate Before/After
wording, which this doc addition had drifted from.

* fix(#4249): clean up doc/comment mismatches and dead fields from opus review

Opus critical-code-reviewer + ponytail-review pass on the final diff:
- assertConfiguredEntrypoints carried finishInstall's old docblock
  ("Apply statusline config, then print completion message") from
  before this function was inserted between comment and callee.
  finishInstall already has its own accurate #4249 comment, so the
  stale docblock is removed rather than moved.
- checked: number on ConfiguredEntrypointValidationResult and
  error.configuredEntrypointValidation on the thrown error: the first
  had zero consumers anywhere in the repo, including its own defining
  file, and is removed. The second matches an existing repo
  convention (bin/install.js's installerMigrationRollbackFailures,
  #4249 predates this PR) of attaching structured diagnostic context
  to a re-thrown Error even before a consumer exists, so it's kept.
- finishInstall's own assertConfiguredEntrypoints call is a redundant
  backstop on the real production path (installAllRuntimes's aggregate
  call already validates the superset first), but its comment read as
  though this call alone provided the before-the-write guarantee.
  Clarified rather than removed — it's the only gate for a caller that
  invokes finishInstall directly.

* chore(#4249): split the manifest-driven rollback engine out into #4544

Issue #4154 asked the installer to consume a validation failure "through
the existing rollback mechanism, without a second transaction mechanism".
The manifest-driven rollback widening added during review (capture every
path the prior gsd-file-manifest.json claims, restore those bytes, remove
what only the failed install introduced) is that second mechanism on a
plain reading. It is a real fix for a #3245-era gap, but an independent
one, so it moves to its own bug report and PR.

Removed here:
- bin/install.js: the pre-install managed-file capture block and
  restoreManagedFileSnapshot, plus its call sites in
  _codexPreConfigRollback and restoreCodexSnapshot (99 lines).
- tests/configured-entrypoint-validation.test.cjs: the five tests that
  exercise the manifest engine.
- CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the
  widened restore. update-gsd.md again documents the #3245 surfaces only.

Kept, because it is #4154's own scope:
- the entrypoint-validation gate itself;
- Codex's install() result binding rollbackInstallerMigrations to
  restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*,
  agents/gsd-*, gsd-core/VERSION);
- the !isMinimalMode gate removal on that snapshot. Binding the closure
  to the result made it reachable for a core/--minimal install, where its
  pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot
  does not claim; an empty minimal-mode snapshot therefore deleted the
  whole surface with nothing to restore.

The surviving aggregate-failure test now asserts on config.toml, a
surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which
only the manifest engine restored.

Refs #4544

* test(#4249): cover configured entrypoints through the packed install path

#4154's scope lists install smoke coverage alongside the installer gate —
"assert representative configured entrypoints resolve for supported runtime
profiles". The gate itself (assertConfiguredEntrypoints /
validateConfiguredEntrypoints) is unit-covered by in-process install() calls;
nothing proved the property survives npm pack -> npm install -g -> install.js.

Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct
config surfaces GSD writes launch paths into (settings.json, and hooks.json +
config.toml) — run the tarball-installed installer into a throwaway HOME, then
re-read that runtime's own written config and return the new
ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a
file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check
becomes a release gate on every matrix host without workflow changes.

The scan re-derives paths from the written config instead of reusing the
installer's own entrypoint list, and test I shows why that matters: a
registration the installer never touched during a run is invisible to the
in-process gate, so the install exits 0 and only reading the config back off
disk catches the dangling launch path.

* ci(#4249): pack a publish-shaped tarball in the install smoke lane

`npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs
build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs
after `npm ci` carries no hook scripts at all — the lane has been smoking a
package that differs from the published one in exactly the artifacts the
lifecycle smoke is supposed to launch.

That went unnoticed because the lane's init runs `--local`, which registers no
statusline and therefore registers no hook whose target is missing. A
`--global` install on the same tarball exits 1 on #4249's own gate
(`gsd-statusline.js (missing)`), which is what the new configured-entrypoint
cycle performs, so without this step the cycle would report INIT_FAILED
instead of checking anything.

Build hooks before packing so the smoked tarball matches prepublishOnly. The
CLI now reports 16 configured entrypoints for claude and 1 for codex instead
of zero.

* fix(#4249): scope Codex's full snapshot restore to entrypoint failures

Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage
exception un-install a Codex install that had already succeeded and already
printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps
the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints`
call.

Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration`
scopes finalize-stage rollback to installer *migrations* ("the executor uses the
journal to restore modified paths"), and this PR's own operator-facing paragraph
in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/
agents/VERSION revert to entrypoint-validation failures specifically ("If a
script is missing, unreadable, ... For Codex, this reverts ..."). The wide
behaviour is also incoherent as a transaction abort: the same doc says Cursor,
Windsurf, Kimi and Cline keep the config they wrote inside install().

Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall
hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install
bytes — on an update, silently downgrading a working Codex install to the
previous version while the user had just been told it was Done.

Select the rollback by error kind instead. `assertConfiguredEntrypoints` already
tags its error with `configuredEntrypointValidation`, so the full snapshot
restore runs for that error (and anything downstream of it, including
finishInstall's per-runtime backstop) and the installer-migrations-only closure
runs for everything else. The codex result now also exposes that narrow closure
as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning
the full restore, so the direct-call contract asserted by
tests/codex-config.test.cjs is unchanged.

Adds a regression test that installs codex+kilo together, injects EACCES on the
Kilo permission write by monkeypatching node:fs (restored in a finally — never
chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps
the bytes the successful install wrote. Verified red against the pre-fix
unconditional path.

Cline cannot host this test: its plan is writesSharedSettings:false +
finishPermissionWriter:null, so its finishInstall performs no write and has no
non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally
(unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded
fs.writeFileSync.

* docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing

CONTEXT.md still described Codex's rollback as an unconditional bind
to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint-
validation failures via rollbackInstallerMigrationsOnly and the
configuredEntrypointValidation error tag. Caught during the round-6
PR body pass.

* fix(#4249): stop rollbackInstallerMigrations meaning its own opposite

Codex's install() result bound `rollbackInstallerMigrations` to
restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the
actual installer-migrations-only closure behind
`rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name
meant the opposite of what it says, and CONTEXT.md had to concede as much
in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure
for every runtime, matching both its name and the meaning it already has
on next, and the snapshot restore gets its own Codex-only field,
`rollbackPreInstallSnapshot`. The selection in
rollbackFinalizedInstallerMigrations collapses to one line and no longer
needs a fallback chain.

Also in this commit, all against the same rollback path:

- Correct the rollbackFinalizedInstallerMigrations comment. It read as if
  the round-6 narrowing prevented any sibling-triggered revert of a Codex
  install the user has already seen "Done!" for. It does not, and is not
  meant to: `wide` is true for ANY entrypoint-validation error from ANY
  runtime, because the aggregate gate is all-or-nothing — an invalid Cline
  entrypoint reverts Codex's snapshot, which
  tests/configured-entrypoint-validation.test.cjs's 'an aggregate
  entrypoint validation failure rolls the Codex install back (#4249)'
  asserts directly. The discriminator is the error's KIND, not which
  runtime owns the failing path. Comment and CONTEXT.md now say that.

- Name the runtime in the "Configured entrypoint validation failed" error.
  ConfiguredEntrypointInvalid already carries `runtime`; the message threw
  it away, leaving an operator of a multi-runtime install unable to tell
  whose entrypoint broke — which matters precisely because the failure can
  revert a runtime that was itself fine.

- Set `configuredEntrypoints: []` explicitly on the copilot-instructions
  early return. Every other branch states the key; this one relied on
  installAllRuntimes' `(result.configuredEntrypoints || [])` defence.
  `[]` is correct, not a workaround: every Copilot hook is an inline
  printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no
  GSD-managed script or interpreter to resolve.

No behaviour change beyond the error-message text.

* docs(#4249): narrow the smoke scan's config-surface claim to what it checks

RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is
told to launch is registered in one of settings.json / hooks.json /
config.toml, and that nothing else in a config dir is runtime
configuration. Both halves are false as stated. Cline registers its hook
at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those
names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native
[[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi),
a directory separate from Kimi's own GSD configDir — the same gap
installer-migration 007 already documents as structurally unreachable.

The scan is in fact correct for what it runs against: entrypointRuntimes
defaults to claude + codex, whose launch paths do all live in those three
top-level files. Restate the docstring at that scope, name the two known
out-of-scope surfaces, and warn that adding either runtime to
entrypointRuntimes without teaching scanConfiguredEntrypoints about its
surface yields a scan that finds zero entrypoints and proves nothing.
The entrypointRuntimes default comment carried the same overgeneralization
("every other runtime reuses one of them") and is corrected with it.

Documentation only; no code change.

* fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json

An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found
that install()'s settings-json early return for an unparseable
settings.local.json omitted configuredEntrypoints and
rollbackInstallerMigrations from its result, unlike every other branch.
rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations
unconditionally, so this branch silently dropped its own installer-migration
rollback on a later finalize-stage failure.

Completed the return shape: configuredEntrypoints: [] (matching Copilot's
equally-early no-entrypoints-yet return) and rollbackInstallerMigrations
(already in closure scope). Red-then-green regression test added.

* test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs

Same internal adversarial review (round 8): this suite's before() packed the
tarball directly, without the ensureHooksDist() guard every sibling
install-test suite (install.test.cjs, install-minimal-hooks.test.cjs,
mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run
in isolation ahead of a suite that builds hooks/dist itself, this suite's
pack would ship a tarball with no hook scripts and fail closed on
SMOKE.INIT_FAILED instead of testing anything.

* fix(#4249): refresh stale test-timings weight for the codex-config split

next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's
heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the
CI shard packer's weight table (tests/test-timings.json) was never updated:
codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the
suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block
this PR extends with its own #4249 install()-pipeline test — had no entry at
all, so the packer would silently underestimate it at the table's median
weight (roughly a 9x underestimate against its real cost).

trek-e's most recent review flagged a Windows shard timeout in-flight on
codex-config.test.cjs, plausibly aggravated by this PR's own addition to that
file before the rebase moved it. Re-measured both files locally (node --test
--test-reporter=tap, max of 3 runs, matching the table's own max-across-streams
methodology) and patched just these two entries — not a full regeneration,
which would need real multi-lane CI data this session doesn't have access to.

* fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists

next's platform-conformance-tier classifier (#4591/#4598) landed after this
branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and
tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's
gen-platform-conformance-tier --check and both the Linux and macOS conformance
suites.

* fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation

next's #4733 (landed after this branch's last rebase) replaced the static
ISOLATED_HEAVY_FILES set with a threshold derived live from
tests/test-timings.json, and pins the current derived result in
EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still
named codex-config.test.cjs, whose own weight this PR already dropped from
127783ms to 189ms (after splitting its heavy install()-pipeline blocks into
codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The
live-computed set correctly no longer includes it; the pinned expectation is
updated to match.

* fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it

trek-e's review flagged two Major gaps: the thrown error read identically
regardless of which of three real outcomes a runtime hit (nothing
persisted / snapshot reverted / config left broken on disk), and no test
proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/
Kimi/Cline — only Codex's revert path was ever asserted.

assertConfiguredEntrypoints now tags each invalid entry with its actual
consequence, mirrored from docs/how-to/update-gsd.md's existing
rollback-matrix disclosure. A new test drives the same aggregate failure
through Cline (whose own entrypoint is the one that fails) and asserts
its hook file is still on disk afterward.

* fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR

One review pass (gemini-3.8-flash-high via the antigravity review lane)
against this PR's full diff against next, findings independently verified
against source before fixing:

- Copilot's install() return object was the only one of 6 runtime branches
  missing rollbackInstallerMigrations — reachable now that this PR's own
  aggregate gate runs rollback across every result on any runtime's
  entrypoint failure, not just Copilot's own.
- buildHookCommand's unresolved-bash early return skipped track() entirely,
  so a win32 install with no Git Bash silently produced an unregistered
  .sh hook instead of the 'unresolved-interpreter' validation failure
  configuredEntrypointsForHook's own comment said it would.
- release-tarball-smoke.cjs reported a Cycle 4 install failure under
  SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing
  SMOKE.INSTALL_FAILED.
- SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's
  trailing args, which also truncated any configDir containing a space
  (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan.
  Anchored the match on the already-known configDir prefix instead of a
  generic absolute-path guess: removes the ambiguity outright rather than
  patching the character class, and stays a raw-text scan on purpose (it
  catches a writer that emits a path without registering it — a
  JSON.parse of the expected schema would miss exactly that case).

One suggested finding (test-timings.json "missing" the new test file) was
verified false — that table only holds measured CI timings, populated
after a file's first real run — and one Ponytail suggestion (a
JSON.stringify dedup key) was rejected as it would reintroduce a real, if
narrow, key-collision risk for no benefit.

* fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment

changeset-lint requires pr: to match the PR it runs on (16 on the fork,
not the eventual upstream number) — rehearsal-branch convention already
established earlier in this PR's history.

lint-docs-guard-registration's quote-pairing heuristic doesn't require the
docs/ path itself to be quoted — it flags a file once ANY quote-delimited
span containing "docs/" appears anywhere in it, alongside any real fs read
call. A comment ending "...update-gsd.md's rollback-matrix paragraph"
supplied the closing quote character (the possessive apostrophe) the
heuristic paired with an unrelated single-quoted string earlier in the
file. Reworded to avoid the unquoted apostrophe next to the path.

* chore(#4249): point the changeset pr field back at the upstream PR

Fork rehearsal (PR #16) is green; the real target for this changeset is
upstream PR #4249.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 04:23:53 -04:00
Tom Boucher
0d6bf19bf1 fix(#4600): explicit --converge overrides the convergence feature gate (#4771)
* test(#4600): pin explicit-flag-overrides-gate precedence

* fix(#4600): explicit --converge overrides the convergence feature gate

PLAN_STRATEGY=converge is set only by an explicit --converge/--cross-ai,
so gating it on workflow.plan_review_convergence made an explicit
operator flag lose to a config default and stop the run with a
question-shaped success. The config remains the default for non-flag
invocation; the flag now wins, and the step says so instead of
gate-and-exit.

* test(#4600): pin the precedence contract sentences as written

* fix(#4600): keep the precedence sentence on one line

The pinned contract phrase wrapped across a line break, so the
writer-contract assertion could not match it.

* fix(#4600): document the flag-overrides-gate precedence on user surfaces

commands/gsd/autonomous.md and docs/COMMANDS.md still said --converge
requires workflow.plan_review_convergence=true; both now state the
override. Changeset typed Changed with the docs update alongside.

* docs(#4600): backfill changeset PR number

* fix(#4600): restore the convergence gate mention in user surfaces

* test(#4600): pin the dispatched convergence run against the config gate

* fix(#4600): override the convergence gate on the dispatched run

Emitted-Drift-Ack-Growth: autonomous.md — #4600: the converge dispatch appends --override-gate inside the PLAN_STRATEGY conditional, and the precedence sentences replace the stale fail-fast instruction
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4600: config gate 1.5 honors an explicit --override-gate dispatch (token-anchored) while the standalone veto and config-get default are preserved

---------

Co-authored-by: sim <sim@local>
2026-09-15 15:53:42 -04:00
Michel Moreira
2f0e99f9e0 fix(#4377): opt in to project-relative includes for local installs (#4425)
* enhance(#4377): opt-in project-relative includes for local installs

A local install wrote the includes that point at GSD's own files as absolute
paths — whatever the installer resolved at install time. For one checkout
that is invisible. Across git worktrees it is not: each worktree gets its own
.claude/ copy, but all of them point back at the checkout that ran the
installer, so a worktree runs its own gsd-tools.cjs while reading workflow
prose from a different checkout. Update that one checkout and every other
worktree is running new instructions against an old engine, with nothing to
stage the update with.

--relative-includes (or GSD_RELATIVE_INCLUDES=1) makes a local install emit
`@.claude/gsd-core/...`. Opt-in, and staying opt-in: absolute works for a
single checkout, which is most people, and flipping the default would change
every existing local install to solve a problem those users do not have.

The prefix is the runtime's own localConfigDir descriptor value, never a
literal — the same value resolveScope joins onto the cwd to produce the
install target, and the same one the rewrite engine already uses for its
./.claude/ -> ./<dir>/ substitutions. Copilot and Antigravity have shipped
this shape for local installs since they were added, with hardcoded .github/
and .agents/; this is that behavior, derived rather than written down.

Six seams compute a path prefix and all six had to be threaded, which is why
the opt-in travels through the environment the way --portable-hooks already
does: one variable they all read cannot fall out of sync the way six
signatures can.

The launcher shim deliberately keeps its ABSOLUTE fallbacks. It probes
gsd-tools through ${CLAUDE_CONFIG_DIR:-$HOME/.claude} and one such default
per runtime; those are shell word expansions, not includes, and a relative
value there resolves against the shell's cwd rather than the project.
Trading an include that points at the wrong checkout for a path that points
at nothing is not a fix. All three rewrite paths mask ${VAR:-default} spans
before substituting and restore them after, and the mask only runs when the
prefix is relative, so an absolute install is byte-for-byte unchanged.

Every unexpressible case falls back to absolute: no opt-in, a global install,
a missing dir name, the configHome.kind === 'none' sentinel, an absolute
descriptor value, or one climbing out of the project with '..'.

* chore(#4377): add changeset for project-relative local includes

* fix(#4377): compare against POSIX-normalized roots in the install e2e arms

The emitted prefix is POSIX-normalized by design — it is substituted into
markdown @-references, which use forward slashes universally, so a backslash
would leak into shipped content (#1615). The e2e arms compared against the
raw temp root, which on Windows is `D:\a\...` and appears in no emitted file.

That reddened the control arm on the windows shard, and it was worse than a
red: the negative arm ("nothing references the checkout") was passing
VACUOUSLY there, because a string that cannot occur is trivially absent. Both
now go through the same normalization, so the Windows lane asserts what the
Linux lane does.

* fix(#4377): tolerate a resolved temp root, and make the e2e diff self-diagnosing

Two changes, one confirmed and one to stop guessing.

Confirmed: the emitted content carries the RESOLVED root, not the spelling
mkdtemp handed back. Reproduced on Linux with a symlinked install root —
236 emitted files carry the realpath, zero carry the link path. macOS has
this structurally, since /var is a symlink to /private/var. Comparisons now
go through both spellings, or the negative arms pass vacuously: "nothing
references the checkout" is trivially true when the string being searched
for cannot occur.

Not confirmed: the macOS shard reported ~every workflow file differing in
the "differ ONLY" arm while the five arms around it passed, and the
assertion printed a list of filenames — which says a difference exists
somewhere across 236 files and leaves the reader to guess which bytes. I
cannot reproduce that platform locally, and guessing turns one CI round-trip
into four. The assertion now reports the first divergence as text: the file,
the byte offset, and a bounded window of both sides.

* fix(#4377): strip the longest root spelling first in the install e2e diff

The macOS failure was my test corrupting its own comparison, not a product
defect. /var/folders/…/X is a SUBSTRING of /private/var/folders/…/X, so
stripping the unresolved spelling first matched inside the resolved one and
left the /private prefix glued to what followed:

  @/private/var/…/X/.claude/gsd-core/…  ->  @/private.claude/gsd-core/…

a string present in neither install, which is why all 236 files "differed".
Sorting the spellings longest-first consumes the whole occurrence, and the
short form then has nothing left to match. Proven in isolation on the exact
macOS shapes: short-first yields @/private.claude/…, longest-first yields
@.claude/….

The self-diagnosing assertion added in the previous commit is what found
this — it named the file, the byte offset, and printed both sides, so the
corrupted string was visible rather than inferred from a list of 236
filenames. Keeping it.

* fix(#4377): address review findings

* test(#4377): scan nested shell defaults without regex backtracking

* fix(#4377): close relative include review gaps

* fix(#4377): preserve root-target runtime includes

* fix(#4377): guard project-root relative includes

* test(#4377): normalize Cline fallback roots

* fix(#4377): persist relative include style

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 03:40:04 -04:00
Tom Boucher
eb49ff98df fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime

#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.

The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:

  - install-on-your-runtime.md  English has NO `### Gemini CLI` section  -> deleted
  - USER-GUIDE.md :843          "…, Antigravity CLI, Kilo)"              -> substituted
  - ARCHITECTURE.md             English has NO Gemini CLI table row      -> row deleted
  - ARCHITECTURE.md :24         English holds `Kimi CLI` in that slot    -> Kimi CLI
  - context-monitor.md :3       "`AfterTool` for Antigravity CLI"        -> substituted
  - spike-and-sketch.md :93     "(Codex, Antigravity CLI, etc.)"         -> substituted
  - configure-model-profiles    "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
  - COMMANDS.md                 English keeps only hyphen + Codex bullets -> colon bullet deleted
  - FEATURES.md                 source docs/features/multi-runtime-support.md:10
                                lists no Gemini CLI                       -> name removed

ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.

The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.

The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.

Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.

Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.

Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.

Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.

Fixes #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate

A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.

1. I committed the exact error I claimed to have avoided. The commit message
   boasted that ARCHITECTURE.md:24 proved the value of reading English rather
   than substituting blind, because Antigravity already appeared later in that
   list. Five hundred lines further down the SAME four files, my
   `Gemini:` -> `Antigravity:` substitution produced TWO consecutive
   `- Antigravity:` bullets, because an Antigravity bullet was already there.
   English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
   locales, reusing each locale's existing words.

2. `--gemini` survived in the runtime-detection CLI flag list in all four
   locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
   already lists `--antigravity` later, so this is another place where
   substituting Antigravity would have duplicated it. Now `--kimi`.

3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
   line I had already corrected -- the "Adaptive (Recommended)" option in
   settings.md:192 and new-project/steps/auto-mode-config.md:95.

4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
   `non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
   lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
   without the word "runtimes", and execute-phase.md:1028 has no parenthetical
   at all. The reviewer proved it by re-introducing Gemini at both lines and
   watching the assertion stay GREEN. That same blind spot is what hid finding 3.

   Replaced with a case-sensitive `/\bGemini\b/` walk over every
   `gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
   reference in that tree is spelled differently and cannot match: Antigravity's
   paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
   are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
   uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
   `Gemini` there means the retired RUNTIME is being named. The walk asserts it
   found at least 50 files so an empty walk cannot pass vacuously, and it now
   covers the nested `new-project/steps/` directory where finding 3 lived.

   Two allowlist entries, both by line CONTENT and both justified:
   reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
   `Known provider` menu. The second was escalated by the agent rather than
   decided: Section 8 of that file says model policy is defined "independently"
   of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
   the same axis as the lowercase model ids -- and must keep working.

   Proven to fail, not just asserted: the predicate reports 0 offenders on the
   real tree and exactly 2 on a /tmp copy with Gemini re-injected at
   health.md:52 and execute-phase.md:1028.

Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.

The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.

Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @ ca8d9d4459) by comparing each tracked file's blob size, which
is one more file than that run reported -- the extra is settings.md, grown again
by fix 3. docs-update.md and map-codebase.md are deliberately NOT acked: they
SHRANK, since there the fix deleted ", Gemini CLI" rather than substituting, and
acking a file no delta consumed is itself an error.

Refs #4728

Emitted-Drift-Ack-Growth: add-tests.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: add-todo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ai-integration-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: check-todos.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: cleanup.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: complete-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: do.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: eval-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-plan.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: health.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: import.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: inbox.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: manager.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: note.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: onboard.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: plant-seed.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: profile-user.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: quick.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: remove-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: secure-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: settings.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ship.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: smart-entry.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: undo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: update.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: validate-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: verify-work.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4728): add the changeset fragment

The PR body claimed one was present and it was not — caught by
scripts/changeset/lint.cjs reporting fail_missing_fragment, not by the
checklist, which is exactly why the lint exists.

Type Fixed: the diff is prose, and a docs-only fix uses Fixed since there is
no Documentation type.

Refs #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:49:52 -04:00
Tom Boucher
b54c1c5848 fix(#4709): retire the Gemini CLI reviewer lane (#4716)
* fix(#4709): retire the Gemini CLI reviewer lane

Google stopped serving Gemini CLI for the free/Pro/Ultra tiers on 2026-06-18 —
the same sunset that removed the gemini RUNTIME in #1928 (shipped 1.8.0). GSD
targets solo developers, so those tiers ARE the user path: the lane spawned
`gemini {{model}} -p -`, a binary that no longer answers for the majority of
users, and five locales documented it as a supported choice.

The lane was re-created after #1928 by the reviewer-lane-as-manifest-data work
(6a9babda69, #2798/#2837, ADR-2782). Per the maintainer that re-creation was an
error in that buildout rather than a considered decision, so this corrects a
mistake and needs no ADR-2782 amendment.

Reviewer roster: 12 lanes / 13 flags -> 11 lanes / 12 flags.

TWO sources of truth had to be removed, not one. Deleting
capabilities/gemini/capability.json left the capability registry at 11 lanes
while src/review-lane-descriptor.cts's hand-maintained REVIEWER_LANES array
still carried its own complete gemini entry at 12 — precisely the disagreement
checkReviewerLaneParity exists to catch. Both are gone; both parity checkers
now run clean against the real tree (lane parity ok/0 violations, docs parity
0 violations).

Surfaces stripped of the dead flag:
- capabilities/gemini/ deleted; registry and capability-matrix regenerated
- src/review-lane-descriptor.cts: REVIEWER_LANES entry, docblock count, and the
  three doc comments that used --gemini as a live example
- commands/gsd/{review,plan-review-convergence,autonomous,progress}.md and the
  four matching skills/*/SKILL.md: argument-hint frontmatter and flag bullets
- gsd-core/workflows/help/modes/{full,full.compact}.md: /gsd-help signatures,
  the detected-CLI list, and the reviewer-title list
- gsd-core/workflows/settings-integrations.md: the integrations wizard no longer
  offers "Gemini" as a model option, and the settable-keys list drops it
- gsd-core/workflows/review.md: the `command -v gemini` probe, the --gemini
  flag, the roster frontmatter, the install pointer to the sunset repo, and the
  jq-less / precedence / self-skip lane lists
- gsd-core/workflows/sync-skills.md: "two runtimes (grok, gemini) resolve to
  ANOTHER runtime's skills root" is now one runtime; gemini never aliased
  anything, it fell through canonicalizeRuntimeName to a fail-closed default
- docs/{CONFIGURATION,COMMANDS,CLI-TOOLS}.md, docs/reference/capability-matrix.md,
  docs/how-to/set-up-cross-ai-review.md — including its `npm install -g
  @google/gemini-cli` instruction and the two rows recommending --gemini
- docs/features/{cross-ai-peer-review,opt-in-parallel-reviewer-lanes}.md as the
  generator inputs behind docs/FEATURES.md, plus the three locale FEATURES.md
  signature lines the docs-parity gate covers (the #2781 class: a flag change
  that never reaches the mirrors)

Counts reconciled against measurement rather than arithmetic: 8 timeout keys of
11 lanes, 11 budget keys, 9 model keys, and four hardcoded literals in
tests/reviewer-lane-declarations.test.cjs (NEW_LANE_ONLY_IDS 5->4, LITERAL_ROSTER
12->11, two roster counts 12->11).

BEHAVIOR CHANGE, accepted deliberately: `gsd config-set review.models.gemini`
now errors with "Unknown config key". An existing key already in
.planning/config.json still parses and is simply never read, so no project fails
to load. This is the repo's own documented policy for exactly this case
(docs/CONFIGURATION.md:327 — "a key left over from a removed reviewer validated
silently and was never read. Such a key is now rejected by config-set"), so no
installer migration ships. Note my first measurement of this was WRONG: I tested
config-get, which reads undeclared keys fine, and generalised. Read and write are
different surfaces and gave different answers.

Antigravity is untouched throughout — its --antigravity/--agy flags,
review.models.agy, ~/.gemini/antigravity configHome, ~/.gemini/config global
skills root (#3738), hookEvents "gemini", GEMINI.md instruction file, and every
gemini-* model id it actually runs on.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): changeset for the reviewer-lane retirement

Type Removed: the --gemini flag and its three config keys are user-visible
surface that no longer exists.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4709): close the 24 test failures and the locale-doc gap the gates found

An adversarial review and a full matrix run between them found substantially
more fallout than inspection had. All of it is this PR's own, and all of it is
fixed rather than waved off.

THE MATRIX RUN FOUND 24 FAILURES ACROSS 6 FILES. Inspection had predicted two.
The dominant class was a test helper that looks up a lane by slug and throws
`no declared lane 'gemini'`:

- tests/feat-2483-review-claude-mds-guard.test.cjs (6) — used gemini as the
  "other declared first-party lane" to contrast against claude's env
  suppression. Now qwen, verified from source as a lane that declares no `env`
  (only claude does), so the contrast still holds.
- tests/review-lane-descriptor.test.cjs (6) — the duplicate-flag and
  duplicate-section fixtures deliberately COLLIDED with a real declared lane to
  prove the parity checker reports a duplicate. `--gemini`/`Gemini` no longer
  collide with anything, so the checker reported
  `descriptor_lane_not_in_registry:acme` instead and the tests proved nothing.
  Now collide with `--codex`/`Codex`, reproduced against the real checker.
- tests/review-reviewer-selection.test.cjs (3) — these distinguish KNOWN-but-
  undetected from UNKNOWN. gemini flipped categories, inverting what they
  proved. The known case now uses qwen; `__nope__` stays the unknown fixture.
- tests/review-default-reviewers-resolution.test.cjs (2), and
  tests/settings-integrations.test.cjs (3) — the wizard now offers three
  reviewer CLIs, not four, so the test and its name say three.
- Two count assertions the earlier sweep missed outright:
  reviewer-lane-declarations.test.cjs:359 (`length, 12`) and
  reviewer-docs-parity.test.cjs:681 (`>= 12`).

THE LOCALE-DOC GAP, and why the parity gate stayed green over it. All four
locale mirrors still documented `--gemini` as a live reviewer flag. The
docs-parity checker asserts the PRESENCE of every current flag and never the
ABSENCE of a retired one, so "0 violations" was never evidence those files were
clean — my earlier reading of it as such was wrong. This is the #2781
locale-drift class in the opposite direction. Fixed across 12 locale files:
COMMANDS.md flag lists and table rows, CONFIGURATION.md `review.models.gemini`
rows and reviewer prose, CLI-TOOLS.md config examples, and
set-up-cross-ai-review.md including its install block and its
which-reviewer-to-choose row, which now recommends Antigravity.

ALSO FOUND, and instructive about my own method: docs/CONFIGURATION.md:297 still
carried a `review.models.gemini` row. My sweep had missed it because my grep
excluded lines matching `gemini-[0-9]` to spare Google's model ids — and that
row's example value is `"gemini-2.5-pro"` on the same line. The exclusion built
to avoid false positives created a false negative.

Remaining comment/example sites: src/review-reviewer-selection.cts:309 and
src/config.cts:598 named the dead flag and key as examples;
gsd-core/references/planning-config.md:269 likewise; and
review-reviewer-selection.cts:22 claimed in the PRESENT tense that gemini is a
lane-only reviewer capability. Line 38 of that same docblock says "Before this
phase the five non-runtime reviewers (gemini, ...)" and is left exactly as is —
that is past-tense history, and rewriting it would falsify the record.

Deliberately still deferred to Phase 4, because it is the RUNTIME axis rather
than the reviewer lane: the locale install-on-your-runtime.md `--gemini --global`
instructions, the USER-GUIDE colon-form notes, and the ARCHITECTURE
runtime-detection flag lists.

Both parity checkers green against the real tree; lint:ci exit 0.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4709): backfill the changeset PR number

pr: 0 -> 4716, now that the PR exists. Never guessed ahead of the number.

Refs #4709

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 03:03:44 -04:00
Tom Boucher
181c4c8659 chore(#4603): retire the test-full CI job (#4604)
* chore(#4603): retire the test-full CI job

Phase 2 (#4591) added test-conformance but left test-full (the pre-existing
full-suite Windows/macOS replay) running unchanged, gated on the same
full_matrix flag, downgraded only from a hard gate to a non-blocking
::warning:: -- framed as "a non-gating safety net for one release cycle."
No phase or issue ever retired it. Result: every full_matrix=true PR ran
10 OS-specific jobs (test-full's 6 + test-conformance's 4, purely
additive) instead of the original 6 -- the epic's own goal (reduce
runner-minutes) was measurably regressing, not improving, for the
majority of PRs.

This phase was missing from the original 4-phase epic decomposition; the
epic (#4589) has been amended to add it as Phase 5 (see its comment
thread), and this issue was filed as the tracked sub-issue.

Deletes the test-full job from .github/workflows/test.yml entirely, along
with every reference to it: required-tests' needs/FULL_TEST_RESULT
warning branch, ci-timeout-report.cjs's JOB_RULES entry,
ci-test-job-timeout-budget.test.cjs's LANE_COSTS/staticLanes/testFullRule
entries, ci-test-scope.test.cjs's test-full-specific tests (preserving
three unrelated tests that were nested in the same describe block, moved
under a renamed describe rather than deleted), and docs mentions.
test-conformance is now the sole gating signal for real-OS coverage.

Two separate defects found and fixed while auditing every test-full
reference:
- tests/ci-pr-mergeability.test.cjs's GATED['test.yml'] safety-critical
  array (jobs that must needs: the mergeability preflight) had test-full
  but was missing test-conformance entirely -- Phase 2 never added it.
  Verified the real workflow wiring was already correct (test-conformance
  does have needs: [changes, preflight]); this was a test-coverage gap,
  not a live defect. Fixed by swapping the array entry.
- docs/TESTING-SUITES.md's "## CI matrix" section was substantially stale
  independent of this phase (predating even #2952's coverage-gate split).
  Rewritten against the real, current job topology, verified directly
  against test.yml rather than trusted from memory.

An isolated code-review pass found and fixed two minor inaccuracies in the
rewritten docs table (two jobs' "Gated on" column didn't match their real
if: condition exactly). An isolated security-review pass found no
qualifying findings -- every compute-provisioning job already carries
needs: preflight directly, unaffected by this deletion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(ci): isolate 7 more heavy test files from chunk-weight packing

`next`'s own push-triggered Tests run failed: `conformance test
(windows-latest, 24, shard 2/3)` chunk 3/6 was killed after 600019ms.
Root cause: state.test.cjs (weight 21.35, measured) was packed alongside
companions by run-tests.cjs's LPT chunk packer, the same failure mode
that previously hit codex-config.test.cjs (weight 17.87) twice and got a
dedicated fix (ISOLATED_HEAVY_FILES, #4497) -- but state.test.cjs was
never added to that set.

This is a direct, unintended consequence of epic #4589 Phase 2: the new
platform-conformance-tier job packs only ~546 files per shard (vs. the
~950-file full suite the packer used to balance against), so the same
absolute-weight outlier now represents a larger share of a smaller, more
homogeneous pool -- the LPT packer has fewer light files to pad around
it with. This was a real, foreseeable side effect of shrinking the
packing pool that nobody checked for when Phase 2 shipped.

A first attempt at this fix hand-picked 4 candidates by eyeballing a
truncated weight list and missed 3 heavier ones -- caught by an isolated
code-review pass (blocker: emitted-attribution.test.cjs at 66.2% of the
Windows chunk budget, install-minimal-hooks.test.cjs at 61.1%,
install.test.cjs at 47.1%, all above codex-config.test.cjs's own
44.7% -- the ratio that already proved dangerous twice). Corrected by
systematically computing weight/budget for every unit-suite file and
isolating everything at or above that same ratio: 7 files total, plus
the pre-existing codex-config.test.cjs (8 total).

Added a durable regression test (tests/run-tests-harness.test.cjs) that
re-derives this exact computation from the live tests/test-timings.json
on every run, so a future heavy file crossing this threshold fails the
test instead of silently reintroducing this failure -- not just a
one-time manual sweep.

Verified end-to-end: simulated the real 3-way windows shard split of the
actual conformance-tier file list with the real packing functions. Max
packable-chunk weight across all 3 shards is now 27.04 / 24.10 / 23.91
(shard 2 is the exact shard that failed on next), comfortably under the
40 budget -- versus 40+ and a 600s kill before this fix.

A second isolated code-review + security-review pass on the corrected
diff found nothing further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 12:46:04 -04:00
Tom Boucher
bcd99696d3 chore(#4591): add platform-conformance-tier classifier + gate CI on it (#4598) 2026-09-10 08:36:13 -04:00
Tom Boucher
5823d2ec7a docs(#4484): correct native-plugin-install's install-time-config parity claim (#4579)
* docs(#4484): correct native-plugin-install's install-time-config parity claim

The doc claimed the plugin path and the npm installer "differ in
namespace and lifecycle only." False: the native plugin path
(claude plugin install, marketplace discovery, and the skills-dir
zero-friction load) materializes the repository tree directly and never
runs GSD's install engine, so install-time config baked into generated
artifact files at install time never applies there -- confirmed for
agent_tools (#4238/#4032, reproduced live in #4484: 35/35 files granted
via npm install, 0/35 via plugin install, even after
`claude plugin update`). model_overrides is the same architectural class
(install-time-only logic on the npm-install call tree, per
src/install-model-override-resolver.cts) but hedged, not claimed
confirmed, matching the issue's own hedging.

Reporter explicitly frames this as a docs-only fix: the code behavior
(zero install step on the plugin path) is presumably intentional design;
the bug is the doc's incorrect parity claim, not the missing
functionality. No code changed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4484): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 10:57:28 -04:00
Tom Boucher
fd4aac5670 fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime

Two documented model-configuration contracts did not hold on the claude
runtime (confirmed-bug scope from the issue triage):

Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of
resolveModelInternal gated runtime-aware tier resolution on
configRuntime !== 'claude', so the key's only reader was never consulted,
while workflows/settings-advanced.md writes it for claude-runtime users.
A new step 4.5 resolves ONLY the user's override entry (never the builtin
claude tier map, so unpinned installs keep resolving aliases). An
override value that maps to a current tier alias collapses to that alias
(byte-equivalent, the #2041 protection); anything else — a pinned older
generation, a bare alias repoint, a non-Anthropic id — resolves verbatim.
It sits after the resolve_model_ids:'omit' gate so an explicit project
omit still wins (#2297) and before the alias return so
resolve_model_ids:true cannot re-materialize the pin to the latest id.

Finding 2 — fully-qualified claude-* ids in model_overrides were
warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable
branch, #2041), while the docs promise any fully-qualified model id is
valid. The unmappable branch now passes the pin through verbatim with a
warn-once breadcrumb (text describes the pass-through). Dropping it
silently unpinned the operator's explicit choice — the exact 'profile
can misrepresent what actually runs' defect of #4192. Mappable ids and
non-claude values behave exactly as before; resolveModelForTier shares
the mapping; the tier honesty signal is unchanged (raw ids still report
'unknown'); the model_policy path is untouched.

Docs updated to the agreed contract (CONFIGURATION.md false 'Claude
example' corrected; how-to + shipped reference document the pin
semantics, the fable alias, and the tier-override composition).

* test(#4192): pin explicit model pin resolution on the claude runtime

28 failing-first rows across the resolver seam and the resolve-model CLI:
pinned-generation fidelity (tier override + per-agent verbatim pins,
object form, explicit runtime), unpinned controls byte-stable (no
override, other runtime/tier, inherit, project omit, precedence),
adversarial rows (prototype-chain keys, malformed values, warn-once
dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through
runGsdTools. The stale #2041 fall-through assertions now pin the
pass-through contract; mappable-id collapse assertions unchanged.

* chore(#4192): add changeset fragment

* chore(#4192): backfill PR number in changeset fragment

---------

Co-authored-by: ZCode <zcode@localhost>
2026-09-06 10:17:50 -04:00
Tom Boucher
03738824de enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) 2026-09-06 05:46:49 -04:00
Tom Boucher
f9f72cb54c enhance(#3777): opt-in concurrent per-plan planners in chunked mode (#4346)
* test(#3777): add failing-first coverage for concurrent per-plan planner dispatch

Extracts and executes the real bash blocks this PR is about to add to
plan-phase.md and chunked-planning-mode.md (CHUNKED_PARALLEL resolution and
the BATCH_PLAN_IDS dedup guard), plus config-set/config-get coverage for the
new planning.chunked_parallel key. Expected RED against the current shipped
workflow text — the extraction anchors do not exist yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3777): dispatch chunked mode's per-plan planners concurrently within a Wave

Adds opt-in planning.chunked_parallel (default false, byte-identical to the
existing serial loop). When true and the runtime's negotiated dispatch
capacity (dispatch-capacity, #3673) is greater than 1, chunked planning's
per-plan Tasks that share one outline Wave are issued together instead of
one at a time; a later Wave still waits for the current one to be verified
on disk and committed. A host with no declared maxConcurrency (most
non-Claude runtimes today) stays serial regardless of the setting.

Resolution and the Plan-ID dedup guard live in chunked-planning-mode.md
itself (gated on the section's own CHUNKED_MODE skip-check) rather than in
plan-phase.md, so a non-chunked run pays no extra gsd_run calls.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#3777): repoint extraction at chunked-planning-mode.md after the move

CHUNKED_PARALLEL resolution moved out of plan-phase.md into
chunked-planning-mode.md itself (see the preceding commit); update the
test's extraction path and header comment to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): relocate the canonical runtime-launcher preamble before its first use

The CHUNKED_PARALLEL resolution block's two gsd_run calls landed earlier in
the file than the sole existing preamble (in the commit step), which
tests/runtime-launcher-parity.test.cjs's (B) check requires to precede every
gsd_run call in the file. Move the preamble (not duplicate it) to the top of
the resolution block; the commit step's fenced block now just calls
gsd_run directly.

Caught by the GREEN checkpoint gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3777): strip the canonical preamble from the extracted resolution block

The CHUNKED_PARALLEL resolution fence now carries the relocated
runtime-launcher preamble as its first line (previous commit). Extracting
the whole fence and running it after the test's own gsd_run stub let the
embedded preamble's own resolver logic `unset -f gsd_run` and exit 1 before
reaching the resolution logic, since no real gsd-tools.cjs exists in the
temp script dir — every test calling runChunkedParallelResolution() failed.

Strip the preamble (sourced from gsd-core/workflows/_runtime-launcher.snippet.sh,
the same file scripts/sync-runtime-launcher.cjs treats as canonical) before
splicing in the stub, so this suite tests only the resolution logic it is
actually about.

Caught by the post-rebase gsd-test run before push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): add the How-To page the phase gate requires

Enablement is 2 commands (config-set, then --chunked), which this repo's
own doc-quadrant gate flags as how-to-owed: a reference table cannot carry
a sequence. Covers enablement, the dispatch-capacity gate's honest
"most runtimes today: no effect" case, and the two accepted trade-offs.

An earlier reasoning pass (recorded in .gsd/phase/.../70-docs.json before
this commit) had incorrectly claimed #3034 shipped with no equivalent
how-to page, as precedent for skipping one here. That claim was false —
docs/how-to/enable-parallel-reviewer-lanes.md exists and is indexed. The
phase gate caught the omission before merge; corrected here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3777): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:57:17 -04:00
Michel Moreira
86b745b48b fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 14:20:26 -04:00
Dennis Alexis Valin Dittrich
5869febb16 enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change

readVerificationStatus() now recomputes a deterministic sha256 fingerprint
over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY,
requirements, implementation files in the verified change set) and returns
stale on any mismatch, fail-closed when a covered file is missing,
unreadable, or escapes the project root. Legacy reports with no fingerprint
metadata keep the prior SUMMARY-mtime staleness check unchanged.

The verifier computes covered_digest via the new verification.fingerprint
CLI command rather than by hand, since a digest is deterministic math, not
an LLM-estimated value.

* chore(#4155): backfill fork PR number in changeset

* fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap

* fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior

Partial fingerprint metadata (one of covered_files/covered_digest present,
the other missing or malformed) now fails closed to stale instead of
silently downgrading to the legacy mtime-only check. computeCoveredDigest
also canonicalizes with realpathSync before re-confining, so an in-root
symlink whose target escapes the project root can no longer produce a
matching digest. gsd-verifier.md restores the completeness requirement and
checklist item trimmed by the earlier size-budget fix, within the LARGE
tier byte cap.

* chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap

* fix(#4155): address gemini adversarial review findings

computeCoveredDigest now threads the caller-supplied opts.fs seam through
its confinement and read paths instead of always using raw node:fs — a
caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP
2, #2790 follow-up) was silently bypassed for covered-input reads. The
project-root anchor itself still canonicalizes through real fs (it is a
trusted value the caller derived, not attacker-influenced covered-input
data); only per-file candidate reads go through the injected seam.

Covered-file paths are now canonicalized (./ prefixes, redundant slashes,
internal .. segments) before becoming dedup/sort/hash keys or confinement
subjects — closes both a spurious-stale false positive (two spellings of
the same file hashing differently) and a confinement gap (an internal ..
segment that doesn't start the string).

gsd-verifier.md now states covered-file paths are project-root-relative,
not phaseDir-relative, closing an ambiguity that would have made a real
verifier agent's first fingerprint invocation fail closed.

defaultFsImpl's methods now late-bind through fs.<method> rather than
capturing function references at module load — the earlier direct-capture
form was invisible to existing tests' t.mock.method(fs, 'statSync', ...)
seams, a real regression caught by the full suite (not the reviewer).

* fix(#4155): catch a plan/summary added to the phase dir after verification but never declared

The content digest only recomputes hashes for paths the verifier actually
declared in covered_files — it had no way to notice a plan or summary
added to the phase directory after verification if that new file was
never declared, silently regressing behind the legacy mtime check it
replaces (which scans the live directory, not a declared list).

findUncoveredCurrentArtifact re-scans the live phase directory for every
current *-PLAN.md/*-SUMMARY.md and requires each to be represented in
covered_files, closing that gap; a directory scan failure fails closed to
stale rather than silently skipping the check.

CONTEXT.md's Verification Module entry corrected to describe the
fingerprint path's stricter fail-closed FS-error contract (routes to
stale) instead of the module's original degrade-to-safe one (missing /
not-stale), which only the legacy path still keeps.

* refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test

computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/
sorted covered_files independently — one shared helper now backs both
(gemini review's ponytail-lens finding).

Adds one CLI-to-readVerificationStatus test against a genuine
.planning/phases/NN-x/ project with an implementation file outside
.planning/ entirely, closing the review finding that prior #4155 unit
fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot
falls back to phaseDir itself there) and never exercised real multi-level
path resolution.

* fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/

Two independent review rounds (opus critical-reviewer + opus ponytail +
agy, run twice) found two instances of the same fail-open class:

- computeCoveredDigest's per-file reads routed through the caller's
  injected fsImpl. planning-inspect.cts passes a `.planning/`-confined
  containment fs into readVerificationStatus's opts.fs, so any covered
  implementation file outside `.planning/` (mandatory per the issue)
  made the confinement wrapper throw, which was caught and turned into
  a stale digest -- reporting every fingerprinted phase permanently
  stale via `planning.inspect`, regardless of actual drift. Per-file
  reads now always use real node:fs, matching the pre-existing
  treatment of root canonicalization; the realRel-vs-realRoot check is
  the real confinement boundary for this data and needs no seam.

- allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans
  reports readdir failures via a `scope` field, it never throws), so
  an unreadable nested plans/ dir was silently treated as "zero
  artifacts, all covered" instead of failing closed. Now branches on
  scope !== SCOPE.COMPLETE.

Also, per ponytail's second-round findings: reverted an unwarranted
FINGERPRINT_VERSION bump and digest length-prefix from the first fix
(no v1 digest has ever existed -- the feature is unreleased -- and the
prefix closed a collision that grants no capability beyond what a
writer of covered_files already has more cheaply); removed a
verifier-facing escape-hatch instruction whose own example was a case
that should trigger staleness, not bypass it; corrected CONTEXT.md
references to the renamed allCurrentArtifactsCovered and a stale
"unconditional" rescan claim; simplified the isStale derivation,
removed dead FsLike members, and tightened test coverage.

Regression tests for both fail-open bugs are included and were each
confirmed to fail against the pre-fix code before the fix landed.

full test suite: 2558/2560 pass, 2 skipped, 0 fail

* fix(#4155): trim gsd-verifier.md under the LARGE size cap

Fork CI caught what my local runs missed: the superseded/nested-plans
instruction added earlier pushed gsd-verifier.md to 49299 bytes,
147 over the LARGE tier's 49152-byte hard cap
(tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's
wording and dropped a redundant inline comment tag; no content lost.

* chore(#4155): point changeset at the upstream PR number

pr: 19 was the fork PR opened for internal review-lane CI; now that
open-gsd/gsd-core#4290 exists, the changeset field must match it per
CONTRIBUTING.md's release-notes convention.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:42:52 -04:00
Dennis Alexis Valin Dittrich
e8800287d5 enhance(#4153): fail closed unresolved update targets (#4237)
* test(#4153): cover unresolved update target

* fix(#4153): fail closed unresolved update target

* test(#4153): require a concrete recovery installer

* fix(#4153): use concrete unresolved recovery command

* chore(#4153): bind changeset to fork PR

* test(#4153): cover portable update diagnostics

* fix(#4153): keep update diagnostics portable

* fix(#4153): harden update version diagnostics

* test(#4153): reject jq in update version checks

* test(#4153): expose step-local parser gap

* fix(#4153): keep JSON parsing step-local

* docs(#4153): align update target guidance

* test(#4153): expose workflow runtime fallback

* test(#4153): expose resolver runtime fallback

* fix(#4153): leave unknown workflow runtime empty

* fix(#4153): stop inferring Claude for unknown targets

* test(#4153): preserve Claude workflow targeting

* test(#4153): preserve known runtime directory identity

* fix(#4153): recognize Claude workflow paths

* fix(#4153): reuse known runtime directory identities

* chore(#4153): acknowledge emitted workflow growth

The fail-closed diagnostic and known-runtime preservation deliberately add 48 emitted bytes.

Emitted-Drift-Ack-Growth: update.md — explicit unresolved-target diagnostics and known-runtime preservation

* test(#4153): expose missing Windsurf workflow contract

* docs(#4153): document Windsurf update targets

* chore(#4153): bind changeset to upstream PR

* fix(#4153): gate unresolved-target exit before the VERSION-missing fallback

The VERSION-missing bullet in get_installed_version sat before the
UPDATE_TARGET_UNRESOLVED exit and shared its trigger condition (version
0.0.0). An LLM agent reading the workflow top-to-bottom could satisfy
"proceed to install" without ever reaching the fail-closed exit this
PR adds, reopening the ill-defined mutating path #4153 closes. Reorder
so the unresolved-target gate runs first and scope the VERSION-missing
bullet to require an already-resolved target.

Also drop two vacuous mutationSpies entries: they checked '--sync'/
'--reapply' (commands/gsd/update.md content) against `step`, a slice of
workflows/update.md — always -1 regardless of correctness. Those routes
bypass get_installed_version entirely and are already covered by
install.test.cjs, reapply-patches.test.cjs, and
skill-frontmatter-contract.test.cjs.

* chore(#4153): point changeset pr field at fork PR #10 for fork CI

* test(#4153): guard RUNTIME_DIRS/update.md table parity, confirm narrowing intent

Nit 1: update.md's PREFERRED_RUNTIME prose and RUNTIME_DIRS
(src/update-context.cts) are two independently maintained copies of the
same runtime->dir mapping with no parity check; add one so a future
edit to either surface without the other fails loudly instead of
silently drifting.

Nit 2: call out in the changeset that a custom --config-dir matching no
known runtime, marker file, or env var now resolves unresolved instead
of silently defaulting to claude -- this narrowing is intentional, it's
the fail-closed behavior #4153 asks for.

* fix(#4153): drop dead $UC fallback in check_latest_version's uc_field, cover unresolved-runtime fast path

agy (gemini-3.8-flash-high) adversarial review of the full PR:

1. check_latest_version's uc_field() copy-pasted get_installed_version's
   `${2:-$UC}` fallback, but every call site here passes $2 explicitly and
   $UC does not exist in this step's scope -- dead, misleading reference.
   Use $2 directly.
2. No unit test covered resolveUpdateContext's preferredConfigDir fast path
   returning runtime: '' for a custom --config-dir matching no RUNTIME_DIRS
   suffix, marker file, or env var (the exact fail-closed case #4153 adds).
   Added.

A third finding (update.md:90 using /gsd:update vs docs using /gsd-update)
was investigated and rejected: /gsd:update is the actual registered
Claude Code command name (commands/gsd/update.md name: gsd:update) and is
locked by this PR's own test (tests/update-workflow.test.cjs); /gsd-update
is a separate, pre-existing, intentional prose convention used in
audience-facing docs (README/INVENTORY/FEATURES). Not a defect.

* chore(#4153): backfill changeset pr field to upstream PR #4237

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:17:21 -04:00
Cody Anderson
77e2472ca0 enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration

Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on
Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets
(the .env.example/.sample/.template/.dist templates stay readable).
Read checks file_path; Grep checks an explicit path and judges the glob
per brace alternative; Bash runs a two-pass token scan (quotes, comments,
redirects with fd digits, separators, $( )/backtick/<( ) recursion,
heredoc bodies never scanned as commands, nested bash -c/eval rescans,
git <ref>:<path> shapes) with a closed non-reading exemption set for
existence checks. Fail-open crash policy; 1 MiB commands are denied as
command-too-large; more than 64 glob alternatives as glob-too-complex.

Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt
for approval whenever any Read() deny rule exists, even in auto mode. A
hook denial is not a permission rule and never arms that check. The
installer-written deny rules are retired in the follow-up commit.

Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks
HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking
guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell),
shell-command-projection managed sets, installer-migration-report,
OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch),
docs tables in five locales, ADR-766 always-on list, regen:derived
fixtures, and a new table-driven unit suite.

* test(#4221): pin the secret-read guard in existing hook gates

Register gsd-secret-read-guard.js in every existing hook gate: the
hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest
REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity
EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal-
hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS,
kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and
typed-payload floors, the OpenCode adapter (grep mapping, include ->
glob, three dispatch tests) and a Kimi TOML matcher assertion.

* fix(#4221): retire installer Read() deny rules (legacy filter)

Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS
and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets)
strings. mergeClaudePermissions now only filters them out of an existing
permissions.deny: an absent deny key stays absent, a malformed one is
still repaired to [], and an array emptied by the filter is deleted so
no `"deny": []` residue is left. Uninstall filters the same legacy list
and, symmetric with the Antigravity branch, drops an emptied allow or
deny key and an emptied permissions object.

Unlike the #2278 allow-side migration there is no surviving current
deny list, so the constant is renamed rather than mirrored. Removal is
byte-exact: a hand-written identical rule is indistinguishable from the
installer's and is removed too (the manifest never recorded permission
strings). USER-GUIDE and CONTEXT.md updated.

* test(#4221): flip install-regressions deny-rule assertions to the retired shape

The fresh-merge, non-destructive merge, idempotency, end-to-end install,
reinstall and uninstall assertions now expect no Read(.env*) deny rules
and no permissions.deny key on a fresh install; the deny:null repair case
is kept. A new describe block covers the legacy filter: retired strings
removed with a user entry kept, partial sets, near-miss strings
untouched, idempotency, GSD-only deny array deleted, a pre-existing
empty deny preserved, and uninstall symmetry for allow/deny/permissions.

* chore(#4221): add changeset fragment for PR #4236

* fix(#4221): case-fold names; scan shell stdin and xargs pipes

Review round 1 (trek-e):

- Blocker: secret-name matching is now case-insensitive in the Read,
  Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a
  case-insensitive filesystem are recognized as the same secret file.
- Major: a shell interpreter's script is now scanned wherever it comes
  from. The tokenizer keeps heredoc bodies as per-segment tokens and
  records separator operators; pass 2 groups by segment id and resolves
  bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined
  `-lc`) scans the script operand, a file operand is checked as a file
  (a `<( )` operand's echo/printf output is reconstructed), otherwise
  stdin is the script and heredocs, here-strings and a piped echo/printf
  source are scanned. `eval` joins all its operands; `source`/`.` handle
  process substitution. Data heredocs (`cat <<EOF`, the commit-message
  shape) stay unscanned.
- Major: `… | xargs <cmd>` checks the upstream segment's operands as
  file names when the sub-command reads (`echo .env | xargs cat`,
  `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the
  inference; a shell sub-command's `-c` script is scanned.

Header, USER-GUIDE bullet and changeset updated; documented gaps now
include piped scripts from non-echo sources and `exec`/`timeout`
wrappers. 60 new suite cases pin the block and allow shapes.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:00:08 -04:00
Tom Boucher
515191f07d feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only)

* test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED)

Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/
merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and
confirms the prior research pass's Open Question 1: a coordinator crash
between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge)
leaves BATCH.json at "pending" with no STATE.md row yet (only written in
Step 9), so --resume's eligibility re-derivation would dispatch a second
executor into a new worktree for the same item, orphaning the first.

This test asserts worktree-dispatch.md's Step 6 excludes an item whose
SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's
existing PLAN.md-existence check one layer earlier. Fails against the
current worktree-dispatch.md, which has no such guard.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the full trace and fix-location rationale.

* fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN)

worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round
via the same quick-batch resume call resume-mode.md uses, but had no check
for "did this item already finish executing" the way planner-wave.md
already checks "did this item already get planned" (PLAN.md existence)
before re-planning. A coordinator crash between Step 6 (executor commits,
SUMMARY.md written) and Step 7 (merge) left the item eligible for a second
dispatch on --resume, orphaning the first worktree's real, already-
committed work and silently losing it once the second executor's SUMMARY.md
write clobbered the first at the same item_dir path.

Adds a SUMMARY.md-existence exclusion before spawn-plan is computed,
symmetric to planner-wave.md's PLAN.md check. The excluded item is not
lost: merge-wave.md's own mergeable-wave criterion (status=pending,
SUMMARY.md on disk, not yet merged) already picks it up independently of
this eligible/spawn list.

Workflow-prose-only fix — touches no already-merged/reviewed .cts module.
See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1
for the fix-location rationale (why not resumeBatch itself).

* test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules

Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677,
epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership
attempts", "scope drift", "submodules"):

- Arbitrary-worktree ownership tampering: a manifest entry naming a
  non-agent branch is silently dropped at normalization before any git
  subprocess runs; a manifest entry naming a plausible agent-branch that
  was never actually created by this repo's own worktree.create (a
  genuinely foreign repo/branch) is blocked via base_mismatch. Both leave
  the foreign location and repoRoot's HEAD provably untouched.

- Advisory scope drift: a committed path outside declared files_modified
  still merges successfully (advisory, never blocking) while surfacing a
  scope_out_of_declared warning naming the drifted path; an exact
  declared-scope match produces zero warnings (boundary case).

- Real .gitmodules submodule integration: a repo containing a real local
  git submodule merges cleanly through executeWorktreeWaveCleanupPlan for
  an unrelated plan; a real gitlink pointer bump (declared) merges cleanly
  with the superproject tree reflecting the new pinned commit; an
  undeclared bump is advisory-only and surfaces a scope warning naming
  vendor/sub, same as any other undeclared modification.

No src/*.cts changes — all three gaps were coverage-only; the underlying
primitives already behaved correctly (independently verified against real
git subprocess output before writing each assertion).

* docs(#3677): document how to diagnose a preserved quick-batch worktree

Extends the one-sentence "worktree is preserved (never deleted)" mention
into a concrete diagnosis procedure: where the preserved directory is, how
to read the executor's real commits/diff against the plan's declared
files_modified, how to read the item's own SUMMARY.md independent of merge
outcome, how to manually merge-and-clean-up or discard, and how to re-run
--resume afterward. Also documents that a SUMMARY.md-written-but-still-
pending item (the crash-window case fixed in this same PR) needs no manual
intervention — --resume routes it straight to the merge step.

* chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only)

* fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable

Orthogonal review (Spec finding): the crash-window regression test added
earlier this phase only asserted readStep('worktree-dispatch.md') + regex
matches against the markdown prose — proving the DOCUMENTATION says the
right thing, never that the runtime condition (pending status + on-disk
SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own
"Alternatives considered" explicitly rejects "document recovery without
fault injection" for exactly this reason.

Extracts the filtering decision into a pure, independently testable
function, filterAlreadyExecuted(eligibleIds, executedIds) in
src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed`
CLI verb (src/quick-batch-command-router.cts) — the same pure-decision-
then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already
establish. worktree-dispatch.md now calls this verb explicitly instead of
only describing the decision in prose. A genuine fixture-based test in
tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch),
writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the
REAL resumeBatch, and proves both that resumeBatch alone still reports the
item eligible AND that filterAlreadyExecuted (fed a real filesystem check)
correctly excludes it. The prior prose-assertion tests are kept — they now
prove the workflow markdown is correctly WIRED to the verb — but are no
longer the only proof.

Self-discovered defect while building that fixture (fixed inline, not
deferred): tracing merge-wave.md against /gsd:quick's own prior art
(QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed
$QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed
coordinator correctly does not re-dispatch an already-executed item (this
fix), but nothing durably recorded that item's worktree_path/branch/base
either — Step 7 in the resumed process would have had no data to build its
cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/
dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT
a reuse of the pre-existing `worktree` field, whose loadBatch validation
requires the path to exist on disk (verified empirically: reusing it made
the batch permanently unloadable the moment a legitimately-merged worktree
was removed). worktree-dispatch.md persists the triple once a worktree is
created; merge-wave.md falls back to it when the ephemeral manifest lacks
an entry, clears it after a successful merge, and fails closed rather than
guessing if no record exists anywhere.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1
and §9.3 for the full trace, empirical verification notes, and rejected
alternatives (reusing `worktree` directly).

* test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees

Orthogonal review (Security finding): the two existing ownership-tampering
tests didn't test ownership — one was trivially rejected by
WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch-
NAME filtering, not ownership), the other pointed at a wholly separate,
never-linked foreign repo, so merge-base failed immediately because the
branch didn't exist as a ref at all. Neither exercised the real scenario:
a manifest entry whose worktree_path/branch are swapped to point at a
DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot,
with a branch name passing the shape check and a base in allowed_bases.

Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts)
directly: this is NOT a reachable gap. Git enforces branch-per-worktree
uniqueness, so a swapped-in entry.branch can only match worktree_path's
ACTUAL checked-out branch if it names that sibling's own real, uniquely-
generated branch name — which manifest tampering confined to one batch's
own record has no way to know (branch names are
agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is
collision-checked GLOBALLY across every existing quick task and batch, not
merely within one batch).

Adds a stronger test that empirically proves this: two REAL, concurrently-
alive sibling worktrees of the same repo (both via real `git worktree add`,
both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with
worktree_path/branch swapped between them in both directions. Both attempts
are blocked via branch_mismatch; both real worktrees, their branches, and
one sibling's real uncommitted-to-main commit survive completely untouched.
Supplements (does not replace) the original two tests, which still prove
distinct, real boundaries.

See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2
for the full trace, including the one explicitly-documented (not fixed)
trust boundary this investigation surfaced: the primitive defends against
fabricated data, not a caller bug that misattributes a real-but-wrong
item's own triple to a different item.

* chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only)

* docs(#3677): add changeset for PR 4240

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-03 09:47:22 -04:00
Tom Boucher
2f64e6230a feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core

Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch
binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the
new pure decision-logic module (arg validation, effective concurrency,
deterministic merge order, spawn backpressure, verification/merge
routing, cleanup-entry construction — design doc rows 3-15,24,26-28,
30-36,39; property rows 51-53). quick-batch-update-items.test.cjs
covers the new updateBatchItems export on src/quick-batch.cts (rows
15,22-23, including the negative cycle-rejection case).
quick-batch-command-router.test.cjs covers the new
gsd-tools quick-batch CLI family (rows 46-47). These reference modules/
exports that do not exist yet.

* feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router

Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision
layer — CLI verbs and pure orchestration logic only; no workflow
markdown, no Agent()/git-worktree I/O.

- src/quick-batch-dispatch.cts (new): pure decision functions consumed
  by the (separate, follow-up) /gsd:quick-batch workflow markdown —
  parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder,
  computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome,
  buildCleanupManifestEntry (the last parses caller-supplied plan text
  via the existing parsePlanDocument; no filesystem access).

- src/quick-batch.cts: adds updateBatchItems, resolving the design
  doc's Open Question 1 as ONE additive export on this module instead
  of the second, independent BATCH.json writer the design doc
  originally proposed. Reuses the same withPlanningLock transaction
  shape, computeWaves, and platformWriteSync call resumeBatch/
  completeQuickItem already use; fails closed without persisting on
  an unknown item, an unknown/self dependency, or an introduced cycle.

- src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI
  family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs)
  as a first-party always-on command (like /gsd:quick), not the opt-in
  capability-registry path graphify uses. Verbs: create/update/resume/
  complete (wrap quick-batch.cts) and effective-concurrency/
  merge-eligible/spawn-plan/verification-routing/merge-routing/
  cleanup-entry/parse-args (wrap quick-batch-dispatch.cts).

Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property
rows 51-53. Rows covering workflow markdown / Agent() dispatch /
`git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the
follow-up markdown-authoring pass, per the phase brief's explicit
scope boundary.

* docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces

New-.cts-module ripple for the two Phase 4 modules (epic #3344,
ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts,
ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source,
not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json
(via `node scripts/gen-inventory-manifest.cjs --write`, after
`npm run build:lib`), and CONTEXT.md glossary entries for
"Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router
Module", plus an update to the existing "Quick-Batch Core Primitives
Module" entry documenting the new updateBatchItems export.

* test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count)

scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs
file under the quick-batch production module by longest-prefix match,
and that module is already at its 2-file cap (quick-batch.test.cjs +
quick-batch.property.test.cjs). The standalone
tests/quick-batch-update-items.test.cjs added in the prior commit
pushed it to 3 and failed `npm run lint:ci`. Fold its content into
quick-batch.test.cjs (append-only — no existing test in that file is
modified) and update the CONTEXT.md glossary reference to match.

Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after
`npm ci` (this worktree previously had no local node_modules, which
also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-*
unable to resolve typescript — resolved by npm ci, no code change
needed there). `npm run lint:ci` and
`npx tsc -p tsconfig.build.json --noEmit` are both green after this
fix.

* test(#3676): add failing tests for the quick-batch command/workflow markdown

Failing-first tests for Phase 4's markdown-authoring pass (epic #3344,
ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs
covers commands/gsd/quick-batch.md's frontmatter/objective/process,
gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR
1610 NEW_FILE_CAP) and step-fragment count, the isolation model
(rows 20-22), the executor single-writer invariant (row 18), merge
validation reusing the existing bounded primitive (row 25), the
optional research/plan-checker/verification leaves (rows 16,17,19,
30,31), planning-failure blocking execution (row 29), the submodule
guard (rows 36,44), and the new agents/gsd-planner.md quick-batch
mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers
row 48 (ordinary /gsd:quick stays byte-identical). Named
`gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's
longest-prefix bucketing doesn't fold these markdown-only tests into
the already-capped quick-batch/quick-batch-dispatch/
quick-batch-command-router production-module buckets from the CORE
pass. These reference files that do not exist yet.

* feat(#3676): author the quick-batch command, workflow, and planner mode

Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch
binding") — the orchestration layer that calls into Pass 1's CLI
verbs (src/quick-batch-command-router.cts).

- commands/gsd/quick-batch.md (new): frontmatter/objective/process,
  delegates argument validation to `quick-batch parse-args`
  (parseQuickBatchArgs) rather than re-deriving the grammar.

- gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR
  1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded
  step fragments under gsd-core/workflows/quick-batch/steps/:
  resume-mode, batch-init, research-phase (flag:--research),
  planner-wave (+ nested plan-checker-loop when --validate),
  worktree-dispatch, merge-wave, verification-wave (flag:--validate),
  completion. Covers design doc rows 3-45: capacity/isolation
  resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG-
  layer planning with full-task-catalog prompts and always-required
  depends_on/files_modified frontmatter, serialized worktree create/
  merge/cleanup via the existing worktree.cleanup-wave primitive,
  deterministic wave-order merging, verification routing
  (human_needed/gaps_found), the executor single-writer invariant,
  submodule fail-loud guard, and #1941 fork-base auto-degrade.

- agents/gsd-planner.md: additive new `load_mode_context` bullet for
  `**Mode:** quick-batch`, pointing at the new
  gsd-core/references/planner-quick-batch.md reference (documents the
  always-required depends_on/files_modified contract, reusing the
  existing frontmatter grammar — no new keys). Existing modes
  byte-identical, only a new bullet added.

- src/init.cts (+init-command-router.cts, +command-aliases.cts):
  cmdInitQuickBatch / `init.quick-batch` — model profiles,
  commit_docs, roadmap/planning existence checks, and the
  section_manifest field gating research-phase/verification-wave
  (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY
  atoms — no new atom needed).

Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the
prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43,
45 covered by construction (verb wiring, single-writer prompt
constraints, crash-window resume via unmodified Phase 3 primitives).

* docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern

npm run regen:derived output for the new command/workflow/reference
(epic #3344, ADR-1239 "Quick-batch binding"):
- skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md)
- docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md,
  planner-quick-batch.md, and the quick-batch-dispatch.cjs/
  quick-batch-command-router.cjs CLI-module rows' now-live
  `/gsd-quick-batch` cross-reference (was "(separate, follow-up)")
  + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`)
- gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`)
  — research-phase/verification-wave gsd:section entries for the new
  quick-batch workflow
- tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) —
  the new command/workflow/skill/reference files now ship to every
  runtime

scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for
gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS
word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional-
flag pattern gsd-core/workflows/quick.md already carries baselined
(e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2);
quoting would break the intended "omit this arg when the flag is
false" splitting.

* fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch

Security review pass findings, both confirmed real:

1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment
   interpolated the raw, attacker-influenced task ${description} (and
   the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw
   description into every planner's prompt in the layer) straight into
   Agent() prompt bodies with no boundary. Fixed by wrapping every such
   interpolation in a <security_context> + DATA_START/DATA_END
   boundary, matching the CONCRETE convention already implemented in
   this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md,
   gsd-core/workflows/debug.md) — commands/gsd/quick.md's own
   <security_notes> only asserts this convention in prose, so the
   debug-agent files are the real precedent followed here. Added a new
   <security_notes> block to commands/gsd/quick-batch.md (it had none)
   documenting both this fix and the one below.

2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection.
   gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both
   ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED,
   causing shell word-splitting and pathname expansion on raw task-list
   text before the parser ever saw it. Fixed at the source: added a
   `--text <string>` form to the `parse-args` verb
   (src/quick-batch-command-router.cts) that accepts the ENTIRE
   $ARGUMENTS as ONE quoted argv element and does the whitespace split
   itself, in Node — which is never glob-aware, unlike the shell.
   Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form
   is kept for direct/test callers that already have a real argv array.

The SC2086 baseline entry added for the original unquoted line is now
stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports
it) and has been removed; the two SC2046 entries for the UNRELATED,
still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style
conditional-flag splitting remain — that line only ever expands to one
of a few known-safe literal strings (never raw user text), matching
quick.md's own already-baselined convention exactly.

Tests: quick-batch-command-router.test.cjs covers the new --text form
(token splitting, glob-shaped text passing through literally
unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs
asserts the DATA_START/DATA_END boundary on every leaf prompt
(research-phase/planner-wave/plan-checker-loop/verification-wave,
including the shared task catalog) and the quoted --text call sites.

* fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35

Spec review pass findings — the test matrix claimed "yes" coverage
these assertions did not actually support:

- Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection
  only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs,
  committed alongside the security fix that touches the same file) that
  .planning/quick-batches/ is never created for any rejected value —
  createBatch is genuinely never reached.
- Row 18 (--resume <unknown-batch-id>): previously only exercised a
  hand-corrupted BATCH.json, never a genuinely nonexistent batch
  directory. Added the real nonexistent-id case (also in
  quick-batch-command-router.test.cjs).
- Row 24 (post-planning updateBatchItems racing a concurrent
  completeQuickItem for a different item, both through
  withPlanningLock): zero test existed. Added a property test
  (tests/quick-batch.property.test.cjs, appended — Phase 3's own file,
  no existing test touched) exercising both call orders and asserting
  no lost update in the final on-disk manifest — the same technique
  Phase 3's own row-15 lock-contention property test uses (sequential
  calls through the real lock; a working mutex makes any interleaving
  equivalent to some serial order, so this is the same claim a literal
  concurrent-thread test would make without OS-level threading).
- Row 34 (worktree preserved on merge_failed) and row 35 (undeclared-
  deletion detection): both were previously asserted only at the pure
  routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs
  using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs
  already establishes for executeWorktreeWaveCleanupPlan (real repo,
  real worktree, a REAL merge conflict / a REAL file deletion diffed
  against declared_deletions) — asserting the actual worktree directory
  survives on disk, not just that a pure function returns a
  preserveWorktree:true field. Named gsd-quick-batch-* so lint-test-
  file-count's bucketing doesn't fold it into any capped module bucket.

Row 48 (/gsd:quick regression) intentionally left as-is per the
reviewer's own framing: the byte-identity claim is already
mechanically proven by the changed-path diff (git diff --name-only
empty on those two paths IS byte-identity), and a genuine execution-
level regression test would require actually running the workflow —
out of scope for this repo's unit-test model (no other quick.md
regression test in this repo does that either).

* docs(#3676): add the changeset and user-facing docs the command needed

Standards review pass findings — both HARD:

- Missing changeset. None of the 6 prior #3676 commits touched
  .changeset/*. /gsd-quick-batch is a new user-facing command;
  CLAUDE.md/CONTRIBUTING.md require one. Added
  .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder —
  backfilled after the PR opens, matching CLAUDE.md's own documented
  convention and Phase 3's own precedent, #4190's
  .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen
  form `/gsd-quick-batch` throughout, never the source-artifact colon
  form (`scripts/lint-docs-command-form.cjs` confirms 0 violations;
  that check scans docs/**, not .changeset/, so it was never actually
  in scope for the fragment itself, but the wording still follows the
  doc convention for consistency, matching how Phase 3's own fragment
  named the not-yet-shipped command).
- Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis
  how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's
  existing convention for /gsd-quick /gsd-fast) covering --jobs,
  --validate, --research, --resume, --file, the capacity/isolation
  interaction, and resume/failure recovery. Cross-linked from
  docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's
  own "Related" section. Added a /gsd-quick-batch section to
  docs/COMMANDS.md (same table format as the existing /gsd-quick
  entry) and docs/features/quick-batch.md (REQ-QB-01..12, same
  frontmatter shape as docs/features/quick-mode.md) — regenerated
  docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md
  via the standard generators.

* fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught

gsd-test's real run against 155e8975b3 found 43 failures, all rooted in
this phase's own new command/workflow never being registered across
~10 independent generated/hand-maintained registries this repo keeps
in parity by convention. Root-caused each, no test weakened or
special-cased.

- help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live-
  registry.test.cjs): added a /gsd:quick-batch entry to
  gsd-core/workflows/help/modes/full.md (the real help.md content;
  gsd-core/workflows/help.md is a thin dispatcher) documenting every
  flag (--file/--jobs/--validate/--research/--resume), matching the
  existing /gsd:quick entry's format.

- gen-section-manifest.test.cjs: quick-batch.md's
  `gsd_run query init.quick-batch` invocation used inline
  `$([ ... ] && echo --flag)` substitutions, which never satisfy the
  test's exact-whitespace-token / assigned-variable detection (the
  trailing `))` glued onto `--research` in the compound substitution
  broke the "exact token" match). Rewrote to the same
  VALIDATE_PARAM/RESEARCH_PARAM two-line pattern
  gsd-core/workflows/quick.md's own Step 2 already uses.

- runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md
  fragments that call gsd_run each needed their OWN embedded copy of
  the canonical shim preamble (every workflow .md that calls gsd_run
  carries its own copy — reading one file does not persist shell state
  into another). Ran `node scripts/sync-runtime-launcher.cjs`, which
  inserted it before each file's first gsd_run call.
  plan-checker-loop.md correctly has none — it never calls gsd_run
  directly.

- Namespace routing (skill-manifest.test.cjs, install-nested-
  layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added
  `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and
  routing table (same namespace `quick` already routes through), and
  to src/clusters.cts's `utility` cluster (same cluster `quick`
  already belongs to). Verified by hand-running installRuntimeArtifacts
  + applySurface for augment/cline against a real temp install: exactly
  6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested
  under gsd-ns-workflow/skills/, never re-flattened.

- mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a
  brand-new command is a real count change, not a bug this test should
  hide).

- model-omit-when-inherit-guard.test.cjs: added the canonical
  `<!-- #2517 model-omit-on-inherit -->` marker block to
  gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/
  researcher/checker/executor/verifier — lives in a steps/ fragment,
  read combined with the host by this test's own readWorkflowCombined,
  same as quick.md's own research-phase.md carries it for its gated
  section). Also fixed a genuine pre-existing inconsistency in the
  test's own "#2711: the guarded set is derived from dispatch sites"
  check: its `nonDispatching` computation read the BARE host file while
  `derived` (the set it's checked against) reads the combined
  host+steps content — inconsistent with that same test file's own
  #2994 doc comment explaining why the combined read is necessary.
  quick-batch.md is the first workflow whose EVERY model="{...}"
  dispatch site lives in a mandatory (never gated) steps/ fragment —
  extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a
  brand-new file — which is what exposed the mismatch. Fixed by using
  the same readWorkflowCombined read in both places.

- skill-frontmatter-contract.test.cjs: shortened
  commands/gsd/quick-batch.md's frontmatter `description` from 107 to
  91 chars (<=100 budget), and added `quick-batch.md` to the hand-
  maintained KNOWN_SKILLS consolidation allowlist with a #3676
  justification comment (a genuinely new first-party command, not a
  consolidation of an existing skill).

- workflow-fragments-emission.install.test.cjs: added `quick-batch.md`
  to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is
  deliberately NOT a no-op for it — its research-phase/verification-
  wave sections are gated).

- Regenerated all downstream artifacts (npm run build:lib && npm run
  regen:derived && npm run gen:plugin-skills -- --write && npm run
  gen:features -- --write): skills/gsd-quick-batch/SKILL.md,
  skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for
  augment/cline/hermes/qwen/trae/zcode.

- emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition
  (one new `load_mode_context` bullet pointing at the new
  gsd-core/references/planner-quick-batch.md reference) grew the file
  124 bytes without an acknowledgment trailer. Acknowledged below —
  the growth is the deliberate, additive, single-bullet change from
  the earlier feat(#3676) commit, not drift.

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green
(includes lint-workflow-shellcheck, lint-test-file-count,
lint-docs-command-form). The deep install/spawn/registry tests gsd-test
actually runs (docs-parity-live-registry, gen-section-manifest,
runtime-launcher-parity, install-nested-layout,
runtime-artifact-layout-surface, skill-manifest, skill-frontmatter-
contract, mcp-server-catalog, model-omit-when-inherit-guard,
workflow-fragments-emission) are not part of lint:ci — each fix above
was independently verified by hand-invoking the exact production
function the failing test calls (installRuntimeArtifacts, applySurface,
composeWorkflow, the CLUSTERS union, the section-manifest forwarding
regex) against the real repo tree and confirming the expected shape.

Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift.

* fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget

skill-frontmatter-contract.test.cjs's "feature #3039: tiered help —
size budgets" enforces a SEPARATE line-count ceiling for
gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines,
tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's
assertTightCeiling) — independent of the skill-frontmatter description-
length budget and consolidation allowlist I touched in the prior round;
those are unrelated checks in the same test FILE, not the same check.

Root cause: the /gsd:quick-batch entry I added to full.md in the
docs-parity fix round was 17 lines, pushing the file from 834 to 851
lines — 7 over the 844 ceiling. Condensed the entry (merged the
per-flag bullet list into one dense "Flags:" line, dropped from 3
Usage examples to 1) to 844 lines exactly — at the ceiling with zero
slack, which assertTightCeiling accepts (it only fails on
actualMax > ceiling, or on slack > grace when the ceiling is too
LOOSE — zero slack triggers neither).

Verified after trimming: full.md still contains a live /gsd:quick-batch
reference (bidirectional parity) and all 5 argument-hint flags
(--jobs/--validate/--research/--resume/--file) still appear as literal
tokens (docs-parity-live-registry.test.cjs's own flag-coverage check,
re-run by hand against the trimmed content).

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green.

* docs(#3676): backfill changeset pr number to 4212

Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's
pr:0 placeholder backfilled with the real PR number now that
gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR
Number Handling convention and Phase 3's own #4190 precedent
(708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt
from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE.

* fix(#3676): resolve prompt-injection-scan false positive on test fixture

tests/quick-batch.test.cjs:232's row 11b regression proves the task-list
parser carries a prompt-injection-shaped task description through
createBatch as inert data, never interpreted. The fixture has to be a
real "ignore all previous instructions..." phrase or the test asserts
nothing, but the full-file --diff scan flagged it once unrelated edits
in the same file pulled it into the changed-file set.

Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the
sanctioned, precedented exemption already used for other legitimate
security-regression fixtures (tests/windsurf-conversion.test.cjs,
tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs)
per DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:31 -04:00
Tom Boucher
bf4485ada2 enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification

Adds unit tests for the not-yet-implemented text_en field on Requirement
(fallback selection, empty/whitespace/non-string rejection, shapes-override
precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX),
and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents
populating text_en for response_language projects. All new tests are RED
until src/edge-probe.cts and the workflow docs are updated.

* feat(#3717): make edge-probe shape classification read an optional text_en field

Requirement gains an optional text_en; classifyShape's own signature stays
untouched (a locked, directly-tested export), and the text_en ?? text
selection is pushed to proposeEdges' single call site instead. text_en is
validated fail-closed: an empty or whitespace-only value throws rather than
silently winning the ?? fallback and degrading classification to zero shapes.

This makes the #2773 doc-only translation convention an explicit,
validatable field instead of an invisible instruction, per the approved
Form-1 scope on #3717.

* docs(#3717): document the text_en field across spec-phase, reference and how-to docs

Updates Step 5.5's response_language instructions, the edge-probe reference
Inputs contract, the FEATURES.md fragment, and the non-English how-to guide
to describe the new text_en field: text keeps the requirement's own wording
in all cases, text_en (when populated) is the engine-only English rendering
the classifier prefers.

* docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550

Updates the Edge Probe Module glossary entry to describe the text_en field
and its fail-closed validation, and appends an ADR-550 amendment recording
why this is additive and does not re-open the #652 LLM-classifier rejection
(text_en is a plain field read by the existing deterministic regex
classifier, not a new model-dependent surface).

* docs(#3717): add changeset fragment and regenerate FEATURES.md

pr:0 placeholder — backfilled with the real PR number after the PR opens.

* docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests

Code-review (Spec axis) finding: the workflow-prose contract tests and the
ADR-550 amendment overclaimed themselves as "the machine check the #2773
doc-only stopgap lacked." That check is actually engine-level
(validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) —
the prose tests are the same style of assertion #2773 already used. Reworded
both to attribute the claim correctly.

* fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line

The #3717 rewrite of Step 5.5's response_language paragraph moved a line
break so "requirement `id`s" ended one physical line and "are never
translated" started the next. The pre-existing #2773 regression test
(tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never
translated" on the SAME line (no \n in between, matching git's own
line-oriented prose), so the reflow silently broke it. Rewrapped so the
sentence lands on one line again, verified against every #2773/#3717
regex assertion in that test file.

Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift.

* chore(#3717): backfill changeset PR number

pr:0 -> pr:4156 now that the PR exists.

---------

Co-authored-by: sim <sim@local>
2026-09-01 21:39:53 -04:00
Tom Boucher
c0fd2e3f4c feat(#3673): add dispatch.maxConcurrency axis and dispatch-capacity query (#4162)
* test(#3673): add failing tests for dispatch.maxConcurrency axis and dispatch-capacity CLI route

Extends tests/host-integration.test.cjs with negotiateHostCapabilities
maxConcurrency negotiation coverage (test matrix rows 1-15, including a
fast-check property test) and a new #3673 dispatch-capacity CLI route
describe block spawning the real gsd-tools.cjs (rows 16-25). Extends
tests/host-integration-validator-parity.test.cjs with an all-19-descriptor
maxConcurrency presence/validity sweep (row 26) and adds a hostile-input
validator test to host-integration.test.cjs (row 27).

The dispatch.maxConcurrency field does not exist yet, so these tests fail.

* feat(#3673): add dispatch.maxConcurrency axis, negotiation, validator parity, and the dispatch-capacity query

Adds a numeric dispatch.maxConcurrency sub-field to the Host-Integration
Interface (ADR-1239 Phase 1), following the existing dispatch.isolation
sub-field pattern: DispatchCapability interface, SAFE_DEFAULTS/PROFILE_BASELINES
floors, and a negotiateHostCapabilities branch that passes through a positive
safe integer and fails closed to 1 otherwise (no engine-side reduction, per
the design doc's explicit rejection of a min(host,engine) rule).

capability-validator.cjs gains parity validation for the new field (optional,
positive safe integer or the "undocumented" sentinel — mirroring isolation's
"added after existing descriptors" treatment).

gsd-tools.cjs gains a new `query dispatch-capacity` route, a pure-read sibling
of `dispatch-isolation` with no side effects: live env
(GSD_DISPATCH_MAX_CONCURRENCY) > descriptor > fallback-to-1 precedence.

All 19 capabilities/*/capability.json descriptors now declare
dispatch.maxConcurrency: claude carries the one cited value (20, per
code.claude.com/docs/en/sub-agents); the other 18 carry "undocumented"
(not yet researched for this axis).

* docs(#3673): document dispatch.maxConcurrency and add its citation row to the capability matrix

Updates docs/reference/host-integration-interface.md's dispatch struct entry
(also backfilling the previously-undocumented isolation/backgroundDispatch
sub-fields found stale in the same table) and adds fail-closed/live-transport
precedence prose for the new maxConcurrency field.

Adds a dispatch.maxConcurrency row (with citation) to all 19 host sections in
docs/reference/host-integration-capability-matrix.md — required for
tests/host-integration-descriptors.test.cjs's kimi-code matrix-parity check,
which asserts every declared dispatch sub-axis is documented there.

Updates docs/how-to/add-or-update-a-host-integration.md's dispatch checklist
and example descriptor block to mention maxConcurrency (and, likewise
backfilling a stale gap, isolation/backgroundDispatch).

* fix(#3673): extract shared maxConcurrency validator, drop dead reserved-name check

Exports isPositiveSafeInteger from src/host-integration.cts as the single
source of truth for the dispatch.maxConcurrency positive-safe-integer
contract; negotiateHostCapabilities and gsd-tools.cjs's routeDispatchCapacity
now both call it instead of independently reimplementing the same predicate.

Also removes the __proto__/constructor/prototype reserved-name branch from
capability-validator.cjs's maxConcurrency check — copy-pasted from the
string-enum fields above it, but unreachable for a numeric field (the
generic positive-safe-integer branch already rejects any string) and absent
from maxDepth, the field the code's own comment claims to mirror.

---------

Co-authored-by: sim <sim@local>
2026-09-01 21:39:12 -04:00
Tom Boucher
62b0d939b6 feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey

Add an optional `timeoutConfigKey` field to the reviewer lane descriptor,
resolved in `resolveLanePlan` at invocation time and falling back to the
frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the
existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes
declare `review.timeouts.<slug>` on both surfaces (the descriptor and their
capability.json manifest), validated by capability-validator.cjs.

For the antigravity lane, the native `agy --print-timeout` flag — previously
a second hardcoded literal (`540s`) independent of the outer cap — is now
derived from the same resolved outer timeout in `antigravityArgv`, preserving
the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is
declared data, the inner one is handler-owned).

The antigravity default timeoutFloorMs stays at 600s per the maintainer's
disposition; users raise it through the new config key instead.

* docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper

Address code-review findings on the timeoutConfigKey change: extract the
inline timeout-resolution logic into a named, exported, directly-tested
resolveTimeoutMs helper (matching the file's existing configString/
normalizeHost convention); document the new review.timeouts.* federated
config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md,
and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment.

* fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner

gsd-test caught two design mistakes in the prior commits:

1. SpawnPlan.argv is documented and tested as fully resolved by
   resolveLanePlan (model/effort/output/prompt already folded in) — leaving
   the antigravity '{{nativeTimeout}}' marker unresolved until the runner's
   antigravityArgv violated that contract and broke tests that read
   plan.argv directly (tests/antigravity-reviewer.test.cjs,
   tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}'
   is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself
   via the new nativeTimeoutToken() helper, exactly like the other four.
   antigravityArgv reverts to its pre-#3274 four-argument form. Also missed
   updating capabilities/antigravity/capability.json's invoke.args to match
   the descriptor, which broke the manifest/descriptor parity test.

2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow
   invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three
   lanes with neither a model flag nor a host — may own no config key beyond
   their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes
   violated it. Fix: those three keep timeoutConfigKey: null and own no
   review.timeouts.* key, matching their existing modelConfigKey: null. The
   other 9 lanes are unaffected.

* chore(#3274): backfill changeset PR number (pr:0 -> 4083)

---------

Co-authored-by: sim <sim@local>
2026-08-30 14:26:02 -04:00
Tom Boucher
370cfc6680 enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90%

Adds two new mechanisms plus an audit-coverage extension:

- scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic
- scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check,
  wired into test/test-full/mutate/smoke as each job's last step
- scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml:
  scheduled REST-API poll that appends new records to
  tests/ci-timeout-budget-history.jsonl and opens a small data-only PR
- tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate
  (mutation.yml) and smoke (install-smoke.yml), which previously had no
  headroom-factor gate coverage at all

Does not change any timeout-minutes value, shard composition, or shard-1
contents — those stay maintainer policy calls per the issue's own scope.

* fix(#4036): address two-orthogonal-review findings

- Parity tests guarding the two hand-duplicated literals this design
  cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES
  vs each job's own timeout-minutes, and ci-timeout-report.cjs's
  JOB_RULES name-prefixes vs each job's actual name: template.
- Thread run.event through as runEvent on every persisted record, so
  PR-context and push-context install-smoke timings (genuinely
  different matrix shape) are distinguishable in the history rather
  than silently conflated under one job name.
- Replace the Windows near-cap start-time step's ambiguous PowerShell
  +/>> precedence with GitHub's documented string-interpolation form.
- Move github.run_id out of direct ${{ }} shell interpolation into an
  env: var in the new scheduled workflow, per this repo's own
  expression-injection-safe convention.

* test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs

npm run gen:install-tree — scripts/ ships wholesale into the installed
package (per ADR/known-defect precedent from #4012's own PR history: a
new scripts/lib/*.cjs file needs its golden entry regenerated or every
runtime's install-tree test fails). Confirmed via gsd-test: this was the
sole cause of the first real verification run's 25 failures (all in
tests/golden-install-tree.test.cjs, one per runtime). Top-level
scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs)
are not individually tracked in these fixtures — consistent with every
other existing top-level scripts/*.cjs file, so no entry was expected
or added for those two.

* fix(#4036): register new lib file with installer, fix H1 shell policy

- bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a
  hand-maintained registry, not generated — tests/install.test.cjs
  asserts every scripts/lib/ file is enumerated here)
- test.yml: replace the two OS-conditional "Record job start time"
  step pairs (test + test-full jobs) with a single unconditional
  `node -e` step. The prior pair's Windows variant declared an
  explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker
  statically flags against every OS a job's matrix can realize,
  independent of the step's own if: gate. A single Node one-liner
  needs no shell override at all — it's syntactically valid and
  behaves identically under bash, zsh, and pwsh — which is both H1
  compliant and removes the last OS-specific shell syntax from this
  change entirely.

Both defects were found by a real gsd-test run, not local gates —
lint:ci and build:lib were clean throughout because neither the
scripts/lib/ install-manifest parity check nor the H1 shell-policy
baseline runs as part of lint:ci; both are gsd-test-only suites.

* docs(#4036): how-to for reading CI timeout budget signals

The phase-gate docs check correctly flagged the enablement sequence as
3 real steps (read the near-cap warning, find the accumulated trend
file, pick the right maintainer lever) — a reference table can't carry
a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from
docs/README.md.

* chore(#4036): backfill changeset PR number (4043)

---------

Co-authored-by: sim <sim@local>
2026-08-29 16:13:15 -04:00
Tom Boucher
dd4f179672 feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam

Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute
plus a new optional `taskContentResolver` capability-manifest field let a
capability resolve a task's action/verify/acceptance-criteria/read_first/done
content from an external issue tracker instead of PLAN.md's inline body.

- src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId`
- src/task-content-resolution.cts: new leaf module — split/find/build/resolve,
  with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed
  resolution, never a silent fallback to possibly-stale inline text
- src/task-command-router.cts: new `task resolve-content --plan --task-id --raw`
  CLI verb wiring the module into a real process exit code
- gsd-core/bin/lib/capability-validator.cjs: validates the new
  `taskContentResolver` manifest field (feature-role only, cross-capability
  trackerPrefix uniqueness)
- gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md,
  docs/reference/capability-manifest.md: wire the seam into the per-task loop
  and document it as a new `execute:task` point outside the existing
  contribution/step/gate vocabulary (unconditional in autonomous mode)

Closes #3970

* fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard

Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646
Phase 1) found three defects:

1. execute-plan.md's task-content-resolution bullet fired on any
   tracker-id-bearing task with no check that it wasn't type="checkpoint:*",
   contradicting ADR-3646 Decision 1 (a checkpoint task must never enter
   resolve-content). plan-document.cts already parses trackerId: null
   unconditionally for checkpoint tasks; only the workflow prose needed the
   fix, so the bullet now explicitly excludes checkpoint tasks.

2. task-content-resolution.cts's parseResolverDeclaration accepted any
   non-empty trackerPrefix with no grammar check, while capability-
   validator.cjs's KEBAB_RE enforces kebab-case at install time — a
   Generative Fix Divergence gap. Added the same grammar (as a literal
   regex, documented as intentionally not shared across the .cts/.cjs build
   boundary) plus a parity test asserting the two surfaces agree across a
   valid/invalid trackerPrefix table.

3. task-command-router.cts's routeResolveContent path-traversal guard on
   --plan had zero test coverage. Added a test exercising a
   ../../../etc/passwit-shaped path and asserting the USAGE rejection names
   the offending path.

* fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs

Two findings caught by an isolated security-review pass on the task
content resolution seam:

- ResolverFailedError/ResolverMalformedOutputError embedded raw,
  unsanitized subprocess stderr/stdout (attacker/model-influenced via
  the tracker-id argv token) into .message. A hostile or buggy resolver
  could smuggle a newline plus a forged "Error: " line, or terminal
  escape sequences, into a diagnostic io.cjs's error() writes verbatim
  to stderr. Fixed at the constructor (task-content-resolution.cts) via
  io.cjs's existing formatDiagnosticToken(), so every caller of
  resolveTaskContent gets a safe .message by construction.

- capability-validator.cjs's validateTaskContentResolverFields had no
  upper bound on taskContentResolver.invoke.timeoutMs, letting a
  manifest declare an effectively unbounded value and defeat the
  "bounded subprocess" design intent. Added a 120000ms ceiling specific
  to this field, without touching the shared isPositiveIntegerMs()
  helper (still used unbounded by the reviewer lane's timeoutFloorMs
  and probe timeoutMs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion

gsd-test (remote dockerized matrix) came back red with 5 failures on this
PR; all five are real defects, fixed here.

- tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's
  execute-plan.md entry pointed at line 415, which ffc190df4's
  checkpoint-exclusion caveat (added near line 221) shifted down by one
  line. The actual "validated downstream by gsd-tools uat
  classify-coverage" descriptive mention now sits at line 416. Updated
  the allowlist entry's line number to match.

- tests/task-command-router-resolve-content.test.cjs: the path-traversal
  test asserted the outside-project-scope diagnostic against the thrown
  ExitError's own .message. io.cts's error() (ADR-3889) writes its
  human-readable message to fd 2 via writeAllSync and then throws a bare
  `new ExitError(1)` with no message argument — by design, so the
  exception carries no duplicate text and the thrown ExitError's message
  defaults to "process exit 1" (cli-exit.cts's ExitError constructor).
  Root cause was the test, not the source: task-command-router.cjs's
  outside-project-scope rejection already calls error() correctly and the
  diagnostic text is genuinely emitted, just on fd 2, not on the
  exception. Fixed the test to capture fd-2 writes (mirroring
  tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this
  same file's own captureStdout for fd 1) and assert against the captured
  stderr text instead of err.message. This was masked locally because a
  manual `node -e` sanity check that only inspects the caught exception's
  .message cannot see what the real node:test run actually failed on.

Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3970): backfill changeset PR number (pr:0 -> pr:4000)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 13:17:04 -04:00
Tom Boucher
d24e22b156 enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1

ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what
remained was the declaration — and the pin that makes it invisible today.

The census corrected two documented figures before any code changed.
ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier
note of mine claiming 23 was wrong and is corrected). And output({error}) is
**64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape
holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3
not 2. That matters because this phase's criterion demands the pin be asserted
over the enumerated population rather than sampled; asserting over a stale 60
would leave four sites unpinned while claiming full coverage, which is the
shape of failure this epic exists to remove.

The issue does not state the fact that shapes the design: output() never
touches the exit code. Confirmed by reading it — it writes fd 1 and returns.
So a declared outcome for those 64 sites had nowhere to be READ. The mapping
was never the work; wiring somewhere for the declaration to land was.

The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol
cells, each because the module is emitted to three locations and a module-level
`let` would let instances disagree, and runMain already maps a code returned by
main(). A third cell inherits that solution. output() records DEGRADED for any
{error} payload — key-order agnostic, which is exactly why the "42 sites"
figure undercounts — and runMain projects the cell only when main() returns
nothing, so an explicit return still wins.

error() maps its reason through a table over the closed 25-member enum, leaving
all 278 call sites untouched; 226 of them pass no reason at all. The version
gate lives in error(), NOT in projectOutcome: registered names are
version-invariant there, so mapping a reason straight through would make USAGE
project to 64 under v1 and break the pin on its first line. projectOutcome is
left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included.

Proven rather than asserted. v1 is byte-identical across three real CLI paths —
config-get plain, config-get --json-errors, and an output({error}) path —
matching exit code and exact bytes against the pre-change build. Under
GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND ->
NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An
anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason,
because without it a mapping where everything projects to 1 under both versions
would satisfy every other assertion and the declaration would be theatre.

A1 iterates all 25 enum members and A3 asserts over the measured 64-site
population, so a 26th reason or a 65th site fails until it is given a mapping —
the drift guard this phase needs, given ADR-2980's own count had drifted +4
unnoticed.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the outcome cell must never lower an exit code

The remote run caught a fail-open that this phase introduced, in the phase
whose entire purpose is removing fail-opens.

`state validate --strict` on a missing STATE.md exited **0** where it must exit
1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned
void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set
non-zero by the command was clobbered down to success. Confirmed live against a
fixture, before and after.

This refutes a review conclusion recorded earlier in this phase, that the cell
was "fail-closed and can never mask a failure as success". It could, and did.
Recording that plainly so the assumption is not repeated: the cell's danger was
never only that it might add a failure — it was that projecting it
unconditionally overwrites whatever decision came before.

Projection is now guarded: it may set a code only when none is set, and an
already-non-zero exit code always wins. The full precedence — explicit `main()`
return, then an existing non-zero exitCode, then the declared outcome — is
written at the projection site. A regression test drives a void return with a
pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix
build.

The second failure was my test encoding the wrong contract, not a code defect.
It asserted `output({found:false, error: undefined})` records DEGRADED because
the KEY is present. `JSON.stringify` drops undefined, so the payload the user
receives is `{"found":false}` — carrying no error at all, and calling that
degraded would hand back exit 80 under v2 for output that reads as clean. The
discriminator is a serializable error VALUE, not key presence. The test now
pins `{error: undefined}` as explicitly NOT degraded, and the design doc's
wording is tightened to match.

Verification runs on the remote runner.

Refs #3912

* docs(#3912): the versioned exit contract, and a flag defect the docs found

Diataxis pass for Phase 8, plus a real fix that only surfaced because writing
the how-to meant running its own examples.

The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned
projection this phase provides, so it gets an amendment naming #3912 /
ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects
DEGRADED to 80. The amendment also records the count drift rather than
restating a stale figure — the ADR ratified 60 output({error}) sites in 9
modules; the AST-measured population is 64 across the same 9 (frontmatter 7
not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the
enumerated 64. json-errors.md gains the outcome-declaration reference,
including the precedence order a review pass got wrong and the suite refuted:
an explicit main() return, then an already-set non-zero process.exitCode, then
the declared outcome. Projection may only ever set a code, never lower one.

A how-to is owed here and is written, not skipped. Under v1 nothing changes,
so the audience is an operator opting into v2 and needing to know what the
codes mean for a CI gate — a migration, which is how-to shaped. It covers
turning v2 on, the code table, why 80 is "ran and reported a condition" rather
than a crash, and how to split a gate that treats any non-zero as fatal. No
tutorial: there is no new entry point to learn, and under the default contract
a reader would be walked through observing nothing.

The defect. Running the how-to's own Step 1 example returned

    $ gsd-tools --exit-contract=v2 state validate --strict
    Error: Unknown command: --exit-contract=v2          (exit 64)

while the same flag trailing the subcommand worked and exited 80. The flag
half-worked, by argv position. resolveContractVersion scans argv
non-destructively, so the token survived into the dispatcher, which treats
argv[2] as the command name. --json-errors had already solved precisely this
at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv
splice must happen here too, otherwise the dispatcher below sees
--json-errors as an unknown command." The later flag never got the same
treatment.

Fixed rather than documented around: the version is resolved first — which
memoizes the cell and makes an invalid value throw early — and then every
--exit-contract= token is spliced out of the dispatcher's argv copy.
--exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The
regression test pins leading position, trailing position, agreement between
the two, and a loud failure on v3 rather than a silent fall back to v1.

Neither review engine would have caught this: the defect is invisible in the
diff, because the diff does not touch argv handling. It surfaced only from
running the documentation's own example. Writing a how-to is an execution pass.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the flag splice has to run before the run-with-timeout return

An isolated review of the previous commit found that the fix did not deliver
what it claimed, and that two of its own tests were weak. All three findings
reproduced by execution before any change was made.

The fix was placed below a return. main() intercepts `run-with-timeout` at
gsd-tools.cjs:4436 and returns from there — above both the --json-errors block
and the --exit-contract splice added in the previous commit. So the flag still
died in leading position for that one command:

    $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..."
    Error: Unknown command: run-with-timeout        (exit 64, child never ran)

The previous commit message and the test's describe-block both claimed
position-independence unconditionally. That was an overclaim, not a gap left
open, and it is the part worth naming: the fix was verified by hand on the
commands I happened to think of, and `run-with-timeout` returns before the
code I was verifying.

Both global-flag blocks now run above the interception, with a comment naming
it so a later edit cannot slide them back down. Moving --json-errors up fixes
the identical pre-existing bug for that flag, verified failing beforehand
(exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect,
same block, and a known-broken twin next to a fixed one is not a resting state.

Two tests were not pulling their weight. The invalid-value test was vacuous —
it passed against the pre-fix build, because `--exit-contract=v3` already
exited 1 there and already printed the resolve error lazily through
error() -> getContractVersion. Both its assertions held before the fix, so it
pinned nothing. The real discriminator is that the pre-fix build emits BOTH
"Unknown command: --exit-contract=v3" and the resolve error, while the fixed
build emits only the latter; the test now asserts that absence.

The leading-position and leading==trailing tests asserted proxies — "not 64",
"no Unknown command", "the two agree" — none of which pin a value, and all of
which would survive both positions being identically broken. With a .planning
directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0
under v1 in both positions. Those numbers are pinned now. The multi-token case
the descending splice loop exists for is covered too, and run-with-timeout has
regression tests for both flags.

The lesson is narrower than "test more". Hand-verifying the production
behavior does not verify that the test would have caught its absence. The
pre-fix binary has to be run against the test's own assertions.

Investigated and deliberately not changed: splicing before --cwd parsing
degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd:
<path>", but that is pre-existing — verified on the pre-fix build via
--json-errors, which already did it. This change joins the pattern rather than
creating it, and both forms exit 64 on malformed input either way.

Verification runs on the remote runner.

Refs #3912

* chore(#3912): backfill changeset pr numbers to 3983

* test(#3912): pin the reason-table invariant as set equality, not a count

A graph-backed review flagged the unchecked lookup in
expectedErrorCode3912. Investigated by execution: the drift guard DOES
hold — for an unmapped reason under v2 the production error() yields 1
while the table yields undefined, so the assertion fails. Not a
correctness defect, and deliberately NOT made tolerant, since a tolerant
lookup would destroy the guard.

Two real problems remained. The guard asserted the wrong invariant: it
counted the TABLE's keys at 25 rather than checking they match the
ENUM's values, so a renamed member keeps the count at 25 and slips past,
and a 26th member leaves the table at 25 and slips past too. Both were
then caught only indirectly, by an undefined mismatch producing 'must
exit undefined'. It is now a sorted set equality, so the failure names
the specific missing or extra reason.

And the comment above it described a '?? FAIL' fallback that does not
exist anywhere in the function. It now states what the code actually
does, verified by running it rather than by reading it.

Refs #3912

---------

Co-authored-by: sim <sim@local>
2026-08-28 08:09:05 -04:00
Tom Boucher
d98b55562c enhance(#3910): the raw terminator is banned by construction (#3980)
* enhance(#3910): move the last src/ terminators onto the seam

Phase 6 bans the raw terminator by construction, which it cannot do while
violations stand. A census found 12 sites the rule would flag; nine of the ten
unsanctioned ones were owned by no phase of the epic at all — a coverage hole
in the decomposition, since P0-P2 are infra, P3 the gate modules, P4 the
scanners, P5 the fragments, P7 the hooks, P8 io.cts, and P6 itself only adds
the rule. `src/**/*.cts` now holds exactly 2 raw exits, both inside
`terminateNow`, the single sanctioned site.

`io.cts`'s `error()` is the interesting one. It was first called substantive on
"dozens of callers, contract risk" — asserted, not measured, and the
measurement refuted it: 289 call sites, zero inside a try whose catch would
swallow a throw. The real obstacle was structural instead: `terminateNow`
cannot emit exit 1, because ADR-3889 §1 makes 0 and 1 unallocatable and
`nameForExitCode(1)` throws. So the only route is `ExitError` under `runMain`,
which sets exitCode and writes stderr only when the error carries a user
message — keeping the existing stderr write and throwing a message-less
ExitError is observably identical.

That census was still too narrow, and running the CLI proved it. It asked
whether the CALL sits in a try/catch; the two regressions that surfaced were
interceptors elsewhere on the stack:

- `command-routing-hub.cts`'s `dispatch()` swallowed the ExitError into a
  HandlerFailure, so the caller emitted a duplicated, wrong stderr line on
  every Hub-routed path. It now rethrows ExitError explicitly — the same shape
  `gsd-tools.cjs` already used at two dispatch sites, so this follows an
  established idiom rather than inventing one.
- the profile-pipeline router's deliberately un-awaited `.catch(e => error(...))`
  turned an ExitError rejection into an uncaught exception; it now mirrors
  runMain's handling.

`edge-probe` and `ui-consideration-probe` gained `runMain` wrappers because
probe-core's new throwing default would otherwise have escaped them.

A follow-up sweep of every dispatcher — 19 command routers, the Hub, the
gsd-tools dispatch seams — found no further swallowing catch. The admitted
bound: ~1260 non-rethrowing catches repo-wide were scanned structurally but not
individually classified. Both real regressions were found by execution, not by
reading, so the suite is the detector that matters here.

`gsd-tools.cjs:253` stays a raw exit deliberately: it is the ensureRuntimeBuild
bootstrap, which runs before cli-exit is required, so the seam does not yet
exist. It needs a second allowlist entry, which means #3910's "single allowlist
entry" criterion is unachievable as written.

Verification runs on the remote runner.

Refs #3910

* enhance(#3910): ban the raw terminator by construction

Adds local/require-registered-exit and registers it on all four globs:
src/**/*.cts, scripts/**/*.cjs, hooks/**/*.js, gsd-core/bin/**/*.cjs.

Registering on the .cts glob is load-bearing, not redundant — the emitted .cjs
mirrors are globally eslint-ignored, so a rule registered only on the emitted
globs is blind to the sources. That is the #3496 lesson, and it is how the
previous guard became invisible: n/no-process-exit was 'error' in one block yet
fired zero times on all three surfaces that mattered.

The dead n/no-process-exit: 'off' block for hooks is deleted in the same PR.
Phase 7 migrated every hook, so the exemption now protects nothing.

Two allowlist entries, not the one #3910 anticipated. terminateNow's body is
detected STRUCTURALLY — a process.exit lexically inside a function of that name
— rather than by a path and line number that rots. The second is
gsd-tools.cjs's ensureRuntimeBuild bootstrap, an inline disable with its reason
at the call site: it runs before ./lib/cli-exit.cjs is required, so the seam
does not exist yet and no migration is possible. #3910's 'single allowlist
entry' criterion is therefore unachievable as written, and is amended with the
measurement rather than quietly missed.

The rule is proven able to FAIL, per glob: four positive controls, one for each
registered glob. A guard that cannot be shown to fire is not a guard. Four
matching negative controls pin process.exitCode as never-flagged — conflating
it with process.exit is what inflated this epic's original census 2x. An
allowlist case and a near-miss (same shape, different function name) fix the
structural detection in place.

Verification runs on the remote runner.

Refs #3910

* fix(#3910): stop the detached catch from throwing, and scope the allowlist

Review findings, one of them a regression the previous fix introduced.

_handlePipelineRejection called error() from inside a DETACHED .catch().
error() now throws, so that throw became an unhandled promise rejection — and
on Node >=15 with --unhandled-rejections=throw, Node dumps a raw stack trace
with absolute paths on top of the clean Error: line. That was impossible before
this branch, because process.exit(1) terminated synchronously before any
rejection machinery could observe it. The handler now writes byte-identical
stderr itself, in both plain and --json-errors form, and sets exitCode in
place. This was the THIRD interceptor found, and like the first two it surfaced
by running the CLI rather than by reading code.

The rule's terminateNow allowlist had no path constraint, so any function
anywhere named terminateNow across all four globs inherited it. It now requires
the structural nesting check AND a cli-exit.cts basename — still no line
numbers to rot.

The four per-glob positive controls only varied a filename inside RuleTester,
which never resolves eslint.config.mjs. Since the rule is filename-agnostic,
all four exercised identical logic and none proved the rule was WIRED — this
epic's own failure mode. A registration test now asserts the rule resolves for
a real path in each glob, and it is proven able to fail: removing one glob's
registration flips the resolved value from [2] to undefined.

Three evasions the rule cannot catch (computed member, aliasing, .call/.apply)
are documented in its header and pinned by tests, labelled as known limits
rather than endorsed, so a future change that starts catching them is a
deliberate diff.

Refs #3910

* docs(#3910): document the raw-terminator ban

Reference and Explanation via a new docs/features fragment (FEATURES.md is
generated from it, not hand-edited). How-To:
docs/how-to/resolve-a-raw-terminator-finding.md, indexed from docs/README.md —
a contributor whose code trips the rule picks among three replacements by
surface (runMain/ExitError for a CLI path, terminateNow for a hook,
process.exitCode where the process should drain), and needs to know why
process.exitCode is correct and never flagged, since conflating the two is what
inflated this epic's original census 2x.

The page also names the three patterns the rule cannot catch and says plainly
that using one to dodge it is a review finding, not a fix — documenting them
without that sentence would read as a sanctioned workaround.

docs/INVENTORY.md deliberately untouched: eslint-rules/ is not a tracked family
in the manifest (verified — a regen produced a zero diff), so a hand-written row
would desync the table from the family it claims to belong to.

Refs #3910

* fix(#3910): a catch that sniffs the message swallows an ExitError

The remote run returned 41 failures, and one of them was a live production
regression rather than a test artifact.

`cmdMilestoneComplete`'s unstarted-phase guard re-threw only when
`e.message.startsWith('Cannot mark milestone complete:')`. `error()` used to
`process.exit(1)`, uncatchable, so the guard always fired. It now throws an
ExitError carrying no message, the string test fails, and the ExitError was
silently swallowed — the guard stopped blocking milestone completion entirely.
Proven against the real CLI: pre-fix, a milestone with an unstarted phase
archived at exit 0; post-fix it is blocked at exit 1 with the intended message.

That is a guard that silently stopped guarding, which is this epic's thesis
appearing inside the phase meant to enforce it. Worth stating plainly: an
earlier census DID examine this site, saw a `throw e`, and classified it as
rethrowing. It was wrong — the rethrow is conditional, and a conditional
rethrow on an inspected message is indistinguishable from an unconditional one
unless you read the predicate.

So the class was swept rather than patched where it was tripped over. An AST
census of every CatchClause across src/, gsd-core/bin/ and scripts/ found 38
conditional rethrows. Two more had the same defect and are fixed the same way:
`config.cts`'s `'No config.json'` sniff and `gsd-tools.cjs`'s
`e.name === 'WindowsError'`. The remaining 25 are provably unreachable — every
one wraps a bare fs, YAML, manifest-require or git-exec primitive that cannot
throw ExitError — and two were scanner false positives, both explained. Each
fix is an unconditional `instanceof ExitError` rethrow placed BEFORE any
inspection, matching the idiom command-routing-hub and gsd-tools already used.

Residual bound, stated rather than implied: zero known-reachable unfixed sites,
contingent only on error() never later being called inside one of those 25
primitive try blocks.

The remaining failures were harness artifacts, and the harnesses were corrected
to the new contract rather than the assertions weakened. Tests that mocked
`process.exit` to observe termination now catch ExitError and assert its code;
tests parsing stderr as a single JSON object still assert exactly that, with
their ad-hoc `node -e` scripts wrapped in runMain so it is true. milestone and
phase-resolution-parity needed no test change — they were correctly written
against the real bug and are what caught it.

Verification runs on the remote runner.

Refs #3910

* chore(#3910): backfill the changeset PR number

Also reframes the fragment to lead with the user-visible change — the
milestone guard blocking again — rather than the narrowest of the three fixes.

Refs #3910

---------

Co-authored-by: sim <sim@local>
2026-08-28 03:15:39 -04:00
Tom Boucher
2ea5efc151 enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build

ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the
epic's 128 terminators and cannot reach `terminateNow` today.

The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as
gsd-agent-isolation-guard.js already does for two other modules — is rejected.
That precedent carries its own warning (#3582): those files are tsc output,
gitignored and absent on a raw plugin-marketplace or git-clone install, so the
hook must first call ensureRuntimeBuild() to self-heal. Making the module a
hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and
its failure mode is precisely the fail-open this phase exists to remove: a
guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam`
already encodes that concern, and Design B would have had to add an
ensureRuntimeBuild() call to all 19 hooks to satisfy it.

So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the
registry, preserving the invariant `src/cli-exit.cts`'s own header states: it
imports nothing but node:fs and its sibling registry, and the generator
dual-emits that sibling alongside each copy so a relative require resolves next
to whichever copy loaded it. Shipping needed no change — build-hooks.js already
declares HOOKS_SUBDIRS_TO_COPY = ['lib'].

Proven, not asserted: the two files are copied into an otherwise-empty tmpdir
and a child process requires them and terminates — PASS exits 0, HOOK_DENY
exits 2 with the payload on both stdout and stderr. That test fails the moment
the hooks copy gains a require reaching outside hooks/lib/.

Also fixed inline: the registry's fifth target let any `--write` test overwrite
the real committed hooks/lib/exit-code-registry.js, because the test helper
derived only three of the other output paths. It now redirects all five, and a
regression test asserts every committed artifact is byte-identical after a
redirected write.

Install-tree goldens pick up the two new shipped paths across 11 runtimes —
insertions only, no removals. lint:ci was green while they were stale, so this
was found by regenerating rather than by a gate.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): declare a crash policy, and migrate the write guard

Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`,
hand-written because the cli-exit copy beside it is generated:

  allow(payload)          exit 0
  deny(payload, stderr?)  exit 2
  crash(onCrash, payload) whichever the hook DECLARED

`crash()` takes the policy as a required argument with no default, which is the
whole mechanism: fail-open by accident stops being expressible. A hook must
name ALLOW or DENY at the call site, and an unrecognized value terminates
INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission
does not.

`gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a
gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by
citing this hook's `emitBlock` — but modeled it as sending the same bytes to
both streams, when `emitBlock` actually sends full JSON to stdout and only the
bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim
back to the model. Migrating as written would have turned a readable sentence
into a JSON blob for Kimi-backed agents.

#3911 requires both "all 19 hooks terminate through terminateNow" and "no
hook's effective default changes". Those are jointly satisfiable only by
teaching the seam to carry a distinct stderr payload, so `terminateNow` gains
an optional third argument: omitted, behavior is byte-for-byte what it was; a
string is written raw, which is exactly the Kimi case. The doc comment's
inaccurate claim about emitBlock is corrected in place.

Proven rather than asserted: the pre-migration file is reconstructed from HEAD
and driven with the same catastrophic-shrink payload as the migrated one —
exit code, stdout and stderr all byte-identical.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): all 19 hooks terminate through the seam

Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk
now reports zero `process.exit(` call sites across every `hooks/*.js` — down
from the 91 the census measured.

Each hook with an outer catch declares its policy once, at module top, with the
reason that policy is right for that specific guard: a read guard that cannot
scan must not block the read; a statusline that renders every prompt must
degrade rather than crash; an injection scanner must not retroactively block a
result already returned. Those sentences are the deliverable — they are what
turns fail-open-by-accident into fail-open-on-purpose. No hook's effective
default changed.

Wiring exposed two defects, both fixed here rather than noted.

A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s
`emitForceAddBlock`, matching the pattern already known from the write guard —
full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the
`stderrPayload` argument added in the previous commit, which is now carrying
its second real caller rather than one special case.

More seriously, `terminateNow` emitted both streams inside ONE try, so a
payload that failed to serialize aborted before the stderr write ever ran. The
two windsurf guards write nothing to stdout on a block and only a reason string
to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny
that silently loses its reason, which is the exact "fails with success" class
this epic exists to close. The streams are now emitted independently, each with
its own guard, and `undefined` means "nothing to write for this stream" rather
than an error. Regression tests inject a throwing write on one fd and assert
the other still receives its payload; they fail against the single-try version.

Byte-identity was proven per hook, not assumed: each pre-change file is
reconstructed from HEAD and driven side by side with the migrated one across
its normal path, its deny path, malformed stdin and empty stdin — exit code,
stdout and stderr compared.

Verification runs on the remote runner.

Refs #3911

* enhance(#3911): harden the three shell hooks, and pin every hook's policy

`gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh`
gain `set -euo pipefail`.

The expected hazard did not materialize, and that is worth recording: every
intentionally-non-zero command in all three is already the condition of an
`if`/`elif`, which `set -e` never fires on, and none of them reads a
possibly-unset variable or pipes through a grep that may legitimately match
nothing. No `|| true` guards were needed. Each hook was still checked
command-by-command before the flags went in rather than after.

Twenty-one before/after cases across the three hooks — disabled and enabled,
planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload
shape, quoted and unquoted `-m`, valid and over-long Conventional Commits —
all match on exit code, stdout and stderr.

The hardening is shown to actually fire, not merely added: with a stubbed
`node` that fails at the JSON-emit step, phase-boundary and session-state go
from silently exiting 0 with empty stdout to failing visibly with the error
surfaced. No such case could be constructed for `gsd-validate-commit.sh`,
whose every statement already sits inside an if-condition — recorded as
unproven rather than claimed.

`tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks
for, table-driven over all 19 hooks rather than 76 hand-written cases: normal
allow, deny where a deny path exists, crash-honors-the-declared-policy, and an
unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The
deny assertions encode each hook's ACTUAL stream split rather than a uniform
shape, since four of the six deliberately differ. A drift guard enumerates
`hooks/*.js` and fails if a terminating hook is ever added without a row.

Writing those tests surfaced two hooks that emit a block decision in their JSON
body and exit 0. Both were checked rather than assumed, and neither is a
fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the
tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js`
follows Cursor's JSON-body protocol. They are deliberately left alone — a
mechanical sweep to `deny()` would have broken exactly these two.

Verification runs on the remote runner.

Refs #3911

* fix(#3838): the commit validator says when it could not validate

#3911 claims to subsume #3838. Measurement said otherwise, so this closes it
for real rather than by assertion.

`set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash
exempts a command used as an `if` condition from `set -e`, and all three of the
hook's swallow-and-pass sites are exactly that shape. Verified against the
hardened hook with a node shim that fails only the classifier call — a
non-conforming commit still exited 0 with empty stdout AND empty stderr,
indistinguishable from "your commit conforms". That is the defect verbatim.

All three sites named in #3838 now capture the real exit status instead of
consuming it as a condition, and each distinguishes its genuine negative from
"could not run":

- the classifier: 0 = is a git commit, 1 = genuinely not one, anything else =
  could not classify. Its `node -e` now wraps the require and the call in
  try/catch and exits 3 on a throw, so a broken require chain can never be
  mistaken for `isGitSubcommand` legitimately returning false — which is the
  arm that matters, since `token-scanner.cjs` is a gitignored build artifact
  and a fresh checkout lands there.
- the opt-in config read and the JSON command extraction get the same
  treatment.

On "could not run" the hook emits a diagnostic to stderr naming which check
failed and why, then exits 0. The issue confirms this is safe — it is a
PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it
the smallest sufficient fix. The gate still fails open, but it can no longer do
so silently, which is the whole complaint: a validator that disables itself
quietly costs more than one that is absent, because it is trusted.

Both controls are unchanged and pinned by tests: a conforming commit still
passes silently, a non-conforming one still exits 2 with its existing block
payload. The defect test asserts stderr is non-empty and names the failure; it
fails against the pre-fix hook.

Verification runs on the remote runner.

Refs #3911, #3838

* docs(#3911): document the hook crash-policy contract

Reference and Explanation via a new docs/features fragment (FEATURES.md is
generated from it), INVENTORY rows for the three new hooks/lib files, and an
ARCHITECTURE note on the hooks section.

How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md
— a hook author now has to choose and declare a crash policy, which is more
than one step and crosses into which harness protocol their hook speaks. It
covers allow/deny/crash, writing an ON_CRASH reason that is actually useful,
when a deny needs a distinct stderr payload, the two hooks whose harness reads
a JSON-body decision and must NOT use deny(), and what to do when a check
cannot run at all — with #3838 as the worked example.

Refs #3911

* test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup

Two review findings.

The acceptance criterion 'hooks/dist/** stays in parity via the build seam
(lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks
something else — that a hook requiring a compiled gsd-core/bin/lib module also
calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib
files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a
recorded ship-blocking bug where a new hook never shipped because a copy list
missed it. The suite now builds dist through the repo's own ensureBuiltHooks(),
byte-compares each shipped copy against its source, and spawns a child that
requires the SHIPPED dist copy and denies — which is what catches a copy that
exists but cannot resolve its sibling registry.

gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent
trap on EXIT replaces them, guarded so cleanup cannot alter the exit status.
Behavior-neutral across five cases, with temp-file counts taken before and
after each run.

Refs #3911

* fix(#3911): stage transitive hook lib requires, not just one level

The remote run returned 7 failures across 3 real causes.

The important one is a PRODUCTION bug this phase exposed rather than caused.
`writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly
one level deep and never re-scanned the lib files it staged for their own
sibling requires. Nothing had a transitive lib dependency before, so the gap
was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js
made real Cursor installs ship a bundle that dies at require time with
MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor
hook runs to completion.

The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so
the injection scanner crashed at require time and its exit-1 was being read as
a policy decision. Migrated to copyScriptWithDeps, which walks the require
graph — the repo's recorded rule for this class, since adding another
copyFileSync keeps it alive for the next person.

The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib
file it expected to be named in the abort message; the same throw now fires for
a different file first. Its assertion is unchanged in substance — staging still
must abort rather than ship a broken hook — only the name is no longer pinned.

The last one was my own test asserting an uppercase reason code. Measured
against origin/next: the pre-change hook emits the same lowercase
'config_unreadable', so the test was wrong, not the migration. Corrected to the
real value rather than making the code match the test.

Verification runs on the remote runner.

Refs #3911

* chore(#3911): regenerate the cursor install-tree golden

The staging fix means a Cursor install now correctly carries the two
transitive lib files it was silently missing. Additive only — no path was
removed. The golden diff is the evidence the packaging defect was real.

Refs #3911

* chore(#3911): backfill the changeset PR number

Refs #3911

* fix(#3911): a git probe that timed out is not a negative

A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just
past the 2000ms budget these hooks give their git probes. The three that passed
took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns,
the hook reads the non-zero result as "not a git repo", and allows with exit 0
and empty stdout AND empty stderr. Under load, the guards silently stop
guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks
this phase is about.

The repo had already recognized the class in one place — gsd-cursor-subagent-start.js
fail-closed-denies on `git_timed_out` (#3045) — but nowhere else.

`hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real
non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than
folding all four into `status !== 0`. Three guards route their eight git probes
through it.

The resolution is the same shape #3838 took, and the same one that issue
endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes
on any path** — a developer on a loaded machine is still not blocked, which
keeps #3911's declaration-pass contract intact for exit codes. What changes is
that the hook now says on stderr which probe could not answer, instead of
presenting silence as a clean verdict.

Scope was checked across every hooks/*.js, not just the three that failed:
gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only
a cosmetic display segment, not an allow/deny decision, and are left alone.

The C2 deny assertion was a real-race test — it demanded exit 2 while a slow
git legitimately yields 0. It now requires the hook to either deny, or allow
with a diagnostic naming the probe that could not run; a silent allow still
fails, so the assertion is not vacuous. A deterministic regression stubs git on
PATH to sleep past the budget rather than waiting for load to reproduce it.

Verification runs on the remote runner.

Refs #3911

* test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows

The deterministic timeout regression stubbed git on PATH and asserted the
guard reports rather than silently allows. It passes on Linux and macOS and
failed on Windows in 83ms and 176ms — the stub was never invoked at all.

Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on
Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The
git.cmd branch could not have worked and is removed rather than left implying
a Windows path that does. Adding shell:true to the hooks to serve a test would
change product behavior and widen an injection surface, so the case is skipped
on win32 only, with the mechanism written into the skip reason so a future
reader does not 'fix' it that way.

Linux and macOS keep the coverage, and macOS is where the underlying fail-open
was actually caught.

Refs #3911

---------

Co-authored-by: sim <sim@local>
2026-08-27 22:21:10 -04:00
Tom Boucher
fa41bfec5c enhance(#3942): the emitted-drift ack is PR-lifetime data — move it to a commit trailer (#3954)
* test(#3942): failing-first suite for the emitted-drift ack commit trailer

Binds 37 input classes from the phase test matrix to the behavior ADR-3942
specifies, before any of it exists. Stubs return benign empty values rather
than throwing, deliberately: several rows assert that something DOES throw
(cap overflow, uncomputable commit range), and a throwing stub would turn
those green for the wrong reason and destroy the red.

The two rows that carry the design's load:

- merge-base semantics. The range is $(git merge-base base HEAD)..HEAD, not
  base..HEAD, because changedPaths comes from `git diff base...HEAD` (three
  dot). Two-dot would let the ack set and the change set disagree about which
  commits are this PR's. The fixture forks a topic branch, puts a trailer on
  each side, and asserts only the topic-side trailer is in range.

- fail-closed on an uncomputable range. With fragments a depth-1 checkout
  passes VACUOUSLY, every fragment reading as brand-new. With trailers the
  range cannot be computed at all, and returning an empty set would silently
  disarm the gate, so it must throw. The fixture builds a genuine shallow
  clone rather than simulating one.

Also covers the self-inflicted case: this change's own documentation quotes
the trailer syntax, so an example landing at the end of a commit message would
arm a live acknowledgment keyed on the literal placeholder text. Keys carrying
angle brackets or whitespace are rejected.

Authored per the phase artifacts 40-design.md and 50-test-matrix.md.
Not yet run on the remote runner — this commit exists to be tested.

Refs #3942

* chore(#3942): move the emitted-drift ack to a commit trailer

Implements ADR-3942, superseding ADR-2719 section 3 and its #2789 amendment.
Sections 1, 2 and 4-7 are retained: the conservation law is unchanged, only the
storage of its escape hatch moved off the working tree.

An acknowledgment explains one PR's ripple, and the moment that PR merges the
ripple is in the base, so it can never clear anything again. It was stored in
permanent shared state anyway, and every consequence of that mismatch had to be
built and then maintained. The chain is #2789 -> #2914 -> #3078 -> #3842 ->
#3823 -> #3875, each fix generating the next defect, ending in a scheduled
sweeper whose own first PR could not merge itself.

Added
  parseAckTrailers + renderAckTrailer (pure) and readAckTrailers (IO shell),
  reading Emitted-Drift-Ack-Hash: / Emitted-Drift-Ack-Growth: trailers over
  the merge-base range. tests/emitted-ack-trailer.test.cjs, 37 cases, written
  failing-first and confirmed red before any of this existed.

Changed
  diffEmitted takes two structurally distinct key-space maps instead of one
  shared paths map. That closes a latent defect: the spaces were separated by
  convention only, so a growth key satisfied a hash lookup by naming
  coincidence. staleAcks now reports which space a key was declared in.
  REMEDIATION teaches the trailer, per space, with its example rendered through
  renderAckTrailer so the taught grammar cannot drift from what the parser
  accepts.

Removed
  the sweep workflow, the guard-no-ack-on-next job, the standalone linter and
  its lint:ci entry, the fragment directory and its three spent fragments, the
  legacy single-file union, and the baseAck/spentAcks mechanism -- spentness is
  now structural, not computed.

Two range properties carry the design and are pinned by tests rather than
asserted: the range is merge-base scoped, matching git diff base...HEAD, so an
already-merged trailer is out of range by construction; and an uncomputable
range throws instead of reading as zero acknowledgments, which is the inverse
of the fragment guard's vacuous pass.

Three deliberate observable changes, each disclosed in the changeset: the
unread runtime field is gone, the legacy file is no longer read, and cross-space
excusal no longer works.

Ten open PRs carry fragments and will meet a modify/delete conflict. Measured
before landing and accepted deliberately; the one-line migration is in the PR
body.

Verified: lint:ci exit 0. Remote runner to follow on this exact sha.

Refs #3942

* fix(#3942): silent trailer collapse, lost coverage, and an unbounded cap

Six findings from the orthogonal review round, all fixed in place.

BLOCKER -- two trailers of the same name on one commit collapsed silently.
readAckTrailers built `separator=1d` where git needs `separator=%x1d`: the
`separator=` value inside a %(trailers:...) placeholder is itself a
pretty-format string, so the bare hex was emitted as two literal characters
and the split on \x1d never matched. Two same-name trailers therefore joined
into one value with errors empty -- the first reason absorbing the second
entry's key. Silent truncation, the exact class MAX_ACK_TRAILERS throws to
prevent. Confirmed with od -c against real git output before and after.

The failing-first matrix did not catch it because its "both spaces coexist"
row uses Hash plus Growth -- different trailer NAMES -- so the value separator
was never exercised. Two regression tests now cover same-name trailers
directly.

Coverage recovered: normalizeAckReason and INVISIBLE stayed on the live path
via parseAckTrailers but lost every test when the old suite was pruned. Back
under test against the current surface -- all six invisible codepoints
individually, whitespace collapse, trim, CRLF, and two seeded fast-check
properties. Dropping any single codepoint now fails.

MAX_ACK_TRAILERS counted raw trailers before de-duplication, so one trailer
carried forward across rebased commits counted once per commit and could throw
on a legitimate branch. Now counts distinct entries; 100 identical repeats
dedupe to one.

diffEmitted validated baseline, current and changedPaths but not the new
ackHash/ackGrowth, so a bad shape raised an unhandled TypeError instead of an
error verdict -- the same defect shape this file documents for #2778.

Docs: CONTRIBUTING and TESTING-SUITES were rewritten only in their first
sections; the later passages still taught fragments, git rm and the deleted
guard, contradicting the new text directly above them. Finished.

Also extends lint-removed-but-needed to exempt docs/adr and docs/research.
That gate fails on any docs mention of a file deleted in the same diff, which
makes it impossible to document a deletion in the PR performing it -- an ADR's
whole job is naming what it retired. Exemption is narrow and comes with a test
proving the gate still fires for a live consumer elsewhere under docs/. A
guard that cannot fail is worse than no guard. Maintainer-approved.

CONTEXT.md names the retired machinery by role rather than by filename: its
generated projection lands in docs/, which that gate does scan.

Adds docs/how-to/acknowledge-emitted-drift.md. The required docs set is
Reference and Explanation, so the task quadrant can be empty with every gate
green -- and this change has a real multi-step journey, including the fragment
migration ten open PRs now need.

lint:ci exit 0.

Refs #3942

* docs(#3942): correct the duplicate-trailer rule in CONTRIBUTING

Both axes of the code review independently flagged the same passage, without
seeing each other's output.

It claimed two declarations of the same key are always "a hard, loudly-reported
error, not a silent last-wins". That is only half true, and the missing half is
the one contributors hit: identical declarations -- same key, same reason --
dedupe silently, because a trailer legitimately survives a rebase and reappears
on every rebased commit. Failing there would red a branch for doing nothing
wrong, which is exactly why the dedup exists.

Only a same-key/different-reason pair errors, and that one is a genuine
ambiguity about which explanation holds.

As written, the paragraph told a contributor that a rebase-carried trailer
breaks the gate -- the opposite of the behavior. CONTEXT.md's parallel entry
already stated it correctly; this brings CONTRIBUTING into line.

Doc-only, root-level markdown.

Refs #3942

* chore(#3942): backfill changeset PR number to 3954

---------

Co-authored-by: sim <sim@local>
2026-08-27 17:28:39 -04:00
Tom Boucher
03b7125293 enhance(#3909): a probe that could not run no longer asserts a verdict (#3944)
* test(#3909): failing-first suite for the fabricated probe fallbacks

Binds the four fabrication sites found by executing the surfaces (ADR-3889
failure class (c)), each with a positive control so an over-firing fix goes red:

- the blocking api-coverage.verify-pre gate certifying "no external-API
  integration" from a zero-byte phase scope
- the assumption-delta query route scanning an unresolvable phase section as
  the empty string and reporting it as an examined negative
- both capability fragments' probe fallbacks, which append a fabricated
  verdict rather than replacing, and fire on the legitimate exit-1 negative

Verification runs on the remote runner.

Refs #3909

* enhance(#3909): a probe that could not run no longer asserts a verdict

ADR-3889 Phase 5. Four sites turned a failed or unexamined probe into a
confident negative; each now reports what it could not establish.

- check api-coverage.verify-pre: a phase with no plan body and no roadmap
  section ran detection over zero bytes and PASSED the blocking seal gate,
  certifying "no external-API integration" from input it never read. It now
  holds with scope_unavailable. The discriminator is bytes examined, never
  signals found, so a phase whose plans are real and simply carry no API
  vocabulary passes exactly as before.
- query assumption-delta scan: an unresolvable phase section was scanned as
  the empty string and reported as an examined negative. It now returns
  {skipped, reason: phase_unresolved}, still at exit 0 — an ADR-2980 degraded
  result in the payload, leaving the gsd-tools exit projection to P8.
- both capability fragments: `|| echo '{"detected":false}'` appended rather
  than replaced, and fired on the legitimate exit-1 negative, so a correct
  answer and an honest skip both arrived as two concatenated objects. They now
  keep the probe's own payload and manufacture only an explicit
  probe_unavailable skip when the probe produced nothing at all.

Every registered outcome is more restrictive on a blocking gate, so this can
turn a false green red and never a red green.

Docs: FEATURES 156, CONFIGURATION (both keys), references/api-coverage.md
seal-time outcome table, and a new how-to for the reason-code vocabulary.

Verification runs on the remote runner.

Closes #3909

* test(#3909): correct the stale unknown-phase assertion

`unknown phase → detected:false, no throw (graceful)` scanned phase 999
against a two-phase roadmap and asserted `detected === false`. That pinned
the fabrication as intended behavior: the phase does not exist, so the
detector was handed the empty string and its "no core assumption changed"
answer described nothing that was ever read.

It now asserts the skipped-with-reason shape. The graceful-degradation
contract the test was actually protecting — the query succeeds and does not
throw on an unknown phase — is unchanged.

Found by code review, not by the author.

Refs #3909

* docs(#3909): author the FEATURES entry in its generator source

`docs/FEATURES.md` is generated by `scripts/gen-features.cjs` from the
per-feature fragments in `docs/features/`. The API-coverage entry was edited
in the generated file, so the next regeneration silently dropped it.

The text now lives in `docs/features/api-coverage-gate.md` and
`docs/FEATURES.md` is regenerated from it, leaving the shipped file
byte-identical and its content actually derivable.

Caught by `lint:generated-sync`.

Refs #3909

* test(#3909): bind the skip to "not found", and pin the discriminator

The first verification run went red on one case, and the case was wrong
rather than the code.

`getRoadmapPhaseWithFallback` returns `null` for an unknown phase and for a
missing ROADMAP.md, but for a section whose body is whitespace-only it returns
the heading line alone — which is not empty. So a body-less section WAS found,
and reporting `detected:false` over its heading is a real negative, not a
fabrication. The test had assumed the resolver yielded `''` there.

Correcting the test rather than the resolver keeps `skipped` bound to the
distinction the issue asks for — found versus not found — and avoids diverging
`assumption-delta scan` from `roadmap.get-phase`, which the fragment documents
as sharing one resolver.

Also adds the seeded property the test matrix had promised: for any plan body,
the scope read back is whitespace-only exactly when the body was. That pins the
gate's discriminator to bytes examined, so it cannot quietly become "no signals
found", across unicode whitespace and CRLF.

`docs/INVENTORY.md` picks up the reference doc's new seal-time outcome table —
surfaced by the co-change gate, not by a lint failure.

Refs #3909

* chore(#3909): backfill the changeset PR number

Refs #3909

---------

Co-authored-by: sim <sim@local>
2026-08-27 15:50:12 -04:00
Tom Boucher
929e02cb2c enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict

ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking
answer. Three families do exactly that today; this commit pins each one RED.

Measured on this tree, 2026-08-27:

  intel query, .planning/intel/file-roles.json nested 12000 deep
    -> exit 1, "Error: Maximum call stack size exceeded"
       searchJsonEntries / matchesInValue carry no depth parameter at all.
       The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage
       (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage
       never received it.

  same fixture nested 48 and 49 deep
    -> both return total=1 at exit 0, truncated=undefined
       Nothing distinguishes "searched to the bottom" from "stopped looking".

  query phase-plan-index, a plan whose depends_on names an unresolvable token
    -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it
                  in wave 1"]
       The token is never mentioned. computeDependencyLevels drops the edge
       with `if (!resolvedDep) continue;`, every plan becomes a root, and the
       tool then reports the author's correct wave: as the thing that is wrong.

  countPhasePlansAndSummaries with fs.readdirSync throwing EACCES
    -> hasContext:false, indistinguishable from a phase that simply has no
       CONTEXT.md. context_read_error is undefined.

The shapes these tests assert against, chosen here so the implementation has a
target rather than inventing one later: `truncated: boolean` on the intel query
result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and
`context_read_error: string | null` per analyzed phase.

Deliberately green, and they must stay that way — each stops the fix from
over-firing:

  depth 48 is found and NOT flagged truncated (the ceiling is inclusive)
  a shallow miss reports no truncation           (noise control, N1)
  10,000 siblings at depth 2 are unaffected      (the bound is DEPTH, N2)
  a genuine wave: mismatch on a fully-resolved DAG still warns (N3)
  a genuinely missing directory is absent, not an error
  the emitted depends_on display mapping still passes an unresolved token
    through verbatim — already pinned by the existing #3785 test, so no
    duplicate was added

T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the
real CLI and reads the emitted JSON, because a unit assertion on
computeDependencyLevels would have passed throughout #3427's life.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3885): no silent swallow, and no verdict manufactured from dropped data

Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed
into an output that reads as authoritative.

The recursion bound, restored but NOT verbatim (src/intel.cts)

  MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and
  matchesInValue, which carried no depth parameter at all. The bound existed in
  the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the
  surviving .cts lineage never received it — §8.3's "a consolidation may not
  delete an invariant along with the surface that held it", demonstrated.

  Measured before: a .planning/intel file nested 12000 deep exits 1 with
  "Error: Maximum call stack size exceeded". Reachable from a project document.

  The original returned a bare `false` at the ceiling. Restoring that verbatim
  would trade a crash for a silent "no match" when the truth is "I stopped
  looking" — the same class this epic exists to close, and ADR-3473 Decision 4
  forbids it. So the bound carries a truncation signal:

    nesting 47 -> found,     truncated false
    nesting 48 -> found,     truncated false      (the ceiling is inclusive)
    nesting 49 -> not found, truncated TRUE
    nesting 12000 -> exit 0, truncated TRUE, no RangeError

  A shallow document that simply has no match reports truncated FALSE — the
  flag means "I stopped early", never "I found nothing", or it would be noise.
  The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected.

The dropped edge is named, and stops being blamed on the author (src/phase.cts)

  computeDependencyLevels dropped every unresolvable depends_on token with a
  bare `continue`. Each drop makes a plan a root, so the whole phase collapses
  to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as
  the thing that was wrong.

  Before:
    warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in
                wave 1"]
  After:
    warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not
                resolve to any plan in this phase — edge dropped, wave placement
                for this plan may be unreliable"]

  The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG
  and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId
  stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is
  deliberately not built here. The emitted depends_on display mapping still
  passes an unresolved token through verbatim (#3785).

No artifact from failed inputs (gsd-core/workflows/review.md, #3352)

  A failed lane leaves no result file, so "every lane failed" is exactly "the
  aggregate JSONL has zero lines" — the gate condition already existed as a
  byproduct. REVIEWS.md is no longer written in that case, and the commit step
  is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT
  counted as a failure. Per-lane output and non-empty .err are preserved to
  .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that
  the lanes failed at all; the commit step names one file, never a glob, so the
  diagnostics are not swept in.

Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2)

  Four callers collapsed an EACCES on a phase directory into [] and reported
  hasContext:false — byte-identical to a phase that simply has no CONTEXT.md.
  Each now names the directory it could not read. A genuinely missing directory
  stays absent rather than becoming an error, which is what keeps the fix from
  over-firing.

Fatal errno folded into a retry set: audited, no defect found

  Reported as a verified negative rather than padded with a change.
  withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776;
  atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by
  construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they
  return or rethrow the final error rather than swallowing it. estimate-cli's
  sole caller surfaces that rethrow as write_error in its JSON output.
  Manufacturing a diff to make the checkbox look worked-on is the Goodhart
  outcome Decision 6 exists to prevent.

Disclosed: R46 (the commit step names one file, never a glob) is a real
regression guard but is NOT independently failing-first — the commit fence is
byte-identical pre- and post-fix, so it only fails pre-fix through its shared
extraction dependency. Recorded rather than claimed as fail-first.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence

Two review findings, both real, both in my own change.

An isolated adversarial review found the evidence-preservation block never
checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in
a SEPARATE fenced block. A disk-full or unwritable phase directory therefore
still destroyed the only copy of the failed lanes' output — reintroducing the
exact #3352 data loss this item exists to stop, inside the fix for it.

Preservation and cleanup are now one block, because each fenced block is a
separate execution and a shell variable cannot carry between them. mkdir -p and
each cp are exit-checked; cleanup runs only when preservation succeeded, and a
failure warns naming the intact run directory. "Nothing to preserve" is not a
failure and still cleans up. Driven three ways: success removes run_dir, failure
leaves it intact with the warning, nothing-to-preserve removes it. The failure is
induced by a file-vs-directory conflict rather than chmod 0o000, which root
bypasses.

The new unresolved-depends_on warning embedded a user-authored token verbatim:

  warnings: ["Plan 03-02: depends_on token \"evil
  Plan 03-01: FORGED WARNING\" does not resolve ..."]

The JSON wire form is safe, and the security reviewer judged it non-exploitable
for that reason. It is escaped anyway through formatDiagnosticToken — the helper
#3884 added one phase earlier for exactly this class. warnings[] is an array a
consumer naturally prints line by line, and not reusing the sibling fix is the
generative-fix-divergence shape this epic exists to close. The same treatment is
applied to context_read_error / phase_dir_read_error, which embed a phase
directory path a repository can choose, and to the fs error message, which
echoes the raw path itself.

Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow
array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per
negative space N2, disclosed rather than left to be discovered.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot"

Blocker from the round-2 isolated review, and it is my own inconsistency:
this phase applied "unreadable is not absent" to phase directories and left it
broken in the file it was already editing.

  chmod 000 .planning/intel/file-roles.json
  gsd-tools intel query <term>
  -> {"matches":[],"total":0,"truncated":false}  exit 0

safeReadJson swallowed every read failure and returned null, so an EACCES was
byte-indistinguishable from an absent file AND from a genuine no-match. Now it
separates three states: ENOENT stays silently absent, because not every project
has every intel file and intelQuery loops over all of them expecting misses;
EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel
file previously read as "no matches" too — same defect, same fix.

Threading that outcome through the other three callers found something worse
than the reported case. intelDiff returned no_baseline:true for a corrupt or
unreadable snapshot — not a silent failure but an actively FALSE verdict, telling
the caller they never took a snapshot when they did. That is §8.5's headline
case, so it is fixed and tested rather than noted. intelStatus and
intelApiSurface collapsed the same way; intelApiSurface additionally printed a
"not yet populated" banner that was simply untrue.

Every row is failing-first, including the absent-file ones — the field is new,
so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin
that the fix does not OVER-fire on the ordinary absent case, which is what would
turn this into noise on every project lacking an intel file. IO failure is
injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root
bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as
a test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): build the pathological intel fixture as text, not by stringifying a nested object

The remote runner came back red on Linux with two failures, both
T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on
macOS. The product was never at fault.

writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then
JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed
the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned.
Linux's container stack is smaller than macOS's, which is the whole of the
platform difference.

Measured, with the same document built as JSON TEXT so nothing in the building
process recurses:

  depth=100    rc=0 truncated=true
  depth=5000   rc=0 truncated=true
  depth=12000  rc=0 truncated=true
  depth=60000  rc=0 truncated=true

V8 parses this shape iteratively; only stringify recurses. The bound works at
every depth tried.

The fixture is now built by string concatenation. That is also the more faithful
input — a real deeply nested JSON document on disk is exactly what the bound
guards, where a stringified object was only ever a way to produce one.

The depth stays 12000. Lowering it would have made the test pass by weakening it
to accommodate a fixture bug, and 12000 is a legitimate pathological input the
product handles. T4 remains a genuine fail-first: rebuilt against the parent of
the commit that added the bound, the string-built depth-12000 fixture still
drives the CLI to rc=1 with "Error: Maximum call stack size exceeded".

A comment records why the fixture is text, so it is not "simplified" back into a
macOS-green / Linux-red test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3885): backfill the changeset PR number

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): normalize path separators before splicing into the workflow's bash

CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the
remote runner were all green.

  AssertionError: commit must name the single REVIEWS.md file; got:
    --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md

Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The
harness spliced an OS-native temp path into the extracted bash, and bash consumes
\U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke
RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory
survived — which is the other two assertions.

This is a fixture defect, not a product one, and that was checked rather than
assumed. In production the phase directory is toPosixPath-normalized at every
call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and
the run directory is created by "mktemp -d" running inside the bash block itself
(gsd-core/workflows/review.md:163), which emits POSIX-style output even under
Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it.

The file's pre-existing #3034 harness splices raw native paths too, but only ever
inside double-quoted assignments, so it never tripped this — my new harness
followed that convention faithfully into the one place where it does not hold.
Both now splice through toPosixPath from shell-command-projection, the
established seam, which is a no-op on POSIX and mirrors what production does.

No assertion was weakened. "commit must name the single REVIEWS.md file" and
"the run dir must still be destroyed" still assert exactly that; only how the
fixture supplies its path changed. Nothing is skipped on Windows — a t.skip()
here would have hidden the question of whether the exposure was real, which is
the question that mattered.

Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI
string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input
produces a byte-identical shape, proving the normalization is idempotent.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): stop the harness making the deleted run dir its own cwd

Windows shard 3/3 stayed red after the separator fix, on two assertions the
separator fix never touched:

  AssertionError: the run dir must still be destroyed
  AssertionError: nothing to preserve is not a failure — run dir must still be removed

The separators were a real bug and fixing them fixed the --files assertion. They
were not this bug, and two CI cycles went into the wrong axis before I stopped
converting path forms and looked at what the harness actually does.

runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's
working directory WAS the directory the block under test then removes with
rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified
locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live
process's working directory cannot be removed. So on Windows the directory
survived and both assertions failed, on macOS and Linux it vanished and they
passed. Nothing to do with slashes.

Harness-only. Production never cd's into the run directory; every reference is by
absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself
(gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged.

Fix: the child now runs with its cwd in an unrelated temp directory that the
block under test never deletes. Neither assertion was weakened, and nothing is
skipped on Windows — the tests in this file carry no platform guard and run
there unconditionally, which is how this surfaced at all.

Honest limit: the Windows failure mode cannot be reproduced on macOS, because
POSIX permits the very thing Windows refuses. The diagnosis is grounded in that
documented divergence and in the fact that only the Windows lane failed, but the
green outcome on windows-latest is unverified until CI runs it.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 04:12:47 -04:00
Tom Boucher
39673ae9ff fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config

Regression tests (RED first): --skills-root and gsd-tools query surfaces,
install-plan dest dirs, and converter skills-path rewrite.

* fix(#3738): antigravity global skills/agents install to ~/.gemini/config

Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents};
the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the
ADR-1239 skills/agents 'home' override on the antigravity global layout — the
same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references
in converted global content to ~/.gemini/config/skills/. configHome, settings,
probe/migration semantics, and the local .agents layout are unchanged.

* fix(#3738): retire deprecated configHome artifacts via installer migration 010

Next install converges an existing antigravity install: manifest-managed
skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does
not scan) are removed — modified files backed up first, unmanifested and
non-gsd entries preserved — and now-empty containers retired. Global scope
only; the local .agents surface is live. Docs + inventory updated.

* fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline

- bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/
  rewrite as src (ADR-1508 dual copy must stay in sync).
- Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so
  antigravity's emitted skills/agents stay differential-visible at their new
  install root; install-tree fixture regen confirms an unchanged key set.
- skills-from-commands rule declares the antigravity converter as a
  runtime-scoped transform; one ack fragment covers the identity-classed
  workflow whose antigravity copy embeds the old skills path.
- Migration 010 checksum baseline + home-override set doc updated; existing
  tests updated to the #3738 contract (global dest, golden parity via layout
  dest, integration expectations).

* fix(#3738): tolerate an absent extra emit root on baseline-side measurement

The base tree's installer predates the home override, so <HOME>/.gemini/config
does not exist there; walk() threw ENOENT and the in-job baseline build failed.
An absent extra root is the legitimate pre-override shape — skip it.

* fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment

- writeManifest resolves the agents-kind home override (_kindDestDirSafe), so
  the manifest records agents at their actual install root and drift detection
  keeps working (isolated review finding 1, major).
- Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing
  slash) divert to ~/.gemini/config/skills instead of falling through to the
  retired configHome path (finding 2).
- real-home-guard comment updated: antigravity's global agents kind is the
  first agents-kind home override (finding 3, doc-only).
- Regression tests for both behavioral findings.

* chore(#3738): changeset fragment (pr number backfilled after PR creation)

* chore(#3738): backfill changeset PR number (3921)

* fix(#3738): sandbox HOME in tests that install antigravity global artifacts

antigravity is the first home-override runtime in the golden-parity and
skills-wrapper suites (codex is not in their runtime lists), so those tests
never needed HOME sandboxing — the real-home guard now (correctly) refuses
their un-sandboxed global installs on CI, where HOME is the passwd home.

* fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans

K3's two back-to-back sandboxHome calls leave HOME pointing at the first
sandbox once the after-hooks restore (each call saves the env as it found
it, so the second saves the first's sandbox as 'original'). On the windows
matrix that leaked gsd-k3-qwen-* home into the L2 property, whose
antigravity/global run then (correctly) refused via the #3712 real-home
guard — antigravity is the runtime that made L2's plan escape into
os.homedir(). K3 now manages the env with a single restore; L2 sandboxes
HOME per run, mirroring L1.

* fix(#3738): L2 property's HOME sandbox must exist on disk

The #3712 guard's sandbox exemption fails closed when identify(effectiveHome)
is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir
under the real home) the antigravity/global run refused even with HOME
sandboxed. Create the per-run sandbox dir and clean it up.

---------

Co-authored-by: sim <sim@local>
2026-08-27 02:24:03 -04:00
Tom Boucher
e20744eacb enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick

ADR-3473 §8.4 says failure is a value. Three families currently encode failure as
success, and this commit pins each one RED before the fix lands.

Measured on this tree, 2026-08-26:

  gsd-tools generate-slug "test" --pick nonexistent
    -> empty stdout, exit 0                                     (#3365)

  gsd-tools audit-open --pick nonexistent_field
    -> dumps the entire human-readable audit report, exit 0

  gsd-tools generate-slug "Hello World" --raw --pick bogus
    -> prints "hello-world", another field's value, exit 0

  gsd-tools query state.planned-phase 3        (positional, no --phase)
    -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to
       "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted
       current_phase_name                                        (#3358)

tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract
("returns empty string for missing field", success === true). That assertion is
replaced by the required behavior rather than deleted.

The new parseNamedArgs block calls the spec-object signature that does not exist
yet, so it fails today by construction. The 11 existing behavior-lock tests are
left untouched here; they are corrected in the implementation commit.

C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180
Decision 4(b). A unit assertion on the parser would have passed throughout this
defect's life.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3884): failure is a value — strict argv, and --pick that signals absence

Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable
ways to say "I could not answer".

parseNamedArgs (src/command-arg-projection.cts)
  Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the
  hub's Result shape instead of a bare Record. Declaring the positional arity is what
  makes #3358's call site unrepresentable rather than merely detectable: an unrecognized
  flag or a token past the declared boundary is now InvalidArgs, naming the offending
  token and listing the accepted flags. The legacy positional-array call shape throws
  a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale
  hand-written .cjs call site fails loudly instead of destructuring undefined off a
  Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a
  projection over the one parser, not a second parser.

  Measured before, against a STATE.md with a populated phase-2 block:
    query state.planned-phase 3        (positional, no --phase)
    -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to
       "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted
       current_phase_name
  After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical.
  The flag form is unchanged and still updates STATE.md.

--pick <field> (gsd-core/bin/gsd-tools.cjs)
  extractField returns {found,value}, and the pick block no longer shares one catch
  between "output was not JSON" and "field was absent". An absent field exits 1 with
  pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1
  with pick_output_not_json instead of dumping the command's entire output. A field that
  is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer,
  not a failure, and it is what keeps `--pick count` printing 0 on a fresh project.

  Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable
  audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" —
  a different field's value, confidently, at exit 0.

  ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The
  sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the
  ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0
  would demote "could not answer" to "the answer is zero" — the hazard
  docs/how-to/resolve-unreachable-guard-findings.md already warns against.

Guard ledger (ADR-3473 Decision 6)
  scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a
  `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the
  correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls,
  a nullglob mechanism this change does not touch) is retained in full, as are the shared
  scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file
  is not deleted.

Call-site audit
  45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an
  if test, && chain, or a pipeline whose status is consumed, and no shell block in
  workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the
  prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind
  a prior found/existence check. No ADR-3409-class "field the command never produces"
  remains.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows

Two review findings, both fixed here rather than recorded as limits.

1. A newline in an untrusted token forged a second stderr line.

   Before, plain-text mode:
     $ gsd-tools query state.planned-phase $'foo\nError: forged second line'
     Error: unexpected positional argument "foo
     Error: forged second line"

   After:
     Error: unexpected positional argument "foo\nError: forged second line"

   --json-errors mode was never affected — io.error runs that payload through
   JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the
   three new InvalidArgs reasons plus the two new --pick diagnostics all
   interpolate a token that comes straight from argv.

   Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at
   every interpolation site — not a copy per site. It is deliberately NOT
   applied inside error() itself: several callers in this tree emit intentional
   multi-line diagnostics, and escaping newlines there would mangle them.

   The available-top-level-keys list needed the same treatment for a reason the
   review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user
   document and echoes that document's own keys into the diagnostic. Verified
   reachable — a frontmatter key containing a newline reaches the key list — so
   formatKeyForDiagnosticList is guarding a live path, not a hypothetical one.
   Ordinary keys still render plain and unquoted; a fix that merely dropped the
   key would also have passed a "one line" assertion, so the test pins the
   escaped key's presence too.

2. Five behavior-table rows were implemented but nothing pinned them:
   B7  a dotted path that dies partway
   B9  bracket syntax on a non-array
   B10 a negative array index, in and out of range
   B14 a JSON root that is not an object
   B17 an @file: payload over 50KB

   B17 is the load-bearing one. output() writes @file:<path> instead of inline
   JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no
   test, a future reordering of those two steps turns every large result into a
   false pick_output_not_json. The fixture seeds 1200 phase directories and
   measures the payload at 62474 characters, asserting the spill actually
   happened rather than assuming it.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): correct the strict-argv surface against a full verification run

The first full run came back with 90 failures across 12 files, none in the new
tests. They were the argv surface telling me what it actually is. Ten root
causes; each classified before anything was changed.

I over-implemented, and that is reverted.

  ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional
  tokens". It says nothing about a value flag whose value is missing. Making
  that an error was my design decision, not the rule, and it broke a
  deliberately recorded contract: `--prd` with no value resolving to null
  (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5;
  tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The
  "requires a value" branch is deleted outright rather than kept behind an
  option — an unused strictness mode is speculative generality. Unknown-flag
  and unexpected-positional rejection, which is what §8.4 actually mandates,
  is unchanged.

--wave needed a third flag kind the original design did not anticipate.

  `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the
  shipped workflow reconstructs and passes it (execute-phase.md:84), while
  #2932 records token-PRESENCE semantics: the CLI cares only that the flag
  appeared, and the value belongs to the workflow layer. That is neither a
  boolean flag nor a value flag, so `optionalValueFlags` now exists —
  presence-only in `data`, and the validation cursor consumes a following
  non-flag token so it is not reported as a stray positional. Every other
  declared boolean flag was checked against every argument-hint and prose
  usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one
  of this shape.

Five tests were pinning forms that never worked.

  tests/adr857-core-without-capabilities.test.cjs passed
  `init plan-phase --phase 01-stub`, but the documented form is positional
  (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form
  is the literal string "--phase". Measured on the pre-fix build against a
  real .planning/phases/01-stub/ directory:

    init plan-phase 01-stub          -> phase_found=true
    init plan-phase --phase 01-stub  -> phase_found=false

  The test asserted only exit 0 and key presence, so it had been green while
  proving nothing about phase resolution. Corrected to the documented form and
  strengthened to assert phase_found === true. Same class in state.test.cjs
  (`--plan-count`, a flag that does not exist; the real one is `--plans`),
  milestone-archive.test.cjs (`init new-milestone --json`, silently ignored),
  and concurrency-safety.test.cjs (a bare positional field name whose
  OR-assertion passed because a whole-document dump happens to contain the
  substring it looked for).

Six handlers had no argv validation at all — the same #3358 shape this phase
exists to close, found while fixing the rest: init verify-work / phase-op /
review / todos / remove-workspace read args[2] with nothing checking the rest,
and validate health read --repair/--backfill through a bare args.includes()
scan that bypassed the parser entirely. All now go through the seam, so the
flag has one owner.

tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must
NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's
local expectation does not override §8, so they are inverted and renamed —
a test still called "ignores an unrecognized flag" while asserting rejection
would be its own defect. Row C6's point is its PWNED canary; that assertion is
kept verbatim and only its exit-status expectation changed, because the
hostile token is now rejected rather than absorbed.

The blast-radius estimate in 40-design.md is corrected rather than quietly
left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was
accurate for what the graph can see — parseNamedArgs's callers. It cannot see
that those callers' handlers accept argv shapes wider than the code reading
args[2] suggests, which is where the real surface was.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert

Second full run: 46 failures, down from 90. Four causes, two of them mine.

Reverted `validate health` entirely — it was scope creep, and it broke a real flag.

  ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`.
  The previous commit routed `validate health` through the parser on the
  reasoning that a flag should have one owner. That was wrong twice over:
  §8.4 names parseNamedArgs and count queries, and `validate health` was never
  a parseNamedArgs call site — it read its flags, just not through the parser,
  so it had no silent-drop defect to fix. Tightening it omitted `--json`, which
  the health-diagnostic suites use heavily. The handler is now byte-for-behaviour
  back to its pre-branch form. `validate context` stays converted: it genuinely
  was a call site, and its `--json` is now declared rather than read by a second
  `args.includes` scan.

  The five handlers that had NO validation at all — init verify-work / phase-op /
  review / todos / remove-workspace — stay fixed. Those read args[2] with nothing
  checking the rest, which is the #3358 shape this phase owns.

Finished the A2/A3 revert. Three tests still encoded the deleted
"a value flag with a missing value is an error" rule, including one added by the
previous commit for that rule. All three now assert the reverted null contract,
and the ones whose titles said "rejected" are renamed — a test named for a
contract it no longer asserts is its own defect.

`--wave=` and `--wave --weird` are correctly rejected. Neither is documented in
commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and
neither is emitted by the shipped prompt layer, so both are unrecognized tokens
that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the
property it exists for — asserted directly now, at the parser, that `--wave` does
not swallow a following flag as its value — and only its exit-status expectation
changed.

A contradiction inside this branch, surfaced by the audit and resolved the safe way.

  Two pre-existing #3573 tests call `state begin-phase '2'` and
  `state planned-phase '2'` with a bare positional, relying on the old permissive
  parser to ignore it. This branch's own #3358 regression test requires that exact
  argv to be REJECTED. The two are mutually exclusive.

  Widening the router to accept a bare positional — mirroring complete-phase —
  would have silently re-opened #3358, and was verified to do exactly that: with
  the widened router, `query state.planned-phase 3` returned exit 0 and wrote
  current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and
  docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the
  two #3573 tests move to it. Their assertions were never about the call shape —
  only that total_phases survives the resync — and both still pass.

  complete-phase is untouched: its bare positional IS documented, and it keeps the
  dynamic boundary and the negative-space note that record why.

The audit that produced this is in the PR body: for every handler whose declaration
changed, the flags it reads anywhere in its body, the flags the shipped surface
documents, and the shapes the suite passes, compared. The `--json` miss was a
pattern, not an accident — declaring a handler's flags from its parseNamedArgs call
alone misses whatever it reads elsewhere.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3884): backfill the changeset PR number

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 00:12:13 -04:00
Tom Boucher
6b7df61938 enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises

ADR-3473 §8.1 carries a blocking open question with a forcing function: it must
be answered before any implementation PR for the rule opens. Answered here as (a),
a string-coercing adapter, with the measurement that settles it.

The sequencing note bet that §8.8's schema would make (b) tractable. Measured
against merged reality it does not: only 33 of extractFrontmatter's 78 non-test
call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists
survive real types, leaving ~31 lines across 3 call sites as the actual prize.

Also corrects three claims verified false while answering it. §8.1's justifying
sentence names #3349 and #3360 as defects a real parser would fix; both are
already fixed on next, confirmed by executing the compiled parser rather than
reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an
expected casualty of this rule, but it guards shell grep idioms in workflow bash
fences and never touches our parser. The same roster calls lint-vendored-deps.cjs
reusable as-is; it is hardcoded to re2js throughout.

The last two were caught by applying the rule this amendment records -- a factual
claim in this ADR is a hypothesis until the implementing phase executes it -- on
its first use.

Refs #3881

* docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable

An adversarial pass on the Phase 4 design established by execution that
extractFrontmatter is not a YAML parser but a line-oriented scanner whose output
is a function of raw source text. Four spellings of the same value collapse to
one js-yaml tree but produce four distinct legacy strings, one of them mangled.
No adapter over a tree can choose among outputs the tree does not distinguish,
so fork (a) -- keep a string-coercing adapter so the existing contract holds --
cannot be built. For any document with a non-scalar value, (a) collapses into
(b); about 26 percent of frontmatter-carrying documents have one.

Also records three design defects and one new attack surface, all confirmed by
execution: catching a parse failure and returning {} would delete the frontmatter
block on the next write at eight call sites that conflate empty with unparseable;
an empty value yields null where legacy yields {}, and reconstructFrontmatter
omits null-valued keys, so the shipped state template's empty progress key would
vanish; the #1882 truncation probe is parseYamlRegion itself rather than a
pre-parse heuristic, so it cannot both stay unchanged and survive that deletion;
and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB.

The rule is not deferred. The measurement is the deliverable and the re-scoping
is recorded as an open question with a forcing function, per section 8's own rule.

Refs #3881

* test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix

Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed.

Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states.

B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today.

B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today.

B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today.

Refs #3881

* chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest

Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496).

gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own.

src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare.

scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored.

docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds.

Refs #3881

* feat(#3881): parse .planning frontmatter with the vendored js-yaml

ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled
line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and
parseQuotedScalar are deleted (not patched); parsing now goes through the
vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA,
json: true }. Everything js-yaml does not do is layered on top, in one
place, carrying the seven design-doc consequences:

1. Empty value: a null js-yaml value is coerced to {} (matching legacy's
   own empty-value contract) so reconstructFrontmatter — which omits
   null-valued keys — still round-trips a bare `key:` line instead of
   deleting it. Verified live: progress: with no value survives
   parse -> reconstruct -> re-parse.

2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE
   Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS
   channel, is carried on the {} returned for malformed/refused YAML. Invisible
   to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that
   never inspect it are unaffected; wiring the 8 hasFrontmatter sites to
   consult it is a separate change, not done here.

3. Non-scalar object-list items (the four spellings of `- test: a b` that
   js-yaml collapses into one tree shape) are rendered as a canonical
   `key: value[, key2: value2]` string per item, keeping the existing
   array-of-strings value SHAPE. A full corpus differential over all 1702
   tracked markdown files found 11 residual divergences from the legacy
   parser (enumerated in the PR/report), most of them the parser now being
   MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key
   defect, a dropped Unicode key).

4. The #1882 truncation probe still runs the one real parser, but derives
   its key count from js-yaml's own thrown error and mark.line when the
   whole region doesn't parse cleanly (the dominant real truncation shape:
   fence opened, well-formed keys, no closing fence). Verified against both
   the clean-parse and the exception-fallback path.

5. The #3257 comment channel now attributes each pending column-0 comment
   against js-yaml's own parsed top-level key list (matched by literal key
   text, in document order) instead of the legacy ASCII-only key regex, so
   a comment above a Unicode key attaches correctly.

6. Anchors, aliases and merge keys are refused outright (a raw-text
   pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences
   today: zero. A 7-line billion-laughs fixture is verified refused rather
   than expanded.

7. A literal U+0000 is swapped for a private-use sentinel before the parse
   and restored in every resulting string afterward, since js-yaml rejects
   NUL unconditionally under every schema.

escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump()
(forced double-quoted style), with control-char hex escapes lowercased to
keep serialized output byte-stable (#1779 emitted lowercase); it keeps its
exported name and signature for its two other call sites (commands.cts,
runtime-artifact-conversion.cts), which need no change.

frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments,
regenerateFrontmatterKey's guard, noOpObjectListSetError and
parseMustHavesBlock are all unchanged — retiring them is fork (b) and is
not this phase.

Refs #3881

* fix(#3881): quote template placeholders and preserve unparseable frontmatter

SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as
bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax
under the vendored js-yaml parser, not the literal placeholder text
intended. Quote them so they parse as strings.

Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the
8 call sites in state.cts/state-transition.cts that compute
hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and
reassemble the document without a frontmatter block when false. That
check conflated 'no frontmatter' with 'unparseable frontmatter' (both
parse to {}), so a document with a merge-conflict marker or refused
alias in its frontmatter had that block silently dropped on write.
Each site now preserves the exact raw bytes stripFrontmatter removed
when the marker is set, leaving the genuinely-empty case unchanged.

Refs #3881

* test(#3881): consequence and boundary coverage for the js-yaml migration

Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix.

Refs #3881

* docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to

Refs #3881

* docs(#3881): correct the frontmatter glossary entry

Two errors in the entry as first written: it named parseYamlRegion as part of
the read path when that function is deleted, and it recorded the eight
hasFrontmatter call sites as unwired follow-on work when they were wired in
e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the
frontmatter block independently, so the marker binds at the transform layer.

Refs #3881

* docs(#3881): record the semantic-migration decision and the counted guard ledger

The maintainer chose the full semantic migration over splitting the rule into
its own epic or patching the scanner, so section 8.1 is answered as "the fork
was ill-posed and the migration is semantic" rather than as (a) or (b).

Also replaces the pre-implementation guess that this phase would shrink the
guard surface with the counted result: excluding vendored third-party lines the
hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines
despite four functions being deleted, because the compatibility layer over
js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is
therefore not delivered as written; what improved is the kind of code
maintained, not the amount. Decision 6 requires recording that rather than
netting it away.

Refs #3881

* chore(#3881): changeset for the vendored YAML parser migration

Refs #3881

* test(#3881): golden parity, round-trip property and packaging coverage

Refs #3881

* fix(#3881): refuse anchors structurally and fold in review findings

ADR-3473 §8.1 review findings, addressed inline:

Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched
only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping
({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor
mechanics while never matching that line shape, so the exact expansion the
guard exists to stop went straight through unrefused (a 303-byte quoted-key
bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener`
callback, which reports `state.anchor` for every event belonging to an
anchored node in every spelling, and throws from inside the callback to abort
before any expansion (~1-2ms vs full expand-then-discard). A merge key with
an alias is still refused (merge always requires a previously anchored node,
so the alias itself trips the listener); a bare merge key with NO alias is no
longer separately refused, documented as intentional: FAILSAFE_SCHEMA never
resolves `!!merge`, so it carries no expansion risk. Table-driven tests added
for all four bypass spellings + merge key, plus a quoted-key-spelled
billion-laughs fixture registered in the adversarial matrix and README.

Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/
aliases were "simply UNREACHABLE from typed code" through the twin. Corrected
to state the truth: anchor/alias resolution is document-level `load`
mechanics reachable through exactly the declared surface, and refusal is
enforced at RUNTIME (Finding 1's listener), not by the type surface.

Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was
non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree
back to NUL, including one the document author legitimately wrote, silently
corrupting it. Now refuses outright whenever the raw region already contains
U+E000 (consistent with the existing anchor/merge-key refusal path), making
the substitution provably injective. Tests added for a real NUL alone
(preserved), a pre-existing U+E000 alone (refused, not corrupted), and both
together (refused, not merged into one byte).

Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead
for a hand-authored row (only read inside the upstream-verbatim branch) —
exactly how Finding 2's stale docblock drifted unnoticed. Added
checkHandAuthoredTwin: every value-level export the twin DECLARES must be an
actual own property of the vendored runtime module at require-time. Tests
added, including a sensor that a declared-but-nonexistent export IS caught.

Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/
unparseableFm/reassemble preamble, copy-pasted at 7 sites in
state-transition.cts plus a sixth hand-inlined copy in state.cts's
cmdStateCompletePhase, is now one exported helper
(beginFrontmatterReassembly) every site routes through, including the
hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore)
keep a literal `body = stripFrontmatter(content)` assignment alongside the
helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward
scan (which does not chase aliases) still sees the strip; stripFrontmatter is
pure/idempotent so the extra call changes nothing observable.

Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a
separate change" claim (the 8 call sites are wired on this branch) and the
changeset's backlink from (#3473) to (#3881).

Finding 7: fixed the lint:ci failures blocking the gate — an
@typescript-eslint/only-throw-error violation from throwing a bare Symbol as
the anchor-detected signal (now a real Error subclass), unused-var warnings
left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by
two migration-specific test files (allowlisted with justification), and the
lint-state-write-path-drift false positive from Finding 5's helper (fixed
above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already
carried an explicit timeout; no change was needed there.

Golden fixture: added a golden entry for the new
anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line
scanner would also produce, since it independently dropped every quoted
top-level key). No other corpus document diverges: real .planning/ documents
carry zero anchors/aliases/merge keys/U+E000 today.

Refs #3881

* fix(#3881): fold in second-round review findings

Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration
ASCII-only key regex for the Unicode fixture; updated to require the 相
key's value now that js-yaml has no such restriction. Audited the rest of
the file for other pre-migration pins (block scalars, quoted keys,
flattened values, empty values, duplicate keys, unclosed blocks, null
bytes) by execution against real fixtures; found none regressed.

Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to
parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts
so no function still answers to the deleted hand-rolled scanner's name
(ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three
external call sites (src/commands.cts, src/runtime-artifact-conversion.cts)
updated in the same change — a mechanical rename, not an ADR-amendment
matter.

Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found
by execution. A top-level key named constructor/__proto__/toString/
valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read
resolving an inherited Object.prototype member); a key literally named
__proto__ was silently DROPPED entirely (bracket assignment on an
ordinary {} invoked the inherited __proto__ setter instead of creating a
data property). Fixed by building every parsed Frontmatter object with
Object.create(null), and replacing an `in` check with hasOwnProperty.call
in propagateCommentChannel. Added round-trip tests for all five hostile
keys, each with its own leading comment.

Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed
full byte-stability across the migration. Verified by execution: BEL/NUL/
NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw-
literal forms. Proved round-trip equivalence (each escape re-parses to the
exact source codepoint) and corrected the docstring. Found and fixed a
related real defect while verifying: a lone UTF-16 surrogate was emitted
BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely
unparseable YAML that silently collapsed to {} on re-read — extended
scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped
path.

Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real
truncation shapes (unquoted colon, open flow collection, mis-indented
sibling key, refused anchor). Root cause: the mark-based prefix recovery
excluded the very line whose key needed counting, and a mark-less refusal
never entered the recovery branch at all. Fixed by taking the max of two
lower bounds: the longest parser-verified line-prefix, and a raw-text
count of key-shaped lines (reusing the same key-shape pattern this file
already uses for isFrontmatterShaped). Extended test-matrix row A5
table-driven over all 4 regressed shapes.

Finding 6: the design doc's claim that no test owned the #3594 adversarial
fixture corpus was false — consolidation epic #1969 had already folded it
into tests/frontmatter.test.cjs. An earlier commit on this branch
re-created a standalone duplicate under that false premise; folded its
genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures,
block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the
duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the
ADR's §8.1 note, including the roadmap-sibling claim (no such file exists).

Finding 7: the golden serializer sorted object keys, making it structurally
blind to the key-order-parity invariant ADR-3473 §8.1 actually claims.
Made it order-preserving and regenerated the golden fixture from a
standalone compile of the legacy (pre-#3881) parser at ddde001af; the
current parser matches it with zero undocumented divergences, confirming
key-order parity genuinely holds. Extended row A2 table-driven across 6 of
the remaining 7 transitionCore kinds (all pass) plus documented, by
execution, a newly-discovered 8th-site regression: state.cts's
cmdStateCompletePhase calls the same preservation helper but its result is
clobbered by a later unconditional resync — filed as a distinct finding
rather than fixed here (touches syncAndPreserveStateMd, outside this
change's verified scope).

Refs #3881

* fix(#3881): preserve unparseable frontmatter through the CLI write path

Characterization (executed, before/after shown): case (b), not (a). The
frontmatter FENCE survives — `state complete-phase` on a conflict-marked
STATE.md returns success and a well-formed, freshly-derived frontmatter
block, not a document with no frontmatter at all. But the block's actual
content (the merge-conflict markers, and with them any signal to a human
that the document was in conflict) is silently discarded and replaced.

Root cause was two clobber sites, not one:

1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved
   `transformedContent` from readModifyWriteStateMd, finds {} + the
   FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh
   frontmatter block from the body anyway.
2. Even after (1) is fixed, applyPostSyncPreservation's own
   postFm/applyStatePreservation/authoritativeFm-reassertion machinery
   re-extracts frontmatter from syncedContent, restores curated fields
   from the pre-write snapshot, and reconstructs a NEW block again —
   confirmed live via `state begin-phase`, which still lost the markers
   after fixing (1) alone.

Both are now guarded by the same predicate (isUnparseableFrontmatter,
checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not
parse and the caller is not on ADR-3408 §8.3's closed "body wins" list,
both functions return their input content unchanged rather than
re-deriving over it. The closed list (cmdStateSync #905,
/gsd-health --repair's REGENERATE_STATE, both routed only through
writeStateMd, which never reaches applyPostSyncPreservation and passes
sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is
untouched — neither widened nor narrowed; verified by execution that
`state sync` still overwrites the conflict-marked block exactly as before.

Other verbs sharing the same readModifyWriteStateMd path were checked and
were equally affected before this fix: state update, query state.patch,
and state begin-phase all lost the conflict markers (RED, shown by
execution), and all three now preserve them (GREEN). Covered table-driven
in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe
block, which drives the real CLI verbs via runGsdTools — not just the pure
transitionCore layer the earlier A2 rows exercised — plus a control
asserting state sync's body-wins contract is unchanged.

Refs #3881

* fix(#3881): restore the parse surface's prototype and fix remote-runner failures

Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched.

Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write.

Refs #3881

* chore(#3881): backfill changeset PR number

Refs #3881

* test(#3881): make golden parity resistant to unrelated tree churn

A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus.

Refs #3881

* test(#3881): make the parser golden hermetic instead of tree-keyed

This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the
exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md,
docs/*.md) the prior golden pinned by tracked path. Any PR editing one of
those files' frontmatter for reasons unrelated to the parser (an
argument-hint addition, an allowed-tools tweak) turned the suite red, and the
reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to
catch a real regression. Excluding .changeset/** was not enough; the design
itself was wrong: a regression fixture must not be keyed to mutable repo
paths, and a single 376-entry JSON every such PR touches is also a
guaranteed merge-conflict surface.

Rebuilt the fixture to carry its own documents: each of 51 entries stores a
stable id, literal documentText (shrunk from a real ddde001af-era corpus
document), and an expectedParse captured independently from the
pre-migration legacy parser (git show ddde001af:src/frontmatter.cts,
compiled standalone against its byte-identical sibling modules). The test
reads no tracked path, shells out to no git command, and enumerates no tree
-- a PR editing commands/gsd/help.md cannot affect it. Every entry's
reconstruction was verified at capture time to reproduce both the current
and legacy parser's output on the original document; 0 of 51 candidates
were dropped by that check (1, the deliberately-unterminated
unclosed-block.md adversarial fixture, has no closing fence to truncate at
and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now
diverges:true entries) and the D2 order-preserving structural serializer
that keeps the comparison from passing vacuously; dropped the
tree-enumeration helpers, the coverage floor, the post-capture-skip logic,
and the vanished-file check -- all artifacts of the path-keyed design.

Refs #3881

* fix(#3881): resolve vendored-deps paths independently of cwd shape

Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on
windows-latest CI: the test passed absolute scratch-file paths into
compareFiles()/checkRow(), whose helpers joined every input onto ROOT
via path.join(ROOT, rel), producing garbage when the input was already
absolute. It surfaced on windows-latest specifically because GitHub's
Windows runners checkout the repo on a different drive than TEMP, so
path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged
(no relative traversal is representable across drives) rather than the
relative form the test assumed. The remote gsd-test runner this repo
gates pushes on is Linux-only and could never have caught this;
GitHub CI's windows-latest job is the only signal that does, and it did.

Fixed the helper itself (scripts/lint-vendored-deps.cjs's new
resolvePath()) to treat an already-absolute input as absolute-in,
absolute-out instead of silently mis-joining it, and updated the test
to pass the scratch file's absolute path directly rather than relying
on a relative conversion that is not always representable. Kept every
mutation-sensor assertion intact and added coverage proving
resolvePath is a no-op for relative inputs and correctly passes
absolute ones through unchanged.

Refs #3881

* fix(#3881): warn when state sync regenerates over unparseable frontmatter

state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly
overwrites an unparseable frontmatter block per its 'body wins'
contract — that overwrite behavior is unchanged here. The defect was
the silence: synced:true/exit 0 gave no signal that the existing
block (including git merge-conflict markers) could not be parsed and
was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be
reported as authoritative when the derivation dropped input it could
not resolve') and §8.4 ('failure is a value').

Adds a gsd: warning — ... (#3881) line on stderr, matching the
existing #3573 precedent, and surfaces the same disclosure in the
JSON result's existing changes[] array so a machine consumer sees it
too. Exit code and synced:true are left unchanged — sync did what its
contract says.

REGENERATE_STATE (/gsd-health --repair's sibling on the same
sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally
refused by applyRepairs's dispatcher before runRepairAction ever runs
(src/health-diagnostic.cts), so it is not a live path today and is not
in scope for this fix.

Refs #3881

* fix(#3881): exit non-zero when a state command returns an error

Refs #3881

* chore(#3881): changeset for the state exit-code fix

Refs #3881

* fix(#3881): honor the documented --project-dir flag

Refs #3881

* revert(#3881): restore exit-0 result envelopes for state errors

Reverts 9638f2936 and its changeset. The change was wrong and the revert is
the correction.

This repo distinguishes two error mechanisms deliberately. error() in
src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure
path. output({error: ...}) writes a JSON result envelope to stdout and returns
normally with exit 0. The reverted commit converted 23 result-envelope sites
into hard failures, which is a different contract, not a bug fix.

tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope
contract directly -- a failing command exits 0 with a JSON error envelope and
must not publish state.json -- and the remote matrix run caught it along with
four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that
the original commit rewrote were encoding that real contract, not the bug it
claimed; they are restored.

Whether an error envelope on stdout with exit 0 is the right CLI design is a
genuine question, and it is section 8.4's rule ('failure is a value') with its
own phase. It is not something to flip inside this PR.

Refs #3881

* chore(#3881): backfill changeset PR number for the project-dir fix

Refs #3881

* test(#3881): keep the frontmatter mutation shard inside its time budget

The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard
cap. Root cause is NOT row-level spawn overhead (contrast the #2790/
core-utils precedent): the three shard test files' own logic runs in
~413ms total (356+30+27ms) with all 392 assertions passing. Instead,
src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to
the vendored YAML parser, proportionally growing the mutant count Stryker
generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner
bills the full 'node --test <3 files>' invocation once per mutant, and
node:test's default per-file process isolation forks a child process for
each of the three files on every one of those invocations — pure fork
overhead multiplied by a much larger mutant population.

Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares
isolation: 'none', and .github/workflows/mutation.yml passes
--test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e.
unchanged behavior — for the other 8 shards, which were not individually
audited for cross-file state leakage under shared-process execution).
Measured locally via node:test's run() API on the exact 3-file set:
isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same
392 passing assertions. The true CI-shard number can only be confirmed
on the GitHub Actions run (Stryker cannot run locally, and 'node --test'
is hard-blocked in this environment).

Refs #3881

* test(#3881): register the vendored-parser tests in the frontmatter mutation shard

stryker.config.mjs's own rule ("Keep this list in sync with the tests
arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew
src/frontmatter.cts from ~825 to 1496 lines but its new tests
(tests/feat-3881-yaml-parser-consequences.test.cjs,
tests/frontmatter-golden-parity.test.cjs,
tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in
tests/frontmatter.test.cjs) were never added to the frontmatter shard's
tests array, so Stryker's mutants in the new vendored-js-yaml adapter had
nothing constraining them. PR #3888 measured 55.8% against the 65 floor
(748 killed / 593 survived / 17 timeout) and the shard was separately
cancelled at 15m04s against the 15-minute per-shard cap.

Registers all four files (each earns its slot on evidence of a unique
constraining assertion, documented inline), gives the shard a
measured/projected 180-minute budget via a new per-module
timeoutMinutes field threaded through mutation.yml's job-level
timeout-minutes the same way isolation is threaded, and removes the
prior isolation:'none' override (re-measured at this file-set size, its
savings are within run-to-run noise, not worth the unaudited
cross-file-state-leakage risk).

Refs #3881

* feat(#3881): derive the mutation test list and ratchet the score floor

Refs #3881

* test(#3881): ratchet five stale mutation floors and close the frontmatter gap

Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1):
config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78,
context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both
scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs
RATCHET_BASELINE in the same diff per the ratchet's own contract.

Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in
tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented
near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order
semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's
leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter),
repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via
extractFrontmatter), and the null-byte sentinel round-trip surviving at region
offset 1. Did not lower minScore.

Refs #3881

* test(#3881): decouple the ratchet test from real module floors

The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour.

Refs #3881

* refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency

Refs #3881

* fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs

A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives.

Closes #2571
Refs #3881

---------

Co-authored-by: sim <sim@local>
2026-08-26 19:29:32 -04:00
Tom Boucher
ddde001af6 enhance(#3873): the STATE.md schema — one owner, generated artifacts (#3880)
* test(#3873): failing-first locale parity, plus tripwires for what must not move

Pins ADR-3473 §8.8 at the artifact a reader actually sees. The English STATE.md
reference carries a Status lifecycle section that is missing from all four
translations — the section documenting the status enum whose clobbering is
#3853. The test derives the heading set rather than hard-coding the missing
one, and names the locale and the heading when it fails.

Two tripwires that must pass today and after. The field-drift guard still
catches a re-derived fallback ladder: §8.8 instructs deleting that script, and
that instruction rests on a wrong premise about what it guards, so the test
stops a future reader from deleting it on the ADR's word. And last_activity's
label resolution is pinned to what ships today, because it is declared in one
of the two tables this phase consolidates and not the other — the
consolidation must not silently pick a side.

The locale test buckets under docs rather than state, which is what it tests;
that bucket is allowlisted with justification rather than folded into an
unrelated docs suite. It reads only markdown, so it carries no allow-test-rule
marker — a marker there would suppress nothing and would grow the unverified
pool against its ceiling.

Refs #3873

* feat(#3873): one schema owns the STATE.md key set, three tables become projections

ADR-3473 §8.8. The key set was declared in four places that had to agree by
hand and already did not: FIELD_CLASSIFICATION, FRONTMATTER_BODY_SOURCE,
FRONTMATTER_KEY_TO_BODY_LABEL and buildStateFrontmatter's emit behavior. One
frozen null-prototype schema now declares each key's type, enum, cardinality,
source, preservation, body source, body label, accepted parse shapes and
whether it is emitted unconditionally; the three tables are derived from it at
module load.

The projections are byte-identical to the literals they replace, key order
included, and the parity tests compare against verbatim copies of today's
tables rather than re-deriving both sides from the schema — a parity test fed
from one source proves nothing, which is how a consolidation ships a changed
policy under a green test.

last_activity was the live disagreement: present in one table, absent from the
other. The schema declares what ships today rather than the tidier answer, and
a test pins it.

The schema is a leaf module and owns the four field-policy types, re-exported
from state-transition so existing importers are untouched — the same split
health-diagnostic-types made to break a CJS require cycle.

Refs #3873

* feat(#3873): generate the schema-derived regions, parity-check the prose tables

ADR-3473 §8.8's generator half. gen-state-md-docs.cjs owns marked regions in
the shipped template and all five reference docs, follows gen-features.cjs's
fail-closed contract, and is wired into regen:derived and lint:generated-sync.

The Status lifecycle section was missing from all four translations — the
section documenting the status enum behind #3853 — and is now generated into
every locale. Field cardinality is a new generated table: pure schema data,
no prose, so nothing to lose.

The Field-reference and Status-values tables are parity-CHECKED rather than
generated. Their Purpose, When-populated and Matched-text columns are
genuinely hand-translated per locale, and §8.8 itself says prose stays
hand-translated; generating them from an English registry would overwrite four
locales' translations on every write. The row set is checked against the schema
instead, so a key added to one and not the other fails, which is what field
drift actually means. Building that check found last_activity_desc
undocumented in all five tables.

Three keys the docs describe are absent from the schema — active_phase,
next_action, next_phases. They are grandfathered by name, not by wildcard, so a
fourth fails: a declared gap with a forcing function rather than a silent one.

Refs #3873

* fix(#3873): declare what the parsers do, and close the shape-parity gap

Two declarations in the new schema described intended behavior rather than
actual — the defect class this epic exists to end, committed inside the epic.
Both were caught by executing the parsers instead of reading their docstrings.

current_plan.acceptedShapes claimed ['N', 'N of M']. Standalone, the hybrid
shape errors; the path that looks like support is parseInt truncating '2 of 5'
to 2 and discarding the rest. Narrowed to ['N']. The parser is deliberately NOT
fixed here: that is #3784 and PR #3791 is already doing it. When #3791 lands
this row must widen, and the shape test will go red until it does — the schema
and the parser cannot drift apart quietly, which is what §8.8's checked-not-
generated rule is for.

STATUS_LIFECYCLE_ENUM claimed to be the closed set status can hold.
normalizeStateStatus passes unrecognized prose through unchanged, so it is not
closed at runtime. The seven members are the canonical values it maps onto; the
docstring now says that and the test asserts the real lenient contract.

Closes the acceptance item that a test asserts the parsers accept exactly the
declared shapes: the check is table-driven over every row carrying
acceptedShapes, guarded against passing vacuously on an empty set, and fails
loudly if a future row has no registered driver. Adds the unwired-label throw
and the fast-check property that every projection agrees with its schema row.

Refs #3873

* fix(#3873): keep the shipped template's frontmatter first, and make row 27 able to fail

The remote matrix caught 12 failures with one cause. Making the template's
frontmatter a generated region wrapped it in its own yaml fence ahead of the
markdown fence, so extractFileTemplate and readShippedStateTemplateBody — which
both match the single markdown block — found the heading first, not the
frontmatter. That breaks the contract every new project's STATE.md is created
from: bug #21 and epic #1969 B8 pin that the File Template block starts with
frontmatter and carries gsd_state_version.

The markers now sit inside the single markdown fence, so the fence opens before
the frontmatter and the region still ends ahead of the heading. Same layout as
before this phase, with markers embedded rather than a second fence.

Row 27 existed to catch exactly this and did not, because it was writer-seeded:
it asserted against the generator's own output shape, so it passed on the broken
template. It now parses the fence the way production does and was verified to
fail against the broken shape before being trusted against the fixed one. A test
that would not have caught the bug it exists to prevent is worse than no test.

The emitted-attribution failure was separate and the fragment was the wrong
remedy: gsd-core/templates/state.md self-attributes under a verbatim-copy
identity rule, so a diff touching it needs no acknowledgment. Fragment deleted
rather than left explaining nothing.

Refs #3873

* docs(#3873): how to change the STATE.md schema

The phase gate was right and my docs artifact was wrong. I listed
lint:generated-sync as the second enablement step, which is a verification
command dressed as one, and then claimed a one-step sequence owed no how-to.

The real sequence is build:lib then regen:derived, and the ordering is a trap:
the generator reads the COMPILED schema, so regenerating before building
regenerates against the previous schema and commits artifacts that look
plausible while disagreeing with the code just written. A reference table
cannot carry an ordering dependency; that is what the how-to test is for.

The page covers adding, changing and removing a key, every reason code the
check emits and what to do about each, what is generated versus hand-translated
and why the two prose-bearing tables are parity-checked instead of generated,
adding a language, and the three grandfathered keys. Indexed from docs/README.md.

Refs #3873

* chore(#3873): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-26 01:57:47 -04:00
Tom Boucher
86fa2917d7 enh(#3866): dispatch step and contribution hooks at verify:pre (#3869)
* test(#3866): pin that verify:pre must dispatch every hook kind

verify-work.md's verify_pre_hooks step dispatches only `kind == "gate"`, so
getWiredKinds reports verify:pre -> {gate} and gen-capability-registry rejects
any capability declaring a step or contribution there. The verify lane is
therefore closed to capabilities that want to contribute to what UAT covers
rather than refuse to let it start.

Failing-first: the step, contribution, and exact-kind-set rows are RED; the
pre-existing gate row is a green regression pin so the new arms cannot orphan
the arm verify:pre already had.

Refs #3866

* feat(#3866): dispatch step and contribution hooks at verify:pre

verify_pre_hooks dispatched `kind == "gate"` only, so getWiredKinds reported
verify:pre -> {gate} and gen-capability-registry's validateHooksWired rejected
any capability declaring a step or contribution there. A capability could
refuse to let UAT start; it could not contribute to what UAT covers.

Add contribution and step arms mirroring execute:wave:post, deferring to
references/loop-hook-dispatch.md and carrying its ref.command in-context
validation guard ahead of any shell-use prose. A verify:pre step is advisory:
it never blocks the start of UAT and an erroring step is routed by its own
onError. The gate arm and its check guard are untouched.

Give extract_tests an additive consumption seam for the artefacts those steps
declare via the existing steps[].produces field -- no new registry field, no
new ordering, no invented filename. Manifest-supplied artefact names are
validated in-context against an allowlist and resolved only inside PHASE_DIR.
With no producing step the derivation is unchanged, pinned by test rather than
asserted in prose.

Review findings folded in: the artefact-name allowlist (isolated adversarial
pass), the artefact-shape contract and the seam-inertness tests (spec axis),
and the reference/how-to split so one constraint has one source of truth
(standards axis).

Closes #3866

* chore(#3866): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-25 18:07:29 -04:00
Tom Boucher
fb2d122d7f feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb

only this package publishes. The path-based branches — a project-local install,
a runtime config directory — had no such guarantee; they trusted their
configured location. This closes them.

Mechanism: once resolution finishes, and before any verb runs, the preamble
probes the tool it picked with `runtime-identity --raw` and matches the answer
with a shell `case` pattern ANCHORED to the start of the compact payload
(`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts
the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which
any colliding package could publish. The outcome is exported as the two-valued
`GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE
rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling:
`unverified` prints one line naming BOTH causes and continues, because
`no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core`
older than the verb, and at rollout the old-version case is the common one.

The blocker was byte budget, not design. The preamble is inlined into 112
shipped files and several sat within single-digit bytes of frozen ceilings
(`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a
first attempt broke five of them. What made room was collapsing the resolver's
twenty near-identical `elif [ -f … ]` arms into one candidate-list helper
(`_gsd_at`), which buys far more than the assertion costs. The preamble is now
2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every
capped file moved away from its ceiling rather than toward it. No cap raised, no
size-budget exception added, no override token emitted.

Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source
fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all
preserved byte-for-byte in substring terms; the snippet still begins with
`_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor
on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal.

Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both
described an `[ -x ]` guard as the load-bearing re-source defense. That guard
was tried and REMOVED in #3831 — it rejected the bare function name, fell
through every branch, and hit `exit 1`, which kills a sourced caller's shell.
`unset -f gsd_run` is the actual mechanism.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3841): pair the anchor's brace by requiring a closed identity payload

The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md
has unbalanced braces: net depth 2" — plus a knock-on report from its parent
`bug #1516` describe, which is the same failure counted once at the child and
once at the block.

Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and
increments on `{`, decrements on `}`, with no awareness of shell quoting. It
scans `new-project.md` PLUS every `new-project/steps/*.md`, and both
`new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy
— hence net 2 from a snippet that was off by exactly one. The unpaired brace was
the `{` inside the single-quoted `case` pattern of the identity anchor, which is
correct shell and invisible to a text scanner.

Fix in the snippet, not the guard. The pattern now anchors at BOTH ends:
`'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that
does real work rather than a cosmetic pair — a truncated payload whose prefix
matches now fails too, where before it verified. Safe for any future additive
field: a JSON object's own closing brace is always the last character, whatever
type the last value has, which is pinned by two negative-space tests (a nested
object and an array-valued last key must both still verify). Cost: +3 bytes,
against the 1,873 the resolver fold already gave back.

The alternative considered and rejected was dropping the literal `{` for a `?`
glob. It balances too, but weakens the anchor from "must be an opening brace" to
"must be any one character", and the anchor is the entire point.

Two guards added so this cannot recur silently:
- runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next
  edit to that pattern fails on the file it broke instead of surfacing three
  files downstream in a test whose name mentions neither the launcher nor this
  issue. It also asserts depth never goes negative, since a `}` preceding its
  `{` nets to zero while being unbalanced at every prefix.
- runtime-identity gains behavioral truncated-payload and trailing-garbage
  fixtures, so the added `}` is proven load-bearing rather than merely present.

Verified: snippet 51/51 braces; new-project combined net depth 0; the seven
other preamble-bearing files with nonzero depth are unchanged from merged next
(their own prose, not the preamble, and not in any guard's scan set); all 112
inlined copies and the resolver reference re-synced byte-equal; sync:launcher
idempotent on the second run.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3841): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 01:05:53 -04:00
Tom Boucher
36375513b9 feat(#3840): generate docs/FEATURES.md from per-feature fragments (#3845)
* feat(#3840): generate docs/FEATURES.md from per-feature fragments

docs/FEATURES.md was hand-maintained, and every feature PR wrote into two
shared mutable cells: the '### N.' heading whose integer was hand-allocated at
authoring time, and the hand-maintained table of contents. Concurrent PRs all
picked the same next integer, and two PRs adding differently numbered features
still collided on the TOC. #3831 was renumbered 165 -> 166 -> 167 -> 168 across
successive rebases, each collision also costing a full matrix verification run
because the sha-keyed pass marker dies with the rebase.

Mechanism: one fragment per feature at docs/features/<slug>.md carrying
id/title/group (and an optional order) in frontmatter, consolidated by
scripts/gen-features.cjs --write|--check into a marker-delimited region of
docs/FEATURES.md that holds BOTH the TOC and every section body. Group headings
and their order are derived too - a group sorts by its lowest-ordered member -
so there is no shared registry to edit either; optional per-group prose lives in
docs/features/_groups/<slug>.md. A contributor adds exactly one new file.
Wired into regen:derived and lint:generated-sync alongside the eight existing
generators, matching gen-adr-index.cjs's CLI shape and typed-REASON reporting.

Migration froze all 168 existing numbers verbatim: identical section set,
identical order, identical bodies. Two defects found in the tree are fixed
inline rather than carried forward - the '## Related' block had been spliced
into the middle of the document, orphaning §142's Reference line, and four
inbound anchors were already broken on next (FEATURES.md#runtime-identity in
two files, and #143-spec-phase-edge-completeness-probe off by one). Since the
repo has no link checker, --check now validates every inbound
FEATURES.md#anchor by resolved target, so that class cannot ship silently
again; locale FEATURES.md files resolve elsewhere and stay out of scope.

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3840): carry upstream §69 delta into its fragment and harden the generator

Review found section 69 missing '[--strict]' and REQ-STATE-05/06 versus
origin/next. Root cause was a stale base, not extraction loss: those lines
landed in 394bf384b (#3844) AFTER this branch forked at 63abcface, and
'git diff 63abcface origin/next -- docs/FEATURES.md' is exactly that hunk.
Merging origin/next auto-applied the hunk into the GENERATED region, which
--check immediately reported as stale; the delta is now carried in
docs/features/statemd-consistency-gates.md and regenerated from there.

--write is now fail-closed. It previously rendered the region even with
violations outstanding, warning only on stderr and exiting 0, so a
'--write && git commit' chain could commit a FEATURES.md carrying two
colliding sections. It now refuses and exits 1; --force is the explicit
override and says so in the report. The test that pinned the old behavior now
pins the refusal, plus the --force override and its scoping.

Marker forgery is rejected at two layers. A fragment body containing
'<!-- FEATURES:START' or '<!-- FEATURES:END' is a typed
body_forges_region_marker violation (fragments and group notes alike), and
spliceIntoFeatures anchors the end boundary with lastIndexOf instead of
indexOf, so a marker that reaches the document by any other route can only
make the generated region grow, never shrink. Matching is on marker PREFIXES,
so a decorated variant comment cannot slip past.

Symlinked corpus entries are refused with a typed dirent_not_regular_file
rather than read. A fork PR could otherwise commit docs/features/evil.md as a
symlink to any readable path and have the generator inline those bytes into
the committed docs/FEATURES.md on the next regen.

Equivalence re-verified with a method that cannot cancel out. The first
check extracted both operands with the same body-normalising helper, so
anything that helper dropped was dropped on both sides. The replacement runs
two independent passes: a global content-line multiset diff with no
per-section logic at all (0 gained, 19 lost, all 19 the stale hand-written
mini-TOC links this change deliberately deletes), and a per-section
byte-exact body diff carrying a coverage assertion that fails loudly per file
when the extractor accounts for fewer lines than the file contains. That
assertion caught two blind spots in the checker itself. 168/168 sections
present, order identical, one intended body difference (§142 regains the
Reference line orphaned by the misplaced '## Related' block).

Refs #3840

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3840): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 22:49:00 -04:00
Tom Boucher
394bf384be fix(#3696): report the last_activity invariant and make the verdict gateable with --strict (#3844)
* test(#3696): failing-first coverage for the last_activity invariant and --strict exit status

* fix(#3696): report the last_activity invariant and make the verdict gateable with --strict

* fix(#3696): agree with the real reader on last_activity, and stop reporting structure as truncation

* chore(#3696): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-24 21:33:48 -04:00
Tom Boucher
63abcface9 feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools

The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git.

The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap.

unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell.

Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true.

An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes.

Closes #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3146): stop sync:launcher relocating a deliberate preamble placement

Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins.

Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture.

Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3146): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3146): document the FEATURES.md section-numbering practice

The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases.

Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set.

Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914).

Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:57:16 -04:00
Tom Boucher
c933184b97 enhance(#3172): require a stated failing direction for every automated acceptance command (#3825)
* test(#3172): failing-first suite for the stated failing-direction probe

Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel
exemption, degraded-read contract, CLI arm and the plan-authoring contract text.
RED by construction: the module exports it requires do not exist yet.
Executed on the remote runner.

* feat(#3172): require a stated failing direction for every automated acceptance command

Every runnable <automated> command now carries a <fails_when> sibling naming
what output constitutes failure. A command with no expressible failure mode is
not an acceptance test: it reads as rigour and is not falsifiable.

- verify-command-grounding gains a failing-direction probe sharing the existing
  <automated> grammar, MISSING sentinel and walk guard rather than copying them
- gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches
  it and hands the JSON to gsd-plan-checker check 8f
- Dimension 8 detail extracted to references to stay under the agent size cap

Verified on the remote runner.

* fix(#3172): close four review findings in the failing-direction probe

- MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so
  a real command was exempted from the new blocking gate. Tightened the SHARED
  constant rather than adding a second copy.
- Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k).
  Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The
  pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too.
- probePhaseFailingDirections reported status 'ok' when one plan was unreadable,
  conflating 'could not look' with 'nothing to report'.
- Extracted the phase-resolution block both check arms had copied verbatim.

Also corrects a docs/AGENTS.md dimension list stale since #2401.
Verified on the remote runner.

* fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping

The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen
under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars
of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled
where such a rule goes: the planner spawn contract in plan-phase.md, beside
<tracked_source_paths>. The agent file is reverted to origin/next verbatim.

- plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the
  contract there and row 30b guards the freeze in both directions
- plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the
  precedent that two ack sources may never name the same path
- install-tree fixtures regenerated for the three new reference files

Verified on the remote runner.

* chore(#3172): backfill PR number into the changeset fragment

pr:0 -> pr:3825 now that the PR exists.

---------

Co-authored-by: sim <sim@local>
2026-08-24 19:05:11 -04:00
Tom Boucher
596540f864 feat(#3227): publish machine-readable state contract at step boundaries (#3824)
* feat(#3227): publish machine-readable state contract at step boundaries

Adds src/state-contract.cts, a best-effort publisher that writes
.planning/state.json (contract 1.0.0) at 11 step-boundary commands, so
external tools read a versioned contract instead of parsing STATE.md and
ROADMAP.md heuristically.

Composes existing owners rather than re-deriving: phase rows come from a
new locateProgressTable extracted from deriveProgressFromRoadmap (so the
snapshot can never disagree with GSD's own progress counters), milestone
identity from getMilestoneInfo, and next from classifyProject. Owners are
required lazily to avoid the state -> state-contract -> smart-entry ->
state require cycle.

Also fixes a pre-existing defect in scripts/lint-test-file-count.cjs
(maintainer-approved as a second concern): testEffectivePrefix never
stripped the suite qualifier, so 65 dotted test files counted against no
module and 9 mis-bucketed into a shorter one. Allowlist re-baselined for
the 74 files the gate can now see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3227): backfill PR number into the changeset fragment

pr:0 -> pr:3824 now that the PR exists. Doc-only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3227): shape hostile-name fixtures away from the scan corpus

The two hostile-input fixtures used a literal phrase from
scripts/prompt-injection-scan.sh's corpus, so CI's Security Scan redded on
this file. These tests assert that an arbitrary phase name round-trips into
state.json as inert data -- the property holds for any string, so the
injection flavor is illustrative, not load-bearing.

Reshaped to a hyphenated fake instruction tag, which stays hostile-looking
while matching none of the scanner's patterns. Allowlisting the file was
rejected: that mechanism is for suites whose subject IS injection defense,
and it would blind the scanner to this whole file permanently.
See DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3227): ratchet the state-contract mutation floor to its measured score

The module was registered at minScore 50, the ratchet's minimum permitted
floor for a newly-registered module whose score had not been measured. This
PR's own Stryker shard measured 66.25% (run 32769289750, job 97565813640),
so the floor moves to floor(measured) - 1 = 65, per the rule the registry
documents.

66.25 is below TARGET_MUTATION_SCORE (80), so this stays a ratchet
candidate: raise as the tests improve, never lower.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 17:56:02 -04:00
Tom Boucher
a2387a0545 feat(#3034): add opt-in parallel reviewer lanes (#3822)
* test(#3034): failing-first coverage for opt-in parallel reviewer lanes

Executes the real invoke_reviewers dispatch block from review.md against a
stubbed gsd_run seam rather than pattern-matching the workflow text, so the
two properties that actually carry risk are observable: that every lane is
joined before aggregation, and that concurrent lanes cannot tear a line in
gsd-review-lane-results.jsonl.

Concurrency is proven by a barrier fixture, not by elapsed time -- each stub
lane blocks until all lanes have checked in, which can only complete if they
overlap.

Red against the current sequential dispatch, by design.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#3034): add opt-in parallel reviewer lanes

Reviewer lanes within one review pass inspect the same immutable plan
snapshot and have no dependency on one another, but were dispatched strictly
one at a time, so a multi-reviewer pass cost roughly the sum of its lanes.
The serialization is a deliberate protection against provider rate limits,
so it stays the default; review.parallel_lanes opts a project out of it.

The loop body is hoisted into run_review_lane so the sequential and
concurrent paths share one body -- two hand-synced dispatch bodies is the
divergence class ADR-2782 spent a phase deleting. Each lane writes a
slug-scoped result file, concatenated in selection order after the join:
concurrent O_APPEND is atomic only below PIPE_BUF, and write_reviews parses
that JSONL to render the models:/model_sources: frontmatter, so a torn line
is a broken REVIEWS.md rather than a cosmetic log defect. Aggregating in
selection order also keeps the artifact byte-identical between the two paths.

The guard is strict equality on "true" and falls back to sequential when
config-get fails -- the opposite polarity from the commit_docs guard,
because failing open here fires the very requests the default prevents.

Also corrects docs/COMMANDS.md and its four locale mirrors, which described
--all as running every configured reviewer in parallel when dispatch was in
fact sequential.

Closes #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3034): de-duplicate dispatch slugs and scope lane locals

Review finding (Standards axis): a slug repeated in SELECTED_REVIEWERS would
put two concurrent background jobs on the same > -truncated per-lane result
file. The shared-append form this replaced could not corrupt itself that way,
so de-duplicating is what keeps the concurrent path no worse than the
sequential one.

Selection de-dupes today -- the roster is a Set and review.default_reviewers
normalizes lowercase-unique -- but reachability analysis is not a contract,
which is the same reason the roster derivation itself is guarded.

Splitting once into DISPATCH_SLUGS also removes the duplicated tr-split the
same review flagged: the dispatch and aggregation loops now share one list,
which is what guarantees they walk the same slugs in the same order. A plain
string accumulator rather than an array, because zsh and bash disagree on
array indexing and this block runs under both.

Also scopes run_review_lane's locals. Not a live fix -- each dispatched call
already forks its own subshell -- but it makes the isolation a property of the
function rather than of the dispatch mechanism happening to fork.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3034): acknowledge review.md growth, drop spent 2295 ack

The differential attribution gate reported review.md growing 4173 bytes
(30712 -> 34885) with no live acknowledgment. Adds the per-PR fragment it
asks for, naming only the one path it reported.

Deleting tests/emitted-drift-acks/2295-resolved-model.json is required, not
opportunistic. That fragment declared review.md and nothing else, and its
ripple is already absorbed into the base, so it is spent -- it can no longer
clear anything, which is why the gate still reported review.md as
unacknowledged. It could not simply be left alone either: two ack sources may
never name the same path, so it blocked this PR's fragment outright.
CONTRIBUTING is explicit that a fragment whose last entry is removed gets
deleted with it, because an empty fragment signals nothing while its presence
reads as a live alarm.

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3034): backfill changeset PR number

Refs #3034

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 15:25:22 -04:00
Tom Boucher
4af59f8dd3 fix(#3662): resolve managed hook node runners at hook-fire time (#3790)
* test(#3662): failing-first suite for runtime-resolving hook runners

* fix(#3662): resolve managed hook node runners at hook-fire time

* fix(#3662): close review findings and document the resolver

* fix(#3662): close adversarial and security review findings

* chore(#3662): backfill changeset pr number

* test(#3662): honor win32 skip return and platform-aware sh runner pin

* test(#3662): pin the bare win32-claude sh-hook shape omitting the bash runner

---------

Co-authored-by: sim <sim@local>
2026-08-24 00:07:06 -04:00
Tom Boucher
cf15682d1c enhance(#3028): responsive Markdown separators instead of fixed-width rules (#3789)
* feat(#3028): responsive Markdown separators instead of fixed-width rules

Stage banners, checkpoints, completion and error panels used fixed-width
runs of box-drawing characters -- a 53-column heavy rule and a 62-column
double-line box. Those runs are ordinary text to a Markdown-rendering
host, so in a narrower pane they wrap and the border comes apart from
the heading it framed.

Shipped content now emits an ATX heading for a titled section and a
blank-line-delimited --- for a break between sections, both of which
adapt to the available width. The same convention is applied to the
three code sites that built these strings at runtime: the UAT
checkpoint renderer, the milestone-close audit report, and the TDD
review checkpoint table.

Removing the box also removes its only reason to exist -- the
east-asian-width padding helpers that kept its right border aligned
(checkpointBoxLine, displayWidth, isWideCodePoint, ZERO_WIDTH_MARK_RE,
CHECKPOINT_BOX_WIDTH). RTL directional isolation is unchanged.

The convention is specified in gsd-core/references/ui-brand.md and
enforced across all shipped content by tests/responsive-separators.test.cjs.

Refs #3028

* test(#3028): pin the heading form in checkpoint and audit-report assertions

These suites asserted the exact box borders and the 62-column padded
banner interior. With the box gone they assert the ### heading form,
the --- break and the bolded instruction line, and each now carries a
positive assertion that no box character remains -- which is what pins
the fix rather than merely tolerating it.

Language coverage is converted, not dropped: Japanese, Chinese, Korean,
Hindi and Arabic all still assert their rendered banner, and the Arabic
case still asserts the RTL directional isolates the box removal must
not disturb. Adds a case for a banner longer than the old inner width,
which previously produced a ragged border and now has none.

Refs #3028

* chore(#3028): acknowledge execute-plan.md growth from the checkpoint display spec

The checkpoint_protocol display spec described the drawn box; it now
describes the heading, the --- break and the bolded action prompt,
which costs 22 bytes (40111 -> 40133, 827 under the cap).

Appended to the existing #3370 fragment rather than filed as a new one:
a growth ack keys on the bare filename and #3370 already declares
execute-plan.md, so a second source naming it would be a hard
duplicate-key error. Same supersede-by-append route #3370 took for the
spent #2652 fragment.

Refs #3028

* docs(#3028): state the load-bearing half of the separator rule, and amend the zh-CN reference

Review found three things.

The rule as first written demanded a blank line above AND below every
---. Only the one above is load-bearing: it is what stops CommonMark
reading the rule as a setext underline for the line above. The one below
is cosmetic, because a thematic break is a leaf block. The rule now says
that, with the reason, instead of asserting a stricter form the content
does not keep.

The zh-CN reference had received the mechanical box-to-heading swap but
none of the prose behind it: it still claimed a 62-character checkpoint
width and still listed --- among forbidden mixed banner styles, so it
contradicted the convention it was translating. It now carries the
separator section, the setext reasoning, the unconditional-vs-per-runtime
rationale and a corrected anti-pattern list, in Chinese.

The user guide asserted that a heading is not a degradation anywhere.
That is an assertion, not a demonstration. It now says what was actually
traded away in a plain terminal, points at the recorded rationale, and
invites the report that would justify the capability flag instead.

Refs #3028

* chore(#3028): backfill changeset PR number

Refs #3028

---------

Co-authored-by: sim <sim@local>
2026-08-23 22:38:12 -04:00
Tom Boucher
004e9dd741 fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability

RED by construction. Binds to behavior renderEffortForRuntime does not yet
have: an optional third `model` argument, a per-model advertised-level table,
`max` passing through instead of clamping to `xhigh`, `minimal` clamping to
`low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/
`reason`) so a downgrade is legible from resolver output rather than silent.

Two of these pin defects that exist on next today:

- `max` is discarded. Both Codex models whose catalog entries are retrievable
  (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing.
- `minimal` is emitted to a model that refuses it. providerPresets.openai.
  haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's
  advertised floor is `low`. GSD is sending a value into a document Codex
  itself validates. The parity test is what pins that fixed, and it names the
  offending path/model/effort when it trips.

Also corrects tests/model-resolver.test.cjs:351, which asserted
renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as
though it were a contract. ADR-443 recorded "Codex has no max" as fact and it
was true when written; Codex has since added both `max` and `ultra`. That is a
stale premise, so the assertion is corrected here rather than worked around.

The property test asserts the invariant the whole change exists for: a rendered
effort is always a level the target model actually advertises, or an explicit
rejection. There is no third outcome.

* fix(#3007): resolve Codex effort per model, and make every clamp visible

Codex declares supported_reasoning_levels per MODEL and validates against it,
so a single per-runtime capability set cannot be right for all of them. GSD's
was wrong in both directions at once.

`max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped
max -> xhigh on that basis; it was accurate when written, and Codex has since
added both `max` and `ultra`. Every Codex model whose catalog entry is
retrievable advertises `max`, so the clamp was discarding a level the provider
supports, silently, on the most-used path.

`minimal` stops reaching Codex. No Codex model advertises it -- both retrievable
entries floor at `low` -- yet providerPresets.openai.haiku.low paired
gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the
receiver validates and refuses into a file the receiver reads. Being
unconservative in what you send is the half of Postel's rule with no defensible
reading, so that preset is corrected and a parity test pins it.

`ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum
reasoning with automatic task delegation": at ultra, effective_multi_agent_mode
returns Proactive and Codex spawns sub-agents on its own initiative, underneath
GSD's orchestration rather than inside it (#2167). It is a mode switch, not a
reasoning depth, so it is not added to the universal ladder -- which stays
provider-agnostic by ADR-443's design -- and it is rejected even for
gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered
and rejected: that silently discards what the user actually asked for.

Clamping is now visible. RenderedEffort carries requested/clamped/reason and
resolve-execution surfaces them. The previous table clamped correctly but
invisibly, so a user asking for `max` on Codex had no way to find out they were
getting `xhigh` -- exactly the failure mode the robustness principle's modern
critique warns about, and why "be liberal" has to mean "liberal and loud".

Also closes a latent trap found while reviewing the implementation: the clamp-up
loop walks the ladder upward, and for a future model advertising `ultra` but not
`max` it would have selected `ultra` as the clamp target -- re-entering by the
back door the mode the rejection above exists to keep out. A clamp may never
produce a value that a direct request for that value would refuse. Unreachable
with today's catalog, which is why no test caught it; a test now asserts the
invariant directly.

Signature stability is preserved: the third `model` argument is optional and the
two-argument form still resolves, against the family baseline. That form's
BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old
answer would have fixed the defect only where a model happened to be threaded
through and left it live everywhere else.

tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract
and is corrected here rather than worked around.

* fix(#3007): close every review finding on the Codex effort alignment

Two isolated reviewers, correctness and security. Both found the same two
blockers, and the per-model work was inert on every surface that matters until
this commit.

BLOCKER — resolve-execution never passed the model and discarded the clamp.
cmdResolveExecution called the two-argument form and emitted only
effort_rendered/effort_param/effort_propagation, so the per-model table was
unreachable from production code (tests were its only caller) and requested/
clamped/reason were computed and thrown away. Requested outcome 3 names "the
effective rendered effort in resolver output" specifically, so the feature was
unmet on the exact surface the issue asks for. Now passes the resolved model and
emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the
existing key convention rather than introducing a nested object.

BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed
a nested {"effort": ...} sample; the real result is flat and those keys were
absent entirely. A reference doc asserting a JSON path a reader can copy is worse
than no doc. Corrected against the actual emitted key set.

MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex
kept minimal in its supported set and still clamped max down to xhigh, so the
invocation-time and install-time channels disagreed about the same runtime's
capability: --host codex with max emitted xhigh while the generated TOML said
max. This is the repo's documented generative-fix-divergence class, so both
tables now cross-reference each other and a parity test fails if they ever
diverge again.

MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null
_baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback
never fired and every effort rendered as null. And a non-array value made the Set
constructor throw at module load — model-catalog.cjs is required across the whole
CLI, so one bad JSON value killed every command, not just codex effort. Guarded
on size and filtered to array values; both degrade to the hardcoded baseline.

MAJOR — value widened to a nullable string with two consumers left behind.
runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a
null effort key in generated frontmatter); install-effort-resolver still declared
a non-nullable return, a structural lie that silently defeated null checking.
Both corrected, both omitting the key on null — the same posture as 'inherit',
where omission means "follow the host default".

MAJOR — the per-model table is inert today, and the docs now say so. All three
shipped models advertise the same usable range and ultra (sol's only
differentiator) is rejected for every model, so no observable output differs by
model. The table stays because Codex declares capability per model and the sets
are free to diverge — a single per-runtime assumption is precisely what went
stale and produced this issue — but overselling it as a visible per-model feature
would have been the same class of error as the doc blocker above.

Tests: three passed under a full revert and are strengthened rather than deleted,
since each guards a real contract (#3533's inherit rule, the undeclared-host
rule, off-ladder handling) — they now also assert the clamp-visibility fields,
which only exist after this change. The fast-check property is kept for its
shrinking, and a deterministic nested loop over the full cross-product now sits
beside it so coverage is exhaustive rather than sampled.

Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg
form and would have written a literal null reasoning effort on the ultra path;
CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT.
The installer defect was found by the co-change gate, not by a reviewer —
install.js is a historical co-change partner of model-catalog.cts that this diff
had not touched.

* test(#3007): correct assertions that pinned Codex's stale effort premise

Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the
shipped commit. Every one is a stale pin, not a defect: each was probed against
the built module before its expectation was changed, and none failed for a
reason other than this premise correction.

Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale
by a production change must not ride inside another commit, because the
release-sdk hotfix cherry-pick filter routes by subject prefix and a correction
buried under the wrong prefix ships a half-state (v1.42.3, #3621).

The most valuable one was tests/model-resolver.test.cjs's cross-provider
validity invariant, which hardcoded the Codex enum as
`minimal|low|medium|high|xhigh` and failed with "real API would 400". That
message is now false in both directions: Codex accepts `max`, and rejects
`minimal`, which no model advertises. The enum is corrected to
`low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the
"would the real API refuse this" check worth having, and it was right to fail
here. It simply carried the stale fact in its own fixture.

Test NAMES were corrected alongside their assertions wherever the name asserted
the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal
passthrough". A renamed test that still claims the old thing is worse than a
failing one, and a green test whose name states a falsehood is how the next
reader inherits the wrong premise.

Both channels are covered: install-time (renderEffortForRuntime, and the
generated .toml in install-runtime-artifacts) and invocation-time argv
(effort-surface-axis). They were deliberately brought into agreement in this
change, so their assertions had to move together.

Each site carries a #3007 comment recording that Codex gained max/ultra and that
capability is declared per model, so a future reader can tell this was a
deliberate premise correction rather than a test bent to fit an implementation.

* test(#3007): separate the effort-precedence case from the clamp case

The previous stale-assertion pass over-corrected one test. It saw
`effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`,
assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to
"max passes through" and changed the expectation to `max`. The remote runner
disagreed.

Reproduced against the real CLI: with that config and `gsd-planner`, the
resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`,
`effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE —
gsd-planner is heavy/opus tier and its routing-tier default outranks
`effort.default` — so `max` never reaches the renderer at all. The test says
nothing about clamping and never did; it only looked like a clamp pin because
both mechanisms happened to yield the same string.

Restored to `xhigh` and renamed to say what it actually tests. It now also
asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is
what makes it impossible to mistake for a clamp pin again: those two fields prove
the value is what the resolver produced rather than something the renderer
downgraded. Before #3007 there was no way to tell the two apart from the output —
which is precisely why the previous pass could not tell them apart either.

Added the test that was actually missing: `effort.agent_overrides`, which
outranks the tier default, so the requested level genuinely reaches the renderer
and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified
by probe before asserting.

One test now pins the precedence rule and the other pins the #3007 behavior, and
neither can be read as the other. That the clamp-visibility fields are what
resolved this is a small argument for having added them.

* chore(#3007): backfill changeset pr number to 3765

* test(#3007): put model-catalog under the mutation gate

The Stryker shard showed as `skipping` on this PR despite the diff rewriting
model-catalog's effort logic. That was legitimate, not a detection bug:
`model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the
whole module — including everything #3007 touches — sat entirely outside
mutation scoring with has_work "false".

Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs
is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no
temp dirs. That shape is not stylistic — it is the #2790 precedent this file
already documents. Stryker's command runner treats a whole `node --test <file>`
invocation as ONE test costing whatever its slowest case costs, and re-runs it
per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses
runGsdTools throughout) would reproduce exactly the 15-minute shard-cap
cancellation #2790 hit. The integration file is unaffected and keeps running in
full in the normal test job.

Coverage spans the module rather than only the diff, because the score is
measured over the whole file: effort rendering across every model and ladder
level in both channels, the prototype-chain host guard, the exported enums and
maps, isAnthropicFlavoredModel's provider namespacings, the profile projections,
nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are
worth naming — every uncovered exported function is score given away, and
mergeEffortTierDefaults turned out to have a genuinely interesting contract
(#3531: a partial override merges over the built-ins rather than replacing them,
and isValid gates the VALUE, not the tier name, so an unknown tier key is still
merged in). Every expectation was probed against the built module before being
asserted.

minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment.
Floors in this repo are measured, not chosen — the existing entries sit at 94, 75
and 56 — and they can only be measured in CI, because mutation shards run
`node --test`, which is hard-blocked locally. The first CI run on this branch
reports the real number and the floor gets ratcheted to it before merge. A
placeholder of 1 reaching `next` would make the gate decorative: it would pass
whether or not a single mutant is ever killed.

Note the target is "never regress from measured", not a fixed 80 — planning-inspect
sits at 56 and is documented as an accepted ratchet candidate.

* test(#3007): bootstrap model-catalog's mutation floor legally

The placeholder floor was structurally illegal and the remote run said so.
tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module
must carry a matching RATCHET_BASELINE entry in the same diff, minScore must
EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three.
That is the ratchet working exactly as intended — a floor nobody can satisfy
accidentally is the point of it.

Bootstrapped at 50 in both places. Fifty is not a measured score and the comment
says so plainly: it is the minimum the guard permits, and it coincides with
Stryker's own configured `break` threshold, so it is the lowest legal starting
point for a module that has never been measured. It still must be ratcheted to
floor(measured) - 1 before this PR merges.

Also corrected a real defect in the file's own instructions. "HOW TO UPDATE"
step 1 read "Run the per-module Stryker shard locally" — which cannot be done
here, and which the same file contradicts eighty lines further down, where the
#2790 scores are recorded as "not a local run; mutation shards run `node --test`,
hard-blocked in this repo's local environment". stryker.config.mjs confirms the
command runner invokes `node --test` once per mutant, and
.claude/hooks/block-local-node-test.sh denies exactly that. So the documented
first step sends the next contributor at a wall. Rewritten to describe the path
that works — push, read the measured score off the CI shard, then set the floor
and its baseline together in one diff — and to say why local measurement is not
available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched.

The two-step is inherent to the environment rather than a shortcut: a floor
cannot be measured before the first CI run exists, and the guard rightly refuses
to accept an unmeasured one below its minimum.

* test(#3007): ratchet model-catalog's mutation floor to its measured score

The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no
timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this
file's own rule, minScore = floor(measured) - 1, which is the same arithmetic
every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94.

Both halves moved together, because the ratchet guard asserts minScore equals its
RATCHET_BASELINE entry and would reject them drifting apart.

The spawn-free unit surface is vindicated by the clock: 57 seconds, against a
15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the
whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing
the shard at tests/model-resolver.test.cjs — #2790 recorded shards being
CANCELLED at that cap when they targeted a runGsdTools-heavy integration file.

The registry comment is rewritten rather than deleted. It previously warned that
the floor was provisional and must not ship that way; leaving that text next to a
measured floor would make the file lie in the other direction. It now records the
measurement the way the sibling entries do, including that 59.62 sits below
TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 —
comfortably clear of its own floor with real room to grow. Raise it as the tests
improve; never lower it.

Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is
an honest floor for a module that had NO mutation coverage at all an hour ago,
and it is now pinned so it cannot silently regress.

---------

Co-authored-by: sim <sim@local>
2026-08-22 20:51:55 -04:00
Tom Boucher
2f86278b5e fix(#3003): opt-in mechanism for intentional deletions in worktree.cleanup-wave (#3757)
* test(#3003): failing-first suite for declared deletions in cleanup-wave

Binds the guard's opt-in before it exists, so the suite is RED against next.

The rows that carry the weight are the over-authorization set: a directory
declaration must not authorize its children, a glob declaration must authorize
nothing, and a declaration must not act as a string prefix of another path.
Each of those BLOCKS, and each would PASS under a prefix, glob, or startsWith
matcher — which is how a path list quietly degrades into the boolean opt-in
#3003 explicitly rejected. The glob row matters most: declaredScopePrefix
already returns null ("matches everything") for a glob-leading pattern, correct
for the advisory it serves and catastrophic for a gate.

Also pinned: a failed deletion check blocks on its own reason rather than being
filtered into a pass; the block detail names only the undeclared residue so the
operator is not misdirected by paths that were fine; an entry with no
declaration blocks exactly as before; junk and non-array declarations do not
authorize; and a blocked entry still isolates rather than aborting the wave
(#2852, which must stay fixed).

Two advisory rows cover an interaction found while designing: git diff
--name-only includes deleted paths, so without unioning the declaration into
the #2596 scope check, authorizing a deletion would raise
SCOPE_OUT_OF_DECLARED against the very path just authorized.

A seeded property states the whole invariant the three over-authorization rows
sample: a deletion merges iff its normalized path is in the declared set.

* feat(#3003): declared deletions opt-in for the cleanup-wave guard

The deletions guard blocked the merge-back of any executor branch whose diff
removed a file, with no way to say a removal was intended. A plan that folded
one test file into a sibling could not be merged by the tool meant to merge it,
forcing a manual --no-ff outside the tool -- strictly less safe than what the
guard protects against.

A plan now declares removals in its own frontmatter (files_deleted), and that
list rides the same path files_modified already travels: plan-document parse ->
phase plan JSON -> the per-plan worktree gate -> record-agent/create
--deletions -> declared_deletions on the manifest entry -> the guard. The guard
blocks only the deletions NOT in that list.

A path list rather than a boolean, per the pinned decision: a boolean disarms
the guard for the whole entry, so an unexpected deletion riding along with a
declared one would pass unnoticed. Matching is exact after the module's shared
normalizer -- never a prefix, never a glob. Both would let one declaration
authorize a whole set, which is the mass-deletion accident the guard exists to
catch. That also means declaredScopePrefix is deliberately NOT reused here: it
returns null ("matches everything") for a glob-leading pattern, which is right
for the advisory it serves and would silently disarm a gate.

The block detail now carries only the undeclared residue, so an operator is not
sent looking at paths that were fine. A failed deletion check still blocks on
its own reason and is never filtered into a pass. A blocked entry still
isolates rather than aborting the wave (#2852).

The #2596 scope advisory unions the declaration into its declared set --
git diff --name-only includes deleted paths, so without that, authorizing a
deletion would immediately warn that the same path was out of declared scope.

Optional and additive throughout: files_deleted is absent from
PLAN_REQUIRED_FIELDS, a manifest entry without declared_deletions keeps the
original unconditional block, and omitting --deletions leaves the on-disk entry
shape untouched.

Supersedes the spent #2856 emitted-drift ack entry for execute-phase.md, the
same supersede that entry performed on #3370 and #3370 on #3324.

* fix(#3003): wire --deletions on every dispatch surface, not just one

Review found the feature inert on two of three dispatch paths. execute-phase.md
(harness inline) passed --deletions, but the orchestrator-worktree path
(executor-isolation-dispatch.md, worktree.create) and the Fleet-parallel batch
path (capabilities/claude-orchestration/fragments/execute-wave-pre.md,
worktree.record-agent) still passed only --files. A plan declaring
files_deleted would have merged on one path and been blocked on the other two
-- the exact bug #3003 exists to fix, left unfixed where most of the isolation
actually runs.

Worse, per-plan-worktree-gate.md already claimed --deletions was passed 'on the
same worktree.record-agent / worktree.create calls', which was false for both
untouched sites. A doc asserting coverage that does not exist is how a gap
survives review.

All four surfaces now pass the flag, verified by sweeping every .md under
gsd-core/, capabilities/, commands/, skills/ and agents/ that invokes
worktree.record-agent or worktree.create: each one that passes --files now also
passes --deletions. The isolation-dispatch note explains why this flag, unlike
--files, is not advisory -- omitting it does not skip a check, it blocks a
merge the plan declared.

Regenerates capability-registry.cjs, which the fragment edit made stale.

Neither newly-grown file needs an emitted-drift ack: executor-isolation-dispatch.md
sits under workflows/execute-phase/steps/ and execute-wave-pre.md under
capabilities/, both outside currentSizes()'s non-recursive scan of
gsd-core/workflows/ and agents/.

* docs(#3003): document files_deleted where a plan author will actually find it

The feature's entire user surface is one plan-frontmatter field, and the
canonical reference for that frontmatter -- docs/reference/plan-md.md, the table
that documents every other key -- never mentioned it. A field nobody can
discover ships as a field nobody uses. Adds the files_deleted row and an example
entry in all five locales (en, ja-JP, zh-CN, ko-KR, pt-BR), stating the property
that makes the opt-in safe: matching is exact per path after separator
normalization, with no globs and no directory prefixes, so a declaration can
never authorize more than it literally lists, and omitting the field keeps the
guard's original unconditional block.

Also corrects two claims in the scope-conformance how-to that this change made
false. Its opening paragraph described the recorded declared scope as
files_modified alone; declared_deletions is now unioned into that comparison.
Its "Renames are not detected specially" bullet asserted the deletions guard
blocks any entry whose diff contains a deletion, full stop -- which was the
whole point of #3003 and is no longer true. Reworked to say what now decides a
rename's fate: declare the old path in files_deleted and both halves become
ordinary paths for the advisory check, which is also why the old path needs no
separate files_modified entry.

Documentation that describes the pre-change behavior of the thing being changed
is worse than no documentation, because a reader trusts it.

* fix(#3003): close every review finding on the declared-deletions opt-in

Two independent isolated reviewers, correctness and security. Neither found a
blocker; both found real defects, and the directive treats a finding at any
severity as blocking. All of them are fixed here.

MAJOR -- the submodule worktree gate could not see a deletion-only plan.
per-plan-worktree-gate.md intersected $SUBMODULE_PATHS against $PLAN_FILES
alone, while $PLAN_DELETIONS was extracted and then never used. Before
files_deleted existed, a path had to appear in files_modified to be planned at
all, so the gate saw it; the new field plus the new docs telling authors a
deleted path needs no files_modified entry opened a hole where a plan whose only
submodule touch is a removal kept worktree isolation on -- the exact case #2772
disabled it for. Both channels now feed the intersection. Note the posture is
deliberately the OPPOSITE of the cleanup-wave guard: there the channels stay
apart because a deletion AUTHORIZATION must never be inferred; here they merge
because a safety fallback must never MISS a touch.

MAJOR -- same-wave conflict detection could not see a deletion. The planner's
implicit-dependency rule compared files_modified only, so plan A editing
src/x.ts and plan B declaring files_deleted: [src/x.ts] scored as conflict-free
and ran in parallel: one branch removing what the other is writing, which is the
sharpest conflict there is. Overlap is now computed across both channels.

MINOR (both reviewers, one root cause) -- the advisory union gave one field two
matching rules. declared_deletions was unioned into the scope list handed to
planWaveScopeConformance, which reads it with prefix-and-glob semantics. So a
field that is exact-match-only at the gate silently became wider at the
advisory: ["*.md"], inert at the gate, yielded a null prefix meaning "matches
everything" and muted the advisory completely, and ["src"] muted all of src/.
The union also activated the advisory on plans that declared no modification
scope at all, warning on every modified path. Replaced with subtraction from the
findings, gated on files_modified alone. One field, one rule, everywhere.

MINOR -- core.quotepath made the feature silently inert for non-ASCII paths.
git emits "tests/\303\251.ts" C-escaped and quoted, which never equals the
declared plain path, so a correctly declared deletion of tests/é.ts would block
forever with nothing pointing at the encoding. Both diffs now pass
-c core.quotepath=false.

NIT -- flag() consumed a following flag as a value, so --deletions --files x
swallowed --files and dropped both. Now treated as a missing declaration, which
fails closed. Fixed at both call sites; the helper is duplicated verbatim in
cmdWorktreeRecordAgent and cmdWorktreeCreate and leaving one would reintroduce it.

TEST -- one test passed for the wrong reason. "a declared deletion is in scope
for the advisory" asserted only that warnings omit the deleted path; under a
full revert the entry blocks first, warnings come back empty, and the negative
assertion passes anyway. It now asserts the entry actually merged, which is the
load-bearing half. Four regressions added, one per fix above.

Docs corrected rather than extended. The rename bullet in the scope-conformance
how-to claimed a rename whose delete side is undeclared never reaches the
advisory. Verified false: git's rename detection is on by default, so a pure
rename is a single R entry that appears in no --diff-filter=D output and was
never gated, before or after #3003. Only a rename that edits enough to fall
below the similarity threshold decomposes into add+delete. The pre-existing
sentence made the same wrong claim; this restates it correctly instead of
sharpening the error. The localized plan-md.md reference edits are reverted:
the PR template requires docs content added here to be English, and the
translations already lag by three fields, so English-only is the repo's
standing posture, not an oversight.

Agent-file size caps respected: gsd-planner.md is XL-tier by bytes but carries a
separate 49152-LF-CHAR cap asserted by four suites, so its edit is deliberately
terse and lands at 49141 with 11 chars of headroom, with the rationale moved to
docs/reference/plan-md.md, which has no cap. gsd-plan-checker.md lands at 49107
bytes, 45 under the LARGE cap. Both acks merged into the existing fragments that
already name those paths, since two ack sources may never name the same path.

* fix(#3003): decode git's path quoting instead of changing the git argv

The previous commit's non-ASCII fix turned the remote suite red: 44 failures,
42 of them "unexpected git call: -c core.quotepath=false diff --diff-filter=D
--name-only ...". The suite's git mocks match on exact argv, so adding two
flags to the deletions diff and the advisory diff invalidated every existing
fixture in tests/worktree-safety.test.cjs. Rewriting dozens of fixtures to
accommodate one flag would be paying a large Hyrum's-law bill to fix a small
defect.

Both execGit calls are reverted to their original argv. The C-quoting is now
decoded in normalizeScopePath instead, via a new decodeGitQuotedPath helper.
That is the better fix on its own merits, not merely the cheaper one: the git
argv is untouched so no fixture moves, the decode lands on the ONE normalizer
already applied to both sides of the comparison so the declared and reported
paths cannot disagree, and it holds regardless of the user's own core.quotepath
setting rather than only when we remember to override it.

A value not wrapped in a leading AND trailing quote is returned completely
untouched, so the plain-ASCII path -- the overwhelmingly common case -- is
byte-identical to before. Escapes decode to BYTES collected into a Buffer and
UTF-8 decoded only at the end, because \303\251 is two bytes forming one
character and decoding them separately yields mojibake. Malformed input never
throws: a trailing lone backslash or a short octal escape degrades to the
literal character, since one bad path must not take down a cleanup wave.

Caught while reviewing the helper: the non-escape branch pushed a UTF-16 code
unit rather than UTF-8 bytes. Git always escapes non-ASCII so its own output was
fine, but this normalizer runs on the DECLARED side too, and an author may write
a quoted path holding a literal é -- pushing 0xE9 alone is invalid UTF-8, so the
declaration would decode to a replacement character and silently stop matching.
That is precisely the failure this change removes, reintroduced on the other
side of the comparison. Now converts whole code points, surrogate pairs intact.

The other 2 failures: tests/parallel-dependent-plans.test.cjs pins the exact
unbackticked substring "files_modified overlap" in gsd-planner.md, and rewording
that comment to "declared-scope overlap" deleted it. The comment is restored
verbatim and the files_deleted change rides in the pseudocode and the Rule
sentence instead. Recorded in the ack fragment so the next contributor does not
rediscover it the same way.

Four regression tests cover the decode through the public cleanup-wave seam
(the helper is module-private): a declared non-ASCII deletion merges against a
C-quoted git report, the symmetric case where the DECLARATION is the quoted
form, an undeclared non-ASCII deletion still blocks with the residue naming the
decoded path an operator can act on, and a path merely containing a quote is
left alone. Plain ASCII was already covered and is not duplicated.

* fix(#3003): revert the leading-dash flag guard, the review nit was wrong

The remote suite came back with 2 failures, down from 44, and both point at the
same thing: tests/worktree-safety.test.cjs:7045 already pins the opposite
contract, deliberately.

  test('a flag-shaped --files value is not re-parsed as a flag', ...)
    recordAgent(['--files', '--branch'])
    -> files_modified === ['--branch']
    -> branch === 'worktree-agent-a1'  ("the real --branch value must be untouched")

So consuming the next argv element positionally, whatever its shape, is the
tested intent of this parser, not an oversight. The security reviewer's nit
claimed --deletions --files x would "swallow --files and drop both". It does
not: each flag runs its own indexOf, so --deletions records the literal
'--files' while --files independently still resolves to x. And that literal is
a path git never reports as deleted, so it authorizes nothing -- already
fail-closed with no guard at all. The guard bought no safety and silently
changed --files behavior along the way, outside this issue's scope.

Reverted at both call sites, which are byte-identical again, along with the test
asserting the reverted behavior and the docs sentence describing it. The nit is
recorded as REJECTED in the review artifact with the reasoning above, rather
than as fixed -- a finding that turns out to be wrong should leave a trace of
why, or the next reviewer files it again.

docs/CLI-TOOLS.md now states the positional-read behavior plainly instead, so
the next person meets it as documented intent rather than rediscovering it
through a red suite.

* chore(#3003): backfill changeset pr number to 3757

* test(#3003): cover parsePlanDocument's filesDeleted branch to clear the mutation gate

CI's Stryker shard for plan-document failed at 73.28 against a break threshold
of 75: 170 killed, 62 survived, 232 total. Eight of those survivors are the
filesDeleted block this issue added to parsePlanDocument, which shipped with no
direct coverage at all -- the field was exercised end to end through the
cleanup-wave tests, but the parser itself was never called with a plan that
declares it, so every mutant in the block lived.

Four tests, each pinned to specific mutants rather than written for coverage
percentage:

- absent key yields exactly [] -- kills the array-literal seed
  (["Stryker was here"]) and the `fmDeleted = true` conditional, which would
  otherwise produce ["true"]
- a scalar underscore `files_deleted:` wraps into a one-element array -- kills
  `fmDeleted = false`, the `&&` logical-operator swap, the `fm[""]` string
  mutation on the first operand, the emptied if-block, and the ternary's
  non-array branch
- an array-valued hyphenated `files-deleted:` maps element-wise -- kills the
  `fm[""]` mutation on the SECOND operand (only reachable when the legacy
  hyphen alias is the one carrying the value) and the ternary's array branch
- an empty list yields [] -- boundary case, and a genuinely distinct one from
  the absent key: [] is truthy in JS so it ENTERS the if, and only
  Array.isArray's true branch mapping over nothing produces the same []

Threshold arithmetic: 174 of 232 are needed for 75%, and these take it to about
178, so the shard clears with margin rather than landing on the line.

Every expected value was confirmed by executing the built parser before being
asserted, not inferred from reading the source.

---------

Co-authored-by: sim <sim@local>
2026-08-22 13:17:51 -04:00
Tom Boucher
dacae92730 docs(#2845): record the inventory-provenance limits where readers meet them (#3746)
The limits shipped with #2845 were disclosed only in the PR body, which is
read once at merge and then buried. They are properties of what the feature
does, so they belong in the documentation.

Three surfaces, each at the point a reader forms an expectation:
docs/how-to/design-a-ui-phase.md gains a 'What this check is and is not'
subsection under the provenance how-to; docs/explanation/security-model.md
gains a residual-risk pair matching the section's existing shape; and
docs/AGENTS.md notes them where gsd-ui-checker's behavior is described.

The substance: a provenance line makes an inventory's origin falsifiable
rather than verified, since nothing re-runs the command or compares the
count; the rule is agent-applied like the other six dimensions, not a schema
check; and 'the checker never runs the recorded command' is an instruction
rather than a capability boundary, because the checker holds a Bash grant it
genuinely needs for the agent-skills bootstrap and tool grants here are not
command-scoped.

Co-authored-by: sim <sim@local>
2026-08-21 13:22:50 -04:00
Tom Boucher
4918c62d76 feat(#2845): require provenance for UI-SPEC component inventories (#3745)
* test(#2845): failing-first suite for UI-SPEC inventory provenance

Binds two shared formats before either exists, so the suite is RED against
next: the gsd-ui-checker dimension roster (asserted independently on twelve
surfaces, eight English and four translated) and the provenance-line grammar
the UI-SPEC template emits and Dimension 7 consumes.

Every parity assertion is paired with a synthetic mutation case, so the guard's
failure branch executes rather than only reading a correct tree: limit-1 (a
surface still declaring 6), limit (7), limit+1 (8), a dropped dimension, a
label that drifts on one surface only, a non-contiguous roster, a duplicated
number, and a surface that stops declaring a count at all. A seeded fast-check
property renders the roster under formatting noise (CRLF, padding, interleaved
sections) and asserts the parse round-trips and is strictly sensitive to a
dropped heading.

Assertions are on parsed typed records, never raw substrings.

* docs: normalize design-a-ui-phase how-to to American English

House style for docs/ is American English (CLAUDE.md). This file carried
colour/initialisation/initialise/artefact throughout. Spelling only — no
content change; kept separate from the #2845 feature commit so the
release-notes classifier and the hotfix cherry-pick filter see it for what
it is.

* feat(#2845): require provenance for UI-SPEC component inventories

A UI-SPEC's component inventory was treated downstream as a closed allowlist
while the document recorded nothing about whether the list had been enumerated
from the installed design system or recalled from memory. A recalled inventory
is indistinguishable from an enumerated one, so an executor complying with the
spec builds against a fraction of what the package offers, and every gate stays
green because they assert semantics rather than composition.

The UI-SPEC template gains a Component Inventory slot carrying one of two
provenance lines: the command that enumerated the list, the count it returned,
the resolved package@version and the date; or a Could not enumerate record with
a real reason. gsd-ui-researcher gains an enumeration ladder and must record
the line rather than write the list from recall.

gsd-ui-checker gains Dimension 7. An inventory with no provenance line, a count
with no command, an empty could-not-enumerate reason, or a line still carrying
the template's unfilled placeholders BLOCKs; a partial line, a line placed below
its table, or an honest negative record FLAGs; a complete line passes, and so
does a spec carrying no inventory at all, which keeps every UI-SPEC predating
the dimension validating unchanged. Whatever the verdict, an unsourced inventory
is reported as a non-exhaustive list of known-good components rather than a
closed allowlist, so the executor is never blocked from a component the spec
merely failed to mention. The checker never runs the recorded command.

The dimension count moved on all thirteen surfaces that assert it, across five
languages. Also corrects the claim in the English, Korean and Portuguese how-tos
that this checker applies a scored six-pillar rubric — that rubric belongs to
/gsd-ui-review's retroactive audit.

* chore(#2845): backfill changeset pr number to 3745

---------

Co-authored-by: sim <sim@local>
2026-08-21 11:59:56 -04:00
Tom Boucher
2b42b28687 fix(#3659): make the worktree base-check trust evidence, not baseRef (#3736)
* test(#3659): baseref-head suppress must be mode-aware regression rows

* fix(#3659): make baseref-head suppress mode-aware and thread isolation mode

* fix(#3659): review fixes - stale advice purge, message pins, mode alias

* fix(#3659): pick-interceptable emit seam, ack merge, writeSync pin

* test(#3659): rewrite set-baseref pin, fix writeSync row stub

* chore(#3659): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 04:58:33 -04:00