Commit Graph

377 Commits

Author SHA1 Message Date
Michel Moreira
76ef60ba25 enhance(#4836): prefer the graphify CLI for planner and researcher graph queries (#4874)
* enhance(#4836): prefer the graphify CLI for planner and researcher graph queries

The planner gets one knowledge-graph query per phase and the researcher two
or three, and that single shot decides which modules the plan treats as
related — and therefore how tasks are ordered into waves. It was spent on
the built-in reader, which seeds by case-insensitive substring match over a
node's label and description and then expands a hardcoded two hops. The
phase "User Authentication" seeds on `author`, `authoring` and
`unauthorized` with the same weight as `authenticate`, and when the
inflated payload exceeds `--budget` the trimmer drops edges by confidence
tier — so the highest-confidence tier can be discarded to fit a payload
that bad seeding inflated in the first place.

The graphify CLI is already a hard dependency of /gsd-graphify build, and
it ranks seeds (IDF weighting, trigram fuzzy matching) and applies context
filters before traversal. Both prompts now prefer it and fall back to the
built-in reader, branching on `command -v graphify` — the same degradation
shape the repo already uses for Context7 to ctx7. Binary presence is a
self-satisfying gate: a graph can only exist if the binary built it, so the
fallback covers edge cases (a CI checkout with a committed graph, a binary
since removed), not the common path. No new config key and no new tool
grant — both agents already have Bash.

The planner additionally runs `graphify affected`. The reference states its
own goal as "which subsystems may be affected by changes in this phase",
which is literally reverse traversal by relation; the built-in reader only
approximates it with undirected two-hop expansion and has no equivalent
verb, so `affected` is skipped on the fallback path.

`graphify status` now reports `graph_path`, the resolved absolute graph
location, on both the present and the missing branch. The CLI takes the
graph location as `--graph`, and the prompts must not re-derive
`.planning/graphs/graph.json` for it: that would point the CLI at a
non-existent local mirror in exactly the umbrella multi-repo setup
`graphify.graph_path` (#1825) exists to serve. For the same reason the
presence gate in both prompts is now the `status` call itself rather than a
bare `ls` of the default location, which was already blind to the override.

Known limit, stated in both prompts rather than implied: the two paths
return different shapes. `graphify query` emits prose and has no `--json`
flag; the built-in emits JSON with per-edge confidence tiers and
budget_met/budget_estimate. `--budget` also counts rendered output on one
and estimated payload bytes on the other (#2738) — same flag name,
different unit. Both are read by a model and nothing machine-parses the
injected block. With graphify absent from PATH the injected context is
byte-identical to before.

Closes #4836

Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — the CLI-first branch, the reason it is preferred, and the output-shape warning are the deliverable; a pointer to a part would not be read at the decision point.
Emitted-Drift-Ack-Growth: gsd-planner.md — one sentence in the load_graph_context step pointer, so it stops naming the default graph path the reference no longer assumes.

* docs(#4836): record the CLI-first graph query in the planner and researcher entries

* chore(#4836): add changeset fragment

* enhance(#4836): name the full domain word in the planner's query-term examples

The reference's own example — phase "User Authentication" → term "auth" — is
the exact collision the CLI-first path exists to avoid, and it stays a
collision whenever the fallback path runs, since that path matches the term as
a substring of label and description.

* fix(#4836): surface graph_path on the unparseable-graph status branch

graphifyStatus() returned graph_path on the exists:true and exists:false
outcomes but not on the third, error, outcome (graph.json present but
unparseable). The planner/researcher prompts gate CLI-first dispatch on
exists, not on this outcome, so a corrupt graph file made them fall
through to the CLI-first branch with the literal <graph> placeholder and
no real path to substitute.

* docs(#4836): note graph_path's trust boundary at the --graph interpolation

graph_path is reflected verbatim into a double-quoted --graph argument the
agent executes via Bash. It comes from graphify.graph_path, a config
surface already trusted elsewhere, so this isn't a new trust boundary --
but it is a new injection site (no --graph flag existed on this call
before). One-line caution for anyone hardening this later.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-22 19:59:45 -04:00
Tom Boucher
6a4984cf69 fix(#4763): surface the displaced session record and pass --phase from the executor decision loop (#4919)
* test(#4763): failing-first — replaced-record payload and executor --phase pins

* fix(#4763): surface the displaced session record and pass --phase from the executor decision loop

state record-session keeps its last-writer-wins write (the recorded single-slot
handoff design) but no longer displaces silently: when a non-empty Stopped At or
authored Resume File record is replaced, the payload carries the full prior text
under replacedRecord. Same-value rewrites, the insert path, and the #944
template-default DWIM are not displacements and report nothing.

The executor decision loop now passes --phase "${PHASE}" to state.add-decision,
matching execute-plan.md, so decisions stop inheriting whichever phase the
global pointer names (#4763 case 2). advance-plan is unchanged (#3311 by-design).

Emitted-Drift-Ack-Growth: gsd-executor.md — the decision loop gained its --phase guard and a comment naming why (#4763)

* docs(#4763): add the changeset fragment

* docs(#4763): backfill the changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-21 12:42:11 -04:00
Tom Boucher
b956bb7c67 fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs (#4902)
* test(#4834): failing-first launcher hijack regressions

* fix(#4834): gate the launcher PATH arm on runtime identity and prefer config-home installs

A gsd_run on PATH that cannot prove it is @opengsd/gsd-core (a foreign package, or a
release older than the runtime-identity verb) is no longer accepted by the launcher
snippet's PATH arm; resolution falls through to the hard error when no path-based
candidate matches. The runtime-config-home arm now precedes the PATH arm, restoring
the documented prefer-local-over-PATH order, so an installer-managed install wins
even against a genuine global. The 16-home probe list is factored into _gsd_homes()
and the identity gate into _gsd_id_ok(), keeping the per-copy delta at +141 bytes.

The files whose frozen ceilings had no headroom (gsd-executor, gsd-plan-checker,
gsd-verifier, gsd-planner, execute-phase, execute-plan) now load the resolver by
@-include from gsd-core/references/gsd-run-resolver.md (the onboard.md pattern)
instead of carrying an inline copy. Propagated to all other inlined workflow/agent
copies via scripts/sync-runtime-launcher.cjs; the resolver reference re-copied
byte-equal (parity B2); the hard-error text, docs/how-to/diagnose-a-foreign-gsd-tools.md,
and the CONTEXT.md launcher predicate updated to match (#4834); the quick-batch row-48
guard gains the canonical-preamble sweep carve-out (#4834, per its own #3730/#2529
precedents); the compact-content benchmark baseline regenerated.

Emitted-Drift-Ack-Growth: add-backlog.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-tests.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: add-todo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ai-integration-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: audit-uat.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: autonomous.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: check-todos.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: cleanup.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: code-review-fix.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: code-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: complete-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: debug.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: diagnose-issues.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: discuss-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: do.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: docs-update.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: edit-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: eval-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: explore.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: extract-learnings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: fast.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: forensics.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: graduation.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-debugger.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-eval-auditor.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-eval-auditor.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-intel-updater.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-intel-updater.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-project-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-project-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-research-synthesizer.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-research-synthesizer.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-ui-researcher.compact.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: health.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: import.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: inbox.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ingest-docs.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: insert-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: list-seeds.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: list-workspaces.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: manager.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: map-codebase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: milestone-summary.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: mvp-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-milestone.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: new-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: next.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: pause-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: plant-seed.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: pr-branch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: profile-user.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: progress.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: quick-batch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: quick.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: remove-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: remove-workspace.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: resume-project.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: scan.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: secure-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings-advanced.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings-integrations.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: settings.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ship.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sketch-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sketch.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: smart-entry.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spec-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spike-wrap-up.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: spike.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: stats.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: sync-skills.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: thread.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: transition.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ui-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ui-review.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: ultraplan-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: undo.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: validate-phase.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)
Emitted-Drift-Ack-Growth: verify-work.md — launcher snippet resolution hardening propagates via sync-runtime-launcher (#4834)

* docs(#4834): backfill the changeset PR number

* test(#4834): regenerate the compact-content benchmark baseline after the rebase

---------

Co-authored-by: sim <sim@local>
2026-09-21 02:27:36 -04:00
Tom Boucher
eea9247c93 enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select

Auto-mode used to auto-select a checkpoint:decision's first <option>
unconditionally, making a decision checkpoint's safety depend on option
presentation order rather than an authored choice. Add an optional
auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode
now escalates to a human exactly like gate="blocking-human" does; present,
it names the option auto-mode selects; naming an id with no matching
<option id> is a hard structural-validation error at plan-parse time
rather than a silent fallback to the first option. gate="blocking-human"
continues to win over everything, unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys

An isolated adversarial review of the auto_select work found that both new
attribute regexes used \b as their left anchor, which is a word boundary,
not a "start of attribute name" boundary. A decoy attribute ending in the
same word (e.g. data-id="...") sitting before the real id="..." on the
same <option> tag matched first, silently corrupting the extracted option
id. Anchor on (?:^|\s) instead so only the real attribute name can match.
Adds a regression test reproducing the exact decoy-attribute shape, plus a
Unicode option-id test from the same review pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane

lint-docs-guard-registration failed: the new test reads docs/reference/
plan-md.md but was not registered, so a future edit to that doc could
silently desync from the test without the guard catching it on the PR
that changed the doc. Registered alongside its direct precedents
(precondition-element.test.cjs, reversibility-tagging.test.cjs), which
read the same file for the same reason.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling

The remote gsd-test run caught what local checks missed: execute-phase.md
carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes
of headroom before this change, and the original checkpoint:decision
wording pushed it to 93772 (over the ceiling). Cascaded into failures in
phase6-capstone-conformance, execute-phase-completion-reconciliation,
claude-orchestration, and the compact-content drift-report test.

Also caught: tests/package-legitimacy-gate.test.cjs anchors a
"decision is conditional, not unconditional" safety check on the literal
phrase "first option" in the decision bullet — which #4095 deliberately
removes, since there is no more unconditional first-option pick. The test
was asserting an assumption this change intentionally makes obsolete;
re-anchored on tokens that still identify the bullet ('decision',
'auto-spawn') without weakening what the test actually verifies (the
bullet must still carry a blocking-human carve-out).

Also fixed a word-order mismatch between my own new test's regex and the
actual doc text it was asserting against (tests/auto-select-attribute.test.cjs).

Regenerated the compact-content benchmark baseline
(tests/fixtures/compact-content-benchmark-baseline.json) to match the new
byte counts.

Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap.
Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4095): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:47:01 -04:00
0xdhx
88b5775dc8 enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI (#4477)
* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI

gsd-ui-auditor is chartered to audit interaction and handed a capture
driver with no interaction verb: `npx playwright screenshot` cannot
click, fill, hover, press or snapshot, so a hover state, an open menu,
a focus ring or a form's validation state never appears in its
evidence and every Experience Design finding degrades to code reading.

Implements the shape approved at triage, not a new capability:

- capabilities/ui/capability.json declares `workflow.ui_interaction_capture`
  (boolean, default false) on the capability that already owns the
  auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated.
- gsd-core/workflows/ui-review.md reads the key through gsd_run and hands
  it to the auditor as `interaction_capture:` in the spawn <config> block —
  the auditor carries no gsd_run resolver, so the key travels by value.
- agents/gsd-ui-auditor.md gains an anchored interaction-capture section
  AFTER the static block. With the key on and a Chrome binary resolved it
  starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on
  an --isolated profile, opens the dev URL the static block reached,
  takes the a11y snapshot for element uids, captures the baseline and a
  Tab focus-ring state, drives the UI-SPEC's interactive components, saves
  console output, and stops the daemon unconditionally. Key off, no dev
  server, or no Chrome: one status line, and the Playwright-only static
  path runs exactly as before — the static fence is untouched.

Needs only Bash: no MCP server, no tools: change. Chromium-only by
nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so
readiness is polled through evaluate_script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* test(#4223): bind the interaction-capture shape and containment

- manifest, generated registry, config schema and config-set/loadConfig
  all know workflow.ui_interaction_capture as a default-off boolean, and
  hand-written non-booleans fall to the slice default
- the orchestrator reads the key and hands it down; the auditor never
  grows a gsd_run dependency
- the static fence stays Playwright-only and the interaction fence
  chrome-devtools-only, so key-off is today's path
- the interaction fence runs under bash with a stub driver on PATH: key
  off / absent / no dev server / no Chrome invoke nothing; the happy path
  starts first and stops last on the [selected] pageId with the documented
  flags; a failed capture is removed and not counted; new_page and start
  failures still honour the stop-only-if-started rule; CHROME_BIN and
  CHROME_DEVTOOLS_MCP_VERSION overrides flow through
- docs/CONFIGURATION.md row shape; registered in the docs-guard lane

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* docs(#4223): document workflow.ui_interaction_capture and its how-to

- docs/CONFIGURATION.md: one row in the workflow.* table, default-off
- docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the
  interaction-capture section adds, skips and never claims
- docs/how-to/enable-ui-interaction-capture.md: turn it on, read the
  `**Interaction captures:**` outcomes, what it does not do, turn it off
- docs/README.md: index the how-to beside live-DOM verification

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* chore(#4223): add changeset

Added-type fragment; pr: carries the issue number until the PR exists.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose

Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace;
the hyphen form is retired there and the slash-command-namespace guard
rejects it. docs/ keep the hyphen form by convention.

Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered.
Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): per-run daemon session, bounded navigation, and step failures that count

Three findings from the pre-file adversarial review of the interaction
fence, folded in:

- `--sessionId <epoch>-<pid>` on every driver call. `start` restarts
  whatever daemon shares its session and --isolated isolates only the
  browser profile, so two concurrent audits — or an audit beside the
  operator's own CLI daemon — would otherwise stop each other. The CLI
  accepts hex and dashes only; the id is validated by the test stub.
- `new_page --timeout 30000`: the one verb that takes a bound, placed
  before every verb that does not, so a hung page is caught first.
- a failed take_snapshot or press_key now increments the failure count
  and is named on stdout; two clean screenshots can no longer read as
  `0 failed` after the step that gives the interactions their uids failed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal

Second review round, both reviewers:

- session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a
  subshell, so two audits forked from one parent in the same second
  shared an id and could stop each other's daemon (driven by the reviewer)
- `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver
  under Git Bash still matches the `$` anchor, and `|| true` on the
  assignment so a failed new_page cannot abort the block under
  `set -e -o pipefail` before the unconditional stop
- a failed take_snapshot removes any snapshot.txt it left or inherited
  from a reused directory, so stale uids never drive the interactions
- `<config>` placeholder is `{interaction_capture}`, lowercase like its
  `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt
  template, not a bash heredoc
- how-to: the `not captured` row no longer claims the daemon started

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges

Third review round:

- a new_page that prints a page line and then exits non-zero is a failed
  navigation, not a page id: the exit status is checked in an `if` before
  the output is parsed (driven by the reviewer against the previous
  `|| true`, which masked exactly that)
- regression cases for what the last two rounds added: CRLF driver
  output, a stale snapshot removed on failure, partial-output new_page
  failure, and the whole fence under `set -e -o pipefail` (both the
  failed-navigation path and the happy path)
- the harness whitelist gains `date`; the session-id assertion now
  requires all three parts, so a silently empty epoch cannot hide again

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder

- the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes,
  over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces;
  the interaction section's comments are tightened to the same content
  in fewer bytes (23559 now). No bash changed — the fence's own tests and
  the real-browser run are unchanged.
- .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which
  the post-create backfill rewrites to the PR's own number.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* chore(#4223): set changeset fragment pr to 4477

* test(#4223): compare the fence's status path with the separator the fence uses

On the windows-latest lane the happy-path case failed on `\interaction` vs
`/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a
literal slash, and the assertion built its expectation with path.join. Every
other case in the file passed on that lane, including the CRLF and
errexit/pipefail ones.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6

* test(#4223): drop the inert file-header allow-test-rule marker

Review round 1 on #4477: the `source-text-is-the-product` marker sat at
line 2, outside no-source-grep's 8-line lookahead of every readFileSync
site (the first is ~60 lines down), so it suppressed nothing. It was also
unnecessary: every read in this file is a .md/.json path, which the rule
does not trigger on. Deleted rather than relocated — there is no site to
relocate it to. Negative control: `eslint` on the file is clean without it.

* chore(#4223): regenerate the platform-conformance-tier lists for the new test

Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier
gate after this branch opened; its two committed lists must name every file
under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never
been in them. Once the branch was updated against next the lists were stale
and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier
--check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's
fragment-single-edit-propagation, which sees the same staleness as regen:derived
touching files beyond the fragment edit under test.

Regenerated with the repo's own generators, no hand-editing. The general tier
goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry;
both --check arms are clean. Verified the red is this PR's own file and not
base drift: at upstream/next both generators report "list matches" (546 / 196),
and our committed copies were byte-identical to next's before this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ

* fix(#4223): bound, confine and trap the chrome-devtools driver fence

Round 4 — three findings in one fence, interleaved on the same lines, so one
commit:

- Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a
  background job in its own process group (`set -m`) under a watchdog that kills
  the whole group at the ceiling — TERM, then KILL two seconds later. One pid is
  not enough: npm forwards SIGTERM only to its direct child, so killing `npx`
  alone leaves the client holding the fence's stdout and a `$(cdt … new_page)`
  capture blocked past the ceiling (driven against a real npx tree by the
  round's adversarial review; the pid-only first cut of this commit had exactly
  that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )`
  subshell: a subshell inherits bash's saved copies of the caller's stdio (the
  fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them
  open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a
  watchdog outlived its kill — measured as the intermittent 30 s test run the
  review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 --
  -pgid`, every 0.1 s) and stands down by itself once the group is empty;
  nothing ever signals it. The group, not the leader pid: a child can outlive
  the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll
  stood down at once and left the substitution open-ended (driven by the
  round's adversarial review at 6× the ceiling; a pgid cannot be reused while
  any member lives, which a bare pid can). The daemon `start` launches is
  spawned detached (its own session) and never in that group. Two platforms
  forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM;
  sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out
  the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git
  Bash a signal to a watchdog still starting up hung the fence's `wait` for it:
  18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs,
  while a fence slowed by xtrace, or three tests run alone, never hit it (a
  startup race; the mechanism is not pinned further). Polling: 0/300 slow calls
  and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3
  full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU,
  msys and busybox sleep; BSD sleep documents it). A clock that cannot launch
  (`sleep … || exit 0`) stands the watchdog down rather than firing at once
  and killing a healthy call — by design that leaves a hung call unbounded,
  the pre-round-4 behaviour, instead of failing a healthy one. A hung call
  returns once its group is gone: at the ceiling, plus up to the 2 s
  TERM-to-KILL grace. The KILL after the grace is sent only to a group that
  is still alive: a pgid freed during the grace can be reused, and an
  unconditional KILL could hit an unrelated group (the round's review).
  `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT
  (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent
  on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog
  rather than either.
- --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may
  write under the run's interaction/ directory and nowhere else. Relative, like
  every --filePath (unchanged from rounds 1-3): the daemon resolves both against
  one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and
  path.resolve()s both), and a relative path needs no dialect translation — an
  absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native
  daemon cannot resolve (CI's windows conformance shard caught the first cut).
  --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified),
  so the documented floor moves from ^1.8.0 to ^1.9.0, where
  --allowUnrestrictedPaths is deprecated.
- `stop` is owed by an EXIT trap after a successful `start`, not by position
  (it replaces any earlier EXIT trap — none exists in this file); the explicit
  call keeps it in order, a flag makes the trap a no-op afterwards, and only the
  shell that installed the trap may act: a subshell copy of the fence state
  carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced
  locally at 3/40 under load: the second `stop` came from a subshell pid, never
  main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo
  "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist
  and CI's macos conformance job showed the guard comparing empty to empty. The
  fence was driven under bash 3.2.57 for the injected-subshell, errexit
  failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy
  paths. A failed resize_page is a counted failed step now, not the one bare
  command an errexit runner could abort on.

Prose in the section is tightened to pay for the mechanism: 23559 -> 24517
bytes against the 24576 DEFAULT-tier cap.

Tests: the stub driver hangs as a real child tree (sh waiting on a child that
holds stdout — never an exec), so a pid-only kill fails the new
aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative-
controlled: it blocks for the harness's whole cap on the old wrapper). A hung
start and a hung capture are cut off within ceiling + grace + slack and still
reach stop; an injected bare failure under errexit reaches stop through the
trap, exactly once; an injected subshell call of cdt_stop issues nothing; the
happy path issues exactly one stop; every driver call site names a ceiling and
the only bare $CDT is the wrapper's own spawn; the start line carries
--workspace with the capture directory, every --filePath lies under it, and no
code line carries --allowUnrestrictedPaths. A driver whose leader exits at
once while a child keeps holding the capture pipe is still cut off at the
ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole
cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone
(negative-controlled: the trap form kills `start` in under 20 ms). The harness
EXPORTS its stub-only PATH — unexported, the exec'd
watchdog fell through to bash's compiled-in default PATH and never saw the stub
dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash,
pid-preserving).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

* fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file

Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the
accessibility tree, with entered form values) and console.txt (which can carry
tokens) were committable by `git add .`. The gate now ignores `interaction/` as
a directory — the next artifact type is covered by construction — and it appends
whatever an existing .gitignore lacks instead of writing once. The write-once
form was the same defect one step later: every project that had already run an
audit would never have received the new pattern at all.

Tests run the gate fence under bash: a fresh file carries every pattern; an
image-only file from an earlier audit gains interaction/ and keeps its own
header without duplicating present lines; a second run appends nothing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

* test(#4223): declare the interaction-capture anchor as a comment marker

The #4324 colon-token gate (slash-command-namespace) landed on next after this
branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an
unconvertible /gsd: command token. It is a section anchor of the same family as
gsd:live-dom-families and gsd:write-continue, so it is declared in
COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added
gates against the merged tree; CI at ca8d2508 predates the gate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
2026-09-20 05:45:15 -04:00
Tom Boucher
c9a5cc3e12 fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head.
2026-09-17 19:06:55 -04:00
Tom Boucher
651511d1e3 fix(#4670): bound the commit-claim window to the plan's own history (#4813)
* test(#4670): add failing-first coverage for the bounded commit-claim window

* fix(#4670): bound the commit-claim window to the plan's own history

The reconciliation measured plan_head_before..HEAD — a window that grows
with every later plan's task and SUMMARY commits plus execute-phase's own
phase-completion commit — so an honest plan flagged commit_claim_mismatch
as soon as anything landed after it (real project: claims 3/5/2 measured
20/10/5).

The executor now also records plan_head_after (HEAD at its measurement
moment, after the last task commit, before the SUMMARY commit), and
verify-work reconciles exactly against plan_head_before..plan_head_after
with a merge-base ancestry check; SUMMARYs without the anchor fall back to
the legacy warning path instead of an unsound BLOCKER. Both #3968 failure
modes (claimed commits never made; task commits lost) still block, driven
by the issue's own fixture scenarios.

Emitted-Drift-Ack-Growth: verify-work.md — reconciliation gains the bounded window and legacy fallback (#4670)
Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670)

* fix(#4670): name the history-rewrite case and keep the executor under its cap

Review findings: the BLOCKER enumeration named only the two #3968 causes,
so an honest plan whose recorded window was rewritten afterwards (rebase,
amend, cherry-pick) got a mislabeled diagnosis — the clause now names that
case with the manual-recount remedy. The executor's growth crossed the
LARGE-tier hard cap (49152), so the plan_head_after documentation is
compressed to the minimal capture + frontmatter write (verify-work.md
carries the semantics), the #2751 PROSE_ALLOWLIST entry is re-pointed at
the shifted line (#4670 moved it from 823 to 825), the compact-content
baseline is regenerated, and the changeset records the two un-established
edges (shared-base waves, subrepo ledgers).

Emitted-Drift-Ack-Growth: gsd-executor.md — plan_head_after anchor documented in the measurement protocol (#4670)

* docs(#4670): backfill changeset PR number

* docs(#4670): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-09-16 22:06:03 -04:00
0xdhx
003d982c83 fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint (#4749)
* fix(#4623): keep repo-wide planning docs out of the verification digest, and accept --files on verification.fingerprint

Two defects in the covered-input fingerprint (#4155), one issue.

1. `computeCoveredDigest` hashed the whole bytes of every declared path
   uniformly, so `.planning/ROADMAP.md` and `.planning/REQUIREMENTS.md` —
   which every phase rewrites as ordinary bookkeeping, and which the closing
   phase's own `phase.complete` / `requirements mark-complete` rewrite AFTER
   the verifier ran — flipped every phase that declared them to `stale` on
   zero implementation change, and from there `isPhaseComplete` →
   `init.manager` → `complete-milestone`'s `ALL_PHASES_VERIFIED` gate.
   Fingerprint v2 leaves any direct child of a planning root out of the
   hash: `.planning/` itself, plus the phase's own planning root (the parent
   of its `phases/`, so `planningDir`'s `<project>/` and `workstreams/<ws>/`
   layouts are covered without the digest knowing what a workstream is —
   `sharedPlanningRoots` / `isSharedPlanningDoc`, defined by position rather
   than a name list so the set cannot drift; a root is accepted only when the
   phase dir sits under a `phases/` directory inside `.planning/`). Such a path is still validated
   exactly as every other covered path (confined, present, a regular file —
   the fail-closed contract is unchanged); only its bytes are ignored, and a
   declaration made only of shared documents fails closed like an empty one.
   A stored digest names its version, and `readVerificationStatus` now
   recomputes under THAT version (`parseFingerprintVersion`,
   `KNOWN_FINGERPRINT_VERSIONS`): a legacy v1 report keeps v1 semantics
   until it is re-fingerprinted, so the upgrade alone stales nothing; a
   version this build cannot recompute fails closed.

2. `verification.fingerprint` received a raw positional slice, so
   `--files a`, `--files "a,b"` and `--files a --files b` all put the literal
   token into the covered set and failed closed as "a covered file is
   missing, unreadable, or escapes the project root" — the message that
   convinced the reporting project the digest was permanently
   unrecomputable. `parseFingerprintFileArgs` accepts every form (plus
   `--files=a,b`, freely mixed with bare positionals), treats any other
   `--flag` and an empty `--files` value as usage errors that say so, and
   the phase-dir argument must now be an existing directory: omitting it
   used to take the first covered file as the phase dir and print a
   plausible digest over the rest at exit 0.

Regression tests (tests/verification-status.test.cjs, #4623 block): the
cross-phase case from the report, the same-phase `requirements
mark-complete` / `phase.complete` cases from the thread, a workstream-scoped
root, v1-preserved / unknown-version-stale, the fail-closed cases (missing,
directory, escaping symlink, all-shared), every `--files` form against the
bare form, the unknown-flag / empty-value / omitted-phase-dir errors, and
AC5's zero-file error. Verified failing against the pre-fix source: 29 of 34
fail, the 7 that pass pin behaviour the fix must leave unchanged.

Docs: CONTEXT.md Verification Module, agents/gsd-verifier.md's
covered_files instruction (rewritten in place — the file sits 21 bytes under
its LARGE hard cap), gsd-core/templates/verification-report.md.

Fixes #4623

Emitted-Drift-Ack-Growth: gsd-verifier.md — the #4155 covered_files instruction now states that planning-root docs are digest-inert (#4623); +18 bytes, under the LARGE cap
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DCMY8P8s6dp4g3Rxu3nNAi

* chore(#4623): set changeset fragment pr to 4749

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 05:26:44 -04:00
0xdhx
48271de43f fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis (#4744)
* test(#4660): pin the letter-axis parity defect across all 6 shell/markdown phase-id sites

Extends tests/nsegment-phase-grammar.test.cjs (#4568) one axis over: for each
of the six sites, reads the live regex off disk and asserts it agrees with
src/phase-id.cts's PHASE_NUMBER_TOKEN_SOURCE on the letter axis in BOTH
directions — accepts `12A` / `3A` / `03A` / `23A.1.2`, still rejects `3a`,
`3AB`, `A3` and the other canonical-invalid shapes — and that the two
extracting sites return the full letter-suffixed token rather than its digit
prefix (or nothing).

Negative control against the unfixed tree: 22 failures, exactly the
"(fails before the fix)" cases; every reject-parity case already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* fix(#4660): widen the 6 shell/markdown phase-id mirrors to the canonical grammar's letter axis

Adds `[A-Z]?` after the leading digit run at all six sites #4568 widened —
the ERE translation of src/phase-id.cts's `\d+[A-Z]?(?:\.\d+)*` — so a
documented, canonical-valid id like `12A` or `23A.1.2` is no longer refused
by the four validating sites (code-review.md, code-review-fix.md,
gsd-code-fixer.md, gsd-code-fixer.compact.md) or truncated to its digit
prefix by the two extracting sites (execute-plan.md's plan-filename grep,
plan-phase.md's --research-phase capture). Behaviour is byte-identical for
every id that matched before; the adjacent comment and error-message text
now names the grammar it mirrors.

Driven: `init code-review 3A` on a fixture with a `03A-slug/` directory and
a `### Phase 3A:` heading emits `padded_phase: "03A"`, which the old regex
rejects and the widened one accepts — nothing upstream of the validator
mangles the id.

At execute-plan.md the trailing `-[0-9]+` is the PLAN number and stays
digit-only; plan and milestone dimensions are out of scope per the brief.
`CASE_FLEXIBLE_PHASE_NUMBER_TOKEN_SOURCE` derives from the canonical source
by a literal `.replaceAll('A-Z', 'A-Za-z')`, so src/phase-id.cts is
deliberately untouched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore(#4634): extend lint-phase-id-drift to ban a letter-less phase-id mirror in workflows/ and agents/

Adds findLetterlessPhaseMirrorDrift — the letter-axis twin of the #4568
single-segment rule — flagging the unbounded-segment shape
`[0-9]+(\.[0-9]+)*` (and its \d / doubled-backslash near-variants) whose
digit run is NOT followed by the `[A-Z]?` class, on any phase-carrying line
across gsd-core/workflows/**/*.md, gsd-core/references/**/*.md and
agents/**/*.md. Sanctioned the same way (`<!-- phase-id-owner: ... -->`),
tolerates the case-flexible `[A-Za-z]?` directory-scanning variant so it
cannot force that separate axis to narrow, and is wired into scanAll.
Confirmed zero violations against the real tree post-#4660 fix, and one
violation when a single site is reverted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* docs(#4660): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore: regenerate conformance-tier manifests for the extended grammar test

tests/nsegment-phase-grammar.test.cjs now requires the compiled
gsd-core/bin/lib/phase-id.cjs (to assert the canonical grammar agrees with
each site's live regex), which moves it to a different platform-conformance
tier; `gen-platform-conformance-tier.cjs --check` in lint:ci flagged the
macOS manifest as stale.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* test(#4660): reword a comment that tripped lint-docs-guard-registration

The comment mentioned `docs/CONFIGURATION.md` between two backticked
tokens, which the lint's template-literal detector read as a docs/ path
expression. The test reads no docs/ file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore(#4660): refresh the compact-content benchmark baseline and acknowledge emitted growth

plan-phase.md grew by 4 bytes (`[A-Z]?`), which moves the committed
compact-content benchmark; refreshed with `benchmark-compact-content.cjs
--write`. The six shipped files below grew by the widened regex literal plus
the comment and error-message text that now names the canonical grammar.

Emitted-Drift-Ack-Growth: code-review.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example
Emitted-Drift-Ack-Growth: code-review-fix.md — #4660: `[A-Z]?` at the PADDED_PHASE validator plus a comment/error message naming the canonical grammar and the `12A` example
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the defense-in-depth comment and error message updated to the canonical grammar
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — #4660: `[A-Z]?` at the padded_phase sink validator plus the comment and error message updated to the canonical grammar
Emitted-Drift-Ack-Growth: execute-plan.md — #4660: `[A-Z]?` in the plan-filename phase extraction (6 bytes)
Emitted-Drift-Ack-Growth: plan-phase.md — #4660: `[A-Z]?` in the --research-phase capture (6 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

* chore(#4660): set changeset fragment pr to 4744

* chore: re-trigger Validate Branch Name

The required check-branch context was cancelled on this head by the
workflow's cancel-in-progress group when the changeset pr-field backfill
push landed three seconds after the PR opened; no completed run exists for
the current head, and a fork contributor cannot re-run it. Empty commit to
re-run it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NLtEbRc1Qfbe95HRMNqwp3

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-14 19:32:12 -04:00
Tom Boucher
6f99e493e7 fix(#4395): make the debug session manager's own gsd-debugger spawn blocking (#4718)
* test(#4395): prove the manager spawns its debugger without blocking

Failing-first regression coverage for #4395.

debug.md:209 mandates the orchestrator to session-manager spawn carry
run_in_background: false, and says why outright: "Claude Code backgrounds
subagents by default, and only that flag makes the spawn return the
compact session summary directly" (#2196).

The session-manager to debugger spawn, one level down, carries no flag.
Measured: run_in_background appears nowhere under agents/ -- only in
gsd-core/workflows/. So by the rule #2196 itself states, that spawn is
backgrounded, Step 3 ("Handle Agent Return") has no return to inspect,
the manager emits CONTINUE_REQUIRED, the orchestrator auto-resumes per
#2257/#3448, and a second detached debugger races the first on
.planning/debug/<slug>.md.

Row 4 is the load-bearing one: it closes the CLASS by requiring every
subagent spawn under agents/ to declare run_in_background explicitly, so
the next agent that spawns one has to decide rather than inherit a silent
host default. It is scoped to agents/ precisely so it cannot misfire on
the workflows that deliberately use true for parallel fan-out.

Rows 5-7 are pins, not fixes: the #2196 mandate one level up, Step 2 as
the single spawn-format source that the eight continuation sites delegate
to, and the survival of CONTINUE_REQUIRED (which has a legitimate trigger
unrelated to this defect).

Red round: 4 of 7 rows fail. Row 3 needed hardening first -- asserting
only that the two variants AGREE passed vacuously, because two missing
flags are also equal; it now asserts each is present before comparing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4395): make the manager's own debugger spawn blocking

The orchestrator-to-manager hop already requires a blocking spawn and says
why (#2196, debug.md:209): Claude Code backgrounds subagents by default,
and only run_in_background=false makes the spawn return its summary. The
manager-to-debugger hop, one level down, carried no flag -- measured,
run_in_background appeared nowhere under agents/ at all.

So that spawn was backgrounded. Step 3 ("Handle Agent Return") opens
"Inspect the return output for the structured return header" -- with
nothing to inspect, the manager correctly declined to fabricate a terminal
summary and returned CONTINUE_REQUIRED; the orchestrator correctly
auto-resumed (#2257/#3448); the resumed manager reached Step 2 and spawned
a SECOND detached debugger. Both then raced on .planning/debug/<slug>.md.

Every observable in the report follows with no further assumption,
including the count: the reporter saw exactly three collisions in one
invocation, and debug.md:251 caps auto-resumes at three per slug -- one
collision per cycle.

Fixed at the cause, in both shipped variants, kept byte-consistent. The
eight continuation sites say "see Step 2 format", so they inherit it.

The issue offered two remedies. The second -- have the auto-resume path
reconcile a still-running debugger before spawning another -- is not taken:
it treats the symptom, and needs machinery that does not exist (no portable
way to enumerate or stop another runtime's live agents, plus an in-flight
sentinel with staleness and recovery rules, or an orphaned marker deadlocks
the session permanently). With the spawn blocking, the manager cannot reach
Step 4 while a debugger is live, so such a guard would also be unreachable.

#2257, #3448, the anti-loop heuristic, the cap of three, and the
CONTINUE_REQUIRED shape are all correct and untouched. CONTINUE_REQUIRED
keeps its legitimate trigger: the manager genuinely exhausting its own turn
budget mid-investigation.

Also corrects the red-round test to the canonical CALL form. debug.md
writes run_in_background=false inside Agent(...) and run_in_background:
false in prose; the first draft asserted the prose form, which the shipped
call would never have matched.

Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — the blocking spawn flag plus the note recording why an unstated flag produced colliding debuggers
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — same change as its full sibling, kept byte-consistent with it

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4395): add changeset fragment

pr:0 placeholder is backfilled with the real number once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4395): refresh the variant benchmark baseline

Two entries move.

gsd-debug-session-manager.md 4766/4477 -> 4938/4649 is this change: the
blocking-spawn flag plus its explanatory note, added to BOTH variants to
keep them byte-consistent, so the compact sibling grows by the same amount
and the pair's reduction ratio dips 6.06 -> 5.85. The compact file remains
strictly smaller than its canonical sibling, which is what the variant
guard's size check actually requires.

gsd-code-fixer.md 10741 -> 10740 is NOT from this branch -- the file is
untouched here. It has scored 10740 since f334f277dd (#4324) reworded a
line without refreshing this fixture, so the stale number is sitting on
next. Fixed here rather than deferred; the #4350 branch carries the
identical one-token correction, so whichever lands first makes the other a
no-op.

Refreshed with the variant script's own --write.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4395): state the blocking rule agent-wide and correct two overclaims

Review round. Three substantive corrections, one of them to a claim I made
in the previous commit message.

1. The blocking rule is now stated AGENT-WIDE, not per-call. My earlier
   claim that "the eight continuation sites say 'see Step 2 format', so
   they inherit it" was false: exactly ONE of them names Step 2 (the
   compact variant says "Step 2 format" without the "see", which is why
   the first draft of the test matched it zero times there). The other
   sites inherit only because Step 2 holds the sole Agent() spawn literal
   in each file. Both variants now say so outright, and the test pins the
   sole-literal invariant in BOTH variants rather than the prose wording
   in one.

2. debug.md's two auto-resume buckets now restate run_in_background=false
   for the re-spawn. "The same session_params" does not carry it --
   session_params is prompt content, not the spawn flag.

3. The test extractor now also matches single-line Agent(...) calls. The
   class guard was blind to exactly the shape a future offender is most
   likely to take.

Also corrects the diagnosis: remedy 2 is NOT unreachable once the spawn
blocks. This agent's own retained CONTINUE_REQUIRED trigger is "turn
budget exhausted WHILE THE DEBUGGER IS STILL INVESTIGATING", so a harness
turn cutoff mid-wait still double-spawns; debug.md:209 names the same
class from the other side. What this fix removes is the SYSTEMATIC case --
every invocation, because the spawn was always backgrounded. The residual
turn-cutoff window survives, bounded by the existing three-resume cap, and
is stated in the PR rather than denied.

Emitted-Drift-Ack-Growth: gsd-debug-session-manager.md — blocking spawn flag, the note recording why an unstated flag produced colliding debuggers, and the agent-wide restatement the per-site inheritance actually depends on
Emitted-Drift-Ack-Growth: gsd-debug-session-manager.compact.md — same change as its full sibling, kept byte-consistent with it
Emitted-Drift-Ack-Growth: debug.md — both auto-resume buckets restate the spawn flag, since session_params does not carry it
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4395): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 02:12:46 -04:00
Tom Boucher
f334f277dd fix(#4324): stop the retired /gsd: prefix reaching users (#4712)
* test(#4324): prove colon tokens the installer cannot convert leak

Failing-first regression coverage for #4324. The install rewrite
(transformContentToHyphen) is gated on an exact match against the
commands/gsd stem list, so any /gsd:<token> whose token is not a
registered stem survives the install and reaches the user as the
deprecated colon form.

The gate is load-bearing -- it is the only thing protecting the
workflow DSL marker family (gsd:section, gsd:protected, gsd:loop-host,
gsd:guard, gsd:dispatch, gsd:plan-revision-conflicts), which
workflow-fragments parses as a literal. So this suite asserts the
shipped text is convertible rather than asserting the transform is
broad, and pins the marker family as explicit negative space.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4324): stop unconvertible colon tokens reaching the user

The install rewrite is gated on an exact match against the commands/gsd
stem list, so a /gsd:<token> whose token is not a registered stem
survives the install and reaches the user as the deprecated colon form.
That gate is load-bearing -- it protects the gsd:section /
gsd:protected / gsd:loop-host marker family -- so the fix is in the
shipped text, and the source stays colon per CONTEXT.md's two-tier rule.

- quick-batch command + skill description: close the command token at a
  boundary so `/gsd:quick`-shaped converts instead of being skipped.
- gsd-code-fixer (both variants): execute-plan and diagnose-issues are
  workflows, not commands, so they never converted and rendered beside
  two hyphenated siblings on the same line. Name them as workflows.
- help topic-mode: the extraction rule hard-coded a colon prefix that
  the converted full.md never ships, so --brief could never match a
  signature line and silently fell back on every topic. Describe the
  signature line without a literal prefix.
- update.md: drop the prefix from prose describing a stale command.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4324): add changeset fragment

pr:0 placeholder is backfilled with the real number once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4324): locate the help summary per reference variant

Adversarial review finding. Restoring the signature-line match (the
#4324 fix) activated a latent defect in the clause next to it: compact
scope emitted "the single non-blank line immediately after" the
signature, and that clause is only correct for full.md.

full.compact.md puts the summary on the signature line itself, after an
em-dash, and its next non-blank line is an unrelated "Usage:" line. Both
variants ship and both are served, so before this commit the compact
variant would have emitted the wrong line as the summary. It was masked
until now only because the stale colon prefix meant no signature line
ever matched at all.

Name the two placements and pick per line, and say explicitly that a
Usage: line is never a summary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4324): de-vacuum the help parity check, narrow the marker waiver

Two adversarial review findings against the #4324 coverage.

The help-parity assertion went vacuous the moment the fix landed: once
topic.md stops spelling a literal prefix, the matched set is empty and
the assertion holds for any rewording, correct or not. It now also
asserts across BOTH served reference variants that each ships signature
lines under the hyphen prefix, that the two genuinely disagree about
where the summary sits, and that topic.md still names both placements
and the Usage: guard.

The marker waiver keyed on "sits inside an HTML comment", which waves
through a real broken reference that happens to be commented out --
`<!-- see /gsd:typo-cmd -->` scored clean. Enumerate the six marker
families instead. Verified the narrowed rule catches that probe and
still passes over the tree; it also surfaced a seventh family,
write-continue, that the broad rule was hiding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4324): normalize the namespace in skill descriptions

Both hyphen-namespace skill converters ran the hyphen transform over the
body but rebuilt the frontmatter description from the raw field, so a
/gsd:<cmd> mention in a command description survived into the installed
SKILL.md -- the exact field the host's skill picker renders, which is
the surface this issue was filed about.

The local flat-command path was already correct because it rewrites the
whole file; only the skills path, used by a global install, was
affected. Confirmed by installing into a fake HOME before and after.

Fixed in both copies: bin/install.js and the src/ source of truth that
compiles into gsd-core/bin/lib.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4324): assert descriptions through the real converters

The previous version of this check called transformContentToHyphen on
the description line itself and passed, while a real install still
shipped the colon form -- the converter never calls that transform on
the description. It asserted a proxy for the behaviour instead of the
behaviour.

Drive convertClaudeCommandToClaudeSkill and
convertClaudeCommandToClineSkill over every registered command and
assert on the emitted description. Verified it fails against the
pre-fix converters and passes against the fixed ones.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4324): regenerate skills after the description change

skills/<name>/SKILL.md is generated by gen-plugin-skills, not
hand-maintained, and lint:generated-sync caught the hand edit. The
regenerated file emits the hyphen form, which also corrects the
assumption behind the scan comment in the namespace test: skills/ is
runtime-emitter output, not colon source.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4324): re-sanction normalizeKimiSkillName's real end line

The description-normalisation fix inserted five lines above
normalizeKimiSkillName in src/runtime-artifact-conversion.cts, moving its
closing brace from 635 to 640. MAJOR-1 pins that line deliberately, so the
planted violation landed INSIDE the exempted body and went unflagged --
0 !== 1.

Re-sanction the value rather than derive it: the array is named
sanctionedRealEndLines, and a pinned line that fails loudly on drift is
the design. Deriving it would remove the human check the name asks for.

Verified by executing all four MAJOR-1 rows against the real tree: each
planted violation is flagged at realEndLine+1 and each unmodified file
stays exempt.

Emitted-Drift-Ack-Growth: gsd-code-fixer.md — names execute-plan and diagnose-issues as workflows rather than as slash commands that do not exist
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — same rewording as its full sibling, kept byte-consistent with it
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4324): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 22:17:48 -04:00
Tom Boucher
a2331c01f1 fix(#4568): widen the phase-number regex to accept N-segment ids at 6 shell/markdown sites (#4646)
* test(#4568): pin the N-segment phase-grammar defect across all 6 shell/markdown sites

Manually traced against the current tree: the validating regex at
code-review.md rejects a 3-segment id (23.1.2), and execute-plan.md's
extraction truncates a 23.1.2-01-PLAN.md filename down to 1.2-01.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4568): widen the phase-number regex to accept N-segment ids at all 6 shell/markdown sites

Widens `?` to `*` on the dotted-segment group at all 6 sites (byte-identical
behavior for 1- and 2-segment ids, character class unchanged): code-review.md,
code-review-fix.md, gsd-code-fixer.md, gsd-code-fixer.compact.md (validating
sites, plus their comment/error-message text), execute-plan.md's plan-filename
extraction, and plan-phase.md's --research-phase flag capture.

Also disambiguates the nsegment-phase-grammar test's plan-phase.md anchor,
which was matching an unrelated earlier `--research-phase` occurrence (line
77's generic-value capture) instead of the targeted site (line 131).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): extend lint-phase-id-drift to ban the single-segment phase regex in workflows/ and agents/

Adds findSingleSegmentPhaseRegexDrift, banning the bounded
`[0-9]+(\.[0-9]+)?` shape (and its \d/doubled-backslash near-variants) on any
phase-carrying line across gsd-core/workflows/**/*.md,
gsd-core/references/**/*.md, and the newly-scanned agents/**/*.md, sanctioned
the same way as the existing shell-arith rule. Wired into scanAll; confirmed
zero violations against the real tree post-#4568 fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4568): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

The emitted-attribution gate also flags 4 files growing: code-review-fix.md
(+21 bytes), code-review.md (+21 bytes), gsd-code-fixer.compact.md (+9
bytes), gsd-code-fixer.md (+6 bytes). The growth is the fix itself: each
site's validation regex widened from a bounded single-optional-dotted-segment
shape to the unbounded form, and the accompanying comment/error-message text
grew by a few characters to mention the new 3-segment example.

Emitted-Drift-Ack-Growth: code-review-fix.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: code-review.md — widens the phase-number validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the error text (#4568)
Emitted-Drift-Ack-Growth: gsd-code-fixer.md — widens the padded_phase validation regex from a bounded single-dotted-segment shape to accept N-segment ids, and adds a 3-segment example to the comment/error text (#4568)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4568): backfill changeset pr number to 4646

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 17:04:28 -04:00
Tom Boucher
37b965c0d1 enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code

ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills
(src/init.cts) now selects between a canonical agents/<name>.md and a
token-minimized agents/<name>.compact.md sibling based on
workflow.compact_content, resolved in code (a real function call with a real
exit code) rather than a prose config-get gate — the same precedent stream 1's
spine/detail split established for a load-bearing seam, applied here because
this seam already runs through TypeScript instead of an eager @-include.

A missing compact sibling falls back to the canonical persona and discloses
the fallback in the served payload itself (a leading HTML-comment provenance
line), so the Done-when contract — compact when on, canonical when off, never
silent or empty — holds even for an agent nobody has compacted yet.

Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md),
each an independent, complete rewrite (not an extraction — nothing is "moved"
the way spine/detail moves text) that preserves frontmatter, every @-include,
every output-format contract, and every guardrail verbatim while cutting
restatement and verbose framing. Verified mechanically: every pair registers
(a canonical sibling exists), every compact file is strictly smaller, and the
full @-include set matches canonical's — including which references are
standalone eager-load lines versus inline prose mentions, since demoting one
to inline changes what the host actually substitutes.

Traced the install path before writing any code (.gsd/phase/.../40-design.md):
stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no
stem filtering under the default full profile, so the new .compact.md files
install for free with zero installer changes — matching issue #4407's stated
scope. A tiered agent profile that doesn't stage a compact sibling degrades
through the same fallback-with-provenance path already required for an
unauthored one, so no installer change is needed there either.

Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export
(deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are
reached by a generic code construction rather than a literal path in prose,
and checkReachability's markdown-search shape has nothing to find there).
Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new
"#4407 compact payload selection" describe block spawns gsd_run agent-skills
against real compact/canonical fixture pairs and asserts on the served
payload, which can only pass if the seam genuinely wires through.

Fixed a pre-existing test whose agents/*.md glob incidentally matched the new
.compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift
guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's
roster, both real, unrelated-to-content defects the new files' mere existence
surfaced.

Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files
under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token
benchmark baseline (npm run benchmark:compact-content-variants --write).

Closes #4407.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): apply orthogonal review findings from the compact-payload seam

Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath)
to collapse the duplicated read-and-empty-check shape between the compact
and canonical branches in cmdAgentSkills, and updated the adjacent comment
enumerating flat JSON extras to name agent_payload_variant alongside
source/degraded (added by the prior commit, comment left stale).

Security review and the Spec axis found no defects requiring a code change;
their non-blocking observations (a pre-existing, unmodified path-construction
pattern; the reasoned, documented substitution of a behavioral test for the
literal reachability check) are recorded in
.gsd/phase/enhance-4407-agent-skill-seam/60-review.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents

Root-caused via a real gsd-test run (93 failures) rather than guessing which
tests glob agents/ naively. Two classes of defect, both genuine:

1. Identity-roster confusion (11 files/areas): many tests and one production
   script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f
   => f.endsWith('.md'))`, which incidentally matched the new .compact.md
   variant siblings too — a compact file is a rendering of an EXISTING agent
   identity, not a new one. Fixed at the shared root
   (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests
   already consolidated on) and at each independent glob that didn't use it:
   agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix
   before checking XL/LARGE membership, so a compact file inherits its
   canonical sibling's tier instead of silently falling through to DEFAULT),
   agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual
   script, not just its test), codex-config.test.cjs (confirmed directly
   against generateCodexAgentToml that a compact role's derived sandbox_mode
   is byte-identical to its canonical sibling's before excluding it — not
   assumed), and copilot-install.test.cjs (two counts that legitimately DO
   need both files — an installed-file count and a full-conversion smoke test
   — fixed to expect 70, not stay pinned to 35).

   no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix:
   two compact files reproduce descriptive prose already allowlisted at their
   canonical file's line number; added matching entries at the compact files'
   own line numbers rather than excluding them from the scan (a genuine bare
   gsd-tools command-position bug in a compact file would be as real a defect
   as in canonical).

2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree
   run): six agents' compact renditions (gsd-debugger, gsd-executor,
   gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed
   the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction —
   confirmed structural, not a compaction-quality gap: each is dominated by
   content this phase's own rules require verbatim (the ~2.6 KB gsd_run
   bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in
   every agent that calls gsd_run, output-format contracts, guardrails).
   ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing
   spot in cmdAgentSkills's single-file synchronous read. Removed these 6
   compact files rather than ship an over-cap file or invent a multi-part
   read mechanism out of scope for this phase; recorded by name with the
   reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and
   50-test-matrix.md, per #4407's own "or explicitly recorded as not worth
   covering" allowance. Their canonical personas are served correctly today
   via the fallback-with-disclosed-provenance path this phase's own Done-when
   #2 already requires — 29 of 35 agents now have a compact variant.

Also fixes an unrelated, genuinely pre-existing defect this gsd-test run
surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own
DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in
this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md
| wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed:
two meaning-preserving trims in the <success_criteria> block (a repeated
parenthetical replaced with a same-exception reference; one redundant
qualifier dropped) bring it to 40,940 bytes.

Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant
benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6
now-orphaned roster rows removed alongside them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): make .compact.md-aware roster checks resilient to partial coverage

Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent
has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception),
breaking once 6 stems legitimately have none.

- tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was
  picking up the "### Compact Payload Variants" subsection's rows as
  phantom/uncounted entries in the primary/advanced/inventory-only
  classification this test validates — a compact row documents an existing
  agent's alternate rendition and never gets its own AGENTS.md heading, so it
  was never meant to participate in that classification. Excluded at the
  parser, not per-assertion.
- tests/copilot-install.test.cjs: the derived expected-file-list generator
  assumed every listAgentFiles() stem has a .compact.md source sibling;
  checks disk per stem now instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4407): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 12:38:59 -04:00
Dennis Alexis Valin Dittrich
18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00
Tom Boucher
c3e2da153b fix(#4134): refuse punctuation-only milestone heading names (#4358)
* test(#4134): fail-first regression — refuse punctuation-fragment milestone names

A first-milestone ROADMAP.md H1 that puts the version after the name
(# Roadmap: Project — Name (v1.13)) leaves exactly ')' after the heading's
own version token, which the ADR-3180 §7.2 pinned name rule returns as a
COMPLETE-scope milestone name. Failing-first coverage:

- getMilestoneInfo: name-then-version H1 (STATE-anchored + ROADMAP-only
  fallback) must yield TRUNCATED {version, name: null}, never ')'
- the refusal is level-agnostic (H2/H3)
- punctuation-family remainders (')', '()', '**', '.,;:', ']}', emoji-only)
- listMilestoneHeadings enumerates the heading with name: null
- init manager CLI reports milestone_name: null and no lone ')' anywhere
- property (seed 20260905, 300 runs): a word-char remainder is always a
  name, a punctuation-only remainder never is
- negative space: canonical delimiter forms, parenthetical names (#3171),
  trailing markers, digit-only names, CRLF headings, version-last-no-parens
  control

* fix(#4134): refuse punctuation-only milestone heading names

extractMilestoneHeadingName returns everything after the heading's own
version token as the name (ADR-3180 §7.2 pinned rule), which assumes
version-then-name. A name-then-version heading — the H1 a first-ever
ROADMAP.md drifts into ('# Roadmap: Project — Name (v1.13)') — leaves
exactly ')' after the token, and that fragment was returned as a
COMPLETE-scope milestone name, propagating into init.* JSON output and
buildStateFrontmatter's STATE.md writes.

A remainder with no letter or digit anywhere (any script) is heading
structure, not a curated name: refuse it as name: null so callers report
the honest §7.2 rule-6 answer (version kept, TRUNCATED scope). Names
that merely contain punctuation are unaffected — '(' stays an ordinary
name character (#3171) — and digit-only names qualify.

Also closes the template gap that lets the shape occur: the roadmapper
agent's output_formats now templates the version-free canonical H1
('# Roadmap: [Project Name]', per templates/roadmap.md) instead of
leaving a first milestone's title line to invention. The new section
shifts the file's existing bare-gsd-tools prose mention from line 647
to 660, so its line-keyed PROSE_ALLOWLIST entry moves with it.

Emitted-Drift-Ack-Growth: gsd-roadmapper.md — deliberate +498 bytes: new '### 0. Top-Level Title (H1)' output_formats section templating the canonical version-free H1, closing the first-milestone template gap that lets an H1 drift into 'Name (vX.Y)' and corrupt milestone_name extraction (#4134)

* chore(#4134): add changeset

* chore(#4134): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-05 23:31:35 -04:00
Tom Boucher
1db726ebbf feat(#3806): canonize the Review Dispositions Ledger contract (#4345)
* test(#3806): add parity tests for the Review Dispositions Ledger contract

Failing-first: asserts references/planner-reviews.md, workflows/plan-phase.md,
and agents/gsd-plan-checker.md agree on a single canonical "Review Dispositions
Ledger" heading, its round-scoping, L##@{sha} anchor format, and append-only
supersession rule. These fail until the canon and its two references are added.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(#3806): canonize the Review Dispositions Ledger contract

Promote the existing planner-reviews.md Step 4 return-payload tables
(Review Feedback Addressed/Deferred) into a canonical `## Review
Dispositions Ledger` PLAN.md section, stated once in planner-reviews.md
and referenced (not restated) from plan-phase.md's
<review_incorporation_contract> and gsd-plan-checker.md's Review
Incorporation dimension. Adds round-scoping (`### Round {N} —
{REVIEWS_sha}`), a `L##@{sha}` line-anchor format so a REVIEWS.md
reference survives the file being rewritten each round, and an
append-only supersession rule. Scoped to part 1 only per the
maintainer's approved-feature verdict — the deterministic lint/check
verb (part 2) is explicitly deferred to a follow-up.

Also: ADR-3806 recording the decision, a docs/features/ fragment
(FEATURES.md is generated), and a changeset fragment.

Closes #3806

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fenced-example count bug and lint findings from review

- tests/plan-review-convergence.test.cjs: the "heading exactly once"
  test counted the canonical heading text globally, so it also matched
  the illustrative fenced-code example in planner-reviews.md that shows
  the same heading as sample content, always failing 2 !== 1. Rewritten
  as a bounded line scanner that skips fenced blocks (found by an
  isolated adversarial review pass). Also bounded an unbounded regex
  quantifier over readFileSync content flagged by
  local/no-unbounded-quantifier.
- docs/features/review-dispositions-ledger.md: match house fragment
  style (bold-lead paragraphs, not #### headings) per the Standards-axis
  review; regenerated docs/FEATURES.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): fit reference-cite fix within size hard caps; ack growth

Trims the plan-phase.md / gsd-plan-checker.md reference-cite text to a
single short clause pointing at gsd-core/references/planner-reviews.md
(also fixes the bare `references/planner-reviews.md` cite the #3576
shipped-reference-cites gate rejects), bringing both files back under
their SIZE hard caps and the plan-phase.md phase6 shrink-only baseline.
Both files still grow slightly versus origin/next, acknowledged below
per ADR-2719's emitted-drift-ack contract.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): correct malformed Emitted-Drift-Ack-Growth trailer block

The previous commit's two Emitted-Drift-Ack-Growth trailers were
separated from the Co-Authored-By trailer by a blank line, so git's
own trailer parser (which tests/helpers/emitted-runtime.cjs reads via
`%(trailers:key=...)`) only recognized the last contiguous block
(Co-Authored-By) and treated the Ack-Growth lines as ordinary body
text — invisible to the emitted-attribution gate, not malformed data.
Restating them here as one contiguous trailer block, git log over the
PR range aggregates trailers from every commit, so this is additive.
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3806): isolate the ack-trailer paragraph as its own trailer block

Git's trailer parser requires the trailer paragraph to be the message's
final paragraph, preceded by a blank line, and to contain nothing but
trailer-shaped lines. The prior commit's blank line before the trailer
lines was missing, which folded the leading Emitted-Drift-Ack-Growth
lines into an ordinary prose paragraph.

Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap.
Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3806): backfill PR #4345 into changeset and ADR

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:50:27 -04:00
Tom Boucher
c20675cc4d fix(#3819): widen executor's pre-commit guard beyond worktree mode (#4343)
* fix(#3819): widen executor's pre-commit guard beyond worktree mode

The pre-commit protected-branch assertion in the executor agent (#2924)
only fired inside a Claude Code worktree and matched a hardcoded
five-name branch list. It never ran in an ordinary checkout and never
covered this repo's own default branch ("next"), so gsd-executor could
commit planning-repo documents directly onto a shared checkout's
default branch with no PR ever created.

Widen the guard to run in every isolation mode, and resolve the
protected branch via the repository's actual default branch (with the
existing five-name list retained as a fallback when the resolver
itself cannot be invoked) plus any configured git.protected_branches.
Add a git.allow_default_branch_commits escape hatch for projects that
intentionally execute on their default branch. Also point the
separate <final_commit> commit helper back at the same guard, so it
cannot be sidestepped by that path.

Emitted-Drift-Ack-Growth: gsd-executor.md — widened pre-commit protected-branch guard (#3819); tightened comments to stay under the size cap.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3819): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 18:35:42 -04:00
Dennis Alexis Valin Dittrich
1017898cb9 fix(#3771): make remediation examples non-binding and surface revision conflicts (#3916)
* fix(#3771): separate the binding property from the advisory remediation

Checker findings fused "what property failed" with "how to fix it" into a
single `fix_hint` and never marked which half binds. The checker rendered
every hint under a "must fix" heading, the orchestrators injected the issues
verbatim and ordered targeted updates, and the shared revision references
mapped each hint to a prescriptive strategy — so a contract-following planner
applied a hint literally even when a smaller mechanism satisfied the same
property, or when the hint contradicted a locked decision. There was no
channel to report that conflict, and every attempt burned a revision
iteration.

Checker side: every issue now carries a binding `required_property` (the
invariant that failed) plus its evidence and severity, and `fix_hint` is
labelled non-binding wherever it appears — including the human-facing blocker
rendering, so "must fix" unambiguously names the property and never the
example.

Planner side: revision re-checks locked decisions, capability guidance and
existing plan constraints before editing; satisfying a blocker through a
smaller valid alternative counts as addressing it; and a hint that conflicts
with any of those returns `REVISION_CONFLICT` carrying the conflict and the
alternatives considered. Orchestrators route that to user choice or the
configured plan-review convergence loop without consuming retry budget.

Also applied to the UI-spec revision loop and the gap-plan hint, and the
generic pattern's stray `suggested_fix` field name is reconciled to the
plan-checker's `fix_hint`.

Nothing legitimately binding is weakened: blockers still block, severity
still gates, iteration caps and stall escalation still fire, and required
task fields and decision coverage still hold.

Refs #3771

* test(#3771): pin the binding/advisory split across the revision chain

Locks the separation at every link that carries it: the checker's issue
schema and blocker rendering, the planner's constraint re-check and
REVISION_CONFLICT return, the generic pattern's reconciled field names, and
each orchestrator's conflict routing without retry-budget consumption. Also
pins what must not have been weakened — blockers, severity gating, iteration
caps and stall escalation.

Red against the pre-fix prose: 32 of 34 assertions fail (the 2 that pass are
the preservation checks, correctly).

Refs #3771

* chore(#3771): add changeset fragment for the remediation-binding fix

* chore(#3771): acknowledge the remediation-binding growth

Five runtime-loaded files grow: the two checkers carry the binding/advisory
split where the model reads it (a `required_property` on every dimension
example, since a schema the examples contradict teaches the examples), and
the three orchestrators carry the REVISION_CONFLICT route, which has to live
with the `iteration_count`/`revision_count` state it declines to spend.

Deletes tests/emitted-drift-acks/3172-stated-failing-direction.json: it is
fully spent on next and still owned plan-phase.md, so it walls off a key it
can no longer clear (#3078). Its removal is the documented remedy for the
duplicate-key collision, not drive-by cleanup.

* fix(#3771): close the review gaps in the conflict contract

Adversarial review (Codex, Antigravity) found four real defects in the first
pass, each confirmed against the source before acting:

- The UI checker's structured return still ordered `Fix: {exact fix required}`
  and "list each BLOCK dimension with exact fix required". The dimension
  examples had been marked non-binding but the rendering the researcher
  actually reads had not — the same omission this issue is about.
- `ui-phase` and the canonical `revision-loop` flow incremented their counter
  BEFORE dispatching the reviser, so "do NOT increment on REVISION_CONFLICT"
  was unreachable prose: the iteration was already spent. The increment now
  sits on the return path in both.
- The conflict gate offered "accept as-is", which is an early exit from a
  still-failing blocker — a weakening the brief explicitly forbids. The three
  options are now adopt an alternative / override the constraint / amend the
  constraint; every one resolves the conflict. Accepting an unaddressed blocker
  remains available only at the unchanged iteration-cap escalation.
- The convergence route was declarative: nothing in
  plan-review-convergence.md could receive a conflict. plan-phase now records
  it in REVIEWS.md — the channel that loop already consumes — convergence
  refuses to declare convergence over an open entry, and routing back into a
  run convergence itself started is explicitly excluded as a cycle. `quick` has
  no REVIEWS.md and no phase, so its convergence branch was dead prose and is
  deleted in favour of asking the user.

Also reconciles the last two drifted field names (`finding`, `affected_field`)
to the plan-checker schema, and repairs a silent no-op: the few-shot
`required_property` insertion never applied because those lines are
blockquoted, and the test's own block filter was anchored on indentation only,
so a vacuous loop passed over zero blocks. Both are fixed and the filter now
asserts it found blocks.

Refs #3771

* chore(#3771): extend the growth acknowledgment for the review round

plan-review-convergence.md joins the list: the conflict route needed a
receiving end, and it lands on the seam that loop already reads (REVIEWS.md)
rather than a new mechanism. The plan-phase, ui-phase and gsd-ui-checker
entries gain the second-pass reasoning — an executable convergence branch, the
increment moved onto the return path, and the structured return that still
ordered an exact fix.

* fix(#3771): make the conflict route bounded, ordered, and owned

Round-2 adversarial review found five more defects, each confirmed in the
source before acting:

- The convergence gate sat AFTER `gsd_run state planned-phase` and the success
  banner, so a run could write and announce convergence over an unresolved
  conflict. OPEN_CONFLICTS is now read from REVIEWS.md and is part of the
  converged CONDITION, evaluated before any write.
- plan-phase's cycle-exclusion ("unless this run was invoked by convergence")
  was not a question the orchestrator can answer at runtime. plan-phase now
  never invokes convergence at all — it records the conflict when a phase
  REVIEWS.md exists and resolves it with the user in-place, which removes the
  cycle instead of describing it.
- Closure had no owner. plan-phase writes the row, so plan-phase strikes it
  resolved; convergence only reads. An open row is a live blocker, never a
  stale artifact.
- Declining to increment the counter removed the only bound on the conflict
  path: an agent returning the same conflict forever would loop unattended. A
  conflict naming the same `required_property` twice in a row is now a stall
  and escalates through the existing gate.
- verify-work's gap-plan revision loop hands `<revision_context>` to
  gsd-planner and so inherits the whole contract, but stated none of it and
  could not handle the conflict return. It is now covered like the others, and
  is in the test's orchestrator table.

Refs #3771

* chore(#3771): acknowledge the round-2 growth

verify-work.md joins the list — the flow the second review found missed — and
the plan-phase, ui-phase and plan-review-convergence entries gain the
round-2 reasoning: the gate moved ahead of the state write, the convergence
hand-off replaced with a runtime-checkable record-and-resolve, and the
recurrence bound that replaces the counter the conflict path stopped spending.

* fix(#3771): make the convergence gate countable and stop the conflict fall-through

Third adversarial round (Antigravity) found three defects:

- The OPEN_CONFLICTS pipeline had no `grep -v '~~'` despite its own comment
  claiming one, and `grep -c '^| '` also counts a markdown table's header and
  separator rows — every resolved conflict would have read as open and
  convergence would have deadlocked instead of converging. plan-phase now
  records each conflict as a `- [ ]` checklist line and flips it to `- [x]`, so
  the gate is an exact fixed-string match with no table parsing.
- "then continue below" fell through to the checker spawn, so a SECOND
  REVISION_CONFLICT would have been handed to the checker as though it were a
  revised plan. plan-phase, quick and verify-work now re-evaluate the return
  from the top of the conflict handler; ui-phase already looped back.
- revision-loop.md still described plan-phase routing a conflict to the
  convergence loop instead of asking — the behaviour round 2 removed. Recording
  is now stated as being in addition to asking, never instead of it.

Refs #3771

* chore(#3771): bring the changeset in line with what shipped

Two review rounds widened the change after the fragment was written:
verify-work's gap-plan revision and the convergence loop are covered, two
more drifted field names are reconciled, and the conflict path carries an
explicit recurrence bound.

* fix(#3771): declare and emit the REVISION_CONFLICT marker

check:contract-drift on CI caught what local lint never reached: four
workflows dispatch on `## REVISION_CONFLICT`, but no agent declared or emitted
it — an orphan consumer, matching a marker nothing produces. The shared
reference (planner-revision.md Step 7b) described the return; the agent
definitions did not carry it.

gsd-planner and gsd-ui-researcher now emit the marker in-fence alongside their
other return markers, and both registry rows in agent-contracts.md declare it.
gsd-planner's Consumed by gains the two workflows that dispatch on it and were
missing from the row.

The gate is right: a return contract belongs where the agent is defined, not
only in a reference the agent happens to load.

Refs #3771

* chore(#3771): acknowledge the return-marker growth

gsd-planner.md and gsd-ui-researcher.md each gain the REVISION_CONFLICT
marker that check:contract-drift requires them to emit.

* fix(#3771): hoist the shared conflict protocol out of the workflows

Two CI failures, both correct gates:

- tests/few-shot-calibration.test.cjs pins the plan-checker calibration file
  at exactly 4 examples (2 positive, 2 negative). The example added in the
  first pass broke that balance — and described PLANNER behaviour in the
  CHECKER's calibration set, which is the wrong surface for it. Removed; the
  smaller-alternative rule is already normative in gsd-plan-checker.md and
  planner-revision.md, and pinned by the regression suite.
- tests/phase6-capstone-conformance.test.cjs (ADR-857 phase 6, #1168) requires
  plan-phase.md to stay BELOW its pre-phase-6 baseline of 94519 bytes. The
  inline conflict block pushed it to 94988.

The fix for the second is the one that should have been made first: the
record/resolve/close protocol and the recurrence bound were identical in four
workflows, and revision-loop.md — which plan-phase already @-imports — is what
a shared contract is for. The protocol now lives there once; plan-phase states
only its bindings (which counter, which artifact, which next step) and points
at it. plan-phase.md: 94988 -> 92739, under the ratchet with headroom, and the
four-way duplication is gone.

quick, ui-phase and verify-work do not import the reference, so they keep their
inline statements. The suite asserts each rule against what the runtime
actually loads for that orchestrator, not against the file in isolation.

Refs #3771

* docs(#3771): state the shared-protocol relationship accurately

Three of the four revision-bearing workflows do not @-import revision-loop.md,
so 'follows it verbatim' overstated the coupling. Only plan-phase defers; the
others restate the rules inline and this section is the authority they must
agree with.

* refactor(#3771): name the authority instead of restating it four times

Self-review finding: the same ~700-byte reviser paragraph was inlined in four
prompts while gsd-planner already loads planner-revision.md whenever
<revision_context> is present (agents/gsd-planner.md:555, :588) — a fifth copy
of the same contract. Four places to edit in lockstep is precisely the drift
class this PR exists to fix; the generic pattern calling the field
suggested_fix while the checker emitted fix_hint is what that looks like after
a year.

Each prompt now carries only the load-bearing clauses and names the authority
it summarises. Deliberately NOT reduced to a bare pointer: these are LLM
prompts, and a contract stated only in a file the reader is supposed to fetch
is the failure mode of this very bug. The saving is modest (~50 bytes each) —
the point is the named source of truth, not the bytes.

Suite: dropped seven assertions that pinned heading text and bold-lead
phrasing a reword would break without changing what the runtime is told; their
neighbours already pin the same contract by content. 519 -> 499 lines, 61
tests. Red gate against origin/next: 56 of 61 fail.

Refs #3771

* fix(#3771): sanitize agent-authored conflict text and bound total conflicts

Cross-AI review (agy/Gemini) raised two security findings. I confirmed the
first by running the gate's own awk against a crafted file and reproduced it
exactly; the second is a real hole in my bound.

1. FAIL-OPEN in the convergence gate. The conflict text is agent-authored and
   is appended to REVIEWS.md, which the gate scans with an awk that stops at
   the next '## ' heading. One line of agent text beginning '## ' ends that
   scan early, so conflicts below it are never counted and convergence declares
   success over a live blocker. Measured: 3 open conflicts, awk returned 2.

   Fixed at the write boundary, which is the trust boundary: every field has
   newlines and tabs collapsed to spaces and a leading '#', '-', '|' or fence
   stripped, so one conflict is exactly one line. Both producing agents now
   declare their fields single-line plain text, and the reader states the
   invariant it depends on so a later edit cannot silently break it. Verified:
   3 open + 1 resolved now counts 3; missing file and absent section count 0.

2. The recurrence bound was 'same required_property twice in a row', which an
   agent alternating property names never trips, leaving the un-incremented
   conflict path unbounded. Now bounded twice: the repeat rule catches the
   common case, and the THIRD conflict return of a loop escalates whatever
   property it names. A conflict still never consumes a revision iteration;
   this cap is separate from and additional to the revision cap.

Rejected from the same review: deleting 'a planner that reaches
required_property by a smaller or different mechanism has addressed the issue
in full' from the CHECKER prompt as misplaced. It is load-bearing exactly
there. A checker that does not know a different mechanism counts will re-flag
the issue on re-check, which is the revision loop that never terminates. The
argument offered for deleting it, that the checker evaluates the new state
independently, describes the failure mode.

Refs #3771

* fix(#3771): fail closed on an unverifiable convergence gate

Second cross-AI pass (agy, this time with the full files rather than the diff)
found two more, both real:

1. The gate read REVIEWS_FILE with `2>/dev/null || echo 0`, so an unreadable or
   empty path counted as ZERO open conflicts and converged. That path is
   resolved a few lines earlier by a pre-existing unquoted
   `ls ${phase_dir}/${padded_phase}-REVIEWS.md` (line 346, not touched by this
   PR), which yields an empty string rather than an error when the path
   contains a space. Unverifiable is not clean: the gate now tests -z and -r
   first and BLOCKS. Verified both branches.

   The unquoted ls itself is left alone deliberately — it predates this change
   and belongs to the reviews lookup, not the conflict gate. Fixing it at my
   own boundary removes its effect on this gate without widening scope.

2. REVIEWS.md is writable by the review agent, which could flip a `- [ ]` to
   `- [x]` or delete the section and forge the state of a blocking gate. The
   section now declares a single writer: /gsd:plan-phase appends and closes,
   every other agent leaves it byte-for-byte alone, readers read.

Also trimmed a clause that explained the increment ordering by reference to
what the file said before this PR. Commit history is not instruction, and
these files are prompts.

Rejected: the claim that quick's conflict gate deadlocks autonomous pipelines
by asking the user. Its existing max-iteration escalation in the same file
already asks the user the same way; this adds no new interaction class.
Noted but out of scope: the per-dimension YAML example blocks and the shim
boilerplate duplicated across agent prompts both predate this change.

Refs #3771

* fix(#3771): count conflicts by line shape, not by section

CodeRabbit review on the rehearsal PR. Five findings, all valid, all applied.

The best one is a deletion. The convergence gate scanned between
'## Plan-Revision Conflicts' and the next '## ' heading, and that scan stops at
the FIRST heading it meets — so one stray '## ' line hid every conflict beneath
it and returned 0, converging over a live blocker. Reproduced: section-scan 0,
shape-scan 1. Sanitizing at the write boundary does not cover a hand-edited,
legacy, or foreign-written REVIEWS.md, so the reader needed its own guarantee.

It now matches the conflict line SHAPE anywhere in the file:

  grep -c '^- \[ \] .*required_property:'

No section bookkeeping, nothing a heading can truncate, and it composes with the
writer's sanitization (which strips a leading '-' from agent text, so agent prose
cannot forge the shape). Verified: injected heading -> 1, all resolved -> 0.

The other four:

- Both checkers told the author never to emit a contradictory fix_hint, then
  offered an escape hatch that put the forbidden route in the hint anyway. They
  now name NO route in that case and state only that the property conflicts with
  the constraint. A hint carrying a forbidden route is applied by anyone who
  trusts hints.
- The REVISION_CONFLICT marker description in gsd-planner.md was narrower than
  planner-revision.md: it covered a contradictory hint but not an unreachable
  required_property. A planner reading only the agent file would have burned
  retry budget on the case the reference routes to a conflict.
- The few-shot calibration examples used uppercase BLOCKER/INFO while the schema
  defines blocker/warning/info. Pre-existing, but it is the same schema-vs-example
  disagreement this PR exists to end, and the file was already being edited.
- verify-work's re-entry instruction existed but sat after the Bounded clause, so
  the paragraph read "re-spawn ... stop re-spawning ... after re-spawning". The
  re-entry now immediately follows the re-spawn, and states that only a
  non-conflict return may reach the checker or increment iteration_count.

Refs #3771

* fix(#3771): resolve the contradictory scope_sanity severity examples

Sixth CodeRabbit finding, posted outside the diff range and missed on my first
read — I had claimed all findings were addressed after reading only the five
inline comments. This one was in the review body.

agents/gsd-plan-checker.md carried TWO scope_sanity examples with identical
metrics (5 tasks, 12 files) and OPPOSITE severities: warning in Dimension 5,
blocker in <examples>. Line 872 states "2-3 tasks/plan good, 4 warning, 5+
blocker" and the severity table lists warning as "Scope 4 tasks (borderline)",
so the warning example contradicted both.

ADR-2629 Decision 5's "over budget is a WARNING, never a blocker" does not
excuse it: that rule governs the smart-zone TOKEN estimate (the estimate-check
verb, lines 299-306), which is a different axis from task count. Verified in
source before touching it.

The contradiction is pre-existing but this PR made it binding and visible:
severity is now declared part of the binding payload, and both examples were
given the same required_property, so they now disagree on the severity of an
identical finding about an identical property.

Deviating from the proposed correction, which was warning -> blocker: that
would duplicate the <examples> entry outright (same tasks, files, severity).
The Dimension 5 example is instead made a genuine 4-task borderline warning, so
the file keeps one worked example per severity and the thresholds, the severity
table and both examples finally agree.

Refs #3771

* fix(#3771): stop laundering a grep error into zero open conflicts

Seventh CodeRabbit finding — from a SECOND review round my own CR-4 push
triggered, which I had not looked for. This one is a regression I introduced
while fixing the previous fail-open.

CR-4 replaced the truncatable section scan with:

  OPEN_CONFLICTS=$(grep -c '^- \[ \] .*required_property:' "$REVIEWS_FILE" || true)

`|| true` masks every grep failure. grep exits 1 for "no matches" (a legitimate
zero) but 2 for a read error, and `|| true` turns both into an empty capture
that `${OPEN_CONFLICTS:-0}` renders as 0. If REVIEWS.md is removed or becomes
unreadable between the -r check and the scan, the gate reports no conflicts and
convergence proceeds. Proven: unreadable file -> captured empty -> 0.

The status is now inspected, and only exit 1 counts as zero; anything else
blocks.

My first attempt at this fix was itself wrong and my own harness caught it: I
wrote `if ! grep ...; then grep_status=$?`, but `!` inverts the status, so `$?`
in that branch is 0 and every failure reads as success — the clean-file case
printed "BLOCKED (grep exit 0)". The status must be read in the ELSE branch of a
non-negated `if`, which is what CodeRabbit proposed. Both traps are now pinned
by tests.

Verified end to end: all resolved -> 0, no conflicts at all -> 0, injected
heading -> 1, unreadable file -> BLOCKED with grep exit 2.

Refs #3771

* test(#3771): execute the conflict gate instead of reading it

CodeRabbit round three: 0 actionable, 1 nitpick — "these assertions inspect
Markdown source only; they do not prove that grep status 1 produces zero
conflicts or that a scan error exits before convergence." Rated Trivial. It is
the most valuable finding of the three rounds.

This gate has been wrong three times: a section scan a heading could truncate, a
`|| true` that laundered grep's error status into zero, and an `if !` whose `$?`
reported the negation rather than the command. Every one of those passed the
text assertions that existed at the time. I proved each fix by hand in a shell,
and none of that proof lived in the suite.

The gate is one self-contained fenced block, so the test now extracts it from
the workflow — located by content, not line number — writes it to a script and
RUNS it against fixtures: two open plus one resolved counts 2; no matches counts
0 and does not fail; a conflict below an injected `## ` heading still counts; an
unreadable path and an empty path both BLOCK with a non-zero status and no zero
count on stdout.

Non-vacuity proven by mutation rather than asserted. Reverting the gate to each
of its three historical broken forms reds the suite:

  section-scan awk  -> 7 failures (5 in the gate cases)
  || true           -> 4 failures (3 in the gate cases)
  if ! (negated $?) -> 4 failures (3 in the gate cases)
  restored          -> 69 pass, 0 fail

The prose assertions stay: they are the right instrument for a prompt. This
covers the one part of the change that is real shell an orchestrator executes.

Refs #3771

* test(#3771): route the gate harness through the shared test helpers

ESLint's project rules caught three violations in the new harness: an unbounded
execFileSync (DEFECT.UNBOUNDED-SUBPROCESS — an unbounded spawn is an indefinite
hang, and on macOS CI that is how a stuck run stops reporting instead of failing)
and two raw fs.rmSync calls, which skip the Windows-EBUSY retry budget that
helpers.cleanup carries.

Now uses createTempDir/cleanup from tests/helpers.cjs and passes an explicit
30s timeout. Suppressing the rules was available and would have been the wrong
call: both exist because of real CI failure modes on platforms I am not testing
on.

* chore(#3771): backfill the changeset PR number

The pr: field is drift-checked against the PR event payload, so it cannot be
written before the PR exists. Set to 3916.

* fix(#3771): close revision conflict persistence gaps

Use the authoritative review artifact, keep conflict and normal retry paths disjoint, and enforce one writer-reader grammar so malformed state fails closed.

Emitted-Drift-Ack-Growth: diagnose-issues.md — #3771 marks the gap-plan remediation hint non-binding while keeping root_cause authoritative
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — #3771 separates binding required_property evidence from advisory fix_hint examples across the checker contract
Emitted-Drift-Ack-Growth: gsd-planner.md — #3771 declares the REVISION_CONFLICT return used when remediation contradicts governing constraints
Emitted-Drift-Ack-Growth: gsd-ui-checker.md — #3771 applies the same binding-property and advisory-hint split to UI review findings
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — #3771 defines the UI revision producer's structured REVISION_CONFLICT return
Emitted-Drift-Ack-Growth: plan-phase.md — #3771 routes and persists bounded revision conflicts before spending the normal retry budget
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3771 adds the fail-closed owned-block parser and prevents convergence over open conflicts
Emitted-Drift-Ack-Growth: review.md — #3771 emits and preserves the canonical writer-owned conflict block across review regeneration
Emitted-Drift-Ack-Growth: ui-phase.md — #3771 routes UI revision conflicts to resolution before consuming revision_count
Emitted-Drift-Ack-Growth: verify-work.md — #3771 gives gap-plan revision the same bounded conflict route before iteration_count

* test(#3916): guard rebases against schema drift

Load the current-base progressive-disclosure examples so every integrated issue remains bound by required_property after branch reconciliation.

* test(#3916): skip the extracted-gate suite's bash spawns on win32

Third review round's sole survivor: runConflictGate()/withReviews() spawn
bash against a Node-native temp path built by createTempDir(), which is
backslash-separated on the Windows CI lane and not a path Git Bash is
guaranteed to accept (DEFECT.WINDOWS-TEST-PORTABILITY, matching the
observed CI failure at revision-remediation-binding.test.cjs:844,
ENOENT on a path Windows read as a directory separator). No eslint rule
catches it since the call has neither a chmod nor a `bash -c` form.

Guards the four call sites with the repo's existing skipOnWin32
convention (describe/test `{ skip: IS_WINDOWS }`) rather than
normalizing the harness path to forward slashes, which would defeat the
one test whose purpose is proving the production gate does NOT rewrite
a literal backslash in a POSIX filename.

* fix(#3916): backfill changeset pr field to the fork validation PR number

* fix(#3771): forbid silently accepting an open plan-revision conflict at max-cycles escalation

The max-cycles escalation prompt only surfaced HIGH_COUNT and ACTIONABLE_COUNT; an open
plan-revision conflict (OPEN_CONFLICTS > 0) was never disclosed there, and "Proceed anyway"
could exit successfully over it — exactly the failure mode this PR exists to close (a success
banner over an unresolved conflict nobody resolved). Blockers still block: withhold "Proceed
anyway" and route to Manual review whenever a conflict is open.

* fix(#3771): do not hard-block REVISION_CONFLICT persistence when no REVIEWS.md exists yet

A phase's first-ever revision cycle can return REVISION_CONFLICT before any REVIEWS.md has been
written — REVIEWS_PATH is then legitimately empty, not a corrupt or deleted file. The persistence
gate's own accompanying prose already says the record channel applies 'when REVIEWS_FILE is
non-empty', but the bash condition never checked that, so it hard-blocked every conflict on a
brand-new phase regardless of whether persistence was even expected to run. Require a non-empty
REVIEWS_FILE before treating a missing file as an error.

* chore(#3771): raise the plan-phase.md ADR-857 host-loop ceiling to 96700

The frozen pre-phase-6 ceiling (94519) collided on rebase: this PR's own
REVISION_CONFLICT persistence/routing gate is core planner control flow, not an
un-extracted optional feature, and landed alongside an unrelated, already-merged
same-file growth (the #4.6 context-drift pre-check) already on next. Same
rationale #1298 already established for execute-phase.md's ceiling.

* chore(#3916): backfill changeset pr field to the upstream PR number

* fix(#3771): make the writer-side REVISION_CONFLICT sanitize step real shell

The Conflict Return record channel sanitized agent-authored fields via a
prose instruction ("Sanitize each agent-authored field before appending")
for the orchestrator LLM to apply by hand, while the reader-side gate in
plan-review-convergence.md parses the same slot with real, executed awk.
Flagged Minor across two review rounds (round 4, round 6) since no code
performed the sanitize anywhere.

plan-phase.md's Conflict Return step now runs a real bash gate: sanitize
each field (collapse newline/tab to space, strip a leading #/-/|/fence),
build the one-line record, skip the append if an identical line already
exists (idempotent), insert before the writer-owned end delimiter, and
fail closed if that delimiter is missing rather than silently dropping
the conflict.

tests/revision-remediation-binding.test.cjs extracts and RUNS the new
fence (matching how the reader gate is already tested), composing it
with the existing reader gate: hostile-field fast-check fuzzing, a
repeated-conflict idempotency check, and a missing-delimiter fail-closed
check that the file is left byte-for-byte unchanged on failure.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 turns the writer-side
REVISION_CONFLICT sanitize+insert step into real, executed shell instead
of a prose instruction, matching the reader gate's existing rigor

* fix(#3916): backfill changeset pr field to the fork validation PR number

Fork CI's changeset-lint reads the real PR number from its own event
payload; the fragment still carried the upstream number from the last
sync, so the DEFECT.CHANGESET-PR-FIELD-DRIFT check failed on this fork
PR. Re-backfill to the upstream number before the final push.

* fix(#3771): close the awk -v forgery and same-session close gaps agy found

Adversarial review (gemini-3.8-flash-high via the internal agy review
lane) on the full PR found two BLOCKERs against the just-added
writer-side conflict gate:

1. `awk -v line="${LINE}"` decodes escape sequences in its argument, so
   a literal two-character `\n` in agent-authored text became a real
   newline inside awk, splitting the appended record across two
   physical lines. `tr` only strips actual control bytes, so it never
   saw this — it defeated the exact forgery the gate exists to
   prevent, both the reader's zero-count and the writer's own
   idempotency check. Fixed by passing LINE/END through awk's
   ENVIRON, which is not escape-decoded.

2. A conflict resolved and re-spawned within the same plan-phase
   session was never flipped from `- [ ]` to `- [x]` — the record
   channel bullet said "plan-phase closes it," but no step did. Only
   a *separate* `--reviews` re-entry (line ~622, still prose-only)
   closes conflicts; the in-session resolve path left them open
   forever, permanently blocking convergence. Fixed by carrying the
   just-written line in `PENDING_CONFLICT` and closing it in the
   `Otherwise` branch before the checker re-spawns.

Also fixed a MAJOR: docs/COMMANDS.md described the `--max-cycles`
escalation gate as uniformly offering "proceed or review manually,"
but the code (this PR's own change) withholds "Proceed anyway"
specifically when a plan-revision conflict is open — only manual
review is offered in that case. Docs now say so.

Not applied: the reviewer's `\r` truncated to plain tr from a MINOR
that also asked for temp-file permission preservation across `mktemp`.
Applying `chmod --reference` is not portable to macOS/BSD `chmod`, so
this is left as a documented low-severity tradeoff — the temp file now
sits alongside REVIEWS.md (same filesystem, atomic `mv`), which was
the same finding's more substantive half. Also not applied: a
suggested `gsd_run review record-conflict` CLI subcommand to
deduplicate the two `awk` blocks — a new command plus wiring is out of
scope for a review-remediation fix.

tests/revision-remediation-binding.test.cjs adds regression coverage
for both BLOCKERs: a literal-backslash-n hostile field composed with
the reader gate, and a close-gate extraction that verifies the flip to
`[x]`, the reader's count dropping to 0, and a fail-closed path when
the pending line is missing.

Emitted-Drift-Ack-Growth: plan-phase.md — #3916 fixes an awk -v escape-
decoding forgery and adds the missing same-session conflict-close step
an adversarial review found in the writer-side gate

* chore(#3916): backfill changeset pr field to the upstream PR number

Fork-validation CI needed pr: 1 to pass its own changeset-lint; restore
pr: 3916 before this push reaches open-gsd/gsd-core.

* fix(#3771): trim plan-phase.md prose back under the XL byte cap

Merging origin/next's unrelated growth pushed plan-phase.md 473 bytes
past the workflow-size-budget XL cap and the ADR-857 phase-6 baseline,
both tripped by CI after review approval. Removed an unpinned inert
bash comment and tightened connective prose in three REVISION_CONFLICT
bullets; no executable shell or test-pinned substring changed.

* chore(rehearsal): pin changeset pr field to fork rehearsal PR #25

Scratch-only commit for the rehearsal branch's own CI. Will not be
carried onto the branch backing upstream #3916 — that keeps pr: 3916.

* fix(#3771): address CodeRabbit findings on the REVISION_CONFLICT protocol

Fork rehearsal PR #25's first CodeRabbit pass surfaced 7 findings against
the already-approved #3916 diff; each verified against current code
before fixing (none hallucinated):

- plan-phase.md: writer-side awk gates now strip a trailing \r before
  comparing lines, matching the reader gate (plan-review-convergence.md)
  -- a CRLF REVIEWS.md previously made both writer gates fail closed.
- plan-phase.md: the close-fence's REVIEWS_FILE/PENDING_CONFLICT/
  CONFLICT_RESOLUTION were read without ever being (re)defined in that
  fence -- shell state does not survive across separate fenced blocks
  (same convention already documented in review.md). Added the explicit
  recompute/set instruction.
- revision-loop.md: previous_conflict_property was never reset after a
  normal (non-conflict) revision, so a later, unrelated conflict on the
  same property could be misread as a repeat and escalate prematurely.
- gsd-plan-checker.md / few-shot-examples/plan-checker.md: two example
  required_property strings were unconditionally binding in a way their
  own dimension's rules aren't (no-analog RESEARCH.md fallback; tasks
  that create no functions), now scoped to match.
- quick/steps/plan-checker-loop.md: added the same disjoint
  "Otherwise (not REVISION_CONFLICT)" branch plan-phase.md already had,
  closing an ambiguity between the conflict and non-conflict return paths.
- revision-remediation-binding.test.cjs: the REVIEWS_PATH init-order
  assertion used indexOf() without checking for -1, so it would pass
  vacuously if either anchor were renamed away.

Also restores an "Export the row's CONFLICT_*" instruction I had cut in
the prior byte-budget trim -- checked non-pinned by tests, but it was the
only text telling the agent to set those vars before the awk block reads
them via ENVIRON.

Net growth from these fixes required reclaiming bytes elsewhere in
plan-phase.md (verified against every pinned substring in
revision-remediation-binding.test.cjs) to stay under the XL tier's
hard 98304-byte cap; final size 98245 bytes.

* fix(#3771): resync the #4079 shrink-only mirror to the current PRE_PHASE6 line

tests/plan-phase-background-wait-wakeup.test.cjs (landed on next via an
unrelated #4079 PR, merged in by this branch's next-sync) mirrored
plan-phase.md's phase6 shrink-only ceiling as a hardcoded local constant
(94519) rather than reading tests/phase6-capstone-conformance.test.cjs's
PRE_PHASE6 value. That value has since been legitimately raised twice
during this PR's own review (94519 -> 96700 -> 98300) to accommodate the
REVISION_CONFLICT persistence/routing gate. The two branches' independent
histories left the mirror stale post-merge -- not a textual git conflict,
but the same class of thing. Resynced to 98300.

* fix(#3771): address round-2 CodeRabbit findings on the conflict gates

CodeRabbit's re-review of the previous remediation commit found two real
issues in what it had already flagged:

- Both writer-side awk CRLF fixes used \`sub(/\r$/, "")\` directly on \`\$0\`,
  which mutates it in place -- \`{ print }\` then emitted the CR-stripped
  copy for every passed-through line, silently rewriting an unrelated
  CRLF REVIEWS.md to LF on any insert or close. Now compares against a
  separate \`cur\` copy and prints the original, untouched \`\$0\`.
- The close-fence's "recompute REVIEWS_FILE/PENDING_CONFLICT" prose
  implied in-fence derivation, but the fence has no such code and the
  test harness (\`runCloseGate\`) deliberately supplies all three as
  pre-set env vars -- matching how the open fence's "Export the row's
  CONFLICT_*" instruction already works. Reworded to "export ... in the
  same invocation", matching that established, test-verified pattern
  instead of promising logic that isn't there.

Added a regression test proving the CRLF fix no longer touches
passthrough lines (red against the mutate-in-place version, green now).

* fix(#3771): use a CRLF-safe check in the new passthrough regression test

local/no-crlf-fragile-split forbids splitting readFileSync content on a
literal \n (Windows git-autocrlf checkouts yield \r\n). My CRLF
passthrough-preservation test from the previous commit did exactly that
to inspect the first line. Replaced with a direct startsWith() check
against the known CRLF-terminated header, which needs no split.

* test(#3771): assert the record itself is inserted in the CRLF passthrough test

CodeRabbit nitpick (round 3): the passthrough-preservation test checked
gate status and the pre-existing line's CRLF ending, but never asserted
the new REVISION_CONFLICT record was actually written.

* fix(#3771): address agy/gemini-3.8-flash-high adversarial review findings

Full-PR adversarial review (internal /gsd-review antigravity lane,
gemini-3.8-flash-high) surfaced 9 findings; each verified against current
code before fixing (none hallucinated):

HIGH:
- quick-batch/steps/plan-checker-loop.md never received the
  required_property/fix_hint binding language or REVISION_CONFLICT
  handling this PR added everywhere else -- a genuinely unmigrated
  producing context. Migrated to match quick/steps/plan-checker-loop.md,
  and added it to the ORCHESTRATORS consistency battery in
  revision-remediation-binding.test.cjs so future drift is caught
  automatically.
- The close-fence's PENDING_CONFLICT was an agent-supplied env var that
  had to exactly reconstruct a five-field sanitized line across a
  multi-minute subagent dispatch -- fragile, and a scalar var also meant
  a second simultaneous conflict silently dropped the first on overwrite.
  Redesigned to match the open conflict by CONFLICT_DIMENSION/
  CONFLICT_PLAN identity instead: the agent re-supplies two short,
  already-tracked identifiers rather than reconstructing the full
  sanitized text, and each conflict resolves independently regardless of
  how many are open. Updated the test harness's runCloseGate contract to
  match, and added a two-open-conflicts regression test.

MEDIUM:
- plan-phase.md's `--reviews` replanning path told the reader to "flip
  the matching line to [x]" in prose only, with no executable path to
  it -- pointed it at the same close gate used in step 12.
- plan-review-convergence.md's reader-gate awk tolerated a blank line
  before the opening delimiter but not before the heading that follows
  it; a formatter or LLM writer inserting one would hard-abort
  convergence on an otherwise well-formed REVIEWS.md. Added the same
  tolerance already granted above it, with a regression test.

LOW:
- Clarified that the escalation destination for a stalled conflict is
  the same iteration/revision-count cap gate already defined in each of
  quick, quick-batch, ui-phase, and verify-work, rather than an
  undefined "stall" concept.
- Clarified "twice in a row" means no successful revision intervened,
  matching revision-loop.md's now-explicit previous_conflict_property
  reset.
- Fixed gsd-ui-researcher.md's stale rationale text, copied verbatim
  from planner-revision.md: ui-phase presents the conflict table
  directly to the user, it does not persist to a shared file scanned by
  heading.

Net growth again required reclaiming bytes in plan-phase.md (verified
against every pinned substring in revision-remediation-binding.test.cjs)
to stay under the XL tier's hard 98304-byte cap; removed a now-dead
PENDING_CONFLICT assignment in the process. Final size 98258 bytes.

* fix(#3771): scope row 48's quick/steps guard away from plan-checker-loop.md

tests/gsd-quick-batch-quick-regression.test.cjs's row 48 (#3676) flagged
this branch's quick-batch/steps/plan-checker-loop.md migration (the agy
HIGH finding) as a violation, because it also edits
quick/steps/plan-checker-loop.md for the same underlying #3771 protocol
fix.

Verified against git history before scoping: 2f64e6230 (#3676's own
landing commit) CREATED quick-batch/steps/plan-checker-loop.md as a new,
independent 119-line file, never a call-site into quick/'s copy. Row
48's "shared primitives, never edits the ordinary quick command" premise
was never about this specific file -- it was always meant to carry its
own per-flow copy of whatever revision-loop contract applies, same as
ui-phase.md/verify-work.md throughout this PR. This is the same
false-positive class the row's own comments already document scoping
away twice (#3730, #2529 round 40); excluded plan-checker-loop.md from
its touched-quick-steps check with the same evidence trail.

* chore(#3771): point changeset pr field at upstream PR 3916

---------

Co-authored-by: davdittrich <davdittrich@gmail.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 15:16:38 -04:00
Dennis Alexis Valin Dittrich
5869febb16 enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change

readVerificationStatus() now recomputes a deterministic sha256 fingerprint
over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY,
requirements, implementation files in the verified change set) and returns
stale on any mismatch, fail-closed when a covered file is missing,
unreadable, or escapes the project root. Legacy reports with no fingerprint
metadata keep the prior SUMMARY-mtime staleness check unchanged.

The verifier computes covered_digest via the new verification.fingerprint
CLI command rather than by hand, since a digest is deterministic math, not
an LLM-estimated value.

* chore(#4155): backfill fork PR number in changeset

* fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap

* fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior

Partial fingerprint metadata (one of covered_files/covered_digest present,
the other missing or malformed) now fails closed to stale instead of
silently downgrading to the legacy mtime-only check. computeCoveredDigest
also canonicalizes with realpathSync before re-confining, so an in-root
symlink whose target escapes the project root can no longer produce a
matching digest. gsd-verifier.md restores the completeness requirement and
checklist item trimmed by the earlier size-budget fix, within the LARGE
tier byte cap.

* chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap

* fix(#4155): address gemini adversarial review findings

computeCoveredDigest now threads the caller-supplied opts.fs seam through
its confinement and read paths instead of always using raw node:fs — a
caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP
2, #2790 follow-up) was silently bypassed for covered-input reads. The
project-root anchor itself still canonicalizes through real fs (it is a
trusted value the caller derived, not attacker-influenced covered-input
data); only per-file candidate reads go through the injected seam.

Covered-file paths are now canonicalized (./ prefixes, redundant slashes,
internal .. segments) before becoming dedup/sort/hash keys or confinement
subjects — closes both a spurious-stale false positive (two spellings of
the same file hashing differently) and a confinement gap (an internal ..
segment that doesn't start the string).

gsd-verifier.md now states covered-file paths are project-root-relative,
not phaseDir-relative, closing an ambiguity that would have made a real
verifier agent's first fingerprint invocation fail closed.

defaultFsImpl's methods now late-bind through fs.<method> rather than
capturing function references at module load — the earlier direct-capture
form was invisible to existing tests' t.mock.method(fs, 'statSync', ...)
seams, a real regression caught by the full suite (not the reviewer).

* fix(#4155): catch a plan/summary added to the phase dir after verification but never declared

The content digest only recomputes hashes for paths the verifier actually
declared in covered_files — it had no way to notice a plan or summary
added to the phase directory after verification if that new file was
never declared, silently regressing behind the legacy mtime check it
replaces (which scans the live directory, not a declared list).

findUncoveredCurrentArtifact re-scans the live phase directory for every
current *-PLAN.md/*-SUMMARY.md and requires each to be represented in
covered_files, closing that gap; a directory scan failure fails closed to
stale rather than silently skipping the check.

CONTEXT.md's Verification Module entry corrected to describe the
fingerprint path's stricter fail-closed FS-error contract (routes to
stale) instead of the module's original degrade-to-safe one (missing /
not-stale), which only the legacy path still keeps.

* refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test

computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/
sorted covered_files independently — one shared helper now backs both
(gemini review's ponytail-lens finding).

Adds one CLI-to-readVerificationStatus test against a genuine
.planning/phases/NN-x/ project with an implementation file outside
.planning/ entirely, closing the review finding that prior #4155 unit
fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot
falls back to phaseDir itself there) and never exercised real multi-level
path resolution.

* fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/

Two independent review rounds (opus critical-reviewer + opus ponytail +
agy, run twice) found two instances of the same fail-open class:

- computeCoveredDigest's per-file reads routed through the caller's
  injected fsImpl. planning-inspect.cts passes a `.planning/`-confined
  containment fs into readVerificationStatus's opts.fs, so any covered
  implementation file outside `.planning/` (mandatory per the issue)
  made the confinement wrapper throw, which was caught and turned into
  a stale digest -- reporting every fingerprinted phase permanently
  stale via `planning.inspect`, regardless of actual drift. Per-file
  reads now always use real node:fs, matching the pre-existing
  treatment of root canonicalization; the realRel-vs-realRoot check is
  the real confinement boundary for this data and needs no seam.

- allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans
  reports readdir failures via a `scope` field, it never throws), so
  an unreadable nested plans/ dir was silently treated as "zero
  artifacts, all covered" instead of failing closed. Now branches on
  scope !== SCOPE.COMPLETE.

Also, per ponytail's second-round findings: reverted an unwarranted
FINGERPRINT_VERSION bump and digest length-prefix from the first fix
(no v1 digest has ever existed -- the feature is unreleased -- and the
prefix closed a collision that grants no capability beyond what a
writer of covered_files already has more cheaply); removed a
verifier-facing escape-hatch instruction whose own example was a case
that should trigger staleness, not bypass it; corrected CONTEXT.md
references to the renamed allCurrentArtifactsCovered and a stale
"unconditional" rescan claim; simplified the isStale derivation,
removed dead FsLike members, and tightened test coverage.

Regression tests for both fail-open bugs are included and were each
confirmed to fail against the pre-fix code before the fix landed.

full test suite: 2558/2560 pass, 2 skipped, 0 fail

* fix(#4155): trim gsd-verifier.md under the LARGE size cap

Fork CI caught what my local runs missed: the superseded/nested-plans
instruction added earlier pushed gsd-verifier.md to 49299 bytes,
147 over the LARGE tier's 49152-byte hard cap
(tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's
wording and dropped a redundant inline comment tag; no content lost.

* chore(#4155): point changeset at the upstream PR number

pr: 19 was the fork PR opened for internal review-lane CI; now that
open-gsd/gsd-core#4290 exists, the changeset field must match it per
CONTRIBUTING.md's release-notes convention.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 05:42:52 -04:00
Dennis Alexis Valin Dittrich
eedb6b5431 enhance(#4107): sequence external review after internal fixes (#4206)
* enhance(#4107): sequence external review after internal fixes

Teach the planner to finish internal review and accepted fixes before opening a PR known to trigger automatic external review. If an open-time property exists, re-check it immediately before opening with nothing intervening; post-open CI, review, changeset, and tracking work may follow.

Emitted-Drift-Ack-Growth: gsd-planner.md — issue #4107 adds the review-before-publish ordering rule

* chore(#4107): add PR #11 changeset

* chore(changeset): link upstream PR 4206

* fix(#4107): ground external-review terms and tighten ordering test

Addresses trek-e review on PR #4206:
- Ground 'known automatic external review' and 'open-time property' with
  concrete anchors (CodeRabbit App / .coderabbit.yaml, not-behind-base).
- Suffix the antipatterns heading with (#4107), matching sibling sections.
- Replace vacuous negative assertion with inverted-order fixtures that
  prove the ordering regexes reject bad phrasing, not just co-occurrence.

* fix(#4107): make directionality fixtures genuinely adversarial

agy (gemini-3.8-flash-high) adversarial review found the two negative
fixtures added in 584ec1cda were vacuous: they proved the ordering regexes
require certain keywords, not that they reject inverted order — the bad
strings simply omitted required tokens rather than reordering them.

- Rebuild both fixtures to contain every required token, reordered/negated,
  so a real reordering would still slip past a weaker regex.
- Drop the unsupported 'changeset' mention from the Wave 4+ antipatterns
  example — gsd-core/workflows/ship.md never references changeset work,
  so naming it here implied a step this rule doesn't actually govern.

* fix(#4107): make the full review-then-fix-then-open sequence explicit

CodeRabbit (fork PR #11) flagged that the planner prose only ordered
accepted fixes before PR open, without explicitly naming 'run internal
review' as its own earlier step, and that no fixture tested the planner
text's own wording for inversion (only the antipatterns example had one).

- Prose now reads 'run internal review and apply the accepted
  internal-review fixes before the final open'.
- Added a planner-text-specific inverted-order fixture alongside the
  existing antipatterns-example one.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 03:14:27 -04:00
Tom Boucher
5214ad5802 fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication (#4295)
* fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication

execute-plan.md and gsd-executor.md each cited "Red-Green-Refactor Cycle"
for three facts (commit-scope contract, fail-fast rule, error handling),
but only the commit-scope contract lives there. Fail-fast is in tdd.md's
"Fail-Fast Rules" subsection (under "Gate Enforcement Rules") and error
handling is in tdd.md's "Error Handling" section — cite each correctly.

gsd-executor.md's "Plan-Level TDD Gate Enforcement" section also fully
restated the gate-sequence rules tdd.md's "Gate Enforcement Rules" already
owns (and covers more thoroughly, including the actual git-log validation
script). Collapse it to a short pointer, matching the treatment already
used by the cycle-steps pointer immediately above it.

Adds tests/tdd-reference-correctness.test.cjs asserting the pointer text
cites the correct section names, that those sections actually carry the
guidance, and that the old gate-sequence restatement is gone from
gsd-executor.md.

Closes #4267
Closes #4269

* docs(#4267): add changeset for tdd.md pointer correctness fix

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4267): acknowledge execute-plan.md growth

execute-plan.md grew 95 bytes (39766 -> 39861) from the corrected
three-section citation in the #3990/#4267 cycle-steps pointer.

Emitted-Drift-Ack-Growth: execute-plan.md — net +95 bytes from citing the "Fail-Fast Rules" and "Error Handling" sections by name instead of a single mis-scoped section (#4267).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4267): update PROSE_ALLOWLIST line numbers shifted by the pointer-citation edit

This branch's edits to agents/gsd-executor.md and gsd-core/workflows/execute-plan.md
shifted line numbers, leaving tests/no-bare-gsd-tools-command-position.test.cjs's
PROSE_ALLOWLIST pointing at stale lines. Update both entries to their new correct
lines (811 and 419 respectively) without changing the underlying prose.

* fix(#4267): restore INVALID_RED citation, fix allowlist line shift after #3770 rebase

The rebase onto next picked up #3770's already-merged fail-fast update to
gsd-executor.md's plan-level gate section, which this branch's own commit
collapses into a pointer. The conflict resolution kept the pointer but
dropped the literal "INVALID_RED" term that tests/tdd-red-evidence.test.cjs
requires gsd-executor.md to name — restored it. Also updates
no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST line number
for gsd-executor.md, shifted again by the rebase.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4267): fix non-matching regression-guard regex in tdd-reference-correctness

The fail-fast regression guard asserted gsd-executor.md no longer contains
"If a test passes unexpectedly during the RED phase" — but the actual old
prose (removed by this branch's pointer-collapse) read "If a test passes
unexpectedly during RED, STOP". The regex never matched the real old text,
so the assertion would have passed even against the unmodified pre-change
file. Caught by an isolated orthogonal review pass. Fixed to match the
actual removed wording, and confirmed (via a direct grep) it is genuinely
absent from the current file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4267): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 19:13:56 -04:00
Tom Boucher
8249ebcf6e fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate

RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do
not exist yet; every row fails on require. Per #3770 only an intentional
target-test failure may authorize GREEN; zero-test discovery, fixture crashes,
unrelated failures, and unexpected green are INVALID_RED.

* fix(3770): require intentional RED evidence before GREEN

Only an intentional failure of the TARGET test (distinctly named, TAP-reported
assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test
discovery, fixture/load crashes (file-named failures), nonzero exits without a
failing test, unrelated failures, unexpected greens, and malformed/missing
records are INVALID_RED and block GREEN.

- src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses
  the prohibition-enforcement TAP primitives; fail-closed, never throws)
- check tdd-red-evidence <record.json>: validates the persisted record
  (command, exit code, failing test, expected, actual)
- gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now
  requires the evidence record + gate verdict, not a nonzero exit or a RED: tag

* chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs

* fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib

- gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B <
  49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid)
- tests: the row-6 fixture used String.replace (first-occurrence), so the
  `not ok` line still named the target test and the classifier was right to
  accept it; replaceAll makes the failure genuinely unrelated
- eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs
  (lint the src/*.cts source, per ADR-457 migration rule)

Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line

* chore(3770): add changeset

* chore(3770): backfill PR number in changeset

---------

Co-authored-by: sim <sim@local>
2026-09-04 16:44:38 -04:00
Tom Boucher
97ce61dee2 fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally

* fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds

Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD

* chore(#3990): changeset for the single-statement TDD cycle

* chore(#3990): backfill changeset pr number

* fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap

* test(#3990): allowlist pin tracks the rebased line

* fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError

* fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError

Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with
empty stdout; the empty string survived the recovery path and surfaced as
'SyntaxError: Unexpected end of JSON input', hiding the captured error. The
recovery path now requires non-empty stdout, and an empty result throws with
the captured stdout/stderr/message so the actual error is on the record.
Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only
stops masking it.

---------

Co-authored-by: sim <sim@local>
2026-09-03 21:41:14 -04:00
Tom Boucher
3639ab0431 fix(#3968): measure commit claims at all three surfaces — ledger, verifier BLOCKER, porcelain HANDOFF (#4230)
* test(#3968): commit claims must be measured against git, never narrated

* fix(#3968): measure commit claims at all three surfaces

Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap)
Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument
Emitted-Drift-Ack-Growth: pause-work.md — #3968 uncommitted_files from git status --porcelain

* fix(#3968): retired slash syntax, allowlist line pin, git-compare test pin

* fix(#3968): persist the ledger on disk and reconcile with the same rev-list instrument

Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap)
Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument

* fix(#3968): hold the gsd-executor size cap with a compact ledger contract

Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap)

* fix(#3968): allowlist pin and HALT regex track the final prose

* chore(#3968): changeset for measured commit claims

* chore(#3968): backfill changeset pr number

* fix(#3968): quote the BASE expansion (SC2086)

* ci: raise the test-lane budget 21 to 32 minutes (measured cost grew past the cap)

---------

Co-authored-by: sim <sim@local>
2026-09-03 05:43:16 -04:00
Tom Boucher
2f64e6230a feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core

Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch
binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the
new pure decision-logic module (arg validation, effective concurrency,
deterministic merge order, spawn backpressure, verification/merge
routing, cleanup-entry construction — design doc rows 3-15,24,26-28,
30-36,39; property rows 51-53). quick-batch-update-items.test.cjs
covers the new updateBatchItems export on src/quick-batch.cts (rows
15,22-23, including the negative cycle-rejection case).
quick-batch-command-router.test.cjs covers the new
gsd-tools quick-batch CLI family (rows 46-47). These reference modules/
exports that do not exist yet.

* feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router

Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision
layer — CLI verbs and pure orchestration logic only; no workflow
markdown, no Agent()/git-worktree I/O.

- src/quick-batch-dispatch.cts (new): pure decision functions consumed
  by the (separate, follow-up) /gsd:quick-batch workflow markdown —
  parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder,
  computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome,
  buildCleanupManifestEntry (the last parses caller-supplied plan text
  via the existing parsePlanDocument; no filesystem access).

- src/quick-batch.cts: adds updateBatchItems, resolving the design
  doc's Open Question 1 as ONE additive export on this module instead
  of the second, independent BATCH.json writer the design doc
  originally proposed. Reuses the same withPlanningLock transaction
  shape, computeWaves, and platformWriteSync call resumeBatch/
  completeQuickItem already use; fails closed without persisting on
  an unknown item, an unknown/self dependency, or an introduced cycle.

- src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI
  family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs)
  as a first-party always-on command (like /gsd:quick), not the opt-in
  capability-registry path graphify uses. Verbs: create/update/resume/
  complete (wrap quick-batch.cts) and effective-concurrency/
  merge-eligible/spawn-plan/verification-routing/merge-routing/
  cleanup-entry/parse-args (wrap quick-batch-dispatch.cts).

Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property
rows 51-53. Rows covering workflow markdown / Agent() dispatch /
`git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the
follow-up markdown-authoring pass, per the phase brief's explicit
scope boundary.

* docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces

New-.cts-module ripple for the two Phase 4 modules (epic #3344,
ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts,
ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source,
not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json
(via `node scripts/gen-inventory-manifest.cjs --write`, after
`npm run build:lib`), and CONTEXT.md glossary entries for
"Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router
Module", plus an update to the existing "Quick-Batch Core Primitives
Module" entry documenting the new updateBatchItems export.

* test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count)

scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs
file under the quick-batch production module by longest-prefix match,
and that module is already at its 2-file cap (quick-batch.test.cjs +
quick-batch.property.test.cjs). The standalone
tests/quick-batch-update-items.test.cjs added in the prior commit
pushed it to 3 and failed `npm run lint:ci`. Fold its content into
quick-batch.test.cjs (append-only — no existing test in that file is
modified) and update the CONTEXT.md glossary reference to match.

Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after
`npm ci` (this worktree previously had no local node_modules, which
also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-*
unable to resolve typescript — resolved by npm ci, no code change
needed there). `npm run lint:ci` and
`npx tsc -p tsconfig.build.json --noEmit` are both green after this
fix.

* test(#3676): add failing tests for the quick-batch command/workflow markdown

Failing-first tests for Phase 4's markdown-authoring pass (epic #3344,
ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs
covers commands/gsd/quick-batch.md's frontmatter/objective/process,
gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR
1610 NEW_FILE_CAP) and step-fragment count, the isolation model
(rows 20-22), the executor single-writer invariant (row 18), merge
validation reusing the existing bounded primitive (row 25), the
optional research/plan-checker/verification leaves (rows 16,17,19,
30,31), planning-failure blocking execution (row 29), the submodule
guard (rows 36,44), and the new agents/gsd-planner.md quick-batch
mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers
row 48 (ordinary /gsd:quick stays byte-identical). Named
`gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's
longest-prefix bucketing doesn't fold these markdown-only tests into
the already-capped quick-batch/quick-batch-dispatch/
quick-batch-command-router production-module buckets from the CORE
pass. These reference files that do not exist yet.

* feat(#3676): author the quick-batch command, workflow, and planner mode

Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch
binding") — the orchestration layer that calls into Pass 1's CLI
verbs (src/quick-batch-command-router.cts).

- commands/gsd/quick-batch.md (new): frontmatter/objective/process,
  delegates argument validation to `quick-batch parse-args`
  (parseQuickBatchArgs) rather than re-deriving the grammar.

- gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR
  1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded
  step fragments under gsd-core/workflows/quick-batch/steps/:
  resume-mode, batch-init, research-phase (flag:--research),
  planner-wave (+ nested plan-checker-loop when --validate),
  worktree-dispatch, merge-wave, verification-wave (flag:--validate),
  completion. Covers design doc rows 3-45: capacity/isolation
  resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG-
  layer planning with full-task-catalog prompts and always-required
  depends_on/files_modified frontmatter, serialized worktree create/
  merge/cleanup via the existing worktree.cleanup-wave primitive,
  deterministic wave-order merging, verification routing
  (human_needed/gaps_found), the executor single-writer invariant,
  submodule fail-loud guard, and #1941 fork-base auto-degrade.

- agents/gsd-planner.md: additive new `load_mode_context` bullet for
  `**Mode:** quick-batch`, pointing at the new
  gsd-core/references/planner-quick-batch.md reference (documents the
  always-required depends_on/files_modified contract, reusing the
  existing frontmatter grammar — no new keys). Existing modes
  byte-identical, only a new bullet added.

- src/init.cts (+init-command-router.cts, +command-aliases.cts):
  cmdInitQuickBatch / `init.quick-batch` — model profiles,
  commit_docs, roadmap/planning existence checks, and the
  section_manifest field gating research-phase/verification-wave
  (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY
  atoms — no new atom needed).

Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the
prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43,
45 covered by construction (verb wiring, single-writer prompt
constraints, crash-window resume via unmodified Phase 3 primitives).

* docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern

npm run regen:derived output for the new command/workflow/reference
(epic #3344, ADR-1239 "Quick-batch binding"):
- skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md)
- docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md,
  planner-quick-batch.md, and the quick-batch-dispatch.cjs/
  quick-batch-command-router.cjs CLI-module rows' now-live
  `/gsd-quick-batch` cross-reference (was "(separate, follow-up)")
  + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`)
- gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`)
  — research-phase/verification-wave gsd:section entries for the new
  quick-batch workflow
- tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) —
  the new command/workflow/skill/reference files now ship to every
  runtime

scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for
gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS
word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional-
flag pattern gsd-core/workflows/quick.md already carries baselined
(e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2);
quoting would break the intended "omit this arg when the flag is
false" splitting.

* fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch

Security review pass findings, both confirmed real:

1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment
   interpolated the raw, attacker-influenced task ${description} (and
   the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw
   description into every planner's prompt in the layer) straight into
   Agent() prompt bodies with no boundary. Fixed by wrapping every such
   interpolation in a <security_context> + DATA_START/DATA_END
   boundary, matching the CONCRETE convention already implemented in
   this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md,
   gsd-core/workflows/debug.md) — commands/gsd/quick.md's own
   <security_notes> only asserts this convention in prose, so the
   debug-agent files are the real precedent followed here. Added a new
   <security_notes> block to commands/gsd/quick-batch.md (it had none)
   documenting both this fix and the one below.

2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection.
   gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both
   ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED,
   causing shell word-splitting and pathname expansion on raw task-list
   text before the parser ever saw it. Fixed at the source: added a
   `--text <string>` form to the `parse-args` verb
   (src/quick-batch-command-router.cts) that accepts the ENTIRE
   $ARGUMENTS as ONE quoted argv element and does the whitespace split
   itself, in Node — which is never glob-aware, unlike the shell.
   Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form
   is kept for direct/test callers that already have a real argv array.

The SC2086 baseline entry added for the original unquoted line is now
stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports
it) and has been removed; the two SC2046 entries for the UNRELATED,
still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style
conditional-flag splitting remain — that line only ever expands to one
of a few known-safe literal strings (never raw user text), matching
quick.md's own already-baselined convention exactly.

Tests: quick-batch-command-router.test.cjs covers the new --text form
(token splitting, glob-shaped text passing through literally
unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs
asserts the DATA_START/DATA_END boundary on every leaf prompt
(research-phase/planner-wave/plan-checker-loop/verification-wave,
including the shared task catalog) and the quoted --text call sites.

* fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35

Spec review pass findings — the test matrix claimed "yes" coverage
these assertions did not actually support:

- Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection
  only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs,
  committed alongside the security fix that touches the same file) that
  .planning/quick-batches/ is never created for any rejected value —
  createBatch is genuinely never reached.
- Row 18 (--resume <unknown-batch-id>): previously only exercised a
  hand-corrupted BATCH.json, never a genuinely nonexistent batch
  directory. Added the real nonexistent-id case (also in
  quick-batch-command-router.test.cjs).
- Row 24 (post-planning updateBatchItems racing a concurrent
  completeQuickItem for a different item, both through
  withPlanningLock): zero test existed. Added a property test
  (tests/quick-batch.property.test.cjs, appended — Phase 3's own file,
  no existing test touched) exercising both call orders and asserting
  no lost update in the final on-disk manifest — the same technique
  Phase 3's own row-15 lock-contention property test uses (sequential
  calls through the real lock; a working mutex makes any interleaving
  equivalent to some serial order, so this is the same claim a literal
  concurrent-thread test would make without OS-level threading).
- Row 34 (worktree preserved on merge_failed) and row 35 (undeclared-
  deletion detection): both were previously asserted only at the pure
  routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs
  using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs
  already establishes for executeWorktreeWaveCleanupPlan (real repo,
  real worktree, a REAL merge conflict / a REAL file deletion diffed
  against declared_deletions) — asserting the actual worktree directory
  survives on disk, not just that a pure function returns a
  preserveWorktree:true field. Named gsd-quick-batch-* so lint-test-
  file-count's bucketing doesn't fold it into any capped module bucket.

Row 48 (/gsd:quick regression) intentionally left as-is per the
reviewer's own framing: the byte-identity claim is already
mechanically proven by the changed-path diff (git diff --name-only
empty on those two paths IS byte-identity), and a genuine execution-
level regression test would require actually running the workflow —
out of scope for this repo's unit-test model (no other quick.md
regression test in this repo does that either).

* docs(#3676): add the changeset and user-facing docs the command needed

Standards review pass findings — both HARD:

- Missing changeset. None of the 6 prior #3676 commits touched
  .changeset/*. /gsd-quick-batch is a new user-facing command;
  CLAUDE.md/CONTRIBUTING.md require one. Added
  .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder —
  backfilled after the PR opens, matching CLAUDE.md's own documented
  convention and Phase 3's own precedent, #4190's
  .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen
  form `/gsd-quick-batch` throughout, never the source-artifact colon
  form (`scripts/lint-docs-command-form.cjs` confirms 0 violations;
  that check scans docs/**, not .changeset/, so it was never actually
  in scope for the fragment itself, but the wording still follows the
  doc convention for consistency, matching how Phase 3's own fragment
  named the not-yet-shipped command).
- Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis
  how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's
  existing convention for /gsd-quick /gsd-fast) covering --jobs,
  --validate, --research, --resume, --file, the capacity/isolation
  interaction, and resume/failure recovery. Cross-linked from
  docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's
  own "Related" section. Added a /gsd-quick-batch section to
  docs/COMMANDS.md (same table format as the existing /gsd-quick
  entry) and docs/features/quick-batch.md (REQ-QB-01..12, same
  frontmatter shape as docs/features/quick-mode.md) — regenerated
  docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md
  via the standard generators.

* fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught

gsd-test's real run against 155e8975b3 found 43 failures, all rooted in
this phase's own new command/workflow never being registered across
~10 independent generated/hand-maintained registries this repo keeps
in parity by convention. Root-caused each, no test weakened or
special-cased.

- help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live-
  registry.test.cjs): added a /gsd:quick-batch entry to
  gsd-core/workflows/help/modes/full.md (the real help.md content;
  gsd-core/workflows/help.md is a thin dispatcher) documenting every
  flag (--file/--jobs/--validate/--research/--resume), matching the
  existing /gsd:quick entry's format.

- gen-section-manifest.test.cjs: quick-batch.md's
  `gsd_run query init.quick-batch` invocation used inline
  `$([ ... ] && echo --flag)` substitutions, which never satisfy the
  test's exact-whitespace-token / assigned-variable detection (the
  trailing `))` glued onto `--research` in the compound substitution
  broke the "exact token" match). Rewrote to the same
  VALIDATE_PARAM/RESEARCH_PARAM two-line pattern
  gsd-core/workflows/quick.md's own Step 2 already uses.

- runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md
  fragments that call gsd_run each needed their OWN embedded copy of
  the canonical shim preamble (every workflow .md that calls gsd_run
  carries its own copy — reading one file does not persist shell state
  into another). Ran `node scripts/sync-runtime-launcher.cjs`, which
  inserted it before each file's first gsd_run call.
  plan-checker-loop.md correctly has none — it never calls gsd_run
  directly.

- Namespace routing (skill-manifest.test.cjs, install-nested-
  layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added
  `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and
  routing table (same namespace `quick` already routes through), and
  to src/clusters.cts's `utility` cluster (same cluster `quick`
  already belongs to). Verified by hand-running installRuntimeArtifacts
  + applySurface for augment/cline against a real temp install: exactly
  6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested
  under gsd-ns-workflow/skills/, never re-flattened.

- mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a
  brand-new command is a real count change, not a bug this test should
  hide).

- model-omit-when-inherit-guard.test.cjs: added the canonical
  `<!-- #2517 model-omit-on-inherit -->` marker block to
  gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/
  researcher/checker/executor/verifier — lives in a steps/ fragment,
  read combined with the host by this test's own readWorkflowCombined,
  same as quick.md's own research-phase.md carries it for its gated
  section). Also fixed a genuine pre-existing inconsistency in the
  test's own "#2711: the guarded set is derived from dispatch sites"
  check: its `nonDispatching` computation read the BARE host file while
  `derived` (the set it's checked against) reads the combined
  host+steps content — inconsistent with that same test file's own
  #2994 doc comment explaining why the combined read is necessary.
  quick-batch.md is the first workflow whose EVERY model="{...}"
  dispatch site lives in a mandatory (never gated) steps/ fragment —
  extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a
  brand-new file — which is what exposed the mismatch. Fixed by using
  the same readWorkflowCombined read in both places.

- skill-frontmatter-contract.test.cjs: shortened
  commands/gsd/quick-batch.md's frontmatter `description` from 107 to
  91 chars (<=100 budget), and added `quick-batch.md` to the hand-
  maintained KNOWN_SKILLS consolidation allowlist with a #3676
  justification comment (a genuinely new first-party command, not a
  consolidation of an existing skill).

- workflow-fragments-emission.install.test.cjs: added `quick-batch.md`
  to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is
  deliberately NOT a no-op for it — its research-phase/verification-
  wave sections are gated).

- Regenerated all downstream artifacts (npm run build:lib && npm run
  regen:derived && npm run gen:plugin-skills -- --write && npm run
  gen:features -- --write): skills/gsd-quick-batch/SKILL.md,
  skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for
  augment/cline/hermes/qwen/trae/zcode.

- emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition
  (one new `load_mode_context` bullet pointing at the new
  gsd-core/references/planner-quick-batch.md reference) grew the file
  124 bytes without an acknowledgment trailer. Acknowledged below —
  the growth is the deliberate, additive, single-bullet change from
  the earlier feat(#3676) commit, not drift.

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green
(includes lint-workflow-shellcheck, lint-test-file-count,
lint-docs-command-form). The deep install/spawn/registry tests gsd-test
actually runs (docs-parity-live-registry, gen-section-manifest,
runtime-launcher-parity, install-nested-layout,
runtime-artifact-layout-surface, skill-manifest, skill-frontmatter-
contract, mcp-server-catalog, model-omit-when-inherit-guard,
workflow-fragments-emission) are not part of lint:ci — each fix above
was independently verified by hand-invoking the exact production
function the failing test calls (installRuntimeArtifacts, applySurface,
composeWorkflow, the CLUSTERS union, the section-manifest forwarding
regex) against the real repo tree and confirming the expected shape.

Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift.

* fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget

skill-frontmatter-contract.test.cjs's "feature #3039: tiered help —
size budgets" enforces a SEPARATE line-count ceiling for
gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines,
tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's
assertTightCeiling) — independent of the skill-frontmatter description-
length budget and consolidation allowlist I touched in the prior round;
those are unrelated checks in the same test FILE, not the same check.

Root cause: the /gsd:quick-batch entry I added to full.md in the
docs-parity fix round was 17 lines, pushing the file from 834 to 851
lines — 7 over the 844 ceiling. Condensed the entry (merged the
per-flag bullet list into one dense "Flags:" line, dropped from 3
Usage examples to 1) to 844 lines exactly — at the ceiling with zero
slack, which assertTightCeiling accepts (it only fails on
actualMax > ceiling, or on slack > grace when the ceiling is too
LOOSE — zero slack triggers neither).

Verified after trimming: full.md still contains a live /gsd:quick-batch
reference (bidirectional parity) and all 5 argument-hint flags
(--jobs/--validate/--research/--resume/--file) still appear as literal
tokens (docs-parity-live-registry.test.cjs's own flag-coverage check,
re-run by hand against the trimmed content).

Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json
--noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green.

* docs(#3676): backfill changeset pr number to 4212

Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's
pr:0 placeholder backfilled with the real PR number now that
gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR
Number Handling convention and Phase 3's own #4190 precedent
(708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt
from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE.

* fix(#3676): resolve prompt-injection-scan false positive on test fixture

tests/quick-batch.test.cjs:232's row 11b regression proves the task-list
parser carries a prompt-injection-shaped task description through
createBatch as inert data, never interpreted. The fixture has to be a
real "ignore all previous instructions..." phrase or the test asserts
nothing, but the full-file --diff scan flagged it once unrelated edits
in the same file pulled it into the changed-file set.

Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the
sanctioned, precedented exemption already used for other legitimate
security-regression fixtures (tests/windsurf-conversion.test.cjs,
tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs)
per DEFECT.PROMPT-INJECTION-SCAN-COLLISION.

---------

Co-authored-by: sim <sim@local>
2026-09-02 22:38:31 -04:00
Tom Boucher
647365faf1 fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection

Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as
a gate condition, the end-of-phase escalation must not require MVP, the
executor agent's gate section triggers on TDD_MODE alone, and the gate
semantics reference loads without MVP_MODE.

* fix(#4011): key the TDD runtime gate on TDD_MODE alone

The RED-commit gate shipped as #76's MVP slice kept the paired
invocation's conjunct, so workflow.tdd_mode=true was silently inert on
every non-MVP phase, contradicting references/tdd.md's own contract.
Drops the MVP conjunct from the per-task gate and the end-of-phase
review escalation; rescopes execute-mvp-tdd.md's load condition,
gsd-executor's gate section, and mvp-concepts' intersection claim.
MVP remains free to imply TDD; the file is not renamed (stated
assumption in the PR body).

* test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing

Review follow-ups: the detector now only inspects if/[ condition lines
so explanatory prose mentioning both flags cannot trip it; remaining
'under/outside MVP+TDD' phrases in execute-phase.md, the gate
reference, and docs/INVENTORY.md now describe TDD-mode semantics.

Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011)
Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011)

* chore(#4011): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 04:24:52 -04:00
Cody Anderson
8c9265d4e5 fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop

Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a
blocker" but tagged severity: warning — the tier plan-phase's revision
loop counts as must-fix — and the planner is never taught the rule, so
every multi-wave phase touching shared mutable state replans at least
once, and intentionally coupled plans re-flag identically every
iteration to the stall prompt.

Three coordinated changes:
- gsd-plan-checker: retag 3b to severity: info, the tier
  references/revision-loop.md already exempts by design; recognize a
  coupling_justified frontmatter declaration in the Do-NOT-flag list so
  deliberate pairs converge. Additions are offset by trimming 3b
  motivation prose — the checker sits 45 bytes under its LARGE hard cap.
- plan-phase step 12: INFO-only accept — an issues block with zero
  BLOCKER/WARNING entries accepts the plan and surfaces the advisories
  instead of re-entering the revision loop. Real blockers and warnings
  still gate unconditionally.
- gsd-planner: slim pointer in assign_waves to the new
  progressive-disclosure reference gsd-core/references/planner-coupling.md
  (the planner sits 19 chars under its own cap), which carries the
  shared-mutable-state rule and the coupling_justified escape hatch so
  first-pass plans avoid the finding when the coupling is unintentional.

Documented the coupling_justified field in docs/reference/plan-md.md.
Growth acks per #2914; inventory manifest and install-tree fixtures
regenerated for the new reference file.

Closes #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin Dimension 3b at severity: info

The severity retag makes the old assertion (severity: warning) stale;
lock the advisory tier from both directions — info must be present,
warning must not — so a future edit cannot silently re-arm the
revision-loop trigger.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* chore(#3724): changeset fragment for PR #3758

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* docs(#3724): roster planner-coupling.md in docs/INVENTORY.md

The new reference was enumerated in the manifest and all 19 install-tree
fixtures but missing its row in the Modular Planner Decomposition table —
the roster half the manifest-sync test cannot check. (Review Blocker.)

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): cover all four acceptance criteria (review round 1)

- plan-checker-coupling: the 3b severity assertion is now a PARITY check
  deriving the exempt tier from revision-loop.md's flow instead of
  hardcoding info — editing either side alone reds the suite. New
  describe pins the other three criteria: plan-phase's INFO-only accept
  clause (proven failing-first), the BLOCKER + WARNING count staying
  intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and
  the planner pointer + planner-coupling.md content.
- ack fragment: $comment's plan-phase figure corrected to +79B; the 2775
  pin note carried forward into the gsd-planner.md entry, updated for
  upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves
  verbatim).

The parallel-dependent-plans re-anchor this commit originally carried was
superseded by upstream #3764 during review; this branch no longer touches
that file.

Refs #3724

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 2 — align the stance enumeration, complete the template contract

MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it
agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it.
Funded by extracting the inline <examples> block to the new progressive-
disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined
from the same spot; #1949 precedent), which also restores the 3b motivation
clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base
(49107 -> 48486) — the extraction the byte pressure was owed.

MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and
the field's shape becomes one 'plan-id: reason' string per coupled peer so a
plan justified against two peers can express it; docs/reference/plan-md.md's
Type column names the shape.

NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow
ratchet an 'XL tier'.

Acks and derived artifacts updated accordingly (checker entry removed — a
shrink needs no ack; INVENTORY roster row + regen:derived for the new file).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): derive the 3b negative severity assertion (review round 2)

Every severity token in the 3b span must BE the tier revision-loop.md exempts,
replacing the hardcoded severity:warning negative — if the loop's exemption
ever moves, the failure names the real conflict instead of blaming the agent
file with a mutually-unsatisfiable pair.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): refit the planner coupling pointer under the char cap

Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the
base, leaving 5 chars of headroom where the +16-char pointer was measured
against 13 more. The pointer prose shortens to 'Non-file coupling:' —
49150 chars, back under the strict 49152-char cap — and the ack figures
follow. The @-path the tests pin is unchanged.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep

Upstream #3078/#3823 deleted all fully-spent ack fragments, including
3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md
+79B append. Per the collision remedy that sweep added: take the deletion
and home the still-live entry in this PR's own fragment. Figures
re-measured at this merge base (90871 -> 90950 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack

Upstream #3825 shipped 3172-stated-failing-direction.json naming only
plan-phase.md, now spent at the base — colliding with this PR's live
plan-phase entry. Per the #3003 pattern the fully-spent single-path
fragment is deleted and this fragment stays the path's one source;
figures re-measured at this base (93073 -> 93152 LF bytes).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 3 — true up the ack figures, restore the wave comment

The fragment's absolute sizes are re-measured and anchored to base
e40e9670 (planner 47259 -> 47330 chars, checker 45537 -> 44916 B,
plan-phase 91186 -> 91265 LF bytes), with a note that absolutes rot as
next moves — the deltas are the durable claims. The round-1 removal of
the '# Implicit dependency: files_modified overlap forces a later wave.'
pseudocode comment offset headroom base drift had already returned, so
it is restored (findings 2-3). Changeset gains the (#3724) backlink
(finding 4).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 4 — close the verify-work surface, harden the boundaries

BLOCKER: verify-work.md's verify_gap_plans is the second multi-plan
consumer of the checker's sentinels, and its ISSUES FOUND handler entered
revision_loop with zero severity parsing — the guaranteed replan #3724
fixed in plan-phase, alive on the gap-closure surface. The handler now
counts BLOCKER + WARNING and accepts INFO-only returns with advisories
displayed. The checker's INFO stance bullet is reworded to the claim that
is true everywhere ('revision gates count only BLOCKER + WARNING').

Minor 1: plan-phase's iteration_count >= 3 arm recounts severities, so an
INFO-only third check accepts instead of halting on a '0 issues remain'
user gate. Minor 2: the coupling_justified exemption now requires the
entry to NAME the other plan, closing the blanket-suppression reading.
Nit 1: INVENTORY row states the extraction buys cap headroom, not context.
Nit 2: the advisory display gains a concrete format on both surfaces.

Ack fragment re-anchored at base ddde001a: verify-work.md +264B (new
entry), plan-phase.md +395B, checker still net negative (-512B).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the verify-work accept and the iteration-cap boundary (review round 4)

Two wiring assertions: verify_gap_plans' ISSUES FOUND handler gates on
BLOCKER + WARNING and accepts INFO-only blocks, and plan-phase's
iteration_count >= 3 arm recounts severities instead of gating advisories
— the limit+1 boundary of the gate this PR fixes.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 5 — fail closed at the gates, surface the advisory

Blocker 1: the checker's step-10 status rule routes an INFO-only result to
## ISSUES FOUND (with a new ### Advisories (info) template section and a
severity-aware recommendation) so the orchestrator receives the block and
displays the advisory instead of silently accepting a bare PASSED.

Blockers 2+3: all three gate surfaces (plan-phase step 12 both arms,
verify-work verify_gap_plans) carry one canonical clause verbatim — an entry
whose severity is missing or unrecognized counts as a BLOCKER (fail closed) —
making the accept condition an explicit-INFO whitelist while keeping
issue_count coherent for stall math.

Major 1: the INFO stance bullet scopes its claim to the plan-phase and
verify-work gates (quick mode's loop still revises on any ISSUES FOUND).
Major 2: INVENTORY row and ack $comment state the extraction's real trade
(readability, +0.6 KB eager runtime context), not a cap remedy.
Minor 1: plan-md.md marks coupling_justified as prompt convention, unvalidated.
Nit 1: ack absolutes re-anchored at base 1e67ec97; checker now +120B and acked.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* test(#3724): pin the round-5 contract — fail-closed parity, INFO-only return shape

New: three-surface verbatim parity test for the fail-closed clause (Blockers
2+3); checker return-contract test for the INFO-only ## ISSUES FOUND route and
advisories section (Blocker 1). All seven newly pinned tokens are absent at
f3a5682d, so each new assertion fails pre-fix.

Updated: accept-clause regexes track the explicit-INFO whitelist wording;
the severity sweep scopes to the span's fenced yaml examples via
yamlSeverityTiers (round-5 Minor 3, applied to the blocker negative too);
the iteration-cap comment states it is a prose pin, not an executed boundary
check (Minor 4); splitLines call sites document the line-pin coupling (Nit 2).

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): adopt next's line wrap in the 3b motivation clause — drops a wrap-only hunk from the diff

Byte-identical content; the wrap difference was an artifact of the round-1
base adaptation predating upstream's #3003 landing.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

* fix(#3724): review round 7 — gate every checker consumer, not just the two audited ones

Blocker: quick/steps/plan-checker-loop.md (issue-named in #3724) gets the
same canonical fail-closed clause and explicit-INFO whitelist accept as
plan-phase/verify-work — an INFO-only result proceeds instead of entering
quick mode's revision loop.
Major: import.md plan_validate handles the checker return by severity
(INFO-only never blocks an import) and is added to agent-contracts.md's
consumer enumeration, which had omitted it.
The checker's INFO stance bullet drops the quick-mode carve-out — the claim
is universally true again now that every consuming gate is severity-aware.
Minor: an applied coupling_justified exemption is surfaced as its own info
advisory so a stale one-sided declaration stays observable.
Nit: plan-phase's revision-iteration Display line is explicitly conditioned
on not having already proceeded to step 13.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM
Emitted-Drift-Ack-Growth: import.md — #3724 round 7: the plan_validate step's checker-return handler becomes severity-aware — counts BLOCKER + WARNING failing closed and accepts an explicitly-INFO-only return with advisories displayed instead of blocking the import

* test(#3724): pin the round-7 surfaces — five-gate parity, quick/import accepts, exemption visibility

The verbatim fail-closed parity test extends to quick/steps/plan-checker-loop.md
and import.md plan_validate; new assertions pin quick mode's INFO-only proceed,
import's never-blocks accept, import.md's presence in agent-contracts.md's
consumer row, and the surfaced coupling_justified exemption advisory. All four
newly pinned token families are absent at the pre-fix head, so each new
assertion fails first.

Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-31 19:08:11 -04:00
Tom Boucher
6beaa66b25 enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate

Content-assertion suite for the Step 7 re-verification evidence gate
(agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md).
Committed before the implementation to prove RED via gsd-test.

* enhance(#3304): gate re-verification blockers on deterministic evidence

Step 7's anti-pattern scan re-runs at full, unbounded scope on every
re-verification pass, independent of the must-haves established in Step 2.
A blocker it finds — other than the self-evidencing debt-marker check —
previously reverted a completed gap-closure round and started another
--gaps cycle on nothing more than the verifier's own new judgment call,
with no bound on how many times that could repeat.

A Step 7 blocker now blocks unconditionally in re-verification mode only
if it is a carried-forward gap (present in the prior VERIFICATION.md's
gaps: list) or the flagged file was git-modified since the prior pass
(a regression; fails closed toward blocking when history is unresolvable).
Otherwise it predates the gap-closure round unflagged and needs
deterministic evidence — a named test run red, or another concrete
reproducible artifact — to stay blocking. Unevidenced, it downgrades to
a new advisory: frontmatter list and report section instead of setting
status: gaps_found, and never reverts a completed must-have.

Maintainer approval was narrowed to this evidence condition only,
explicitly rejecting the broader "advisory whenever untraceable to a
requirement/decision/prior-gap" proposal — implemented and pinned by
tests/verifier-evidence-gate.test.cjs and documented as rejected in
gsd-core/references/verifier-evidence-gate.md so it can't silently
re-expand.

Closes #3304

* fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests

gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself
(not the production prose): a {0,600} match window was shorter than the
724-char paragraph it was scanning (the "exclude from Step 9 Rule 1"
phrase starts at offset 662), and two regexes assumed no indentation
after a markdown list-continuation line break. All three phrases are
confirmed unique across agents/gsd-verifier.md, so the windowed
submatches are replaced with direct whole-string assertions instead of
just widening the window.

Also acknowledges the deliberate byte growth in agents/gsd-verifier.md
that the differential-attribution check (ADR-2719) correctly flagged.

Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152).

* docs(#3304): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-30 16:51:52 -04:00
Tom Boucher
529480b4a5 fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model (#4048)
* test(#3895): no shipped agent may hardcode a model frontmatter pin (failing first)

* fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model

Exactly one of the 34 shipped agents carried 'model: sonnet' in its
frontmatter; every other agent resolves through the model-profile
system. The ship:post dispatch (#2684) resolves per-hook and — per
#2517 — deliberately OMITS model= on inherit so the agent inherits the
orchestrator's model; the frontmatter pin intercepted that inherit
case, silently forcing sonnet where all 33 siblings would inherit, and
operators could not durably remove it (install rewrites live copies
wholesale).

Deleting the line changes nothing for default profiles — the catalog
entry (model-catalog.json agents.gsd-mempalace-curator:
golden/balanced sonnet, budget haiku) preserves today's behavior —
while restoring model_overrides and inherit authority. Pinned by a new
agent-frontmatter guard: no shipped agent may hardcode a model pin,
and the catalog entry must keep existing so the pin's deletion can
never orphan the agent.

* chore(#3895): changeset fragment (pr number backfilled after PR creation)

* chore(#3895): backfill changeset PR number (4048)

---------

Co-authored-by: sim <sim@local>
2026-08-29 14:11:39 -04:00
Tom Boucher
192eb1dfbd fix(#3886): git commit timeout reported as commit_timeout; stale lock surfaced; 30s band (#4046)
* test(#3886): a timed-out git commit reports commit_timeout, not commit_failed (failing first)

* fix(#3886): git commit timeout reported as commit_timeout; 30s band; stale-lock surfaced

cmdCommit's git commit invocation did not distinguish a spawnSync
timeout from a real non-zero exit (#2608 fixed this for the staging
loop only): a slow pre-commit hook crossing the 10s cap was
SIGTERM'd mid-hook and reported as reason commit_failed with whatever
partial stderr git had flushed (in the reporter's case an incidental
CRLF warning), while the kill left a stale .git/index.lock blocking
the next attempt.

All three commit sites now check isSpawnTimeout before the
nothing-to-commit/ordinary-failure branches: cmdCommit reports
reason commit_timeout + timed_out:true and names the stale lock's
path (surfaced, not auto-deleted — deleting a lock a live git holds
is destructive; the caller recovers deliberately); the subrepo
counterparts do the same within their per-repo result / rollback
error. The commit calls also move to the 30s band the push call
already uses — husky+lint-staged alone idles ~4s on Windows before
any task runs.

* fix(#3886): review fold-ins — git-path lock resolution, shared band constant, executor contract row, precedence pin

- The stale-lock path is resolved via git rev-parse --git-path
  index.lock, never a literal .git/index.lock join (#3588 row 8's
  class: a linked worktree's .git is a FILE, so the literal path cannot
  exist there while the real lock — under <gitdir>/worktrees/<name>/ —
  blocks the next commit; this repo leans on linked worktrees).
- COMMIT_TIMEOUT_MS hoisted; all three sites and their messages build
  from it (the subrepo variant also regains the stdout fallback the
  primary site had).
- agents/gsd-executor.md's commit-result contract gains the
  commit_timeout row with the OPPOSITE retry advice from
  staging_timeout (remove the stale lock, then retry once) — an
  executor matching the doc previously had no handling for the new
  reason.
- Precedence pin: a timeout whose partial output contains 'nothing to
  commit' must still read as a timeout (branch-reorder mutant).

Emitted-Drift-Ack-Growth: gsd-executor.md — #3886: +commit_timeout row to the commit-result contract with the retry guidance OPPOSITE staging_timeout's (remove the stale lock, then retry once); the executor previously had no handling for the new reason.

* chore(#3886): changeset fragment (pr number backfilled after PR creation)

* chore(#3886): backfill changeset PR number (4046)

---------

Co-authored-by: sim <sim@local>
2026-08-29 13:27:14 -04:00
Tom Boucher
ac3668e4b7 fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow (#4008)
* test(#3797): the roadmapper must follow one write-first contract

* fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow

The roadmapper contradicted itself: role blurb, output format, and
completion checklist described an approve-first flow while its execution
flow said "Write Files Immediately" with reactive-only revision (#3797).
The approval gate belongs to the ORCHESTRATOR — both callers read the
written ROADMAP.md, present it, and gate on approval (with an auto-mode
bypass a subagent cannot host) — so write-first is the contract.

All approve-first text now describes the write-then-return reality, the
old "Draft Presentation Format" (whose ## ROADMAP DRAFT header matched no
orchestrator branch) is folded into the ## ROADMAP CREATED structured
return as a preview block, and the duplicate checklist lines are merged.
A structural guard pins the single contract.

Emitted-Drift-Ack-Growth: gsd-roadmapper.md — #3797: +bytes — approve-first wording replaced with write-first descriptions; the DRAFT presentation template folded into the ROADMAP CREATED return as a preview block

* chore(#3797): changeset fragment (pr number backfilled after PR creation)

* chore(#3797): backfill changeset PR number (4008)

---------

Co-authored-by: sim <sim@local>
2026-08-28 15:55:55 -04:00
Tom Boucher
52b11ee811 fix(#3763): pass --raw at every shipped config-get bash call site (#3961)
* test(#3763): guard every shipped config-get substitution on --raw

* fix(#3763): pass --raw at every shipped config-get bash call site

config-get without --raw prints JSON.stringify(value), so string-typed values
reach bash with literal quotes and every string comparison silently never
matches (#3763). --raw added at 75 command-substitution sites across shipped
content; four JSON consumers (default_reviewers, sub_repos, pr_body_sections,
code_review_depth_overrides) deliberately keep default JSON output.

Emitted-Drift-Ack-Growth: ai-integration-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: audit-fix.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: autonomous.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: cleanup.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: code-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: complete-milestone.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: do.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: eval-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: execute-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: execute-plan.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: fast.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: graduation.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: gsd-executor.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: health.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: import.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: inbox.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ingest-docs.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: mvp-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: new-milestone.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: next.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: plan-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: plant-seed.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: profile-user.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: progress.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: quick.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: remove-workspace.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: secure-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: settings-integrations.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: settings.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ship.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: sketch-wrap-up.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: sketch.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: smart-entry.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: spike-wrap-up.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: spike.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ui-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ui-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: undo.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: validate-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted

* chore(#3763): changeset fragment (pr number backfilled after PR creation)

* chore(#3763): backfill changeset PR number (3961)

---------

Co-authored-by: sim <sim@local>
2026-08-27 19:55:34 -04:00
Tom Boucher
fb2d122d7f feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb

only this package publishes. The path-based branches — a project-local install,
a runtime config directory — had no such guarantee; they trusted their
configured location. This closes them.

Mechanism: once resolution finishes, and before any verb runs, the preamble
probes the tool it picked with `runtime-identity --raw` and matches the answer
with a shell `case` pattern ANCHORED to the start of the compact payload
(`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts
the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which
any colliding package could publish. The outcome is exported as the two-valued
`GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE
rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling:
`unverified` prints one line naming BOTH causes and continues, because
`no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core`
older than the verb, and at rollout the old-version case is the common one.

The blocker was byte budget, not design. The preamble is inlined into 112
shipped files and several sat within single-digit bytes of frozen ceilings
(`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a
first attempt broke five of them. What made room was collapsing the resolver's
twenty near-identical `elif [ -f … ]` arms into one candidate-list helper
(`_gsd_at`), which buys far more than the assertion costs. The preamble is now
2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every
capped file moved away from its ceiling rather than toward it. No cap raised, no
size-budget exception added, no override token emitted.

Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source
fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all
preserved byte-for-byte in substring terms; the snippet still begins with
`_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor
on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal.

Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both
described an `[ -x ]` guard as the load-bearing re-source defense. That guard
was tried and REMOVED in #3831 — it rejected the bare function name, fell
through every branch, and hit `exit 1`, which kills a sourced caller's shell.
`unset -f gsd_run` is the actual mechanism.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3841): pair the anchor's brace by requiring a closed identity payload

The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md
has unbalanced braces: net depth 2" — plus a knock-on report from its parent
`bug #1516` describe, which is the same failure counted once at the child and
once at the block.

Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and
increments on `{`, decrements on `}`, with no awareness of shell quoting. It
scans `new-project.md` PLUS every `new-project/steps/*.md`, and both
`new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy
— hence net 2 from a snippet that was off by exactly one. The unpaired brace was
the `{` inside the single-quoted `case` pattern of the identity anchor, which is
correct shell and invisible to a text scanner.

Fix in the snippet, not the guard. The pattern now anchors at BOTH ends:
`'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that
does real work rather than a cosmetic pair — a truncated payload whose prefix
matches now fails too, where before it verified. Safe for any future additive
field: a JSON object's own closing brace is always the last character, whatever
type the last value has, which is pinned by two negative-space tests (a nested
object and an array-valued last key must both still verify). Cost: +3 bytes,
against the 1,873 the resolver fold already gave back.

The alternative considered and rejected was dropping the literal `{` for a `?`
glob. It balances too, but weakens the anchor from "must be an opening brace" to
"must be any one character", and the anchor is the entire point.

Two guards added so this cannot recur silently:
- runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next
  edit to that pattern fails on the file it broke instead of surfacing three
  files downstream in a test whose name mentions neither the launcher nor this
  issue. It also asserts depth never goes negative, since a `}` preceding its
  `{` nets to zero while being unbalanced at every prefix.
- runtime-identity gains behavioral truncated-payload and trailing-garbage
  fixtures, so the added `}` is proven load-bearing rather than merely present.

Verified: snippet 51/51 braces; new-project combined net depth 0; the seven
other preamble-bearing files with nonzero depth are unchanged from merged next
(their own prose, not the preamble, and not in any guard's scan set); all 112
inlined copies and the resolver reference re-synced byte-equal; sync:launcher
idempotent on the second run.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3841): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 01:05:53 -04:00
Tom Boucher
63abcface9 feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools

The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git.

The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap.

unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell.

Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true.

An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes.

Closes #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3146): stop sync:launcher relocating a deliberate preamble placement

Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins.

Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture.

Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3146): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3146): document the FEATURES.md section-numbering practice

The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases.

Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set.

Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914).

Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:57:16 -04:00
Tom Boucher
c933184b97 enhance(#3172): require a stated failing direction for every automated acceptance command (#3825)
* test(#3172): failing-first suite for the stated failing-direction probe

Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel
exemption, degraded-read contract, CLI arm and the plan-authoring contract text.
RED by construction: the module exports it requires do not exist yet.
Executed on the remote runner.

* feat(#3172): require a stated failing direction for every automated acceptance command

Every runnable <automated> command now carries a <fails_when> sibling naming
what output constitutes failure. A command with no expressible failure mode is
not an acceptance test: it reads as rigour and is not falsifiable.

- verify-command-grounding gains a failing-direction probe sharing the existing
  <automated> grammar, MISSING sentinel and walk guard rather than copying them
- gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches
  it and hands the JSON to gsd-plan-checker check 8f
- Dimension 8 detail extracted to references to stay under the agent size cap

Verified on the remote runner.

* fix(#3172): close four review findings in the failing-direction probe

- MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so
  a real command was exempted from the new blocking gate. Tightened the SHARED
  constant rather than adding a second copy.
- Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k).
  Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The
  pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too.
- probePhaseFailingDirections reported status 'ok' when one plan was unreadable,
  conflating 'could not look' with 'nothing to report'.
- Extracted the phase-resolution block both check arms had copied verbatim.

Also corrects a docs/AGENTS.md dimension list stale since #2401.
Verified on the remote runner.

* fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping

The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen
under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars
of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled
where such a rule goes: the planner spawn contract in plan-phase.md, beside
<tracked_source_paths>. The agent file is reverted to origin/next verbatim.

- plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the
  contract there and row 30b guards the freeze in both directions
- plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the
  precedent that two ack sources may never name the same path
- install-tree fixtures regenerated for the three new reference files

Verified on the remote runner.

* chore(#3172): backfill PR number into the changeset fragment

pr:0 -> pr:3825 now that the PR exists.

---------

Co-authored-by: sim <sim@local>
2026-08-24 19:05:11 -04:00
Tom Boucher
8442d984b9 fix(#3809): route runtime-loaded markdown through the gsd_run launcher (#3815)
* test(#3809): generalize dead-ref guard into a rule table (failing first)

The #2020 guard hardcoded `sdk/(src|dist|handlers)/` — the three dead paths
that had caused that storm. That proved those three paths were gone and said
nothing about the class, so #3809 reproduced the identical Windows find.exe
storm under a different token and the guard could not see it.

Replaces the single regex with a rule table over the same runtime-loaded
markdown surface, adds `commands/` to the scan set (previously uncovered),
and adds rule B: the runtime shim filename must never appear in command
position, because it is not a PATH command and an agent that meets it falls
back to locating the file.

Rule B's matcher is deliberately lenient — the launcher's own resolver
assignment, `node <path>/<shim>` calls, bare paths, and prose that names the
file all stay unflagged, each pinned by a negative-space row.

This commit is expected to FAIL: 50 offenders across 23 files remain in the
tree. The remediation lands next.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): route every workflow call through the gsd_run launcher

50 places across 23 runtime-loaded workflow, agent, reference, and command
files instructed the agent to run the runtime shim by filename. That filename
is not on PATH under any name -- package.json ships gsd-core, gsd-tools,
gsd_run and gsd-mcp-server -- so the call exited 127, the file-shaped token
sent the agent looking for the file, and on Git Bash for Windows the resulting
`find /` walked the entire drive (7268 CPU-seconds in the report) until
somebody killed it by hand.

CONTEXT.md -> Runtime Launcher Module already makes gsd_run the single entry
point: "Canonical space-safe shell preamble (`gsd_run`) used by every workflow
bash block to invoke the GSD runtime CLI." These sites predate that rule --
they trace to 0e6907050 (docs(#195): migrate workflow markdown off gsd-sdk
query), which swapped one non-PATH token for another.

Two further instances of the same class surfaced during remediation and are
fixed here rather than left for later:

  - references/model-profiles.md prescribed `node <shim> effort sync` with no
    path at all; node resolves a bare filename against cwd, so it fails the
    same way.
  - references/universal-anti-patterns.md rule 25 instructed every agent to
    "use <shim>" when shelling out. That rule did not contain the defect, it
    prescribed it repo-wide.

Five "(or legacy <shim>)" parentheticals left dangling by the substitution are
removed; after the rewrite they offered the non-resolving form as an
alternative.

The guard from the previous commit now passes. Its node-prefix exemption was
tightened to require a path separator, which is what exposed model-profiles.

Fixes #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): key the guard on the CLI's whole verb roster, not observed usage

Review found the first cut of rule B repeating the very mistake it exists to
prevent. Its verb set held query, commit and effort -- the verbs that happened
to appear in the tree -- so it could not see `<shim> phase add`,
`<shim> state load`, `<shim> verify ...` or twenty-odd other real single-word
subcommands. A guard that only recognises yesterday's offenders is not a guard.

The set is now the CLI's full advertised roster, unioned from the usage banner
and HOST_COMMAND_ROUTERS (which carries verification, planning, uat, stats,
todo and windows, all absent from the banner).

Widening it immediately caught a live offender the first pass had missed:
references/planning-config.md prescribed `node <shim> worktree set-baseref`
with no path. Fixed here.

Also drops the "a hyphen or a dot means subcommand" heuristic, which was
unsound for prose -- it flagged `built-in` and `v1.2`. Detection now keys
entirely on the roster, testing the first dot-segment so that phase.add and
state.patch still match while prose does not. Both false positives are pinned
as negative-space rows.

Guard verified against the pre-fix tree at origin/next: 52 offenders across 25
files, and 0 after this branch's remediation.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): derive the verb roster from the router; repair launcher parity

Standards review caught the guard repeating the defect it exists to prevent.
Its verb list was a hand-copied literal -- and worse, transcribed from an
INSTALLED older binary, so it was missing 22 verbs this tree actually ships
(websearch, windows, state-snapshot, context-predicates and the dispatch-*
family among them). gsd-tools.cjs already carries three hand-maintained
rosters whose drift is a named defect pinned by the parity test in
tests/commands.test.cjs; a hand-copied fourth was that same defect wearing a
guard's clothes.

The roster is now derived from HOST_COMMAND_ROUTERS + TOP_LEVEL_USAGE, lazily
and memoised, with `query` supplemented explicitly -- it dispatches through
the routing hub ahead of the host-router table, so it appears in neither
export, yet 45 of the 50 offenders used it. A parity test pins the derivation.

Two regressions this branch introduced, both caught by the remote runner:

  - runtime-launcher-parity: rewriting a comment in gsd-research-synthesizer.md
    put a `gsd_run` token at line 65 while the canonical preamble sits at 158,
    breaking "exactly ONE preamble, before the first gsd_run call". The comment
    is descriptive and needs no command token at all; it now names none.
  - The #2751 guard's PROSE_ALLOWLIST entry for that same line went stale once
    the line stopped carrying a bare mention. Pruned, exactly as that guard's
    own stale-entry test instructs.

Also corrects git-planning-commit.md, where the first pass rewrote only the
trailing "legacy" clause and left the sentence reading backwards.

Note the #2751 guard and this one are complementary, not duplicates: its regex
requires whitespace immediately after `gsd-tools`, so it cannot match the
`.cjs` form, and this one only matches the `.cjs` form.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2751): extend the bare-command guard to references/ and commands/

The #2751 guard has only ever scanned agents/ and gsd-core/workflows/. Two
runtime-loaded directories were never in its scan set, and 47 bare
`gsd-tools <verb>` calls had accumulated there unseen -- the same defect that
guard exists to catch, in the rooms it never entered.

  - gsd-core/references/: 37 calls, all rewritten to gsd_run. references are
    fragments inlined into a parent that defines the launcher, which is why 21
    of the 22 files already using gsd_run carry no local preamble.
  - commands/gsd/: 10 operative calls rewritten. The remaining 10 are
    descriptive prose ("resolved inside the workflow via ...") and are
    allowlisted with reasons, bringing PROSE_ALLOWLIST to 15.

commands/ also came under launcher propagation. sync-runtime-launcher.cjs
walked only WORKFLOWS_DIR and AGENTS_DIR, so every preamble under commands/
was a hand-pasted copy nothing propagated and no test checked -- graphify.md
had accumulated five. It now walks COMMANDS_DIR too, which collapses those
five to the canonical one-per-file, and runtime-launcher-parity gains a
(B-commands) arm mirroring (B-agents) exactly so the placement stays honest.

The parity arm keys on shell blocks, so commands/gsd/workstreams.md and
config.md -- which name gsd_run only in inline backtick prose -- are exempt,
as they should be. gsd_run is itself a shipped npm bin, so those inline
instructions resolve from PATH exactly as the gsd-tools form they replace did.

skills/ is deliberately NOT added to either guard's scan set: it is generated
from commands/ and pinned by lint:generated-sync, so guarding the source
guards both, and scanning the mirror would double-report every future
offender. Regenerated here.

Refs #2751, #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3809): acknowledge the one emitted file this change grows

The emitted-attribution gate failed on the previous sha: gsd-research-synthesizer.md
grew 3 bytes (13847 -> 13850) with no acknowledgment. The substitution SHRANK the
other 19 emitted files, which is why the growth arm was not expected to fire at all.

The 3 bytes are unavoidable. Line 65 is a descriptive comment inside a fenced block;
naming any command there puts a gsd_run token ahead of the file's canonical preamble
at line 158, which runtime-launcher-parity's (B-agents) arm correctly rejects. So the
comment names no command and says where the config is actually loaded instead, which
reads longer than the token it replaced.

Acks only the path the gate reported, per the fragment rules.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* revert(#2751): drop the commands/ half — three contracts pin it in place

The remote runner refuted the commands/ extension outright. Reverting it and
keeping the references/ conversion, which passed.

What broke, all of it caused by bringing commands/ under launcher propagation:

  - graphify.md's five per-block preambles are LOAD-BEARING, not accumulated
    drift. tests/graphify-visualization.test.cjs extracts individual Step-3
    shell chains and executes them standalone, so each fenced block needs its
    own definition of gsd_run. Collapsing them to the canonical one-per-file
    produced `bash: gsd_run: command not found`, exit 127, across four tests.
    The "define once per file" contract holds for workflows and agents because
    nothing extracts their blocks in isolation; commands/ is not like that.
  - explore.md broke "the preamble that DEFINES gsd_run must appear before the
    first USE of gsd_run anywhere in the file".
  - tests/gsd-tools-path-refs.test.cjs (#1766) ASSERTS that
    commands/gsd/workstreams.md contains the literal string
    `gsd-tools query workstream.list`. Rewriting it to gsd_run contradicts a
    test that pins the opposite, so the two guards disagree about that file by
    construction.

So commands/ is not a scan-set widening. It needs those contracts reconciled
first, and that is its own change. SCAN_DIRS keeps gsd-core/references/ and
drops commands/, the ten commands/ allowlist entries go with it (back to 5),
and the reasoning is recorded in the guard itself so the next person does not
rediscover it by burning a matrix run.

commands/gsd/import.md keeps its #3809 fix — that one is the .cjs form this
PR exists to remove, and it is untouched by any of the above.

Refs #2751, #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* revert(#3809): restore explore.md's Step 1 preamble placement

Running the launcher sync script processed workflows/ and agents/ too, not
just the commands/ directory the run was aimed at, and it MOVED
gsd-core/workflows/explore.md's preamble from Step 1 down to Step 3.

The script inserts into the first bash block that USES gsd_run. explore.md's
Step 1 block only DEFINES it, and that placement is deliberate -- the file
says so on the line above: "Placed in Step 1 rather than Step 3 so declining
the research offer cannot leave Step 5's commit call unbootstrapped."
tests/explore-command.test.cjs pins it.

explore.md carried no #3809 offender, so reverting it costs this fix nothing.
This was collateral from invoking the sync script at all, not from the
COMMANDS_DIR change, which is why the earlier commands/ revert did not catch it.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3809): backfill PR number into changeset fragments

pr:0 -> pr:3815 for both fragments now that the PR exists.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): drop the hand-rolled regex escaper CodeQL flagged

CodeQL raised js/incomplete-sanitization (HIGH) on the guard's pattern build:
`SHIM.replace(/\./g, '\\.')` escapes the dot and nothing else, so it does not
escape backslashes. It blocked PR #3815.

The repo already bans this shape -- local/no-adhoc-regex-escape exists exactly
to stop hand-rolled escapers, with the canonical one in src/pattern.cts. Rather
than reach for that helper, the pattern now carries no escaping logic at all:
SHIM is a compile-time constant whose only metacharacter is the dot, so the
regex source is spelled out literally. The generated source string is
byte-identical to what the replace() produced, verified before and after --
0 offenders on this tree, 52 against origin/next, unchanged.

A drift pin asserts SHIM_PATTERN still matches SHIM exactly, and that the dot
is escaped rather than acting as a wildcard, so the two cannot separate.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 11:47:37 -04:00
Tom Boucher
cf15682d1c enhance(#3028): responsive Markdown separators instead of fixed-width rules (#3789)
* feat(#3028): responsive Markdown separators instead of fixed-width rules

Stage banners, checkpoints, completion and error panels used fixed-width
runs of box-drawing characters -- a 53-column heavy rule and a 62-column
double-line box. Those runs are ordinary text to a Markdown-rendering
host, so in a narrower pane they wrap and the border comes apart from
the heading it framed.

Shipped content now emits an ATX heading for a titled section and a
blank-line-delimited --- for a break between sections, both of which
adapt to the available width. The same convention is applied to the
three code sites that built these strings at runtime: the UAT
checkpoint renderer, the milestone-close audit report, and the TDD
review checkpoint table.

Removing the box also removes its only reason to exist -- the
east-asian-width padding helpers that kept its right border aligned
(checkpointBoxLine, displayWidth, isWideCodePoint, ZERO_WIDTH_MARK_RE,
CHECKPOINT_BOX_WIDTH). RTL directional isolation is unchanged.

The convention is specified in gsd-core/references/ui-brand.md and
enforced across all shipped content by tests/responsive-separators.test.cjs.

Refs #3028

* test(#3028): pin the heading form in checkpoint and audit-report assertions

These suites asserted the exact box borders and the 62-column padded
banner interior. With the box gone they assert the ### heading form,
the --- break and the bolded instruction line, and each now carries a
positive assertion that no box character remains -- which is what pins
the fix rather than merely tolerating it.

Language coverage is converted, not dropped: Japanese, Chinese, Korean,
Hindi and Arabic all still assert their rendered banner, and the Arabic
case still asserts the RTL directional isolates the box removal must
not disturb. Adds a case for a banner longer than the old inner width,
which previously produced a ragged border and now has none.

Refs #3028

* chore(#3028): acknowledge execute-plan.md growth from the checkpoint display spec

The checkpoint_protocol display spec described the drawn box; it now
describes the heading, the --- break and the bolded action prompt,
which costs 22 bytes (40111 -> 40133, 827 under the cap).

Appended to the existing #3370 fragment rather than filed as a new one:
a growth ack keys on the bare filename and #3370 already declares
execute-plan.md, so a second source naming it would be a hard
duplicate-key error. Same supersede-by-append route #3370 took for the
spent #2652 fragment.

Refs #3028

* docs(#3028): state the load-bearing half of the separator rule, and amend the zh-CN reference

Review found three things.

The rule as first written demanded a blank line above AND below every
---. Only the one above is load-bearing: it is what stops CommonMark
reading the rule as a setext underline for the line above. The one below
is cosmetic, because a thematic break is a leaf block. The rule now says
that, with the reason, instead of asserting a stricter form the content
does not keep.

The zh-CN reference had received the mechanical box-to-heading swap but
none of the prose behind it: it still claimed a 62-character checkpoint
width and still listed --- among forbidden mixed banner styles, so it
contradicted the convention it was translating. It now carries the
separator section, the setext reasoning, the unconditional-vs-per-runtime
rationale and a corrected anti-pattern list, in Chinese.

The user guide asserted that a heading is not a degradation anywhere.
That is an assertion, not a demonstration. It now says what was actually
traded away in a plain terminal, points at the recorded rationale, and
invites the report that would justify the capability flag instead.

Refs #3028

* chore(#3028): backfill changeset PR number

Refs #3028

---------

Co-authored-by: sim <sim@local>
2026-08-23 22:38:12 -04:00
Behruz Nassre Esfahani
622f43353c fix(#3299): tracer feedback gate honors workflow.human_verify_mode (#3390)
* fix(#3299): tracer feedback gate honors workflow.human_verify_mode

The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.

Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.

The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.

`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.

Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.

Fixes #3299

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): add changeset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): reconcile the canonical schema table and the stale acceptance test

Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.

1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
   schema reference for the tracer task-type contract, and its Task-types row
   still claimed interactive runs unconditionally present a
   checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
   first round; this one was missed, so the authoritative reference was the
   wrong answer. The row now carries the human_verify_mode-conditional
   behavior and points at the canonical precedence chain.

2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
   tracer ROW EXISTS, never its content, which is why CI could not see the
   drift. It now asserts the row's actual claims and rejects the pre-#3299
   wording. Separately, the #1945 acceptance test named 'interactive run emits
   checkpoint:human-verify after the tracer' kept passing only because its
   substrings still occur in the fallback clause, while its name asserted the
   opposite of shipped behavior. Renamed and narrowed to what #1945 still
   guarantees, plus a new interactiveIsConditional pin so the unconditional
   prose cannot be restored under a passing substring check.

3. plan-md.md's <verify> row now documents that the legacy bare-text form
   (valid, and still shown at :179) does not reach the #3299 auto-continue —
   only a <verify> carrying <automated> does — so the benefit is silently
   unreachable for tracers using that format.

Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions

Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.

MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.

MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:

- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
  'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
  same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
  appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
  presence. Deliberately brittle: CONTEXT.md names that table the canonical
  schema reference, so a wording change must be a conscious edit in both
  places.

Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs

Peer review round 4. Blacklisting did not hold, twice over:

- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
  forms in the auto-continue clause. Round 4 defeated that by appending
  'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
  before expansion' — none of the banned tokens, same restored interruption
  after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
  '<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
  legacy bare form operative. 107/107 passed across tracer, planner and the
  three size-cap suites.

Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.

These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.

Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): strip comments, require uniqueness, pin whole regions

Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):

- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
  wrong copy: every extractor selected the commented decoy. Worked against the
  planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
  checkpoint_protocol before expansion, regardless of the mode-specific rules
  below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
  approval' below the canonical table. The pinned text was untouched, so
  equality held while the shipped meaning inverted.

The shape that holds, applied to every operative surface:
  1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
  2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
     cannot hide behind a correct first one;
  3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
     suffix can override what the pin proves.

Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.

Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.

Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): drop the superseded exact-placeholder planner assertion

Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.

Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.

Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner

Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:

- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
  internally uses the repo's interleaved fence/comment scanner. Instrumenting
  candidate lines as throwaway predicate declarations borrows that scanner with
  no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
  edits silently turned the guards into decoy checks:
    * a forgotten '-->' comments the live rule through to EOF, and the
      balanced-only stripper still saw and accepted the commented rule;
    * a normal fenced documentation example of the rule, plus a whitespace-only
      reformat of the live list item, made the selector choose the example.
  Neither needs intent. A dangling comment is a typo; a fenced example is good
  documentation. Together they reproduce exactly the accidental drift #3299 came
  from — with CI green.

The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.

Verified both ordinary-edit scenarios now fail the suite (each was green before).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): close the operative-selection gaps the maintainer blocked on

trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.

1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
   replaced a matched candidate with an UNINDENTED marker regardless of the
   original line's indentation. A 4-space-indented CommonMark code block is not
   skipped by parsePredicates (it accepts indented declarations by design), so
   stripping the indent PROMOTED an indented decoy to operative — the exact
   inversion of the guard's purpose. The marker now preserves the original
   indent, and a candidate that is itself indented 4+ spaces is never injected.

2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
   pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
   pollute the count. Now filters on a Set of the indexes actually injected on
   this call.

3. RAW FENCE SELECTION (planner). The template test matched the first raw
   ```xml fence after the marker with no fence/comment awareness — the one
   selection in the suite that was not operative-aware — so a commented-out
   decoy template between the marker and the real one would be selected while
   the live template regressed. The opener must now be operative AND the first
   non-blank line after the marker.

4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
   a fenced example containing a ### / <type line truncated the pinned region
   early — a false FAILURE on a legitimate doc edit. End anchors now go through
   the same operative filter as start anchors.

Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.

Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#3299): allow-list operative indentation; pin marker provenance

Review round 9.

BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", "  \t" and "   \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.

MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.

MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.

Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.

KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge

The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in 362d0434b (#3370) once #2728's entries were spent.
#3370's own fragment now owns execute-plan.md at the base, so a new
3299-*.json naming that path would collide — mergeAckSources rejects a
duplicate key across fragments rather than silently last-winning.

Re-arms #3370's entry instead, the mechanism the gate is built for (a spent
ack whose reason changes in the diff is live again), carrying #3370's own
reason forward verbatim so the base growth keeps its account.

Verified: emitted-attribution 175/175 against origin/next@be9329b10.

* fix(#3299): honor golden rule 6 in the tracer gate, extract the chain

Addresses the review on #3390 (B1-B3, M1-M4, minors).

B3 — checkpoints.md asserted two incompatible rules about the same gate.
Golden rule 6 says gate="blocking-human" stops for a human in every mode;
the precedence table scoped row 1 to interactive runs, so a first-match
chain let an auto-mode tracer carrying that gate fall to row 2 and
auto-continue. Rule 6 wins: row 1 is now "Any run, any mode", the
justification sentence it falsified is gone, and the STOP is evaluated
before the auto-mode branch at all three dispatch sites — gsd-executor.md,
execute-plan.md and the plan-md.md schema row. Unreachable by our planner
is not unreachable: src/verify.cts parses only `type` and never consults
`gate` on non-checkpoint tasks, so an imported PLAN.md can carry it.

B1 — the LARGE-tier cap. gsd-executor.md is 49150 on next against a 49152
cap, so this PR could not add a byte. Extracted rather than trimmed: the
precedence chain now lives only in checkpoints.md (already @-imported by
<checkpoint_protocol>, so no new load), and the duplicate summary inside
that protocol section is a pointer. The rationale the earlier trim
deleted is restored — "production-quality, never a throwaway" and
"Pouring more layers onto a broken foundation...". Result 49097: 55 bytes
under the cap and a net 53-byte REDUCTION against next, so the PR returns
headroom instead of consuming it.

B2 — merged upstream/next and resolved all three drift-ack conflicts.
2775 changed shape upstream (string -> {reason}); adopted the new form.

M1 — the 2775 ack claimed the Nyquist Rule sat "twelve lines earlier"; it
is ~75 lines. Corrected to "earlier in the file".
M2 — ack arithmetic restated from measurement, not from a stale base. The
2943 #3299 append is DELETED: with gsd-executor.md now shrinking there is
no ripple to acknowledge, and emitted-attribution correctly flagged the
entry as stale.
M3 — changeset rewritten to the documented bold-lead + em-dash one-liner.
M4 — the two self-defeated shapes are gone. The planner-human-verify-mode
presence checks now go through operativeLineIndexes. The config-get check
does NOT: all three reads live inside ```bash fences, which is their
correct executable form, and that selector excludes fenced lines by
design. It instead pins exactly one live, uncommented, fenced read per
file — mutation-tested against both a commented-out read and a duplicate.

Minors — dangling colon lead-in dropped, a "below" pointer that pointed
above corrected, and the `(default)` asymmetry between the two dispatch
copies aligned.

Two defects the merge surfaced, both caught only by the full suite:
the new #3576 gate rejected this PR's own bare `references/checkpoints.md`
cite in planner-human-verify-mode.md (rewritten to the canonical
gsd-core/ form), and the line-keyed PROSE_ALLOWLIST entry for
gsd-executor.md needed 794 -> 795 after this change shifted the line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): correct the size record the 08-22 merge falsified

Review round: one Major, four Minors.

Major — the #3299 arm's arithmetic was measured before the merge and is
now wrong in a document whose whole purpose is to be an accurate size
record. Re-measured at head: execute-plan.md is 39315 B on next and
40111 B here, so the 796-byte delta was right but the endpoints and the
headroom were not (849 bytes against DEFAULT_CAP 40960, not 1003). The
superseded figures are named rather than silently replaced. Confirmed
the workflow cap counts LF BYTES while the agent cap counts CHARACTERS —
two caps in two units, one per file.

Minor 1 — 2943-context7-tool-name.json reverted to next. JSON.parse of
both sides was already identical; the diff was an em-dash/times-sign
re-serialization left over from adding and then removing the #3299 arm.
No business in this PR.

Minor 2 — the duplicated `tracer row Autonomy cell` test is gone. Both
copies were new here and carried the same ~8-line canonical string; the
one removed selected its row with a raw startsWith find, the shape this
suite records at :477 as defeated in round 1. Its rationale — why the
cell is pinned EXACTLY, and the append-a-contradiction attack that
defeated keyword matching — is carried onto the surviving fence-aware
copy rather than deleted with it.

Minor 3 — the executor's condensed interactive clause said only "re-run,
continue", which does not distinguish pass from fail; read in isolation
it invites expansion onto a broken slice, the outcome the gate exists to
prevent. Now "re-run; fails → HALT as above, passes → continue, no
checkpoint". The pinned expected string moved with it. Executor at
48,905 chars, 247 under the cap.

Minor 4 — 2775 asserted two different current sizes for gsd-planner.md.
The stale half is next's own text taken wholesale, so the contradiction
was inherited; it now reads as a before-figure rather than a current one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): cite the plan-md example by section, not by a drifting line

Review round 7, Nit N-1. The 2775 ack fragment justified its one-line
formatting with "matching docs/reference/plan-md.md:207's own example
style". At head, :207 is prose; the one-line <verify><automated>
example it means is at :222. The citation was accurate when written
(77c2fda, f23205c) and drifted with a later merge of next.

Re-pointed by section rather than by line — it has already drifted
once, and the fragment's whole purpose is to be an accurate record —
and the drift itself is recorded inline so the correction does not
quietly overwrite what the earlier number said.

Also narrows the changeset's "any task with gate=blocking-human" to
"any tracer carrying gate=blocking-human" (found by Codex in the
whole-PR pass). Golden rule 6 and the #3299 decision table both scope
that gate to checkpoints and to the tracer feedback gate; the normal
type="auto" branch never inspects `gate`, so the wider claim promised
behavior the implementation does not have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): answer fence-delimiter liveness by insertion, not replacement

Review round 9. The round-8 fence-awareness fix was itself unsound, in the same
class it was added to close.

`operativeLineIndexes` detects operative lines by REPLACING each candidate with
a throwaway predicate declaration and asking `parsePredicates` which survived.
Sound for ordinary content lines. Not sound for a fence DELIMITER, which is
exactly what the tracer-template selection passed it: deleting every ```xml
OPENER leaves each matching closer to become an opener, and since
`computeSkippedLineFlags` is a strict FORWARD state machine, fence parity
inverts for the whole remainder of the document.

Measured against the real file rather than argued:

  agents/gsd-planner.md has 3 live top-level ```xml openers — 0-based 180, 232,
  262. operativeLineIndexes reported 180 and 262. Line 232, the "Task-level TDD"
  example, read NON-OPERATIVE — a wrong answer from a helper whose only job is
  that question.

It passed only by parity coincidence, and one extra live example anywhere
earlier flipped it to a false FAILURE blaming a decoy that does not exist:

  HEAD as-is                  | anchor 260 | openIdx 262 | ASSERTION PASSES
  +1 unrelated ```xml example | anchor 265 | openIdx 267 | ASSERTION *** FAILS ***

Fixed by asking the question a way that perturbs nothing. `isOperativePosition`
INSERTS a marker on its own line immediately before the candidate instead of
replacing it. Insertion preserves every delimiter, and because the skip-state
machine runs strictly forward, a line inserted at `idx` observes exactly the
fence/comment state the candidate observes, with nothing but the marker between
them — so marker-operative IS the candidate's position-liveness.

The review's suggested direction (substitute a same-shaped opener that still
opens a fence) cannot work here: the marker would then be inside the fence and
would never parse as a predicate at all.

Position-liveness is not content-liveness, so the helper also rejects a line
that is entirely comment (`<!-- ```xml -->`), rather than leaving that to each
caller's own shape test to happen to exclude.

`operativeLineIndexes` now THROWS when its candidate regex matches a fence
delimiter, so the unsound route cannot be reached again by a future caller
rather than only being fixed at the one site that got it wrong.

Verified with the same extra-example scenario above: with the fix, all 35 rows
stay green. Teeth: reverting the call site to `operativeLineSet` turns the
tracer-template row red on the new guard. The regression row pins both live
openers (the second is the one the deletion route lost), the block-commented
and same-line-commented openers, a line inside a fence, and re-checks both
openers after unrelated lines shift above them.

Only tests/tracer-bullet.test.cjs changes — no agent file is touched, so the
5-char gsd-planner.md and 19-byte gsd-executor.md headroom are unaffected.

Verified: `npm run lint:ci` exit 0; full `npm test` 31307 tests / 31292 pass /
0 fail / 14 skipped, TMPDIR unset, against a freshly synced origin/next.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#3299): guard the delimiter class, match the scanner, pin the assignment

Codex full-PR review of #3390, run against the round-9 head. Three defects,
two of them in the code that round added.

1. The mode-read pin survived the regression it exists to catch.
   `READ` matched the config-get substring only, so rewriting the shipped line
   as `IGNORED_MODE=$(gsd_run query config-get ...)` kept the row green while
   nothing defined HUMAN_VERIFY_MODE — the gate falls through to STOP and #3299
   is back with the suite passing. The regex now requires the assignment. A
   lookahead after `end-of-phase` closes the other half: the bare prefix also
   accepted `--default end-of-phase-wrong`. Proven by mutation: renaming the
   variable in agents/gsd-executor.md now turns that row red, and did not before.

2. The round-9 fence-delimiter guard was a SAMPLE of the class, not the class.
   It probed a fixed list of five delimiter strings. `~~~xml`, ```json, `~~~~`
   and arbitrary info strings all walk past any list short enough to write down
   — the guard was added precisely because one such regex had already slipped
   through. Now matched against the lines the regex actually selects in the
   document, which cannot go stale and cannot miss a spelling nobody thought of.
   Four such spellings pinned as rows.

3. `isOperativePosition` disagreed with the scanner it delegates to.
   For `<!-- closed --> real content` it stripped the span, found surviving
   content, and answered "live". `computeSkippedLineFlags` skips an ENTIRE line
   whose trimmed text starts with `<!--`, balanced or not, before it considers
   fences at all. Verified directly against parsePredicates. It now applies the
   scanner's own rule instead of out-reasoning it. Latent for the present caller
   (its anchored ```xml shape cannot match a comment-prefixed line), real in
   general.

Disclosed rather than fixed, and raised with the maintainer: the exact executor
region pin ends before the second operative tracer-gate paragraph at
agents/gsd-executor.md:327, which is only heading-checked — so contradictory
later instructions could ship. How much of that file to pin is a call for its
owner.

Independently probed isOperativePosition across 19 edge cases before the review
(line 0, CRLF, tab / 4-space / mixed " \t" indentation, 0-3 space fences, nested
fences, ~~~ fences, info strings, bounds); all correct. That probe is what
surfaced finding 2, which the review then confirmed from the other direction.

Verified: `npm run lint:ci` exit 0; full `npm test` 31296 tests / 31281 pass /
0 fail / 14 skipped, TMPDIR unset. One caveat stated rather than smoothed over:
in that run tests/planning-snapshot.test.cjs was truncated by concurrency after
row A5 — 11 tests did not execute, which a 0-fail aggregate cannot show. Re-run
in isolation it is 87 tests / 87 pass / 0 fail, and it is untouched by this
change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-23 18:43:53 -04:00
Tom Boucher
2f86278b5e fix(#3003): opt-in mechanism for intentional deletions in worktree.cleanup-wave (#3757)
* test(#3003): failing-first suite for declared deletions in cleanup-wave

Binds the guard's opt-in before it exists, so the suite is RED against next.

The rows that carry the weight are the over-authorization set: a directory
declaration must not authorize its children, a glob declaration must authorize
nothing, and a declaration must not act as a string prefix of another path.
Each of those BLOCKS, and each would PASS under a prefix, glob, or startsWith
matcher — which is how a path list quietly degrades into the boolean opt-in
#3003 explicitly rejected. The glob row matters most: declaredScopePrefix
already returns null ("matches everything") for a glob-leading pattern, correct
for the advisory it serves and catastrophic for a gate.

Also pinned: a failed deletion check blocks on its own reason rather than being
filtered into a pass; the block detail names only the undeclared residue so the
operator is not misdirected by paths that were fine; an entry with no
declaration blocks exactly as before; junk and non-array declarations do not
authorize; and a blocked entry still isolates rather than aborting the wave
(#2852, which must stay fixed).

Two advisory rows cover an interaction found while designing: git diff
--name-only includes deleted paths, so without unioning the declaration into
the #2596 scope check, authorizing a deletion would raise
SCOPE_OUT_OF_DECLARED against the very path just authorized.

A seeded property states the whole invariant the three over-authorization rows
sample: a deletion merges iff its normalized path is in the declared set.

* feat(#3003): declared deletions opt-in for the cleanup-wave guard

The deletions guard blocked the merge-back of any executor branch whose diff
removed a file, with no way to say a removal was intended. A plan that folded
one test file into a sibling could not be merged by the tool meant to merge it,
forcing a manual --no-ff outside the tool -- strictly less safe than what the
guard protects against.

A plan now declares removals in its own frontmatter (files_deleted), and that
list rides the same path files_modified already travels: plan-document parse ->
phase plan JSON -> the per-plan worktree gate -> record-agent/create
--deletions -> declared_deletions on the manifest entry -> the guard. The guard
blocks only the deletions NOT in that list.

A path list rather than a boolean, per the pinned decision: a boolean disarms
the guard for the whole entry, so an unexpected deletion riding along with a
declared one would pass unnoticed. Matching is exact after the module's shared
normalizer -- never a prefix, never a glob. Both would let one declaration
authorize a whole set, which is the mass-deletion accident the guard exists to
catch. That also means declaredScopePrefix is deliberately NOT reused here: it
returns null ("matches everything") for a glob-leading pattern, which is right
for the advisory it serves and would silently disarm a gate.

The block detail now carries only the undeclared residue, so an operator is not
sent looking at paths that were fine. A failed deletion check still blocks on
its own reason and is never filtered into a pass. A blocked entry still
isolates rather than aborting the wave (#2852).

The #2596 scope advisory unions the declaration into its declared set --
git diff --name-only includes deleted paths, so without that, authorizing a
deletion would immediately warn that the same path was out of declared scope.

Optional and additive throughout: files_deleted is absent from
PLAN_REQUIRED_FIELDS, a manifest entry without declared_deletions keeps the
original unconditional block, and omitting --deletions leaves the on-disk entry
shape untouched.

Supersedes the spent #2856 emitted-drift ack entry for execute-phase.md, the
same supersede that entry performed on #3370 and #3370 on #3324.

* fix(#3003): wire --deletions on every dispatch surface, not just one

Review found the feature inert on two of three dispatch paths. execute-phase.md
(harness inline) passed --deletions, but the orchestrator-worktree path
(executor-isolation-dispatch.md, worktree.create) and the Fleet-parallel batch
path (capabilities/claude-orchestration/fragments/execute-wave-pre.md,
worktree.record-agent) still passed only --files. A plan declaring
files_deleted would have merged on one path and been blocked on the other two
-- the exact bug #3003 exists to fix, left unfixed where most of the isolation
actually runs.

Worse, per-plan-worktree-gate.md already claimed --deletions was passed 'on the
same worktree.record-agent / worktree.create calls', which was false for both
untouched sites. A doc asserting coverage that does not exist is how a gap
survives review.

All four surfaces now pass the flag, verified by sweeping every .md under
gsd-core/, capabilities/, commands/, skills/ and agents/ that invokes
worktree.record-agent or worktree.create: each one that passes --files now also
passes --deletions. The isolation-dispatch note explains why this flag, unlike
--files, is not advisory -- omitting it does not skip a check, it blocks a
merge the plan declared.

Regenerates capability-registry.cjs, which the fragment edit made stale.

Neither newly-grown file needs an emitted-drift ack: executor-isolation-dispatch.md
sits under workflows/execute-phase/steps/ and execute-wave-pre.md under
capabilities/, both outside currentSizes()'s non-recursive scan of
gsd-core/workflows/ and agents/.

* docs(#3003): document files_deleted where a plan author will actually find it

The feature's entire user surface is one plan-frontmatter field, and the
canonical reference for that frontmatter -- docs/reference/plan-md.md, the table
that documents every other key -- never mentioned it. A field nobody can
discover ships as a field nobody uses. Adds the files_deleted row and an example
entry in all five locales (en, ja-JP, zh-CN, ko-KR, pt-BR), stating the property
that makes the opt-in safe: matching is exact per path after separator
normalization, with no globs and no directory prefixes, so a declaration can
never authorize more than it literally lists, and omitting the field keeps the
guard's original unconditional block.

Also corrects two claims in the scope-conformance how-to that this change made
false. Its opening paragraph described the recorded declared scope as
files_modified alone; declared_deletions is now unioned into that comparison.
Its "Renames are not detected specially" bullet asserted the deletions guard
blocks any entry whose diff contains a deletion, full stop -- which was the
whole point of #3003 and is no longer true. Reworked to say what now decides a
rename's fate: declare the old path in files_deleted and both halves become
ordinary paths for the advisory check, which is also why the old path needs no
separate files_modified entry.

Documentation that describes the pre-change behavior of the thing being changed
is worse than no documentation, because a reader trusts it.

* fix(#3003): close every review finding on the declared-deletions opt-in

Two independent isolated reviewers, correctness and security. Neither found a
blocker; both found real defects, and the directive treats a finding at any
severity as blocking. All of them are fixed here.

MAJOR -- the submodule worktree gate could not see a deletion-only plan.
per-plan-worktree-gate.md intersected $SUBMODULE_PATHS against $PLAN_FILES
alone, while $PLAN_DELETIONS was extracted and then never used. Before
files_deleted existed, a path had to appear in files_modified to be planned at
all, so the gate saw it; the new field plus the new docs telling authors a
deleted path needs no files_modified entry opened a hole where a plan whose only
submodule touch is a removal kept worktree isolation on -- the exact case #2772
disabled it for. Both channels now feed the intersection. Note the posture is
deliberately the OPPOSITE of the cleanup-wave guard: there the channels stay
apart because a deletion AUTHORIZATION must never be inferred; here they merge
because a safety fallback must never MISS a touch.

MAJOR -- same-wave conflict detection could not see a deletion. The planner's
implicit-dependency rule compared files_modified only, so plan A editing
src/x.ts and plan B declaring files_deleted: [src/x.ts] scored as conflict-free
and ran in parallel: one branch removing what the other is writing, which is the
sharpest conflict there is. Overlap is now computed across both channels.

MINOR (both reviewers, one root cause) -- the advisory union gave one field two
matching rules. declared_deletions was unioned into the scope list handed to
planWaveScopeConformance, which reads it with prefix-and-glob semantics. So a
field that is exact-match-only at the gate silently became wider at the
advisory: ["*.md"], inert at the gate, yielded a null prefix meaning "matches
everything" and muted the advisory completely, and ["src"] muted all of src/.
The union also activated the advisory on plans that declared no modification
scope at all, warning on every modified path. Replaced with subtraction from the
findings, gated on files_modified alone. One field, one rule, everywhere.

MINOR -- core.quotepath made the feature silently inert for non-ASCII paths.
git emits "tests/\303\251.ts" C-escaped and quoted, which never equals the
declared plain path, so a correctly declared deletion of tests/é.ts would block
forever with nothing pointing at the encoding. Both diffs now pass
-c core.quotepath=false.

NIT -- flag() consumed a following flag as a value, so --deletions --files x
swallowed --files and dropped both. Now treated as a missing declaration, which
fails closed. Fixed at both call sites; the helper is duplicated verbatim in
cmdWorktreeRecordAgent and cmdWorktreeCreate and leaving one would reintroduce it.

TEST -- one test passed for the wrong reason. "a declared deletion is in scope
for the advisory" asserted only that warnings omit the deleted path; under a
full revert the entry blocks first, warnings come back empty, and the negative
assertion passes anyway. It now asserts the entry actually merged, which is the
load-bearing half. Four regressions added, one per fix above.

Docs corrected rather than extended. The rename bullet in the scope-conformance
how-to claimed a rename whose delete side is undeclared never reaches the
advisory. Verified false: git's rename detection is on by default, so a pure
rename is a single R entry that appears in no --diff-filter=D output and was
never gated, before or after #3003. Only a rename that edits enough to fall
below the similarity threshold decomposes into add+delete. The pre-existing
sentence made the same wrong claim; this restates it correctly instead of
sharpening the error. The localized plan-md.md reference edits are reverted:
the PR template requires docs content added here to be English, and the
translations already lag by three fields, so English-only is the repo's
standing posture, not an oversight.

Agent-file size caps respected: gsd-planner.md is XL-tier by bytes but carries a
separate 49152-LF-CHAR cap asserted by four suites, so its edit is deliberately
terse and lands at 49141 with 11 chars of headroom, with the rationale moved to
docs/reference/plan-md.md, which has no cap. gsd-plan-checker.md lands at 49107
bytes, 45 under the LARGE cap. Both acks merged into the existing fragments that
already name those paths, since two ack sources may never name the same path.

* fix(#3003): decode git's path quoting instead of changing the git argv

The previous commit's non-ASCII fix turned the remote suite red: 44 failures,
42 of them "unexpected git call: -c core.quotepath=false diff --diff-filter=D
--name-only ...". The suite's git mocks match on exact argv, so adding two
flags to the deletions diff and the advisory diff invalidated every existing
fixture in tests/worktree-safety.test.cjs. Rewriting dozens of fixtures to
accommodate one flag would be paying a large Hyrum's-law bill to fix a small
defect.

Both execGit calls are reverted to their original argv. The C-quoting is now
decoded in normalizeScopePath instead, via a new decodeGitQuotedPath helper.
That is the better fix on its own merits, not merely the cheaper one: the git
argv is untouched so no fixture moves, the decode lands on the ONE normalizer
already applied to both sides of the comparison so the declared and reported
paths cannot disagree, and it holds regardless of the user's own core.quotepath
setting rather than only when we remember to override it.

A value not wrapped in a leading AND trailing quote is returned completely
untouched, so the plain-ASCII path -- the overwhelmingly common case -- is
byte-identical to before. Escapes decode to BYTES collected into a Buffer and
UTF-8 decoded only at the end, because \303\251 is two bytes forming one
character and decoding them separately yields mojibake. Malformed input never
throws: a trailing lone backslash or a short octal escape degrades to the
literal character, since one bad path must not take down a cleanup wave.

Caught while reviewing the helper: the non-escape branch pushed a UTF-16 code
unit rather than UTF-8 bytes. Git always escapes non-ASCII so its own output was
fine, but this normalizer runs on the DECLARED side too, and an author may write
a quoted path holding a literal é -- pushing 0xE9 alone is invalid UTF-8, so the
declaration would decode to a replacement character and silently stop matching.
That is precisely the failure this change removes, reintroduced on the other
side of the comparison. Now converts whole code points, surrogate pairs intact.

The other 2 failures: tests/parallel-dependent-plans.test.cjs pins the exact
unbackticked substring "files_modified overlap" in gsd-planner.md, and rewording
that comment to "declared-scope overlap" deleted it. The comment is restored
verbatim and the files_deleted change rides in the pseudocode and the Rule
sentence instead. Recorded in the ack fragment so the next contributor does not
rediscover it the same way.

Four regression tests cover the decode through the public cleanup-wave seam
(the helper is module-private): a declared non-ASCII deletion merges against a
C-quoted git report, the symmetric case where the DECLARATION is the quoted
form, an undeclared non-ASCII deletion still blocks with the residue naming the
decoded path an operator can act on, and a path merely containing a quote is
left alone. Plain ASCII was already covered and is not duplicated.

* fix(#3003): revert the leading-dash flag guard, the review nit was wrong

The remote suite came back with 2 failures, down from 44, and both point at the
same thing: tests/worktree-safety.test.cjs:7045 already pins the opposite
contract, deliberately.

  test('a flag-shaped --files value is not re-parsed as a flag', ...)
    recordAgent(['--files', '--branch'])
    -> files_modified === ['--branch']
    -> branch === 'worktree-agent-a1'  ("the real --branch value must be untouched")

So consuming the next argv element positionally, whatever its shape, is the
tested intent of this parser, not an oversight. The security reviewer's nit
claimed --deletions --files x would "swallow --files and drop both". It does
not: each flag runs its own indexOf, so --deletions records the literal
'--files' while --files independently still resolves to x. And that literal is
a path git never reports as deleted, so it authorizes nothing -- already
fail-closed with no guard at all. The guard bought no safety and silently
changed --files behavior along the way, outside this issue's scope.

Reverted at both call sites, which are byte-identical again, along with the test
asserting the reverted behavior and the docs sentence describing it. The nit is
recorded as REJECTED in the review artifact with the reasoning above, rather
than as fixed -- a finding that turns out to be wrong should leave a trace of
why, or the next reviewer files it again.

docs/CLI-TOOLS.md now states the positional-read behavior plainly instead, so
the next person meets it as documented intent rather than rediscovering it
through a red suite.

* chore(#3003): backfill changeset pr number to 3757

* test(#3003): cover parsePlanDocument's filesDeleted branch to clear the mutation gate

CI's Stryker shard for plan-document failed at 73.28 against a break threshold
of 75: 170 killed, 62 survived, 232 total. Eight of those survivors are the
filesDeleted block this issue added to parsePlanDocument, which shipped with no
direct coverage at all -- the field was exercised end to end through the
cleanup-wave tests, but the parser itself was never called with a plan that
declares it, so every mutant in the block lived.

Four tests, each pinned to specific mutants rather than written for coverage
percentage:

- absent key yields exactly [] -- kills the array-literal seed
  (["Stryker was here"]) and the `fmDeleted = true` conditional, which would
  otherwise produce ["true"]
- a scalar underscore `files_deleted:` wraps into a one-element array -- kills
  `fmDeleted = false`, the `&&` logical-operator swap, the `fm[""]` string
  mutation on the first operand, the emptied if-block, and the ternary's
  non-array branch
- an array-valued hyphenated `files-deleted:` maps element-wise -- kills the
  `fm[""]` mutation on the SECOND operand (only reachable when the legacy
  hyphen alias is the one carrying the value) and the ternary's array branch
- an empty list yields [] -- boundary case, and a genuinely distinct one from
  the absent key: [] is truthy in JS so it ENTERS the if, and only
  Array.isArray's true branch mapping over nothing produces the same []

Threshold arithmetic: 174 of 232 are needed for 75%, and these take it to about
178, so the shard clears with margin rather than landing on the line.

Every expected value was confirmed by executing the built parser before being
asserted, not inferred from reading the source.

---------

Co-authored-by: sim <sim@local>
2026-08-22 13:17:51 -04:00
Tom Boucher
4918c62d76 feat(#2845): require provenance for UI-SPEC component inventories (#3745)
* test(#2845): failing-first suite for UI-SPEC inventory provenance

Binds two shared formats before either exists, so the suite is RED against
next: the gsd-ui-checker dimension roster (asserted independently on twelve
surfaces, eight English and four translated) and the provenance-line grammar
the UI-SPEC template emits and Dimension 7 consumes.

Every parity assertion is paired with a synthetic mutation case, so the guard's
failure branch executes rather than only reading a correct tree: limit-1 (a
surface still declaring 6), limit (7), limit+1 (8), a dropped dimension, a
label that drifts on one surface only, a non-contiguous roster, a duplicated
number, and a surface that stops declaring a count at all. A seeded fast-check
property renders the roster under formatting noise (CRLF, padding, interleaved
sections) and asserts the parse round-trips and is strictly sensitive to a
dropped heading.

Assertions are on parsed typed records, never raw substrings.

* docs: normalize design-a-ui-phase how-to to American English

House style for docs/ is American English (CLAUDE.md). This file carried
colour/initialisation/initialise/artefact throughout. Spelling only — no
content change; kept separate from the #2845 feature commit so the
release-notes classifier and the hotfix cherry-pick filter see it for what
it is.

* feat(#2845): require provenance for UI-SPEC component inventories

A UI-SPEC's component inventory was treated downstream as a closed allowlist
while the document recorded nothing about whether the list had been enumerated
from the installed design system or recalled from memory. A recalled inventory
is indistinguishable from an enumerated one, so an executor complying with the
spec builds against a fraction of what the package offers, and every gate stays
green because they assert semantics rather than composition.

The UI-SPEC template gains a Component Inventory slot carrying one of two
provenance lines: the command that enumerated the list, the count it returned,
the resolved package@version and the date; or a Could not enumerate record with
a real reason. gsd-ui-researcher gains an enumeration ladder and must record
the line rather than write the list from recall.

gsd-ui-checker gains Dimension 7. An inventory with no provenance line, a count
with no command, an empty could-not-enumerate reason, or a line still carrying
the template's unfilled placeholders BLOCKs; a partial line, a line placed below
its table, or an honest negative record FLAGs; a complete line passes, and so
does a spec carrying no inventory at all, which keeps every UI-SPEC predating
the dimension validating unchanged. Whatever the verdict, an unsourced inventory
is reported as a non-exhaustive list of known-good components rather than a
closed allowlist, so the executor is never blocked from a component the spec
merely failed to mention. The checker never runs the recorded command.

The dimension count moved on all thirteen surfaces that assert it, across five
languages. Also corrects the claim in the English, Korean and Portuguese how-tos
that this checker applies a scored six-pillar rubric — that rubric belongs to
/gsd-ui-review's retroactive audit.

* chore(#2845): backfill changeset pr number to 3745

---------

Co-authored-by: sim <sim@local>
2026-08-21 11:59:56 -04:00
Tom Boucher
94bc492f57 fix(#3645): tracked-source rule for planner/pattern-mapper path resolution (#3728)
* test(#3645): failing-first agent tracked-source contract rows

* fix(#3645): tracked-source rule for planner and pattern-mapper paths

* Revert "fix(#3645): tracked-source rule for planner and pattern-mapper paths"

This reverts commit 61f05e947bbdaf3b4897240c3819d349215744fb.

* fix(#3645): tracked-source rule at the spawn seam and mapper gate

* fix(#3645): review fixes - bounded block, ack merge assertion, git wording

* fix(#3645): fit the tracked-source block under the 1168 ceiling

* chore(#3645): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-20 23:24:42 -04:00
Tom Boucher
14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00
Tom Boucher
8df5cb36c2 enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata (#3718)
* test(#2951): pin the absent-evidence provenance contract (failing first)

17 tests / 22 anchors on the deployed agent text. Measured against the parent
commit: 20 anchors fail, 2 pass. The two that pass are the sibling-integrity
guards on the package-name and in-repo-value rules -- green before and after is
their intended signature.

Refs #2951

* enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata

A claim of the form "X does not support Y" drawn from MISSING metadata -- no
python_requires, no engines field, no per-version classifier, no changelog entry,
no matching support-matrix row -- no longer earns [VERIFIED] however
authoritative the source consulted. An absence is silence about every value, so
the same evidence would "prove" both the version being ruled out and the version
being standardized on. The only route from an absence to [VERIFIED] is a positive
falsification attempt with its failing output pasted; everything short of that is
[ASSUMED], which the file already routes to "needs user confirmation before
becoming a locked decision".

Third member of the family beside the package-name and in-repo-value provenance
rules, mirroring PR #2768's shape. A present declared constraint and an
affirmatively documented incompatibility are untouched.

Closes #2951

* fix(#2951): close the allow-list ambiguity and the mutation gap review found

Findings from the isolated adversarial pass and the two-axis review, all fixed:

MAJOR (x2, one root cause) -- the absence clause and the present-constraint
carve-out gave opposite verdicts on the same evidence for the commonest real
case: a classifier list enumerating :: 3.9 through :: 3.13 with no :: 3.14. A
researcher could read the enumerated list as a "declared" positive constraint
and re-earn [VERIFIED], which is also the evasion vector. The rule now states
the decision procedure -- does the declaration bound EVERY value or only the
ones it names -- and closes the positive-reframing restatement explicitly. New
contract test pins all four clauses.

MAJOR -- 'licenses a positive falsification attempt as the route to [VERIFIED]'
asserted two independent substrings and never that the route lands on
[VERIFIED]. A mutant swapping the tag for [CITED] or [ASSUMED] inverted the
rule and survived all 17 tests. Now pinned as one joined sentence.

MINOR -- the attributable-failure test regex-matched illustrative examples
("a missing certificate, a wrong host"), so a copy-edit would break it for no
reason; relaxed to the substantive clause. The no-paraphrase guard counted only
the heading, missing the drift mode in its own name; it now also pins the core
proposition to one occurrence, and the test name matches what it checks. An
off-by-one in the new allow-list regex bound (141 actual vs 140) is fixed.

MINOR -- docs/AGENTS.md listed four of the five governed absence forms while
the agent prose, docs/COMMANDS.md and the changeset listed five; three copies
disagreeing on list membership is the drift this repo treats as a defect.

SCOPE -- removed docs/how-to/verify-a-dependency-compatibility-claim.md and its
docs/README.md index line. Both reviewers flagged them as a seventh and eighth
surface beyond the six the requester capped, and CONTRIBUTING's "Agent or skill
change" row requires only docs/AGENTS.md. The actionable four-case guidance is
retained in docs/COMMANDS.md, which is inside the approved scope.

Ack byte figures corrected for the final size: 44250 -> 46602 (+2352), 2550
bytes headroom under the LARGE cap of 49152.

Refs #2951

* docs(#2951): restore the how-to the phase gate requires

Reverses the removal in 6404b43d3. Both /code-review axes had flagged
docs/how-to/verify-a-dependency-compatibility-claim.md and its docs/README.md
index line as a seventh and eighth surface beyond the six the requester capped,
and CONTRIBUTING.md's required-docs row for an "Agent or skill change" names
only docs/AGENTS.md, so they were dropped.

gsd-phase-gate.cjs then denied gh pr create: it refuses when the recorded
enablement sequence has more than one step and the how-to quadrant is empty.
The sequence here is genuinely four steps -- run plan-phase, read the [ASSUMED]
claim, probe or cite or accept it unlocked, then answer discuss-phase's
checkpoint -- and the last step lands on a different capability's surface, so a
reference table cannot carry it. Compressing the sequence to one step to unlock
howToSkipReason would be gaming the gate, which is the same Goodhart failure
this whole change exists to close.

A machine-enforced repo gate outranks two reviewers' scope preference and my own
reading, so the page is restored and the PR body discloses the two extra
surfaces instead of hiding them. Reverting is a one-file change if a maintainer
prefers the tighter scope.

Refs #2951

* chore(#2951): backfill the changeset PR number

pr: 0 -> 3718 now that the real PR exists. The placeholder fails
scripts/changeset/lint.cjs with fail_invalid_fragment, which also blocks
lint-docs-required from consuming the fragment.

Refs #2951

---------

Co-authored-by: sim <sim@local>
2026-08-20 14:35:14 -04:00
Tom Boucher
79781e68eb enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands

Adds a deterministic resolvability probe over each PLAN.md <automated> verify
command and surfaces the nearest prior phase's proven commands to the planner
at every context window.

- src/verify-command-grounding.cts: recognizer (not a shell interpreter) that
  grounds a leading cd <literal> chain and npm --prefix <literal>, and reports
  unresolvable rather than guessing. Never executes command text.
- gsd-tools check verify-command-paths <N>: per-phase probe, wired into
  plan-phase.md before the plan-check pass.
- init.plan-phase gains prior_verify_commands, ungated by context_window.
- gsd-plan-checker: new Verify Command Path Resolvability dimension that
  reports the failing target and never prescribes a replacement.

Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs
(readdir order is not stable across platforms, so a module whose name extends
another's with a hyphen bucketed differently on Linux than on macOS).

Closes #2401

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets

Independent review found three defects in the recognizer:

- npm --prefix DIR run SCRIPT never reached the script-existence check,
  because the pattern required npm and run to be adjacent. That is the
  form the docs tell planners to prefer, so script_missing never fired
  for it. The prefix flag and its value are now stripped before matching.
- --prefix captured with \S+, so a quoted path containing a space was
  truncated to a stray opening quote and reported as a missing directory
  - a false blocker, worse than the bug this feature fixes. The capture
  is now quote-aware.
- A chained cd whose later segment was absolute concatenated instead of
  resetting, producing a nonsense path and another false blocker. The
  fold now resets on an absolute segment.

Also replaces the bespoke phase-directory regex with the canonical
phase-id helpers. Real phase directories are NN-slug, not phase-N-slug,
so the prior-command harvest matched nothing outside its own fixtures
and the planner-inheritance half of this feature was dead code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(#2401): source task blocks from the canonical sectionizer

The module carried its own copy of the <task>-block grammar - a fourth
hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps
its copy only because it needs the type= attribute the canonical helper
discards; this module never reads that attribute, so it can share the
owner outright instead of adding a test around a copy.

extractAutomatedCommands now takes task bodies from extractTaggedBlocks
and the out-of-task remainder from stripTaggedBlocks. A task-grammar
parity test pins the attributed task-name set against the canonical
helper across six awkward task shapes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2401): extract agent-file overflow to references and repair the property arbitrary

The remote matrix run came back red with 19 failures, four root causes:

- agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the
  49152 agent cap. Their bodies move to gsd-core/references/, leaving
  @-reference stubs, per the documented overflow pattern.
- The new checker dimension invoked gsd_run before the canonical
  preamble that defines it. The call is deleted outright: plan-phase.md
  already runs the probe and hands the result in as {VERIFY_PATHS}, so
  the dimension consumes that rather than re-running anything.
- fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with
  fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range.
- Three runtime-loaded files grew; acknowledged in the existing ack
  fragments that already own those bare filenames, since two ack sources
  may never name the same path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2401): regenerate golden install-tree fixtures for the new references

Adding two files under gsd-core/references/ changes what the installer
emits into every runtime's tree, so all 19 golden install-parity
fixtures went stale. Regenerated with npm run gen:install-tree; the
delta is exactly the two new reference paths per runtime, no removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2401): backfill changeset pr number to 3678

* fix(#2401): treat ~ as a home expansion only at the start of a path

Windows CI caught this on both shards; the Linux-only remote matrix
cannot see it. The dynamic-path refusal rejected ~ anywhere, and a
GitHub Windows runner's tmpdir is an 8.3 short name -
C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows
path came back unresolvable/dynamic_path.

This was a production bug, not a test artifact: any Windows user whose
project path carries an 8.3 short name, or any literal ~, silently lost
the probe entirely - every command degrading to unresolvable with no
explanation.

~ is a home expansion only at the start of a path; elsewhere it is an
ordinary literal. The check is now split: $, backtick, *, ? and newline
stay refused anywhere (substitution and globs, and the glob characters
are illegal in Windows path components regardless), while ~ is refused
only leading, tolerating one leading quote since the check runs before
quote stripping.

The prior tests only caught this on Windows because only Windows puts a
~ in tmpdir. Four new tests pin it on every platform via a fixture
directory literally named RUNNER~1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 15:21:15 -04:00
Tom Boucher
ec7e49a64c fix(#3576): repair all 43 dead references/ cites and gate the canonical resolvable form (#3596)
* test(#3576): gate shipped reference citations on the canonical resolvable form

Failing-first gate for #3576: a backticked bare references/<name>.md cite
resolves from no install location (agents, workflows, and references all
install where a bare relative references/ path is dead). The gate walks the
runtime-loaded trees the issue prescribes, strips @~/ include tokens
PER-TOKEN (a line-skip guard would miss a bare cite sharing a line with an
include — the issue-named trap), pins the genuinely relative ../ href and
canonical forms as non-offenders, and checks canonical cite targets exist.
43 offenders today across 19 files.

* fix(#3576): repair all 43 dead references/ cites to the canonical resolvable form

Every backticked bare references/<name>.md cite across the 19 shipped files
rewritten to gsd-core/references/<name>.md — the form every required_reading
block and @~/ include already uses, and the only form that resolves from any
install location. All 20 cited targets verified to exist; the one genuinely
relative href (plan-phase.md's ../references/mvp-concepts.md) is untouched
(the repair is backtick-anchored). Growth acks: new fragment for the three
first-time paths, #3206-pattern appends to the five fragments already naming
the other grown files (two ack sources may never name the same path).
execute-phase.md lands at 93,391/93,400 and gsd-executor.md at 49,150/49,152
— exactly the issue's projections; every repair fits.

* fix(#3576): drop stale default.md growth ack (nested modes file is hash-attributed, not growth-ratcheted)

Review finding: the emitted-attribution ratchet covers only top-level
workflows/ + agents/ files; discuss-phase/modes/default.md's delta is
source-attributed, so acknowledging its growth is a stale entry the
differential lane fails on.

* chore(#3576): add changeset fragment

* chore(#3576): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-17 15:24:24 -04:00
0xdhx
285cd41be0 fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites (#3435)
* fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites

Step 3 item 5b abstained on non-inferable (backstop) truths "unless
confirmed by explicit evidence" with the term undefined — its definition
lived only in the non-included gsd-core/references/honest-verifier.md,
behind a stale bare `references/` cite that 404s. Undefined, the term
falls back to presence + wiring, the exact false-pass the #1154
abstention protocol refuses.

- 5b: inline the compressed definition (a passing wired
  held-out/property-based test or directly observed behavior; presence +
  wiring never qualifies) and fix the cite. +84 B on the rewritten line;
  file lands at 49,151 of the 49,152 LARGE cap.
- verifier-phase-gates.md (already <required_reading>): gains the
  backstop-abstention reporting contract — AFK completion line
  ("complete with N unverified non-inferable checks", never silent,
  never a halt) and reason-distinctness (insufficient_spec vs manual-UAT
  human_needed). New content, no relocation of measured prose.
- 5c (line 204) and MVP-mode (line 644) bare cites repaired to
  gsd-core/references/ (+9 B each).
- Drift acks per ADR-2719 §4; two entries merge-appended into existing
  fragments (two ack sources may never name the same path).
- Changeset fragment with the sanctioned pr: 0 placeholder (post-create
  backfill).

Sibling census at next@7976b1ca0: 7 bare-cite instances in 4 agent
files; the 3 in gsd-verifier.md are fixed here, gsd-executor.md:429,439
and gsd-doc-synthesizer.md:20,176 stay with the epic #1891 follow-up.

Refs #1891

* chore(#3206): set changeset fragment pr to 3435

* fix(#3206): drop stale emitted-drift-ack entries that trip the ADR-2719 ratchet

The round's ack bookkeeping explained ripples that were already
self-attributed, so `tests/emitted-attribution.test.cjs` failed
deterministically on the PR head with 5 stale acknowledgments.

`agents/gsd-verifier.md` and `gsd-core/references/verifier-phase-gates.md`
appear directly in `git diff --name-only`, so PROVENANCE_RULES attributes
their emitted deltas without an ack; `agents/gsd-verifier.agent.md`,
`agents/gsd-verifier.toml` and `agents/subagents/gsd-verifier.md` are
derived emissions of a changed source and are attributed the same way.
None of the five entries could ever be consumed, so all five were stale.

Removed: the whole `3206-verifier-explicit-evidence.json` fragment (all
four entries) and the `#3206 append` to `0000-legacy-migration.json`.

Deliberately KEPT: the `#3206 append` to
`1955-verifier-coincidental-reliance.json`. Its `gsd-verifier.md` entry is
consumed by the size-growth ratchet, not the hash pass — the agent grew
49049 -> 49151 bytes, and `diffEmitted` treats a base-identical ack as
spent and excludes it from `ackEntries`. Reverting that append as well
turns the stale-ack failure into `1 file(s) grew without an
acknowledgment` (verified both ways locally).

* fix(#3206): compress 5b and re-acknowledge growth after rebase onto next

The rebase onto next (159145435, PR #3558) invalidated two things at once:

- #3558 grew agents/gsd-verifier.md to 49,098 bytes, leaving 54 bytes of
  LARGE-cap headroom where this PR's +102 no longer fits. Compressed the
  5b rewrite to +34 net by dropping the trailing honest-verifier cite —
  superseded by the now-inline definition; honest-verifier.md stays
  cited at the adjacent 5c line. File lands at 49,150 (2 under the cap).
- #3558 deleted tests/emitted-drift-acks/1955-verifier-coincidental-
  reliance.json, which carried this PR's growth acknowledgment, and its
  own 3409 fragment now names gsd-verifier.md. Re-homed the #3206 growth
  ack as an append to that entry (two ack sources may never name the
  same path).

The changeset is updated to match: two bare references/ cites repaired
(5c honest-verifier.md, MVP-mode verify-mvp-mode.md), the third (5b's)
superseded by the inline definition rather than repaired.

Reversion controls: restoring the uncompressed 5b fails
agent-size-budget.test.cjs (LARGE hard cap); reverting the ack append
fails emitted-attribution.test.cjs (differential attribution over the
real tree, GSD_EMITTED_BASE=upstream/next). Both re-verified green at
this tree: 213/213 (size + attribution), 992/992 across the 14 suites
reading the touched files.

Refs #3206

* test(#3206): pin the 5b explicit-evidence definition and cite resolution

* fix(#3206): pin regression tests to the shipped contract text

---------

Co-authored-by: sim <sim@local>
2026-08-16 14:14:27 -04:00
Tom Boucher
3c61b4a838 enh(#3565): sentinel/contract registry + check:contract-drift lint (#3571)
* enh(#3565): sentinel/contract registry + check:contract-drift lint

* fix(#3565): report artifact-row markers once and dedupe per marker

* docs(#3565): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-16 12:59:08 -04:00
Tom Boucher
1591454357 feat(#3409): reject shell guards that cannot observe their own failure arm (#3558)
* test(#3409): failing-first regression tests for unreachable shell guard arms

Drives the three live defects fail-first, executing the shipped workflow
snippets rather than a re-typed copy:

- G1/G2 plan-phase.md Walking Skeleton gate reads `--pick summaries_total`,
  a field that does not exist, so PRIOR_SUMMARIES is always "" and the gate
  has never fired (#3365). G2 is the load-bearing negative-space case: it
  rejects a fix that treats "no answer" as "zero" and fires unconditionally.
- G3 plan-phase.md PHASE_REQ_IDS resolves "" instead of the TBD sentinel on
  a phase with zero requirements.
- G4 complete-milestone.md's bare `cat <glob>` blocks on stdin under a
  nullglob left set by an earlier block (measured hang).

Skipped on Windows for G4 only: the FIFO-blocked-stdin mechanism is POSIX
only, and a weakened assertion there would pass vacuously.

Refs #3409

* fix(#3409): make nine shell guards observe their own failure arm

`--pick` coerces a missing field to empty string and exits 0, so the
`|| echo <default>` fallback after it fires only on a verb typo, never on
the field absence it was written for. Nine sites relied on that arm.

- plan-phase.md walking-skeleton gate: `--pick summaries_total` names a
  field that does not exist under any flag combination, so the gate has
  never fired on any project (#3365). Repointed at the existing single
  owner, `phases.list --type summaries --pick count`, which returns a real
  integer in every case including a project with no `.planning` directory.
  No new counter is added: a second one would duplicate the ownership
  ADR-3180 Decision 1 forbids. The gate now fires only on a literal "0",
  so an unanswerable query fails safe instead of entering skeleton mode.
- plan-phase.md phase_req_ids: now falls back to the documented TBD.
- The remaining seven convert to an explicit empty test.
- complete-milestone.md read all phase summaries through a bare
  `cat <glob>`; under a nullglob left set by an earlier block that is zero
  operands, so cat blocks on stdin. Guarded with the array shape the
  #3300 fix already established in review.md.

Refs #3409

* fix(#3409): guard eleven more globs that defeat their own fallback arm

The nullglob audit this issue asks for turned up the same class in files
#3300 never touched.

- Eight bare `cat <glob>` reads (transition, complete-milestone, planner x4,
  verifier, phase-researcher). With nullglob set that is zero operands, so
  cat reads stdin and blocks; measured rc=137 at 3s.
- Three `ls <glob> || echo "<message>"` sites (session-report,
  review-backlog and its generated skill). nullglob makes ls succeed
  listing the cwd, so the message never prints and the user gets a
  directory listing instead.

Guarded with `[ -e "${_ARR[0]}" ]` rather than `[ ${#_ARR[@]} -gt 0 ]`.
The count form is correct only when nullglob is set, and six of these
seven files never set it: without it the array holds the unmatched literal
pattern, so the count is 1 and the guard passes wrongly. `-e` is correct
in both worlds. review.md keeps its count guards — that block sets
nullglob two lines above them.

skills/gsd-review-backlog regenerated from commands/, never hand-edited.

Refs #3409

* feat(#3409): add the unreachable-shell-guard drift lint

A sibling of lint-planning-prompt-drift.cjs, consuming the shared
scripts/lib/drift-scan.cjs rather than copying it, wired into lint:ci.

Both detectors are one shape — a fallback arm defeated by a legitimate
success-on-empty:

- Detector A: `--pick` and `|| echo` on one line. `--pick` is the
  discriminator because "missing field renders empty at exit 0" is a
  documented CLI contract, not a heuristic. A rule keyed on gsd_run
  matched 111 lines, ~132 of them legitimate, and was rejected.
- Detector B: `cat <glob>` in command position, and `ls <glob>` whose
  exit code feeds a real fallback or an if/while head. Informational
  `ls <glob>` whose stdout is consumed (97 sites) and `|| true` failure
  suppression (~15) are not guards and never fire.

Shrink-only ratchet keyed on (file, trimmed text) with a per-pair count,
POSIX-normalized unconditionally so Windows CI cannot report everything
fresh and stale at once. Ships with a ZERO-entry baseline: every site it
can find is fixed. Exemption is the per-line `# gsd-scan-ignore: #NNN`
marker whose reason must name an issue or URL; a malformed reason reports
a distinct error rather than silently exempting. No file allowlists.

ADR-3409 records the invariant, the measurements behind both detectors,
and why the upstream `--pick` contract fix belongs to #3473.

Refs #3409

* fix(#3409): resolve review findings — typed surface, sanitized reports, tighter marker

Standards axis (blocker): the guard's tests asserted on human-readable
stdout/stderr and on free-form baseline-load prose, which CONTRIBUTING
prohibits by name. Added the typed surface it prescribes instead of
weakening the tests: a frozen REASON enum, a --json report mode,
structured loadBaseline errors, and a test locking Object.keys(REASON)
so a new reason stays three coordinated changes.

Security axis: sanitizeForReport covered every violation field but not
the baseline-load error path, which embeds raw JSON.stringify output --
that escapes nothing above 0x1f, so bidi and C1 controls reached CI logs
unfiltered. Routed through the sanitizer at the output seam.

Security axis: the scan-ignore marker accepted `#0` and a bare
`http://`. Tightened to a positive issue number and a URL with a host.
This diverges deliberately from the sibling in
tests/commit-files-pathspec.test.cjs, whose looser form was copied
verbatim; the header now records the divergence.

Security axis: G4 built its FIFO with `mktemp -u`, reserving a name
without creating it. Now created inside a `mktemp -d` directory.

Spec axis: ADR-3409 claimed a ninth site landed after the issue was
filed. git blame disproves it -- all nine predate it; the issue's hand
count missed one. Corrected. The design and test matrix still specified
B9 as a FLAG after implementation reversed it to PASS; both now record
the reversal and why.

Refs #3409

* docs(#3409): add the how-to for resolving unreachable-guard findings

Reference and Explanation are carried by ADR-3409; this is the
task-oriented quadrant CI cannot check for.

The page exists mainly for one thing the lint structurally cannot catch:
both `[ -e "${_ARR[0]}" ]` and `[ ${#_ARR[@]} -gt 0 ]` remove the glob
from the command and therefore both pass, but the count form is correct
only when nullglob is set — and nullglob is usually set in a different
block of the same file. A reference table cannot carry that; a how-to can.

Also documents the reason codes, so a reader can tell "nothing to report"
from "could not look".

No tutorial: this is a gate inside an existing CI loop, not a new entry
point a newcomer starts from.

Refs #3409

* fix(#3409): bring the touched prompt files back under their size gates

The remote run was red on 14 tests, all size/attribution, none of them
the regression suite.

- agents/gsd-planner.md was 194 chars over a 49152 cap enforced by four
  separate tests, each of which says the remedy is extraction, not a bump.
  It had 41 chars of headroom before this branch. Its `## Checkpoint
  Types` section was an unlinked, condensed duplicate of
  references/checkpoints.md, which already carries all three types and
  their XML shapes; the section now points there and keeps the three
  names and percentages inline. Net -969, margin 1010.
- gsd-core/workflows/execute-phase.md sat 2 chars under a comfortable
  margin assertion. Dropped the AUTO_MODE default: the `|| echo "false"`
  it replaced was unreachable, so the value was already sometimes empty
  on next, and its only consumer compares against `true`. Net -16.
  Left plan-phase.md's AUTO_CHAIN default alone -- that file names an
  explicit `false` branch, so empty would match neither branch.
- Acknowledged the seven prompt files that genuinely grew, one specific
  reason each. Five of those paths were already claimed by spent
  fragments identical to next, which blocks a second source naming the
  same path; removed just the colliding key from each, deleting the two
  that this emptied.

Refs #3409

* test(#3409): extract the whole PHASE_REQ_IDS block, not just its first line

G3 failed on the remote runner with '' !== 'TBD'. The test was wrong, not
the workflow.

The shipped contract is now two consecutive lines -- the capture and the
`${PHASE_REQ_IDS:-TBD}` default -- but the helper's `^PREFIX=.*$` regex
returns only the first match, so the test executed half the contract and
correctly observed the empty string. Renamed to extractAssignmentBlockFor
and taught it to consume the contiguous run of lines sharing the prefix.

The assertion is untouched: TBD is the right expectation, and weakening
it to accept the empty string would have reinstated exactly the class
this suite exists to catch -- a check that cannot observe the thing it
is checking.

extractFencedBashAfterAnchor is unaffected: it is fence-delimited rather
than line-anchored, so G1/G2/G4 still capture their full blocks.

Refs #3409

* chore(#3409): drop a spent ack fragment that collided on complete-milestone.md

#3458 landed on next while this branch was in flight and its fragment
claims complete-milestone.md, which this branch also grows. Two ack
sources may never name the same path.

Its entry is spent: the +9163 it explains is already absorbed at base, so
it can no longer clear anything, and the checker's own guidance for spent
entries is to delete them. Removing the key emptied the fragment, so the
file goes too -- an empty one signals nothing.

Refs #3409

* chore(#3409): backfill changeset pr number 3558

* test(#3409): hoist a regex subject out of exec() to clear the injection scan

CI's prompt-injection scan flagged `MARKER_RE.exec('# gsd-scan-ignore: ...')`.
The pattern `exec[[:space:]]*\(["']` is receiver-blind on purpose, so it
catches `require('child_process').exec('...')` -- and the scanner's own
header records that RegExp.prototype.exec is collateral, to be handled by
its allowlist.

Allowlisting the file would blind it to the real exec vector permanently,
so the subject is hoisted into a const instead: same assertion, scanner
left at full strength, no security surface widened.

Refs #3409

---------

Co-authored-by: sim <sim@local>
2026-08-15 21:09:05 -04:00
Tom Boucher
8fc88f663d fix(#3210): gate unmet preconditions as blocking-human; cap blocker retries at needs_human (#3528)
* fix(#3210): gate unmet preconditions as blocking-human and cap blocker retries at needs_human

* chore(#3210): add changeset fragment for PR #3528

* fix(#3210): restore blocking-human carve-out and CRLF-safe split

---------

Co-authored-by: sim <sim@local>
2026-08-14 23:01:30 -04:00
Tom Boucher
d2fa696a30 fix(#3479): treat absent default-true mempalace keys as enabled in every prose gate (#3527)
* fix(#3479): treat absent default-true mempalace keys as enabled in every prose gate

The #2982 absent-key fix (capture_artifacts === false) was applied to only
one of the sibling gates. Five more hand-written gates in the mempalace
skill/command mirrors and the curator agent still used positive presence
('when <key> is true'), silently skipping default-enabled behavior
(mirror_kg, diary_journal) whenever the key was absent from
.planning/config.json — inverted from the registry-declared defaults.

Corrected sites, each now disabled only on an explicit false:
- skills/gsd-mempalace-capture/SKILL.md step 3 (mirror_kg)
- commands/gsd/mempalace-capture.md step 3 (mirror_kg)
- skills/gsd-mempalace-recall/SKILL.md step 3 (mirror_kg)
- commands/gsd/mempalace-recall.md step 3 (mirror_kg)
- agents/gsd-mempalace-curator.md tasks 1+2 (diary_journal, mirror_kg)

Default-false keys (mempalace.enabled, cross_project_tunnels) keep their
positive-presence gates. New #3479 regression cases in
tests/mempalace-capture-gate-default.test.cjs lock each site's absent/explicit-false
boundary and add a registry-parity guard: no gate file may positively gate
any mempalace boolean whose registry-declared default is true.

* chore(#3479): acknowledge curator size growth from the gate rewording

* chore(#3479): add changeset fragment for PR #3527

---------

Co-authored-by: sim <sim@local>
2026-08-14 23:01:15 -04:00