Commit Graph

223 Commits

Author SHA1 Message Date
Tom Boucher
e48eb44003 fix(#1856): give the executor-worktree refusal a handoff instead of a dead end (#2727)
* test(#1856): failing-first contract for the orchestrator cwd-drift guard handoff

The #48 guard correctly refuses to execute waves from an agent worktree, but the
refusal is a dead end: that worktree can hold committed fixes AND uncommitted
work, and "re-run from the orchestrator's own worktree" silently means
abandoning them. The reporter was left choosing between continuing from a
blocked worktree and losing the work.

The guard is shell embedded in execute-phase.md, so these tests extract the
block by a stable marker and EXECUTE it against real git fixtures — the shipped
text is the runtime contract. Covers the stranded-commit and dirty-tree report,
the integration commands, both agent- namespaces, commit-count boundaries 0/1/2,
and the constraints the guard's own comment records: it must NOT fire on an
ordinary branch, on 'agentic-refactor', or on a legitimate feature worktree under
.claude/worktrees/, and must degrade cleanly with no resolvable base or a
detached HEAD.

RED expected: the marker does not exist, so extraction fails and every case errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#1856): give the executor-worktree refusal a handoff instead of a dead end

The #48 cwd-drift guard correctly refuses to execute waves from an agent
worktree — its comment records why ("this is how a wrong-base merge nearly
shipped ~1000 files"), and that refusal is untouched here. The defect is that it
was a dead end.

At the moment it fires, the worktree can hold committed product work, uncommitted
product and planning changes, and the live gap-planning context. Telling the user
to "re-run from the orchestrator's own worktree" silently means abandoning all of
it, because the orchestrator worktree cannot see commits that live only on the
agent branch. The reporter was left choosing between continuing from a blocked
worktree and losing five commits plus uncommitted work.

The refusal now reports what is actually stranded — the commit count and log
against the resolved base, and the uncommitted files — followed by the concrete
integration sequence (commit here, switch to an orchestrator-safe checkout,
merge or cherry-pick, re-run) and a verify command. Nothing is claimed that is
not there: a clean worktree with no commits ahead prints the plain refusal with
no empty sections.

Every added command is diagnostic and `|| true`-guarded, so a failure degrades to
the original refusal rather than crashing before the message prints. Verified: an
unresolvable base still refuses cleanly.

Deliberately NOT done: auto-merging or auto-cherry-picking the agent branch. That
is precisely the operation #48 exists to stop the orchestrator performing from a
drifted cwd, at the moment it has least confidence about which tree is which.
Reporting beats acting here.

The guard block carries a `gsd:guard=orchestrator-cwd-drift` marker so the new
contract test can extract and EXECUTE the shipped shell against real git
fixtures rather than asserting on its characters.

Fixes #1856

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#1856): report true counts, add the changeset, and document the seam

Review findings, all from the isolated adversarial pass:

- The dirty-file list was capped at 20 with no indication, so a worktree with 27
  uncommitted files reported 20 — under-informing the user about exactly what is
  stranded, which is the entire point of this change. Both lists now count BEFORE
  truncating and print "… and N more". Verified with 25 commits / 27 dirty files.
- The has-commits condition was written out twice and could drift on a future
  edit. Collapsed to a single _WT_HAS_COMMITS flag.
- The changeset fragment existed but was untracked, so it was in neither commit
  on this branch and the PR gate would have failed against real history.
- CONTEXT.md:122 documents this exact seam ("the orchestrator runs a cwd-drift
  guard at execute_waves entry…") and was not extended. Now records the handoff
  report, that the refusal condition and exit code are unchanged, and that every
  added command is diagnostic and || true-guarded.

Verified NOT a defect, correcting the review's premise: the guard block does break
when its line endings are CRLF, but .gitattributes:2 is `* text=auto eol=lf`,
which OVERRIDES core.autocrlf and forces LF on checkout on every platform
including Windows — so the shipped file is LF there too, and the installer copies
it through Node without translating endings. The reproduction (mine and the
reviewer's) required injecting CRLF by hand. It is also not fixable from inside
the script: a \r breaks the shell parse at the block's first line, before any
#1856 code runs. Neither introduced nor amplified by this change.

Also noted and left as-is by design: the review flagged that #1856's "offer an
explicit recovery option" could be read as requiring an interactive/automated
integration rather than printed instructions. Deliberate — see the commit that
added the block: auto-merging is the exact operation #48 exists to prevent the
orchestrator performing from a drifted cwd. Called out in the PR body for a
maintainer decision rather than silently chosen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* chore(#1856): backfill changeset PR number

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 18:32:00 -04:00
Tom Boucher
1c93df04db fix(#2711): propagate the #2517 omit-on-inherit rule to all 15 unguarded workflows (#2713)
* test(#2711): derive the omit-rule guarded set from the corpus instead of a hand list

The GUARDED array was a Goodhart metric: it reported green across 15
non-compliant workflows for no better reason than that nobody had added them to
it. The guard now derives its set — every workflow emitting a model="{…}"
dispatch site must state the omit-on-inherit/empty rule — and asserts the
derivation is non-empty so a broken scan fails rather than passes.

Rule detection stays a PROPERTY check, not a template match: plan-phase.md and
execute-phase.md state it in different words and both are correct.

RED expected on 15 workflows: audit-milestone, code-review, code-review-fix,
debug, discuss-phase-assumptions, docs-update, map-codebase, new-milestone,
new-project, quick, secure-phase, ui-phase, ui-review, validate-phase,
verify-work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2711): propagate the #2517 omit-on-inherit rule to all 15 unguarded workflows

15 of the 19 model=-dispatching workflows carried no omit-on-inherit/empty
guidance — 43 unguarded dispatch sites. Each would emit model="" whenever the
bound *_model resolved empty, which is the DEFAULT state on non-Claude runtimes:
the installer writes resolve_model_ids:"omit" into ~/.gsd/defaults.json for every
one of them (references/model-profiles.md:101), and resolveModelInternal returns
"" for that case (src/model-resolver.cts:383-386) and "inherit" for opus-tier
agents and the inherit profile (:395). Both 404 on runtimes without native tier
aliases — the failure #2517 documented and fixed in one file.

Each file now carries a `<!-- #2517 model-omit-on-inherit -->` blockquote naming
its own bound placeholders and linking the canonical statement in
references/model-profile-resolution.md, mirroring the `<!-- #2508
runtime-aware-dispatch -->` block already present in all 15. The rule text lives
in the reference; the workflows carry a pointer plus the one-line instruction, so
the next revision edits one file rather than fifteen.

plan-phase.md and execute-phase.md are deliberately untouched — they already
state the rule in their own wording, and the guard checks the property rather
than a template string.

No dispatch site is edited and no placeholder renamed: the #2684 binding guard
reports the same 19 files / 60 placeholders / 0 findings before and after, which
is the independence proof that this change is additive prose only. There is no
Hyrum's-Law routing change to disclose.

Placement is span-aware. An initial pass anchored to the #2508 marker, but in six
files that marker sits INSIDE the Agent(prompt="…") string, so the new paragraph's
literal model= landed in a dispatch call span and tripped the #2284 fail-closed
Hermes projection guard (bin/install.js:3704), refusing the install. Blocks are
now anchored before the opening Agent( of the span owning the first dispatch, and
verified to fall inside no span. gen:golden exits 0 across all 19 runtimes.

Fixes #2711

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2711): cite the issue number in the changeset body and tidy block placement

Review findings from the two orthogonal passes:

- The changeset body ended (#0). Repo convention across every prior fragment
  (e.g. #2617/#2693, #2608, #2605) is that the trailing (#NNN) is the ISSUE
  number, known at authoring time; only the frontmatter pr: field carries the 0
  placeholder pending backfill. (#0) would have rendered a dead link in the
  published release notes.
- new-milestone.md glued the inserted block directly under the preceding
  paragraph with no blank line, inconsistent with the other 14 insertions.
- The derived-guard non-vacuity floor was >=17 against an actual derived count
  of 19, tolerating a silent two-file regression. Tightened to >=19.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2711): reword the omit block so it survives Hermes projection, and exempt quick.md by size

The first block wording regressed two suites on the full matrix (4 failures on
both linux-node22 and linux-node24). gen:golden passing was not sufficient
evidence — it exercises the installer's own fail-closed guard, which is
narrower than the dedicated tests.

1. tests/fix-2284-hermes-agent-delegate-task-projection.test.cjs asserts that
   the INSTALLED code-review-fix.md contains no `model=` anywhere outside a
   string literal — masked whole-file, not merely inside call spans. The block's
   backticked `model=` survived the mask. The assertion is right: on Hermes the
   projection strips the parameter because delegate_task has no per-call model
   at all, so instructing the orchestrator to "omit the model= parameter" is
   advice about a parameter that does not exist there. The block now says "the
   `model` parameter" and carries no bare `model=` token.

2. tests/prompt-injection-scan.security.test.cjs flagged quick.md at 50,164
   normalized chars against a 50,000 prompt-stuffing threshold. quick.md sits
   just under the line on next, so any insertion trips it — the situation
   review.md is already documented for in SIZE_ONLY_WORKFLOWS ("sat at 49,971
   chars — 29 below the threshold — so it was going to trip on whatever was
   added to it next"). quick.md joins it with the same justification. This is a
   size-finding exemption only: the file is still fully injection scanned, and
   every other security check still runs on it.

Because the canonical block can no longer carry a literal `model=`, the guard's
detector now accepts the `<!-- #2517 model-omit-on-inherit -->` marker as the
canonical signal, falling back to the inline-prose property for the four files
that predate it (plan-phase, execute-phase, scan, ship — all four match the
legacy branch). That is strictly stronger than word-proximity matching, and it
keeps the guard a property check rather than a template match.

Verified: derived guard 19/19 with 0 missing; the #2684 binding guard unchanged
at 19 files / 60 placeholders / 0 findings; no inserted block contains a bare
model= token; the masked-projection assertion passes for code-review-fix.md;
gen:golden exits 0 across all 19 runtimes; lint:ci exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* chore(#2711): backfill changeset PR number

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:02:26 -04:00
Tom Boucher
6fb832a93d fix(#2684): bind scan.md and ship.md model dispatch to fields their workflows resolve (#2710)
* test(#2684): failing-first guard for unbound model= dispatch placeholders

Extends the #2517 omit-on-inherit bucket with a behavioral binding guard: every
model="{X}" in a workflow must name a field that workflow actually binds — an
init-payload key (queried for real), a shell assignment, or a declared parse
field. scan.md ({resolved_model}) and ship.md ({balanced_model}) substitute
names nothing emits, so the orchestrator invents the value (ADR-1411).

Also pins the shipped reference that seeded the placeholder and instructs the
#2517-forbidden model="inherit".

RED expected on scan.md, ship.md, and references/model-profile-resolution.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2684): bind scan.md and ship.md model dispatch to fields their workflows resolve

scan.md:85 passed model="{resolved_model}" and ship.md:486 passed
model="{balanced_model}" — names neither workflow's init payload emits.
init.map-codebase emits mapper_model; init.phase-op emits no model field at
all. With no source for the substitution the orchestrator invents a value, so
model_overrides/model_policy are silently inert at both sites — the invisible
partial application ADR-1411 prohibits.

- scan.md: declare the init fields in a Parse JSON line and substitute
  {mapper_model}, matching its sibling map-codebase.md.
- ship.md: ref.agent is only known at runtime, so resolve it per hook via
  query resolve-model and dispatch with {HOOK_AGENT_MODEL}, following the
  same NAME=$(...) → model="{NAME}" convention every other shell-resolved
  dispatch in the corpus uses (code-review.md, secure-phase.md, ui-phase.md).
- Both sites now carry the #2517 rule: omit model= entirely when the resolved
  value is "inherit" or empty. A bare rename would have traded a dangling
  placeholder for model="", which 404s on non-Claude runtimes — an agent type
  absent from the profile table resolves to the empty string, which is exactly
  ship.md's case.
- references/model-profile-resolution.md was the seam: it shipped the
  copy-pasteable {resolved_model} snippet scan.md inherited, used the stale
  Task( spelling, and instructed passing model="inherit" outright. Rewritten
  to teach the real binding convention and the omit rule.

Two defects surfaced while fixing this and fixed inline rather than deferred:

1. ref.agent originates in a capability manifest, which may be third-party.
   Resolving it at runtime made this the first place that value reaches a
   shell command, so ship.md now validates its shape before interpolating and
   skips the hook when it fails — matching code-review.md's existing
   defense-in-depth pattern. Covered by a test that runs the shipped regex
   against real agent names and injection payloads.
2. The reference doc's omit example first placed a literal model= inside an
   Agent(...) comment, which tripped the #2284 fail-closed Hermes projection
   guard (bin/install.js:3704) and refused the install outright. Moved out of
   the call span; gen:golden is green across all 19 runtimes.

Guard tests extend the existing #2517 bucket: every model="{X}" in a workflow
must name a field that workflow binds — an init-payload key queried for real,
a shell assignment, or a declared parse field.

Behavior change (Hyrum's Law): the scan mapper now runs on the catalog-resolved
model rather than the session model. The omit path is unchanged. Disclosed in
the changeset body.

Fixes #2684

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* fix(#2684): validate capability-supplied ref.agent in-context, not in the shell

The first cut of the ref.agent guard was placed after the injection point it
was meant to close. It instructed the orchestrator to substitute the untrusted
manifest value into a shell assignment and test it there:

    HOOK_AGENT="<the ref.agent value>"
    if [[ "$HOOK_AGENT" =~ ^[A-Za-z0-9][A-Za-z0-9._-]*$ ]]; then …

Substitution happens before bash parses anything, so a manifest supplying
`x"; touch /tmp/pwned; echo "` yields three statements and runs the middle one
unconditionally — the regex fires afterwards and protects nothing.

ship.md now requires the check to run in-context, the same way the workflow
already reads activeHooks ("do NOT pipe it through a shell parser"), and to
skip the hook outright on a mismatch. Only a value that has already matched
^[A-Za-z0-9][A-Za-z0-9._-]*$ ever reaches a command line. The guard test asserts
the ordering — the in-context requirement and the absence of any raw shell
assignment — not just that the pattern rejects metacharacters, since a pattern
alone was exactly what gave false assurance here.

Also corrects the empty-string explanation in references/model-profile-
resolution.md. model_profile:"inherit" resolves to the literal "inherit", not
"" (model-resolver.cts:395), and an unknown agent takes the empty-string path
via resolve_model_ids:"omit" rather than by absence alone — the doc claimed all
three produced "". The workflow instructions were already correct (omit on
"inherit" OR empty); only the rationale in the citable reference was wrong.

Both found by the isolated adversarial review pass for #2684.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

* chore(#2684): backfill changeset PR number

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 13:51:16 -04:00
Tom Boucher
2c44241a0b fix(#2605): make dropped local-server reviewer lanes loud instead of silent (#2689)
* fix(#2605): make dropped lm_studio / llama_cpp reviewer lanes loud, not silent

The `lm_studio` and `llama_cpp` reviewer legs in `gsd-core/workflows/review.md`
carried the same empty-output defect as the claude/gemini legs (#2494, fixed in
#2592) in a worse variant: when the local endpoint was unreachable or returned
empty content, they wrote NOTHING to `{run_dir}/gsd-review-<leg>.md`. There was
no `[ ! -s … ]` stub at all, so the file never existed, `write_reviews` omitted
that reviewer's section, and the outcome was indistinguishable from the reviewer
never having been selected.

Two diagnostic holes are closed, because an OpenAI-compatible server fails in two
ways that leave evidence in different places:

- Transport failure (endpoint unreachable): curl writes to stderr and exits
  non-zero. Both legs used `curl -s`, which suppresses curl's ERROR text as well
  as the progress meter, and then discarded stderr to `/dev/null` — so there was
  nothing to capture even in principle. Now `-sS` with stderr to a `.err`
  sidecar, matching the claude/gemini/codex legs.
- Application failure (HTTP 4xx/5xx): curl exits 0 and the error JSON is in the
  response BODY, so stderr is empty and only the body is evidence. The stub
  appends the raw response. The `llama_cpp` leg additionally piped curl straight
  into `jq`, throwing the body away before anything could inspect it; the
  response is now captured to a variable first, as the `lm_studio` leg already
  did.

The existing `>&2` warning is kept and now points at the stub file.

Failing-first verified by extracting both shipped blocks and running them under a
real bash against a stubbed curl: pre-fix, all three failure modes produce NO
review file and no `.err` sidecar for both legs; post-fix, each produces a
diagnosable stub, and a successful review still passes through untouched.

Closes #2605

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* fix(#2605): close the remaining silent-drop paths the review surfaced

Four defects found by the orthogonal review of the first commit, all fixed here
rather than deferred (they are the same defect class the issue is about, and
three of them sit inside the code that commit touched).

1. CodeRabbit leg was still unguarded. `coderabbit review --prompt-only
   2>/dev/null > file` with no `[ ! -s … ]` stub — the last leg still shaped like
   pre-#2494 code. A missing or unauthenticated binary left a zero-byte file that
   write_reviews rendered as "ran cleanly, nothing to report". It now captures
   stderr to a `.err` sidecar and emits the same diagnosable stub as every other
   leg. This is the leg the next issue in the #2494 -> #2592 -> #2605 series
   would have been about.

2. The budget-skip path dropped the lane just as silently. When
   `prepare_trimmed_prompt_for_reviewer` fails, `*_SKIP=1` bypasses the entire
   block — guard included — so no file was written and the only trace was a
   stderr warning nothing persists. All three local-server legs now write a
   "review skipped: prompt budget too small" stub on that path.

3. Whitespace-only replies evaded the guard. `[ ! -s … ]` counts BYTES, and
   command substitution strips trailing newlines but not spaces, so a reply of
   `"   "` was written out and passed as a "successful" but vacuous review — the
   same indistinguishable-from-success outcome the guard exists to prevent. A
   `case` glob now normalizes whitespace-only content to empty.

4. `echo "$VAR"` swallowed option-like content. bash's builtin `echo` treats a
   value of exactly `-n`/`-e`/`-E` as a flag and writes 0 bytes, which would trip
   the empty guard and DISCARD a genuine reply. Switched to `printf '%s\n'`, the
   idiom the OpenCode leg in this same file already uses for this reason.

Also brings the Ollama leg to parity while it is in hand: it always emitted a
stub so it never silently vanished, but it was the least diagnosable of the three
local-server legs — bare `-s`, stderr to /dev/null, and the response piped
straight into jq so the error body was discarded unread.

Verified by extracting all four shipped blocks and running them under a real bash
against stubbed CLIs: 22 cases (7 per local-server leg x 3, plus CodeRabbit) all
produce the contracted output, and a successful review still passes through
untouched on every leg.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* test(#2605): regenerate golden install-parity fixtures for the review.md edit

The remote test run on the prior commit failed with 19 "golden parity —
<runtime>" mismatches. review.md is installed into every runtime's tree, so
editing it changes its content hash in all 19 golden fixtures.

Regenerated with `npm run gen:golden`; the diff is exactly one line per fixture
— the gsd-core/workflows/review.md hash — and nothing else.

This is the second ripple of a workflow edit, alongside
tests/workflow-size-baseline.json which the first commit already updated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* chore(#2605): backfill changeset PR number (#2689)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 00:31:17 -04:00
Tom Boucher
c87f6f358e enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#1854): offer restore for user-added files backed up on update

Adds a restore-custom-files gsd-tools verb and wires it into update.md as a
restore_custom_files step: plan, compatibility-check against the newly
installed release, then restore only on explicit opt-in. The backup is never
deleted, a shipped path is never overwritten, and a single unwritable entry
does not abort the rest.

Also drops the jq pipe from update-context field extraction (#2589 class,
missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): reject symlinked restore destinations and backup roots

Self-review of the restore path found two write-through holes: copyFileSync
follows a symlinked destination, so a link planted at the restore target wrote
outside the config dir with every ancestor still a real directory; and statSync
on the backup root followed a link, letting the walk read arbitrary files and
present them as the user's own backup. Both now lstat.

Also marks the report's path/detail strings as untrusted data in update.md so
the rendered step cannot carry instructions into the runtime model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1854): move the update-context jq guard into the #2589 sweep

update.md joins the AUDITED list rather than carrying a duplicate assertion in
the backup-restore suite, and the guard gains a negative-proof companion so
'no jq pipe' cannot pass by the fields simply no longer being read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): validate manifest files map shape before trusting it

Security review flagged that Object.keys on a non-plain-object files field
yields numeric-index keys matching nothing, so the managed-path check dies
silently while manifest_found still reports true. Shape, not just type
(ADR-227): an array or scalar files map is now an unusable manifest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): size the restore prompt by eligible_count

Spec review found the prompt was driven by entries.length, so a backup holding
only blocked entries asked "Restore 1 file(s)?" when accepting would restore
zero. The question now reads eligible_count, and an all-blocked backup reports
its reasons instead of offering a choice that cannot be honored. The decline
path names the resolved backup_dir rather than the bare directory name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1854): use t.skip on hosts without symlink support

A bare return in a node:test body registers as a PASS, so the four symlink
guards silently reported green on unprivileged Windows instead of skipping.
Adds the dangling-link destination case the security review called out, and
moves outside-dir teardown to t.after so a failing assert cannot leak it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): unfence restore hint, regen goldens, widen install timeout

Three gate failures from the c99d612a5 run, all root-caused:

1. capability-registry (3): update.md's decline message put an instructional
   'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged
   fences as shell blocks. Retagged both display blocks as text and switched
   the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users
   can actually paste.

2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every
   runtime fixture moved. Regenerated; the diff is exactly those two hashes
   per fixture, no other drift.

3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's
   before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22
   lane passed the SAME commit in 12.7s. A full install measures 13-30s idle,
   so 60s was under 2x headroom and shrinks with every file added to the
   payload. Raised to 120s, matching the heavy case already in this file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1854): backfill changeset pr number to 2679

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:50:56 -04:00
Tom Boucher
bd570618d4 feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop

* fix(#2632): calibrate against the raw projection so the loop converges

* test(#2632): add closed-loop convergence guard and codify the feedback-loop rule

* fix(#2632): pair calibration samples per plan; atomic write; amend adr

* chore(#2632): backfill changeset pr to 2672

* fix(#2632): retry renameSync on transient windows errnos and clean up the temp
2026-07-26 16:28:56 -04:00
Tom Boucher
920a5f3f06 fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep (#2673)
* fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep

The reviewer/workflow config lookups resolved scalars and object fields with a
`gsd_run query <cmd> … | jq … 2>/dev/null || <default>` shape. On any machine
without jq (the default on Windows/Git-Bash) the jq stage fails with exit 127,
the failure is swallowed by 2>/dev/null + the trailing || default, and the
variable comes back EMPTY — the configured per-lane model/host/budget is
silently dropped and the lane falls back to CLI defaults with no diagnostic.

gsd-tools ships native flags that do the same job with no external dep:
  config-get <key> --raw          (strips JSON quotes off a scalar)
  resolve-model <id> --pick model (descends an object)
  resolve-execution … --pick <f>  (same)
  verification.status … --pick status

Replaced every jq-piped config/model/verify lookup across review.md (×23),
plan-phase.md, ship.md, debug.md (incl. the redundant boolean coercion — --raw
returns true/false as bare tokens natively), autonomous.md (×2),
ai-integration-phase.md (×4), and eval-review.md. The legitimate structured-JSON
jq sites that parse HTTP curl responses (.choices[0], jq -rs, jq -n --rawfile)
are untouched — only the jq-replaceable lookups moved to the native flags.

Adds tests/fix-2589-config-get-no-jq.test.cjs: a source-invariant guard asserting
no audited workflow pipes config-get/resolve-model/resolve-execution/verification.status
to jq (fails-first on the pre-fix text, passes after).

* test(#2589): update autonomous-converge jq assertion to --pick; regen golden fixtures

Two test consequences of the workflow-doc edits in the prior commit:

1. tests/autonomous-converge.test.cjs pinned the OLD jq-dependent shape
   (`verification.status … | jq -r '.status//empty'`) as the canonical routing
   contract. The test's INTENT is correct (route human validation through
   canonical verification.status) but it over-specified the MECHANISM (the jq
   pipe). Updated the assertion to match the new native --pick status shape;
   the contract being guarded (canonical verification.status read before the
   human_needed branch) is unchanged.

2. The golden-install-parity fixtures (19 runtimes) record a content hash of
   every installed workflow .md; the 7 edited workflows changed those hashes.
   Regenerated via `npm run gen:golden` (the test's own failure message
   instructs this). Only the 7 edited workflow hashes changed in each fixture.

* fix(#2589): declare jq a prerequisite for the lanes that still need it; repair test file

Three defects in the first cut of the #2589 fix:

1. tests/autonomous-converge.test.cjs was a JavaScript syntax error. The regex
   literal /...2>\/dev/null .../ left the second slash unescaped, terminating the
   literal early and parsing `null` as regex flags:
     SyntaxError: Invalid regular expression flags
   The whole file failed to load, so every assertion in it — including the #1522
   and #1526 guards — silently stopped running. Replaced with the string-compare
   form already used at line 202 for the sibling shell-snippet assertion.

2. lint:ci failed. tests/fix-2589-config-get-no-jq.test.cjs buckets into the
   capped `config` production module via its `config-get-...` effective prefix,
   making it a novel offender against the 2-file cap. The test is about workflow
   documents, not the config module, so it is renamed to
   fix-2589-workflow-jq-dependency.test.cjs (free prefix) rather than growing the
   allowlist with a module that does not actually need a 5th test file.

3. The fix deleted the repo's only jq-prerequisite declaration. review.md:244
   ("install jq if missing") was the anchor plan-review-convergence.md cites by
   line number, and it went away with the jq pipes — while the ollama, lm_studio,
   llama_cpp, opencode, and agy lanes still hard-require jq to parse HTTP
   /v1/chat/completions responses, opencode's JSONL event stream, and agy's
   conversation cache. On a jq-less host those five lanes swallow exit 127 into
   empty output: the same silent-degradation class #2589 exists to close.

   detect_clis now probes jq alongside the other prerequisites and emits
   jq:available / jq:missing, and the five dependent lanes are treated as
   undetected when it is absent, with an install hint. The six lanes that do not
   need jq stay selectable. plan-review-convergence.md now cites the section by
   name instead of a line number that moves.

Regression guards added to the renamed test file: review.md must keep the jq
probe and must name all five dependent lanes, and no workflow may cite review.md
by line number. Workflow-size baseline and the 19 install-parity goldens
regenerated for the review.md / plan-review-convergence.md edits.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* fix(#2589): decide the ship verification gate on a single verification.status read

Isolated review finding (medium). Pre-fix, ship.md captured verification.status
ONCE into $VERIFICATION and picked status / next_action / next_command off that
cached JSON with three jq calls. --pick takes a single dot-path field, so the
mechanical conversion issued three separate queries up front: three node spawns
that each re-read the phase VERIFICATION.md and re-derive the commit-time vs
mtime staleness comparison, on every ship — including the common passing path
that never uses the two message fields. It also meant the gate's verdict and the
message shown to the user were derived from three reads with no guarantee they
observed the same state.

The gate now reads `status` once and decides. The two message-only fields are
read on the blocking path only, after PHASE_VERIFICATION_INCOMPLETE is already
determined — so the passing path costs one query instead of three, and a
concurrent write between reads can no longer make the gate and its message
disagree, because the block/allow decision no longer depends on them.

Adding a multi-field --pick to gsd-tools would have collapsed this to one query,
but that changes the flag's output contract and belongs in its own change.

Regression guard in tests/fix-2589-workflow-jq-dependency.test.cjs: ship.md must
read verification.status exactly three times total, the block decision must
follow the status read, and next_action / next_command must both appear after
the blocking prose so they cannot drift back onto the passing path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* docs(#2589): document the jq prerequisite for the five reviewer lanes that need it

/gsd-review's ollama, lm_studio, llama_cpp, opencode, and agy lanes parse JSON
GSD does not produce (OpenAI-compatible /v1/chat/completions responses,
OpenCode's JSONL event stream, Antigravity's conversation cache), so they require
jq on PATH. Nothing in docs/ said so. Records which five lanes need it, which six
do not, that reading configured models/hosts/budgets no longer requires jq at
all, and what /gsd-review now does when jq is absent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ

* chore(#2589): backfill changeset pr number (#2673)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 16:17:02 -04:00
Behruz Nassre Esfahani
a40ee8a5e7 fix(#2479): drop the codex hook-trust bypass flag and its capability probe (#2536)
* fix(#2479): drop the codex hook-trust bypass flag and its capability probe

The /gsd-review codex lane emitted a hook-trust bypass flag via a
capability-probed variable (#1115). Host-harness safety classifiers deny
commands carrying the flag (23/33 sampled invocations), one denial citing
the probe itself as intent, while flagless retries succeeded 32/32 — the
flag only bypasses persisted hook trust, a first-run condition with no
steady-state value. Remove both the flag and the probe per maintainer
direction (no config key). #1115's diagnosability half — stderr to .err,
folded into the lane on empty output — is untouched; its version-gate
becomes vacuous with no flag to gate. The regression test inverts: the
literal flag is now banned file-wide in review.md (covers continuation
lines, carrier variables, and probes), alongside bans on the carrier
variable and any codex help-grep probe shape.

Fixes #2479

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2479): add changeset for PR #2536

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(#2479): changeset body in house format (bold lead) + post-rebase fixture regen

Review round 2 Minor: wrap the changeset lead clause in the required
**bold** span (reviewer-supplied text, applied verbatim). Rebased onto
current next; golden-install-parity (19 runtimes) + size baseline
regenerated — delta confined to review.md's hash/size.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 20:15:49 -04:00
Tom Boucher
6ee4349272 fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) (#2642)
* fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored)

* chore(#2537): backfill changeset pr to 2642
2026-07-25 05:30:20 -04:00
Tom Boucher
6ad30f74b6 feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler.

harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run.

Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard.

Closes #2627

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-25 01:50:20 -04:00
Tom Boucher
b97b3dad21 fix(#2517): plan-phase omits model= when *_model is inherit/empty (port execute-phase fix) (#2634)
* test(#2517): guard plan/execute-phase omit model= when *_model inherit/empty

* fix(#2517): plan-phase omits model= when *_model is inherit/empty (port execute-phase fix)

* chore(#2517): backfill changeset pr to 2634
2026-07-24 23:44:12 -04:00
Tom Boucher
c2b498e452 fix(#2498): pass --no-track when branching off origin/$DEFAULT_BRANCH (#2628)
* test(#2498): guard branch-creation uses --no-track (all workflows)

* fix(#2498): pass --no-track when branching off origin/$DEFAULT_BRANCH

* chore(#2498): backfill changeset pr to 2628
2026-07-24 22:49:29 -04:00
Tom Boucher
57b2bd8368 fix(#2491): finish todos/done -> todos/completed rename (14 stale refs) + guard (#2626)
* test(#2491): add todos/done rename under-sweep guard

* fix(#2491): finish todos/done -> todos/completed rename (14 stale refs)

* chore(#2491): backfill changeset pr to 2626

* test(#2491): fix lint-legacy-dir-name + allow-test-rule-refs (split legacy token, add issue ref)
2026-07-24 21:36:33 -04:00
0xdhx
bd18a593a9 fix(#2494): capture stderr and guard empty output on the claude and gemini reviewer legs (#2592)
* fix(#2494): capture stderr and guard empty output on the claude and gemini reviewer legs

The gemini and claude blocks in gsd-core/workflows/review.md were the only
two of the ten prompt-fed reviewer legs with both `2>/dev/null` and no
empty-output guard. Any failure that wrote no stdout — CLI missing,
unauthenticated, rate-limited, crashed — left a zero-byte review file with
the only diagnostic evidence already discarded.

write_reviews substitutes each file's raw content into a `## <Reviewer>
Review` section with no marker separating "empty/failed" from "ran cleanly,
nothing to report", so a failed lane silently degraded the advertised
N-reviewer consensus to N-1 while present_results reported success. Both
/gsd:review and /gsd:plan-review-convergence share the invoke_reviewers
step, so both were affected.

Both legs now redirect stderr to a `.err` sidecar and write a diagnostic
stub with the captured stderr appended when the review file comes back
empty — the same shape the codex and cursor legs already use.

Scope limit, noted in the block comment: the guard only runs if the block
itself completes. A host Bash-tool timeout that kills the whole block skips
it, the same hard bound the OpenCode block already documents. The existing
timeout guidance (#2194) covers that case and is unchanged.

Regression test extracts the two dispatch blocks verbatim from the workflow
and runs them under bash against a failing CLI stub, asserting the review
file is non-empty and carries a diagnosable message plus the captured
stderr. It fails against pre-fix review.md (4 of 5 cases).

Golden install-parity fixtures and the workflow size baseline are
regenerated: one review.md hash per fixture, one size entry.

Fixes #2494

* fix(#2494): stamp the changeset fragment with the filed PR number

The fragment was authored with the documented `pr: 0` placeholder because
the PR number does not exist until `gh pr create` returns. Now that this
PR is #2592, stamp it — scripts/changeset/parse.cjs requires pr > 0, so
the placeholder would fail the Changeset Required check.
2026-07-24 12:50:34 -04:00
Tom Boucher
3b15a1e3cc fix(#2576): normalize padded vs unpadded resolves_phase in close_phase_todos (#2597)
* test(#2576): add failing regression for padded resolves_phase compare

* fix(#2576): normalize padded vs unpadded resolves_phase in close_phase_todos

* chore(#2576): backfill changeset pr to 2597

* test(#2576): drop fast-check property test (cross-platform-fragile on Windows CI)
2026-07-24 12:10:32 -04:00
Tom Boucher
7d298d6d4d fix(#2474): gate worktree dispatch on project-level USE_WORKTREES too (#2561)
* test(#2474): update dispatch gate test for dual-gate behavior

The #2772 test asserted the gate reads USE_WORKTREES_FOR_PLAN only.
Update to accept the dual-gate (USE_WORKTREES + USE_WORKTREES_FOR_PLAN).

* fix(#2474): gate worktree dispatch on project-level USE_WORKTREES too

The per-plan dispatch condition checked only USE_WORKTREES_FOR_PLAN
(submodule-derived), ignoring the project-level USE_WORKTREES flag.
Add USE_WORKTREES to the gate. Net-negative edit: compress two
nearby prose lines to offset the added shell condition (93353 bytes,
down from 93368).

Closes #2474

* docs(#2474): backfill changeset PR number (2561)

* fix: merge coverage gate into single-process check (#2474)

The test:coverage:unit script chained two c8 invocations with &&:
the first ran tests and wrote coverage data to .nyc_output/, the
second read that data for per-file branch checks. On fast CI runners
(ubuntu/24), the second process started before the filesystem flushed
the first process's writes — a classic TOCTOU race that caused
intermittent coverage gate failures.

Replace the two-process chain with a single c8 invocation that
generates both text and json-summary reports, followed by a Node
script (scripts/check-coverage-gate.cjs) that reads the JSON summary
once and checks both overall and per-file thresholds. No filesystem
race is possible because the JSON report is fully written before the
check script reads it.
2026-07-23 09:14:02 -04:00
Tom Boucher
0ad3c5dd14 fix(#2279): refresh date stamps on map-codebase Update runs (#2550)
* fix(#2279): reword date stamping to overwrite existing dates on Update runs

The map-codebase agent and workflow instructions only said to replace
[YYYY-MM-DD] placeholders, but Update-path files already contain concrete
dates from the prior run. Reword to SET the date stamps unconditionally,
overwriting whatever date is already there.

Closes #2279

* docs(#2279): backfill changeset PR number (2550)
2026-07-23 07:37:37 -04:00
Tom Boucher
200daa456c fix(#2269): add --files to three unscoped query commit call sites (#2549)
* test(#2269): regression test for query commit --files scoping

Verify secure-phase.md, validate-phase.md, and next.md all pass --files
to their query commit calls.

* fix(#2269): add --files to three unscoped query commit call sites

secure-phase.md, validate-phase.md, and next.md were the only 3 of 65
query commit call sites that omitted --files, causing blanket staging
of .planning/ and committing unrelated files. Add --files with the
specific artifact path to each.

Closes #2269

* docs(#2269): backfill changeset PR number (2549)
2026-07-23 07:37:15 -04:00
Tom Boucher
77bf21b3a6 fix(#1995): widen worktree branch regex to accept agent-<id> namespace (#2548)
* test(#1995): regression test for agent-<id> branch namespace

Add failing-first tests proving that normalizeCleanupManifestEntry and
planWorktreeRecordAgent reject Claude Code's current agent-<id> isolation
branches (only worktree-agent-<id> is accepted). Boundary tests cover both
namespaces plus rejection cases.

* fix(#1995): widen worktree branch regex to accept agent-<id> namespace

Claude Code's isolation="worktree" branch naming changed from
worktree-agent-<id> to agent-<id>. Widen the regex in all 7 locations
from ^worktree-agent-[A-Za-z0-9._/-]+$ to ^(worktree-)?agent-[A-Za-z0-9._/-]+$
so both namespaces are accepted. Introduce a shared WORKTREE_AGENT_BRANCH_RE
constant in src/worktree-safety.cts to prevent future drift.

Closes #1995

* fix(#1995): update workflow guards, test assertions, and baselines

Widen the branch-check regex in execute-phase.md and execute-plan.md.
Update all test assertions that checked for ^worktree-agent- to expect
the widened ^(worktree-)?agent- pattern. Regenerate golden-install-parity
fixtures, agent-size-baseline, and workflow-size-baseline.

Closes #1995

* fix(#1995): update extractCwdGuardBash sanity check for widened regex

The e2e test's sanity check verified the extracted bash block contained
'worktree-agent-'. After widening to '(worktree-)?agent-', update the
check to match the new pattern.

* fix(#1995): widen missed workflow-guard branch check + changeset + lint fixes

- hooks/gsd-workflow-guard.js: widen startsWith('worktree-agent-') to
  /^(worktree-)?agent-/ regex — same defect class, was missed in prior commit
- tests/worktree.test.cjs: fix indentation regression from prior edit
- Add .changeset/1995-worktree-agent-branch-namespace.md (pr:0 placeholder)

Found by orthogonal code review (Step 4).

* fix(#1995): regenerate golden + size baselines for workflow-guard change

* docs(#1995): backfill changeset PR number (2548)
2026-07-23 07:36:53 -04:00
Tom Boucher
f654c24a3e feat(#2505): Phase 4 — runtime-aware subagent dispatch (Option A; resolve-dispatch-type query) (#2525)
* feat(#2508): Phase 4 Option A — runtime-aware subagent dispatch via resolve-dispatch-type query (#2505)

* fix(#2508): prose-variant preamble (avoid scanner-tripping literals) + namedDispatch===false-only mapping

* fix(#2508): remove leftover old-preamble lines (keep prose variant only)

* fix #2508: prose-only reference file

* test #2508: regen golden install parity after workflow preamble additions

* fix #2508: remove preamble from plan-phase.md (Phase 6 capstone ceiling); regen size+golden baselines

* docs(changeset): backfill PR #2525 for Phase 4 (#2508)
2026-07-22 10:22:42 -04:00
Tom Boucher
09b535ac00 feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.

Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.

Every per-host value is documentation-sourced, never inferred:
- claude   argv  -- verified via 'claude --help' (--effort <level>)
- opencode argv  -- verified via 'opencode run --help' (--variant)
- codex    argv  -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
                    flag (config.toml key only), so the global -c override is the
                    only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
                    fails closed rather than inheriting a profile baseline

No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by 8f2ebbe9b (#1928, PR #1996),
and neither Antigravity CLI nor ZCode documents a reasoning setting.

Closes #2481

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 19:10:22 -04:00
Tom Boucher
c5e0371775 feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging

Red phase for issue #1951 (reversibility tagging: classify decisions by
undo cost, gate one-way doors behind a checkpoint:decision).

Tests assert, per the issue's acceptance criteria:
- discuss-phase CONTEXT.md template records a **Reversibility:** field with
  a rationale on captured decisions, and states it is optional
- gsd-planner @-references planner-reversibility.md and stays under the
  49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs)
- a one-way rating inserts a checkpoint:decision before the dependent task;
  reversible inserts none; costly is flagged but never blocks
- the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard)
  and inserting a checkpoint implies autonomous: false
- docs/reference/plan-md.md documents <reversibility> as optional with all
  three ratings
- --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected
  into the planner prompt, and is advertised in the command argument-hint
  and help full mode (argument-hint parity)
- the override suppresses the gate but still persists the rating
- cmdVerifyPlanStructure accepts every rating and the absent case
  (additive-validator guarantee, behavioral via runGsdTools)
- parity: thinking-models-planning.md #4 adopts the canonical three-level
  taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone
- no content loss from the planner extraction made to fit under the cap

Prose-contract assertions are Red until the implementation lands. The
behavioral validator assertions pass immediately — regression guards
proving the validator already accepts unknown optional tags.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1951): reversibility tagging — gate one-way-door decisions

Classify planning decisions by what undoing them would cost, and give a
one-way door a human beat before the agent walks through it (issue #1951,
The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way
door framing).

Acceptance criteria met:
- discuss-phase records an optional reversibility rating with a rationale
  on <decisions> entries in the phase CONTEXT.md template. Unrated
  decisions are treated as reversible, so existing phases are unaffected.
- a one-way rating makes gsd-planner insert a checkpoint:decision before
  the task that implements the decision, reusing the existing checkpoint
  mechanism -- no new checkpoint machinery.
- reversible ratings trigger no checkpoint; costly ratings are flagged in
  the plan but never block.
- the rating persists on the task as the optional <reversibility rating=>
  element. cmdVerifyPlanStructure accepts every rating and the absent
  case; the structural validator does not reject unknown optional tags.
- --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses
  checkpoint insertion for intentionally-unattended runs while still
  recording ratings -- the override changes what stops the run, not what
  the plan remembers.

Single taxonomy, not two: references/thinking-models-planning.md #4
already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is
loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the
canonical three-level vocabulary and now points at planner-reversibility.md
as the taxonomy owner, with a parity test that fails if the surfaces
diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE).

agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the
checkpoint DO/DON'T guidance was relocated verbatim into
planner-antipatterns.md -- already @-referenced from the same section for
the same topic, so the planner still loads it and nothing was dropped. A
test guards the relocation against content loss.

Files: gsd-core/references/planner-reversibility.md (NEW, canonical
taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase
workflow/command/help (flag wiring + parity), plan-md.md schema,
discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest,
size baselines, install goldens, plugin skills regen, changeset.

Closes #1951

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1951): address orthogonal review findings

Two isolated reviewers (correctness + security), neither of which authored
the change. Every finding fixed:

Security — the rationale is untrusted input (ADR-1577). It originates in
conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each
hop an LLM reading the previous hop's output, with no validation on the
path. planner-reversibility.md and the discuss-phase template now state
it is data and never instructions, and name the </reversibility>
early-termination hazard explicitly -- a rationale that closes its own
element injects sibling structure the executor reads as real tasks.
Four tests guard it.

Correctness 1 — nothing machine-enforced the feature's own promise: a task
rated one-way with no preceding checkpoint:decision validated as fully
clean, so a planner error silently reopened the gap this feature exists to
close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A
warning, not an error: <reversibility> stays additive and the plan stays
valid. Four tests cover ungated (warns), gated (silent), still-valid, and
reversible/costly never flagged.

Correctness 2 — pass-always test. The --no-reversibility-gates parse test
substring-matched the whole workflow file, and plan-phase.md prose mentions
both tokens in one sentence, so it passed with the bash conditional
deleted: it was testing the documentation, not the parser. Now scoped to
the fenced bash blocks and matched as one physical line, with a negative
control confirming prose alone cannot satisfy it.

Correctness 3 — costly had no itemized emission rule, only one-way did, so
two agents could diverge on whether to tag costly at all.

Correctness 4 — template convention break: the example ratings were bare
while every sibling field uses [...] to signal substitution, inviting an
LLM to copy one-way/costly forward as boilerplate. Now bracketed.

Correctness 5 — latent false-green: .includes('reversible') also matches
inside irreversible/irreversibility, which appear in anti-pattern
prose, so a surface that dropped the real taxonomy entry would still pass.
Now word-boundary matched.

ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md
1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on
next). The ceiling may only rise for privileged host machinery, and
reversibility gating is optional-feature logic, so the wiring was slimmed
to its minimum and the explanatory prose moved to the reference files the
planner already loads. plan-phase.md is now 94400 bytes -- 119 under the
ceiling and 70 bytes SMALLER than on next, so the host loop shrank while
gaining the feature, which is what phase 6 ratchets toward. The tracer
contract (tests/tracer-bullet.test.cjs) is unchanged.

Lint — fixed an unnecessary non-null assertion in verify.cts and a
CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST-
PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture must carry the common task elements

The gated-one-way fixture built a checkpoint:decision task from the
abbreviated skeleton in gsd-planner.md, which shows only the
checkpoint-specific elements (<decision>/<context>/<resume-signal>).
cmdVerifyPlanStructure requires <name> and <action> on EVERY task
regardless of type, so the fixture failed validation for reasons that had
nothing to do with reversibility:

  errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"]

Caught by gsd-test on 14d14a39 (2 failures, both this fixture).

The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task
carries <name>/<files>/<action>/<verify> like any other. Fixture corrected
to match. Verified behaviorally against the real gsd-tools CLI across all
four cases: gated one-way (valid, silent), ungated one-way (valid, warns),
costly (valid, silent), absent (valid, silent).

Not a product defect: the validator's every-task contract is intentional
and pre-existing, and docs/reference/plan-md.md scopes its required-element
list to type=auto/tracer only because those are the elements a planner must
author, not because checkpoints are exempt from <name>.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1951): backfill changeset pr number to 2471

* fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision

Both CI failures were real defects in code this PR added, not false
positives.

CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 —
the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`,
which escapes the hyphen but not backslash, so the escape was incomplete.
It was also unnecessary: `-` carries no special meaning outside a character
class. Replaced with a complete metacharacter escape (backslash included).
Word-boundary behavior verified unchanged across all three ratings — notably
that "irreversible" prose still does not satisfy a "reversible" match, which
is the false-green this helper exists to prevent.

Prompt injection scan — the checkpoint fixture used the human-verification
child element inside <verify>. That tag name is a fake-instruction-boundary
pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed
files, so copying the shape from tests/verify.test.cjs (unflagged only
because it is not in this diff) tripped the gate. Switched to the documented
plain-prose <verify> form.

The first attempt at that fix failed the same gate a second time: the
comment explaining the collision quoted the offending tag literally. The
comment now names it in prose instead — the scanner does not care whether a
match is code or commentary, which is the whole point of the
DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md.

Verified locally before push: scan reports 0 findings across 57 changed
files, eslint clean, and both fixtures still validate as designed (gated
one-way silent, ungated one-way warns, neither errors).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): record measured cost and halve gsd-tools spawns

The Windows shard 1/3 job timeout was traced to the sharding layer, not to
this PR's assertions — see #2472. Two contributing factors were this file's
own, and are fixed here.

1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs,
   so scripts/run-tests.cjs weighted it at the table's median fallback
   (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x
   under-weight. Recorded the measured value from the green gsd-test run
   (max across the node22/node24 lanes, per gen-test-timings.cjs's
   convention). Only this one entry: a full regen churns 634 entries of
   run-to-run drift, and the table is explicitly advisory and un-gated, so
   a 637-line diff does not belong in a feature PR.

2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost.
   Spawns cut from 9 to 6 with no coverage lost:
   - the ungated-one-way warning and its stays-valid assertion now share
     one plan instead of building the same plan twice;
   - the reversible/costly never-flagged-as-ungated test was strictly
     subsumed by the additive suite, which already runs those two ratings
     ungated and asserts no /reversibilit/ warning at all — and the gate
     warning's text contains both "reversibility" and "one-way", so the
     broader assertion catches it. It only re-spawned gsd-tools twice to
     prove the same thing.

Both are symptom fixes. The shard imbalance itself (19/11/10 minutes
against a 20-minute cap, from a cost-blind round-robin partition that also
reshuffles downstream files whenever one is inserted) is tracked in #2472.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1951): checkpoint fixture adopts the #2444 type-branched contract

Surfaced by rebasing onto next, which gained #2444 (branch plan-structure
validation on task type=checkpoint:*) while this PR was in review.

cmdVerifyPlanStructure no longer applies one required-element set to every
task. A checkpoint:decision now requires <name> + <resume-signal> +
<decision> + <options>, and is exempt from the <action>/<verify>/<done>/
<files> set that auto and tracer tasks carry. The gated-one-way fixture
predated that split and failed on the new requirement:

  errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"]

Fixture rewritten to mirror the checkpoint:decision contract exactly — real
<options> with two <option> children — rather than padding it with fields
checkpoints no longer need. That also drops the plain-prose <verify> the
earlier revision carried purely to dodge the prompt-injection scan; a
checkpoint task has no <verify> requirement at all, so the workaround is
moot.

Verified against the real gsd-tools CLI across all four cases: gated one-way
(valid, silent), ungated one-way (valid, warns), costly (valid, silent),
absent (valid, silent).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 10:44:55 -04:00
Tom Boucher
bb97ffb5aa fix(#2431): self-suppress TDD Audit section when all commits are missing (#2467)
* fix(#2431): self-suppress TDD Audit section when all commits are missing

The TDD Audit section in ship.md step 8 was always emitted — but the
execute pipeline only writes gate_status: git trailers when TDD mode is
active. Without TDD mode (the default), every commit's trailer is absent
and the section normalizes to 100% missing, producing a noise table with
no way to disable it.

Fix: add a self-suppress instruction at the point where gate_status
values are normalized. When every commit in the scan normalizes to
'missing', skip both step 8 (TDD Audit section) and step 9 (aggregate
gate_status trailer) entirely. Only emit when at least one commit carries
a real value (skill, fallback, or exempt).

This is data-driven, NOT config-gated. An earlier iteration used inline
'gsd_run query config-get workflow.tdd_mode' — but workflow.tdd_mode is
owned by the tdd capability, and ADR-857 Phase 6 forbids host loop
workflows from reading capability-owned keys via inline config-get. The
self-suppress approach avoids any config-get entirely; it checks the
actual trailer data and skips when there's nothing real to report.

Matches the triage's suggested approach: 'have the audit gracefully
degrade (skip the section)' when there is no real signal.

Tests: tests/workflow-compat.test.cjs gains 3 #2431 assertions:
- documents self-suppress when every commit is missing
- step 9 (aggregate trailer) is also gated on real values existing
- does NOT read workflow.tdd_mode inline (ADR-857 Phase 6 compliant)

The existing feat-41 assertions still pass — the section content is
preserved, only gated by the self-suppress instruction.

References: #2431; PR #585 (consumer shipped, producer never wired);
ADR-857 Phase 6 (capability-owned config keys must not be read inline by
host loop workflows).

* chore(#2431): backfill pr:2467 in .changeset/clever-moles-frolic.md
2026-07-20 19:24:08 -04:00
Tom Boucher
455ad49ae3 feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded

Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.

Red until the resolver, CLI flag, and manifest key land.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#2296): config-gated provider escalation on quota-exceeded

The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.

- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
  capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
  Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
  the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
  validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
  passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
  every model tried once the ladder is spent.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2296): extract quota recovery to a reference fragment; regen goldens

The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.

That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.

Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2351): make the C1 orphan-reaping test load-independent

tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.

The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.

Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md

* chore(#2296): regenerate fixtures after rebase onto #2402

The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (b6e6a22fc) regenerated the same
artifacts. Conflict resolution picked a side to unblock the rebase; a true
regeneration on the combined tree then produced further drift, confirming the
resolved content was stale and would have dropped #2402's fixture changes.

Regenerated goldens, install-tree, size baseline, and INVENTORY-MANIFEST from
the merged tree. docs/INVENTORY.md keeps BOTH new reference rows.

execute-phase.md is 92782 LF bytes with both #2402's and this PR's extractions
applied — under the frozen ceiling (hard <93600, margin <=93400).

Refs #2296

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 14:59:30 -04:00
Tom Boucher
b6e6a22fce fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer

Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage
(seven commits, never pushed) onto current origin/next as a single squashed commit.
The original work was substantial and correct; this commit preserves its full scope,
trimmed where rebase conflicts + workflow size budgets required it.

Three independent layers where response_language was being dropped are closed:

Layer 1 — orchestrator-facing directives across workflows. Adds the strong
"All user-facing output in this workflow MUST be presented in {response_language};
technical terms, code, paths, and subagent prompts stay in English" directive
to ~40 workflows that previously either lacked it entirely (verify-work,
new-project, new-milestone, quick, manager, and ~35 more) or carried only
the weak subagent-prompt-only form (plan-phase, execute-phase). The directive
covers narration between tool calls and banner output, not just the
AskUserQuestion prompts.

Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts
an optional responseLanguage parameter and renders the frame strings
("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.")
in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/
Chinese/Korean/Italian) with an alias table covering ~30 input variants
(en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads
config.response_language via loadConfig(cwd) and passes it through, so the
byte-for-byte block verify-work.md reprints verbatim is already localized
when written — preserving the anti-injection hygiene rule at verify-work.md
(the model is forbidden to translate after the fact). CJK display width is
computed by East Asian Width property ranges (W/F) so the right ║ border of
the banner stays aligned for full-width characters. English fallback is
byte-identical to the pre-fix behavior when response_language is unset or
unrecognized.

Layer 3 — literal English report templates in execute-phase. The top-of-
workflow directive covers all template sites (templates are a structural
source, not literal output). Inline render-language notes that previously
sat at each template site were removed during the squash because they
pushed execute-phase.md over its frozen pre-phase-6 byte ceiling
(93600 — ADR-857 Phase 6 capstone). The single top directive covers the
same surface with fewer bytes.

Also extends src/docs.cts and src/init.cts to propagate response_language
into the init JSON bundle of the additional workflows so the directive can
read it.

Tests added:
- tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls
  back to English default; recognized language swaps only the two frame
  strings while structural lines stay untouched; CJK display-width regression
  (independent recomputation of East Asian Width W/F ranges).
- tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language
  wiring through docs.cts/init.cts.

References: #2402; reporter's three-layer triage + Layer-4 follow-up; the
byte-for-byte anti-injection hygiene rule at verify-work.md (the reason
Layer 2 must be renderer-side, not model-translated).

This is a squash of the in-flight bot branch — seven commits representing
the original implementation plus its subsequent fix/CJK-padding/test/
changeset/regen cycles, none of which were ever pushed or PR'd. The squash
captures the final coherent state.

* chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md

* chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451)

Rebase conflicts were entirely in generated artifacts (golden-install-parity
fixtures + workflow-size-baseline.json). After taking theirs during rebase,
regenerated cleanly against the merged source tree.
2026-07-20 14:22:27 -04:00
Tom Boucher
352876ff0c fix(#2315): respect review.default_reviewers in bare convergence invocation (#2451)
* test(#2315): regression test for review.default_reviewers precedence

A bare /gsd-plan-review-convergence invocation (no reviewer flags) is supposed
to let users configure a persistent reviewer lineup via review.default_reviewers
and just run the loop. Instead, the orchestrator's argument parser silently
discards that configuration and forces --codex on every no-flag invocation —
with no warning that the configured reviewers were ignored.

This commit adds a regression test that fails against the pre-fix workflow
(the buggy unconditional --codex fallback is still present at this commit)
and passes after the fix lands:

- Structural: the workflow must NOT contain an unconditional
  'if [ -z "$REVIEWER_FLAGS" ]; then REVIEWER_FLAGS="--codex"; fi' line
  before the workflow.plan_review_convergence config gate.
- Structural: the workflow must query review.default_reviewers AFTER the config
  gate and document that empty REVIEWER_FLAGS lets gsd-review apply the default.
- Behavioral: matrix across {configured, unset, empty-array, explicit-flag}
  invoking the actual deployed parse + resolution blocks with a stubbed gsd_run.

Also updates two existing tests whose assertions the fix makes stale:
- #2293 behavioral: endMarker was the buggy unconditional fallback line; the
  bare invocation assertion was 'run("5") === "--codex"'. Both flip post-fix
  (endMarker is now the last --all grep line; bare invocation returns empty
  from the parse block, default applied later in step 1.5).
- command-default-claim: the pre-fix command documented '--codex (default if
  no reviewer specified)' which was the user-facing mirror of the bug. The
  assertion now requires the command to document the review.default_reviewers
  precedence.

* fix(#2315): respect review.default_reviewers in bare convergence invocation

Root cause: plan-review-convergence.md step 1 (Parse and Normalize Arguments)
contained an unconditional fallback that set REVIEWER_FLAGS=\"--codex\" whenever
no explicit reviewer flag was supplied. This value was then interpolated
verbatim into the gsd-review args, so gsd-review saw --codex as an explicit
flag (precedence rule 1) and never reached rule 3 (review.default_reviewers).
The same path silently dropped any configured review.reviewer_instances
(instances participate ONLY via review.default_reviewers per ADR-1517).

Fix:
- Remove the unconditional --codex fallback from step 1.
- Add a config-gated resolution in step 1.5 (after CONVERGENCE_ENABLED check)
  that queries review.default_reviewers and either leaves REVIEWER_FLAGS empty
  (letting gsd-review apply its own rule-3 default) or falls back to --codex
  when no default is configured — preserving the pre-fix default for
  unconfigured users (AC3).
- Replace the banner {REVIEWER_FLAGS} token with {REVIEWER_DISPLAY} so the
  startup banner reflects what will actually run (AC4), not a hardcoded value.

The fix upholds the documented precedence contract (ADR-0011, ADR-0015) that
the bug was actively violating. Explicit-flag invocations (--gemini, --all,
etc.) are unaffected (AC5).

References: #2315; ADR-0011 (review.default_reviewers precedence);
ADR-0015 (autonomous cross-AI convergence); ADR-1517 (reviewer instances).

* chore(#2315): bump plan-review-convergence.md baseline + changeset

- Bump plan-review-convergence.md size baseline 23713 → 25536 (the new
  step-1.5 default-resolution block).
- Add .changeset/plucky-yaks-roar.md documenting the user-visible change.

* fix(#2315): restore /gsd: colon syntax + skip behavioral test when jq missing

Two follow-ups to the #2315 fix discovered by gsd-test:

1. The fix commit accidentally regressed the slash-command namespace in the
   disabled-feature exit message: /gsd:plan-review-convergence (correct,
   from PR #3452) became /gsd-plan-review-convergence (retired dash syntax).
   The slash-command-namespace invariant test caught this. Restored the
   colon form.

2. The behavioral test exercises the deployed reviewer-resolution block,
   which pipes through jq. jq is a documented production dependency
   (review.md:244 "install jq if missing") and is present in every
   production deployment, but is NOT on PATH in the gsd-test linux-node{22,24}
   containers (same constraint as tests/opencode-review-reconstruction.
   property.test.cjs). Without jq, the printf|jq pipeline fails silently,
   the ||echo 0 fallback yields DEFAULT_REVIEWERS_COUNT=0, and the resolution
   falls through to the --codex branch — producing a false negative.
   Added a jqAvailable guard at module load (matching the existing pattern)
   and skip the behavioral test when jq is absent. The structural tests
   (no bash execution) still run and validate the fix.

* chore(#2315): regenerate golden-install-parity fixtures

Source changes to plan-review-convergence.md (workflow + skill mirror +
command doc) changed the install-tree hashes. Regenerated via
'npm run gen:golden' after rebuilding gsd-core/bin/lib/install-engine.cjs
('npm run build:lib') — the on-disk lib was stale relative to
src/install-engine.cts (isSymlinkedDestOptIn) and blocked fixture gen.

* test(#2315): address review findings — strengthen structural tests + property test

Code-review + security-review (isolated subagent passes) surfaced Low/Nit
findings; this commit addresses the test-side findings:

- Structural test 1 ("unconditional --codex one-liner") now asserts the
  buggy line is absent EVERYWHERE, not just before the config gate. The
  earlier assertion allowed a maintainer to re-add the line after the gate
  (passing the structural test) while the bug would still bite at runtime
  before the gate runs.
- Structural test 4 ("banner uses REVIEWER_DISPLAY") now asserts the
  LITERAL banner placeholder "Reviewers: {REVIEWER_DISPLAY}" and forbids
  "Reviewers: {REVIEWER_FLAGS}". The earlier workflow.includes(
  "REVIEWER_DISPLAY") was satisfied by a comment mention.
- New structural test for command/skill content parity: both files must
  document the review.default_reviewers precedence on the --codex flag
  (catches a manual edit to one that the gen:plugin-skills mirror misses).
- Behavioral test stub now passes default_reviewers via env var
  ($GSD_TEST_DEFAULT_REVIEWERS) instead of inline-interpolating into a
  bash single-quoted string. Removes the (currently-safe) fragility where
  a future test input containing a single quote would close the bash
  quote and execute as bash under execFileSync.
- New property test (fast-check, numRuns=25) for the JSON-classification
  contract: non-empty arrays of slugs -> empty REVIEWER_FLAGS; empty
  array / scalar JSON / malformed JSON -> --codex fallback. Locks the
  parser contract per CLAUDE.md mandate.

* fix(#2315): address review findings — defensive jq-missing warning + banner cleanup

Code-review surfaced two Low-severity workflow-side findings:

- Defensive jq-missing warning: if jq is not on PATH in production (it is
  a documented dependency per review.md:244, but the dependency can be
  absent in degraded environments), the printf|jq pipeline fails silently
  to "0" and a user with review.default_reviewers configured gets --codex
  with no indication their configured default was unreadable. Added a
  command -v jq guard at the top of the resolution that falls back to
  --codex AND emits a stderr warning explaining the reason. This makes the
  failure diagnosable instead of silently reproducing the #2315 override.

- Banner leading-space cleanup: REVIEWER_FLAGS accumulates with a leading
  space ("$REVIEWER_FLAGS --gemini" from ""), so the explicit-flag branch
  of REVIEWER_DISPLAY="$REVIEWER_FLAGS" rendered "Reviewers:  --gemini"
  (double space). Pre-existing but worth fixing alongside the AC4 banner
  work. Strip one leading space with the ${VAR# } parameter expansion in
  the explicit-flag branch only (the configured-default and --codex
  branches already produce clean strings).

* chore(#2315): bump plan-review-convergence.md baseline + regen golden fixtures

The defensive jq-missing warning grew plan-review-convergence.md
(25536 -> 26285 bytes). Bumps the per-file baseline snapshot and
regenerates the golden-install-parity fixtures for the resulting
install-tree hash changes.

* chore(#2315): backfill pr:2451 in .changeset/plucky-yaks-roar.md
2026-07-20 10:55:28 -04:00
Tom Boucher
eb45fc0e8c fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/ (#2447)
* fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/

Bug: the workflow step moved resolved todos from .planning/todos/pending/
to .planning/todos/completed/ with a plain 'mv', then committed listing
ONLY the destination directory in --files:

    mv "$TODO_FILE" "$COMPLETED_DIR/"
    gsd_run query commit '...' --files .planning/todos/completed/ .planning/STATE.md

Git's index still tracked the moved file at its old pending/<name>.md path.
The commit therefore only staged the new completed/<name>.md copy — the
deletion at pending/ was never staged, never committed, and lingered as an
unstaged deletion in git status indefinitely until some later broad
'git add -A' caught it. The phase genuinely closed the todo, but the
working tree was never clean.

Fix: add .planning/todos/pending/ to the --files list. 'git add' of that
directory (since git 2.0) stages deletions of tracked files in the pathspec,
so the moved-away file is staged as a deletion atomically with the new
completed/ copy in the same commit.

Chose plain mv + two-dir --files over 'git mv' because git mv FAILS on:
  - untracked todos (new todo file not yet committed)
  - non-git .planning dirs (worktree safety / pre-init projects)
Plain mv has neither failure mode.

Regression tests in tests/close-phase-todos-stage-deletion.test.cjs
(source-text-is-the-product: workflow .md text IS what the runtime loads)
cover:
  - the commit --files list includes BOTH completed/ AND pending/
  - the move uses plain 'mv' (not 'git mv') so untracked + non-git cases work

* chore(#2415): trim commit subject to keep execute-phase.md under byte ceiling

The fix added '.planning/todos/pending/' (~26 bytes) to the commit --files
list. To stay under the ADR-857 Phase 6 pre-phase-6 byte ceiling margin
(93400 bytes, hard ceiling 93600), shortened the commit subject from
'auto-close N todo(s) resolved by this phase' to 'close N resolved todo(s)'.
Net change vs origin/next: +6 bytes (93384 → 93390), well under the margin.

Also regenerates the golden-install-parity fixtures (execute-phase.md
content-hash update across all runtimes).

* chore(#2415): bump execute-phase.md workflow-size baseline (93384 → 93390)

The +pending/ fix added 6 net bytes (93384 → 93390), still well under the
ADR-857 Phase 6 pre-phase-6 byte ceiling margin (93400).

* chore(changeset): backfill pr:2447 in .changeset/sturdy-wasps-run.md

* fix(#2415): add issue ref to allow-test-rule annotation (ADR-456)

CI lint-allow-test-rule-refs failed on the prior commit — ADR-456 requires
'// allow-test-rule: <category> see #NNNN' so every exemption is traceable
to an issue. Added 'see #2415' to the source-text-is-the-product annotation.
2026-07-20 08:22:23 -04:00
Tom Boucher
d16a66479a feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship

Adds a new  capability (#1950) that operationalizes GSD's
no-defer discipline as a tracked, enforced artifact:
accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths
across phases, and /gsd-ship blocks while any entry is open.

Implementation:
- src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR +
  I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed
  + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed
  error assertions. Windows-safe atomic rename with retry on transient
  EPERM/EBUSY/EACCES.
- gsd-tools.cjs: new  subcommand (status | append | waive | fixed),
  wired via routeWindows + HOST_COMMAND_ROUTERS.windows.
- capabilities/broken-windows/capability.json: one ship:pre gate with
  artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0.
  activationKey windows.enabled (default true) + sibling windows.enforce
  (default true, separate so tracking can precede enforcement).
- gsd-core/workflows/ship.md: capId==broken-windows branch in preflight,
  sibling to security — reads gsd_run windows status --raw, fails closed
  on open_count > 0 or unreadable ledger.
- agents/gsd-executor.md: extends the existing ## Known Stubs instruction
  to also append to WINDOWS.md via gsd_run windows append (best-effort,
  never blocks execution).
- agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify
  items in WINDOWS.md.
- gsd-core/workflows/progress.md: surfaces open + waived counts.
- docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md:
  document the gate, waiver mechanism, and new module.
- tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check
  roundtrip property; fail-closed on malformed ledger; security boundary on
  path traversal in --file.

Backward-compatible: a project with no .planning/WINDOWS.md reports
open_count: 0 and ships cleanly. Disable enforcement per-project with
gsd config-set windows.enforce false (tracking continues, gate stays open).

* chore(#1950): ratchet size baselines, defer verifier integration

- Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632
  (broken-windows preflight branch + open-windows surface).
- Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also
  appends to WINDOWS.md). gsd-verifier.md unchanged.
- LARGE_CAP (49152) preempted the planned verifier integration
  (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the
  documented 'real headroom'). Verifier integration deferred to a follow-up
  PR that extracts the VERIFICATION.md template (lines 739-859) to
  gsd-core/references/ — a pre-existing cap-tightness defect this PR
  exposed but does not expand scope to fix. Verifier integration is not in
  the issue's acceptance criteria (executor writes is; unmet-truths
  recording was an enhancement, not a gate).

* fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens

Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44
cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all
caps off → empty hooks'):

- capability manifest: rename windows.enabled+windows.enforce (default
  true) → single federated key workflow.windows_enforce (default FALSE,
  opt-in). Matches security's workflow.security_enforce convention and
  makes the adr857 all-caps-off test pass without modification (the test's
  buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps
  the gate out of the registry's default ship:pre resolution so existing
  loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId
  'security') stay valid; users opt in via
  gsd config-set workflow.windows_enforce true.
- drop activationKey (security doesn't have one either; workflow.* key
  doubles as the activation toggle).
- regenerate docs/reference/capability-matrix.md to include broken-windows
  (capability-matrix-sync test).
- regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) —
  installer now emits the new capability + lib file.
- update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md,
  agents/gsd-executor.md to use the new key name and /gsd:colon slash
  syntax (slash-command-namespace test).
- restore accidentally-regressed /gsd:capture in progress.md.

Tracking-only by default; enforcement is opt-in. Acceptance criterion
'/gsd-ship fails while any ledger entry is open' is met when
workflow.windows_enforce=true (test fixture enables it).

* test(#1950): update ship:pre structural invariants for 2-gate registry

- loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre
  (security + broken-windows), regardless of activation. Activation tests
  above still pin security-only or empty behavior via fixtures; these
  structural tests pin the REGISTRY shape, which has 2 gates as of #1950.
- workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce
  rename added 17 bytes).

* fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker

Adversarial isolated review (Step 6.3) found 2 HIGH findings that block
the PR and 3 mediums. All addressed:

H1 (HIGH): description containing the markdown 3-backtick fence would
terminate the ledger's JSON code block early inside JSON.stringify output
(JSON doesn't escape backticks), corrupting the file and bricking the
next parse. Fix: use a 4-backtick fence (json ... ) which
JSON.stringify cannot produce on its own, AND validate that no entry
text field contains a 4-backtick run (reject at append time with new
WINDOWS_INVALID_TEXT reason code). Locked by a regression test.

H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger',
silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate
would then pass on an unreadable ledger — the precise vector the
workflow doc claims is impossible. Fix: only ENOENT returns null;
every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the
gate blocks and the operator sees a real diagnostic. Locked by a
regression test that chmod 000s a ledger with open_count=1 and
asserts the result is never a false-green 0.

M1: writeLedgerAtomic left an orphaned .tmp file on rename failure.
Wrapped renameWithRetry in try/catch with best-effort unlink.

M2: validateLine silently coerced 'abc' → NaN → null, hiding type
drift. Removed the line === 0 special case (was undocumented) and
made the error message match the strict check. Now any non-positive-
integer line value throws, including strings.

M3: tests/broken-windows.test.cjs (with its fast-check property test)
was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate
src/broken-windows.cts but no test would catch the mutations,
producing false surviving-mutant scores. Added to the list.

L1 (dead throw e after error()), L7 (line boundary tests, H1/H2
regression tests, 4-backtick CLI test) also addressed.

* docs(#1950): inline concurrency + busy-wait notes (review L2+L3)

* fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test

gsd-test v4 caught two issues:
- goldens I regenerated earlier (commit 526682084) predated the L1
  routeWindows catch-block cleanup (commit dd844d565). Regenerated
  via 'npm run gen:golden' against current HEAD so the install
  parity hash for gsd-tools.cjs matches.
- 'append --line boundary' test expected --line 0 to succeed with
  null entry.line, but the M2 fix correctly rejects 0 (lines are
  1-indexed; 0 is not a valid source line). Updated the boundary
  test to assert --line 0 fails alongside -1 and 'abc'.

* chore(#1950): regen goldens after rebase onto next

* chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase)

* chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md

* fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization)

CodeQL flagged the markdown-table cell escaper:
  String(s ?? '').replace(/\|/g, '\\|')
— it escapes pipe but not backslash first. A description containing '\|'
would render as '\\|' which markdown parses as 'literal backslash' +
'cell separator', splitting the column.

Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|).
Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped
pipe), which markdown renders as a single '\|' inside the cell. The JSON
code block (the parse source-of-truth) was already correctly escaped via
JSON.stringify; only the display-only table was affected.

Locked by a regression test that:
1. Verifies the JSON block reparses with the description intact.
2. Walks the rendered table row counting unescaped pipes — must be
   exactly 11 (the row separators for 10 cells), proving no in-cell
   pipe added a split.
2026-07-19 20:24:21 -04:00
Tom Boucher
1a46bc068a fix(#2376): emit absolute subagent-facing paths from init/state, convert workflow literals (#2428)
* fix(#2376): emit absolute subagent-facing init/state paths

Make init.* and state.* path fields absolute rather than cwd-relative
so subagent prompts resolve correctly regardless of working directory.
Adds intel_dir/conflicts_path/requirements_path/roadmap_path/state_path
to cmdInitIngestDocs, an absolute debug_dir to cmdStateLoad, and
replaces bare .planning/... literals in 12 workflow Agent() prompt
blocks with the absolute init-JSON path fields. Includes decoy-cwd
regression tests and realpath'd tmpdir fixtures for macOS.

Squashed rebase of the #2376 commit series onto a fresh origin/next
(previous merge ee25543a1 was against a now-stale next).

* chore(#2376): add changeset

* chore(#2376): regenerate golden fixtures + workflow size baseline

Regenerated after rebasing the absolute-path fix onto current next
(picks up #2351's run-with-timeout content in execute-phase.md too).

* fix(#2376): trim execute-phase.md redundancy to stay under the size margin

* chore(#2376): regenerate golden/size baseline after rebase onto next
2026-07-19 15:43:48 -04:00
Tom Boucher
d0bacc2517 fix(#2351): replace hardcoded timeout with portable run-with-timeout (#2426)
* fix(#2351): replace hardcoded gnu timeout with portable run-with-timeout

Stock macOS ships neither `timeout` nor `gtimeout` (GNU coreutils). The 10
hardcoded `timeout <n> <cmd>` calls across the workflow/agent/reference gates
exited 127 ("command not found") on such hosts, and the gates — which only
distinguish 0/124/other — misreported a passing build or test as a FAILURE.

Fix: a single Node-based `gsd_run run-with-timeout <secs> [--] <cmd…>` verb in
gsd-tools.cjs. Coreutils-independent (stock macOS AND Windows), keeps GNU
`timeout`'s exit-code contract (124 timeout, passthrough, 127/126 ENOENT/EACCES,
128+signum on signal), inherits stdio so pipes/redirects work, and reaps the
whole process group so a watch-mode runner cannot outlive its budget. Runs
before gsd-tools' flag parsing so the wrapped argv stays opaque.

Hardened per adversarial review:
- On timeout, SIGKILL the group SYNCHRONOUSLY before resolving — a descendant
  that traps SIGTERM was otherwise orphaned holding stdout, hanging captured
  gates (the exact watch-mode hang the feature prevents).
- Forward SIGINT/SIGTERM to the child tree instead of dying and orphaning it.
- Reject blank/whitespace <seconds> (was a silent unbounded run); clamp the
  timer to the 32-bit setTimeout ceiling (was a spurious immediate timeout).
- Lint detector: catch GNU long options / `-k5` / `$((...))`; anchor to command
  position so prose "timeout 30 seconds" no longer false-positives.

Resolution lives once in the CLI; all 10 sites call the shared verb. A parity
guard (scripts/lint-portable-timeout.cjs, wired into lint:ci) fails the build if
a bare `timeout`/`gtimeout` execution reappears (the portable `command -v
timeout` probe form is intentionally allowed). Also fixes the identical bug in
the zh-CN checkpoints translation, updates the tests that asserted the old
strings, trims a redundant phrase in gsd-verifier.md to keep it under its size
hard cap, and refreshes the size baselines + golden install-parity fixtures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2351): add changeset (#2426)

* chore: regenerate golden/size baseline after rebase onto next

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 14:37:15 -04:00
Tom Boucher
40ce95f882 fix(#2358): scope review.md and ship.md temp files to a per-run mktemp directory (#2433)
* fix(#2358): scope review workflow temp files to a per-run mktemp dir

/gsd-review wrote every prompt/section/output temp file to a hardcoded
/tmp path keyed only on the phase number, so two GSD projects sharing
a small phase number collide on the exact same path and a crashed
run's leftover file becomes bait a later, unrelated run can silently
read. ship.md's external peer-review stderr capture was strictly
worse — one shared, unqualified path across every project/phase/run.

Thread a single mktemp -d "${TMPDIR:-/tmp}/gsd-review.XXXXXX" run
directory through every review.md temp path (67 sites) via a new
{run_dir}/$RUN_DIR placeholder, mirroring the existing {phase}
substitution mechanism, and clean it up at the end of the run. Route
ship.md's stderr capture through a per-run mktemp file the same way.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2358): regenerate fixtures + lint gate-prep

* fix(#2358): repair failing tests after gate verification

* fix(#2358): thread RUN_DIR scoping into reviewer-instances.md (#1517)

review.md's own invoke_reviewers step lazily loads
gsd-core/references/reviewer-instances.md for the review.reviewer_instances
codepath, but that doc was missed when review.md and ship.md were moved to
the run-scoped {run_dir} temp directory. It still read the combined prompt
from the old /tmp/gsd-review-prompt-{phase}.md (which build_prompt no longer
writes, breaking reviewer-instances functionality outright) and wrote each
instance's output to the old unscoped /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md,
leaving the exact cross-project temp-file collision bug open for that code
path. Both paths now thread through {run_dir}, matching every other reviewer
block in review.md.

Extends the existing #2358 regression test with assertions pinning
reviewer-instances.md's prompt read and output write to {run_dir}, and adds
the Fixed changeset fragment. Regenerated the golden-install-parity content
hashes for reviewer-instances.md via `npm run gen:golden` (paths unchanged;
only the modified file's hash moved).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2358): backfill changeset pr (#2433)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 13:12:03 -04:00
Tom Boucher
a7d83dc234 fix(#2390): warn on goal-shaped phase.add titles, correct auto-detect docs (#2425)
* fix(#2390): phase.add title warning + auto-detect doc fix

phase.add now returns a `warning` field when a description reads as
goal-shaped (>80 chars and/or multi-sentence) rather than title-shaped,
instead of silently writing the whole paragraph verbatim as the
`### Phase N:` header. The CLI still creates the phase as-is (the
strict two-layer slash-vs-CLI interface is unchanged); the warning
just surfaces the gap.

Also clarifies six doc sites (command argument hints, workflow
detection steps, and how-to/reference docs) that described the
phase-number argument as "auto-detecting" the next unplanned phase --
that detection is an orchestrating-workflow/LLM step reading
ROADMAP.md (concretely: `query roadmap.analyze`'s `next_phase`
field), not a `gsd-tools.cjs` CLI feature.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2390): regenerate fixtures + lint gate-prep

* fix(#2390): repair failing tests after gate verification

* chore(#2390): add changeset (#2425)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 07:53:04 -04:00
Tom Boucher
8d2f8bcb23 fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found

Adds requirements.ready-ids (execute-plan.md's update_requirements step)
so a requirement ID declared by multiple plans in a phase only marks
Complete once every declaring plan has produced a SUMMARY.md, and
requirements.revert-phase (execute-phase.md's gaps_found branch) so a
gaps_found verdict reverts the phase's own prematurely-Complete IDs
before the gap report renders. Single-plan IDs still mark immediately.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2388): regenerate fixtures + lint gate-prep

* fix(#2388): repair failing tests after gate verification

* chore(#2388): add changeset (#2424)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 07:52:17 -04:00
Tom Boucher
873bdf51e5 fix(#2352): expand tilde paths in review scope before the deleted-file filter (#2419)
* fix(#2352): tilde-expand SUMMARY.md key-files paths before deleted-file filter

compute_file_scope's "Filter deleted files" step tested the literal `~/...`
value from SUMMARY.md key-files entries with `[ -f "$file" ]`, which bash
never tilde-expands (only a literal `~` in source text expands, not one
arriving as an already-expanded variable value). Real files recorded with a
`~/...` path were silently misclassified as deleted and dropped from
REVIEW_FILES, and a phase whose every recorded file used a tilde path hit the
empty-scope skip as a false negative.

Adds a tilde-normalization loop as step 1 of post-processing (all tiers),
before the deleted-file filter, rewriting a leading `~/` to `${HOME}/...` so
downstream existence checks, the empty-scope short-circuit, and the
FILES_TO_READ/CONFIG_FILES construction all see a real, openable path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2352): regenerate fixtures + lint gate-prep

* chore(#2352): add Fixed changeset fragment (pr 2419)

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 21:29:58 -04:00
Tom Boucher
863a54ec82 fix(#2350): pass --raw to config-get in every build/test gate (#2399)
Adds --raw to config-get workflow.build_command|test_command reads in the post-merge, regression, verify-phase, and audit-fix gates so an unset key is a genuinely empty string, not the literal "" — restoring the auto-detect cascade and graceful skip instead of a false exit-127 failure. Regression guard sweeps all four gate files. Fixes #2350.
2026-07-18 01:14:45 -04:00
Tom Boucher
b0f672f88c fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity.

Closes #2337. Admin-merged (self-review bypass) with full green CI.
2026-07-17 14:01:38 -04:00
Tom Boucher
ada79bee97 fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md

Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md
unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and
per-workstream milestone state already lives in the workstream's own STATE.md /
ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design —
whichever workstream ran new-milestone last silently won the shared heading.
Step 4 is now skipped when a workstream is active; step 6 no longer stages
PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing
when a staged path is unchanged).

Also fixes a second defect found while diagnosing this, same root cause (the
workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and
the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at
the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x`
suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped,
violating the routing-propagation contract. Step 1 now parses --ws using the
established idiom from verify-work.md.

Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not
export the latter and it is only priority 2 of 5 in resolution, so it would miss
the --ws flag case that is the actual repro.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2308): regenerate install goldens for the new-milestone workflow change

gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content
hash is pinned in all 18 runtime golden fixtures. Only that hash changed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests

Independent review found the first pass was partly cosmetic:

1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in
   step 1's shell and each step's bash block runs in its own shell — this file
   already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file
   for exactly that reason (#2288). The guard read an unset variable, always took
   the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS
   in step 6, the branch is removed entirely: step 4 Part A's guard is what
   protects the shared heading, so post-guard the only content PROJECT.md can
   carry is Part B's idempotent Evolution backfill — which must be staged, not
   stranded. A regression test now asserts no cross-step GSD_WS branch returns.

2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a
   shared, idempotent backfill that is not workstream state. A pre-Evolution
   project running only `--ws` would never get the section that transition and
   complete-milestone expect. Step 4 is now split: Part A (milestone-state write)
   is workstream-guarded; Part B (Evolution) always runs.

3. The tests were tautological prose-pinning — including one asserting a comment
   mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it
   passed on the inert guard it existed to catch. Replaced with executable tests
   that extract the step-1 and step-6 fences and run them under bash with stubbed
   gsd_run, asserting real parse and --files behavior.

4. --ws is now stripped from the milestone name (step 1 previously left
   "--ws search" in the remaining text), and documented in argument-hint,
   help/modes/full.md, and docs/COMMANDS.md.

5. Changeset no longer overstates: --ws reaches the prose guard and routing hints
   only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone
   still take no ${GSD_WS} — out of scope here).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change

skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md,
so documenting --ws in the argument-hint made it stale (caught by lint:ci's
gen-plugin-skills --check). Regenerated it plus the install goldens and workflow
size baseline, since commands/, skills/, and gsd-core/workflows/ all ship.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2308): backfill PR number 2338 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:48:13 -04:00
Tom Boucher
1bb724048a fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist (#2325)
* fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist

The convergence reviewer-flag whitelist predated the 1.7.0 Antigravity CLI
adapter and silently dropped --agy/--antigravity, so convergence fell back to
--codex only and the working adapter was unreachable (worse after Gemini CLI's
upstream shutdown). Add both flags to the workflow grep whitelist, the command
argument-hint + flag docs, and the regenerated SKILL.md; they pass through to
/gsd-review unchanged. --gemini behavior is untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#2293): backfill PR number 2325 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 01:55:45 -04:00
Tom Boucher
52fab7d9d7 fix(#2288): archive phase history under the outgoing milestone version (#2323)
* fix(#2288): archive phase history under the outgoing milestone version

phases.clear derived its archive directory from a live getMilestoneInfo()
read, but new-milestone.md switches the milestone BEFORE phases.clear runs,
so phase history was filed under the NEW milestone's <version>-phases/ dir.

Add a --archive-version override (threaded from new-milestone.md, captured
before the switch) with precedence override -> live read -> dated label.
Harden the version label against path traversal on both phases.clear and the
sibling milestone-complete sink (the label is a moved directory name), and
persist the outgoing version via a file + quoted shell expansion so untrusted
STATE.md content is never re-parsed by the shell.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#2288): backfill PR number 2323 into changesets

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 22:51:19 -04:00
Tom Boucher
b041f101fb fix(#2287): surface unresolved deferred-items.md entries in progress + audit-uat (#2318)
The executor SCOPE BOUNDARY convention (agents/gsd-executor.md) logs
out-of-scope discoveries to a phase directory's deferred-items.md, but no
reader ever consumed it — forensic_audit, cmdAuditUat, and capture --list
all skipped it — so deferred items were permanently invisible.

cmdAuditUat (src/uat.cts) now scans each phase dir's deferred-items.md via
a new parseDeferredItems (reusing the collectSection/splitGapsEntries/
extractGapEntryFields seams) and surfaces entries whose status != resolved
(fail-safe: a missing/garbled status surfaces rather than hides, matching
the false-negative-averse posture of #2286). forensic_audit
(gsd-core/workflows/progress.md) gains Check 7 that globs
.planning/phases/*/deferred-items.md and reports unresolved entries.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 21:40:59 -04:00
Tom Boucher
ff9cb6069f fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active'
but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller
outside their own CLI router, and execute-phase.md declared an
execute:wave:pre hook point that the workflow body never rendered — so
claude_orchestration.enabled:true had zero effect on real runs.

Approach B (maintainer-chosen):
- execute-phase.md now renders the execute:wave:pre hook
  (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75,
  immediately before each wave's Agent() dispatch — fixing the latent
  dead-hook gap for any pre-wave capability.
- Move the claude-orchestration contribution execute:wave:post ->
  execute:wave:pre (a pre-wave backend selector belongs before dispatch,
  not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md
  with prose instructing the orchestrator to call resolve-wave-dispatch
  before step 3. Unrelated wave:post contributions (ui.safety-gate, drift,
  external-job, mempalace) untouched.
- New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend
  + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result;
  exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is
  a real non-CLI-router, non-test caller of both functions.

Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool
absent, SDK below floor, execution_backend:inline, malformed input) or an
emit failure resolves to inline with a byte-identical result shape — no
regression to the default-off execute-phase path.

Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor
BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a
fast-check composition property, capability.json contribution assertions,
and a source-contract guard that execute:wave:pre is now actually rendered.
Dependent registry-shape assertions updated in-scope.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 19:18:28 -04:00
Tom Boucher
f74442310d fix(#2257): auto-resume debug on non-terminal session-manager return (#2300)
The /gsd-debug orchestrator handled the gsd-debug-session-manager return
with only two literal-string checks (DEBUG SESSION COMPLETE, ABANDONED)
and no else branch, so a usable-but-non-terminal progress summary (the
manager's own turn/context budget exhausted mid-loop, with a valid
on-disk checkpoint) fell through to the user as if the debug were
complete. Same gap at the continue subcommand.

Callee side (agents/gsd-debug-session-manager.md): add an explicit
non-terminal CONTINUE_REQUIRED return marker, distinct from the two
terminal shapes and from a genuine user-input checkpoint.

Orchestrator (gsd-core/workflows/debug.md Sections 4 and 1c): classify
returns exhaustively — recognized terminal markers behave as before,
anything else is non-terminal and auto-resumes by re-spawning the
session manager from the same slug/checkpoint. Anti-loop guard: after
two consecutive no-progress resumes (unchanged next_action/updated),
emit a blocker report instead of looping.

Regression test (source-text contract guard, fix-2196 idiom) asserts
both sections' non-terminal/auto-resume branch, the CONTINUE_REQUIRED
marker, and the anti-loop bound.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 12:55:38 -04:00
Tom Boucher
315d94f6d4 feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate

Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode.

- gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top.
- gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer.
- --no-tracer flag wired through plan-phase workflow/command/help/skill.
- CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled.
- tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1945): backfill changeset PR number to 2294

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 09:41:36 -04:00
Tom Boucher
8b70db343b fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 (#2259)
* fix(#2204): phase-completion writes 'All phases complete' per ADR-2207

completePhaseCore was writing the overloaded bare 'Milestone complete' on the
last phase — the same string space the milestone-close verb owns for terminal
state. Per ADR-2207, phase-completion now writes the existing intermediate
value 'All phases complete' (already used in gsd2-import.cts). Milestone
termination ('<version> milestone complete' / 'Awaiting next milestone')
remains solely with milestoneCompleteCore.

Status lifecycle: Ready to plan → All phases complete → <version> milestone
complete → Awaiting next milestone.

Changes:
- src/state-transition.cts: completePhaseCore status value
- src/phase.cts: #2028 guard comment
- tests/state-transition.test.cjs: assertion + test name
- tests/phase.test.cjs: 8 assertion updates (positive + negative)
- tests/state.test.cjs: normalizeStateStatus test case + reset regex
- tests/workstream.test.cjs: fixture status to terminal value
- gsd-core/workflows/progress.md: Route D label
- gsd-core/workflows/transition.md: Route B label
- CONTEXT.md: Status lifecycle glossary entry (ADR-2207)
- .changeset/brave-geese-jump.md

* test(#2204): regenerate golden-install-parity fixtures + workflow-size baseline

Workflow file edits (progress.md, transition.md) changed install payload
hashes and pushed past the committed workflow-size baseline. Regenerated
all 17 golden-install-parity fixtures + claude-local via the standalone gen
script (which now also covers the local-scope claude layout). Updated
workflow-size-baseline.json and agent-size-baseline.json via size:baseline.

* fix(#2204): correct claude-local golden hashes + document gen-script limitation

The gen-script's claude-local generation produces macOS-specific hashes
incompatible with Linux CI (local-scope install embeds platform-varying
node-runner paths). Reverted to manual update using Linux FAILURES.md
+actual hashes for the 2 changed workflow files. Added explanatory
comment in the gen script.

* test(#2204): add isCompletedInventory coverage + clarify CONTEXT.md glossary

Addresses orthogonal code-review findings (Medium #1 + #2):
- Add isCompletedInventory test cases for ADR-2207 status lifecycle
  (terminal 'milestone complete' → true; intermediate 'All phases
  complete' → false; archived → true; active statuses → false)
- Clarify CONTEXT.md glossary: note that isCompletedInventory
  intentionally excludes the intermediate value

* docs: backfill changeset PR number (#2259)

* docs(#2204): add Status lifecycle table to state-md reference (ADR-2207)
2026-07-14 14:47:03 -04:00
Tom Boucher
d49ac81306 chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1

Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a
canonical seam and migrate the pilot reader.

- Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed
  {columns, rows} addressed by column NAME; ragged rows are typed parse
  errors, not silent), a single-source TABLE_SCHEMAS registry
  (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with
  variants under one id), matchTableSchema, and findTableBySchema. Result<T>
  is scoped to this seam (distinct from the dispatch Result).
- Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the
  position-anchored regex to name-based resolution via the seam — fixes #2137
  (the 5-column milestone-grouped Progress table previously returned all-null).
- Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route
  fast.md's log_to_state through it, retiring the inline `awk NF-2` column
  arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped
  (| and newlines) and the STATE.md read-modify-write is atomic under
  readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230).
- Writer/reader/template parity test guards TABLE_SCHEMAS against drift
  (ADR-2143 §3 Generative-Fix-Divergence).

Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md +
INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md.

Behaviour-preserving for the canonical 4-column Progress table; the named
bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#2242): backfill changeset PR number (#2248)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2242): escape backslash before pipe in markdown-table cell escaping

CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not
the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow
unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl.
literal backslashes) round-trip exactly. Added backslash round-trip tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan

Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself
"pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via
a new seam helper findTableWithColumns (first table whose header is a superset of
Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads
cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact
TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while
staying seam-based and preserving its `## Progress` scoping (#2012/#1445).
Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale
state.test.cjs assertion that predated the Phase-1 migration.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:36:14 -04:00
Cody Anderson
98e4233ce9 fix(#2176): ground the Antigravity reviewer in the repo under review (#2184)
* fix(#2176): ground the Antigravity reviewer in the repo under review

- capability-probe --add-dir (mirrors the Codex bypass-flag probe) and pass
  the repo root on both invocation arms
- anchor _AGY_PROMPT to the absolute repo root; mandate a
  REVIEWED-WITHOUT-REPO-ACCESS self-report when the repo is unreadable
- stamp a [reviewed-without-repo-access] marker on self-reported or
  scratch-anchored output; Consensus Summary down-weights marked reviews
- apply the same absolute-root anchor to the cursor-agent prompt (AC5)

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2176): changeset fragment for PR #2184

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): review fixes — size baseline, cursor root anchor, anchored blind tells

- regenerate tests/workflow-size-baseline.json for review.md's growth
- cursor anchor uses git rev-parse --show-toplevel (bare pwd resolved the
  wrong root from a repo subdirectory)
- blind-review tells anchored: self-report to the first lines of output,
  scratch tell to a workspace-declaration phrasing — a grounded review
  quoting either string is no longer mis-stamped

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): round-2 review fixes — scratch-tell bridge, behavioral test, changeset

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the review.md change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): pass the transcript path to bash with forward slashes

The behavioral detection test substitutes a mkdtemp path into the bash
compound; on Windows runners that path contains backslashes, which bash
strips, so the transcript is never found and the first assertion fails
(windows-latest/24 lane). Git Bash accepts D:/-style paths.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): use /gsd:review namespace syntax in workflow comment

The slash-command namespace invariant (#3443) bans retired /gsd-<cmd>
references in Claude-facing sources; a cursor-anchor comment used
/gsd-review. Size baseline + golden fixtures regenerated for the byte
change.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): derive the POSIX path via path.sep, not a hardcoded separator

Review finding: out.replaceAll('\\', '/') hardcodes both separators;
use the separator-safe out.split(path.sep).join(path.posix.sep) idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): use the merged toPosixPath seam for the bash path

Per maintainer note: #2247's shell-command-projection now centralizes
running-OS → POSIX path conversion; import it instead of the inline
split/join idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 15:54:19 -04:00
Adnan
3592697bed fix(#2107): orchestrator honors gate="blocking-human" checkpoints in auto-mode (#2113)
* fix(execute-phase): honor gate="blocking-human" in auto-mode checkpoint handling

The package-legitimacy gate (#2827) spans two layers. gsd-executor refuses to
auto-approve a gate="blocking-human" checkpoint and escalates it so a human can
vet the package. execute-phase's checkpoint_handling step then dispatched purely
on checkpoint *type* and never read gate -- so under --auto/--chain it
auto-approved the checkpoint the executor had just refused to auto-approve.

Net effect: the slopsquatting defence was inert in exactly the unattended mode
where it matters. An [ASSUMED]/[SUS] package reached install with no human ever
seeing the prompt.

- gsd-core/workflows/execute-phase.md: carve out gate="blocking-human" (and the
  package-legitimacy what-built markers) ahead of every auto-mode branch.
- gsd-core/references/checkpoints.md: document the gate attribute and its two
  values. blocking-human previously appeared nowhere outside gsd-executor.md,
  so no planner had a documented way to author a non-auto-approvable checkpoint.
- tests/package-legitimacy-gate.test.cjs: the existing regression test asserted
  the executor half only, which is why it stayed green while the gate was open.
  Now asserts the orchestrator half too.

* chore(changeset): link to issue #2107

* chore(changeset): backfill PR number 2113

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* test(#2107): refresh golden-install-parity hashes for edited gsd-core files

The golden fixtures pin content hashes for gsd-core/references/checkpoints.md
and gsd-core/workflows/execute-phase.md, both edited by this fix. Regenerated
via UPDATE_GOLDEN=1; only those two keys change across all 17 runtime fixtures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): keep the carve-out inside the ADR-857 host-loop budget

The ADR-857 phase-6 ratchet pins execute-phase.md below 93600 LF bytes so
optional-feature logic keeps migrating out of the host loop. The carve-out
first landed 623 bytes over that ceiling.

Move the two-layer rationale (why gsd-executor escalates these checkpoints)
into references/checkpoints.md, where the gate is now documented, and reduce
the workflow to the operative rule. execute-phase.md is 93589 bytes, under
the ceiling; the gate token and both <what-built> marker strings are kept
because the orchestrator matches on them.

Refresh the two baselines the edit invalidates: golden-install-parity
fixtures (only the checkpoints.md and execute-phase.md hashes move) and
workflow-size-baseline.json (one line). The ADR-857 ceiling itself is
untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): executor honors blocking-human on the decision branch + gate transport

Review found the fix incomplete one layer down. Two executor-layer gaps:

1. Blocker — agents/gsd-executor.md auto-mode dispatch gated
   checkpoint:human-verify on gate="blocking-human" but the checkpoint:decision
   branch below auto-selected the first option with no gate check. The executor
   resolves a decision itself (auto-selects and continues) without returning it,
   so the orchestrator carve-out never runs for it. A planner following the new
   checkpoints.md rule 6 ("gate a decision whose default would be wrong to
   assume") would have it silently auto-selected under --auto/--chain — the exact
   #2107 harm, one checkpoint type over. The decision branch now STOPs and
   returns for an explicit human decision when gate="blocking-human".

2. Major (transport) — checkpoint_return_format carried no field conveying the
   gate to the freshly-spawned orchestrator, so recognition of the proactive
   pre-install checkpoint rested on freeform prose. Added a **Gate:** field to
   the return format and re-pointed the execute-phase carve-out at it
   ("If the returned Gate: is blocking-human"). Net byte-negative: execute-phase.md
   drops 93589 -> 93583, widening ADR-857 headroom from 11 to 17 bytes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): cover decision carve-out + gate transport, de-vacuum conditional tests

- New: 'auto mode does not auto-select a blocking-human decision checkpoint'
  asserts the executor decision branch STOPs on blocking-human. Verified red on
  the pre-fix executor (2 fail), green with the fix (27 pass).
- New: 'checkpoint_return_format transports the gate ...' asserts the **Gate:**
  field carries blocking-human across the executor->orchestrator boundary.
- New: 'auto-select rule for decision is conditional' — orchestrator-side mirror
  of the human-verify conditional test, for the execute-phase decision branch.
- Fix vacuous test: both conditional tests now assert the anchor matched
  (length > 0) before iterating, so anchor drift can no longer pass with zero
  assertions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): refresh golden + size baselines for executor + execute-phase edits

Regenerated via UPDATE_GOLDEN=1 and update-size-baseline.cjs. Only the
gsd-executor.md and gsd-core/workflows/execute-phase.md hashes move across the
runtime fixtures (35 ins / 35 del, no keys added or removed); checkpoints.md is
unchanged this round. Size baselines: gsd-executor.md 43607 -> 43973,
execute-phase.md 93589 -> 93583 (still under the ADR-857 ceiling).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:43:29 -04:00
Tom Boucher
b5ce72f729 fix(#2119): single SECURITY.md writer — auditor is return-only (#2154)
* fix(2119): single SECURITY.md writer — auditor is return-only

The gsd-security-auditor held Write/Edit and was instructed to write
SECURITY.md (no <N>- prefix, no template frontmatter), while the
orchestrator's Step 6 also wrote the correct padded <N>-SECURITY.md
from templates/SECURITY.md. Two writers, two naming conventions, two
shapes — the auditor's unprefixed file was invisible to the workflow's
*-SECURITY.md glob detector and unparseable for the threats_open gate.

Fix (option 1 from the issue): make the auditor return-only.
- Remove Write/Edit from auditor's tools
- Rewrite all 'Write SECURITY.md' instructions to 'Return structured
  verdict' with threats_open count
- Add explicit constraint in workflow Step 5 spawn prompt
- Update existing test (was asserting Write in tools — now asserts absence)
- Add new regression test for single-writer contract
- Update docs/AGENTS.md stale Tools/Produces rows
- Regenerate golden fixtures + agent size baseline

* docs(changeset): backfill PR number (#2154)

* chore(#2119): regenerate pi/qwen golden fixtures after next merge

The single-writer change edits gsd-core/workflows/secure-phase.md and
agents/gsd-security-auditor.md; pi.json (added on next) and qwen.json (merge
straggler) were the only runtime fixtures still holding pre-change hashes for
those files. All other runtimes already reflect the change. Regenerated via
the sanctioned gen-golden-install-parity script.

* merge origin/next — regenerate goldens + baseline for merged state

* fix slash-command syntax: /gsd-secure-phase → /gsd:secure-phase (#2154 CI fix)
2026-07-13 00:47:15 -04:00
Tom Boucher
4bb846b67a fix(#2112): scope commit to --files pathspec, not entire index (#2148)
* fix(2112): scope commit to --files pathspec, not entire index

cmdCommit/cmdCommitToSubrepo/cmdPrSubrepo staged exactly the files
named in --files but then ran a bare 'git commit' with no pathspec,
absorbing anything else in the index into a commit whose message
described only the named files (#2112).

Fix: append '-- ...stagedPaths' to the commit args when the caller
declared a scope. Three guards are load-bearing:
- stagedPaths (not filesToStage) excludes skipped missing files (#2014)
- explicitFiles gate keeps the default .planning/ path byte-identical
- MERGE_HEAD check via 'git rev-parse' falls back to bare commit during merge
- --amend is left without pathspec (different operation)

cmdPrSubrepo pathspec uses changedFiles (old+new for renames) so the
full rename is captured atomically.

Also fixes workflow markdown in spec-phase.md and add-tests.md.

All-files-missing now short-circuits to nothing_to_commit instead of
absorbing the entire index under a message describing files that
were not committed.

* docs(changeset): backfill PR number (#2148)

* test: update golden-install-parity fixtures for workflow markdown changes (#2112)

* test: update golden fixtures + workflow baselines for #2112 changes

- claude-local.json golden fixture (now generated via gen script)
- workflow-size-baseline.json (add-tests.md +16, spec-phase.md +42 bytes)
- Extended gen-golden-install-parity-zcode.cjs to also regenerate the
  claude local-layout fixture
2026-07-13 00:21:46 -04:00
Tom Boucher
f8c5c1590f fix(#2196): declare the debug session-manager spawn foreground + no-TaskOutput + recovery (#2227) 2026-07-12 19:47:42 -04:00