Commit Graph

113 Commits

Author SHA1 Message Date
Cody Anderson
98e4233ce9 fix(#2176): ground the Antigravity reviewer in the repo under review (#2184)
* fix(#2176): ground the Antigravity reviewer in the repo under review

- capability-probe --add-dir (mirrors the Codex bypass-flag probe) and pass
  the repo root on both invocation arms
- anchor _AGY_PROMPT to the absolute repo root; mandate a
  REVIEWED-WITHOUT-REPO-ACCESS self-report when the repo is unreadable
- stamp a [reviewed-without-repo-access] marker on self-reported or
  scratch-anchored output; Consensus Summary down-weights marked reviews
- apply the same absolute-root anchor to the cursor-agent prompt (AC5)

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2176): changeset fragment for PR #2184

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): review fixes — size baseline, cursor root anchor, anchored blind tells

- regenerate tests/workflow-size-baseline.json for review.md's growth
- cursor anchor uses git rev-parse --show-toplevel (bare pwd resolved the
  wrong root from a repo subdirectory)
- blind-review tells anchored: self-report to the first lines of output,
  scratch tell to a workspace-declaration phrasing — a grounded review
  quoting either string is no longer mis-stamped

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): round-2 review fixes — scratch-tell bridge, behavioral test, changeset

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the review.md change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): pass the transcript path to bash with forward slashes

The behavioral detection test substitutes a mkdtemp path into the bash
compound; on Windows runners that path contains backslashes, which bash
strips, so the transcript is never found and the first assertion fails
(windows-latest/24 lane). Git Bash accepts D:/-style paths.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): use /gsd:review namespace syntax in workflow comment

The slash-command namespace invariant (#3443) bans retired /gsd-<cmd>
references in Claude-facing sources; a cursor-anchor comment used
/gsd-review. Size baseline + golden fixtures regenerated for the byte
change.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): derive the POSIX path via path.sep, not a hardcoded separator

Review finding: out.replaceAll('\\', '/') hardcodes both separators;
use the separator-safe out.split(path.sep).join(path.posix.sep) idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): use the merged toPosixPath seam for the bash path

Per maintainer note: #2247's shell-command-projection now centralizes
running-OS → POSIX path conversion; import it instead of the inline
split/join idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 15:54:19 -04:00
Adnan
3592697bed fix(#2107): orchestrator honors gate="blocking-human" checkpoints in auto-mode (#2113)
* fix(execute-phase): honor gate="blocking-human" in auto-mode checkpoint handling

The package-legitimacy gate (#2827) spans two layers. gsd-executor refuses to
auto-approve a gate="blocking-human" checkpoint and escalates it so a human can
vet the package. execute-phase's checkpoint_handling step then dispatched purely
on checkpoint *type* and never read gate -- so under --auto/--chain it
auto-approved the checkpoint the executor had just refused to auto-approve.

Net effect: the slopsquatting defence was inert in exactly the unattended mode
where it matters. An [ASSUMED]/[SUS] package reached install with no human ever
seeing the prompt.

- gsd-core/workflows/execute-phase.md: carve out gate="blocking-human" (and the
  package-legitimacy what-built markers) ahead of every auto-mode branch.
- gsd-core/references/checkpoints.md: document the gate attribute and its two
  values. blocking-human previously appeared nowhere outside gsd-executor.md,
  so no planner had a documented way to author a non-auto-approvable checkpoint.
- tests/package-legitimacy-gate.test.cjs: the existing regression test asserted
  the executor half only, which is why it stayed green while the gate was open.
  Now asserts the orchestrator half too.

* chore(changeset): link to issue #2107

* chore(changeset): backfill PR number 2113

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* test(#2107): refresh golden-install-parity hashes for edited gsd-core files

The golden fixtures pin content hashes for gsd-core/references/checkpoints.md
and gsd-core/workflows/execute-phase.md, both edited by this fix. Regenerated
via UPDATE_GOLDEN=1; only those two keys change across all 17 runtime fixtures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): keep the carve-out inside the ADR-857 host-loop budget

The ADR-857 phase-6 ratchet pins execute-phase.md below 93600 LF bytes so
optional-feature logic keeps migrating out of the host loop. The carve-out
first landed 623 bytes over that ceiling.

Move the two-layer rationale (why gsd-executor escalates these checkpoints)
into references/checkpoints.md, where the gate is now documented, and reduce
the workflow to the operative rule. execute-phase.md is 93589 bytes, under
the ceiling; the gate token and both <what-built> marker strings are kept
because the orchestrator matches on them.

Refresh the two baselines the edit invalidates: golden-install-parity
fixtures (only the checkpoints.md and execute-phase.md hashes move) and
workflow-size-baseline.json (one line). The ADR-857 ceiling itself is
untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): executor honors blocking-human on the decision branch + gate transport

Review found the fix incomplete one layer down. Two executor-layer gaps:

1. Blocker — agents/gsd-executor.md auto-mode dispatch gated
   checkpoint:human-verify on gate="blocking-human" but the checkpoint:decision
   branch below auto-selected the first option with no gate check. The executor
   resolves a decision itself (auto-selects and continues) without returning it,
   so the orchestrator carve-out never runs for it. A planner following the new
   checkpoints.md rule 6 ("gate a decision whose default would be wrong to
   assume") would have it silently auto-selected under --auto/--chain — the exact
   #2107 harm, one checkpoint type over. The decision branch now STOPs and
   returns for an explicit human decision when gate="blocking-human".

2. Major (transport) — checkpoint_return_format carried no field conveying the
   gate to the freshly-spawned orchestrator, so recognition of the proactive
   pre-install checkpoint rested on freeform prose. Added a **Gate:** field to
   the return format and re-pointed the execute-phase carve-out at it
   ("If the returned Gate: is blocking-human"). Net byte-negative: execute-phase.md
   drops 93589 -> 93583, widening ADR-857 headroom from 11 to 17 bytes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): cover decision carve-out + gate transport, de-vacuum conditional tests

- New: 'auto mode does not auto-select a blocking-human decision checkpoint'
  asserts the executor decision branch STOPs on blocking-human. Verified red on
  the pre-fix executor (2 fail), green with the fix (27 pass).
- New: 'checkpoint_return_format transports the gate ...' asserts the **Gate:**
  field carries blocking-human across the executor->orchestrator boundary.
- New: 'auto-select rule for decision is conditional' — orchestrator-side mirror
  of the human-verify conditional test, for the execute-phase decision branch.
- Fix vacuous test: both conditional tests now assert the anchor matched
  (length > 0) before iterating, so anchor drift can no longer pass with zero
  assertions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): refresh golden + size baselines for executor + execute-phase edits

Regenerated via UPDATE_GOLDEN=1 and update-size-baseline.cjs. Only the
gsd-executor.md and gsd-core/workflows/execute-phase.md hashes move across the
runtime fixtures (35 ins / 35 del, no keys added or removed); checkpoints.md is
unchanged this round. Size baselines: gsd-executor.md 43607 -> 43973,
execute-phase.md 93589 -> 93583 (still under the ADR-857 ceiling).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:43:29 -04:00
Tom Boucher
b5ce72f729 fix(#2119): single SECURITY.md writer — auditor is return-only (#2154)
* fix(2119): single SECURITY.md writer — auditor is return-only

The gsd-security-auditor held Write/Edit and was instructed to write
SECURITY.md (no <N>- prefix, no template frontmatter), while the
orchestrator's Step 6 also wrote the correct padded <N>-SECURITY.md
from templates/SECURITY.md. Two writers, two naming conventions, two
shapes — the auditor's unprefixed file was invisible to the workflow's
*-SECURITY.md glob detector and unparseable for the threats_open gate.

Fix (option 1 from the issue): make the auditor return-only.
- Remove Write/Edit from auditor's tools
- Rewrite all 'Write SECURITY.md' instructions to 'Return structured
  verdict' with threats_open count
- Add explicit constraint in workflow Step 5 spawn prompt
- Update existing test (was asserting Write in tools — now asserts absence)
- Add new regression test for single-writer contract
- Update docs/AGENTS.md stale Tools/Produces rows
- Regenerate golden fixtures + agent size baseline

* docs(changeset): backfill PR number (#2154)

* chore(#2119): regenerate pi/qwen golden fixtures after next merge

The single-writer change edits gsd-core/workflows/secure-phase.md and
agents/gsd-security-auditor.md; pi.json (added on next) and qwen.json (merge
straggler) were the only runtime fixtures still holding pre-change hashes for
those files. All other runtimes already reflect the change. Regenerated via
the sanctioned gen-golden-install-parity script.

* merge origin/next — regenerate goldens + baseline for merged state

* fix slash-command syntax: /gsd-secure-phase → /gsd:secure-phase (#2154 CI fix)
2026-07-13 00:47:15 -04:00
Tom Boucher
4bb846b67a fix(#2112): scope commit to --files pathspec, not entire index (#2148)
* fix(2112): scope commit to --files pathspec, not entire index

cmdCommit/cmdCommitToSubrepo/cmdPrSubrepo staged exactly the files
named in --files but then ran a bare 'git commit' with no pathspec,
absorbing anything else in the index into a commit whose message
described only the named files (#2112).

Fix: append '-- ...stagedPaths' to the commit args when the caller
declared a scope. Three guards are load-bearing:
- stagedPaths (not filesToStage) excludes skipped missing files (#2014)
- explicitFiles gate keeps the default .planning/ path byte-identical
- MERGE_HEAD check via 'git rev-parse' falls back to bare commit during merge
- --amend is left without pathspec (different operation)

cmdPrSubrepo pathspec uses changedFiles (old+new for renames) so the
full rename is captured atomically.

Also fixes workflow markdown in spec-phase.md and add-tests.md.

All-files-missing now short-circuits to nothing_to_commit instead of
absorbing the entire index under a message describing files that
were not committed.

* docs(changeset): backfill PR number (#2148)

* test: update golden-install-parity fixtures for workflow markdown changes (#2112)

* test: update golden fixtures + workflow baselines for #2112 changes

- claude-local.json golden fixture (now generated via gen script)
- workflow-size-baseline.json (add-tests.md +16, spec-phase.md +42 bytes)
- Extended gen-golden-install-parity-zcode.cjs to also regenerate the
  claude local-layout fixture
2026-07-13 00:21:46 -04:00
Tom Boucher
f8c5c1590f fix(#2196): declare the debug session-manager spawn foreground + no-TaskOutput + recovery (#2227) 2026-07-12 19:47:42 -04:00
Tom Boucher
24c324cbfa fix(#2194): add Bash timeout guidance for prompt-fed reviewers in review.md
The Gemini, Claude, and Codex reviewer blocks invoked the CLIs with no explicit
timeout, so each inherited the host default (~2 min on Claude Code). A source-
grounded review of a large plan set takes ~570s (Codex xhigh) / ~525s (headless
Claude) — both exceed that window, so the lane is killed mid-review, its output
is empty, and the cross-AI review silently proceeds with fewer lanes. CodeRabbit
and OpenCode already documented a timeout; the four main lanes did not.

Add a shared timeout-guidance note directing a high Bash timeout (>= 900000;
1200000 for Codex xhigh / headless Claude), referencing BASH_MAX_TIMEOUT_MS for
the Claude Code host cap, and framing a slow-lane empty output as a timeout kill
(not the 0xc0000142 crash it gets misdiagnosed as) so operators re-run with more
time instead of diagnosing a CLI failure.

Closes #2194

Recaptures the 18 golden-install-parity fixtures + workflow-size baseline (only
the review.md entry changed in each; review is LARGE-tier, 47168 < 61440).
2026-07-12 18:53:49 -04:00
Tom Boucher
924ff6822e fix(#2138): surface track_shipping push failures (review)
Review (MEDIUM): `git push ... 2>&1` did not check exit code, so a silent push
failure (auth-token expiry, network blip, non-fast-forward) would proceed to the
report step and declare success — silently reproducing the exact #2138 defect.
Add a fallback warning naming the rerun command so a failed push is visible
(best-effort: the PR already exists, so we still report it rather than abort).
Recaptures goldens + size baseline.
2026-07-12 15:55:39 -04:00
Tom Boucher
aaf74878f7 fix(#2138): push the track_shipping ship-note onto the PR branch [ci skip]
track_shipping committed the STATE ship-note ('Phase N shipped — PR #N') AFTER
create_pr and never pushed it, so the commit stayed local-only. When the GitHub
PR merged (especially fast/auto-merge) the ship-note was not in the source branch
and never reached the default branch — STATE's ship-status was silently lost,
recoverable only by STATE self-heal on the next /gsd-start.

Push the ship-note commit onto the PR branch with a [ci skip] trailer. GitHub
honors [ci skip]/[skip ci], so this lands the note on merge without triggering a
redundant pipeline, and preserves the PR number in STATE.

Recaptures the 18 golden-install-parity fixtures + the workflow-size baseline
(only the ship.md entry changed in each).

Closes #2138
2026-07-12 15:55:39 -04:00
Tom Boucher
c6b4304bab test(#2133): regenerate golden + workflow-size baselines for fast.md
The fast.md log_to_state fix changes an installed workflow file, so the per-
runtime golden-install-parity fixtures (18 runtimes) and the workflow-size
baseline are recaptured. Diff is exactly one entry per fixture (the fast.md
hash) and one baseline byte count — no spurious drift.
2026-07-12 10:48:43 -04:00
Tom Boucher
b55a2e7655 fix(#2117): distinguish not-yet-validated phase from validated failure in audit-milestone
audit-milestone's Nyquist scan classified a phase from `nyquist_compliant`
alone, so a phase seeded by plan-phase but never run through validate-phase
read PARTIAL — identical to a phase that validated and genuinely failed. The
template's `status` field could discriminate the two, but no workflow ever
promoted it off `draft`, so it was dead.

Make `status` live and read it:
- validate-phase.md §6: set `status: validated` in both the create (State B)
  and update (State A) VALIDATION.md paths.
- audit-milestone.md §5.5: parse `status`; add a distinct NOT-VALIDATED bucket
  keyed on `status: draft`, gate COMPLIANT/PARTIAL on `status: validated`, and
  report `not_validated_phases` in the audit YAML.
- VALIDATION.md template: document the draft → validated lifecycle.

Tests & generated artifacts:
- Regression test folded into policy-138 (owning workflow-contract file);
  fail-first verified vs origin/next (0 matches pre-fix).
- Regenerate golden-install-parity fixtures cleanly: adds the previously-missed
  qwen.json and removes a contaminated `settings.local.json` entry that had
  leaked into claude-local.json (the harness excludes hook-config files).
- Correct a stale validate-phase.md workflow-size-baseline entry.

Closes #2117

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 22:47:51 -04:00
Tom Boucher
79d7657eff feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.

Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
  slash-programmatic / active-model / native-extension / bun) + hostBehaviors
  {nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
  RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
  payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
  by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
  EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
  surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
  from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
  (global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
  markdown); the 16 other fixtures + claude-local change only by the shared
  model-catalog hash line.

Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
  subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
  in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
  (every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
  gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
  a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
  shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
  query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
  consumes params; getArgumentCompletions; before_provider_request active-model
  steering (fail-open on null resolution); functional session_start /
  before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
  hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
  vocabulary.

Docs (host-integration matrix + how-to) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 21:07:38 -04:00
Tom Boucher
f014ec83bd feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads:
finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks
copy), and a skills converter-name registry (the artifactLayout.converter
field is now load-bearing, not decorative). frontmatterDialect stays the
documented dispatch key for frontmatter (no descriptor field for it). Dead
isKilo destructure bindings removed. Byte-identical golden parity for all 16
runtimes (opencode, which shares kilo's combined-family path, verified clean).

UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin +
extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus).
UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread
modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped
from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's
mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is
the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented'
per AC so dispatch degrades to 'degraded' by design.

Model-catalog single-source edit ripples the shared model-catalog.json hash
into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale-
bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/
codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the
connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error.

Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/
hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades
(plugin parity+load, model-override converter, agents dispatch surface, MCP
doc). Matrix + how-to + config docs updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 19:55:15 -04:00
Dave
a8ff8fbb30 docs(#1867): specify propose-then-confirm --auto behavior in ui-phase Step 9.5 (review Minor) 2026-07-10 14:41:12 -04:00
Dave
31a500b970 docs(#1867): replace stale plan-phase.md:921 line-pointer with section-name reference (review #6) 2026-07-10 14:40:03 -04:00
Dave
20608cba90 chore(#1867): regen goldens + size baselines after merge onto next
next advanced to a3b9cbae (incl. #1820 specless-rail merge); regenerate
golden-install-parity fixtures + workflow-size-baseline authoritatively.
Manifest + agent baseline already in sync. plan-phase.md conflict
hand-merged to keep both EDGE_ABSENT/PROHIB_ABSENT and UI Considerations.
2026-07-10 14:40:03 -04:00
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
dbc730d8de fix(#2073): capability-probe external killer (timeout/gtimeout) for macOS
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock
macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for
'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout
alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s)
stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor
the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both
branches. Update the agy + #687 tests to assert the probe + bound + fallback,
regen the 17 goldens + size baseline, refresh the maintainer-note version stamp
to 1.0.16.
2026-07-08 17:53:28 -04:00
Tom Boucher
7abe6e34ac chore(#2073): regen workflow-size baseline for review.md growth
The agy block grew (~file-reference prompt instruction, external-timeout
rationale, --model wiring, richer Step 3 diagnostic) — all load-bearing
content fixing 3 production failure modes (#2073), not bloat. Growth
justified in the PR.
2026-07-08 16:47:02 -04:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
Tom Boucher
d7129222c0 Merge branch 'next' into codex/gsd-onboard 2026-07-06 23:49:49 -04:00
Tom Boucher
080bacdb4b Merge branch 'next' into codex/gsd-onboard 2026-07-06 22:43:32 -04:00
Dave
91998dcbca chore(#1820): regen goldens + workflow size baseline after rebase onto next 2026-07-06 14:41:00 -04:00
Dave
7ef834cabc feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them 2026-07-06 14:40:38 -04:00
Joe Slitzker
7d4fc3a519 test(#1941): regenerate golden fixtures + size baseline after rebase onto next
The rebase onto origin/next pulled in the runtime-launcher preamble resync
(applied repo-wide on next) alongside this branch's quick.md change; both
together shift every runtime's install hashes and workflow sizes, so the
fixtures from the pre-rebase regen were stale.
2026-07-06 13:01:57 -05:00
Joe Slitzker
3fe9a81428 fix(#1941): degrade /gsd-quick worktree dispatch when fork base is stale
Claude Code's isolation="worktree" forks new worktrees from origin/HEAD, not
the live local HEAD. When prior local commits (e.g. an earlier quick task in
the same session, or this task's own Step 5.6 pre-dispatch plan commit)
advance local HEAD without an intervening push, origin/HEAD stays pinned to a
stale ancestor and the executor's worktree_branch_check guard halts with a
base-mismatch fatal that can be many commits behind, not just one.

Port the worktree.base-check auto-degrade pattern already used by
execute-phase (#683/#1369) into quick.md's single-dispatch path, run
immediately before EXPECTED_BASE is captured in Step 6.
2026-07-06 13:01:47 -05:00
Tom Boucher
9f0d785b61 fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows (#2027)
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows

gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).

- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
  file refs in agents/workflows/references markdown.

Closes #2020

* docs(#2020): backfill changeset pr 2027
2026-07-05 16:43:26 -04:00
Tom Boucher
1bfadec2d0 fix(#1921): preserve verify-work state across gap-closure + defer follow-ups (#2025)
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups

Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.

- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
  addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
  status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
  are not re-diagnosed and do not spawn new gap plans; a re-reported break
  is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
  'next version', 'out of scope', ...) is captured to UAT ## Deferred
  Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).

Closes #1921

* docs(#1921): backfill changeset pr 2025

* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence

The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 15:42:10 -04:00
Codesmith
a979cfd3f4 chore(#1990): regenerate golden fixtures and size baseline after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:18:28 +00:00
Tom Boucher
ef8a3e27d4 fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy (#2014)
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy

§8 Model Policy ended with </step> but had no matching opening tag (5 opens /
6 closes), leaving it as loose inter-step content. Add the missing
<step name="model_policy"> opener so the section is a proper step.

- gsd-core/workflows/settings-advanced.md: add <step name="model_policy">
- tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level
  workflow must have balanced <step>/</step> (fenced code stripped), plus a
  focused assertion that §8 is wrapped in model_policy.
- goldens + workflow-size baseline recaptured.

Closes #1864

* docs(#1864): backfill changeset pr 2014
2026-07-05 14:38:24 -04:00
Tom Boucher
a62079b2da fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR (#2024)
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR

The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').

The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.

- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
  + explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.

Closes #1865

* docs(#1865): backfill changeset pr 2024
2026-07-05 14:20:52 -04:00
Behruz Nassre Esfahani
23254ca5a7 fix(#1936): reconstruct OpenCode review from JSON events (#1992)
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub

On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.

Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).

review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1936): add changeset

* test(#1936): property-test the OpenCode review jq reconstruction

Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.

Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.

Verified the invariant has teeth (a comma-join jq fails the property).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test when jq is absent

The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).

Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent

The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.

Verified: macOS → 7 pass; simulated win32 → skips.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 14:07:07 -04:00
Tom Boucher
8de2ff9121 feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator

Third-party capability gates declared via check.predicate were rendered for
display but never evaluated (only built-in check.query gates fired; the
security capability's gate worked solely via a hard-coded ship.md branch).

Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts)
that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a
bounded sh -c command at the project root (via shell-command-projection.execTool),
inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed.

Wire a 'check predicate' subcommand into check-command-router.cts and extend
the three generic workflow gate-dispatch sites (execute:wave:post, execute:post,
plan:post) to route check.predicate gates to the new evaluator. The two-step
gate contract (command-failure => onError; block => halt) is unchanged.

- src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible
- src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags
- docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md
- tests: 38 unit + integration tests (exit mapping, timeout, interpolation,
  property-based bijection, malformed-predicate fail-closed, real subprocess e2e)

Closes #2008

* docs(#2008): backfill changeset pr number 2011
2026-07-05 14:04:29 -04:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Rezolv
55604e9124 fix(#1906): require node-test clean-fixture causation control (#2001)
* fix(#1906): require node-test clean-fixture causation control

The node-test fail-first proof accepted a deceptive content-independent
negative test — one that reds merely because GSD_PROHIB_SUBJECT is set,
ignoring the subject's content — whenever no cleanFixture was supplied,
because #1346's causation control was opt-in. The proof's observed signal
(RED) thus diverged from its target (RED caused by content) by default.

Make the causation control mandatory for the node-test kind: a descriptor
that omits cleanFixture is un-provable (fail-closed), never accepted under
the weaker violation-only proof. When a clean fixture is present, fail-first
is proven exactly as before (RED on violation AND non-vacuous GREEN on clean).
The lint-rule kind is unchanged (its subject IS the linted file; no
GSD_PROHIB_SUBJECT indirection).

Breaking (Hyrum): a previously-green node-test prohibition with no clean
fixture now hard-gates — blast radius is zero in-tree (no node-test
prohibition ships today; only the lint-rule local/no-source-grep dogfood).

Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in.

Closes #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ

* docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test)

Record the node-test mandatory-causation-control supersede across the
governing surfaces:

- ADR-1606 (the enforcement decision-of-record): addendum + Decision 4
  annotated + the "Mandatory causation control — REJECTED" alternative
  flipped to accepted (premise no longer holds: zero in-tree node-test
  consumers).
- ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph
  marked SUPERSEDED, pointing at ADR-1606.
- spec-phase.md: check_clean_fixture is now REQUIRED for node-test
  (was "optional").
- CONTEXT.md: PROHIB.enforce.causation predicate updated.

Regenerated the shipped-artifact cascade from the spec-phase.md edit
(+149 B, well under the 40960 cap): 16 golden-install-parity fixtures
and the workflow size baseline.

Refs #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
2026-07-03 23:21:55 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Rezolv
9d7c046eae refactor(#1852): lazy-split plan-phase.md into steps/ (#1934)
* refactor(#1852): lazy-split plan-phase.md into steps/

Extract 3 self-contained, rarely-hit sections into gsd-core/workflows/plan-phase/steps/
via lazy 'Read and execute' pointers (mirrors execute-phase/steps/, ADR-1610 progressive
disclosure — no eager @-import): closed-phase-gate (1.5), prd-express-path (3.5),
windows-troubleshooting. Byte-invariant: plan-phase.md 94459 -> 89775 (-4684), each step
< 32 KiB anchor, resolved instructions unchanged.

Scope note: only 3 sections were extractable. plan-phase.md is guarded by a dense net of
content-presence tests (plan-bounce/enh-3209/phase6-planning-capabilities assert specific
sections/flags inline) that block extracting the larger blocks without also refactoring
those tests — durable low-80s headroom is deferred pending maintainer re-scope (issue #1852).

Cascade: size:baseline regen, 16 golden-install-parity fixtures regen, INVENTORY in sync.
Full plan-phase test surface green (4129/4130, 0 fail); lint:ci exit 0.

* docs(changeset): Changed fragment for #1934 (plan-phase lazy-split)

* docs(changeset): mark #1934 fragment docs-exempt (internal workflow refactor)

* fix(#1852): add gsd_run launcher preamble to prd-express-path step

The extracted prd-express-path.md calls gsd_run but the canonical launcher
preamble lived in the parent plan-phase.md — runtime-launcher-parity (#373)
walks workflows/ recursively and requires every .md using gsd_run to carry
exactly one preamble + the $HOME/.claude fallback arm (same as the existing
execute-phase/steps/ files). Injected via scripts/sync-runtime-launcher.cjs;
golden fixtures + size baseline regenerated. plan-phase.md unchanged (89775).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 23:30:19 -04:00
Behruz Nassre Esfahani
7bef6a6496 fix(#1863): use named flags for state.* calls in executor + workflows (#1873)
* fix(#1863): use named flags for state.* calls in executor + workflows

The named-only state-command router (parseNamedArgs) silently drops
positional args, so state.cjs threw its required-arg error and
metrics/decisions/blockers/session continuity were never recorded.

Convert record-metric / add-decision / add-blocker / record-session in
agents/gsd-executor.md to the named-flag form (mirroring execute-plan.md),
and fix the two remaining positional record-session calls in
gsd-core/workflows/milestone-summary.md and forensics.md. Recapture the
golden-install-parity fixtures and size baselines for the edited files.

Also fix a pre-existing detached-rebuild handle leak in
tests/graphify-auto-update.slow.test.cjs: three dispatch tests returned
after observing only the synchronous "running" status without awaiting the
detached rebuild's terminal state. That leak was latent until the new
#1863 regression block's added runtime shifted --test-force-exit timing
and surfaced it as a non-zero chunk exit. The three tests now await
terminal status via the file's existing waitForBuildStatus helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1863): add changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 13:01:10 -04:00
jecanore
5657994702 fix(#1716): route resume_from_file to complete_session when no pending tests remain (#1722)
* fix(#1716): route resume_from_file to complete_session when no pending tests remain

When a UAT session has status:partial with blocked_count>0 and pending_count==0 (all remaining tests are blocked, none are pending), resume_from_file found no [pending] test and terminated silently — never routing to complete_session. This blocked the issues==0 auto-transition path even when there were zero code defects.

Guard clause added immediately after the find-pending step: if no [pending] test is found, route to complete_session. complete_session then correctly sets status:partial (because blocked_count>0) without presenting further tests.

Closes #1716

* chore(#1716): add changeset fragment and regenerate golden-install-parity fixtures

Changeset fragment for PR #1722 (type: Fixed).

Golden-install-parity fixtures regenerated for all 16 runtimes — the workflow fix shifts verify-work.md's byte-stable hash in the golden manifest. Regenerated via UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
2026-07-02 11:58:06 -04:00
Tom Boucher
92091d71f2 fix(#1871): wire phase archival end-to-end (phases archive cmd + default + atomic) (#1924)
Follow-up to #1919 (archive-then-remove core). Closes the remaining #1871
acceptance criteria so phase history is preserved across the full milestone
lifecycle, not just at phases.clear:

- #2 src/milestone.cts + src/phases-command-router.cts: extract shared
  archivePhaseDirectories() helper; add cmdPhasesArchive (the previously
  half-wired phases.archive alias now routes instead of erroring Unknown).
- #4 gsd-tools.cjs + src/milestone.cts: milestone complete archives phase
  dirs by default (--no-archive-phases opts out; --archive-phases is now a
  harmless no-op). complete-milestone.md updated to drop the redundant manual
  Yes/Skip archive prompt.
- #3 gsd-core/workflows/new-milestone.md: §6 stages the archive move + source
  removal (git add .planning/milestones/ .planning/phases/) in the same commit
  as the milestone start, so the archive lands atomically — no orphaned
  uncommitted deletions, no un-archived dirs inherited.
- docs/CLI-TOOLS.md (+ ja/zh/ko/pt) + help/modes/full.md: flag accuracy.
- tests: phases archive command (#2) + milestone complete default archive /
  --no-archive-phases opt-out (#4). Goldens + workflow size baseline refreshed.

Closes #1871
2026-07-02 11:14:01 -04:00
Tom Boucher
3c13903dcd feat(#1866): agent-side self-load of configured agent_skills
Each of the 22 consumer agents now self-loads its configured agent_skills
in its mandatory init step, so .planning/config.json agent_skills.<type>
reaches the agent on every runtime — including Cursor and /gsd-autonomous,
where Skill()-delegated workflow bash init did not reliably execute.

- gsd-core/references/agent-skills-bootstrap.md: shared contract
  (query + Read + dedup guard that skips when <agent_skills> is already
  in the prompt, so Claude's orchestrator-side injection never doubles)
- 22 agents/gsd-*.md: one self-load line naming the agent's own type
- gsd-core/workflows/autonomous.md: note that delegated agents self-load
- tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS
  bijection + fast-check property) — Generative-Fix-Divergence guard
- docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY
  row, Changed changeset

Closes #1866
2026-07-01 20:09:01 -04:00
Tom Boucher
3cc4d1608c test: regenerate golden fixtures + size baselines for the example/doc edits
execute-phase.md and gsd-ai-researcher.md are installed artifacts; their edits
shift install hashes and file sizes. Diff is scoped to those two files' hashes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 23:04:13 -04:00
Tom Boucher
da37986cd0 fix(#1847): resolve standard tier to claude sonnet 5
Point the sonnet/standard tier at Claude Sonnet 5 (`claude-sonnet-5`,
GA 2026-06-30) across the Anthropic-backed runtimes and provider presets,
replacing the superseded `claude-sonnet-4-6`. Mirrors the change into the
CONFIGURATION.md and settings-advanced.md runtime-defaults tables (the
#3229 catalog↔docs parity gate) plus the pt-BR/zh-CN translations, and
updates the tests that pin the old ID. Regenerates the workflow size
baseline for the (smaller) settings-advanced.md.

Scope is Sonnet only — opus/haiku IDs are untouched. Prepared as a 1.6.1
hotfix off the v1.6.0 tag.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 33260555b4)
2026-06-30 22:21:18 -04:00
Tom Boucher
93e5d2dd84 fix(#1525): skip deferred phases on autonomous reruns (#1846)
* fix(#1525): skip deferred phases on autonomous reruns

* chore(#1525): add changeset fragment

* chore(#1525): fix changeset body

* test(#1525): refresh install parity fixtures

* test(#1525): shrink autonomous workflow

* test(#1525): refresh autonomous baselines

* test(#1525): tolerate Windows temp cleanup flake
2026-06-30 21:29:38 -04:00
Behruz Nassre Esfahani
dae7f81482 fix(#1528): drop next-phase guidance from security-blocked verify-work presentation (#1687)
* fix(#1528): drop next-phase guidance from security-blocked verify-work presentation

When security enforcement blocks phase advancement (no SECURITY.md produced),
the verify-work presentation told the user advancement was blocked but still
offered `/gsd:plan-phase {next}` and `/gsd:execute-phase {next}`, competing
with the current-phase fix. Remove those two next-phase lines so the blocked
state routes only to the current-phase resolution (secure-phase, ui-review).
The post-transition presentation — reached only after the completion contract
passes — still offers next-phase planning, which is the correct place for it.

Regression coverage added to tests/ui-review-next-guidance.test.cjs: the
security-blocked block must not offer next-phase actions, and the
post-completion block must still offer them. Regenerated workflow size baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1528): add changeset for security-blocked next-phase fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1528): recapture golden-install-parity fixtures for verify-work.md change

Rebased onto next; verify-work.md's installed hash changed across all 16
runtime fixtures. Diff confined to the single gsd-core/workflows/verify-work.md
key per runtime. Assert mode 16/16 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-30 10:52:30 -04:00
Tom Boucher
1f649838b8 Merge branch 'next' into fix/1698-codex-output-last-message 2026-06-30 10:04:59 -04:00
Rezolv
18995380ce feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154)

Carry the edge-probe's existing `backstop` (non-inferable) tier through the
plan-phase projection as a structured flat-scalar marker instead of a prose
parenthetical, and make verify-phase abstain -> human_needed (never silent-pass)
on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror
of #644's prohibition judgment-tier (ADR-550 D4).

Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict):
- src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths
  (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence
  -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable
  -> green, the over-abstention guard).
- src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form
  backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers).

Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550
#1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md;
FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror).

Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a
distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity
test; abstain-on-unconfirmed-backstop regression test red-first.

Implementation notes (deviations from the issue's proposed file list, verified live):
- frontmatter.cts needs no change — its flat parser already round-trips object-form truths.
- verify.cts needs no change — it grades artifacts/key_links structurally; truths are
  LLM-graded at the workflow layer, so consumption lives there + the deterministic helper.
- No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source.

Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines.

* chore(#1154): add changeset (Changed) for honest verifier

User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per
trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs
(a confident silent `passed` becomes `human_needed`), which is user-visible even
though the schema marker is additive.

* docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1)

trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED
truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained
`insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes
to human_needed). Behavior was already correct; this tightens the wording.
Regenerated golden-install-parity fixtures + workflow-size baseline for the touched
verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is
intentionally not taken: the current design is ADR-550-D4-conformant, the abstain
cause rides as a distinguishable report reason, and adding it would exceed the
approved scope.)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-29 00:14:32 -04:00