Commit Graph

161 Commits

Author SHA1 Message Date
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
dbc730d8de fix(#2073): capability-probe external killer (timeout/gtimeout) for macOS
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock
macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for
'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout
alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s)
stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor
the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both
branches. Update the agy + #687 tests to assert the probe + bound + fallback,
regen the 17 goldens + size baseline, refresh the maintainer-note version stamp
to 1.0.16.
2026-07-08 17:53:28 -04:00
Tom Boucher
af9069f865 fix(#2073): harden agy reviewer block (arg overflow, 404, pre-session stall)
Three failure modes on agy 1.0.16, all fixed by mirroring the Cursor block's
invocation discipline:
  * file-reference prompt instead of inline "$(cat)" — a large review prompt
    overflowed the exec arg list (rc 126).
  * external 'timeout 600' wrapper — --print-timeout cannot fire before agy
    creates a session, so a pre-session stall hung unbounded.
  * --model from review.models.agy when set — escape hatch for a pinned model
    that 404s (exit 0, empty stdout + transcript).
  * stdin </dev/null so agy never blocks on a tty.
Also enrich the Step 3 empty-output stub to grep agy cli.log for a
model-availability diagnostic, and correct the stale 'no --model flag' note
plus the 'review.models.agy reserved for future' comment (the config key was
already read but never passed through).
2026-07-08 16:47:02 -04:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
Tom Boucher
d7129222c0 Merge branch 'next' into codex/gsd-onboard 2026-07-06 23:49:49 -04:00
Tom Boucher
080bacdb4b Merge branch 'next' into codex/gsd-onboard 2026-07-06 22:43:32 -04:00
Dave
7ef834cabc feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them 2026-07-06 14:40:38 -04:00
Joe Slitzker
3fe9a81428 fix(#1941): degrade /gsd-quick worktree dispatch when fork base is stale
Claude Code's isolation="worktree" forks new worktrees from origin/HEAD, not
the live local HEAD. When prior local commits (e.g. an earlier quick task in
the same session, or this task's own Step 5.6 pre-dispatch plan commit)
advance local HEAD without an intervening push, origin/HEAD stays pinned to a
stale ancestor and the executor's worktree_branch_check guard halts with a
base-mismatch fatal that can be many commits behind, not just one.

Port the worktree.base-check auto-degrade pattern already used by
execute-phase (#683/#1369) into quick.md's single-dispatch path, run
immediately before EXPECTED_BASE is captured in Step 6.
2026-07-06 13:01:47 -05:00
Tom Boucher
9f0d785b61 fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows (#2027)
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows

gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).

- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
  file refs in agents/workflows/references markdown.

Closes #2020

* docs(#2020): backfill changeset pr 2027
2026-07-05 16:43:26 -04:00
Tom Boucher
1bfadec2d0 fix(#1921): preserve verify-work state across gap-closure + defer follow-ups (#2025)
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups

Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.

- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
  addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
  status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
  are not re-diagnosed and do not spawn new gap plans; a re-reported break
  is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
  'next version', 'out of scope', ...) is captured to UAT ## Deferred
  Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).

Closes #1921

* docs(#1921): backfill changeset pr 2025

* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence

The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 15:42:10 -04:00
Codesmith
425d2f2c8b fix(#1990): keep onboard resolver delegation in sync with canonical launcher
onboard.md delegates the gsd_run preamble to
gsd-core/references/gsd-run-resolver.md via @-include, but two guards
regressed once the canonical launcher snippet advanced to
${CLAUDE_CONFIG_DIR:-$HOME/.claude} (#2024):

- runtime-launcher-parity (B2): the resolver reference still shipped the
  old $HOME/.claude arm. references/ is not covered by
  sync-runtime-launcher.cjs, so refresh the resolver bash block to be
  byte-equal to _runtime-launcher.snippet.sh.
- /gsd:onboard command contract: sync-runtime-launcher.cjs had re-inlined
  the preamble into onboard.md (a delegating file). Teach the sync
  transform to strip-but-never-inline files that @-include the resolver,
  mirroring the exemption already in the parity test (B/B2).

Regenerate golden install fixtures and the workflow size baseline for the
smaller onboard.md.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:10 +00:00
Codesmith
4c673e51f3 chore(#1990): resync runtime launcher and regenerate artifacts after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:10 +00:00
jeremymcs
7734dee051 fix(onboard): resolve Tests-lane failures for onboard workflow
- runtime-launcher-parity: recognize workflows that delegate gsd_run to references/gsd-run-resolver.md (onboard.md) and exempt them from the inline-preamble checks; add a compensating byte-equality guard (B2) asserting the reference bash block matches _runtime-launcher.snippet.sh. - onboard.md: document the TEXT_MODE plain-text/numbered-list fallback for AskUserQuestion on non-Claude runtimes (fixes ask-user-questions-fallback, #2012). - Regenerate golden-install-parity fixtures, docs/INVENTORY-MANIFEST.json (add gsd-run-resolver.md + onboard-projection.cjs), and tests/workflow-size-baseline.json.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:09 +00:00
Cursor Agent
83da2e1ca9 fix(onboard): enforce complete map gate in fast mode and add skip-ingest rerun
Remove the projectExists guard so fast mode with a partial map still routes
to complete-map-before-new-project even after project planning exists.

Add the missing onboard rerun instruction to the skip-mapping docs-ingest
handoff so the onboarding loop can continue after ingest.
2026-07-05 19:16:09 +00:00
Cursor Agent
ee2ddbae53 fix(onboard): add onboard rerun handoffs and correct summary next step
Add missing 'Then rerun onboard' instructions after new-project handoffs
so brownfield onboarding returns to create SUMMARY.md per REQ-ONBOARD-05.

Use handoff_commands.manager instead of next_action.reason in the summary
template so persisted SUMMARY.md recommends the correct post-onboarding step.
2026-07-05 19:16:09 +00:00
Jeremy McSpadden
2f2d33aea7 no-mistakes(review): Format onboard handoffs by runtime 2026-07-05 19:16:09 +00:00
Jeremy McSpadden
a845a7d231 no-mistakes(review): Preserve partial-planning skip guard 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
c9d964243a no-mistakes(review): Fix onboard fast-map and skip routing 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
cd5971fd3f no-mistakes(review): Fix onboard skip handoffs 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
a5298c1fc0 refactor: project onboard routing in init 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
afc3309b61 no-mistakes(review): Forward onboard fast init flag 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
1c992bf57d no-mistakes(review): Stop fast onboard dead-end 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
41015129e1 no-mistakes(review): Derive onboard map status before summary prompt 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
3e0dbf6cf4 no-mistakes(review): Label fast onboard partial maps 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
6f2c2c2fd6 no-mistakes(review): Anchor onboard summary writes 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
b3555b103d no-mistakes(review): fix onboard fast root anchoring 2026-07-05 19:16:07 +00:00
jeremymcs
1797207280 fix(onboard): route onboard under ns-project and drop from core profile
Integrate the brownfield /gsd:onboard skill into the skill subsystems so the
full CI suite passes:

- Route onboard under commands/gsd/ns-project.md (requires + routing row) so it
  nests as gsd-ns-project/skills/onboard on nested-layout runtimes instead of
  leaking as a 7th top-level skill dir (fixes install-nested-layout + issue-69).
- Remove onboard from PROFILES.core (src/install-profiles.cts) so the frozen
  main-loop core stays at 8 skills; onboard remains in standard/full.
- Add the TEXT_MODE plain-text fallback note to gsd-core/workflows/onboard.md
  for non-Claude runtimes (#2012).
- Allowlist onboard.md as a user-invocable skill (enh-2790 ratchet).
- Regenerate docs/INVENTORY-MANIFEST.json, golden-install-parity fixtures, and
  the workflow size baseline to match.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:07 +00:00
Cursor Agent
fb5d3abb57 fix(onboard): suggest new-milestone instead of blocked new-project
When PROJECT.md exists but planning files are incomplete, onboarding
previously listed /gsd:new-project as a remediation option. That command
errors when the project is already initialized, leaving users at a dead
end. Route partial planning to /gsd:new-milestone instead.
2026-07-05 19:16:07 +00:00
Cursor Agent
7b3bf9be3f fix: align new-project map gate and guard onboarding summary overwrite
init new-project now uses the same seven-file codebase map completeness
check as init onboard, so partial .planning/codebase/ directories no longer
skip the brownfield mapping offer after onboarding warns about an incomplete
map.

The onboard workflow now branches on onboarding_summary_exists and asks for
confirmation before regenerating SUMMARY.md on repeat runs.
2026-07-05 19:16:06 +00:00
Jeremy McSpadden
a3cca0704d no-mistakes(document): Sync onboard documentation 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
b006b1a23e no-mistakes(review): Fix doc-only onboarding route 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
66ff521c71 no-mistakes(review): Fix onboard runtime and doc detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
f29f981486 no-mistakes(review): Fix onboarding docs gate detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
6b0b5f1b2c no-mistakes(review): Fix onboard planning detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
e0196d5369 no-mistakes(review): Fix onboarding artifact status 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
896c2740d3 feat(#1990): add onboard command for brownfield setup 2026-07-05 19:16:06 +00:00
Tom Boucher
ef8a3e27d4 fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy (#2014)
* fix(#1864): balance <step> tags in settings-advanced.md §8 Model Policy

§8 Model Policy ended with </step> but had no matching opening tag (5 opens /
6 closes), leaving it as loose inter-step content. Add the missing
<step name="model_policy"> opener so the section is a proper step.

- gsd-core/workflows/settings-advanced.md: add <step name="model_policy">
- tests/workflow-step-tag-balance.test.cjs: regression guard — every top-level
  workflow must have balanced <step>/</step> (fenced code stripped), plus a
  focused assertion that §8 is wrapped in model_policy.
- goldens + workflow-size baseline recaptured.

Closes #1864

* docs(#1864): backfill changeset pr 2014
2026-07-05 14:38:24 -04:00
Tom Boucher
a62079b2da fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR (#2024)
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR

The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').

The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.

- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
  + explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.

Closes #1865

* docs(#1865): backfill changeset pr 2024
2026-07-05 14:20:52 -04:00
Behruz Nassre Esfahani
23254ca5a7 fix(#1936): reconstruct OpenCode review from JSON events (#1992)
* fix(#1936): reconstruct OpenCode review from JSON events; diagnosable empty-output stub

On a large review prompt, OpenCode's default `build` agent runs a few read
tool calls then ends its turn with zero output tokens (reason:"stop",
output:0), so `opencode run --format default` emits empty stdout. The reviewer
block redirected stderr to /dev/null and wrote a generic "failed or returned
empty output" stub — so the phase silently lost its second independent reviewer
with no diagnostic and no timeout.

Rewrite the OpenCode reviewer block to invoke `--format json` as the primary
call and reconstruct the review from the assistant `text` parts (jq). Capture
stderr to a `.err` sidecar (mirrors the Codex block). When the agent emits no
text, surface the stop reason, output-token count, and stderr so the failure is
diagnosable. Gate the stub on the extracted CONTENT, not the output file size —
an empty jq extraction still prints a lone newline that a `[ -s file ]` check
would treat as populated. Document the wall-clock timeout as a Bash-tool param
(macOS lacks GNU timeout; opencode has no native timeout flag).

review.md was already at the DEFAULT size-tier ceiling (40956/40960), so the
fix cannot fit without reclassifying it into the LARGE tier (it is a
multi-reviewer orchestration file that outgrew "focused single-purpose"; 43.4 KB
sits well under the LARGE high-water mark). Recapture the 16 golden-install
fixtures — the diff is exactly one review.md hash per runtime. Regression block
folded into review-default-reviewers-workflow.test.cjs (new bug-NNNN test files
are not accepted).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1936): add changeset

* test(#1936): property-test the OpenCode review jq reconstruction

Address the re-review's one actionable finding: the jq JSON-event → text
reconstruction had no fast-check property test.

Add tests/opencode-review-reconstruction.property.test.cjs. It extracts the two
shipped jq programs (OPENCODE_REVIEW, OPENCODE_DIAG) verbatim from
gsd-core/workflows/review.md and runs the real jq — not a reimplementation — so
the shipped logic is what gets tested. Properties: the reconstructed review
equals the newline-join of every assistant text part (order preserved); a stream
with no text part reconstructs to empty (drives the #1936 stub); null/absent text
parts are dropped, never rendered as "null". Plus example-based coverage of the
diagnostic edges the reviewer cited: missing .tokens.output and no step_finish
degrade to "?"; non-JSON stdout makes jq fail rather than masquerade as a review.

Verified the invariant has teeth (a comma-join jq fails the property).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test when jq is absent

The property test shells out to `jq`, which GitHub's windows-latest runners do
not ship (macOS/Linux runners do). `execFileSync('jq')` therefore ENOENT-failed
the whole file on `test (windows-latest, *)`. Probe `jq --version` at load and
skip the suite when jq is not on PATH — the reconstruction logic is
platform-independent, so the assertions still run in full on every jq-present
runner (mirrors how golden-install-parity skips on win32).

Verified: jq present → 7 pass; jq removed from PATH → 7 skipped, 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(#1936): skip jq reconstruction property test on Windows, not just when jq is absent

The prior guard skipped only when `jq` was absent from PATH — but the
windows-latest runners DO ship jq, so the suite still ran there and failed with
`jq: parse error: Invalid numeric literal` (confirmed from the CI job log). Root
cause is Node's child_process argument quoting mangling the jq program (it embeds
double quotes) on Windows, not the shipped review.md logic — the macOS/Linux legs
pass. Gate the suite on `process.platform === 'win32'` (still also skipping when
jq is absent), mirroring golden-install-parity's win32 skip. Logic is
platform-independent and fully asserted on every macOS/Linux CI leg.

Verified: macOS → 7 pass; simulated win32 → skips.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 14:07:07 -04:00
Tom Boucher
8de2ff9121 feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator

Third-party capability gates declared via check.predicate were rendered for
display but never evaluated (only built-in check.query gates fired; the
security capability's gate worked solely via a hard-coded ship.md branch).

Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts)
that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a
bounded sh -c command at the project root (via shell-command-projection.execTool),
inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed.

Wire a 'check predicate' subcommand into check-command-router.cts and extend
the three generic workflow gate-dispatch sites (execute:wave:post, execute:post,
plan:post) to route check.predicate gates to the new evaluator. The two-step
gate contract (command-failure => onError; block => halt) is unchanged.

- src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible
- src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags
- docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md
- tests: 38 unit + integration tests (exit mapping, timeout, interpolation,
  property-based bijection, malformed-predicate fail-closed, real subprocess e2e)

Closes #2008

* docs(#2008): backfill changeset pr number 2011
2026-07-05 14:04:29 -04:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Rezolv
55604e9124 fix(#1906): require node-test clean-fixture causation control (#2001)
* fix(#1906): require node-test clean-fixture causation control

The node-test fail-first proof accepted a deceptive content-independent
negative test — one that reds merely because GSD_PROHIB_SUBJECT is set,
ignoring the subject's content — whenever no cleanFixture was supplied,
because #1346's causation control was opt-in. The proof's observed signal
(RED) thus diverged from its target (RED caused by content) by default.

Make the causation control mandatory for the node-test kind: a descriptor
that omits cleanFixture is un-provable (fail-closed), never accepted under
the weaker violation-only proof. When a clean fixture is present, fail-first
is proven exactly as before (RED on violation AND non-vacuous GREEN on clean).
The lint-rule kind is unchanged (its subject IS the linted file; no
GSD_PROHIB_SUBJECT indirection).

Breaking (Hyrum): a previously-green node-test prohibition with no clean
fixture now hard-gates — blast radius is zero in-tree (no node-test
prohibition ships today; only the lint-rule local/no-source-grep dogfood).

Supersedes ADR-1606 Decision 4 / ADR-550 #1346 addendum's opt-in.

Closes #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ

* docs(#1906): supersede the #1346 opt-in causation control (mandatory for node-test)

Record the node-test mandatory-causation-control supersede across the
governing surfaces:

- ADR-1606 (the enforcement decision-of-record): addendum + Decision 4
  annotated + the "Mandatory causation control — REJECTED" alternative
  flipped to accepted (premise no longer holds: zero in-tree node-test
  consumers).
- ADR-550: the 2026-06-21 #1346 "Why opt-in, not required" paragraph
  marked SUPERSEDED, pointing at ADR-1606.
- spec-phase.md: check_clean_fixture is now REQUIRED for node-test
  (was "optional").
- CONTEXT.md: PROHIB.enforce.causation predicate updated.

Regenerated the shipped-artifact cascade from the spec-phase.md edit
(+149 B, well under the 40960 cap): 16 golden-install-parity fixtures
and the workflow size baseline.

Refs #1906

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
2026-07-03 23:21:55 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Rezolv
9d7c046eae refactor(#1852): lazy-split plan-phase.md into steps/ (#1934)
* refactor(#1852): lazy-split plan-phase.md into steps/

Extract 3 self-contained, rarely-hit sections into gsd-core/workflows/plan-phase/steps/
via lazy 'Read and execute' pointers (mirrors execute-phase/steps/, ADR-1610 progressive
disclosure — no eager @-import): closed-phase-gate (1.5), prd-express-path (3.5),
windows-troubleshooting. Byte-invariant: plan-phase.md 94459 -> 89775 (-4684), each step
< 32 KiB anchor, resolved instructions unchanged.

Scope note: only 3 sections were extractable. plan-phase.md is guarded by a dense net of
content-presence tests (plan-bounce/enh-3209/phase6-planning-capabilities assert specific
sections/flags inline) that block extracting the larger blocks without also refactoring
those tests — durable low-80s headroom is deferred pending maintainer re-scope (issue #1852).

Cascade: size:baseline regen, 16 golden-install-parity fixtures regen, INVENTORY in sync.
Full plan-phase test surface green (4129/4130, 0 fail); lint:ci exit 0.

* docs(changeset): Changed fragment for #1934 (plan-phase lazy-split)

* docs(changeset): mark #1934 fragment docs-exempt (internal workflow refactor)

* fix(#1852): add gsd_run launcher preamble to prd-express-path step

The extracted prd-express-path.md calls gsd_run but the canonical launcher
preamble lived in the parent plan-phase.md — runtime-launcher-parity (#373)
walks workflows/ recursively and requires every .md using gsd_run to carry
exactly one preamble + the $HOME/.claude fallback arm (same as the existing
execute-phase/steps/ files). Injected via scripts/sync-runtime-launcher.cjs;
golden fixtures + size baseline regenerated. plan-phase.md unchanged (89775).

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 23:30:19 -04:00
Behruz Nassre Esfahani
7bef6a6496 fix(#1863): use named flags for state.* calls in executor + workflows (#1873)
* fix(#1863): use named flags for state.* calls in executor + workflows

The named-only state-command router (parseNamedArgs) silently drops
positional args, so state.cjs threw its required-arg error and
metrics/decisions/blockers/session continuity were never recorded.

Convert record-metric / add-decision / add-blocker / record-session in
agents/gsd-executor.md to the named-flag form (mirroring execute-plan.md),
and fix the two remaining positional record-session calls in
gsd-core/workflows/milestone-summary.md and forensics.md. Recapture the
golden-install-parity fixtures and size baselines for the edited files.

Also fix a pre-existing detached-rebuild handle leak in
tests/graphify-auto-update.slow.test.cjs: three dispatch tests returned
after observing only the synchronous "running" status without awaiting the
detached rebuild's terminal state. That leak was latent until the new
#1863 regression block's added runtime shifted --test-force-exit timing
and surfaced it as a non-zero chunk exit. The three tests now await
terminal status via the file's existing waitForBuildStatus helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1863): add changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-02 13:01:10 -04:00
jecanore
5657994702 fix(#1716): route resume_from_file to complete_session when no pending tests remain (#1722)
* fix(#1716): route resume_from_file to complete_session when no pending tests remain

When a UAT session has status:partial with blocked_count>0 and pending_count==0 (all remaining tests are blocked, none are pending), resume_from_file found no [pending] test and terminated silently — never routing to complete_session. This blocked the issues==0 auto-transition path even when there were zero code defects.

Guard clause added immediately after the find-pending step: if no [pending] test is found, route to complete_session. complete_session then correctly sets status:partial (because blocked_count>0) without presenting further tests.

Closes #1716

* chore(#1716): add changeset fragment and regenerate golden-install-parity fixtures

Changeset fragment for PR #1722 (type: Fixed).

Golden-install-parity fixtures regenerated for all 16 runtimes — the workflow fix shifts verify-work.md's byte-stable hash in the golden manifest. Regenerated via UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
2026-07-02 11:58:06 -04:00