Commit Graph

178 Commits

Author SHA1 Message Date
Cody Anderson
98e4233ce9 fix(#2176): ground the Antigravity reviewer in the repo under review (#2184)
* fix(#2176): ground the Antigravity reviewer in the repo under review

- capability-probe --add-dir (mirrors the Codex bypass-flag probe) and pass
  the repo root on both invocation arms
- anchor _AGY_PROMPT to the absolute repo root; mandate a
  REVIEWED-WITHOUT-REPO-ACCESS self-report when the repo is unreadable
- stamp a [reviewed-without-repo-access] marker on self-reported or
  scratch-anchored output; Consensus Summary down-weights marked reviews
- apply the same absolute-root anchor to the cursor-agent prompt (AC5)

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2176): changeset fragment for PR #2184

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): review fixes — size baseline, cursor root anchor, anchored blind tells

- regenerate tests/workflow-size-baseline.json for review.md's growth
- cursor anchor uses git rev-parse --show-toplevel (bare pwd resolved the
  wrong root from a repo subdirectory)
- blind-review tells anchored: self-report to the first lines of output,
  scratch tell to a workspace-declaration phrasing — a grounded review
  quoting either string is no longer mis-stamped

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): round-2 review fixes — scratch-tell bridge, behavioral test, changeset

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the review.md change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): pass the transcript path to bash with forward slashes

The behavioral detection test substitutes a mkdtemp path into the bash
compound; on Windows runners that path contains backslashes, which bash
strips, so the transcript is never found and the first assertion fails
(windows-latest/24 lane). Git Bash accepts D:/-style paths.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2176): use /gsd:review namespace syntax in workflow comment

The slash-command namespace invariant (#3443) bans retired /gsd-<cmd>
references in Claude-facing sources; a cursor-anchor comment used
/gsd-review. Size baseline + golden fixtures regenerated for the byte
change.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): derive the POSIX path via path.sep, not a hardcoded separator

Review finding: out.replaceAll('\\', '/') hardcodes both separators;
use the separator-safe out.split(path.sep).join(path.posix.sep) idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test(#2176): use the merged toPosixPath seam for the bash path

Per maintainer note: #2247's shell-command-projection now centralizes
running-OS → POSIX path conversion; import it instead of the inline
split/join idiom.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 15:54:19 -04:00
Adnan
3592697bed fix(#2107): orchestrator honors gate="blocking-human" checkpoints in auto-mode (#2113)
* fix(execute-phase): honor gate="blocking-human" in auto-mode checkpoint handling

The package-legitimacy gate (#2827) spans two layers. gsd-executor refuses to
auto-approve a gate="blocking-human" checkpoint and escalates it so a human can
vet the package. execute-phase's checkpoint_handling step then dispatched purely
on checkpoint *type* and never read gate -- so under --auto/--chain it
auto-approved the checkpoint the executor had just refused to auto-approve.

Net effect: the slopsquatting defence was inert in exactly the unattended mode
where it matters. An [ASSUMED]/[SUS] package reached install with no human ever
seeing the prompt.

- gsd-core/workflows/execute-phase.md: carve out gate="blocking-human" (and the
  package-legitimacy what-built markers) ahead of every auto-mode branch.
- gsd-core/references/checkpoints.md: document the gate attribute and its two
  values. blocking-human previously appeared nowhere outside gsd-executor.md,
  so no planner had a documented way to author a non-auto-approvable checkpoint.
- tests/package-legitimacy-gate.test.cjs: the existing regression test asserted
  the executor half only, which is why it stayed green while the gate was open.
  Now asserts the orchestrator half too.

* chore(changeset): link to issue #2107

* chore(changeset): backfill PR number 2113

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* test(#2107): refresh golden-install-parity hashes for edited gsd-core files

The golden fixtures pin content hashes for gsd-core/references/checkpoints.md
and gsd-core/workflows/execute-phase.md, both edited by this fix. Regenerated
via UPDATE_GOLDEN=1; only those two keys change across all 17 runtime fixtures.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): keep the carve-out inside the ADR-857 host-loop budget

The ADR-857 phase-6 ratchet pins execute-phase.md below 93600 LF bytes so
optional-feature logic keeps migrating out of the host loop. The carve-out
first landed 623 bytes over that ceiling.

Move the two-layer rationale (why gsd-executor escalates these checkpoints)
into references/checkpoints.md, where the gate is now documented, and reduce
the workflow to the operative rule. execute-phase.md is 93589 bytes, under
the ceiling; the gate token and both <what-built> marker strings are kept
because the orchestrator matches on them.

Refresh the two baselines the edit invalidates: golden-install-parity
fixtures (only the checkpoints.md and execute-phase.md hashes move) and
workflow-size-baseline.json (one line). The ADR-857 ceiling itself is
untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa

* fix(#2107): executor honors blocking-human on the decision branch + gate transport

Review found the fix incomplete one layer down. Two executor-layer gaps:

1. Blocker — agents/gsd-executor.md auto-mode dispatch gated
   checkpoint:human-verify on gate="blocking-human" but the checkpoint:decision
   branch below auto-selected the first option with no gate check. The executor
   resolves a decision itself (auto-selects and continues) without returning it,
   so the orchestrator carve-out never runs for it. A planner following the new
   checkpoints.md rule 6 ("gate a decision whose default would be wrong to
   assume") would have it silently auto-selected under --auto/--chain — the exact
   #2107 harm, one checkpoint type over. The decision branch now STOPs and
   returns for an explicit human decision when gate="blocking-human".

2. Major (transport) — checkpoint_return_format carried no field conveying the
   gate to the freshly-spawned orchestrator, so recognition of the proactive
   pre-install checkpoint rested on freeform prose. Added a **Gate:** field to
   the return format and re-pointed the execute-phase carve-out at it
   ("If the returned Gate: is blocking-human"). Net byte-negative: execute-phase.md
   drops 93589 -> 93583, widening ADR-857 headroom from 11 to 17 bytes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): cover decision carve-out + gate transport, de-vacuum conditional tests

- New: 'auto mode does not auto-select a blocking-human decision checkpoint'
  asserts the executor decision branch STOPs on blocking-human. Verified red on
  the pre-fix executor (2 fail), green with the fix (27 pass).
- New: 'checkpoint_return_format transports the gate ...' asserts the **Gate:**
  field carries blocking-human across the executor->orchestrator boundary.
- New: 'auto-select rule for decision is conditional' — orchestrator-side mirror
  of the human-verify conditional test, for the execute-phase decision branch.
- Fix vacuous test: both conditional tests now assert the anchor matched
  (length > 0) before iterating, so anchor drift can no longer pass with zero
  assertions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#2107): refresh golden + size baselines for executor + execute-phase edits

Regenerated via UPDATE_GOLDEN=1 and update-size-baseline.cjs. Only the
gsd-executor.md and gsd-core/workflows/execute-phase.md hashes move across the
runtime fixtures (35 ins / 35 del, no keys added or removed); checkpoints.md is
unchanged this round. Size baselines: gsd-executor.md 43607 -> 43973,
execute-phase.md 93589 -> 93583 (still under the ADR-857 ceiling).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:43:29 -04:00
Tom Boucher
b5ce72f729 fix(#2119): single SECURITY.md writer — auditor is return-only (#2154)
* fix(2119): single SECURITY.md writer — auditor is return-only

The gsd-security-auditor held Write/Edit and was instructed to write
SECURITY.md (no <N>- prefix, no template frontmatter), while the
orchestrator's Step 6 also wrote the correct padded <N>-SECURITY.md
from templates/SECURITY.md. Two writers, two naming conventions, two
shapes — the auditor's unprefixed file was invisible to the workflow's
*-SECURITY.md glob detector and unparseable for the threats_open gate.

Fix (option 1 from the issue): make the auditor return-only.
- Remove Write/Edit from auditor's tools
- Rewrite all 'Write SECURITY.md' instructions to 'Return structured
  verdict' with threats_open count
- Add explicit constraint in workflow Step 5 spawn prompt
- Update existing test (was asserting Write in tools — now asserts absence)
- Add new regression test for single-writer contract
- Update docs/AGENTS.md stale Tools/Produces rows
- Regenerate golden fixtures + agent size baseline

* docs(changeset): backfill PR number (#2154)

* chore(#2119): regenerate pi/qwen golden fixtures after next merge

The single-writer change edits gsd-core/workflows/secure-phase.md and
agents/gsd-security-auditor.md; pi.json (added on next) and qwen.json (merge
straggler) were the only runtime fixtures still holding pre-change hashes for
those files. All other runtimes already reflect the change. Regenerated via
the sanctioned gen-golden-install-parity script.

* merge origin/next — regenerate goldens + baseline for merged state

* fix slash-command syntax: /gsd-secure-phase → /gsd:secure-phase (#2154 CI fix)
2026-07-13 00:47:15 -04:00
Tom Boucher
4bb846b67a fix(#2112): scope commit to --files pathspec, not entire index (#2148)
* fix(2112): scope commit to --files pathspec, not entire index

cmdCommit/cmdCommitToSubrepo/cmdPrSubrepo staged exactly the files
named in --files but then ran a bare 'git commit' with no pathspec,
absorbing anything else in the index into a commit whose message
described only the named files (#2112).

Fix: append '-- ...stagedPaths' to the commit args when the caller
declared a scope. Three guards are load-bearing:
- stagedPaths (not filesToStage) excludes skipped missing files (#2014)
- explicitFiles gate keeps the default .planning/ path byte-identical
- MERGE_HEAD check via 'git rev-parse' falls back to bare commit during merge
- --amend is left without pathspec (different operation)

cmdPrSubrepo pathspec uses changedFiles (old+new for renames) so the
full rename is captured atomically.

Also fixes workflow markdown in spec-phase.md and add-tests.md.

All-files-missing now short-circuits to nothing_to_commit instead of
absorbing the entire index under a message describing files that
were not committed.

* docs(changeset): backfill PR number (#2148)

* test: update golden-install-parity fixtures for workflow markdown changes (#2112)

* test: update golden fixtures + workflow baselines for #2112 changes

- claude-local.json golden fixture (now generated via gen script)
- workflow-size-baseline.json (add-tests.md +16, spec-phase.md +42 bytes)
- Extended gen-golden-install-parity-zcode.cjs to also regenerate the
  claude local-layout fixture
2026-07-13 00:21:46 -04:00
Tom Boucher
f8c5c1590f fix(#2196): declare the debug session-manager spawn foreground + no-TaskOutput + recovery (#2227) 2026-07-12 19:47:42 -04:00
Tom Boucher
24c324cbfa fix(#2194): add Bash timeout guidance for prompt-fed reviewers in review.md
The Gemini, Claude, and Codex reviewer blocks invoked the CLIs with no explicit
timeout, so each inherited the host default (~2 min on Claude Code). A source-
grounded review of a large plan set takes ~570s (Codex xhigh) / ~525s (headless
Claude) — both exceed that window, so the lane is killed mid-review, its output
is empty, and the cross-AI review silently proceeds with fewer lanes. CodeRabbit
and OpenCode already documented a timeout; the four main lanes did not.

Add a shared timeout-guidance note directing a high Bash timeout (>= 900000;
1200000 for Codex xhigh / headless Claude), referencing BASH_MAX_TIMEOUT_MS for
the Claude Code host cap, and framing a slow-lane empty output as a timeout kill
(not the 0xc0000142 crash it gets misdiagnosed as) so operators re-run with more
time instead of diagnosing a CLI failure.

Closes #2194

Recaptures the 18 golden-install-parity fixtures + workflow-size baseline (only
the review.md entry changed in each; review is LARGE-tier, 47168 < 61440).
2026-07-12 18:53:49 -04:00
Tom Boucher
924ff6822e fix(#2138): surface track_shipping push failures (review)
Review (MEDIUM): `git push ... 2>&1` did not check exit code, so a silent push
failure (auth-token expiry, network blip, non-fast-forward) would proceed to the
report step and declare success — silently reproducing the exact #2138 defect.
Add a fallback warning naming the rerun command so a failed push is visible
(best-effort: the PR already exists, so we still report it rather than abort).
Recaptures goldens + size baseline.
2026-07-12 15:55:39 -04:00
Tom Boucher
aaf74878f7 fix(#2138): push the track_shipping ship-note onto the PR branch [ci skip]
track_shipping committed the STATE ship-note ('Phase N shipped — PR #N') AFTER
create_pr and never pushed it, so the commit stayed local-only. When the GitHub
PR merged (especially fast/auto-merge) the ship-note was not in the source branch
and never reached the default branch — STATE's ship-status was silently lost,
recoverable only by STATE self-heal on the next /gsd-start.

Push the ship-note commit onto the PR branch with a [ci skip] trailer. GitHub
honors [ci skip]/[skip ci], so this lands the note on merge without triggering a
redundant pipeline, and preserves the PR number in STATE.

Recaptures the 18 golden-install-parity fixtures + the workflow-size baseline
(only the ship.md entry changed in each).

Closes #2138
2026-07-12 15:55:39 -04:00
Tom Boucher
78b9b04f1d fix(#2133): correct fast.md log_to_state column-count gate (NF-2)
The guard used `awk -F'|' '{print NF-1}'` but a markdown header has a leading
and trailing pipe, so NF counts (real columns + 2); NF-1 is always one too
high. The `-eq 5` test was therefore unsatisfiable for the very 5-column
header quick.md writes, so /gsd-fast has never appended a Quick Task row since
PR #85 (regression closing #27).

- Count real columns with NF-2 (5 for the 5-col header, 6 for the 6-col).
- Accept 5 OR 6 columns (quick.md Step 7b writes both shapes).
- Select the appended row template by the detected count so its cell count
  always matches the header — keeps #27 fixed for the validate-mode table.

Closes #2133
2026-07-12 10:27:55 -04:00
Tom Boucher
b55a2e7655 fix(#2117): distinguish not-yet-validated phase from validated failure in audit-milestone
audit-milestone's Nyquist scan classified a phase from `nyquist_compliant`
alone, so a phase seeded by plan-phase but never run through validate-phase
read PARTIAL — identical to a phase that validated and genuinely failed. The
template's `status` field could discriminate the two, but no workflow ever
promoted it off `draft`, so it was dead.

Make `status` live and read it:
- validate-phase.md §6: set `status: validated` in both the create (State B)
  and update (State A) VALIDATION.md paths.
- audit-milestone.md §5.5: parse `status`; add a distinct NOT-VALIDATED bucket
  keyed on `status: draft`, gate COMPLIANT/PARTIAL on `status: validated`, and
  report `not_validated_phases` in the audit YAML.
- VALIDATION.md template: document the draft → validated lifecycle.

Tests & generated artifacts:
- Regression test folded into policy-138 (owning workflow-contract file);
  fail-first verified vs origin/next (0 matches pre-fix).
- Regenerate golden-install-parity fixtures cleanly: adds the previously-missed
  qwen.json and removes a contaminated `settings.local.json` entry that had
  leaked into claude-local.json (the harness excludes hook-config files).
- Correct a stale validate-phase.md workflow-size-baseline entry.

Closes #2117

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 22:47:51 -04:00
Tom Boucher
79d7657eff feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.

Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
  slash-programmatic / active-model / native-extension / bun) + hostBehaviors
  {nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
  RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
  payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
  by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
  EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
  surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
  from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
  (global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
  markdown); the 16 other fixtures + claude-local change only by the shared
  model-catalog hash line.

Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
  subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
  in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
  (every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
  gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
  a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
  shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
  query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
  consumes params; getArgumentCompletions; before_provider_request active-model
  steering (fail-open on null resolution); functional session_start /
  before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
  hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
  vocabulary.

Docs (host-integration matrix + how-to) + changeset (Added).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 21:07:38 -04:00
Tom Boucher
f014ec83bd feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads:
finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks
copy), and a skills converter-name registry (the artifactLayout.converter
field is now load-bearing, not decorative). frontmatterDialect stays the
documented dispatch key for frontmatter (no descriptor field for it). Dead
isKilo destructure bindings removed. Byte-identical golden parity for all 16
runtimes (opencode, which shares kilo's combined-family path, verified clean).

UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin +
extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus).
UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread
modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped
from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's
mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is
the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented'
per AC so dispatch degrades to 'degraded' by design.

Model-catalog single-source edit ripples the shared model-catalog.json hash
into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale-
bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/
codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the
connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error.

Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/
hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades
(plugin parity+load, model-override converter, agents dispatch surface, MCP
doc). Matrix + how-to + config docs updated; changeset added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 19:55:15 -04:00
Dave
a8ff8fbb30 docs(#1867): specify propose-then-confirm --auto behavior in ui-phase Step 9.5 (review Minor) 2026-07-10 14:41:12 -04:00
Dave
31a500b970 docs(#1867): replace stale plan-phase.md:921 line-pointer with section-name reference (review #6) 2026-07-10 14:40:03 -04:00
Dave
830670f288 docs(#1867): clarify Step 9.5 manual element-extraction is by-design (review Minor 1)
trek-e's re-review asked to either mechanize the Step 9.5 element
extraction or add an inline note justifying why it stays manual, so a
future maintainer doesn't read it as an oversight.

Verified the premise: the requirement-side edge-probe path (spec-phase
Step 5.5) it mirrors is ALSO a hand-populated heredoc + fail-loud
<replace:> placeholder guard — not a mechanical parse. The UI probe
mirrors that idiom verbatim. Mechanizing would be worse: a UI-SPEC has no
single machine-parseable "elements" column (surfaces are spread across the
design-token tables, Copywriting, and researcher-named prose), so a
regex/table parse would fail-OPEN (miss a prose-named surface, or feed a
design-token row as a bogus element).

Expanded the inline comment to state the parity + the fail-open rationale
explicitly. No logic change; regenerated the workflow size baseline and
the 16 golden-install-parity fixtures for the +847 B comment.

Refs #1867

Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
2026-07-10 14:40:03 -04:00
Dave
bea7196c4c feat(#1867): wire ui-consideration probe into ui-phase (WIRE-01)
Add the live ui-phase producer path for the UI-consideration probe (Phase 2,
WIRE-01). Two small exports on the Phase-1 adapter — proposeElements (the
propose-then-confirm view of detected kinds + applicable categories) and
autoResolve (the deterministic --auto floor that never dismisses and never
auto-backstops an unclassified item, #1110) — plus a post-verification
'## 9.5 UI-Consideration Probe' step in ui-phase.md mirroring spec-phase 5.5's
RUNTIME_DIR shim + fatal-invoke/malformed-report/zero-applicable fail-closed
guards, propose-then-confirm (the partial-cue recall mitigation), and the
'## UI Considerations' write-back in the shipped plan-phase.md:921 lift format.

autoResolve is the CODE floor; the covered-upgrade stays workflow prose (the
two-layer --auto). Un-upgraded backstops route to insufficient_spec ->
human_needed at verify, never a silent pass (#1154).

Tests: +8 typed (proposeElements shape/determinism, autoResolve never-dismiss,
partial-cue strict-subset) — structured-value only. ui-phase.md size baseline
ratcheted 15477->24447 (under DEFAULT cap). The plan-phase.md PRE_PHASE6 ceiling
stays RED pending #1852 (unchanged from Phase 1).

Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS
2026-07-10 14:40:03 -04:00
Dave
6d7450c465 feat(#1867): lift UI Considerations into plan-phase must_haves (LIFT-01)
Separate *-UI-SPEC.md glob (edge-coverage glob exclusion at :746 left intact
- Hyrum/D-08); a terse lift bullet reusing the ## Edge Coverage rule verbatim
(covered -> truths string, backstop -> flat scalar {statement, verification:
backstop}, unresolved -> assumption; no new verb - ADR-550 #1278/#1154); and a
no-silent-drop checklist line. Lift logic lives ONLY in the plan-phase
workflow, never gsd-planner.md (D-10, agent-size cap).

NOTE: this grows plan-phase.md +862B, over the #1168 PRE_PHASE6 ceiling (94519)
by ~802B, so the workflow-size gates are RED until #1852's lazy-split lands the
-21KB headroom. Deliberately does NOT raise the ceiling constant (would collide
with #1820/#1835's in-flight raise). #1867 sequences after #1852; rebase +
regenerate the size baseline then.
2026-07-10 14:40:03 -04:00
Tom Boucher
185abe2d66 feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.

Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium

Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.

Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.

Closes #2122
2026-07-10 12:11:55 -04:00
Tom Boucher
dbc730d8de fix(#2073): capability-probe external killer (timeout/gtimeout) for macOS
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock
macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for
'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout
alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s)
stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor
the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both
branches. Update the agy + #687 tests to assert the probe + bound + fallback,
regen the 17 goldens + size baseline, refresh the maintainer-note version stamp
to 1.0.16.
2026-07-08 17:53:28 -04:00
Tom Boucher
af9069f865 fix(#2073): harden agy reviewer block (arg overflow, 404, pre-session stall)
Three failure modes on agy 1.0.16, all fixed by mirroring the Cursor block's
invocation discipline:
  * file-reference prompt instead of inline "$(cat)" — a large review prompt
    overflowed the exec arg list (rc 126).
  * external 'timeout 600' wrapper — --print-timeout cannot fire before agy
    creates a session, so a pre-session stall hung unbounded.
  * --model from review.models.agy when set — escape hatch for a pinned model
    that 404s (exit 0, empty stdout + transcript).
  * stdin </dev/null so agy never blocks on a tty.
Also enrich the Step 3 empty-output stub to grep agy cli.log for a
model-availability diagnostic, and correct the stale 'no --model flag' note
plus the 'review.models.agy reserved for future' comment (the config key was
already read but never passed through).
2026-07-08 16:47:02 -04:00
Tom Boucher
1ee00320c7 Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns 2026-07-07 22:03:47 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Tom Boucher
603593d41d fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".

Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
  normalize-test-command` verb: rewrites a resolved command to a best-effort
  one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
  a package-manager `test` script whose package.json runner is watch-vitest →
  `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
  never double-flagged). Named `normalize-test-command` (not `test-*`) so the
  file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
  and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
  (new config key, default 600s): the regression gate (extracted to
  execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
  it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
  audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
  it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
  on 124, staying under its frozen 40960-byte tier cap.

Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).

Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.

Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 12:14:43 -04:00
Tom Boucher
ea4378063f Merge branch 'next' into feat/1820-specless-predicate-rail 2026-07-07 07:49:07 -04:00
Tom Boucher
d7129222c0 Merge branch 'next' into codex/gsd-onboard 2026-07-06 23:49:49 -04:00
Tom Boucher
080bacdb4b Merge branch 'next' into codex/gsd-onboard 2026-07-06 22:43:32 -04:00
Dave
7ef834cabc feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them 2026-07-06 14:40:38 -04:00
Joe Slitzker
3fe9a81428 fix(#1941): degrade /gsd-quick worktree dispatch when fork base is stale
Claude Code's isolation="worktree" forks new worktrees from origin/HEAD, not
the live local HEAD. When prior local commits (e.g. an earlier quick task in
the same session, or this task's own Step 5.6 pre-dispatch plan commit)
advance local HEAD without an intervening push, origin/HEAD stays pinned to a
stale ancestor and the executor's worktree_branch_check guard halts with a
base-mismatch fatal that can be many commits behind, not just one.

Port the worktree.base-check auto-degrade pattern already used by
execute-phase (#683/#1369) into quick.md's single-dispatch path, run
immediately before EXPECTED_BASE is captured in Step 6.
2026-07-06 13:01:47 -05:00
Tom Boucher
9f0d785b61 fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows (#2027)
* fix(#2020): remove dead SDK file refs that triggered infinite find.exe on Windows

gsd-executor.md referenced sdk/src/query/QUERY-HANDLERS.md and reapply-patches.md
referenced sdk/dist/cli.js — both retired with the SDK (ADR-0174). AI runtimes
that locate doc refs via filesystem search ran find /, which on Git Bash for
Windows traverses the whole drive (14h+, orphaned find.exe, 4M+ handles).

- agents/gsd-executor.md: drop dead QUERY-HANDLERS.md ref.
- workflows/reapply-patches.md: drop dead sdk/dist/cli.js clause.
- tests/no-dead-sdk-refs.test.cjs: regression guard — no sdk/src|dist|handlers
  file refs in agents/workflows/references markdown.

Closes #2020

* docs(#2020): backfill changeset pr 2027
2026-07-05 16:43:26 -04:00
Tom Boucher
1bfadec2d0 fix(#1921): preserve verify-work state across gap-closure + defer follow-ups (#2025)
* fix(#1921): preserve verify-work state across gap-closure + defer follow-ups

Resuming /gsd:verify-work after /gsd:execute-phase --gaps-only lost the
verification state: UAT ## Gaps still read status: failed even after their
fix plans executed, so verify-work re-diagnosed them as fresh blockers,
spawned a new gap plan, and reported only the new plan verified. A new
state contract links each gap to its fix plan so fixed gaps are recognized.

- verify-work.md gap YAML: stable gap_id (G-{phase}-{N}) per gap.
- plan_gap_closure planner prompt: each *-PLAN.md tags the gap_ids it
  addresses in its frontmatter (gap_closure: true, gap_ids: [...]).
- new reconcile_gaps step (run at resume_from_file entry): marks a gap
  status: resolved when its plan has a matching *-SUMMARY.md — fixed gaps
  are not re-diagnosed and do not spawn new gap plans; a re-reported break
  is treated as a fresh regression with a new gap_id.
- deferred-follow-up branch: a future-work idea (signals: 'later',
  'next version', 'out of scope', ...) is captured to UAT ## Deferred
  Follow-Ups instead of becoming a blocking gap/plan.
- workflow-size baseline recaptured (verify-work LARGE tier, 38221/61440).

Closes #1921

* docs(#1921): backfill changeset pr 2025

* fix(#1921): balance <step> tags in verify-work.md via 4-backtick outer fence

The plan_gap_closure step nested a ```yaml example inside a ``` block;
same-length fences made stripFencedCode close the outer block early,
swallowing the step's </step> (14 opens / 13 closes). Widen the outer
fence to 4 backticks so the nested yaml is contained. Regenerate golden
fixtures + size baselines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 15:42:10 -04:00
Codesmith
425d2f2c8b fix(#1990): keep onboard resolver delegation in sync with canonical launcher
onboard.md delegates the gsd_run preamble to
gsd-core/references/gsd-run-resolver.md via @-include, but two guards
regressed once the canonical launcher snippet advanced to
${CLAUDE_CONFIG_DIR:-$HOME/.claude} (#2024):

- runtime-launcher-parity (B2): the resolver reference still shipped the
  old $HOME/.claude arm. references/ is not covered by
  sync-runtime-launcher.cjs, so refresh the resolver bash block to be
  byte-equal to _runtime-launcher.snippet.sh.
- /gsd:onboard command contract: sync-runtime-launcher.cjs had re-inlined
  the preamble into onboard.md (a delegating file). Teach the sync
  transform to strip-but-never-inline files that @-include the resolver,
  mirroring the exemption already in the parity test (B/B2).

Regenerate golden install fixtures and the workflow size baseline for the
smaller onboard.md.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:10 +00:00
Codesmith
4c673e51f3 chore(#1990): resync runtime launcher and regenerate artifacts after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:10 +00:00
jeremymcs
7734dee051 fix(onboard): resolve Tests-lane failures for onboard workflow
- runtime-launcher-parity: recognize workflows that delegate gsd_run to references/gsd-run-resolver.md (onboard.md) and exempt them from the inline-preamble checks; add a compensating byte-equality guard (B2) asserting the reference bash block matches _runtime-launcher.snippet.sh. - onboard.md: document the TEXT_MODE plain-text/numbered-list fallback for AskUserQuestion on non-Claude runtimes (fixes ask-user-questions-fallback, #2012). - Regenerate golden-install-parity fixtures, docs/INVENTORY-MANIFEST.json (add gsd-run-resolver.md + onboard-projection.cjs), and tests/workflow-size-baseline.json.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:09 +00:00
Cursor Agent
83da2e1ca9 fix(onboard): enforce complete map gate in fast mode and add skip-ingest rerun
Remove the projectExists guard so fast mode with a partial map still routes
to complete-map-before-new-project even after project planning exists.

Add the missing onboard rerun instruction to the skip-mapping docs-ingest
handoff so the onboarding loop can continue after ingest.
2026-07-05 19:16:09 +00:00
Cursor Agent
ee2ddbae53 fix(onboard): add onboard rerun handoffs and correct summary next step
Add missing 'Then rerun onboard' instructions after new-project handoffs
so brownfield onboarding returns to create SUMMARY.md per REQ-ONBOARD-05.

Use handoff_commands.manager instead of next_action.reason in the summary
template so persisted SUMMARY.md recommends the correct post-onboarding step.
2026-07-05 19:16:09 +00:00
Jeremy McSpadden
2f2d33aea7 no-mistakes(review): Format onboard handoffs by runtime 2026-07-05 19:16:09 +00:00
Jeremy McSpadden
a845a7d231 no-mistakes(review): Preserve partial-planning skip guard 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
c9d964243a no-mistakes(review): Fix onboard fast-map and skip routing 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
cd5971fd3f no-mistakes(review): Fix onboard skip handoffs 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
a5298c1fc0 refactor: project onboard routing in init 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
afc3309b61 no-mistakes(review): Forward onboard fast init flag 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
1c992bf57d no-mistakes(review): Stop fast onboard dead-end 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
41015129e1 no-mistakes(review): Derive onboard map status before summary prompt 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
3e0dbf6cf4 no-mistakes(review): Label fast onboard partial maps 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
6f2c2c2fd6 no-mistakes(review): Anchor onboard summary writes 2026-07-05 19:16:07 +00:00
Jeremy McSpadden
b3555b103d no-mistakes(review): fix onboard fast root anchoring 2026-07-05 19:16:07 +00:00
jeremymcs
1797207280 fix(onboard): route onboard under ns-project and drop from core profile
Integrate the brownfield /gsd:onboard skill into the skill subsystems so the
full CI suite passes:

- Route onboard under commands/gsd/ns-project.md (requires + routing row) so it
  nests as gsd-ns-project/skills/onboard on nested-layout runtimes instead of
  leaking as a 7th top-level skill dir (fixes install-nested-layout + issue-69).
- Remove onboard from PROFILES.core (src/install-profiles.cts) so the frozen
  main-loop core stays at 8 skills; onboard remains in standard/full.
- Add the TEXT_MODE plain-text fallback note to gsd-core/workflows/onboard.md
  for non-Claude runtimes (#2012).
- Allowlist onboard.md as a user-invocable skill (enh-2790 ratchet).
- Regenerate docs/INVENTORY-MANIFEST.json, golden-install-parity fixtures, and
  the workflow size baseline to match.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
2026-07-05 19:16:07 +00:00
Cursor Agent
fb5d3abb57 fix(onboard): suggest new-milestone instead of blocked new-project
When PROJECT.md exists but planning files are incomplete, onboarding
previously listed /gsd:new-project as a remediation option. That command
errors when the project is already initialized, leaving users at a dead
end. Route partial planning to /gsd:new-milestone instead.
2026-07-05 19:16:07 +00:00
Cursor Agent
7b3bf9be3f fix: align new-project map gate and guard onboarding summary overwrite
init new-project now uses the same seven-file codebase map completeness
check as init onboard, so partial .planning/codebase/ directories no longer
skip the brownfield mapping offer after onboarding warns about an incomplete
map.

The onboard workflow now branches on onboarding_summary_exists and asks for
confirmation before regenerating SUMMARY.md on repeat runs.
2026-07-05 19:16:06 +00:00