* test(#1963): add failing-first prevention/postmortem contract tests
Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless
5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent
error as 'why was that possible?'), the 'why wasn't this caught?' question,
the recurrence-guard taxonomy (regression test / assertion / lint rule / KB
pattern), the KB-entry why_not_caught + recurrence_guard fields with backward
compat, the session-manager prevention summary line, and the Zawinski
scope-boundary (a block, not a subsystem).
Failing-first: reference, archive_session edit, KB schema extension, and
session-manager summary do not yet exist.
* feat(#1963): emit blameless-postmortem Prevention block at resolution
Epic #1957 Phase 3B. At archive_session the debugger now produces a
Prevention block with three blame-free components: a branching 5-Whys causal
chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts
'why was that possible?', never blame), a 'why wasn't this caught?' answer
naming the missed gate (test/typecheck/lint/review/verify), and a concrete
recurrence guard (regression test / assertion / lint rule / KB pattern).
The knowledge-base entry gains two structured fields (why_not_caught +
recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not
just the prior fix. Additive: old entries without the fields still load. The
session-manager compact summary surfaces a one-line prevention summary.
Full rules extracted to gsd-core/references/debugger-prevention.md (slim
archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md updated.
* fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test)
- CRITICAL: the archive_session KB append template omitted Why not caught +
Recurrence guard (only the Entry Format had them) — the feature's core
deliverable silently did not happen. Added both fields to the append template
the agent actually follows (nearest-instruction wins).
- HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were
dead data. Extended the Phase 0 Evidence line to consume why_not_caught +
recurrence_guard when present (absent on old entries — backward compat holds).
- MEDIUM: added a cross-section parity test (every Entry-Format field must also
appear in the append template — the guard that would have caught the
Critical) + a Phase-0-consumption assertion.
- MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses
reasoning_checkpoint.candidate_causes across the four categories.
- MEDIUM: recurrence-guard taxonomy gains type refinement + config-default
change; LOW: added 'build' gate to both surfaces for parity.
- NIT: compact-summary fallback shape ('no gate existed'); verify the guard
artifact exists before recording it.
* test(#1963): anchor Phase-0 consumption test on the specific heading
The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file
(in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop.
Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars.
* chore(#1963): backfill changeset pr number (PR #2410)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests
Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone
>=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning
checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause
note, DEBUG template) plus behavioral schema-invariant checks on two fixtures:
two contributing causes (AND-gate yes) -> both recorded; single-cause
(AND-gate no) -> one root_cause, identical to today.
Failing-first: reference, agent edits, and template note do not yet exist.
* feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger
Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing
root_cause, the debugger enumerates candidate causes across >=2 Ishikawa
categories (code/config/environment/data) and explicitly answers an AND-gate
question. When the AND-gate fires, every contributing cause is recorded, so a
multi-cause fix no longer recurs via the unaddressed second cause.
Resolution.root_cause may hold one OR a small set (additive; single-cause
sessions are byte-identical to today). The Structured Reasoning Checkpoint gains
candidate_causes + and_gate fields; debugger-philosophy.md adds the
single-cause-bias trap.
Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim
Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest +
agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated.
* fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples)
- Reference: the collapse rule now enforces AND-gate self-consistency —
and_gate=yes with a single confirmed cause is flagged as incomplete
(return to Phase 3); a race/timing note clarifies such bugs bridge
categories; the 'byte-identical' backward-compat claim narrowed to
'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every
session'.
- DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface
drift the reviewer flagged); new debug-session-management parity test pins
the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe).
- Scalar-assuming consumers of set-valued root_cause updated: session-manager
compact summaries (319/332), diagnose-only return (1062), archive entry
(1216), ROOT CAUSE FOUND return (1322).
- Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the
fixture describe block honestly as a schema-invariant specification.
- Phase 2 bullet phrasing clarified ('at hypothesis formation, before the
Phase 4 commit').
* test(#1960): parity regex accepts word-form count ('seven-field' or '7-field')
* test(#1960): parity regex counts array-valued YAML keys (no inline value)
* chore(#1960): backfill changeset pr number (PR #2405)
* test(#1958): add failing-first guardrail contract tests
Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the
5-signal fix-acceptance guardrail contract (target test, mutation check,
no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful
degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file
recording, and subprocess bounding.
Failing-first: reference file and agent sections do not yet exist.
* feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger
Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test
(Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix
acceptance: target test, mutation check (Stryker), no-op/behavior-deleting
detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully
when Stryker or a test suite is absent (each skip logged, never a silent pass),
records per-signal results under Resolution.verification, and returns a
FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for
revise / accept-as-debt / abandon.
Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim
routing kept in the agent to respect the agent-size cap). Debug template +
INVENTORY + manifest + agent-size baseline + AGENTS.md updated.
* test(#1958): correct newline-tolerant assertion + regen install-parity goldens
The revert-and-reconfirm assertion collapsed whitespace before matching so
markdown line-wrapping does not break it. Regenerated the golden-install-parity
and install-tree fixtures (npm run gen:golden) to absorb the intentional
gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file
changes to the installed artifact tree.
* fix(#1958): tighten guardrail per orthogonal review
Addresses the isolated reviewer's findings:
- signal 5 now states its recorded-repro dependency and routes the no-repro
case to the degradation row; revert mechanism specified (git stash / git
revert -n); minimality flag tied to diff structure, not revert-ability.
- bounded-subprocesses section now bounds the git subprocess (5-30s) too,
requires argv-array argument passing, and scopes Stryker to the driving
regression test (a mutant killed only by a non-driving test is a finding).
- new test-provenance (security) clause: the driving test must be
agent-authored; bug-report repro scripts are DATA, never executed verbatim.
- tightened 3 contract assertions to bind to specific clauses
(guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding).
- Goodhart framing softened to 'partially-independent'; DEBUG.md template
verification field notes the nested map shape.
* chore(#1958): backfill changeset pr number (PR #2396)
* fix(#1958): add issue ref to allow-test-rule annotation (ADR-456)
CI lint-allow-test-rule-refs requires every allow-test-rule exemption to
carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation
lacked it; this adds (see #1958).
The /gsd-debug orchestrator handled the gsd-debug-session-manager return
with only two literal-string checks (DEBUG SESSION COMPLETE, ABANDONED)
and no else branch, so a usable-but-non-terminal progress summary (the
manager's own turn/context budget exhausted mid-loop, with a valid
on-disk checkpoint) fell through to the user as if the debug were
complete. Same gap at the continue subcommand.
Callee side (agents/gsd-debug-session-manager.md): add an explicit
non-terminal CONTINUE_REQUIRED return marker, distinct from the two
terminal shapes and from a genuine user-input checkpoint.
Orchestrator (gsd-core/workflows/debug.md Sections 4 and 1c): classify
returns exhaustively — recognized terminal markers behave as before,
anything else is non-terminal and auto-resumes by re-spawning the
session manager from the same slug/checkpoint. Anti-loop guard: after
two consecutive no-progress resumes (unchanged next_action/updated),
emit a blocker report instead of looping.
Regression test (source-text contract guard, fix-2196 idiom) asserts
both sections' non-terminal/auto-resume branch, the CONTINUE_REQUIRED
marker, and the anti-loop bound.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1865): runtime launcher honors CLAUDE_CONFIG_DIR
The gsd_run preamble resolved the Claude global install only at
$HOME/.claude/gsd-core/bin/, but the installer honors CLAUDE_CONFIG_DIR —
so a global install redirected via CLAUDE_CONFIG_DIR was invisible to every
gsd_run call (every command failed with 'gsd-tools.cjs not found').
The Claude resolver arm now uses ${CLAUDE_CONFIG_DIR:-$HOME/.claude},
matching the installer + the other runtimes' ${VAR:-default} pattern.
Default $HOME/.claude behavior is unchanged.
- _runtime-launcher.snippet.sh: Claude arm honors CLAUDE_CONFIG_DIR.
- sync-runtime-launcher.cjs re-run: 95 workflows/agents re-synced.
- review.md / discuss-phase.md: trimmed to stay under their byte budgets.
- runtime-launcher-parity.test.cjs: (A) substring updated for the new form
+ explicit #1865 assertion that the snippet honors CLAUDE_CONFIG_DIR.
- goldens + size baselines recaptured.
Closes#1865
* docs(#1865): backfill changeset pr 2024
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes
On runtimes that execute each fenced bash block in a separate shell process
(e.g. Claude Code — documented behavior: each Bash command is a separate
process; inline shell functions and exported vars do not persist between
calls), the once-per-file gsd_run() function was undefined in every block
after the preamble block, and the call was swallowed by
`2>/dev/null || echo "{}"` into silent empty state.
Fix (budget-neutral session-level resolution):
- Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own
location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm
`bin` field (global installs) and shipped to local installs via the recursive
gsd-core/ copy.
- The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"`
to the file named by $CLAUDE_ENV_FILE (Claude Code's documented
env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from
PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline
gsd_run() definition remains the fallback for all other runtimes. The
single-quoted dir neutralizes shell metacharacters at source time.
- Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files.
- XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md
to 93135; legitimate content growth, ratchet-up per #717).
Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper
delegation and end-to-end PATH persistence (sourcing the env file with a
space-bearing install path).
Known limitation: an install path containing a literal single-quote yields a
malformed env-file line and falls back to the status quo (no regression);
rare on sanitized home directories.
Closes#381
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#381): add changeset for gsd_run fresh-shell reachability fix
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit)
Windows Git Bash (msys2) does not honor Node's chmod exec bit for
PATH-executing extension-less scripts, so the bare `gsd_run` command lookup
failed there even though the env-file PATH persistence was correct. The
env-file content assertions (the fix's actual cross-platform logic) still run
on every platform; only the final source-and-execute sub-step is gated to
non-win32. Global installs on Windows are covered by npm's generated bin shim.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver
Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker,
gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks.
On a shim-only install — where gsd-tools.cjs exists under the runtime home but
gsd-tools is NOT on PATH — those calls fail with "command not found" and the
agent silently skips init/state/validate/commit ceremony, deferring to the
orchestrator or bypassing GSD bookkeeping entirely.
were never migrated, so it persisted on Claude Code and every other runtime that
consumes the source agents directly. Only gsd-phase-researcher.md carried a
resolver — and a stale, claude-only truncated one.
Fix (all runtimes):
- Inject the canonical multi-runtime gsd_run preamble (byte-equal to
_runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/
augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of
the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every
command-position bare gsd-tools to gsd_run.
- Upgrade gsd-phase-researcher.md's stale resolver to the canonical one.
- Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the
sync caught and corrected a mis-placed preamble during development).
- Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards
to agents/ so no runtime can silently regress.
Closes#1041
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1041): backfill changeset PR number to 1045
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#581): add Edit to six writer agents' tools so Edit-only discipline is enforceable
Six writer agents (gsd-eval-planner, gsd-ai-researcher, gsd-domain-researcher,
gsd-phase-researcher, gsd-ui-researcher, gsd-debug-session-manager) shipped with
Write but no Edit in their tools: frontmatter. Their spawn prompts instruct
surgical in-place section edits on existing/shared files (notably the AI-SPEC.md
trio writing disjoint sections of the same file), but with no Edit tool they
fall back to whole-file Write — silently clobbering sibling sections
(last-writer-wins) while still reporting success.
Same bug class as #571, fixed for gsd-doc-writer in #575. This adds Edit
alongside the existing Write for all six (Edit placed adjacent to Write, mirroring
the gsd-doc-writer fix). Write is retained; no prompt-body changes; no other agents
touched.
Adds a regression test (tests/agent-frontmatter.test.cjs) asserting each of the
six section-writer agents carries both Write and Edit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#581): add changeset fragment for writer-agent Edit fix
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#581): regenerate changeset via npm run changeset
Replace hand-authored fragment with one generated by the official
scripts/changeset/new.cjs script (correct <adjective>-<noun>-<noun>
filename convention).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fixes#3168
The Claude Code subagent dispatcher tool is named `Agent` (with `subagent_type`
parameter). The `Task*` namespace (TaskCreate, TaskList, TaskGet, TaskUpdate,
TaskOutput, TaskStop) is the separate task-tracker. GSD's commands, workflows,
and agents were partially migrated and still referenced `- Task` / `Task(` in
55 files, causing orchestrators to silently fall back to inline execution when
no `Task` tool appeared on their tool surface.
Changes:
- `commands/gsd/*.md` allowed-tools: replaced `- Task` with `- Agent` in 24
files; removed duplicate `- Task` from autonomous.md (already had `- Agent`)
- `get-shit-done/workflows/*.md`: replaced dispatcher `Task(` → `Agent(` in
29 workflow files (~133 call sites); TaskCreate/List/Get/Update/Output/Stop
left untouched
- `agents/gsd-debug-session-manager.md`: replaced `Task` → `Agent` in tools
frontmatter (the only remaining agent with the wrong name)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(2148): add specialist_hint to ROOT CAUSE FOUND and skill dispatch to /gsd-debug
- Add specialist_hint field to ROOT CAUSE FOUND return format in gsd-debugger structured_returns section
- Add derivation guidance in return_diagnosis step (file extensions → hint mapping)
- Add Step 4.5 specialist skill dispatch block to debug.md with security-hardened DATA_START/DATA_END prompt
- Map specialist_hint values to skills: typescript-expert, swift-concurrency, python-expert-best-practices-code-review, ios-debugger-agent, engineering:debug
- Session manager now handles specialist dispatch internally; debug.md documents delegation intent
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(2151): add gsd-debug-session-manager agent and refactor debug command as thin bootstrap
- Create agents/gsd-debug-session-manager.md: handles full checkpoint/continuation loop in isolated context
- Agent spawns gsd-debugger, handles ROOT CAUSE FOUND/TDD CHECKPOINT/DEBUG COMPLETE/CHECKPOINT REACHED/INVESTIGATION INCONCLUSIVE returns
- Specialist dispatch via AskUserQuestion before fix options; user responses wrapped in DATA_START/DATA_END
- Returns compact ≤2K DEBUG SESSION COMPLETE summary to keep main context lean
- Refactor commands/gsd/debug.md: Steps 3-5 replaced with thin bootstrap that spawns session manager
- Update available_agent_types to include gsd-debug-session-manager
- Continue subcommand also delegates to session manager
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(2148,2151): add tests for skill dispatch and session manager
- Add 8 new tests in debug-session-management.test.cjs covering specialist_hint field,
skill dispatch mapping in debug.md, DATA_START/DATA_END security boundaries,
session manager tools, compact summary format, anti-heredoc rule, and delegation check
- Update copilot-install.test.cjs expected agent list to include gsd-debug-session-manager
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>