From e95186b2a74c6ae5b58450244afa7ea317d89444 Mon Sep 17 00:00:00 2001 From: Lex Christopherson Date: Thu, 15 Jan 2026 09:34:32 -0600 Subject: [PATCH] refactor: update execute commands to spawn gsd-executor subagent - execute-plan.md: spawn gsd-executor instead of general-purpose with template - execute-phase.md: spawn gsd-executor for wave-based parallel execution - Remove template filling, subagent has all logic baked in Co-Authored-By: Claude --- commands/gsd/execute-phase.md | 21 +- commands/gsd/execute-plan.md | 70 ++++-- get-shit-done/workflows/execute-phase.md | 262 ++++++++++++++++++----- 3 files changed, 274 insertions(+), 79 deletions(-) diff --git a/commands/gsd/execute-phase.md b/commands/gsd/execute-phase.md index dd587b4af..14be4e377 100644 --- a/commands/gsd/execute-phase.md +++ b/commands/gsd/execute-phase.md @@ -26,6 +26,8 @@ Context budget: ~15% orchestrator, 100% fresh per subagent. @~/.claude/get-shit-done/references/principles.md @~/.claude/get-shit-done/workflows/execute-phase.md @~/.claude/get-shit-done/templates/subagent-task-prompt.md +@~/.claude/get-shit-done/templates/subagent-verify-prompt.md +@~/.claude/get-shit-done/workflows/verify-phase.md @@ -62,9 +64,18 @@ Phase: $ARGUMENTS 5. **Aggregate results** - Collect summaries from all plans - Report phase completion status + +6. **Verify phase goal** + - Spawn verification subagent (uses subagent-verify-prompt.md) + - Verify must_haves from plan frontmatter against actual codebase + - If gaps found: create fix plans, execute, re-verify (max 3 cycles) + - If human verification needed: present items to user + - Block until verification passes or user approves + +7. **Update roadmap and state** - Update ROADMAP.md, STATE.md -6. **Update requirements** +8. **Update requirements** Mark phase requirements as Complete: - Read ROADMAP.md, find this phase's `Requirements:` line (e.g., "AUTH-01, AUTH-02") - Read REQUIREMENTS.md traceability table @@ -72,14 +83,14 @@ Phase: $ARGUMENTS - Write updated REQUIREMENTS.md - Skip if: REQUIREMENTS.md doesn't exist, or phase has no Requirements line -7. **Commit phase completion** +9. **Commit phase completion** Bundle all phase metadata updates in one commit: - Stage: `git add .planning/ROADMAP.md .planning/STATE.md` - Stage REQUIREMENTS.md if updated: `git add .planning/REQUIREMENTS.md` - Commit: `docs({phase}): complete {phase-name} phase` -8. **Offer next steps** - - Route to next action (see ``) +10. **Offer next steps** + - Route to next action (see ``) @@ -226,6 +237,8 @@ After all plans in phase complete (step 7): - [ ] All incomplete plans in phase executed - [ ] Each plan has SUMMARY.md +- [ ] Phase goal verified (must_haves checked against codebase) +- [ ] VERIFICATION.md created in phase directory - [ ] STATE.md reflects phase completion - [ ] ROADMAP.md updated - [ ] REQUIREMENTS.md updated (phase requirements marked Complete) diff --git a/commands/gsd/execute-plan.md b/commands/gsd/execute-plan.md index 9a1f9e7c7..53205a4ed 100644 --- a/commands/gsd/execute-plan.md +++ b/commands/gsd/execute-plan.md @@ -13,16 +13,15 @@ allowed-tools: --- -Execute a single PLAN.md file by spawning a subagent. +Execute a single PLAN.md file by spawning the `gsd-executor` subagent. -Orchestrator stays lean: validate plan, spawn subagent, handle checkpoints, report completion. Subagent loads full execute-plan workflow and handles all execution details. +Orchestrator stays lean: validate plan, spawn subagent, handle checkpoints, report completion. The `gsd-executor` has all execution logic baked in. Context budget: ~15% orchestrator, 100% fresh for subagent. @~/.claude/get-shit-done/references/principles.md -@~/.claude/get-shit-done/templates/subagent-task-prompt.md @@ -78,9 +77,26 @@ Plan path: $ARGUMENTS ⚡ Executing {phase_number}-{plan_number}: {objective one-liner} ``` -5. **Fill and spawn subagent** - - Fill subagent-task-prompt template with extracted values - - Spawn: `Task(prompt=filled_template, subagent_type="general-purpose")` +5. **Spawn gsd-executor subagent** + + ``` + Task( + prompt="Execute plan at {plan_path} + +Plan: @{plan_path} +Project state: @.planning/STATE.md +Config: @.planning/config.json (if exists)", + subagent_type="gsd-executor", + description="Execute {phase}-{plan}" + ) + ``` + + The `gsd-executor` subagent has all execution logic baked in: + - Deviation rules (auto-fix bugs, critical gaps, blockers; ask for architectural) + - Checkpoint protocols (human-verify, decision, human-action) + - Commit formatting (per-task atomic commits) + - Summary creation + - State updates 6. **Handle subagent return** - If contains "## CHECKPOINT REACHED": Execute checkpoint_handling @@ -205,9 +221,11 @@ All {N} phases finished. -When subagent returns with checkpoint: +When `gsd-executor` returns with checkpoint: **1. Parse return:** + +The subagent returns a structured checkpoint: ``` ## CHECKPOINT REACHED @@ -306,32 +324,38 @@ Wait for user input: **4. Spawn fresh continuation agent:** -Fill continuation-prompt template with: -- completed_tasks_table: From checkpoint return -- resume_task_number: Current task number -- resume_task_name: Current task name -- resume_status: Derived from checkpoint type and user response -- user_response: What user provided -- resume_instructions: Type-specific guidance (see template) +Spawn fresh `gsd-executor` with continuation context: ``` -Task(prompt=filled_continuation_template, subagent_type="general-purpose") +Task( + prompt="Continue executing plan at {plan_path} + + +{completed_tasks_table from checkpoint return} + + + +Resume from: Task {N} - {task_name} +User response: {user_response} +{resume_instructions based on checkpoint type} + + +Plan: @{plan_path} +Project state: @.planning/STATE.md", + subagent_type="gsd-executor", + description="Continue {phase}-{plan}" +) ``` +The `gsd-executor` has continuation handling baked in — it will verify previous commits and resume correctly. + **Why fresh agent, not resume:** -Task tool resume fails after multiple tool calls (presenting to user, waiting for response). Fresh agent with state handoff via continuation-prompt.md is the correct pattern. +Task tool resume fails after multiple tool calls (presenting to user, waiting for response). Fresh agent with state handoff is the correct pattern. **5. Repeat:** Continue handling returns until "## PLAN COMPLETE" or user stops. - -Templates for checkpoint handling: - -- `@~/.claude/get-shit-done/templates/checkpoint-return.md` - Subagent return format -- `@~/.claude/get-shit-done/templates/continuation-prompt.md` - Fresh agent spawn template - - - [ ] Plan executed (SUMMARY.md created) - [ ] All checkpoints handled diff --git a/get-shit-done/workflows/execute-phase.md b/get-shit-done/workflows/execute-phase.md index edec65c0e..b2bfc5b4b 100644 --- a/get-shit-done/workflows/execute-phase.md +++ b/get-shit-done/workflows/execute-phase.md @@ -3,7 +3,9 @@ Execute all plans in a phase using wave-based parallel execution. Orchestrator s -The orchestrator's job is coordination, not execution. Each subagent loads the full execute-plan context itself. Orchestrator discovers plans, analyzes dependencies, groups into waves, spawns agents, handles checkpoints, collects results. +The orchestrator's job is coordination, not execution. Orchestrator discovers plans, groups into waves, spawns `gsd-executor` agents, handles checkpoints, collects results. + +**Subagent:** `gsd-executor` — dedicated plan execution agent with all execution logic baked in. @@ -150,43 +152,34 @@ Execute each wave in sequence. Autonomous plans within a wave run in parallel. - Bad: "Executing terrain generation plan" - Good: "Procedural terrain generator using Perlin noise — creates height maps, biome zones, and collision meshes. Required before vehicle physics can interact with ground." -2. **Spawn all autonomous agents in wave simultaneously:** +2. **Spawn all agents in wave simultaneously using gsd-executor:** - Use Task tool with multiple parallel calls. Each agent gets prompt from subagent-task-prompt template: + Use Task tool with multiple parallel calls. Each agent is `gsd-executor` with minimal prompt: ``` - - Execute plan {plan_number} of phase {phase_number}-{phase_name}. + Task( + prompt="Execute plan at {plan_path} - Commit each task atomically. Create SUMMARY.md. Update STATE.md. - - - - @~/.claude/get-shit-done/workflows/execute-plan.md - @~/.claude/get-shit-done/templates/summary.md - @~/.claude/get-shit-done/references/checkpoints.md - @~/.claude/get-shit-done/references/tdd.md - - - - Plan: @{plan_path} - Project state: @.planning/STATE.md - Config: @.planning/config.json (if exists) - - - - - [ ] All tasks executed - - [ ] Each task committed individually - - [ ] SUMMARY.md created in plan directory - - [ ] STATE.md updated with position and decisions - +Plan: @{plan_path} +Project state: @.planning/STATE.md +Config: @.planning/config.json (if exists)", + subagent_type="gsd-executor", + description="Execute {phase}-{plan}" + ) ``` -2. **Wait for all agents in wave to complete:** + The `gsd-executor` subagent has all execution logic baked in: + - Deviation rules + - Checkpoint protocols + - Commit formatting + - Summary creation + - State updates + + No template filling needed. Just pass the plan path. Task tool blocks until each agent finishes. All parallel agents return together. -3. **Report completion and what was built:** +4. **Report completion and what was built:** For each completed agent: - Verify SUMMARY.md exists at expected path @@ -215,7 +208,7 @@ Execute each wave in sequence. Autonomous plans within a wave run in parallel. - Bad: "Wave 2 complete. Proceeding to Wave 3." - Good: "Terrain system complete — 3 biome types, height-based texturing, physics collision meshes. Vehicle physics (Wave 3) can now reference ground surfaces." -4. **Handle failures:** +5. **Handle failures:** If any agent in wave fails: - Report which plan failed and why @@ -223,11 +216,11 @@ Execute each wave in sequence. Autonomous plans within a wave run in parallel. - If continue: proceed to next wave (dependent plans may also fail) - If stop: exit with partial completion report -5. **Execute checkpoint plans between waves:** +6. **Execute checkpoint plans between waves:** See `` for details. -6. **Proceed to next wave** +7. **Proceed to next wave** @@ -238,15 +231,22 @@ Plans with `autonomous: false` require user interaction. **Execution flow for checkpoint plans:** -1. **Spawn agent for checkpoint plan:** +1. **Spawn gsd-executor for checkpoint plan:** ``` - Task(prompt="{subagent-task-prompt}", subagent_type="general-purpose") + Task( + prompt="Execute plan at {plan_path} + +Plan: @{plan_path} +Project state: @.planning/STATE.md", + subagent_type="gsd-executor", + description="Execute {phase}-{plan}" + ) ``` 2. **Agent runs until checkpoint:** - Executes auto tasks normally - - Reaches checkpoint task (e.g., `type="checkpoint:human-verify"`) or auth gate - - Agent returns with structured checkpoint (see checkpoint-return.md template) + - Reaches checkpoint task or auth gate + - Agent returns with structured checkpoint (format baked into gsd-executor) 3. **Agent return includes (structured format):** - Completed Tasks table with commit hashes and files @@ -275,20 +275,29 @@ Plans with `autonomous: false` require user interaction. 6. **Spawn continuation agent (NOT resume):** - Use the continuation-prompt.md template: + Spawn fresh `gsd-executor` with continuation context: ``` Task( - prompt=filled_continuation_template, - subagent_type="general-purpose" + prompt="Continue executing plan at {plan_path} + + +{completed_tasks_table from checkpoint return} + + + +Resume from: Task {N} - {task_name} +User response: {user_response} +{resume_instructions based on checkpoint type} + + +Plan: @{plan_path} +Project state: @.planning/STATE.md", + subagent_type="gsd-executor", + description="Continue {phase}-{plan}" ) ``` - Fill template with: - - `{completed_tasks_table}`: From agent's checkpoint return - - `{resume_task_number}`: Current task from checkpoint - - `{resume_task_name}`: Current task name from checkpoint - - `{user_response}`: What user provided - - `{resume_instructions}`: Based on checkpoint type (see continuation-prompt.md) + The `gsd-executor` has continuation handling baked in — it will verify previous commits and resume correctly. 7. **Continuation agent executes:** - Verifies previous commits exist @@ -341,6 +350,154 @@ After all waves complete, aggregate results: ``` + +**Verify the phase GOAL was achieved, not just that tasks completed.** + +This step catches the common failure: tasks done but goal not met (stubs, placeholders, unwired code). + +**1. Spawn verification subagent:** + +Use the subagent-verify-prompt template: + +``` +Task( + prompt: filled_subagent_verify_prompt, + subagent_type: "general-purpose", + description: "Verify phase {X} goal achievement" +) +``` + +Template variables: +- `{phase_number}`: Current phase +- `{phase_name}`: From ROADMAP.md +- `{phase_goal_from_roadmap}`: Phase description +- `{phase_dir}`: Filesystem directory +- `{must_haves_yaml}`: From PLAN.md frontmatter (or "derive from goal") + +**2. Verification subagent runs:** + +The subagent loads `workflows/verify-phase.md` and: +- Establishes must-haves (from frontmatter or derived) +- Verifies observable truths against codebase +- Checks artifacts exist and are substantive (not stubs) +- Traces key links (wiring between components) +- Scans for anti-patterns +- Creates VERIFICATION.md report +- Returns status to orchestrator + +**3. Handle verification result:** + +**If status = "passed":** +``` +## ✓ Phase Verification Passed + +All {N} must-haves verified: +- {truth 1} ✓ +- {truth 2} ✓ +- {truth 3} ✓ + +Phase goal achieved. Proceeding to update roadmap. +``` + +Continue to update_roadmap step. + +**If status = "gaps_found":** +``` +## ⚠️ Phase Verification Found Gaps + +{M} of {N} must-haves incomplete: + +| Must-Have | Status | Issue | +|-----------|--------|-------| +| {truth 1} | ✓ VERIFIED | - | +| {truth 2} | ✗ FAILED | API route returns placeholder | +| {truth 3} | ✗ FAILED | Component not wired to API | + +### Recommended Fix Plans + +1. **{phase}-{next}-PLAN.md**: {description} +2. **{phase}-{next+1}-PLAN.md**: {description} + +Creating fix plans and executing... +``` + +Then: +1. Generate fix PLAN.md files from recommendations +2. Execute fix plans (loop back to execute_waves) +3. Re-verify (loop back to verify_phase_goal) +4. Repeat until all must-haves pass + +**If status = "human_needed":** +``` +## 👤 Human Verification Required + +Automated checks passed. These items need manual testing: + +### 1. {Test Name} +**Test:** {what to do} +**Expected:** {what should happen} + +### 2. {Test Name} +**Test:** {what to do} +**Expected:** {what should happen} + +After testing, type "verified" or describe issues found. +``` + +Wait for user response: +- "verified" / "pass" / "ok" → Continue to update_roadmap +- Description of issues → Generate fix plans, execute, re-verify + +**4. Fix plan generation:** + +When gaps are found, generate fix plans: + +```bash +# Create fix plan from verification recommendations +NEXT_PLAN_NUM=$(ls "$PHASE_DIR"/*-PLAN.md | wc -l) +NEXT_PLAN_NUM=$((NEXT_PLAN_NUM + 1)) +``` + +Fix plans: +- Use standard PLAN.md template +- Include must_haves (same as original, for re-verification) +- Tasks target specific gaps (not entire feature) +- Wave 99 (runs after all original plans) + +**5. Re-verification loop:** + +After fix plans execute: +1. Spawn verification subagent again +2. Check same must-haves +3. If still gaps → more fix plans +4. If passed → continue + +Limit: 3 fix cycles. If still failing after 3 rounds, present to user: +``` +## ⚠️ Verification Still Failing After 3 Fix Attempts + +Remaining gaps: +- {gap 1} +- {gap 2} + +Options: +1. Continue anyway (manual fixes later) +2. Stop and investigate +``` + +**Why this matters:** + +Without verification: +- Phase 3 "complete" but chat doesn't work +- Phase 4 builds on broken foundation +- Phase 8: "nothing works, start over" + +With verification: +- Phase 3 verified before moving on +- Gaps caught and fixed immediately +- Each phase delivers real value + + Update ROADMAP.md to reflect phase completion: @@ -388,20 +545,21 @@ All {N} phases executed. Orchestrator context usage: ~10-15% - Read plan frontmatter (small) -- Analyze dependencies (logic, no heavy reads) -- Fill template strings -- Spawn Task calls +- Group by wave (logic, no heavy reads) +- Spawn Task calls with minimal prompts - Collect results -Each subagent: Fresh 200k context -- Loads full execute-plan workflow -- Loads templates, references +Each `gsd-executor` subagent: Fresh 200k context +- Execution logic baked into subagent prompt (cached) +- Only plan-specific context varies - Executes plan with full capacity - Creates SUMMARY, commits +**Prompt caching benefit:** The `gsd-executor` subagent prompt is stable across all invocations. Only the plan path varies. 90% cost reduction on cached portion. + **No polling.** Task tool blocks until completion. No TaskOutput loops. -**No context bleed.** Orchestrator never reads workflow internals. Just paths and results. +**No context bleed.** Orchestrator never reads execution internals. Just paths and results.