feat: add subagent resume capability

Add ability to resume interrupted subagent executions using Task tool's
resume parameter. Tracks agent IDs during execution and enables seamless
continuation after session timeout, rate limits, or crashes.

- New /gsd:resume-task command for resuming interrupted agents
- Agent ID tracking in execute-phase workflow
- Dual storage: current-agent-id.txt (fast) + agent-history.json (audit)
- Auto-detection of interrupted agents in resume-project workflow
- File conflict warnings before resume

Closes #37

Co-Authored-By: davesienkowski <davesienkowski@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
Lex Christopherson
2026-01-10 14:19:01 -06:00
parent 992d034df2
commit 6ba850712f
5 changed files with 647 additions and 5 deletions

View File

@@ -0,0 +1,87 @@
---
name: gsd:resume-task
description: Resume an interrupted subagent execution
argument-hint: "[agent-id]"
allowed-tools:
- Read
- Write
- Edit
- Bash
- Task
- AskUserQuestion
---
<objective>
Resume an interrupted subagent execution using the Task tool's resume parameter.
When a session ends mid-execution, subagents may be left in an incomplete state. This command allows users to continue that work without starting over.
Uses the agent ID tracking infrastructure from execute-phase to identify and resume agents.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/resume-task.md
</execution_context>
<context>
Agent ID: $ARGUMENTS (optional - defaults to most recent)
**Load project state:**
@.planning/STATE.md
**Load agent tracking:**
@.planning/current-agent-id.txt
@.planning/agent-history.json
</context>
<process>
1. Check .planning/ directory exists (error if not)
2. Parse agent ID from arguments or current-agent-id.txt
3. Validate agent exists in history and is resumable
4. Check for file conflicts since spawn
5. Follow resume-task.md workflow:
- Update agent status to "interrupted"
- Resume via Task tool resume parameter
- Update history on completion
- Clear current-agent-id.txt
</process>
<usage>
**Resume most recent interrupted agent:**
```
/gsd:resume-task
```
**Resume specific agent by ID:**
```
/gsd:resume-task agent_01HXYZ123
```
**Find available agents to resume:**
Check `.planning/agent-history.json` for entries with status "spawned" or "interrupted".
</usage>
<error_handling>
**No agent to resume:**
- current-agent-id.txt empty or missing
- Solution: Run /gsd:progress to check project status
**Agent already completed:**
- Agent finished successfully, nothing to resume
- Solution: Continue with next plan
**Agent not found:**
- Provided ID not in history
- Solution: Check agent-history.json for valid IDs
**Resume failed:**
- Agent context expired or invalidated
- Solution: Start fresh with /gsd:execute-plan
</error_handling>
<success_criteria>
- [ ] Agent resumed via Task tool resume parameter
- [ ] Agent-history.json updated with completion
- [ ] current-agent-id.txt cleared
- [ ] User informed of result
</success_criteria>

View File

@@ -0,0 +1,161 @@
# Agent History Template
Template for `.planning/agent-history.json` - tracks subagent spawns during plan execution for resume capability.
---
## File Template
```json
{
"version": "1.0",
"max_entries": 50,
"entries": []
}
```
## Entry Schema
Each entry tracks a subagent spawn or status change:
```json
{
"agent_id": "agent_01HXXXX...",
"task_description": "Execute tasks 1-3 from plan 02-01",
"phase": "02",
"plan": "01",
"segment": 1,
"timestamp": "2026-01-15T14:22:10Z",
"status": "spawned",
"completion_timestamp": null
}
```
### Field Definitions
| Field | Type | Description |
|-------|------|-------------|
| agent_id | string | Unique ID returned by Task tool |
| task_description | string | Brief description of what agent is executing |
| phase | string | Phase number (e.g., "02", "02.1") |
| plan | string | Plan number within phase |
| segment | number | Segment number (1-based) for segmented plans, null for Pattern A |
| timestamp | string | ISO 8601 timestamp when agent was spawned |
| status | string | Current status: spawned, completed, interrupted, resumed |
| completion_timestamp | string/null | ISO 8601 timestamp when completed, null if pending |
### Status Lifecycle
```
spawned ──────────────────────────> completed
│ ^
│ │
└──> interrupted ──> resumed ───────┘
```
- **spawned**: Agent created via Task tool, execution in progress
- **completed**: Agent finished successfully, results received
- **interrupted**: Session ended before agent completed (detected on resume)
- **resumed**: Previously interrupted agent resumed via resume parameter
## Usage
### When to Create File
Create `.planning/agent-history.json` from this template when:
- First subagent spawn in execute-phase workflow
- File doesn't exist yet
### When to Add Entry
Add new entry immediately after Task tool returns with agent_id:
```
1. Task tool spawns subagent
2. Response includes agent_id
3. Write agent_id to .planning/current-agent-id.txt
4. Append entry to agent-history.json with status "spawned"
```
### When to Update Entry
Update existing entry when:
**On successful completion:**
```json
{
"status": "completed",
"completion_timestamp": "2026-01-15T14:45:33Z"
}
```
**On resume detection (interrupted agent found):**
```json
{
"status": "interrupted"
}
```
Then add new entry with resumed status:
```json
{
"agent_id": "agent_01HXXXX...",
"status": "resumed",
"timestamp": "2026-01-15T15:00:00Z"
}
```
### Entry Retention
- Keep maximum 50 entries (configurable via max_entries)
- On exceeding limit, remove oldest completed entries first
- Never remove entries with status "spawned" (may need resume)
- Prune during init_agent_tracking step
## Example File
```json
{
"version": "1.0",
"max_entries": 50,
"entries": [
{
"agent_id": "agent_01HXY123ABC",
"task_description": "Execute full plan 02-01 (autonomous)",
"phase": "02",
"plan": "01",
"segment": null,
"timestamp": "2026-01-15T14:22:10Z",
"status": "completed",
"completion_timestamp": "2026-01-15T14:45:33Z"
},
{
"agent_id": "agent_01HXY456DEF",
"task_description": "Execute tasks 1-3 from plan 02-02",
"phase": "02",
"plan": "02",
"segment": 1,
"timestamp": "2026-01-15T15:00:00Z",
"status": "spawned",
"completion_timestamp": null
}
]
}
```
## Related Files
- `.planning/current-agent-id.txt`: Single line with currently active agent ID (for quick resume lookup)
- `.planning/STATE.md`: Project state including session continuity info
---
## Template Notes
**When to create:** First subagent spawn during execute-phase workflow.
**Location:** `.planning/agent-history.json`
**Companion file:** `.planning/current-agent-id.txt` (single agent ID, overwritten on each spawn)
**Purpose:** Enable resume capability for interrupted subagent executions via Task tool's resume parameter.

View File

@@ -204,15 +204,48 @@ No segmentation benefit - execute entirely in main
**For fully autonomous plans:**
```
Use Task tool with subagent_type="general-purpose":
1. Run init_agent_tracking step first (see step below)
Prompt: "Execute plan at .planning/phases/{phase}-{plan}-PLAN.md
2. Use Task tool with subagent_type="general-purpose":
This is an autonomous plan (no checkpoints). Execute all tasks, create SUMMARY.md in phase directory, commit with message following plan's commit guidance.
Prompt: "Execute plan at .planning/phases/{phase}-{plan}-PLAN.md
Follow all deviation rules and authentication gate protocols from the plan.
This is an autonomous plan (no checkpoints). Execute all tasks, create SUMMARY.md in phase directory, commit with message following plan's commit guidance.
When complete, report: plan name, tasks completed, SUMMARY path, commit hash."
Follow all deviation rules and authentication gate protocols from the plan.
When complete, report: plan name, tasks completed, SUMMARY path, commit hash."
3. After Task tool returns with agent_id:
a. Write agent_id to current-agent-id.txt:
echo "[agent_id]" > .planning/current-agent-id.txt
b. Append spawn entry to agent-history.json:
{
"agent_id": "[agent_id from Task response]",
"task_description": "Execute full plan {phase}-{plan} (autonomous)",
"phase": "{phase}",
"plan": "{plan}",
"segment": null,
"timestamp": "[ISO timestamp]",
"status": "spawned",
"completion_timestamp": null
}
4. Wait for subagent to complete
5. After subagent completes successfully:
a. Update agent-history.json entry:
- Find entry with matching agent_id
- Set status: "completed"
- Set completion_timestamp: "[ISO timestamp]"
b. Clear current-agent-id.txt:
rm .planning/current-agent-id.txt
6. Report completion to user
```
**For segmented plans (has verify-only checkpoints):**
@@ -247,6 +280,54 @@ Quality maintained through small scope (2-3 tasks per plan)
See step name="segment_execution" for detailed segment execution loop.
</step>
<step name="init_agent_tracking">
**Initialize agent tracking for subagent resume capability.**
Before spawning any subagents, set up tracking infrastructure:
**1. Create/verify tracking files:**
```bash
# Create agent history file if doesn't exist
if [ ! -f .planning/agent-history.json ]; then
echo '{"version":"1.0","max_entries":50,"entries":[]}' > .planning/agent-history.json
fi
# Clear any stale current-agent-id (from interrupted sessions)
# Will be populated when subagent spawns
rm -f .planning/current-agent-id.txt
```
**2. Check for interrupted agents (resume detection):**
```bash
# Check if current-agent-id.txt exists from previous interrupted session
if [ -f .planning/current-agent-id.txt ]; then
INTERRUPTED_ID=$(cat .planning/current-agent-id.txt)
echo "Found interrupted agent: $INTERRUPTED_ID"
fi
```
**If interrupted agent found:**
- The agent ID file exists from a previous session that didn't complete
- This agent can potentially be resumed using Task tool's `resume` parameter
- Present to user: "Previous session was interrupted. Resume agent [ID] or start fresh?"
- If resume: Use Task tool with `resume` parameter set to the interrupted ID
- If fresh: Clear the file and proceed normally
**3. Prune old entries (housekeeping):**
If agent-history.json has more than `max_entries`:
- Remove oldest entries with status "completed"
- Never remove entries with status "spawned" (may need resume)
- Keep file under size limit for fast reads
**When to run this step:**
- Pattern A (fully autonomous): Before spawning the single subagent
- Pattern B (segmented): Before the segment execution loop
- Pattern C (main context): Skip - no subagents spawned
</step>
<step name="segment_execution">
**Detailed segment execution loop for segmented plans.**
@@ -299,8 +380,36 @@ For Pattern A (fully autonomous) and Pattern C (decision-dependent), skip this s
- Deviations encountered
- Any issues or blockers"
**After Task tool returns with agent_id:**
1. Write agent_id to current-agent-id.txt:
echo "[agent_id]" > .planning/current-agent-id.txt
2. Append spawn entry to agent-history.json:
{
"agent_id": "[agent_id from Task response]",
"task_description": "Execute tasks [X-Y] from plan {phase}-{plan}",
"phase": "{phase}",
"plan": "{plan}",
"segment": [segment_number],
"timestamp": "[ISO timestamp]",
"status": "spawned",
"completion_timestamp": null
}
Wait for subagent to complete
Capture results (files changed, deviations, etc.)
**After subagent completes successfully:**
1. Update agent-history.json entry:
- Find entry with matching agent_id
- Set status: "completed"
- Set completion_timestamp: "[ISO timestamp]"
2. Clear current-agent-id.txt:
rm .planning/current-agent-id.txt
```
C. If routing = Main context:

View File

@@ -69,6 +69,12 @@ for plan in .planning/phases/*/*-PLAN.md; do
summary="${plan/PLAN/SUMMARY}"
[ ! -f "$summary" ] && echo "Incomplete: $plan"
done 2>/dev/null
# Check for interrupted agents
if [ -f .planning/current-agent-id.txt ] && [ -s .planning/current-agent-id.txt ]; then
AGENT_ID=$(cat .planning/current-agent-id.txt | tr -d '\n')
echo "Interrupted agent: $AGENT_ID"
fi
```
**If .continue-here file exists:**
@@ -81,6 +87,12 @@ done 2>/dev/null
- Execution was started but not completed
- Flag: "Found incomplete plan execution"
**If interrupted agent found:**
- Subagent was spawned but session ended before completion
- Read agent-history.json for task details
- Flag: "Found interrupted agent"
</step>
<step name="present_status">
@@ -103,6 +115,14 @@ Present complete project status to user:
⚠️ Incomplete work detected:
- [.continue-here file or incomplete plan]
[If interrupted agent found:]
⚠️ Interrupted agent detected:
Agent ID: [id]
Task: [task description from agent-history.json]
Interrupted: [timestamp]
Resume with: /gsd:resume-task
[If deferred issues exist:]
📋 [N] deferred issues awaiting attention
@@ -120,6 +140,10 @@ Present complete project status to user:
<step name="determine_next_action">
Based on project state, determine the most logical next action:
**If interrupted agent exists:**
→ Primary: Resume interrupted agent (/gsd:resume-task)
→ Option: Start fresh (abandon agent work)
**If .continue-here file exists:**
→ Primary: Resume from checkpoint
→ Option: Start fresh on current plan
@@ -154,6 +178,8 @@ Present contextual options based on project state:
What would you like to do?
[Primary action based on state - e.g.:]
1. Resume interrupted agent (/gsd:resume-task) [if interrupted agent found]
OR
1. Resume from checkpoint (/gsd:execute-plan .planning/phases/XX-name/.continue-here-02-01.md)
OR
1. Execute next plan (/gsd:execute-plan .planning/phases/XX-name/02-02-PLAN.md)

View File

@@ -0,0 +1,259 @@
<trigger>
Use this workflow when:
- User runs /gsd:resume-task
- Need to continue interrupted subagent execution
- Resuming work after session ended mid-plan
</trigger>
<purpose>
Resume an interrupted subagent execution using the Task tool's resume parameter.
Enables seamless continuation of autonomous work that was interrupted by session timeout, user exit, or crash.
</purpose>
<process>
<step name="parse_arguments">
Parse the optional agent_id argument:
```bash
# Check if argument provided
if [ -n "$ARGUMENTS" ]; then
AGENT_ID="$ARGUMENTS"
echo "Using provided agent ID: $AGENT_ID"
else
# Read from current-agent-id.txt
if [ -f .planning/current-agent-id.txt ] && [ -s .planning/current-agent-id.txt ]; then
AGENT_ID=$(cat .planning/current-agent-id.txt | tr -d '\n')
echo "Using current agent ID: $AGENT_ID"
else
echo "ERROR: No agent to resume"
exit 1
fi
fi
```
**If no agent ID found:**
Present error:
```
No active agent to resume.
There's no interrupted agent recorded. This could mean:
- No subagent was spawned in the current plan
- The last agent completed successfully
- .planning/current-agent-id.txt was cleared
Use /gsd:progress to check project status.
```
**If agent ID found:** Continue to validate_agent step.
</step>
<step name="validate_agent">
Validate the agent exists and is resumable:
```bash
# Check if agent-history.json exists
if [ ! -f .planning/agent-history.json ]; then
echo "ERROR: No agent history found"
exit 1
fi
# Read history and find the agent
cat .planning/agent-history.json
```
**Parse agent-history.json and find entry matching AGENT_ID:**
Check the entry status:
- **If status = "completed":** Error - agent already completed
- **If status = "spawned" or "interrupted":** Valid for resume
- **If not found:** Error - agent ID not in history
**If agent already completed:**
```
Agent already completed.
Agent ID: [id]
Completed: [completion_timestamp]
Task: [task_description]
This agent finished successfully. No resume needed.
```
**If agent not found:**
```
Agent ID not found in history.
ID: [provided_id]
Available agents:
- [list most recent 5 agents with status]
Did you mean one of these?
```
**If valid for resume:** Continue to check_conflicts step.
</step>
<step name="check_conflicts">
Check for file modifications since the agent was spawned.
**Read the agent entry to get context:**
- phase, plan, segment information
- timestamp of spawn
**Check for git changes since spawn:**
```bash
# Get files modified since agent spawn
# Note: This is a best-effort check - relies on git status
git status --short
```
**If modifications detected:**
Use AskUserQuestion to warn user:
```
Files modified since agent was interrupted:
- [list files]
These changes may conflict with the agent's work.
Options:
1. Continue anyway - Agent will resume with current files
2. Abort - Review changes first
Select option:
```
Wait for user response.
**If user selects "Abort":**
```
Resume aborted. Review your changes and run /gsd:resume-task when ready.
```
End workflow.
**If user selects "Continue anyway" or no conflicts found:**
Continue to update_status step.
</step>
<step name="update_status">
Update the agent status to "interrupted" if it was "spawned" (marking the interruption point):
```bash
# Read current history
HISTORY=$(cat .planning/agent-history.json)
```
Update the entry in agent-history.json:
- If status was "spawned", change to "interrupted"
- Add note about resume attempt
This provides audit trail of the interruption before resume.
</step>
<step name="resume_agent">
Resume the agent using Task tool's resume parameter:
```
Resuming agent: [agent_id]
Task: [task_description from history]
Phase: [phase]-[plan]
The agent will continue from where it left off...
```
**Use Task tool with resume parameter:**
```
Task(
description: "Resume interrupted agent",
prompt: "Continue your previous work. You were executing [task_description].",
subagent_type: "general-purpose",
resume: "[AGENT_ID]"
)
```
Wait for agent completion.
**On agent completion:**
- Capture any results returned
- Continue to completion_update step
</step>
<step name="completion_update">
Update tracking files on successful completion:
**1. Update agent-history.json:**
Add new entry marking the resume completion:
```json
{
"agent_id": "[AGENT_ID]",
"task_description": "[original task] (resumed)",
"phase": "[phase]",
"plan": "[plan]",
"segment": [segment or null],
"timestamp": "[now]",
"status": "completed",
"completion_timestamp": "[now]"
}
```
**2. Clear current-agent-id.txt:**
```bash
# Clear the current agent file
> .planning/current-agent-id.txt
```
**3. Present completion message:**
```
Agent resumed and completed successfully.
Agent ID: [id]
Task: [task_description]
Original spawn: [original_timestamp]
Completed: [now]
The agent's work has been incorporated. Check git status for changes.
```
</step>
<step name="handle_errors">
Error handling for resume failures:
**If Task tool returns error on resume:**
```
Failed to resume agent.
Agent ID: [id]
Error: [error message]
Possible causes:
- Agent context may have expired
- Agent may have been invalidated
Options:
1. Start fresh - Execute plan from beginning
2. Check status - Review what was completed
Run /gsd:execute-plan to start fresh if needed.
```
Do NOT clear current-agent-id.txt on error - allow retry.
</step>
</process>
<success_criteria>
Resume is complete when:
- [ ] Agent resumed successfully via Task tool resume parameter
- [ ] Agent-history.json updated with completion status
- [ ] Current-agent-id.txt cleared
- [ ] User informed of completion
</success_criteria>