* test(#4112): add failing-first regression coverage for pause-work.md Context Detection Extracts and executes the shipped Context Detection bash block to prove the $(( arithmetic-context misparse (SC1102/SC1106/SC2205) as a hard syntax error, plus boundary/independence coverage for the phase/spike/sketch/ deliberation resolution the fix must preserve. * test(#4112): fix false-green regression test — assert dash/sh failure, not bash -n The prior version asserted bash -n syntax validity and wrapped bash -c execution, neither of which detects this bug: bash/zsh silently fall back to a working command-substitution reading of the ambiguous $(( construct, and bash -n never evaluates arithmetic-context content at parse time. A gsd-test run against the still-broken file passed 41594/41594 with this test in place — a false green. POSIX sh/dash does not implement that fallback and throws a real "Syntax error: Missing '))'", which is what this version asserts against. * fix(#4112): remove leaked tool-call tags corrupting the regression test file A subagent's Write introduced trailing </content>/</invoke> markup at the end of tests/pause-work-context-detection.test.cjs. gsd-test's own tests/portability-rule-disable-ban.test.cjs (which parses every test file) correctly caught this as a parse error. Strips the garbage lines; no behavioral change. * fix(#4112): drop redundant $(( subshell wrapping in pause-work.md Context Detection phase=/spike=/sketch= opened with $((, which POSIX sh/dash parses as arithmetic expansion and rejects (Syntax error: Missing '))') since the enclosed text isn't valid arithmetic. bash/zsh silently retry it as command substitution, which masked the defect there. Dropping the inner grouping parens (matching the deliberation= line already in the same block) makes the construct valid under bash, zsh, and dash alike, with identical fallback-to-empty-string behavior when nothing matches. Once the surrounding syntax parses, ShellCheck can now also analyze the inner ls -lt calls it previously couldn't see past the parse failure, newly surfacing the same pre-existing SC2012 ("use find instead of ls") suggestion already baselined for the file's other ls usage. Baseline updated to reflect the true current count; no ls-vs-find behavior change made, as that is a separate, pre-existing, out-of-scope question. * fix(#4112): verify shell discrimination instead of trusting the binary name Code review finding: falling back from dash to sh could silently produce a non-discriminating test on a host where /bin/sh is bash-compatible (e.g. macOS) — such a shell never throws on the broken construct either, so the test would pass whether the bug were present or not. resolvePosixShell now verifies the candidate actually rejects a known-ambiguous $(( snippet before using it, and skips with an explicit reason when none does. The gsd-test Linux bench (dash as /bin/sh) is unaffected either way. * fix(#4112): route regression test through process-seam, splitLines, cleanup npm run lint:ci flagged 7 violations gsd-test's node:test run doesn't check: local/no-adhoc-markdown-parsing (single fence-spanning regex), local/no-crlf-fragile-split (bare \n split), local/no-unbounded-spawn (x4, hand-rolled execFileSync with no timeout), and local/no-raw-rmsync-in-tests. Rewrites the test to route every subprocess call through tests/helpers/process-seam.cjs's runHook() (bounded by construction), tests/helpers.cjs's cleanup() for temp-dir removal, and a line-by-line fence scan using the text-lines.cts splitLines() seam, matching this repo's established extraction idiom (tests/no-hardcoded-home-gsd-tools.test.cjs). No behavioral change to what is asserted. * docs(#4112): add changeset for the pause-work Context Detection fix * docs(#4112): backfill changeset PR number (pr:0 -> pr:4140) --------- Co-authored-by: sim <sim@local>
12 KiB
<required_reading> Read all files referenced by the invoking prompt's execution_context before starting. </required_reading>
## Context DetectionDetermine what kind of work is being paused and set the handoff destination accordingly:
# Check for active phase
phase=$(ls -lt .planning/phases/*/PLAN.md 2>/dev/null | head -1 | grep -oP 'phases/\K[^/]+' || true)
# Check for active spike
spike=$(ls -lt .planning/spikes/*/SPIKE.md .planning/spikes/*/DESIGN.md .planning/spikes/*/README.md 2>/dev/null | head -1 | grep -oP 'spikes/\K[^/]+' || true)
# Check for active sketch
sketch=$(ls -lt .planning/sketches/*/README.md .planning/sketches/*/index.html 2>/dev/null | head -1 | grep -oP 'sketches/\K[^/]+' || true)
# Check for active deliberation
deliberation=$(ls .planning/deliberations/*.md 2>/dev/null | head -1 || true)
- Phase work: active phase directory → handoff to
.planning/phases/XX-name/.continue-here.md - Spike work: active spike directory or spike-related files (no active phase) → handoff to
.planning/spikes/SPIKE-NNN/.continue-here.md(create directory if needed) - Sketch work: active sketch directory (no active phase/spike) → handoff to
.planning/sketches/.continue-here.md - Deliberation work: active deliberation file (no phase/spike/sketch) → handoff to
.planning/deliberations/.continue-here.md - Research work: research notes exist but no phase/spike/sketch/deliberation → handoff to
.planning/.continue-here.md - Default: no detectable context → handoff to
.planning/.continue-here.md, note the ambiguity in<current_state>
If phase is detected, proceed with phase handoff path. Otherwise use the first matching non-phase path above.
**Collect complete state for handoff:**- Current position: Which phase, which plan, which task
- Work completed: What got done this session
- Work remaining: What's left in current plan/phase
- Decisions made: Key decisions and rationale
- Blockers/issues: Anything stuck
- Human actions pending: Things that need manual intervention (MCP setup, API keys, approvals, manual testing)
- Background processes: Any running servers/watchers that were part of the workflow
- Files modified: What's changed but not committed
- Outstanding async external jobs: any
.planning/async-jobs/*.jsonmanifests for non-terminal jobs — record job id, backend, status, expected artifacts, verification + resume commands, and any watcher/daemon state. Do NOT cancel the external job; it keeps running across the pause. - Blocking constraints: Anti-patterns or methodological failures encountered during this session that a resuming agent MUST be aware of before proceeding. Only include items discovered through actual failure — not warnings or predictions. Assign each constraint a
severity:
blocking— The resuming agent MUST demonstrate understanding before proceeding. The discuss-phase and execute-phase workflows will enforce a mandatory understanding check.advisory— Important context but does not gate resumption.
Ask user for clarifications if needed via conversational questions.
Also inspect SUMMARY.md files for false completions:
# Check for placeholder content in existing summaries
grep -l "To be filled\|placeholder\|TBD" .planning/phases/*/*.md 2>/dev/null || true
Report any summaries with placeholder content as incomplete items.
**Write structured handoff to `.planning/HANDOFF.json`:**_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; _gsd_at() { for _p; do if [ -f "$_p" ]; then GSD_TOOLS="$_p"; return 0; fi; done; return 1; }; if _gsd_at "${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif unset -f gsd_run; _G="$(command -v gsd_run)"; then GSD_TOOLS="$_G"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif _gsd_at "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd_run is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; GSD_IDENTITY_STATUS=unverified; case "$(gsd_run runtime-identity --raw 2>/dev/null || true)" in '{"packageName":"@opengsd/gsd-core"'*'}') GSD_IDENTITY_STATUS=ok;; esac; export GSD_IDENTITY_STATUS; [ "$GSD_IDENTITY_STATUS" = ok ] || echo "WARNING: \"$GSD_TOOLS\" did not prove it is @opengsd/gsd-core - it is either a different package or an @opengsd/gsd-core older than the runtime-identity verb. See docs/how-to/diagnose-a-foreign-gsd-tools.md" >&2; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
timestamp=$(gsd_run query current-timestamp full --raw)
{
"version": "1.0",
"timestamp": "{timestamp}",
"phase": "{phase_number}",
"phase_name": "{phase_name}",
"phase_dir": "{phase_dir}",
"plan": {current_plan_number},
"task": {current_task_number},
"total_tasks": {total_task_count},
"status": "paused",
"completed_tasks": [
{"id": 1, "name": "{task_name}", "status": "done", "commit": "{short_hash}"},
{"id": 2, "name": "{task_name}", "status": "done", "commit": "{short_hash}"},
{"id": 3, "name": "{task_name}", "status": "in_progress", "progress": "{what_done}"}
],
"remaining_tasks": [
{"id": 4, "name": "{task_name}", "status": "not_started"},
{"id": 5, "name": "{task_name}", "status": "not_started"}
],
"blockers": [
{"description": "{blocker}", "type": "technical|human_action|external", "workaround": "{if any}"}
],
"async_jobs": [
{"manifest": ".planning/async-jobs/{job}.json", "job_id": "{id}", "backend": "{backend}", "status": "running", "submit_command": "{cmd}", "submitted_at": "{iso8601}", "expected_artifacts": ["..."], "verification_command": "{cmd}", "resume_command": "{cmd}"}
],
"human_actions_pending": [
{"action": "{what needs to be done}", "context": "{why}", "blocking": true}
],
"decisions": [
{"decision": "{what}", "rationale": "{why}", "phase": "{phase_number}"}
],
"uncommitted_files": [],
"next_action": "{specific first action when resuming}",
"context_notes": "{mental state, approach, what you were thinking}"
}
Any recorded async_jobs entries are the primary resume context on the next session — check them first before treating a PLAN-without-SUMMARY as incomplete work.
---
context: [phase|spike|sketch|deliberation|research|default]
phase: XX-name
task: 3
total_tasks: 7
status: in_progress
last_updated: [timestamp from current-timestamp]
---
# BLOCKING CONSTRAINTS — Read Before Anything Else
> These are not suggestions. Each constraint below was discovered through failure.
> Acknowledge each one explicitly before proceeding.
- [ ] CONSTRAINT: [name] — [what it is] — [structural mitigation required]
**Do not proceed until all boxes are checked.**
_If no constraints have been identified yet, remove this section._
## Critical Anti-Patterns
| Pattern | Description | Severity | Prevention Mechanism |
|---------|-------------|----------|---------------------|
| [pattern name] | [what it is and how it manifested] | blocking | [structural step that prevents recurrence — not acknowledgment] |
| [pattern name] | [what it is and how it manifested] | advisory | [guidance for avoiding it] |
**Severity values:** `blocking` — resuming agent must pass understanding check before proceeding. `advisory` — important context, does not gate resumption.
_Remove rows that do not apply. The discuss-phase and execute-phase workflows parse this table and enforce a mandatory understanding check for any `blocking` rows._
<current_state>
[Where exactly are we? Immediate context]
</current_state>
<completed_work>
Completed Tasks:
- Task 1: [name] - Done
- Task 2: [name] - Done
- Task 3: [name] - In progress, [what's done]
</completed_work>
<remaining_work>
- Task 3: [what's left]
- Task 4: Not started
- Task 5: Not started
</remaining_work>
<decisions_made>
- Decided to use [X] because [reason]
- Chose [approach] over [alternative] because [reason]
</decisions_made>
<blockers>
- [Blocker 1]: [status/workaround]
</blockers>
## Required Reading (in order)
<!-- List documents the resuming agent must read before acting -->
1. [document] — [why it matters]
1. `.planning/METHODOLOGY.md` (if it exists) — project analytical lenses; apply before any assumption analysis
## Critical Anti-Patterns (do NOT repeat these)
<!-- Mistakes discovered this session that must be structurally avoided -->
- [ANTI-PATTERN]: [what it is] → [structural mitigation]
## Infrastructure State
<!-- Running services, external state, environment specifics -->
- [service/env]: [current state]
## Pre-Execution Critique Required
<!-- Fill in ONLY if pausing between design and execution (e.g. spike design done, not yet run) -->
- Design artifact: [path]
- Critique focus: [key questions the critic should probe]
- Gate: Do NOT begin execution until critique is complete and design is revised
<context>
[Mental state, what were you thinking, the plan]
</context>
<next_action>
Start with: [specific first action when resuming]
</next_action>
Be specific enough for a fresh Claude to understand immediately.
Use current-timestamp for last_updated field. You can use init todos (which provides timestamps) or call directly:
timestamp=$(gsd_run query current-timestamp full --raw)
Current state:
- Context: [phase|spike|deliberation|research]
- Location: [XX-name or SPIKE-NNN]
- Task: [X] of [Y]
- Status: [in_progress/blocked]
- Blockers: [count] ({human_actions_pending count} need human action)
- Committed as WIP
To resume: /gsd:resume-work
</step>
</process>
<success_criteria>
- [ ] Context detected (phase/spike/deliberation/research/default)
- [ ] .continue-here.md created at correct path for detected context
- [ ] Required Reading, Anti-Patterns, and Infrastructure State sections filled
- [ ] Pre-Execution Critique section filled if pausing between design and execution
- [ ] Committed as WIP
- [ ] User knows location and how to resume
</success_criteria>