* test(3309): red — workflow.human_verify_mode contract New behavioral test file covers: - workflow.human_verify_mode is a recognized config key (VALID_CONFIG_KEYS) - defaults to 'mid-flight' (preserves current behavior) - config-set / config-get round-trips for both values - persists in config.json as string - planner agent file references the flag with canonical wording, couples end-of-phase mode with the rule that checkpoint:human-verify is not emitted, and documents the <verify><human-check> deferred-item shape - verifier agent file references harvesting <verify><human-check> blocks - references/checkpoints.md documents the cost-control alternative Source-text assertions on agent .md files are exempted via allow-test-rule: source-text-is-the-product — those files ARE the runtime contract loaded by AI runtimes, so asserting their wording is the only way to verify the agents will respect the flag. Fails 10/11 against current source. Will pass after the fix. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(3309): add workflow.human_verify_mode = end-of-phase opt-out Each mid-flight checkpoint:human-verify halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on every respawn) because subagent context is discarded across the pause. A plan with N human-verify checkpoints pays the cold-start cost N+1 times. The reporter (rentanything-nb) measured this at "tens of thousands of tokens" per round-trip and "hundreds of thousands per week." This adds workflow.human_verify_mode (default 'mid-flight') with an 'end-of-phase' value that: - instructs gsd-planner to NOT emit <task type="checkpoint:human-verify"> tasks; verification details go into a <verify><human-check> sub-block on the relevant auto task instead - instructs gsd-verifier (Step 8) to harvest those <verify><human-check> blocks at end-of-phase and merge them into its own human-verification list - the existing human_needed → HUMAN-UAT.md flow in execute-phase.md is the single sink — no new file/writer is created checkpoint:decision and checkpoint:human-action are unaffected — those gate the work itself, not post-hoc verification. Surfaces touched: - bin/lib/config-schema.cjs, bin/lib/config.cjs — register key + default - sdk/src/config.ts, sdk/src/query/config-schema.ts — SDK parity - agents/gsd-planner.md — slim Detection section + reference link - agents/gsd-verifier.md — Step 8 harvest instruction - get-shit-done/references/planner-human-verify-mode.md — full rules, loaded conditionally to keep planner.md under its size budget - get-shit-done/references/checkpoints.md — surface the alternative - docs/CONFIGURATION.md — config table row - docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json — track new reference Tag name <human-check> chosen instead of <human> to avoid the prompt-injection scan pattern that flags <system|assistant|human> tags. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(3309): align changeset pr: to actual PR number The pr: field was authored as 3319 (a guess at the next number) before the PR was opened. Actual PR is #3325. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(3309): flip workflow.human_verify_mode default to end-of-phase Per maintainer direction on PR #3325, end-of-phase is the new project default. Mid-flight checkpoint:human-verify halts cost a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per round-trip — reported at "tens of thousands of tokens" per round-trip, "hundreds of thousands per week" on real projects. The cost-control mode is what new projects should get out of the box. mid-flight remains a one-line opt-back-in via: gsd config-set workflow.human_verify_mode mid-flight Behavior change for existing projects: the new default takes effect when .planning/config.json is rewritten (config-set, fresh project). Existing in-flight PLAN.md files with checkpoint:human-verify tasks continue to work in either mode — the flag only changes what the planner emits next time it runs. Surfaces updated: - bin/lib/config.cjs, sdk/src/config.ts — default flipped - sdk/src/config.ts docstring — describes new default + opt-back-in - agents/gsd-planner.md — Detection section explains new default - references/planner-human-verify-mode.md — reordered modes; added guidance on when to opt back into mid-flight - references/checkpoints.md — surface the default flip and the why - docs/CONFIGURATION.md — table row reflects new default + reason - tests/feat-3309-human-verify-mode.test.cjs — default test asserts end-of-phase - .changeset/fierce-geese-march.md — describes the default flip and the migration semantics Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: address human verify mode review --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
GET SHIT DONE
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Gemini CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
Solves context rot — the quality degradation that happens as your AI fills its context window.
npx get-shit-done-cc@latest
Works on Mac, Windows, and Linux.
"If you know clearly what you want, this WILL build it for you. No bs."
"I've done SpecKit, OpenSpec and Taskmaster — this has produced the best results for me."
"By far the most powerful addition to my Claude Code. Nothing over-engineered. Literally just gets shit done."
Trusted by engineers at Amazon, Google, Shopify, and Webflow.
Important
Returning to GSD?
Run
/gsd-map-codebaseto re-index your codebase, then/gsd-new-projectto rebuild GSD's planning context. Your code is fine — GSD just needs its context rebuilt. See the CHANGELOG for what's new.
Why I Built This
I'm a solo developer. I don't write code — Claude Code does.
Other spec-driven tools exist, but they're all built for 50-person engineering orgs — sprint ceremonies, story points, stakeholder syncs, Jira workflows. I'm not that. I'm a creative person trying to build great things consistently.
So I built GSD. The complexity is in the system, not in your workflow. Behind the scenes: context engineering, XML prompt formatting, subagent orchestration, state management. What you see: a few commands that just work.
The system gives Claude everything it needs to do the work and verify it. I trust the workflow. It just does a good job.
— TÂCHES
How It Works
The loop is six commands. Each one does exactly one thing.
1. Initialize
/gsd-new-project
Questions → research → requirements → roadmap. You approve it, then you're ready to build.
Already have code? Run
/gsd-map-codebasefirst. It analyzes your stack, architecture, and conventions so/gsd-new-projectasks the right questions.
2. Discuss
/gsd-discuss-phase 1
Your roadmap has a sentence per phase. That's not enough to build it the way you imagine it. Discuss captures your decisions before anything gets planned: layouts, API shapes, error handling, data structures — whatever gray areas exist for this specific phase.
The output feeds directly into research and planning. Skip it, get reasonable defaults. Use it, get your vision.
3. Plan
/gsd-plan-phase 1
Research → plan → verify, in a loop until the plans pass. Each plan is small enough to execute in a fresh context window.
4. Execute
/gsd-execute-phase 1
Plans run in parallel waves. Each executor gets a fresh 200k-token context. Each task gets its own atomic commit. Walk away, come back to completed work with a clean git history.
Your main context window stays at 30–40%. The work happens in the subagents.
5. Verify
/gsd-verify-work 1
Walk through what was built. Anything broken gets a diagnosed fix plan — ready for immediate re-execution. You don't debug manually; you just run execute again.
6. Repeat → Ship
/gsd-ship 1
/gsd-complete-milestone
/gsd-new-milestone
Loop discuss → plan → execute → verify → ship until the milestone is done. Then archive, tag, and start the next one fresh.
Getting Started
npx get-shit-done-cc@latest
The installer prompts for your runtime (Claude Code, OpenCode, Gemini CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally.
claude --dangerously-skip-permissions
GSD is built for frictionless automation. Skip-permissions is how it's intended to run.
See docs/USER-GUIDE.md for the full walkthrough, non-interactive install flags for all 15 runtimes, minimal install (--minimal), Docker setup, and permissions configuration.
Commands
The main loop:
| Command | What it does |
|---|---|
/gsd-new-project |
Questions → research → requirements → roadmap |
/gsd-discuss-phase [N] |
Capture implementation decisions before planning |
/gsd-plan-phase [N] |
Research + plan + verify |
/gsd-execute-phase <N> |
Execute plans in parallel waves |
/gsd-verify-work [N] |
Manual acceptance testing |
/gsd-ship [N] |
Create PR from verified phase work |
/gsd-progress --next |
Auto-detect and run the next step |
/gsd-complete-milestone |
Archive milestone and tag release |
/gsd-new-milestone |
Start next version |
For ad-hoc tasks, autonomous mode, codebase analysis, forensics, and the full command surface — see docs/COMMANDS.md.
Why It Works
Three things most AI-coding setups get wrong:
1. Context bloat. As a session grows, quality degrades. GSD keeps your main context clean by doing the heavy work in fresh subagent contexts. Researchers, planners, and executors each start fresh with exactly what they need.
2. No shared memory. GSD maintains structured artifacts that survive session boundaries: PROJECT.md (vision), REQUIREMENTS.md (scope), ROADMAP.md (where you're going), STATE.md (current position and decisions), CONTEXT.md (per-phase implementation decisions). Every new session loads these and knows exactly where things stand.
3. No verification. Code that "runs" isn't code that "works." GSD's verify step walks you through what was built, diagnoses failures with dedicated debug agents, and generates fix plans before you declare a phase done.
See docs/ARCHITECTURE.md for how the multi-agent orchestration and context engineering work in detail.
Configuration
Settings live in .planning/config.json. Configure during /gsd-new-project or update with /gsd-settings.
Key dials:
| Setting | What it controls |
|---|---|
mode |
interactive (confirm each step) or yolo (auto-approve) |
| Model profiles | quality / balanced / budget — controls which model each agent uses |
workflow.research / plan_check / verifier |
Toggle the quality agents that add tokens and time |
parallelization.enabled |
Run independent plans simultaneously |
For the full configuration reference — all settings, git branching strategies, per-runtime model overrides, workstream config inheritance, agent skills injection — see docs/CONFIGURATION.md.
Documentation
| Doc | What's in it |
|---|---|
| User Guide | End-to-end walkthrough, install options, all runtime flags, configuration reference |
| Commands | Every command with flags and examples |
| Configuration | Full config schema, model profiles, git branching |
| Architecture | How the multi-agent orchestration works |
| CLI Tools | gsd-sdk query and programmatic SDK dispatch seams |
| Features | Complete feature index |
| Changelog | What changed in each release |
Troubleshooting
Commands not showing up? Restart your runtime after install. GSD installs to ~/.claude/skills/gsd-*/ (Claude Code), ~/.codex/skills/gsd-*/ (Codex), or the equivalent for your runtime.
Something broken? Re-run the installer — it's idempotent:
npx get-shit-done-cc@latest
Containers or Docker? Set CLAUDE_CONFIG_DIR before installing to avoid tilde-expansion issues:
CLAUDE_CONFIG_DIR=/home/youruser/.claude npx get-shit-done-cc --global
Full troubleshooting and uninstall instructions in docs/USER-GUIDE.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD makes it reliable.