Cristian Uibar c2dc9e2532 Fix W002 false positives on archived phase references in STATE.md body (#3655)
* W002 health check: cross-reference milestones archive for STATE.md phase refs

After /gsd-complete-milestone, phase dirs move into milestones/vX.Y-phases/
and their `#### Phase N:` headings in ROADMAP.md are collapsed inside
<details> blocks. The ROADMAP heading scan misses them, so W002 fired for
every archived phase number mentioned in STATE.md's historical narrative body
("Recent", "Decisions", "Deferred Items") — leaving every project that ever
ran /gsd-complete-milestone permanently degraded with proportional W002 noise.

Union archived milestone phase directories into the validPhases set used by
the W002 check, mirroring the same milestone-archive lookup that W006
already uses (sdk/src/query/validate.ts:723-736).

Closes #3652

* Address review: use shared regex constants and mirror W002 archive fix into CJS path

CodeRabbit (PR #3655 review 1): the ad-hoc /^(\d+[A-Z]?(?:\.\d+)*)/i used to
extract phase tokens from archived phase dirs would skip project-code-prefixed
names like `CK-64-...`. Switch the new SDK block to the shared
PHASE_TOKEN_FROM_DIR_RE / MILESTONE_ARCHIVE_DIR_RE constants (defined at
sdk/src/query/validate.ts:32-33) so prefixed archives are recognised. Added a
companion regression test using `CK-`-prefixed dirs.

Codex review: the shipped CJS health command path (get-shit-done/bin/lib/verify.cjs
cmdValidateHealth, routed by validate-command-router.cjs) only unions
collectDiskPhases (active archive only) plus ROADMAP heading scan — same bug as
the SDK had. Port the all-archive scan into the CJS path via listMilestoneArchiveDirs
+ PHASE_TOKEN_FROM_DIR_RE (already declared at verify.cjs:401-402). Added a CJS
regression test covering the multi-sub-milestone (v1.3a + v1.3b) scenario from
the issue report.

Adds .changeset/lucky-lynx-wave.md.

* Address Gemini findings: shared regex + helper reuse + cross-platform path

P1 #2 — refactor the new W002 archive-scan block to reuse the existing
listMilestoneArchiveDirs helper (sdk/src/query/validate.ts:40) instead of
re-implementing readdir + filter inline. Eliminates duplication and prevents
the two call sites from drifting apart.

P2 #4 — listMilestoneArchiveDirs sorted by `a.slice(a.lastIndexOf('/') + 1)`,
which returns the full path on Windows where path.join produces backslashes.
Switch to path.basename(a) so the numeric version sort works cross-platform.
Brings the SDK helper in line with the CJS sibling at
get-shit-done/bin/lib/verify.cjs:411 which already uses path.basename.

P1 #1 / P2 #5 — the pre-existing W006/W007 archive + active phase scans
(Check 8) used an ad-hoc `^(\d+[A-Z]?(?:\.\d+)*)` regex that silently skipped
project-code-prefixed phase dirs like `CK-64-foo`, so W006 fired for a
correctly-archived phase and W007 fired for a correctly-on-disk phase. Switch
both scans to the shared PHASE_TOKEN_FROM_DIR_RE constant declared at
sdk/src/query/validate.ts:32. The W006 archive loop also now reuses
listMilestoneArchiveDirs for consistency.

P2 #6 — strengthen the CK-prefix regression test to also assert no W006
fires for `#### Phase 64: Prior shipped` (placed inside <details> so the
heading scan picks it up while the on-disk scan does not), pinning the
shared-regex behaviour in the W006 path.

* CI: switch retired /gsd-<cmd> comment syntax to canonical /gsd:<cmd>

The bug-2543 slash-namespace invariant lint scans get-shit-done/bin/lib/**
for /gsd-<cmd> patterns and fails CI when one slips into a comment. Use the
canonical /gsd:complete-milestone form in the new verify.cjs comment (and
mirror the change in the SDK + tests + changeset entry so all docstrings
referencing the milestone-completion command share one spelling).

Also: extract a small forEachArchivedPhaseToken helper in validate.ts (Gemini
P3 finding from review pass 2) so Check 4 (W002) and Check 8 (W006) share the
archive-walking loop instead of inlining it twice.

* Address rev3 review: shared regex parity, numeric sort, drop dead try/catch

Gemini P1 — Check 4's flat phases/ scan still used the ad-hoc
/^(\d+[A-Z]?(?:\.\d+)*)/ regex while the archive scan used the shared
PHASE_TOKEN_FROM_DIR_RE. Project-code-prefixed dirs (e.g. CK-65-current) on
the flat layout would have slipped past validity, so the W002 check could
still mis-classify them. Use PHASE_TOKEN_FROM_DIR_RE here too.

Gemini P3 — `[...validPhases].sort()` ordered tokens alphabetically, producing
error messages like "phases 1, 10, 19, 2, 20" instead of "1, 2, 10, 19, 20".
Switch both SDK and CJS to numeric localeCompare so the displayed list is
human-readable. Mirrored in both paths.

Grok P3 — the CJS Check 4 archive block wrapped listMilestoneArchiveDirs in
an outer try/catch even though the helper already swallows ENOENT/EACCES into
[]. The outer catch was unreachable. Removed; only the per-archive readdir
needs a catch.

Grok P2 / Gemini P1 (CJS Check 8 archive scan) — the assertion that CJS
Check 8 needs the same archive union as Check 4 was repeatedly raised across
review passes. It is incorrect: Check 8 filters ROADMAP.md through
extractCurrentMilestone() before scanning headings, which strips shipped
milestones (collapsed in <details> or not) so archived phase numbers never
reach `roadmapPhases`. Added an inline note documenting this and a positive
regression assertion in the CJS test that W006 does NOT fire for the
archived phases in the multi-sub-milestone fixture. (Skipped a parallel
W007 assertion because the active-archive fallback in
getActiveMilestoneArchiveDir is pre-existing behavior unrelated to #3652.)

* Port forEachArchivedPhaseToken helper to verify.cjs for SDK parity

Gemini rev5 P2 — the CJS Check 4 inlined the archive-walking loop while
the SDK already factored it into forEachArchivedPhaseToken(). Add a
mirror helper in verify.cjs so both seams use the same primitive,
matching the cooperating-sibling pattern documented in
scripts/shared-module-handsync-allowlist.json.

* ci: retrigger to clear unrelated TOCTOU flake

The previous CI run failed at the pre-existing #1925 concurrency test
(state add-blocker concurrent calls) on macos-24 and ubuntu-22 but
passed on ubuntu-24 — and the same test passed on the prior CI run of
this branch (commit 8a246916). The state add-blocker code path is
completely independent of the W002 archive-union changes in this PR.
2026-05-18 00:14:10 -04:00
2026-05-15 18:16:33 -04:00

GET SHIT DONE

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Gemini CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

Solves context rot — the quality degradation that happens as your AI fills its context window.

npm version npm downloads Tests Discord X (Twitter) $GSD Token GitHub stars License


npx get-shit-done-cc@latest

Works on Mac, Windows, and Linux.


GSD Install


"If you know clearly what you want, this WILL build it for you. No bs."

"I've done SpecKit, OpenSpec and Taskmaster — this has produced the best results for me."

"By far the most powerful addition to my Claude Code. Nothing over-engineered. Literally just gets shit done."


Trusted by engineers at Amazon, Google, Shopify, and Webflow.


Important

Returning to GSD?

Run /gsd-map-codebase to re-index your codebase, then /gsd-new-project to rebuild GSD's planning context. Your code is fine — GSD just needs its context rebuilt. See the CHANGELOG for what's new.


Why I Built This

I'm a solo developer. I don't write code — Claude Code does.

Other spec-driven tools exist, but they're all built for 50-person engineering orgs — sprint ceremonies, story points, stakeholder syncs, Jira workflows. I'm not that. I'm a creative person trying to build great things consistently.

So I built GSD. The complexity is in the system, not in your workflow. Behind the scenes: context engineering, XML prompt formatting, subagent orchestration, state management. What you see: a few commands that just work.

The system gives Claude everything it needs to do the work and verify it. I trust the workflow. It just does a good job.

— TÂCHES


How It Works

The loop is six commands. Each one does exactly one thing.

1. Initialize

/gsd-new-project

Questions → research → requirements → roadmap. You approve it, then you're ready to build.

Already have code? Run /gsd-map-codebase first. It analyzes your stack, architecture, and conventions so /gsd-new-project asks the right questions.

2. Discuss

/gsd-discuss-phase 1

Your roadmap has a sentence per phase. That's not enough to build it the way you imagine it. Discuss captures your decisions before anything gets planned: layouts, API shapes, error handling, data structures — whatever gray areas exist for this specific phase.

The output feeds directly into research and planning. Skip it, get reasonable defaults. Use it, get your vision.

3. Plan

/gsd-plan-phase 1

Research → plan → verify, in a loop until the plans pass. Each plan is small enough to execute in a fresh context window.

4. Execute

/gsd-execute-phase 1

Plans run in parallel waves. Each executor gets a fresh 200k-token context. Each task gets its own atomic commit. Walk away, come back to completed work with a clean git history.

Your main context window stays at 30–40%. The work happens in the subagents.

5. Verify

/gsd-verify-work 1

Walk through what was built. Anything broken gets a diagnosed fix plan — ready for immediate re-execution. You don't debug manually; you just run execute again.

6. Repeat → Ship

/gsd-ship 1
/gsd-complete-milestone
/gsd-new-milestone

Loop discuss → plan → execute → verify → ship until the milestone is done. Then archive, tag, and start the next one fresh.


Getting Started

npx get-shit-done-cc@latest

The installer prompts for your runtime (Claude Code, OpenCode, Gemini CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally.

claude --dangerously-skip-permissions

GSD is built for frictionless automation. Skip-permissions is how it's intended to run.

Install only the skills you need with --profile=core (six core-loop skills), --profile=standard (core + phase management), or the default full install. Profiles compose: --profile=core,audit. --minimal is an alias for --profile=core. See docs/USER-GUIDE.md for the full walkthrough, non-interactive install flags for all 15 runtimes, and permissions configuration. See ADR-0011 for the profile model and runtime surface control.

Current release highlights are in docs/RELEASE-v1.42.1.md: package legitimacy checks, safer installer migrations, runtime surface control, custom ship PR sections, reviewer defaults, fallow structural review, and quota-aware execution recovery.


Commands

The main loop:

Command What it does
/gsd-new-project Questions → research → requirements → roadmap
/gsd-discuss-phase [N] Capture implementation decisions before planning
/gsd-plan-phase [N] Research + plan + verify
/gsd-execute-phase <N> Execute plans in parallel waves
/gsd-verify-work [N] Manual acceptance testing
/gsd-ship [N] Create PR from verified phase work
/gsd-progress --next Auto-detect and run the next step
/gsd-complete-milestone Archive milestone and tag release
/gsd-new-milestone Start next version
/gsd:surface Enable/disable skill clusters at runtime without reinstall

For ad-hoc tasks, autonomous mode, codebase analysis, forensics, and the full command surface — see docs/COMMANDS.md.


Why It Works

Three things most AI-coding setups get wrong:

1. Context bloat. As a session grows, quality degrades. GSD keeps your main context clean by doing the heavy work in fresh subagent contexts. Researchers, planners, and executors each start fresh with exactly what they need.

2. No shared memory. GSD maintains structured artifacts that survive session boundaries: PROJECT.md (vision), REQUIREMENTS.md (scope), ROADMAP.md (where you're going), STATE.md (current position and decisions), CONTEXT.md (per-phase implementation decisions). Every new session loads these and knows exactly where things stand.

3. No verification. Code that "runs" isn't code that "works." GSD's verify step walks you through what was built, diagnoses failures with dedicated debug agents, and generates fix plans before you declare a phase done.

See docs/ARCHITECTURE.md for how the multi-agent orchestration and context engineering work in detail.


Configuration

Settings live in .planning/config.json. Configure during /gsd-new-project or update with /gsd-settings.

Key dials:

Setting What it controls
mode interactive (confirm each step) or yolo (auto-approve)
Model profiles quality / balanced / budget — controls which model each agent uses
workflow.research / plan_check / verifier Toggle the quality agents that add tokens and time
parallelization.enabled Run independent plans simultaneously

Optional structural review: set code_quality.fallow.enabled to true to add a fallow pre-pass to /gsd-code-review. GSD writes .planning/phases/<phase>/FALLOW.json and surfaces a Structural Findings (fallow) section in REVIEW.md. Install with npm install -D fallow@^2.70.0 (or system-wide via cargo install fallow; note that the Rust binary's JSON schema must match the documented v2.70+ contract — older versions may produce silent zero-finding output).

Package legitimacy checks are built into the research, planning, and execution path: recommended dependencies get audited, unverified packages require a human checkpoint, and failed installs stop instead of trying similarly named alternatives.

For the full configuration reference — all settings, git branching strategies, per-runtime model overrides, workstream config inheritance, agent skills injection — see docs/CONFIGURATION.md.


Documentation

Doc What's in it
User Guide End-to-end walkthrough, install options, all runtime flags, configuration reference
Commands Every command with flags and examples
Configuration Full config schema, model profiles, git branching
Architecture How the multi-agent orchestration works
CLI Tools gsd-sdk query and programmatic SDK dispatch seams
Features Complete feature index
Changelog What changed in each release

Troubleshooting

Commands not showing up? Restart your runtime after install. GSD installs to ~/.claude/skills/gsd-*/ (Claude Code), ~/.codex/skills/gsd-*/ (Codex), or the equivalent for your runtime.

Codex users — minimum supported CLI version is 0.130.0. Codex CLI 0.130.0 (release notes) removed extra-skill-roots discovery via openai/codex#21485; from that version onward Codex discovers skills from standard roots (including ~/.codex/skills/<name>/SKILL.md). GSD installs there directly. Earlier Codex CLI versions may still discover additional roots, which can surface duplicate gsd-* entries (one from extra-roots discovery, one from ~/.codex/skills/); restart Codex after install and either upgrade or accept the duplicate listing.

Something broken? Re-run the installer — it's idempotent:

npx get-shit-done-cc@latest

Containers or Docker? Set CLAUDE_CONFIG_DIR before installing to avoid tilde-expansion issues:

CLAUDE_CONFIG_DIR=/home/youruser/.claude npx get-shit-done-cc --global

Full troubleshooting and uninstall instructions in docs/USER-GUIDE.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%