* test(#2444): failing-first regression for checkpoint:* plan-structure validation Add acceptance-criteria tests covering the three canonical checkpoint task types (human-verify, decision, human-action) plus an unknown-subtype forward-compat case. Each canonical type must pass verify plan-structure when it carries its type-specific required fields (per gsd-core/references/checkpoints.md), and must be flagged when those fields are missing. Non-checkpoint tasks keep the existing <action>/<verify>/<done>/<files> requirements unchanged (AC3 regression guards). The existing 'errors when checkpoint task but autonomous is true' fixture is updated to use the canonical checkpoint:human-verify triple (<what-built>/<how-to-verify>/<resume-signal>) so it does not collide with the new per-type validator; the assertion (autonomous is not false) is unchanged. * fix(#2444): branch plan-structure validation on task type=checkpoint:* cmdVerifyPlanStructure unconditionally required <action>/<verify>/<done>/ <files> on every task, so every checkpoint:* task — which uses the checkpoint convention's type-specific fields instead — was reported as a structural error. Checkpoint-heavy phases produced walls of false findings. The fix introduces two pure helpers in verify.cts: - extractPlanTaskInfos(content): single ReDoS-safe pass over <task ...>...</task> blocks that captures BOTH the opening-tag attribute string (so the type= selector is not lost, as it is with extractTaggedBlocks) and the body, returning a typed PlanTaskInfo. - validatePlanTaskStructure(task): branches on the task's type. checkpoint:human-verify requires <what-built>/<how-to-verify>/ <resume-signal> (the canonical triple). checkpoint:decision requires <decision>/<options>/<resume-signal>. checkpoint:human-action requires <action>/<instructions>/ <verification>/<resume-signal>. Unknown checkpoint:* subtypes require only the universal <resume-signal> (forward-compat). All other types keep the historical <action>/<verify>/<done>/<files> requirements unchanged. Canonical reference: gsd-core/references/checkpoints.md. Per-type field sets validated against the documented templates in agents/gsd-planner.md and gsd-core/templates/phase-prompt.md. * fix(#2444): re-resolve body-parser to 2.3.0 in lockfile (GHSA-v422-hmwv-36x6) GHSA-v422-hmwv-36x6 (body-parser DoS via invalid limit value, low severity, published 2026-07-20T23:23:26Z) made tests/npm-integrity-gate.test.cjs (#3588: root workspace production tree has no advisories) fail any subsequent npm audit --omit=dev. The advisory affects body-parser >=2.0.0 <2.3.0 pulled transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol/sdk -> express -> body-parser@2.2.2. express@5.2.1 already declares body-parser as ^2.2.1, so 2.3.0 is a valid re-resolution within express's own compatibility range — no override needed. Regenerated the lockfile via 'npm audit fix --omit=dev' which re-resolves transitive deps within their declared ranges; package.json is unchanged. Verified: npm audit --omit=dev reports 0/0/0/0/0 advisories; body-parser now reads as 2.3.0 in 'npm ls body-parser --omit=dev'. * test(#2444): close review gap-closure tests + harden type-attr charset Orthogonal review (code-review + security-review subagents) returned APPROVE on Standards and Spec. Per the playbook's zero-tolerance policy, address every Low finding: Spec gap-closures: - AC3 verbatim: add explicit <done> and <files> regression tests for non-checkpoint tasks (pre-existing tests only covered <action> and <verify>). - AC2: add checkpoint:decision missing <decision>, checkpoint:human-action missing <action>, checkpoint:human-action missing <verification> cases (the implementation enforces all of these; only one missing-field case per type was previously tested). - Remove the duplicate 'returns error for nonexistent file' test that leaked into the new describe block from the insertion edit. Security hardening (Low-sev, defense-in-depth): - Tighten the task type= attribute extractor in src/verify.cts from [^"'>\s]+ to [\w:-]+ so a hostile type= attribute cannot carry markup fragments (e.g. type=evil<fragment) into the verifier's typed JSON output. All legitimate type values (auto, tracer, manual, checkpoint:human-verify, checkpoint:decision, checkpoint:human-action, checkpoint:tdd-review) match the tighter charset. - Add adversarial regression test asserting type=evil<fragment surfaces as 'evil' (capture stops at '<'), with no markup chars (< > ( ) &) in the surfaced type field. * docs(changeset): add Fixed fragments for #2444 PR Two fragments: - sturdy-jays-tumble.md: the verify plan-structure checkpoint fix - witty-badgers-hum.md: the body-parser 2.3.0 re-resolution PR number backfilled to 0 placeholder per CLAUDE.md 'PR Number Handling'; will backfill to the real PR number immediately after gh pr create returns. * docs(changeset): backfill PR number to 2473 Per CLAUDE.md 'PR Number Handling': backfill the placeholder pr:0 with the real PR number returned by gh pr create.
GSD Core
Git. Ship. Done.
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
What is GSD Core
GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.
How it works
Each milestone repeats the same five-step loop, one phase at a time:
- Discuss — capture implementation decisions before anything is planned
- Plan — research, decompose, and verify the plan fits a fresh context window
- Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
- Verify — walk through what was built; diagnose and fix before declaring done
- Ship — create the PR, archive the phase, repeat for the next one
Quickstart
npx @opengsd/gsd-core@latest
The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.
On another runtime or without Node.js? See Install on your runtime.
Once installed, start a new project or onboard an existing repo:
/gsd-new-project # greenfield project
/gsd-onboard # existing codebase
New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.
Documentation
What's new in 1.7.0 → docs/whats-new-1.7.0.md
Tutorials — learning by doing:
How-to guides — task-focused recipes:
Reference — authoritative facts:
Explanation — concepts and design decisions:
Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.
Why it works
Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.
Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD Core makes it reliable.