* feat(#1579): deterministic gsd-tools query eval.score verb Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb (coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain; gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt. Non-breaking — additive only. arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR). * fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline - C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives) - C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws - C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface) - C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access) - inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows) - size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step) - eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage) * fix(#1579): register eval family in alias-drift gates Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs families and to familyArrayKeys in the manifest-coverage test, so the eval family lands under the same drift guard as every sibling family (state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review. check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10. * docs(#1579): use half-open verdict band ranges in CLI-TOOLS overall_score is fractional and thresholds are >=80/>=60/>=40, so a score in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit. * fix(#1579): validate eval.score CLI inputs Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior.
GSD Core
Git. Ship. Done.
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Gemini CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
What is GSD Core
GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Gemini CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.
How it works
Each milestone repeats the same five-step loop, one phase at a time:
- Discuss — capture implementation decisions before anything is planned
- Plan — research, decompose, and verify the plan fits a fresh context window
- Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
- Verify — walk through what was built; diagnose and fix before declaring done
- Ship — create the PR, archive the phase, repeat for the next one
Quickstart
npx @opengsd/gsd-core@latest
The installer prompts for your runtime (Claude Code, OpenCode, Gemini CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.
On another runtime or without Node.js? See Install on your runtime.
Once installed, start your first project:
/gsd-new-project
New here? Follow Your first project for a guided walkthrough from install to first shipped phase.
Documentation
Tutorials — learning by doing:
How-to guides — task-focused recipes:
Reference — authoritative facts:
Explanation — concepts and design decisions:
Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.
Why it works
Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.
Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD Core makes it reliable.