* test(#2398): failing-first suite for the CYCLE_SUMMARY consensus gate Binds the gate before it exists, so the suite is RED against next. The load-bearing rows are the two the closed PR #2417 did not have. The B2 regression row asserts a judgment-class lone HIGH counts WITHOUT corroboration when its raiser is unmarked — if anyone re-couples that class to corroboration, more reviewers again produce a weaker gate than one, which is what closed #2417. The parity row asserts every marker literal the gate names is one review-lane-runner actually emits, so the gate cannot key on a signal nothing produces; a mutation row and a seeded fast-check property prove that guard runs its failure branch rather than only reading a correct tree. Also pinned: gate position before Counting rules, the untouched CYCLE_SUMMARY line shape the orchestrator greps, fence balance, the single-reviewer no-op, classification by what a claim asserts rather than by citation presence, the all-marked fail-open, current_actionable staying out of scope, and the leading-marker requirement that stops a review which merely quotes a marker from suppressing its own findings. * feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs With review.reviewer_instances running several reviewer identities off one adapter, any single instance's fabricated HIGH could force a full replan cycle on its own. Across ~9 real cycles on two projects each of four instances fabricated at least once, and each was also the most accurate reviewer in some other cycle, so dropping to fewer reviewers trades away real signal. The gate engages only when 2+ reviewers actually ran, and weighs a lone HIGH by what the claim asserts rather than by whether anyone agreed with it. An existence claim -- a symbol, file, flag, commit or ID exists, is absent, or says something specific -- counts only if source-grounding confirms it or another reviewer raised the same concern. A judgment claim -- a design or correctness property -- counts unless that reviewer's own section opens with an evidence-quality discount marker the review lane already stamps ([reviewed-without-source-citations] #3194, [reviewed-without-repo-access] #2176, or a diff-only lane). That split is what resolves B2, the finding that closed PR #2417. B2 showed the approved wording made more reviewers produce a WEAKER gate than one: condition (a) pointed at the source-grounding pass, which verifies every symbol THE PLAN cites and never takes reviewer claims as input, so a genuine architectural HIGH that one reviewer caught and another missed was neither groundable nor corroborated and stopped gating. Judgment-class findings are therefore exempt from corroboration entirely -- reviewers catch materially different classes of issue, and demanding two of them independently raise the same architectural concern suppresses exactly what a multi-reviewer setup exists to surface. Guards on the gate itself: an all-marked cycle disengages it, so a cycle in which nothing was verified can never be counted as converged; the marker must OPEN a reviewer's section, so a review that merely quotes a marker does not suppress its own findings; a suppressed HIGH stays listed and tagged rather than dropped; current_actionable is untouched; and a single-reviewer run is unchanged. No new command, config key, or dependency -- the gate reads signals that already exist. The CYCLE_SUMMARY line shape the orchestrator greps is unchanged; only the integer it computes moves, and only for 2+ reviewers. Known limit, inherited rather than introduced: SOURCE_CITATION_RE checks citation presence, not resolution, which src/review-lane-runner.cts records as a deliberate #3194 scope boundary. A fabricated but plausible file:line still gates. Scope revised and re-approved on the issue before any code was written. * test(#2398): make marker parity behavioral, and stop overclaiming the gate Review found the parity tests were vacuous: they asserted a marker STRING appeared in review-lane-runner.cjs's source text, never requiring the module or calling the stampers, so they would pass even if stampUngroundedReview were broken or never invoked. They now invoke the real exported functions and assert what those functions PRODUCE — that an uncited review gains a leading marker blockquote, that a review carrying a file:line does not, that a self-reported blind review is stamped, and that stamping is idempotent. Removing the source read also removes an incidental no-source-grep evasion via a parameterized path. Review also found the changeset headline false for the class it matters most in. The discount markers detect 'cited nothing' and 'had no repo access'; they cannot detect 'drew a wrong conclusion from a real citation', so a judgment-class finding invented by an evidence-bearing reviewer still counts alone. That is the deliberate side of the tradeoff jags-faith named when closing #2417 — the alternative is requiring corroboration for design findings, which is B2 — but the changeset claimed lone hallucinations no longer force a cycle, full stop. Corrected there, and stated plainly in docs/COMMANDS.md and the design record. Also dropped the reviewer-instances.md entry from the emitted-drift ack: the growth ratchet's currentSizes() scans only gsd-core/workflows/ and agents/ (tests/helpers/emitted-runtime.cjs:916-929), so references/ is outside it and that entry acknowledged a delta the gate cannot see. * chore(#2398): backfill changeset pr number to 3755 --------- Co-authored-by: sim <sim@local>
GSD Core
Git. Ship. Done.
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
What is GSD Core
GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.
How it works
Each milestone repeats the same five-step loop, one phase at a time:
- Discuss — capture implementation decisions before anything is planned
- Plan — research, decompose, and verify the plan fits a fresh context window
- Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
- Verify — walk through what was built; diagnose and fix before declaring done
- Ship — create the PR, archive the phase, repeat for the next one
Quickstart
npx @opengsd/gsd-core@latest
The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.
On another runtime or without Node.js? See Install on your runtime.
Once installed, start a new project or onboard an existing repo:
/gsd-new-project # greenfield project
/gsd-onboard # existing codebase
New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.
Documentation
What's new in 1.7.0 → docs/whats-new-1.7.0.md
Tutorials — learning by doing:
How-to guides — task-focused recipes:
Reference — authoritative facts:
Explanation — concepts and design decisions:
Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.
Why it works
Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.
Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD Core makes it reliable.