* fix(#2605): make dropped lm_studio / llama_cpp reviewer lanes loud, not silent The `lm_studio` and `llama_cpp` reviewer legs in `gsd-core/workflows/review.md` carried the same empty-output defect as the claude/gemini legs (#2494, fixed in #2592) in a worse variant: when the local endpoint was unreachable or returned empty content, they wrote NOTHING to `{run_dir}/gsd-review-<leg>.md`. There was no `[ ! -s … ]` stub at all, so the file never existed, `write_reviews` omitted that reviewer's section, and the outcome was indistinguishable from the reviewer never having been selected. Two diagnostic holes are closed, because an OpenAI-compatible server fails in two ways that leave evidence in different places: - Transport failure (endpoint unreachable): curl writes to stderr and exits non-zero. Both legs used `curl -s`, which suppresses curl's ERROR text as well as the progress meter, and then discarded stderr to `/dev/null` — so there was nothing to capture even in principle. Now `-sS` with stderr to a `.err` sidecar, matching the claude/gemini/codex legs. - Application failure (HTTP 4xx/5xx): curl exits 0 and the error JSON is in the response BODY, so stderr is empty and only the body is evidence. The stub appends the raw response. The `llama_cpp` leg additionally piped curl straight into `jq`, throwing the body away before anything could inspect it; the response is now captured to a variable first, as the `lm_studio` leg already did. The existing `>&2` warning is kept and now points at the stub file. Failing-first verified by extracting both shipped blocks and running them under a real bash against a stubbed curl: pre-fix, all three failure modes produce NO review file and no `.err` sidecar for both legs; post-fix, each produces a diagnosable stub, and a successful review still passes through untouched. Closes #2605 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2605): close the remaining silent-drop paths the review surfaced Four defects found by the orthogonal review of the first commit, all fixed here rather than deferred (they are the same defect class the issue is about, and three of them sit inside the code that commit touched). 1. CodeRabbit leg was still unguarded. `coderabbit review --prompt-only 2>/dev/null > file` with no `[ ! -s … ]` stub — the last leg still shaped like pre-#2494 code. A missing or unauthenticated binary left a zero-byte file that write_reviews rendered as "ran cleanly, nothing to report". It now captures stderr to a `.err` sidecar and emits the same diagnosable stub as every other leg. This is the leg the next issue in the #2494 -> #2592 -> #2605 series would have been about. 2. The budget-skip path dropped the lane just as silently. When `prepare_trimmed_prompt_for_reviewer` fails, `*_SKIP=1` bypasses the entire block — guard included — so no file was written and the only trace was a stderr warning nothing persists. All three local-server legs now write a "review skipped: prompt budget too small" stub on that path. 3. Whitespace-only replies evaded the guard. `[ ! -s … ]` counts BYTES, and command substitution strips trailing newlines but not spaces, so a reply of `" "` was written out and passed as a "successful" but vacuous review — the same indistinguishable-from-success outcome the guard exists to prevent. A `case` glob now normalizes whitespace-only content to empty. 4. `echo "$VAR"` swallowed option-like content. bash's builtin `echo` treats a value of exactly `-n`/`-e`/`-E` as a flag and writes 0 bytes, which would trip the empty guard and DISCARD a genuine reply. Switched to `printf '%s\n'`, the idiom the OpenCode leg in this same file already uses for this reason. Also brings the Ollama leg to parity while it is in hand: it always emitted a stub so it never silently vanished, but it was the least diagnosable of the three local-server legs — bare `-s`, stderr to /dev/null, and the response piped straight into jq so the error body was discarded unread. Verified by extracting all four shipped blocks and running them under a real bash against stubbed CLIs: 22 cases (7 per local-server leg x 3, plus CodeRabbit) all produce the contracted output, and a successful review still passes through untouched on every leg. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2605): regenerate golden install-parity fixtures for the review.md edit The remote test run on the prior commit failed with 19 "golden parity — <runtime>" mismatches. review.md is installed into every runtime's tree, so editing it changes its content hash in all 19 golden fixtures. Regenerated with `npm run gen:golden`; the diff is exactly one line per fixture — the gsd-core/workflows/review.md hash — and nothing else. This is the second ripple of a workflow edit, alongside tests/workflow-size-baseline.json which the first commit already updated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2605): backfill changeset PR number (#2689) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
GSD Core
Git. Ship. Done.
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
What is GSD Core
GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.
How it works
Each milestone repeats the same five-step loop, one phase at a time:
- Discuss — capture implementation decisions before anything is planned
- Plan — research, decompose, and verify the plan fits a fresh context window
- Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
- Verify — walk through what was built; diagnose and fix before declaring done
- Ship — create the PR, archive the phase, repeat for the next one
Quickstart
npx @opengsd/gsd-core@latest
The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.
On another runtime or without Node.js? See Install on your runtime.
Once installed, start a new project or onboard an existing repo:
/gsd-new-project # greenfield project
/gsd-onboard # existing codebase
New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.
Documentation
What's new in 1.7.0 → docs/whats-new-1.7.0.md
Tutorials — learning by doing:
How-to guides — task-focused recipes:
Reference — authoritative facts:
Explanation — concepts and design decisions:
Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.
Why it works
Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.
Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD Core makes it reliable.