Tom Boucher 2c44241a0b fix(#2605): make dropped local-server reviewer lanes loud instead of silent (#2689)
* fix(#2605): make dropped lm_studio / llama_cpp reviewer lanes loud, not silent

The `lm_studio` and `llama_cpp` reviewer legs in `gsd-core/workflows/review.md`
carried the same empty-output defect as the claude/gemini legs (#2494, fixed in
#2592) in a worse variant: when the local endpoint was unreachable or returned
empty content, they wrote NOTHING to `{run_dir}/gsd-review-<leg>.md`. There was
no `[ ! -s … ]` stub at all, so the file never existed, `write_reviews` omitted
that reviewer's section, and the outcome was indistinguishable from the reviewer
never having been selected.

Two diagnostic holes are closed, because an OpenAI-compatible server fails in two
ways that leave evidence in different places:

- Transport failure (endpoint unreachable): curl writes to stderr and exits
  non-zero. Both legs used `curl -s`, which suppresses curl's ERROR text as well
  as the progress meter, and then discarded stderr to `/dev/null` — so there was
  nothing to capture even in principle. Now `-sS` with stderr to a `.err`
  sidecar, matching the claude/gemini/codex legs.
- Application failure (HTTP 4xx/5xx): curl exits 0 and the error JSON is in the
  response BODY, so stderr is empty and only the body is evidence. The stub
  appends the raw response. The `llama_cpp` leg additionally piped curl straight
  into `jq`, throwing the body away before anything could inspect it; the
  response is now captured to a variable first, as the `lm_studio` leg already
  did.

The existing `>&2` warning is kept and now points at the stub file.

Failing-first verified by extracting both shipped blocks and running them under a
real bash against a stubbed curl: pre-fix, all three failure modes produce NO
review file and no `.err` sidecar for both legs; post-fix, each produces a
diagnosable stub, and a successful review still passes through untouched.

Closes #2605

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* fix(#2605): close the remaining silent-drop paths the review surfaced

Four defects found by the orthogonal review of the first commit, all fixed here
rather than deferred (they are the same defect class the issue is about, and
three of them sit inside the code that commit touched).

1. CodeRabbit leg was still unguarded. `coderabbit review --prompt-only
   2>/dev/null > file` with no `[ ! -s … ]` stub — the last leg still shaped like
   pre-#2494 code. A missing or unauthenticated binary left a zero-byte file that
   write_reviews rendered as "ran cleanly, nothing to report". It now captures
   stderr to a `.err` sidecar and emits the same diagnosable stub as every other
   leg. This is the leg the next issue in the #2494 -> #2592 -> #2605 series
   would have been about.

2. The budget-skip path dropped the lane just as silently. When
   `prepare_trimmed_prompt_for_reviewer` fails, `*_SKIP=1` bypasses the entire
   block — guard included — so no file was written and the only trace was a
   stderr warning nothing persists. All three local-server legs now write a
   "review skipped: prompt budget too small" stub on that path.

3. Whitespace-only replies evaded the guard. `[ ! -s … ]` counts BYTES, and
   command substitution strips trailing newlines but not spaces, so a reply of
   `"   "` was written out and passed as a "successful" but vacuous review — the
   same indistinguishable-from-success outcome the guard exists to prevent. A
   `case` glob now normalizes whitespace-only content to empty.

4. `echo "$VAR"` swallowed option-like content. bash's builtin `echo` treats a
   value of exactly `-n`/`-e`/`-E` as a flag and writes 0 bytes, which would trip
   the empty guard and DISCARD a genuine reply. Switched to `printf '%s\n'`, the
   idiom the OpenCode leg in this same file already uses for this reason.

Also brings the Ollama leg to parity while it is in hand: it always emitted a
stub so it never silently vanished, but it was the least diagnosable of the three
local-server legs — bare `-s`, stderr to /dev/null, and the response piped
straight into jq so the error body was discarded unread.

Verified by extracting all four shipped blocks and running them under a real bash
against stubbed CLIs: 22 cases (7 per local-server leg x 3, plus CodeRabbit) all
produce the contracted output, and a successful review still passes through
untouched on every leg.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* test(#2605): regenerate golden install-parity fixtures for the review.md edit

The remote test run on the prior commit failed with 19 "golden parity —
<runtime>" mismatches. review.md is installed into every runtime's tree, so
editing it changes its content hash in all 19 golden fixtures.

Regenerated with `npm run gen:golden`; the diff is exactly one line per fixture
— the gsd-core/workflows/review.md hash — and nothing else.

This is the second ripple of a workflow edit, alongside
tests/workflow-size-baseline.json which the first commit already updated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* chore(#2605): backfill changeset PR number (#2689)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 00:31:17 -04:00
2026-07-21 23:54:43 +00:00
2026-07-21 23:54:43 +00:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%