Tom Boucher ec22377a9d fix(#4351): run-scope preserved review evidence (#4713)
* fix(#4351): run-scope preserved review evidence

The preserve block copied lane output into a flat .review-diagnostics/
using each file's source basename. A lane slug is stable across runs, so
the destination was stable across runs too -- and cp over an existing
file is a success, so a second review of the same phase destroyed the
first run's evidence with no error and no warning, in the one directory
that exists to outlive the rm -rf beside it.

Copy into one subdirectory per run instead. $RUN_DIR is mktemp -d, so
its basename is already unique per run by construction; the UTC stamp in
front is only a sort key and is omitted if date fails. Destination-only:
nothing inside $RUN_DIR is renamed, because both
prepare_trimmed_prompt_for_reviewer and the lane invocation resolver
depend on those exact basenames.

Verified by extracting the real fence and running it twice against one
phase dir: before, one report survived and it was run 2's; after, both.

Emitted-Drift-Ack-Growth: review.md — per-run diagnostics subdirectory plus the comment explaining why uniqueness cannot come from the clock
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#4351): cover repeated runs, resolve through the run subdir

Adds the two-run rows the issue asks for: both runs' reports recoverable,
and a lane failing identically twice leaving two stubs. Both assert on
CONTENT, not a file count -- a clobber producing the same number of
files would pass a count-only check, and the defect is that run 1's
bytes were replaced.

runWriteReviewsFlow gains an optional phaseDir so a caller can run the
flow twice against one phase directory, which is the only arrangement
that can observe the overwrite. Omitted, it mints a fresh one as before.

The existing rows asserted a flat readdir of the diagnostics root, which
the fix makes stale. They now resolve through preservedPath/preservedNames
so they keep asserting WHICH files were preserved rather than silently
becoming assertions about the layout; the layout is pinned once,
explicitly, by oneRunSubdirectoryPerRun_4351.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4351): add changeset fragment

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4351): name the run, not the clock, in the changeset

Adversarial review: the fragment said each run gets its own "timestamped
subdirectory", which reads as though the timestamp provides the
separation. It does not -- collision-safety is mktemp's random basename,
and the stamp is a sort key that is dropped entirely when date fails.
The code comment already said so; the user-facing text now agrees.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4351): backfill the changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 23:58:38 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%