Tom Boucher 57437071e7 test(#3244): codex smoke test across the emit, validate and repair seam (#3298)
* test(#3244): end-to-end codex smoke test across the emit and validate seam

Test-only. The only place in this epic where the emitter (Phase 1) and
the validator (Phase 2) meet the same bytes.

Every phase so far tested its own half against fixtures it authored, and
per #2371 a fixture written by the gate's own author can only confirm
what that author already believed. Rows 7 and 10 feed a REAL emitted
tree to Phase 2's checkCodexModelPosture: if the emitter and the
validator disagree about what "clean" means, nothing else in the suite
can see it.

Scope corrected from the epic, which scoped this to model_profile:
inherit. That profile is the one this epic does NOT change —
readGsdRuntimeProfileResolver returns null for it, so a Codex install
under inherit omitted the model before Phase 1 too, and a test scoped
only to it would pass identically before and after the change it exists
to prove. Row 1 uses `balanced`, the default and the path that actually
lost its pin; inherit is row 4, the unchanged control.

Honest classification, because this is a smoke test and claiming
otherwise would be false: NO row here is red-first. Phases 1 and 2 are
merged and working, so every row passes today. Rows 1-6 and 9 are
regression guards; rows 7 and 10 are cross-phase integration. Their
value is that nothing else can catch the drift they cover, not that they
are red now. The RED checkpoint is therefore skipped deliberately for
this phase rather than spent proving a tautology.

Rows 8 and 11 — the same cross-checks against Phase 3's repairer — are
deliberately ABSENT, not forgotten. Phase 3 (#3243) is still in CI and
its module does not exist on next yet. They land once it merges; writing
them against a surface that does not exist would have meant inventing
the assertion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3244): close the emit-validate-repair loop with rows 8 and 11

These were deliberately absent from the previous commit because Phase 3
had not merged; writing them against a module that did not exist would
have meant inventing the assertion. It has merged, so they land now.

Row 8: a freshly-installed tree reports ZERO changes from the sync's dry
run. If a brand-new install needs repair, the emitter and the repairer
disagree about what a correct file looks like.

Row 11 is the sharpest row in the phase: after installing with an
explicit real-Codex model_overrides pin, the sync reports that agent
skipped rather than synced, and the file is byte-identical afterwards.
An over-eager stripper that removes any `model` line passes every Phase
3 unit test and fails only here.

Byte-identity is asserted by comparing file CONTENTS, not mtime — a
rewrite with identical bytes still moves mtime, so an mtime check would
pass exactly the implementation this row exists to catch.

With these, all four cross-phase rows are in place, and this suite is
the only place in the epic where the emitter, the validator and the
repairer meet the same bytes. Every phase before this tested its own
half against fixtures it authored, and per #2371 those can only confirm
what their author already believed.

Both rows are cross-phase integration and pass today, like the rest of
this suite. No row here is red-first and none is described as such.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3244): extract withAgentsDir instead of try/finally in test bodies

Review finding. Rows 7 and 10 put try/finally directly inside the test
callback to save and restore GSD_AGENTS_DIR. CONTRIBUTING prohibits that
in test bodies — it masks failures — and permits it only inside
standalone helpers. The phase's own test matrix restated the rule and it
was violated anyway.

lint:ci does not catch this; it is a prose standard, which is precisely
why it survived to review.

It was also inconsistent with this file's own conventions:
runCodexInstall, runEffortSyncDryRun and captureStderr already factor
save/restore into helpers. withAgentsDir now follows the same shape and
both sites use it.

No assertion changed — both rows still drive the real
checkCodexModelPosture against the real installed tree, which is the
whole point of them. Swept the rest of the file: the only other try
blocks are inside the two pre-existing helpers and were already
compliant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 02:11:57 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%