Tom Boucher 062f3fda90 chore(#2371): representative gate-fixture corpus + document-shaped property test (#2380)
* chore(#2371): representative gate-fixture corpus + document-shaped property test

Adds tests/fixtures/representative/ — a permanent corpus of verbatim,
incident-sourced fixtures (never author-invented) from #2286, #2347,
#2365, #2366, each labeled with its expected gate verdict in a
MANIFEST.json and driven through the real CLI gate entrypoint via
tests/representative-corpus.test.cjs.

Adds a document-shaped fast-check property test alongside the existing
writer-seeded bijection test in tests/api-coverage.test.cjs: the existing
generator produces rows and renders them through the writer, so the
document shape is a constant and it cannot fail against a decoy table;
the new one generates the document space instead.

Two gates (#2365, #2347) are still open, so their corpus/property
assertions are marked with node:test's official `todo` option — the test
executes and reports its failure without affecting the process exit code
(https://nodejs.org/api/test.html#test-options). The audit-uat corpus
(#2286, fixed by #2317) is a normal passing assertion, proving the
methodology works end to end and not just cataloguing gaps.

Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures
may not be derived from the gate's own writer, grammar, or docstring
examples; a negative fixture must come from a source that doesn't know
the gate exists.

No production src/*.cts changes — validation only, per #2371's scope.

Closes #2371

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields

Standards-axis review findings, all fixed:

- Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen)
  that were copy-pasted between the parse/render bijection test and the
  new document-shaped property test in tests/api-coverage.test.cjs — a
  future edit to one could have silently desynced the two properties.
  Hoisted to a single module-scope declaration both tests reference.

- Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome ->
  expectedReason. It asserted against the gate's `reason` field, but this
  codebase already has a real, different `outcome` field at parser
  altitude (extractDecisions' DecisionOutcome) — naming the manifest
  field after the wrong altitude's term was exactly the ambiguity the
  "Fixture provenance" rule this PR adds exists to eliminate.

- Removed the unused `role` field from three MANIFEST.json files (never
  read by any test) and wired the previously-dead per-fixture
  `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file
  assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's
  `results` array — catches a regression that moves items between the
  two fixture files while preserving the aggregate total, which the
  existing total_items check alone would miss.

No changes to test intent or coverage — same assertions, correctly named
and fully wired.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs

Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while
preparing this PR — unrelated to #2371's own changes, but a defect
found while working is fixed in place rather than deferred.

Leftover from #2368/#2370 (merged just before this branch rebased onto
it): the case 'capability' arm that needed these two requires was
relocated to bin/lib/capability-command-router.cjs, which already
requires both modules directly (lines 24-25) and is their only real
consumer (cmdCapabilityState, resolveCapabilityRuntimeState,
cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had
zero other references in the file and were never re-exported —
confirmed via grep across the file and its module.exports.

Behavior-preserving: Node's require cache means the underlying modules
still load exactly once via capability-command-router.cjs's own
requires; gsd-tools.cjs never used its now-removed local bindings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2371): replace todo-marked assertions with characterization tests

gsd-test's own JSONL result parser (gsd-test-runner's
internal/pipeline/parse.go, verified directly against that repo's
source) has no concept of node:test's `todo` option — it only
recognizes kind:"pass"|"fail" and hard-errors on anything else. A
{ todo: true } test whose body throws is counted as a real failure in
gsd-test's own verdict, exactly as if it weren't marked todo — proven
by an actual gsd-test run against this branch, which reported
outcome:"failed" with all six todo-marked assertions (the property
test plus five representative-corpus fixtures) in the failure list,
each carrying the correct raw node:test `todo` field the tool's parser
simply doesn't read.

Replaces todo with characterization: MANIFEST.json now carries both
the correct target verdict (expected*) and the exact current observed
verdict (currentBuggyOutput, directly verified against live CLI
output for all five fixtures). Tests assert currentBuggyOutput — an
honest, non-vacuous pin of today's known-broken reality that passes
today and will fail loudly the moment the referenced fix changes the
observed output, at which point the assertion should be flipped to
expected* and currentBuggyOutput deleted.

The document-shaped property test switches from throwing fc.assert to
non-throwing fc.check (returns RunDetails per fast-check's own docs)
and asserts report.failed === true directly, for the same reason.

Updates all prose (CONTRIBUTING.md, the fixture READMEs) that
previously claimed todo would be respected — that claim was
factually wrong for this repo's actual tooling and must not ship.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: regenerate golden-install-parity fixtures after rebase onto next

Rebasing onto the current next (which now includes #2381's
todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real
conflicts in all 18 golden-install-parity fixtures — expected, since
both branches changed the same gsd-tools.cjs hash entry. Resolved by
taking one side to unblock the rebase, then regenerating fresh from
source via npm run gen:golden and verifying the result; every file's
diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs,
correcting a stale intermediate hash from the arbitrary conflict pick.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 15:23:09 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%