* test(#4070): failing-first regression for shard1 aux-suite budget imbalance - selectShard has no way to reserve virtual weight on a bin, so the LPT unit-test packer cannot account for shard 1's fixed aux-suite cost (integration/security/install/slow all pinned to shard 1/3). - test.yml wires no such reserve into the workflow. - ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test` was stale (7m12s from run 30677442953, predating the aux-suite growth); corrected to the real evidence cited in #4070 (13m48s / cancelled at ~14m51s), which now honestly fails the file's own 1.5x headroom policy against the current 15-minute cap. All three are expected RED on this commit; see .gsd/bug/fix-4070-shard1-aux-suite-budget/50-test-matrix.md. * fix(#4070): reserve shard 1's aux-suite cost out of the LPT unit-test packer selectShard now accepts an optional initialWeights array giving one or more bins a virtual head start before any file is placed, so LPT converges each bin's FINAL total (assigned weight + head start) toward equal instead of raw assigned weight alone. test.yml wires RUN_TESTS_SHARD_RESERVE=1:77 into the full-scope unit-test step (gated on matrix.scope == 'full', so the unrelated windows lane is unaffected) -- 77 weight units is the empirical conversion of the aux suites' ~220s measured fixed cost, derived against the real tests/test-timings.json (see the diagnosis artifact for the full computation). Also corrects two pieces of now-stale bookkeeping this issue exposed: - ci-test-job-timeout-budget.test.cjs's LANE_COSTS entry for job `test` carried a 7m12s figure that predated the aux-suite growth; corrected to the real pre-fix evidence (13m48s / cancelled at ~14m51s, issue #4070), which requires raising timeout-minutes from 15 to 21 (1.5x headroom over the real worst-case measurement) to satisfy the file's own policy. - test.yml's job-header and matrix comments, which still claimed the aux suites cost "~1m35s combined" (they now measure ~216-224s). Closes the gap the previous commit's failing-first tests proved: selectShard had no reserve-capacity mechanism and test.yml wired none in. * fix(#4070): correct the reserved-weight property bound; cover main()'s reserve bounds check Isolated code review found a genuine gap and gsd-test's real run confirmed a real test bug it exposed: - The fast-check property "no shard exceeds average(+reserve) + heaviest file" was falsified by gsd-test itself (weights=[1,1,1], total=2, reserve=6 on bin 0): selectShard is correct, the BOUND was wrong. A reserve large enough that its bin never receives a real item stays at exactly that reserve forever -- no amount of routing real items elsewhere can dilute a fixed head start below itself -- so the true bound is max(reserve, the classic Graham term), not the Graham term alone. Verified the corrected bound against the exact counterexample plus 20,000 additional random trials (zero violations) before re-running gsd-test. - Isolated review (MAJOR): the shard-total bounds check on RUN_TESTS_SHARD_RESERVE and its console.error fallback in main() were untested end-to-end -- parseShardReserve itself has no concept of the shard total, so only main() enforces that guard, and nothing exercised it through the subprocess seam. Added an E2E harness test that sets RUN_TESTS_SHARD_RESERVE to an out-of-range index via the real CLI, asserts the fallback warning fires, AND asserts the resulting file selection is byte-identical to a no-reserve control run against the same injected timings table -- proving the reserve was actually ignored, not just that a warning printed. * chore(#4070): backfill changeset PR number pr:0 -> pr:4072 * fix(#4070): strip leaked RUN_TESTS_SHARD_RESERVE from the harness test's child env Real GH Actions CI on this PR (run 33288554040, ubuntu shard 2/3) failed 7 tests in the shard-partitioning describe block, all with the same symptom: `run-tests: no tests in suite "all"` where a real file count was expected. gsd-test's own dockerized bench run never showed this, and ubuntu shards 1/3 and 3/3 (which run the same test.yml step) passed clean -- the discrepancy is the tell: only shard 2/3 happened to schedule this specific test FILE for that run, and the outer CI job's own environment is where the leak lives. Root cause: test.yml's "Run unit tests" step now sets RUN_TESTS_SHARD_RESERVE=1:77 (this issue's own reserve mechanism) on the OUTER job that runs `npm run test:coverage:unit:raw -- --shard N`. The harness test file's runHarness() helper spawns run-tests.cjs as a CHILD of that same job via `{...process.env, ...extraEnv}`, so every pre-existing --shard test in this describe block silently inherited the ambient reserve -- even though none of them know it exists. A reserve of 77 weight units utterly dwarfs the ~0.3 total weight of the 9-file synthetic fixtures these tests use (none are in the real timings table, so all fall back to the same tiny median weight), so shard index 1 is routed zero files every time -- exactly the observed "no tests" failures, and exactly the skewed 5/4 split observed on the shard-2 test that expected a plain 3/3/3 round-robin. Reproduced locally end to end (set RUN_TESTS_SHARD_RESERVE=1:77, spawn the old runHarness against a synthetic 9-file fixture, --shard 1/3 -- reproduces the exact "no tests in suite \"all\"" stderr) and confirmed the fix (env stripped unless a test opts in via extraEnv, as the #4070 E2E bounds-check test already does) resolves it, before re-running gsd-test. This is a genuine bug this PR introduced -- a new ambient env var that a pre-existing subprocess-spawning test helper didn't know to isolate against -- not a pre-existing flake and not resource contention. --------- Co-authored-by: sim <sim@local>
GSD Core
Git. Ship. Done.
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
What is GSD Core
GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.
How it works
Each milestone repeats the same five-step loop, one phase at a time:
- Discuss — capture implementation decisions before anything is planned
- Plan — research, decompose, and verify the plan fits a fresh context window
- Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
- Verify — walk through what was built; diagnose and fix before declaring done
- Ship — create the PR, archive the phase, repeat for the next one
Quickstart
npx @opengsd/gsd-core@latest
The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.
On another runtime or without Node.js? See Install on your runtime.
Once installed, start a new project or onboard an existing repo:
/gsd-new-project # greenfield project
/gsd-onboard # existing codebase
New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.
Documentation
What's new in 1.7.0 → docs/whats-new-1.7.0.md
Tutorials — learning by doing:
How-to guides — task-focused recipes:
Reference — authoritative facts:
Explanation — concepts and design decisions:
Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.
Why it works
Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.
Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD Core makes it reliable.