* ci(#2952): budget CI job timeouts by headroom over measured cost `origin/next` was red. The only failing check was `Required tests`, and its sole cause was `test (ubuntu-latest, 24)` reported `cancelled` — GitHub's conclusion for a job that exceeds its `timeout-minutes`, not a button press. That job's `ubuntu-latest / 24` entry is the only `scope: full` matrix entry: it runs the whole unit suite under c8 coverage, then the scripts/ coverage floor, integration, security, install and slow, serially on one runner. On next@5a0a9f097 it ran 15m16s against `timeout-minutes: 15` and was axed 23s into `npm run test:slow`. Projected to a completed slow step (27s on the last green run) the lane costs ~15m20s. Confirmed hypothesis: the budget, not the suite. The lane had been riding the ceiling all day — 12m03s, 11m34s, 11m50s, 14m25s, 14m51s — and crossed on three of the last four full-lane runs (d2d2f7c08,07603df8f,5a0a9f097). There is no pathological test: the unit run is cost-first bin-packed into 12 chunks, chunk 1 is gated by run-tests-harness.test.cjs at 169s (expensive by design — it spawns real harness subprocesses, one of which exercises the per-chunk timeout), and the remaining chunks are 32-94s. 769s is the honest cost of 689 files under c8. Re-running could not have helped; the work exceeded the budget. Review of the first cut surfaced the same defect one runner away: `full test (windows-latest, 22, shard 3/3)` reached 18m59s against its own 20-minute cap on05b170e44(94%) and 18m14s on81eeb8a53(91%). That lane has already blown its cap twice (#1051, #1212). Fixed here rather than deferred. The first cut also asserted `test >= test-full`, which is unsound — those two budgets are dominated by different platforms, so their ordering carries no meaning. Replaced with the invariant that actually generalises: every lane is held to a headroom FACTOR over its own measured cost. `test` 15 -> 25 (1.5x of 16m), `test-full` 20 -> 30 (1.5x of 19m), `test-inert` unchanged at 15. tests/ci-test-job-timeout-budget.test.cjs locks that rule. No unit test can prove a lane still FITS its budget — only a real run measures that — but a budget can no longer be lowered back beneath what its lane is known to need, and a lane that gets slower must be re-measured rather than excused. Separately: the earlier `failure` at05b170e44was an unrelated, already-fixed CONTEXT-INDEX.json drift (07603df8fre-synced it; lint-tests is green at HEAD). 07603df8f's own run hit this same timeout, which is why it never reported green. Refs #869, #1051, #1212 * ci(#2952): shard the full test lane instead of widening its cap The `scope: full` lane was the only unsharded lane in this file. It ran the entire unit suite under c8 on one runner, grew past a 15-minute cap, and reddened `next`. Raising the cap bought room; it did not change the shape, and the same lane would have walked back into the ceiling. Shard it, the way #1212 answered this for the Windows lane. Balance comes from measurement, not file counts. scripts/run-tests.cjs already partitions by measured per-file duration using LPT (#2472); the table it reads was 10 days stale — 638 of 695 files timed, 64 missing, including the whole context-predicates group. Regenerated from a verified matrix run: 700 files, 0 missing. On that table the 685-file unit suite splits 19.37m / 19.37m / 19.37m — 0.0% spread — and the split is a total, disjoint cover with 0 files dropped. Completeness, disjointness, balance and determinism of the partition itself are already pinned against selectShard in run-tests-harness.test.cjs, including a fast-check property, so this change does not restate them. Sharding a COVERAGE run is the part that needs care. A per-shard percentage is meaningless — shard 2 never executes shard 1's files, so those read 0% — and leaving the gate on the shards would have quietly measured a third of the tree. Each shard now renders no report and only leaves raw V8 dumps; a new `coverage-gate` job merges all three into one coverage/tmp and runs the gate there. c8's default temp directory is where the download lands, so the ≥70% lines / ≥60% branches gate and the ≥55% scripts floor run unmodified against merged data. Both surfaces call the same npm scripts rather than inlining c8 into YAML, so the include/exclude globs and both thresholds stay defined once in package.json. The workflow holding its own copy is the divergence this repo has a rule against, and the new test cross-checks package.json so an inline reintroduction fails rather than drifts. tests/ci-full-lane-sharding.test.cjs covers the two ways this stays GREEN while being wrong: an incomplete shard set (declare 1/3 and 2/3, never 3/3, and a third of the suite silently stops running) and a coverage gate that stops being required. required-tests now depends on coverage-gate and fails on it, while still tolerating `skipped` so docs-only PRs are not blocked. `timeout-minutes: 25` on the lane is deliberately left alone. The budget test requires a real measurement before a lane's declared cost changes, and the sharded cost is not measured until this PR's own CI run. Refs #1212, #2472 * ci(#2952): tighten the sharded lane's budget to its measured cost The sharding commit deliberately left `timeout-minutes: 25` alone, because tests/ci-test-job-timeout-budget.test.cjs requires a real measurement before a lane's declared cost changes and the sharded cost did not exist yet. It exists now. Run 30677442953: shard 1/3 7m12s, 2/3 4m32s, 3/3 3m59s, and coverage-gate 1m20s. Shard 1 is the long pole because the unsharded aux suites ride along on it, which is deliberate — they total ~1m35s and sharding them would cost more than it saves. So the lane's budget is 15 against a slowest measured shard of 8 minutes (~1.9x), and coverage-gate joins LANE_COSTS at 2 minutes. 15 is the same number the lane blew before sharding; the work behind it is now a third the size. Merged coverage was checked against the pre-shard single-runner baseline rather than assumed from a green check: 94.36 stmts / 96.3 funcs / 94.36 lines identical, branches 84.22 vs 84.21 — one branch across two different trees, noise rather than a regression. --------- Co-authored-by: sim <sim@local>
GSD Core
Git. Ship. Done.
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
What is GSD Core
GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.
How it works
Each milestone repeats the same five-step loop, one phase at a time:
- Discuss — capture implementation decisions before anything is planned
- Plan — research, decompose, and verify the plan fits a fresh context window
- Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
- Verify — walk through what was built; diagnose and fix before declaring done
- Ship — create the PR, archive the phase, repeat for the next one
Quickstart
npx @opengsd/gsd-core@latest
The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.
On another runtime or without Node.js? See Install on your runtime.
Once installed, start a new project or onboard an existing repo:
/gsd-new-project # greenfield project
/gsd-onboard # existing codebase
New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.
Documentation
What's new in 1.7.0 → docs/whats-new-1.7.0.md
Tutorials — learning by doing:
How-to guides — task-focused recipes:
Reference — authoritative facts:
Explanation — concepts and design decisions:
Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.
Why it works
Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.
Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD Core makes it reliable.