* test(#3271): guard against a folded suite appearing twice in one host Adds local/no-duplicate-fold-marker, an AST rule that reports the second and every subsequent `folded:<name>` marker in a host file, plus RuleTester cases and a tree-wide regression assertion. Failing-first on purpose: the rule is registered at error and the 25 duplicated regions are still present, so eslint and the new tree-wide test are RED. The deletions land in the next commit. The marker key is the whitespace-delimited token after `folded:` — not the issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and feat-443-effort-fast-mode are two distinct folded suites. Refs #3271 * fix(#3271): delete 25 duplicated folded suites from three install hosts Three consolidated install suites each carried a verbatim second copy of a contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and green — each duplicated block registered and ran twice on every lane. tests/install.test.cjs 5981-9937 (3957 lines, 18 blocks) tests/install-minimal-hooks.test.cjs 2734-4015 (1282 lines, 5 blocks) tests/install-write-confinement.test.cjs 1754-2321 ( 568 lines, 2 blocks) Introduced by6d072435d(#1975 re-applying #1970's hunks on a tree that already had them, 2026-07-03) — one stale-base re-application, three files, one commit. Verified by marker-count bisect: 1 at4f779eda4and0cc7a1a42, 2 from6d072435donward. The later copy is deleted in each case, so every file returns to what its authoring batch produced and blame on the surviving lines stays accurate. local/no-duplicate-fold-marker, red on the previous commit, is now green. tests/model-resolver.test.cjs is untouched: the issue lists it, but its two blocks are folded from two different files and are not identical. It is a false positive of the issue's own grep, whose `[a-z0-9-]*` key truncates at `.`. Fixes #3271 * test(#3271): property-test marker identity and pin the alias non-goal Three review findings, all fixed inline: 1. foldMarkerOf is a parser and carried no fast-check property test. Raised independently by the /code-review standards axis and the isolated adversarial pass; the file already establishes the fc.property-driving-ruleTester idiom for a sibling rule. Added, two arms over markers generated from [a-z0-9-._]: the same marker twice always reports exactly once against firstLine 1, and two distinct markers never collide. The alphabet includes `.` on purpose — an implementation keyed on the issue's [a-z0-9-]* slice passes arm 1 and fails arm 2, which is exactly the model-resolver false positive. 2. meta.docs.category was the novel value 'Test hygiene'; all 16 sibling local rules use 'Best Practices', 'Portability' or 'Reliability'. Now 'Best Practices'. 3. A call through a further alias (const d = __foldDescribe) was unreported and undocumented — accidental rather than deliberate. It is now the fourth entry in the rule's documented non-goals, with the reason, and pinned by a valid RuleTester case so it cannot drift silently. Refs #3271 * test(#3271): name the step and elapsed time when a baseline build fails buildBaselineAtRef runs four bounded steps and, when one exceeded its bound, threw a bare "spawnSync ETIMEDOUT" naming neither the step nor how long anything took. Diagnosing one real failure took four separate experiments to recover information the throw already had. Each step is now timed, and any throw carries the breakdown: which step failed, its elapsed time, the timings of every step that completed before it, all three bounds, and the tail of the child's captured stdout/stderr. The failure message is deliberately the carrier. On the remote runner the captured output field comes back empty in failures.json while error and stack survive verbatim, so the message is the only channel that reaches a reader of a remote verdict. Refs #3271 * fix(#3271): size the baseline generator bound for the machine it runs on Instrumentation from a real remote-runner failure gave the breakdown: git-worktree-add=15.1s npm-run-build-lib=19.8s gen-emitted-baseline=FAILED@300.1s Steps 1 and 2 are comfortable. Only the generator exceeds its bound, and it is not hung — it needs more than 300s there. Measured ladder for that step: ~22s idle in a container, ~39s end-to-end in a clean container, ~142s with 8 CPU burners on 8 cores, and >300s under the real suite. Its cost is 19 sequential installer spawns, and spawn latency is exactly where a container degrades worst (3.9x slower than host, against 1.1x for file IO) — which is why a CPU-only load test did not reproduce it and why four earlier hypotheses (container slowness, network, shallow clone, CPU contention) all measured clean. The 300s bound was sized on an idle machine for a step that never runs on one. Under the remote runner the on-disk baseline cache is structurally absent — CI restores it via actions/cache keyed on github.event.pull_request.base.sha, a key that exists only inside GitHub Actions — so this slow path runs on every remote verification. The result: this gate has passed 0 times in 754 runs, failing 80 times and never once executing successfully. Raised to the 600000ms ceiling that local/no-unbounded-spawn treats as the largest meaningful bound; the other two bounds are untouched. This makes the gate RUN, which is the point: the alternative considered and rejected was degrading the timeout to a skip, and that was measured to turn the suite green with the gate silently not running at all. The real remedy is making the cache reachable from the remote runner so the in-job build returns to being the rare fallback ADR-2719 §5 describes. That is a gsd-test-runner change, not one this repo can make. Refs #3271 * fix(#3271): tolerate an overlay source that vanishes mid-walk Observed on the remote runner, three runs across three different branches: ENOENT: no such file or directory, link '/work/hooks/dist/gsd-config-reload.js' -> '/tmp/gsd-2930-overlay-6nOZay/hooks/dist/gsd-config-reload.js' buildOverlayRepo enumerates names with readdirSync and then acts on each one, so statSync, copyFileSync and linkSync all sit in a TOCTOU window. hooks/dist is regenerated by an ATOMIC REPLACE (scripts/build-hooks.js unlinks and renames), so any concurrently running test that rebuilds hooks retires a just-listed name mid-walk and the overlay dies on it. linkOrCopyFile already tolerated EXDEV and EPERM; ENOENT went straight through. On ENOENT the source is now re-examined ONCE rather than slept on. An atomic rename is a single syscall, so by the time the failure surfaces the successor is either already in place (the retry succeeds) or the path has genuinely left the tree, in which case there is nothing to mirror and the leaf is skipped. No sleep and no spin: a timing-based wait here would be the very flake being fixed. Every other errno still propagates untouched, so a real permission or IO fault stays a hard failure. Five tests hold the boundary: gone-for-good skips without retrying, mid-replace retries exactly once and places the file, EACCES still throws, a real linkSync ENOENT is injected by monkeypatching fs and restoring it in a finally (never a mode-bit trick, which root bypasses), and isMissingPath accepts only ENOENT. Refs #3271 * fix(#3271): order the timeout ladder inward-out and lock it Two review blockers, both real. The generator bound had been raised to 600000ms — exactly the whole-chunk timeout in scripts/run-tests.cjs:973. A step bound equal to the chunk ceiling loses the race: the chunk is killed first and the failure arrives as an opaque "no failed step" kill, so the per-step diagnostic added a commit earlier was built and then made unreachable in the same change. Separately the #2767 test declared a per-test timeout of 300000ms, BELOW the inner bound it was meant to permit, so it could still die at the exact 300s ceiling this was supposed to lift — via node:test's timeout rather than spawnSync's. Its sibling declared 900000ms, above the chunk ceiling, which is the same opaque-kill hazard from the other direction. The three bounds only produce a useful failure if they fire inward-out, so they now do: step 360s, per-test 480s, chunk 600s. 360s is ~3x the passing observation (91.6s / 115.8s) and 20% above the censored 300.1s timeout, while leaving 240s of chunk headroom for every other file sharing it. Four tests lock the ordering, including a drift guard on the exported values — without it, editing a call site's literal timeout would leave the ordering assertions passing while the real ladder inverted. Also from review: - err.gsdBaselineStep and err.gsdBaselineTimings were written and never read anywhere in the tree; only the rewritten message is consumed. Removed rather than kept as speculative surface. - buildOverlayRepo discarded placeVanishableLeaf's boolean at both call sites, so a vanished leaf left the overlay with no accounting at all. It now collects the skipped paths and warns once. Not thrown: a source that left the tree really is not part of the snapshot, and throwing would reintroduce the crash the tolerance removes — but silence would let a dropped leaf resurface later as an unrelated missing-file assertion. - The instrumentation commit shipped no test. One now drives a real failure and asserts the message names the step, its elapsed time, and the bounds. Refs #3271 * chore(#3271): backfill the changeset PR number * fix(#3271): bound a hook fan-out as its own class, not as a bare probe CI failure on PR #3285, job full test (windows-latest, 22, shard 2/3) — every other lane green, including windows-latest node 24 across all three shards: not ok 1 - blocks push when any to-be-pushed commit matches local blocked regex error: bash .githooks\pre-push failed — outcome=timed_out exitCode=null stderr= duration_ms: 15040.2168 A bound, not a hang: the test supplies stdin via input:, so the hook is not blocked reading its ref list, and the duration lands exactly on the 15000ms bound. The site used PROBE_TIMEOUT_MS, which tests/helpers/timeouts.cjs documents as "a single short CLI query or node -e probe against a temp fixture". This is not that. It spawns bash running .githooks/pre-push, and the hook then invokes a MOCK git that is itself a bash script, so one runHook is roughly four Git Bash spawns. On Windows each is Defender-scanned and the first hook test in a file pays cold start on top. That module's own docstring warns against precisely this: a call site that differs from its class must not be forced onto a shared value that does not describe it. HOOK_FANOUT_TIMEOUT_MS is that missing class — 60000ms, 4x the bound that failed and half INSTALL_TIMEOUT_MS, which is the right order: a hook fan-out is much lighter than a full installer run and far heavier than reading back a version string. Two tests lock the ordering against both neighbours, including one asserting real margin over the censored 15040ms observation, since a bound that merely matched what was measured would be the same defect again. Scoped deliberately: the other ~360 runHook sites keep their current bounds. This adds the norm and applies it where a real failure demonstrated the need, rather than sweeping a value across sites with no evidence for any of them. Refs #3271 --------- Co-authored-by: sim <sim@local>
GSD Core documentation
Documentation is organised into four quadrants: tutorials help you learn by doing, how-to guides solve specific tasks, reference states authoritative facts, and explanation explores concepts and design decisions.
Language versions: English · Português (pt-BR) · 日本語 · 简体中文
Tutorials
- Your first project — install to first shipped phase, one guaranteed path
- Onboarding an existing codebase — bring GSD Core to a brownfield repo
- Build your first capability — author a tiny declarative capability and watch it act in the loop
- Install your first capability — install a third-party capability end-to-end: consent, verify, check for updates, remove
How-to guides
- Install on your runtime — runtime-specific install steps for all 16 supported runtimes
- Install a minimal GSD and add skills later — install only the core skills, then grow the surface with profiles and
/gsd-surface - Attach a plugin-provided skill to a GSD agent — use the
global:plugin:skillentry form to load Claude Code plugin skills into agent prompts - Discuss a phase — capture implementation decisions before planning begins
- Resolve edge-coverage findings — turn the spec phase's surfaced domain-boundary edges into covered, dismissed, or backstopped spec decisions
- Resolve prohibition findings — turn the spec phase's surfaced must-NOT constraints into resolved, dismissed, or deferred spec decisions
- Resolve an ESLint glob-coverage finding — bring a source file that matches no lint rule under coverage, or record a reasoned exemption
- Plan a phase — run research, decompose work, and verify plan quality
- Execute a phase — run plans in parallel waves with fresh-context subagents
- Verify and ship — walk through completed work, diagnose failures, and create the PR
- Catch complexity before it compounds — enable the post-execute refactor hook, read a proposal's score vs. anchor delta, and accept or decline it
- Run phases autonomously — use autonomous mode for unattended phase execution
- Handle quick and fast tasks — use
/gsd-quickand/gsd-fastfor ad-hoc work outside the phase loop - Configure model profiles — switch between quality, balanced, and budget model tiers
- Set up cross-AI review — configure a second AI to review code produced by the primary agent
- Work in parallel with workstreams — run independent lines of work simultaneously using workstreams
- Isolate work with workspaces — use workspaces to sandbox experimental or risky changes
- Debug a failed execution — diagnose and recover from broken or incomplete phase execution
- Interpret scope-conformance warnings — read the advisory the worktree-wave merge emits when a plan branch commits outside its declared scope
- Interpret
state validateresults — read thescopereason codes and tell "nothing to report" apart from "could not look" - Spike and sketch — use
/gsd-spikeand/gsd-sketchfor exploratory work before committing to a plan - Design a UI phase — use the UI phase loop for frontend and visual work
- Develop a Capability for GSD 1.5+ — add feature Capabilities, hook fragments, and registry entries
- Ship a reviewer lane in your capability — declare a
reviewerbody so/gsd-reviewdiscovers, invokes, and renders your external review CLI or model endpoint - List your reviewer lane in the registry — publish a lane you have built to the Reviewer Lane Registry so other people can find and install it
- Take over a capability or EoS integration — assume maintainership of an existing third-party capability, reviewer lane, or EoS host integration through a handoff, an adoption fork, first-party absorption, or a de-listing
- Add or update a host's integration — set a host's documentation-sourced
runtime.hostIntegrationaxes (ADR-1239 Phase A), with theundocumentedsentinel rule - Turn a capability off (and keep it off) — disable a capability via the surface, or gate individual hooks off without removing the capability
- Drive GSD from a tracker issue — start a phase from a GitHub, Linear, or Jira issue
- Migrate from GSD 2 — upgrade an existing GSD 2 project to GSD Core
- Update GSD — re-run the installer to pick up the latest release
- Clean up get-shit-done-cc — remove leftover old-package artifacts that cause a spurious
⬆ /gsd-updateindicator after migrating to@opengsd/gsd-core - Fix the worktree base-mismatch (exit 42) error — resolve the branch-divergence condition that halts parallel phase execution
- Recover and troubleshoot — fix common problems, rebuild context, and uninstall
Reference
- Commands — every command with flags and examples
- Configuration — full config schema, model profiles, git branching strategies
- CLI tools —
gsd-tools.cjsprogrammatic API for workflows and agents - JSON error mode —
gsd-toolsfailure channels: faults (stderr, exit 1) vs degraded results (stdout, exit 0), and the reason-code taxonomy - Features — complete feature index
- Inventory — installed skills and surface map
- STATE.md schema — field-by-field reference for
.planning/STATE.md - CONTEXT.md schema — field-by-field reference for
.planning/phases/<N>/CONTEXT.md - PLAN.md schema — field-by-field reference for
.planning/phases/<N>/PLAN.md - Planning artifacts — all
.planning/files and their roles - Review and verification capabilities — code review, security, and Nyquist capability ownership and hook contracts
- Gate predicates — canonical specification of the phase-gate predicate vocabulary
- Capability matrix — generated catalogue of every capability's role, tier, extension points, hook kinds, and
engines.gsd - Capability manifest — the full
capability.jsonschema and validation rules gsd capabilitycommand — install / update / remove / list reference for third-party capabilities- Workflow fragments — in-file
<!-- gsd:section -->marker grammar for fragmentizing workflow markdown at emission time - Reviewer Lane Registry — generated catalogue of third-party reviewer lanes, with their flags, transport, and install commands
Explanation
- Context engineering — how context rot forms and how GSD Core prevents it
- The phase loop — design rationale for the Discuss → Plan → Execute → Verify → Ship cycle
- Multi-agent orchestration — how subagents are spawned, scoped, and coordinated
- Security model — trust boundaries, permissions, and safe automation
- The capability trust model — why third-party capabilities are gated by consent + integrity + reversibility, not a sandbox
- How overlay capabilities compose — why first-party always wins and how the loader resolves precedence, conflicts, and fail-open load-failure warnings
- Architecture — system architecture, agent model, and data flow
- The Embeddable Orchestration System — one public, versioned contract for embedding GSD across many hosts
- Discuss modes — assumptions mode vs interview mode for
/gsd-discuss-phase - Context monitoring — context window monitoring hook architecture
- Issue-driven orchestration — recipe for driving GSD from a tracker issue using existing primitives
Related
- What's new in 1.7.0 — curated highlights of the 1.7.0 release
- Root README — landing page, quickstart, and documentation overview
- Changelog — release history