* fix(#4881): trust worktree.baseRef:"head" in harness mode and keep the #4868 observation as what restores it under a WorktreeCreate hook The pre-dispatch base check still derived its harness-mode verdict from the retired #48 premise that the harness never reads worktree.baseRef: with "head" set and HEAD diverged from origin/HEAD it degraded every wave with baseref-head-ignored-by-harness. #4868 inserted an observation of a clean prior harness worktree at HEAD ahead of that comparison, but an execute-phase run never has one at the moment it checks — the base-check runs before any dispatch, a degraded wave creates no worktrees, and a wave that did run in worktrees has them removed and HEAD moved before the next check — so the common case was unchanged (#4881 repro states 1 and 4). Re-scope of the closed #4752 onto current next, with #4868 kept: - branch a trusts "head" in both isolation modes, the way the harness is measured to behave (#4588: three settings layers, three OSes), and the spawn-time exit-42 guard stays the observation-based backstop; - a Claude Code WorktreeCreate hook in any settings file the check reads, or a file that does not parse, withholds that trust — the hook creates the worktree without applying the setting — and the inferred comparison runs, degrading with baseref-head-bypassed-by-hook; - the #4868 observation (b2) now sits behind branch a: it is reached only when "head" was not trusted outright, and on a hook host it is what restores the trust — a hook that forks from HEAD leaves exactly that evidence, one that forks elsewhere never does. Its per-HEAD cache is unchanged. It is skipped under an explicit --observed-fork-base, which outranks an inference from a prior worktree; - --observed-fork-base <sha> threads a measured fork base through the evaluation (strict full-hex, TypeError otherwise). Tests: the #4868 rows are unchanged and still reachable (they run with the setting unset); the #4752 rows re-land, with the one exit-128 row updated to the degrade #4734 pinned since; a new #4881 block pins the start-of-run state (baseref-head, git never consulted), the hook + observation composition in both directions, and that an explicit observation skips the probe. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS * docs(#4881): rewrite the eight surfaces that still state the harness ignores worktree.baseRef:"head" Every prose surface #4868 left untouched still asserted the retired #48 premise as verified fact, starting with the step file the orchestrator reads. Each now describes the measured behaviour, the WorktreeCreate-hook exception, the --observed-fork-base input, and the #4868 observation as what lifts the hook degrade; docs/CLI-TOOLS.md gains the fork-from-head-observed reason row #4868 did not document. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS * chore(#4881): set changeset fragment pr to 4921 * fix(#4881): withhold the #4868 observation under the hook interlock The hook interlock this PR added withheld the worktree.baseRef:"head" trust but still let b2's prior-worktree observation restore it, and that observation cannot be attributed to the hook. The evidence is a clean agent worktree sitting at the orchestrator HEAD; nothing on disk records which creator left it there, so one the plain harness created BEFORE a WorktreeCreate hook was configured — with HEAD unmoved since — reads as evidence for the hook. It was the one fail-open branch in a mechanism documented as fail-closed. Keying the observation cache by hook configuration does not close it. observeHarnessForkFromHead has two legs: a HEAD-keyed cache and a live probe over .claude/worktrees/agent-*. A hook-keyed cache simply misses, and the miss falls through to the probe, which re-finds the same stale worktree and re-confirms. The probe takes no hook input at all. So the observation is not consulted under the interlock rather than re-keyed: on a hook host the only admissible positive signal is an explicit --observed-fork-base measurement of the dispatch in hand, and absent one the inferred comparison runs and a mismatch degrades with baseref-head-bypassed-by-hook, leaving the spawn-time exit-42 guard as the backstop. Scoped to the case branch a. declined to trust: "head" set AND a hook (or an unparseable layer) in the harness's path. With no "head" setting b2 is #4868's own arm and is unchanged, hook or not — re-scoping that trust is a separate question this PR does not open, and a test pins the boundary. Cost, stated: a hook host with "head" set, on a branch diverged from origin/HEAD and passing no observation, now runs sequentially. It still runs parallel when HEAD matches origin/HEAD. No workflow threads an observation today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ * refactor(#4881): drop the now-unused forkRef message-builder parameter buildMsgBaserefHeadIgnored stopped reading forkRef when the message was made mode-neutral, and the parameter was retained with `void forkRef;` for symmetry with its two sibling builders, which do read it. Symmetry is not reason enough to keep a dead parameter on a module-private function with one caller, so drop it (#4921 review). Behaviour is unchanged; the message text is pinned by an existing full-string assertion, which is what covers the only real risk here — transposing the two remaining arguments at the call site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ * test(#4881): exercise both FULL_SHA_RE alternatives at their boundaries The invalid-observation list pinned 39 and 41 hex around the 40-hex SHA-1 arm but left the 64-hex SHA-256 arm's own +/-1 boundary unexercised, which the repo's boundary-coverage convention asks for (#4921 review). Adds 63, 65 and a 64-length non-hex string. The regex already rejected all three -- this is coverage of correct behaviour, not a fix -- so it carries no negative control against a pre-fix base. That it is not vacuous was shown instead by widening the arm to {63,65}, under which the row fails by name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ * docs(#4881): finish the surface sweep the hook interlock owes Self-found by this round's own pre-push adversarial review, over six passes. Eight sites, four classes. FIVE were prose still asserting the stance the interlock overturned, which reads as live canon to anyone arriving cold: docs/CLI-TOOLS.md, docs/CONFIGURATION.md and gsd-core/references/planning-config.md each still said the hook degrade is lifted "unless/until a clean prior harness worktree is observed"; observeHarnessForkFromHead's own header still said a qualifying worktree "can only exist if the harness forked from HEAD"; and a test name still called the observation "required on a hook host" when it is now inadmissible there. My own sweep had grepped for "restores"/"lifts" and missed every one -- the ordinary failure of a grep, which returns what you thought to search for. The SIXTH is the same class one step worse: gsd-core/workflows/execute-plan.md still said flatly that Claude Code's isolation="worktree" "forks from origin/HEAD, not live local HEAD" -- in a paragraph THIS PR already edits, a few sentences after the clause it corrected. A tombstone makes only its own line clean; adjoining text asserting the dead stance is the other half of the same defect. Now qualified on the setting, with the hook exception named. The SEVENTH is a proof-strength overstatement that predates this PR, with a driven counterexample: a worktree created from an older base and since `git checkout --detach`ed onto HEAD is clean, sits at HEAD, and satisfies the probe identically, so "can only exist" was false. The worktree's own reflog does retain that original checkout -- the information is not lost, the probe simply does not consult it. The EIGHTH is that the header described only one of the function's two legs. A cache hit returns the prior conclusion without reading any worktree, so "the probe reads a worktree's present state" was true of the live probe and false of the cache. The header now separates them, and names the cache's blindness as a third reason the observation is inadmissible under a hook. Gaps seven and eight are inherited from #4868 and accepted there for the no-hook case. Nothing about the mechanism changes here; only what the header claims for it. No behavioural change -- comments, prose, and one test's registered name. Emitted-Drift-Ack-Growth: execute-plan.md — the Pattern A paragraph gained a qualifying clause: it stated flatly that Claude Code forks from origin/HEAD, which is the premise this PR retires, a few sentences after the clause the PR had already corrected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
GSD Core
Git. Ship. Done.
English · Português · 简体中文 · 日本語 · 한국어
A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.
What is GSD Core
GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.
How it works
Each milestone repeats the same five-step loop, one phase at a time:
- Discuss — capture implementation decisions before anything is planned
- Plan — research, decompose, and verify the plan fits a fresh context window
- Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
- Verify — walk through what was built; diagnose and fix before declaring done
- Ship — create the PR, archive the phase, repeat for the next one
Quickstart
npx @opengsd/gsd-core@latest
The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.
On another runtime or without Node.js? See Install on your runtime.
Once installed, start a new project or onboard an existing repo:
/gsd-new-project # greenfield project
/gsd-onboard # existing codebase
New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.
Documentation
What's new in 1.7.0 → docs/whats-new-1.7.0.md
Tutorials — learning by doing:
How-to guides — task-focused recipes:
Reference — authoritative facts:
Explanation — concepts and design decisions:
Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.
Why it works
Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.
Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.
Community
| Project | Platform |
|---|---|
| gsd-opencode | Original OpenCode port |
| Discord | Community support |
Star History
License
MIT License. See LICENSE for details.
Claude Code is powerful. GSD Core makes it reliable.