0xdhx 238bee7b03 fix(#4881): trust worktree.baseRef:"head" in harness mode; a WorktreeCreate hook withholds it and only a measured fork base restores it (#4921)
* fix(#4881): trust worktree.baseRef:"head" in harness mode and keep the #4868 observation as what restores it under a WorktreeCreate hook

The pre-dispatch base check still derived its harness-mode verdict from the
retired #48 premise that the harness never reads worktree.baseRef: with
"head" set and HEAD diverged from origin/HEAD it degraded every wave with
baseref-head-ignored-by-harness. #4868 inserted an observation of a clean
prior harness worktree at HEAD ahead of that comparison, but an
execute-phase run never has one at the moment it checks — the base-check
runs before any dispatch, a degraded wave creates no worktrees, and a wave
that did run in worktrees has them removed and HEAD moved before the next
check — so the common case was unchanged (#4881 repro states 1 and 4).

Re-scope of the closed #4752 onto current next, with #4868 kept:

- branch a trusts "head" in both isolation modes, the way the harness is
  measured to behave (#4588: three settings layers, three OSes), and the
  spawn-time exit-42 guard stays the observation-based backstop;
- a Claude Code WorktreeCreate hook in any settings file the check reads,
  or a file that does not parse, withholds that trust — the hook creates
  the worktree without applying the setting — and the inferred comparison
  runs, degrading with baseref-head-bypassed-by-hook;
- the #4868 observation (b2) now sits behind branch a: it is reached only
  when "head" was not trusted outright, and on a hook host it is what
  restores the trust — a hook that forks from HEAD leaves exactly that
  evidence, one that forks elsewhere never does. Its per-HEAD cache is
  unchanged. It is skipped under an explicit --observed-fork-base, which
  outranks an inference from a prior worktree;
- --observed-fork-base <sha> threads a measured fork base through the
  evaluation (strict full-hex, TypeError otherwise).

Tests: the #4868 rows are unchanged and still reachable (they run with the
setting unset); the #4752 rows re-land, with the one exit-128 row updated
to the degrade #4734 pinned since; a new #4881 block pins the start-of-run
state (baseref-head, git never consulted), the hook + observation
composition in both directions, and that an explicit observation skips
the probe.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS

* docs(#4881): rewrite the eight surfaces that still state the harness ignores worktree.baseRef:"head"

Every prose surface #4868 left untouched still asserted the retired #48
premise as verified fact, starting with the step file the orchestrator
reads. Each now describes the measured behaviour, the WorktreeCreate-hook
exception, the --observed-fork-base input, and the #4868 observation as
what lifts the hook degrade; docs/CLI-TOOLS.md gains the
fork-from-head-observed reason row #4868 did not document.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RPXQQPzHGintWtbhBoQCRS

* chore(#4881): set changeset fragment pr to 4921

* fix(#4881): withhold the #4868 observation under the hook interlock

The hook interlock this PR added withheld the worktree.baseRef:"head"
trust but still let b2's prior-worktree observation restore it, and that
observation cannot be attributed to the hook. The evidence is a clean
agent worktree sitting at the orchestrator HEAD; nothing on disk records
which creator left it there, so one the plain harness created BEFORE a
WorktreeCreate hook was configured — with HEAD unmoved since — reads as
evidence for the hook. It was the one fail-open branch in a mechanism
documented as fail-closed.

Keying the observation cache by hook configuration does not close it.
observeHarnessForkFromHead has two legs: a HEAD-keyed cache and a live
probe over .claude/worktrees/agent-*. A hook-keyed cache simply misses,
and the miss falls through to the probe, which re-finds the same stale
worktree and re-confirms. The probe takes no hook input at all. So the
observation is not consulted under the interlock rather than re-keyed:
on a hook host the only admissible positive signal is an explicit
--observed-fork-base measurement of the dispatch in hand, and absent one
the inferred comparison runs and a mismatch degrades with
baseref-head-bypassed-by-hook, leaving the spawn-time exit-42 guard as
the backstop.

Scoped to the case branch a. declined to trust: "head" set AND a hook
(or an unparseable layer) in the harness's path. With no "head" setting
b2 is #4868's own arm and is unchanged, hook or not — re-scoping that
trust is a separate question this PR does not open, and a test pins the
boundary.

Cost, stated: a hook host with "head" set, on a branch diverged from
origin/HEAD and passing no observation, now runs sequentially. It still
runs parallel when HEAD matches origin/HEAD. No workflow threads an
observation today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* refactor(#4881): drop the now-unused forkRef message-builder parameter

buildMsgBaserefHeadIgnored stopped reading forkRef when the message was
made mode-neutral, and the parameter was retained with `void forkRef;`
for symmetry with its two sibling builders, which do read it. Symmetry
is not reason enough to keep a dead parameter on a module-private
function with one caller, so drop it (#4921 review).

Behaviour is unchanged; the message text is pinned by an existing
full-string assertion, which is what covers the only real risk here —
transposing the two remaining arguments at the call site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* test(#4881): exercise both FULL_SHA_RE alternatives at their boundaries

The invalid-observation list pinned 39 and 41 hex around the 40-hex
SHA-1 arm but left the 64-hex SHA-256 arm's own +/-1 boundary
unexercised, which the repo's boundary-coverage convention asks for
(#4921 review). Adds 63, 65 and a 64-length non-hex string.

The regex already rejected all three -- this is coverage of correct
behaviour, not a fix -- so it carries no negative control against a
pre-fix base. That it is not vacuous was shown instead by widening the
arm to {63,65}, under which the row fails by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

* docs(#4881): finish the surface sweep the hook interlock owes

Self-found by this round's own pre-push adversarial review, over six passes.
Eight sites, four classes.

FIVE were prose still asserting the stance the interlock overturned, which
reads as live canon to anyone arriving cold: docs/CLI-TOOLS.md,
docs/CONFIGURATION.md and gsd-core/references/planning-config.md each still
said the hook degrade is lifted "unless/until a clean prior harness worktree is
observed"; observeHarnessForkFromHead's own header still said a qualifying
worktree "can only exist if the harness forked from HEAD"; and a test name
still called the observation "required on a hook host" when it is now
inadmissible there. My own sweep had grepped for "restores"/"lifts" and missed
every one -- the ordinary failure of a grep, which returns what you thought to
search for.

The SIXTH is the same class one step worse: gsd-core/workflows/execute-plan.md
still said flatly that Claude Code's isolation="worktree" "forks from
origin/HEAD, not live local HEAD" -- in a paragraph THIS PR already edits, a
few sentences after the clause it corrected. A tombstone makes only its own
line clean; adjoining text asserting the dead stance is the other half of the
same defect. Now qualified on the setting, with the hook exception named.

The SEVENTH is a proof-strength overstatement that predates this PR, with a
driven counterexample: a worktree created from an older base and since `git
checkout --detach`ed onto HEAD is clean, sits at HEAD, and satisfies the probe
identically, so "can only exist" was false. The worktree's own reflog does
retain that original checkout -- the information is not lost, the probe simply
does not consult it.

The EIGHTH is that the header described only one of the function's two legs. A
cache hit returns the prior conclusion without reading any worktree, so "the
probe reads a worktree's present state" was true of the live probe and false of
the cache. The header now separates them, and names the cache's blindness as a
third reason the observation is inadmissible under a hook.

Gaps seven and eight are inherited from #4868 and accepted there for the
no-hook case. Nothing about the mechanism changes here; only what the header
claims for it.

No behavioural change -- comments, prose, and one test's registered name.

Emitted-Drift-Ack-Growth: execute-plan.md — the Pattern A paragraph gained a qualifying
 clause: it stated flatly that Claude Code forks from origin/HEAD, which is the premise
 this PR retires, a few sentences after the clause the PR had already corrected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NuT9wTyrMaxjAPevqeH4YZ

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-23 20:15:28 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%