* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
GSD Core 문서
문서는 네 가지 유형으로 구성됩니다. 튜토리얼은 직접 해보며 배우고, how-to 가이드는 특정 작업을 해결하며, 레퍼런스는 권위 있는 사실을 제시하고, 설명은 개념과 설계 결정을 탐구합니다.
언어 버전: English · Português (pt-BR) · 日本語 · 简体中文 · 한국어
튜토리얼
- 첫 번째 프로젝트 — 설치부터 첫 단계 출시까지, 확실한 한 가지 경로
- 기존 코드베이스 온보딩 — 기존 저장소에 GSD Core 적용하기
How-to guides
- 런타임에 설치하기 — 지원하는 15개 런타임 각각의 설치 단계
- 단계 논의하기 — 기획 시작 전 구현 결정 사항 정리
- 단계 기획하기 — 리서치 실행, 작업 분해, 플랜 품질 검증
- 단계 실행하기 — 새 컨텍스트 서브에이전트로 병렬 웨이브 실행
- 검증 및 출시 — 완료된 작업 검토, 오류 진단, PR 생성
- 단계 자율 실행하기 — 무인 단계 실행을 위한 자율 모드 사용
- 빠른 임시 작업 처리 — 단계 루프 외 임시 작업에
/gsd-quick과/gsd-fast활용 - 모델 프로필 설정 — 고품질, 균형, 예산 모델 티어 전환
- 크로스 AI 리뷰 설정 — 주 에이전트가 생성한 코드를 두 번째 AI가 검토하도록 설정
- 워크스트림으로 병렬 작업 — 워크스트림을 사용해 독립적인 작업 라인 동시 실행
- 워크스페이스로 작업 격리 — 워크스페이스로 실험적이거나 위험한 변경 사항 샌드박스 처리
- 실패한 실행 디버깅 — 깨지거나 불완전한 단계 실행 진단 및 복구
- 스파이크와 스케치 — 플랜 확정 전 탐색 작업에
/gsd-spike와/gsd-sketch활용 - UI 단계 설계 — 프론트엔드 및 시각적 작업에 UI 단계 루프 활용
- 트래커 이슈로 GSD 구동 — GitHub, Linear, Jira 이슈에서 단계 시작
- GSD 2에서 마이그레이션 — 기존 GSD 2 프로젝트를 GSD Core로 업그레이드
- GSD 업데이트 — 설치 프로그램을 재실행해 최신 릴리스 적용
- 복구 및 문제 해결 — 일반적인 문제 해결, 컨텍스트 재구축, 제거
레퍼런스
- 명령어 — 플래그와 예제가 포함된 모든 명령어
- 설정 — 전체 설정 스키마, 모델 프로필, git 브랜칭 전략
- CLI 도구 — 워크플로우와 에이전트를 위한
gsd-tools.cjs프로그래밍 API - 기능 — 전체 기능 색인
- 인벤토리 — 설치된 스킬과 서피스 맵
- STATE.md 스키마 —
.planning/STATE.md필드별 레퍼런스 - CONTEXT.md 스키마 —
.planning/phases/<N>/CONTEXT.md필드별 레퍼런스 - PLAN.md 스키마 —
.planning/phases/<N>/PLAN.md필드별 레퍼런스 - 기획 아티팩트 — 모든
.planning/파일과 역할
설명
- 컨텍스트 엔지니어링 — 컨텍스트 rot가 형성되는 방식과 GSD Core의 방지 방법
- 단계 루프 — 논의 → 기획 → 실행 → 검증 → 출시 사이클의 설계 근거
- 멀티 에이전트 오케스트레이션 — 서브에이전트의 생성, 범위 지정, 조율 방식
- 보안 모델 — 신뢰 경계, 권한, 안전한 자동화
- 아키텍처 — 시스템 아키텍처, 에이전트 모델, 데이터 흐름
- 논의 모드 —
/gsd-discuss-phase의 가정 모드와 인터뷰 모드 - 컨텍스트 모니터링 — 컨텍스트 창 모니터링 훅 아키텍처
- 이슈 기반 오케스트레이션 — 기존 프리미티브를 사용해 트래커 이슈로 GSD를 구동하는 레시피