* test(#4014): add failing-first coverage for unreadable-vs-empty directory scope (epic #3473 B4) * fix(#4014): an unreadable directory must not report as an empty one (epic #3473 B4) * test(#4014): update hardcoded generateSlugInternal closing-brace line after import shift src/core-utils.cts's new #4014 import block shifted every subsequent line by 6, moving generateSlugInternal's real closing brace from line 193 to 199. tests/slug-derivation-drift-guard.test.cjs's MAJOR-1 fixture hardcodes that line number to plant a synthetic violation immediately after the function's real body; the guard script itself locates the boundary dynamically via brace-matching and needed no change. * docs(#4014): document the unreadable-directory scope signal and add changeset * docs(#4014): backfill changeset PR number to #4163 * test(#4014): kill pre-existing core-utils.cjs mutation-score gap, unrelated to this issue's diff --------- Co-authored-by: sim <sim@local>
256 KiB
GSD Feature Reference
Feature index and reference for GSD Core. For architecture details, see Architecture. For command syntax, see Command Reference. Return to docs index.
Table of Contents
- Core Features
- Planning Features
- Quality Assurance Features
- Context Engineering Features
- Brownfield Features
- Utility Features
- Infrastructure Features
- v1.27 Features
- v1.28 Features
- v1.29 Features
- v1.31 Features
- v1.32 Features
- STATE.md Consistency Gates
- Autonomous
--to NFlag - Research Gate
- Verifier Milestone Scope Filtering
- Read-Before-Edit Guard Hook
- Context Reduction
- Discuss-Phase
--powerFlag - Debug
--diagnoseFlag - Phase Dependency Analysis
- Anti-Pattern Severity Levels
- Methodology Artifact Type
- Planner Reachability Check
- Playwright-MCP UI Verification
- Pause-Work Expansion
- Response Language Config
- Manual Update Procedure
- New Runtime Support (Trae, Cline, Augment Code)
- Autonomous
--interactiveFlag - Commit-Docs Guard Hook
- Community Hooks Opt-In
- v1.34.0 Features
- Global Learnings Store
- Queryable Codebase Intelligence
- Execution Context Profiles
- Gates Taxonomy
- Code Review Pipeline
- Socratic Exploration
- Safe Undo
- Plan Import
- Rapid Codebase Scan
- Autonomous Audit-to-Fix
- Improved Prompt Injection Scanner
- Stall Detection in Plan-Phase
- Hard Stop Safety Gates in /gsd-progress --next
- Adaptive Model Preset
- Post-Merge Hunk Verification
- v1.35.0 Features
- v1.36.0 Features
- v1.37.0 Features
- v1.40.0 Features
- v1.41.0 Features
- v1.42.1 Features
- Package Legitimacy Gate
- Skill Surface Budgeting
- Installer Migrations
- Custom Ship PR Body Sections
- Review Default Reviewers
- Fallow Structural Review Pre-Pass
- End-of-Phase Human Verification Mode
- Quota and Rate-Limit Failure Classification
- Statusline Context Position
- Milestone Tag Creation Toggle
- Structured JSON Error Mode
- UAT-Passed Predicate
- Spec-Phase Edge-Completeness Probe
- v1.43.0 Features
- v1.7.0 Features
- Embeddable Orchestration System (Host-Integration Interface)
- Discoverability Registries
- Companion MCP Server
- Statusline Token Count & Git Segment
- Model Catalog Advances
- Claude Orchestration Capability (BETA)
- External-Job Capability
- API-Coverage Gate
- State Rebuild & Configurable Graph Path
- Broken-Windows Ledger
- Complexity-Triggered Refactor
- Archive Quick Tasks at Milestone Close
- Verify-Command Path Grounding
- Statusline STATE.md Freshness Marker
- Read-Only Planning Snapshot (
planning inspect) - Live-DOM UAT Capability
- Opt-In Parallel Reviewer Lanes
- Machine-Readable State Contract (
.planning/state.json) - Stated Failing Direction
- Runtime Identity
- Context Drift Gate
- "Failure Is a Value" — Strict Argv Rejection and the
--pickAbsence Contract - No Silent Swallow, No Verdict From Dropped Data
- Runtime Marker Resolution, Derived Codex Sandbox, and In-Phase Short-Form Dependencies
- The Raw Terminator Is Banned by Construction
- Hooks Declare Their Crash Policy
- gsd-tools Declares Outcomes, Pinned at v1
- Reachable Lint Rules and a Non-Destructive Quick-Task Append
- Per-Task External-Tracker Content-Resolution Seam
- Unreadable-Directory Scope Signal
Core Features
1. Project Initialization
Command: /gsd-new-project [--auto @file.md]
Purpose: Transform a user's idea into a fully structured project with research, scoped requirements, and a phased roadmap.
Requirements:
- REQ-INIT-01: System MUST conduct adaptive questioning until project scope is fully understood
- REQ-INIT-02: System MUST spawn parallel research agents to investigate the domain ecosystem
- REQ-INIT-03: System MUST extract requirements into v1 (must-have), v2 (future), and out-of-scope categories
- REQ-INIT-04: System MUST generate a phased roadmap with requirement traceability
- REQ-INIT-05: System MUST require user approval of the roadmap before proceeding
- REQ-INIT-06: System MUST prevent re-initialization when
.planning/PROJECT.mdalready exists - REQ-INIT-07: System MUST support
--auto @file.mdflag to skip interactive questions and extract from a document
Produces:
| Artifact | Description |
|---|---|
PROJECT.md |
Project vision, constraints, technical decisions, evolution rules |
REQUIREMENTS.md |
Scoped requirements with unique IDs (REQ-XX) |
ROADMAP.md |
Phase breakdown with status tracking and requirement mapping |
STATE.md |
Initial project state with position, decisions, metrics |
config.json |
Workflow configuration |
research/SUMMARY.md |
Synthesized domain research |
research/STACK.md |
Technology stack investigation |
research/FEATURES.md |
Feature implementation patterns |
research/ARCHITECTURE.md |
Architecture patterns and trade-offs |
research/PITFALLS.md |
Common failure modes and mitigations |
Process:
- Questions — Adaptive questioning guided by the "dream extraction" philosophy (not requirements gathering)
- Research — 4 parallel researcher agents investigate stack, features, architecture, and pitfalls
- Synthesis — Research synthesizer combines findings into SUMMARY.md
- Requirements — Extracted from user responses + research, categorized by scope
- Roadmap — Phase breakdown mapped to requirements, with granularity setting controlling phase count
Functional Requirements:
- Questions adapt based on detected project type (web app, CLI, mobile, API, etc.)
- Research agents have web search capability for current ecosystem information
- Granularity setting controls phase count:
coarse(2-4),standard(4-6),fine(6-10) --automode extracts all information from the provided document without interactive questioning- Existing codebase context (from
/gsd-map-codebase) is loaded if present
2. Phase Discussion
Command: /gsd-discuss-phase [N] [--auto] [--batch]
Purpose: Capture user's implementation preferences and decisions before research and planning begin. Eliminates the gray areas that cause AI to guess.
Requirements:
- REQ-DISC-01: System MUST analyze the phase scope and identify decision areas (gray areas)
- REQ-DISC-02: System MUST categorize gray areas by type (visual, API, content, organization, etc.)
- REQ-DISC-03: System MUST ask only questions not already answered in prior CONTEXT.md files
- REQ-DISC-04: System MUST persist decisions in
{phase}-CONTEXT.mdwith canonical references - REQ-DISC-05: System MUST support
--autoflag to auto-select recommended defaults - REQ-DISC-06: System MUST support
--batchflag for grouped question intake - REQ-DISC-07: System MUST scout relevant source files before identifying gray areas (code-aware discussion)
- REQ-DISC-08: System MUST adapt gray area language to product-outcome terms when USER-PROFILE.md indicates a non-technical owner (learning_style: guided, jargon in frustration_triggers, or high-level explanation depth)
- REQ-DISC-09: When REQ-DISC-08 applies, advisor_research rationale paragraphs MUST be rewritten in plain language — same decisions, translated framing
Produces: {padded_phase}-CONTEXT.md — User preferences that feed into research and planning
Gray Area Categories:
| Category | Example Decisions |
|---|---|
| Visual features | Layout, density, interactions, empty states |
| APIs/CLIs | Response format, flags, error handling, verbosity |
| Content systems | Structure, tone, depth, flow |
| Organization | Grouping criteria, naming, duplicates, exceptions |
3. UI Design Contract
Command: /gsd-ui-phase [N]
Purpose: Lock design decisions before planning so that all components in a phase share consistent visual standards.
Requirements:
- REQ-UI-01: System MUST detect existing design system state (shadcn components.json, Tailwind config, tokens)
- REQ-UI-02: System MUST ask only unanswered design contract questions
- REQ-UI-03: System MUST validate against 7 dimensions (Copywriting, Visuals, Color, Typography, Spacing, Registry Safety, Inventory Provenance)
- REQ-UI-04: System MUST enter revision loop if validation returns BLOCKED (max 2 iterations)
- REQ-UI-05: System MUST offer shadcn initialization for React/Next.js/Vite projects without
components.json - REQ-UI-06: System MUST enforce registry safety gate for third-party shadcn registries
Produces: {padded_phase}-UI-SPEC.md — Design contract consumed by executors
7 Validation Dimensions:
- Copywriting — CTA labels, empty states, error messages
- Visuals — Focal points, visual hierarchy, icon accessibility
- Color — Accent usage discipline, 60/30/10 compliance
- Typography — Font size/weight constraint adherence
- Spacing — Grid alignment, token consistency
- Registry Safety — Third-party component inspection requirements
- Inventory Provenance — Component inventory enumerated from the installed design system, not recalled
shadcn Integration:
- Detects missing
components.jsonin React/Next.js/Vite projects - Guides user through
ui.shadcn.com/createpreset configuration - Preset string becomes a planning artifact reproducible across phases
- Safety gate requires
npx shadcn viewandnpx shadcn diffbefore third-party components
4. Phase Planning
Command: /gsd-plan-phase [N] [--auto] [--skip-research] [--skip-verify]
Purpose: Research the implementation domain and produce verified, atomic execution plans.
Requirements:
- REQ-PLAN-01: System MUST spawn a phase researcher to investigate implementation approaches
- REQ-PLAN-02: System MUST produce plans with 2-3 tasks each, sized for a single context window
- REQ-PLAN-03: System MUST structure plans as XML with
<task>elements containingname,files,action,verify, anddonefields - REQ-PLAN-04: System MUST include
read_firstandacceptance_criteriasections in every plan - REQ-PLAN-05: System MUST run plan checker verification loop (up to 3 iterations) unless
--skip-verifyis set - REQ-PLAN-06: System MUST support
--skip-researchflag to bypass research phase - REQ-PLAN-07: System MUST prompt user to run
/gsd-ui-phaseif frontend phase detected and no UI-SPEC.md exists (UI safety gate) - REQ-PLAN-08: System MUST include Nyquist validation mapping when
workflow.nyquist_validationis enabled - REQ-PLAN-09: System MUST verify all phase requirements are covered by at least one plan before planning completes (requirements coverage gate)
- REQ-PLAN-10: System MUST support an optional
<reversibility rating="reversible|costly|one-way">element recording how costly a decision would be to undo, and MUST insert acheckpoint:decisionbefore the task implementing aone-waydecision unless--no-reversibility-gatesis set (costlyis flagged without blocking;reversibleand unrated flow normally)
Produces:
| Artifact | Description |
|---|---|
{phase}-RESEARCH.md |
Ecosystem research findings |
{phase}-{N}-PLAN.md |
Atomic execution plans (2-3 tasks each) |
{phase}-VALIDATION.md |
Test coverage mapping (Nyquist layer) |
Plan Structure (XML):
<task type="auto">
<name>Create login endpoint</name>
<files>src/app/api/auth/login/route.ts</files>
<action>
Use jose for JWT. Validate credentials against users table.
Return httpOnly cookie on success.
</action>
<verify>curl -X POST localhost:3000/api/auth/login returns 200 + Set-Cookie</verify>
<done>Valid credentials return cookie, invalid return 401</done>
</task>
Plan Checker Verification (8 Dimensions):
- Requirement coverage — Plans address all phase requirements
- Task atomicity — Each task is independently committable
- Dependency ordering — Tasks sequence correctly
- File scope — No excessive file overlap between plans
- Verification commands — Each task has testable done criteria
- Context fit — Tasks fit within a single context window
- Gap detection — No missing implementation steps
- Nyquist compliance — Tasks have automated verify commands (when enabled)
5. Phase Execution
Command: /gsd-execute-phase <N>
Purpose: Execute all plans in a phase using wave-based parallelization with fresh context windows per executor.
Requirements:
- REQ-EXEC-01: System MUST analyze plan dependencies and group into execution waves
- REQ-EXEC-02: System MUST spawn independent plans in parallel within each wave
- REQ-EXEC-03: System MUST give each executor a fresh context window (200K tokens)
- REQ-EXEC-04: System MUST produce atomic git commits per task
- REQ-EXEC-05: System MUST produce a SUMMARY.md for each completed plan
- REQ-EXEC-06: System MUST run post-execution verifier to check phase goals were met
- REQ-EXEC-07: System MUST support git branching strategies (
none,phase,milestone) - REQ-EXEC-08: System MUST invoke node repair operator on task verification failure (when enabled)
- REQ-EXEC-09: System MUST run prior phases' test suites before verification to catch cross-phase regressions
Produces:
| Artifact | Description |
|---|---|
{phase}-{N}-SUMMARY.md |
Execution outcomes per plan |
{phase}-VERIFICATION.md |
Post-execution verification report |
| Git commits | Atomic commits per task |
Wave Execution:
- Plans with no dependencies → Wave 1 (parallel)
- Plans depending on Wave 1 → Wave 2 (parallel, waits for Wave 1)
- Continues until all plans complete
- File conflicts force sequential execution within same wave
Executor Capabilities:
- Reads PLAN.md with full task instructions
- Has access to PROJECT.md, STATE.md, CONTEXT.md, RESEARCH.md
- Commits each task atomically with structured commit messages
- Uses
--no-verifyon commits during parallel execution to avoid build lock contention - Handles checkpoint types:
auto,checkpoint:human-verify,checkpoint:decision,checkpoint:human-action - Reports deviations from plan in SUMMARY.md
Parallel Safety:
- Pre-commit hooks: Skipped by parallel agents (
--no-verify), run once by orchestrator after each wave - STATE.md locking: File-level lockfile prevents concurrent write corruption across agents
6. Work Verification
Command: /gsd-verify-work [N]
Purpose: User acceptance testing — walk the user through testing each deliverable and auto-diagnose failures.
Requirements:
- REQ-VERIFY-01: System MUST extract testable deliverables from the phase
- REQ-VERIFY-02: System MUST present deliverables one at a time for user confirmation
- REQ-VERIFY-03: System MUST spawn debug agents to diagnose failures automatically
- REQ-VERIFY-04: System MUST create fix plans for identified issues
- REQ-VERIFY-05: System MUST inject cold-start smoke test for phases modifying server/database/seed/startup files
- REQ-VERIFY-06: System MUST produce UAT.md with pass/fail results
Produces: {phase}-UAT.md — User acceptance test results, plus fix plans if issues found
6.5. Ship
Command: /gsd-ship [N] [--draft]
Purpose: Bridge local completion → merged PR. After verification passes, push branch, create PR with auto-generated body from planning artifacts, optionally trigger review, and track in STATE.md.
Requirements:
- REQ-SHIP-01: System MUST verify phase has passed verification before shipping
- REQ-SHIP-02: System MUST push branch and create PR via
ghCLI - REQ-SHIP-03: System MUST auto-generate PR body from SUMMARY.md, VERIFICATION.md, and REQUIREMENTS.md
- REQ-SHIP-04: System MUST update STATE.md with shipping status and PR number
- REQ-SHIP-05: System MUST support
--draftflag for draft PRs - REQ-SHIP-06: System MUST support append-only project PR body sections configured with
ship.pr_body_sections
Prerequisites: Phase verified, gh CLI installed and authenticated, work on feature branch
Produces: GitHub PR with rich body, optional configured PRD-style sections, STATE.md updated
User documentation: Custom PR Body Sections
7. UI Review
Command: /gsd-ui-review [N]
Purpose: Retroactive 6-pillar visual audit of implemented frontend code. Works standalone on any project.
Requirements:
- REQ-UIREVIEW-01: System MUST score each of the 6 pillars on a 1-4 scale
- REQ-UIREVIEW-02: System MUST capture screenshots via Playwright CLI to
.planning/ui-reviews/ - REQ-UIREVIEW-03: System MUST create
.gitignorefor screenshot directory - REQ-UIREVIEW-04: System MUST identify top 3 priority fixes
- REQ-UIREVIEW-05: System MUST work standalone (without UI-SPEC.md) using abstract quality standards
6 Audit Pillars (scored 1-4):
- Copywriting — CTA labels, empty states, error states
- Visuals — Focal points, visual hierarchy, icon accessibility
- Color — Accent usage discipline, 60/30/10 compliance
- Typography — Font size/weight constraint adherence
- Spacing — Grid alignment, token consistency
- Experience Design — Loading/error/empty state coverage
Produces: {padded_phase}-UI-REVIEW.md — Scores and prioritized fixes
8. Milestone Management
Commands: /gsd-audit-milestone, /gsd-complete-milestone, /gsd-new-milestone [name]
Purpose: Verify milestone completion, archive, tag release, and start the next development cycle.
Requirements:
- REQ-MILE-01: Audit MUST verify all milestone requirements are met
- REQ-MILE-02: Audit MUST detect stubs, placeholder implementations, and untested code
- REQ-MILE-03: Audit MUST check Nyquist validation compliance across phases
- REQ-MILE-04: Complete MUST archive milestone data to MILESTONES.md
- REQ-MILE-05: Complete MUST offer git tag creation for the release
- REQ-MILE-06: Complete MUST offer squash merge or merge with history for branching strategies
- REQ-MILE-07: Complete MUST clean up UI review screenshots
- REQ-MILE-08: New milestone MUST follow same flow as new-project (questions → research → requirements → roadmap)
- REQ-MILE-09: New milestone MUST NOT reset existing workflow configuration
Planning Features
9. Phase Management
Commands: /gsd-phase, /gsd-phase --insert [N], /gsd-phase --remove [N]
Purpose: Dynamic roadmap modification during development.
Requirements:
- REQ-PHASE-01: Add MUST append a new phase to the end of the current roadmap
- REQ-PHASE-02: Insert MUST use decimal numbering (e.g., 3.1) between existing phases
- REQ-PHASE-03: Remove MUST renumber all subsequent phases
- REQ-PHASE-04: Remove MUST prevent removing phases that have been executed
- REQ-PHASE-05: All operations MUST update ROADMAP.md and create/remove phase directories
- REQ-PHASE-06: Bare-number phase lookup MUST resolve digit-leading slug names consistently across phase verbs, preserve project-code-prefixed result shaping, and fail loudly when multiple directories match
10. Quick Mode
Command: /gsd-quick [--full] [--discuss] [--research]
Purpose: Ad-hoc task execution with GSD guarantees but a faster path.
Requirements:
- REQ-QUICK-01: System MUST accept freeform task description
- REQ-QUICK-02: System MUST use same planner + executor agents as full workflow
- REQ-QUICK-03: System MUST skip research, plan checker, and verifier by default
- REQ-QUICK-04:
--fullflag MUST enable plan checking (max 2 iterations) and post-execution verification - REQ-QUICK-05:
--discussflag MUST run lightweight pre-planning discussion - REQ-QUICK-06:
--researchflag MUST spawn focused research agent before planning - REQ-QUICK-07: Flags MUST be composable (
--discuss --research --full) - REQ-QUICK-08: System MUST track quick tasks in
.planning/quick/YYMMDD-xxx-slug/ - REQ-QUICK-09: System MUST produce atomic commits for quick task execution
11. Autonomous Mode
Command: /gsd-autonomous [--from N]
Purpose: Run all remaining phases autonomously — discuss → plan → execute per phase.
Requirements:
- REQ-AUTO-01: System MUST iterate through all incomplete phases in roadmap order
- REQ-AUTO-02: System MUST run discuss → plan → execute for each phase
- REQ-AUTO-03: System MUST pause for explicit user decisions (gray area acceptance, blockers, validation)
- REQ-AUTO-04: System MUST re-read ROADMAP.md after each phase to catch dynamically inserted phases
- REQ-AUTO-05:
--from Nflag MUST start from a specific phase number
12. Freeform Routing
Command: /gsd-progress --do (see also /gsd-manager for interactive routing)
Purpose: Analyze freeform text and route to the appropriate GSD command.
Requirements:
- REQ-DO-01: System MUST parse user intent from natural language input
- REQ-DO-02: System MUST map intent to the best matching GSD command
- REQ-DO-03: System MUST confirm the routing with the user before executing
- REQ-DO-04: System MUST handle project-exists vs no-project contexts differently
13. Note Capture
Command: /gsd-capture
Purpose: Zero-friction idea capture without interrupting workflow. Append timestamped notes, list all notes, or promote notes to structured todos.
Requirements:
- REQ-NOTE-01: System MUST save timestamped note files with a single Write call
- REQ-NOTE-02: System MUST support
listsubcommand to show all notes from project and global scopes - REQ-NOTE-03: System MUST support
promote Nsubcommand to convert a note into a structured todo - REQ-NOTE-04: System MUST support
--globalflag for global scope operations - REQ-NOTE-05: System MUST NOT use Task, AskUserQuestion, or Bash — runs inline only
14. Auto-Advance (Next)
Command: /gsd-progress --next
Purpose: Automatically detect current project state and advance to the next logical workflow step, eliminating the need to remember which phase/step you're on.
Requirements:
- REQ-NEXT-01: System MUST read STATE.md, ROADMAP.md, and phase directories to determine current position
- REQ-NEXT-02: System MUST detect whether discuss, plan, execute, or verify is needed
- REQ-NEXT-03: System MUST invoke the correct command automatically
- REQ-NEXT-04: System MUST suggest
/gsd-new-projectif no project exists - REQ-NEXT-05: System MUST suggest
/gsd-complete-milestonewhen all phases are complete
State Detection Logic:
| State | Action |
|---|---|
No .planning/ directory |
Suggest /gsd-new-project |
| Phase has no CONTEXT.md | Run /gsd-discuss-phase |
| Phase has no PLAN.md files | Run /gsd-plan-phase |
| Phase has plans but no SUMMARY.md | Run /gsd-execute-phase |
| Phase executed but no VERIFICATION.md | Run /gsd-verify-work |
| All phases complete | Suggest /gsd-complete-milestone |
Quality Assurance Features
15. Nyquist Validation
Purpose: Map automated test coverage to phase requirements before any code is written. Named after the Nyquist sampling theorem — ensures a feedback signal exists for every requirement.
Requirements:
- REQ-NYQ-01: System MUST detect existing test infrastructure during plan-phase research
- REQ-NYQ-02: System MUST map each requirement to a specific test command
- REQ-NYQ-03: System MUST identify Wave 0 tasks (test scaffolding needed before implementation)
- REQ-NYQ-04: Plan checker MUST enforce Nyquist compliance as 8th verification dimension
- REQ-NYQ-05: System MUST support retroactive validation via
/gsd-validate-phase - REQ-NYQ-06: System MUST be disableable via
workflow.nyquist_validation: false
Produces: {phase}-VALIDATION.md — Test coverage contract
Retroactive Validation (/gsd-validate-phase [N]):
- Scans implementation and maps requirements to tests
- Identifies gaps where requirements lack automated verification
- Spawns auditor to generate tests (max 3 attempts)
- Never modifies implementation code — only test files and VALIDATION.md
- Flags implementation bugs as escalations for user to address
16. Plan Checking
Purpose: Goal-backward verification that plans will achieve phase objectives before execution.
Requirements:
- REQ-PLANCK-01: System MUST verify plans against 8 quality dimensions
- REQ-PLANCK-02: System MUST loop up to 3 iterations until plans pass
- REQ-PLANCK-03: System MUST produce specific, actionable feedback on failures
- REQ-PLANCK-04: System MUST be disableable via
workflow.plan_check: false
17. Post-Execution Verification
Purpose: Automated check that the codebase delivers what the phase promised.
Requirements:
- REQ-POSTVER-01: System MUST check against phase goals, not just task completion
- REQ-POSTVER-02: System MUST produce VERIFICATION.md with pass/fail analysis
- REQ-POSTVER-03: System MUST log issues for
/gsd-verify-workto address - REQ-POSTVER-04: System MUST be disableable via
workflow.verifier: false
18. Node Repair
Purpose: Autonomous recovery when task verification fails during execution.
Requirements:
- REQ-REPAIR-01: System MUST analyze failure and choose one strategy: RETRY, DECOMPOSE, or PRUNE
- REQ-REPAIR-02: RETRY MUST attempt with a concrete adjustment
- REQ-REPAIR-03: DECOMPOSE MUST break task into smaller verifiable sub-steps
- REQ-REPAIR-04: PRUNE MUST remove unachievable tasks and escalate to user
- REQ-REPAIR-05: System MUST respect repair budget (default: 2 attempts per task)
- REQ-REPAIR-06: System MUST be configurable via
workflow.node_repair_budgetandworkflow.node_repair
19. Health Validation
Command: /gsd-health [--repair] [--backfill]
Purpose: Validate .planning/ directory integrity and auto-repair issues.
Requirements:
- REQ-HEALTH-01: System MUST check for missing required files
- REQ-HEALTH-02: System MUST validate configuration consistency
- REQ-HEALTH-03: System MUST detect orphaned plans without summaries
- REQ-HEALTH-04: System MUST check phase numbering and roadmap sync
- REQ-HEALTH-05:
--repairflag MUST auto-fix recoverable issues except DESTRUCTIVE-risk ones, which it MUST report but never auto-apply - REQ-HEALTH-06:
--backfillflag MUST synthesize missing MILESTONES.md entries from archived milestone snapshots
20. Cross-Phase Regression Gate
Purpose: Prevent regressions from compounding across phases by running prior phases' test suites after execution.
Requirements:
- REQ-REGR-01: System MUST run test suites from all completed prior phases after phase execution
- REQ-REGR-02: System MUST report any test failures as cross-phase regressions
- REQ-REGR-03: Regressions MUST be surfaced before post-execution verification
- REQ-REGR-04: System MUST identify which prior phase's tests were broken
When: Runs automatically during /gsd-execute-phase before the verifier step.
21. Requirements Coverage Gate
Purpose: Ensure all phase requirements are covered by at least one plan before planning completes.
Requirements:
- REQ-COVGATE-01: System MUST extract all requirement IDs assigned to the phase from ROADMAP.md
- REQ-COVGATE-02: System MUST verify each requirement appears in at least one PLAN.md
- REQ-COVGATE-03: Uncovered requirements MUST block planning completion
- REQ-COVGATE-04: System MUST report which specific requirements lack plan coverage
When: Runs automatically at the end of /gsd-plan-phase after the plan checker loop.
Context Engineering Features
22. Context Window Monitoring
Purpose: Prevent context rot by alerting both user and agent when context is running low.
Requirements:
- REQ-CTX-01: Statusline MUST display context usage percentage to user
- REQ-CTX-02: Context monitor MUST inject agent-facing warnings at ≤35% remaining (WARNING)
- REQ-CTX-03: Context monitor MUST inject agent-facing warnings at ≤25% remaining (CRITICAL)
- REQ-CTX-04: Warnings MUST debounce (5 tool uses between repeated warnings)
- REQ-CTX-05: Severity escalation (WARNING→CRITICAL) MUST bypass debounce
- REQ-CTX-06: Context monitor MUST differentiate GSD-active vs non-GSD-active projects
- REQ-CTX-07: Warnings MUST be advisory, never imperative commands that override user preferences
- REQ-CTX-08: All hooks MUST fail silently and never block tool execution
Architecture: Two-part bridge system:
- Statusline writes metrics to
/tmp/claude-ctx-{session}.json - Context monitor reads metrics and injects
additionalContextwarnings
23. Session Management
Commands: /gsd-pause-work, /gsd-resume-work, /gsd-progress
Purpose: Maintain project continuity across context resets and sessions.
Requirements:
- REQ-SESSION-01: Pause MUST save current position and next steps to
continue-here.mdand structuredHANDOFF.json - REQ-SESSION-02: Resume MUST restore full project context from HANDOFF.json (preferred) or state files (fallback)
- REQ-SESSION-03: Progress MUST show current position, next action, and overall completion
- REQ-SESSION-04: Progress MUST read all state files (STATE.md, ROADMAP.md, phase directories)
- REQ-SESSION-05: All session operations MUST work after
/clear(context reset) - REQ-SESSION-06: HANDOFF.json MUST include blockers, human actions pending, and in-progress task state
- REQ-SESSION-07: Resume MUST surface human actions and blockers immediately on session start
24. Session Reporting
Command: /gsd-pause-work --report
Purpose: Generate a structured post-session summary document capturing work performed, outcomes achieved, and estimated resource usage.
Requirements:
- REQ-REPORT-01: System MUST gather data from STATE.md, git log, and plan/summary files
- REQ-REPORT-02: System MUST include commits made, plans executed, and phases progressed
- REQ-REPORT-03: System MUST estimate token usage and cost based on session activity
- REQ-REPORT-04: System MUST include active blockers and decisions made
- REQ-REPORT-05: System MUST recommend next steps
Produces: .planning/reports/SESSION_REPORT.md
Report Sections:
- Session overview (duration, milestone, phase)
- Work performed (commits, plans, phases)
- Outcomes and deliverables
- Blockers and decisions
- Resource estimates (tokens, cost)
- Next steps recommendation
25. Multi-Agent Orchestration
Purpose: Coordinate specialized agents with fresh context windows for each task.
Requirements:
- REQ-ORCH-01: Each agent MUST receive a fresh context window
- REQ-ORCH-02: Orchestrators MUST be thin — spawn agents, collect results, route next
- REQ-ORCH-03: Context payload MUST include all relevant project artifacts
- REQ-ORCH-04: Parallel agents MUST be truly independent (no shared mutable state)
- REQ-ORCH-05: Agent results MUST be written to disk before orchestrator processes them
- REQ-ORCH-06: Failed agents MUST be detected (spot-check actual output vs reported failure)
26. Model Profiles
Command: /gsd-config --profile <quality|balanced|budget|adaptive|inherit>
Purpose: Control which AI model each agent uses, balancing quality vs cost.
Requirements:
- REQ-MODEL-01: System MUST support 4 profiles:
quality,balanced,budget,inherit - REQ-MODEL-02: Each profile MUST define model tier per agent (see profile table)
- REQ-MODEL-03: Per-agent overrides MUST take precedence over profile
- REQ-MODEL-04:
inheritprofile MUST defer to runtime's current model selection - REQ-MODEL-04a:
inheritprofile MUST be used when running non-Anthropic providers (OpenRouter, local models) to avoid unexpected API costs - REQ-MODEL-05: Profile switch MUST be programmatic (script, not LLM-driven)
- REQ-MODEL-06: Model resolution MUST happen once per orchestration, not per spawn
Profile Assignments:
| Agent | quality |
balanced |
budget |
inherit |
|---|---|---|---|---|
| gsd-planner | Opus | Opus | Sonnet | Inherit |
| gsd-roadmapper | Opus | Sonnet | Sonnet | Inherit |
| gsd-executor | Opus | Sonnet | Sonnet | Inherit |
| gsd-phase-researcher | Opus | Sonnet | Haiku | Inherit |
| gsd-project-researcher | Opus | Sonnet | Haiku | Inherit |
| gsd-research-synthesizer | Sonnet | Sonnet | Haiku | Inherit |
| gsd-debugger | Opus | Sonnet | Sonnet | Inherit |
| gsd-codebase-mapper | Sonnet | Haiku | Haiku | Inherit |
| gsd-verifier | Sonnet | Sonnet | Haiku | Inherit |
| gsd-plan-checker | Sonnet | Sonnet | Haiku | Inherit |
| gsd-integration-checker | Sonnet | Sonnet | Haiku | Inherit |
| gsd-nyquist-auditor | Sonnet | Sonnet | Haiku | Inherit |
Brownfield Features
27. Codebase Mapping
Command: /gsd-map-codebase [area]
Purpose: Analyze an existing codebase before starting a new project or as the mapping handoff from /gsd-onboard, so GSD understands what exists.
Requirements:
- REQ-MAP-01: System MUST spawn parallel mapper agents for each analysis area
- REQ-MAP-02: System MUST produce structured documents in
.planning/codebase/ - REQ-MAP-03: System MUST detect: tech stack, architecture patterns, coding conventions, concerns
- REQ-MAP-04: Subsequent
/gsd-new-projectMUST load codebase mapping and focus questions on what's being added - REQ-MAP-05: Optional
[area]argument MUST scope mapping to a specific area
Produces:
| Document | Content |
|---|---|
STACK.md |
Languages, frameworks, databases, infrastructure |
ARCHITECTURE.md |
Patterns, layers, data flow, boundaries |
CONVENTIONS.md |
Naming, file organization, code style, testing patterns |
CONCERNS.md |
Technical debt, security issues, performance bottlenecks |
STRUCTURE.md |
Directory layout and file organization |
TESTING.md |
Test infrastructure, coverage, patterns |
INTEGRATIONS.md |
External services, APIs, third-party dependencies |
Incremental remap — --paths (#2003): The mapper accepts an optional
--paths <p1,p2,...> scope hint. When provided, it restricts exploration
to the listed repo-relative prefixes instead of scanning the whole tree.
This is the pathway used by the post-execute codebase-drift gate to refresh
only the subtrees the phase actually changed. Each produced document carries
last_mapped_commit in its YAML frontmatter so drift can be measured
against the mapping point, not HEAD.
27b. Existing Codebase Onboarding
Command: /gsd-onboard [--fast] [--text]
Purpose: Guide first-time setup for an existing repository by checking brownfield state, routing through codebase mapping and docs ingest, then handing off to project initialization without silently overwriting planning artifacts.
Requirements:
- REQ-ONBOARD-01: System MUST detect existing code, package manifests, planning documents, partial
.planning/state, and complete or missing codebase-map files. - REQ-ONBOARD-02: System MUST hand off to
/gsd-map-codebaseor/gsd-map-codebase --fastwhen brownfield code lacks the required.planning/codebase/map files; fast-map readiness is partial and MUST NOT be treated as sufficient for/gsd-new-project. - REQ-ONBOARD-03: System MUST offer
/gsd-ingest-docsbefore/gsd-new-projectwhen ADR/PRD/SPEC/RFC candidates exist and no project exists. - REQ-ONBOARD-04: System MUST refuse to report onboarding complete until
PROJECT.md,REQUIREMENTS.md,ROADMAP.md, andSTATE.mdall exist. - REQ-ONBOARD-05: System MUST create or confirm
.planning/onboarding/SUMMARY.mdonly after project setup exists. - REQ-ONBOARD-06: System MUST support
--textfor numbered plain-text gates on runtimes without interactive menus.
Produces:
| Artifact | Description |
|---|---|
.planning/codebase/ |
Codebase map produced by the /gsd-map-codebase handoff |
.planning/PROJECT.md, REQUIREMENTS.md, ROADMAP.md, STATE.md |
Planning setup produced by /gsd-new-project or /gsd-ingest-docs |
.planning/onboarding/SUMMARY.md |
Onboarding status, artifact index, and next-command summary |
27a. Post-Execute Codebase Drift Detection
Introduced by: #2003
Trigger: Runs automatically at the end of every /gsd-execute-phase
Configuration:
workflow.drift_threshold(integer, default3) — minimum new structural elements before the gate acts.workflow.drift_action(warn|auto-remap, defaultwarn) — warn-only or spawngsd-codebase-mapperwith--pathsscoped to affected subtrees.
What counts as drift:
- New directory outside mapped paths
- New barrel export at
(packages|apps)/*/src/index.* - New migration file (supabase/prisma/drizzle/src/migrations/…)
- New route module under
routes/orapi/
Non-blocking guarantee: any internal failure (missing STRUCTURE.md, git errors, mapper spawn failure) logs a single line and the phase continues. Drift detection cannot fail verification.
Requirements:
- REQ-DRIFT-01: System MUST detect the four drift categories from
git diff --name-status last_mapped_commit..HEAD - REQ-DRIFT-02: Action fires only when element count ≥
workflow.drift_threshold - REQ-DRIFT-03:
warnaction MUST NOT spawn any agent - REQ-DRIFT-04:
auto-remapaction MUST pass sanitized--pathsto the mapper - REQ-DRIFT-05: Detection/remap failure MUST be non-blocking for
/gsd-execute-phase - REQ-DRIFT-06:
last_mapped_commitround-trip through YAML frontmatter on each.planning/codebase/*.mdfile
Utility Features
28. Debug System
Command: /gsd-debug [description]
Purpose: Systematic debugging with persistent state across context resets.
Requirements:
- REQ-DEBUG-01: System MUST create debug session file in
.planning/debug/ - REQ-DEBUG-02: System MUST track hypotheses, evidence, and eliminated theories
- REQ-DEBUG-03: System MUST persist state so debugging survives context resets
- REQ-DEBUG-04: System MUST require human verification before marking resolved
- REQ-DEBUG-05: Resolved sessions MUST append to
.planning/debug/knowledge-base.md - REQ-DEBUG-06: Knowledge base MUST be consulted on new debug sessions to prevent re-investigation
Debug Session States: gathering → investigating → fixing → verifying → awaiting_human_verify → resolved
29. Todo Management
Commands: /gsd-capture [desc], /gsd-capture --list
Purpose: Capture ideas and tasks during sessions for later work.
Requirements:
- REQ-TODO-01: System MUST capture todo from current conversation context
- REQ-TODO-02: Todos MUST be stored in
.planning/todos/pending/ - REQ-TODO-03: Completed todos MUST move to
.planning/todos/completed/ - REQ-TODO-04: Check-todos MUST list all pending items with selection to work on one
30. Statistics Dashboard
Command: /gsd-stats
Purpose: Display project metrics — phases, plans, requirements, git history, and timeline.
Requirements:
- REQ-STATS-01: System MUST show phase/plan completion counts
- REQ-STATS-02: System MUST show requirement coverage
- REQ-STATS-03: System MUST show git commit metrics
- REQ-STATS-04: System MUST support multiple output formats (json, table, bar)
31. Update System
Command: /gsd-update
Purpose: Update GSD to the latest version with changelog preview.
Requirements:
- REQ-UPDATE-01: System MUST check for new versions via npm
- REQ-UPDATE-02: System MUST display changelog for new version before updating
- REQ-UPDATE-03: System MUST be runtime-aware and target the correct directory
- REQ-UPDATE-04: System MUST back up locally modified files to
gsd-local-patches/ - REQ-UPDATE-05:
/gsd-update --reapplyMUST restore local modifications after update - REQ-UPDATE-06:
/gsd-update --next(alias--rc) MUST target the@nextRC dist-tag for version check and install; omitting the flag MUST keep@latestbehavior unchanged (ADR #660) - REQ-UPDATE-07: System MUST back up user-added files found inside GSD-managed directories to
gsd-user-files-backup/before the clean install - REQ-UPDATE-08: When that backup is non-empty, the update MUST offer an explicit restore choice before finishing, and MUST leave the backup intact whichever way the user answers
- REQ-UPDATE-09: A restore MUST NOT overwrite a path the newly installed release ships, MUST NOT overwrite a different file already on disk, and MUST report best-effort compatibility warnings for restored files without blocking on them
32. Settings Management
Command: /gsd-settings
Purpose: Interactive configuration of workflow toggles and model profile.
Requirements:
- REQ-SETTINGS-01: System MUST present current settings with toggle options
- REQ-SETTINGS-02: System MUST update
.planning/config.json - REQ-SETTINGS-03: System MUST support saving as global defaults (
~/.gsd/defaults.json)
Configurable Settings:
| Setting | Type | Default | Description |
|---|---|---|---|
mode |
enum | interactive |
interactive or yolo (auto-approve) |
granularity |
enum | standard |
coarse, standard, or fine |
model_profile |
enum | balanced |
quality, balanced, budget, or inherit |
models.<phase_type> |
enum | (none) | Per-phase-type tier override (planning, discuss, research, execution, verification, completion). Values: opus, sonnet, haiku, inherit. Coarse phase-level tuning that wins over model_profile but loses to per-agent model_overrides. See CONFIGURATION.md. Added in v1.40 |
granularities.<phase_type> |
enum | (none) | Per-phase-type granularity override (planning, discuss, research, execution, verification, completion). Values: coarse, standard, fine. Mirrors models.<phase_type> for granularity. See CONFIGURATION.md. Added in v1.43 (#68). /gsd-plan-phase --granularity <coarse|standard|fine> overrides all config-based granularity for a single invocation (takes precedence over granularities.planning, top-level granularity, and planning.granularity). (#703) |
dynamic_routing.enabled |
boolean | false |
Master switch for failure-tier escalation. When true, agents resolve to tier_models[default_tier] and escalate one tier on orchestrator-detected soft failure. Capped by max_escalations. See CONFIGURATION.md. Added in v1.40 |
workflow.research |
boolean | true |
Domain research before planning |
workflow.plan_check |
boolean | true |
Plan verification loop |
workflow.verifier |
boolean | true |
Post-execution verification |
workflow.auto_advance |
boolean | false |
Auto-chain discuss→plan→execute |
workflow.nyquist_validation |
boolean | true |
Nyquist test coverage mapping |
workflow.ui_phase |
boolean | true |
UI design contract generation |
workflow.ui_safety_gate |
boolean | true |
Prompt for ui-phase on frontend phases |
workflow.node_repair |
boolean | true |
Autonomous task repair |
workflow.node_repair_budget |
number | 2 |
Max repair attempts per task |
planning.commit_docs |
boolean | true |
Commit .planning/ files to git |
planning.search_gitignored |
boolean | false |
Include gitignored files in searches |
parallelization.enabled |
boolean | true |
Run independent plans simultaneously |
git.branching_strategy |
enum | none |
none, phase, or milestone |
33. Test Generation
Command: /gsd-add-tests [N]
Purpose: Generate tests for a completed phase based on UAT criteria and implementation.
Requirements:
- REQ-TEST-01: System MUST analyze completed phase implementation
- REQ-TEST-02: System MUST generate tests based on UAT criteria and acceptance criteria
- REQ-TEST-03: System MUST use existing test infrastructure patterns
Infrastructure Features
Looking for a third-party add-on instead? See the GSD Community Capability Registry & EoS Registry — non-endorsing discoverability catalogs for community-contributed Capabilities and EoS host integrations.
34. Git Integration
Purpose: Atomic commits, branching strategies, and clean history management.
Requirements:
- REQ-GIT-01: Each task MUST get its own atomic commit
- REQ-GIT-02: Commit messages MUST follow structured format:
type(scope): description - REQ-GIT-03: System MUST support 3 branching strategies:
none,phase,milestone - REQ-GIT-04: Phase strategy MUST create one branch per phase
- REQ-GIT-05: Milestone strategy MUST create one branch per milestone
- REQ-GIT-06: Complete-milestone MUST offer squash merge (recommended) or merge with history
- REQ-GIT-07: System MUST respect
commit_docssetting for.planning/files - REQ-GIT-08: System MUST auto-detect
.planning/in.gitignoreand skip commits
Commit Format:
type(phase-plan): description
# Examples:
docs(08-02): complete user registration plan
feat(08-02): add email confirmation flow
fix(03-01): correct auth token expiry
35. CLI Tools
Purpose: Programmatic utilities for workflows and agents, replacing repetitive inline bash patterns.
Requirements:
- REQ-CLI-01: System MUST provide atomic commands for state, config, phase, roadmap operations
- REQ-CLI-02: System MUST provide compound
initcommands that load all context for each workflow - REQ-CLI-03: System MUST support
--rawflag for machine-readable output - REQ-CLI-04: System MUST support
--cwdflag for sandboxed subagent operation - REQ-CLI-05: All operations MUST use forward-slash paths on Windows
Command Categories: State (11 subcommands), Phase (5), Roadmap (3), Verify (8), Template (2), Frontmatter (4), Scaffold (4), Init (12), Validate (2), Progress, Stats, Todo
36. Multi-Runtime Support
Purpose: Run GSD across multiple AI coding agent runtimes.
Requirements:
- REQ-RUNTIME-01: System MUST support Claude Code, OpenCode, Kilo, Codex, Copilot, Antigravity, Trae, Cline, Augment Code, CodeBuddy, Qwen Code
- REQ-RUNTIME-02: Installer MUST transform content per runtime (tool names, paths, frontmatter)
- REQ-RUNTIME-03: Installer MUST support interactive and non-interactive (
--claude --global) modes - REQ-RUNTIME-04: Installer MUST support both global and local installation
- REQ-RUNTIME-05: Uninstall MUST cleanly remove all GSD files without affecting other configurations
- REQ-RUNTIME-06: Installer MUST handle platform differences (Windows, macOS, Linux, WSL, Docker)
- REQ-RUNTIME-07: Runtimes with lifecycle hook support MUST register per-turn context-headroom tracking events at install time
- REQ-RUNTIME-08: Native packaging manifests MUST be version-stamped and enable runtime-native install/update/uninstall flows
Runtime Transformations:
| Aspect | Claude Code | OpenCode | Kilo | Codex | Copilot | Antigravity | Cursor | Trae | Cline | Augment | CodeBuddy | Qwen Code |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Commands | Slash commands | Slash commands | Slash commands | Skills (TOML) | Slash commands | Skills | Skills + Slash commands | Skills | Rules | Skills + Slash commands | Slash commands | Skills |
| Agent format | Claude native | mode: subagent |
mode: subagent |
Skills | Tool mapping | Skills | Skills | Skills | Rules | Skills | Skills | Skills |
| Skills emission | N/A | On-demand SKILL.md (1.4.0) | On-demand SKILL.md (1.4.0) | /skills picker (1.4.0) |
N/A | N/A | SKILL.md | N/A | On-demand SKILL.md (1.4.0) | N/A | N/A | N/A |
| Hook events | SessionStart, PreToolUse, PostToolUse, SubagentStop, Stop, PreCompact, FileChanged |
N/A | N/A | SessionStart, SubagentStart, Stop, PostToolUse |
sessionStart |
N/A | sessionStart, postToolUse |
N/A | PreToolUse |
N/A | N/A | SessionStart, PreToolUse, PostToolUse, SubagentStop, Stop, PreCompact |
| Config | settings.json |
opencode.json(c) |
kilo.json(c) |
TOML | Instructions | Config | Config | Config | .clinerules |
Config | Config | Config |
Cursor artifact surfaces: gsd install --cursor writes two artifact kinds:
~/.cursor/skills/gsd-<name>/SKILL.md— rich skills with YAML frontmatter, Cursor tool-name mapping, and adapter context header (existing surface)~/.cursor/commands/gsd-<name>.md— plain markdown slash commands (no frontmatter) invocable via/in the Agent input (Cursor 1.6+)
Native skills emission (1.4.0): Three runtimes now emit GSD as on-demand native skills (skills/<name>/SKILL.md) at install time, in addition to their existing command and agent surfaces. Skills respect the active install profile and are removed on uninstall.
- Cline (global installs, Cline >= v3.48.0) — emits skills alongside the existing
.clinerules/directory - Kilo — emits skills alongside
command/andagents/ - OpenCode — emits skills alongside its existing surfaces
New slash-command surfaces (1.4.0):
- CodeBuddy —
/gsd-*slash commands written to~/.codebuddy/commands/ - Augment —
commands/gsd-<name>.mdwritten to~/.augment/commands/ - Cursor (Cursor >= 1.6) —
.cursor/commands/gsd-<name>.mdso GSD appears in the/command menu
Cross-runtime lifecycle hooks (1.4.0): Each supported runtime registers lifecycle hook events for per-turn context-headroom tracking and workflow state management. Notable registrations:
- Claude Code:
SubagentStop,Stop,PreCompact(context-headroom warnings),FileChanged(hot-reloads.planning/config.jsonmid-session) - Qwen Code:
SubagentStop,Stop,PreCompact - Codex:
SubagentStart,Stop,PostToolUse(new in 1.4.0); on Windows theSessionStarthook entry gains acommandWindowsfield so the.cmdshim is used for native execution - Cline:
PreToolUse - Cursor:
sessionStart(injects workflow state),postToolUse(nudges.planningupdates) - Copilot:
sessionStart
Runtime-specific enrichments (1.4.0):
- Codex emits
service_tier: flexfor light-tier agents; GSD skills appear in the Codex/skillspicker viaSKILL.md(noagents/openai.yamlsidecar is emitted — doing so caused duplicate autocomplete entries, #1326)
Native packaging:
- Claude Code: GSD Core ships a
.claude-plugin/plugin.jsonmanifest, enabling installation and lifecycle management viaclaude plugin install|enable|disable|update gsd-core. Commands load under the/gsd-core:namespace (e.g./gsd-core:plan-phase), avoiding slash-command collisions with the classic npm installer which uses/gsd:. Always-on guard and update hooks are wired automatically viahooks/hooks.json. The plugin path is additive — the npm installer (npx @opengsd/gsd-core) remains fully supported.
37. Hook System
Purpose: Runtime event hooks for context monitoring, status display, and update checking.
Requirements:
- REQ-HOOK-01: Statusline MUST display model, current task, directory, and context usage
- REQ-HOOK-02: Context monitor MUST inject agent-facing warnings at threshold levels
- REQ-HOOK-03: Update checker MUST run in background on session start
- REQ-HOOK-04: All hooks MUST respect
CLAUDE_CONFIG_DIRenv var - REQ-HOOK-05: All hooks MUST include 3-second stdin timeout guard
- REQ-HOOK-06: All hooks MUST fail silently on any error
- REQ-HOOK-07: Context usage MUST normalize for autocompact buffer (16.5% reserved)
- REQ-HOOK-08: Update banner MUST be opt-in and silent unless an update is available (PR #2795)
Statusline Display:
[⬆ /gsd-update │] model │ [current task │] directory [█████░░░░░ 50%]
Color coding: <50% green, <65% yellow, <80% orange, ≥80% red with skull emoji
Update Banner (opt-in, when GSD statusline isn't used):
When the user declines (or keeps a non-GSD) statusline, the installer offers a SessionStart banner that surfaces update availability without occupying statusline real estate. The banner reads ~/.cache/gsd/gsd-update-check.json (written by gsd-check-update-worker.js) and emits one line only when an update is available:
GSD update available: 1.39.0 → 1.40.0. Run /gsd-update.
The banner is silent when up-to-date and rate-limits "check failed" diagnostics to once per 24 hours. Removed cleanly by npx @opengsd/gsd-core --uninstall or by deleting the SessionStart entry that references gsd-update-banner.js.
38. Developer Profiling
Command: /gsd-profile-user [--questionnaire] [--refresh]
Purpose: Analyze Claude Code session history to build behavioral profiles across 8 dimensions, generating artifacts that personalize Claude's responses to the developer's style.
Dimensions:
- Communication style (terse vs verbose, formal vs casual)
- Decision patterns (rapid vs deliberate, risk tolerance)
- Debugging approach (systematic vs intuitive, log preference)
- UX preferences (design sensibility, accessibility awareness)
- Vendor/technology choices (framework preferences, ecosystem familiarity)
- Frustration triggers (what causes friction in workflows)
- Learning style (documentation vs examples, depth preference)
- Explanation depth (high-level vs implementation detail)
Generated Artifacts:
USER-PROFILE.md— Full behavioral profile with evidence citationsCLAUDE.mdprofile section — Auto-discovered by Claude Code
Flags:
--questionnaire— Interactive questionnaire fallback when session history is unavailable--refresh— Re-analyze sessions and regenerate profile
Pipeline Modules:
profile-pipeline.cjs— Session scanning, message extraction, samplingprofile-output.cjs— Profile rendering, questionnaire, artifact generationgsd-user-profileragent — Behavioral analysis from session data
Requirements:
- REQ-PROF-01: Session analysis MUST cover at least 8 behavioral dimensions
- REQ-PROF-02: Profile MUST cite evidence from actual session messages
- REQ-PROF-03: Questionnaire MUST be available as fallback when no session history exists
- REQ-PROF-04: Generated artifacts MUST be discoverable by Claude Code (CLAUDE.md integration)
39. Execution Hardening
Purpose: Three additive quality improvements to the execution pipeline that catch cross-plan failures before they cascade.
Components:
1. Pre-Wave Dependency Check (execute-phase) Before spawning wave N+1, verify key-links from prior wave artifacts exist and are wired correctly. Catches cross-plan dependency gaps before they cascade into downstream failures.
2. Cross-Plan Data Contracts — Dimension 9 (plan-checker) New analysis dimension that checks plans sharing data pipelines have compatible transformations. Flags when one plan strips data that another plan needs in its original form.
3. Export-Level Spot Check (verify-phase) After Level 3 wiring verification passes, spot-check individual exports for actual usage. Catches dead stores that exist in wired files but are never called.
Requirements:
- REQ-HARD-01: Pre-wave check MUST verify key-links from all prior wave artifacts before spawning next wave
- REQ-HARD-02: Cross-plan contract check MUST detect incompatible data transformations between plans
- REQ-HARD-03: Export spot-check MUST identify dead stores in wired files
40. Verification Debt Tracking
Command: /gsd-audit-uat
Purpose: Prevent silent loss of UAT/verification items when projects advance past phases with outstanding tests. Surfaces verification debt across all prior phases so items are never forgotten.
Components:
1. Cross-Phase Health Check (progress.md Step 1.6)
Every /gsd-progress call scans ALL phases in the current milestone for outstanding items (pending, skipped, blocked, human_needed). Displays a non-blocking warning section with actionable links.
2. status: partial (verify-work.md, UAT.md)
New UAT status that distinguishes between "session ended" and "all tests resolved". Prevents status: complete when tests are still pending, blocked, or skipped without reason.
3. result: blocked with blocked_by tag (verify-work.md, UAT.md)
New test result type for tests blocked by external dependencies (server, physical device, release build, third-party services). Categorized separately from skipped tests.
4. HUMAN-UAT.md Persistence (execute-phase.md)
When verification returns human_needed, items are persisted as a trackable HUMAN-UAT.md file with status: partial. Feeds into the cross-phase health check and audit systems.
5. Phase Completion Warnings (phase.cjs, transition.md)
phase complete CLI returns verification debt warnings in its JSON output. Transition workflow surfaces outstanding items before confirmation.
Requirements:
- REQ-DEBT-01: System MUST surface outstanding UAT/verification items from ALL prior phases in
/gsd-progress - REQ-DEBT-02: System MUST distinguish incomplete testing (partial) from completed testing (complete)
- REQ-DEBT-03: System MUST categorize blocked tests with
blocked_bytags - REQ-DEBT-04: System MUST persist human_needed verification items as trackable UAT files
- REQ-DEBT-05: System MUST warn (non-blocking) during phase completion and transition when verification debt exists
- REQ-DEBT-06:
/gsd-audit-uatMUST scan all phases, categorize items by testability, and produce a human test plan
v1.27 Features
41. Fast Mode
Command: /gsd-fast [task description]
Purpose: Execute trivial tasks inline without spawning subagents or generating PLAN.md files. For tasks too small to justify planning overhead: typo fixes, config changes, small refactors, forgotten commits, simple additions.
Requirements:
- REQ-FAST-01: System MUST execute the task directly in the current context without subagents
- REQ-FAST-02: System MUST produce an atomic git commit for the change
- REQ-FAST-03: System MUST track the task in
.planning/quick/for state consistency - REQ-FAST-04: System MUST NOT be used for tasks requiring research, multi-step planning, or verification
When to use vs /gsd-quick:
/gsd-fast— One-sentence tasks executable in under 2 minutes (typo, config change, small addition)/gsd-quick— Anything needing research, multi-step planning, or verification
42. Cross-AI Peer Review
Command: /gsd-review --phase N [--gemini] [--claude] [--codex] [--coderabbit] [--opencode] [--qwen] [--cursor] [--agy] [--antigravity] [--ollama] [--lm-studio] [--llama-cpp] [--kimi-code] [--all]
Purpose: Invoke external AI CLIs (Gemini, Claude, Codex, CodeRabbit, OpenCode, Qwen Code, Cursor, Antigravity, Kimi Code) and local OpenAI-compatible servers (Ollama, LM Studio, llama.cpp) to independently review phase plans. Produces structured REVIEWS.md with per-reviewer feedback.
Each reviewer is a declared lane: its binary, prompt and output channels, timeout, availability probe, and empty-output policy come from a capability manifest rather than hand-written per-CLI logic, so a reviewer can be shipped as an installable capability instead of a core change.
Requirements:
- REQ-REVIEW-01: System MUST detect available AI CLIs on the system
- REQ-REVIEW-02: System MUST build a structured review prompt from phase plans
- REQ-REVIEW-03: System MUST invoke each selected CLI independently
- REQ-REVIEW-04: System MUST collect responses and produce
REVIEWS.md - REQ-REVIEW-05: Reviews MUST be consumable by
/gsd-plan-phase --reviews - REQ-REVIEW-06: System MUST support project-level no-flag defaults via
review.default_reviewers - REQ-REVIEW-07: Reviewer precedence MUST be explicit flags >
--all>review.default_reviewers> all detected reviewers
Produces: {phase}-REVIEWS.md — Per-reviewer structured feedback
User configuration note:
- Set
review.default_reviewersin.planning/config.json(or viagsd config-set) to control no-flag/gsd-reviewfan-out. review.default_reviewersmay include configuredreview.reviewer_instancesnames; each instance runs as an independent reviewer identity backed by its configured adapter/model. Instance names are not CLI flags.- Use
--allfor a full pre-merge sweep without changing project defaults. - For local model servers with small context windows, set
review.max_prompt_tokens_per_reviewerto auto-trim prompts per reviewer — see Prompt budgets for small-context reviewers in CONFIGURATION.md.
Why record which model produced a review (#2295): reviewers: in the frontmatter recorded which CLIs ran, but not which model each one resolved to. Without a pin, the model is whatever the CLI's own config or internal default happens to pick, so a "Codex vs Antigravity" comparison could quietly be a frontier model against a cheap-tier default with nothing in the record to say so — and a CLI update, or an unrelated config edit, could silently make past and future reviews incomparable.
The fix records the model and its provenance. Provenance is what makes the value trustworthy: pinned (from review.models.<slug>) is certain, while banner and transcript are recovered from third-party CLI output this project does not own — a startup banner or an undocumented session log.
That third-party dependence is a real trade-off, held honestly rather than papered over: the banner and transcript arms read formats GSD does not control, so they are best-effort by design and degrade to unknown rather than guessing or failing the run. A recorded unknown is a real answer — a wrong model name attributed to a review would be worse than none.
43. Backlog Parking Lot
Commands: /gsd-capture --backlog <description>, /gsd-review-backlog, /gsd-capture --seed <idea>, /gsd-capture --list-seeds [status]
Purpose: Capture ideas that aren't ready for active planning. Backlog items use 999.x numbering to stay outside the active phase sequence. Seeds are forward-looking ideas with trigger conditions that surface automatically at the right milestone. --list-seeds provides a read-only audit of all parked seeds (with optional status filter) without waiting for the next milestone.
Requirements:
- REQ-BACKLOG-01: Backlog items MUST use 999.x numbering to stay outside active phase sequence
- REQ-BACKLOG-02: Phase directories MUST be created immediately so
/gsd-discuss-phaseand/gsd-plan-phasework on them - REQ-BACKLOG-03:
/gsd-review-backlogMUST support promote, keep, and remove actions per item - REQ-BACKLOG-04: Promoted items MUST be renumbered into the active milestone sequence
- REQ-SEED-01: Seeds MUST capture the full WHY and WHEN to surface conditions
- REQ-SEED-02:
/gsd-new-milestoneMUST scan seeds and present matches - REQ-SEED-03:
/gsd-capture --list-seedsMUST list seeds with status, scope, and trigger for audit, with optional status filtering
Produces:
| Artifact | Description |
|---|---|
.planning/phases/999.x-slug/ |
Backlog item directory |
.planning/seeds/SEED-NNN-slug.md |
Seed with trigger conditions |
44. Persistent Context Threads
Command: /gsd-thread [name | description]
Purpose: Lightweight cross-session knowledge stores for work that spans multiple sessions but doesn't belong to any specific phase. Lighter weight than /gsd-pause-work — no phase state, no plan context.
Requirements:
- REQ-THREAD-01: System MUST support create, list, and resume modes
- REQ-THREAD-02: Threads MUST be stored in
.planning/threads/as markdown files - REQ-THREAD-03: Thread files MUST include Goal, Context, References, and Next Steps sections
- REQ-THREAD-04: Resuming a thread MUST load its full context into the current session
- REQ-THREAD-05: Threads MUST be promotable to phases or backlog items
Produces: .planning/threads/{slug}.md — Persistent context thread
45. PR Branch Filtering
Command: /gsd-pr-branch [target branch]
Purpose: Create a clean branch suitable for pull requests by filtering out .planning/ commits. Reviewers see only code changes, not GSD planning artifacts.
Requirements:
- REQ-PRBRANCH-01: System MUST identify commits that only modify
.planning/files - REQ-PRBRANCH-02: System MUST create a new branch with planning commits filtered out
- REQ-PRBRANCH-03: Code changes MUST be preserved exactly as committed
- REQ-PRBRANCH-04: System MUST NOT delete a
.planning/path the target branch already tracks - REQ-PRBRANCH-05: Verification MUST assert against the active filter mode's contract, not an unconditional zero
Filter modes. planning.pr_strict selects what "filtered" means. The default mode treats .planning/ as two populations: structural state that belongs in review (STATE.md, ROADMAP.md, MILESTONES.md, PROJECT.md, REQUIREMENTS.md, milestones/**) and transient per-phase artifacts that do not (phases/, quick/, research/, threads/, todos/, debug/, seeds/, codebase/, ui-reviews/). Strict mode collapses that distinction: nothing under .planning/ reaches the PR branch, and a commit is carried over only when it touches at least one file outside .planning/.
Strict mode exists because the two ways to keep planning private are not equivalent. Turning off planning.commit_docs keeps .planning/ out of git, which also takes parallel executor worktrees with it — a worktree is checked out from a commit, so an untracked planning tree is simply absent inside it and the executor has no PLAN.md to read. Strict mode leaves planning committed, so worktrees and revert paths keep working, and moves the guarantee to the publication boundary instead. See Publish PRs without planning artifacts.
Both modes filter by forcing the excluded paths back to whatever the target branch already tracks, in the index and the working tree. Un-staging alone would record a deletion of any planning file the target branch carries, and would leave the picked file untracked on disk, where it makes a later cherry-pick of the same path abort.
46. Security Hardening
Purpose: Defense-in-depth security for GSD's planning artifacts. Because GSD generates markdown files that become LLM system prompts, user-controlled text flowing into these files is a potential indirect prompt injection vector.
Components:
1. Centralized Security Module (security.cjs)
- Path traversal prevention — validates file paths resolve within the project directory
- Prompt injection detection — scans for known injection patterns in user-supplied text
- Safe JSON parsing — catches malformed input before state corruption
- Field name validation — prevents injection through config field names
- Shell argument validation — sanitizes user text before shell interpolation
2. Prompt Injection Guard Hook (gsd-prompt-guard.js)
PreToolUse hook that scans Write/Edit calls targeting .planning/ for injection patterns. Advisory-only — logs detection for awareness without blocking legitimate operations.
3. Workflow Guard Hook (gsd-workflow-guard.js)
PreToolUse hook that detects when Claude attempts file edits outside a GSD workflow context. Advises using /gsd-quick or /gsd-fast instead of direct edits. Configurable via hooks.workflow_guard (default: false).
4. CI-Ready Injection Scanner (prompt-injection-scan.security.test.cjs)
Test suite that scans all agent, workflow, and command files for embedded injection vectors.
Requirements:
- REQ-SEC-01: All user-supplied file paths MUST be validated against the project directory
- REQ-SEC-02: Prompt injection patterns MUST be detected before text enters planning artifacts
- REQ-SEC-03: Security hooks MUST be advisory-only (never block legitimate operations)
- REQ-SEC-04: JSON parsing of user input MUST catch malformed data gracefully
- REQ-SEC-05: macOS
/var→/private/varsymlink resolution MUST be handled in path validation
47. Multi-Repo Workspace Support
Purpose: Auto-detection and project root resolution for monorepos and multi-repo setups. Supports workspaces where .planning/ may need to resolve across repository boundaries.
Requirements:
- REQ-MULTIREPO-01: System MUST auto-detect multi-repo workspace configuration
- REQ-MULTIREPO-02: System MUST resolve project root across repository boundaries
- REQ-MULTIREPO-03: Executor MUST record per-repo commit hashes in multi-repo mode
48. Discussion Audit Trail
Purpose: Auto-generate DISCUSSION-LOG.md during /gsd-discuss-phase for full audit trail of decisions made during discussion.
Requirements:
- REQ-DISCLOG-01: System MUST auto-generate DISCUSSION-LOG.md during discuss-phase
- REQ-DISCLOG-02: Log MUST capture questions asked, options presented, and decisions made
- REQ-DISCLOG-03: Decision IDs MUST enable traceability from discuss-phase to plan-phase
v1.28 Features
49. Forensics
Command: /gsd-forensics [description]
Purpose: Post-mortem investigation of failed or stuck GSD workflows.
Requirements:
- REQ-FORENSICS-01: System MUST analyze git history for anomalies (stuck loops, long gaps, repeated commits)
- REQ-FORENSICS-02: System MUST check artifact integrity (completed phases have expected files)
- REQ-FORENSICS-03: System MUST generate a markdown report saved to
.planning/forensics/ - REQ-FORENSICS-04: System MUST offer to create a GitHub issue with findings
- REQ-FORENSICS-05: System MUST NOT modify project files (read-only investigation)
Produces:
| Artifact | Description |
|---|---|
.planning/forensics/report-{timestamp}.md |
Post-mortem investigation report |
Process:
- Scan — Analyze git history for anomalies: stuck loops, long gaps between commits, repeated identical commits
- Integrity Check — Verify completed phases have expected artifact files
- Report — Generate markdown report with findings, saved to
.planning/forensics/ - Issue — Offer to create a GitHub issue with findings for team visibility
50. Milestone Summary
Command: /gsd-milestone-summary [version]
Purpose: Generate comprehensive project summary from milestone artifacts for team onboarding.
Requirements:
- REQ-SUMMARY-01: System MUST aggregate phase plans, summaries, and verification results
- REQ-SUMMARY-02: System MUST work for both current and archived milestones
- REQ-SUMMARY-03: System MUST produce a single navigable document
Produces:
| Artifact | Description |
|---|---|
MILESTONE-SUMMARY.md |
Comprehensive navigable summary of milestone artifacts |
Process:
- Collect — Aggregate phase plans, summaries, and verification results from the target milestone
- Synthesize — Combine artifacts into a single navigable document with cross-references
- Output — Write
MILESTONE-SUMMARY.mdsuitable for team onboarding and stakeholder review
51. Workstream Namespacing
Command: /gsd-workstreams
Purpose: Parallel workstreams for concurrent work on different milestone areas.
Requirements:
- REQ-WS-01: System MUST isolate workstream state in separate
.planning/workstreams/{name}/directories - REQ-WS-02: System MUST validate workstream names (alphanumeric + hyphens only, no path traversal)
- REQ-WS-03: System MUST support list, create, switch, status, progress, complete, resume subcommands
Produces:
| Artifact | Description |
|---|---|
.planning/workstreams/{name}/ |
Isolated workstream directory structure |
Process:
- Create — Initialize a named workstream with isolated
.planning/workstreams/{name}/directory - Switch — Change active workstream context for subsequent GSD commands
- Manage — List, check status, track progress, complete, or resume workstreams
52. Manager Dashboard
Command: /gsd-manager
Purpose: Interactive command center for managing multiple phases from one terminal.
Requirements:
- REQ-MGR-01: System MUST show overview of all phases with status
- REQ-MGR-02: System MUST filter to current milestone scope
- REQ-MGR-03: System MUST show phase dependencies and conflicts
Produces: Interactive terminal output
Process:
- Scan — Load all phases in the current milestone with their statuses
- Display — Render overview showing phase dependencies, conflicts, and progress
- Interact — Accept commands to navigate, inspect, or act on individual phases
53. Assumptions Discussion Mode
Command: /gsd-discuss-phase with workflow.discuss_mode: 'assumptions'
Purpose: Replace interview-style questioning with codebase-first assumption analysis.
Requirements:
- REQ-ASSUME-01: System MUST analyze codebase to generate structured assumptions before asking questions
- REQ-ASSUME-02: System MUST classify assumptions by confidence level (Confident/Likely/Unclear)
- REQ-ASSUME-03: System MUST produce identical CONTEXT.md format as default discuss mode
- REQ-ASSUME-04: System MUST support confidence-based skip gate (all HIGH = no questions)
Produces:
| Artifact | Description |
|---|---|
{phase}-CONTEXT.md |
Same format as default discuss mode |
Process:
- Analyze — Scan codebase to generate structured assumptions about implementation approach
- Classify — Categorize assumptions by confidence level: Confident, Likely, Unclear
- Gate — If all assumptions are HIGH confidence, skip questioning entirely
- Confirm — Present unclear assumptions as targeted questions to the user
- Output — Produce
{phase}-CONTEXT.mdin identical format to default discuss mode
54. UI Phase Auto-Detection
Part of: /gsd-new-project and /gsd-progress
Purpose: Automatically detect UI-heavy projects and surface /gsd-ui-phase recommendation.
Requirements:
- REQ-UI-DETECT-01: System MUST detect UI signals in project description (keywords, framework references)
- REQ-UI-DETECT-02: System MUST annotate ROADMAP.md phases with
ui_hintwhen applicable - REQ-UI-DETECT-03: System MUST suggest
/gsd-ui-phasein next steps for UI-heavy phases - REQ-UI-DETECT-04: System MUST NOT make
/gsd-ui-phasemandatory
Process:
- Detect — Scan project description and tech stack for UI signals (keywords, framework references)
- Annotate — Add
ui_hintmarkers to applicable phases in ROADMAP.md - Surface — Include
/gsd-ui-phaserecommendation in next steps for UI-heavy phases
55. Multi-Runtime Installer Selection
Part of: npx @opengsd/gsd-core
Purpose: Select multiple runtimes in a single interactive install session.
Requirements:
- REQ-MULTI-RT-01: Interactive prompt MUST support multi-select (e.g., Claude Code + Antigravity)
- REQ-MULTI-RT-02: CLI flags MUST continue to work for non-interactive installs
Process:
- Detect — Identify available AI CLI runtimes on the system
- Prompt — Present multi-select interface for runtime selection
- Install — Configure GSD for all selected runtimes in a single session
v1.29 Features
56. Windsurf Runtime Support
Part of: npx @opengsd/gsd-core
Purpose: Add Windsurf as a supported AI CLI runtime for GSD installation and execution.
Requirements:
- REQ-WINDSURF-01: Installer MUST detect Windsurf runtime and offer it as a target
- REQ-WINDSURF-02: GSD commands MUST function correctly within Windsurf sessions
Process:
- Detect — Identify Windsurf runtime availability on the system
- Install — Configure GSD skills and hooks for the Windsurf environment
57. Internationalized Documentation
Part of: docs/
Purpose: Provide GSD documentation in Portuguese, Korean, and Japanese.
Requirements:
- REQ-I18N-01: Documentation MUST be available in Portuguese (pt), Korean (ko), and Japanese (ja)
- REQ-I18N-02: Translations MUST stay synchronized with English source documents
Process:
- Translate — Convert core documentation into target languages
- Publish — Make translated documentation accessible alongside English originals
v1.31 Features
59. Schema Drift Detection
Command: Automatic during /gsd-execute-phase
Purpose: Detect when ORM schema files are modified without corresponding migration or push commands, preventing false-positive verification.
Requirements:
- REQ-SCHEMA-01: System MUST detect modifications to ORM schema files (Prisma, Drizzle, Payload, Sanity, Mongoose)
- REQ-SCHEMA-02: System MUST verify corresponding migration/push commands exist when schema changes are detected
- REQ-SCHEMA-03: System MUST implement two-layer defense: plan-time injection and execute-time gate
- REQ-SCHEMA-04: System MUST support
GSD_SKIP_SCHEMA_CHECKenv var to override detection - REQ-SCHEMA-05: System MUST prevent false-positive verification when schema is modified without migration
Process:
- Detect — Monitor ORM schema file modifications during plan execution
- Verify — Check that corresponding migration/push commands are present in the plan
- Gate — Block execution if schema drift is detected without migration (execute-time gate)
- Inject — Add migration reminders during plan generation (plan-time injection)
Config: GSD_SKIP_SCHEMA_CHECK environment variable to bypass detection.
60. Security Enforcement
Command: /gsd-secure-phase <N>
Purpose: Threat-model-anchored security verification for phase implementations.
Requirements:
- REQ-SEC-01: System MUST perform threat-model-anchored verification (not blind scanning)
- REQ-SEC-02: System MUST support configurable OWASP ASVS verification levels (1-3)
- REQ-SEC-03: System MUST block phase advancement based on configurable severity threshold
- REQ-SEC-04: System MUST spawn
gsd-security-auditoragent for analysis
Produces:
| Artifact | Description |
|---|---|
| Security audit report | Threat-model-anchored findings with severity classification |
Process:
- Model — Build threat model from phase implementation context
- Audit — Spawn
gsd-security-auditorto verify against threat model - Gate — Block phase advancement if findings meet or exceed
security_block_onseverity
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
security_enforcement |
boolean | true |
Enable threat-model security verification |
security_asvs_level |
number (1-3) | 1 |
OWASP ASVS verification level |
security_block_on |
string | "high" |
Minimum severity to block phase advancement |
61. Documentation Generation
Command: /gsd-docs-update
Purpose: Generate and verify project documentation with accuracy checks.
Requirements:
- REQ-DOCS-01: System MUST spawn
gsd-doc-writeragent to generate documentation - REQ-DOCS-02: System MUST spawn
gsd-doc-verifieragent to check accuracy - REQ-DOCS-03: System MUST verify generated documentation against actual implementation
Produces:
| Artifact | Description |
|---|---|
| Updated project documentation | Generated and verified documentation files |
Process:
- Generate — Spawn
gsd-doc-writerto create or update documentation from implementation - Verify — Spawn
gsd-doc-verifierto check documentation accuracy against codebase - Output — Produce verified documentation with accuracy annotations
62. Discuss Chain Mode
Flag: /gsd-discuss-phase <N> --chain
Purpose: Auto-chain discuss, plan, and execute phases in one flow to reduce manual command sequencing.
Requirements:
- REQ-CHAIN-01: System MUST auto-chain discuss → plan → execute when
--chainflag is provided - REQ-CHAIN-02: System MUST respect all gate settings between chained phases
- REQ-CHAIN-03: System MUST halt the chain if any phase fails
Process:
- Discuss — Run discuss-phase to gather context
- Plan — Automatically invoke plan-phase with gathered context
- Execute — Automatically invoke execute-phase with generated plan
63. Single-Phase Autonomous
Flag: /gsd-autonomous --only N
Purpose: Execute just one phase autonomously instead of all remaining phases.
Requirements:
- REQ-ONLY-01: System MUST execute only the specified phase number when
--only Nis provided - REQ-ONLY-02: System MUST follow the same discuss → plan → execute flow as full autonomous mode
- REQ-ONLY-03: System MUST stop after the specified phase completes
Process:
- Select — Identify the target phase from
--only Nargument - Execute — Run full autonomous flow (discuss → plan → execute) for that single phase
- Stop — Halt after the phase completes instead of advancing to the next
64. Scope Reduction Detection
Part of: /gsd-plan-phase
Purpose: Prevent silent requirement dropping during plan generation with three-layer defense.
Requirements:
- REQ-SCOPE-01: System MUST prohibit planners from reducing scope without explicit justification
- REQ-SCOPE-02: System MUST have plan-checker verify requirement dimension coverage
- REQ-SCOPE-03: System MUST have orchestrator recover dropped requirements and re-inject them
- REQ-SCOPE-04: System MUST implement three-layer defense: planner prohibition, checker dimension, orchestrator recovery
Process:
- Prohibit — Planner instructions explicitly forbid scope reduction
- Check — Plan-checker verifies all phase requirements are covered in the plan
- Recover — Orchestrator detects dropped requirements and re-injects them into the planning loop
65. Claim Provenance Tagging
Part of: /gsd-plan-phase --research-phase <N>
Purpose: Ensure research claims are tagged with source evidence and assumptions are logged separately.
Requirements:
- REQ-PROVENANCE-01: Researcher MUST mark claims with source evidence references
- REQ-PROVENANCE-02: Assumptions MUST be logged separately from sourced claims
- REQ-PROVENANCE-03: System MUST distinguish between evidenced facts and inferred assumptions
Process:
- Research — Researcher gathers information from codebase and domain sources
- Tag — Each claim is annotated with its source (file path, documentation, API response)
- Separate — Assumptions without direct evidence are logged in a distinct section
66. Worktree Toggle
Config: workflow.use_worktrees: false
Purpose: Disable git worktree isolation for users who prefer sequential execution.
Requirements:
- REQ-WORKTREE-01: System MUST respect
workflow.use_worktreessetting when deciding isolation strategy - REQ-WORKTREE-02: System MUST default to
true(worktrees enabled) for backward compatibility - REQ-WORKTREE-03: System MUST fall back to sequential execution when worktrees are disabled
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
workflow.use_worktrees |
boolean | true |
When false, disables git worktree isolation |
67. Project Code Prefixing
Config: project_code: "ABC"
Purpose: Prefix phase directory names with a project code for multi-project disambiguation.
Requirements:
- REQ-PREFIX-01: System MUST prefix phase directories with project code when configured (e.g.,
ABC-01-setup/) - REQ-PREFIX-02: System MUST use standard naming when
project_codeis not set - REQ-PREFIX-03: System MUST apply prefix consistently across all phase operations
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
project_code |
string | (none) | Prefix for phase directory names |
68. Claude Code Skills Migration
Part of: npx @opengsd/gsd-core
Purpose: Migrate GSD commands to Claude Code 2.1.88+ skills format with backward compatibility.
Requirements:
- REQ-SKILLS-01: Installer MUST write
skills/gsd-*/SKILL.mdfor Claude Code 2.1.88+ - REQ-SKILLS-02: Installer MUST auto-clean legacy
commands/gsd/directory - REQ-SKILLS-03: Installer MUST maintain backward compatibility with older Claude Code versions via the legacy
commands/gsd/path
Process:
- Detect — Check Claude Code version to determine skills support
- Migrate — Write
skills/gsd-*/SKILL.mdfiles for each GSD command - Clean — Remove legacy
commands/gsd/directory if skills are installed - Fallback — Maintain legacy
commands/gsd/path compatibility for older Claude Code versions
v1.32 Features
69. STATE.md Consistency Gates
Commands: state validate [--strict], state sync [--verify], state planned-phase --phase N --plans N
Purpose: Detect and repair drift between STATE.md and the actual filesystem, preventing cascading errors from stale state.
Requirements:
- REQ-STATE-01:
state validateMUST detect drift between STATE.md fields and filesystem reality - REQ-STATE-02:
state syncMUST reconstruct STATE.md from actual project state on disk - REQ-STATE-03:
state sync --verifyMUST perform a dry-run showing proposed changes without writing - REQ-STATE-04:
state planned-phaseMUST record the state transition after plan-phase completes (Planned/Ready to execute) - REQ-STATE-05:
state validateMUST report aLast activityvalue that no reader can parse, rather than validating clean - REQ-STATE-06:
state validate --strictMUST reflectvalidin the process exit status, leaving the default exit status unchanged
Produces:
| Artifact | Description |
|---|---|
Updated STATE.md |
Corrected state reflecting filesystem reality |
Process:
- Validate — Compare STATE.md fields against filesystem (phase directories, plan files, summaries)
- Sync — Reconstruct STATE.md from disk when drift is detected
- Transition — Record post-planning state with plan count for execute-phase readiness
70. Autonomous --to N Flag
Flag: /gsd-autonomous --to N
Purpose: Stop autonomous execution after completing a specific phase, allowing partial autonomous runs.
Requirements:
- REQ-TO-01: System MUST stop execution after the specified phase number completes
- REQ-TO-02: System MUST follow the same discuss -> plan -> execute flow for each phase up to N
- REQ-TO-03:
--to NMUST be combinable with--from Nfor bounded autonomous ranges
Process:
- Bound — Set the upper phase limit from
--to Nargument - Execute — Run autonomous flow for each phase up to and including phase N
- Stop — Halt after phase N completes
71. Research Gate
Part of: /gsd-plan-phase
Purpose: Block planning when RESEARCH.md has unresolved open questions, preventing plans built on incomplete information.
Requirements:
- REQ-RESGATE-01: System MUST scan RESEARCH.md for unresolved open questions before planning begins
- REQ-RESGATE-02: System MUST block plan-phase entry when open questions exist
- REQ-RESGATE-03: System MUST surface the specific unresolved questions to the user
Process:
- Scan — Check RESEARCH.md for open questions section with unresolved items
- Gate — Block planning if unresolved questions are found
- Surface — Display the specific open questions requiring resolution
72. Verifier Milestone Scope Filtering
Part of: /gsd-execute-phase (verifier step)
Purpose: Distinguish between genuine gaps and items deferred to later phases, reducing false negatives in verification.
Requirements:
- REQ-VSCOPE-01: Verifier MUST check whether a gap is addressed in a later milestone phase
- REQ-VSCOPE-02: Gaps addressed in later phases MUST be marked as "deferred", not "gap"
- REQ-VSCOPE-03: Only genuine gaps (not covered by any future phase) MUST be reported as failures
Process:
- Verify — Run standard goal-backward verification
- Filter — Cross-reference detected gaps against later milestone phases
- Classify — Mark deferred items separately from genuine gaps
73. Read-Before-Edit Guard Hook
Part of: Hooks (PreToolUse)
Purpose: Prevent infinite retry loops in non-Claude runtimes by ensuring files are read before editing.
Requirements:
- REQ-RBE-01: Hook MUST detect Edit/Write tool calls that target files not previously read in the session
- REQ-RBE-02: Hook MUST advise reading the file first (advisory, non-blocking)
- REQ-RBE-03: Hook MUST prevent infinite retry loops common in runtimes without built-in read-before-edit enforcement
74. Context Reduction
Part of: prompt assembly pipeline
Purpose: Reduce context prompt sizes through markdown truncation and cache-friendly prompt ordering.
Requirements:
- REQ-CTXRED-01: System MUST truncate oversized markdown artifacts to fit within context budgets
- REQ-CTXRED-02: System MUST order prompts for cache-friendly assembly (stable prefixes first)
- REQ-CTXRED-03: Reduction MUST preserve essential information (headings, requirements, task structure)
- REQ-CTXRED-04: Skill
description:fields MUST be ≤ 100 chars; enforced bynpm run lint:descriptions(seescripts/lint-descriptions.cjsandtests/skill-frontmatter-contract.test.cjs)
Process:
- Measure — Calculate total prompt size for the workflow
- Truncate — Apply markdown-aware truncation to oversized artifacts
- Order — Arrange prompt sections for optimal KV-cache reuse
75. Discuss-Phase --power Flag
Flag: /gsd-discuss-phase --power
Purpose: File-based bulk question answering for discuss-phase, enabling batch input from a prepared answers file.
Requirements:
- REQ-POWER-01: System MUST accept a file containing pre-written answers to discussion questions
- REQ-POWER-02: System MUST map answers to the corresponding gray area questions
- REQ-POWER-03: System MUST produce CONTEXT.md identical to interactive discuss-phase
76. Debug --diagnose Flag
Flag: /gsd-debug --diagnose
Purpose: Diagnosis-only mode that investigates without attempting fixes.
Requirements:
- REQ-DIAG-01: System MUST perform full debug investigation (hypotheses, evidence, root cause)
- REQ-DIAG-02: System MUST NOT attempt any code modifications
- REQ-DIAG-03: System MUST produce a diagnostic report with findings and recommended fixes
77. Phase Dependency Analysis
Command: /gsd-manager --analyze-deps
Purpose: Detect phase dependencies and suggest Depends on entries for ROADMAP.md before running /gsd-manager.
Requirements:
- REQ-DEP-01: System MUST detect file overlap between phases
- REQ-DEP-02: System MUST detect semantic dependencies (API/schema producers and consumers)
- REQ-DEP-03: System MUST detect data flow dependencies (output producers and readers)
- REQ-DEP-04: System MUST suggest dependency entries with user confirmation before writing
Produces: Dependency suggestion table; optionally updates ROADMAP.md Depends on fields
78. Anti-Pattern Severity Levels
Part of: /gsd-resume-work
Purpose: Mandatory understanding checks at resume with severity-based anti-pattern enforcement.
Requirements:
- REQ-ANTI-01: System MUST classify anti-patterns by severity level
- REQ-ANTI-02: System MUST enforce mandatory understanding checks at session resume
- REQ-ANTI-03: Higher severity anti-patterns MUST block workflow progression until acknowledged
79. Methodology Artifact Type
Part of: Planning artifacts
Purpose: Define consumption mechanisms for methodology documents, ensuring they are consumed correctly by agents.
Requirements:
- REQ-METHOD-01: System MUST support methodology as a distinct artifact type
- REQ-METHOD-02: Methodology artifacts MUST have defined consumption mechanisms for agents
80. Planner Reachability Check
Part of: /gsd-plan-phase
Purpose: Validate that plan steps are achievable before committing to execution.
Requirements:
- REQ-REACH-01: Planner MUST validate that each plan step references reachable files and APIs
- REQ-REACH-02: Unreachable steps MUST be flagged during planning, not discovered during execution
81. Playwright-MCP UI Verification
Part of: /gsd-verify-work (optional)
Purpose: Automated visual verification using Playwright-MCP during verify-phase.
Requirements:
- REQ-PLAY-01: System MUST support optional Playwright-MCP visual verification during verify-phase
- REQ-PLAY-02: Visual verification MUST be opt-in, not mandatory
- REQ-PLAY-03: System MUST capture and compare visual state against UI-SPEC.md expectations
82. Pause-Work Expansion
Part of: /gsd-pause-work
Purpose: Support non-phase contexts with richer handoff data for broader pause-work applicability.
Requirements:
- REQ-PAUSE-01: System MUST support pausing in non-phase contexts (quick tasks, debug sessions, threads)
- REQ-PAUSE-02: Handoff data MUST include richer context appropriate to the current work type
83. Response Language Config
Config: response_language
Purpose: Cross-phase language consistency for non-English users.
Requirements:
- REQ-LANG-01: System MUST respect
response_languagesetting across all phases and agents - REQ-LANG-02: Setting MUST propagate to all spawned agents for consistent language output
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
response_language |
string | (none) | Language code for agent responses (e.g., "pt", "ko", "ja") |
84. Manual Update Procedure
Part of: docs/manual-update.md
Purpose: Document a manual update path for environments where npx is unavailable or npm publish is experiencing outages.
Requirements:
- REQ-MANUAL-01: Documentation MUST describe step-by-step manual update procedure
- REQ-MANUAL-02: Procedure MUST work without npm access
85. New Runtime Support (Trae, Cline, Augment Code)
Part of: npx @opengsd/gsd-core
Purpose: Extend GSD installation to Trae IDE, Cline, and Augment Code runtimes.
Requirements:
- REQ-TRAE-01: Installer MUST support
--traeflag for Trae IDE installation - REQ-CLINE-01: Installer MUST support Cline via
.clinerulesconfiguration - REQ-AUGMENT-01: Installer MUST support Augment Code with skill conversion and config management
86. Autonomous --interactive Flag
Flag: /gsd-autonomous --interactive
Purpose: Lean-context autonomous mode that keeps discuss-phase interactive (user answers questions) while dispatching plan and execute as background agents on runtimes that support nested background dispatch; on Claude Code, plan and execute run inline to preserve worktree isolation and independent verification.
Requirements:
- REQ-INTERACT-01:
--interactiveMUST run discuss-phase inline with interactive questions (not auto-answered) - REQ-INTERACT-02:
--interactiveMUST dispatch plan-phase and execute-phase as background agents for context isolation on runtimes where a backgrounded agent can spawn subagents; on Claude Code, plan and execute run inline - REQ-INTERACT-03:
--interactiveMUST enable pipeline parallelism — discuss Phase N+1 while Phase N builds (applies on runtimes that support nested background dispatch; on Claude Code, discuss does not overlap planning/execution) - REQ-INTERACT-04: Main context MUST only accumulate discuss conversations (lean context) on runtimes that support nested background dispatch; on Claude Code, inline plan/execute also accumulate in the main context
Process:
- Discuss inline — Run discuss-phase in the main context with user interaction
- Dispatch — On runtimes that support nested background dispatch: send plan and execute to background agents with fresh context windows. On Claude Code: run plan and execute inline.
- Pipeline — On runtimes with background dispatch: while background agents build Phase N, begin discussing Phase N+1. On Claude Code: phases run sequentially.
87. Commit-Docs Guard Hook
Hook: gsd-commit-docs.js
Purpose: PreToolUse hook that enforces the commit_docs configuration, preventing .planning/ files from being committed when planning.commit_docs is false.
Requirements:
- REQ-COMMITDOCS-01: Hook MUST intercept git commit commands that stage
.planning/files - REQ-COMMITDOCS-02: Hook MUST block commits containing
.planning/files whencommit_docsisfalse - REQ-COMMITDOCS-03: Hook MUST be advisory — does not block when
commit_docsistrueor absent
88. Community Hooks Opt-In
Hooks: gsd-validate-commit.sh, gsd-session-state.sh, gsd-phase-boundary.sh
Purpose: Optional git and session hooks for GSD projects, gated behind hooks.community: true in config.
Requirements:
- REQ-COMMUNITY-01: All community hooks MUST be no-ops unless
hooks.communityistruein.planning/config.json - REQ-COMMUNITY-02:
gsd-validate-commit.shMUST enforce Conventional Commits format on git commit messages - REQ-COMMUNITY-03:
gsd-session-state.shMUST track session state transitions - REQ-COMMUNITY-04:
gsd-phase-boundary.shMUST enforce phase boundary checks
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
hooks.community |
boolean | false |
Enable optional community hooks for commit validation, session state, and phase boundaries |
v1.34.0 Features
89. Global Learnings Store
Commands: Auto-triggered at phase completion; consumed by planner
Config: features.global_learnings
Purpose: Persist cross-session, cross-project learnings in a global store so the planner agent can learn from patterns across the entire project history — not just the current session.
Requirements:
- REQ-LEARN-01: Learnings MUST be auto-copied from
.planning/to the global store at phase completion - REQ-LEARN-02: The planner agent MUST receive relevant learnings at spawn time via injection
- REQ-LEARN-03: Injection MUST be capped by
learnings.max_injectto avoid context bloat - REQ-LEARN-04: Feature MUST be opt-in via
features.global_learnings: true
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
features.global_learnings |
boolean | false |
Enable cross-project learnings pipeline |
learnings.max_inject |
number | (system default) | Maximum learnings entries injected into planner |
90. Queryable Codebase Intelligence
Command: /gsd-map-codebase --query [<term>|status|diff|refresh]
Config: intel.enabled
Purpose: Maintain a queryable JSON index of codebase structure, API surface, dependency graph, file roles, and architecture decisions in .planning/intel/. Enables targeted lookups without reading the entire codebase.
Requirements:
- REQ-INTEL-01: Intel files MUST be stored as JSON in
.planning/intel/ - REQ-INTEL-02:
querymode MUST search across all intel files for a term and group results by file - REQ-INTEL-03:
statusmode MUST report freshness (FRESH/STALE, stale threshold: 24 hours) - REQ-INTEL-04:
diffmode MUST compare current intel state to the last snapshot - REQ-INTEL-05:
refreshmode MUST spawn the intel-updater agent to rebuild all files - REQ-INTEL-06: Feature MUST be opt-in via
intel.enabled: true
Intel files produced:
| File | Contents |
|---|---|
stack.json |
Technology stack and dependencies |
api-map.json |
Exported functions and API surface |
dependency-graph.json |
Inter-module dependency relationships |
file-roles.json |
Role classification for each source file |
arch-decisions.json |
Detected architecture decisions |
91. Execution Context Profiles
Config: context_profile
Purpose: Select a pre-configured execution context (mode, model, workflow settings) tuned for a specific type of work without manually adjusting individual settings.
Requirements:
- REQ-CTX-01:
devprofile MUST optimize for iterative development (balanced model, plan_check enabled) - REQ-CTX-02:
researchprofile MUST optimize for research-heavy work (higher model tier, research enabled) - REQ-CTX-03:
reviewprofile MUST optimize for code review work (verifier and code_review enabled)
Available profiles: dev, research, review
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
context_profile |
string | (none) | Execution context preset: dev, research, or review |
92. Gates Taxonomy
References: gsd-core/references/gates.md
Agents: plan-checker, verifier
Purpose: Define 4 canonical gate types that structure all workflow decision points, enabling plan-checker and verifier agents to apply consistent gate logic.
Gate types:
| Type | Description |
|---|---|
| Confirm | User approves before proceeding (e.g., roadmap review) |
| Quality | Automated quality check must pass (e.g., plan verification loop) |
| Safety | Hard stop on detected risk or policy violation |
| Transition | Phase or milestone boundary acknowledgment |
Requirements:
- REQ-GATES-01: plan-checker MUST classify each checkpoint as one of the 4 gate types
- REQ-GATES-02: verifier MUST apply gate logic appropriate to the gate type
- REQ-GATES-03: Hard stop safety gates MUST never be bypassed by
--autoflags
93. Code Review Pipeline
Commands: /gsd-code-review, /gsd-code-review --fix
Purpose: Structured review of source files changed during a phase, with a separate auto-fix pass that commits each fix atomically.
Requirements:
- REQ-REVIEW-01:
gsd-code-reviewMUST scope files to the phase using SUMMARY.md and git diff fallback - REQ-REVIEW-02: Review MUST support three depth levels:
quick,standard,deep - REQ-REVIEW-03: Findings MUST be severity-classified: Critical, Warning, Info
- REQ-REVIEW-04:
gsd-code-review --fixMUST read REVIEW.md and fix Critical + Warning findings by default - REQ-REVIEW-05: Each fix MUST be committed atomically with a descriptive message
- REQ-REVIEW-06:
--autoflag MUST enable fix + re-review iteration loop, capped at 3 iterations - REQ-REVIEW-07: Feature MUST be gated by
workflow.code_reviewconfig flag
Config:
| Setting | Type | Default | Description |
|---|---|---|---|
workflow.code_review |
boolean | true |
Enable code review commands |
workflow.code_review_depth |
string | standard |
Default review depth: quick, standard, or deep |
workflow.code_review_depth_overrides |
array | [] |
Ordered { paths, depth } rules that escalate depth for directories matched by path prefix against the changed-file set (#2554). See below. |
Path-scoped code review depth overrides
workflow.code_review_depth_overrides matches rules against the review's changed-file set by whole-segment directory-path prefix — src/auth matches src/auth/token.ts and src/auth itself, never src/authfoo/x.ts or docs/src/auth/x.ts — and is case-sensitive, following git.
Escalation is whole-review, not per-file: depth is a single scalar handed to the reviewer agent, not a per-file setting, so the strongest matching tier across the whole rule set applies to every file in the review — a sensitive file is never reviewed shallowly because it shared a review with an unrelated one.
v1 supports directory-prefix matching only, not glob syntax: no glob engine (minimatch, picomatch, fast-glob) exists in this project and none was added for this feature. A path containing * or ? (e.g. src/auth/**) is a configuration error rather than a silent near-miss, because accepting it as sugar for a prefix would make unsupported patterns look armed when they match nothing. Every use case in the issue is expressible as a directory prefix. See Scope code review depth by path for the resolution order, error table, and a worked example.
94. Socratic Exploration
Command: /gsd-explore [topic]
Purpose: Guide a developer through exploring an idea via Socratic probing questions before committing to a plan. Routes outputs to the appropriate GSD artifact: notes, todos, seeds, research questions, requirements updates, or a new phase.
Requirements:
- REQ-EXPLORE-01: Exploration MUST use Socratic probing — ask questions before proposing solutions
- REQ-EXPLORE-02: Session MUST offer to route outputs to the appropriate GSD artifact
- REQ-EXPLORE-03: An optional topic argument MUST prime the first question
- REQ-EXPLORE-04: Exploration MUST optionally spawn a research agent for technical feasibility
- REQ-EXPLORE-05: A research pass MUST disposition each surfaced claim (admit / refute / abstain) and route every abstention to a visible Unresolved Ledger — never smoothing an ungrounded claim into the narrative as confident prose
95. Safe Undo
Command: /gsd-undo --last N | --phase NN | --plan NN-MM
Purpose: Roll back GSD phase or plan commits safely using the phase manifest and git log, with dependency checks and a hard confirmation gate before any revert is applied.
Requirements:
- REQ-UNDO-01:
--phasemode MUST identify all commits for the phase via manifest and git log fallback - REQ-UNDO-02:
--planmode MUST identify all commits for a specific plan - REQ-UNDO-03:
--last Nmode MUST display recent GSD commits for interactive selection - REQ-UNDO-04: System MUST check for dependent phases/plans before reverting
- REQ-UNDO-05: A confirmation gate MUST be shown before any git revert is executed
96. Plan Import
Command: /gsd-import --from <filepath>
Purpose: Ingest an external plan file into the GSD planning system with conflict detection against PROJECT.md decisions, converting it to a valid GSD PLAN.md and validating it through the plan-checker.
Requirements:
- REQ-IMPORT-01: Importer MUST detect conflicts between the external plan and existing PROJECT.md decisions
- REQ-IMPORT-02: All detected conflicts MUST be presented to the user for resolution before writing
- REQ-IMPORT-03: Imported plan MUST be written as a valid GSD PLAN.md format
- REQ-IMPORT-04: Written plan MUST pass
gsd-plan-checkervalidation
97. Rapid Codebase Scan
Command: /gsd-map-codebase --fast [--focus tech|arch|quality|concerns]
Purpose: Lightweight alternative to /gsd-map-codebase that spawns a single mapper agent for one or two combined focus areas, producing targeted output in .planning/codebase/ without the overhead of 4 parallel agents.
Requirements:
- REQ-SCAN-01: Scan MUST spawn exactly one mapper agent (not four parallel agents)
- REQ-SCAN-02: Focus area MUST be one of:
tech,arch,quality,concerns, or the combinedtech+archshorthand (default:tech+arch); combined focus runs as a single agent covering both areas in one pass - REQ-SCAN-03: Output MUST be written to
.planning/codebase/in the same format as/gsd-map-codebase
98. Autonomous Audit-to-Fix
Command: /gsd-audit-fix [--source <audit>] [--severity high|medium|all] [--max N] [--dry-run]
Purpose: End-to-end pipeline that runs an audit, classifies findings as auto-fixable vs. manual-only, then autonomously fixes auto-fixable issues with test verification and atomic commits.
Requirements:
- REQ-AUDITFIX-01: Findings MUST be classified as auto-fixable or manual-only before any changes
- REQ-AUDITFIX-02: Each fix MUST be verified with tests before committing
- REQ-AUDITFIX-03: Each fix MUST be committed atomically
- REQ-AUDITFIX-04:
--dry-runMUST show classification table without applying any fixes - REQ-AUDITFIX-05:
--max NMUST limit the number of fixes applied in one run (default: 5)
99. Improved Prompt Injection Scanner
Hook: gsd-prompt-guard.js, gsd-read-injection-scanner.js
Script: scripts/prompt-injection-scan.sh, scripts/base64-scan.sh
Purpose: Defense-in-depth detection of prompt injection attempts in planning artifacts and ingested content. Live hooks inline their own pattern subsets for hook independence (they do not import from security.cts). The CI scanner (scanForInjection in security.cts) provides a centralized engine for codebase-wide scanning in tests.
Requirements:
-
REQ-SCAN-INJ-01: Live hooks MUST detect invisible Unicode characters (zero-width spaces, soft hyphens, Unicode tag block U+E0000–E007F)
-
REQ-SCAN-INJ-02: Live hooks MUST detect known injection patterns (instruction override, role manipulation, system-prompt extraction, fake message boundaries). Base64-decode scanning is a CI-time control (
scripts/base64-scan.sh), not a live hook — live hooks match a base64-exfiltration phrase regex only, they do not decode. -
REQ-SCAN-INJ-03:
Scanner MUST apply entropy analysis— Entropy analysis (scanEntropyAnomalies) was removed in #2198 as dead code (zero production callers; live hooks do not perform entropy analysis). This requirement is deferred pending a maintainable live implementation. -
REQ-SCAN-INJ-04: Scanner MUST remain advisory-only — detection is logged, not blocking
-
REQ-SCAN-INJ-05: A scanner that could not establish its file list MUST NOT report clean (#3908). The CI scanners (
prompt-injection-scan.sh,base64-scan.sh,secret-scan.sh) distinguish four outcomes rather than collapsing them into exit 0:Outcome Exit Meaning scanned, no findings 0files were in scope and none matched findings 1the scan's own verdict nothing in scope NO_INPUTthe diff resolved and was genuinely empty — e.g. a docs-only PR could not scan UNAVAILABLEthe file list was never established: a bad ref, no repository, or a repository with no commits Codes come from the exit-code registry (ADR-3889), sourced from
gsd-core/bin/shared/exit-codes.sh, never written into the scripts. Every one is non-zero, so a caller writtenif ! scanner; thenbehaves identically for a clean scan and trips for everything else — this can turn a false green red, never a red green..github/workflows/security-scan.ymltreats nothing in scope as a pass and could not scan as a failure; previously the latter passed silently, having scanned nothing.
100. Stall Detection in Plan-Phase
Command: /gsd-plan-phase
Purpose: Detect when the planner revision loop has stalled — producing the same output across multiple iterations — and break the cycle by escalating to a different strategy or exiting with a clear diagnostic.
Requirements:
- REQ-STALL-01: Revision loop MUST detect identical plan output across consecutive iterations
- REQ-STALL-02: On stall detection, system MUST escalate strategy before retrying
- REQ-STALL-03: Maximum stall retries MUST be bounded (capped at the existing max 3 iterations)
101. Hard Stop Safety Gates in /gsd-progress --next
Command: /gsd-progress --next
Purpose: Prevent /gsd-progress --next from entering runaway loops by adding hard stop safety gates and a consecutive-call guard that interrupts autonomous chaining when repeated identical steps are detected.
Requirements:
- REQ-NEXT-GATE-01:
/gsd-progress --nextMUST track consecutive same-step calls - REQ-NEXT-GATE-02: On repeated same-step, system MUST present a hard stop gate to the user
- REQ-NEXT-GATE-03: User MUST explicitly confirm to continue past a hard stop gate
102. Adaptive Model Preset
Config: model_profile: "adaptive"
Purpose: Role-based model assignment that automatically selects the appropriate model tier based on the current agent's role, rather than applying a single tier to all agents.
Requirements:
- REQ-ADAPTIVE-01:
adaptivepreset MUST assign model tiers based on agent role (planner → quality tier, executor → balanced tier, etc.) - REQ-ADAPTIVE-02:
adaptiveMUST be selectable via/gsd-config --profile adaptive
103. Post-Merge Hunk Verification
Command: /gsd-update --reapply
Purpose: After applying local patches post-update, verify that all hunks were actually applied by comparing the expected patch content against the live filesystem. Surface any dropped or partial hunks immediately rather than silently accepting incomplete merges.
Requirements:
- REQ-PATCH-VERIFY-01: Reapply-patches MUST verify each hunk was applied after the merge
- REQ-PATCH-VERIFY-02: Dropped or partial hunks MUST be reported to the user with file and line context
- REQ-PATCH-VERIFY-03: Verification MUST run after all patches are applied, not per-patch
v1.35.0 Features
104. New Runtime Support (Cline, CodeBuddy, Qwen Code)
Part of: npx @opengsd/gsd-core
Purpose: Extend GSD installation to Cline, CodeBuddy, and Qwen Code runtimes.
Requirements:
- REQ-CLINE-02: Cline install MUST write
.clinerulesto~/.cline/(global) or./.cline/(local). No custom slash commands — rules-based integration only. Flag:--cline. - REQ-CODEBUDDY-01: CodeBuddy install MUST deploy skills to
~/.codebuddy/skills/gsd-*/SKILL.md(emitteduser-invocable: false),/gsd-*slash commands to~/.codebuddy/commands/gsd-*.md, and subagents to~/.codebuddy/agents/gsd-*.md. The commands surface is the sole/menu entry point. Nomcp.jsonis written (gsd ships no MCP server). Flag:--codebuddy. - REQ-QWEN-01: Qwen Code install MUST deploy skills to
~/.qwen/skills/gsd-*/SKILL.md, following the open standard used by Claude Code 2.1.88+.QWEN_CONFIG_DIRenv var overrides the default path. Flag:--qwen.
Runtime summary:
| Runtime | Install Format | Config Path | Flag |
|---|---|---|---|
| Cline | .clinerules |
~/.cline/ or ./.cline/ |
--cline |
| CodeBuddy | Skills (SKILL.md) |
~/.codebuddy/skills/ |
--codebuddy |
| Qwen Code | Skills (SKILL.md) |
~/.qwen/skills/ |
--qwen |
105. GSD-2 Reverse Migration
Command: /gsd-import --from-gsd2 [--dry-run] [--force] [--path <dir>]
Purpose: Migrate a project from GSD-2 format (.gsd/ directory with Milestone→Slice→Task hierarchy) back to the v1 .planning/ format, restoring full compatibility with all GSD v1 commands.
Requirements:
- REQ-FROM-GSD2-01: Importer MUST read
.gsd/from the specified or current directory - REQ-FROM-GSD2-02: Milestone→Slice hierarchy MUST be flattened to sequential phase numbers (M001/S01→phase 01, M001/S02→phase 02, M002/S01→phase 03, etc.)
- REQ-FROM-GSD2-03: System MUST guard against overwriting an existing
.planning/directory without--force - REQ-FROM-GSD2-04:
--dry-runMUST preview all changes without writing any files - REQ-FROM-GSD2-05: Migration MUST produce
PROJECT.md,REQUIREMENTS.md,ROADMAP.md,STATE.md, and sequential phase directories
Flags:
| Flag | Description |
|---|---|
--dry-run |
Preview migration output without writing files |
--force |
Overwrite an existing .planning/ directory |
--path <dir> |
Specify the GSD-2 root directory |
106. AI Integration Phase Wizard
Command: /gsd-ai-integration-phase [N]
Purpose: Guide developers through selecting, integrating, and planning evaluation for AI/LLM capabilities in a project phase. Produces a structured AI-SPEC.md that feeds into planning and verification.
Requirements:
- REQ-AISPEC-01: Wizard MUST present an interactive decision matrix covering framework selection, model choice, and integration approach
- REQ-AISPEC-02: System MUST surface domain-specific failure modes and eval criteria relevant to the project type
- REQ-AISPEC-03: System MUST spawn 3 parallel specialist agents: domain-researcher, framework-selector, and eval-planner
- REQ-AISPEC-04: Output MUST produce
{phase}-AI-SPEC.mdwith framework recommendation, implementation guidance, and evaluation strategy
Produces: {phase}-AI-SPEC.md in the phase directory
107. AI Eval Review
Command: /gsd-eval-review [N]
Purpose: Retroactively audit an executed AI phase's evaluation coverage against the AI-SPEC.md plan. Identifies gaps between planned and implemented evaluation before the phase is closed.
Requirements:
- REQ-EVALREVIEW-01: Review MUST read
AI-SPEC.mdfrom the specified phase - REQ-EVALREVIEW-02: Each eval dimension MUST be scored as COVERED, PARTIAL, or MISSING
- REQ-EVALREVIEW-03: Output MUST include findings, gap descriptions, and remediation guidance
- REQ-EVALREVIEW-04:
EVAL-REVIEW.mdMUST be written to the phase directory
Produces: {phase}-EVAL-REVIEW.md with scored eval dimensions, gap analysis, and remediation steps
v1.36.0 Features
108. Plan Bounce
Command: /gsd-plan-phase N --bounce
Purpose: After plans pass the checker, optionally refine them through an external script (a second AI, a linter, a custom validator). The bounce step backs up each plan, runs the script, validates YAML frontmatter integrity on the result, re-runs the plan checker, and restores the original if anything fails.
Requirements:
- REQ-BOUNCE-01:
--bounceflag orworkflow.plan_bounce: trueactivates the step;--skip-bouncealways disables it - REQ-BOUNCE-02:
workflow.plan_bounce_scriptmust point to a valid executable; missing script produces a warning and skips - REQ-BOUNCE-03: Each plan is backed up to
*-PLAN.pre-bounce.mdbefore the script runs - REQ-BOUNCE-04: Bounced plans with broken YAML frontmatter or that fail the plan checker are restored from backup
- REQ-BOUNCE-05:
workflow.plan_bounce_passes(default: 2) controls how many refinement passes the script receives
Configuration: workflow.plan_bounce, workflow.plan_bounce_script, workflow.plan_bounce_passes
109. External Code Review Command
Command: /gsd-ship (enhanced)
Purpose: Before the manual review step in /gsd-ship, automatically run an external code review command if configured. The command receives the diff and phase context via stdin and returns a JSON verdict (APPROVED or REVISE). Falls through to the existing manual review flow regardless of outcome.
Requirements:
- REQ-EXTREVIEW-01:
workflow.code_review_commandmust be set to a command string; null means skip - REQ-EXTREVIEW-02: Diff is generated against
BASE_BRANCHwith--statsummary included - REQ-EXTREVIEW-03: Review prompt is piped via stdin (never shell-interpolated)
- REQ-EXTREVIEW-04: 120-second timeout; stderr captured on failure
- REQ-EXTREVIEW-05: JSON output parsed for
verdict,confidence,summary,issuesfields
Configuration: workflow.code_review_command
110. Cross-AI Execution Delegation
Command: /gsd-execute-phase N --cross-ai
Purpose: Delegate individual plans to an external AI runtime for execution. Plans with cross_ai: true in their frontmatter (or all plans when --cross-ai is used) are sent to the configured command via stdin. Successfully handled plans are removed from the normal executor queue.
Requirements:
- REQ-CROSSAI-01:
--cross-aiforces all plans through cross-AI;--no-cross-aidisables it - REQ-CROSSAI-02:
workflow.cross_ai_execution: trueand plan frontmattercross_ai: truerequired for per-plan activation - REQ-CROSSAI-03: Task prompt is piped via stdin to prevent injection
- REQ-CROSSAI-04: Dirty working tree produces a warning before execution
- REQ-CROSSAI-05: On failure, user chooses: retry, skip (fall back to normal executor), or abort
Configuration: workflow.cross_ai_execution, workflow.cross_ai_command, workflow.cross_ai_timeout
111. Architectural Responsibility Mapping
Command: /gsd-plan-phase (enhanced research step)
Purpose: During phase research, the phase-researcher now maps each capability to its architectural tier owner (browser, frontend server, API, CDN/static, database). The planner cross-references tasks against this map, and the plan-checker enforces tier compliance as Dimension 7c.
Requirements:
- REQ-ARM-01: Phase researcher produces an Architectural Responsibility Map table in RESEARCH.md (Step 1.5)
- REQ-ARM-02: Planner sanity-checks task-to-tier assignments against the map
- REQ-ARM-03: Plan checker validates tier compliance as Dimension 7c (WARNING for general mismatches, BLOCKER for security-sensitive ones)
Produces: ## Architectural Responsibility Map section in {phase}-RESEARCH.md
112. Extract Learnings
Command: /gsd-extract-learnings N
Purpose: Extract structured knowledge from completed phase artifacts. Reads PLAN.md and SUMMARY.md (required) plus VERIFICATION.md, UAT.md, and STATE.md (optional) to produce four categories of learnings: decisions, lessons, patterns, and surprises. Optionally captures each item to an external knowledge base via capture_thought tool.
Requirements:
- REQ-LEARN-01: Requires PLAN.md and SUMMARY.md; exits with clear error if missing
- REQ-LEARN-02: Each extracted item includes source attribution (artifact and section)
- REQ-LEARN-03: If
capture_thoughttool is available, captures items withsource,project, andphasemetadata - REQ-LEARN-04: If
capture_thoughtis unavailable, completes successfully and logs that external capture was skipped - REQ-LEARN-05: Running twice overwrites the previous
LEARNINGS.md
Produces: {phase}-LEARNINGS.md with YAML frontmatter (phase, project, counts per category, missing_artifacts)
Optional integration — capture_thought: capture_thought is a convention, not a bundled tool. GSD does not ship one and does not require one. The workflow checks whether any MCP server in the current session exposes a tool named capture_thought and, if so, calls it once per extracted learning with the signature below. If no such tool is present, the step is skipped silently and LEARNINGS.md remains the primary output.
Expected tool signature:
capture_thought({
category: "decision" | "lesson" | "pattern" | "surprise",
phase: <phase_number>,
content: <learning_text>,
source: <artifact_name>
})
Users who run a memory / knowledge-base MCP server (for example, ExoCortex-style servers, claude-mem, or mem0-style servers) can implement this tool name to have learnings routed into their knowledge base automatically with project, phase, and source metadata. Everyone else can use /gsd-extract-learnings without any extra setup — the LEARNINGS.md artifact is the feature.
With features.global_learnings: true, phase completion runs the extraction for the just-completed phase automatically and copies the artifact to the global store at ~/.gsd/knowledge/ (#3683) — extraction and copy failures never block completion. With the gate off (the default), extraction stays fully manual.
114. Context-Window-Aware Prompt Thinning
Purpose: Reduce static prompt overhead by ~40% for models with context windows under 200K tokens. Extended examples and anti-pattern lists are extracted from agent definitions into reference files loaded on demand via @ required_reading.
Requirements:
- REQ-THIN-01: When
CONTEXT_WINDOW < 200000, executor and planner agent prompts omit inline examples - REQ-THIN-02: Extracted content lives in
references/executor-examples.mdandreferences/planner-antipatterns.md - REQ-THIN-03: Standard (200K-500K) and enriched (500K+) tiers are unaffected
- REQ-THIN-04: Core rules and decision logic remain inline; only verbose examples are extracted
Reference files: executor-examples.md, planner-antipatterns.md
115. Configurable CLAUDE.md Path
Purpose: Allow projects to store their CLAUDE.md in a non-root location. The claude_md_path config key controls where /gsd-profile-user and related commands write the generated CLAUDE.md file.
Requirements:
- REQ-CMDPATH-01:
claude_md_pathdefaults to./.claude/CLAUDE.md(a valid project-scoped memory location; changed from./CLAUDE.mdin v1.5 per #1098 so generated content does not pollute a hand-crafted repo-rootCLAUDE.md) - REQ-CMDPATH-02: Profile generation commands read the path from config and write to the specified location
- REQ-CMDPATH-03: Relative paths are resolved from the project root
- REQ-CMDPATH-04:
generate-claude-mdnever overwrites an existing instruction file that lacks GSD section markers (a hand-crafted file) unless--forceis passed
Configuration: claude_md_path
116. TDD Pipeline Mode
Purpose: Opt-in TDD (red-green-refactor) as a first-class phase execution mode. When enabled, the planner aggressively selects type: tdd for eligible tasks and the executor enforces RED/GREEN/REFACTOR gate sequence with fail-fast on unexpected GREEN before RED.
Requirements:
- REQ-TDD-01:
workflow.tdd_modeconfig key (boolean, defaultfalse) - REQ-TDD-02: When enabled, planner applies TDD heuristics from
references/tdd.mdto all eligible tasks (business logic, APIs, validations, algorithms, state machines) - REQ-TDD-03: Executor enforces gate sequence for
type: tddplans — RED commit (test(...)) must precede GREEN commit (feat(...)) - REQ-TDD-04: Executor fails fast if tests pass unexpectedly during RED phase (feature already exists or test is wrong)
- REQ-TDD-05: End-of-phase collaborative review checkpoint verifies gate compliance across all TDD plans (advisory, non-blocking)
- REQ-TDD-06: Gate violations surfaced in SUMMARY.md under
## TDD Gate Compliancesection
Configuration: workflow.tdd_mode
Reference files: tdd.md, checkpoints.md
v1.37.0 Features
117. Spike Command
Command: /gsd-spike [idea] [--quick]
Purpose: Run 2–5 focused feasibility experiments before committing to an implementation approach. Each experiment uses Given/When/Then framing, produces executable code, and returns a VALIDATED / INVALIDATED / PARTIAL verdict. Companion /gsd-spike --wrap-up packages findings into a project-local skill.
Requirements:
- REQ-SPIKE-01: Each experiment MUST produce a Given/When/Then hypothesis before any code is written
- REQ-SPIKE-02: Each experiment MUST include working code or a minimal reproduction
- REQ-SPIKE-03: Each experiment MUST return one of: VALIDATED, INVALIDATED, or PARTIAL verdict with evidence
- REQ-SPIKE-04: Results MUST be stored in
.planning/spikes/NNN-experiment-name/with a README and MANIFEST.md - REQ-SPIKE-05:
--quickflag skips intake conversation and uses the argument text as the experiment direction - REQ-SPIKE-06:
/gsd-spike --wrap-upMUST package findings into.claude/skills/spike-findings-[project]/
Produces:
| Artifact | Description |
|---|---|
.planning/spikes/NNN-name/README.md |
Hypothesis, experiment code, verdict, and evidence |
.planning/spikes/MANIFEST.md |
Index of all spikes with verdicts |
.claude/skills/spike-findings-[project]/ |
Packaged findings (via /gsd-spike --wrap-up) |
118. Sketch Command
Command: /gsd-sketch [idea] [--quick] [--text]
Purpose: Explore design directions through throwaway HTML mockups before committing to implementation. Produces 2–3 interactive variants per design question, all viewable directly in a browser with no build step. Companion /gsd-sketch --wrap-up packages winning decisions into a project-local skill.
Requirements:
- REQ-SKETCH-01: Each sketch MUST answer one specific visual design question
- REQ-SKETCH-02: Each sketch MUST include 2–3 meaningfully different variants in a single
index.htmlwith tab navigation - REQ-SKETCH-03: All interactive elements (hover, click, transitions) MUST be functional
- REQ-SKETCH-04: Sketches MUST use real-ish content, not lorem ipsum
- REQ-SKETCH-05: A shared
themes/default.cssMUST provide CSS variables adapted to the agreed aesthetic - REQ-SKETCH-06:
--quickflag skips mood intake;--textflag replacesAskUserQuestionwith numbered lists for non-Claude runtimes - REQ-SKETCH-07: The winning variant MUST be marked in the README frontmatter and with a ★ in the HTML tab
- REQ-SKETCH-08:
/gsd-sketch --wrap-upMUST package winning decisions into.claude/skills/sketch-findings-[project]/
Produces:
| Artifact | Description |
|---|---|
.planning/sketches/NNN-name/index.html |
2–3 interactive HTML variants |
.planning/sketches/NNN-name/README.md |
Design question, variants, winner, what to look for |
.planning/sketches/themes/default.css |
Shared CSS theme variables |
.planning/sketches/MANIFEST.md |
Index of all sketches with winners |
.claude/skills/sketch-findings-[project]/ |
Packaged decisions (via /gsd-sketch --wrap-up) |
119. Agent Size-Budget Enforcement
Purpose: Keep agent prompt files lean with tiered line-count limits enforced in CI. Oversized agents are caught before they bloat context windows in production.
Requirements:
- REQ-BUDGET-01:
agents/gsd-*.mdfiles are classified into three tiers: XL (≤ 1 600 lines), Large (≤ 1 000 lines), Default (≤ 500 lines) - REQ-BUDGET-02: Tier assignment is declared in the file's YAML frontmatter (
size: xl | large | default) - REQ-BUDGET-03:
tests/agent-size-budget.test.cjsenforces limits and fails CI on violation - REQ-BUDGET-04: Files without a
sizefrontmatter key default to the Default (500-line) limit
Test file: tests/agent-size-budget.test.cjs
120. Shared Boilerplate Extraction
Purpose: Reduce duplication across agents by extracting two common boilerplate blocks into shared reference files loaded on demand. Keeps agent files within size budget and makes boilerplate updates a single-file change.
Requirements:
- REQ-BOILER-01: Mandatory-initial-read instructions extracted to
references/mandatory-initial-read.md - REQ-BOILER-02: Project-skills-discovery instructions extracted to
references/project-skills-discovery.md - REQ-BOILER-03: Agents that previously inlined these blocks MUST now reference them via
@required_reading
Reference files: references/mandatory-initial-read.md, references/project-skills-discovery.md
121. Knowledge Graph Integration
Purpose: Build, query, and inspect a lightweight knowledge graph of the project in .planning/graphs/. Opt-in per project. Exposed as the /gsd-graphify user-facing command and the gsd-tools.cjs graphify … programmatic verb family. Complements /gsd-map-codebase --query (snapshot-oriented) with a graph-oriented view of nodes and edges across commands, agents, workflows, and phases.
Requirements:
- REQ-GRAPH-01: Opt-in via
graphify.enabled: truein.planning/config.json. When disabled,/gsd-graphifyprints an activation hint and stops without writing. - REQ-GRAPH-02: Slash-command
/gsd-graphifyexposes subcommandsbuild,query <term>,status,diff. The programmatic CLInode gsd-tools.cjs graphify …additionally exposessnapshot, which is also invoked automatically as the final step ofgraphify build. - REQ-GRAPH-03: Build runs within the configurable
graphify.build_timeout(seconds); exceeding the timeout aborts cleanly without leaving a partial graph. - REQ-GRAPH-04:
graphify.cjsfalls back tograph.linkswhengraph.edgesis absent so older graph artifacts keep rendering. - REQ-GRAPH-05: Graphify is invoked through
gsd-tools.cjs graphify ...command handlers. - REQ-GRAPH-06: The knowledge-graph location is configurable via
graphify.graph_path(issue #1825) so one umbrella-level cross-repo graph can serve multiple sibling projects;query/status/diffread the configured graph (relative to project root), with a byte-identical.planning/graphs/default when unset.
Configuration: graphify.enabled, graphify.build_timeout, graphify.graph_path
Reference files: commands/gsd/graphify.md, bin/lib/graphify.cjs
v1.40.0 Features
122. Skill Surface Consolidation
Purpose: Cut the eager skill-listing overhead by folding 31 micro-skills into 4 new grouped parents and 6 existing parents that absorb sub-operations as flags. Zero functional loss — every removed micro-skill's behavior survives via a flag on a consolidated parent. After consolidation, commands/gsd/*.md ships 60 sub-skills (plus 6 namespace meta-skills, see #123).
Requirements:
- REQ-CONSOLIDATE-01: Four new grouped skills replace clusters of micro-skills:
/gsd-capture— folds add-todo (default), note (--note), add-backlog (--backlog), plant-seed (--seed), check-todos (--list)/gsd-phase— folds add-phase (default), insert-phase (--insert), remove-phase (--remove), edit-phase (--edit)/gsd-config— folds settings-advanced (--advanced), settings-integrations (--integrations), set-profile (--profile)/gsd-workspace— folds new-workspace (--new), list-workspaces (--list), remove-workspace (--remove)
- REQ-CONSOLIDATE-02: Six existing parents absorb wrap-up / sub-operations as flags:
/gsd-update --sync,/gsd-update --reapply,/gsd-sketch --wrap-up,/gsd-spike --wrap-up,/gsd-map-codebase --fast,/gsd-map-codebase --query,/gsd-code-review --fix,/gsd-progress --do,/gsd-progress --next. - REQ-CONSOLIDATE-03:
/gsd-nextis not the retired workflow-advance command; it is reserved for the state-aware smart-entry launcher. Workflow advancement remains under/gsd-progress --next. - REQ-CONSOLIDATE-04: Deleted micro-skill slash forms (the bare
gsd-add-todo,gsd-add-backlog,gsd-plant-seed,gsd-check-todos,gsd-add-phase,gsd-insert-phase,gsd-remove-phase,gsd-edit-phase,gsd-new-workspace,gsd-list-workspaces,gsd-remove-workspace,gsd-settings-advanced,gsd-settings-integrations,gsd-set-profile,gsd-sketch-wrap-up,gsd-spike-wrap-up,gsd-reapply-patches,gsd-code-review-fix, …) MUST resolve to "Unknown command" — no shadow stubs. - REQ-CONSOLIDATE-05:
autonomous.mdinvokes/gsd-code-review --fix(was previously calling the deletedgsd-code-review-fix).
Reference issue: #2790
123. Namespace Meta-Skills (Two-Stage Routing)
Purpose: Replace the flat eager skill listing with a two-stage hierarchical routing layer. The model sees 6 namespace routers instead of 86 entries, selects a namespace, then routes to the sub-skill. Descriptions use pipe-separated keyword tags (≤ 60 chars) for routing density.
Commands:
/gsd-workflow— phase pipeline router (discuss / plan / execute / verify / phase / progress / next)/gsd-project— project lifecycle (milestones, audits, summary)/gsd-quality— quality gates (code review, debug, audit, security, eval, ui)/gsd-context— codebase intelligence (map, graphify, docs, learnings)/gsd-manage— config / workspace / workstreams / thread / update / ship / inbox/gsd-ideate— exploration & capture (explore, sketch, spike, spec, capture)
Token cost:
| Entries | Approx tokens | |
|---|---|---|
| Pre-1.40 full install | 86 | ~2,150 |
| Namespace meta-skills | 6 | ~120 |
Requirements:
- REQ-NS-01: Six
commands/gsd/ns-*.mdnamespace routers ship with pipe-separated keyword-tag descriptions (≤ 60 chars). - REQ-NS-02: Existing sub-skills are unchanged and still invocable directly — namespace skills are additive, not a replacement for direct slash forms.
- REQ-NS-03: The body of each namespace router contains a routing table that maps user intent to the correct concrete sub-skill on the post-#2790 consolidated surface.
- REQ-NS-04: Tests validate namespace files exist, include matching command
requires, and reference only existing sub-skill files.
Reference issue: #2792
124. Context-Window Utilization Guard
Command: /gsd-health --context
Purpose: Quality guard against context-window saturation. Two thresholds: 60 % utilization warns ("consider /gsd-thread"), 70 % is critical ("reasoning quality may degrade"; matches the fracture-point per recent context-attention research).
Requirements:
- REQ-CTX-GUARD-01:
/gsd-health --contextprints a structured status line with current utilization, threshold tier (ok/warn/critical), and a remediation suggestion. - REQ-CTX-GUARD-02: The same triage is exposed as
gsd-tools.cjs validate context --tokens-used <int> --context-window <int>— a structured envelope for status-line and hook callers (#125). Both flags are required; the handler returns the same{ percent, state }envelope as the pure classifier in REQ-CTX-GUARD-03. - REQ-CTX-GUARD-03: The classifier (
bin/lib/context-utilization.cjs) is pure: input(tokensUsed, contextWindow), output{ percent, state }. Easy to unit-test, easy to reuse from any caller.
Reference issue: #2792
125. Phase-Lifecycle Status-Line Read-Side
Purpose: Surface phase orchestration state on the status-line. parseStateMd() reads four new STATE.md frontmatter fields and formatGsdState() renders in-flight, idle, and progress scenes. Write-side wiring follows in a later RC.
Requirements:
- REQ-LIFECYCLE-01:
parseStateMd()reads four optional fields:active_phase— phase number when an orchestrator is in flightnext_action— recommended next command when idlenext_phases— YAML flow array of next phase numbersprogress— nestedtotal_phases/completed_phases/percentblock
- REQ-LIFECYCLE-02:
formatGsdState()checks the lifecycle fields in priority order and emits the first matching scene (Phase active → Idle next-recommended → Milestone complete → Default fallback). - REQ-LIFECYCLE-03: All four fields default to undefined; existing STATE.md files render byte-for-byte identically.
Reference issue: #2833 — see docs/STATE-MD-LIFECYCLE.md for the full field reference and rendering rules.
v1.41.0 Features
126. Per-Phase-Type Model Selection
Purpose: Express model tuning at the phase level (planning, research, execution, verification) without learning the full agent taxonomy. Sits between per-agent model_overrides (precise, verbose) and the global model_profile tier (coarse, uniform).
Config key: models in .planning/config.json
Phase-type slots:
| Slot | Agents assigned |
|---|---|
planning |
gsd-planner, gsd-roadmapper, gsd-pattern-mapper |
discuss |
gsd-assumptions-analyzer |
research |
gsd-phase-researcher, gsd-project-researcher, gsd-research-synthesizer, gsd-codebase-mapper, gsd-ui-researcher |
execution |
gsd-executor, gsd-debugger, gsd-doc-writer |
verification |
gsd-verifier, gsd-plan-checker, gsd-integration-checker, gsd-nyquist-auditor, gsd-ui-checker, gsd-ui-auditor, gsd-doc-verifier, gsd-code-reviewer |
completion |
(reserved for future subagent) |
Accepted values: "opus" / "sonnet" / "haiku" / "inherit"
Resolution precedence (highest → lowest):
1. model_overrides[<agent>]
2. dynamic_routing.tier_models[<tier>] (when enabled)
3. models[<phase_type>] (this feature)
4. model_profile
5. Runtime default
Requirements:
- REQ-PHASE-MODELS-01: Six named
models.*slots accepted byconfig-schema.cjsandconfig-schema.ts;config-setrejects unknown phase-types. - REQ-PHASE-MODELS-02: Configs without a
modelsblock behave byte-for-byte identically to pre-v1.41 behavior. - REQ-PHASE-MODELS-03:
discussandcompletionare accepted by the schema for forward compatibility; setting them today is a no-op until a subagent maps to each.
Reference issue: #3023
127. Dynamic Routing with Failure-Tier Escalation
Purpose: Pay for the cheap tier by default; escalate to a more capable model automatically when the orchestrator detects a soft failure (verification inconclusive, plan-check FLAG, etc.).
Config key: dynamic_routing in .planning/config.json
Behavior:
enabled: false(default) — feature is off; all agents use the precedence chain unchanged.enabled: true— the resolver pickstier_models[default_tier]for the first spawn and escalates one tier up on orchestrator-detected soft failure, capped bymax_escalations.
Composition: model_overrides always wins; dynamic_routing.tier_models[<tier>] resolves above models.<phase_type> and model_profile.
Requirements:
- REQ-DYNROUTE-01:
dynamic_routing.enabledacts as a master switch; whenfalseor block is absent, zero behavior change. - REQ-DYNROUTE-02: New resolver
resolveModelForTier(cwd, agent, attempt)incore.cjsis the single call-site for orchestrator integration. - REQ-DYNROUTE-03:
max_escalationscaps the escalation chain to prevent runaway cost.
Reference issue: #3024
128. Update Banner Opt-In
Purpose: Surface update availability to users who have declined or bypassed the GSD statusline, without requiring the statusline.
Behavior:
- At install time, if the installer detects no GSD statusline, it offers an opt-in
SessionStarthook. - The hook reads the existing
~/.cache/gsd/gsd-update-check.jsoncache — the same cache used by the statusline — and prints a banner only when an update is available. - Silent when up-to-date.
- Failure diagnostics rate-limited to once per 24 h.
- Cleanly removed by
npx @opengsd/gsd-core --uninstall.
Requirements:
- REQ-BANNER-01: Banner does not install without explicit opt-in.
- REQ-BANNER-02: No additional network requests — reuses the existing background update-check cache.
- REQ-BANNER-03: Uninstall path removes the banner hook.
Reference issue: #2795
129. Issue-Driven Orchestration Guide
Purpose: Document a recipe for driving the full GSD workflow from a GitHub / Linear / Jira issue, mapping tracker-centric concepts onto existing GSD primitives.
Document: docs/issue-driven-orchestration.md
Covered workflow:
- Create an isolated workspace per issue (
/gsd-workspace --new) - Run the manager dashboard to get oriented (
/gsd-manager) - Execute autonomously (
/gsd-autonomous) - Verify and review (
/gsd-verify-work,/gsd-review) - Ship and close the issue (
/gsd-ship)
No new commands or daemon process — purely a documentation artifact that maps existing primitives onto a tracker-driven workflow.
Reference issue: #2840
130. Graphify Commit-Based Staleness
Purpose: Surface whether the architecture graph was built from the current commit or an older one, complementing the existing mtime-based stale signal.
Command: /gsd-graphify status
New fields returned (graphify v0.7+ graphs):
| Field | Type | Description |
|---|---|---|
built_at_commit |
string | Commit SHA the graph was built from |
current_commit |
string | Current git HEAD |
commits_behind |
number | How many commits behind HEAD the graph is |
commit_stale |
boolean | null | true=stale, false=current, null=unavailable (pre-v0.7, non-git) |
Rendered output (when signal is available):
Source commit: abc1234 (3 commits behind HEAD)
Security: built_at_commit validated as 4–40 hex chars before reaching git — a hostile graph.json cannot inject dashed options into argv.
Fallback: pre-v0.7 graphs and non-git checkouts return commit_stale: null; callers fall back to the existing mtime-based stale flag. No behavior change for existing users.
Reference issue: #3170
v1.42.1 Features
132. Package Legitimacy Gate
Purpose: Stop hallucinated, suspicious, or slopsquatting package names before they reach a shell install command.
Behavior:
- Phase research writes a
## Package Legitimacy Audittable for recommended packages. - Packages verified only through search are treated as
[ASSUMED], not trusted. [SLOP]packages are removed from recommendations.- Plans that need
[ASSUMED]or suspicious packages add a human verification checkpoint. - Executor install failures stop for human verification instead of auto-trying similarly named packages.
Requirements:
- REQ-PKG-GATE-01: Research MUST record package registry, age, download/source signals, legitimacy verdict, and disposition.
- REQ-PKG-GATE-02: Planner MUST gate unverified or suspicious package installs before execution.
- REQ-PKG-GATE-03: Executor MUST NOT auto-substitute package names after failed package-manager installs.
Reference: v1.42.1 Release Notes
133. Skill Surface Budgeting
Purpose: Let users reduce installed skill and agent surface area when context budget matters.
Install profiles:
| Profile | Purpose |
|---|---|
core |
Minimal main-loop surface |
standard |
Core plus common phase-management commands |
full |
Complete surface; default |
Runtime control: /gsd-surface lists profile state and enables, disables, or resets skill clusters without reinstalling.
Requirements:
- REQ-SURFACE-01: Installer MUST resolve
--profile=<name>and persist the active profile in.gsd-profile. - REQ-SURFACE-02:
--minimaland--core-onlyMUST remain aliases for--profile=core. - REQ-SURFACE-03: Runtime surface state MUST persist outside the install profile marker.
Reference: ADR-0011
134. Installer Migrations
Purpose: Make runtime config cleanup explicit, auditable, and rollback-aware during installs and updates.
Capabilities:
- First-time baseline migration records managed files.
- Legacy stale-file cleanup uses ownership evidence before deleting or rewriting.
- User-owned artifacts are preserved.
- Ambiguous GSD-looking files block with a clear report instead of being silently overwritten.
- Migration plans support dry-run reporting and rollback protection.
Requirements:
- REQ-INSTALL-MIGRATION-01: Migration records MUST include metadata, install scope, and ownership evidence.
- REQ-INSTALL-MIGRATION-02: Destructive actions MUST fail closed when ownership is ambiguous.
- REQ-INSTALL-MIGRATION-03: Install failures MUST restore the pre-install state when rollback data exists.
Reference: Installer Migrations
135. Custom Ship PR Body Sections
Command: /gsd-ship
Config key: ship.pr_body_sections
Purpose: Add project-specific PRD-style sections to generated PR bodies without editing GSD workflow files.
Behavior: Configured sections append after the required Summary, Changes, Requirements Addressed, Verification, and Key Decisions sections. They can copy from artifact headings, render templates, or fall back to static text.
Requirements:
- REQ-SHIP-SECTIONS-01: Custom sections MUST NOT replace, remove, or reorder required PR sections.
- REQ-SHIP-SECTIONS-02: Unknown template tokens MUST be rejected by config validation.
- REQ-SHIP-SECTIONS-03: Disabled sections MUST stay in config without appearing in PR output.
Reference: Custom PR Body Sections
136. Review Default Reviewers
Command: /gsd-review
Config key: review.default_reviewers
Purpose: Let teams choose the default reviewer subset for no-flag /gsd-review runs.
Precedence:
explicit reviewer flags -> --all -> review.default_reviewers -> all detected reviewers
Requirements:
- REQ-REVIEW-DEFAULTS-01: Missing
review.default_reviewersMUST preserve the previous all-detected behavior. - REQ-REVIEW-DEFAULTS-02: Empty arrays MUST be rejected; remove the key to restore all-detected behavior.
- REQ-REVIEW-DEFAULTS-03: Known but unavailable reviewers MUST be skipped with diagnostics rather than hard-failing the run.
Reference: Configuration Reference
137. Fallow Structural Review Pre-Pass
Command: /gsd-code-review
Config keys: code_quality.fallow.*
Purpose: Add an optional structural analysis pass before the agent review.
Behavior: When enabled, GSD resolves a fallow binary, runs a bounded audit, writes FALLOW.json, and embeds structural findings in REVIEW.md.
Requirements:
- REQ-FALLOW-01: Fallow MUST be opt-in and disabled by default.
- REQ-FALLOW-02: Missing or failing fallow runs MUST produce clear diagnostics.
- REQ-FALLOW-03: Findings larger than the embed budget MUST be skipped with a warning, preserving the raw JSON artifact.
Reference: Configuration Reference
138. End-of-Phase Human Verification Mode
Config key: workflow.human_verify_mode
Purpose: Reduce mid-flight human checkpoint interruptions while preserving human verification requirements.
Behavior: The default "end-of-phase" mode embeds human checks into <verify><human-check> blocks for phase review. "mid-flight" restores blocking checkpoint:human-verify tasks.
Requirements:
- REQ-HUMAN-VERIFY-01:
checkpoint:decisionandcheckpoint:human-actionMUST remain blocking regardless of mode. - REQ-HUMAN-VERIFY-02: Human-needed verification MUST remain pending until the end-of-phase review resolves it.
- REQ-HUMAN-VERIFY-03: Configs without the key MUST use
"end-of-phase".
Reference: Checkpoints Reference
139. Quota and Rate-Limit Failure Classification
Command: /gsd-execute-phase
Purpose: Treat provider quota and rate-limit failures as wait-and-resume conditions, not normal executor failures.
Behavior: Agent output is classified for signals such as 429, rate limit, usage limit, RESOURCE_EXHAUSTED, and usage_limit_reached. Matching failures present a wait-for-reset recovery path.
Requirements:
- REQ-QUOTA-01: Quota failures MUST NOT offer immediate retry as the primary recovery.
- REQ-QUOTA-02: Classification MUST cover Claude, Copilot, Codex, and generic provider sentinels.
- REQ-QUOTA-03: Non-quota failures MUST continue through the normal execution failure path.
Reference: Provider Rate Limit Signals
140. Statusline Context Position
Config key: statusline.context_position
Purpose: Keep the context meter visible in narrow terminals.
Options:
| Value | Behavior |
|---|---|
"end" |
Default; render context meter near the line tail |
"front" |
Render context meter immediately after the model name |
Requirements:
- REQ-STATUSLINE-POS-01: Invalid values MUST be rejected by config validation.
- REQ-STATUSLINE-POS-02: Missing config MUST preserve existing end-position rendering.
Reference: Configuration Reference
141. Milestone Tag Creation Toggle
Command: /gsd-complete-milestone
Config key: git.create_tag
Purpose: Let projects with external release automation complete milestones without creating local git tags.
Behavior: git.create_tag: false skips milestone tag creation. The workflow still updates milestone artifacts and state.
Requirements:
- REQ-MILESTONE-TAG-01: Missing config MUST preserve automatic tag creation.
- REQ-MILESTONE-TAG-02: Existing tag collisions MUST fail clearly instead of overwriting tags.
- REQ-MILESTONE-TAG-03: Disabling tag creation MUST NOT skip milestone archival.
Reference: Configuration Reference
142. Structured JSON Error Mode
CLI: gsd-tools --json-errors
Purpose: Give automation callers stable machine-readable error envelopes.
Behavior: Commands that fail under --json-errors return structured ok: false payloads with error kind, message, command context, and exit mapping instead of prose-only stderr.
Requirements:
- REQ-JSON-ERRORS-01: Unknown commands, validation errors, timeouts, native failures, fallback failures, and internal errors MUST map to canonical error kinds.
- REQ-JSON-ERRORS-02: CLI exit code mapping MUST remain stable for automation callers.
- REQ-JSON-ERRORS-03: Human-readable output MUST remain the default when
--json-errorsis absent.
Reference: JSON Error Mode
143. UAT-Passed Predicate
CLI: node gsd-tools.cjs phase uat-passed <N> [--require-verification]
Purpose: Provide a runtime-neutral, automatable predicate that evaluates HUMAN-UAT results for a phase and returns a structured pass/fail verdict with full diagnostic detail.
Behavior: Locates *-UAT.md and optionally *-VERIFICATION.md files for the given phase, parses UAT test blocks (heading-block parser, column-0 result lines) with a markdown-aware stripper that removes false-positive contexts (YAML frontmatter, fenced code blocks, HTML comments, and blockquotes). Returns passed: true only when at least one check exists AND all checks pass AND no blockers — fail-closed, no vacuous pass. The --require-verification flag requires at least one *-VERIFICATION.md with an allowlisted passing status; the command fails without one.
Output envelope: { passed, uat_files[], verification_files[], checks[], blockers[], no_uat_artifacts, policy: { require_verification } }
| Field | Type | Description |
|---|---|---|
passed |
boolean |
true only when ≥1 check exists AND all passing AND no blockers |
uat_files |
string[] |
Filenames of *-UAT.md files evaluated |
verification_files |
string[] |
Filenames of *-VERIFICATION.md files evaluated |
checks[] |
{ file, test, name, result, passing }[] |
Per-item results from heading blocks |
blockers[] |
string[] |
Human-readable failure reasons (frontmatter, failing/missing items, policy, malformed markdown) |
no_uat_artifacts |
boolean |
true when no test items were parsed; passed is always false when true |
policy.require_verification |
boolean |
Whether --require-verification was active |
Requirements:
- REQ-UAT-PRED-01: The predicate MUST ignore result lines inside YAML frontmatter, fenced code blocks, HTML comments, and blockquotes.
- REQ-UAT-PRED-02:
passed: trueMUST require at least one check AND all checks passing AND no blockers (fail-closed, no vacuous pass). - REQ-UAT-PRED-03:
--require-verificationMUST cause the command to fail when no*-VERIFICATION.mdfile with an allowlisted passing status is found. - REQ-UAT-PRED-04:
blockers[]contains all human-readable failure reasons including frontmatter issues, policy violations, and malformed markdown — NOT limited to a subset ofchecks[]. - REQ-UAT-PRED-05: The module MUST be runtime-neutral (no runtime-specific env checks or exit shortcuts).
- REQ-UAT-PRED-06: A heading block with no column-0
result:line emitsresult:'missing'(blocker); test items are never silently dropped.
Reference: Phase Management Commands
144. Spec-Phase Edge-Completeness Probe
Command: /gsd-spec-phase
Purpose: Surface the omitted domain-boundary edges that silently invalidate a requirement — touching intervals, empty inputs, rounding ties, grapheme truncation — before they become production defects. Runs as Step 5.5 of spec-phase, after the ambiguity gate.
Behavior: For each SPEC requirement the probe classifies its data/behavior shape, then raises only the applicable categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency) via a relevance filter. Each raised category proposes one concrete candidate edge, which the author resolves to exactly one of four states:
| State | Meaning | Downstream effect |
|---|---|---|
covered |
An acceptance criterion handles the edge | Pass/fail line written into the SPEC Acceptance Criteria block; lifted into plan-phase must_haves.truths |
dismissed |
The edge cannot occur (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
backstop |
Intent recorded, needs a held-out/property-based test | Lifted into must_haves.truths as a non-inferable check |
unresolved |
Deferred | Soft-gates the spec; row stamped ⚠ Edge unresolved — planner must treat as assumption |
When a requirement's prose matches no shape cue, the probe does not silently drop it (#1110): it emits a single unclassified — review manually candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block.
Non-English projects: the probe reads English via text_en, the SPEC does not have to (#2773, durable fix #3717). The shape cues are English word-boundary patterns, so a project running with response_language set would otherwise have every requirement match nothing, classify to zero shapes, and land in unclassified — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. spec-phase Step 5.5 therefore populates an optional text_en field alongside each requirement's text with a faithful English translation: text_en is engine input, never user-facing output, so it is translated while text keeps the requirement's own wording (the SPEC stays in the original language), requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to response_language. Translation makes the classifier applicable; it does not make it omniscient. A requirement carrying no shape cue in any language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit shapes array on the requirement instead of relying on prose classification.
The resolved edges populate a ## Edge Coverage section in SPEC.md. Unresolved applicable edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under --auto, the probe never auto-dismisses — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs [auto] edge coverage: C covered, B backstop, U unresolved. The one exception is an unclassified candidate: --auto leaves it unresolved (surfaced as a flagged assumption), never auto-backstop — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim.
The load-bearing wire is the plan-phase lift: covered and backstop edges become must_haves.truths the verifier can check, so the section is not merely documentation. A backstop edge is lifted as a structured non-inferable marker ({ statement, verification: backstop }, a flat scalar — not a prose note), which the honest verifier then consumes (see below) — closing the loop the edge-probe opened.
Honest verifier — abstention on non-inferable checks (#1154). A non-inferable (backstop) truth is one whose correct behavior is not derivable from the spec alone, so the verifier cannot self-detect the gap and would confidently false-pass it (~100% of the time). At verify time, a backstop truth the verifier cannot confirm with explicit evidence (a passing wired held-out/property-based test, or a directly-observed behavior) abstains → human_needed with reason insufficient_spec (reported as unverified — held-out test recommended), never a silent passed. This is the verify-time, truth-axis mirror of the prohibition judgment-tier disposition (ADR-550 D4): exogenous (driven by the backstop tag, never a self-judged "abstain if unsure"), routing-not-diagnosis (the held-out test carries the omitted rule), and capable-tier dependent (reliable on sonnet+; the budget haiku tier degrades toward current behavior). An inferable truth is never abstained (the over-abstention guard). Reference: Honest Verifier.
Requirements:
- REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a
## Edge CoverageSPEC section. - REQ-EDGE-02: The relevance filter MUST raise only applicable categories; each raised edge resolves to exactly one of covered/dismissed/backstop/unresolved.
- REQ-EDGE-03: A
dismissedresolution MUST require a non-empty reason. - REQ-EDGE-04: An unresolved applicable edge MUST trigger the soft gate; write-anyway stamps the row as a planner assumption.
- REQ-EDGE-05:
--autoMUST never auto-dismiss — auto-cover or auto-backstop only. - REQ-EDGE-06:
plan-phaseMUST liftcoveredcriteria andbackstopnotes intomust_haves.truths. - REQ-EDGE-07: A requirement whose prose matches no shape cue MUST surface an
unclassified — review manuallycandidate (never silently dropped);--autoMUST leave itunresolved, never auto-backstop. - REQ-EDGE-08:
plan-phaseMUST lift abackstopedge intomust_haves.truthsas a structured flat-scalar marker ({ statement, verification: backstop }), never a prose parenthetical. - REQ-HONEST-01: At verify time a
backstoptruth that cannot be confirmed with explicit evidence MUST abstain →human_needed(reasoninsufficient_spec), neverpassed; an inferable truth MUST never be abstained (over-abstention guard); abstention MUST be exogenous (driven by thebackstoptag, not self-judgment).
Reference: Edge Probe
v1.43.0 Features
145. MemPalace Memory Capability
Purpose: Opt-in cross-session and cross-project memory via the MemPalace external service (local-first, MCP + CLI). Wires deliberate recall before discuss/plan and verbatim capture + temporal-KG sync at phase boundaries through the ADR-857 capability mechanism. Default-resilient: disabled by default, every hook is onError: skip, and an absent MemPalace installation leaves the loop unchanged.
Commands: /gsd-mempalace-recall, /gsd-mempalace-capture
Requirements:
- REQ-MP-01: Opt-in via
mempalace.enabled: true. Defaultfalse— the loop is unchanged when unset. - REQ-MP-02: At
plan:pre, skillmempalace-recallproducesMEMORY-RECALL.mdfrom prior decisions, patterns, and surprises retrieved via wake-up + semantic search + KG timeline. When MemPalace is unreachable, writes an "unavailable" stub and continues. - REQ-MP-03: At
discuss:post,plan:post, andverify:post, skillmempalace-capturefiles the phase artifact verbatim into the appropriate MemPalace room (decisions,planning,milestones). Capture is idempotent viamempalace_check_duplicate. - REQ-MP-04: At
ship:post, agentgsd-mempalace-curatorwrites a diary entry, proposes cross-project tunnels (whenmempalace.cross_project_tunnels: true), and runs wing-scoped sync pruning. - REQ-MP-05:
mempalace.memory_modehas three wired values:augment(default — palace is an additive recall layer alongside GSD native memory, which stays authoritative),kg_backend(knowledge-graph queries resolve against the palace's temporal KG as the primary source,.planning/graphs/as fallback; non-KG drawer recall stays additive),replace(recall resolves through the palace as the source of truth, native memory as fallback). Every mode isonError:skipand default-resilient — an unreachable palace degrades to native memory and GSD keeps writing.planning/graphs/, so no mode loses memory. Cross-mode migration of existing.planning/graphs/into the palace is out of scope (not yet implemented). - REQ-MP-06: Every hook is
onError: skip. No hook carriesblocking: true. Memory never halts or fails a phase. - REQ-MP-07: Interactive runs prefer MCP tools; headless/cron runs prefer the MemPalace CLI (
mempalace wake-up,mempalace search,mempalace mine,mempalace sync). - REQ-MP-08:
mempalace.auto_capture_hooksis forward-declared and not yet functional. No native Claude Code hooks (stop,precompact,session-start) are installed by this key; the capability's hooks array is empty. This key is reserved for the future "Connected Capability" phase. Defaultfalse.
Configuration: mempalace.enabled, mempalace.memory_mode, mempalace.wing, mempalace.recall_on_discuss, mempalace.recall_on_plan, mempalace.capture_artifacts, mempalace.mirror_kg, mempalace.cross_project_tunnels, mempalace.diary_journal, mempalace.auto_capture_hooks
See Configuration Reference for full schema and How to enable cross-session memory with MemPalace for a setup walkthrough.
146. Spec-Phase Prohibition Probe
Command: /gsd-spec-phase
Purpose: Surface the unwritten must-NOT constraints — the values/safety/ethics interpretations a feature could silently become that the author would never want but the spec does not forbid — before any code is written. The edge probe reaches data-shape edges; it structurally cannot reach prohibitions. This is the missing instrument, running as Step 5.6 of spec-phase, after the edge probe.
Behavior: A two-stage, prose-orchestrated pass per requirement (no compiled recall engine — recall is inherently model-driven, ADR-550 D7b):
- Recall (adversarial probe): "What could this feature silently become that the author would NOT want, but the spec does not forbid?" — model-robust open-vocabulary elicitation across values/safety/ethics.
- Precision (one-pass classifier): drop routine-engineering items, keep genuine values/safety/ethics prohibitions — collapses the raw list to the load-bearing few.
Each surfaced prohibition is resolved to exactly one of three states:
| State | Meaning | Downstream effect |
|---|---|---|
resolved |
Confirmed a real must-NOT | NEGATIVE acceptance criterion written into the SPEC ## Prohibitions (must-NOT) section; lifted into plan-phase must_haves.prohibitions (its own sibling block, never truths) |
dismissed |
Not a genuine prohibition (requires a non-empty reason) | Recorded with its reason; empty dismissals are rejected |
unresolved |
Deferred | Soft-gates the spec; surfaced as a planner assumption |
Each resolved prohibition carries a verification tier — test (a negative test can enforce it) or judgment (only human/LLM judgment can). At verify time, judgment-tier prohibitions route to a never-silent / never-hard-halt soft gate (autonomous emits an unverified-prohibition — human review recommended flag); test-tier prohibitions are enforced via the deterministic check prohibition-enforcement gate — green when the wired negative test / lint rule passes, hard-gate (flagged, non-green) when missing or failing, in both interactive and autonomous modes (#1259, ADR-550 D5d). Under --auto, the probe never auto-dismisses. Canon-bound concerns (OWASP / GDPR / fairness) are referred to /gsd-secure-phase rather than minting SPEC prohibitions (ADR-550 D6).
The load-bearing wire is the plan-phase lift into must_haves.prohibitions, so the section is not merely documentation.
Deterministic prohibition-check descriptor source (#1278). A resolved test-tier prohibition MAY carry an optional check descriptor — the flat-scalar keys check_kind (node-test | lint-rule), check_target, and check_rule (lint-rule only) — authored at spec-phase. projectProhibitions projects these scalars deterministically and verify-phase reads them back to locate the check handed to check prohibition-enforcement, so a wired, passing test closes the gap with zero manual descriptor authoring (previously the verify-phase LLM had to invent {kind, target, rule} each run, #1259). The descriptor is optional and backward-compatible — a descriptor-less prohibition parses and disposes byte-identically to today — and fail-closed: a partial, invalid, or absent descriptor falls through to the producer's existing fail-closed locate, never a silent green. failFirst stays a verify-time caller attestation (machine-proven fail-first is tracked in #1279).
Requirements:
- REQ-PROHIB-01: The prohibition pass MUST run after the edge probe and emit a
## Prohibitions (must-NOT)SPEC section. - REQ-PROHIB-02: Stage 1 MUST ask the adversarial recall question; Stage 2 MUST drop routine-engineering items and keep values/safety/ethics prohibitions.
- REQ-PROHIB-03: A
dismissedresolution MUST require a non-empty reason. - REQ-PROHIB-04:
--autoMUST never auto-dismiss. - REQ-PROHIB-05:
plan-phaseMUST lift resolved prohibitions intomust_haves.prohibitions(nevertruths). - REQ-PROHIB-06: A well-formed but unwired
test-tier prohibition MUST fail closed at verify time — never a silent pass. - REQ-PROHIB-07: A
test-tier prohibition with a machine-proven-fail-first, genuinely-passing (non-vacuous) wired mechanical check (anode --testnegative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing, un-provable, or non-passing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes. Fail-first is machine-proven, not caller-attested (#1279, ADR-550 D5d): before a clean pass greens, the producer independently runs the wired check against a known violation (the descriptor'sviolationFixture) and confirms it goes RED — a lint rule via the violating fixture, a node test via the violating subject injected through theGSD_PROHIB_SUBJECTconvention; absent a violation source it fails closed, never falling back to attestation. (Enforcement half shipped #1259; deterministic descriptor auto-locate in #1278.)
Reference: Prohibition Probe
147. Capability Management Command
Command: gsd capability install | update | remove | list | outdated | disable | enable
Purpose: The user-facing CLI for the ADR-1244 capability ecosystem — install, upgrade, remove, list, check for updates, and toggle GSD capabilities (first-party and third-party overlays) from a registry / git / npm / tarball / local source. Wires the Phase-3/4 lifecycle library (source resolver, install ledger, trust gate) to a command users actually run.
Behavior:
install <spec> [--integrity sha512-…] [--scope global|project] [--yes] [--shared-file <rel>]…— resolve (copy-only) → verify integrity / SHA pin →engines.gsdgate → disclose executable surfaces → consent (--yesgrants; without it an executable install aborts after printing the disclosure and writes nothing) → validate → extract → record the ledger.update [<id> | --all] [--scope] [--yes]— re-resolve the capability's recorded source and upgrade via atomic stage-then-swap; re-consent when the executable set changed;--allreports a per-capability outcome and exits non-zero on any partial failure.remove <id> [--purge-data] [--scope]— strip the ledger-recorded files + marker-isolated shared edits; first-party capabilities are rejected (use the product uninstaller).list [--json]— first-party + installed overlay capabilities (both scopes) as a JSON array.outdated [--json] [--scope]— light remote peek of each installed overlay's recorded source (ADR-1244 D6 per-source matrix: gitls-remote --tags, npmview … versionresolving the highest version matching the recorded range, local re-read; tarball →manual, registry →unknown) reportingoutdated/current/pinned/manual/unknownper capability. A source pinned to an immutable ref (git#sha:or#tag:, or an exact npm version) is reportedpinned. A bare git#<ref>is classified at the remote: if it resolves exclusively underrefs/tags/it is an immutable tag →pinned; if it resolves to a mutable branch (or is ambiguous) it isunknown. Bounded subprocesses (git ≤30s, npm ≤60s) and a failing peek degrades that row tounknownwithout crashing the command.--jsonfor machine output, default for a table.disable | enable <id>— toggle activation state (equivalent togsd capability set <id> --off/--on).
Trust boundary: install never executes capability code (copy-only staging); executable surfaces require explicit consent; sources are gated by the project-scoped capabilities.strict_known_registries policy (fail-closed on a malformed/unparseable value); every shared-config write/delete is realpath-confined to the scope root, and a name collision with a user's mcpServers entry is never clobbered.
Reference: gsd capability command reference · ADR-1244
148. Smart Entry Launcher
Command: /gsd-next
Tool: gsd-tools smart-entry [--json]
Purpose: Provide a state-aware front door that reads project/workflow state, classifies the user's situation, presents a short menu, and dispatches exactly one existing GSD command.
Requirements:
- REQ-SMART-ENTRY-01: Detection MUST be read-only and deterministic; classification lives in
gsd-tools smart-entry. - REQ-SMART-ENTRY-02: The launcher MUST never perform project work directly; it only displays a menu and dispatches one command.
- REQ-SMART-ENTRY-03: The workflow MUST fall back to
/gsd-progressif detection fails. - REQ-SMART-ENTRY-04: Each classified situation MUST provide exactly one recommended action and valid slash commands.
- REQ-SMART-ENTRY-05: Text-mode runtimes MUST receive a numbered-list fallback instead of being stranded by interactive UI assumptions.
Situations: no project, paused, blocked, verify failed, needs first phase, planning, executing, verify pending, idle stranded, complete, unknown.
Reference: Smart Entry Design
v1.7.0 Features
These are features new to @opengsd/gsd-core 1.7.0 (the current release line: 1.0.0 → 1.2.0 → … → 1.6.1 → 1.7.0). The preceding
v1.27–v1.43.0sections use the retired get-shit-done-cc / get-shit-done-redux feature numbering and are not gsd-core releases — see Legacy Release Notes.
149. Embeddable Orchestration System (Host-Integration Interface)
Purpose: Express every host integration against one public, versioned contract (ADR-1239 Phase A, #1690) instead of bespoke per-host wiring, so onboarding a new host becomes additive descriptor work.
Behavior: The interface exposes six interface points (command, dispatch, model, hooks, state, artifact), eight negotiated axes, and a PROTOCOL_VERSION handshake that negotiates down to min(host, engine). In 1.7.0, 14 runtimes were migrated onto the interface via imperative adapters (OpenCode #2087, Cursor #2089, Cline #2090, Hermes #2091, Qwen #2092, Kilo #2093, Trae #2094, Kimi #2095, Antigravity #2096, Augment #2097), a declarative adapter (Codex #2088), plus full lifecycle-hook wiring for CodeBuddy (#2098), GitHub Copilot (#2099), and Windsurf (#2100). Descriptors gained an extensionEvents vocabulary (#1946), and /gsd-surface now reproduces a runtime's agent output byte-for-byte from the installer's descriptors (#1575).
New runtimes: ZCode (Z.ai — Agentic Development Environment for GLM-5.2, #1925), pi (npx @opengsd/gsd-core --pi, #2102), and a repo-local VS Code extension driven through the adapter (#2103). The retired Gemini CLI now redirects to Antigravity CLI, its official successor (#1928).
Reference: The Embeddable Orchestration System · Host-Integration Interface · Interface versioning policy
150. Discoverability Registries
Purpose: Two non-endorsing catalogs for third-party extensions (#2182).
Behavior: The Community Capability Registry (#2188) lists third-party Feature Capabilities installed with gsd capability install; the EoS Registry (#2193) lists third-party host integrations built on the ADR-1239 interface. Every entry embeds a live release badge and links to a GitHub Discussion. Registration is a documentation PR, regenerated with npm run gen:registry.
Reference: GSD Registries
151. Companion MCP Server
Command: gsd-mcp-server
Purpose: A companion MCP server exposing GSD over stdio JSON-RPC 2.0, covering interface points 1 and 5 (#1681).
Behavior: OpenCode installs auto-register it as mcp.gsd (#1682). OpenCode also gained the opencode-subset hook dialect plus session.idle handling (#1682) and now runs GSD's lifecycle safety hooks — prompt-injection guard, read-before-edit guard, and injection scanner (#1923).
152. Statusline Token Count & Git Segment
Purpose: Opt-in statusline additions surfacing more session context.
Behavior: An absolute token count on the context meter (#2161) and a git branch + working-state segment (#2163), both opt-in. A companion opt-in compact GSD-state format condenses the GSD state segment (#2162).
Configuration: statusline.*
153. Model Catalog Advances
Purpose: Refresh the default model tiers and how models are surfaced.
Behavior: Codex/OpenAI defaults advance to the GPT-5.6 family (Sol / Terra / Luna) (#2122); the verbose (1M context) model suffix collapses to a compact (1M) badge (#2160). GSD warns when model config changes without re-running the installer on static-frontmatter runtimes such as Codex and OpenCode (#1688).
Reference: Configuration · Configure model profiles
154. Claude Orchestration Capability (BETA)
Purpose: A default-off, BETA, Claude-only capability that adopts Claude Code's Workflow tool for parallel sub-agent orchestration (#1143).
Reference: The Claude orchestration capability
155. External-Job Capability
Purpose: A default-off capability that externalizes long-running compute as asynchronous external jobs, e.g. SLURM submission (#1165).
Configuration: external_job.submit_timeout_ms, external_job.poll_timeout_ms, external_job.artifact_dir (#1164)
156. API-Coverage Gate
Command: /gsd-verify-work
Purpose: A phase that integrates an external API, SDK, or service can no longer seal verification without a decided coverage matrix (#1562).
Behavior: At seal time the gate reads the phase scope — the plan bodies, falling back to this phase's ROADMAP section — and runs the deterministic detector over it. An integration signal without a COVERAGE.md matrix blocks the seal; no signal passes.
Unestablished scope is not a negative verdict (#3909). A phase with no plan body and no roadmap section gives the detector nothing to examine. The gate used to run detection over zero bytes and pass, certifying "no external-API integration" from a probe that never looked. It now holds the seal instead, reporting scope_unavailable: true. A phase whose plans are real and simply contain no API vocabulary is unaffected — the discriminator is bytes examined, never signals found.
Breaking change: a phase that previously sealed because its detector could not establish a scope is now correctly held. Add the phase plan, or record a reasoned No external API integration: <reason> declaration in COVERAGE.md. See Resolve a skipped capability probe.
157. State Rebuild & Configurable Graph Path
Behavior: A new gsd-tools state rebuild subcommand re-derives STATE.md from source (#1830). The new graphify.graph_path setting makes the knowledge-graph location configurable, so a single umbrella graph can serve several projects (#1825).
158. Broken-Windows Ledger
Behavior: A cross-phase defect register at .planning/WINDOWS.md accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths (#1950). /gsd-ship blocks while any entry is open; an entry can be waived only with a recorded reason (auditable) or marked fixed (removed from the blocking set). /gsd-progress surfaces the open + waived counts.
Commands: gsd-tools windows status | append | waive | fixed.
Config: workflow.windows_enforce (gate active, default false — opt-in enforcement). Enable with gsd config-set workflow.windows_enforce true. Tracking (the ledger itself, populated by the executor) is always on; only the ship gate is opt-in.
Backward compatibility: A project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly; the gate only activates once windows are recorded.
Configuration: graphify.graph_path
159. Complexity-Triggered Refactor
Behavior: An execute:post step measures the complexity of the files a phase touched (decision-point counting over comment- and literal-stripped source, no external dependency) and surfaces a scoped refactor proposal at .planning/phases/<N>/<NN>-REFACTOR.md when a function's score exceeds refactor.complexity_threshold or its growth over its recorded anchor exceeds refactor.complexity_jump_delta — whichever trips first, both reported. Trigger semantics are strictly greater (ESLint's complexity: {max: N} convention), so a score exactly equal to the threshold does not trigger. The anchor is set the first time a function is observed and moves only when the proposal is dispositioned via refactor accept or refactor decline — never when the score alone improves — so the jump delta is cumulative growth since the last conscious decision about that function, not the change made in a single phase. Advisory by default: the proposal is informational only, never edits code, and never blocks. Opt-in refactor.trigger_strict records an untriaged proposal as an open deviation entry in the broken-windows ledger (#1950) instead — it does not block on its own; ship-blocking is broken-windows' existing ship:pre gate, enabled separately with workflow.windows_enforce. Without broken-windows installed, strict mode still records the proposal locally and says so. Enabling refactor.trigger_strict without workflow.windows_enforce also on (or with broken-windows absent) surfaces a typed refactor_strict_not_enforcing warning on every triggering evaluate, naming the exact remediation, so this enforcement gap is never silent. A declined proposal resolves its ledger entry as waived with the recorded reason; an accepted one resolves as fixed. The metric is approximate by construction: biased against a flat switch, blind to nesting depth, JS/TS-family only, and a renamed function loses its anchor (issue #1953).
Commands: gsd-tools refactor evaluate | status | accept | decline.
Config: refactor.trigger_enabled (master gate, default false), refactor.complexity_threshold (default 15), refactor.complexity_jump_delta (default 5), refactor.trigger_strict (default false). See Configuration Reference.
Backward compatibility: Off by default. When refactor.trigger_enabled is false the hook never runs and writes nothing; a project that never enables it is completely unaffected.
160. Archive Quick Tasks at Milestone Close
Command: /gsd-complete-milestone (forward path), /gsd-cleanup (retroactive path), gsd-tools milestone complete --archive-quick / gsd-tools milestone archive-quick <version> (#2142)
Behavior: .planning/quick/ otherwise accumulates one directory per /gsd-quick task forever. /gsd-complete-milestone now offers a Yes/Skip prompt — when accepted, it moves every directory under .planning/quick/ into .planning/milestones/<version>-quick/, (re)writes that archive directory's README.md (an index built by scanning the archive directory, one entry per task, linked to its SUMMARY.md when one exists), and clears the data rows of STATE.md's ### Quick Tasks Completed table while preserving its header and detected column variant. /gsd-cleanup offers the same archival retroactively, for milestones that were already closed before their quick tasks were swept, via the narrower milestone archive-quick <version> command — identical move/index/reset behavior, but without touching ROADMAP.md, REQUIREMENTS.md, MILESTONES.md, or milestone-completion guards, so it can be re-run safely against an already-completed milestone.
Why opt-in. Phase-directory archival is default-ON (#1871) — omitting a phase directory from an archive would silently leave stale execution history in the way of the next milestone's roadmap. Quick tasks carry no such downstream conflict, so archival here defaults OFF: a user who never passes --archive-quick sees zero behavior change. This is a deliberate asymmetry with phase archival, not an oversight.
Why bucket-all, not per-milestone. .planning/quick/ is a flat directory with no on-disk record of which milestone a given task belongs to. Splitting tasks per milestone was considered and rejected — inferring provenance from dates (creation time vs. a milestone's shipped date) is a proxy, not a fact, and a wrong inference on a one-way mv is silently irreversible. Archival instead buckets everything currently in .planning/quick/ into the one milestone being completed (or, on the retroactive path, the one milestone chosen), and says so in the confirmation prompt.
Why the index is built from disk, not from STATE.md's table. The ### Quick Tasks Completed table is a running log a workflow step appends to — it demonstrably drifts from what's actually in .planning/quick/ (the motivating case: 53 rows against 49 directories, ~22 rows pointing at directories that no longer existed, 18 directories with no row at all). Building the archive's README.md index by scanning the archive directory itself, rather than trusting the table, means the index can never inherit that drift; a re-run's index also naturally includes entries a prior run already archived, since it's re-derived from what's physically present.
Known limits:
- No per-milestone provenance — bucket-all is the only option (see above).
- A
### Quick Tasks Completedtable whose columns match neither registered variant (with/without a Status column) is left untouched with a warning rather than reset, since clearing it would risk destroying rows under a schema GSD doesn't recognize. - A
STATE.mdwith no### Quick Tasks Completedsection at all is a normal, silent no-op for the reset step — the section is created lazily by/gsd-quick, not present in the project template.
See Archiving quick tasks for the full walkthrough.
161. Verify-Command Path Grounding
Command: /gsd-plan-phase (automatic), gsd-tools check verify-command-paths <N> (#2401)
Behavior: A planner authoring a per-task <automated> verify command has no line of sight to whether the path it just wrote actually resolves, and gsd-plan-checker had no deterministic way to check — so it hand-reasoned the filesystem and, in the motivating case, prescribed two successively-wrong replacement paths (the second citing a package.json that did not exist). Two changes close that:
- Prior-command inheritance. The nearest prior phase's
<automated>commands are surfaced to the planner asprior_verify_commands, at every context window. Cross-phase enrichment was previously gated oncontext_window >= 500000; at 200k the planner re-invented the command and got it wrong. This payload is a handful of one-liners, so it is never gated. - A deterministic probe.
gsd-tools check verify-command-paths <N>resolves each<automated>command's target directory and reports whether it exists and holds the manifest the command needs./gsd-plan-phaseruns it before the plan-check pass and hands the JSON to the checker, which acts onseverityinstead of guessing.
It never executes command text. PLAN.md is model-authored, so running it from the checker would be arbitrary code execution — and would trigger the real lint/build as a side effect. The probe only resolves paths and stats directories; a package.json it finds is read for script names only.
Why a recognizer, not a shell parser. Interpreting shell would mean maintaining a bad shell. Exactly two forms are grounded — a leading cd <literal> chain and npm --prefix <literal> — and any path carrying a variable, glob, substitution, or ~ returns unresolvable, which is a warning and never a blocker. The parser's incompleteness is the specification: it degrades to "cannot prove" rather than growing features. Refusing to guess is the fix, not a limitation of it.
It reports, it never prescribes. The payload carries the target that failed and what was missing; there is deliberately no suggestion field. Choosing the replacement is the planner's job — and the planner now has the prior phase's proven command to reach for.
Not findings: a target an earlier task in this phase creates (pending_creation), a command with no cd/--prefix at all, and the Nyquist MISSING — Wave 0 … sentinel, which Dimension 8 owns.
Known limits:
- Only
cd <literal>andnpm --prefix <literal>are recognized.pushd,make -C,yarn --cwd,pnpm -C, andcargo --manifest-pathreportunresolvable. - Verdicts are relative to the checker's project root. Under parallel worktree execution the executor's root differs, so a bare ancestor climb (
cd ../..) is reportedoutside_rootas a warning rather than asserted about. script_missingis advisory only — this phase may be adding the script — so a genuinely mistyped npm script still reaches the executor.
See Resolve verify-command path findings and gsd-tools check verify-command-paths.
162. Statusline STATE.md Freshness Marker
Config key: statusline.show_state_freshness (default false)
Purpose: A solo developer returning to a project after time away reads "Phase 4, executing" in STATE.md and acts on it — without noticing the codebase has moved 40 commits since that line was written. /gsd-health reports this as W024, but only if the user thinks to run it. The statusline is the one surface seen continuously without asking (#2734).
Behavior: Renders state ~N commits back inside the GSD-state segment when STATE.md carries a state_head stamp (#2573) and HEAD is at least STATE_HEAD_ADVISORY_COMMITS (20) commits past it. Both statusline formats carry it — the default renderer and the compact statusline.state_format one.
The threshold is 20, deliberately not 1. With commit_docs: true (the default) the commit carrying a STATE.md sync advances HEAD by one, so a > 0 threshold would render state ~1 commits back permanently on a project that is by construction fresh — alarm fatigue on the one always-visible surface.
It degrades to silence rather than to a wrong answer. The marker is absent — never "fresh" — when the stamp is malformed, when the project root does not own its .git (an enclosing unrelated repo would otherwise answer), in a planning.sub_repos workspace (the outer HEAD never advances when code lands in children), when history was rewound past the stamp, and when git is unavailable or slow. A freshness claim the project cannot substantiate degrades to unknown.
Cost: exactly one bounded git rev-list call per render, and only when enabled and a stamp is present — rev-list --left-right --count answers ancestry and distance together, and repo pinning is a filesystem check rather than a subprocess. Disabled (the default) it adds none.
A proxy, never a drift measurement. The count includes commits that touched nothing STATE.md describes, and the stamp restamps on every state write — so a low count means "something wrote STATE recently", not "STATE is accurate". Rendered with a ~; never gate on it.
Reference: Configuration · Read the statusline freshness marker · ADR-2164
163. Read-Only Planning Snapshot (planning inspect)
Command: gsd-tools query planning inspect
Purpose: Give downstream consumers — harness UIs, mission-control surfaces, dashboards, bots — one schema-versioned JSON document describing everything .planning/ knows, so nothing outside gsd-core has to parse ROADMAP.md / REQUIREMENTS.md / *-PLAN.md / *-SUMMARY.md a second time. gsd-core is the single source of .planning/ truth; a second parser is a second answer.
Requirements:
- REQ-INSP-01:
PLANNING_INSPECT_SCHEMA_VERSION = 1is emitted asschema_version. Consumers MUST reject any other value rather than best-effort-parse an unknown shape. - REQ-INSP-02: Read-only. The command mutates no planning state, and mutates nothing on disk, under any input.
- REQ-INSP-03: Unknown or conflicting evidence serializes as
null/"unknown"with a coded entry indiagnostics[]— never inferred, reconciled, or defaulted. Every key is always present; a key is never omitted to signal absence. - REQ-INSP-04: Argument errors fail loud (non-zero exit, typed
ERROR_REASON); data gaps do not. v1 takes no arguments, and a stray positional or unknown flag is a usage error rather than a silently-ignored one. - REQ-INSP-05: Roadmap acceptance, verification status, and UAT items are reported side by side per phase and are never folded into a single verdict. A ROADMAP checkbox carries
authoritative: false— completion is derived from disk state. - REQ-INSP-06:
accepted_phasesandcompleted_plansare independent fractions.percentisnullwhenever the scope is notcomplete, per the same rule the roadmap and progress surfaces follow. - REQ-INSP-07: Payloads over ~50 KB use the existing
@file:spill channel, resolved transparently before stdout.
Why it does not simply serialize the internal snapshot. PlanningSnapshot (the diagnostic-rule subject introduced by ADR-3180 §8.1) is deliberately additive and still growing — four fields at Phase 10, twenty-plus by Phase 12. Handing that shape to external consumers would freeze an internal contract by accident. planning inspect declares its own flat schema and maps into it, so a field added to PlanningSnapshot never changes what this command emits.
Composed, never re-derived. Milestone identity and phase enumeration arrive via buildPlanningSnapshot; completion from isPhaseComplete (disk-strict); live-plan counting from scanPhasePlans; the percentage arithmetic from clampPercent; STATE fields from stateFieldValue; plan bodies from the Plan Document Module; requirement IDs from parseRequirements; UAT items from parseUatItems. Markdown structure is read through the Markdown Sectionizer and Markdown Table Model seams, so the Traceability table is resolved by column name against its registered schema rather than by a position-anchored regex.
Known limit — task-scoped file provenance. A <task> declares the files it plans to touch, but SUMMARY.md's ## Files Created/Modified describes the whole plan. Spreading that list across a plan's tasks would be inference, so a task's changed_files is populated only where the summary attributes files to that specific task; otherwise it is null with provenance: "plan_scoped". Closing this needs a change to the SUMMARY format, not to the reader.
Reference: CLI Tools · Consume the planning snapshot
164. Live-DOM UAT Capability
Config key: workflow.live_dom_uat (default false)
Purpose: A phase with a live-UI acceptance criterion could not be finished by the agent that executed it. gsd-executor carries no browser tools, so it correctly returned a checkpoint:human-action — even though the work was not human-only, just tool-less. Every such phase quietly degraded from executed by the executor to executed, then finished by hand in the orchestrator, and the plan's autonomous: false marker could not distinguish "a human must judge this" from "the executor lacks the tool" (#2856).
Behavior: A default-off capability owns one boolean key, one agent, and one additive step. When the key is on, gsd-dom-verifier runs at execute:wave:post and writes {phase}-DOM-VERIFY.md; the orchestrator's automated_ui_verification step additionally considers mcp__chrome-devtools__* / mcp__claude-in-chrome__* when present.
The executor's tool surface is unchanged in every configuration. Widening it was the reported proposal and was refused: for a first-party agent the static tools: list is the only control that exists — no capability can grant tools to one (ADR-1244 D2), no hook kind grants tool permissions (ADR-857 D4), and there is no per-dispatch override. Browser reach lives in one purpose-built agent that carries no Bash.
Two independent gates, both fail-closed. The capability's activationKey makes it resolve inactive when the key is off — resolveLoopHooks renders a hook only on state.active === true — and the step carries its own when guard. Tool presence alone never activates it: a browser MCP configured for unrelated work is not driven by default.
The pre-existing Playwright path is untouched. mcp__playwright__* keeps the gating it already had (presence plus an active UI phase). Pulling it behind a new default-off key would have silently removed working behavior from current users on upgrade; the key gates only the newly added families.
It tolerates the browser-profile lock rather than coordinating it. chrome-devtools-mcp holds an exclusive lock on its profile, so parallel waves collide. --isolated is a flag on the operator's own MCP server registration — GSD neither launches that server nor passes its arguments — so the verifier reports could_not_look / profile_locked, names the flag, and stops. No retry, no held-up wave.
nothing_to_report is never conflated with could_not_look. A report claiming no issues when it never opened a browser is worse than no report; the artifact carries a closed reason enum so the two are always distinguishable.
Known limits: no sandbox — once enabled, nothing constrains which origins are reached (ADR-1244 D5); DOM observation only, no screenshot diffing, accessibility audit, or performance tracing.
Reference: Configuration · Enable live-DOM verification · Explanation · Agents
165. Opt-In Parallel Reviewer Lanes
Command: /gsd-review, /gsd-plan-review-convergence
Config key: review.parallel_lanes (default false)
Purpose: Reviewer lanes within one review pass have no data dependency on each other — they all inspect the same immutable plan snapshot — but were dispatched strictly one at a time, so a pass with Codex, Gemini and Claude cost roughly the sum of three long reviewer calls. The serialization was a deliberate, unconditional protection against provider rate limits, which made it a global policy imposed on users whose providers could comfortably take concurrent requests, or who run local model servers with no limits at all (#3034).
Behavior: With the key enabled, the invoke_reviewers step dispatches each selected lane as a background job and joins all of them before REVIEWS.md and consensus are rendered. Wall-clock cost falls toward the slowest lane rather than the sum. Default remains false, preserving the existing sequential dispatch and its rate-limit protection.
The guard is strict equality, and it fails safe. Only the exact value true opts in — "1", "yes" and "TRUE" all stay sequential, so a mistyped config gets the conservative behavior rather than concurrent requests at a rate-limited provider. A failure to read the config falls back to sequential too. This polarity is deliberately the opposite of the commit_docs guard, which fails open: there, failing open preserves user intent; here it would fire the very requests the default exists to prevent.
Result ordering is unchanged in both modes. Per-lane results are written to slug-scoped files and concatenated in reviewer-selection order after the join, so gsd-review-lane-results.jsonl reads identically whether lanes ran sequentially or concurrently. Completion order never reaches the artifact. This also means concurrent lanes never share an append handle — a lane result larger than the pipe-atomicity bound cannot interleave and corrupt the models: / model_sources: frontmatter that write_reviews renders from that file.
Per-lane semantics are untouched. Timeouts, prompt budgets, the diagnostic stub for an empty or failed lane, explicit-lane failure (ADR-2782 D4), trust/egress checks and result-file layout all behave exactly as they do sequentially. A failing lane does not abort its siblings.
Known limits: convergence cycles stay sequential by design (review → replan → re-review has a genuine data dependency), so this speeds up each pass rather than reducing the number of passes; there is no concurrency bound, so every selected lane dispatches at once; and reviewer instances sharing one adapter dispatch concurrently against that single provider, which is the most likely way to hit a limit.
Reference: Configuration · Enable parallel reviewer lanes · Commands
166. Machine-Readable State Contract (.planning/state.json)
Purpose: External tools that display GSD project state — a workbench, a dashboard, an editor extension — had to parse STATE.md and ROADMAP.md heuristically. Those are human surfaces: their shape drifts as the templates evolve, and every consumer ends up carrying a brittle second parser that silently reports wrong numbers after an upgrade. GSD now publishes a small, versioned JSON snapshot instead, so the reader binds to a contract rather than to markdown (#3227).
Behavior: At every step boundary, GSD writes .planning/state.json — contract, flavor, milestone, phases[], next, updated_at. The boundaries are state begin-phase / planned-phase / advance-plan / complete-phase / milestone-switch, phase add / add-batch / insert / remove / complete, and milestone complete. The write is best-effort and completely invisible to the command that triggered it: it cannot change an exit code, cannot change stdout, and cannot fail a workflow. Readers prefer the file when it is present and fall back to markdown when it is not.
Requirements:
- REQ-SC-01:
contractis semver,1.0.0at introduction. Consumers gate on the MAJOR version;1.xchanges are additive only. Every key is ALWAYS present — an unknown value isnull, never an omitted key, because an omitted key is itself an observable a consumer would bind to. - REQ-SC-02:
phases[]carries{number, name, status}per phase,statusdrawn from exactlycomplete | in_progress | pending.numberis a string ("01"and"2.1"are both real ids and neither survives a number cast);nameisnullwhen the roadmap gives a phase no name, never a fabricated placeholder. - REQ-SC-03:
nextis the same recommended action the/gsdfront door routes, derived from the smart-entry classifier itself rather than from a second copy of its routing table. - REQ-SC-04: A missing
ROADMAP.md, a missing or unreadable.planning/, an unwritable target, or any other failure NEVER errors the parent command. A directory that is not a GSD project stays untouched — the publisher will not create.planning/in order to publish into it. - REQ-SC-05: The skills own the file; readers never write it. It is a derived cache — safe to delete, regenerated at the next boundary.
Composed, never re-derived. Milestone identity comes from getMilestoneInfo; phase rows from locateProgressTable, the same ## Progress locator the progress counters use, so state.json can never disagree with the rest of GSD about which phases are complete; the recommended action from classifyProject. This module introduces no second answer to any question GSD already answers.
Why it does not reuse planning inspect's schema. The two surfaces answer different questions and have opposite shapes. planning inspect is a rich, diagnostic-carrying pull query a consumer runs; this is a small push artifact a consumer watches. Publishing planning inspect's payload at every phase add would mean opening every plan, summary and requirements document on a hot path, and freezing a much larger surface as a contract.
It costs up to three bounded git calls per boundary. Deriving next from the smart-entry classifier means inheriting its git signals — git status --porcelain, and git log @{u}..HEAD. Each is timeout-bounded and swallows every error, so nothing can hang or fail because of it, but a command like phase add did not previously touch git at all. "Invisible to the parent command" is exact about exit code and output; it is not a claim about latency.
Known limits: an empty phases: [] cannot be told apart from "no ROADMAP.md" or "roadmap unreadable" — the 1.0 schema carries no diagnostic channel, and planning inspect is the surface that does. A roadmap phase marked Deferred is reported as pending, because the roadmap vocabulary has four values and this contract has three; inventing a fourth wire value would break every existing reader. phases[] is not milestone-scoped, so a long-running project lists every phase it has ever had.
Reference: Consume the state contract · Consume the planning snapshot
167. Stated Failing Direction
Command: /gsd-plan-phase (automatic), gsd-tools check verify-failure-directions <N> (#3172)
Behavior: A plan's <automated> block is the thing that decides whether work is done, and nothing checked that the command inside it could fail. In the motivating case six plans shipped 21 commands that could not run at all — cargo test -p <pkg> --lib against a package with no library target. They read as rigour and were not falsifiable, so three separate executors each rediscovered the defect and improvised a substitute at execution time. Every runnable <automated> command now needs a <fails_when> sibling naming what output constitutes failure:
<verify>
<automated>npm --prefix apps/api test -- auth.spec.ts</automated>
<fails_when>non-zero exit, or "0 passed" in the summary line</fails_when>
</verify>
gsd-planner emits it; gsd-tools check verify-failure-directions <N> verifies it deterministically; /gsd-plan-phase runs the probe before the plan-check pass and hands the JSON to gsd-plan-checker, whose check 8f blocks on severity.
Why this shape and not the two obvious alternatives. Validating command shape — teaching the checker Cargo's --lib/--bin target resolution, then pytest's node-ids, then the next one — always trails the newest toolchain. Executing each command at plan time is the strongest signal but means running planner-invented commands, with whatever side effects they carry, during planning. Requiring a stated failing direction needs no toolchain knowledge at all, and it is the only one of the three that catches the dangerous case: the motivating command exited non-zero, so it failed loudly, but the same class of error with a command that exits 0 on a no-op passes green and silently. Naming the failure signal is what makes that visible.
Presence, not quality — deliberately split. The probe is deterministic and owns the blockers: a statement is missing, blank, or a whole-value placeholder (TBD, TODO, N/A, NA, none, unknown, TBA, ?, -). Whether the statement names the right signal is prose judgment, so gsd-plan-checker raises a vacuous statement ("the command fails") as a WARNING only. Every BLOCKER stays reproducible; judgment stays advisory.
It reports, it never prescribes. The payload names the command with no stated failure mode and stops there. A prescribed statement would be copied verbatim and carry zero information — reproducing the original defect one level up.
Not findings: the Nyquist MISSING — Wave 0 … sentinel (not runnable, so it has no failure mode to state), an empty <automated> body (check 8a owns command presence), and a <verify> with no <automated> at all.
Known limits:
- Presence only. A statement that is present and specific can still name the wrong signal; that is caught, if at all, by judgment rather than by the probe.
- Breaking for plans authored before this shipped. A phase planned earlier has no
<fails_when>anywhere and blocks on re-check until statements are added or the phase is re-planned. - The adjacent vacuous pass — a command that runs successfully and asserts nothing, such as a test-name filter matching zero tests and exiting 0 — is a distinct problem and is explicitly out of scope.
See State a failing direction and gsd-tools check verify-failure-directions.
168. Runtime Identity
Purpose: The predecessor package get-shit-done-cc publishes a binary named gsd-tools, and so does this one. They answer some of the same verb names with different semantics. #3129 is the worked example: phases.clear archives here and deletes there. Both print success-shaped output, and .planning/ is gitignored by default, so a user lost 43 phase directories with no error, no warning, and nothing recoverable from git. The failure was silent in both directions — the workflow could not tell it had reached the wrong handler, and the handler could not tell it had been called by a workflow written for a different contract (#3146).
Behavior: two independent defenses, one structural and one asserted.
Structural — the PATH branch. The launcher's PATH resolution branch looks for gsd_run instead of gsd-tools. Only this package publishes gsd_run; the predecessor publishes gsd-tools and gsd-sdk. Our gsd_run follows its own symlink chain and executes the gsd-tools.cjs sitting beside it, so resolving it cannot land on a foreign handler.
Asserted — every other branch. The path-based branches (a project-local install, a runtime config directory) have no such guarantee: they trust their configured location. So once resolution finishes, and before any verb runs, the preamble probes the tool it picked with runtime-identity --raw and matches the answer anchored against the compact payload. It exports the result as a two-valued GSD_IDENTITY_STATUS (ok / unverified) and, when it is unverified, prints one actionable line naming both plausible causes. The same gsd-tools runtime-identity verb remains available by hand, so a human or a support thread can settle "which tool am I actually running?" in one command.
The match is anchored at both ends, not a substring. A substring search for @opengsd/gsd-core accepts the decoy {"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}, which any colliding package could publish. The preamble instead requires the payload to begin with {"packageName":"@opengsd/gsd-core" and to end with a closing brace, so a truncated answer fails as well. Closing on } costs nothing in future-proofing: a JSON object's own brace is always the last character, whatever type the last value has.
The status is a value, not prose. GSD_IDENTITY_STATUS exists so the gate can be tested — and read by a later step — without anyone parsing the warning text.
The byte budget is why the assertion arrived second. The preamble is inlined into 113 shipped files, several of which sat within single-digit bytes of frozen size ceilings — agents/gsd-verifier.md had 2 bytes of headroom — and those caps are red lines, not budgets. A first attempt to inline an assertion broke five of them. What made it fit was collapsing the resolver's twenty near-identical elif [ -f … ] arms into a single candidate-list helper, which is worth far more bytes than the assertion costs: the preamble is now 1,876 bytes smaller than the version that carried no assertion at all, so every one of the 113 files moved away from its ceiling.
It fails closed. If no gsd_run is reachable, the resolver falls through its remaining path-based branches and finally errors with an install command. It does not fall back to executing whatever gsd-tools happens to be on PATH — that fallback was the vulnerability.
A doubly-sourced preamble cannot build a recursive launcher. command -v gsd_run finds the shell function on a second source and would return the bare string gsd_run, defining the function in terms of itself. unset -f gsd_run leads that branch, so the second source resolves exactly as the first did. (An executability guard was tried here instead and removed: it rejected the bare name, fell through every branch, and hit the resolver's exit 1 — which, in a sourced script, kills the caller's shell.)
Known limits:
- The assertion warns; it does not yet stop the run. The rollout is warn-then-fail. It cannot hard-fail yet because an
@opengsd/gsd-coreolder than theruntime-identityverb answers exactly as a foreign package does — neither answers — and at rollout the old-version case is the common one. The warning therefore names both causes. A later release turnsunverifiedinto a refusal. - An installation old enough to predate
bin/gsd_run(#381) is not reachable through thePATHbranch and must be upgraded or invoked through one of the path-based branches. - The probe costs one extra process launch per preamble source. It is a pure local read of baked coordinates, deliberately kept off the SDK bridge for that reason.
Reference: runtime-identity · Diagnose which gsd-tools is running
Generated by scripts/gen-features.cjs — add a fragment under docs/features/ and run --write.
3348. Context Drift Gate
Purpose: Warns (or optionally blocks) before /gsd-plan-phase reuses an existing
RESEARCH.md, PATTERNS.md, VALIDATION.md, or SPEC.md that predates a decision added to the
phase's CONTEXT.md after that artifact was derived from it. Deterministic — compares git commit
time (falling back to mtime for uncommitted edits), no model call. Sibling to the existing
codebase-drift and schema-drift gates in the drift capability. Configure with
workflow.context_drift_precheck (on/off) and workflow.context_drift_action (warn/block).
3884. "Failure Is a Value" — Strict Argv Rejection and the --pick Absence Contract
Purpose: ADR-3473 §8.4 states the rule directly: absence, emptiness, and
failure are three different things, and a routine that cannot tell them
apart eventually reports the wrong one. Before this change, gsd-tools had
two silent instances of exactly that collapse.
Half one — a stray positional corrupted state, silently (#3358).
parseNamedArgs read only the flags it recognized and dropped everything
else — an unrecognized --flag or an extra positional argument (for
example, a stray phase number appended after state.planned-phase) was
silently discarded rather than rejected. The caller's own positional read
(args[2], etc.) still worked, so the command ran anyway, on the wrong
phase, and overwrote the previously-current phase block with no error at
all. The fix makes parseNamedArgs(args, spec) return the command-routing
hub's own Result shape ({ok:true,data} | {ok:false,kind:'InvalidArgs',...})
and requires every call site to declare positionals: number | 'rest' — the
count of leading argv slots the caller itself reads directly. An unknown
flag or an unexpected positional past that boundary is now a loud,
non-zero-exit InvalidArgs failure instead of a token quietly falling on
the floor. A duplicate flag, a negative-number value (--plans -1), and a
documented free-text tail (init quick <description>) are deliberately
left alone — none of them are the defect this closes, and forbidding them
would just break working call sites for no gain.
Half two — --pick on an absent field answered '' at exit 0, exactly
like a present-but-empty one (#3365). --pick <field> extracted one field
from a command's JSON output, but a missing key, an out-of-range array
index, a partially-missing dotted path, or non-JSON output (including a
--raw command's output) all rendered the same way: empty stdout, exit
0. That is indistinguishable from a field that genuinely holds null or
'' — a real answer. The shell idiom X=$(… --pick F) || X=default could
therefore never observe the failure it was written to react to; only a typo
in the verb name would ever make it exit non-zero. --pick now exits 1
with a diagnostic on stderr (pick_field_absent naming the field and the
available top-level keys, or pick_output_not_json when the output could
not be parsed as JSON at all) whenever the field cannot be resolved. A
present field's value — including 0, false, null, and '' — is
unchanged: those are answers, not failures, and remain exit 0. See
CLI-TOOLS.md's --pick <field> contract
for the full outcome table and json-errors.md for the two
new reason codes.
Why not just default to zero for an absent count. The sub-issue's own Done-when checkbox suggested treating an absent field the same as a zero-valued one. That is rejected on the merits: it demotes "I could not answer" to "the answer is zero", which would make a count-gated shell guard fire unconditionally on the very projects that could never resolve the count in the first place — the opposite of what a gate is for.
Consequence for scripts/lint-unreachable-guard-drift.cjs. That guard's
Detector A existed specifically because the old --pick behavior made a
--pick … || echo <default> line's fallback arm permanently unreachable.
Once --pick exits non-zero on absence, that premise is false and the
shape the detector forbade becomes the correct idiom — so Detector A was
retired rather than kept. Detector B (the unrelated cat/ls-over-a-glob
nullglob hazard) is untouched. See
Resolve unreachable-guard findings, Shape A.
Known limits:
- A value token beginning with
--still cannot be passed to a declared value flag (--summary "--force is now default"now fails loudly instead of silently dropping the value) — strictly better, but no--flag=valueescape was added. --pickstill cannot distinguish an absent field from anullone on stdout alone — the distinction is carried entirely by exit code.- The ~10
gsd-tools.cjscall sites ofparseNamedArgsget no compile-time check (that file is hand-written JavaScript, not.cts); enforcement there is the runtime throw on a stale legacy call shape plus behavioral tests. - This phase does not sweep every routine in
gsd-coreforResultconformance — it applies the rule to the argument-projection seam and the--pickextractor it names, not the whole codebase.
3885. No Silent Swallow, No Verdict From Dropped Data
Purpose: ADR-3473 §8.5 states the rule directly: a failure or a gap in the
input must not be absorbed into an output that reads as authoritative. A
routine that drops data it could not read or could not resolve, and then
reports a clean result anyway, turns a diagnosable gap into a confidently
wrong answer. This closes four instances of that collapse found across
gsd-tools.
intel query no longer crashes past ~12000 levels of nesting (#3427).
searchJsonEntries / matchesInValue recursed with no depth bound at all —
an intel JSON file nested deeply enough overflowed the call stack with an
uncaught RangeError instead of a diagnosis. The original SDK-era bound
(MAX_JSON_SEARCH_DEPTH = 48, lost in the ADR-0174 consolidation) is
restored, paired with a truncated result field: a match at or above the
ceiling is not returned, and the result says so rather than reporting a bare
"not found" that is indistinguishable from a genuine miss. A match at depth
48 (inclusive) or shallower is unaffected; the bound is on nesting depth, not
breadth or total node count, so a shallow object with many siblings still
works unchanged.
phase-plan-index no longer blames the author for an edge the tool itself
dropped (#3427). A depends_on: token that resolves to no plan in the
phase (typo, or a stale cross-phase reference) silently dropped that edge,
making the dependent plan a DAG root — its own docstring recorded the intent
as "ignore this edge, never a throw." The tool then compared the resulting
degraded wave against the plan's declared wave: and reported the author's
correct declaration as a mismatch. The unresolved token is now named in its
own warnings[] entry (plan and token together), and the wave-mismatch
warning is suppressed for that plan only — a plan with no dropped edges and a
genuinely wrong wave: still warns as before. The token is escaped
(quoted, control characters and embedded newlines backslash-escaped) before
it is embedded in the warning text, so a depends_on value crafted to
contain a newline or a quote cannot forge a second, fabricated warning entry.
A code-review run where every lane failed no longer writes REVIEWS.md
from nothing (#3352). review.md's aggregation step wrote REVIEWS.md
regardless of whether any lane actually produced results — a run where every
lane failed still emitted a completed-looking review artifact, and the
per-lane outputs and .err files that would have explained the failure were
then destroyed by the run's own cleanup. REVIEWS.md is now withheld when
the aggregate has zero lines (every lane failed, not merely skipped under a
lower budget), the run reports the failure instead, and per-lane outputs and
non-empty .err files are preserved beside the phase's artifacts before
cleanup runs.
Unreadable directories are distinguished from absent ones (#3473 B5).
Four call sites collapsed an EACCES/EIO on a phase directory into the
same "nothing here" result as a directory that genuinely does not exist —
countPhasePlansAndSummaries (hasContext:false), runGapAnalysis, and two
guarded blocks in init.cts (context_path absent). Each now distinguishes
"could not read" from "does not exist" and names the discarded path and
error in a dedicated field (context_read_error / phase_dir_read_error)
rather than silently reading as absent.
Audited, no defect found: every retry-set / swallowed-catch call site
this phase's rule covers that had not already been fixed by a prior PR
(withPlanningLock, acquireStateLock, atomicRenameWithRetry,
renameWithRetry) was reviewed and found to already fail loudly on a fatal
errno rather than folding it into a retry.
Known limits:
- A match deeper than 48 levels is still not surfaced by
intel query— it is reported as truncated rather than as absent, but the value itself is not returned. Raising the ceiling is a separate decision. phase-plan-index'swaves/wavefields remain computed from the degraded DAG when an edge is dropped — this phase stops the tool from manufacturing a false verdict about it, but does not invent the missing edge. A consumer that schedules work fromwave(e.g.--wave Nfiltering) is still working from the degraded assignment.review.md's evidence preservation is bounded by what a lane actually wrote — a lane that produced no output at all leaves nothing to preserve.
3897. Runtime Marker Resolution, Derived Codex Sandbox, and In-Phase Short-Form Dependencies
Purpose: ADR-3473 §8.3 states the rule directly — one implementation per
invariant, not a hand-maintained copy that quietly drifts from the rule it
stands in for. This closes three instances across gsd-tools: a resolver
that never read the install-time signal it was documented to read, a
Codex sandbox map that was fully redundant with the tool contract it stood
in for, and a dependency-resolution tier lost when the SDK lineage was
retired.
A non-Claude install now resolves its own runtime with no config
needed (#3897). resolveRuntime's ladder was GSD_RUNTIME env var →
project config.runtime → 'claude' — the per-install .gsd-runtime
marker the installer has written beside VERSION since #2297 was read by
four separate hand-rolled copies (model-resolver.cts, and two more inside
gsd-cursor-subagent-start.js), but never by resolveRuntime itself. A
Codex, Cursor, or other non-Claude install with no GSD_RUNTIME set and no
runtime key in .planning/config.json therefore still resolved claude
everywhere resolveRuntime is consulted (slash-command style, query teams-status, validate agents's agent-directory selection, and 19 other
call sites). The marker is now the third rung — GSD_RUNTIME → config.runtime
→ install marker → 'claude' — so those installs resolve their own runtime
by default. The marker's contents are never trusted verbatim: they are routed
through the same name-normalization the env rung already uses, so a marker
holding an unexpected or hostile value degrades exactly like an unexpected
GSD_RUNTIME value would.
Codex sandbox permissions are derived from each agent's own tool contract,
not a hand-maintained map (#3897). generateCodexAgentToml looked up
sandbox_mode in an 11-entry CODEX_AGENT_SANDBOX map, falling back to
read-only — silently — for every role the map didn't name. Measured against
all 35 shipped roles, the map's 11 entries agree with deriving sandbox_mode
from each role's declared tools: frontmatter (workspace-write when it
declares Write or Edit, read-only otherwise) with zero disagreements,
so the map is deleted rather than clamped. The fallback, however, was
under-granting: 16 of the 24 roles that hit it declare Write/Edit and
would derive workspace-write. Pending a decision on whether Codex actually
enforces sandbox_mode (a question the derivation can't answer on its own),
those 16 are held at read-only by an explicit, self-invalidating hold
list — a hold whose role no longer derives broader, or that names a role
that no longer exists, fails loudly instead of being silently honored.
Every one of the 35 emitted .toml files is byte-identical to before this
change — the fix is in provenance (an explicit, reviewable rule instead of
a silent default), not in any installed agent's actual permissions today.
validate agents now reports Codex sandbox drift (#3897). A new
sandbox_posture field — report-only, exit 0, same shape as the existing
codex_posture — flags any installed Codex .toml whose sandbox_mode
disagrees with what its role's tool contract derives. Populated only when
the active runtime is codex.
depends_on accepts the bare plan number (#3897). A plan's frontmatter
could already reference a dependency by its full id ("03-01-auth-hardening")
or its canonical phase-plan prefix ("03-01"). A third form — the bare
plan number alone ("01") — existed in the retired SDK lineage but was lost
when that lineage was consolidated; a plan written with it silently dropped
the edge entirely, collapsing into wave 1 regardless of its declared
dependency. That form is restored, scoped to the same phase only: "01"
resolves to the sibling plan whose canonical id ends -01. This is an
observable behavior change — a phase whose plans used the bare form and had
silently collapsed into a single wave will now execute in its actual declared
waves. Two plans in the same phase sharing a bare form resolve first-write-wins,
by sorted plan-file order — deterministic, but arbitrary where the collision
happens, matching the retired behavior exactly.
Known limits:
- The 17 held Codex roles are pinned at
read-only, not widened. A faithful derivation from the tool contract would widen them, because they declareWriteorEdit; the previous hand-maintained map never listed them and they fell through a silent|| 'read-only'default instead. Deriving and holding keeps emitted TOML byte-identical for all 35 roles today while the derivation becomes the single owner of the rule. Widening them is a follow-up once Codex's actual enforcement ofsandbox_modeis confirmed — until then a hold is reversible and a widened sandbox is not. The hold list is self-invalidating: an entry naming a role that no longer derives broader, or that has no file in the shipped roster, fails rather than rotting into the subset map this change deletes. - The bare plan-number form is ambiguous by construction across two plans in the same phase that share a short form; first-write-wins is deterministic but not a conflict warning. Prefer the full or canonical id when a phase's plan numbering risks a short-form collision.
- The install marker never feeds model-tier resolution (
model_profile_overrides,model_policy.runtime_tiers) — that still readsconfig.runtimealone, and reporting-only host detection (agent_runtime) is a separate, pre-existing ladder this change does not touch.
3910. The Raw Terminator Is Banned by Construction
Purpose: Make a bare process.exit(...) a lint error everywhere it matters, so the
"nothing fails with success" defect class ADR-3889 exists to close cannot silently reopen
through a new call site.
What changed (ADR-3889 Phase 6, #3910):
- New rule
local/require-registered-exit(eslint-rules/require-registered-exit.cjs) flags anyCallExpressionshaped exactly likeprocess.exit(...). It does not flagprocess.exitCode = N— that assignment is the correct drain-then-exit patternrunMainitself uses, and the two are structurally distinct (an assignment target is never aCallExpression). - Registered on four globs:
src/**/*.cts,scripts/**/*.cjs,hooks/**/*.js,gsd-core/bin/**/*.cjs(eslint.config.mjs:420-426,545-547,574-576,601). Registering onsrc/**/*.cts— not only the emittedgsd-core/bin/lib/*.cjsmirrors, which are globally eslint-ignored (ADR-457) — is load-bearing: a rule registered only on the emitted surface is blind to the real sources, the same wayn/no-process-exitwent invisible (#3496). - The dead
n/no-process-exit: 'off'carve-out forhooks/**is deleted: Phase 7 (#3911) migrated every enforcement hook ontoterminateNow, so it protected nothing. - Exactly two allowlist entries, repo-wide:
- The body of
terminateNowinsrc/cli-exit.cts— detected structurally (anyprocess.exit()lexically nested inside a function namedterminateNow, and the file's basename iscli-exit.cts), not by path+line, so it does not rot when the function moves. gsd-core/bin/gsd-tools.cjs'sensureRuntimeBuildbootstrap-failure path, via an inline// eslint-disable-next-line local/require-registered-exitwith a stated reason — it runs before./lib/cli-exit.cjsis even required, so the registered-exit seam does not exist yet at that point in the process's lifetime.
- The body of
- The last raw terminators in
src/**/*.ctswere migrated onto the seam, most notablysrc/io.cts'serror(): it changed from an uncatchableprocess.exit(1)to a catchablethrow new ExitError(1)(stderr output is byte-identical;runMainprojects the exit code).terminateNowcould not serve this site — ADR-3889 §1 makes exit codes 0 and 1 unallocatable, sonameForExitCode(1)throws. That control-flow change required three interceptor fixes so anExitErrorreachesrunMain:command-routing-hub'sdispatch()now rethrows it, and the profile-pipeline router's detached.catch()no longer callserror()— it writes stderr and setsexitCodein place. - Known limits (documented and test-pinned, not endorsed): the rule matches the literal
process.exit(...)shape only, with no scope/flow analysis. It does not catchprocess['exit'](0)(computed member access),const e = process.exit; e(1)(aliasing to a local binding before calling), orprocess.exit.call(...)/.apply(...)(indirect invocation). Catching these needs binding/scope-aware analysis, out of scope for this issue; pinning tests intests/eslint-rules.test.cjsassert today's non-detection so a future widening is a visible choice, not a silent one.
See Resolve a raw-terminator finding for what to do when this rule fires, and ADR-3889 for the exit-code registry the seam is layered over.
3911. Hooks Declare Their Crash Policy
Purpose: Give every shipped enforcement hook (hooks/*.js, hooks/*.sh) a
named, auditable termination vocabulary instead of a bare process.exit(N)
scattered per file — and make a hook's fail-open/fail-closed choice a
declaration a reviewer can see, rather than an inference from which literal
integer follows process.exit( in its outer catch.
What changed (ADR-3889 Phase 7, #3911):
hooks/lib/hook-exit.js(hand-written) exposesallow(payload)→ exit 0,deny(payload, stderrPayload?)→ exit 2, andcrash(onCrash, payload), which dispatches toallow/denyper aHOOK_ON_CRASHpolicy the caller must supply —crash()has no default policy, so a hook cannot fail open by omission.- Every one of the 19 enforcement hooks under
hooks/*.jsnow declaresconst ON_CRASH = HOOK_ON_CRASH.ALLOW(orDENY) once, with a hook-specific comment naming why, and callscrash(ON_CRASH, payload)from its outer catch instead of a bareprocess.exit(0)/process.exit(2). No hook's effective exit code changed — this is a naming-and-declaration migration, not a behavior change. hooks/lib/cli-exit.jsandhooks/lib/exit-code-registry.jsare new, generated, git-tracked copies of the exit-code seam (src/cli-exit.cts/gsd-core/bin/shared/exit-codes.json), so a shipped hook can terminate correctly on a raw, unbuilt clone without depending ongsd-core/bin/lib/tsc output. Generated byscripts/gen-hooks-cli-exit.cjsandscripts/gen-exit-code-registry.cjs, both--checked bynpm run lint:generated-sync.terminateNowgained an optional third argument,stderrPayload, so a deny can send a full JSON body to stdout and a distinct plain-text reason to stderr — needed becausegsd-write-guard.js(Kimi's native hook bus reads stderr verbatim back to the model) always sent only the bare reason string on fd 2. The two streams are now written in independent try/catch blocks: previously a payload that failed to serialize on fd 1 aborted before fd 2 ever wrote, producing a deny with an empty stderr reason.- Two hooks are deliberately not migrated to
deny():gsd-read-injection-scanner.js(PostToolUse — its harness reads the block decision from the JSON response body, not the exit code) andgsd-cursor-subagent-start.js(follows Cursor's ownsubagentStartprotocol, which readspermission: "deny"from the JSON body at exit 0). Both still useallow()/crash()for their no-op and crash paths. gsd-phase-boundary.sh,gsd-session-state.sh, andgsd-validate-commit.shgainedset -euo pipefail, andgsd-validate-commit.sh's three swallow-and-pass sites (the opt-in config read, JSON command extraction, and theisGitSubcommandclassifier) now distinguish a genuine negative from "could not run" — on the latter they emit a stderr diagnostic and exit 0 instead of silently allowing every commit (#3838).
See Declare a hook's crash policy for the full how-to, and ADR-3889 for the exit-code registry this vocabulary is layered over.
3912. gsd-tools Declares Outcomes, Pinned at v1
Purpose: Give every gsd-tools terminating path a declared outcome name, and project that
declaration through the versioned exit contract (ADR-3889
§4) — without changing a single exit code for a caller that has not opted in.
Reference — what changed (ADR-3889 Phase 8, #3912):
error(message, reason)now maps itsreasonargument onto a declared outcome name (USAGE,NO_INPUT,UNAVAILABLE,INTERNAL,FAIL) via a fixed table closed over all 25ERROR_REASONmembers (src/io.cts'sREASON_TO_OUTCOME). Under the default contract versionv1, the declaration is recorded buterror()still throwsExitError(1)unconditionally, byte-identical to every prior release. Underv2(--exit-contract=v2/GSD_EXIT_CONTRACT=v2), it throwsExitError(projectOutcome(outcome, 'v2'))instead — e.g.SDK_MISSING_ARG/SDK_UNKNOWN_COMMANDproject to64(USAGE),CONFIG_KEY_NOT_FOUNDto66(NO_INPUT). All 278 call sites are untouched; 226 pass no reason and default toUNKNOWN->FAIL-> exit1under both versions.output()now declaresDEGRADEDwhenever its payload carries a serializableerrorvalue (any key order —{found:false, error}counts the same as{error, found:false}). The discriminator is survives-JSON.stringify, not mere key presence:{ error: undefined }does not declareDEGRADED, becauseJSON.stringifydrops anundefined-valued property before the payload reaches the wire.- A third
globalThiscell (src/cli-exit.cts'sPENDING_OUTCOME_KEY) holds the pending declared outcome betweenoutput()andrunMain. Semantics: last declaration wins, cleared on consumption — a later cleanoutput()call in the same invocation undoes an earlier degraded one, andrunMainclears the cell on every exit so a secondrunMainin the same process never inherits a stale declaration. - Precedence for the code a void-returning
main()ends up with, highest first: (1) an explicitmain()return, (2) a non-zeroprocess.exitCodemain()already set directly, (3) the pending declared outcome, (4) otherwise0. Projection may only ever set a code, never lower one — a review pass wrongly concluded the cell was fail-closed by construction; without rule 2,state validate --strictbriefly exited0where it must exit1. v1is byte-identical.DEGRADEDprojects to0underv1and to80(exitCodeFor('DEGRADED')) underv2— that asymmetry is ADR-2980's compatibility boundary, deliberately preserved, not a bug to reconcile.
Explanation — why this is the shape it is:
ADR-2980 ratified output({error})'s exit-0 population on measured blast radius (output has 170
direct callers) and Hyrum's Law grounds — a CLI exit code has no /v2/ of its own, so normalizing
it in place would have broken every caller already treating exit 0 as a soft signal. Its own
"Revisit if" clause named the missing piece: "a future gsd-tools major version provides a
compatibility boundary that a CLI exit code otherwise lacks." ADR-3889 §4 built exactly that
boundary — a versioned projection selected by flag or env var, defaulting to today's behavior — and
this phase is what wires error() and output() onto it. Declaring an outcome is unconditional and
immediate; only its projection onto an integer is deferred behind the version switch, so the
population ADR-2980 ratified keeps exiting 0 until a caller explicitly asks for something else.
The count matters here too: an AST re-measure for this phase found 64 output({error}) call
sites across the same nine modules ADR-2980 named — not the 60 that ADR itself recorded, the drift
concentrated in frontmatter.cts, phase.cts, and roadmap.cts. The v2 projection is asserted
over the enumerated 64, not a restated 60; see ADR-2980's amendment for the module-by-module
breakdown.
See Adopt the v2 exit contract for how to opt in and what
it means for a CI gate, docs/json-errors.md
for the full reference, and
ADR-2980 /
ADR-3889 for the decisions.
3951. Reachable Lint Rules and a Non-Destructive Quick-Task Append
Purpose: Make two ESLint rules cover the code they were written to govern, and stop
quick-tasks-append from overwriting curated progress.* values on a body-only write.
What changed:
local/no-adhoc-markdown-parsingreaches its whole registered surface. The rule short-circuited unless a file's path matched a flatsrc/*.ctspattern, so it self-gated on its own filename. Two consequences: 28.ctsfiles insrc/subdirectories sat inside thesrc/**/*.ctsglob it was registered on and were silently skipped, and the rule could not be extended by configuration at all — widening the glob alone left it inert. Both halves now move together, and a test pins that the gate and the registration agree in both directions.- The rule now also covers
tests/**andscripts/**, which surfaced 80 hand-rolled markdown parses across 43 test files. Seventy are routed through the existingmarkdown-sectionizerandmarkdown-tableseams; ten are suppressed with a stated reason (six of those are a shell-pipe detector whose regex merely resembles a table). local/no-adhoc-regex-escapesees property access. Its unsafe-new RegExparm examined only bare identifiers, sonew RegExp(obj['key'])— the shape runtime data actually arrives in — was invisible. That is why it never fired on a known ReDoS. It now inspectsMemberExpression, with an exemption keyed strictly on the property beingsource(18 safe sites), plus provenance exemptions for_SOURCEconstants reached through a required module (3 sites).quick-tasks-appendcan write the canonical row. Optional--quick-id,--slugand--directorylet a caller that has a real quick task emit the same row/gsd-quickrenders. Omit them — asfast.mddoes, having neither an id nor a directory — and the row is byte-identical to before.- A body-only append no longer re-derives progress. The route was the only body-only STATE.md
writer not passing
{ resync: false }, so appending one row triggered a full re-derive of the disk-derivedprogress.*frontmatter and replaced curated values. Reproduced: a project with two real phase directories and a curatedtotal_phases: 25collapsed to2on append.
Found by the widening: tests/config-field-docs.test.cjs asserted that
workflow.subagent_timeout's documented default is not 600 — but read the Type column instead
of Default, so it compared 'number' against '600' and could never fail. The guard against
regressing to the old seconds-based default had been inert. It is now row-scoped and real.
Known limits:
- #3426 and #3239 are not closed by this. Their hand-rolled scans in
tests/package-legitimacy-gate.test.cjsare built from line filters andsplit('|'), not the regex-literal fingerprints this rule detects — measured at zero violations even with the gate bypassed. They need new detectors, which is a separate design. - The 10 suppressions are suppressions, not fixes. Each names why the raw markdown text is the subject of that assertion.
- The
src/subdirectory hole was latent — zero violations existed there when it was fixed. It is closed because "no violations today" is not a property that keeps holding, not because it was hiding anything.
3970. Per-Task External-Tracker Content-Resolution Seam
Purpose: Let a capability declare that an external issue tracker — beads, Linear, Jira,
GitHub Issues — owns a task's content (<action>/<verify>/<acceptance_criteria>/
<read_first>/<done>), not just its status, so execute-plan.md can resolve that content
from the tracker at execution time instead of reading it inline out of PLAN.md.
What changed (ADR-3646, #3970):
- A new optional feature-body manifest field,
taskContentResolver, declares atrackerPrefix(matched against a task's<task tracker-id="beads:GSD-42">attribute — everything before the first:) and a boundedinvoke(binary,argscarrying the{{id}}placeholder,timeoutMs). execute-plan.md's per-task loop gains one new, unconditional call before that task'sread_firstgate:gsd_run task resolve-content --plan <path> --task-id <tracker-id> --raw. A task with notracker-idattribute is unaffected — the call is only made when the attribute is present, and resolves instantly to a no-op for every project that declares none.- The safety property is a real process exit code, not a prose dispatch. No capability
registered for the tracker, or resolution succeeds with empty content, exits
0withresolved: falseand falls back to inlinePLAN.md— the one legitimate pre-migration boundary case. Resolution succeeding with non-empty content exits0withresolved: trueand itscontentsupersedes the task's inline fields for every downstream gate in the execute step. A resolver that is declared but fails — tracker unreachable, id not found, timeout, malformed JSON — makestask resolve-contentitself exit non-zero, whichexecute-plan.mdtreats as a hard halt: stop, surface the tracker-id/prefix/stderr, never fall back to stalePLAN.mdcontent. execute:taskis a new dispatch shape below wave granularity, deliberately not one of the 12 existing loop extension points (discuss:pre…ship:post) and not routed throughgsd_run loop render-hooks <point>/activeHooks. It exists because the existingstep/gateprose-dispatch mechanism cannot deliver a hard-halt guarantee while dispatch reliability at that layer is an open concern (#3647) — see ADR-3646's Context and Rejected Alternatives for the full reasoning.
See Develop a task-content resolver capability
for the authoring walkthrough, Capability manifest → taskContentResolver
for the field reference, and
loop-hook-dispatch.md
for how execute:task differs from the twelve prose-dispatched points.
4014. Unreadable-Directory Scope Signal
Purpose: ADR-3473 §8.4 ("failure is a value") applies to filesystem
listings, not only command argv. #3885 (B5) gave roadmap analyze,
gap-checker, and init's JSON bundles a context_read_error /
phase_dir_read_error string naming an unreadable phase directory — but the
underlying has_context / hasContext boolean stayed false either way,
so a consumer branching on that boolean alone still cannot tell "genuinely
no context file" from "could not read the directory at all." This closes
that gap with a typed signal, reusing ADR-3180's existing frozen SCOPE
enum rather than a new vocabulary.
findContextMdIn (src/planning-workspace.cts) now reports its own
scope. Called with a directory path, it returns { file, files, scope }
instead of a bare filename-or-null, and never throws — an unreadable
directory reports scope: 'unreadable' (previously it threw, forcing every
caller to hand-roll its own try/catch); a genuinely absent directory
(ENOENT) reports scope: 'complete', the same "real empty" answer as
today. The array-input call form (an already-read listing) is unchanged.
Five downstream call sites gain an additive scope field, none renamed
or removed: roadmap analyze's AnalyzePhase.context_scope,
gap-checker's phase_dir_scope, and init's context_scope on all three
JSON bundles (init plan-phase, init phase-op, init manager) —
including cmdInitManager, whose own read failure previously vanished into
a bare empty catch {} with no signal of any kind. getPhaseFileStats
(src/core-utils.cts) — the shared listing owner behind roadmap analyze
and init's has_context — no longer lets its own failed read get masked
by an unrelated, already-successful scanPhasePlans scope on the same
phase directory.
Known limits:
context_read_error/phase_dir_read_error's message text is now a fixed "Could not read phase directory<path>" rather than embedding the underlying OS errno text —findContextMdIn's directory-string form reports only theSCOPEdiscriminator, not the raw caught error. The field's presence and type are unchanged; only its message detail is coarser than before #4014.init.cts's three call sites callfindContextMdInfor the scope signal and then still run their own, pre-existingfs.readdirSyncon the same path for the rest of their output — an intentional, additive-only choice to avoid altering already-complex failure control-flow at those sites, not a performance optimization.
Generated by scripts/gen-features.cjs — add a fragment under docs/features/ and run --write.