diff --git a/get-shit-done/references/goal-backward.md b/get-shit-done/references/goal-backward.md
index 03cd6ec9e..f103fb92a 100644
--- a/get-shit-done/references/goal-backward.md
+++ b/get-shit-done/references/goal-backward.md
@@ -1,286 +1,33 @@
-# Goal-Backward Planning
+# DEPRECATED: Goal-Backward Planning Reference
-How to derive requirements by working backwards from the goal, not forwards from tasks.
+**This reference has been consolidated into the gsd-planner agent.**
-
-**Forward planning asks:** "What should we build?"
-**Goal-backward planning asks:** "What must be TRUE for the goal to be achieved?"
+## Migration
-Forward planning produces tasks. Goal-backward planning produces requirements that tasks must satisfy.
-
+Planning expertise is now baked into:
+- `agents/gsd-planner.md` - Section: ``
-
-Forward planning fails silently. A task like "create chat component" can be marked complete when the component is a placeholder. The task was done—a component was created—but the goal "working chat interface" was not achieved.
+## Why This Changed
-Goal-backward planning starts from "user can chat" and works backwards:
-- What must be TRUE for a user to chat?
-- What must EXIST for those truths to hold?
-- What must be WIRED for those artifacts to function?
+The thin orchestrator pattern consolidates all planning methodology into the agent:
+- Before: Reference files loaded separately (~287 lines)
+- After: Agent has expertise baked in, orchestrator is thin
-This produces must-haves that are verifiable. Either a user CAN chat, or they can't. No ambiguity.
-
+## Historical Reference
-
+This file previously contained:
+- Goal-backward vs forward planning distinction
+- Must-haves derivation process (5 steps)
+- Observable truths from user perspective
+- Required artifacts mapping
+- Required wiring analysis
+- Key links identification
+- must_haves YAML structure for PLAN.md frontmatter
+- Examples (e-commerce, settings, notifications)
+- Common failures and anti-patterns
-## Step 1: State the Goal
-
-Take the phase goal from ROADMAP.md. This is the outcome, not the work.
-
-**Examples:**
-- "Working chat interface" (not "build chat components")
-- "Users can authenticate" (not "implement auth system")
-- "Products display with prices" (not "create product pages")
-
-If the roadmap goal is task-shaped ("implement X"), reframe it as outcome-shaped ("X works").
-
-## Step 2: Derive Observable Truths
-
-Ask: **"What must be TRUE for this goal to be achieved?"**
-
-List 3-7 truths from the USER's perspective. These are observable behaviors, not implementation details.
-
-**For "working chat interface":**
-- User can see existing messages
-- User can type a new message
-- User can send the message
-- Sent message appears in the list
-- Messages persist across page refresh
-
-**For "users can authenticate":**
-- User can reach login page
-- User can enter credentials
-- Valid credentials grant access
-- Invalid credentials show error
-- Session persists across refresh
-- User can log out
-
-**Test:** Each truth should be verifiable by a human using the application. If you can't test it by clicking around, it's not observable.
-
-## Step 3: Derive Required Artifacts
-
-For each truth, ask: **"What must EXIST for this to be true?"**
-
-Map truths to concrete artifacts (files, routes, schemas, components).
-
-**"User can see existing messages" requires:**
-- Message list component (renders Message[])
-- Messages state (loaded from somewhere)
-- API route or data source (provides messages)
-- Message type definition (shapes the data)
-
-**"Valid credentials grant access" requires:**
-- Login form component (captures credentials)
-- Auth API route (validates credentials)
-- Session/token mechanism (persists auth state)
-- User record in database (to validate against)
-
-**Test:** Each artifact should be a specific file or database object. If you can't point to where it lives, it's too abstract.
-
-## Step 4: Derive Required Wiring
-
-For each artifact, ask: **"What must be CONNECTED for this artifact to function?"**
-
-Wiring is where most failures hide. The pieces exist but don't talk to each other.
-
-**Message list component wiring:**
-- Imports Message type (not using `any`)
-- Receives messages prop or fetches from API
-- Maps over messages to render (not hardcoded)
-- Handles empty state (not just crashes)
-
-**Auth API route wiring:**
-- Imports database client
-- Queries users table (not placeholder)
-- Compares password hash (not plaintext)
-- Returns session token (not empty response)
-
-**Test:** Wiring is verified by tracing data flow. Does A actually call B? Does B actually return to A? Does A actually use what B returned?
-
-## Step 5: Identify Key Links
-
-Ask: **"Where is this most likely to break?"**
-
-Key links are the critical connections that, if missing, cause cascading failures.
-
-**For chat interface:**
-- Input onSubmit → API call (if broken: typing works but sending doesn't)
-- API save → database (if broken: appears to send but doesn't persist)
-- Component → real data (if broken: shows placeholder, not messages)
-
-**For authentication:**
-- Form submit → API (if broken: form works but auth doesn't)
-- API → database query (if broken: accepts any password)
-- Session → protected routes (if broken: logged in but can't access anything)
-
-Key links get extra verification attention. These are where stubs and placeholders hide.
-
-
-
-
-The derive_must_haves step produces a structured list for PLAN.md frontmatter:
-
-```yaml
-must_haves:
- truths:
- - "User can see existing messages"
- - "User can send a message"
- - "Messages persist across refresh"
- artifacts:
- - path: "src/components/Chat.tsx"
- provides: "Message list rendering"
- - path: "src/app/api/chat/route.ts"
- provides: "Message CRUD operations"
- - path: "prisma/schema.prisma"
- provides: "Message model"
- key_links:
- - from: "Chat.tsx"
- to: "api/chat"
- via: "fetch in useEffect"
- - from: "api/chat POST"
- to: "database"
- via: "prisma.message.create"
-```
-
-This structure is machine-readable for verification after execution.
-
-
-
-
-## Example 1: E-commerce Product Page
-
-**Goal:** "Products display with prices and add-to-cart"
-
-**Truths:**
-- User can see product image
-- User can see product name and description
-- User can see product price
-- User can click "Add to Cart"
-- Cart updates when product added
-
-**Artifacts:**
-- `src/components/ProductCard.tsx` - displays product info
-- `src/components/AddToCart.tsx` - button with cart logic
-- `src/app/products/[id]/page.tsx` - product detail page
-- `src/hooks/useCart.ts` - cart state management
-- `prisma/schema.prisma` - Product model with price field
-
-**Key Links:**
-- ProductCard receives product data (not hardcoded)
-- AddToCart calls cart hook (not just console.log)
-- useCart persists to localStorage or API (not just memory)
-- Price displays from product.price (not placeholder "$XX.XX")
+All content preserved in `agents/gsd-planner.md`.
---
-
-## Example 2: User Settings Page
-
-**Goal:** "Users can update their profile settings"
-
-**Truths:**
-- User can see current settings values
-- User can edit each setting field
-- User can save changes
-- Saved changes persist
-- User sees confirmation of save
-
-**Artifacts:**
-- `src/app/settings/page.tsx` - settings page
-- `src/components/SettingsForm.tsx` - form with fields
-- `src/app/api/settings/route.ts` - GET and PUT endpoints
-- `prisma/schema.prisma` - User model with settings fields
-
-**Key Links:**
-- Form loads current values on mount (not empty defaults)
-- Submit calls API with form data (not console.log)
-- API updates database (not just returns success)
-- Success triggers UI feedback (not silent)
-
----
-
-## Example 3: Real-time Notifications
-
-**Goal:** "Users receive notifications in real-time"
-
-**Truths:**
-- User sees notification badge/indicator
-- New notifications appear without refresh
-- User can view notification list
-- User can mark notifications as read
-- Read state persists
-
-**Artifacts:**
-- `src/components/NotificationBell.tsx` - badge/indicator
-- `src/components/NotificationList.tsx` - dropdown/panel
-- `src/app/api/notifications/route.ts` - CRUD endpoints
-- `src/hooks/useNotifications.ts` - real-time subscription
-- `prisma/schema.prisma` - Notification model
-
-**Key Links:**
-- useNotifications connects to WebSocket/SSE (not polling placeholder)
-- NotificationBell shows actual unread count (not hardcoded)
-- Mark-as-read calls API (not just local state)
-- API broadcasts to other clients (if multi-device)
-
-
-
-
-
-## Failure: Truths Too Vague
-
-**Bad:** "User can use chat"
-**Good:** "User can see messages", "User can send message", "Messages persist"
-
-Vague truths can't be verified. Break them into specific, observable behaviors.
-
-## Failure: Artifacts Too Abstract
-
-**Bad:** "Chat system", "Auth module"
-**Good:** "src/components/Chat.tsx", "src/app/api/auth/login/route.ts"
-
-Abstract artifacts can't be checked. Point to specific files.
-
-## Failure: Missing Wiring
-
-**Bad:** Listing components without how they connect
-**Good:** "Chat.tsx fetches from /api/chat via useEffect on mount"
-
-Artifacts existing isn't enough. The connections between them are where stubs hide.
-
-## Failure: Skipping Key Links
-
-**Bad:** Assuming "if files exist, it works"
-**Good:** Identifying the 2-3 critical connections that make-or-break the goal
-
-Key links are verification priorities. Without them, you check everything equally (inefficient) or check nothing deeply (ineffective).
-
-
-
-
-
-## In plan-phase.md
-
-The `derive_must_haves` step runs after gathering context, before breaking into tasks.
-
-Output: `must_haves` structure written to PLAN.md frontmatter.
-
-Tasks are then designed to CREATE the artifacts and ESTABLISH the wiring.
-
-## In execute-phase.md
-
-The `verify_phase_goal` step runs after all plans execute, before updating roadmap.
-
-Input: `must_haves` from PLAN.md frontmatter (or derived from goal if missing).
-
-Process: Check each truth against codebase, verify artifacts exist and aren't stubs, trace key links.
-
-Output: VERIFICATION.md with pass/fail per item, fix recommendations if gaps found.
-
-## The Loop
-
-```
-derive_must_haves → tasks → execute → verify → [gaps?] → fix plans → execute → verify → pass
-```
-
-Must-haves are derived once, verified as many times as needed until all pass.
-
-
+*Deprecated: 2026-01-16*
+*Replaced by: agents/gsd-planner.md*
diff --git a/get-shit-done/references/plan-format.md b/get-shit-done/references/plan-format.md
index f7111448a..5f58d0161 100644
--- a/get-shit-done/references/plan-format.md
+++ b/get-shit-done/references/plan-format.md
@@ -1,473 +1,32 @@
-
-Claude-executable plans have a specific format that enables Claude to implement without interpretation. This reference defines what makes a plan executable vs. vague.
+# DEPRECATED: Plan Format Reference
-**Key insight:** PLAN.md IS the executable prompt. It contains everything Claude needs to execute the phase, including objective, context references, tasks, verification, success criteria, and output specification.
-
+**This reference has been consolidated into the gsd-planner agent.**
-
-A plan is Claude-executable when Claude can read the PLAN.md and immediately start implementing without asking clarifying questions.
+## Migration
-If Claude has to guess, interpret, or make assumptions - the task is too vague.
-
+Planning expertise is now baked into:
+- `agents/gsd-planner.md` - Section: ``
-
-Every PLAN.md starts with YAML frontmatter:
+## Why This Changed
+
+The thin orchestrator pattern consolidates all planning methodology into the agent:
+- Before: Reference files loaded separately (~474 lines)
+- After: Agent has expertise baked in, orchestrator is thin
+
+## Historical Reference
+
+This file previously contained:
+- PLAN.md frontmatter structure
+- XML prompt structure
+- Task anatomy (files, action, verify, done)
+- Task types (auto, checkpoint:*)
+- TDD plans guidance
+- Context references and anti-patterns
+- Specificity levels (too vague vs just right)
+- Task sizing guidance
+
+All content preserved in `agents/gsd-planner.md`.
-```yaml
---
-phase: XX-name
-plan: NN
-type: execute
-wave: N # Execution wave (1, 2, 3...). Pre-computed at plan time.
-depends_on: [] # Plan IDs this plan requires (e.g., ["01-01"])
-files_modified: [] # Files this plan modifies
-autonomous: true # false if plan has checkpoints
----
-```
-
-| Field | Required | Purpose |
-|-------|----------|---------|
-| `phase` | Yes | Phase identifier (e.g., `01-foundation`) |
-| `plan` | Yes | Plan number within phase (e.g., `01`, `02`) |
-| `type` | Yes | `execute` for standard plans, `tdd` for TDD plans |
-| `wave` | Yes | Execution wave number (1, 2, 3...). Pre-computed during planning. |
-| `depends_on` | Yes | Array of plan IDs this plan requires. |
-| `files_modified` | Yes | Files this plan touches. |
-| `autonomous` | Yes | `true` if no checkpoints, `false` if has checkpoints |
-
-**Wave is pre-computed:** `/gsd:plan-phase` assigns wave numbers based on `depends_on`. `/gsd:execute-phase` reads `wave` directly from frontmatter and groups plans by wave number. No runtime dependency analysis needed.
-
-**Checkpoint handling:** Plans with `autonomous: false` require user interaction. They run in their assigned wave but pause at checkpoints.
-
-
-
-Every PLAN.md follows this XML structure:
-
-```markdown
----
-phase: XX-name
-plan: NN
-type: execute
-wave: N
-depends_on: []
-files_modified: [path/to/file.ts]
-autonomous: true
----
-
-
-[What and why]
-Purpose: [...]
-Output: [...]
-
-
-
-@~/.claude/get-shit-done/workflows/execute-plan.md
-@~/.claude/get-shit-done/templates/summary.md
-[If checkpoints exist:]
-@~/.claude/get-shit-done/references/checkpoints.md
-
-
-
-@.planning/PROJECT.md
-@.planning/ROADMAP.md
-@.planning/STATE.md
-[Only if genuinely needed:]
-@.planning/phases/XX-name/XX-YY-SUMMARY.md
-@relevant/source/files.ts
-
-
-
-
- Task N: [Name]
- [paths]
- [what to do, what to avoid and WHY]
- [command/check]
- [criteria]
-
-
-
- [what Claude automated]
- [numbered verification steps]
- [how to continue - "approved" or describe issues]
-
-
-
- [what needs deciding]
- [why this matters]
-
-
-
-
- [how to indicate choice]
-
-
-
-
-[Overall phase checks]
-
-
-
-[Measurable completion]
-
-
-
-```
-
-
-
-
-Every task has four required fields:
-
-
-**What it is**: Exact file paths that will be created or modified.
-
-**Good**: `src/app/api/auth/login/route.ts`, `prisma/schema.prisma`
-**Bad**: "the auth files", "relevant components"
-
-Be specific. If you don't know the file path, figure it out first.
-
-
-
-**What it is**: Specific implementation instructions, including what to avoid and WHY.
-
-**Good**: "Create POST endpoint that accepts {email, password}, validates using bcrypt against User table, returns JWT in httpOnly cookie with 15-min expiry. Use jose library (not jsonwebtoken - CommonJS issues with Next.js Edge runtime)."
-
-**Bad**: "Add authentication", "Make login work"
-
-Include: technology choices, data structures, behavior details, pitfalls to avoid.
-
-
-
-**What it is**: How to prove the task is complete.
-
-**Good**:
-
-- `npm test` passes
-- `curl -X POST /api/auth/login` returns 200 with Set-Cookie header
-- Build completes without errors
-
-**Bad**: "It works", "Looks good", "User can log in"
-
-Must be executable - a command, a test, an observable behavior.
-
-
-
-**What it is**: Acceptance criteria - the measurable state of completion.
-
-**Good**: "Valid credentials return 200 + JWT cookie, invalid credentials return 401"
-
-**Bad**: "Authentication is complete"
-
-Should be testable without subjective judgment.
-
-
-
-
-Tasks have a `type` attribute that determines how they execute:
-
-
-**Default task type** - Claude executes autonomously.
-
-**Structure:**
-
-```xml
-
- Task 3: Create login endpoint with JWT
- src/app/api/auth/login/route.ts
- POST endpoint accepting {email, password}. Query User by email, compare password with bcrypt. On match, create JWT with jose library, set as httpOnly cookie (15-min expiry). Return 200. On mismatch, return 401.
- curl -X POST localhost:3000/api/auth/login returns 200 with Set-Cookie header
- Valid credentials → 200 + cookie. Invalid → 401.
-
-```
-
-Use for: Everything Claude can do independently (code, tests, builds, file operations).
-
-
-
-**RARELY USED** - Only for actions with NO CLI/API. Claude automates everything possible first.
-
-**Structure:**
-
-```xml
-
- [Unavoidable manual step - email link, 2FA code]
-
- [What Claude already automated]
- [The ONE thing requiring human action]
-
- [What Claude can check afterward]
- [How to continue]
-
-```
-
-Use ONLY for: Email verification links, SMS 2FA codes, manual approvals with no API, 3D Secure payment flows.
-
-Do NOT use for: Anything with a CLI (Vercel, Stripe, Upstash, Railway, GitHub), builds, tests, file creation, deployments.
-
-**Execution:** Claude automates everything with CLI/API, stops only for truly unavoidable manual steps.
-
-
-
-**Human must verify Claude's work** - Visual checks, UX testing.
-
-**Structure:**
-
-```xml
-
- Responsive dashboard layout
-
- 1. Run: npm run dev
- 2. Visit: http://localhost:3000/dashboard
- 3. Desktop (>1024px): Verify sidebar left, content right
- 4. Tablet (768px): Verify sidebar collapses to hamburger
- 5. Mobile (375px): Verify single column, bottom nav
- 6. Check: No layout shift, no horizontal scroll
-
- Type "approved" or describe issues
-
-```
-
-Use for: UI/UX verification, visual design checks, animation smoothness, accessibility testing.
-
-**Execution:** Claude builds the feature, stops, provides testing instructions, waits for approval/feedback.
-
-
-
-**Human must make implementation choice** - Direction-setting decisions.
-
-**Structure:**
-
-```xml
-
- Select authentication provider
- We need user authentication. Three approaches with different tradeoffs:
-
-
-
-
-
- Select: supabase, clerk, or nextauth
-
-```
-
-Use for: Technology selection, architecture decisions, design choices, feature prioritization.
-
-**Execution:** Claude presents options with balanced pros/cons, waits for decision, proceeds with chosen direction.
-
-
-**When to use checkpoints:**
-
-- Visual/UX verification (after Claude builds) → `checkpoint:human-verify`
-- Implementation direction choice → `checkpoint:decision`
-- Truly unavoidable manual actions (email links, 2FA) → `checkpoint:human-action` (rare)
-
-**When NOT to use checkpoints:**
-
-- Anything with CLI/API (Claude automates it) → `type="auto"`
-- Deployments (Vercel, Railway, Fly) → `type="auto"` with CLI
-- Creating resources (Upstash, Stripe, GitHub) → `type="auto"` with CLI/API
-- File operations, tests, builds → `type="auto"`
-
-**Golden rule:** If Claude CAN automate it, Claude MUST automate it.
-
-**Checkpoint impact on parallelization:**
-- Plans with checkpoints set `autonomous: false` in frontmatter
-- Non-autonomous plans execute after parallel wave or in main context
-- Subagent pauses at checkpoint, returns to orchestrator
-- Orchestrator presents checkpoint to user
-- User responds
-- Orchestrator resumes agent with `resume: agent_id`
-
-See `./checkpoints.md` for comprehensive checkpoint guidance.
-
-
-
-**TDD work uses dedicated plans.**
-
-TDD features require 2-3 execution cycles (RED → GREEN → REFACTOR), each with file reads, test runs, and potential debugging. This is fundamentally heavier than standard tasks and would consume 50-60% of context if embedded in a multi-task plan.
-
-**When to create a TDD plan:**
-- Business logic with defined inputs/outputs
-- API endpoints with request/response contracts
-- Data transformations and parsing
-- Validation rules
-- Algorithms with testable behavior
-
-**When to use standard plans (skip TDD):**
-- UI layout and styling
-- Configuration changes
-- Glue code connecting existing components
-- One-off scripts
-
-**Heuristic:** Can you write `expect(fn(input)).toBe(output)` before writing `fn`?
-→ Yes: Create a TDD plan (one feature per plan)
-→ No: Use standard plan, add tests after if needed
-
-See `./tdd.md` for TDD plan structure and execution guidance.
-
-
-
-Use @file references to load context for the prompt:
-
-```markdown
-
-@.planning/PROJECT.md # Project vision
-@.planning/ROADMAP.md # Phase structure
-@.planning/STATE.md # Current position
-
-# Only include prior SUMMARY if genuinely needed:
-# - This plan imports types from prior plan
-# - Prior plan made decision affecting this plan
-# Independent plans need NO prior SUMMARY references.
-
-@src/lib/db.ts # Existing database setup
-@src/types/user.ts # Existing type definitions
-
-```
-
-Reference files that Claude needs to understand before implementing.
-
-**Anti-pattern:** Reflexive chaining (02 refs 01, 03 refs 02). Only reference what you actually need.
-
-
-
-Overall phase verification (beyond individual task verification):
-
-```markdown
-
-Before declaring phase complete:
-- [ ] `npm run build` succeeds without errors
-- [ ] `npm test` passes all tests
-- [ ] No TypeScript errors
-- [ ] Feature works end-to-end manually
-
-```
-
-
-
-
-Measurable criteria for phase completion:
-
-```markdown
-
-
-- All tasks completed
-- All verification checks pass
-- No errors or warnings introduced
-- JWT auth flow works end-to-end
-- Protected routes redirect unauthenticated users
-
-```
-
-
-
-
-Specify the SUMMARY.md structure:
-
-```markdown
-
-```
-
-
-
-
-
-
-```xml
-
- Task 1: Add authentication
- ???
- Implement auth
- ???
- Users can authenticate
-
-```
-
-Claude: "How? What type? What library? Where?"
-
-
-
-
-```xml
-
- Task 1: Create login endpoint with JWT
- src/app/api/auth/login/route.ts
- POST endpoint accepting {email, password}. Query User by email, compare password with bcrypt. On match, create JWT with jose library, set as httpOnly cookie (15-min expiry). Return 200. On mismatch, return 401. Use jose instead of jsonwebtoken (CommonJS issues with Edge).
- curl -X POST localhost:3000/api/auth/login -H "Content-Type: application/json" -d '{"email":"test@test.com","password":"test123"}' returns 200 with Set-Cookie header containing JWT
- Valid credentials → 200 + cookie. Invalid → 401. Missing fields → 400.
-
-```
-
-Claude can implement this immediately.
-
-
-
-**TDD candidates get dedicated plans.**
-
-If email validation warrants TDD, create a TDD plan for it. See `./tdd.md` for TDD plan structure.
-
-
-
-Writing the actual code in the plan. Trust Claude to implement from clear instructions.
-
-
-
-
-
-
-- "Set up the infrastructure"
-- "Handle edge cases"
-- "Make it production-ready"
-- "Add proper error handling"
-
-These require Claude to decide WHAT to do. Specify it.
-
-
-
-
-- "It works correctly"
-- "User experience is good"
-- "Code is clean"
-- "Tests pass" (which tests? do they exist?)
-
-These require subjective judgment. Make it objective.
-
-
-
-
-- "Use the standard approach"
-- "Follow best practices"
-- "Like the other endpoints"
-
-Claude doesn't know your standards. Be explicit.
-
-
-
-
-Good task size: 15-60 minutes of Claude work.
-
-**Too small**: "Add import statement for bcrypt" (combine with related task)
-**Just right**: "Create login endpoint with JWT validation" (focused, specific)
-**Too big**: "Implement full authentication system" (split into multiple plans)
-
-If a task takes multiple sessions, break it down.
-If a task is trivial, combine with related tasks.
-
-**Note on scope:** If a phase has >3 tasks or spans multiple subsystems, split into multiple plans using the naming convention `{phase}-{plan}-PLAN.md`. See `./scope-estimation.md` for guidance.
-
+*Deprecated: 2026-01-16*
+*Replaced by: agents/gsd-planner.md*
diff --git a/get-shit-done/references/principles.md b/get-shit-done/references/principles.md
index 52fd4129e..49e033e09 100644
--- a/get-shit-done/references/principles.md
+++ b/get-shit-done/references/principles.md
@@ -1,73 +1,29 @@
-
+# DEPRECATED: GSD Principles
-Core principles for the GSD planning system.
+**This reference has been consolidated into the gsd-planner agent.**
-
+## Migration
-You are planning for ONE person (the user) and ONE implementer (Claude).
-- No teams, stakeholders, ceremonies, coordination overhead
-- User is the visionary/product owner
-- Claude is the builder
-- Estimate effort in Claude execution time, not human dev time
-
+Planning expertise is now baked into:
+- `agents/gsd-planner.md` - Section: ``
-
+## Why This Changed
-PLAN.md is not a document that gets transformed into a prompt.
-PLAN.md IS the prompt. It contains:
-- Objective (what and why)
-- Context (@file references)
-- Tasks (with verification criteria)
-- Success criteria (measurable)
+The thin orchestrator pattern consolidates all planning methodology into the agent:
+- Before: Reference files loaded separately (~74 lines)
+- After: Agent has expertise baked in, orchestrator is thin
-When planning a phase, you are writing the prompt that will execute it.
-
+## Historical Reference
-
+This file previously contained:
+- Solo developer + Claude workflow philosophy
+- "Plans are prompts" principle
+- Scope control and quality degradation curve
+- "Claude automates" and "ship fast" principles
+- Anti-enterprise patterns
-Plans must complete within reasonable context usage.
+All content preserved in `agents/gsd-planner.md`.
-**Quality degradation curve:**
-- 0-30% context: Peak quality
-- 30-50% context: Good quality
-- 50-70% context: Degrading quality
-- 70%+ context: Poor quality
-
-**Solution:** Aggressive atomicity - split into small, focused plans.
-- 2-3 tasks per plan maximum
-- Each plan independently executable
-- Better to have many small plans than few large ones
-
-
-
-
-If Claude CAN do it via CLI/API/tool, Claude MUST do it.
-
-Checkpoints are for:
-- **Verification** - Human confirms Claude's work (visual, UX)
-- **Decision** - Human makes implementation choice
-
-
-
-
-No enterprise process. No approval gates.
-
-Plan → Execute → Ship → Learn → Repeat
-
-Milestones mark shipped versions (v1.0 → v1.1 → v2.0).
-
-
-
-
-NEVER include:
-- Team structures, RACI matrices
-- Stakeholder management
-- Sprint ceremonies
-- Human dev time estimates (hours, days, weeks—Claude works differently)
-- Change management processes
-- Documentation for documentation's sake
-
-If it sounds like corporate PM theater, delete it.
-
-
-
+---
+*Deprecated: 2026-01-16*
+*Replaced by: agents/gsd-planner.md*
diff --git a/get-shit-done/references/scope-estimation.md b/get-shit-done/references/scope-estimation.md
index 3db1308ed..967fa69be 100644
--- a/get-shit-done/references/scope-estimation.md
+++ b/get-shit-done/references/scope-estimation.md
@@ -1,256 +1,32 @@
-
-Plans must maintain consistent quality from first task to last. This requires understanding quality degradation and splitting aggressively.
+# DEPRECATED: Scope Estimation Reference
-
-Claude degrades when it *perceives* context pressure and enters "completion mode."
+**This reference has been consolidated into the gsd-planner agent.**
-| Context Usage | Quality | Claude's State |
-|---------------|---------|----------------|
-| 0-30% | PEAK | Thorough, comprehensive |
-| 30-50% | GOOD | Confident, solid work |
-| 50-70% | DEGRADING | Efficiency mode begins |
-| 70%+ | POOR | Rushed, minimal |
+## Migration
-**The 40-50% inflection point:** Claude sees context mounting and thinks "I'd better conserve now." Result: "I'll complete the remaining tasks more concisely" = quality crash.
+Planning expertise is now baked into:
+- `agents/gsd-planner.md` - Section: ``
-**The rule:** Stop BEFORE quality degrades, not at context limit.
-
+## Why This Changed
-
-**Plans should complete within ~50% of context usage.**
+The thin orchestrator pattern consolidates all planning methodology into the agent:
+- Before: Reference files loaded separately (~257 lines)
+- After: Agent has expertise baked in, orchestrator is thin
-Why 50% not 80%?
-- No context anxiety possible
-- Quality maintained start to finish
-- Room for unexpected complexity
-- If you target 80%, you've already spent 40% in degradation mode
-
+## Historical Reference
-
-**Each plan: 2-3 tasks maximum. Stay under 50% context.**
+This file previously contained:
+- Quality degradation curve (0-30%, 30-50%, 50-70%, 70%+)
+- Context budget targets (~50%)
+- Task-per-plan rules (2-3 tasks)
+- Split signals (always split, consider splitting)
+- Splitting strategies (vertical slices preferred)
+- Dependency awareness and wave assignment
+- File ownership for parallel execution
+- Depth calibration (quick, standard, comprehensive)
-| Task Complexity | Tasks/Plan | Context/Task | Total |
-|-----------------|------------|--------------|-------|
-| Simple (CRUD, config) | 3 | ~10-15% | ~30-45% |
-| Complex (auth, payments) | 2 | ~20-30% | ~40-50% |
-| Very complex (migrations, refactors) | 1-2 | ~30-40% | ~30-50% |
+All content preserved in `agents/gsd-planner.md`.
-**When in doubt: Default to 2 tasks.** Better to have an extra plan than degraded quality.
-
-
-
-**TDD features get their own plans. Target ~40% context.**
-
-TDD requires 2-3 execution cycles (RED → GREEN → REFACTOR), each with file reads, test runs, and potential debugging. This is fundamentally heavier than linear task execution.
-
-| TDD Feature Complexity | Context Usage |
-|------------------------|---------------|
-| Simple utility function | ~25-30% |
-| Business logic with edge cases | ~35-40% |
-| Complex algorithm | ~40-50% |
-
-**One feature per TDD plan.** If features are trivial enough to batch, they're trivial enough to skip TDD.
-
-**Why TDD plans are separate:**
-- TDD consumes 40-50% context for a single feature
-- Dedicated plans ensure full quality throughout RED-GREEN-REFACTOR
-- Each TDD feature gets fresh context, peak quality
-
-See `~/.claude/get-shit-done/references/tdd.md` for TDD plan structure.
-
-
-
-
-
-- **More than 3 tasks** - Even if tasks seem small
-- **Multiple subsystems** - DB + API + UI = separate plans
-- **Any task with >5 file modifications** - Split by file groups
-- **Checkpoint + implementation work** - Checkpoints in one plan, implementation after in separate plan
-- **Discovery + implementation** - DISCOVERY.md in one plan, implementation in another
-
-
-
-- Estimated >5 files modified total
-- Complex domains (auth, payments, data modeling)
-- Any uncertainty about approach
-- Natural semantic boundaries (Setup -> Core -> Features)
-
-
-
-
-**Vertical slices (default):** Group by feature, not by layer.
-
-```
-PREFER: Plan 01 = User (model + API + UI)
- Plan 02 = Product (model + API + UI)
- Plan 03 = Order (model + API + UI)
-
-AVOID: Plan 01 = All models
- Plan 02 = All APIs (depends on 01)
- Plan 03 = All UIs (depends on 02)
-```
-
-Vertical slices maximize parallelism: [01, 02, 03] run simultaneously.
-Horizontal layers force sequential execution: 01 → 02 → 03.
-
-**By dependency:** Only when genuine dependencies exist.
-```
-Plan 01: Auth foundation (middleware, JWT utils)
-Plan 02: Protected features (uses auth from 01)
-```
-
-**By complexity:** When one slice is much heavier.
-```
-Plan 01: Dashboard layout shell
-Plan 02: Data fetching and state
-Plan 03: Visualization components
-```
-
-
-
-**Plans declare dependencies explicitly via frontmatter.**
-
-```yaml
-# Independent plan (Wave 1 candidate)
-depends_on: []
-files_modified: [src/features/user/model.ts, src/features/user/api.ts]
-autonomous: true
-
-# Dependent plan (later wave)
-depends_on: ["03-01"]
-files_modified: [src/integration/stripe.ts]
-autonomous: true
-```
-
-**Wave assignment rules:**
-- `depends_on: []` + no file conflicts → Wave 1 (parallel)
-- `depends_on: ["XX"]` → runs after plan XX completes
-- Shared `files_modified` with sibling → sequential (by plan number)
-
-**SUMMARY references:**
-- Only reference prior SUMMARY if genuinely needed (imported types, decisions affecting this plan)
-- Independent plans need NO prior SUMMARY references
-- Reflexive chaining (02 refs 01, 03 refs 02) is an anti-pattern
-
-
-
-**Exclusive file ownership prevents conflicts:**
-
-```yaml
-# Plan 01 frontmatter
-files_modified: [src/models/user.ts, src/api/users.ts, src/components/UserList.tsx]
-
-# Plan 02 frontmatter
-files_modified: [src/models/product.ts, src/api/products.ts, src/components/ProductList.tsx]
-```
-
-No overlap → can run parallel.
-
-**If file appears in multiple plans:** Later plan depends on earlier (by plan number).
-**If file cannot be split:** Plans must be sequential for that file.
-
-
-
-**Bad - Comprehensive plan:**
-```
-Plan: "Complete Authentication System"
-Tasks: 8 (models, migrations, API, JWT, middleware, hashing, login form, register form)
-Result: Task 1-3 good, Task 4-5 degrading, Task 6-8 rushed
-```
-
-**Good - Atomic plans:**
-```
-Plan 1: "Auth Database Models" (2 tasks)
-Plan 2: "Auth API Core" (3 tasks)
-Plan 3: "Auth API Protection" (2 tasks)
-Plan 4: "Auth UI Components" (2 tasks)
-Each: 30-40% context, peak quality, atomic commits
-```
-
-**Bad - Horizontal layers (sequential):**
-```
-Plan 01: Create User model, Product model, Order model
-Plan 02: Create /api/users, /api/products, /api/orders
-Plan 03: Create UserList UI, ProductList UI, OrderList UI
-```
-Result: 02 depends on 01, 03 depends on 02
-Waves: [01] → [02] → [03] (fully sequential)
-
-**Good - Vertical slices (parallel):**
-```
-Plan 01: User feature (model + API + UI)
-Plan 02: Product feature (model + API + UI)
-Plan 03: Order feature (model + API + UI)
-```
-Result: Each plan self-contained, no file overlap
-Waves: [01, 02, 03] (all parallel)
-
-
-
-| Files Modified | Context Impact |
-|----------------|----------------|
-| 0-3 files | ~10-15% (small) |
-| 4-6 files | ~20-30% (medium) |
-| 7+ files | ~40%+ (large - split) |
-
-| Complexity | Context/Task |
-|------------|--------------|
-| Simple CRUD | ~15% |
-| Business logic | ~25% |
-| Complex algorithms | ~40% |
-| Domain modeling | ~35% |
-
-**2 tasks:** Simple ~30%, Medium ~50%, Complex ~80% (split)
-**3 tasks:** Simple ~45%, Medium ~75% (risky), Complex 120% (impossible)
-
-
-
-**Depth controls compression tolerance, not artificial inflation.**
-
-| Depth | Typical Phases | Typical Plans/Phase | Tasks/Plan |
-|-------|----------------|---------------------|------------|
-| Quick | 3-5 | 1-3 | 2-3 |
-| Standard | 5-8 | 3-5 | 2-3 |
-| Comprehensive | 8-12 | 5-10 | 2-3 |
-
-Tasks/plan is CONSTANT at 2-3. The 50% context rule applies universally.
-
-**Key principle:** Derive from actual work. Depth determines how aggressively you combine things, not a target to hit.
-
-- Comprehensive auth = 8 plans (because auth genuinely has 8 concerns)
-- Comprehensive "add favicon" = 1 plan (because that's all it is)
-
-Don't pad small work to hit a number. Don't compress complex work to look efficient.
-
-**Comprehensive depth example:**
-Auth system at comprehensive depth = 8 plans (not 3 big ones):
-- 01: DB models (2 tasks)
-- 02: Password hashing (2 tasks)
-- 03: JWT generation (2 tasks)
-- 04: JWT validation middleware (2 tasks)
-- 05: Login endpoint (2 tasks)
-- 06: Register endpoint (2 tasks)
-- 07: Protected route patterns (2 tasks)
-- 08: Auth UI components (3 tasks)
-
-Each plan: fresh context, peak quality. More plans = more thoroughness, same quality per plan.
-
-
-
-**2-3 tasks, 50% context target:**
-- All tasks: Peak quality
-- Git: Atomic per-task commits
-- Parallel by default: Fresh context per subagent
-
-**The principle:** Aggressive atomicity. More plans, smaller scope, consistent quality.
-
-**The rules:**
-- If in doubt, split. Quality over consolidation.
-- Depth increases plan COUNT, never plan SIZE.
-- Vertical slices over horizontal layers.
-- Explicit dependencies via `depends_on` frontmatter.
-- Autonomous plans get parallel execution.
-
-**Commit rule:** Each plan produces 3-4 commits total (2-3 task commits + 1 docs commit).
-
-
+---
+*Deprecated: 2026-01-16*
+*Replaced by: agents/gsd-planner.md*