* fix: Created 10 headless prompt files (5 workflows + 5 agents) in sdk/p… - "sdk/prompts/workflows/execute-plan.md" - "sdk/prompts/workflows/research-phase.md" - "sdk/prompts/workflows/plan-phase.md" - "sdk/prompts/workflows/verify-phase.md" - "sdk/prompts/workflows/discuss-phase.md" - "sdk/prompts/agents/gsd-executor.md" - "sdk/prompts/agents/gsd-phase-researcher.md" - "sdk/prompts/agents/gsd-planner.md" GSD-Task: S01/T02 * feat: Created prompt-sanitizer.ts, wired headless prompt loading into P… - "sdk/src/prompt-sanitizer.ts" - "sdk/src/phase-prompt.ts" - "sdk/src/gsd-tools.ts" - "sdk/src/gsd-tools.test.ts" - "sdk/src/phase-runner-types.test.ts" GSD-Task: S01/T01 * test: Added 111 unit tests covering sanitizePrompt(), headless prompt l… - "sdk/src/prompt-sanitizer.test.ts" - "sdk/src/headless-prompts.test.ts" - "sdk/src/phase-prompt.test.ts" GSD-Task: S01/T03 * feat: Wired sdkPromptsDir preference and sanitizePrompt into InitRunner… - "sdk/src/init-runner.ts" - "sdk/package.json" GSD-Task: S02/T01 * feat: add --init flag to auto command for single-command PRD-to-execution gsd-sdk auto --init @path/to/prd.md now bootstraps the project (init) then immediately runs the autonomous phase execution loop. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: add remaining headless prompt files and templates Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test: Extended init-runner.test.ts with 7 sdkPromptsDir preference and… - "sdk/src/init-runner.test.ts" GSD-Task: S02/T03 --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
4.6 KiB
<core_principle> Task completion does not equal goal achievement.
A task "create chat component" can be marked complete when the component is a placeholder. The task was done — but the goal "working chat interface" was not achieved.
Goal-backward verification:
- What must be TRUE for the goal to be achieved?
- What must EXIST for those truths to hold?
- What must be WIRED for those artifacts to function?
Then verify each level against the actual codebase. </core_principle>
Load phase operation context from injected context files. Extract: phase directory, phase number, phase name, plan count.Load phase details, plans, and summaries. Extract the phase goal from the roadmap (the outcome to verify, not tasks) and requirements if they exist.
**Option A: Must-haves in PLAN frontmatter**Extract must_haves from each PLAN: { truths: [...], artifacts: [...], key_links: [...] }
Aggregate all must_haves across plans for phase-level verification.
Option B: Use Success Criteria from roadmap
If no must_haves in frontmatter, use Success Criteria directly as truths. Derive artifacts and key links from there.
Option C: Derive from phase goal (fallback)
If neither source available: state the goal, derive 3-7 observable truths, derive artifacts, derive key links.
For each observable truth, determine if the codebase enables it.Status: VERIFIED (all supporting artifacts pass) | FAILED (artifact missing/stub/unwired) | UNCERTAIN (needs investigation)
For each truth: identify supporting artifacts, check artifact status, check wiring, determine truth status.
Three-level verification:Level 1 — Exists: File exists on disk. Level 2 — Substantive: File has real content (not stub/placeholder). Check line count, expected patterns. Level 3 — Wired: File is imported AND used by other code.
| Exists | Substantive | Wired | Status |
|---|---|---|---|
| Yes | Yes | Yes | VERIFIED |
| Yes | Yes | No | ORPHANED |
| Yes | No | - | STUB |
| No | - | - | MISSING |
Verify each key link by checking imports, usage patterns, fetch calls, database queries, form handlers, and state rendering.
For each requirement mapped to this phase: identify supporting truths/artifacts, determine status (SATISFIED / BLOCKED / UNCERTAIN). Scan files modified in this phase for:| Pattern | Severity |
|---|---|
| TODO/FIXME/XXX/HACK | Warning |
| Placeholder content | Blocker |
| Empty returns | Warning |
| Log-only functions | Warning |
Categorize: Blocker (prevents goal) | Warning (incomplete) | Info (notable).
**passed:** All truths VERIFIED, all artifacts pass levels 1-3, all key links WIRED, no blocker anti-patterns.gaps_found: Any truth FAILED, artifact MISSING/STUB, key link NOT_WIRED, or blocker found.
Score: verified_truths / total_truths
If gaps_found: 1. Cluster related gaps by concern 2. Generate plan per cluster: objective, 2-3 tasks, re-verify step 3. Order by dependency: fix missing, fix stubs, fix wiring, verify Create VERIFICATION.md with: frontmatter (phase/timestamp/status/score), goal achievement, artifact table, wiring table, requirements coverage, anti-patterns, gaps summary, fix plans (if gaps_found). Return status (passed | gaps_found), score (N/M must-haves), report path.If gaps_found: list gaps and recommended fix plan names.
<success_criteria>
- Must-haves established (from frontmatter or derived)
- All truths verified with status and evidence
- All artifacts checked at all three levels
- All key links verified
- Requirements coverage assessed
- Anti-patterns scanned and categorized
- Overall status determined
- Fix plans generated (if gaps_found)
- VERIFICATION.md created with complete report
- Results returned to orchestrator </success_criteria>