diff --git a/agents/gsd-phase-researcher.md b/agents/gsd-phase-researcher.md index f0df62543..461bc3482 100644 --- a/agents/gsd-phase-researcher.md +++ b/agents/gsd-phase-researcher.md @@ -307,23 +307,21 @@ Verified patterns from official sources: | Config file | {path or "none — see Wave 0"} | | Quick run command | `{command}` | | Full suite command | `{command}` | -| Estimated runtime | ~{N} seconds | ### Phase Requirements → Test Map | Req ID | Behavior | Test Type | Automated Command | File Exists? | |--------|----------|-----------|-------------------|-------------| -| REQ-XX | {behavior description} | unit | `pytest tests/test_{module}.py::test_{name} -x` | ✅ yes / ❌ Wave 0 gap | +| REQ-XX | {behavior} | unit | `pytest tests/test_{module}.py::test_{name} -x` | ✅ / ❌ Wave 0 | -### Nyquist Sampling Rate -- **Minimum sample interval:** After every committed task → run: `{quick run command}` -- **Full suite trigger:** Before merging final task of any plan wave -- **Phase-complete gate:** Full suite green before `/gsd:verify-work` runs -- **Estimated feedback latency per task:** ~{N} seconds +### Sampling Rate +- **Per task commit:** `{quick run command}` +- **Per wave merge:** `{full suite command}` +- **Phase gate:** Full suite green before `/gsd:verify-work` -### Wave 0 Gaps (must be created before implementation) +### Wave 0 Gaps - [ ] `{tests/test_file.py}` — covers REQ-{XX} -- [ ] `{tests/conftest.py}` — shared fixtures for phase {N} -- [ ] Framework install: `{command}` — if no framework detected +- [ ] `{tests/conftest.py}` — shared fixtures +- [ ] Framework install: `{command}` — if none detected *(If no gaps: "None — existing test infrastructure covers all phase requirements")* @@ -366,7 +364,7 @@ INIT=$(node ~/.claude/get-shit-done/bin/gsd-tools.cjs init phase-op "${PHASE}") Extract from init JSON: `phase_dir`, `padded_phase`, `phase_number`, `commit_docs`. -Also check Nyquist validation config — read `.planning/config.json` and check if `workflow.nyquist_validation` is `true`. If `true`, include the Validation Architecture section in RESEARCH.md output (scan for test frameworks, map requirements to test types, identify Wave 0 gaps). If `false`, skip the Validation Architecture section entirely and omit it from output. +Also read `.planning/config.json` — if `workflow.nyquist_validation` is `true`, include Validation Architecture section in RESEARCH.md. If `false`, skip it. Then read CONTEXT.md if exists: ```bash @@ -402,29 +400,16 @@ For each domain: Context7 first → Official docs → WebSearch → Cross-verify ## Step 4: Validation Architecture Research (if nyquist_validation enabled) -**Skip this step if** workflow.nyquist_validation is false in config. - -This step answers: "How will Claude's executor know, within seconds of committing each task, whether the output is correct?" +**Skip if** workflow.nyquist_validation is false. ### Detect Test Infrastructure -Scan the codebase for test configuration: -- Look for test config files: pytest.ini, pyproject.toml, jest.config.*, vitest.config.*, etc. -- Look for test directories: test/, tests/, __tests__/ -- Look for test files: *.test.*, *.spec.* -- Check package.json scripts for test commands +Scan for: test config files (pytest.ini, jest.config.*, vitest.config.*), test directories (test/, tests/, __tests__/), test files (*.test.*, *.spec.*), package.json test scripts. ### Map Requirements to Tests -For each requirement in : -- Identify the behavior to verify -- Determine test type: unit / integration / contract / smoke / e2e / manual-only -- Specify the automated command to run that test in < 30 seconds -- Flag if only verifiable manually (justify why) +For each phase requirement: identify behavior, determine test type (unit/integration/smoke/e2e/manual-only), specify automated command runnable in < 30 seconds, flag manual-only with justification. ### Identify Wave 0 Gaps -List test files, fixtures, or utilities that must be created BEFORE implementation: -- Missing test files for phase requirements -- Missing test framework configuration -- Missing shared fixtures or test utilities +List missing test files, framework config, or shared fixtures needed before implementation. ## Step 5: Quality Check diff --git a/agents/gsd-plan-checker.md b/agents/gsd-plan-checker.md index feada2e65..3ef73ea36 100644 --- a/agents/gsd-plan-checker.md +++ b/agents/gsd-plan-checker.md @@ -314,102 +314,48 @@ issue: ## Dimension 8: Nyquist Compliance - -Skip this entire dimension if: -- workflow.nyquist_validation is false in .planning/config.json -- The phase being checked has no RESEARCH.md (researcher was skipped) -- The RESEARCH.md has no "Validation Architecture" section (researcher ran without Nyquist) - -If skipped, output: "Dimension 8: SKIPPED (nyquist_validation disabled or not applicable)" - - - -This dimension enforces the Nyquist-Shannon Sampling Theorem for AI code generation: -if Claude's executor produces output at high frequency (one task per commit), feedback -must run at equally high frequency. A plan that produces code without pre-defined -automated verification is under-sampled — errors will be statistically missed. - -The gsd-phase-researcher already determined WHAT to test. This dimension verifies -that the planner correctly incorporated that information into the actual task plans. - +Skip if: `workflow.nyquist_validation` is false, phase has no RESEARCH.md, or RESEARCH.md has no "Validation Architecture" section. Output: "Dimension 8: SKIPPED (nyquist_validation disabled or not applicable)" ### Check 8a — Automated Verify Presence -For EACH `` element in EACH plan file for this phase: - -1. Does `` contain an `` command (or structured equivalent)? -2. If `` is absent or empty: - - Is there a Wave 0 dependency that creates the test before this task runs? - - If no Wave 0 dependency exists → **BLOCKING FAIL** -3. If `` says "MISSING": - - A Wave 0 task must reference the same test file path → verify this link is present - - If the link is broken → **BLOCKING FAIL** - -**PASS criteria:** Every task either has an `` verify command, OR explicitly -references a Wave 0 task that creates the test scaffold it depends on. +For each `` in each plan: +- `` must contain `` command, OR a Wave 0 dependency that creates the test first +- If `` is absent with no Wave 0 dependency → **BLOCKING FAIL** +- If `` says "MISSING", a Wave 0 task must reference the same test file path → **BLOCKING FAIL** if link broken ### Check 8b — Feedback Latency Assessment -Review each `` command in the plans: - -1. Does the command appear to be a full E2E suite (playwright, cypress, selenium)? - - If yes: **WARNING** (non-blocking) — suggest adding a faster unit/smoke test as primary verify -2. Does the command include `--watchAll` or equivalent watch mode flags? - - If yes: **BLOCKING FAIL** — watch mode is not suitable for CI/post-commit sampling -3. Does the command include `sleep`, `wait`, or arbitrary delays > 30 seconds? - - If yes: **WARNING** — flag as latency risk +For each `` command: +- Full E2E suite (playwright, cypress, selenium) → **WARNING** — suggest faster unit/smoke test +- Watch mode flags (`--watchAll`) → **BLOCKING FAIL** +- Delays > 30 seconds → **WARNING** ### Check 8c — Sampling Continuity -Review ALL tasks across ALL plans for this phase in wave order: - -1. Map each task to its wave number -2. For each consecutive window of 3 tasks in the same wave: at least 2 must have - an `` verify command (not just Wave 0 scaffolding) -3. If any 3 consecutive implementation tasks all lack automated verify: **BLOCKING FAIL** +Map tasks to waves. Per wave, any consecutive window of 3 implementation tasks must have ≥2 with `` verify. 3 consecutive without → **BLOCKING FAIL**. ### Check 8d — Wave 0 Completeness -If any plan contains `MISSING` or references Wave 0: +For each `MISSING` reference: +- Wave 0 task must exist with matching `` path +- Wave 0 plan must execute before dependent task +- Missing match → **BLOCKING FAIL** -1. Does a Wave 0 task exist for every MISSING reference? -2. Does the Wave 0 task's `` match the path referenced in the MISSING automated command? -3. Is the Wave 0 task in a plan that executes BEFORE the dependent task? - -**FAIL condition:** Any MISSING automated verify without a matching Wave 0 task. - -### Dimension 8 Output Block - -Include this block in the plan-checker report: +### Dimension 8 Output ``` ## Dimension 8: Nyquist Compliance -### Automated Verify Coverage -| Task | Plan | Wave | Automated Command | Latency | Status | -|------|------|------|-------------------|---------|--------| -| {task name} | {plan} | {wave} | `{command}` | ~{N}s | ✅ PASS / ❌ FAIL | +| Task | Plan | Wave | Automated Command | Status | +|------|------|------|-------------------|--------| +| {task} | {plan} | {wave} | `{command}` | ✅ / ❌ | -### Sampling Continuity Check -Wave {N}: {X}/{Y} tasks verified → ✅ PASS / ❌ FAIL - -### Wave 0 Completeness -- {test file} → Wave 0 task present ✅ / MISSING ❌ - -### Overall Nyquist Status: ✅ PASS / ❌ FAIL - -### Revision Instructions (if FAIL) -Return to planner with the following required changes: -{list of specific fixes needed} +Sampling: Wave {N}: {X}/{Y} verified → ✅ / ❌ +Wave 0: {test file} → ✅ present / ❌ MISSING +Overall: ✅ PASS / ❌ FAIL ``` -### Revision Loop Behavior - -If Dimension 8 FAILS: -- Return to `gsd-planner` with the specific revision instructions above -- The planner must address ALL failing checks before returning -- This follows the same loop behavior as existing dimensions -- Maximum 3 revision loops for Dimension 8 before escalating to user +If FAIL: return to planner with specific fixes. Same revision loop as other dimensions (max 3 loops). diff --git a/agents/gsd-planner.md b/agents/gsd-planner.md index 1612f19f6..b576531f6 100644 --- a/agents/gsd-planner.md +++ b/agents/gsd-planner.md @@ -157,21 +157,19 @@ Every task has four required fields: - Good: "Create POST endpoint accepting {email, password}, validates using bcrypt against User table, returns JWT in httpOnly cookie with 15-min expiry. Use jose library (not jsonwebtoken - CommonJS issues with Edge runtime)." - Bad: "Add authentication", "Make login work" -**:** How to prove the task is complete. Supports structured format: +**:** How to prove the task is complete. ```xml pytest tests/test_module.py::test_behavior -x - Optional: human-readable description of what to check - run after this task commits, before next task begins ``` - Good: Specific automated command that runs in < 60 seconds - Bad: "It works", "Looks good", manual-only verification -- Simple format also accepted: `npm test` passes, `curl -X POST /api/auth/login` returns 200 with Set-Cookie header +- Simple format also accepted: `npm test` passes, `curl -X POST /api/auth/login` returns 200 -**Nyquist Rule:** Every `` must include an `` command. If no test exists yet for this behavior, set `MISSING — Wave 0 must create {test_file} first` and create a Wave 0 task that generates the test scaffold. +**Nyquist Rule:** Every `` must include an `` command. If no test exists yet, set `MISSING — Wave 0 must create {test_file} first` and create a Wave 0 task that generates the test scaffold. **:** Acceptance criteria - measurable state of completion. - Good: "Valid credentials return 200 + JWT cookie, invalid credentials return 401" diff --git a/get-shit-done/templates/VALIDATION.md b/get-shit-done/templates/VALIDATION.md index 1fae33a69..d569841ee 100644 --- a/get-shit-done/templates/VALIDATION.md +++ b/get-shit-done/templates/VALIDATION.md @@ -9,9 +9,7 @@ created: {date} # Phase {N} — Validation Strategy -> Generated by `gsd-phase-researcher` during `/gsd:plan-phase {N}`. -> Updated by `gsd-plan-checker` after plan approval. -> Governs feedback sampling during `/gsd:execute-phase {N}`. +> Per-phase validation contract for feedback sampling during execution. --- @@ -20,22 +18,19 @@ created: {date} | Property | Value | |----------|-------| | **Framework** | {pytest 7.x / jest 29.x / vitest / go test / other} | -| **Config file** | {path/to/pytest.ini or "none — Wave 0 installs"} | -| **Quick run command** | `{e.g., pytest -x --tb=short}` | -| **Full suite command** | `{e.g., pytest tests/ --tb=short}` | +| **Config file** | {path or "none — Wave 0 installs"} | +| **Quick run command** | `{quick command}` | +| **Full suite command** | `{full command}` | | **Estimated runtime** | ~{N} seconds | -| **CI pipeline** | {.github/workflows/test.yml — exists / needs creation} | --- -## Nyquist Sampling Rate - -> The minimum feedback frequency required to reliably catch errors in this phase. +## Sampling Rate - **After every task commit:** Run `{quick run command}` - **After every plan wave:** Run `{full suite command}` - **Before `/gsd:verify-work`:** Full suite must be green -- **Maximum acceptable task feedback latency:** {N} seconds +- **Max feedback latency:** {N} seconds --- @@ -43,62 +38,39 @@ created: {date} | Task ID | Plan | Wave | Requirement | Test Type | Automated Command | File Exists | Status | |---------|------|------|-------------|-----------|-------------------|-------------|--------| -| {N}-01-01 | 01 | 1 | REQ-{XX} | unit | `pytest tests/test_{module}.py::test_{name} -x` | ✅ / ❌ W0 | ⬜ pending | -| {N}-01-02 | 01 | 1 | REQ-{XX} | integration | `pytest tests/test_{flow}.py -x` | ✅ / ❌ W0 | ⬜ pending | -| {N}-02-01 | 02 | 2 | REQ-{XX} | smoke | `curl -s {endpoint} \| grep {expected}` | ✅ N/A | ⬜ pending | +| {N}-01-01 | 01 | 1 | REQ-{XX} | unit | `{command}` | ✅ / ❌ W0 | ⬜ pending | -*Status values: ⬜ pending · ✅ green · ❌ red · ⚠️ flaky* +*Status: ⬜ pending · ✅ green · ❌ red · ⚠️ flaky* --- ## Wave 0 Requirements -> Test scaffolding committed BEFORE any implementation task. Executor runs Wave 0 first. - -- [ ] `{tests/test_file.py}` — stubs for REQ-{XX}, REQ-{XX} +- [ ] `{tests/test_file.py}` — stubs for REQ-{XX} - [ ] `{tests/conftest.py}` — shared fixtures - [ ] `{framework install}` — if no framework detected -*If none required: "Existing infrastructure covers all phase requirements — no Wave 0 test tasks needed."* +*If none: "Existing infrastructure covers all phase requirements."* --- ## Manual-Only Verifications -> Behaviors that genuinely cannot be automated, with justification. -> These are surfaced during `/gsd:verify-work` UAT. - | Behavior | Requirement | Why Manual | Test Instructions | |----------|-------------|------------|-------------------| -| {behavior} | REQ-{XX} | {reason: visual, third-party auth, physical device...} | {step-by-step} | +| {behavior} | REQ-{XX} | {reason} | {steps} | -*If none: "All phase behaviors have automated verification coverage."* +*If none: "All phase behaviors have automated verification."* --- ## Validation Sign-Off -Updated by `gsd-plan-checker` when plans are approved: - -- [ ] All tasks have `` verify commands or Wave 0 dependencies -- [ ] No 3 consecutive implementation tasks without automated verify (sampling continuity) -- [ ] Wave 0 test files cover all MISSING references -- [ ] No watch-mode flags in any automated command -- [ ] Feedback latency per task: < {N}s ✅ +- [ ] All tasks have `` verify or Wave 0 dependencies +- [ ] Sampling continuity: no 3 consecutive tasks without automated verify +- [ ] Wave 0 covers all MISSING references +- [ ] No watch-mode flags +- [ ] Feedback latency < {N}s - [ ] `nyquist_compliant: true` set in frontmatter -**Plan-checker approval:** {pending / approved on YYYY-MM-DD} - ---- - -## Execution Tracking - -Updated during `/gsd:execute-phase {N}`: - -| Wave | Tasks | Tests Run | Pass | Fail | Sampling Status | -|------|-------|-----------|------|------|-----------------| -| 0 | {N} | — | — | — | scaffold | -| 1 | {N} | {command} | {N} | {N} | ✅ sampled | -| 2 | {N} | {command} | {N} | {N} | ✅ sampled | - -**Phase validation complete:** {pending / YYYY-MM-DD HH:MM} +**Approval:** {pending / approved YYYY-MM-DD}