From e7a6d9ef2ea33c26f67bb550efc06b227e789d1b Mon Sep 17 00:00:00 2001 From: Lex Christopherson Date: Mon, 20 Apr 2026 14:04:31 -0600 Subject: [PATCH] fix: sync spike/sketch workflows with upstream skill improvements MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Spike workflow: - Add prior spike check — skips already-validated questions - Add comparison spikes (NNN-a/NNN-b) for head-to-head evaluation - Add research-before-building step (context7 + web search) - Add forensic logging/observability for runtime-interactive spikes - Add Type column to MANIFEST, type/Research/Observability to README Sketch workflow: - Add research-the-target-stack step — check component availability, framework constraints, and idiomatic patterns before building Spike wrap-up workflow: - Replace per-spike curation with auto-include-all (every spike carries signal: VALIDATED=patterns, PARTIAL=constraints, INVALIDATED=landmines) - Add Step 10 intelligent routing — integration spike candidates, frontier spike candidates, and standard next-step options Commands updated with context7/WebSearch tools and --text flag. Co-Authored-By: Claude Opus 4.6 (1M context) --- commands/gsd/sketch.md | 8 +- commands/gsd/spike-wrap-up.md | 2 +- commands/gsd/spike.md | 9 +- get-shit-done/workflows/sketch.md | 25 +++- get-shit-done/workflows/spike-wrap-up.md | 130 +++++++++++---------- get-shit-done/workflows/spike.md | 140 +++++++++++++++++++---- 6 files changed, 223 insertions(+), 91 deletions(-) diff --git a/commands/gsd/sketch.md b/commands/gsd/sketch.md index 0f9ae883b..60614cc7d 100644 --- a/commands/gsd/sketch.md +++ b/commands/gsd/sketch.md @@ -1,7 +1,7 @@ --- name: gsd:sketch description: Rapidly sketch UI/design ideas using throwaway HTML mockups with multi-variant exploration -argument-hint: " [--quick]" +argument-hint: " [--quick] [--text]" allowed-tools: - Read - Write @@ -10,6 +10,10 @@ allowed-tools: - Grep - Glob - AskUserQuestion + - WebSearch + - WebFetch + - mcp__context7__resolve-library-id + - mcp__context7__query-docs --- Explore design directions through throwaway HTML mockups before committing to implementation. @@ -41,5 +45,5 @@ Design idea: $ARGUMENTS Execute the sketch workflow from @~/.claude/get-shit-done/workflows/sketch.md end-to-end. -Preserve all workflow gates (intake, decomposition, variant evaluation, MANIFEST updates, commit patterns). +Preserve all workflow gates (intake, decomposition, target stack research, variant evaluation, MANIFEST updates, commit patterns). diff --git a/commands/gsd/spike-wrap-up.md b/commands/gsd/spike-wrap-up.md index 5baaf6591..486d5377c 100644 --- a/commands/gsd/spike-wrap-up.md +++ b/commands/gsd/spike-wrap-up.md @@ -27,5 +27,5 @@ project history. Output skill goes to `./.claude/skills/spike-findings-[project] Execute the spike-wrap-up workflow from @~/.claude/get-shit-done/workflows/spike-wrap-up.md end-to-end. -Preserve all curation gates (per-spike review, grouping approval, CLAUDE.md routing line). +Preserve all workflow gates (auto-include, feature-area grouping, skill synthesis, CLAUDE.md routing line, intelligent next-step routing). diff --git a/commands/gsd/spike.md b/commands/gsd/spike.md index 05f5de96d..c99403e4a 100644 --- a/commands/gsd/spike.md +++ b/commands/gsd/spike.md @@ -1,7 +1,7 @@ --- name: gsd:spike description: Rapidly spike an idea with throwaway experiments to validate feasibility before planning -argument-hint: " [--quick]" +argument-hint: " [--quick] [--text]" allowed-tools: - Read - Write @@ -10,6 +10,10 @@ allowed-tools: - Grep - Glob - AskUserQuestion + - WebSearch + - WebFetch + - mcp__context7__resolve-library-id + - mcp__context7__query-docs --- Rapid feasibility validation through focused, throwaway experiments. Each spike answers one @@ -33,9 +37,10 @@ Idea: $ARGUMENTS **Available flags:** - `--quick` — Skip decomposition/alignment, jump straight to building. Use when you already know what to spike. +- `--text` — Use plain-text numbered lists instead of AskUserQuestion (for non-Claude runtimes). Execute the spike workflow from @~/.claude/get-shit-done/workflows/spike.md end-to-end. -Preserve all workflow gates (decomposition, risk ordering, verification, MANIFEST updates, commit patterns). +Preserve all workflow gates (prior spike check, decomposition, research, risk ordering, observability assessment, verification, MANIFEST updates, commit patterns). diff --git a/get-shit-done/workflows/sketch.md b/get-shit-done/workflows/sketch.md index 0a3d66419..e624835bd 100644 --- a/get-shit-done/workflows/sketch.md +++ b/get-shit-done/workflows/sketch.md @@ -87,6 +87,28 @@ Bad sketches: Present the table and get alignment before building. + +## Research the Target Stack + +Before sketching, ground the design in what's actually buildable. Sketches are HTML, but they should reflect real constraints of the target implementation. + +**a. Identify the target stack.** Check for package.json, Cargo.toml, etc. If the user mentioned a framework (React, SwiftUI, Flutter, etc.), note it. + +**b. Check component/pattern availability.** Use context7 (resolve-library-id → query-docs) or web search to answer: +- What layout primitives does the target framework provide? (grid systems, nav patterns, panel components) +- Are there existing component libraries in use? (shadcn, Material UI, etc.) What components are available? +- What interaction patterns are idiomatic? (e.g., sheet vs modal vs dialog in mobile) + +**c. Note constraints that affect design.** Some things that look great in HTML are painful or impossible in certain stacks: +- Platform conventions (iOS nav patterns, desktop menu bars, terminal grid constraints) +- Framework limitations (what's easy vs requires custom work) +- Existing design tokens or theme systems already in the project + +**d. Let research inform variants.** Use findings to make variants that are actually buildable — at least one variant should follow the path of least resistance for the target stack. + +**Skip when unnecessary.** If it's a greenfield project with no stack chosen, or the user explicitly says "just explore visually, don't worry about implementation," skip this step entirely. The point is grounding, not gatekeeping. + + Create or update `.planning/sketches/MANIFEST.md`: @@ -250,7 +272,8 @@ After all sketches complete, present the summary: - [ ] `.planning/sketches/` created (auto-creates if needed, no project init required) - [ ] Design direction explored conversationally before any code (unless --quick) -- [ ] Each sketch has 2-3 variants for comparison +- [ ] Target stack researched — component availability, constraints, and idioms noted (unless greenfield/skipped) +- [ ] Each sketch has 2-3 variants for comparison (at least one follows path of least resistance for target stack) - [ ] User can open and interact with sketches in a browser - [ ] Winning variant selected and marked for each sketch - [ ] All variants preserved (winner marked, not others deleted) diff --git a/get-shit-done/workflows/spike-wrap-up.md b/get-shit-done/workflows/spike-wrap-up.md index ce543cc39..000bfcc42 100644 --- a/get-shit-done/workflows/spike-wrap-up.md +++ b/get-shit-done/workflows/spike-wrap-up.md @@ -41,53 +41,28 @@ COMMIT_DOCS=$(gsd-sdk query config-get commit_docs 2>/dev/null || echo "true") ``` - -## Curate Spikes One-at-a-Time + +## Auto-Include All Spikes -Present each unprocessed spike in ascending order. For each spike, show: +Include all unprocessed spikes automatically. Present a brief inventory showing what's being processed: -- **Spike number and name** -- **Validates:** the Given/When/Then from frontmatter -- **Verdict:** VALIDATED / INVALIDATED / PARTIAL -- **Tags:** from frontmatter -- **Key findings:** summarize the Results section from the README -- **Grey areas:** anything uncertain or partially proven +``` +Processing N spikes: + 001 — name (VALIDATED) + 002 — name (PARTIAL) + 003 — name (INVALIDATED) +``` -Then ask the user: - -╔══════════════════════════════════════════════════════════════╗ -║ CHECKPOINT: Decision Required ║ -╚══════════════════════════════════════════════════════════════╝ - -Spike {NNN}: {name} — {verdict} - -{key findings summary} - -────────────────────────────────────────────────────────────── -→ Include / Exclude / Partial / Help me UAT this -────────────────────────────────────────────────────────────── - -**If "Help me UAT this":** -1. Read the spike's README "How to Run" and "What to Expect" sections -2. Present step-by-step instructions -3. Ask: "Does this match what you expected?" -4. After UAT, return to the include/exclude/partial decision - -**If "Partial":** -Ask what specifically to include or exclude. Record their notes alongside the spike. +Every spike carries forward: +- **VALIDATED** spikes provide proven patterns +- **PARTIAL** spikes provide constrained patterns +- **INVALIDATED** spikes provide landmines and dead ends ## Auto-Group by Feature Area -After all spikes are curated: - -1. Read all included spikes' tags, names, `related` fields, and content -2. Propose feature-area groupings, e.g.: - - "**WebSocket Streaming** — spikes 001, 004, 007" - - "**Foo API Integration** — spikes 002, 003" - - "**PDF Parsing** — spike 005" -3. Present the grouping for approval — user may merge, split, rename, or rearrange +Group spikes by feature area based on tags, names, `related` fields, and content. Proceed directly into synthesis. Each group becomes one reference file in the generated skill. @@ -193,13 +168,9 @@ Write `.planning/spikes/WRAP-UP-SUMMARY.md` for project history: **Feature areas:** [list] **Skill output:** `./.claude/skills/spike-findings-[project]/` -## Included Spikes -| # | Name | Verdict | Feature Area | -|---|------|---------|--------------| - -## Excluded Spikes -| # | Name | Reason | -|---|------|--------| +## Processed Spikes +| # | Name | Type | Verdict | Feature Area | +|---|------|------|---------|--------------| ## Key Findings [consolidated findings summary] @@ -232,7 +203,7 @@ gsd-sdk query commit "docs(spike-wrap-up): package [N] spike findings into proje GSD ► SPIKE WRAP-UP COMPLETE ✓ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ -**Curated:** {N} spikes ({included} included, {excluded} excluded) +**Processed:** {N} spikes **Feature areas:** {list} **Skill:** `./.claude/skills/spike-findings-[project]/` **Summary:** `.planning/spikes/WRAP-UP-SUMMARY.md` @@ -240,34 +211,71 @@ gsd-sdk query commit "docs(spike-wrap-up): package [N] spike findings into proje The spike-findings skill will auto-load in future build conversations. ``` + -─────────────────────────────────────────────────────────────── + +## What's Next — Intelligent Spike Routing -## ▶ Next Up +Analyze the full spike landscape (MANIFEST.md, all curated findings, feature-area groupings, validated/invalidated/partial verdicts) and present three categories of next-step options: -**Start building** — plan the real implementation +### Category A: Integration Spikes — "Do any validated spikes need to be tested together?" -`/gsd-plan-phase` +Review every pair and cluster of VALIDATED spikes. Look for: -─────────────────────────────────────────────────────────────── +- **Shared resources:** Two spikes that both touch the same API, database, state, or data format but were tested independently. Will they conflict, race, or step on each other? +- **Data handoffs:** Spike A produces output that Spike B consumes. The formats were assumed compatible but never proven. +- **Timing/ordering:** Spikes that work in isolation but have sequencing dependencies in the real flow (e.g., auth must complete before streaming starts). +- **Resource contention:** Spikes that individually work but may compete for connections, memory, rate limits, or tokens when combined. -**Also available:** -- `/gsd-add-phase` — add a phase based on spike findings -- `/gsd-spike` — spike additional ideas -- `/gsd-explore` — continue exploring +If integration risks exist, present them as concrete proposed spikes: -─────────────────────────────────────────────────────────────── +> **Integration spike candidates:** +> - "Spikes 001 + 003 together: streaming through the authenticated connection" — these were tested separately but the real app needs both at once +> - "Spikes 002 + 005 data handoff: does the parser output match what the renderer expects?" + +If no meaningful integration risks exist, say so and skip this category. + +### Category B: Frontier Spikes — "What else should we spike?" + +Think laterally about the overall idea from MANIFEST.md and what's been proven so far. Consider: + +- **Gaps in the vision:** What does the user's idea need that hasn't been spiked yet? Look at the MANIFEST.md idea description and identify capabilities that are assumed but unproven. +- **Discovered dependencies:** Findings from completed spikes that reveal new questions. A spike that validated "X works" may imply "but we'd also need Y" — surface those implied needs. +- **Alternative approaches:** If any spike was PARTIAL or INVALIDATED, suggest a different angle to achieve the same goal. +- **Adjacent capabilities:** Things that aren't strictly required but would meaningfully improve the idea if feasible — worth a quick spike to find out. +- **Comparison opportunities:** If a spike used one library/approach and it worked but felt heavy or awkward, suggest a comparison spike with an alternative. + +Present frontier spikes as concrete proposals with names, validation questions (Given/When/Then), and risk-ordering: + +> **Frontier spike candidates:** +> 1. `NNN-descriptive-name` — Given [X], when [Y], then [Z]. *Why now: [reason this is the logical next thing to explore]* +> 2. `NNN-descriptive-name` — Given [X], when [Y], then [Z]. *Why now: [reason]* + +Number them continuing from the highest existing spike number. + +### Category C: Standard Options + +- `/gsd-plan-phase` — Start planning the real implementation +- `/gsd-add-phase` — Add a phase based on spike findings +- `/gsd-spike` — Spike additional ideas +- `/gsd-explore` — Continue exploring +- Other + +### Presenting the Options + +Present all applicable categories, then ask the user which direction to go. If the user picks a frontier or integration spike, write the spike definitions directly into `.planning/spikes/MANIFEST.md` (appending to the existing table) and kick off `/gsd-spike` with those spikes pre-defined — the user shouldn't have to re-describe what was just proposed. -- [ ] Every unprocessed spike presented for individual curation -- [ ] Feature-area grouping proposed and approved +- [ ] All unprocessed spikes auto-included and processed +- [ ] Spikes grouped by feature area - [ ] Spike-findings skill exists at `./.claude/skills/` with SKILL.md, references/, sources/ -- [ ] Core source files from included spikes copied into sources/ +- [ ] Core source files from all spikes copied into sources/ - [ ] Reference files contain validated patterns, code snippets, landmines, constraints - [ ] `.planning/spikes/WRAP-UP-SUMMARY.md` written for project history - [ ] Project CLAUDE.md has auto-load routing line -- [ ] Summary presented with next-step routing +- [ ] Summary presented +- [ ] Intelligent next-step analysis presented with integration spike candidates, frontier spike candidates, and standard options diff --git a/get-shit-done/workflows/spike.md b/get-shit-done/workflows/spike.md index f2db5752f..8d2178290 100644 --- a/get-shit-done/workflows/spike.md +++ b/get-shit-done/workflows/spike.md @@ -19,7 +19,10 @@ Read all files referenced by the invoking prompt's execution_context before star Parse `$ARGUMENTS` for: - `--quick` flag → set `QUICK_MODE=true` +- `--text` flag → set `TEXT_MODE=true` - Remaining text → the idea to spike + +**Text mode:** If TEXT_MODE is enabled, replace AskUserQuestion calls with plain-text numbered lists — emit the options and ask the user to type the number of their choice. @@ -56,31 +59,48 @@ Avoid unless the spike specifically requires it: - Env files or config systems — hardcode everything + +If `.planning/spikes/MANIFEST.md` exists, read it. Scan the verdicts, names, and validation questions of all prior spikes. When decomposing the new idea, cross-reference against this history: + +- **Skip already-validated questions.** If a prior spike proved "WebSocket streaming works" with a VALIDATED verdict, don't re-spike it. Note the prior spike number and move on. +- **Build on prior findings.** If a prior spike was INVALIDATED or PARTIAL, factor that into the new decomposition — don't repeat the same approach, and flag the constraint to the user. +- **Call out relevant prior art.** When presenting the decomposition, mention any prior spikes that overlap: "Spike 003 already validated X, so we can skip that and focus on Y." + +If no `.planning/spikes/MANIFEST.md` exists, skip this step. + + -**If `QUICK_MODE` is true:** Skip decomposition and alignment. Take the user's idea as a single spike question. Assign it spike number `001` (or next available). Jump to `build_spikes`. +**If `QUICK_MODE` is true:** Skip decomposition and alignment. Take the user's idea as a single spike question. Assign it spike number `001` (or next available). Jump to `research`. **Otherwise:** Break the idea into 2-5 independent questions that each prove something specific. Frame each as an informal Given/When/Then. Present as a table: ``` -| # | Spike | Validates (Given/When/Then) | Risk | -|---|-------|-----------------------------|------| -| 001 | websocket-streaming | Given a WS connection, when LLM streams tokens, then client receives chunks < 100ms | **High** | -| 002 | pdf-extraction | Given a multi-page PDF, when parsed with pdfjs, then structured text is extractable | Medium | +| # | Spike | Type | Validates (Given/When/Then) | Risk | +|---|-------|------|-----------------------------|------| +| 001 | websocket-streaming | standard | Given a WS connection, when LLM streams tokens, then client receives chunks < 100ms | **High** | +| 002a | pdf-parse-pdfjs | comparison | Given a multi-page PDF, when parsed with pdfjs, then structured text is extractable | Medium | +| 002b | pdf-parse-camelot | comparison | Given a multi-page PDF, when parsed with camelot, then structured text is extractable | Medium | ``` +**Spike types:** +- **standard** — one approach answering one question +- **comparison** — same question, different approaches. Use a shared number with lettered variants: `NNN-a-name` and `NNN-b-name`. Both built back-to-back, then head-to-head comparison. + Good spikes answer one specific feasibility question: - "Can we parse X format and extract Y?" — script that does it on a sample file - "How fast is X approach?" — benchmark with real-ish data - "Can we get X and Y to talk to each other?" — thinnest integration - "What does X feel like as a UI?" — minimal interactive prototype - "Does X API actually support Y?" — script that calls it and shows the response +- "Should we use X or Y for this?" — **comparison spike**: same thin proof built with both Bad spikes are too broad or don't produce observable output: - "Set up the project" — not a question, just busywork - "Design the architecture" — planning, not spiking - "Build the backend" — too broad, no specific question +- "Research best practices" — open-ended reading with no runnable output Order by risk — the spike most likely to kill the idea runs first. @@ -103,6 +123,33 @@ Present the ordered spike list and ask which to build: The user may reorder, merge, split, or skip spikes. Wait for alignment. + +## Research Before Building + +Before writing any spike code, ground each spike in reality. This prevents building against outdated APIs, picking the wrong library, or discovering mid-spike that the approach is impossible. + +For each spike about to be built: + +**a. Identify unknowns.** What libraries, APIs, protocols, or techniques does this spike depend on? What assumptions are you making about how they work? + +**b. Check current docs.** Use context7 (resolve-library-id → query-docs) for any library or framework involved. Use web search for APIs, services, or techniques without a context7 entry. Read actual documentation — not training data, which may be stale. + +**c. Validate feasibility before coding.** Specifically check: +- Does the API/library actually support what the spike assumes? (Check endpoints, methods, capabilities) +- What's the current recommended approach? (The "right way" changes — what was learned in training may be deprecated) +- Are there version constraints, breaking changes, or migration gotchas? +- Are there rate limits, auth requirements, or platform restrictions that would block the spike? + +**d. Pick the right tool.** If multiple libraries could solve the problem, briefly compare them on: current maintenance status, API fit for the specific spike question, and complexity. Pick the one that gets to a runnable answer fastest with the fewest surprises. + +**e. Capture research findings.** Add a `## Research` section to the spike's README (before `## How to Run`) with: +- Which docs were checked and key findings +- The chosen approach and why +- Any gotchas or constraints discovered + +**Skip research when unnecessary.** If the spike uses only well-known, stable tools already verified in this session, or if the entire spike is pure logic with no external dependencies, skip this step. The goal is grounding in reality, not busywork. + + Create or update `.planning/spikes/MANIFEST.md`: @@ -114,8 +161,11 @@ Create or update `.planning/spikes/MANIFEST.md`: ## Spikes -| # | Name | Validates | Verdict | Tags | -|---|------|-----------|---------|------| +| # | Name | Type | Validates | Verdict | Tags | +|---|------|------|-----------|---------|------| +| 001 | websocket-streaming | standard | WS connections can stream LLM output | VALIDATED | websocket, real-time | +| 002a | pdf-parse-pdfjs | comparison | PDF table extraction | WINNER | pdf, parsing | +| 002b | pdf-parse-camelot | comparison | PDF table extraction | — | pdf, parsing | ``` If MANIFEST.md already exists, append new spikes to the existing table. @@ -124,21 +174,50 @@ If MANIFEST.md already exists, append new spikes to the existing table. Build each spike sequentially, highest-risk first. +**Comparison spikes** use a shared number with lettered variants: `NNN-a-descriptive-name` and `NNN-b-descriptive-name`. Both answer the same question using different approaches. Build them back-to-back, then report a head-to-head comparison before moving on. Judge on criteria that matter for the real build: API ergonomics, output quality, complexity, performance, or whatever the user cares about. The comparison spike's verdict names the winner and why. + ### For Each Spike: **a.** Find next available number by checking existing `.planning/spikes/NNN-*/` directories. -Format: three-digit zero-padded + hyphenated descriptive name. +Format: three-digit zero-padded + hyphenated descriptive name. Comparison spikes: same number with letter suffix — `002a-pdf-parse-pdfjs`, `002b-pdf-parse-camelot`. **b.** Create the spike directory: `.planning/spikes/NNN-descriptive-name/` -**c.** Build the minimum code that answers the spike's question. Every line must serve the question — nothing incidental. If auth isn't the question, hardcode a token. If the database isn't the question, use a JSON file. Strip everything that doesn't directly answer "does X work?" +**c.** Assess observability needs before writing code. Ask: **can Claude fully verify this spike's outcome by running a command and reading stdout, or does it require human interaction with a runtime?** -**d.** Write `README.md` with YAML frontmatter: +Spikes that need runtime observability: +- **UI spikes** — anything with a browser, clicks, visual feedback +- **Streaming spikes** — WebSockets, SSE, real-time data flow +- **Multi-process spikes** — client/server, IPC, subprocess orchestration +- **Timing-sensitive spikes** — race conditions, debounce, polling, reconnection +- **External API spikes** — where the API response shape, latency, or error behavior matters for the verdict + +Spikes that do NOT need it: +- Pure computation (parse this file, transform this data) +- Single-run scripts with deterministic stdout +- Anything Claude can run and check the output of directly + +**If the spike needs runtime observability,** build a forensic log layer into the spike: + +1. **An event log array** at module level that captures every meaningful event with an ISO timestamp and a direction/category tag (e.g., `"user_input"`, `"api_response"`, `"sse_frame"`, `"error"`, `"state_change"`) +2. **A log export mechanism** appropriate to the spike's runtime: + - For server spikes: a `GET /api/export-log` endpoint returning downloadable JSON + - For CLI spikes: write `spike-log-{timestamp}.json` to the spike directory on exit or on signal + - For browser spikes: a visible "Export Log" button that triggers a JSON download +3. **A log summary** included in the export: total event counts by category, duration, errors detected, environment metadata +4. **Analysis helpers** if the event volume warrants it: a small script (bash/python) in the spike directory that extracts the signal from the log. Name it `analyze-log.sh` or similar. + +Keep the logging lightweight — an array push per event, not a logging framework. Inline it in the spike code. + +**d.** Build the minimum code that answers the spike's question (with the observability layer from step c if applicable). Every line must serve the question — nothing incidental. + +**e.** Write `README.md` with YAML frontmatter: ```markdown --- spike: NNN name: descriptive-name +type: standard validates: "Given [precondition], when [action], then [expected outcome]" verdict: PENDING related: [] @@ -150,19 +229,25 @@ tags: [tag1, tag2] ## What This Validates [The specific feasibility question, framed as Given/When/Then] +## Research +[Docs checked, key findings, chosen approach and why, gotchas discovered. Omit if no external dependencies.] + ## How to Run [Single command or short sequence to run the spike] ## What to Expect [Concrete observable outcomes: "When you click X, you should see Y within Z seconds"] +## Observability +[If this spike has a forensic log layer: describe what's captured, how to export the log, and how to analyze it. Omit for spikes without runtime observability.] + ## Results -[Filled in after running — verdict, evidence, surprises] +[Filled in after running — verdict, evidence, surprises. If a forensic log was exported, include key findings from the log analysis here.] ``` -**e.** Auto-link related spikes: read existing spike READMEs and infer relationships from tags, names, and descriptions. Write the `related` field silently. +**f.** Auto-link related spikes: read existing spike READMEs and infer relationships from tags, names, and descriptions. Write the `related` field silently. -**f.** Run and verify: +**g.** Run and verify: - If self-verifiable: run it, check output, update README verdict and Results section - If needs human judgment: run it, present instructions using a checkpoint box: @@ -179,16 +264,18 @@ tags: [tag1, tag2] → Does this match what you expected? Describe what you see. ────────────────────────────────────────────────────────────── -**g.** Update verdict to VALIDATED / INVALIDATED / PARTIAL. Update Results section with evidence. +- If the spike has a forensic log layer: after verification, export the log and include key findings in the Results section. If something went wrong, ask the user to export the log and provide it for diagnosis. -**h.** Update `.planning/spikes/MANIFEST.md` with the spike's row. +**h.** Update verdict to VALIDATED / INVALIDATED / PARTIAL (or WINNER for comparison spike winners). Update Results section with evidence. -**i.** Commit (if `COMMIT_DOCS` is true): +**i.** Update `.planning/spikes/MANIFEST.md` with the spike's row. + +**j.** Commit (if `COMMIT_DOCS` is true): ```bash gsd-sdk query commit "docs(spike-NNN): [VERDICT] — [key finding in one sentence]" .planning/spikes/NNN-descriptive-name/ .planning/spikes/MANIFEST.md ``` -**j.** Report before moving to next spike: +**k.** Report before moving to next spike: ``` ◆ Spike NNN: {name} Verdict: {VALIDATED ✓ / INVALIDATED ✗ / PARTIAL ⚠} @@ -196,7 +283,7 @@ gsd-sdk query commit "docs(spike-NNN): [VERDICT] — [key finding in one sentenc Impact: {effect on remaining spikes, if any} ``` -**k.** If a spike invalidates a core assumption: stop and present: +**l.** If a spike invalidates a core assumption: stop and present: ╔══════════════════════════════════════════════════════════════╗ ║ CHECKPOINT: Decision Required ║ @@ -223,10 +310,11 @@ After all spikes complete, present the consolidated report: ## Verdicts -| # | Name | Verdict | -|---|------|---------| -| 001 | {name} | ✓ VALIDATED | -| 002 | {name} | ✗ INVALIDATED | +| # | Name | Type | Verdict | +|---|------|------|---------| +| 001 | {name} | standard | ✓ VALIDATED | +| 002a | {name} | comparison | ✓ WINNER | +| 002b | {name} | comparison | — | ## Key Discoveries {surprises, gotchas, things that weren't expected} @@ -260,10 +348,14 @@ After all spikes complete, present the consolidated report: - [ ] `.planning/spikes/` created (auto-creates if needed, no project init required) +- [ ] Prior spikes checked — already-validated questions skipped, prior findings factored in +- [ ] Research grounded each spike in current docs before coding (unless pure logic/no deps) +- [ ] Comparison spikes built back-to-back with head-to-head verdict +- [ ] Spikes needing human interaction have forensic log layer (event capture, export, analysis) - [ ] Each spike answers one specific question with observable evidence -- [ ] Each spike README has complete frontmatter, run instructions, and results +- [ ] Each spike README has complete frontmatter (including type), run instructions, and results - [ ] User verified each spike (self-verified or human checkpoint) -- [ ] MANIFEST.md is current +- [ ] MANIFEST.md is current (with Type column) - [ ] Commits use `docs(spike-NNN): [VERDICT]` format - [ ] Consolidated report presented with next-step routing - [ ] If core assumption invalidated, execution stopped and user consulted