fix: sync spike/sketch workflows with upstream skill improvements
Spike workflow: - Add prior spike check — skips already-validated questions - Add comparison spikes (NNN-a/NNN-b) for head-to-head evaluation - Add research-before-building step (context7 + web search) - Add forensic logging/observability for runtime-interactive spikes - Add Type column to MANIFEST, type/Research/Observability to README Sketch workflow: - Add research-the-target-stack step — check component availability, framework constraints, and idiomatic patterns before building Spike wrap-up workflow: - Replace per-spike curation with auto-include-all (every spike carries signal: VALIDATED=patterns, PARTIAL=constraints, INVALIDATED=landmines) - Add Step 10 intelligent routing — integration spike candidates, frontier spike candidates, and standard next-step options Commands updated with context7/WebSearch tools and --text flag. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: gsd:sketch
|
||||
description: Rapidly sketch UI/design ideas using throwaway HTML mockups with multi-variant exploration
|
||||
argument-hint: "<design idea to explore> [--quick]"
|
||||
argument-hint: "<design idea to explore> [--quick] [--text]"
|
||||
allowed-tools:
|
||||
- Read
|
||||
- Write
|
||||
@@ -10,6 +10,10 @@ allowed-tools:
|
||||
- Grep
|
||||
- Glob
|
||||
- AskUserQuestion
|
||||
- WebSearch
|
||||
- WebFetch
|
||||
- mcp__context7__resolve-library-id
|
||||
- mcp__context7__query-docs
|
||||
---
|
||||
<objective>
|
||||
Explore design directions through throwaway HTML mockups before committing to implementation.
|
||||
@@ -41,5 +45,5 @@ Design idea: $ARGUMENTS
|
||||
|
||||
<process>
|
||||
Execute the sketch workflow from @~/.claude/get-shit-done/workflows/sketch.md end-to-end.
|
||||
Preserve all workflow gates (intake, decomposition, variant evaluation, MANIFEST updates, commit patterns).
|
||||
Preserve all workflow gates (intake, decomposition, target stack research, variant evaluation, MANIFEST updates, commit patterns).
|
||||
</process>
|
||||
|
||||
@@ -27,5 +27,5 @@ project history. Output skill goes to `./.claude/skills/spike-findings-[project]
|
||||
|
||||
<process>
|
||||
Execute the spike-wrap-up workflow from @~/.claude/get-shit-done/workflows/spike-wrap-up.md end-to-end.
|
||||
Preserve all curation gates (per-spike review, grouping approval, CLAUDE.md routing line).
|
||||
Preserve all workflow gates (auto-include, feature-area grouping, skill synthesis, CLAUDE.md routing line, intelligent next-step routing).
|
||||
</process>
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: gsd:spike
|
||||
description: Rapidly spike an idea with throwaway experiments to validate feasibility before planning
|
||||
argument-hint: "<idea to validate> [--quick]"
|
||||
argument-hint: "<idea to validate> [--quick] [--text]"
|
||||
allowed-tools:
|
||||
- Read
|
||||
- Write
|
||||
@@ -10,6 +10,10 @@ allowed-tools:
|
||||
- Grep
|
||||
- Glob
|
||||
- AskUserQuestion
|
||||
- WebSearch
|
||||
- WebFetch
|
||||
- mcp__context7__resolve-library-id
|
||||
- mcp__context7__query-docs
|
||||
---
|
||||
<objective>
|
||||
Rapid feasibility validation through focused, throwaway experiments. Each spike answers one
|
||||
@@ -33,9 +37,10 @@ Idea: $ARGUMENTS
|
||||
|
||||
**Available flags:**
|
||||
- `--quick` — Skip decomposition/alignment, jump straight to building. Use when you already know what to spike.
|
||||
- `--text` — Use plain-text numbered lists instead of AskUserQuestion (for non-Claude runtimes).
|
||||
</context>
|
||||
|
||||
<process>
|
||||
Execute the spike workflow from @~/.claude/get-shit-done/workflows/spike.md end-to-end.
|
||||
Preserve all workflow gates (decomposition, risk ordering, verification, MANIFEST updates, commit patterns).
|
||||
Preserve all workflow gates (prior spike check, decomposition, research, risk ordering, observability assessment, verification, MANIFEST updates, commit patterns).
|
||||
</process>
|
||||
|
||||
@@ -92,6 +92,28 @@ Bad sketches:
|
||||
Present the table and get alignment before building.
|
||||
</step>
|
||||
|
||||
<step name="research_stack">
|
||||
## Research the Target Stack
|
||||
|
||||
Before sketching, ground the design in what's actually buildable. Sketches are HTML, but they should reflect real constraints of the target implementation.
|
||||
|
||||
**a. Identify the target stack.** Check for package.json, Cargo.toml, etc. If the user mentioned a framework (React, SwiftUI, Flutter, etc.), note it.
|
||||
|
||||
**b. Check component/pattern availability.** Use context7 (resolve-library-id → query-docs) or web search to answer:
|
||||
- What layout primitives does the target framework provide? (grid systems, nav patterns, panel components)
|
||||
- Are there existing component libraries in use? (shadcn, Material UI, etc.) What components are available?
|
||||
- What interaction patterns are idiomatic? (e.g., sheet vs modal vs dialog in mobile)
|
||||
|
||||
**c. Note constraints that affect design.** Some things that look great in HTML are painful or impossible in certain stacks:
|
||||
- Platform conventions (iOS nav patterns, desktop menu bars, terminal grid constraints)
|
||||
- Framework limitations (what's easy vs requires custom work)
|
||||
- Existing design tokens or theme systems already in the project
|
||||
|
||||
**d. Let research inform variants.** Use findings to make variants that are actually buildable — at least one variant should follow the path of least resistance for the target stack.
|
||||
|
||||
**Skip when unnecessary.** If it's a greenfield project with no stack chosen, or the user explicitly says "just explore visually, don't worry about implementation," skip this step entirely. The point is grounding, not gatekeeping.
|
||||
</step>
|
||||
|
||||
<step name="create_manifest">
|
||||
Create or update `.planning/sketches/MANIFEST.md`:
|
||||
|
||||
@@ -255,7 +277,8 @@ After all sketches complete, present the summary:
|
||||
<success_criteria>
|
||||
- [ ] `.planning/sketches/` created (auto-creates if needed, no project init required)
|
||||
- [ ] Design direction explored conversationally before any code (unless --quick)
|
||||
- [ ] Each sketch has 2-3 variants for comparison
|
||||
- [ ] Target stack researched — component availability, constraints, and idioms noted (unless greenfield/skipped)
|
||||
- [ ] Each sketch has 2-3 variants for comparison (at least one follows path of least resistance for target stack)
|
||||
- [ ] User can open and interact with sketches in a browser
|
||||
- [ ] Winning variant selected and marked for each sketch
|
||||
- [ ] All variants preserved (winner marked, not others deleted)
|
||||
|
||||
@@ -41,53 +41,28 @@ COMMIT_DOCS=$(gsd-sdk query config-get commit_docs 2>/dev/null || echo "true")
|
||||
```
|
||||
</step>
|
||||
|
||||
<step name="curate">
|
||||
## Curate Spikes One-at-a-Time
|
||||
<step name="auto_include">
|
||||
## Auto-Include All Spikes
|
||||
|
||||
Present each unprocessed spike in ascending order. For each spike, show:
|
||||
Include all unprocessed spikes automatically. Present a brief inventory showing what's being processed:
|
||||
|
||||
- **Spike number and name**
|
||||
- **Validates:** the Given/When/Then from frontmatter
|
||||
- **Verdict:** VALIDATED / INVALIDATED / PARTIAL
|
||||
- **Tags:** from frontmatter
|
||||
- **Key findings:** summarize the Results section from the README
|
||||
- **Grey areas:** anything uncertain or partially proven
|
||||
```
|
||||
Processing N spikes:
|
||||
001 — name (VALIDATED)
|
||||
002 — name (PARTIAL)
|
||||
003 — name (INVALIDATED)
|
||||
```
|
||||
|
||||
Then ask the user:
|
||||
|
||||
╔══════════════════════════════════════════════════════════════╗
|
||||
║ CHECKPOINT: Decision Required ║
|
||||
╚══════════════════════════════════════════════════════════════╝
|
||||
|
||||
Spike {NNN}: {name} — {verdict}
|
||||
|
||||
{key findings summary}
|
||||
|
||||
──────────────────────────────────────────────────────────────
|
||||
→ Include / Exclude / Partial / Help me UAT this
|
||||
──────────────────────────────────────────────────────────────
|
||||
|
||||
**If "Help me UAT this":**
|
||||
1. Read the spike's README "How to Run" and "What to Expect" sections
|
||||
2. Present step-by-step instructions
|
||||
3. Ask: "Does this match what you expected?"
|
||||
4. After UAT, return to the include/exclude/partial decision
|
||||
|
||||
**If "Partial":**
|
||||
Ask what specifically to include or exclude. Record their notes alongside the spike.
|
||||
Every spike carries forward:
|
||||
- **VALIDATED** spikes provide proven patterns
|
||||
- **PARTIAL** spikes provide constrained patterns
|
||||
- **INVALIDATED** spikes provide landmines and dead ends
|
||||
</step>
|
||||
|
||||
<step name="group">
|
||||
## Auto-Group by Feature Area
|
||||
|
||||
After all spikes are curated:
|
||||
|
||||
1. Read all included spikes' tags, names, `related` fields, and content
|
||||
2. Propose feature-area groupings, e.g.:
|
||||
- "**WebSocket Streaming** — spikes 001, 004, 007"
|
||||
- "**Foo API Integration** — spikes 002, 003"
|
||||
- "**PDF Parsing** — spike 005"
|
||||
3. Present the grouping for approval — user may merge, split, rename, or rearrange
|
||||
Group spikes by feature area based on tags, names, `related` fields, and content. Proceed directly into synthesis.
|
||||
|
||||
Each group becomes one reference file in the generated skill.
|
||||
</step>
|
||||
@@ -193,13 +168,9 @@ Write `.planning/spikes/WRAP-UP-SUMMARY.md` for project history:
|
||||
**Feature areas:** [list]
|
||||
**Skill output:** `./.claude/skills/spike-findings-[project]/`
|
||||
|
||||
## Included Spikes
|
||||
| # | Name | Verdict | Feature Area |
|
||||
|---|------|---------|--------------|
|
||||
|
||||
## Excluded Spikes
|
||||
| # | Name | Reason |
|
||||
|---|------|--------|
|
||||
## Processed Spikes
|
||||
| # | Name | Type | Verdict | Feature Area |
|
||||
|---|------|------|---------|--------------|
|
||||
|
||||
## Key Findings
|
||||
[consolidated findings summary]
|
||||
@@ -232,7 +203,7 @@ gsd-sdk query commit "docs(spike-wrap-up): package [N] spike findings into proje
|
||||
GSD ► SPIKE WRAP-UP COMPLETE ✓
|
||||
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
**Curated:** {N} spikes ({included} included, {excluded} excluded)
|
||||
**Processed:** {N} spikes
|
||||
**Feature areas:** {list}
|
||||
**Skill:** `./.claude/skills/spike-findings-[project]/`
|
||||
**Summary:** `.planning/spikes/WRAP-UP-SUMMARY.md`
|
||||
@@ -240,34 +211,71 @@ gsd-sdk query commit "docs(spike-wrap-up): package [N] spike findings into proje
|
||||
|
||||
The spike-findings skill will auto-load in future build conversations.
|
||||
```
|
||||
</step>
|
||||
|
||||
───────────────────────────────────────────────────────────────
|
||||
<step name="whats_next">
|
||||
## What's Next — Intelligent Spike Routing
|
||||
|
||||
## ▶ Next Up
|
||||
Analyze the full spike landscape (MANIFEST.md, all curated findings, feature-area groupings, validated/invalidated/partial verdicts) and present three categories of next-step options:
|
||||
|
||||
**Start building** — plan the real implementation
|
||||
### Category A: Integration Spikes — "Do any validated spikes need to be tested together?"
|
||||
|
||||
`/gsd-plan-phase`
|
||||
Review every pair and cluster of VALIDATED spikes. Look for:
|
||||
|
||||
───────────────────────────────────────────────────────────────
|
||||
- **Shared resources:** Two spikes that both touch the same API, database, state, or data format but were tested independently. Will they conflict, race, or step on each other?
|
||||
- **Data handoffs:** Spike A produces output that Spike B consumes. The formats were assumed compatible but never proven.
|
||||
- **Timing/ordering:** Spikes that work in isolation but have sequencing dependencies in the real flow (e.g., auth must complete before streaming starts).
|
||||
- **Resource contention:** Spikes that individually work but may compete for connections, memory, rate limits, or tokens when combined.
|
||||
|
||||
**Also available:**
|
||||
- `/gsd-add-phase` — add a phase based on spike findings
|
||||
- `/gsd-spike` — spike additional ideas
|
||||
- `/gsd-explore` — continue exploring
|
||||
If integration risks exist, present them as concrete proposed spikes:
|
||||
|
||||
───────────────────────────────────────────────────────────────
|
||||
> **Integration spike candidates:**
|
||||
> - "Spikes 001 + 003 together: streaming through the authenticated connection" — these were tested separately but the real app needs both at once
|
||||
> - "Spikes 002 + 005 data handoff: does the parser output match what the renderer expects?"
|
||||
|
||||
If no meaningful integration risks exist, say so and skip this category.
|
||||
|
||||
### Category B: Frontier Spikes — "What else should we spike?"
|
||||
|
||||
Think laterally about the overall idea from MANIFEST.md and what's been proven so far. Consider:
|
||||
|
||||
- **Gaps in the vision:** What does the user's idea need that hasn't been spiked yet? Look at the MANIFEST.md idea description and identify capabilities that are assumed but unproven.
|
||||
- **Discovered dependencies:** Findings from completed spikes that reveal new questions. A spike that validated "X works" may imply "but we'd also need Y" — surface those implied needs.
|
||||
- **Alternative approaches:** If any spike was PARTIAL or INVALIDATED, suggest a different angle to achieve the same goal.
|
||||
- **Adjacent capabilities:** Things that aren't strictly required but would meaningfully improve the idea if feasible — worth a quick spike to find out.
|
||||
- **Comparison opportunities:** If a spike used one library/approach and it worked but felt heavy or awkward, suggest a comparison spike with an alternative.
|
||||
|
||||
Present frontier spikes as concrete proposals with names, validation questions (Given/When/Then), and risk-ordering:
|
||||
|
||||
> **Frontier spike candidates:**
|
||||
> 1. `NNN-descriptive-name` — Given [X], when [Y], then [Z]. *Why now: [reason this is the logical next thing to explore]*
|
||||
> 2. `NNN-descriptive-name` — Given [X], when [Y], then [Z]. *Why now: [reason]*
|
||||
|
||||
Number them continuing from the highest existing spike number.
|
||||
|
||||
### Category C: Standard Options
|
||||
|
||||
- `/gsd-plan-phase` — Start planning the real implementation
|
||||
- `/gsd-add-phase` — Add a phase based on spike findings
|
||||
- `/gsd-spike` — Spike additional ideas
|
||||
- `/gsd-explore` — Continue exploring
|
||||
- Other
|
||||
|
||||
### Presenting the Options
|
||||
|
||||
Present all applicable categories, then ask the user which direction to go. If the user picks a frontier or integration spike, write the spike definitions directly into `.planning/spikes/MANIFEST.md` (appending to the existing table) and kick off `/gsd-spike` with those spikes pre-defined — the user shouldn't have to re-describe what was just proposed.
|
||||
</step>
|
||||
|
||||
</process>
|
||||
|
||||
<success_criteria>
|
||||
- [ ] Every unprocessed spike presented for individual curation
|
||||
- [ ] Feature-area grouping proposed and approved
|
||||
- [ ] All unprocessed spikes auto-included and processed
|
||||
- [ ] Spikes grouped by feature area
|
||||
- [ ] Spike-findings skill exists at `./.claude/skills/` with SKILL.md, references/, sources/
|
||||
- [ ] Core source files from included spikes copied into sources/
|
||||
- [ ] Core source files from all spikes copied into sources/
|
||||
- [ ] Reference files contain validated patterns, code snippets, landmines, constraints
|
||||
- [ ] `.planning/spikes/WRAP-UP-SUMMARY.md` written for project history
|
||||
- [ ] Project CLAUDE.md has auto-load routing line
|
||||
- [ ] Summary presented with next-step routing
|
||||
- [ ] Summary presented
|
||||
- [ ] Intelligent next-step analysis presented with integration spike candidates, frontier spike candidates, and standard options
|
||||
</success_criteria>
|
||||
|
||||
@@ -19,7 +19,10 @@ Read all files referenced by the invoking prompt's execution_context before star
|
||||
|
||||
Parse `$ARGUMENTS` for:
|
||||
- `--quick` flag → set `QUICK_MODE=true`
|
||||
- `--text` flag → set `TEXT_MODE=true`
|
||||
- Remaining text → the idea to spike
|
||||
|
||||
**Text mode:** If TEXT_MODE is enabled, replace AskUserQuestion calls with plain-text numbered lists — emit the options and ask the user to type the number of their choice.
|
||||
</step>
|
||||
|
||||
<step name="setup_directory">
|
||||
@@ -56,31 +59,48 @@ Avoid unless the spike specifically requires it:
|
||||
- Env files or config systems — hardcode everything
|
||||
</step>
|
||||
|
||||
<step name="check_prior_spikes">
|
||||
If `.planning/spikes/MANIFEST.md` exists, read it. Scan the verdicts, names, and validation questions of all prior spikes. When decomposing the new idea, cross-reference against this history:
|
||||
|
||||
- **Skip already-validated questions.** If a prior spike proved "WebSocket streaming works" with a VALIDATED verdict, don't re-spike it. Note the prior spike number and move on.
|
||||
- **Build on prior findings.** If a prior spike was INVALIDATED or PARTIAL, factor that into the new decomposition — don't repeat the same approach, and flag the constraint to the user.
|
||||
- **Call out relevant prior art.** When presenting the decomposition, mention any prior spikes that overlap: "Spike 003 already validated X, so we can skip that and focus on Y."
|
||||
|
||||
If no `.planning/spikes/MANIFEST.md` exists, skip this step.
|
||||
</step>
|
||||
|
||||
<step name="decompose">
|
||||
**If `QUICK_MODE` is true:** Skip decomposition and alignment. Take the user's idea as a single spike question. Assign it spike number `001` (or next available). Jump to `build_spikes`.
|
||||
**If `QUICK_MODE` is true:** Skip decomposition and alignment. Take the user's idea as a single spike question. Assign it spike number `001` (or next available). Jump to `research`.
|
||||
|
||||
**Otherwise:**
|
||||
|
||||
Break the idea into 2-5 independent questions that each prove something specific. Frame each as an informal Given/When/Then. Present as a table:
|
||||
|
||||
```
|
||||
| # | Spike | Validates (Given/When/Then) | Risk |
|
||||
|---|-------|-----------------------------|------|
|
||||
| 001 | websocket-streaming | Given a WS connection, when LLM streams tokens, then client receives chunks < 100ms | **High** |
|
||||
| 002 | pdf-extraction | Given a multi-page PDF, when parsed with pdfjs, then structured text is extractable | Medium |
|
||||
| # | Spike | Type | Validates (Given/When/Then) | Risk |
|
||||
|---|-------|------|-----------------------------|------|
|
||||
| 001 | websocket-streaming | standard | Given a WS connection, when LLM streams tokens, then client receives chunks < 100ms | **High** |
|
||||
| 002a | pdf-parse-pdfjs | comparison | Given a multi-page PDF, when parsed with pdfjs, then structured text is extractable | Medium |
|
||||
| 002b | pdf-parse-camelot | comparison | Given a multi-page PDF, when parsed with camelot, then structured text is extractable | Medium |
|
||||
```
|
||||
|
||||
**Spike types:**
|
||||
- **standard** — one approach answering one question
|
||||
- **comparison** — same question, different approaches. Use a shared number with lettered variants: `NNN-a-name` and `NNN-b-name`. Both built back-to-back, then head-to-head comparison.
|
||||
|
||||
Good spikes answer one specific feasibility question:
|
||||
- "Can we parse X format and extract Y?" — script that does it on a sample file
|
||||
- "How fast is X approach?" — benchmark with real-ish data
|
||||
- "Can we get X and Y to talk to each other?" — thinnest integration
|
||||
- "What does X feel like as a UI?" — minimal interactive prototype
|
||||
- "Does X API actually support Y?" — script that calls it and shows the response
|
||||
- "Should we use X or Y for this?" — **comparison spike**: same thin proof built with both
|
||||
|
||||
Bad spikes are too broad or don't produce observable output:
|
||||
- "Set up the project" — not a question, just busywork
|
||||
- "Design the architecture" — planning, not spiking
|
||||
- "Build the backend" — too broad, no specific question
|
||||
- "Research best practices" — open-ended reading with no runnable output
|
||||
|
||||
Order by risk — the spike most likely to kill the idea runs first.
|
||||
</step>
|
||||
@@ -103,6 +123,33 @@ Present the ordered spike list and ask which to build:
|
||||
The user may reorder, merge, split, or skip spikes. Wait for alignment.
|
||||
</step>
|
||||
|
||||
<step name="research">
|
||||
## Research Before Building
|
||||
|
||||
Before writing any spike code, ground each spike in reality. This prevents building against outdated APIs, picking the wrong library, or discovering mid-spike that the approach is impossible.
|
||||
|
||||
For each spike about to be built:
|
||||
|
||||
**a. Identify unknowns.** What libraries, APIs, protocols, or techniques does this spike depend on? What assumptions are you making about how they work?
|
||||
|
||||
**b. Check current docs.** Use context7 (resolve-library-id → query-docs) for any library or framework involved. Use web search for APIs, services, or techniques without a context7 entry. Read actual documentation — not training data, which may be stale.
|
||||
|
||||
**c. Validate feasibility before coding.** Specifically check:
|
||||
- Does the API/library actually support what the spike assumes? (Check endpoints, methods, capabilities)
|
||||
- What's the current recommended approach? (The "right way" changes — what was learned in training may be deprecated)
|
||||
- Are there version constraints, breaking changes, or migration gotchas?
|
||||
- Are there rate limits, auth requirements, or platform restrictions that would block the spike?
|
||||
|
||||
**d. Pick the right tool.** If multiple libraries could solve the problem, briefly compare them on: current maintenance status, API fit for the specific spike question, and complexity. Pick the one that gets to a runnable answer fastest with the fewest surprises.
|
||||
|
||||
**e. Capture research findings.** Add a `## Research` section to the spike's README (before `## How to Run`) with:
|
||||
- Which docs were checked and key findings
|
||||
- The chosen approach and why
|
||||
- Any gotchas or constraints discovered
|
||||
|
||||
**Skip research when unnecessary.** If the spike uses only well-known, stable tools already verified in this session, or if the entire spike is pure logic with no external dependencies, skip this step. The goal is grounding in reality, not busywork.
|
||||
</step>
|
||||
|
||||
<step name="create_manifest">
|
||||
Create or update `.planning/spikes/MANIFEST.md`:
|
||||
|
||||
@@ -114,8 +161,11 @@ Create or update `.planning/spikes/MANIFEST.md`:
|
||||
|
||||
## Spikes
|
||||
|
||||
| # | Name | Validates | Verdict | Tags |
|
||||
|---|------|-----------|---------|------|
|
||||
| # | Name | Type | Validates | Verdict | Tags |
|
||||
|---|------|------|-----------|---------|------|
|
||||
| 001 | websocket-streaming | standard | WS connections can stream LLM output | VALIDATED | websocket, real-time |
|
||||
| 002a | pdf-parse-pdfjs | comparison | PDF table extraction | WINNER | pdf, parsing |
|
||||
| 002b | pdf-parse-camelot | comparison | PDF table extraction | — | pdf, parsing |
|
||||
```
|
||||
|
||||
If MANIFEST.md already exists, append new spikes to the existing table.
|
||||
@@ -124,21 +174,50 @@ If MANIFEST.md already exists, append new spikes to the existing table.
|
||||
<step name="build_spikes">
|
||||
Build each spike sequentially, highest-risk first.
|
||||
|
||||
**Comparison spikes** use a shared number with lettered variants: `NNN-a-descriptive-name` and `NNN-b-descriptive-name`. Both answer the same question using different approaches. Build them back-to-back, then report a head-to-head comparison before moving on. Judge on criteria that matter for the real build: API ergonomics, output quality, complexity, performance, or whatever the user cares about. The comparison spike's verdict names the winner and why.
|
||||
|
||||
### For Each Spike:
|
||||
|
||||
**a.** Find next available number by checking existing `.planning/spikes/NNN-*/` directories.
|
||||
Format: three-digit zero-padded + hyphenated descriptive name.
|
||||
Format: three-digit zero-padded + hyphenated descriptive name. Comparison spikes: same number with letter suffix — `002a-pdf-parse-pdfjs`, `002b-pdf-parse-camelot`.
|
||||
|
||||
**b.** Create the spike directory: `.planning/spikes/NNN-descriptive-name/`
|
||||
|
||||
**c.** Build the minimum code that answers the spike's question. Every line must serve the question — nothing incidental. If auth isn't the question, hardcode a token. If the database isn't the question, use a JSON file. Strip everything that doesn't directly answer "does X work?"
|
||||
**c.** Assess observability needs before writing code. Ask: **can Claude fully verify this spike's outcome by running a command and reading stdout, or does it require human interaction with a runtime?**
|
||||
|
||||
**d.** Write `README.md` with YAML frontmatter:
|
||||
Spikes that need runtime observability:
|
||||
- **UI spikes** — anything with a browser, clicks, visual feedback
|
||||
- **Streaming spikes** — WebSockets, SSE, real-time data flow
|
||||
- **Multi-process spikes** — client/server, IPC, subprocess orchestration
|
||||
- **Timing-sensitive spikes** — race conditions, debounce, polling, reconnection
|
||||
- **External API spikes** — where the API response shape, latency, or error behavior matters for the verdict
|
||||
|
||||
Spikes that do NOT need it:
|
||||
- Pure computation (parse this file, transform this data)
|
||||
- Single-run scripts with deterministic stdout
|
||||
- Anything Claude can run and check the output of directly
|
||||
|
||||
**If the spike needs runtime observability,** build a forensic log layer into the spike:
|
||||
|
||||
1. **An event log array** at module level that captures every meaningful event with an ISO timestamp and a direction/category tag (e.g., `"user_input"`, `"api_response"`, `"sse_frame"`, `"error"`, `"state_change"`)
|
||||
2. **A log export mechanism** appropriate to the spike's runtime:
|
||||
- For server spikes: a `GET /api/export-log` endpoint returning downloadable JSON
|
||||
- For CLI spikes: write `spike-log-{timestamp}.json` to the spike directory on exit or on signal
|
||||
- For browser spikes: a visible "Export Log" button that triggers a JSON download
|
||||
3. **A log summary** included in the export: total event counts by category, duration, errors detected, environment metadata
|
||||
4. **Analysis helpers** if the event volume warrants it: a small script (bash/python) in the spike directory that extracts the signal from the log. Name it `analyze-log.sh` or similar.
|
||||
|
||||
Keep the logging lightweight — an array push per event, not a logging framework. Inline it in the spike code.
|
||||
|
||||
**d.** Build the minimum code that answers the spike's question (with the observability layer from step c if applicable). Every line must serve the question — nothing incidental.
|
||||
|
||||
**e.** Write `README.md` with YAML frontmatter:
|
||||
|
||||
```markdown
|
||||
---
|
||||
spike: NNN
|
||||
name: descriptive-name
|
||||
type: standard
|
||||
validates: "Given [precondition], when [action], then [expected outcome]"
|
||||
verdict: PENDING
|
||||
related: []
|
||||
@@ -150,19 +229,25 @@ tags: [tag1, tag2]
|
||||
## What This Validates
|
||||
[The specific feasibility question, framed as Given/When/Then]
|
||||
|
||||
## Research
|
||||
[Docs checked, key findings, chosen approach and why, gotchas discovered. Omit if no external dependencies.]
|
||||
|
||||
## How to Run
|
||||
[Single command or short sequence to run the spike]
|
||||
|
||||
## What to Expect
|
||||
[Concrete observable outcomes: "When you click X, you should see Y within Z seconds"]
|
||||
|
||||
## Observability
|
||||
[If this spike has a forensic log layer: describe what's captured, how to export the log, and how to analyze it. Omit for spikes without runtime observability.]
|
||||
|
||||
## Results
|
||||
[Filled in after running — verdict, evidence, surprises]
|
||||
[Filled in after running — verdict, evidence, surprises. If a forensic log was exported, include key findings from the log analysis here.]
|
||||
```
|
||||
|
||||
**e.** Auto-link related spikes: read existing spike READMEs and infer relationships from tags, names, and descriptions. Write the `related` field silently.
|
||||
**f.** Auto-link related spikes: read existing spike READMEs and infer relationships from tags, names, and descriptions. Write the `related` field silently.
|
||||
|
||||
**f.** Run and verify:
|
||||
**g.** Run and verify:
|
||||
- If self-verifiable: run it, check output, update README verdict and Results section
|
||||
- If needs human judgment: run it, present instructions using a checkpoint box:
|
||||
|
||||
@@ -179,16 +264,18 @@ tags: [tag1, tag2]
|
||||
→ Does this match what you expected? Describe what you see.
|
||||
──────────────────────────────────────────────────────────────
|
||||
|
||||
**g.** Update verdict to VALIDATED / INVALIDATED / PARTIAL. Update Results section with evidence.
|
||||
- If the spike has a forensic log layer: after verification, export the log and include key findings in the Results section. If something went wrong, ask the user to export the log and provide it for diagnosis.
|
||||
|
||||
**h.** Update `.planning/spikes/MANIFEST.md` with the spike's row.
|
||||
**h.** Update verdict to VALIDATED / INVALIDATED / PARTIAL (or WINNER for comparison spike winners). Update Results section with evidence.
|
||||
|
||||
**i.** Commit (if `COMMIT_DOCS` is true):
|
||||
**i.** Update `.planning/spikes/MANIFEST.md` with the spike's row.
|
||||
|
||||
**j.** Commit (if `COMMIT_DOCS` is true):
|
||||
```bash
|
||||
gsd-sdk query commit "docs(spike-NNN): [VERDICT] — [key finding in one sentence]" .planning/spikes/NNN-descriptive-name/ .planning/spikes/MANIFEST.md
|
||||
```
|
||||
|
||||
**j.** Report before moving to next spike:
|
||||
**k.** Report before moving to next spike:
|
||||
```
|
||||
◆ Spike NNN: {name}
|
||||
Verdict: {VALIDATED ✓ / INVALIDATED ✗ / PARTIAL ⚠}
|
||||
@@ -196,7 +283,7 @@ gsd-sdk query commit "docs(spike-NNN): [VERDICT] — [key finding in one sentenc
|
||||
Impact: {effect on remaining spikes, if any}
|
||||
```
|
||||
|
||||
**k.** If a spike invalidates a core assumption: stop and present:
|
||||
**l.** If a spike invalidates a core assumption: stop and present:
|
||||
|
||||
╔══════════════════════════════════════════════════════════════╗
|
||||
║ CHECKPOINT: Decision Required ║
|
||||
@@ -223,10 +310,11 @@ After all spikes complete, present the consolidated report:
|
||||
|
||||
## Verdicts
|
||||
|
||||
| # | Name | Verdict |
|
||||
|---|------|---------|
|
||||
| 001 | {name} | ✓ VALIDATED |
|
||||
| 002 | {name} | ✗ INVALIDATED |
|
||||
| # | Name | Type | Verdict |
|
||||
|---|------|------|---------|
|
||||
| 001 | {name} | standard | ✓ VALIDATED |
|
||||
| 002a | {name} | comparison | ✓ WINNER |
|
||||
| 002b | {name} | comparison | — |
|
||||
|
||||
## Key Discoveries
|
||||
{surprises, gotchas, things that weren't expected}
|
||||
@@ -260,10 +348,14 @@ After all spikes complete, present the consolidated report:
|
||||
|
||||
<success_criteria>
|
||||
- [ ] `.planning/spikes/` created (auto-creates if needed, no project init required)
|
||||
- [ ] Prior spikes checked — already-validated questions skipped, prior findings factored in
|
||||
- [ ] Research grounded each spike in current docs before coding (unless pure logic/no deps)
|
||||
- [ ] Comparison spikes built back-to-back with head-to-head verdict
|
||||
- [ ] Spikes needing human interaction have forensic log layer (event capture, export, analysis)
|
||||
- [ ] Each spike answers one specific question with observable evidence
|
||||
- [ ] Each spike README has complete frontmatter, run instructions, and results
|
||||
- [ ] Each spike README has complete frontmatter (including type), run instructions, and results
|
||||
- [ ] User verified each spike (self-verified or human checkpoint)
|
||||
- [ ] MANIFEST.md is current
|
||||
- [ ] MANIFEST.md is current (with Type column)
|
||||
- [ ] Commits use `docs(spike-NNN): [VERDICT]` format
|
||||
- [ ] Consolidated report presented with next-step routing
|
||||
- [ ] If core assumption invalidated, execution stopped and user consulted
|
||||
|
||||
Reference in New Issue
Block a user