chore: remove old planning files

This commit is contained in:
Lex Christopherson
2026-01-21 10:17:14 -06:00
parent e7ceaf6483
commit 93b963c3fe
20 changed files with 0 additions and 6215 deletions

View File

@@ -1,135 +0,0 @@
# Roadmap: v1.9.0 Codebase Intelligence System
**Goal:** Make GSD feel intelligent and automagical in how it navigates and understands both greenfield and brownfield projects.
**Phases:** 5 (4 complete)
---
## Current Milestone: v1.9.0
### Phase 1: Foundation & Learning ✓
**Goal:** Establish index schema and incremental learning via PostToolUse hook
**Status:** Complete
**Plans:** 2/2
### Phase 2: Context Injection ✓
**Goal:** Inject codebase awareness into every session via SessionStart hook
**Status:** Complete
**Plans:** 2/2
### Phase 3: Brownfield & Integration ✓
**Goal:** Deep analysis command for existing codebases, workflow integration
**Status:** Complete
**Plans:** 3/3
### Phase 4: Semantic Intelligence & Scale ✓
**Goal:** Transform syntax-only indexing into semantic understanding with graph-based relationships
**Status:** Complete
**Plans:** 5/5
Plans:
- [x] 04-01-PLAN.md — SQLite graph layer with sql.js (Wave 1)
- [x] 04-02-PLAN.md — Graph-backed rich summary generation (Wave 2)
- [x] 04-03-PLAN.md — Semantic entity generation via Claude API (Wave 2)
- [x] 04-04-PLAN.md — CLI query interface for getDependents (Wave 3)
- [x] 04-05-PLAN.md — Wire plan-phase.md to inject intel into planner (Wave 3)
**Wave Structure:**
- Wave 1: 04-01 (SQLite foundation)
- Wave 2: 04-02, 04-03 (parallel - both depend only on 04-01)
- Wave 3: 04-04, 04-05 (parallel - consumption layer)
**Why this phase:**
- Current system provides "2-3 ls commands worth of information" (Claude's own assessment)
- Missing: what files actually DO, who uses them, blast radius of changes
- Senior engineers at top companies need real intelligence, not file counts
**Delivers:**
- SQLite graph layer (sql.js - zero native deps) for relationship queries
- Entity-based semantic documentation (Claude writes understanding, not just syntax)
- Semantic `/gsd:analyze-codebase` that creates initial entities
- Rich summary generation from accumulated semantic knowledge
- CLI query interface for "what uses this file?" queries
**Requirements:**
- INTEL-04: Entity files capture semantic understanding (purpose, what exports do)
- INTEL-05: Relationships queryable ("what uses this file?", "blast radius")
- INTEL-06: `/gsd:analyze-codebase` creates initial entity docs via Claude
- INTEL-07: Summary reflects accumulated semantic knowledge
**Success Criteria:**
1. Claude can answer "what uses src/lib/db.ts?" from SessionStart context
2. Summary includes file purposes, not just file counts
3. Transitive dependency queries work (blast radius)
4. Works at scale (500+ file codebases)
### Phase 5: Subagent Codebase Analysis
**Goal:** Prevent context exhaustion on large codebases by delegating analysis to subagents
**Depends on:** Phase 4
**Status:** Gap closure in progress
**Plans:** 4 plans (2 complete, 2 gap closure)
Plans:
- [x] 05-01-PLAN.md — Create gsd-entity-generator subagent (Wave 1)
- [x] 05-02-PLAN.md — Refactor Step 9 for subagent delegation (Wave 2) — partial, gap found
- [ ] 05-03-PLAN.md — Create gsd-indexer subagent for Steps 2-3 (Wave 3) — gap closure
- [ ] 05-04-PLAN.md — Refactor Steps 2-3 for subagent delegation (Wave 4) — gap closure
**Wave Structure:**
- Wave 1: 05-01 (entity generator agent)
- Wave 2: 05-02 (Step 9 refactor)
- Wave 3: 05-03 (indexer agent) — gap closure
- Wave 4: 05-04 (Steps 2-3 refactor + verification) — gap closure
**Why this phase:**
- Current entity generation loads file contents in orchestrator context
- On large codebases (500+ files), orchestrator exhausts context during file selection and batching
- Subagent delegation gives fresh 200k context for file processing
**Gap Found (05-VERIFICATION.md):**
- Original scope only addressed Step 9 (entity generation)
- Actual context exhaustion occurs during Steps 2-3 (indexing)
- Orchestrator reads ALL file contents during indexing, not just entity generation
- Need additional gsd-indexer subagent for Steps 2-3
**Delivers:**
- `gsd-entity-generator` subagent following gsd-codebase-mapper pattern
- `gsd-indexer` subagent for file reading and export/import extraction
- Refactored `/gsd:analyze-codebase` with full subagent delegation
- Preserved orchestrator context for large codebase analysis
**Requirements:**
- INTEL-08: Entity generation delegated to subagent (not inline) ✓
- INTEL-09: Subagent writes entities directly, returns statistics only ✓
- INTEL-10: Orchestrator passes file paths, not file contents — BLOCKED (needs 05-03, 05-04)
- INTEL-11: Indexing phase delegated to subagent (gap closure)
**Success Criteria:**
1. Entity generation works via subagent spawn ✓
2. Indexing works via subagent spawn (gap closure)
3. Orchestrator context preserved (no file contents loaded in orchestrator)
4. Entities correctly formatted and graph.db updated ✓
5. Works on 500+ file codebases without context exhaustion
---
## Traceability
| Requirement | Phase | Status |
|-------------|-------|--------|
| INTEL-01 | Phase 1 | ✓ Complete |
| INTEL-02 | Phase 2 | ✓ Complete |
| INTEL-03 | Phase 3 | ✓ Complete |
| INTEL-04 | Phase 4 | ✓ Complete (04-03) |
| INTEL-05 | Phase 4 | ✓ Complete (04-04) |
| INTEL-06 | Phase 4 | ✓ Complete (04-03) |
| INTEL-07 | Phase 4 | ✓ Complete (04-02) |
| INTEL-08 | Phase 5 | ✓ Complete (05-01, 05-02) |
| INTEL-09 | Phase 5 | ✓ Complete (05-01) |
| INTEL-10 | Phase 5 | Gap closure (05-03, 05-04) |
| INTEL-11 | Phase 5 | Gap closure (05-03, 05-04) |
---
*Created: 2026-01-19*
*Updated: 2026-01-20 — Phase 5 gap closure plans added (05-03, 05-04)*

View File

@@ -1,109 +0,0 @@
# Project State
## Project Reference
See: .planning/PROJECT.md (updated 2026-01-19)
**Core value:** Claude understands your codebase structure and conventions before it starts working — automatically
**Current focus:** v1.9.0 Codebase Intelligence System
## Current Position
Phase: 5 of 5 (Subagent Codebase Analysis)
Plan: 3 of 4 (gap closure in progress)
Status: In progress
Last activity: 2026-01-21 — Completed 05-03-PLAN.md (gsd-indexer agent definition)
Progress: [████████░░] 85%
## Performance Metrics
**Velocity:**
- Total plans completed: 15
- Average duration: 2.4 min
- Total execution time: 36 min
**By Phase:**
| Phase | Plans | Total | Avg/Plan |
|-------|-------|-------|----------|
| 1. Foundation & Learning | 2/2 | 7 min | 3.5 min |
| 2. Context Injection | 2/2 | 4 min | 2.0 min |
| 3. Brownfield & Integration | 3/3 | 6 min | 2.0 min |
| 4. Semantic Intelligence | 5/5 | 13 min | 2.6 min |
| 5. Subagent Analysis | 3/4 | 6 min | 2.0 min |
*Updated after each plan completion*
## Accumulated Context
### Decisions
| Decision | Phase | Rationale |
|----------|-------|-----------|
| index.json keyed by absolute path | 01-01 | O(1) lookup for file entries |
| JSON schema with version field | 01-01 | Enables future schema migrations |
| updated=null for initialization | 01-01 | Distinguishes init from update |
| Use heredoc for stdin testing | 01-02 | Pipe chaining has timing issues with async stdin |
| Extract 'default' as export name | 01-02 | Both 'default' and identifier recorded for default exports |
| Read file from disk for Edit tool | 01-02 | Edit only provides old_string/new_string, not full content |
| Regenerate conventions every index update | 02-01 | Detection is fast, avoids staleness issues |
| Skip 'default' in case detection | 02-01 | Keyword, not naming convention indicator |
| Single lowercase words as camelCase | 02-01 | Follows camelCase rules (e.g., 'main', 'app') |
| Use lookup tables for purposes | 02-01 | More maintainable than regex patterns |
| Target < 500 tokens for summary | 02-02 | Minimize context window usage |
| Top 5 directories, top 3 suffixes | 02-02 | Keep output concise |
| Command documents same regex as hook | 03-01 | Consistency between bulk scan and incremental updates |
| generateSummary in intel-index.js | 03-01 | Co-locate all intel generation; regenerate on every update |
| No FK constraints in graph schema | 04-01 | Entities can reference before target indexed |
| Virtual id from JSON body | 04-01 | Flexible node structure with unique constraint |
| Delete-then-insert for edges | 04-01 | Clean replacement removes stale links |
| Singleton WASM instance | 04-01 | Avoids repeated sql.js init overhead |
| 50 file limit per entity run | 04-03 | Prevents context window exhaustion |
| Batches of 10 for Task tool | 04-03 | Balances parallelization with overhead |
| Entity slug: path--segments--file-ext.md | 04-03 | Flat directory with reversible identification |
| LEFT JOIN allows forward references | 04-02 | Edges can exist before target nodes indexed |
| UNION in recursive CTE | 04-02 | Prevents infinite loops in cyclic graphs |
| maxDepth default of 5 | 04-02 | Prevents runaway queries on deep dependencies |
| Query mode read-only | 04-04 | Query actions don't persist to disk, safe operations |
| Default limit 10 for dependents | 04-04 | Prevents huge output for files with many dependents |
| Query routing before Write/Edit | 04-04 | Clean separation between query and indexing modes |
| Intel read in Step 7 with others | 04-05 | Keep all context file reads in one place |
| 2>/dev/null for missing intel | 04-05 | Graceful degradation when summary.md doesn't exist |
| Skip existing entities by default | 05-01 | Prevents overwriting manual edits to entities |
| Absolute path keys in gsd-indexer | 05-03 | Consistent with index.json schema, O(1) lookup |
### Pending Todos
- `/gsd:resume-work` decimal phase handling (deferred from v1.8.0)
### Roadmap Evolution
- Phase 5 added: Subagent Codebase Analysis
- Phase 5 gaps found: Need gsd-indexer subagent for Steps 2-3
- Phase 5 gap closure: Plans 05-03, 05-04 created for indexer agent and integration
### Blockers/Concerns
- `.planning/` is gitignored in GSD repo - intel files created but not committed (expected for project-local data)
- **RESOLVED:** Phase 5 context bottleneck addressed with gsd-indexer agent (05-03 complete, 05-04 pending)
## Session Continuity
Last session: 2026-01-21
Stopped at: Completed 05-03-PLAN.md (gsd-indexer agent definition)
Resume file: None
## Phase Progress
- Phase 1: Foundation & Learning ✓
- Phase 2: Context Injection ✓
- Phase 3: Brownfield & Integration ✓
- Phase 4: Semantic Intelligence & Scale ✓
- Phase 5: Subagent Analysis — IN PROGRESS (gap closure)
**Phase 5 status:**
- 05-01: gsd-entity-generator subagent ✓
- 05-02: analyze-codebase Step 9 refactor ✓
- 05-03: gsd-indexer subagent ✓
- 05-04: analyze-codebase Steps 2-3 refactor — PENDING

View File

@@ -1,334 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 01
type: execute
wave: 1
depends_on: []
files_modified:
- hooks/gsd-intel-index.js
- package.json
autonomous: true
must_haves:
truths:
- "Entity files sync to SQLite graph database on write"
- "Graph persists across hook invocations via graph.db file"
- "Wiki-links become edges in the graph"
artifacts:
- path: "hooks/gsd-intel-index.js"
provides: "SQLite graph sync on entity write"
contains: "initSqlJs"
- path: ".planning/intel/graph.db"
provides: "Persistent SQLite database"
key_links:
- from: "hooks/gsd-intel-index.js"
to: ".planning/intel/graph.db"
via: "sql.js export/import"
pattern: "db\\.export\\(\\)"
---
<objective>
Add SQLite graph layer to the codebase intelligence system using sql.js (WASM).
Purpose: Enable relationship queries ("what uses this file?", "blast radius") by storing entity relationships in a queryable graph database.
Output: Modified gsd-intel-index.js with SQLite sync, updated package.json with sql.js dependency.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/04-semantic-intelligence/04-RESEARCH.md
Key implementation details from research:
- sql.js is WASM SQLite (zero native deps)
- Schema: nodes table (JSON body with virtual id), edges table (source/target)
- Must export() and persist after every write
- Async init but sync operations
- Use ON CONFLICT REPLACE for upserts
Existing code to modify:
- hooks/gsd-intel-index.js already has: parseEntityFrontmatter(), extractWikiLinks(), regenerateEntitySummary()
- Entity files trigger regenerateEntitySummary() when written
</context>
<tasks>
<task type="auto">
<name>Task 1: Add sql.js dependency and graph schema</name>
<files>package.json, hooks/gsd-intel-index.js</files>
<action>
1. Add sql.js dependency to package.json:
```json
"dependencies": {
"sql.js": "^1.12.0"
}
```
2. At top of hooks/gsd-intel-index.js, add require and schema constant:
```javascript
const initSqlJs = require('sql.js');
// Graph database schema (simple-graph pattern)
const GRAPH_SCHEMA = `
CREATE TABLE IF NOT EXISTS nodes (
body TEXT,
id TEXT GENERATED ALWAYS AS (json_extract(body, '$.id')) VIRTUAL NOT NULL UNIQUE
);
CREATE INDEX IF NOT EXISTS id_idx ON nodes(id);
CREATE TABLE IF NOT EXISTS edges (
source TEXT NOT NULL,
target TEXT NOT NULL,
relationship TEXT DEFAULT 'depends_on',
UNIQUE(source, target, relationship) ON CONFLICT REPLACE
);
CREATE INDEX IF NOT EXISTS source_idx ON edges(source);
CREATE INDEX IF NOT EXISTS target_idx ON edges(target);
`;
```
Note: Intentionally no FOREIGN KEY constraints - entity A can reference entity B before B is indexed. Orphan edges are acceptable.
</action>
<verify>
- `grep -q "sql.js" package.json` returns 0
- `grep -q "GRAPH_SCHEMA" hooks/gsd-intel-index.js` returns 0
</verify>
<done>package.json has sql.js dependency, gsd-intel-index.js has schema constant</done>
</task>
<task type="auto">
<name>Task 2: Implement graph database helpers</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add helper functions after the GRAPH_SCHEMA constant:
```javascript
// Singleton SQL instance (async init, reuse across calls)
let sqlInstance = null;
/**
* Get or initialize sql.js instance
* Caches the SQL constructor for reuse
*/
async function getSQL() {
if (!sqlInstance) {
sqlInstance = await initSqlJs();
}
return sqlInstance;
}
/**
* Load or create the graph database
* Returns { db, dbPath } for operations and persistence
*/
async function loadGraphDatabase() {
const SQL = await getSQL();
const dbPath = path.join(process.cwd(), '.planning', 'intel', 'graph.db');
let db;
if (fs.existsSync(dbPath)) {
const buffer = fs.readFileSync(dbPath);
db = new SQL.Database(buffer);
} else {
db = new SQL.Database();
db.run(GRAPH_SCHEMA);
}
return { db, dbPath };
}
/**
* Persist database to disk
* Must call after every write operation
*/
function persistDatabase(db, dbPath) {
const data = db.export();
const buffer = Buffer.from(data);
fs.writeFileSync(dbPath, buffer);
}
```
Key design notes:
- getSQL() caches the WASM instance (expensive to init)
- loadGraphDatabase() handles both create and load
- persistDatabase() called after EVERY write (sql.js is in-memory only)
</action>
<verify>`grep -q "loadGraphDatabase" hooks/gsd-intel-index.js` returns 0</verify>
<done>Graph database helper functions exist in hook</done>
</task>
<task type="auto">
<name>Task 3: Sync entity to graph on write</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add syncEntityToGraph() function and integrate with existing entity handling:
```javascript
/**
* Sync entity file to graph database
* Called when an entity .md file is written
*
* @param {string} entityPath - Path to entity file
*/
async function syncEntityToGraph(entityPath) {
const intelDir = path.join(process.cwd(), '.planning', 'intel');
// Opt-in check (same as updateIndex)
if (!fs.existsSync(intelDir)) {
return;
}
try {
const { db, dbPath } = await loadGraphDatabase();
// Read entity file
const content = fs.readFileSync(entityPath, 'utf8');
const entityId = path.basename(entityPath, '.md').toLowerCase();
const frontmatter = parseEntityFrontmatter(content);
const links = extractWikiLinks(content);
// Build node JSON
const nodeBody = JSON.stringify({
id: entityId,
path: frontmatter.path || entityPath,
type: frontmatter.type || 'unknown',
updated: frontmatter.updated || new Date().toISOString().split('T')[0],
status: frontmatter.status || 'active'
});
// Upsert node (ON CONFLICT handled by schema)
db.run(
`INSERT INTO nodes (body) VALUES (?)
ON CONFLICT(id) DO UPDATE SET body = excluded.body`,
[nodeBody]
);
// Delete old edges for this source, insert new ones
db.run('DELETE FROM edges WHERE source = ?', [entityId]);
if (links.length > 0) {
const stmt = db.prepare('INSERT INTO edges (source, target) VALUES (?, ?)');
for (const target of links) {
stmt.run([entityId, target.toLowerCase()]);
}
stmt.free();
}
// Persist to disk (critical - sql.js is in-memory)
persistDatabase(db, dbPath);
db.close();
} catch (e) {
// Silent failure - never block Claude
// Graph sync is best-effort enhancement
}
}
```
Then modify the entity file handling in the stdin handler:
Find this section:
```javascript
// Handle entity file writes - regenerate summary
if (isEntityFile(filePath)) {
regenerateEntitySummary();
process.exit(0);
}
```
Change to:
```javascript
// Handle entity file writes - sync to graph, regenerate summary
if (isEntityFile(filePath)) {
syncEntityToGraph(filePath).then(() => {
regenerateEntitySummary();
process.exit(0);
}).catch(() => {
// Silent failure
process.exit(0);
});
return; // Don't exit synchronously, wait for async
}
```
Note the return statement - we need to wait for async graph sync before exiting.
</action>
<verify>
Run manual test:
```bash
cd /Users/lexchristopherson/Developer/claude-code-resources/get-shit-done
mkdir -p .planning/intel/entities
echo '---
path: /test/example.ts
type: util
updated: 2026-01-20
status: active
---
# example.ts
## Purpose
Test file for graph sync.
## Dependencies
- [[src-lib-db]]
## Used By
TBD
' > .planning/intel/entities/test-example.md
# Simulate hook execution
echo '{"tool_name":"Write","tool_input":{"file_path":".planning/intel/entities/test-example.md"}}' | node hooks/gsd-intel-index.js
# Check graph.db was created
ls -la .planning/intel/graph.db
# Cleanup
rm .planning/intel/entities/test-example.md
rm .planning/intel/graph.db 2>/dev/null
```
</verify>
<done>Entity writes sync to SQLite graph database, graph.db persists across invocations</done>
</task>
</tasks>
<verification>
After all tasks complete:
1. Dependency installed:
```bash
grep -q '"sql.js"' package.json && echo "PASS: sql.js in package.json"
```
2. Schema and helpers exist:
```bash
grep -q "GRAPH_SCHEMA" hooks/gsd-intel-index.js && echo "PASS: Schema defined"
grep -q "loadGraphDatabase" hooks/gsd-intel-index.js && echo "PASS: Helpers exist"
```
3. Graph sync works:
- Create test entity file with [[wiki-link]]
- Simulate Write hook
- Verify graph.db created
- Verify node and edge inserted (use sqlite3 CLI if available, or just check file size > 0)
</verification>
<success_criteria>
- [ ] sql.js added to package.json dependencies
- [ ] GRAPH_SCHEMA constant defines nodes and edges tables
- [ ] loadGraphDatabase() handles create and load
- [ ] persistDatabase() saves after writes
- [ ] syncEntityToGraph() upserts nodes and edges
- [ ] Entity file writes trigger graph sync before summary regeneration
- [ ] Silent failures don't block Claude
</success_criteria>
<output>
After completion, create `.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md`
</output>

View File

@@ -1,99 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 01
subsystem: database
tags: [sqlite, sql.js, wasm, graph-database, entity-relationships]
requires:
- phase: 03-brownfield-integration
provides: entity file system and summary regeneration
provides:
- SQLite graph database for entity relationships
- Node/edge schema for semantic queries
- Automatic graph sync on entity writes
affects: [04-02 query interface, future blast-radius queries]
tech-stack:
added: [sql.js ^1.12.0]
patterns: [simple-graph schema, WASM singleton, async-then-sync operations]
key-files:
created: []
modified: [hooks/gsd-intel-index.js, package.json]
key-decisions:
- "No FOREIGN KEY constraints - entities can reference before target indexed"
- "Virtual id column from JSON body for flexible node structure"
- "Delete-then-insert for edge updates (clean replacement)"
- "Singleton WASM instance to avoid repeated init overhead"
patterns-established:
- "Graph sync before summary regeneration"
- "Silent failure pattern for non-blocking operations"
duration: 4min
completed: 2026-01-20
---
# Phase 4 Plan 1: SQLite Graph Layer Summary
**SQLite graph database using sql.js WASM for entity relationship storage and querying**
## Performance
- **Duration:** 4 min
- **Started:** 2026-01-20T09:45:00Z
- **Completed:** 2026-01-20T09:53:00Z
- **Tasks:** 3
- **Files modified:** 2
## Accomplishments
- Added sql.js WASM SQLite dependency for zero-native-dependency graph storage
- Implemented simple-graph schema with nodes (JSON body) and edges (source/target/relationship)
- Graph database helpers for load, persist, and singleton WASM management
- Entity files now sync to graph database on every write, creating nodes and edges from wiki-links
## Task Commits
Each task was committed atomically:
1. **Task 1: Add sql.js dependency and graph schema** - `11ab4a9` (feat)
2. **Task 2: Implement graph database helpers** - `ec038e7` (feat)
3. **Task 3: Sync entity to graph on write** - `a39988e` (feat)
## Files Created/Modified
- `package.json` - Added sql.js ^1.12.0 dependency
- `hooks/gsd-intel-index.js` - GRAPH_SCHEMA constant, loadGraphDatabase(), persistDatabase(), getSQL(), syncEntityToGraph()
## Decisions Made
- **No FOREIGN KEY constraints:** Entity A can reference entity B before B is indexed. Orphan edges are acceptable and expected in incremental indexing workflows.
- **Virtual id column:** Uses `json_extract(body, '$.id')` for flexible node structure while maintaining unique constraint for upserts.
- **Delete-then-insert for edges:** Clean replacement of all outgoing edges on entity update ensures stale links are removed.
- **Singleton WASM instance:** sql.js WASM init is expensive; caching the SQL constructor avoids repeated overhead across hook invocations within same process.
## Deviations from Plan
None - plan executed exactly as written.
## Issues Encountered
None.
## User Setup Required
None - no external service configuration required.
## Next Phase Readiness
- Graph database infrastructure complete
- Ready for query interface implementation (Plan 04-02)
- Schema supports "what depends on X" and "what does X depend on" queries
---
*Phase: 04-semantic-intelligence*
*Completed: 2026-01-20*

View File

@@ -1,465 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 02
type: execute
wave: 2
depends_on: [04-01]
files_modified:
- hooks/gsd-intel-index.js
autonomous: true
must_haves:
truths:
- "Summary includes dependency hotspots queried from SQLite"
- "Summary shows file purposes, not just file counts"
- "Transitive dependents queryable via recursive CTE"
- "SessionStart hook injects graph-backed summary into context"
artifacts:
- path: "hooks/gsd-intel-index.js"
provides: "Graph-backed summary generation"
contains: "generateGraphSummary"
- path: ".planning/intel/summary.md"
provides: "Rich semantic summary for context injection"
key_links:
- from: "hooks/gsd-intel-index.js"
to: ".planning/intel/graph.db"
via: "SQL queries for hotspots"
pattern: "SELECT.*FROM edges.*GROUP BY"
- from: "hooks/gsd-intel-session.js"
to: ".planning/intel/summary.md"
via: "fs.readFileSync on startup/resume"
pattern: "readFileSync.*summary\\.md"
---
<objective>
Generate rich summaries from SQLite graph instead of simple file counts.
Purpose: Provide Claude with actionable intelligence - dependency hotspots, file purposes, and relationship awareness at session start.
Output: Updated gsd-intel-index.js with graph-backed summary generation.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/04-semantic-intelligence/04-RESEARCH.md
@.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md
From 04-01: SQLite graph layer with nodes (entity metadata) and edges (wiki-links).
Summary generation requirements (from research):
- Query hotspots: most-depended-on files
- Group by type from node body
- Include file purposes from entity content
- Target < 500 tokens for context injection
Existing wiring (from Phase 2):
- hooks/gsd-intel-session.js reads summary.md on startup/resume
- Injects content as <codebase-intelligence> tag into Claude's context
- This wiring already exists - we verify it works with new graph format
</context>
<tasks>
<task type="auto">
<name>Task 1: Add graph query helpers</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add graph query functions after the existing graph helpers (loadGraphDatabase, persistDatabase):
```javascript
/**
* Get dependency hotspots from graph
* Returns top N files by number of dependents
*
* @param {object} db - sql.js database instance
* @param {number} limit - Max results (default 5)
* @returns {Array<{id: string, count: number, path: string, type: string}>}
*/
function getHotspots(db, limit = 5) {
const results = db.exec(`
SELECT
e.target as id,
COUNT(*) as count,
json_extract(n.body, '$.path') as path,
json_extract(n.body, '$.type') as type
FROM edges e
LEFT JOIN nodes n ON e.target = n.id
GROUP BY e.target
ORDER BY count DESC
LIMIT ?
`, [limit]);
if (!results[0]?.values) return [];
return results[0].values.map(([id, count, path, type]) => ({
id,
count,
path: path || id,
type: type || 'unknown'
}));
}
/**
* Get nodes grouped by type
* Returns type -> count mapping
*
* @param {object} db - sql.js database instance
* @returns {Array<{type: string, count: number}>}
*/
function getNodesByType(db) {
const results = db.exec(`
SELECT
json_extract(body, '$.type') as type,
COUNT(*) as count
FROM nodes
GROUP BY type
ORDER BY count DESC
`);
if (!results[0]?.values) return [];
return results[0].values.map(([type, count]) => ({
type: type || 'other',
count
}));
}
/**
* Get all dependents of a file (transitive)
* Uses recursive CTE for graph traversal
*
* @param {object} db - sql.js database instance
* @param {string} entityId - Starting entity
* @param {number} maxDepth - Max recursion depth (default 5)
* @returns {Array<{id: string, depth: number, path: string}>}
*/
function getDependents(db, entityId, maxDepth = 5) {
const results = db.exec(`
WITH RECURSIVE dependents(id, depth) AS (
SELECT ?, 0
UNION
SELECT e.source, d.depth + 1
FROM edges e
JOIN dependents d ON e.target = d.id
WHERE d.depth < ?
)
SELECT DISTINCT
d.id,
d.depth,
json_extract(n.body, '$.path') as path
FROM dependents d
LEFT JOIN nodes n ON d.id = n.id
WHERE d.id != ?
ORDER BY d.depth, d.id
`, [entityId.toLowerCase(), maxDepth, entityId.toLowerCase()]);
if (!results[0]?.values) return [];
return results[0].values.map(([id, depth, path]) => ({
id,
depth,
path: path || id
}));
}
```
Key design notes:
- LEFT JOIN on nodes allows edges to exist even if target node doesn't exist yet
- UNION (not UNION ALL) prevents infinite loops in cyclic graphs
- maxDepth limit prevents runaway queries
- All IDs lowercased for consistency
</action>
<verify>`grep -q "getHotspots" hooks/gsd-intel-index.js && grep -q "getDependents" hooks/gsd-intel-index.js`</verify>
<done>Graph query helpers exist: getHotspots, getNodesByType, getDependents</done>
</task>
<task type="auto">
<name>Task 2: Create graph-backed summary generator</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add generateGraphSummary() function that queries the graph database:
```javascript
/**
* Generate semantic summary from graph database
* Called when graph.db exists (Phase 4+)
* Falls back to entity-based summary if no graph
*
* Target: < 500 tokens for context injection
*/
async function generateGraphSummary() {
const intelDir = path.join(process.cwd(), '.planning', 'intel');
const dbPath = path.join(intelDir, 'graph.db');
const summaryPath = path.join(intelDir, 'summary.md');
const entitiesDir = path.join(intelDir, 'entities');
// Require graph.db to exist
if (!fs.existsSync(dbPath)) {
return null; // Caller should fall back to entity summary
}
try {
const { db } = await loadGraphDatabase();
const lines = [];
// Header
lines.push('# Codebase Intelligence');
lines.push('');
// File count from nodes
const countResult = db.exec('SELECT COUNT(*) FROM nodes');
const fileCount = countResult[0]?.values[0]?.[0] || 0;
lines.push(`**Indexed entities:** ${fileCount}`);
lines.push(`**Last updated:** ${new Date().toISOString().split('T')[0]}`);
lines.push('');
// Dependency hotspots (most impactful files)
const hotspots = getHotspots(db, 5);
if (hotspots.length > 0) {
lines.push('## Dependency Hotspots');
lines.push('');
lines.push('Files with most dependents (change carefully):');
for (const { path: filePath, count, type } of hotspots) {
const typeLabel = type !== 'unknown' ? ` [${type}]` : '';
lines.push(`1. \`${filePath}\` (${count} dependents)${typeLabel}`);
}
lines.push('');
}
// Group by type
const byType = getNodesByType(db);
if (byType.length > 0) {
lines.push('## Module Types');
lines.push('');
for (const { type, count } of byType) {
const label = type.charAt(0).toUpperCase() + type.slice(1);
lines.push(`- **${label}**: ${count} files`);
}
lines.push('');
}
// Edge count (relationship density)
const edgeResult = db.exec('SELECT COUNT(*) FROM edges');
const edgeCount = edgeResult[0]?.values[0]?.[0] || 0;
if (edgeCount > 0) {
lines.push(`**Relationships tracked:** ${edgeCount}`);
lines.push('');
}
db.close();
// Write summary
const summary = lines.join('\n');
fs.writeFileSync(summaryPath, summary);
return summary;
} catch (e) {
// Graph query failed, return null to fall back
return null;
}
}
```
Key design notes:
- Returns null if graph doesn't exist or query fails (allows fallback)
- Hotspots show files that cause most downstream impact
- Module types provide quick orientation
- Edge count indicates relationship density
- Targets < 500 tokens (no verbose lists)
</action>
<verify>`grep -q "generateGraphSummary" hooks/gsd-intel-index.js`</verify>
<done>generateGraphSummary() function queries graph and writes summary.md</done>
</task>
<task type="auto">
<name>Task 3: Integrate graph summary into regeneration flow and verify SessionStart wiring</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Modify regenerateEntitySummary() to prefer graph summary when available.
Find the existing regenerateEntitySummary() function and update it:
```javascript
/**
* Regenerate summary.md from all entity files
* Uses graph database if available (Phase 4+), falls back to file-based
*/
async function regenerateEntitySummary() {
const intelDir = path.join(process.cwd(), '.planning', 'intel');
const entitiesDir = path.join(intelDir, 'entities');
const summaryPath = path.join(intelDir, 'summary.md');
const dbPath = path.join(intelDir, 'graph.db');
// Check directories exist
if (!fs.existsSync(entitiesDir)) {
return;
}
// Try graph-based summary first (Phase 4+)
if (fs.existsSync(dbPath)) {
try {
const graphSummary = await generateGraphSummary();
if (graphSummary) {
return; // Graph summary written, done
}
} catch (e) {
// Fall through to file-based summary
}
}
// Fall back to existing file-based entity summary
// (Keep all existing regenerateEntitySummary logic here)
```
The key change: Check for graph.db first, try generateGraphSummary(), only fall back to existing logic if graph unavailable or fails.
Also update the stdin handler to use async regenerateEntitySummary:
Find:
```javascript
if (isEntityFile(filePath)) {
syncEntityToGraph(filePath).then(() => {
regenerateEntitySummary();
process.exit(0);
})
```
Change to:
```javascript
if (isEntityFile(filePath)) {
syncEntityToGraph(filePath).then(async () => {
await regenerateEntitySummary();
process.exit(0);
})
```
Note: regenerateEntitySummary becomes async because it calls generateGraphSummary.
</action>
<verify>
Test the full flow including SessionStart injection:
```bash
cd /Users/lexchristopherson/Developer/claude-code-resources/get-shit-done
# Create test entities with dependencies
mkdir -p .planning/intel/entities
echo '---
path: /test/db.ts
type: util
updated: 2026-01-20
status: active
---
# db.ts
## Purpose
Database client.
## Dependencies
None
## Used By
TBD
' > .planning/intel/entities/test-db.md
echo '---
path: /test/auth.ts
type: util
updated: 2026-01-20
status: active
---
# auth.ts
## Purpose
Auth utilities.
## Dependencies
- [[test-db]]
## Used By
TBD
' > .planning/intel/entities/test-auth.md
echo '---
path: /test/api.ts
type: api
updated: 2026-01-20
status: active
---
# api.ts
## Purpose
API routes.
## Dependencies
- [[test-db]]
- [[test-auth]]
## Used By
TBD
' > .planning/intel/entities/test-api.md
# Sync all to graph
for f in .planning/intel/entities/test-*.md; do
echo "{\"tool_name\":\"Write\",\"tool_input\":{\"file_path\":\"$f\"}}" | node hooks/gsd-intel-index.js
done
# Check summary.md has graph-based content
cat .planning/intel/summary.md
# Should show:
# - "Dependency Hotspots" section
# - test-db with 2 dependents (auth and api both depend on it)
# Verify SessionStart hook reads new summary format
echo '{"source":"startup"}' | node hooks/gsd-intel-session.js
# Should output <codebase-intelligence>...</codebase-intelligence> with hotspots
# Cleanup
rm .planning/intel/entities/test-*.md
rm .planning/intel/graph.db
rm .planning/intel/summary.md
```
</verify>
<done>Summary generation prefers graph when available, SessionStart hook confirmed to inject graph-backed summary</done>
</task>
</tasks>
<verification>
After all tasks complete:
1. Query helpers exist:
```bash
grep -q "getHotspots" hooks/gsd-intel-index.js && echo "PASS"
grep -q "getDependents" hooks/gsd-intel-index.js && echo "PASS"
```
2. Graph summary generator exists:
```bash
grep -q "generateGraphSummary" hooks/gsd-intel-index.js && echo "PASS"
```
3. Summary prefers graph:
- Create entities with [[wiki-links]]
- Simulate entity writes
- Check summary.md has "Dependency Hotspots" section
- Hotspot counts are accurate
4. SessionStart wiring verified:
```bash
# Verify SessionStart reads and injects the new format
echo '{"source":"startup"}' | node hooks/gsd-intel-session.js | grep -q "Dependency Hotspots" && echo "PASS: SessionStart injects graph summary"
```
</verification>
<success_criteria>
- [ ] getHotspots() queries top N most-depended files
- [ ] getNodesByType() groups entities by type
- [ ] getDependents() uses recursive CTE for transitive queries
- [ ] generateGraphSummary() produces < 500 token summary
- [ ] regenerateEntitySummary() prefers graph when graph.db exists
- [ ] Falls back gracefully to file-based summary
- [ ] Summary includes dependency hotspots with accurate counts
- [ ] SessionStart hook (gsd-intel-session.js) correctly injects graph-backed summary into context
</success_criteria>
<output>
After completion, create `.planning/phases/04-semantic-intelligence/04-02-SUMMARY.md`
</output>

View File

@@ -1,99 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 02
subsystem: intel
tags: [sqlite, sql-js, graph-query, recursive-cte, summary-generation]
requires:
- phase: 04-01
provides: SQLite graph layer with nodes/edges schema
provides:
- Graph query helpers (getHotspots, getNodesByType, getDependents)
- Graph-backed summary generation (generateGraphSummary)
- Transitive dependent queries via recursive CTE
affects: [04-03, context-injection, session-start]
tech-stack:
added: []
patterns:
- Recursive CTE for transitive graph traversal
- Graph-backed summary with hotspot analysis
key-files:
created: []
modified:
- hooks/gsd-intel-index.js
key-decisions:
- "LEFT JOIN allows edges to exist before target nodes indexed"
- "UNION (not UNION ALL) prevents infinite loops in cyclic graphs"
- "maxDepth limit on recursive CTE prevents runaway queries"
patterns-established:
- "Graph queries return structured arrays with id/path/type"
- "Summary prefers graph when available, falls back to entity-file-based"
duration: 2min
completed: 2026-01-20
---
# Phase 4 Plan 2: Query Interface Summary
**Graph-backed summary generation with dependency hotspots, type grouping, and recursive CTE for transitive dependents**
## Performance
- **Duration:** 2 min
- **Started:** 2026-01-20T15:56:42Z
- **Completed:** 2026-01-20T15:59:07Z
- **Tasks:** 3
- **Files modified:** 1
## Accomplishments
- Added graph query helpers (getHotspots, getNodesByType, getDependents)
- Created generateGraphSummary() that queries SQLite for rich semantic summaries
- Integrated graph summary into regeneration flow with entity-file fallback
- Verified SessionStart hook correctly injects graph-backed summary into context
## Task Commits
Each task was committed atomically:
1. **Task 1: Add graph query helpers** - `5824196` (feat)
2. **Task 2: Create graph-backed summary generator** - `3d8cf70` (feat)
3. **Task 3: Integrate graph summary into regeneration flow** - `101bc58` (feat)
## Files Created/Modified
- `hooks/gsd-intel-index.js` - Added graph query helpers and summary generator
## Decisions Made
- **LEFT JOIN on nodes**: Allows edges to exist even if target node not yet indexed (forward references)
- **UNION vs UNION ALL**: Using UNION in recursive CTE prevents infinite loops in cyclic dependency graphs
- **maxDepth default of 5**: Prevents runaway queries on deeply nested dependencies
- **All entity IDs lowercased**: Ensures consistent matching across queries
## Deviations from Plan
None - plan executed exactly as written.
## Issues Encountered
None
## User Setup Required
None - no external service configuration required.
## Next Phase Readiness
- Graph query interface complete
- Summary generation produces < 500 token output (47 words in test)
- SessionStart hook verified working with new graph-backed format
- Ready for 04-03: Entity generation instructions integration
---
*Phase: 04-semantic-intelligence*
*Completed: 2026-01-20*

View File

@@ -1,376 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 03
type: execute
wave: 2
depends_on: [04-01]
files_modified:
- commands/gsd/analyze-codebase.md
- package.json
autonomous: true
must_haves:
truths:
- "Claude creates entity files with semantic understanding via /gsd:analyze-codebase"
- "Entity files include purpose, not just syntax"
- "Batch processing handles 100+ files efficiently"
artifacts:
- path: "commands/gsd/analyze-codebase.md"
provides: "Semantic entity generation instructions for Claude"
contains: "semantic entities"
- path: "package.json"
provides: "Anthropic SDK dependency"
contains: "@anthropic-ai/sdk"
key_links:
- from: "commands/gsd/analyze-codebase.md"
to: "Task tool"
via: "Subagent spawning for entity batch processing"
pattern: "Task.*entity"
---
<objective>
Enhance /gsd:analyze-codebase to create semantic entity files using Claude.
Purpose: Generate entity documentation that captures file PURPOSE (what it does, why it exists), not just syntax (exports/imports). This transforms "2-3 ls commands" of information into genuine semantic understanding.
Output: Updated analyze-codebase.md command with entity generation instructions, @anthropic-ai/sdk dependency (for future direct API use).
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/04-semantic-intelligence/04-RESEARCH.md
@.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md
From research:
- Use @anthropic-ai/sdk for Claude API calls
- claude-sonnet-4-5-20250929 for entity generation (fast, cost-effective)
- Process files in batches to avoid rate limits
- Entity template format already exists
Current analyze-codebase.md:
- Steps 1-8 for bulk codebase scanning
- Creates index.json, conventions.json, summary.md
- Does NOT create entity files
New requirement:
- After indexing, optionally create entity .md files
- Claude (executing the command) reads file content and generates semantic documentation
- No embedded JavaScript in command markdown - Claude IS the executor
Execution model clarification:
- GSD command .md files contain INSTRUCTIONS for Claude to follow
- Claude reads the markdown and executes the instructions using its tools
- Commands cannot contain executable JavaScript - Claude interprets and acts on the instructions
- For batch processing, Claude uses the Task tool to spawn subagents
</context>
<tasks>
<task type="auto">
<name>Task 1: Add Anthropic SDK dependency</name>
<files>package.json</files>
<action>
Add @anthropic-ai/sdk to package.json dependencies:
```json
"dependencies": {
"sql.js": "^1.12.0",
"@anthropic-ai/sdk": "^0.52.0"
}
```
Note: Version 0.52.0+ includes Messages API with proper TypeScript support. This dependency enables future direct API integration (e.g., hooks that call Claude API directly). For the /gsd:analyze-codebase command, Claude itself generates the entity content.
</action>
<verify>`grep -q "@anthropic-ai/sdk" package.json`</verify>
<done>@anthropic-ai/sdk added to package.json</done>
</task>
<task type="auto">
<name>Task 2: Add semantic entity generation to analyze-codebase</name>
<files>commands/gsd/analyze-codebase.md</files>
<action>
Update the analyze-codebase.md command to add entity generation after index creation.
1. Update the objective to mention entity generation:
```markdown
<objective>
Scan codebase to populate .planning/intel/ with file index, conventions, and semantic entity documentation.
Works standalone (without /gsd:new-project) for brownfield codebases. Creates:
- index.json for file index
- conventions.json for naming patterns
- summary.md for context injection
- entities/*.md for semantic file documentation (optional)
Output: .planning/intel/index.json, conventions.json, summary.md, entities/*.md
</objective>
```
2. Add Task to allowed-tools (for entity generation via subagent):
```yaml
allowed-tools:
- Read
- Bash
- Glob
- Write
- Task
```
3. Add new Step 9 after Step 8 (before completion report). This step provides INSTRUCTIONS for Claude to follow:
```markdown
## Step 9: Generate semantic entities (optional)
After indexing, generate semantic entity files for key codebase files.
### 9a: Select key files for entity generation
From the index, select files for entity generation using these criteria:
- Files with 3+ exports (significant modules)
- Files imported by 5+ other files (dependency hotspots)
- Files in key directories: api/, lib/, utils/, services/, models/
- Limit to 50 files maximum per run (context management)
Skip:
- Test files (*.test.*, *.spec.*)
- Generated files (*.generated.*, *.d.ts)
- Config files (*.config.*)
- Files already with entities in .planning/intel/entities/
### 9b: Create entity directory
```bash
mkdir -p .planning/intel/entities
```
### 9c: Generate entities in batches
For efficient processing, use the Task tool to spawn a subagent for batch entity generation.
**Subagent prompt template:**
```
Generate semantic entity documentation for the following files.
For each file:
1. Read the file content
2. Analyze its purpose, exports, and dependencies
3. Write an entity file to .planning/intel/entities/{slug}.md
Entity template (use EXACTLY this format):
---
path: {file_path}
type: [module|component|util|config|api|hook|service|model]
updated: {today's date YYYY-MM-DD}
status: active
---
# {filename}
## Purpose
[1-3 sentences: What does this file do? Why does it exist? What problem does it solve?]
## Exports
[List each export with signature and brief description]
- `exportName(args): ReturnType` - What it does
## Dependencies
[Internal deps use wiki-links, external use plain text]
- [[slugified-internal-path]] - Why needed
- external-package - Why needed
## Used By
TBD
## Notes
[Optional: patterns, gotchas, or important context]
---
Slug convention: `src/lib/db.ts` -> `src-lib-db` (replace / and . with -, remove extension)
Files to process:
{list of file paths, max 10 per batch}
```
### 9d: Process in batches of 10
For codebases with many key files:
1. Split the selected files into batches of 10
2. Use Task tool for each batch with the prompt template above
3. Wait for each batch to complete before starting the next
4. This prevents context exhaustion and allows progress tracking
### 9e: Verify entity creation
After batch processing:
- Count entities created: `ls .planning/intel/entities/*.md | wc -l`
- Verify they have semantic content (Purpose section, not just syntax)
- The PostToolUse hook will automatically sync new entities to graph.db
```
4. Update Step 10 (renumber from Step 8) to include entity stats:
```markdown
## Step 10: Report completion
Display summary statistics:
```
Codebase Analysis Complete
Files indexed: [N]
Exports found: [N]
Imports found: [N]
Conventions detected:
- Naming: [dominant case] ([percentage]%)
- Directories: [list]
- Patterns: [list]
Entities created: [N] (if entity generation ran)
- Skipped: [N] (already existed or filtered)
Files created:
- .planning/intel/index.json
- .planning/intel/conventions.json
- .planning/intel/summary.md
- .planning/intel/entities/*.md (if entities generated)
Next: Intel hooks will continue incremental learning as you code.
```
```
5. Update success criteria to include entity generation:
```markdown
<success_criteria>
- [ ] .planning/intel/ directory created
- [ ] All JS/TS files scanned (excluding node_modules, dist, build, .git, vendor, coverage)
- [ ] index.json populated with exports and imports for each file
- [ ] conventions.json has detected patterns (naming, directories, suffixes)
- [ ] summary.md is concise (< 500 tokens)
- [ ] entities/*.md created for key files (if Step 9 executed)
- [ ] Entity files have semantic Purpose sections (not just syntax extraction)
- [ ] Statistics reported to user
</success_criteria>
```
</action>
<verify>
```bash
# Check command has entity generation step
grep -q "Generate semantic entities" commands/gsd/analyze-codebase.md && echo "PASS: Step 9 exists"
# Check mentions Task tool for batching
grep -q "Task tool" commands/gsd/analyze-codebase.md && echo "PASS: Task batching documented"
# Check entity template is present
grep -q "## Purpose" commands/gsd/analyze-codebase.md && echo "PASS: Entity template included"
# Check batch size documented
grep -q "batches of 10" commands/gsd/analyze-codebase.md && echo "PASS: Batch processing"
```
</verify>
<done>analyze-codebase.md includes semantic entity generation via Claude + Task tool batching</done>
</task>
<task type="auto">
<name>Task 3: Add context section explaining execution model</name>
<files>commands/gsd/analyze-codebase.md</files>
<action>
Update the context section to explain how entity generation works.
In the context section, add:
```markdown
**Entity generation:**
Step 9 generates semantic entity files using Claude's understanding of code purpose. Unlike regex-based extraction (exports/imports), this captures WHY code exists.
For large codebases (50+ key files), entity generation uses the Task tool to spawn subagents that process files in batches of 10. This:
- Prevents context exhaustion
- Allows progress tracking
- Enables parallel processing
Entity files are written to `.planning/intel/entities/` and automatically synced to the graph database by the PostToolUse hook.
**When to skip entity generation:**
- Quick index-only scan: Stop after Step 8
- Already have entities: Existing entities won't be overwritten
- Small codebase: May not need formal entities
```
This clarifies:
1. Claude generates entity content (not embedded JavaScript)
2. Task tool handles batching for large codebases
3. Users can skip Step 9 if they just want the index
</action>
<verify>
```bash
grep -q "Task tool to spawn subagents" commands/gsd/analyze-codebase.md && echo "PASS: Execution model explained"
```
</verify>
<done>Command explains entity generation execution model clearly</done>
</task>
</tasks>
<verification>
After all tasks complete:
1. SDK dependency added:
```bash
grep -q "@anthropic-ai/sdk" package.json && echo "PASS"
```
2. Command has entity generation:
```bash
grep -q "Step 9" commands/gsd/analyze-codebase.md && echo "PASS"
grep -q "semantic entities" commands/gsd/analyze-codebase.md && echo "PASS"
```
3. Task tool batching:
```bash
grep -q "Task tool" commands/gsd/analyze-codebase.md && echo "PASS"
grep -q "batches of 10" commands/gsd/analyze-codebase.md && echo "PASS"
```
4. Entity template present:
```bash
grep -q "## Purpose" commands/gsd/analyze-codebase.md && echo "PASS"
```
Manual test:
```bash
# Run /gsd:analyze-codebase on a test project
# Verify Steps 1-8 produce index.json, conventions.json, summary.md
# If Step 9 runs, verify .planning/intel/entities/*.md created with semantic content
```
</verification>
<success_criteria>
- [ ] @anthropic-ai/sdk added to package.json
- [ ] Step 9 added for entity generation (instructions for Claude, not embedded JS)
- [ ] Task tool documented for batch processing subagents
- [ ] File selection criteria documented (3+ exports, 5+ dependents, key dirs)
- [ ] 50 file limit per run to manage context
- [ ] Batch processing with batches of 10 via Task tool
- [ ] Entity slug convention documented
- [ ] Execution model explained (Claude generates content, not script execution)
- [ ] Updated success criteria includes entity generation
</success_criteria>
<output>
After completion, create `.planning/phases/04-semantic-intelligence/04-03-SUMMARY.md`
</output>

View File

@@ -1,101 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 03
subsystem: intel
tags: [entity-generation, semantic-analysis, task-batching, anthropic-sdk]
# Dependency graph
requires:
- phase: 04-01
provides: SQLite graph layer for relationship storage
provides:
- Semantic entity generation instructions in analyze-codebase command
- Task tool batching pattern for 100+ file processing
- Entity template with Purpose-focused documentation
affects: [04-02, future-intel-queries]
# Tech tracking
tech-stack:
added: ["@anthropic-ai/sdk ^0.52.0"]
patterns: [Task tool batching, entity slug convention]
key-files:
created: []
modified:
- commands/gsd/analyze-codebase.md
- package.json
key-decisions:
- "50 file limit per run to manage context"
- "Batches of 10 files via Task tool for parallelization"
- "Entity slug convention: path--segments--filename-ext.md"
- "Selection criteria: 3+ exports, 5+ dependents, key directories"
patterns-established:
- "Task tool batching: spawn subagents for batch processing large file sets"
- "Entity template: Purpose > Exports > Dependencies > Used By"
# Metrics
duration: 2min
completed: 2026-01-20
---
# Phase 4 Plan 3: Entity Generation Instructions Summary
**Semantic entity generation via Task tool batching in analyze-codebase command**
## Performance
- **Duration:** 2 min
- **Started:** 2026-01-20T15:56:34Z
- **Completed:** 2026-01-20T15:58:41Z
- **Tasks:** 3
- **Files modified:** 2
## Accomplishments
- Added @anthropic-ai/sdk dependency for future API-based entity generation
- Created Step 9 in analyze-codebase for semantic entity file generation
- Documented Task tool batching pattern for 100+ file codebases
- Established entity template with Purpose-focused semantic documentation
## Task Commits
Each task was committed atomically:
1. **Task 1: Add Anthropic SDK dependency** - `8d33ae7` (chore)
2. **Task 2 & 3: Add entity generation + execution model** - `b3db2ff` (feat)
## Files Created/Modified
- `package.json` - Added @anthropic-ai/sdk ^0.52.0 dependency
- `commands/gsd/analyze-codebase.md` - Added Step 9 for entity generation, Task tool in allowed-tools, execution model context
## Decisions Made
- **50 file limit per run:** Prevents context window exhaustion during entity generation
- **Batches of 10:** Balances parallelization with Task tool overhead
- **Entity slug convention (path--segments--filename-ext.md):** Flat directory structure with reversible file identification
- **Selection criteria priority:** High-export files first, then hub files, then key directories
## Deviations from Plan
None - plan executed exactly as written.
## Issues Encountered
None.
## User Setup Required
None - no external service configuration required.
## Next Phase Readiness
- Entity generation instructions complete in analyze-codebase command
- Ready for 04-02 (Query Interface) to query entity relationships
- Graph layer from 04-01 available for entity relationship storage
---
*Phase: 04-semantic-intelligence*
*Completed: 2026-01-20*

View File

@@ -1,250 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 04
type: execute
wave: 1
depends_on: []
files_modified:
- hooks/gsd-intel-index.js
autonomous: true
gap_closure: true
must_haves:
truths:
- "Claude can query 'what uses src/lib/db.ts?' via CLI"
- "getDependents() is callable through stdin query action"
- "Query results return as JSON to stdout"
artifacts:
- path: "hooks/gsd-intel-index.js"
provides: "CLI query interface for graph database"
contains: "action.*query"
key_links:
- from: "stdin handler"
to: "getDependents()"
via: "query action routing"
pattern: "action.*query.*getDependents"
---
<objective>
Expose the orphaned getDependents() function via CLI query interface so Claude can answer "what uses this file?" questions during sessions.
Purpose: Close the verification gap - INTEL-05 is blocked because getDependents() exists but has no interface. The infrastructure is complete; this adds the "last mile" wiring.
Output: Modified gsd-intel-index.js that accepts query actions via stdin and returns JSON results to stdout.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/STATE.md
# Gap source
@.planning/phases/04-semantic-intelligence/04-VERIFICATION.md
# Target file (contains getDependents at line 125)
@hooks/gsd-intel-index.js
</context>
<tasks>
<task type="auto">
<name>Task 1: Add query action routing to stdin handler</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Modify the stdin 'end' handler (starting at line 930) to detect and route query actions before the existing Write/Edit handling.
Add this routing logic after `const data = JSON.parse(input);` (line 932):
```javascript
// Handle query actions (graph queries)
if (data.action === 'query') {
handleQuery(data).then(result => {
console.log(JSON.stringify(result));
process.exit(0);
}).catch(err => {
console.log(JSON.stringify({ error: err.message }));
process.exit(1);
});
return; // Don't fall through to Write/Edit handling
}
```
Then add the `handleQuery` async function before the stdin handler (around line 925):
```javascript
/**
* Handle CLI query actions
* Routes to appropriate graph query function based on query type
*
* @param {Object} data - Query action data
* @param {string} data.action - Must be 'query'
* @param {string} data.type - Query type: 'dependents' | 'hotspots'
* @param {string} [data.target] - Entity ID for dependents query (e.g., 'src-lib-db')
* @param {number} [data.limit] - Max results (default: 10 for dependents, 5 for hotspots)
* @param {number} [data.maxDepth] - Max traversal depth for dependents (default: 5)
* @returns {Promise<Object>} Query results
*/
async function handleQuery(data) {
const { db, dbPath } = await loadGraphDatabase();
try {
switch (data.type) {
case 'dependents': {
if (!data.target) {
return { error: 'target is required for dependents query' };
}
const results = getDependents(db, data.target, data.maxDepth || 5);
const limited = data.limit ? results.slice(0, data.limit) : results.slice(0, 10);
return {
query: 'dependents',
target: data.target,
count: results.length,
results: limited
};
}
case 'hotspots': {
const results = getHotspots(db, data.limit || 5);
return {
query: 'hotspots',
count: results.length,
results
};
}
default:
return { error: `Unknown query type: ${data.type}. Valid types: dependents, hotspots` };
}
} finally {
db.close();
}
}
```
Key implementation notes:
- Query mode does NOT persist to disk (read-only queries)
- db.close() in finally block ensures cleanup
- Default limit of 10 prevents huge output for files with many dependents
- Output goes to stdout as JSON for Claude to parse
</action>
<verify>
Test query interface with heredoc:
```bash
# Test dependents query (should return empty or results depending on graph state)
node hooks/gsd-intel-index.js <<'EOF'
{"action":"query","type":"dependents","target":"src-lib-db"}
EOF
# Test hotspots query
node hooks/gsd-intel-index.js <<'EOF'
{"action":"query","type":"hotspots","limit":3}
EOF
# Test error handling (missing target)
node hooks/gsd-intel-index.js <<'EOF'
{"action":"query","type":"dependents"}
EOF
# Test unknown query type
node hooks/gsd-intel-index.js <<'EOF'
{"action":"query","type":"invalid"}
EOF
```
All commands should output valid JSON to stdout.
</verify>
<done>
Query actions route to handleQuery(), which returns JSON results via stdout. Claude can invoke `echo '{"action":"query",...}' | node hooks/gsd-intel-index.js` to query the graph.
</done>
</task>
<task type="auto">
<name>Task 2: Add usage documentation as code comment</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add a documentation comment block at the top of the file (after line 4) explaining the CLI query interface:
```javascript
/**
* CLI Query Interface
*
* In addition to PostToolUse indexing (Write/Edit actions), this hook supports
* direct graph queries via stdin:
*
* Query dependents (what uses this file?):
* echo '{"action":"query","type":"dependents","target":"src-lib-db"}' | node hooks/gsd-intel-index.js
*
* Query hotspots (most depended-on files):
* echo '{"action":"query","type":"hotspots","limit":5}' | node hooks/gsd-intel-index.js
*
* Options:
* - target: Entity ID (required for dependents, e.g., 'src-lib-db' for src/lib/db.ts)
* - limit: Max results (default: 10 for dependents, 5 for hotspots)
* - maxDepth: Traversal depth for dependents (default: 5)
*
* Output: JSON to stdout with query results
*/
```
This makes the CLI interface discoverable for future Claude sessions.
</action>
<verify>
Read first 30 lines of file to confirm documentation is present:
```bash
head -30 hooks/gsd-intel-index.js | grep -q "CLI Query Interface" && echo "Documentation added"
```
</verify>
<done>
CLI query interface is documented in the file header for discoverability.
</done>
</task>
</tasks>
<verification>
After both tasks complete, verify the gap is closed:
1. **Query interface works:**
```bash
# Create test graph.db if needed (empty is fine)
mkdir -p .planning/intel
node hooks/gsd-intel-index.js <<'EOF'
{"action":"query","type":"hotspots"}
EOF
# Should output: {"query":"hotspots","count":0,"results":[]}
```
2. **getDependents is no longer orphaned:**
```bash
grep -n "getDependents" hooks/gsd-intel-index.js
# Should show: definition (line ~125) AND call in handleQuery
```
3. **Error handling works:**
```bash
node hooks/gsd-intel-index.js <<'EOF'
{"action":"query","type":"dependents"}
EOF
# Should output: {"error":"target is required for dependents query"}
```
</verification>
<success_criteria>
- [ ] `handleQuery()` function added and routes query actions
- [ ] getDependents() called from handleQuery() (no longer orphaned)
- [ ] Query results output as JSON to stdout
- [ ] CLI documentation added to file header
- [ ] Error handling returns JSON error objects
- [ ] INTEL-05 requirement unblocked: Claude can query "what uses this file?"
</success_criteria>
<output>
After completion, create `.planning/phases/04-semantic-intelligence/04-04-SUMMARY.md`
</output>

View File

@@ -1,101 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 04
subsystem: codebase-intelligence
tags: [sql.js, graph-queries, cli-interface, dependency-analysis]
# Dependency graph
requires:
- phase: 04-01
provides: SQLite graph database with nodes and edges tables
- phase: 04-02
provides: getDependents() and getHotspots() query functions
provides:
- CLI query interface via stdin for graph database queries
- handleQuery() routing function for dependents and hotspots queries
- JSON output to stdout for Claude consumption
affects: [codebase-analysis, dependency-tracking, refactoring-safety]
# Tech tracking
tech-stack:
added: []
patterns: [stdin-json-routing, query-action-pattern]
key-files:
created: []
modified: [hooks/gsd-intel-index.js]
key-decisions:
- "Query mode does not persist to disk (read-only operations)"
- "Default limit of 10 for dependents prevents huge output"
- "Query action routing occurs before Write/Edit handling in stdin handler"
patterns-established:
- "Query action pattern: {action: 'query', type: 'dependents'|'hotspots', ...options}"
- "JSON error objects returned to stdout for graceful error handling"
# Metrics
duration: 4min
completed: 2026-01-20
---
# Phase 04 Plan 04: CLI Query Interface Summary
**CLI query interface exposes getDependents() and getHotspots() via stdin, enabling Claude to answer "what uses this file?" questions**
## Performance
- **Duration:** 4 min
- **Started:** 2026-01-20T16:10:00Z
- **Completed:** 2026-01-20T16:14:00Z
- **Tasks:** 2
- **Files modified:** 1
## Accomplishments
- getDependents() function now accessible via CLI (no longer orphaned)
- Query interface accepts JSON actions via stdin and returns results to stdout
- Support for both dependents queries (transitive "what uses this?") and hotspots queries (most depended-on files)
- Complete error handling for missing parameters and unknown query types
## Task Commits
Each task was committed atomically:
1. **Task 1: Add query action routing to stdin handler** - `f46327a` (feat)
2. **Task 2: Add usage documentation as code comment** - `791265a` (docs)
## Files Created/Modified
- `hooks/gsd-intel-index.js` - Added handleQuery() function and stdin routing for query actions, plus CLI usage documentation
## Decisions Made
**Query mode read-only:** Query actions do not persist to disk. The database is opened, queried, and closed without writes. This keeps query operations safe and lightweight.
**Default limits:** Dependents queries default to 10 results, hotspots to 5. This prevents overwhelming output when a file has many dependents. Clients can override via `limit` parameter.
**Routing priority:** Query actions are handled before Write/Edit tool processing in the stdin handler. This ensures clean separation between query mode and indexing mode.
## Deviations from Plan
None - plan executed exactly as written.
## Issues Encountered
None
## User Setup Required
None - no external service configuration required.
## Next Phase Readiness
**Gap closure complete.** INTEL-05 verification is now unblocked:
- getDependents() is callable via CLI query interface
- Graph database queries work end-to-end (stdin → query → stdout)
- Claude can query dependency information during sessions
The semantic intelligence system is complete and ready for production use.
---
*Phase: 04-semantic-intelligence*
*Completed: 2026-01-20*

View File

@@ -1,101 +0,0 @@
---
phase: 04-semantic-intelligence
plan: 05
subsystem: codebase-intelligence
tags: [planner-integration, context-injection, dependency-awareness]
# Dependency graph
requires:
- phase: 04-02
provides: Query functions and summary.md generation
provides:
- Intel injection into planner prompt via {intel_content} placeholder
- Planner receives dependency hotspots and module types at planning time
affects: [planning, context-assembly, plan-quality]
# Tech tracking
tech-stack:
added: []
patterns: [context-file-injection, graceful-degradation]
key-files:
created: []
modified: [commands/gsd/plan-phase.md]
key-decisions:
- "Read intel at Step 7 alongside other context files"
- "Use 2>/dev/null for graceful degradation when summary.md doesn't exist"
- "Add intel section after gap closure section in planning_context"
patterns-established:
- "Context file injection: read into variable in Step 7, inject via placeholder in Step 8"
- "Optional context: empty string if file doesn't exist, planner handles gracefully"
# Metrics
duration: 1min
completed: 2026-01-20
---
# Phase 04 Plan 05: Intel Injection into Planner Summary
**Wired plan-phase.md to read and inject .planning/intel/summary.md into planner prompt, giving planners awareness of dependency hotspots and module composition**
## Performance
- **Duration:** 1 min
- **Started:** 2026-01-20T16:23:37Z
- **Completed:** 2026-01-20T16:24:14Z
- **Tasks:** 2
- **Files modified:** 1
## Accomplishments
- Planner now receives codebase intelligence (dependency hotspots, module types, relationship counts)
- Phase 4 infrastructure is fully connected: index → graph → summary → planner
- Graceful degradation: planners work fine when intel doesn't exist yet
## Task Commits
Each task was committed atomically:
1. **Task 1 & 2: Add intel read and injection** - `61d1e91` (feat)
## Files Created/Modified
- `commands/gsd/plan-phase.md` - Added INTEL_CONTENT read in Step 7, {intel_content} placeholder in Step 8 planning_context
## Decisions Made
**Read location:** Added intel read in Step 7 alongside other context files (STATE, ROADMAP, RESEARCH, etc.) rather than a new step. Keeps all context file reads in one place.
**Template location:** Added intel section after gap closure section and before `</planning_context>`. This puts codebase-level context after phase-specific context.
**Empty string fallback:** Used `2>/dev/null` so INTEL_CONTENT is empty string when summary.md doesn't exist. Planner template handles this gracefully.
## Deviations from Plan
None - plan executed exactly as written.
## Issues Encountered
None
## User Setup Required
None - no external service configuration required.
## Next Phase Readiness
**Last mile wiring complete.** The codebase intelligence system is now fully operational:
- `gsd-intel-index` hook captures file changes and builds graph database
- Summary generation creates human-readable intel from graph queries
- Planner receives intel automatically when planning phases
When a project has been indexed via `/gsd:scan-codebase` or through hook triggers, the planner will see:
- Dependency hotspots (files with most dependents - change carefully)
- Module type breakdown (utils, APIs, components, etc.)
- Total relationship count
This closes the Phase 4 objective: Claude understands codebase structure and conventions before it starts working.
---
*Phase: 04-semantic-intelligence*
*Completed: 2026-01-20*

View File

@@ -1,164 +0,0 @@
---
phase: 05-subagent-codebase-analysis
plan: 01
type: execute
wave: 1
depends_on: []
files_modified:
- agents/gsd-entity-generator.md
autonomous: true
must_haves:
truths:
- "Entity generator agent definition exists following gsd-codebase-mapper pattern"
- "Agent reads files, generates semantic entities, writes directly to disk"
- "Agent returns statistics only (not entity contents)"
artifacts:
- path: "agents/gsd-entity-generator.md"
provides: "Subagent definition for semantic entity generation"
contains: "gsd-entity-generator"
key_links:
- from: "agents/gsd-entity-generator.md"
to: ".planning/intel/entities/"
via: "Write tool calls"
pattern: "Write.*entities"
---
<objective>
Create gsd-entity-generator subagent definition.
Purpose: Define the subagent that generates semantic entity documentation for codebase files. This agent is spawned by `/gsd:analyze-codebase` with a list of file paths, reads each file, creates entity markdown, and writes directly to `.planning/intel/entities/`.
Output: `agents/gsd-entity-generator.md`
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/STATE.md
@.planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md
@agents/gsd-codebase-mapper.md
</context>
<tasks>
<task type="auto">
<name>Task 1: Create gsd-entity-generator agent definition</name>
<files>agents/gsd-entity-generator.md</files>
<action>
Create agent definition following gsd-codebase-mapper.md structure:
**Frontmatter:**
```yaml
---
name: gsd-entity-generator
description: Generates semantic entity documentation for codebase files. Spawned by analyze-codebase with file list. Writes entities directly to disk.
tools: Read, Write, Bash
color: cyan
---
```
**Role section:**
- Spawned by `/gsd:analyze-codebase` with file paths
- Reads source files, analyzes purpose/exports/dependencies
- Writes entity markdown to `.planning/intel/entities/{slug}.md`
- Returns statistics only (NOT entity contents)
**Process sections:**
1. `parse_file_list` - Extract file paths from prompt
- Expect: total count, output directory, slug convention, template, file list
2. `process_each_file` - For each file path:
- Read file using Read tool
- Analyze: purpose (why exists), exports (signatures), dependencies (internal [[wiki-links]], external plain text), module type
- Generate slug: `src/lib/db.ts` -> `src-lib-db` (replace / and . with -, remove extension, lowercase)
- Write entity to `.planning/intel/entities/{slug}.md`
- Track statistics (created, skipped, errors)
3. `return_statistics` - Return ONLY:
```
## ENTITY GENERATION COMPLETE
**Files processed:** {N}
**Entities created:** {M}
**Already existed:** {K}
**Errors:** {E}
Entities written to: .planning/intel/entities/
```
**Entity template section:**
Include the full entity template from 05-RESEARCH.md (frontmatter with path/type/updated/status, Purpose, Exports, Dependencies with [[wiki-links]], Used By = TBD, Notes optional).
**Type heuristics table:**
| Type | Indicators |
|------|-----------|
| api | api/, routes/, endpoints/, route handlers |
| component | components/, React/Vue exports |
| util | utils/, lib/, helpers/ |
| config | config/, *.config.* |
| hook | hooks/, use* functions |
| service | services/ |
| model | models/, types/ |
| test | *.test.*, *.spec.* |
| module | default |
**Wiki-link rules:**
- Internal (starts with `.` or `@/`): convert to slug, wrap in [[brackets]]
- External (package name): plain text, no brackets
**Critical rules:**
- WRITE entities directly (never return contents)
- PostToolUse hook syncs to graph.db automatically
- Use EXACT template format (hook parses frontmatter + [[links]])
**Success criteria checklist** (from research):
- All file paths processed
- Each entity written to correct path
- Frontmatter is valid YAML
- Purpose section is substantive (not "exports X")
- Internal deps use [[wiki-links]]
- Statistics returned (not entity contents)
</action>
<verify>
File exists and contains:
- Frontmatter with name, description, tools, color
- Role section explaining spawn context
- Process steps (parse, process, return)
- Entity template
- Type heuristics
- Wiki-link rules
- Critical rules matching gsd-codebase-mapper pattern
</verify>
<done>
`agents/gsd-entity-generator.md` exists with complete agent definition following gsd-codebase-mapper pattern. Agent is ready to be spawned by analyze-codebase.
</done>
</task>
</tasks>
<verification>
- [ ] File exists at `agents/gsd-entity-generator.md`
- [ ] Frontmatter valid YAML
- [ ] Role section explains subagent purpose
- [ ] Process has 3 steps: parse, process, return
- [ ] Entity template included in full
- [ ] Type heuristics table present
- [ ] Wiki-link rules specified
- [ ] Critical rules section matches gsd-codebase-mapper style
- [ ] Returns statistics only (not entity contents)
</verification>
<success_criteria>
gsd-entity-generator agent definition complete. Agent can be spawned with file paths and will generate semantic entities, writing directly to disk.
</success_criteria>
<output>
After completion, create `.planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md`
</output>

View File

@@ -1,122 +0,0 @@
---
phase: 05-subagent-codebase-analysis
plan: 01
title: gsd-entity-generator Agent Definition
subsystem: agents
tags: [subagent, entity-generation, semantic-analysis]
dependency-graph:
requires:
- 04-03 (entity generation instructions)
- gsd-codebase-mapper.md (pattern reference)
provides:
- gsd-entity-generator subagent definition
- Entity template specification
- Type heuristics table
affects:
- 05-02 (analyze-codebase integration)
- Future entity generation workflows
tech-stack:
added: []
patterns:
- Subagent direct-write pattern
- Statistics-only return pattern
- Wiki-link dependency notation
key-files:
created:
- agents/gsd-entity-generator.md
modified: []
decisions:
- id: skip-existing-entities
choice: Check for existing entity file before writing
rationale: Prevents overwriting manual edits to entities
metrics:
duration: 1 min 13 sec
completed: 2026-01-20
---
# Phase 05 Plan 01: gsd-entity-generator Agent Definition Summary
**One-liner:** Subagent definition for semantic entity generation with direct disk writes and statistics-only returns.
## What Was Built
Created `agents/gsd-entity-generator.md` following the established `gsd-codebase-mapper.md` pattern:
**Agent structure:**
- Frontmatter with name, description, tools (Read, Write, Bash), color
- Role section explaining spawn context from `/gsd:analyze-codebase`
- `<why_this_matters>` section explaining how entities feed the intelligence system
- 3-step process: parse_file_list, process_each_file, return_statistics
**Entity template specification:**
- YAML frontmatter: path, type, updated, status
- Sections: Purpose, Exports, Dependencies, Used By, Notes
- Internal dependencies use [[wiki-links]] for graph edge creation
- External dependencies as plain text
**Type heuristics table:**
| Type | Indicators |
|------|-----------|
| api | api/, routes/, endpoints/ |
| component | components/, React/Vue exports |
| util | utils/, lib/, helpers/ |
| config | config/, *.config.* |
| hook | hooks/, use* functions |
| service | services/ |
| model | models/, types/ |
| test | *.test.*, *.spec.* |
| module | default |
**Wiki-link rules:**
- Internal (starts with `.` or `@/`) -> convert to slug, wrap in [[brackets]]
- External (package name) -> plain text, no brackets
**Critical rules matching gsd-codebase-mapper:**
- WRITE entities directly (never return contents)
- PostToolUse hook syncs to graph.db automatically
- Use EXACT template format (hook parses frontmatter + [[links]])
- Skip existing entities (don't overwrite)
- Return statistics only (~10 lines)
## Tasks Completed
| Task | Name | Commit | Key Files |
|------|------|--------|-----------|
| 1 | Create gsd-entity-generator agent definition | f4c5817 | agents/gsd-entity-generator.md |
## Deviations from Plan
None - plan executed exactly as written.
## Decisions Made
1. **Skip existing entities by default**
- Prevents accidental overwrite of manually edited entities
- Aligns with research recommendation in 05-RESEARCH.md
- Agent checks `ls .planning/intel/entities/{slug}.md` before writing
## Verification Results
- [x] File exists at `agents/gsd-entity-generator.md`
- [x] Frontmatter valid YAML (name, description, tools, color)
- [x] Role section explains subagent purpose
- [x] Process has 3 steps: parse, process, return
- [x] Entity template included in full
- [x] Type heuristics table present
- [x] Wiki-link rules specified
- [x] Critical rules section matches gsd-codebase-mapper style
- [x] Returns statistics only (not entity contents)
## Next Phase Readiness
**Prerequisites for 05-02:**
- [x] gsd-entity-generator.md exists
- [x] Agent follows expected patterns (direct write, stats return)
- [x] Entity template matches what PostToolUse hook expects
**Ready for:** Integration into `/gsd:analyze-codebase` command (Plan 05-02)

View File

@@ -1,232 +0,0 @@
---
phase: 05-subagent-codebase-analysis
plan: 02
type: execute
wave: 2
depends_on: ["05-01"]
files_modified:
- commands/gsd/analyze-codebase.md
autonomous: false
must_haves:
truths:
- "Entity generation scales to 500+ files without context exhaustion"
- "Subagent receives only file paths, preserving orchestrator context"
- "Entity files appear in .planning/intel/entities/ after generation"
- "User can run /gsd:analyze-codebase on large codebases without degradation"
artifacts:
- path: "commands/gsd/analyze-codebase.md"
provides: "Refactored command with subagent delegation"
contains: "gsd-entity-generator"
key_links:
- from: "commands/gsd/analyze-codebase.md"
to: "agents/gsd-entity-generator.md"
via: "Task tool spawn"
pattern: "Task.*gsd-entity-generator"
---
<objective>
Refactor analyze-codebase Step 9 to spawn subagent instead of inline batching.
Purpose: Replace the current "batches of 10 via Task tool" approach with a single subagent spawn. The orchestrator selects files (Step 9.2) then spawns `gsd-entity-generator` with the file list. Subagent reads files, generates entities, writes to disk. This preserves orchestrator context for large codebases.
Output: Updated `commands/gsd/analyze-codebase.md`
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/STATE.md
@.planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md
@.planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md
@commands/gsd/analyze-codebase.md
@agents/gsd-entity-generator.md
</context>
<tasks>
<task type="auto">
<name>Task 1: Refactor Step 9 to use subagent delegation</name>
<files>commands/gsd/analyze-codebase.md</files>
<action>
Modify Step 9 of analyze-codebase.md. Keep Steps 9.1 (create directory) and 9.2 (select files) unchanged. Replace Steps 9.3-9.5 with subagent spawn.
**Remove:** Step 9.3 "Generate entities via Task tool batching" (the batch-of-10 pattern)
**Replace with:** Step 9.3 "Spawn entity generator subagent"
New Step 9.3 content:
```markdown
### 9.3 Spawn entity generator subagent
Spawn `gsd-entity-generator` with the selected file list.
**Pass to subagent:**
- Total file count
- Output directory: `.planning/intel/entities/`
- Slug convention: `src/lib/db.ts` -> `src-lib-db` (replace / with -, remove extension, lowercase)
- Entity template (include full template)
- List of absolute file paths (one per line)
**Task tool invocation:**
```python
Task(
prompt=f"""Generate semantic entity documentation for key codebase files.
You are a GSD entity generator. Read source files and create semantic documentation that captures PURPOSE (what/why), not just syntax.
**Parameters:**
- Files to process: {len(selected_files)}
- Output directory: .planning/intel/entities/
- Date: {today}
**Slug convention:**
- src/lib/db.ts -> src-lib-db
- Replace / with -, remove extension, lowercase
**Entity template:**
[Include full template from gsd-entity-generator.md]
**Process:**
For each file path below:
1. Read file content using Read tool
2. Analyze purpose, exports, dependencies
3. Write entity to .planning/intel/entities/{{slug}}.md
4. PostToolUse hook syncs to graph.db automatically
**Files:**
{file_list}
**Return format:**
When complete, return ONLY statistics:
## ENTITY GENERATION COMPLETE
**Files processed:** {{N}}
**Entities created:** {{M}}
**Already existed:** {{K}}
**Errors:** {{E}}
Entities written to: .planning/intel/entities/
Do NOT include entity contents in your response.
""",
subagent_type="gsd-entity-generator"
)
```
**Wait for completion:** Task() blocks until subagent finishes.
**Parse result:** Extract entities_created count from response for final report.
```
**Update Step 9.4:** Rename from "Verify entity generation" to just verify count:
```markdown
### 9.4 Verify entity generation
Confirm entities written:
```bash
ls .planning/intel/entities/*.md 2>/dev/null | wc -l
```
```
**Update Step 9.5:** Keep "Report entity statistics" but update text:
```markdown
### 9.5 Report entity statistics
```
Entity Generation Complete
Entity files created: [N] (from subagent response)
Location: .planning/intel/entities/
Graph database: Updated automatically via PostToolUse hook
Next: Intel hooks will continue incremental updates as you code.
```
```
**Update context section** (line ~32) - add subagent reference:
```markdown
**Execution model (Step 9 - Entity Generation):**
- Orchestrator selects files for entity generation (up to 50 based on priority)
- Spawns `gsd-entity-generator` subagent with file list (paths only, not contents)
- Subagent reads files in fresh 200k context, generates entities, writes to disk
- PostToolUse hook automatically syncs entities to graph.db
- Subagent returns statistics only (not entity contents)
- This preserves orchestrator context for large codebases (500+ files)
```
**Fix outdated slug documentation** (around line 286):
The entity filename convention example is outdated. Change:
```markdown
- Example: src/utils/auth.js -> src--utils--auth-js.md
```
To:
```markdown
- Example: src/utils/auth.js -> src-utils-auth-js.md
```
(Single hyphen, not double hyphen. The hook `gsd-intel-index.js:generateSlug` already uses single hyphen format.)
**Important:** Do NOT pass file contents to subagent. Pass file PATHS only. Subagent reads files itself (fresh context).
</action>
<verify>
Read updated commands/gsd/analyze-codebase.md and confirm:
- Step 9.3 spawns gsd-entity-generator via Task tool
- No batch-of-10 pattern remains
- File paths passed, not file contents
- Context section mentions subagent delegation
- Steps 9.4-9.5 updated for new flow
- Slug example at line ~286 uses single hyphen (src-utils-auth-js.md)
</verify>
<done>
analyze-codebase.md Step 9 refactored. Entity generation now delegates to gsd-entity-generator subagent instead of inline Task batching. Outdated slug documentation corrected.
</done>
</task>
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Subagent delegation for entity generation in /gsd:analyze-codebase</what-built>
<how-to-verify>
1. Navigate to a test project (not this repo)
2. Run `/gsd:analyze-codebase`
3. Observe:
- Steps 1-8 complete (index, conventions, summary)
- Step 9.2 selects files (should see file list)
- Step 9.3 spawns subagent (Task tool call with "gsd-entity-generator")
- Subagent generates entities (should see Read/Write calls in subagent)
- Subagent returns statistics only (not entity contents)
4. Verify `.planning/intel/entities/` contains entity files
5. Verify entity files follow template (frontmatter, Purpose, Exports, Dependencies with [[links]])
6. Verify graph.db was updated (entities appear in summary.md)
</how-to-verify>
<resume-signal>Type "approved" if entity generation works via subagent, or describe issues</resume-signal>
</task>
</tasks>
<verification>
- [ ] Step 9.3 spawns gsd-entity-generator subagent
- [ ] File paths passed to subagent (not file contents)
- [ ] No batch-of-10 pattern in command
- [ ] Context section documents subagent model
- [ ] Slug example uses single hyphen format (matches hook)
- [ ] Entity files created in test project
- [ ] Entities follow template format
- [ ] Graph database updated (via hook)
- [ ] User verification checkpoint passed
</verification>
<success_criteria>
analyze-codebase refactored to use subagent delegation. Entity generation works end-to-end on a test project, with orchestrator context preserved and entities written correctly.
</success_criteria>
<output>
After completion, create `.planning/phases/05-subagent-codebase-analysis/05-02-SUMMARY.md`
</output>

View File

@@ -1,192 +0,0 @@
---
phase: 05-subagent-codebase-analysis
plan: 03
type: execute
wave: 3
depends_on: []
files_modified:
- agents/gsd-indexer.md
autonomous: true
gap_closure: true
must_haves:
truths:
- "Indexer agent definition exists following gsd-entity-generator pattern"
- "Agent reads files and extracts exports/imports using same regex as Step 3"
- "Agent writes index.json directly to disk"
- "Agent returns statistics only (not file contents)"
artifacts:
- path: "agents/gsd-indexer.md"
provides: "Subagent definition for file indexing"
contains: "gsd-indexer"
key_links:
- from: "agents/gsd-indexer.md"
to: ".planning/intel/index.json"
via: "Write tool call"
pattern: "Write.*index\\.json"
---
<objective>
Create gsd-indexer subagent definition for Steps 2-3 file indexing.
Purpose: Define the subagent that reads files and extracts exports/imports. This agent is spawned by `/gsd:analyze-codebase` with a list of file paths (from Glob), reads each file, applies the same regex patterns as current Step 3, and writes the complete index.json to disk. Returns statistics only to preserve orchestrator context.
Output: `agents/gsd-indexer.md`
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/STATE.md
@.planning/phases/05-subagent-codebase-analysis/05-VERIFICATION.md
@agents/gsd-entity-generator.md
@agents/gsd-codebase-mapper.md
@commands/gsd/analyze-codebase.md
</context>
<tasks>
<task type="auto">
<name>Task 1: Create gsd-indexer agent definition</name>
<files>agents/gsd-indexer.md</files>
<action>
Create agent definition following gsd-entity-generator.md structure:
**Frontmatter:**
```yaml
---
name: gsd-indexer
description: Indexes codebase files by extracting exports and imports. Spawned by analyze-codebase with file list. Writes index.json directly to disk.
tools: Read, Write, Bash
color: cyan
---
```
**Role section:**
- Spawned by `/gsd:analyze-codebase` with file paths (from Glob results)
- Reads source files using Read tool
- Extracts exports and imports using regex patterns
- Writes complete index.json to `.planning/intel/index.json`
- Returns statistics only (NOT file contents or index data)
**Why this matters section:**
- Index.json is consumed by convention detection (Step 4)
- Entity generation uses index to find hub files (Step 9.2)
- PostToolUse hook uses index for incremental updates
- Orchestrator MUST NOT load file contents (context exhaustion on 500+ files)
**Process sections:**
1. `parse_input` - Extract from prompt:
- Output path: `.planning/intel/index.json`
- List of absolute file paths (one per line)
- Initialize counters: files_processed=0, exports_found=0, imports_found=0, errors=0
2. `process_each_file` - For each file path:
a. Read file content using Read tool
b. Extract exports using these patterns (EXACTLY as Step 3):
- Named exports: `export\s*\{([^}]+)\}`
- Declaration exports: `export\s+(?:const|let|var|function\*?|async\s+function|class)\s+(\w+)`
- Default exports: `export\s+default\s+(?:function\s*\*?\s*|class\s+)?(\w+)?`
- CommonJS object: `module\.exports\s*=\s*\{([^}]+)\}`
- CommonJS single: `module\.exports\s*=\s*(\w+)\s*[;\n]`
- TypeScript: `export\s+(?:type|interface)\s+(\w+)`
c. Extract imports using these patterns (EXACTLY as Step 3):
- ES6: `import\s+(?:\{[^}]*\}|\*\s+as\s+\w+|\w+)\s+from\s+['"]([^'"]+)['"]`
- Side-effect: `import\s+['"]([^'"]+)['"]` (not preceded by 'from')
- CommonJS: `require\s*\(\s*['"]([^'"]+)['"]\s*\)`
d. Store in index structure:
```javascript
index.files[absolutePath] = {
exports: [], // Array of export names
imports: [], // Array of import sources
indexed: Date.now()
}
```
e. Track statistics: increment files_processed, add to exports_found/imports_found
f. Handle errors: if file can't be read, increment errors and continue
3. `write_index` - Write complete index to disk:
```javascript
{
version: 1,
updated: Date.now(),
files: { /* all file entries */ }
}
```
Write to `.planning/intel/index.json` using Write tool.
4. `return_statistics` - Return ONLY:
```
## INDEXING COMPLETE
**Files processed:** {files_processed}
**Exports found:** {exports_found}
**Imports found:** {imports_found}
**Errors:** {errors}
Index written to: .planning/intel/index.json
```
**Critical rules section:**
- WRITE index.json directly (never return index contents)
- Use EXACT regex patterns from Step 3 (patterns are validated)
- Handle read errors gracefully (log path, continue)
- Return only statistics (~10 lines)
- DO NOT commit (orchestrator handles git)
- File paths in index must be ABSOLUTE paths (key for O(1) lookup)
**Success criteria checklist:**
- All file paths processed
- index.json written with version, updated, files
- Each file entry has exports, imports, indexed timestamp
- Absolute paths as keys (not relative)
- Statistics returned (not index contents)
</action>
<verify>
File exists and contains:
- Frontmatter with name: gsd-indexer, tools: Read/Write/Bash
- Role section explaining spawn context from analyze-codebase
- Why this matters section explaining index consumers
- Process steps: parse_input, process_each_file, write_index, return_statistics
- EXACT regex patterns matching Step 3 of analyze-codebase
- Critical rules section (write directly, return stats only)
- Success criteria checklist
</verify>
<done>
`agents/gsd-indexer.md` exists with complete agent definition. Agent uses same regex patterns as current Step 3. Ready to be spawned by analyze-codebase.
</done>
</task>
</tasks>
<verification>
- [ ] File exists at `agents/gsd-indexer.md`
- [ ] Frontmatter valid YAML (name, description, tools, color)
- [ ] Role section explains subagent purpose (spawned by analyze-codebase)
- [ ] Process has 4 steps: parse, process, write, return
- [ ] Export regex patterns match Step 3 exactly (6 patterns)
- [ ] Import regex patterns match Step 3 exactly (3 patterns)
- [ ] Index schema matches Step 5 format (version, updated, files)
- [ ] Critical rules section matches gsd-entity-generator style
- [ ] Returns statistics only (not index contents)
</verification>
<success_criteria>
gsd-indexer agent definition complete. Agent can be spawned with file paths and will read files, extract exports/imports using validated regex patterns, and write index.json directly to disk.
</success_criteria>
<output>
After completion, create `.planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md`
</output>

View File

@@ -1,135 +0,0 @@
---
phase: 05-subagent-codebase-analysis
plan: 03
title: gsd-indexer Agent Definition
subsystem: agents
tags: [subagent, indexing, file-analysis, context-preservation]
dependency-graph:
requires:
- 05-01 (gsd-entity-generator pattern reference)
- gsd-codebase-mapper.md (agent structure pattern)
- analyze-codebase.md Step 3 (regex patterns)
provides:
- gsd-indexer subagent definition
- File indexing via subagent delegation
- Context-preserving index generation
affects:
- 05-04 (analyze-codebase integration)
- Steps 2-3 context exhaustion fix
tech-stack:
added: []
patterns:
- Subagent direct-write pattern
- Statistics-only return pattern
- Regex-based export/import extraction
key-files:
created:
- agents/gsd-indexer.md
modified: []
decisions:
- id: absolute-path-keys
choice: Use absolute paths as index keys
rationale: O(1) lookup, matches existing index.json schema
metrics:
duration: 1 min 23 sec
completed: 2026-01-21
---
# Phase 05 Plan 03: gsd-indexer Agent Definition Summary
**One-liner:** Subagent definition for file indexing that extracts exports/imports using validated regex patterns and writes index.json directly to disk.
## What Was Built
Created `agents/gsd-indexer.md` following the `gsd-entity-generator.md` pattern:
**Agent structure:**
- Frontmatter with name, description, tools (Read, Write, Bash), color
- Role section explaining spawn context from `/gsd:analyze-codebase`
- `<why_this_matters>` section explaining index consumers (convention detection, entity generation, PostToolUse hook)
- 4-step process: parse_input, process_each_file, write_index, return_statistics
**Regex patterns (exact match with analyze-codebase Step 3):**
Export patterns:
| Pattern | Regex | Purpose |
|---------|-------|---------|
| Named exports | `export\s*\{([^}]+)\}` | `export { a, b }` |
| Declaration | `export\s+(?:const\|let\|var\|function\*?\|async\s+function\|class)\s+(\w+)` | `export const foo` |
| Default | `export\s+default\s+(?:function\s*\*?\s*\|class\s+)?(\w+)?` | `export default` |
| CommonJS object | `module\.exports\s*=\s*\{([^}]+)\}` | `module.exports = { }` |
| CommonJS single | `module\.exports\s*=\s*(\w+)\s*[;\n]` | `module.exports = X` |
| TypeScript | `export\s+(?:type\|interface)\s+(\w+)` | `export type/interface` |
Import patterns:
| Pattern | Regex | Purpose |
|---------|-------|---------|
| ES6 | `import\s+(?:\{[^}]*\}\|\*\s+as\s+\w+\|\w+)\s+from\s+['"]([^'"]+)['"]` | `import X from 'y'` |
| Side-effect | `import\s+['"]([^'"]+)['"]` | `import 'styles.css'` |
| CommonJS | `require\s*\(\s*['"]([^'"]+)['"]\s*\)` | `require('x')` |
**Index schema (matches Step 5):**
```javascript
{
version: 1,
updated: Date.now(),
files: {
"/absolute/path": {
exports: [],
imports: [],
indexed: Date.now()
}
}
}
```
**Critical rules:**
- Write index.json directly (never return contents)
- Use exact regex patterns from Step 3
- Absolute paths as keys for O(1) lookup
- Handle read errors gracefully (log, continue)
- Return statistics only (~10 lines)
## Tasks Completed
| Task | Name | Commit | Key Files |
|------|------|--------|-----------|
| 1 | Create gsd-indexer agent definition | 5d03e14 | agents/gsd-indexer.md |
## Deviations from Plan
None - plan executed exactly as written.
## Decisions Made
1. **Absolute path keys**
- Continues existing index.json schema decision from 01-01
- Enables O(1) lookup for any file path
- Consistent with PostToolUse hook expectations
## Verification Results
- [x] File exists at `agents/gsd-indexer.md`
- [x] Frontmatter valid YAML (name, description, tools, color)
- [x] Role section explains subagent purpose (spawned by analyze-codebase)
- [x] Process has 4 steps: parse, process, write, return
- [x] Export regex patterns match Step 3 exactly (6 patterns)
- [x] Import regex patterns match Step 3 exactly (3 patterns)
- [x] Index schema matches Step 5 format (version, updated, files)
- [x] Critical rules section matches gsd-entity-generator style
- [x] Returns statistics only (not index contents)
## Next Phase Readiness
**Prerequisites for 05-04:**
- [x] gsd-indexer.md exists
- [x] Agent follows expected patterns (direct write, stats return)
- [x] Regex patterns validated against analyze-codebase Step 3
- [x] Index schema matches existing expectations
**Ready for:** Integration into `/gsd:analyze-codebase` command Steps 2-3 (Plan 05-04)

View File

@@ -1,272 +0,0 @@
---
phase: 05-subagent-codebase-analysis
plan: 04
type: execute
wave: 4
depends_on: ["05-03"]
files_modified:
- commands/gsd/analyze-codebase.md
autonomous: false
gap_closure: true
must_haves:
truths:
- "Orchestrator never reads file contents during Steps 2-3"
- "Indexer subagent is spawned with file paths from Glob"
- "Subagent writes index.json, orchestrator receives statistics only"
- "User can run /gsd:analyze-codebase on 500+ file codebases without context exhaustion"
artifacts:
- path: "commands/gsd/analyze-codebase.md"
provides: "Refactored command with subagent delegation for indexing"
contains: "gsd-indexer"
key_links:
- from: "commands/gsd/analyze-codebase.md"
to: "agents/gsd-indexer.md"
via: "Task tool spawn"
pattern: "Task.*gsd-indexer"
- from: "Step 2 Glob"
to: "Step 3 subagent spawn"
via: "file paths array"
pattern: "Glob.*file_paths"
---
<objective>
Refactor analyze-codebase Steps 2-3 to spawn indexer subagent.
Purpose: Replace the current inline file reading (Step 3: "Read file content using Read tool") with a single subagent spawn. The orchestrator runs Glob to find files (Step 2), spawns `gsd-indexer` with the file list, and receives index.json back. This prevents context exhaustion on large codebases where 250+ file reads were causing the orchestrator to die before reaching Step 9.
Output: Updated `commands/gsd/analyze-codebase.md`
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/STATE.md
@.planning/phases/05-subagent-codebase-analysis/05-VERIFICATION.md
@.planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md
@commands/gsd/analyze-codebase.md
@agents/gsd-indexer.md
</context>
<tasks>
<task type="auto">
<name>Task 1: Refactor Steps 2-3 to use indexer subagent</name>
<files>commands/gsd/analyze-codebase.md</files>
<action>
Modify Steps 2 and 3 of analyze-codebase.md. Step 2 stays mostly the same (Glob). Step 3 is completely replaced with subagent spawn.
**Update Step 2: Find all indexable files**
Keep the Glob pattern and exclusion list, but change the output description:
```markdown
## Step 2: Find all indexable files
Use Glob tool with pattern: `**/*.{js,ts,jsx,tsx,mjs,cjs}`
Exclude directories (skip any path containing):
- node_modules
- dist
- build
- .git
- vendor
- coverage
- .next
- __pycache__
Filter results to remove excluded paths. Store as `file_paths` array.
**Output:** List of absolute file paths for indexing. Do NOT read file contents.
```
**Replace Step 3: Process each file**
Remove the entire current Step 3 (lines ~67-100) which contains:
- "Read file content using Read tool"
- Export regex patterns inline
- Import regex patterns inline
- Store in index structure
Replace with:
```markdown
## Step 3: Spawn indexer subagent
Spawn `gsd-indexer` subagent with the file paths from Step 2.
**Why subagent delegation:**
- Orchestrator would exhaust context reading 500+ files inline
- Subagent gets fresh 200k context for file reading
- Orchestrator only handles file paths (small)
- Subagent writes index.json directly (large)
**Task tool invocation:**
```python
file_list = "\n".join(file_paths) # From Step 2 Glob results
Task(
prompt=f"""Index codebase files by extracting exports and imports.
You are a GSD indexer. Read source files and extract exports/imports using regex patterns.
**Parameters:**
- Files to process: {len(file_paths)}
- Output path: .planning/intel/index.json
**Export patterns:**
- Named: export\\s*\\{{([^}}]+)\\}}
- Declaration: export\\s+(?:const|let|var|function\\*?|async\\s+function|class)\\s+(\\w+)
- Default: export\\s+default\\s+(?:function\\s*\\*?\\s*|class\\s+)?(\\w+)?
- CommonJS object: module\\.exports\\s*=\\s*\\{{([^}}]+)\\}}
- CommonJS single: module\\.exports\\s*=\\s*(\\w+)\\s*[;\\n]
- TypeScript: export\\s+(?:type|interface)\\s+(\\w+)
**Import patterns:**
- ES6: import\\s+(?:\\{{[^}}]*\\}}|\\*\\s+as\\s+\\w+|\\w+)\\s+from\\s+['\"]([^'\"]+)['\"]
- Side-effect: import\\s+['\"]([^'\"]+)['\"]
- CommonJS: require\\s*\\(\\s*['\"]([^'\"]+)['\"]\\s*\\)
**Index schema:**
```json
{{
"version": 1,
"updated": {timestamp},
"files": {{
"/absolute/path/file.js": {{
"exports": ["name1", "name2"],
"imports": ["source1", "source2"],
"indexed": {timestamp}
}}
}}
}}
```
**Process:**
For each file path below:
1. Read file content using Read tool
2. Apply export regex patterns, collect export names
3. Apply import regex patterns, collect import sources
4. Store in index structure with absolute path as key
**Files:**
{file_list}
**Return format:**
When complete, return ONLY statistics:
## INDEXING COMPLETE
**Files processed:** {{N}}
**Exports found:** {{M}}
**Imports found:** {{K}}
**Errors:** {{E}}
Index written to: .planning/intel/index.json
Do NOT return index contents.
""",
subagent_type="gsd-indexer"
)
```
**Wait for completion:** Task() blocks until subagent finishes.
**Verify index created:**
```bash
ls -la .planning/intel/index.json
```
```
**Update context section** (around lines 21-38):
Add to the context section, before the existing "Execution model (Step 9)" section:
```markdown
**Execution model (Steps 2-3 - Indexing):**
- Orchestrator finds file paths via Glob (Step 2)
- Spawns `gsd-indexer` subagent with file paths only (Step 3)
- Subagent reads files in fresh 200k context, applies regex patterns
- Subagent writes index.json directly to disk
- Subagent returns statistics only (not file contents or index data)
- This prevents context exhaustion on large codebases (500+ files)
```
**Remove inline regex documentation:**
The regex patterns are now documented in the subagent prompt. Remove any duplicate documentation that described inline processing. The command should make clear that file reading is DELEGATED, not performed inline.
**Verify Step 4 still works:**
Step 4 "Detect conventions" reads from index.json that the subagent wrote. This should work unchanged since the index format is the same.
</action>
<verify>
Read updated commands/gsd/analyze-codebase.md and confirm:
- Step 2 mentions storing as "file_paths" array, NOT reading contents
- Step 3 spawns gsd-indexer via Task tool
- No "Read file content" instruction in orchestrator steps
- Context section documents subagent delegation for Steps 2-3
- Step 4+ remain unchanged (read index.json from disk)
- Regex patterns appear in subagent prompt (not inline)
</verify>
<done>
analyze-codebase.md Steps 2-3 refactored. Orchestrator only Globs for paths, then spawns indexer subagent. File reading is fully delegated to subagent with fresh context.
</done>
</task>
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Subagent delegation for indexing in /gsd:analyze-codebase (Steps 2-3)</what-built>
<how-to-verify>
1. Navigate to a LARGE test project (100+ files, ideally 250+)
2. Delete existing intel if any: `rm -rf .planning/intel`
3. Run `/gsd:analyze-codebase`
4. Observe execution:
- Step 1: Creates directory
- Step 2: Runs Glob, collects file paths (does NOT read files)
- Step 3: Spawns gsd-indexer subagent with file list
- Subagent reads files (should see many Read tool calls in subagent output)
- Subagent writes index.json directly
- Subagent returns statistics only
- Step 4+: Orchestrator reads index.json, continues with conventions
- Step 9: Spawns gsd-entity-generator (existing flow)
5. Verify:
- Orchestrator context NOT exhausted (no early death)
- `.planning/intel/index.json` exists with file entries
- `.planning/intel/conventions.json` has detected patterns
- `.planning/intel/summary.md` exists
- Entity generation works (if Step 9 executed)
6. Compare orchestrator context usage before/after refactor
- Before: Orchestrator read all files -> context exhaustion at ~250 files
- After: Orchestrator only handles paths -> survives 500+ files
</how-to-verify>
<resume-signal>Type "approved" if /gsd:analyze-codebase completes on large codebase without context exhaustion, or describe issues</resume-signal>
</task>
</tasks>
<verification>
- [ ] Step 2 outputs file_paths array (no file reading)
- [ ] Step 3 spawns gsd-indexer subagent
- [ ] Subagent prompt includes all 6 export patterns
- [ ] Subagent prompt includes all 3 import patterns
- [ ] Context section documents indexing subagent delegation
- [ ] No "Read file content" in orchestrator steps
- [ ] Step 4+ reads index.json from disk (unchanged)
- [ ] Command runs successfully on 250+ file codebase
- [ ] Orchestrator context preserved (no early death)
- [ ] User verification checkpoint passed
</verification>
<success_criteria>
analyze-codebase refactored with full subagent delegation. Both indexing (Steps 2-3) and entity generation (Step 9) now use subagents. Command works on 500+ file codebases without context exhaustion.
</success_criteria>
<output>
After completion, create `.planning/phases/05-subagent-codebase-analysis/05-04-SUMMARY.md`
</output>

View File

@@ -1,774 +0,0 @@
# Phase 5: Subagent Codebase Analysis - Research
**Researched:** 2026-01-20
**Domain:** Subagent orchestration, context window management, Claude Code Task tool
**Confidence:** HIGH
## Summary
Phase 5 refactors the current `/gsd:analyze-codebase` entity generation from main-context execution to subagent delegation. The current implementation (Phase 4) has Claude executing the command generate entity content directly via the Task tool with batches of 10 files, but ALL exploration and decision-making happens in the main orchestrator context.
**The problem:** On large codebases (500+ files), the orchestrator exhausts context by:
1. Reading all files during selection (identifying which 50 files to generate entities for)
2. Orchestrating batch splits and subagent spawns
3. Collecting and validating subagent results
**The solution:** Delegate the entire entity generation phase to a subagent, following the `gsd-codebase-mapper.md` pattern:
- Orchestrator provides minimal instructions + file list
- Subagent operates in fresh 200k context
- Subagent reads files, generates entities, writes directly to disk
- Subagent returns only confirmation (not entity contents)
**Current baseline:** `gsd-codebase-mapper.md` demonstrates successful subagent delegation for analysis tasks. It spawns with a focus area, explores thoroughly, writes documents directly, and returns ~10 lines of confirmation. This pattern scales because the orchestrator never loads document contents.
**Primary recommendation:** Extract entity generation (Step 9) from `/gsd:analyze-codebase` into a dedicated `gsd-entity-generator` subagent. Orchestrator handles Steps 1-8 (indexing), then spawns subagent with file list. Subagent generates all entities and returns statistics only.
## Standard Stack
### Core
| Library | Version | Purpose | Why Standard |
|---------|---------|---------|--------------|
| Claude Code Task tool | Built-in | Subagent spawning with model selection | Official Claude Code orchestration primitive |
| sql.js | 1.12.0+ | Graph database in subagent context | Subagent needs graph access to resolve [[wiki-links]] |
### Supporting
| Component | Version | Purpose | When to Use |
|---------|---------|---------|-------------|
| gsd-codebase-mapper.md | Current | Reference pattern for subagent delegation | Template for entity-generator architecture |
| gsd-executor.md | Current | Reference for Task tool usage | Shows how to spawn with model profile |
### Installation
No new dependencies required. Phase 5 refactors existing architecture.
### Alternatives Considered
| Instead of | Could Use | Tradeoff |
|------------|-----------|----------|
| Single subagent for all files | Multiple subagents (1 per batch) | Multiple subagents = parallel processing but complex orchestration and potential race conditions in graph.db writes |
| Subagent with file list | Subagent discovers files itself | Discovery requires reading index.json anyway, orchestrator already has this data |
| Inline in hook | Keep current approach | Hooks have strict execution limits, can't spawn subagents or handle 500+ files |
## Architecture Patterns
### Recommended Execution Flow
```
User: /gsd:analyze-codebase
Orchestrator (main context):
├── Steps 1-8: Index creation (existing)
│ ├── Scan files with Glob
│ ├── Extract exports/imports
│ ├── Write index.json, conventions.json, summary.md
│ └── Identify 50 key files for entity generation
│
└── Step 9: Spawn entity-generator subagent
├── Pass: file list, index data, config
├── Subagent operates in fresh 200k context
└── Returns: { entities_created: N, skipped: M }
Subagent (gsd-entity-generator):
├── Load file list from prompt
├── For each file:
│ ├── Read file content
│ ├── Generate entity markdown (Claude's own analysis)
│ ├── Write to .planning/intel/entities/{slug}.md
│ └── PostToolUse hook syncs to graph.db
└── Return confirmation statistics
```
### Pattern 1: Minimal Orchestrator Handoff
**What:** Orchestrator passes only essential data, not full file contents
**When to use:** When subagent needs to read files itself (maintains fresh context)
**Example:**
```markdown
Task(
prompt=f"""Generate semantic entity files for key codebase files.
Entity generation parameters:
- Total files to process: {len(selected_files)}
- Output directory: .planning/intel/entities/
- Slug convention: src/lib/db.ts -> src-lib-db
Files to process:
{chr(10).join(selected_files)}
For each file:
1. Read the file content using Read tool
2. Analyze purpose, exports, dependencies
3. Generate entity markdown following template
4. Write to .planning/intel/entities/{{slug}}.md
Entity template:
---
path: {{file_path}}
type: [module|component|util|config|api|hook|service|model]
updated: {today}
status: active
---
# {{filename}}
## Purpose
[1-3 sentences: What does this file do? Why does it exist?]
[... rest of template ...]
After all files processed, return statistics:
- Entities created: N
- Files skipped: M (already existed)
- Errors: K (if any)
""",
subagent_type="gsd-entity-generator",
model="{model}"
)
```
**Why minimal:** Passing file contents in prompt exhausts orchestrator context (defeats purpose of subagent delegation).
### Pattern 2: Subagent Direct Write
**What:** Subagent writes entity files directly, doesn't return contents to orchestrator
**When to use:** Always (learned from gsd-codebase-mapper.md)
**Anti-pattern:**
```python
# DON'T: Return entity contents to orchestrator
entities = []
for file in files:
entity = generate_entity(file) # Claude generates
entities.append(entity) # Accumulates in context
return entities # Passes back to orchestrator
```
**Correct pattern:**
```python
# DO: Write directly, return only confirmation
for file in files:
entity_content = generate_entity(file) # Claude generates
Write(path=f".planning/intel/entities/{slug}.md", content=entity_content)
# PostToolUse hook automatically syncs to graph.db
return {
"entities_created": len(files),
"location": ".planning/intel/entities/"
}
```
### Pattern 3: Model Profile Resolution
**What:** Orchestrator resolves model profile, passes specific model to Task tool
**When to use:** Every subagent spawn (ensures consistent model selection)
**Example:**
```bash
# Read model profile from config
MODEL_PROFILE=$(cat .planning/config.json 2>/dev/null | \
grep -o '"model_profile"[[:space:]]*:[[:space:]]*"[^"]*"' | \
grep -o '"[^"]*"$' | tr -d '"' || echo "balanced")
```
```python
# Model lookup table for gsd-entity-generator
model_map = {
"quality": "claude-opus-4-5-20251101",
"balanced": "claude-sonnet-4-5-20250929",
"budget": "claude-sonnet-4-5-20250929"
}
model = model_map.get(MODEL_PROFILE, "claude-sonnet-4-5-20250929")
```
**Why this matters:** Entity generation is semantic analysis (requires strong reasoning). Sonnet 4.5 adequate for balanced/budget, Opus 4.5 for quality profile.
### Pattern 4: PostToolUse Hook Integration
**What:** Subagent writes entities, hook syncs to graph.db automatically
**When to use:** Always (no explicit sync needed in subagent)
**Flow:**
```
Subagent writes: .planning/intel/entities/src-lib-db.md
↓
PostToolUse hook (gsd-intel-index.js) detects Write tool
↓
isEntityFile(path) returns true
↓
syncEntityToGraph(path):
- Extracts frontmatter
- Extracts [[wiki-links]]
- Upserts node to graph.db
- Inserts edges
- Persists database
↓
regenerateEntitySummary():
- Generates new summary from graph
- Writes summary.md
```
**Critical insight:** Subagent doesn't need graph access for writes. Hook handles all graph operations. Subagent only needs to write well-formed entity markdown.
### Anti-Patterns to Avoid
**Anti-pattern 1: Orchestrator reads files for subagent**
```python
# DON'T: Load all file contents in orchestrator
file_contents = {}
for file_path in selected_files:
file_contents[file_path] = Read(file_path) # Exhausts orchestrator context
Task(prompt=f"Generate entities for: {file_contents}", ...) # Too late, context blown
```
**Why it fails:** Defeats purpose of subagent delegation (500 files × 5KB avg = 2.5MB of code in orchestrator context).
**Anti-pattern 2: Multiple parallel subagents**
```python
# DON'T: Spawn one subagent per file (or per small batch)
for file in selected_files:
Task(prompt=f"Generate entity for {file}", ...) # 50 subagents = chaos
```
**Why it fails:**
- 50 concurrent subagents writing to `.planning/intel/entities/`
- 50 concurrent PostToolUse hooks writing to `graph.db`
- Race conditions in sql.js export/import (no file locking)
- Excessive API calls (50 × context overhead)
**Anti-pattern 3: Subagent returns generated content**
```markdown
## ENTITY GENERATION COMPLETE
**Files processed:** 50
**Generated entities:**
[... 50 × 500 lines of entity content ...] <!-- Explodes orchestrator context -->
```
**Why it fails:** Orchestrator asked for entity generation, not entity contents. Return confirmation only.
## Don't Hand-Roll
| Problem | Don't Build | Use Instead | Why |
|---------|-------------|-------------|-----|
| Subagent orchestration | Custom process spawning | Claude Code Task tool | Built-in, handles model selection, output capture, error handling |
| File batching | Complex batch scheduling | Single subagent with full file list | Subagent has 200k context (enough for 500 file paths), no coordination overhead |
| Entity template | Embedded in subagent logic | Pass template in prompt | Template may evolve, keep it in prompt not in subagent code |
| Graph database access | Direct sql.js in subagent | PostToolUse hook handles sync | Hooks are designed for this, no need to duplicate graph logic |
| Progress tracking | Real-time updates to orchestrator | Batch completion confirmation | Subagent works in isolation, returns final stats |
**Key insight:** The Task tool is designed for this exact pattern. Don't try to implement manual process forking, IPC, or result aggregation. Task() blocks until subagent completes, returns output directly.
## Common Pitfalls
### Pitfall 1: Orchestrator Context Bloat
**What goes wrong:** Orchestrator loads too much data before spawning subagent, exhausts context anyway.
**Why it happens:** Instinct to "prepare everything" for the subagent (read files, validate, format).
**How to avoid:**
- Orchestrator's job: Identify WHICH files (paths only)
- Subagent's job: Read files and process
- Pass file paths, not file contents
- Trust subagent to handle file reading
**Warning signs:** Orchestrator context usage >50% before Task() call.
### Pitfall 2: Subagent Result Explosion
**What goes wrong:** Subagent returns all generated entity content to orchestrator.
**Why it happens:** Thinking orchestrator needs to "validate" or "display" results.
**How to avoid:**
- Subagent writes directly to `.planning/intel/entities/`
- Return only: `{ entities_created: N, errors: [] }`
- Orchestrator reports to user: "Created N entities in .planning/intel/entities/"
- User can inspect entity files themselves
**Warning signs:** Task() return value contains >1000 lines of text.
### Pitfall 3: Race Conditions in graph.db
**What goes wrong:** Multiple concurrent writes to graph.db corrupt database.
**Why it happens:** PostToolUse hook fires for every Write. 50 entity writes = 50 hooks attempting sql.js export/import simultaneously.
**How to avoid:**
- Use single subagent (not parallel subagents)
- PostToolUse hooks run sequentially per Write operation
- sql.js persistence is synchronous (no async race conditions within single process)
- If implementing parallel subagents in future: file locking or write queue required
**Warning signs:** graph.db corruption, missing edges, "database is locked" errors.
### Pitfall 4: Missing Entity Template
**What goes wrong:** Subagent generates entities in wrong format, hook can't parse frontmatter/links.
**Why it happens:** Template not provided in subagent prompt, or template is incomplete.
**How to avoid:**
- Include EXACT entity template in Task prompt
- Show examples with all sections (frontmatter, Purpose, Exports, Dependencies, Used By, Notes)
- Specify [[wiki-link]] format for internal dependencies
- Test generated entities: hook should extract frontmatter + links successfully
**Warning signs:** Entities created but not appearing in graph queries, summary.md unchanged.
### Pitfall 5: No Progress Visibility
**What goes wrong:** Subagent processes 500 files silently, user has no idea if it's working or stuck.
**Why it happens:** Task() blocks until completion, no intermediate output.
**How to avoid:**
- For large operations (100+ files), consider chunking:
- Orchestrator: Split 500 files into 10 batches of 50
- Spawn subagent per batch (sequential, not parallel)
- Report progress: "Batch 3/10 complete (150/500 files)"
- For moderate operations (50-100 files): Single subagent is fine, document expected duration
**Warning signs:** User cancels command thinking it's frozen (actually still processing).
## Code Examples
### Complete Subagent Spawn (Orchestrator)
```markdown
## Step 9: Generate Semantic Entities via Subagent
Read model profile from config:
```bash
MODEL_PROFILE=$(cat .planning/config.json 2>/dev/null | \
grep -o '"model_profile"[[:space:]]*:[[:space:]]*"[^"]*"' | \
grep -o '"[^"]*"$' | tr -d '"' || echo "balanced")
```
Resolve model for gsd-entity-generator:
| Profile | Model |
|---------|-------|
| quality | claude-opus-4-5-20251101 |
| balanced | claude-sonnet-4-5-20250929 |
| budget | claude-sonnet-4-5-20250929 |
Spawn entity generator subagent:
```python
# Build file list from Step 9a selection
file_list = "\n".join(selected_files)
today = date.today().isoformat()
Task(
prompt=f"""Generate semantic entity documentation for key codebase files.
You are a GSD entity generator. You read source files and create semantic documentation that captures PURPOSE (what/why), not just syntax.
**Parameters:**
- Files to process: {len(selected_files)}
- Output directory: .planning/intel/entities/
- Date: {today}
**Slug convention:**
- src/lib/db.ts → src-lib-db
- Replace / with -, remove extension
**Entity template (use EXACTLY this format):**
---
path: {{absolute_file_path}}
type: [module|component|util|config|api|hook|service|model|test]
updated: {today}
status: active
---
# {{filename}}
## Purpose
[1-3 sentences explaining what this file does and why it exists. Focus on the problem it solves, not implementation details.]
## Exports
[For each export, provide signature and brief purpose:]
- `functionName(params): ReturnType` - What it does
- `ClassName` - What it represents
If no exports: "None"
## Dependencies
[Internal dependencies as [[wiki-links]], external as plain text:]
- [[internal-file-slug]] - Why this dependency exists
- external-package - What it provides
If no dependencies: "None"
## Used By
TBD
## Notes
[Optional: Important patterns, gotchas, or context. Omit if nothing notable.]
**Process:**
For each file path below:
1. Read file content using Read tool
2. Analyze exports, imports, purpose
3. Write entity to .planning/intel/entities/{{slug}}.md
4. PostToolUse hook will sync to graph.db automatically
**Files:**
{file_list}
**Return format:**
When all files processed, return ONLY this structure:
```
## ENTITY GENERATION COMPLETE
**Files processed:** {{N}}
**Entities created:** {{M}}
**Already existed:** {{K}}
**Errors:** {{E}} (if any, list file paths)
Entities written to: .planning/intel/entities/
```
Do NOT include entity contents in your response.
""",
subagent_type="gsd-entity-generator",
model=model
)
```
Wait for subagent completion. Task() blocks until done.
Parse result for statistics:
- Extract entities_created count
- Report to user
```
### Subagent Response Format (gsd-entity-generator)
```markdown
## ENTITY GENERATION COMPLETE
**Files processed:** 47
**Entities created:** 47
**Already existed:** 0
**Errors:** 0
Entities written to: .planning/intel/entities/
```
**Critical:** Subagent does NOT return entity contents. Only statistics.
### Orchestrator Final Report
```markdown
Codebase Analysis Complete
Files indexed: 347
Exports found: 1,423
Imports found: 2,891
Conventions detected:
- Naming: camelCase (87%)
- Directories: components/ (23 files), lib/ (12 files), api/ (8 files)
- Patterns: *.test.ts (34 files), *.config.ts (5 files)
**Entities created: 47**
- Location: .planning/intel/entities/
- Graph database: Updated automatically
- Summary: Regenerated with dependency hotspots
Files created:
- .planning/intel/index.json
- .planning/intel/conventions.json
- .planning/intel/summary.md
- .planning/intel/entities/*.md (47 files)
- .planning/intel/graph.db
Next: Intel hooks will continue incremental updates as you code.
```
### Entity Generator Subagent Definition (New File)
```markdown
---
name: gsd-entity-generator
description: Generates semantic entity documentation for codebase files. Spawned by analyze-codebase with file list. Writes entities directly to disk.
tools: Read, Write, Bash
color: cyan
---
<role>
You are a GSD entity generator. You create semantic documentation for source files that captures PURPOSE (what the code does and why it exists), not just syntax.
You are spawned by `/gsd:analyze-codebase` with a list of file paths.
Your job: Read each file, analyze its purpose, write entity markdown to `.planning/intel/entities/`, return statistics only.
</role>
<process>
<step name="parse_file_list">
Extract file paths from your prompt. You'll receive:
- Total file count
- Output directory path
- Slug convention rules
- Entity template
- List of absolute file paths
Parse file paths into an array for processing.
</step>
<step name="process_each_file">
For each file path:
1. **Read file content:**
```bash
Read(file_path)
```
2. **Analyze the file:**
- What is the purpose? (Why does this file exist?)
- What does it export? (Functions, classes, types)
- What does it import? (Dependencies and why)
- What type of module is it? (api, component, util, service, etc.)
3. **Generate slug:**
- Remove leading slashes
- Remove file extension
- Replace / and . with -
- Lowercase everything
- Example: `src/lib/db.ts` → `src-lib-db`
4. **Build entity content using template:**
- Frontmatter with path, type, date, status
- Purpose section (1-3 sentences)
- Exports section (signatures + descriptions)
- Dependencies section ([[wiki-links]] for internal, plain text for external)
- Used By: Always "TBD" (graph analysis fills this later)
- Notes: Optional (only if important context)
5. **Write entity file:**
```bash
Write(
file_path=f".planning/intel/entities/{slug}.md",
content=entity_markdown
)
```
6. **Track statistics:**
- Count files processed
- Count entities created
- Track any errors
**Important:** PostToolUse hook automatically syncs entity to graph.db. You don't need to touch the graph.
</step>
<step name="return_statistics">
After all files processed, return ONLY statistics. Do NOT include entity contents.
Format:
```
## ENTITY GENERATION COMPLETE
**Files processed:** {N}
**Entities created:** {M}
**Already existed:** {K}
**Errors:** {E}
Entities written to: .planning/intel/entities/
```
If errors occurred, list file paths that failed (not the error messages themselves).
</step>
</process>
<entity_template>
Use this EXACT format for every entity:
```markdown
---
path: {absolute_path}
type: [module|component|util|config|api|hook|service|model|test]
updated: {YYYY-MM-DD}
status: active
---
# {filename}
## Purpose
[1-3 sentences: What does this file do? Why does it exist? What problem does it solve? Focus on the "why", not implementation details.]
## Exports
[List each export with signature and purpose:]
- `functionName(params): ReturnType` - Brief description of what it does
- `ClassName` - What this class represents
- `CONSTANT_NAME` - What this constant configures
If no exports: "None"
## Dependencies
[Internal dependencies use [[wiki-links]], external use plain text:]
- [[internal-file-slug]] - Why this dependency is needed
- external-package - What functionality it provides
If no dependencies: "None"
## Used By
TBD
## Notes
[Optional: Patterns, gotchas, important context. Omit section if nothing notable.]
```
</entity_template>
<type_heuristics>
Determine entity type from file path and content:
| Type | Indicators |
|------|-----------|
| api | In api/, routes/, endpoints/ directory, exports route handlers |
| component | In components/, exports React/Vue/etc components |
| util | In utils/, lib/, helpers/, exports utility functions |
| config | In config/, *.config.*, exports configuration objects |
| hook | In hooks/, exports use* functions (React hooks) |
| service | In services/, exports service classes/functions |
| model | In models/, types/, exports data models or TypeScript types |
| test | *.test.*, *.spec.*, contains test suites |
| module | Default if unclear, general-purpose module |
</type_heuristics>
<wiki_link_rules>
**Internal dependencies** (files in the codebase):
- Convert to slug format
- Wrap in [[double brackets]]
- Example: Import from `../../lib/db.ts` → Dependency: `[[src-lib-db]]`
**External dependencies** (npm packages):
- Plain text, no brackets
- Example: `import { z } from 'zod'` → Dependency: `zod - Schema validation`
**When unsure if internal/external:**
- If import path starts with `.` or `@/` → internal (wiki-link)
- If import path is package name → external (plain text)
</wiki_link_rules>
<success_criteria>
Entity generation complete when:
- [ ] All file paths processed
- [ ] Each entity file written to `.planning/intel/entities/`
- [ ] Entity markdown follows template exactly
- [ ] Frontmatter is valid YAML
- [ ] Purpose section is substantive (not just "This file exports X")
- [ ] Internal dependencies use [[wiki-links]]
- [ ] Statistics returned (not entity contents)
</success_criteria>
```
## State of the Art
| Old Approach | Current Approach | When Changed | Impact |
|--------------|------------------|--------------|--------|
| Orchestrator generates entities | Subagent delegation | Phase 5 (planned) | Prevents context exhaustion on large codebases |
| Sequential processing in main context | Fresh 200k subagent context | Phase 5 (planned) | Scales to 500+ files without orchestrator bloat |
| Batches of 10 with multiple Task calls | Single subagent, all files | Phase 5 (planned) | Simpler orchestration, no batch coordination |
| Hook generates entities via `claude -p` | Subagent generates entities | Phase 5 (planned) | Richer semantic analysis (subagent has full context vs hook's one-shot) |
**Deprecated/outdated:**
- **Hook-based entity generation** (Phase 4 uses `execSync('claude -p')` in hook): Limited to 30s timeout, no retry, crude prompt passing. Phase 5 moves to proper subagent pattern.
- **Multiple parallel subagents** for entity batches: Overcomplicated, race condition risks, excessive overhead. Single subagent with full file list is cleaner.
## Open Questions
### 1. **Optimal File Count per Subagent**
- **What we know:** Subagent has 200k context. File path = ~50 chars avg. 500 paths = 25KB (negligible).
- **What's unclear:** At what file count does entity generation hit subagent context limits? 500 files? 1000?
- **Recommendation:** Start with single subagent for up to 500 files. If codebases >500 common, implement chunking (spawn 5 subagents of 100 files each, sequentially).
### 2. **Should Orchestrator Pre-read Index Data?**
- **What we know:** Subagent needs to understand codebase conventions (naming patterns, directory purposes) for better entity generation.
- **What's unclear:** Should orchestrator pass conventions.json content in prompt, or should subagent read it?
- **Recommendation:** Pass conventions in prompt (it's <2KB JSON). Saves subagent a Read operation, provides useful context for entity type classification.
### 3. **Entity Regeneration Strategy**
- **What we know:** Hook regenerates entities when file signature changes (exports/imports differ).
- **What's unclear:** Should bulk regeneration (via /gsd:analyze-codebase) skip existing entities or overwrite?
- **Recommendation:** Skip existing entities by default (check if `.planning/intel/entities/{slug}.md` exists). Add flag: `--force-regenerate` to overwrite all. This prevents destroying manual edits to entities.
### 4. **Error Handling for Unparseable Files**
- **What we know:** Some files might be binary, corrupted, or have syntax errors.
- **What's unclear:** Should subagent skip silently, or return error list?
- **Recommendation:** Try-catch around file reading. Skip unparseable files, track in errors list, report at end. Don't let one bad file block entire batch.
## Sources
### Primary (HIGH confidence)
- [commands/gsd/analyze-codebase.md](file://./commands/gsd/analyze-codebase.md) - Current entity generation implementation (Phase 4)
- [agents/gsd-codebase-mapper.md](file://./agents/gsd-codebase-mapper.md) - Subagent delegation pattern reference
- [agents/gsd-executor.md](file://./agents/gsd-executor.md) - Task tool usage with model profiles
- [commands/gsd/execute-phase.md](file://./commands/gsd/execute-phase.md) - Wave-based parallel execution pattern
- [hooks/gsd-intel-index.js](file://./hooks/gsd-intel-index.js) - PostToolUse hook that syncs entities to graph
### Secondary (MEDIUM confidence)
- [Claude Code documentation on Task tool](https://docs.anthropic.com/claude/docs/claude-code) - Official docs on subagent spawning
- [Phase 4 research findings](.planning/phases/04-semantic-intelligence/04-RESEARCH.md) - Context window management patterns
### Tertiary (LOW confidence)
- None - all findings based on existing codebase analysis
## Metadata
**Confidence breakdown:**
- Standard stack: HIGH - Task tool is official Claude Code primitive, sql.js already in use
- Architecture: HIGH - Pattern directly mirrors gsd-codebase-mapper.md (proven)
- Pitfalls: HIGH - Based on direct code analysis and understanding of context limits
- Open questions: MEDIUM - Edge cases identifiable but not yet tested at scale
**Research date:** 2026-01-20
**Valid until:** 60 days (stable Claude Code APIs, no fast-moving dependencies)
**Critical insight:** Phase 5 is an architectural refactor, not a feature addition. The goal is context preservation, not new capabilities. Success = same output with less orchestrator context usage.

File diff suppressed because it is too large Load Diff

View File

@@ -1,360 +0,0 @@
# Technical Research: Self-Evolving Codebase Intelligence for GSD
## Strategic Summary
The most mind-blowing approach combines **semantic code indexing** (tree-sitter AST + voyage-code-3 embeddings), **adaptive pattern learning** (conventions extracted and refined through feedback), and **self-improving compliance** (the system gets smarter about YOUR codebase over time). Instead of static indices that go stale, this creates a **living knowledge graph** that understands not just what code exists, but WHY it's structured that way—and enforces that understanding during execution.
**Recommendation:** Approach 3 (Self-Evolving Intelligence) - because "mindblowing" means the system should feel like it genuinely understands your codebase and gets better over time.
## Requirements
- Zero external services (local-first, works offline)
- Sub-second queries (can't slow down Claude's planning/execution)
- Adaptive learning (improves from corrections without manual intervention)
- Works across language ecosystems (not just TypeScript)
- Integrates seamlessly with existing GSD workflow
- Minimal storage overhead (< 50MB for typical project)
---
## Approach 1: Static JSON Indices with Stale Detection
**How it works:** Generate JSON indices during `map-codebase`, store commit hash in metadata, refresh lazily when HEAD moves.
**Libraries/tools:**
- tree-sitter (AST parsing) - `npm install tree-sitter tree-sitter-typescript tree-sitter-python`
- Node.js built-in fs for JSON persistence
- Git hooks via husky or manual `.git/hooks/post-commit`
**Pros:**
- Simple to implement (~500 lines)
- No external dependencies
- Fast queries (JSON parse + lookup)
- Familiar format (developers can hand-edit)
**Cons:**
- Indices go stale between sessions
- No semantic understanding (just structural data)
- Manual schema maintenance
- Doesn't learn from corrections
**Best when:** You want quick wins with minimal complexity
**Complexity:** S
---
## Approach 2: Semantic Search with Local Embeddings
**How it works:** Embed code chunks using voyage-code-3, store in sqlite-vec, enable semantic queries like "find authentication logic" or "code that validates user input."
**Libraries/tools:**
- `voyage-code-3` via API (200M tokens free, $0.06/1M after)
- `sqlite-vec` - pure C, runs anywhere SQLite runs
- `tree-sitter` for intelligent chunking (function-level, not file-level)
- `better-sqlite3` for Node.js bindings
**Architecture:**
```
Code Changes → Tree-sitter chunks → voyage-code-3 embeds → sqlite-vec stores
↓
Query ("auth logic") → embed query → sqlite-vec similarity → ranked results
```
**Pros:**
- Semantic understanding ("find error handling" works even with varied naming)
- Hybrid search possible (FTS5 keywords + vector similarity)
- Local-first with cloud embeddings (best of both)
- Works across languages (voyage-code-3 trained on 80+ languages)
**Cons:**
- Requires API calls for embedding (latency, cost at scale)
- More complex setup
- Still doesn't learn conventions automatically
- Embedding drift as model versions change
**Best when:** You need powerful search across large/unfamiliar codebases
**Complexity:** M
---
## Approach 3: Self-Evolving Codebase Intelligence (The Mind-Blowing One)
**How it works:** Three-layer system that builds understanding incrementally and learns from corrections.
### Layer 1: Structural Index (Tree-sitter AST)
Fast, deterministic extraction of code structure:
- Functions, classes, exports, imports
- File patterns and naming conventions
- Directory structure and module boundaries
### Layer 2: Semantic Memory (Embeddings + Patterns)
Understanding what code DOES, not just what it IS:
- Function-level embeddings via voyage-code-3
- Pattern clusters detected via embedding similarity
- Convention rules extracted from consistent patterns
### Layer 3: Adaptive Learning (The Magic)
The system improves from every interaction:
- When Claude deviates from conventions and you correct it → system learns
- When Claude asks "should this be a service or util?" and you answer → system remembers
- Confidence scores that increase with consistent patterns, decrease with exceptions
**Architecture:**
```
┌─────────────────────────────────────────────────────────────────┐
│ Codebase Intelligence │
├─────────────────────────────────────────────────────────────────┤
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────────────────┐ │
│ │ STRUCTURE │ │ SEMANTICS │ │ LEARNED RULES │ │
│ │ (AST/JSON) │ │ (Embeddings)│ │ (Adaptive/Scored) │ │
│ ├─────────────┤ ├─────────────┤ ├─────────────────────────┤ │
│ │ symbols.json│ │ sqlite-vec │ │ conventions.json │ │
│ │ patterns.json │ + FTS5 │ │ ├─ rule: "services/*" │ │
│ │ structure.json │ │ │ │ confidence: 0.95 │ │
│ │ │ │ │ │ │ examples: [...] │ │
│ │ │ │ │ │ │ exceptions: [...] │ │
│ └─────────────┘ └─────────────┘ └─────────────────────────┘ │
├─────────────────────────────────────────────────────────────────┤
│ FEEDBACK LOOP │
│ ┌──────────┐ ┌──────────┐ ┌──────────────────────────┐ │
│ │ Executor │ → │ Deviation │ → │ User Correction/Approval │ │
│ │ writes │ │ detected │ │ → Update confidence │ │
│ │ code │ │ │ │ → Add to examples │ │
│ └──────────┘ └──────────┘ └──────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
```
**Libraries/tools:**
```bash
# Core parsing
npm install tree-sitter tree-sitter-typescript tree-sitter-python tree-sitter-go tree-sitter-rust
# Vector storage (local)
npm install better-sqlite3
# sqlite-vec compiled extension (downloaded during install)
# Embeddings (API)
npm install voyageai # or call API directly
# Optional: Local embeddings fallback
# ollama pull nomic-embed-text (for offline mode)
```
**The "Mind-Blowing" Features:**
1. **Convention Inference Engine**
- Analyzes existing code to detect patterns automatically
- "I see 15 files matching `src/services/*.service.ts`, all exporting classes with `@Injectable()`"
- Confidence scores: 15 examples = high confidence, 2 examples = tentative
2. **Deviation Detection During Execution**
- Before Claude writes `src/utils/auth.ts`, check conventions
- "This looks like a service (has `@Injectable`, depends on repository). Convention suggests `src/services/auth.service.ts`"
- NOT blocking—advisory with reasoning
3. **Correction Learning**
- User says "no, utils is correct here because X"
- System adds exception: `{pattern: "auth*", location: "utils", reason: "X", confidence: 0.8}`
- Future similar cases consider this exception
4. **Semantic "Why" Queries**
- Claude can ask the index: "Why is UserRepository in `src/data` not `src/repositories`?"
- Index returns: historical context, similar patterns, any recorded exceptions
5. **Auto-Refresh with Minimal Recomputation**
- Git hook triggers on commit
- Only re-index changed files (incremental)
- Only re-embed functions that changed (hash comparison)
- Full re-scan weekly or on major refactors
**Pros:**
- Feels magical ("it knows my codebase")
- Gets smarter over time (adaptive)
- Handles edge cases gracefully (exceptions are learned, not errors)
- Works across languages
- Local-first with optional cloud embeddings
**Cons:**
- Most complex to implement (~2000 lines + iteration)
- Requires careful feedback loop design
- Cold start problem (needs usage to learn)
- More storage (embeddings + history)
**Best when:** You want GSD to feel like a senior engineer who truly knows the codebase
**Complexity:** L
---
## Approach 4: MCP Server Integration (Leverage Existing Tools)
**How it works:** Use existing code-index-mcp or claude-context MCP servers, integrate with GSD's planning/execution.
**Libraries/tools:**
- `code-index-mcp` or `claude-context` MCP server
- GSD adds MCP configuration to project setup
- Planner/executor query MCP tools instead of custom indices
**Pros:**
- Leverage battle-tested implementations
- Active community development
- Already handles tree-sitter, embeddings, incremental updates
- MCP is Claude Code native
**Cons:**
- Less control over index structure
- May not support adaptive learning
- Another dependency to manage
- May not align perfectly with GSD's workflow
**Best when:** You want proven tooling without building from scratch
**Complexity:** M
---
## Comparison
| Aspect | Static JSON | Semantic Search | Self-Evolving | MCP Server |
|--------|-------------|-----------------|---------------|------------|
| Complexity | S | M | L | M |
| Query Speed | Instant | ~100ms | ~150ms | Varies |
| Semantic Understanding | None | Good | Excellent | Good |
| Learns from Corrections | No | No | Yes | No |
| Offline Capable | Yes | Partial | Partial | Depends |
| Cross-language | Manual | Yes | Yes | Yes |
| Maintenance | Manual | Medium | Self-maintaining | External |
| "Wow Factor" | Low | Medium | High | Medium |
---
## Recommendation
**Go with Approach 3: Self-Evolving Intelligence**, implemented in phases:
**Phase 1 (MVP):** Static JSON indices with tree-sitter extraction
- Get basic pattern detection working
- Prove value before adding complexity
**Phase 2 (Semantic):** Add sqlite-vec + voyage-code-3
- Enable "find code that does X" queries
- Hybrid search for maximum flexibility
**Phase 3 (Adaptive):** Add feedback loop and confidence scoring
- Convention rules with confidence
- Learn from corrections
- The "magic" emerges here
**Phase 4 (Polish):** Auto-refresh, git hooks, exception handling
- Incremental updates
- Graceful degradation when offline
---
## Implementation Context
<claude_context>
<chosen_approach>
- name: Self-Evolving Codebase Intelligence (phased)
- libraries:
- tree-sitter + language grammars (parsing)
- better-sqlite3 + sqlite-vec extension (storage)
- voyageai SDK or direct API (embeddings)
- chokidar (file watching, optional)
- install: |
npm install tree-sitter tree-sitter-typescript tree-sitter-python tree-sitter-go
npm install better-sqlite3
# sqlite-vec: download prebuilt from https://github.com/asg017/sqlite-vec/releases
npm install voyageai
</chosen_approach>
<architecture>
- pattern: Three-layer knowledge graph (Structure → Semantics → Learned Rules)
- components:
- CodebaseIndexer: Orchestrates tree-sitter parsing, embedding generation, storage
- StructureExtractor: AST → JSON symbols, patterns, structure
- SemanticMemory: sqlite-vec for embeddings, FTS5 for keywords
- ConventionEngine: Pattern detection, rule inference, confidence scoring
- FeedbackCollector: Captures corrections, updates confidence, logs exceptions
- QueryInterface: Unified API for planner/executor to query knowledge
- data_flow: |
Init: codebase → tree-sitter → symbols.json + patterns.json
→ voyage-code-3 → sqlite-vec
→ ConventionEngine → conventions.json
Query: planner asks "where should auth service go?"
→ QueryInterface checks conventions.json (high confidence rules)
→ Falls back to semantic search if no rule
→ Returns recommendation with reasoning
Feedback: executor writes code → deviation detected → user corrects
→ FeedbackCollector updates rule confidence or adds exception
</architecture>
<files>
- create:
- `.planning/indices/symbols.json` - Extracted code symbols
- `.planning/indices/patterns.json` - Detected architectural patterns
- `.planning/indices/conventions.json` - Learned rules with confidence
- `.planning/indices/codebase.db` - sqlite-vec embeddings + FTS5
- `.planning/indices/meta.json` - Commit hash, last update, stats
- `get-shit-done/lib/indexer/` - Index generation code
- `get-shit-done/lib/query/` - Query interface for planner/executor
- structure: |
.planning/
indices/
symbols.json # {exports, imports, classes, functions}
patterns.json # {detected patterns with locations}
conventions.json # {rules with confidence, examples, exceptions}
codebase.db # sqlite-vec + FTS5 hybrid storage
meta.json # {lastCommit, lastUpdate, stats}
- reference:
- `.planning/codebase/ARCHITECTURE.md` - Use as input for initial pattern detection
- `.planning/codebase/CONVENTIONS.md` - Seed initial convention rules
- `commands/gsd/map-codebase.md` - Extend to also generate indices
</files>
<implementation>
- start_with: StructureExtractor using tree-sitter (symbols.json output)
- order:
1. Tree-sitter parsing → symbols.json (Phase 1)
2. Pattern detection → patterns.json (Phase 1)
3. sqlite-vec setup + embedding generation (Phase 2)
4. ConventionEngine + confidence scoring (Phase 3)
5. FeedbackCollector + correction learning (Phase 3)
6. Git hooks + incremental updates (Phase 4)
7. Query interface integration with planner/executor (Phase 4)
- gotchas:
- Tree-sitter grammars are separate packages per language
- sqlite-vec is a C extension, needs platform-specific binary
- voyage-code-3 has 32K context but 120K token batch limit
- Chunking strategy matters: function-level > file-level for embeddings
- Confidence scores need tuning (start conservative, adjust based on feedback)
- Don't over-index: only public exports, key patterns, not every variable
- testing:
- StructureExtractor: Parse known files, verify symbol extraction
- SemanticMemory: Embed test functions, verify similarity search
- ConventionEngine: Feed patterns, verify rule inference
- FeedbackCollector: Simulate corrections, verify confidence updates
- Integration: Full flow from code change → query → recommendation
</implementation>
</claude_context>
**Next Action:** Start with Phase 1 - build the tree-sitter StructureExtractor that outputs `symbols.json` for a test project. This proves the parsing pipeline works before adding embeddings.
---
## Sources
- [Semantic Code Indexing with AST and Tree-sitter](https://medium.com/@email2dineshkuppan/semantic-code-indexing-with-ast-and-tree-sitter-for-ai-agents-part-1-of-3-eb5237ba687a) - Tree-sitter fundamentals
- [mcp-server-tree-sitter](https://github.com/wrale/mcp-server-tree-sitter) - Reference implementation
- [code-index-mcp](https://github.com/johnhuang316/code-index-mcp) - 7-language tree-sitter integration
- [claude-context (Zilliz)](https://github.com/zilliztech/claude-context) - 40% token reduction with semantic search
- [voyage-code-3 announcement](https://blog.voyageai.com/2024/12/04/voyage-code-3/) - State-of-art code embeddings
- [sqlite-vec](https://github.com/asg017/sqlite-vec) - Local vector search for SQLite
- [Hybrid search with sqlite-vec](https://alexgarcia.xyz/blog/2024/sqlite-vec-hybrid-search/index.html) - FTS5 + vector combination
- [Codebases are uniquely hard to search semantically](https://www.greptile.com/blog/semantic-codebase-search) - Why chunking strategy matters
- [Self-Improving Coding Agent (arxiv)](https://arxiv.org/abs/2504.15228) - Research on adaptive code agents
- [Building Self-Improving AI Agents](https://yoheinakajima.com/better-ways-to-build-self-improving-ai-agents/) - Feedback loop architecture
- [State of AI code quality 2025](https://www.qodo.ai/reports/state-of-ai-code-quality/) - Context-aware AI expectations
- [Git Hooks Guide 2025](https://dev.to/arasosman/git-hooks-for-automated-code-quality-checks-guide-2025-372f) - Modern hook practices
- [pre-commit framework](https://pre-commit.com/) - Hook management