docs(04): create phase 4 plans for semantic intelligence

Phase 4: Semantic Intelligence & Scale
- 3 plans in 2 waves
- Wave 1: 04-01 SQLite graph layer (sql.js WASM)
- Wave 2: 04-02 Rich summary generation, 04-03 Semantic entity generation

Key additions:
- sql.js for zero-native-deps graph database
- @anthropic-ai/sdk for Claude API entity generation
- Recursive CTEs for dependency queries ("what uses X")
- Graph-backed summaries with hotspots

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
Lex Christopherson
2026-01-20 09:39:22 -06:00
parent 3a82930579
commit e3d39c09e6
4 changed files with 1249 additions and 0 deletions

80
.planning/ROADMAP.md Normal file
View File

@@ -0,0 +1,80 @@
# Roadmap: v1.9.0 Codebase Intelligence System
**Goal:** Make GSD feel intelligent and automagical in how it navigates and understands both greenfield and brownfield projects.
**Phases:** 4 (3 complete, 1 remaining)
---
## Current Milestone: v1.9.0
### Phase 1: Foundation & Learning ✓
**Goal:** Establish index schema and incremental learning via PostToolUse hook
**Status:** Complete
**Plans:** 2/2
### Phase 2: Context Injection ✓
**Goal:** Inject codebase awareness into every session via SessionStart hook
**Status:** Complete
**Plans:** 2/2
### Phase 3: Brownfield & Integration ✓
**Goal:** Deep analysis command for existing codebases, workflow integration
**Status:** Complete
**Plans:** 3/3
### Phase 4: Semantic Intelligence & Scale
**Goal:** Transform syntax-only indexing into semantic understanding with graph-based relationships
**Depends on:** Phase 3
**Plans:** 3 plans
Plans:
- [ ] 04-01-PLAN.md — SQLite graph layer with sql.js (Wave 1)
- [ ] 04-02-PLAN.md — Graph-backed rich summary generation (Wave 2)
- [ ] 04-03-PLAN.md — Semantic entity generation via Claude API (Wave 2)
**Wave Structure:**
- Wave 1: 04-01 (SQLite foundation)
- Wave 2: 04-02, 04-03 (parallel - both depend only on 04-01)
**Why this phase:**
- Current system provides "2-3 ls commands worth of information" (Claude's own assessment)
- Missing: what files actually DO, who uses them, blast radius of changes
- Senior engineers at top companies need real intelligence, not file counts
**Delivers:**
- SQLite graph layer (sql.js - zero native deps) for relationship queries
- Entity-based semantic documentation (Claude writes understanding, not just syntax)
- Semantic `/gsd:analyze-codebase` that creates initial entities
- Rich summary generation from accumulated semantic knowledge
**Requirements:**
- INTEL-04: Entity files capture semantic understanding (purpose, what exports do)
- INTEL-05: Relationships queryable ("what uses this file?", "blast radius")
- INTEL-06: `/gsd:analyze-codebase` creates initial entity docs via Claude
- INTEL-07: Summary reflects accumulated semantic knowledge
**Success Criteria:**
1. Claude can answer "what uses src/lib/db.ts?" from SessionStart context
2. Summary includes file purposes, not just file counts
3. Transitive dependency queries work (blast radius)
4. Works at scale (500+ file codebases)
---
## Traceability
| Requirement | Phase | Status |
|-------------|-------|--------|
| INTEL-01 | Phase 1 | ✓ Complete |
| INTEL-02 | Phase 2 | ✓ Complete |
| INTEL-03 | Phase 3 | ✓ Complete |
| INTEL-04 | Phase 4 | Pending |
| INTEL-05 | Phase 4 | Pending |
| INTEL-06 | Phase 4 | Pending |
| INTEL-07 | Phase 4 | Pending |
---
*Created: 2026-01-19*
*Updated: 2026-01-20 — Phase 4 planned (3 plans in 2 waves)*

View File

@@ -0,0 +1,334 @@
---
phase: 04-semantic-intelligence
plan: 01
type: execute
wave: 1
depends_on: []
files_modified:
- hooks/gsd-intel-index.js
- package.json
autonomous: true
must_haves:
truths:
- "Entity files sync to SQLite graph database on write"
- "Graph persists across hook invocations via graph.db file"
- "Wiki-links become edges in the graph"
artifacts:
- path: "hooks/gsd-intel-index.js"
provides: "SQLite graph sync on entity write"
contains: "initSqlJs"
- path: ".planning/intel/graph.db"
provides: "Persistent SQLite database"
key_links:
- from: "hooks/gsd-intel-index.js"
to: ".planning/intel/graph.db"
via: "sql.js export/import"
pattern: "db\\.export\\(\\)"
---
<objective>
Add SQLite graph layer to the codebase intelligence system using sql.js (WASM).
Purpose: Enable relationship queries ("what uses this file?", "blast radius") by storing entity relationships in a queryable graph database.
Output: Modified gsd-intel-index.js with SQLite sync, updated package.json with sql.js dependency.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/04-semantic-intelligence/04-RESEARCH.md
Key implementation details from research:
- sql.js is WASM SQLite (zero native deps)
- Schema: nodes table (JSON body with virtual id), edges table (source/target)
- Must export() and persist after every write
- Async init but sync operations
- Use ON CONFLICT REPLACE for upserts
Existing code to modify:
- hooks/gsd-intel-index.js already has: parseEntityFrontmatter(), extractWikiLinks(), regenerateEntitySummary()
- Entity files trigger regenerateEntitySummary() when written
</context>
<tasks>
<task type="auto">
<name>Task 1: Add sql.js dependency and graph schema</name>
<files>package.json, hooks/gsd-intel-index.js</files>
<action>
1. Add sql.js dependency to package.json:
```json
"dependencies": {
"sql.js": "^1.12.0"
}
```
2. At top of hooks/gsd-intel-index.js, add require and schema constant:
```javascript
const initSqlJs = require('sql.js');
// Graph database schema (simple-graph pattern)
const GRAPH_SCHEMA = `
CREATE TABLE IF NOT EXISTS nodes (
body TEXT,
id TEXT GENERATED ALWAYS AS (json_extract(body, '$.id')) VIRTUAL NOT NULL UNIQUE
);
CREATE INDEX IF NOT EXISTS id_idx ON nodes(id);
CREATE TABLE IF NOT EXISTS edges (
source TEXT NOT NULL,
target TEXT NOT NULL,
relationship TEXT DEFAULT 'depends_on',
UNIQUE(source, target, relationship) ON CONFLICT REPLACE
);
CREATE INDEX IF NOT EXISTS source_idx ON edges(source);
CREATE INDEX IF NOT EXISTS target_idx ON edges(target);
`;
```
Note: Intentionally no FOREIGN KEY constraints - entity A can reference entity B before B is indexed. Orphan edges are acceptable.
</action>
<verify>
- `grep -q "sql.js" package.json` returns 0
- `grep -q "GRAPH_SCHEMA" hooks/gsd-intel-index.js` returns 0
</verify>
<done>package.json has sql.js dependency, gsd-intel-index.js has schema constant</done>
</task>
<task type="auto">
<name>Task 2: Implement graph database helpers</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add helper functions after the GRAPH_SCHEMA constant:
```javascript
// Singleton SQL instance (async init, reuse across calls)
let sqlInstance = null;
/**
* Get or initialize sql.js instance
* Caches the SQL constructor for reuse
*/
async function getSQL() {
if (!sqlInstance) {
sqlInstance = await initSqlJs();
}
return sqlInstance;
}
/**
* Load or create the graph database
* Returns { db, dbPath } for operations and persistence
*/
async function loadGraphDatabase() {
const SQL = await getSQL();
const dbPath = path.join(process.cwd(), '.planning', 'intel', 'graph.db');
let db;
if (fs.existsSync(dbPath)) {
const buffer = fs.readFileSync(dbPath);
db = new SQL.Database(buffer);
} else {
db = new SQL.Database();
db.run(GRAPH_SCHEMA);
}
return { db, dbPath };
}
/**
* Persist database to disk
* Must call after every write operation
*/
function persistDatabase(db, dbPath) {
const data = db.export();
const buffer = Buffer.from(data);
fs.writeFileSync(dbPath, buffer);
}
```
Key design notes:
- getSQL() caches the WASM instance (expensive to init)
- loadGraphDatabase() handles both create and load
- persistDatabase() called after EVERY write (sql.js is in-memory only)
</action>
<verify>`grep -q "loadGraphDatabase" hooks/gsd-intel-index.js` returns 0</verify>
<done>Graph database helper functions exist in hook</done>
</task>
<task type="auto">
<name>Task 3: Sync entity to graph on write</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add syncEntityToGraph() function and integrate with existing entity handling:
```javascript
/**
* Sync entity file to graph database
* Called when an entity .md file is written
*
* @param {string} entityPath - Path to entity file
*/
async function syncEntityToGraph(entityPath) {
const intelDir = path.join(process.cwd(), '.planning', 'intel');
// Opt-in check (same as updateIndex)
if (!fs.existsSync(intelDir)) {
return;
}
try {
const { db, dbPath } = await loadGraphDatabase();
// Read entity file
const content = fs.readFileSync(entityPath, 'utf8');
const entityId = path.basename(entityPath, '.md').toLowerCase();
const frontmatter = parseEntityFrontmatter(content);
const links = extractWikiLinks(content);
// Build node JSON
const nodeBody = JSON.stringify({
id: entityId,
path: frontmatter.path || entityPath,
type: frontmatter.type || 'unknown',
updated: frontmatter.updated || new Date().toISOString().split('T')[0],
status: frontmatter.status || 'active'
});
// Upsert node (ON CONFLICT handled by schema)
db.run(
`INSERT INTO nodes (body) VALUES (?)
ON CONFLICT(id) DO UPDATE SET body = excluded.body`,
[nodeBody]
);
// Delete old edges for this source, insert new ones
db.run('DELETE FROM edges WHERE source = ?', [entityId]);
if (links.length > 0) {
const stmt = db.prepare('INSERT INTO edges (source, target) VALUES (?, ?)');
for (const target of links) {
stmt.run([entityId, target.toLowerCase()]);
}
stmt.free();
}
// Persist to disk (critical - sql.js is in-memory)
persistDatabase(db, dbPath);
db.close();
} catch (e) {
// Silent failure - never block Claude
// Graph sync is best-effort enhancement
}
}
```
Then modify the entity file handling in the stdin handler:
Find this section:
```javascript
// Handle entity file writes - regenerate summary
if (isEntityFile(filePath)) {
regenerateEntitySummary();
process.exit(0);
}
```
Change to:
```javascript
// Handle entity file writes - sync to graph, regenerate summary
if (isEntityFile(filePath)) {
syncEntityToGraph(filePath).then(() => {
regenerateEntitySummary();
process.exit(0);
}).catch(() => {
// Silent failure
process.exit(0);
});
return; // Don't exit synchronously, wait for async
}
```
Note the return statement - we need to wait for async graph sync before exiting.
</action>
<verify>
Run manual test:
```bash
cd /Users/lexchristopherson/Developer/claude-code-resources/get-shit-done
mkdir -p .planning/intel/entities
echo '---
path: /test/example.ts
type: util
updated: 2026-01-20
status: active
---
# example.ts
## Purpose
Test file for graph sync.
## Dependencies
- [[src-lib-db]]
## Used By
TBD
' > .planning/intel/entities/test-example.md
# Simulate hook execution
echo '{"tool_name":"Write","tool_input":{"file_path":".planning/intel/entities/test-example.md"}}' | node hooks/gsd-intel-index.js
# Check graph.db was created
ls -la .planning/intel/graph.db
# Cleanup
rm .planning/intel/entities/test-example.md
rm .planning/intel/graph.db 2>/dev/null
```
</verify>
<done>Entity writes sync to SQLite graph database, graph.db persists across invocations</done>
</task>
</tasks>
<verification>
After all tasks complete:
1. Dependency installed:
```bash
grep -q '"sql.js"' package.json && echo "PASS: sql.js in package.json"
```
2. Schema and helpers exist:
```bash
grep -q "GRAPH_SCHEMA" hooks/gsd-intel-index.js && echo "PASS: Schema defined"
grep -q "loadGraphDatabase" hooks/gsd-intel-index.js && echo "PASS: Helpers exist"
```
3. Graph sync works:
- Create test entity file with [[wiki-link]]
- Simulate Write hook
- Verify graph.db created
- Verify node and edge inserted (use sqlite3 CLI if available, or just check file size > 0)
</verification>
<success_criteria>
- [ ] sql.js added to package.json dependencies
- [ ] GRAPH_SCHEMA constant defines nodes and edges tables
- [ ] loadGraphDatabase() handles create and load
- [ ] persistDatabase() saves after writes
- [ ] syncEntityToGraph() upserts nodes and edges
- [ ] Entity file writes trigger graph sync before summary regeneration
- [ ] Silent failures don't block Claude
</success_criteria>
<output>
After completion, create `.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md`
</output>

View File

@@ -0,0 +1,445 @@
---
phase: 04-semantic-intelligence
plan: 02
type: execute
wave: 2
depends_on: [04-01]
files_modified:
- hooks/gsd-intel-index.js
autonomous: true
must_haves:
truths:
- "Summary includes dependency hotspots queried from SQLite"
- "Summary shows file purposes, not just file counts"
- "Transitive dependents queryable via recursive CTE"
artifacts:
- path: "hooks/gsd-intel-index.js"
provides: "Graph-backed summary generation"
contains: "generateGraphSummary"
- path: ".planning/intel/summary.md"
provides: "Rich semantic summary for context injection"
key_links:
- from: "hooks/gsd-intel-index.js"
to: ".planning/intel/graph.db"
via: "SQL queries for hotspots"
pattern: "SELECT.*FROM edges.*GROUP BY"
---
<objective>
Generate rich summaries from SQLite graph instead of simple file counts.
Purpose: Provide Claude with actionable intelligence - dependency hotspots, file purposes, and relationship awareness at session start.
Output: Updated gsd-intel-index.js with graph-backed summary generation.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/04-semantic-intelligence/04-RESEARCH.md
@.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md
From 04-01: SQLite graph layer with nodes (entity metadata) and edges (wiki-links).
Summary generation requirements (from research):
- Query hotspots: most-depended-on files
- Group by type from node body
- Include file purposes from entity content
- Target < 500 tokens for context injection
</context>
<tasks>
<task type="auto">
<name>Task 1: Add graph query helpers</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add graph query functions after the existing graph helpers (loadGraphDatabase, persistDatabase):
```javascript
/**
* Get dependency hotspots from graph
* Returns top N files by number of dependents
*
* @param {object} db - sql.js database instance
* @param {number} limit - Max results (default 5)
* @returns {Array<{id: string, count: number, path: string, type: string}>}
*/
function getHotspots(db, limit = 5) {
const results = db.exec(`
SELECT
e.target as id,
COUNT(*) as count,
json_extract(n.body, '$.path') as path,
json_extract(n.body, '$.type') as type
FROM edges e
LEFT JOIN nodes n ON e.target = n.id
GROUP BY e.target
ORDER BY count DESC
LIMIT ?
`, [limit]);
if (!results[0]?.values) return [];
return results[0].values.map(([id, count, path, type]) => ({
id,
count,
path: path || id,
type: type || 'unknown'
}));
}
/**
* Get nodes grouped by type
* Returns type -> count mapping
*
* @param {object} db - sql.js database instance
* @returns {Array<{type: string, count: number}>}
*/
function getNodesByType(db) {
const results = db.exec(`
SELECT
json_extract(body, '$.type') as type,
COUNT(*) as count
FROM nodes
GROUP BY type
ORDER BY count DESC
`);
if (!results[0]?.values) return [];
return results[0].values.map(([type, count]) => ({
type: type || 'other',
count
}));
}
/**
* Get all dependents of a file (transitive)
* Uses recursive CTE for graph traversal
*
* @param {object} db - sql.js database instance
* @param {string} entityId - Starting entity
* @param {number} maxDepth - Max recursion depth (default 5)
* @returns {Array<{id: string, depth: number, path: string}>}
*/
function getDependents(db, entityId, maxDepth = 5) {
const results = db.exec(`
WITH RECURSIVE dependents(id, depth) AS (
SELECT ?, 0
UNION
SELECT e.source, d.depth + 1
FROM edges e
JOIN dependents d ON e.target = d.id
WHERE d.depth < ?
)
SELECT DISTINCT
d.id,
d.depth,
json_extract(n.body, '$.path') as path
FROM dependents d
LEFT JOIN nodes n ON d.id = n.id
WHERE d.id != ?
ORDER BY d.depth, d.id
`, [entityId.toLowerCase(), maxDepth, entityId.toLowerCase()]);
if (!results[0]?.values) return [];
return results[0].values.map(([id, depth, path]) => ({
id,
depth,
path: path || id
}));
}
```
Key design notes:
- LEFT JOIN on nodes allows edges to exist even if target node doesn't exist yet
- UNION (not UNION ALL) prevents infinite loops in cyclic graphs
- maxDepth limit prevents runaway queries
- All IDs lowercased for consistency
</action>
<verify>`grep -q "getHotspots" hooks/gsd-intel-index.js && grep -q "getDependents" hooks/gsd-intel-index.js`</verify>
<done>Graph query helpers exist: getHotspots, getNodesByType, getDependents</done>
</task>
<task type="auto">
<name>Task 2: Create graph-backed summary generator</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Add generateGraphSummary() function that queries the graph database:
```javascript
/**
* Generate semantic summary from graph database
* Called when graph.db exists (Phase 4+)
* Falls back to entity-based summary if no graph
*
* Target: < 500 tokens for context injection
*/
async function generateGraphSummary() {
const intelDir = path.join(process.cwd(), '.planning', 'intel');
const dbPath = path.join(intelDir, 'graph.db');
const summaryPath = path.join(intelDir, 'summary.md');
const entitiesDir = path.join(intelDir, 'entities');
// Require graph.db to exist
if (!fs.existsSync(dbPath)) {
return null; // Caller should fall back to entity summary
}
try {
const { db } = await loadGraphDatabase();
const lines = [];
// Header
lines.push('# Codebase Intelligence');
lines.push('');
// File count from nodes
const countResult = db.exec('SELECT COUNT(*) FROM nodes');
const fileCount = countResult[0]?.values[0]?.[0] || 0;
lines.push(`**Indexed entities:** ${fileCount}`);
lines.push(`**Last updated:** ${new Date().toISOString().split('T')[0]}`);
lines.push('');
// Dependency hotspots (most impactful files)
const hotspots = getHotspots(db, 5);
if (hotspots.length > 0) {
lines.push('## Dependency Hotspots');
lines.push('');
lines.push('Files with most dependents (change carefully):');
for (const { path: filePath, count, type } of hotspots) {
const typeLabel = type !== 'unknown' ? ` [${type}]` : '';
lines.push(`1. \`${filePath}\` (${count} dependents)${typeLabel}`);
}
lines.push('');
}
// Group by type
const byType = getNodesByType(db);
if (byType.length > 0) {
lines.push('## Module Types');
lines.push('');
for (const { type, count } of byType) {
const label = type.charAt(0).toUpperCase() + type.slice(1);
lines.push(`- **${label}**: ${count} files`);
}
lines.push('');
}
// Edge count (relationship density)
const edgeResult = db.exec('SELECT COUNT(*) FROM edges');
const edgeCount = edgeResult[0]?.values[0]?.[0] || 0;
if (edgeCount > 0) {
lines.push(`**Relationships tracked:** ${edgeCount}`);
lines.push('');
}
db.close();
// Write summary
const summary = lines.join('\n');
fs.writeFileSync(summaryPath, summary);
return summary;
} catch (e) {
// Graph query failed, return null to fall back
return null;
}
}
```
Key design notes:
- Returns null if graph doesn't exist or query fails (allows fallback)
- Hotspots show files that cause most downstream impact
- Module types provide quick orientation
- Edge count indicates relationship density
- Targets < 500 tokens (no verbose lists)
</action>
<verify>`grep -q "generateGraphSummary" hooks/gsd-intel-index.js`</verify>
<done>generateGraphSummary() function queries graph and writes summary.md</done>
</task>
<task type="auto">
<name>Task 3: Integrate graph summary into regeneration flow</name>
<files>hooks/gsd-intel-index.js</files>
<action>
Modify regenerateEntitySummary() to prefer graph summary when available.
Find the existing regenerateEntitySummary() function and update it:
```javascript
/**
* Regenerate summary.md from all entity files
* Uses graph database if available (Phase 4+), falls back to file-based
*/
async function regenerateEntitySummary() {
const intelDir = path.join(process.cwd(), '.planning', 'intel');
const entitiesDir = path.join(intelDir, 'entities');
const summaryPath = path.join(intelDir, 'summary.md');
const dbPath = path.join(intelDir, 'graph.db');
// Check directories exist
if (!fs.existsSync(entitiesDir)) {
return;
}
// Try graph-based summary first (Phase 4+)
if (fs.existsSync(dbPath)) {
try {
const graphSummary = await generateGraphSummary();
if (graphSummary) {
return; // Graph summary written, done
}
} catch (e) {
// Fall through to file-based summary
}
}
// Fall back to existing file-based entity summary
// (Keep all existing regenerateEntitySummary logic here)
```
The key change: Check for graph.db first, try generateGraphSummary(), only fall back to existing logic if graph unavailable or fails.
Also update the stdin handler to use async regenerateEntitySummary:
Find:
```javascript
if (isEntityFile(filePath)) {
syncEntityToGraph(filePath).then(() => {
regenerateEntitySummary();
process.exit(0);
})
```
Change to:
```javascript
if (isEntityFile(filePath)) {
syncEntityToGraph(filePath).then(async () => {
await regenerateEntitySummary();
process.exit(0);
})
```
Note: regenerateEntitySummary becomes async because it calls generateGraphSummary.
</action>
<verify>
Test the full flow:
```bash
cd /Users/lexchristopherson/Developer/claude-code-resources/get-shit-done
# Create test entities with dependencies
mkdir -p .planning/intel/entities
echo '---
path: /test/db.ts
type: util
updated: 2026-01-20
status: active
---
# db.ts
## Purpose
Database client.
## Dependencies
None
## Used By
TBD
' > .planning/intel/entities/test-db.md
echo '---
path: /test/auth.ts
type: util
updated: 2026-01-20
status: active
---
# auth.ts
## Purpose
Auth utilities.
## Dependencies
- [[test-db]]
## Used By
TBD
' > .planning/intel/entities/test-auth.md
echo '---
path: /test/api.ts
type: api
updated: 2026-01-20
status: active
---
# api.ts
## Purpose
API routes.
## Dependencies
- [[test-db]]
- [[test-auth]]
## Used By
TBD
' > .planning/intel/entities/test-api.md
# Sync all to graph
for f in .planning/intel/entities/test-*.md; do
echo "{\"tool_name\":\"Write\",\"tool_input\":{\"file_path\":\"$f\"}}" | node hooks/gsd-intel-index.js
done
# Check summary.md has graph-based content
cat .planning/intel/summary.md
# Should show:
# - "Dependency Hotspots" section
# - test-db with 2 dependents (auth and api both depend on it)
# Cleanup
rm .planning/intel/entities/test-*.md
rm .planning/intel/graph.db
rm .planning/intel/summary.md
```
</verify>
<done>Summary generation prefers graph when available, falls back to file-based</done>
</task>
</tasks>
<verification>
After all tasks complete:
1. Query helpers exist:
```bash
grep -q "getHotspots" hooks/gsd-intel-index.js && echo "PASS"
grep -q "getDependents" hooks/gsd-intel-index.js && echo "PASS"
```
2. Graph summary generator exists:
```bash
grep -q "generateGraphSummary" hooks/gsd-intel-index.js && echo "PASS"
```
3. Summary prefers graph:
- Create entities with [[wiki-links]]
- Simulate entity writes
- Check summary.md has "Dependency Hotspots" section
- Hotspot counts are accurate
</verification>
<success_criteria>
- [ ] getHotspots() queries top N most-depended files
- [ ] getNodesByType() groups entities by type
- [ ] getDependents() uses recursive CTE for transitive queries
- [ ] generateGraphSummary() produces < 500 token summary
- [ ] regenerateEntitySummary() prefers graph when graph.db exists
- [ ] Falls back gracefully to file-based summary
- [ ] Summary includes dependency hotspots with accurate counts
</success_criteria>
<output>
After completion, create `.planning/phases/04-semantic-intelligence/04-02-SUMMARY.md`
</output>

View File

@@ -0,0 +1,390 @@
---
phase: 04-semantic-intelligence
plan: 03
type: execute
wave: 2
depends_on: [04-01]
files_modified:
- commands/gsd/analyze-codebase.md
- package.json
autonomous: true
must_haves:
truths:
- "Claude creates entity files with semantic understanding via /gsd:analyze-codebase"
- "Entity files include purpose, not just syntax"
- "Batch processing handles 100+ files efficiently"
artifacts:
- path: "commands/gsd/analyze-codebase.md"
provides: "Semantic entity generation via Claude API"
contains: "@anthropic-ai/sdk"
- path: "package.json"
provides: "Anthropic SDK dependency"
contains: "@anthropic-ai/sdk"
key_links:
- from: "commands/gsd/analyze-codebase.md"
to: "Anthropic Messages API"
via: "client.messages.create"
pattern: "messages\\.create"
---
<objective>
Enhance /gsd:analyze-codebase to create semantic entity files using Claude API.
Purpose: Generate entity documentation that captures file PURPOSE (what it does, why it exists), not just syntax (exports/imports). This transforms "2-3 ls commands" of information into genuine semantic understanding.
Output: Updated analyze-codebase.md command with Claude API integration, @anthropic-ai/sdk dependency.
</objective>
<execution_context>
@~/.claude/get-shit-done/workflows/execute-plan.md
@~/.claude/get-shit-done/templates/summary.md
</execution_context>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/phases/04-semantic-intelligence/04-RESEARCH.md
@.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md
From research:
- Use @anthropic-ai/sdk for Claude API calls
- claude-sonnet-4-5-20250929 for entity generation (fast, cost-effective)
- Process files in batches to avoid rate limits
- Entity template format already exists
Current analyze-codebase.md:
- Steps 1-8 for bulk codebase scanning
- Creates index.json, conventions.json, summary.md
- Does NOT create entity files
New requirement:
- After indexing, optionally create entity .md files
- Use Claude to write semantic purpose, not just regex extraction
</context>
<tasks>
<task type="auto">
<name>Task 1: Add Anthropic SDK dependency</name>
<files>package.json</files>
<action>
Add @anthropic-ai/sdk to package.json dependencies:
```json
"dependencies": {
"sql.js": "^1.12.0",
"@anthropic-ai/sdk": "^0.52.0"
}
```
Note: Version 0.52.0+ includes Messages API with proper TypeScript support.
</action>
<verify>`grep -q "@anthropic-ai/sdk" package.json`</verify>
<done>@anthropic-ai/sdk added to package.json</done>
</task>
<task type="auto">
<name>Task 2: Add semantic entity generation to analyze-codebase</name>
<files>commands/gsd/analyze-codebase.md</files>
<action>
Update the analyze-codebase.md command to add entity generation after index creation.
1. Update the objective to mention entity generation:
```markdown
<objective>
Scan codebase to populate .planning/intel/ with file index, conventions, and semantic entity documentation.
Works standalone (without /gsd:new-project) for brownfield codebases. Creates:
- index.json for file index
- conventions.json for naming patterns
- summary.md for context injection
- entities/*.md for semantic file documentation (optional, requires ANTHROPIC_API_KEY)
Output: .planning/intel/index.json, conventions.json, summary.md, entities/*.md
</objective>
```
2. Add allowed-tools: Task (for entity generation subagent)
3. Add new Step 9 after Step 8 (before completion report):
```markdown
## Step 9: Generate semantic entities (optional)
If `ANTHROPIC_API_KEY` environment variable is set, generate semantic entity files.
### 9a: Select key files for entity generation
From the index, select files for entity generation using these criteria:
- Files with 3+ exports (significant modules)
- Files imported by 5+ other files (dependency hotspots)
- Files in key directories: api/, lib/, utils/, services/, models/
- Limit to 50 files maximum per run (avoid excessive API costs)
Skip:
- Test files (*.test.*, *.spec.*)
- Generated files (*.generated.*, *.d.ts)
- Config files (*.config.*)
- Files already with entities in .planning/intel/entities/
### 9b: Create entity directory
```bash
mkdir -p .planning/intel/entities
```
### 9c: Generate entities using Claude API
For each selected file, use the Anthropic SDK to generate entity content:
```javascript
const Anthropic = require('@anthropic-ai/sdk');
const client = new Anthropic(); // Uses ANTHROPIC_API_KEY env var
async function generateEntityContent(filePath, fileContent) {
const response = await client.messages.create({
model: 'claude-sonnet-4-5-20250929',
max_tokens: 1500,
system: `You are a senior engineer documenting a codebase. Create entity documentation following this exact template format. Be concise - focus on PURPOSE and key relationships.
Output ONLY the markdown content, no explanations or commentary.`,
messages: [{
role: 'user',
content: `Create entity documentation for this file.
Path: ${filePath}
Content:
\`\`\`
${fileContent}
\`\`\`
Follow this template EXACTLY:
---
path: ${filePath}
type: [module|component|util|config|test|api|hook|service|model]
updated: ${new Date().toISOString().split('T')[0]}
status: active
---
# [filename]
## Purpose
[1-3 sentences: What does this file do? Why does it exist? What problem does it solve?]
## Exports
[List each export with signature and brief description]
- \`exportName(args): ReturnType\` - What it does
## Dependencies
[Internal deps use wiki-links, external use plain text]
- [[slugified-path]] - Why needed
- external-package - Why needed
## Used By
TBD
## Notes
[Optional: patterns, gotchas, or important context]`
}]
});
return response.content[0].text;
}
```
### 9d: Write entity files
For each generated entity:
1. Create slug from path: `src/lib/db.ts` -> `src-lib-db`
2. Write to `.planning/intel/entities/{slug}.md`
3. The PostToolUse hook will automatically sync to graph.db
### 9e: Process in batches
Process files in batches of 5 with 1 second delay between batches to respect rate limits.
```javascript
async function processEntities(files) {
const batchSize = 5;
for (let i = 0; i < files.length; i += batchSize) {
const batch = files.slice(i, i + batchSize);
await Promise.all(batch.map(async (filePath) => {
const content = fs.readFileSync(filePath, 'utf8');
const entityContent = await generateEntityContent(filePath, content);
const slug = filePath.replace(/^\//, '').replace(/[\/\.]/g, '-').replace(/-[jt]sx?$/, '');
fs.writeFileSync(`.planning/intel/entities/${slug}.md`, entityContent);
}));
if (i + batchSize < files.length) {
await new Promise(r => setTimeout(r, 1000)); // Rate limit
}
}
}
```
```
4. Update Step 10 (completion report) to include entity stats:
```markdown
## Step 10: Report completion
Display summary statistics:
\`\`\`
Codebase Analysis Complete
Files indexed: [N]
Exports found: [N]
Imports found: [N]
Conventions detected:
- Naming: [dominant case] ([percentage]%)
- Directories: [list]
- Patterns: [list]
Entities created: [N] (if ANTHROPIC_API_KEY set)
- Skipped: [N] (already existed or filtered)
Files created:
- .planning/intel/index.json
- .planning/intel/conventions.json
- .planning/intel/summary.md
- .planning/intel/entities/*.md (if API key set)
Next: Intel hooks will continue incremental learning as you code.
\`\`\`
```
5. Update success criteria to include entity generation:
```markdown
<success_criteria>
- [ ] .planning/intel/ directory created
- [ ] All JS/TS files scanned (excluding node_modules, dist, build, .git, vendor, coverage)
- [ ] index.json populated with exports and imports for each file
- [ ] conventions.json has detected patterns (naming, directories, suffixes)
- [ ] summary.md is concise (< 500 tokens)
- [ ] entities/*.md created for key files (if ANTHROPIC_API_KEY set)
- [ ] Statistics reported to user
</success_criteria>
```
</action>
<verify>
```bash
# Check command has entity generation step
grep -q "Generate semantic entities" commands/gsd/analyze-codebase.md && echo "PASS: Step 9 exists"
# Check mentions Anthropic SDK
grep -q "@anthropic-ai/sdk" commands/gsd/analyze-codebase.md && echo "PASS: SDK mentioned"
# Check has batch processing
grep -q "batchSize" commands/gsd/analyze-codebase.md && echo "PASS: Batch processing"
```
</verify>
<done>analyze-codebase.md includes semantic entity generation via Claude API</done>
</task>
<task type="auto">
<name>Task 3: Add fallback messaging for missing API key</name>
<files>commands/gsd/analyze-codebase.md</files>
<action>
Add clear messaging when ANTHROPIC_API_KEY is not set.
In the context section, add:
```markdown
**Entity generation (optional):**
Requires `ANTHROPIC_API_KEY` environment variable. If not set, only index/conventions/summary are created. Set with:
```bash
export ANTHROPIC_API_KEY=sk-ant-...
```
Entity generation costs approximately $0.01-0.02 per file (using claude-sonnet-4-5-20250929).
```
In Step 9, add at the beginning:
```markdown
## Step 9: Generate semantic entities (optional)
**Check for API key:**
```javascript
if (!process.env.ANTHROPIC_API_KEY) {
console.log('Skipping entity generation: ANTHROPIC_API_KEY not set');
console.log('To enable, run: export ANTHROPIC_API_KEY=sk-ant-...');
// Skip to Step 10
}
```
If no API key, skip directly to Step 10 with message:
```
Entity generation skipped (no ANTHROPIC_API_KEY).
To enable semantic entities, set ANTHROPIC_API_KEY and re-run.
```
```
This ensures the command still works without API key (backwards compatible) while clearly explaining how to enable entity generation.
</action>
<verify>
```bash
grep -q "ANTHROPIC_API_KEY not set" commands/gsd/analyze-codebase.md && echo "PASS: Fallback messaging exists"
```
</verify>
<done>Command gracefully handles missing API key with clear instructions</done>
</task>
</tasks>
<verification>
After all tasks complete:
1. SDK dependency added:
```bash
grep -q "@anthropic-ai/sdk" package.json && echo "PASS"
```
2. Command has entity generation:
```bash
grep -q "Step 9" commands/gsd/analyze-codebase.md && echo "PASS"
grep -q "generateEntityContent" commands/gsd/analyze-codebase.md && echo "PASS"
```
3. Fallback works:
```bash
grep -q "ANTHROPIC_API_KEY not set" commands/gsd/analyze-codebase.md && echo "PASS"
```
4. Batch processing:
```bash
grep -q "batchSize" commands/gsd/analyze-codebase.md && echo "PASS"
```
Manual test (requires API key):
```bash
export ANTHROPIC_API_KEY=sk-ant-...
# Run /gsd:analyze-codebase on a test project
# Verify .planning/intel/entities/*.md created with semantic content
```
</verification>
<success_criteria>
- [ ] @anthropic-ai/sdk added to package.json
- [ ] Step 9 added for entity generation
- [ ] File selection criteria documented (3+ exports, 5+ dependents, key dirs)
- [ ] 50 file limit per run to control costs
- [ ] Batch processing with rate limiting (5 files, 1s delay)
- [ ] Entity slug convention documented
- [ ] Graceful fallback when ANTHROPIC_API_KEY not set
- [ ] Cost estimate included (~$0.01-0.02 per file)
- [ ] Updated success criteria includes entity generation
</success_criteria>
<output>
After completion, create `.planning/phases/04-semantic-intelligence/04-03-SUMMARY.md`
</output>