From 93b963c3feb4fd8750c666bfd95e00ac13f702dc Mon Sep 17 00:00:00 2001 From: Lex Christopherson Date: Wed, 21 Jan 2026 10:17:14 -0600 Subject: [PATCH] chore: remove old planning files --- .planning/ROADMAP.md | 135 -- .planning/STATE.md | 109 - .../04-semantic-intelligence/04-01-PLAN.md | 334 --- .../04-semantic-intelligence/04-01-SUMMARY.md | 99 - .../04-semantic-intelligence/04-02-PLAN.md | 465 ----- .../04-semantic-intelligence/04-02-SUMMARY.md | 99 - .../04-semantic-intelligence/04-03-PLAN.md | 376 ---- .../04-semantic-intelligence/04-03-SUMMARY.md | 101 - .../04-semantic-intelligence/04-04-PLAN.md | 250 --- .../04-semantic-intelligence/04-04-SUMMARY.md | 101 - .../04-semantic-intelligence/04-05-SUMMARY.md | 101 - .../05-01-PLAN.md | 164 -- .../05-01-SUMMARY.md | 122 -- .../05-02-PLAN.md | 232 --- .../05-03-PLAN.md | 192 -- .../05-03-SUMMARY.md | 135 -- .../05-04-PLAN.md | 272 --- .../05-RESEARCH.md | 774 ------- ...2025-01-19-codebase-intelligence-system.md | 1794 ----------------- ...25-01-19-static-code-indexing-technical.md | 360 ---- 20 files changed, 6215 deletions(-) delete mode 100644 .planning/ROADMAP.md delete mode 100644 .planning/STATE.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-01-PLAN.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-01-SUMMARY.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-02-PLAN.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-02-SUMMARY.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-03-PLAN.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-03-SUMMARY.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-04-PLAN.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-04-SUMMARY.md delete mode 100644 .planning/phases/04-semantic-intelligence/04-05-SUMMARY.md delete mode 100644 .planning/phases/05-subagent-codebase-analysis/05-01-PLAN.md delete mode 100644 .planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md delete mode 100644 .planning/phases/05-subagent-codebase-analysis/05-02-PLAN.md delete mode 100644 .planning/phases/05-subagent-codebase-analysis/05-03-PLAN.md delete mode 100644 .planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md delete mode 100644 .planning/phases/05-subagent-codebase-analysis/05-04-PLAN.md delete mode 100644 .planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md delete mode 100644 artifacts/research/2025-01-19-codebase-intelligence-system.md delete mode 100644 artifacts/research/2025-01-19-static-code-indexing-technical.md diff --git a/.planning/ROADMAP.md b/.planning/ROADMAP.md deleted file mode 100644 index 3aa598c1f..000000000 --- a/.planning/ROADMAP.md +++ /dev/null @@ -1,135 +0,0 @@ -# Roadmap: v1.9.0 Codebase Intelligence System - -**Goal:** Make GSD feel intelligent and automagical in how it navigates and understands both greenfield and brownfield projects. - -**Phases:** 5 (4 complete) - ---- - -## Current Milestone: v1.9.0 - -### Phase 1: Foundation & Learning ✓ -**Goal:** Establish index schema and incremental learning via PostToolUse hook -**Status:** Complete -**Plans:** 2/2 - -### Phase 2: Context Injection ✓ -**Goal:** Inject codebase awareness into every session via SessionStart hook -**Status:** Complete -**Plans:** 2/2 - -### Phase 3: Brownfield & Integration ✓ -**Goal:** Deep analysis command for existing codebases, workflow integration -**Status:** Complete -**Plans:** 3/3 - -### Phase 4: Semantic Intelligence & Scale ✓ -**Goal:** Transform syntax-only indexing into semantic understanding with graph-based relationships -**Status:** Complete -**Plans:** 5/5 - -Plans: -- [x] 04-01-PLAN.md — SQLite graph layer with sql.js (Wave 1) -- [x] 04-02-PLAN.md — Graph-backed rich summary generation (Wave 2) -- [x] 04-03-PLAN.md — Semantic entity generation via Claude API (Wave 2) -- [x] 04-04-PLAN.md — CLI query interface for getDependents (Wave 3) -- [x] 04-05-PLAN.md — Wire plan-phase.md to inject intel into planner (Wave 3) - -**Wave Structure:** -- Wave 1: 04-01 (SQLite foundation) -- Wave 2: 04-02, 04-03 (parallel - both depend only on 04-01) -- Wave 3: 04-04, 04-05 (parallel - consumption layer) - -**Why this phase:** -- Current system provides "2-3 ls commands worth of information" (Claude's own assessment) -- Missing: what files actually DO, who uses them, blast radius of changes -- Senior engineers at top companies need real intelligence, not file counts - -**Delivers:** -- SQLite graph layer (sql.js - zero native deps) for relationship queries -- Entity-based semantic documentation (Claude writes understanding, not just syntax) -- Semantic `/gsd:analyze-codebase` that creates initial entities -- Rich summary generation from accumulated semantic knowledge -- CLI query interface for "what uses this file?" queries - -**Requirements:** -- INTEL-04: Entity files capture semantic understanding (purpose, what exports do) -- INTEL-05: Relationships queryable ("what uses this file?", "blast radius") -- INTEL-06: `/gsd:analyze-codebase` creates initial entity docs via Claude -- INTEL-07: Summary reflects accumulated semantic knowledge - -**Success Criteria:** -1. Claude can answer "what uses src/lib/db.ts?" from SessionStart context -2. Summary includes file purposes, not just file counts -3. Transitive dependency queries work (blast radius) -4. Works at scale (500+ file codebases) - -### Phase 5: Subagent Codebase Analysis -**Goal:** Prevent context exhaustion on large codebases by delegating analysis to subagents -**Depends on:** Phase 4 -**Status:** Gap closure in progress -**Plans:** 4 plans (2 complete, 2 gap closure) - -Plans: -- [x] 05-01-PLAN.md — Create gsd-entity-generator subagent (Wave 1) -- [x] 05-02-PLAN.md — Refactor Step 9 for subagent delegation (Wave 2) — partial, gap found -- [ ] 05-03-PLAN.md — Create gsd-indexer subagent for Steps 2-3 (Wave 3) — gap closure -- [ ] 05-04-PLAN.md — Refactor Steps 2-3 for subagent delegation (Wave 4) — gap closure - -**Wave Structure:** -- Wave 1: 05-01 (entity generator agent) -- Wave 2: 05-02 (Step 9 refactor) -- Wave 3: 05-03 (indexer agent) — gap closure -- Wave 4: 05-04 (Steps 2-3 refactor + verification) — gap closure - -**Why this phase:** -- Current entity generation loads file contents in orchestrator context -- On large codebases (500+ files), orchestrator exhausts context during file selection and batching -- Subagent delegation gives fresh 200k context for file processing - -**Gap Found (05-VERIFICATION.md):** -- Original scope only addressed Step 9 (entity generation) -- Actual context exhaustion occurs during Steps 2-3 (indexing) -- Orchestrator reads ALL file contents during indexing, not just entity generation -- Need additional gsd-indexer subagent for Steps 2-3 - -**Delivers:** -- `gsd-entity-generator` subagent following gsd-codebase-mapper pattern -- `gsd-indexer` subagent for file reading and export/import extraction -- Refactored `/gsd:analyze-codebase` with full subagent delegation -- Preserved orchestrator context for large codebase analysis - -**Requirements:** -- INTEL-08: Entity generation delegated to subagent (not inline) ✓ -- INTEL-09: Subagent writes entities directly, returns statistics only ✓ -- INTEL-10: Orchestrator passes file paths, not file contents — BLOCKED (needs 05-03, 05-04) -- INTEL-11: Indexing phase delegated to subagent (gap closure) - -**Success Criteria:** -1. Entity generation works via subagent spawn ✓ -2. Indexing works via subagent spawn (gap closure) -3. Orchestrator context preserved (no file contents loaded in orchestrator) -4. Entities correctly formatted and graph.db updated ✓ -5. Works on 500+ file codebases without context exhaustion - ---- - -## Traceability - -| Requirement | Phase | Status | -|-------------|-------|--------| -| INTEL-01 | Phase 1 | ✓ Complete | -| INTEL-02 | Phase 2 | ✓ Complete | -| INTEL-03 | Phase 3 | ✓ Complete | -| INTEL-04 | Phase 4 | ✓ Complete (04-03) | -| INTEL-05 | Phase 4 | ✓ Complete (04-04) | -| INTEL-06 | Phase 4 | ✓ Complete (04-03) | -| INTEL-07 | Phase 4 | ✓ Complete (04-02) | -| INTEL-08 | Phase 5 | ✓ Complete (05-01, 05-02) | -| INTEL-09 | Phase 5 | ✓ Complete (05-01) | -| INTEL-10 | Phase 5 | Gap closure (05-03, 05-04) | -| INTEL-11 | Phase 5 | Gap closure (05-03, 05-04) | - ---- -*Created: 2026-01-19* -*Updated: 2026-01-20 — Phase 5 gap closure plans added (05-03, 05-04)* diff --git a/.planning/STATE.md b/.planning/STATE.md deleted file mode 100644 index 947d8ee57..000000000 --- a/.planning/STATE.md +++ /dev/null @@ -1,109 +0,0 @@ -# Project State - -## Project Reference - -See: .planning/PROJECT.md (updated 2026-01-19) - -**Core value:** Claude understands your codebase structure and conventions before it starts working — automatically -**Current focus:** v1.9.0 Codebase Intelligence System - -## Current Position - -Phase: 5 of 5 (Subagent Codebase Analysis) -Plan: 3 of 4 (gap closure in progress) -Status: In progress -Last activity: 2026-01-21 — Completed 05-03-PLAN.md (gsd-indexer agent definition) - -Progress: [████████░░] 85% - -## Performance Metrics - -**Velocity:** -- Total plans completed: 15 -- Average duration: 2.4 min -- Total execution time: 36 min - -**By Phase:** - -| Phase | Plans | Total | Avg/Plan | -|-------|-------|-------|----------| -| 1. Foundation & Learning | 2/2 | 7 min | 3.5 min | -| 2. Context Injection | 2/2 | 4 min | 2.0 min | -| 3. Brownfield & Integration | 3/3 | 6 min | 2.0 min | -| 4. Semantic Intelligence | 5/5 | 13 min | 2.6 min | -| 5. Subagent Analysis | 3/4 | 6 min | 2.0 min | - -*Updated after each plan completion* - -## Accumulated Context - -### Decisions - -| Decision | Phase | Rationale | -|----------|-------|-----------| -| index.json keyed by absolute path | 01-01 | O(1) lookup for file entries | -| JSON schema with version field | 01-01 | Enables future schema migrations | -| updated=null for initialization | 01-01 | Distinguishes init from update | -| Use heredoc for stdin testing | 01-02 | Pipe chaining has timing issues with async stdin | -| Extract 'default' as export name | 01-02 | Both 'default' and identifier recorded for default exports | -| Read file from disk for Edit tool | 01-02 | Edit only provides old_string/new_string, not full content | -| Regenerate conventions every index update | 02-01 | Detection is fast, avoids staleness issues | -| Skip 'default' in case detection | 02-01 | Keyword, not naming convention indicator | -| Single lowercase words as camelCase | 02-01 | Follows camelCase rules (e.g., 'main', 'app') | -| Use lookup tables for purposes | 02-01 | More maintainable than regex patterns | -| Target < 500 tokens for summary | 02-02 | Minimize context window usage | -| Top 5 directories, top 3 suffixes | 02-02 | Keep output concise | -| Command documents same regex as hook | 03-01 | Consistency between bulk scan and incremental updates | -| generateSummary in intel-index.js | 03-01 | Co-locate all intel generation; regenerate on every update | -| No FK constraints in graph schema | 04-01 | Entities can reference before target indexed | -| Virtual id from JSON body | 04-01 | Flexible node structure with unique constraint | -| Delete-then-insert for edges | 04-01 | Clean replacement removes stale links | -| Singleton WASM instance | 04-01 | Avoids repeated sql.js init overhead | -| 50 file limit per entity run | 04-03 | Prevents context window exhaustion | -| Batches of 10 for Task tool | 04-03 | Balances parallelization with overhead | -| Entity slug: path--segments--file-ext.md | 04-03 | Flat directory with reversible identification | -| LEFT JOIN allows forward references | 04-02 | Edges can exist before target nodes indexed | -| UNION in recursive CTE | 04-02 | Prevents infinite loops in cyclic graphs | -| maxDepth default of 5 | 04-02 | Prevents runaway queries on deep dependencies | -| Query mode read-only | 04-04 | Query actions don't persist to disk, safe operations | -| Default limit 10 for dependents | 04-04 | Prevents huge output for files with many dependents | -| Query routing before Write/Edit | 04-04 | Clean separation between query and indexing modes | -| Intel read in Step 7 with others | 04-05 | Keep all context file reads in one place | -| 2>/dev/null for missing intel | 04-05 | Graceful degradation when summary.md doesn't exist | -| Skip existing entities by default | 05-01 | Prevents overwriting manual edits to entities | -| Absolute path keys in gsd-indexer | 05-03 | Consistent with index.json schema, O(1) lookup | - -### Pending Todos - -- `/gsd:resume-work` decimal phase handling (deferred from v1.8.0) - -### Roadmap Evolution - -- Phase 5 added: Subagent Codebase Analysis -- Phase 5 gaps found: Need gsd-indexer subagent for Steps 2-3 -- Phase 5 gap closure: Plans 05-03, 05-04 created for indexer agent and integration - -### Blockers/Concerns - -- `.planning/` is gitignored in GSD repo - intel files created but not committed (expected for project-local data) -- **RESOLVED:** Phase 5 context bottleneck addressed with gsd-indexer agent (05-03 complete, 05-04 pending) - -## Session Continuity - -Last session: 2026-01-21 -Stopped at: Completed 05-03-PLAN.md (gsd-indexer agent definition) -Resume file: None - -## Phase Progress - -- Phase 1: Foundation & Learning ✓ -- Phase 2: Context Injection ✓ -- Phase 3: Brownfield & Integration ✓ -- Phase 4: Semantic Intelligence & Scale ✓ -- Phase 5: Subagent Analysis — IN PROGRESS (gap closure) - -**Phase 5 status:** -- 05-01: gsd-entity-generator subagent ✓ -- 05-02: analyze-codebase Step 9 refactor ✓ -- 05-03: gsd-indexer subagent ✓ -- 05-04: analyze-codebase Steps 2-3 refactor — PENDING diff --git a/.planning/phases/04-semantic-intelligence/04-01-PLAN.md b/.planning/phases/04-semantic-intelligence/04-01-PLAN.md deleted file mode 100644 index 71617282a..000000000 --- a/.planning/phases/04-semantic-intelligence/04-01-PLAN.md +++ /dev/null @@ -1,334 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 01 -type: execute -wave: 1 -depends_on: [] -files_modified: - - hooks/gsd-intel-index.js - - package.json -autonomous: true - -must_haves: - truths: - - "Entity files sync to SQLite graph database on write" - - "Graph persists across hook invocations via graph.db file" - - "Wiki-links become edges in the graph" - artifacts: - - path: "hooks/gsd-intel-index.js" - provides: "SQLite graph sync on entity write" - contains: "initSqlJs" - - path: ".planning/intel/graph.db" - provides: "Persistent SQLite database" - key_links: - - from: "hooks/gsd-intel-index.js" - to: ".planning/intel/graph.db" - via: "sql.js export/import" - pattern: "db\\.export\\(\\)" ---- - - -Add SQLite graph layer to the codebase intelligence system using sql.js (WASM). - -Purpose: Enable relationship queries ("what uses this file?", "blast radius") by storing entity relationships in a queryable graph database. - -Output: Modified gsd-intel-index.js with SQLite sync, updated package.json with sql.js dependency. - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/phases/04-semantic-intelligence/04-RESEARCH.md - -Key implementation details from research: -- sql.js is WASM SQLite (zero native deps) -- Schema: nodes table (JSON body with virtual id), edges table (source/target) -- Must export() and persist after every write -- Async init but sync operations -- Use ON CONFLICT REPLACE for upserts - -Existing code to modify: -- hooks/gsd-intel-index.js already has: parseEntityFrontmatter(), extractWikiLinks(), regenerateEntitySummary() -- Entity files trigger regenerateEntitySummary() when written - - - - - - Task 1: Add sql.js dependency and graph schema - package.json, hooks/gsd-intel-index.js - -1. Add sql.js dependency to package.json: - ```json - "dependencies": { - "sql.js": "^1.12.0" - } - ``` - -2. At top of hooks/gsd-intel-index.js, add require and schema constant: - ```javascript - const initSqlJs = require('sql.js'); - - // Graph database schema (simple-graph pattern) - const GRAPH_SCHEMA = ` - CREATE TABLE IF NOT EXISTS nodes ( - body TEXT, - id TEXT GENERATED ALWAYS AS (json_extract(body, '$.id')) VIRTUAL NOT NULL UNIQUE - ); - CREATE INDEX IF NOT EXISTS id_idx ON nodes(id); - - CREATE TABLE IF NOT EXISTS edges ( - source TEXT NOT NULL, - target TEXT NOT NULL, - relationship TEXT DEFAULT 'depends_on', - UNIQUE(source, target, relationship) ON CONFLICT REPLACE - ); - CREATE INDEX IF NOT EXISTS source_idx ON edges(source); - CREATE INDEX IF NOT EXISTS target_idx ON edges(target); - `; - ``` - -Note: Intentionally no FOREIGN KEY constraints - entity A can reference entity B before B is indexed. Orphan edges are acceptable. - - - - `grep -q "sql.js" package.json` returns 0 - - `grep -q "GRAPH_SCHEMA" hooks/gsd-intel-index.js` returns 0 - - package.json has sql.js dependency, gsd-intel-index.js has schema constant - - - - Task 2: Implement graph database helpers - hooks/gsd-intel-index.js - -Add helper functions after the GRAPH_SCHEMA constant: - -```javascript -// Singleton SQL instance (async init, reuse across calls) -let sqlInstance = null; - -/** - * Get or initialize sql.js instance - * Caches the SQL constructor for reuse - */ -async function getSQL() { - if (!sqlInstance) { - sqlInstance = await initSqlJs(); - } - return sqlInstance; -} - -/** - * Load or create the graph database - * Returns { db, dbPath } for operations and persistence - */ -async function loadGraphDatabase() { - const SQL = await getSQL(); - const dbPath = path.join(process.cwd(), '.planning', 'intel', 'graph.db'); - - let db; - if (fs.existsSync(dbPath)) { - const buffer = fs.readFileSync(dbPath); - db = new SQL.Database(buffer); - } else { - db = new SQL.Database(); - db.run(GRAPH_SCHEMA); - } - - return { db, dbPath }; -} - -/** - * Persist database to disk - * Must call after every write operation - */ -function persistDatabase(db, dbPath) { - const data = db.export(); - const buffer = Buffer.from(data); - fs.writeFileSync(dbPath, buffer); -} -``` - -Key design notes: -- getSQL() caches the WASM instance (expensive to init) -- loadGraphDatabase() handles both create and load -- persistDatabase() called after EVERY write (sql.js is in-memory only) - - `grep -q "loadGraphDatabase" hooks/gsd-intel-index.js` returns 0 - Graph database helper functions exist in hook - - - - Task 3: Sync entity to graph on write - hooks/gsd-intel-index.js - -Add syncEntityToGraph() function and integrate with existing entity handling: - -```javascript -/** - * Sync entity file to graph database - * Called when an entity .md file is written - * - * @param {string} entityPath - Path to entity file - */ -async function syncEntityToGraph(entityPath) { - const intelDir = path.join(process.cwd(), '.planning', 'intel'); - - // Opt-in check (same as updateIndex) - if (!fs.existsSync(intelDir)) { - return; - } - - try { - const { db, dbPath } = await loadGraphDatabase(); - - // Read entity file - const content = fs.readFileSync(entityPath, 'utf8'); - const entityId = path.basename(entityPath, '.md').toLowerCase(); - const frontmatter = parseEntityFrontmatter(content); - const links = extractWikiLinks(content); - - // Build node JSON - const nodeBody = JSON.stringify({ - id: entityId, - path: frontmatter.path || entityPath, - type: frontmatter.type || 'unknown', - updated: frontmatter.updated || new Date().toISOString().split('T')[0], - status: frontmatter.status || 'active' - }); - - // Upsert node (ON CONFLICT handled by schema) - db.run( - `INSERT INTO nodes (body) VALUES (?) - ON CONFLICT(id) DO UPDATE SET body = excluded.body`, - [nodeBody] - ); - - // Delete old edges for this source, insert new ones - db.run('DELETE FROM edges WHERE source = ?', [entityId]); - - if (links.length > 0) { - const stmt = db.prepare('INSERT INTO edges (source, target) VALUES (?, ?)'); - for (const target of links) { - stmt.run([entityId, target.toLowerCase()]); - } - stmt.free(); - } - - // Persist to disk (critical - sql.js is in-memory) - persistDatabase(db, dbPath); - db.close(); - } catch (e) { - // Silent failure - never block Claude - // Graph sync is best-effort enhancement - } -} -``` - -Then modify the entity file handling in the stdin handler: - -Find this section: -```javascript -// Handle entity file writes - regenerate summary -if (isEntityFile(filePath)) { - regenerateEntitySummary(); - process.exit(0); -} -``` - -Change to: -```javascript -// Handle entity file writes - sync to graph, regenerate summary -if (isEntityFile(filePath)) { - syncEntityToGraph(filePath).then(() => { - regenerateEntitySummary(); - process.exit(0); - }).catch(() => { - // Silent failure - process.exit(0); - }); - return; // Don't exit synchronously, wait for async -} -``` - -Note the return statement - we need to wait for async graph sync before exiting. - - -Run manual test: -```bash -cd /Users/lexchristopherson/Developer/claude-code-resources/get-shit-done -mkdir -p .planning/intel/entities -echo '--- -path: /test/example.ts -type: util -updated: 2026-01-20 -status: active ---- - -# example.ts - -## Purpose -Test file for graph sync. - -## Dependencies -- [[src-lib-db]] - -## Used By -TBD -' > .planning/intel/entities/test-example.md - -# Simulate hook execution -echo '{"tool_name":"Write","tool_input":{"file_path":".planning/intel/entities/test-example.md"}}' | node hooks/gsd-intel-index.js - -# Check graph.db was created -ls -la .planning/intel/graph.db - -# Cleanup -rm .planning/intel/entities/test-example.md -rm .planning/intel/graph.db 2>/dev/null -``` - - Entity writes sync to SQLite graph database, graph.db persists across invocations - - - - - -After all tasks complete: - -1. Dependency installed: - ```bash - grep -q '"sql.js"' package.json && echo "PASS: sql.js in package.json" - ``` - -2. Schema and helpers exist: - ```bash - grep -q "GRAPH_SCHEMA" hooks/gsd-intel-index.js && echo "PASS: Schema defined" - grep -q "loadGraphDatabase" hooks/gsd-intel-index.js && echo "PASS: Helpers exist" - ``` - -3. Graph sync works: - - Create test entity file with [[wiki-link]] - - Simulate Write hook - - Verify graph.db created - - Verify node and edge inserted (use sqlite3 CLI if available, or just check file size > 0) - - - -- [ ] sql.js added to package.json dependencies -- [ ] GRAPH_SCHEMA constant defines nodes and edges tables -- [ ] loadGraphDatabase() handles create and load -- [ ] persistDatabase() saves after writes -- [ ] syncEntityToGraph() upserts nodes and edges -- [ ] Entity file writes trigger graph sync before summary regeneration -- [ ] Silent failures don't block Claude - - - -After completion, create `.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md` - diff --git a/.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md b/.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md deleted file mode 100644 index a62310f39..000000000 --- a/.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md +++ /dev/null @@ -1,99 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 01 -subsystem: database -tags: [sqlite, sql.js, wasm, graph-database, entity-relationships] - -requires: - - phase: 03-brownfield-integration - provides: entity file system and summary regeneration - -provides: - - SQLite graph database for entity relationships - - Node/edge schema for semantic queries - - Automatic graph sync on entity writes - -affects: [04-02 query interface, future blast-radius queries] - -tech-stack: - added: [sql.js ^1.12.0] - patterns: [simple-graph schema, WASM singleton, async-then-sync operations] - -key-files: - created: [] - modified: [hooks/gsd-intel-index.js, package.json] - -key-decisions: - - "No FOREIGN KEY constraints - entities can reference before target indexed" - - "Virtual id column from JSON body for flexible node structure" - - "Delete-then-insert for edge updates (clean replacement)" - - "Singleton WASM instance to avoid repeated init overhead" - -patterns-established: - - "Graph sync before summary regeneration" - - "Silent failure pattern for non-blocking operations" - -duration: 4min -completed: 2026-01-20 ---- - -# Phase 4 Plan 1: SQLite Graph Layer Summary - -**SQLite graph database using sql.js WASM for entity relationship storage and querying** - -## Performance - -- **Duration:** 4 min -- **Started:** 2026-01-20T09:45:00Z -- **Completed:** 2026-01-20T09:53:00Z -- **Tasks:** 3 -- **Files modified:** 2 - -## Accomplishments - -- Added sql.js WASM SQLite dependency for zero-native-dependency graph storage -- Implemented simple-graph schema with nodes (JSON body) and edges (source/target/relationship) -- Graph database helpers for load, persist, and singleton WASM management -- Entity files now sync to graph database on every write, creating nodes and edges from wiki-links - -## Task Commits - -Each task was committed atomically: - -1. **Task 1: Add sql.js dependency and graph schema** - `11ab4a9` (feat) -2. **Task 2: Implement graph database helpers** - `ec038e7` (feat) -3. **Task 3: Sync entity to graph on write** - `a39988e` (feat) - -## Files Created/Modified - -- `package.json` - Added sql.js ^1.12.0 dependency -- `hooks/gsd-intel-index.js` - GRAPH_SCHEMA constant, loadGraphDatabase(), persistDatabase(), getSQL(), syncEntityToGraph() - -## Decisions Made - -- **No FOREIGN KEY constraints:** Entity A can reference entity B before B is indexed. Orphan edges are acceptable and expected in incremental indexing workflows. -- **Virtual id column:** Uses `json_extract(body, '$.id')` for flexible node structure while maintaining unique constraint for upserts. -- **Delete-then-insert for edges:** Clean replacement of all outgoing edges on entity update ensures stale links are removed. -- **Singleton WASM instance:** sql.js WASM init is expensive; caching the SQL constructor avoids repeated overhead across hook invocations within same process. - -## Deviations from Plan - -None - plan executed exactly as written. - -## Issues Encountered - -None. - -## User Setup Required - -None - no external service configuration required. - -## Next Phase Readiness - -- Graph database infrastructure complete -- Ready for query interface implementation (Plan 04-02) -- Schema supports "what depends on X" and "what does X depend on" queries - ---- -*Phase: 04-semantic-intelligence* -*Completed: 2026-01-20* diff --git a/.planning/phases/04-semantic-intelligence/04-02-PLAN.md b/.planning/phases/04-semantic-intelligence/04-02-PLAN.md deleted file mode 100644 index 23074e0dc..000000000 --- a/.planning/phases/04-semantic-intelligence/04-02-PLAN.md +++ /dev/null @@ -1,465 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 02 -type: execute -wave: 2 -depends_on: [04-01] -files_modified: - - hooks/gsd-intel-index.js -autonomous: true - -must_haves: - truths: - - "Summary includes dependency hotspots queried from SQLite" - - "Summary shows file purposes, not just file counts" - - "Transitive dependents queryable via recursive CTE" - - "SessionStart hook injects graph-backed summary into context" - artifacts: - - path: "hooks/gsd-intel-index.js" - provides: "Graph-backed summary generation" - contains: "generateGraphSummary" - - path: ".planning/intel/summary.md" - provides: "Rich semantic summary for context injection" - key_links: - - from: "hooks/gsd-intel-index.js" - to: ".planning/intel/graph.db" - via: "SQL queries for hotspots" - pattern: "SELECT.*FROM edges.*GROUP BY" - - from: "hooks/gsd-intel-session.js" - to: ".planning/intel/summary.md" - via: "fs.readFileSync on startup/resume" - pattern: "readFileSync.*summary\\.md" ---- - - -Generate rich summaries from SQLite graph instead of simple file counts. - -Purpose: Provide Claude with actionable intelligence - dependency hotspots, file purposes, and relationship awareness at session start. - -Output: Updated gsd-intel-index.js with graph-backed summary generation. - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/phases/04-semantic-intelligence/04-RESEARCH.md -@.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md - -From 04-01: SQLite graph layer with nodes (entity metadata) and edges (wiki-links). - -Summary generation requirements (from research): -- Query hotspots: most-depended-on files -- Group by type from node body -- Include file purposes from entity content -- Target < 500 tokens for context injection - -Existing wiring (from Phase 2): -- hooks/gsd-intel-session.js reads summary.md on startup/resume -- Injects content as tag into Claude's context -- This wiring already exists - we verify it works with new graph format - - - - - - Task 1: Add graph query helpers - hooks/gsd-intel-index.js - -Add graph query functions after the existing graph helpers (loadGraphDatabase, persistDatabase): - -```javascript -/** - * Get dependency hotspots from graph - * Returns top N files by number of dependents - * - * @param {object} db - sql.js database instance - * @param {number} limit - Max results (default 5) - * @returns {Array<{id: string, count: number, path: string, type: string}>} - */ -function getHotspots(db, limit = 5) { - const results = db.exec(` - SELECT - e.target as id, - COUNT(*) as count, - json_extract(n.body, '$.path') as path, - json_extract(n.body, '$.type') as type - FROM edges e - LEFT JOIN nodes n ON e.target = n.id - GROUP BY e.target - ORDER BY count DESC - LIMIT ? - `, [limit]); - - if (!results[0]?.values) return []; - - return results[0].values.map(([id, count, path, type]) => ({ - id, - count, - path: path || id, - type: type || 'unknown' - })); -} - -/** - * Get nodes grouped by type - * Returns type -> count mapping - * - * @param {object} db - sql.js database instance - * @returns {Array<{type: string, count: number}>} - */ -function getNodesByType(db) { - const results = db.exec(` - SELECT - json_extract(body, '$.type') as type, - COUNT(*) as count - FROM nodes - GROUP BY type - ORDER BY count DESC - `); - - if (!results[0]?.values) return []; - - return results[0].values.map(([type, count]) => ({ - type: type || 'other', - count - })); -} - -/** - * Get all dependents of a file (transitive) - * Uses recursive CTE for graph traversal - * - * @param {object} db - sql.js database instance - * @param {string} entityId - Starting entity - * @param {number} maxDepth - Max recursion depth (default 5) - * @returns {Array<{id: string, depth: number, path: string}>} - */ -function getDependents(db, entityId, maxDepth = 5) { - const results = db.exec(` - WITH RECURSIVE dependents(id, depth) AS ( - SELECT ?, 0 - UNION - SELECT e.source, d.depth + 1 - FROM edges e - JOIN dependents d ON e.target = d.id - WHERE d.depth < ? - ) - SELECT DISTINCT - d.id, - d.depth, - json_extract(n.body, '$.path') as path - FROM dependents d - LEFT JOIN nodes n ON d.id = n.id - WHERE d.id != ? - ORDER BY d.depth, d.id - `, [entityId.toLowerCase(), maxDepth, entityId.toLowerCase()]); - - if (!results[0]?.values) return []; - - return results[0].values.map(([id, depth, path]) => ({ - id, - depth, - path: path || id - })); -} -``` - -Key design notes: -- LEFT JOIN on nodes allows edges to exist even if target node doesn't exist yet -- UNION (not UNION ALL) prevents infinite loops in cyclic graphs -- maxDepth limit prevents runaway queries -- All IDs lowercased for consistency - - `grep -q "getHotspots" hooks/gsd-intel-index.js && grep -q "getDependents" hooks/gsd-intel-index.js` - Graph query helpers exist: getHotspots, getNodesByType, getDependents - - - - Task 2: Create graph-backed summary generator - hooks/gsd-intel-index.js - -Add generateGraphSummary() function that queries the graph database: - -```javascript -/** - * Generate semantic summary from graph database - * Called when graph.db exists (Phase 4+) - * Falls back to entity-based summary if no graph - * - * Target: < 500 tokens for context injection - */ -async function generateGraphSummary() { - const intelDir = path.join(process.cwd(), '.planning', 'intel'); - const dbPath = path.join(intelDir, 'graph.db'); - const summaryPath = path.join(intelDir, 'summary.md'); - const entitiesDir = path.join(intelDir, 'entities'); - - // Require graph.db to exist - if (!fs.existsSync(dbPath)) { - return null; // Caller should fall back to entity summary - } - - try { - const { db } = await loadGraphDatabase(); - - const lines = []; - - // Header - lines.push('# Codebase Intelligence'); - lines.push(''); - - // File count from nodes - const countResult = db.exec('SELECT COUNT(*) FROM nodes'); - const fileCount = countResult[0]?.values[0]?.[0] || 0; - lines.push(`**Indexed entities:** ${fileCount}`); - lines.push(`**Last updated:** ${new Date().toISOString().split('T')[0]}`); - lines.push(''); - - // Dependency hotspots (most impactful files) - const hotspots = getHotspots(db, 5); - if (hotspots.length > 0) { - lines.push('## Dependency Hotspots'); - lines.push(''); - lines.push('Files with most dependents (change carefully):'); - for (const { path: filePath, count, type } of hotspots) { - const typeLabel = type !== 'unknown' ? ` [${type}]` : ''; - lines.push(`1. \`${filePath}\` (${count} dependents)${typeLabel}`); - } - lines.push(''); - } - - // Group by type - const byType = getNodesByType(db); - if (byType.length > 0) { - lines.push('## Module Types'); - lines.push(''); - for (const { type, count } of byType) { - const label = type.charAt(0).toUpperCase() + type.slice(1); - lines.push(`- **${label}**: ${count} files`); - } - lines.push(''); - } - - // Edge count (relationship density) - const edgeResult = db.exec('SELECT COUNT(*) FROM edges'); - const edgeCount = edgeResult[0]?.values[0]?.[0] || 0; - if (edgeCount > 0) { - lines.push(`**Relationships tracked:** ${edgeCount}`); - lines.push(''); - } - - db.close(); - - // Write summary - const summary = lines.join('\n'); - fs.writeFileSync(summaryPath, summary); - - return summary; - } catch (e) { - // Graph query failed, return null to fall back - return null; - } -} -``` - -Key design notes: -- Returns null if graph doesn't exist or query fails (allows fallback) -- Hotspots show files that cause most downstream impact -- Module types provide quick orientation -- Edge count indicates relationship density -- Targets < 500 tokens (no verbose lists) - - `grep -q "generateGraphSummary" hooks/gsd-intel-index.js` - generateGraphSummary() function queries graph and writes summary.md - - - - Task 3: Integrate graph summary into regeneration flow and verify SessionStart wiring - hooks/gsd-intel-index.js - -Modify regenerateEntitySummary() to prefer graph summary when available. - -Find the existing regenerateEntitySummary() function and update it: - -```javascript -/** - * Regenerate summary.md from all entity files - * Uses graph database if available (Phase 4+), falls back to file-based - */ -async function regenerateEntitySummary() { - const intelDir = path.join(process.cwd(), '.planning', 'intel'); - const entitiesDir = path.join(intelDir, 'entities'); - const summaryPath = path.join(intelDir, 'summary.md'); - const dbPath = path.join(intelDir, 'graph.db'); - - // Check directories exist - if (!fs.existsSync(entitiesDir)) { - return; - } - - // Try graph-based summary first (Phase 4+) - if (fs.existsSync(dbPath)) { - try { - const graphSummary = await generateGraphSummary(); - if (graphSummary) { - return; // Graph summary written, done - } - } catch (e) { - // Fall through to file-based summary - } - } - - // Fall back to existing file-based entity summary - // (Keep all existing regenerateEntitySummary logic here) -``` - -The key change: Check for graph.db first, try generateGraphSummary(), only fall back to existing logic if graph unavailable or fails. - -Also update the stdin handler to use async regenerateEntitySummary: - -Find: -```javascript -if (isEntityFile(filePath)) { - syncEntityToGraph(filePath).then(() => { - regenerateEntitySummary(); - process.exit(0); - }) -``` - -Change to: -```javascript -if (isEntityFile(filePath)) { - syncEntityToGraph(filePath).then(async () => { - await regenerateEntitySummary(); - process.exit(0); - }) -``` - -Note: regenerateEntitySummary becomes async because it calls generateGraphSummary. - - -Test the full flow including SessionStart injection: -```bash -cd /Users/lexchristopherson/Developer/claude-code-resources/get-shit-done - -# Create test entities with dependencies -mkdir -p .planning/intel/entities - -echo '--- -path: /test/db.ts -type: util -updated: 2026-01-20 -status: active ---- -# db.ts -## Purpose -Database client. -## Dependencies -None -## Used By -TBD -' > .planning/intel/entities/test-db.md - -echo '--- -path: /test/auth.ts -type: util -updated: 2026-01-20 -status: active ---- -# auth.ts -## Purpose -Auth utilities. -## Dependencies -- [[test-db]] -## Used By -TBD -' > .planning/intel/entities/test-auth.md - -echo '--- -path: /test/api.ts -type: api -updated: 2026-01-20 -status: active ---- -# api.ts -## Purpose -API routes. -## Dependencies -- [[test-db]] -- [[test-auth]] -## Used By -TBD -' > .planning/intel/entities/test-api.md - -# Sync all to graph -for f in .planning/intel/entities/test-*.md; do - echo "{\"tool_name\":\"Write\",\"tool_input\":{\"file_path\":\"$f\"}}" | node hooks/gsd-intel-index.js -done - -# Check summary.md has graph-based content -cat .planning/intel/summary.md -# Should show: -# - "Dependency Hotspots" section -# - test-db with 2 dependents (auth and api both depend on it) - -# Verify SessionStart hook reads new summary format -echo '{"source":"startup"}' | node hooks/gsd-intel-session.js -# Should output ... with hotspots - -# Cleanup -rm .planning/intel/entities/test-*.md -rm .planning/intel/graph.db -rm .planning/intel/summary.md -``` - - Summary generation prefers graph when available, SessionStart hook confirmed to inject graph-backed summary - - - - - -After all tasks complete: - -1. Query helpers exist: - ```bash - grep -q "getHotspots" hooks/gsd-intel-index.js && echo "PASS" - grep -q "getDependents" hooks/gsd-intel-index.js && echo "PASS" - ``` - -2. Graph summary generator exists: - ```bash - grep -q "generateGraphSummary" hooks/gsd-intel-index.js && echo "PASS" - ``` - -3. Summary prefers graph: - - Create entities with [[wiki-links]] - - Simulate entity writes - - Check summary.md has "Dependency Hotspots" section - - Hotspot counts are accurate - -4. SessionStart wiring verified: - ```bash - # Verify SessionStart reads and injects the new format - echo '{"source":"startup"}' | node hooks/gsd-intel-session.js | grep -q "Dependency Hotspots" && echo "PASS: SessionStart injects graph summary" - ``` - - - -- [ ] getHotspots() queries top N most-depended files -- [ ] getNodesByType() groups entities by type -- [ ] getDependents() uses recursive CTE for transitive queries -- [ ] generateGraphSummary() produces < 500 token summary -- [ ] regenerateEntitySummary() prefers graph when graph.db exists -- [ ] Falls back gracefully to file-based summary -- [ ] Summary includes dependency hotspots with accurate counts -- [ ] SessionStart hook (gsd-intel-session.js) correctly injects graph-backed summary into context - - - -After completion, create `.planning/phases/04-semantic-intelligence/04-02-SUMMARY.md` - diff --git a/.planning/phases/04-semantic-intelligence/04-02-SUMMARY.md b/.planning/phases/04-semantic-intelligence/04-02-SUMMARY.md deleted file mode 100644 index a2e657369..000000000 --- a/.planning/phases/04-semantic-intelligence/04-02-SUMMARY.md +++ /dev/null @@ -1,99 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 02 -subsystem: intel -tags: [sqlite, sql-js, graph-query, recursive-cte, summary-generation] - -requires: - - phase: 04-01 - provides: SQLite graph layer with nodes/edges schema -provides: - - Graph query helpers (getHotspots, getNodesByType, getDependents) - - Graph-backed summary generation (generateGraphSummary) - - Transitive dependent queries via recursive CTE -affects: [04-03, context-injection, session-start] - -tech-stack: - added: [] - patterns: - - Recursive CTE for transitive graph traversal - - Graph-backed summary with hotspot analysis - -key-files: - created: [] - modified: - - hooks/gsd-intel-index.js - -key-decisions: - - "LEFT JOIN allows edges to exist before target nodes indexed" - - "UNION (not UNION ALL) prevents infinite loops in cyclic graphs" - - "maxDepth limit on recursive CTE prevents runaway queries" - -patterns-established: - - "Graph queries return structured arrays with id/path/type" - - "Summary prefers graph when available, falls back to entity-file-based" - -duration: 2min -completed: 2026-01-20 ---- - -# Phase 4 Plan 2: Query Interface Summary - -**Graph-backed summary generation with dependency hotspots, type grouping, and recursive CTE for transitive dependents** - -## Performance - -- **Duration:** 2 min -- **Started:** 2026-01-20T15:56:42Z -- **Completed:** 2026-01-20T15:59:07Z -- **Tasks:** 3 -- **Files modified:** 1 - -## Accomplishments - -- Added graph query helpers (getHotspots, getNodesByType, getDependents) -- Created generateGraphSummary() that queries SQLite for rich semantic summaries -- Integrated graph summary into regeneration flow with entity-file fallback -- Verified SessionStart hook correctly injects graph-backed summary into context - -## Task Commits - -Each task was committed atomically: - -1. **Task 1: Add graph query helpers** - `5824196` (feat) -2. **Task 2: Create graph-backed summary generator** - `3d8cf70` (feat) -3. **Task 3: Integrate graph summary into regeneration flow** - `101bc58` (feat) - -## Files Created/Modified - -- `hooks/gsd-intel-index.js` - Added graph query helpers and summary generator - -## Decisions Made - -- **LEFT JOIN on nodes**: Allows edges to exist even if target node not yet indexed (forward references) -- **UNION vs UNION ALL**: Using UNION in recursive CTE prevents infinite loops in cyclic dependency graphs -- **maxDepth default of 5**: Prevents runaway queries on deeply nested dependencies -- **All entity IDs lowercased**: Ensures consistent matching across queries - -## Deviations from Plan - -None - plan executed exactly as written. - -## Issues Encountered - -None - -## User Setup Required - -None - no external service configuration required. - -## Next Phase Readiness - -- Graph query interface complete -- Summary generation produces < 500 token output (47 words in test) -- SessionStart hook verified working with new graph-backed format -- Ready for 04-03: Entity generation instructions integration - ---- -*Phase: 04-semantic-intelligence* -*Completed: 2026-01-20* diff --git a/.planning/phases/04-semantic-intelligence/04-03-PLAN.md b/.planning/phases/04-semantic-intelligence/04-03-PLAN.md deleted file mode 100644 index 4104e6ce6..000000000 --- a/.planning/phases/04-semantic-intelligence/04-03-PLAN.md +++ /dev/null @@ -1,376 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 03 -type: execute -wave: 2 -depends_on: [04-01] -files_modified: - - commands/gsd/analyze-codebase.md - - package.json -autonomous: true - -must_haves: - truths: - - "Claude creates entity files with semantic understanding via /gsd:analyze-codebase" - - "Entity files include purpose, not just syntax" - - "Batch processing handles 100+ files efficiently" - artifacts: - - path: "commands/gsd/analyze-codebase.md" - provides: "Semantic entity generation instructions for Claude" - contains: "semantic entities" - - path: "package.json" - provides: "Anthropic SDK dependency" - contains: "@anthropic-ai/sdk" - key_links: - - from: "commands/gsd/analyze-codebase.md" - to: "Task tool" - via: "Subagent spawning for entity batch processing" - pattern: "Task.*entity" ---- - - -Enhance /gsd:analyze-codebase to create semantic entity files using Claude. - -Purpose: Generate entity documentation that captures file PURPOSE (what it does, why it exists), not just syntax (exports/imports). This transforms "2-3 ls commands" of information into genuine semantic understanding. - -Output: Updated analyze-codebase.md command with entity generation instructions, @anthropic-ai/sdk dependency (for future direct API use). - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/phases/04-semantic-intelligence/04-RESEARCH.md -@.planning/phases/04-semantic-intelligence/04-01-SUMMARY.md - -From research: -- Use @anthropic-ai/sdk for Claude API calls -- claude-sonnet-4-5-20250929 for entity generation (fast, cost-effective) -- Process files in batches to avoid rate limits -- Entity template format already exists - -Current analyze-codebase.md: -- Steps 1-8 for bulk codebase scanning -- Creates index.json, conventions.json, summary.md -- Does NOT create entity files - -New requirement: -- After indexing, optionally create entity .md files -- Claude (executing the command) reads file content and generates semantic documentation -- No embedded JavaScript in command markdown - Claude IS the executor - -Execution model clarification: -- GSD command .md files contain INSTRUCTIONS for Claude to follow -- Claude reads the markdown and executes the instructions using its tools -- Commands cannot contain executable JavaScript - Claude interprets and acts on the instructions -- For batch processing, Claude uses the Task tool to spawn subagents - - - - - - Task 1: Add Anthropic SDK dependency - package.json - -Add @anthropic-ai/sdk to package.json dependencies: - -```json -"dependencies": { - "sql.js": "^1.12.0", - "@anthropic-ai/sdk": "^0.52.0" -} -``` - -Note: Version 0.52.0+ includes Messages API with proper TypeScript support. This dependency enables future direct API integration (e.g., hooks that call Claude API directly). For the /gsd:analyze-codebase command, Claude itself generates the entity content. - - `grep -q "@anthropic-ai/sdk" package.json` - @anthropic-ai/sdk added to package.json - - - - Task 2: Add semantic entity generation to analyze-codebase - commands/gsd/analyze-codebase.md - -Update the analyze-codebase.md command to add entity generation after index creation. - -1. Update the objective to mention entity generation: -```markdown - -Scan codebase to populate .planning/intel/ with file index, conventions, and semantic entity documentation. - -Works standalone (without /gsd:new-project) for brownfield codebases. Creates: -- index.json for file index -- conventions.json for naming patterns -- summary.md for context injection -- entities/*.md for semantic file documentation (optional) - -Output: .planning/intel/index.json, conventions.json, summary.md, entities/*.md - -``` - -2. Add Task to allowed-tools (for entity generation via subagent): -```yaml -allowed-tools: - - Read - - Bash - - Glob - - Write - - Task -``` - -3. Add new Step 9 after Step 8 (before completion report). This step provides INSTRUCTIONS for Claude to follow: - -```markdown -## Step 9: Generate semantic entities (optional) - -After indexing, generate semantic entity files for key codebase files. - -### 9a: Select key files for entity generation - -From the index, select files for entity generation using these criteria: -- Files with 3+ exports (significant modules) -- Files imported by 5+ other files (dependency hotspots) -- Files in key directories: api/, lib/, utils/, services/, models/ -- Limit to 50 files maximum per run (context management) - -Skip: -- Test files (*.test.*, *.spec.*) -- Generated files (*.generated.*, *.d.ts) -- Config files (*.config.*) -- Files already with entities in .planning/intel/entities/ - -### 9b: Create entity directory - -```bash -mkdir -p .planning/intel/entities -``` - -### 9c: Generate entities in batches - -For efficient processing, use the Task tool to spawn a subagent for batch entity generation. - -**Subagent prompt template:** - -``` -Generate semantic entity documentation for the following files. - -For each file: -1. Read the file content -2. Analyze its purpose, exports, and dependencies -3. Write an entity file to .planning/intel/entities/{slug}.md - -Entity template (use EXACTLY this format): - ---- -path: {file_path} -type: [module|component|util|config|api|hook|service|model] -updated: {today's date YYYY-MM-DD} -status: active ---- - -# {filename} - -## Purpose - -[1-3 sentences: What does this file do? Why does it exist? What problem does it solve?] - -## Exports - -[List each export with signature and brief description] -- `exportName(args): ReturnType` - What it does - -## Dependencies - -[Internal deps use wiki-links, external use plain text] -- [[slugified-internal-path]] - Why needed -- external-package - Why needed - -## Used By - -TBD - -## Notes - -[Optional: patterns, gotchas, or important context] - ---- - -Slug convention: `src/lib/db.ts` -> `src-lib-db` (replace / and . with -, remove extension) - -Files to process: -{list of file paths, max 10 per batch} -``` - -### 9d: Process in batches of 10 - -For codebases with many key files: -1. Split the selected files into batches of 10 -2. Use Task tool for each batch with the prompt template above -3. Wait for each batch to complete before starting the next -4. This prevents context exhaustion and allows progress tracking - -### 9e: Verify entity creation - -After batch processing: -- Count entities created: `ls .planning/intel/entities/*.md | wc -l` -- Verify they have semantic content (Purpose section, not just syntax) -- The PostToolUse hook will automatically sync new entities to graph.db -``` - -4. Update Step 10 (renumber from Step 8) to include entity stats: - -```markdown -## Step 10: Report completion - -Display summary statistics: - -``` -Codebase Analysis Complete - -Files indexed: [N] -Exports found: [N] -Imports found: [N] - -Conventions detected: -- Naming: [dominant case] ([percentage]%) -- Directories: [list] -- Patterns: [list] - -Entities created: [N] (if entity generation ran) -- Skipped: [N] (already existed or filtered) - -Files created: -- .planning/intel/index.json -- .planning/intel/conventions.json -- .planning/intel/summary.md -- .planning/intel/entities/*.md (if entities generated) - -Next: Intel hooks will continue incremental learning as you code. -``` -``` - -5. Update success criteria to include entity generation: - -```markdown - -- [ ] .planning/intel/ directory created -- [ ] All JS/TS files scanned (excluding node_modules, dist, build, .git, vendor, coverage) -- [ ] index.json populated with exports and imports for each file -- [ ] conventions.json has detected patterns (naming, directories, suffixes) -- [ ] summary.md is concise (< 500 tokens) -- [ ] entities/*.md created for key files (if Step 9 executed) -- [ ] Entity files have semantic Purpose sections (not just syntax extraction) -- [ ] Statistics reported to user - -``` - - -```bash -# Check command has entity generation step -grep -q "Generate semantic entities" commands/gsd/analyze-codebase.md && echo "PASS: Step 9 exists" - -# Check mentions Task tool for batching -grep -q "Task tool" commands/gsd/analyze-codebase.md && echo "PASS: Task batching documented" - -# Check entity template is present -grep -q "## Purpose" commands/gsd/analyze-codebase.md && echo "PASS: Entity template included" - -# Check batch size documented -grep -q "batches of 10" commands/gsd/analyze-codebase.md && echo "PASS: Batch processing" -``` - - analyze-codebase.md includes semantic entity generation via Claude + Task tool batching - - - - Task 3: Add context section explaining execution model - commands/gsd/analyze-codebase.md - -Update the context section to explain how entity generation works. - -In the context section, add: - -```markdown -**Entity generation:** -Step 9 generates semantic entity files using Claude's understanding of code purpose. Unlike regex-based extraction (exports/imports), this captures WHY code exists. - -For large codebases (50+ key files), entity generation uses the Task tool to spawn subagents that process files in batches of 10. This: -- Prevents context exhaustion -- Allows progress tracking -- Enables parallel processing - -Entity files are written to `.planning/intel/entities/` and automatically synced to the graph database by the PostToolUse hook. - -**When to skip entity generation:** -- Quick index-only scan: Stop after Step 8 -- Already have entities: Existing entities won't be overwritten -- Small codebase: May not need formal entities -``` - -This clarifies: -1. Claude generates entity content (not embedded JavaScript) -2. Task tool handles batching for large codebases -3. Users can skip Step 9 if they just want the index - - -```bash -grep -q "Task tool to spawn subagents" commands/gsd/analyze-codebase.md && echo "PASS: Execution model explained" -``` - - Command explains entity generation execution model clearly - - - - - -After all tasks complete: - -1. SDK dependency added: - ```bash - grep -q "@anthropic-ai/sdk" package.json && echo "PASS" - ``` - -2. Command has entity generation: - ```bash - grep -q "Step 9" commands/gsd/analyze-codebase.md && echo "PASS" - grep -q "semantic entities" commands/gsd/analyze-codebase.md && echo "PASS" - ``` - -3. Task tool batching: - ```bash - grep -q "Task tool" commands/gsd/analyze-codebase.md && echo "PASS" - grep -q "batches of 10" commands/gsd/analyze-codebase.md && echo "PASS" - ``` - -4. Entity template present: - ```bash - grep -q "## Purpose" commands/gsd/analyze-codebase.md && echo "PASS" - ``` - -Manual test: -```bash -# Run /gsd:analyze-codebase on a test project -# Verify Steps 1-8 produce index.json, conventions.json, summary.md -# If Step 9 runs, verify .planning/intel/entities/*.md created with semantic content -``` - - - -- [ ] @anthropic-ai/sdk added to package.json -- [ ] Step 9 added for entity generation (instructions for Claude, not embedded JS) -- [ ] Task tool documented for batch processing subagents -- [ ] File selection criteria documented (3+ exports, 5+ dependents, key dirs) -- [ ] 50 file limit per run to manage context -- [ ] Batch processing with batches of 10 via Task tool -- [ ] Entity slug convention documented -- [ ] Execution model explained (Claude generates content, not script execution) -- [ ] Updated success criteria includes entity generation - - - -After completion, create `.planning/phases/04-semantic-intelligence/04-03-SUMMARY.md` - diff --git a/.planning/phases/04-semantic-intelligence/04-03-SUMMARY.md b/.planning/phases/04-semantic-intelligence/04-03-SUMMARY.md deleted file mode 100644 index 5484fc309..000000000 --- a/.planning/phases/04-semantic-intelligence/04-03-SUMMARY.md +++ /dev/null @@ -1,101 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 03 -subsystem: intel -tags: [entity-generation, semantic-analysis, task-batching, anthropic-sdk] - -# Dependency graph -requires: - - phase: 04-01 - provides: SQLite graph layer for relationship storage -provides: - - Semantic entity generation instructions in analyze-codebase command - - Task tool batching pattern for 100+ file processing - - Entity template with Purpose-focused documentation -affects: [04-02, future-intel-queries] - -# Tech tracking -tech-stack: - added: ["@anthropic-ai/sdk ^0.52.0"] - patterns: [Task tool batching, entity slug convention] - -key-files: - created: [] - modified: - - commands/gsd/analyze-codebase.md - - package.json - -key-decisions: - - "50 file limit per run to manage context" - - "Batches of 10 files via Task tool for parallelization" - - "Entity slug convention: path--segments--filename-ext.md" - - "Selection criteria: 3+ exports, 5+ dependents, key directories" - -patterns-established: - - "Task tool batching: spawn subagents for batch processing large file sets" - - "Entity template: Purpose > Exports > Dependencies > Used By" - -# Metrics -duration: 2min -completed: 2026-01-20 ---- - -# Phase 4 Plan 3: Entity Generation Instructions Summary - -**Semantic entity generation via Task tool batching in analyze-codebase command** - -## Performance - -- **Duration:** 2 min -- **Started:** 2026-01-20T15:56:34Z -- **Completed:** 2026-01-20T15:58:41Z -- **Tasks:** 3 -- **Files modified:** 2 - -## Accomplishments - -- Added @anthropic-ai/sdk dependency for future API-based entity generation -- Created Step 9 in analyze-codebase for semantic entity file generation -- Documented Task tool batching pattern for 100+ file codebases -- Established entity template with Purpose-focused semantic documentation - -## Task Commits - -Each task was committed atomically: - -1. **Task 1: Add Anthropic SDK dependency** - `8d33ae7` (chore) -2. **Task 2 & 3: Add entity generation + execution model** - `b3db2ff` (feat) - -## Files Created/Modified - -- `package.json` - Added @anthropic-ai/sdk ^0.52.0 dependency -- `commands/gsd/analyze-codebase.md` - Added Step 9 for entity generation, Task tool in allowed-tools, execution model context - -## Decisions Made - -- **50 file limit per run:** Prevents context window exhaustion during entity generation -- **Batches of 10:** Balances parallelization with Task tool overhead -- **Entity slug convention (path--segments--filename-ext.md):** Flat directory structure with reversible file identification -- **Selection criteria priority:** High-export files first, then hub files, then key directories - -## Deviations from Plan - -None - plan executed exactly as written. - -## Issues Encountered - -None. - -## User Setup Required - -None - no external service configuration required. - -## Next Phase Readiness - -- Entity generation instructions complete in analyze-codebase command -- Ready for 04-02 (Query Interface) to query entity relationships -- Graph layer from 04-01 available for entity relationship storage - ---- -*Phase: 04-semantic-intelligence* -*Completed: 2026-01-20* diff --git a/.planning/phases/04-semantic-intelligence/04-04-PLAN.md b/.planning/phases/04-semantic-intelligence/04-04-PLAN.md deleted file mode 100644 index 413534b18..000000000 --- a/.planning/phases/04-semantic-intelligence/04-04-PLAN.md +++ /dev/null @@ -1,250 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 04 -type: execute -wave: 1 -depends_on: [] -files_modified: - - hooks/gsd-intel-index.js -autonomous: true -gap_closure: true - -must_haves: - truths: - - "Claude can query 'what uses src/lib/db.ts?' via CLI" - - "getDependents() is callable through stdin query action" - - "Query results return as JSON to stdout" - artifacts: - - path: "hooks/gsd-intel-index.js" - provides: "CLI query interface for graph database" - contains: "action.*query" - key_links: - - from: "stdin handler" - to: "getDependents()" - via: "query action routing" - pattern: "action.*query.*getDependents" ---- - - -Expose the orphaned getDependents() function via CLI query interface so Claude can answer "what uses this file?" questions during sessions. - -Purpose: Close the verification gap - INTEL-05 is blocked because getDependents() exists but has no interface. The infrastructure is complete; this adds the "last mile" wiring. - -Output: Modified gsd-intel-index.js that accepts query actions via stdin and returns JSON results to stdout. - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/STATE.md - -# Gap source -@.planning/phases/04-semantic-intelligence/04-VERIFICATION.md - -# Target file (contains getDependents at line 125) -@hooks/gsd-intel-index.js - - - - - - Task 1: Add query action routing to stdin handler - hooks/gsd-intel-index.js - -Modify the stdin 'end' handler (starting at line 930) to detect and route query actions before the existing Write/Edit handling. - -Add this routing logic after `const data = JSON.parse(input);` (line 932): - -```javascript -// Handle query actions (graph queries) -if (data.action === 'query') { - handleQuery(data).then(result => { - console.log(JSON.stringify(result)); - process.exit(0); - }).catch(err => { - console.log(JSON.stringify({ error: err.message })); - process.exit(1); - }); - return; // Don't fall through to Write/Edit handling -} -``` - -Then add the `handleQuery` async function before the stdin handler (around line 925): - -```javascript -/** - * Handle CLI query actions - * Routes to appropriate graph query function based on query type - * - * @param {Object} data - Query action data - * @param {string} data.action - Must be 'query' - * @param {string} data.type - Query type: 'dependents' | 'hotspots' - * @param {string} [data.target] - Entity ID for dependents query (e.g., 'src-lib-db') - * @param {number} [data.limit] - Max results (default: 10 for dependents, 5 for hotspots) - * @param {number} [data.maxDepth] - Max traversal depth for dependents (default: 5) - * @returns {Promise} Query results - */ -async function handleQuery(data) { - const { db, dbPath } = await loadGraphDatabase(); - - try { - switch (data.type) { - case 'dependents': { - if (!data.target) { - return { error: 'target is required for dependents query' }; - } - const results = getDependents(db, data.target, data.maxDepth || 5); - const limited = data.limit ? results.slice(0, data.limit) : results.slice(0, 10); - return { - query: 'dependents', - target: data.target, - count: results.length, - results: limited - }; - } - - case 'hotspots': { - const results = getHotspots(db, data.limit || 5); - return { - query: 'hotspots', - count: results.length, - results - }; - } - - default: - return { error: `Unknown query type: ${data.type}. Valid types: dependents, hotspots` }; - } - } finally { - db.close(); - } -} -``` - -Key implementation notes: -- Query mode does NOT persist to disk (read-only queries) -- db.close() in finally block ensures cleanup -- Default limit of 10 prevents huge output for files with many dependents -- Output goes to stdout as JSON for Claude to parse - - -Test query interface with heredoc: - -```bash -# Test dependents query (should return empty or results depending on graph state) -node hooks/gsd-intel-index.js <<'EOF' -{"action":"query","type":"dependents","target":"src-lib-db"} -EOF - -# Test hotspots query -node hooks/gsd-intel-index.js <<'EOF' -{"action":"query","type":"hotspots","limit":3} -EOF - -# Test error handling (missing target) -node hooks/gsd-intel-index.js <<'EOF' -{"action":"query","type":"dependents"} -EOF - -# Test unknown query type -node hooks/gsd-intel-index.js <<'EOF' -{"action":"query","type":"invalid"} -EOF -``` - -All commands should output valid JSON to stdout. - - -Query actions route to handleQuery(), which returns JSON results via stdout. Claude can invoke `echo '{"action":"query",...}' | node hooks/gsd-intel-index.js` to query the graph. - - - - - Task 2: Add usage documentation as code comment - hooks/gsd-intel-index.js - -Add a documentation comment block at the top of the file (after line 4) explaining the CLI query interface: - -```javascript -/** - * CLI Query Interface - * - * In addition to PostToolUse indexing (Write/Edit actions), this hook supports - * direct graph queries via stdin: - * - * Query dependents (what uses this file?): - * echo '{"action":"query","type":"dependents","target":"src-lib-db"}' | node hooks/gsd-intel-index.js - * - * Query hotspots (most depended-on files): - * echo '{"action":"query","type":"hotspots","limit":5}' | node hooks/gsd-intel-index.js - * - * Options: - * - target: Entity ID (required for dependents, e.g., 'src-lib-db' for src/lib/db.ts) - * - limit: Max results (default: 10 for dependents, 5 for hotspots) - * - maxDepth: Traversal depth for dependents (default: 5) - * - * Output: JSON to stdout with query results - */ -``` - -This makes the CLI interface discoverable for future Claude sessions. - - -Read first 30 lines of file to confirm documentation is present: - -```bash -head -30 hooks/gsd-intel-index.js | grep -q "CLI Query Interface" && echo "Documentation added" -``` - - -CLI query interface is documented in the file header for discoverability. - - - - - - -After both tasks complete, verify the gap is closed: - -1. **Query interface works:** -```bash -# Create test graph.db if needed (empty is fine) -mkdir -p .planning/intel -node hooks/gsd-intel-index.js <<'EOF' -{"action":"query","type":"hotspots"} -EOF -# Should output: {"query":"hotspots","count":0,"results":[]} -``` - -2. **getDependents is no longer orphaned:** -```bash -grep -n "getDependents" hooks/gsd-intel-index.js -# Should show: definition (line ~125) AND call in handleQuery -``` - -3. **Error handling works:** -```bash -node hooks/gsd-intel-index.js <<'EOF' -{"action":"query","type":"dependents"} -EOF -# Should output: {"error":"target is required for dependents query"} -``` - - - -- [ ] `handleQuery()` function added and routes query actions -- [ ] getDependents() called from handleQuery() (no longer orphaned) -- [ ] Query results output as JSON to stdout -- [ ] CLI documentation added to file header -- [ ] Error handling returns JSON error objects -- [ ] INTEL-05 requirement unblocked: Claude can query "what uses this file?" - - - -After completion, create `.planning/phases/04-semantic-intelligence/04-04-SUMMARY.md` - diff --git a/.planning/phases/04-semantic-intelligence/04-04-SUMMARY.md b/.planning/phases/04-semantic-intelligence/04-04-SUMMARY.md deleted file mode 100644 index 014998b37..000000000 --- a/.planning/phases/04-semantic-intelligence/04-04-SUMMARY.md +++ /dev/null @@ -1,101 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 04 -subsystem: codebase-intelligence -tags: [sql.js, graph-queries, cli-interface, dependency-analysis] - -# Dependency graph -requires: - - phase: 04-01 - provides: SQLite graph database with nodes and edges tables - - phase: 04-02 - provides: getDependents() and getHotspots() query functions -provides: - - CLI query interface via stdin for graph database queries - - handleQuery() routing function for dependents and hotspots queries - - JSON output to stdout for Claude consumption -affects: [codebase-analysis, dependency-tracking, refactoring-safety] - -# Tech tracking -tech-stack: - added: [] - patterns: [stdin-json-routing, query-action-pattern] - -key-files: - created: [] - modified: [hooks/gsd-intel-index.js] - -key-decisions: - - "Query mode does not persist to disk (read-only operations)" - - "Default limit of 10 for dependents prevents huge output" - - "Query action routing occurs before Write/Edit handling in stdin handler" - -patterns-established: - - "Query action pattern: {action: 'query', type: 'dependents'|'hotspots', ...options}" - - "JSON error objects returned to stdout for graceful error handling" - -# Metrics -duration: 4min -completed: 2026-01-20 ---- - -# Phase 04 Plan 04: CLI Query Interface Summary - -**CLI query interface exposes getDependents() and getHotspots() via stdin, enabling Claude to answer "what uses this file?" questions** - -## Performance - -- **Duration:** 4 min -- **Started:** 2026-01-20T16:10:00Z -- **Completed:** 2026-01-20T16:14:00Z -- **Tasks:** 2 -- **Files modified:** 1 - -## Accomplishments -- getDependents() function now accessible via CLI (no longer orphaned) -- Query interface accepts JSON actions via stdin and returns results to stdout -- Support for both dependents queries (transitive "what uses this?") and hotspots queries (most depended-on files) -- Complete error handling for missing parameters and unknown query types - -## Task Commits - -Each task was committed atomically: - -1. **Task 1: Add query action routing to stdin handler** - `f46327a` (feat) -2. **Task 2: Add usage documentation as code comment** - `791265a` (docs) - -## Files Created/Modified -- `hooks/gsd-intel-index.js` - Added handleQuery() function and stdin routing for query actions, plus CLI usage documentation - -## Decisions Made - -**Query mode read-only:** Query actions do not persist to disk. The database is opened, queried, and closed without writes. This keeps query operations safe and lightweight. - -**Default limits:** Dependents queries default to 10 results, hotspots to 5. This prevents overwhelming output when a file has many dependents. Clients can override via `limit` parameter. - -**Routing priority:** Query actions are handled before Write/Edit tool processing in the stdin handler. This ensures clean separation between query mode and indexing mode. - -## Deviations from Plan - -None - plan executed exactly as written. - -## Issues Encountered - -None - -## User Setup Required - -None - no external service configuration required. - -## Next Phase Readiness - -**Gap closure complete.** INTEL-05 verification is now unblocked: -- getDependents() is callable via CLI query interface -- Graph database queries work end-to-end (stdin → query → stdout) -- Claude can query dependency information during sessions - -The semantic intelligence system is complete and ready for production use. - ---- -*Phase: 04-semantic-intelligence* -*Completed: 2026-01-20* diff --git a/.planning/phases/04-semantic-intelligence/04-05-SUMMARY.md b/.planning/phases/04-semantic-intelligence/04-05-SUMMARY.md deleted file mode 100644 index 54abfba2c..000000000 --- a/.planning/phases/04-semantic-intelligence/04-05-SUMMARY.md +++ /dev/null @@ -1,101 +0,0 @@ ---- -phase: 04-semantic-intelligence -plan: 05 -subsystem: codebase-intelligence -tags: [planner-integration, context-injection, dependency-awareness] - -# Dependency graph -requires: - - phase: 04-02 - provides: Query functions and summary.md generation -provides: - - Intel injection into planner prompt via {intel_content} placeholder - - Planner receives dependency hotspots and module types at planning time -affects: [planning, context-assembly, plan-quality] - -# Tech tracking -tech-stack: - added: [] - patterns: [context-file-injection, graceful-degradation] - -key-files: - created: [] - modified: [commands/gsd/plan-phase.md] - -key-decisions: - - "Read intel at Step 7 alongside other context files" - - "Use 2>/dev/null for graceful degradation when summary.md doesn't exist" - - "Add intel section after gap closure section in planning_context" - -patterns-established: - - "Context file injection: read into variable in Step 7, inject via placeholder in Step 8" - - "Optional context: empty string if file doesn't exist, planner handles gracefully" - -# Metrics -duration: 1min -completed: 2026-01-20 ---- - -# Phase 04 Plan 05: Intel Injection into Planner Summary - -**Wired plan-phase.md to read and inject .planning/intel/summary.md into planner prompt, giving planners awareness of dependency hotspots and module composition** - -## Performance - -- **Duration:** 1 min -- **Started:** 2026-01-20T16:23:37Z -- **Completed:** 2026-01-20T16:24:14Z -- **Tasks:** 2 -- **Files modified:** 1 - -## Accomplishments -- Planner now receives codebase intelligence (dependency hotspots, module types, relationship counts) -- Phase 4 infrastructure is fully connected: index → graph → summary → planner -- Graceful degradation: planners work fine when intel doesn't exist yet - -## Task Commits - -Each task was committed atomically: - -1. **Task 1 & 2: Add intel read and injection** - `61d1e91` (feat) - -## Files Created/Modified -- `commands/gsd/plan-phase.md` - Added INTEL_CONTENT read in Step 7, {intel_content} placeholder in Step 8 planning_context - -## Decisions Made - -**Read location:** Added intel read in Step 7 alongside other context files (STATE, ROADMAP, RESEARCH, etc.) rather than a new step. Keeps all context file reads in one place. - -**Template location:** Added intel section after gap closure section and before ``. This puts codebase-level context after phase-specific context. - -**Empty string fallback:** Used `2>/dev/null` so INTEL_CONTENT is empty string when summary.md doesn't exist. Planner template handles this gracefully. - -## Deviations from Plan - -None - plan executed exactly as written. - -## Issues Encountered - -None - -## User Setup Required - -None - no external service configuration required. - -## Next Phase Readiness - -**Last mile wiring complete.** The codebase intelligence system is now fully operational: -- `gsd-intel-index` hook captures file changes and builds graph database -- Summary generation creates human-readable intel from graph queries -- Planner receives intel automatically when planning phases - -When a project has been indexed via `/gsd:scan-codebase` or through hook triggers, the planner will see: -- Dependency hotspots (files with most dependents - change carefully) -- Module type breakdown (utils, APIs, components, etc.) -- Total relationship count - -This closes the Phase 4 objective: Claude understands codebase structure and conventions before it starts working. - ---- -*Phase: 04-semantic-intelligence* -*Completed: 2026-01-20* diff --git a/.planning/phases/05-subagent-codebase-analysis/05-01-PLAN.md b/.planning/phases/05-subagent-codebase-analysis/05-01-PLAN.md deleted file mode 100644 index 68fa465bc..000000000 --- a/.planning/phases/05-subagent-codebase-analysis/05-01-PLAN.md +++ /dev/null @@ -1,164 +0,0 @@ ---- -phase: 05-subagent-codebase-analysis -plan: 01 -type: execute -wave: 1 -depends_on: [] -files_modified: - - agents/gsd-entity-generator.md -autonomous: true - -must_haves: - truths: - - "Entity generator agent definition exists following gsd-codebase-mapper pattern" - - "Agent reads files, generates semantic entities, writes directly to disk" - - "Agent returns statistics only (not entity contents)" - artifacts: - - path: "agents/gsd-entity-generator.md" - provides: "Subagent definition for semantic entity generation" - contains: "gsd-entity-generator" - key_links: - - from: "agents/gsd-entity-generator.md" - to: ".planning/intel/entities/" - via: "Write tool calls" - pattern: "Write.*entities" ---- - - -Create gsd-entity-generator subagent definition. - -Purpose: Define the subagent that generates semantic entity documentation for codebase files. This agent is spawned by `/gsd:analyze-codebase` with a list of file paths, reads each file, creates entity markdown, and writes directly to `.planning/intel/entities/`. - -Output: `agents/gsd-entity-generator.md` - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/STATE.md -@.planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md -@agents/gsd-codebase-mapper.md - - - - - - Task 1: Create gsd-entity-generator agent definition - agents/gsd-entity-generator.md - -Create agent definition following gsd-codebase-mapper.md structure: - -**Frontmatter:** -```yaml ---- -name: gsd-entity-generator -description: Generates semantic entity documentation for codebase files. Spawned by analyze-codebase with file list. Writes entities directly to disk. -tools: Read, Write, Bash -color: cyan ---- -``` - -**Role section:** -- Spawned by `/gsd:analyze-codebase` with file paths -- Reads source files, analyzes purpose/exports/dependencies -- Writes entity markdown to `.planning/intel/entities/{slug}.md` -- Returns statistics only (NOT entity contents) - -**Process sections:** - -1. `parse_file_list` - Extract file paths from prompt - - Expect: total count, output directory, slug convention, template, file list - -2. `process_each_file` - For each file path: - - Read file using Read tool - - Analyze: purpose (why exists), exports (signatures), dependencies (internal [[wiki-links]], external plain text), module type - - Generate slug: `src/lib/db.ts` -> `src-lib-db` (replace / and . with -, remove extension, lowercase) - - Write entity to `.planning/intel/entities/{slug}.md` - - Track statistics (created, skipped, errors) - -3. `return_statistics` - Return ONLY: - ``` - ## ENTITY GENERATION COMPLETE - - **Files processed:** {N} - **Entities created:** {M} - **Already existed:** {K} - **Errors:** {E} - - Entities written to: .planning/intel/entities/ - ``` - -**Entity template section:** -Include the full entity template from 05-RESEARCH.md (frontmatter with path/type/updated/status, Purpose, Exports, Dependencies with [[wiki-links]], Used By = TBD, Notes optional). - -**Type heuristics table:** -| Type | Indicators | -|------|-----------| -| api | api/, routes/, endpoints/, route handlers | -| component | components/, React/Vue exports | -| util | utils/, lib/, helpers/ | -| config | config/, *.config.* | -| hook | hooks/, use* functions | -| service | services/ | -| model | models/, types/ | -| test | *.test.*, *.spec.* | -| module | default | - -**Wiki-link rules:** -- Internal (starts with `.` or `@/`): convert to slug, wrap in [[brackets]] -- External (package name): plain text, no brackets - -**Critical rules:** -- WRITE entities directly (never return contents) -- PostToolUse hook syncs to graph.db automatically -- Use EXACT template format (hook parses frontmatter + [[links]]) - -**Success criteria checklist** (from research): -- All file paths processed -- Each entity written to correct path -- Frontmatter is valid YAML -- Purpose section is substantive (not "exports X") -- Internal deps use [[wiki-links]] -- Statistics returned (not entity contents) - - -File exists and contains: -- Frontmatter with name, description, tools, color -- Role section explaining spawn context -- Process steps (parse, process, return) -- Entity template -- Type heuristics -- Wiki-link rules -- Critical rules matching gsd-codebase-mapper pattern - - -`agents/gsd-entity-generator.md` exists with complete agent definition following gsd-codebase-mapper pattern. Agent is ready to be spawned by analyze-codebase. - - - - - - -- [ ] File exists at `agents/gsd-entity-generator.md` -- [ ] Frontmatter valid YAML -- [ ] Role section explains subagent purpose -- [ ] Process has 3 steps: parse, process, return -- [ ] Entity template included in full -- [ ] Type heuristics table present -- [ ] Wiki-link rules specified -- [ ] Critical rules section matches gsd-codebase-mapper style -- [ ] Returns statistics only (not entity contents) - - - -gsd-entity-generator agent definition complete. Agent can be spawned with file paths and will generate semantic entities, writing directly to disk. - - - -After completion, create `.planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md` - diff --git a/.planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md b/.planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md deleted file mode 100644 index 32842dfaa..000000000 --- a/.planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md +++ /dev/null @@ -1,122 +0,0 @@ ---- -phase: 05-subagent-codebase-analysis -plan: 01 -title: gsd-entity-generator Agent Definition -subsystem: agents -tags: [subagent, entity-generation, semantic-analysis] - -dependency-graph: - requires: - - 04-03 (entity generation instructions) - - gsd-codebase-mapper.md (pattern reference) - provides: - - gsd-entity-generator subagent definition - - Entity template specification - - Type heuristics table - affects: - - 05-02 (analyze-codebase integration) - - Future entity generation workflows - -tech-stack: - added: [] - patterns: - - Subagent direct-write pattern - - Statistics-only return pattern - - Wiki-link dependency notation - -key-files: - created: - - agents/gsd-entity-generator.md - modified: [] - -decisions: - - id: skip-existing-entities - choice: Check for existing entity file before writing - rationale: Prevents overwriting manual edits to entities - -metrics: - duration: 1 min 13 sec - completed: 2026-01-20 ---- - -# Phase 05 Plan 01: gsd-entity-generator Agent Definition Summary - -**One-liner:** Subagent definition for semantic entity generation with direct disk writes and statistics-only returns. - -## What Was Built - -Created `agents/gsd-entity-generator.md` following the established `gsd-codebase-mapper.md` pattern: - -**Agent structure:** -- Frontmatter with name, description, tools (Read, Write, Bash), color -- Role section explaining spawn context from `/gsd:analyze-codebase` -- `` section explaining how entities feed the intelligence system -- 3-step process: parse_file_list, process_each_file, return_statistics - -**Entity template specification:** -- YAML frontmatter: path, type, updated, status -- Sections: Purpose, Exports, Dependencies, Used By, Notes -- Internal dependencies use [[wiki-links]] for graph edge creation -- External dependencies as plain text - -**Type heuristics table:** -| Type | Indicators | -|------|-----------| -| api | api/, routes/, endpoints/ | -| component | components/, React/Vue exports | -| util | utils/, lib/, helpers/ | -| config | config/, *.config.* | -| hook | hooks/, use* functions | -| service | services/ | -| model | models/, types/ | -| test | *.test.*, *.spec.* | -| module | default | - -**Wiki-link rules:** -- Internal (starts with `.` or `@/`) -> convert to slug, wrap in [[brackets]] -- External (package name) -> plain text, no brackets - -**Critical rules matching gsd-codebase-mapper:** -- WRITE entities directly (never return contents) -- PostToolUse hook syncs to graph.db automatically -- Use EXACT template format (hook parses frontmatter + [[links]]) -- Skip existing entities (don't overwrite) -- Return statistics only (~10 lines) - -## Tasks Completed - -| Task | Name | Commit | Key Files | -|------|------|--------|-----------| -| 1 | Create gsd-entity-generator agent definition | f4c5817 | agents/gsd-entity-generator.md | - -## Deviations from Plan - -None - plan executed exactly as written. - -## Decisions Made - -1. **Skip existing entities by default** - - Prevents accidental overwrite of manually edited entities - - Aligns with research recommendation in 05-RESEARCH.md - - Agent checks `ls .planning/intel/entities/{slug}.md` before writing - -## Verification Results - -- [x] File exists at `agents/gsd-entity-generator.md` -- [x] Frontmatter valid YAML (name, description, tools, color) -- [x] Role section explains subagent purpose -- [x] Process has 3 steps: parse, process, return -- [x] Entity template included in full -- [x] Type heuristics table present -- [x] Wiki-link rules specified -- [x] Critical rules section matches gsd-codebase-mapper style -- [x] Returns statistics only (not entity contents) - -## Next Phase Readiness - -**Prerequisites for 05-02:** -- [x] gsd-entity-generator.md exists -- [x] Agent follows expected patterns (direct write, stats return) -- [x] Entity template matches what PostToolUse hook expects - -**Ready for:** Integration into `/gsd:analyze-codebase` command (Plan 05-02) diff --git a/.planning/phases/05-subagent-codebase-analysis/05-02-PLAN.md b/.planning/phases/05-subagent-codebase-analysis/05-02-PLAN.md deleted file mode 100644 index 770c4a59d..000000000 --- a/.planning/phases/05-subagent-codebase-analysis/05-02-PLAN.md +++ /dev/null @@ -1,232 +0,0 @@ ---- -phase: 05-subagent-codebase-analysis -plan: 02 -type: execute -wave: 2 -depends_on: ["05-01"] -files_modified: - - commands/gsd/analyze-codebase.md -autonomous: false - -must_haves: - truths: - - "Entity generation scales to 500+ files without context exhaustion" - - "Subagent receives only file paths, preserving orchestrator context" - - "Entity files appear in .planning/intel/entities/ after generation" - - "User can run /gsd:analyze-codebase on large codebases without degradation" - artifacts: - - path: "commands/gsd/analyze-codebase.md" - provides: "Refactored command with subagent delegation" - contains: "gsd-entity-generator" - key_links: - - from: "commands/gsd/analyze-codebase.md" - to: "agents/gsd-entity-generator.md" - via: "Task tool spawn" - pattern: "Task.*gsd-entity-generator" ---- - - -Refactor analyze-codebase Step 9 to spawn subagent instead of inline batching. - -Purpose: Replace the current "batches of 10 via Task tool" approach with a single subagent spawn. The orchestrator selects files (Step 9.2) then spawns `gsd-entity-generator` with the file list. Subagent reads files, generates entities, writes to disk. This preserves orchestrator context for large codebases. - -Output: Updated `commands/gsd/analyze-codebase.md` - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/STATE.md -@.planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md -@.planning/phases/05-subagent-codebase-analysis/05-01-SUMMARY.md -@commands/gsd/analyze-codebase.md -@agents/gsd-entity-generator.md - - - - - - Task 1: Refactor Step 9 to use subagent delegation - commands/gsd/analyze-codebase.md - -Modify Step 9 of analyze-codebase.md. Keep Steps 9.1 (create directory) and 9.2 (select files) unchanged. Replace Steps 9.3-9.5 with subagent spawn. - -**Remove:** Step 9.3 "Generate entities via Task tool batching" (the batch-of-10 pattern) - -**Replace with:** Step 9.3 "Spawn entity generator subagent" - -New Step 9.3 content: - -```markdown -### 9.3 Spawn entity generator subagent - -Spawn `gsd-entity-generator` with the selected file list. - -**Pass to subagent:** -- Total file count -- Output directory: `.planning/intel/entities/` -- Slug convention: `src/lib/db.ts` -> `src-lib-db` (replace / with -, remove extension, lowercase) -- Entity template (include full template) -- List of absolute file paths (one per line) - -**Task tool invocation:** - -```python -Task( - prompt=f"""Generate semantic entity documentation for key codebase files. - -You are a GSD entity generator. Read source files and create semantic documentation that captures PURPOSE (what/why), not just syntax. - -**Parameters:** -- Files to process: {len(selected_files)} -- Output directory: .planning/intel/entities/ -- Date: {today} - -**Slug convention:** -- src/lib/db.ts -> src-lib-db -- Replace / with -, remove extension, lowercase - -**Entity template:** -[Include full template from gsd-entity-generator.md] - -**Process:** -For each file path below: -1. Read file content using Read tool -2. Analyze purpose, exports, dependencies -3. Write entity to .planning/intel/entities/{{slug}}.md -4. PostToolUse hook syncs to graph.db automatically - -**Files:** -{file_list} - -**Return format:** -When complete, return ONLY statistics: - -## ENTITY GENERATION COMPLETE - -**Files processed:** {{N}} -**Entities created:** {{M}} -**Already existed:** {{K}} -**Errors:** {{E}} - -Entities written to: .planning/intel/entities/ - -Do NOT include entity contents in your response. -""", - subagent_type="gsd-entity-generator" -) -``` - -**Wait for completion:** Task() blocks until subagent finishes. - -**Parse result:** Extract entities_created count from response for final report. -``` - -**Update Step 9.4:** Rename from "Verify entity generation" to just verify count: -```markdown -### 9.4 Verify entity generation - -Confirm entities written: - -```bash -ls .planning/intel/entities/*.md 2>/dev/null | wc -l -``` -``` - -**Update Step 9.5:** Keep "Report entity statistics" but update text: -```markdown -### 9.5 Report entity statistics - -``` -Entity Generation Complete - -Entity files created: [N] (from subagent response) -Location: .planning/intel/entities/ -Graph database: Updated automatically via PostToolUse hook - -Next: Intel hooks will continue incremental updates as you code. -``` -``` - -**Update context section** (line ~32) - add subagent reference: -```markdown -**Execution model (Step 9 - Entity Generation):** -- Orchestrator selects files for entity generation (up to 50 based on priority) -- Spawns `gsd-entity-generator` subagent with file list (paths only, not contents) -- Subagent reads files in fresh 200k context, generates entities, writes to disk -- PostToolUse hook automatically syncs entities to graph.db -- Subagent returns statistics only (not entity contents) -- This preserves orchestrator context for large codebases (500+ files) -``` - -**Fix outdated slug documentation** (around line 286): -The entity filename convention example is outdated. Change: -```markdown -- Example: src/utils/auth.js -> src--utils--auth-js.md -``` -To: -```markdown -- Example: src/utils/auth.js -> src-utils-auth-js.md -``` -(Single hyphen, not double hyphen. The hook `gsd-intel-index.js:generateSlug` already uses single hyphen format.) - -**Important:** Do NOT pass file contents to subagent. Pass file PATHS only. Subagent reads files itself (fresh context). - - -Read updated commands/gsd/analyze-codebase.md and confirm: -- Step 9.3 spawns gsd-entity-generator via Task tool -- No batch-of-10 pattern remains -- File paths passed, not file contents -- Context section mentions subagent delegation -- Steps 9.4-9.5 updated for new flow -- Slug example at line ~286 uses single hyphen (src-utils-auth-js.md) - - -analyze-codebase.md Step 9 refactored. Entity generation now delegates to gsd-entity-generator subagent instead of inline Task batching. Outdated slug documentation corrected. - - - - - Subagent delegation for entity generation in /gsd:analyze-codebase - -1. Navigate to a test project (not this repo) -2. Run `/gsd:analyze-codebase` -3. Observe: - - Steps 1-8 complete (index, conventions, summary) - - Step 9.2 selects files (should see file list) - - Step 9.3 spawns subagent (Task tool call with "gsd-entity-generator") - - Subagent generates entities (should see Read/Write calls in subagent) - - Subagent returns statistics only (not entity contents) -4. Verify `.planning/intel/entities/` contains entity files -5. Verify entity files follow template (frontmatter, Purpose, Exports, Dependencies with [[links]]) -6. Verify graph.db was updated (entities appear in summary.md) - - Type "approved" if entity generation works via subagent, or describe issues - - - - - -- [ ] Step 9.3 spawns gsd-entity-generator subagent -- [ ] File paths passed to subagent (not file contents) -- [ ] No batch-of-10 pattern in command -- [ ] Context section documents subagent model -- [ ] Slug example uses single hyphen format (matches hook) -- [ ] Entity files created in test project -- [ ] Entities follow template format -- [ ] Graph database updated (via hook) -- [ ] User verification checkpoint passed - - - -analyze-codebase refactored to use subagent delegation. Entity generation works end-to-end on a test project, with orchestrator context preserved and entities written correctly. - - - -After completion, create `.planning/phases/05-subagent-codebase-analysis/05-02-SUMMARY.md` - diff --git a/.planning/phases/05-subagent-codebase-analysis/05-03-PLAN.md b/.planning/phases/05-subagent-codebase-analysis/05-03-PLAN.md deleted file mode 100644 index be46579bb..000000000 --- a/.planning/phases/05-subagent-codebase-analysis/05-03-PLAN.md +++ /dev/null @@ -1,192 +0,0 @@ ---- -phase: 05-subagent-codebase-analysis -plan: 03 -type: execute -wave: 3 -depends_on: [] -files_modified: - - agents/gsd-indexer.md -autonomous: true -gap_closure: true - -must_haves: - truths: - - "Indexer agent definition exists following gsd-entity-generator pattern" - - "Agent reads files and extracts exports/imports using same regex as Step 3" - - "Agent writes index.json directly to disk" - - "Agent returns statistics only (not file contents)" - artifacts: - - path: "agents/gsd-indexer.md" - provides: "Subagent definition for file indexing" - contains: "gsd-indexer" - key_links: - - from: "agents/gsd-indexer.md" - to: ".planning/intel/index.json" - via: "Write tool call" - pattern: "Write.*index\\.json" ---- - - -Create gsd-indexer subagent definition for Steps 2-3 file indexing. - -Purpose: Define the subagent that reads files and extracts exports/imports. This agent is spawned by `/gsd:analyze-codebase` with a list of file paths (from Glob), reads each file, applies the same regex patterns as current Step 3, and writes the complete index.json to disk. Returns statistics only to preserve orchestrator context. - -Output: `agents/gsd-indexer.md` - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/STATE.md -@.planning/phases/05-subagent-codebase-analysis/05-VERIFICATION.md -@agents/gsd-entity-generator.md -@agents/gsd-codebase-mapper.md -@commands/gsd/analyze-codebase.md - - - - - - Task 1: Create gsd-indexer agent definition - agents/gsd-indexer.md - -Create agent definition following gsd-entity-generator.md structure: - -**Frontmatter:** -```yaml ---- -name: gsd-indexer -description: Indexes codebase files by extracting exports and imports. Spawned by analyze-codebase with file list. Writes index.json directly to disk. -tools: Read, Write, Bash -color: cyan ---- -``` - -**Role section:** -- Spawned by `/gsd:analyze-codebase` with file paths (from Glob results) -- Reads source files using Read tool -- Extracts exports and imports using regex patterns -- Writes complete index.json to `.planning/intel/index.json` -- Returns statistics only (NOT file contents or index data) - -**Why this matters section:** -- Index.json is consumed by convention detection (Step 4) -- Entity generation uses index to find hub files (Step 9.2) -- PostToolUse hook uses index for incremental updates -- Orchestrator MUST NOT load file contents (context exhaustion on 500+ files) - -**Process sections:** - -1. `parse_input` - Extract from prompt: - - Output path: `.planning/intel/index.json` - - List of absolute file paths (one per line) - - Initialize counters: files_processed=0, exports_found=0, imports_found=0, errors=0 - -2. `process_each_file` - For each file path: - - a. Read file content using Read tool - - b. Extract exports using these patterns (EXACTLY as Step 3): - - Named exports: `export\s*\{([^}]+)\}` - - Declaration exports: `export\s+(?:const|let|var|function\*?|async\s+function|class)\s+(\w+)` - - Default exports: `export\s+default\s+(?:function\s*\*?\s*|class\s+)?(\w+)?` - - CommonJS object: `module\.exports\s*=\s*\{([^}]+)\}` - - CommonJS single: `module\.exports\s*=\s*(\w+)\s*[;\n]` - - TypeScript: `export\s+(?:type|interface)\s+(\w+)` - - c. Extract imports using these patterns (EXACTLY as Step 3): - - ES6: `import\s+(?:\{[^}]*\}|\*\s+as\s+\w+|\w+)\s+from\s+['"]([^'"]+)['"]` - - Side-effect: `import\s+['"]([^'"]+)['"]` (not preceded by 'from') - - CommonJS: `require\s*\(\s*['"]([^'"]+)['"]\s*\)` - - d. Store in index structure: - ```javascript - index.files[absolutePath] = { - exports: [], // Array of export names - imports: [], // Array of import sources - indexed: Date.now() - } - ``` - - e. Track statistics: increment files_processed, add to exports_found/imports_found - - f. Handle errors: if file can't be read, increment errors and continue - -3. `write_index` - Write complete index to disk: - ```javascript - { - version: 1, - updated: Date.now(), - files: { /* all file entries */ } - } - ``` - Write to `.planning/intel/index.json` using Write tool. - -4. `return_statistics` - Return ONLY: - ``` - ## INDEXING COMPLETE - - **Files processed:** {files_processed} - **Exports found:** {exports_found} - **Imports found:** {imports_found} - **Errors:** {errors} - - Index written to: .planning/intel/index.json - ``` - -**Critical rules section:** -- WRITE index.json directly (never return index contents) -- Use EXACT regex patterns from Step 3 (patterns are validated) -- Handle read errors gracefully (log path, continue) -- Return only statistics (~10 lines) -- DO NOT commit (orchestrator handles git) -- File paths in index must be ABSOLUTE paths (key for O(1) lookup) - -**Success criteria checklist:** -- All file paths processed -- index.json written with version, updated, files -- Each file entry has exports, imports, indexed timestamp -- Absolute paths as keys (not relative) -- Statistics returned (not index contents) - - -File exists and contains: -- Frontmatter with name: gsd-indexer, tools: Read/Write/Bash -- Role section explaining spawn context from analyze-codebase -- Why this matters section explaining index consumers -- Process steps: parse_input, process_each_file, write_index, return_statistics -- EXACT regex patterns matching Step 3 of analyze-codebase -- Critical rules section (write directly, return stats only) -- Success criteria checklist - - -`agents/gsd-indexer.md` exists with complete agent definition. Agent uses same regex patterns as current Step 3. Ready to be spawned by analyze-codebase. - - - - - - -- [ ] File exists at `agents/gsd-indexer.md` -- [ ] Frontmatter valid YAML (name, description, tools, color) -- [ ] Role section explains subagent purpose (spawned by analyze-codebase) -- [ ] Process has 4 steps: parse, process, write, return -- [ ] Export regex patterns match Step 3 exactly (6 patterns) -- [ ] Import regex patterns match Step 3 exactly (3 patterns) -- [ ] Index schema matches Step 5 format (version, updated, files) -- [ ] Critical rules section matches gsd-entity-generator style -- [ ] Returns statistics only (not index contents) - - - -gsd-indexer agent definition complete. Agent can be spawned with file paths and will read files, extract exports/imports using validated regex patterns, and write index.json directly to disk. - - - -After completion, create `.planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md` - diff --git a/.planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md b/.planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md deleted file mode 100644 index d19fe6aeb..000000000 --- a/.planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md +++ /dev/null @@ -1,135 +0,0 @@ ---- -phase: 05-subagent-codebase-analysis -plan: 03 -title: gsd-indexer Agent Definition -subsystem: agents -tags: [subagent, indexing, file-analysis, context-preservation] - -dependency-graph: - requires: - - 05-01 (gsd-entity-generator pattern reference) - - gsd-codebase-mapper.md (agent structure pattern) - - analyze-codebase.md Step 3 (regex patterns) - provides: - - gsd-indexer subagent definition - - File indexing via subagent delegation - - Context-preserving index generation - affects: - - 05-04 (analyze-codebase integration) - - Steps 2-3 context exhaustion fix - -tech-stack: - added: [] - patterns: - - Subagent direct-write pattern - - Statistics-only return pattern - - Regex-based export/import extraction - -key-files: - created: - - agents/gsd-indexer.md - modified: [] - -decisions: - - id: absolute-path-keys - choice: Use absolute paths as index keys - rationale: O(1) lookup, matches existing index.json schema - -metrics: - duration: 1 min 23 sec - completed: 2026-01-21 ---- - -# Phase 05 Plan 03: gsd-indexer Agent Definition Summary - -**One-liner:** Subagent definition for file indexing that extracts exports/imports using validated regex patterns and writes index.json directly to disk. - -## What Was Built - -Created `agents/gsd-indexer.md` following the `gsd-entity-generator.md` pattern: - -**Agent structure:** -- Frontmatter with name, description, tools (Read, Write, Bash), color -- Role section explaining spawn context from `/gsd:analyze-codebase` -- `` section explaining index consumers (convention detection, entity generation, PostToolUse hook) -- 4-step process: parse_input, process_each_file, write_index, return_statistics - -**Regex patterns (exact match with analyze-codebase Step 3):** - -Export patterns: -| Pattern | Regex | Purpose | -|---------|-------|---------| -| Named exports | `export\s*\{([^}]+)\}` | `export { a, b }` | -| Declaration | `export\s+(?:const\|let\|var\|function\*?\|async\s+function\|class)\s+(\w+)` | `export const foo` | -| Default | `export\s+default\s+(?:function\s*\*?\s*\|class\s+)?(\w+)?` | `export default` | -| CommonJS object | `module\.exports\s*=\s*\{([^}]+)\}` | `module.exports = { }` | -| CommonJS single | `module\.exports\s*=\s*(\w+)\s*[;\n]` | `module.exports = X` | -| TypeScript | `export\s+(?:type\|interface)\s+(\w+)` | `export type/interface` | - -Import patterns: -| Pattern | Regex | Purpose | -|---------|-------|---------| -| ES6 | `import\s+(?:\{[^}]*\}\|\*\s+as\s+\w+\|\w+)\s+from\s+['"]([^'"]+)['"]` | `import X from 'y'` | -| Side-effect | `import\s+['"]([^'"]+)['"]` | `import 'styles.css'` | -| CommonJS | `require\s*\(\s*['"]([^'"]+)['"]\s*\)` | `require('x')` | - -**Index schema (matches Step 5):** -```javascript -{ - version: 1, - updated: Date.now(), - files: { - "/absolute/path": { - exports: [], - imports: [], - indexed: Date.now() - } - } -} -``` - -**Critical rules:** -- Write index.json directly (never return contents) -- Use exact regex patterns from Step 3 -- Absolute paths as keys for O(1) lookup -- Handle read errors gracefully (log, continue) -- Return statistics only (~10 lines) - -## Tasks Completed - -| Task | Name | Commit | Key Files | -|------|------|--------|-----------| -| 1 | Create gsd-indexer agent definition | 5d03e14 | agents/gsd-indexer.md | - -## Deviations from Plan - -None - plan executed exactly as written. - -## Decisions Made - -1. **Absolute path keys** - - Continues existing index.json schema decision from 01-01 - - Enables O(1) lookup for any file path - - Consistent with PostToolUse hook expectations - -## Verification Results - -- [x] File exists at `agents/gsd-indexer.md` -- [x] Frontmatter valid YAML (name, description, tools, color) -- [x] Role section explains subagent purpose (spawned by analyze-codebase) -- [x] Process has 4 steps: parse, process, write, return -- [x] Export regex patterns match Step 3 exactly (6 patterns) -- [x] Import regex patterns match Step 3 exactly (3 patterns) -- [x] Index schema matches Step 5 format (version, updated, files) -- [x] Critical rules section matches gsd-entity-generator style -- [x] Returns statistics only (not index contents) - -## Next Phase Readiness - -**Prerequisites for 05-04:** -- [x] gsd-indexer.md exists -- [x] Agent follows expected patterns (direct write, stats return) -- [x] Regex patterns validated against analyze-codebase Step 3 -- [x] Index schema matches existing expectations - -**Ready for:** Integration into `/gsd:analyze-codebase` command Steps 2-3 (Plan 05-04) diff --git a/.planning/phases/05-subagent-codebase-analysis/05-04-PLAN.md b/.planning/phases/05-subagent-codebase-analysis/05-04-PLAN.md deleted file mode 100644 index fcdb0daf4..000000000 --- a/.planning/phases/05-subagent-codebase-analysis/05-04-PLAN.md +++ /dev/null @@ -1,272 +0,0 @@ ---- -phase: 05-subagent-codebase-analysis -plan: 04 -type: execute -wave: 4 -depends_on: ["05-03"] -files_modified: - - commands/gsd/analyze-codebase.md -autonomous: false -gap_closure: true - -must_haves: - truths: - - "Orchestrator never reads file contents during Steps 2-3" - - "Indexer subagent is spawned with file paths from Glob" - - "Subagent writes index.json, orchestrator receives statistics only" - - "User can run /gsd:analyze-codebase on 500+ file codebases without context exhaustion" - artifacts: - - path: "commands/gsd/analyze-codebase.md" - provides: "Refactored command with subagent delegation for indexing" - contains: "gsd-indexer" - key_links: - - from: "commands/gsd/analyze-codebase.md" - to: "agents/gsd-indexer.md" - via: "Task tool spawn" - pattern: "Task.*gsd-indexer" - - from: "Step 2 Glob" - to: "Step 3 subagent spawn" - via: "file paths array" - pattern: "Glob.*file_paths" ---- - - -Refactor analyze-codebase Steps 2-3 to spawn indexer subagent. - -Purpose: Replace the current inline file reading (Step 3: "Read file content using Read tool") with a single subagent spawn. The orchestrator runs Glob to find files (Step 2), spawns `gsd-indexer` with the file list, and receives index.json back. This prevents context exhaustion on large codebases where 250+ file reads were causing the orchestrator to die before reaching Step 9. - -Output: Updated `commands/gsd/analyze-codebase.md` - - - -@~/.claude/get-shit-done/workflows/execute-plan.md -@~/.claude/get-shit-done/templates/summary.md - - - -@.planning/PROJECT.md -@.planning/ROADMAP.md -@.planning/STATE.md -@.planning/phases/05-subagent-codebase-analysis/05-VERIFICATION.md -@.planning/phases/05-subagent-codebase-analysis/05-03-SUMMARY.md -@commands/gsd/analyze-codebase.md -@agents/gsd-indexer.md - - - - - - Task 1: Refactor Steps 2-3 to use indexer subagent - commands/gsd/analyze-codebase.md - -Modify Steps 2 and 3 of analyze-codebase.md. Step 2 stays mostly the same (Glob). Step 3 is completely replaced with subagent spawn. - -**Update Step 2: Find all indexable files** - -Keep the Glob pattern and exclusion list, but change the output description: - -```markdown -## Step 2: Find all indexable files - -Use Glob tool with pattern: `**/*.{js,ts,jsx,tsx,mjs,cjs}` - -Exclude directories (skip any path containing): -- node_modules -- dist -- build -- .git -- vendor -- coverage -- .next -- __pycache__ - -Filter results to remove excluded paths. Store as `file_paths` array. - -**Output:** List of absolute file paths for indexing. Do NOT read file contents. -``` - -**Replace Step 3: Process each file** - -Remove the entire current Step 3 (lines ~67-100) which contains: -- "Read file content using Read tool" -- Export regex patterns inline -- Import regex patterns inline -- Store in index structure - -Replace with: - -```markdown -## Step 3: Spawn indexer subagent - -Spawn `gsd-indexer` subagent with the file paths from Step 2. - -**Why subagent delegation:** -- Orchestrator would exhaust context reading 500+ files inline -- Subagent gets fresh 200k context for file reading -- Orchestrator only handles file paths (small) -- Subagent writes index.json directly (large) - -**Task tool invocation:** - -```python -file_list = "\n".join(file_paths) # From Step 2 Glob results - -Task( - prompt=f"""Index codebase files by extracting exports and imports. - -You are a GSD indexer. Read source files and extract exports/imports using regex patterns. - -**Parameters:** -- Files to process: {len(file_paths)} -- Output path: .planning/intel/index.json - -**Export patterns:** -- Named: export\\s*\\{{([^}}]+)\\}} -- Declaration: export\\s+(?:const|let|var|function\\*?|async\\s+function|class)\\s+(\\w+) -- Default: export\\s+default\\s+(?:function\\s*\\*?\\s*|class\\s+)?(\\w+)? -- CommonJS object: module\\.exports\\s*=\\s*\\{{([^}}]+)\\}} -- CommonJS single: module\\.exports\\s*=\\s*(\\w+)\\s*[;\\n] -- TypeScript: export\\s+(?:type|interface)\\s+(\\w+) - -**Import patterns:** -- ES6: import\\s+(?:\\{{[^}}]*\\}}|\\*\\s+as\\s+\\w+|\\w+)\\s+from\\s+['\"]([^'\"]+)['\"] -- Side-effect: import\\s+['\"]([^'\"]+)['\"] -- CommonJS: require\\s*\\(\\s*['\"]([^'\"]+)['\"]\\s*\\) - -**Index schema:** -```json -{{ - "version": 1, - "updated": {timestamp}, - "files": {{ - "/absolute/path/file.js": {{ - "exports": ["name1", "name2"], - "imports": ["source1", "source2"], - "indexed": {timestamp} - }} - }} -}} -``` - -**Process:** -For each file path below: -1. Read file content using Read tool -2. Apply export regex patterns, collect export names -3. Apply import regex patterns, collect import sources -4. Store in index structure with absolute path as key - -**Files:** -{file_list} - -**Return format:** -When complete, return ONLY statistics: - -## INDEXING COMPLETE - -**Files processed:** {{N}} -**Exports found:** {{M}} -**Imports found:** {{K}} -**Errors:** {{E}} - -Index written to: .planning/intel/index.json - -Do NOT return index contents. -""", - subagent_type="gsd-indexer" -) -``` - -**Wait for completion:** Task() blocks until subagent finishes. - -**Verify index created:** -```bash -ls -la .planning/intel/index.json -``` -``` - -**Update context section** (around lines 21-38): - -Add to the context section, before the existing "Execution model (Step 9)" section: - -```markdown -**Execution model (Steps 2-3 - Indexing):** -- Orchestrator finds file paths via Glob (Step 2) -- Spawns `gsd-indexer` subagent with file paths only (Step 3) -- Subagent reads files in fresh 200k context, applies regex patterns -- Subagent writes index.json directly to disk -- Subagent returns statistics only (not file contents or index data) -- This prevents context exhaustion on large codebases (500+ files) -``` - -**Remove inline regex documentation:** - -The regex patterns are now documented in the subagent prompt. Remove any duplicate documentation that described inline processing. The command should make clear that file reading is DELEGATED, not performed inline. - -**Verify Step 4 still works:** - -Step 4 "Detect conventions" reads from index.json that the subagent wrote. This should work unchanged since the index format is the same. - - -Read updated commands/gsd/analyze-codebase.md and confirm: -- Step 2 mentions storing as "file_paths" array, NOT reading contents -- Step 3 spawns gsd-indexer via Task tool -- No "Read file content" instruction in orchestrator steps -- Context section documents subagent delegation for Steps 2-3 -- Step 4+ remain unchanged (read index.json from disk) -- Regex patterns appear in subagent prompt (not inline) - - -analyze-codebase.md Steps 2-3 refactored. Orchestrator only Globs for paths, then spawns indexer subagent. File reading is fully delegated to subagent with fresh context. - - - - - Subagent delegation for indexing in /gsd:analyze-codebase (Steps 2-3) - -1. Navigate to a LARGE test project (100+ files, ideally 250+) -2. Delete existing intel if any: `rm -rf .planning/intel` -3. Run `/gsd:analyze-codebase` -4. Observe execution: - - Step 1: Creates directory - - Step 2: Runs Glob, collects file paths (does NOT read files) - - Step 3: Spawns gsd-indexer subagent with file list - - Subagent reads files (should see many Read tool calls in subagent output) - - Subagent writes index.json directly - - Subagent returns statistics only - - Step 4+: Orchestrator reads index.json, continues with conventions - - Step 9: Spawns gsd-entity-generator (existing flow) -5. Verify: - - Orchestrator context NOT exhausted (no early death) - - `.planning/intel/index.json` exists with file entries - - `.planning/intel/conventions.json` has detected patterns - - `.planning/intel/summary.md` exists - - Entity generation works (if Step 9 executed) -6. Compare orchestrator context usage before/after refactor - - Before: Orchestrator read all files -> context exhaustion at ~250 files - - After: Orchestrator only handles paths -> survives 500+ files - - Type "approved" if /gsd:analyze-codebase completes on large codebase without context exhaustion, or describe issues - - - - - -- [ ] Step 2 outputs file_paths array (no file reading) -- [ ] Step 3 spawns gsd-indexer subagent -- [ ] Subagent prompt includes all 6 export patterns -- [ ] Subagent prompt includes all 3 import patterns -- [ ] Context section documents indexing subagent delegation -- [ ] No "Read file content" in orchestrator steps -- [ ] Step 4+ reads index.json from disk (unchanged) -- [ ] Command runs successfully on 250+ file codebase -- [ ] Orchestrator context preserved (no early death) -- [ ] User verification checkpoint passed - - - -analyze-codebase refactored with full subagent delegation. Both indexing (Steps 2-3) and entity generation (Step 9) now use subagents. Command works on 500+ file codebases without context exhaustion. - - - -After completion, create `.planning/phases/05-subagent-codebase-analysis/05-04-SUMMARY.md` - diff --git a/.planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md b/.planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md deleted file mode 100644 index 9aa0cb0f3..000000000 --- a/.planning/phases/05-subagent-codebase-analysis/05-RESEARCH.md +++ /dev/null @@ -1,774 +0,0 @@ -# Phase 5: Subagent Codebase Analysis - Research - -**Researched:** 2026-01-20 -**Domain:** Subagent orchestration, context window management, Claude Code Task tool -**Confidence:** HIGH - -## Summary - -Phase 5 refactors the current `/gsd:analyze-codebase` entity generation from main-context execution to subagent delegation. The current implementation (Phase 4) has Claude executing the command generate entity content directly via the Task tool with batches of 10 files, but ALL exploration and decision-making happens in the main orchestrator context. - -**The problem:** On large codebases (500+ files), the orchestrator exhausts context by: -1. Reading all files during selection (identifying which 50 files to generate entities for) -2. Orchestrating batch splits and subagent spawns -3. Collecting and validating subagent results - -**The solution:** Delegate the entire entity generation phase to a subagent, following the `gsd-codebase-mapper.md` pattern: -- Orchestrator provides minimal instructions + file list -- Subagent operates in fresh 200k context -- Subagent reads files, generates entities, writes directly to disk -- Subagent returns only confirmation (not entity contents) - -**Current baseline:** `gsd-codebase-mapper.md` demonstrates successful subagent delegation for analysis tasks. It spawns with a focus area, explores thoroughly, writes documents directly, and returns ~10 lines of confirmation. This pattern scales because the orchestrator never loads document contents. - -**Primary recommendation:** Extract entity generation (Step 9) from `/gsd:analyze-codebase` into a dedicated `gsd-entity-generator` subagent. Orchestrator handles Steps 1-8 (indexing), then spawns subagent with file list. Subagent generates all entities and returns statistics only. - -## Standard Stack - -### Core - -| Library | Version | Purpose | Why Standard | -|---------|---------|---------|--------------| -| Claude Code Task tool | Built-in | Subagent spawning with model selection | Official Claude Code orchestration primitive | -| sql.js | 1.12.0+ | Graph database in subagent context | Subagent needs graph access to resolve [[wiki-links]] | - -### Supporting - -| Component | Version | Purpose | When to Use | -|---------|---------|---------|-------------| -| gsd-codebase-mapper.md | Current | Reference pattern for subagent delegation | Template for entity-generator architecture | -| gsd-executor.md | Current | Reference for Task tool usage | Shows how to spawn with model profile | - -### Installation - -No new dependencies required. Phase 5 refactors existing architecture. - -### Alternatives Considered - -| Instead of | Could Use | Tradeoff | -|------------|-----------|----------| -| Single subagent for all files | Multiple subagents (1 per batch) | Multiple subagents = parallel processing but complex orchestration and potential race conditions in graph.db writes | -| Subagent with file list | Subagent discovers files itself | Discovery requires reading index.json anyway, orchestrator already has this data | -| Inline in hook | Keep current approach | Hooks have strict execution limits, can't spawn subagents or handle 500+ files | - -## Architecture Patterns - -### Recommended Execution Flow - -``` -User: /gsd:analyze-codebase - -Orchestrator (main context): -├── Steps 1-8: Index creation (existing) -│ ├── Scan files with Glob -│ ├── Extract exports/imports -│ ├── Write index.json, conventions.json, summary.md -│ └── Identify 50 key files for entity generation -│ -└── Step 9: Spawn entity-generator subagent - ├── Pass: file list, index data, config - ├── Subagent operates in fresh 200k context - └── Returns: { entities_created: N, skipped: M } - -Subagent (gsd-entity-generator): -├── Load file list from prompt -├── For each file: -│ ├── Read file content -│ ├── Generate entity markdown (Claude's own analysis) -│ ├── Write to .planning/intel/entities/{slug}.md -│ └── PostToolUse hook syncs to graph.db -└── Return confirmation statistics -``` - -### Pattern 1: Minimal Orchestrator Handoff - -**What:** Orchestrator passes only essential data, not full file contents - -**When to use:** When subagent needs to read files itself (maintains fresh context) - -**Example:** - -```markdown -Task( - prompt=f"""Generate semantic entity files for key codebase files. - -Entity generation parameters: -- Total files to process: {len(selected_files)} -- Output directory: .planning/intel/entities/ -- Slug convention: src/lib/db.ts -> src-lib-db - -Files to process: -{chr(10).join(selected_files)} - -For each file: -1. Read the file content using Read tool -2. Analyze purpose, exports, dependencies -3. Generate entity markdown following template -4. Write to .planning/intel/entities/{{slug}}.md - -Entity template: ---- -path: {{file_path}} -type: [module|component|util|config|api|hook|service|model] -updated: {today} -status: active ---- - -# {{filename}} - -## Purpose - -[1-3 sentences: What does this file do? Why does it exist?] - -[... rest of template ...] - -After all files processed, return statistics: -- Entities created: N -- Files skipped: M (already existed) -- Errors: K (if any) -""", - subagent_type="gsd-entity-generator", - model="{model}" -) -``` - -**Why minimal:** Passing file contents in prompt exhausts orchestrator context (defeats purpose of subagent delegation). - -### Pattern 2: Subagent Direct Write - -**What:** Subagent writes entity files directly, doesn't return contents to orchestrator - -**When to use:** Always (learned from gsd-codebase-mapper.md) - -**Anti-pattern:** - -```python -# DON'T: Return entity contents to orchestrator -entities = [] -for file in files: - entity = generate_entity(file) # Claude generates - entities.append(entity) # Accumulates in context -return entities # Passes back to orchestrator -``` - -**Correct pattern:** - -```python -# DO: Write directly, return only confirmation -for file in files: - entity_content = generate_entity(file) # Claude generates - Write(path=f".planning/intel/entities/{slug}.md", content=entity_content) - # PostToolUse hook automatically syncs to graph.db - -return { - "entities_created": len(files), - "location": ".planning/intel/entities/" -} -``` - -### Pattern 3: Model Profile Resolution - -**What:** Orchestrator resolves model profile, passes specific model to Task tool - -**When to use:** Every subagent spawn (ensures consistent model selection) - -**Example:** - -```bash -# Read model profile from config -MODEL_PROFILE=$(cat .planning/config.json 2>/dev/null | \ - grep -o '"model_profile"[[:space:]]*:[[:space:]]*"[^"]*"' | \ - grep -o '"[^"]*"$' | tr -d '"' || echo "balanced") -``` - -```python -# Model lookup table for gsd-entity-generator -model_map = { - "quality": "claude-opus-4-5-20251101", - "balanced": "claude-sonnet-4-5-20250929", - "budget": "claude-sonnet-4-5-20250929" -} -model = model_map.get(MODEL_PROFILE, "claude-sonnet-4-5-20250929") -``` - -**Why this matters:** Entity generation is semantic analysis (requires strong reasoning). Sonnet 4.5 adequate for balanced/budget, Opus 4.5 for quality profile. - -### Pattern 4: PostToolUse Hook Integration - -**What:** Subagent writes entities, hook syncs to graph.db automatically - -**When to use:** Always (no explicit sync needed in subagent) - -**Flow:** - -``` -Subagent writes: .planning/intel/entities/src-lib-db.md - ↓ -PostToolUse hook (gsd-intel-index.js) detects Write tool - ↓ -isEntityFile(path) returns true - ↓ -syncEntityToGraph(path): - - Extracts frontmatter - - Extracts [[wiki-links]] - - Upserts node to graph.db - - Inserts edges - - Persists database - ↓ -regenerateEntitySummary(): - - Generates new summary from graph - - Writes summary.md -``` - -**Critical insight:** Subagent doesn't need graph access for writes. Hook handles all graph operations. Subagent only needs to write well-formed entity markdown. - -### Anti-Patterns to Avoid - -**Anti-pattern 1: Orchestrator reads files for subagent** - -```python -# DON'T: Load all file contents in orchestrator -file_contents = {} -for file_path in selected_files: - file_contents[file_path] = Read(file_path) # Exhausts orchestrator context - -Task(prompt=f"Generate entities for: {file_contents}", ...) # Too late, context blown -``` - -**Why it fails:** Defeats purpose of subagent delegation (500 files × 5KB avg = 2.5MB of code in orchestrator context). - -**Anti-pattern 2: Multiple parallel subagents** - -```python -# DON'T: Spawn one subagent per file (or per small batch) -for file in selected_files: - Task(prompt=f"Generate entity for {file}", ...) # 50 subagents = chaos -``` - -**Why it fails:** -- 50 concurrent subagents writing to `.planning/intel/entities/` -- 50 concurrent PostToolUse hooks writing to `graph.db` -- Race conditions in sql.js export/import (no file locking) -- Excessive API calls (50 × context overhead) - -**Anti-pattern 3: Subagent returns generated content** - -```markdown -## ENTITY GENERATION COMPLETE - -**Files processed:** 50 - -**Generated entities:** - -[... 50 × 500 lines of entity content ...] -``` - -**Why it fails:** Orchestrator asked for entity generation, not entity contents. Return confirmation only. - -## Don't Hand-Roll - -| Problem | Don't Build | Use Instead | Why | -|---------|-------------|-------------|-----| -| Subagent orchestration | Custom process spawning | Claude Code Task tool | Built-in, handles model selection, output capture, error handling | -| File batching | Complex batch scheduling | Single subagent with full file list | Subagent has 200k context (enough for 500 file paths), no coordination overhead | -| Entity template | Embedded in subagent logic | Pass template in prompt | Template may evolve, keep it in prompt not in subagent code | -| Graph database access | Direct sql.js in subagent | PostToolUse hook handles sync | Hooks are designed for this, no need to duplicate graph logic | -| Progress tracking | Real-time updates to orchestrator | Batch completion confirmation | Subagent works in isolation, returns final stats | - -**Key insight:** The Task tool is designed for this exact pattern. Don't try to implement manual process forking, IPC, or result aggregation. Task() blocks until subagent completes, returns output directly. - -## Common Pitfalls - -### Pitfall 1: Orchestrator Context Bloat - -**What goes wrong:** Orchestrator loads too much data before spawning subagent, exhausts context anyway. - -**Why it happens:** Instinct to "prepare everything" for the subagent (read files, validate, format). - -**How to avoid:** -- Orchestrator's job: Identify WHICH files (paths only) -- Subagent's job: Read files and process -- Pass file paths, not file contents -- Trust subagent to handle file reading - -**Warning signs:** Orchestrator context usage >50% before Task() call. - -### Pitfall 2: Subagent Result Explosion - -**What goes wrong:** Subagent returns all generated entity content to orchestrator. - -**Why it happens:** Thinking orchestrator needs to "validate" or "display" results. - -**How to avoid:** -- Subagent writes directly to `.planning/intel/entities/` -- Return only: `{ entities_created: N, errors: [] }` -- Orchestrator reports to user: "Created N entities in .planning/intel/entities/" -- User can inspect entity files themselves - -**Warning signs:** Task() return value contains >1000 lines of text. - -### Pitfall 3: Race Conditions in graph.db - -**What goes wrong:** Multiple concurrent writes to graph.db corrupt database. - -**Why it happens:** PostToolUse hook fires for every Write. 50 entity writes = 50 hooks attempting sql.js export/import simultaneously. - -**How to avoid:** -- Use single subagent (not parallel subagents) -- PostToolUse hooks run sequentially per Write operation -- sql.js persistence is synchronous (no async race conditions within single process) -- If implementing parallel subagents in future: file locking or write queue required - -**Warning signs:** graph.db corruption, missing edges, "database is locked" errors. - -### Pitfall 4: Missing Entity Template - -**What goes wrong:** Subagent generates entities in wrong format, hook can't parse frontmatter/links. - -**Why it happens:** Template not provided in subagent prompt, or template is incomplete. - -**How to avoid:** -- Include EXACT entity template in Task prompt -- Show examples with all sections (frontmatter, Purpose, Exports, Dependencies, Used By, Notes) -- Specify [[wiki-link]] format for internal dependencies -- Test generated entities: hook should extract frontmatter + links successfully - -**Warning signs:** Entities created but not appearing in graph queries, summary.md unchanged. - -### Pitfall 5: No Progress Visibility - -**What goes wrong:** Subagent processes 500 files silently, user has no idea if it's working or stuck. - -**Why it happens:** Task() blocks until completion, no intermediate output. - -**How to avoid:** -- For large operations (100+ files), consider chunking: - - Orchestrator: Split 500 files into 10 batches of 50 - - Spawn subagent per batch (sequential, not parallel) - - Report progress: "Batch 3/10 complete (150/500 files)" -- For moderate operations (50-100 files): Single subagent is fine, document expected duration - -**Warning signs:** User cancels command thinking it's frozen (actually still processing). - -## Code Examples - -### Complete Subagent Spawn (Orchestrator) - -```markdown -## Step 9: Generate Semantic Entities via Subagent - -Read model profile from config: - -```bash -MODEL_PROFILE=$(cat .planning/config.json 2>/dev/null | \ - grep -o '"model_profile"[[:space:]]*:[[:space:]]*"[^"]*"' | \ - grep -o '"[^"]*"$' | tr -d '"' || echo "balanced") -``` - -Resolve model for gsd-entity-generator: - -| Profile | Model | -|---------|-------| -| quality | claude-opus-4-5-20251101 | -| balanced | claude-sonnet-4-5-20250929 | -| budget | claude-sonnet-4-5-20250929 | - -Spawn entity generator subagent: - -```python -# Build file list from Step 9a selection -file_list = "\n".join(selected_files) -today = date.today().isoformat() - -Task( - prompt=f"""Generate semantic entity documentation for key codebase files. - -You are a GSD entity generator. You read source files and create semantic documentation that captures PURPOSE (what/why), not just syntax. - -**Parameters:** -- Files to process: {len(selected_files)} -- Output directory: .planning/intel/entities/ -- Date: {today} - -**Slug convention:** -- src/lib/db.ts → src-lib-db -- Replace / with -, remove extension - -**Entity template (use EXACTLY this format):** - ---- -path: {{absolute_file_path}} -type: [module|component|util|config|api|hook|service|model|test] -updated: {today} -status: active ---- - -# {{filename}} - -## Purpose - -[1-3 sentences explaining what this file does and why it exists. Focus on the problem it solves, not implementation details.] - -## Exports - -[For each export, provide signature and brief purpose:] -- `functionName(params): ReturnType` - What it does -- `ClassName` - What it represents - -If no exports: "None" - -## Dependencies - -[Internal dependencies as [[wiki-links]], external as plain text:] -- [[internal-file-slug]] - Why this dependency exists -- external-package - What it provides - -If no dependencies: "None" - -## Used By - -TBD - -## Notes - -[Optional: Important patterns, gotchas, or context. Omit if nothing notable.] - -**Process:** - -For each file path below: -1. Read file content using Read tool -2. Analyze exports, imports, purpose -3. Write entity to .planning/intel/entities/{{slug}}.md -4. PostToolUse hook will sync to graph.db automatically - -**Files:** - -{file_list} - -**Return format:** - -When all files processed, return ONLY this structure: - -``` -## ENTITY GENERATION COMPLETE - -**Files processed:** {{N}} -**Entities created:** {{M}} -**Already existed:** {{K}} -**Errors:** {{E}} (if any, list file paths) - -Entities written to: .planning/intel/entities/ -``` - -Do NOT include entity contents in your response. -""", - subagent_type="gsd-entity-generator", - model=model -) -``` - -Wait for subagent completion. Task() blocks until done. - -Parse result for statistics: -- Extract entities_created count -- Report to user -``` - -### Subagent Response Format (gsd-entity-generator) - -```markdown -## ENTITY GENERATION COMPLETE - -**Files processed:** 47 -**Entities created:** 47 -**Already existed:** 0 -**Errors:** 0 - -Entities written to: .planning/intel/entities/ -``` - -**Critical:** Subagent does NOT return entity contents. Only statistics. - -### Orchestrator Final Report - -```markdown -Codebase Analysis Complete - -Files indexed: 347 -Exports found: 1,423 -Imports found: 2,891 - -Conventions detected: -- Naming: camelCase (87%) -- Directories: components/ (23 files), lib/ (12 files), api/ (8 files) -- Patterns: *.test.ts (34 files), *.config.ts (5 files) - -**Entities created: 47** -- Location: .planning/intel/entities/ -- Graph database: Updated automatically -- Summary: Regenerated with dependency hotspots - -Files created: -- .planning/intel/index.json -- .planning/intel/conventions.json -- .planning/intel/summary.md -- .planning/intel/entities/*.md (47 files) -- .planning/intel/graph.db - -Next: Intel hooks will continue incremental updates as you code. -``` - -### Entity Generator Subagent Definition (New File) - -```markdown ---- -name: gsd-entity-generator -description: Generates semantic entity documentation for codebase files. Spawned by analyze-codebase with file list. Writes entities directly to disk. -tools: Read, Write, Bash -color: cyan ---- - - -You are a GSD entity generator. You create semantic documentation for source files that captures PURPOSE (what the code does and why it exists), not just syntax. - -You are spawned by `/gsd:analyze-codebase` with a list of file paths. - -Your job: Read each file, analyze its purpose, write entity markdown to `.planning/intel/entities/`, return statistics only. - - - - - -Extract file paths from your prompt. You'll receive: -- Total file count -- Output directory path -- Slug convention rules -- Entity template -- List of absolute file paths - -Parse file paths into an array for processing. - - - -For each file path: - -1. **Read file content:** - ```bash - Read(file_path) - ``` - -2. **Analyze the file:** - - What is the purpose? (Why does this file exist?) - - What does it export? (Functions, classes, types) - - What does it import? (Dependencies and why) - - What type of module is it? (api, component, util, service, etc.) - -3. **Generate slug:** - - Remove leading slashes - - Remove file extension - - Replace / and . with - - - Lowercase everything - - Example: `src/lib/db.ts` → `src-lib-db` - -4. **Build entity content using template:** - - Frontmatter with path, type, date, status - - Purpose section (1-3 sentences) - - Exports section (signatures + descriptions) - - Dependencies section ([[wiki-links]] for internal, plain text for external) - - Used By: Always "TBD" (graph analysis fills this later) - - Notes: Optional (only if important context) - -5. **Write entity file:** - ```bash - Write( - file_path=f".planning/intel/entities/{slug}.md", - content=entity_markdown - ) - ``` - -6. **Track statistics:** - - Count files processed - - Count entities created - - Track any errors - -**Important:** PostToolUse hook automatically syncs entity to graph.db. You don't need to touch the graph. - - - -After all files processed, return ONLY statistics. Do NOT include entity contents. - -Format: -``` -## ENTITY GENERATION COMPLETE - -**Files processed:** {N} -**Entities created:** {M} -**Already existed:** {K} -**Errors:** {E} - -Entities written to: .planning/intel/entities/ -``` - -If errors occurred, list file paths that failed (not the error messages themselves). - - - - - -Use this EXACT format for every entity: - -```markdown ---- -path: {absolute_path} -type: [module|component|util|config|api|hook|service|model|test] -updated: {YYYY-MM-DD} -status: active ---- - -# {filename} - -## Purpose - -[1-3 sentences: What does this file do? Why does it exist? What problem does it solve? Focus on the "why", not implementation details.] - -## Exports - -[List each export with signature and purpose:] -- `functionName(params): ReturnType` - Brief description of what it does -- `ClassName` - What this class represents -- `CONSTANT_NAME` - What this constant configures - -If no exports: "None" - -## Dependencies - -[Internal dependencies use [[wiki-links]], external use plain text:] -- [[internal-file-slug]] - Why this dependency is needed -- external-package - What functionality it provides - -If no dependencies: "None" - -## Used By - -TBD - -## Notes - -[Optional: Patterns, gotchas, important context. Omit section if nothing notable.] -``` - - - -Determine entity type from file path and content: - -| Type | Indicators | -|------|-----------| -| api | In api/, routes/, endpoints/ directory, exports route handlers | -| component | In components/, exports React/Vue/etc components | -| util | In utils/, lib/, helpers/, exports utility functions | -| config | In config/, *.config.*, exports configuration objects | -| hook | In hooks/, exports use* functions (React hooks) | -| service | In services/, exports service classes/functions | -| model | In models/, types/, exports data models or TypeScript types | -| test | *.test.*, *.spec.*, contains test suites | -| module | Default if unclear, general-purpose module | - - - -**Internal dependencies** (files in the codebase): -- Convert to slug format -- Wrap in [[double brackets]] -- Example: Import from `../../lib/db.ts` → Dependency: `[[src-lib-db]]` - -**External dependencies** (npm packages): -- Plain text, no brackets -- Example: `import { z } from 'zod'` → Dependency: `zod - Schema validation` - -**When unsure if internal/external:** -- If import path starts with `.` or `@/` → internal (wiki-link) -- If import path is package name → external (plain text) - - - -Entity generation complete when: - -- [ ] All file paths processed -- [ ] Each entity file written to `.planning/intel/entities/` -- [ ] Entity markdown follows template exactly -- [ ] Frontmatter is valid YAML -- [ ] Purpose section is substantive (not just "This file exports X") -- [ ] Internal dependencies use [[wiki-links]] -- [ ] Statistics returned (not entity contents) - -``` - -## State of the Art - -| Old Approach | Current Approach | When Changed | Impact | -|--------------|------------------|--------------|--------| -| Orchestrator generates entities | Subagent delegation | Phase 5 (planned) | Prevents context exhaustion on large codebases | -| Sequential processing in main context | Fresh 200k subagent context | Phase 5 (planned) | Scales to 500+ files without orchestrator bloat | -| Batches of 10 with multiple Task calls | Single subagent, all files | Phase 5 (planned) | Simpler orchestration, no batch coordination | -| Hook generates entities via `claude -p` | Subagent generates entities | Phase 5 (planned) | Richer semantic analysis (subagent has full context vs hook's one-shot) | - -**Deprecated/outdated:** -- **Hook-based entity generation** (Phase 4 uses `execSync('claude -p')` in hook): Limited to 30s timeout, no retry, crude prompt passing. Phase 5 moves to proper subagent pattern. -- **Multiple parallel subagents** for entity batches: Overcomplicated, race condition risks, excessive overhead. Single subagent with full file list is cleaner. - -## Open Questions - -### 1. **Optimal File Count per Subagent** - -- **What we know:** Subagent has 200k context. File path = ~50 chars avg. 500 paths = 25KB (negligible). -- **What's unclear:** At what file count does entity generation hit subagent context limits? 500 files? 1000? -- **Recommendation:** Start with single subagent for up to 500 files. If codebases >500 common, implement chunking (spawn 5 subagents of 100 files each, sequentially). - -### 2. **Should Orchestrator Pre-read Index Data?** - -- **What we know:** Subagent needs to understand codebase conventions (naming patterns, directory purposes) for better entity generation. -- **What's unclear:** Should orchestrator pass conventions.json content in prompt, or should subagent read it? -- **Recommendation:** Pass conventions in prompt (it's <2KB JSON). Saves subagent a Read operation, provides useful context for entity type classification. - -### 3. **Entity Regeneration Strategy** - -- **What we know:** Hook regenerates entities when file signature changes (exports/imports differ). -- **What's unclear:** Should bulk regeneration (via /gsd:analyze-codebase) skip existing entities or overwrite? -- **Recommendation:** Skip existing entities by default (check if `.planning/intel/entities/{slug}.md` exists). Add flag: `--force-regenerate` to overwrite all. This prevents destroying manual edits to entities. - -### 4. **Error Handling for Unparseable Files** - -- **What we know:** Some files might be binary, corrupted, or have syntax errors. -- **What's unclear:** Should subagent skip silently, or return error list? -- **Recommendation:** Try-catch around file reading. Skip unparseable files, track in errors list, report at end. Don't let one bad file block entire batch. - -## Sources - -### Primary (HIGH confidence) - -- [commands/gsd/analyze-codebase.md](file://./commands/gsd/analyze-codebase.md) - Current entity generation implementation (Phase 4) -- [agents/gsd-codebase-mapper.md](file://./agents/gsd-codebase-mapper.md) - Subagent delegation pattern reference -- [agents/gsd-executor.md](file://./agents/gsd-executor.md) - Task tool usage with model profiles -- [commands/gsd/execute-phase.md](file://./commands/gsd/execute-phase.md) - Wave-based parallel execution pattern -- [hooks/gsd-intel-index.js](file://./hooks/gsd-intel-index.js) - PostToolUse hook that syncs entities to graph - -### Secondary (MEDIUM confidence) - -- [Claude Code documentation on Task tool](https://docs.anthropic.com/claude/docs/claude-code) - Official docs on subagent spawning -- [Phase 4 research findings](.planning/phases/04-semantic-intelligence/04-RESEARCH.md) - Context window management patterns - -### Tertiary (LOW confidence) - -- None - all findings based on existing codebase analysis - -## Metadata - -**Confidence breakdown:** -- Standard stack: HIGH - Task tool is official Claude Code primitive, sql.js already in use -- Architecture: HIGH - Pattern directly mirrors gsd-codebase-mapper.md (proven) -- Pitfalls: HIGH - Based on direct code analysis and understanding of context limits -- Open questions: MEDIUM - Edge cases identifiable but not yet tested at scale - -**Research date:** 2026-01-20 -**Valid until:** 60 days (stable Claude Code APIs, no fast-moving dependencies) - -**Critical insight:** Phase 5 is an architectural refactor, not a feature addition. The goal is context preservation, not new capabilities. Success = same output with less orchestrator context usage. diff --git a/artifacts/research/2025-01-19-codebase-intelligence-system.md b/artifacts/research/2025-01-19-codebase-intelligence-system.md deleted file mode 100644 index b82edb07c..000000000 --- a/artifacts/research/2025-01-19-codebase-intelligence-system.md +++ /dev/null @@ -1,1794 +0,0 @@ -# Codebase Intelligence System - -## Vision - -A living, breathing knowledge system that learns your codebase patterns as you build, provides real-time advisory during execution, and injects accumulated wisdom into every Claude session—without API calls, without manual maintenance, without going stale. - -**The experience:** By session 5, Claude knows where services go, what naming conventions you use, and what patterns have emerged—because it learned from watching you build. - ---- - -## Table of Contents - -1. [Problem Statement](#problem-statement) -2. [Design Principles](#design-principles) -3. [Architecture Overview](#architecture-overview) -4. [Hook System](#hook-system) -5. [Data Structures](#data-structures) -6. [Pattern Detection Engine](#pattern-detection-engine) -7. [Convention Learning](#convention-learning) -8. [GSD Integration](#gsd-integration) -9. [User Experience](#user-experience) -10. [Implementation Plan](#implementation-plan) -11. [Future Extensions](#future-extensions) - ---- - -## Problem Statement - -### Current State: `map-codebase` - -The existing GSD approach to codebase understanding has fundamental limitations: - -| Issue | Impact | -|-------|--------| -| **Point-in-time snapshots** | 7 markdown docs generated once, immediately stale | -| **Human-readable but not machine-actionable** | Claude re-parses prose every time | -| **Passive consumption** | Docs sit idle; planners/executors may not load them | -| **No feedback loop** | Learnings during execution don't flow back | -| **Heavy upfront cost** | 4 parallel agents do deep analysis before any code exists | -| **Brownfield-first design** | Assumes existing code to analyze; awkward for greenfield | - -### The Mental Model Problem - -`map-codebase` treats codebase understanding as a **document generation task**—a one-time survey that produces static artifacts. - -The correct mental model: codebase understanding is a **living knowledge system** that evolves with the code. - -### User Feedback - -Users report `map-codebase` isn't effective because: -- Output is verbose but not actionable -- Planners don't consistently use the generated docs -- No way to correct or refine the understanding -- Manual refresh required after significant changes -- Greenfield projects get no benefit until substantial code exists - ---- - -## Design Principles - -### 1. Build Incrementally, Not Upfront - -For greenfield projects (primary GSD use case), there's nothing to map at start. Intelligence should build as code is written: - -``` -Phase 1: Creates src/services/user.service.ts -Phase 2: Creates src/services/auth.service.ts -Phase 3: Claude should KNOW "services go in src/services/*.service.ts" -``` - -### 2. Zero API Calls for Core Functionality - -The intelligence system must work: -- Offline -- Without cost -- Without latency -- Without external dependencies - -Core operations use only: -- Tree-sitter (local AST parsing) -- JSON file manipulation -- Algorithmic pattern detection - -### 3. Hooks as Infrastructure - -Claude Code's hook system provides the integration points: -- **PostToolUse**: Observe what Claude writes -- **PreToolUse**: Advise before Claude writes -- **SessionStart**: Inject accumulated knowledge - -No custom tooling required—leverage existing infrastructure. - -### 4. Advisory, Not Blocking - -Following GSD philosophy ("no checkpoints for automatable work"), convention checks are advisory: -- Show the deviation -- Let Claude proceed -- Record the decision -- Learn from exceptions - -### 5. Confidence-Based Learning - -Not all patterns are equal: -- 2 files matching: Coincidence (no rule) -- 3 files matching: Tentative pattern (60% confidence) -- 5+ files matching: Established convention (80%+ confidence) -- Exceptions reduce confidence; consistency increases it - -### 6. Greenfield-First, Brownfield-Compatible - -Optimize for projects starting from scratch. Support existing codebases through optional deep scan. - ---- - -## Architecture Overview - -``` -┌─────────────────────────────────────────────────────────────────────────┐ -│ Codebase Intelligence System │ -├─────────────────────────────────────────────────────────────────────────┤ -│ │ -│ ┌────────────────┐ ┌────────────────────┐ │ -│ │ Claude │ │ intel-index.js │ │ -│ │ Session │ │ (PostToolUse) │ │ -│ │ │ │ │ │ -│ │ Write tool ───┼─── PostToolUse hook ────► │ • Parse AST │ │ -│ │ executes │ (async) │ • Extract exports │ │ -│ │ │ │ • Update indices │ │ -│ │ │ │ • Detect patterns │ │ -│ └────────────────┘ └─────────┬──────────┘ │ -│ │ │ │ -│ │ ▼ │ -│ │ ┌───────────────────────────┐ │ -│ │ │ .planning/intel/ │ │ -│ │ │ │ │ -│ │ │ ┌─────────────────────┐ │ │ -│ │ │ │ structure.json │ │ │ -│ │ │ │ (files, exports) │ │ │ -│ │ │ └─────────────────────┘ │ │ -│ │ │ │ │ -│ │ Query │ ┌─────────────────────┐ │ │ -│ │ │ │ conventions.json │ │ │ -│ │ │ │ (rules+confidence) │ │ │ -│ │ │ └─────────────────────┘ │ │ -│ │ │ │ │ -│ │ │ ┌─────────────────────┐ │ │ -│ │ │ │ summary.md │ │ │ -│ │ │ │ (human-readable) │ │ │ -│ │ │ └─────────────────────┘ │ │ -│ │ │ │ │ -│ │ │ ┌─────────────────────┐ │ │ -│ │ │ │ semantic.db │ │ │ -│ │ │ │ (optional) │ │ │ -│ │ │ └─────────────────────┘ │ │ -│ │ └───────────────────────────┘ │ -│ │ ▲ │ -│ │ │ │ -│ ┌───────┴────────┐ ┌────────┴───────────┐ │ -│ │ Claude │ │ intel-advise.js │ │ -│ │ Session │ │ (PreToolUse) │ │ -│ │ │ │ │ │ -│ │ Write tool ◄──┼─── PreToolUse hook ────── │ • Check path │ │ -│ │ about to │ (sync) │ • Match patterns │ │ -│ │ execute │ │ • Return advisory │ │ -│ │ │ │ │ │ -│ └────────────────┘ └────────────────────┘ │ -│ ▲ │ -│ │ │ -│ ┌───────┴────────┐ ┌────────────────────┐ │ -│ │ Session │ │ intel-init.js │ │ -│ │ Starts │◄── SessionStart hook ──── │ (SessionStart) │ │ -│ │ │ │ │ │ -│ │ Context now │ injects │ • Read summary.md │ │ -│ │ includes │ │ • Output to ctx │ │ -│ │ conventions │ │ │ │ -│ └────────────────┘ └────────────────────┘ │ -│ │ -└─────────────────────────────────────────────────────────────────────────┘ -``` - -### Data Flow - -``` -┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ -│ Write │ │ Parse │ │ Update │ │ Regenerate │ -│ File │────►│ AST │────►│ Indices │────►│ Summary │ -└─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘ - │ │ - │ │ - ▼ ▼ - ┌─────────────┐ ┌─────────────┐ - │ Extract │ │ Detect │ - │ Exports │ │ Patterns │ - └─────────────┘ └─────────────┘ - │ - ▼ - ┌─────────────┐ - │ Update │ - │ Confidence │ - └─────────────┘ -``` - ---- - -## Hook System - -### Overview - -Three hooks work together to create the intelligence loop: - -| Hook | Trigger | Mode | Purpose | -|------|---------|------|---------| -| `intel-index.js` | PostToolUse (Write/Edit) | Async | Index new code, detect patterns | -| `intel-advise.js` | PreToolUse (Write) | Sync | Advisory on convention deviations | -| `intel-init.js` | SessionStart | Sync | Inject knowledge into context | - -### Hook 1: intel-index.js (PostToolUse) - -**Trigger:** After every Write or Edit tool execution - -**Purpose:** Extract structure and update indices - -**Process:** -1. Receive tool execution details via stdin (JSON) -2. Skip non-code files -3. Parse file content with tree-sitter -4. Extract exports, imports, patterns -5. Update `structure.json` -6. Run pattern detection -7. Update `conventions.json` if new patterns found -8. Regenerate `summary.md` - -**Performance:** Async, non-blocking. Can take 100-500ms without affecting Claude. - -```javascript -#!/usr/bin/env node -// ~/.claude/hooks/intel-index.js - -const fs = require('fs'); -const path = require('path'); -const Parser = require('tree-sitter'); -const TypeScript = require('tree-sitter-typescript').typescript; -const JavaScript = require('tree-sitter-javascript'); -const Python = require('tree-sitter-python'); - -// Language configuration -const LANGUAGES = { - '.ts': TypeScript, - '.tsx': TypeScript, - '.js': JavaScript, - '.jsx': JavaScript, - '.py': Python, -}; - -const CODE_EXTENSIONS = Object.keys(LANGUAGES); - -// Main execution -let input = ''; -process.stdin.setEncoding('utf8'); -process.stdin.on('data', chunk => input += chunk); -process.stdin.on('end', async () => { - try { - await processHookEvent(JSON.parse(input)); - } catch (err) { - // Silent failure - never break Claude's workflow - if (process.env.INTEL_DEBUG) { - console.error('Intel indexer error:', err.message); - } - } - process.exit(0); -}); - -async function processHookEvent(event) { - // Only process Write and Edit tools - if (!['Write', 'Edit'].includes(event.tool_name)) { - return; - } - - const filePath = event.tool_input.file_path; - const content = event.tool_input.content || event.tool_input.new_string; - - // Skip non-code files - if (!isCodeFile(filePath)) { - return; - } - - // Find project root - const projectRoot = findProjectRoot(filePath); - if (!projectRoot) { - return; // Not a GSD project - } - - const intelDir = path.join(projectRoot, '.planning', 'intel'); - ensureIntelDir(intelDir); - - // For Edit tool, we need to read the full file content - let fullContent = content; - if (event.tool_name === 'Edit') { - try { - fullContent = fs.readFileSync(filePath, 'utf8'); - } catch { - return; // File doesn't exist yet or read error - } - } - - // Extract information from the file - const fileInfo = extractFileInfo(filePath, fullContent); - - // Update indices - updateStructure(intelDir, filePath, fileInfo); - updateConventions(intelDir); - regenerateSummary(intelDir); -} - -function isCodeFile(filePath) { - return CODE_EXTENSIONS.includes(path.extname(filePath)); -} - -function findProjectRoot(filePath) { - let dir = path.isAbsolute(filePath) ? path.dirname(filePath) : process.cwd(); - const maxDepth = 10; - let depth = 0; - - while (dir !== '/' && depth < maxDepth) { - if (fs.existsSync(path.join(dir, '.planning'))) { - return dir; - } - dir = path.dirname(dir); - depth++; - } - return null; -} - -function ensureIntelDir(intelDir) { - if (!fs.existsSync(intelDir)) { - fs.mkdirSync(intelDir, { recursive: true }); - } -} - -function extractFileInfo(filePath, content) { - const ext = path.extname(filePath); - const language = LANGUAGES[ext]; - - if (!language) { - return { - exports: [], - imports: [], - directory: path.dirname(filePath), - filename: path.basename(filePath), - relativePath: filePath, - }; - } - - const parser = new Parser(); - parser.setLanguage(language); - - let tree; - try { - tree = parser.parse(content); - } catch { - return { - exports: [], - imports: [], - directory: path.dirname(filePath), - filename: path.basename(filePath), - relativePath: filePath, - }; - } - - return { - exports: extractExports(tree.rootNode, ext), - imports: extractImports(tree.rootNode, ext), - directory: path.dirname(filePath), - filename: path.basename(filePath), - relativePath: filePath, - lineCount: content.split('\n').length, - }; -} - -function extractExports(rootNode, ext) { - const exports = []; - - // TypeScript/JavaScript exports - if (['.ts', '.tsx', '.js', '.jsx'].includes(ext)) { - walkTree(rootNode, node => { - // export class Foo - if (node.type === 'export_statement') { - const declaration = node.childForFieldName('declaration'); - if (declaration) { - if (declaration.type === 'class_declaration') { - const name = declaration.childForFieldName('name'); - if (name) { - exports.push({ - name: name.text, - type: 'class', - line: node.startPosition.row + 1, - }); - } - } else if (declaration.type === 'function_declaration') { - const name = declaration.childForFieldName('name'); - if (name) { - exports.push({ - name: name.text, - type: 'function', - line: node.startPosition.row + 1, - }); - } - } else if (declaration.type === 'lexical_declaration') { - // export const foo = ... - for (const child of declaration.children) { - if (child.type === 'variable_declarator') { - const name = child.childForFieldName('name'); - if (name) { - exports.push({ - name: name.text, - type: 'const', - line: node.startPosition.row + 1, - }); - } - } - } - } - } - } - - // export default - if (node.type === 'export_default_declaration') { - exports.push({ - name: 'default', - type: 'default', - line: node.startPosition.row + 1, - }); - } - }); - } - - // Python exports (functions and classes at module level) - if (ext === '.py') { - for (const child of rootNode.children) { - if (child.type === 'class_definition') { - const name = child.childForFieldName('name'); - if (name && !name.text.startsWith('_')) { - exports.push({ - name: name.text, - type: 'class', - line: child.startPosition.row + 1, - }); - } - } else if (child.type === 'function_definition') { - const name = child.childForFieldName('name'); - if (name && !name.text.startsWith('_')) { - exports.push({ - name: name.text, - type: 'function', - line: child.startPosition.row + 1, - }); - } - } - } - } - - return exports; -} - -function extractImports(rootNode, ext) { - const imports = []; - - if (['.ts', '.tsx', '.js', '.jsx'].includes(ext)) { - walkTree(rootNode, node => { - if (node.type === 'import_statement') { - const source = node.childForFieldName('source'); - if (source) { - imports.push({ - source: source.text.replace(/['"]/g, ''), - line: node.startPosition.row + 1, - }); - } - } - }); - } - - if (ext === '.py') { - walkTree(rootNode, node => { - if (node.type === 'import_statement' || node.type === 'import_from_statement') { - const moduleName = node.childForFieldName('module_name'); - if (moduleName) { - imports.push({ - source: moduleName.text, - line: node.startPosition.row + 1, - }); - } - } - }); - } - - return imports; -} - -function walkTree(node, callback) { - callback(node); - for (const child of node.children) { - walkTree(child, callback); - } -} - -function updateStructure(intelDir, filePath, fileInfo) { - const structurePath = path.join(intelDir, 'structure.json'); - let structure = loadJSON(structurePath, { - version: 1, - lastUpdated: null, - directories: {}, - files: {}, - exports: {}, - }); - - // Get relative path from project root - const projectRoot = path.dirname(path.dirname(intelDir)); - const relativePath = path.relative(projectRoot, filePath); - const relativeDir = path.dirname(relativePath); - - // Update directory tracking - if (!structure.directories[relativeDir]) { - structure.directories[relativeDir] = { - count: 0, - files: [], - patterns: [], - }; - } - - const dirInfo = structure.directories[relativeDir]; - if (!dirInfo.files.includes(fileInfo.filename)) { - dirInfo.files.push(fileInfo.filename); - dirInfo.count = dirInfo.files.length; - } - - // Update file tracking - structure.files[relativePath] = { - exports: fileInfo.exports.map(e => e.name), - imports: fileInfo.imports.map(i => i.source), - lineCount: fileInfo.lineCount, - lastModified: new Date().toISOString(), - }; - - // Update exports index - for (const exp of fileInfo.exports) { - structure.exports[exp.name] = { - file: relativePath, - type: exp.type, - line: exp.line, - }; - } - - structure.lastUpdated = new Date().toISOString(); - saveJSON(structurePath, structure); -} - -function updateConventions(intelDir) { - const structurePath = path.join(intelDir, 'structure.json'); - const conventionsPath = path.join(intelDir, 'conventions.json'); - - const structure = loadJSON(structurePath, { directories: {} }); - let conventions = loadJSON(conventionsPath, { - version: 1, - lastUpdated: null, - rules: [], - exceptions: [], - }); - - // Detect patterns in each directory - for (const [dir, info] of Object.entries(structure.directories)) { - if (info.count < 3) continue; // Need 3+ files for pattern - - const pattern = detectNamingPattern(info.files); - if (!pattern) continue; - - const ruleId = generateRuleId(dir, pattern); - const existingIndex = conventions.rules.findIndex(r => r.id === ruleId); - - const rule = { - id: ruleId, - directory: dir, - pattern: pattern.pattern, - patternType: pattern.type, - confidence: calculateConfidence(info.files.length, pattern), - examples: info.files.slice(0, 10), // Keep max 10 examples - fileCount: info.count, - lastUpdated: new Date().toISOString(), - }; - - if (existingIndex >= 0) { - // Preserve exceptions from existing rule - rule.exceptions = conventions.rules[existingIndex].exceptions || []; - conventions.rules[existingIndex] = rule; - } else { - rule.exceptions = []; - rule.learned = new Date().toISOString(); - conventions.rules.push(rule); - } - } - - // Sort by confidence (highest first) - conventions.rules.sort((a, b) => b.confidence - a.confidence); - conventions.lastUpdated = new Date().toISOString(); - - saveJSON(conventionsPath, conventions); -} - -function detectNamingPattern(files) { - if (files.length < 3) return null; - - // Analyze extensions - const extCounts = {}; - for (const file of files) { - const ext = path.extname(file); - extCounts[ext] = (extCounts[ext] || 0) + 1; - } - - // Find dominant extension - const dominantExt = Object.entries(extCounts) - .sort((a, b) => b[1] - a[1])[0]; - - if (!dominantExt || dominantExt[1] < files.length * 0.7) { - return null; // No consistent extension - } - - const ext = dominantExt[0]; - const filesWithExt = files.filter(f => f.endsWith(ext)); - const bases = filesWithExt.map(f => path.basename(f, ext)); - - // Check for suffix patterns (*.service, *.controller, *.test, etc.) - const suffixPattern = detectSuffixPattern(bases, ext); - if (suffixPattern) return suffixPattern; - - // Check for casing patterns - const casingPattern = detectCasingPattern(bases, ext); - if (casingPattern) return casingPattern; - - // Fallback: just the extension - return { - type: 'extension', - pattern: `*${ext}`, - matchRate: filesWithExt.length / files.length, - }; -} - -function detectSuffixPattern(bases, ext) { - // Look for common suffixes like .service, .controller, .test, .spec - const suffixCounts = {}; - - for (const base of bases) { - const parts = base.split('.'); - if (parts.length >= 2) { - const suffix = parts[parts.length - 1]; - suffixCounts[suffix] = (suffixCounts[suffix] || 0) + 1; - } - } - - const dominantSuffix = Object.entries(suffixCounts) - .filter(([suffix]) => suffix.length > 1) // Ignore single-char suffixes - .sort((a, b) => b[1] - a[1])[0]; - - if (dominantSuffix && dominantSuffix[1] >= bases.length * 0.7) { - return { - type: 'suffix', - pattern: `*.${dominantSuffix[0]}${ext}`, - suffix: dominantSuffix[0], - matchRate: dominantSuffix[1] / bases.length, - }; - } - - return null; -} - -function detectCasingPattern(bases, ext) { - const patterns = { - PascalCase: /^[A-Z][a-zA-Z0-9]*$/, - camelCase: /^[a-z][a-zA-Z0-9]*$/, - 'kebab-case': /^[a-z][a-z0-9]*(-[a-z0-9]+)*$/, - snake_case: /^[a-z][a-z0-9]*(_[a-z0-9]+)*$/, - SCREAMING_SNAKE: /^[A-Z][A-Z0-9]*(_[A-Z0-9]+)*$/, - }; - - for (const [name, regex] of Object.entries(patterns)) { - const matches = bases.filter(b => regex.test(b)); - if (matches.length >= bases.length * 0.8) { - return { - type: 'casing', - pattern: `${name}${ext}`, - casing: name, - matchRate: matches.length / bases.length, - }; - } - } - - return null; -} - -function generateRuleId(dir, pattern) { - const dirSlug = dir.replace(/[^a-zA-Z0-9]/g, '-').replace(/-+/g, '-'); - const patternSlug = pattern.pattern.replace(/[^a-zA-Z0-9]/g, '-').replace(/-+/g, '-'); - return `${dirSlug}-${patternSlug}`.toLowerCase(); -} - -function calculateConfidence(fileCount, pattern) { - // Base confidence from file count - // 3 files: 0.60, 5 files: 0.70, 10 files: 0.85, 20+: 0.95 - let confidence = Math.min(0.95, 0.50 + (fileCount * 0.025)); - - // Adjust for pattern match rate - confidence *= pattern.matchRate; - - // Suffix patterns are more reliable than casing patterns - if (pattern.type === 'suffix') { - confidence = Math.min(0.98, confidence * 1.1); - } - - return Math.round(confidence * 100) / 100; -} - -function regenerateSummary(intelDir) { - const structurePath = path.join(intelDir, 'structure.json'); - const conventionsPath = path.join(intelDir, 'conventions.json'); - const summaryPath = path.join(intelDir, 'summary.md'); - - const structure = loadJSON(structurePath, { directories: {}, exports: {} }); - const conventions = loadJSON(conventionsPath, { rules: [] }); - - let summary = `# Codebase Intelligence Summary\n\n`; - summary += `*Auto-generated by GSD. Last updated: ${new Date().toISOString()}*\n\n`; - - // High-confidence conventions - const highConfidence = conventions.rules.filter(r => r.confidence >= 0.70); - if (highConfidence.length > 0) { - summary += `## Detected Conventions\n\n`; - for (const rule of highConfidence) { - const pct = Math.round(rule.confidence * 100); - summary += `- **\`${rule.directory}/\`**: \`${rule.pattern}\` `; - summary += `(${rule.fileCount} files, ${pct}% confidence)\n`; - - if (rule.exceptions && rule.exceptions.length > 0) { - summary += ` - Exceptions: ${rule.exceptions.map(e => e.file).join(', ')}\n`; - } - } - summary += '\n'; - } - - // Lower confidence patterns (for awareness) - const lowConfidence = conventions.rules.filter(r => r.confidence >= 0.50 && r.confidence < 0.70); - if (lowConfidence.length > 0) { - summary += `## Emerging Patterns\n\n`; - for (const rule of lowConfidence) { - const pct = Math.round(rule.confidence * 100); - summary += `- \`${rule.directory}/\`: \`${rule.pattern}\` (${rule.fileCount} files, ${pct}%)\n`; - } - summary += '\n'; - } - - // Key exports (grouped by type) - const exportsByType = {}; - for (const [name, info] of Object.entries(structure.exports)) { - if (name === 'default') continue; - if (!exportsByType[info.type]) { - exportsByType[info.type] = []; - } - exportsByType[info.type].push(name); - } - - if (Object.keys(exportsByType).length > 0) { - summary += `## Key Exports\n\n`; - for (const [type, names] of Object.entries(exportsByType)) { - const displayNames = names.slice(0, 15).join(', '); - const more = names.length > 15 ? ` (+${names.length - 15} more)` : ''; - summary += `- **${type}s**: ${displayNames}${more}\n`; - } - summary += '\n'; - } - - // Active directories - const activeDirs = Object.entries(structure.directories) - .filter(([_, info]) => info.count >= 2) - .sort((a, b) => b[1].count - a[1].count) - .slice(0, 15); - - if (activeDirs.length > 0) { - summary += `## Directory Structure\n\n`; - for (const [dir, info] of activeDirs) { - summary += `- \`${dir}/\` — ${info.count} files\n`; - } - summary += '\n'; - } - - // Statistics - const totalFiles = Object.keys(structure.files).length; - const totalExports = Object.keys(structure.exports).length; - const totalDirs = Object.keys(structure.directories).length; - - summary += `## Statistics\n\n`; - summary += `- Files indexed: ${totalFiles}\n`; - summary += `- Exports tracked: ${totalExports}\n`; - summary += `- Directories: ${totalDirs}\n`; - summary += `- Conventions detected: ${conventions.rules.length}\n`; - - fs.writeFileSync(summaryPath, summary); -} - -function loadJSON(filePath, defaultValue) { - try { - if (fs.existsSync(filePath)) { - return JSON.parse(fs.readFileSync(filePath, 'utf8')); - } - } catch { - // Corrupted file, reset to default - } - return defaultValue; -} - -function saveJSON(filePath, data) { - fs.writeFileSync(filePath, JSON.stringify(data, null, 2)); -} -``` - -### Hook 2: intel-advise.js (PreToolUse) - -**Trigger:** Before Write tool execution - -**Purpose:** Provide advisory if file placement deviates from conventions - -**Process:** -1. Receive proposed file path -2. Check against high-confidence conventions -3. If deviation detected, return advisory message -4. Always allow proceed (never block) - -**Performance:** Sync, must complete in <1000ms. Simple JSON lookup. - -```javascript -#!/usr/bin/env node -// ~/.claude/hooks/intel-advise.js - -const fs = require('fs'); -const path = require('path'); - -let input = ''; -process.stdin.setEncoding('utf8'); -process.stdin.on('data', chunk => input += chunk); -process.stdin.on('end', () => { - try { - const result = processAdvisory(JSON.parse(input)); - console.log(JSON.stringify(result)); - } catch (err) { - // On any error, just allow proceed - console.log(JSON.stringify({ proceed: true })); - } - process.exit(0); -}); - -function processAdvisory(event) { - // Only advise on Write tool - if (event.tool_name !== 'Write') { - return { proceed: true }; - } - - const filePath = event.tool_input.file_path; - - // Skip non-code files - if (!isCodeFile(filePath)) { - return { proceed: true }; - } - - const projectRoot = findProjectRoot(filePath); - if (!projectRoot) { - return { proceed: true }; - } - - const conventionsPath = path.join(projectRoot, '.planning', 'intel', 'conventions.json'); - if (!fs.existsSync(conventionsPath)) { - return { proceed: true }; - } - - let conventions; - try { - conventions = JSON.parse(fs.readFileSync(conventionsPath, 'utf8')); - } catch { - return { proceed: true }; - } - - const advisory = checkConventions(filePath, projectRoot, conventions); - - if (advisory) { - return { - proceed: true, // Never block - message: advisory, - }; - } - - return { proceed: true }; -} - -function checkConventions(filePath, projectRoot, conventions) { - const relativePath = path.relative(projectRoot, filePath); - const relativeDir = path.dirname(relativePath); - const filename = path.basename(filePath); - - // Only check high-confidence rules - const activeRules = conventions.rules.filter(r => r.confidence >= 0.70); - - for (const rule of activeRules) { - // Check if file SHOULD follow this rule but is in wrong location - if (fileMatchesPattern(filename, rule)) { - if (relativeDir !== rule.directory && !relativeDir.startsWith(rule.directory + '/')) { - // Check if this is a known exception - const isException = (rule.exceptions || []).some( - e => e.file === relativePath || e.pattern && filename.match(new RegExp(e.pattern)) - ); - - if (!isException) { - const pct = Math.round(rule.confidence * 100); - return `📍 Convention advisory: Files matching '${rule.pattern}' are typically in '${rule.directory}/' ` + - `(${rule.fileCount} files, ${pct}% confidence). ` + - `Current path: '${relativeDir}/'. Proceeding—this will be noted if intentional.`; - } - } - } - } - - // Check if writing to a directory with existing conventions - const dirRule = activeRules.find(r => - relativeDir === r.directory || relativeDir.startsWith(r.directory + '/') - ); - - if (dirRule && !fileMatchesPattern(filename, dirRule)) { - // File is in a convention directory but doesn't match pattern - const isException = (dirRule.exceptions || []).some( - e => e.file === relativePath || e.pattern && filename.match(new RegExp(e.pattern)) - ); - - if (!isException) { - const pct = Math.round(dirRule.confidence * 100); - return `📍 Convention advisory: Files in '${dirRule.directory}/' typically match '${dirRule.pattern}' ` + - `(${dirRule.fileCount} files, ${pct}% confidence). ` + - `'${filename}' doesn't match. Proceeding—this will be noted if intentional.`; - } - } - - return null; -} - -function fileMatchesPattern(filename, rule) { - const pattern = rule.pattern; - - // Suffix pattern: *.service.ts - if (pattern.startsWith('*.') && pattern.includes('.', 2)) { - const suffix = pattern.slice(1); // .service.ts - return filename.endsWith(suffix); - } - - // Extension pattern: *.ts - if (pattern.startsWith('*.')) { - const ext = pattern.slice(1); - return filename.endsWith(ext); - } - - // Casing pattern: PascalCase.tsx - if (rule.patternType === 'casing') { - const ext = path.extname(filename); - const base = path.basename(filename, ext); - - switch (rule.casing) { - case 'PascalCase': - return /^[A-Z][a-zA-Z0-9]*$/.test(base); - case 'camelCase': - return /^[a-z][a-zA-Z0-9]*$/.test(base); - case 'kebab-case': - return /^[a-z][a-z0-9]*(-[a-z0-9]+)*$/.test(base); - case 'snake_case': - return /^[a-z][a-z0-9]*(_[a-z0-9]+)*$/.test(base); - } - } - - return false; -} - -function isCodeFile(filePath) { - const codeExts = ['.ts', '.tsx', '.js', '.jsx', '.py', '.go', '.rs', '.swift', '.java', '.kt']; - return codeExts.includes(path.extname(filePath)); -} - -function findProjectRoot(filePath) { - let dir = path.isAbsolute(filePath) ? path.dirname(filePath) : process.cwd(); - const maxDepth = 10; - let depth = 0; - - while (dir !== '/' && depth < maxDepth) { - if (fs.existsSync(path.join(dir, '.planning'))) { - return dir; - } - dir = path.dirname(dir); - depth++; - } - return null; -} -``` - -### Hook 3: intel-init.js (SessionStart) - -**Trigger:** When a Claude Code session starts - -**Purpose:** Inject codebase intelligence into session context - -**Process:** -1. Find project root from cwd -2. Read `summary.md` -3. Output wrapped in `` tags -4. Claude sees this in system context - -**Performance:** Sync, should complete in <500ms. Simple file read. - -```javascript -#!/usr/bin/env node -// ~/.claude/hooks/intel-init.js - -const fs = require('fs'); -const path = require('path'); - -const projectRoot = findProjectRoot(process.cwd()); - -if (!projectRoot) { - // Not a GSD project - process.exit(0); -} - -const summaryPath = path.join(projectRoot, '.planning', 'intel', 'summary.md'); - -if (!fs.existsSync(summaryPath)) { - // No intelligence gathered yet - process.exit(0); -} - -try { - const summary = fs.readFileSync(summaryPath, 'utf8'); - - // Check if summary has meaningful content - const lines = summary.split('\n').filter(l => l.trim()); - if (lines.length < 5) { - // Summary too sparse, skip injection - process.exit(0); - } - - // Output to Claude's context - console.log(` -${summary} -`); - -} catch (err) { - // Silent failure - if (process.env.INTEL_DEBUG) { - console.error('Intel init error:', err.message); - } -} - -process.exit(0); - -function findProjectRoot(dir) { - const maxDepth = 10; - let depth = 0; - - while (dir !== '/' && depth < maxDepth) { - if (fs.existsSync(path.join(dir, '.planning'))) { - return dir; - } - dir = path.dirname(dir); - depth++; - } - return null; -} -``` - ---- - -## Data Structures - -### Directory: `.planning/intel/` - -``` -.planning/ -└── intel/ - ├── structure.json # File/export index - ├── conventions.json # Detected patterns with confidence - ├── summary.md # Human-readable summary (auto-generated) - └── semantic.db # Optional: embeddings for semantic search -``` - -### structure.json - -Tracks files, exports, and directory contents. - -```json -{ - "version": 1, - "lastUpdated": "2025-01-19T14:30:00Z", - "directories": { - "src/services": { - "count": 5, - "files": [ - "user.service.ts", - "auth.service.ts", - "email.service.ts", - "payment.service.ts", - "notification.service.ts" - ], - "patterns": [] - }, - "src/components": { - "count": 12, - "files": ["Button.tsx", "Card.tsx", "Modal.tsx", "..."], - "patterns": [] - } - }, - "files": { - "src/services/user.service.ts": { - "exports": ["UserService", "CreateUserDTO", "UpdateUserDTO"], - "imports": ["@prisma/client", "./base.service"], - "lineCount": 145, - "lastModified": "2025-01-19T14:30:00Z" - } - }, - "exports": { - "UserService": { - "file": "src/services/user.service.ts", - "type": "class", - "line": 15 - }, - "Button": { - "file": "src/components/Button.tsx", - "type": "function", - "line": 8 - } - } -} -``` - -### conventions.json - -Detected patterns with confidence scores and learning history. - -```json -{ - "version": 1, - "lastUpdated": "2025-01-19T14:30:00Z", - "rules": [ - { - "id": "src-services-service-ts", - "directory": "src/services", - "pattern": "*.service.ts", - "patternType": "suffix", - "suffix": "service", - "confidence": 0.88, - "examples": [ - "user.service.ts", - "auth.service.ts", - "email.service.ts", - "payment.service.ts", - "notification.service.ts" - ], - "fileCount": 5, - "exceptions": [], - "learned": "2025-01-15T10:00:00Z", - "lastUpdated": "2025-01-19T14:30:00Z" - }, - { - "id": "src-components-pascalcase-tsx", - "directory": "src/components", - "pattern": "PascalCase.tsx", - "patternType": "casing", - "casing": "PascalCase", - "confidence": 0.92, - "examples": ["Button.tsx", "Card.tsx", "Modal.tsx"], - "fileCount": 12, - "exceptions": [ - { - "file": "src/components/index.ts", - "reason": "barrel export file", - "added": "2025-01-16T11:00:00Z" - } - ], - "learned": "2025-01-14T09:00:00Z", - "lastUpdated": "2025-01-19T14:30:00Z" - } - ], - "exceptions": [] -} -``` - -### summary.md (Auto-generated) - -Human-readable summary injected into Claude's context. - -```markdown -# Codebase Intelligence Summary - -*Auto-generated by GSD. Last updated: 2025-01-19T14:30:00Z* - -## Detected Conventions - -- **`src/services/`**: `*.service.ts` (5 files, 88% confidence) -- **`src/components/`**: `PascalCase.tsx` (12 files, 92% confidence) - - Exceptions: index.ts -- **`src/app/api/`**: `*/route.ts` (8 files, 85% confidence) -- **`src/lib/`**: `kebab-case.ts` (6 files, 78% confidence) - -## Emerging Patterns - -- `src/hooks/`: `use*.ts` (2 files, 55%) -- `src/types/`: `*.types.ts` (2 files, 52%) - -## Key Exports - -- **classes**: UserService, AuthService, EmailService, PaymentService -- **functions**: Button, Card, Modal, Form, Input, useAuth, useUser -- **consts**: API_URL, DEFAULT_THEME, ROUTES - -## Directory Structure - -- `src/components/` — 12 files -- `src/app/api/` — 8 files -- `src/lib/` — 6 files -- `src/services/` — 5 files -- `src/hooks/` — 2 files - -## Statistics - -- Files indexed: 45 -- Exports tracked: 78 -- Directories: 12 -- Conventions detected: 6 -``` - ---- - -## Pattern Detection Engine - -### Detection Algorithm - -``` -For each directory with 3+ files: - 1. Analyze file extensions - → Find dominant extension (70%+ of files) - - 2. Check for suffix patterns - → *.service.ts, *.controller.ts, *.test.ts - → Look at filename segments split by '.' - - 3. Check for casing patterns - → PascalCase, camelCase, kebab-case, snake_case - → Analyze base filename (without extension) - - 4. Calculate confidence - → Base: 0.50 + (fileCount * 0.025) - → Adjust for pattern match rate - → Suffix patterns get 1.1x multiplier - → Cap at 0.95 (never 100% certain) -``` - -### Confidence Scoring - -| File Count | Base Confidence | -|------------|-----------------| -| 3 | 0.575 | -| 4 | 0.600 | -| 5 | 0.625 | -| 10 | 0.750 | -| 15 | 0.875 | -| 20+ | 0.950 (capped) | - -Adjustments: -- Multiply by pattern match rate (e.g., 4/5 files match = 0.8) -- Suffix patterns: +10% (more reliable) -- Exceptions: Each exception reduces confidence by 0.02 - -### Supported Patterns - -| Pattern Type | Example | Detection | -|--------------|---------|-----------| -| Suffix | `*.service.ts` | Dot-separated segments | -| Extension | `*.ts` | File extension only | -| PascalCase | `Button.tsx` | Regex: `/^[A-Z][a-zA-Z0-9]*$/` | -| camelCase | `userUtils.ts` | Regex: `/^[a-z][a-zA-Z0-9]*$/` | -| kebab-case | `user-service.ts` | Regex: `/^[a-z][a-z0-9]*(-[a-z0-9]+)*$/` | -| snake_case | `user_service.py` | Regex: `/^[a-z][a-z0-9]*(_[a-z0-9]+)*$/` | - ---- - -## Convention Learning - -### Learning Loop - -``` -┌─────────────────────────────────────────────────────────────────┐ -│ Convention Learning Loop │ -├─────────────────────────────────────────────────────────────────┤ -│ │ -│ Write file ──► PostToolUse ──► Update structure ──► Detect │ -│ │ patterns │ -│ │ │ │ -│ │ ┌──────────────────────────────────────┘ │ -│ │ ▼ │ -│ │ 3+ files match? │ -│ │ │ │ -│ │ Yes │ No │ -│ │ ▼ └──► (no rule yet) │ -│ │ │ -│ │ Create/update rule │ -│ │ │ │ -│ │ ▼ │ -│ │ Calculate confidence │ -│ │ │ │ -│ │ ▼ │ -│ │ Update conventions.json │ -│ │ │ │ -│ │ ▼ │ -│ │ Regenerate summary.md │ -│ │ │ -│ └──────────────────────────────────────────────────────── │ -│ │ -│ Next session starts ──► SessionStart ──► Inject summary │ -│ │ -│ Write new file ──► PreToolUse ──► Check conventions │ -│ │ │ │ -│ │ Advisory? │ -│ │ │ │ │ -│ │ Yes No │ -│ │ │ │ │ -│ │ Show message └──► Proceed │ -│ │ │ │ -│ │ Claude decides │ -│ │ │ │ │ -│ │ Adjust Proceed │ -│ │ path anyway │ -│ │ │ │ │ -│ │ │ PostToolUse notes │ -│ │ │ potential exception │ -│ │ │ │ │ -│ └─────────────────────┴───────────┘ │ -│ │ -└─────────────────────────────────────────────────────────────────┘ -``` - -### Exception Handling - -When a file deviates from a convention: - -1. **PreToolUse shows advisory** -2. **Claude proceeds anyway** -3. **PostToolUse records the file** -4. **If pattern persists** (same deviation 2+ times): - - Option A: Lower rule confidence - - Option B: Add as exception with inferred reason - -Future: `/gsd:intel-correct` command to manually add exceptions with reasons. - -### Confidence Evolution - -``` -Day 1: user.service.ts created - → No pattern yet (need 3+ files) - -Day 2: auth.service.ts, email.service.ts created - → Pattern detected: *.service.ts - → Confidence: 0.575 (3 files) - -Day 3: payment.service.ts created - → Confidence: 0.60 (4 files) - -Day 5: notification.service.ts created - → Confidence: 0.625 (5 files) - -Day 7: crypto.ts created in services/ (deviation) - → Advisory shown - → Claude proceeds - → Exception recorded - → Confidence: 0.605 (5/6 files match) - -Day 8: utils.ts created in services/ (another deviation) - → Advisory shown - → User manually corrects via /gsd:intel-correct - → Exception added: "utility files don't need .service suffix" - → Confidence stabilizes at 0.625 (5 services + 2 exceptions) -``` - ---- - -## GSD Integration - -### Changes to Existing Commands - -| Command | Change | -|---------|--------| -| `/gsd:new-project` | Creates `.planning/intel/` structure | -| `/gsd:execute-phase` | Hooks automatically index created files | -| `/gsd:plan-phase` | Planner sees conventions in context | -| `/gsd:map-codebase` | **Deprecated** → becomes `/gsd:analyze-codebase` | - -### New Commands - -| Command | Purpose | -|---------|---------| -| `/gsd:intel-status` | Show intelligence state, conventions, confidence | -| `/gsd:intel-refresh` | Force full re-index (after major refactor) | -| `/gsd:intel-correct` | Manually add/modify convention rules | -| `/gsd:analyze-codebase` | Deep scan for brownfield (optional) | - -### Install Script Changes - -`bin/install.js` registers hooks: - -```javascript -// Hook registration added to settings.json -{ - "hooks": { - "PostToolUse": [ - { - "matcher": "Write|Edit", - "hooks": [ - { - "type": "command", - "command": "node ~/.claude/hooks/intel-index.js", - "timeout": 5000 - } - ] - } - ], - "PreToolUse": [ - { - "matcher": "Write", - "hooks": [ - { - "type": "command", - "command": "node ~/.claude/hooks/intel-advise.js", - "timeout": 1000 - } - ] - } - ], - "SessionStart": [ - { - "hooks": [ - { - "type": "command", - "command": "node ~/.claude/hooks/intel-init.js", - "timeout": 2000 - } - ] - } - ] - } -} -``` - -### Dependencies Added - -```json -{ - "dependencies": { - "tree-sitter": "^0.21.0", - "tree-sitter-typescript": "^0.21.0", - "tree-sitter-javascript": "^0.21.0", - "tree-sitter-python": "^0.21.0" - } -} -``` - -Note: tree-sitter requires native compilation. May need to ship prebuilt binaries or use WASM versions for portability. - ---- - -## User Experience - -### Session 1 (New Project) - -``` -You: /gsd:new-project -...creates PROJECT.md, ROADMAP.md, etc... -...creates empty .planning/intel/ structure... - -You: /gsd:execute-phase 1 -...Claude creates files... - -PostToolUse (silent): - → Indexes src/services/user.service.ts - → Indexes src/services/auth.service.ts - → No patterns yet (only 2 files) - -End of session: - → structure.json has 2 services tracked - → No conventions.json rules yet -``` - -### Session 2 (Building) - -``` -You: /gsd:execute-phase 2 - -SessionStart: - → intel-init finds no summary yet (too sparse) - → Nothing injected - -...Claude creates more files... - -PostToolUse (silent): - → Indexes src/services/email.service.ts - → Pattern detected! 3 files match *.service.ts - → Rule created: src/services/*.service.ts (57% confidence) - → Summary regenerated - -End of session: - → conventions.json has first rule - → summary.md has "Detected Conventions" section -``` - -### Session 3 (Intelligence Active) - -``` -You: /gsd:execute-phase 3 - -SessionStart: - → intel-init reads summary.md - → Injects into Claude's context: - - -# Codebase Intelligence Summary - -## Detected Conventions -- **`src/services/`**: `*.service.ts` (3 files, 58% confidence) -- **`src/components/`**: `PascalCase.tsx` (5 files, 63% confidence) - -## Key Exports -- classes: UserService, AuthService, EmailService -... - - -Claude (planning): -"I need to create a payment service. Based on codebase conventions, -I'll create src/services/payment.service.ts" - -...Claude creates file... - -PostToolUse: - → Indexes payment.service.ts - → Confidence for services rule: 58% → 65% -``` - -### Session 5 (Deviation Scenario) - -``` -Claude about to write: src/services/crypto.ts - -PreToolUse: - → Checks: crypto.ts doesn't match *.service.ts - → Returns advisory - -Claude sees: -"📍 Convention advisory: Files in 'src/services/' typically match -'*.service.ts' (5 files, 78% confidence). 'crypto.ts' doesn't match. -Proceeding—this will be noted if intentional." - -Claude: "I'm placing this here because crypto utilities are used by -multiple services. Proceeding with the deviation." - -...file created... - -PostToolUse: - → Notes crypto.ts as potential exception - → Confidence slightly reduced: 78% → 76% -``` - -### Session 10 (Mature Intelligence) - -``` -SessionStart injects: - - -# Codebase Intelligence Summary - -## Detected Conventions - -- **`src/services/`**: `*.service.ts` (8 files, 91% confidence) - - Exceptions: crypto.ts, constants.ts -- **`src/components/`**: `PascalCase.tsx` (15 files, 94% confidence) -- **`src/app/api/`**: `*/route.ts` (12 files, 89% confidence) -- **`src/lib/`**: `kebab-case.ts` (6 files, 82% confidence) -- **`src/hooks/`**: `use*.ts` (4 files, 75% confidence) - -## Key Exports - -- **classes**: UserService, AuthService, EmailService, PaymentService, - OrderService, NotificationService, AnalyticsService, ReportService -- **functions**: Button, Card, Modal, Form, Input, Table, useAuth, - useUser, useCart, useOrders -- **consts**: API_URL, ROUTES, THEMES, LIMITS - -## Directory Structure - -- `src/components/` — 15 files -- `src/app/api/` — 12 files -- `src/services/` — 8 files -- `src/lib/` — 6 files -- `src/hooks/` — 4 files -- `src/types/` — 3 files - -## Statistics - -- Files indexed: 67 -- Exports tracked: 124 -- Directories: 15 -- Conventions detected: 8 - - -Claude now KNOWS: -- Where services go (and what naming to use) -- Component naming conventions -- API route structure -- That crypto.ts is a known exception - -No manual configuration. No stale documentation. -The codebase taught Claude its own patterns. -``` - ---- - -## Implementation Plan - -### Phase 1: Core Hooks (MVP) - -**Deliverables:** -- `intel-index.js` with tree-sitter parsing -- `intel-init.js` for session injection -- `structure.json` and `summary.md` generation -- Basic pattern detection (suffix + casing) - -**GSD changes:** -- Create `.planning/intel/` in `/gsd:new-project` -- Register hooks in `install.js` -- Add tree-sitter dependencies - -**Timeline:** Foundation for all other features - -### Phase 2: Advisory System - -**Deliverables:** -- `intel-advise.js` PreToolUse hook -- `conventions.json` with confidence scoring -- Deviation detection and advisory messages - -**GSD changes:** -- Register PreToolUse hook -- Update summary generation to include conventions - -**Timeline:** Adds real-time guidance - -### Phase 3: Brownfield Support - -**Deliverables:** -- `/gsd:analyze-codebase` command (replaces map-codebase) -- Batch indexing for existing codebases -- Handle large codebases efficiently - -**GSD changes:** -- Deprecate `/gsd:map-codebase` -- Add new command to `commands/gsd/` - -**Timeline:** Supports existing projects - -### Phase 4: Learning Refinement - -**Deliverables:** -- `/gsd:intel-correct` command -- Manual exception management -- Confidence adjustment from corrections - -**GSD changes:** -- Add new command -- Update conventions.json schema for manual entries - -**Timeline:** User control over learning - -### Phase 5: Semantic Search (Optional) - -**Deliverables:** -- `semantic.db` with embeddings (sqlite-vec) -- voyage-code-3 integration (or local ollama) -- "Find code that handles X" queries - -**GSD changes:** -- Optional `--deep` flag for analyze-codebase -- Query interface for planners - -**Timeline:** Advanced capability, optional - ---- - -## Future Extensions - -### 1. Cross-Project Learning - -Share conventions across related projects: -- Export conventions to `~/.gsd/conventions/` -- Import proven patterns into new projects -- "This looks like a Next.js project, applying common conventions" - -### 2. Framework Detection - -Automatically detect and apply framework conventions: -- Detect Next.js → suggest app router patterns -- Detect NestJS → suggest module/controller/service patterns -- Detect FastAPI → suggest router patterns - -### 3. Team Conventions - -For team settings: -- Central conventions repository -- Merge team rules with project-specific rules -- Higher confidence for team-wide patterns - -### 4. Architectural Patterns - -Beyond file naming: -- Detect dependency injection patterns -- Detect state management patterns -- Detect API design patterns - -### 5. Refactoring Suggestions - -Proactive improvement suggestions: -- "5 files in src/utils/ might be better as services" -- "Component X has grown large, consider splitting" -- "Inconsistent naming: authService.ts vs user.service.ts" - ---- - -## Technical Considerations - -### Performance - -| Operation | Target | Approach | -|-----------|--------|----------| -| PostToolUse indexing | <500ms | Async, non-blocking | -| PreToolUse advisory | <100ms | Simple JSON lookup | -| SessionStart injection | <200ms | File read only | -| Pattern detection | <50ms | In-memory, algorithmic | - -### Portability - -Tree-sitter requires native binaries. Options: -1. **NPM postinstall** - compile on install -2. **Prebuilt binaries** - ship for common platforms -3. **WASM** - use web-tree-sitter (slower but universal) - -Recommendation: Start with npm native, fall back to WASM. - -### Storage - -Typical project after 6 months: -- `structure.json`: ~50KB -- `conventions.json`: ~5KB -- `summary.md`: ~3KB -- Total: ~60KB (negligible) - -With semantic.db (optional): -- ~1KB per function embedded -- 500 functions = ~500KB -- Still reasonable for local storage - -### Error Handling - -All hooks follow silent failure principle: -- Never break Claude's workflow -- Log errors only if `INTEL_DEBUG=1` -- Return `{ proceed: true }` on any error -- Corrupted files reset to defaults - ---- - -## Summary - -The Codebase Intelligence System transforms GSD from a planning/execution framework into a **learning system** that understands your codebase and gets smarter over time. - -**Key innovations:** -1. **Hook-driven** - Uses Claude Code's native infrastructure -2. **Zero API calls** - Pure local computation -3. **Incremental** - Builds knowledge as you build code -4. **Advisory** - Guides without blocking -5. **Self-improving** - Learns from every session - -**The result:** By session 5, Claude knows your codebase conventions without being told. By session 10, it's an expert in your project's patterns. No manual documentation. No stale snapshots. Just accumulated wisdom that grows with your code. diff --git a/artifacts/research/2025-01-19-static-code-indexing-technical.md b/artifacts/research/2025-01-19-static-code-indexing-technical.md deleted file mode 100644 index cc2828063..000000000 --- a/artifacts/research/2025-01-19-static-code-indexing-technical.md +++ /dev/null @@ -1,360 +0,0 @@ -# Technical Research: Self-Evolving Codebase Intelligence for GSD - -## Strategic Summary - -The most mind-blowing approach combines **semantic code indexing** (tree-sitter AST + voyage-code-3 embeddings), **adaptive pattern learning** (conventions extracted and refined through feedback), and **self-improving compliance** (the system gets smarter about YOUR codebase over time). Instead of static indices that go stale, this creates a **living knowledge graph** that understands not just what code exists, but WHY it's structured that way—and enforces that understanding during execution. - -**Recommendation:** Approach 3 (Self-Evolving Intelligence) - because "mindblowing" means the system should feel like it genuinely understands your codebase and gets better over time. - -## Requirements - -- Zero external services (local-first, works offline) -- Sub-second queries (can't slow down Claude's planning/execution) -- Adaptive learning (improves from corrections without manual intervention) -- Works across language ecosystems (not just TypeScript) -- Integrates seamlessly with existing GSD workflow -- Minimal storage overhead (< 50MB for typical project) - ---- - -## Approach 1: Static JSON Indices with Stale Detection - -**How it works:** Generate JSON indices during `map-codebase`, store commit hash in metadata, refresh lazily when HEAD moves. - -**Libraries/tools:** -- tree-sitter (AST parsing) - `npm install tree-sitter tree-sitter-typescript tree-sitter-python` -- Node.js built-in fs for JSON persistence -- Git hooks via husky or manual `.git/hooks/post-commit` - -**Pros:** -- Simple to implement (~500 lines) -- No external dependencies -- Fast queries (JSON parse + lookup) -- Familiar format (developers can hand-edit) - -**Cons:** -- Indices go stale between sessions -- No semantic understanding (just structural data) -- Manual schema maintenance -- Doesn't learn from corrections - -**Best when:** You want quick wins with minimal complexity - -**Complexity:** S - ---- - -## Approach 2: Semantic Search with Local Embeddings - -**How it works:** Embed code chunks using voyage-code-3, store in sqlite-vec, enable semantic queries like "find authentication logic" or "code that validates user input." - -**Libraries/tools:** -- `voyage-code-3` via API (200M tokens free, $0.06/1M after) -- `sqlite-vec` - pure C, runs anywhere SQLite runs -- `tree-sitter` for intelligent chunking (function-level, not file-level) -- `better-sqlite3` for Node.js bindings - -**Architecture:** -``` -Code Changes → Tree-sitter chunks → voyage-code-3 embeds → sqlite-vec stores - ↓ -Query ("auth logic") → embed query → sqlite-vec similarity → ranked results -``` - -**Pros:** -- Semantic understanding ("find error handling" works even with varied naming) -- Hybrid search possible (FTS5 keywords + vector similarity) -- Local-first with cloud embeddings (best of both) -- Works across languages (voyage-code-3 trained on 80+ languages) - -**Cons:** -- Requires API calls for embedding (latency, cost at scale) -- More complex setup -- Still doesn't learn conventions automatically -- Embedding drift as model versions change - -**Best when:** You need powerful search across large/unfamiliar codebases - -**Complexity:** M - ---- - -## Approach 3: Self-Evolving Codebase Intelligence (The Mind-Blowing One) - -**How it works:** Three-layer system that builds understanding incrementally and learns from corrections. - -### Layer 1: Structural Index (Tree-sitter AST) -Fast, deterministic extraction of code structure: -- Functions, classes, exports, imports -- File patterns and naming conventions -- Directory structure and module boundaries - -### Layer 2: Semantic Memory (Embeddings + Patterns) -Understanding what code DOES, not just what it IS: -- Function-level embeddings via voyage-code-3 -- Pattern clusters detected via embedding similarity -- Convention rules extracted from consistent patterns - -### Layer 3: Adaptive Learning (The Magic) -The system improves from every interaction: -- When Claude deviates from conventions and you correct it → system learns -- When Claude asks "should this be a service or util?" and you answer → system remembers -- Confidence scores that increase with consistent patterns, decrease with exceptions - -**Architecture:** -``` -┌─────────────────────────────────────────────────────────────────┐ -│ Codebase Intelligence │ -├─────────────────────────────────────────────────────────────────┤ -│ ┌─────────────┐ ┌─────────────┐ ┌─────────────────────────┐ │ -│ │ STRUCTURE │ │ SEMANTICS │ │ LEARNED RULES │ │ -│ │ (AST/JSON) │ │ (Embeddings)│ │ (Adaptive/Scored) │ │ -│ ├─────────────┤ ├─────────────┤ ├─────────────────────────┤ │ -│ │ symbols.json│ │ sqlite-vec │ │ conventions.json │ │ -│ │ patterns.json │ + FTS5 │ │ ├─ rule: "services/*" │ │ -│ │ structure.json │ │ │ │ confidence: 0.95 │ │ -│ │ │ │ │ │ │ examples: [...] │ │ -│ │ │ │ │ │ │ exceptions: [...] │ │ -│ └─────────────┘ └─────────────┘ └─────────────────────────┘ │ -├─────────────────────────────────────────────────────────────────┤ -│ FEEDBACK LOOP │ -│ ┌──────────┐ ┌──────────┐ ┌──────────────────────────┐ │ -│ │ Executor │ → │ Deviation │ → │ User Correction/Approval │ │ -│ │ writes │ │ detected │ │ → Update confidence │ │ -│ │ code │ │ │ │ → Add to examples │ │ -│ └──────────┘ └──────────┘ └──────────────────────────┘ │ -└─────────────────────────────────────────────────────────────────┘ -``` - -**Libraries/tools:** -```bash -# Core parsing -npm install tree-sitter tree-sitter-typescript tree-sitter-python tree-sitter-go tree-sitter-rust - -# Vector storage (local) -npm install better-sqlite3 -# sqlite-vec compiled extension (downloaded during install) - -# Embeddings (API) -npm install voyageai # or call API directly - -# Optional: Local embeddings fallback -# ollama pull nomic-embed-text (for offline mode) -``` - -**The "Mind-Blowing" Features:** - -1. **Convention Inference Engine** - - Analyzes existing code to detect patterns automatically - - "I see 15 files matching `src/services/*.service.ts`, all exporting classes with `@Injectable()`" - - Confidence scores: 15 examples = high confidence, 2 examples = tentative - -2. **Deviation Detection During Execution** - - Before Claude writes `src/utils/auth.ts`, check conventions - - "This looks like a service (has `@Injectable`, depends on repository). Convention suggests `src/services/auth.service.ts`" - - NOT blocking—advisory with reasoning - -3. **Correction Learning** - - User says "no, utils is correct here because X" - - System adds exception: `{pattern: "auth*", location: "utils", reason: "X", confidence: 0.8}` - - Future similar cases consider this exception - -4. **Semantic "Why" Queries** - - Claude can ask the index: "Why is UserRepository in `src/data` not `src/repositories`?" - - Index returns: historical context, similar patterns, any recorded exceptions - -5. **Auto-Refresh with Minimal Recomputation** - - Git hook triggers on commit - - Only re-index changed files (incremental) - - Only re-embed functions that changed (hash comparison) - - Full re-scan weekly or on major refactors - -**Pros:** -- Feels magical ("it knows my codebase") -- Gets smarter over time (adaptive) -- Handles edge cases gracefully (exceptions are learned, not errors) -- Works across languages -- Local-first with optional cloud embeddings - -**Cons:** -- Most complex to implement (~2000 lines + iteration) -- Requires careful feedback loop design -- Cold start problem (needs usage to learn) -- More storage (embeddings + history) - -**Best when:** You want GSD to feel like a senior engineer who truly knows the codebase - -**Complexity:** L - ---- - -## Approach 4: MCP Server Integration (Leverage Existing Tools) - -**How it works:** Use existing code-index-mcp or claude-context MCP servers, integrate with GSD's planning/execution. - -**Libraries/tools:** -- `code-index-mcp` or `claude-context` MCP server -- GSD adds MCP configuration to project setup -- Planner/executor query MCP tools instead of custom indices - -**Pros:** -- Leverage battle-tested implementations -- Active community development -- Already handles tree-sitter, embeddings, incremental updates -- MCP is Claude Code native - -**Cons:** -- Less control over index structure -- May not support adaptive learning -- Another dependency to manage -- May not align perfectly with GSD's workflow - -**Best when:** You want proven tooling without building from scratch - -**Complexity:** M - ---- - -## Comparison - -| Aspect | Static JSON | Semantic Search | Self-Evolving | MCP Server | -|--------|-------------|-----------------|---------------|------------| -| Complexity | S | M | L | M | -| Query Speed | Instant | ~100ms | ~150ms | Varies | -| Semantic Understanding | None | Good | Excellent | Good | -| Learns from Corrections | No | No | Yes | No | -| Offline Capable | Yes | Partial | Partial | Depends | -| Cross-language | Manual | Yes | Yes | Yes | -| Maintenance | Manual | Medium | Self-maintaining | External | -| "Wow Factor" | Low | Medium | High | Medium | - ---- - -## Recommendation - -**Go with Approach 3: Self-Evolving Intelligence**, implemented in phases: - -**Phase 1 (MVP):** Static JSON indices with tree-sitter extraction -- Get basic pattern detection working -- Prove value before adding complexity - -**Phase 2 (Semantic):** Add sqlite-vec + voyage-code-3 -- Enable "find code that does X" queries -- Hybrid search for maximum flexibility - -**Phase 3 (Adaptive):** Add feedback loop and confidence scoring -- Convention rules with confidence -- Learn from corrections -- The "magic" emerges here - -**Phase 4 (Polish):** Auto-refresh, git hooks, exception handling -- Incremental updates -- Graceful degradation when offline - ---- - -## Implementation Context - - - -- name: Self-Evolving Codebase Intelligence (phased) -- libraries: - - tree-sitter + language grammars (parsing) - - better-sqlite3 + sqlite-vec extension (storage) - - voyageai SDK or direct API (embeddings) - - chokidar (file watching, optional) -- install: | - npm install tree-sitter tree-sitter-typescript tree-sitter-python tree-sitter-go - npm install better-sqlite3 - # sqlite-vec: download prebuilt from https://github.com/asg017/sqlite-vec/releases - npm install voyageai - - -- pattern: Three-layer knowledge graph (Structure → Semantics → Learned Rules) -- components: - - CodebaseIndexer: Orchestrates tree-sitter parsing, embedding generation, storage - - StructureExtractor: AST → JSON symbols, patterns, structure - - SemanticMemory: sqlite-vec for embeddings, FTS5 for keywords - - ConventionEngine: Pattern detection, rule inference, confidence scoring - - FeedbackCollector: Captures corrections, updates confidence, logs exceptions - - QueryInterface: Unified API for planner/executor to query knowledge -- data_flow: | - Init: codebase → tree-sitter → symbols.json + patterns.json - → voyage-code-3 → sqlite-vec - → ConventionEngine → conventions.json - - Query: planner asks "where should auth service go?" - → QueryInterface checks conventions.json (high confidence rules) - → Falls back to semantic search if no rule - → Returns recommendation with reasoning - - Feedback: executor writes code → deviation detected → user corrects - → FeedbackCollector updates rule confidence or adds exception - - -- create: - - `.planning/indices/symbols.json` - Extracted code symbols - - `.planning/indices/patterns.json` - Detected architectural patterns - - `.planning/indices/conventions.json` - Learned rules with confidence - - `.planning/indices/codebase.db` - sqlite-vec embeddings + FTS5 - - `.planning/indices/meta.json` - Commit hash, last update, stats - - `get-shit-done/lib/indexer/` - Index generation code - - `get-shit-done/lib/query/` - Query interface for planner/executor -- structure: | - .planning/ - indices/ - symbols.json # {exports, imports, classes, functions} - patterns.json # {detected patterns with locations} - conventions.json # {rules with confidence, examples, exceptions} - codebase.db # sqlite-vec + FTS5 hybrid storage - meta.json # {lastCommit, lastUpdate, stats} -- reference: - - `.planning/codebase/ARCHITECTURE.md` - Use as input for initial pattern detection - - `.planning/codebase/CONVENTIONS.md` - Seed initial convention rules - - `commands/gsd/map-codebase.md` - Extend to also generate indices - - -- start_with: StructureExtractor using tree-sitter (symbols.json output) -- order: - 1. Tree-sitter parsing → symbols.json (Phase 1) - 2. Pattern detection → patterns.json (Phase 1) - 3. sqlite-vec setup + embedding generation (Phase 2) - 4. ConventionEngine + confidence scoring (Phase 3) - 5. FeedbackCollector + correction learning (Phase 3) - 6. Git hooks + incremental updates (Phase 4) - 7. Query interface integration with planner/executor (Phase 4) -- gotchas: - - Tree-sitter grammars are separate packages per language - - sqlite-vec is a C extension, needs platform-specific binary - - voyage-code-3 has 32K context but 120K token batch limit - - Chunking strategy matters: function-level > file-level for embeddings - - Confidence scores need tuning (start conservative, adjust based on feedback) - - Don't over-index: only public exports, key patterns, not every variable -- testing: - - StructureExtractor: Parse known files, verify symbol extraction - - SemanticMemory: Embed test functions, verify similarity search - - ConventionEngine: Feed patterns, verify rule inference - - FeedbackCollector: Simulate corrections, verify confidence updates - - Integration: Full flow from code change → query → recommendation - - - -**Next Action:** Start with Phase 1 - build the tree-sitter StructureExtractor that outputs `symbols.json` for a test project. This proves the parsing pipeline works before adding embeddings. - ---- - -## Sources - -- [Semantic Code Indexing with AST and Tree-sitter](https://medium.com/@email2dineshkuppan/semantic-code-indexing-with-ast-and-tree-sitter-for-ai-agents-part-1-of-3-eb5237ba687a) - Tree-sitter fundamentals -- [mcp-server-tree-sitter](https://github.com/wrale/mcp-server-tree-sitter) - Reference implementation -- [code-index-mcp](https://github.com/johnhuang316/code-index-mcp) - 7-language tree-sitter integration -- [claude-context (Zilliz)](https://github.com/zilliztech/claude-context) - 40% token reduction with semantic search -- [voyage-code-3 announcement](https://blog.voyageai.com/2024/12/04/voyage-code-3/) - State-of-art code embeddings -- [sqlite-vec](https://github.com/asg017/sqlite-vec) - Local vector search for SQLite -- [Hybrid search with sqlite-vec](https://alexgarcia.xyz/blog/2024/sqlite-vec-hybrid-search/index.html) - FTS5 + vector combination -- [Codebases are uniquely hard to search semantically](https://www.greptile.com/blog/semantic-codebase-search) - Why chunking strategy matters -- [Self-Improving Coding Agent (arxiv)](https://arxiv.org/abs/2504.15228) - Research on adaptive code agents -- [Building Self-Improving AI Agents](https://yoheinakajima.com/better-ways-to-build-self-improving-ai-agents/) - Feedback loop architecture -- [State of AI code quality 2025](https://www.qodo.ai/reports/state-of-ai-code-quality/) - Context-aware AI expectations -- [Git Hooks Guide 2025](https://dev.to/arasosman/git-hooks-for-automated-code-quality-checks-guide-2025-372f) - Modern hook practices -- [pre-commit framework](https://pre-commit.com/) - Hook management