90e60b414595f510d6f519e5087ce7b7943ae38c
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d0f916728b |
feat(skill-surface): install-time profiles + runtime /gsd:surface (#3408) (#3456)
* feat(skill-deps): add requires: frontmatter to all 51 skills with cross-skill references Mechanical migration from docs/research/data/2026-05-12-skill-audit.json. Every skill whose body references another GSD skill now declares those dependencies in `requires:` YAML frontmatter (flow-style array). Notable: discuss-phase, plan-phase, and execute-phase all reference `phase`, which confirms the latent gap in MINIMAL_SKILL_ALLOWLIST — `phase` is pulled by the core loop but was never in the allowlist. The profile closure model (ADR-0010 Phase 1) resolves this automatically. Closes part of #3408. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(skill-surface-budget): add PROFILES map, resolveProfile, loadSkillsManifest, staging, marker IO Implements the Skill Surface Budget Module core (ADR-0010, Phase 1): - PROFILES Object.freeze map: core (6 skills), standard (~13), full ('*') - loadSkillsManifest: parses requires: frontmatter from commands/gsd/*.md into a Map<stem, string[]> without external YAML dep - resolveProfile({modes, manifest}): computes transitive closure over the requires: graph; composable (modes=['core','audit'] unions closures) - stageSkillsForProfile / stageAgentsForProfile: filesystem staging with same exit-cleanup machinery as the legacy stageSkillsForMode - readActiveProfile / writeActiveProfile: .gsd-profile marker round-trip - Back-compat shims preserved: MINIMAL_SKILL_ALLOWLIST, isMinimalMode, shouldInstallSkill (overloaded), stageSkillsForMode — all legacy tests pass The phase latent bug is now resolved by closure: discuss-phase, plan-phase, and execute-phase all require phase, so any profile including any of them automatically includes phase via transitive closure. Tests: 22 manifest+resolve, 9 stage, 10 marker (41 new tests, all green). Back-compat anchor: 80/80 passing. Closes part of #3408. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(skill-surface-budget): add lint-skill-deps.cjs CI gate and fix 19 missed requires: entries Two lint checks (scripts/lint-skill-deps.cjs): a) Frontmatter-body consistency: skill body references must appear in requires: b) Profile closure: every requires: dep of any profile skill must be in closure Running the lint revealed 19 body references missed by the audit JSON (the audit used static analysis; some bodies have conditional references). Fixed: complete-milestone: +audit-milestone, discuss-phase, plan-phase, execute-phase, new-milestone fast: +quick health: +thread map-codebase: +new-project, plan-phase new-milestone, new-project, review, ultraplan-phase: +plan-phase ship: +verify-work sketch, spike: +new-project verify-work: +execute-phase workstreams: +new-milestone, resume-work Wired into package.json as lint:skill-deps and added to pretest. 8 fixture-based tests: all green. Closes part of #3408. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(skill-surface-budget): wire --profile= arg, profile marker write/read in bin/install.js - Add --profile=<name> / --profile=<n1>,<n2> arg parsing (composable). Mutually exclusive with --minimal / --core-only (aliases for --profile=core). Default (no flag): full. - Import readActiveProfile / writeActiveProfile from install-profiles.cjs. - After writeManifest: persist active profile to .gsd-profile marker. - gsd update path: if no --profile flag given, read existing .gsd-profile marker so non-full profiles are not silently re-expanded to full (ADR-0010). - Update --help block to document --profile= with per-tier token costs. New test: install-minimal-backcompat.test.cjs (6 tests): - PROFILES.core === MINIMAL_SKILL_ALLOWLIST (contract) - --minimal writes .gsd-profile marker "core" - --profile=core, --profile=standard write correct markers - default install writes marker "full" Closes part of #3408. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(changeset): add feat-3408-skill-profiles changelog fragment Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(install-profiles): derive agents from skill body refs and wire into resolveProfile Deviation 1 of ADR-0010 phase 1b: tiered profiles (core, standard) now produce a non-empty agents Set instead of always returning empty. resolveProfile() scans each skill body for gsd-* agent name references (via new parseCallsAgents()), stores them in _calls_agents_<stem> manifest entries, and unions them across the resolved skill closure. stageAgentsForProfile() already checked resolvedProfile.agents — it now gets real data so tiered profiles install the correct subset of agents instead of zero. Closes #3408 (partial — Deviation 1 only) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(install): honor .gsd-profile marker on update, add resolveEffectiveProfile/mostRestrictiveProfile Deviation 2 of ADR-0010 phase 1b: the marker written during installation is now actually honored when re-running without explicit flags (e.g. gsd update). The dead-end logging block is replaced by resolveEffectiveProfile(), which picks the marker profile over 'full' when no explicit --profile= flag was given. The resolved profile is piped through to all 13 stageSkillsForMode dispatch sites (now _stageSkills) so updates install only the previously-chosen skill subset. --minimal retains its back-compat behavior (strict 6-skill allowlist, no closure) while writing 'core' to the marker. mostRestrictiveProfile() is exported for callers that need to reconcile disagreeing markers across runtimes (smallest skill set wins). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(surface): add CLUSTERS data + state IO module Add clusters.cjs with 10 named skill groups covering all 66 skills (verified by surface-clusters.test.cjs). Add surface.cjs with readSurface/ writeSurface atomic IO, resolveSurface, applySurface, and listSurface. Tests: 17 passing (11 state IO + 6 cluster integrity). Closes #3408 * docs(adr): add ADR-0011 Skill Surface Budget Module (Phase 1 accepted, Phase 2 amendment) Records the install-time profile staging decision (Phase 1, landed) and the runtime /gsd:surface cluster-toggle decision (Phase 2, in flight) as an amendment. Updates the ADR README index. Closes #3408 * docs(install-profiles): update module docblock for Phase 2 and ADR-0011 Corrects the ADR reference from 0010 to 0011, documents the three-profile model and back-compat aliases, adds resolveEffectiveProfile precedence rule, and notes the companion surface.cjs Phase 2 engine. * docs(context): add Skill Surface Budget Module canonical entry Adds the Domain terms entry for the Skill Surface Budget Module covering both Phase 1 (install-time profiles, .gsd-profile marker) and Phase 2 (runtime /gsd:surface cluster toggles, clusters.cjs, .gsd-surface.json), per ADR-0011 Consequences requirement. * feat(surface): add resolveSurface and applySurface engine + tests Tests cover: profile → surface equivalence, cluster disable/enable, explicitAdds transitive closure, applySurface file sync (add missing, remove superseded, preserve non-gsd files), listSurface token cost. 16 new tests passing. * docs(readme): document --profile= flag and /gsd:surface command Brief user-facing mention of install profiles (core/standard/full) and the /gsd:surface slash command in the Commands table. Points to ADR-0011 for details. * feat(surface): add /gsd:surface slash command runbook New skill: gsd:surface — runtime profile/cluster toggle without reinstall. Sub-commands: list, status, profile <name>, disable/enable <cluster>, reset. Persists state to .gsd-surface.json (independent of .gsd-profile). Description 96 chars (≤100 limit). lint:descriptions + lint:skill-deps: 0 violations. * feat(surface): add changeset fragment for /gsd:surface runtime toggle * feat(surface): add surface skill stem to utility cluster surface.md is a new skill; add it to the utility cluster so the surface-clusters.test.cjs coverage invariant stays satisfied. * docs(adr): fix ADR references to 0011 and record Phase 2 as shipped ADR-0010 number was already claimed by the file-operation-engine ADR; this ADR landed as 0011-skill-surface-budget-module.md. Update inline ADR references in clusters.cjs, surface.cjs, install-profiles.cjs, and the Phase 2 changeset to ADR-0011. Update the ADR Status section to record Phase 2 artifacts as shipped on this branch rather than "in progress". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(research): port skill-surface-budget memo and audit data ADR-0011 references docs/research/2026-05-12-skill-surface-budget.md and docs/research/data/2026-05-12-skill-audit.json, which only existed in the research worktree. Port both onto this branch so the ADR's References section resolves and reviewers can read the cluster taxonomy (§3.2), dependency topology (§3.1), and option grading (§4) that justify Phase 1 and Phase 2 decisions. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(registration): register surface/clusters in INVENTORY, COMMANDS, and help.md - surface.md: convert allowed-tools from inline YAML array to block style (was parsed as a single tool name "[Read, Write, Bash]" by test harness) - docs/INVENTORY.md: add CLI module rows for clusters.cjs and surface.cjs; add Commands row for /gsd-surface; bump CLI Modules count 55→57, Commands 66→67 - docs/INVENTORY-MANIFEST.json: add entries for clusters.cjs, surface.cjs, and /gsd-surface (filename-based command key) - docs/COMMANDS.md: add ### `/gsd-surface` heading in Configuration Commands - get-shit-done/workflows/help.md: add /gsd:surface entry in Configuration section Fixes registration failures introduced by Phase 2 of #3408. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(surface,docs): scrub .claude leakage and escape hypothetical slash tokens Two PR regressions introduced earlier on this branch: 1. surface.cjs JSDoc comments contained the canonical paths (~/.claude/commands/gsd, ~/.claude/agents) as example values, which the cline-install leak regex (~\/\.claude\/(?:get-shit-done|commands|agents |hooks)) flagged as install-time path leaks. Reworded the docblocks to describe runtime-resolved paths without literal ~/.claude tokens. 2. The ported research memo proposed hypothetical Option C dispatchers using slash syntax (/gsd:milestone, /gsd:research). The docs-parity-live-registry test enforces that every slash-command token in docs/ resolves to a real command. Rewrote the Option C sketch without the slash prefix and added a clarifying note that the dispatchers are illustrative, not shipped. Targeted tests now pass: tests/cline-install.test.cjs and tests/docs-parity-live-registry.test.cjs both green. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test: remove raw output/source grep in lint tests * fix: close coderabbit profile and requires issues * test: align surface token-cost assertion wording * fix(install): align core profile alias and defer profile marker write --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
1452b1275b |
fix(dispatcher): rename Task→Agent in allowed-tools, workflow prose, and agent tools frontmatter
Fixes #3168 The Claude Code subagent dispatcher tool is named `Agent` (with `subagent_type` parameter). The `Task*` namespace (TaskCreate, TaskList, TaskGet, TaskUpdate, TaskOutput, TaskStop) is the separate task-tracker. GSD's commands, workflows, and agents were partially migrated and still referenced `- Task` / `Task(` in 55 files, causing orchestrators to silently fall back to inline execution when no `Task` tool appeared on their tool surface. Changes: - `commands/gsd/*.md` allowed-tools: replaced `- Task` with `- Agent` in 24 files; removed duplicate `- Task` from autonomous.md (already had `- Agent`) - `get-shit-done/workflows/*.md`: replaced dispatcher `Task(` → `Agent(` in 29 workflow files (~133 call sites); TaskCreate/List/Get/Update/Output/Stop left untouched - `agents/gsd-debug-session-manager.md`: replaced `Task` → `Agent` in tools frontmatter (the only remaining agent with the wrong name) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
81f9534b5a |
feat(adr-0002): command contract validation module + prose @-ref cleanup + workflow extraction
ADR-0002: commands/gsd/*.md contract now enforced at two layers: LINT (scripts/lint-command-contract.cjs — new CI step): - name: present, starts with gsd: or gsd- - description: non-empty - allowed-tools: non-empty, all entries canonical - execution_context @-refs: resolve on disk, no trailing prose on same line - handles both @~/ and $HOME/ path prefixes TEST (tests/command-contract.test.cjs — 361 assertions): - Behavioral contract for all 65 command files - Replaces scattered coverage in enh-2790 + bug-3135 - Per-command per-rule test — one failure names the exact file + rule CI (.github/workflows/test.yml): - 'Lint — command contract (ADR-0002)' step added to lint-tests job PROSE @-REF CLEANUP (39 command files, ~900 tokens/invocation recovered): - Removed redundant @~/.claude/get-shit-done/... paths from <process> prose - execution_context block is now the single authoritative load declaration - Routing commands (sketch, spike, update, pause-work, etc.) keep routing instructions; only the inert path token is stripped WORKFLOW EXTRACTION (debug.md + thread.md, ~15,000 chars / ~3,750 tokens): - get-shit-done/workflows/debug.md: full process extracted from commands/gsd/debug.md - get-shit-done/workflows/thread.md: full process extracted from commands/gsd/thread.md - Command files reduced to frontmatter + objective + execution_context + context - debug.md: 9,603 → 1,703 chars; thread.md: 7,868 → 585 chars RENAME: - get-shit-done/workflows/extract_learnings.md → extract-learnings.md (aligns with hyphen convention of all other workflow files) DOCS: - docs/INVENTORY.md: count 85→87, new rows, rename row, fix add-todo --backlog attribution - docs/INVENTORY-MANIFEST.json: +debug.md +thread.md +extract-learnings.md -extract_learnings.md Closes ADR-0002 implementation. |
||
|
|
e81592878e |
feat(#2789): trim skill description anti-patterns; enforce 100-char budget (#2823)
* feat(#2789): trim skill description anti-patterns; enforce 100-char budget - Trim descriptions in all commands/gsd/*.md files over 100 chars - Remove flag documentation from descriptions (belongs in argument-hint) - Remove Triggers: keyword stuffing - Add scripts/lint-descriptions.cjs — fails on descriptions > 100 chars - Add npm script: lint:descriptions - Add tests/enh-2789-description-budget.test.cjs Closes #2789 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(#2789): add CHANGELOG entry for description budget lint * docs(#2789): update COMMANDS.md descriptions; add skill description standards note Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
33575ba91d |
feat: /gsd-ai-integration-phase + /gsd-eval-review — AI framework selection and eval coverage layer (#1971)
* feat: /gsd:ai-phase + /gsd:eval-review — AI evals and framework selection layer Adds a structured AI development layer to GSD with 5 new agents, 2 new commands, 2 new workflows, 2 reference files, and 1 template. Commands: - /gsd:ai-phase [N] — pre-planning AI design contract (inserts between discuss-phase and plan-phase). Orchestrates 4 agents in sequence: framework-selector → ai-researcher → domain-researcher → eval-planner. Output: AI-SPEC.md with framework decision, implementation guidance, domain expert context, and evaluation strategy. - /gsd:eval-review [N] — retroactive eval coverage audit. Scores each planned eval dimension as COVERED/PARTIAL/MISSING. Output: EVAL-REVIEW.md with 0-100 score, verdict, and remediation plan. Agents: - gsd-framework-selector: interactive decision matrix (6 questions) → scored framework recommendation for CrewAI, LlamaIndex, LangChain, LangGraph, OpenAI Agents SDK, Claude Agent SDK, AutoGen/AG2, Haystack - gsd-ai-researcher: fetches official framework docs + writes AI systems best practices (Pydantic structured outputs, async-first, prompt discipline, context window management, cost/latency budget) - gsd-domain-researcher: researches business domain and use-case context — surfaces domain expert evaluation criteria, industry failure modes, regulatory constraints, and practitioner rubric ingredients before eval-planner writes measurable criteria - gsd-eval-planner: designs evaluation strategy grounded in domain context; defaults to Arize Phoenix (tracing) + RAGAS (RAG eval) with detect-first guard for existing tooling - gsd-eval-auditor: retroactive codebase scan → scores eval coverage Integration points: - plan-phase: non-blocking nudge (step 4.5) when AI keywords detected and no AI-SPEC.md present - settings: new workflow.ai_phase toggle (default on) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: refine ai-integration-phase layer — rename, house style, consistency fixes Amends the ai-evals framework layer (df8cb6c) with post-review improvements before opening upstream PR. Rename /gsd:ai-phase → /gsd:ai-integration-phase: - Renamed commands/gsd/ai-phase.md → ai-integration-phase.md - Renamed get-shit-done/workflows/ai-phase.md → ai-integration-phase.md - Updated config key: workflow.ai_phase → workflow.ai_integration_phase - Updated repair action: addAiPhaseKey → addAiIntegrationPhaseKey - Updated all 84 cross-references across agents, workflows, templates, tests Consistency fixes (same class as PR #1380 review): - commands/gsd: objective described 3-agent chain, missing gsd-domain-researcher - workflows/ai-integration-phase: purpose tag described 3-agent chain + "locks three things" — updated to 4 agents + 4 outputs - workflows/ai-integration-phase: missing DOMAIN_MODEL resolve-model call in step 1 (domain-researcher was spawned in step 7.5 with no model variable) - workflows/ai-integration-phase: fractional step ## 7.5 renumbered to integers (steps 8–12 shifted) Agent house style (GSD meta-prompting conformance): - All 5 new agents refactored to execution_flow + step name="" structure - Role blocks compressed to 2 lines (removed verbose "Core responsibilities") - Added skills: frontmatter to all 5 agents (agent-frontmatter tests) - Added # hooks: commented pattern to file-writing agents - Added ALWAYS use Write tool anti-heredoc instruction to file-writing agents - Line reductions: ai-researcher −41%, domain-researcher −25%, eval-planner −26%, eval-auditor −25%, framework-selector −9% Test coverage (tests/ai-evals.test.cjs — 48 tests): - CONFIG: workflow.ai_integration_phase defaults and config-set/get - HEALTH: W010 warning emission and addAiIntegrationPhaseKey repair - TEMPLATE: AI-SPEC.md section completeness (10 sections) - COMMAND: ai-integration-phase + eval-review frontmatter validity - AGENTS: all 5 new agent files exist - REFERENCES: ai-evals.md + ai-frameworks.md exist and are non-empty - WORKFLOW: plan-phase nudge integration, workflow files exist + agent coverage 603/603 tests passing. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: add Google ADK to framework selector and reference matrix Google ADK (released March 2025) was missing from the framework options. Adds Python + Java multi-agent framework optimised for Gemini / Vertex AI. - get-shit-done/references/ai-frameworks.md: add Google ADK profile (type, language, model support, best for, avoid if, strengths, weaknesses, eval concerns); update Quick Picks, By System Type, and By Model Commitment tables - agents/gsd-framework-selector.md: add "Google (Gemini)" to model provider interview question - agents/gsd-ai-researcher.md: add Google ADK docs URL to documentation_sources Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions post-rebase - Remove skills: frontmatter from all 5 new agents (upstream changed convention — skills: breaks Gemini CLI and must not be present) - Add workflow.ai_integration_phase to VALID_CONFIG_KEYS whitelist in config.cjs (config-set blocked unknown keys) - Add ai_integration_phase: true to CONFIG_DEFAULTS in core.cjs Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: rephrase 4b.1 line to avoid false-positive in prompt-injection scan "contract as a Pydantic model" matched the `act as a` pattern case-insensitively. Rephrased to "output schema using a Pydantic model". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions (W016, colon refs, config docs) - Replace verify.cjs from upstream to restore W010-W015 + cmdValidateAgents, lost when rebase conflict was resolved with --theirs - Add W016 (workflow.ai_integration_phase absent) inside the config try block, avoids collision with upstream's W010 agent-installation check - Add addAiIntegrationPhaseKey repair case mirroring addNyquistKey pattern - Replace /gsd: colon format with /gsd- hyphen format across all new files (agents, workflows, templates, verify.cjs) per stale-colon-refs guard (#1748) - Add workflow.ai_integration_phase to planning-config.md reference table - Add ai_integration_phase → workflow.ai_integration_phase to NAMESPACE_MAP in config-field-docs.test.cjs so CONFIG_DEFAULTS coverage check passes - Update ai-evals tests to use W016 instead of W010 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: add 5 new agents to E2E Copilot install expected list gsd-ai-researcher, gsd-domain-researcher, gsd-eval-auditor, gsd-eval-planner, gsd-framework-selector added to the hardcoded expected agent list in copilot-install.test.cjs (#1890). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |