Commit Graph

5 Commits

Author SHA1 Message Date
Tom Boucher
d0f916728b feat(skill-surface): install-time profiles + runtime /gsd:surface (#3408) (#3456)
* feat(skill-deps): add requires: frontmatter to all 51 skills with cross-skill references

Mechanical migration from docs/research/data/2026-05-12-skill-audit.json.
Every skill whose body references another GSD skill now declares those
dependencies in `requires:` YAML frontmatter (flow-style array).

Notable: discuss-phase, plan-phase, and execute-phase all reference `phase`,
which confirms the latent gap in MINIMAL_SKILL_ALLOWLIST — `phase` is pulled
by the core loop but was never in the allowlist. The profile closure model
(ADR-0010 Phase 1) resolves this automatically.

Closes part of #3408.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(skill-surface-budget): add PROFILES map, resolveProfile, loadSkillsManifest, staging, marker IO

Implements the Skill Surface Budget Module core (ADR-0010, Phase 1):

- PROFILES Object.freeze map: core (6 skills), standard (~13), full ('*')
- loadSkillsManifest: parses requires: frontmatter from commands/gsd/*.md
  into a Map<stem, string[]> without external YAML dep
- resolveProfile({modes, manifest}): computes transitive closure over the
  requires: graph; composable (modes=['core','audit'] unions closures)
- stageSkillsForProfile / stageAgentsForProfile: filesystem staging with
  same exit-cleanup machinery as the legacy stageSkillsForMode
- readActiveProfile / writeActiveProfile: .gsd-profile marker round-trip
- Back-compat shims preserved: MINIMAL_SKILL_ALLOWLIST, isMinimalMode,
  shouldInstallSkill (overloaded), stageSkillsForMode — all legacy tests pass

The phase latent bug is now resolved by closure: discuss-phase, plan-phase,
and execute-phase all require phase, so any profile including any of them
automatically includes phase via transitive closure.

Tests: 22 manifest+resolve, 9 stage, 10 marker (41 new tests, all green).
Back-compat anchor: 80/80 passing.

Closes part of #3408.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(skill-surface-budget): add lint-skill-deps.cjs CI gate and fix 19 missed requires: entries

Two lint checks (scripts/lint-skill-deps.cjs):
  a) Frontmatter-body consistency: skill body references must appear in requires:
  b) Profile closure: every requires: dep of any profile skill must be in closure

Running the lint revealed 19 body references missed by the audit JSON (the
audit used static analysis; some bodies have conditional references). Fixed:
  complete-milestone: +audit-milestone, discuss-phase, plan-phase, execute-phase, new-milestone
  fast: +quick
  health: +thread
  map-codebase: +new-project, plan-phase
  new-milestone, new-project, review, ultraplan-phase: +plan-phase
  ship: +verify-work
  sketch, spike: +new-project
  verify-work: +execute-phase
  workstreams: +new-milestone, resume-work

Wired into package.json as lint:skill-deps and added to pretest.
8 fixture-based tests: all green.

Closes part of #3408.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(skill-surface-budget): wire --profile= arg, profile marker write/read in bin/install.js

- Add --profile=<name> / --profile=<n1>,<n2> arg parsing (composable).
  Mutually exclusive with --minimal / --core-only (aliases for --profile=core).
  Default (no flag): full.
- Import readActiveProfile / writeActiveProfile from install-profiles.cjs.
- After writeManifest: persist active profile to .gsd-profile marker.
- gsd update path: if no --profile flag given, read existing .gsd-profile
  marker so non-full profiles are not silently re-expanded to full (ADR-0010).
- Update --help block to document --profile= with per-tier token costs.

New test: install-minimal-backcompat.test.cjs (6 tests):
  - PROFILES.core === MINIMAL_SKILL_ALLOWLIST (contract)
  - --minimal writes .gsd-profile marker "core"
  - --profile=core, --profile=standard write correct markers
  - default install writes marker "full"

Closes part of #3408.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): add feat-3408-skill-profiles changelog fragment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(install-profiles): derive agents from skill body refs and wire into resolveProfile

Deviation 1 of ADR-0010 phase 1b: tiered profiles (core, standard) now produce
a non-empty agents Set instead of always returning empty. resolveProfile() scans
each skill body for gsd-* agent name references (via new parseCallsAgents()),
stores them in _calls_agents_<stem> manifest entries, and unions them across the
resolved skill closure. stageAgentsForProfile() already checked resolvedProfile.agents
— it now gets real data so tiered profiles install the correct subset of agents
instead of zero.

Closes #3408 (partial — Deviation 1 only)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(install): honor .gsd-profile marker on update, add resolveEffectiveProfile/mostRestrictiveProfile

Deviation 2 of ADR-0010 phase 1b: the marker written during installation is now
actually honored when re-running without explicit flags (e.g. gsd update). The
dead-end logging block is replaced by resolveEffectiveProfile(), which picks the
marker profile over 'full' when no explicit --profile= flag was given. The resolved
profile is piped through to all 13 stageSkillsForMode dispatch sites (now _stageSkills)
so updates install only the previously-chosen skill subset.

--minimal retains its back-compat behavior (strict 6-skill allowlist, no closure)
while writing 'core' to the marker. mostRestrictiveProfile() is exported for callers
that need to reconcile disagreeing markers across runtimes (smallest skill set wins).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(surface): add CLUSTERS data + state IO module

Add clusters.cjs with 10 named skill groups covering all 66 skills
(verified by surface-clusters.test.cjs). Add surface.cjs with readSurface/
writeSurface atomic IO, resolveSurface, applySurface, and listSurface.
Tests: 17 passing (11 state IO + 6 cluster integrity).

Closes #3408

* docs(adr): add ADR-0011 Skill Surface Budget Module (Phase 1 accepted, Phase 2 amendment)

Records the install-time profile staging decision (Phase 1, landed) and the
runtime /gsd:surface cluster-toggle decision (Phase 2, in flight) as an
amendment. Updates the ADR README index.

Closes #3408

* docs(install-profiles): update module docblock for Phase 2 and ADR-0011

Corrects the ADR reference from 0010 to 0011, documents the three-profile
model and back-compat aliases, adds resolveEffectiveProfile precedence rule,
and notes the companion surface.cjs Phase 2 engine.

* docs(context): add Skill Surface Budget Module canonical entry

Adds the Domain terms entry for the Skill Surface Budget Module covering
both Phase 1 (install-time profiles, .gsd-profile marker) and Phase 2
(runtime /gsd:surface cluster toggles, clusters.cjs, .gsd-surface.json),
per ADR-0011 Consequences requirement.

* feat(surface): add resolveSurface and applySurface engine + tests

Tests cover: profile → surface equivalence, cluster disable/enable,
explicitAdds transitive closure, applySurface file sync (add missing,
remove superseded, preserve non-gsd files), listSurface token cost.
16 new tests passing.

* docs(readme): document --profile= flag and /gsd:surface command

Brief user-facing mention of install profiles (core/standard/full) and the
/gsd:surface slash command in the Commands table. Points to ADR-0011 for details.

* feat(surface): add /gsd:surface slash command runbook

New skill: gsd:surface — runtime profile/cluster toggle without reinstall.
Sub-commands: list, status, profile <name>, disable/enable <cluster>, reset.
Persists state to .gsd-surface.json (independent of .gsd-profile).
Description 96 chars (≤100 limit). lint:descriptions + lint:skill-deps: 0 violations.

* feat(surface): add changeset fragment for /gsd:surface runtime toggle

* feat(surface): add surface skill stem to utility cluster

surface.md is a new skill; add it to the utility cluster so the
surface-clusters.test.cjs coverage invariant stays satisfied.

* docs(adr): fix ADR references to 0011 and record Phase 2 as shipped

ADR-0010 number was already claimed by the file-operation-engine ADR; this
ADR landed as 0011-skill-surface-budget-module.md. Update inline ADR
references in clusters.cjs, surface.cjs, install-profiles.cjs, and the
Phase 2 changeset to ADR-0011. Update the ADR Status section to record
Phase 2 artifacts as shipped on this branch rather than "in progress".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(research): port skill-surface-budget memo and audit data

ADR-0011 references docs/research/2026-05-12-skill-surface-budget.md and
docs/research/data/2026-05-12-skill-audit.json, which only existed in the
research worktree. Port both onto this branch so the ADR's References
section resolves and reviewers can read the cluster taxonomy (§3.2),
dependency topology (§3.1), and option grading (§4) that justify Phase 1
and Phase 2 decisions.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(registration): register surface/clusters in INVENTORY, COMMANDS, and help.md

- surface.md: convert allowed-tools from inline YAML array to block style
  (was parsed as a single tool name "[Read, Write, Bash]" by test harness)
- docs/INVENTORY.md: add CLI module rows for clusters.cjs and surface.cjs;
  add Commands row for /gsd-surface; bump CLI Modules count 55→57, Commands 66→67
- docs/INVENTORY-MANIFEST.json: add entries for clusters.cjs, surface.cjs,
  and /gsd-surface (filename-based command key)
- docs/COMMANDS.md: add ### `/gsd-surface` heading in Configuration Commands
- get-shit-done/workflows/help.md: add /gsd:surface entry in Configuration section

Fixes registration failures introduced by Phase 2 of #3408.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(surface,docs): scrub .claude leakage and escape hypothetical slash tokens

Two PR regressions introduced earlier on this branch:

1. surface.cjs JSDoc comments contained the canonical paths
   (~/.claude/commands/gsd, ~/.claude/agents) as example values, which the
   cline-install leak regex (~\/\.claude\/(?:get-shit-done|commands|agents
   |hooks)) flagged as install-time path leaks. Reworded the docblocks to
   describe runtime-resolved paths without literal ~/.claude tokens.

2. The ported research memo proposed hypothetical Option C dispatchers
   using slash syntax (/gsd:milestone, /gsd:research). The
   docs-parity-live-registry test enforces that every slash-command token
   in docs/ resolves to a real command. Rewrote the Option C sketch
   without the slash prefix and added a clarifying note that the
   dispatchers are illustrative, not shipped.

Targeted tests now pass: tests/cline-install.test.cjs and
tests/docs-parity-live-registry.test.cjs both green.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test: remove raw output/source grep in lint tests

* fix: close coderabbit profile and requires issues

* test: align surface token-cost assertion wording

* fix(install): align core profile alias and defer profile marker write

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-13 12:45:16 -04:00
Tom Boucher
1452b1275b fix(dispatcher): rename Task→Agent in allowed-tools, workflow prose, and agent tools frontmatter
Fixes #3168

The Claude Code subagent dispatcher tool is named `Agent` (with `subagent_type`
parameter). The `Task*` namespace (TaskCreate, TaskList, TaskGet, TaskUpdate,
TaskOutput, TaskStop) is the separate task-tracker. GSD's commands, workflows,
and agents were partially migrated and still referenced `- Task` / `Task(` in
55 files, causing orchestrators to silently fall back to inline execution when
no `Task` tool appeared on their tool surface.

Changes:
- `commands/gsd/*.md` allowed-tools: replaced `- Task` with `- Agent` in 24
  files; removed duplicate `- Task` from autonomous.md (already had `- Agent`)
- `get-shit-done/workflows/*.md`: replaced dispatcher `Task(` → `Agent(` in
  29 workflow files (~133 call sites); TaskCreate/List/Get/Update/Output/Stop
  left untouched
- `agents/gsd-debug-session-manager.md`: replaced `Task` → `Agent` in tools
  frontmatter (the only remaining agent with the wrong name)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-06 15:00:08 -04:00
Tom Boucher
81f9534b5a feat(adr-0002): command contract validation module + prose @-ref cleanup + workflow extraction
ADR-0002: commands/gsd/*.md contract now enforced at two layers:

LINT (scripts/lint-command-contract.cjs — new CI step):
- name: present, starts with gsd: or gsd-
- description: non-empty
- allowed-tools: non-empty, all entries canonical
- execution_context @-refs: resolve on disk, no trailing prose on same line
- handles both @~/ and $HOME/ path prefixes

TEST (tests/command-contract.test.cjs — 361 assertions):
- Behavioral contract for all 65 command files
- Replaces scattered coverage in enh-2790 + bug-3135
- Per-command per-rule test — one failure names the exact file + rule

CI (.github/workflows/test.yml):
- 'Lint — command contract (ADR-0002)' step added to lint-tests job

PROSE @-REF CLEANUP (39 command files, ~900 tokens/invocation recovered):
- Removed redundant @~/.claude/get-shit-done/... paths from <process> prose
- execution_context block is now the single authoritative load declaration
- Routing commands (sketch, spike, update, pause-work, etc.) keep routing
  instructions; only the inert path token is stripped

WORKFLOW EXTRACTION (debug.md + thread.md, ~15,000 chars / ~3,750 tokens):
- get-shit-done/workflows/debug.md: full process extracted from commands/gsd/debug.md
- get-shit-done/workflows/thread.md: full process extracted from commands/gsd/thread.md
- Command files reduced to frontmatter + objective + execution_context + context
- debug.md: 9,603 → 1,703 chars; thread.md: 7,868 → 585 chars

RENAME:
- get-shit-done/workflows/extract_learnings.md → extract-learnings.md
  (aligns with hyphen convention of all other workflow files)

DOCS:
- docs/INVENTORY.md: count 85→87, new rows, rename row, fix add-todo --backlog attribution
- docs/INVENTORY-MANIFEST.json: +debug.md +thread.md +extract-learnings.md -extract_learnings.md

Closes ADR-0002 implementation.
2026-05-05 15:18:13 -04:00
Tom Boucher
e81592878e feat(#2789): trim skill description anti-patterns; enforce 100-char budget (#2823)
* feat(#2789): trim skill description anti-patterns; enforce 100-char budget

- Trim descriptions in all commands/gsd/*.md files over 100 chars
- Remove flag documentation from descriptions (belongs in argument-hint)
- Remove Triggers: keyword stuffing
- Add scripts/lint-descriptions.cjs — fails on descriptions > 100 chars
- Add npm script: lint:descriptions
- Add tests/enh-2789-description-budget.test.cjs

Closes #2789

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#2789): add CHANGELOG entry for description budget lint

* docs(#2789): update COMMANDS.md descriptions; add skill description standards note

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-29 08:14:11 -04:00
Fana
33575ba91d feat: /gsd-ai-integration-phase + /gsd-eval-review — AI framework selection and eval coverage layer (#1971)
* feat: /gsd:ai-phase + /gsd:eval-review — AI evals and framework selection layer

Adds a structured AI development layer to GSD with 5 new agents, 2 new
commands, 2 new workflows, 2 reference files, and 1 template.

Commands:
- /gsd:ai-phase [N] — pre-planning AI design contract (inserts between
  discuss-phase and plan-phase). Orchestrates 4 agents in sequence:
  framework-selector → ai-researcher → domain-researcher → eval-planner.
  Output: AI-SPEC.md with framework decision, implementation guidance,
  domain expert context, and evaluation strategy.
- /gsd:eval-review [N] — retroactive eval coverage audit. Scores each
  planned eval dimension as COVERED/PARTIAL/MISSING. Output: EVAL-REVIEW.md
  with 0-100 score, verdict, and remediation plan.

Agents:
- gsd-framework-selector: interactive decision matrix (6 questions) →
  scored framework recommendation for CrewAI, LlamaIndex, LangChain,
  LangGraph, OpenAI Agents SDK, Claude Agent SDK, AutoGen/AG2, Haystack
- gsd-ai-researcher: fetches official framework docs + writes AI systems
  best practices (Pydantic structured outputs, async-first, prompt
  discipline, context window management, cost/latency budget)
- gsd-domain-researcher: researches business domain and use-case context —
  surfaces domain expert evaluation criteria, industry failure modes,
  regulatory constraints, and practitioner rubric ingredients before
  eval-planner writes measurable criteria
- gsd-eval-planner: designs evaluation strategy grounded in domain context;
  defaults to Arize Phoenix (tracing) + RAGAS (RAG eval) with detect-first
  guard for existing tooling
- gsd-eval-auditor: retroactive codebase scan → scores eval coverage

Integration points:
- plan-phase: non-blocking nudge (step 4.5) when AI keywords detected and
  no AI-SPEC.md present
- settings: new workflow.ai_phase toggle (default on)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: refine ai-integration-phase layer — rename, house style, consistency fixes

Amends the ai-evals framework layer (df8cb6c) with post-review improvements
before opening upstream PR.

Rename /gsd:ai-phase → /gsd:ai-integration-phase:
- Renamed commands/gsd/ai-phase.md → ai-integration-phase.md
- Renamed get-shit-done/workflows/ai-phase.md → ai-integration-phase.md
- Updated config key: workflow.ai_phase → workflow.ai_integration_phase
- Updated repair action: addAiPhaseKey → addAiIntegrationPhaseKey
- Updated all 84 cross-references across agents, workflows, templates, tests

Consistency fixes (same class as PR #1380 review):
- commands/gsd: objective described 3-agent chain, missing gsd-domain-researcher
- workflows/ai-integration-phase: purpose tag described 3-agent chain + "locks
  three things" — updated to 4 agents + 4 outputs
- workflows/ai-integration-phase: missing DOMAIN_MODEL resolve-model call in
  step 1 (domain-researcher was spawned in step 7.5 with no model variable)
- workflows/ai-integration-phase: fractional step ## 7.5 renumbered to integers
  (steps 8–12 shifted)

Agent house style (GSD meta-prompting conformance):
- All 5 new agents refactored to execution_flow + step name="" structure
- Role blocks compressed to 2 lines (removed verbose "Core responsibilities")
- Added skills: frontmatter to all 5 agents (agent-frontmatter tests)
- Added # hooks: commented pattern to file-writing agents
- Added ALWAYS use Write tool anti-heredoc instruction to file-writing agents
- Line reductions: ai-researcher −41%, domain-researcher −25%, eval-planner −26%,
  eval-auditor −25%, framework-selector −9%

Test coverage (tests/ai-evals.test.cjs — 48 tests):
- CONFIG: workflow.ai_integration_phase defaults and config-set/get
- HEALTH: W010 warning emission and addAiIntegrationPhaseKey repair
- TEMPLATE: AI-SPEC.md section completeness (10 sections)
- COMMAND: ai-integration-phase + eval-review frontmatter validity
- AGENTS: all 5 new agent files exist
- REFERENCES: ai-evals.md + ai-frameworks.md exist and are non-empty
- WORKFLOW: plan-phase nudge integration, workflow files exist + agent coverage

603/603 tests passing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: add Google ADK to framework selector and reference matrix

Google ADK (released March 2025) was missing from the framework options.
Adds Python + Java multi-agent framework optimised for Gemini / Vertex AI.

- get-shit-done/references/ai-frameworks.md: add Google ADK profile (type,
  language, model support, best for, avoid if, strengths, weaknesses, eval
  concerns); update Quick Picks, By System Type, and By Model Commitment tables
- agents/gsd-framework-selector.md: add "Google (Gemini)" to model provider
  interview question
- agents/gsd-ai-researcher.md: add Google ADK docs URL to documentation_sources

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: adapt to upstream conventions post-rebase

- Remove skills: frontmatter from all 5 new agents (upstream changed
  convention — skills: breaks Gemini CLI and must not be present)
- Add workflow.ai_integration_phase to VALID_CONFIG_KEYS whitelist in
  config.cjs (config-set blocked unknown keys)
- Add ai_integration_phase: true to CONFIG_DEFAULTS in core.cjs

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: rephrase 4b.1 line to avoid false-positive in prompt-injection scan

"contract as a Pydantic model" matched the `act as a` pattern case-insensitively.
Rephrased to "output schema using a Pydantic model".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: adapt to upstream conventions (W016, colon refs, config docs)

- Replace verify.cjs from upstream to restore W010-W015 + cmdValidateAgents,
  lost when rebase conflict was resolved with --theirs
- Add W016 (workflow.ai_integration_phase absent) inside the config try block,
  avoids collision with upstream's W010 agent-installation check
- Add addAiIntegrationPhaseKey repair case mirroring addNyquistKey pattern
- Replace /gsd: colon format with /gsd- hyphen format across all new files
  (agents, workflows, templates, verify.cjs) per stale-colon-refs guard (#1748)
- Add workflow.ai_integration_phase to planning-config.md reference table
- Add ai_integration_phase → workflow.ai_integration_phase to NAMESPACE_MAP
  in config-field-docs.test.cjs so CONFIG_DEFAULTS coverage check passes
- Update ai-evals tests to use W016 instead of W010

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: add 5 new agents to E2E Copilot install expected list

gsd-ai-researcher, gsd-domain-researcher, gsd-eval-auditor,
gsd-eval-planner, gsd-framework-selector added to the hardcoded
expected agent list in copilot-install.test.cjs (#1890).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-10 10:49:00 -04:00