* feat: /gsd:ai-phase + /gsd:eval-review — AI evals and framework selection layer Adds a structured AI development layer to GSD with 5 new agents, 2 new commands, 2 new workflows, 2 reference files, and 1 template. Commands: - /gsd:ai-phase [N] — pre-planning AI design contract (inserts between discuss-phase and plan-phase). Orchestrates 4 agents in sequence: framework-selector → ai-researcher → domain-researcher → eval-planner. Output: AI-SPEC.md with framework decision, implementation guidance, domain expert context, and evaluation strategy. - /gsd:eval-review [N] — retroactive eval coverage audit. Scores each planned eval dimension as COVERED/PARTIAL/MISSING. Output: EVAL-REVIEW.md with 0-100 score, verdict, and remediation plan. Agents: - gsd-framework-selector: interactive decision matrix (6 questions) → scored framework recommendation for CrewAI, LlamaIndex, LangChain, LangGraph, OpenAI Agents SDK, Claude Agent SDK, AutoGen/AG2, Haystack - gsd-ai-researcher: fetches official framework docs + writes AI systems best practices (Pydantic structured outputs, async-first, prompt discipline, context window management, cost/latency budget) - gsd-domain-researcher: researches business domain and use-case context — surfaces domain expert evaluation criteria, industry failure modes, regulatory constraints, and practitioner rubric ingredients before eval-planner writes measurable criteria - gsd-eval-planner: designs evaluation strategy grounded in domain context; defaults to Arize Phoenix (tracing) + RAGAS (RAG eval) with detect-first guard for existing tooling - gsd-eval-auditor: retroactive codebase scan → scores eval coverage Integration points: - plan-phase: non-blocking nudge (step 4.5) when AI keywords detected and no AI-SPEC.md present - settings: new workflow.ai_phase toggle (default on) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: refine ai-integration-phase layer — rename, house style, consistency fixes Amends the ai-evals framework layer (df8cb6c) with post-review improvements before opening upstream PR. Rename /gsd:ai-phase → /gsd:ai-integration-phase: - Renamed commands/gsd/ai-phase.md → ai-integration-phase.md - Renamed get-shit-done/workflows/ai-phase.md → ai-integration-phase.md - Updated config key: workflow.ai_phase → workflow.ai_integration_phase - Updated repair action: addAiPhaseKey → addAiIntegrationPhaseKey - Updated all 84 cross-references across agents, workflows, templates, tests Consistency fixes (same class as PR #1380 review): - commands/gsd: objective described 3-agent chain, missing gsd-domain-researcher - workflows/ai-integration-phase: purpose tag described 3-agent chain + "locks three things" — updated to 4 agents + 4 outputs - workflows/ai-integration-phase: missing DOMAIN_MODEL resolve-model call in step 1 (domain-researcher was spawned in step 7.5 with no model variable) - workflows/ai-integration-phase: fractional step ## 7.5 renumbered to integers (steps 8–12 shifted) Agent house style (GSD meta-prompting conformance): - All 5 new agents refactored to execution_flow + step name="" structure - Role blocks compressed to 2 lines (removed verbose "Core responsibilities") - Added skills: frontmatter to all 5 agents (agent-frontmatter tests) - Added # hooks: commented pattern to file-writing agents - Added ALWAYS use Write tool anti-heredoc instruction to file-writing agents - Line reductions: ai-researcher −41%, domain-researcher −25%, eval-planner −26%, eval-auditor −25%, framework-selector −9% Test coverage (tests/ai-evals.test.cjs — 48 tests): - CONFIG: workflow.ai_integration_phase defaults and config-set/get - HEALTH: W010 warning emission and addAiIntegrationPhaseKey repair - TEMPLATE: AI-SPEC.md section completeness (10 sections) - COMMAND: ai-integration-phase + eval-review frontmatter validity - AGENTS: all 5 new agent files exist - REFERENCES: ai-evals.md + ai-frameworks.md exist and are non-empty - WORKFLOW: plan-phase nudge integration, workflow files exist + agent coverage 603/603 tests passing. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: add Google ADK to framework selector and reference matrix Google ADK (released March 2025) was missing from the framework options. Adds Python + Java multi-agent framework optimised for Gemini / Vertex AI. - get-shit-done/references/ai-frameworks.md: add Google ADK profile (type, language, model support, best for, avoid if, strengths, weaknesses, eval concerns); update Quick Picks, By System Type, and By Model Commitment tables - agents/gsd-framework-selector.md: add "Google (Gemini)" to model provider interview question - agents/gsd-ai-researcher.md: add Google ADK docs URL to documentation_sources Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions post-rebase - Remove skills: frontmatter from all 5 new agents (upstream changed convention — skills: breaks Gemini CLI and must not be present) - Add workflow.ai_integration_phase to VALID_CONFIG_KEYS whitelist in config.cjs (config-set blocked unknown keys) - Add ai_integration_phase: true to CONFIG_DEFAULTS in core.cjs Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: rephrase 4b.1 line to avoid false-positive in prompt-injection scan "contract as a Pydantic model" matched the `act as a` pattern case-insensitively. Rephrased to "output schema using a Pydantic model". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions (W016, colon refs, config docs) - Replace verify.cjs from upstream to restore W010-W015 + cmdValidateAgents, lost when rebase conflict was resolved with --theirs - Add W016 (workflow.ai_integration_phase absent) inside the config try block, avoids collision with upstream's W010 agent-installation check - Add addAiIntegrationPhaseKey repair case mirroring addNyquistKey pattern - Replace /gsd: colon format with /gsd- hyphen format across all new files (agents, workflows, templates, verify.cjs) per stale-colon-refs guard (#1748) - Add workflow.ai_integration_phase to planning-config.md reference table - Add ai_integration_phase → workflow.ai_integration_phase to NAMESPACE_MAP in config-field-docs.test.cjs so CONFIG_DEFAULTS coverage check passes - Update ai-evals tests to use W016 instead of W010 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: add 5 new agents to E2E Copilot install expected list gsd-ai-researcher, gsd-domain-researcher, gsd-eval-auditor, gsd-eval-planner, gsd-framework-selector added to the hardcoded expected agent list in copilot-install.test.cjs (#1890). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
11 KiB
11 KiB
AI Framework Decision Matrix
Reference used by
gsd-framework-selectorandgsd-ai-researcher. Distilled from official docs, benchmarks, and developer reports (2026).
Quick Picks
| Situation | Pick |
|---|---|
| Simplest path to a working agent (OpenAI) | OpenAI Agents SDK |
| Simplest path to a working agent (model-agnostic) | CrewAI |
| Production RAG / document Q&A | LlamaIndex |
| Complex stateful workflows with branching | LangGraph |
| Multi-agent teams with defined roles | CrewAI |
| Code-aware autonomous agents (Anthropic) | Claude Agent SDK |
| "I don't know my requirements yet" | LangChain |
| Regulated / audit-trail required | LangGraph |
| Enterprise Microsoft/.NET shops | AutoGen/AG2 |
| Google Cloud / Gemini-committed teams | Google ADK |
| Pure NLP pipelines with explicit control | Haystack |
Framework Profiles
CrewAI
- Type: Multi-agent orchestration
- Language: Python only
- Model support: Model-agnostic
- Learning curve: Beginner (role/task/crew maps to real teams)
- Best for: Content pipelines, research automation, business process workflows, rapid prototyping
- Avoid if: Fine-grained state management, TypeScript, fault-tolerant checkpointing, complex conditional branching
- Strengths: Fastest multi-agent prototyping, 5.76x faster than LangGraph on QA tasks, built-in memory (short/long/entity/contextual), Flows architecture, standalone (no LangChain dep)
- Weaknesses: Limited checkpointing, coarse error handling, Python only
- Eval concerns: Task decomposition accuracy, inter-agent handoff, goal completion rate, loop detection
LlamaIndex
- Type: RAG and data ingestion
- Language: Python + TypeScript
- Model support: Model-agnostic
- Learning curve: Intermediate
- Best for: Legal research, internal knowledge assistants, enterprise document search, any system where retrieval quality is the #1 priority
- Avoid if: Primary need is agent orchestration, multi-agent collaboration, or chatbot conversation flow
- Strengths: Best-in-class document parsing (LlamaParse), 35% retrieval accuracy improvement, 20-30% faster queries, mixed retrieval strategies (vector + graph + reranker)
- Weaknesses: Data framework first — agent orchestration is secondary
- Eval concerns: Context faithfulness, hallucination, answer relevance, retrieval precision/recall
LangChain
- Type: General-purpose LLM framework
- Language: Python + TypeScript
- Model support: Model-agnostic (widest ecosystem)
- Learning curve: Intermediate–Advanced
- Best for: Evolving requirements, many third-party integrations, teams wanting one framework for everything, RAG + agents + chains
- Avoid if: Simple well-defined use case, RAG-primary (use LlamaIndex), complex stateful workflows (use LangGraph), performance at scale is critical
- Strengths: Largest community and integration ecosystem, 25% faster development vs scratch, covers RAG/agents/chains/memory
- Weaknesses: Abstraction overhead, p99 latency degrades under load, complexity creep risk
- Eval concerns: End-to-end task completion, chain correctness, retrieval quality
LangGraph
- Type: Stateful agent workflows (graph-based)
- Language: Python + TypeScript (full parity)
- Model support: Model-agnostic (inherits LangChain integrations)
- Learning curve: Intermediate–Advanced (graph mental model)
- Best for: Production-grade stateful workflows, regulated industries, audit trails, human-in-the-loop flows, fault-tolerant multi-step agents
- Avoid if: Simple chatbot, purely linear workflow, rapid prototyping
- Strengths: Best checkpointing (every node), time-travel debugging, native Postgres/Redis persistence, streaming support, chosen by 62% of developers for stateful agent work (2026)
- Weaknesses: More upfront scaffolding, steeper curve, overkill for simple cases
- Eval concerns: State transition correctness, goal completion rate, tool use accuracy, safety guardrails
OpenAI Agents SDK
- Type: Native OpenAI agent framework
- Language: Python + TypeScript
- Model support: Optimized for OpenAI (supports 100+ via Chat Completions compatibility)
- Learning curve: Beginner (4 primitives: Agents, Handoffs, Guardrails, Tracing)
- Best for: OpenAI-committed teams, rapid agent prototyping, voice agents (gpt-realtime), teams wanting visual builder (AgentKit)
- Avoid if: Model flexibility needed, complex multi-agent collaboration, persistent state management required, vendor lock-in concern
- Strengths: Simplest mental model, built-in tracing and guardrails, Handoffs for agent delegation, Realtime Agents for voice
- Weaknesses: OpenAI vendor lock-in, no built-in persistent state, younger ecosystem
- Eval concerns: Instruction following, safety guardrails, escalation accuracy, tone consistency
Claude Agent SDK (Anthropic)
- Type: Code-aware autonomous agent framework
- Language: Python + TypeScript
- Model support: Claude models only
- Learning curve: Intermediate (18 hook events, MCP, tool decorators)
- Best for: Developer tooling, code generation/review agents, autonomous coding assistants, MCP-heavy architectures, safety-critical applications
- Avoid if: Model flexibility needed, stable/mature API required, use case unrelated to code/tool-use
- Strengths: Deepest MCP integration, built-in filesystem/shell access, 18 lifecycle hooks, automatic context compaction, extended thinking, safety-first design
- Weaknesses: Claude-only vendor lock-in, newer/evolving API, smaller community
- Eval concerns: Tool use correctness, safety, code quality, instruction following
AutoGen / AG2 / Microsoft Agent Framework
- Type: Multi-agent conversational framework
- Language: Python (AG2), Python + .NET (Microsoft Agent Framework)
- Model support: Model-agnostic
- Learning curve: Intermediate–Advanced
- Best for: Research applications, conversational problem-solving, code generation + execution loops, Microsoft/.NET shops
- Avoid if: You want ecosystem stability, deterministic workflows, or "safest long-term bet" (fragmentation risk)
- Strengths: Most sophisticated conversational agent patterns, code generation + execution loop, async event-driven (v0.4+), cross-language interop (Microsoft Agent Framework)
- Weaknesses: Ecosystem fragmented (AutoGen maintenance mode, AG2 fork, Microsoft Agent Framework preview) — genuine long-term risk
- Eval concerns: Conversation goal completion, consensus quality, code execution correctness
Google ADK (Agent Development Kit)
- Type: Multi-agent orchestration framework
- Language: Python + Java
- Model support: Optimized for Gemini; supports other models via LiteLLM
- Learning curve: Intermediate (agent/tool/session model, familiar if you know LangGraph)
- Best for: Google Cloud / Vertex AI shops, multi-agent workflows needing built-in session management and memory, teams already committed to Gemini, agent pipelines that need Google Search / BigQuery tool integration
- Avoid if: Model flexibility is required beyond Gemini, no Google Cloud dependency acceptable, TypeScript-only stack
- Strengths: First-party Google support, built-in session/memory/artifact management, tight Vertex AI and Google Search integration, own eval framework (RAGAS-compatible), multi-agent by design (sequential, parallel, loop patterns), Java SDK for enterprise teams
- Weaknesses: Gemini vendor lock-in in practice, younger community than LangChain/LlamaIndex, less third-party integration depth
- Eval concerns: Multi-agent task decomposition, tool use correctness, session state consistency, goal completion rate
Haystack
- Type: NLP pipeline framework
- Language: Python
- Model support: Model-agnostic
- Learning curve: Intermediate
- Best for: Explicit, auditable NLP pipelines, document processing with fine-grained control, enterprise search, regulated industries needing transparency
- Avoid if: Rapid prototyping, multi-agent workflows, or you want a large community
- Strengths: Explicit pipeline control, strong for structured data pipelines, good documentation
- Weaknesses: Smaller community, less agent-oriented than alternatives
- Eval concerns: Extraction accuracy, pipeline output validity, retrieval quality
Decision Dimensions
By System Type
| System Type | Primary Framework(s) | Key Eval Concerns |
|---|---|---|
| RAG / Knowledge Q&A | LlamaIndex, LangChain | Context faithfulness, hallucination, retrieval precision/recall |
| Multi-agent orchestration | CrewAI, LangGraph, Google ADK | Task decomposition, handoff quality, goal completion |
| Conversational assistants | OpenAI Agents SDK, Claude Agent SDK | Tone, safety, instruction following, escalation |
| Structured data extraction | LangChain, LlamaIndex | Schema compliance, extraction accuracy |
| Autonomous task agents | LangGraph, OpenAI Agents SDK | Safety guardrails, tool correctness, cost adherence |
| Content generation | Claude Agent SDK, OpenAI Agents SDK | Brand voice, factual accuracy, tone |
| Code automation | Claude Agent SDK | Code correctness, safety, test pass rate |
By Team Size and Stage
| Context | Recommendation |
|---|---|
| Solo dev, prototyping | OpenAI Agents SDK or CrewAI (fastest to running) |
| Solo dev, RAG | LlamaIndex (batteries included) |
| Team, production, stateful | LangGraph (best fault tolerance) |
| Team, evolving requirements | LangChain (broadest escape hatches) |
| Team, multi-agent | CrewAI (simplest role abstraction) |
| Enterprise, .NET | AutoGen AG2 / Microsoft Agent Framework |
By Model Commitment
| Preference | Framework |
|---|---|
| OpenAI-only | OpenAI Agents SDK |
| Anthropic/Claude-only | Claude Agent SDK |
| Google/Gemini-committed | Google ADK |
| Model-agnostic (full flexibility) | LangChain, LlamaIndex, CrewAI, LangGraph, Haystack |
Anti-Patterns
- Using LangChain for simple chatbots — Direct SDK call is less code, faster, and easier to debug
- Using CrewAI for complex stateful workflows — Checkpointing gaps will bite you in production
- Using OpenAI Agents SDK with non-OpenAI models — Loses the integration benefits you chose it for
- Using LlamaIndex as a multi-agent framework — It can do agents, but that's not its strength
- Defaulting to LangChain without evaluating alternatives — "Everyone uses it" ≠ right for your use case
- Starting a new project on AutoGen (not AG2) — AutoGen is in maintenance mode; use AG2 or wait for Microsoft Agent Framework GA
- Choosing LangGraph for simple linear flows — The graph overhead is not worth it; use LangChain chains instead
- Ignoring vendor lock-in — Provider-native SDKs (OpenAI, Claude) trade flexibility for integration depth; decide consciously
Combination Plays (Multi-Framework Stacks)
| Production Pattern | Stack |
|---|---|
| RAG with observability | LlamaIndex + LangSmith or Langfuse |
| Stateful agent with RAG | LangGraph + LlamaIndex |
| Multi-agent with tracing | CrewAI + Langfuse |
| OpenAI agents with evals | OpenAI Agents SDK + Promptfoo or Braintrust |
| Claude agents with MCP | Claude Agent SDK + LangSmith or Arize Phoenix |