* feat: /gsd:ai-phase + /gsd:eval-review — AI evals and framework selection layer Adds a structured AI development layer to GSD with 5 new agents, 2 new commands, 2 new workflows, 2 reference files, and 1 template. Commands: - /gsd:ai-phase [N] — pre-planning AI design contract (inserts between discuss-phase and plan-phase). Orchestrates 4 agents in sequence: framework-selector → ai-researcher → domain-researcher → eval-planner. Output: AI-SPEC.md with framework decision, implementation guidance, domain expert context, and evaluation strategy. - /gsd:eval-review [N] — retroactive eval coverage audit. Scores each planned eval dimension as COVERED/PARTIAL/MISSING. Output: EVAL-REVIEW.md with 0-100 score, verdict, and remediation plan. Agents: - gsd-framework-selector: interactive decision matrix (6 questions) → scored framework recommendation for CrewAI, LlamaIndex, LangChain, LangGraph, OpenAI Agents SDK, Claude Agent SDK, AutoGen/AG2, Haystack - gsd-ai-researcher: fetches official framework docs + writes AI systems best practices (Pydantic structured outputs, async-first, prompt discipline, context window management, cost/latency budget) - gsd-domain-researcher: researches business domain and use-case context — surfaces domain expert evaluation criteria, industry failure modes, regulatory constraints, and practitioner rubric ingredients before eval-planner writes measurable criteria - gsd-eval-planner: designs evaluation strategy grounded in domain context; defaults to Arize Phoenix (tracing) + RAGAS (RAG eval) with detect-first guard for existing tooling - gsd-eval-auditor: retroactive codebase scan → scores eval coverage Integration points: - plan-phase: non-blocking nudge (step 4.5) when AI keywords detected and no AI-SPEC.md present - settings: new workflow.ai_phase toggle (default on) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: refine ai-integration-phase layer — rename, house style, consistency fixes Amends the ai-evals framework layer (df8cb6c) with post-review improvements before opening upstream PR. Rename /gsd:ai-phase → /gsd:ai-integration-phase: - Renamed commands/gsd/ai-phase.md → ai-integration-phase.md - Renamed get-shit-done/workflows/ai-phase.md → ai-integration-phase.md - Updated config key: workflow.ai_phase → workflow.ai_integration_phase - Updated repair action: addAiPhaseKey → addAiIntegrationPhaseKey - Updated all 84 cross-references across agents, workflows, templates, tests Consistency fixes (same class as PR #1380 review): - commands/gsd: objective described 3-agent chain, missing gsd-domain-researcher - workflows/ai-integration-phase: purpose tag described 3-agent chain + "locks three things" — updated to 4 agents + 4 outputs - workflows/ai-integration-phase: missing DOMAIN_MODEL resolve-model call in step 1 (domain-researcher was spawned in step 7.5 with no model variable) - workflows/ai-integration-phase: fractional step ## 7.5 renumbered to integers (steps 8–12 shifted) Agent house style (GSD meta-prompting conformance): - All 5 new agents refactored to execution_flow + step name="" structure - Role blocks compressed to 2 lines (removed verbose "Core responsibilities") - Added skills: frontmatter to all 5 agents (agent-frontmatter tests) - Added # hooks: commented pattern to file-writing agents - Added ALWAYS use Write tool anti-heredoc instruction to file-writing agents - Line reductions: ai-researcher −41%, domain-researcher −25%, eval-planner −26%, eval-auditor −25%, framework-selector −9% Test coverage (tests/ai-evals.test.cjs — 48 tests): - CONFIG: workflow.ai_integration_phase defaults and config-set/get - HEALTH: W010 warning emission and addAiIntegrationPhaseKey repair - TEMPLATE: AI-SPEC.md section completeness (10 sections) - COMMAND: ai-integration-phase + eval-review frontmatter validity - AGENTS: all 5 new agent files exist - REFERENCES: ai-evals.md + ai-frameworks.md exist and are non-empty - WORKFLOW: plan-phase nudge integration, workflow files exist + agent coverage 603/603 tests passing. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: add Google ADK to framework selector and reference matrix Google ADK (released March 2025) was missing from the framework options. Adds Python + Java multi-agent framework optimised for Gemini / Vertex AI. - get-shit-done/references/ai-frameworks.md: add Google ADK profile (type, language, model support, best for, avoid if, strengths, weaknesses, eval concerns); update Quick Picks, By System Type, and By Model Commitment tables - agents/gsd-framework-selector.md: add "Google (Gemini)" to model provider interview question - agents/gsd-ai-researcher.md: add Google ADK docs URL to documentation_sources Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions post-rebase - Remove skills: frontmatter from all 5 new agents (upstream changed convention — skills: breaks Gemini CLI and must not be present) - Add workflow.ai_integration_phase to VALID_CONFIG_KEYS whitelist in config.cjs (config-set blocked unknown keys) - Add ai_integration_phase: true to CONFIG_DEFAULTS in core.cjs Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: rephrase 4b.1 line to avoid false-positive in prompt-injection scan "contract as a Pydantic model" matched the `act as a` pattern case-insensitively. Rephrased to "output schema using a Pydantic model". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions (W016, colon refs, config docs) - Replace verify.cjs from upstream to restore W010-W015 + cmdValidateAgents, lost when rebase conflict was resolved with --theirs - Add W016 (workflow.ai_integration_phase absent) inside the config try block, avoids collision with upstream's W010 agent-installation check - Add addAiIntegrationPhaseKey repair case mirroring addNyquistKey pattern - Replace /gsd: colon format with /gsd- hyphen format across all new files (agents, workflows, templates, verify.cjs) per stale-colon-refs guard (#1748) - Add workflow.ai_integration_phase to planning-config.md reference table - Add ai_integration_phase → workflow.ai_integration_phase to NAMESPACE_MAP in config-field-docs.test.cjs so CONFIG_DEFAULTS coverage check passes - Update ai-evals tests to use W016 instead of W010 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: add 5 new agents to E2E Copilot install expected list gsd-ai-researcher, gsd-domain-researcher, gsd-eval-auditor, gsd-eval-planner, gsd-framework-selector added to the hardcoded expected agent list in copilot-install.test.cjs (#1890). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
8.5 KiB
AI Evaluation Reference
Reference used by
gsd-eval-plannerandgsd-eval-auditor. Based on "AI Evals for Everyone" course (Reganti & Badam) + industry practice.
Core Concepts
Why Evals Exist
AI systems are non-deterministic. Input X does not reliably produce output Y across runs, users, or edge cases. Evals are the continuous process of assessing whether your system's behavior meets expectations under real-world conditions — unit tests and integration tests alone are insufficient.
Model vs. Product Evaluation
- Model evals (MMLU, HumanEval, GSM8K) — measure general capability in standardized conditions. Use as initial filter only.
- Product evals — measure behavior inside your specific system, with your data, your users, your domain rules. This is where 80% of eval effort belongs.
The Three Components of Every Eval
- Input — everything affecting the system: query, history, retrieved docs, system prompt, config
- Expected — what good behavior looks like, defined through rubrics
- Actual — what the system produced, including intermediate steps, tool calls, and reasoning traces
Three Measurement Approaches
- Code-based metrics — deterministic checks: JSON validation, required disclaimers, performance thresholds, classification flags. Fast, cheap, reliable. Use first.
- LLM judges — one model evaluates another against a rubric. Powerful for subjective qualities (tone, reasoning, escalation). Requires calibration against human judgment before trusting.
- Human evaluation — gold standard for nuanced judgment. Doesn't scale. Use for calibration, edge cases, periodic sampling, and high-stakes decisions.
Most effective systems combine all three.
Evaluation Dimensions
Pre-Deployment (Development Phase)
| Dimension | What It Measures | When It Matters |
|---|---|---|
| Factual accuracy | Correctness of claims against ground truth | RAG, knowledge bases, any factual assertions |
| Context faithfulness | Response grounded in provided context vs. fabricated | RAG pipelines, document Q&A, retrieval-augmented systems |
| Hallucination detection | Plausible but unsupported claims | All generative systems, high-stakes domains |
| Escalation accuracy | Correct identification of when human intervention needed | Customer service, healthcare, financial advisory |
| Policy compliance | Adherence to business rules, legal requirements, disclaimers | Regulated industries, enterprise deployments |
| Tone/style appropriateness | Match with brand voice, audience expectations, emotional context | Customer-facing systems, content generation |
| Output structure validity | Schema compliance, required fields, format correctness | Structured extraction, API integrations, data pipelines |
| Task completion | Whether the system accomplished the stated goal | Agentic workflows, multi-step tasks |
| Tool use correctness | Correct selection and invocation of tools | Agent systems with tool calls |
| Safety | Absence of harmful, biased, or inappropriate outputs | All user-facing systems |
Production Monitoring
| Dimension | Monitoring Approach |
|---|---|
| Safety violations | Online guardrail — real-time, immediate intervention |
| Compliance failures | Online guardrail — block or escalate before user sees output |
| Quality degradation trends | Offline flywheel — batch analysis of sampled interactions |
| Emerging failure modes | Signal-metric divergence — when user behavior signals diverge from metric scores, investigate manually |
| Cost/latency drift | Code-based metrics — automated threshold alerts |
The Guardrail vs. Flywheel Decision
Ask: "If this behavior goes wrong, would it be catastrophic for my business?"
- Yes → Guardrail — run online, real-time, with immediate intervention (block, escalate, hand off). Be selective: guardrails add latency.
- No → Flywheel — run offline as batch analysis feeding system refinements over time.
Rubric Design
Generic metrics are meaningless without context. "Helpfulness" in real estate means summarizing listings clearly. In healthcare it means knowing when not to answer.
A rubric must define:
- The dimension being measured
- What scores 1, 3, and 5 on a 5-point scale (or pass/fail criteria)
- Domain-specific examples of acceptable vs. unacceptable behavior
Without rubrics, LLM judges produce noise rather than signal.
Reference Dataset Guidelines
- Start with 10-20 high-quality examples — not 200 mediocre ones
- Cover: critical success scenarios, common user workflows, known edge cases, historical failure modes
- Have domain experts label the examples (not just engineers)
- Expand based on what you learn in production — don't build for hypothetical coverage
Eval Tooling Guide
| Tool | Type | Best For | Key Strength |
|---|---|---|---|
| RAGAS | Python library | RAG evaluation | Purpose-built metrics: faithfulness, answer relevance, context precision/recall |
| Langfuse | Platform (open-source, self-hostable) | All system types | Strong tracing, prompt management, good for teams wanting infrastructure control |
| LangSmith | Platform (commercial) | LangChain/LangGraph ecosystems | Tightest integration with LangChain; best if already in that ecosystem |
| Arize Phoenix | Platform (open-source + hosted) | RAG + multi-agent tracing | Strong RAG eval + trace visualization; open-source with hosted option |
| Braintrust | Platform (commercial) | Model-agnostic evaluation | Dataset and experiment management; good for comparing across frameworks |
| Promptfoo | CLI tool (open-source) | Prompt testing, CI/CD | CLI-first, excellent for CI/CD prompt regression testing |
Tool Selection by System Type
| System Type | Recommended Tooling |
|---|---|
| RAG / Knowledge Q&A | RAGAS + Arize Phoenix or Braintrust |
| Multi-agent systems | Langfuse + Arize Phoenix |
| Conversational / single-model | Promptfoo + Braintrust |
| Structured extraction | Promptfoo + code-based validators |
| LangChain/LangGraph projects | LangSmith (native integration) |
| Production monitoring (all types) | Langfuse, Arize Phoenix, or LangSmith |
Evals in the Development Lifecycle
Plan Phase (Evaluation-Aware Design)
Before writing code, define:
- What type of AI system is being built → determines framework and dominant eval concerns
- Critical failure modes (3-5 behaviors that cannot go wrong)
- Rubrics — explicit definitions of acceptable/unacceptable behavior per dimension
- Evaluation strategy — which dimensions use code metrics, LLM judges, or human review
- Reference dataset requirements — size, composition, labeling approach
- Eval tooling selection
Output: EVALS-SPEC section of AI-SPEC.md
Execute Phase (Instrument While Building)
- Add tracing from day one (Langfuse, Arize Phoenix, or LangSmith)
- Build reference dataset concurrently with implementation
- Implement code-based checks first; add LLM judges only for subjective dimensions
- Run evals in CI/CD via Promptfoo or Braintrust
Verify Phase (Pre-Deployment Validation)
- Run full reference dataset against all metrics
- Conduct human review of edge cases and LLM judge disagreements
- Calibrate LLM judges against human scores (target ≥ 0.7 correlation before trusting)
- Define and configure production guardrails
- Establish monitoring baseline
Monitor Phase (Production Evaluation Loop)
- Smart sampling — weight toward interactions with concerning signals (retries, unusual length, explicit escalations)
- Online guardrails on every interaction
- Offline flywheel on sampled batch
- Watch for signal-metric divergence — the early warning system for evaluation gaps
Common Pitfalls
- Assuming benchmarks predict product success — they don't; model evals are a filter, not a verdict
- Engineering evals in isolation — domain experts must co-define rubrics; engineers alone miss critical nuances
- Building comprehensive coverage on day one — start small (10-20 examples), expand from real failure modes
- Trusting uncalibrated LLM judges — validate against human judgment before relying on them
- Measuring everything — only track metrics that drive decisions; "collect it all" produces noise
- Treating evaluation as one-time setup — user behavior evolves, requirements change, failure modes emerge; evaluation is continuous