* feat: /gsd:ai-phase + /gsd:eval-review — AI evals and framework selection layer Adds a structured AI development layer to GSD with 5 new agents, 2 new commands, 2 new workflows, 2 reference files, and 1 template. Commands: - /gsd:ai-phase [N] — pre-planning AI design contract (inserts between discuss-phase and plan-phase). Orchestrates 4 agents in sequence: framework-selector → ai-researcher → domain-researcher → eval-planner. Output: AI-SPEC.md with framework decision, implementation guidance, domain expert context, and evaluation strategy. - /gsd:eval-review [N] — retroactive eval coverage audit. Scores each planned eval dimension as COVERED/PARTIAL/MISSING. Output: EVAL-REVIEW.md with 0-100 score, verdict, and remediation plan. Agents: - gsd-framework-selector: interactive decision matrix (6 questions) → scored framework recommendation for CrewAI, LlamaIndex, LangChain, LangGraph, OpenAI Agents SDK, Claude Agent SDK, AutoGen/AG2, Haystack - gsd-ai-researcher: fetches official framework docs + writes AI systems best practices (Pydantic structured outputs, async-first, prompt discipline, context window management, cost/latency budget) - gsd-domain-researcher: researches business domain and use-case context — surfaces domain expert evaluation criteria, industry failure modes, regulatory constraints, and practitioner rubric ingredients before eval-planner writes measurable criteria - gsd-eval-planner: designs evaluation strategy grounded in domain context; defaults to Arize Phoenix (tracing) + RAGAS (RAG eval) with detect-first guard for existing tooling - gsd-eval-auditor: retroactive codebase scan → scores eval coverage Integration points: - plan-phase: non-blocking nudge (step 4.5) when AI keywords detected and no AI-SPEC.md present - settings: new workflow.ai_phase toggle (default on) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: refine ai-integration-phase layer — rename, house style, consistency fixes Amends the ai-evals framework layer (df8cb6c) with post-review improvements before opening upstream PR. Rename /gsd:ai-phase → /gsd:ai-integration-phase: - Renamed commands/gsd/ai-phase.md → ai-integration-phase.md - Renamed get-shit-done/workflows/ai-phase.md → ai-integration-phase.md - Updated config key: workflow.ai_phase → workflow.ai_integration_phase - Updated repair action: addAiPhaseKey → addAiIntegrationPhaseKey - Updated all 84 cross-references across agents, workflows, templates, tests Consistency fixes (same class as PR #1380 review): - commands/gsd: objective described 3-agent chain, missing gsd-domain-researcher - workflows/ai-integration-phase: purpose tag described 3-agent chain + "locks three things" — updated to 4 agents + 4 outputs - workflows/ai-integration-phase: missing DOMAIN_MODEL resolve-model call in step 1 (domain-researcher was spawned in step 7.5 with no model variable) - workflows/ai-integration-phase: fractional step ## 7.5 renumbered to integers (steps 8–12 shifted) Agent house style (GSD meta-prompting conformance): - All 5 new agents refactored to execution_flow + step name="" structure - Role blocks compressed to 2 lines (removed verbose "Core responsibilities") - Added skills: frontmatter to all 5 agents (agent-frontmatter tests) - Added # hooks: commented pattern to file-writing agents - Added ALWAYS use Write tool anti-heredoc instruction to file-writing agents - Line reductions: ai-researcher −41%, domain-researcher −25%, eval-planner −26%, eval-auditor −25%, framework-selector −9% Test coverage (tests/ai-evals.test.cjs — 48 tests): - CONFIG: workflow.ai_integration_phase defaults and config-set/get - HEALTH: W010 warning emission and addAiIntegrationPhaseKey repair - TEMPLATE: AI-SPEC.md section completeness (10 sections) - COMMAND: ai-integration-phase + eval-review frontmatter validity - AGENTS: all 5 new agent files exist - REFERENCES: ai-evals.md + ai-frameworks.md exist and are non-empty - WORKFLOW: plan-phase nudge integration, workflow files exist + agent coverage 603/603 tests passing. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: add Google ADK to framework selector and reference matrix Google ADK (released March 2025) was missing from the framework options. Adds Python + Java multi-agent framework optimised for Gemini / Vertex AI. - get-shit-done/references/ai-frameworks.md: add Google ADK profile (type, language, model support, best for, avoid if, strengths, weaknesses, eval concerns); update Quick Picks, By System Type, and By Model Commitment tables - agents/gsd-framework-selector.md: add "Google (Gemini)" to model provider interview question - agents/gsd-ai-researcher.md: add Google ADK docs URL to documentation_sources Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions post-rebase - Remove skills: frontmatter from all 5 new agents (upstream changed convention — skills: breaks Gemini CLI and must not be present) - Add workflow.ai_integration_phase to VALID_CONFIG_KEYS whitelist in config.cjs (config-set blocked unknown keys) - Add ai_integration_phase: true to CONFIG_DEFAULTS in core.cjs Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: rephrase 4b.1 line to avoid false-positive in prompt-injection scan "contract as a Pydantic model" matched the `act as a` pattern case-insensitively. Rephrased to "output schema using a Pydantic model". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: adapt to upstream conventions (W016, colon refs, config docs) - Replace verify.cjs from upstream to restore W010-W015 + cmdValidateAgents, lost when rebase conflict was resolved with --theirs - Add W016 (workflow.ai_integration_phase absent) inside the config try block, avoids collision with upstream's W010 agent-installation check - Add addAiIntegrationPhaseKey repair case mirroring addNyquistKey pattern - Replace /gsd: colon format with /gsd- hyphen format across all new files (agents, workflows, templates, verify.cjs) per stale-colon-refs guard (#1748) - Add workflow.ai_integration_phase to planning-config.md reference table - Add ai_integration_phase → workflow.ai_integration_phase to NAMESPACE_MAP in config-field-docs.test.cjs so CONFIG_DEFAULTS coverage check passes - Update ai-evals tests to use W016 instead of W010 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: add 5 new agents to E2E Copilot install expected list gsd-ai-researcher, gsd-domain-researcher, gsd-eval-auditor, gsd-eval-planner, gsd-framework-selector added to the hardcoded expected agent list in copilot-install.test.cjs (#1890). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
157 lines
8.5 KiB
Markdown
157 lines
8.5 KiB
Markdown
# AI Evaluation Reference
|
|
|
|
> Reference used by `gsd-eval-planner` and `gsd-eval-auditor`.
|
|
> Based on "AI Evals for Everyone" course (Reganti & Badam) + industry practice.
|
|
|
|
---
|
|
|
|
## Core Concepts
|
|
|
|
### Why Evals Exist
|
|
AI systems are non-deterministic. Input X does not reliably produce output Y across runs, users, or edge cases. Evals are the continuous process of assessing whether your system's behavior meets expectations under real-world conditions — unit tests and integration tests alone are insufficient.
|
|
|
|
### Model vs. Product Evaluation
|
|
- **Model evals** (MMLU, HumanEval, GSM8K) — measure general capability in standardized conditions. Use as initial filter only.
|
|
- **Product evals** — measure behavior inside your specific system, with your data, your users, your domain rules. This is where 80% of eval effort belongs.
|
|
|
|
### The Three Components of Every Eval
|
|
- **Input** — everything affecting the system: query, history, retrieved docs, system prompt, config
|
|
- **Expected** — what good behavior looks like, defined through rubrics
|
|
- **Actual** — what the system produced, including intermediate steps, tool calls, and reasoning traces
|
|
|
|
### Three Measurement Approaches
|
|
1. **Code-based metrics** — deterministic checks: JSON validation, required disclaimers, performance thresholds, classification flags. Fast, cheap, reliable. Use first.
|
|
2. **LLM judges** — one model evaluates another against a rubric. Powerful for subjective qualities (tone, reasoning, escalation). Requires calibration against human judgment before trusting.
|
|
3. **Human evaluation** — gold standard for nuanced judgment. Doesn't scale. Use for calibration, edge cases, periodic sampling, and high-stakes decisions.
|
|
|
|
Most effective systems combine all three.
|
|
|
|
---
|
|
|
|
## Evaluation Dimensions
|
|
|
|
### Pre-Deployment (Development Phase)
|
|
|
|
| Dimension | What It Measures | When It Matters |
|
|
|-----------|-----------------|-----------------|
|
|
| **Factual accuracy** | Correctness of claims against ground truth | RAG, knowledge bases, any factual assertions |
|
|
| **Context faithfulness** | Response grounded in provided context vs. fabricated | RAG pipelines, document Q&A, retrieval-augmented systems |
|
|
| **Hallucination detection** | Plausible but unsupported claims | All generative systems, high-stakes domains |
|
|
| **Escalation accuracy** | Correct identification of when human intervention needed | Customer service, healthcare, financial advisory |
|
|
| **Policy compliance** | Adherence to business rules, legal requirements, disclaimers | Regulated industries, enterprise deployments |
|
|
| **Tone/style appropriateness** | Match with brand voice, audience expectations, emotional context | Customer-facing systems, content generation |
|
|
| **Output structure validity** | Schema compliance, required fields, format correctness | Structured extraction, API integrations, data pipelines |
|
|
| **Task completion** | Whether the system accomplished the stated goal | Agentic workflows, multi-step tasks |
|
|
| **Tool use correctness** | Correct selection and invocation of tools | Agent systems with tool calls |
|
|
| **Safety** | Absence of harmful, biased, or inappropriate outputs | All user-facing systems |
|
|
|
|
### Production Monitoring
|
|
|
|
| Dimension | Monitoring Approach |
|
|
|-----------|---------------------|
|
|
| **Safety violations** | Online guardrail — real-time, immediate intervention |
|
|
| **Compliance failures** | Online guardrail — block or escalate before user sees output |
|
|
| **Quality degradation trends** | Offline flywheel — batch analysis of sampled interactions |
|
|
| **Emerging failure modes** | Signal-metric divergence — when user behavior signals diverge from metric scores, investigate manually |
|
|
| **Cost/latency drift** | Code-based metrics — automated threshold alerts |
|
|
|
|
---
|
|
|
|
## The Guardrail vs. Flywheel Decision
|
|
|
|
Ask: "If this behavior goes wrong, would it be catastrophic for my business?"
|
|
|
|
- **Yes → Guardrail** — run online, real-time, with immediate intervention (block, escalate, hand off). Be selective: guardrails add latency.
|
|
- **No → Flywheel** — run offline as batch analysis feeding system refinements over time.
|
|
|
|
---
|
|
|
|
## Rubric Design
|
|
|
|
Generic metrics are meaningless without context. "Helpfulness" in real estate means summarizing listings clearly. In healthcare it means knowing when *not* to answer.
|
|
|
|
A rubric must define:
|
|
1. The dimension being measured
|
|
2. What scores 1, 3, and 5 on a 5-point scale (or pass/fail criteria)
|
|
3. Domain-specific examples of acceptable vs. unacceptable behavior
|
|
|
|
Without rubrics, LLM judges produce noise rather than signal.
|
|
|
|
---
|
|
|
|
## Reference Dataset Guidelines
|
|
|
|
- Start with **10-20 high-quality examples** — not 200 mediocre ones
|
|
- Cover: critical success scenarios, common user workflows, known edge cases, historical failure modes
|
|
- Have domain experts label the examples (not just engineers)
|
|
- Expand based on what you learn in production — don't build for hypothetical coverage
|
|
|
|
---
|
|
|
|
## Eval Tooling Guide
|
|
|
|
| Tool | Type | Best For | Key Strength |
|
|
|------|------|----------|-------------|
|
|
| **RAGAS** | Python library | RAG evaluation | Purpose-built metrics: faithfulness, answer relevance, context precision/recall |
|
|
| **Langfuse** | Platform (open-source, self-hostable) | All system types | Strong tracing, prompt management, good for teams wanting infrastructure control |
|
|
| **LangSmith** | Platform (commercial) | LangChain/LangGraph ecosystems | Tightest integration with LangChain; best if already in that ecosystem |
|
|
| **Arize Phoenix** | Platform (open-source + hosted) | RAG + multi-agent tracing | Strong RAG eval + trace visualization; open-source with hosted option |
|
|
| **Braintrust** | Platform (commercial) | Model-agnostic evaluation | Dataset and experiment management; good for comparing across frameworks |
|
|
| **Promptfoo** | CLI tool (open-source) | Prompt testing, CI/CD | CLI-first, excellent for CI/CD prompt regression testing |
|
|
|
|
### Tool Selection by System Type
|
|
|
|
| System Type | Recommended Tooling |
|
|
|-------------|---------------------|
|
|
| RAG / Knowledge Q&A | RAGAS + Arize Phoenix or Braintrust |
|
|
| Multi-agent systems | Langfuse + Arize Phoenix |
|
|
| Conversational / single-model | Promptfoo + Braintrust |
|
|
| Structured extraction | Promptfoo + code-based validators |
|
|
| LangChain/LangGraph projects | LangSmith (native integration) |
|
|
| Production monitoring (all types) | Langfuse, Arize Phoenix, or LangSmith |
|
|
|
|
---
|
|
|
|
## Evals in the Development Lifecycle
|
|
|
|
### Plan Phase (Evaluation-Aware Design)
|
|
Before writing code, define:
|
|
1. What type of AI system is being built → determines framework and dominant eval concerns
|
|
2. Critical failure modes (3-5 behaviors that cannot go wrong)
|
|
3. Rubrics — explicit definitions of acceptable/unacceptable behavior per dimension
|
|
4. Evaluation strategy — which dimensions use code metrics, LLM judges, or human review
|
|
5. Reference dataset requirements — size, composition, labeling approach
|
|
6. Eval tooling selection
|
|
|
|
Output: EVALS-SPEC section of AI-SPEC.md
|
|
|
|
### Execute Phase (Instrument While Building)
|
|
- Add tracing from day one (Langfuse, Arize Phoenix, or LangSmith)
|
|
- Build reference dataset concurrently with implementation
|
|
- Implement code-based checks first; add LLM judges only for subjective dimensions
|
|
- Run evals in CI/CD via Promptfoo or Braintrust
|
|
|
|
### Verify Phase (Pre-Deployment Validation)
|
|
- Run full reference dataset against all metrics
|
|
- Conduct human review of edge cases and LLM judge disagreements
|
|
- Calibrate LLM judges against human scores (target ≥ 0.7 correlation before trusting)
|
|
- Define and configure production guardrails
|
|
- Establish monitoring baseline
|
|
|
|
### Monitor Phase (Production Evaluation Loop)
|
|
- Smart sampling — weight toward interactions with concerning signals (retries, unusual length, explicit escalations)
|
|
- Online guardrails on every interaction
|
|
- Offline flywheel on sampled batch
|
|
- Watch for signal-metric divergence — the early warning system for evaluation gaps
|
|
|
|
---
|
|
|
|
## Common Pitfalls
|
|
|
|
1. **Assuming benchmarks predict product success** — they don't; model evals are a filter, not a verdict
|
|
2. **Engineering evals in isolation** — domain experts must co-define rubrics; engineers alone miss critical nuances
|
|
3. **Building comprehensive coverage on day one** — start small (10-20 examples), expand from real failure modes
|
|
4. **Trusting uncalibrated LLM judges** — validate against human judgment before relying on them
|
|
5. **Measuring everything** — only track metrics that drive decisions; "collect it all" produces noise
|
|
6. **Treating evaluation as one-time setup** — user behavior evolves, requirements change, failure modes emerge; evaluation is continuous
|