Files
msd-core/get-shit-done/references/ai-evals.md
Fana 33575ba91d feat: /gsd-ai-integration-phase + /gsd-eval-review — AI framework selection and eval coverage layer (#1971)
* feat: /gsd:ai-phase + /gsd:eval-review — AI evals and framework selection layer

Adds a structured AI development layer to GSD with 5 new agents, 2 new
commands, 2 new workflows, 2 reference files, and 1 template.

Commands:
- /gsd:ai-phase [N] — pre-planning AI design contract (inserts between
  discuss-phase and plan-phase). Orchestrates 4 agents in sequence:
  framework-selector → ai-researcher → domain-researcher → eval-planner.
  Output: AI-SPEC.md with framework decision, implementation guidance,
  domain expert context, and evaluation strategy.
- /gsd:eval-review [N] — retroactive eval coverage audit. Scores each
  planned eval dimension as COVERED/PARTIAL/MISSING. Output: EVAL-REVIEW.md
  with 0-100 score, verdict, and remediation plan.

Agents:
- gsd-framework-selector: interactive decision matrix (6 questions) →
  scored framework recommendation for CrewAI, LlamaIndex, LangChain,
  LangGraph, OpenAI Agents SDK, Claude Agent SDK, AutoGen/AG2, Haystack
- gsd-ai-researcher: fetches official framework docs + writes AI systems
  best practices (Pydantic structured outputs, async-first, prompt
  discipline, context window management, cost/latency budget)
- gsd-domain-researcher: researches business domain and use-case context —
  surfaces domain expert evaluation criteria, industry failure modes,
  regulatory constraints, and practitioner rubric ingredients before
  eval-planner writes measurable criteria
- gsd-eval-planner: designs evaluation strategy grounded in domain context;
  defaults to Arize Phoenix (tracing) + RAGAS (RAG eval) with detect-first
  guard for existing tooling
- gsd-eval-auditor: retroactive codebase scan → scores eval coverage

Integration points:
- plan-phase: non-blocking nudge (step 4.5) when AI keywords detected and
  no AI-SPEC.md present
- settings: new workflow.ai_phase toggle (default on)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: refine ai-integration-phase layer — rename, house style, consistency fixes

Amends the ai-evals framework layer (df8cb6c) with post-review improvements
before opening upstream PR.

Rename /gsd:ai-phase → /gsd:ai-integration-phase:
- Renamed commands/gsd/ai-phase.md → ai-integration-phase.md
- Renamed get-shit-done/workflows/ai-phase.md → ai-integration-phase.md
- Updated config key: workflow.ai_phase → workflow.ai_integration_phase
- Updated repair action: addAiPhaseKey → addAiIntegrationPhaseKey
- Updated all 84 cross-references across agents, workflows, templates, tests

Consistency fixes (same class as PR #1380 review):
- commands/gsd: objective described 3-agent chain, missing gsd-domain-researcher
- workflows/ai-integration-phase: purpose tag described 3-agent chain + "locks
  three things" — updated to 4 agents + 4 outputs
- workflows/ai-integration-phase: missing DOMAIN_MODEL resolve-model call in
  step 1 (domain-researcher was spawned in step 7.5 with no model variable)
- workflows/ai-integration-phase: fractional step ## 7.5 renumbered to integers
  (steps 8–12 shifted)

Agent house style (GSD meta-prompting conformance):
- All 5 new agents refactored to execution_flow + step name="" structure
- Role blocks compressed to 2 lines (removed verbose "Core responsibilities")
- Added skills: frontmatter to all 5 agents (agent-frontmatter tests)
- Added # hooks: commented pattern to file-writing agents
- Added ALWAYS use Write tool anti-heredoc instruction to file-writing agents
- Line reductions: ai-researcher −41%, domain-researcher −25%, eval-planner −26%,
  eval-auditor −25%, framework-selector −9%

Test coverage (tests/ai-evals.test.cjs — 48 tests):
- CONFIG: workflow.ai_integration_phase defaults and config-set/get
- HEALTH: W010 warning emission and addAiIntegrationPhaseKey repair
- TEMPLATE: AI-SPEC.md section completeness (10 sections)
- COMMAND: ai-integration-phase + eval-review frontmatter validity
- AGENTS: all 5 new agent files exist
- REFERENCES: ai-evals.md + ai-frameworks.md exist and are non-empty
- WORKFLOW: plan-phase nudge integration, workflow files exist + agent coverage

603/603 tests passing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: add Google ADK to framework selector and reference matrix

Google ADK (released March 2025) was missing from the framework options.
Adds Python + Java multi-agent framework optimised for Gemini / Vertex AI.

- get-shit-done/references/ai-frameworks.md: add Google ADK profile (type,
  language, model support, best for, avoid if, strengths, weaknesses, eval
  concerns); update Quick Picks, By System Type, and By Model Commitment tables
- agents/gsd-framework-selector.md: add "Google (Gemini)" to model provider
  interview question
- agents/gsd-ai-researcher.md: add Google ADK docs URL to documentation_sources

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: adapt to upstream conventions post-rebase

- Remove skills: frontmatter from all 5 new agents (upstream changed
  convention — skills: breaks Gemini CLI and must not be present)
- Add workflow.ai_integration_phase to VALID_CONFIG_KEYS whitelist in
  config.cjs (config-set blocked unknown keys)
- Add ai_integration_phase: true to CONFIG_DEFAULTS in core.cjs

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: rephrase 4b.1 line to avoid false-positive in prompt-injection scan

"contract as a Pydantic model" matched the `act as a` pattern case-insensitively.
Rephrased to "output schema using a Pydantic model".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: adapt to upstream conventions (W016, colon refs, config docs)

- Replace verify.cjs from upstream to restore W010-W015 + cmdValidateAgents,
  lost when rebase conflict was resolved with --theirs
- Add W016 (workflow.ai_integration_phase absent) inside the config try block,
  avoids collision with upstream's W010 agent-installation check
- Add addAiIntegrationPhaseKey repair case mirroring addNyquistKey pattern
- Replace /gsd: colon format with /gsd- hyphen format across all new files
  (agents, workflows, templates, verify.cjs) per stale-colon-refs guard (#1748)
- Add workflow.ai_integration_phase to planning-config.md reference table
- Add ai_integration_phase → workflow.ai_integration_phase to NAMESPACE_MAP
  in config-field-docs.test.cjs so CONFIG_DEFAULTS coverage check passes
- Update ai-evals tests to use W016 instead of W010

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: add 5 new agents to E2E Copilot install expected list

gsd-ai-researcher, gsd-domain-researcher, gsd-eval-auditor,
gsd-eval-planner, gsd-framework-selector added to the hardcoded
expected agent list in copilot-install.test.cjs (#1890).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-10 10:49:00 -04:00

8.5 KiB

AI Evaluation Reference

Reference used by gsd-eval-planner and gsd-eval-auditor. Based on "AI Evals for Everyone" course (Reganti & Badam) + industry practice.


Core Concepts

Why Evals Exist

AI systems are non-deterministic. Input X does not reliably produce output Y across runs, users, or edge cases. Evals are the continuous process of assessing whether your system's behavior meets expectations under real-world conditions — unit tests and integration tests alone are insufficient.

Model vs. Product Evaluation

  • Model evals (MMLU, HumanEval, GSM8K) — measure general capability in standardized conditions. Use as initial filter only.
  • Product evals — measure behavior inside your specific system, with your data, your users, your domain rules. This is where 80% of eval effort belongs.

The Three Components of Every Eval

  • Input — everything affecting the system: query, history, retrieved docs, system prompt, config
  • Expected — what good behavior looks like, defined through rubrics
  • Actual — what the system produced, including intermediate steps, tool calls, and reasoning traces

Three Measurement Approaches

  1. Code-based metrics — deterministic checks: JSON validation, required disclaimers, performance thresholds, classification flags. Fast, cheap, reliable. Use first.
  2. LLM judges — one model evaluates another against a rubric. Powerful for subjective qualities (tone, reasoning, escalation). Requires calibration against human judgment before trusting.
  3. Human evaluation — gold standard for nuanced judgment. Doesn't scale. Use for calibration, edge cases, periodic sampling, and high-stakes decisions.

Most effective systems combine all three.


Evaluation Dimensions

Pre-Deployment (Development Phase)

Dimension What It Measures When It Matters
Factual accuracy Correctness of claims against ground truth RAG, knowledge bases, any factual assertions
Context faithfulness Response grounded in provided context vs. fabricated RAG pipelines, document Q&A, retrieval-augmented systems
Hallucination detection Plausible but unsupported claims All generative systems, high-stakes domains
Escalation accuracy Correct identification of when human intervention needed Customer service, healthcare, financial advisory
Policy compliance Adherence to business rules, legal requirements, disclaimers Regulated industries, enterprise deployments
Tone/style appropriateness Match with brand voice, audience expectations, emotional context Customer-facing systems, content generation
Output structure validity Schema compliance, required fields, format correctness Structured extraction, API integrations, data pipelines
Task completion Whether the system accomplished the stated goal Agentic workflows, multi-step tasks
Tool use correctness Correct selection and invocation of tools Agent systems with tool calls
Safety Absence of harmful, biased, or inappropriate outputs All user-facing systems

Production Monitoring

Dimension Monitoring Approach
Safety violations Online guardrail — real-time, immediate intervention
Compliance failures Online guardrail — block or escalate before user sees output
Quality degradation trends Offline flywheel — batch analysis of sampled interactions
Emerging failure modes Signal-metric divergence — when user behavior signals diverge from metric scores, investigate manually
Cost/latency drift Code-based metrics — automated threshold alerts

The Guardrail vs. Flywheel Decision

Ask: "If this behavior goes wrong, would it be catastrophic for my business?"

  • Yes → Guardrail — run online, real-time, with immediate intervention (block, escalate, hand off). Be selective: guardrails add latency.
  • No → Flywheel — run offline as batch analysis feeding system refinements over time.

Rubric Design

Generic metrics are meaningless without context. "Helpfulness" in real estate means summarizing listings clearly. In healthcare it means knowing when not to answer.

A rubric must define:

  1. The dimension being measured
  2. What scores 1, 3, and 5 on a 5-point scale (or pass/fail criteria)
  3. Domain-specific examples of acceptable vs. unacceptable behavior

Without rubrics, LLM judges produce noise rather than signal.


Reference Dataset Guidelines

  • Start with 10-20 high-quality examples — not 200 mediocre ones
  • Cover: critical success scenarios, common user workflows, known edge cases, historical failure modes
  • Have domain experts label the examples (not just engineers)
  • Expand based on what you learn in production — don't build for hypothetical coverage

Eval Tooling Guide

Tool Type Best For Key Strength
RAGAS Python library RAG evaluation Purpose-built metrics: faithfulness, answer relevance, context precision/recall
Langfuse Platform (open-source, self-hostable) All system types Strong tracing, prompt management, good for teams wanting infrastructure control
LangSmith Platform (commercial) LangChain/LangGraph ecosystems Tightest integration with LangChain; best if already in that ecosystem
Arize Phoenix Platform (open-source + hosted) RAG + multi-agent tracing Strong RAG eval + trace visualization; open-source with hosted option
Braintrust Platform (commercial) Model-agnostic evaluation Dataset and experiment management; good for comparing across frameworks
Promptfoo CLI tool (open-source) Prompt testing, CI/CD CLI-first, excellent for CI/CD prompt regression testing

Tool Selection by System Type

System Type Recommended Tooling
RAG / Knowledge Q&A RAGAS + Arize Phoenix or Braintrust
Multi-agent systems Langfuse + Arize Phoenix
Conversational / single-model Promptfoo + Braintrust
Structured extraction Promptfoo + code-based validators
LangChain/LangGraph projects LangSmith (native integration)
Production monitoring (all types) Langfuse, Arize Phoenix, or LangSmith

Evals in the Development Lifecycle

Plan Phase (Evaluation-Aware Design)

Before writing code, define:

  1. What type of AI system is being built → determines framework and dominant eval concerns
  2. Critical failure modes (3-5 behaviors that cannot go wrong)
  3. Rubrics — explicit definitions of acceptable/unacceptable behavior per dimension
  4. Evaluation strategy — which dimensions use code metrics, LLM judges, or human review
  5. Reference dataset requirements — size, composition, labeling approach
  6. Eval tooling selection

Output: EVALS-SPEC section of AI-SPEC.md

Execute Phase (Instrument While Building)

  • Add tracing from day one (Langfuse, Arize Phoenix, or LangSmith)
  • Build reference dataset concurrently with implementation
  • Implement code-based checks first; add LLM judges only for subjective dimensions
  • Run evals in CI/CD via Promptfoo or Braintrust

Verify Phase (Pre-Deployment Validation)

  • Run full reference dataset against all metrics
  • Conduct human review of edge cases and LLM judge disagreements
  • Calibrate LLM judges against human scores (target ≥ 0.7 correlation before trusting)
  • Define and configure production guardrails
  • Establish monitoring baseline

Monitor Phase (Production Evaluation Loop)

  • Smart sampling — weight toward interactions with concerning signals (retries, unusual length, explicit escalations)
  • Online guardrails on every interaction
  • Offline flywheel on sampled batch
  • Watch for signal-metric divergence — the early warning system for evaluation gaps

Common Pitfalls

  1. Assuming benchmarks predict product success — they don't; model evals are a filter, not a verdict
  2. Engineering evals in isolation — domain experts must co-define rubrics; engineers alone miss critical nuances
  3. Building comprehensive coverage on day one — start small (10-20 examples), expand from real failure modes
  4. Trusting uncalibrated LLM judges — validate against human judgment before relying on them
  5. Measuring everything — only track metrics that drive decisions; "collect it all" produces noise
  6. Treating evaluation as one-time setup — user behavior evolves, requirements change, failure modes emerge; evaluation is continuous