Files
msd-core/gsd-core/references/ai-evals.md
Tom Boucher 463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00

8.5 KiB

AI Evaluation Reference

Reference used by gsd-eval-planner and gsd-eval-auditor. Based on "AI Evals for Everyone" course (Reganti & Badam) + industry practice.


Core Concepts

Why Evals Exist

AI systems are non-deterministic. Input X does not reliably produce output Y across runs, users, or edge cases. Evals are the continuous process of assessing whether your system's behavior meets expectations under real-world conditions — unit tests and integration tests alone are insufficient.

Model vs. Product Evaluation

  • Model evals (MMLU, HumanEval, GSM8K) — measure general capability in standardized conditions. Use as initial filter only.
  • Product evals — measure behavior inside your specific system, with your data, your users, your domain rules. This is where 80% of eval effort belongs.

The Three Components of Every Eval

  • Input — everything affecting the system: query, history, retrieved docs, system prompt, config
  • Expected — what good behavior looks like, defined through rubrics
  • Actual — what the system produced, including intermediate steps, tool calls, and reasoning traces

Three Measurement Approaches

  1. Code-based metrics — deterministic checks: JSON validation, required disclaimers, performance thresholds, classification flags. Fast, cheap, reliable. Use first.
  2. LLM judges — one model evaluates another against a rubric. Powerful for subjective qualities (tone, reasoning, escalation). Requires calibration against human judgment before trusting.
  3. Human evaluation — gold standard for nuanced judgment. Doesn't scale. Use for calibration, edge cases, periodic sampling, and high-stakes decisions.

Most effective systems combine all three.


Evaluation Dimensions

Pre-Deployment (Development Phase)

Dimension What It Measures When It Matters
Factual accuracy Correctness of claims against ground truth RAG, knowledge bases, any factual assertions
Context faithfulness Response grounded in provided context vs. fabricated RAG pipelines, document Q&A, retrieval-augmented systems
Hallucination detection Plausible but unsupported claims All generative systems, high-stakes domains
Escalation accuracy Correct identification of when human intervention needed Customer service, healthcare, financial advisory
Policy compliance Adherence to business rules, legal requirements, disclaimers Regulated industries, enterprise deployments
Tone/style appropriateness Match with brand voice, audience expectations, emotional context Customer-facing systems, content generation
Output structure validity Schema compliance, required fields, format correctness Structured extraction, API integrations, data pipelines
Task completion Whether the system accomplished the stated goal Agentic workflows, multi-step tasks
Tool use correctness Correct selection and invocation of tools Agent systems with tool calls
Safety Absence of harmful, biased, or inappropriate outputs All user-facing systems

Production Monitoring

Dimension Monitoring Approach
Safety violations Online guardrail — real-time, immediate intervention
Compliance failures Online guardrail — block or escalate before user sees output
Quality degradation trends Offline flywheel — batch analysis of sampled interactions
Emerging failure modes Signal-metric divergence — when user behavior signals diverge from metric scores, investigate manually
Cost/latency drift Code-based metrics — automated threshold alerts

The Guardrail vs. Flywheel Decision

Ask: "If this behavior goes wrong, would it be catastrophic for my business?"

  • Yes → Guardrail — run online, real-time, with immediate intervention (block, escalate, hand off). Be selective: guardrails add latency.
  • No → Flywheel — run offline as batch analysis feeding system refinements over time.

Rubric Design

Generic metrics are meaningless without context. "Helpfulness" in real estate means summarizing listings clearly. In healthcare it means knowing when not to answer.

A rubric must define:

  1. The dimension being measured
  2. What scores 1, 3, and 5 on a 5-point scale (or pass/fail criteria)
  3. Domain-specific examples of acceptable vs. unacceptable behavior

Without rubrics, LLM judges produce noise rather than signal.


Reference Dataset Guidelines

  • Start with 10-20 high-quality examples — not 200 mediocre ones
  • Cover: critical success scenarios, common user workflows, known edge cases, historical failure modes
  • Have domain experts label the examples (not just engineers)
  • Expand based on what you learn in production — don't build for hypothetical coverage

Eval Tooling Guide

Tool Type Best For Key Strength
RAGAS Python library RAG evaluation Purpose-built metrics: faithfulness, answer relevance, context precision/recall
Langfuse Platform (open-source, self-hostable) All system types Strong tracing, prompt management, good for teams wanting infrastructure control
LangSmith Platform (commercial) LangChain/LangGraph ecosystems Tightest integration with LangChain; best if already in that ecosystem
Arize Phoenix Platform (open-source + hosted) RAG + multi-agent tracing Strong RAG eval + trace visualization; open-source with hosted option
Braintrust Platform (commercial) Model-agnostic evaluation Dataset and experiment management; good for comparing across frameworks
Promptfoo CLI tool (open-source) Prompt testing, CI/CD CLI-first, excellent for CI/CD prompt regression testing

Tool Selection by System Type

System Type Recommended Tooling
RAG / Knowledge Q&A RAGAS + Arize Phoenix or Braintrust
Multi-agent systems Langfuse + Arize Phoenix
Conversational / single-model Promptfoo + Braintrust
Structured extraction Promptfoo + code-based validators
LangChain/LangGraph projects LangSmith (native integration)
Production monitoring (all types) Langfuse, Arize Phoenix, or LangSmith

Evals in the Development Lifecycle

Plan Phase (Evaluation-Aware Design)

Before writing code, define:

  1. What type of AI system is being built → determines framework and dominant eval concerns
  2. Critical failure modes (3-5 behaviors that cannot go wrong)
  3. Rubrics — explicit definitions of acceptable/unacceptable behavior per dimension
  4. Evaluation strategy — which dimensions use code metrics, LLM judges, or human review
  5. Reference dataset requirements — size, composition, labeling approach
  6. Eval tooling selection

Output: EVALS-SPEC section of AI-SPEC.md

Execute Phase (Instrument While Building)

  • Add tracing from day one (Langfuse, Arize Phoenix, or LangSmith)
  • Build reference dataset concurrently with implementation
  • Implement code-based checks first; add LLM judges only for subjective dimensions
  • Run evals in CI/CD via Promptfoo or Braintrust

Verify Phase (Pre-Deployment Validation)

  • Run full reference dataset against all metrics
  • Conduct human review of edge cases and LLM judge disagreements
  • Calibrate LLM judges against human scores (target ≥ 0.7 correlation before trusting)
  • Define and configure production guardrails
  • Establish monitoring baseline

Monitor Phase (Production Evaluation Loop)

  • Smart sampling — weight toward interactions with concerning signals (retries, unusual length, explicit escalations)
  • Online guardrails on every interaction
  • Offline flywheel on sampled batch
  • Watch for signal-metric divergence — the early warning system for evaluation gaps

Common Pitfalls

  1. Assuming benchmarks predict product success — they don't; model evals are a filter, not a verdict
  2. Engineering evals in isolation — domain experts must co-define rubrics; engineers alone miss critical nuances
  3. Building comprehensive coverage on day one — start small (10-20 examples), expand from real failure modes
  4. Trusting uncalibrated LLM judges — validate against human judgment before relying on them
  5. Measuring everything — only track metrics that drive decisions; "collect it all" produces noise
  6. Treating evaluation as one-time setup — user behavior evolves, requirements change, failure modes emerge; evaluation is continuous