* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
157 lines
8.5 KiB
Markdown
157 lines
8.5 KiB
Markdown
# AI Evaluation Reference
|
|
|
|
> Reference used by `gsd-eval-planner` and `gsd-eval-auditor`.
|
|
> Based on "AI Evals for Everyone" course (Reganti & Badam) + industry practice.
|
|
|
|
---
|
|
|
|
## Core Concepts
|
|
|
|
### Why Evals Exist
|
|
AI systems are non-deterministic. Input X does not reliably produce output Y across runs, users, or edge cases. Evals are the continuous process of assessing whether your system's behavior meets expectations under real-world conditions — unit tests and integration tests alone are insufficient.
|
|
|
|
### Model vs. Product Evaluation
|
|
- **Model evals** (MMLU, HumanEval, GSM8K) — measure general capability in standardized conditions. Use as initial filter only.
|
|
- **Product evals** — measure behavior inside your specific system, with your data, your users, your domain rules. This is where 80% of eval effort belongs.
|
|
|
|
### The Three Components of Every Eval
|
|
- **Input** — everything affecting the system: query, history, retrieved docs, system prompt, config
|
|
- **Expected** — what good behavior looks like, defined through rubrics
|
|
- **Actual** — what the system produced, including intermediate steps, tool calls, and reasoning traces
|
|
|
|
### Three Measurement Approaches
|
|
1. **Code-based metrics** — deterministic checks: JSON validation, required disclaimers, performance thresholds, classification flags. Fast, cheap, reliable. Use first.
|
|
2. **LLM judges** — one model evaluates another against a rubric. Powerful for subjective qualities (tone, reasoning, escalation). Requires calibration against human judgment before trusting.
|
|
3. **Human evaluation** — gold standard for nuanced judgment. Doesn't scale. Use for calibration, edge cases, periodic sampling, and high-stakes decisions.
|
|
|
|
Most effective systems combine all three.
|
|
|
|
---
|
|
|
|
## Evaluation Dimensions
|
|
|
|
### Pre-Deployment (Development Phase)
|
|
|
|
| Dimension | What It Measures | When It Matters |
|
|
|-----------|-----------------|-----------------|
|
|
| **Factual accuracy** | Correctness of claims against ground truth | RAG, knowledge bases, any factual assertions |
|
|
| **Context faithfulness** | Response grounded in provided context vs. fabricated | RAG pipelines, document Q&A, retrieval-augmented systems |
|
|
| **Hallucination detection** | Plausible but unsupported claims | All generative systems, high-stakes domains |
|
|
| **Escalation accuracy** | Correct identification of when human intervention needed | Customer service, healthcare, financial advisory |
|
|
| **Policy compliance** | Adherence to business rules, legal requirements, disclaimers | Regulated industries, enterprise deployments |
|
|
| **Tone/style appropriateness** | Match with brand voice, audience expectations, emotional context | Customer-facing systems, content generation |
|
|
| **Output structure validity** | Schema compliance, required fields, format correctness | Structured extraction, API integrations, data pipelines |
|
|
| **Task completion** | Whether the system accomplished the stated goal | Agentic workflows, multi-step tasks |
|
|
| **Tool use correctness** | Correct selection and invocation of tools | Agent systems with tool calls |
|
|
| **Safety** | Absence of harmful, biased, or inappropriate outputs | All user-facing systems |
|
|
|
|
### Production Monitoring
|
|
|
|
| Dimension | Monitoring Approach |
|
|
|-----------|---------------------|
|
|
| **Safety violations** | Online guardrail — real-time, immediate intervention |
|
|
| **Compliance failures** | Online guardrail — block or escalate before user sees output |
|
|
| **Quality degradation trends** | Offline flywheel — batch analysis of sampled interactions |
|
|
| **Emerging failure modes** | Signal-metric divergence — when user behavior signals diverge from metric scores, investigate manually |
|
|
| **Cost/latency drift** | Code-based metrics — automated threshold alerts |
|
|
|
|
---
|
|
|
|
## The Guardrail vs. Flywheel Decision
|
|
|
|
Ask: "If this behavior goes wrong, would it be catastrophic for my business?"
|
|
|
|
- **Yes → Guardrail** — run online, real-time, with immediate intervention (block, escalate, hand off). Be selective: guardrails add latency.
|
|
- **No → Flywheel** — run offline as batch analysis feeding system refinements over time.
|
|
|
|
---
|
|
|
|
## Rubric Design
|
|
|
|
Generic metrics are meaningless without context. "Helpfulness" in real estate means summarizing listings clearly. In healthcare it means knowing when *not* to answer.
|
|
|
|
A rubric must define:
|
|
1. The dimension being measured
|
|
2. What scores 1, 3, and 5 on a 5-point scale (or pass/fail criteria)
|
|
3. Domain-specific examples of acceptable vs. unacceptable behavior
|
|
|
|
Without rubrics, LLM judges produce noise rather than signal.
|
|
|
|
---
|
|
|
|
## Reference Dataset Guidelines
|
|
|
|
- Start with **10-20 high-quality examples** — not 200 mediocre ones
|
|
- Cover: critical success scenarios, common user workflows, known edge cases, historical failure modes
|
|
- Have domain experts label the examples (not just engineers)
|
|
- Expand based on what you learn in production — don't build for hypothetical coverage
|
|
|
|
---
|
|
|
|
## Eval Tooling Guide
|
|
|
|
| Tool | Type | Best For | Key Strength |
|
|
|------|------|----------|-------------|
|
|
| **RAGAS** | Python library | RAG evaluation | Purpose-built metrics: faithfulness, answer relevance, context precision/recall |
|
|
| **Langfuse** | Platform (open-source, self-hostable) | All system types | Strong tracing, prompt management, good for teams wanting infrastructure control |
|
|
| **LangSmith** | Platform (commercial) | LangChain/LangGraph ecosystems | Tightest integration with LangChain; best if already in that ecosystem |
|
|
| **Arize Phoenix** | Platform (open-source + hosted) | RAG + multi-agent tracing | Strong RAG eval + trace visualization; open-source with hosted option |
|
|
| **Braintrust** | Platform (commercial) | Model-agnostic evaluation | Dataset and experiment management; good for comparing across frameworks |
|
|
| **Promptfoo** | CLI tool (open-source) | Prompt testing, CI/CD | CLI-first, excellent for CI/CD prompt regression testing |
|
|
|
|
### Tool Selection by System Type
|
|
|
|
| System Type | Recommended Tooling |
|
|
|-------------|---------------------|
|
|
| RAG / Knowledge Q&A | RAGAS + Arize Phoenix or Braintrust |
|
|
| Multi-agent systems | Langfuse + Arize Phoenix |
|
|
| Conversational / single-model | Promptfoo + Braintrust |
|
|
| Structured extraction | Promptfoo + code-based validators |
|
|
| LangChain/LangGraph projects | LangSmith (native integration) |
|
|
| Production monitoring (all types) | Langfuse, Arize Phoenix, or LangSmith |
|
|
|
|
---
|
|
|
|
## Evals in the Development Lifecycle
|
|
|
|
### Plan Phase (Evaluation-Aware Design)
|
|
Before writing code, define:
|
|
1. What type of AI system is being built → determines framework and dominant eval concerns
|
|
2. Critical failure modes (3-5 behaviors that cannot go wrong)
|
|
3. Rubrics — explicit definitions of acceptable/unacceptable behavior per dimension
|
|
4. Evaluation strategy — which dimensions use code metrics, LLM judges, or human review
|
|
5. Reference dataset requirements — size, composition, labeling approach
|
|
6. Eval tooling selection
|
|
|
|
Output: EVALS-SPEC section of AI-SPEC.md
|
|
|
|
### Execute Phase (Instrument While Building)
|
|
- Add tracing from day one (Langfuse, Arize Phoenix, or LangSmith)
|
|
- Build reference dataset concurrently with implementation
|
|
- Implement code-based checks first; add LLM judges only for subjective dimensions
|
|
- Run evals in CI/CD via Promptfoo or Braintrust
|
|
|
|
### Verify Phase (Pre-Deployment Validation)
|
|
- Run full reference dataset against all metrics
|
|
- Conduct human review of edge cases and LLM judge disagreements
|
|
- Calibrate LLM judges against human scores (target ≥ 0.7 correlation before trusting)
|
|
- Define and configure production guardrails
|
|
- Establish monitoring baseline
|
|
|
|
### Monitor Phase (Production Evaluation Loop)
|
|
- Smart sampling — weight toward interactions with concerning signals (retries, unusual length, explicit escalations)
|
|
- Online guardrails on every interaction
|
|
- Offline flywheel on sampled batch
|
|
- Watch for signal-metric divergence — the early warning system for evaluation gaps
|
|
|
|
---
|
|
|
|
## Common Pitfalls
|
|
|
|
1. **Assuming benchmarks predict product success** — they don't; model evals are a filter, not a verdict
|
|
2. **Engineering evals in isolation** — domain experts must co-define rubrics; engineers alone miss critical nuances
|
|
3. **Building comprehensive coverage on day one** — start small (10-20 examples), expand from real failure modes
|
|
4. **Trusting uncalibrated LLM judges** — validate against human judgment before relying on them
|
|
5. **Measuring everything** — only track metrics that drive decisions; "collect it all" produces noise
|
|
6. **Treating evaluation as one-time setup** — user behavior evolves, requirements change, failure modes emerge; evaluation is continuous
|