Retroactive audit of an implemented AI phase's evaluation coverage. Standalone command that works on any GSD-managed AI phase. Produces a scored EVAL-REVIEW.md with gap analysis and remediation plan. Use after /gsd:execute-phase to verify that the evaluation strategy from AI-SPEC.md was actually implemented. Mirrors the pattern of /gsd:ui-review and /gsd:validate-phase. @~/.claude/gsd-core/references/ai-evals.md ## 0. Initialize ```bash _GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; _gsd_at() { for _p; do if [ -f "$_p" ]; then GSD_TOOLS="$_p"; return 0; fi; done; return 1; }; if _gsd_at "${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif unset -f gsd_run; _G="$(command -v gsd_run)"; then GSD_TOOLS="$_G"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif _gsd_at "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd_run is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; GSD_IDENTITY_STATUS=unverified; case "$(gsd_run runtime-identity --raw 2>/dev/null || true)" in '{"packageName":"@opengsd/gsd-core"'*'}') GSD_IDENTITY_STATUS=ok;; esac; export GSD_IDENTITY_STATUS; [ "$GSD_IDENTITY_STATUS" = ok ] || echo "WARNING: \"$GSD_TOOLS\" did not prove it is @opengsd/gsd-core - it is either a different package or an @opengsd/gsd-core older than the runtime-identity verb. See docs/how-to/diagnose-a-foreign-gsd-tools.md" >&2; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi RESPONSE_LANGUAGE=$(gsd_run query config-get response_language --default "" 2>/dev/null || echo "") INIT=$(gsd_run query init.phase-op "${PHASE_ARG}") if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi ``` **If `response_language` is set:** All user-facing questions, prompts, and explanations in this workflow MUST be presented in `{response_language}`. Technical terms, code, file paths, and subagent prompts stay in English — only user-facing output is translated. Parse: `phase_dir`, `phase_number`, `phase_name`, `phase_slug`, `padded_phase`, `commit_docs`. ```bash AUDITOR_MODEL=$(gsd_run query resolve-model gsd-eval-auditor --pick model 2>/dev/null || true) AGENT_SKILLS_AUDITOR=$(gsd_run query agent-skills gsd-eval-auditor) ``` Display banner: ``` ### GSD ► EVAL AUDIT — PHASE {N}: {name} ``` ## 1. Detect Input State ```bash SUMMARY_FILES=$(ls "${PHASE_DIR}"/*-SUMMARY.md 2>/dev/null) AI_SPEC_FILE=$(ls "${PHASE_DIR}"/*-AI-SPEC.md 2>/dev/null | head -1) EVAL_REVIEW_FILE=$(ls "${PHASE_DIR}"/*-EVAL-REVIEW.md 2>/dev/null | head -1) ``` **State A** — AI-SPEC.md + SUMMARY.md exist: Full audit against spec **State B** — SUMMARY.md exists, no AI-SPEC.md: Audit against general best practices **State C** — No SUMMARY.md: Exit — "Phase {N} not executed. Run /gsd:execute-phase {N} first." **Text mode (`workflow.text_mode: true` in config or `--text` flag):** Set `TEXT_MODE=true` if `--text` is present in `$ARGUMENTS` OR `text_mode` from init JSON is `true`. When TEXT_MODE is active, replace every `AskUserQuestion` call with a plain-text numbered list and ask the user to type their choice number. This is required for non-Claude runtimes (OpenAI Codex, Gemini CLI, etc.) where `AskUserQuestion` is not available. **If `EVAL_REVIEW_FILE` non-empty:** Use AskUserQuestion: - header: "Existing Eval Review" - question: "EVAL-REVIEW.md already exists for Phase {N}." - options: - "Re-audit — run fresh audit" - "View — display current review and exit" If "View": display file, exit. If "Re-audit": continue. **If State B (no AI-SPEC.md):** Warn: ``` No AI-SPEC.md found for Phase {N}. Audit will evaluate against general AI eval best practices rather than a phase-specific plan. Consider running /gsd:ai-integration-phase {N} before implementation next time. ``` Continue (non-blocking). ## 2. Gather Context Paths Build file list for auditor: - AI-SPEC.md (if exists — the planned eval strategy) - All SUMMARY.md files in phase dir - All PLAN.md files in phase dir ## 3. Spawn gsd-eval-auditor ``` ◆ Spawning eval auditor... (runs in a subagent — no output until it returns, ~1–5 min; expected, not a freeze) ``` Build prompt: ```markdown Read ~/.claude/agents/gsd-eval-auditor.md for instructions. Conduct evaluation coverage audit of Phase {phase_number}: {phase_name} {If AI-SPEC exists: "Audit against AI-SPEC.md evaluation plan."} {If no AI-SPEC: "Audit against general AI eval best practices."} - {summary_paths} - {plan_paths} - {ai_spec_path if exists} ai_spec_path: {ai_spec_path or "none"} phase_dir: {phase_dir} phase_number: {phase_number} phase_name: {phase_name} padded_phase: {padded_phase} state: {A or B} ${AGENT_SKILLS_AUDITOR} ``` Spawn as Task with model `AUDITOR_MODEL`. ## 4. Parse Auditor Result Read the written EVAL-REVIEW.md. Extract: - `overall_score` - `verdict` (PRODUCTION READY | NEEDS WORK | SIGNIFICANT GAPS | NOT IMPLEMENTED) - `critical_gap_count` ## 5. Display Summary ``` ### GSD ► EVAL AUDIT COMPLETE — PHASE {N}: {name} ◆ Score: {overall_score}/100 ◆ Verdict: {verdict} ◆ Critical Gaps: {critical_gap_count} ◆ Output: {eval_review_path} {If PRODUCTION READY:} Next step: /gsd:plan-phase (next phase) or deploy {If NEEDS WORK:} Address critical gaps in EVAL-REVIEW.md, then re-run /gsd:eval-review {N} {If SIGNIFICANT GAPS or NOT IMPLEMENTED:} Review AI-SPEC.md evaluation plan. Critical eval dimensions are not implemented. Do not deploy until gaps are addressed. ``` ## 6. Commit ```bash gsd_run query commit "docs({phase_slug}): add EVAL-REVIEW.md — score {overall_score}/100 ({verdict})" --files "${EVAL_REVIEW_FILE}" ``` - [ ] Phase execution state detected correctly - [ ] AI-SPEC.md presence handled (with or without) - [ ] gsd-eval-auditor spawned with correct context - [ ] EVAL-REVIEW.md written (by auditor) - [ ] Score and verdict displayed to user - [ ] Appropriate next steps surfaced based on verdict - [ ] Committed if commit_docs enabled