Files
msd-core/gsd-core/workflows/eval-review.md
Michel Moreira af822a8024 fix(#4776): resolve the artifact-exists prompt under --auto (#4832)
* fix(#4776): resolve the artifact-exists prompt under --auto

/gsd-ui-phase <phase> --auto stopped at 'UI-SPEC.md already exists for
Phase {N}. What would you like to do?' whenever the file was on disk —
which is most often after an earlier run wrote the contract as a draft
and ended before its checker ran, exactly the state a re-run exists to
verify. Step 4 had no --auto arm; step 9.5 in the same file has had one
since it was added, which is how the drift went unnoticed.

Step 4 now auto-selects Skip: the existing UI-SPEC is left untouched and
the run proceeds to the checker. Skip is the only non-destructive choice
— Update re-runs the researcher, which rewrites the whole contract and
drops answers a person already recorded in it, and View exits without
verifying anything.

spec-phase.md's artifact-exists arm auto-selected 'Update it', which is
the same defect with the opposite sign: an unattended run regenerating a
spec nobody is watching. Per the decision recorded on #4776 — an
unattended run reuses an existing artifact rather than regenerating it —
it now auto-selects Skip and leaves the spec unchanged.

The max-revision-iterations escalation (Force approve / Edit manually /
Abandon) is deliberately untouched and pinned by a test: accepting
blocking findings is a decision a person makes.

Closes #4776

* chore(#4776): add changeset fragment

Emitted-Drift-Ack-Growth: ui-phase.md — --auto arm added to the existing-UI-SPEC branch (#4776)
Emitted-Drift-Ack-Growth: spec-phase.md — --auto arm reworded to reuse the existing SPEC (#4776)

* fix(#4776): extend the reuse-as-is --auto fix to the 3 sibling files

The PR's original scope claim -- that ai-integration-phase.md,
eval-review.md and ui-review.md were "scoped by triage to their own
issues" -- was false; no such issues existed, and it contradicted #4776's
own most recent (2026-09-16) triage comment, which explicitly widened the
fix to require all 5 files under one recommended fix.

Applies the same reuse-as-is pattern: ai-integration-phase.md mirrors
ui-phase.md's 3-way Update/View/Skip shape (auto-selects Skip);
eval-review.md and ui-review.md have only Re-audit/View (auto-selects
View, the only non-regenerating choice). None of the three had any prior
--auto handling at all -- each has exactly one AskUserQuestion call site
total, and it is the one this fix resolves, so an --auto run through any
of them no longer stalls anywhere.

Emitted-Drift-Ack-Growth: ai-integration-phase.md — the --auto arm reusing an existing AI-SPEC is the deliverable (#4776)
Emitted-Drift-Ack-Growth: eval-review.md — the --auto arm reusing an existing EVAL-REVIEW is the deliverable (#4776)
Emitted-Drift-Ack-Growth: ui-review.md — the --auto arm reusing an existing UI-REVIEW is the deliverable (#4776)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-23 11:13:34 -04:00

8.1 KiB
Raw Blame History

Retroactive audit of an implemented AI phase's evaluation coverage. Standalone command that works on any GSD-managed AI phase. Produces a scored EVAL-REVIEW.md with gap analysis and remediation plan.

Use after /gsd:execute-phase to verify that the evaluation strategy from AI-SPEC.md was actually implemented. Mirrors the pattern of /gsd:ui-review and /gsd:validate-phase.

<required_reading> @~/.claude/gsd-core/references/ai-evals.md </required_reading>

0. Initialize

_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; _gsd_at() { for _p; do if [ -f "$_p" ]; then GSD_TOOLS="$_p"; return 0; fi; done; return 1; }; _gsd_id_ok() { case "$("$1" runtime-identity --raw 2>/dev/null || true)" in '{"packageName":"@opengsd/gsd-core"'*'}') return 0;; *) return 1;; esac; }; _gsd_homes() { _gsd_at "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; }; if _gsd_at "${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif _gsd_homes; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif unset -f gsd_run; _G="$(command -v gsd_run)"; [ -n "$_G" ] && _gsd_id_ok "$_G"; then GSD_TOOLS="$_G"; gsd_run() { "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and no identity-proving gsd_run is on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; GSD_IDENTITY_STATUS=unverified; _gsd_id_ok gsd_run && GSD_IDENTITY_STATUS=ok; export GSD_IDENTITY_STATUS; [ "$GSD_IDENTITY_STATUS" = ok ] || echo "WARNING: \"$GSD_TOOLS\" did not prove it is @opengsd/gsd-core - it is either a different package or an @opengsd/gsd-core older than the runtime-identity verb. See docs/how-to/diagnose-a-foreign-gsd-tools.md" >&2; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
RESPONSE_LANGUAGE=$(gsd_run query config-get response_language --raw --default "" 2>/dev/null || echo "")
INIT=$(gsd_run query init.phase-op "${PHASE_ARG}")
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi

If response_language is set: All user-facing output of this workflow — narration between tool calls, status updates, progress notes, findings, questions, prompts, and explanations — MUST be presented in {response_language}. Technical terms, code, file paths, and subagent prompts stay in English — only user-facing output is translated.

Parse: phase_dir, phase_number, phase_name, phase_slug, padded_phase, commit_docs.

AUDITOR_MODEL=$(gsd_run query resolve-model gsd-eval-auditor --pick model 2>/dev/null || true)
AGENT_SKILLS_AUDITOR=$(gsd_run query agent-skills gsd-eval-auditor)

Display banner:

### GSD ► EVAL AUDIT — PHASE {N}: {name}

1. Detect Input State

SUMMARY_FILES=$(ls "${PHASE_DIR}"/*-SUMMARY.md 2>/dev/null)
AI_SPEC_FILE=$(ls "${PHASE_DIR}"/*-AI-SPEC.md 2>/dev/null | head -1)
EVAL_REVIEW_FILE=$(ls "${PHASE_DIR}"/*-EVAL-REVIEW.md 2>/dev/null | head -1)

State A — AI-SPEC.md + SUMMARY.md exist: Full audit against spec State B — SUMMARY.md exists, no AI-SPEC.md: Audit against general best practices State C — No SUMMARY.md: Exit — "Phase {N} not executed. Run /gsd:execute-phase {N} first."

Text mode (workflow.text_mode: true in config or --text flag): Set TEXT_MODE=true if --text is present in $ARGUMENTS OR text_mode from init JSON is true. When TEXT_MODE is active, replace every AskUserQuestion call with a plain-text numbered list and ask the user to type their choice number. This is required for non-Claude runtimes (OpenAI Codex, Antigravity, etc.) where AskUserQuestion is not available. If EVAL_REVIEW_FILE non-empty:

If --auto: Auto-select "View" — keep the existing EVAL-REVIEW.md untouched and exit without re-auditing. Log: [auto] EVAL-REVIEW.md exists — reusing as-is. "Re-audit" is the regenerating choice; an unattended run reuses an existing artifact instead (#4776).

Otherwise: Use AskUserQuestion:

  • header: "Existing Eval Review"
  • question: "EVAL-REVIEW.md already exists for Phase {N}."
  • options:
    • "Re-audit — run fresh audit"
    • "View — display current review and exit"

If "View": display file, exit. If "Re-audit": continue.

If State B (no AI-SPEC.md): Warn:

No AI-SPEC.md found for Phase {N}.
Audit will evaluate against general AI eval best practices rather than a phase-specific plan.
Consider running /gsd:ai-integration-phase {N} before implementation next time.

Continue (non-blocking).

2. Gather Context Paths

Build file list for auditor:

  • AI-SPEC.md (if exists — the planned eval strategy)
  • All SUMMARY.md files in phase dir
  • All PLAN.md files in phase dir

3. Spawn gsd-eval-auditor

◆ Spawning eval auditor... (runs in a subagent — no output until it returns, ~1–5 min; expected, not a freeze)

Build prompt:

Read ~/.claude/agents/gsd-eval-auditor.md for instructions.

<objective>
Conduct evaluation coverage audit of Phase {phase_number}: {phase_name}
{If AI-SPEC exists: "Audit against AI-SPEC.md evaluation plan."}
{If no AI-SPEC: "Audit against general AI eval best practices."}
</objective>

<required_reading>
- {summary_paths}
- {plan_paths}
- {ai_spec_path if exists}
</required_reading>

<input>
ai_spec_path: {ai_spec_path or "none"}
phase_dir: {phase_dir}
phase_number: {phase_number}
phase_name: {phase_name}
padded_phase: {padded_phase}
state: {A or B}
</input>

${AGENT_SKILLS_AUDITOR}

Spawn as Task with model AUDITOR_MODEL.

4. Parse Auditor Result

Read the written EVAL-REVIEW.md. Extract:

  • overall_score
  • verdict (PRODUCTION READY | NEEDS WORK | SIGNIFICANT GAPS | NOT IMPLEMENTED)
  • critical_gap_count

5. Display Summary

### GSD ► EVAL AUDIT COMPLETE — PHASE {N}: {name}

◆ Score: {overall_score}/100
◆ Verdict: {verdict}
◆ Critical Gaps: {critical_gap_count}
◆ Output: {eval_review_path}

{If PRODUCTION READY:}
  Next step: /gsd:plan-phase (next phase) or deploy

{If NEEDS WORK:}
  Address critical gaps in EVAL-REVIEW.md, then re-run /gsd:eval-review {N}

{If SIGNIFICANT GAPS or NOT IMPLEMENTED:}
  Review AI-SPEC.md evaluation plan. Critical eval dimensions are not implemented.
  Do not deploy until gaps are addressed.

6. Commit

gsd_run query commit "docs({phase_slug}): add EVAL-REVIEW.md — score {overall_score}/100 ({verdict})" --files "${EVAL_REVIEW_FILE}"

<success_criteria>

  • Phase execution state detected correctly
  • AI-SPEC.md presence handled (with or without)
  • gsd-eval-auditor spawned with correct context
  • EVAL-REVIEW.md written (by auditor)
  • Score and verdict displayed to user
  • Appropriate next steps surfaced based on verdict
  • Committed if commit_docs enabled </success_criteria>