Files
msd-core/get-shit-done/workflows/eval-review.md
Tom Boucher 79002a00cb chore(#518): rename npm package + bin to @opengsd/gsd-core (#519)
* chore: rename npm package + bin to @opengsd/gsd-core (functional)

- package.json: name @opengsd/get-shit-done-redux → @opengsd/gsd-core,
  bin key get-shit-done-redux → gsd-core, repository/homepage/bugs URLs
- package-lock.json: regenerated (npm install --package-lock-only)
- tests/**, scripts/**, bin/**, .github/**, agents/**, commands/**,
  get-shit-done/bin/**, get-shit-done/workflows/**:
  applied the 4-rule replacement (scoped npm ref, GitHub repo path,
  bin/clone invocations) per #505 single-source refactor

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: sweep live references to @opengsd/gsd-core

Update all live documentation (README.md + translations, docs/**,
CONTRIBUTING.md, VERSIONING.md, SECURITY.md, CONTEXT.md,
docs/CANARY.md) to reflect the renamed package and repository.

Rules applied:
- @opengsd/get-shit-done-redux → @opengsd/gsd-core (scoped npm name)
- open-gsd/get-shit-done-redux → open-gsd/gsd-core (GitHub repo)
- GSD-redux/get-shit-done-redux → open-gsd/gsd-core (stale badge org)
- bare bin/clone refs → gsd-core

CHANGELOG.md, docs/adr/**, docs/RELEASE-*.md, docs/research/**,
and .changeset/** are preserved byte-identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: add negative lookbehind to slash-command regex in bug-2954 test

The extractSlashReferences regex matched /gsd-core inside npm package
URLs (@opengsd/gsd-core), producing a false /gsd:core command reference.
Adding a negative lookbehind (?<![a-z]) excludes matches preceded by a
letter, so only standalone /gsd-<cmd> and /gsd:<cmd> tokens are found.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#518): add changeset for package rename

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#518): update package-identity expectations to the renamed coordinates

The rebase regenerated the seam to @opengsd/gsd-core (bin gsd-core, repo
open-gsd/gsd-core). The #498 seam tests assert deriveIdentity against the REAL
package.json, so their expected literals must follow the rename. The drift-lint
unit test is left as-is — its SEAM is a self-consistent fixture and its
stale-literal detection cases would shift if altered; the live-repo scan in it
already passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 17:25:02 -04:00

6.0 KiB

Retroactive audit of an implemented AI phase's evaluation coverage. Standalone command that works on any GSD-managed AI phase. Produces a scored EVAL-REVIEW.md with gap analysis and remediation plan.

Use after /gsd:execute-phase to verify that the evaluation strategy from AI-SPEC.md was actually implemented. Mirrors the pattern of /gsd:ui-review and /gsd:validate-phase.

<required_reading> @~/.claude/get-shit-done/references/ai-evals.md </required_reading>

0. Initialize

_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/get-shit-done/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/get-shit-done/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/get-shit-done/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "$HOME/.claude/get-shit-done/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="$HOME/.claude/get-shit-done/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi
INIT=$(gsd_run query init.phase-op "${PHASE_ARG}")
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi

Parse: phase_dir, phase_number, phase_name, phase_slug, padded_phase, commit_docs.

AUDITOR_MODEL=$(gsd_run query resolve-model gsd-eval-auditor 2>/dev/null | jq -r '.model' 2>/dev/null || true)

Display banner:

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 GSD ► EVAL AUDIT — PHASE {N}: {name}
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

1. Detect Input State

SUMMARY_FILES=$(ls "${PHASE_DIR}"/*-SUMMARY.md 2>/dev/null)
AI_SPEC_FILE=$(ls "${PHASE_DIR}"/*-AI-SPEC.md 2>/dev/null | head -1)
EVAL_REVIEW_FILE=$(ls "${PHASE_DIR}"/*-EVAL-REVIEW.md 2>/dev/null | head -1)

State A — AI-SPEC.md + SUMMARY.md exist: Full audit against spec State B — SUMMARY.md exists, no AI-SPEC.md: Audit against general best practices State C — No SUMMARY.md: Exit — "Phase {N} not executed. Run /gsd:execute-phase {N} first."

Text mode (workflow.text_mode: true in config or --text flag): Set TEXT_MODE=true if --text is present in $ARGUMENTS OR text_mode from init JSON is true. When TEXT_MODE is active, replace every AskUserQuestion call with a plain-text numbered list and ask the user to type their choice number. This is required for non-Claude runtimes (OpenAI Codex, Gemini CLI, etc.) where AskUserQuestion is not available. If EVAL_REVIEW_FILE non-empty: Use AskUserQuestion:

  • header: "Existing Eval Review"
  • question: "EVAL-REVIEW.md already exists for Phase {N}."
  • options:
    • "Re-audit — run fresh audit"
    • "View — display current review and exit"

If "View": display file, exit. If "Re-audit": continue.

If State B (no AI-SPEC.md): Warn:

No AI-SPEC.md found for Phase {N}.
Audit will evaluate against general AI eval best practices rather than a phase-specific plan.
Consider running /gsd:ai-integration-phase {N} before implementation next time.

Continue (non-blocking).

2. Gather Context Paths

Build file list for auditor:

  • AI-SPEC.md (if exists — the planned eval strategy)
  • All SUMMARY.md files in phase dir
  • All PLAN.md files in phase dir

3. Spawn gsd-eval-auditor

◆ Spawning eval auditor...

Build prompt:

Read ~/.claude/agents/gsd-eval-auditor.md for instructions.

<objective>
Conduct evaluation coverage audit of Phase {phase_number}: {phase_name}
{If AI-SPEC exists: "Audit against AI-SPEC.md evaluation plan."}
{If no AI-SPEC: "Audit against general AI eval best practices."}
</objective>

<files_to_read>
- {summary_paths}
- {plan_paths}
- {ai_spec_path if exists}
</files_to_read>

<input>
ai_spec_path: {ai_spec_path or "none"}
phase_dir: {phase_dir}
phase_number: {phase_number}
phase_name: {phase_name}
padded_phase: {padded_phase}
state: {A or B}
</input>

Spawn as Task with model AUDITOR_MODEL.

4. Parse Auditor Result

Read the written EVAL-REVIEW.md. Extract:

  • overall_score
  • verdict (PRODUCTION READY | NEEDS WORK | SIGNIFICANT GAPS | NOT IMPLEMENTED)
  • critical_gap_count

5. Display Summary

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 GSD ► EVAL AUDIT COMPLETE — PHASE {N}: {name}
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

◆ Score: {overall_score}/100
◆ Verdict: {verdict}
◆ Critical Gaps: {critical_gap_count}
◆ Output: {eval_review_path}

{If PRODUCTION READY:}
  Next step: /gsd:plan-phase (next phase) or deploy

{If NEEDS WORK:}
  Address critical gaps in EVAL-REVIEW.md, then re-run /gsd:eval-review {N}

{If SIGNIFICANT GAPS or NOT IMPLEMENTED:}
  Review AI-SPEC.md evaluation plan. Critical eval dimensions are not implemented.
  Do not deploy until gaps are addressed.

6. Commit

If commit_docs is true:

git add "${EVAL_REVIEW_FILE}"
git commit -m "docs({phase_slug}): add EVAL-REVIEW.md — score {overall_score}/100 ({verdict})"

<success_criteria>

  • Phase execution state detected correctly
  • AI-SPEC.md presence handled (with or without)
  • gsd-eval-auditor spawned with correct context
  • EVAL-REVIEW.md written (by auditor)
  • Score and verdict displayed to user
  • Appropriate next steps surfaced based on verdict
  • Committed if commit_docs enabled </success_criteria>