* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) BSD/macOS mktemp only substitutes the XXXXXX template when it is the final path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md` return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent workflow runs collide on the same temp manifest/body file — one run can overwrite or consume another's. Reproduced on macOS: the second call to the suffixed template fails `mkstemp: File exists`. Fix: use a suffixless `XXXXXX` template (so it IS the final component), then rename to add the intended extension — portable across BSD + GNU userlands, no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at every site. Affected workflow temp files: - execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest) - quick.md: gsd-quick-worktree-*.json - spec-phase.md: edge-probe-reqs-*.json - ship.md: gsd-pr-body-*.md - profile-user.md: gsd-profile-answers-*.json, gsd-profile-analysis-*.json The execute-phase.md edit uses a compact intermediate var + trailing comment to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the workflow size baseline accordingly. Validated on macOS: 20 concurrent calls yield 20 unique randomized paths. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): add changeset fragment (Fixed) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp template whose XXXXXX run is followed by a filename suffix (the BSD/macOS non-randomizing form). Fails on the six pre-fix instances and passes on the fix, and locks the copy-paste-prone idiom out of future workflows. Mirrors the bug-637 hardcoded-$HOME workflow guard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): rename regression test to fix- prefix (regression-test-names lint) New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names ratchet; use the fix- prefix (matches the fix-1445 precedent). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint) lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to carry a #NNN reference (don't allowlist). Add (#1520) to the source-text exemption. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1520): abort touched mktemp chains on failure (|| exit 1) Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's suggested failure guard. If mktemp fails, $VAR is empty and the subsequent mv/write lands on an unintended relative path. Add `|| exit 1` to all six touched chains so a mktemp failure aborts the snippet. Regenerated the workflow size baseline for the slightly longer lines. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): rebase onto next — regen size baseline + describe rename Resolve the workflow-size-baseline.json conflict from next advancing by regenerating from the current workflow sizes. Also rename the test describe from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit) The two profile-user.md temp sites this PR already rewrites kept a hardcoded /tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for consistency and macOS-correctness (some sandboxes have no writable /tmp). Regenerated the workflow size baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): regen size baseline after rebase onto next --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
31 KiB
This workflow handles "what" and "why" — discuss-phase handles "how".
<ambiguity_model> Score each dimension 0.0 (completely unclear) to 1.0 (crystal clear):
| Dimension | Weight | Minimum | What it measures |
|---|---|---|---|
| Goal Clarity | 35% | 0.75 | Is the outcome specific and measurable? |
| Boundary Clarity | 25% | 0.70 | What's in scope vs out of scope? |
| Constraint Clarity | 20% | 0.65 | Performance, compatibility, data requirements? |
| Acceptance Criteria | 20% | 0.70 | How do we know it's done? |
Ambiguity score = 1.0 − (0.35×goal + 0.25×boundary + 0.20×constraint + 0.20×acceptance)
Gate: ambiguity ≤ 0.20 AND all dimensions ≥ their minimums → ready to write SPEC.md.
A score of 0.20 means 80% weighted clarity — enough precision that the planner won't silently make wrong assumptions. </ambiguity_model>
<interview_perspectives> Rotate through these perspectives — each naturally surfaces different blindspots:
Researcher (rounds 1–2): Ground the discussion in current reality.
- "What exists in the codebase today related to this phase?"
- "What's the delta between today and the target state?"
- "What triggers this work — what's broken or missing?"
Simplifier (round 2): Surface minimum viable scope.
- "What's the simplest version that solves the core problem?"
- "If you had to cut 50%, what's the irreducible core?"
- "What would make this phase a success even without the nice-to-haves?"
Boundary Keeper (round 3): Lock the perimeter.
- "What explicitly will NOT be done in this phase?"
- "What adjacent problems is it tempting to solve but shouldn't?"
- "What does 'done' look like — what's the final deliverable?"
Failure Analyst (round 4): Find the edge cases that invalidate requirements.
- "What's the worst thing that could go wrong if we get the requirements wrong?"
- "What does a broken version of this look like?"
- "What would cause a verifier to reject the output?"
Seed Closer (rounds 5–6): Lock remaining undecided territory.
- "We have [dimension] at [score] — what would make it completely clear?"
- "The remaining ambiguity is in [area] — can we make a decision now?"
- "Is there anything you'd regret not specifying before planning starts?" </interview_perspectives>
Step 1: Initialize
_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
INIT=$(gsd_run init phase-op "${PHASE}")
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi
Parse JSON for: phase_found, phase_dir, phase_number, phase_name, phase_slug, padded_phase, state_path, requirements_path, roadmap_path, planning_path, response_language, commit_docs.
If response_language is set: All user-facing text in this workflow MUST be in {response_language}. Technical terms, code, and file paths stay in English.
If phase_found is false:
Phase [X] not found in roadmap.
Use /gsd:progress to see available phases.
Exit.
Check for existing SPEC.md:
ls ${phase_dir}/*-SPEC.md 2>/dev/null | grep -v AI-SPEC | head -1 || true
If SPEC.md already exists:
If --auto: Auto-select "Update it". Log: [auto] SPEC.md exists — updating.
Otherwise: Use AskUserQuestion:
- header: "Spec"
- question: "Phase [X] already has a SPEC.md. What do you want to do?"
- options:
- "Update it" — Revise and re-score
- "View it" — Show current spec
- "Skip" — Exit (use existing spec as-is)
If "View": Display SPEC.md, then offer Update/Skip. If "Skip": Exit with message: "Existing SPEC.md unchanged. Run /gsd:discuss-phase [X] to continue." If "Update": Load existing SPEC.md, continue to Step 3.
Step 2: Scout Codebase
Read these files before any questions:
{requirements_path}— Project requirements{state_path}— Decisions already made, current phase, blockers- ROADMAP.md phase entry — Phase description, goals, canonical refs
Grep the codebase for code/files relevant to this phase goal. Look for:
- Existing implementations of similar functionality
- Integration points where new code will connect
- Test coverage gaps relevant to the phase
- Prior phase artifacts (SUMMARY.md, VERIFICATION.md) that inform current state
Synthesize current state — the grounded baseline for the interview:
- What exists today related to this phase
- The gap between current state and the phase goal
- The primary deliverable: what file/behavior/capability does NOT exist yet?
Confirm your current state synthesis internally. Do not present it to the user yet — you'll use it to ask precise, grounded questions.
Step 3: First Ambiguity Assessment
Before questioning begins, score the phase's current ambiguity based only on what ROADMAP.md and REQUIREMENTS.md say:
Goal Clarity: [score 0.0–1.0]
Boundary Clarity: [score 0.0–1.0]
Constraint Clarity: [score 0.0–1.0]
Acceptance Criteria:[score 0.0–1.0]
Ambiguity: [score] ([calculate])
If --auto and initial ambiguity already ≤ 0.20 with all minimums met: Skip interview — derive SPEC.md directly from roadmap + requirements. Log: [auto] Phase requirements are already sufficiently clear — generating SPEC.md from existing context. Jump to Step 6.
Otherwise: Continue to Step 4.
Step 4: Socratic Interview Loop
Max 6 rounds. Each round: 2–3 questions max. End round after user responds.
Round selection by perspective:
- Round 1: Researcher
- Round 2: Researcher + Simplifier
- Round 3: Boundary Keeper
- Round 4: Failure Analyst
- Rounds 5–6: Seed Closer (focus on lowest-scoring dimensions)
After each round:
- Update all 4 dimension scores from the user's answers
- Calculate new ambiguity score
- Display the updated scoring:
After round [N]:
Goal Clarity: [score] (min 0.75) [✓ or ↑ needed]
Boundary Clarity: [score] (min 0.70) [✓ or ↑ needed]
Constraint Clarity: [score] (min 0.65) [✓ or ↑ needed]
Acceptance Criteria:[score] (min 0.70) [✓ or ↑ needed]
Ambiguity: [score] (gate: ≤ 0.20)
Gate check after each round:
If gate passes (ambiguity ≤ 0.20 AND all minimums met):
If --auto: Jump to Step 6.
Otherwise: AskUserQuestion:
- header: "Spec Gate Passed"
- question: "Ambiguity is [score] — requirements are clear enough to write SPEC.md. Proceed?"
- options:
- "Yes — write SPEC.md" → Jump to Step 6
- "One more round" → Continue interview
- "Done talking — write it" → Jump to Step 6
If max rounds reached (6) and gate not passed:
If --auto: Write SPEC.md anyway — flag unresolved dimensions. Log: [auto] Max rounds reached. Writing SPEC.md with [N] dimensions below minimum. Planner will need to treat these as assumptions.
Otherwise: AskUserQuestion:
- header: "Max Rounds"
- question: "After 6 rounds, ambiguity is [score]. [List dimensions still below minimum.] What would you like to do?"
- options:
- "Write SPEC.md anyway — flag gaps" → Write SPEC.md, mark unresolved dimensions in Ambiguity Report
- "Keep talking" → Continue (no round limit from here)
- "Abandon" → Exit without writing
If --auto mode throughout: Replace all AskUserQuestion calls above with Claude's recommended choice. Log decisions inline. Apply the same logic as --auto in discuss-phase.
Text mode (workflow.text_mode: true or --text flag): Use plain-text numbered lists instead of AskUserQuestion TUI menus.
Step 5: (covered inline — ambiguity scoring is per-round)
Step 5.5: Edge-Completeness Probe
Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague ones). Reference: @~/.claude/gsd-core/references/edge-probe.md.
Runtime coverage compute — resolve and invoke edge-probe.cjs:
# Resolve the compiled edge-probe.cjs against the GSD install dir via RUNTIME_DIR (#448)
# — NOT the consuming project's git root — falling back to git toplevel / $HOME/.claude.
# Mirrors the ui-safety-gate.cjs resolution idiom at autonomous.md:290 / plan-phase.md:631.
_GSD_RT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"
EDGE_PROBE_JS=$(for _c in \
"$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
"$_GSD_RT/bin/lib/edge-probe.cjs" \
"$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
"$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
"$HOME/.claude/bin/lib/edge-probe.cjs"; do
[ -f "$_c" ] && { echo "$_c"; break; }
done)
# Graceful degradation — never silent skip (RR-04). Build ONLY when $_GSD_RT is a verified
# GSD source checkout (has tsconfig.build.json + src/edge-probe.cts), and pin npm to it with
# --prefix so we never trigger the CONSUMING project's own build:lib (its cwd package scripts:
# codegen/migrations/writes) during a spec workflow. Real installs ship the compiled .cjs via
# prepublishOnly, so this build path only matters in a GSD dev checkout (review High).
if [ -z "$EDGE_PROBE_JS" ]; then
if [ -f "$_GSD_RT/tsconfig.build.json" ] && [ -f "$_GSD_RT/src/edge-probe.cts" ]; then
npm --prefix "$_GSD_RT" run build:lib 2>/dev/null || true
EDGE_PROBE_JS=$(for _c in \
"$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
"$_GSD_RT/bin/lib/edge-probe.cjs" \
"$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
"$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
"$HOME/.claude/bin/lib/edge-probe.cjs"; do
[ -f "$_c" ] && { echo "$_c"; break; }
done)
fi
if [ -z "$EDGE_PROBE_JS" ]; then
echo "ERROR: edge-probe.cjs not found — reinstall GSD or run \`npm run build:lib\` in your GSD checkout." >&2
exit 1
fi
fi
# Write the Requirements gathered in THIS spec session to a temp JSON, then invoke the
# canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object
# per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes
# the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04).
# BSD/macOS mktemp only randomizes XXXXXX when it is the final path component, so make a
# suffixless temp then append the extension — portable across BSD + GNU (#1520).
REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX") && mv "$REQS_JSON" "${REQS_JSON}.json" && REQS_JSON="${REQS_JSON}.json" || exit 1
cat > "$REQS_JSON" <<'JSON'
[
{ "id": "R1", "text": "<replace: requirement text from the SPEC>" }
]
JSON
# Guard — never invoke on an empty/invalid array, OR one still holding the heredoc
# `<replace: …>` placeholder (a forgotten substitution would otherwise yield a
# meaningful-looking but bogus coverage report). Fail loud, not silent no-op.
if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===0)process.exit(1);if(a.some(r=>typeof r.text!=="string"||!r.text.trim()||r.text.includes("<replace:")))process.exit(1)' "$REQS_JSON" 2>/dev/null; then
echo "ERROR: edge-probe requirements JSON is empty/invalid or still holds the <replace: …> placeholder — populate \$REQS_JSON from the SPEC Requirements before Step 5.5 runs." >&2
exit 1
fi
# Invoke the compiled engine and CAPTURE its report — it computes which categories apply per
# requirement. The covered/backstop/dismissed/unresolved rows in $COVERAGE drive the
# resolution loop below (canonical taxonomy compute, NOT LLM re-derivation from prose).
# The engine FAILS CLOSED (exit 2) on an invalid authored shape or bad input — so the capture
# MUST be exit-checked. A bare `COVERAGE=$(node …)` swallows that exit code, leaves $COVERAGE
# empty, and lets the workflow fall through to prose re-derivation: fail-OPEN at the boundary
# the engine validation exists to protect. Make the run fatal, then validate the captured
# report is well-formed JSON before the resolution loop consumes it.
if ! COVERAGE=$(node "$EDGE_PROBE_JS" "$REQS_JSON"); then
rm -f "$REQS_JSON"
echo "ERROR: edge-probe engine failed (invalid shapes or bad input) — fix the requirement(s) and re-run; never proceed with empty coverage." >&2
exit 1
fi
rm -f "$REQS_JSON"
# Exit-0-but-garbage guard: the report must parse as JSON with the expected { items[], coverage{} } shape.
if ! printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let r;try{r=JSON.parse(s)}catch{process.exit(1)}if(!r||!Array.isArray(r.items)||typeof r.coverage!=="object"||r.coverage===null)process.exit(1)})'; then
echo "ERROR: edge-probe produced an unparseable or malformed coverage report — refusing to proceed with the resolution loop." >&2
exit 1
fi
# Zero-applicable guard: a report where the engine proposed NO applicable edge across ANY
# requirement is far more likely a shape-classification miss (or malformed requirements) than
# a genuinely edge-free spec — the same fail-open shape as an invalid shape yielding
# applicable:0. Surface it loudly; the author must explicitly confirm "no applicable edges"
# below rather than silently emitting a green empty ## Edge Coverage section.
APPLICABLE=$(printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let n=0;try{n=JSON.parse(s).coverage.applicable}catch{n=0}process.stdout.write(String(n))})')
if [ "$APPLICABLE" = "0" ]; then
echo "WARNING: edge-probe proposed ZERO applicable edges across all requirements — likely a classification miss or malformed requirements, not a genuinely edge-free spec. Do NOT silently write an empty Edge Coverage section." >&2
fi
If $APPLICABLE is 0, do NOT proceed silently: ask the author to confirm via AskUserQuestion
("The edge probe found no applicable edges for any requirement — is this genuinely an
edge-free spec, or should we revisit the requirement wording / authored shapes?"). Only write
an empty ## Edge Coverage section after explicit confirmation.
For each Requirement gathered so far:
- Classify its shape and raise only applicable edge categories (relevance filter — see
the taxonomy in the reference). Reuse any edges the Round-4 Failure Analyst already
surfaced as pre-
covered. - For each raised category, propose a CONCRETE candidate edge (not "consider
boundaries" — e.g. "R2 merges intervals; what about
[[1,2],[2,3]]that only touch?"). - Resolve each with the user (AskUserQuestion; text mode → numbered list):
- Specify it → write a new pass/fail line into Acceptance Criteria AND mark the
edge
covered. - Dismiss (reason) → mark
dismissedwith a required non-empty reason. - Backstop with a test → mark
backstop; note "held-out edge test" for plan-phase. - Defer → leave
unresolved. - An
unclassifiedrow (probeunclassified — review manually) means the requirement's prose matched no shape cue (#1110) — treat it like any other candidate (Specify, Dismiss (reason), or Defer). A manual-review nudge, not a hard block.
- Specify it → write a new pass/fail line into Acceptance Criteria AND mark the
edge
Soft gate (after resolving):
- All applicable edges resolved → proceed to Step 6.
- Any
unresolved→ AskUserQuestion:- header: "Edge Coverage"
- question: "[N] edge(s) are unresolved: [list]. What do you want to do?"
- options: "Resolve now" (loop back) / "Write SPEC.md anyway — flag unresolved" / "Keep probing"
- On "anyway": write SPEC.md with those rows marked
⚠ Edge unresolved — planner must treat as assumption.
--auto mode: auto-covered where a defensible acceptance criterion can be written;
otherwise auto-backstop (never auto-dismiss — a wrong dismissal is the exact silent
failure being eliminated). Log: [auto] edge coverage: C covered, B backstop, U unresolved.
unclassified exception (#1110): --auto leaves an unclassified candidate
unresolved (the soft gate surfaces it as a flagged planner assumption) — it never
auto-backstops it. A missing shape is not evidence an edge exists, so minting a held-out
edge obligation on a requirement that may be genuinely edge-free would be a false claim and
risks a vacuous edge test. Leaving it unresolved keeps the zero-cue requirement visible
(never a silent drop) without fabricating an edge — which is exactly #1110's purpose: surface
it for review, do not auto-handle it.
Populate the ## Edge Coverage section of SPEC.md from the resolved edges.
Step 5.6: Prohibition-Completeness Probe (must-NOT)
Run AFTER Step 5.5 (you probe the must-NOT axis of clear requirements, over the same requirement list). Reference: @~/.claude/gsd-core/references/prohibition-probe.md — the portable two-stage protocol, the canon-referral rule, and the status×verification schema live there (size-cap discipline; keep this step lean).
D1 — no compiled engine (ADR-550 D7b). Unlike Step 5.5, the prohibition probe has NO
compiled recall engine and runs NO node invocation here. The recall stage is an LLM prose
pass: the closed eight-category edge taxonomy a classifier can apply does not exist for the
open values/safety/ethics must-NOT axis. Do NOT copy the Step 5.5 engine-resolution block.
Only the schema/projection layer is real code; the recall is prose.
For each Requirement gathered so far, run the two-stage recall→precision pass:
- Stage 1 — Recall (adversarial prose probe). Ask the single adversarial question of the requirement: "What could this feature silently become that the author would NOT want, but the spec does not forbid?" Over-produce (~10 raw must-NOT candidates) — recall first.
- Stage 2 — Precision (one-pass classifier). Filter the raw list in a single pass: DROP routine-engineering items (normal correctness/hygiene — "must not mutate input", "must not throw on empty" — owned by the edge probe or code review); KEEP values / safety / ethics items (manipulative framing, protected-attribute proxies, raw PII in plaintext). This collapses ~10 → ~2–3 genuine prohibitions.
- Canon-referral (ADR-550 D6, PROB-13). A kept candidate that is canon security/compliance (OWASP / prototype-pollution / path-traversal / injection / GDPR / generic fairness) is NOT minted here — emit a one-line breadcrumb ("prototype-pollution is canon — owned by /gsd:secure-phase + eslint; not minted here") and DROP it. Minting canon items duplicates /gsd:secure-phase and drowns the bespoke signal.
- Resolve each surfaced (non-canon) prohibition (AskUserQuestion; text mode → numbered list):
- Keep it → write a NEGATIVE acceptance criterion (a must-NOT line) into Acceptance
Criteria AND mark the prohibition
resolvedwith a verification tier:test(a mechanical negative test/lint/assertion exists) orjudgment(real but not mechanically checkable — routes to judgment review).- Capture the wired-check descriptor on
test-tier (#1278, SOFT). When a prohibition is resolvedverification: test, ALSO capture the descriptor of the wired check soverify-phasecan LOCATE it deterministically (no verifier invention at verify time). Capture the flat scalars — persisted into SPEC and projected onto themust_haves.prohibitionsitem byprojectProhibitions:check_kind—node-test|lint-rule.check_target— the negative-test file path (fornode-test), or the path to lint (forlint-rule).check_rule— the eslint rule id (e.g.local/no-source-grep);lint-ruleonly.check_violation_fixture(#1279) — path to a KNOWN-BAD subject the wired check is run against to machine-prove fail-first; rides BOTH kinds. Capture it to let the item green end-to-end with zero hand-authoring at verify time; fornode-testthe negative test should read its subject from theGSD_PROHIB_SUBJECTenv var so the prover can inject this fixture.check_clean_fixture(#1346) — optional path to a KNOWN-CLEAN control subject. When captured, thenode-testprover also runs the check against it and requires GREEN — proving the violation's RED is caused by the subject's content, not byGSD_PROHIB_SUBJECTmerely being set. Capture it for a stronger guarantee; omit it and the check still proves fail-first on the violation alone (the content-causation residual stays documented for that case). This is a SOFT capture (CHK-04): atest-tier prohibition WITHOUT a descriptor is still allowed — if the author cannot yet name the wired check, leave the descriptor empty and proceed. It is NOT a hard authoring block; the item simply stays fail-closed/flagged downstream (an absent/partial descriptor — or one with nocheck_violation_fixture— →descriptorFromProjectionnull/under-specified/fixture-less → producer fail-closed locate-or-unprovable, never green). Do NOT capturefailFirsthere — it is a verify-time caller attestation, not a spec-authored field (#1279).
- Capture the wired-check descriptor on
- Dismiss (reason) → mark
dismissedwith a REQUIRED non-empty reason (PROB-05). The reason string is the audit trail; silence is not a valid dismissal. - Defer → leave
unresolved.
- Keep it → write a NEGATIVE acceptance criterion (a must-NOT line) into Acceptance
Criteria AND mark the prohibition
Soft gate (after resolving) — PROB-06:
- All applicable prohibitions resolved → proceed to Step 6.
- Any
unresolved→ AskUserQuestion:- header: "Prohibitions"
- question: "[N] prohibition(s) are unresolved: [list]. What do you want to do?"
- options: "Resolve now" (loop back) / "Write SPEC.md anyway — flag unresolved" / "Keep probing"
- On "anyway": write SPEC.md with those rows marked
⚠ Prohibition unresolved — planner must treat as assumption. This is a soft gate (write-anyway-with-flags), never a silent skip — the soft gate IS the control.
--auto mode: auto-resolved where a defensible negative acceptance criterion can be
written (test or judgment tier); otherwise leave unresolved. --auto NEVER auto-dismisses
a prohibition — a wrong dismissal is the exact silent failure this probe eliminates (PROB-06,
the load-bearing safety property). On a test-tier auto-resolution, capture the check_kind /
check_target / check_rule / check_violation_fixture / check_clean_fixture descriptor only when a wired check is unambiguous; otherwise
leave it empty — --auto NEVER fabricates a check path or fixture (a wrong locate is re-validated and
fails closed at the producer, but a fabricated path is still noise to avoid). Log:
[auto] prohibitions: R resolved, U unresolved.
Text mode (PROB-09): per Step 5's text-mode rule, replace the AskUserQuestion menus above with plain-text numbered lists — there is NO hard AskUserQuestion dependency, so the probe runs identically for non-Claude / text-mode hosts.
Populate the ## Prohibitions section of SPEC.md from the resolved prohibitions (each
resolved/test row is a checkable negative acceptance criterion; resolved/judgment
rows route to judgment review; ⚠ UNRESOLVED rows are flagged as assumptions). A
resolved/test row ALSO carries its captured check_kind / check_target / check_rule /
check_violation_fixture / check_clean_fixture descriptor when present (so the projection feeds verify-phase's deterministic locate + machine-proof + causation control, #1278 + #1279 + #1346);
a test row with no captured descriptor is still valid — it stays fail-closed/flagged
downstream rather than blocking authoring.
Step 6: Generate SPEC.md
Use the SPEC.md template from @~/.claude/gsd-core/templates/spec.md.
- Populate the Edge Coverage section from Step 5.5 (covered/dismissed/backstop/unresolved rows).
- Populate the Prohibitions section from Step 5.6 (resolved/dismissed/unresolved rows with the test|judgment tier).
Requirements for every requirement entry:
- One specific, testable statement
- Current state (what exists now)
- Target state (what it should become)
- Acceptance criterion (how to verify it was met)
Vague requirements are rejected:
- ✗ "The system should be fast"
- ✗ "Improve user experience"
- ✓ "API endpoint responds in < 200ms at p95 under 100 concurrent requests"
- ✓ "CLI command exits with code 1 and prints to stderr on invalid input"
Count requirements. The display in discuss-phase reads: "Found SPEC.md — {N} requirements locked."
Boundaries must be explicit lists:
- "In scope" — what this phase produces
- "Out of scope" — what it explicitly does NOT do (with brief reasoning)
Acceptance criteria must be pass/fail checkboxes — no "should feel good" or "looks reasonable."
If any dimensions are below minimum, mark them in the Ambiguity Report with: ⚠ Below minimum — planner must treat as assumption.
Write to: {phase_dir}/{padded_phase}-SPEC.md
Step 7: Commit
git add "${phase_dir}/${padded_phase}-SPEC.md"
git commit -m "spec(phase-${phase_number}): add SPEC.md for ${phase_name} — ${requirement_count} requirements (#2213)"
If commit_docs is false: Skip commit. Note that SPEC.md was written but not committed.
Step 8: Wrap Up
Display:
SPEC.md written — {N} requirements locked.
Phase {X}: {name}
Ambiguity: {final_score} (gate: ≤ 0.20)
Next: /gsd:discuss-phase {X}
discuss-phase will detect SPEC.md and focus on implementation decisions only.
<critical_rules>
- Every requirement MUST have current state, target state, and acceptance criterion
- Boundaries section is MANDATORY — cannot be empty
- "In scope" and "Out of scope" must be explicit lists, not narrative prose
- Acceptance criteria must be pass/fail — no subjective criteria
- SPEC.md is NEVER written if the user selects "Abandon"
- Do NOT ask about HOW to implement — that is discuss-phase territory
- Scout the codebase BEFORE the first question — grounded questions only
- Max 2–3 questions per round — do not frontload all questions at once
- Step 5.5 edge probe runs after the ambiguity gate; dismissals require a reason; --auto never auto-dismisses
- Step 5.6 prohibition probe runs after the edge probe; dismissals require a reason; --auto never auto-dismisses a prohibition </critical_rules>
<success_criteria>
- Codebase scouted and current state understood before questioning
- All 4 dimensions scored after every round
- Gate passed OR user explicitly chose to write despite gaps
- SPEC.md contains only falsifiable requirements
- Boundaries are explicit (in scope / out of scope with reasoning)
- Acceptance criteria are pass/fail checkboxes
- SPEC.md committed atomically (when commit_docs is true)
- User directed to /gsd:discuss-phase as next step
- Edge-completeness probe run; Edge Coverage section populated; unresolved edges flagged as assumptions
- Prohibition-completeness probe run; Prohibitions section populated; unresolved prohibitions flagged as assumptions </success_criteria>