Files
msd-core/gsd-core/workflows/spec-phase.md
Rezolv 3e836fef0d feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/

Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.

Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.

Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
  planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
  source-checkout-gated build fallback) instead of LLM re-derivation; the engine
  capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
  resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
  per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
  validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.

Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.

* test(#550): RED — status×verification re-cut + probe-core engine specs

Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):

  status: resolved | dismissed | unresolved   (lifecycle, shared)
  verification: explicit | backstop | null     (only when resolved)

- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
  to be extracted — validateResolution(r, validators), validateRequirement,
  analyzeCoverage(items, resolutions?, validators), byVerification rollup,
  runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
  {resolved,backstop}; coverage gains byVerification.{explicit,backstop};
  proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
  COUNT preserved on every fixture (closed set = resolved+dismissed; doc
  line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
  rewritten to the two-axis model.

Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).

* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)

Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.

probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
  verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
  items[] (core never assumes propose is deterministic — edge resolves via
  LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
  count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
  {categories, verification, requiredFieldsByVerification} (ADR-550 #5)

edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.

* chore(#550): register probe-core.cjs artifact in ledgers + inventory

New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:

- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
  never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
  is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
  edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).

probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.

* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]

trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.

Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.

* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard

Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:

- validateResolution now enforces the 'verification is null unless resolved'
  invariant for EVERY status (not just resolved): a dismissed/unresolved
  resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
  (was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
  it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
  contract; a new test locks that an all-dismissed run is NOT affirmatively covered
  (byVerification is the honest gate).

Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.

* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics

Re-review #5 (trek-e) clarity edits:

- Decision 5: annotate that only contract item (a) ships on #584 (the edge
  adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
  parenthetical — the blessed/implemented semantics are count-preserved = the
  CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
  carrying the per-tier resolved-status breakdown. The old parenthetical
  contradicted the shipped count.

* test(#550): cover runProbeCli structural-guard numeric-count branch

Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.

* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)

The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.

* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)

A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.

* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)

templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.

* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)

The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.

* docs(#550): add how-to for resolving edge-coverage findings (B1)

Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.

* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)

trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.

* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)

trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.

* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)

trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.

* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe

Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:

  - spec-phase.md  15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
  - plan-phase.md  93135 -> 94253 (+1118): covered/backstop edge lift into
    must_haves.truths (the live <downstream_consumer> block)

Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.

* chore(#550): reconcile INVENTORY headline counts after rebase onto next

Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
  - References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
  - CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs

Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
2026-06-12 11:05:31 -04:00

23 KiB
Raw Blame History

Clarify WHAT a phase delivers through a Socratic interview loop with quantitative ambiguity scoring. Produces a SPEC.md with falsifiable requirements that discuss-phase treats as locked decisions.

This workflow handles "what" and "why" — discuss-phase handles "how".

<ambiguity_model> Score each dimension 0.0 (completely unclear) to 1.0 (crystal clear):

Dimension Weight Minimum What it measures
Goal Clarity 35% 0.75 Is the outcome specific and measurable?
Boundary Clarity 25% 0.70 What's in scope vs out of scope?
Constraint Clarity 20% 0.65 Performance, compatibility, data requirements?
Acceptance Criteria 20% 0.70 How do we know it's done?

Ambiguity score = 1.0 − (0.35×goal + 0.25×boundary + 0.20×constraint + 0.20×acceptance)

Gate: ambiguity ≤ 0.20 AND all dimensions ≥ their minimums → ready to write SPEC.md.

A score of 0.20 means 80% weighted clarity — enough precision that the planner won't silently make wrong assumptions. </ambiguity_model>

<interview_perspectives> Rotate through these perspectives — each naturally surfaces different blindspots:

Researcher (rounds 1–2): Ground the discussion in current reality.

  • "What exists in the codebase today related to this phase?"
  • "What's the delta between today and the target state?"
  • "What triggers this work — what's broken or missing?"

Simplifier (round 2): Surface minimum viable scope.

  • "What's the simplest version that solves the core problem?"
  • "If you had to cut 50%, what's the irreducible core?"
  • "What would make this phase a success even without the nice-to-haves?"

Boundary Keeper (round 3): Lock the perimeter.

  • "What explicitly will NOT be done in this phase?"
  • "What adjacent problems is it tempting to solve but shouldn't?"
  • "What does 'done' look like — what's the final deliverable?"

Failure Analyst (round 4): Find the edge cases that invalidate requirements.

  • "What's the worst thing that could go wrong if we get the requirements wrong?"
  • "What does a broken version of this look like?"
  • "What would cause a verifier to reject the output?"

Seed Closer (rounds 5–6): Lock remaining undecided territory.

  • "We have [dimension] at [score] — what would make it completely clear?"
  • "The remaining ambiguity is in [area] — can we make a decision now?"
  • "Is there anything you'd regret not specifying before planning starts?" </interview_perspectives>

Step 1: Initialize

_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
INIT=$(gsd_run init phase-op "${PHASE}")
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi

Parse JSON for: phase_found, phase_dir, phase_number, phase_name, phase_slug, padded_phase, state_path, requirements_path, roadmap_path, planning_path, response_language, commit_docs.

If response_language is set: All user-facing text in this workflow MUST be in {response_language}. Technical terms, code, and file paths stay in English.

If phase_found is false:

Phase [X] not found in roadmap.
Use /gsd:progress to see available phases.

Exit.

Check for existing SPEC.md:

ls ${phase_dir}/*-SPEC.md 2>/dev/null | grep -v AI-SPEC | head -1 || true

If SPEC.md already exists:

If --auto: Auto-select "Update it". Log: [auto] SPEC.md exists — updating.

Otherwise: Use AskUserQuestion:

  • header: "Spec"
  • question: "Phase [X] already has a SPEC.md. What do you want to do?"
  • options:
    • "Update it" — Revise and re-score
    • "View it" — Show current spec
    • "Skip" — Exit (use existing spec as-is)

If "View": Display SPEC.md, then offer Update/Skip. If "Skip": Exit with message: "Existing SPEC.md unchanged. Run /gsd:discuss-phase [X] to continue." If "Update": Load existing SPEC.md, continue to Step 3.

Step 2: Scout Codebase

Read these files before any questions:

  • {requirements_path} — Project requirements
  • {state_path} — Decisions already made, current phase, blockers
  • ROADMAP.md phase entry — Phase description, goals, canonical refs

Grep the codebase for code/files relevant to this phase goal. Look for:

  • Existing implementations of similar functionality
  • Integration points where new code will connect
  • Test coverage gaps relevant to the phase
  • Prior phase artifacts (SUMMARY.md, VERIFICATION.md) that inform current state

Synthesize current state — the grounded baseline for the interview:

  • What exists today related to this phase
  • The gap between current state and the phase goal
  • The primary deliverable: what file/behavior/capability does NOT exist yet?

Confirm your current state synthesis internally. Do not present it to the user yet — you'll use it to ask precise, grounded questions.

Step 3: First Ambiguity Assessment

Before questioning begins, score the phase's current ambiguity based only on what ROADMAP.md and REQUIREMENTS.md say:

Goal Clarity:       [score 0.0–1.0]
Boundary Clarity:   [score 0.0–1.0]
Constraint Clarity: [score 0.0–1.0]
Acceptance Criteria:[score 0.0–1.0]

Ambiguity: [score] ([calculate])

If --auto and initial ambiguity already ≤ 0.20 with all minimums met: Skip interview — derive SPEC.md directly from roadmap + requirements. Log: [auto] Phase requirements are already sufficiently clear — generating SPEC.md from existing context. Jump to Step 6.

Otherwise: Continue to Step 4.

Step 4: Socratic Interview Loop

Max 6 rounds. Each round: 2–3 questions max. End round after user responds.

Round selection by perspective:

  • Round 1: Researcher
  • Round 2: Researcher + Simplifier
  • Round 3: Boundary Keeper
  • Round 4: Failure Analyst
  • Rounds 5–6: Seed Closer (focus on lowest-scoring dimensions)

After each round:

  1. Update all 4 dimension scores from the user's answers
  2. Calculate new ambiguity score
  3. Display the updated scoring:
After round [N]:
  Goal Clarity:       [score] (min 0.75) [✓ or ↑ needed]
  Boundary Clarity:   [score] (min 0.70) [✓ or ↑ needed]
  Constraint Clarity: [score] (min 0.65) [✓ or ↑ needed]
  Acceptance Criteria:[score] (min 0.70) [✓ or ↑ needed]
  Ambiguity: [score] (gate: ≤ 0.20)

Gate check after each round:

If gate passes (ambiguity ≤ 0.20 AND all minimums met):

If --auto: Jump to Step 6.

Otherwise: AskUserQuestion:

  • header: "Spec Gate Passed"
  • question: "Ambiguity is [score] — requirements are clear enough to write SPEC.md. Proceed?"
  • options:
    • "Yes — write SPEC.md" → Jump to Step 6
    • "One more round" → Continue interview
    • "Done talking — write it" → Jump to Step 6

If max rounds reached (6) and gate not passed:

If --auto: Write SPEC.md anyway — flag unresolved dimensions. Log: [auto] Max rounds reached. Writing SPEC.md with [N] dimensions below minimum. Planner will need to treat these as assumptions.

Otherwise: AskUserQuestion:

  • header: "Max Rounds"
  • question: "After 6 rounds, ambiguity is [score]. [List dimensions still below minimum.] What would you like to do?"
  • options:
    • "Write SPEC.md anyway — flag gaps" → Write SPEC.md, mark unresolved dimensions in Ambiguity Report
    • "Keep talking" → Continue (no round limit from here)
    • "Abandon" → Exit without writing

If --auto mode throughout: Replace all AskUserQuestion calls above with Claude's recommended choice. Log decisions inline. Apply the same logic as --auto in discuss-phase.

Text mode (workflow.text_mode: true or --text flag): Use plain-text numbered lists instead of AskUserQuestion TUI menus.

Step 5: (covered inline — ambiguity scoring is per-round)

Step 5.5: Edge-Completeness Probe

Run AFTER the ambiguity gate passes (you probe edges of clear requirements, not vague ones). Reference: @~/.claude/gsd-core/references/edge-probe.md.

Runtime coverage compute — resolve and invoke edge-probe.cjs:

# Resolve the compiled edge-probe.cjs against the GSD install dir via RUNTIME_DIR (#448)
# — NOT the consuming project's git root — falling back to git toplevel / $HOME/.claude.
# Mirrors the ui-safety-gate.cjs resolution idiom at autonomous.md:290 / plan-phase.md:631.
_GSD_RT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"
EDGE_PROBE_JS=$(for _c in \
  "$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
  "$_GSD_RT/bin/lib/edge-probe.cjs" \
  "$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
  "$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
  "$HOME/.claude/bin/lib/edge-probe.cjs"; do
  [ -f "$_c" ] && { echo "$_c"; break; }
done)

# Graceful degradation — never silent skip (RR-04). Build ONLY when $_GSD_RT is a verified
# GSD source checkout (has tsconfig.build.json + src/edge-probe.cts), and pin npm to it with
# --prefix so we never trigger the CONSUMING project's own build:lib (its cwd package scripts:
# codegen/migrations/writes) during a spec workflow. Real installs ship the compiled .cjs via
# prepublishOnly, so this build path only matters in a GSD dev checkout (review High).
if [ -z "$EDGE_PROBE_JS" ]; then
  if [ -f "$_GSD_RT/tsconfig.build.json" ] && [ -f "$_GSD_RT/src/edge-probe.cts" ]; then
    npm --prefix "$_GSD_RT" run build:lib 2>/dev/null || true
    EDGE_PROBE_JS=$(for _c in \
      "$_GSD_RT/gsd-core/bin/lib/edge-probe.cjs" \
      "$_GSD_RT/bin/lib/edge-probe.cjs" \
      "$_GSD_RT/.claude/bin/lib/edge-probe.cjs" \
      "$HOME/.claude/gsd-core/bin/lib/edge-probe.cjs" \
      "$HOME/.claude/bin/lib/edge-probe.cjs"; do
      [ -f "$_c" ] && { echo "$_c"; break; }
    done)
  fi
  if [ -z "$EDGE_PROBE_JS" ]; then
    echo "ERROR: edge-probe.cjs not found — reinstall GSD or run \`npm run build:lib\` in your GSD checkout." >&2
    exit 1
  fi
fi

# Write the Requirements gathered in THIS spec session to a temp JSON, then invoke the
# canonical coverage compute. Populate the heredoc from the SPEC's Requirements — one object
# per requirement: {"id","text","shapes"?}. This is the load-bearing step: an empty file makes
# the probe a no-op, so the guard below fails loud rather than silently skipping (RR-04).
REQS_JSON=$(mktemp "${TMPDIR:-/tmp}/edge-probe-reqs-XXXXXX.json")
cat > "$REQS_JSON" <<'JSON'
[
  { "id": "R1", "text": "<replace: requirement text from the SPEC>" }
]
JSON
# Guard — never invoke on an empty/invalid array, OR one still holding the heredoc
# `<replace: …>` placeholder (a forgotten substitution would otherwise yield a
# meaningful-looking but bogus coverage report). Fail loud, not silent no-op.
if ! node -e 'const a=require(process.argv[1]);if(!Array.isArray(a)||a.length===0)process.exit(1);if(a.some(r=>typeof r.text!=="string"||!r.text.trim()||r.text.includes("<replace:")))process.exit(1)' "$REQS_JSON" 2>/dev/null; then
  echo "ERROR: edge-probe requirements JSON is empty/invalid or still holds the <replace: …> placeholder — populate \$REQS_JSON from the SPEC Requirements before Step 5.5 runs." >&2
  exit 1
fi
# Invoke the compiled engine and CAPTURE its report — it computes which categories apply per
# requirement. The covered/backstop/dismissed/unresolved rows in $COVERAGE drive the
# resolution loop below (canonical taxonomy compute, NOT LLM re-derivation from prose).
# The engine FAILS CLOSED (exit 2) on an invalid authored shape or bad input — so the capture
# MUST be exit-checked. A bare `COVERAGE=$(node …)` swallows that exit code, leaves $COVERAGE
# empty, and lets the workflow fall through to prose re-derivation: fail-OPEN at the boundary
# the engine validation exists to protect. Make the run fatal, then validate the captured
# report is well-formed JSON before the resolution loop consumes it.
if ! COVERAGE=$(node "$EDGE_PROBE_JS" "$REQS_JSON"); then
  rm -f "$REQS_JSON"
  echo "ERROR: edge-probe engine failed (invalid shapes or bad input) — fix the requirement(s) and re-run; never proceed with empty coverage." >&2
  exit 1
fi
rm -f "$REQS_JSON"
# Exit-0-but-garbage guard: the report must parse as JSON with the expected { items[], coverage{} } shape.
if ! printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let r;try{r=JSON.parse(s)}catch{process.exit(1)}if(!r||!Array.isArray(r.items)||typeof r.coverage!=="object"||r.coverage===null)process.exit(1)})'; then
  echo "ERROR: edge-probe produced an unparseable or malformed coverage report — refusing to proceed with the resolution loop." >&2
  exit 1
fi
# Zero-applicable guard: a report where the engine proposed NO applicable edge across ANY
# requirement is far more likely a shape-classification miss (or malformed requirements) than
# a genuinely edge-free spec — the same fail-open shape as an invalid shape yielding
# applicable:0. Surface it loudly; the author must explicitly confirm "no applicable edges"
# below rather than silently emitting a green empty ## Edge Coverage section.
APPLICABLE=$(printf '%s' "$COVERAGE" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{let n=0;try{n=JSON.parse(s).coverage.applicable}catch{n=0}process.stdout.write(String(n))})')
if [ "$APPLICABLE" = "0" ]; then
  echo "WARNING: edge-probe proposed ZERO applicable edges across all requirements — likely a classification miss or malformed requirements, not a genuinely edge-free spec. Do NOT silently write an empty Edge Coverage section." >&2
fi

If $APPLICABLE is 0, do NOT proceed silently: ask the author to confirm via AskUserQuestion ("The edge probe found no applicable edges for any requirement — is this genuinely an edge-free spec, or should we revisit the requirement wording / authored shapes?"). Only write an empty ## Edge Coverage section after explicit confirmation.

For each Requirement gathered so far:

  1. Classify its shape and raise only applicable edge categories (relevance filter — see the taxonomy in the reference). Reuse any edges the Round-4 Failure Analyst already surfaced as pre-covered.
  2. For each raised category, propose a CONCRETE candidate edge (not "consider boundaries" — e.g. "R2 merges intervals; what about [[1,2],[2,3]] that only touch?").
  3. Resolve each with the user (AskUserQuestion; text mode → numbered list):
    • Specify it → write a new pass/fail line into Acceptance Criteria AND mark the edge covered.
    • Dismiss (reason) → mark dismissed with a required non-empty reason.
    • Backstop with a test → mark backstop; note "held-out edge test" for plan-phase.
    • Defer → leave unresolved.

Soft gate (after resolving):

  • All applicable edges resolved → proceed to Step 6.
  • Any unresolved → AskUserQuestion:
    • header: "Edge Coverage"
    • question: "[N] edge(s) are unresolved: [list]. What do you want to do?"
    • options: "Resolve now" (loop back) / "Write SPEC.md anyway — flag unresolved" / "Keep probing"
    • On "anyway": write SPEC.md with those rows marked ⚠ Edge unresolved — planner must treat as assumption.

--auto mode: auto-covered where a defensible acceptance criterion can be written; otherwise auto-backstop (never auto-dismiss — a wrong dismissal is the exact silent failure being eliminated). Log: [auto] edge coverage: C covered, B backstop, U unresolved.

Populate the ## Edge Coverage section of SPEC.md from the resolved edges.

Step 6: Generate SPEC.md

Use the SPEC.md template from @~/.claude/gsd-core/templates/spec.md.

  • Populate the Edge Coverage section from Step 5.5 (covered/dismissed/backstop/unresolved rows).

Requirements for every requirement entry:

  • One specific, testable statement
  • Current state (what exists now)
  • Target state (what it should become)
  • Acceptance criterion (how to verify it was met)

Vague requirements are rejected:

  • ✗ "The system should be fast"
  • ✗ "Improve user experience"
  • ✓ "API endpoint responds in < 200ms at p95 under 100 concurrent requests"
  • ✓ "CLI command exits with code 1 and prints to stderr on invalid input"

Count requirements. The display in discuss-phase reads: "Found SPEC.md — {N} requirements locked."

Boundaries must be explicit lists:

  • "In scope" — what this phase produces
  • "Out of scope" — what it explicitly does NOT do (with brief reasoning)

Acceptance criteria must be pass/fail checkboxes — no "should feel good" or "looks reasonable."

If any dimensions are below minimum, mark them in the Ambiguity Report with: ⚠ Below minimum — planner must treat as assumption.

Write to: {phase_dir}/{padded_phase}-SPEC.md

Step 7: Commit

git add "${phase_dir}/${padded_phase}-SPEC.md"
git commit -m "spec(phase-${phase_number}): add SPEC.md for ${phase_name} — ${requirement_count} requirements (#2213)"

If commit_docs is false: Skip commit. Note that SPEC.md was written but not committed.

Step 8: Wrap Up

Display:

SPEC.md written — {N} requirements locked.

  Phase {X}: {name}
  Ambiguity: {final_score} (gate: ≤ 0.20)

Next: /gsd:discuss-phase {X}
  discuss-phase will detect SPEC.md and focus on implementation decisions only.

<critical_rules>

  • Every requirement MUST have current state, target state, and acceptance criterion
  • Boundaries section is MANDATORY — cannot be empty
  • "In scope" and "Out of scope" must be explicit lists, not narrative prose
  • Acceptance criteria must be pass/fail — no subjective criteria
  • SPEC.md is NEVER written if the user selects "Abandon"
  • Do NOT ask about HOW to implement — that is discuss-phase territory
  • Scout the codebase BEFORE the first question — grounded questions only
  • Max 2–3 questions per round — do not frontload all questions at once
  • Step 5.5 edge probe runs after the ambiguity gate; dismissals require a reason; --auto never auto-dismisses </critical_rules>

<success_criteria>

  • Codebase scouted and current state understood before questioning
  • All 4 dimensions scored after every round
  • Gate passed OR user explicitly chose to write despite gaps
  • SPEC.md contains only falsifiable requirements
  • Boundaries are explicit (in scope / out of scope with reasoning)
  • Acceptance criteria are pass/fail checkboxes
  • SPEC.md committed atomically (when commit_docs is true)
  • User directed to /gsd:discuss-phase as next step
  • Edge-completeness probe run; Edge Coverage section populated; unresolved edges flagged as assumptions </success_criteria>