Files
msd-core/gsd-core/workflows/new-milestone.md
Tom Boucher 63abcface9 feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools

The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git.

The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap.

unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell.

Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true.

An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes.

Closes #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3146): stop sync:launcher relocating a deliberate preamble placement

Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins.

Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture.

Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3146): backfill changeset PR number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#3146): document the FEATURES.md section-numbering practice

The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases.

Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set.

Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914).

Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight.

Refs #3146

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 20:57:16 -04:00

37 KiB
Raw Blame History

Start a new milestone cycle for an existing project. Loads project context, gathers milestone goals (from MILESTONE-CONTEXT.md or conversation), updates PROJECT.md and STATE.md, optionally runs parallel research, defines scoped requirements with REQ-IDs, spawns the roadmapper to create phased execution plan, and commits all artifacts. Brownfield equivalent of new-project.

<required_reading>

Read all files referenced by the invoking prompt's execution_context before starting.

</required_reading>

<available_agent_types> Valid GSD subagent types (use exact names — do not fall back to 'general-purpose'):

  • gsd-project-researcher — Researches project-level technical decisions
  • gsd-research-synthesizer — Synthesizes findings from parallel research agents
  • gsd-roadmapper — Creates phased execution roadmaps </available_agent_types>

1. Load Context

Parse $ARGUMENTS before doing anything else:

  • --reset-phase-numbers flag → opt into restarting roadmap phase numbering at 1. If absent, keep the current behavior of continuing phase numbering from the previous milestone.
  • --ws <name> flag → active workstream scope, parsed into GSD_WS
  • remaining text, with --ws <name> stripped → use as milestone name if present, captured into MILESTONE_ARG

Parse GSD_WS and MILESTONE_ARG using the established idiom (see verify-work.md):

_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif unset -f gsd_run; _G="$(command -v gsd_run)"; then GSD_TOOLS="$_G"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd_run is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
GSD_WS=""
echo "$ARGUMENTS" | grep -qE -- '--ws[[:space:]]+[A-Za-z0-9._-]+' && GSD_WS=$(echo "$ARGUMENTS" | grep -oE -- '--ws[[:space:]]+[A-Za-z0-9._-]+')
MILESTONE_ARG=$(echo "$ARGUMENTS" | sed -E 's/--ws[[:space:]]+[A-Za-z0-9._-]+//g' | xargs)
RESPONSE_LANGUAGE=$(gsd_run query config-get response_language --default "" 2>/dev/null || echo "")
# #2994: EARLY, section-manifest-only init.new-milestone call — needed here
# (before Step 4) to gate the project-md-milestone-write section. This is
# DELIBERATELY separate from Step 7's full init.new-milestone call below,
# which must stay AFTER Step 6's phase archival/phases.clear so its
# phase_dir_count / roadmap_exists / latest_completed_milestone fields
# reflect POST-archival state — moving that call here would compute those
# fields too early and corrupt the roadmapper's phase-numbering context.
# init.new-milestone is a pure read (no mutation), so calling it twice is
# safe; only `section_manifest` is consumed from this early call.
INIT_EARLY=$(gsd_run query init.new-milestone)
if [[ "$INIT_EARLY" == @file:* ]]; then INIT_EARLY=$(cat "${INIT_EARLY#@file:}"); fi

GSD_WS must chain to every downstream routing suggestion in this workflow (Step 4's shared-file guard, and the /gsd:discuss-phase//gsd:plan-phase routing hints below) per the routing-propagation contract in gsd-core/references/workstream-flag.md — never let it silently drop.

If response_language is set: All user-facing questions, prompts, and explanations in this workflow (including the "What do you want to build next?" prompt and seed-selection questions below) MUST be presented in {response_language}. Technical terms, code, file paths, and subagent prompts stay in English — only user-facing output is translated.

  • Read PROJECT.md (existing project, validated requirements, decisions)
  • Read MILESTONES.md (what shipped previously)
  • Read STATE.md (pending todos, blockers)
  • Check for MILESTONE-CONTEXT.md (from /gsd-discuss-milestone)

2. Gather Milestone Goals

If MILESTONE-CONTEXT.md exists:

  • Use features and scope from discuss-milestone
  • Present summary for confirmation

If no context file:

  • Present what shipped in last milestone

Text mode (workflow.text_mode: true in config or --text flag): Set TEXT_MODE=true if --text is present in $ARGUMENTS OR text_mode from init JSON is true. When TEXT_MODE is active, replace every AskUserQuestion call with a plain-text numbered list and ask the user to type their choice number. This is required for non-Claude runtimes (OpenAI Codex, Gemini CLI, etc.) where AskUserQuestion is not available.

  • Ask inline (freeform, NOT AskUserQuestion): "What do you want to build next?"
  • Wait for their response, then use AskUserQuestion to probe specifics
  • If user selects "Other" at any point to provide freeform input, ask follow-up as plain text — not another AskUserQuestion

2.5. Scan Planted Seeds

Check .planning/seeds/ for seed files that match the milestone goals gathered in step 2.

ls .planning/seeds/SEED-*.md 2>/dev/null

If no seed files exist: Skip this step silently — do not print any message or prompt.

If seed files exist: Read each SEED-*.md file and extract from its frontmatter and body:

  • Idea — the seed title (heading after frontmatter, e.g. # SEED-001: <idea>)
  • Trigger conditions — the trigger_when frontmatter field and the "When to Surface" section's bullet list
  • Planted during — the planted_during frontmatter field (for context)

Compare each seed's trigger conditions against the milestone goals from step 2. A seed matches when its trigger conditions are relevant to any of the milestone's target features or goals.

If no seeds match: Skip silently — do not prompt the user.

If matching seeds found:

--auto mode: Auto-select ALL matching seeds. Log: [auto] Selected N matching seed(s): [list seed names]

Text mode (TEXT_MODE=true): Present matching seeds as a plain-text numbered list:

Seeds that match your milestone goals:
1. SEED-001: <idea> (trigger: <trigger_when>)
2. SEED-003: <idea> (trigger: <trigger_when>)

Enter numbers to include (comma-separated), or "none" to skip:

Normal mode: Present via AskUserQuestion:

AskUserQuestion(
  header: "Seeds",
  question: "These planted seeds match your milestone goals. Include any in this milestone's scope?",
  multiSelect: true,
  options: [
    { label: "SEED-001: <idea>", description: "Trigger: <trigger_when> | Planted during: <planted_during>" },
    ...
  ]
)

After selection:

  • Selected seeds become additional context for requirement definition in step 9. Store them in an accumulator (e.g. $SELECTED_SEEDS) so step 9 can reference the ideas and their "Why This Matters" sections when defining requirements.
  • Unselected seeds remain untouched in .planning/seeds/ — never delete or modify seed files during this workflow.

3. Determine Milestone Version

  • Parse last version from MILESTONES.md
  • Suggest next version (v1.0 → v1.1, or v2.0 for major)
  • Confirm with user

3.5. Verify Milestone Understanding

Before writing any files, present a summary of what was gathered and ask for confirmation.

### GSD ► MILESTONE SUMMARY

**Milestone v[X.Y]: [Name]**

**Goal:** [One sentence]

**Target features:**
- [Feature 1]
- [Feature 2]
- [Feature 3]

**Key context:** [Any important constraints, decisions, or notes from questioning]

AskUserQuestion:

  • header: "Confirm?"
  • question: "Does this capture what you want to build in this milestone?"
  • options:
    • "Looks good" — Proceed to write PROJECT.md
    • "Adjust" — Let me correct or add details

If "Adjust": Ask what needs changing (plain text, NOT AskUserQuestion). Incorporate changes, re-present the summary. Loop until "Looks good" is selected.

If "Looks good": Proceed to Step 4.

4. Update PROJECT.md

PROJECT.md is shared across workstreams (gsd-core/references/workstream-flag.md marks it # Shared in the directory diagram). This step has two independently-scoped parts — only Part A is workstream-guarded.

If section_manifest (from INIT_EARLY) is null or "project-md-milestone-write" is in its included list: read and execute gsd-core/workflows/new-milestone/steps/project-md-milestone-write.md. Otherwise (a workstream is active) skip — do not read the file; Part B below still runs regardless of GSD_WS.

Part B — Evolution structural repair (always runs, regardless of GSD_WS). ## Evolution is a shared, idempotent structural section, not workstream state — a pre-Evolution project must be backfilled whether or not a workstream is active, so this part is NOT covered by Part A's skip. Ensure the ## Evolution section exists in PROJECT.md. If missing (projects created before this feature), add it before the footer:

## Evolution

This document evolves at phase transitions and milestone boundaries.

**After each phase transition** (via `/gsd-transition`):
1. Requirements invalidated? → Move to Out of Scope with reason
2. Requirements validated? → Move to Validated with phase reference
3. New requirements emerged? → Add to Active
4. Decisions to log? → Add to Key Decisions
5. "What This Is" still accurate? → Update if drifted

**After each milestone** (via `/gsd:complete-milestone`):
1. Full review of all sections
2. Core Value check — still the right priority?
3. Audit Out of Scope — reasons still valid?
4. Update Context with current state

5. Update STATE.md

Reset STATE.md frontmatter AND body atomically via the SDK. This writes the new milestone version/name into the YAML frontmatter, resets status to planning, zeroes progress.* counters, and rewrites the ## Current Position section to the new-milestone template. Accumulated Context (decisions, blockers, todos) is preserved across the switch — symmetric with milestone.complete.

OUTGOING_MILESTONE=$(gsd_run query state.get milestone --raw 2>/dev/null || true)
printf '%s' "$OUTGOING_MILESTONE" > .planning/.gsd-outgoing-milestone 2>/dev/null || true
echo "Outgoing milestone (phase history archives under THIS version in step 6): ${OUTGOING_MILESTONE:-<unknown>}"
gsd_run query state.milestone-switch --milestone "v[X.Y]" --name "[Name]"

Capture the outgoing version now. The lines above read the current (previous) milestone version BEFORE the switch flips STATE.md's milestone: field to the new one, and persist it to .planning/.gsd-outgoing-milestone so Step 6 can consume it via a shell variable — do NOT transcribe the echoed value into a later command by hand. Step 6 reads that file back into --archive-version so the previous milestone's phase directories archive under <outgoing-version>-phases/, not the new one (#2288). Once state.milestone-switch runs, current-milestone state no longer holds the outgoing version, which is why it is captured here.

The resulting Current Position section looks like:

## Current Position

Phase: Not started (defining requirements)
Plan: —
Status: Defining requirements
Last activity: [today] — Milestone v[X.Y] started

Bug #2630: a prior version of this workflow rewrote the Current Position body manually but left the frontmatter pointing at the previous milestone, so every downstream reader (state.json, getMilestoneInfo, progress bars) reported the stale milestone until the first phase advance forced a resync. Always use the SDK handler above — do not hand-edit STATE.md here.

6. Cleanup and Commit

Delete MILESTONE-CONTEXT.md if exists (consumed).

Clear leftover phase directories from the previous milestone. Read the outgoing version persisted in Step 5 back into a shell variable and pass it as --archive-version so the archive lands under the previous milestone's label — the switch in Step 5 has already advanced current-milestone state, so without this override the archive would be mislabeled with the new version (#2288). Use the shell variable directly (quoted) — never hand-retype the captured value into the command, so untrusted STATE.md content cannot be re-parsed by the shell:

OUTGOING_MILESTONE=$(cat .planning/.gsd-outgoing-milestone 2>/dev/null || true)
if [ -n "$OUTGOING_MILESTONE" ]; then
  gsd_run query phases.clear --confirm --archive-version "$OUTGOING_MILESTONE"
else
  gsd_run query phases.clear --confirm
fi
rm -f .planning/.gsd-outgoing-milestone 2>/dev/null || true

If the captured file is empty or absent (a fresh project with no prior milestone), the fallback branch runs phases.clear --confirm with no override — it then uses current-milestone state, and a dated archive label only if no version label is resolvable at all. phases.clear rejects any --archive-version value that is not a plain version token (no path separators or ..), so a malformed capture fails loudly rather than writing outside the archive directory.

Stage the phase archive move + source removal so they land in the same commit as the milestone start (atomic — no orphaned uncommitted deletions, no un-archived dirs carried forward). phases.clear archives each non-999 dir to milestones/<version>-phases/; staging both dirs captures the new archive and the removals together (#1871).

COMMIT_DOCS=$(gsd_run query config-get commit_docs 2>/dev/null || echo "true")
if [ "$COMMIT_DOCS" != "false" ]; then
  git add .planning/milestones/ .planning/phases/ 2>/dev/null || true
fi

When commit_docs is false, the archive move and phase removals are deliberately left unstaged here — not a bug — since Step 6's commit is skipped too.

Stage PROJECT.md in both modes. Step 4's Part A guard — not this commit — is what protects the shared ## Current Milestone heading (#2308): when a workstream is active Part A never writes it, so the only change PROJECT.md can carry here is Part B's idempotent ## Evolution backfill, which must be committed rather than stranded as a dangling edit. Do NOT reintroduce a [ -n "$GSD_WS" ] branch around this commit: GSD_WS is set in Step 1's shell and each step's bash block runs in its own shell (the same reason Step 5 round-trips OUTGOING_MILESTONE through a file), so such a guard reads an unset variable, always takes the flat-mode branch, and only appears to work.

gsd_run query commit "docs: start milestone v[X.Y] [Name]" --files .planning/PROJECT.md .planning/STATE.md

7. Load Context and Resolve Models

RESET_PHASE_NUMBERS_PARAM=""; if [[ "$ARGUMENTS" =~ (^|[[:space:]])--reset-phase-numbers([[:space:]]|$) ]]; then RESET_PHASE_NUMBERS_PARAM="--reset-phase-numbers"; fi
INIT=$(gsd_run query init.new-milestone $RESET_PHASE_NUMBERS_PARAM)
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi
AGENT_SKILLS_RESEARCHER=$(gsd_run query agent-skills gsd-project-researcher)
AGENT_SKILLS_SYNTHESIZER=$(gsd_run query agent-skills gsd-research-synthesizer)
AGENT_SKILLS_ROADMAPPER=$(gsd_run query agent-skills gsd-roadmapper)

Extract from init JSON: researcher_model, synthesizer_model, roadmapper_model, commit_docs, research_enabled, current_milestone, project_exists, roadmap_exists, latest_completed_milestone, phase_dir_count, phase_archive_path, agents_installed, missing_agents, project_path, roadmap_path, requirements_path, config_path, research_dir, milestones_path.

If agents_installed is false: Display a warning before proceeding:

⚠ GSD agents not installed. The following agents are missing from your agents directory:
  {missing_agents joined with newline}

Subagent spawns (gsd-project-researcher, gsd-research-synthesizer, gsd-roadmapper) will fail
with "agent type not found". Run the installer with --global to make agents available:

  npx @opengsd/gsd-core@latest --global

Proceeding without research subagents — roadmap will be generated inline.

Skip the parallel research spawn step and generate the roadmap inline.

If section_manifest is null or "reset-phase-safety" is in its included list: read and execute gsd-core/workflows/new-milestone/steps/reset-phase-safety.md. Otherwise skip — do not read the file.

8. Research Decision

Check research_enabled from init JSON (loaded from config).

If research_enabled is true:

AskUserQuestion: "Research the domain ecosystem for new features before defining requirements?"

  • "Research first (Recommended)" — Discover patterns, features, architecture for NEW capabilities
  • "Skip research for this milestone" — Go straight to requirements (does not change your default)

If research_enabled is false:

AskUserQuestion: "Research the domain ecosystem for new features before defining requirements?"

  • "Skip research (current default)" — Go straight to requirements
  • "Research first" — Discover patterns, features, architecture for NEW capabilities

IMPORTANT: Do NOT persist this choice to config.json. The workflow.research setting is a persistent user preference that controls plan-phase behavior across the project. Changing it here would silently alter future /gsd:plan-phase behavior. To change the default, use /gsd:settings.

If user chose "Research first":

### GSD ► RESEARCHING

◆ Spawning 4 researchers in parallel... (each runs in a subagent — no output until they return, ~1–5 min; expected, not a freeze)
  → Stack, Features, Architecture, Pitfalls
mkdir -p .planning/research

Spawn 4 parallel gsd-project-researcher agents. Each uses this template with dimension-specific fields:

Common structure for all 4 researchers:

Model omission (#2517). Omit the model parameter entirely when the value it would carry (researcher_model, synthesizer_model, roadmapper_model) is "inherit" or empty. An empty value 404s on runtimes without native tier aliases — the default on non-Claude runtimes. Omitting it inherits the orchestrator's model. See @gsd-core/references/model-profile-resolution.md.

Agent(prompt="
<research_type>Project Research — {DIMENSION} for [new features].</research_type>

<milestone_context>
SUBSEQUENT MILESTONE — Adding [target features] to existing app.
{EXISTING_CONTEXT}
Focus ONLY on what's needed for the NEW features.
</milestone_context>

<question>{QUESTION}</question>

<required_reading>
- {project_path} (Project context)
</required_reading>

${AGENT_SKILLS_RESEARCHER}

<downstream_consumer>{CONSUMER}</downstream_consumer>

<quality_gate>{GATES}</quality_gate>

<!-- #2508 runtime-aware-dispatch -->

> **Runtime-aware dispatch (#2508 Phase 4).** GSD workflows dispatch specialized subagents by role. Before dispatching on a built-in-only runtime (kimi-code — three built-ins only), resolve the role to a built-in via `gsd_run query resolve-dispatch-type --requested <role> --raw`. On named-dispatch runtimes (Claude/OpenCode/…) the role is returned unchanged; on kimi-code it maps to `coder`/`explore`/`plan` by role-suffix. The persona rides `${AGENT_SKILLS_<ROLE>}` (Phase 3) regardless. See @gsd-core/references/runtime-aware-dispatch.md.

<output>
Write to: {research_dir}/{FILE}
Use template: ~/.claude/gsd-core/templates/research-project/{FILE}
</output>
", subagent_type="gsd-project-researcher", model="{researcher_model}", description="{DIMENSION} research")

Dimension-specific fields:

Field Stack Features Architecture Pitfalls
EXISTING_CONTEXT Existing validated capabilities (DO NOT re-research): [from PROJECT.md] Existing features (already built): [from PROJECT.md] Existing architecture: [from PROJECT.md or codebase map] Focus on common mistakes when ADDING these features to existing system
QUESTION What stack additions/changes are needed for [new features]? How do [target features] typically work? Expected behavior? How do [target features] integrate with existing architecture? Common mistakes when adding [target features] to [domain]?
CONSUMER Specific libraries with versions for NEW capabilities, integration points, what NOT to add Table stakes vs differentiators vs anti-features, complexity noted, dependencies on existing Integration points, new components, data flow changes, suggested build order Warning signs, prevention strategy, which phase should address it
GATES Versions current (verify with Context7), rationale explains WHY, integration considered Categories clear, complexity noted, dependencies identified Integration points identified, new vs modified explicit, build order considers deps Pitfalls specific to adding these features, integration pitfalls covered, prevention actionable
FILE STACK.md FEATURES.md ARCHITECTURE.md PITFALLS.md

ORCHESTRATOR RULE — CODEX RUNTIME: After calling all 4 researcher Agent() calls above, do NOT read research files or synthesize content independently while the subagents are active. Wait for all 4 researchers to complete before spawning the synthesizer. This prevents duplicate work and wasted context.

After all 4 complete, spawn synthesizer:

Agent(prompt="
Synthesize research outputs into SUMMARY.md.

<required_reading>
- {research_dir}/STACK.md
- {research_dir}/FEATURES.md
- {research_dir}/ARCHITECTURE.md
- {research_dir}/PITFALLS.md
</required_reading>

${AGENT_SKILLS_SYNTHESIZER}

Write to: {research_dir}/SUMMARY.md
Use template: ~/.claude/gsd-core/templates/research-project/SUMMARY.md
Commit after writing.
", subagent_type="gsd-research-synthesizer", model="{synthesizer_model}", description="Synthesize research")

ORCHESTRATOR RULE — CODEX RUNTIME: After calling Agent() above, stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available.

Synthesizer output self-heal (#222) — verify SUMMARY.md materialized: The synthesizer's canonical output is .planning/research/SUMMARY.md on disk; its brief structured return (## SYNTHESIS COMPLETE plus a few ### confirmation lines) is NOT the file content. A known LLM false-refusal (issue #222) sometimes makes the agent return the full SUMMARY.md document inline — fabricating a write restriction (e.g. "the runtime is blocking file writes") — instead of writing the file. Prompt hardening alone does not fully eliminate it, so the orchestrator MUST absorb the failure deterministically before spawning gsd-roadmapper:

  1. Verify .planning/research/SUMMARY.md exists AND is substantive — non-empty, and free of any leftover <!-- gsd:write-continue --> continuation sentinel (which marks a truncated/incomplete write). You may validate with gsd_run verify-summary .planning/research/SUMMARY.md — it exits 0 regardless, so check its JSON passed field ("passed": false means missing or invalid), not the process exit code. If it passes, continue normally.
  2. If it is MISSING or invalid AND the synthesizer's return message contains the FULL SUMMARY.md document — recognizable by the template's top-level markers # Project Research Summary, ## Key Findings, ## Implications for Roadmap, and ## Sources, not merely the brief ## SYNTHESIS COMPLETE confirmation — the false-refusal fired: write that returned document to .planning/research/SUMMARY.md with the Write tool, then commit ALL research artifacts the synthesizer owns (it commits on behalf of the four researchers) with gsd_run query commit "docs: complete project research" --files .planning/research/ unless they are already committed. Log ⚠ #222 self-heal: synthesizer returned SUMMARY.md inline without writing it; orchestrator persisted the file.
  3. If it is MISSING or invalid AND the return is only a brief confirmation (no full SUMMARY document to recover), the synthesizer genuinely failed — surface the error and stop; do NOT spawn gsd-roadmapper against a missing or incomplete SUMMARY.md.

This guarantees gsd-roadmapper (which lists SUMMARY.md as required reading) never runs against a missing or truncated SUMMARY.md.

Display key findings from SUMMARY.md:

### GSD ► RESEARCH COMPLETE ✓

**Stack additions:** [from SUMMARY.md]
**Feature table stakes:** [from SUMMARY.md]
**Watch Out For:** [from SUMMARY.md]

If "Skip research": Continue to Step 9.

9. Define Requirements

### GSD ► DEFINING REQUIREMENTS

Read PROJECT.md: core value, current milestone goals, validated requirements (what exists).

If $SELECTED_SEEDS is non-empty (from step 2.5): Include selected seed ideas and their "Why This Matters" sections as additional input when defining requirements. Seeds provide user-validated feature ideas that should be incorporated into the requirement categories alongside research findings or conversation-gathered features.

If research exists: Read FEATURES.md, extract feature categories.

Present features by category:

## [Category 1]
**Table stakes:** Feature A, Feature B
**Differentiators:** Feature C, Feature D
**Research notes:** [any relevant notes]

If no research: Gather requirements through conversation. Ask: "What are the main things users need to do with [new features]?" Clarify, probe for related capabilities, group into categories.

Scope each category via AskUserQuestion (multiSelect: true, header max 12 chars):

  • "[Feature 1]" — [brief description]
  • "[Feature 2]" — [brief description]
  • "None for this milestone" — Defer entire category

Track: Selected → this milestone. Unselected table stakes → future. Unselected differentiators → out of scope.

Identify gaps via AskUserQuestion:

  • "No, research covered it" — Proceed
  • "Yes, let me add some" — Capture additions

Generate REQUIREMENTS.md:

  • v1 Requirements grouped by category (checkboxes, REQ-IDs)
  • Future Requirements (deferred)
  • Out of Scope (explicit exclusions with reasoning)
  • Traceability section (empty, filled by roadmap)

REQ-ID format: [CATEGORY]-[NUMBER] (AUTH-01, NOTIF-02). Continue numbering from existing.

Requirement quality criteria:

Good requirements are:

  • Specific and testable: "User can reset password via email link" (not "Handle password reset")
  • User-centric: "User can X" (not "System does Y")
  • Atomic: One capability per requirement (not "User can login and manage profile")
  • Independent: Minimal dependencies on other requirements

Present FULL requirements list for confirmation:

## Milestone v[X.Y] Requirements

### [Category 1]
- [ ] **CAT1-01**: User can do X
- [ ] **CAT1-02**: User can do Y

### [Category 2]
- [ ] **CAT2-01**: User can do Z

Does this capture what you're building? (yes / adjust)

If "adjust": Return to scoping.

Commit requirements:

gsd_run query commit "docs: define milestone v[X.Y] requirements" --files .planning/REQUIREMENTS.md

10. Create Roadmap

### GSD ► CREATING ROADMAP

◆ Spawning roadmapper... (runs in a subagent — no output until it returns, ~1–5 min; expected, not a freeze)

Starting phase number:

  • If --reset-phase-numbers is active, start at Phase 1
  • Otherwise, continue from the previous milestone's last phase number (v1.0 ended at phase 5 → v1.1 starts at phase 6)
Agent(prompt="
<planning_context>
<required_reading>
- {project_path}
- {requirements_path}
- {research_dir}/SUMMARY.md (if exists)
- {config_path}
- {milestones_path}
</required_reading>

${AGENT_SKILLS_ROADMAPPER}

</planning_context>

<instructions>
Create roadmap for milestone v[X.Y]:
1. Respect the selected numbering mode:
   - `--reset-phase-numbers` → start at Phase 1
   - default behavior → continue from the previous milestone's last phase number
2. Derive phases from THIS MILESTONE's requirements only
3. Map every requirement to exactly one phase
4. Derive 2-5 success criteria per phase (observable user behaviors)
5. Validate 100% coverage
6. Write files immediately (ROADMAP.md, STATE.md, update REQUIREMENTS.md traceability)
7. Return ROADMAP CREATED with summary

Write files first, then return.
</instructions>
", subagent_type="gsd-roadmapper", model="{roadmapper_model}", description="Create roadmap")

ORCHESTRATOR RULE — CODEX RUNTIME: After calling Agent() above, stop working on this task immediately. Do not read more files, edit code, or run tests related to this task while the subagent is active. Wait for the subagent to return its result. This prevents duplicate work, conflicting edits, and wasted context. Only resume when the subagent result is available.

Handle return:

If ## ROADMAP BLOCKED: Present blocker, work with user, re-spawn.

If ## ROADMAP CREATED: Read ROADMAP.md, present inline:

## Proposed Roadmap

**[N] phases** | **[X] requirements mapped** | All covered ✓

| # | Phase | Goal | Requirements | Success Criteria |
|---|-------|------|--------------|------------------|
| [N] | [Name] | [Goal] | [REQ-IDs] | [count] |

### Phase Details

**Phase [N]: [Name]**
Goal: [goal]
Requirements: [REQ-IDs]
Success criteria:
1. [criterion]
2. [criterion]

Ask for approval via AskUserQuestion:

  • "Approve" — Commit and continue
  • "Adjust phases" — Tell me what to change
  • "Review full file" — Show raw ROADMAP.md

If "Adjust": Get notes, re-spawn roadmapper with revision context, loop until approved. If "Review": Display raw ROADMAP.md, re-ask.

Commit roadmap (after approval):

gsd_run query commit "docs: create milestone v[X.Y] roadmap ([N] phases)" --files .planning/ROADMAP.md .planning/STATE.md .planning/REQUIREMENTS.md

After roadmap approval, scan pending todos against the newly approved phases. For each todo whose scope matches a phase, tag it with resolves_phase: N in its YAML frontmatter.

Check for pending todos:

PENDING_TODOS=$(ls .planning/todos/pending/*.md 2>/dev/null | head -50)

If no pending todos exist: Skip this step silently.

If pending todos exist:

Read the approved ROADMAP.md and extract the phase list: phase number, phase name, goal, and requirement IDs.

For each pending todo, compare:

  • The todo's title and area frontmatter fields
  • The todo body (Problem and Solution sections)

Against each phase's:

  • Phase goal
  • Requirement IDs and descriptions

Match criteria (best-effort — do not over-match): A todo is considered resolved by a phase if the phase's goal or requirements directly describe implementing the same feature, area, or capability as the todo. Narrow, specific todos with concrete scopes are the best candidates. Vague or cross-cutting todos should be left unlinked.

For each matched todo, add resolves_phase: [N] to the YAML frontmatter block (after the existing fields):

---
created: [existing]
title: [existing]
area: [existing]
resolves_phase: [N]
files: [existing]
---

Only modify todos that have a clear, confident match. Leave unmatched todos unmodified.

If any todos were linked:

gsd_run query commit "docs: tag [count] pending todos with resolves_phase after milestone v[X.Y] roadmap" --files .planning/todos/pending/*.md

Print a summary:

◆ Linked [N] pending todos to roadmap phases:
  → [todo title] → Phase [N]: [Phase Name]
  (Leave [M] unmatched todos in pending/)

11. Done

### GSD ► MILESTONE INITIALIZED ✓

**Milestone v[X.Y]: [Name]**

| Artifact       | Location                    |
|----------------|-----------------------------|
| Project        | `.planning/PROJECT.md`      |
| Research       | `.planning/research/`       |
| Requirements   | `.planning/REQUIREMENTS.md` |
| Roadmap        | `.planning/ROADMAP.md`      |

**[N] phases** | **[X] requirements** | Ready to build ✓

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Phase [N]: [Phase Name]** — [Goal]

`/clear` then:

`/gsd:discuss-phase [N] ${GSD_WS}` — gather context and clarify approach

Also: `/gsd:plan-phase [N] ${GSD_WS}` — skip discussion, plan directly

<success_criteria>

  • PROJECT.md updated with Current Milestone section (skipped when a workstream is active — shared file, see Step 4)
  • STATE.md reset for new milestone
  • MILESTONE-CONTEXT.md consumed and deleted (if existed)
  • Research completed (if selected) — 4 parallel agents, milestone-aware
  • Requirements gathered and scoped per category
  • REQUIREMENTS.md created with REQ-IDs
  • gsd-roadmapper spawned with phase numbering context
  • Roadmap files written immediately (not draft)
  • User feedback incorporated (if any)
  • Phase numbering mode respected (continued or reset)
  • All commits made (if planning docs committed)
  • Pending todos scanned for phase matches; matched todos tagged with resolves_phase: N
  • User knows next step: /gsd:discuss-phase [N] ${GSD_WS}

Atomic commits: Each phase commits its artifacts immediately. </success_criteria>