* test(#3409): failing-first regression tests for unreachable shell guard arms Drives the three live defects fail-first, executing the shipped workflow snippets rather than a re-typed copy: - G1/G2 plan-phase.md Walking Skeleton gate reads `--pick summaries_total`, a field that does not exist, so PRIOR_SUMMARIES is always "" and the gate has never fired (#3365). G2 is the load-bearing negative-space case: it rejects a fix that treats "no answer" as "zero" and fires unconditionally. - G3 plan-phase.md PHASE_REQ_IDS resolves "" instead of the TBD sentinel on a phase with zero requirements. - G4 complete-milestone.md's bare `cat <glob>` blocks on stdin under a nullglob left set by an earlier block (measured hang). Skipped on Windows for G4 only: the FIFO-blocked-stdin mechanism is POSIX only, and a weakened assertion there would pass vacuously. Refs #3409 * fix(#3409): make nine shell guards observe their own failure arm `--pick` coerces a missing field to empty string and exits 0, so the `|| echo <default>` fallback after it fires only on a verb typo, never on the field absence it was written for. Nine sites relied on that arm. - plan-phase.md walking-skeleton gate: `--pick summaries_total` names a field that does not exist under any flag combination, so the gate has never fired on any project (#3365). Repointed at the existing single owner, `phases.list --type summaries --pick count`, which returns a real integer in every case including a project with no `.planning` directory. No new counter is added: a second one would duplicate the ownership ADR-3180 Decision 1 forbids. The gate now fires only on a literal "0", so an unanswerable query fails safe instead of entering skeleton mode. - plan-phase.md phase_req_ids: now falls back to the documented TBD. - The remaining seven convert to an explicit empty test. - complete-milestone.md read all phase summaries through a bare `cat <glob>`; under a nullglob left set by an earlier block that is zero operands, so cat blocks on stdin. Guarded with the array shape the #3300 fix already established in review.md. Refs #3409 * fix(#3409): guard eleven more globs that defeat their own fallback arm The nullglob audit this issue asks for turned up the same class in files #3300 never touched. - Eight bare `cat <glob>` reads (transition, complete-milestone, planner x4, verifier, phase-researcher). With nullglob set that is zero operands, so cat reads stdin and blocks; measured rc=137 at 3s. - Three `ls <glob> || echo "<message>"` sites (session-report, review-backlog and its generated skill). nullglob makes ls succeed listing the cwd, so the message never prints and the user gets a directory listing instead. Guarded with `[ -e "${_ARR[0]}" ]` rather than `[ ${#_ARR[@]} -gt 0 ]`. The count form is correct only when nullglob is set, and six of these seven files never set it: without it the array holds the unmatched literal pattern, so the count is 1 and the guard passes wrongly. `-e` is correct in both worlds. review.md keeps its count guards — that block sets nullglob two lines above them. skills/gsd-review-backlog regenerated from commands/, never hand-edited. Refs #3409 * feat(#3409): add the unreachable-shell-guard drift lint A sibling of lint-planning-prompt-drift.cjs, consuming the shared scripts/lib/drift-scan.cjs rather than copying it, wired into lint:ci. Both detectors are one shape — a fallback arm defeated by a legitimate success-on-empty: - Detector A: `--pick` and `|| echo` on one line. `--pick` is the discriminator because "missing field renders empty at exit 0" is a documented CLI contract, not a heuristic. A rule keyed on gsd_run matched 111 lines, ~132 of them legitimate, and was rejected. - Detector B: `cat <glob>` in command position, and `ls <glob>` whose exit code feeds a real fallback or an if/while head. Informational `ls <glob>` whose stdout is consumed (97 sites) and `|| true` failure suppression (~15) are not guards and never fire. Shrink-only ratchet keyed on (file, trimmed text) with a per-pair count, POSIX-normalized unconditionally so Windows CI cannot report everything fresh and stale at once. Ships with a ZERO-entry baseline: every site it can find is fixed. Exemption is the per-line `# gsd-scan-ignore: #NNN` marker whose reason must name an issue or URL; a malformed reason reports a distinct error rather than silently exempting. No file allowlists. ADR-3409 records the invariant, the measurements behind both detectors, and why the upstream `--pick` contract fix belongs to #3473. Refs #3409 * fix(#3409): resolve review findings — typed surface, sanitized reports, tighter marker Standards axis (blocker): the guard's tests asserted on human-readable stdout/stderr and on free-form baseline-load prose, which CONTRIBUTING prohibits by name. Added the typed surface it prescribes instead of weakening the tests: a frozen REASON enum, a --json report mode, structured loadBaseline errors, and a test locking Object.keys(REASON) so a new reason stays three coordinated changes. Security axis: sanitizeForReport covered every violation field but not the baseline-load error path, which embeds raw JSON.stringify output -- that escapes nothing above 0x1f, so bidi and C1 controls reached CI logs unfiltered. Routed through the sanitizer at the output seam. Security axis: the scan-ignore marker accepted `#0` and a bare `http://`. Tightened to a positive issue number and a URL with a host. This diverges deliberately from the sibling in tests/commit-files-pathspec.test.cjs, whose looser form was copied verbatim; the header now records the divergence. Security axis: G4 built its FIFO with `mktemp -u`, reserving a name without creating it. Now created inside a `mktemp -d` directory. Spec axis: ADR-3409 claimed a ninth site landed after the issue was filed. git blame disproves it -- all nine predate it; the issue's hand count missed one. Corrected. The design and test matrix still specified B9 as a FLAG after implementation reversed it to PASS; both now record the reversal and why. Refs #3409 * docs(#3409): add the how-to for resolving unreachable-guard findings Reference and Explanation are carried by ADR-3409; this is the task-oriented quadrant CI cannot check for. The page exists mainly for one thing the lint structurally cannot catch: both `[ -e "${_ARR[0]}" ]` and `[ ${#_ARR[@]} -gt 0 ]` remove the glob from the command and therefore both pass, but the count form is correct only when nullglob is set — and nullglob is usually set in a different block of the same file. A reference table cannot carry that; a how-to can. Also documents the reason codes, so a reader can tell "nothing to report" from "could not look". No tutorial: this is a gate inside an existing CI loop, not a new entry point a newcomer starts from. Refs #3409 * fix(#3409): bring the touched prompt files back under their size gates The remote run was red on 14 tests, all size/attribution, none of them the regression suite. - agents/gsd-planner.md was 194 chars over a 49152 cap enforced by four separate tests, each of which says the remedy is extraction, not a bump. It had 41 chars of headroom before this branch. Its `## Checkpoint Types` section was an unlinked, condensed duplicate of references/checkpoints.md, which already carries all three types and their XML shapes; the section now points there and keeps the three names and percentages inline. Net -969, margin 1010. - gsd-core/workflows/execute-phase.md sat 2 chars under a comfortable margin assertion. Dropped the AUTO_MODE default: the `|| echo "false"` it replaced was unreachable, so the value was already sometimes empty on next, and its only consumer compares against `true`. Net -16. Left plan-phase.md's AUTO_CHAIN default alone -- that file names an explicit `false` branch, so empty would match neither branch. - Acknowledged the seven prompt files that genuinely grew, one specific reason each. Five of those paths were already claimed by spent fragments identical to next, which blocks a second source naming the same path; removed just the colliding key from each, deleting the two that this emptied. Refs #3409 * test(#3409): extract the whole PHASE_REQ_IDS block, not just its first line G3 failed on the remote runner with '' !== 'TBD'. The test was wrong, not the workflow. The shipped contract is now two consecutive lines -- the capture and the `${PHASE_REQ_IDS:-TBD}` default -- but the helper's `^PREFIX=.*$` regex returns only the first match, so the test executed half the contract and correctly observed the empty string. Renamed to extractAssignmentBlockFor and taught it to consume the contiguous run of lines sharing the prefix. The assertion is untouched: TBD is the right expectation, and weakening it to accept the empty string would have reinstated exactly the class this suite exists to catch -- a check that cannot observe the thing it is checking. extractFencedBashAfterAnchor is unaffected: it is fence-delimited rather than line-anchored, so G1/G2/G4 still capture their full blocks. Refs #3409 * chore(#3409): drop a spent ack fragment that collided on complete-milestone.md #3458 landed on next while this branch was in flight and its fragment claims complete-milestone.md, which this branch also grows. Two ack sources may never name the same path. Its entry is spent: the +9163 it explains is already absorbed at base, so it can no longer clear anything, and the checker's own guidance for spent entries is to delete them. Removing the key emptied the fragment, so the file goes too -- an empty one signals nothing. Refs #3409 * chore(#3409): backfill changeset pr number 3558 * test(#3409): hoist a regex subject out of exec() to clear the injection scan CI's prompt-injection scan flagged `MARKER_RE.exec('# gsd-scan-ignore: ...')`. The pattern `exec[[:space:]]*\(["']` is receiver-blind on purpose, so it catches `require('child_process').exec('...')` -- and the scanner's own header records that RegExp.prototype.exec is collateral, to be handled by its allowlist. Allowlisting the file would blind it to the real exec vector permanently, so the subject is hoisted into a const instead: same assertion, scanner left at full strength, no security surface widened. Refs #3409 --------- Co-authored-by: sim <sim@local>
33 KiB
<required_reading> Read all files referenced by the invoking prompt's execution_context before starting. </required_reading>
<available_agent_types> Valid GSD subagent types (use exact names — do not fall back to 'general-purpose'):
- gsd-mempalace-curator — Ship-time MemPalace curation (diary, KG mirror, cross-project tunnels, wing-scoped prune); dispatched at ship:post when the mempalace capability is enabled. </available_agent_types>
_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
RESPONSE_LANGUAGE=$(gsd_run query config-get response_language --default "" 2>/dev/null || echo "")
INIT=$(gsd_run query init.phase-op "${PHASE_ARG}")
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi
If response_language is set: All user-facing questions, prompts, and explanations in this workflow MUST be presented in {response_language}. Technical terms, code, file paths, and subagent prompts stay in English — only user-facing output is translated.
Parse from init JSON: phase_found, phase_dir, phase_number, phase_name, padded_phase, commit_docs.
Also load config for branching strategy:
CONFIG=$(gsd_run query state.load)
Extract: branching_strategy, branch_name.
Detect base branch for PRs and merges:
BASE_BRANCH=$(gsd_run query git.base-branch)
-
Verification passed?
# The gate decides on ONE read. --pick takes a single field, so the two # human-facing fields are read only on the blocking path below — never on the # passing path — rather than issuing three queries up front (#2589). STATUS=$(gsd_run query verification.status "${PHASE_DIR}" --pick status 2>/dev/null)Only
passedmay ship. If$STATUSispassed, verification is complete — continue to the next preflight check; do not read any further verification field.Any other value (including
gaps_found,human_needed,missing, andunknown) blocks withPHASE_VERIFICATION_INCOMPLETE. Only then, read the two message fields:NEXT_ACTION=$(gsd_run query verification.status "${PHASE_DIR}" --pick next_action 2>/dev/null) NEXT_COMMAND=$(gsd_run query verification.status "${PHASE_DIR}" --pick next_command 2>/dev/null)Present
$NEXT_ACTIONto the user and, when$NEXT_COMMANDis non-empty, show it as the command to run next. These two are message text only — the block/allow decision has already been made from$STATUS, so a concurrent write between the reads cannot change the gate's verdict. The query already handles missing files and unexpected values, so no per-status arm is needed. -
Clean working tree?
git status --shortIf uncommitted changes exist: ask user to commit or stash first.
-
On correct branch?
CURRENT_BRANCH=$(git branch --show-current)If on
${BASE_BRANCH}: warn — should be on a feature branch. If branching_strategy isnone: offer to create a branch now. -
Remote configured?
git remote -v | head -2Detect
originremote. If no remote: error — can't create PR. -
ghCLI available?which gh && gh auth status 2>&1If
ghnot found or not authenticated: provide setup instructions and exit. -
Security ship gate (capability-driven).
Resolve active
ship:pregate hooks from the capability registry — the registry evaluates each hook'swhencondition, so do not readworkflow.security_enforcementdirectly:SHIP_PRE_HOOKS_JSON=$(gsd_run loop render-hooks ship:pre --raw) SECURITY_FILE=$(ls "${PHASE_DIR}"/*-SECURITY.md 2>/dev/null | head -1)Read the
activeHooksarray fromSHIP_PRE_HOOKS_JSONin-context (do NOT pipe it through a shell parser).If an active entry exists with
kind == "gate",capId == "security", andblocking == true, enforce its predicate (SECURITY.mdfrontmatterthreats_open == 0) before shipping:SECURITY_FILEis empty → block withSECURITY_SHIP_GATE_NO_REVIEW:⚠ Security enforcement is enabled but no SECURITY.md exists for this phase. Run /gsd:secure-phase {phase} and resolve findings before shipping.SECURITY_FILEexists → read its frontmatterthreats_open. The gate passes only whenthreats_openis exactly0. For any other value —threats_open> 0, or a missing / non-numeric / unparsable field — fail closed and block withSECURITY_SHIP_GATE_OPEN_THREATS(the predicate is strict equality to0; never ship on an ambiguous value):⚠ Security ship gate: SECURITY.md does not assert threats_open == 0 (found: {threats_open|unset}). Resolve open threats (or re-run /gsd:secure-phase {phase}) before shipping.
If no active security
ship:pregate hook is present (security enforcement off), skip this check silently. -
Broken-windows ship gate (capability-driven, issue #1950).
The
SHIP_PRE_HOOKS_JSONresolved in step 6 already includes anybroken-windowsgate. InspectactiveHooksfor an entry withcapId == "broken-windows"andkind == "gate":WINDOWS_GATE_ACTIVE=$(printf '%s' "$SHIP_PRE_HOOKS_JSON" | jq -r \ '.activeHooks[]? | select(.capId == "broken-windows" and .kind == "gate" and .blocking == true) | .capId' \ 2>/dev/null | head -1)If
$WINDOWS_GATE_ACTIVEis non-empty, enforce the gate by reading the ledger's typed status. The ledger lives at the project root (cross-phase, not phase-scoped):WINDOWS_STATUS_JSON=$(gsd_run windows status --raw 2>/dev/null || echo '') WINDOWS_OPEN_COUNT=$(printf '%s' "$WINDOWS_STATUS_JSON" | jq -r '.ledger.open_count // "?"' 2>/dev/null || echo '?')WINDOWS_OPEN_COUNT == "0"→ gate passes; continue to the next preflight check.WINDOWS_OPEN_COUNTis a positive integer → block withWINDOWS_SHIP_GATE_OPEN:⚠ Broken-windows ship gate: WINDOWS.md has {WINDOWS_OPEN_COUNT} open window(s). Resolve each entry before shipping, or explicitly waive with a recorded reason: gsd_run windows fixed <id> # defect resolved gsd_run windows waive <id> "<reason>" # justified deferral (reason required) Then re-run /gsd:ship.WINDOWS_OPEN_COUNTis"?", empty, or non-numeric → fail closed and block withWINDOWS_SHIP_GATE_READ_FAILED(the gate is strict equality to0; never ship on an unreadable ledger):⚠ Broken-windows ship gate: could not read open_count from .planning/WINDOWS.md. Inspect the file or run `gsd_run windows status --raw` to diagnose. The ledger may be malformed; fix it before shipping (an unparseable ledger is a broken window).
The ledger is optional and backward-compatible: on a project where
gsd_run windows statusreturnsopen_count: 0(no.planning/WINDOWS.mdyet, or an empty ledger), the gate passes silently. The gate only blocks when at least one entry isopen.If no active
broken-windowsship:pregate hook is present (gate disabled viaworkflow.windows_enforce=false, the default — tracking continues but the gate is opt-in), skip this check silently.
git push origin ${CURRENT_BRANCH} 2>&1
If push fails (e.g., no upstream): set upstream:
git push --set-upstream origin ${CURRENT_BRANCH} 2>&1
Report: "Pushed {branch} to origin ({commit_count} commits ahead of ${BASE_BRANCH})"
1. Title:
Phase {phase_number}: {phase_name}
Or for milestone: Milestone {version}: {name}
2. Summary section: Read ROADMAP.md for phase goal. Read VERIFICATION.md for verification status.
## Summary
**Phase {N}: {Name}**
**Goal:** {goal from ROADMAP.md}
**Status:** Verified ✓
{One paragraph synthesized from SUMMARY.md files — what was built}
3. Changes section: For each SUMMARY.md in the phase directory:
## Changes
### Plan {plan_id}: {plan_name}
{one_liner from SUMMARY.md frontmatter}
**Key files:**
{key-files.created and key-files.modified from SUMMARY.md frontmatter}
4. Requirements section:
## Requirements Addressed
{REQ-IDs from plan frontmatter, linked to REQUIREMENTS.md descriptions}
5. Testing section:
## Verification
- [x] Automated verification: {pass/fail from VERIFICATION.md}
- {human verification items from VERIFICATION.md, if any}
6. Decisions section:
## Key Decisions
{Decisions from STATE.md accumulated context relevant to this phase}
7. Configured project sections: Read append-only project-specific PRD/PR body sections from config:
CUSTOM_PR_SECTIONS=$(gsd_run query config-get ship.pr_body_sections --default '[]' 2>/dev/null || echo '[]')
ship.pr_body_sections is an onboarding-time extension point for teams that need extra PRD-style sections such as User Stories & Acceptance Criteria, Risks & Dependencies, Success Metrics, Release Criteria, or Stakeholder Review & Approval.
Use these sections for lean/agile PRD material that should travel with the PR without making the core /gsd:ship body configurable:
- User stories and acceptance criteria that explain the functional increment from the user's point of view.
- Definition of Done or release criteria that make the completion standard explicit.
- Risks, dependencies, stakeholder review, and traceability notes needed by regulated or approval-heavy projects.
Rules:
- Treat configured sections as append-only. They are rendered after
Key Decisionsand cannot replace, remove, or reorder the required core sections:Summary,Changes,Requirements Addressed,Verification, andKey Decisions. - Each entry must have
headingplus at least one ofsource,template, orfallback. enableddefaults totrue; whenenabledisfalse, skip the section without warning. This lets onboarding seed optional sections that a project can enable later.sourceis a fallback chain of planning artifact headings:PLAN.md ## Risks || VERIFICATION.md ## Manual Checks. Allowed artifacts areROADMAP.md,PLAN.md,SUMMARY.md,VERIFICATION.md,STATE.md,REQUIREMENTS.md, andCONTEXT.md.templateis literal Markdown with a closed token namespace only:{phase_number},{phase_name},{phase_dir},{base_branch},{padded_phase}.fallbackis literal Markdown used whensourcefinds no content and notemplateis present.- Omit sections whose final rendered body is empty after trimming.
Example configured sections:
[
{
"heading": "User Stories & Acceptance Criteria",
"enabled": true,
"source": "REQUIREMENTS.md ## User Stories || REQUIREMENTS.md ## Acceptance Criteria",
"fallback": "- Acceptance criteria are covered by the linked requirements and verification evidence."
},
{
"heading": "Risks & Dependencies",
"enabled": true,
"source": "PLAN.md ## Risks || PLAN.md ## Dependencies",
"fallback": "- No known high-risk rollout dependencies."
},
{
"heading": "Stakeholder Review & Approval",
"enabled": false,
"template": "- Product owner approval pending for {phase_name}."
}
]
8. TDD Audit section:
Reconstruct the per-commit TDD gate trail before squash-merge discards it. Walk the PR branch's own commits (merges excluded) and read each commit's gate_status: trailer with Git's native trailer machinery — never a raw %B grep, which would also match the string written in prose:
# Anchor on the merge-base so a stale local ${BASE_BRANCH} ref cannot over-count.
RANGE_BASE=$(git merge-base "${BASE_BRANCH}" HEAD)
git log "${RANGE_BASE}..HEAD" --no-merges --reverse \
--format='%H%x1f%s%x1f%(trailers:key=gate_status,valueonly,separator=%x2c)%x1e'
Records are separated by \x1e; the fields inside each are \x1f-separated — <sha>, <subject>, <gate_status value>.
Pair commits by their conventional-commit type (the type: prefix of the subject):
- A
test:commit is the RED row. Pair it with the next following implementation commit — afeat:orfix:— as its Impl commit (the GREEN step), skipping over any interveningrefactor:,docs:, orchore:commits so they are never mistaken for the GREEN step. - A
refactor:,docs:, orchore:commit that is not consumed as an Impl pairing is a standalone row with Impl commit—. - A
feat:/fix:commit with no preceding unpairedtest:is a standalone row.
Surface each commit's gate_status: value, normalized to exactly one of skill, fallback, exempt, or missing — never the raw trailer text. A commit whose trailer is absent, whose value is none of the first three, or which carries more than one gate_status: trailer (ambiguous) is counted as missing and still listed. This section is informational; it never blocks the ship.
Self-suppress when every commit is missing (#2431): the execute pipeline only writes gate_status: trailers when TDD mode is active. If every commit in the scan normalizes to missing, skip this section and the aggregate trailer (step 9) entirely — a 100%-missing table is pure noise. Only emit when at least one commit carries a real value (skill, fallback, or exempt).
Harden every table cell against injection, not just subjects: escape | as \| and strip \r/\n from both commit subjects and the rendered gate_status value. Prefer NUL (-z / %x00) record separation, and reject any record whose fields contain the \x1f/\x1e delimiters, so an adversarial commit message cannot corrupt record or field boundaries.
## TDD Audit
| Test commit | Impl commit | gate_status |
|---|---|---|
| `a1b2c3d` test: failing parser test | `e4f5g6h` feat: implement parser | skill |
| `i7j8k9l` test: failing export test | `m0n1o2p` feat: implement export | fallback |
| `q3r4s5t` refactor: extract helper | — | exempt |
Aggregate: 2 skill, 1 fallback, 1 exempt — 0 missing.
This ## TDD Audit section is the final body section — it renders after the configured pr_body_sections, immediately before the aggregate trailer — so the frozen core sections and the append-only configured sections both keep their existing order.
9. Aggregate gate_status trailer (final line) (only when step 8 was emitted — i.e., at least one real gate_status value exists):
After every other section — including any configured pr_body_sections — emit the audit aggregate as a single Git trailer on the final line of the PR body, preceded by a blank line so it parses as a valid trailer:
gate_status: skill=2, fallback=1, exempt=1, missing=0
Use the exact key order skill=, fallback=, exempt=, missing= so downstream tooling parses it stably. Keeping it last means a GitHub squash-merge that defaults its commit message to the PR description carries the aggregate into ${BASE_BRANCH}, preserving the audit footprint in git log after the PR branch is deleted. (Best-effort: it depends on the repo's squash-message default; the in-body ## TDD Audit section is the source of truth regardless.)
# BSD/macOS mktemp only randomizes XXXXXX when it is the final path component, so make a
# suffixless temp then append the extension — portable across BSD + GNU (#1520).
PR_BODY_FILE=$(mktemp "${TMPDIR:-/tmp}/gsd-pr-body-XXXXXX") && mv "$PR_BODY_FILE" "${PR_BODY_FILE}.md" && PR_BODY_FILE="${PR_BODY_FILE}.md" || exit 1
trap 'rm -f "${PR_BODY_FILE:-}"' EXIT
printf '%s\n' "${PR_BODY}" > "${PR_BODY_FILE}"
gh pr create \
--title "Phase ${PHASE_NUMBER}: ${PHASE_NAME}" \
--body-file "${PR_BODY_FILE}" \
--base "${BASE_BRANCH}"
If --draft flag was passed: add --draft.
Report: "PR #{number} created: {url}"
External code review command (automated sub-step):
Before prompting the user, check if an external review command is configured:
REVIEW_CMD=$(gsd_run query config-get workflow.code_review_command --raw 2>/dev/null || echo "")
If REVIEW_CMD is non-empty and not "null", run the external review:
-
Generate diff and stats:
DIFF=$(git diff ${BASE_BRANCH}...HEAD) DIFF_STATS=$(git diff --stat ${BASE_BRANCH}...HEAD) -
Load phase context from STATE.md:
STATE_STATUS=$(gsd_run query state.load 2>/dev/null | head -20) -
Build review prompt and pipe to command via stdin: Construct a review prompt containing the diff, diff stats, and phase context, then pipe it to the configured command:
REVIEW_PROMPT="You are reviewing a pull request.\n\nDiff stats:\n${DIFF_STATS}\n\nPhase context:\n${STATE_STATUS}\n\nFull diff:\n${DIFF}\n\nRespond with JSON: { \"verdict\": \"APPROVED\" or \"REVISE\", \"confidence\": 0-100, \"summary\": \"...\", \"issues\": [{\"severity\": \"...\", \"file\": \"...\", \"line_range\": \"...\", \"description\": \"...\", \"suggestion\": \"...\"}] }" # #2358: a per-run temp file (not a shared, unqualified path) so concurrent # ship runs — same or different phase, same or different project — never # clobber or read each other's stderr. Portable via ${TMPDIR:-/tmp}. REVIEW_STDERR_FILE=$(mktemp "${TMPDIR:-/tmp}/gsd-review-stderr-XXXXXX") REVIEW_OUTPUT=$(echo "${REVIEW_PROMPT}" | gsd_run run-with-timeout 120 -- ${REVIEW_CMD} 2>"${REVIEW_STDERR_FILE}") REVIEW_EXIT=$? -
Handle timeout (120s) and failure: If
REVIEW_EXITis non-zero or the command times out:if [ $REVIEW_EXIT -ne 0 ]; then REVIEW_STDERR=$(cat "${REVIEW_STDERR_FILE}" 2>/dev/null) echo "WARNING: External review command failed (exit ${REVIEW_EXIT}). stderr: ${REVIEW_STDERR}" echo "Continuing with manual review flow..." fi rm -f "${REVIEW_STDERR_FILE}"On failure, warn with stderr output and fall through to the manual review flow below.
-
Parse JSON result: If the command succeeded, parse the JSON output and report the verdict:
# Parse verdict and summary from REVIEW_OUTPUT JSON VERDICT=$(echo "${REVIEW_OUTPUT}" | node -e " let d=''; process.stdin.on('data',c=>d+=c); process.stdin.on('end',()=>{ try { const r=JSON.parse(d); console.log(r.verdict); } catch(e) { console.log('INVALID_JSON'); } }); ")- If
verdictis"APPROVED": report approval with confidence and summary. - If
verdictis"REVISE": report issues found, list each issue with severity, file, line_range, description, and suggestion. - If JSON is invalid (
INVALID_JSON): warn "External review returned invalid JSON" with stderr and continue.
Regardless of the external review result, fall through to the manual review options below.
- If
Manual review options:
Ask if user wants to trigger a code review:
Text mode (workflow.text_mode: true in config or --text flag): Set TEXT_MODE=true if --text is present in $ARGUMENTS OR text_mode from init JSON is true. When TEXT_MODE is active, replace every AskUserQuestion call with a plain-text numbered list and ask the user to type their choice number. This is required for non-Claude runtimes (OpenAI Codex, Gemini CLI, etc.) where AskUserQuestion is not available.
AskUserQuestion:
question: "PR created. Run a code review before merge?"
options:
- label: "Skip review"
description: "PR is ready — merge when CI passes"
- label: "Self-review"
description: "I'll review the diff in the PR myself"
- label: "Request review"
description: "Request review from a teammate"
If "Request review":
gh pr edit ${PR_NUMBER} --add-reviewer "${REVIEWER}"
If "Self-review": Report the PR URL and suggest: "Review the diff at {url}/files"
Update STATE.md to reflect the shipping action:gsd_run query state.update "Last Activity" "$(date +%Y-%m-%d)"
gsd_run query state.update "Status" "Phase ${PHASE_NUMBER} shipped — PR #${PR_NUMBER}"
If commit_docs is true, commit the ship-note AND push it onto the PR branch so
it reaches the default branch when the PR merges. Without this push the ship-note
commit stays local-only and is silently discarded when the branch is deleted on
merge (#2138). The [ci skip] trailer suppresses the redundant pipeline the push
would otherwise trigger (GitHub honors [ci skip] / [skip ci]):
gsd_run query commit "docs(${padded_phase}): ship phase ${PHASE_NUMBER} — PR #${PR_NUMBER} [ci skip]" --files .planning/STATE.md
SHIP_NOTE_SHA=$(git rev-parse HEAD)
git push origin ${CURRENT_BRANCH} 2>&1 || echo "⚠ track_shipping: ship-note push failed — it is local-only; rerun: git push origin ${CURRENT_BRANCH}"
# Preserve the skip-token optimization for repositories without a required-check
# wedge; only synthesize a second CI-triggering commit when GitHub reports one (#2783).
# Poll mergeStateStatus with backoff to avoid racing GitHub's async state computation.
# Note: Skip tokens recognized by GitHub Actions are [skip ci], [ci skip], [no ci], [skip actions], [actions skip], and skip-checks:true.
# The recovery commit message MUST NOT contain any of these tokens.
STATUS="UNKNOWN"
CHECKS=0
REVIEW_DECISION=""
for i in {1..5}; do
PR_STATE=$(gh pr view ${PR_NUMBER} --json headRefOid,mergeStateStatus,statusCheckRollup,reviewDecision -q '{head: .headRefOid, status: .mergeStateStatus, checks: ((.statusCheckRollup // []) | length), review: (.reviewDecision // "")}' 2>/dev/null || echo '{"head":"","status":"UNKNOWN","checks":0,"review":""}')
HEAD_OID=$(echo "$PR_STATE" | jq -r .head)
if [ "$HEAD_OID" = "$SHIP_NOTE_SHA" ]; then
STATUS=$(echo "$PR_STATE" | jq -r .status)
CHECKS=$(echo "$PR_STATE" | jq -r .checks)
REVIEW_DECISION=$(echo "$PR_STATE" | jq -r .review)
fi
if [ "$HEAD_OID" = "$SHIP_NOTE_SHA" ] && [ "$STATUS" != "UNKNOWN" ]; then
break
fi
sleep 3
done
if [ "$STATUS" = "BLOCKED" ] && [ "$CHECKS" = "0" ] && [ "$REVIEW_DECISION" != "REVIEW_REQUIRED" ] && [ "$REVIEW_DECISION" != "CHANGES_REQUESTED" ] && git log -1 --format=%B "$SHIP_NOTE_SHA" | grep -q '\[ci skip\]'; then
echo "⚠ PR is BLOCKED with zero checks. The [ci skip] trailer wedged the PR due to required checks."
echo "Pushing an empty commit to trigger the required pipelines..."
# gsd_run query commit requires a file list; use git directly for this intentionally empty commit.
git commit --allow-empty -m "chore: trigger CI (recover from ship-note skip-token)"
git push origin ${CURRENT_BRANCH} 2>&1 || echo "⚠ track_shipping: recovery push failed — rerun: git push origin ${CURRENT_BRANCH}"
elif [ "$STATUS" = "UNKNOWN" ]; then
echo "⚠ track_shipping: PR mergeStateStatus is UNKNOWN after polling; PR may require manual check re-trigger."
fi
Capability-driven dispatch. Resolves active
ship:posthooks via the capability registry; each hook'swhenis evaluated by the registry — no inlineconfig-get. Allship:posthooks are post-ship and additive (onError: skip); a failure here never affects the already-created PR.
SHIP_POST_HOOKS_JSON=$(gsd_run loop render-hooks ship:post --raw)
Read the activeHooks array directly from SHIP_POST_HOOKS_JSON in-context (do NOT pipe it through a shell parser).
Branch 1 — no active ship:post step hooks (activeHooks has no entry with kind == "step"): Skip silently to the report.
Generic step hook dispatch contract: For each active entry where kind == "step":
-
Honor
consumes: if it listsUAT.md, resolvels "${PHASE_DIR}"/*-UAT.md 2>/dev/null | head -1and pass it to the dispatch; if a consumed artifact is absent, skip that hook. -
If
ref.agentis set, first show the spawn banner, then dispatch the agent named byref.agent(use the exactref.agentvalue as the subagent type — e.g.gsd-mempalace-curator— nevergeneral-purpose):◆ Spawning ship:post capability agent... (runs in a subagent — no output until it returns, ~1–2 min; expected, not a freeze)
Runtime-aware dispatch (#2508 Phase 4). GSD workflows dispatch specialized subagents by role. Before dispatching on a built-in-only runtime (kimi-code — three built-ins only), resolve the role to a built-in via
gsd_run query resolve-dispatch-type --requested <role> --raw. On named-dispatch runtimes (Claude/OpenCode/…) the role is returned unchanged; on kimi-code it maps tocoder/explore/planby role-suffix. The persona rides${AGENT_SKILLS_<ROLE>}(Phase 3) regardless. See @gsd-core/references/runtime-aware-dispatch.md.
#2684 model resolution. init.phase-op emits no model field, and ref.agent is only known at runtime, so resolve it per hook before dispatching.
Input validation (defense-in-depth) — do this IN-CONTEXT, before any shell use. ref.agent originates in a capability manifest, which may be third-party. Check the value you read from activeHooks against ^[A-Za-z0-9][A-Za-z0-9._-]*$ yourself, the same way you read activeHooks itself — never by pasting it into a shell command to be tested there. A value carrying a quote, ;, `, $(, or a newline would terminate the assignment and run as its own statement before any shell-side check could execute, so a shell-side check is no protection at all.
A value that fails the check is a malformed manifest: record a warning, skip that hook entirely, and move to the next activeHooks entry. Do not dispatch it and do not place it in a command line.
Only once the value has passed, resolve its model — substituting the validated value for <agent>:
HOOK_AGENT_MODEL=$(gsd_run query resolve-model "<agent>" --raw 2>/dev/null || true)
#2517: omit the model= parameter entirely when HOOK_AGENT_MODEL is inherit or empty — a capability may name an agent absent from the model-profile table, which resolves to the empty string, and passing an empty model 404s on non-Claude runtimes. Omitting inherits the orchestrator's model.
With a resolved model ({HOOK_AGENT_MODEL} is the value the command above printed; ${…} are bound shell variables):
Agent(subagent_type=ref.agent, prompt="Ship-time capability hook for phase ${PHASE_NUMBER}. Phase dir: ${PHASE_DIR}. Consume: ${consumed_files}. Follow your agent instructions.", model="{HOOK_AGENT_MODEL}")
When it resolved to inherit or empty, drop the parameter:
Agent(subagent_type=ref.agent, prompt="Ship-time capability hook for phase ${PHASE_NUMBER}. Phase dir: ${PHASE_DIR}. Consume: ${consumed_files}. Follow your agent instructions.")
- If
ref.skillis set, dispatch withSkill(skill="gsd-${ref.skill}", args="${PHASE_NUMBER} --auto ${GSD_WS}")(prependgsd-toref.skill).
Each dispatch is best-effort: if it errors, record a warning and continue — never re-raise (onError: skip).
✓ Phase {X}: {Name} — Shipped
PR: #{number} ({url}) Branch: {branch} → ${BASE_BRANCH} Commits: {count} Verification: ✓ Passed Requirements: {N} REQ-IDs addressed
Next steps:
- Review/approve PR
- Merge when CI passes
- /gsd:complete-milestone (if last phase in milestone)
- /gsd:progress (to see what's next)
───────────────────────────────────────────────────────────────
</step>
</process>
<offer_next>
After shipping:
- /gsd:complete-milestone — if all phases in milestone are done
- /gsd:progress — see overall project state
- /gsd:execute-phase {next} — continue to next phase
</offer_next>
<success_criteria>
- [ ] Preflight checks passed (verification, clean tree, branch, remote, gh)
- [ ] Branch pushed to remote
- [ ] PR created with rich auto-generated body
- [ ] STATE.md updated with shipping status
- [ ] User knows PR number and next steps
</success_criteria>