chore: promote CHANGELOG for v1.9.0

This commit is contained in:
github-actions[bot]
2026-07-31 03:13:23 +00:00
parent 63503b2fc1
commit 7d270a205c
104 changed files with 120 additions and 521 deletions

View File

@@ -6,6 +6,126 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
## [Unreleased]
## [1.9.0] - 2026-07-31
### Added
- **New `kimi-code` runtime (Node Kimi Code CLI) registered as a distinct EoS capability** — Kimi Code users running `--kimi --global` were silently installing the Python kimi-cli agent YAMLs (which Kimi Code ignores) and getting an empty `gsd-tools query agent-skills` response. The split adds a `kimi-code` descriptor with `runtime: "node"`, `dispatch.namedDispatch: false`, `builtInSubagents: [coder, explore, plan]`, and registers it across every drift-guarded surface (allRuntimes, runtimeMap, FALLBACK_ALIASES, RUNTIME_LABELS, model-catalog, runtime-aliases manifest, capability-registry, capability-matrix, CONTEXT.md glossary). `runtimeFlags('kimi-code').isKimiCode === true`; `--kimi-code` selects kimi-code without interactive prompt; existing `kimi` (Python kimi-cli) users see no behavior change beyond the corrected `localConfigDir: ".kimi"`. (#2511) (#2519)
- **`--kimi-code --global` now installs a working Agent Skills surface at `~/.kimi-code/skills/gsd-*/SKILL.md`** — previously the kimi-code descriptor (Phase 1) carried an empty `artifactLayout`, so the install produced zero skills and Kimi Code's `merge_all_available_skills = true` auto-discovery found nothing. Phase 2 adds the `convertClaudeCommandToKimiCodeSkill` converter, fills the descriptor's `artifactLayout.global` with the skills kind entry, and removes the Phase 1 `SKIP_INSTALL_CONTRACT` skip by setting the install contract surface to `flat-skills` (NOT `kimi-skills-agents` — Kimi Code has no custom agents). Kimi Code auto-discovers the skills on next launch; no `agents/gsd.yaml` or `subagents/*.yaml` installed. (#2509) (#2520)
- **`gsd-tools query agent-skills <name>` returns the installed agent's prompt content on non-Claude runtimes** — previously, when a non-Claude runtime (kimi, kimi-code, opencode, kilo, etc.) had no explicit `agent_skills` config entry, `buildAgentSkillsBlock` returned empty and the `${AGENT_SKILLS_*}` workflow injection carried no persona. Phase 3 adds a fallback in `cmdAgentSkills`: on non-Claude runtimes, when the configured block is empty, resolve the runtime's agents directory via `checkAgentsInstalled(runtime)` and read `<agentsDir>/<agentType>.md` as the block. Gated to `runtime !== 'claude'` (Claude supports named dispatch and its `${AGENT_SKILLS_*}` contract is a skills-injection path, not a persona fallback). (#2510) (#2521)
- **Runtime-aware subagent dispatch for built-in-only runtimes (kimi-code)** — workflows calling `Agent(subagent_type="gsd-*")` now resolve the type for the current runtime via `gsd_run query resolve-dispatch-type --requested <name> --raw` before dispatch. On named-dispatch runtimes (Claude/OpenCode/…) the `gsd-*` name is returned unchanged; on built-in-only runtimes (kimi-code — three built-in subagents `coder`/`explore`/`plan`, no custom registration) it maps to the closest built-in by role-suffix heuristic (`-planner`→`plan`, `-researcher`/`-checker`/`-auditor`→`explore`, everything else→`coder`). The persona rides `${AGENT_SKILLS_<ROLE>}` (Phase 3) regardless of the resolved type. Adds the `resolveDispatchType` function to host-integration, the query to gsd-tools, a reference doc, and the resolution preamble to 26 workflow files. Pivot from the epic's original Option B (PreToolUse hook remap) after research confirmed Kimi Code's hook API supports only allow/deny, not tool_input rewriting. (#2508) (#2525)
- **The installer now distinguishes Kimi CLI (Python) from Kimi Code (Node) at install time** — running `--kimi` or `--kimi-code` prints a one-line description of each product, and if the selected variant doesn't match the detected `~/.kimi/config.toml` vs `~/.kimi-code/config.toml`, the installer warns with the correct `--kimi-code` / `--kimi` re-run command. Catches the "ran `--kimi --global` but actually on Kimi Code" mistake that produced inert YAMLs and empty agent-skills before the Phase 1 descriptor split. (#2513) (#2535)
- **New `docs/migration/kimi-to-kimi-code.md` migration guide + `built-in-only` subagent-toolkit enum value** — users who installed via `--kimi` but are actually on Kimi Code (Node CLI) now have a step-by-step migration path (re-install with `--kimi-code`, remove inert YAMLs, verify skills, verify agent-skills query). The `built-in-only` enum value replaces the `undocumented` sentinel on the kimi-code descriptor's `subagentToolkit` axis, making the descriptor self-documenting: Kimi Code's three built-in subagents (coder/explore/plan) are now a first-class negotiated value rather than an escape hatch. (#2512) (#2538)
- **`npm run regen:derived` regenerates every derived artifact in one command** — replacing several separate invocations (`build`, `gen:registry`, `gen-adr-index`, `gen-capability-matrix`, `gen-inventory-manifest`, `sync-manifest-versions`, `gen:install-tree`) with one dependency-ordered command. (#2721) (#2730)
- **`/gsd:review --kimi-code` reviews your plans with Kimi Code CLI** — the new lane joins the cross-AI reviewer roster and is included by `--all` when detected. Detection distinguishes Kimi Code from the legacy Python kimi-cli, which ships a binary of the same name, so a host with only the legacy tool reports the lane unavailable instead of registering a reviewer that cannot serve it. (#2718) (#2861)
- **`/gsd:update` now offers to restore the user-added files it backs up** — files you added inside GSD-managed directories were copied to `gsd-user-files-backup/` before the clean install and then left there forever; only `--reapply` (a different bucket, `gsd-local-patches/`) had a restore path. The update now lists what it backed up, runs a compatibility pass against the newly installed release, and offers to put the files back. Declining leaves the backup untouched, and the backup is never deleted. (#1854) (#2679)
- **Phase effort estimation against a calibrated smart-zone budget** — plans can now be sized against a configurable token budget (`workflow.smart_zone_tokens`, default 100000) instead of a static heuristic, and the estimate self-corrects against measured reality. Adds the `estimate-check` and `estimate-calibration` query verbs. (#2630) (#2661)
- **Bracket phase-ID core grammar lands behind an opt-in flag** — `parsePhaseId`/`renderPhaseId`/`toDir` add one pure round-trippable `PhaseId` model inside the ADR-2121 canonical owner (`src/phase-id.cts`), gated on `phase_id_convention: 'bracket'`, with generative round-trip properties; legacy `null`/`milestone-prefixed` paths stay byte-untouched (epic #612 PR-1). (#2249) (#2258)
- **Reviewer CLIs now honor GSD's configured reasoning effort instead of silently inheriting your global CLI default** — cross-AI review runs previously picked up whatever `effort` sat in your own `~/.codex`/Claude/OpenCode config, so the same project produced 1-3 minute review cycles on one machine and 12-15+ minute cycles on another with no in-project way to influence it. GSD now resolves one effort value from the `effort.*` cascade and passes it to each reviewer in that CLI's own syntax; a host with no documented reasoning setting is left untouched rather than given a guessed flag. (#2481) (#2490)
- **List the `gsd-cursor` EoS host integration in the registry** — six phase-aware Cursor profiles (max / hybrid / value / budget / frontier / openweight), added to `docs/registries/eos.json` with a versioned v1.1.0 install command. (#2581)
- **Plans now carry a calibrated effort estimate** — every generated PLAN.md includes an `estimate` block, and `/gsd-plan-phase` flags a phase projected to exceed the smart-zone budget with a concrete split recommendation. Advisory only; it never blocks planning. (#2631) (#2670)
- **Reviewer lanes can be declared as capability manifest data** — a capability may now carry a `reviewer` body describing a cross-AI review lane (slug, flags, transport, probe, invocation shape, timeout floor, output policy), and a new `role: "reviewer"` declares a lane that is not an install target. The registry validates the body against closed vocabularies and enforces slug, flag, and section uniqueness across first-party and installed capabilities, so two lanes can no longer silently share a REVIEWS.md heading. A capability with no reviewer body is unaffected. (#2795) (#2823)
- **Reviewer lanes now ship as capability declarations** — the eleven cross-AI reviewer lanes are declared as manifest data instead of a half-derived, half-hardcoded roster. Five reviewers GSD never installs into (Gemini, CodeRabbit, Ollama, LM Studio, llama.cpp) become lane-only capabilities with no install surface, and the six hosts that are also reviewers gain a reviewer body alongside their runtime descriptor. `gsd capability list` shows the five new lanes. The roster itself is unchanged — the same eleven reviewers, derived rather than hardcoded — and `runtime.hostBehaviors.reviewerCli` keeps working for one release. (#2798) (#2837)
- **Parallel execute-phase waves now run on Codex, OpenCode, Kimi and Kimi Code** — previously only Claude Code could execute a wave's independent plans concurrently, because worktree isolation relied on its harness-native `isolation="worktree"` primitive and every other runtime failed closed to sequential. Executor isolation is now a negotiated capability: runtimes whose harness isolates executors (Claude Code, Cursor) use their own flag, and runtimes exposing a headless exec with a working directory (Codex, OpenCode, Kimi, Kimi Code) get worktrees that GSD creates, validates and merges itself. Runtimes with no isolation primitive still run sequentially, and an unknown declaration always degrades to sequential rather than to an unisolated parallel run. (#2627) (#2635)
- **Codex host-plugin binding + negotiated executor-worktree isolation** — ADR-1239 gains a Codex worked-binding amendment and a new `dispatch.isolation` capability (harness- vs orchestrator-managed git worktrees) enabling parallel execute-phase waves on non-Claude runtimes. (#2600) (#2600)
- **Estimates now calibrate against reality** — the executor records what a phase actually cost into SUMMARY.md, and `/gsd:extract-learnings` computes the estimate-vs-actual correction so future plan estimates improve for your project. (#2632) (#2672)
### Changed
- **The phase researcher must now read and cite in-repo values before calling them verified** — an enum, schema or type union, error code, status constant, or filesystem path earns a `[VERIFIED: path:line-range]` tag only if the researcher opened the source-of-truth file with `Read` during the run and quoted the values verbatim in the `<interfaces>` block; every value used in a code skeleton must appear in that quote, and anything else stays `[ASSUMED]`. Previously the tag could be earned from training memory or a web search alone, so a plausible-but-drifted enum could pass into RESEARCH.md, get copied into PLAN.md, and fail only at the executor's `parse()`/typecheck — a mid-execution deviation, the most expensive place to discover it. (#1699) (#2768)
- **Completing a phase now warns when its SUMMARY claims files that never landed** — `phase complete` runs the artifact check that `verify-summary` has always applied to the research SUMMARY against the completing phase's own `SUMMARY.md` files, and reports any referenced path that is not on disk through its existing `warnings[]` channel. Previously the check was wired to exactly two call sites, both pointed at `.planning/research/SUMMARY.md`, so the summaries that actually assert "I created these files" were never verified and an interrupted phase counted toward 100% silently. Advisory only: it never blocks completion. Paths are recovered heuristically from the SUMMARY body, so globs, URLs, bare hostnames, and paths resolving outside the project are skipped rather than reported; the `key-files:` frontmatter block and commit hashes are deliberately not read. (#2572) (#2685)
- **The UI consideration probe now asks about loading and error states for interactive controls** — a UI surface classified only as an interactive control (a button, toggle, switch, or slider, with no accompanying form or list) previously had only its long-text state probed, so a spec could omit what the control shows while its action is in flight or when it fails and still pass. Control-only surfaces are now probed for their in-flight and failure states too. (#2151) (#2575)
- **Reviewer lane flags and section titles are now gated across every documentation surface** — `/gsd:review` reviewer flags were hand-enumerated in five docs and three workflow files that had silently drifted apart: `--kimi-code` was missing from all four translated `COMMANDS.md` mirrors, `--coderabbit` from every workflow forwarding list, and `--antigravity` from `FEATURES.md` entirely. The lane roster is now the single declared source: workflows derive their flag lists from a new `review-lane flags` query, and a parity gate fails the build when any documented flag or reviewer section title diverges from it. The capability manifest reference also gains the previously undocumented `reviewer` body and `hostBehaviors` field. (#2800) (#2882)
- **Cross-AI reviewer lanes are now declared data rather than hand-written per-CLI blocks** — every lane's binary, prompt and output channel, timeout, probe and empty-output policy comes from its capability manifest, so a reviewer can be shipped as a plugin instead of a core patch. Two user-visible consequences: a reviewer that returns only whitespace is now reported as a failed lane on every reviewer (previously only on LM Studio and llama.cpp, so elsewhere a blank reply was rendered as a clean review), and an OpenAI-compatible lane whose configured host has changed since you consented to it is blocked with an explanation rather than silently sending your plans to the new destination. `jq`, `curl` and GNU `timeout` are no longer required on PATH for any lane. (#2782) (#2861)
- **Reviewer config keys are now owned by their reviewer-lane capabilities** — `review.models.<lane>`, `review.<lane>_host` and `review.max_prompt_tokens_per_reviewer.<lane>` moved from the central config schema to federated slices on the lanes that use them. Key names and existing `.planning/config.json` files are unchanged and no migration is needed. Two consequences are user-visible: a `review.models.<x>` or `review.max_prompt_tokens_per_reviewer.<x>` key naming something that is not a declared lane is now rejected by `config-set` where it was previously accepted and silently ignored; and clearing one of these keys now reads back as its declared default rather than reporting key-not-found, because a federated key always resolves — an empty string for the model and host keys, and `-1` for a per-lane token budget (a deliberate sentinel, since `0` already means "do not trim this lane" and must stay distinguishable from unset). `review.max_prompt_tokens`, `review.default_reviewers` and `review.reviewer_instances` describe policy across lanes and deliberately remain central. (#2797) (#2841)
- **Reviewer lanes are disclosed and consent-gated before install** — a capability that declares a reviewer lane now discloses what it will run and what it will be sent, and blocks on consent before any file is promoted. A spawned lane discloses its binary and its full arguments; an OpenAI-compatible lane discloses its destination host and the config key naming it, including a localhost destination. Both name the egress payload classes — plan text, requirements, research findings, and CONTEXT.md decisions. Changing a lane's binary, arguments, destination, prompt channel, or handler forces re-consent on update; a capability with no reviewer lane is unaffected and its consent record is unchanged. (#2796) (#2826)
- **Raw and calibrated phase-estimate token counts are now distinct types** — the two states of an estimate (the planner's uncorrected projection and the same figure with the project's calibration factor applied) could previously be swapped at any seam without complaint, because both are plain positive integers. That produced two shipped defects in epic #1952: a doubly-applied correction (factor squared) and a calibration loop that measured against its own output and never converged. Both are now compile errors. No behavior, output, or schema change. (#2671) (#2676)
- **The emitted-attribution size ratchet now tells you how to clear it** — a PR that only grew a workflow or agent file used to fail with a byte delta and the word "acknowledgment", without naming `tests/emitted-drift-ack.json`, saying it does not exist yet, giving its schema, or stating that the key is the bare filename. All three failing branches now print a minimal valid document and repeat that nothing is regenerated. (#2778) (#2780)
### Removed
- **`npm run gen:golden`, `UPDATE_GOLDEN`, `npm run size:baseline`, and `npm run setup:merge-driver` are removed** — the committed golden-install-parity fixtures and the two per-file size baselines they regenerated are deleted. The differential attribution check (`tests/emitted-attribution.test.cjs`) is now the sole gate for both emitted-content propagation and workflow/agent size growth; editing shipped content requires zero manual fixture regeneration. `npm run regen:derived` and `npm run gen:install-tree` are unaffected. (#2724) (#2767)
### Fixed
- **Permission errors on phase and milestone directories now surface instead of looking empty** — an unreadable phase directory used to be silently reported as "no CONTEXT.md" (so the discuss/plan gates wrongly skipped context) and an unreadable `milestones/` directory as "no archives" (so active-milestone resolution and archived-phase filtering misbehaved), because both scans treated a permission or I-O failure the same as a genuinely empty directory. (#1883) (#2802)
- **Worktree branch guards now accept Claude Code's `agent-<id>` namespace** — the `worktree record-agent` command, the spawn-time branch check, the cleanup-wave manifest reader, and the force-add/path/workflow guards all accept both the current `agent-<id>` and the legacy `worktree-agent-<id>` branch naming. Previously, Claude Code's rename from `worktree-agent-<id>` to `agent-<id>` caused every executor sub-agent to fail its branch check (false-positive FATAL / exit 42) and silently dropped valid cleanup-manifest entries (`empty_manifest`), blocking merge-back. (#1995) (#2548)
- **`secure-phase`, `validate-phase`, and `next` workflows now scope their `query commit` calls** — all three pass `--files` with the specific artifact path, preventing the blanket `git add .planning/` default branch from sweeping unrelated staged or unstaged files into a commit whose message describes a single artifact. Previously, these three call sites (out of 65 total) were the only ones omitting `--files`, causing #2112's commit-scoping fix to never reach them. (#2269) (#2549)
- **`/gsd-map-codebase` Update mode now refreshes all date stamps** — the `**Analysis Date:**` line, the `*... analysis: ...*` footer, and the `<!-- refreshed: ... -->` header are set to the current date on every run, overwriting any prior date. Previously, Update runs only replaced `[YYYY-MM-DD]` placeholder tokens, which don't exist in already-generated files (they contain concrete dates from the prior run), so stamps silently retained the original mapping date. (#2279) (#2550)
- **All seven guard hooks now normalize Kimi's payload shape** — the five JS guards (`gsd-prompt-guard`, `gsd-read-guard`, `gsd-worktree-path-guard`, `gsd-read-injection-scanner`, `gsd-workflow-guard`) and the two shell hooks (`gsd-graphify-update.sh`, `gsd-phase-boundary.sh`) normalize Kimi's native payload shape before their checks: the tool name (`WriteFile` → `Write`, `StrReplaceFile` → `Edit`, `ReadFile` → `Read`, `Shell` → `Bash`, bare or module-qualified), the tool-input fields (`path` → `file_path`, `edit.old`/`edit.new` — single or list — → `old_string`/`new_string`), and the PostToolUse `tool_output` field → `tool_response`, matching kimi-cli's actual tool and hook-event schemas. The two blocking guards (worktree path and workflow) also write their block reason to stderr, which is what Kimi feeds back to the model on exit 2. Previously the Kimi `[[hooks]]` matcher was translated to Kimi's vocabulary but the scripts' payload checks were not, leaving every guard — including the prompt-injection read scanner — dormant on Kimi while appearing registered. (#2304) (#2518)
- **`parseCoverageMatrix` now scopes table parsing to recognized coverage matrices** — pipe-tables outside the matrix (e.g., summary tables) are ignored instead of being silently parsed as data rows, multi-section matrices with repeated headers are supported, and inline markdown emphasis (`**OPT-OUT**`) on decision cells is stripped before validation. Previously, the parser scanned every `|`-prefixed line file-wide with a latching header flag, causing silent phantom-capability corruption from unrelated tables, false rejection of multi-section matrices, and rejection of bold-emphasized decisions. (#2366) (#2551)
- **`state.planned-phase` now warns on no-op transitions and syncs `progress.total_plans`** — when STATE.md's Current Position has no recognized labels (narrative prose), the command emits a `warning` field so the workflow can detect the no-op instead of continuing with stale state. When a plan count is provided, `progress.total_plans` in the YAML frontmatter is updated alongside the body `Total Plans in Phase` field, preventing contradictory state between the two representations. Previously, the command silently returned success with an empty `updated` array and zero bytes written, and left `progress.total_plans` at 0 while the body reported the actual count. (#2400) (#2552)
- **Codex `--local` installation no longer writes skills to `$HOME/.agents/skills`** — the skills-kind `home` override (which redirects skills to the user-global `.agents` directory) is now only applied for `--global` scope. When `--local` is specified, skills are installed under the project-local config directory, matching the scope the user selected. Previously, a `--local` Codex install created a split installation: project-local config but user-global skills. (#2429) (#2553)
- **`use_worktrees: false` is now honored at the worktree dispatch gate** — the per-plan dispatch condition checks BOTH the project-level `USE_WORKTREES` flag AND the per-plan `USE_WORKTREES_FOR_PLAN` variable. Previously, the dispatch gate checked only the per-plan variable (derived from submodule intersection), so plans that didn't touch submodules would still fork `isolation="worktree"` agents even when the project-level setting disabled worktrees entirely. The fix is net-negative in file size (prose compression offsets the added shell condition). (#2474) (#2561)
- **The Gemini and Claude reviewer legs now fail loudly instead of silently dropping out of the cross-AI review** — both blocks capture stderr to a `.err` sidecar instead of discarding it to `/dev/null`, and write a diagnostic stub with the captured error when the lane produces no output. Previously they were the only two of the ten prompt-fed reviewer legs with neither guard, so any failure that wrote no stdout (CLI missing, unauthenticated, rate-limited, crashed) left a zero-byte review file that `write_reviews` rendered as a reviewer that had run cleanly with nothing to report — quietly degrading an N-reviewer consensus to N-1 while `present_results` reported success. The guard matches the shape the Codex and Cursor legs already use. (#2494) (#2592)
- **`gsd-ui-auditor` no longer documents an uncallable Playwright-MCP capture path** — the agent's `tools:` allowlist grants no MCP namespace, so the `<playwright_mcp_approach>` block it presented as "preferred" could never dispatch: the availability check had a fixed answer, the three `mcp__playwright__*` calls were unreachable, and the CLI fallback was the only branch that ever ran. The dead block is removed, leaving the CLI screenshot path as the sole documented approach, and a new consistency test fails any `agents/*.md` that documents an `mcp__<server>__*` namespace its own `tools:` line withholds. Session-level Playwright-MCP capture in `/gsd-ui-review` is unaffected — that path is genuinely runtime-detected. The same documented-vs-granted drift is corrected one layer out in `docs/AGENTS.md`, where 26 of 34 per-agent **Tools** rows disagreed with the agent's frontmatter — 22 omitting `Skill`, 7 omitting `Edit`, 8 omitting MCP grants entirely (7 of them abbreviating up to eight distinct servers as "mcp (context7)"), and one still naming `Task`, a tool that no longer exists — with a parity guard added so the role cards and the frontmatter cannot drift apart again. (#2526) (#2594)
- **`query commit --files` no longer silently checks out the wrong phase branch mid-commit** — the phase-token extraction is now anchored to the directory segment under `.planning/phases/` and reuses the project-code-aware `extractPhaseToken` helper instead of an unanchored regex, so a `project_code` ending in a digit (e.g. `PROJECT_V2`) no longer makes `…/PROJECT_V2-07-name/…` match the `2-` inside `V2-` and resolve to the wrong phase. The commit-path branch auto-switch also no longer silently force-switches an already-checked-out working branch onto a different existing phase branch (it creates-if-absent only, per the original `#1278` intent); the only prior trace of the silent switch was a `git reflog` entry. (#2539) (#2669)
- **Reviewer/workflow config lookups no longer silently drop the configured value on machines without `jq`** — `review.md`, `plan-phase.md`, `ship.md`, `debug.md`, `autonomous.md`, `ai-integration-phase.md`, and `eval-review.md` now resolve `config-get` scalars with the native `--raw` flag and `resolve-model` / `resolve-execution` / `verification.status` object fields with `--pick`, instead of piping through `jq`. Previously, on a stock Windows/Git-Bash box with no `jq` on PATH, the `… | jq …` stage failed (exit 127), the failure was swallowed by `2>/dev/null || <default>`, and the configured per-lane model/host/budget came back empty — so the lane fell back to CLI defaults (e.g. `~/.codex/config.toml` instead of the configured `review.models.codex`) with no diagnostic, and the `autonomous.md` verify gate could misroute on an empty status. The legitimate structured-JSON `jq` sites that parse HTTP `curl` responses (`.choices[0]`, `jq -rs`, `jq -n --rawfile`) are untouched — only the jq-replaceable config/model/verify lookups moved to the native flags. Because those sites remain, `/gsd-review` now probes for `jq` up front and reports the `ollama`, `lm_studio`, `llama_cpp`, `opencode`, and `antigravity` lanes as unavailable with an install hint when it is missing, instead of running them into empty output; the `gemini`, `claude`, `codex`, `coderabbit`, `qwen`, and `cursor` lanes stay selectable with no `jq` installed. (#2589) (#2673)
- **Upgrading a Claude-global GSD install now uses the new version's skill content instead of the previous version's** — the installer read a `.gsd-source` marker that still pointed at the prior install's source location before rewriting it, so on an upgrade every converted skill was generated from the old version's command definitions (while the file manifest faithfully recorded the stale content's hash as correct). The marker is now written before anything reads it. (#2624) (#2811)
- **`phase complete` and `state begin-phase` no longer rewrite `current_phase_name` to the name's own parenthetical** — transitions that already hold the exact display name now pass it to `syncStateFrontmatter` as an authoritative override, so the lossy body-prose re-derivation never runs the final word on a field the transition just resolved. Previously, completing into a phase named `Closer-ruling measurement (D1a)` wrote `current_phase_name: D1a` (the prose parser's paren-over-dash preference harvested the name's own parenthetical), and every downstream consumer of the scalar inherited the mangled name. `parsePhaseFromProse` also gains status-keyword-aware precedence (the #1695 AC #3 residual) for genuinely unknown prose: the em-dash name wins when it is not a status keyword or `Milestone:` tail, so `48 — Closer-ruling measurement (D1a)` now parses to `Closer-ruling measurement` instead of `D1a`. (#2736) (#2821)
- **Seven dangling references in the ADR corpus and contributor docs now resolve** — (1) `docs/adr/1239-gsd-embeddable-orchestration-engine.md` linked the host-integration capability matrix as `reference/…` from inside `docs/adr/`, resolving to the nonexistent `docs/adr/reference/`; all three occurrences now use `../reference/…`, and the two whose link text promises `§codex` now carry the matching `#codex` fragment. (2) `src/plan-drift-guard.cts` cited `docs/adr/0022-source-grounding-drift-guard.md`, a path that has never existed — corrected to the real `docs/adr/22-plan-drift-guard.md`; because the file is compiled into the shipped payload, the bad citation was shipping to users. (3) `CONTRIBUTING.md` and `docs/contributor-standards.md` illustrated the ADR naming convention with issue `#3485`, a pre-rename number from `get-shit-done-redux` that does not resolve in `open-gsd/gsd-core` — the worked example now uses `#2264`, which does, and the one genuinely historical `#3485` reference is annotated rather than rewritten. (4) `docs/adr/857-capability-system.md`'s H1 still carried a `[Proposed]` status bracket contradicting its `Accepted — ratified 2026-07-17` Status field; the ADR index generator strips the bracket for display, so the contradiction was invisible to the gate. (5) `scripts/gen-adr-index.cjs`'s back-link comment still described ADR-857 as `Proposed` and its claim over ADR-0011/ADR-58 as a supersession — both restated at the 2026-07-17 ratification, when the claim became `Subsumes` and the reciprocal back-links were added. (6) `docs/how-to/install-on-your-runtime.md` linked that same capability matrix as a bare `host-integration-capability-matrix.md` from inside `docs/how-to/` in its ZCode and pi sections — the identical defect as (1), so both now use `../reference/…`. (7) `docs/CONFIGURATION.md` cited ADR-1244 as `adr/1244-runtime-capability-registry-overlay.md`; the file is `adr/1244-capability-ecosystem.md`. (#2691) (#2692)
- **`roadmap get-phase` no longer drops success criteria that wrap onto a second line** — the parser broke the criteria run at any indented continuation line, truncating the wrapped criterion (losing its trailing `[REQ-ID]` tag) and silently dropping every criterion below it. `verify-work` and `plan-phase` consumed the shortened list, so a phase could be planned and certified complete against a strict subset of its own success criteria with nothing reporting the gap. Continuation lines now fold into their criterion; blank-line-separated criteria still parse. (#2522) (#2637)
- **The host-integration capability matrix now documents the `kimi-code` runtime** — kimi-code shipped as a distinct runtime but its section was never added, so its `hostIntegration` axes had no cited source. Sourcing each axis against Kimi Code CLI's own docs also corrected three values that had been inherited from the unrelated Python `kimi` CLI: `embeddingMode` is `declarative` (plugins are a manifest plus markdown Skills, with no in-process API), `dispatch.nested` is `true` (the `coder` built-in dispatches nested sub-agents), and `dispatch.maxDepth` is `undocumented` (no depth bound is published). (#2603) (#2687)
- **`/gsd-profile-user` now writes the runtime-native instruction file on Codex and other AGENTS-native runtimes** — `generate-claude-profile` hardcoded `.claude/CLAUDE.md` for both project and global scope, ignoring the runtime-aware resolution that #3163 wired into the sibling `generate-claude-md` handler. The #3163 fix diverged when it didn't propagate here, so running `$gsd-profile-user --refresh` on a Codex install created/modified Claude configuration instead of producing a Codex `AGENTS.md` profile. The command now resolves its target through the shared runtime policy: project scope uses `getProjectInstructionFile(runtime)` and global scope derives `~/.<config-home>/<instruction-basename>`, so codex lands at `~/.codex/AGENTS.md`. Claude behaviour is preserved. A parity test guards against future re-divergence between the two handlers. (#2659) (#2659)
- **The plan-phase decision-coverage gate can no longer silently pass when its context-path argument is missing** — the handler now fails closed on an empty/missing argument (a caller error), and the plan-phase workflow recomputes the CONTEXT.md path in the same Bash block that runs the gate (the variable set in the init block did not survive into the gate block). A genuinely-absent CONTEXT.md still produces the legitimate green skip. Previously the gate reported `passed` without ever checking coverage. (#2770) (#2881)
- **`/gsd-code-review` no longer silently drops CRLF-saved artifacts** — the Tier-2 file-scope extractor (and every REVIEW/REVIEW-FIX frontmatter reader in the code-review and code-review-fix workflows) used a literal `\n` to find the YAML block, so any SUMMARY.md/REVIEW.md saved with CRLF line endings (default on Windows) contributed zero files with no warning. The boundary now normalizes CRLF first, so a mixed CRLF/LF phase reviews the union of its files instead of an incomplete set. (#2694) (#2839)
- **Dev-dependency `brace-expansion` bumped to patched versions (1.1.18 / 5.0.9), resolving the high-severity DoS/OOM advisories** — the lockfile now pins the 2026-07-30 patch backports reachable via eslint and stryker. A non-breaking in-range bump (no overrides, no major bumps); production `npm audit --omit=dev` is unaffected (devDependency only). (#2765) (#2888)
- **The markdown-parsing lint rule now catches the stricter cell-regex spelling it previously missed** — a hand-rolled table scan written as `[^|\n]` (excluding both the pipe and the newline, which is the more correct form) slipped past the guard entirely, so `STATE.md` field replacement kept parsing tables with a local regex and rewriting the whole document. The rule now flags any pipe-excluding character class, and the STATE.md field writer edits a bounded byte range instead. (#2880) (#2889)
- **Workstream-scoped config reads now inherit from the project root config** — `config-get` under an active workstream (`GSD_WORKSTREAM`) now resolves a key absent from the workstream's own config to the project-root value before falling back to schema defaults, instead of reporting 'Key not found'. A workstream config still overrides root for any key it sets; root only fills gaps. Previously a key set only at root was silently lost under a workstream, causing shipped workflow boolean guards (e.g. use_worktrees, plan_review_convergence) to apply their hardcoded fallback and silently invert the user's setting. (#2833)
- **The Claude-orchestration Workflow backend can now actually dispatch a wave** — every script `emitWorkflowScript` generated was rejected by the Workflow tool. It omitted the required `export const meta = {…}` first statement (fatal on its own), called `resumeFromRunId()` and `budget()` which are a tool input parameter and a read-only object rather than script functions, and passed `parallel(agent(…), agent(…))` where an array of thunks is required. Two further defects meant the script was never even reached: nothing resolved the Agent SDK version, so the gate ladder returned `agent_sdk_version_unknown` on every automated run while `capability state` still reported the capability active; and the runtime fallback diverged from the canonical `GSD_RUNTIME > config.runtime > 'claude'` chain, so any invocation without `--runtime` reported `runtime_not_claude`. The router now resolves the installed SDK version itself and defers to the canonical runtime resolver, and the emitted script is valid ES module syntax with `phase()` titles matching `meta.phases`. (#2590) (#2681)
- **Releases no longer fail their own emitted-parity gate** — cutting any release ran the differential attribution check against a baseline built at a different version, so the install-time hook version stamp made all 364 emitted hook paths look like unexplained drift and every `finalize`/`rc` run hard-failed before tagging or publishing. (#2891) (#2894)
- **Merging an emitted-drift acknowledgment no longer turns the mainline red.** An acknowledgment is now scoped to the diff that introduced it, so once its ripple is absorbed into the base it goes inert instead of reporting as stale — which had reddened `next` for five consecutive commits and every pull request branching off it. (#2789) (#2803)
- **Discuss-phase advisor mode now spawns the registered `gsd-advisor-researcher` subagent instead of `general-purpose`** — resolving a contradiction with the universal-anti-patterns rule (injected into the same context) that forbids non-GSD agent types. The manual "read the agent def" prompt line is dropped (spawning by type auto-loads it). (#2771; the sibling assumptions-site needs a design decision — filed as #2883) (#2886)
- **Subagent spawns no longer fail on non-Claude runtimes when no model resolves** — 15 workflows told the orchestrator to pass a model parameter without saying to drop it when nothing resolved, so 43 dispatch sites sent an empty model and the spawn 404'd. That was the default state on Codex, OpenCode, Gemini CLI, Kilo, Qwen and Hermes, where GSD sets `resolve_model_ids: "omit"` on install. Every dispatching workflow now carries the rule. (#2711) (#2713)
- **The statusline now renders GSD state correctly on Windows-authored (CRLF) STATE.md** — `parseStateMd` no longer drops the entire frontmatter block on CRLF input. The fence regex and downstream splits now accept CRLF line endings, matching the canonical `extractFrontmatter` parser. Previously a CRLF STATE.md silently produced an empty GSD-state segment (no status, phase, or milestone) with no error. (#2754) (#2865)
- **`api-coverage` now ships the #2366 coverage-matrix fix** — the tracked `gsd-core/bin/lib/api-coverage.cjs` build artifact had drifted four days behind `src/api-coverage.cts`, so the module that actually ships still parsed non-coverage tables as data, mishandled multi-section matrices with repeated headers, and failed to parse `**OPT-OUT**`. Regenerated, plus a new `lint:generated-sync` check that fails when any tracked compiled artifact no longer matches its source. Also prunes two stale entries from the `no-phantom-issue-refs` guard: GitHub numbers issues and PRs from one shared counter, so both had since become real merged PRs, and the guard was rejecting accurate citations of them. (#2653) (#2656)
- **OpenCode/Kilo no longer spawn the context-monitor subprocess on every tool call when context warnings are disabled** — the adapter now reads the existing `hooks.context_warnings` toggle in-process and skips the child-process spawn entirely when it is set to `false`, instead of paying a Node boot per tool call only to read the flag and exit inside the child. Behavior is unchanged when the toggle is absent or enabled (the default). (#2824)
- **Editing `src/` no longer trips an undocumented changeset-lint failure** — CONTRIBUTING.md listed the Changeset Required triggers without `src/`, the path that compiles into every `gsd-core/bin/lib/*.cjs`, so contributors touching it hit a CI failure the docs said could not happen — and a local run of the lint reported success regardless, because it silently requires `GITHUB_BASE_REF` to see the branch at all. Both are now documented, and the config-loader test-helper that reset only one of its two warning-dedup sets now resets both. (#2674) (#2678)
- **The .planning/ write reminder can no longer be suppressed or fabricated by a model-supplied file_path** — the phase-boundary hook now treats `tool_input.path` (the field kimi-cli actually executes on) as authoritative and `file_path` as the fallback, reaching the same "path authoritative" outcome the JS guards establish via upstream normalization (#2595). Previously a model-controlled decoy `file_path` could silence the reminder for a genuine .planning/ write or raise one naming a file never touched. (#2752) (#2860)
- **test:/chore:/ci:/docs:/refactor:/perf:/revert: PRs no longer publish under the user-facing Enhancement heading in release notes** — the release-notes classifier now routes recognized non-user-facing conventional-commit types to an Internal bucket and omits them from the published GitHub release notes (and the Discord announcement's user-facing sections). Previously these internal-work PRs rendered as Enhancements alongside genuinely user-facing changes. feat:/fix: classification is unchanged, and untyped or anchor-defeated titles still fall back to Enhancement. (#2838)
- **`/gsd-execute-phase` and `/gsd-quick` branches no longer auto-track `origin/master`** — the branch-creation `git checkout -b <branch> origin/$DEFAULT_BRANCH` omitted `--no-track`, so with the default `branch.autoSetupMerge=true` git wired the new branch's upstream to `refs/heads/$DEFAULT_BRANCH`. A subsequent GUI sync (GitHub Desktop, VS Code) then pushed the branch's commits straight onto `origin/$DEFAULT_BRANCH`, bypassing PR review — in one project every commit of a 7-plan phase landed on `origin/master`. `--no-track` is now passed; the first `git push -u origin <branch>` sets up correct same-name tracking. (#2498) (#2628)
- **Cursor CLI sessions now detect `.planning/`** — the `sessionStart` and `stop` hooks resolved the project from `process.cwd()`, which under the `cursor-agent` CLI is the Cursor config dir (`~/.cursor`), not the workspace. Every CLI session therefore reported "no .planning/ workflow found" even with `.planning/STATE.md` present, and the stop hook's verify-work reminder could never fire. Both hooks now read `workspace_roots` from the hook payload they already buffered but never parsed, preferring the root that actually carries `.planning/STATE.md` (multi-root workspaces) and falling back to the first root, then `cwd` so IDE invocations are unchanged. (#2587) (#2680)
- **Plan, summary, verification, and state validators now reject NUL-corrupted files** — `frontmatter validate`, `verify plan-structure`, and `state validate` now fail loud (valid:false) when a file contains embedded NUL bytes, with an error naming the encoding problem and its downstream consequence. Previously such a file passed as valid:true but was silently skipped by recursive/binary-skipping search tools (rg, grep -I), reading downstream as 'file absent' rather than 'file corrupt.' (#2829)
- **OpenCode no longer declares background subagent dispatch it does not have** — `capabilities/opencode/capability.json` advertised `dispatch.background` and `dispatch.backgroundDispatch` as `true`, but OpenCode's native subagent dispatch is synchronous: the Task tool's `background` parameter is hidden from the model behind the opt-in `OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS` flag, which defaults to false, and the session loop still handles one subtask at a time. Since `negotiateHostCapabilities` and every `degradationFor` consumer trusts these per-field values, declaring an absent capability overstated it — the opposite of the fail-closed posture the negotiation exists to enforce. Both fields are now `false`, and the host-integration capability matrix carries the corrected values with current upstream citations. (#2598) (#2682)
- **Debug sessions now commit their session docs** — with `commit_docs: true`, finishing a `/gsd:debug` session left the session doc (and sometimes the fix's own code changes) sitting untracked in the working tree. The session manager, which owns the end of a debug session, never had a commit step — only the single-spawn debugger path did. Terminal sessions now commit the doc and any uncommitted in-session fix code, still respecting `commit_docs`; sessions that pause mid-investigation deliberately do not. (#2568) (#2731)
- **Codex installs now ship the complete update-check hook set** — the `--codex` installer (both `--profile=core` and `--profile=full`) now installs and refreshes all four hook files the update-check/context-monitor feature needs (`gsd-check-update.js`, `gsd-check-update-worker.js`, `managed-hooks-registry.cjs`, `gsd-context-monitor.js`) together, instead of only the two parent scripts. Previously a registered parent hook pointed at a worker and registry the same installer never delivered. (#2695) (#2822)
- **Discuss-phase no longer carries four internal text contradictions** — auto-mode removed a dead `max_discuss_passes` config read that contradicted its single-pass rule; the gate-prompts reference now matches the actual context-handling options and drops the 'Let Claude decide' cop-out that conflicted with the workflow's no-skip rule; the auto_advance fallback no longer routes back to the already-run confirm_creation step; and the assumptions workflow's answer_validation is re-synced to the canonical parent block. (#2886)
- **Corrected the legacy ADR range documentation** — the legacy zero-padded ADR range is now stated once (in docs/adr/README.md, as 0001–0012) and referenced rather than restated by docs/contributor-standards.md, so the two can no longer drift. The two zero-padded files that look legacy but are not (0174, 0656) are now identified as modern, mis-padded issue-numbered ADRs. Previously the two documents disagreed and neither matched disk. (#2836)
- **Agents and workflows no longer instruct a bare `gsd-tools` that fails on a shim-only install** — command-position `gsd-tools` invocations in the shipped agent/workflow source are now the portable `gsd_run` resolver (already defined in those files), so they resolve the runtime-local shim on installs with no `gsd-tools` binary on PATH. Previously only the Codex install-conversion pipeline rewrote these; the Claude-facing source shipped them verbatim and failed with `command not found`. (#2751) (#2851)
- **`query commit --files` now accepts absolute paths** — `cmdCommit` used `path.join(cwd, file)`, which concatenates instead of resetting on an absolute path, so absolute `--files` entries (e.g. the absolute `phase_dir` emitted by `init phase-op` since #2428) were joined to `cwd+absPath` (non-existent) and silently dropped as `nothing_to_commit` — and a mixed relative/absolute list committed the relative entries while reporting `committed:true`. Absolute paths are now normalized to repo-relative before staging/branch-detection, so they commit correctly and the phase-branch detection no longer matches digit-hyphen runs in the absolute prefix. (#2523) (#2638)
- **`/gsd:review` no longer silently drops a reviewer you asked for** — naming a reviewer with an explicit flag (`--gemini --qwen`) on a host where that lane could not run reported an info note and reviewed with a thinner set, while the run reported success; a cross-AI review that quietly loses a lane is blind in one eye. An explicitly-named lane that cannot run — CLI absent, `jq` missing, or local server unreachable — is now an error. `--all` and `review.default_reviewers` are unchanged and still skip undetected lanes with an info note. The Qwen lane also now captures stderr to a sidecar and includes it in its failure stub, matching every other lane, so a missing binary and an auth prompt are no longer indistinguishable from an empty review. (#2794) (#2820)
- **EoS Registry entries carrying the documented `effortSurface` axis are no longer rejected** — the registry validator required an exact eight-key axes object, so an entry that faithfully mirrored its upstream descriptor's optional ninth `effortSurface` key (`argv` or `none`, added by ADR-1239 amendment #2481) failed validation outright. (#2810) (#2813)
- **The claude-orchestration Workflow backend now honors your model settings** — with that BETA capability enabled, every plan was dispatched with no model at all, so `model_overrides`, `model_policy` and `model_profile` were silently ignored and each agent ran on whatever the session happened to be using. Plans now run on the same model the normal dispatch path would have used, and the generated script states which model was applied. Two consequences to expect: agents that were inheriting the session model will now run on the model your profile selects, and the first run after upgrading re-executes any in-flight resumable run, because the dispatch options changed. (#2686) (#2715)
- **A truncated or half-written frontmatter file is no longer silently read as "no metadata"** — a document whose `---` fence was opened and never closed used to return exactly the same empty result as a file that legitimately has no frontmatter, so a crash mid-write left every phase/state reader proceeding with empty contracts and no signal. GSD now names the offending file on stderr while returning the same value as before, so nothing that consumed the old result changes. A Markdown horizontal rule at the top of a document — including one above a labelled line such as `Note:` or `Author:` — is not mistaken for a truncated fence. (#1882) (#2712)
- **Refusing to run a phase from an executor worktree now tells you how to recover your work** — when GSD stopped because the session had drifted into an executor worktree, it only said to re-run from the orchestrator's worktree. If that worktree held commits or uncommitted changes, following that advice silently abandoned them. The refusal now lists the commits and files that exist only there, and gives the exact steps to integrate them before continuing. (#1856) (#2727)
- **An unreadable ROADMAP.md is no longer reported as a brand-new project** — a permission or I/O error reading `.planning/ROADMAP.md` used to return the same "phase not found" and `v1.0 / milestone` values as a project that simply has no roadmap yet, so workflows synthesized a blank phase or skipped requirement extraction with no signal. GSD now names the unreadable file on stderr while returning exactly what it returned before. A project that genuinely has no ROADMAP.md stays silent. (#1881) (#2729)
- **A corrupt `.planning/config.json` no longer silently discards your entire configuration** — a single trailing comma used to fall back to built-in defaults with no signal, indistinguishable from having no config file at all, so a project could run for weeks on defaults while its model profile, workflow toggles and branching strategy sat unread on disk. GSD now tells you the file could not be used and that its settings were not applied, and reports the cause (`config_unparseable` / `config_unreadable`) distinctly from genuine absence. The same applies to an unreadable file and to the global `~/.gsd/defaults.json`. (#1880) (#2688)
- **`--validate` is no longer documented for `/gsd-plan-phase` and `/gsd-execute-phase`** — both commands silently ignored the flag (only `/gsd-quick` implements it), so the docs promised a state-validation step that never ran. The false flag-table rows, CLI examples, and the `manager.flags.execute: "--validate"` config example are removed across the English docs and the ja-JP/zh-CN/ko-KR/pt-BR mirrors; the config example now shows `--cross-ai` (a flag execute-phase actually parses). `/gsd-quick`'s `--validate` docs are unchanged. (#2197) (#2574)
- **`/gsd-plan-phase` no longer 404s on non-Claude runtimes with `model_profile:"inherit"` + `resolve_model_ids:"omit"`** — the workflow passed `model="{planner_model}"` (and researcher_model/checker_model) verbatim into Agent() calls, so when the resolved model was empty it sent `model=""` and the runtime fell back to an unavailable Claude model → 404. plan-phase now mirrors execute-phase: when a `*_model` is "inherit" or empty, the `model=` param is omitted so the subagent inherits the orchestrator model. (#2517) (#2634)
- **Stale `todos/done` references in workflows and docs now read `todos/completed`** — the todos/done → todos/completed rename (commit 447d17a9) under-swept 14 descriptive lines across check-todos.md, the /gsd-help tree, ARCHITECTURE.md, and USER-GUIDE.md (en + 4 locales). Those stale references steered agents and users to archive closed todos into `done/` — a directory nothing in gsd-core reads — so closed todos became invisible to ID sequencing and to anything that inventories closed work. All 14 sites now read `completed/`, matching the canonical code path (cmdTodoComplete). A CI guard now blocks future under-sweeps. (#2491) (#2626)
- **Verification-status next-step commands now use the command surface each runtime actually installs** — on a Codex project, a phase blocked on verification suggested `/gsd:execute-phase`, which Codex does not install; the correct form is `$gsd-execute-phase`. The routing table stored hard-coded, deprecated colon-form strings with no runtime context, so `phase complete` and `query verification.status` relayed them verbatim to every runtime. All four routed states (missing, unknown, gaps_found, stale) now project through the shared runtime formatter. (#2617) (#2700)
- **A failed LM Studio or llama.cpp reviewer leg is now visible instead of silently dropped** — when a local OpenAI-compatible endpoint was unreachable or returned empty content, `/gsd-review` wrote no review file at all, so the reviewer's section was omitted from the final review and the result was indistinguishable from that reviewer never having been selected. Both legs now emit a diagnosable stub carrying curl's stderr and the raw response body, matching the guard the claude/gemini/codex legs already had. (#2605) (#2689)
- **`/gsd-execute-phase` now auto-closes pending todos for single-digit phases** — the close_phase_todos step normalizes both the phase number and each todo's `resolves_phase` value before comparing, so a todo tagged `resolves_phase: 5` is recognized when phase `05` completes. Previously the step compared the zero-padded `PHASE_NUMBER` (e.g. "05") against the unpadded value new-milestone wrote (e.g. "5") as literal strings, so every single-digit phase (1-9) silently failed to auto-close its todos — they stayed stuck in `pending/` forever despite their resolving phase completing. Decimal sub-phases (4.1 vs 04.1), letter suffixes, and quoted YAML values are now handled too. (#2576) (#2597)
- **The host-integration capability matrix now documents the `effortSurface` axis for every runtime** — the axis shipped in #2481 with real values in 19 runtime descriptors, but the matrix that ADR-1239 designates its cited source of truth had no legend entry and not one per-runtime row, so every committed value was undocumented in the one place meant to explain it. (#2615) (#2698)
- **`STATE.md` frontmatter is no longer silently overwritten by stale field lines in archive sections** — `buildStateFrontmatter` extracted Last Activity, Paused At, and the other current-state fields from the entire `STATE.md` body via `stateExtractField`, which matches the first `Field:` line anywhere. A historical line in an archive section further down the file silently overwrote the correct frontmatter value on every sync, and because the poisoning line stayed in the body it regressed again on the next write — so each repair looked successful and then silently reverted, with the offending line hundreds of lines away from the frontmatter. Field extraction is now scoped: current-state fields read from the body preamble before the first `##` heading, and session fields read from `## Session`. This generalizes the #2444 fix, which scoped `Stopped At` to `## Session` but did not propagate to the sibling fields. (#2660) (#2660)
- **A commit whose `git add` fails now says so, instead of partially committing or reporting "nothing to commit"** — when staging failed (an unwritable index in a linked worktree, permissions, or a timeout), GSD discarded the error: a multi-file request silently committed only the paths that happened to stage, and a total failure surfaced as `nothing_to_commit` or a downstream pathspec error naming an innocent file. Staging failures are now collected and reported as `staging_failed` (or `staging_timeout`) with the offending file and git's original stderr, before any commit is attempted, and the index is rolled back to its prior state. Applies to scoped (`--files`) commits, default `.planning/` commits, and sub-repo commits alike. (#2608) (#2693)
- **Cursor, Windsurf, and Codex hooks no longer fail with `require is not defined` under an ESM config root** — GSD now writes the `{"type":"commonjs"}` marker into the hooks directory alongside the staged `.js` scripts for these three runtimes (it already did for every other runtime), so Node loads them as CommonJS regardless of the runtime config's `"type"`. (#2717) (#2846)
- **The portability linter now catches Windows-path failures in membership and substring assertions** — `no-path-literal-in-assert` flags `.includes`/`.indexOf`/`.startsWith`/`.endsWith`/`.match` over a path-returning receiver (including through a `.map()` hop), not just equality assertions. Previously these passed lint and failed on Windows CI; the rule now surfaces them at lint time. (#2764) (#2879)
- **`/gsd-review`'s codex lane no longer passes the hook-trust bypass flag or runs its capability probe** — host-harness safety classifiers denied invocations carrying them, and flagless invocations work in steady state. A genuine untrusted-hook failure still surfaces as a dropped lane with diagnosable stderr. (#2479) (#2536)
- **`/gsd-plan-phase --reviews` now actually replans in chunked mode instead of silently skipping every plan** — the per-plan resume-check skips existing plans for crash-resume, but now exempts `--reviews` (whose purpose is to replan with review feedback). Also fixed the outline resume-check, which looked for a marker the agent only returned (never wrote to the file), so the outline always re-ran. (#2762) (#2887)
- **pi no longer silently hijacks non-Anthropic providers' model choices** — `pi/gsd.cjs`'s `before_provider_request` handler unconditionally rewrote `payload.model` to the built-in pi/sonnet tier default (`claude-sonnet-5`) via the model-catalog fallback, breaking every outgoing request for pi users on non-Anthropic providers (kimi-coding, zai, openrouter, openai-codex, minimax). The handler now inspects `model_profile_overrides.pi[tier]` explicitly *before* calling `resolveTierEntry` (whose catalog fallback previously masked the "user did not opt in" signal) and fail-opens (`return undefined`) when the user has not set an override — including explicit `null` and `''` (clearing a previously-set value). An explicit opt-in via `model_profile_overrides.pi[tier]` still steers, preserving the legitimate use case. (#2460) (#2499)
- **`GSD_AUDIT=1` now actually produces an audit trail** — the reference dispatch logger is wired onto the live command seam, so opting in yields the documented structured stderr line and the `.planning/.gsd-trace.jsonl` audit trail. Previously the seam built its dispatch hub without a logger, so it fell back to a no-op and the opt-in signal was inert with no indication why. With observability off, dispatch output is byte-for-byte unchanged. (#2620) (#2621)
- **Codebase scan and ship-time capability hooks now honor your model settings** — /gsd:scan dispatched its mapper agent with a model placeholder nothing resolved, and ship-time capability hooks did the same, so `model_overrides` and `model_policy` were silently ignored at both and the agent ran on whatever the session happened to be using. Both now resolve a real model, and omit the model parameter entirely when it resolves to "inherit" or empty rather than passing an empty value that fails on non-Claude runtimes. Note: the scan mapper now runs on the model your profile selects rather than inheriting the session's. (#2684) (#2710)
- **State sync now reports the correct total phase count on a flat unmilestoned roadmap** — `progress.total_phases` no longer falls back to the on-disk phase-directory count when the roadmap has no versioned milestone heading; it uses the authoritative roadmap count, matching the write-path and resolving the contradiction between smart-entry's `total_phases` and `roadmap_total_phases`. (#2828) (#2892)
- **Worktree cleanup-wave now rescues uncommitted SUMMARY.md** — the rescue step's `git cat-file -e HEAD:<path>` check assumed an absent path returns exit 1, but git returns 128, so rescue never fired: the executor's uncommitted `<id>-SUMMARY.md` blocked cleanup as `worktree_dirty` and risked silent loss on `worktree remove --force`. Rescue now fires on any non-zero exit (only exit 0 = committed → skip), so uncommitted SUMMARYs are copied into the main tree before the dirty check. (#2556) (#2611)
- **Code-review now scopes repository-root and extensionless build files (Dockerfile, Makefile, .gitlab-ci.yml, renovate.json, AGENTS.md)** — the SUMMARY.md file extractor no longer silently drops every root-level path and every extensionless build file, and a partial SUMMARY scope is now cross-checked against `git diff` with a warning naming any changed files it missed. (#2666) (#2895)
- **`execute-phase.md` now has ~3.3 KB of byte-budget headroom** — the `offer_next` step body (terminal reporting + next-phase routing prose) was extracted to `gsd-core/references/offer-next.md` and eagerly `@`-referenced, restoring the headroom the frozen size ceiling exists to provide. Previously the ceiling had only ~32-137 bytes of margin, so any bugfix touching `execute-phase.md` had to extract unrelated content or raise the ceiling. Runtime behavior is unchanged (the `@`-reference loads eagerly). (#2537) (#2642)
### Security
- **Malformed and shadowing Kimi payloads no longer disarm the guards that block** — `normalizeKimiPayload` (inlined in all five PreToolUse/PostToolUse guard hooks) rebuilt `old_string`/`new_string` with `String(e.old ?? '')`. Two inputs crashed it, and because normalization runs before any tool dispatch, both crashes landed in each guard's outer `catch { process.exit(0) }` — which emits the same exit code as "nothing to report", turning a should-**block** call into a silent **allow**. First, `??` guards the value and not the dereference, so a nullish entry (`edit: [null]`) threw on the property read. Second, coercion itself can throw: `{"toString": null}` is valid JSON that raises `Cannot convert object to primitive value`, so even a well-formed edit object could crash normalization. Two hard blocks were bypassable through either route: `gsd-worktree-path-guard`'s cross-git-root write block (the same write is correctly blocked with a well-formed edit list), and `gsd-workflow-guard`'s force-add block on `agent-*` branches (via a `Shell` payload carrying a spurious `edit` field the Bash path never even reads). Fixed with `e?.old` / `e?.new` plus a guarded coercion, landed identically across all five copies; the coercion is wrapped rather than type-tested so that stringification is unchanged for every value that can coerce. **Three model-supplied fields are now authoritative rather than merely defaulted.** Normalization used to fill `file_path`, `old_string` and `new_string` only when the key was `=== undefined`, so any value the model chose to include won — while kimi-cli executes on `path` and `edit`. Its `StrReplaceFile` schema is `path` + `edit` only (`src/kimi_cli/tools/file/replace.py` @ `4a550ef`) and carries none of those three keys, so each one appearing in a Kimi payload is always model-supplied. A cross-root `path` paired with a spurious `file_path: ""` left `gsd-worktree-path-guard` reading an empty string and exiting 0 while the identical write without the extra key blocked; likewise a `new_string: ""` — or any benign non-empty decoy, which a type test would not have caught — left `gsd-prompt-guard`'s injection scan reading empty content and returning at its `if (!content)` guard before it ever saw the real `edit[].new`. All three are now reconstructed unconditionally, which can only ever narrow what a guard inspects to what will actually be written. Reachability is not speculative: kimi-cli's `soul/toolset.py` json-parses the model's raw tool arguments and passes the dict verbatim as `tool_input` to `PreToolUse`, doing typed validation only later inside `tool.call()` — so the model controls extra keys at the moment the hook decides. **Separately, the guards now read payload path fields typed.** A non-string `file_path` (`[]`, `{}`) is truthy, so it survived each guard's `if (!filePath)` early-out and then threw inside `path.isAbsolute()` / `.includes()` / `.replace()`, reaching the same fail-open catch — crash-to-allow through the guard's own read rather than through normalization, and live on **native Claude Code payloads** too, since normalization returns early for non-Kimi tool names and so never masked the bad value there. Previously this was closed only as a side effect of a valid string `path` overwriting `file_path`; it is now closed unconditionally at all six read sites (the five normalized guards plus `gsd-windsurf-pre-write`, which already read typed), and a source-level invariant (`tests/kimi-guard-typed-payload-reads.test.cjs`) fails if any hook regresses to an untyped read. The native Claude Code contract (`file_path` governs) is unchanged. **Scope on Kimi:** normalization makes each guard's *checks* run; it does not make every guard *enforceable*. What can actually block on Kimi is what runs at PreToolUse — the worktree cross-root write block and the workflow force-add block. `gsd-read-injection-scanner` is a PostToolUse hook, and kimi-cli's dispatch never inspects PostToolUse hook results (`soul/toolset.py` fires them as a detached task and returns the tool result without awaiting it), so no output shape the scanner emits can block or flag a Kimi tool call; its prompt-injection block is not enforceable on Kimi under Kimi's current hook architecture. Regression coverage is negative-controlled against the pre-fix guards, and a property test (`tests/kimi-normalize-payload.property.test.cjs`) backs the totality claim generatively. `next`-only — released versions carry no Kimi normalization at all. (#2547) (#2595)
- **Dev-tooling `js-yaml` bumped past the merge-key DoS advisory** — `js-yaml` was pinned `^4.2.0`, inside the vulnerable `4.0.0 - 4.2.0` range of GHSA-52cp-r559-cp3m (quadratic CPU on YAML merge-key chains). It is a devDependency with no shipped-runtime reachability, but `scripts/workflow-policy.cjs` parses workflow frontmatter in CI, which is attacker-controlled on a fork PR. Now `^4.2.1`. (#2654) (#2655)
## [1.8.0] - 2026-07-22
### Added